lm-sensors reads the hardware monitoring chips on a motherboard and CPU through the kernel's hwmon interface, giving you temperatures, fan speeds and voltages from the command line. It is the quickest way to find out whether a server is running hot or a fan has stopped. In this tutorial you will install lm-sensors on Ubuntu 24.04, detect the sensor chips, read and label their values, and export them to Prometheus with node_exporter so you can alert on high temperatures.

Prerequisites

To follow this guide you need:

  • A bare metal server or physical machine running Ubuntu 24.04 LTS. Virtual machines, including VPS, do not expose the host's sensors, so sensors shows little or nothing there.
  • A non-root user with sudo privileges.
  • Optional, for Step 5: a Prometheus server that can reach port 9100 on this machine.

Step 1 - Installing lm-sensors

The package is in the Ubuntu repositories:

sudo apt update
sudo apt install lm-sensors

Check that the sensors command is available:

sensors -v
sensors version 3.6.0 with libsensors version 3.6.0

On many modern systems the kernel already loads the right drivers at boot, so try sensors right away. If it prints CPU temperatures you can skip to Step 3.

Step 2 - Detecting sensor chips

When sensors only shows a few values or prints No sensors found!, run sensors-detect. It probes the system for known monitoring chips and tells you which kernel drivers they need. The --auto option accepts the default answer to every question, which only runs the safe probes:

sudo sensors-detect --auto

The end of the output lists the drivers it found:

Driver `coretemp':
  * Chip `Intel digital thermal sensor' (confidence: 9)

Driver `nct6775':
  * ISA bus, address 0x290
    Chip `Nuvoton NCT6798D Super IO Sensors' (confidence: 9)

To load everything that is needed, add this to /etc/modules:
#----cut here----
# Chip drivers
coretemp
nct6775
#----cut here----

Create a file in /etc/modules-load.d/ with the drivers from your own output so they load at every boot, then load them now with modprobe -a:

printf 'coretemp\nnct6775\n' | sudo tee /etc/modules-load.d/lm-sensors.conf
sudo modprobe -a coretemp nct6775

Confirm the drivers are loaded. The usual CPU drivers are coretemp for Intel and k10temp for AMD:

lsmod | grep -E 'coretemp|k10temp|nct6775|it87'
nct6775                94208  0
hwmon_vid              12288  1 nct6775
coretemp               24576  0

Step 3 - Reading temperatures, fans and voltages

Run sensors to print every reading:

sensors
coretemp-isa-0000
Adapter: ISA adapter
Package id 0:  +46.0°C  (high = +80.0°C, crit = +100.0°C)
Core 0:        +43.0°C  (high = +80.0°C, crit = +100.0°C)
Core 1:        +45.0°C  (high = +80.0°C, crit = +100.0°C)

nct6798-isa-0290
Adapter: ISA adapter
in0:                   1.06 V  (min =  +0.00 V, max =  +1.74 V)
fan1:                1150 RPM  (min =    0 RPM)
fan2:                 820 RPM  (min =    0 RPM)
SYSTIN:               +34.0°C  (high = +80.0°C, hyst = +75.0°C)
CPUTIN:               +41.5°C  (high = +80.0°C, hyst = +75.0°C)

nvme-pci-0100
Adapter: PCI adapter
Composite:    +38.9°C  (low  = -273.1°C, high = +84.8°C)

Each block is a chip. The first line is its name (for example coretemp-isa-0000), and each row shows the current value followed by the limits the chip reports. The high and crit values of the CPU come from the processor itself and are the ones to compare against.

On AMD systems the CPU block is called k10temp-pci-00c3 and the main value is Tctl (and Tccd1, Tccd2 per core complex).

A few options are useful:

sensors coretemp-isa-0000

Shows a single chip.

sensors -j

Prints JSON, which is easier to parse in scripts than the text output.

watch -n 2 sensors

Refreshes the readings every two seconds, handy while you run a load test. Press Ctrl+C to stop.

Step 4 - Labeling sensors and ignoring bogus readings

Motherboard chips often show generic names (in0, fan2, SYSTIN) and inputs that are not connected, which report nonsense values. You can rename and hide them in a configuration file. Do not edit /etc/sensors3.conf, which the package owns; add your own file in /etc/sensors.d/ instead:

sudo nano /etc/sensors.d/local.conf

This example targets the Nuvoton chip from Step 3. Use the chip name and input names from your own sensors output:

chip "nct6798-isa-*"
    label in0 "Vcore"
    label fan1 "CPU fan"
    label fan2 "Rear fan"
    label temp1 "Motherboard"
    ignore fan3
    ignore fan4
    set fan1_min 400
  • label changes the displayed name. The input names (temp1, fan1, in0) are the kernel names; SYSTIN in the output is the driver's own label for temp1.
  • ignore hides an input that is not connected.
  • set writes a limit into the chip. Here the chip will flag the CPU fan if it drops below 400 RPM.

Labels and ignore take effect the next time you run sensors. set statements are only written to the hardware when you run:

sudo sensors -s

Check the result:

sensors nct6798-isa-0290
nct6798-isa-0290
Adapter: ISA adapter
Vcore:                 1.06 V  (min =  +0.00 V, max =  +1.74 V)
CPU fan:             1150 RPM  (min =  400 RPM)
Rear fan:             820 RPM  (min =    0 RPM)
Motherboard:          +34.0°C  (high = +80.0°C, hyst = +75.0°C)

If sensors reports a parse error, it names the file and line; fix it and run the command again. The CPU temperature limits from coretemp and k10temp are read-only, so set does not work on them.

Step 5 - Exporting sensors to Prometheus

Reading values by hand is fine for troubleshooting, but you want alerts when a server overheats at night. Prometheus node_exporter reads the same hwmon data as lm-sensors through its hwmon collector, which is enabled by default. Install the Ubuntu package, which runs it as a systemd service on port 9100:

sudo apt install prometheus-node-exporter

Check that the service is running and exposes temperature metrics:

systemctl status prometheus-node-exporter --no-pager
curl -s http://localhost:9100/metrics | grep '^node_hwmon_temp_celsius'
node_hwmon_temp_celsius{chip="platform_coretemp_0",sensor="temp1"} 46
node_hwmon_temp_celsius{chip="platform_coretemp_0",sensor="temp2"} 43
node_hwmon_temp_celsius{chip="platform_coretemp_0",sensor="temp3"} 45

The human-readable names are in a separate metric, node_hwmon_sensor_label, and fan speeds are in node_hwmon_fan_rpm.

node_exporter has no authentication, so do not publish port 9100 to the internet. If UFW is enabled, allow it only from your Prometheus server, replacing your_prometheus_ip:

sudo ufw allow from your_prometheus_ip to any port 9100 proto tcp

Add the server as a target in your Prometheus configuration, then create an alerting rule. This rule fires when any temperature stays within 5 degrees of the critical limit reported by the chip for five minutes:

groups:
  - name: hardware
    rules:
      - alert: HardwareTemperatureNearCritical
        expr: node_hwmon_temp_celsius > on(instance, chip, sensor) (node_hwmon_temp_crit_celsius - 5)
        for: 5m
        labels:
          severity: critical
        annotations:
          summary: "{{ $labels.chip }} {{ $labels.sensor }} on {{ $labels.instance }} is near its critical temperature"

Using the chip's own critical limit avoids hardcoding a number that is wrong for half of your CPUs. For sensors that do not report a critical limit, add a second rule with a fixed threshold, for example node_hwmon_temp_celsius{chip=~".*coretemp.*"} > 85.

Troubleshooting

No sensors found! No hardware monitoring driver is loaded. On a VPS this is expected. On bare metal, run sudo sensors-detect --auto again and load the drivers it suggests with sudo modprobe coretemp (Intel) or sudo modprobe k10temp (AMD).

The Super I/O chip driver fails to load with resource conflicts with ACPI region. Check with sudo dmesg | grep -i acpi. The firmware reserves the chip for its own use, and Linux refuses to access it to avoid conflicts. On many servers the BMC is the better source for board sensors in that case: read them with ipmitool sdr type Temperature.

Readings look wrong (for example 0 RPM fans or -128 °C). The input is not connected or the driver guesses the wiring. Compare with the BMC or firmware readings and ignore the inputs that do not match anything.

AMD Tctl looks 10 to 20 degrees higher than expected. On some AMD processors Tctl includes an offset used for fan control. Use Tccd values, when present, to judge the real die temperature.

Conclusion

You installed lm-sensors, detected the monitoring chips on the server, read and labeled temperatures, fans and voltages, and exported them to Prometheus with an alert that uses each chip's own critical limit. As next steps, build a Grafana panel from node_hwmon_temp_celsius to spot cooling problems over time, add disk health monitoring with smartmontools, and on servers with a BMC, cross-check readings and fan status with IPMI.