Hard drives, SSDs and NVMe drives keep internal health counters called SMART (Self-Monitoring, Analysis and Reporting Technology). Read regularly, those counters often warn you days or weeks before a drive fails. In this tutorial you will install smartmontools on Ubuntu 24.04, read the health of every drive with smartctl, learn which attributes actually predict failure, run self-tests, and configure the smartd daemon to test drives on a schedule and email you when something goes wrong.
Prerequisites
To follow this guide you need:
- A bare metal server or physical machine running Ubuntu 24.04 LTS. Virtual machines, including VPS, see virtual disks that do not report real SMART data.
- A non-root user with
sudoprivileges. - For email alerts (Step 6): a working way to send mail from the server, such as Postfix configured as a relay through your SMTP provider.
Step 1 - Installing smartmontools
smartmontools is in the Ubuntu repositories. By default apt also pulls in a mail server as a recommended package; install without recommendations so you decide separately how the server sends mail:
sudo apt update
sudo apt install --no-install-recommends smartmontools
Check the version:
smartctl --version | head -1
smartctl 7.4 2023-08-01 r5530 [x86_64-linux-6.8.0-45-generic] (local build)
List the drives smartmontools can see and the device type it will use for each:
sudo smartctl --scan
/dev/sda -d scsi # /dev/sda, SCSI device
/dev/sdb -d scsi # /dev/sdb, SCSI device
/dev/nvme0 -d nvme # /dev/nvme0, NVMe device
SATA drives show up as scsi here; smartctl detects that they are ATA drives when you query them.
Step 2 - Checking a drive's overall health
Start with the drive's identity to confirm what you are looking at and that SMART is enabled:
sudo smartctl -i /dev/sda
Device Model: Samsung SSD 870 EVO 1TB
Serial Number: S6PUNX0T123456
Firmware Version: SVT02B6Q
User Capacity: 1,000,204,886,016 bytes [1.00 TB]
Rotation Rate: Solid State Device
SMART support is: Available - device has SMART capability.
SMART support is: Enabled
If the last line says Disabled, enable SMART on the drive:
sudo smartctl -s on /dev/sda
Now ask for the drive's own verdict:
sudo smartctl -H /dev/sda
SMART overall-health self-assessment test result: PASSED
FAILED means the drive itself predicts failure: back up the data and replace it. PASSED is not a guarantee, though. The overall check only fails when an attribute crosses a threshold set by the manufacturer, which often happens late. The raw attributes in the next step give an earlier warning.
To see everything smartctl knows about a drive in one output (identity, health, attributes, error log and self-test log), use sudo smartctl -x /dev/sda.
Step 3 - Reading the SMART attributes that matter
For SATA and SAS drives, list the vendor attributes:
sudo smartctl -A /dev/sda
ID# ATTRIBUTE_NAME FLAG VALUE WORST THRESH TYPE UPDATED WHEN_FAILED RAW_VALUE
5 Reallocated_Sector_Ct 0x0033 100 100 010 Pre-fail Always - 0
9 Power_On_Hours 0x0032 095 095 000 Old_age Always - 21034
12 Power_Cycle_Count 0x0032 099 099 000 Old_age Always - 48
177 Wear_Leveling_Count 0x0013 097 097 000 Pre-fail Always - 31
187 Reported_Uncorrect 0x0032 100 100 000 Old_age Always - 0
194 Temperature_Celsius 0x0022 066 052 000 Old_age Always - 34
197 Current_Pending_Sector 0x0012 100 100 000 Old_age Always - 0
198 Offline_Uncorrectable 0x0010 100 100 000 Old_age Always - 0
199 UDMA_CRC_Error_Count 0x003e 100 100 000 Old_age Always - 0
VALUE, WORST and THRESH are normalized scores where lower is worse; the drive reports FAILED when VALUE drops to THRESH. RAW_VALUE is the real counter and is what you should watch. The attribute list differs between vendors, but these are the ones most closely linked to failure:
| ID | Attribute | What a non-zero or rising raw value means |
|---|---|---|
| 5 | Reallocated_Sector_Ct | Bad sectors were replaced with spares. Any increase is a warning. |
| 187 | Reported_Uncorrect | Reads the drive could not recover. Data was lost or had to come from RAID. |
| 197 | Current_Pending_Sector | Sectors that could not be read and wait to be reallocated. Serious on HDDs. |
| 198 | Offline_Uncorrectable | Sectors found unreadable during offline scans. |
| 199 | UDMA_CRC_Error_Count | Transfer errors between drive and controller. Usually a bad cable or backplane, not the drive. |
| 177 / 231 / 233 | Wear indicators (name varies by SSD vendor) | Remaining SSD endurance. Watch the normalized VALUE fall towards the threshold. |
A single reallocated sector on an old drive can stay stable for years. What matters is the trend: if attributes 5, 187, 197 or 198 grow from one week to the next, plan the replacement now.
Step 4 - Checking NVMe drives
NVMe drives do not use the attribute table. They report a standard health log instead:
sudo smartctl -a /dev/nvme0
The relevant part of the output looks like this:
SMART overall-health self-assessment test result: PASSED
SMART/Health Information (NVMe Log 0x02)
Critical Warning: 0x00
Temperature: 41 Celsius
Available Spare: 100%
Available Spare Threshold: 10%
Percentage Used: 3%
Data Units Written: 48,211,934 [24.6 TB]
Power On Hours: 11,872
Unsafe Shutdowns: 21
Media and Data Integrity Errors: 0
Error Information Log Entries: 0
Watch these fields:
- Critical Warning: must be
0x00. Any other value means the controller flagged a problem (spare space low, temperature, read-only mode or reliability degraded). - Available Spare: must stay above Available Spare Threshold.
- Percentage Used: the manufacturer's estimate of consumed endurance. At 100% the drive has reached its rated life; it may keep working, but plan a replacement.
- Media and Data Integrity Errors: must be
0. Unrecovered data errors.
Error Information Log Entries often counts harmless errors caused by the OS probing unsupported commands, so do not panic if it is non-zero while everything else is clean.
Step 5 - Running self-tests
Self-tests make the drive check itself while it keeps serving I/O. A short test takes a couple of minutes; an extended (long) test reads the whole surface and can take hours on large HDDs:
sudo smartctl -t short /dev/sda
Testing has begun.
Please wait 2 minutes for test to complete.
Test will complete after Wed Sep 24 10:14:52 2026 UTC
Use smartctl -X to abort test.
Start an extended test the same way with -t long. NVMe drives that support self-tests accept the same commands with smartmontools 7.4 (sudo smartctl -t short /dev/nvme0).
When the time has passed, read the self-test log:
sudo smartctl -l selftest /dev/sda
SMART Self-test log structure revision number 1
Num Test_Description Status Remaining LifeTime(hours) LBA_of_first_error
# 1 Short offline Completed without error 00% 21034 -
# 2 Extended offline Completed without error 00% 20867 -
A status such as Completed: read failure with an LBA_of_first_error means the drive found a sector it cannot read. Back up the data and replace the drive.
Step 6 - Automating checks and alerts with smartd
Running smartctl by hand does not scale. The smartd daemon, included in the package, polls every drive every 30 minutes, runs self-tests on a schedule and sends an alert when a check fails.
On Ubuntu, the service is called smartmontools and reads /etc/smartd.conf. Open it:
sudo nano /etc/smartd.conf
The file contains a single active DEVICESCAN line by default. Comment it out by adding # at the start, then add this line at the end of the file. A directive must stay on one line (or use \ continuations with no comments in between):
DEVICESCAN -a -n standby,q -W 4,45,55 -s (S/../.././02|L/../../6/03) -m [email protected] -M exec /usr/share/smartmontools/smartd-runner
Replace [email protected] with your address. Here is what each option does:
DEVICESCAN: monitor every drive smartd finds.-a: check overall health, failing attributes, error and self-test logs, and pending/offline uncorrectable sectors.-n standby,q: do not wake HDDs that are spun down just to check them.-W 4,45,55: log temperature changes of 4 degrees or more, log when the drive reaches 45 C and alert at 55 C.-s (S/../.././02|L/../../6/03): run a short test every day at 02:00 and a long test every Saturday at 03:00.-mand-M exec: send alerts to that address through Ubuntu'ssmartd-runner, which runs the scripts in/etc/smartmontools/run.d/(including one that sends mail).
If a controller hides drives from DEVICESCAN (for example a MegaRAID card), list them one per line instead, such as /dev/sda -d megaraid,0 -a -m [email protected].
Before relying on the configuration, add -M test to the line temporarily. With it, smartd sends one test alert per drive when it starts:
DEVICESCAN -a -n standby,q -W 4,45,55 -s (S/../.././02|L/../../6/03) -m [email protected] -M test -M exec /usr/share/smartmontools/smartd-runner
Restart the service and check its log:
sudo systemctl restart smartmontools
sudo journalctl -u smartmontools -n 30
smartd[1204]: Device: /dev/sda [SAT], opened
smartd[1204]: Device: /dev/sda [SAT], Samsung SSD 870 EVO 1TB, S/N:S6PUNX0T123456, FW:SVT02B6Q, 1.00 TB
smartd[1204]: Device: /dev/sda [SAT], is SMART capable. Adding to "monitor" list.
smartd[1204]: Device: /dev/nvme0, is SMART capable. Adding to "monitor" list.
smartd[1204]: Monitoring 1 ATA/SATA, 0 SCSI/SAS and 1 NVMe devices
smartd[1204]: Sending warning via /usr/share/smartmontools/smartd-runner to [email protected] ...
smartd[1204]: Warning via /usr/share/smartmontools/smartd-runner to [email protected]: successful
If the log shows the warning was sent successfully and it arrives in your mailbox, remove -M test from the line and restart the service again. If the log reports a mail error, fix your mail setup first (see Troubleshooting).
Make sure the service starts at boot:
sudo systemctl is-enabled smartmontools
enabled
When a scheduled test is due, smartd logs Starting scheduled Short Self-Test in the same journal, so you can confirm the schedule works the next day.
Troubleshooting
SMART support is: Unavailable or Unknown USB bridge. The drive sits behind a USB adapter or RAID controller that hides it. Tell smartctl how to reach it: -d sat for most USB enclosures, -d megaraid,N for disks behind LSI/Broadcom MegaRAID controllers (N is the disk's device ID, visible in the controller tool), and -d cciss,N for older HPE Smart Array controllers.
smartctl shows no attributes on a VPS. Virtual disks do not have SMART data. Monitor disk health at the hypervisor or, on a bare metal server, on the physical drives.
No email arrives. Test mail on its own with echo test | mail -s test [email protected] (install mailutils if the mail command is missing). If that fails, the problem is the mail setup, not smartd. Then look for mail errors with sudo journalctl -u smartmontools | grep -i mail.
UDMA_CRC_Error_Count keeps increasing. Reseat or replace the SATA cable or check the backplane slot. The drive itself is usually fine.
Want a one-off check of every drive without waiting for the daemon. Run sudo smartd -q onecheck: smartd registers all drives, checks them once and exits, printing any problem to the terminal.
Conclusion
You installed smartmontools, learned to read overall health, the SMART attributes and the NVMe health log, ran self-tests, and set up smartd to test drives on a schedule and email you when a drive starts to fail. As next steps, export SMART metrics to Prometheus with smartctl_exporter to graph trends across servers, pair disk monitoring with RAID or ZFS so a single failure never loses data, and keep verified backups of anything stored on the drives.
