The load average is the first number most administrators check when a Linux server feels slow, and also one of the most misread. It does not measure CPU usage: it counts processes that are running, waiting for a CPU, or blocked waiting for disk or network I/O. In this tutorial you will learn to read the load average correctly, work out whether a high value comes from CPU, disk I/O or memory pressure, identify the process responsible and apply the right fix. The commands are for Ubuntu 24.04 and work the same on Debian 12 and Rocky Linux 9.

Prerequisites

To follow this guide you need:

  • A server running Ubuntu 24.04 LTS, for example a CubePath VPS.
  • A non-root user with sudo privileges.
  • The sysstat and iotop packages, which you will install in Step 3.

Understanding the load average

Linux reports three load averages: the average number of tasks in the runnable state (R, using or waiting for a CPU) or the uninterruptible sleep state (D, usually waiting for disk I/O) over the last 1, 5 and 15 minutes.

uptime
 11:02:17 up 34 days,  2:10,  2 users,  load average: 7.84, 5.12, 2.31

The numbers only make sense compared with the number of CPUs:

nproc
4

On this 4-vCPU server:

LoadMeaning
Below 4There is spare CPU capacity; tasks rarely wait.
Around 4All CPUs are busy, but tasks are not queuing much.
Well above 4Tasks are waiting, either for a CPU or for I/O.

Read the three values as a trend. In the example above, 7.84, 5.12, 2.31 means the load is rising: something started a few minutes ago and is getting worse. 2.3, 5.1, 7.8 would mean the server is recovering from an earlier spike.

Because D state tasks count too, a server can show a load of 20 with its CPUs almost idle if many processes are stuck waiting for a slow disk or an unresponsive NFS mount. That is why the next step is always to find out which kind of load you have.

Step 1 - Checking CPU, I/O wait and steal time

Run vmstat with a one-second interval for five samples. Ignore the first line, which is an average since boot:

vmstat 1 5
procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
 r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa st
 9  0      0 412304  61220 2104880    0    0     2    35  812 1420 91  7  2  0  0
 8  0      0 411980  61220 2104912    0    0     0    48  790 1388 93  6  1  0  0  0
 9  0      0 411700  61228 2104936    0    0     0    12  803 1402 92  7  1  0  0
 8  0      0 411512  61228 2104960    0    0     0    20  799 1395 92  7  1  0  0  0

The columns that tell you what kind of load you have:

ColumnMeaningWhat a high value points to
rTasks running or waiting for a CPUCPU-bound load if consistently above nproc
bTasks blocked in uninterruptible sleepI/O-bound load
us, syCPU time in user and kernel spaceBusy processes (user) or heavy system calls (kernel)
waCPU idle while waiting for I/OSlow or saturated storage
stTime stolen by the hypervisorThe host is busy; common on shared VPS plans
si, soMemory swapped in and out per secondMemory pressure

The example shows a CPU-bound server: r is around 8 on 4 CPUs, us is above 90 and wa is 0.

Linux also exposes pressure stall information (PSI), which tells you directly what percentage of time tasks were delayed waiting for each resource. It is enabled on the Ubuntu 24.04 kernel:

cat /proc/pressure/cpu /proc/pressure/io /proc/pressure/memory
some avg10=68.41 avg60=55.02 avg300=31.77 total=924871652
full avg10=0.00 avg60=0.00 avg300=0.00 total=0
some avg10=0.12 avg60=0.30 avg300=0.21 total=31044211
full avg10=0.05 avg60=0.11 avg300=0.09 total=20511234
some avg10=0.00 avg60=0.00 avg300=0.00 total=1204511
full avg10=0.00 avg60=0.00 avg300=0.00 total=920144

The three blocks are CPU, I/O and memory, in that order. some avg10=68.41 for CPU means that in the last 10 seconds, at least one task was waiting for a CPU 68% of the time. Whichever resource has the highest some values is your bottleneck.

Step 2 - Listing the tasks behind the load

Since the load counts R and D tasks, list exactly those:

ps -eo state,pid,user,%cpu,wchan:30,cmd --sort=state | awk '$1 ~ /^[RD]/'
D  1893 mysql     4.1 io_schedule                    /usr/sbin/mysqld
R 20411 www-data 98.7 -                              php-fpm: pool www
R 20417 www-data 97.9 -                              php-fpm: pool www
R 20422 www-data 97.2 -                              php-fpm: pool www

R processes are competing for CPU. For D processes, the wchan column shows the kernel function they are blocked in: names containing io_schedule, blk, ext4 or jbd2 point to local disk, while nfs or rpc point to a network file system.

Step 3 - Installing diagnostic tools

sysstat provides iostat, pidstat and sar; iotop shows disk I/O per process:

sudo apt update
sudo apt install sysstat iotop

On Rocky Linux 9, install them with sudo dnf install sysstat iotop.

Step 4 - Diagnosing CPU-bound load

If r is high and us or sy dominate, find which processes use the CPU. pidstat averages over an interval, which is more reliable than a single ps snapshot:

pidstat -u 5 1
Average:      UID       PID    %usr %system  %guest   %wait    %CPU   CPU  Command
Average:       33     20411   95.20    2.40    0.00   12.60   97.60     -  php-fpm8.3
Average:       33     20417   94.80    2.80    0.00   13.00   97.60     -  php-fpm8.3
Average:       33     20422   94.00    3.00    0.00   12.80   97.00     -  php-fpm8.3
Average:      109      1893    3.40    0.80    0.00    1.20    4.20     -  mysqld

The %wait column is the time a process was ready to run but waiting for a CPU, which is the direct cost of the overload.

For an interactive view, run top, press P to sort by CPU and 1 to show each CPU separately. A single process pinned at 100% on one CPU while the others are idle points to single-threaded code that cannot be fixed by adding CPUs.

Also look at st in the vmstat output. If steal time is consistently above 10% while your own processes are not especially busy, the physical host is overcommitted and the fix is on the provider side or a plan with dedicated CPUs.

Step 5 - Diagnosing I/O-bound load

If b and wa are high, check the storage devices with iostat in extended mode. -z hides idle devices:

iostat -xz 5 2
Device            r/s     w/s     rkB/s     wkB/s  r_await  w_await  aqu-sz  %util
vda            412.00  1210.40  13184.00  48416.00    18.42    41.97   52.31  99.60

The output above is trimmed to the most useful columns. Look at the second report, since the first is an average since boot. The important fields are:

  • %util near 100: the device is busy all the time.
  • r_await and w_await: average milliseconds per read and write, including queue time. On SSD or NVMe storage, values consistently above 10 ms mean the device is saturated.
  • aqu-sz: the average number of requests waiting in the queue.

Find which process generates the I/O. iotop -o shows only processes that are currently reading or writing:

sudo iotop -o

pidstat -d gives the same information without an interactive screen:

pidstat -d 5 1
Average:      UID       PID   kB_rd/s   kB_wr/s kB_ccwr/s iodelay  Command
Average:      109      1893  13020.40  46211.20      0.00     4210  mysqld
Average:        0      2240      0.00   1840.60      0.00       12  rsyslogd

Common causes are an unindexed database query scanning a large table, a backup or rsync job running during peak hours, or heavy swapping (next step).

Step 6 - Checking memory pressure

When RAM runs out, the kernel pushes memory pages to swap and reads them back as they are needed. That I/O raises the load and makes everything slow, even though the root cause is memory:

free -h
               total        used        free      shared  buff/cache   available
Mem:           3.8Gi       3.6Gi        98Mi        12Mi       210Mi       112Mi
Swap:          2.0Gi       1.7Gi       300Mi

A low available value together with non-zero si and so in vmstat means the server is swapping actively. Find the largest memory users:

ps -eo pid,user,rss,cmd --sort=-rss | head -n 6

Also check whether the kernel has already started killing processes to free memory:

sudo journalctl -k -b | grep -i "out of memory"

Step 7 - Fixing the cause

Once you know the kind of load and the process behind it, apply a fix that matches.

CPU-bound load

If the busy process is a batch job that can run slower, lower its CPU priority so interactive services win. Replace 20411 with the PID you found:

sudo renice -n 10 -p 20411

For a systemd service that should never take the whole server, set a CPU limit. CPUQuota=200% allows it to use at most two CPUs' worth of time:

sudo systemctl set-property your_service CPUQuota=200%

The change applies immediately and persists across reboots. Check it with:

systemctl show -p CPUQuotaPerSecUSec your_service
CPUQuotaPerSecUSec=2s

If the load is legitimate traffic, tune the application (for example, the number of PHP-FPM workers or caching of expensive pages) or move to a plan with more vCPUs.

I/O-bound load

For a backup or maintenance job, lower its I/O priority so it only uses the disk when nothing else needs it:

sudo ionice -c 3 -p 2240

For recurring jobs, start them with both nice and ionice from the beginning:

nice -n 10 ionice -c 3 /usr/local/bin/your_backup_script

For a database, find and index the slow queries. On MySQL and MariaDB, enable the slow query log; on PostgreSQL, use pg_stat_statements. Faster storage or more RAM for the database cache are the hardware fixes.

Memory pressure

Restart or reconfigure the process that grew too large, reduce the memory settings of caches and worker pools (for example, innodb_buffer_pool_size in MySQL or pm.max_children in PHP-FPM), or add RAM. Adding swap only hides the problem and keeps the load high.

Step 8 - Recording load history with sar

When the load was high overnight and is normal now, you need history. sysstat can record system statistics every 10 minutes. On Ubuntu, collection is disabled by default. Open its configuration file:

sudo nano /etc/default/sysstat

Set ENABLED to true:

ENABLED="true"

Restart the service so the collectors start:

sudo systemctl enable sysstat
sudo systemctl restart sysstat

After some time has passed, view the run queue and load averages for today:

sar -q
10:00:01      runq-sz  plist-sz   ldavg-1   ldavg-5  ldavg-15   blocked
10:10:01            2       412      1.84      1.62      1.40         0
10:20:01            9       431      7.91      5.20      2.44         0

sar -u shows CPU usage (including %iowait and %steal) and sar -d disk activity for the same intervals, which lets you match a past load spike to its cause.

Troubleshooting

The load is high but CPU, wa and PSI are all low. Look for D state processes stuck on a network file system (wchan containing nfs or rpc). Check the mount with mount -t nfs,nfs4 and the server it points to; those processes cannot be killed until the mount responds.

iostat shows %util at 100% but low latency. On NVMe and SSD devices, which handle many requests in parallel, %util can read 100% long before the device is saturated. Rely on r_await, w_await and the I/O pressure values instead.

The load drops as soon as you log in to investigate. Periodic jobs are the usual cause. Check systemctl list-timers and crontab -l for each user, and compare their schedules with the history from sar -q.

Conclusion

You learned that the load average counts runnable and I/O-blocked tasks, compared it with the CPU count, used vmstat, PSI, pidstat and iostat to classify it as CPU, I/O or memory load, and applied a matching fix with renice, ionice or a systemd CPUQuota. As next steps, set up alerts on load relative to nproc and on PSI values, tune your database's slow queries, and review how to find and remove zombie processes.