High CPU usage makes a server slow to respond, but the number alone does not say whether the cause is one runaway process, too little capacity, disk I/O or a noisy neighbor on the hypervisor. In this tutorial you will read the system-wide CPU picture, find the responsible process and thread with top, ps and pidstat, map it to its systemd service and apply a limit. The commands are for Ubuntu 24.04 and work the same on Debian 12.

Prerequisites

  • A server running Ubuntu 24.04 LTS, for example a CubePath VPS.
  • A non-root user with sudo privileges.
  • The sysstat package, which provides pidstat, mpstat and sar. You will install it in Step 1.

Step 1 - Installing sysstat

top and ps are part of procps, which is always installed. pidstat and mpstat come from sysstat:

sudo apt update
sudo apt install sysstat

Check that it works:

pidstat -V
sysstat version 12.6.1

Step 2 - Reading load average and CPU breakdown

Start with the number of CPUs and the load average:

nproc
uptime
4
 10:42:13 up 12 days,  3:01,  1 user,  load average: 6.12, 5.80, 3.95

The three values are the average number of tasks running or waiting over 1, 5 and 15 minutes. Compare them with nproc: a load of 6 on 4 CPUs means that, on average, two tasks were always waiting. The load rising from 3.95 (15 min) to 6.12 (1 min) means the problem is recent and growing.

Load average also counts tasks waiting on disk, so it does not prove the CPU is busy. Look at how CPU time is spent with mpstat, one line per CPU, every second, five times:

mpstat -P ALL 1 5
CPU    %usr   %nice    %sys %iowait    %irq   %soft  %steal  %guest  %gnice   %idle
all   71.40    0.00    6.10    0.50    0.00    0.80    1.20    0.00    0.00   20.00
  0   98.00    0.00    2.00    0.00    0.00    0.00    0.00    0.00    0.00    0.00
  1   64.00    0.00    9.00    1.00    0.00    1.00    2.00    0.00    0.00   23.00

Read the columns this way:

ColumnHigh value means
%usrApplication code is busy. Find the process in the next steps.
%sysTime in the kernel: many system calls, network interrupts, or memory pressure.
%iowaitCPUs are idle waiting for disk. This is a storage problem, not a CPU problem. Check with iostat -x 1.
%stealThe hypervisor gives your VM less CPU than it asks for. Sustained values above 10% point to the host, not your workload.
%idleFree capacity.

A single CPU at 100% while others idle, like CPU 0 above, usually means a single-threaded process is maxed out. Adding more vCPUs will not help it.

Step 3 - Finding the process with top

top shows live usage, sorted by CPU by default:

top

Useful keys inside top:

  • 1: toggle the per-CPU view.
  • P: sort by CPU, M: sort by memory.
  • c: show the full command line, which helps tell apart several php-fpm or python3 processes.
  • H: show threads instead of processes.
  • k: kill a process by PID (asks for confirmation).
  • q: quit.

The header line %Cpu(s): 71.4 us, 6.1 sy, 0.0 ni, 20.0 id, 0.5 wa, 0.0 hi, 0.8 si, 1.2 st is the same breakdown as mpstat, averaged across all CPUs.

In the process list, %CPU is relative to one CPU, so a multithreaded process can show 350% on a 4-CPU server. For a one-shot snapshot you can paste into a ticket, use batch mode:

top -b -n 1 -o %CPU | head -20

Step 4 - Listing processes with ps

ps gives a sortable, scriptable list. Show the top CPU consumers with their parent, owner and running time:

ps -eo pid,ppid,user,%cpu,%mem,etime,cmd --sort=-%cpu | head -10
    PID    PPID USER     %CPU %MEM     ELAPSED CMD
  18342       1 www-data 187.3  4.1    01:12:09 /usr/bin/php8.3 /var/www/app/artisan queue:work
   1043       1 mysql     22.6 18.9 12-03:01:44 /usr/sbin/mysqld

Keep in mind that %CPU in ps is the average over the whole lifetime of the process, not the current usage. A database running for 12 days with a spike in the last minute will show a low value. Use ps to identify candidates and pidstat (next step) to measure what they are doing right now.

To see if a process spawns many short-lived children, such as a cron job or a shell loop, look at the process tree:

ps -ef --forest | less

Step 5 - Measuring current usage with pidstat

pidstat reports per-process usage over an interval, which makes it the most accurate tool for "what is using CPU now". Sample every 2 seconds, 5 times:

pidstat -u 2 5
Average:      UID       PID    %usr %system  %guest   %wait    %CPU   CPU  Command
Average:       33     18342  176.20    9.40    0.00    2.10  185.60     -  php8.3
Average:      112      1043   18.00    4.30    0.00    0.60   22.30     -  mysqld

The Average: lines at the end summarize the whole sample. %wait is the time the process was ready to run but waiting for a CPU; high values mean the server has more work than CPUs.

Drill into a single process by its PID. With -t, pidstat shows each thread, which is useful for Java, Node.js worker pools or database servers:

pidstat -u -t -p 18342 1 5

Context switches reveal processes that spin or thrash between threads. cswch/s are voluntary switches (waiting for I/O or locks), nvcswch/s are involuntary (the scheduler preempted the process because its time slice ran out):

pidstat -w -p 18342 1 5

A very high nvcswch/s together with high %CPU confirms a CPU-bound process competing for cores.

Step 6 - Mapping the process to its service

Before changing anything, find out which service owns the process. systemctl status accepts a PID:

systemctl status 18342
● app-queue.service - Laravel queue worker
     Loaded: loaded (/etc/systemd/system/app-queue.service; enabled; preset: enabled)
     Active: active (running) since Thu 2026-09-25 09:30:04 UTC; 1h 12min ago
   Main PID: 18342 (php8.3)

systemd-cgtop shows CPU usage aggregated per service, which is useful when the load is spread across many worker processes:

sudo systemd-cgtop

Then check the service log for the reason it is busy: a retry loop, an error being repeated thousands of times, or an unusual volume of jobs:

sudo journalctl -u app-queue.service --since "30 minutes ago" | tail -50

Step 7 - Reducing the impact

The right fix depends on the cause you found: a bug, a stuck job, a missing database index or simply insufficient capacity. While you work on it, you can protect the rest of the system.

Lower the priority of a running process so interactive work and other services get CPU first (19 is the lowest priority):

sudo renice -n 10 -p 18342

For a systemd service, set a hard CPU limit. CPUQuota=150% allows at most one and a half CPUs:

sudo systemctl set-property app-queue.service CPUQuota=150%

The setting takes effect immediately and persists across reboots in a drop-in file. Verify it:

systemctl show app-queue.service -p CPUQuotaPerSecUSec
CPUQuotaPerSecUSec=1.500000s

To remove the limit later, run sudo systemctl set-property app-queue.service CPUQuota=.

If a process is stuck and must go, stop it through its service rather than killing the PID, so systemd does not restart it in a broken state:

sudo systemctl restart app-queue.service

Step 8 - Keeping CPU history with sar

Many CPU problems happen at night or disappear before you log in. sysstat can record CPU usage every 10 minutes so you can look back. Enable collection:

sudo sed -i 's/^ENABLED="false"/ENABLED="true"/' /etc/default/sysstat
sudo systemctl enable --now sysstat

After some time, show today's CPU history:

sar -u

To read a previous day, pass its data file. Files are stored in /var/log/sysstat/ and named after the day of the month, for example sa24 for the 24th:

sar -u -f /var/log/sysstat/sa24

Conclusion

You read the load average against the CPU count, separated real CPU work from I/O wait and steal time, identified the busy process and thread with top, ps and pidstat, traced it to its systemd service and limited it with CPUQuota. For recurring problems, the sar history shows when spikes begin. As next steps, look at memory and disk I/O with the same method, and set up alerting on load and CPU so you get notified before users notice.