A zombie process (shown as <defunct>) is a process that has already exited but still has an entry in the process table because its parent has not yet read its exit status. Zombies use no CPU or memory, but each one holds a PID, and a buggy parent that leaks thousands of them can eventually stop the system from starting new processes. In this tutorial you will create a zombie on purpose, find it and its parent, clear it the right way, and learn how to stop your services from producing them. The commands were tested on Ubuntu 24.04 and work the same on Debian 12 and Rocky Linux 9.
Prerequisites
To follow this guide you need:
- A server running Ubuntu 24.04 LTS (or any modern Linux distribution), for example a CubePath VPS.
- A non-root user with
sudoprivileges. - The
procpsandpsmiscpackages, which provideps,topandpstree. Both are installed by default on Ubuntu.
How a process becomes a zombie
Every process on Linux is created by a parent with fork(). When the child exits, the kernel frees its memory and open files but keeps a small entry in the process table with its PID and exit code. The entry stays there until the parent collects it by calling wait() or waitpid(), an action known as reaping.
The kernel notifies the parent with a SIGCHLD signal when a child exits. A well-written parent reacts to that signal and reaps the child within milliseconds, so you almost never see a zombie. A zombie that stays around means the parent is:
- not calling
wait()at all (a programming bug), - blocked or hung, so it never gets to its
wait()call, - ignoring
SIGCHLDwithout having told the kernel to auto-reap its children.
If the parent itself dies, its children (living or zombie) are adopted by PID 1 (systemd) or by the closest subreaper process, which reaps them automatically.
A zombie is different from an orphan: an orphan is a child that is still running after its parent died. Orphans are adopted and keep working normally; they are not a problem.
Step 1 - Creating a test zombie
To practice safely, create a zombie with standard tools. The following command starts a subshell that launches sleep 1 in the background and then replaces itself with sleep 300 using exec. The new sleep 300 process inherits the child but never calls wait(), so when sleep 1 exits it becomes a zombie for five minutes:
(sleep 1 & exec sleep 300) &
Wait two seconds and move on to the next step. The zombie disappears on its own when sleep 300 ends, so nothing is left behind.
Step 2 - Finding zombie processes
The quickest way to see whether the system has zombies is the Tasks line at the top of top:
top -bn1 | head -n 3
top - 10:14:02 up 12 days, 3:41, 1 user, load average: 0.02, 0.05, 0.01
Tasks: 118 total, 1 running, 116 sleeping, 0 stopped, 1 zombie
%Cpu(s): 0.0 us, 3.1 sy, 0.0 ni, 96.9 id, 0.0 wa, 0.0 hi, 0.0 si, 0.0 st
To list the zombies, filter ps output by the process state. Zombies have a STAT value starting with Z:
ps -eo pid,ppid,stat,etime,cmd | awk 'NR == 1 || $3 ~ /^Z/'
PID PPID STAT ELAPSED CMD
24519 24518 Z 00:41 [sleep] <defunct>
The columns that matter are:
PID: the zombie itself. You cannot do anything with this number directly.PPID: the parent that failed to reap it. This is the process you have to deal with.ELAPSED: how long it has existed. A zombie that is seconds old may be reaped any moment; one that is hours old is stuck.
To count zombies, which is useful when a parent leaks many of them:
ps -eo stat= | grep -c '^Z'
1
Step 3 - Identifying the parent process
Show the full details of the parent using the PPID from the previous step. Replace 24518 with your own value:
ps -o pid,user,stat,etime,cmd -p 24518
PID USER STAT ELAPSED CMD
24518 sammy S 00:52 sleep 300
To see where the parent sits in the process tree, and therefore which service or session started it, use pstree with the zombie's PID:
pstree -ps 24519
systemd(1)───sshd(812)───sshd(24301)───sshd(24388)───bash(24389)───sleep(24518)───sleep(24519)
If the parent belongs to a systemd service, systemctl status with the PID tells you the unit name directly:
systemctl status 24518
For a real service you would see a line such as ● php8.3-fpm.service - The PHP 8.3 FastCGI Process Manager at the top of the output. For the test zombie, the output shows your SSH session scope instead.
When many zombies share the same parent, list the parents ranked by the number of zombies they hold:
ps -eo ppid=,stat= | awk '$2 ~ /^Z/ {print $1}' | sort | uniq -c | sort -rn
1 24518
Step 4 - Removing the zombie
A zombie is already dead, so kill and even kill -9 on its PID have no effect. The only way to remove it is to get its entry reaped, which leaves three options, from least to most disruptive.
Option 1: Ask the parent to reap its children
If the parent handles SIGCHLD but missed a notification, sending the signal again can trigger the reap:
sudo kill -s SIGCHLD 24518
Check whether the zombie is gone:
ps -o pid,stat,cmd -p 24519
If the output still shows the Z state, the parent does not handle SIGCHLD (the case with the test sleep) and you need the next option.
Option 2: Restart the parent service
When the parent is part of a systemd service, restart the service. systemd stops every process in the unit's control group and reaps any leftovers itself:
sudo systemctl restart your_service
Replace your_service with the unit name you found in Step 3, for example php8.3-fpm. This is the normal fix in production because it gives you a clean process that starts reaping properly again.
Option 3: Terminate the parent process
If the parent is not a service, or you cannot restart the whole unit, terminate the parent. Its zombies are then adopted by systemd (or the nearest subreaper) and reaped immediately. Try a normal SIGTERM first so the parent can shut down cleanly:
kill 24518
Use sudo if the parent belongs to another user, and kill -9 only if the process ignores SIGTERM. Verify that the zombie count is back to zero:
ps -eo stat= | grep -c '^Z'
0
WarningNever terminate PID 1 or a critical parent such as
sshdor your database server just to clear a handful of zombies. A few zombies are harmless; plan the restart during a maintenance window instead.
Step 5 - Checking whether zombies are a real problem
A handful of short-lived zombies is normal. The real risk is running out of PIDs. Check the system-wide limit:
cat /proc/sys/kernel/pid_max
4194304
Services also have their own task limit set by systemd (TasksMax), which is usually much lower than pid_max. Zombies count against it, so a leaking service can hit its own limit long before the system does:
systemctl show -p TasksMax -p TasksCurrent your_service
TasksMax=4915
TasksCurrent=37
If TasksCurrent keeps growing while the service does the same amount of work, and most of the extra tasks are zombies, the service has a reaping bug. Restart it as a short-term fix and report or fix the bug in the application.
Step 6 - Preventing zombies
Zombies are a symptom of a parent that does not clean up after its children. The fix belongs in the application or in how it is run.
In your own code
Always wait for the child processes you start. In Python, call wait() or communicate() on every subprocess.Popen object, or use subprocess.run(), which waits for you:
import subprocess
result = subprocess.run(["/usr/local/bin/generate-report"], check=True)
In C, install a SIGCHLD handler that reaps every finished child in a loop, or set SIGCHLD to SIG_IGN if you never need the exit codes, which tells the kernel to reap children automatically:
#include <signal.h>
#include <sys/wait.h>
#include <errno.h>
static void reap_children(int sig) {
int saved_errno = errno;
(void)sig;
while (waitpid(-1, NULL, WNOHANG) > 0)
;
errno = saved_errno;
}
/* in main(): */
struct sigaction sa = {0};
sa.sa_handler = reap_children;
sa.sa_flags = SA_RESTART | SA_NOCLDSTOP;
sigaction(SIGCHLD, &sa, NULL);
In containers
Inside a container, your application often runs as PID 1. A normal application does not reap adopted orphans the way systemd does, so zombies pile up inside the container. Run a minimal init process in front of it. With Docker, add the --init flag:
docker run --init -d your_image
With Docker Compose, set init: true on the service:
services:
app:
image: your_image
init: true
Both options start tini as PID 1, which forwards signals to your application and reaps zombies.
For systemd services
Run long-lived daemons as systemd services rather than from nohup or a screen session. systemd tracks every process a service starts, cleans them all up on stop or restart, and lets you cap the damage of a leak with TasksMax. You can set a lower limit with a drop-in:
sudo systemctl edit your_service
[Service]
TasksMax=512
Apply it and confirm:
sudo systemctl daemon-reload
sudo systemctl restart your_service
systemctl show -p TasksMax your_service
Troubleshooting
The zombie's parent is PID 1. systemd reaps adopted children as soon as it receives SIGCHLD, so a zombie that stays attached to PID 1 for more than a few seconds is very unusual. Check journalctl -b -p err for systemd errors and consider a reboot during a maintenance window.
Zombies come back right after a restart. The bug is in the parent's code or in a plugin or script it runs. Use ps -eo pid,ppid,stat,lstart,cmd to see which child commands turn into zombies; that usually points to the code path that forgets to wait.
fork: Resource temporarily unavailable errors. The system or a service ran out of tasks. Count zombies per parent with the command from Step 3, restart the worst offender, and check the service's TasksMax.
Conclusion
You created a zombie process, found it with ps and top, traced its parent with pstree and systemctl, and cleared it by restarting or terminating that parent, the only method that works. You also saw how to prevent zombies by reaping children in code, using an init process in containers and running daemons under systemd. As next steps, learn to diagnose a high load average, set up monitoring that alerts on the zombie count, and review how your application starts subprocesses.
