systemd starts and supervises every service on a modern Linux server, so it already knows whether a service is running, how often it crashed, how much memory it uses and what it logged. In this tutorial you will use those built-in tools on Ubuntu 24.04 to check service health, read service logs, measure resource usage, restart crashed services automatically and send an alert when a service fails for good. Nginx is used as the example service, but every command works with any unit.

Prerequisites

To follow this tutorial you need:

  • A server running Ubuntu 24.04 LTS, such as a CubePath VPS, and a non-root user with sudo privileges.
  • Nginx installed as a sample service:
sudo apt update
sudo apt install nginx
  • For the alerting step: a webhook URL that accepts a JSON POST, for example an incoming webhook from Slack, Discord or Mattermost. You can skip it and test the alert with a log message only.

Step 1 - Checking the state of a service

systemctl status gives a summary of a unit: whether it is loaded and enabled, its current state, its main PID, its resource usage and its last log lines.

systemctl status nginx
● nginx.service - A high performance web server and a reverse proxy server
     Loaded: loaded (/usr/lib/systemd/system/nginx.service; enabled; preset: enabled)
     Active: active (running) since Thu 2026-09-25 10:02:14 UTC; 2h 5min ago
       Docs: man:nginx(8)
   Main PID: 1422 (nginx)
      Tasks: 3 (limit: 4553)
     Memory: 3.9M (peak: 4.4M)
        CPU: 61ms
     CGroup: /system.slice/nginx.service
             ├─1422 "nginx: master process /usr/sbin/nginx -g daemon on; master_process on;"
             ├─1423 "nginx: worker process"
             └─1424 "nginx: worker process"

The important parts are the Loaded line (enabled means it starts at boot) and the Active line. The most common states are:

StateMeaning
active (running)The service is running
active (exited)A one-shot service finished successfully
inactive (dead)The service is stopped
activating (auto-restart)The service crashed and systemd is about to restart it
failedThe service stopped with an error and will not be restarted

For scripts, use the is-* commands. They print a single word and set the exit code, so they can be used directly in if statements:

systemctl is-active nginx
systemctl is-enabled nginx
systemctl is-failed nginx
active
enabled
active

is-failed prints the current state and returns exit code 0 only when the unit is failed.

To see every failed unit on the server, which is the first thing to check after a reboot or an incident:

systemctl --failed
  UNIT LOAD ACTIVE SUB DESCRIPTION

0 loaded units listed.

You can also list all running services:

systemctl list-units --type=service --state=running

systemctl show prints machine-readable properties. These are the most useful ones for monitoring:

systemctl show nginx -p ActiveState -p SubState -p MainPID -p NRestarts -p ActiveEnterTimestamp
MainPID=1422
NRestarts=0
ActiveState=active
SubState=running
ActiveEnterTimestamp=Thu 2026-09-25 10:02:14 UTC

NRestarts counts how many times systemd restarted the service automatically since it was last started by hand, which makes it a good crash counter.

Step 2 - Reading service logs with journalctl

Everything a service writes to standard output or standard error, plus its syslog messages, ends up in the systemd journal tagged with the unit name. Filter by unit with -u:

sudo journalctl -u nginx -n 20 --no-pager

Follow new messages in real time, like tail -f:

sudo journalctl -u nginx -f

Narrow the output by time and priority. -p err shows only messages with priority error or worse, and -b limits the output to the current boot:

sudo journalctl -u nginx --since "1 hour ago"
sudo journalctl -u nginx -p err -b

When a service fails, systemctl status and journalctl -u together usually tell you why. The status shows the exit code, and the journal shows the error message the program printed just before it stopped.

Step 3 - Monitoring CPU and memory per service

Each service runs in its own control group (cgroup), so systemd can account CPU, memory and tasks per service. Ubuntu 24.04 uses cgroup v2 with this accounting enabled by default, which is where the Memory and CPU lines in systemctl status come from.

For a live view of all services, similar to top, use systemd-cgtop:

sudo systemd-cgtop --order=memory
Control Group                            Tasks   %CPU   Memory  Input/s Output/s
/                                          142    3.1     1.1G        -        -
system.slice                                58    1.9   612.4M        -        -
system.slice/mysql.service                  38    1.2   402.7M        -        -
system.slice/nginx.service                   3    0.0     3.9M        -        -

Press q to exit. Inside the tool, c, m and t sort by CPU, memory and tasks.

To read the numbers for one service, for example from a monitoring script, query the properties directly:

systemctl show nginx -p MemoryCurrent -p CPUUsageNSec -p TasksCurrent
MemoryCurrent=4071424
CPUUsageNSec=61234000
TasksCurrent=3

Memory is in bytes and CPU time in nanoseconds.

You can also cap a service so it cannot starve the rest of the server. systemctl set-property writes a persistent drop-in file and applies it immediately. This limits Nginx to 512 MB of RAM and half a CPU core:

sudo systemctl set-property nginx.service MemoryMax=512M CPUQuota=50%
systemctl show nginx -p MemoryMax -p CPUQuotaPerSecUSec
MemoryMax=536870912
CPUQuotaPerSecUSec=500ms

If a service exceeds MemoryMax, the kernel kills it inside its cgroup, and the journal shows an out-of-memory message for that unit. Remove the limits by running sudo systemctl revert nginx, which deletes all local drop-ins for the unit.

Step 4 - Restarting crashed services automatically

By default, Ubuntu's nginx.service is not restarted if it crashes. The Restart= setting tells systemd to start it again, and a start limit prevents an endless restart loop when the service cannot start at all.

Instead of editing the packaged unit file under /usr/lib/systemd/system, which would be overwritten on upgrades, create a drop-in override:

sudo systemctl edit nginx

The command opens an editor on /etc/systemd/system/nginx.service.d/override.conf. Add the following between the comment lines at the top of the file:

[Unit]
StartLimitIntervalSec=300
StartLimitBurst=5

[Service]
Restart=on-failure
RestartSec=5s

These settings mean:

  • Restart=on-failure restarts the service when it exits with a non-zero code, is killed by a signal, times out or trips its watchdog. A clean stop with systemctl stop is not restarted.
  • RestartSec=5s waits 5 seconds before each restart.
  • StartLimitIntervalSec=300 and StartLimitBurst=5 allow at most 5 starts in 5 minutes. After that, the unit goes to the failed state and stays there.

Save and close the editor. systemctl edit reloads the systemd configuration automatically. Check that the new values are active:

systemctl show nginx -p Restart -p RestartUSec -p StartLimitBurst
Restart=on-failure
RestartUSec=5s
StartLimitBurst=5

Now simulate a crash by killing the main process with SIGKILL:

sudo kill -9 "$(systemctl show -p MainPID --value nginx)"
sleep 7
systemctl show nginx -p ActiveState -p NRestarts
ActiveState=active
NRestarts=1

The journal records the crash and the restart:

sudo journalctl -u nginx -n 5 --no-pager
Sep 25 12:10:31 web01 systemd[1]: nginx.service: Main process exited, code=killed, status=9/KILL
Sep 25 12:10:31 web01 systemd[1]: nginx.service: Failed with result 'signal'.
Sep 25 12:10:36 web01 systemd[1]: nginx.service: Scheduled restart job, restart counter is at 1.
Sep 25 12:10:36 web01 systemd[1]: Starting nginx.service - A high performance web server and a reverse proxy server...
Sep 25 12:10:36 web01 systemd[1]: Started nginx.service - A high performance web server and a reverse proxy server.

Step 5 - Sending an alert when a service fails

Automatic restarts hide short crashes, but you still want to know when a service is down for good. The OnFailure= setting starts another unit when a unit enters the failed state. For a service with Restart=, that happens only after the start limit from Step 4 is reached, so you get one alert when restarting did not help instead of one per crash.

Start by storing the webhook URL in a file that only root can read:

sudo install -m 0600 /dev/null /etc/notify-failure.env
sudo nano /etc/notify-failure.env
WEBHOOK_URL=https://hooks.example.com/your_webhook_path

Next, create the script that sends the alert. It logs the failure to the journal and, if a webhook is configured, posts a short JSON message:

sudo nano /usr/local/bin/notify-failure
#!/usr/bin/env bash
set -euo pipefail

unit="$1"
host="$(hostname)"
message="Service ${unit} failed on ${host}"

logger -t notify-failure -p daemon.err "${message}"

if [[ -n "${WEBHOOK_URL:-}" ]]; then
    curl -fsS --max-time 10 -H 'Content-Type: application/json' \
        -d "{\"text\": \"${message}\"}" "${WEBHOOK_URL}"
fi

The text key works for Slack and Mattermost; for Discord, use content instead. Make the script executable:

sudo chmod 0755 /usr/local/bin/notify-failure

Create a template unit that runs the script. The %i specifier is replaced by whatever comes after the @ when the unit is started, which will be the name of the failed unit:

sudo nano /etc/systemd/system/[email protected]
[Unit]
Description=Send failure alert for %i

[Service]
Type=oneshot
EnvironmentFile=-/etc/notify-failure.env
ExecStart=/usr/local/bin/notify-failure %i

The - in front of the EnvironmentFile path makes the file optional. Reload systemd and test the alert directly, without breaking anything:

sudo systemctl daemon-reload
sudo systemctl start [email protected]
sudo journalctl -t notify-failure -n 1 --no-pager
Sep 25 12:20:03 web01 notify-failure[5120]: Service nginx failed on web01

If you configured a webhook, the message also appears in your chat channel.

Finally, attach the alert to Nginx. Open the override again:

sudo systemctl edit nginx

Add OnFailure= to the existing [Unit] section, so the file looks like this:

[Unit]
StartLimitIntervalSec=300
StartLimitBurst=5
OnFailure=notify-failure@%N.service

[Service]
Restart=on-failure
RestartSec=5s

%N expands to the unit name without its suffix, nginx, so the alert unit becomes [email protected]. Confirm the dependency:

systemctl show nginx -p OnFailure

You can add the same OnFailure= line to the override of any other service you care about, such as your database or application.

Step 6 - Checking service startup times

Slow services delay the boot and the recovery after a restart. systemd-analyze shows how long each unit took to start during the last boot:

systemd-analyze blame | head -n 10
5.112s apt-daily-upgrade.service
2.301s cloud-init.service
1.204s mysql.service
 311ms nginx.service

To see what a specific service waited for before it could start, use critical-chain:

systemd-analyze critical-chain nginx.service

The time after @ is when the unit became active, and the time after + is how long it took to start.

Troubleshooting

  • Start request repeated too quickly. The unit hit its start limit and is in the failed state. Fix the cause shown by journalctl -u, then run sudo systemctl reset-failed nginx and start it again.
  • Changes in the override have no effect. Check what systemd actually loaded with systemctl cat nginx, which prints the unit file followed by every drop-in. A setting in the wrong section, such as StartLimitBurst= under [Service], is ignored with a warning in the journal.
  • The alert unit fails. Run sudo systemctl status [email protected] and sudo journalctl -u [email protected]. A curl error usually means the webhook URL is wrong or blocked by an outbound firewall.
  • systemd-cgtop shows - for memory or CPU. Accounting is off for that unit. Enable it with sudo systemctl set-property your_service MemoryAccounting=yes CPUAccounting=yes.

Conclusion

You can now check the state of any service with systemctl, read its logs with journalctl, measure and limit its CPU and memory, have systemd restart it after a crash and receive an alert when restarting is not enough. As next steps, apply the same restart and alert override to every critical service on your server, learn more journalctl filters for log analysis, or ship these metrics to an external monitoring system such as Prometheus with the node exporter's systemd collector.