A full monitoring stack such as Prometheus or Zabbix is the right tool for a fleet, but a single server often only needs a few answers: is a disk almost full, is memory running out, is the load too high, and are the important services running. A small, well-written Bash script can check all of that without installing an agent. In this tutorial you will write one health check script on Ubuntu 24.04, schedule it with a systemd timer, and receive an alert by email or webhook only when the state of the server changes.

Prerequisites

To follow this guide you need:

  • A server running Ubuntu 24.04 LTS, for example a CubePath VPS. Debian 12 works the same way.
  • A non-root user with sudo privileges.
  • Optionally, a working mail command if you want email alerts (a send-only Postfix relay is enough), or a Slack or Mattermost incoming webhook URL if you want chat alerts.

Step 1 - Installing the tools the script uses

The checks rely only on files in /proc and on standard commands (df, nproc, systemctl) that are already installed. For webhook alerts you also need curl and jq, which the script uses to build valid JSON:

sudo apt update
sudo apt install curl jq

Verify that both are available:

jq --version
jq-1.7.1

Step 2 - Writing the health check script

The script collects every problem it finds in a list, compares the result with the previous run, and only sends a notification when something changed: a new problem appears, a problem disappears, or everything recovers. This avoids receiving the same alert every five minutes.

Create the script in /usr/local/sbin, the standard place for administrator scripts that run as root:

sudo nano /usr/local/sbin/healthcheck

Add the following content:

#!/usr/bin/env bash
# Simple server health check: disk, memory, load and systemd services.
# Settings can be overridden in /etc/default/healthcheck.
set -euo pipefail

DISK_MAX=85            # alert when a filesystem is at or above this % used
MEM_MIN_AVAIL=10       # alert when available memory drops below this %
LOAD_MAX_PER_CPU=2     # alert when the 5-minute load exceeds CPUs x this value
SERVICES=(cron)        # systemd units that must be active
ALERT_EMAIL=""         # empty = no email
WEBHOOK_URL=""         # empty = no webhook (Slack/Mattermost compatible)

if [[ -r /etc/default/healthcheck ]]; then
    # shellcheck source=/dev/null
    . /etc/default/healthcheck
fi

STATE_FILE="${STATE_DIRECTORY:-/var/lib/healthcheck}/state"
HOST="$(hostname)"
problems=()

# Disk usage on real filesystems
while read -r pcent target; do
    used="${pcent%\%}"
    if (( used >= DISK_MAX )); then
        problems+=("Disk ${target} is ${used}% full (limit ${DISK_MAX}%)")
    fi
done < <(df --output=pcent,target -x tmpfs -x devtmpfs -x squashfs -x overlay | tail -n +2)

# Available memory
mem_total="$(awk '/^MemTotal:/ {print $2}' /proc/meminfo)"
mem_avail="$(awk '/^MemAvailable:/ {print $2}' /proc/meminfo)"
mem_avail_pct=$(( mem_avail * 100 / mem_total ))
if (( mem_avail_pct < MEM_MIN_AVAIL )); then
    problems+=("Available memory is ${mem_avail_pct}% (minimum ${MEM_MIN_AVAIL}%)")
fi

# Load average (5 minutes) compared with the number of CPUs
read -r _ load5 _ < /proc/loadavg
load_limit=$(( $(nproc) * LOAD_MAX_PER_CPU ))
if awk -v l="$load5" -v m="$load_limit" 'BEGIN { exit !(l >= m) }'; then
    problems+=("5-minute load is ${load5} (limit ${load_limit})")
fi

# systemd services
for svc in "${SERVICES[@]}"; do
    if ! systemctl is-active --quiet "$svc"; then
        problems+=("Service ${svc} is not active")
    fi
done

if (( ${#problems[@]} > 0 )); then
    status="$(printf '%s\n' "${problems[@]}")"
    subject="[ALERT] ${HOST}: ${#problems[@]} problem(s)"
else
    status="OK"
    subject="[OK] ${HOST}: all checks passed"
fi

# Always log the current result (goes to the journal under systemd)
printf '%s\n' "$status"

notify() {
    local subject="$1" body="$2"
    if [[ -n "$ALERT_EMAIL" ]]; then
        printf '%s\n' "$body" | mail -s "$subject" "$ALERT_EMAIL"
    fi
    if [[ -n "$WEBHOOK_URL" ]]; then
        jq -n --arg text "${subject}"$'\n'"${body}" '{text: $text}' \
            | curl -fsS -m 10 -H 'Content-Type: application/json' --data @- "$WEBHOOK_URL" > /dev/null
    fi
}

# Notify only when the result differs from the previous run
previous="$(cat "$STATE_FILE" 2>/dev/null || true)"
if [[ "$status" != "$previous" ]]; then
    mkdir -p "$(dirname "$STATE_FILE")"
    notify "$subject" "$status"
    printf '%s\n' "$status" > "$STATE_FILE"
fi

(( ${#problems[@]} == 0 ))

A few details worth knowing:

  • df -x excludes virtual filesystems such as tmpfs and container overlays, which would otherwise produce false alerts.
  • MemAvailable is the kernel's estimate of memory that can be used without swapping. It is a better signal than "free" memory, which is always low on a healthy server because Linux uses spare RAM as cache.
  • The load average is compared with the number of CPUs, so the same threshold works on a 2 vCPU and a 16 vCPU server. awk does the comparison because Bash cannot compare decimals.
  • The state file is written only after the notification succeeds. If the email or webhook fails, set -e stops the script and the next run tries again.
  • The last line makes the script exit with status 1 when there are problems, so the systemd unit shows as failed and systemctl --failed lists it.

Save the file and make it executable:

sudo chmod 0755 /usr/local/sbin/healthcheck

Step 3 - Configuring thresholds and alerts

Keep your settings in /etc/default/healthcheck so you can change them without editing the script:

sudo nano /etc/default/healthcheck

Set the values that apply to your server. Replace the service names, the email address and the webhook URL with your own, and leave a variable empty to disable that channel:

DISK_MAX=85
MEM_MIN_AVAIL=10
LOAD_MAX_PER_CPU=2
SERVICES=(cron nginx)
ALERT_EMAIL="[email protected]"
WEBHOOK_URL=""

The webhook URL is a secret, so restrict the file to root:

sudo chmod 0600 /etc/default/healthcheck

Step 4 - Running the script by hand

Run the script once as root:

sudo /usr/local/sbin/healthcheck; echo "exit code: $?"

On a healthy server the output is:

OK
exit code: 0

Because this is the first run, the result differs from the (missing) previous state, so you also receive an [OK] notification. This confirms that email or webhook delivery works.

Now simulate a failure by adding a unit that does not exist to the service list:

sudo sed -i 's/^SERVICES=(\(.*\))/SERVICES=(\1 does-not-exist)/' /etc/default/healthcheck
sudo /usr/local/sbin/healthcheck; echo "exit code: $?"
Service does-not-exist is not active
exit code: 1

You should receive an [ALERT] message. Run the script again and no new alert is sent, because the state has not changed. Remove the fake service to restore the configuration:

sudo sed -i 's/ does-not-exist)/)/' /etc/default/healthcheck
sudo /usr/local/sbin/healthcheck

The next run reports OK and sends a recovery notification.

Step 5 - Scheduling the check with a systemd timer

A systemd timer is preferable to cron here: the output lands in the journal, runs never overlap, and systemctl shows the last result. First create the service unit:

sudo nano /etc/systemd/system/healthcheck.service
[Unit]
Description=Server health check

[Service]
Type=oneshot
ExecStart=/usr/local/sbin/healthcheck
StateDirectory=healthcheck

StateDirectory=healthcheck makes systemd create /var/lib/healthcheck and pass its path to the script in $STATE_DIRECTORY. Next, create the timer that starts the service every five minutes:

sudo nano /etc/systemd/system/healthcheck.timer
[Unit]
Description=Run the server health check every 5 minutes

[Timer]
OnCalendar=*:0/5
AccuracySec=30s

[Install]
WantedBy=timers.target

Load the new units and enable the timer:

sudo systemctl daemon-reload
sudo systemctl enable --now healthcheck.timer

Confirm that the timer is scheduled:

systemctl list-timers healthcheck.timer
NEXT                        LEFT       LAST PASSED UNIT              ACTIVATES
Thu 2026-09-25 14:35:00 UTC 2min 1s left -    -      healthcheck.timer healthcheck.service

After the first run, read the result in the journal:

journalctl -u healthcheck.service -n 10 --no-pager
Sep 25 14:35:00 web01 systemd[1]: Starting healthcheck.service - Server health check...
Sep 25 14:35:00 web01 healthcheck[4120]: OK
Sep 25 14:35:00 web01 systemd[1]: healthcheck.service: Deactivated successfully.
Sep 25 14:35:00 web01 systemd[1]: Finished healthcheck.service - Server health check.

Troubleshooting

No email arrives: run echo test | mail -s test [email protected] and check journalctl -t postfix/smtp -n 20. If mail is not installed or the relay is not configured, the script fails at the notification step and the journal shows the error.

The webhook returns an error: curl -f makes the script fail on any HTTP error, and the response code appears in the journal. Test the URL by hand with the same jq ... | curl command. Discord webhooks expect a content field instead of text; change the jq filter to {content: $text} if you use Discord.

Disk alerts for a filesystem you do not care about: add another -x type option to the df line, for example -x vfat to ignore /boot/efi.

The unit shows as failed: this is expected while there are problems, since the script exits with status 1. Read journalctl -u healthcheck.service to see which check failed.

Conclusion

You now have a dependency-free health check that watches disk, memory, load and services, runs every five minutes from a systemd timer, and alerts you only when the state changes. From here you can add checks that matter to your workload (certificate expiry, a local HTTP endpoint with curl -f), keep the script in version control, and move to Prometheus with Node Exporter when you manage more than a handful of servers.