Healthchecks is an open source "dead man's switch" for scheduled jobs: every job pings a unique URL when it runs, and if a ping is late or reports a failure, Healthchecks sends you an alert. Uptime monitors cannot catch a backup that silently stopped running; Healthchecks can. In this tutorial you will run Healthchecks with Docker Compose and PostgreSQL on Ubuntu 24.04, publish it over HTTPS with Caddy, and wire it into a cron job and a systemd timer.
Prerequisites
To follow this guide you need:
- A server running Ubuntu 24.04 LTS with at least 1 GB of RAM, for example a CubePath VPS.
- A non-root user with
sudoprivileges. - Docker Engine and the Docker Compose plugin installed from Docker's official repository.
- A domain name with a DNS
Arecord pointing to your server, for examplehc.your_domain. Replaceyour_domainwith your own domain throughout the guide. - Ports 80 and 443 open in your firewall.
- SMTP credentials (host, port, user and password) from a mail provider, so Healthchecks can send email alerts.
Step 1 - Creating the project directory and secrets
Keep the Compose file and its environment file together in one directory:
sudo mkdir -p /opt/healthchecks
cd /opt/healthchecks
Healthchecks needs a Django SECRET_KEY, and PostgreSQL needs a password. Generate two random values and keep them for the next step:
openssl rand -hex 32
openssl rand -hex 16
3f9c1d0b8e6a4f27c55d2e9b1a7c0e84d9f3b6a2c1e5d7f8a9b0c1d2e3f4a5b6
b71e2c9d4a8f3e6015c7d9a2b4e6f801
Step 2 - Writing the environment file
Put all configuration in an .env file so the Compose file stays free of secrets:
sudo nano /opt/healthchecks/.env
Paste the following, replacing the placeholder values with your domain, the two random strings from Step 1 and your SMTP details:
# PostgreSQL
DB=postgres
DB_HOST=db
DB_PORT=5432
DB_NAME=hc
DB_USER=hc
DB_PASSWORD=your_db_password
# Healthchecks
SECRET_KEY=your_secret_key
DEBUG=False
ALLOWED_HOSTS=hc.your_domain
SITE_ROOT=https://hc.your_domain
SITE_NAME=Healthchecks
REGISTRATION_OPEN=False
# Outgoing email
DEFAULT_FROM_EMAIL=healthchecks@your_domain
EMAIL_HOST=smtp.your_provider.com
EMAIL_PORT=587
EMAIL_USE_TLS=True
EMAIL_HOST_USER=your_smtp_user
EMAIL_HOST_PASSWORD=your_smtp_password
A few settings matter more than the rest:
SITE_ROOTmust be the public HTTPS URL. Healthchecks uses it to build ping URLs and links in alert emails.ALLOWED_HOSTSmust contain the hostname you will use in the browser, or Django rejects the request with a400 Bad Request.REGISTRATION_OPEN=Falsestops strangers from creating accounts on your instance. You will create the first account from the command line.
Restrict the file, since it contains passwords:
sudo chmod 600 /opt/healthchecks/.env
Step 3 - Starting Healthchecks with Docker Compose
Create the Compose file:
sudo nano /opt/healthchecks/compose.yaml
services:
db:
image: postgres:16
restart: unless-stopped
environment:
POSTGRES_DB: ${DB_NAME}
POSTGRES_USER: ${DB_USER}
POSTGRES_PASSWORD: ${DB_PASSWORD}
volumes:
- db_data:/var/lib/postgresql/data
healthcheck:
test: ["CMD-SHELL", "pg_isready -U ${DB_USER} -d ${DB_NAME}"]
interval: 5s
retries: 10
web:
image: healthchecks/healthchecks:latest
restart: unless-stopped
env_file: .env
depends_on:
db:
condition: service_healthy
ports:
- "127.0.0.1:8000:8000"
volumes:
db_data:
The official image runs the web application and the background process that sends alerts (sendalerts) in the same container, and it applies database migrations when it starts, so you do not need a separate worker service. The port is bound to 127.0.0.1 because Docker publishes ports around UFW; only the reverse proxy should reach it.
Start the stack:
sudo docker compose up -d
Check that both containers are running:
sudo docker compose ps
NAME IMAGE SERVICE STATUS PORTS
healthchecks-db-1 postgres:16 db Up 40 seconds (healthy) 5432/tcp
healthchecks-web-1 healthchecks/healthchecks:latest web Up 30 seconds 127.0.0.1:8000->8000/tcp
If web keeps restarting, read its logs with sudo docker compose logs web --tail 50. A wrong database password or a missing variable is the usual cause.
Step 4 - Creating the administrator account
Because registration is closed, create your first user with Django's createsuperuser command inside the container:
sudo docker compose exec web /opt/healthchecks/manage.py createsuperuser
Enter an email address and a strong password when prompted. You will log in with that email address.
Step 5 - Publishing Healthchecks over HTTPS with Caddy
Caddy obtains and renews a Let's Encrypt certificate automatically. Install it from the Ubuntu repository:
sudo apt update
sudo apt install caddy
Replace the default site configuration:
sudo nano /etc/caddy/Caddyfile
hc.your_domain {
reverse_proxy 127.0.0.1:8000
}
Allow web traffic through UFW and reload Caddy:
sudo ufw allow 80/tcp
sudo ufw allow 443/tcp
sudo systemctl reload caddy
Confirm that the site answers over HTTPS with a valid certificate:
curl -sI https://hc.your_domain/accounts/login/ | head -n 1
HTTP/2 200
Open https://hc.your_domain in your browser and log in with the account from Step 4.
Step 6 - Creating a check
A check describes one job and how often it should report in. In the dashboard, click Add Check, then Change Schedule. You can choose between two kinds of schedule:
- Simple: a period (for example, 1 day) plus a grace time. The check goes down if no ping arrives within period + grace.
- Cron: a cron expression and time zone, identical to your crontab line. This is more accurate for jobs that run at fixed times.
For a nightly backup at 03:00, choose Cron, enter 0 3 * * *, pick your time zone and set a grace time slightly longer than the job usually takes, for example 1 hour. Copy the ping URL shown on the check page. It looks like this:
https://hc.your_domain/ping/5c1f0a4e-2d7b-4c33-9a0e-8f6b2d1e7a90
You can also create checks with the Management API. Create an API key in Settings > API Access, then:
curl -s -X POST https://hc.your_domain/api/v3/checks/ \
-H "X-Api-Key: your_api_key" \
-H "Content-Type: application/json" \
-d '{"name": "nightly-backup", "schedule": "0 3 * * *", "tz": "Europe/Madrid", "grace": 3600, "unique": ["name"]}'
The response is a JSON object that includes ping_url. The unique field makes the call idempotent: running it twice returns the existing check instead of creating a duplicate.
Test the ping URL from the server:
curl -fsS -m 10 --retry 3 https://hc.your_domain/ping/your_check_uuid
OK
The check changes from New to Up in the dashboard.
Step 7 - Monitoring a cron job
Healthchecks understands several ping endpoints appended to the check URL:
| Endpoint | Meaning |
|---|---|
/ping/<uuid> | Job finished successfully |
/ping/<uuid>/start | Job started (lets Healthchecks measure run time and detect hung jobs) |
/ping/<uuid>/fail | Job failed, alert now |
/ping/<uuid>/<exit-code> | 0 means success, any other number means failure |
The body of a POST request is stored with the ping (up to 10 KB by default, configurable with PING_BODY_LIMIT), which is a convenient way to attach the job output to the alert.
Instead of writing long one-liners in the crontab, install a small wrapper that sends the start signal, runs the command, and reports its exit code and the last part of its output:
sudo nano /usr/local/bin/hc-run
#!/usr/bin/env bash
# Usage: hc-run <ping-url> <command> [args...]
set -uo pipefail
url="$1"
shift
curl -fsS -m 10 --retry 3 -o /dev/null "${url}/start" || true
output="$("$@" 2>&1)"
status=$?
printf '%s\n' "$output" | tail -c 10000 | \
curl -fsS -m 10 --retry 3 -o /dev/null --data-binary @- "${url}/${status}" || true
exit "$status"
The script does not use set -e on purpose: it must keep going after the job fails so it can report the failure. Failed pings (|| true) never change the job's own exit code.
Make it executable:
sudo chmod 755 /usr/local/bin/hc-run
Now use it in root's crontab:
sudo crontab -e
0 3 * * * /usr/local/bin/hc-run https://hc.your_domain/ping/your_check_uuid /usr/local/bin/backup.sh
Test the wrapper by hand with a command that fails:
/usr/local/bin/hc-run https://hc.your_domain/ping/your_check_uuid false; echo "exit: $?"
exit: 1
The check goes Down in the dashboard and an alert is sent. Run it again with true to bring it back Up.
Step 8 - Monitoring a systemd timer
For jobs driven by systemd timers, let systemd send the pings. ExecStopPost= always runs, whether the job succeeded or not, and systemd exposes the outcome in the SERVICE_RESULT variable.
Edit the service that your timer starts, for example /etc/systemd/system/backup.service:
sudo nano /etc/systemd/system/backup.service
[Unit]
Description=Nightly backup
Wants=network-online.target
After=network-online.target
[Service]
Type=oneshot
Environment=HC_URL=https://hc.your_domain/ping/your_check_uuid
ExecStartPre=-/usr/bin/curl -fsS -m 10 --retry 3 -o /dev/null ${HC_URL}/start
ExecStart=/usr/local/bin/backup.sh
ExecStopPost=/bin/sh -c 'if [ "$SERVICE_RESULT" = success ]; then curl -fsS -m 10 --retry 3 -o /dev/null "$HC_URL"; else curl -fsS -m 10 --retry 3 -o /dev/null "$HC_URL/fail"; fi'
The leading - on ExecStartPre= makes the job run even if Healthchecks is unreachable. Reload systemd and run the service once to test it:
sudo systemctl daemon-reload
sudo systemctl start backup.service
sudo systemctl status backup.service --no-pager
The check should show a new successful ping with the job duration.
Step 9 - Adding notification channels
Open Integrations in your project. The email address of your account is added automatically as the first channel. Common additions:
- Email: add extra recipients such as a team mailbox. Each address must confirm the subscription.
- Webhook: send a request to any URL on up and down events, for example your chat system or an internal API.
- Slack, Discord, Telegram, ntfy and others: follow the instructions shown on each integration page.
Every new check is assigned to all channels by default. You can change that per check on the check page.
To confirm that email works end to end, use Django's built-in test command:
sudo docker compose exec web /opt/healthchecks/manage.py sendtestemail you@your_domain
If the message does not arrive, check the EMAIL_* values in .env and recreate the container with sudo docker compose up -d --force-recreate web.
Troubleshooting
The browser shows "Bad Request (400)". The hostname is missing from ALLOWED_HOSTS. Fix .env and run sudo docker compose up -d --force-recreate web.
Checks stay "New". No ping has arrived yet. Run the curl command from Step 6 on the server and look for errors such as DNS failures or certificate problems.
Alerts arrive for jobs that ran fine. The grace time is shorter than the job's run time, or the cron expression's time zone does not match the server's crontab time zone. Increase the grace time or fix the time zone on the check.
No alerts at all. Look for errors from the alert sender in the web container logs:
sudo docker compose logs web --tail 100
Upgrading
Pull the new image and recreate the container. Migrations run automatically on start:
cd /opt/healthchecks
sudo docker compose pull
sudo docker compose up -d
Back up the db_data volume or run pg_dump inside the db container before major upgrades.
Conclusion
You now run your own Healthchecks instance over HTTPS, with checks that alert you when a cron job or systemd timer fails, hangs or stops running. From here you can add a check for every backup and maintenance job on your servers, create one project per team or environment, and use the Management API to create checks automatically from your configuration management tool.
