A running container is not necessarily a working one: the process can be alive while the application is deadlocked, out of database connections or still loading. A Docker health check is a command that Docker runs inside the container at a fixed interval to decide whether it is healthy or unhealthy. In this tutorial you will add a health check to an Nginx image, watch it change state, configure health checks in Docker Compose so a web app waits for its database, and learn what Docker does, and does not do, with an unhealthy container.

Prerequisites

To follow this tutorial you need:

  • A server running Ubuntu 24.04 LTS, for example a CubePath VPS.
  • Docker Engine 25.0 or later with the Docker Compose plugin, installed from Docker's official repository.
  • A non-root user with sudo privileges who is a member of the docker group.
  • The jq tool to read JSON output: sudo apt install jq.

How health checks work

When a container has a health check, Docker runs the check command inside it and reads the exit code:

  • 0 means healthy.
  • 1 means unhealthy.

The container starts in the starting state. After the first successful check it becomes healthy. It becomes unhealthy after the check fails a number of times in a row (retries). Containers without a health check have no health status at all.

These options control the timing:

OptionDefaultMeaning
--interval30sTime between checks once the container is running
--timeout30sA check that takes longer than this counts as failed
--start-period0sGrace period after start: failures do not count, a success marks it healthy
--start-interval5sTime between checks during the start period (Docker 25.0 and later)
--retries3Consecutive failures needed to become unhealthy

Step 1 - Adding a HEALTHCHECK to a Dockerfile

Create a project directory for an Nginx image with a health check:

mkdir -p ~/healthcheck-demo && cd ~/healthcheck-demo
nano Dockerfile
FROM nginx:alpine

HEALTHCHECK --interval=10s --timeout=3s --start-period=10s --retries=3 \
  CMD wget -q --spider http://127.0.0.1/ || exit 1

A few details matter here:

  • The check runs inside the container, so it can only use tools present in the image. nginx:alpine ships BusyBox wget but not curl. Always pick a command that exists in your image, or install one.
  • wget --spider requests the page without saving it and exits with a non-zero code on connection errors and HTTP error responses.
  • || exit 1 normalizes any failure to exit code 1, which is what Docker expects (other non-zero codes are reserved).
  • 127.0.0.1 avoids surprises when localhost resolves to the IPv6 address ::1 and the server only listens on IPv4.

Build the image and start a container:

docker build -t nginx-health .
docker run -d --name web -p 8080:80 nginx-health

Step 2 - Watching the health status

Right after starting, the status column shows the starting state:

docker ps --filter name=web
CONTAINER ID   IMAGE          COMMAND                  CREATED         STATUS                            PORTS                  NAMES
b3c4d5e6f7a8   nginx-health   "/docker-entrypoint.…"   3 seconds ago   Up 2 seconds (health: starting)   0.0.0.0:8080->80/tcp   web

A few seconds later it becomes healthy:

... Up 15 seconds (healthy) ...

Docker keeps the last five check results, including their output, in the container state. Inspect them with jq:

docker inspect --format '{{json .State.Health}}' web | jq
{
  "Status": "healthy",
  "FailingStreak": 0,
  "Log": [
    {
      "Start": "2026-09-25T10:00:21.412Z",
      "End": "2026-09-25T10:00:21.468Z",
      "ExitCode": 0,
      "Output": ""
    }
  ]
}

Step 3 - Simulating a failure

Break the application without stopping the Nginx process. Removing the index page makes Nginx answer / with 403 Forbidden, so the check starts to fail while the container keeps running:

docker exec web rm /usr/share/nginx/html/index.html

In a second terminal, follow the health events Docker emits:

docker events --filter container=web --filter event=health_status

After three failed checks (about 30 seconds with a 10 second interval), the event appears:

2026-09-25T10:02:04.118Z container health_status: unhealthy b3c4d5e6f7a8... (image=nginx-health, name=web)

Press Ctrl+C to stop following events. The log explains why the check failed:

docker inspect --format '{{json .State.Health}}' web | jq '.Status, .Log[-1].Output'
"unhealthy"
"wget: server returned error: HTTP/1.1 403 Forbidden\n"

Notice that the container is still running. Standalone Docker does not restart unhealthy containers. Restart policies such as --restart unless-stopped only react when the main process exits. The health status is information that other tools act on: Docker Swarm replaces unhealthy tasks, Docker Compose uses it to order startup, and load balancers and monitoring can read it through the API.

Remove the test container:

docker rm -f web

Step 4 - Setting or overriding a health check at run time

You can add a health check to an image that has none, or replace the image's check, with docker run flags. This is useful for third-party images:

docker run -d --name redis \
  --health-cmd "redis-cli ping | grep -q PONG" \
  --health-interval 10s \
  --health-timeout 3s \
  --health-retries 3 \
  redis:7-alpine

--health-cmd runs through the container's shell, so the pipe works. After a few seconds, docker ps shows (healthy) for the redis container. To disable a health check defined in an image, use --no-healthcheck.

docker rm -f redis

Step 5 - Health checks in Docker Compose

In Compose, health checks are most useful together with depends_on. Without them, Compose only waits for a dependency container to be started, not for the service inside it to be ready, so your app crashes on its first database connection.

Create a Compose project with PostgreSQL and a web service that waits for it:

mkdir -p ~/compose-health && cd ~/compose-health
nano compose.yaml
services:
  db:
    image: postgres:17
    environment:
      POSTGRES_DB: app
      POSTGRES_USER: app
      POSTGRES_PASSWORD: your_strong_password
    volumes:
      - db-data:/var/lib/postgresql/data
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U app -d app"]
      interval: 10s
      timeout: 5s
      retries: 5
      start_period: 30s
      start_interval: 2s

  web:
    image: nginx:alpine
    ports:
      - "8080:80"
    depends_on:
      db:
        condition: service_healthy
    healthcheck:
      test: ["CMD", "wget", "-q", "--spider", "http://127.0.0.1/"]
      interval: 15s
      timeout: 3s
      retries: 3

volumes:
  db-data:

Replace your_strong_password with a real password. The test key accepts two forms:

  • ["CMD", "wget", ...] runs the command directly, without a shell.
  • ["CMD-SHELL", "..."] runs the string through /bin/sh -c, which you need for pipes, || or environment variable expansion.

pg_isready ships with the PostgreSQL image and returns 0 once the server accepts connections. The long start_period gives the first initialization of the database time to finish, while start_interval: 2s still detects readiness quickly.

Start the project and let Compose wait until every service is healthy:

docker compose up -d --wait

Compose starts db, holds web back until db reports healthy, and returns only when both are healthy:

 ✔ Network compose-health_default    Created
 ✔ Volume "compose-health_db-data"   Created
 ✔ Container compose-health-db-1     Healthy
 ✔ Container compose-health-web-1    Healthy

Check the status of both services:

docker compose ps
NAME                   IMAGE          SERVICE   STATUS                    PORTS
compose-health-db-1    postgres:17    db        Up 25 seconds (healthy)   5432/tcp
compose-health-web-1   nginx:alpine   web       Up 12 seconds (healthy)   0.0.0.0:8080->80/tcp

The --wait flag is also useful in CI: the command fails if a service becomes unhealthy, so tests never run against a half-started stack.

Writing good health checks

A health check should answer one question: can this container do its job right now?

  • Check the application, not the process. Requesting an HTTP endpoint or running the database's own readiness tool catches deadlocks and misconfiguration that a process check misses.
  • Keep it cheap. The check runs every interval for the lifetime of the container. Avoid endpoints that run heavy queries.
  • Avoid checking dependencies. If every web container fails its check when the database is down, an orchestrator restarts all of them, which does not fix the database. A dedicated /healthz endpoint that verifies only the process itself is the usual pattern.
  • Use tools that exist in the image. For images without curl or wget, use the runtime you already have, for example in a Python image:
HEALTHCHECK --interval=30s --timeout=5s --start-period=15s \
  CMD python -c "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/healthz', timeout=4)" || exit 1

Troubleshooting

  • Container stays in health: starting forever: the check never succeeds. Read the output with docker inspect --format '{{json .State.Health}}' container_name | jq and run the same command manually with docker exec.
  • exec: "curl": executable file not found: the tool is not in the image. Use wget, a language runtime, or install curl in the Dockerfile.
  • Service is healthy but unreachable from the host: the check runs inside the container, so it does not test port publishing or firewalls. Test from outside with curl http://your_server_ip:8080.
  • dependency failed to start: container ... is unhealthy: Compose gave up waiting. Increase start_period or retries for slow services, and check the dependency's logs with docker compose logs db.

Conclusion

You added health checks to a Dockerfile, a docker run command and a Compose project, saw how a container moves between starting, healthy and unhealthy, and used depends_on with service_healthy to start services in the right order. Next, deploy the same images to Docker Swarm, where unhealthy tasks are replaced automatically, and send the health_status events to your monitoring system so failures raise an alert.