Docker Swarm is the orchestration mode built into Docker Engine. It turns several Docker hosts into one cluster where you declare services (image, replicas, ports) and the managers keep that state running, rescheduling containers when a node fails. In this tutorial you will build a three-node Swarm on Ubuntu 24.04, deploy a replicated service behind the routing mesh, deploy a multi-service stack, and practice rolling updates, rollbacks and node maintenance.

Prerequisites

To follow this tutorial you need:

  • Three servers running Ubuntu 24.04 LTS, for example three CubePath VPS instances. This guide calls them manager1, worker1 and worker2.
  • A private network between the three servers. The examples use 10.0.0.11, 10.0.0.12 and 10.0.0.13; replace them with your own private IPs.
  • Docker Engine installed from Docker's official repository on every node, with your non-root user in the docker group so you can run docker without sudo.
  • A non-root user with sudo privileges and UFW enabled on each node.
  • Unique hostnames on each node (sudo hostnamectl set-hostname manager1) and time synchronization active, which Ubuntu provides by default with systemd-timesyncd.

Confirm Docker works on each node before you continue:

docker version --format '{{.Server.Version}}'

Any current Docker Engine release prints its version, for example:

28.4.0

Step 1 - Opening the Swarm ports

Swarm nodes talk to each other on three ports. Allow them only from your private network, never from the public internet:

PortProtocolPurpose
2377TCPCluster management (managers only need to accept it, but allowing it everywhere lets you promote workers later)
7946TCP and UDPNode discovery and gossip
4789UDPOverlay network traffic (VXLAN)

Run these commands on all three nodes, replacing 10.0.0.0/24 with your private subnet:

sudo ufw allow from 10.0.0.0/24 to any port 2377 proto tcp
sudo ufw allow from 10.0.0.0/24 to any port 7946 proto tcp
sudo ufw allow from 10.0.0.0/24 to any port 7946 proto udp
sudo ufw allow from 10.0.0.0/24 to any port 4789 proto udp

If you plan to use encrypted overlay networks (Step 5), also allow the ESP protocol (IP protocol 50), which UFW cannot express with ufw allow. Add this line to /etc/ufw/before.rules just before the final COMMIT:

sudo nano /etc/ufw/before.rules
-A ufw-before-input -p esp -s 10.0.0.0/24 -j ACCEPT

Reload UFW and check the rules:

sudo ufw reload
sudo ufw status

You should see the four port rules limited to your subnet:

2377/tcp                   ALLOW       10.0.0.0/24
7946/tcp                   ALLOW       10.0.0.0/24
7946/udp                   ALLOW       10.0.0.0/24
4789/udp                   ALLOW       10.0.0.0/24

Step 2 - Initializing the Swarm

On manager1, initialize the cluster. The --advertise-addr flag tells the other nodes which address to use, which matters because your server has both a public and a private IP:

docker swarm init --advertise-addr 10.0.0.11

Docker creates the cluster certificate authority, makes this node the leader and prints the worker join command:

Swarm initialized: current node (k2f3x0wz1q8e9o7yq2n6m5l4a) is now a manager.

To add a worker to this swarm, run the following command:

    docker swarm join --token SWMTKN-1-3pu6hszjas19xyp7ghgosyx9k8atbfcr8p2is99znpy26u2lkl-1awxwuwd3z9j1z3puu7rcgdbx 10.0.0.11:2377

To add a manager to this swarm, run 'docker swarm join-token manager' and follow the instructions.

Treat the token like a password: anyone with it and network access to port 2377 can join a node to your cluster. You can print it again at any time with docker swarm join-token worker.

Step 3 - Joining the worker nodes

On worker1 and worker2, run the join command that docker swarm init printed:

docker swarm join --token SWMTKN-1-your_worker_token 10.0.0.11:2377

Each worker confirms it joined:

This node joined a swarm as a worker.

Back on manager1, list the nodes. Node commands only work on managers:

docker node ls

All three nodes should be Ready and Active, with manager1 marked as Leader:

ID                            HOSTNAME   STATUS    AVAILABILITY   MANAGER STATUS   ENGINE VERSION
k2f3x0wz1q8e9o7yq2n6m5l4a *   manager1   Ready     Active         Leader           28.4.0
p9d8c7b6a5z4y3x2w1v0u9t8s     worker1    Ready     Active                          28.4.0
h1g2f3e4d5c6b7a8z9y0x1w2v     worker2    Ready     Active                          28.4.0

Planning for manager high availability

A single manager is a single point of failure for the control plane: running containers keep working if it dies, but you cannot change or reschedule anything. Swarm managers use the Raft consensus algorithm, which needs a majority of managers online. Three managers tolerate the loss of one, five tolerate two. Always use an odd number.

For a production cluster, promote the two workers so all three nodes are managers (they still run workloads by default):

docker node promote worker1 worker2

Run docker node ls again and the MANAGER STATUS column shows Reachable for the promoted nodes. The rest of this tutorial works either way.

Step 4 - Deploying a replicated service

A service describes the desired state: which image to run, how many replicas and which ports to publish. Create a service with three replicas of traefik/whoami, a tiny web server that replies with the hostname of the container that served the request:

docker service create --name whoami --replicas 3 --publish published=8080,target=80 traefik/whoami

Check the service and where its tasks were placed:

docker service ls
docker service ps whoami

The service should report 3/3 replicas, spread across the nodes:

ID             NAME      MODE         REPLICAS   IMAGE                   PORTS
x1y2z3a4b5c6   whoami    replicated   3/3        traefik/whoami:latest   *:8080->80/tcp

ID             NAME       IMAGE                   NODE       DESIRED STATE   CURRENT STATE
q1w2e3r4t5y6   whoami.1   traefik/whoami:latest   worker1    Running         Running 20 seconds ago
a1s2d3f4g5h6   whoami.2   traefik/whoami:latest   worker2    Running         Running 20 seconds ago
z1x2c3v4b5n6   whoami.3   traefik/whoami:latest   manager1   Running         Running 20 seconds ago

Swarm's routing mesh publishes port 8080 on every node, and load balances requests across all replicas regardless of which node receives them. Send a few requests to one node:

for i in 1 2 3; do curl -s http://10.0.0.11:8080 | grep Hostname; done

Each request is answered by a different container:

Hostname: 4f1c2d3e4a5b
Hostname: 9a8b7c6d5e4f
Hostname: 1a2b3c4d5e6f

Scale the service up or down with a single command:

docker service scale whoami=5

Swarm starts two more tasks and docker service ls shows 5/5 when they are running.

Step 5 - Connecting services with an overlay network

Services that need to talk to each other should share a user-defined overlay network. On an overlay network, every service is reachable by its name through Swarm's internal DNS, which resolves to a virtual IP that load balances across the replicas.

Create an overlay network. The encrypted option enables IPsec encryption of traffic between nodes, which is why Step 1 allowed ESP:

docker network create --driver overlay --opt encrypted --attachable appnet

The --attachable flag also allows standalone containers to join the network, which is handy for debugging. Attach the whoami service to it:

docker service update --network-add appnet whoami

Now run a temporary container on the same network and call the service by name:

docker run --rm --network appnet curlimages/curl -s http://whoami | grep Hostname
Hostname: 9a8b7c6d5e4f

The name whoami resolved inside the overlay network without any published port.

Step 6 - Deploying a stack from a Compose file

For applications with several services, describe everything in a Compose file and deploy it as a stack. Remove the test service first:

docker service rm whoami

Create a directory and a stack file on manager1:

mkdir -p ~/stacks/demo
nano ~/stacks/demo/stack.yml

This stack runs a web tier with three replicas and a Redis instance pinned to nodes labelled for data, both on a private overlay network:

services:
  web:
    image: traefik/whoami
    ports:
      - "8080:80"
    networks:
      - backend
    deploy:
      replicas: 3
      update_config:
        parallelism: 1
        delay: 10s
        failure_action: rollback
        order: start-first
      restart_policy:
        condition: on-failure
      resources:
        limits:
          cpus: "0.50"
          memory: 128M

  redis:
    image: redis:7-alpine
    networks:
      - backend
    volumes:
      - redis-data:/data
    deploy:
      replicas: 1
      placement:
        constraints:
          - node.labels.tier == data

networks:
  backend:
    driver: overlay

volumes:
  redis-data:

The placement constraint only matches nodes with the label tier=data, so add it to one node. Pinning a stateful service keeps its local volume on the same host:

docker node update --label-add tier=data worker2

Deploy the stack:

docker stack deploy -c ~/stacks/demo/stack.yml demo

Swarm prefixes every resource with the stack name. Verify the services:

docker stack services demo
ID             NAME         MODE         REPLICAS   IMAGE                   PORTS
m1n2b3v4c5x6   demo_redis   replicated   1/1        redis:7-alpine
l1k2j3h4g5f6   demo_web     replicated   3/3        traefik/whoami:latest   *:8080->80/tcp

Run docker stack ps demo to confirm that demo_redis is running on worker2. To change the stack later, edit stack.yml and run the same docker stack deploy command again; Swarm only updates the services that changed.

Step 7 - Rolling updates and rollbacks

The update_config block in the stack controls how Swarm replaces tasks: one at a time (parallelism: 1), waiting 10 seconds between them, starting the new task before stopping the old one (start-first), and rolling back automatically if the update fails.

Trigger an update by changing the service definition. Here you set the WHOAMI_NAME environment variable, which whoami includes in its response; changing the image tag with --image works the same way:

docker service update --env-add WHOAMI_NAME=v2 demo_web

Watch the tasks being replaced one by one:

docker service ps demo_web

Old tasks appear with the Shutdown desired state and the new ones as Running. If the new version misbehaves, return to the previous service definition with a single command:

docker service rollback demo_web
demo_web
rollback: manually requested rollback
overall progress: rolling back update: 3 out of 3 tasks
verify: Service demo_web converged

Step 8 - Draining a node for maintenance

Before rebooting or upgrading a node, drain it so Swarm moves its tasks elsewhere:

docker node update --availability drain worker1

Check that no tasks are left running on it:

docker node ps worker1

All tasks listed for worker1 should have the Shutdown desired state, and docker stack ps demo shows their replacements on other nodes. Do your maintenance, then bring the node back:

docker node update --availability active worker1

Swarm does not move existing tasks back automatically. New tasks, updates and failures will use the node again; to rebalance immediately, force a redeploy of a service:

docker service update --force demo_web

To remove a node permanently, drain it, run docker swarm leave on the node itself, and then remove it from a manager with docker node rm worker1. Demote managers with docker node demote before removing them so the Raft quorum stays healthy.

Step 9 - Backing up the Swarm state

Every manager stores the cluster state (services, networks, secrets, configs) in /var/lib/docker/swarm. Back it up from a manager periodically. Docker must be stopped while you copy it, so do this on a non-leader manager if you have three:

sudo systemctl stop docker
sudo tar -czf /root/swarm-backup-$(date +%F).tar.gz -C /var/lib/docker swarm
sudo systemctl start docker

Confirm the node rejoined the cluster:

docker node ls

To restore after losing quorum, you extract the archive into /var/lib/docker/ on a fresh node with the same IP and run docker swarm init --force-new-cluster, then add the other managers again.

Troubleshooting

  • docker swarm join hangs or times out: the worker cannot reach 10.0.0.11:2377. Test with nc -zv 10.0.0.11 2377 from the worker and review the UFW rules on the manager.
  • Services on different nodes cannot reach each other: UDP 4789 or 7946 is blocked, or ESP is blocked for an encrypted network. Check sudo ufw status on every node.
  • Tasks stuck in Pending: no node satisfies the placement constraints or resource reservations. Run docker service ps --no-trunc demo_redis to see the reason, for example no suitable node (scheduling constraints not satisfied on 3 nodes).
  • This node is not a swarm manager: you ran a cluster command on a worker. Run it on a manager.

Conclusion

You now have a three-node Docker Swarm cluster with firewalled control ports, an overlay network with service discovery, and a stack that you can update with rolling deployments and roll back when needed. From here, store credentials with Docker secrets instead of environment variables, put a reverse proxy such as Traefik or Nginx in front of your published services, and host your own images in a private Docker registry so every node can pull them.