Docker Swarm is the orchestration mode built into Docker Engine. It turns several Docker hosts into one cluster where you declare services (image, replicas, ports) and the managers keep that state running, rescheduling containers when a node fails. In this tutorial you will build a three-node Swarm on Ubuntu 24.04, deploy a replicated service behind the routing mesh, deploy a multi-service stack, and practice rolling updates, rollbacks and node maintenance.
Prerequisites
To follow this tutorial you need:
- Three servers running Ubuntu 24.04 LTS, for example three CubePath VPS instances. This guide calls them
manager1,worker1andworker2. - A private network between the three servers. The examples use
10.0.0.11,10.0.0.12and10.0.0.13; replace them with your own private IPs. - Docker Engine installed from Docker's official repository on every node, with your non-root user in the
dockergroup so you can rundockerwithoutsudo. - A non-root user with
sudoprivileges and UFW enabled on each node. - Unique hostnames on each node (
sudo hostnamectl set-hostname manager1) and time synchronization active, which Ubuntu provides by default withsystemd-timesyncd.
Confirm Docker works on each node before you continue:
docker version --format '{{.Server.Version}}'
Any current Docker Engine release prints its version, for example:
28.4.0
Step 1 - Opening the Swarm ports
Swarm nodes talk to each other on three ports. Allow them only from your private network, never from the public internet:
| Port | Protocol | Purpose |
|---|---|---|
| 2377 | TCP | Cluster management (managers only need to accept it, but allowing it everywhere lets you promote workers later) |
| 7946 | TCP and UDP | Node discovery and gossip |
| 4789 | UDP | Overlay network traffic (VXLAN) |
Run these commands on all three nodes, replacing 10.0.0.0/24 with your private subnet:
sudo ufw allow from 10.0.0.0/24 to any port 2377 proto tcp
sudo ufw allow from 10.0.0.0/24 to any port 7946 proto tcp
sudo ufw allow from 10.0.0.0/24 to any port 7946 proto udp
sudo ufw allow from 10.0.0.0/24 to any port 4789 proto udp
If you plan to use encrypted overlay networks (Step 5), also allow the ESP protocol (IP protocol 50), which UFW cannot express with ufw allow. Add this line to /etc/ufw/before.rules just before the final COMMIT:
sudo nano /etc/ufw/before.rules
-A ufw-before-input -p esp -s 10.0.0.0/24 -j ACCEPT
Reload UFW and check the rules:
sudo ufw reload
sudo ufw status
You should see the four port rules limited to your subnet:
2377/tcp ALLOW 10.0.0.0/24
7946/tcp ALLOW 10.0.0.0/24
7946/udp ALLOW 10.0.0.0/24
4789/udp ALLOW 10.0.0.0/24
Importantports that a Swarm service publishes (for example
--publish 80:80) are programmed by Docker directly in iptables and bypass UFW rules. Only publish ports you intend to expose, and use a firewall in front of the nodes if you need stricter filtering.
Step 2 - Initializing the Swarm
On manager1, initialize the cluster. The --advertise-addr flag tells the other nodes which address to use, which matters because your server has both a public and a private IP:
docker swarm init --advertise-addr 10.0.0.11
Docker creates the cluster certificate authority, makes this node the leader and prints the worker join command:
Swarm initialized: current node (k2f3x0wz1q8e9o7yq2n6m5l4a) is now a manager.
To add a worker to this swarm, run the following command:
docker swarm join --token SWMTKN-1-3pu6hszjas19xyp7ghgosyx9k8atbfcr8p2is99znpy26u2lkl-1awxwuwd3z9j1z3puu7rcgdbx 10.0.0.11:2377
To add a manager to this swarm, run 'docker swarm join-token manager' and follow the instructions.
Treat the token like a password: anyone with it and network access to port 2377 can join a node to your cluster. You can print it again at any time with docker swarm join-token worker.
Step 3 - Joining the worker nodes
On worker1 and worker2, run the join command that docker swarm init printed:
docker swarm join --token SWMTKN-1-your_worker_token 10.0.0.11:2377
Each worker confirms it joined:
This node joined a swarm as a worker.
Back on manager1, list the nodes. Node commands only work on managers:
docker node ls
All three nodes should be Ready and Active, with manager1 marked as Leader:
ID HOSTNAME STATUS AVAILABILITY MANAGER STATUS ENGINE VERSION
k2f3x0wz1q8e9o7yq2n6m5l4a * manager1 Ready Active Leader 28.4.0
p9d8c7b6a5z4y3x2w1v0u9t8s worker1 Ready Active 28.4.0
h1g2f3e4d5c6b7a8z9y0x1w2v worker2 Ready Active 28.4.0
Planning for manager high availability
A single manager is a single point of failure for the control plane: running containers keep working if it dies, but you cannot change or reschedule anything. Swarm managers use the Raft consensus algorithm, which needs a majority of managers online. Three managers tolerate the loss of one, five tolerate two. Always use an odd number.
For a production cluster, promote the two workers so all three nodes are managers (they still run workloads by default):
docker node promote worker1 worker2
Run docker node ls again and the MANAGER STATUS column shows Reachable for the promoted nodes. The rest of this tutorial works either way.
Step 4 - Deploying a replicated service
A service describes the desired state: which image to run, how many replicas and which ports to publish. Create a service with three replicas of traefik/whoami, a tiny web server that replies with the hostname of the container that served the request:
docker service create --name whoami --replicas 3 --publish published=8080,target=80 traefik/whoami
Check the service and where its tasks were placed:
docker service ls
docker service ps whoami
The service should report 3/3 replicas, spread across the nodes:
ID NAME MODE REPLICAS IMAGE PORTS
x1y2z3a4b5c6 whoami replicated 3/3 traefik/whoami:latest *:8080->80/tcp
ID NAME IMAGE NODE DESIRED STATE CURRENT STATE
q1w2e3r4t5y6 whoami.1 traefik/whoami:latest worker1 Running Running 20 seconds ago
a1s2d3f4g5h6 whoami.2 traefik/whoami:latest worker2 Running Running 20 seconds ago
z1x2c3v4b5n6 whoami.3 traefik/whoami:latest manager1 Running Running 20 seconds ago
Swarm's routing mesh publishes port 8080 on every node, and load balances requests across all replicas regardless of which node receives them. Send a few requests to one node:
for i in 1 2 3; do curl -s http://10.0.0.11:8080 | grep Hostname; done
Each request is answered by a different container:
Hostname: 4f1c2d3e4a5b
Hostname: 9a8b7c6d5e4f
Hostname: 1a2b3c4d5e6f
Scale the service up or down with a single command:
docker service scale whoami=5
Swarm starts two more tasks and docker service ls shows 5/5 when they are running.
Step 5 - Connecting services with an overlay network
Services that need to talk to each other should share a user-defined overlay network. On an overlay network, every service is reachable by its name through Swarm's internal DNS, which resolves to a virtual IP that load balances across the replicas.
Create an overlay network. The encrypted option enables IPsec encryption of traffic between nodes, which is why Step 1 allowed ESP:
docker network create --driver overlay --opt encrypted --attachable appnet
The --attachable flag also allows standalone containers to join the network, which is handy for debugging. Attach the whoami service to it:
docker service update --network-add appnet whoami
Now run a temporary container on the same network and call the service by name:
docker run --rm --network appnet curlimages/curl -s http://whoami | grep Hostname
Hostname: 9a8b7c6d5e4f
The name whoami resolved inside the overlay network without any published port.
Step 6 - Deploying a stack from a Compose file
For applications with several services, describe everything in a Compose file and deploy it as a stack. Remove the test service first:
docker service rm whoami
Create a directory and a stack file on manager1:
mkdir -p ~/stacks/demo
nano ~/stacks/demo/stack.yml
This stack runs a web tier with three replicas and a Redis instance pinned to nodes labelled for data, both on a private overlay network:
services:
web:
image: traefik/whoami
ports:
- "8080:80"
networks:
- backend
deploy:
replicas: 3
update_config:
parallelism: 1
delay: 10s
failure_action: rollback
order: start-first
restart_policy:
condition: on-failure
resources:
limits:
cpus: "0.50"
memory: 128M
redis:
image: redis:7-alpine
networks:
- backend
volumes:
- redis-data:/data
deploy:
replicas: 1
placement:
constraints:
- node.labels.tier == data
networks:
backend:
driver: overlay
volumes:
redis-data:
The placement constraint only matches nodes with the label tier=data, so add it to one node. Pinning a stateful service keeps its local volume on the same host:
docker node update --label-add tier=data worker2
Deploy the stack:
docker stack deploy -c ~/stacks/demo/stack.yml demo
Swarm prefixes every resource with the stack name. Verify the services:
docker stack services demo
ID NAME MODE REPLICAS IMAGE PORTS
m1n2b3v4c5x6 demo_redis replicated 1/1 redis:7-alpine
l1k2j3h4g5f6 demo_web replicated 3/3 traefik/whoami:latest *:8080->80/tcp
Run docker stack ps demo to confirm that demo_redis is running on worker2. To change the stack later, edit stack.yml and run the same docker stack deploy command again; Swarm only updates the services that changed.
Step 7 - Rolling updates and rollbacks
The update_config block in the stack controls how Swarm replaces tasks: one at a time (parallelism: 1), waiting 10 seconds between them, starting the new task before stopping the old one (start-first), and rolling back automatically if the update fails.
Trigger an update by changing the service definition. Here you set the WHOAMI_NAME environment variable, which whoami includes in its response; changing the image tag with --image works the same way:
docker service update --env-add WHOAMI_NAME=v2 demo_web
Watch the tasks being replaced one by one:
docker service ps demo_web
Old tasks appear with the Shutdown desired state and the new ones as Running. If the new version misbehaves, return to the previous service definition with a single command:
docker service rollback demo_web
demo_web
rollback: manually requested rollback
overall progress: rolling back update: 3 out of 3 tasks
verify: Service demo_web converged
Note
docker service updatechanges the running service, but the nextdocker stack deployapplies whateverstack.ymlsays. Keep the file as your source of truth and put your changes there once you settle on them.
Step 8 - Draining a node for maintenance
Before rebooting or upgrading a node, drain it so Swarm moves its tasks elsewhere:
docker node update --availability drain worker1
Check that no tasks are left running on it:
docker node ps worker1
All tasks listed for worker1 should have the Shutdown desired state, and docker stack ps demo shows their replacements on other nodes. Do your maintenance, then bring the node back:
docker node update --availability active worker1
Swarm does not move existing tasks back automatically. New tasks, updates and failures will use the node again; to rebalance immediately, force a redeploy of a service:
docker service update --force demo_web
To remove a node permanently, drain it, run docker swarm leave on the node itself, and then remove it from a manager with docker node rm worker1. Demote managers with docker node demote before removing them so the Raft quorum stays healthy.
Step 9 - Backing up the Swarm state
Every manager stores the cluster state (services, networks, secrets, configs) in /var/lib/docker/swarm. Back it up from a manager periodically. Docker must be stopped while you copy it, so do this on a non-leader manager if you have three:
sudo systemctl stop docker
sudo tar -czf /root/swarm-backup-$(date +%F).tar.gz -C /var/lib/docker swarm
sudo systemctl start docker
Confirm the node rejoined the cluster:
docker node ls
To restore after losing quorum, you extract the archive into /var/lib/docker/ on a fresh node with the same IP and run docker swarm init --force-new-cluster, then add the other managers again.
Troubleshooting
docker swarm joinhangs or times out: the worker cannot reach10.0.0.11:2377. Test withnc -zv 10.0.0.11 2377from the worker and review the UFW rules on the manager.- Services on different nodes cannot reach each other: UDP
4789or7946is blocked, or ESP is blocked for an encrypted network. Checksudo ufw statuson every node. - Tasks stuck in
Pending: no node satisfies the placement constraints or resource reservations. Rundocker service ps --no-trunc demo_redisto see the reason, for exampleno suitable node (scheduling constraints not satisfied on 3 nodes). This node is not a swarm manager: you ran a cluster command on a worker. Run it on a manager.
Conclusion
You now have a three-node Docker Swarm cluster with firewalled control ports, an overlay network with service discovery, and a stack that you can update with rolling deployments and roll back when needed. From here, store credentials with Docker secrets instead of environment variables, put a reverse proxy such as Traefik or Nginx in front of your published services, and host your own images in a private Docker registry so every node can pull them.
