RabbitMQ is a widely used message broker for AMQP and other protocols. A single RabbitMQ node is a single point of failure: if it goes down, producers cannot publish and the queued messages are unavailable until it returns. In this tutorial you will build a three-node RabbitMQ cluster on Ubuntu 24.04, create quorum queues that replicate every message to all three nodes, test what happens when a node fails, and learn how to restart nodes safely during maintenance.
NoteClassic mirrored queues (the old
ha-modepolicies) were removed in RabbitMQ 4.0. Quorum queues are the supported way to replicate queue data, and they are what this guide uses.
Prerequisites
To follow this tutorial you need:
- Three servers running Ubuntu 24.04 LTS on x86_64, for example three CubePath VPS, each with at least 2 GB of RAM.
- A private network between them. This guide uses
10.10.0.21,10.10.0.22and10.10.0.23; replace them with your own private IPs. - A non-root user with
sudoprivileges on every server.
Unless a step says otherwise, run the commands on all three nodes.
| Hostname | Private IP | RabbitMQ node name |
|---|---|---|
| rabbit1 | 10.10.0.21 | rabbit@rabbit1 |
| rabbit2 | 10.10.0.22 | rabbit@rabbit2 |
| rabbit3 | 10.10.0.23 | rabbit@rabbit3 |
Step 1 - Preparing hostnames and the firewall
RabbitMQ nodes identify each other by name (rabbit@<short hostname>), so every node must resolve the others' short hostnames. Set the hostname on each server (use rabbit2 and rabbit3 on the other nodes):
sudo hostnamectl set-hostname rabbit1
Add all three nodes to /etc/hosts on every server:
sudo nano /etc/hosts
10.10.0.21 rabbit1
10.10.0.22 rabbit2
10.10.0.23 rabbit3
Check that the names resolve:
getent hosts rabbit1 rabbit2 rabbit3
Clustered nodes talk to each other on several ports. Allow them only from the private network (replace 10.10.0.0/24 with your subnet):
4369: epmd, the Erlang port mapper.25672: inter-node communication.35672-35682: connections from CLI tools on other nodes.5672: AMQP clients.15672: management UI and HTTP API.
sudo ufw allow from 10.10.0.0/24 to any port 4369,25672,5672,15672 proto tcp
sudo ufw allow from 10.10.0.0/24 to any port 35672:35682 proto tcp
Step 2 - Installing Erlang and RabbitMQ
Ubuntu's own packages lag behind the RabbitMQ release series, so install both Erlang and RabbitMQ from the repositories maintained by the RabbitMQ team. First install the prerequisites:
sudo apt update
sudo apt install -y curl gnupg apt-transport-https
Download the team's signing key into a keyring file:
sudo install -m 0755 -d /etc/apt/keyrings
curl -1sLf "https://keys.openpgp.org/vks/v1/by-fingerprint/0A9AF2115F4687BD29803A206B73A36E6026DFCA" \
| sudo gpg --dearmor -o /etc/apt/keyrings/com.rabbitmq.team.gpg
Add the Erlang and RabbitMQ repositories for Ubuntu 24.04 (noble). Each one is served from two mirrors:
sudo nano /etc/apt/sources.list.d/rabbitmq.list
deb [arch=amd64 signed-by=/etc/apt/keyrings/com.rabbitmq.team.gpg] https://deb1.rabbitmq.com/rabbitmq-erlang/ubuntu/noble noble main
deb [arch=amd64 signed-by=/etc/apt/keyrings/com.rabbitmq.team.gpg] https://deb2.rabbitmq.com/rabbitmq-erlang/ubuntu/noble noble main
deb [arch=amd64 signed-by=/etc/apt/keyrings/com.rabbitmq.team.gpg] https://deb1.rabbitmq.com/rabbitmq-server/ubuntu/noble noble main
deb [arch=amd64 signed-by=/etc/apt/keyrings/com.rabbitmq.team.gpg] https://deb2.rabbitmq.com/rabbitmq-server/ubuntu/noble noble main
Install the Erlang packages RabbitMQ needs, then the server itself:
sudo apt update
sudo apt install -y erlang-base erlang-asn1 erlang-crypto erlang-eldap erlang-ftp \
erlang-inets erlang-mnesia erlang-os-mon erlang-parsetools erlang-public-key \
erlang-runtime-tools erlang-snmp erlang-ssl erlang-syntax-tools erlang-tftp \
erlang-tools erlang-xmerl
sudo apt install -y rabbitmq-server
The package starts the service. Confirm the node is running and check its version:
sudo rabbitmq-diagnostics status | grep -E "RabbitMQ version|Erlang"
RabbitMQ version: 4.x.x
Erlang configuration: Erlang/OTP 27 ...
Enable the management plugin, which provides the web UI and the HTTP API used later in this guide. Plugins are enabled per node:
sudo rabbitmq-plugins enable rabbitmq_management
Step 3 - Configuring partition handling
When the network between nodes breaks, a cluster can split into groups that each believe the others are down. The pause_minority strategy makes nodes in the smaller side pause themselves, so only the majority keeps serving clients and no conflicting state is created. With three nodes this is the recommended setting. Create the configuration file on every node:
sudo nano /etc/rabbitmq/rabbitmq.conf
cluster_partition_handling = pause_minority
You will restart RabbitMQ in the next step, which applies it.
Step 4 - Sharing the Erlang cookie
Nodes and CLI tools authenticate to each other with a shared secret called the Erlang cookie, stored in /var/lib/rabbitmq/.erlang.cookie. Each node generated its own on first start, so they must all be made identical. On rabbit1, print the cookie:
sudo cat /var/lib/rabbitmq/.erlang.cookie
QWERTYUIOPASDFGHJKLZ
On rabbit2 and rabbit3, stop RabbitMQ, write the same value (without a trailing newline) and restore the strict permissions RabbitMQ requires. Replace cookie_from_rabbit1 with the value you just printed:
sudo systemctl stop rabbitmq-server
printf '%s' 'cookie_from_rabbit1' | sudo tee /var/lib/rabbitmq/.erlang.cookie > /dev/null
sudo chown rabbitmq:rabbitmq /var/lib/rabbitmq/.erlang.cookie
sudo chmod 400 /var/lib/rabbitmq/.erlang.cookie
sudo systemctl start rabbitmq-server
On rabbit1, restart RabbitMQ so it loads the new configuration file:
sudo systemctl restart rabbitmq-server
Verify the cookie is identical everywhere by comparing a hash on each node:
sudo sha256sum /var/lib/rabbitmq/.erlang.cookie
Step 5 - Forming the cluster
A node joins a cluster by resetting its own (empty) state and then contacting an existing member. Run these commands on rabbit2, then repeat them on rabbit3:
sudo rabbitmqctl stop_app
sudo rabbitmqctl reset
sudo rabbitmqctl join_cluster rabbit@rabbit1
sudo rabbitmqctl start_app
Stopping rabbit application on node rabbit@rabbit2 ...
Resetting node rabbit@rabbit2 ...
Clustering node rabbit@rabbit2 with rabbit@rabbit1
Starting node rabbit@rabbit2 ...
Warning
resetdeletes all data on the node where you run it. Only run it on a new node that is joining the cluster, never on rabbit1.
Check the cluster from any node:
sudo rabbitmqctl cluster_status
Cluster status of node rabbit@rabbit1 ...
Basics
Cluster name: rabbit@rabbit1
...
Disk Nodes
rabbit@rabbit1
rabbit@rabbit2
rabbit@rabbit3
Running Nodes
rabbit@rabbit1
rabbit@rabbit2
rabbit@rabbit3
All three nodes must appear under Running Nodes.
Step 6 - Creating users and a virtual host
Users, virtual hosts and permissions are cluster-wide, so run the following commands on one node only. The built-in guest user can only connect from localhost; delete it and create an administrator and an application user. Replace the passwords with strong ones:
sudo rabbitmqctl delete_user guest
sudo rabbitmqctl add_user admin 'your_admin_password'
sudo rabbitmqctl set_user_tags admin administrator
sudo rabbitmqctl add_vhost app
sudo rabbitmqctl set_permissions -p app admin ".*" ".*" ".*"
sudo rabbitmqctl add_user orders 'your_orders_password'
sudo rabbitmqctl set_permissions -p app orders ".*" ".*" ".*"
The three regular expressions in set_permissions grant configure, write and read access to all resources in the app virtual host. List the users to confirm:
sudo rabbitmqctl list_users
Listing users ...
user tags
admin [administrator]
orders []
Step 7 - Creating a quorum queue
A quorum queue is replicated with the Raft consensus algorithm. One replica is the leader; a message is confirmed to the publisher only after a majority of replicas (two out of three) have written it to disk. The queue type is chosen when the queue is declared and cannot be changed later, so applications must declare it with the x-queue-type: quorum argument.
To test without writing code, declare a queue through the HTTP API from any node:
curl -s -u admin:'your_admin_password' -X PUT \
-H 'Content-Type: application/json' \
-d '{"durable": true, "arguments": {"x-queue-type": "quorum"}}' \
http://localhost:15672/api/queues/app/orders
Check where its replicas live:
sudo rabbitmq-queues quorum_status orders --vhost app
Status of quorum queue orders on node rabbit@rabbit1 ...
Node Name Raft State Log Index Commit Index Snapshot Index Term
rabbit@rabbit1 leader 1 1 undefined 1
rabbit@rabbit2 follower 1 1 undefined 1
rabbit@rabbit3 follower 1 1 undefined 1
One leader and two followers: the queue has a full copy on every node. Publish a test message through the default exchange, which routes by queue name:
curl -s -u admin:'your_admin_password' -X POST \
-H 'Content-Type: application/json' \
-d '{"properties": {"delivery_mode": 2}, "routing_key": "orders", "payload": "order 1001", "payload_encoding": "string"}' \
http://localhost:15672/api/exchanges/app/amq.default/publish
{"routed":true}
Confirm the queue holds it:
sudo rabbitmqctl list_queues -p app name type messages
name type messages
orders quorum 1
In your application code, declare queues the same way. For example, with the Python pika library:
channel.queue_declare(queue="orders", durable=True, arguments={"x-queue-type": "quorum"})
Enable publisher confirms in the client (channel.confirm_delivery() in pika) so the application knows when a majority of replicas has stored each message.
Step 8 - Testing a node failure
Find the leader of the orders queue from step 7 and stop that node. In this example it is rabbit1, so run this on rabbit1:
sudo systemctl stop rabbitmq-server
On rabbit2, check the cluster and the queue:
sudo rabbitmqctl cluster_status | sed -n '/Running Nodes/,/Versions/p'
sudo rabbitmq-queues quorum_status orders --vhost app
rabbit1 is missing from the running nodes, and one of the remaining followers has been elected leader within a few seconds. The message is still there and the queue keeps accepting new ones, because two of three replicas are still a majority:
sudo rabbitmqctl list_queues -p app name messages
name messages
orders 1
Start rabbit1 again:
sudo systemctl start rabbitmq-server
Its replica rejoins as a follower and catches up automatically. Leadership does not move back on its own. To spread queue leaders evenly across the nodes again, run:
sudo rabbitmq-queues rebalance quorum
ImportantA three-node cluster tolerates the loss of one node. With two nodes down, quorum queues lose their majority and stop accepting messages until a second node returns, which is the intended behavior to prevent data loss.
Step 9 - Performing rolling maintenance
To upgrade or reboot nodes without downtime, restart them one at a time and only after checking that doing so will not leave any quorum queue without a majority. On the node you plan to stop, run:
sudo rabbitmq-queues check_if_node_is_quorum_critical
Checking if node rabbit@rabbit2 is critical for quorum of any queues ...
If the command exits successfully, the node can be stopped safely. If it lists queues, wait until their replicas on other nodes have caught up. Then put the node into maintenance mode, which moves leadership away and stops accepting client connections before shutdown:
sudo rabbitmq-upgrade drain
sudo systemctl restart rabbitmq-server
sudo rabbitmq-upgrade revive
Before moving to the next node, confirm the node is running and part of the cluster again:
sudo rabbitmq-diagnostics check_running
sudo rabbitmqctl cluster_status | sed -n '/Running Nodes/,/Versions/p'
Removing a node permanently
To decommission a node, first stop it, then remove it from any other node:
sudo rabbitmqctl forget_cluster_node rabbit@rabbit3
Quorum queues that had a replica on the removed node now run with fewer members. When you add a replacement node, extend all quorum queues to it with:
sudo rabbitmq-queues grow rabbit@rabbit4 all
Step 10 - Connecting clients to the cluster
Clients can connect to any node: RabbitMQ routes operations to the queue leader internally. For resilience, configure your clients with all three node addresses (most client libraries accept a list and reconnect to the next one), or put a TCP load balancer in front of port 5672. A typical connection URI for one node is:
amqp://orders:[email protected]:5672/app
The management UI is available on port 15672 of any node, for example http://10.10.0.21:15672, where you can log in as admin from the private network. It shows all nodes, their memory and disk alarms, and the replica state of every queue.
Troubleshooting
join_cluster fails with unable to connect to node rabbit@rabbit1: nodedown. Usually the cookies differ or a port is blocked. Compare the cookie hashes from step 4 and check that ports 4369 and 25672 are open between the nodes (nc -zv rabbit1 25672).
CLI commands fail with an authentication or cookie error. A CLI tool is using a different cookie from the server. Always run rabbitmqctl with sudo, so it uses the rabbitmq user's cookie.
A node stays paused after a network problem. With pause_minority, a node that cannot see the majority pauses itself on purpose and resumes when connectivity returns. Check the network between nodes and sudo journalctl -u rabbitmq-server on the paused node.
The node fails to start after a hostname change. RabbitMQ stores data under the node name. Set the final hostname before installing RabbitMQ; changing it later requires removing the node from the cluster and joining it again.
Conclusion
You now have a three-node RabbitMQ cluster on Ubuntu 24.04 where quorum queues keep a copy of every message on each node, survive the loss of any single node, and can be maintained one node at a time without downtime. Next, enable TLS for client and inter-node traffic, monitor the cluster with the built-in Prometheus plugin (rabbitmq_prometheus), and look at the Federation or Shovel plugins if you need to move messages between clusters in different locations.
