Redis Cluster splits your data across several primary nodes, so the dataset and the write load can grow beyond a single server, and keeps a replica of each primary to fail over automatically. In this tutorial you will build a cluster of three primaries and three replicas on three Ubuntu 24.04 servers, check how keys are distributed, test a failover and add a fourth shard.

Prerequisites

To follow this tutorial, you need:

  • Three servers running Ubuntu 24.04 LTS, for example three CubePath VPS, connected by a private network. This guide uses:
HostnamePrivate IPRedis instances
node110.0.0.11ports 7000 and 7001
node210.0.0.12ports 7000 and 7001
node310.0.0.13ports 7000 and 7001
  • A non-root user with sudo privileges on each server.
  • UFW enabled, with SSH allowed.

Each server runs two Redis instances. When the cluster is created, redis-cli places every replica on a different server than its primary, so losing one server never loses a shard.

How Redis Cluster shards data

The key space is divided into 16,384 hash slots. Redis computes CRC16(key) mod 16384 to find the slot of each key, and every primary owns a range of slots. A client can send a command to any node; if the key lives elsewhere, the node answers with a MOVED redirection and cluster-aware clients update their slot map.

Two consequences matter when you design your keys:

  • Commands that touch several keys (MGET, transactions, Lua scripts) only work when all keys are in the same slot. Use hash tags to force this: only the part inside {} is hashed, so {user:42}:profile and {user:42}:cart share a slot.
  • Redis Cluster supports only database 0; SELECT is not available.

Each node also opens a cluster bus port, the data port plus 10000 (17000 and 17001 here), which the nodes use to exchange state and detect failures.

Step 1 - Installing Redis

Run this step on all three servers:

sudo apt update
sudo apt install redis-server

The package starts a standalone instance on port 6379 that you will not use. Stop and disable it:

sudo systemctl disable --now redis-server

Confirm the version:

redis-server --version
Redis server v=7.0.15 sha=00000000:0 malloc=jemalloc-5.3.0 bits=64 build=...

Step 2 - Opening the firewall

Allow the data ports and the cluster bus ports only from the private subnet, on every server:

sudo ufw allow from 10.0.0.0/24 to any port 7000:7001 proto tcp
sudo ufw allow from 10.0.0.0/24 to any port 17000:17001 proto tcp

Check with sudo ufw status. Both ranges should appear with ALLOW from 10.0.0.0/24.

Step 3 - Configuring the cluster instances

Run this step on all three servers. Create a data directory for each instance:

sudo mkdir -p /var/lib/redis-cluster/7000 /var/lib/redis-cluster/7001
sudo chown -R redis:redis /var/lib/redis-cluster
sudo chmod 750 /var/lib/redis-cluster/7000 /var/lib/redis-cluster/7001

Generate one password for the whole cluster, run once and reuse it on every server:

openssl rand -base64 32

Create the configuration for the instance on port 7000:

sudo mkdir -p /etc/redis-cluster
sudo nano /etc/redis-cluster/7000.conf

Replace 10.0.0.11 with the private IP of the server you are on and your_cluster_password with the generated password:

port 7000
bind 10.0.0.11 127.0.0.1
daemonize no
dir /var/lib/redis-cluster/7000

cluster-enabled yes
cluster-config-file nodes-7000.conf
cluster-node-timeout 5000

requirepass your_cluster_password
masterauth your_cluster_password

appendonly yes

The key directives:

  • cluster-enabled yes: start the instance in cluster mode.
  • cluster-config-file: a state file that Redis writes and updates itself. Never edit it by hand.
  • cluster-node-timeout 5000: after 5 seconds without a reply, a node is considered failing and a replica can take over.
  • requirepass and masterauth: the same password on every node, because any replica may need to authenticate to any primary.

Copy the file for the second instance, replacing every 7000 with 7001:

sudo sed 's/7000/7001/g' /etc/redis-cluster/7000.conf | sudo tee /etc/redis-cluster/7001.conf > /dev/null
sudo chown redis:redis /etc/redis-cluster/*.conf
sudo chmod 640 /etc/redis-cluster/*.conf

Step 4 - Running the instances with systemd

A systemd template unit runs any number of instances from one file, with the port as the instance name:

sudo nano /etc/systemd/system/[email protected]
[Unit]
Description=Redis Cluster node on port %i
After=network-online.target
Wants=network-online.target

[Service]
User=redis
Group=redis
ExecStart=/usr/bin/redis-server /etc/redis-cluster/%i.conf
Restart=on-failure
LimitNOFILE=65535

[Install]
WantedBy=multi-user.target

Start both instances and enable them at boot:

sudo systemctl daemon-reload
sudo systemctl enable --now redis-cluster@7000 redis-cluster@7001

Export the password so redis-cli uses it without exposing it on the command line, then check that both instances answer:

export REDISCLI_AUTH='your_cluster_password'
redis-cli -p 7000 ping
redis-cli -p 7001 ping
PONG
PONG

At this point the instances run in cluster mode but do not know each other yet. If an instance fails to start, check sudo journalctl -u redis-cluster@7000 -n 50.

Step 5 - Creating the cluster

Run this step once, from node1. List the three primaries first and then the three replica candidates; --cluster-replicas 1 assigns one replica to each primary:

redis-cli --cluster create \
  10.0.0.11:7000 10.0.0.12:7000 10.0.0.13:7000 \
  10.0.0.11:7001 10.0.0.12:7001 10.0.0.13:7001 \
  --cluster-replicas 1

redis-cli shows the proposed layout and asks for confirmation:

>>> Performing hash slots allocation on 6 nodes...
Master[0] -> Slots 0 - 5460
Master[1] -> Slots 5461 - 10922
Master[2] -> Slots 10923 - 16383
Adding replica 10.0.0.12:7001 to 10.0.0.11:7000
Adding replica 10.0.0.13:7001 to 10.0.0.12:7000
Adding replica 10.0.0.11:7001 to 10.0.0.13:7000
...
Can I set the above configuration? (type 'yes' to accept):

Check that no replica sits on the same server as its primary, then type yes. The command ends with:

[OK] All nodes agree about slots configuration.
>>> Check for open slots...
>>> Check slots coverage...
[OK] All 16384 slots covered.

Step 6 - Verifying the cluster

Check the overall state from any node:

redis-cli -p 7000 cluster info | head -n 7
cluster_state:ok
cluster_slots_assigned:16384
cluster_slots_ok:16384
cluster_slots_pfail:0
cluster_slots_fail:0
cluster_known_nodes:6
cluster_size:3

cluster_state:ok means every slot is served. For a readable summary of primaries, replicas and slot counts, use:

redis-cli --cluster check 10.0.0.11:7000

Now write some keys. The -c flag makes redis-cli follow redirections like a cluster-aware client:

redis-cli -c -h 10.0.0.11 -p 7000 set user:1 alice
redis-cli -c -h 10.0.0.11 -p 7000 set user:2 bob
-> Redirected to slot [10778] located at 10.0.0.12:7000
OK
-> Redirected to slot [6777] located at 10.0.0.12:7000
OK

Both keys hash to slots in the range owned by node2, so each command was redirected there. You can ask which slot a key belongs to with cluster keyslot, and see hash tags in action:

redis-cli -p 7000 cluster keyslot '{user:1}:profile'
redis-cli -p 7000 cluster keyslot '{user:1}:cart'

Both commands return the same slot, so these two keys can be used together in one MGET or transaction.

Step 7 - Testing automatic failover

Find the node ID and role of each instance:

redis-cli -p 7000 cluster nodes

Each line shows the node ID, address, flags (master, slave, myself) and, for primaries, their slot ranges. Pick the primary on node1 (10.0.0.11:7000) and stop it:

sudo systemctl stop redis-cluster@7000

Wait a few seconds longer than cluster-node-timeout, then check the cluster from another server, for example node2:

redis-cli -h 10.0.0.12 -p 7000 cluster nodes | grep -E 'master|fail'

The stopped node is flagged master,fail and its former replica (on node2, port 7001) is now a master that owns slots 0-5460. cluster info reports cluster_state:ok again, and the keys in those slots are still readable.

Start the failed instance again on node1:

sudo systemctl start redis-cluster@7000

It rejoins as a replica of the promoted node. To return the original primary to its role, run a manual failover on it; this is coordinated with its primary and loses no writes:

redis-cli -p 7000 cluster failover

Step 8 - Adding a shard and resharding

To grow the cluster, prepare two new instances as in Steps 3 and 4 (for example ports 7000 and 7001 on a fourth server, 10.0.0.14, with the same firewall rules). Then add the new primary, pointing to any existing node:

redis-cli --cluster add-node 10.0.0.14:7000 10.0.0.11:7000

Add its replica, telling redis-cli which primary to follow. Get the new primary's node ID from cluster nodes first:

redis-cli --cluster add-node 10.0.0.14:7001 10.0.0.11:7000 \
  --cluster-slave --cluster-master-id <new_primary_node_id>

The new primary has no slots yet. Move a fair share of slots to it from the existing primaries:

redis-cli --cluster rebalance 10.0.0.11:7000 --cluster-use-empty-masters

Keys move with their slots while the cluster keeps serving traffic. When it finishes, redis-cli --cluster check 10.0.0.11:7000 shows four primaries with about 4,096 slots each.

To remove a primary later, first move all its slots away with redis-cli --cluster reshard, then remove the empty node with redis-cli --cluster del-node 10.0.0.11:7000 <node_id>.

Step 9 - Connecting an application

Your application needs a cluster-aware client that follows MOVED redirections and refreshes the slot map after failovers. With Python on a machine in the private network:

sudo apt install python3-redis
nano cluster_test.py
from redis.cluster import RedisCluster

rc = RedisCluster(
    host="10.0.0.11",
    port=7000,
    password="your_cluster_password",
    decode_responses=True,
)

rc.set("{user:1}:profile", "alice")
rc.set("{user:1}:cart", "3 items")
print(rc.mget("{user:1}:profile", "{user:1}:cart"))
python3 cluster_test.py
['alice', '3 items']

The client discovers the rest of the nodes from the first one, so a single reachable address is enough to start, but list several in production if your client supports it.

Troubleshooting

  • redis-cli --cluster create hangs on Waiting for the cluster to join: the cluster bus ports (17000 and 17001) are blocked between servers. Check the UFW rules on every node.
  • [ERR] Node ... is not empty: the instance already has data or a nodes-*.conf from a previous attempt. On that node, stop the instance, delete the files in its directory under /var/lib/redis-cluster/, and start it again.
  • CROSSSLOT Keys in request don't hash to the same slot: a multi-key command used keys from different slots. Use a common hash tag for keys that must be accessed together.
  • cluster_state:fail after a server outage: a primary and its replica were both lost, so some slots have no owner. Bring one of them back; that is why replicas must live on a different server than their primary.

Conclusion

You built a Redis Cluster with three shards and one replica each, spread across three servers, verified slot distribution, survived the loss of a primary and grew the cluster with a new shard. Next, size maxmemory on each instance to leave headroom for the server, schedule backups of the RDB and AOF files on every primary, and monitor cluster_state and replication lag. If your dataset fits on one server and you only need failover, Redis Sentinel is a simpler alternative.