Apache Cassandra is a distributed wide-column NoSQL database built for high write throughput and no single point of failure: every node is equal, and data is replicated across nodes according to the replication factor you choose per keyspace. In this tutorial you will install Cassandra 5.0 from the official Apache repository on three Ubuntu 24.04 servers, join them into one cluster, turn on password authentication, and create a keyspace that keeps three copies of every row.

Prerequisites

To follow this guide you need:

  • Three servers running Ubuntu 24.04 LTS, for example three CubePath VPS, each with at least 4 GB of RAM (8 GB or more for production) and 2 vCPUs.
  • A private network between them. This guide uses 10.0.0.11, 10.0.0.12 and 10.0.0.13; replace them with your own private IPs.
  • A non-root user with sudo privileges on each server.

Unless a step says otherwise, run the commands on all three nodes.

Step 1 - Installing Java

Cassandra 5.0 runs on Java 11 or Java 17. Install the headless OpenJDK 17 runtime from the Ubuntu repositories:

sudo apt update
sudo apt install openjdk-17-jre-headless

Check the version:

java -version
openjdk version "17.0.x" 2025-xx-xx
OpenJDK Runtime Environment (build 17.0.x+x-Ubuntu-...)
OpenJDK 64-Bit Server VM (build 17.0.x+x-Ubuntu-..., mixed mode, sharing)

Step 2 - Installing Cassandra from the official repository

Apache publishes Debian packages for each release series. Download the Apache Cassandra signing keys into /etc/apt/keyrings:

sudo install -m 0755 -d /etc/apt/keyrings
sudo curl -fsSL -o /etc/apt/keyrings/apache-cassandra.asc https://downloads.apache.org/cassandra/KEYS

Add the repository for the 5.0 series (50x), pinned to that key:

echo "deb [signed-by=/etc/apt/keyrings/apache-cassandra.asc] https://debian.cassandra.apache.org 50x main" | sudo tee /etc/apt/sources.list.d/cassandra.list

Install the package:

sudo apt update
sudo apt install cassandra

The package creates the cassandra system user, the data directory /var/lib/cassandra, the configuration directory /etc/cassandra, the file limits in /etc/security/limits.d/cassandra.conf, and starts the service right away. After about 30 seconds, when the node has finished booting, confirm the installed version:

nodetool version
ReleaseVersion: 5.0.x

Step 3 - Resetting the default single-node data

On first start every node boots as a standalone cluster called Test Cluster. Cassandra stores the cluster name in its system tables and refuses to start if you change it later, so stop the service and wipe the data it created before configuring the real cluster:

sudo systemctl stop cassandra
sudo rm -rf /var/lib/cassandra/data/* /var/lib/cassandra/commitlog/* /var/lib/cassandra/saved_caches/* /var/lib/cassandra/hints/*

Cassandra performs badly when the JVM is swapped out. If your servers have swap enabled, turn it off and comment out the swap line in /etc/fstab:

sudo swapoff -a

Step 4 - Configuring cassandra.yaml

Open the main configuration file:

sudo nano /etc/cassandra/cassandra.yaml

Find and change the following keys. The values below are for node 1 (10.0.0.11); the file is long, so use Ctrl+W in nano to search for each key:

cluster_name: 'ProductionCluster'

num_tokens: 16

seed_provider:
  - class_name: org.apache.cassandra.locator.SimpleSeedProvider
    parameters:
      - seeds: "10.0.0.11:7000,10.0.0.12:7000"

listen_address: 10.0.0.11
rpc_address: 10.0.0.11

endpoint_snitch: GossipingPropertyFileSnitch

What each setting does:

  • cluster_name must be identical on every node. Nodes with a different name will not join.
  • seeds lists the nodes that new nodes contact to learn the cluster topology. Use two or three nodes per datacenter, and the same list everywhere. Not every node should be a seed.
  • listen_address is the IP used for node-to-node traffic (port 7000). rpc_address is the IP where clients connect over CQL (port 9042). Use the node's own private IP for both.
  • GossipingPropertyFileSnitch reads the datacenter and rack name from a local file and is the recommended snitch for production, even with a single datacenter.

On node 2 and node 3, use the same values but set listen_address and rpc_address to 10.0.0.12 and 10.0.0.13.

Next, set the datacenter and rack name in cassandra-rackdc.properties:

sudo nano /etc/cassandra/cassandra-rackdc.properties
dc=dc1
rack=rack1

Use the same dc on all three nodes. You will reference this name when you create keyspaces.

Setting the heap size

By default Cassandra sizes its heap automatically. To set it explicitly, edit cassandra-env.sh:

sudo nano /etc/cassandra/cassandra-env.sh

Uncomment and adjust these two lines. They must be set together, and the heap should stay at or below half of the server's RAM:

MAX_HEAP_SIZE="4G"
HEAP_NEWSIZE="800M"

Step 5 - Opening the firewall between nodes

Cassandra uses port 7000 for internode communication and port 9042 for CQL clients. Port 7199 (JMX, used by nodetool) only listens on localhost by default and does not need to be opened. Allow the cluster ports only from your private network:

sudo ufw allow OpenSSH
sudo ufw allow from 10.0.0.0/24 to any port 7000 proto tcp
sudo ufw allow from 10.0.0.0/24 to any port 9042 proto tcp
sudo ufw enable

If your application servers live outside that subnet, add a 9042 rule for their IPs as well. Never expose 7000 or 9042 to the internet.

Step 6 - Starting the cluster

Start the seed nodes first, one at a time. On node 1:

sudo systemctl start cassandra

Wait until the node reports that it is ready for clients:

sudo journalctl -u cassandra -f

You can also watch /var/log/cassandra/system.log. When you see a line containing Starting listening for CQL clients on /10.0.0.11:9042, press Ctrl+C and start node 2, then node 3, waiting for each to finish joining before starting the next one.

Check the ring from any node:

nodetool status
Datacenter: dc1
===============
Status=Up/Down
|/ State=Normal/Leaving/Joining/Moving
--  Address    Load        Tokens  Owns (effective)  Host ID                               Rack
UN  10.0.0.11  110.2 KiB   16      64.7%             3f1c2a9e-...                          rack1
UN  10.0.0.12  98.5 KiB    16      68.1%             8b7d4e21-...                          rack1
UN  10.0.0.13  104.9 KiB   16      67.2%             c02e9f53-...                          rack1

UN means Up and Normal. A node showing UJ is still joining; wait and run the command again. Confirm that all nodes agree on the schema:

nodetool describecluster

The Schema versions section should list a single version with all three IPs.

Make sure the service starts at boot:

sudo systemctl enable cassandra

Step 7 - Enabling authentication

Out of the box Cassandra accepts any client without a password. Switch to password authentication and role-based authorization on every node:

sudo nano /etc/cassandra/cassandra.yaml
authenticator: PasswordAuthenticator
authorizer: CassandraAuthorizer

Credentials are stored in the system_auth keyspace. Before restarting, raise its replication so that losing one node does not lock you out. Connect to node 1 (still without authentication) and run:

cqlsh 10.0.0.11
ALTER KEYSPACE system_auth WITH replication = {'class': 'NetworkTopologyStrategy', 'dc1': 3};
EXIT;

Now restart the nodes one by one, waiting for each to show UN in nodetool status before moving on:

sudo systemctl restart cassandra

Log in with the default superuser cassandra (password cassandra) and create your own administrator. Replace your_strong_password with a real password:

cqlsh 10.0.0.11 -u cassandra -p cassandra
CREATE ROLE admin WITH SUPERUSER = true AND LOGIN = true AND PASSWORD = 'your_strong_password';
EXIT;

Log back in as admin and disable the default account:

cqlsh 10.0.0.11 -u admin
ALTER ROLE cassandra WITH SUPERUSER = false AND LOGIN = false;

Finally, run a repair of system_auth on each node so the new roles are copied to every replica:

nodetool repair system_auth

Step 8 - Creating a replicated keyspace and table

A keyspace defines how many copies of the data exist in each datacenter. With three nodes, a replication factor of 3 stores every row on every node and lets you lose one node while still reading and writing at QUORUM. In the cqlsh session as admin, create the keyspace:

CREATE KEYSPACE app WITH replication = {'class': 'NetworkTopologyStrategy', 'dc1': 3};

Cassandra tables are modeled around the queries you run. The table below stores events per user; user_id is the partition key (it decides which nodes hold the row) and event_time is a clustering column that keeps each user's events sorted, newest first:

USE app;

CREATE TABLE events_by_user (
  user_id    uuid,
  event_time timestamp,
  event_id   timeuuid,
  event_type text,
  payload    text,
  PRIMARY KEY ((user_id), event_time, event_id)
) WITH CLUSTERING ORDER BY (event_time DESC, event_id DESC);

Create an application role that can only read and write this keyspace:

CREATE ROLE app_user WITH LOGIN = true AND PASSWORD = 'another_strong_password';
GRANT SELECT ON KEYSPACE app TO app_user;
GRANT MODIFY ON KEYSPACE app TO app_user;

Insert and read a row at QUORUM consistency, which requires two of the three replicas to answer:

CONSISTENCY QUORUM;
INSERT INTO events_by_user (user_id, event_time, event_id, event_type, payload)
VALUES (5b6962dd-3f90-4c93-8f61-eabfa4a803e2, toTimestamp(now()), now(), 'login', '{"ip":"203.0.113.10"}');

SELECT event_time, event_type FROM events_by_user
WHERE user_id = 5b6962dd-3f90-4c93-8f61-eabfa4a803e2;
 event_time                      | event_type
---------------------------------+------------
 2026-09-25 10:14:03.512000+0000 |      login

(1 rows)

Check which nodes hold that partition:

nodetool getendpoints app events_by_user 5b6962dd-3f90-4c93-8f61-eabfa4a803e2

All three IPs should be listed, since the replication factor equals the number of nodes.

Step 9 - Routine maintenance

Repairs. Replicas can drift after node outages. Run a primary-range repair on each node, one node at a time, at least once within every gc_grace_seconds window (10 days by default):

nodetool repair -pr app

Snapshots. A snapshot creates hard links to the current SSTables, so it is instant and uses no extra space until data changes. Take one on every node at the same time:

nodetool snapshot -t before_upgrade app
nodetool listsnapshots

Snapshots live under /var/lib/cassandra/data/app/<table>/snapshots/. Copy them off the server for a real backup, and remove old ones with nodetool clearsnapshot -t before_upgrade. Save the schema too, since snapshots only contain data:

cqlsh 10.0.0.11 -u admin -e "DESCRIBE KEYSPACE app" > app_schema.cql

Removing a node. To take a healthy node out of the cluster permanently, run this on that node and wait for it to finish streaming its data:

nodetool decommission

Troubleshooting

Saved cluster name Test Cluster != configured name: the node started before you changed cluster_name. Stop Cassandra and clear the data directories as shown in Step 3.

Node stays DN or never appears in nodetool status: check that port 7000 is reachable between nodes (nc -zv 10.0.0.12 7000), that cluster_name matches everywhere, and that listen_address is the node's own IP, not localhost.

Connection refused from cqlsh: cqlsh connects to localhost by default, but rpc_address is the private IP. Pass the IP explicitly: cqlsh 10.0.0.11.

Out of memory or long GC pauses: check /var/log/cassandra/gc.log* and system.log, and make sure the heap is not larger than half the RAM and that swap is off.

Conclusion

You now have a three-node Cassandra 5.0 cluster with password authentication, a keyspace replicated to every node and an application role with limited permissions. From here you can connect your application with an official driver such as the DataStax Python or Java driver, schedule repairs with a tool like Cassandra Reaper, and add a second datacenter by giving new nodes a different dc name and extending the keyspace replication.