ScyllaDB is a wide-column NoSQL database that speaks the Cassandra Query Language (CQL) and the Cassandra wire protocol, but is written in C++ with a shard-per-core design: each CPU core owns its own slice of data and memory, so there are no garbage collection pauses and throughput scales with the number of cores. In this tutorial you will install ScyllaDB on three Ubuntu 24.04 servers, join them into a cluster, turn on password authentication, create a replicated keyspace and query it from cqlsh and from a Python application.
Prerequisites
To follow this tutorial you need:
- Three servers running Ubuntu 24.04 LTS (x86_64 or arm64), for example three CubePath VPS in the same location. A single server also works if you only want to test; skip Step 4 in that case.
- A non-root user with
sudoprivileges on each server. - At least 4 CPU cores and 8 GB of RAM per node for anything beyond testing. ScyllaDB uses all the memory and cores it is given.
- SSD or NVMe storage. For production, ScyllaDB expects the data directory on an XFS file system, ideally on a dedicated disk.
- A private network between the nodes. This guide uses
10.0.0.11,10.0.0.12and10.0.0.13; replace them with your nodes' private IPs.
Run Steps 1 to 3 on every node.
Step 1 - Installing ScyllaDB from the official repository
ScyllaDB publishes its own APT repository per release. Create the keyring directory and import the repository signing key:
sudo mkdir -p /etc/apt/keyrings
sudo gpg --homedir /tmp --no-default-keyring --keyring /etc/apt/keyrings/scylladb.gpg \
--keyserver hkp://keyserver.ubuntu.com:80 --recv-keys a43e06657bac99e3
Download the repository definition. Each release has its own list file; open the installation page at https://docs.scylladb.com to confirm the current release number and replace 2025.1 if a newer one is available. All nodes in a cluster must run the same release.
sudo wget -O /etc/apt/sources.list.d/scylla.list \
https://downloads.scylladb.com/deb/debian/scylla-2025.1.list
Install the package, which also provides cqlsh and nodetool:
sudo apt update
sudo apt install -y scylla
Confirm the installed version:
scylla --version
2025.1.x-0.2025xxxx.xxxxxxxxxxxx
Step 2 - Tuning the system with scylla_setup
ScyllaDB gets its performance from direct control of the disks, network queues and CPUs, so it ships a setup script that tunes the kernel, checks the file system, configures NTP and benchmarks the disks (the result goes to /etc/scylla.d/io.conf). Run it and answer the questions:
sudo scylla_setup
Useful answers on a typical VPS:
- Answer yes to the kernel, NTP and I/O setup questions.
- Answer no to RAID setup unless you attached extra empty disks that ScyllaDB may format as an XFS RAID 0 array for
/var/lib/scylla. - Answer yes to enabling the
scylla-serverservice at boot.
The I/O benchmark takes a few minutes.
NoteIf you are only testing on a small VPS with an ext4 root disk,
scylla_setupwill refuse to continue. Put ScyllaDB in developer mode instead, which skips the hardware checks:sudo scylla_dev_mode_setup --developer-mode 1. Never run a production node in developer mode.
Do not start the service yet. The cluster name and addresses must be set before the first start, because ScyllaDB stores them in its system tables.
Step 3 - Configuring each node
The main configuration file is /etc/scylla/scylla.yaml. Open it on the first node:
sudo nano /etc/scylla/scylla.yaml
Find and set the following keys. Most are already present, some commented out:
cluster_name: 'cubepath-cluster'
seed_provider:
- class_name: org.apache.cassandra.locator.SimpleSeedProvider
parameters:
- seeds: "10.0.0.11,10.0.0.12"
listen_address: 10.0.0.11
rpc_address: 10.0.0.11
endpoint_snitch: GossipingPropertyFileSnitch
authenticator: PasswordAuthenticator
authorizer: CassandraAuthorizer
What each setting does:
cluster_name: must be identical on all nodes. Nodes with a different name refuse to join.seeds: nodes that a new node contacts to discover the cluster. Use the same list everywhere, two seeds is enough for three nodes.listen_address: the private IP used for traffic between nodes.rpc_address: the IP where clients connect with CQL on port9042. Keep it on the private network.endpoint_snitch:GossipingPropertyFileSnitchreads the data center and rack name from a local file, which lets you add a second data center later without rebuilding.authenticatorandauthorizer: require a user and password and enableGRANTpermissions.
On the second and third node, use the same file but change listen_address and rpc_address to 10.0.0.12 and 10.0.0.13.
Next, set the data center and rack name on every node:
sudo nano /etc/scylla/cassandra-rackdc.properties
dc=dc1
rack=rack1
Finally, allow the cluster ports only from the private network. Port 7000 carries internode traffic, 7001 is internode traffic over TLS, 9042 is CQL and 19042 is the shard-aware CQL port that modern drivers use. Replace 10.0.0.0/24 with your private subnet:
sudo ufw allow from 10.0.0.0/24 to any port 7000,7001,9042,19042 proto tcp
sudo ufw status
Status: active
To Action From
-- ------ ----
OpenSSH ALLOW Anywhere
7000,7001,9042,19042/tcp ALLOW 10.0.0.0/24
Step 4 - Starting the cluster
Start the nodes one at a time, beginning with a seed node. On 10.0.0.11:
sudo systemctl start scylla-server
Follow the log until the node reports that it is serving CQL:
sudo journalctl -u scylla-server -f
... init - Scylla version ... initialization completed.
Press Ctrl+C to stop following the log, then check the node state:
nodetool status
The node should appear as UN (Up, Normal). Only then start scylla-server on 10.0.0.12, wait for it to become UN, and finally start it on 10.0.0.13. Starting several nodes at once makes them compete to join and one will fail.
When the third node has joined, nodetool status on any node lists all three:
Datacenter: dc1
===============
Status=Up/Down
|/ State=Normal/Leaving/Joining/Moving
-- Address Load Tokens Owns Host ID Rack
UN 10.0.0.11 1.02 MB 256 ? 0b3d8c52-5a1e-4f4e-9d4a-1c2f0e8a7b11 rack1
UN 10.0.0.12 998.4 KB 256 ? 6f7a1e0c-2d3b-4c5a-8e9f-0a1b2c3d4e5f rack1
UN 10.0.0.13 1.01 MB 256 ? 9c8b7a6d-5e4f-4a3b-2c1d-0e9f8a7b6c5d rack1
Both services are enabled at boot, so the cluster comes back after a reboot.
Step 5 - Securing the default superuser
With PasswordAuthenticator enabled, ScyllaDB creates a default superuser named cassandra with the password cassandra. Replace it with your own superuser straight away. Connect from any node:
cqlsh 10.0.0.11 -u cassandra -p cassandra
Create a new superuser role. Replace your_strong_password with a long random password:
CREATE ROLE dbadmin WITH PASSWORD = 'your_strong_password' AND SUPERUSER = true AND LOGIN = true;
Exit with exit, log in as the new role, and remove the default one:
cqlsh 10.0.0.11 -u dbadmin
DROP ROLE cassandra;
LIST ROLES;
role | super | login | options
---------+-------+-------+---------
dbadmin | True | True | {}
Authentication data is replicated across the cluster automatically, so the change applies to every node.
Step 6 - Creating a keyspace, a table and an application role
A keyspace defines how many copies of each row the cluster keeps. With three nodes, a replication factor of 3 stores every row on every node and survives the loss of one node at QUORUM consistency. Still in cqlsh as dbadmin:
CREATE KEYSPACE IF NOT EXISTS myapp
WITH replication = {'class': 'NetworkTopologyStrategy', 'dc1': 3};
Create a time-series table. The partition key (host) decides which nodes store the row; the clustering key (ts) sorts rows inside the partition, newest first:
CREATE TABLE IF NOT EXISTS myapp.metrics (
host text,
ts timestamp,
cpu_usage float,
memory_mb int,
PRIMARY KEY ((host), ts)
) WITH CLUSTERING ORDER BY (ts DESC);
Insert a row and read it back:
INSERT INTO myapp.metrics (host, ts, cpu_usage, memory_mb)
VALUES ('web01', toTimestamp(now()), 45.2, 2048);
SELECT * FROM myapp.metrics WHERE host = 'web01' LIMIT 10;
host | ts | cpu_usage | memory_mb
-------+---------------------------------+-----------+-----------
web01 | 2026-09-25 10:14:03.512000+0000 | 45.2 | 2048
Applications should not use a superuser. Create a role limited to this keyspace, replacing your_app_password:
CREATE ROLE app WITH PASSWORD = 'your_app_password' AND LOGIN = true;
GRANT SELECT ON KEYSPACE myapp TO app;
GRANT MODIFY ON KEYSPACE myapp TO app;
SELECT allows reads and MODIFY allows INSERT, UPDATE and DELETE.
Step 7 - Connecting from a Python application
ScyllaDB maintains a fork of the Cassandra Python driver, scylla-driver, that is shard-aware: it opens a connection to each CPU shard and sends every request straight to the shard that owns the data. It is imported with the same cassandra module name. Install it in a virtual environment on a client machine inside the private network:
sudo apt install -y python3-venv
python3 -m venv ~/scylla-client
~/scylla-client/bin/pip install scylla-driver
Create a test script:
nano ~/scylla_test.py
from cassandra.auth import PlainTextAuthProvider
from cassandra.cluster import Cluster
from cassandra.policies import DCAwareRoundRobinPolicy, TokenAwarePolicy
auth = PlainTextAuthProvider(username="app", password="your_app_password")
cluster = Cluster(
contact_points=["10.0.0.11", "10.0.0.12", "10.0.0.13"],
auth_provider=auth,
load_balancing_policy=TokenAwarePolicy(DCAwareRoundRobinPolicy(local_dc="dc1")),
)
session = cluster.connect("myapp")
for row in session.execute("SELECT host, ts, cpu_usage FROM metrics WHERE host = %s LIMIT 5", ["web01"]):
print(row.host, row.ts, row.cpu_usage)
cluster.shutdown()
Run it:
~/scylla-client/bin/python ~/scylla_test.py
web01 2026-09-25 10:14:03.512000 45.2
Step 8 - Checking cluster health
nodetool talks to the local node's REST API (bound to 127.0.0.1:10000) and is the first tool to reach for:
nodetool status
nodetool info
nodetool tablestats myapp.metrics
nodetool compactionstats
Each node also exposes Prometheus metrics on port 9180:
curl -s http://localhost:9180/metrics | grep -m 5 '^scylla_reactor_utilization'
scylla_reactor_utilization{shard="0"} 1.23
scylla_reactor_utilization{shard="1"} 0.87
There is one line per shard, which is a quick way to see how many shards (cores) ScyllaDB is using. For dashboards and alerts, ScyllaDB publishes the ScyllaDB Monitoring Stack (Prometheus and Grafana with ready-made dashboards); point it at port 9180 on each node and allow that port from the monitoring server with UFW.
Troubleshooting
The service fails at start with an I/O or file system error. ScyllaDB refuses to run without the I/O configuration from scylla_setup or on unsupported file systems in production mode. Read the reason with sudo journalctl -u scylla-server -n 50, then rerun sudo scylla_setup or, on a test box only, enable developer mode.
A node never leaves the joining state or cannot see the others. Check that cluster_name and seeds are identical on all nodes, that listen_address is the private IP and not 127.0.0.1, and that port 7000 is reachable:
nc -zv 10.0.0.11 7000
If you started a node with the wrong cluster_name, stop it and clear its data before starting again (this deletes everything on that node):
sudo systemctl stop scylla-server
sudo rm -rf /var/lib/scylla/data/* /var/lib/scylla/commitlog/* /var/lib/scylla/hints/* /var/lib/scylla/view_hints/*
cqlsh reports Connection refused. Connect to the rpc_address of the node, not localhost, because CQL only listens on that IP.
Conclusion
You now have a three-node ScyllaDB cluster on Ubuntu 24.04 with password authentication, a keyspace replicated to every node and a least-privilege role used from a shard-aware Python client. As next steps, deploy the ScyllaDB Monitoring Stack, enable client-to-node and internode TLS in scylla.yaml, and schedule regular snapshots with nodetool snapshot copied off the servers.
