Apache Kafka is a distributed event streaming platform: producers append records to partitioned, replicated logs and consumers read them at their own pace. Since Kafka 4.0, ZooKeeper is gone and the cluster metadata is managed by Kafka itself through the KRaft protocol. In this tutorial you will build a three-node Kafka cluster on Ubuntu 24.04 where every node acts as both broker and controller, create a replicated topic, test a node failure, and learn the day-to-day management commands for consumer groups and partition reassignment.
Prerequisites
To follow this tutorial you need:
- Three servers running Ubuntu 24.04 LTS, for example three CubePath VPS, each with at least 2 vCPU, 4 GB of RAM and a dedicated disk or enough free space for the log data.
- A private network between the three servers. This guide uses
10.10.0.11,10.10.0.12and10.10.0.13; replace them with your own private IPs. - A non-root user with
sudoprivileges on every server.
Unless a step says otherwise, run the commands on all three nodes. The only per-node differences are the hostname and the node.id.
| Hostname | Private IP | node.id |
|---|---|---|
| kafka1 | 10.10.0.11 | 1 |
| kafka2 | 10.10.0.12 | 2 |
| kafka3 | 10.10.0.13 | 3 |
Step 1 - Preparing the nodes
Kafka brokers advertise their hostnames to clients and to each other, so every node must resolve every other node. Set the hostname on each server (use kafka2 and kafka3 on the other nodes):
sudo hostnamectl set-hostname kafka1
Add the three nodes to /etc/hosts on all servers:
sudo nano /etc/hosts
10.10.0.11 kafka1
10.10.0.12 kafka2
10.10.0.13 kafka3
Kafka 4.x brokers require Java 17 or newer. Install the headless OpenJDK 21 runtime from the Ubuntu repositories:
sudo apt update
sudo apt install -y openjdk-21-jre-headless
Check the Java version:
java -version
openjdk version "21.0.x" 2025-xx-xx
OpenJDK Runtime Environment (build 21.0.x+x-Ubuntu-...)
Confirm that the names resolve from each node:
getent hosts kafka1 kafka2 kafka3
Step 2 - Installing Kafka
Kafka is distributed as a tarball. Create a dedicated system user that will own the data and run the service:
sudo useradd --system --no-create-home --shell /usr/sbin/nologin kafka
Check the current release on the Apache Kafka downloads page and set it in a variable. The example uses 4.1.0 with Scala 2.13; the Apache archive keeps every release at a stable URL:
KAFKA_VERSION=4.1.0
cd /tmp
curl -fLO "https://archive.apache.org/dist/kafka/${KAFKA_VERSION}/kafka_2.13-${KAFKA_VERSION}.tgz"
curl -fLO "https://archive.apache.org/dist/kafka/${KAFKA_VERSION}/kafka_2.13-${KAFKA_VERSION}.tgz.sha512"
Compare the checksum of the tarball with the published one before extracting it:
sha512sum "kafka_2.13-${KAFKA_VERSION}.tgz"
cat "kafka_2.13-${KAFKA_VERSION}.tgz.sha512"
Both commands must show the same hash (the .sha512 file prints it in groups of 8 characters). Extract Kafka to /opt and create a version-independent symlink, which makes future upgrades a matter of switching the link:
sudo tar -xzf "kafka_2.13-${KAFKA_VERSION}.tgz" -C /opt
sudo ln -sfn "/opt/kafka_2.13-${KAFKA_VERSION}" /opt/kafka
Create the configuration, data and log directories:
sudo mkdir -p /etc/kafka /var/lib/kafka/data /var/log/kafka
sudo chown -R kafka:kafka /var/lib/kafka /var/log/kafka
Add the Kafka tools to your PATH for the current shell so the management commands later are shorter:
echo 'export PATH="$PATH:/opt/kafka/bin"' >> ~/.bashrc
source ~/.bashrc
kafka-topics.sh --version
4.1.0
Step 3 - Configuring the brokers
Each node runs in combined mode (process.roles=broker,controller): it stores data and also takes part in the Raft quorum that holds the cluster metadata. Three controllers tolerate the loss of one node. The file below is the configuration for kafka1:
sudo nano /etc/kafka/server.properties
# Identity and roles
process.roles=broker,controller
node.id=1
# KRaft quorum: node.id@host:controller_port for every controller
controller.quorum.voters=1@kafka1:9093,2@kafka2:9093,3@kafka3:9093
controller.listener.names=CONTROLLER
# Listeners: 9092 for clients and replication, 9093 for the controller quorum
listeners=PLAINTEXT://:9092,CONTROLLER://:9093
advertised.listeners=PLAINTEXT://kafka1:9092
inter.broker.listener.name=PLAINTEXT
listener.security.protocol.map=PLAINTEXT:PLAINTEXT,CONTROLLER:PLAINTEXT
# Storage
log.dirs=/var/lib/kafka/data
# Durability defaults for new topics and internal topics
num.partitions=6
default.replication.factor=3
min.insync.replicas=2
offsets.topic.replication.factor=3
transaction.state.log.replication.factor=3
transaction.state.log.min.isr=2
unclean.leader.election.enable=false
auto.create.topics.enable=false
# Retention: 7 days
log.retention.hours=168
On kafka2 and kafka3 use the same file and change only two lines:
node.id=2andadvertised.listeners=PLAINTEXT://kafka2:9092on kafka2.node.id=3andadvertised.listeners=PLAINTEXT://kafka3:9092on kafka3.
What the durability settings do:
default.replication.factor=3: each partition is stored on all three nodes.min.insync.replicas=2: a write withacks=all(the producer default since Kafka 3.0) succeeds only when at least two replicas have it. You can lose one node without losing acknowledged data or availability.unclean.leader.election.enable=false: an out-of-date replica is never promoted to leader, which would silently drop records.auto.create.topics.enable=false: topics are created on purpose, with the right partition count, instead of by a typo in a producer.
Make the file readable by the service user:
sudo chown root:kafka /etc/kafka/server.properties
sudo chmod 640 /etc/kafka/server.properties
Step 4 - Formatting the storage
In KRaft mode every node's data directory must be formatted with the same cluster ID before the first start. Generate the ID once, on kafka1:
kafka-storage.sh random-uuid
q1Sh-9_ISia_zwGINzRvyQ
Run the format command on each node, using that same ID:
sudo -u kafka /opt/kafka/bin/kafka-storage.sh format \
--cluster-id q1Sh-9_ISia_zwGINzRvyQ \
--config /etc/kafka/server.properties
Formatting metadata directory /var/lib/kafka/data with metadata.version 4.1-IV1.
If a node is formatted with a different cluster ID it will refuse to join; see the troubleshooting section below.
Step 5 - Running Kafka with systemd
Create a systemd unit so Kafka starts at boot and restarts if it crashes:
sudo nano /etc/systemd/system/kafka.service
[Unit]
Description=Apache Kafka (KRaft)
Wants=network-online.target
After=network-online.target
[Service]
Type=simple
User=kafka
Group=kafka
Environment="KAFKA_HEAP_OPTS=-Xms1g -Xmx1g"
Environment="LOG_DIR=/var/log/kafka"
ExecStart=/opt/kafka/bin/kafka-server-start.sh /etc/kafka/server.properties
ExecStop=/opt/kafka/bin/kafka-server-stop.sh
Restart=on-failure
RestartSec=10
LimitNOFILE=100000
TimeoutStopSec=120
[Install]
WantedBy=multi-user.target
A 1 GB heap is plenty for most clusters: Kafka relies on the operating system page cache, not the JVM heap, to serve data. Leave the rest of the RAM to the OS. LimitNOFILE matters because every log segment and client connection uses a file descriptor.
Allow the Kafka ports only from the private network. Replace 10.10.0.0/24 with your subnet:
sudo ufw allow from 10.10.0.0/24 to any port 9092 proto tcp
sudo ufw allow from 10.10.0.0/24 to any port 9093 proto tcp
Start the service on all three nodes, roughly at the same time, since the controllers need a majority to elect a leader:
sudo systemctl daemon-reload
sudo systemctl enable --now kafka
Check that it is running:
sudo systemctl status kafka --no-pager
● kafka.service - Apache Kafka (KRaft)
Loaded: loaded (/etc/systemd/system/kafka.service; enabled; preset: enabled)
Active: active (running) since ...
The broker log is written to /var/log/kafka/server.log. Look for the line that confirms the broker started:
grep "started" /var/log/kafka/server.log | tail -n 3
Step 6 - Verifying the cluster
The kafka-metadata-quorum.sh tool shows the state of the KRaft quorum. Run it from any node:
kafka-metadata-quorum.sh --bootstrap-server kafka1:9092 describe --status
ClusterId: q1Sh-9_ISia_zwGINzRvyQ
LeaderId: 2
LeaderEpoch: 3
HighWatermark: 812
MaxFollowerLag: 0
MaxFollowerLagTimeMs: 0
CurrentVoters: [{"id": 1, ...}, {"id": 2, ...}, {"id": 3, ...}]
CurrentObservers: []
Three current voters and a leader mean the quorum is healthy. To see per-node replication lag of the metadata log:
kafka-metadata-quorum.sh --bootstrap-server kafka1:9092 describe --replication
The Lag column should be 0 or close to it for all three nodes.
Step 7 - Creating a replicated topic
Create a topic with 6 partitions and 3 replicas. The partition count caps how many consumers in one group can read in parallel, so size it for your peak consumer count; it can be increased later but never decreased:
kafka-topics.sh --bootstrap-server kafka1:9092 \
--create --topic orders \
--partitions 6 --replication-factor 3
Created topic orders.
Describe it to see where each partition lives:
kafka-topics.sh --bootstrap-server kafka1:9092 --describe --topic orders
Topic: orders TopicId: ... PartitionCount: 6 ReplicationFactor: 3 Configs: min.insync.replicas=2,...
Topic: orders Partition: 0 Leader: 1 Replicas: 1,2,3 Isr: 1,2,3
Topic: orders Partition: 1 Leader: 2 Replicas: 2,3,1 Isr: 2,3,1
Topic: orders Partition: 2 Leader: 3 Replicas: 3,1,2 Isr: 3,1,2
...
Leader: the node that serves reads and writes for that partition.Replicas: all nodes holding a copy. The first one is the preferred leader.Isr: in-sync replicas. If this list is shorter thanReplicas, a node is down or lagging.
Produce a few records from one node. Type one message per line and press CTRL+D to finish:
kafka-console-producer.sh --bootstrap-server kafka1:9092 --topic orders
>order 1001 created
>order 1002 created
Read them from a different node to prove replication and client routing work:
kafka-console-consumer.sh --bootstrap-server kafka3:9092 \
--topic orders --from-beginning --max-messages 2
order 1001 created
order 1002 created
Processed a total of 2 messages
Step 8 - Testing a node failure
Stop Kafka on kafka3:
sudo systemctl stop kafka
From kafka1, list the partitions that are now under-replicated:
kafka-topics.sh --bootstrap-server kafka1:9092 --describe --under-replicated-partitions
Topic: orders Partition: 0 Leader: 1 Replicas: 1,2,3 Isr: 1,2
Topic: orders Partition: 2 Leader: 1 Replicas: 3,1,2 Isr: 1,2
...
Node 3 dropped out of every ISR and leadership moved to the surviving nodes. With two replicas still in sync, min.insync.replicas=2 is satisfied, so producing to orders keeps working. Start kafka3 again:
sudo systemctl start kafka
After a few seconds the same command returns no output, which means every partition is fully replicated again. Leadership stays where it moved during the outage; Kafka rebalances it back to the preferred leaders automatically (auto.leader.rebalance.enable is on by default, checked every 5 minutes). To do it immediately:
kafka-leader-election.sh --bootstrap-server kafka1:9092 \
--election-type PREFERRED --all-topic-partitions
ImportantNever stop two nodes of a three-node cluster at the same time. You lose the controller majority and every partition falls below
min.insync.replicas, so producers usingacks=allwill fail. Restart nodes one at a time and wait for the under-replicated list to be empty before moving on.
Step 9 - Managing consumer groups
Consumers that share a group.id split the partitions of a topic between them, and Kafka stores each group's committed offsets. Start a consumer in a group:
kafka-console-consumer.sh --bootstrap-server kafka1:9092 \
--topic orders --group billing
In another terminal, list groups and inspect the lag of billing:
kafka-consumer-groups.sh --bootstrap-server kafka1:9092 --list
kafka-consumer-groups.sh --bootstrap-server kafka1:9092 --describe --group billing
GROUP TOPIC PARTITION CURRENT-OFFSET LOG-END-OFFSET LAG CONSUMER-ID HOST CLIENT-ID
billing orders 0 14 14 0 console-... /10.10.0.11 console-consumer
billing orders 1 9 12 3 console-... /10.10.0.11 console-consumer
...
LAG is the number of records the group has not processed yet. Lag that keeps growing means your consumers cannot keep up: add consumers (up to the partition count) or make processing faster.
To reprocess data, reset the group offsets. The group must have no active members, so stop the consumer first. Always preview with --dry-run:
kafka-consumer-groups.sh --bootstrap-server kafka1:9092 \
--group billing --topic orders \
--reset-offsets --to-datetime 2026-09-01T00:00:00.000 --dry-run
If the new offsets look right, run the same command with --execute instead of --dry-run. Other useful targets are --to-earliest, --to-latest and --shift-by -100.
Step 10 - Adding a broker and rebalancing partitions
When you add a node, Kafka does not move existing partitions to it; only new topics use it. You move data explicitly with kafka-reassign-partitions.sh.
Prepare the fourth server as in steps 1 and 2, then give it a broker-only configuration. It does not join the controller quorum, which stays at three voters:
process.roles=broker
node.id=4
controller.quorum.voters=1@kafka1:9093,2@kafka2:9093,3@kafka3:9093
controller.listener.names=CONTROLLER
listeners=PLAINTEXT://:9092
advertised.listeners=PLAINTEXT://kafka4:9092
inter.broker.listener.name=PLAINTEXT
listener.security.protocol.map=PLAINTEXT:PLAINTEXT,CONTROLLER:PLAINTEXT
log.dirs=/var/lib/kafka/data
default.replication.factor=3
min.insync.replicas=2
unclean.leader.election.enable=false
auto.create.topics.enable=false
Add kafka4 to /etc/hosts on every node, format its storage with the same cluster ID (step 4) and start it (step 5). Then, on kafka1, list the topics to spread over the four brokers:
nano ~/topics.json
{
"version": 1,
"topics": [
{ "topic": "orders" }
]
}
Ask Kafka for a proposed plan:
kafka-reassign-partitions.sh --bootstrap-server kafka1:9092 \
--topics-to-move-json-file ~/topics.json \
--broker-list "1,2,3,4" --generate
The output contains two JSON documents: Current partition replica assignment (save it as a rollback plan) and Proposed partition reassignment configuration. Copy the proposed JSON into ~/reassign.json, then execute it with a replication throttle in bytes per second so the data copy does not saturate the network:
kafka-reassign-partitions.sh --bootstrap-server kafka1:9092 \
--reassignment-json-file ~/reassign.json \
--execute --throttle 50000000
Check progress until every partition reports completion. --verify also removes the throttle once the reassignment is done:
kafka-reassign-partitions.sh --bootstrap-server kafka1:9092 \
--reassignment-json-file ~/reassign.json --verify
Status of partition reassignment:
Reassignment of partition orders-0 is completed.
...
Clearing broker-level throttles on brokers 1,2,3,4
Troubleshooting
A node does not join and logs Invalid cluster.id. The data directory was formatted with a different ID. On that node only, stop Kafka, empty /var/lib/kafka/data and run the format command from step 4 again with the correct ID. Never do this on a node whose data you need.
Clients connect but then fail with UnknownHostException: kafka2. The client received the advertised hostname and cannot resolve it. Add the brokers to the client's DNS or /etc/hosts, or advertise names that the clients can resolve.
Producers fail with NotEnoughReplicasException. Fewer than min.insync.replicas replicas are in sync, usually because two nodes are down or lagging. Run the --under-replicated-partitions check from step 8 and look at /var/log/kafka/server.log on the missing nodes.
The quorum has no leader (LeaderId: -1). Fewer than two controllers are reachable. Check that port 9093 is open between nodes and that controller.quorum.voters is identical on every node.
Conclusion
You now have a three-node Kafka cluster running in KRaft mode, with topics replicated to every node, protection against data loss when a single node fails, and the commands to manage consumer lag and move partitions onto new brokers. As next steps, enable TLS and SASL authentication on the client listener before exposing it beyond your private network, export broker metrics to your monitoring system, and set per-topic retention with kafka-configs.sh to match how long each stream really needs to be kept.
