Apache Kafka is a distributed event streaming platform: producers append records to topics, Kafka stores them durably on disk split into partitions, and consumers read them at their own pace. Since Kafka 4.0, clusters are coordinated by Kafka's built-in KRaft protocol and ZooKeeper is no longer supported. In this tutorial you will install the current Kafka 4 release on Ubuntu 24.04 as a single-node broker and controller, run it under systemd with its own user, and create topics, produce and consume messages, and inspect consumer groups.
Prerequisites
To follow this guide you need:
- A server running Ubuntu 24.04 LTS with at least 2 CPU cores and 4 GB of RAM, for example a CubePath VPS. Kafka's default heap is 1 GB, and it relies heavily on the operating system page cache.
- Enough disk for the data you plan to keep. Kafka keeps records for 7 days by default.
- A non-root user with
sudoprivileges. - UFW enabled if other machines will connect to the broker.
Step 1 - Installing Java
Kafka 4 brokers require Java 17 or newer. Install the headless OpenJDK 21 runtime from the Ubuntu archive:
sudo apt update
sudo apt install -y openjdk-21-jre-headless
Verify the version:
java -version
openjdk version "21.0.x" ...
OpenJDK Runtime Environment (build 21.0.x+...-Ubuntu-...)
OpenJDK 64-Bit Server VM (build 21.0.x+...-Ubuntu-..., mixed mode, sharing)
Step 2 - Creating a Kafka user and downloading Kafka
Running Kafka as root is unnecessary and risky. Create a system user without a login shell:
sudo useradd --system --no-create-home --shell /usr/sbin/nologin kafka
Kafka is distributed as a binary tarball. Check the Kafka downloads page for the latest release; this guide uses 4.3.1 built for Scala 2.13. Set the version once so the next commands stay consistent:
KAFKA_VERSION=4.3.1
cd /tmp
curl -fLO "https://downloads.apache.org/kafka/${KAFKA_VERSION}/kafka_2.13-${KAFKA_VERSION}.tgz"
curl -fLO "https://downloads.apache.org/kafka/${KAFKA_VERSION}/kafka_2.13-${KAFKA_VERSION}.tgz.sha512"
Note
downloads.apache.orgonly keeps the newest releases. If the download returns a 404, the version has been superseded: use the current version from the downloads page, or fetch the old one fromhttps://archive.apache.org/dist/kafka/.
Compare the checksum of the file with the published one. The .sha512 file is formatted for humans, so print both and compare them:
sha512sum "kafka_2.13-${KAFKA_VERSION}.tgz"
cat "kafka_2.13-${KAFKA_VERSION}.tgz.sha512"
The hex digits must match (the published file splits them into groups separated by spaces). Then extract Kafka into /opt/kafka and create directories for data and logs:
sudo mkdir -p /opt/kafka
sudo tar -xzf "kafka_2.13-${KAFKA_VERSION}.tgz" -C /opt/kafka --strip-components=1
sudo mkdir -p /var/lib/kafka /var/log/kafka
sudo chown -R kafka:kafka /opt/kafka /var/lib/kafka /var/log/kafka
Check that the scripts are in place:
ls /opt/kafka/bin | grep -E "kafka-(server-start|storage|topics)"
kafka-server-start.sh
kafka-storage.sh
kafka-topics.sh
Step 3 - Configuring the broker in KRaft mode
In KRaft mode, each node has one or both roles: broker (stores data and serves clients) and controller (manages cluster metadata). For a single server, one process takes both roles. The default config/server.properties already describes such a combined node, so you only need to change where data lives and how clients reach the broker.
Open the file:
sudo nano /opt/kafka/config/server.properties
Find and change these settings, leaving the others as they are:
process.roles=broker,controller
node.id=1
controller.quorum.bootstrap.servers=localhost:9093
listeners=PLAINTEXT://:9092,CONTROLLER://:9093
advertised.listeners=PLAINTEXT://your_server_ip:9092,CONTROLLER://localhost:9093
# Keep data out of /tmp, which is cleared on reboot
log.dirs=/var/lib/kafka
num.partitions=3
log.retention.hours=168
What these settings do:
listeners: Kafka accepts client connections on port 9092 and controller traffic on 9093.advertised.listeners: the address Kafka returns to clients after the first connection. Clients then connect to this address, so it must be reachable from them. Replaceyour_server_ipwith the server's private IP if clients are on the same private network, or keeplocalhostif only local processes will use Kafka.log.dirs: where Kafka stores partition data (called logs). The default in/tmpwould lose all data on reboot.num.partitions: default partition count for topics created without an explicit value.
Before the first start, KRaft needs the storage directory formatted with a cluster ID. Generate the ID and format the directory as the kafka user. The --standalone flag initializes a single-controller quorum:
KAFKA_CLUSTER_ID="$(sudo -u kafka /opt/kafka/bin/kafka-storage.sh random-uuid)"
sudo -u kafka /opt/kafka/bin/kafka-storage.sh format --standalone -t "$KAFKA_CLUSTER_ID" -c /opt/kafka/config/server.properties
Formatting metadata directory /var/lib/kafka with metadata.version 4.x-IV...
The directory now contains a meta.properties file with the cluster and node IDs:
sudo cat /var/lib/kafka/meta.properties
cluster.id=...
directory.id=...
node.id=1
version=1
Step 4 - Running Kafka as a systemd service
Create a unit file so Kafka starts at boot and restarts after crashes:
sudo nano /etc/systemd/system/kafka.service
[Unit]
Description=Apache Kafka (KRaft)
After=network-online.target
Wants=network-online.target
[Service]
Type=simple
User=kafka
Group=kafka
Environment="KAFKA_HEAP_OPTS=-Xms1g -Xmx1g"
Environment="LOG_DIR=/var/log/kafka"
ExecStart=/opt/kafka/bin/kafka-server-start.sh /opt/kafka/config/server.properties
ExecStop=/opt/kafka/bin/kafka-server-stop.sh
Restart=on-failure
RestartSec=10
LimitNOFILE=100000
TimeoutStopSec=120
[Install]
WantedBy=multi-user.target
LOG_DIR moves Kafka's application logs (server.log, controller.log) out of /opt/kafka/logs. LimitNOFILE raises the open-file limit, because Kafka keeps several files open per partition segment. Size the heap to your server: 1 GB is fine for up to around 8 GB of RAM; leave the rest for the page cache.
Load the unit and start Kafka:
sudo systemctl daemon-reload
sudo systemctl enable --now kafka
sudo systemctl status kafka --no-pager
● kafka.service - Apache Kafka (KRaft)
Loaded: loaded (/etc/systemd/system/kafka.service; enabled; preset: enabled)
Active: active (running) since ...
Confirm that the broker finished starting:
sudo journalctl -u kafka --no-pager | grep "Kafka Server started"
... INFO [KafkaRaftServer nodeId=1] Kafka Server started (kafka.server.KafkaRaftServer)
Check the KRaft quorum. The single node should be the leader:
/opt/kafka/bin/kafka-metadata-quorum.sh --bootstrap-server localhost:9092 describe --status
ClusterId: ...
LeaderId: 1
LeaderEpoch: 1
HighWatermark: ...
MaxFollowerLag: 0
MaxFollowerLagTimeMs: 0
CurrentVoters: [{"id": 1, ...}]
CurrentObservers: []
Step 5 - Creating and inspecting topics
A topic is a named stream of records, split into partitions. Partitions are the unit of parallelism: within a consumer group, each partition is read by at most one consumer, and ordering is guaranteed only within a partition.
Create a topic named orders with three partitions. On a single broker the replication factor must be 1:
/opt/kafka/bin/kafka-topics.sh --bootstrap-server localhost:9092 \
--create --topic orders --partitions 3 --replication-factor 1
Created topic orders.
List topics and describe the new one:
/opt/kafka/bin/kafka-topics.sh --bootstrap-server localhost:9092 --list
/opt/kafka/bin/kafka-topics.sh --bootstrap-server localhost:9092 --describe --topic orders
orders
Topic: orders TopicId: ... PartitionCount: 3 ReplicationFactor: 1 Configs:
Topic: orders Partition: 0 Leader: 1 Replicas: 1 Isr: 1 ...
Topic: orders Partition: 1 Leader: 1 Replicas: 1 Isr: 1 ...
Topic: orders Partition: 2 Leader: 1 Replicas: 1 Isr: 1 ...
Step 6 - Producing and consuming messages
Kafka ships console clients that are handy for testing. Start a producer and type a few lines, pressing ENTER after each one. Each line becomes one record:
/opt/kafka/bin/kafka-console-producer.sh --bootstrap-server localhost:9092 --topic orders
>order 1
>order 2
>order 3
Press CTRL+C to exit. Now read the topic from the beginning with a consumer that belongs to the group billing:
/opt/kafka/bin/kafka-console-consumer.sh --bootstrap-server localhost:9092 \
--topic orders --group billing --from-beginning
order 1
order 2
order 3
Records without a key can land in any partition, and Kafka only guarantees order within a partition, so with several partitions the output order can differ. Press CTRL+C to stop the consumer. Kafka stores the group's committed position (offset) per partition, so you can check how far it has read:
/opt/kafka/bin/kafka-consumer-groups.sh --bootstrap-server localhost:9092 --describe --group billing
Consumer group 'billing' has no active members.
GROUP TOPIC PARTITION CURRENT-OFFSET LOG-END-OFFSET LAG CONSUMER-ID HOST CLIENT-ID
billing orders 0 3 3 0 - - -
billing orders 1 0 0 0 - - -
billing orders 2 0 0 0 - - -
Which partitions received the records depends on the producer's partitioner, so your offsets may be spread differently. LAG is the number of records the group has not consumed yet. Monitoring lag is the simplest way to see whether consumers keep up with producers.
Step 7 - Adjusting retention per topic
Kafka deletes old data by time or size, per topic. The broker default set by log.retention.hours is 7 days. To keep orders for 3 days and cap each partition at 5 GB, override the topic configuration:
/opt/kafka/bin/kafka-configs.sh --bootstrap-server localhost:9092 \
--alter --entity-type topics --entity-name orders \
--add-config retention.ms=259200000,retention.bytes=5368709120
Completed updating config for topic orders.
Verify the override:
/opt/kafka/bin/kafka-configs.sh --bootstrap-server localhost:9092 \
--describe --entity-type topics --entity-name orders
Dynamic configs for topic orders are:
retention.bytes=5368709120 sensitive=false synonyms={DYNAMIC_TOPIC_CONFIG:retention.bytes=5368709120}
retention.ms=259200000 sensitive=false synonyms={DYNAMIC_TOPIC_CONFIG:retention.ms=259200000}
Retention applies to whole segments, so Kafka deletes data a little after the limit is reached, not at the exact byte or millisecond.
Step 8 - Allowing remote clients
The PLAINTEXT listener has no encryption and no authentication. Never expose port 9092 to the internet. Allow it only from the private network your applications use (replace 10.0.0.0/24 with your subnet), and keep the controller port 9093 closed:
sudo ufw allow from 10.0.0.0/24 to any port 9092 proto tcp
sudo ufw status
From an application server in that network, clients connect with your_server_ip:9092 as the bootstrap server. If a remote client connects but then times out, advertised.listeners still points to localhost.
Troubleshooting
The service exits right after starting. Check sudo journalctl -u kafka -n 50 and /var/log/kafka/server.log. A message like No readable meta.properties files found means the storage directory was not formatted: repeat the kafka-storage.sh format command from Step 3.
UnsupportedClassVersionError at startup. An older Java is first in the path. Kafka 4 needs Java 17 or newer; check java -version and remove or switch away from older JDKs with sudo update-alternatives --config java.
AccessDeniedException on /var/lib/kafka or /var/log/kafka. The directories are owned by another user, usually because the format command ran with plain sudo. Fix ownership with sudo chown -R kafka:kafka /var/lib/kafka /var/log/kafka.
Clients connect but cannot produce or consume. Clients use the address in advertised.listeners after the first request. Make sure it resolves and is reachable from the client, then restart Kafka with sudo systemctl restart kafka.
Replication factor: 3 larger than available brokers: 1. A client or tool requested more replicas than you have brokers. Use --replication-factor 1 on a single node.
Conclusion
You installed Apache Kafka 4 on Ubuntu 24.04 in KRaft mode, ran it as a hardened systemd service, and worked with topics, console producers and consumers, consumer group lag and retention. For production, run at least three nodes so partitions can be replicated with replication.factor=3 and min.insync.replicas=2, enable TLS and SASL authentication on the client listener, and export broker metrics to your monitoring system.
