etcd is a strongly consistent, distributed key-value store built on the Raft consensus algorithm. Kubernetes stores its entire state in it, and many other systems use it for configuration, leader election and locks. In this tutorial you will deploy a three-node etcd 3.6 cluster on Ubuntu 24.04 with mutual TLS for both peer and client traffic, run it under systemd, work with keys through etcdctl, schedule daily snapshots and restore the cluster from one.
Prerequisites
To follow this guide you need:
- Three servers running Ubuntu 24.04 LTS, for example CubePath VPS instances on the same private network, each with 2 vCPUs, 2 GB of RAM or more and SSD or NVMe storage. etcd is very sensitive to disk write latency.
- A non-root user with
sudoprivileges on each server, and SSH access between your workstation and the servers. - Static private IP addresses. This guide uses:
| Hostname | Private IP |
|---|---|
| etcd-01 | 10.0.0.31 |
| etcd-02 | 10.0.0.32 |
| etcd-03 | 10.0.0.33 |
A three-member cluster keeps working if one member fails. Always use an odd number of members (3 or 5).
Step 1 - Installing the etcd binaries
Ubuntu's etcd-server package is an older 3.4 release, so install the official binaries from the etcd project instead. Run this step on all three servers. Check the releases page for the latest 3.6 patch release and set it in the variable:
ETCD_VER=v3.6.4
cd /tmp
curl -fsSLO "https://github.com/etcd-io/etcd/releases/download/${ETCD_VER}/etcd-${ETCD_VER}-linux-amd64.tar.gz"
curl -fsSLO "https://github.com/etcd-io/etcd/releases/download/${ETCD_VER}/SHA256SUMS"
sha256sum --check --ignore-missing SHA256SUMS
etcd-v3.6.4-linux-amd64.tar.gz: OK
Extract the archive and install the three binaries: etcd (the server), etcdctl (the client) and etcdutl (offline tool for snapshots and data directories):
tar xzf "etcd-${ETCD_VER}-linux-amd64.tar.gz"
sudo install -m 0755 "etcd-${ETCD_VER}-linux-amd64/etcd" "etcd-${ETCD_VER}-linux-amd64/etcdctl" "etcd-${ETCD_VER}-linux-amd64/etcdutl" /usr/local/bin/
etcd --version
etcd Version: 3.6.4
Git SHA: ...
Go Version: go1.23.x
Go OS/Arch: linux/amd64
On arm64 servers, download the linux-arm64 archive instead. Create a system user and the directories for configuration and data. etcd requires its data directory to be accessible only by its owner:
sudo useradd --system --home-dir /var/lib/etcd --shell /usr/sbin/nologin etcd
sudo install -d -m 0755 /etc/etcd
sudo install -d -m 0700 -o etcd -g etcd /var/lib/etcd
Step 2 - Creating the TLS certificates
etcd has no authentication by default, so anyone who can reach its ports can read and write every key. Mutual TLS fixes that: every member and every client must present a certificate signed by your own CA. Run this step once, on your workstation or on etcd-01, with openssl (installed by default on Ubuntu).
Create a working directory and a CA valid for ten years:
mkdir -p ~/etcd-pki && cd ~/etcd-pki
openssl req -x509 -new -nodes -newkey rsa:4096 -sha256 -days 3650 -keyout ca.key -out ca.crt -subj "/CN=etcd-ca"
Each member needs a certificate that includes its hostname and IP address in the Subject Alternative Name, and that is valid both as a server certificate (for clients and peers connecting to it) and as a client certificate (when it connects to other peers). Issue one per member with a loop:
for member in etcd-01:10.0.0.31 etcd-02:10.0.0.32 etcd-03:10.0.0.33; do
name="${member%%:*}"
ip="${member##*:}"
openssl req -new -nodes -newkey rsa:2048 -keyout "${name}.key" -out "${name}.csr" -subj "/CN=${name}"
printf 'subjectAltName=DNS:%s,DNS:localhost,IP:%s,IP:127.0.0.1\nextendedKeyUsage=serverAuth,clientAuth\nkeyUsage=digitalSignature,keyEncipherment\n' "$name" "$ip" > "${name}.ext"
openssl x509 -req -in "${name}.csr" -CA ca.crt -CAkey ca.key -CAcreateserial -sha256 -days 825 -extfile "${name}.ext" -out "${name}.crt"
done
Also create a client certificate for yourself, used by etcdctl:
openssl req -new -nodes -newkey rsa:2048 -keyout admin.key -out admin.csr -subj "/CN=admin"
printf 'extendedKeyUsage=clientAuth\nkeyUsage=digitalSignature,keyEncipherment\n' > admin.ext
openssl x509 -req -in admin.csr -CA ca.crt -CAkey ca.key -CAcreateserial -sha256 -days 825 -extfile admin.ext -out admin.crt
Check that a member certificate contains the right names:
openssl x509 -in etcd-01.crt -noout -ext subjectAltName
X509v3 Subject Alternative Name:
DNS:etcd-01, DNS:localhost, IP Address:10.0.0.31, IP Address:127.0.0.1
Copy to each server its own certificate and key, the CA certificate (not the CA key) and the admin client files. For etcd-01:
scp ca.crt etcd-01.crt etcd-01.key admin.crt admin.key [email protected]:~
Repeat for etcd-02 and etcd-03 with their own files. Then, on each server, install the files in place (shown for etcd-01):
sudo install -m 0644 ~/ca.crt /etc/etcd/ca.crt
sudo install -m 0644 ~/etcd-01.crt /etc/etcd/server.crt
sudo install -m 0600 -o etcd -g etcd ~/etcd-01.key /etc/etcd/server.key
mkdir -p ~/.etcd && mv ~/admin.crt ~/admin.key ~/.etcd/ && cp ~/ca.crt ~/.etcd/ && chmod 600 ~/.etcd/admin.key
rm ~/etcd-01.crt ~/etcd-01.key ~/ca.crt
Keep ca.key offline in a safe place: whoever has it can issue certificates that the cluster trusts.
Step 3 - Configuring the members
etcd reads every command-line flag from an environment variable with the ETCD_ prefix, which keeps the configuration in a simple file. Create it on each server:
sudo nano /etc/etcd/etcd.env
This is the file for etcd-01. On the other members, change ETCD_NAME and the four URLs that contain the member's own IP:
ETCD_NAME=etcd-01
ETCD_DATA_DIR=/var/lib/etcd
ETCD_LISTEN_PEER_URLS=https://10.0.0.31:2380
ETCD_INITIAL_ADVERTISE_PEER_URLS=https://10.0.0.31:2380
ETCD_LISTEN_CLIENT_URLS=https://10.0.0.31:2379,https://127.0.0.1:2379
ETCD_ADVERTISE_CLIENT_URLS=https://10.0.0.31:2379
ETCD_LISTEN_METRICS_URLS=http://127.0.0.1:2381
ETCD_INITIAL_CLUSTER=etcd-01=https://10.0.0.31:2380,etcd-02=https://10.0.0.32:2380,etcd-03=https://10.0.0.33:2380
ETCD_INITIAL_CLUSTER_TOKEN=etcd-cluster-1
ETCD_INITIAL_CLUSTER_STATE=new
ETCD_CERT_FILE=/etc/etcd/server.crt
ETCD_KEY_FILE=/etc/etcd/server.key
ETCD_TRUSTED_CA_FILE=/etc/etcd/ca.crt
ETCD_CLIENT_CERT_AUTH=true
ETCD_PEER_CERT_FILE=/etc/etcd/server.crt
ETCD_PEER_KEY_FILE=/etc/etcd/server.key
ETCD_PEER_TRUSTED_CA_FILE=/etc/etcd/ca.crt
ETCD_PEER_CLIENT_CERT_AUTH=true
ETCD_AUTO_COMPACTION_MODE=periodic
ETCD_AUTO_COMPACTION_RETENTION=1
ETCD_QUOTA_BACKEND_BYTES=8589934592
The key settings:
ETCD_INITIAL_CLUSTERlists every member and must be identical on all three. It is only used the first time, when the data directory is empty.ETCD_CLIENT_CERT_AUTHandETCD_PEER_CLIENT_CERT_AUTHreject any connection without a certificate signed by your CA.ETCD_LISTEN_METRICS_URLSexposes the Prometheus metrics and/healthover plain HTTP on localhost only, so monitoring does not need client certificates.- Periodic auto-compaction with a retention of
1keeps one hour of key history and discards older revisions. ETCD_QUOTA_BACKEND_BYTESraises the database size limit from the default 2 GiB to 8 GiB, the maximum recommended value.
Create the systemd unit:
sudo nano /etc/systemd/system/etcd.service
[Unit]
Description=etcd key-value store
Documentation=https://etcd.io/docs/
Wants=network-online.target
After=network-online.target
[Service]
Type=notify
User=etcd
Group=etcd
EnvironmentFile=/etc/etcd/etcd.env
ExecStart=/usr/local/bin/etcd
Restart=on-failure
RestartSec=5
LimitNOFILE=65536
[Install]
WantedBy=multi-user.target
Step 4 - Opening the firewall and starting the cluster
Members use port 2380 to talk to each other, and clients use port 2379. Allow both only from the private network, on every server:
sudo ufw allow OpenSSH
sudo ufw allow from 10.0.0.0/24 to any port 2379,2380 proto tcp
sudo ufw enable
Start etcd on all three servers within a short time of each other. The first member waits for the others before it reports ready, so systemctl may appear to hang on it until the second one starts:
sudo systemctl daemon-reload
sudo systemctl enable --now etcd
Check the service and the local health endpoint:
systemctl is-active etcd
curl -s http://127.0.0.1:2381/health
active
{"health":"true","reason":""}
Step 5 - Connecting with etcdctl
etcdctl reads its connection settings from ETCDCTL_* environment variables. Save them in a file on each server so you can load them in any shell:
nano ~/.etcd/etcdctl.env
export ETCDCTL_ENDPOINTS=https://10.0.0.31:2379,https://10.0.0.32:2379,https://10.0.0.33:2379
export ETCDCTL_CACERT="$HOME/.etcd/ca.crt"
export ETCDCTL_CERT="$HOME/.etcd/admin.crt"
export ETCDCTL_KEY="$HOME/.etcd/admin.key"
Load it and inspect the cluster:
source ~/.etcd/etcdctl.env
etcdctl member list -w table
etcdctl endpoint status -w table
+------------------+---------+---------+------------------------+------------------------+------------+
| ID | STATUS | NAME | PEER ADDRS | CLIENT ADDRS | IS LEARNER |
+------------------+---------+---------+------------------------+------------------------+------------+
| 3a57933972cb5131 | started | etcd-01 | https://10.0.0.31:2380 | https://10.0.0.31:2379 | false |
| 8e9e05c52164694d | started | etcd-02 | https://10.0.0.32:2380 | https://10.0.0.32:2379 | false |
| e1b9c2c8f0a6d2a4 | started | etcd-03 | https://10.0.0.33:2380 | https://10.0.0.33:2379 | false |
+------------------+---------+---------+------------------------+------------------------+------------+
The endpoint status table shows the version, database size, and which member IS LEADER. Confirm that TLS is enforced by trying without a client certificate; the request must fail:
curl -s --cacert ~/.etcd/ca.crt https://10.0.0.31:2379/version || echo "rejected"
rejected
Step 6 - Working with keys
Write, read and delete keys. Keys are plain byte strings; a path-like naming scheme makes prefix queries easy:
etcdctl put /config/app/db_host db.internal
etcdctl put /config/app/pool_size 20
etcdctl get /config/app/ --prefix
etcdctl del /config/app/pool_size
OK
OK
/config/app/db_host
db.internal
/config/app/pool_size
20
1
Leases attach a time to live to keys, which is how services announce themselves and hold locks. The key disappears when the lease expires unless it is renewed:
etcdctl lease grant 60
lease 694d91b3c3f1e20a granted with TTL(60s)
etcdctl put /services/web/10.0.0.50 up --lease=694d91b3c3f1e20a
etcdctl lease timetolive 694d91b3c3f1e20a --keys
To follow changes live, run etcdctl watch /config/ --prefix in one terminal and change a key in another. Stop the watch with Ctrl+C.
Step 7 - Scheduling snapshot backups
A snapshot taken from one healthy member contains the whole keyspace and is all you need to rebuild the cluster. Create a backup script on one member, for example etcd-01:
sudo nano /usr/local/bin/etcd-backup
#!/usr/bin/env bash
set -euo pipefail
backup_dir="/var/backups/etcd"
keep_days=7
snapshot="${backup_dir}/snapshot-$(date +%Y%m%d-%H%M%S).db"
mkdir -p "$backup_dir"
chmod 700 "$backup_dir"
etcdctl \
--endpoints=https://127.0.0.1:2379 \
--cacert=/etc/etcd/ca.crt \
--cert=/etc/etcd/server.crt \
--key=/etc/etcd/server.key \
snapshot save "$snapshot"
etcdutl snapshot status "$snapshot" -w table
find "$backup_dir" -name 'snapshot-*.db' -mtime +"$keep_days" -delete
Make it executable and run it once as root, which can read the member key:
sudo chmod 750 /usr/local/bin/etcd-backup
sudo /usr/local/bin/etcd-backup
... Snapshot saved at /var/backups/etcd/snapshot-20260925-101500.db
+----------+----------+------------+------------+---------+
| HASH | REVISION | TOTAL KEYS | TOTAL SIZE | VERSION |
+----------+----------+------------+------------+---------+
| 7c3fd5e1 | 42 | 12 | 25 kB | 3.6.0 |
+----------+----------+------------+------------+---------+
Schedule it daily with a systemd service and timer:
sudo nano /etc/systemd/system/etcd-backup.service
[Unit]
Description=etcd snapshot backup
After=etcd.service
[Service]
Type=oneshot
ExecStart=/usr/local/bin/etcd-backup
sudo nano /etc/systemd/system/etcd-backup.timer
[Unit]
Description=Daily etcd snapshot backup
[Timer]
OnCalendar=*-*-* 02:00:00
RandomizedDelaySec=15m
Persistent=true
[Install]
WantedBy=timers.target
sudo systemctl daemon-reload
sudo systemctl enable --now etcd-backup.timer
systemctl list-timers etcd-backup.timer
Copy the snapshots off the server regularly, for example to object storage. A backup that lives only on a cluster member is lost with it.
Step 8 - Restoring the cluster from a snapshot
Restoring creates a new cluster from the snapshot, so you restore the same file on every member. Copy the snapshot to all three servers first, for example to /tmp/snapshot.db. Then stop etcd on all members:
sudo systemctl stop etcd
On each member, move the old data directory aside and restore with etcdutl, adjusting --name and --initial-advertise-peer-urls for that member (shown for etcd-01):
sudo mv /var/lib/etcd /var/lib/etcd.old
sudo etcdutl snapshot restore /tmp/snapshot.db \
--name etcd-01 \
--initial-cluster etcd-01=https://10.0.0.31:2380,etcd-02=https://10.0.0.32:2380,etcd-03=https://10.0.0.33:2380 \
--initial-cluster-token etcd-cluster-restored \
--initial-advertise-peer-urls https://10.0.0.31:2380 \
--data-dir /var/lib/etcd
sudo chown -R etcd:etcd /var/lib/etcd
sudo chmod 700 /var/lib/etcd
Start etcd on all three members and verify with etcdctl endpoint status -w table and etcdctl get /config/app/ --prefix. Once the cluster is healthy, remove /var/lib/etcd.old.
Step 9 - Routine maintenance
Auto-compaction discards old revisions, but the database file does not shrink until you defragment it. Check the size per member in the DB SIZE column:
etcdctl endpoint status -w table
Defragment one member at a time during a quiet period, because a member blocks reads and writes while it defragments:
etcdctl defrag --endpoints=https://10.0.0.31:2379
etcdctl defrag --endpoints=https://10.0.0.32:2379
etcdctl defrag --endpoints=https://10.0.0.33:2379
Finished defragmenting etcd member[https://10.0.0.31:2379]. took ...
For monitoring, point Prometheus or your metrics collector at http://127.0.0.1:2381/metrics on each member. The most useful series are etcd_server_has_leader, etcd_server_leader_changes_seen_total, etcd_disk_wal_fsync_duration_seconds and etcd_mvcc_db_total_size_in_bytes.
Troubleshooting
The first member never becomes ready. It is waiting for a quorum. Check that the other members are running and that port 2380 is open between them: nc -zv 10.0.0.32 2380. Look for TLS errors with sudo journalctl -u etcd -n 50 --no-pager; certificate is valid for ..., not ... means the member's IP is missing from its certificate.
etcdserver: mvcc: database space exceeded. The database reached the quota and the cluster only accepts reads and deletes. Compact and defragment (Step 9), then clear the alarm with etcdctl alarm disarm and check it is gone with etcdctl alarm list.
Frequent leader elections and slow requests. The disks are too slow. Check etcd_disk_wal_fsync_duration_seconds (the 99th percentile should stay under 10 ms), and move etcd to faster storage. Increasing ETCD_HEARTBEAT_INTERVAL and ETCD_ELECTION_TIMEOUT only helps when members are far apart on the network.
member ... has already been bootstrapped. The data directory already contains a cluster. Either keep it, or move /var/lib/etcd aside if you really want to bootstrap from scratch.
Conclusion
You now have a three-member etcd 3.6 cluster on Ubuntu 24.04 with mutual TLS on peer and client traffic, daily snapshots scheduled by systemd and a tested restore procedure. Next, enable etcd's role-based access control (etcdctl user add and etcdctl auth enable) to give each application its own permissions, set up alerts on the metrics listed above, or use the cluster as the external etcd for a Kubernetes control plane.
