Thanos adds long-term storage and a global query view to Prometheus. A sidecar next to Prometheus uploads each completed two-hour block of data to S3-compatible object storage, a Store Gateway serves those blocks back, and a Querier answers PromQL across both, so Grafana sees recent and historical data as one source. In this tutorial you will set up Thanos on Ubuntu 24.04 next to an existing Prometheus server: sidecar, Store Gateway, Querier and Compactor, each as a systemd service.
Prerequisites
To follow this tutorial you need:
- A server running Ubuntu 24.04 LTS (x86_64), for example a CubePath VPS, with at least 2 GB of RAM.
- A non-root user with
sudoprivileges. - Prometheus installed from the Ubuntu package (
sudo apt install prometheus) and scraping at least one target. If you installed Prometheus another way, adapt the paths in Step 2. - An empty bucket on an S3-compatible object storage service, with an access key and secret key that can read, write, list and delete objects in it.
All components run on the same server in this guide. Here is what each one does and the ports it will use:
| Component | Role | HTTP port | gRPC port |
|---|---|---|---|
| Sidecar | Uploads Prometheus blocks, serves recent data | 10902 | 10901 |
| Store Gateway | Serves historical blocks from the bucket | 10912 | 10911 |
| Querier | PromQL API and web UI over all stores | 10904 | 10903 |
| Compactor | Compacts, downsamples and applies retention | 10922 | none |
Step 1 - Installing the Thanos binary
Thanos is a single binary; each component is a subcommand (thanos sidecar, thanos store, and so on). Download the release archive and its checksums. Check the releases page for the latest version:
cd /tmp
THANOS_VERSION=0.42.4
curl -fLO "https://github.com/thanos-io/thanos/releases/download/v${THANOS_VERSION}/thanos-${THANOS_VERSION}.linux-amd64.tar.gz"
curl -fLO "https://github.com/thanos-io/thanos/releases/download/v${THANOS_VERSION}/sha256sums.txt"
sha256sum --ignore-missing -c sha256sums.txt
thanos-0.42.4.linux-amd64.tar.gz: OK
Extract and install the binary:
tar -xzf "thanos-${THANOS_VERSION}.linux-amd64.tar.gz"
sudo install -m 0755 "thanos-${THANOS_VERSION}.linux-amd64/thanos" /usr/local/bin/thanos
thanos --version
thanos, version 0.42.4 (branch: HEAD, revision: 45c4f39487e1653e606307d3952a6259fb96acae)
Create a system user for the Store Gateway, Querier and Compactor, plus their directories:
sudo useradd --system --no-create-home --shell /usr/sbin/nologin thanos
sudo install -d -o thanos -g thanos -m 0750 /var/lib/thanos
sudo install -d -m 0755 /etc/thanos
The sidecar is different: it reads and writes the Prometheus data directory, so it will run as the prometheus user.
Step 2 - Preparing Prometheus
The sidecar has two requirements on the Prometheus side.
First, every Prometheus server in a Thanos setup needs a unique set of external labels. Thanos stores them with each block and uses them to tell sources apart. Open the Prometheus configuration:
sudo nano /etc/prometheus/prometheus.yml
Add external_labels to the global section, replacing the values with your own:
global:
scrape_interval: 15s
evaluation_interval: 15s
external_labels:
cluster: "prod"
replica: "prometheus-01"
Second, Prometheus must stop compacting its own blocks, because the Thanos Compactor will do that in the bucket. Setting the minimum and maximum block durations to the same value (2h) disables local compaction. On Ubuntu, Prometheus reads its extra flags from /etc/default/prometheus:
sudo nano /etc/default/prometheus
Set the ARGS line to:
ARGS="--storage.tsdb.min-block-duration=2h --storage.tsdb.max-block-duration=2h --storage.tsdb.retention.time=7d"
--storage.tsdb.retention.time=7d keeps one week locally. Older data will be served from object storage, so you no longer need long local retention.
Check the configuration and restart Prometheus:
promtool check config /etc/prometheus/prometheus.yml
sudo systemctl restart prometheus
Confirm the flags are active:
curl -s http://localhost:9090/api/v1/status/flags | python3 -m json.tool | grep block-duration
"storage.tsdb.max-block-duration": "2h",
"storage.tsdb.min-block-duration": "2h",
The Ubuntu package stores the TSDB in /var/lib/prometheus/metrics2. Verify it:
sudo ls /var/lib/prometheus/metrics2
chunks_head lock queries.active wal
After a few hours you will also see directories with long names such as 01J8Z3K4...: those are completed blocks.
Step 3 - Configuring object storage
All Thanos components that touch the bucket read the same object storage file. Create it:
sudo nano /etc/thanos/objstore.yml
type: S3
config:
bucket: "your_bucket"
endpoint: "your_s3_endpoint"
region: "your_region"
access_key: "your_access_key"
secret_key: "your_secret_key"
insecure: false
endpoint is the host name of the S3 API without https://, for example s3.eu-central-1.amazonaws.com or your provider's S3 host. Keep insecure: false so the connection uses TLS.
The file holds credentials. Make it readable only by root, prometheus (for the sidecar) and thanos:
sudo groupadd --system thanos-objstore
sudo usermod -aG thanos-objstore prometheus
sudo usermod -aG thanos-objstore thanos
sudo chown root:thanos-objstore /etc/thanos/objstore.yml
sudo chmod 640 /etc/thanos/objstore.yml
Test the credentials by listing the bucket:
sudo -u thanos thanos tools bucket ls --objstore.config-file=/etc/thanos/objstore.yml
An empty bucket prints only a few log lines and no block IDs. An authentication or endpoint error appears immediately.
Step 4 - Running the sidecar
Create the sidecar service:
sudo nano /etc/systemd/system/thanos-sidecar.service
[Unit]
Description=Thanos Sidecar
After=network-online.target prometheus.service
Wants=network-online.target
[Service]
User=prometheus
Group=prometheus
ExecStart=/usr/local/bin/thanos sidecar \
--tsdb.path=/var/lib/prometheus/metrics2 \
--prometheus.url=http://localhost:9090 \
--objstore.config-file=/etc/thanos/objstore.yml \
--http-address=127.0.0.1:10902 \
--grpc-address=127.0.0.1:10901
Restart=on-failure
RestartSec=5
[Install]
WantedBy=multi-user.target
Start it:
sudo systemctl daemon-reload
sudo systemctl enable --now thanos-sidecar
Check that it is ready:
curl http://127.0.0.1:10902/-/ready
OK
The sidecar uploads a block as soon as Prometheus finishes one, which happens every two hours. You can follow uploads in the log:
sudo journalctl -u thanos-sidecar -f
Look for a line with msg="upload new block". Afterwards, thanos tools bucket ls from Step 3 lists the block ID.
Step 5 - Running the Store Gateway
The Store Gateway makes the blocks in the bucket queryable. It caches a small part of each block (the index headers) on local disk.
sudo nano /etc/systemd/system/thanos-store.service
[Unit]
Description=Thanos Store Gateway
After=network-online.target
Wants=network-online.target
[Service]
User=thanos
Group=thanos
ExecStart=/usr/local/bin/thanos store \
--data-dir=/var/lib/thanos/store \
--objstore.config-file=/etc/thanos/objstore.yml \
--http-address=127.0.0.1:10912 \
--grpc-address=127.0.0.1:10911
Restart=on-failure
RestartSec=5
[Install]
WantedBy=multi-user.target
sudo systemctl daemon-reload
sudo systemctl enable --now thanos-store
curl http://127.0.0.1:10912/-/ready
OK
Step 6 - Running the Querier
The Querier fans out each PromQL query to the sidecar (recent data) and the Store Gateway (historical data) and merges the results. The list of stores lives in an endpoint file, which the Querier reloads periodically:
sudo nano /etc/thanos/endpoints.yml
endpoints:
- address: "127.0.0.1:10901"
- address: "127.0.0.1:10911"
Create the service:
sudo nano /etc/systemd/system/thanos-query.service
[Unit]
Description=Thanos Querier
After=network-online.target thanos-sidecar.service thanos-store.service
Wants=network-online.target
[Service]
User=thanos
Group=thanos
ExecStart=/usr/local/bin/thanos query \
--http-address=127.0.0.1:10904 \
--grpc-address=127.0.0.1:10903 \
--endpoint.sd-config-file=/etc/thanos/endpoints.yml \
--query.replica-label=replica
Restart=on-failure
RestartSec=5
[Install]
WantedBy=multi-user.target
--query.replica-label=replica tells the Querier that series differing only in the replica label are copies of each other. If you later run two identical Prometheus servers for high availability (with replica: "prometheus-01" and replica: "prometheus-02"), the Querier deduplicates their data automatically.
Start it:
sudo systemctl daemon-reload
sudo systemctl enable --now thanos-query
Check that both stores are connected:
curl -s http://127.0.0.1:10904/api/v1/stores | python3 -m json.tool
The output lists a sidecar entry for 127.0.0.1:10901 and a store entry for 127.0.0.1:10911, each with "lastError": null. The sidecar entry shows your external labels (cluster="prod", replica="prometheus-01").
Run a query through the Querier:
curl -s http://127.0.0.1:10904/api/v1/query --data-urlencode 'query=up' | python3 -m json.tool | head -20
Every series now carries the cluster external label. The Querier also has a web UI similar to Prometheus's. Open it through an SSH tunnel from your workstation:
ssh -L 10904:127.0.0.1:10904 your_user@your_server_ip
Then browse to http://localhost:10904.
To use Thanos from Grafana, add a Prometheus data source with the URL http://localhost:10904 (or the tunnelled address) instead of Prometheus itself. Dashboards then work across the full retention period.
Step 7 - Running the Compactor
Without the Compactor, the bucket fills up with thousands of two-hour blocks, and queries over months become slow. The Compactor merges blocks into larger ones, creates downsampled copies at 5-minute and 1-hour resolution, and deletes data that falls outside your retention.
sudo nano /etc/systemd/system/thanos-compact.service
[Unit]
Description=Thanos Compactor
After=network-online.target
Wants=network-online.target
[Service]
User=thanos
Group=thanos
ExecStart=/usr/local/bin/thanos compact \
--wait \
--data-dir=/var/lib/thanos/compact \
--objstore.config-file=/etc/thanos/objstore.yml \
--http-address=127.0.0.1:10922 \
--retention.resolution-raw=30d \
--retention.resolution-5m=180d \
--retention.resolution-1h=2y
Restart=on-failure
RestartSec=30
[Install]
WantedBy=multi-user.target
The retention flags keep full-resolution data for 30 days, 5-minute data for 180 days and 1-hour data for two years. Queries over long ranges automatically use the downsampled data. The default for each flag is 0d, which means keep forever. --wait keeps the process running and re-checks the bucket periodically instead of exiting after one pass.
ImportantRun exactly one Compactor per bucket. Two Compactors working on the same bucket can corrupt data.
Start it:
sudo systemctl daemon-reload
sudo systemctl enable --now thanos-compact
The Compactor needs local disk space in /var/lib/thanos/compact roughly equal to the size of the largest blocks it merges. Check its progress with:
sudo journalctl -u thanos-compact -n 30 --no-pager
Troubleshooting
The sidecar logs "no external labels configured on Prometheus server". Add external_labels to prometheus.yml (Step 2) and restart Prometheus, then the sidecar.
The sidecar fails with "permission denied" on the TSDB path. It must run as the prometheus user, and --tsdb.path must match the directory Prometheus writes to.
No blocks appear in the bucket. Blocks are only uploaded after Prometheus completes one, so wait at least two hours after the first start. If the log shows errors from the S3 client, test the credentials again with thanos tools bucket ls.
/api/v1/stores shows an entry with lastError. The component on that address is not running or listens on a different port. Compare the --grpc-address flags with /etc/thanos/endpoints.yml.
The Compactor stops with a "halt" error. It halts on problems such as overlapping blocks, which usually come from two Prometheus servers with identical external labels or from local compaction still being enabled. Fix the cause, then restart the service.
Conclusion
Your Prometheus server now ships its data to object storage through the Thanos sidecar, and the Querier combines recent and historical data behind a single Prometheus-compatible API, with the Compactor handling compaction, downsampling and retention. As next steps, point Grafana at the Querier, add more Prometheus servers with their own sidecars and unique external labels to endpoints.yml, and consider thanos rule if you need alerting rules evaluated over the global view.
