A complete metrics stack needs four pieces: exporters that expose metrics, a server that collects and evaluates them, something that routes alerts to people, and a UI for dashboards. In this tutorial you will build that stack on Ubuntu 24.04 with Node Exporter, Prometheus, Alertmanager and Grafana, monitor both the monitoring server and additional hosts, define alert rules for down hosts, CPU, memory and disks, and receive notifications by email.
Prerequisites
To follow this tutorial you need:
- A monitoring server running Ubuntu 24.04 LTS, for example a CubePath VPS, with at least 2 vCPUs, 4 GB of RAM and 40 GB of disk.
- One or more servers to monitor, running Ubuntu 24.04 or Debian 12, reachable from the monitoring server over a private network or the internet.
- A non-root user with
sudoprivileges on every server, and UFW enabled with SSH allowed. - An SMTP account to send alert emails (host, port, username and password).
- The public IP address of your workstation, referred to as
your_admin_ip.
How the components fit together
| Component | Port | Role |
|---|---|---|
| Node Exporter | 9100 | Runs on every host and exposes CPU, memory, disk and network metrics. |
| Prometheus | 9090 | Scrapes exporters every 15 seconds, stores the data and evaluates alert rules. |
| Alertmanager | 9093 | Receives firing alerts from Prometheus, groups and deduplicates them, and sends notifications. |
| Grafana | 3000 | Queries Prometheus and shows dashboards. |
Prometheus pulls from exporters; alerts flow from Prometheus to Alertmanager; Grafana only reads. In this setup Prometheus and Alertmanager listen on 127.0.0.1 only, and Grafana is the only service reachable from your workstation.
Step 1 - Installing Node Exporter on every host
Ubuntu and Debian package Node Exporter with a systemd service and security updates. Run this on the monitoring server and on each host you want to monitor:
sudo apt update
sudo apt install prometheus-node-exporter
Verify that it answers locally:
curl -s http://localhost:9100/metrics | grep '^node_uname_info'
node_uname_info{domainname="(none)",machine="x86_64",nodename="web-01",release="6.8.0-79-generic",sysname="Linux",version="#79-Ubuntu SMP ..."} 1
Node Exporter has no authentication. On each monitored host (not the monitoring server), allow port 9100 only from the monitoring server, replacing your_monitoring_ip with its address:
sudo ufw allow from your_monitoring_ip to any port 9100 proto tcp
From the monitoring server, confirm you can reach each host:
curl -s http://web01_ip:9100/metrics | head -n 3
If the command hangs, check the firewall rule on the host and any cloud firewall in front of it.
Step 2 - Installing Prometheus
The rest of the steps run on the monitoring server. Create a system user and the directories Prometheus needs:
sudo useradd --system --no-create-home --shell /usr/sbin/nologin prometheus
sudo mkdir -p /etc/prometheus/rules /var/lib/prometheus
sudo chown prometheus:prometheus /var/lib/prometheus
Download the latest release from the Prometheus download page and verify its checksum. Adjust the version to the current one:
PROM_VERSION=3.15.0
cd /tmp
curl -LO "https://github.com/prometheus/prometheus/releases/download/v${PROM_VERSION}/prometheus-${PROM_VERSION}.linux-amd64.tar.gz"
curl -Lo prometheus-sha256sums.txt "https://github.com/prometheus/prometheus/releases/download/v${PROM_VERSION}/sha256sums.txt"
sha256sum --check --ignore-missing prometheus-sha256sums.txt
prometheus-3.15.0.linux-amd64.tar.gz: OK
Install the binaries:
tar xzf "prometheus-${PROM_VERSION}.linux-amd64.tar.gz"
sudo install -m 0755 "prometheus-${PROM_VERSION}.linux-amd64/prometheus" "prometheus-${PROM_VERSION}.linux-amd64/promtool" /usr/local/bin/
prometheus --version | head -n 1
prometheus, version 3.15.0 (branch: HEAD, revision: ...)
You will write its configuration in Step 5, after Alertmanager and the alert rules exist.
Step 3 - Installing Alertmanager
Alertmanager is a separate binary with its own user and data directory:
sudo useradd --system --no-create-home --shell /usr/sbin/nologin alertmanager
sudo mkdir -p /etc/alertmanager /var/lib/alertmanager
sudo chown alertmanager:alertmanager /var/lib/alertmanager
Download and verify it, using the current version from the same download page:
AM_VERSION=0.34.1
cd /tmp
curl -LO "https://github.com/prometheus/alertmanager/releases/download/v${AM_VERSION}/alertmanager-${AM_VERSION}.linux-amd64.tar.gz"
curl -Lo alertmanager-sha256sums.txt "https://github.com/prometheus/alertmanager/releases/download/v${AM_VERSION}/sha256sums.txt"
sha256sum --check --ignore-missing alertmanager-sha256sums.txt
tar xzf "alertmanager-${AM_VERSION}.linux-amd64.tar.gz"
sudo install -m 0755 "alertmanager-${AM_VERSION}.linux-amd64/alertmanager" "alertmanager-${AM_VERSION}.linux-amd64/amtool" /usr/local/bin/
Create the configuration. It sends every alert by email, groups alerts that share an alertname and instance into one message, and repeats unresolved alerts every 4 hours. Replace the SMTP values and addresses with your own:
sudo nano /etc/alertmanager/alertmanager.yml
global:
smtp_smarthost: "smtp.example.com:587"
smtp_from: "[email protected]"
smtp_auth_username: "[email protected]"
smtp_auth_password: "your_smtp_password"
smtp_require_tls: true
route:
receiver: ops-email
group_by: ["alertname", "instance"]
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
receivers:
- name: ops-email
email_configs:
- to: "[email protected]"
send_resolved: true
inhibit_rules:
- source_matchers: ['severity="critical"']
target_matchers: ['severity="warning"']
equal: ["alertname", "instance"]
The inhibit_rules block silences a warning while a critical alert with the same name is firing on the same host, so you do not get two emails about one problem.
The file contains the SMTP password, so restrict it and validate it:
sudo chown root:alertmanager /etc/alertmanager/alertmanager.yml
sudo chmod 640 /etc/alertmanager/alertmanager.yml
sudo amtool check-config /etc/alertmanager/alertmanager.yml
Checking '/etc/alertmanager/alertmanager.yml' SUCCESS
Found:
- global config
- route
- 1 inhibit rules
- 1 receivers
- 0 templates
Create the systemd unit. --cluster.listen-address= with an empty value disables the high-availability gossip port, which a single instance does not need:
sudo nano /etc/systemd/system/alertmanager.service
[Unit]
Description=Prometheus Alertmanager
Wants=network-online.target
After=network-online.target
[Service]
User=alertmanager
Group=alertmanager
Type=simple
ExecStart=/usr/local/bin/alertmanager \
--config.file=/etc/alertmanager/alertmanager.yml \
--storage.path=/var/lib/alertmanager \
--web.listen-address=127.0.0.1:9093 \
--cluster.listen-address=
ExecReload=/bin/kill -HUP $MAINPID
Restart=on-failure
NoNewPrivileges=true
ProtectSystem=full
ProtectHome=true
[Install]
WantedBy=multi-user.target
Start it and check that it is ready:
sudo systemctl daemon-reload
sudo systemctl enable --now alertmanager
curl -s http://127.0.0.1:9093/-/ready
OK
Step 4 - Writing alert rules
Alert rules are PromQL expressions that Prometheus evaluates on a schedule. When an expression returns results for longer than for, the alert fires and is sent to Alertmanager. Create a rules file for host alerts:
sudo nano /etc/prometheus/rules/node.yml
groups:
- name: node
rules:
- alert: InstanceDown
expr: up == 0
for: 2m
labels:
severity: critical
annotations:
summary: "{{ $labels.instance }} is down"
description: "Prometheus has not been able to scrape {{ $labels.instance }} (job {{ $labels.job }}) for 2 minutes."
- alert: HostHighCpu
expr: 100 * (1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m]))) > 90
for: 10m
labels:
severity: warning
annotations:
summary: "High CPU on {{ $labels.instance }}"
description: "CPU usage has been above 90% for 10 minutes (current: {{ $value | printf \"%.1f\" }}%)."
- alert: HostLowMemory
expr: 100 * (1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) > 90
for: 5m
labels:
severity: warning
annotations:
summary: "Low memory on {{ $labels.instance }}"
description: "Memory usage is {{ $value | printf \"%.1f\" }}%."
- alert: HostDiskAlmostFull
expr: 100 * node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"} / node_filesystem_size_bytes{fstype!~"tmpfs|overlay"} < 10
for: 5m
labels:
severity: warning
annotations:
summary: "Disk almost full on {{ $labels.instance }}"
description: "{{ $labels.mountpoint }} has {{ $value | printf \"%.1f\" }}% free space left."
- alert: HostDiskWillFillIn4Hours
expr: predict_linear(node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"}[1h], 4 * 3600) < 0
for: 10m
labels:
severity: critical
annotations:
summary: "Disk on {{ $labels.instance }} will fill within 4 hours"
description: "At the current rate, {{ $labels.mountpoint }} will run out of space in less than 4 hours."
The severity label is what the inhibit rule in Alertmanager matches on, and the annotations become the body of the email. Validate the file:
promtool check rules /etc/prometheus/rules/node.yml
Checking /etc/prometheus/rules/node.yml
SUCCESS: 5 rules found
Step 5 - Configuring and starting Prometheus
Now write the Prometheus configuration. It scrapes Prometheus, Alertmanager and every Node Exporter, loads the rules, and sends alerts to Alertmanager. Replace the example IPs and names with your hosts:
sudo nano /etc/prometheus/prometheus.yml
global:
scrape_interval: 15s
evaluation_interval: 15s
alerting:
alertmanagers:
- static_configs:
- targets: ["127.0.0.1:9093"]
rule_files:
- /etc/prometheus/rules/*.yml
scrape_configs:
- job_name: prometheus
static_configs:
- targets: ["127.0.0.1:9090"]
- job_name: alertmanager
static_configs:
- targets: ["127.0.0.1:9093"]
- job_name: node
static_configs:
- targets: ["127.0.0.1:9100"]
labels:
host: monitoring-01
- targets: ["10.0.0.11:9100"]
labels:
host: web-01
- targets: ["10.0.0.12:9100"]
labels:
host: db-01
Validate the configuration. promtool also checks every rules file it references:
promtool check config /etc/prometheus/prometheus.yml
Checking /etc/prometheus/prometheus.yml
SUCCESS: 1 rule files found
SUCCESS: /etc/prometheus/prometheus.yml is valid prometheus config file syntax
Checking /etc/prometheus/rules/node.yml
SUCCESS: 5 rules found
Create the systemd unit. Prometheus keeps 30 days of data or 20 GB, whichever limit is reached first:
sudo nano /etc/systemd/system/prometheus.service
[Unit]
Description=Prometheus monitoring system
Wants=network-online.target
After=network-online.target
[Service]
User=prometheus
Group=prometheus
Type=simple
ExecStart=/usr/local/bin/prometheus \
--config.file=/etc/prometheus/prometheus.yml \
--storage.tsdb.path=/var/lib/prometheus \
--storage.tsdb.retention.time=30d \
--storage.tsdb.retention.size=20GB \
--web.listen-address=127.0.0.1:9090
ExecReload=/bin/kill -HUP $MAINPID
Restart=on-failure
NoNewPrivileges=true
ProtectSystem=full
ProtectHome=true
[Install]
WantedBy=multi-user.target
Start Prometheus:
sudo systemctl daemon-reload
sudo systemctl enable --now prometheus
curl -s http://127.0.0.1:9090/-/ready
Prometheus Server is Ready.
Check that every target is up. The up metric is 1 for each target whose last scrape succeeded, so this query should return the number of targets you configured (5 in this example):
curl -s 'http://127.0.0.1:9090/api/v1/query?query=count(up==1)'
{"status":"success","data":{"resultType":"vector","result":[{"metric":{},"value":[1790336120.456,"5"]}]}}
You can also use the web UI. Because Prometheus listens on localhost only, open an SSH tunnel from your workstation:
ssh -L 9090:127.0.0.1:9090 -L 9093:127.0.0.1:9093 your_user@your_server_ip
With the tunnel open, browse to http://localhost:9090/targets and confirm every target is UP. The Alerts page lists the five rules, all Inactive. http://localhost:9093 shows the Alertmanager UI.
Step 6 - Installing Grafana
Add the Grafana APT repository and install Grafana OSS:
sudo apt install apt-transport-https wget
sudo mkdir -p /etc/apt/keyrings
sudo wget -O /etc/apt/keyrings/grafana.asc https://apt.grafana.com/gpg-full.key
sudo chmod 644 /etc/apt/keyrings/grafana.asc
echo "deb [signed-by=/etc/apt/keyrings/grafana.asc] https://apt.grafana.com stable main" | sudo tee /etc/apt/sources.list.d/grafana.list
sudo apt update
sudo apt install grafana
Provision Prometheus and Alertmanager as data sources so they exist the moment Grafana starts:
sudo nano /etc/grafana/provisioning/datasources/stack.yaml
apiVersion: 1
datasources:
- name: Prometheus
uid: prometheus
type: prometheus
access: proxy
url: http://127.0.0.1:9090
isDefault: true
jsonData:
timeInterval: 15s
- name: Alertmanager
uid: alertmanager
type: alertmanager
access: proxy
url: http://127.0.0.1:9093
jsonData:
implementation: prometheus
Start Grafana and allow port 3000 from your workstation only:
sudo systemctl daemon-reload
sudo systemctl enable --now grafana-server
sudo ufw allow from your_admin_ip to any port 3000 proto tcp
Browse to http://your_server_ip:3000, log in with admin / admin, and set a new password when prompted. Go to Connections > Data sources, open Prometheus and click Test; you should see Successfully queried the Prometheus API.
ImportantPlain HTTP sends your Grafana password unencrypted. For anything beyond a test, put Grafana behind Nginx with a Let's Encrypt certificate and set
root_urlin/etc/grafana/grafana.ini, as described in the tutorial "How to Install Grafana on Ubuntu 24.04 and Build Dashboards".
Step 7 - Importing dashboards
Instead of building host dashboards from scratch, import the community standard:
- Go to Dashboards > New > Import.
- Enter
1860(Node Exporter Full) and click Load. - Select the Prometheus data source and click Import.
Use the Host drop-down at the top to switch between monitoring-01, web-01 and db-01. The CPU, memory, disk and network panels should all show data within a minute.
Because Alertmanager is also a data source, Alerting > Alert rules in Grafana lists the Prometheus rules, and Alerting > Silences (with the Alertmanager data source selected) lets you create silences for planned maintenance without leaving Grafana.
Step 8 - Testing the alerting pipeline
Test the path in two stages. First, confirm that Alertmanager can deliver email by injecting a fake alert with amtool:
amtool alert add alertname=TestAlert severity=warning instance=test-host \
--annotation=summary="Test alert from amtool" \
--alertmanager.url=http://127.0.0.1:9093
After the 30-second group_wait, the message arrives at the address in alertmanager.yml. If it does not, the Alertmanager log shows the SMTP error:
sudo journalctl -u alertmanager -n 30 --no-pager
Next, test a real alert end to end. On one of the monitored hosts, stop Node Exporter:
sudo systemctl stop prometheus-node-exporter
Within about 15 seconds the target turns DOWN in Prometheus and InstanceDown becomes Pending. After the 2-minute for period it turns Firing, and you can list it from the monitoring server:
amtool alert query --alertmanager.url=http://127.0.0.1:9093
Alertname Starts At Summary
InstanceDown 2026-09-25 11:32:10 UTC 10.0.0.11:9100 is down
The email follows. Start the exporter again and, because send_resolved is true, a RESOLVED email arrives a few minutes later:
sudo systemctl start prometheus-node-exporter
Step 9 - Adding more hosts later
To monitor a new server, install prometheus-node-exporter on it and open port 9100 to the monitoring server as in Step 1. Then add its target to the node job in /etc/prometheus/prometheus.yml and reload Prometheus without a restart:
promtool check config /etc/prometheus/prometheus.yml && sudo systemctl reload prometheus
With many hosts, editing prometheus.yml becomes tedious. Move the targets to a file and use file-based service discovery instead: replace static_configs in the node job with:
file_sd_configs:
- files:
- /etc/prometheus/targets/*.yml
Each file in /etc/prometheus/targets/ holds a list of targets with labels, in the same shape as a static_configs entry, and Prometheus picks up changes automatically without a reload.
Troubleshooting
Alerts show as firing in Prometheus but no email arrives. Check http://localhost:9090/status (through the tunnel) under Alertmanagers for 127.0.0.1:9093, then read sudo journalctl -u alertmanager for SMTP authentication or TLS errors. With port 587, keep smtp_require_tls: true and make sure the username and password are correct.
A remote target is DOWN with context deadline exceeded. The monitoring server cannot reach port 9100. Check the UFW rule on the target and any provider firewall between the two hosts.
HostDiskAlmostFull fires for /boot/efi or snap mounts. Exclude those mount points by adding mountpoint!~"/boot/efi|/snap/.*" to the selector in the rule.
Grafana panels show No data. The Prometheus data source URL must be http://127.0.0.1:9090, since Prometheus listens only on localhost. Test the data source and check that the node job is UP.
Conclusion
You now have a working monitoring stack: Node Exporter on every host, Prometheus collecting and evaluating alert rules, Alertmanager sending grouped email notifications with resolution messages, and Grafana dashboards for day-to-day visibility. From here, add exporters for your services (Blackbox Exporter for HTTP and TLS checks, MySQL or PostgreSQL exporters), route critical alerts to Slack or PagerDuty with a second Alertmanager receiver, and back up /etc/prometheus, /etc/alertmanager and /var/lib/grafana regularly.
