A complete metrics stack needs four pieces: exporters that expose metrics, a server that collects and evaluates them, something that routes alerts to people, and a UI for dashboards. In this tutorial you will build that stack on Ubuntu 24.04 with Node Exporter, Prometheus, Alertmanager and Grafana, monitor both the monitoring server and additional hosts, define alert rules for down hosts, CPU, memory and disks, and receive notifications by email.

Prerequisites

To follow this tutorial you need:

  • A monitoring server running Ubuntu 24.04 LTS, for example a CubePath VPS, with at least 2 vCPUs, 4 GB of RAM and 40 GB of disk.
  • One or more servers to monitor, running Ubuntu 24.04 or Debian 12, reachable from the monitoring server over a private network or the internet.
  • A non-root user with sudo privileges on every server, and UFW enabled with SSH allowed.
  • An SMTP account to send alert emails (host, port, username and password).
  • The public IP address of your workstation, referred to as your_admin_ip.

How the components fit together

ComponentPortRole
Node Exporter9100Runs on every host and exposes CPU, memory, disk and network metrics.
Prometheus9090Scrapes exporters every 15 seconds, stores the data and evaluates alert rules.
Alertmanager9093Receives firing alerts from Prometheus, groups and deduplicates them, and sends notifications.
Grafana3000Queries Prometheus and shows dashboards.

Prometheus pulls from exporters; alerts flow from Prometheus to Alertmanager; Grafana only reads. In this setup Prometheus and Alertmanager listen on 127.0.0.1 only, and Grafana is the only service reachable from your workstation.

Step 1 - Installing Node Exporter on every host

Ubuntu and Debian package Node Exporter with a systemd service and security updates. Run this on the monitoring server and on each host you want to monitor:

sudo apt update
sudo apt install prometheus-node-exporter

Verify that it answers locally:

curl -s http://localhost:9100/metrics | grep '^node_uname_info'
node_uname_info{domainname="(none)",machine="x86_64",nodename="web-01",release="6.8.0-79-generic",sysname="Linux",version="#79-Ubuntu SMP ..."} 1

Node Exporter has no authentication. On each monitored host (not the monitoring server), allow port 9100 only from the monitoring server, replacing your_monitoring_ip with its address:

sudo ufw allow from your_monitoring_ip to any port 9100 proto tcp

From the monitoring server, confirm you can reach each host:

curl -s http://web01_ip:9100/metrics | head -n 3

If the command hangs, check the firewall rule on the host and any cloud firewall in front of it.

Step 2 - Installing Prometheus

The rest of the steps run on the monitoring server. Create a system user and the directories Prometheus needs:

sudo useradd --system --no-create-home --shell /usr/sbin/nologin prometheus
sudo mkdir -p /etc/prometheus/rules /var/lib/prometheus
sudo chown prometheus:prometheus /var/lib/prometheus

Download the latest release from the Prometheus download page and verify its checksum. Adjust the version to the current one:

PROM_VERSION=3.15.0
cd /tmp
curl -LO "https://github.com/prometheus/prometheus/releases/download/v${PROM_VERSION}/prometheus-${PROM_VERSION}.linux-amd64.tar.gz"
curl -Lo prometheus-sha256sums.txt "https://github.com/prometheus/prometheus/releases/download/v${PROM_VERSION}/sha256sums.txt"
sha256sum --check --ignore-missing prometheus-sha256sums.txt
prometheus-3.15.0.linux-amd64.tar.gz: OK

Install the binaries:

tar xzf "prometheus-${PROM_VERSION}.linux-amd64.tar.gz"
sudo install -m 0755 "prometheus-${PROM_VERSION}.linux-amd64/prometheus" "prometheus-${PROM_VERSION}.linux-amd64/promtool" /usr/local/bin/
prometheus --version | head -n 1
prometheus, version 3.15.0 (branch: HEAD, revision: ...)

You will write its configuration in Step 5, after Alertmanager and the alert rules exist.

Step 3 - Installing Alertmanager

Alertmanager is a separate binary with its own user and data directory:

sudo useradd --system --no-create-home --shell /usr/sbin/nologin alertmanager
sudo mkdir -p /etc/alertmanager /var/lib/alertmanager
sudo chown alertmanager:alertmanager /var/lib/alertmanager

Download and verify it, using the current version from the same download page:

AM_VERSION=0.34.1
cd /tmp
curl -LO "https://github.com/prometheus/alertmanager/releases/download/v${AM_VERSION}/alertmanager-${AM_VERSION}.linux-amd64.tar.gz"
curl -Lo alertmanager-sha256sums.txt "https://github.com/prometheus/alertmanager/releases/download/v${AM_VERSION}/sha256sums.txt"
sha256sum --check --ignore-missing alertmanager-sha256sums.txt
tar xzf "alertmanager-${AM_VERSION}.linux-amd64.tar.gz"
sudo install -m 0755 "alertmanager-${AM_VERSION}.linux-amd64/alertmanager" "alertmanager-${AM_VERSION}.linux-amd64/amtool" /usr/local/bin/

Create the configuration. It sends every alert by email, groups alerts that share an alertname and instance into one message, and repeats unresolved alerts every 4 hours. Replace the SMTP values and addresses with your own:

sudo nano /etc/alertmanager/alertmanager.yml
global:
  smtp_smarthost: "smtp.example.com:587"
  smtp_from: "[email protected]"
  smtp_auth_username: "[email protected]"
  smtp_auth_password: "your_smtp_password"
  smtp_require_tls: true

route:
  receiver: ops-email
  group_by: ["alertname", "instance"]
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 4h

receivers:
  - name: ops-email
    email_configs:
      - to: "[email protected]"
        send_resolved: true

inhibit_rules:
  - source_matchers: ['severity="critical"']
    target_matchers: ['severity="warning"']
    equal: ["alertname", "instance"]

The inhibit_rules block silences a warning while a critical alert with the same name is firing on the same host, so you do not get two emails about one problem.

The file contains the SMTP password, so restrict it and validate it:

sudo chown root:alertmanager /etc/alertmanager/alertmanager.yml
sudo chmod 640 /etc/alertmanager/alertmanager.yml
sudo amtool check-config /etc/alertmanager/alertmanager.yml
Checking '/etc/alertmanager/alertmanager.yml'  SUCCESS
Found:
 - global config
 - route
 - 1 inhibit rules
 - 1 receivers
 - 0 templates

Create the systemd unit. --cluster.listen-address= with an empty value disables the high-availability gossip port, which a single instance does not need:

sudo nano /etc/systemd/system/alertmanager.service
[Unit]
Description=Prometheus Alertmanager
Wants=network-online.target
After=network-online.target

[Service]
User=alertmanager
Group=alertmanager
Type=simple
ExecStart=/usr/local/bin/alertmanager \
  --config.file=/etc/alertmanager/alertmanager.yml \
  --storage.path=/var/lib/alertmanager \
  --web.listen-address=127.0.0.1:9093 \
  --cluster.listen-address=
ExecReload=/bin/kill -HUP $MAINPID
Restart=on-failure
NoNewPrivileges=true
ProtectSystem=full
ProtectHome=true

[Install]
WantedBy=multi-user.target

Start it and check that it is ready:

sudo systemctl daemon-reload
sudo systemctl enable --now alertmanager
curl -s http://127.0.0.1:9093/-/ready
OK

Step 4 - Writing alert rules

Alert rules are PromQL expressions that Prometheus evaluates on a schedule. When an expression returns results for longer than for, the alert fires and is sent to Alertmanager. Create a rules file for host alerts:

sudo nano /etc/prometheus/rules/node.yml
groups:
  - name: node
    rules:
      - alert: InstanceDown
        expr: up == 0
        for: 2m
        labels:
          severity: critical
        annotations:
          summary: "{{ $labels.instance }} is down"
          description: "Prometheus has not been able to scrape {{ $labels.instance }} (job {{ $labels.job }}) for 2 minutes."

      - alert: HostHighCpu
        expr: 100 * (1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m]))) > 90
        for: 10m
        labels:
          severity: warning
        annotations:
          summary: "High CPU on {{ $labels.instance }}"
          description: "CPU usage has been above 90% for 10 minutes (current: {{ $value | printf \"%.1f\" }}%)."

      - alert: HostLowMemory
        expr: 100 * (1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) > 90
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "Low memory on {{ $labels.instance }}"
          description: "Memory usage is {{ $value | printf \"%.1f\" }}%."

      - alert: HostDiskAlmostFull
        expr: 100 * node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"} / node_filesystem_size_bytes{fstype!~"tmpfs|overlay"} < 10
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "Disk almost full on {{ $labels.instance }}"
          description: "{{ $labels.mountpoint }} has {{ $value | printf \"%.1f\" }}% free space left."

      - alert: HostDiskWillFillIn4Hours
        expr: predict_linear(node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"}[1h], 4 * 3600) < 0
        for: 10m
        labels:
          severity: critical
        annotations:
          summary: "Disk on {{ $labels.instance }} will fill within 4 hours"
          description: "At the current rate, {{ $labels.mountpoint }} will run out of space in less than 4 hours."

The severity label is what the inhibit rule in Alertmanager matches on, and the annotations become the body of the email. Validate the file:

promtool check rules /etc/prometheus/rules/node.yml
Checking /etc/prometheus/rules/node.yml
  SUCCESS: 5 rules found

Step 5 - Configuring and starting Prometheus

Now write the Prometheus configuration. It scrapes Prometheus, Alertmanager and every Node Exporter, loads the rules, and sends alerts to Alertmanager. Replace the example IPs and names with your hosts:

sudo nano /etc/prometheus/prometheus.yml
global:
  scrape_interval: 15s
  evaluation_interval: 15s

alerting:
  alertmanagers:
    - static_configs:
        - targets: ["127.0.0.1:9093"]

rule_files:
  - /etc/prometheus/rules/*.yml

scrape_configs:
  - job_name: prometheus
    static_configs:
      - targets: ["127.0.0.1:9090"]

  - job_name: alertmanager
    static_configs:
      - targets: ["127.0.0.1:9093"]

  - job_name: node
    static_configs:
      - targets: ["127.0.0.1:9100"]
        labels:
          host: monitoring-01
      - targets: ["10.0.0.11:9100"]
        labels:
          host: web-01
      - targets: ["10.0.0.12:9100"]
        labels:
          host: db-01

Validate the configuration. promtool also checks every rules file it references:

promtool check config /etc/prometheus/prometheus.yml
Checking /etc/prometheus/prometheus.yml
  SUCCESS: 1 rule files found
 SUCCESS: /etc/prometheus/prometheus.yml is valid prometheus config file syntax

Checking /etc/prometheus/rules/node.yml
  SUCCESS: 5 rules found

Create the systemd unit. Prometheus keeps 30 days of data or 20 GB, whichever limit is reached first:

sudo nano /etc/systemd/system/prometheus.service
[Unit]
Description=Prometheus monitoring system
Wants=network-online.target
After=network-online.target

[Service]
User=prometheus
Group=prometheus
Type=simple
ExecStart=/usr/local/bin/prometheus \
  --config.file=/etc/prometheus/prometheus.yml \
  --storage.tsdb.path=/var/lib/prometheus \
  --storage.tsdb.retention.time=30d \
  --storage.tsdb.retention.size=20GB \
  --web.listen-address=127.0.0.1:9090
ExecReload=/bin/kill -HUP $MAINPID
Restart=on-failure
NoNewPrivileges=true
ProtectSystem=full
ProtectHome=true

[Install]
WantedBy=multi-user.target

Start Prometheus:

sudo systemctl daemon-reload
sudo systemctl enable --now prometheus
curl -s http://127.0.0.1:9090/-/ready
Prometheus Server is Ready.

Check that every target is up. The up metric is 1 for each target whose last scrape succeeded, so this query should return the number of targets you configured (5 in this example):

curl -s 'http://127.0.0.1:9090/api/v1/query?query=count(up==1)'
{"status":"success","data":{"resultType":"vector","result":[{"metric":{},"value":[1790336120.456,"5"]}]}}

You can also use the web UI. Because Prometheus listens on localhost only, open an SSH tunnel from your workstation:

ssh -L 9090:127.0.0.1:9090 -L 9093:127.0.0.1:9093 your_user@your_server_ip

With the tunnel open, browse to http://localhost:9090/targets and confirm every target is UP. The Alerts page lists the five rules, all Inactive. http://localhost:9093 shows the Alertmanager UI.

Step 6 - Installing Grafana

Add the Grafana APT repository and install Grafana OSS:

sudo apt install apt-transport-https wget
sudo mkdir -p /etc/apt/keyrings
sudo wget -O /etc/apt/keyrings/grafana.asc https://apt.grafana.com/gpg-full.key
sudo chmod 644 /etc/apt/keyrings/grafana.asc
echo "deb [signed-by=/etc/apt/keyrings/grafana.asc] https://apt.grafana.com stable main" | sudo tee /etc/apt/sources.list.d/grafana.list
sudo apt update
sudo apt install grafana

Provision Prometheus and Alertmanager as data sources so they exist the moment Grafana starts:

sudo nano /etc/grafana/provisioning/datasources/stack.yaml
apiVersion: 1

datasources:
  - name: Prometheus
    uid: prometheus
    type: prometheus
    access: proxy
    url: http://127.0.0.1:9090
    isDefault: true
    jsonData:
      timeInterval: 15s

  - name: Alertmanager
    uid: alertmanager
    type: alertmanager
    access: proxy
    url: http://127.0.0.1:9093
    jsonData:
      implementation: prometheus

Start Grafana and allow port 3000 from your workstation only:

sudo systemctl daemon-reload
sudo systemctl enable --now grafana-server
sudo ufw allow from your_admin_ip to any port 3000 proto tcp

Browse to http://your_server_ip:3000, log in with admin / admin, and set a new password when prompted. Go to Connections > Data sources, open Prometheus and click Test; you should see Successfully queried the Prometheus API.

Step 7 - Importing dashboards

Instead of building host dashboards from scratch, import the community standard:

  1. Go to Dashboards > New > Import.
  2. Enter 1860 (Node Exporter Full) and click Load.
  3. Select the Prometheus data source and click Import.

Use the Host drop-down at the top to switch between monitoring-01, web-01 and db-01. The CPU, memory, disk and network panels should all show data within a minute.

Because Alertmanager is also a data source, Alerting > Alert rules in Grafana lists the Prometheus rules, and Alerting > Silences (with the Alertmanager data source selected) lets you create silences for planned maintenance without leaving Grafana.

Step 8 - Testing the alerting pipeline

Test the path in two stages. First, confirm that Alertmanager can deliver email by injecting a fake alert with amtool:

amtool alert add alertname=TestAlert severity=warning instance=test-host \
  --annotation=summary="Test alert from amtool" \
  --alertmanager.url=http://127.0.0.1:9093

After the 30-second group_wait, the message arrives at the address in alertmanager.yml. If it does not, the Alertmanager log shows the SMTP error:

sudo journalctl -u alertmanager -n 30 --no-pager

Next, test a real alert end to end. On one of the monitored hosts, stop Node Exporter:

sudo systemctl stop prometheus-node-exporter

Within about 15 seconds the target turns DOWN in Prometheus and InstanceDown becomes Pending. After the 2-minute for period it turns Firing, and you can list it from the monitoring server:

amtool alert query --alertmanager.url=http://127.0.0.1:9093
Alertname     Starts At                Summary
InstanceDown  2026-09-25 11:32:10 UTC  10.0.0.11:9100 is down

The email follows. Start the exporter again and, because send_resolved is true, a RESOLVED email arrives a few minutes later:

sudo systemctl start prometheus-node-exporter

Step 9 - Adding more hosts later

To monitor a new server, install prometheus-node-exporter on it and open port 9100 to the monitoring server as in Step 1. Then add its target to the node job in /etc/prometheus/prometheus.yml and reload Prometheus without a restart:

promtool check config /etc/prometheus/prometheus.yml && sudo systemctl reload prometheus

With many hosts, editing prometheus.yml becomes tedious. Move the targets to a file and use file-based service discovery instead: replace static_configs in the node job with:

    file_sd_configs:
      - files:
          - /etc/prometheus/targets/*.yml

Each file in /etc/prometheus/targets/ holds a list of targets with labels, in the same shape as a static_configs entry, and Prometheus picks up changes automatically without a reload.

Troubleshooting

Alerts show as firing in Prometheus but no email arrives. Check http://localhost:9090/status (through the tunnel) under Alertmanagers for 127.0.0.1:9093, then read sudo journalctl -u alertmanager for SMTP authentication or TLS errors. With port 587, keep smtp_require_tls: true and make sure the username and password are correct.

A remote target is DOWN with context deadline exceeded. The monitoring server cannot reach port 9100. Check the UFW rule on the target and any provider firewall between the two hosts.

HostDiskAlmostFull fires for /boot/efi or snap mounts. Exclude those mount points by adding mountpoint!~"/boot/efi|/snap/.*" to the selector in the rule.

Grafana panels show No data. The Prometheus data source URL must be http://127.0.0.1:9090, since Prometheus listens only on localhost. Test the data source and check that the node job is UP.

Conclusion

You now have a working monitoring stack: Node Exporter on every host, Prometheus collecting and evaluating alert rules, Alertmanager sending grouped email notifications with resolution messages, and Grafana dashboards for day-to-day visibility. From here, add exporters for your services (Blackbox Exporter for HTTP and TLS checks, MySQL or PostgreSQL exporters), route critical alerts to Slack or PagerDuty with a second Alertmanager receiver, and back up /etc/prometheus, /etc/alertmanager and /var/lib/grafana regularly.