Blackbox Exporter is the official Prometheus exporter for black-box monitoring: instead of reading metrics from inside a service, it probes the service from the outside over HTTP, HTTPS, TCP, ICMP or DNS, the same way a user would. It tells you whether an endpoint answers, how long it took, which status code it returned and when its TLS certificate expires. In this tutorial you will install Blackbox Exporter on Ubuntu 24.04 next to Prometheus, define probe modules, scrape them with Prometheus and add alerts for downtime and certificate expiry.

Prerequisites

To follow this tutorial, you will need:

  • A server running Ubuntu 24.04 LTS, such as a CubePath VPS, with a non-root user with sudo privileges.
  • Prometheus installed on the same server with its configuration in /etc/prometheus/prometheus.yml. The Ubuntu package (sudo apt install prometheus) is enough for this guide.
  • One or more endpoints to probe, for example https://your_domain. The examples also use public targets so you can test without your own site.

Blackbox Exporter only needs outbound access to the targets. It will listen on 127.0.0.1:9115, so no inbound firewall rule is required.

Step 1 - Installing Blackbox Exporter

Ubuntu packages an older release of the exporter, so this guide uses the official binary from the project's GitHub releases. Check the latest version on the releases page and set it in a shell variable:

VERSION=0.27.0
cd /tmp
curl -LO "https://github.com/prometheus/blackbox_exporter/releases/download/v${VERSION}/blackbox_exporter-${VERSION}.linux-amd64.tar.gz"
curl -LO "https://github.com/prometheus/blackbox_exporter/releases/download/v${VERSION}/sha256sums.txt"

On an ARM server, replace linux-amd64 with linux-arm64. Verify the checksum before extracting the archive:

sha256sum -c --ignore-missing sha256sums.txt
blackbox_exporter-0.27.0.linux-amd64.tar.gz: OK

Extract the archive and install the binary:

tar xzf "blackbox_exporter-${VERSION}.linux-amd64.tar.gz"
sudo install -m 0755 "blackbox_exporter-${VERSION}.linux-amd64/blackbox_exporter" /usr/local/bin/

Create a dedicated system user without a login shell to run the service:

sudo useradd --system --no-create-home --shell /usr/sbin/nologin blackbox

Confirm the binary works:

blackbox_exporter --version
blackbox_exporter, version 0.27.0 (branch: HEAD, revision: ...)

Step 2 - Defining probe modules

A module describes how to probe a target: which prober to use, the timeout and what counts as success. Prometheus later chooses a module per scrape job and passes the target as a URL parameter.

Create the configuration directory and file:

sudo mkdir -p /etc/blackbox_exporter
sudo nano /etc/blackbox_exporter/blackbox.yml

Add the following modules:

modules:
  http_2xx:
    prober: http
    timeout: 5s
    http:
      method: GET
      preferred_ip_protocol: ip4
      follow_redirects: true

  http_keyword:
    prober: http
    timeout: 5s
    http:
      preferred_ip_protocol: ip4
      fail_if_body_not_matches_regexp:
        - "Example Domain"

  tcp_connect:
    prober: tcp
    timeout: 5s

  icmp:
    prober: icmp
    timeout: 5s
    icmp:
      preferred_ip_protocol: ip4

  dns_a:
    prober: dns
    timeout: 5s
    dns:
      query_name: example.com
      query_type: A
      valid_rcodes:
        - NOERROR

What each module does:

  • http_2xx succeeds when the URL returns a 2xx status code after following redirects. For HTTPS targets it also records the certificate expiry date.
  • http_keyword also requires the response body to match a regular expression, which catches error pages served with a 200 status. Change Example Domain to a string that only appears on a healthy page of your site.
  • tcp_connect succeeds when a TCP connection opens, useful for SSH, databases or mail ports.
  • icmp sends a ping. It needs the CAP_NET_RAW capability, which the systemd unit in the next step grants.
  • dns_a asks the target DNS server for the A record of example.com and expects a NOERROR answer. Replace the name with your own domain.

Set the ownership and validate the file. The --config.check flag parses the configuration and exits:

sudo chown -R root:blackbox /etc/blackbox_exporter
sudo chmod 0750 /etc/blackbox_exporter
sudo chmod 0640 /etc/blackbox_exporter/blackbox.yml
blackbox_exporter --config.check --config.file=/etc/blackbox_exporter/blackbox.yml

A valid file produces a line containing Config file is ok exiting.... Any YAML or key error is reported with its line number.

Step 3 - Running Blackbox Exporter with systemd

Create a unit file so the exporter starts on boot and restarts if it fails:

sudo nano /etc/systemd/system/blackbox_exporter.service
[Unit]
Description=Prometheus Blackbox Exporter
Wants=network-online.target
After=network-online.target

[Service]
User=blackbox
Group=blackbox
ExecStart=/usr/local/bin/blackbox_exporter \
  --config.file=/etc/blackbox_exporter/blackbox.yml \
  --web.listen-address=127.0.0.1:9115
ExecReload=/bin/kill -HUP $MAINPID
AmbientCapabilities=CAP_NET_RAW
CapabilityBoundingSet=CAP_NET_RAW
NoNewPrivileges=true
Restart=on-failure

[Install]
WantedBy=multi-user.target

AmbientCapabilities=CAP_NET_RAW lets the unprivileged blackbox user open raw sockets for ICMP without running as root. The listen address keeps the exporter on localhost, because anyone who can reach /probe could use your server to probe arbitrary hosts.

Reload systemd and start the service:

sudo systemctl daemon-reload
sudo systemctl enable --now blackbox_exporter
sudo systemctl status blackbox_exporter
● blackbox_exporter.service - Prometheus Blackbox Exporter
     Loaded: loaded (/etc/systemd/system/blackbox_exporter.service; enabled; preset: enabled)
     Active: active (running) since ...

If the status is failed, read the log with sudo journalctl -u blackbox_exporter -n 50.

Step 4 - Testing probes manually

Before involving Prometheus, call the /probe endpoint directly. The target and module parameters are exactly what Prometheus will send:

curl -s "http://127.0.0.1:9115/probe?target=https://example.com&module=http_2xx" | grep -E '^probe_(success|http_status_code|duration_seconds|ssl_earliest_cert_expiry) '
probe_duration_seconds 0.183412
probe_http_status_code 200
probe_ssl_earliest_cert_expiry 1.7704224e+09
probe_success 1

probe_success 1 means the probe passed. probe_ssl_earliest_cert_expiry is a Unix timestamp of the first certificate in the chain to expire. Test the other modules the same way:

curl -s "http://127.0.0.1:9115/probe?target=example.com:443&module=tcp_connect" | grep '^probe_success'
curl -s "http://127.0.0.1:9115/probe?target=1.1.1.1&module=icmp" | grep '^probe_success'
curl -s "http://127.0.0.1:9115/probe?target=1.1.1.1:53&module=dns_a" | grep '^probe_success'

Each should print probe_success 1. When a probe fails, add &debug=true to the URL: the exporter returns a step-by-step log of the probe (DNS resolution, connection, TLS handshake, status code) instead of metrics, which usually shows the cause immediately.

Step 5 - Scraping probes from Prometheus

Prometheus does not scrape the targets directly. Each scrape job sends requests to Blackbox Exporter and uses relabeling to move the real target into the target URL parameter and the instance label.

Open the Prometheus configuration:

sudo nano /etc/prometheus/prometheus.yml

Add these jobs under the existing scrape_configs: key, keeping the indentation of the other jobs:

  - job_name: blackbox_http
    metrics_path: /probe
    params:
      module: [http_2xx]
    static_configs:
      - targets:
          - https://example.com
          - https://your_domain
    relabel_configs:
      - source_labels: [__address__]
        target_label: __param_target
      - source_labels: [__param_target]
        target_label: instance
      - target_label: __address__
        replacement: 127.0.0.1:9115

  - job_name: blackbox_icmp
    metrics_path: /probe
    params:
      module: [icmp]
    static_configs:
      - targets:
          - 1.1.1.1
          - your_server_ip
    relabel_configs:
      - source_labels: [__address__]
        target_label: __param_target
      - source_labels: [__param_target]
        target_label: instance
      - target_label: __address__
        replacement: 127.0.0.1:9115

  - job_name: blackbox_exporter
    static_configs:
      - targets: ['127.0.0.1:9115']

The three relabel rules copy the listed target into __param_target (the ?target= parameter), keep it as the instance label so graphs show the real endpoint, and finally point the scrape at the exporter. The last job scrapes the exporter's own metrics, such as its configuration reload status.

Validate the configuration and restart Prometheus:

promtool check config /etc/prometheus/prometheus.yml
sudo systemctl restart prometheus
Checking /etc/prometheus/prometheus.yml
 SUCCESS: /etc/prometheus/prometheus.yml is valid prometheus config file syntax

After one scrape interval, query the results from the Prometheus API:

curl -s 'http://localhost:9090/api/v1/query' --data-urlencode 'query=probe_success{job="blackbox_http"}'

The response contains one result per target, with "instance":"https://example.com" and a value of "1" when the probe succeeds.

You can also open the Prometheus web interface, go to Status > Targets and confirm that the blackbox_http and blackbox_icmp targets are UP. A target shows as UP even when the probe fails, because the scrape of the exporter itself worked: always alert on probe_success, not on up.

Step 6 - Adding alerting rules

Create a directory for rule files and a rules file for the probes:

sudo mkdir -p /etc/prometheus/rules
sudo nano /etc/prometheus/rules/blackbox.yml
groups:
  - name: blackbox
    rules:
      - alert: EndpointDown
        expr: probe_success == 0
        for: 2m
        labels:
          severity: critical
        annotations:
          summary: "{{ $labels.instance }} is not responding"
          description: "The {{ $labels.job }} probe for {{ $labels.instance }} has failed for more than 2 minutes."

      - alert: EndpointSlow
        expr: avg_over_time(probe_duration_seconds[5m]) > 2
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "{{ $labels.instance }} is slow"
          description: "Average probe duration has been above 2 seconds for 5 minutes."

      - alert: TLSCertificateExpiringSoon
        expr: probe_ssl_earliest_cert_expiry - time() < 14 * 86400
        for: 1h
        labels:
          severity: warning
        annotations:
          summary: "TLS certificate for {{ $labels.instance }} expires soon"
          description: "The certificate expires in {{ $value | humanizeDuration }}."

The for clause prevents a single failed probe from firing an alert. Fourteen days before expiry gives you time to react if automatic renewal (for example with Certbot) is broken.

Tell Prometheus to load the file. In /etc/prometheus/prometheus.yml, set the top-level rule_files key (replace the commented example if your file has one):

rule_files:
  - /etc/prometheus/rules/*.yml

Check the rules and the configuration, then restart Prometheus:

promtool check rules /etc/prometheus/rules/blackbox.yml
promtool check config /etc/prometheus/prometheus.yml
sudo systemctl restart prometheus
Checking /etc/prometheus/rules/blackbox.yml
  SUCCESS: 3 rules found

The rules now appear under Alerts in the Prometheus web interface. To send notifications by email or chat, point Prometheus at an Alertmanager instance or create equivalent alert rules in Grafana.

To test the EndpointDown alert, add a target that does not exist, such as https://does-not-exist.invalid, to the blackbox_http job and restart Prometheus. The alert goes to Pending on the next evaluation and to Firing two minutes later.

Troubleshooting

  • ICMP probes fail with permission denied or socket: operation not permitted: the service is missing the raw socket capability. Confirm the AmbientCapabilities=CAP_NET_RAW line is in the unit, then run sudo systemctl daemon-reload and sudo systemctl restart blackbox_exporter.
  • HTTP probes fail but curl from the server works: add &debug=true to the probe URL. Common causes are an IPv6 address being tried first (keep preferred_ip_protocol: ip4 if the host has no IPv6 route) and redirects to another host that fails.
  • context deadline exceeded: the module timeout is longer than the Prometheus scrape_timeout (10 seconds by default) or the target is slow. Keep the module timeout below the scrape timeout.
  • Configuration changes are ignored: after editing blackbox.yml, run sudo systemctl reload blackbox_exporter. The exporter re-reads the file on SIGHUP and logs an error if it is invalid.

Conclusion

Blackbox Exporter now probes your endpoints over HTTP, ICMP, TCP and DNS, Prometheus stores the results, and alert rules catch outages, slow responses and certificates close to expiry. Because the probes run from outside the application, they detect problems that internal metrics miss, such as DNS or TLS errors.

As next steps, run a second Blackbox Exporter from a different location so you can tell a real outage from a network problem at one site, send the alerts through Alertmanager or Grafana alerting, and build a Grafana dashboard on probe_success and probe_duration_seconds for an availability overview.