Blackbox Exporter is the official Prometheus exporter for black-box monitoring: instead of reading metrics from inside a service, it probes the service from the outside over HTTP, HTTPS, TCP, ICMP or DNS, the same way a user would. It tells you whether an endpoint answers, how long it took, which status code it returned and when its TLS certificate expires. In this tutorial you will install Blackbox Exporter on Ubuntu 24.04 next to Prometheus, define probe modules, scrape them with Prometheus and add alerts for downtime and certificate expiry.
Prerequisites
To follow this tutorial, you will need:
- A server running Ubuntu 24.04 LTS, such as a CubePath VPS, with a non-root user with
sudoprivileges. - Prometheus installed on the same server with its configuration in
/etc/prometheus/prometheus.yml. The Ubuntu package (sudo apt install prometheus) is enough for this guide. - One or more endpoints to probe, for example
https://your_domain. The examples also use public targets so you can test without your own site.
Blackbox Exporter only needs outbound access to the targets. It will listen on 127.0.0.1:9115, so no inbound firewall rule is required.
Step 1 - Installing Blackbox Exporter
Ubuntu packages an older release of the exporter, so this guide uses the official binary from the project's GitHub releases. Check the latest version on the releases page and set it in a shell variable:
VERSION=0.27.0
cd /tmp
curl -LO "https://github.com/prometheus/blackbox_exporter/releases/download/v${VERSION}/blackbox_exporter-${VERSION}.linux-amd64.tar.gz"
curl -LO "https://github.com/prometheus/blackbox_exporter/releases/download/v${VERSION}/sha256sums.txt"
On an ARM server, replace linux-amd64 with linux-arm64. Verify the checksum before extracting the archive:
sha256sum -c --ignore-missing sha256sums.txt
blackbox_exporter-0.27.0.linux-amd64.tar.gz: OK
Extract the archive and install the binary:
tar xzf "blackbox_exporter-${VERSION}.linux-amd64.tar.gz"
sudo install -m 0755 "blackbox_exporter-${VERSION}.linux-amd64/blackbox_exporter" /usr/local/bin/
Create a dedicated system user without a login shell to run the service:
sudo useradd --system --no-create-home --shell /usr/sbin/nologin blackbox
Confirm the binary works:
blackbox_exporter --version
blackbox_exporter, version 0.27.0 (branch: HEAD, revision: ...)
Step 2 - Defining probe modules
A module describes how to probe a target: which prober to use, the timeout and what counts as success. Prometheus later chooses a module per scrape job and passes the target as a URL parameter.
Create the configuration directory and file:
sudo mkdir -p /etc/blackbox_exporter
sudo nano /etc/blackbox_exporter/blackbox.yml
Add the following modules:
modules:
http_2xx:
prober: http
timeout: 5s
http:
method: GET
preferred_ip_protocol: ip4
follow_redirects: true
http_keyword:
prober: http
timeout: 5s
http:
preferred_ip_protocol: ip4
fail_if_body_not_matches_regexp:
- "Example Domain"
tcp_connect:
prober: tcp
timeout: 5s
icmp:
prober: icmp
timeout: 5s
icmp:
preferred_ip_protocol: ip4
dns_a:
prober: dns
timeout: 5s
dns:
query_name: example.com
query_type: A
valid_rcodes:
- NOERROR
What each module does:
http_2xxsucceeds when the URL returns a 2xx status code after following redirects. For HTTPS targets it also records the certificate expiry date.http_keywordalso requires the response body to match a regular expression, which catches error pages served with a 200 status. ChangeExample Domainto a string that only appears on a healthy page of your site.tcp_connectsucceeds when a TCP connection opens, useful for SSH, databases or mail ports.icmpsends a ping. It needs theCAP_NET_RAWcapability, which the systemd unit in the next step grants.dns_aasks the target DNS server for the A record ofexample.comand expects aNOERRORanswer. Replace the name with your own domain.
Set the ownership and validate the file. The --config.check flag parses the configuration and exits:
sudo chown -R root:blackbox /etc/blackbox_exporter
sudo chmod 0750 /etc/blackbox_exporter
sudo chmod 0640 /etc/blackbox_exporter/blackbox.yml
blackbox_exporter --config.check --config.file=/etc/blackbox_exporter/blackbox.yml
A valid file produces a line containing Config file is ok exiting.... Any YAML or key error is reported with its line number.
Step 3 - Running Blackbox Exporter with systemd
Create a unit file so the exporter starts on boot and restarts if it fails:
sudo nano /etc/systemd/system/blackbox_exporter.service
[Unit]
Description=Prometheus Blackbox Exporter
Wants=network-online.target
After=network-online.target
[Service]
User=blackbox
Group=blackbox
ExecStart=/usr/local/bin/blackbox_exporter \
--config.file=/etc/blackbox_exporter/blackbox.yml \
--web.listen-address=127.0.0.1:9115
ExecReload=/bin/kill -HUP $MAINPID
AmbientCapabilities=CAP_NET_RAW
CapabilityBoundingSet=CAP_NET_RAW
NoNewPrivileges=true
Restart=on-failure
[Install]
WantedBy=multi-user.target
AmbientCapabilities=CAP_NET_RAW lets the unprivileged blackbox user open raw sockets for ICMP without running as root. The listen address keeps the exporter on localhost, because anyone who can reach /probe could use your server to probe arbitrary hosts.
Reload systemd and start the service:
sudo systemctl daemon-reload
sudo systemctl enable --now blackbox_exporter
sudo systemctl status blackbox_exporter
● blackbox_exporter.service - Prometheus Blackbox Exporter
Loaded: loaded (/etc/systemd/system/blackbox_exporter.service; enabled; preset: enabled)
Active: active (running) since ...
If the status is failed, read the log with sudo journalctl -u blackbox_exporter -n 50.
Step 4 - Testing probes manually
Before involving Prometheus, call the /probe endpoint directly. The target and module parameters are exactly what Prometheus will send:
curl -s "http://127.0.0.1:9115/probe?target=https://example.com&module=http_2xx" | grep -E '^probe_(success|http_status_code|duration_seconds|ssl_earliest_cert_expiry) '
probe_duration_seconds 0.183412
probe_http_status_code 200
probe_ssl_earliest_cert_expiry 1.7704224e+09
probe_success 1
probe_success 1 means the probe passed. probe_ssl_earliest_cert_expiry is a Unix timestamp of the first certificate in the chain to expire. Test the other modules the same way:
curl -s "http://127.0.0.1:9115/probe?target=example.com:443&module=tcp_connect" | grep '^probe_success'
curl -s "http://127.0.0.1:9115/probe?target=1.1.1.1&module=icmp" | grep '^probe_success'
curl -s "http://127.0.0.1:9115/probe?target=1.1.1.1:53&module=dns_a" | grep '^probe_success'
Each should print probe_success 1. When a probe fails, add &debug=true to the URL: the exporter returns a step-by-step log of the probe (DNS resolution, connection, TLS handshake, status code) instead of metrics, which usually shows the cause immediately.
Step 5 - Scraping probes from Prometheus
Prometheus does not scrape the targets directly. Each scrape job sends requests to Blackbox Exporter and uses relabeling to move the real target into the target URL parameter and the instance label.
Open the Prometheus configuration:
sudo nano /etc/prometheus/prometheus.yml
Add these jobs under the existing scrape_configs: key, keeping the indentation of the other jobs:
- job_name: blackbox_http
metrics_path: /probe
params:
module: [http_2xx]
static_configs:
- targets:
- https://example.com
- https://your_domain
relabel_configs:
- source_labels: [__address__]
target_label: __param_target
- source_labels: [__param_target]
target_label: instance
- target_label: __address__
replacement: 127.0.0.1:9115
- job_name: blackbox_icmp
metrics_path: /probe
params:
module: [icmp]
static_configs:
- targets:
- 1.1.1.1
- your_server_ip
relabel_configs:
- source_labels: [__address__]
target_label: __param_target
- source_labels: [__param_target]
target_label: instance
- target_label: __address__
replacement: 127.0.0.1:9115
- job_name: blackbox_exporter
static_configs:
- targets: ['127.0.0.1:9115']
The three relabel rules copy the listed target into __param_target (the ?target= parameter), keep it as the instance label so graphs show the real endpoint, and finally point the scrape at the exporter. The last job scrapes the exporter's own metrics, such as its configuration reload status.
Validate the configuration and restart Prometheus:
promtool check config /etc/prometheus/prometheus.yml
sudo systemctl restart prometheus
Checking /etc/prometheus/prometheus.yml
SUCCESS: /etc/prometheus/prometheus.yml is valid prometheus config file syntax
After one scrape interval, query the results from the Prometheus API:
curl -s 'http://localhost:9090/api/v1/query' --data-urlencode 'query=probe_success{job="blackbox_http"}'
The response contains one result per target, with "instance":"https://example.com" and a value of "1" when the probe succeeds.
You can also open the Prometheus web interface, go to Status > Targets and confirm that the blackbox_http and blackbox_icmp targets are UP. A target shows as UP even when the probe fails, because the scrape of the exporter itself worked: always alert on probe_success, not on up.
Step 6 - Adding alerting rules
Create a directory for rule files and a rules file for the probes:
sudo mkdir -p /etc/prometheus/rules
sudo nano /etc/prometheus/rules/blackbox.yml
groups:
- name: blackbox
rules:
- alert: EndpointDown
expr: probe_success == 0
for: 2m
labels:
severity: critical
annotations:
summary: "{{ $labels.instance }} is not responding"
description: "The {{ $labels.job }} probe for {{ $labels.instance }} has failed for more than 2 minutes."
- alert: EndpointSlow
expr: avg_over_time(probe_duration_seconds[5m]) > 2
for: 5m
labels:
severity: warning
annotations:
summary: "{{ $labels.instance }} is slow"
description: "Average probe duration has been above 2 seconds for 5 minutes."
- alert: TLSCertificateExpiringSoon
expr: probe_ssl_earliest_cert_expiry - time() < 14 * 86400
for: 1h
labels:
severity: warning
annotations:
summary: "TLS certificate for {{ $labels.instance }} expires soon"
description: "The certificate expires in {{ $value | humanizeDuration }}."
The for clause prevents a single failed probe from firing an alert. Fourteen days before expiry gives you time to react if automatic renewal (for example with Certbot) is broken.
Tell Prometheus to load the file. In /etc/prometheus/prometheus.yml, set the top-level rule_files key (replace the commented example if your file has one):
rule_files:
- /etc/prometheus/rules/*.yml
Check the rules and the configuration, then restart Prometheus:
promtool check rules /etc/prometheus/rules/blackbox.yml
promtool check config /etc/prometheus/prometheus.yml
sudo systemctl restart prometheus
Checking /etc/prometheus/rules/blackbox.yml
SUCCESS: 3 rules found
The rules now appear under Alerts in the Prometheus web interface. To send notifications by email or chat, point Prometheus at an Alertmanager instance or create equivalent alert rules in Grafana.
To test the EndpointDown alert, add a target that does not exist, such as https://does-not-exist.invalid, to the blackbox_http job and restart Prometheus. The alert goes to Pending on the next evaluation and to Firing two minutes later.
Troubleshooting
- ICMP probes fail with
permission deniedorsocket: operation not permitted: the service is missing the raw socket capability. Confirm theAmbientCapabilities=CAP_NET_RAWline is in the unit, then runsudo systemctl daemon-reloadandsudo systemctl restart blackbox_exporter. - HTTP probes fail but
curlfrom the server works: add&debug=trueto the probe URL. Common causes are an IPv6 address being tried first (keeppreferred_ip_protocol: ip4if the host has no IPv6 route) and redirects to another host that fails. context deadline exceeded: the module timeout is longer than the Prometheusscrape_timeout(10 seconds by default) or the target is slow. Keep the module timeout below the scrape timeout.- Configuration changes are ignored: after editing
blackbox.yml, runsudo systemctl reload blackbox_exporter. The exporter re-reads the file onSIGHUPand logs an error if it is invalid.
Conclusion
Blackbox Exporter now probes your endpoints over HTTP, ICMP, TCP and DNS, Prometheus stores the results, and alert rules catch outages, slow responses and certificates close to expiry. Because the probes run from outside the application, they detect problems that internal metrics miss, such as DNS or TLS errors.
As next steps, run a second Blackbox Exporter from a different location so you can tell a real outage from a network problem at one site, send the alerts through Alertmanager or Grafana alerting, and build a Grafana dashboard on probe_success and probe_duration_seconds for an availability overview.
