When no existing exporter covers the system you want to monitor, you can write your own with one of the official Prometheus client libraries. An exporter is just a small HTTP server that gathers values when Prometheus scrapes it and returns them in the Prometheus text format. In this tutorial you will build a real exporter that reports the state of a backup directory (number of files, age and size of the newest backup), first in Python and then in Go, run it as a systemd service on Ubuntu 24.04 and alert when backups stop arriving.

Prerequisites

To follow this tutorial, you will need:

  • A server running Ubuntu 24.04 LTS, such as a CubePath VPS, with a non-root user with sudo privileges.
  • Prometheus installed and configured in /etc/prometheus/prometheus.yml.
  • For the Go section only: Go 1.23 or later installed from go.dev/dl.

Step 1 - Choosing metric names and types

Design the metrics before writing code. Prometheus has four metric types:

TypeUse it forExample
CounterValues that only go up and reset on restarthttp_requests_total
GaugeValues that go up and downqueue_length, temperature_celsius
HistogramDistributions bucketed on the server side, aggregatablerequest_duration_seconds
SummaryClient-side quantiles, not aggregatable across instancesrarely the best choice

An exporter that reads the state of another system usually produces gauges, because it reports a current value rather than counting events itself. For the backup directory, the exporter will expose:

  • backup_dir_up: 1 if the directory could be read, 0 otherwise.
  • backup_files: number of regular files in the directory.
  • backup_last_modified_timestamp_seconds: modification time of the newest file, as a Unix timestamp.
  • backup_last_size_bytes: size of the newest file.

These names follow the Prometheus naming conventions: a common prefix for the exporter, base units (seconds, bytes) as suffixes and _total reserved for counters. Exposing a timestamp instead of an "age" lets PromQL compute the age with time() - backup_last_modified_timestamp_seconds, which stays correct even if a scrape is delayed.

Create a test directory with a sample file so you have something to measure:

sudo mkdir -p /var/backups/app
sudo dd if=/dev/zero of=/var/backups/app/db-test.sql.gz bs=1M count=5

Step 2 - Writing the exporter in Python

The Python client library supports custom collectors: classes with a collect() method that the library calls on every scrape. This is the right pattern for exporters, because the values are read fresh at scrape time and nothing runs between scrapes.

Create a directory for the exporter and a virtual environment with the prometheus-client package:

sudo apt update
sudo apt install -y python3-venv
sudo mkdir -p /opt/backup-exporter
sudo python3 -m venv /opt/backup-exporter/venv
sudo /opt/backup-exporter/venv/bin/pip install prometheus-client

Create the exporter script:

sudo nano /opt/backup-exporter/backup_exporter.py
#!/usr/bin/env python3
"""Prometheus exporter that reports the state of a backup directory."""

import argparse
import os
import time

from prometheus_client import start_http_server
from prometheus_client.core import REGISTRY, GaugeMetricFamily


class BackupCollector:
    def __init__(self, directory):
        self.directory = directory

    def collect(self):
        up = GaugeMetricFamily(
            "backup_dir_up",
            "1 if the backup directory could be read, 0 otherwise.",
        )
        try:
            files = [
                entry for entry in os.scandir(self.directory)
                if entry.is_file(follow_symlinks=False)
            ]
        except OSError:
            up.add_metric([], 0)
            yield up
            return

        up.add_metric([], 1)
        yield up

        yield GaugeMetricFamily(
            "backup_files",
            "Number of files in the backup directory.",
            value=len(files),
        )

        if files:
            newest = max(files, key=lambda entry: entry.stat().st_mtime).stat()
            yield GaugeMetricFamily(
                "backup_last_modified_timestamp_seconds",
                "Modification time of the newest backup file.",
                value=newest.st_mtime,
            )
            yield GaugeMetricFamily(
                "backup_last_size_bytes",
                "Size of the newest backup file in bytes.",
                value=newest.st_size,
            )


def main():
    parser = argparse.ArgumentParser(description=__doc__)
    parser.add_argument("--directory", default="/var/backups/app")
    parser.add_argument("--listen-address", default="127.0.0.1")
    parser.add_argument("--port", type=int, default=9900)
    args = parser.parse_args()

    REGISTRY.register(BackupCollector(args.directory))
    start_http_server(args.port, addr=args.listen_address)

    # start_http_server serves metrics from a background thread.
    while True:
        time.sleep(3600)


if __name__ == "__main__":
    main()

Key points of this code:

  • collect() is a generator that yields metric families. If the directory cannot be read, it still returns backup_dir_up 0 instead of failing the scrape, so you can alert on it.
  • Metrics that do not apply (there is no newest file in an empty directory) are left out instead of being reported as 0, which would look like a backup from 1970.
  • The default registry also exports process_* and python_* metrics about the exporter itself.

Run it in the foreground to test it:

sudo /opt/backup-exporter/venv/bin/python /opt/backup-exporter/backup_exporter.py

In a second terminal, request the metrics:

curl -s http://127.0.0.1:9900/metrics | grep '^backup_'
backup_dir_up 1.0
backup_files 1.0
backup_last_modified_timestamp_seconds 1.7587872e+09
backup_last_size_bytes 5.24288e+06

Press CTRL+C in the first terminal to stop the exporter.

Step 3 - Running the exporter as a systemd service

Create an unprivileged user for the exporter. It only needs read access to the backup directory:

sudo useradd --system --no-create-home --shell /usr/sbin/nologin backup_exporter

If the directory is not world-readable, grant read access with a group or an ACL instead of running the exporter as root, for example sudo setfacl -m u:backup_exporter:rx /var/backups/app (install acl if the command is missing).

Create the unit file:

sudo nano /etc/systemd/system/backup_exporter.service
[Unit]
Description=Prometheus backup directory exporter
After=network-online.target
Wants=network-online.target

[Service]
User=backup_exporter
Group=backup_exporter
ExecStart=/opt/backup-exporter/venv/bin/python /opt/backup-exporter/backup_exporter.py \
  --directory /var/backups/app --port 9900
Restart=on-failure
NoNewPrivileges=true
ProtectSystem=strict
ProtectHome=true

[Install]
WantedBy=multi-user.target

ProtectSystem=strict mounts the file system read-only for the service, which is fine because the exporter never writes anything. Start the service:

sudo systemctl daemon-reload
sudo systemctl enable --now backup_exporter
sudo systemctl status backup_exporter
● backup_exporter.service - Prometheus backup directory exporter
     Loaded: loaded (/etc/systemd/system/backup_exporter.service; enabled; preset: enabled)
     Active: active (running) since ...

Check again with curl -s http://127.0.0.1:9900/metrics | grep '^backup_'. If backup_dir_up is 0, the service user cannot read the directory.

Step 4 - Writing the same exporter in Go

Go produces a single static binary with no runtime dependencies, which is convenient when you deploy the exporter to many servers. The official client_golang library uses the same collector pattern: Describe announces the metrics and Collect produces their values at scrape time.

Create a module:

mkdir -p ~/backup-exporter && cd ~/backup-exporter
go mod init example.com/backup-exporter
nano main.go
package main

import (
	"flag"
	"log"
	"net/http"
	"os"
	"time"

	"github.com/prometheus/client_golang/prometheus"
	"github.com/prometheus/client_golang/prometheus/collectors"
	"github.com/prometheus/client_golang/prometheus/promhttp"
)

type backupCollector struct {
	dir      string
	up       *prometheus.Desc
	files    *prometheus.Desc
	lastMod  *prometheus.Desc
	lastSize *prometheus.Desc
}

func newBackupCollector(dir string) *backupCollector {
	return &backupCollector{
		dir:      dir,
		up:       prometheus.NewDesc("backup_dir_up", "1 if the backup directory could be read, 0 otherwise.", nil, nil),
		files:    prometheus.NewDesc("backup_files", "Number of files in the backup directory.", nil, nil),
		lastMod:  prometheus.NewDesc("backup_last_modified_timestamp_seconds", "Modification time of the newest backup file.", nil, nil),
		lastSize: prometheus.NewDesc("backup_last_size_bytes", "Size of the newest backup file in bytes.", nil, nil),
	}
}

func (c *backupCollector) Describe(ch chan<- *prometheus.Desc) {
	ch <- c.up
	ch <- c.files
	ch <- c.lastMod
	ch <- c.lastSize
}

func (c *backupCollector) Collect(ch chan<- prometheus.Metric) {
	entries, err := os.ReadDir(c.dir)
	if err != nil {
		ch <- prometheus.MustNewConstMetric(c.up, prometheus.GaugeValue, 0)
		return
	}
	ch <- prometheus.MustNewConstMetric(c.up, prometheus.GaugeValue, 1)

	var count int
	var newest time.Time
	var newestSize int64
	for _, e := range entries {
		if !e.Type().IsRegular() {
			continue
		}
		info, err := e.Info()
		if err != nil {
			continue
		}
		count++
		if info.ModTime().After(newest) {
			newest = info.ModTime()
			newestSize = info.Size()
		}
	}

	ch <- prometheus.MustNewConstMetric(c.files, prometheus.GaugeValue, float64(count))
	if count > 0 {
		ch <- prometheus.MustNewConstMetric(c.lastMod, prometheus.GaugeValue, float64(newest.Unix()))
		ch <- prometheus.MustNewConstMetric(c.lastSize, prometheus.GaugeValue, float64(newestSize))
	}
}

func main() {
	dir := flag.String("directory", "/var/backups/app", "Backup directory to inspect.")
	addr := flag.String("listen-address", "127.0.0.1:9900", "Address to expose metrics on.")
	flag.Parse()

	reg := prometheus.NewRegistry()
	reg.MustRegister(
		collectors.NewGoCollector(),
		collectors.NewProcessCollector(collectors.ProcessCollectorOpts{}),
		newBackupCollector(*dir),
	)

	http.Handle("/metrics", promhttp.HandlerFor(reg, promhttp.HandlerOpts{}))
	log.Printf("listening on %s", *addr)
	log.Fatal(http.ListenAndServe(*addr, nil))
}

Using a dedicated registry instead of the global default keeps the exported metrics explicit. Download the dependencies and build the binary:

go mod tidy
CGO_ENABLED=0 go build -o backup_exporter .

Stop the Python service first so the port is free, then test the Go binary:

sudo systemctl stop backup_exporter
./backup_exporter &
curl -s http://127.0.0.1:9900/metrics | grep '^backup_'
kill %1
backup_dir_up 1
backup_files 1
backup_last_modified_timestamp_seconds 1.758787215e+09
backup_last_size_bytes 5.24288e+06

The output matches the Python version. To use the Go binary in production, copy it to /usr/local/bin/backup_exporter and change ExecStart in the unit to /usr/local/bin/backup_exporter --directory /var/backups/app. Otherwise, start the Python service again with sudo systemctl start backup_exporter.

Step 5 - Scraping the exporter and alerting on it

Add a job to /etc/prometheus/prometheus.yml under scrape_configs::

sudo nano /etc/prometheus/prometheus.yml
  - job_name: backup
    static_configs:
      - targets: ['127.0.0.1:9900']

The exporter listens on localhost because Prometheus runs on the same server here. If Prometheus runs elsewhere, start the exporter with --listen-address 0.0.0.0 and allow the port only from the Prometheus server with sudo ufw allow from prometheus_server_ip to any port 9900 proto tcp.

Create a rules file:

sudo mkdir -p /etc/prometheus/rules
sudo nano /etc/prometheus/rules/backup.yml
groups:
  - name: backup
    rules:
      - alert: BackupTooOld
        expr: time() - backup_last_modified_timestamp_seconds > 26 * 3600
        for: 10m
        labels:
          severity: critical
        annotations:
          summary: "No new backup on {{ $labels.instance }} for more than 26 hours"

      - alert: BackupDirectoryUnreadable
        expr: backup_dir_up == 0
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "Backup exporter on {{ $labels.instance }} cannot read the directory"

The 26-hour threshold allows a daily backup to run a little late without firing. Make sure rule_files in prometheus.yml includes /etc/prometheus/rules/*.yml, then validate and restart:

promtool check rules /etc/prometheus/rules/backup.yml
promtool check config /etc/prometheus/prometheus.yml
sudo systemctl restart prometheus

To test the alert without waiting a day, set the sample file's modification time to two days ago:

sudo touch -d '2 days ago' /var/backups/app/db-test.sql.gz

After the next scrape, the query time() - backup_last_modified_timestamp_seconds returns about 172800, and BackupTooOld goes to Pending and then Firing in the Prometheus Alerts page.

Step 6 - Checking the output with promtool

Prometheus parses the exposition format strictly. promtool can lint the output of your exporter for naming problems, such as a counter without _total or a missing HELP line:

curl -s http://127.0.0.1:9900/metrics | promtool check metrics

No output means no problems were found. Run this check whenever you add metrics.

Troubleshooting

  • OSError: [Errno 98] Address already in use: another process uses port 9900. Find it with sudo ss -ltnp | grep 9900 and pick another port.
  • Metrics missing after a code change: the Python process only reloads code on restart. Run sudo systemctl restart backup_exporter.
  • Scrapes are slow or time out: collect() runs on every scrape, so expensive work (scanning millions of files, calling slow APIs) delays the response. Cache the result in a background thread and return the cached values, or increase scrape_interval for that job.
  • Cardinality grows without limit: never put unbounded values such as file names, user IDs or full URLs in labels. Each label value creates a new time series.

Conclusion

You built a custom Prometheus exporter with the collector pattern in both Python and Go, ran it as an unprivileged systemd service, scraped it with Prometheus and alerted when backups stop arriving. The same structure works for any system you can query: read the values in collect(), expose them as gauges with clear names and base units, and let PromQL do the calculations.

For short-lived batch jobs that cannot be scraped, consider writing metrics to a file for the node_exporter textfile collector instead. For your own applications, instrument the code directly with counters and histograms rather than building a separate exporter.