Python is a good fit for server automation once a task outgrows a few lines of Bash: it has real data structures, clear error handling and libraries for almost everything. In this tutorial you will set up an isolated Python environment on Ubuntu 24.04 and write three practical scripts: a resource checker, a website and TLS certificate monitor, and an Nginx access log summarizer. You will then run the resource checker every five minutes with a systemd timer.

Prerequisites

To follow this tutorial you need:

  • A server running Ubuntu 24.04 LTS, for example a CubePath VPS. Ubuntu 24.04 ships Python 3.12.
  • A non-root user with sudo privileges.
  • For the log summarizer in Step 5, Nginx serving traffic with its default access log at /var/log/nginx/access.log. The other scripts work without it.
  • Basic Python knowledge: functions, lists and dictionaries.

Step 1 - Creating a virtual environment for your tools

Ubuntu 24.04 marks the system Python as "externally managed" (PEP 668), so sudo pip install is refused to protect packages that the operating system depends on. The correct approach is a virtual environment that belongs to your scripts.

Install the venv module:

sudo apt update
sudo apt install -y python3-venv

Create a directory for the tools and a virtual environment inside it:

sudo mkdir -p /opt/ops-tools
sudo python3 -m venv /opt/ops-tools/venv

Install the two third-party libraries the scripts use: psutil for system metrics and requests for HTTP checks:

sudo /opt/ops-tools/venv/bin/pip install psutil requests

Verify that the environment can import them:

/opt/ops-tools/venv/bin/python -c 'import psutil, requests; print(psutil.__version__, requests.__version__)'
7.0.0 2.32.4

Your version numbers may differ. Because each script calls the virtual environment's interpreter directly, you never need to activate the environment, which also makes them easy to run from systemd.

Step 2 - Writing a resource checker

The first script reports memory usage, disk usage for each mount point and load average compared to the number of CPUs. It exits with status 1 if any value is over its threshold, so both a person and systemd can tell a warning from a normal run.

sudo nano /opt/ops-tools/check_resources.py
#!/opt/ops-tools/venv/bin/python
"""Report memory, disk and load usage and exit 1 if a threshold is exceeded."""

import argparse
import os
import sys

import psutil

IGNORED_FS = {"tmpfs", "devtmpfs", "squashfs", "overlay"}


def parse_args():
    parser = argparse.ArgumentParser(description=__doc__)
    parser.add_argument("--memory", type=float, default=90.0,
                        help="memory usage threshold in percent (default: 90)")
    parser.add_argument("--disk", type=float, default=85.0,
                        help="disk usage threshold in percent (default: 85)")
    parser.add_argument("--load", type=float, default=1.5,
                        help="5-minute load per CPU threshold (default: 1.5)")
    return parser.parse_args()


def main():
    args = parse_args()
    problems = []

    memory = psutil.virtual_memory().percent
    print(f"memory: {memory:.1f}%")
    if memory >= args.memory:
        problems.append(f"memory at {memory:.1f}%")

    for part in psutil.disk_partitions():
        if part.fstype in IGNORED_FS:
            continue
        usage = psutil.disk_usage(part.mountpoint).percent
        print(f"disk {part.mountpoint}: {usage:.1f}%")
        if usage >= args.disk:
            problems.append(f"disk {part.mountpoint} at {usage:.1f}%")

    cpus = os.cpu_count() or 1
    load_per_cpu = os.getloadavg()[1] / cpus
    print(f"load (5 min) per CPU: {load_per_cpu:.2f}")
    if load_per_cpu >= args.load:
        problems.append(f"load per CPU at {load_per_cpu:.2f}")

    if problems:
        print("WARNING: " + "; ".join(problems), file=sys.stderr)
        return 1
    return 0


if __name__ == "__main__":
    sys.exit(main())

A few design choices:

  • The shebang points to the virtual environment, so psutil is always available when the script is executed directly.
  • psutil.virtual_memory().percent accounts for reclaimable cache, unlike naive "used / total" calculations that make every Linux server look full.
  • Load average is divided by the CPU count; a load of 4 is fine on 8 cores and a problem on 1.

Make the script executable and run it:

sudo chmod 755 /opt/ops-tools/check_resources.py
/opt/ops-tools/check_resources.py; echo "exit code: $?"
memory: 31.4%
disk /: 23.0%
disk /boot: 14.2%
load (5 min) per CPU: 0.04
exit code: 0

Force a warning by lowering a threshold:

/opt/ops-tools/check_resources.py --disk 5; echo "exit code: $?"
memory: 31.4%
disk /: 23.0%
disk /boot: 14.2%
load (5 min) per CPU: 0.03
WARNING: disk / at 23.0%; disk /boot at 14.2%
exit code: 1

Step 3 - Scheduling the checker with a systemd timer

A systemd timer runs the script on a schedule, records every run in the journal and marks failed runs, which here means "a threshold was exceeded".

Create the service unit. DynamicUser=yes runs the script as a throwaway unprivileged user, since reading system metrics does not require root:

sudo nano /etc/systemd/system/check-resources.service
[Unit]
Description=Check memory, disk and load thresholds

[Service]
Type=oneshot
DynamicUser=yes
ExecStart=/opt/ops-tools/check_resources.py

Create the timer:

sudo nano /etc/systemd/system/check-resources.timer
[Unit]
Description=Run the resource check every 5 minutes

[Timer]
OnCalendar=*:0/5
Persistent=true

[Install]
WantedBy=timers.target

Enable the timer:

sudo systemctl daemon-reload
sudo systemctl enable --now check-resources.timer

Trigger a run manually and read the result:

sudo systemctl start check-resources.service
journalctl -u check-resources.service -n 6 --no-pager
Sep 25 11:05:02 web01 systemd[1]: Starting check-resources.service - Check memory, disk and load thresholds...
Sep 25 11:05:02 web01 check_resources.py[6120]: memory: 31.5%
Sep 25 11:05:02 web01 check_resources.py[6120]: disk /: 23.0%
Sep 25 11:05:02 web01 check_resources.py[6120]: disk /boot: 14.2%
Sep 25 11:05:02 web01 check_resources.py[6120]: load (5 min) per CPU: 0.02
Sep 25 11:05:02 web01 systemd[1]: Finished check-resources.service - Check memory, disk and load thresholds.

When a threshold is exceeded, the unit ends in a failed state and appears in systemctl --failed.

Step 4 - Monitoring websites and TLS certificates

The second script checks a list of URLs for a successful HTTP status and response time, and warns when the TLS certificate of an HTTPS site expires soon. Expired certificates are one of the most avoidable outages.

sudo nano /opt/ops-tools/check_sites.py
#!/opt/ops-tools/venv/bin/python
"""Check HTTP status, response time and TLS certificate expiry for URLs."""

import argparse
import socket
import ssl
import sys
import time
from urllib.parse import urlsplit

import requests


def cert_days_left(host, port=443, timeout=10):
    context = ssl.create_default_context()
    with socket.create_connection((host, port), timeout=timeout) as sock:
        with context.wrap_socket(sock, server_hostname=host) as tls:
            cert = tls.getpeercert()
    expires = ssl.cert_time_to_seconds(cert["notAfter"])
    return int((expires - time.time()) // 86400)


def check(url, timeout, min_days):
    problems = []
    try:
        response = requests.get(url, timeout=timeout)
        elapsed_ms = response.elapsed.total_seconds() * 1000
        print(f"{url}: HTTP {response.status_code} in {elapsed_ms:.0f} ms")
        if response.status_code >= 400:
            problems.append(f"{url} returned HTTP {response.status_code}")
    except requests.RequestException as exc:
        problems.append(f"{url} request failed: {exc}")
        return problems

    parts = urlsplit(url)
    if parts.scheme == "https":
        try:
            days = cert_days_left(parts.hostname, parts.port or 443, timeout)
            print(f"{url}: certificate expires in {days} days")
            if days < min_days:
                problems.append(f"{url} certificate expires in {days} days")
        except (OSError, ssl.SSLError) as exc:
            problems.append(f"{url} TLS check failed: {exc}")
    return problems


def main():
    parser = argparse.ArgumentParser(description=__doc__)
    parser.add_argument("urls", nargs="+", help="URLs to check")
    parser.add_argument("--timeout", type=float, default=10.0)
    parser.add_argument("--min-days", type=int, default=14,
                        help="warn when a certificate expires sooner (default: 14)")
    args = parser.parse_args()

    problems = []
    for url in args.urls:
        problems.extend(check(url, args.timeout, args.min_days))

    for problem in problems:
        print(f"WARNING: {problem}", file=sys.stderr)
    return 1 if problems else 0


if __name__ == "__main__":
    sys.exit(main())

requests follows redirects and validates certificates by default, so an invalid certificate shows up as a failed request. The separate cert_days_left function uses only the standard library ssl module to read the expiry date of the certificate the server actually presents.

Make it executable and check a couple of sites, replacing your_domain with a site you run:

sudo chmod 755 /opt/ops-tools/check_sites.py
/opt/ops-tools/check_sites.py https://your_domain https://example.com
https://your_domain: HTTP 200 in 84 ms
https://your_domain: certificate expires in 61 days
https://example.com: HTTP 200 in 112 ms
https://example.com: certificate expires in 143 days

You can schedule this script the same way as in Step 3, passing your URLs on the ExecStart= line.

Step 5 - Summarizing Nginx access logs

The third script reads an Nginx access log in the default combined format, including rotated .gz files, and prints the status code distribution, the busiest client IPs and the most requested paths. It uses only the standard library.

sudo nano /opt/ops-tools/nginx_summary.py
#!/opt/ops-tools/venv/bin/python
"""Summarize an Nginx access log in the combined format."""

import argparse
import gzip
import re
from collections import Counter

LINE_RE = re.compile(
    r'(?P<ip>\S+) \S+ \S+ \[[^\]]+\] '
    r'"(?P<method>[A-Z]+) (?P<path>\S+) [^"]*" '
    r'(?P<status>\d{3}) \S+'
)


def open_log(path):
    if path.endswith(".gz"):
        return gzip.open(path, "rt", encoding="utf-8", errors="replace")
    return open(path, encoding="utf-8", errors="replace")


def main():
    parser = argparse.ArgumentParser(description=__doc__)
    parser.add_argument("logfiles", nargs="+", help="access log files (.log or .gz)")
    parser.add_argument("--top", type=int, default=5, help="rows per table (default: 5)")
    args = parser.parse_args()

    statuses, ips, paths = Counter(), Counter(), Counter()
    total = skipped = 0

    for logfile in args.logfiles:
        with open_log(logfile) as handle:
            for line in handle:
                match = LINE_RE.match(line)
                if not match:
                    skipped += 1
                    continue
                total += 1
                statuses[match["status"]] += 1
                ips[match["ip"]] += 1
                paths[match["path"].split("?", 1)[0]] += 1

    print(f"Requests: {total} (unparsed lines: {skipped})\n")
    print("Status codes:")
    for status, count in sorted(statuses.items()):
        print(f"  {status}  {count:>8}")
    print(f"\nTop {args.top} client IPs:")
    for ip, count in ips.most_common(args.top):
        print(f"  {count:>8}  {ip}")
    print(f"\nTop {args.top} paths:")
    for path, count in paths.most_common(args.top):
        print(f"  {count:>8}  {path}")


if __name__ == "__main__":
    main()

Query strings are stripped from paths so that /search?q=a and /search?q=b are counted together. Lines that do not match the pattern are counted as unparsed instead of crashing the script.

Nginx logs are readable by root and the adm group, so run it with sudo:

sudo chmod 755 /opt/ops-tools/nginx_summary.py
sudo /opt/ops-tools/nginx_summary.py /var/log/nginx/access.log
Requests: 18423 (unparsed lines: 0)

Status codes:
  200     15870
  301       912
  304       604
  404       985
  499        52

Top 5 client IPs:
      2210  203.0.113.24
      1187  198.51.100.7
       954  192.0.2.61
       610  203.0.113.90
       488  198.51.100.33

Top 5 paths:
      4120  /
      2033  /wp-login.php
      1650  /assets/app.css
      1402  /api/health
       988  /favicon.ico

To analyze the last few days, pass the rotated files too, for example sudo /opt/ops-tools/nginx_summary.py /var/log/nginx/access.log*. A high count on paths like /wp-login.php on a site that is not WordPress is a sign of scanners worth blocking.

Troubleshooting

  • error: externally-managed-environment: you ran pip from the system Python. Use /opt/ops-tools/venv/bin/pip instead.
  • ModuleNotFoundError: No module named 'psutil': the script was started with python3 script.py, which uses the system interpreter. Run the script directly so its shebang is used, or call /opt/ops-tools/venv/bin/python.
  • Permission denied reading /var/log/nginx/access.log: run the script with sudo or add your user to the adm group.
  • The venv stops working after a release upgrade: virtual environments are tied to the Python version they were created with. Recreate it with sudo python3 -m venv --clear /opt/ops-tools/venv and reinstall the packages.

Conclusion

You created an isolated Python environment for operations tooling and wrote three scripts that check server resources, watch websites and certificates, and summarize web traffic, each with clear output and meaningful exit codes. The resource checker already runs every five minutes under systemd.

Good next steps are to keep /opt/ops-tools in a Git repository with a requirements.txt, add an OnFailure= unit that sends you a notification when a check fails, and extend check_sites.py to read its URL list from a configuration file.