Python is a good fit for server automation once a task outgrows a few lines of Bash: it has real data structures, clear error handling and libraries for almost everything. In this tutorial you will set up an isolated Python environment on Ubuntu 24.04 and write three practical scripts: a resource checker, a website and TLS certificate monitor, and an Nginx access log summarizer. You will then run the resource checker every five minutes with a systemd timer.
Prerequisites
To follow this tutorial you need:
- A server running Ubuntu 24.04 LTS, for example a CubePath VPS. Ubuntu 24.04 ships Python 3.12.
- A non-root user with
sudoprivileges. - For the log summarizer in Step 5, Nginx serving traffic with its default access log at
/var/log/nginx/access.log. The other scripts work without it. - Basic Python knowledge: functions, lists and dictionaries.
Step 1 - Creating a virtual environment for your tools
Ubuntu 24.04 marks the system Python as "externally managed" (PEP 668), so sudo pip install is refused to protect packages that the operating system depends on. The correct approach is a virtual environment that belongs to your scripts.
Install the venv module:
sudo apt update
sudo apt install -y python3-venv
Create a directory for the tools and a virtual environment inside it:
sudo mkdir -p /opt/ops-tools
sudo python3 -m venv /opt/ops-tools/venv
Install the two third-party libraries the scripts use: psutil for system metrics and requests for HTTP checks:
sudo /opt/ops-tools/venv/bin/pip install psutil requests
Verify that the environment can import them:
/opt/ops-tools/venv/bin/python -c 'import psutil, requests; print(psutil.__version__, requests.__version__)'
7.0.0 2.32.4
Your version numbers may differ. Because each script calls the virtual environment's interpreter directly, you never need to activate the environment, which also makes them easy to run from systemd.
Step 2 - Writing a resource checker
The first script reports memory usage, disk usage for each mount point and load average compared to the number of CPUs. It exits with status 1 if any value is over its threshold, so both a person and systemd can tell a warning from a normal run.
sudo nano /opt/ops-tools/check_resources.py
#!/opt/ops-tools/venv/bin/python
"""Report memory, disk and load usage and exit 1 if a threshold is exceeded."""
import argparse
import os
import sys
import psutil
IGNORED_FS = {"tmpfs", "devtmpfs", "squashfs", "overlay"}
def parse_args():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--memory", type=float, default=90.0,
help="memory usage threshold in percent (default: 90)")
parser.add_argument("--disk", type=float, default=85.0,
help="disk usage threshold in percent (default: 85)")
parser.add_argument("--load", type=float, default=1.5,
help="5-minute load per CPU threshold (default: 1.5)")
return parser.parse_args()
def main():
args = parse_args()
problems = []
memory = psutil.virtual_memory().percent
print(f"memory: {memory:.1f}%")
if memory >= args.memory:
problems.append(f"memory at {memory:.1f}%")
for part in psutil.disk_partitions():
if part.fstype in IGNORED_FS:
continue
usage = psutil.disk_usage(part.mountpoint).percent
print(f"disk {part.mountpoint}: {usage:.1f}%")
if usage >= args.disk:
problems.append(f"disk {part.mountpoint} at {usage:.1f}%")
cpus = os.cpu_count() or 1
load_per_cpu = os.getloadavg()[1] / cpus
print(f"load (5 min) per CPU: {load_per_cpu:.2f}")
if load_per_cpu >= args.load:
problems.append(f"load per CPU at {load_per_cpu:.2f}")
if problems:
print("WARNING: " + "; ".join(problems), file=sys.stderr)
return 1
return 0
if __name__ == "__main__":
sys.exit(main())
A few design choices:
- The shebang points to the virtual environment, so
psutilis always available when the script is executed directly. psutil.virtual_memory().percentaccounts for reclaimable cache, unlike naive "used / total" calculations that make every Linux server look full.- Load average is divided by the CPU count; a load of 4 is fine on 8 cores and a problem on 1.
Make the script executable and run it:
sudo chmod 755 /opt/ops-tools/check_resources.py
/opt/ops-tools/check_resources.py; echo "exit code: $?"
memory: 31.4%
disk /: 23.0%
disk /boot: 14.2%
load (5 min) per CPU: 0.04
exit code: 0
Force a warning by lowering a threshold:
/opt/ops-tools/check_resources.py --disk 5; echo "exit code: $?"
memory: 31.4%
disk /: 23.0%
disk /boot: 14.2%
load (5 min) per CPU: 0.03
WARNING: disk / at 23.0%; disk /boot at 14.2%
exit code: 1
Step 3 - Scheduling the checker with a systemd timer
A systemd timer runs the script on a schedule, records every run in the journal and marks failed runs, which here means "a threshold was exceeded".
Create the service unit. DynamicUser=yes runs the script as a throwaway unprivileged user, since reading system metrics does not require root:
sudo nano /etc/systemd/system/check-resources.service
[Unit]
Description=Check memory, disk and load thresholds
[Service]
Type=oneshot
DynamicUser=yes
ExecStart=/opt/ops-tools/check_resources.py
Create the timer:
sudo nano /etc/systemd/system/check-resources.timer
[Unit]
Description=Run the resource check every 5 minutes
[Timer]
OnCalendar=*:0/5
Persistent=true
[Install]
WantedBy=timers.target
Enable the timer:
sudo systemctl daemon-reload
sudo systemctl enable --now check-resources.timer
Trigger a run manually and read the result:
sudo systemctl start check-resources.service
journalctl -u check-resources.service -n 6 --no-pager
Sep 25 11:05:02 web01 systemd[1]: Starting check-resources.service - Check memory, disk and load thresholds...
Sep 25 11:05:02 web01 check_resources.py[6120]: memory: 31.5%
Sep 25 11:05:02 web01 check_resources.py[6120]: disk /: 23.0%
Sep 25 11:05:02 web01 check_resources.py[6120]: disk /boot: 14.2%
Sep 25 11:05:02 web01 check_resources.py[6120]: load (5 min) per CPU: 0.02
Sep 25 11:05:02 web01 systemd[1]: Finished check-resources.service - Check memory, disk and load thresholds.
When a threshold is exceeded, the unit ends in a failed state and appears in systemctl --failed.
Step 4 - Monitoring websites and TLS certificates
The second script checks a list of URLs for a successful HTTP status and response time, and warns when the TLS certificate of an HTTPS site expires soon. Expired certificates are one of the most avoidable outages.
sudo nano /opt/ops-tools/check_sites.py
#!/opt/ops-tools/venv/bin/python
"""Check HTTP status, response time and TLS certificate expiry for URLs."""
import argparse
import socket
import ssl
import sys
import time
from urllib.parse import urlsplit
import requests
def cert_days_left(host, port=443, timeout=10):
context = ssl.create_default_context()
with socket.create_connection((host, port), timeout=timeout) as sock:
with context.wrap_socket(sock, server_hostname=host) as tls:
cert = tls.getpeercert()
expires = ssl.cert_time_to_seconds(cert["notAfter"])
return int((expires - time.time()) // 86400)
def check(url, timeout, min_days):
problems = []
try:
response = requests.get(url, timeout=timeout)
elapsed_ms = response.elapsed.total_seconds() * 1000
print(f"{url}: HTTP {response.status_code} in {elapsed_ms:.0f} ms")
if response.status_code >= 400:
problems.append(f"{url} returned HTTP {response.status_code}")
except requests.RequestException as exc:
problems.append(f"{url} request failed: {exc}")
return problems
parts = urlsplit(url)
if parts.scheme == "https":
try:
days = cert_days_left(parts.hostname, parts.port or 443, timeout)
print(f"{url}: certificate expires in {days} days")
if days < min_days:
problems.append(f"{url} certificate expires in {days} days")
except (OSError, ssl.SSLError) as exc:
problems.append(f"{url} TLS check failed: {exc}")
return problems
def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("urls", nargs="+", help="URLs to check")
parser.add_argument("--timeout", type=float, default=10.0)
parser.add_argument("--min-days", type=int, default=14,
help="warn when a certificate expires sooner (default: 14)")
args = parser.parse_args()
problems = []
for url in args.urls:
problems.extend(check(url, args.timeout, args.min_days))
for problem in problems:
print(f"WARNING: {problem}", file=sys.stderr)
return 1 if problems else 0
if __name__ == "__main__":
sys.exit(main())
requests follows redirects and validates certificates by default, so an invalid certificate shows up as a failed request. The separate cert_days_left function uses only the standard library ssl module to read the expiry date of the certificate the server actually presents.
Make it executable and check a couple of sites, replacing your_domain with a site you run:
sudo chmod 755 /opt/ops-tools/check_sites.py
/opt/ops-tools/check_sites.py https://your_domain https://example.com
https://your_domain: HTTP 200 in 84 ms
https://your_domain: certificate expires in 61 days
https://example.com: HTTP 200 in 112 ms
https://example.com: certificate expires in 143 days
You can schedule this script the same way as in Step 3, passing your URLs on the ExecStart= line.
Step 5 - Summarizing Nginx access logs
The third script reads an Nginx access log in the default combined format, including rotated .gz files, and prints the status code distribution, the busiest client IPs and the most requested paths. It uses only the standard library.
sudo nano /opt/ops-tools/nginx_summary.py
#!/opt/ops-tools/venv/bin/python
"""Summarize an Nginx access log in the combined format."""
import argparse
import gzip
import re
from collections import Counter
LINE_RE = re.compile(
r'(?P<ip>\S+) \S+ \S+ \[[^\]]+\] '
r'"(?P<method>[A-Z]+) (?P<path>\S+) [^"]*" '
r'(?P<status>\d{3}) \S+'
)
def open_log(path):
if path.endswith(".gz"):
return gzip.open(path, "rt", encoding="utf-8", errors="replace")
return open(path, encoding="utf-8", errors="replace")
def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("logfiles", nargs="+", help="access log files (.log or .gz)")
parser.add_argument("--top", type=int, default=5, help="rows per table (default: 5)")
args = parser.parse_args()
statuses, ips, paths = Counter(), Counter(), Counter()
total = skipped = 0
for logfile in args.logfiles:
with open_log(logfile) as handle:
for line in handle:
match = LINE_RE.match(line)
if not match:
skipped += 1
continue
total += 1
statuses[match["status"]] += 1
ips[match["ip"]] += 1
paths[match["path"].split("?", 1)[0]] += 1
print(f"Requests: {total} (unparsed lines: {skipped})\n")
print("Status codes:")
for status, count in sorted(statuses.items()):
print(f" {status} {count:>8}")
print(f"\nTop {args.top} client IPs:")
for ip, count in ips.most_common(args.top):
print(f" {count:>8} {ip}")
print(f"\nTop {args.top} paths:")
for path, count in paths.most_common(args.top):
print(f" {count:>8} {path}")
if __name__ == "__main__":
main()
Query strings are stripped from paths so that /search?q=a and /search?q=b are counted together. Lines that do not match the pattern are counted as unparsed instead of crashing the script.
Nginx logs are readable by root and the adm group, so run it with sudo:
sudo chmod 755 /opt/ops-tools/nginx_summary.py
sudo /opt/ops-tools/nginx_summary.py /var/log/nginx/access.log
Requests: 18423 (unparsed lines: 0)
Status codes:
200 15870
301 912
304 604
404 985
499 52
Top 5 client IPs:
2210 203.0.113.24
1187 198.51.100.7
954 192.0.2.61
610 203.0.113.90
488 198.51.100.33
Top 5 paths:
4120 /
2033 /wp-login.php
1650 /assets/app.css
1402 /api/health
988 /favicon.ico
To analyze the last few days, pass the rotated files too, for example sudo /opt/ops-tools/nginx_summary.py /var/log/nginx/access.log*. A high count on paths like /wp-login.php on a site that is not WordPress is a sign of scanners worth blocking.
Troubleshooting
error: externally-managed-environment: you ranpipfrom the system Python. Use/opt/ops-tools/venv/bin/pipinstead.ModuleNotFoundError: No module named 'psutil': the script was started withpython3 script.py, which uses the system interpreter. Run the script directly so its shebang is used, or call/opt/ops-tools/venv/bin/python.Permission deniedreading/var/log/nginx/access.log: run the script withsudoor add your user to theadmgroup.- The venv stops working after a release upgrade: virtual environments are tied to the Python version they were created with. Recreate it with
sudo python3 -m venv --clear /opt/ops-tools/venvand reinstall the packages.
Conclusion
You created an isolated Python environment for operations tooling and wrote three scripts that check server resources, watch websites and certificates, and summarize web traffic, each with clear output and meaningful exit codes. The resource checker already runs every five minutes under systemd.
Good next steps are to keep /opt/ops-tools in a Git repository with a requirements.txt, add an OnFailure= unit that sends you a notification when a check fails, and extend check_sites.py to read its URL list from a configuration file.
