Structured logging means writing every log entry as a machine-readable object (usually one JSON document per line) with consistent field names, instead of free-form text. Aggregators such as Loki, Elasticsearch or Graylog can then filter and graph on fields like status or request_id without fragile regular expressions. In this tutorial you will define a small log schema, make Nginx write JSON access logs, build a Python service that logs JSON with structlog, and link both with a shared request ID on Ubuntu 24.04.
Prerequisites
To follow this tutorial you need:
- A server running Ubuntu 24.04 LTS, for example a CubePath VPS.
- A non-root user with
sudoprivileges. - Nginx installed (
sudo apt install nginx) with port 80 reachable, or at least usable locally withcurl. - Basic familiarity with Python and the command line.
The same approach works for any language: the Python service is only the example. Equivalent libraries are pino for Node.js and the standard library log/slog package for Go.
Step 1 - Defining a log schema
The value of structured logs comes from consistency. If one service writes status and another writes http_status_code, cross-service queries break. Before touching any configuration, agree on a small set of field names and use them everywhere.
| Field | Example | Notes |
|---|---|---|
timestamp | 2026-09-25T10:30:45.123Z | ISO 8601, always UTC |
level | info | debug, info, warning, error, critical (lowercase) |
message | request completed | Short, constant text; put variable data in fields |
service | orders-api | Name of the application |
request_id | 9f1c2b... | Same value in every log line of one request, in every service |
method, path, status | GET, /orders/42, 200 | HTTP fields; status as a number |
duration_ms | 12.4 | Durations with the unit in the name |
A few rules that save pain later:
- Keep
messageconstant ("user login failed") and put the variable parts in fields ("user_id": 42). Constant messages can be grouped and counted. - Use
infoas the default level in production. Enabledebugtemporarily, not permanently. - Never log passwords, tokens, full card numbers or session cookies. Log an ID that lets you find the record instead.
- One JSON object per line. Pretty-printed, multi-line JSON breaks most log shippers.
Step 2 - Writing Nginx access logs as JSON
Nginx can write JSON natively with log_format ... escape=json, which escapes quotes and control characters so every line stays valid JSON. Nginx also generates a random $request_id per request. The map below reuses an incoming X-Request-ID header when a client or upstream proxy already set one, and falls back to Nginx's own ID otherwise.
Files in /etc/nginx/conf.d/ are included inside the http block, which is where map and log_format belong. Create one:
sudo nano /etc/nginx/conf.d/json-logging.conf
Add the following content:
map $http_x_request_id $req_id {
default $http_x_request_id;
"" $request_id;
}
log_format json_access escape=json
'{'
'"timestamp":"$time_iso8601",'
'"level":"info",'
'"message":"http request",'
'"service":"nginx",'
'"request_id":"$req_id",'
'"remote_addr":"$remote_addr",'
'"method":"$request_method",'
'"path":"$request_uri",'
'"status":$status,'
'"bytes_sent":$body_bytes_sent,'
'"request_time_s":$request_time,'
'"upstream_time_s":"$upstream_response_time",'
'"user_agent":"$http_user_agent"'
'}';
$status, $body_bytes_sent and $request_time are always numeric, so they are written without quotes. $upstream_response_time can be - or a comma-separated list when several upstreams were tried, so it stays a string.
Next, create a site that proxies to the Python service you will build in Step 3, passes the request ID upstream, returns it to the client and writes the JSON log:
sudo nano /etc/nginx/sites-available/logdemo
server {
listen 80;
server_name your_domain;
access_log /var/log/nginx/access_json.log json_access;
location / {
proxy_set_header Host $host;
proxy_set_header X-Request-ID $req_id;
add_header X-Request-ID $req_id always;
proxy_pass http://127.0.0.1:5000;
}
}
Replace your_domain with your domain or the server's IP address. The log file ends in .log on purpose: the logrotate configuration shipped with Ubuntu's Nginx package only rotates /var/log/nginx/*.log.
Enable the site, test the configuration and reload Nginx:
sudo ln -s /etc/nginx/sites-available/logdemo /etc/nginx/sites-enabled/
sudo nginx -t
sudo systemctl reload nginx
nginx: the configuration file /etc/nginx/nginx.conf syntax is ok
nginx: configuration file /etc/nginx/nginx.conf test is successful
Until the backend exists, requests return 502 Bad Gateway, which is still logged. Send one and read the last log line:
curl -s -o /dev/null http://your_domain/
sudo tail -n 1 /var/log/nginx/access_json.log
{"timestamp":"2026-09-25T10:30:45+00:00","level":"info","message":"http request","service":"nginx","request_id":"5d2f8a1c9b0e4f7da3c6e1b2f4a8d9c0","remote_addr":"203.0.113.10","method":"GET","path":"/","status":502,"bytes_sent":166,"request_time_s":0.001,"upstream_time_s":"0.000","user_agent":"curl/8.5.0"}
Step 3 - Logging JSON from a Python service with structlog
structlog builds log entries as dictionaries and renders them as JSON at the end of a processor chain. Its contextvars support lets you bind request_id once at the start of a request and have it added to every log line produced while handling that request.
Install the tools to create a virtual environment and a project directory:
sudo apt install python3-venv jq
mkdir ~/logdemo && cd ~/logdemo
python3 -m venv venv
./venv/bin/pip install flask gunicorn structlog
Create the application:
nano ~/logdemo/app.py
import logging
import time
import uuid
import structlog
from flask import Flask, g, request
structlog.configure(
processors=[
structlog.contextvars.merge_contextvars,
structlog.processors.add_log_level,
structlog.processors.TimeStamper(fmt="iso", utc=True),
structlog.processors.dict_tracebacks,
structlog.processors.EventRenamer("message"),
structlog.processors.JSONRenderer(),
],
wrapper_class=structlog.make_filtering_bound_logger(logging.INFO),
)
log = structlog.get_logger().bind(service="orders-api")
app = Flask(__name__)
@app.before_request
def bind_request_context():
g.start = time.perf_counter()
request_id = request.headers.get("X-Request-ID") or uuid.uuid4().hex
structlog.contextvars.clear_contextvars()
structlog.contextvars.bind_contextvars(
request_id=request_id,
method=request.method,
path=request.path,
)
@app.after_request
def log_request(response):
duration_ms = round((time.perf_counter() - g.start) * 1000, 2)
log.info("request completed", status=response.status_code, duration_ms=duration_ms)
response.headers["X-Request-ID"] = structlog.contextvars.get_contextvars()["request_id"]
return response
@app.get("/orders/<int:order_id>")
def get_order(order_id):
log.info("fetching order", order_id=order_id)
if order_id == 0:
try:
1 / 0
except ZeroDivisionError:
log.exception("order lookup failed", order_id=order_id)
return {"detail": "internal error"}, 500
return {"id": order_id, "state": "shipped"}
What each processor does:
merge_contextvarsadds the fields bound inbind_request_context(request_id,method,path).add_log_levelandTimeStamper(fmt="iso", utc=True)addleveland a UTCtimestamp.dict_tracebacksturns exceptions logged withlog.exception()into a structuredexceptionfield instead of a multi-line string.EventRenamer("message")renames structlog's defaulteventkey tomessage, matching the schema from Step 1.make_filtering_bound_logger(logging.INFO)dropsdebugcalls cheaply.
Start the app in the foreground to test it:
cd ~/logdemo
./venv/bin/gunicorn --bind 127.0.0.1:5000 app:app
In a second terminal, send a request through Nginx:
curl -i http://your_domain/orders/42
The response carries the X-Request-ID header, and the Gunicorn terminal prints one JSON line per log call:
{"service": "orders-api", "order_id": 42, "request_id": "8b1e0f2a6c7d4e9fa1b2c3d4e5f60718", "method": "GET", "path": "/orders/42", "level": "info", "timestamp": "2026-09-25T10:31:02.418733Z", "message": "fetching order"}
{"service": "orders-api", "status": 200, "duration_ms": 0.41, "request_id": "8b1e0f2a6c7d4e9fa1b2c3d4e5f60718", "method": "GET", "path": "/orders/42", "level": "info", "timestamp": "2026-09-25T10:31:02.419120Z", "message": "request completed"}
Stop Gunicorn with CTRL+C.
Step 4 - Running the service under systemd
On a server, applications should log to standard output and let systemd's journal collect it, rather than managing their own log files. The journal adds its own metadata (unit, PID, host) and a shipper can read from it.
Create a unit file:
sudo nano /etc/systemd/system/logdemo.service
[Unit]
Description=Structured logging demo (orders-api)
After=network.target
[Service]
User=your_user
WorkingDirectory=/home/your_user/logdemo
ExecStart=/home/your_user/logdemo/venv/bin/gunicorn --bind 127.0.0.1:5000 app:app
Environment=PYTHONUNBUFFERED=1
Restart=on-failure
[Install]
WantedBy=multi-user.target
Replace your_user with your username, then start the service:
sudo systemctl daemon-reload
sudo systemctl enable --now logdemo
systemctl status logdemo --no-pager
The status output should show Active: active (running).
Step 5 - Querying and correlating logs with jq
Generate a successful request and a failing one:
curl -s http://your_domain/orders/7 > /dev/null
curl -s -D - -o /dev/null http://your_domain/orders/0 | grep -i x-request-id
X-Request-ID: 3e9a41c07b5d4b2e8f6a1c9d0e2b7f54
Find all server errors in the Nginx log:
jq -c 'select(.status >= 500) | {timestamp, request_id, path, status}' /var/log/nginx/access_json.log
Now take that request_id and pull every application log line for the same request from the journal. -o cat prints only the message, and fromjson? skips lines that are not JSON (for example Gunicorn's own startup messages):
journalctl -u logdemo -o cat --since "10 min ago" \
| jq -R -c 'fromjson? | select(.request_id == "3e9a41c07b5d4b2e8f6a1c9d0e2b7f54")'
You get the fetching order line, the order lookup failed line with a structured exception field and the request completed line with "status": 500. The same ID links the proxy log and the application log, which is exactly what an aggregator does at scale when you search for one field value.
A few more useful queries:
# Slowest 5 requests seen by Nginx
jq -s -c 'sort_by(-.request_time_s) | .[:5][] | {path, request_time_s}' /var/log/nginx/access_json.log
# Count application log lines per level
journalctl -u logdemo -o cat | jq -R -r 'fromjson? | .level' | sort | uniq -c
The journal itself can also emit JSON, which is what shippers such as Promtail, Grafana Alloy, Vector or Filebeat read when they collect from journald:
journalctl -u logdemo -o json -n 1 | jq '{_SYSTEMD_UNIT, _PID, MESSAGE}'
Troubleshooting
jq reports parse error: Invalid numeric literal. A non-JSON line is mixed into the stream. Use jq -R 'fromjson? | ...' as shown above, and make sure your application writes only JSON to standard output.
Nginx fails with unknown "req_id" variable. The map in /etc/nginx/conf.d/json-logging.conf is not loaded. Confirm that /etc/nginx/nginx.conf still contains include /etc/nginx/conf.d/*.conf; inside the http block.
Log lines appear late in the journal. Python buffers standard output when it is not a terminal. Keep Environment=PYTHONUNBUFFERED=1 in the unit file.
Fields have different types across services. If one service writes "status":"200" and another "status":200, Elasticsearch and similar systems reject or misindex one of them. Fix the type at the source rather than in the aggregator.
Conclusion
You now have JSON access logs from Nginx and JSON application logs from a Python service, both carrying the same request_id, and you can filter and correlate them with jq. The key discipline is the schema: the same field names and types in every service. From here, ship these logs to a central system, for example with Filebeat to Elasticsearch or with Fluentd, and propagate X-Request-ID in calls between your own services so one ID follows a request end to end.
