grep, awk and sed are installed on every Linux server, work on files of any size and need no setup, which makes them the fastest way to answer questions such as "which IP is hammering my site?" or "who is trying to log in over SSH?". In this tutorial you will use them on two real log files of an Ubuntu 24.04 server: the Nginx access log and /var/log/auth.log. Each step builds on the previous one, and the last one combines the three tools into pipelines you can reuse during an incident.
Prerequisites
To follow this tutorial you need:
- A server running Ubuntu 24.04 LTS, such as a CubePath VPS, and a non-root user with
sudoprivileges. - Nginx installed with its default logging (
sudo apt install nginx), and ideally some real traffic in/var/log/nginx/access.log. The same commands work for Apache'scombinedformat. - rsyslog running, which is the default on Ubuntu Server, so that
/var/log/auth.logexists.
Log files under /var/log are readable by the adm group. Instead of prefixing every command with sudo, add your user to that group and log in again:
sudo usermod -aG adm your_user
After logging back in, id should list adm among your groups.
Step 1 - Understanding the log formats
Before filtering anything, look at a few lines to know which field holds what:
tail -n 2 /var/log/nginx/access.log
203.0.113.45 - - [25/Sep/2026:14:03:11 +0000] "GET /index.html HTTP/1.1" 200 615 "-" "Mozilla/5.0 (X11; Linux x86_64)"
198.51.100.7 - - [25/Sep/2026:14:03:12 +0000] "GET /wp-login.php HTTP/1.1" 404 162 "-" "Mozilla/5.0"
This is the combined format. When awk splits the line on whitespace, the fields you will use most are:
| Field | Content | Example |
|---|---|---|
$1 | Client IP | 203.0.113.45 |
$4 | Date and time, with a leading [ | [25/Sep/2026:14:03:11 |
$6 | Method, with a leading " | "GET |
$7 | Requested path | /index.html |
$9 | HTTP status | 200 |
$10 | Response size in bytes | 615 |
The user agent contains spaces, so it spans several fields and is easier to extract by splitting on double quotes, as shown later.
Now look at the authentication log:
grep sshd /var/log/auth.log | tail -n 2
2026-09-25T14:05:21.339112+00:00 web01 sshd[18234]: Failed password for invalid user admin from 198.51.100.7 port 51522 ssh2
2026-09-25T14:05:40.102771+00:00 web01 sshd[18240]: Accepted publickey for deploy from 203.0.113.45 port 40112 ssh2: ED25519 SHA256:...
Ubuntu 24.04 writes syslog files with an ISO 8601 timestamp, so $1 is the timestamp, $2 the hostname and $3 the program with its PID. Older releases used the Sep 25 14:05:21 format, which takes three fields.
Step 2 - Filtering lines with grep
grep prints the lines that match a pattern. It is the first tool in almost every pipeline because it quickly reduces a large file to the lines you care about.
Find every failed SSH password attempt and count them:
grep 'Failed password' /var/log/auth.log
grep -c 'Failed password' /var/log/auth.log
1284
The options you will use most often are:
| Option | Effect |
|---|---|
-i | Case-insensitive match |
-v | Print lines that do NOT match |
-c | Print only the number of matching lines |
-E | Extended regular expressions: alternation, + and {n} without backslashes |
-F | Treat the pattern as a fixed string, faster and no escaping needed |
-o | Print only the matched part of the line |
-A n, -B n, -C n | Show n lines after, before or around each match |
-w | Match whole words only |
Some practical examples:
# Requests that returned 500, 502, 503 or 504
grep -E '" 50[0-4] ' /var/log/nginx/access.log
# Everything except health checks and static files
grep -vE '/healthz|\.(css|js|png|jpg|svg|ico) ' /var/log/nginx/access.log
# All requests from one IP, matched literally
grep -F '198.51.100.7 ' /var/log/nginx/access.log
# Requests during one hour: 14:00 to 14:59 on 25 September
grep '25/Sep/2026:14:' /var/log/nginx/access.log
# Show 2 lines of context around each Nginx error
grep -C 2 -i 'error' /var/log/nginx/error.log
Logs are rotated and older files are compressed, for example access.log.2.gz. Use zgrep, which accepts the same options, to search them without decompressing:
zgrep -c 'wp-login.php' /var/log/nginx/access.log.*.gz
Step 3 - Extracting and counting fields with awk
awk processes a file line by line, splits each line into fields and runs a small program on them. The syntax is awk 'condition { action }' file. Ubuntu ships mawk as awk, which is fast and supports everything used here.
Printing and filtering by field
Print the IP and path of every request:
awk '{print $1, $7}' /var/log/nginx/access.log | head -n 3
203.0.113.45 /index.html
198.51.100.7 /wp-login.php
203.0.113.45 /about/
Filter on a field value instead of on the whole line. This is more precise than grep ' 404 ', which would also match a response of 404 bytes:
# Only 404 responses
awk '$9 == 404' /var/log/nginx/access.log
# All 5xx responses
awk '$9 >= 500' /var/log/nginx/access.log
# Responses larger than 1 MB
awk '$10 > 1048576 {print $1, $7, $10}' /var/log/nginx/access.log
Counting with sort and uniq
The classic pattern to rank anything is: print one field, sort it, count identical lines with uniq -c, then sort the counts in reverse numeric order.
Top 10 client IPs:
awk '{print $1}' /var/log/nginx/access.log | sort | uniq -c | sort -rn | head -n 10
4821 203.0.113.45
1203 198.51.100.7
388 192.0.2.19
Requests per status code:
awk '{print $9}' /var/log/nginx/access.log | sort | uniq -c | sort -rn
15230 200
1904 404
612 301
12 502
Most requested paths that returned 404:
awk '$9 == 404 {print $7}' /var/log/nginx/access.log | sort | uniq -c | sort -rn | head -n 10
Aggregating with arrays
awk arrays let you count and sum without calling sort and uniq. The END block runs once after the last line.
Total bytes sent, in MiB:
awk '{bytes += $10} END {printf "%.1f MiB\n", bytes / 1024 / 1024}' /var/log/nginx/access.log
842.3 MiB
Requests per hour. The time field $4 looks like [25/Sep/2026:14:03:11, so splitting it on : puts the hour in the second element:
awk '{split($4, t, ":"); hits[t[2]]++} END {for (h in hits) print h, hits[h]}' /var/log/nginx/access.log | sort
13 1840
14 2291
15 1975
Bytes served per IP, largest first:
awk '{bytes[$1] += $10} END {for (ip in bytes) print bytes[ip], ip}' /var/log/nginx/access.log | sort -rn | head -n 5
Using a different field separator
The -F option changes the separator. Splitting the combined format on double quotes gives the request in $2 and the user agent in $6, which is the easiest way to rank user agents:
awk -F'"' '{print $6}' /var/log/nginx/access.log | sort | uniq -c | sort -rn | head -n 5
Step 4 - Transforming lines with sed
sed edits text as it streams through. In log analysis it is mostly used to print ranges of lines, to clean up lines and to rewrite them before handing them to another tool.
Print a range of lines by number, which is handy for large files:
sed -n '1000,1020p' /var/log/nginx/access.log
Print everything between two patterns. This prints from the first line that contains 14:00: on that day up to the first line that contains 14:30::
sed -n '/25\/Sep\/2026:14:00:/,/25\/Sep\/2026:14:30:/p' /var/log/nginx/access.log
The range only starts if a line matches the first pattern, so pick a start time where you know there was traffic. For an exact window, awk on the time field is more robust.
Replace text with s/pattern/replacement/. Anonymize the last octet of client IPs before sharing a log excerpt:
sed -E 's/^([0-9]+\.[0-9]+\.[0-9]+)\.[0-9]+/\1.0/' /var/log/nginx/access.log | head -n 2
203.0.113.0 - - [25/Sep/2026:14:03:11 +0000] "GET /index.html HTTP/1.1" 200 615 "-" "Mozilla/5.0 (X11; Linux x86_64)"
198.51.100.0 - - [25/Sep/2026:14:03:12 +0000] "GET /wp-login.php HTTP/1.1" 404 162 "-" "Mozilla/5.0"
Strip query strings so that /search?q=a and /search?q=b count as the same path:
awk '{print $7}' /var/log/nginx/access.log | sed 's/?.*//' | sort | uniq -c | sort -rn | head -n 10
Remove ANSI color codes from application logs that were written for a terminal:
sed 's/\x1b\[[0-9;]*m//g' /var/log/myapp/app.log > app-clean.log
Without -i, sed never modifies the original file; it writes the result to standard output. Avoid sed -i on live log files, because the application is still writing to them.
Step 5 - Combining the tools in pipelines
The real power comes from chaining the three tools: grep narrows the input, sed normalizes it and awk counts. These are pipelines worth keeping at hand.
IPs with the most failed SSH logins. Instead of relying on a field position, the loop finds the word from and prints the next field, which works for both valid and invalid users:
grep 'Failed password' /var/log/auth.log \
| awk '{for (i = 1; i <= NF; i++) if ($i == "from") print $(i + 1)}' \
| sort | uniq -c | sort -rn | head -n 10
512 198.51.100.7
301 192.0.2.88
44 203.0.113.200
Usernames tried by attackers that do not exist on the system:
grep -o 'Invalid user [^ ]*' /var/log/auth.log | awk '{print $3}' | sort | uniq -c | sort -rn | head -n 10
Successful logins, with user and source IP:
grep 'Accepted ' /var/log/auth.log | awk '{print $1, $5, $7, $9}'
2026-09-25T14:05:40.102771+00:00 publickey deploy 203.0.113.45
Commands run with sudo:
grep -F 'COMMAND=' /var/log/auth.log | sed -E 's/.*sudo: +([^ ]+) .*COMMAND=(.*)/\1: \2/' | tail -n 10
IPs that are probing for WordPress or PHP files on a site that does not use them:
grep -E '(wp-login|xmlrpc|\.php)' /var/log/nginx/access.log | awk '$9 == 404 {print $1}' | sort | uniq -c | sort -rn | head
Error rate per hour, as a percentage of all requests:
awk '{split($4, t, ":"); total[t[2]]++; if ($9 >= 500) err[t[2]]++}
END {for (h in total) printf "%s %d requests %.2f%% 5xx\n", h, total[h], 100 * err[h] / total[h]}' \
/var/log/nginx/access.log | sort
To analyze today's log together with the rotated ones, zcat -f reads compressed and uncompressed files alike:
zcat -f /var/log/nginx/access.log* | awk '{print $1}' | sort | uniq -c | sort -rn | head -n 10
Step 6 - Watching logs in real time
During an incident you often want to watch new lines as they arrive. tail -F follows a file and reopens it after rotation. When piping it into another tool, make sure the output is not buffered, otherwise nothing appears until a large block has accumulated:
# Only server errors, as they happen
tail -F /var/log/nginx/access.log | awk '$9 >= 500 {print; fflush()}'
# Failed SSH logins, as they happen
tail -F /var/log/auth.log | grep --line-buffered 'Failed password'
Press CTRL+C to stop.
Services that only log to the systemd journal have no file under /var/log. Send the journal to the same tools with journalctl:
journalctl -u nginx --since "1 hour ago" --no-pager | grep -i error
Troubleshooting
grepprintsbinary file matchesinstead of the lines. The log contains a few non-text bytes. Add-ato treat it as text:grep -a 'pattern' file.Permission deniedwhen reading/var/log/auth.log. Your user is not in theadmgroup yet, or you have not logged in again after adding it. Check withid.- awk results look shifted. A different log format moves the fields. Number the fields of one line to see where each value really is:
head -n 1 file | awk '{for (i = 1; i <= NF; i++) print i, $i}'. - Pipelines are slow on multi-GB files. Filter as early as possible with
grep -F, and run the pipeline withLC_ALL=Cin front ofgrepandsortto avoid slow locale-aware matching.
Conclusion
You can now filter logs with grep, extract and aggregate fields with awk, clean and reshape lines with sed, and chain them into pipelines that answer operational and security questions in seconds. As next steps, turn the pipelines you use most into shell aliases, add a custom Nginx log format with $request_time to analyze slow requests the same way, or centralize logs from several servers with rsyslog so you can run these commands across your whole fleet from one place.
