grep, awk and sed are installed on every Linux server, work on files of any size and need no setup, which makes them the fastest way to answer questions such as "which IP is hammering my site?" or "who is trying to log in over SSH?". In this tutorial you will use them on two real log files of an Ubuntu 24.04 server: the Nginx access log and /var/log/auth.log. Each step builds on the previous one, and the last one combines the three tools into pipelines you can reuse during an incident.

Prerequisites

To follow this tutorial you need:

  • A server running Ubuntu 24.04 LTS, such as a CubePath VPS, and a non-root user with sudo privileges.
  • Nginx installed with its default logging (sudo apt install nginx), and ideally some real traffic in /var/log/nginx/access.log. The same commands work for Apache's combined format.
  • rsyslog running, which is the default on Ubuntu Server, so that /var/log/auth.log exists.

Log files under /var/log are readable by the adm group. Instead of prefixing every command with sudo, add your user to that group and log in again:

sudo usermod -aG adm your_user

After logging back in, id should list adm among your groups.

Step 1 - Understanding the log formats

Before filtering anything, look at a few lines to know which field holds what:

tail -n 2 /var/log/nginx/access.log
203.0.113.45 - - [25/Sep/2026:14:03:11 +0000] "GET /index.html HTTP/1.1" 200 615 "-" "Mozilla/5.0 (X11; Linux x86_64)"
198.51.100.7 - - [25/Sep/2026:14:03:12 +0000] "GET /wp-login.php HTTP/1.1" 404 162 "-" "Mozilla/5.0"

This is the combined format. When awk splits the line on whitespace, the fields you will use most are:

FieldContentExample
$1Client IP203.0.113.45
$4Date and time, with a leading [[25/Sep/2026:14:03:11
$6Method, with a leading ""GET
$7Requested path/index.html
$9HTTP status200
$10Response size in bytes615

The user agent contains spaces, so it spans several fields and is easier to extract by splitting on double quotes, as shown later.

Now look at the authentication log:

grep sshd /var/log/auth.log | tail -n 2
2026-09-25T14:05:21.339112+00:00 web01 sshd[18234]: Failed password for invalid user admin from 198.51.100.7 port 51522 ssh2
2026-09-25T14:05:40.102771+00:00 web01 sshd[18240]: Accepted publickey for deploy from 203.0.113.45 port 40112 ssh2: ED25519 SHA256:...

Ubuntu 24.04 writes syslog files with an ISO 8601 timestamp, so $1 is the timestamp, $2 the hostname and $3 the program with its PID. Older releases used the Sep 25 14:05:21 format, which takes three fields.

Step 2 - Filtering lines with grep

grep prints the lines that match a pattern. It is the first tool in almost every pipeline because it quickly reduces a large file to the lines you care about.

Find every failed SSH password attempt and count them:

grep 'Failed password' /var/log/auth.log
grep -c 'Failed password' /var/log/auth.log
1284

The options you will use most often are:

OptionEffect
-iCase-insensitive match
-vPrint lines that do NOT match
-cPrint only the number of matching lines
-EExtended regular expressions: alternation, + and {n} without backslashes
-FTreat the pattern as a fixed string, faster and no escaping needed
-oPrint only the matched part of the line
-A n, -B n, -C nShow n lines after, before or around each match
-wMatch whole words only

Some practical examples:

# Requests that returned 500, 502, 503 or 504
grep -E '" 50[0-4] ' /var/log/nginx/access.log

# Everything except health checks and static files
grep -vE '/healthz|\.(css|js|png|jpg|svg|ico) ' /var/log/nginx/access.log

# All requests from one IP, matched literally
grep -F '198.51.100.7 ' /var/log/nginx/access.log

# Requests during one hour: 14:00 to 14:59 on 25 September
grep '25/Sep/2026:14:' /var/log/nginx/access.log

# Show 2 lines of context around each Nginx error
grep -C 2 -i 'error' /var/log/nginx/error.log

Logs are rotated and older files are compressed, for example access.log.2.gz. Use zgrep, which accepts the same options, to search them without decompressing:

zgrep -c 'wp-login.php' /var/log/nginx/access.log.*.gz

Step 3 - Extracting and counting fields with awk

awk processes a file line by line, splits each line into fields and runs a small program on them. The syntax is awk 'condition { action }' file. Ubuntu ships mawk as awk, which is fast and supports everything used here.

Printing and filtering by field

Print the IP and path of every request:

awk '{print $1, $7}' /var/log/nginx/access.log | head -n 3
203.0.113.45 /index.html
198.51.100.7 /wp-login.php
203.0.113.45 /about/

Filter on a field value instead of on the whole line. This is more precise than grep ' 404 ', which would also match a response of 404 bytes:

# Only 404 responses
awk '$9 == 404' /var/log/nginx/access.log

# All 5xx responses
awk '$9 >= 500' /var/log/nginx/access.log

# Responses larger than 1 MB
awk '$10 > 1048576 {print $1, $7, $10}' /var/log/nginx/access.log

Counting with sort and uniq

The classic pattern to rank anything is: print one field, sort it, count identical lines with uniq -c, then sort the counts in reverse numeric order.

Top 10 client IPs:

awk '{print $1}' /var/log/nginx/access.log | sort | uniq -c | sort -rn | head -n 10
   4821 203.0.113.45
   1203 198.51.100.7
    388 192.0.2.19

Requests per status code:

awk '{print $9}' /var/log/nginx/access.log | sort | uniq -c | sort -rn
  15230 200
   1904 404
    612 301
     12 502

Most requested paths that returned 404:

awk '$9 == 404 {print $7}' /var/log/nginx/access.log | sort | uniq -c | sort -rn | head -n 10

Aggregating with arrays

awk arrays let you count and sum without calling sort and uniq. The END block runs once after the last line.

Total bytes sent, in MiB:

awk '{bytes += $10} END {printf "%.1f MiB\n", bytes / 1024 / 1024}' /var/log/nginx/access.log
842.3 MiB

Requests per hour. The time field $4 looks like [25/Sep/2026:14:03:11, so splitting it on : puts the hour in the second element:

awk '{split($4, t, ":"); hits[t[2]]++} END {for (h in hits) print h, hits[h]}' /var/log/nginx/access.log | sort
13 1840
14 2291
15 1975

Bytes served per IP, largest first:

awk '{bytes[$1] += $10} END {for (ip in bytes) print bytes[ip], ip}' /var/log/nginx/access.log | sort -rn | head -n 5

Using a different field separator

The -F option changes the separator. Splitting the combined format on double quotes gives the request in $2 and the user agent in $6, which is the easiest way to rank user agents:

awk -F'"' '{print $6}' /var/log/nginx/access.log | sort | uniq -c | sort -rn | head -n 5

Step 4 - Transforming lines with sed

sed edits text as it streams through. In log analysis it is mostly used to print ranges of lines, to clean up lines and to rewrite them before handing them to another tool.

Print a range of lines by number, which is handy for large files:

sed -n '1000,1020p' /var/log/nginx/access.log

Print everything between two patterns. This prints from the first line that contains 14:00: on that day up to the first line that contains 14:30::

sed -n '/25\/Sep\/2026:14:00:/,/25\/Sep\/2026:14:30:/p' /var/log/nginx/access.log

The range only starts if a line matches the first pattern, so pick a start time where you know there was traffic. For an exact window, awk on the time field is more robust.

Replace text with s/pattern/replacement/. Anonymize the last octet of client IPs before sharing a log excerpt:

sed -E 's/^([0-9]+\.[0-9]+\.[0-9]+)\.[0-9]+/\1.0/' /var/log/nginx/access.log | head -n 2
203.0.113.0 - - [25/Sep/2026:14:03:11 +0000] "GET /index.html HTTP/1.1" 200 615 "-" "Mozilla/5.0 (X11; Linux x86_64)"
198.51.100.0 - - [25/Sep/2026:14:03:12 +0000] "GET /wp-login.php HTTP/1.1" 404 162 "-" "Mozilla/5.0"

Strip query strings so that /search?q=a and /search?q=b count as the same path:

awk '{print $7}' /var/log/nginx/access.log | sed 's/?.*//' | sort | uniq -c | sort -rn | head -n 10

Remove ANSI color codes from application logs that were written for a terminal:

sed 's/\x1b\[[0-9;]*m//g' /var/log/myapp/app.log > app-clean.log

Without -i, sed never modifies the original file; it writes the result to standard output. Avoid sed -i on live log files, because the application is still writing to them.

Step 5 - Combining the tools in pipelines

The real power comes from chaining the three tools: grep narrows the input, sed normalizes it and awk counts. These are pipelines worth keeping at hand.

IPs with the most failed SSH logins. Instead of relying on a field position, the loop finds the word from and prints the next field, which works for both valid and invalid users:

grep 'Failed password' /var/log/auth.log \
  | awk '{for (i = 1; i <= NF; i++) if ($i == "from") print $(i + 1)}' \
  | sort | uniq -c | sort -rn | head -n 10
    512 198.51.100.7
    301 192.0.2.88
     44 203.0.113.200

Usernames tried by attackers that do not exist on the system:

grep -o 'Invalid user [^ ]*' /var/log/auth.log | awk '{print $3}' | sort | uniq -c | sort -rn | head -n 10

Successful logins, with user and source IP:

grep 'Accepted ' /var/log/auth.log | awk '{print $1, $5, $7, $9}'
2026-09-25T14:05:40.102771+00:00 publickey deploy 203.0.113.45

Commands run with sudo:

grep -F 'COMMAND=' /var/log/auth.log | sed -E 's/.*sudo: +([^ ]+) .*COMMAND=(.*)/\1: \2/' | tail -n 10

IPs that are probing for WordPress or PHP files on a site that does not use them:

grep -E '(wp-login|xmlrpc|\.php)' /var/log/nginx/access.log | awk '$9 == 404 {print $1}' | sort | uniq -c | sort -rn | head

Error rate per hour, as a percentage of all requests:

awk '{split($4, t, ":"); total[t[2]]++; if ($9 >= 500) err[t[2]]++}
     END {for (h in total) printf "%s %d requests %.2f%% 5xx\n", h, total[h], 100 * err[h] / total[h]}' \
  /var/log/nginx/access.log | sort

To analyze today's log together with the rotated ones, zcat -f reads compressed and uncompressed files alike:

zcat -f /var/log/nginx/access.log* | awk '{print $1}' | sort | uniq -c | sort -rn | head -n 10

Step 6 - Watching logs in real time

During an incident you often want to watch new lines as they arrive. tail -F follows a file and reopens it after rotation. When piping it into another tool, make sure the output is not buffered, otherwise nothing appears until a large block has accumulated:

# Only server errors, as they happen
tail -F /var/log/nginx/access.log | awk '$9 >= 500 {print; fflush()}'

# Failed SSH logins, as they happen
tail -F /var/log/auth.log | grep --line-buffered 'Failed password'

Press CTRL+C to stop.

Services that only log to the systemd journal have no file under /var/log. Send the journal to the same tools with journalctl:

journalctl -u nginx --since "1 hour ago" --no-pager | grep -i error

Troubleshooting

  • grep prints binary file matches instead of the lines. The log contains a few non-text bytes. Add -a to treat it as text: grep -a 'pattern' file.
  • Permission denied when reading /var/log/auth.log. Your user is not in the adm group yet, or you have not logged in again after adding it. Check with id.
  • awk results look shifted. A different log format moves the fields. Number the fields of one line to see where each value really is: head -n 1 file | awk '{for (i = 1; i <= NF; i++) print i, $i}'.
  • Pipelines are slow on multi-GB files. Filter as early as possible with grep -F, and run the pipeline with LC_ALL=C in front of grep and sort to avoid slow locale-aware matching.

Conclusion

You can now filter logs with grep, extract and aggregate fields with awk, clean and reshape lines with sed, and chain them into pipelines that answer operational and security questions in seconds. As next steps, turn the pipelines you use most into shell aliases, add a custom Nginx log format with $request_time to analyze slow requests the same way, or centralize logs from several servers with rsyslog so you can run these commands across your whole fleet from one place.