bpftrace is a high-level tracing language for Linux built on eBPF. It lets you attach small programs to kernel functions, tracepoints and user-space functions on a running system, aggregate the results in the kernel and print them, all with low overhead and without restarting anything. In this tutorial you will install bpftrace on Ubuntu 24.04, learn how probes and maps work, run the one-liners that answer the most common performance questions, and save a reusable script.
Prerequisites
To follow this tutorial you need:
- A server running Ubuntu 24.04 LTS, for example a CubePath VPS. The stock 6.8 kernel ships with BTF type information, so no kernel headers are required.
- A non-root user with
sudoprivileges. bpftrace needs root (orCAP_BPFplusCAP_PERFMON) to load programs. - Basic familiarity with Linux processes and system calls.
Step 1 - Installing bpftrace
The Ubuntu archive includes a maintained bpftrace package, which is the simplest option and is updated with the rest of the system:
sudo apt update
sudo apt install bpftrace
Check the installed version:
bpftrace --version
bpftrace v0.20.2
Now run a minimal program to confirm that the kernel accepts BPF programs. The BEGIN probe fires once when the program starts, and exit() stops it:
sudo bpftrace -e 'BEGIN { printf("bpftrace is working\n"); exit(); }'
Attaching 1 probe...
bpftrace is working
If you see Attaching 1 probe... followed by your message, bpftrace is ready.
Step 2 - Understanding probes, filters and maps
Every bpftrace program has the same shape:
probe /filter/ { action }
- The probe says where to attach, for example
tracepoint:syscalls:sys_enter_openat. - The optional filter decides whether the action runs, for example
/comm == "nginx"/. - The action is the code that runs, usually printing or updating a map.
These are the probe types you will use most:
| Probe type | Example | Fires on |
|---|---|---|
tracepoint | tracepoint:syscalls:sys_enter_openat | A stable kernel tracepoint |
kprobe / kretprobe | kprobe:tcp_connect | Entry or return of a kernel function |
uprobe / uretprobe | uretprobe:/bin/bash:readline | Entry or return of a user-space function |
profile | profile:hz:99 | A timer on every CPU, for sampling |
interval | interval:s:5 | A timer on one CPU, for periodic output |
BEGIN / END | END | Program start and exit |
Prefer tracepoints when one exists: their names and arguments are a stable interface, while kernel function names used by kprobes can change or be inlined between kernel versions.
Common built-in variables are pid, tid, comm (process name), uid, cpu, nsecs (timestamp in nanoseconds), retval (return value in kretprobe/uretprobe), kstack/ustack (stack traces) and args (tracepoint arguments).
Variables that start with @ are maps. They live in the kernel and are printed automatically when the program exits. Functions such as count(), sum() and hist() aggregate into them.
To find probes, use -l with a wildcard. For example, list the tracepoints for open system calls:
sudo bpftrace -l 'tracepoint:syscalls:sys_enter_open*'
tracepoint:syscalls:sys_enter_open
tracepoint:syscalls:sys_enter_open_by_handle_at
tracepoint:syscalls:sys_enter_open_tree
tracepoint:syscalls:sys_enter_openat
tracepoint:syscalls:sys_enter_openat2
Add -v to see the arguments a tracepoint exposes:
sudo bpftrace -lv 'tracepoint:syscalls:sys_enter_openat'
tracepoint:syscalls:sys_enter_openat
int __syscall_nr
int dfd
const char * filename
int flags
umode_t mode
Those field names are what you access through args in the next step.
Step 3 - Running essential one-liners
Each one-liner below runs until you press Ctrl+C, then prints its maps. Run them while the server is doing real work so the output is meaningful.
Count system calls by process
This answers "which process is hammering the kernel?":
sudo bpftrace -e 'tracepoint:raw_syscalls:sys_enter { @[comm] = count(); }'
Attaching 1 probe...
^C
@[sshd]: 214
@[systemd-journal]: 388
@[bpftrace]: 1022
@[nginx]: 18734
Trace which files are being opened
Print the process and file name for every openat() call:
sudo bpftrace -e 'tracepoint:syscalls:sys_enter_openat { printf("%-16s %s\n", comm, str(args->filename)); }'
Attaching 1 probe...
cron /etc/passwd
nginx /var/www/html/index.html
systemd-journal /proc/412/cmdline
str() copies the string from the traced process into bpftrace. Add a filter such as /comm == "nginx"/ before the braces to watch a single program.
Notenewer bpftrace releases also accept
args.filename. Theargs->filenameform used here works with the 0.20 package in Ubuntu 24.04.
Show disk I/O latency as a histogram
Block I/O tracepoints record when a request is issued to the device and when it completes. Storing the start time in a map keyed by device and sector lets you compute the latency of each request:
sudo bpftrace -e '
tracepoint:block:block_rq_issue { @start[args->dev, args->sector] = nsecs; }
tracepoint:block:block_rq_complete /@start[args->dev, args->sector]/ {
@usecs = hist((nsecs - @start[args->dev, args->sector]) / 1000);
delete(@start[args->dev, args->sector]);
}
END { clear(@start); }'
@usecs:
[64, 128) 12 |@@@ |
[128, 256) 184 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@|
[256, 512) 97 |@@@@@@@@@@@@@@@@@@@@@@@@@@@ |
[512, 1K) 21 |@@@@@ |
[1K, 2K) 3 | |
The END block clears the helper map so only the histogram is printed. Most requests here complete in 128 to 512 microseconds; a long tail in the millisecond buckets points to a saturated disk.
Count outgoing TCP connections by process
tcp_connect is the kernel function that starts an outgoing TCP connection:
sudo bpftrace -e 'kprobe:tcp_connect { @[comm] = count(); }'
@[curl]: 3
@[apt-get]: 8
@[php-fpm8.3]: 41
Capture commands typed in bash
A uretprobe on bash's readline() function returns each line users type in interactive bash sessions on the server:
sudo bpftrace -e 'uretprobe:/bin/bash:readline { printf("%-6d %s\n", pid, str(retval)); }'
Attaching 1 probe...
2311 ls -la /etc/nginx
2311 sudo systemctl reload nginx
This works because readline is an exported symbol in /bin/bash. The same technique works on any binary that exposes the function you care about; list candidates with sudo bpftrace -l 'uprobe:/path/to/binary:*'.
Step 4 - Sampling CPU stacks for a flame graph
The profile probe fires at a fixed frequency on every CPU. Recording the stack at each sample shows where CPU time goes. The following command samples kernel and user stacks of one process at 99 Hz for 30 seconds and writes the result to a file. Replace your_pid with the process ID you want to profile (for example from pgrep -o nginx):
sudo bpftrace -o /tmp/stacks.txt -e '
profile:hz:99 /pid == your_pid/ { @[kstack, ustack, comm] = count(); }
interval:s:30 { exit(); }'
Sampling at 99 Hz instead of 100 Hz avoids running in lockstep with other periodic activity. When the command returns, turn the samples into an SVG with Brendan Gregg's FlameGraph scripts:
git clone https://github.com/brendangregg/FlameGraph ~/FlameGraph
~/FlameGraph/stackcollapse-bpftrace.pl /tmp/stacks.txt | ~/FlameGraph/flamegraph.pl > ~/flamegraph.svg
Copy flamegraph.svg to your workstation (for example with scp) and open it in a browser. The widest boxes are the functions that consumed the most CPU.
Tipuser-space stacks are only readable when the program keeps frame pointers or debug symbols. Ubuntu 24.04 builds its packages with frame pointers, so stacks for distribution binaries are usually complete.
Step 5 - Writing a reusable bpftrace script
When a one-liner grows, save it as a .bt file. The script below measures read() latency per process and reports reads slower than 10 ms as they happen.
Create the file:
sudo nano /usr/local/bin/readlatency.bt
Add the following content:
#!/usr/bin/env bpftrace
/*
* readlatency.bt: read() latency per process, reports reads slower than 10 ms.
*/
BEGIN
{
printf("Tracing read() latency. Press Ctrl+C to stop.\n");
}
tracepoint:syscalls:sys_enter_read
{
@start[tid] = nsecs;
}
tracepoint:syscalls:sys_exit_read
/@start[tid]/
{
$lat = nsecs - @start[tid];
delete(@start[tid]);
if (args->ret > 0) {
@usecs[comm] = hist($lat / 1000);
@bytes[comm] = sum(args->ret);
}
if ($lat > 10000000) {
printf("slow read: pid=%d comm=%s %d ms\n", pid, comm, $lat / 1000000);
}
}
END
{
clear(@start);
}
On a syscall exit tracepoint the return value is args->ret, not retval, which only exists for kretprobe and uretprobe. Keying @start by tid keeps concurrent threads from overwriting each other.
Make the script executable and run it:
sudo chmod +x /usr/local/bin/readlatency.bt
sudo /usr/local/bin/readlatency.bt
Attaching 3 probes...
Tracing read() latency. Press Ctrl+C to stop.
slow read: pid=5120 comm=mysqld 14 ms
^C
@bytes[sshd]: 18240
@bytes[mysqld]: 9437184
@usecs[mysqld]:
[1] 802 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@|
[2, 4) 311 |@@@@@@@@@@@@@@@@@@@@ |
...
The Ubuntu package also ships ready-made tools such as opensnoop.bt, biolatency.bt and tcpconnect.bt. List them with:
dpkg -L bpftrace | grep '\.bt$'
Reading those scripts is one of the fastest ways to learn idiomatic bpftrace.
Troubleshooting
ERROR: kprobe target 'function_name' not found or a failure to attach. The kernel function does not exist in your kernel or was inlined by the compiler. Search for it with sudo bpftrace -l 'kprobe:*name*' and, whenever possible, switch to a tracepoint from sudo bpftrace -l 'tracepoint:*', which is stable across kernel versions.
ERROR: Could not resolve symbol on a uprobe. The binary is stripped or does not export the function. Check the available symbols with sudo bpftrace -l 'uprobe:/path/to/binary:*' and install the matching debug symbols package if you need internal functions.
Lost N events in the output. You are printing more events per second than bpftrace can read. Add a filter (/pid == 1234/, /comm == "nginx"/) or aggregate into a map with count() or hist() instead of calling printf() for every event.
Permission errors. Run bpftrace with sudo. Inside containers, the container needs the CAP_BPF, CAP_PERFMON and CAP_SYS_RESOURCE capabilities, or you should trace from the host.
Conclusion
You installed bpftrace on Ubuntu 24.04, learned how probes, filters and maps fit together, and used one-liners to find busy processes, opened files, disk latency, outgoing connections and CPU hot spots. You also wrote a reusable script with per-thread state. As next steps, read the bundled .bt tools under the package directory, trace scheduler behaviour with the tracepoint:sched:* probes, and combine bpftrace with perf or flame graphs when you investigate production latency.
