perf is the profiler built into the Linux kernel. It samples what the CPU is executing many times per second, with very low overhead, and tells you which functions, in your application or in the kernel, consume the most CPU time. It can also count hardware events such as instructions, cache misses and branch mispredictions. In this tutorial you will install perf on Ubuntu 24.04, measure a program with perf stat, find its hot function with perf record and perf report, profile a running service, and turn the result into a flame graph.
Prerequisites
To follow this tutorial, you will need:
- A server running Ubuntu 24.04 LTS, for example a CubePath VPS.
- A non-root user with
sudoprivileges. - About 200 MB of free memory for the example program.
- Git, which is installed by default on Ubuntu Server, for the flame graph tools in Step 6.
Step 1 - Installing perf
On Ubuntu, perf is shipped in the kernel tools packages, and the version must match the running kernel. Install the common package and the one for your exact kernel:
sudo apt update
sudo apt install linux-tools-common linux-tools-$(uname -r)
Verify the installation:
perf --version
perf version 6.8.12
Ubuntu sets kernel.perf_event_paranoid to 4, which blocks profiling for regular users:
cat /proc/sys/kernel/perf_event_paranoid
4
Run perf with sudo throughout this tutorial rather than lowering that setting on a shared server.
NoteAfter a kernel upgrade and reboot,
perfprintsWARNING: perf not found for kernel .... Run theapt installcommand above again to get the tools for the new kernel.
Step 2 - Building an example program
To learn the workflow it helps to profile a program whose bottleneck you know. This one sums the same 128 MB array twice: strided_sum jumps 64 bytes (one cache line) between reads, so almost every read misses the CPU cache, while linear_sum reads memory in order.
nano hotspot.c
#include <stdio.h>
#include <stdlib.h>
__attribute__((noinline))
static unsigned long strided_sum(const int *a, size_t n) {
unsigned long s = 0;
for (size_t start = 0; start < 16; start++)
for (size_t i = start; i < n; i += 16)
s += a[i];
return s;
}
__attribute__((noinline))
static unsigned long linear_sum(const int *a, size_t n) {
unsigned long s = 0;
for (size_t i = 0; i < n; i++)
s += a[i];
return s;
}
int main(void) {
size_t n = 32UL * 1024 * 1024; /* 32M ints = 128 MB */
int *a = malloc(n * sizeof *a);
if (!a)
return 1;
for (size_t i = 0; i < n; i++)
a[i] = (int)(i & 0xff);
unsigned long r = 0;
for (int k = 0; k < 5; k++) {
r += strided_sum(a, n);
r += linear_sum(a, n);
}
printf("%lu\n", r);
free(a);
return 0;
}
Compile it with debug symbols (-g) and frame pointers, so perf can show function names and full call stacks:
sudo apt install build-essential
gcc -O2 -g -fno-omit-frame-pointer -o hotspot hotspot.c
The noinline attributes keep the two functions separate at -O2, so they show up individually in the profile.
Step 3 - Counting events with perf stat
perf stat runs a command and reports counters for the whole run. It is the quickest way to see whether a program is CPU-bound and how efficiently it uses the CPU:
sudo perf stat ./hotspot
Performance counter stats for './hotspot':
2,487.31 msec task-clock # 0.999 CPUs utilized
9 context-switches # 3.618 /sec
0 cpu-migrations # 0.000 /sec
32,832 page-faults # 13.200 K/sec
8,912,405,127 cycles # 3.583 GHz
5,873,112,904 instructions # 0.66 insn per cycle
763,528,410 branches # 306.970 M/sec
342,118 branch-misses # 0.04% of all branches
2.489603847 seconds time elapsed
Key lines:
CPUs utilizedclose to 1.0 means the program was busy on one CPU the whole time (CPU-bound). A low value means it was mostly waiting, andstraceis the better tool.insn per cycle(IPC) below 1 on a modern CPU often means the CPU is stalled waiting on memory.
Ask for cache events explicitly to confirm the memory problem:
sudo perf stat -e cache-references,cache-misses ./hotspot
412,907,331 cache-references
291,530,016 cache-misses # 70.60% of all cache refs
NoteMany virtual machines do not expose hardware counters to the guest. If
cycles,instructionsorcache-missesshow<not supported>, the hypervisor is hiding the CPU's performance counters. Software events such astask-clock,context-switchesandpage-faultsstill work, and CPU profiling in the next steps works too becauseperf recordfalls back to thecpu-clocktimer.
Step 4 - Finding the hot function with perf record
perf record samples the call stack at a fixed frequency and writes the samples to perf.data. Use -F 99 (99 samples per second per CPU, which avoids aliasing with timers that fire at round frequencies) and -g to record call stacks:
sudo perf record -F 99 -g -- ./hotspot
[ perf record: Woken up 1 times to write data ]
[ perf record: Captured and wrote 0.061 MB perf.data (247 samples) ]
Show the report as text. --no-children attributes each sample only to the function that was actually executing, which is what you want to find hot spots:
sudo perf report --stdio --no-children | grep -v '^$' | head -n 20
# Overhead Command Shared Object Symbol
82.19% hotspot hotspot [.] strided_sum
|
---__libc_start_call_main
main
strided_sum
9.31% hotspot hotspot [.] linear_sum
6.07% hotspot hotspot [.] main
1.62% hotspot [kernel.kallsyms] [k] clear_page_erms
...
strided_sum takes about 82% of the CPU time even though it does exactly the same additions as linear_sum. [.] marks user-space code and [k] marks kernel code; here the kernel time is page zeroing for the 128 MB allocation (main fills the array).
Run sudo perf report without --stdio for an interactive view: use the arrow keys to select a function, Enter to expand it, and a to annotate it and see which source lines and instructions are hot.
Once you have found the hotspot, the fix follows: rewrite the loop to read memory sequentially, and the program runs several times faster.
Step 5 - Profiling a running service
The same technique works on any process that is already running, such as a web server, database or application worker. Find its process ID first, for example for MySQL:
pgrep -x mysqld
1523
Record 30 seconds of samples from that process while it is under its normal (or benchmark) load. sleep 30 only defines how long the recording lasts:
sudo perf record -F 99 -g -p 1523 -- sleep 30
sudo perf report --stdio --no-children | head -n 40
To see everything the server is doing, including the kernel and all processes, use -a (all CPUs) instead of -p:
sudo perf record -F 99 -g -a -- sleep 30
For a live view that updates every few seconds, similar to top but per function, use perf top:
sudo perf top
A few things to know when profiling real software:
- Ubuntu 24.04 builds its packages with frame pointers enabled, so call stacks of packaged software (Nginx, MySQL, PostgreSQL, Python) are usually complete without extra work.
- For your own compiled code, build with
-g -fno-omit-frame-pointer, or record with--call-graph dwarf, which works without frame pointers but produces much largerperf.datafiles. - Interpreted and JIT-compiled languages need help to show their own function names. Python 3.12 supports
python3 -X perf your_script.py, and Node.js supportsnode --perf-basic-prof your_app.js. Without these, you only see the interpreter's C functions.
Step 6 - Generating a flame graph
A flame graph shows all sampled stacks at once: each box is a function, its width is the share of CPU time, and callers sit below callees. It is the fastest way to understand a profile with many code paths.
Clone Brendan Gregg's FlameGraph scripts:
git clone --depth 1 https://github.com/brendangregg/FlameGraph.git ~/FlameGraph
From the directory that contains the perf.data file from Step 4 (or Step 5), convert it into text stacks, fold them, and render an SVG:
sudo perf script > out.perf
~/FlameGraph/stackcollapse-perf.pl out.perf > out.folded
~/FlameGraph/flamegraph.pl out.folded > flame.svg
Copy flame.svg to your computer and open it in a web browser. Run this on your local machine, adjusting the path if you did not work in your home directory:
scp your_user@your_server_ip:~/flame.svg .
In the browser, click a box to zoom into that part of the stack and use the search box at the top right to highlight a function. Look for wide plateaus at the top of the graph: those are the functions where the CPU actually spends its time. For the example program you will see one wide strided_sum tower on top of main.
Clean up the temporary files when you are done:
sudo rm -f perf.data perf.data.old out.perf out.folded
Troubleshooting
Access to performance monitoring and observability operations is limited:perfwas run withoutsudowhileperf_event_paranoidis4. Run it withsudo.WARNING: perf not found for kernel 6.8.0-xx: the tools package for the running kernel is missing. Installlinux-tools-$(uname -r).- Functions appear as
[unknown]or hex addresses: the binary has no symbols or no frame pointers. Rebuild with-g -fno-omit-frame-pointer, record with--call-graph dwarf, or install the package's debug symbols from Ubuntu'sddebsrepository. <not supported>for hardware events: the virtual machine does not expose the CPU's performance counters. Use software events (task-clock,cpu-clock) for counting; sampling withperf recordstill works.- The profile is dominated by idle or
sleep: the process was not busy while you recorded. Generate load (for example withwrk) during the recording, or profile the whole system with-a.
Conclusion
You installed perf, used perf stat to confirm a CPU-bound workload with poor cache behavior, located the exact function responsible with perf record and perf report, profiled a running service, and visualized the profile as a flame graph. Next, combine perf with a load generator such as wrk to profile your web stack under realistic traffic, use strace when a process is slow but not busy, and compare flame graphs before and after each optimization to confirm it helped.
