perf is the profiler built into the Linux kernel. It samples what the CPU is executing many times per second, with very low overhead, and tells you which functions, in your application or in the kernel, consume the most CPU time. It can also count hardware events such as instructions, cache misses and branch mispredictions. In this tutorial you will install perf on Ubuntu 24.04, measure a program with perf stat, find its hot function with perf record and perf report, profile a running service, and turn the result into a flame graph.

Prerequisites

To follow this tutorial, you will need:

  • A server running Ubuntu 24.04 LTS, for example a CubePath VPS.
  • A non-root user with sudo privileges.
  • About 200 MB of free memory for the example program.
  • Git, which is installed by default on Ubuntu Server, for the flame graph tools in Step 6.

Step 1 - Installing perf

On Ubuntu, perf is shipped in the kernel tools packages, and the version must match the running kernel. Install the common package and the one for your exact kernel:

sudo apt update
sudo apt install linux-tools-common linux-tools-$(uname -r)

Verify the installation:

perf --version
perf version 6.8.12

Ubuntu sets kernel.perf_event_paranoid to 4, which blocks profiling for regular users:

cat /proc/sys/kernel/perf_event_paranoid
4

Run perf with sudo throughout this tutorial rather than lowering that setting on a shared server.

Step 2 - Building an example program

To learn the workflow it helps to profile a program whose bottleneck you know. This one sums the same 128 MB array twice: strided_sum jumps 64 bytes (one cache line) between reads, so almost every read misses the CPU cache, while linear_sum reads memory in order.

nano hotspot.c
#include <stdio.h>
#include <stdlib.h>

__attribute__((noinline))
static unsigned long strided_sum(const int *a, size_t n) {
    unsigned long s = 0;
    for (size_t start = 0; start < 16; start++)
        for (size_t i = start; i < n; i += 16)
            s += a[i];
    return s;
}

__attribute__((noinline))
static unsigned long linear_sum(const int *a, size_t n) {
    unsigned long s = 0;
    for (size_t i = 0; i < n; i++)
        s += a[i];
    return s;
}

int main(void) {
    size_t n = 32UL * 1024 * 1024;   /* 32M ints = 128 MB */
    int *a = malloc(n * sizeof *a);
    if (!a)
        return 1;
    for (size_t i = 0; i < n; i++)
        a[i] = (int)(i & 0xff);

    unsigned long r = 0;
    for (int k = 0; k < 5; k++) {
        r += strided_sum(a, n);
        r += linear_sum(a, n);
    }
    printf("%lu\n", r);
    free(a);
    return 0;
}

Compile it with debug symbols (-g) and frame pointers, so perf can show function names and full call stacks:

sudo apt install build-essential
gcc -O2 -g -fno-omit-frame-pointer -o hotspot hotspot.c

The noinline attributes keep the two functions separate at -O2, so they show up individually in the profile.

Step 3 - Counting events with perf stat

perf stat runs a command and reports counters for the whole run. It is the quickest way to see whether a program is CPU-bound and how efficiently it uses the CPU:

sudo perf stat ./hotspot
 Performance counter stats for './hotspot':

          2,487.31 msec task-clock                #    0.999 CPUs utilized
                 9      context-switches          #    3.618 /sec
                 0      cpu-migrations            #    0.000 /sec
            32,832      page-faults               #   13.200 K/sec
     8,912,405,127      cycles                    #    3.583 GHz
     5,873,112,904      instructions              #    0.66  insn per cycle
       763,528,410      branches                  #  306.970 M/sec
           342,118      branch-misses             #    0.04% of all branches

       2.489603847 seconds time elapsed

Key lines:

  • CPUs utilized close to 1.0 means the program was busy on one CPU the whole time (CPU-bound). A low value means it was mostly waiting, and strace is the better tool.
  • insn per cycle (IPC) below 1 on a modern CPU often means the CPU is stalled waiting on memory.

Ask for cache events explicitly to confirm the memory problem:

sudo perf stat -e cache-references,cache-misses ./hotspot
       412,907,331      cache-references
       291,530,016      cache-misses              #   70.60% of all cache refs

Step 4 - Finding the hot function with perf record

perf record samples the call stack at a fixed frequency and writes the samples to perf.data. Use -F 99 (99 samples per second per CPU, which avoids aliasing with timers that fire at round frequencies) and -g to record call stacks:

sudo perf record -F 99 -g -- ./hotspot
[ perf record: Woken up 1 times to write data ]
[ perf record: Captured and wrote 0.061 MB perf.data (247 samples) ]

Show the report as text. --no-children attributes each sample only to the function that was actually executing, which is what you want to find hot spots:

sudo perf report --stdio --no-children | grep -v '^$' | head -n 20
# Overhead  Command  Shared Object      Symbol
    82.19%  hotspot  hotspot            [.] strided_sum
            |
            ---__libc_start_call_main
               main
               strided_sum
     9.31%  hotspot  hotspot            [.] linear_sum
     6.07%  hotspot  hotspot            [.] main
     1.62%  hotspot  [kernel.kallsyms]  [k] clear_page_erms
...

strided_sum takes about 82% of the CPU time even though it does exactly the same additions as linear_sum. [.] marks user-space code and [k] marks kernel code; here the kernel time is page zeroing for the 128 MB allocation (main fills the array).

Run sudo perf report without --stdio for an interactive view: use the arrow keys to select a function, Enter to expand it, and a to annotate it and see which source lines and instructions are hot.

Once you have found the hotspot, the fix follows: rewrite the loop to read memory sequentially, and the program runs several times faster.

Step 5 - Profiling a running service

The same technique works on any process that is already running, such as a web server, database or application worker. Find its process ID first, for example for MySQL:

pgrep -x mysqld
1523

Record 30 seconds of samples from that process while it is under its normal (or benchmark) load. sleep 30 only defines how long the recording lasts:

sudo perf record -F 99 -g -p 1523 -- sleep 30
sudo perf report --stdio --no-children | head -n 40

To see everything the server is doing, including the kernel and all processes, use -a (all CPUs) instead of -p:

sudo perf record -F 99 -g -a -- sleep 30

For a live view that updates every few seconds, similar to top but per function, use perf top:

sudo perf top

A few things to know when profiling real software:

  • Ubuntu 24.04 builds its packages with frame pointers enabled, so call stacks of packaged software (Nginx, MySQL, PostgreSQL, Python) are usually complete without extra work.
  • For your own compiled code, build with -g -fno-omit-frame-pointer, or record with --call-graph dwarf, which works without frame pointers but produces much larger perf.data files.
  • Interpreted and JIT-compiled languages need help to show their own function names. Python 3.12 supports python3 -X perf your_script.py, and Node.js supports node --perf-basic-prof your_app.js. Without these, you only see the interpreter's C functions.

Step 6 - Generating a flame graph

A flame graph shows all sampled stacks at once: each box is a function, its width is the share of CPU time, and callers sit below callees. It is the fastest way to understand a profile with many code paths.

Clone Brendan Gregg's FlameGraph scripts:

git clone --depth 1 https://github.com/brendangregg/FlameGraph.git ~/FlameGraph

From the directory that contains the perf.data file from Step 4 (or Step 5), convert it into text stacks, fold them, and render an SVG:

sudo perf script > out.perf
~/FlameGraph/stackcollapse-perf.pl out.perf > out.folded
~/FlameGraph/flamegraph.pl out.folded > flame.svg

Copy flame.svg to your computer and open it in a web browser. Run this on your local machine, adjusting the path if you did not work in your home directory:

scp your_user@your_server_ip:~/flame.svg .

In the browser, click a box to zoom into that part of the stack and use the search box at the top right to highlight a function. Look for wide plateaus at the top of the graph: those are the functions where the CPU actually spends its time. For the example program you will see one wide strided_sum tower on top of main.

Clean up the temporary files when you are done:

sudo rm -f perf.data perf.data.old out.perf out.folded

Troubleshooting

  • Access to performance monitoring and observability operations is limited: perf was run without sudo while perf_event_paranoid is 4. Run it with sudo.
  • WARNING: perf not found for kernel 6.8.0-xx: the tools package for the running kernel is missing. Install linux-tools-$(uname -r).
  • Functions appear as [unknown] or hex addresses: the binary has no symbols or no frame pointers. Rebuild with -g -fno-omit-frame-pointer, record with --call-graph dwarf, or install the package's debug symbols from Ubuntu's ddebs repository.
  • <not supported> for hardware events: the virtual machine does not expose the CPU's performance counters. Use software events (task-clock, cpu-clock) for counting; sampling with perf record still works.
  • The profile is dominated by idle or sleep: the process was not busy while you recorded. Generate load (for example with wrk) during the recording, or profile the whole system with -a.

Conclusion

You installed perf, used perf stat to confirm a CPU-bound workload with poor cache behavior, located the exact function responsible with perf record and perf report, profiled a running service, and visualized the profile as a flame graph. Next, combine perf with a load generator such as wrk to profile your web stack under realistic traffic, use strace when a process is slow but not busy, and compare flame graphs before and after each optimization to confirm it helped.