When a program is slow but CPU usage looks normal, it is usually waiting on something: disk, network, locks or files that do not exist. strace shows every system call a process makes (opening files, reading sockets, sleeping) and how long each one takes, and ltrace does the same for calls into shared libraries such as malloc or strlen. In this tutorial you will use both on Ubuntu 24.04 to count system calls, time individual calls, attach to a running service, and understand where ltrace works and where it does not.

Prerequisites

To follow this tutorial, you will need:

  • A server running Ubuntu 24.04 LTS, for example a CubePath VPS.
  • A non-root user with sudo privileges.
  • Optionally, a running service to inspect, such as Nginx or PHP-FPM.

Step 1 - Installing strace and ltrace

Install both tools, plus a compiler you will use in Step 6 to build a small test program:

sudo apt update
sudo apt install strace ltrace build-essential

Verify the installation:

strace -V | head -n 1
strace -- version 6.8

Ubuntu restricts which processes a user may trace (the Yama ptrace_scope setting is 1). You can trace programs you start yourself without sudo, but attaching to an already running process needs sudo.

Step 2 - Counting system calls with strace -c

The best first step is a summary: which system calls does the program make, how many, and how much time do they take? The -c option prints a table instead of every call.

This example shows a classic performance problem, tiny I/O operations. Copy 1 MB with dd using a 1-byte block size:

strace -c dd if=/dev/zero of=/tmp/test.bin bs=1 count=1000000

The output of dd is followed by the summary (trimmed here):

% time     seconds  usecs/call     calls    errors syscall
------ ----------- ----------- --------- --------- ----------------
 50.89    1.935213           1   1000000           write
 49.08    1.866402           1   1000003           read
  0.01    0.000401          13        30        10 openat
...
------ ----------- ----------- --------- --------- ----------------
100.00    3.802511           1   2000140        14 total

Two million system calls to move 1 MB. Now repeat with a 1 MB block size:

strace -c dd if=/dev/zero of=/tmp/test.bin bs=1M count=1
% time     seconds  usecs/call     calls    errors syscall
------ ----------- ----------- --------- --------- ----------------
 71.24    0.000703         703         1           write
 ...
------ ----------- ----------- --------- --------- ----------------
100.00    0.000987           7       136        14 total

The same work now needs about 140 calls. The same pattern (a program reading or writing in tiny chunks, or calling stat thousands of times) is behind many real slowdowns, and strace -c exposes it in seconds.

By default -c measures time spent inside the kernel. Add -w to measure wall-clock time instead, which includes time the process spent blocked, for example waiting on a slow network or disk:

strace -c -w curl -s -o /dev/null https://www.ubuntu.com/

Remove the test file:

rm /tmp/test.bin

Step 3 - Timing individual calls

Once you know which calls matter, look at them one by one. These options are the most useful for performance work:

OptionEffect
-fAlso trace child processes and threads. Almost always needed for servers.
-ttPrefix every line with a wall-clock timestamp with microseconds.
-TAppend the time spent in each call, as <seconds>.
-e trace=...Only trace a class of calls, for example %file, %network, %process, or a list such as openat,read,write.
-yShow the file or socket behind each file descriptor.
-o fileWrite the trace to a file instead of the terminal.

For example, trace only network calls made by curl and see how long each takes:

strace -f -tt -T -e trace=%network curl -s -o /dev/null https://www.ubuntu.com/
10:21:04.118203 socket(AF_INET, SOCK_STREAM, IPPROTO_TCP) = 5 <0.000031>
10:21:04.118412 connect(5, {sa_family=AF_INET, sin_port=htons(443), sin_addr=inet_addr("185.125.190.20")}, 16) = -1 EINPROGRESS (Operation now in progress) <0.000089>
10:21:04.142950 getsockopt(5, SOL_SOCKET, SO_ERROR, [0], [4]) = 0 <0.000012>
...

The gap between the connect timestamp and the next call is the TCP handshake time. A long gap before connect usually points to slow DNS resolution.

Step 4 - Finding failed and wasted calls

Failed calls cost time too. A frequent case is a program searching many directories for a file that is only in the last one. The -Z option prints only calls that failed:

strace -f -Z -e trace=%file python3 -c 'import json' 2>&1 | head -n 5
openat(AT_FDCWD, "/usr/lib/python312.zip", O_RDONLY|O_CLOEXEC) = -1 ENOENT (No such file or directory)
newfstatat(AT_FDCWD, "/usr/lib/python312.zip", 0x7ffc3b8e1d40, 0) = -1 ENOENT (No such file or directory)
...

Count them to judge the scale:

strace -f -Z -e trace=%file -o /tmp/failed.trace python3 -c 'import json'
wc -l /tmp/failed.trace

A handful of failures is normal. Thousands usually mean a long include path, an autoloader probing many directories, or a missing cache (for example a PHP application without OPcache enabled), and they are worth fixing.

Step 5 - Attaching to a running process

To see what a running service is doing right now, attach to it with -p. First find the process ID. For Nginx, trace a worker process, not the master:

pgrep -f 'nginx: worker'
1234
1235

Attach to one worker for about ten seconds while it handles traffic, recording timing for each call. timeout sends a signal to strace after 10 seconds, which detaches it cleanly and leaves the worker running:

sudo timeout 10 strace -f -tt -T -y -p 1234 -o /tmp/worker.trace

List the 10 slowest calls. The grep keeps only complete lines that end with a duration, and sort orders them by that duration:

grep -E '<[0-9]+\.[0-9]+>$' /tmp/worker.trace \
  | awk -F'<' '{ t = $NF; sub(/>$/, "", t); print t, $0 }' \
  | sort -rn | head -n 10
0.431022 1234 10:30:12.004113 epoll_wait(8, [{events=EPOLLIN, data={u32=..., u64=...}}], 512, 60000) = 1 <0.431022>
0.052117 1234 10:30:14.212906 read(15</var/www/html/big.json>, "..."..., 65536) = 65536 <0.052117>
...

Interpret the results carefully. Long epoll_wait, poll or accept calls are an idle worker waiting for work, which is normal. A slow read from a file on disk, a slow connect to a backend, or a futex that blocks for a long time are the interesting ones: they tell you the service is waiting on storage, on an upstream, or on a lock.

For a quick overview of the same process instead of a full log, use the summary mode:

sudo timeout 10 strace -c -f -p 1234

Step 6 - Tracing library calls with ltrace

ltrace intercepts calls from a program into shared libraries, for example malloc, free, strlen or snprintf in the C library. It is useful when a program spends CPU time in user space rather than in the kernel.

Create a small program that allocates, formats and frees a string 100,000 times:

nano demo.c
#include <stdio.h>
#include <stdlib.h>
#include <string.h>

int main(void) {
    size_t total = 0;
    for (int i = 0; i < 100000; i++) {
        char *buf = malloc(64);
        snprintf(buf, 64, "item-%d", i);
        total += strlen(buf);
        free(buf);
    }
    printf("%zu\n", total);
    return 0;
}

Compile it without CET and with lazy binding, so ltrace can hook its library calls:

gcc -O0 -fcf-protection=none -Wl,-z,lazy -o demo demo.c

Get a summary of library calls with -c:

ltrace -c ./demo
988890
% time     seconds  usecs/call     calls      function
------ ----------- ----------- --------- --------------------
 27.61    1.519784          15    100000 snprintf
 24.34    1.339871          13    100000 malloc
 24.11    1.327218          13    100000 free
 23.93    1.317266          13    100000 strlen
  0.01    0.000392         392         1 printf
------ ----------- ----------- --------- --------------------
100.00    5.504531                400001 total

The program itself finishes in a few milliseconds, but takes several seconds under ltrace. That overhead is why absolute times from ltrace are not meaningful: use the call counts and relative percentages. Here, 100,000 malloc/free pairs for a buffer that could be allocated once is the obvious fix.

Other useful options:

  • ltrace -e malloc+free ./demo traces only the listed functions.
  • ltrace -T ./demo shows the time spent in each call.
  • ltrace -S ./demo shows system calls together with library calls.
  • sudo ltrace -c -p PID attaches to a running process, subject to the same limitations described above.

Step 7 - Choosing the right tool

strace and ltrace answer "what is the program asking for", not "where is the CPU going". Use this table to pick a tool:

SymptomTool
Slow, but low CPU usage (waiting on disk, network, locks)strace -c -w, then strace -T on the calls that dominate
Errors or missing files at startupstrace -f -Z -e trace=%file
High CPU in the kernel (sy in top)strace -c to find which calls
High CPU in user space (us in top)perf top / perf record, or ltrace -c for your own builds
Tracing a busy production service for longer periodsperf trace, which has much lower overhead than strace

Troubleshooting

  • strace: attach: ptrace(PTRACE_SEIZE, 1234): Operation not permitted: attaching to a process you did not start requires sudo on Ubuntu. Inside containers, ptrace may also be blocked by the container runtime's security profile.
  • Lines ending in <unfinished ...> and <... read resumed>: with -f, a call from one thread was interrupted by output from another. Both lines belong to the same call; the duration appears on the resumed line.
  • ltrace prints only the program's output and no calls: the binary was built with CET or immediate binding, as explained in Step 6. Use perf for that binary.
  • The traced service became slow or timed out: tracing overhead. Filter with -e trace=..., trace a single worker, and keep sessions short with timeout.

Conclusion

You used strace -c to spot an excessive number of system calls, timed individual network and file calls, found failed lookups, attached safely to a running service, and traced library calls with ltrace on a program built for it. When the problem turns out to be CPU time rather than waiting, continue with perf to profile hot functions and generate flame graphs, and use a load testing tool such as wrk to reproduce the slowdown while you trace.