When a program is slow but CPU usage looks normal, it is usually waiting on something: disk, network, locks or files that do not exist. strace shows every system call a process makes (opening files, reading sockets, sleeping) and how long each one takes, and ltrace does the same for calls into shared libraries such as malloc or strlen. In this tutorial you will use both on Ubuntu 24.04 to count system calls, time individual calls, attach to a running service, and understand where ltrace works and where it does not.
Prerequisites
To follow this tutorial, you will need:
- A server running Ubuntu 24.04 LTS, for example a CubePath VPS.
- A non-root user with
sudoprivileges. - Optionally, a running service to inspect, such as Nginx or PHP-FPM.
WarningBoth tools stop the traced process on every call they record, which can slow it down by 10x or more. Trace production processes only for a few seconds at a time, and filter to the calls you care about.
Step 1 - Installing strace and ltrace
Install both tools, plus a compiler you will use in Step 6 to build a small test program:
sudo apt update
sudo apt install strace ltrace build-essential
Verify the installation:
strace -V | head -n 1
strace -- version 6.8
Ubuntu restricts which processes a user may trace (the Yama ptrace_scope setting is 1). You can trace programs you start yourself without sudo, but attaching to an already running process needs sudo.
Step 2 - Counting system calls with strace -c
The best first step is a summary: which system calls does the program make, how many, and how much time do they take? The -c option prints a table instead of every call.
This example shows a classic performance problem, tiny I/O operations. Copy 1 MB with dd using a 1-byte block size:
strace -c dd if=/dev/zero of=/tmp/test.bin bs=1 count=1000000
The output of dd is followed by the summary (trimmed here):
% time seconds usecs/call calls errors syscall
------ ----------- ----------- --------- --------- ----------------
50.89 1.935213 1 1000000 write
49.08 1.866402 1 1000003 read
0.01 0.000401 13 30 10 openat
...
------ ----------- ----------- --------- --------- ----------------
100.00 3.802511 1 2000140 14 total
Two million system calls to move 1 MB. Now repeat with a 1 MB block size:
strace -c dd if=/dev/zero of=/tmp/test.bin bs=1M count=1
% time seconds usecs/call calls errors syscall
------ ----------- ----------- --------- --------- ----------------
71.24 0.000703 703 1 write
...
------ ----------- ----------- --------- --------- ----------------
100.00 0.000987 7 136 14 total
The same work now needs about 140 calls. The same pattern (a program reading or writing in tiny chunks, or calling stat thousands of times) is behind many real slowdowns, and strace -c exposes it in seconds.
By default -c measures time spent inside the kernel. Add -w to measure wall-clock time instead, which includes time the process spent blocked, for example waiting on a slow network or disk:
strace -c -w curl -s -o /dev/null https://www.ubuntu.com/
Remove the test file:
rm /tmp/test.bin
Step 3 - Timing individual calls
Once you know which calls matter, look at them one by one. These options are the most useful for performance work:
| Option | Effect |
|---|---|
-f | Also trace child processes and threads. Almost always needed for servers. |
-tt | Prefix every line with a wall-clock timestamp with microseconds. |
-T | Append the time spent in each call, as <seconds>. |
-e trace=... | Only trace a class of calls, for example %file, %network, %process, or a list such as openat,read,write. |
-y | Show the file or socket behind each file descriptor. |
-o file | Write the trace to a file instead of the terminal. |
For example, trace only network calls made by curl and see how long each takes:
strace -f -tt -T -e trace=%network curl -s -o /dev/null https://www.ubuntu.com/
10:21:04.118203 socket(AF_INET, SOCK_STREAM, IPPROTO_TCP) = 5 <0.000031>
10:21:04.118412 connect(5, {sa_family=AF_INET, sin_port=htons(443), sin_addr=inet_addr("185.125.190.20")}, 16) = -1 EINPROGRESS (Operation now in progress) <0.000089>
10:21:04.142950 getsockopt(5, SOL_SOCKET, SO_ERROR, [0], [4]) = 0 <0.000012>
...
The gap between the connect timestamp and the next call is the TCP handshake time. A long gap before connect usually points to slow DNS resolution.
Step 4 - Finding failed and wasted calls
Failed calls cost time too. A frequent case is a program searching many directories for a file that is only in the last one. The -Z option prints only calls that failed:
strace -f -Z -e trace=%file python3 -c 'import json' 2>&1 | head -n 5
openat(AT_FDCWD, "/usr/lib/python312.zip", O_RDONLY|O_CLOEXEC) = -1 ENOENT (No such file or directory)
newfstatat(AT_FDCWD, "/usr/lib/python312.zip", 0x7ffc3b8e1d40, 0) = -1 ENOENT (No such file or directory)
...
Count them to judge the scale:
strace -f -Z -e trace=%file -o /tmp/failed.trace python3 -c 'import json'
wc -l /tmp/failed.trace
A handful of failures is normal. Thousands usually mean a long include path, an autoloader probing many directories, or a missing cache (for example a PHP application without OPcache enabled), and they are worth fixing.
Step 5 - Attaching to a running process
To see what a running service is doing right now, attach to it with -p. First find the process ID. For Nginx, trace a worker process, not the master:
pgrep -f 'nginx: worker'
1234
1235
Attach to one worker for about ten seconds while it handles traffic, recording timing for each call. timeout sends a signal to strace after 10 seconds, which detaches it cleanly and leaves the worker running:
sudo timeout 10 strace -f -tt -T -y -p 1234 -o /tmp/worker.trace
List the 10 slowest calls. The grep keeps only complete lines that end with a duration, and sort orders them by that duration:
grep -E '<[0-9]+\.[0-9]+>$' /tmp/worker.trace \
| awk -F'<' '{ t = $NF; sub(/>$/, "", t); print t, $0 }' \
| sort -rn | head -n 10
0.431022 1234 10:30:12.004113 epoll_wait(8, [{events=EPOLLIN, data={u32=..., u64=...}}], 512, 60000) = 1 <0.431022>
0.052117 1234 10:30:14.212906 read(15</var/www/html/big.json>, "..."..., 65536) = 65536 <0.052117>
...
Interpret the results carefully. Long epoll_wait, poll or accept calls are an idle worker waiting for work, which is normal. A slow read from a file on disk, a slow connect to a backend, or a futex that blocks for a long time are the interesting ones: they tell you the service is waiting on storage, on an upstream, or on a lock.
For a quick overview of the same process instead of a full log, use the summary mode:
sudo timeout 10 strace -c -f -p 1234
Step 6 - Tracing library calls with ltrace
ltrace intercepts calls from a program into shared libraries, for example malloc, free, strlen or snprintf in the C library. It is useful when a program spends CPU time in user space rather than in the kernel.
ImportantOn Ubuntu, most system binaries are built with Intel CET (
-fcf-protection) and immediate symbol binding.ltrace0.7.3 often cannot hook calls in those binaries and prints little or nothing. It works reliably on programs you build yourself with the flags shown below. For profiling packaged software, useperfinstead.
Create a small program that allocates, formats and frees a string 100,000 times:
nano demo.c
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
int main(void) {
size_t total = 0;
for (int i = 0; i < 100000; i++) {
char *buf = malloc(64);
snprintf(buf, 64, "item-%d", i);
total += strlen(buf);
free(buf);
}
printf("%zu\n", total);
return 0;
}
Compile it without CET and with lazy binding, so ltrace can hook its library calls:
gcc -O0 -fcf-protection=none -Wl,-z,lazy -o demo demo.c
Get a summary of library calls with -c:
ltrace -c ./demo
988890
% time seconds usecs/call calls function
------ ----------- ----------- --------- --------------------
27.61 1.519784 15 100000 snprintf
24.34 1.339871 13 100000 malloc
24.11 1.327218 13 100000 free
23.93 1.317266 13 100000 strlen
0.01 0.000392 392 1 printf
------ ----------- ----------- --------- --------------------
100.00 5.504531 400001 total
The program itself finishes in a few milliseconds, but takes several seconds under ltrace. That overhead is why absolute times from ltrace are not meaningful: use the call counts and relative percentages. Here, 100,000 malloc/free pairs for a buffer that could be allocated once is the obvious fix.
Other useful options:
ltrace -e malloc+free ./demotraces only the listed functions.ltrace -T ./demoshows the time spent in each call.ltrace -S ./demoshows system calls together with library calls.sudo ltrace -c -p PIDattaches to a running process, subject to the same limitations described above.
Step 7 - Choosing the right tool
strace and ltrace answer "what is the program asking for", not "where is the CPU going". Use this table to pick a tool:
| Symptom | Tool |
|---|---|
| Slow, but low CPU usage (waiting on disk, network, locks) | strace -c -w, then strace -T on the calls that dominate |
| Errors or missing files at startup | strace -f -Z -e trace=%file |
High CPU in the kernel (sy in top) | strace -c to find which calls |
High CPU in user space (us in top) | perf top / perf record, or ltrace -c for your own builds |
| Tracing a busy production service for longer periods | perf trace, which has much lower overhead than strace |
Troubleshooting
strace: attach: ptrace(PTRACE_SEIZE, 1234): Operation not permitted: attaching to a process you did not start requiressudoon Ubuntu. Inside containers,ptracemay also be blocked by the container runtime's security profile.- Lines ending in
<unfinished ...>and<... read resumed>: with-f, a call from one thread was interrupted by output from another. Both lines belong to the same call; the duration appears on theresumedline. ltraceprints only the program's output and no calls: the binary was built with CET or immediate binding, as explained in Step 6. Useperffor that binary.- The traced service became slow or timed out: tracing overhead. Filter with
-e trace=..., trace a single worker, and keep sessions short withtimeout.
Conclusion
You used strace -c to spot an excessive number of system calls, timed individual network and file calls, found failed lookups, attached safely to a running service, and traced library calls with ltrace on a program built for it. When the problem turns out to be CPU time rather than waiting, continue with perf to profile hot functions and generate flame graphs, and use a load testing tool such as wrk to reproduce the slowdown while you trace.
