On servers with more than one CPU socket, and on some single-socket CPUs such as AMD EPYC configured with several NUMA nodes per socket, memory is split into NUMA (Non-Uniform Memory Access) nodes. Each CPU core reaches the memory of its own node faster than the memory attached to another node. In this tutorial you will inspect the NUMA topology of an Ubuntu 24.04 server, measure the difference between local and remote memory, run processes bound to one node with numactl, pin a systemd service to a node and monitor where memory is really allocated.
Prerequisites
To follow this guide you need:
- A dedicated or bare-metal server running Ubuntu 24.04 LTS with at least two NUMA nodes, for example a dual-socket CubePath dedicated server.
- A non-root user with
sudoprivileges.
NoteMost virtual machines, including typical VPS plans, expose a single NUMA node. On such a system all memory is local, the commands in this guide run but show only
node 0, and NUMA pinning has no benefit.
Step 1 - Installing the NUMA tools
Install numactl (which also provides numastat), the text-only hwloc tools and mbw, a small memory bandwidth benchmark:
sudo apt update
sudo apt install numactl hwloc-nox mbw
Check that numactl works:
numactl --version
numactl 2.0.18
Step 2 - Inspecting the NUMA topology
Start with lscpu, which shows how many NUMA nodes exist and which CPUs belong to each:
lscpu | grep -i numa
NUMA node(s): 2
NUMA node0 CPU(s): 0-15,32-47
NUMA node1 CPU(s): 16-31,48-63
In this example each node has 16 physical cores with Hyper-Threading, so CPU 0 and CPU 32 are two threads of the same core. Note the CPU lists for your server: you will use them later.
numactl --hardware adds the memory size of each node and the relative distance between nodes:
numactl --hardware
available: 2 nodes (0-1)
node 0 cpus: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47
node 0 size: 128614 MB
node 0 free: 119305 MB
node 1 cpus: 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63
node 1 size: 129018 MB
node 1 free: 121877 MB
node distances:
node 0 1
0: 10 21
1: 21 10
The distance table is relative: 10 is local access, and 21 means accessing the other node costs roughly twice as much. For a fuller picture, including shared caches, use lstopo-no-graphics from hwloc:
lstopo-no-graphics --no-io
Machine (251GB total)
Package L#0
NUMANode L#0 (P#0 126GB)
L3 L#0 (40MB)
L2 L#0 (1024KB) + L1d L#0 (48KB) + L1i L#0 (32KB) + Core L#0
PU L#0 (P#0)
PU L#1 (P#32)
...
Step 3 - Measuring local and remote memory bandwidth
Before changing how applications run, see how much remote memory costs on your hardware. numactl --cpunodebind restricts a command to the CPUs of a node, and --membind forces its memory allocations onto a node.
Run mbw on node 0 with memory from node 0 (local). The arguments run the memcpy test (-t0) five times (-n 5) on 1024 MiB arrays, large enough to not fit in any CPU cache:
numactl --cpunodebind=0 --membind=0 mbw -q -n 5 -t0 1024 | grep AVG
AVG Method: MEMCPY Elapsed: 0.10412 MiB: 1024.00000 Copy: 9834.806 MiB/s
Now run it on node 0 with memory from node 1 (remote):
numactl --cpunodebind=0 --membind=1 mbw -q -n 5 -t0 1024 | grep AVG
AVG Method: MEMCPY Elapsed: 0.16873 MiB: 1024.00000 Copy: 6068.884 MiB/s
In this example remote access delivers about 40% less bandwidth, and latency is also higher. The exact gap depends on the CPU generation and the interconnect between sockets. This is the penalty that NUMA-aware deployment avoids.
Step 4 - Running a process with a NUMA policy
numactl starts a command with a CPU binding and a memory policy. The four memory policies are:
| Option | Behaviour |
|---|---|
--membind=N | Allocate only from node N. If it runs out of memory, the process can be killed by the OOM killer. |
--preferred=N | Prefer node N, fall back to other nodes when it is full. |
--interleave=all | Spread pages round-robin across all nodes, for even bandwidth. |
--localalloc | Allocate on the node of the CPU the thread is running on. This is the kernel default. |
You can see the effect of any combination by running numactl --show under it, which prints the policy it inherited:
numactl --cpunodebind=1 --preferred=1 numactl --show
policy: preferred
preferred node: 1
physcpubind: 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63
cpubind: 1
nodebind: 1
membind: 0 1
To run your own program on node 1 with local memory, prefix its command in the same way, replacing /usr/local/bin/your_app with the real path:
numactl --cpunodebind=1 --membind=1 /usr/local/bin/your_app
Which policy to choose depends on the workload:
- A process that fits in one node (a cache, an API worker, a game server): bind CPUs and memory to the same node. Running one instance per node, each bound to its own node, often scales better than one large instance spread across both.
- A process that needs more memory than one node has, or that shares a large memory area between all its threads (a database buffer pool, an in-memory analytics engine): use
--interleave=all, so no single node becomes a bandwidth hotspot and one node does not run out of memory first. - Everything else: leave the default. The kernel's automatic NUMA balancing moves pages closer to the threads that use them.
Step 5 - Pinning a systemd service to a NUMA node
For services, configure the binding in systemd instead of wrapping ExecStart in numactl. systemd 255 on Ubuntu 24.04 supports CPUAffinity=, NUMAPolicy= and NUMAMask=. Open a drop-in for your service, replacing your_app with the unit name:
sudo systemctl edit your_app.service
Add the following lines in the editor. CPUAffinity= takes the CPU list of node 0 you found in Step 2, and NUMAPolicy=bind with NUMAMask=0 restricts memory allocations to node 0:
[Service]
CPUAffinity=0-15 32-47
NUMAPolicy=bind
NUMAMask=0
For a database that should interleave its memory across all nodes instead, use:
[Service]
NUMAPolicy=interleave
NUMAMask=all
Restart the service to apply the drop-in:
sudo systemctl restart your_app.service
Check the CPU affinity of the main process:
PID=$(systemctl show -p MainPID --value your_app.service)
taskset -cp "$PID"
pid 21874's current affinity list: 0-15,32-47
Check the memory policy. Each line of /proc/PID/numa_maps describes one memory mapping and starts with the policy that applies to it:
sudo head -3 /proc/"$PID"/numa_maps
55d2c3a00000 bind:0 file=/usr/local/bin/your_app mapped=312 active=0 N0=312 kernelpagesize_kB=4
55d2c3b4a000 bind:0 anon=1024 dirty=1024 active=0 N0=1024 kernelpagesize_kB=4
7f9a10000000 bind:0 anon=65536 dirty=65536 active=0 N0=65536 kernelpagesize_kB=4
bind:0 confirms the policy, and N0= shows that the pages are on node 0. With interleaving you would see interleave:0-1 and pages split between N0= and N1=.
MySQL has its own switch for the InnoDB buffer pool, innodb_numa_interleave. It only exists when MySQL is built with NUMA support; check with sudo mysql -e "SHOW VARIABLES LIKE 'innodb_numa_interleave';". If the variable is not listed, use the systemd NUMAPolicy=interleave drop-in above for the mysql service instead.
Step 6 - Monitoring NUMA allocation
numastat without arguments shows system-wide counters per node since boot:
numastat
node0 node1
numa_hit 9812345678 8723456789
numa_miss 1234567 2345678
numa_foreign 2345678 1234567
interleave_hit 65432 65420
local_node 9811000000 8722000000
other_node 2579245 3802345
numa_hitcounts allocations that landed on the intended node.numa_misscounts allocations that landed on this node although another node was intended, usually because that node was full.other_nodecounts allocations on this node made by a process running on another node.
What matters is the trend: run numastat twice, a few minutes apart during normal load, and compare. A steadily growing numa_miss means a node is running out of memory and allocations are spilling over.
To see how the memory of one process is spread across nodes, use numastat -p with a PID or a process name:
numastat -p your_app
Per-node process memory usage (in MBs) for PID 21874 (your_app)
Node 0 Node 1 Total
--------------- --------------- ---------------
Huge 0.00 0.00 0.00
Heap 256.31 0.00 256.31
Stack 0.13 0.00 0.13
Private 15890.22 4.02 15894.24
---------------- --------------- --------------- ---------------
Total 16146.66 4.02 16150.68
For a service bound to node 0, almost all memory should be under Node 0. For the whole system per node, including free memory and page cache, use numastat -m.
Step 7 - Deciding on automatic NUMA balancing
The kernel's automatic NUMA balancing periodically samples memory accesses and migrates pages to the node of the thread that uses them. Check whether it is enabled:
cat /proc/sys/kernel/numa_balancing
1
Its activity is visible in /proc/vmstat:
grep -E 'numa_pages_migrated|numa_hint_faults ' /proc/vmstat
numa_hint_faults 48213377
numa_pages_migrated 3218842
For general workloads leave it enabled. If all latency-sensitive services are explicitly pinned as in Step 5, the sampling adds overhead without benefit, and you can disable it persistently:
echo 'kernel.numa_balancing = 0' | sudo tee /etc/sysctl.d/90-numa.conf
sudo sysctl --system
Measure your application before and after this change and keep the setting that performs better.
Troubleshooting
numactl --hardware shows available: 1 nodes (0): the server has a single NUMA node, or the BIOS has node interleaving enabled, which hides the topology from the OS. On dual-socket servers, disable "Node Interleaving" in the BIOS to expose both nodes. On AMD EPYC, the "NUMA nodes per socket" (NPS) setting controls how many nodes each socket exposes.
A process bound with --membind is killed by the OOM killer while the other node has free memory: strict binding does not fall back to other nodes. Use --preferred (or NUMAPolicy=preferred in systemd), or give the process less memory than the node has.
numastat -p reports Can't read /proc/PID/numa_maps: the process belongs to another user. Run the command with sudo.
Conclusion
You inspected the NUMA topology with lscpu, numactl and hwloc, measured the cost of remote memory, ran processes with explicit NUMA policies, pinned a systemd service to one node and checked where its memory really lives with numa_maps and numastat. As next steps, benchmark your real application with and without pinning, consider running one instance per NUMA node behind a load balancer, and check which node your network card is attached to with cat /sys/class/net/your_interface/device/numa_node so network-heavy services run on that node.
