On servers with more than one CPU socket, and on some single-socket CPUs such as AMD EPYC configured with several NUMA nodes per socket, memory is split into NUMA (Non-Uniform Memory Access) nodes. Each CPU core reaches the memory of its own node faster than the memory attached to another node. In this tutorial you will inspect the NUMA topology of an Ubuntu 24.04 server, measure the difference between local and remote memory, run processes bound to one node with numactl, pin a systemd service to a node and monitor where memory is really allocated.

Prerequisites

To follow this guide you need:

  • A dedicated or bare-metal server running Ubuntu 24.04 LTS with at least two NUMA nodes, for example a dual-socket CubePath dedicated server.
  • A non-root user with sudo privileges.

Step 1 - Installing the NUMA tools

Install numactl (which also provides numastat), the text-only hwloc tools and mbw, a small memory bandwidth benchmark:

sudo apt update
sudo apt install numactl hwloc-nox mbw

Check that numactl works:

numactl --version
numactl 2.0.18

Step 2 - Inspecting the NUMA topology

Start with lscpu, which shows how many NUMA nodes exist and which CPUs belong to each:

lscpu | grep -i numa
NUMA node(s):                         2
NUMA node0 CPU(s):                    0-15,32-47
NUMA node1 CPU(s):                    16-31,48-63

In this example each node has 16 physical cores with Hyper-Threading, so CPU 0 and CPU 32 are two threads of the same core. Note the CPU lists for your server: you will use them later.

numactl --hardware adds the memory size of each node and the relative distance between nodes:

numactl --hardware
available: 2 nodes (0-1)
node 0 cpus: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47
node 0 size: 128614 MB
node 0 free: 119305 MB
node 1 cpus: 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63
node 1 size: 129018 MB
node 1 free: 121877 MB
node distances:
node   0   1
  0:  10  21
  1:  21  10

The distance table is relative: 10 is local access, and 21 means accessing the other node costs roughly twice as much. For a fuller picture, including shared caches, use lstopo-no-graphics from hwloc:

lstopo-no-graphics --no-io
Machine (251GB total)
  Package L#0
    NUMANode L#0 (P#0 126GB)
    L3 L#0 (40MB)
      L2 L#0 (1024KB) + L1d L#0 (48KB) + L1i L#0 (32KB) + Core L#0
        PU L#0 (P#0)
        PU L#1 (P#32)
...

Step 3 - Measuring local and remote memory bandwidth

Before changing how applications run, see how much remote memory costs on your hardware. numactl --cpunodebind restricts a command to the CPUs of a node, and --membind forces its memory allocations onto a node.

Run mbw on node 0 with memory from node 0 (local). The arguments run the memcpy test (-t0) five times (-n 5) on 1024 MiB arrays, large enough to not fit in any CPU cache:

numactl --cpunodebind=0 --membind=0 mbw -q -n 5 -t0 1024 | grep AVG
AVG	Method: MEMCPY	Elapsed: 0.10412	MiB: 1024.00000	Copy: 9834.806 MiB/s

Now run it on node 0 with memory from node 1 (remote):

numactl --cpunodebind=0 --membind=1 mbw -q -n 5 -t0 1024 | grep AVG
AVG	Method: MEMCPY	Elapsed: 0.16873	MiB: 1024.00000	Copy: 6068.884 MiB/s

In this example remote access delivers about 40% less bandwidth, and latency is also higher. The exact gap depends on the CPU generation and the interconnect between sockets. This is the penalty that NUMA-aware deployment avoids.

Step 4 - Running a process with a NUMA policy

numactl starts a command with a CPU binding and a memory policy. The four memory policies are:

OptionBehaviour
--membind=NAllocate only from node N. If it runs out of memory, the process can be killed by the OOM killer.
--preferred=NPrefer node N, fall back to other nodes when it is full.
--interleave=allSpread pages round-robin across all nodes, for even bandwidth.
--localallocAllocate on the node of the CPU the thread is running on. This is the kernel default.

You can see the effect of any combination by running numactl --show under it, which prints the policy it inherited:

numactl --cpunodebind=1 --preferred=1 numactl --show
policy: preferred
preferred node: 1
physcpubind: 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63
cpubind: 1
nodebind: 1
membind: 0 1

To run your own program on node 1 with local memory, prefix its command in the same way, replacing /usr/local/bin/your_app with the real path:

numactl --cpunodebind=1 --membind=1 /usr/local/bin/your_app

Which policy to choose depends on the workload:

  • A process that fits in one node (a cache, an API worker, a game server): bind CPUs and memory to the same node. Running one instance per node, each bound to its own node, often scales better than one large instance spread across both.
  • A process that needs more memory than one node has, or that shares a large memory area between all its threads (a database buffer pool, an in-memory analytics engine): use --interleave=all, so no single node becomes a bandwidth hotspot and one node does not run out of memory first.
  • Everything else: leave the default. The kernel's automatic NUMA balancing moves pages closer to the threads that use them.

Step 5 - Pinning a systemd service to a NUMA node

For services, configure the binding in systemd instead of wrapping ExecStart in numactl. systemd 255 on Ubuntu 24.04 supports CPUAffinity=, NUMAPolicy= and NUMAMask=. Open a drop-in for your service, replacing your_app with the unit name:

sudo systemctl edit your_app.service

Add the following lines in the editor. CPUAffinity= takes the CPU list of node 0 you found in Step 2, and NUMAPolicy=bind with NUMAMask=0 restricts memory allocations to node 0:

[Service]
CPUAffinity=0-15 32-47
NUMAPolicy=bind
NUMAMask=0

For a database that should interleave its memory across all nodes instead, use:

[Service]
NUMAPolicy=interleave
NUMAMask=all

Restart the service to apply the drop-in:

sudo systemctl restart your_app.service

Check the CPU affinity of the main process:

PID=$(systemctl show -p MainPID --value your_app.service)
taskset -cp "$PID"
pid 21874's current affinity list: 0-15,32-47

Check the memory policy. Each line of /proc/PID/numa_maps describes one memory mapping and starts with the policy that applies to it:

sudo head -3 /proc/"$PID"/numa_maps
55d2c3a00000 bind:0 file=/usr/local/bin/your_app mapped=312 active=0 N0=312 kernelpagesize_kB=4
55d2c3b4a000 bind:0 anon=1024 dirty=1024 active=0 N0=1024 kernelpagesize_kB=4
7f9a10000000 bind:0 anon=65536 dirty=65536 active=0 N0=65536 kernelpagesize_kB=4

bind:0 confirms the policy, and N0= shows that the pages are on node 0. With interleaving you would see interleave:0-1 and pages split between N0= and N1=.

MySQL has its own switch for the InnoDB buffer pool, innodb_numa_interleave. It only exists when MySQL is built with NUMA support; check with sudo mysql -e "SHOW VARIABLES LIKE 'innodb_numa_interleave';". If the variable is not listed, use the systemd NUMAPolicy=interleave drop-in above for the mysql service instead.

Step 6 - Monitoring NUMA allocation

numastat without arguments shows system-wide counters per node since boot:

numastat
                           node0           node1
numa_hit              9812345678      8723456789
numa_miss                1234567         2345678
numa_foreign             2345678         1234567
interleave_hit             65432           65420
local_node            9811000000      8722000000
other_node               2579245         3802345
  • numa_hit counts allocations that landed on the intended node.
  • numa_miss counts allocations that landed on this node although another node was intended, usually because that node was full.
  • other_node counts allocations on this node made by a process running on another node.

What matters is the trend: run numastat twice, a few minutes apart during normal load, and compare. A steadily growing numa_miss means a node is running out of memory and allocations are spilling over.

To see how the memory of one process is spread across nodes, use numastat -p with a PID or a process name:

numastat -p your_app
Per-node process memory usage (in MBs) for PID 21874 (your_app)
                           Node 0          Node 1           Total
                  --------------- --------------- ---------------
Huge                         0.00            0.00            0.00
Heap                       256.31            0.00          256.31
Stack                        0.13            0.00            0.13
Private                  15890.22            4.02        15894.24
----------------  --------------- --------------- ---------------
Total                    16146.66            4.02        16150.68

For a service bound to node 0, almost all memory should be under Node 0. For the whole system per node, including free memory and page cache, use numastat -m.

Step 7 - Deciding on automatic NUMA balancing

The kernel's automatic NUMA balancing periodically samples memory accesses and migrates pages to the node of the thread that uses them. Check whether it is enabled:

cat /proc/sys/kernel/numa_balancing
1

Its activity is visible in /proc/vmstat:

grep -E 'numa_pages_migrated|numa_hint_faults ' /proc/vmstat
numa_hint_faults 48213377
numa_pages_migrated 3218842

For general workloads leave it enabled. If all latency-sensitive services are explicitly pinned as in Step 5, the sampling adds overhead without benefit, and you can disable it persistently:

echo 'kernel.numa_balancing = 0' | sudo tee /etc/sysctl.d/90-numa.conf
sudo sysctl --system

Measure your application before and after this change and keep the setting that performs better.

Troubleshooting

numactl --hardware shows available: 1 nodes (0): the server has a single NUMA node, or the BIOS has node interleaving enabled, which hides the topology from the OS. On dual-socket servers, disable "Node Interleaving" in the BIOS to expose both nodes. On AMD EPYC, the "NUMA nodes per socket" (NPS) setting controls how many nodes each socket exposes.

A process bound with --membind is killed by the OOM killer while the other node has free memory: strict binding does not fall back to other nodes. Use --preferred (or NUMAPolicy=preferred in systemd), or give the process less memory than the node has.

numastat -p reports Can't read /proc/PID/numa_maps: the process belongs to another user. Run the command with sudo.

Conclusion

You inspected the NUMA topology with lscpu, numactl and hwloc, measured the cost of remote memory, ran processes with explicit NUMA policies, pinned a systemd service to one node and checked where its memory really lives with numa_maps and numastat. As next steps, benchmark your real application with and without pinning, consider running one instance per NUMA node behind a load balancer, and check which node your network card is attached to with cat /sys/class/net/your_interface/device/numa_node so network-heavy services run on that node.