By default, the Linux scheduler moves a KVM guest's vCPU threads freely across all host cores, and guest memory can land on any NUMA node. For most workloads that is fine, but databases, packet processing and latency-sensitive services lose performance when their vCPUs share cores with other guests or read memory attached to the other CPU socket. In this tutorial you will read the host's CPU and NUMA topology, pin a guest's vCPUs and emulator threads to dedicated cores, bind its memory to one NUMA node, and back it with huge pages, using libvirt on Ubuntu 24.04.

Prerequisites

To follow this guide you need:

  • A physical or bare metal server running Ubuntu 24.04 LTS with KVM and libvirt installed (qemu-kvm, libvirt-daemon-system, libvirt-clients). A multi-socket server shows the NUMA effects best, but pinning also helps on single-socket machines.
  • A non-root user with sudo privileges.
  • A guest called vm1 with 4 vCPUs and 8 GiB of RAM. Adapt the numbers to your VM.

Install the NUMA tools:

sudo apt update
sudo apt install numactl

Step 1 - Reading the host CPU topology

You need to know which logical CPUs are hyperthread siblings of the same core and which NUMA node each belongs to. Show one line per logical CPU:

lscpu -e=CPU,NODE,SOCKET,CORE
CPU NODE SOCKET CORE
  0    0      0    0
  1    0      0    1
  2    0      0    2
...
 15    0      0   15
 16    1      1   16
...
 31    1      1   31
 32    0      0    0
 33    0      0    1
...
 63    1      1   31

In this example there are two sockets, one NUMA node each, 16 cores per socket, and two threads per core. CPUs 0 and 32 share core 0; CPUs 2 and 34 share core 2, and so on. The numbering scheme differs between vendors and BIOS versions, so always read it from your own host.

Confirm the NUMA layout and how much memory each node has:

numactl --hardware
available: 2 nodes (0-1)
node 0 cpus: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47
node 0 size: 128695 MB
node 0 free: 97310 MB
node 1 cpus: 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63
node 1 size: 129019 MB
node 1 free: 101844 MB
node distances:
node   0   1
  0:  10  21
  1:  21  10

The distance table shows that accessing memory on the other node costs about twice as much as local access. The goal is to keep each guest's vCPUs and memory on a single node.

Step 2 - Planning the CPU allocation

Reserve the first core of each socket (and its sibling) for the host: the kernel, libvirt and QEMU's I/O threads. For vm1 with 4 vCPUs on node 0, give it two full physical cores including their hyperthreads, and present them to the guest as 2 cores with 2 threads each so the guest scheduler knows which vCPUs are siblings:

Guest vCPUHost CPUPhysical core
02core 2
134core 2 (sibling)
23core 3
335core 3 (sibling)
emulator threads1, 33core 1

No other guest should be pinned to CPUs 2, 3, 34 or 35.

Step 3 - Pinning vCPUs and emulator threads

Check the current placement:

sudo virsh vcpupin vm1
 VCPU   CPU Affinity
----------------------
 0      0-63
 1      0-63
 2      0-63
 3      0-63

Every vCPU may run anywhere. Edit the domain to add a persistent pinning configuration:

sudo virsh edit vm1

Add a <cputune> block after the <vcpu> line, and replace the existing <cpu> block with one that defines the topology:

<vcpu placement='static'>4</vcpu>
<cputune>
  <vcpupin vcpu='0' cpuset='2'/>
  <vcpupin vcpu='1' cpuset='34'/>
  <vcpupin vcpu='2' cpuset='3'/>
  <vcpupin vcpu='3' cpuset='35'/>
  <emulatorpin cpuset='1,33'/>
</cputune>
<cpu mode='host-passthrough' check='none'>
  <topology sockets='1' dies='1' cores='2' threads='2'/>
</cpu>

sockets x dies x cores x threads must equal the <vcpu> count. host-passthrough gives the guest the full host CPU feature set, which is the best choice for a pinned guest that will not migrate. emulatorpin keeps QEMU's own threads off the vCPU cores.

Pinning changes in the XML apply at the next start. Restart the guest fully (a reboot from inside the guest is not enough):

sudo virsh shutdown vm1
sudo virsh start vm1

Verify:

sudo virsh vcpupin vm1
sudo virsh emulatorpin vm1
 VCPU   CPU Affinity
----------------------
 0      2
 1      34
 2      3
 3      35

 emulator: CPU Affinity
----------------------------------
       *: 1,33

virsh vcpupin vm1 0 2 --live --config changes a single vCPU on a running guest if you need to adjust without a restart. Inside the guest, lscpu should now report Thread(s) per core: 2 and Core(s) per socket: 2.

Step 4 - Binding guest memory to a NUMA node

The vCPUs are on node 0, so the memory must be allocated there too. Edit the domain again:

sudo virsh edit vm1

Add a <numatune> block right after </cputune>:

<numatune>
  <memory mode='strict' nodeset='0'/>
</numatune>

With strict, allocations fail rather than fall back to node 1. Use preferred if you would rather accept remote memory than risk a failure when node 0 is full. Restart the guest:

sudo virsh shutdown vm1
sudo virsh start vm1

Check where the QEMU process's memory lives. numastat -p takes a process name or PID; the guest's QEMU process includes guest=vm1 in its command line:

sudo numastat -p "$(pgrep -f 'guest=vm1')"
Per-node process memory usage (in MBs) for PID 48213 (qemu-system-x86)
                           Node 0          Node 1           Total
                  --------------- --------------- ---------------
Huge                         0.00            0.00            0.00
Heap                        35.12            0.00           35.12
Stack                        0.08            0.00            0.08
Private                   8254.37            0.00         8254.37
----------------  --------------- --------------- ---------------
Total                     8289.57            0.00         8289.57

All memory on node 0 means the binding works.

Disabling automatic NUMA balancing

The kernel's automatic NUMA balancing periodically unmaps pages to discover access patterns and migrates them. On a host where guests are already pinned, that adds overhead with no benefit. Check it:

cat /proc/sys/kernel/numa_balancing

If it prints 1, disable it persistently:

echo 'kernel.numa_balancing = 0' | sudo tee /etc/sysctl.d/60-numa-balancing.conf
sudo sysctl --system

Leave it enabled on hosts where most guests are not pinned.

Step 5 - Backing guest memory with huge pages

With normal 4 KiB pages, a guest with 8 GiB of RAM needs about two million page table entries on the host, and TLB misses add latency. Static 2 MiB huge pages cut that by a factor of 512 and are never swapped.

Reserve 4096 pages of 2 MiB (8 GiB) on node 0, which is where vm1 runs:

echo 4096 | sudo tee /sys/devices/system/node/node0/hugepages/hugepages-2048kB/nr_hugepages

Check that the kernel could allocate them all. Memory fragmentation on a busy host can prevent it:

cat /sys/devices/system/node/node0/hugepages/hugepages-2048kB/nr_hugepages
grep -i hugepages_ /proc/meminfo
4096
HugePages_Total:    4096
HugePages_Free:     4096
HugePages_Rsvd:        0
HugePages_Surp:        0

The sysfs setting does not survive a reboot. To make it persistent, reserve the pages at boot through the kernel command line, which also avoids fragmentation. Edit the GRUB defaults:

sudo nano /etc/default/grub

Append the parameters to the existing GRUB_CMDLINE_LINUX_DEFAULT value:

GRUB_CMDLINE_LINUX_DEFAULT="hugepagesz=2M hugepages=8192"

Keep any options that were already there. On the kernel command line, hugepages is split evenly across nodes, so 8192 gives about 4096 per node on this two-node host. Apply it:

sudo update-grub

Now tell libvirt to use huge pages for vm1:

sudo virsh edit vm1

Add after the <currentMemory> line:

<memoryBacking>
  <hugepages/>
</memoryBacking>

Restart the guest and confirm that huge pages were consumed:

sudo virsh shutdown vm1
sudo virsh start vm1
grep -i hugepages_free /proc/meminfo
HugePages_Free:        0

The free count drops by the guest's memory size. If the guest fails to start with an error about huge pages, there are not enough free pages on the node set in <numatune>.

Step 6 - Keeping other workloads off the pinned cores

Pinning keeps vm1 on cores 2 and 3, but other guests and host processes can still run there. Pin every other guest to a separate set of CPUs so none of them overlap. For host services, systemd can restrict the system and user slices to the housekeeping cores. Run on the host:

sudo systemctl set-property --runtime system.slice AllowedCPUs=0-1,32-33
sudo systemctl set-property --runtime user.slice AllowedCPUs=0-1,32-33

The --runtime flag lets you test the change until the next reboot. Libvirt guests run in machine.slice, which is not affected. Drop --runtime to make it persistent once you are satisfied. Adjust the CPU list to leave enough capacity for the host.

Confirm that no unrelated processes are running on a pinned CPU. The psr column shows the CPU each thread last ran on:

ps -eLo psr,comm | awk '$1 == 2 || $1 == 34'

Only kernel threads and the vm1 vCPU threads (CPU 0/KVM, CPU 1/KVM) should appear.

Troubleshooting

Guest does not start: cannot set CPU affinity or Invalid value for number of CPUs. A cpuset refers to a CPU that does not exist or is offline on this host. Compare it with lscpu -e.

Performance is worse after pinning. Check that the vCPUs of different guests do not share physical cores, and that the guest topology matches the siblings you assigned. Pinning two guests to the two threads of the same core makes them compete.

numastat shows memory on both nodes. The guest was started before <numatune> was added, or mode='preferred' fell back because node 0 was full. Check free memory per node with numactl --hardware.

Huge pages not allocated after reboot. Confirm that cat /proc/cmdline shows the hugepages= options. If not, update-grub was not run or another file in /etc/default/grub.d/ overrides the value.

Conclusion

Your guest now runs on dedicated physical cores with the correct thread topology, its emulator threads are kept apart, its memory is bound to the local NUMA node and backed by huge pages. Measure the effect with your real workload before and after the change. Good next steps are tuning memory overcommit on the rest of the host, using isolcpus or nohz_full kernel parameters for hard real-time guests, and setting the CPU frequency governor to performance.