Control groups (cgroups) are the Linux kernel feature that groups processes and limits the CPU, memory, disk I/O and number of tasks they can use. Version 2 replaces the separate per-controller trees of v1 with a single unified hierarchy under /sys/fs/cgroup, and it is the only mode enabled by default on Ubuntu 24.04, Debian 12 and Rocky Linux 9. In this tutorial you will inspect the hierarchy, apply CPU, memory and I/O limits through systemd, work with the raw cgroup files to see what systemd does underneath, and check the limits Docker applies to containers.

Prerequisites

To follow this tutorial you need:

  • A server running Ubuntu 24.04 LTS, for example a CubePath VPS. Debian 12 and Rocky Linux 9 behave the same way.
  • A non-root user with sudo privileges.
  • Optional: Docker Engine installed from the official repository, for the container section.

Step 1 - Confirming cgroups v2 is active

Check the filesystem type mounted at /sys/fs/cgroup:

stat -fc %T /sys/fs/cgroup/
cgroup2fs

cgroup2fs means the unified v2 hierarchy. If you see tmpfs, the system is using v1 or hybrid mode (see Troubleshooting). Next, list the controllers the kernel offers:

cat /sys/fs/cgroup/cgroup.controllers
cpuset cpu io memory hugetlb pids rdma misc

Each controller manages one resource. The most useful are cpu, memory, io and pids.

Step 2 - Exploring the hierarchy

systemd owns the cgroup tree and creates one cgroup per unit. Services live under system.slice, user sessions under user.slice, and virtual machines or containers under machine.slice. Display the tree:

systemd-cgls --no-pager | head -n 15
Control group /:
-.slice
├─user.slice
│ └─user-1000.slice
│   ├─[email protected]
│   └─session-3.scope
│     ├─1450 sshd: your_user [priv]
│     └─1502 -bash
└─system.slice
  ├─ssh.service
  │ └─880 "sshd: /usr/sbin/sshd -D [listener] 0 of 10-100 startups"
  ├─cron.service
  │ └─640 /usr/sbin/cron -f

Find which cgroup a process belongs to. In v2 there is a single line starting with 0:::

cat /proc/self/cgroup
0::/user.slice/user-1000.slice/session-3.scope

Every cgroup is a directory. The files inside it are the interface: cgroup.procs lists its processes, cgroup.subtree_control says which controllers are enabled for its children, and files such as cpu.max or memory.max hold the limits. Look at the files of the SSH service:

ls /sys/fs/cgroup/system.slice/ssh.service/ | head -n 12

Two rules of v2 explain most errors you will meet:

  • A controller is available in a cgroup only if the parent lists it in cgroup.subtree_control.
  • Except for the root, a cgroup that distributes resources to children (has controllers in cgroup.subtree_control) cannot contain processes itself. This is called the "no internal processes" rule.

Step 3 - Limiting CPU

The cpu controller offers two mechanisms:

  • cpu.weight (systemd: CPUWeight=): a relative share from 1 to 10000, default 100. It only matters when CPUs are busy; a group with weight 200 gets twice the CPU time of a sibling with 100.
  • cpu.max (systemd: CPUQuota=): a hard cap written as quota period in microseconds. 50000 100000 means 50 ms of CPU every 100 ms, which is half of one CPU. 200000 100000 allows two full CPUs.

Start a CPU-bound test process in a transient service limited to 50% of one CPU. sha256sum /dev/zero keeps one core busy forever:

sudo systemd-run --unit=cpu-test -p CPUQuota=50% /usr/bin/sha256sum /dev/zero

Read the limit systemd wrote into the cgroup:

cat /sys/fs/cgroup/system.slice/cpu-test.service/cpu.max
50000 100000

Look at the throttling statistics. nr_throttled and throttled_usec grow every time the quota runs out:

grep -E 'usage_usec|nr_throttled|throttled_usec' /sys/fs/cgroup/system.slice/cpu-test.service/cpu.stat
usage_usec 5021377
nr_throttled 98
throttled_usec 4950120

top shows the sha256sum process at about 50% CPU. Stop the test:

sudo systemctl stop cpu-test.service

To pin a service to specific cores, use AllowedCPUs=0-1, which writes to cpuset.cpus.

Step 4 - Limiting and protecting memory

The memory controller has four main knobs, from softest to hardest:

Filesystemd propertyBehavior
memory.minMemoryMin=Guaranteed memory that is never reclaimed
memory.lowMemoryLow=Best-effort protection from reclaim
memory.highMemoryHigh=Above this, the kernel throttles the group and reclaims aggressively
memory.maxMemoryMax=Hard limit: the OOM killer runs inside the group when it is exceeded

Setting MemoryHigh a little below MemoryMax gives an application time to slow down before it is killed. Test the hard limit with a Python process that tries to allocate 300 MB in a group limited to 100 MB, with swap disabled for the group:

sudo systemd-run --unit=mem-test -p MemoryMax=100M -p MemorySwapMax=0 /usr/bin/python3 -c "b = bytearray(300 * 1024 * 1024)"

The process is killed almost immediately. Check the unit:

systemctl status mem-test.service --no-pager | head -n 5
× mem-test.service - /usr/bin/python3 -c "b = bytearray(300 * 1024 * 1024)"
     Loaded: loaded (/run/systemd/transient/mem-test.service; transient)
  Transient: yes
     Active: failed (Result: oom-kill) since Fri 2026-09-25 10:41:07 UTC; 4s ago

Only that group was affected; the rest of the system kept running. Clear the failed unit:

sudo systemctl reset-failed mem-test.service

For running services, memory.current shows usage and memory.events counts how often limits were hit. The oom_kill line is the one to alert on:

cat /sys/fs/cgroup/system.slice/ssh.service/memory.current
cat /sys/fs/cgroup/system.slice/ssh.service/memory.events
4538368
low 0
high 0
max 0
oom 0
oom_kill 0
oom_group_kill 0

Step 5 - Throttling disk I/O

The io controller limits bandwidth and operations per second for each block device, identified by its major:minor number. Find yours:

lsblk -d -o NAME,MAJ:MIN,TYPE
NAME MAJ:MIN TYPE
vda  253:0   disk

Run a write test limited to 20 MB/s. systemd accepts the device path and resolves it to the number. oflag=direct bypasses the page cache so the limit is visible immediately. Replace /dev/vda with your disk:

sudo systemd-run --wait --unit=io-test -p IOWriteBandwidthMax="/dev/vda 20M" /usr/bin/dd if=/dev/zero of=/var/tmp/io-test bs=1M count=200 oflag=direct

Read the result from the journal:

journalctl -u io-test.service -n 3 --no-pager
200+0 records in
200+0 records out
209715200 bytes (210 MB, 200 MiB) copied, 10.0012 s, 21.0 MB/s

The same test without the limit finishes much faster. Remove the test file:

sudo rm /var/tmp/io-test

The related properties are IOReadBandwidthMax=, IOReadIOPSMax= and IOWriteIOPSMax=, which write to io.max. IOWeight= (the io.weight file) sets a relative share, but it only takes effect with an I/O scheduler that supports it, such as BFQ; many virtual disks use the none scheduler, where only hard limits apply. Check with cat /sys/block/vda/queue/scheduler.

Step 6 - Making limits persistent

The systemd-run tests were transient. For real services you have three persistent options.

Set properties on an existing service from the command line. systemd stores them in /etc/systemd/system.control/ and applies them immediately, without a restart:

sudo systemctl set-property cron.service CPUWeight=50 MemoryMax=256M
systemctl show cron.service -p CPUWeight -p MemoryMax
CPUWeight=50
MemoryMax=268435456

Or add the same directives to the [Service] section of a unit or a drop-in created with sudo systemctl edit name.service.

The third option is a slice, which limits a whole group of services together. Create one for batch jobs:

sudo nano /etc/systemd/system/batch.slice
[Unit]
Description=Slice for batch jobs

[Slice]
CPUWeight=20
MemoryMax=2G
TasksMax=512

Assign a service to it by adding Slice=batch.slice to its [Service] section, or test it with a transient unit:

sudo systemctl daemon-reload
sudo systemd-run --slice=batch.slice --unit=batch-test /usr/bin/sleep 300
systemd-cgls --no-pager /batch.slice
Control group /batch.slice:
└─batch-test.service
  └─3120 /usr/bin/sleep 300

All services in batch.slice now share 2 GB of memory and get a low CPU share when the server is busy. Stop the test with sudo systemctl stop batch-test.service.

Step 7 - Using the cgroup filesystem directly

Understanding the raw interface helps when you debug container runtimes or write tools. Create a cgroup directly under the root. This is fine for experiments, but leave production cgroups to systemd:

sudo mkdir /sys/fs/cgroup/demo
cat /sys/fs/cgroup/demo/cgroup.controllers
cpuset cpu io memory pids

These are the controllers the root has enabled in its cgroup.subtree_control. Set a CPU cap of 20% and a memory limit of 50 MB:

echo "20000 100000" | sudo tee /sys/fs/cgroup/demo/cpu.max
echo 50M | sudo tee /sys/fs/cgroup/demo/memory.max

Start a process inside the group. The shell writes its own PID into cgroup.procs, then replaces itself with a CPU-bound command that stops after 30 seconds:

sudo sh -c 'echo $$ > /sys/fs/cgroup/demo/cgroup.procs; exec timeout 30 sha256sum /dev/zero' &

Confirm the process is in the group and being throttled:

cat /sys/fs/cgroup/demo/cgroup.procs
grep nr_throttled /sys/fs/cgroup/demo/cpu.stat

After 30 seconds the process exits and the group is empty. Remove it (you use rmdir, not rm -r, because the files are virtual):

sudo rmdir /sys/fs/cgroup/demo

Step 8 - Watching pressure and usage

Pressure Stall Information (PSI) tells you how much time tasks spent waiting for a resource, which is a better saturation signal than utilization. System-wide values are in /proc/pressure:

cat /proc/pressure/memory
some avg10=0.00 avg60=0.00 avg300=0.00 total=0
full avg10=0.00 avg60=0.00 avg300=0.00 total=0

some is the share of time at least one task was stalled, and full is the share when all non-idle tasks were stalled, averaged over 10, 60 and 300 seconds. Every cgroup has its own cpu.pressure, memory.pressure and io.pressure files, so you can see which service is starving. For a live overview sorted by usage, run:

sudo systemd-cgtop

Press m to sort by memory, c by CPU and q to quit.

Step 9 - Delegation and container limits

Delegation hands a subtree to a non-root manager, which can then create its own child cgroups. systemd delegates the cpu, memory and pids controllers to each user manager, which is what rootless Podman and Docker rely on. Check your user's delegated controllers:

cat /sys/fs/cgroup/user.slice/user-$(id -u).slice/user@$(id -u).service/cgroup.controllers
cpu memory pids

For your own service that manages child processes in sub-cgroups, add Delegate=yes to its [Service] section.

Docker uses the same kernel mechanism. Confirm it runs on cgroups v2 with the systemd driver:

docker info --format '{{.CgroupVersion}} {{.CgroupDriver}}'
2 systemd

Start a container with limits:

docker run -d --name limited --cpus 1.5 --memory 512m --pids-limit 100 nginx:stable

Docker creates a scope unit under system.slice named after the full container ID. Read the values the kernel enforces:

cid=$(docker inspect -f '{{.Id}}' limited)
cat /sys/fs/cgroup/system.slice/docker-"$cid".scope/cpu.max
cat /sys/fs/cgroup/system.slice/docker-"$cid".scope/memory.max
150000 100000
536870912

docker stats --no-stream limited shows the same limits from Docker's side. Remove the container with docker rm -f limited.

Troubleshooting

stat prints tmpfs instead of cgroup2fs. The system booted in v1 or hybrid mode, which happens on older releases such as Ubuntu 20.04 or Rocky Linux 8. Add systemd.unified_cgroup_hierarchy=1 to GRUB_CMDLINE_LINUX in /etc/default/grub, run sudo update-grub (or sudo grub2-mkconfig -o /boot/grub2/grub.cfg on Rocky) and reboot.

cpu.max or memory.max does not exist in a cgroup. The controller is not enabled in the parent. Check the parent's cgroup.subtree_control and enable it, for example echo "+memory" | sudo tee /sys/fs/cgroup/parent/cgroup.subtree_control.

write error: Device or resource busy when adding a process. You are writing a PID into a cgroup that has child cgroups with controllers enabled, which violates the no internal processes rule. Create a leaf cgroup and put the process there.

A limit disappears after a service restart. It was written directly into /sys/fs/cgroup. Use systemctl set-property or a drop-in instead.

Conclusion

You confirmed cgroups v2 is active, explored how systemd lays out the hierarchy, applied and verified CPU, memory and I/O limits, grouped services in a slice, worked with the raw cgroup files, read pressure metrics and checked the limits behind a Docker container. Next, add MemoryHigh and MemoryMax to services that have caused out-of-memory incidents, alert on oom_kill in memory.events, and read man systemd.resource-control for every available property.