Control groups (cgroups) are the Linux kernel feature that groups processes and limits the CPU, memory, disk I/O and number of tasks they can use. Version 2 replaces the separate per-controller trees of v1 with a single unified hierarchy under /sys/fs/cgroup, and it is the only mode enabled by default on Ubuntu 24.04, Debian 12 and Rocky Linux 9. In this tutorial you will inspect the hierarchy, apply CPU, memory and I/O limits through systemd, work with the raw cgroup files to see what systemd does underneath, and check the limits Docker applies to containers.
Prerequisites
To follow this tutorial you need:
- A server running Ubuntu 24.04 LTS, for example a CubePath VPS. Debian 12 and Rocky Linux 9 behave the same way.
- A non-root user with
sudoprivileges. - Optional: Docker Engine installed from the official repository, for the container section.
Step 1 - Confirming cgroups v2 is active
Check the filesystem type mounted at /sys/fs/cgroup:
stat -fc %T /sys/fs/cgroup/
cgroup2fs
cgroup2fs means the unified v2 hierarchy. If you see tmpfs, the system is using v1 or hybrid mode (see Troubleshooting). Next, list the controllers the kernel offers:
cat /sys/fs/cgroup/cgroup.controllers
cpuset cpu io memory hugetlb pids rdma misc
Each controller manages one resource. The most useful are cpu, memory, io and pids.
Step 2 - Exploring the hierarchy
systemd owns the cgroup tree and creates one cgroup per unit. Services live under system.slice, user sessions under user.slice, and virtual machines or containers under machine.slice. Display the tree:
systemd-cgls --no-pager | head -n 15
Control group /:
-.slice
├─user.slice
│ └─user-1000.slice
│ ├─[email protected]
│ └─session-3.scope
│ ├─1450 sshd: your_user [priv]
│ └─1502 -bash
└─system.slice
├─ssh.service
│ └─880 "sshd: /usr/sbin/sshd -D [listener] 0 of 10-100 startups"
├─cron.service
│ └─640 /usr/sbin/cron -f
Find which cgroup a process belongs to. In v2 there is a single line starting with 0:::
cat /proc/self/cgroup
0::/user.slice/user-1000.slice/session-3.scope
Every cgroup is a directory. The files inside it are the interface: cgroup.procs lists its processes, cgroup.subtree_control says which controllers are enabled for its children, and files such as cpu.max or memory.max hold the limits. Look at the files of the SSH service:
ls /sys/fs/cgroup/system.slice/ssh.service/ | head -n 12
Two rules of v2 explain most errors you will meet:
- A controller is available in a cgroup only if the parent lists it in
cgroup.subtree_control. - Except for the root, a cgroup that distributes resources to children (has controllers in
cgroup.subtree_control) cannot contain processes itself. This is called the "no internal processes" rule.
Step 3 - Limiting CPU
The cpu controller offers two mechanisms:
cpu.weight(systemd:CPUWeight=): a relative share from 1 to 10000, default 100. It only matters when CPUs are busy; a group with weight 200 gets twice the CPU time of a sibling with 100.cpu.max(systemd:CPUQuota=): a hard cap written asquota periodin microseconds.50000 100000means 50 ms of CPU every 100 ms, which is half of one CPU.200000 100000allows two full CPUs.
Start a CPU-bound test process in a transient service limited to 50% of one CPU. sha256sum /dev/zero keeps one core busy forever:
sudo systemd-run --unit=cpu-test -p CPUQuota=50% /usr/bin/sha256sum /dev/zero
Read the limit systemd wrote into the cgroup:
cat /sys/fs/cgroup/system.slice/cpu-test.service/cpu.max
50000 100000
Look at the throttling statistics. nr_throttled and throttled_usec grow every time the quota runs out:
grep -E 'usage_usec|nr_throttled|throttled_usec' /sys/fs/cgroup/system.slice/cpu-test.service/cpu.stat
usage_usec 5021377
nr_throttled 98
throttled_usec 4950120
top shows the sha256sum process at about 50% CPU. Stop the test:
sudo systemctl stop cpu-test.service
To pin a service to specific cores, use AllowedCPUs=0-1, which writes to cpuset.cpus.
Step 4 - Limiting and protecting memory
The memory controller has four main knobs, from softest to hardest:
| File | systemd property | Behavior |
|---|---|---|
memory.min | MemoryMin= | Guaranteed memory that is never reclaimed |
memory.low | MemoryLow= | Best-effort protection from reclaim |
memory.high | MemoryHigh= | Above this, the kernel throttles the group and reclaims aggressively |
memory.max | MemoryMax= | Hard limit: the OOM killer runs inside the group when it is exceeded |
Setting MemoryHigh a little below MemoryMax gives an application time to slow down before it is killed. Test the hard limit with a Python process that tries to allocate 300 MB in a group limited to 100 MB, with swap disabled for the group:
sudo systemd-run --unit=mem-test -p MemoryMax=100M -p MemorySwapMax=0 /usr/bin/python3 -c "b = bytearray(300 * 1024 * 1024)"
The process is killed almost immediately. Check the unit:
systemctl status mem-test.service --no-pager | head -n 5
× mem-test.service - /usr/bin/python3 -c "b = bytearray(300 * 1024 * 1024)"
Loaded: loaded (/run/systemd/transient/mem-test.service; transient)
Transient: yes
Active: failed (Result: oom-kill) since Fri 2026-09-25 10:41:07 UTC; 4s ago
Only that group was affected; the rest of the system kept running. Clear the failed unit:
sudo systemctl reset-failed mem-test.service
For running services, memory.current shows usage and memory.events counts how often limits were hit. The oom_kill line is the one to alert on:
cat /sys/fs/cgroup/system.slice/ssh.service/memory.current
cat /sys/fs/cgroup/system.slice/ssh.service/memory.events
4538368
low 0
high 0
max 0
oom 0
oom_kill 0
oom_group_kill 0
Step 5 - Throttling disk I/O
The io controller limits bandwidth and operations per second for each block device, identified by its major:minor number. Find yours:
lsblk -d -o NAME,MAJ:MIN,TYPE
NAME MAJ:MIN TYPE
vda 253:0 disk
Run a write test limited to 20 MB/s. systemd accepts the device path and resolves it to the number. oflag=direct bypasses the page cache so the limit is visible immediately. Replace /dev/vda with your disk:
sudo systemd-run --wait --unit=io-test -p IOWriteBandwidthMax="/dev/vda 20M" /usr/bin/dd if=/dev/zero of=/var/tmp/io-test bs=1M count=200 oflag=direct
Read the result from the journal:
journalctl -u io-test.service -n 3 --no-pager
200+0 records in
200+0 records out
209715200 bytes (210 MB, 200 MiB) copied, 10.0012 s, 21.0 MB/s
The same test without the limit finishes much faster. Remove the test file:
sudo rm /var/tmp/io-test
The related properties are IOReadBandwidthMax=, IOReadIOPSMax= and IOWriteIOPSMax=, which write to io.max. IOWeight= (the io.weight file) sets a relative share, but it only takes effect with an I/O scheduler that supports it, such as BFQ; many virtual disks use the none scheduler, where only hard limits apply. Check with cat /sys/block/vda/queue/scheduler.
Step 6 - Making limits persistent
The systemd-run tests were transient. For real services you have three persistent options.
Set properties on an existing service from the command line. systemd stores them in /etc/systemd/system.control/ and applies them immediately, without a restart:
sudo systemctl set-property cron.service CPUWeight=50 MemoryMax=256M
systemctl show cron.service -p CPUWeight -p MemoryMax
CPUWeight=50
MemoryMax=268435456
Or add the same directives to the [Service] section of a unit or a drop-in created with sudo systemctl edit name.service.
The third option is a slice, which limits a whole group of services together. Create one for batch jobs:
sudo nano /etc/systemd/system/batch.slice
[Unit]
Description=Slice for batch jobs
[Slice]
CPUWeight=20
MemoryMax=2G
TasksMax=512
Assign a service to it by adding Slice=batch.slice to its [Service] section, or test it with a transient unit:
sudo systemctl daemon-reload
sudo systemd-run --slice=batch.slice --unit=batch-test /usr/bin/sleep 300
systemd-cgls --no-pager /batch.slice
Control group /batch.slice:
└─batch-test.service
└─3120 /usr/bin/sleep 300
All services in batch.slice now share 2 GB of memory and get a low CPU share when the server is busy. Stop the test with sudo systemctl stop batch-test.service.
ImportantDo not write limits directly into
/sys/fs/cgroupfor systemd services. systemd rewrites those files when the unit restarts or reloads, so your change silently disappears. Useset-property, unit files or slices.
Step 7 - Using the cgroup filesystem directly
Understanding the raw interface helps when you debug container runtimes or write tools. Create a cgroup directly under the root. This is fine for experiments, but leave production cgroups to systemd:
sudo mkdir /sys/fs/cgroup/demo
cat /sys/fs/cgroup/demo/cgroup.controllers
cpuset cpu io memory pids
These are the controllers the root has enabled in its cgroup.subtree_control. Set a CPU cap of 20% and a memory limit of 50 MB:
echo "20000 100000" | sudo tee /sys/fs/cgroup/demo/cpu.max
echo 50M | sudo tee /sys/fs/cgroup/demo/memory.max
Start a process inside the group. The shell writes its own PID into cgroup.procs, then replaces itself with a CPU-bound command that stops after 30 seconds:
sudo sh -c 'echo $$ > /sys/fs/cgroup/demo/cgroup.procs; exec timeout 30 sha256sum /dev/zero' &
Confirm the process is in the group and being throttled:
cat /sys/fs/cgroup/demo/cgroup.procs
grep nr_throttled /sys/fs/cgroup/demo/cpu.stat
After 30 seconds the process exits and the group is empty. Remove it (you use rmdir, not rm -r, because the files are virtual):
sudo rmdir /sys/fs/cgroup/demo
Step 8 - Watching pressure and usage
Pressure Stall Information (PSI) tells you how much time tasks spent waiting for a resource, which is a better saturation signal than utilization. System-wide values are in /proc/pressure:
cat /proc/pressure/memory
some avg10=0.00 avg60=0.00 avg300=0.00 total=0
full avg10=0.00 avg60=0.00 avg300=0.00 total=0
some is the share of time at least one task was stalled, and full is the share when all non-idle tasks were stalled, averaged over 10, 60 and 300 seconds. Every cgroup has its own cpu.pressure, memory.pressure and io.pressure files, so you can see which service is starving. For a live overview sorted by usage, run:
sudo systemd-cgtop
Press m to sort by memory, c by CPU and q to quit.
Step 9 - Delegation and container limits
Delegation hands a subtree to a non-root manager, which can then create its own child cgroups. systemd delegates the cpu, memory and pids controllers to each user manager, which is what rootless Podman and Docker rely on. Check your user's delegated controllers:
cat /sys/fs/cgroup/user.slice/user-$(id -u).slice/user@$(id -u).service/cgroup.controllers
cpu memory pids
For your own service that manages child processes in sub-cgroups, add Delegate=yes to its [Service] section.
Docker uses the same kernel mechanism. Confirm it runs on cgroups v2 with the systemd driver:
docker info --format '{{.CgroupVersion}} {{.CgroupDriver}}'
2 systemd
Start a container with limits:
docker run -d --name limited --cpus 1.5 --memory 512m --pids-limit 100 nginx:stable
Docker creates a scope unit under system.slice named after the full container ID. Read the values the kernel enforces:
cid=$(docker inspect -f '{{.Id}}' limited)
cat /sys/fs/cgroup/system.slice/docker-"$cid".scope/cpu.max
cat /sys/fs/cgroup/system.slice/docker-"$cid".scope/memory.max
150000 100000
536870912
docker stats --no-stream limited shows the same limits from Docker's side. Remove the container with docker rm -f limited.
Troubleshooting
stat prints tmpfs instead of cgroup2fs. The system booted in v1 or hybrid mode, which happens on older releases such as Ubuntu 20.04 or Rocky Linux 8. Add systemd.unified_cgroup_hierarchy=1 to GRUB_CMDLINE_LINUX in /etc/default/grub, run sudo update-grub (or sudo grub2-mkconfig -o /boot/grub2/grub.cfg on Rocky) and reboot.
cpu.max or memory.max does not exist in a cgroup. The controller is not enabled in the parent. Check the parent's cgroup.subtree_control and enable it, for example echo "+memory" | sudo tee /sys/fs/cgroup/parent/cgroup.subtree_control.
write error: Device or resource busy when adding a process. You are writing a PID into a cgroup that has child cgroups with controllers enabled, which violates the no internal processes rule. Create a leaf cgroup and put the process there.
A limit disappears after a service restart. It was written directly into /sys/fs/cgroup. Use systemctl set-property or a drop-in instead.
Conclusion
You confirmed cgroups v2 is active, explored how systemd lays out the hierarchy, applied and verified CPU, memory and I/O limits, grouped services in a slice, worked with the raw cgroup files, read pressure metrics and checked the limits behind a Docker container. Next, add MemoryHigh and MemoryMax to services that have caused out-of-memory incidents, alert on oom_kill in memory.events, and read man systemd.resource-control for every available property.
