B3. cgroups and the OOM killer
Every “2 CPU, 512 MiB” limit in Docker, Kubernetes or systemd
comes down to a few files in /sys/fs/cgroup. In this lab you create
such a group by hand, without any container tooling, and push
a process in it to its memory limit. That way you see what
module 12 describes in words: malloc almost never
returns NULL, and running out of memory ends with a process being
terminated unexpectedly.
After this lab you will be able to:
- limit CPU, memory and I/O for a single process or service and check that the limit is in effect;
- tell a CPU quota (
cpu.max) from a priority (nice) and explain when you need which; - read
memory.eventsand the kernel log after the OOM killer fires and say who was killed and why; - explain the difference between
memory.maxandmemory.highand pick the right one for a service that must not go down.
This is the same work people do when investigating an incident like “the container restarted with OOMKilled”: find the group, look at the counters and figure out whether it is the limit or a leak.
On a virtual machine or in a container with cgroup nesting enabled (see the overview), create your own cgroup in the cgroups v2 hierarchy and demonstrate three kinds of limits on it.
What to do:
- A group with the
cpu,memoryandiocontrollers enabled, in which a process you started is running. - A CPU limit by quota: a looping process gets the specified share,
and this is visible in
top. - A memory limit via
memory.maxfor a program that allocates and writes memory until the OOM killer kills it in your group. - The same with
memory.high: the process stays alive but slows down. - An I/O limit via
io.max, with addmeasurement before and after.
Write the memory test program yourself, in C or Python: a loop that allocates 10 MiB at a time and writes at least one byte per page into each block.
What goes in the report:
- the commands that create the group, and the contents of
cgroup.controllersandcgroup.subtree_controlat each level; - the CPU share from
topwith the givencpu.max, compared withnice 19on an idle machine; memory.eventsafter the OOM and the matching line fromjournalctl -k;- your explanation of how
memory.highdiffers frommemory.maxand which type of service you would choose each one for.
Done when you can confirm each claim in the “What you must be able to prove” table with a command, and the report contains the four items above.
Before you start
Section titled “Before you start”- Read the sections “Demand paging” and “How it actually works in Linux” (the part about overcommit and the OOM killer) in module 12, and the section “What a container is made of” in module 17.
- You need a system with cgroups v2 and root access. The Vagrant machine
from the course archive works; it already has
stress-ngandcgroup-tools. - In a container the lab only works with cgroup nesting enabled
and sufficient privileges. If
cgroup.subtree_controlcannot be written, switch to a virtual machine.
Stages
Section titled “Stages”-
Reconnaissance.
Terminal window mount | grep cgroup2cat /sys/fs/cgroup/cgroup.controllerssystemd-cgls | head -20Make sure this really is v2 (a single hierarchy), and look at how
systemdhas already sorted all the services into groups. -
Your own group.
Create a directory in
/sys/fs/cgroup, enable the controllers you need in the parent’scgroup.subtree_control, and put a process into the group by writing its PID tocgroup.procs.Note the no internal processes rule: in v2 processes can live only in the leaves of the tree. Trying to put a process into a group that has children gives an error, and that is the expected behavior.
-
CPU limit.
cpu.maxin the format “quota period”. Start a looping process and check intopthat it gets exactly the specified share.Compare with
nice:nicesets a weight relative to competitors, whilecpu.maxsets an absolute quota (module 7). A process withnice 19on an idle machine takes 100%, a process withcpu.max 50000 100000takes exactly 50%, even if nobody else is there. -
Memory limit and OOM.
memory.maxat 64 MiB, then a program that allocates and writes memory 10 MiB at a time in a loop. The write is mandatory: without it the pages do not physically exist (module 12).Watch
memory.current,memory.eventsand the kernel log. -
Pressure instead of killing.
Now the same, but with
memory.highinstead ofmemory.max. The process stays alive: it is throttled and forced to free memory. Look atmemory.pressureand explain the difference in the report. -
I/O.
io.maxfor a specific device, thenddand a throughput measurement before and after.
What you must be able to prove
Section titled “What you must be able to prove”| Claim | How to prove it |
|---|---|
| The system uses cgroups v2 | mount | grep cgroup2 |
| Your group exists and contains the process | cat <group>/cgroup.procs |
| CPU is limited by a quota, not by priority | cat <group>/cpu.max + top |
| Memory is limited | cat <group>/memory.max |
| The OOM fired in your group specifically | cat <group>/memory.events (the oom_kill field) |
| The kernel recorded it | journalctl -k --grep 'Out of memory' |
You understand the difference between high and max |
a written explanation in the report |
Common mistakes
Section titled “Common mistakes”Controllers not enabled. The cpu.max and memory.max files simply do not exist
until the controller is allowed in the parent group’s cgroup.subtree_control.
The process does not end up in the group. cgroup.procs takes one PID at a time,
and threads are moved via cgroup.threads.
systemd takes the group back. You can create groups by hand next to the
systemd tree, but it is more reliable to do it via systemd-run --scope.
The OOM killed the wrong process. If the limit is set on a parent group, the victim is chosen from all its descendants. Check that you limited the level you meant to.
Going further
Section titled “Going further”Limit the number of processes with pids.max and see how
the fork bomb :(){ :|:& };: stops being dangerous.