Skip to content

B3. cgroups and the OOM killer

intermediatebuilds on module 12, module 17

Every “2 CPU, 512 MiB” limit in Docker, Kubernetes or systemd comes down to a few files in /sys/fs/cgroup. In this lab you create such a group by hand, without any container tooling, and push a process in it to its memory limit. That way you see what module 12 describes in words: malloc almost never returns NULL, and running out of memory ends with a process being terminated unexpectedly.

After this lab you will be able to:

  • limit CPU, memory and I/O for a single process or service and check that the limit is in effect;
  • tell a CPU quota (cpu.max) from a priority (nice) and explain when you need which;
  • read memory.events and the kernel log after the OOM killer fires and say who was killed and why;
  • explain the difference between memory.max and memory.high and pick the right one for a service that must not go down.

This is the same work people do when investigating an incident like “the container restarted with OOMKilled”: find the group, look at the counters and figure out whether it is the limit or a leak.

On a virtual machine or in a container with cgroup nesting enabled (see the overview), create your own cgroup in the cgroups v2 hierarchy and demonstrate three kinds of limits on it.

What to do:

  1. A group with the cpu, memory and io controllers enabled, in which a process you started is running.
  2. A CPU limit by quota: a looping process gets the specified share, and this is visible in top.
  3. A memory limit via memory.max for a program that allocates and writes memory until the OOM killer kills it in your group.
  4. The same with memory.high: the process stays alive but slows down.
  5. An I/O limit via io.max, with a dd measurement before and after.

Write the memory test program yourself, in C or Python: a loop that allocates 10 MiB at a time and writes at least one byte per page into each block.

What goes in the report:

  • the commands that create the group, and the contents of cgroup.controllers and cgroup.subtree_control at each level;
  • the CPU share from top with the given cpu.max, compared with nice 19 on an idle machine;
  • memory.events after the OOM and the matching line from journalctl -k;
  • your explanation of how memory.high differs from memory.max and which type of service you would choose each one for.

Done when you can confirm each claim in the “What you must be able to prove” table with a command, and the report contains the four items above.

  • Read the sections “Demand paging” and “How it actually works in Linux” (the part about overcommit and the OOM killer) in module 12, and the section “What a container is made of” in module 17.
  • You need a system with cgroups v2 and root access. The Vagrant machine from the course archive works; it already has stress-ng and cgroup-tools.
  • In a container the lab only works with cgroup nesting enabled and sufficient privileges. If cgroup.subtree_control cannot be written, switch to a virtual machine.
  1. Reconnaissance.

    Terminal window
    mount | grep cgroup2
    cat /sys/fs/cgroup/cgroup.controllers
    systemd-cgls | head -20

    Make sure this really is v2 (a single hierarchy), and look at how systemd has already sorted all the services into groups.

  2. Your own group.

    Create a directory in /sys/fs/cgroup, enable the controllers you need in the parent’s cgroup.subtree_control, and put a process into the group by writing its PID to cgroup.procs.

    Note the no internal processes rule: in v2 processes can live only in the leaves of the tree. Trying to put a process into a group that has children gives an error, and that is the expected behavior.

  3. CPU limit.

    cpu.max in the format “quota period”. Start a looping process and check in top that it gets exactly the specified share.

    Compare with nice: nice sets a weight relative to competitors, while cpu.max sets an absolute quota (module 7). A process with nice 19 on an idle machine takes 100%, a process with cpu.max 50000 100000 takes exactly 50%, even if nobody else is there.

  4. Memory limit and OOM.

    memory.max at 64 MiB, then a program that allocates and writes memory 10 MiB at a time in a loop. The write is mandatory: without it the pages do not physically exist (module 12).

    Watch memory.current, memory.events and the kernel log.

  5. Pressure instead of killing.

    Now the same, but with memory.high instead of memory.max. The process stays alive: it is throttled and forced to free memory. Look at memory.pressure and explain the difference in the report.

  6. I/O.

    io.max for a specific device, then dd and a throughput measurement before and after.

Claim How to prove it
The system uses cgroups v2 mount | grep cgroup2
Your group exists and contains the process cat <group>/cgroup.procs
CPU is limited by a quota, not by priority cat <group>/cpu.max + top
Memory is limited cat <group>/memory.max
The OOM fired in your group specifically cat <group>/memory.events (the oom_kill field)
The kernel recorded it journalctl -k --grep 'Out of memory'
You understand the difference between high and max a written explanation in the report

Controllers not enabled. The cpu.max and memory.max files simply do not exist until the controller is allowed in the parent group’s cgroup.subtree_control.

The process does not end up in the group. cgroup.procs takes one PID at a time, and threads are moved via cgroup.threads.

systemd takes the group back. You can create groups by hand next to the systemd tree, but it is more reliable to do it via systemd-run --scope.

The OOM killed the wrong process. If the limit is set on a parent group, the victim is chosen from all its descendants. Check that you limited the level you meant to.

Limit the number of processes with pids.max and see how the fork bomb :(){ :|:& };: stops being dangerous.