Virtualization and containers
Why this matters
Section titled “Why this matters”A container starts in milliseconds, a virtual machine in seconds.
A container from a gigabyte image takes up megabytes. docker run alpine
on an Ubuntu machine gives you a working Alpine, even though the kernel stays the same.
There is no separate “container” technology. Containers are a combination of several kernel mechanisms, each of which existed long before Docker. Once you take them apart, you can build a container by hand and stop treating it as a black box.
Prerequisites. The kernel and system calls (module 3), processes (module 6), overlayfs (module 15), capabilities and seccomp (module 16).
Two different ideas
Section titled “Two different ideas”Virtualization creates the illusion of a whole computer. The guest OS has its own kernel and believes it controls the hardware.
Containerization creates the illusion of a separate system within a single kernel. The container’s processes are ordinary processes of the host, just with a limited view of the world.
The consequences are opposite on every count. A virtual machine is isolated more strongly, but heavier. A container is light, but shares the kernel, so a vulnerability in the kernel affects all containers at once.
Hypervisors
Section titled “Hypervisors”A type 1 hypervisor, bare metal, runs directly on the hardware and is itself a minimal operating system. VMware ESXi, Xen and Hyper-V are of this kind, and clouds are built on them.
Type 2, hosted, runs as a program inside an ordinary OS; this is VirtualBox or VMware Workstation. Easier to install, more overhead.
The line between the types is blurry. KVM, for example, turns the Linux kernel into a type 1 hypervisor while remaining an ordinary module of an ordinary system.
Hardware virtualization (Intel VT-x, AMD-V) added a mode in which the guest kernel executes privileged instructions directly, and the hypervisor intercepts only what really needs intervention. Before it, guest code had to be translated on the fly, and that was expensive.
EPT does the same for memory: the hardware translates a guest physical address into a real physical one and takes the most expensive part of the work off the hypervisor (module 11).
Paravirtualization takes a different route. The guest OS knows it is virtual,
and instead of pretending, it calls the hypervisor directly. The virtio drivers
work this way, which is why they are faster than emulating real devices.
What a container is made of
Section titled “What a container is made of”The kernel has no “container” entity. There is a process to which four independent mechanisms have been applied.
-
Namespaces limit what a process sees. There are separate namespaces for PIDs, mount points, network, users, hostname, interprocess communication and cgroups. In its own PID namespace a process sees itself as number 1 and does not see the host’s processes.
-
Control groups (cgroups) limit what a process consumes: CPU time, memory, I/O bandwidth. Thanks to cgroups, the OOM killer becomes predictable for a specific group (module 12).
-
overlayfs provides a root file system made of shared layers (module 15). That is why containers from the same image do not take up space twice.
-
Capabilities and seccomp remove unneeded privileges and system calls (module 16).
Docker, Podman and containerd add nothing fundamental to this. They package images and manage networking and the lifecycle, while the isolation is done by the same four mechanisms, available to anyone from the command line.
Images and OCI
Section titled “Images and OCI”An image consists of file system layers and a manifest with metadata, where each layer is a diff against the previous one. Layers are addressed by content hash, so identical ones are stored once for everyone.
OCI standardizes the image format and the runtime, so an image built with Docker runs under Podman or containerd.
The layered architecture has a very practical consequence: the order of commands in a Dockerfile determines build speed. If you copy the code before installing dependencies, the cache of the dependency layer will be invalidated on every code change.
microVMs and WSL2
Section titled “microVMs and WSL2”Containers are fast, but share the kernel. Virtual machines are isolated, but heavy. A microVM tries to get both advantages: Firecracker starts a virtual machine with a minimal set of emulated devices in tens of milliseconds. This is how cloud functions work, where code from different customers runs on the same machine and a shared kernel is unacceptable.
WSL2 runs a full Linux kernel in a lightweight virtual machine on top of Hyper-V, integrated with the Windows file system and network. The first version of WSL tried to translate Linux system calls into Windows calls and ran into how completely a foreign ABI has to be reproduced.
How it actually works in Linux
Section titled “How it actually works in Linux”lsns 2>/dev/null | headThe namespaces on the system: type, number of processes, and who is in each. On a machine with containers there will be many.
ls -l /proc/self/ns/The namespaces of the current process. The numbers in brackets are identifiers; two processes in the same namespace have the same number.
unshare --pid --fork --mount-proc sh -c 'echo "my PID: $$"; ps -e'A PID namespace of its own: the shell sees itself as number 1 and does not see a single host process. This is the same mechanism that gives a container its PID 1.
unshare --net sh -c 'ip link show'A network namespace of its own: only the lo interface, no network to the outside.
Docker connects such a namespace to the host through a pair of virtual interfaces.
cat /sys/fs/cgroup/cgroup.controllerssystemd-cgls --no-pager 2>/dev/null | head -15The available cgroups v2 controllers and the group tree. systemd puts every
service in its own group; containers use the same technology.
systemd-run --user --scope -p MemoryMax=64M sh -c 'echo limited to 64 MiB; cat /sys/fs/cgroup/$(cat /proc/self/cgroup | cut -d: -f3)/memory.max 2>/dev/null'Running a process with a memory limit without any container.
lsmod | grep -E '^kvm'; ls -l /dev/kvm 2>/dev/nullgrep -o -E 'vmx|svm' /proc/cpuinfo | sort -uHardware virtualization support: vmx for Intel, svm for AMD.
systemd-detect-virtWhether the system itself runs in a virtual environment, and which one.
Common misconceptions
Section titled “Common misconceptions”“A container is a lightweight virtual machine.” It is a host process with a limited view of the world. It has no kernel of its own, which is why its isolation is weaker than a virtual machine’s.
“Docker is an isolation technology.” Isolation comes from namespaces, cgroups, overlayfs, seccomp and capabilities. Docker wraps them in a convenient interface.
“A container runs a different operating system.” What differs there is user space. There is one kernel, and it is the kernel of the distribution installed on the host.
“Root in a container is safe.” Without a user namespace it is the same
host root, just with trimmed capabilities. A hole in the kernel
turns it into full root on the machine.
“Containers isolate you from kernel vulnerabilities.” If anything, the opposite: the kernel is shared, so a vulnerability in it affects all containers at once. That is exactly what microVMs are for.
“The order of commands in a Dockerfile is a matter of taste.” Layers are cached in order, so copying the code before installing dependencies invalidates the cache on every code change.
Check yourself
B4: a container by hand. unshare for namespaces, pivot_root
for the root, overlayfs for layers, cgroups for limits, seccomp
for the call filter, no Docker. The goal is to show that isolation
is made of separate mechanisms, each of which can be turned on and off.
B3: cgroups and OOM. Limit memory and watch the OOM killer fire within a group.
Sources
Section titled “Sources”- Silberschatz, Operating System Concepts, chapter 18
man 7 namespaces,man 7 cgroups,man 1 unshare,man 2 pivot_root- Documentation/admin-guide/cgroup-v2
- OCI Image Specification
- Firecracker: Lightweight Virtualization, NSDI 2020