B4. A container by hand
Module 17 claims that a container is not a kernel entity
but just a combination of independent mechanisms: namespaces, cgroups,
overlayfs, capabilities and seccomp. In this lab you turn them on one by one,
without Docker, and after each one look at what changed for the process inside.
At the end you will have a script of about a hundred lines that does the same thing as
docker run, and you will be able to say what it does not do.
After this lab you will be able to:
- build, from
unshare,pivot_rootandmount, an environment in which a shell has PID 1, its own hostname and an empty network stack; - explain how
pivot_rootdiffers fromchrootand why you can escape from the latter; - show on the host a process that believes itself alone inside,
and prove via
/proc/self/ns/*that the namespaces differ; - strip a process’s privileges with capabilities and seccomp and check that a forbidden call does not go through;
- explain what a user namespace gives you and why Ubuntu forbids it to unprivileged processes by default.
You need this kind of analysis when a container behaves differently from what the image promises:
a process sees someone else’s PIDs, writes go to the wrong layer, mount inside is suddenly
allowed. Then you check which namespaces and limits the process got.
The result is a script mycontainer.sh that starts a shell
in a prepared root directory with all isolation mechanisms enabled,
and a report with a “mechanism → what got hidden” table.
sudo ./mycontainer.sh ./rootfs /bin/shThe script prints step by step which mechanism is being enabled and what it changes.
What to do:
- A root image: a directory of files via
debootstrapor an unpackedalpine-minirootfs. - A mount namespace of its own, with
pivot_rootonto that directory and/proc,/sys,/devmounted. - PID, UTS, IPC and network namespaces; for the network a
vethpair so thatpingto the outside works. - A root via overlayfs: a shared read-only
lowerand a separateupperfor each run. - A cgroup with a memory limit on the container process, as in B3.
- Stripped privileges:
capsh --dropand a seccomp profile that forbids at leastmount,rebootandkexec_load. - A user namespace:
rootinside maps to an ordinary user outside.
What not to do. An image format with hashed layers, a network bridge with NAT, an image registry and lifecycle management are not part of the lab: that is what the runtime contributes, and in the report you will explain why it is needed.
Constraints. No Docker, Podman or any other ready-made runtime: each mechanism is turned on by a separate command, so you can see what it changes.
What goes in the report:
- the “mechanism → what got hidden” table: the output of
ps -e,hostname,ip link,ls /,idafter each stage; - why you can escape from
chrootbut not frompivot_root; - the value of
readlink /proc/self/ns/pidinside and on the host for one and the same process; - the user namespace trade-off: why Ubuntu forbids it via AppArmor and what you risk by lifting the restriction;
- a list of what your container does not do compared with
docker run, in your own words.
Done when:
- the script starts a shell with the command above, and inside
ps -eshows only the container’s processes,findmnt /showsoverlay, andhostnameis its own; - you can confirm with a command each of the eight claims in the “What you must be able to prove” table;
- the report contains the five items above.
Before you start
Section titled “Before you start”- Read the sections “What a container is made of” and “Images and OCI” in module 17, the section “How many checks a call goes through” (the part about capabilities and seccomp) in module 16, and the section “overlayfs and FUSE” in module 15.
- Do B3 first: the resource limit stage relies on the group from there.
- You need root access. A virtual machine works, or a container
with namespace nesting enabled (see the overview).
The Vagrant machine from the archive already has
debootstrap,uidmap,libcap2-binandlibseccomp-dev; there is no starter code orcheck.shfor this lab.
Stages
Section titled “Stages”After each stage run the same commands inside
(ps -e, hostname, ip link, ls /, id) and write down what changed.
The “mechanism → what got hidden” table is the main result of the lab.
-
Root image.
debootstrap --variant=minbaseor an unpackedalpine-minirootfs. It is just a directory of files. -
Mount namespace and
pivot_root.unshare --mount, thenpivot_rootonto your directory, and mount/proc,/sys,/dev.Try
chrootfirst instead ofpivot_rootand find out why you can escape from it; the difference between them is entirely practical. -
PID namespace.
unshare --pid --fork --mount-proc. Nowps -eshows one process, and it is yourshwith number 1.Check the reverse: from the host the same process is visible under an ordinary PID. Isolation changes only the view: the process exists just as it did before.
-
UTS, IPC and network namespaces.
unshare --utslets you changehostnamewithout touching the host.unshare --netgives an empty stack with onlylo.Connect the container to the host with a
vethpair and getpingto the outside working: this is what Docker does under the hood. -
overlayfs.
Mount the root as a read-only
lowerplus a separateupperfor writes. Create a file inside and find it in theupperdirectory on the host (module 15).Start two containers from the same
lowerand make sure they do not see each other’s changes. -
Resource limits.
The cgroup from B3 on the container process.
-
Stripping privileges.
capsh --drop=cap_sys_admin,cap_net_admin,...and aseccompprofile that forbids at leastmount,rebootandkexec_load(module 16).Check that the forbidden call really does not go through now.
-
User namespace.
unshare --user --map-root-user: inside you areroot, outside you are an ordinary user. This is what separates an unprivileged container from a privileged one, and it is the most important layer of protection of all those listed.
What you must be able to prove
Section titled “What you must be able to prove”For reference: in a correctly built container
hostname is its own, ps -e sees 4 processes instead of 144 on the host,
the shell has PID 1, findmnt / shows overlay, and there are zero network interfaces.
| Claim | How to prove it |
|---|---|
| A PID namespace of its own inside | ps -e shows a few processes instead of hundreds on the host |
| The same process is visible from the host | ps -ef | grep on the host |
| The namespaces differ | compare readlink /proc/self/ns/pid on both sides |
The root was replaced, not chrooted |
findmnt / inside |
Writes go to upper |
find the created file on the host |
| Resources are limited | cat <group>/memory.max |
| Privileges are stripped | capsh --print inside |
root inside ≠ root outside |
id inside, ps -o user outside |
Common mistakes
Section titled “Common mistakes”/proc not remounted. ps shows the host’s processes even though your PID
namespace is already your own. The fix is the --mount-proc flag or mounting it
by hand.
pivot_root failed. The new root must be a mount point,
and often a mount --bind of the directory onto itself is enough.
The network does not work after --net. That is how it should be: the namespace is empty.
You need a veth pair, addresses on both ends and a route.
Everything runs as the host’s root. Then you have built a privileged
container. The eighth stage is mandatory here; it is the point of the lab.
Going further
Section titled “Going further”Run a real image from Docker Hub inside by unpacking its layers
by hand. Or compare your script with runc, the OCI reference implementation.