Skip to content

B4. A container by hand

intermediatebuilds on module 16, module 17

Module 17 claims that a container is not a kernel entity but just a combination of independent mechanisms: namespaces, cgroups, overlayfs, capabilities and seccomp. In this lab you turn them on one by one, without Docker, and after each one look at what changed for the process inside. At the end you will have a script of about a hundred lines that does the same thing as docker run, and you will be able to say what it does not do.

After this lab you will be able to:

  • build, from unshare, pivot_root and mount, an environment in which a shell has PID 1, its own hostname and an empty network stack;
  • explain how pivot_root differs from chroot and why you can escape from the latter;
  • show on the host a process that believes itself alone inside, and prove via /proc/self/ns/* that the namespaces differ;
  • strip a process’s privileges with capabilities and seccomp and check that a forbidden call does not go through;
  • explain what a user namespace gives you and why Ubuntu forbids it to unprivileged processes by default.

You need this kind of analysis when a container behaves differently from what the image promises: a process sees someone else’s PIDs, writes go to the wrong layer, mount inside is suddenly allowed. Then you check which namespaces and limits the process got.

The result is a script mycontainer.sh that starts a shell in a prepared root directory with all isolation mechanisms enabled, and a report with a “mechanism → what got hidden” table.

Terminal window
sudo ./mycontainer.sh ./rootfs /bin/sh

The script prints step by step which mechanism is being enabled and what it changes.

What to do:

  1. A root image: a directory of files via debootstrap or an unpacked alpine-minirootfs.
  2. A mount namespace of its own, with pivot_root onto that directory and /proc, /sys, /dev mounted.
  3. PID, UTS, IPC and network namespaces; for the network a veth pair so that ping to the outside works.
  4. A root via overlayfs: a shared read-only lower and a separate upper for each run.
  5. A cgroup with a memory limit on the container process, as in B3.
  6. Stripped privileges: capsh --drop and a seccomp profile that forbids at least mount, reboot and kexec_load.
  7. A user namespace: root inside maps to an ordinary user outside.

What not to do. An image format with hashed layers, a network bridge with NAT, an image registry and lifecycle management are not part of the lab: that is what the runtime contributes, and in the report you will explain why it is needed.

Constraints. No Docker, Podman or any other ready-made runtime: each mechanism is turned on by a separate command, so you can see what it changes.

What goes in the report:

  • the “mechanism → what got hidden” table: the output of ps -e, hostname, ip link, ls /, id after each stage;
  • why you can escape from chroot but not from pivot_root;
  • the value of readlink /proc/self/ns/pid inside and on the host for one and the same process;
  • the user namespace trade-off: why Ubuntu forbids it via AppArmor and what you risk by lifting the restriction;
  • a list of what your container does not do compared with docker run, in your own words.

Done when:

  • the script starts a shell with the command above, and inside ps -e shows only the container’s processes, findmnt / shows overlay, and hostname is its own;
  • you can confirm with a command each of the eight claims in the “What you must be able to prove” table;
  • the report contains the five items above.
  • Read the sections “What a container is made of” and “Images and OCI” in module 17, the section “How many checks a call goes through” (the part about capabilities and seccomp) in module 16, and the section “overlayfs and FUSE” in module 15.
  • Do B3 first: the resource limit stage relies on the group from there.
  • You need root access. A virtual machine works, or a container with namespace nesting enabled (see the overview). The Vagrant machine from the archive already has debootstrap, uidmap, libcap2-bin and libseccomp-dev; there is no starter code or check.sh for this lab.

After each stage run the same commands inside (ps -e, hostname, ip link, ls /, id) and write down what changed. The “mechanism → what got hidden” table is the main result of the lab.

  1. Root image.

    debootstrap --variant=minbase or an unpacked alpine-minirootfs. It is just a directory of files.

  2. Mount namespace and pivot_root.

    unshare --mount, then pivot_root onto your directory, and mount /proc, /sys, /dev.

    Try chroot first instead of pivot_root and find out why you can escape from it; the difference between them is entirely practical.

  3. PID namespace.

    unshare --pid --fork --mount-proc. Now ps -e shows one process, and it is your sh with number 1.

    Check the reverse: from the host the same process is visible under an ordinary PID. Isolation changes only the view: the process exists just as it did before.

  4. UTS, IPC and network namespaces.

    unshare --uts lets you change hostname without touching the host. unshare --net gives an empty stack with only lo.

    Connect the container to the host with a veth pair and get ping to the outside working: this is what Docker does under the hood.

  5. overlayfs.

    Mount the root as a read-only lower plus a separate upper for writes. Create a file inside and find it in the upper directory on the host (module 15).

    Start two containers from the same lower and make sure they do not see each other’s changes.

  6. Resource limits.

    The cgroup from B3 on the container process.

  7. Stripping privileges.

    capsh --drop=cap_sys_admin,cap_net_admin,... and a seccomp profile that forbids at least mount, reboot and kexec_load (module 16).

    Check that the forbidden call really does not go through now.

  8. User namespace.

    unshare --user --map-root-user: inside you are root, outside you are an ordinary user. This is what separates an unprivileged container from a privileged one, and it is the most important layer of protection of all those listed.

For reference: in a correctly built container hostname is its own, ps -e sees 4 processes instead of 144 on the host, the shell has PID 1, findmnt / shows overlay, and there are zero network interfaces.

Claim How to prove it
A PID namespace of its own inside ps -e shows a few processes instead of hundreds on the host
The same process is visible from the host ps -ef | grep on the host
The namespaces differ compare readlink /proc/self/ns/pid on both sides
The root was replaced, not chrooted findmnt / inside
Writes go to upper find the created file on the host
Resources are limited cat <group>/memory.max
Privileges are stripped capsh --print inside
root inside ≠ root outside id inside, ps -o user outside

/proc not remounted. ps shows the host’s processes even though your PID namespace is already your own. The fix is the --mount-proc flag or mounting it by hand.

pivot_root failed. The new root must be a mount point, and often a mount --bind of the directory onto itself is enough.

The network does not work after --net. That is how it should be: the namespace is empty. You need a veth pair, addresses on both ends and a route.

Everything runs as the host’s root. Then you have built a privileged container. The eighth stage is mandatory here; it is the point of the lab.

Run a real image from Docker Hub inside by unpacking its layers by hand. Or compare your script with runc, the OCI reference implementation.