Skip to content

B1. Installing the system

basicbuilds on module 5

After a kernel update the server did not come up, the console says emergency mode, and you need to figure out which link in the chain the boot reached. The chain from firmware to multi-user.target is described in module 5. In this lab you assemble it by hand: partition the disk, install the kernel, initramfs and bootloader, write a unit. Then you break the boot three times and fix it from the emergency console.

After this lab you will be able to:

  • partition a disk for UEFI and explain what the firmware uses to recognize the ESP;
  • show with two boots why a system needs initramfs even when the disk driver is built into the kernel;
  • write a .service that depends on the network and restarts after a failure, and find in journalctl -u why it failed;
  • read systemd-analyze critical-chain and say what delayed the boot;
  • tell a system that booted properly from one that lets you in over ssh but has broken units.

This kind of repair comes up after a bad kernel update, after moving a disk to another machine, or after changing partitions: the system does not come up or comes up degraded, and the cause is tracked down from the emergency console.

The result: a system installed by hand in a virtual machine with UEFI that boots from GPT through the ESP, mounts the root by UUID and starts your unit, plus a report on three failure scenarios.

What to do:

  1. Partition the disk with GPT: a 512 MiB ESP of type EFI System with FAT32, the rest for the root.
  2. Install the base system step by step: debootstrap for Debian or Ubuntu, pacstrap for Arch.
  3. Write fstab by UUID.
  4. Install the kernel, build the initramfs, install systemd-boot or GRUB in UEFI mode. Use two boots, one with root=UUID= and one with root=/dev/…, to find out what initramfs is for here.
  5. Right after the first boot, capture systemd-analyze and systemd-analyze critical-chain.
  6. Write a .service that runs a script, depends on the network and restarts after a failure. Once, make the script rely on an environment variable the unit does not have, watch it fail, and fix it.
  7. Break and fix the boot in three scenarios, with a VM snapshot before each: a corrupted root UUID in fstab, a deleted initramfs for the running kernel, a unit that blocks multi-user.target.

What not to do. The graphical installer will not do: it handles partitioning, fstab and the bootloader for you, and you will not see the chain. Secure Boot and root encryption are left for “Going further”, and multiple disks, LVM and RAID belong to lab B5.

Constraints. UEFI and GPT only, no BIOS mode and no MBR. fstab by UUID only. A VM snapshot is mandatory before each failure scenario.

What goes in the report:

  • the output of sgdisk -p with the ESP type GUID, and findmnt -o SOURCE,TARGET,UUID /;
  • the result of the two boots without initramfs, with root=UUID= and with root=/dev/…, and your explanation of why they differ;
  • systemd-analyze and critical-chain after the first boot, and what turned out to be the longest link in the chain;
  • how the unit without the environment variable failed: the line from journalctl -u and what you fixed;
  • for each of the three scenarios: how the system behaved, what the emergency console showed, which commands you used to find the cause, and how you fixed it.

Done when:

  • you can prove each of the eight claims in the “What you must be able to prove” table with a single command on your system;
  • systemctl is-system-running says running after all three repairs;
  • the report has the five items above.
  • Read the sections “The chain”, “systemd” and “How it actually works in Linux” in module 5. Why a container has no bootloader of its own is explained in module 17, section “What a container is made of”.
  • You need a virtual machine with UEFI enabled, an empty disk and the ability to take snapshots. WSL2 will not do: Windows controls the boot there. On a cloud machine you need to be able to add a disk (overview).
  • You will need sgdisk (package gdisk), mkfs.fat (dosfstools), debootstrap or pacstrap, lsinitramfs and systemd-analyze.
  • The course archive has no files for this lab: no starter code and no check.sh. The claims table and the report are the check.
  1. GPT partitioning.

    Two partitions: a 512 MiB ESP of type EFI System with a FAT32 file system, the rest for the root.

    Look at the table with sgdisk -p and find the partition type GUID in it: that is what the firmware uses to recognize the ESP. The partition’s name and size do not matter to it.

  2. Base system.

    debootstrap for Debian or Ubuntu, pacstrap for Arch. Any step-by-step method works, except the graphical installer.

  3. fstab by UUID.

    Names like /dev/sda2 will not do: they depend on the order in which devices are detected and change when you add a disk in B5.

  4. Kernel, initramfs and bootloader.

    Install the kernel, build the initramfs, install systemd-boot or GRUB in UEFI mode.

    Then find out why initramfs is needed here. First check whether your controller’s driver is a separate module at all:

    Terminal window
    grep -E '^CONFIG_(SCSI_VIRTIO|VIRTIO_BLK|BLK_DEV_NVME)=' /boot/config-$(uname -r)
    lsinitramfs /boot/initrd.img | grep -c 'kernel/drivers'

    If the answer says =y, the driver is built into the kernel itself, and looking for it among the initramfs modules is pointless (module 5).

  5. First boot.

    systemd-analyze and systemd-analyze critical-chain right after logging in, so you have something to compare against later.

  6. Your own unit.

    Write a .service that runs a simple script, depends on the network and restarts after a failure. Check systemctl status, journalctl -u.

    Make sure the script relies on an environment variable the unit does not have. Watch how it fails, and fix it: this is the most common real-world problem (module 5).

  7. Break and fix.

    Three scenarios, each on its own, with a VM snapshot before each:

    • corrupt the root UUID in fstab;
    • delete the initramfs for the running kernel;
    • create a unit that blocks multi-user.target.

    For each: how the system behaves, what the emergency console shows, which commands you used to find it, how you fixed it.

    Do not expect all three to give you a black screen. Only one of them stops the system from booting at all. The other two do something worse: the machine comes up, lets you in over ssh and looks healthy. The point is to learn the difference between “works” and “booted properly”, and systemctl is-system-running is more useful for that than any external monitoring.

The lab is accepted if you can prove each of the claims on your system with a single command:

Claim How to prove it
The system booted via UEFI, not BIOS [ -d /sys/firmware/efi ] && echo UEFI
The partition table is GPT, with an ESP sgdisk -p /dev/… or lsblk -o NAME,PARTTYPENAME
The root is mounted by UUID findmnt -o SOURCE,TARGET,UUID /
The bootloader knows about the kernel cat /proc/cmdline
You know why initramfs is needed here two boots: with root=UUID= and with root=/dev/…
Your unit is active and restarts systemctl status name.service
You know what slowed the boot down systemd-analyze critical-chain
After all the repairs the system really is intact systemctl is-system-running says running, not degraded

The ESP is not formatted as FAT32. The firmware simply will not see it.

GRUB installed in BIOS mode on a UEFI machine. A classic: the command ran without a single error, and the system did not boot.

fstab by device name. It works until the first new disk.

A unit without After=. The service starts before the network and fails, and in the journal it looks like an application error.

Repairs without a snapshot. By the third attempt you no longer remember what you broke yourself and what was broken before you. A snapshot before each scenario is mandatory.

mount -o remount,rw / does not rescue a broken fstab. The command itself reads fstab to find out which device to mount at /, and trips over the same wrong UUID: can't find UUID=…. Name the device explicitly: mount -o remount,rw /dev/sda2 /, and then you can fix the file.

The bootloader was installed on a different machine. If you built the system on one computer and moved the disk to another, grub-install may have written an image that lacks the driver for the root file system. The firmware starts GRUB, GRUB cannot find its own config and drops you into grub> with error: file '/boot/' not found. You can get out of this state on the spot:

grub> insmod ext2
grub> search --fs-uuid --set=root <root UUID>
grub> set prefix=($root)/boot/grub
grub> insmod normal
grub> normal

To fix it for good, rerun grub-install from the booted system, on its own machine.

Enable Secure Boot and sign your own kernel. Or encrypt the root with LUKS and see how the role of initramfs changes.