B1. Installing the system
After a kernel update the server did not come up, the console says emergency mode,
and you need to figure out which link in the chain the boot reached. The chain from firmware
to multi-user.target is described in module 5. In this lab you
assemble it by hand: partition the disk, install the kernel, initramfs
and bootloader, write a unit. Then you break the boot three times
and fix it from the emergency console.
After this lab you will be able to:
- partition a disk for UEFI and explain what the firmware uses to recognize the ESP;
- show with two boots why a system needs initramfs even when the disk driver is built into the kernel;
- write a
.servicethat depends on the network and restarts after a failure, and find injournalctl -uwhy it failed; - read
systemd-analyze critical-chainand say what delayed the boot; - tell a system that booted properly from one that lets you in over ssh but has broken units.
This kind of repair comes up after a bad kernel update, after moving
a disk to another machine, or after changing partitions: the system does not come up
or comes up degraded, and the cause is tracked down from the emergency console.
The result: a system installed by hand in a virtual machine with UEFI that boots from GPT through the ESP, mounts the root by UUID and starts your unit, plus a report on three failure scenarios.
What to do:
- Partition the disk with GPT: a 512 MiB ESP of type
EFI Systemwith FAT32, the rest for the root. - Install the base system step by step:
debootstrapfor Debian or Ubuntu,pacstrapfor Arch. - Write
fstabby UUID. - Install the kernel, build the initramfs, install
systemd-bootor GRUB in UEFI mode. Use two boots, one withroot=UUID=and one withroot=/dev/…, to find out what initramfs is for here. - Right after the first boot, capture
systemd-analyzeandsystemd-analyze critical-chain. - Write a
.servicethat runs a script, depends on the network and restarts after a failure. Once, make the script rely on an environment variable the unit does not have, watch it fail, and fix it. - Break and fix the boot in three scenarios, with a VM snapshot
before each: a corrupted root UUID in
fstab, a deleted initramfs for the running kernel, a unit that blocksmulti-user.target.
What not to do. The graphical installer will not do: it handles
partitioning, fstab and the bootloader for you, and you will not see the chain.
Secure Boot and root encryption are left for “Going further”,
and multiple disks, LVM and RAID belong to lab B5.
Constraints. UEFI and GPT only, no BIOS mode and no MBR. fstab by
UUID only. A VM snapshot is mandatory before each failure scenario.
What goes in the report:
- the output of
sgdisk -pwith the ESP type GUID, andfindmnt -o SOURCE,TARGET,UUID /; - the result of the two boots without initramfs, with
root=UUID=and withroot=/dev/…, and your explanation of why they differ; systemd-analyzeandcritical-chainafter the first boot, and what turned out to be the longest link in the chain;- how the unit without the environment variable failed: the line from
journalctl -uand what you fixed; - for each of the three scenarios: how the system behaved, what the emergency console showed, which commands you used to find the cause, and how you fixed it.
Done when:
- you can prove each of the eight claims in the “What you must be able to prove” table with a single command on your system;
systemctl is-system-runningsaysrunningafter all three repairs;- the report has the five items above.
Before you start
Section titled “Before you start”- Read the sections “The chain”, “systemd” and “How it actually works in Linux” in module 5. Why a container has no bootloader of its own is explained in module 17, section “What a container is made of”.
- You need a virtual machine with UEFI enabled, an empty disk and the ability to take snapshots. WSL2 will not do: Windows controls the boot there. On a cloud machine you need to be able to add a disk (overview).
- You will need
sgdisk(packagegdisk),mkfs.fat(dosfstools),debootstraporpacstrap,lsinitramfsandsystemd-analyze. - The course archive has no files for this lab:
no starter code and no
check.sh. The claims table and the report are the check.
Stages
Section titled “Stages”-
GPT partitioning.
Two partitions: a 512 MiB ESP of type
EFI Systemwith a FAT32 file system, the rest for the root.Look at the table with
sgdisk -pand find the partition type GUID in it: that is what the firmware uses to recognize the ESP. The partition’s name and size do not matter to it. -
Base system.
debootstrapfor Debian or Ubuntu,pacstrapfor Arch. Any step-by-step method works, except the graphical installer. -
fstabby UUID.Names like
/dev/sda2will not do: they depend on the order in which devices are detected and change when you add a disk in B5. -
Kernel, initramfs and bootloader.
Install the kernel, build the initramfs, install
systemd-bootor GRUB in UEFI mode.Then find out why initramfs is needed here. First check whether your controller’s driver is a separate module at all:
Terminal window grep -E '^CONFIG_(SCSI_VIRTIO|VIRTIO_BLK|BLK_DEV_NVME)=' /boot/config-$(uname -r)lsinitramfs /boot/initrd.img | grep -c 'kernel/drivers'If the answer says
=y, the driver is built into the kernel itself, and looking for it among the initramfs modules is pointless (module 5). -
First boot.
systemd-analyzeandsystemd-analyze critical-chainright after logging in, so you have something to compare against later. -
Your own unit.
Write a
.servicethat runs a simple script, depends on the network and restarts after a failure. Checksystemctl status,journalctl -u.Make sure the script relies on an environment variable the unit does not have. Watch how it fails, and fix it: this is the most common real-world problem (module 5).
-
Break and fix.
Three scenarios, each on its own, with a VM snapshot before each:
- corrupt the root UUID in
fstab; - delete the initramfs for the running kernel;
- create a unit that blocks
multi-user.target.
For each: how the system behaves, what the emergency console shows, which commands you used to find it, how you fixed it.
Do not expect all three to give you a black screen. Only one of them stops the system from booting at all. The other two do something worse: the machine comes up, lets you in over ssh and looks healthy. The point is to learn the difference between “works” and “booted properly”, and
systemctl is-system-runningis more useful for that than any external monitoring. - corrupt the root UUID in
What you must be able to prove
Section titled “What you must be able to prove”The lab is accepted if you can prove each of the claims on your system with a single command:
| Claim | How to prove it |
|---|---|
| The system booted via UEFI, not BIOS | [ -d /sys/firmware/efi ] && echo UEFI |
| The partition table is GPT, with an ESP | sgdisk -p /dev/… or lsblk -o NAME,PARTTYPENAME |
| The root is mounted by UUID | findmnt -o SOURCE,TARGET,UUID / |
| The bootloader knows about the kernel | cat /proc/cmdline |
| You know why initramfs is needed here | two boots: with root=UUID= and with root=/dev/… |
| Your unit is active and restarts | systemctl status name.service |
| You know what slowed the boot down | systemd-analyze critical-chain |
| After all the repairs the system really is intact | systemctl is-system-running says running, not degraded |
Common mistakes
Section titled “Common mistakes”The ESP is not formatted as FAT32. The firmware simply will not see it.
GRUB installed in BIOS mode on a UEFI machine. A classic: the command ran without a single error, and the system did not boot.
fstab by device name. It works until the first new disk.
A unit without After=. The service starts before the network and fails, and in the journal
it looks like an application error.
Repairs without a snapshot. By the third attempt you no longer remember what you broke yourself and what was broken before you. A snapshot before each scenario is mandatory.
mount -o remount,rw / does not rescue a broken fstab. The command itself reads
fstab to find out which device to mount at /, and trips over
the same wrong UUID: can't find UUID=…. Name the device explicitly:
mount -o remount,rw /dev/sda2 /, and then you can fix the file.
The bootloader was installed on a different machine. If you built the system
on one computer and moved the disk to another, grub-install may have
written an image that lacks the driver for the root file system. The firmware
starts GRUB, GRUB cannot find its own config and drops you
into grub> with error: file '/boot/' not found. You can get out of this state
on the spot:
grub> insmod ext2grub> search --fs-uuid --set=root <root UUID>grub> set prefix=($root)/boot/grubgrub> insmod normalgrub> normalTo fix it for good, rerun grub-install from the booted
system, on its own machine.
Going further
Section titled “Going further”Enable Secure Boot and sign your own kernel. Or encrypt the root with LUKS and see how the role of initramfs changes.