Skip to content

B5. LVM and RAID

advancedbuilds on module 14

Module 14 claims that RAID 5 on large disks is risky: the rebuild takes a long time, the whole time the array runs with no redundancy to spare, and a second failure destroys everything. Here you build an array from 1 GiB disks, break it and measure the rebuild yourself. The scale fits in a VM; the mechanics are the same as on 16 TB. On top of the array you carve out logical volumes and see an LVM snapshot as the copy-on-write from module 12, at the block level.

After this lab you will be able to:

  • build RAID 1 and RAID 5 with mdadm, read /proc/mdstat and say what state the array is in and what it is doing right now;
  • fail a disk, replace it and prove with a checksum that the data survived the failure;
  • measure the rebuild time and calculate how long it would take on 16 TB disks;
  • grow a logical volume together with its file system without unmounting;
  • explain why a snapshot takes no space at first and what makes it grow, and why RAID is not a substitute for a backup.

This is the job of whoever is responsible for a server with disks: an alert about a degraded array, replacing a disk, estimating whether the array will finish rebuilding before the next failure, and the conversation after an rm that the mirror dutifully duplicated.

The result is a report with measurements: an array that survived a disk failure, its rebuild time, a calculation for large disks, and logical volumes with a snapshot on top of it.

What to do:

  1. RAID 1 from two disks, a file system, a file with a known checksum. Failure of one disk (mdadm --fail), a check that the data can be read, disk replacement and resync.
  2. RAID 5 from three disks, filled at least halfway. A disk failure and a measurement of the full rebuild time from /proc/mdstat.
  3. A second failure while the rebuild is running: what happened to the array and the data.
  4. LVM on top of /dev/md0: a physical volume, a volume group, two logical volumes, one of them grown on the fly together with its file system.
  5. A snapshot of a logical volume: change the original, mount the snapshot, the old state in it, the snapshot’s usage growing.
  6. Deleting a file on a healthy array and checking that redundancy did not bring it back.

What not to do. Booting from the array and writing to /etc/mdadm/mdadm.conf are not needed: all stages are done in a single VM session. RAID 6 and RAID 10 are not part of the lab. Do not include the VM’s system disk in the array.

Constraints. Software RAID with mdadm and volumes with lvm2. btrfs with its own RAID is left for the “Going further” section.

What goes in the report:

  • the file’s checksum before the failure and after the disk replacement;
  • the RAID 5 rebuild time and speed from /proc/mdstat, and the calculation for 16 TB disks;
  • what happened to the array and the data after the second failure during the rebuild;
  • lvs and df -h before and after growing the volume;
  • one sentence about the deleted file: did the array bring it back.

Done when:

  • you can confirm with a command each of the eight claims in the “What you must be able to prove” table;
  • the report contains the five items above;
  • the calculation for 16 TB is of the same order as the reference figure from stage 3, or you have explained why yours differs.
  • Read the sections “RAID”, “Logical volumes” and “How it actually works in Linux” (the mdadm, vgs, lvs, iostat commands) in module 14, and the section “Copy-on-write” in module 12.
  • You need a virtual machine with root access and four additional disks. A container will not do, because it shares block devices with the host (see the overview); a cloud machine must let you add disks.
  • The Vagrant machine from setup/Vagrantfile creates four 1 GiB disks by itself (Vagrant 2.2.14 or newer is required), and it already has mdadm, lvm2, parted and gdisk. There is no starter code or check.sh for this lab.
  1. Preparing the devices.

    Four 1 GiB disks. Check with lsblk that the kernel sees them.

  2. RAID 1 and the first failure.

    Build a mirror from two disks, create a file system, write a file with a known checksum.

    Mark one disk as faulty (mdadm --fail), make sure the data can be read, replace the disk and wait for the resync. Watch /proc/mdstat during the process.

  3. RAID 5 and measuring the rebuild.

    Build it from three disks, fill it at least halfway, fail a disk and measure the full rebuild time.

    Calculate: if 1 GiB took N seconds, how long would the rebuild of an array of 16 TB disks take? That number is the argument from module 14.

    For reference, measured on virtual disks backed by an SSD: the initial sync of a 2 GiB array took 6 s, the rebuild after a disk replacement 7 s at about 150 MB/s. Extrapolated to 16 TB, that gives about 30 hours. On real spinning disks the speed is of the same order, so the conclusion is the same.

    During the rebuild, read the data and make sure the array works, even though it is degraded. Watch what happens to performance.

  4. A second failure during the rebuild.

    Fail a second disk while the rebuild is running. Record exactly what happened to the data. This is the most important experiment in the lab.

  5. LVM on top of the array.

    A physical volume on /dev/md0, a volume group, two logical volumes. Grow one of them on the fly together with its file system.

  6. Snapshot.

    Take a snapshot of a logical volume, change the data in the original, mount the snapshot and make sure it holds the old state.

    Watch how the snapshot’s usage grows as the original changes: this is copy-on-write in its purest form (module 12).

  7. RAID is not a backup.

    On a working, healthy array, delete a file. Make sure that no redundancy brought it back. Write this down in the report in one sentence.

Claim How to prove it
The array is built and working cat /proc/mdstat
The array survived a disk failure mdadm --detail /dev/md0 in the degraded state
The data was not harmed the file’s checksum before and after
You measured the rebuild time and speed from /proc/mdstat
You know what happens at 16 TB the calculation in the report
The logical volume was grown without downtime lvs, df -h before and after
The snapshot holds the old state diff of the original’s and the snapshot’s contents
RAID does not save you from deletion a description of experiment 7

The array “disappeared” after a reboot. There is no entry in /etc/mdadm/mdadm.conf, and the initramfs was not rebuilt (module 5).

The rebuild is “instant”. You measured the sync of an empty array. Fill it with data, because classic md rebuilds the whole space, not just the used part.

The snapshot overflowed. An LVM snapshot has its own limited size, and when more changes accumulate than fit, it becomes invalid. This is the expected behavior.

The file system was not grown. lvextend grows the volume but not the file system on it. You need resize2fs, or lvextend -r from the start.

Compare with btrfs in RAID 1 mode: there only the used blocks are rebuilt, and on a half-empty array the difference in time will be striking (module 15).