B5. LVM and RAID
Module 14 claims that RAID 5 on large disks is risky: the rebuild takes a long time, the whole time the array runs with no redundancy to spare, and a second failure destroys everything. Here you build an array from 1 GiB disks, break it and measure the rebuild yourself. The scale fits in a VM; the mechanics are the same as on 16 TB. On top of the array you carve out logical volumes and see an LVM snapshot as the copy-on-write from module 12, at the block level.
After this lab you will be able to:
- build RAID 1 and RAID 5 with
mdadm, read/proc/mdstatand say what state the array is in and what it is doing right now; - fail a disk, replace it and prove with a checksum that the data survived the failure;
- measure the rebuild time and calculate how long it would take on 16 TB disks;
- grow a logical volume together with its file system without unmounting;
- explain why a snapshot takes no space at first and what makes it grow, and why RAID is not a substitute for a backup.
This is the job of whoever is responsible for a server with disks: an alert about
a degraded array, replacing a disk, estimating whether the array will finish rebuilding before
the next failure, and the conversation after an rm that the mirror dutifully duplicated.
The result is a report with measurements: an array that survived a disk failure, its rebuild time, a calculation for large disks, and logical volumes with a snapshot on top of it.
What to do:
- RAID 1 from two disks, a file system, a file with a known
checksum. Failure of one disk (
mdadm --fail), a check that the data can be read, disk replacement and resync. - RAID 5 from three disks, filled at least halfway. A disk
failure and a measurement of the full rebuild time from
/proc/mdstat. - A second failure while the rebuild is running: what happened to the array and the data.
- LVM on top of
/dev/md0: a physical volume, a volume group, two logical volumes, one of them grown on the fly together with its file system. - A snapshot of a logical volume: change the original, mount the snapshot, the old state in it, the snapshot’s usage growing.
- Deleting a file on a healthy array and checking that redundancy did not bring it back.
What not to do. Booting from the array and writing to
/etc/mdadm/mdadm.conf are not needed: all stages are done in a single VM
session. RAID 6 and RAID 10 are not part of the lab. Do not include the VM’s system disk
in the array.
Constraints. Software RAID with mdadm and volumes with lvm2.
btrfs with its own RAID is left for the “Going further” section.
What goes in the report:
- the file’s checksum before the failure and after the disk replacement;
- the RAID 5 rebuild time and speed from
/proc/mdstat, and the calculation for 16 TB disks; - what happened to the array and the data after the second failure during the rebuild;
lvsanddf -hbefore and after growing the volume;- one sentence about the deleted file: did the array bring it back.
Done when:
- you can confirm with a command each of the eight claims in the “What you must be able to prove” table;
- the report contains the five items above;
- the calculation for 16 TB is of the same order as the reference figure from stage 3, or you have explained why yours differs.
Before you start
Section titled “Before you start”- Read the sections “RAID”, “Logical volumes”
and “How it actually works in Linux” (the
mdadm,vgs,lvs,iostatcommands) in module 14, and the section “Copy-on-write” in module 12. - You need a virtual machine with root access and four additional disks. A container will not do, because it shares block devices with the host (see the overview); a cloud machine must let you add disks.
- The Vagrant machine from
setup/Vagrantfilecreates four 1 GiB disks by itself (Vagrant 2.2.14 or newer is required), and it already hasmdadm,lvm2,partedandgdisk. There is no starter code orcheck.shfor this lab.
Stages
Section titled “Stages”-
Preparing the devices.
Four 1 GiB disks. Check with
lsblkthat the kernel sees them. -
RAID 1 and the first failure.
Build a mirror from two disks, create a file system, write a file with a known checksum.
Mark one disk as faulty (
mdadm --fail), make sure the data can be read, replace the disk and wait for the resync. Watch/proc/mdstatduring the process. -
RAID 5 and measuring the rebuild.
Build it from three disks, fill it at least halfway, fail a disk and measure the full rebuild time.
Calculate: if 1 GiB took N seconds, how long would the rebuild of an array of 16 TB disks take? That number is the argument from module 14.
For reference, measured on virtual disks backed by an SSD: the initial sync of a 2 GiB array took 6 s, the rebuild after a disk replacement 7 s at about 150 MB/s. Extrapolated to 16 TB, that gives about 30 hours. On real spinning disks the speed is of the same order, so the conclusion is the same.
During the rebuild, read the data and make sure the array works, even though it is
degraded. Watch what happens to performance. -
A second failure during the rebuild.
Fail a second disk while the rebuild is running. Record exactly what happened to the data. This is the most important experiment in the lab.
-
LVM on top of the array.
A physical volume on
/dev/md0, a volume group, two logical volumes. Grow one of them on the fly together with its file system. -
Snapshot.
Take a snapshot of a logical volume, change the data in the original, mount the snapshot and make sure it holds the old state.
Watch how the snapshot’s usage grows as the original changes: this is copy-on-write in its purest form (module 12).
-
RAID is not a backup.
On a working, healthy array, delete a file. Make sure that no redundancy brought it back. Write this down in the report in one sentence.
What you must be able to prove
Section titled “What you must be able to prove”| Claim | How to prove it |
|---|---|
| The array is built and working | cat /proc/mdstat |
| The array survived a disk failure | mdadm --detail /dev/md0 in the degraded state |
| The data was not harmed | the file’s checksum before and after |
| You measured the rebuild | time and speed from /proc/mdstat |
| You know what happens at 16 TB | the calculation in the report |
| The logical volume was grown without downtime | lvs, df -h before and after |
| The snapshot holds the old state | diff of the original’s and the snapshot’s contents |
| RAID does not save you from deletion | a description of experiment 7 |
Common mistakes
Section titled “Common mistakes”The array “disappeared” after a reboot. There is no entry
in /etc/mdadm/mdadm.conf, and the initramfs was not rebuilt
(module 5).
The rebuild is “instant”. You measured the sync of an empty array.
Fill it with data, because classic md rebuilds the whole space,
not just the used part.
The snapshot overflowed. An LVM snapshot has its own limited size, and when more changes accumulate than fit, it becomes invalid. This is the expected behavior.
The file system was not grown. lvextend grows the volume but not the file
system on it. You need resize2fs, or lvextend -r from the start.
Going further
Section titled “Going further”Compare with btrfs in RAID 1 mode: there only the used
blocks are rebuilt, and on a half-empty array the difference in time will be striking
(module 15).