Disks and the block layer
Why this matters
Section titled “Why this matters”A file system sees a numbered array of blocks. Behind that abstraction are devices with radically different physics, and a decision that is right for one turns out to be harmful for another.
Defragmentation helps on a spinning disk and hurts on an SSD. A request scheduler that is optimal for an HDD lowers NVMe performance. RAID 5 on large modern disks is more dangerous than it looks. All of this follows from what is inside.
Prerequisites. Drivers, DMA and request queues (module 13).
The spinning disk
Section titled “The spinning disk”The platters spin, the head moves across them. The time to read a sector has three parts:
- seek, moving the head, takes a few milliseconds and is the most expensive part;
- rotational delay, waiting for the right sector to come around under the head; at 7200 RPM half a revolution takes about 4 ms;
- transfer, the actual read, takes tens of microseconds.
The first two parts are mechanical and don’t improve from one generation to the next. Hence the defining property of the HDD: sequential reads are two orders of magnitude faster than random ones. For thirty years, all disk I/O optimization came down to turning random access into sequential access.
Disk scheduling does the same thing at the queue level: it reorders requests so the head moves in one direction. The classic algorithm here is the “elevator” (SCAN): go one way, serving everything along the way, then go back.
There is no mechanics here, so a random read costs almost as much as a sequential one. In exchange, there are constraints the HDD never had.
Flash memory is read in pages and erased in blocks, and a block is hundreds of times larger than a page. A page can’t be overwritten in place; the whole block has to be erased.
That is why an SSD runs a flash translation layer, the FTL, which is effectively the controller’s own file system. Writes go to free space, and the old location is marked invalid. Logical block 100 today and logical block 100 tomorrow live in different physical cells.
The consequences:
- Wear leveling. A cell survives a limited number of rewrites, so the FTL spreads writes evenly. A file you “overwrite in place” physically lands in a new location every time.
- Garbage collection and TRIM. The controller needs to know which blocks are no longer
needed. The
TRIMcommand passes this on from the file system. Without it the drive slows down over time, because it has to erase blocks at the moment of writing. - Write amplification. One logical write can cause several physical ones: move live pages, erase a block, write the new data.
- Defragmentation is harmful. It doesn’t speed up access (there is no mechanics), but it uses up write endurance.
NVMe is not a type of memory, it’s a protocol. SATA was designed for a disk with one head, so it has one queue of 32 commands. NVMe provides thousands of queues of thousands of commands, one per CPU core, and without that there is nowhere to put the parallelism of flash memory.
Several physical disks that look like one. For speed, for reliability, or for both at once.
RAID 0 spreads blocks across disks; this is called striping. Speed goes up, capacity is full, but reliability ends up worse than with a single disk: the failure of any disk destroys the whole array. Despite the name, there is no redundancy here.
RAID 1 mirrors. Half the capacity, faster reads, survives the failure of one disk.
RAID 5 stores data blocks plus parity, the XOR of the other blocks in the stripe, distributed across all disks. A lost block is recovered by computation. Capacity is (n−1)/n, it survives one disk.
RAID 10 stripes across mirrors. Half the capacity, the best write performance, survives one disk from each pair.
Logical volumes
Section titled “Logical volumes”RAID solves reliability, but not flexibility: partitions are as fixed as they always were. LVM adds a level of indirection between the physical disks and what the file system sees.
Physical disks are combined into a volume group, and logical volumes are carved out of it. A logical volume can be grown on the fly, moved to another disk without downtime, or snapshotted, which gives an instant image of its state at a specific moment.
A snapshot works through the same copy-on-write as in module 12: at first it takes up no space, and it grows as the original changes. The classic scenario: take a snapshot, make a backup of a consistent state from it, and delete the snapshot.
How it actually works in Linux
Section titled “How it actually works in Linux”lsblk -o NAME,TYPE,SIZE,ROTA,SCHED,RQ-SIZE,MODELROTA=1 means a spinning disk, 0 means solid state. SCHED shows
the current request scheduler.
cat /sys/block/*/queue/scheduler 2>/dev/nullThe available schedulers, with the current one in brackets. For NVMe, [none] is normal.
lsblk -DTRIM support: non-zero DISC-GRAN and DISC-MAX mean
the device accepts notifications about freed blocks.
systemctl status fstrim.timer 2>/dev/null | head -3Periodic TRIM. Modern distributions run it once a week
from a timer instead of using the discard mount option, because it’s cheaper.
sudo smartctl -A /dev/sda 2>/dev/null | grep -Ei 'Power_On|Reallocated|Wear|Total_LBAs'Drive health: power-on hours, reallocated sectors, remaining write endurance for an SSD.
cat /proc/mdstat 2>/dev/null; sudo vgs; sudo lvs 2>/dev/nullThe state of software RAID and logical volumes.
iostat -xz 1 3 2>/dev/null | head -20Device load: r_await/w_await show latency,
aqu-sz shows the average queue length. High latency with a short queue
means a slow device, not an overloaded one.
Common misconceptions
Section titled “Common misconceptions”“RAID is a backup.” RAID protects against hardware failure. A deleted file, corrupted data and the work of ransomware get duplicated to every disk just as faithfully.
“RAID 0 is more reliable than a single disk.” It’s worse than a single disk: the failure of any disk destroys the whole array, and the failure probabilities add up.
“SSDs need defragmenting.” They don’t. There is nothing to gain, because there is no mechanics, and write endurance gets used up.
“Logical block 100 is always the same place on an SSD.” The FTL changes the physical location on every write. Because of this, you can’t reliably erase data from an SSD by overwriting it.
“NVMe is fast flash memory.” NVMe is a protocol. The speed comes from the fact that it removes SATA’s limit of one queue of 32 commands.
“A request scheduler always helps.” For an HDD, yes; for NVMe it gets turned off: reordering costs more than it saves.
Check yourself
B5: LVM and RAID. Build an array from file-backed devices, carve out logical volumes, take a snapshot, “break” a disk and follow the rebuild. The main part: measure how long the rebuild takes and what happens to the array the whole time.
Sources
Section titled “Sources”- OSTEP: Hard Disk Drives, RAID, Flash-based SSDs
- Silberschatz, Operating System Concepts, chapter 11
man 8 lvm,man 8 mdadm,man 8 smartctl,man 8 fstrim- Documentation/block, the Linux block layer