Skip to content

Disks and the block layer

A file system sees a numbered array of blocks. Behind that abstraction are devices with radically different physics, and a decision that is right for one turns out to be harmful for another.

Defragmentation helps on a spinning disk and hurts on an SSD. A request scheduler that is optimal for an HDD lowers NVMe performance. RAID 5 on large modern disks is more dangerous than it looks. All of this follows from what is inside.

Prerequisites. Drivers, DMA and request queues (module 13).

The platters spin, the head moves across them. The time to read a sector has three parts:

  • seek, moving the head, takes a few milliseconds and is the most expensive part;
  • rotational delay, waiting for the right sector to come around under the head; at 7200 RPM half a revolution takes about 4 ms;
  • transfer, the actual read, takes tens of microseconds.

The first two parts are mechanical and don’t improve from one generation to the next. Hence the defining property of the HDD: sequential reads are two orders of magnitude faster than random ones. For thirty years, all disk I/O optimization came down to turning random access into sequential access.

Disk scheduling does the same thing at the queue level: it reorders requests so the head moves in one direction. The classic algorithm here is the “elevator” (SCAN): go one way, serving everything along the way, then go back.

There is no mechanics here, so a random read costs almost as much as a sequential one. In exchange, there are constraints the HDD never had.

Flash memory is read in pages and erased in blocks, and a block is hundreds of times larger than a page. A page can’t be overwritten in place; the whole block has to be erased.

That is why an SSD runs a flash translation layer, the FTL, which is effectively the controller’s own file system. Writes go to free space, and the old location is marked invalid. Logical block 100 today and logical block 100 tomorrow live in different physical cells.

The consequences:

  • Wear leveling. A cell survives a limited number of rewrites, so the FTL spreads writes evenly. A file you “overwrite in place” physically lands in a new location every time.
  • Garbage collection and TRIM. The controller needs to know which blocks are no longer needed. The TRIM command passes this on from the file system. Without it the drive slows down over time, because it has to erase blocks at the moment of writing.
  • Write amplification. One logical write can cause several physical ones: move live pages, erase a block, write the new data.
  • Defragmentation is harmful. It doesn’t speed up access (there is no mechanics), but it uses up write endurance.

NVMe is not a type of memory, it’s a protocol. SATA was designed for a disk with one head, so it has one queue of 32 commands. NVMe provides thousands of queues of thousands of commands, one per CPU core, and without that there is nowhere to put the parallelism of flash memory.

Several physical disks that look like one. For speed, for reliability, or for both at once.

Layout of data and parity blocks for RAID 0, 1, 5, and 10RAID 0D1D2A1A2A3A4A5A6capacity 100%losing any disk meanslosing the whole arrayRAID 1D1D2A1A1A2A2A3A3capacity 50%survives losing 1 diskRAID 5D1D2D3A1A2P1A3P2A4P3A5A6capacity (n−1)/nsurvives losing 1 diskRAID 10D1D2D3D4A1A1A2A2A3A3A4A4A5A5A6A6capacity 50%survives 1 from each pairdata blocksparity
Choosing a level means choosing between capacity, write speed and how many disks you can lose. There is no level that is right for everything.

RAID 0 spreads blocks across disks; this is called striping. Speed goes up, capacity is full, but reliability ends up worse than with a single disk: the failure of any disk destroys the whole array. Despite the name, there is no redundancy here.

RAID 1 mirrors. Half the capacity, faster reads, survives the failure of one disk.

RAID 5 stores data blocks plus parity, the XOR of the other blocks in the stripe, distributed across all disks. A lost block is recovered by computation. Capacity is (n−1)/n, it survives one disk.

RAID 10 stripes across mirrors. Half the capacity, the best write performance, survives one disk from each pair.

RAID solves reliability, but not flexibility: partitions are as fixed as they always were. LVM adds a level of indirection between the physical disks and what the file system sees.

Physical disks are combined into a volume group, and logical volumes are carved out of it. A logical volume can be grown on the fly, moved to another disk without downtime, or snapshotted, which gives an instant image of its state at a specific moment.

A snapshot works through the same copy-on-write as in module 12: at first it takes up no space, and it grows as the original changes. The classic scenario: take a snapshot, make a backup of a consistent state from it, and delete the snapshot.

Terminal window
lsblk -o NAME,TYPE,SIZE,ROTA,SCHED,RQ-SIZE,MODEL

ROTA=1 means a spinning disk, 0 means solid state. SCHED shows the current request scheduler.

Terminal window
cat /sys/block/*/queue/scheduler 2>/dev/null

The available schedulers, with the current one in brackets. For NVMe, [none] is normal.

Terminal window
lsblk -D

TRIM support: non-zero DISC-GRAN and DISC-MAX mean the device accepts notifications about freed blocks.

Terminal window
systemctl status fstrim.timer 2>/dev/null | head -3

Periodic TRIM. Modern distributions run it once a week from a timer instead of using the discard mount option, because it’s cheaper.

Terminal window
sudo smartctl -A /dev/sda 2>/dev/null | grep -Ei 'Power_On|Reallocated|Wear|Total_LBAs'

Drive health: power-on hours, reallocated sectors, remaining write endurance for an SSD.

Terminal window
cat /proc/mdstat 2>/dev/null; sudo vgs; sudo lvs 2>/dev/null

The state of software RAID and logical volumes.

Terminal window
iostat -xz 1 3 2>/dev/null | head -20

Device load: r_await/w_await show latency, aqu-sz shows the average queue length. High latency with a short queue means a slow device, not an overloaded one.

“RAID is a backup.” RAID protects against hardware failure. A deleted file, corrupted data and the work of ransomware get duplicated to every disk just as faithfully.

“RAID 0 is more reliable than a single disk.” It’s worse than a single disk: the failure of any disk destroys the whole array, and the failure probabilities add up.

“SSDs need defragmenting.” They don’t. There is nothing to gain, because there is no mechanics, and write endurance gets used up.

“Logical block 100 is always the same place on an SSD.” The FTL changes the physical location on every write. Because of this, you can’t reliably erase data from an SSD by overwriting it.

“NVMe is fast flash memory.” NVMe is a protocol. The speed comes from the fact that it removes SATA’s limit of one queue of 32 commands.

“A request scheduler always helps.” For an HDD, yes; for NVMe it gets turned off: reordering costs more than it saves.

Check yourself

1. Why are sequential reads on a spinning disk two orders of magnitude faster than random ones?
2. Why does an SSD need the TRIM command?
3. A RAID 0 array of four disks. One of them fails. What happens to the data?
4. Why is RAID 5 considered risky on 16–20 TB disks?
5. What is NVMe?
6. Why is an LVM snapshot created instantly and takes no space at first?

B5: LVM and RAID. Build an array from file-backed devices, carve out logical volumes, take a snapshot, “break” a disk and follow the rebuild. The main part: measure how long the rebuild takes and what happens to the array the whole time.