File systems
Why this matters
Section titled “Why this matters”The file system is the longest-lived abstraction in operating systems. It turns a numbered array of blocks (module 14) into a tree of names that both people and programs use.
Plenty of practical questions have their answers here. Why didn’t a deleted
file free up space? Why do du and df show different numbers? Why is copying
a million small files hundreds of times slower than copying one
large file of the same total size? Why does a container start instantly
when its image weighs a gigabyte?
Prerequisites. Block devices (module 14), file descriptors (module 6), the page cache (module 13).
A system runs ext4 on disk, tmpfs in memory, NFS over the network and procfs with no storage at all, all at the same time. And a program accesses all of them the same way.
The virtual file system is a layer inside the kernel that defines
a common interface: open, read, write, get attributes.
Each concrete file system implements these operations in its own way, and read()
stays the same.
It’s the same trick as with drivers: one interface on top, different
implementations underneath. VFS is what lets /proc look like files without
being files (module 6).
An inode stores everything about a file: type, permissions, owner, size, timestamps, link count and the location of the data. It is identified by a number within the file system.
Direct pointers give access to the start of a file in one lookup, and most files are small. Large files pay with extra lookups through the indirect levels. The scheme is lopsided on purpose, because it optimizes the common case.
A directory is a file that contains “name → inode number” pairs. Everything else follows from that:
- A hard link is one more such pair pointing to the same inode. Both names are fully equal; neither is a copy or a shortcut. The file exists until the link count drops to zero.
- A symbolic link is a separate file that contains a path as text. It can point to nothing, and it can cross file system boundaries.
- The right to delete a file belongs to the directory, not to the file itself. So you can easily delete a file you have no right to write to.
Block allocation
Section titled “Block allocation”There are several ways to record where a file’s data lives.
A list of blocks, the classic inode, works well for small files and bloats for large ones: a 10 GiB file would need millions of entries.
Extents instead store a range: a start and 20,000 contiguous blocks, one entry instead of thousands. That’s how ext4, XFS and NTFS work. The price is that a fragmented file loses this advantage, because it ends up with many ranges.
Copy-on-write trees, as in btrfs and ZFS, never change a block in place: new content is written to a new location, and pointers are updated from the bottom up. Hence nearly free snapshots, because the old blocks simply stay where they were.
Journaling
Section titled “Journaling”Writing a file takes several operations: update the inode, mark blocks as used, add an entry to the directory, write the data itself. If power is lost halfway, the file system is left in an inconsistent state: blocks are marked as used, and nothing points to them.
Checking the whole file system after every crash is unacceptable, because on a multi-terabyte volume that takes hours.
A journal solves the problem like this: first write the intent to the journal, and only then carry it out. After a crash it’s enough to read the journal and either finish the operations that were written to it completely, or discard the unfinished ones.
Journaling modes differ in what exactly goes into that journal:
| Mode | In the journal | Consequence |
|---|---|---|
writeback |
metadata only | fastest; after a crash the metadata is intact, but file contents may be garbage |
ordered |
metadata, but data is written first | the ext4 default: a reasonable balance |
journal |
both metadata and data | most reliable, everything is written twice |
Modern file systems
Section titled “Modern file systems”ext4 is the default in most distributions. Extents, journaling, predictable behavior. It has no snapshots and no data checksums.
XFS is built for large files and high parallelism, and is typical for RHEL. It can be grown but not shrunk.
btrfs provides copy-on-write, snapshots, data checksums, built-in RAID, compression and subvolumes.
ZFS does the same and also takes over volume management: a file system and a volume manager in one. Historically strong on data integrity.
The most important thing about the last two is data checksums. ext4 and XFS check metadata but not contents, so silent corruption of a block will go unnoticed. btrfs and ZFS will notice it, and with redundancy they will also fix it.
overlayfs and FUSE
Section titled “overlayfs and FUSE”overlayfs merges several directories into one: the lower ones are read-only, the upper one takes writes. Writing to a file from a lower layer causes that file to be copied up.
Containers are built on this (module 17): image layers sit in the lower layers, shared by all containers, and the changes of a particular container go into the upper one. That’s why a container from a gigabyte image starts instantly and takes almost no space: nothing is copied until something changes.
FUSE lets you implement a file system as an ordinary program in user
mode. It runs slower than an in-kernel one because of the boundary crossings
(module 3), but a bug doesn’t bring down the system. That’s how
sshfs, cloud storage mounts and most experimental file systems are built.
Permissions
Section titled “Permissions”The classic Unix model has three sets of permissions, for the owner, the group and everyone else,
with three bits r, w, x in each.
For a directory these bits mean something entirely different than for a file. r allows
reading the list of names, w allows creating and deleting entries, and x allows entering
the directory and accessing its contents by name. A directory with r but without x
lets you see the names and do nothing with them.
There are also special bits. setuid makes a program run as
the file’s owner; that’s how passwd works. setgid does the same
for the group. sticky on the /tmp directory allows a file to be deleted only by
its owner.
ACLs provide what the three-set model can’t hold: separate permissions
for specific users. A + in the output of ls -l shows that an ACL is present.
How it actually works in Linux
Section titled “How it actually works in Linux”stat /etc/hostnameEverything stored in the inode: number, size, block count, permissions, link count and three timestamps.
df -h /; df -i /Usage in bytes and in inodes, separately. You can run out of space on inodes while terabytes are free; this is a classic situation on a file system with millions of small files.
du -sh /var/log 2>/dev/null; df -h /vardu counts what has names; df counts what is actually used. The difference
is deleted files that are still open.
sudo lsof +L1 2>/dev/null | headFiles with a link count of zero that are still open. They explain where the space went after deleting log files.
echo test > /tmp/a; ln /tmp/a /tmp/b; ln -s /tmp/a /tmp/cstat -c '%i %h %n' /tmp/a /tmp/b /tmp/c; rm /tmp/a; cat /tmp/b; cat /tmp/cA hard link and a symbolic link side by side. After the original is deleted,
b can be read and c can’t: the first is a full name of the same inode,
the second only stored a path.
findmnt -t ext4,xfs,btrfs,overlay -o TARGET,SOURCE,FSTYPE,OPTIONSMounted file systems with their options. On a system with containers
you’ll see overlay here.
filefrag -v /var/log/syslog 2>/dev/null | head -8The extents of a particular file: how many contiguous ranges it occupies.
Common misconceptions
Section titled “Common misconceptions”“A file’s name is stored in the file.” The name lives in the directory, and the inode doesn’t contain it. That’s why one inode can have several names.
“I deleted the file, so the space is free.” It will be freed when the link count reaches zero and nobody holds the file open.
“A hard link is a shortcut.” It’s an equal name for the same inode. The shortcut is the symbolic link: it stores a path and breaks when the target disappears.
“I can’t write to the file, so I can’t delete it.” Deletion changes the directory, so the directory’s permissions decide.
“Journaling protects data.” In the default ordered mode, what gets journaled is
metadata. The file system will stay consistent, but a file’s contents after
a crash may turn out incomplete.
“There’s free space but I can’t write, so the file system is broken.”
Check df -i: most likely you’ve run out of inodes.
“btrfs and ext4 differ only in speed.” The main difference is data checksums and snapshots. ext4 won’t notice silent corruption of contents.
Check yourself
A6: a file system in a file. Your own format with a superblock, an inode table and a block bitmap, mounted through FUSE. Required part: implement hard links and confirm that deleting one name doesn’t destroy the file.
Sources
Section titled “Sources”- OSTEP: File System Implementation, Journaling
- Silberschatz, Operating System Concepts, chapters 13–14
man 7 inode,man 2 unlink,man 5 ext4,man 8 mount- Documentation/filesystems, VFS and specific file systems