Protection and security
Why this matters
Section titled “Why this matters”An OS has no separate security subsystem. Security is a property spread across everything the course has already covered. The MMU provides memory isolation (module 2), execution modes provide the boundary (module 3), and the permissions in an inode provide access control (module 15).
Here we put all of that together and add what came later: fine-grained
privileges instead of an all-powerful root, mandatory access control,
system call filtering, and protection against the thing that gets around
all of these mechanisms.
Prerequisites. The user/kernel boundary (module 3), the address space (module 10), permissions (module 15).
What this is about
Section titled “What this is about”Talking about security without a threat model is pointless: the word “secure” means nothing until you say who you are secure against.
A general-purpose operating system guards at least four boundaries:
| Boundary | Against what | Mechanism |
|---|---|---|
| process ↔ process | reading another process’s memory | MMU, separate address spaces |
| process ↔ kernel | arbitrary code in kernel mode | execution modes, argument checks |
| user ↔ user | access to other users’ files | permissions, ACL |
| program ↔ its own data | exploitation of bugs | ASLR, NX, canaries |
The last row stands apart from the rest. The first three boundaries protect against an outside attacker. The fourth protects against your own program being forced, through its own bug, to run someone else’s code.
How many checks a call goes through
Section titled “How many checks a call goes through”DAC, discretionary access control, is the classic Unix model. Discretionary here means “at the owner’s discretion”: the owner of a file decides who gets access. It is simple, but not enough, because a compromised process inherits every right its user has.
Capabilities split the all-powerful root into separate privileges.
A web server only needs to bind a port below 1024, which is
CAP_NET_BIND_SERVICE, and it has no use for the rest of root’s rights.
The idea is right, though it is used less often than it should be.
MAC, mandatory access control, is set by the administrator, and the owner
of a file cannot change that policy. SELinux and AppArmor implement it through
LSM, check points inside the kernel. A rule reads roughly like “the web server
process may read only the site directory”, and it applies even when
the process runs as root.
seccomp restricts actions rather than objects, that is, the list of allowed
system calls. A program that only computes needs neither socket nor execve,
and by forbidding them you make whole classes of exploits pointless. This is
how browser tabs are isolated, and container profiles are built on the same thing.
Protection from your own bugs
Section titled “Protection from your own bugs”A buffer overflow lets an attacker overwrite the return address on the stack and make the program jump to the attacker’s code. Three mechanisms make this harder.
NX, a bit in the page table entry (module 11), forbids execution. The stack and the heap are marked as data, so code written there simply will not run.
Attackers answered with ROP: instead of bringing their own code, they assemble what they need from pieces already present in the program. NX does not stop that.
ASLR randomly shifts the base addresses of code, libraries, heap and stack on
every run. Building a ROP chain requires knowing the addresses, and they are
different every time. This is why two copies of the same program show different
addresses in /proc/*/maps (module 10).
A stack canary places a random value between the local variables and the return address. It is checked before the function returns, and an overflow that reached the return address has inevitably corrupted the canary on the way.
None of the three mechanisms removes the bug. They make exploitation more expensive, and together they raise the bar high enough that reliable exploits become rare.
When the hardware itself breaks
Section titled “When the hardware itself breaks”Meltdown and Spectre in 2018 took advantage of speculative execution. The processor executes instructions ahead of time, and even though the result is later discarded, a trace remains in the cache. By measuring access time, you can recover data you had no right to read.
They broke the hardware isolation that everything else relied on. Mitigations had to go into the kernel: split the page tables so that kernel memory is not mapped into the process at all.
The price was the cost of a system call, because every transition now requires switching page tables and flushing part of the TLB. This is the same performance regression everyone noticed at the time (module 3).
The lesson is broader than these particular vulnerabilities: abstractions leak. The model in which isolation is guaranteed by hardware turned out to be incomplete, because it did not treat execution time as a channel for passing information.
How it actually works in Linux
Section titled “How it actually works in Linux”id; capsh --print 2>/dev/null | head -4Your user and groups, and the capabilities of the current process.
getcap -r /usr/bin 2>/dev/null | headPrograms with individual privileges instead of setuid. A modern ping
is usually already here, and it is missing from the next list.
find /usr/bin -perm -4000 -type f 2>/dev/null | headPrograms with the setuid bit. The shorter this list, the better.
sestatus 2>/dev/null || aa-status 2>/dev/null | head -5The state of SELinux or AppArmor. enforcing means the policy is applied,
permissive means it is only logged.
grep Seccomp /proc/self/status; grep -E 'NoNewPrivs' /proc/self/statusThe seccomp mode of the current process: 0 is off, 2 is an active filter.
cat /proc/sys/kernel/randomize_va_spacefor i in 1 2; do awk '/\[stack\]/{print $1}' /proc/self/maps; doneThe ASLR mode (2 is full) and two different stack addresses in two runs.
checksec --file=/bin/ls 2>/dev/null || readelf -lW /bin/ls | grep -E 'GNU_STACK|GNU_RELRO'Which protections were built into the executable: NX, canaries, PIE, RELRO.
grep . /sys/devices/system/cpu/vulnerabilities/* 2>/dev/null | headThe processor’s hardware vulnerabilities and the mitigations applied to them.
Common misconceptions
Section titled “Common misconceptions”“Root can do anything.” Not in modern Linux. SELinux can forbid root
an action, seccomp can forbid a system call, and root in a container is not
the host’s root at all (module 17).
“777 permissions are just convenient.” They let any process on the system overwrite the file. A compromised process of any user gets write access.
“ASLR protects against buffer overflows.” It does not; the bug is still there. It makes exploitation unreliable, because the addresses differ every time.
“NX makes an overflow harmless.” It does not, because ROP builds the attack from the program’s existing code without writing anything anywhere.
“SELinux gets in the way, I’ll turn it off.” Setting permissive instead of
enforcing keeps the logging and lets you see exactly what is being blocked.
Turning it off completely removes a whole layer of protection to save an hour
of configuration.
“Hardware isolation is absolute.” Meltdown and Spectre showed that execution time is also a channel for passing information, and no hardware boundary accounted for it.
Check yourself
B4: a container by hand (module 17) directly
uses capabilities and seccomp. Extra part: build a program with
-fno-stack-protector and without PIE, and compare the checksec output
and the behavior on a buffer overflow with a protected build.
Sources
Section titled “Sources”- Silberschatz, Operating System Concepts, chapters 16–17
man 7 capabilities,man 2 seccomp,man 8 selinux,man 7 apparmor- Smashing The Stack For Fun And Profit by Aleph One, 1996
- Meltdown and Spectre: the original papers
- Documentation/admin-guide/hw-vuln: mitigations in the kernel