Skip to content

Protection and security

An OS has no separate security subsystem. Security is a property spread across everything the course has already covered. The MMU provides memory isolation (module 2), execution modes provide the boundary (module 3), and the permissions in an inode provide access control (module 15).

Here we put all of that together and add what came later: fine-grained privileges instead of an all-powerful root, mandatory access control, system call filtering, and protection against the thing that gets around all of these mechanisms.

Prerequisites. The user/kernel boundary (module 3), the address space (module 10), permissions (module 15).

Talking about security without a threat model is pointless: the word “secure” means nothing until you say who you are secure against.

A general-purpose operating system guards at least four boundaries:

Boundary Against what Mechanism
process ↔ process reading another process’s memory MMU, separate address spaces
process ↔ kernel arbitrary code in kernel mode execution modes, argument checks
user ↔ user access to other users’ files permissions, ACL
program ↔ its own data exploitation of bugs ASLR, NX, canaries

The last row stands apart from the rest. The first three boundaries protect against an outside attacker. The fourth protects against your own program being forced, through its own bug, to run someone else’s code.

A system call passes DAC, capabilities, security module, and seccomp checks in turnprocessopen()DACrwx bits, ownerEACCEScapabilitiesgranular root rightsEPERMLSMSELinux, AppArmorEACCESseccompfilter by syscall numberSIGSYS✓a denial at any step stops the call; root bypasses DAC, but not seccomp or LSM
The mechanisms complement each other. Root skips the permission check, but it does not get through seccomp and does not bypass SELinux.

DAC, discretionary access control, is the classic Unix model. Discretionary here means “at the owner’s discretion”: the owner of a file decides who gets access. It is simple, but not enough, because a compromised process inherits every right its user has.

Capabilities split the all-powerful root into separate privileges. A web server only needs to bind a port below 1024, which is CAP_NET_BIND_SERVICE, and it has no use for the rest of root’s rights. The idea is right, though it is used less often than it should be.

MAC, mandatory access control, is set by the administrator, and the owner of a file cannot change that policy. SELinux and AppArmor implement it through LSM, check points inside the kernel. A rule reads roughly like “the web server process may read only the site directory”, and it applies even when the process runs as root.

seccomp restricts actions rather than objects, that is, the list of allowed system calls. A program that only computes needs neither socket nor execve, and by forbidding them you make whole classes of exploits pointless. This is how browser tabs are isolated, and container profiles are built on the same thing.

A buffer overflow lets an attacker overwrite the return address on the stack and make the program jump to the attacker’s code. Three mechanisms make this harder.

NX, a bit in the page table entry (module 11), forbids execution. The stack and the heap are marked as data, so code written there simply will not run.

Attackers answered with ROP: instead of bringing their own code, they assemble what they need from pieces already present in the program. NX does not stop that.

ASLR randomly shifts the base addresses of code, libraries, heap and stack on every run. Building a ROP chain requires knowing the addresses, and they are different every time. This is why two copies of the same program show different addresses in /proc/*/maps (module 10).

A stack canary places a random value between the local variables and the return address. It is checked before the function returns, and an overflow that reached the return address has inevitably corrupted the canary on the way.

None of the three mechanisms removes the bug. They make exploitation more expensive, and together they raise the bar high enough that reliable exploits become rare.

Meltdown and Spectre in 2018 took advantage of speculative execution. The processor executes instructions ahead of time, and even though the result is later discarded, a trace remains in the cache. By measuring access time, you can recover data you had no right to read.

They broke the hardware isolation that everything else relied on. Mitigations had to go into the kernel: split the page tables so that kernel memory is not mapped into the process at all.

The price was the cost of a system call, because every transition now requires switching page tables and flushing part of the TLB. This is the same performance regression everyone noticed at the time (module 3).

The lesson is broader than these particular vulnerabilities: abstractions leak. The model in which isolation is guaranteed by hardware turned out to be incomplete, because it did not treat execution time as a channel for passing information.

Terminal window
id; capsh --print 2>/dev/null | head -4

Your user and groups, and the capabilities of the current process.

Terminal window
getcap -r /usr/bin 2>/dev/null | head

Programs with individual privileges instead of setuid. A modern ping is usually already here, and it is missing from the next list.

Terminal window
find /usr/bin -perm -4000 -type f 2>/dev/null | head

Programs with the setuid bit. The shorter this list, the better.

Terminal window
sestatus 2>/dev/null || aa-status 2>/dev/null | head -5

The state of SELinux or AppArmor. enforcing means the policy is applied, permissive means it is only logged.

Terminal window
grep Seccomp /proc/self/status; grep -E 'NoNewPrivs' /proc/self/status

The seccomp mode of the current process: 0 is off, 2 is an active filter.

Terminal window
cat /proc/sys/kernel/randomize_va_space
for i in 1 2; do awk '/\[stack\]/{print $1}' /proc/self/maps; done

The ASLR mode (2 is full) and two different stack addresses in two runs.

Terminal window
checksec --file=/bin/ls 2>/dev/null || readelf -lW /bin/ls | grep -E 'GNU_STACK|GNU_RELRO'

Which protections were built into the executable: NX, canaries, PIE, RELRO.

Terminal window
grep . /sys/devices/system/cpu/vulnerabilities/* 2>/dev/null | head

The processor’s hardware vulnerabilities and the mitigations applied to them.

“Root can do anything.” Not in modern Linux. SELinux can forbid root an action, seccomp can forbid a system call, and root in a container is not the host’s root at all (module 17).

“777 permissions are just convenient.” They let any process on the system overwrite the file. A compromised process of any user gets write access.

“ASLR protects against buffer overflows.” It does not; the bug is still there. It makes exploitation unreliable, because the addresses differ every time.

“NX makes an overflow harmless.” It does not, because ROP builds the attack from the program’s existing code without writing anything anywhere.

“SELinux gets in the way, I’ll turn it off.” Setting permissive instead of enforcing keeps the logging and lets you see exactly what is being blocked. Turning it off completely removes a whole layer of protection to save an hour of configuration.

“Hardware isolation is absolute.” Meltdown and Spectre showed that execution time is also a channel for passing information, and no hardware boundary accounted for it.

Check yourself

1. A process runs as root, but SELinux in enforcing mode forbids it to read a file. Who wins?
2. Why do we need POSIX capabilities if we have setuid?
3. NX forbids executing code on the stack. Why is that not enough?
4. What does seccomp do?
5. Why did the Meltdown mitigations slow down system calls?
6. Two copies of the same program show different stack addresses in /proc/*/maps. What is this?

B4: a container by hand (module 17) directly uses capabilities and seccomp. Extra part: build a program with -fno-stack-protector and without PIE, and compare the checksec output and the behavior on a buffer overflow with a protected build.