Where operating systems are heading
Why this matters
Section titled “Why this matters”Everything earlier in the course settled decades ago. This module is about what is changing right now, and about which assumptions of the classic theory are no longer true.
Most of the items share a theme: the boundary between the kernel and user space, so clear in module 3, is gradually blurring. User code gets inside the kernel, the kernel hands work off through shared memory, and the hardware takes over some of the guarantees.
eBPF: the kernel became programmable
Section titled “eBPF: the kernel became programmable”It used to be that you could change the kernel’s behavior in two ways: write a module, which is allowed to bring the system down, or wait for the next release.
eBPF lets you load a small program into the kernel and attach it to an event: a system call, a kernel function, a network packet, a scheduling decision. Before loading, a verifier checks the program: it proves that the program terminates, does not access memory out of bounds, and does not loop forever. If this cannot be proven, the program is not loaded.
That is the key difference from a kernel module. An eBPF program cannot bring down the kernel by construction, not by convention.
| Use | What it does |
|---|---|
| Tracing | bpftrace, profiling the kernel without rebuilding it and without stopping it |
| Networking | XDP handles a packet before the network stack: filtering and load balancing at millions of packets per second |
| Security | per-call policies with full context, which seccomp cannot do |
| Scheduling | sched_ext: a CPU scheduler as an eBPF program |
The last row of the table is worth rereading carefully. sched_ext, which appeared
in kernel 6.12, lets you replace the scheduler described
in module 7 with a loadable program, and switch back
if it misbehaves. An experiment that used to require
rebuilding the kernel and rebooting now takes seconds.
io_uring: fewer boundaries
Section titled “io_uring: fewer boundaries”Every system call costs hundreds of nanoseconds (module 3), and at a million operations per second that cost becomes the main expense.
io_uring takes the call off the hot path. The program and the kernel share
two ring queues in shared memory: requests go into one,
results are picked up from the other, and in polling mode the boundary need not
be crossed at all.
Behind this is a broader principle, visible both here and in XDP: do not make the transition faster, remove it.
Rust in the kernel
Section titled “Rust in the kernel”Most critical vulnerabilities in the Linux kernel are memory errors: use after free, out-of-bounds access, data races. C does not catch them.
Since kernel 6.1, Rust has been an officially supported language for drivers. Compile-time checks remove a whole class of bugs and cost nothing at run time. Nobody is going to rewrite the kernel that already exists, but new drivers can be safe by construction.
This change is slow and comes with arguments, and the reason is not technical. Mixing two languages in a project of thirty million lines with two thousand developers is expensive purely in organizational terms.
Unikernels
Section titled “Unikernels”When a machine runs a single application, which is common in the cloud, splitting into modes protects against nothing. A unikernel builds the application together with the kernel parts it needs into a single image that runs in a single address space.
There are no system calls and no context switches, the image is measured in megabytes, and startup in milliseconds. The hypervisor takes care of isolation (module 17).
You pay with everything else: debugging, familiar tools, the ability to SSH in and look at what is going on. That is why the niche stays narrow.
Real time and embedded systems
Section titled “Real time and embedded systems”Devices with no OS at all are becoming rarer: even a cheap microcontroller today gets Zephyr or FreeRTOS, that is, a scheduler, drivers and a network stack in a few tens of kilobytes.
At the same time, Linux itself has become better suited to real time. PREEMPT_RT,
which lived as a set of patches for decades, was merged into the mainline kernel in version 6.12.
It makes almost all parts of the kernel preemptible and brings the worst-case latency
down to a predictable few microseconds, at the cost of a small loss
of throughput.
Confidential computing
Section titled “Confidential computing”The classic security model (module 16) assumes that the kernel and the hypervisor can be trusted. In the cloud this assumption looks doubtful, because your data is processed on someone else’s machine under someone else’s hypervisor.
Confidential computing removes the host from the circle of trust. AMD SEV-SNP and Intel TDX encrypt the virtual machine’s memory with a key the hypervisor cannot access, and make attestation possible: cryptographic proof that exactly the image you expected is running.
The trust boundary thus moves from software into hardware. The limitations are openly acknowledged: timing side channels (module 16) do not go away, and you have to trust the processor vendor.
AI workloads
Section titled “AI workloads”Accelerators broke several assumptions of the classic OS at once.
Start with scheduling. The operating system scheduler manages CPU time (module 7), while a GPU has its own queue that the kernel has little to do with. Who decides whose chunk of computation goes first, and what to do about fairness, remains a question without a settled answer.
Next, memory. On an accelerator it is a separate address space. Unified
memory and HMM try to hide this the same way virtual memory
hides the disk (module 12), with the same consequences:
what looks like a memory access may well mean a transfer
over a bus.
And finally, scale. A model that does not fit on one machine makes the cluster, not the process, the unit of scheduling. At that level classic OS theory has no answers at all.
What will remain
Section titled “What will remain”It is worth separating what changes from what does not.
What does not change: the process as the unit of isolation, virtual memory, file abstractions, the need for synchronization, and the trade-off between throughput and response time. These follow from physics and mathematics.
What changes are other things: where the kernel boundary lies, how much it costs, whom you have to trust, and what counts as the unit of scheduling.
That is why the course is built around mechanisms. Interfaces will become obsolete, but the question “who waits, who pays, and why exactly this way” will remain.
Check yourself
Where to go next
Section titled “Where to go next”The course ends here, and there are three natural directions to continue in:
- Deeper into the kernel. Build your own kernel with
PREEMPT_RT, write a simple module, work through some subsystem using its documentation and source code. - Toward distributed systems. Where the classic OS ends: consensus, replication, fault tolerance beyond a single machine.
- Toward performance.
perf,bpftrace, latency analysis, that is, applying the mechanisms from this course to real systems.
Sources
Section titled “Sources”- Brendan Gregg, BPF Performance Tools
- Documentation/bpf and sched_ext
- Efficient IO with io_uring
- Rust for Linux
- LWN.net: the best source on what is happening in the kernel right now