Skip to content

The kernel and the user/kernel boundary

printf("hi") looks like an ordinary function call, but inside it there is a transition between two worlds with different privileges, with the context saved, the arguments checked and a return trip back, and it costs roughly a hundred times more.

Once you understand this boundary, a lot of things that otherwise seem arbitrary become clear: why buffered input is faster than unbuffered, why io_uring came into existence at all, why microkernels lost, and why strace shows something quite different from what is written in the source code.

Prerequisites. Execution modes and interrupts (module 2).

Path of a system call from printf in user mode across the boundary into the kernel and backuser modekernel modeprintf("hi")your codewrite() in libcstill user modesyscall instructioncall number in a registerentry pointcontext is savedsys_writeargument checksreturn: context restored, result in a registerthe boundary is crossed twice; there is no other way into kernel mode
libc still runs in user mode. The real boundary is the syscall instruction, and it can be crossed only through the kernel's predefined entry point.

The program puts the call number and arguments into designated registers and executes the syscall instruction. The CPU switches mode and transfers control to an address the kernel registered at boot, not one specified by the program: otherwise all of isolation would be decorative.

Next, the kernel saves the context, uses the number to find the handler in a table and checks the arguments. The check cannot be skipped: a pointer passed by a process may well point into kernel memory, and the kernel must catch that.

On return, the context is restored, user mode comes back, and the result goes into a register.

An ordinary function call takes a few nanoseconds; a system call takes hundreds. After the Meltdown and Spectre mitigations (module 16) even more, because switching page tables became mandatory.

This hundredfold difference explains a lot:

  • printf buffers output and makes one write instead of a hundred, which is why output to a file and output to a terminal run at different speeds;
  • epoll instead of select saves, above all, the calls themselves;
  • io_uring came into existence precisely to remove the system call from every operation (module 13);
  • the vDSO is a small piece of kernel code mapped into every process so that cheap requests such as gettimeofday can be served without crossing the boundary.

Whatever the architecture, a kernel contains roughly the same subsystems:

Subsystem Responsible for Module
Process management creation, states, scheduling 6, 7
Memory management page tables, allocation, replacement 10–12
Virtual file system a single interface to different file systems 15
I/O subsystem drivers, request queues 13
Network stack sockets, protocols, routing —
Security permissions, isolation, auditing 16

The architecture question is not “what does the kernel do” but which of this runs in kernel mode and which is moved outside.

A monolithic kernel keeps all subsystems in a single kernel-mode address space, so a call between them is an ordinary function call. This is fast, but there is no isolation inside, and a bug in any driver can bring down the whole system.

A microkernel keeps only the minimum in kernel mode: scheduling, basic memory management and message passing. Drivers, file systems and the network stack are moved into separate user-mode processes, so a driver crash kills one process, which can be restarted.

The price is that what used to be a function call becomes a message exchange through the kernel. A single file read now costs several context switches. Historically, this cost is what settled the debate in favor of monolithic kernels.

Hybrid kernels, such as Windows NT or XNU in macOS, claim a microkernel structure but keep critical subsystems in kernel mode for the sake of speed.

The opposite extreme is the unikernel, where the application and the kernel are built into one image and run in a single address space with no split into modes. There is no isolation at all, because a single application has no one to be isolated from, but there is no transition cost either. Their niche is narrow: virtual machines, where the hypervisor provides the isolation.

Terminal window
strace -c -- cat /etc/hostname

A summary of the system calls from one run. You can see that most of the time goes to a few calls, and the rest is startup overhead.

Terminal window
strace -e trace=write -- sh -c 'printf "a"; printf "b"; printf "c"' > /dev/null

Three printf calls produce three write calls to a terminal and only one to a file, because libc buffers output going to a file. This makes the cost of the boundary visible to the naked eye.

Terminal window
ltrace -e 'printf' -- /bin/echo hi 2>&1 | head -3

Library calls belong to a different level. Comparing the output of ltrace and strace shows exactly where the boundary lies.

Terminal window
grep -c . /usr/include/asm/unistd_64.h 2>/dev/null || \
ausyscall --dump 2>/dev/null | wc -l

How many system calls exist in total. The number grows with every kernel and almost never shrinks, because of the commitment not to break the ABI.

Terminal window
cat /proc/self/maps | grep -E 'vdso|vsyscall'

The vDSO, mapped into every process. This is where gettimeofday and clock_gettime come from without a switch into kernel mode.

Terminal window
lsmod | head; modinfo $(lsmod | awk 'NR==2{print $1}') 2>/dev/null | head -5

Loaded kernel modules and the description of one of them.

Terminal window
dmesg --level=err,warn | tail -5

Kernel messages. A bug in a module shows up here and, unlike a bug in a program, can affect the whole system.

“Calling a function from libc is a system call.” libc runs in user mode, and printf may make no system call at all if the data fit into the buffer.

“A system call is just a call to a kernel function.” It is a mode switch with the context saved and the arguments checked, and it costs roughly a hundred times more than an ordinary call.

“Root runs code in kernel mode.” A process running as root runs in user mode just like any other; the kernel simply allows it more in its system calls.

“Modules make Linux a microkernel.” A module runs in the kernel address space with full privileges. A microkernel moves drivers into separate user-mode processes, which is a fundamentally different model.

“Microkernels are more reliable, so they should be used everywhere.” They are more reliable, and they pay for it with the cost of message passing. Where reliability matters more than throughput, that is exactly where they are used: QNX in cars, seL4 in aviation.

Check yourself

1. Where does the CPU transfer control when it executes the syscall instruction?
2. Three printf calls in a program produced a single write system call. Why?
3. A process is started as root. In which mode does its code run?
4. Why did microkernels lose to monolithic kernels in general-purpose systems?
5. What is the vDSO and why is it needed?
6. Why does Linux have both open and openat, and both clone and clone3?

A1 — your own shell builds on this material: every action of the shell is a system call, and running strace on your own program shows this clearly.