Skip to content

Processes

You run ./server, it crashes, you restart it and get bind: address already in use, even though ps shows no server at all. You press Ctrl+C and the process doesn’t die. You run kill -9, and it still hangs around in the list with state Z.

Each of these has an exact explanation, and all three come down to the same thing: what the operating system considers a process, what states it can be in, and who cleans up after it dies. That is what this module covers: from fork to a table entry that somebody has to read.

Prerequisites. A picture of interrupts and the timer (module 2) and of the boundary between user space and kernel space (module 3).

A program is a file on disk, a sequence of bytes in a format the kernel knows how to load; in Linux that format is ELF. It is entirely passive and does nothing on its own.

A process is a program being executed, together with everything the kernel set up for it: an address space, open files, a current directory, credentials, a program counter, a stack. Forty processes can run from the single file /usr/bin/python3 at once, and all forty will be independent.

Each process sees its own memory and cannot see anyone else’s. What it sees is divided into sections with different access permissions.

Sections of a process address space: code, data, heap, free addresses, stack0x0000…0xFFFF…text (code)read + executedata, bssglobals and staticsheapmalloc, grows down ↓unmapped addressesstackcall frames, grows up ↑heapstackif they meet: stack overflow or malloc fails
The classic section layout. A real address space is more complicated: there are pages with no access permissions between the sections, and the base addresses are randomly shifted (ASLR, module 16).
  • Code (text) holds instructions. It can be read and executed but not written: an attempt to overwrite its own code kills the process.
  • Data (data, bss) holds global and static variables: data holds the ones with an initial value, bss the ones that start at zero and therefore take no space in the file.
  • The heap holds dynamic memory handed out by malloc and grows toward higher addresses.
  • The stack holds function call frames and local variables, and grows toward the heap.

The gap between the heap and the stack is not free memory belonging to the process: these are addresses with nothing mapped to them, and touching them causes a page fault that the kernel cannot service, so the process gets SIGSEGV.

Everything the kernel knows about a process lives in one structure, the process control block. In Linux it is struct task_struct, and it is a sizable structure, several hundred fields. The main groups are:

Group What it holds Why
Identification PID, PPID, UID, GID who it is and on whose behalf it runs
State current state, exit code what to do with it next
Context program counter, registers, stack pointer to resume execution after a switch
Memory pointer to the address space, section bounds so the MMU knows what to translate
I/O open file table, working directory so read(3, …) knows what “3” means
Scheduling priority, nice, CPU time used so the scheduler knows whom to let run
Relationships parent, children, process group to build the tree and deliver signals

Process control blocks live in kernel memory. A process can neither read its own block directly nor tweak its priority there; that is what system calls are for.

Process state diagram: new, ready, running, blocked, terminatedadmitteddispatchexit()preemption: time quantum expiredI/O requestevent occurrednewreadyrunningterminatedblockedon a single core, exactly one process is in the “running” state
Ready and blocked wait in fundamentally different ways. A ready process is waiting for the CPU and will run as soon as it is let in. A blocked one won't, even if the CPU is idle.

The pair most often confused here is “ready” and “blocked”, even though the difference is entirely practical. If the system is slow and processes sit in the “ready” state, there isn’t enough CPU. If they are blocked, the CPU has nothing to do with it, and you should look for the bottleneck in the disk or the network.

The transition from “running” to “ready” is called preemption: the timer fired, the quantum ran out, the kernel took the CPU away. The process did nothing wrong and isn’t waiting for anything; it was simply put back in the queue.

To take one process off a CPU core and put another on, the kernel saves the first one’s registers into its process control block, loads the second one’s registers and switches the address space. This is a context switch. It has a cost, both direct and indirect.

The direct cost is measured in single microseconds, but the indirect one is larger. The new process arrives to cold caches and an empty TLB, so its first few thousand instructions run slower than they could. This is why the time quantum isn’t made very small: the more often you switch, the larger the share of the CPU spent on the switching itself.

In Unix a new process is always cloned from an existing one, and then, if needed, replaces its image with a different program.

  1. fork() creates a child process, an almost exact copy of the parent: the same open files, the same address space, the same point of execution. The only differences are the PID and what fork itself returns: the child’s PID to the parent, zero to the child.

  2. execve() replaces the process’s address space with a new program. The PID stays the same, because it is the same process, now running different code.

  3. exit() terminates the process. The kernel frees its memory and closes its files, but keeps the process control block for now: it holds the exit code, which nobody has read yet.

  4. wait() in the parent reads that code and removes the block for good.

Splitting this into fork and exec looks redundant until you notice the gap between them, where the child already exists but hasn’t yet become a different program. That is where the shell does redirection: it opens the file, replaces descriptor 1, and only then calls exec. So ./program > out.txt doesn’t require any redirection support from program itself; it all happens before it starts.

These two get confused because both sound like “a process without a proper parent”.

A zombie process has already exited, but its parent hasn’t called wait yet. It uses no memory; all that is left of it is one entry in the process table with an exit code. One zombie bothers nobody, but a thousand means a bug in the parent, which isn’t cleaning up after its children, and they will eventually exhaust the PID table.

You can’t kill a zombie, it is already dead, so kill -9 has no effect on it at all. The only way to remove it is to make the parent call wait or to kill the parent itself.

An orphan process, on the other hand, is alive: it is a process whose parent died first. The kernel immediately assigns it a new parent, usually PID 1. Since init always calls wait, orphans get cleaned up correctly, and that is exactly why killing the parent cures a buildup of zombies.

A signal is the simplest way to tell a process something: one number, no data. A process can install a handler, ignore the signal, or let the default action happen.

Signal Default action When
SIGINT (2) terminate Ctrl+C
SIGTERM (15) terminate a polite request to stop, what kill sends with no arguments
SIGKILL (9) terminate forced; cannot be caught or ignored
SIGSEGV (11) terminate with core dump access to an invalid address
SIGCHLD (17) ignore a child exited, the signal for wait
SIGSTOP / SIGCONT stop / continue Ctrl+Z and fg

The kill command is badly named: it doesn’t kill anyone, it sends a signal, SIGTERM by default. A process can catch SIGTERM and shut down cleanly: finish writing a file, close connections, release locks. SIGKILL can’t be caught, so it leaves behind half-written files and locks that were never released. So always start with SIGTERM.

Everything the kernel knows about processes is available as files in /proc. These files don’t exist on disk: procfs generates their contents at the moment you read them.

Terminal window
ps -eo pid,ppid,stat,wchan:20,comm --sort=-pcpu | head

STAT shows the process state: R running or ready, S waiting (interruptible), D waiting (uninterruptible), Z zombie, T stopped. WCHAN shows which kernel function the process went to sleep in; for state D this names the cause directly.

Terminal window
cat /proc/self/status | head -20

/proc/self always points to the process doing the reading. Here you can see the PID, PPID, state, number of threads and memory used.

Terminal window
cat /proc/$$/maps

The address space sections of the current shell: the address range, permissions (r-xp, rw-p), and what exactly is mapped at those addresses. You can see the code, the heap ([heap]), the stack ([stack]) and every loaded library.

Terminal window
ls -l /proc/$$/fd

The open file table: 0, 1, 2 are the standard streams. These are the numbers the shell swaps during redirection.

Terminal window
strace -f -e trace=clone,execve,wait4 -- sh -c 'ls > /dev/null'

The full cycle in three calls: clone (this is how fork is implemented), execve in the child, wait4 in the parent. -f is needed to trace the child as well.

Create a zombie and look at it:

Terminal window
sh -c 'sleep 0.3 & exec sleep 5' & sleep 1; ps -eo pid,ppid,stat,comm | grep -w sleep

The line with state Z is the zombie: the first sleep has exited, and there is nobody to clean it up.

The exec is essential here: without it the shell stays alive, and most shells manage to reap the finished background job themselves, so no zombie appears. exec replaces the shell with the sleep process, which knows nothing about wait and can’t reap anything. After five seconds it exits, the zombie becomes an orphan and goes to PID 1, which reaps it.

“fork copies the process’s memory.” The pages are only marked read-only and are copied one at a time when they are written to. That is why a fork of a large process costs almost as much as a fork of a small one.

“A zombie eats resources, it must be killed right away.” A zombie has no memory and no CPU time, just a table entry. The problem is the parent that keeps piling them up. And you can’t kill a zombie anyway, it is already dead.

“The process in state D is hung, I’ll kill -9 it.” SIGKILL is also delivered through a signal check, and a process in uninterruptible sleep never reaches that check. It will leave state D when the I/O operation it is waiting for completes, or never, if the disk doesn’t respond.

“ps doesn’t show the process, so it doesn’t exist.” The port may be held by a child you didn’t notice, or by a process in a state your filter missed. ss -lptn shows which PID holds the socket, and that is more reliable than ps | grep.

“A PID is unique forever.” PIDs are reused after a process exits, so storing a PID and sending a signal to it a minute later is a race: that number may already belong to someone else. cgroups and pidfd exist for exactly these cases.

Check yourself

1. A process in the "ready" state and a process in the "blocked" state are both not running. What is the practical difference?
2. What does fork() return in the child process?
3. Why does the shell use fork and exec as separate calls instead of one?
4. A process sits in ps with state Z. What will kill -9 do?
5. What is the fundamental difference between SIGKILL and SIGTERM?
6. Why does a context switch cost more than saving and restoring registers?

A1: your own shell. fork, exec, wait, pipelines, redirection, job control. It applies everything in this module directly: the gap between fork and exec, swapping descriptors 0/1/2, reaping children with wait.

B2: mini-top. Reading /proc, parsing status and stat, computing CPU usage from the difference in counters between two samples.