Processes
Why this matters
Section titled “Why this matters”You run ./server, it crashes, you restart it and get
bind: address already in use, even though ps shows no server at all.
You press Ctrl+C and the process doesn’t die. You run kill -9,
and it still hangs around in the list with state Z.
Each of these has an exact explanation, and all three come down to the same
thing: what the operating system considers a process, what states it can be in,
and who cleans up after it dies. That is what this module covers: from fork
to a table entry that somebody has to read.
Prerequisites. A picture of interrupts and the timer (module 2) and of the boundary between user space and kernel space (module 3).
A program is not a process
Section titled “A program is not a process”A program is a file on disk, a sequence of bytes in a format the kernel knows how to load; in Linux that format is ELF. It is entirely passive and does nothing on its own.
A process is a program being executed, together with everything the kernel
set up for it: an address space, open files, a current directory, credentials,
a program counter, a stack. Forty processes can run from the single file
/usr/bin/python3 at once, and all forty will be independent.
Address space
Section titled “Address space”Each process sees its own memory and cannot see anyone else’s. What it sees is divided into sections with different access permissions.
- Code (text) holds instructions. It can be read and executed but not written: an attempt to overwrite its own code kills the process.
- Data (data, bss) holds global and static variables:
dataholds the ones with an initial value,bssthe ones that start at zero and therefore take no space in the file. - The heap holds dynamic memory handed out by
mallocand grows toward higher addresses. - The stack holds function call frames and local variables, and grows toward the heap.
The gap between the heap and the stack is not free memory belonging to the
process: these are addresses with nothing mapped to them, and touching them
causes a page fault that the kernel cannot service,
so the process gets SIGSEGV.
Process control block
Section titled “Process control block”Everything the kernel knows about a process lives in one structure, the
process control block. In Linux it is struct task_struct, and it is a
sizable structure, several hundred fields. The main groups are:
| Group | What it holds | Why |
|---|---|---|
| Identification | PID, PPID, UID, GID | who it is and on whose behalf it runs |
| State | current state, exit code | what to do with it next |
| Context | program counter, registers, stack pointer | to resume execution after a switch |
| Memory | pointer to the address space, section bounds | so the MMU knows what to translate |
| I/O | open file table, working directory | so read(3, …) knows what “3” means |
| Scheduling | priority, nice, CPU time used |
so the scheduler knows whom to let run |
| Relationships | parent, children, process group | to build the tree and deliver signals |
Process control blocks live in kernel memory. A process can neither read its own block directly nor tweak its priority there; that is what system calls are for.
States and transitions
Section titled “States and transitions”The pair most often confused here is “ready” and “blocked”, even though the difference is entirely practical. If the system is slow and processes sit in the “ready” state, there isn’t enough CPU. If they are blocked, the CPU has nothing to do with it, and you should look for the bottleneck in the disk or the network.
The transition from “running” to “ready” is called preemption: the timer fired, the quantum ran out, the kernel took the CPU away. The process did nothing wrong and isn’t waiting for anything; it was simply put back in the queue.
Context switch
Section titled “Context switch”To take one process off a CPU core and put another on, the kernel saves the first one’s registers into its process control block, loads the second one’s registers and switches the address space. This is a context switch. It has a cost, both direct and indirect.
The direct cost is measured in single microseconds, but the indirect one is larger. The new process arrives to cold caches and an empty TLB, so its first few thousand instructions run slower than they could. This is why the time quantum isn’t made very small: the more often you switch, the larger the share of the CPU spent on the switching itself.
Creating a process
Section titled “Creating a process”In Unix a new process is always cloned from an existing one, and then, if needed, replaces its image with a different program.
-
fork()creates a child process, an almost exact copy of the parent: the same open files, the same address space, the same point of execution. The only differences are the PID and whatforkitself returns: the child’s PID to the parent, zero to the child. -
execve()replaces the process’s address space with a new program. The PID stays the same, because it is the same process, now running different code. -
exit()terminates the process. The kernel frees its memory and closes its files, but keeps the process control block for now: it holds the exit code, which nobody has read yet. -
wait()in the parent reads that code and removes the block for good.
Splitting this into fork and exec looks redundant until you notice the gap
between them, where the child already exists but hasn’t yet become a different
program. That is where the shell does redirection: it opens the file,
replaces descriptor 1, and only then calls exec. So ./program > out.txt
doesn’t require any redirection support from program itself; it all
happens before it starts.
Zombies and orphans
Section titled “Zombies and orphans”These two get confused because both sound like “a process without a proper parent”.
A zombie process has already exited, but its parent hasn’t called wait yet.
It uses no memory; all that is left of it is one entry in the process table
with an exit code. One zombie bothers nobody, but a thousand means a bug in
the parent, which isn’t cleaning up after its children, and they will
eventually exhaust the PID table.
You can’t kill a zombie, it is already dead, so kill -9 has no effect on it
at all. The only way to remove it is to make the parent call wait
or to kill the parent itself.
An orphan process, on the other hand, is alive: it is a process whose parent
died first. The kernel immediately assigns it a new parent, usually PID 1. Since
init always calls wait, orphans get cleaned up correctly, and that is exactly
why killing the parent cures a buildup of zombies.
Signals
Section titled “Signals”A signal is the simplest way to tell a process something: one number, no data. A process can install a handler, ignore the signal, or let the default action happen.
| Signal | Default action | When |
|---|---|---|
SIGINT (2) |
terminate | Ctrl+C |
SIGTERM (15) |
terminate | a polite request to stop, what kill sends with no arguments |
SIGKILL (9) |
terminate | forced; cannot be caught or ignored |
SIGSEGV (11) |
terminate with core dump | access to an invalid address |
SIGCHLD (17) |
ignore | a child exited, the signal for wait |
SIGSTOP / SIGCONT |
stop / continue | Ctrl+Z and fg |
The kill command is badly named: it doesn’t kill anyone, it sends a signal,
SIGTERM by default. A process can catch SIGTERM and shut down cleanly:
finish writing a file, close connections, release locks. SIGKILL can’t be
caught, so it leaves behind half-written files and locks that were never
released. So always start with SIGTERM.
How it actually works in Linux
Section titled “How it actually works in Linux”Everything the kernel knows about processes is available as files in /proc.
These files don’t exist on disk: procfs generates their contents at the moment
you read them.
ps -eo pid,ppid,stat,wchan:20,comm --sort=-pcpu | headSTAT shows the process state: R running or ready, S waiting (interruptible),
D waiting (uninterruptible), Z zombie, T stopped. WCHAN shows which kernel
function the process went to sleep in; for state D this names the cause directly.
cat /proc/self/status | head -20/proc/self always points to the process doing the reading. Here you can see
the PID, PPID, state, number of threads and memory used.
cat /proc/$$/mapsThe address space sections of the current shell: the address range, permissions
(r-xp, rw-p), and what exactly is mapped at those addresses. You can see the
code, the heap ([heap]), the stack ([stack]) and every loaded library.
ls -l /proc/$$/fdThe open file table: 0, 1, 2 are the standard streams. These are the
numbers the shell swaps during redirection.
strace -f -e trace=clone,execve,wait4 -- sh -c 'ls > /dev/null'The full cycle in three calls: clone (this is how fork is implemented),
execve in the child, wait4 in the parent. -f is needed to trace the child as well.
Create a zombie and look at it:
sh -c 'sleep 0.3 & exec sleep 5' & sleep 1; ps -eo pid,ppid,stat,comm | grep -w sleepThe line with state Z is the zombie: the first sleep has exited, and there
is nobody to clean it up.
The exec is essential here: without it the shell stays alive, and most shells
manage to reap the finished background job themselves, so no zombie appears.
exec replaces the shell with the sleep process, which knows nothing about
wait and can’t reap anything. After five seconds it exits, the zombie
becomes an orphan and goes to PID 1, which reaps it.
Common misconceptions
Section titled “Common misconceptions”“fork copies the process’s memory.” The pages are only marked read-only
and are copied one at a time when they are written to. That is why a fork of
a large process costs almost as much as a fork of a small one.
“A zombie eats resources, it must be killed right away.” A zombie has no memory and no CPU time, just a table entry. The problem is the parent that keeps piling them up. And you can’t kill a zombie anyway, it is already dead.
“The process in state D is hung, I’ll kill -9 it.” SIGKILL is also
delivered through a signal check, and a process in uninterruptible sleep never
reaches that check. It will leave state D when the I/O operation it is
waiting for completes, or never, if the disk doesn’t respond.
“ps doesn’t show the process, so it doesn’t exist.” The port may be held
by a child you didn’t notice, or by a process in a state your filter missed.
ss -lptn shows which PID holds the socket, and that is more reliable than ps | grep.
“A PID is unique forever.” PIDs are reused after a process exits, so storing
a PID and sending a signal to it a minute later is a race: that number may
already belong to someone else. cgroups and pidfd exist for exactly these cases.
Check yourself
A1: your own shell. fork, exec, wait, pipelines, redirection,
job control. It applies everything in this module directly: the gap between fork
and exec, swapping descriptors 0/1/2, reaping children with wait.
B2: mini-top. Reading /proc, parsing status and stat, computing
CPU usage from the difference in counters between two samples.
Sources
Section titled “Sources”- OSTEP, chapters 4–5: Abstraction: The Process, Process API
- Silberschatz, Operating System Concepts, chapter 3
- Tanenbaum, Modern Operating Systems, section 2.1
man 2 fork,man 2 execve,man 2 wait,man 7 signal,man 5 proc- Documentation/filesystems/proc.rst: what each field means