Skip to content

B2. Your own mini-top

basicbuilds on module 2, module 6

The machine is slow, you open top, and the first question is whether to trust the %CPU column. There is nothing magic about top: it reads text files from /proc twice and divides the difference by the time. In this lab you write such a monitor yourself and see what module 6 says about process states and /proc, and what module 2 says about the timer that counts ticks.

After this lab you will be able to:

  • walk /proc and collect each process’s name, state, user, CPU time used and RSS;
  • compute %CPU from the difference between two measurements and explain why a single measurement gives an average over the process’s whole lifetime;
  • read the cpu line of /proc/stat and tell “not enough CPU” from “waiting on the disk” by iowait;
  • parse /proc/<pid>/stat for a process whose name contains a space or a parenthesis, without the fields shifting;
  • explain from the state summary what a growing number of D or Z means.

This is the same work people do when top shows 100% on a core and the reason is not obvious: check where the number came from, how many processes are waiting on the disk, and whether zombies are piling up.

Write a program minitop that reads /proc and shows processes with their current CPU load, in three modes:

Terminal window
./minitop # refresh once a second
./minitop --once # one snapshot and exit, handy for checking
./minitop -n 10 # the ten heaviest processes

What it must do:

  • walk /proc without crashing when a process disappears between readdir and open;
  • show PID, user, state, %CPU, RSS and name for each process;
  • compute %CPU from the difference in utime + stime between two measurements, in --once mode too: a process that was just loading the CPU and then went quiet should show something close to zero;
  • show overall load from the cpu line in /proc/stat, with iowait separately;
  • sort by %CPU and show the first N with -n;
  • summarize states: how many processes are in R, S, D, Z;
  • parse a name with spaces and parentheses without the fields shifting;
  • in --once mode exit on its own: the automated check waits no more than 20 seconds for the snapshot.

The --once format is fixed because the automated check relies on it: one header line, then a line per process with six fields separated by spaces, and at the end a line summarizing the states.

PID USER STATE CPU RSS NAME
1 root S 0.0 12484 systemd
842 michael R 99.3 3120 sh
...
states R=2 S=181 D=0 Z=1

CPU is a percentage of one core, RSS is in kilobytes. NAME comes last and may contain spaces, so a line can have more than six fields: the first five are fixed, the rest of the line is the name. In interactive mode the format is free: there you are producing output for a human.

What not to do. Threads, a process tree, colors and keyboard control are not part of this lab. The automated check only looks at the --once output.

Constraints. Any language. The only thing forbidden is calling top or ps and parsing their output: all the information comes from /proc.

What goes in the report:

  • how you divide %CPU: per core or across all of them, and what two looping processes on one core show;
  • the value of getconf CLK_TCK on your system and where it enters the formula;
  • how many processes in state D you saw under heavy disk load compared with an idle system;
  • what happened to the first version of your program when a process disappeared during the walk, and how you handled it.

Done when:

  • ./check.sh ./minitop passes all fifteen checks;
  • a looping sh -c 'while :; do :; done' shows about 100% in your monitor, and the same process shows about 0% after the loop stops;
  • the report has the four items above.
  • Read the sections “States and transitions”, “Zombies and orphans” and “How it actually works in Linux” (the part about /proc/self/status) in module 6, and the sections “Interrupts and exceptions” and “How it actually works in Linux” (the part about /proc/interrupts) in module 2.
  • Unpack the course archive: the check is in labs/b2-mini-top/check.sh.
  • You will need getconf CLK_TCK, man 5 proc and top for comparison. Any Linux will do, including a container and WSL2.
What File Field
Process list /proc numeric directories
Name, state, PPID /proc/<pid>/stat 2, 3, 4
CPU time used /proc/<pid>/stat utime (14), stime (15)
Memory /proc/<pid>/status VmRSS
User /proc/<pid>/status Uid → /etc/passwd
System time /proc/stat the cpu line
Ticks per second — getconf CLK_TCK
  1. One snapshot. Walk /proc, collect names and states, print a table.

    Plan from the start for a process disappearing between the moment you saw its directory and the moment you opened a file in it. This is a normal situation.

  2. CPU load.

    Read utime + stime twice with an interval Δt and compute:

    %CPU = (Δ(utime + stime) / CLK_TCK) / Δt × 100

    You need two measurements because /proc holds the accumulated time since the process started, not the current load.

  3. Overall system load.

    The same for the cpu line of /proc/stat: the difference for each field, the share of non-idle in the total. Print the iowait column separately: it tells “not enough CPU” from “waiting on the disk” (module 13).

  4. Sorting and refresh. Top N by %CPU, refreshed once a second.

  5. State summary. How many processes are in R, S, D, Z. Run something disk-heavy alongside and watch D grow.

The check is in the archive with the course files, and the commands below are run from the unpacked directory.

Terminal window
cd labs/b2-mini-top
./check.sh ./minitop

The script starts a load with known behavior and checks that your monitor sees it: finds the process, shows a %CPU close to the expected value, notices zombies and does not crash when a process disappears during the walk.

Crashing on a vanished process. /proc/1234/stat stops existing between readdir and open, so treat ENOENT as a normal situation.

The process name is parsed wrong. The second field of /proc/<pid>/stat is in parentheses and may well contain both spaces and parentheses. You cannot split the line on spaces; look for the last ).

%CPU computed from a single measurement. Then you show the average over the process’s whole lifetime instead of its current load, and it is almost always close to zero.

CLK_TCK forgotten. Values in /proc are in ticks, not seconds, and a hundred ticks per second is not guaranteed by any standard.

VSZ instead of RSS. It shows the address space, not the memory in use (module 10).

Add a -H mode that shows individual threads from /proc/<pid>/task, and look at a multithreaded process (module 8).