Skip to content

A7. An HTTP server in three versions

advancedbuilds on module 13, module 8

A server for a hundred connections can be written any way you like. At ten thousand, the I/O model becomes the main decision: thread per connection runs into stack memory and context switches, an event loop on epoll hands thousands of connections to a single thread, and io_uring removes the system call for every operation. Module 13 compares these mechanisms in a table, and module 8 explains why a thread per client doesn’t scale. Here you’ll write all three servers, load them the same way, and see on a graph where each one breaks.

After this lab you will be able to:

  • write an HTTP server in three models: thread per connection, an event loop on epoll, and io_uring via liburing;
  • calculate how much memory thread stacks take for 10,000 connections, and verify it by measuring RSS;
  • explain why a single blocking call in an event loop stalls all clients, and show it with the slow-client check;
  • measure requests per second, p50, p99, RSS, and the number of context switches under load, and tell a limit of the model from a limit of the machine;
  • read three graphs and say at how many connections each model broke and why.

On the job, this is the postmortem of “the server stopped responding under load even though the CPU was idle”, and choosing a model for a new service: after this lab you’ll know what each one runs into.

Write three HTTP servers with identical behavior and a report with measurements under the same load.

Terminal window
./srv-threads 8080
./srv-epoll 8080
./srv-uring 8080

What it must do, each of the three:

  • listen on the port given as the first argument and answer GET / with a fixed HTTP/1.x response with status 200;
  • serve 20 sequential and 100 concurrent connections without losses;
  • not crash on malformed requests: an empty one, garbage instead of HTTP, a 40,000-byte request line, and a disconnect in the middle of the headers;
  • assemble a request that arrives in pieces with pauses, up to the empty line;
  • answer other clients while one has connected and sends nothing;
  • close descriptors: after 200 requests, their count in /proc/<pid>/fd doesn’t grow.

Stage 2 adds a thread pool variant to the first model; it’s measured separately.

What you don’t need to do. You don’t need to parse HTTP fully, serve files, or support keep-alive, HTTPS, or other methods: the response is the same for any valid GET /, and the connection can be closed after it.

Constraints. C and system calls: sockets, pthread, epoll, liburing. Don’t use ready-made HTTP libraries: the I/O model has to be written by you. The automated check needs python3.

What to write in the report:

  • the stack size from ulimit -s multiplied by 10,000 connections, and the actual RSS of the first server at that point;
  • the ulimit -n and net.ipv4.ip_local_port_range you set for the server and the client;
  • for each model at 100, 1,000, 5,000, and 10,000 connections: requests per second, p50, p99, RSS, voluntary_ctxt_switches, and nonvoluntary_ctxt_switches;
  • three graphs: “connections → requests/s”, “connections → p99”, “connections → memory”;
  • where each model broke and why; if the client ran on the same machine as the server, say so too.

Done when:

  • ./check.sh ./srv-epoll 8080 passes all twelve checks in seven groups, and the same for ./srv-threads and ./srv-uring;
  • measurements are taken at all four load points for each model;
  • the report covers the five points above.
  • Read the section “Blocking and non-blocking I/O” in module 13, including the Aside “Non-blocking doesn’t mean asynchronous”, and the sections “Concurrency without threads” and “How many threads” in module 8.
  • Unpack the course archive: the check is in labs/a7-http-server/check.sh; without a second argument, the port is 18080.
  • You’ll need liburing-dev, ab from apache2-utils or wrk, and a raised ulimit -n: setup/provision.sh sets it to 65536, after which you need to log in again. Any Linux will do, including a container and WSL2; for absolute numbers, it’s better to run the client on another machine.
  1. Thread per connection.

    accept in a loop, pthread_create for each connection. The simplest and easiest to understand option, so we start with it.

    Right away, calculate how much memory the stacks will take for 10,000 connections with your ulimit -s. You’ll need this number in the report.

  2. Thread pool.

    The same model, but threads aren’t created for every connection. Measure it separately: some of the overhead goes away, but the limit remains.

  3. epoll.

    One thread, non-blocking sockets, an event loop. Don’t forget that accept must be non-blocking too, and that data can arrive in pieces.

    Compare it with the first version not only on throughput but also on memory per connection.

  4. io_uring.

    Requests for accept, recv, and send are submitted to one ring, and results are collected from the other. Use liburing; you don’t have to write raw system calls.

  5. Measurement.

    Generate load with wrk, ab, or your own client. Four points are required: 100, 1,000, 5,000, and 10,000 concurrent connections.

    For each point, record: requests per second, p50 and p99 latency, the server’s RSS, and the number of context switches (/proc/<pid>/status, the fields voluntary_ctxt_switches and nonvoluntary_ctxt_switches).

  6. Report.

    Three graphs: “connections → requests/s”, “connections → p99”, “connections → memory”. And text: where each model broke and why.

The check is in the archive with the course files, and the commands below are run from the unpacked directory.

Terminal window
cd labs/a7-http-server
./check.sh ./srv-epoll 8080

The check starts the server, sends valid and deliberately malformed requests, checks the responses, holds a slow connection open, and verifies that other clients are served in the meantime.

Reading with a single read. TCP isn’t obliged to deliver the request in one piece, so you need a loop up to the empty line.

Descriptors leak. A close is missing on an error path. You can see it with ls /proc/<pid>/fd | wc -l under load.

EAGAIN is treated as an error. On a non-blocking socket it’s the normal answer “no data right now”.

One slow client stalls everyone. In the epoll version this means you made a blocking call somewhere inside the event loop.

Measuring on the same machine. The client will compete with the server for the CPU. That’s acceptable if you say so honestly in the report, but it won’t do for absolute numbers.

Add SO_REUSEPORT and several processes, each with its own epoll: this is the model nginx uses. Compare it with a single event loop on all cores.