Networking in C · intermediate · ~20 min

HTTP/1.1 protocol structure

- By the end you can read an HTTP/1.1 request or response frame byte-by-byte and name every part: request/status line, headers, blank line, and body. - By the end you can hand-write a defensive parser that splits on `CRLF` exactly, rejects bare `LF`, and caps line and header counts. - By the end you can explain how the body is framed by `Content-Length` versus `Transfer-Encoding: chunked`, and why accepting both at once is a security bug. - By the end you can describe request smuggling as parser *divergence* between two servers and state the defensive rule that stops it. - By the end you can point `curl -v` and `nc` at a running server to see the raw bytes on the wire.

Overview

You already know from sockets-intro how to open a TCP connection with socket(), connect()/accept(), and move raw bytes with read() and write(). TCP hands you an undelimited stream of bytes; it says nothing about where one message stops and the next begins. HTTP/1.1 is the text protocol that adds that structure on top of the stream. This lesson is entirely about the shape of those bytes — the framing — not about sockets themselves.

Every HTTP/1.1 message, request or response, has the same four-part layout: a start line, zero or more header lines, a blank line, and an optional body. Everything up to the blank line is US-ASCII text with lines ended by \r\n (carriage return + line feed, together called CRLF). Once you can find those boundaries in a byte buffer, you can parse HTTP.

Why it matters

Almost every web server, proxy, load balancer, and WAF you will ever read contains a hand-written HTTP parser, and framing bugs in those parsers have caused some of the worst production security failures on record — request smuggling, response splitting, and header-injection classes all live here. When a front-end proxy and a back-end server disagree by even one byte about where a request ends, an attacker can smuggle a hidden second request past your access controls. Getting the byte-level framing exactly right, and rejecting anything ambiguous, is the whole game. A parser that is merely lenient is a parser that is exploitable.

Core concepts

The four-part frame

Every HTTP/1.1 message looks like this on the wire. Spaces are shown as · and line ends as CRLF; the blank line is a CRLF on a line by itself:

 GET·/index.html·HTTP/1.1CRLF     <- start line (request line)
 Host:·example.comCRLF            <- header
 Content-Length:·5CRLF            <- header
 CRLF                             <- blank line: END of header block
 hello                            <- body (5 bytes, framed by Content-Length)

A response has the identical shape; only the start line differs — it carries a status instead of a method:

 HTTP/1.1·404·Not·FoundCRLF       <- status line: version SP code SP reason
 Content-Length:·0CRLF
 CRLF

The single most important structural fact: an empty line (\r\n with nothing before it) marks the end of the headers. Everything after those two bytes is body.

The request line: METHOD SP TARGET SP HTTP/1.x CRLF

Exactly three tokens separated by single spaces (SP), ended by CRLF:

Field Example Notes
Method GET Case-sensitive, must be uppercase per RFC 7230. Allow-list what you accept.
Target /index.html The request path (origin-form). No spaces allowed inside it.
Version HTTP/1.1 Literal HTTP/1. then 0 or 1.

Common methods: GET, HEAD, POST, PUT, DELETE, OPTIONS, PATCH. Because the fields are space-delimited, you parse by finding the first two spaces — but only within the request line, never past the terminating CRLF.

Headers: Name : OWS Value CRLF

Each header is a name, a colon, optional whitespace, and a value, ended by CRLF:

  • Names are case-insensitive — Host, host, and HOST are the same field. Compare with strncasecmp, never strncmp.
  • No space is allowed between the name and the colon. Host : x (space before :) is malformed and must be rejected — leniency here is a smuggling vector.
  • The optional whitespace around the value (OWS = spaces/tabs) is trimmed and is not part of the value.

Knowledge check: a client sends Host: example.com with three spaces on each side of the value. What is the header name, and what is the value?

Name is Host (case-insensitive, so it matches host too). The value is example.com — the surrounding OWS (optional whitespace) is stripped. The interior of the value is preserved verbatim, but leading/trailing spaces and tabs are not part of it.

Body framing: Content-Length vs Transfer-Encoding

Headers end at the blank line; the body that follows must be delimited somehow, because TCP won't tell you where it ends. HTTP/1.1 gives exactly two mechanisms:

Mechanism Header How the body length is found
Fixed length Content-Length: 5 Read exactly N bytes after the blank line.
Chunked Transfer-Encoding: chunked Read size-prefixed chunks (<hexlen>CRLF<data>CRLF), ending with a 0-length chunk.

They must never both appear in one message, and neither may appear twice. If both arrive, or either is duplicated, the frame is ambiguous — refuse it with 400 and close the connection.

Why 'refuse ambiguity' is the whole defence: request smuggling

Request smuggling exploits parser divergence between two servers in a chain (say a CDN front-end and an origin back-end). Suppose one honours Content-Length and the other honours Transfer-Encoding:

 Client --> [Front-end: uses CL] --> [Back-end: uses TE] --> app

 POST / HTTP/1.1
 Content-Length: 6      <- front-end thinks body is 6 bytes
 Transfer-Encoding: chunked   <- back-end thinks body is chunked

 0            <- back-end: chunk length 0 => body ends HERE

 GPUT /admin  <- front-end thinks this is still body;
              back-end thinks it is a NEW smuggled request

The two servers disagree about where the request ends, so the back-end treats the leftover bytes as a separate request that bypassed the front-end's checks. The fix is not clever parsing — it is a flat rule: any message with both Content-Length and Transfer-Encoding, or with a duplicated framing header, is rejected before either is trusted. Prefer Transfer-Encoding only after you've already decided the frame is unambiguous.

Defensive coding habits (do these every time)

  • Allow-list the methods you accept; reject the rest with 405/400.
  • Require CRLF. A bare LF line terminator is not HTTP/1.1 — reject it rather than 'helpfully' accepting it, because a lenient peer downstream may split it differently.
  • Cap the maximum line length and the maximum header count, and cap total header-block size. Pathological 64 KB header lines are a classic memory-exhaustion and overflow trigger.
  • Reject a space before the header colon, empty header names, and any control byte in the request line.
  • On any doubt, return 400 and close the connection — do not try to recover mid-stream.

Syntax notes

Standard library helpers you lean on for byte-accurate, allocation-free parsing:

/* Locate a byte within a bounded region (does NOT stop at NUL, unlike strchr). */
void *memchr(const void *s, int c, size_t n);
/*   s: start pointer; c: byte to find (as int); n: bytes to scan.
 *   Returns pointer to the first match, or NULL. Use this, not strchr,
 *   because network buffers are not NUL-terminated and may contain NUL. */

/* Case-insensitive bounded compare — for case-insensitive header names. */
int strncasecmp(const char *a, const char *b, size_t n);   /* <strings.h> */
/*   Returns 0 when the first n bytes match ignoring ASCII case. */

/* Exact bounded byte compare — for the literal 'HTTP/1.' version token. */
int memcmp(const void *a, const void *b, size_t n);
/*   Returns 0 when the first n bytes are byte-identical. */

Parsing conventions used below:

  • Represent every token as a (const char *ptr, size_t len) slice into the original buffer — never copy, never assume NUL-termination. Print with %.*s and an (int)len.
  • A helper find_crlf(p, end) returns a pointer to the next \r\n in [p, end), or NULL. The header block ends when find_crlf returns a pointer equal to the current position (i.e. an empty line).
  • Bounds first: check every pointer stays < end before dereferencing. The buffer is untrusted input.

Lesson

HTTP/1.1 is ASCII text carried over TCP.

A request has three parts:

  1. A line with the method, path, and version.
  2. Header lines.
  3. An empty line, then an optional body.

Responses have the same shape, except the first line carries a status code instead of a method.

Code examples

/* Hermetic HTTP/1.1 request parser: no sockets, no root.
 * We feed a fixed in-memory byte buffer (exactly what would arrive on a TCP
 * stream) and parse it defensively, the way a hardened server front-end must. */
#include <stdio.h>
#include <string.h>
#include <strings.h>   /* strncasecmp */
#include <stddef.h>

#define MAX_HEADERS   32
#define MAX_LINE      1024   /* refuse pathological long lines */

enum { PARSE_OK = 0, PARSE_BAD = 400, PARSE_SMUGGLE = 4001 };

struct header { const char *name; size_t nlen; const char *val; size_t vlen; };

struct request {
    const char *method; size_t mlen;
    const char *target; size_t tlen;
    int minor_version;                 /* the x in HTTP/1.x */
    struct header hdr[MAX_HEADERS];
    size_t nhdr;
    const char *body; size_t body_len; /* pointer into the buffer, not a copy */
};

/* Find the next CRLF at or after p, within [p,end). Returns NULL if none. */
static const char *find_crlf(const char *p, const char *end) {
    for (; p + 1 < end; p++)
        if (p[0] == '\r' && p[1] == '\n') return p;
    return NULL;
}

/* Trim leading/trailing SP and HTAB from [*s, *e). RFC 7230 OWS. */
static void trim_ows(const char **s, const char **e) {
    while (*s < *e && (**s == ' ' || **s == '\t')) (*s)++;
    while (*e > *s && ((*e)[-1] == ' ' || (*e)[-1] == '\t')) (*e)--;
}

static int parse_request(const char *buf, size_t len, struct request *r) {
    memset(r, 0, sizeof *r);
    const char *end = buf + len;

    /* --- 1. Request line: METHOD SP TARGET SP HTTP/1.x CRLF --- */
    const char *eol = find_crlf(buf, end);
    if (!eol || eol - buf > MAX_LINE) return PARSE_BAD;
    const char *sp1 = memchr(buf, ' ', (size_t)(eol - buf));
    if (!sp1) return PARSE_BAD;
    const char *sp2 = memchr(sp1 + 1, ' ', (size_t)(eol - (sp1 + 1)));
    if (!sp2) return PARSE_BAD;

    r->method = buf;          r->mlen = (size_t)(sp1 - buf);
    r->target = sp1 + 1;      r->tlen = (size_t)(sp2 - (sp1 + 1));
    const char *ver = sp2 + 1;
    if ((size_t)(eol - ver) != 8 || memcmp(ver, "HTTP/1.", 7) != 0)
        return PARSE_BAD;
    if (ver[7] != '0' && ver[7] != '1') return PARSE_BAD;
    r->minor_version = ver[7] - '0';

    /* --- 2. Header lines until a blank line (CRLFCRLF) --- */
    const char *p = eol + 2;
    int seen_cl = 0, seen_te = 0;
    for (;;) {
        const char *hend = find_crlf(p, end);
        if (!hend) return PARSE_BAD;          /* headers never terminated */
        if (hend == p) { p += 2; break; }     /* blank line ends the block */
        if (hend - p > MAX_LINE) return PARSE_BAD;
        if (r->nhdr >= MAX_HEADERS) return PARSE_BAD;

        const char *colon = memchr(p, ':', (size_t)(hend - p));
        if (!colon || colon == p) return PARSE_BAD;  /* no name */
        const char *ns = p,     *ne = colon;         /* name  */
        const char *vs = colon + 1, *ve = hend;      /* value */
        trim_ows(&vs, &ve);

        struct header *h = &r->hdr[r->nhdr++];
        h->name = ns; h->nlen = (size_t)(ne - ns);
        h->val  = vs; h->vlen = (size_t)(ve - vs);

        if (h->nlen == 14 && strncasecmp(ns, "Content-Length", 14) == 0)
            seen_cl++;
        if (h->nlen == 17 && strncasecmp(ns, "Transfer-Encoding", 17) == 0)
            seen_te++;

        p = hend + 2;
    }

    /* --- 3. Reject ambiguous framing BEFORE trusting either header --- */
    if (seen_te && seen_cl) return PARSE_SMUGGLE;  /* classic CL.TE / TE.CL */
    if (seen_cl > 1 || seen_te > 1) return PARSE_SMUGGLE; /* dup headers */

    r->body = p;
    r->body_len = (size_t)(end - p);
    return PARSE_OK;
}

static void dump(const char *label, const char *buf, size_t len) {
    struct request r;
    int rc = parse_request(buf, len, &r);
    printf("=== %s ===\n", label);
    if (rc == PARSE_SMUGGLE) { printf("  REJECTED 400: ambiguous framing (smuggling)\n\n"); return; }
    if (rc != PARSE_OK)      { printf("  REJECTED 400: malformed frame\n\n"); return; }
    printf("  method=%.*s target=%.*s version=HTTP/1.%d\n",
           (int)r.mlen, r.method, (int)r.tlen, r.target, r.minor_version);
    for (size_t i = 0; i < r.nhdr; i++)
        printf("  header[%zu] '%.*s' = '%.*s'\n", i,
               (int)r.hdr[i].nlen, r.hdr[i].name,
               (int)r.hdr[i].vlen, r.hdr[i].val);
    printf("  body_len=%zu body='%.*s'\n\n", r.body_len,
           (int)r.body_len, r.body);
}

int main(void) {
    /* Note the explicit \r\n; a bare \n would (correctly) fail to parse. */
    const char good[] =
        "GET /index.html HTTP/1.1\r\n"
        "Host:   example.com  \r\n"
        "User-Agent: demo/1.0\r\n"
        "\r\n";

    const char with_body[] =
        "POST /submit HTTP/1.1\r\n"
        "Host: example.com\r\n"
        "Content-Length: 5\r\n"
        "\r\n"
        "hello";

    const char smuggle[] =
        "POST /x HTTP/1.1\r\n"
        "Host: example.com\r\n"
        "Content-Length: 6\r\n"
        "Transfer-Encoding: chunked\r\n"
        "\r\n"
        "0\r\n\r\n";

    const char bare_lf[] =
        "GET / HTTP/1.1\n"        /* bare LF, not CRLF -> malformed */
        "Host: example.com\n"
        "\n";

    dump("well-formed GET",           good,      sizeof good - 1);
    dump("POST with Content-Length",  with_body, sizeof with_body - 1);
    dump("CL + TE (smuggling probe)", smuggle,   sizeof smuggle - 1);
    dump("bare LF terminators",       bare_lf,   sizeof bare_lf - 1);
    return 0;
}

Line by line

find_crlf(p, end) — the workhorse. It scans [p, end) for the two-byte sequence \r\n. The loop condition p + 1 < end guarantees p[1] is in bounds before we read it — this is the bounds-safety habit for untrusted buffers. Returning NULL means 'no line terminator found', which the caller treats as malformed.

trim_ows — strips leading and trailing spaces/tabs from a (start, end) slice by advancing *s and retreating *e. It never copies; it just narrows the window into the original buffer, which is how we keep the parser allocation-free.

parse_request step 1 (request line) — find_crlf locates the end of the first line. eol - buf > MAX_LINE caps the line length so a giant first line can't be processed. memchr(buf, ' ', ...) finds the first space (end of method); a second memchr starting past it finds the second space (end of target). Crucially both searches are bounded by eol, so we can never scan into the headers looking for a space. The method/target are recorded as slices. The version check is exact: the remaining eight bytes must be HTTP/1. followed by 0 or 1 — memcmp of 7 bytes plus one explicit digit check. Anything else is 400.

Step 2 (header loop) — starting just past the request line's CRLF (eol + 2), we repeatedly find the next line end. hend == p means the line is empty — that is the blank line, so we skip its CRLF and break out; everything after is body. Otherwise we enforce the length and count caps, then memchr for the colon. colon == p (colon at the very start) means an empty name — rejected. Name is [p, colon), value is (colon+1, hend) after trim_ows. We tally how many times Content-Length and Transfer-Encoding appear, comparing names case-insensitively with strncasecmp and guarding on exact length first so Content-Length-Foo can't match.

Step 3 (anti-smuggling gate) — before the parser would ever trust a length, it checks the tallies. Both present, or either duplicated, returns PARSE_SMUGGLE. Only then do we set body/body_len to the remaining bytes. This ordering is the point: reject ambiguity first, interpret second.

dump / main — feed four fixed buffers (well-formed, a POST with a body, a CL+TE smuggling probe, and a bare-LF message) and print the parse result. sizeof buf - 1 passes the byte length without the compiler's trailing NUL, matching what a socket read() would hand you.

Common mistakes

1. Splitting on bare \n instead of \r\n.

char *nl = strchr(line, '\n');   /* WRONG: accepts bare LF */

Why it breaks: HTTP/1.1 lines end with CRLF. Accepting a bare LF makes your parser disagree with a strict peer about line boundaries — the exact divergence that enables smuggling and header injection. Fix:

const char *crlf = find_crlf(line, end);   /* require BOTH bytes */
if (!crlf) return PARSE_BAD;

2. Trusting Content-Length and Transfer-Encoding together.

if (has("Content-Length")) body_len = cl;      /* WRONG: never checks TE too */

Why it breaks: if the message also carries Transfer-Encoding: chunked, a downstream server may frame the body differently and a smuggled request slips through. Fix — reject before trusting either:

if (seen_te && seen_cl) return PARSE_SMUGGLE;
if (seen_cl > 1 || seen_te > 1) return PARSE_SMUGGLE;

3. Scanning past the line end when finding delimiters.

const char *sp1 = strchr(buf, ' ');   /* WRONG: may run into the headers or off the buffer */

Why it breaks: strchr stops only at a NUL, but a network buffer isn't NUL-terminated and the space might be found far past the request line. Fix — bound every search:

const char *sp1 = memchr(buf, ' ', (size_t)(eol - buf));

4. Case-sensitive header-name comparison.

if (strncmp(name, "Host", 4) == 0) ...   /* WRONG: misses 'host', 'HOST' */

Why it breaks: header names are case-insensitive; strncmp will miss legitimate variants and can be tricked into skipping a security-relevant header. Fix:

if (nlen == 4 && strncasecmp(name, "Host", 4) == 0) ...

Debugging tips

  • See the real bytes. curl -v http://localhost:8080/ prints the exact request lines it sends (> prefix) and the response lines it receives (< ), including every header and the blank line. This is the fastest way to check your framing against a reference client.
  • Be the server, cheaply. nc -l 8080 (or ncat -l 8080) listens on a port and dumps whatever a client sends verbatim, so you can inspect an exact byte stream. Pipe through xxd or hexdump -C to see the 0d 0a (\r\n) pairs and confirm you aren't getting bare 0a.
  • printf-trace your slices. Because tokens are (ptr, len) slices, a bug often shows up as a length off by one or two (the CRLF). Print each slice with %.*s between markers, e.g. printf("[%.*s]\n", (int)len, ptr), so trailing whitespace or a stray \r becomes visible inside the brackets.
  • Valgrind for the out-of-bounds read. valgrind ./m flags any dereference past the buffer end — the classic failure when a delimiter search isn't bounded by end. An 'Invalid read of size 1' points straight at an unbounded memchr/loop.
  • gdb the boundary. Set a breakpoint at the header loop and inspect p, hend, and end; watch hend == p fire exactly at the blank line. If it never fires, your buffer is missing the terminating CRLFCRLF.

Memory safety

This parser touches untrusted network bytes, so every hazard is a read-past-the-end waiting to happen:

  • Never assume NUL-termination. A TCP buffer is raw bytes and may contain embedded \0. Use memchr/memcmp (length-bounded) rather than strchr/strcmp (NUL-bounded), and carry an explicit end pointer everywhere.
  • Check bounds before every dereference. find_crlf's p + 1 < end guard is not optional — reading p[1] one byte past the buffer is undefined behaviour and a real exploit primitive. Any two-byte lookahead needs the same guard.
  • Cap sizes to stop resource exhaustion. MAX_LINE and MAX_HEADERS bound the work an attacker can force. Without them a 64 KB header line or thousands of headers become a denial-of-service or an overflow of any fixed-size storage you copy into.
  • Slices must not outlive the buffer. Every (ptr, len) in struct request points into buf. If buf is freed or reused (e.g. the next read() overwrites it), those pointers dangle. Copy out anything you need to keep, or keep the buffer alive for the request's lifetime.
  • Integer/size_t care. Pointer differences are cast to size_t; make sure the ordering guarantees the subtraction is non-negative (we always find the later pointer first). A reversed subtraction wraps to a huge size_t and turns a bounded scan into an out-of-bounds one.

Real-world uses

Hand-written HTTP/1.1 parsers sit at the core of nginx, Apache httpd, HAProxy, Envoy, every CDN edge node, every WAF, and embedded servers in routers and IoT firmware. Client-side, libcurl and language HTTP libraries build and parse these same frames. Best practice mirrors what the demo does: treat the input as untrusted bytes, require strict CRLF, allow-list methods, cap line/header sizes, normalise header-name case for lookups, and — above all — reject ambiguous framing rather than guessing. In a proxy chain, make sure your front-end and back-end agree on framing rules; the security research on request smuggling (Kettle's classic work, and CVEs across major proxies) all traces back to two hops parsing the same bytes differently. When performance matters, production parsers use SIMD or state-machine designs (e.g. picohttpparser, the http-parser/llhttp family), but they enforce the same strictness — speed is never a reason to become lenient.

Practice tasks

  1. Write a function that takes the request line as a (ptr, len) slice and fills a struct with method, target, and version slices, returning an error if there aren't exactly two spaces or the version isn't HTTP/1.0/HTTP/1.1.

  2. Build (not parse) a valid HTTP GET request string for a given host and path, with the correct Host header and a terminating blank line, ensuring every line ends in \r\n.

  3. Extend the parser to look up a header by name case-insensitively and return its value slice, correctly returning 'not found' versus 'present but empty'.

  4. Add a validator that rejects a request whenever Content-Length and Transfer-Encoding both appear, OR when either header appears more than once, returning a distinct 'ambiguous framing' error code.

  5. Implement Content-Length body handling: parse the numeric value (rejecting non-digits, leading +/-, and overflow), then verify the buffer actually contains that many body bytes, returning 'incomplete' if the frame is short.

Summary

  • An HTTP/1.1 message is US-ASCII text: start line, header lines, a blank line, then an optional body — all lines ended by \r\n (CRLF).
  • The request line is METHOD SP TARGET SP HTTP/1.x; headers are Name: Value with case-insensitive names and trimmed surrounding whitespace.
  • The blank line ends the headers; the body is framed by either Content-Length or Transfer-Encoding: chunked — never both, and never duplicated.
  • Request smuggling is parser divergence between two servers; the defence is a flat rule: reject ambiguous framing before trusting any length.
  • Parse untrusted bytes defensively: bound every delimiter search with an end pointer, require strict CRLF, cap line and header limits, and return 400 + close on anything malformed.

Practice with these exercises