Networking in C · intermediate · ~15 min
- By the end you can split an HTTP/1.1 request line into its three tokens (method, request-target, version) using pointer arithmetic instead of unsafe scanf. - By the end you can copy each token into a fixed-size struct field with an explicit bounds check, rejecting anything that would overflow. - By the end you can enforce the `\r\n` (CRLF) line terminator and refuse input that lacks it. - By the end you can apply an allow-list to the request-target so hostile bytes never reach the rest of your program. - By the end you can explain why a defensive request-line parser is the first line of defence in proxies, WAFs, and access loggers.
You already know how to work with C strings (NUL-terminated char arrays, strlen, strchr, memcpy) and how to group related data into a struct. This lesson puts both to work on a real-world job: reading the very first line of an HTTP request and turning it into a small, well-bounded record you can trust.
An HTTP/1.1 request begins with one plain-ASCII line — the request line — made of three space-separated tokens ending in \r\n. That is the entire input to this lesson. We do not touch sockets, headers, or bodies; we parse a fixed in-memory buffer, exactly the kind of buffer a proxy or firewall would hand you after recv(). The skill you build is turning untrusted text into a validated struct without ever writing past the end of an array.
Every web proxy, WAF, reverse proxy, and access logger starts by parsing this one line — it is the front door of the entire protocol. A parser that copies tokens with strcpy or sscanf("%s", ...) and no length limit is a classic stack buffer overflow: an attacker sends a 5,000-byte path and overwrites your return address. A parser that skips the CRLF check or accepts arbitrary bytes in the path opens the door to log injection and request smuggling. Getting the bounded, allow-list version right here means the easy attacks die at the door, before any routing or logging logic ever runs.
A raw request that lands in your receive buffer looks like this (the \r\n are real bytes, shown here as escapes):
G E T / i n d e x . h t m l H T T P / 1 . 1 \r \n
^--method--^ ^-------- request-target --------^ ^----- version -----^ ^CRLF^
sp sp
The request line is only the first line: three tokens separated by a single space (0x20) each, terminated by carriage-return + line-feed (0x0D 0x0A, written \r\n). Everything after that first CRLF — Host: and other headers, then a blank line, then any body — is out of scope here.
The tempting one-liner sscanf(buf, "%s %s %s", m, p, v) is a trap: %s has no length limit, so a long token writes past the end of m, p, or v. You can bound it with field widths (%7s %255s %15s), but scanf's whitespace handling also silently eats the CRLF and any run of spaces, so you lose the ability to reject malformed input. We parse by hand with pointers so every rule is explicit and every rejection is deliberate.
The algorithm is four steps, all pointer arithmetic — no allocation, no copying until the very end:
strstr(buf, "\r\n") finds the end of the request line. No CRLF -> reject.memchr finds the first space; memchr again (bounded by the line end) finds the second. We require exactly two spaces, which gives us three tokens.Working with half-open ranges [start, end) (a start pointer and a one-past-the-end pointer) keeps the length arithmetic honest: len = end - start, and we never rely on a NUL being in the right place inside the buffer.
Knowledge check: why require exactly two spaces rather than "at least two"?
The grammar is
method SP request-target SP version. A path token in this exercise's allow-list can never contain a space, so a third space means malformed (or smuggled) input. Splitting on the first and last space, or tolerating extra spaces, is exactly how parser mismatches between a proxy and a backend create request-smuggling bugs. Being strict — exactly two — keeps your interpretation unambiguous.
Each struct field is a fixed array. A copy is safe only if the token, plus its NUL terminator, fits:
dst[cap]: [ ][ ][ ][ ][ ] cap = 5, so max token length = 4
token: G E T len = 3 -> 3 < 5 OK, dst = "GET\0"
token: P A T C H len = 5 -> 5 >= 5 REJECT (no room for NUL)
The rule is if (len >= cap) reject; — note the >=, not >, because index cap-1 must hold the '\0'. This single check is what turns an overflow into a clean rejection.
| Field | Size | Max content | Holds after success |
|---|---|---|---|
method |
8 | 7 chars | e.g. GET, POST |
path |
256 | 255 chars | e.g. /index.html |
version |
16 | 15 chars | e.g. HTTP/1.1 |
Once parse_request_line returns 0, the rest of your program can treat out->method, out->path, and out->version as valid, NUL-terminated, length-bounded C strings. The parser is the trust boundary: untrusted bytes go in, a validated struct comes out, and everything downstream is simpler because of it.
For the path we accept only a-z A-Z 0-9 and / ? & = . - _, and reject everything else. An allow-list is safer than a block-list because you cannot forget to ban a character you never thought of (control bytes, spaces, <, \r, \n, high-bit bytes). Bytes like \n in a logged path are how log-injection forges fake log lines; the allow-list stops them before they are ever stored.
char *strstr(const char *hay, const char *needle);
Returns a pointer to the first occurrence of needle inside hay, or NULL if absent. We use it to find "\r\n". It scans a NUL-terminated string, so buf must be NUL-terminated.
void *memchr(const void *s, int c, size_t n);
Searches the first n bytes of s for byte c; returns a pointer to it or NULL. Unlike strchr, it is length-bounded — ideal for searching only up to the line end, not into the headers.
void *memcpy(void *dst, const void *src, size_t n);
Copies exactly n bytes. It does not add a NUL and does not check bounds — you must guarantee dst has room. We always call it after verifying len < cap, then write the '\0' ourselves.
int parse_request_line(const char *buf, http_req_t *out);
Parses the request line at the start of buf into *out. Returns 0 on success (all three fields are valid NUL-terminated strings), -1 on any malformed input. On failure the contents of *out are unspecified — callers must check the return value before reading fields. No heap allocation: nothing to free, nothing to close.
Every web proxy, every WAF (web application firewall), and every reverse-proxy access log begins the same way. They read the first line of an HTTP request and pull out three tokens:
GET)/index.html)HTTP/1.1)This first line is plain ASCII text and ends in \r\n (a carriage return followed by a newline). Because the whole protocol is text-driven, a C parser for it is small and worth reading closely.
This is the parser side of tools like Burp, mitmproxy, and nginx's access log. Here, we just write it ourselves.
A raw request arrives like this:
GET /index.html HTTP/1.1\r\n
Host: example.com\r\n
\r\n
The request line is the first line: three space-separated tokens, followed by \r\n.
Implement this function:
int parse_request_line(const char *buf, http_req_t *out);
The http_req_t struct holds three fixed-size (bounded) char arrays:
method[8]path[256]version[16]Return 0 on success, or -1 on any malformed input.
strcpy without a bounds check.\r\n./, alphanumerics, ?, &, =, ., -, and _. Reject any other character for this exercise.sscanf("%s %s %s", ...) without length specifiers. That is an uncontrolled write into memory.\r\n terminator.parse-http-smuggling-defence.#include <stdio.h>
#include <string.h>
#include <stdbool.h>
/* Bounded, allow-list HTTP/1.1 request-line parser.
* Lab-only: parses fixed in-memory buffers, never touches a real socket. */
typedef struct {
char method[8]; /* e.g. "GET" (max 7 chars + NUL) */
char path[256]; /* e.g. "/a?b=c" (max 255 chars + NUL) */
char version[16]; /* e.g. "HTTP/1.1" (max 15 chars + NUL) */
} http_req_t;
/* Which bytes are legal inside the request-target for this exercise. */
static bool path_char_ok(unsigned char c) {
if (c >= 'a' && c <= 'z') return true;
if (c >= 'A' && c <= 'Z') return true;
if (c >= '0' && c <= '9') return true;
return strchr("/?&=.-_", c) != NULL;
}
/* Copy [start, end) into dst[cap], NUL-terminating.
* Returns 0 on success, -1 if the token does not fit. */
static int bounded_copy(char *dst, size_t cap, const char *start, const char *end) {
size_t len = (size_t)(end - start);
if (len >= cap) return -1; /* leave room for the NUL */
memcpy(dst, start, len);
dst[len] = '\0';
return 0;
}
int parse_request_line(const char *buf, http_req_t *out) {
if (!buf || !out) return -1;
/* 1. Locate the CRLF that terminates the request line. */
const char *crlf = strstr(buf, "\r\n");
if (!crlf) return -1; /* no terminator -> reject */
const char *line_end = crlf; /* one past the last real char */
/* 2. Split on the two single spaces. Exactly two are required. */
const char *sp1 = memchr(buf, ' ', (size_t)(line_end - buf));
if (!sp1) return -1;
const char *sp2 = memchr(sp1 + 1, ' ', (size_t)(line_end - (sp1 + 1)));
if (!sp2) return -1;
const char *method_s = buf, *method_e = sp1;
const char *path_s = sp1 + 1, *path_e = sp2;
const char *version_s = sp2 + 1, *version_e = line_end;
/* 3. No empty tokens. */
if (method_e == method_s || path_e == path_s || version_e == version_s)
return -1;
/* 4. Allow-list the path bytes. */
for (const char *p = path_s; p < path_e; p++)
if (!path_char_ok((unsigned char)*p)) return -1;
/* 5. Bounded copies into the fixed-size struct fields. */
if (bounded_copy(out->method, sizeof out->method, method_s, method_e)) return -1;
if (bounded_copy(out->path, sizeof out->path, path_s, path_e)) return -1;
if (bounded_copy(out->version, sizeof out->version, version_s, version_e)) return -1;
return 0;
}
int main(void) {
const char *cases[] = {
"GET /index.html HTTP/1.1\r\nHost: example.com\r\n\r\n", /* ok */
"POST /api?id=42&x=1 HTTP/1.1\r\n", /* ok */
"GET /index.html HTTP/1.1\n", /* bad: LF only */
"GET HTTP/1.1\r\n", /* bad: empty path */
"GET /a<script> HTTP/1.1\r\n", /* bad: '<' */
"VERYLONGMETHODNAME / HTTP/1.1\r\n", /* bad: method overflow */
};
for (size_t i = 0; i < sizeof cases / sizeof cases[0]; i++) {
http_req_t r;
int rc = parse_request_line(cases[i], &r);
if (rc == 0)
printf("case %zu: OK method=%-4s path=%-18s version=%s\n",
i, r.method, r.path, r.version);
else
printf("case %zu: REJECT\n", i);
}
return 0;
}
typedef struct { char method[8]; char path[256]; char version[16]; } http_req_t; — the trust boundary. Fixed-size arrays mean the sizes are known at compile time, so sizeof gives us the caps for free. Comments record that one slot is always reserved for the NUL.path_char_ok — the allow-list, isolated in one small function so the policy is easy to audit and extend. Note the parameter is unsigned char: passing a plain char with the high bit set to functions like this can produce a negative value and undefined behaviour, so we cast at the call site.bounded_copy — the reusable safe-copy primitive. len = end - start is the token length; if (len >= cap) return -1; is the overflow guard (>= because index cap-1 must hold '\0'); memcpy moves the bytes; dst[len] = '\0' terminates. Every field goes through this one function, so there is a single place to get the check right.if (!buf || !out) return -1; — reject NULL pointers before dereferencing anything.strstr(buf, "\r\n") — finds the CRLF. If it is missing we reject immediately; this is the terminator check that scanf would have swallowed silently.memchr calls — locate exactly two spaces, each search bounded by line_end so we never scan into the headers. Two spaces yield three tokens._s / _e pointer pairs — half-open ranges [start, end) for each token. No mutation of the buffer, no reliance on embedded NULs.GET HTTP/1.1 (double space) has an empty path; we reject it here.bounded_copy calls — only now do we write into the struct, and only after every validation has passed. Any failure returns -1.main — six hermetic test buffers (two valid, four malformed) exercise every rejection path and print OK or REJECT, so the behaviour is observable without a network.1. Unbounded scanf
/* WRONG */
sscanf(buf, "%s %s %s", req.method, req.path, req.version);
Why it breaks: %s has no length limit; a 300-byte path overflows path[256] and smashes the stack — a textbook remote code execution bug.
/* FIXED: parse by hand with bounded copies, or at minimum bound every field */
sscanf(buf, "%7s %255s %15s", req.method, req.path, req.version);
(The hand-written parser in this lesson is still better, because it can also reject malformed input rather than silently truncating.)
2. Forgetting the CRLF check
/* WRONG */
const char *sp1 = strchr(buf, ' '); /* assumes a well-formed line exists */
Why it breaks: without confirming a \r\n, a partial read ("GET /ind) is parsed as if complete, and strchr may scan far past the intended line.
/* FIXED */
const char *crlf = strstr(buf, "\r\n");
if (!crlf) return -1;
3. Off-by-one in the bound (> instead of >=)
/* WRONG */
if (len > cap) return -1; /* allows len == cap */
memcpy(dst, start, len);
dst[len] = '\0'; /* writes at index cap -> one past the end */
Why it breaks: when len == cap there is no room for the terminator, so dst[len] writes one byte past the array.
/* FIXED */
if (len >= cap) return -1;
4. Block-list instead of allow-list
/* WRONG */
if (strchr(path, '<') || strchr(path, '>')) return -1; /* ban a few bad chars */
Why it breaks: you will always miss something — \n, \r, %00, spaces, high-bit bytes. Log injection and smuggling slip through the gaps.
/* FIXED: accept only known-good bytes */
for (const char *p = path_s; p < path_e; p++)
if (!path_char_ok((unsigned char)*p)) return -1;
printf("method len=%ld\n", (long)(method_e - method_s));. A negative or huge length means your pointer arithmetic is wrong.-fsanitize=address,undefined -g and run. AddressSanitizer pinpoints any one-byte overflow in bounded_copy instantly; UBSan catches the negative-char mistake in the allow-list. valgrind ./m is the fallback when sanitizers are unavailable.break parse_request_line, then next through the steps and print sp1, print crlf to see exactly which check fired. If a valid input is rejected, one of these pointers is NULL unexpectedly.\r, \n, \x00) and pipe them through xxd to confirm the bytes are what you think — many "it works on my string" bugs are really "my test string had a \n where I meant \r\n".printf("%s", out->path) prints garbage past the token, bounded_copy wrote the bytes but not the terminator — check that dst[len] = '\0' runs on every success path.len < cap (strictly less), because bounded_copy writes dst[len] = '\0'. Using <= or > here is a classic off-by-one overflow.buf must be NUL-terminated. strstr scans until it finds the needle or a NUL. If you hand it raw recv() bytes with no terminator, it can read past your buffer — a heap over-read. In real code, either NUL-terminate the received data yourself or switch to a length-bounded scan (memmem).unsigned char before char classification. Passing a char that is negative (high-bit-set byte) where a function expects a small non-negative int is undefined behaviour; the (unsigned char) cast in the allow-list loop avoids it.-1, *out is left in an unspecified state. Callers must not read the fields unless the return was 0; treat the struct as valid only inside the success branch.free and no ownership to track — a deliberate design choice that removes an entire class of memory bugs.\n-injected log lines out of your logs.scanf/strcpy on network input; parse into fixed-size fields with explicit length checks; prefer allow-lists over block-lists; validate the full grammar (exactly two spaces, a real CRLF) rather than the happy path only; and cap the total request-line length so a client cannot force you to scan megabytes looking for a CRLF.parse_request_line to take an extra const char **why out-parameter and set it to a short reason string ("no CRLF", "path too long", "bad path char") before each return -1. Print the reason in main for every rejected case.GET, POST, HEAD, PUT, DELETE. Reject anything else even if it fits in method[8]. Add a test for a well-formed but unknown method like BREW.size_t buflen parameter and use memmem (or a bounded loop) so the parser never reads past buflen bytes even if the buffer is not NUL-terminated. Add a test where the CRLF is absent and the buffer is exactly buflen bytes.%20 becomes a space and %2F becomes /, rejecting malformed escapes (%2, %GZ). Keep the decoded result in its own bounded buffer and prove it never overflows.http://evil.example/ or a target containing a second :// — since those are a common smuggling vector against naive path routers. Add both a valid /api case and a hostile absolute-URI case to your tests.\r\n.strstr/memchr and half-open [start, end) ranges; never use unbounded sscanf("%s", ...) or strcpy on network input.len < cap (strict <, to leave room for the NUL).