Networking in C · advanced · ~15 min

A tiny HTTP GET client

- By the end you can hand-write a valid HTTP/1.0 or HTTP/1.1 request as a plain byte string and send it over a connected TCP socket. - By the end you can read a full response back and parse the **status line** into version, status code, and reason phrase. - By the end you can explain why HTTP is line-oriented text, why every line ends in `\r\n`, and why the blank line matters. - By the end you can loop `read()` correctly until end-of-message and NUL-terminate a byte buffer before treating it as a C string. - By the end you can name what changes (and what does not) when you move from HTTP/1.0 to 1.1, and why HTTP/2 and /3 need a library.

Overview

You already built a TCP client in the previous lesson: you called socket(), filled a sockaddr_in, connect()ed, and moved bytes with read() and write(). TCP gives you a reliable, ordered byte stream — but a raw stream says nothing about what the bytes mean. HTTP is one agreement, layered on top, for what those bytes should say.

This lesson adds exactly that agreement. HTTP/1.x is a text protocol: your "request" is just a formatted string you write() into the socket you already know how to open, and the "response" is more text you read() back out. Everything hard about networking (addresses, the three-way handshake, retransmission) was solved in the TCP lesson; here you only learn the shape of the message.

Why it matters

Almost every service you touch — REST APIs, health-check probes, package registries, webhooks, OAuth token endpoints — speaks HTTP. Being able to form a request by hand means you can debug when a library misbehaves, write a tiny liveness checker with zero dependencies, or read a raw capture and know exactly which byte is wrong. On the security side, HTTP parsing is a classic source of bugs: request smuggling, response splitting, and buffer overflows all come from mishandling \r\n boundaries, Content-Length, or unbounded reads. Writing the parser yourself once, defensively, is the fastest way to understand why those attacks work and how careful length handling defeats them.

Core concepts

HTTP is text on top of TCP

TCP hands you an unstructured stream of bytes. HTTP/1.x adds a simple rule: a message is a sequence of CRLF-terminated lines, followed by a blank line, optionally followed by a body. "CRLF" means the two bytes carriage-return (\r, 0x0D) then line-feed (\n, 0x0A). Because it is plain text, you can build a request with snprintf and inspect a reply with printf.

A minimal GET request is four logical pieces:

GET /index.html HTTP/1.1\r\n   <- request line: METHOD SP TARGET SP VERSION
Host: example.com\r\n           <- one or more header lines: Name: value
Connection: close\r\n
\r\n                            <- BLANK line: "headers are done"

The blank line (an "empty" line that is just its own \r\n) is the part beginners forget. Without it the server keeps waiting for more headers and your read() blocks forever.

The response and its status line

The server answers with the same structure, but the first line is a status line instead of a request line:

HTTP/1.1 200 OK\r\n              <- VERSION SP CODE SP REASON
Content-Type: text/plain\r\n
Content-Length: 17\r\n
Connection: close\r\n
\r\n                            <- blank line ends the headers
Hello over HTTP!\n              <- the body (exactly Content-Length bytes)

The three fields of the status line are the first thing any client parses:

Field Example Meaning
version HTTP/1.1 protocol the server replied with
status code 200 3-digit machine-readable result
reason phrase OK human hint; ignore it in logic

Status codes come in classes by their first digit:

Class Range Meaning Common example
1xx 100–199 informational 100 Continue
2xx 200–299 success 200 OK
3xx 300–399 redirect 301 Moved Permanently
4xx 400–499 client error 404 Not Found
5xx 500–599 server error 503 Service Unavailable

Knowledge check: your request wrote the four header lines but the program hangs at the first read() and never returns. What single byte sequence is most likely missing?

The trailing blank line — the final \r\n that turns the last header into "headers complete." The server is still waiting for more headers, so it never starts its reply, so your read() blocks. Your request string must end with ...Connection: close\r\n\r\n — note the doubled CRLF.

Framing: how do you know the reply is finished?

TCP will not tell you where a "message" ends — that is HTTP's job. There are two common framing rules:

  • Connection: close: the server sends the body then closes the socket. Your read() eventually returns 0 (EOF). Read in a loop until then. This is the simplest and what our demo uses.
  • Content-Length: N: read exactly N body bytes after the blank line. Required when the connection stays open (HTTP/1.1 keep-alive) so you know when one response ends and the next begins.

A robust client reads in a loop regardless, because a single read() may return only part of the response even when everything is on the way:

time ->
 write(request) ........
 read() -> "HTTP/1.1 200 OK\r\nContent-Ty"   (partial! TCP split it)
 read() -> "pe: text/plain\r\n...\r\n\r\nHello"
 read() -> 0                                   (EOF, done)

Never assume one read() equals one message. That assumption is the single most common HTTP-client bug.

HTTP/1.0 vs 1.1 vs 2/3

Version Wire format Default connection Host required? Parse by hand?
HTTP/1.0 text lines closes after one reply no yes
HTTP/1.1 text lines keep-alive by default yes yes
HTTP/2 binary frames, multiplexed persistent (pseudo-headers) no — use a library
HTTP/3 binary over QUIC/UDP persistent (pseudo-headers) no — use a library

HTTP/1.1 is the line-protocol baseline: it is still readable text, so it is what you write by hand. It made the Host header mandatory (so one IP can serve many sites) and made keep-alive the default (so you should send Connection: close in a throwaway client, or handle Content-Length framing). HTTP/2 and /3 pack headers into binary frames and multiplex many requests on one connection — you cannot reasonably hand-roll them, so you reach for a library there.

Stay defensive and lab-only

Every example here targets 127.0.0.1 (loopback) and, in the demo, a mock server the program starts itself in a background thread — no real network, no third party. Point a hand-written client only at servers you own or are explicitly authorized to test. When parsing responses, treat every incoming byte as hostile: cap field widths, never index past what you actually read, and always NUL-terminate before using string functions.

Syntax notes

#include <sys/socket.h>
#include <netinet/in.h>
#include <arpa/inet.h>
#include <unistd.h>

int socket(int domain, int type, int protocol);
// AF_INET + SOCK_STREAM = a TCP socket. Returns a file descriptor >= 0, or -1 on error (sets errno).

int connect(int fd, const struct sockaddr *addr, socklen_t len);
// Opens the TCP connection to addr. 0 on success, -1 on error. Blocks until the handshake completes.

ssize_t write(int fd, const void *buf, size_t n);
// Sends up to n bytes. Returns bytes actually written (may be < n: a SHORT WRITE) or -1. Loop until all sent.

ssize_t read(int fd, void *buf, size_t n);
// Receives up to n bytes. Returns count read, 0 at EOF (peer closed), or -1 on error. May return < n: a SHORT READ.

uint32_t htonl(uint32_t x);  uint16_t htons(uint16_t x);
// Host-to-network byte order for addresses/ports. INADDR_LOOPBACK is 127.0.0.1; port 0 asks the kernel to pick.

int getsockname(int fd, struct sockaddr *addr, socklen_t *len);
// After bind(fd, ...port 0), reveals the ephemeral port the kernel assigned. 0 / -1.

int sscanf(const char *s, const char *fmt, ...);
// Parse the status line. Use WIDTH LIMITS: "%15s %d %63[^\r\n]" caps fields so they cannot overflow.
// Returns the number of fields successfully assigned.

Ownership and cleanup: every fd from socket()/accept() must be close()d. There is no memory to free() in this demo (all buffers are fixed stack arrays), but the socket file descriptors are a resource — leaking them exhausts the process fd table.

Lesson

How HTTP/1.0 works

HTTP/1.0 is a text-based protocol that runs on top of TCP. Because it is text, you can write a request as a plain string and read the reply the same way.

The exchange has three steps:

  1. Connect a TCP socket to the server.

  2. Write the request. A minimal GET request looks like this:

    GET /path HTTP/1.0\r\nHost: example\r\n\r\n
    

    Each line ends with \r\n (carriage return + line feed). The blank line at the end (an extra \r\n) tells the server the request is finished.

  3. Read the response from the same socket.

Stay on localhost

For the C-platform exercises, only target localhost. The test harness starts a local server in a separate thread for you to connect to.

Never point such a client at third-party hosts without explicit permission.

Code examples

// A tiny HTTP GET client — hermetic loopback demo.
// A background thread runs a mock HTTP server on 127.0.0.1:<ephemeral>.
// main() connects, sends one HTTP/1.1 GET, reads the whole reply,
// and parses the status line into version / code / reason.
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <unistd.h>
#include <errno.h>
#include <arpa/inet.h>
#include <sys/socket.h>
#include <netinet/in.h>
#include <pthread.h>

// Fixed body the mock server always returns.
static const char *BODY = "Hello over HTTP!\n";

// The server thread receives the listening fd via this struct.
struct srv { int listen_fd; };

static void *server_main(void *arg) {
    struct srv *s = arg;
    int c = accept(s->listen_fd, NULL, NULL);   // block until client connects
    if (c < 0) return NULL;

    // Drain the request (read until the blank line \r\n\r\n).
    char req[2048];
    size_t used = 0;
    while (used < sizeof req - 1) {
        ssize_t n = read(c, req + used, sizeof req - 1 - used);
        if (n <= 0) break;
        used += (size_t)n;
        req[used] = '\0';
        if (strstr(req, "\r\n\r\n")) break;      // end of request headers
    }

    // Build a well-formed HTTP/1.1 response with Content-Length.
    char resp[512];
    int len = snprintf(resp, sizeof resp,
        "HTTP/1.1 200 OK\r\n"
        "Content-Type: text/plain\r\n"
        "Content-Length: %zu\r\n"
        "Connection: close\r\n"
        "\r\n"
        "%s",
        strlen(BODY), BODY);

    // Write the whole response, handling short writes.
    ssize_t off = 0;
    while (off < len) {
        ssize_t w = write(c, resp + off, (size_t)(len - off));
        if (w <= 0) break;
        off += w;
    }
    close(c);
    return NULL;
}

int main(void) {
    // 1. Create a listening socket bound to an ephemeral loopback port.
    int lfd = socket(AF_INET, SOCK_STREAM, 0);
    if (lfd < 0) { perror("socket"); return 1; }

    struct sockaddr_in addr;
    memset(&addr, 0, sizeof addr);
    addr.sin_family = AF_INET;
    addr.sin_addr.s_addr = htonl(INADDR_LOOPBACK); // 127.0.0.1 only
    addr.sin_port = 0;                             // 0 = kernel picks a free port
    if (bind(lfd, (struct sockaddr *)&addr, sizeof addr) < 0) { perror("bind"); return 1; }
    if (listen(lfd, 1) < 0) { perror("listen"); return 1; }

    // Ask the kernel which port it actually assigned.
    socklen_t alen = sizeof addr;
    if (getsockname(lfd, (struct sockaddr *)&addr, &alen) < 0) { perror("getsockname"); return 1; }
    int port = ntohs(addr.sin_port);
    printf("mock server listening on 127.0.0.1:%d\n", port);

    // 2. Start the server thread.
    struct srv s = { lfd };
    pthread_t tid;
    if (pthread_create(&tid, NULL, server_main, &s) != 0) { perror("pthread_create"); return 1; }

    // 3. Client: connect a fresh socket to that same port.
    int cfd = socket(AF_INET, SOCK_STREAM, 0);
    if (cfd < 0) { perror("socket"); return 1; }
    if (connect(cfd, (struct sockaddr *)&addr, sizeof addr) < 0) { perror("connect"); return 1; }

    // 4. Send exactly one HTTP/1.1 GET request. Note the required blank line.
    const char *req =
        "GET / HTTP/1.1\r\n"
        "Host: localhost\r\n"
        "Connection: close\r\n"
        "\r\n";
    size_t reqlen = strlen(req), sent = 0;
    while (sent < reqlen) {
        ssize_t w = write(cfd, req + sent, reqlen - sent);
        if (w <= 0) { perror("write"); return 1; }
        sent += (size_t)w;
    }

    // 5. Read the full response. read() returns 0 at EOF (server closed).
    char buf[4096];
    size_t total = 0;
    for (;;) {
        ssize_t n = read(cfd, buf + total, sizeof buf - 1 - total);
        if (n < 0) { perror("read"); return 1; }
        if (n == 0) break;                 // EOF: peer closed the connection
        total += (size_t)n;
        if (total >= sizeof buf - 1) break; // buffer full; stop to stay safe
    }
    buf[total] = '\0';                      // NUL-terminate for string parsing
    printf("received %zu bytes\n", total);

    // 6. Parse the status line: "HTTP/1.1 200 OK\r\n".
    char version[16] = {0}, reason[64] = {0};
    int code = 0;
    // %15s / %63s cap the fields so they can't overflow the buffers.
    if (sscanf(buf, "%15s %d %63[^\r\n]", version, &code, reason) >= 2) {
        printf("status: version=%s code=%d reason=%s\n", version, code, reason);
    } else {
        printf("could not parse a status line\n");
    }

    // 7. Find the body: it starts right after the blank line.
    char *body = strstr(buf, "\r\n\r\n");
    if (body) printf("body: %s", body + 4);

    close(cfd);
    pthread_join(tid, NULL);
    close(lfd);
    return 0;
}

Line by line

  • BODY and struct srv: the mock server always returns one fixed string; the struct just hands the listening fd to the thread so it knows which socket to accept() on.
  • server_main / accept: the background thread blocks in accept() until the client connects, then gets a new fd (c) for that one conversation — the listening fd keeps listening.
  • The server's read loop: it drains the request into req, stopping as soon as it sees \r\n\r\n. This is real HTTP framing — the server must know the request headers are complete before it replies. It also NUL-terminates after each read so strstr is safe.
  • snprintf response: builds a correct response — status line, headers, Content-Length set to strlen(BODY), then the blank line, then the body. snprintf bounds the write to sizeof resp so it can never overflow.
  • The server's write loop: write() may send fewer bytes than asked (short write), so it loops on an offset until the whole response is out.
  • Step 1 (socket/bind/listen): INADDR_LOOPBACK restricts the listener to 127.0.0.1 — nothing off the machine can reach it. sin_port = 0 tells the kernel to pick any free port, which avoids "address already in use" in a test harness.
  • Step 1 getsockname: because we asked for port 0, we must ask back which port we actually got, so the client can connect to it.
  • Step 2: launch the server thread before the client connects, so someone is listening when connect() fires.
  • Step 3–4: the client opens its own fresh socket, connect()s to the discovered port, then sends the request. The request string ends in \r\n\r\n — three header lines and then the mandatory blank line.
  • Step 4 write loop: same short-write discipline as the server; keep writing until sent == reqlen.
  • Step 5 read loop: the heart of the lesson. It calls read() repeatedly, appending into buf, until read() returns 0 (the server closed after Connection: close). It also stops if the buffer is nearly full, leaving one byte for the terminator so it can never overflow.
  • buf[total] = '\0': turns the raw bytes into a valid C string so sscanf/strstr are safe.
  • Step 6 sscanf: parses the status line with width-limited conversions — %15s for the version, %d for the code, %63[^\r\n] for the reason (everything up to the CR/LF). The width caps guarantee the fields cannot overflow their arrays even on a malicious reply.
  • Step 7 strstr: locates the header/body separator and prints everything after it — the body begins 4 bytes past the start of \r\n\r\n.
  • Cleanup: close(cfd), pthread_join (wait for the server thread to finish so we don't tear down mid-reply), then close(lfd).

Common mistakes

1. Forgetting the trailing blank line.

// WRONG: request "ends" after the Host header
const char *req = "GET / HTTP/1.1\r\nHost: localhost\r\n";

Why it breaks: the server never sees \r\n\r\n, so it thinks more headers are coming and never replies. Your read() blocks forever.

// FIXED: doubled CRLF marks "headers complete"
const char *req = "GET / HTTP/1.1\r\nHost: localhost\r\nConnection: close\r\n\r\n";

2. Assuming one read() returns the whole response.

// WRONG: a single read may get only part of the reply
ssize_t n = read(cfd, buf, sizeof buf - 1);
buf[n] = 0;

Why it breaks: TCP can split the response across many read()s (and n can even be -1; buf[-1] is out-of-bounds). You parse a truncated message.

// FIXED: loop until EOF, tracking the total
size_t total = 0; ssize_t n;
while ((n = read(cfd, buf + total, sizeof buf - 1 - total)) > 0) total += (size_t)n;
buf[total] = 0;

3. Using \n instead of \r\n.

const char *req = "GET / HTTP/1.1\nHost: localhost\n\n"; // WRONG line endings

Why it breaks: HTTP lines are defined as CRLF-terminated. Strict servers reject or mis-frame a bare-LF request, and bare-LF handling differences are exactly where request-smuggling bugs live.

const char *req = "GET / HTTP/1.1\r\nHost: localhost\r\n\r\n"; // FIXED

4. Parsing the status line with an unbounded %s.

char version[16], reason[64];
sscanf(buf, "%s %d %[^\r\n]", version, &code, reason); // WRONG: no width caps

Why it breaks: a hostile or malformed reply with a very long version/reason overflows the stack buffers — a classic remote memory-corruption bug.

sscanf(buf, "%15s %d %63[^\r\n]", version, &code, reason); // FIXED: capped fields

Debugging tips

  • See the actual bytes. Print the buffer with visible escapes, or pipe through od -c / cat -A, so you can literally see whether your line endings are \r\n or a bare \n, and whether the blank line is present.
  • read() hangs forever. Almost always a missing final \r\n in the request, or a server that uses Content-Length framing while you wait for EOF. Check the request string ends in a doubled CRLF.
  • connect() fails with ECONNREFUSED. Nothing is listening on that port. In the demo, make sure the server thread starts before the client connects; in general, confirm the port with getsockname and that the server is up.
  • Watch the syscalls. On Linux, strace -e trace=network,read,write ./m shows every read/write and their return counts — you will see short reads and the final read() = 0 at EOF. On macOS use dtruss.
  • Check for buffer bugs. Run under valgrind ./m (Linux) or compile with -fsanitize=address,undefined to catch off-by-one indexing, missing NUL terminators, or reads past total.
  • Confirm your parse. After sscanf, print each field with delimiters like [%s] so a stray trailing \r in the reason phrase is obvious.

Memory safety

  • Always leave room for the NUL. Read into sizeof buf - 1 and terminate with buf[total] = '\0' before any str*/sscanf call. Treating un-terminated network bytes as a C string reads past the buffer — undefined behaviour and a common CVE pattern.
  • read() can return -1. Never index buf[n] with the raw return value; check n > 0 first. buf[-1] = 0 is an out-of-bounds write.
  • Cap every parsed field. Width-limited conversions (%15s, %63[^\r\n]) turn attacker-controlled response bytes into a bounded copy. An unbounded %s on network input is a remote stack-smash.
  • Never trust Content-Length blindly. If you frame by length, validate it against how many bytes you are willing to buffer; a huge or negative value must not drive an allocation or a copy size. Bounds-check before you read.
  • Concurrency in the demo. The server thread only reads s.listen_fd, which main sets before pthread_create and never mutates afterward, so there is no data race. pthread_join before close(lfd) ensures the thread has finished using the fd. If you extend it to share mutable state, add a mutex.
  • Close every fd. cfd, the accepted c, and lfd are all closed; leaking descriptors eventually exhausts the table and new socket()/accept() calls fail with EMFILE.

Real-world uses

  • Health checks and probes. Load balancers, Kubernetes liveness/readiness probes, and monitoring agents send a bare GET /healthz HTTP/1.1 and only look at the status code — exactly what you just built.
  • Minimal API clients on constrained devices. Embedded and IoT firmware without room for a full HTTP library often hand-roll HTTP/1.1 over a socket to hit a REST endpoint.
  • Tooling and diagnostics. curl -v, nc, and telnet let engineers type raw requests; understanding the wire format lets you reproduce and debug what a library does under the hood.
  • Security testing (authorized). Fuzzers and scanners craft deliberately malformed requests to probe how servers frame \r\n, Content-Length, and chunked bodies — the root of request-smuggling research. Best practice: only against systems you own or are contracted to test.
  • Best practice going forward. For anything production-facing use a maintained library (libcurl, and TLS via OpenSSL) that handles chunked encoding, redirects, timeouts, and certificate validation. Hand-rolled HTTP is for learning, health checks, and controlled tools — not for talking to the open internet with secrets.

Practice tasks

  1. Change the target. Modify the request to GET /status HTTP/1.1 and have the mock server reply 404 Not Found with a short body. Confirm your parser prints code=404.
  2. Extract one header. After reading the full response, scan the header block and print the value of Content-Type: (case-insensitively). Handle the header being absent.
  3. Frame by Content-Length. Parse the Content-Length header, then read exactly that many body bytes after the blank line instead of relying on EOF. Print how many body bytes you got.
  4. Robust status-line parser. Write a function int parse_status(const char *buf, int *code) that returns 0 on success and -1 if the first line is not a valid HTTP/x.y NNN ... status line. Reject a missing version, a non-numeric code, or a code outside 100–599.
  5. Timeout guard. Set a receive timeout on the client socket with setsockopt(SO_RCVTIMEO) so a server that never replies makes read() fail with EAGAIN instead of hanging, and report the timeout cleanly. Keep everything on loopback.

Summary

  • HTTP/1.x is plain text carried over the TCP socket you already know how to open: format a request string, write() it, read() the reply.
  • A request is METHOD TARGET VERSION\r\n, header lines, then a mandatory blank line (\r\n\r\n). Forgetting it hangs the client.
  • The response begins with a status line — version, 3-digit code, reason — which you parse first; codes group by leading digit (2xx ok, 4xx client error, 5xx server error).
  • Never assume one read() = one message. Loop until EOF (with Connection: close) or until you have read Content-Length body bytes.
  • Parse defensively: width-cap every field, always NUL-terminate, check for read() returning 0 or -1, and close every fd. Stay on loopback / authorized hosts.
  • HTTP/1.1 is the hand-writable line-protocol baseline; HTTP/2 and /3 are binary — use a library.

Practice with these exercises