Networking in C · intermediate · ~12 min

TLS handshake bytes — where SNI lives

- By the end you can describe why the SNI hostname in a TLS ClientHello travels in **plaintext** while the rest of the session is encrypted. - By the end you can read the SNI extension's length-prefixed layout byte by byte and decode its **big-endian** length fields by hand. - By the end you can write a `extract_sni()` parser that validates every length against both the input size and the output buffer before copying. - By the end you can spot and reject malformed or hostile SNI payloads (lying lengths, wrong type byte, truncated headers) instead of over-reading or over-writing. - By the end you can explain, from a defender's seat, how passive SNI monitors and filters use this field — and why bounds-checking the parser is the real security lesson.

Overview

You already know three things this lesson leans on directly: pointers (walking a byte buffer with an offset), endianness (multi-byte integers have a byte order), and bounded copy (never write more than the destination can hold). This lesson puts all three to work on one small, concrete artifact: the Server Name Indication (SNI) extension inside a TLS ClientHello. It is a short, length-prefixed byte structure — exactly the kind of thing pointer arithmetic and bounds checks were made for.

The scope here is narrow on purpose. We are not building a TLS client, not decrypting anything, and not walking the whole ClientHello. We take a captured SNI extension payload (a plain const uint8_t * buffer) and pull the hostname out of it safely. Everything runs on fixed in-memory buffers — no network, no root, no live traffic.

Why it matters

TLS encrypts almost the entire connection, but the very first message — the ClientHello — goes out before any keys exist, so it is in the clear. The SNI extension inside it names the host the client wants, in plaintext. That single field is what passive TLS monitors, SNI-based filters, CDNs, and traffic classifiers read to route or observe a connection without breaking encryption. Because the payload is attacker-controlled length-prefixed data, a naive parser that trusts the declared lengths is a textbook buffer-overflow and over-read bug — so the defensive skill (validate every length before you touch memory) is the whole point.

Core concepts

1. Why SNI is plaintext

TLS establishes encryption keys during the handshake. But the client has to tell the server which site it wants before those keys exist — a single IP can host thousands of virtual sites, and the server needs the name to pick the right certificate. So the hostname rides in the ClientHello, the first packet, unencrypted.

Client                                Server
  |  ClientHello  (PLAINTEXT) ------->  |   <- SNI lives here
  |  <------- ServerHello, Certificate  |
  |  ... key exchange ...               |
  |  ===== everything below encrypted = |

The SNI extension is buried inside the ClientHello's extension list. A full ClientHello walker (out of scope here) would locate that extension and hand you its payload; this lesson starts from that payload.

2. The SNI payload layout

The payload is a sequence of length-prefixed fields. Every multi-byte length is big-endian (network byte order): most-significant byte first.

offset  field                    size  notes
------  ----------------------   ----  -----------------------------
  0     server_name_list_len      2    big-endian, bytes that follow
  2     name_type                 1    0x00 = host_name
  3     name_len                  2    big-endian, length of host
  5     host bytes ...          name_len  the actual hostname (no NUL)

As raw bytes, www.a.com (a host_name) looks like:

00 0C   00   00 09   77 77 77 2e 61 2e 63 6f 6d
\---/   \/   \---/   \-------------------------/
list=12 type name=9   'w''w''w''.''a''.''c''o''m'

Big-endian decode reminder: the two bytes 01 2C are 0x01 * 256 + 0x2C = 256 + 44 = 300. Reading them in the wrong order gives 0x2C01 = 11265 — a completely different (and dangerous) length.

Field Bytes Meaning Must satisfy
list_len 2 count of bytes after it list_len <= n - 2
name_type 1 entry kind == 0x00 (host_name)
name_len 2 hostname length <= list_len - 3 and <= n - 5
host name_len the hostname name_len < cap (leave room for NUL)

Knowledge check: the payload starts with the bytes 00 09 00 00 03 61 62 63. What hostname does it carry, and is it valid?

list_len = 0x0009 = 9. name_type = 0x00 (host_name — good). name_len = 0x0003 = 3. The next 3 bytes are 61 62 63 = "abc". It is valid: 9 bytes follow the list length, and 3 ≤ 9-3 = 6, and 3 fits the buffer. Hostname is "abc".

3. Two independent trust boundaries

The payload declares its own lengths — but a hostile or corrupt buffer can lie. There are two sizes you must respect and they come from different places:

  • n — how many bytes you actually received. This is ground truth. No declared length may push your read offset past n.
  • cap — how big your output buffer is. The hostname must fit with room for a terminating '\0'.

A declared name_len of 300 means nothing if you only hold 6 bytes, or if out is a 64-byte array. Validate against both before the memcpy. This is the bounded-copy discipline applied to network data.

input buffer (n bytes)         output buffer (cap bytes)
[00 0C 00 00 09 w w w . a . c o m]   [ _ _ _ _ ... _ ]
 |<---------- must not read past ---->|   |<- must not write past ->|

Knowledge check: the header says name_len = 5 but the buffer only has 3 host bytes left. What must the parser do?

Reject it. name_len (5) > n - pos (3), so copying 5 bytes would read past the end of the input — an out-of-bounds read (classic Heartbleed-shaped bug). Return an error; never clamp silently and pretend you got a name.

4. Defensive parsing as the takeaway

Every check in the parser exists to stop a specific memory bug: reading uninitialized/adjacent memory (info leak), or writing past out (stack smash). The order matters — check sizes before dereferencing. This is exactly how a real SNI-reading monitor or filter must be written, because it parses bytes chosen by whoever is on the other end of the wire.

Syntax notes

int extract_sni(const uint8_t *ext, size_t n, char *out, size_t cap);
  • ext — pointer to the captured SNI extension payload. const because parsing never modifies input.
  • n — number of valid bytes at ext. The hard read limit; treat it as untrusted-but-real.
  • out — destination for the extracted hostname.
  • cap — total capacity of out in bytes. Room for the NUL terminator must come out of this.
  • returns — the hostname length (>= 0) on success, or -1 on any malformed/oversized input. On success out is NUL-terminated. On failure the contents of out are unspecified — callers must check the return value before using out.
uint16_t be16(const uint8_t *p);      // combine 2 bytes, MSB first: (p[0]<<8)|p[1]
void *memcpy(void *dst, const void *src, size_t n);  // no overlap, no NUL added
size_t strlen(const char *s);          // used only to build test input, not to trust network data

Key conventions: uint8_t/uint16_t from <stdint.h> give exact widths; size_t for all sizes and offsets so comparisons stay unsigned. There is nothing to free() or close() here — the buffers are caller-owned. Never call strlen() on network bytes: the hostname is length-prefixed, not NUL-terminated.

Lesson

Why this matters

TLS encrypts almost everything in a connection. The one exception is the ClientHello — the very first message the client sends.

The ClientHello goes out before any encryption keys exist, so it travels in the clear. Inside it sits the SNI extension (Server Name Indication), which names the host the client wants to reach. That hostname is plaintext.

This is the field every passive TLS monitor, SNI-based filter, and traffic classifier reads.

What the SNI extension payload looks like

The payload is a short, fixed sequence of length-prefixed fields:

[server_name_list_len : 2]   big-endian
[name_type : 1]              0x00 = host_name
[name_len : 2]               big-endian
[host bytes ...]

All multi-byte lengths are big-endian — also called network byte order. Big-endian means the most significant byte comes first. For example, the two bytes 0x01 0x2C represent the value 300 (1 * 256 + 44).

Your job

Implement this function:

int extract_sni(const uint8_t *ext, size_t n, char *out, size_t cap);

It must:

  • Read the list length.
  • Check that the type is 0x00 (host_name).
  • Read the host length.
  • Copy the host into out.

Validate every length against n (the input size) and the host length against cap (the output buffer size). Never read past the input. Never write past the output.

What this is NOT

  • It is not a TLS client or a man-in-the-middle tool. We only parse a captured extension payload.
  • It is not a full ClientHello walker. That walker wraps this same field — here we focus on the SNI field alone.

Code examples

#include <stdint.h>
#include <stddef.h>
#include <stdio.h>
#include <string.h>

/* Read a 16-bit big-endian value from p. Caller guarantees 2 bytes exist. */
static uint16_t be16(const uint8_t *p) {
    return (uint16_t)((p[0] << 8) | p[1]);
}

/*
 * Parse an SNI extension payload and copy the host_name into out.
 * Layout: [list_len:2][name_type:1][name_len:2][host...]
 * Returns host length on success, or -1 on any malformed/oversized input.
 * On success out is always NUL-terminated.
 */
static int extract_sni(const uint8_t *ext, size_t n, char *out, size_t cap) {
    size_t pos = 0;

    if (cap == 0) return -1;                 /* no room even for a NUL */
    if (n < 2) return -1;                     /* need list_len */
    uint16_t list_len = be16(ext + pos);
    pos += 2;

    /* The declared list must fit inside the bytes we actually hold. */
    if ((size_t)list_len > n - pos) return -1;

    if (list_len < 3) return -1;              /* need type(1)+name_len(2) */
    uint8_t name_type = ext[pos];
    pos += 1;
    if (name_type != 0x00) return -1;         /* 0x00 = host_name */

    uint16_t name_len = be16(ext + pos);
    pos += 2;

    /* name_len must fit both the declared list and the real buffer. */
    if ((size_t)name_len > (size_t)list_len - 3) return -1;
    if ((size_t)name_len > n - pos) return -1;

    /* Leave one byte for the terminating NUL. */
    if ((size_t)name_len >= cap) return -1;

    memcpy(out, ext + pos, name_len);
    out[name_len] = '\0';
    return (int)name_len;
}

static void try_parse(const char *label, const uint8_t *ext, size_t n) {
    char host[256];
    int r = extract_sni(ext, n, host, sizeof host);
    if (r >= 0)
        printf("%-22s -> OK, host=\"%s\" (len=%d)\n", label, host, r);
    else
        printf("%-22s -> REJECTED\n", label);
}

int main(void) {
    /* A well-formed SNI payload naming "cmagic.dev" (10 bytes). */
    const char *name = "cmagic.dev";
    size_t nlen = strlen(name);
    uint8_t good[64];
    size_t i = 0;
    uint16_t list_len = (uint16_t)(1 + 2 + nlen);   /* type + name_len + host */
    good[i++] = (uint8_t)(list_len >> 8);
    good[i++] = (uint8_t)(list_len & 0xFF);
    good[i++] = 0x00;                                /* host_name */
    good[i++] = (uint8_t)(nlen >> 8);
    good[i++] = (uint8_t)(nlen & 0xFF);
    memcpy(good + i, name, nlen);
    i += nlen;
    try_parse("well-formed", good, i);

    /* Truncated: claims a huge name but buffer is tiny. 0x012F = 303. */
    uint8_t lying[] = { 0x01, 0x2F, 0x00, 0x01, 0x2C, 'X' };
    try_parse("lying length", lying, sizeof lying);

    /* Wrong name_type (0x01 instead of 0x00). */
    uint8_t badtype[] = { 0x00, 0x06, 0x01, 0x00, 0x03, 'a', 'b', 'c' };
    try_parse("wrong name_type", badtype, sizeof badtype);

    /* Too short to hold even the list length. */
    uint8_t stub[] = { 0x00 };
    try_parse("truncated header", stub, sizeof stub);

    return 0;
}

Line by line

  • be16() — combines two bytes MSB-first into a uint16_t. This is the endianness step: the wire is big-endian, so p[0] is the high byte. The (uint16_t) cast documents the width and silences integer-promotion warnings.
  • if (cap == 0) return -1; — a zero-size output can't even hold a NUL; bail before touching anything.
  • if (n < 2) return -1; — you can't read a 2-byte list_len from fewer than 2 bytes. This is the first n guard.
  • list_len = be16(ext + pos); pos += 2; — decode the declared list length and advance the offset. pos is our single source of truth for "where am I in the buffer."
  • if (list_len > n - pos) return -1; — the declared list must fit in the bytes we actually hold. n - pos is safe because we already ensured n >= 2 == pos.
  • if (list_len < 3) return -1; — a valid entry needs at least type(1) + name_len(2); a smaller list is malformed.
  • name_type = ext[pos]; pos += 1; then if (name_type != 0x00) — only 0x00 is host_name. Any other type isn't the field we want, so reject.
  • name_len = be16(ext + pos); pos += 2; — decode the hostname length (big-endian again).
  • if (name_len > list_len - 3) return -1; — the name must fit inside what the list claimed (list_len minus the 3 bytes we consumed for type + name_len).
  • if (name_len > n - pos) return -1; — the name must fit inside the real buffer. This is the check that stops an over-read (Heartbleed-shaped bug). Both this and the previous check are needed: they defend different boundaries.
  • if (name_len >= cap) return -1; — leave one byte for '\0'; >= (not >) reserves that byte. This is the bounded-copy check protecting out.
  • memcpy(out, ext + pos, name_len); out[name_len] = '\0'; — only now, with every length verified, do we copy and terminate.
  • try_parse() / main() — build four payloads in memory (one valid, three hostile/broken) and print the verdict, proving the parser copies the good one and rejects the rest.

Common mistakes

1. Reading lengths in the wrong byte order.

uint16_t name_len = p[0] | (p[1] << 8);   // WRONG: little-endian

Why it breaks: the wire is big-endian. 00 09 becomes 0x0900 = 2304 instead of 9, so you try to copy 2304 bytes and blow past the buffer.

uint16_t name_len = (p[0] << 8) | p[1];   // FIXED: MSB first

2. Trusting the declared length without checking the real buffer.

memcpy(out, ext + 5, name_len);           // WRONG: name_len came from the wire

Why it breaks: a hostile payload sets name_len far larger than the bytes actually present — an out-of-bounds read that can leak adjacent memory.

if (name_len > n - pos) return -1;        // FIXED: bound against n first
memcpy(out, ext + pos, name_len);

3. Forgetting room for the NUL terminator.

if (name_len > cap) return -1;            // WRONG: off-by-one
memcpy(out, src, name_len);
out[name_len] = '\0';                     // writes out[cap] -> overflow

Why it breaks: when name_len == cap, out[name_len] writes one byte past the end.

if (name_len >= cap) return -1;           // FIXED: reserve the NUL byte

4. Using signed types or strlen on network bytes.

int name_len = ...; if (name_len < n) ...  // WRONG: mixed signedness
size_t len = strlen((char *)host_bytes);   // WRONG: not NUL-terminated

Why it breaks: signed/unsigned comparisons can flip; and the hostname on the wire has no '\0', so strlen runs off the end.

size_t name_len = ...;                      // FIXED: unsigned everywhere
memcpy(out, host_bytes, name_len);          // use the length prefix, not strlen

Debugging tips

  • Print the bytes. When a parse fails, dump the payload in hex (for (size_t k=0;k<n;k++) printf("%02x ", ext[k]);) and decode the first five bytes by hand: list_len, type, name_len. Compare with what the parser computed.
  • AddressSanitizer catches the exact bugs this lesson prevents: cc -std=c11 -fsanitize=address,undefined -g m.c && ./a.out. A missing bound check turns into a clear heap-buffer-overflow / stack-buffer-overflow report with the offending line.
  • valgrind (valgrind ./a.out) flags reads/writes past your buffers and use of uninitialized bytes if you ever forget to terminate out.
  • gdb: set a breakpoint on extract_sni, then p list_len, p name_len, p n, p cap at each check to see which guard should have fired. x/8xb ext shows the raw bytes.
  • Feed it garbage on purpose. Fuzz with random lengths and short buffers; a correct parser only ever returns a valid length or -1, never crashes. AFL++ or a simple loop of rand() payloads works.

Memory safety

The whole topic is a memory-safety exercise on untrusted input. The two classic hazards: out-of-bounds read (copying name_len bytes when fewer exist — leaks adjacent memory, the shape of the Heartbleed bug) and out-of-bounds write (copying more than cap into out — stack/heap smash). Defend both by validating every declared length against n (read limit) and cap (write limit) before any dereference or memcpy, and reserve one byte for the NUL with >= cap, not > cap. Keep every size and offset in size_t so comparisons are unsigned — a signed name_len could go negative and defeat a < check. Compute n - pos only after proving pos <= n, so the unsigned subtraction never wraps to a huge value. There is no concurrency and nothing to free here: the buffers are caller-owned and the function is pure with respect to its input. On the failure path, treat out as garbage — always branch on the return value.

Real-world uses

SNI parsing shows up in CDNs and load balancers (route by hostname before TLS terminates), passive network monitors and IDS/IPS, SNI-based firewalls and captive portals, and TLS libraries themselves (the server reads SNI to select a certificate). In every one of these, the parser runs on attacker-controlled bytes, so the same defensive discipline applies. Best practice: never trust a length from the wire; bound it against the actual data and the destination; prefer fixed-width stdint.h types; fail closed (reject) rather than clamp-and-continue; and if you must build the ClientHello walker that feeds this function, apply the identical checks at every nested length prefix. Real-world hardening also means limiting the maximum hostname you accept (a legal DNS name is ≤ 253 bytes) so a valid-but-absurd length can't force a huge allocation.

Practice tasks

  1. Decode by hand. Given the bytes 00 08 00 00 05 6d 61 69 6c 73, write out list_len, name_type, name_len, and the hostname. Say whether it is valid and why.
  2. Add a max-length policy. Extend extract_sni to reject any hostname longer than 253 bytes (the DNS limit) even when it would fit in out. Return a distinct error code so callers can tell "too long" from "malformed."
  3. Harden against wrap-around. Write a test that passes n = 1 and confirm the parser never computes n - pos when pos > n. Add whatever guard is missing so the unsigned subtraction can never wrap.
  4. Multiple entries. The server_name_list can technically hold more than one entry. Loop over the list, skipping non-0x00 types by their name_len, and return the first host_name found — still bounding every step against n and list_len.
  5. Fuzz it. Write a loop that generates 100000 random payloads (random length 0–32, random bytes) and asserts extract_sni either returns a length in [0, cap-1] or -1, and never crashes under -fsanitize=address.

Summary

  • The ClientHello is sent before encryption, so the SNI hostname inside it is plaintext — the one name a passive TLS monitor or filter can read.
  • The SNI payload is length-prefixed: list_len (2B) → name_type (1B, must be 0x00) → name_len (2B) → host bytes. All multi-byte lengths are big-endian.
  • Decode big-endian as (hi << 8) | lo; 01 2C = 300, not 11265.
  • Enforce two independent bounds: no read past n (real bytes) and no write past cap (output buffer), reserving one byte for the NUL with >= cap.
  • Keep sizes in size_t, check before you dereference, never strlen network bytes, and fail closed — return -1 on anything malformed rather than clamping.
  • This is the defensive core of every real SNI reader: it parses attacker-controlled bytes, so the bounds checks are the feature.

Practice with these exercises