Networking in C · intermediate · ~12 min

Pull the method out of a SIP message

- By the end you can identify the request line of a SIP message and locate the method token within it. - By the end you can write a bounded parser that copies the method into a caller-supplied buffer without overflowing it. - By the end you can validate an untrusted token character-by-character and reject anything that is not an uppercase ASCII letter. - By the end you can design a scan that has a hard upper bound, so a message with no delimiter can never make your loop run away. - By the end you can choose sensible return values (bytes written vs. -1) that let a caller distinguish success from every failure mode.

Overview

SIP (Session Initiation Protocol) is a text protocol that looks almost identical to HTTP: a request line, headers, a blank line, then an optional body. A SIP request line begins with a method — INVITE, REGISTER, OPTIONS, BYE, and friends — followed by a space, a URI, and the version. Your job in this lesson is narrow and defensive: pull out just that first token, safely, from a buffer you do not trust.

This builds directly on your two prerequisites. From strings you already know that C text is a run of char ending in a '\0', that there is no length stored anywhere, and that you must walk bytes yourself. From bounded-copy you already know the golden rule: never write more bytes into a destination than its capacity, and always leave room for the terminator. Here we apply both ideas to a real, attacker-reachable input — the very first bytes of a packet arriving from the network.

Why it matters

SIP gateways and PBXes sit on the public internet and receive UDP and TCP from anyone who can reach port 5060. The request line is the first thing your code touches, so it is the first place a malformed or hostile message can bite. A parser that scans for a space without a hard limit can be walked off the end of a buffer by a packet that simply omits the space; one that copies the token without checking capacity is a classic stack-smash. Rejecting a bad method in a dozen well-bounded lines is cheap insurance against a remote crash or worse.

Core concepts

The shape of a SIP request line

A SIP request starts with three space-separated fields on one line, terminated by CRLF (\r\n):

INVITE sip:alice@example.com SIP/2.0\r\n
^^^^^^ ^^^^^^^^^^^^^^^^^^^^^^^ ^^^^^^^
method        request-URI      version

The method is everything from the first byte up to (but not including) the first space. That is the entire span we care about in this lesson. Everything after the first space — the URI, the version, and every header below — is out of scope.

Laid out byte-by-byte, the start of the buffer looks like this:

index:  0   1   2   3   4   5   6   7 ...
byte : 'I''N''V''I''T''E'' ''s' ...
                          ^
                          first space => method is bytes [0,6)

A bounded scan, not an open-ended search

The naive version says "loop until you hit a space." But what if the attacker never sends a space? Then the loop reads past the end of your buffer — undefined behaviour, and on a network daemon a remotely triggerable one. The fix is a hard cap on how far you are willing to look. The longest real SIP method (SUBSCRIBE) is 9 characters, so a scan window of 16 bytes is generous and still safe:

for (i = 0; i < 16; i++) {
    if (msg[i] == ' ') break;   // found delimiter
    if (!is_upper(msg[i])) return -1;   // reject bad byte
}
if (i == 16) return -1;   // no space within the window

Notice the loop can only exit three ways: it hits a space (good), it hits a bad byte (rejected), or it exhausts the window (rejected). There is no path where i runs unbounded.

Validate every byte before you trust it

Real SIP methods are uppercase ASCII letters only. So we do not merely find the space — we assert that every byte before it is in the range A–Z. This single check rejects a huge class of garbage in one stroke: embedded NUL bytes, control characters, digits, lowercase, UTF-8 sequences, and the CR of an empty request line. Whitelisting (accept only what you expect) is far safer than blacklisting (try to enumerate the bad bytes).

Input first token Accepted? Why
INVITE yes all A-Z, space follows
BYE yes all A-Z, space follows
invite no lowercase byte
INV1TE no digit is not a letter
INVITE\r (no space) no CR is not a letter
(16+ bytes, no space) no scan window exhausted
empty ( sip:...) no zero-length method

Knowledge check: why prefer a whitelist (A-Z only) over a blacklist that rejects, say, spaces and control characters?

A blacklist is a promise you can never fully keep: you have to think of every dangerous byte, and the ones you forget become bugs. A whitelist inverts the burden — you state the small set you understand and trust, and everything else is rejected by default. New or unusual inputs (multibyte UTF-8, high-bit bytes, NULs) fail closed instead of slipping through.

Respect the caller's buffer

The function signature is int parse_sip_method(const char *msg, char *out, size_t cap). cap is the total size of out, including space for the terminator. If the method is i bytes long, we need i + 1 <= cap before we copy. Get this comparison right and you can never overflow out, regardless of what the network sends. After the bounded memcpy, we write the '\0' ourselves at out[i] — the source bytes contain no terminator, so producing one is our responsibility.

Knowledge check: the method is exactly 8 bytes (REGISTER) and the caller passes cap == 8. Should you copy?

No. You need 8 bytes for the letters plus 1 for the '\0', so 9 bytes total. With cap == 8 the result would not be NUL-terminated (or the terminator would overflow). Reject with -1; only cap >= 9 is safe here.

Syntax notes

int parse_sip_method(const char *msg, char *out, size_t cap);
  • msg — pointer to the untrusted message bytes. const because we never modify them. May legitimately be non-NUL-terminated network data, so we rely on our own bounded scan, not on strlen.
  • out — caller-owned destination buffer. We write the method and a trailing '\0' here. The caller owns and frees/reuses it; we never allocate.
  • cap — total capacity of out in bytes, terminator included. size_t because it is a size.
  • returns — number of method bytes written, excluding the terminator, on success; -1 on any failure (NULL args, cap == 0, no space in the scan window, a non-A-Z byte, an empty method, or a method that would not fit).
void *memcpy(void *dst, const void *src, size_t n);
  • Copies exactly n bytes. It does not stop at a NUL and does not add one — you must size n yourself and terminate afterward. Source and destination must not overlap (use memmove if they might).

Lesson

Why this matters

SIP is the protocol behind every VoIP call and WebRTC signalling exchange. (SIP stands for Session Initiation Protocol; it sets up and tears down voice and video sessions.)

SIP looks almost exactly like HTTP. Both have:

  • A request line
  • Headers
  • A blank line
  • An optional body

The bug surface is similar too. The first place to get input validation right is the request line.

Here we extract just the method. Later modules can pull out the URI and the Via chain.

What the wire looks like

INVITE sip:alice@example.com SIP/2.0\r\n
Via: SIP/2.0/UDP 10.0.0.1:5060;branch=z9hG4bK1234\r\n
...

The method is the first token, INVITE, ending at the first space.

Your job

Implement this function:

int parse_sip_method(const char *msg, char *out, size_t cap);

It should:

  • Copy the leading method (everything up to the first space) into out.
  • NUL-terminate the result.
  • Return the number of bytes written.

Return -1 on any failure. Failure cases:

  • msg or out is NULL, or cap == 0.
  • No space found within the first 16 bytes.
  • The method would overflow cap.
  • Any non-letter character appears before the first space.

Common mistakes

  • Allowing lowercase. Real SIP methods are uppercase ASCII.
  • Forgetting the NUL terminator on the bounded output.
  • Treating arbitrary bytes as a method. Reject anything outside A-Z.

What this is NOT

  • A full SIP parser. Headers, SDP, and dialog tracking are all out of scope.
  • A signalling stack. We are only classifying the method.

Code examples

#include <stdio.h>
#include <string.h>
#include <stddef.h>

/* Longest real SIP method is "SUBSCRIBE" (9). We scan at most 16 bytes
 * before giving up, so a line with no space can never run away. */
#define SIP_METHOD_SCAN_MAX 16

/* Copy the leading method token (up to the first space) of a SIP request
 * line into out. Returns bytes written (excluding the NUL), or -1 on error. */
int parse_sip_method(const char *msg, char *out, size_t cap)
{
    if (msg == NULL || out == NULL || cap == 0)
        return -1;

    size_t i = 0;
    for (i = 0; i < SIP_METHOD_SCAN_MAX; i++) {
        char c = msg[i];
        if (c == ' ')
            break;                 /* found the method delimiter */
        if (c < 'A' || c > 'Z')
            return -1;             /* NUL, digit, lowercase, CR: all rejected */
    }
    if (i == SIP_METHOD_SCAN_MAX)  /* no space in the scan window */
        return -1;
    if (i == 0)                    /* empty method (" sip:...") */
        return -1;
    if (i + 1 > cap)               /* i letters + NUL must fit */
        return -1;

    memcpy(out, msg, i);
    out[i] = '\0';
    return (int)i;
}

static void try_case(const char *label, const char *msg, size_t cap)
{
    char buf[32];
    int n = parse_sip_method(msg, buf, cap);
    if (n < 0)
        printf("%-14s cap=%2zu -> REJECTED\n", label, cap);
    else
        printf("%-14s cap=%2zu -> \"%s\" (%d bytes)\n", label, cap, buf, n);
}

int main(void)
{
    try_case("INVITE",     "INVITE sip:alice@example.com SIP/2.0\r\n", 16);
    try_case("REGISTER",   "REGISTER sip:example.com SIP/2.0\r\n",     16);
    try_case("OPTIONS",    "OPTIONS sip:bob@host SIP/2.0\r\n",         16);
    try_case("BYE",        "BYE sip:alice@host SIP/2.0\r\n",           16);
    try_case("lowercase",  "invite sip:alice SIP/2.0\r\n",             16); /* reject */
    try_case("digit",      "INV1TE sip:alice SIP/2.0\r\n",             16); /* reject */
    try_case("no-space",   "INVITEINVITEINVITEINVITE",                 16); /* reject */
    try_case("empty",      " sip:alice SIP/2.0\r\n",                   16); /* reject */
    try_case("tight-buf",  "REGISTER sip:x SIP/2.0\r\n",               8);  /* reject */
    try_case("fits-buf",   "REGISTER sip:x SIP/2.0\r\n",               9);  /* ok */
    return 0;
}

Line by line

  • #define SIP_METHOD_SCAN_MAX 16 — the hard ceiling on the scan. It turns "search for a space" into "search for a space within 16 bytes," which is what keeps the loop bounded even on hostile input.
  • if (msg == NULL || out == NULL || cap == 0) return -1; — guard the contract up front. A NULL pointer or a zero-size buffer means there is nothing we can safely do, so fail before touching memory.
  • for (i = 0; i < SIP_METHOD_SCAN_MAX; i++) — the bounded scan. i is size_t to match cap and index arithmetic. The loop bound is the safety net.
  • if (c == ' ') break; — the normal exit: we found the delimiter, so the method is bytes [0, i).
  • if (c < 'A' || c > 'Z') return -1; — the whitelist check on every byte. This one line rejects NULs, CR, digits, lowercase, and high-bit bytes.
  • if (i == SIP_METHOD_SCAN_MAX) return -1; — we ran the window to the end without a space. That is a malformed (or malicious) line; reject.
  • if (i == 0) return -1; — the very first byte was a space, i.e. an empty method. Reject.
  • if (i + 1 > cap) return -1; — the capacity check: i letters plus one terminator must fit. This is the line that makes the copy overflow-proof.
  • memcpy(out, msg, i); — copy exactly the method bytes, no more. We deliberately do not use strcpy/strncpy, because msg may not be NUL-terminated.
  • out[i] = '\0'; — we add the terminator ourselves, since the copied bytes did not include one.
  • return (int)i; — hand back the length so the caller knows how many bytes are valid.
  • try_case(...) and main — a hermetic harness: fixed in-memory strings, a range of good and bad inputs, and two cap values that straddle the 9-byte boundary for REGISTER to prove the capacity check fires exactly where it should.

Common mistakes

1. Unbounded scan for the space

// WRONG
size_t i = 0;
while (msg[i] != ' ') i++;   // walks off the end if there is no space

Why it breaks: a packet with no space (or no space in the first many KB) makes i index past the buffer — an out-of-bounds read, and on network input a remotely triggerable one.

// FIXED
for (i = 0; i < SIP_METHOD_SCAN_MAX && msg[i] != ' '; i++) { ... }
if (i == SIP_METHOD_SCAN_MAX) return -1;

2. Copying without checking capacity

// WRONG
memcpy(out, msg, i);   // i could be larger than cap
out[i] = '\0';

Why it breaks: if the method (or the run before the space) is longer than out, you overflow the caller's buffer — a classic stack smash.

// FIXED
if (i + 1 > cap) return -1;
memcpy(out, msg, i);
out[i] = '\0';

3. Forgetting the terminator (or off-by-one on it)

// WRONG
if (i > cap) return -1;   // leaves no room for the NUL
memcpy(out, msg, i);
// no out[i] = '\0';

Why it breaks: the caller does printf("%s", out) and reads past the method into garbage, or the terminator itself overflows when i == cap.

// FIXED
if (i + 1 > cap) return -1;
memcpy(out, msg, i);
out[i] = '\0';

4. Blacklisting instead of whitelisting

// WRONG
if (c == '\0' || c == '\r' || c == '\n') return -1;   // misses digits, lowercase, high bytes...

Why it breaks: you will never enumerate every bad byte; the ones you miss (a lowercase method, a UTF-8 lead byte, 0x7f) slip through as a "valid" method.

// FIXED
if (c < 'A' || c > 'Z') return -1;   // accept only what you understand

Debugging tips

  • Build with the sanitizers. cc -std=c11 -Wall -Wextra -fsanitize=address,undefined m.c -o m will trap an out-of-bounds read the instant an unbounded scan walks off a buffer, and print the exact offending line. This is the fastest way to catch the classic mistakes in this lesson.
  • Feed it degenerate inputs. Test the empty string, a string that is all letters with no space, a 1000-byte line, a method that is exactly cap-1, cap, and cap+1 bytes, and a buffer whose bytes are not NUL-terminated (allocate a non-terminated array so any accidental strlen reads past it — ASan will flag it).
  • Print the return value and the buffer separately. A common confusion is a positive return but a blank out; print n and out on their own lines so you can see whether the copy or the terminator is the problem.
  • gdb for the boundary. Set a breakpoint on the capacity check, then p i, p cap, and p i + 1 to confirm the comparison fires exactly at the boundary you expect. x/16bx msg dumps the first 16 raw bytes so you can see stray control characters that print as blanks.
  • valgrind ./m confirms there are no invalid reads/writes and no leaks; for this function it should report zero errors since we allocate nothing.

Memory safety

  • Never trust the input to be NUL-terminated. Network bytes are just bytes. Any function that stops at '\0' (strlen, strcpy, strchr) can read past the buffer if the sender omits the terminator. This is why we scan with an explicit index bound and copy with memcpy sized by that bound.
  • The scan bound and the capacity bound are two different guarantees. SIP_METHOD_SCAN_MAX protects reads from msg; i + 1 <= cap protects writes to out. You need both — bounding one does not bound the other.
  • i + 1 cannot overflow here because i < 16, but in code that adds two attacker-influenced sizes, a + b can wrap around size_t and pass a length check while still being huge. Prefer comparisons like i > cap - 1 only when you have already proven cap != 0 (we check cap == 0 first).
  • Always terminate what you copy. memcpy never adds a '\0'; forgetting out[i] = '\0' turns a correct copy into an information leak when the caller treats out as a C string.
  • Concurrency: the function is reentrant and thread-safe as written because it holds no shared state — every buffer is caller-supplied. It stays safe only if each thread passes its own out; sharing one out across threads reintroduces a data race that has nothing to do with this function.

Real-world uses

  • SIP proxies and PBXes (Asterisk, FreeSWITCH, Kamailio) parse the method on every packet to route requests and to drop obviously malformed ones early. Kamailio in particular front-ends the whole flood with cheap request-line checks before any expensive processing.
  • VoIP-aware firewalls and IDS/IPS classify traffic by method to apply policy (e.g. rate-limit REGISTER floods, which are a common brute-force and DoS vector).
  • WebRTC signalling servers speak SIP-over-WebSocket; the same first-token discipline applies.
  • Best practice mirrors this lesson: parse defensively at the very edge, whitelist the tokens you accept, bound every scan, size every copy against the destination, and reject early. Where you can, hand the raw bytes to a well-tested, fuzzed parser rather than rolling your own for production — but understanding the bounded-copy discipline here is exactly what lets you audit that parser.

Practice tasks

  1. Case-insensitive report. Without changing the accept rule (still A-Z only), extend main to also try "invite ..." and "Invite ..." and confirm both are rejected. Explain in a comment which byte triggered the rejection in each case.
  2. Return the delimiter position too. Add a variant int parse_sip_method2(const char *msg, char *out, size_t cap, size_t *space_idx) that, on success, also writes the index of the first space through space_idx. Guard against a NULL space_idx.
  3. Table-driven validity check. Replace the c < 'A' || c > 'Z' test with a 256-entry lookup table is_method_char[256] you initialise once. Verify the behaviour is identical, then discuss when a lookup table is worth it.
  4. Known-method whitelist. After extracting the token, compare it against the fixed set {INVITE,ACK,BYE,CANCEL,REGISTER,OPTIONS,SUBSCRIBE,NOTIFY,INFO,PRACK,UPDATE,REFER,MESSAGE} and return -1 for a well-formed but unknown method like HACKME. Keep the comparison bounded.
  5. Fuzz it. Write a loop that generates random byte buffers of random length (including zero), passes each to parse_sip_method with a small cap, and asserts the function never reads or writes out of bounds. Build with -fsanitize=address,undefined and let it run a few million iterations; report any crash with the seed that produced it.

Summary

  • A SIP request line is METHOD SP request-URI SP version CRLF; the method is the first token, ending at the first space.
  • Scan with a hard upper bound (16 bytes) so a message with no space can never walk your loop off the buffer.
  • Whitelist the bytes: accept only A-Z, and reject the whole token if any other byte appears before the space (this catches NULs, CR, digits, lowercase, and high-bit bytes at once).
  • Check method_len + 1 <= cap before copying, copy exactly method_len bytes with memcpy, then write the '\0' yourself — memcpy never adds one.
  • Return the number of bytes written, or -1 on any failure (NULL args, cap == 0, no space, empty method, bad byte, or would-not-fit).
  • Never rely on network input being NUL-terminated; bound your reads and your writes separately.

Practice with these exercises