Networking in C · beginner · ~8 min
- By the end you can create a TCP or UDP socket in C and explain what `SOCK_STREAM` vs `SOCK_DGRAM` actually change. - By the end you can explain reliability, ordering, and flow control, and say exactly which of these the kernel gives you for free with TCP and which you must build yourself on UDP. - By the end you can predict how `recv()` behaves on a byte stream vs how `recvfrom()` behaves on datagrams, and why TCP has no message boundaries. - By the end you can choose the right transport for a given job (bulk transfer, request/reply, real-time media, service discovery) and justify the trade-off. - By the end you can name the security-relevant differences (spoofing, amplification, connection state) and the defensive habits that follow from them.
You already opened a socket in the sockets intro lesson and saw socket(AF_INET, SOCK_STREAM, 0) return a file descriptor. That third-from-the-parameters detail — the socket type — is the whole subject of this lesson. Changing SOCK_STREAM to SOCK_DGRAM swaps the transport protocol underneath your file descriptor from TCP to UDP, and that one word changes almost everything about how your data behaves on the wire.
This lesson builds directly on that intro: same socket() call, same file-descriptor mental model, same AF_INET address family. What's new is what the kernel promises to do for you once bytes leave your process. TCP promises a lot (a reliable, ordered stream) and UDP promises almost nothing (a thin envelope over IP). Understanding that difference is what lets you pick correctly instead of cargo-culting SOCK_STREAM into every program.
Picking the wrong transport is a bug you feel much later, in production, under load. Reach for UDP where you needed ordering and you will ship a protocol that silently drops and reorders records; reach for TCP where you needed low latency and every lost packet stalls the whole stream (head-of-line blocking) — this is exactly why gaming, voice, and QUIC/HTTP-3 avoid raw TCP. The choice is also a security boundary: UDP is connectionless and trivially source-spoofed, which is why it is the workhorse of reflection/amplification DDoS (DNS, NTP, memcached), while TCP's three-way handshake gives you a weak proof the peer can receive at its claimed address — and its own abuse, the SYN flood. Knowing which guarantees you actually have tells you which attacks you must defend against and which validation you cannot skip.
The transport is chosen by the socket type argument to socket():
int tcp = socket(AF_INET, SOCK_STREAM, 0); /* -> TCP */
int udp = socket(AF_INET, SOCK_DGRAM, 0); /* -> UDP */
AF_INET says "IPv4 addresses." The type says how those bytes travel. SOCK_STREAM selects TCP: a connection-oriented, reliable byte stream. SOCK_DGRAM selects UDP: connectionless, best-effort datagrams. The 0 means "default protocol for this type," which is why you rarely name IPPROTO_TCP/IPPROTO_UDP explicitly.
| Property | TCP (SOCK_STREAM) |
UDP (SOCK_DGRAM) |
|---|---|---|
| Connection | Yes — 3-way handshake before data | No — just send |
| Reliability | Kernel retransmits lost data | None — lost is lost |
| Ordering | Bytes delivered in send order | May arrive reordered |
| Duplicates | Removed by kernel | Possible |
| Flow control | Yes (receiver window) | No |
| Congestion control | Yes (backs off on loss) | No (you can flood the link) |
| Message boundaries | None — a byte stream | Preserved — one send = one recv |
| Header overhead | 20+ bytes/segment | 8 bytes/datagram |
| Typical calls | connect/accept, send/recv |
sendto/recvfrom |
| Powers | HTTP/1–2, SSH, SMTP, Git, TLS | DNS, DHCP, VoIP, games, QUIC/HTTP-3 |
The mental shortcut: TCP is a phone call (you dial, a circuit is established, words arrive in order, and if the line garbles you both repeat), while UDP is postcards (you drop each one in the box, most arrive, some don't, and they can arrive out of order — but each postcard is whole).
This is the single most common beginner trap, so internalize it now. TCP does not preserve message boundaries. send() is not "send one message"; it is "append these bytes to the stream." The receiver's recv() returns whatever bytes are available right now — which may be part of one message, exactly one message, or several messages stuck together ("coalesced").
Sender: send("AB") send("CD") send("EF") three calls
Wire (TCP stream): A B C D E F boundaries erased
Receiver: recv() -> "ABCDEF" one call, all six bytes
(or "AB" then "CDEF", or "ABC" then "DEF" — any split!)
UDP is the opposite. Each sendto() becomes exactly one datagram, and each recvfrom() returns exactly one datagram — never a partial one, never two joined:
Sender: sendto("AB") sendto("CD") sendto("EF")
Wire (UDP): [AB] [CD] [EF] three envelopes
Receiver: recvfrom() -> "AB"
recvfrom() -> "CD" boundaries preserved
recvfrom() -> "EF" (but any of them may be lost)
Because TCP has no boundaries, you must add them back when you send structured messages. The two standard techniques are a length prefix (write a 4-byte length, then that many payload bytes; the reader loops until it has the full length) or a delimiter (e.g. \n for line-based protocols). This is called framing.
Knowledge check: You call
send()twice over one TCP connection, 100 bytes then 50 bytes. On the other side you callrecv()once with a 4096-byte buffer. How many bytes might it return?Anywhere from 1 to 150. TCP is a byte stream with no boundaries: the two sends may have coalesced in the socket buffer (up to 150 bytes in one recv), or only part may have arrived so far (as few as 1 byte). You must never assume one recv equals one logical message — loop and reassemble using your own framing.
UDP preserving boundaries does not mean UDP is easy. A datagram may vanish, may be duplicated, or may arrive after a later one. And there is a sharp edge: if your recvfrom() buffer is smaller than the datagram, UDP delivers the bytes that fit and discards the rest — the extra bytes are gone, not saved for the next call (the MSG_TRUNC flag reports this on Linux). With TCP, an undersized buffer is harmless: leftover bytes simply wait in the stream for your next recv(). So on UDP, always size the receive buffer to your protocol's maximum datagram.
Because UDP is connectionless, the receiver has no handshake proving the sender is really at the source IP in the packet. That makes UDP the classic vehicle for reflection/amplification attacks: an attacker sends a small query with a forged source address (the victim's), and the server dutifully sends a much larger reply to the victim. DNS, NTP, and memcached have all been abused this way. Defensive practice for anyone writing UDP services on a lab/localhost box: keep responses no larger than requests where possible, require a lightweight challenge/cookie before sending big replies, rate-limit per source, and never trust the source IP for authorization. TCP resists source spoofing because the handshake's server sequence number must be echoed back — but it has its own resource-exhaustion attack, the SYN flood (half-open connections filling the accept queue), defended with SYN cookies and connection limits. The general rule: with UDP you must build in your own validation and reliability; with TCP the kernel gives you more, but you still validate every byte of payload before you trust it.
#include <sys/socket.h>
#include <netinet/in.h>
#include <arpa/inet.h>
int socket(int domain, int type, int protocol);
/* domain: AF_INET (IPv4) or AF_INET6.
* type: SOCK_STREAM -> TCP, SOCK_DGRAM -> UDP.
* protocol: 0 = default for the type (TCP or UDP respectively).
* returns: a file descriptor (>= 0), or -1 with errno set. close() it. */
/* --- TCP path --- */
int connect(int fd, const struct sockaddr *addr, socklen_t len);
/* client: establishes the connection (does the 3-way handshake). 0 or -1/errno. */
int listen(int fd, int backlog); /* server: mark socket passive. */
int accept(int fd, struct sockaddr *addr, socklen_t *len);
/* server: returns a NEW fd for one accepted connection, or -1/errno. */
ssize_t send(int fd, const void *buf, size_t n, int flags);
ssize_t recv(int fd, void *buf, size_t n, int flags);
/* stream I/O. Return value is the count actually transferred and MAY be less
* than n (a "short" transfer). recv() returning 0 means the peer closed.
* -1 with errno on error. You must loop to send/receive a full message. */
/* --- UDP path --- */
ssize_t sendto(int fd, const void *buf, size_t n, int flags,
const struct sockaddr *dst, socklen_t dstlen);
ssize_t recvfrom(int fd, void *buf, size_t n, int flags,
struct sockaddr *src, socklen_t *srclen);
/* datagram I/O. One sendto == one datagram; one recvfrom == one datagram.
* srclen is in/out: set it to sizeof(your addr) before the call. Pass NULL
* for src if you don't need the sender's address. If the datagram is larger
* than n, the excess is discarded (data loss). -1/errno on error. */
/* Helpers used to fill sockaddr_in: */
htons(port); /* host -> network byte order, 16-bit */
htonl(addr); /* host -> network byte order, 32-bit; INADDR_LOOPBACK = 127.0.0.1 */
Every socket fd must be close()d. accept() returns a separate fd you must also close. Always check for -1 and read errno (via perror); network calls fail routinely.
When you write networked C, you choose one of two transport protocols. Each handles delivery very differently.
TCP gives you a reliable, ordered stream of bytes:
The kernel does the hard work for you. It handles retransmits (resending lost data), ordering, and flow control (matching the sender's speed to what the receiver can handle).
TCP powers almost everything you use directly: HTTP, SSH, SMTP, and Git.
UDP is simpler and does far less for you. Each sendto() call becomes exactly one packet on the wire. That packet:
If you need reliability, you must build it yourself on top of UDP.
UDP is used by DNS, video conferencing, online gaming, and QUIC.
A message boundary marks where one message ends and the next begins. TCP and UDP treat these very differently.
send() calls of 4 bytes each might arrive as a single recv() of 8 bytes.sendto() arrives as one packet at the receiver. But delivery is not guaranteed.For everything in this course, use TCP unless an exercise specifically calls for UDP.
#include <stdio.h>
#include <string.h>
#include <stdlib.h>
#include <unistd.h>
#include <arpa/inet.h>
#include <sys/socket.h>
#include <netinet/in.h>
/* Demonstrates the core TCP-vs-UDP difference on loopback (no root needed):
* TCP is a byte STREAM (message boundaries are erased),
* UDP is DATAGRAM based (one sendto == one recvfrom). */
static void die(const char *what) { perror(what); exit(1); }
/* Send three small messages, then show how the receiver sees them. */
static const char *MSGS[3] = { "AB", "CD", "EF" };
static void tcp_demo(void) {
/* 1. Listening socket on an OS-chosen (ephemeral) port. */
int lst = socket(AF_INET, SOCK_STREAM, 0);
if (lst < 0) die("socket");
struct sockaddr_in addr = {0};
addr.sin_family = AF_INET;
addr.sin_addr.s_addr = htonl(INADDR_LOOPBACK); /* 127.0.0.1 */
addr.sin_port = 0; /* let kernel pick */
if (bind(lst, (struct sockaddr *)&addr, sizeof addr) < 0) die("bind");
if (listen(lst, 1) < 0) die("listen");
socklen_t alen = sizeof addr;
if (getsockname(lst, (struct sockaddr *)&addr, &alen) < 0) die("getsockname");
/* 2. Client connects to that port. */
int cli = socket(AF_INET, SOCK_STREAM, 0);
if (cli < 0) die("socket");
if (connect(cli, (struct sockaddr *)&addr, sizeof addr) < 0) die("connect");
int srv = accept(lst, NULL, NULL);
if (srv < 0) die("accept");
/* 3. Three separate send() calls. */
for (int i = 0; i < 3; i++)
if (send(cli, MSGS[i], strlen(MSGS[i]), 0) < 0) die("send");
/* 4. ONE recv() often returns all of them joined: boundaries are gone. */
char buf[64];
ssize_t n = recv(srv, buf, sizeof buf - 1, 0);
if (n < 0) die("recv");
buf[n] = '\0';
printf("TCP : 3 sends of 2 bytes -> 1 recv got %zd bytes: \"%s\"\n", n, buf);
close(cli); close(srv); close(lst);
}
static void udp_demo(void) {
/* Receiver socket bound to an ephemeral loopback port. */
int rx = socket(AF_INET, SOCK_DGRAM, 0);
if (rx < 0) die("socket");
struct sockaddr_in addr = {0};
addr.sin_family = AF_INET;
addr.sin_addr.s_addr = htonl(INADDR_LOOPBACK);
addr.sin_port = 0;
if (bind(rx, (struct sockaddr *)&addr, sizeof addr) < 0) die("bind");
socklen_t alen = sizeof addr;
if (getsockname(rx, (struct sockaddr *)&addr, &alen) < 0) die("getsockname");
int tx = socket(AF_INET, SOCK_DGRAM, 0);
if (tx < 0) die("socket");
/* Three datagrams. */
for (int i = 0; i < 3; i++)
if (sendto(tx, MSGS[i], strlen(MSGS[i]), 0,
(struct sockaddr *)&addr, sizeof addr) < 0) die("sendto");
/* Each recvfrom() returns EXACTLY one datagram: boundaries preserved. */
for (int i = 0; i < 3; i++) {
char buf[64];
ssize_t n = recvfrom(rx, buf, sizeof buf - 1, 0, NULL, NULL);
if (n < 0) die("recvfrom");
buf[n] = '\0';
printf("UDP : recvfrom #%d got %zd bytes: \"%s\"\n", i + 1, n, buf);
}
close(tx); close(rx);
}
int main(void) {
tcp_demo();
udp_demo();
return 0;
}
Expected output:
TCP : 3 sends of 2 bytes -> 1 recv got 6 bytes: "ABCDEF"
UDP : recvfrom #1 got 2 bytes: "AB"
UDP : recvfrom #2 got 2 bytes: "CD"
UDP : recvfrom #3 got 2 bytes: "EF"
Compile and run (no root, no external network — everything is on 127.0.0.1):
cc -std=c11 -Wall -Wextra m.c -o m && ./m
die() — a tiny helper: perror() prints the failing call plus a human-readable errno string, then we exit(1). Every socket call can fail, so we check all of them.MSGS[3] — the same three 2-byte payloads are sent over both transports so you can compare how the receiver perceives them.tcp_demo step 1 — socket(AF_INET, SOCK_STREAM, 0) gives us a TCP socket. We fill a sockaddr_in with loopback (INADDR_LOOPBACK, wrapped in htonl for network byte order) and port 0, which tells the kernel "pick any free port." bind() claims it; listen() turns the socket passive so it can accept connections.getsockname — because we asked for port 0, we don't yet know which port we got. getsockname() writes the actual chosen address (including port) back into addr so the client can connect to it. This is what makes the demo hermetic — no hard-coded port to collide with.cli calls connect() to that address (this performs TCP's three-way handshake, all within the kernel on loopback). The server side accept()s and gets srv, a brand-new fd representing that one connection. Note there are now three fds: the listener, the client end, and the server end.send() calls, 2 bytes each. On the wire these become part of one continuous byte stream.recv() with a large buffer. Because TCP has no message boundaries and all six bytes are already buffered, one recv() returns "ABCDEF". We NUL-terminate at buf[n] (that's why the buffer is sizeof buf - 1) so %s is safe. This line is the whole lesson in one printout.udp_demo — mirrors the setup but with SOCK_DGRAM. rx is bound to an ephemeral loopback port; getsockname retrieves it; tx is an unbound sender.sendto() calls each produce one datagram addressed to rx's port.recvfrom() calls each return exactly one datagram: "AB", then "CD", then "EF". Boundaries are preserved — the opposite of TCP. (On a real network any of these could be lost or reordered; on loopback they aren't, which keeps the demo deterministic.)close() calls — every fd, including the accepted srv, is closed. Leaking socket fds is a real resource leak in long-running servers.1. Treating one recv() as one message.
char buf[128];
recv(fd, buf, sizeof buf, 0); /* WRONG: assumes a whole message arrived */
process(buf);
Why it breaks: TCP is a byte stream. recv() may return a partial message, or several messages coalesced. process() then parses garbage or half a record.
/* FIX: frame it. Send a length prefix, then loop until you have that many bytes. */
uint32_t need; read_exact(fd, &need, 4); need = ntohl(need);
char *msg = malloc(need);
read_exact(fd, msg, need); /* read_exact loops over recv() until 'need' bytes */
2. Ignoring short writes/reads.
send(fd, buf, len, 0); /* WRONG: return value ignored */
Why it breaks: send/recv may transfer fewer bytes than requested and return the count. Ignoring it silently truncates data.
/* FIX: loop until everything is sent. */
size_t off = 0;
while (off < len) {
ssize_t w = send(fd, buf + off, len - off, 0);
if (w < 0) { perror("send"); break; }
off += (size_t)w;
}
3. Using an undersized UDP receive buffer.
char buf[8];
recvfrom(fd, buf, sizeof buf, 0, NULL, NULL); /* WRONG for a 40-byte datagram */
Why it breaks: unlike TCP, UDP discards the bytes that don't fit — they are gone forever, not saved for the next call.
/* FIX: size the buffer to your protocol's max datagram (e.g. 65507 max for UDP/IPv4). */
char buf[2048]; /* >= your largest expected message */
ssize_t n = recvfrom(fd, buf, sizeof buf, 0, NULL, NULL);
4. Picking the wrong transport for the workload.
/* WRONG: raw UDP for a file transfer that must arrive intact */
sendto(fd, chunk, len, 0, dst, dstlen); /* lost/reordered chunks corrupt the file */
Why it breaks: UDP won't retransmit or reorder; you'd silently corrupt data. (And the mirror mistake — TCP for real-time voice — adds latency via retransmits you don't want.)
/* FIX: use TCP when you need reliability+order; reserve UDP for loss-tolerant,
latency-sensitive data or when you implement your own reliability layer. */
int fd = socket(AF_INET, SOCK_STREAM, 0);
ss -tunap (t=TCP, u=UDP, n=numeric, a=all, p=process) or netstat -anp shows whether your program opened a TCP or UDP socket and which port. A UDP socket where you expected TCP is an instant tell that your SOCK_* constant is wrong.tcpdump -i lo -X port <p> (or Wireshark) shows the real packets on loopback. For TCP you'll see the SYN/SYN-ACK/ACK handshake and data segments; for UDP just the datagrams. If you "lose" a UDP message, tcpdump tells you whether it left the sender at all.strace -e trace=network ./prog (Linux) or dtruss/ktrace (macOS) prints every socket, bind, connect, send, recv with arguments and return values. This is the fastest way to catch an ignored short read or a recv() returning 0 (peer closed).n = recv(...) every time. n == 0 means the peer closed the connection; n < 0 with errno == EWOULDBLOCK/EAGAIN on a non-blocking socket just means "nothing yet"; a small n proves you must loop.send() calls, or send one byte at a time, to force TCP to split messages across recv()s. If your parser only works when everything arrives in one chunk, this exposes it immediately.getsockname/getpeername print the local/remote address the kernel actually assigned — useful when you bound to port 0 and don't know the port, or to confirm you're on 127.0.0.1 and not a real interface.recv/recvfrom do not add a \0. Read into buf[sizeof buf - 1] at most and set buf[n] = '\0' (as the demo does) before any printf("%s"), strlen, or strcpy — otherwise you read past the buffer (UB) or leak adjacent memory.len <= MAX_MSG before malloc(len) or before copying, or you get a giant allocation, an integer overflow in len + 1, or a heap overflow.MSG_TRUNC (Linux) if you need to detect over-length datagrams.recvfrom's source address after a truncated/short read as if it were fully populated. Initialize srclen = sizeof(addr) before every call; it is an in/out parameter and stale values cause the kernel to write the wrong amount.close() can close a different fd that was recycled to the same number (a real, exploitable class of bug in concurrent servers). Set the variable to -1 after closing if the code path might reach it again.send()s, corrupting your framing. Give each connection its own thread/fd or serialize writes with a mutex.recv boundaries, and handle partial reads/writes with loops.(Warm-up) Modify the demo so both tcp_demo and udp_demo send five messages of your choice. Print the total byte count each side received and confirm the TCP total equals the sum of the sends while each UDP recvfrom returns one message.
(Framing) Add a length-prefix framing layer to the TCP path: before each payload, send() a 4-byte big-endian length (use htonl), then the payload. On the receiver, write a read_exact() helper that loops over recv() until it has the full count, and use it to read the length then the body. Verify you can recover the three original messages separately from the byte stream.
(Short-read robustness) Insert a deliberate 100 ms delay (e.g. usleep) between the three TCP send() calls, and change the receiver to loop calling recv() and print each chunk. Observe how the stream now splits differently across calls, and confirm your task-2 framing still reassembles the messages correctly.
(UDP reliability) Simulate loss on the UDP path by skipping one sendto() (don't send message #2), and give each datagram a 1-byte sequence number prefix. On the receiver, detect the gap in sequence numbers and print "missing message N" — this is the first step toward building reliability on UDP.
(Defensive challenge) Write a tiny UDP echo responder bound to loopback that only echoes a datagram back if the payload begins with a fixed 4-byte cookie the client must include; drop and count anything else, and cap the reply size to the request size. Explain in a comment how these two rules (cookie + no amplification) blunt reflection/amplification abuse.
SOCK_STREAM = TCP, SOCK_DGRAM = UDP.sendto = one recvfrom, so boundaries are preserved, but delivery, ordering, and dedup are your problem, and an undersized buffer silently truncates.