Networking in C · advanced · ~10 min

Socket timeouts

- By the end you can explain why a blocking `recv`/`read` on a TCP socket can hang forever, and name what causes it. - By the end you can put a hard deadline on a receive using `setsockopt` with `SO_RCVTIMEO` and correctly detect the timeout via `errno`. - By the end you can bound a blocking operation with `select`/`poll` and understand when to prefer it over `SO_RCVTIMEO`. - By the end you can distinguish a timeout (`EAGAIN`/`EWOULDBLOCK`) from a real error and from a clean peer close (`recv` returns 0). - By the end you can reason about connect timeouts and send timeouts, not just receive timeouts.

Overview

In the TCP client lesson you learned to socket(), connect(), and then recv()/send() bytes over a connection. Every one of those blocking calls quietly assumes the peer will cooperate. This lesson removes that assumption: it shows how to put a deadline on the reading side so a silent or slow peer cannot freeze your program.

You already know how to open a connection and read from it. Here we keep the same client shape but add one thing: a bound on how long any single blocking call is allowed to wait. We will use a fully local (127.0.0.1) TCP pair so you can watch a real timeout fire without any network or root access.

Why it matters

A blocking recv with no timeout is one of the most common ways real services wedge themselves: a peer crashes without sending a FIN, a firewall silently drops the connection, or a malicious client opens a socket and then just sits there. Without a deadline, that one stuck connection ties up a thread or file descriptor forever, and enough of them is a denial-of-service — this is exactly how slowloris-style attacks starve a server. Setting timeouts is a defensive baseline: it turns an unbounded hang into a bounded, recoverable error you can log, retry, or close.

Core concepts

Why a blocking socket can wait forever

By default a TCP socket is in blocking mode. When you call recv() and no data has arrived, the kernel parks your thread until either bytes show up or the connection ends. If the peer is alive but simply never sends, and never closes, there is nothing to wake you. TCP itself will not time out a merely idle connection quickly — keepalives are off by default and measured in hours. So "the other side went quiet" and "the other side is gone" look identical to a plain recv: both just block.

  your thread                     peer
  -----------                     ----
  recv(fd, ...)  ------------->    (silent: never sends, never closes)
     |
     | blocked... 1s
     | blocked... 10s
     | blocked... forever   <-- no timeout = permanent hang
     v

The cure is to attach a deadline to the wait. There are three standard ways to do that, summarized here:

Technique What it bounds Granularity Best when
SO_RCVTIMEO / SO_SNDTIMEO via setsockopt one recv / one send on that socket per-socket, persists across calls you just want each blocking call to give up after N ms
select / poll + non-blocking or blocking fd waiting for readiness across one or many fds per-call timeout you watch many sockets at once, or want a single deadline across several ops
alarm() + SIGALRM interrupts the syscall with EINTR seconds only, whole-process legacy code; rarely the right modern choice

Technique 1 — SO_RCVTIMEO: a deadline baked into the socket

setsockopt(fd, SOL_SOCKET, SO_RCVTIMEO, &tv, sizeof tv) takes a struct timeval and tells the kernel: any blocking recv/read on this socket may wait at most that long. If the time expires with no data, the call returns -1 and sets errno to EAGAIN (equal to EWOULDBLOCK on Linux/macOS). Crucially this is not an error you should treat like a failed connection — it is the expected "deadline reached" signal. SO_SNDTIMEO is the mirror image for send, useful when the peer's receive window is full and your send would otherwise block.

 recv() outcomes
 +--------------------------+-------------------------------+
 | return value             | meaning                       |
 +--------------------------+-------------------------------+
 | > 0                      | that many bytes received      |
 | 0                        | peer closed cleanly (FIN)     |
 | -1, errno=EAGAIN/EWOULD  | timeout fired (SO_RCVTIMEO)   |
 | -1, errno=EINTR          | interrupted by a signal       |
 | -1, other errno          | real error (ECONNRESET, ...)  |
 +--------------------------+-------------------------------+

Knowledge check: after you set SO_RCVTIMEO to 5 seconds and recv returns -1 with errno == EWOULDBLOCK, is the connection broken?

No. The timeout simply means no data arrived within 5 seconds. The TCP connection is still open and valid; you may call recv again, or decide the peer is too slow and close it yourself. EWOULDBLOCK/EAGAIN here means "deadline reached," not "connection failed." A real failure shows up as a different errno (like ECONNRESET) or as a return of 0 for a clean close.

Technique 2 — select/poll: wait for readiness, then read

Instead of blocking inside recv, you can ask the kernel "tell me when this fd has data, but wait no longer than T." select(nfds, &readset, ..., &timeout) returns 0 on timeout, >0 with the fd flagged when it is readable, or -1 on error. You then do a recv you already know will not block. The big win is that one select/poll can watch many sockets and apply one deadline across all of them — which is why servers handling thousands of connections are built on poll/epoll, not on one SO_RCVTIMEO per socket. A subtlety: select's timeout struct may be modified on return on Linux, so re-initialize it before each call.

Technique 3 — alarm() + SIGALRM (know it, rarely reach for it)

alarm(3) schedules SIGALRM in 3 seconds; the signal interrupts the blocked syscall, which returns -1 with errno == EINTR. It only offers whole-second granularity, is process-global (one pending alarm at a time), and is fiddly to make race-free. It predates the other two and is mostly of historical interest; prefer SO_RCVTIMEO or poll.

Don't forget connect timeouts

SO_RCVTIMEO/SO_SNDTIMEO do not bound connect(). A connect to an unreachable host can hang for the kernel's SYN-retry period (often ~75s+). To bound it, put the socket in non-blocking mode (O_NONBLOCK), call connect (it returns -1 with errno == EINPROGRESS), then select/poll for writability with your own timeout, and check SO_ERROR via getsockopt to see whether the connection actually succeeded.

Syntax notes

#include <sys/socket.h>
#include <sys/select.h>

// Set a receive deadline on a socket.
// level = SOL_SOCKET; optname = SO_RCVTIMEO (or SO_SNDTIMEO for sends).
// optval points to a struct timeval {tv_sec, tv_usec}; a zero timeval = block forever.
// Returns 0 on success, -1 on error (errno set). Setting persists for the socket's life.
int setsockopt(int fd, int level, int optname, const void *optval, socklen_t optlen);

// struct timeval: tv_sec = whole seconds, tv_usec = microseconds (0..999999).
struct timeval { time_t tv_sec; suseconds_t tv_usec; };

// recv: like read but socket-specific. flags = 0 for normal use.
// Returns >0 bytes read, 0 on clean peer close (FIN), -1 on error/timeout (check errno).
// On timeout via SO_RCVTIMEO: errno == EAGAIN (== EWOULDBLOCK on Linux/macOS).
ssize_t recv(int fd, void *buf, size_t len, int flags);

// select: block until an fd is ready or timeout elapses.
// nfds = highest fd number + 1. readfds/writefds/exceptfds are fd_set bitmaps (may be NULL).
// timeout: NULL = wait forever; {0,0} = poll and return immediately; else a deadline.
// Returns count of ready fds, 0 on timeout, -1 on error. NOTE: on Linux it may modify *timeout.
int select(int nfds, fd_set *readfds, fd_set *writefds,
           fd_set *exceptfds, struct timeval *timeout);

void FD_ZERO(fd_set *set);         // clear the set
void FD_SET(int fd, fd_set *set);  // add fd
int  FD_ISSET(int fd, fd_set *set); // is fd flagged ready?

Must-do notes: nothing here allocates heap, so there is nothing to free. You still must close() every socket fd. Always check setsockopt's return value — a failed option-set can silently leave you with an unbounded socket.

Lesson

The problem

By default, read or recv on a blocking socket (one that pauses your program until data arrives) can wait forever. If the other side never sends anything, your program hangs.

The fix is to set a deadline. There are three common ways to do this.

Three ways to bound a blocking operation

  • setsockopt with SO_RCVTIMEO — Sets a read timeout directly on the socket. There is also SO_SNDTIMEO for send timeouts. This is the simplest option.
  • Non-blocking socket plus select or poll — Make the socket non-blocking, then wait for it to become ready using select or poll with a timeout. This gives you the most control.
  • SIGALRM with alarm(seconds) — Schedule a signal that interrupts the blocked call after a set number of seconds. This is the older approach and is less flexible.

Code examples

#include <stdio.h>
#include <string.h>
#include <errno.h>
#include <time.h>
#include <unistd.h>
#include <sys/socket.h>
#include <sys/select.h>
#include <netinet/in.h>
#include <arpa/inet.h>

/* Milliseconds between two CLOCK_MONOTONIC samples. */
static double ms_since(struct timespec a, struct timespec b) {
    return (b.tv_sec - a.tv_sec) * 1e3 + (b.tv_nsec - a.tv_nsec) / 1e6;
}

int main(void) {
    /* ---- Build a hermetic loopback TCP pair: everything on 127.0.0.1 ---- */
    int lst = socket(AF_INET, SOCK_STREAM, 0);
    if (lst < 0) { perror("socket"); return 1; }

    struct sockaddr_in addr = {0};
    addr.sin_family = AF_INET;
    addr.sin_addr.s_addr = htonl(INADDR_LOOPBACK); /* 127.0.0.1 */
    addr.sin_port = 0;                             /* kernel picks a free port */

    if (bind(lst, (struct sockaddr *)&addr, sizeof addr) < 0) { perror("bind"); return 1; }
    if (listen(lst, 1) < 0) { perror("listen"); return 1; }

    /* Discover the ephemeral port the kernel chose. */
    socklen_t alen = sizeof addr;
    if (getsockname(lst, (struct sockaddr *)&addr, &alen) < 0) { perror("getsockname"); return 1; }
    printf("listening on 127.0.0.1:%d\n", ntohs(addr.sin_port));

    /* Client connects. On loopback the connection lands in the accept queue,
       so a blocking connect() completes before we ever call accept(). */
    int cli = socket(AF_INET, SOCK_STREAM, 0);
    if (connect(cli, (struct sockaddr *)&addr, sizeof addr) < 0) { perror("connect"); return 1; }
    int srv = accept(lst, NULL, NULL);
    if (srv < 0) { perror("accept"); return 1; }

    /* The server deliberately sends nothing yet, so the client will stall. */

    /* ---- Technique 1: SO_RCVTIMEO on the client socket ---- */
    struct timeval tv = { .tv_sec = 0, .tv_usec = 300000 }; /* 300 ms */
    if (setsockopt(cli, SOL_SOCKET, SO_RCVTIMEO, &tv, sizeof tv) < 0) {
        perror("setsockopt"); return 1;
    }

    char buf[64];
    struct timespec t0, t1;
    clock_gettime(CLOCK_MONOTONIC, &t0);
    ssize_t n = recv(cli, buf, sizeof buf, 0);
    clock_gettime(CLOCK_MONOTONIC, &t1);

    if (n < 0 && (errno == EAGAIN || errno == EWOULDBLOCK)) {
        printf("SO_RCVTIMEO: recv timed out after ~%.0f ms (errno=%s)\n",
               ms_since(t0, t1), strerror(errno));
    } else {
        printf("SO_RCVTIMEO: unexpected recv result n=%zd\n", n);
    }

    /* ---- Technique 2: select() with a timeout (socket left blocking) ---- */
    fd_set rfds;
    FD_ZERO(&rfds);
    FD_SET(cli, &rfds);
    struct timeval sel_tv = { .tv_sec = 0, .tv_usec = 300000 };
    clock_gettime(CLOCK_MONOTONIC, &t0);
    int r = select(cli + 1, &rfds, NULL, NULL, &sel_tv);
    clock_gettime(CLOCK_MONOTONIC, &t1);
    if (r == 0) {
        printf("select: readiness wait timed out after ~%.0f ms\n", ms_since(t0, t1));
    } else if (r > 0 && FD_ISSET(cli, &rfds)) {
        printf("select: socket became readable\n");
    } else {
        perror("select");
    }

    /* ---- Now the server sends: the next recv returns promptly ---- */
    const char *msg = "hi";
    send(srv, msg, strlen(msg), 0);

    clock_gettime(CLOCK_MONOTONIC, &t0);
    n = recv(cli, buf, sizeof buf, 0);
    clock_gettime(CLOCK_MONOTONIC, &t1);
    if (n > 0) {
        buf[n] = '\0';
        printf("after send: recv got \"%s\" in ~%.0f ms\n", buf, ms_since(t0, t1));
    }

    close(cli);
    close(srv);
    close(lst);
    return 0;
}

Line by line

  • ms_since(...): a tiny helper that turns two CLOCK_MONOTONIC timestamps into elapsed milliseconds so we can observe that the timeout really fired at ~300 ms. CLOCK_MONOTONIC is used (not wall-clock) because it never jumps backward.
  • socket/bind/listen on INADDR_LOOPBACK with sin_port = 0: builds a listener bound to 127.0.0.1 on a kernel-chosen ephemeral port. Using port 0 avoids clashing with anything already running — the demo is fully self-contained and needs no privileges.
  • getsockname: after bind, this reads back which port the kernel actually assigned, purely so we can print it.
  • connect then accept: because both ends are on loopback with a listen backlog, the client's blocking connect completes as soon as the SYN is queued, before we call accept. That lets a single thread own both ends of the connection — no second process or thread needed.
  • "server sends nothing yet": this is the whole point — it simulates a silent peer, the situation that would otherwise hang us.
  • setsockopt(cli, SOL_SOCKET, SO_RCVTIMEO, &tv, ...): arms a 300 ms receive deadline on the client socket, and we check its return value.
  • first recv: with the peer silent, this blocks until the deadline and returns -1. We test errno == EAGAIN || EWOULDBLOCK to confirm it was a timeout, not a broken connection, and print the measured elapsed time (~300 ms).
  • FD_ZERO/FD_SET/select: the second technique. We ask select to wait up to 300 ms for the client fd to become readable. Since nothing arrives, select returns 0 (timeout) and we print the elapsed time. Note sel_tv is freshly initialized right before the call because select may modify it.
  • send(srv, "hi", ...): now the server finally sends two bytes over the same connection.
  • final recv: this time data is waiting, so recv returns promptly with n > 0; we NUL-terminate and print the payload. The near-zero elapsed time contrasts with the two timeouts above.
  • close(...) on all three fds: sockets are file descriptors and must be closed to release kernel resources.

Common mistakes

1. Treating a timeout as a fatal error and killing the connection unnecessarily.

if (recv(fd, buf, len, 0) < 0) { close(fd); return -1; } // WRONG

Why it breaks: recv returns -1 both for genuine errors and for an SO_RCVTIMEO timeout (errno == EAGAIN). This code tears down a perfectly healthy but momentarily-idle connection.

ssize_t n = recv(fd, buf, len, 0);
if (n < 0 && (errno == EAGAIN || errno == EWOULDBLOCK)) { /* just a timeout: retry or give up gracefully */ }
else if (n < 0) { /* real error */ }
else if (n == 0) { /* peer closed */ }

2. Passing the timeout as an int of seconds instead of a struct timeval.

int secs = 5;
setsockopt(fd, SOL_SOCKET, SO_RCVTIMEO, &secs, sizeof secs); // WRONG

Why it breaks: the kernel reinterprets the int's bytes as a timeval, giving a garbage (often zero, i.e. "never time out") deadline. Always use the right struct:

struct timeval tv = { .tv_sec = 5, .tv_usec = 0 };
setsockopt(fd, SOL_SOCKET, SO_RCVTIMEO, &tv, sizeof tv);

3. Expecting SO_RCVTIMEO to bound connect().

setsockopt(fd, SOL_SOCKET, SO_RCVTIMEO, &tv, sizeof tv);
connect(fd, addr, len); // still hangs ~75s on an unreachable host

Why it breaks: the receive timeout only governs recv/read, not connection establishment. Use non-blocking connect + select/poll on writability with your own timeout, then check SO_ERROR.

4. Reusing a select timeout struct across calls on Linux.

struct timeval tv = {2,0};
while (running) select(nfds, &rs, 0, 0, &tv); // WRONG: tv may be decremented to 0

Why it breaks: Linux updates *timeout to the remaining time, so after the first timeout the loop busy-spins with a zero timeout. Re-initialize tv before every select call.

Debugging tips

  • Confirm it's actually a timeout: right after a -1 return, print strerror(errno). EAGAIN/EWOULDBLOCK = your deadline fired; ECONNRESET = peer reset; ETIMEDOUT = TCP-level timeout; EINTR = a signal interrupted you.
  • Measure the wait with clock_gettime(CLOCK_MONOTONIC, ...) around the call (as in the demo). If a "5 second" timeout returns in 0 ms, your timeval is malformed (see mistake #2).
  • strace -e trace=network ./prog (Linux) shows each recvfrom/setsockopt/select syscall and its return; you can literally watch recvfrom return -1 EAGAIN at the deadline. On macOS use dtruss.
  • Verify the option stuck with getsockopt(fd, SOL_SOCKET, SO_RCVTIMEO, &tv, &len) and print tv.
  • If it hangs anyway, attach with gdb -p <pid> and bt: a stack sitting in recv/__libc_recv confirms an unbounded blocking read, meaning the option was never applied to that fd.
  • Reproduce a silent peer locally with nc -l 12345 (listens, never sends) and point your client at it.

Memory safety

  • recv does not NUL-terminate. If you plan to treat the buffer as a C string, you must write buf[n] = '\0' using the returned length n, and your buffer must have room for that extra byte (read at most sizeof buf - 1). Printing an un-terminated buffer with %s reads out of bounds — undefined behavior.
  • Never index the buffer with an unchecked n: on -1 (timeout/error) n is negative, and using it as a length or index is UB. Branch on n first.
  • Pass the exact sizeof(struct timeval) as optlen; a wrong size makes setsockopt read past or short of the struct.
  • fd_set has a fixed capacity (FD_SETSIZE, typically 1024). Calling FD_SET with an fd >= FD_SETSIZE writes out of bounds — a classic bug when many connections are open; prefer poll/epoll at scale.
  • Concurrency: SO_RCVTIMEO is a property of the socket, so if two threads call recv on the same fd they share the deadline and can race on the buffer. Give each connection its own fd/thread or serialize access. Also handle EINTR in signal-heavy programs by retrying the call.

Real-world uses

Every robust network client and server sets timeouts. HTTP libraries (curl's CURLOPT_TIMEOUT, Go's net/http Client.Timeout) are ultimately built on these same socket options and readiness waits. Databases and RPC frameworks (gRPC deadlines, Redis client timeout) bound reads so a stalled backend can't hang the caller. Load balancers and reverse proxies (nginx proxy_read_timeout) close idle upstream connections. On the defensive side, servers set aggressive read timeouts and connection caps specifically to blunt slowloris and other slow-drip resource-exhaustion attacks. Best practice: pick timeouts from real latency budgets (a little above the p99 you expect), always distinguish timeout from error from clean-close in your handling, build poll/epoll event loops rather than one-blocking-call-per-thread when you scale, and treat a connect timeout as separate from a read timeout.

Practice tasks

  1. Detect the timeout precisely. Modify the demo so that after the SO_RCVTIMEO recv returns, it prints one of three lines depending on the outcome: "timeout", "peer closed", or "error: ". Verify each branch by adjusting whether/when the server sends or closes.
  2. Send timeout. Set SO_SNDTIMEO to 200 ms on the server socket, then have the server send a large buffer to a client that never reads. Observe and report the send returning -1 with EAGAIN once the socket buffer fills.
  3. Retry with a total budget. Wrap recv in a loop that retries on EAGAIN but gives up after a total wall-clock budget of 1 second (using CLOCK_MONOTONIC), even though each individual recv times out at 200 ms. Print how many attempts it made.
  4. poll-based version. Reimplement Technique 2 using poll() (with a struct pollfd and a millisecond timeout) instead of select, and confirm you get the same timeout behavior. Note why poll avoids the FD_SETSIZE limit.
  5. Bounded connect. Put a client socket in non-blocking mode, connect to a deliberately unreachable address (e.g. 10.255.255.1:80), then use select/poll on writability with a 1-second timeout and read SO_ERROR via getsockopt to report success, timeout, or connection error.

Summary

  • A blocking recv/read has no deadline by default — a silent peer hangs your program indefinitely; TCP won't rescue you quickly.
  • SO_RCVTIMEO (and SO_SNDTIMEO) via setsockopt with a struct timeval is the simplest fix: each blocking call gives up after the set time, returning -1 with errno == EAGAIN/EWOULDBLOCK.
  • A timeout is not a broken connection: distinguish -1+EAGAIN (timeout), 0 (clean close), and -1+other-errno (real error).
  • select/poll waits for readiness with a per-call timeout and scales to many sockets — the foundation of real event loops.
  • Receive/send timeouts do not bound connect(); use non-blocking connect + readiness wait + SO_ERROR for that.
  • Watch the details: use a struct timeval (not an int), re-init select's timeout each call, NUL-terminate before %s, and always close() your fds.

Practice with these exercises