Networking in C · intermediate · ~8 min

Common socket errors

- Recognize the `errno` values you meet most often (`ECONNREFUSED`, `EADDRINUSE`, `ECONNRESET`, `EPIPE`, `EAGAIN`, `EINTR`) and say what each one means and which call produces it. - Read every socket call's return value correctly: `-1` means failure with `errno` set, `0` from `recv` means a clean peer close, and a positive value is a partial or full count. - Classify a failure as retryable versus fatal, and write a loop that retries `EINTR`/`EAGAIN` while surfacing the rest. - Tame `SIGPIPE` so a dead peer returns `EPIPE` instead of silently killing your process. - Turn a raw `errno` into a human-readable diagnosis with `perror`/`strerror` and reproduce these errors on localhost in a lab.

Overview

You already know from send/recv that send, recv, connect, bind, and friends return a count on success and -1 on failure, and that a short recv is normal. This lesson picks up exactly where that left off: it answers the question you hit the first time a call returns -1 — why did it fail, and what should I do about it? The answer lives in the global variable errno, which the kernel sets whenever a system call fails.

Rather than memorising dozens of codes, you will learn the roughly nine values that account for almost all real socket failures, where each one comes from, and the single most important decision they drive: retry, or give up and report. That decision is the difference between a server that survives a signal or a slow client and one that crashes on the first hiccup.

Why it matters

Networked code fails constantly and normally — ports get reused, clients vanish mid-request, signals interrupt syscalls, and non-blocking sockets say "not yet" all day long. A program that treats every -1 as fatal will crash under exactly the conditions a server must survive; one that treats every -1 as retryable will spin forever on a genuinely dead connection. In security terms, mishandled errno is a classic source of denial-of-service and of confusing, exploitable state: an unignored SIGPIPE lets any client that hangs up at the wrong moment kill your process, and swallowing ECONNRESET/EPIPE silently can leave you reading stale or attacker-influenced buffers. Getting these codes right is table stakes for robust, hard-to-crash network services.

Core concepts

errno: how a failed syscall tells you why

Every socket system call follows the same contract: on success it returns a count or 0, and on failure it returns -1 and stores a reason code in the thread-local global errno (from <errno.h>). errno is only meaningful immediately after a call that reported failure — a successful call may leave it at any value, so never test errno unless the return value already told you something went wrong.

    rc = connect(fd, ...);
            |
     +------+------+
     |             |
   rc == 0      rc == -1
  success    read errno NOW
             (before any other
              libc call clobbers it)

The trap: almost any library call (even printf) can overwrite errno. If you need the value later, copy it into a local int the instant the call returns.

The codes worth memorising

errno Typical call Meaning First response
ECONNREFUSED connect Nothing is listening on that address:port (peer sent an RST). Fatal for this attempt; report or try another endpoint.
EADDRINUSE bind The port is already bound. Set SO_REUSEADDR, or pick another port.
EACCES bind Tried a privileged port (<1024) without privilege. Use a high port or gain privilege.
ETIMEDOUT connect The SYN got no answer in time. Fatal for this attempt; maybe retry with backoff.
ECONNRESET recv/send Peer crashed or closed abruptly (RST). Fatal for this connection; clean up.
EPIPE send You wrote to a peer that already closed. Fatal for this connection; stop writing.
EAGAIN (== EWOULDBLOCK) non-blocking recv/send Operation would block right now. Retry later (e.g. after poll).
EINTR any blocking syscall A signal interrupted the call before any data moved. Retry the call immediately.
ENETUNREACH connect No route to the destination. Fatal; common in --network=none sandboxes.

A few of these lean on TCP control packets: SYN starts a connection, FIN closes it cleanly, and RST tears it down abruptly. ECONNREFUSED and ECONNRESET both arrive as an RST — the first when no one is listening, the second when an established peer dies.

Knowledge check — you call connect("127.0.0.1", 9) and there is no server on port 9. Which errno do you get, and why is it not ETIMEDOUT?

You get ECONNREFUSED. The host is up and reachable, so its kernel immediately answers your SYN with an RST saying "no one is home on that port." ETIMEDOUT is for the opposite situation — the SYN is sent but nothing at all comes back before the timer expires (host down, packet dropped, firewall blackhole).

Return value first, errno second

Getting the return-value semantics right matters as much as the code:

  • recv returning 0 is not an error — it means the peer performed an orderly shutdown (FIN). errno is irrelevant here.
  • recv/send returning a positive number can be less than you asked for; that is a short transfer, not a failure (covered in send/recv).
  • Only a -1 return means "consult errno."
 recv() result:
   n > 0   -> got n bytes (maybe fewer than asked)
   n == 0  -> peer closed cleanly  (done, not an error)
   n < 0   -> failure; look at errno
              EINTR/EAGAIN -> retry
              everything else -> fatal for this connection

Retryable vs. fatal: the one decision that matters

Every robust socket loop splits -1 into two buckets. Retryable errors — EINTR (a signal interrupted a blocking call) and EAGAIN/EWOULDBLOCK (a non-blocking call has no data yet) — mean nothing went wrong, just run the call again (for EAGAIN, after the fd is ready). Fatal errors — everything else — mean this connection or operation is over: report it and clean up. Note EAGAIN and EWOULDBLOCK may be the same numeric value or two different ones depending on the platform, so test for both when portability matters.

SIGPIPE: the error that kills you before you can read errno

There is one nasty special case. When you send to a socket whose peer has closed, the kernel would set errno = EPIPE — but it also raises the SIGPIPE signal, whose default action is to terminate your process. So a naive server can be killed simply because a client hung up at the wrong instant. The fix is to disable that behaviour once, at startup, and then let the ordinary -1/EPIPE path handle it:

Approach Scope Notes
signal(SIGPIPE, SIG_IGN) Whole process Simplest and portable; send then returns -1/EPIPE.
send(..., MSG_NOSIGNAL) Per call Linux; no SIGPIPE for that call.
setsockopt(SO_NOSIGPIPE) Per socket BSD/macOS equivalent.

Ignoring SIGPIPE at startup is the standard move for any server. It does not hide the error — it converts an unignorable process kill into a normal, checkable EPIPE return.

Syntax notes

#include <errno.h>
extern int errno;  /* actually a thread-local lvalue macro; set on failure, valid only then */

#include <string.h>
char *strerror(int errnum);   /* returns a human string for an errno value; not thread-safe */

#include <stdio.h>
void perror(const char *s);   /* prints: "s: <strerror(errno)>\n" to stderr */

#include <sys/socket.h>
/* All return -1 and set errno on failure. */
int     connect(int fd, const struct sockaddr *addr, socklen_t len);   /* 0 on success */
int     bind(int fd, const struct sockaddr *addr, socklen_t len);      /* 0 on success */
ssize_t recv(int fd, void *buf, size_t n, int flags);   /* >0 bytes, 0 = peer closed, -1 = error */
ssize_t send(int fd, const void *buf, size_t n, int flags); /* >=0 bytes sent, -1 = error */
int     setsockopt(int fd, int level, int opt, const void *val, socklen_t len); /* 0 on success */

#include <signal.h>
/* Install SIG_IGN for SIGPIPE so a dead peer yields EPIPE instead of killing you. */
sig_t signal(int sig, sig_t handler);   /* returns previous handler, SIG_ERR on failure */

Notes: errno is thread-local — each thread has its own copy, so a failure in one thread does not corrupt another's diagnosis. strerror returns a pointer to a static buffer (use strerror_r for thread-safe formatting). perror(NULL) prints just the error string. SO_REUSEADDR is passed to setsockopt with level = SOL_SOCKET to avoid EADDRINUSE on quick restarts. None of these functions allocate memory you must free; the fds you create with socket/accept/socketpair must be close()d.

Lesson

When a socket call fails, it returns -1 and sets the global variable errno. The value of errno tells you why it failed. Knowing the common values lets you decide whether to retry, give up, or report the problem.

Common errno values

errno Where you see it What it means
EADDRINUSE bind The port is already taken. Use SO_REUSEADDR or pick another port.
EACCES bind You tried to use a privileged port (below 1024) without root.
ECONNREFUSED connect Nothing is listening on that address and port.
ETIMEDOUT connect The connection request (SYN) got no answer.
ECONNRESET recv / send The peer crashed or closed abruptly (it sent an RST instead of a clean FIN).
EPIPE send The peer already closed. SIGPIPE was raised (unless you used MSG_NOSIGNAL).
EAGAIN non-blocking sockets The call would block right now. Try again later.
EINTR any syscall A signal interrupted the call. Retry it, or set SA_RESTART.
ENETUNREACH connect No route to the destination exists. You will see this in --network=none sandboxes.

A quick note on the terms above:

  • SYN, FIN, RST are TCP control packets. SYN starts a connection, FIN closes it cleanly, and RST tears it down abruptly.
  • SIGPIPE is a signal the system sends when you write to a closed connection. By default it kills your program.

Pattern for robust code

Wrap every socket call in an error check:

if (rc < 0) {
    switch (errno) {
        /* handle each case */
    }
}

The key idea is that your loop must know the difference between two kinds of failure:

  • Retryable errors (EINTR, EAGAIN) — try the call again.
  • Fatal errors (everything else) — stop and report.

Code examples

#include <stdio.h>
#include <string.h>
#include <errno.h>
#include <signal.h>
#include <unistd.h>
#include <sys/socket.h>
#include <netinet/in.h>
#include <arpa/inet.h>

/* Map the errno values you meet most often to a printable name. */
static const char *errno_name(int e) {
    switch (e) {
        case EADDRINUSE:   return "EADDRINUSE";
        case EACCES:       return "EACCES";
        case ECONNREFUSED: return "ECONNREFUSED";
        case ETIMEDOUT:    return "ETIMEDOUT";
        case ECONNRESET:   return "ECONNRESET";
        case EPIPE:        return "EPIPE";
        case EAGAIN:       return "EAGAIN";
        case EINTR:        return "EINTR";
        case ENETUNREACH:  return "ENETUNREACH";
        default:           return "other";
    }
}

/* A recv() that transparently retries the retryable errors. Returns the
   byte count, 0 on clean peer close, or -1 (errno set) on a fatal error. */
static ssize_t recv_retry(int fd, void *buf, size_t cap) {
    for (;;) {
        ssize_t n = recv(fd, buf, cap, 0);
        if (n >= 0) return n;                 /* data, or 0 = peer closed */
        if (errno == EINTR) continue;         /* signal: just try again  */
        if (errno == EAGAIN) continue;        /* would block: try again  */
        return -1;                            /* everything else: fatal  */
    }
}

/* Demo 1: connect to a loopback port with no listener -> ECONNREFUSED. */
static void demo_connection_refused(void) {
    /* Grab a free port, then close it so nobody is listening. */
    int probe = socket(AF_INET, SOCK_STREAM, 0);
    struct sockaddr_in a = {0};
    a.sin_family = AF_INET;
    a.sin_addr.s_addr = htonl(INADDR_LOOPBACK);
    a.sin_port = 0;                           /* let the kernel pick one */
    bind(probe, (struct sockaddr *)&a, sizeof a);
    socklen_t alen = sizeof a;
    getsockname(probe, (struct sockaddr *)&a, &alen);
    close(probe);                             /* port is now unused      */

    int fd = socket(AF_INET, SOCK_STREAM, 0);
    int rc = connect(fd, (struct sockaddr *)&a, sizeof a);
    printf("connect(closed port %d): rc=%d errno=%s\n",
           ntohs(a.sin_port), rc, rc < 0 ? errno_name(errno) : "-");
    close(fd);
}

/* Demo 2: bind the same port twice without SO_REUSEADDR -> EADDRINUSE. */
static void demo_address_in_use(void) {
    struct sockaddr_in a = {0};
    a.sin_family = AF_INET;
    a.sin_addr.s_addr = htonl(INADDR_LOOPBACK);
    a.sin_port = 0;

    int s1 = socket(AF_INET, SOCK_STREAM, 0);
    bind(s1, (struct sockaddr *)&a, sizeof a);
    listen(s1, 1);
    socklen_t alen = sizeof a;
    getsockname(s1, (struct sockaddr *)&a, &alen);   /* learn the port   */

    int s2 = socket(AF_INET, SOCK_STREAM, 0);
    int rc = bind(s2, (struct sockaddr *)&a, sizeof a);
    printf("bind(port %d again): rc=%d errno=%s\n",
           ntohs(a.sin_port), rc, rc < 0 ? errno_name(errno) : "-");
    close(s1);
    close(s2);
}

/* Demo 3: write to a peer that has closed -> EPIPE (SIGPIPE ignored). */
static void demo_broken_pipe(void) {
    int sv[2];
    socketpair(AF_UNIX, SOCK_STREAM, 0, sv);
    close(sv[1]);                             /* peer is gone            */
    ssize_t rc = send(sv[0], "hi", 2, 0);
    printf("send(closed peer): rc=%zd errno=%s\n",
           rc, rc < 0 ? errno_name(errno) : "-");
    close(sv[0]);
}

/* Demo 4: the retry wrapper succeeds on a live connection. */
static void demo_retry_ok(void) {
    int sv[2];
    socketpair(AF_UNIX, SOCK_STREAM, 0, sv);
    send(sv[1], "pong", 4, 0);
    char buf[16];
    ssize_t n = recv_retry(sv[0], buf, sizeof buf);
    if (n > 0)      printf("recv_retry: got %zd bytes '%.*s'\n", n, (int)n, buf);
    else if (n == 0) printf("recv_retry: peer closed\n");
    else            printf("recv_retry: fatal %s\n", errno_name(errno));
    close(sv[0]);
    close(sv[1]);
}

int main(void) {
    signal(SIGPIPE, SIG_IGN);   /* turn the SIGPIPE kill into an EPIPE return */
    demo_connection_refused();
    demo_address_in_use();
    demo_broken_pipe();
    demo_retry_ok();
    return 0;
}

Line by line

  • errno_name — a tiny lookup that turns numeric errno values into names for readable output. This is the manual version of what strerror does with human sentences; it exists so the demo can label each result.
  • recv_retry — the heart of the lesson. It loops calling recv: n >= 0 (data or the clean-close 0) is returned straight to the caller; EINTR and EAGAIN continue the loop (retryable); any other errno returns -1 with errno still set (fatal). This is the retry/fatal split from the concepts, in code.
  • demo_connection_refused — binds a throwaway socket to an ephemeral port (sin_port = 0 asks the kernel to choose), reads the chosen port back with getsockname, then closes it so the port is guaranteed empty. Connecting there produces ECONNREFUSED because the loopback host is up and immediately answers with an RST. Everything stays on 127.0.0.1, so it works even in a no-network sandbox.
  • demo_address_in_use — binds and listens on one socket, discovers its port, then tries to bind a second socket to the same port. Without SO_REUSEADDR the kernel rejects it with EADDRINUSE — the classic "address already in use" you hit restarting a server too quickly.
  • demo_broken_pipe — socketpair gives two connected fds. Closing sv[1] kills the peer; sending on sv[0] would raise SIGPIPE, but because main installed SIG_IGN the call instead returns -1 with EPIPE, which we can print and handle.
  • demo_retry_ok — proves the wrapper on the happy path: one end sends "pong", recv_retry on the other returns 4 and the bytes, showing that the retry logic is invisible when nothing goes wrong.
  • main — the single most important line is signal(SIGPIPE, SIG_IGN): install it before any send so a dead peer becomes a checkable EPIPE rather than a process kill. Then each demo runs and prints its rc/errno.

Common mistakes

  • Reading errno after a success (or too late).
recv(fd, buf, n, 0);
if (errno == ECONNRESET) { /* WRONG */ }

Why it breaks: errno is only meaningful when the call returned -1; a successful call may leave any stale value there, and even printf can overwrite it. Fixed:

ssize_t n = recv(fd, buf, cap, 0);
if (n < 0) { int e = errno; if (e == ECONNRESET) {/* ... */} }
  • Treating recv() == 0 as an error.
if (recv(fd, buf, cap, 0) <= 0) { perror("recv"); die(); } /* WRONG */

Why it breaks: 0 is a clean peer shutdown (FIN), not a failure — and perror will print a garbage reason from a stale errno. Fixed:

ssize_t n = recv(fd, buf, cap, 0);
if (n == 0)      { /* peer closed, done */ }
else if (n < 0)  { perror("recv"); }
  • Retrying everything (or nothing).
do { n = recv(fd, buf, cap, 0); } while (n < 0); /* WRONG: spins on ECONNRESET */

Why it breaks: fatal errors like ECONNRESET never clear, so the loop spins forever burning CPU. Fixed: retry only the retryable codes:

do { n = recv(fd, buf, cap, 0); } while (n < 0 && (errno == EINTR || errno == EAGAIN));
  • Forgetting SIGPIPE, so a hung-up client kills the server.
send(fd, buf, len, 0);   /* WRONG if peer closed: process is killed */

Why it breaks: writing to a closed peer raises SIGPIPE, whose default action terminates the process — you never even reach your error check. Fixed: neutralise it once at startup:

signal(SIGPIPE, SIG_IGN);
if (send(fd, buf, len, 0) < 0 && errno == EPIPE) { /* peer gone */ }

Debugging tips

  • perror("context") / strerror(errno) first. The fastest diagnosis is to print perror("connect") right after a -1; it turns the numeric code into connect: Connection refused. Always copy errno to a local before doing anything else if you need it twice.
  • errno -l / errno <name> (from the moreutils errno tool) lists every code and its meaning when you have a raw number.
  • strace -e trace=network ./prog (Linux) or dtruss/ktrace on macOS shows each syscall with its return value and errno symbol, e.g. connect(...) = -1 ECONNREFUSED — invaluable when you cannot tell which call failed.
  • ss -ltnp / netstat -an reveals whether something is actually listening on the port you expect — the quickest way to confirm an ECONNREFUSED or diagnose EADDRINUSE (look for TIME_WAIT sockets holding the port).
  • In gdb, set catch signal SIGPIPE to see whether an unignored SIGPIPE is what is killing you, and print errno directly after a failing call.
  • printf-debug the split: log the name of the errno (as this lesson's errno_name does) at each -1 so you can see, in production logs, whether failures are retryable noise or genuine faults.

Memory safety

  • Only touch the buffer after a positive count. After recv, the bytes buf[0 .. n-1] are valid only when n > 0. Reading buf after a -1 or 0 return reads uninitialised or stale memory — a real information-disclosure risk if you then log or forward it. Guard every buffer use behind n > 0, and if you NUL-terminate for string use, ensure n < cap so buf[n] = 0 stays in bounds.
  • errno is thread-local, your buffers are not. In a threaded server each thread sees its own errno, so diagnoses do not cross-contaminate — but a shared receive buffer absolutely can, so give each connection its own buffer or protect it.
  • strerror is not thread-safe. Two threads calling strerror can clobber the shared static string. Use strerror_r (or per-thread buffers) in concurrent code.
  • Don't call async-unsafe functions from a real SIGPIPE handler. If you install an actual handler instead of SIG_IGN, it must only use async-signal-safe calls; SIG_IGN sidesteps this entirely, which is why it is preferred.
  • Signals interrupt syscalls (EINTR). Any blocking call can return early with no data moved; treating that partial state as if data arrived leads to reading past what was actually received. The retry wrapper is the safe pattern.

Real-world uses

  • Servers and daemons (nginx, redis, postgres) ignore SIGPIPE at startup and branch on EPIPE/ECONNRESET so a client that hangs up mid-response never takes the server down.
  • Connection retry with backoff: clients treat ETIMEDOUT/ECONNREFUSED as "try again later" with exponential backoff, while treating ENETUNREACH as "give up, the route is gone."
  • Event loops (libevent, libuv, epoll/kqueue reactors) are built entirely around EAGAIN: a non-blocking recv returning EAGAIN is the normal signal to go back to poll/epoll_wait and wait for readiness, not an error to log.
  • SO_REUSEADDR on restart: virtually every long-running listener sets it so a redeploy does not fail with EADDRINUSE while old sockets linger in TIME_WAIT.
  • Best practice: wrap syscalls in thin helpers (recv_retry, send_all) that centralise the retryable/fatal split, always log the errno name, and never trust a buffer you did not receive positive bytes into.

Practice tasks

  1. Reproduce ECONNREFUSED. Write a program that binds a loopback socket to an ephemeral port, reads the port with getsockname, closes it, then connects to that port and prints the errno name. Confirm you see ECONNREFUSED.

  2. Robust recv wrapper. Implement ssize_t recv_all(int fd, void *buf, size_t n) that loops until n bytes arrive, retries EINTR/EAGAIN, returns 0 on clean peer close, and -1 on any fatal errno. Test it against a socketpair that sends the data in two chunks.

  3. Defeat SIGPIPE. Build a socketpair, close one end, then send on the other both with and without signal(SIGPIPE, SIG_IGN). Show that the ignored version returns -1/EPIPE while the default version terminates the process (check the exit status).

  4. Cure EADDRINUSE. Reproduce EADDRINUSE by binding the same port twice, then fix the second bind by setting SO_REUSEADDR with setsockopt before binding. Print the results of both attempts.

  5. errno dispatcher. Write a function that, given a failing socket call's errno, classifies it as RETRY, FATAL_CONNECTION, or FATAL_STARTUP, and returns a suggested action string. Cover at least EINTR, EAGAIN, ECONNRESET, EPIPE, ECONNREFUSED, EADDRINUSE, and EACCES, then drive it with a small table of inputs.

Summary

  • A failed socket call returns -1 and sets errno; read errno only after a -1, and copy it to a local before any other call can clobber it.
  • Learn the core codes: ECONNREFUSED (no listener), EADDRINUSE (port taken), ETIMEDOUT/ENETUNREACH (can't reach), ECONNRESET/EPIPE (peer died), EAGAIN (would block), EINTR (signal).
  • recv() == 0 is a clean peer close, not an error; a positive count may be short; only -1 means consult errno.
  • Split every -1 into retryable (EINTR, EAGAIN/EWOULDBLOCK) — run the call again — versus fatal — report and clean up.
  • Install signal(SIGPIPE, SIG_IGN) at startup so writing to a dead peer yields a checkable EPIPE instead of killing your process.
  • Use perror/strerror, and tools like strace/ss, to diagnose; touch a receive buffer only after a positive count.

Practice with these exercises