Networking in C · advanced · ~10 min

Localhost port checking

- By the end you can explain what a TCP connect-scan does and how `connect()` success vs. `ECONNREFUSED` reveals whether a port is listening. - By the end you can write a loopback-only port checker that opens a **fresh socket per port** and reports open / refused / timeout. - By the end you can apply the non-blocking-`connect()` + `select()` timeout pattern (from the prerequisite) so a filtered or dead port does not hang your scan. - By the end you can read `SO_ERROR` with `getsockopt()` to learn the real outcome of an asynchronous connect. - By the end you can state, in engineering terms, why a checker must be pinned to `127.0.0.1` and why scanning hosts you do not own is off-limits.

Overview

You already know how to keep a single TCP connect() from blocking forever: put the socket in non-blocking mode, start the connect, and use select() with a timeval to bound the wait (the socket-timeouts prerequisite). A localhost port check is that exact skill applied in a loop. Instead of connecting to one server and reading data, you connect to a range of ports on the loopback address and record, for each one, whether the connection was accepted, refused, or timed out.

This lesson builds directly on the timeout pattern and adds three new ideas: (1) the meaning of a connection result as a signal about what is listening, (2) the discipline of a fresh socket per attempt, and (3) reading SO_ERROR to interpret an asynchronous connect. Everything here is strictly loopback — the address is always 127.0.0.1, which the kernel routes back to your own machine and never onto a wire.

Why it matters

Every service on a machine — a database, a web server, an SSH daemon — is reachable through a TCP port, and knowing which ports are open is the first thing both defenders and attackers establish. Defensively, a localhost port check is how a health check confirms "is my server actually accepting connections yet?", how a container start-up script waits for a dependency, and how a security audit verifies that a box is not exposing something it shouldn't. Getting the timeout and per-socket handling right is what separates a checker that finishes in a second from one that hangs for minutes on a filtered port. And pinning to loopback is not politeness: scanning hosts you don't own without written authorisation is unauthorised access under laws like the US CFAA and the UK Computer Misuse Act.

Core concepts

A "scan" is just connect() in a loop

A TCP connect-scan carries no secret technique. For each port you want to check, you create a socket and call connect() to 127.0.0.1:port. The kernel performs the TCP three-way handshake on your behalf, and the result is the information:

  • Accepted — some process called listen()/accept() on that port. The handshake completes; connect() succeeds. The port is open.
  • Refused (ECONNREFUSED) — nothing is listening. The host's kernel replies with a TCP RST. The port is closed.
  • No answer (timeout) — a firewall silently dropped the packet, or the host is slow. On raw loopback you rarely see this, but a real network does. The port is filtered/unknown.
  your probe            127.0.0.1 kernel
     |   SYN  -------------->  |
     |   <-------- SYN-ACK     |   port is LISTENing  => connect() == 0  (OPEN)
     |   ACK  -------------->  |

     |   SYN  -------------->  |
     |   <---------- RST       |   nobody listening   => ECONNREFUSED     (CLOSED)

     |   SYN  -------------->  |
     |        (silence)        |   packet dropped     => select() times out (FILTERED)

One socket per port — descriptors are single-use here

A socket file descriptor is the small integer the OS hands you to name one connection endpoint. Once you've called connect() on it, that descriptor is committed — you cannot rewind it and point it at a different port. So the loop must socket() a new fd for every port and close() it afterward. Forgetting the close() is a file-descriptor leak: each process has a limit (see ulimit -n), and a long scan that never closes will eventually fail with EMFILE ("too many open files").

Connect result errno Meaning Port state
returns 0 — handshake completed immediately open
returns -1 EINPROGRESS non-blocking connect still in flight pending — wait with select()
returns -1 ECONNREFUSED kernel got a RST closed
returns -1 ETIMEDOUT no reply within the OS limit filtered
returns -1 EMFILE/ENFILE out of file descriptors bug: you leaked sockets

Knowledge check: your scan of 5 ports returns EMFILE on the 4th socket even though the machine is idle. What did you forget?

You forgot to close(fd) after each probe. The descriptors from the first three attempts are still open, and combined with a low ulimit -n you ran out. Every probe must close its socket on every return path — success, refused, and timeout alike.

The non-blocking connect + select() timeout (built on the prerequisite)

A blocking connect() to a filtered port can stall for the OS default (often ~75 s). That's unacceptable in a loop. Reusing the prerequisite pattern:

  1. fcntl(fd, F_SETFL, flags | O_NONBLOCK) — make the socket non-blocking.
  2. connect() — it returns -1/EINPROGRESS immediately instead of waiting.
  3. select(fd+1, NULL, &wset, NULL, &tv) — wait for the fd to become writable, bounded by your timeval. A completed connect (success or failure) makes the socket writable.
  4. If select() returns 0, the timeout fired — treat the port as filtered and move on.

Reading the real outcome with SO_ERROR

Here's the subtlety beginners miss: when select() reports the socket is writable, that only means the connect finished — not that it succeeded. A refused connection also makes the socket writable. To learn which, ask the socket for its pending error:

getsockopt(fd, SOL_SOCKET, SO_ERROR, &so_err, &len);
//  so_err == 0            -> connect succeeded  (OPEN)
//  so_err == ECONNREFUSED -> connect refused    (CLOSED)
//  so_err == other        -> some other failure

SO_ERROR also clears the pending error as a side effect. Skipping this step is the classic bug: you assume "writable == open" and report every closed port as open.

Keep it strictly local — an engineering constraint, not just etiquette

The address must be INADDR_LOOPBACK (127.0.0.1). The defensive design is to make the tool incapable of targeting anything else: validate the host argument and refuse anything that isn't loopback, rather than trusting the caller. That way a copy-paste or a hostile config can't turn your health-checker into a network scanner. Scanning third-party hosts without explicit written authorisation is treated as unauthorised access in many jurisdictions.

Syntax notes

int socket(int domain, int type, int protocol); — AF_INET (IPv4), SOCK_STREAM (TCP), 0 (default protocol). Returns a new fd, or -1/sets errno. Must be close()d.

int connect(int fd, const struct sockaddr *addr, socklen_t len); — initiates the TCP handshake to addr. On a blocking socket, returns 0 on success or -1 with errno (ECONNREFUSED, ETIMEDOUT). On a non-blocking socket, usually returns -1/EINPROGRESS and completes later.

int fcntl(int fd, int cmd, ...); — F_GETFL reads the flag word, F_SETFL writes it. OR in O_NONBLOCK to make the socket non-blocking. Returns -1/errno on failure.

int select(int nfds, fd_set *r, fd_set *w, fd_set *e, struct timeval *tv); — nfds is the highest fd plus one. For a pending connect, watch the write set. tv bounds the wait: returns >0 (ready), 0 (timeout), -1/errno. tv may be modified — reinitialise it each call.

int getsockopt(int fd, int level, int optname, void *val, socklen_t *len); — with SOL_SOCKET/SO_ERROR, writes the socket's pending error (0 = none) into *val and clears it. Returns 0/-1.

int getsockname(int fd, struct sockaddr *addr, socklen_t *len); — fills in the local address a socket is bound to. After bind(...port 0) it tells you the kernel-chosen ephemeral port. Set *len to the buffer size before calling.

uint32_t htonl(uint32_t) / uint16_t htons(uint16_t) / ntohs — convert between host and network byte order for addresses and ports. Always wrap the port and address you put in a sockaddr_in.

Errors are reported via return value + errno; check the return before trusting errno. Every socket you open must be closed on every path.

Lesson

How a port scan works

A "scan" is simple: it tries to connect() to each port, one at a time.

  • If the connection succeeds, something is listening on that port.
  • If the connection is refused (ECONNREFUSED), nothing is listening there.

That is the whole idea. You learn what is running by seeing which connection attempts go through.

Keep it strictly local

Scanning hosts you do not own is risky, and not just technically.

Scanning third-party hosts without explicit, written authorisation counts as unauthorised access in many jurisdictions. In other words, it can be illegal.

The platform's exercises only ever target 127.0.0.1 (the loopback address, which always points back to your own machine). They run against test servers ("fixtures") that the exercise harness opens itself.

Code examples

#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <errno.h>
#include <fcntl.h>
#include <unistd.h>
#include <sys/socket.h>
#include <sys/select.h>
#include <netinet/in.h>
#include <arpa/inet.h>

/* Open a listening TCP socket on 127.0.0.1, kernel-chosen port.
   Returns the fd and writes the chosen port into *out_port. */
static int open_listener(uint16_t *out_port) {
    int fd = socket(AF_INET, SOCK_STREAM, 0);
    if (fd < 0) { perror("socket"); return -1; }

    struct sockaddr_in addr;
    memset(&addr, 0, sizeof addr);
    addr.sin_family = AF_INET;
    addr.sin_addr.s_addr = htonl(INADDR_LOOPBACK); /* 127.0.0.1 only */
    addr.sin_port = 0;                             /* 0 = pick a free port */

    if (bind(fd, (struct sockaddr *)&addr, sizeof addr) < 0) {
        perror("bind"); close(fd); return -1;
    }
    if (listen(fd, 1) < 0) { perror("listen"); close(fd); return -1; }

    /* Ask the kernel which port it actually gave us. */
    socklen_t len = sizeof addr;
    if (getsockname(fd, (struct sockaddr *)&addr, &len) < 0) {
        perror("getsockname"); close(fd); return -1;
    }
    *out_port = ntohs(addr.sin_port);
    return fd;
}

/* Try a non-blocking connect to 127.0.0.1:port with a millisecond timeout.
   Returns  1 = open (accepted), 0 = refused, -1 = timeout/error. */
static int probe_port(uint16_t port, int timeout_ms) {
    int fd = socket(AF_INET, SOCK_STREAM, 0);   /* fresh socket per probe */
    if (fd < 0) return -1;

    /* Switch to non-blocking so connect() returns immediately. */
    int flags = fcntl(fd, F_GETFL, 0);
    fcntl(fd, F_SETFL, flags | O_NONBLOCK);

    struct sockaddr_in target;
    memset(&target, 0, sizeof target);
    target.sin_family = AF_INET;
    target.sin_addr.s_addr = htonl(INADDR_LOOPBACK);
    target.sin_port = htons(port);

    int rc = connect(fd, (struct sockaddr *)&target, sizeof target);
    if (rc == 0) { close(fd); return 1; }        /* connected instantly */
    if (errno != EINPROGRESS) {                  /* e.g. ECONNREFUSED now */
        int result = (errno == ECONNREFUSED) ? 0 : -1;
        close(fd);
        return result;
    }

    /* Connect is in progress: wait for the socket to become writable. */
    fd_set wset;
    FD_ZERO(&wset);
    FD_SET(fd, &wset);
    struct timeval tv = { timeout_ms / 1000, (timeout_ms % 1000) * 1000 };

    rc = select(fd + 1, NULL, &wset, NULL, &tv);
    if (rc <= 0) { close(fd); return -1; }       /* 0 = timeout, <0 = error */

    /* Writable: read SO_ERROR to learn if the connect actually succeeded. */
    int so_err = 0;
    socklen_t elen = sizeof so_err;
    getsockopt(fd, SOL_SOCKET, SO_ERROR, &so_err, &elen);
    close(fd);
    if (so_err == 0)            return 1;         /* open */
    if (so_err == ECONNREFUSED) return 0;         /* refused */
    return -1;                                    /* other error */
}

int main(void) {
    uint16_t open_port;
    int listener = open_listener(&open_port);
    if (listener < 0) return 1;

    printf("Listener is up on 127.0.0.1:%u\n", open_port);
    printf("Scanning a small loopback range (lab-only)...\n\n");

    uint16_t lo = open_port - 2, hi = open_port + 2;
    for (uint16_t p = lo; p <= hi; p++) {
        int r = probe_port(p, 200);
        const char *label = (r == 1) ? "OPEN    "
                          : (r == 0) ? "refused "
                          :            "timeout ";
        printf("  port %-6u %s%s\n", p, label,
               (p == open_port) ? "<- our listener" : "");
    }

    close(listener);
    return 0;
}

Line by line

  • open_listener() builds the target the scan will find, so the demo is fully self-contained — no external server needed.
  • htonl(INADDR_LOOPBACK) sets the address to 127.0.0.1 in network byte order. Everything in this program is pinned to loopback.
  • addr.sin_port = 0 asks the kernel to choose a free ephemeral port. This avoids clashing with whatever else is already running.
  • bind then listen register the socket so the kernel will accept handshakes on it — this is what makes one port "open" for the scan to find.
  • getsockname reads back the actual port the kernel picked (since we asked for 0). ntohs converts it from network to host byte order for printing.
  • probe_port() — socket() at the top makes a fresh fd for this one attempt; the matching close(fd) appears on every return path so nothing leaks.
  • fcntl(... O_NONBLOCK) flips the socket to non-blocking, the prerequisite timeout pattern's first step, so connect() cannot stall the loop.
  • connect() returns 0 — rare but possible on fast loopback: the port was open and the handshake finished instantly. Return 1.
  • errno != EINPROGRESS — the connect already failed synchronously; if it's ECONNREFUSED the port is closed (return 0), otherwise an error (return -1).
  • FD_ZERO/FD_SET + select(fd+1, NULL, &wset, NULL, &tv) — wait, bounded by tv, for the socket to become writable. nfds is fd+1. select returning 0 means the timeout fired → filtered.
  • getsockopt(... SO_ERROR ...) — the crucial step: writable only means finished. Reading SO_ERROR reveals whether it finished with success (0 → open) or refusal (ECONNREFUSED → closed).
  • main() scans a tiny window around the listener's own port so you can see exactly one OPEN (our listener) surrounded by refused — demonstrating both signals in one run.

Common mistakes

1. Reusing one socket for every port.

int fd = socket(AF_INET, SOCK_STREAM, 0);
for (uint16_t p = lo; p <= hi; p++) connect(fd, addr_for(p), len); // WRONG

Why it breaks: a descriptor is committed after the first connect(); subsequent calls fail with EISCONN/EALREADY and every port looks wrong.

for (uint16_t p = lo; p <= hi; p++) {
    int fd = socket(AF_INET, SOCK_STREAM, 0); // fresh fd each time
    /* ...probe... */ close(fd);
}

2. Treating "writable" as "open".

if (select(fd+1, NULL, &wset, NULL, &tv) > 0) return 1; // WRONG: refused is also writable

Why it breaks: a refused connect also makes the socket writable, so every closed port is reported open.

int e = 0; socklen_t l = sizeof e;
getsockopt(fd, SOL_SOCKET, SO_ERROR, &e, &l);
return e == 0 ? 1 : 0; // ask the socket what really happened

3. Forgetting to close on the timeout/refused path.

if (select(...) <= 0) return -1; // WRONG: fd leaks

Why it breaks: after enough probes you hit EMFILE and every further socket() fails.

if (select(...) <= 0) { close(fd); return -1; }

4. Trusting the caller's host instead of pinning loopback.

connect_to(host, port); // WRONG: will happily scan any host given to it

Why it breaks: a stray config or copy-paste turns a health-checker into an unauthorised network scanner.

if (strcmp(host, "127.0.0.1") != 0) return -2; // refuse anything non-loopback
connect_to(host, port);

Debugging tips

  • Confirm what's actually listening before blaming your code: lsof -iTCP -sTCP:LISTEN -P -n (macOS) or ss -ltnp (Linux). If your scan says a port is closed but ss shows a listener, your address/byte-order is likely wrong.
  • Watch the syscalls: strace -e trace=network ./m (Linux) or sudo dtruss ./m (macOS) shows each socket/connect/select/getsockopt and the exact errno. This instantly reveals a leaked fd (socket called more times than close) or an unexpected EISCONN.
  • Check errno immediately after a failing call, before any other libc call can overwrite it. Print it with perror() or strerror(errno).
  • See real numbers: temporarily print so_err for every probe. If closed ports show so_err == 0, you skipped getsockopt somewhere.
  • File-descriptor pressure: run ulimit -n 8 then your scan — if it dies with EMFILE early, you have a close leak to fix.
  • Reproduce a timeout deliberately by probing a port behind a firewall DROP rule; verify your select() returns 0 and the loop keeps going instead of hanging.

Memory safety

  • struct sockaddr_in must be zeroed with memset before use; leftover stack bytes in unset fields can produce a malformed address. Set sin_family, sin_addr, and sin_port explicitly.
  • getsockopt/getsockname take an in/out length: initialise the socklen_t to sizeof(buffer) before the call, or the kernel may read/write the wrong size.
  • select reinitialises nothing for you: fd_set and the timeval are modified by the call. Re-FD_ZERO/FD_SET and rebuild tv before each select, or a subsequent wait may use a garbage timeout.
  • Never FD_SET an fd ≥ FD_SETSIZE (typically 1024). A high fd overflows the fd_set bitmap — undefined behaviour. For large scans prefer poll().
  • No dynamic memory here, so no leaks of the malloc kind — but fd leaks are the analogue. Every socket() needs a matching close() on all paths; a leaked descriptor is a leaked kernel resource.
  • Single-threaded and hermetic: this demo has no shared state, so no data races. If you parallelise a scan across threads, give each thread its own fd_set/timeval and never share a socket.

Real-world uses

  • Service health checks / readiness gates: orchestration and CI scripts probe 127.0.0.1:PORT in a loop to wait until a just-started database or web server is accepting connections before proceeding. Best practice: bounded timeout + capped retry count with backoff.
  • Container & test harness startup: tools like wait-for-it/dockerize are exactly this connect-scan-with-timeout, ensuring dependency ordering.
  • Local security auditing: verifying a hardened host is not exposing unexpected services; the defensive use is confirming a small, known set of ports and flagging anything extra.
  • Well-known scanners (nmap) generalise this, but responsibly used only against authorised targets. For your own code, best practice is: pin to loopback by construction, use non-blocking connect with a short timeout, one fd per probe closed on every path, and never widen the target host from an untrusted input.

Practice tasks

  1. Range report. Modify main to scan ports lo..hi given as two argv numbers (still forced to 127.0.0.1) and print a one-line summary: counts of open / refused / timeout.
  2. Second listener. Open a second loopback listener on another ephemeral port before scanning, and confirm your scan reports exactly two OPEN ports.
  3. Configurable timeout. Add a --timeout-ms option and demonstrate that probing a firewalled/dropped port returns timeout after roughly that many milliseconds, not the OS default.
  4. Fd-leak guard. Deliberately remove one close(fd), run under ulimit -n 6, observe the EMFILE failure, then reinstate the close and show the scan completes — and explain in a comment why.
  5. Defensive host gate. Add a probe_host(const char *host, ...) that refuses any host other than 127.0.0.1 / localhost (returning a distinct sentinel), with a unit-style main that asserts a non-loopback host is rejected before any socket is created.

Summary

  • A TCP connect-scan is just connect() per port: success = open, ECONNREFUSED = closed, timeout = filtered.
  • Use a fresh socket for every port and close() it on every return path — leaked fds cause EMFILE.
  • Reuse the timeout pattern: non-blocking socket → connect() (EINPROGRESS) → select() on the write set with a timeval.
  • After select() says writable, read SO_ERROR with getsockopt — writable means finished, not succeeded.
  • Pin to 127.0.0.1 by construction and refuse non-loopback hosts; scanning others without authorisation is unauthorised access.

Practice with these exercises