Networking in C · beginner · ~6 min

socket() — creating an endpoint

- By the end you can call `socket()` and read back the file descriptor it returns. - You can choose the right `family`, `type`, and `proto` arguments for a TCP or UDP endpoint over IPv4. - You can check `socket()` for failure the C way — test for `-1`, then read `errno`. - You can explain why a fresh socket is *unbound* and what steps must follow before it can carry data. - You can close a socket correctly and explain why a leaked socket is exactly a leaked file descriptor.

Overview

You already know from the TCP-vs-UDP lesson that TCP gives you a reliable, connection-oriented byte stream while UDP gives you cheap, unordered, connectionless datagrams. That was the networking decision. This lesson is where you make that decision concrete in C: socket() is the one call that turns "I want a TCP endpoint" or "I want a UDP endpoint" into an actual kernel object your program can hold onto.

Think of socket() as malloc() for the network: it asks the kernel to allocate a communication endpoint and hands you back a small integer — a file descriptor — that names it. Everything else in the networking chapter (bind, listen, connect, send, recv) operates on that descriptor. Nothing here re-teaches TCP vs UDP; instead you translate that choice into the exact arguments socket() expects.

Why it matters

Every network program in the world — your browser, ssh, a database driver, a game server — starts with a socket() call; if you get the arguments or the error handling wrong here, nothing downstream can work. Because a socket is a file descriptor, mishandling it causes the same real bugs as any FD leak: a long-running server that never closes failed sockets slowly exhausts its descriptor table and starts refusing connections with EMFILE. From a defensive standpoint, disciplined error checking and prompt close() on every path (including error paths) is the difference between a server that degrades gracefully under load or attack and one that silently falls over.

Core concepts

A socket is a kernel object you reach through a file descriptor

When you call socket(), the kernel allocates the data structures for a communication endpoint (send/receive buffers, protocol state, and so on) and records them in your process's file descriptor table. You never touch those structures directly — you get back an int index into that table, and every later call passes that int back to the kernel.

  Your process                     Kernel
  ------------                     ------
  int fd = 3; ------------+
                          |   FD table (per process)
                          v   +-----+---------------------------+
                              |  0  | stdin                     |
                              |  1  | stdout                    |
                              |  2  | stderr                    |
                              |  3  | ---> [ TCP socket object ] |  <- socket() gave you this
                              |  4  | ---> [ UDP socket object ] |
                              +-----+---------------------------+

That is why FDs 3 and 4 show up first in the demo: 0, 1, 2 are already taken by the standard streams, and the kernel hands out the lowest free index.

The three arguments, decoded

socket(family, type, proto) is small but every argument matters.

Argument Common values Meaning
family AF_INET, AF_INET6 Address family. AF_INET = IPv4 (this course), AF_INET6 = IPv6. Decides what an address looks like.
type SOCK_STREAM, SOCK_DGRAM Communication semantics. SOCK_STREAM = reliable ordered byte stream (TCP-style). SOCK_DGRAM = independent messages, best-effort (UDP-style).
proto 0 (usual), IPPROTO_TCP, IPPROTO_UDP The specific protocol. 0 tells the kernel "pick the default for this family+type" — TCP for stream, UDP for datagram.

The family/type pair almost always determines the protocol, so passing 0 for proto is idiomatic and correct for the sockets you will build in this course.

Knowledge check: You want a UDP endpoint over IPv4. What do you pass?

socket(AF_INET, SOCK_DGRAM, 0). AF_INET selects IPv4, SOCK_DGRAM selects datagram (message) semantics, and 0 lets the kernel fill in the default protocol for that pair, which is UDP. You could write IPPROTO_UDP explicitly, but 0 is the conventional choice.

The return value and the -1 / errno convention

socket() follows the classic POSIX system-call contract:

Outcome Return value What to do
Success A non-negative FD (e.g. 3) Use it, and remember to close() it later.
Failure -1 Read errno (via perror/strerror) to learn why.

Common failure reasons you should recognize: EAFNOSUPPORT (family not supported), EPROTONOSUPPORT (bad/unsupported protocol), and EMFILE/ENFILE (too many open descriptors — the FD-leak symptom). Because a valid FD can legitimately be 0, the correct failure test is < 0 (equivalently == -1), never "is it zero?".

A fresh socket is unbound and peerless

Right after socket() succeeds you have an endpoint that is not tied to any local address or port and is not connected to anyone. It cannot yet send or receive application data. What comes next depends on the role:

              socket()
                 |
         +-------+--------+
         |                |
      SERVER            CLIENT
         |                |
      bind()           connect()  --> then send()/recv()
      listen()
      accept()  --> then send()/recv()

This lesson stops at the very first box; the later lessons fill in the rest. The key mental model: socket() creates the endpoint; other calls give it an identity and a partner.

Syntax notes

#include <sys/socket.h>   /* socket(), AF_*, SOCK_* */
#include <netinet/in.h>   /* IPPROTO_* constants, sockaddr_in */
#include <unistd.h>       /* close() */

int socket(int family, int type, int protocol);
//  |         |           |          |
//  |         |           |          +-- 0 = default protocol for family/type;
//  |         |           |              or IPPROTO_TCP / IPPROTO_UDP explicitly
//  |         |           +------------- SOCK_STREAM (TCP) | SOCK_DGRAM (UDP) | ...
//  |         +------------------------- AF_INET (IPv4) | AF_INET6 (IPv6) | ...
//  +----------------------------------- returns: >=0 file descriptor on success,
//                                                -1 on error (errno set)

int close(int fd);   // returns 0 on success, -1 on error; releases the FD
  • Return: a non-negative file descriptor on success, -1 on failure. Test with < 0.
  • errno: only meaningful after a -1 return; read it immediately, before any other call can overwrite it.
  • Must close: every successfully created socket must eventually be passed to close(). The FD is a finite kernel resource; leaking it is a real bug.
  • No malloc involved: you did not allocate any heap memory, so there is nothing to free() — the resource to release is the descriptor, via close().

Lesson

What socket() does

int socket(int family, int type, int proto) asks the kernel to allocate a new socket. It returns a file descriptor (an integer the kernel uses to identify an open resource) for that socket.

The socket starts out empty: it is not yet tied to an address or a peer.

The three arguments

  • family — the address family.
    • AF_INET for IPv4.
    • AF_INET6 for IPv6.
    • Use AF_INET throughout this course.
  • type — the kind of communication.
    • SOCK_STREAM for TCP (reliable, connection-based).
    • SOCK_DGRAM for UDP (message-based, no connection).
  • proto — the protocol.
    • Pass 0 to let the kernel choose the default for the family/type pair.
    • That default is TCP for SOCK_STREAM and UDP for SOCK_DGRAM.

Return value

socket() returns the new file descriptor on success.

On failure it returns -1 and sets errno. Always check the result before using the descriptor.

Code examples

#include <stdio.h>
#include <string.h>
#include <errno.h>
#include <unistd.h>
#include <sys/socket.h>
#include <netinet/in.h>

/* Print a compact description of one socket file descriptor. */
static void describe(const char *label, int fd) {
    if (fd < 0) {
        printf("%-24s -> FAILED: %s (errno=%d)\n", label, strerror(errno), errno);
    } else {
        printf("%-24s -> fd=%d (open, unbound)\n", label, fd);
    }
}

int main(void) {
    /* 1. A TCP (stream) socket over IPv4. proto=0 => kernel picks TCP. */
    int tcp = socket(AF_INET, SOCK_STREAM, 0);
    describe("TCP  IPv4 SOCK_STREAM", tcp);

    /* 2. A UDP (datagram) socket over IPv4. proto=0 => kernel picks UDP. */
    int udp = socket(AF_INET, SOCK_DGRAM, 0);
    describe("UDP  IPv4 SOCK_DGRAM", udp);

    /* 3. Fresh descriptors are small non-negative ints handed out lowest-first. */
    printf("\nEach socket() call returned a distinct file descriptor.\n");
    printf("They start unbound: no local address, no peer yet.\n");

    /* 4. A deliberately invalid call to show the -1 / errno convention. */
    int bad = socket(AF_INET, SOCK_STREAM, 999);   /* bogus protocol */
    describe("bogus proto (expected)", bad);
    if (bad >= 0) close(bad);   /* just in case a platform accepts it */

    /* 5. A socket is a file descriptor: not closing it leaks a kernel resource. */
    if (tcp >= 0) close(tcp);
    if (udp >= 0) close(udp);
    printf("\nClosed all valid sockets. No descriptors leaked.\n");
    return 0;
}

Line by line

  • The includes. <sys/socket.h> declares socket() and the AF_*/SOCK_* constants; <netinet/in.h> brings in the IPPROTO_* protocol constants; <unistd.h> declares close(); <errno.h>/<string.h> give us errno and strerror for readable error messages.
  • describe() helper. Centralizes the success/failure branch so every call site checks fd < 0 the same way. On failure it prints strerror(errno) — the human-readable text for the error code — right after the failing call, before anything can clobber errno.
  • Line marked 1 (tcp = socket(AF_INET, SOCK_STREAM, 0)). Creates an IPv4 stream endpoint. proto=0 means "default for AF_INET + SOCK_STREAM," which is TCP. On success you get the lowest free FD, typically 3.
  • Line marked 2 (udp = socket(AF_INET, SOCK_DGRAM, 0)). Same idea but SOCK_DGRAM selects datagram semantics; proto=0 resolves to UDP. It gets the next FD, typically 4, proving each call yields a distinct descriptor.
  • Block 3 (the two printfs). Reinforces the core idea: the calls succeeded, the FDs differ, and both sockets are still unbound — no address, no peer.
  • Line marked 4 (socket(AF_INET, SOCK_STREAM, 999)). Passes a nonsense protocol number so the call fails on purpose. This demonstrates the -1 return and a real errno (EPROTONOSUPPORT), so you see the error path instead of just reading about it. The guarded close(bad) is defensive in case some platform tolerates it.
  • Block 5 (close(tcp), close(udp)). Releases each valid descriptor back to the kernel. Guarding with >= 0 avoids calling close() on a -1 that was never a real FD. The final message confirms nothing leaked.

Common mistakes

1. Testing the wrong failure condition.

int fd = socket(AF_INET, SOCK_STREAM, 0);
if (fd == 0) { /* handle error */ }   // WRONG

Why it breaks: 0 is a perfectly valid file descriptor (it's just stdin's slot when free). The error sentinel is -1, so this test misses real failures and can misfire on a legitimate FD.

if (fd < 0) { perror("socket"); return 1; }   // FIXED

2. Leaking the socket on the error path.

int fd = socket(AF_INET, SOCK_STREAM, 0);
if (bind(fd, ...) < 0) return 1;   // WRONG: fd never closed

Why it breaks: the socket is an open descriptor; returning without close() leaks it. In a loop or long-running server this eventually hits EMFILE and new sockets fail.

if (bind(fd, ...) < 0) { close(fd); return 1; }   // FIXED

3. Reading errno too late.

int fd = socket(AF_INET, SOCK_STREAM, 0);
printf("trying...\n");                 // this call can change errno
if (fd < 0) perror("socket");         // WRONG: errno may be stale

Why it breaks: any intervening library call may overwrite errno, so the reported reason can be wrong.

int fd = socket(AF_INET, SOCK_STREAM, 0);
if (fd < 0) { perror("socket"); return 1; }   // FIXED: check before anything else

4. Mismatching type and protocol.

int fd = socket(AF_INET, SOCK_STREAM, IPPROTO_UDP);   // WRONG

Why it breaks: asking for stream semantics but UDP as the protocol is inconsistent and fails with EPROTONOSUPPORT.

int fd = socket(AF_INET, SOCK_STREAM, 0);   // FIXED: let 0 pick TCP

Debugging tips

  • perror("socket") first. The moment a call returns -1, print perror (or strerror(errno)) — it turns an opaque -1 into Protocol not supported, Address family not supported, etc.
  • Watch for EMFILE/ENFILE. If socket() starts failing only after the program has run a while, you are leaking descriptors. Check ulimit -n for the per-process cap.
  • List open FDs while running (Linux): ls -l /proc/<pid>/fd shows every open descriptor, including sockets (they appear as socket:[...]). A steadily growing list confirms a leak. On macOS use lsof -p <pid>.
  • strace ./prog (Linux) / dtruss ./prog (macOS). Traces the actual socket(...) = 3 / = -1 EPROTONOSUPPORT return so you can see exactly which argument the kernel rejected.
  • Under gdb, set break socket and inspect the arguments, then finish and print $eax/$rax (or just the assigned variable) to see the returned FD.
  • Print the FD. A quick printf("fd=%d\n", fd) early on distinguishes "never created" (-1) from "created but misused later."

Memory safety

  • The resource here is a descriptor, not heap memory. socket() allocates no memory you own, so there is nothing to free(). The leak hazard is failing to close() the FD — a resource leak, not a heap leak, but just as real. Tools like valgrind --track-fds=yes report descriptors still open at exit.
  • Never use an FD after close(). Once closed, the integer is dead; using it is a use-after-close. Worse, the kernel may reuse that number for the next socket()/open(), so a stale FD can silently point at an unrelated resource — the descriptor analogue of a use-after-free dangling pointer. Set the variable to -1 after closing if it might be reused.
  • Don't double-close. Calling close() twice on the same FD can close a different resource that reused the number in between. Guard closes and null out the FD.
  • errno is thread-local but volatile across calls. Read it immediately after the failing call; do not let any other library or system call run in between.

Real-world uses

  • Servers and clients everywhere. Web servers (nginx, Apache), databases (PostgreSQL, Redis), and language runtimes all begin every connection path with socket(). Higher-level libraries (libcurl, Python's socket module, Go's net package) wrap this exact syscall.
  • Choosing type by workload. SOCK_STREAM/TCP for anything needing reliability and ordering (HTTP, SSH, SQL); SOCK_DGRAM/UDP for latency-sensitive or broadcast-style traffic (DNS queries, game state, metrics, DHCP).
  • Best practice — RAII-style discipline in C. Treat every successful socket() like an acquired resource: pair it with a close() on every exit path (success and error), initialize FD variables to -1, and reset to -1 after closing. On Linux, prefer SOCK_CLOEXEC (socket(AF_INET, SOCK_STREAM | SOCK_CLOEXEC, 0)) so the descriptor doesn't leak across an exec() into a child process — a small but important hardening step in security-sensitive services.
  • Capacity planning. Because each socket consumes a descriptor, high-concurrency servers raise ulimit -n and audit for leaks; a descriptor leak is a classic slow-burn availability bug and a denial-of-service risk under load.

Practice tasks

  1. Create and report. Write a program that creates a single IPv4 TCP socket, prints the returned file descriptor, closes it, and prints a confirmation. Verify the FD is 3.
  2. Two flavors. Create one SOCK_STREAM and one SOCK_DGRAM socket over AF_INET, print both FDs, and confirm they differ. Explain in a comment which protocol each resolves to when proto is 0.
  3. Force a failure. Deliberately call socket() with an unsupported argument (e.g. a bogus protocol or AF_INET6, SOCK_DGRAM, IPPROTO_TCP), detect the -1, and print the reason with perror. Confirm the reported errno name.
  4. Leak detector. Loop calling socket() without closing until it returns -1; print how many succeeded before failure and the errno. Then fix the loop by closing each socket and show it now runs indefinitely (cap the count). Relate the failure to ulimit -n.
  5. Harden it. Write a make_socket(int type) helper that returns an FD or -1, uses SOCK_CLOEXEC where available, checks the return, and never leaks on failure. Call it for both TCP and UDP and close everything on all paths.

Summary

  • socket(family, type, proto) asks the kernel for a new communication endpoint and returns a file descriptor (a small non-negative int) naming it.
  • AF_INET = IPv4, SOCK_STREAM = TCP-style stream, SOCK_DGRAM = UDP-style datagram; proto = 0 picks the default protocol for the family/type pair.
  • Failure is signalled by -1 with errno set — always test fd < 0 and read errno immediately.
  • A fresh socket is unbound and peerless: you still need bind/listen/accept (server) or connect (client) before data flows.
  • A socket is a file descriptor: close() it on every path, initialize FD vars to -1, and never use or double-close a stale FD.

Practice with these exercises