Linux System Programming · intermediate · ~25 min
- By the end you can spawn threads with pthread_create and collect their results with pthread_join. - By the end you can identify a critical section and protect shared state with a pthread_mutex_t so updates don't corrupt each other. - By the end you can coordinate threads with a condition variable using the mandatory lock + while-loop wait pattern. - By the end you can build a working bounded producer/consumer queue from these three primitives. - By the end you can name the classic threading bugs — data races, lost wakeups, deadlock — and explain how each is prevented.
You already write functions and pass data around with pointers. A thread is simply one of your functions running as an independent flow of execution, handed its argument as a void * (which is exactly the pointer generality you learned earlier). Everything new here is a consequence of one fact: threads in a process share the same address space — the same heap, the same globals, the same file descriptors.
That shared memory is the whole point (threads communicate for free by touching the same variables) and also the whole danger (two threads touching the same variable at the same time is undefined behaviour). This lesson gives you the three POSIX primitives that turn that danger back into something correct: threads (create/join), the mutex that serialises access to shared data, and the condition variable that lets one thread sleep until another says "now".
Almost every server, database driver, and GUI toolkit written in C is multi-threaded — a worker-pool web server, redis's background I/O, an audio callback running alongside your main loop. Get the locking wrong and you don't get a clean crash; you get a data race, which is undefined behaviour that usually works in testing and corrupts memory in production under load. Concurrency bugs are also a security surface: a shared "check the permission, then use the resource" sequence with no lock is a TOCTOU (time-of-check-to-time-of-use) race, and real CVEs (privilege escalations, double-frees) come from exactly that pattern. Knowing how to make shared state safe is not optional polish — it is the difference between a program that is correct and one that merely hasn't failed yet.
When you fork(), the child gets a copy of memory. When you create a thread, there is no copy — the new thread runs inside the same process and sees the same bytes.
PROCESS (one address space)
+-------------------------------------------------+
| heap / globals (SHARED by all threads) |
| int counter = 0; queue_t q; |
+-------------------------------------------------+
^ ^ ^
| | |
+-----+ +-----+ +-----+
| T0 | | T1 | | T2 | <- each thread
| own | | own | | own | has its OWN
|stack| |stack| |stack| stack + registers
+-----+ +-----+ +-----+
Each thread has a private stack (its locals and call frames) and private registers, including its own program counter. Everything else — heap, globals, open files — is common ground. So a pointer to a heap object handed to another thread is valid there; a pointer to this thread's local variable is a landmine once this function returns.
pthread_create(&tid, attr, fn, arg) starts fn running concurrently, passing it arg. fn must have the signature void *(*)(void *). pthread_join(tid, &ret) blocks until that thread finishes and hands you back whatever it returned. Join also reclaims the thread's bookkeeping — a joinable thread you never join is a resource leak, the thread equivalent of a zombie.
A subtle but critical detail: pthreads functions return 0 on success and a positive error number on failure. They do not set errno the way most syscalls do. Always check the return value.
counter++ looks atomic. It isn't. It compiles to roughly load, add one, store. Interleave two threads:
T0: load counter (0)
T1: load counter (0)
T0: add 1 -> 1
T1: add 1 -> 1
T0: store 1
T1: store 1 result: 1, not 2 -- one increment vanished
Two threads accessing the same memory where at least one is a write, with no synchronisation, is a data race: undefined behaviour in C. Not "a wrong number" — UB, meaning the compiler is free to do anything. A pthread_mutex_t fixes this by making the read-modify-write a critical section that only one thread can be inside at a time.
pthread_mutex_lock(&m);
counter++; /* critical section: exclusive */
pthread_mutex_unlock(&m);
Knowledge check: is volatile int counter; enough to make counter++ thread-safe?
No.
volatileonly tells the compiler not to cache the value in a register across accesses — it was designed for memory-mapped hardware and signal handlers. It gives you no atomicity (the load/add/store can still interleave) and no memory ordering between threads. Use a mutex, or a_Atomic int/atomic_fetch_add, for real thread safety.
A mutex answers "who may touch the data." It does not answer "wait until the data looks a certain way" (e.g. a consumer waiting for the queue to be non-empty). Busy-spinning on while (empty) {} burns a CPU. A condition variable lets a thread sleep, releasing the lock while it sleeps, and be woken when another thread signals.
The non-negotiable pattern:
pthread_mutex_lock(&m);
while (!condition_is_true) /* WHILE, never if */
pthread_cond_wait(&cond, &m); /* atomically unlock + sleep; re-lock on wake */
/* ... use the data ... */
pthread_mutex_unlock(&m);
pthread_cond_wait does three things atomically: unlocks the mutex, puts the thread to sleep, and — when woken — re-acquires the mutex before returning. You must re-check the condition in a while loop, because (a) spurious wakeups are allowed by the standard, and (b) by the time you re-acquire the lock another thread may have already consumed the thing you were woken for. The waker holds the same mutex, changes the state, then calls pthread_cond_signal (wake one) or pthread_cond_broadcast (wake all).
If you never need a thread's return value, pthread_detach(tid) (or creating it with the detached attribute) tells the library to reclaim its resources automatically the instant it exits — no join required. A thread is either joinable (must be joined) or detached (must not be joined); doing the wrong one is a bug.
| Primitive | Type | Static init | Question it answers |
|---|---|---|---|
| Thread | pthread_t |
(created, not initialised) | "run this concurrently" |
| Mutex | pthread_mutex_t |
PTHREAD_MUTEX_INITIALIZER |
"who may touch the data" (mutual exclusion) |
| Condition variable | pthread_cond_t |
PTHREAD_COND_INITIALIZER |
"sleep until the data is ready" |
pthread_* return value.lock has exactly one matching unlock on every path (including early return/goto error exits).while loop, never an if.#include <pthread.h>
/* Start a thread. Returns 0 on success, else an error number (NOT errno).
tid -> out-param, the thread's handle (use with join/detach)
attr -> NULL for defaults (joinable, default stack size)
fn -> the thread body: void *fn(void *)
arg -> passed verbatim to fn as its void* argument */
int pthread_create(pthread_t *tid, const pthread_attr_t *attr,
void *(*fn)(void *), void *arg);
/* Wait for tid to finish; reclaims its resources.
retval -> if non-NULL, receives the void* the thread returned. */
int pthread_join(pthread_t tid, void **retval);
/* Reclaim automatically on exit; do NOT join a detached thread. */
int pthread_detach(pthread_t tid);
/* Mutex: acquire / release the lock. lock() blocks until free. */
int pthread_mutex_lock(pthread_mutex_t *m);
int pthread_mutex_unlock(pthread_mutex_t *m);
int pthread_mutex_destroy(pthread_mutex_t *m); /* only when unlocked & unused */
/* Condition variable. wait() MUST be called with m held; it atomically
releases m, sleeps, and re-acquires m before returning. */
int pthread_cond_wait(pthread_cond_t *c, pthread_mutex_t *m);
int pthread_cond_signal(pthread_cond_t *c); /* wake one waiter */
int pthread_cond_broadcast(pthread_cond_t *c); /* wake all waiters */
int pthread_cond_destroy(pthread_cond_t *c);
Compile with the threads library: cc -std=c11 -Wall -Wextra prog.c -o prog -lpthread. Statically-initialised mutexes/conds (the *_INITIALIZER macros) need no runtime init and no destroy strictly speaking, but destroying them on teardown is harmless and good hygiene.
POSIX threads (pthreads) are the standard C concurrency API on Linux.
The core operations are:
pthread_create.pthread_join.pthread_mutex_t.pthread_cond_t.#include <pthread.h>
#include <stdio.h>
#include <stdlib.h>
#define NPROD 3 /* producer threads */
#define PER_PROD 4 /* items each producer generates */
#define CAP 2 /* bounded-buffer capacity */
#define TOTAL (NPROD * PER_PROD)
/* A bounded FIFO shared by producers and one consumer. */
typedef struct {
int buf[CAP];
int count; /* items currently in buf */
int head, tail; /* ring-buffer indices */
int produced; /* how many items ever enqueued */
pthread_mutex_t lock; /* guards ALL fields above */
pthread_cond_t not_empty; /* consumer waits here */
pthread_cond_t not_full; /* producers wait here */
} queue_t;
static queue_t q = {
.lock = PTHREAD_MUTEX_INITIALIZER,
.not_empty = PTHREAD_COND_INITIALIZER,
.not_full = PTHREAD_COND_INITIALIZER,
};
static void die(const char *what, int err) {
fprintf(stderr, "%s: %d\n", what, err);
exit(1);
}
static void *producer(void *arg) {
long id = (long)arg;
for (int i = 0; i < PER_PROD; i++) {
int item = (int)(id * 100 + i);
pthread_mutex_lock(&q.lock);
while (q.count == CAP) /* buffer full: wait */
pthread_cond_wait(&q.not_full, &q.lock);
q.buf[q.tail] = item;
q.tail = (q.tail + 1) % CAP;
q.count++;
q.produced++;
pthread_cond_signal(&q.not_empty); /* wake the consumer */
pthread_mutex_unlock(&q.lock);
}
return NULL;
}
int main(void) {
pthread_t prod[NPROD];
for (long i = 0; i < NPROD; i++) {
int rc = pthread_create(&prod[i], NULL, producer, (void *)i);
if (rc) die("pthread_create", rc);
}
/* main thread is the consumer; it pops exactly TOTAL items */
long sum = 0;
for (int got = 0; got < TOTAL; got++) {
pthread_mutex_lock(&q.lock);
while (q.count == 0) /* buffer empty: wait */
pthread_cond_wait(&q.not_empty, &q.lock);
int item = q.buf[q.head];
q.head = (q.head + 1) % CAP;
q.count--;
pthread_cond_signal(&q.not_full); /* wake a producer */
pthread_mutex_unlock(&q.lock);
sum += item;
}
for (int i = 0; i < NPROD; i++) {
int rc = pthread_join(prod[i], NULL);
if (rc) die("pthread_join", rc);
}
pthread_mutex_destroy(&q.lock);
pthread_cond_destroy(&q.not_empty);
pthread_cond_destroy(&q.not_full);
printf("consumed %d items, produced=%d, checksum=%ld\n",
TOTAL, q.produced, sum);
return 0;
}
Build and run:
$ cc -std=c11 -Wall -Wextra prog.c -o prog -lpthread && ./prog
consumed 12 items, produced=12, checksum=1218
typedef struct { ... } queue_t; — one struct bundles the shared state and the primitives that guard it. Keeping the lock next to the data it protects is a habit worth forming; it documents exactly what q.lock covers (every other field).buf, count, head, tail — a classic ring buffer. head is where the consumer reads, tail is where producers write, and count tells us how full it is. All three are shared, so all three are touched only while holding lock.static queue_t q = { .lock = PTHREAD_MUTEX_INITIALIZER, ... }; — designated initialisers set the mutex and both condition variables to their static-init values at program start. No runtime pthread_mutex_init needed.die(...) — pthreads calls return an error number, not -1/errno, so we print the returned code directly. Every create/join is checked through this.producer — the thread body. Its void *arg is cast back to the long id we passed in. Note we pass the integer by value through the pointer ((void *)i), never a pointer to a loop variable — that variable would change under us.pthread_mutex_lock(&q.lock); — enter the critical section. From here until unlock, this thread has exclusive access to every q field.while (q.count == CAP) pthread_cond_wait(&q.not_full, &q.lock); — if the buffer is full, sleep on not_full, atomically releasing the lock so the consumer can drain it. On wake we re-hold the lock and re-test count (the while, not if, handles spurious wakeups and races).q.buf[q.tail] = item; q.tail = (q.tail+1)%CAP; q.count++; — the actual enqueue, safe because we hold the lock.pthread_cond_signal(&q.not_empty); — tell a possibly-sleeping consumer that data is available. We still hold the lock here; the consumer won't actually proceed until we unlock.pthread_mutex_unlock(&q.lock); — leave the critical section. Exactly one unlock for the one lock, on the only path.main — the mirror image: wait on not_empty while count == 0, pop from head, signal not_full. main itself is a thread, so using it as the consumer is free.pthread_join loop — wait for all producers to finish and reclaim them. By construction the consumer has already taken all TOTAL items, so the producers are done.pthread_*_destroy — teardown once no thread can touch the primitives.printf — produced == 12 and a stable checksum every run prove nothing was lost or double-counted despite three producers racing to write.1. Reading/writing shared state without the lock
/* WRONG */
if (q.count > 0) { int x = q.buf[q.head]; q.head++; q.count--; }
Another thread can change count between the check and the pop — a data race and a TOCTOU bug. Undefined behaviour, corrupted indices.
/* FIXED */
pthread_mutex_lock(&q.lock);
while (q.count == 0) pthread_cond_wait(&q.not_empty, &q.lock);
int x = q.buf[q.head]; q.head = (q.head+1)%CAP; q.count--;
pthread_mutex_unlock(&q.lock);
2. Waiting with if instead of while
/* WRONG */
if (q.count == 0) pthread_cond_wait(&q.not_empty, &q.lock);
int x = q.buf[q.head]; /* may run with count still 0! */
Spurious wakeups are legal, and another consumer may grab the item first. You proceed on a false condition.
/* FIXED */
while (q.count == 0) pthread_cond_wait(&q.not_empty, &q.lock);
3. Passing the address of a loop variable to every thread
/* WRONG */
for (int i = 0; i < N; i++)
pthread_create(&t[i], NULL, worker, &i); /* all share one i */
All threads read the same i, which is mutating (and may be gone). They see garbage or duplicates.
/* FIXED: pass the value, or a per-thread slot */
for (long i = 0; i < N; i++)
pthread_create(&t[i], NULL, worker, (void *)i); /* by value */
4. Forgetting to join (or joining a detached thread)
/* WRONG */
pthread_create(&t, NULL, worker, NULL);
/* main returns; t leaked, or its result never collected */
A joinable thread you never join leaks resources; if main returns, the process may exit before the thread runs.
/* FIXED */
pthread_join(t, NULL); /* wait + reclaim */
/* or, if you truly don't care about it: */
pthread_detach(t);
cc -std=c11 -fsanitize=thread -g prog.c -lpthread. Run the program and TSan reports the exact two stack traces of any data race, even one that didn't visibly misbehave this run. Treat every report as a real bug.valgrind --tool=helgrind ./prog also finds races and lock-order violations (potential deadlocks) without recompiling.gdb -p <pid> (or run under gdb), then info threads to list all threads and thread N + bt to see where each is stuck. Two threads each blocked in pthread_mutex_lock on the lock the other holds = deadlock (fix lock ordering). A thread stuck forever in pthread_cond_wait while the state is actually ready = a lost wakeup, usually from signalling before waiting or from using if instead of while.sched_yield() or micro-sleeps between the sub-steps of a critical section in a debug build to widen the window and make the race show up reliably.printf itself is a synchronisation point (it locks stdout), so adding it can hide the very race you're chasing. Prefer a sanitizer.volatile is not a synchronisation primitive. It provides neither atomicity nor cross-thread ordering. Use pthread_mutex_t, or C11 _Atomic/<stdatomic.h> for simple counters and flags.free (or let go out of scope) an object while another thread might still touch it. Join first, or use reference counting / hand ownership across cleanly.pthread_mutex_destroy a mutex that is unlocked and that no thread will use again.pthread_once for one-time init, read/write locks for read-heavy data) over hand-rolling; and always test under ThreadSanitizer in CI. When a simple atomic counter suffices, use <stdatomic.h> instead of a mutex.Write a program with two threads that each increment a shared int one million times under a single mutex, then join and print the total. Confirm it is exactly 2,000,000; then remove the mutex and observe it come out too low.
Replace the mutex-guarded counter from task 1 with a _Atomic long and atomic_fetch_add, and confirm you get the same correct total with no mutex.
Extend the producer/consumer demo so multiple consumer threads share the same queue. Ensure each item is consumed exactly once and the program terminates cleanly (hint: a sentinel/poison value or a 'done' flag broadcast to all consumers).
Use pthread_once with a pthread_once_t to lazily initialise a shared resource (e.g. a lookup table) exactly once, even when many threads call the getter simultaneously. Print a line inside the init function to prove it runs only once.
Deliberately construct a two-mutex deadlock: thread A locks m1 then m2, thread B locks m2 then m1. Reproduce the hang, confirm it with gdb's info threads, then fix it by enforcing a global lock ordering so both threads take m1 before m2.
pthread_create starts a void *(*)(void *) function concurrently; pthread_join waits for it and reclaims it. pthreads calls return an error number, not errno — check it.volatile does not.while loop (spurious wakeups + races), and the waker changes state under the same lock before signalling.