Linux System Programming · intermediate · ~10 min
- By the end you can define thread safety precisely and explain why 'works in one thread' does not imply 'works in many'. - By the end you can classify a function as reentrant, thread-safe-via-locking, or unsafe, and pick the right `_r` variant when one exists. - By the end you can spot the three ways shared state leaks into a function: globals, `static` locals, and returned pointers into shared buffers. - By the end you can make your own helpers safe by choosing pure functions, caller-owned state, `_Atomic`, or a mutex — and know which to reach for. - By the end you can recognise a data race as undefined behaviour and diagnose one with ThreadSanitizer.
You already know from mutex-basics that a mutex serialises access to shared data so only one thread touches it at a time. That skill answers how to protect a critical section. This lesson answers the prior question: which code needs protecting, and how to design functions so callers do not have to think about it. Thread safety is a property of a function's contract, not just a lock you sprinkle on afterward.
The core idea is simple: a function is thread-safe if two or more threads can call it at overlapping moments and every call still returns its documented result. The danger is always the same — hidden shared state. This lesson teaches you to find that hidden state and either eliminate it (pure/reentrant functions) or guard it (locks, atomics), building directly on the mutex you already know how to use.
Almost every real program is multithreaded today — web servers handle each request on a worker thread, GUIs run work off the UI thread, and even a simple logger may be written to from several threads. A single unsafe call like strtok or an unguarded count++ produces a data race, which the C standard declares undefined behaviour: the program may crash, corrupt memory, or silently return wrong numbers that only appear under load in production. In security terms, races cause TOCTOU bugs and heap corruption that attackers turn into exploits, so knowing which functions are safe to share is a defensive baseline, not a nicety.
Concurrency means two threads may be inside the same function at the same instant. A thread-safe function tolerates that. The enemy is shared mutable state — any memory more than one thread can read and write without coordination. Three sources hide in ordinary C code:
Function's view of memory
-------------------------------------------------
arguments / local vars -> private per call (SAFE)
static local variable -> ONE copy, all calls share it (DANGER)
global variable -> ONE copy, all threads share it (DANGER)
pointer returned into a -> caller aliases shared memory (DANGER)
shared/static buffer
If a function touches only its arguments and locals, nothing is shared and it is safe for free. The moment it reaches for a static buffer or a global, multiple threads collide on one piece of memory.
| Category | Shared state? | Who locks? | Examples |
|---|---|---|---|
| Reentrant | none — args + locals only | nobody needs to | strlen, memcpy, snprintf into a caller buffer |
| Thread-safe via internal locking | yes, but self-guarded | the function, invisibly | malloc/free on glibc, printf to stdout |
| Not thread-safe | yes, unprotected static |
the caller must serialise, or avoid it | strtok, gmtime, asctime, ctime, rand, strerror |
Reentrant is the strongest guarantee: the function can even be re-entered by a signal handler mid-execution, because it holds no state between or across calls. Every reentrant function is thread-safe, but not every thread-safe function is reentrant — one that takes an internal lock is thread-safe yet not reentrant (a signal handler re-entering it could deadlock on the lock it already holds).
_r familyThe classic unsafe functions keep a hidden static buffer and hand you a pointer into it. Two threads calling gmtime both get the same pointer, and the second call overwrites what the first is still reading. The POSIX fix is a reentrant twin with an _r (reentrant) suffix that takes a caller-owned buffer:
| Unsafe (static state) | Reentrant twin | Extra parameter |
|---|---|---|
char *strtok(s, delim) |
strtok_r(s, delim, &saveptr) |
you own the save pointer |
struct tm *gmtime(t) |
gmtime_r(t, &tm) |
you own the struct tm |
struct tm *localtime(t) |
localtime_r(t, &tm) |
you own the struct tm |
char *asctime(tm) |
asctime_r(tm, buf) |
you own the char buffer |
int rand(void) |
rand_r(&seedp) |
you own the seed |
Knowledge check: what does the _r suffix stand for, and why does it make the function safe to call from many threads?
_rmeans reentrant. The twin removes the hiddenstaticstate by requiring the caller to pass in the buffer or save-pointer that the state lives in. Because each thread passes its own storage (typically a stack variable), no two threads share memory, so there is nothing to race on. Note:_rfunctions are safe only if each thread supplies distinct storage — passing one shared buffer to two threads reintroduces the race.
int is not safeA tempting myth is "a global counter is fine if it's just an int." It is not. count++ is really three steps — load, add, store — and two threads can interleave so that both read the same old value and one increment is lost:
Thread A Thread B count
load count=41 41
load count=41 41
add ->42 41
add ->42 41
store 42 42
store 42 42 <- lost an increment!
The program below reproduces exactly this: four threads each add 100,000, the correct total is 400,000, but the plain counter comes up short every run. Two clean fixes exist. Make the variable _Atomic and use atomic_fetch_add (the whole read-modify-write becomes one indivisible operation), or wrap the ++ in a mutex you already know how to use. Atomics are lighter for a single scalar; a mutex is right when you must update several fields together as one consistent unit.
Four rules, in order of preference:
char *buf, size_t cap, or a context struct). Document that the caller must not share one instance across threads without their own lock._Atomic when a lone counter or flag must be global.#include <stdatomic.h>
_Atomic long counter = 0; // whole-object atomic; RMW is indivisible
long old = atomic_fetch_add(&counter, 1); // atomically counter+=1, returns prior value
long now = atomic_load(&counter); // atomic read (never torn)
atomic_store(&counter, 0); // atomic write
#include <pthread.h>
// PTHREAD_MUTEX_INITIALIZER: static init, no destroy needed for a file-scope mutex
pthread_mutex_t m = PTHREAD_MUTEX_INITIALIZER;
int pthread_mutex_lock(pthread_mutex_t *m); // blocks until held; 0 on success
int pthread_mutex_unlock(pthread_mutex_t *m); // must be called by the locking thread
// Reentrant twins take caller-owned storage:
char *strtok_r(char *str, const char *delim, char **saveptr); // saveptr is yours
struct tm *gmtime_r(const time_t *t, struct tm *result); // result is yours; returns it or NULL
int rand_r(unsigned int *seedp); // seed is yours
// snprintf is reentrant when the destination buffer is caller-owned:
int snprintf(char *buf, size_t cap, const char *fmt, ...); // never writes past cap; NUL-terminates
Error conventions: pthread functions return 0 on success and a positive errno on failure (they do not set errno). gmtime_r returns NULL on failure. Ownership: for every _r function, the storage you pass in is yours to allocate and free/scope — the function keeps no reference after it returns.
A thread-safe function can be called from several threads at the same time and still produce the documented result.
In multithreaded code, concurrently means two or more threads may run the same function at overlapping moments. A thread-safe function behaves correctly even then.
Functions fall into three rough groups.
A reentrant function uses only its arguments and local variables. It shares no state at all.
strlen(const char *s).These functions do have shared state, but they take a lock internally to protect it. The caller does not need to do anything extra.
malloc on glibc.These functions use static state (data that persists between calls and is shared by all callers). The caller must serialize access to them.
strtok, gmtime, asctime, rand._r variant you can use instead: strtok_r, gmtime_r, rand_r.When you write your own functions, follow these rules.
_Atomic.#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <stdatomic.h>
#include <pthread.h>
#define THREADS 4
#define BUMPS 100000
/* --- UNSAFE: shared static buffer, returned to every caller --- */
static char *to_hex_unsafe(unsigned x) {
static char buf[16]; /* one buffer for ALL threads */
snprintf(buf, sizeof buf, "%x", x);
return buf; /* pointer aliases across threads */
}
/* --- SAFE (reentrant): caller owns the buffer --- */
static char *to_hex_r(unsigned x, char *buf, size_t cap) {
snprintf(buf, cap, "%x", x);
return buf;
}
/* --- Three counters bumped concurrently --- */
static long plain_counter = 0; /* racy: read-modify-write */
static _Atomic long atomic_counter = 0; /* safe: atomic RMW */
static long locked_counter = 0; /* safe: mutex-protected */
static pthread_mutex_t lock = PTHREAD_MUTEX_INITIALIZER;
static void *worker(void *arg) {
(void)arg;
char local[16]; /* per-thread stack buffer */
for (int i = 0; i < BUMPS; i++) {
plain_counter++; /* DATA RACE (undefined behaviour) */
atomic_fetch_add(&atomic_counter, 1);
pthread_mutex_lock(&lock);
locked_counter++;
pthread_mutex_unlock(&lock);
}
/* The reentrant helper never collides: each thread uses its own buffer */
to_hex_r(0xC0FFEE, local, sizeof local);
if (strcmp(local, "c0ffee") != 0) {
fprintf(stderr, "reentrant helper produced wrong output!\n");
exit(1);
}
return NULL;
}
int main(void) {
printf("one-shot unsafe helper: 255 -> %s\n", to_hex_unsafe(255));
pthread_t th[THREADS];
for (int i = 0; i < THREADS; i++)
pthread_create(&th[i], NULL, worker, NULL);
for (int i = 0; i < THREADS; i++)
pthread_join(th[i], NULL);
long expected = (long)THREADS * BUMPS;
printf("expected total : %ld\n", expected);
printf("plain_counter : %ld (racy -- usually WRONG)\n", plain_counter);
printf("atomic_counter : %ld (correct)\n", (long)atomic_counter);
printf("locked_counter : %ld (correct)\n", locked_counter);
return 0;
}
to_hex_unsafe declares static char buf[16]. static means there is exactly one buf for the entire program, shared by every caller. Returning buf hands each caller a pointer into that one buffer — safe when called once from main, disastrous if two threads call it at once. It is here as the anti-pattern to recognise.to_hex_r does the same formatting but writes into a buffer the caller supplies. No shared state, so it is reentrant. This is the fix for the pattern above.plain_counter is an ordinary long, atomic_counter is _Atomic long, locked_counter is a plain long guarded by lock. PTHREAD_MUTEX_INITIALIZER statically initialises the mutex so no runtime pthread_mutex_init/destroy is needed for this file-scope object.worker is the body every thread runs. char local[16] lives on each thread's own stack, so the reentrant call cannot collide. plain_counter++ is the deliberate bug: a load-add-store that two threads interleave, losing increments. atomic_fetch_add(&atomic_counter, 1) performs the same increment as one indivisible hardware operation. The mutex-guarded block does the load-add-store while holding lock, so only one thread is inside at a time.main prints the unsafe helper once (safe because single-threaded here), spawns four workers with pthread_create, and pthread_joins them so all increments finish before we read totals. The printout shows plain_counter below the expected 400000 while the atomic and locked counters land exactly on it — the race made visible.1. Returning a pointer into a static buffer.
char *to_hex(unsigned x){ static char b[16]; snprintf(b,sizeof b,"%x",x); return b; }
Why it breaks: all threads share b; one thread overwrites it while another reads it — garbage output and UB.
char *to_hex_r(unsigned x, char *buf, size_t cap){ snprintf(buf,cap,"%x",x); return buf; }
2. Using strtok (or gmtime, asctime) in threads.
for (char *t = strtok(s, ","); t; t = strtok(NULL, ",")) { ... } // hidden static state
Why it breaks: strtok stores its position in a hidden static; a concurrent call in another thread clobbers your parse mid-loop.
char *save; for (char *t = strtok_r(s, ",", &save); t; t = strtok_r(NULL, ",", &save)) { ... }
3. Assuming int++ is atomic.
global_count++; // load, add, store -- two threads lose updates
Why it breaks: the read-modify-write is three steps; interleaving drops increments (a data race = UB).
_Atomic int global_count; atomic_fetch_add(&global_count, 1); // or lock/++/unlock
4. Sharing one _r buffer across threads.
static char save; // one saveptr shared by all threads -- back to square one
Why it breaks: _r functions are only safe if each thread owns distinct storage; a shared save-pointer reintroduces the race.
void *worker(void *a){ char *save; strtok_r(mine, ",", &save); ... } // save is per-thread on the stack
cc -std=c11 -fsanitize=thread -g prog.c -lpthread and run. TSan reports each data race with both stack traces (the read and the write) and the variable involved. In the example above it flags plain_counter immediately and stays silent on the atomic and locked ones.for i in $(seq 100); do ./prog; done) and watch the number wobble.--tool=helgrind also detects races and lock-ordering problems if TSan is unavailable.printf debugging lies here — adding prints changes timing and often hides the race (a "heisenbug"). Prefer the sanitizers, which detect the race regardless of whether it manifested on that run.gdb, thread apply all bt, and look for two threads each blocked in pthread_mutex_lock on the other's mutex.A data race — two threads accessing the same memory concurrently, at least one writing, with no synchronisation — is explicitly undefined behaviour in C11. That is not a warning about wrong numbers only: the compiler is permitted to assume races never happen, so it may cache a shared variable in a register, reorder your loads and stores, or optimise a spin loop into an infinite one. "It printed the right number on my machine" proves nothing. Use _Atomic for any scalar shared without a lock; a mutex establishes the happens-before ordering that makes plain reads and writes inside the critical section well-defined. The returned-static-buffer pattern is doubly dangerous: besides the race, the pointer may be overwritten before the caller finishes reading it, which is a use-after-clobber. Prefer caller-owned buffers so lifetime and ownership are explicit. Finally, remember reentrancy ≠ thread safety for signal handlers: a thread-safe function that takes a lock can deadlock if a signal interrupts it and the handler calls the same function — only truly reentrant (lock-free, state-free) functions are async-signal-safe.
strtok/ctime.localtime_r/gmtime_r, never the static-state originals._Atomic use — cheap, lock-free increments.Take the to_hex_unsafe function and rewrite it as a reentrant version that the caller can safely call from many threads. Prove it by calling it from four threads with per-thread buffers and asserting each result.
Write a program with a global long incremented by 8 threads a million times each. Run it several times and record how far below the expected total the plain counter lands; then fix it with _Atomic and confirm the total is exact every run.
Convert a strtok parsing loop into a strtok_r loop, and demonstrate the difference by having two threads tokenise different strings at the same time — show the strtok version produces garbled tokens while strtok_r does not.
Implement a thread-safe bank-account type where deposit and withdraw must update a balance and a transaction count together; explain why _Atomic on each field individually is not enough and use a mutex to keep them consistent.
Build a small struct rng { unsigned seed; } context type wrapping rand_r, so each thread owns its own generator with no shared state. Then intentionally share one context between two threads, run under -fsanitize=thread, and capture the reported data race.
static locals, globals, and pointers returned into shared buffers.static).strtok/gmtime/localtime/asctime/rand with their caller-owned _r twins — and give each thread its own storage.count++ is never atomic; use _Atomic + atomic_fetch_add for a lone scalar, or a mutex when several fields must stay consistent.-fsanitize=thread, not by eyeballing output.