Linux System Programming · intermediate · ~10 min

Writing thread-safe code

- By the end you can define thread safety precisely and explain why 'works in one thread' does not imply 'works in many'. - By the end you can classify a function as reentrant, thread-safe-via-locking, or unsafe, and pick the right `_r` variant when one exists. - By the end you can spot the three ways shared state leaks into a function: globals, `static` locals, and returned pointers into shared buffers. - By the end you can make your own helpers safe by choosing pure functions, caller-owned state, `_Atomic`, or a mutex — and know which to reach for. - By the end you can recognise a data race as undefined behaviour and diagnose one with ThreadSanitizer.

Overview

You already know from mutex-basics that a mutex serialises access to shared data so only one thread touches it at a time. That skill answers how to protect a critical section. This lesson answers the prior question: which code needs protecting, and how to design functions so callers do not have to think about it. Thread safety is a property of a function's contract, not just a lock you sprinkle on afterward.

The core idea is simple: a function is thread-safe if two or more threads can call it at overlapping moments and every call still returns its documented result. The danger is always the same — hidden shared state. This lesson teaches you to find that hidden state and either eliminate it (pure/reentrant functions) or guard it (locks, atomics), building directly on the mutex you already know how to use.

Why it matters

Almost every real program is multithreaded today — web servers handle each request on a worker thread, GUIs run work off the UI thread, and even a simple logger may be written to from several threads. A single unsafe call like strtok or an unguarded count++ produces a data race, which the C standard declares undefined behaviour: the program may crash, corrupt memory, or silently return wrong numbers that only appear under load in production. In security terms, races cause TOCTOU bugs and heap corruption that attackers turn into exploits, so knowing which functions are safe to share is a defensive baseline, not a nicety.

Core concepts

What "thread-safe" actually means

Concurrency means two threads may be inside the same function at the same instant. A thread-safe function tolerates that. The enemy is shared mutable state — any memory more than one thread can read and write without coordination. Three sources hide in ordinary C code:

  Function's view of memory
  -------------------------------------------------
  arguments / local vars  -> private per call  (SAFE)
  static local variable   -> ONE copy, all calls share it   (DANGER)
  global variable         -> ONE copy, all threads share it  (DANGER)
  pointer returned into a  -> caller aliases shared memory    (DANGER)
     shared/static buffer

If a function touches only its arguments and locals, nothing is shared and it is safe for free. The moment it reaches for a static buffer or a global, multiple threads collide on one piece of memory.

The three categories of function

Category Shared state? Who locks? Examples
Reentrant none — args + locals only nobody needs to strlen, memcpy, snprintf into a caller buffer
Thread-safe via internal locking yes, but self-guarded the function, invisibly malloc/free on glibc, printf to stdout
Not thread-safe yes, unprotected static the caller must serialise, or avoid it strtok, gmtime, asctime, ctime, rand, strerror

Reentrant is the strongest guarantee: the function can even be re-entered by a signal handler mid-execution, because it holds no state between or across calls. Every reentrant function is thread-safe, but not every thread-safe function is reentrant — one that takes an internal lock is thread-safe yet not reentrant (a signal handler re-entering it could deadlock on the lock it already holds).

The _r family

The classic unsafe functions keep a hidden static buffer and hand you a pointer into it. Two threads calling gmtime both get the same pointer, and the second call overwrites what the first is still reading. The POSIX fix is a reentrant twin with an _r (reentrant) suffix that takes a caller-owned buffer:

Unsafe (static state) Reentrant twin Extra parameter
char *strtok(s, delim) strtok_r(s, delim, &saveptr) you own the save pointer
struct tm *gmtime(t) gmtime_r(t, &tm) you own the struct tm
struct tm *localtime(t) localtime_r(t, &tm) you own the struct tm
char *asctime(tm) asctime_r(tm, buf) you own the char buffer
int rand(void) rand_r(&seedp) you own the seed

Knowledge check: what does the _r suffix stand for, and why does it make the function safe to call from many threads?

_r means reentrant. The twin removes the hidden static state by requiring the caller to pass in the buffer or save-pointer that the state lives in. Because each thread passes its own storage (typically a stack variable), no two threads share memory, so there is nothing to race on. Note: _r functions are safe only if each thread supplies distinct storage — passing one shared buffer to two threads reintroduces the race.

Even an int is not safe

A tempting myth is "a global counter is fine if it's just an int." It is not. count++ is really three steps — load, add, store — and two threads can interleave so that both read the same old value and one increment is lost:

Thread A            Thread B            count
  load  count=41                          41
                     load  count=41       41
  add   ->42                               41
                     add   ->42            41
  store 42                                 42
                     store 42              42   <- lost an increment!

The program below reproduces exactly this: four threads each add 100,000, the correct total is 400,000, but the plain counter comes up short every run. Two clean fixes exist. Make the variable _Atomic and use atomic_fetch_add (the whole read-modify-write becomes one indivisible operation), or wrap the ++ in a mutex you already know how to use. Atomics are lighter for a single scalar; a mutex is right when you must update several fields together as one consistent unit.

Designing your own thread-safe functions

Four rules, in order of preference:

  1. Prefer pure functions — output depends only on input, no side effects. Automatically safe and reentrant.
  2. Let the caller own the state. If you need scratch space or persistent state, take it as a parameter (a char *buf, size_t cap, or a context struct). Document that the caller must not share one instance across threads without their own lock.
  3. Make single scalars _Atomic when a lone counter or flag must be global.
  4. Guard everything else with a mutex — especially when several variables must change together and stay mutually consistent.

Syntax notes

#include <stdatomic.h>
_Atomic long counter = 0;              // whole-object atomic; RMW is indivisible
long old = atomic_fetch_add(&counter, 1); // atomically counter+=1, returns prior value
long now = atomic_load(&counter);      // atomic read (never torn)
atomic_store(&counter, 0);             // atomic write

#include <pthread.h>
// PTHREAD_MUTEX_INITIALIZER: static init, no destroy needed for a file-scope mutex
pthread_mutex_t m = PTHREAD_MUTEX_INITIALIZER;
int pthread_mutex_lock(pthread_mutex_t *m);   // blocks until held; 0 on success
int pthread_mutex_unlock(pthread_mutex_t *m); // must be called by the locking thread

// Reentrant twins take caller-owned storage:
char *strtok_r(char *str, const char *delim, char **saveptr); // saveptr is yours
struct tm *gmtime_r(const time_t *t, struct tm *result);      // result is yours; returns it or NULL
int rand_r(unsigned int *seedp);                              // seed is yours

// snprintf is reentrant when the destination buffer is caller-owned:
int snprintf(char *buf, size_t cap, const char *fmt, ...);    // never writes past cap; NUL-terminates

Error conventions: pthread functions return 0 on success and a positive errno on failure (they do not set errno). gmtime_r returns NULL on failure. Ownership: for every _r function, the storage you pass in is yours to allocate and free/scope — the function keeps no reference after it returns.

Lesson

What "thread-safe" means

A thread-safe function can be called from several threads at the same time and still produce the documented result.

In multithreaded code, concurrently means two or more threads may run the same function at overlapping moments. A thread-safe function behaves correctly even then.

Three categories

Functions fall into three rough groups.

1. Reentrant

A reentrant function uses only its arguments and local variables. It shares no state at all.

  • Example: strlen(const char *s).
  • These are automatically thread-safe.

2. Thread-safe with internal locking

These functions do have shared state, but they take a lock internally to protect it. The caller does not need to do anything extra.

  • Example: malloc on glibc.

3. Not thread-safe

These functions use static state (data that persists between calls and is shared by all callers). The caller must serialize access to them.

  • Classic examples: strtok, gmtime, asctime, rand.
  • Each has a reentrant _r variant you can use instead: strtok_r, gmtime_r, rand_r.

Writing your own helpers

When you write your own functions, follow these rules.

  • Prefer pure functions. Avoid globals and static buffers.
  • Let the caller own any state. If you need state, put it in a struct the caller passes in, and document that the caller must serialize access to it.
  • Protect any globals you must use. Wrap them in a mutex, or make them _Atomic.

Code examples

#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <stdatomic.h>
#include <pthread.h>

#define THREADS 4
#define BUMPS   100000

/* --- UNSAFE: shared static buffer, returned to every caller --- */
static char *to_hex_unsafe(unsigned x) {
    static char buf[16];               /* one buffer for ALL threads */
    snprintf(buf, sizeof buf, "%x", x);
    return buf;                         /* pointer aliases across threads */
}

/* --- SAFE (reentrant): caller owns the buffer --- */
static char *to_hex_r(unsigned x, char *buf, size_t cap) {
    snprintf(buf, cap, "%x", x);
    return buf;
}

/* --- Three counters bumped concurrently --- */
static long           plain_counter  = 0;  /* racy: read-modify-write */
static _Atomic long   atomic_counter = 0;  /* safe: atomic RMW        */
static long           locked_counter = 0;  /* safe: mutex-protected   */
static pthread_mutex_t lock = PTHREAD_MUTEX_INITIALIZER;

static void *worker(void *arg) {
    (void)arg;
    char local[16];                        /* per-thread stack buffer */
    for (int i = 0; i < BUMPS; i++) {
        plain_counter++;                   /* DATA RACE (undefined behaviour) */
        atomic_fetch_add(&atomic_counter, 1);
        pthread_mutex_lock(&lock);
        locked_counter++;
        pthread_mutex_unlock(&lock);
    }
    /* The reentrant helper never collides: each thread uses its own buffer */
    to_hex_r(0xC0FFEE, local, sizeof local);
    if (strcmp(local, "c0ffee") != 0) {
        fprintf(stderr, "reentrant helper produced wrong output!\n");
        exit(1);
    }
    return NULL;
}

int main(void) {
    printf("one-shot unsafe helper: 255 -> %s\n", to_hex_unsafe(255));

    pthread_t th[THREADS];
    for (int i = 0; i < THREADS; i++)
        pthread_create(&th[i], NULL, worker, NULL);
    for (int i = 0; i < THREADS; i++)
        pthread_join(th[i], NULL);

    long expected = (long)THREADS * BUMPS;
    printf("expected total : %ld\n", expected);
    printf("plain_counter  : %ld  (racy -- usually WRONG)\n", plain_counter);
    printf("atomic_counter : %ld  (correct)\n", (long)atomic_counter);
    printf("locked_counter : %ld  (correct)\n", locked_counter);
    return 0;
}

Line by line

  • to_hex_unsafe declares static char buf[16]. static means there is exactly one buf for the entire program, shared by every caller. Returning buf hands each caller a pointer into that one buffer — safe when called once from main, disastrous if two threads call it at once. It is here as the anti-pattern to recognise.
  • to_hex_r does the same formatting but writes into a buffer the caller supplies. No shared state, so it is reentrant. This is the fix for the pattern above.
  • The three counters show the spectrum: plain_counter is an ordinary long, atomic_counter is _Atomic long, locked_counter is a plain long guarded by lock. PTHREAD_MUTEX_INITIALIZER statically initialises the mutex so no runtime pthread_mutex_init/destroy is needed for this file-scope object.
  • worker is the body every thread runs. char local[16] lives on each thread's own stack, so the reentrant call cannot collide. plain_counter++ is the deliberate bug: a load-add-store that two threads interleave, losing increments. atomic_fetch_add(&atomic_counter, 1) performs the same increment as one indivisible hardware operation. The mutex-guarded block does the load-add-store while holding lock, so only one thread is inside at a time.
  • main prints the unsafe helper once (safe because single-threaded here), spawns four workers with pthread_create, and pthread_joins them so all increments finish before we read totals. The printout shows plain_counter below the expected 400000 while the atomic and locked counters land exactly on it — the race made visible.

Common mistakes

1. Returning a pointer into a static buffer.

char *to_hex(unsigned x){ static char b[16]; snprintf(b,sizeof b,"%x",x); return b; }

Why it breaks: all threads share b; one thread overwrites it while another reads it — garbage output and UB.

char *to_hex_r(unsigned x, char *buf, size_t cap){ snprintf(buf,cap,"%x",x); return buf; }

2. Using strtok (or gmtime, asctime) in threads.

for (char *t = strtok(s, ","); t; t = strtok(NULL, ",")) { ... }  // hidden static state

Why it breaks: strtok stores its position in a hidden static; a concurrent call in another thread clobbers your parse mid-loop.

char *save; for (char *t = strtok_r(s, ",", &save); t; t = strtok_r(NULL, ",", &save)) { ... }

3. Assuming int++ is atomic.

global_count++;  // load, add, store -- two threads lose updates

Why it breaks: the read-modify-write is three steps; interleaving drops increments (a data race = UB).

_Atomic int global_count; atomic_fetch_add(&global_count, 1);  // or lock/++/unlock

4. Sharing one _r buffer across threads.

static char save;  // one saveptr shared by all threads -- back to square one

Why it breaks: _r functions are only safe if each thread owns distinct storage; a shared save-pointer reintroduces the race.

void *worker(void *a){ char *save; strtok_r(mine, ",", &save); ... }  // save is per-thread on the stack

Debugging tips

  • ThreadSanitizer is the primary tool. Compile with cc -std=c11 -fsanitize=thread -g prog.c -lpthread and run. TSan reports each data race with both stack traces (the read and the write) and the variable involved. In the example above it flags plain_counter immediately and stays silent on the atomic and locked ones.
  • Non-determinism is the tell. If a bug appears only under load, only on multi-core, or the wrong total changes run to run, suspect a race — not a logic error. Run the program in a loop (for i in $(seq 100); do ./prog; done) and watch the number wobble.
  • Valgrind's --tool=helgrind also detects races and lock-ordering problems if TSan is unavailable.
  • printf debugging lies here — adding prints changes timing and often hides the race (a "heisenbug"). Prefer the sanitizers, which detect the race regardless of whether it manifested on that run.
  • For deadlocks (over-locking), attach gdb, thread apply all bt, and look for two threads each blocked in pthread_mutex_lock on the other's mutex.

Memory safety

A data race — two threads accessing the same memory concurrently, at least one writing, with no synchronisation — is explicitly undefined behaviour in C11. That is not a warning about wrong numbers only: the compiler is permitted to assume races never happen, so it may cache a shared variable in a register, reorder your loads and stores, or optimise a spin loop into an infinite one. "It printed the right number on my machine" proves nothing. Use _Atomic for any scalar shared without a lock; a mutex establishes the happens-before ordering that makes plain reads and writes inside the critical section well-defined. The returned-static-buffer pattern is doubly dangerous: besides the race, the pointer may be overwritten before the caller finishes reading it, which is a use-after-clobber. Prefer caller-owned buffers so lifetime and ownership are explicit. Finally, remember reentrancy ≠ thread safety for signal handlers: a thread-safe function that takes a lock can deadlock if a signal interrupts it and the handler calls the same function — only truly reentrant (lock-free, state-free) functions are async-signal-safe.

Real-world uses

  • Web and RPC servers dispatch each connection to a worker thread; every helper on that path — parsers, formatters, allocators — must be thread-safe or the server corrupts responses under concurrency.
  • Logging libraries are written from many threads at once; good ones either use a mutex around the write or a lock-free ring buffer, and they never rely on strtok/ctime.
  • Time and date formatting is a classic bug source: production code uses localtime_r/gmtime_r, never the static-state originals.
  • Statistics and metrics counters (requests served, bytes sent) are textbook _Atomic use — cheap, lock-free increments.
  • Best practice: document each public function's thread-safety contract in its header comment ("reentrant", "thread-safe", or "caller must serialise"), prefer immutable/const data across threads, and keep critical sections tiny so locks are held briefly.

Practice tasks

  1. Take the to_hex_unsafe function and rewrite it as a reentrant version that the caller can safely call from many threads. Prove it by calling it from four threads with per-thread buffers and asserting each result.

  2. Write a program with a global long incremented by 8 threads a million times each. Run it several times and record how far below the expected total the plain counter lands; then fix it with _Atomic and confirm the total is exact every run.

  3. Convert a strtok parsing loop into a strtok_r loop, and demonstrate the difference by having two threads tokenise different strings at the same time — show the strtok version produces garbled tokens while strtok_r does not.

  4. Implement a thread-safe bank-account type where deposit and withdraw must update a balance and a transaction count together; explain why _Atomic on each field individually is not enough and use a mutex to keep them consistent.

  5. Build a small struct rng { unsigned seed; } context type wrapping rand_r, so each thread owns its own generator with no shared state. Then intentionally share one context between two threads, run under -fsanitize=thread, and capture the reported data race.

Summary

  • Thread-safe = correct when called concurrently; the only enemy is unprotected shared mutable state.
  • Shared state hides in three places: static locals, globals, and pointers returned into shared buffers.
  • Functions are reentrant (args+locals only, always safe), thread-safe via internal locking, or unsafe (hidden static).
  • Replace strtok/gmtime/localtime/asctime/rand with their caller-owned _r twins — and give each thread its own storage.
  • count++ is never atomic; use _Atomic + atomic_fetch_add for a lone scalar, or a mutex when several fields must stay consistent.
  • A data race is undefined behaviour — verify with -fsanitize=thread, not by eyeballing output.
  • When writing your own: prefer pure functions, then caller-owned state, then atomics, then a mutex.

Practice with these exercises