Linux System Programming · intermediate · ~10 min
- By the end you can pass one or many values into a worker thread through the single `void *` argument slot. - By the end you can explain why passing `&i` (a loop variable) to `pthread_create` is a data race, and predict what the threads actually read. - By the end you can choose correctly between the two safe idioms: pass a small integer *by value* via `intptr_t`, or hand each thread its *own* heap-allocated struct. - By the end you can decide who owns per-thread memory and when it is safe to `free` it (hint: after `pthread_join`). - By the end you can return a per-thread result without extra synchronisation by writing it into that thread's private struct.
You already know from the prerequisite (pthread-create-join) how to spawn a thread with pthread_create and wait for it with pthread_join. But every worker you launched so far started from a fixed function with no inputs. Real workers need data: which row to process, which socket to serve, which slice of an array to sum. This lesson is about the one channel pthread_create gives you to deliver that data — the final void *arg parameter — and how to use it without introducing a bug on the very first line of your first concurrent program.
The whole subject fits in one signature constraint: a thread entry function must look like void *fn(void *arg). Exactly one pointer in, one pointer out. Everything below is about packing what you need into that pointer safely, given that the new thread runs concurrently with the loop that created it.
The single most common beginner threading bug — passing the address of the loop counter — compiles without a warning, appears to work on a lightly loaded machine, and then silently corrupts results or duplicates work under real load. In a web server that means two threads handling the same connection; in a data pipeline it means chunks skipped or double-counted. Because the failure is a race, it is timing-dependent and often invisible in a debugger, which is exactly the class of bug that survives to production. Getting argument passing right is also the foundation of clean ownership: knowing who allocates and who frees per-thread data is what keeps a threaded program free of use-after-free and leaks.
pthread_create accepts a function pointer of type void *(*)(void *) and a single void *arg. When the new thread starts, the library calls your_fn(arg) passing back the exact pointer value you supplied. That is the only built-in way to get data in. There is no variadic form, no second argument. So the design question is always: how do I encode everything this thread needs into one pointer?
There are two encodings, and choosing between them is the core skill of this lesson:
intptr_t, and never dereference anything.struct. Now the lifetime and ownership of that struct become your responsibility.&i is a race, step by stepHere is the classic wrong loop:
for (int i = 0; i < N; i++)
pthread_create(&t[i], NULL, worker, &i);
pthread_create returns as soon as the thread is scheduled, not when it runs. So the parent loop keeps spinning — incrementing i — while the children may not have read &i yet. Two threads reading and one thread writing the same int with no synchronisation is a textbook data race (undefined behaviour). In practice, the loop is so much faster than thread startup that by the time the workers dereference &i, the loop has finished and i == N. Every worker prints N.
Picture the timeline:
parent: create(&i) i++ create(&i) i++ ... i==N (loop done)
| | | | |
child 0: +--- scheduled --------------------------+--> reads *(&i) == N
child 1: +--- scheduled --------------------> reads *(&i) == N
(all children share the SAME address &i, which is now N)
The address &i is a single storage cell that the parent keeps overwriting. Sharing one mutable cell is the bug; the fix is to give each thread something that will not change out from under it.
Knowledge check: your loop passes &i and you see every thread print the same number. Is adding a printf inside the loop a real fix?
No. A
printfafterpthread_createjust slows the parent down enough that some threads happen to readibefore it is incremented — it hides the race on your machine but does not remove it. The value each thread sees is still undefined and will change under different load, optimisation, or CPU count. The real fixes are passing the value by copy (intptr_t) or giving each thread its own storage.
If all a thread needs is one integer (an index, an id, an fd), you do not need an address at all. Cast the int up to intptr_t (an integer type guaranteed wide enough to hold a pointer), then to void *. Inside the thread, reverse the casts. Nothing is shared, nothing is freed:
pthread_create(&t[i], NULL, worker, (void *)(intptr_t)i); /* out */
int id = (int)(intptr_t)arg; /* in */
The void * here is not a pointer to anything — you must never dereference it. It is just a box carrying the number i.
When a thread needs more than one value, define a struct and give each thread a separate instance. Two placement choices:
| Where the struct lives | Safe? | Notes |
|---|---|---|
malloc'd, one per thread |
Yes | Thread (or joiner) frees it. The robust default. |
A distinct slot in a parent array, e.g. args[i], kept alive until join |
Yes | No malloc, but the array must outlive every thread. |
| One shared struct reused each iteration | No | Same cell overwritten — same race as &i. |
| A struct local to the creating function that returns early | No | Stack frame vanishes → use-after-free. |
The malloc-per-thread pattern is the one to reach for by default because it makes lifetime explicit and decouples the data from the parent's stack frame.
Because each thread owns a private struct, it can also write its answer back into that struct. You do not need a mutex for this: pthread_join establishes a happens-before relationship, so once join returns, all writes the joined thread made are visible to you. Read the result after join, never during. This is the cheapest way to collect per-thread output.
args[0] ┌─────────────┐ worker 0 writes result ──┐
│ id, in, out │ │ after join,
args[1] ├─────────────┤ worker 1 writes result ──┼─ parent reads
│ id, in, out │ │ every .out
args[2] └─────────────┘ worker 2 writes result ──┘
int pthread_create(pthread_t *thread, const pthread_attr_t *attr,
void *(*start_routine)(void *), void *arg);
thread — out-param; receives the new thread's handle. Must stay alive until you join.attr — NULL for defaults.start_routine — must have exactly the type void *(*)(void *).arg — the single pointer handed verbatim to start_routine. May encode a value or an address; the library never inspects it.0 on success, or a positive error number (e.g. EAGAIN). It does not set errno; use the return value directly, e.g. with strerror(rc).int pthread_join(pthread_t thread, void **retval);
thread finishes; *retval receives that thread's return value (or pass NULL to ignore it). Also a synchronisation barrier: after it returns, the joined thread's memory writes are visible.#include <stdint.h>
intptr_t /* signed integer type wide enough to round-trip a void* */
(void *)(intptr_t)x to carry an integer through a pointer, and (int)(intptr_t)arg to recover it. Never free or dereference such a void *.Ownership rule of thumb: whoever mallocs the per-thread struct must ensure exactly one free. Either the worker frees its own struct just before returning, or the joiner frees it after pthread_join — pick one, document it, never both.
A thread function always has this signature:
void *fn(void *arg);
It receives exactly one pointer. To send more than one value, pack the values into a struct and pass the struct's address.
The classic mistake is passing the address of a loop variable (the counter i):
/* WRONG — every thread reads the same `i` */
for (int i = 0; i < N; i++)
pthread_create(&tids[i], NULL, worker, &i);
Here is the problem. The new threads do not run instantly. By the time they actually read i, the loop has usually finished and i == N. So every thread reads N, not the value you intended.
int into intptr_t, then into void *, and reverse the cast inside the thread. (intptr_t is an integer type guaranteed to be wide enough to hold a pointer.) No shared variable, no race.malloc'd struct so nothing is shared or reassigned.#include <pthread.h>
#include <stdio.h>
#include <stdlib.h>
#include <stdint.h>
#include <string.h>
/* Per-thread argument: each thread gets its OWN copy of this struct. */
typedef struct {
int id; /* which worker this is */
const char *name; /* a shared read-only string */
long chunk_start; /* slice of work for this thread */
long chunk_len;
long result; /* thread writes its answer here */
} task_t;
/* The one legal thread signature: void *fn(void *). */
static void *worker(void *arg)
{
task_t *t = (task_t *)arg; /* recover the typed pointer */
long sum = 0;
for (long k = 0; k < t->chunk_len; k++)
sum += t->chunk_start + k; /* add up this thread's slice */
t->result = sum; /* publish result into our own struct */
printf("worker %d (%s): summed [%ld..%ld) = %ld\n",
t->id, t->name,
t->chunk_start, t->chunk_start + t->chunk_len, sum);
return NULL;
}
/* Demonstrate the tiny-value trick too: pass an int BY VALUE, no pointer. */
static void *ping(void *arg)
{
int my_id = (int)(intptr_t)arg; /* reverse the value cast */
printf("ping thread carrying value %d\n", my_id);
return NULL;
}
#define N 4
int main(void)
{
pthread_t tids[N];
task_t *tasks[N]; /* one heap struct per thread */
const long total = 40;
const long per = total / N;
/* ---- Method 1: heap-allocated per-thread struct ---- */
for (int i = 0; i < N; i++) {
task_t *t = malloc(sizeof *t);
if (!t) { perror("malloc"); return 1; }
t->id = i;
t->name = "sum";
t->chunk_start = (long)i * per;
t->chunk_len = per;
t->result = 0;
tasks[i] = t; /* keep the pointer so we can free later */
if (pthread_create(&tids[i], NULL, worker, t) != 0) {
perror("pthread_create");
free(t);
return 1;
}
}
long grand = 0;
for (int i = 0; i < N; i++) {
pthread_join(tids[i], NULL); /* wait, THEN read the result */
grand += tasks[i]->result; /* safe: join is a happens-before barrier */
free(tasks[i]); /* owner frees after the thread is done */
}
printf("grand total = %ld\n", grand);
/* ---- Method 2: pass a small int by value (no allocation) ---- */
pthread_t pt[N];
for (int i = 0; i < N; i++)
pthread_create(&pt[i], NULL, ping, (void *)(intptr_t)i);
for (int i = 0; i < N; i++)
pthread_join(pt[i], NULL);
return 0;
}
typedef struct { ... } task_t; — the argument bundle. It carries inputs (id, name, chunk_start, chunk_len) and an output slot (result). One pointer to this struct delivers everything a worker needs.static void *worker(void *arg) — the mandatory signature. The first thing it does is task_t *t = (task_t *)arg;, turning the opaque pointer back into a typed one. This is the mirror of what the parent passed.for (long k ...) sum += t->chunk_start + k; — each worker touches only its slice, computed from its struct. No two workers read the same field, so there is no shared mutable state and no lock needed.t->result = sum; — the worker writes its answer into its own struct. Because the parent won't read it until after join, this is race-free.ping / (int)(intptr_t)arg — the by-value path: the void * is not an address, it's a number in disguise, recovered with the reverse cast. Notice ping never dereferences arg.task_t *t = malloc(sizeof *t); — one allocation per iteration, so every thread gets a distinct struct. sizeof *t (not sizeof(task_t)) keeps the size tied to the pointer's type.tasks[i] = t; — the parent stashes each pointer so it can free them later. If it didn't, the pointers would leak.if (pthread_create(...) != 0) { ... free(t); ... } — always check the return; on failure the thread never ran, so the parent must free the struct it just allocated to avoid a leak.pthread_join(tids[i], NULL); grand += tasks[i]->result; — join first (the barrier), then read the result. Reading before join would be a race.free(tasks[i]); — the joiner owns cleanup here. Exactly one free per struct, after the thread that used it is guaranteed finished.1. Passing the address of the loop variable.
/* WRONG */
for (int i = 0; i < N; i++)
pthread_create(&t[i], NULL, run, &i);
All threads share the single cell &i, which the loop keeps incrementing; they read whatever it holds when they finally run — usually N. It's a data race (UB), not just a wrong value.
/* FIXED — value by copy */
for (int i = 0; i < N; i++)
pthread_create(&t[i], NULL, run, (void *)(intptr_t)i);
2. Pointing at a stack local that dies too soon.
/* WRONG */
void spawn(void) {
task_t a = { .id = 1 };
pthread_create(&t, NULL, run, &a);
} /* a's storage is reclaimed here; thread reads freed stack */
When spawn returns, a no longer exists — use-after-free.
/* FIXED — heap outlives the frame */
task_t *a = malloc(sizeof *a);
a->id = 1;
pthread_create(&t, NULL, run, a); /* free after join */
3. Reusing one shared struct across iterations.
/* WRONG */
task_t a;
for (int i = 0; i < N; i++) {
a.id = i; /* overwrites what earlier threads may still read */
pthread_create(&t[i], NULL, run, &a);
}
Same single cell, same race as &i.
/* FIXED — distinct storage per thread */
task_t a[N];
for (int i = 0; i < N; i++) {
a[i].id = i;
pthread_create(&t[i], NULL, run, &a[i]); /* a[] must outlive the threads */
}
4. Freeing the argument before (or twice around) join.
/* WRONG */
pthread_create(&t, NULL, run, p);
free(p); /* thread may still be reading *p */
Free only after the reader is done, and exactly once.
/* FIXED */
pthread_create(&t, NULL, run, p);
pthread_join(t, NULL);
free(p);
sched_yield() or a short usleep right after pthread_create — if the printed values change, you have a sharing bug (a correct program is insensitive to timing).valgrind --tool=helgrind ./prog) will name the exact conflicting accesses to a shared int — perfect for catching the &i pattern.valgrind --leak-check=full) catches the mirror-image ownership bugs: leaked per-thread structs (forgot to free) and use-after-free (freed before join).cc -fsanitize=address,undefined -g) flags use-after-return when you pass &stack_local from a function that returns, and UBSan flags dubious casts.info threads lists live threads; thread apply all bt shows every stack. Print the argument pointer in the worker (p arg) — if every thread shows the same address, they're sharing one cell.%p, arg) and the value inside each worker. Identical addresses across threads is the smoking gun.&i, a reused struct) to multiple threads while the parent still writes it is undefined behaviour. Give each thread private data or pass by value.void *arg pointing into a stack frame that returns before the thread reads it is a use-after-free. Heap-allocate, or guarantee the frame outlives every thread (e.g. the creating function joins before returning).free. The struct must live until the last read by the thread. Freeing before join, or double-freeing (worker frees AND joiner frees), corrupts the heap. Choose one owner.pthread_join (or another synchronisation point) is a race — the write may not be visible. Join is a happens-before barrier; rely on it.intptr_t only carries integers. A void * built with (void *)(intptr_t)i must never be dereferenced or freed; it holds a value, not an address.pthread_detach instead of joining, you cannot free the arg from the parent safely — the thread must own and free its own struct.malloc'd struct holding the fd and client address, freed by the handler when the connection closes.malloc'd task structs that idle workers pop and execute; ownership transfers from producer to worker, which frees the task when done.intptr_t only for a single small integer; always check pthread_create's return; and collect results through join rather than shared globals whenever you can, to keep the data flow lock-free and easy to reason about.Launch 5 threads, each of which prints its own index 0..4 by receiving the index by value through intptr_t. Confirm the output shows each number exactly once (in any order).
Rewrite task 1 to pass each index through a heap-allocated int instead, and free each allocation after joining. Run it under valgrind --leak-check=full and confirm zero leaks.
Deliberately introduce the &i bug, run the program 20 times, and record how the printed indices vary. Then add sched_yield() after pthread_create and observe the behaviour change — write one sentence explaining why neither version is correct.
Give each thread a struct containing a slice of an integer array [start, len]; have every thread sum its slice into a result field, then have main add the per-thread results after joining. Verify the total matches a single-threaded sum.
Convert task 4 to detached threads (pthread_detach): the parent no longer joins, so make each worker responsible for freeing its own struct. Use a separate mechanism (e.g. a counter protected by a mutex, or writing results to a pre-sized shared array indexed by id) so main can still read all results safely before exiting.
void *fn(void *arg) — one pointer in, one pointer out.&i or any single mutable cell shared across threads; the loop reassigns it before the threads read, which is a data race, and they typically all see N.(void *)(intptr_t)i and recover it with (int)(intptr_t)arg — never dereference that pointer.pthread_join, which acts as a happens-before barrier — no mutex required.pthread_create's return value, and free per-thread memory exactly once (never before join, never twice).