Linux System Programming · beginner · ~10 min

pthread_create() and pthread_join()

- By the end you can start a new thread with pthread_create() and pass it a typed argument safely. - By the end you can wait for a thread with pthread_join() and collect its return value. - By the end you can explain why every thread must be either joined or detached, and what leaks if you do neither. - By the end you can check pthread errors correctly (return value, not errno) and report them with strerror(). - By the end you can reason about the non-deterministic order in which concurrent threads run.

Overview

You already met threads conceptually in threads-intro: a thread is an independent flow of execution that shares one process's address space (its heap, globals, and file descriptors) with every other thread. This lesson turns that idea into two concrete calls — the two you will use in almost every threaded program you ever write: pthread_create() to spawn a worker, and pthread_join() to wait for it and harvest its result.

Think of pthread_create/pthread_join as the threading equivalent of fork/wait you may have seen with processes — except threads share memory instead of copying it, which makes them faster to start and communicate through, but also means a mistake in one thread can corrupt another. We keep this first lesson deliberately simple: create, do work, join, read the result. Locking and coordination between running threads come later; here every thread writes to its own slice of memory so we never need a mutex.

Why it matters

Threads are how real programs stay responsive and use more than one CPU core: a web server handles many clients at once, a game runs physics while rendering, a build tool compiles files in parallel. pthread_create/pthread_join are the POSIX foundation those systems are built on. Getting them wrong is a classic source of security and reliability bugs — a thread that reads a stack variable after it has gone out of scope, a main that returns while workers are still touching freed memory, or unjoined threads that pile up until the process exhausts its handles. Because threads share one address space, a single out-of-bounds write in a worker can silently smash data another thread depends on, so disciplined create/join hygiene is a defensive habit, not just tidiness.

Core concepts

A thread is a flow of execution, not a copy of the program

When you call pthread_create, the kernel schedules a new stack and instruction pointer that begins running inside your existing process. It does not copy your memory the way fork copies a process. Every thread — the original one running main (often called the main thread) and each worker — sees the same heap, the same globals, and the same open file descriptors.

            Process (one address space)
  +-------------------------------------------------+
  |  heap  |  globals  |  file descriptor table     |   <- shared by all threads
  +-------------------------------------------------+
     ^              ^                 ^
     |              |                 |
  +--+----+     +---+---+         +---+---+
  | main  |     |worker1|         |worker2|   <- each has its OWN stack + registers
  | stack |     | stack |         | stack |
  +-------+     +-------+         +-------+

The practical consequence: a pointer created in one thread is valid in another (same address space), but a stack variable belongs to the thread whose stack it lives on and vanishes when that function returns. That distinction drives most of the mistakes later.

pthread_create: launch and return immediately

int pthread_create(pthread_t *thread, const pthread_attr_t *attr,
                   void *(*start_routine)(void *), void *arg);

You hand it four things: a place to store the new thread's handle (thread), attributes (NULL = defaults, which is what you want at first), the function the thread should run (start_routine), and a single void * argument (arg) that gets passed to that function. The call returns right away — it does not wait for the work to happen. From that instant, two flows of execution run concurrently: your caller keeps going on the next line, and the new thread starts at the top of start_routine.

Because both run at once and the OS scheduler decides who gets the CPU when, you cannot predict the order in which threads print or finish. Running the demo below several times will show worker lines in different orders. Any program that assumes a particular order without synchronising is buggy.

pthread_join: wait and collect the result

int pthread_join(pthread_t thread, void **retval);

pthread_join blocks — pauses the calling thread — until thread finishes. When the target thread returns (or calls pthread_exit), join wakes up. If you pass a non-NULL retval, the void * that the thread returned is stored into *retval. Join also reaps the thread: it releases the bookkeeping and stack the system was holding for it. This is the moment it becomes safe to read anything the thread promised to produce — before join, the thread might still be mid-write.

Knowledge check: After pthread_create returns 0, is the worker's result ready to read?

No. pthread_create only starts the thread and returns immediately; the worker may not have run a single instruction yet. The result is guaranteed ready only after a successful pthread_join on that thread (or after some other synchronisation you set up). Reading it before joining is a data race.

Every thread must be joined or detached

A thread you create holds resources (its stack, a kernel task, and a small bookkeeping record). Those are freed in exactly one of two ways:

You do... Then... Can you get its return value?
pthread_join(t, &ret) You block until it ends, then reap it Yes, into ret
pthread_detach(t) It cleans itself up automatically when it ends No
Neither It stays joinable — a zombie whose resources leak until the process exits No

Join when you need the result or need to know it finished; detach a fire-and-forget background thread you never wait on. Doing neither is a leak that, in a long-running server spawning threads in a loop, eventually exhausts memory or the thread limit.

Errors: check the return value, not errno

The pthreads API is unusual. Most system calls return -1 and set the global errno. The pthread functions instead return the error number directly and leave errno untouched. So pthread_create returns 0 on success or a positive error code (like EAGAIN if you've hit the thread limit) on failure. Convert that code to text with strerror(rc) — never inspect errno for these calls.

Syntax notes

int pthread_create(pthread_t *thread, const pthread_attr_t *attr, void *(*start_routine)(void *), void *arg);

  • thread — out-parameter; on success the opaque handle of the new thread is written here. Pass &mytid.
  • attr — thread attributes (stack size, detach state, scheduling). Pass NULL for sensible defaults.
  • start_routine — the function the thread runs. It must have the signature void *(*)(void *): takes one void *, returns a void *.
  • arg — the single argument delivered to start_routine. Package multiple values in a struct and pass its address. The pointee must outlive the thread's use of it.
  • Returns 0 on success, or a positive errno-style code on failure. Does not set errno.

int pthread_join(pthread_t thread, void **retval);

  • thread — the handle returned by a successful pthread_create.
  • retval — out-parameter; if non-NULL, receives the void * the thread returned (or PTHREAD_CANCELED). Pass NULL if you don't care about the return value.
  • Blocks until thread terminates, then reaps it. Returns 0 on success, or an errno-style code (e.g. ESRCH for no such thread, EINVAL if not joinable, EDEADLK if you try to join yourself).
  • Joining an already-joined or detached thread is undefined behaviour — join each thread exactly once.

int pthread_detach(pthread_t thread); — marks a thread so its resources are reclaimed automatically at exit; you then must not join it.

The start routine returns a void *. Return NULL if there's nothing to hand back. Never return a pointer to one of the routine's own local (stack) variables — that memory is gone once the routine returns. Return a pointer into the heap or into the caller-provided argument struct instead.

Compile and link with the pthreads library: cc -std=c11 -Wall -Wextra prog.c -o prog -lpthread (on some systems -pthread both defines the right macros and links the library).

Lesson

Starting a thread

A thread is a separate flow of execution inside the same process. Use pthread_create to start one:

pthread_create(pthread_t *t, attr, fn, arg);

It launches a new thread that runs fn(arg). The call returns right away. The new thread then runs concurrently with the caller, so both can make progress at the same time.

Waiting for a thread

Use pthread_join to wait for a thread to complete:

pthread_join(t, &retval);

This call blocks (pauses the caller) until thread t finishes. When the thread is done, its return value (the pointer it returned) is stored in *retval.

Every thread you create must be handled in one of two ways:

  • Join it with pthread_join, or
  • Detach it with pthread_detach.

If you do neither, the thread's resources leak until the process exits.

Checking for errors

pthread_create returns 0 on success, or an error number (an errno value) on failure.

This is different from most system calls. It does not set the global errno variable. Always check the return value directly.

Code examples

#include <pthread.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>

/* A tiny per-thread job: sum the integers in [lo, hi). */
struct job {
    int  id;
    long lo, hi;
    long result;   /* filled in by the thread, read after join */
};

static void *worker(void *arg)
{
    struct job *j = arg;              /* recover our typed argument */
    long sum = 0;
    for (long i = j->lo; i < j->hi; i++)
        sum += i;
    j->result = sum;                 /* publish result into shared struct */
    printf("thread %d summed [%ld,%ld) = %ld\n", j->id, j->lo, j->hi, sum);
    return &j->result;               /* also return it via the void* channel */
}

int main(void)
{
    enum { N = 4 };
    pthread_t   tid[N];
    struct job  jobs[N];
    const long  span = 1000;

    /* Spawn N worker threads, each handling a disjoint slice. */
    for (int i = 0; i < N; i++) {
        jobs[i] = (struct job){ .id = i,
                                .lo = (long)i * span,
                                .hi = (long)(i + 1) * span,
                                .result = 0 };
        int rc = pthread_create(&tid[i], NULL, worker, &jobs[i]);
        if (rc != 0) {
            fprintf(stderr, "pthread_create: %s\n", strerror(rc));
            return 1;
        }
    }

    /* Join every thread; only after join is jobs[i].result safe to read. */
    long total = 0;
    for (int i = 0; i < N; i++) {
        void *ret = NULL;
        int rc = pthread_join(tid[i], &ret);
        if (rc != 0) {
            fprintf(stderr, "pthread_join: %s\n", strerror(rc));
            return 1;
        }
        long via_return = *(long *)ret;   /* same as jobs[i].result */
        total += via_return;
    }

    printf("grand total of [0,%ld) = %ld\n", (long)N * span, total);
    return 0;
}

Line by line

  • struct job — the argument-and-result package. Passing four separate values through one void * is impossible, so we bundle them. Crucially, jobs[] lives in main's stack and stays alive for the whole run, so each thread's arg points at memory that outlives the thread — no dangling pointer.
  • worker(void *arg) — has the mandatory void *(*)(void *) signature. Line one casts arg back to struct job * (an implicit conversion from void * in C). It sums its private slice into a local sum, then writes it into j->result.
  • return &j->result; — we hand the result back two ways for teaching: written into the shared struct and returned via the thread's void * channel. Returning &j->result is safe because it points into main's long-lived jobs[], not into worker's own stack.
  • The create loop — fills each jobs[i] with a disjoint range [i*span, (i+1)*span) so no two threads touch the same memory, then spawns the thread. We pass &jobs[i] — a distinct address per thread. Passing &i here would be a classic bug (all threads would share the one changing i). We check rc and print with strerror(rc), not errno.
  • The join loop — pthread_join(tid[i], &ret) blocks until worker i is done and stores its returned pointer in ret. Only now is it safe to dereference: *(long *)ret reads the very long the worker wrote. We accumulate into total.
  • Non-determinism — the four "thread N summed..." lines may print in any order between runs, because the scheduler interleaves the running threads. The final grand total is always the same (7998000) because each thread owns disjoint data and we read only after joining.

Common mistakes

1. Passing the address of the loop variable to every thread.

for (int i = 0; i < N; i++)
    pthread_create(&tid[i], NULL, worker, &i);   // WRONG

Why it breaks: all threads share the single i, which keeps changing (and is destroyed when the loop ends). Threads race to read it and see garbage or the same value. Fix: give each thread its own storage.

struct job jobs[N];
for (int i = 0; i < N; i++) { jobs[i].id = i; pthread_create(&tid[i], NULL, worker, &jobs[i]); }

2. Returning a pointer to a local variable.

void *worker(void *arg) { int sum = compute(); return &sum; }   // WRONG

Why it breaks: sum lives on the worker's stack, which is torn down when the routine returns; the caller's pthread_join receives a dangling pointer. Fix: return a heap pointer (malloc'd, freed after join) or write into caller-owned memory.

void *worker(void *arg) { long *r = malloc(sizeof *r); *r = compute(); return r; }  // free after join

3. Letting main return (or reading results) without joining.

pthread_create(&t, NULL, worker, &job);
printf("%ld\n", job.result);   // WRONG: worker may not have run yet
return 0;                       // WRONG: kills the still-running thread

Why it breaks: a data race on job.result, and return/exit from main terminates the whole process, cutting live threads off mid-work. Fix: pthread_join(t, NULL); before reading and before returning.

4. Checking errno instead of the return value.

pthread_create(&t, NULL, worker, &job);
if (errno != 0) { ... }         // WRONG: pthread calls don't set errno

Why it breaks: pthread_create returns the error code and leaves errno alone, so this test is meaningless. Fix: int rc = pthread_create(...); if (rc != 0) fprintf(stderr, "%s\n", strerror(rc));

Debugging tips

  • Compile with warnings and the thread sanitizer. -Wall -Wextra catches signature mismatches; cc -fsanitize=thread (ThreadSanitizer) instruments the binary to detect data races at runtime — it will point at the exact line where two threads touch the same memory without synchronisation. Invaluable once you move beyond disjoint-data programs.
  • Valgrind. valgrind --tool=helgrind ./prog finds lock-order and race errors; plain valgrind ./prog (memcheck) finds leaks, including threads you forgot to join or detach ("possibly lost" blocks tied to thread stacks).
  • gdb. Run under gdb ./prog; info threads lists all threads, thread N switches to one, and bt shows that thread's backtrace. Great for seeing which thread is blocked in pthread_join and which is still running.
  • strace. strace -f ./prog follows threads (-f) and shows the underlying clone() calls that create them and the futex calls used to block/wake — useful to confirm a thread actually started or to see a join waiting.
  • printf debugging. Print a unique id at thread entry and exit. If you never see an exit line, that thread hung or was killed early. Remember prints from different threads can interleave, so include the id.

Memory safety

  • Argument lifetime. The pointee you pass as arg must stay valid until the thread is done using it. Point at heap memory, a global, or a stack frame (like main's) that outlives the thread — never at a short-lived local that will be reused or destroyed.
  • Return-value lifetime. The void * a thread returns must not point at the thread's own stack (destroyed on return). Use heap or caller-owned storage. If it's heap, the joiner owns it and must free it.
  • Read-after-join, not read-during-run. Reading a value another thread is still writing is a data race — undefined behaviour, and on multicore CPUs you may read a torn or stale value even for a single long. This demo is safe only because each thread owns disjoint data and results are read strictly after pthread_join. The moment two threads share writable data you need a mutex or atomics (later lessons).
  • One join per thread. Joining a thread twice, or joining a detached thread, is undefined behaviour. Track which handles you've already joined.
  • No implicit ordering. Do not assume thread A runs before thread B just because you created A first. Without synchronisation there is no ordering guarantee, and code that "works" by luck on your machine can fail under a different scheduler or load.

Real-world uses

  • Servers and services. Thread-per-request or worker-pool designs use pthread_create to handle clients concurrently and pthread_join during graceful shutdown to let in-flight work finish. Nginx, databases, and language runtimes all build on these primitives.
  • Parallel computation. Splitting a large array, image, or dataset into slices and summing/transforming each on its own thread — exactly the shape of this lesson's demo — is how CPU-bound work scales across cores. Higher-level frameworks (OpenMP, thread pools) sit on top of pthreads.
  • Background tasks. Long operations (I/O, compression, a network fetch) run on a detached thread so a UI or main loop stays responsive; you detach because you never need to join back.
  • Best practice. Prefer a fixed pool of long-lived worker threads over creating one per tiny task (creation has real cost). Always pair creation with a definite join-or-detach decision. Keep shared mutable state minimal; when threads must share, protect it. Check every return code and fail loudly with strerror.

Practice tasks

  1. Single worker, single result. Write a program that creates one thread running a worker that computes 10! (factorial), returns the result via a heap long *, and have main join, print the value, and free it. Confirm no leak with valgrind.

  2. Argument struct. Pass a struct { const char *name; int times; } to a thread that prints name times times, then joins. Prove you pass a distinct struct per thread by launching three threads with different names and counts.

  3. Parallel sum with a twist. Extend the lesson demo to N=8 threads and have each thread also report how many numbers it summed; verify the grand total matches the closed-form N*span*(N*span-1)/2.

  4. Error handling drill. Deliberately trigger a pthread_create failure (e.g. spawn threads in an unbounded loop, or set a tiny stack via pthread_attr_setstacksize to an invalid size) and print the failing error with strerror(rc). Show that errno is unchanged.

  5. Detach vs join. Write two versions of a background "ticker" thread that prints once and exits: one you pthread_join, one you pthread_detach. Run both under valgrind and explain why the un-joined-but-detached version does not leak while a never-handled thread would.

Summary

  • pthread_create(&t, NULL, fn, arg) starts a new thread that runs fn(arg) concurrently and returns immediately — the result is not ready yet.
  • Threads share the process's heap, globals, and file descriptors but each has its own stack; a pointer into another thread's stack is a trap.
  • pthread_join(t, &ret) blocks until t finishes, reaps it, and hands back its returned void *. Read a thread's results only after joining.
  • Every thread must be joined (to wait/collect) or detached (fire-and-forget); neither = leaked resources.
  • Give each thread its own argument storage; never return a pointer to a local variable.
  • pthread functions return an errno-style code and do not set errno — check the return value and print it with strerror.
  • Thread run order is non-deterministic; never rely on it without synchronisation.
  • Compile with -lpthread (or -pthread); debug races with ThreadSanitizer/helgrind.

Practice with these exercises