Linux System Programming · intermediate · ~12 min

Writing safe signal handlers — async-signal-safe rules

- By the end you can name the small set of functions that are legal to call from inside a signal handler, and explain why `printf`, `malloc`, and most of libc are not. - By the end you can describe *reentrancy* and *async-signal-safety* precisely, and predict the deadlocks and corruption that follow from breaking them. - By the end you can write the standard defensive handler: set a `volatile sig_atomic_t` flag (and preserve `errno`), then do the real work back in the main loop. - By the end you can install a handler with `sigaction` and choose flags (`SA_RESTART`, mask) deliberately rather than by copy-paste. - By the end you can debug a hung or crashing handler and reason about the concurrency hazards it introduces.

Overview

You already know from signal-vs-sigaction how to install a handler — that sigaction is the portable, well-defined way to register a callback for a signal, and that it gives you control over the signal mask and flags that signal() leaves ambiguous. This lesson answers the next, harder question: once your handler runs, what is it actually allowed to do? The answer is surprisingly restrictive, and getting it wrong produces bugs that are intermittent, timing-dependent, and nearly impossible to reproduce.

A signal handler is not a normal function call. It is an asynchronous interruption: the kernel can deliver a signal at any machine instruction, freezing your normal code mid-step and running the handler on the same thread. If the code you interrupted was halfway through updating a shared data structure, and your handler touches that same structure, you have re-entered a function that assumed it would finish uninterrupted. This lesson builds the mental model for what is and isn't safe, and the one pattern that sidesteps the whole problem.

Why it matters

Almost every long-running program — servers, daemons, CLIs that need clean Ctrl-C — must handle signals, and the naive handler (printf("caught!\n"); cleanup();) is a latent bug that ships to production and then hangs one time in ten thousand under load. The failure mode is the worst kind: a deadlock inside malloc's arena lock or corrupted stdio buffers that only manifests when a signal happens to land at exactly the wrong instruction. In security-sensitive code the stakes are higher: CVE-2006-5051 in OpenSSH was an async-signal-unsafe SIGALRM handler that called into non-reentrant code, and its 2024 re-introduction ("regreSSHion", CVE-2024-6387) became a remote-code-execution race precisely because a handler touched functions that were not async-signal-safe. Knowing these rules is the difference between a robust daemon and an exploitable one.

Core concepts

1. A handler is an interruption, not a call

When you call a function normally, control flows into it and back out in a predictable order. A signal handler is different: the kernel suspends whatever your thread was doing and injects the handler on top of the current stack. The interrupted code has no idea it was paused.

  Normal thread timeline (main is inside malloc, holding the arena lock):

  main:   ... call malloc() ── lock arena ── update free-list ─┐  X  ┌─ unlock ...
                                                               │     │
                              signal delivered HERE ───────────┘     │
                                                                     │
  handler:                    on_signal() runs on same stack ...     │
                              if it calls malloc()  → tries to lock  │
                              the arena AGAIN → the lock is already   │
                              held by the very thread we interrupted  │
                              → DEADLOCK (main can never resume to     
                              release it because we are on its stack) ─┘

That picture is the whole danger in one image: the handler runs on the interrupted thread, so anything the interrupted code was in the middle of is now in an inconsistent, half-finished state.

2. Reentrancy and async-signal-safety

Two related terms:

  • Reentrant — a function that can be safely called again before a previous invocation finishes. It keeps no static/global state and takes no locks, so a second concurrent entry can't trample the first.
  • Async-signal-safe — a stronger, signal-specific guarantee: the function is safe to call from a signal handler even if it interrupted itself or any other function at an arbitrary point. POSIX publishes the exact list in signal-safety(7). A function qualifies only if the standard promises it.

Why is printf unsafe? It writes into a FILE* buffer that is shared process-wide and guarded by an internal lock. If a signal lands while printf holds that buffer lock and the handler calls printf too, you deadlock or corrupt the buffer. Why is malloc unsafe? It maintains a global heap protected by locks and invariants; re-entering it mid-update corrupts the free-list. Why is write safe? It is a thin wrapper over a single syscall with no user-space buffer or lock of its own.

Function Async-signal-safe? Why
write, read Yes Direct syscall, no libc buffer/lock
_exit, abort Yes Terminate without running atexit/flush
kill, raise, sigaction, sigprocmask Yes Simple syscalls, defined safe by POSIX
signalfd, waitpid (with care) Yes Syscall-level
printf, fprintf, puts, snprintf No stdio buffer + lock, not reentrant
malloc, free, calloc, realloc No Global heap lock/invariants
strsignal, localtime, getenv No Static internal buffers
pthread_mutex_lock No Can deadlock against interrupted holder

Rule of thumb: if a function allocates, buffers, or locks anything shared, assume it is unsafe until signal-safety(7) says otherwise.

Knowledge check: Your handler calls snprintf(buf, n, "sig=%d", sig) into a local stack buffer — no heap, no FILE*. Is that safe?

No. snprintf is not on the async-signal-safe list — the destination being a local buffer doesn't matter, because the formatting machinery inside stdio may touch locale data and shared internal state that is not reentrant. Format the number by hand with a small integer-to-ASCII loop and emit it with write, as the code example does with safe_put_int.

3. The one pattern that always works: flip a flag

Because the safe function list is so small, the durable strategy is to do almost nothing in the handler. Set a flag and return; let the main loop notice the flag and do the real, unrestricted work in normal control flow.

   ┌── signal arrives ──▶ handler: g_want_quit = 1;  (one atomic store)  return
   │
   │                     main loop:  while (!g_want_quit) { work / pause(); }
   └──────────────────────────────────▶ sees flag ▶ cleanup(), printf(), free(), exit()

The flag must have type volatile sig_atomic_t. sig_atomic_t is the only type C guarantees can be read and written in one indivisible step with respect to signal delivery, so the main loop never observes a half-written value. volatile stops the compiler from caching the flag in a register or optimizing the while (!g_want_quit) loop away — without it, the compiler is entitled to assume the flag never changes and loop forever.

4. Two more handler hygiene rules

  • Save and restore errno. A handler that calls write (or any syscall) can clobber the global errno, and the interrupted code may have been about to read it. Snapshot errno on entry and restore it on exit.
  • Beware interrupted syscalls. A blocking call like read/sleep/accept returns -1 with errno == EINTR when a signal handler runs during it. Either install with SA_RESTART (kernel auto-restarts many syscalls) or check for EINTR and retry. In the flag pattern you often want the interruption — it wakes pause()/select() so the loop can re-check the flag — so leaving SA_RESTART off is a deliberate choice.

Syntax notes

#include <signal.h>
int sigaction(int signum, const struct sigaction *act, struct sigaction *oldact);
//   signum : which signal (SIGTERM, SIGINT, ...). Cannot catch SIGKILL/SIGSTOP.
//   act    : new disposition; if non-NULL, installed for signum.
//   oldact : if non-NULL, receives the previous disposition (for restore).
//   returns 0 on success, -1 + errno on failure.

struct sigaction {
    void     (*sa_handler)(int);      // simple handler: gets the signal number
    void     (*sa_sigaction)(int, siginfo_t *, void *); // used if SA_SIGINFO set
    sigset_t   sa_mask;               // extra signals blocked WHILE handler runs
    int        sa_flags;              // SA_RESTART, SA_SIGINFO, SA_RESETHAND, ...
};

int sigemptyset(sigset_t *set);       // start with an empty mask; MUST init before use
int sigfillset(sigset_t *set);        // block everything

typedef /* impl-defined */ sig_atomic_t; // integer type; atomic wrt signals
// Handler flags MUST be: static volatile sig_atomic_t

ssize_t write(int fd, const void *buf, size_t n); // async-signal-safe; use in handlers
void    _exit(int status);            // async-signal-safe immediate exit (no flush)
int     raise(int sig);               // async-signal-safe; send signal to self
int     pause(void);                  // block until a signal is delivered/handled

Key conventions: always sigemptyset(&sa.sa_mask) (or zero the whole struct) before sigaction — an uninitialized mask is garbage. Prefer sigaction over signal() for portability. Nothing here is heap-allocated, so there is nothing to free; the struct sigaction lives on the stack.

Lesson

What a signal handler interrupts

A signal handler can run at any moment. It interrupts your normal code at an arbitrary point.

That point might be in the middle of:

  • malloc() — allocating memory
  • printf() — writing buffered output
  • pthread_mutex_lock() — holding a lock

If your handler then calls one of those same functions, you can hit serious problems:

  • Deadlock — the lock is already held by the interrupted thread, so the handler waits forever.
  • Corruption — you re-enter code that was halfway through updating its internal state.

Async-signal-safe functions

The Linux manual page signal-safety(7) lists the async-signal-safe functions. These are the functions you are allowed to call from a handler. "Async-signal-safe" means a function is safe to call even when it interrupts other code.

The main safe functions include:

  • write, read
  • _exit
  • kill, signal, sigaction, sigprocmask
  • most simple system calls

Functions that are not safe include:

  • printf and anything else that uses stdio buffers
  • malloc, free, and anything that allocates memory
  • anything that takes a libc lock

The defensive pattern

Keep the handler tiny. Let the main program do the real work.

  • The handler sets a flag.
  • The main loop checks the flag and does the work.

Use a volatile sig_atomic_t for the flag. This is the only type the C standard guarantees can be read and written atomically with respect to signal delivery, so the value is never seen half-updated.

Code examples

#define _POSIX_C_SOURCE 200809L
#include <signal.h>
#include <unistd.h>
#include <string.h>
#include <stdio.h>
#include <errno.h>

/* Flag written from the handler, read from main. sig_atomic_t + volatile is
   the ONLY type the C standard guarantees is safe to touch across a signal. */
static volatile sig_atomic_t g_want_quit = 0;
static volatile sig_atomic_t g_last_sig  = 0;

/* async-signal-safe: prints an integer with write(2) only, no stdio. */
static void safe_put_int(int fd, int v) {
    char buf[16];
    int i = (int)sizeof buf;
    int neg = v < 0;
    unsigned u = neg ? (unsigned)(-(v + 1)) + 1u : (unsigned)v;
    if (u == 0) buf[--i] = '0';
    while (u) { buf[--i] = (char)('0' + (u % 10)); u /= 10; }
    if (neg) buf[--i] = '-';
    (void)write(fd, buf + i, (size_t)((int)sizeof buf - i));
}

static void on_signal(int sig) {
    int saved = errno;                 /* handler must restore errno */
    g_last_sig  = sig;                 /* atomic store */
    g_want_quit = 1;                   /* atomic store */
    static const char msg[] = "[handler] signal received, requesting shutdown\n";
    (void)write(STDOUT_FILENO, msg, sizeof msg - 1); /* write() is safe */
    errno = saved;
}

static int install(int sig, void (*fn)(int)) {
    struct sigaction sa;
    memset(&sa, 0, sizeof sa);
    sa.sa_handler = fn;
    sigemptyset(&sa.sa_mask);          /* no extra signals blocked in handler */
    sa.sa_flags = 0;                   /* no SA_RESTART: let pause() return */
    return sigaction(sig, &sa, NULL);
}

int main(void) {
    if (install(SIGTERM, on_signal) != 0 || install(SIGINT, on_signal) != 0) {
        perror("sigaction");
        return 1;
    }

    printf("pid=%ld  send me SIGTERM/SIGINT (Ctrl-C) to shut down.\n",
           (long)getpid());
    fflush(stdout); /* flush from MAIN, never from the handler */

    /* For a hermetic demo: raise SIGTERM ourselves instead of waiting. */
    raise(SIGTERM);

    /* Main loop: the handler only flipped a flag; the real work is here. */
    while (!g_want_quit) {
        pause(); /* sleeps until any signal arrives */
    }

    /* Safe to use stdio again — we are back in normal control flow. */
    printf("[main] observed shutdown flag; last signal was ");
    safe_put_int(STDOUT_FILENO, g_last_sig);
    printf(" (%s)\n", strsignal(g_last_sig));
    printf("[main] flushing buffers, freeing resources, exiting cleanly.\n");
    return 0;
}

Line by line

  • #define _POSIX_C_SOURCE 200809L — exposes sigaction, strsignal, and friends from the POSIX headers; without it a strict -std=c11 build may not declare them.
  • static volatile sig_atomic_t g_want_quit / g_last_sig — the shared flags. volatile forbids the compiler from caching them (so the while loop actually re-reads them); sig_atomic_t guarantees each read/write is indivisible across signal delivery.
  • safe_put_int(...) — a hand-rolled integer-to-decimal routine using only a stack buffer and write. It exists to demonstrate how you emit a number from signal-restricted code without touching printf/snprintf.
  • on_signal: int saved = errno; snapshots errno on entry; errno = saved; restores it before returning — the write call could otherwise clobber a value the interrupted code needed.
  • g_last_sig = sig; g_want_quit = 1; — the entire real payload of the handler: two atomic stores. Everything else is deferred.
  • static const char msg[] = ...; write(STDOUT_FILENO, msg, sizeof msg - 1); — the only I/O the handler does, via the async-signal-safe write. sizeof msg - 1 drops the trailing NUL. static const keeps the string out of any per-call setup.
  • install(...): memset(&sa, 0, sizeof sa) zeroes the struct; sigemptyset(&sa.sa_mask) makes the mask well-defined; sa.sa_flags = 0 deliberately omits SA_RESTART so pause() returns after the handler runs.
  • raise(SIGTERM) — makes the demo hermetic and self-contained: instead of waiting for an external kill, the process signals itself so the output is reproducible.
  • while (!g_want_quit) pause(); — the main loop. pause() blocks until any signal fires; when the handler returns, pause() returns, the loop re-checks the flag, and exits.
  • The closing printf/strsignal block runs in normal control flow, where full stdio and non-reentrant functions like strsignal are perfectly fine again.

Common mistakes

1. Calling printf (or any stdio) inside the handler

void on_term(int s){ printf("caught %d\n", s); }   // WRONG

Why it breaks: printf locks and writes a shared stdio buffer. If the signal interrupted another printf, the handler deadlocks on the buffer lock or corrupts the buffer — intermittently, under load.

void on_term(int s){ (void)s; write(STDOUT_FILENO, "caught\n", 7); } // FIXED

2. malloc/free/logging frameworks in the handler

void on_term(int s){ char *m = malloc(64); log_event(m); } // WRONG

Why it breaks: malloc takes the heap arena lock. Interrupt a malloc in progress and the handler re-locking the same arena deadlocks, or corrupts the free-list. Most logging libraries allocate internally too.

void on_term(int s){ (void)s; g_want_quit = 1; } // FIXED: defer to main loop

3. Non-atomic / non-volatile flag

int want_quit = 0;               // WRONG: plain int, no volatile
while (!want_quit) { /* ... */ } // compiler may cache it → infinite loop

Why it breaks: the compiler can hoist want_quit into a register and never re-read it, or tear a multi-word write. The loop may spin forever or read a half-written value.

static volatile sig_atomic_t want_quit = 0; // FIXED

4. Forgetting to preserve errno

void on_term(int s){ (void)s; write(2,"x",1); } // clobbers errno silently

Why it breaks: if the handler interrupted code between a failing syscall and its errno check, write overwrites errno and the interrupted code misdiagnoses the error.

void on_term(int s){ int e=errno; (void)s; write(2,"x",1); errno=e; } // FIXED

Debugging tips

  • Reproduce the deadlock with gdb. If a program hangs after a signal, attach with gdb -p <pid> and run bt. A backtrace showing your handler stacked on top of malloc/_int_malloc or _IO_* (stdio) internals, both waiting on a lock, is the classic async-signal-safety deadlock.
  • strace -f ./prog shows signal delivery (--- SIGTERM ---) interleaved with syscalls, and shows EINTR returns from interrupted blocking calls — useful for diagnosing loops that exit early or spin.
  • valgrind --tool=helgrind / --tool=drd can flag lock-order and reentrancy hazards; ThreadSanitizer (-fsanitize=thread) similarly reports data races on flags that aren't properly atomic.
  • Static analysis: compile with -Wall -Wextra; some linters and clang-tidy (bugprone-signal-handler, cert-sig30-c) flag non-async-signal-safe calls inside handlers directly.
  • When in doubt, consult man 7 signal-safety — it is the authoritative list. If a function isn't on it, treat it as unsafe.
  • printf debugging inside a handler is itself a trap — the very function you'd reach for is unsafe. Debug with write of a fixed string, or better, set the flag and print from the main loop.

Memory safety

The hazards here are concurrency and undefined behavior, not classic buffer overflows. (1) Data races: the handler and main thread share the flag; only volatile sig_atomic_t access is defined — a plain int or a struct flag is UB and may tear. (2) Reentrancy corruption: calling non-async-signal-safe functions can corrupt the heap free-list or stdio buffers, which later surfaces as a crash or overwrite far from the signal — a heisenbug. (3) errno clobbering is a subtle correctness bug: unrelated code misreads an error. (4) Longjmp out of a handler (siglongjmp) is legal only with great care — jumping out while libc holds a lock leaves that lock held forever; prefer the flag pattern. (5) In multithreaded programs, remember the handler runs on whichever thread the signal was delivered to, and it must never touch a non-async-signal-safe library on that thread; the usual design is to block signals in all threads and dedicate one thread to sigwait, or use signalfd on Linux, so no code runs in async context at all.

Real-world uses

  • Graceful shutdown in daemons and servers (nginx, PostgreSQL, Redis): SIGTERM sets a shutdown flag; the main loop finishes in-flight requests, flushes, and exits. This is the flag pattern at production scale.
  • Config reload on SIGHUP (most Unix daemons): handler sets a reload flag; the main loop re-reads config where allocation and file I/O are safe.
  • Clean Ctrl-C in CLIs: catch SIGINT to restore the terminal, remove temp files, and exit — instead of leaving the terminal in raw mode.
  • Security-critical correctness: the OpenSSH regreSSHion (CVE-2024-6387) RCE came from a SIGALRM handler calling async-signal-unsafe functions (syslog/malloc paths). Best practice in security code: handlers do the absolute minimum, or signals are handled synchronously via signalfd/sigwait so no restricted context exists. Prefer signalfd (Linux) or a self-pipe/eventfd to convert async signals into readable file descriptors your normal event loop drains.

Practice tasks

  1. (Warm-up) Modify the example to also catch SIGHUP. On SIGHUP, set a separate volatile sig_atomic_t g_reload flag and, in the main loop, print a "reloading config" line — without ever calling stdio from the handler.

  2. (Correctness) Remove the volatile keyword and rebuild with -O2. Replace raise with a real pause() loop and send the signal from another terminal (kill -TERM <pid>). Observe whether the loop still exits, then explain in a comment why volatile mattered (or why the optimizer got lucky).

  3. (Reentrancy) Write a deliberately-broken version whose handler calls printf in a tight loop while main also printfs in a tight loop, run it under signal spam (while true; do kill -INT <pid>; done), and capture a hang. Then attach with gdb and paste the backtrace that proves the stdio-lock deadlock.

  4. (Self-pipe) Implement the self-pipe trick: create a pipe(), have the handler write one byte to the write end (async-signal-safe), and have the main loop select/read the read end so signals become normal I/O events. Handle short reads and EINTR.

  5. (Design, Linux) Rewrite the program to avoid async handlers entirely: block SIGTERM/SIGINT process-wide with sigprocmask, then either call sigwait in a dedicated thread or read events from signalfd, doing all shutdown work in fully normal (unrestricted) code. Explain why this removes the async-signal-safety constraint completely.

Summary

  • A signal handler is an interruption that runs on the interrupted thread, so anything that code was mid-way through is in an inconsistent state.
  • Only async-signal-safe functions (write, _exit, kill, raise, sigaction, sigprocmask, …) may be called from a handler; printf, malloc/free, snprintf, and strsignal are not — they buffer, allocate, or lock shared state.
  • The durable pattern: handler sets a volatile sig_atomic_t flag (and preserves errno), and the main loop does all real work in normal control flow.
  • volatile sig_atomic_t is mandatory for handler flags: sig_atomic_t for indivisible access, volatile so the compiler actually re-reads it.
  • Install with sigaction (init the mask!), choose SA_RESTART deliberately, and for the strongest safety convert signals to file-descriptor events via signalfd/self-pipe/sigwait so no restricted async context ever runs.
  • Real stakes: async-signal-unsafe handlers caused the OpenSSH regreSSHion RCE — this is a security rule, not just a style rule.

Practice with these exercises