Linux System Programming · intermediate · ~8 min

Common Mistakes with Signals

- Explain why only async-signal-safe functions may run inside a handler, and use write() instead of printf. - Share state with a handler correctly using volatile sig_atomic_t flags. - Reap children with a SIGCHLD handler and a waitpid() loop so zombies never pile up. - Handle EINTR on interrupted syscalls, and decide between retrying and SA_RESTART. - Install handlers portably with sigaction() and know which signals (SIGKILL, SIGSTOP) can never be caught.

Overview

You already know how to catch Ctrl-C: install a handler for SIGINT and flip a flag so your program can shut down cleanly instead of dying instantly. That single pattern hides a surprising number of traps. A handler is not a normal function call you make — it is code the kernel injects at an unpredictable instant, possibly in the middle of malloc() or printf(). This lesson catalogs the mistakes that catch almost everyone the first time and shows the correct pattern for each.

Building directly on Ctrl-C handling, we keep the same core idea — a handler that sets a flag while the main loop does the real work — but harden it against the concurrency, portability, and lifecycle bugs that turn a tidy handler into an intermittent crash.

Why it matters

Signal bugs are the classic 'works on my machine, crashes in production once a week' failure: they depend on the exact microsecond a signal lands, so they hide during testing and surface under load. A printf() in a handler can deadlock a server; a missing SIGCHLD reaper leaks a zombie per request until the process table fills and the box stops forking anything. Getting EINTR wrong makes robust programs abort on a harmless Ctrl-Z. And because handlers run with the privileges of your process, a handler that calls unsafe code is a real avenue for state corruption an attacker can try to trigger.

Core concepts

Signals are asynchronous. When one is delivered, the kernel suspends whatever your thread was doing, runs the handler, then resumes. The interrupted code has no idea it was paused. Every mistake below flows from that one fact.

1. Only async-signal-safe functions inside a handler

A handler can fire while the main program is halfway through malloc, printf, or free. Those functions keep global state (heap bookkeeping, stdio buffers) protected by locks or invariants that are momentarily inconsistent. Re-entering them from the handler can deadlock or corrupt memory. POSIX defines a small whitelist of async-signal-safe functions guaranteed to be re-entrant — write, _exit, signal, kill, sigaction, waitpid, and a few dozen more. printf, malloc, free, and most of the C library are not on it.

main thread:  ... printf() holds stdio lock ......[resumes, deadlock]
                                    |
              SIGINT delivered ->   v
handler:                           printf()  <-- tries to take same lock -> hang

Two safe strategies:

Strategy How Use when
Write directly write(STDOUT_FILENO, msg, len) You must emit output from the handler
Flag and defer handler sets a flag; main loop acts Almost always the cleaner choice

2. Share state only through volatile sig_atomic_t

A plain int flag has two problems. First, the compiler may cache it in a register, so the main loop never notices the handler's write — volatile forces a real memory read each time. Second, a non-atomic type could be read while half-written. sig_atomic_t is the one integer type the standard guarantees you can read and write in a single uninterruptible step from a handler.

static volatile sig_atomic_t g_stop = 0;   /* correct */
static int g_stop = 0;                       /* wrong: may be cached / torn */

This does not make it a general concurrency tool — volatile sig_atomic_t is safe only for the single-writer (handler) / single-reader (main) signal case, not for threads. For threads you need atomics or a mutex.

Knowledge check: your handler sets g_stop = 1 but the while (!g_stop) loop in main spins forever. The type is int. What is the most likely cause?

The compiler hoisted g_stop into a register because nothing in the loop appears to change it, so the loop never re-reads memory. Declaring it volatile sig_atomic_t forces a fresh load on every iteration and fixes the hang.

3. SIGKILL and SIGSTOP cannot be caught

sigaction() (and signal()) will fail with EINVAL if you try to install a handler, ignore, or block SIGKILL (9) or SIGSTOP. This is deliberate: the kernel must always retain a way to terminate or freeze a runaway process. Do not design cleanup logic that depends on intercepting kill -9 — it will never run. Put durable cleanup in restart logic or a supervisor, not a SIGKILL handler.

4. Reap children or drown in zombies

When a child exits, it becomes a zombie: a tiny kernel record kept so the parent can read its exit status with wait/waitpid. Until the parent reaps it, that PID slot is consumed. A long-running server that forks per task but never reaps will exhaust the process table.

fork() -> child runs -> child exits -> [ZOMBIE] --waitpid()--> gone
                                          ^
                        stays here until parent reaps it

Install a SIGCHLD handler that reaps in a loop, because signals are not queued — several children exiting close together may coalesce into one SIGCHLD:

while (waitpid(-1, &status, WNOHANG) > 0) { /* reap all ready children */ }

WNOHANG makes waitpid return 0 immediately when no more children are ready, so the loop terminates instead of blocking.

Knowledge check: three children exit within the same millisecond. How many SIGCHLD deliveries are you guaranteed?

Possibly just one. Standard (non-realtime) signals are not counted or queued — they collapse into a single pending bit. That is exactly why the handler must loop on waitpid(..., WNOHANG) until it returns 0, rather than reaping a single child per signal.

5. Handle EINTR on interrupted syscalls

A blocking call like read, accept, or sleep interrupted by a delivered signal returns -1 with errno == EINTR. Treating that as a fatal error is a bug — nothing actually failed. You have two fixes:

Option Mechanism Trade-off
Manual retry wrap the call in do { r = read(...); } while (r < 0 && errno == EINTR); explicit, works everywhere
SA_RESTART pass the flag to sigaction; kernel restarts the call convenient, but not every call is restartable (e.g. some timeouts still return EINTR)

Even with SA_RESTART, a few interfaces still surface EINTR, so robust code keeps the retry wrapper on the calls that matter.

6. Use sigaction(), not signal()

The old signal() has implementation-defined semantics: on some historical systems it reset the disposition to default after the first delivery, forcing you to re-install inside the handler — itself a race. sigaction() gives you portable, stable behavior plus control over the mask and flags (SA_RESTART, SA_NOCLDSTOP, SA_SIGINFO). Prefer it everywhere.

7. alarm(0) and other cancellations still race

alarm(0) cancels a pending alarm — but if the timer already expired and the SIGALRM is already queued for delivery, cancelling the alarm does not un-deliver the signal. There is a small window where the handler still runs. The lesson generalizes: never assume a cancel/stop request retroactively prevents an in-flight signal. Always check the operation's real result rather than trusting that cancellation won.

Syntax notes

int sigaction(int signum, const struct sigaction *act,
              struct sigaction *oldact);
  • signum: the signal to configure (e.g. SIGINT). SIGKILL/SIGSTOP -> fails with EINVAL.
  • act: new disposition; oldact (may be NULL) receives the previous one.
  • Returns 0 on success, -1 with errno set on error.
struct sigaction {
    void     (*sa_handler)(int);          /* handler, or SIG_IGN / SIG_DFL */
    sigset_t   sa_mask;                    /* extra signals blocked during handler */
    int        sa_flags;                   /* SA_RESTART, SA_NOCLDSTOP, SA_SIGINFO ... */
};
  • Always memset the struct to 0 (or zero-init) and call sigemptyset(&sa.sa_mask) before use — uninitialized flags cause erratic behavior.
  • SA_RESTART: auto-restart interrupted slow syscalls.
  • SA_NOCLDSTOP: with SIGCHLD, don't fire when a child merely stops (only on exit).
pid_t waitpid(pid_t pid, int *wstatus, int options);
  • pid == -1: wait for any child. options = WNOHANG: return 0 immediately if none ready (non-blocking).
  • Returns child PID (reaped), 0 (none ready, with WNOHANG), or -1/errno.
  • Inspect status with WIFEXITED(st), WEXITSTATUS(st), WIFSIGNALED(st), WTERMSIG(st).
ssize_t write(int fd, const void *buf, size_t count);   /* async-signal-safe */
  • The safe way to print from a handler. Returns bytes written or -1; in a handler, typically ignore the return.
volatile sig_atomic_t flag;   /* the only integer type safe to touch in a handler */

Lesson

Signals are easy to install but easy to misuse. Here are the mistakes that bite most people, and how to avoid each one.

1. Calling printf or malloc from a handler

A signal handler runs at an unpredictable moment, possibly in the middle of another call. Functions like printf and malloc are not async-signal-safe (safe to call from a handler), so calling them there can corrupt state or deadlock.

  • Use write() to print.
  • Or set a flag in the handler and do the real work later in your main loop.

2. Using a plain int flag

A signal can arrive between the read and the write of an ordinary variable, leaving it in an inconsistent state.

  • Declare the flag as volatile sig_atomic_t. This type is guaranteed safe to read and write atomically across a signal.

3. Trying to catch SIGKILL or SIGSTOP

This cannot be done. The kernel refuses to install a handler for these two signals.

4. No SIGCHLD handler in a server

When a child process exits, it stays as a zombie (a finished process the parent has not yet collected) until the parent reaps it. Without a SIGCHLD handler, zombies accumulate forever.

5. Forgetting EINTR

When a signal interrupts a blocking system call, that call returns -1 with errno == EINTR.

  • Retry the call, or
  • Install the handler with the SA_RESTART flag so the kernel restarts the call automatically.

6. Race between alarm(0) and the signal

alarm(0) cancels a pending alarm. But if the kernel has already queued the delivery, there is a tiny window where the signal still arrives.

  • Do not assume cancellation succeeded. Always check the operation's actual result.

7. Re-installing the handler inside the handler

This is only a concern with the legacy signal() function. Use sigaction() instead and the problem disappears.

Code examples

/* signals_done_right.c -- the safe way to do the things people get wrong.
 * Hermetic: no network, forks one short-lived child, signals itself.
 * Build: cc -std=c11 -Wall -Wextra signals_done_right.c -o demo
 */
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <errno.h>
#include <unistd.h>
#include <signal.h>
#include <sys/types.h>
#include <sys/wait.h>

/* Flags touched from a handler MUST be volatile sig_atomic_t. */
static volatile sig_atomic_t g_stop       = 0;  /* set by SIGTERM/SIGINT */
static volatile sig_atomic_t g_child_gone = 0;  /* set by SIGCHLD        */

/* async-signal-safe: only write(), never printf/malloc, inside a handler. */
static void on_stop(int sig) {
    (void)sig;
    g_stop = 1;
    static const char msg[] = "[handler] stop requested\n";
    write(STDOUT_FILENO, msg, sizeof msg - 1);
}

static void on_child(int sig) {
    (void)sig;
    g_child_gone = 1;  /* just flag it; real reaping happens in main loop */
}

/* Install a handler the portable way, choosing restart behaviour explicitly. */
static int install(int sig, void (*fn)(int), int flags) {
    struct sigaction sa;
    memset(&sa, 0, sizeof sa);
    sa.sa_handler = fn;
    sigemptyset(&sa.sa_mask);
    sa.sa_flags = flags;
    return sigaction(sig, &sa, NULL);
}

/* EINTR-safe wrapper: retry when a signal interrupts the read. */
static ssize_t read_retry(int fd, void *buf, size_t n) {
    ssize_t r;
    do { r = read(fd, buf, n); } while (r < 0 && errno == EINTR);
    return r;
}

int main(void) {
    /* SA_RESTART auto-restarts slow syscalls hit by these signals. */
    if (install(SIGINT,  on_stop,  SA_RESTART) < 0 ||
        install(SIGTERM, on_stop,  SA_RESTART) < 0 ||
        install(SIGCHLD, on_child, SA_RESTART | SA_NOCLDSTOP) < 0) {
        perror("sigaction");
        return 1;
    }

    /* Proof you cannot catch SIGKILL: sigaction refuses it. */
    if (install(SIGKILL, on_stop, 0) < 0)
        printf("[main] as expected, cannot install SIGKILL: %s\n",
               strerror(errno));

    pid_t pid = fork();
    if (pid < 0) { perror("fork"); return 1; }
    if (pid == 0) { _exit(42); }              /* child: exit immediately */
    printf("[main] forked child pid=%d\n", (int)pid);

    /* A self-pipe: the classic way to wake a blocking read from a handler.
     * We use it here only to exercise read_retry() against a real EINTR. */
    kill(getpid(), SIGTERM);                  /* deterministic shutdown */

    int reaped = 0;
    while (!(g_stop && reaped)) {
        if (g_child_gone) {
            g_child_gone = 0;
            /* Reap EVERY finished child; one SIGCHLD can cover several. */
            int status; pid_t w;
            while ((w = waitpid(-1, &status, WNOHANG)) > 0) {
                if (WIFEXITED(status))
                    printf("[main] reaped pid=%d exit=%d\n",
                           (int)w, WEXITSTATUS(status));
                reaped = 1;
            }
        }
        if (g_stop && reaped) break;

        /* If nothing is ready yet, block briefly and stay EINTR-safe. */
        char b;
        int fds[2];
        if (pipe(fds) == 0) {
            close(fds[1]);                     /* write end closed -> EOF */
            (void)read_retry(fds[0], &b, 1);   /* returns 0 at EOF */
            close(fds[0]);
        }
    }

    printf("[main] clean shutdown (stop=%d reaped=%d)\n",
           (int)g_stop, reaped);
    return 0;
}

Line by line

  • g_stop / g_child_gone as volatile sig_atomic_t: the only kind of variable safe to write in a handler and re-read in main without a torn value or a stale cached copy.
  • on_stop: sets the stop flag and prints via write() with a static const buffer — no printf, no heap, fully async-signal-safe. sizeof msg - 1 drops the trailing NUL.
  • on_child: does the minimum — just flags that a child changed state. All the real reaping (which calls waitpid, itself safe, but we keep handlers tiny) is deferred to the main loop.
  • install(): zeroes a struct sigaction, clears the mask with sigemptyset, sets flags explicitly, and calls sigaction. This is the portable replacement for signal().
  • Installing SIGINT/SIGTERM/SIGCHLD with SA_RESTART: interrupted slow syscalls restart automatically; SA_NOCLDSTOP keeps SIGCHLD from firing when a child is merely stopped.
  • The SIGKILL attempt: install(SIGKILL, ...) deliberately fails; we print the EINVAL message to prove the kernel refuses it.
  • fork() + _exit(42): the child exits immediately with a distinctive code so we can watch it get reaped. _exit (not exit) avoids running the parent's stdio flushers in the child.
  • kill(getpid(), SIGTERM): makes the demo deterministic — we signal ourselves so the shutdown path runs without waiting for a human Ctrl-C.
  • The main loop while (!(g_stop && reaped)): exits only once we've been asked to stop and collected the child — the flag-and-defer pattern in action.
  • waitpid(-1, &status, WNOHANG) loop: reaps every ready child, not just one, because a single SIGCHLD may represent several exits.
  • read_retry over a closed pipe: a genuine EINTR-safe blocking read (the write end is closed so it returns EOF); it shows the do { } while (errno == EINTR) idiom against a real syscall.

Common mistakes

1. printf() inside a handler

void h(int s){ (void)s; printf("caught %d\n", s); }   /* WRONG */

Why it breaks: printf takes the stdio lock and touches buffered state; if the signal landed mid-printf in the main thread, the handler deadlocks or corrupts the buffer.

void h(int s){ (void)s; char c='!'; write(STDOUT_FILENO,&c,1); }  /* FIXED */

2. Plain int flag

static int stop = 0;                 /* WRONG: may be cached / torn */
while (!stop) { /* ... */ }

Why it breaks: the compiler can keep stop in a register and never see the handler's write; the loop spins forever.

static volatile sig_atomic_t stop = 0;   /* FIXED */

3. Reaping only one child per SIGCHLD

void h(int s){ (void)s; waitpid(-1,NULL,WNOHANG); }   /* WRONG */

Why it breaks: multiple children exiting together produce a single SIGCHLD; the others stay zombies.

void h(int s){ (void)s; while (waitpid(-1,NULL,WNOHANG) > 0){} }  /* FIXED */

4. Treating EINTR as failure

if (read(fd,buf,n) < 0) { perror("read"); exit(1); }   /* WRONG */

Why it breaks: a harmless signal (e.g. SIGCHLD) interrupts the read; it returns EINTR and the program aborts though nothing failed.

ssize_t r; do { r = read(fd,buf,n); } while (r < 0 && errno == EINTR);  /* FIXED */

Debugging tips

  • strace / ltrace: strace -f ./prog shows every signal delivery (--- SIGINT ---), every sigaction install, and syscalls returning EINTR or ERESTARTSYS. This is the fastest way to see whether your handler even fired.
  • ps / ps aux | grep defunct: zombies show as <defunct> or state Z. A growing count means your SIGCHLD reaper is missing or not looping.
  • gdb: signals can be tricky under gdb because it intercepts them. Use handle SIGINT nostop pass to let your program receive them; catch signal SIGCHLD to break on delivery; bt inside the handler to see where it interrupted.
  • printf debugging done safely: never add printf to a handler to debug it — the debug print itself introduces the async-safety bug. Set a flag and print from the main loop, or use write with a fixed string.
  • valgrind --tool=helgrind / DRD: catches the case where a handler touches data the main code also uses without proper synchronization.
  • Reproduce EINTR on demand: send a harmless signal (kill -USR1 <pid>, or a repeating SIGALRM) while the program blocks in a syscall, and confirm it recovers instead of aborting.

Memory safety

  • Re-entrancy is the core hazard. A handler that calls non-async-signal-safe code (malloc, free, printf, most of libc) can re-enter a function that is mid-update, corrupting the heap or stdio state — classic undefined behavior that manifests as random crashes far from the real cause.
  • Torn / cached reads. Sharing anything other than volatile sig_atomic_t (or a proper atomic) between handler and main code is a data race and undefined behavior. Reads may be stale or half-written.
  • Async-cancel of allocations. If a signal arrives between malloc and storing the pointer, and the handler longjmps out, you leak. Prefer flag-and-defer over longjmp/siglongjmp out of handlers; if you must, be extremely disciplined.
  • errno clobbering. A handler that calls any syscall (even write or waitpid) can overwrite errno, confusing the interrupted code. Save and restore it: int e = errno; ...; errno = e;.
  • Handler vs threads. In a multithreaded program a signal may run on any thread; volatile sig_atomic_t is not a thread-sync primitive. Use pthread_sigmask to steer delivery and C11 atomics for cross-thread flags.
  • Fork safety. After fork, only async-signal-safe functions are legal in the child until exec; the demo uses _exit, not exit, for exactly this reason.

Real-world uses

  • Graceful shutdown: servers (nginx, PostgreSQL, most daemons) catch SIGTERM to finish in-flight requests, flush state, and close sockets before exiting — the flag-and-defer pattern at scale.
  • Config reload: SIGHUP is the conventional 're-read your config' signal; the handler sets a flag and the main loop reloads at a safe point.
  • Child management: shells, init systems, and process supervisors reap children via SIGCHLD loops to keep the process table clean.
  • Timeouts: alarm/setitimer + SIGALRM guard operations that might hang — with awareness of the cancellation race covered above.
  • Best practice: keep handlers to a few lines (set a flag, maybe write); always use sigaction with an explicit mask and flags; prefer modern alternatives when available — signalfd/pidfd (Linux) or a pselect/ppoll self-pipe turn asynchronous signals into ordinary readable file descriptors you can handle in your normal event loop, sidestepping most of these traps entirely.

Practice tasks

  1. Rewrite a handler that currently calls printf("got signal\n") so it is async-signal-safe, using a static const buffer and a single write() to STDOUT. Verify the output still appears.

  2. Take a program with static int running = 1; shared with a SIGINT handler and a while (running) loop that hangs on exit. Change the type to volatile sig_atomic_t and confirm Ctrl-C now stops it. Explain in a comment why the original hung.

  3. Write a program that forks 5 children which each _exit(i), then reaps them all from a single SIGCHLD handler using a waitpid(-1, &st, WNOHANG) loop. Print each reaped PID and exit code, and use ps to confirm no <defunct> processes remain.

  4. Wrap a blocking read() from a pipe in an EINTR-safe retry loop. Then send the process a harmless signal (e.g. SIGUSR1 with an empty handler) mid-read and show the read completes instead of failing. Repeat with SA_RESTART set and compare behavior.

  5. Convert a program that uses signal() to sigaction(): install SIGTERM and SIGCHLD with explicit sa_mask and appropriate flags (SA_RESTART, SA_NOCLDSTOP), save errno inside the handler, and attempt to install a SIGKILL handler to confirm it fails with EINVAL.

Summary

  • Handlers are interruptions, not calls. Only async-signal-safe functions (write, _exit, waitpid, sigaction, ...) are legal inside one — never printf/malloc/free.
  • Keep handlers tiny: set a volatile sig_atomic_t flag and let the main loop do the real work. Save/restore errno if the handler makes any syscall.
  • volatile sig_atomic_t is the only integer type safe to share with a handler — a plain int can be cached or torn.
  • SIGKILL and SIGSTOP can never be caught; don't build cleanup that depends on them.
  • Reap children in a waitpid(..., WNOHANG) loop from SIGCHLD — signals don't queue, so one delivery may cover many exits, and unreaped children become zombies.
  • EINTR is not failure: retry the call, or set SA_RESTART (but keep retry wrappers where it matters).
  • Always use sigaction(), not signal(), for portable, well-defined behavior; and remember that cancellations like alarm(0) can still race an in-flight signal.

Practice with these exercises