Linux System Programming · intermediate · ~8 min
- Explain why only async-signal-safe functions may run inside a handler, and use write() instead of printf. - Share state with a handler correctly using volatile sig_atomic_t flags. - Reap children with a SIGCHLD handler and a waitpid() loop so zombies never pile up. - Handle EINTR on interrupted syscalls, and decide between retrying and SA_RESTART. - Install handlers portably with sigaction() and know which signals (SIGKILL, SIGSTOP) can never be caught.
You already know how to catch Ctrl-C: install a handler for SIGINT and flip a flag so your program can shut down cleanly instead of dying instantly. That single pattern hides a surprising number of traps. A handler is not a normal function call you make — it is code the kernel injects at an unpredictable instant, possibly in the middle of malloc() or printf(). This lesson catalogs the mistakes that catch almost everyone the first time and shows the correct pattern for each.
Building directly on Ctrl-C handling, we keep the same core idea — a handler that sets a flag while the main loop does the real work — but harden it against the concurrency, portability, and lifecycle bugs that turn a tidy handler into an intermittent crash.
Signal bugs are the classic 'works on my machine, crashes in production once a week' failure: they depend on the exact microsecond a signal lands, so they hide during testing and surface under load. A printf() in a handler can deadlock a server; a missing SIGCHLD reaper leaks a zombie per request until the process table fills and the box stops forking anything. Getting EINTR wrong makes robust programs abort on a harmless Ctrl-Z. And because handlers run with the privileges of your process, a handler that calls unsafe code is a real avenue for state corruption an attacker can try to trigger.
Signals are asynchronous. When one is delivered, the kernel suspends whatever your thread was doing, runs the handler, then resumes. The interrupted code has no idea it was paused. Every mistake below flows from that one fact.
A handler can fire while the main program is halfway through malloc, printf, or free. Those functions keep global state (heap bookkeeping, stdio buffers) protected by locks or invariants that are momentarily inconsistent. Re-entering them from the handler can deadlock or corrupt memory. POSIX defines a small whitelist of async-signal-safe functions guaranteed to be re-entrant — write, _exit, signal, kill, sigaction, waitpid, and a few dozen more. printf, malloc, free, and most of the C library are not on it.
main thread: ... printf() holds stdio lock ......[resumes, deadlock]
|
SIGINT delivered -> v
handler: printf() <-- tries to take same lock -> hang
Two safe strategies:
| Strategy | How | Use when |
|---|---|---|
| Write directly | write(STDOUT_FILENO, msg, len) |
You must emit output from the handler |
| Flag and defer | handler sets a flag; main loop acts | Almost always the cleaner choice |
volatile sig_atomic_tA plain int flag has two problems. First, the compiler may cache it in a register, so the main loop never notices the handler's write — volatile forces a real memory read each time. Second, a non-atomic type could be read while half-written. sig_atomic_t is the one integer type the standard guarantees you can read and write in a single uninterruptible step from a handler.
static volatile sig_atomic_t g_stop = 0; /* correct */
static int g_stop = 0; /* wrong: may be cached / torn */
This does not make it a general concurrency tool — volatile sig_atomic_t is safe only for the single-writer (handler) / single-reader (main) signal case, not for threads. For threads you need atomics or a mutex.
Knowledge check: your handler sets g_stop = 1 but the while (!g_stop) loop in main spins forever. The type is int. What is the most likely cause?
The compiler hoisted
g_stopinto a register because nothing in the loop appears to change it, so the loop never re-reads memory. Declaring itvolatile sig_atomic_tforces a fresh load on every iteration and fixes the hang.
sigaction() (and signal()) will fail with EINVAL if you try to install a handler, ignore, or block SIGKILL (9) or SIGSTOP. This is deliberate: the kernel must always retain a way to terminate or freeze a runaway process. Do not design cleanup logic that depends on intercepting kill -9 — it will never run. Put durable cleanup in restart logic or a supervisor, not a SIGKILL handler.
When a child exits, it becomes a zombie: a tiny kernel record kept so the parent can read its exit status with wait/waitpid. Until the parent reaps it, that PID slot is consumed. A long-running server that forks per task but never reaps will exhaust the process table.
fork() -> child runs -> child exits -> [ZOMBIE] --waitpid()--> gone
^
stays here until parent reaps it
Install a SIGCHLD handler that reaps in a loop, because signals are not queued — several children exiting close together may coalesce into one SIGCHLD:
while (waitpid(-1, &status, WNOHANG) > 0) { /* reap all ready children */ }
WNOHANG makes waitpid return 0 immediately when no more children are ready, so the loop terminates instead of blocking.
Knowledge check: three children exit within the same millisecond. How many SIGCHLD deliveries are you guaranteed?
Possibly just one. Standard (non-realtime) signals are not counted or queued — they collapse into a single pending bit. That is exactly why the handler must loop on
waitpid(..., WNOHANG)until it returns 0, rather than reaping a single child per signal.
A blocking call like read, accept, or sleep interrupted by a delivered signal returns -1 with errno == EINTR. Treating that as a fatal error is a bug — nothing actually failed. You have two fixes:
| Option | Mechanism | Trade-off |
|---|---|---|
| Manual retry | wrap the call in do { r = read(...); } while (r < 0 && errno == EINTR); |
explicit, works everywhere |
SA_RESTART |
pass the flag to sigaction; kernel restarts the call |
convenient, but not every call is restartable (e.g. some timeouts still return EINTR) |
Even with SA_RESTART, a few interfaces still surface EINTR, so robust code keeps the retry wrapper on the calls that matter.
sigaction(), not signal()The old signal() has implementation-defined semantics: on some historical systems it reset the disposition to default after the first delivery, forcing you to re-install inside the handler — itself a race. sigaction() gives you portable, stable behavior plus control over the mask and flags (SA_RESTART, SA_NOCLDSTOP, SA_SIGINFO). Prefer it everywhere.
alarm(0) and other cancellations still racealarm(0) cancels a pending alarm — but if the timer already expired and the SIGALRM is already queued for delivery, cancelling the alarm does not un-deliver the signal. There is a small window where the handler still runs. The lesson generalizes: never assume a cancel/stop request retroactively prevents an in-flight signal. Always check the operation's real result rather than trusting that cancellation won.
int sigaction(int signum, const struct sigaction *act,
struct sigaction *oldact);
signum: the signal to configure (e.g. SIGINT). SIGKILL/SIGSTOP -> fails with EINVAL.act: new disposition; oldact (may be NULL) receives the previous one.0 on success, -1 with errno set on error.struct sigaction {
void (*sa_handler)(int); /* handler, or SIG_IGN / SIG_DFL */
sigset_t sa_mask; /* extra signals blocked during handler */
int sa_flags; /* SA_RESTART, SA_NOCLDSTOP, SA_SIGINFO ... */
};
memset the struct to 0 (or zero-init) and call sigemptyset(&sa.sa_mask) before use — uninitialized flags cause erratic behavior.SA_RESTART: auto-restart interrupted slow syscalls.SA_NOCLDSTOP: with SIGCHLD, don't fire when a child merely stops (only on exit).pid_t waitpid(pid_t pid, int *wstatus, int options);
pid == -1: wait for any child. options = WNOHANG: return 0 immediately if none ready (non-blocking).0 (none ready, with WNOHANG), or -1/errno.WIFEXITED(st), WEXITSTATUS(st), WIFSIGNALED(st), WTERMSIG(st).ssize_t write(int fd, const void *buf, size_t count); /* async-signal-safe */
-1; in a handler, typically ignore the return.volatile sig_atomic_t flag; /* the only integer type safe to touch in a handler */
Signals are easy to install but easy to misuse. Here are the mistakes that bite most people, and how to avoid each one.
printf or malloc from a handlerA signal handler runs at an unpredictable moment, possibly in the middle of another call. Functions like printf and malloc are not async-signal-safe (safe to call from a handler), so calling them there can corrupt state or deadlock.
write() to print.int flagA signal can arrive between the read and the write of an ordinary variable, leaving it in an inconsistent state.
volatile sig_atomic_t. This type is guaranteed safe to read and write atomically across a signal.SIGKILL or SIGSTOPThis cannot be done. The kernel refuses to install a handler for these two signals.
SIGCHLD handler in a serverWhen a child process exits, it stays as a zombie (a finished process the parent has not yet collected) until the parent reaps it. Without a SIGCHLD handler, zombies accumulate forever.
EINTRWhen a signal interrupts a blocking system call, that call returns -1 with errno == EINTR.
SA_RESTART flag so the kernel restarts the call automatically.alarm(0) and the signalalarm(0) cancels a pending alarm. But if the kernel has already queued the delivery, there is a tiny window where the signal still arrives.
This is only a concern with the legacy signal() function. Use sigaction() instead and the problem disappears.
/* signals_done_right.c -- the safe way to do the things people get wrong.
* Hermetic: no network, forks one short-lived child, signals itself.
* Build: cc -std=c11 -Wall -Wextra signals_done_right.c -o demo
*/
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <errno.h>
#include <unistd.h>
#include <signal.h>
#include <sys/types.h>
#include <sys/wait.h>
/* Flags touched from a handler MUST be volatile sig_atomic_t. */
static volatile sig_atomic_t g_stop = 0; /* set by SIGTERM/SIGINT */
static volatile sig_atomic_t g_child_gone = 0; /* set by SIGCHLD */
/* async-signal-safe: only write(), never printf/malloc, inside a handler. */
static void on_stop(int sig) {
(void)sig;
g_stop = 1;
static const char msg[] = "[handler] stop requested\n";
write(STDOUT_FILENO, msg, sizeof msg - 1);
}
static void on_child(int sig) {
(void)sig;
g_child_gone = 1; /* just flag it; real reaping happens in main loop */
}
/* Install a handler the portable way, choosing restart behaviour explicitly. */
static int install(int sig, void (*fn)(int), int flags) {
struct sigaction sa;
memset(&sa, 0, sizeof sa);
sa.sa_handler = fn;
sigemptyset(&sa.sa_mask);
sa.sa_flags = flags;
return sigaction(sig, &sa, NULL);
}
/* EINTR-safe wrapper: retry when a signal interrupts the read. */
static ssize_t read_retry(int fd, void *buf, size_t n) {
ssize_t r;
do { r = read(fd, buf, n); } while (r < 0 && errno == EINTR);
return r;
}
int main(void) {
/* SA_RESTART auto-restarts slow syscalls hit by these signals. */
if (install(SIGINT, on_stop, SA_RESTART) < 0 ||
install(SIGTERM, on_stop, SA_RESTART) < 0 ||
install(SIGCHLD, on_child, SA_RESTART | SA_NOCLDSTOP) < 0) {
perror("sigaction");
return 1;
}
/* Proof you cannot catch SIGKILL: sigaction refuses it. */
if (install(SIGKILL, on_stop, 0) < 0)
printf("[main] as expected, cannot install SIGKILL: %s\n",
strerror(errno));
pid_t pid = fork();
if (pid < 0) { perror("fork"); return 1; }
if (pid == 0) { _exit(42); } /* child: exit immediately */
printf("[main] forked child pid=%d\n", (int)pid);
/* A self-pipe: the classic way to wake a blocking read from a handler.
* We use it here only to exercise read_retry() against a real EINTR. */
kill(getpid(), SIGTERM); /* deterministic shutdown */
int reaped = 0;
while (!(g_stop && reaped)) {
if (g_child_gone) {
g_child_gone = 0;
/* Reap EVERY finished child; one SIGCHLD can cover several. */
int status; pid_t w;
while ((w = waitpid(-1, &status, WNOHANG)) > 0) {
if (WIFEXITED(status))
printf("[main] reaped pid=%d exit=%d\n",
(int)w, WEXITSTATUS(status));
reaped = 1;
}
}
if (g_stop && reaped) break;
/* If nothing is ready yet, block briefly and stay EINTR-safe. */
char b;
int fds[2];
if (pipe(fds) == 0) {
close(fds[1]); /* write end closed -> EOF */
(void)read_retry(fds[0], &b, 1); /* returns 0 at EOF */
close(fds[0]);
}
}
printf("[main] clean shutdown (stop=%d reaped=%d)\n",
(int)g_stop, reaped);
return 0;
}
g_stop / g_child_gone as volatile sig_atomic_t: the only kind of variable safe to write in a handler and re-read in main without a torn value or a stale cached copy.on_stop: sets the stop flag and prints via write() with a static const buffer — no printf, no heap, fully async-signal-safe. sizeof msg - 1 drops the trailing NUL.on_child: does the minimum — just flags that a child changed state. All the real reaping (which calls waitpid, itself safe, but we keep handlers tiny) is deferred to the main loop.install(): zeroes a struct sigaction, clears the mask with sigemptyset, sets flags explicitly, and calls sigaction. This is the portable replacement for signal().SA_RESTART: interrupted slow syscalls restart automatically; SA_NOCLDSTOP keeps SIGCHLD from firing when a child is merely stopped.install(SIGKILL, ...) deliberately fails; we print the EINVAL message to prove the kernel refuses it.fork() + _exit(42): the child exits immediately with a distinctive code so we can watch it get reaped. _exit (not exit) avoids running the parent's stdio flushers in the child.kill(getpid(), SIGTERM): makes the demo deterministic — we signal ourselves so the shutdown path runs without waiting for a human Ctrl-C.while (!(g_stop && reaped)): exits only once we've been asked to stop and collected the child — the flag-and-defer pattern in action.waitpid(-1, &status, WNOHANG) loop: reaps every ready child, not just one, because a single SIGCHLD may represent several exits.read_retry over a closed pipe: a genuine EINTR-safe blocking read (the write end is closed so it returns EOF); it shows the do { } while (errno == EINTR) idiom against a real syscall.1. printf() inside a handler
void h(int s){ (void)s; printf("caught %d\n", s); } /* WRONG */
Why it breaks: printf takes the stdio lock and touches buffered state; if the signal landed mid-printf in the main thread, the handler deadlocks or corrupts the buffer.
void h(int s){ (void)s; char c='!'; write(STDOUT_FILENO,&c,1); } /* FIXED */
2. Plain int flag
static int stop = 0; /* WRONG: may be cached / torn */
while (!stop) { /* ... */ }
Why it breaks: the compiler can keep stop in a register and never see the handler's write; the loop spins forever.
static volatile sig_atomic_t stop = 0; /* FIXED */
3. Reaping only one child per SIGCHLD
void h(int s){ (void)s; waitpid(-1,NULL,WNOHANG); } /* WRONG */
Why it breaks: multiple children exiting together produce a single SIGCHLD; the others stay zombies.
void h(int s){ (void)s; while (waitpid(-1,NULL,WNOHANG) > 0){} } /* FIXED */
4. Treating EINTR as failure
if (read(fd,buf,n) < 0) { perror("read"); exit(1); } /* WRONG */
Why it breaks: a harmless signal (e.g. SIGCHLD) interrupts the read; it returns EINTR and the program aborts though nothing failed.
ssize_t r; do { r = read(fd,buf,n); } while (r < 0 && errno == EINTR); /* FIXED */
strace -f ./prog shows every signal delivery (--- SIGINT ---), every sigaction install, and syscalls returning EINTR or ERESTARTSYS. This is the fastest way to see whether your handler even fired.ps / ps aux | grep defunct: zombies show as <defunct> or state Z. A growing count means your SIGCHLD reaper is missing or not looping.handle SIGINT nostop pass to let your program receive them; catch signal SIGCHLD to break on delivery; bt inside the handler to see where it interrupted.printf to a handler to debug it — the debug print itself introduces the async-safety bug. Set a flag and print from the main loop, or use write with a fixed string.kill -USR1 <pid>, or a repeating SIGALRM) while the program blocks in a syscall, and confirm it recovers instead of aborting.malloc, free, printf, most of libc) can re-enter a function that is mid-update, corrupting the heap or stdio state — classic undefined behavior that manifests as random crashes far from the real cause.volatile sig_atomic_t (or a proper atomic) between handler and main code is a data race and undefined behavior. Reads may be stale or half-written.malloc and storing the pointer, and the handler longjmps out, you leak. Prefer flag-and-defer over longjmp/siglongjmp out of handlers; if you must, be extremely disciplined.errno clobbering. A handler that calls any syscall (even write or waitpid) can overwrite errno, confusing the interrupted code. Save and restore it: int e = errno; ...; errno = e;.volatile sig_atomic_t is not a thread-sync primitive. Use pthread_sigmask to steer delivery and C11 atomics for cross-thread flags.fork, only async-signal-safe functions are legal in the child until exec; the demo uses _exit, not exit, for exactly this reason.alarm/setitimer + SIGALRM guard operations that might hang — with awareness of the cancellation race covered above.write); always use sigaction with an explicit mask and flags; prefer modern alternatives when available — signalfd/pidfd (Linux) or a pselect/ppoll self-pipe turn asynchronous signals into ordinary readable file descriptors you can handle in your normal event loop, sidestepping most of these traps entirely.Rewrite a handler that currently calls printf("got signal\n") so it is async-signal-safe, using a static const buffer and a single write() to STDOUT. Verify the output still appears.
Take a program with static int running = 1; shared with a SIGINT handler and a while (running) loop that hangs on exit. Change the type to volatile sig_atomic_t and confirm Ctrl-C now stops it. Explain in a comment why the original hung.
Write a program that forks 5 children which each _exit(i), then reaps them all from a single SIGCHLD handler using a waitpid(-1, &st, WNOHANG) loop. Print each reaped PID and exit code, and use ps to confirm no <defunct> processes remain.
Wrap a blocking read() from a pipe in an EINTR-safe retry loop. Then send the process a harmless signal (e.g. SIGUSR1 with an empty handler) mid-read and show the read completes instead of failing. Repeat with SA_RESTART set and compare behavior.
Convert a program that uses signal() to sigaction(): install SIGTERM and SIGCHLD with explicit sa_mask and appropriate flags (SA_RESTART, SA_NOCLDSTOP), save errno inside the handler, and attempt to install a SIGKILL handler to confirm it fails with EINVAL.
write, _exit, waitpid, sigaction, ...) are legal inside one — never printf/malloc/free.volatile sig_atomic_t flag and let the main loop do the real work. Save/restore errno if the handler makes any syscall.volatile sig_atomic_t is the only integer type safe to share with a handler — a plain int can be cached or torn.waitpid(..., WNOHANG) loop from SIGCHLD — signals don't queue, so one delivery may cover many exits, and unreaped children become zombies.SA_RESTART (but keep retry wrappers where it matters).sigaction(), not signal(), for portable, well-defined behavior; and remember that cancellations like alarm(0) can still race an in-flight signal.