Linux System Programming · intermediate · ~20 min

eventfd, timerfd, signalfd — synthetic fds

By the end you can: - Explain what **eventfd**, **timerfd**, and **signalfd** are and why Linux exposes asynchronous events (counters, timers, signals) as ordinary readable file descriptors. - Create each fd with the correct flags (`EFD_/TFD_/SFD_CLOEXEC`, `*_NONBLOCK`) and read its fixed-size payload correctly. - Register all three in a single **epoll** loop and dispatch on `data.fd`, so one thread multiplexes timers, signals, and cross-thread wakeups alongside sockets. - Handle signals *synchronously* by blocking them with `sigprocmask` first, avoiding the classic async-signal-safety traps. - Recognise the common bugs (short reads, forgetting to drain, forgetting to block signals, fd leaks) and diagnose them.

Overview

You already know two things this lesson depends on: file descriptors (an int handle into the kernel's per-process fd table, backing sockets, pipes, and files) and the epoll event loop (epoll_create1 + epoll_ctl + epoll_wait, which tells you which fds are ready without polling each one). Those work great for sockets and pipes — but timers, signals, and "another thread wants me to wake up" are traditionally handled by different mechanisms (alarm/setitimer, async signal handlers, pthread_cond_t) that do not fit into an epoll loop.

Linux fixes this with three syscalls that turn those out-of-band events into plain readable fds. eventfd is a 64-bit counter you can write to wake a reader; timerfd makes a timer's expirations readable; signalfd delivers signals as data you read() in normal code. Once each event source is an fd, your existing epoll knowledge just works — you add them with epoll_ctl exactly like a socket, and one epoll_wait loop owns everything.

Why it matters

Real Linux servers are built around a single event loop, and "everything is an fd" is what makes that loop complete. Without these fds you are forced back into fragile patterns: a signal handler that can only legally touch a volatile sig_atomic_t, a self-pipe hack to wake select, or a separate timer thread with its own locking. signalfd in particular eliminates a whole bug class — async-signal-safety — because you handle SIGINT/SIGTERM as ordinary data instead of inside a handler that may interrupt malloc. For defensive/security-sensitive daemons this matters: signal handlers are a notorious source of re-entrancy and race bugs, and a correctly blocked-then-signalfd shutdown path is far easier to audit and reason about than handler-based teardown. systemd, container runtimes, databases, and language runtimes all rely on these fds.

Core concepts

These three syscalls share one idea: convert an asynchronous event into a readable file descriptor so it can be multiplexed with epoll just like a socket. Each has a fixed-size read payload — read the wrong number of bytes and you get an error, not a partial result.

        asynchronous world            ->        "just an fd" world
   ---------------------------                ------------------------
   another thread: "wake up!"   --eventfd-->  read() an 8-byte counter
   100ms elapsed                --timerfd-->  read() an 8-byte expiry count
   SIGINT delivered             --signalfd->  read() a signalfd_siginfo record

                     +-----------------------------+
   sfd tfd efd  -->  |          epoll_wait         |  --> dispatch on data.fd
   socket sock  -->  |  (one loop, one thread)     |
                     +-----------------------------+
fd type header create call read payload what a read means
eventfd <sys/eventfd.h> eventfd() uint64_t (8 bytes) current counter value (then reset to 0)
timerfd <sys/timerfd.h> timerfd_create() uint64_t (8 bytes) number of expirations since last read
signalfd <sys/signalfd.h> signalfd() struct signalfd_siginfo one queued signal's metadata

eventfd — a counter you can wake a reader with

An eventfd wraps a single 64-bit unsigned counter in the kernel.

  • write(fd, &n, 8) adds n to the counter (must be 8 bytes; n must not overflow to UINT64_MAX).
  • read(fd, &out, 8) in the default mode returns the whole counter and resets it to 0.
  • In EFD_SEMAPHORE mode, each read returns 1 and decrements the counter by 1 (so N writes of 1 = N successful reads) — useful for counting-semaphore semantics.
  • If the counter is 0 and the fd is EFD_NONBLOCK, read returns -1/EAGAIN instead of blocking. With epoll you only read after EPOLLIN, so it will be non-zero.

This is the idiomatic "wake the loop" primitive: a worker thread does write(efd, &one, 8), the epoll loop sees EPOLLIN and drains it. No mutex, no condition variable, no lost-wakeup window.

timerfd — timer expirations as readable data

timerfd_create(clockid, flags) makes a disarmed timer; timerfd_settime arms it with a struct itimerspec:

  • it_value = when it first fires (relative to now, unless TFD_TIMER_ABSTIME). If both fields are 0, the timer is disarmed.
  • it_interval = period for repeats. Zero interval = one-shot.

When the timer fires, the fd becomes readable; read returns a uint64_t = how many times it expired since your last read. That count matters: if your loop was busy and the timer ticked 3 times, one read returns 3. You must read to re-arm readability — if you never read, epoll keeps reporting it ready (level-triggered).

Knowledge check: your 100ms periodic timerfd is registered with epoll, but your loop stalled for 350ms doing other work. When you finally read the timerfd, what value comes back, and why does that matter?

It returns 3 (three expirations coalesced into one readable event). It matters because if you assumed "one read = one tick" you would under-count elapsed periods and drift; correct code adds exp to its tick counter rather than incrementing by 1.

signalfd — signals without a signal handler

Normally a signal interrupts your program and runs a handler in async context, where you may only call async-signal-safe functions. signalfd instead makes signals readable:

  1. Block the signals first with sigprocmask(SIG_BLOCK, &mask, NULL). Blocking means the kernel keeps them pending rather than delivering them to a handler or default action.
  2. Call signalfd(-1, &mask, flags) to get an fd that becomes readable whenever one of those (now-blocked) signals is pending.
  3. read() yields one or more struct signalfd_siginfo records; si.ssi_signo tells you which signal.

Because you read in ordinary code, there are no async-signal-safety constraints — you can printf, malloc, flush buffers, and shut down cleanly.

Defensive note: block before you create the signalfd

The ordering is a real correctness/security property, not a style choice:

CORRECT:  sigprocmask(BLOCK, {SIGINT})  -->  signalfd(...)  -->  epoll  -->  read
WRONG:    signalfd(...) without blocking -->  SIGINT still runs the default
          action (process dies) OR an installed handler races your read

If you skip the block, the signal follows its normal disposition: SIGINT/SIGTERM terminate the process before your loop ever sees the fd. For a daemon, that means shutdown logic (flush logs, close connections, release locks) silently never runs. An audit rule for any signalfd-based service: the signal set passed to signalfd must be a subset of the currently blocked set, established before the fd is created, and (in a multithreaded program) blocked in every thread — signal masks are per-thread, so block in main before spawning workers, and threads inherit the mask.

flag applies to effect
EFD_CLOEXEC / TFD_CLOEXEC / SFD_CLOEXEC all three fd is closed automatically on exec* (no leak into child programs)
EFD_NONBLOCK / TFD_NONBLOCK / SFD_NONBLOCK all three reads return EAGAIN instead of blocking when nothing is ready
EFD_SEMAPHORE eventfd read returns 1 and decrements by 1, instead of returning+zeroing
TFD_TIMER_ABSTIME timerfd_settime interpret it_value as an absolute time on the chosen clock

Always pass *_CLOEXEC: an fd that leaks across exec into a child process is both a resource leak and a potential information/capability leak.

Syntax notes

#include <sys/eventfd.h>
int eventfd(unsigned int initval, int flags);
// initval: starting counter value. flags: EFD_CLOEXEC | EFD_NONBLOCK | EFD_SEMAPHORE.
// returns: a new fd, or -1 (errno set). read/write use a uint64_t (exactly 8 bytes).

#include <sys/timerfd.h>
int timerfd_create(int clockid, int flags);
// clockid: usually CLOCK_MONOTONIC (immune to wall-clock jumps) or CLOCK_REALTIME.
// flags: TFD_CLOEXEC | TFD_NONBLOCK. returns: fd, or -1.

int timerfd_settime(int fd, int flags,
                    const struct itimerspec *new_value,
                    struct itimerspec *old_value);
// flags: 0, or TFD_TIMER_ABSTIME. new_value->it_value arms it (0/0 disarms);
// it_interval sets the repeat period. old_value (may be NULL) receives the
// previous setting. returns: 0, or -1.
// struct itimerspec { struct timespec it_interval; struct timespec it_value; };
// struct timespec  { time_t tv_sec; long tv_nsec; };  // tv_nsec in [0, 999999999]

#include <sys/signalfd.h>
int signalfd(int fd, const sigset_t *mask, int flags);
// fd: -1 to create a new fd, or an existing signalfd to update its mask.
// mask: the set of signals to receive (must already be BLOCKED via sigprocmask).
// flags: SFD_CLOEXEC | SFD_NONBLOCK. returns: fd, or -1.
// read yields struct signalfd_siginfo; key field: uint32_t ssi_signo.

#include <sys/epoll.h>
int epoll_create1(int flags);                                   // EPOLL_CLOEXEC
int epoll_ctl(int epfd, int op, int fd, struct epoll_event *ev);// ADD/MOD/DEL
int epoll_wait(int epfd, struct epoll_event *evs, int max, int timeout_ms);

Lifetime rules: every one of these is a real fd — close() it when done (or rely on _CLOEXEC across exec, but still close on normal teardown). Reads/writes are all-or-nothing at the fixed payload size; a buffer smaller than the payload fails with EINVAL.

Lesson

Three Linux-specific syscalls turn normally-asynchronous events into readable file descriptors:

  • a timer firing
  • a signal arriving
  • another thread saying "wake up"

Combined with epoll, this gives you one unified event loop. No signal handlers, no pselect, no condition variables.

Code examples

#define _GNU_SOURCE
#include <stdio.h>
#include <stdlib.h>
#include <stdint.h>
#include <string.h>
#include <unistd.h>
#include <errno.h>
#include <time.h>
#include <pthread.h>
#include <signal.h>
#include <sys/epoll.h>
#include <sys/eventfd.h>
#include <sys/timerfd.h>
#include <sys/signalfd.h>

/* Abort the program if a syscall returns -1, printing why. */
static int must(int rc, const char *what) {
    if (rc == -1) { perror(what); exit(EXIT_FAILURE); }
    return rc;
}

/* A worker thread: after a short pause it "wakes" the event loop by
   writing to the eventfd. This models a background job signalling
   the main loop without shared locks or condition variables. */
static void *worker(void *arg) {
    int evfd = *(int *)arg;
    struct timespec nap = { .tv_sec = 0, .tv_nsec = 250 * 1000 * 1000 };
    nanosleep(&nap, NULL);
    uint64_t one = 1;
    if (write(evfd, &one, sizeof one) != (ssize_t)sizeof one)
        perror("worker write");
    return NULL;
}

int main(void) {
    /* 1. Block the signals we want to receive via signalfd, so the
          kernel queues them for us instead of running a handler. */
    sigset_t mask;
    sigemptyset(&mask);
    sigaddset(&mask, SIGINT);
    must(sigprocmask(SIG_BLOCK, &mask, NULL), "sigprocmask");

    int sfd = must(signalfd(-1, &mask, SFD_CLOEXEC | SFD_NONBLOCK), "signalfd");

    /* 2. A periodic timer: first fire at 100ms, then every 100ms. */
    int tfd = must(timerfd_create(CLOCK_MONOTONIC, TFD_CLOEXEC | TFD_NONBLOCK),
                   "timerfd_create");
    struct itimerspec spec = {
        .it_value    = { .tv_sec = 0, .tv_nsec = 100 * 1000 * 1000 },
        .it_interval = { .tv_sec = 0, .tv_nsec = 100 * 1000 * 1000 },
    };
    must(timerfd_settime(tfd, 0, &spec, NULL), "timerfd_settime");

    /* 3. An eventfd used as a cross-thread wakeup counter. */
    int efd = must(eventfd(0, EFD_CLOEXEC | EFD_NONBLOCK), "eventfd");

    /* 4. One epoll instance watches all three synthetic fds at once. */
    int ep = must(epoll_create1(EPOLL_CLOEXEC), "epoll_create1");
    struct epoll_event ev;
    ev.events = EPOLLIN;
    ev.data.fd = sfd; must(epoll_ctl(ep, EPOLL_CTL_ADD, sfd, &ev), "add sfd");
    ev.data.fd = tfd; must(epoll_ctl(ep, EPOLL_CTL_ADD, tfd, &ev), "add tfd");
    ev.data.fd = efd; must(epoll_ctl(ep, EPOLL_CTL_ADD, efd, &ev), "add efd");

    pthread_t th;
    pthread_create(&th, NULL, worker, &efd);

    int ticks = 0, running = 1;
    while (running) {
        struct epoll_event out[8];
        int n = epoll_wait(ep, out, 8, -1);
        if (n == -1) { if (errno == EINTR) continue; must(-1, "epoll_wait"); }

        for (int i = 0; i < n; i++) {
            int fd = out[i].data.fd;
            if (fd == tfd) {
                uint64_t exp;
                if (read(tfd, &exp, sizeof exp) == (ssize_t)sizeof exp) {
                    ticks += (int)exp;
                    printf("timer  : fired (%llu expiration[s], total=%d)\n",
                           (unsigned long long)exp, ticks);
                }
                if (ticks >= 3) {   /* end the demo deterministically */
                    printf("timer  : reached 3 ticks -> raising SIGINT\n");
                    raise(SIGINT);
                }
            } else if (fd == efd) {
                uint64_t cnt;
                if (read(efd, &cnt, sizeof cnt) == (ssize_t)sizeof cnt)
                    printf("eventfd: worker woke the loop (counter was %llu)\n",
                           (unsigned long long)cnt);
            } else if (fd == sfd) {
                struct signalfd_siginfo si;
                if (read(sfd, &si, sizeof si) == (ssize_t)sizeof si) {
                    printf("signal : got %s -> shutting down cleanly\n",
                           strsignal((int)si.ssi_signo));
                    running = 0;
                }
            }
        }
    }

    pthread_join(th, NULL);
    close(sfd); close(tfd); close(efd); close(ep);
    puts("done.");
    return 0;
}

Line by line

  • #define _GNU_SOURCE (line 1): required before includes to expose epoll_create1, signalfd, strsignal, and friends.
  • must() (17-20): a tiny helper — if a syscall returns -1, print perror and exit. Keeps the demo readable without a check after every line; real code would handle errors per call.
  • worker() (25-33): runs in a second thread. It sleeps 250ms, then write(evfd, &one, 8) — this is the cross-thread wakeup. Note the write is exactly 8 bytes; anything else would fail.
  • Block the signal (38-41): build a sigset_t containing SIGINT and sigprocmask(SIG_BLOCK, ...). This must happen before signalfd and before spawning the worker, so the mask is inherited and SIGINT is delivered to us as data, not as a process-killing default action.
  • signalfd(-1, &mask, ...) (43): create the fd that becomes readable when a blocked signal in mask is pending. -1 means "new fd".
  • timerfd_create + timerfd_settime (46-52): a monotonic timer, first firing at 100ms and repeating every 100ms. it_value non-zero arms it; equal it_interval makes it periodic.
  • eventfd(0, ...) (55): counter starts at 0; EFD_NONBLOCK so a spurious read can't hang us.
  • epoll_create1 + three epoll_ctl ADDs (58-63): one epoll instance now watches all three fds. We stash each fd in ev.data.fd so the loop knows which source fired. This is identical to how you'd add a socket — that's the whole point.
  • pthread_create (66): start the worker, passing &efd.
  • The loop (69-101): epoll_wait(ep, out, 8, -1) blocks until at least one fd is ready. EINTR is retried (defensive habit even though our signals go via signalfd).
  • timer branch (76-86): read the uint64_t expiration count and add it (not +1) to ticks, correctly handling coalesced ticks. After 3 ticks we raise(SIGINT) to end the demo deterministically — since SIGINT is blocked, this makes the signalfd readable rather than killing us.
  • eventfd branch (87-91): read+reset the counter; the value is whatever the worker accumulated.
  • signalfd branch (92-98): read one signalfd_siginfo, report the signal name via strsignal, and set running = 0 for a clean exit — done in normal code, so printf is perfectly safe here.
  • Teardown (103-106): join the worker and close() every fd. Each is a real descriptor and must be released.

Common mistakes

1. Reading an eventfd/timerfd with the wrong buffer size

uint32_t got;                 // WRONG: 4 bytes
read(ev, &got, sizeof got);   // returns -1, errno == EINVAL

Why it breaks: these fds require an exact 8-byte (uint64_t) transfer; a 4-byte read is rejected, and if you ignore the error you read garbage.

uint64_t got;
if (read(ev, &got, sizeof got) == 8) { /* use got */ }   // FIXED

2. Creating a signalfd without blocking the signal first

int sfd = signalfd(-1, &mask, SFD_CLOEXEC);   // WRONG: SIGINT not blocked
// Ctrl-C still triggers the default action -> process dies before the loop reads sfd

Why it breaks: signalfd only sees signals that are pending; an unblocked signal takes its normal disposition instead. Your shutdown path never runs.

sigprocmask(SIG_BLOCK, &mask, NULL);          // block first
int sfd = signalfd(-1, &mask, SFD_CLOEXEC);   // FIXED

3. Counting timer ticks by +1 instead of by the expiration count

uint64_t exp; read(tfd, &exp, 8);
ticks += 1;                    // WRONG: loses coalesced expirations

Why it breaks: if the loop was busy, one readable event can represent several expirations; +1 under-counts and the timer effectively drifts slow.

uint64_t exp; read(tfd, &exp, 8);
ticks += exp;                  // FIXED: honor the real count

4. Forgetting to drain (and thinking it's a hot-loop bug)

// handle EPOLLIN on tfd but never read() it

Why it breaks: epoll is level-triggered by default, so an un-drained fd stays readable forever — epoll_wait returns instantly every time and you spin at 100% CPU.

uint64_t exp; read(tfd, &exp, 8);   // FIXED: always consume the payload

Debugging tips

  • See the fds themselves: ls -l /proc/<pid>/fd shows each descriptor. Synthetic fds appear as anon_inode:[eventfd], anon_inode:[timerfd], anon_inode:[signalfd] — great for confirming they were created and not leaked.
  • Watch the syscalls: strace -e trace=eventfd2,timerfd_create,timerfd_settime,signalfd4,epoll_ctl,epoll_wait,read,write ./prog. You'll see the exact flags, the itimerspec you armed, and every read's byte count — a 4-byte read returning EINVAL jumps right out.
  • Timer facts: cat /proc/<pid>/fdinfo/<n> for a timerfd shows the clock, interval, and remaining time — useful to confirm it's actually armed.
  • 100% CPU / instant epoll_wait returns: you forgot to read() (drain) a ready fd; level-triggered epoll keeps reporting it. Add a printf of n and data.fd to see which fd never gets consumed.
  • SIGINT kills the program instead of hitting signalfd: you didn't block it, or you blocked it in main but the signal was delivered to another thread that had it unblocked. Verify the mask with pthread_sigmask in each thread.
  • read returns EAGAIN: normal for *_NONBLOCK fds when nothing is pending — only read after EPOLLIN, and treat EAGAIN as "nothing there", not an error.

Memory safety

  • Fixed payload sizes are a UB trap. Reading a timerfd/eventfd into anything but a uint64_t, or a signalfd into anything smaller than struct signalfd_siginfo, either fails with EINVAL or (if you lie about the size) reads out of bounds. Always use the exact type and sizeof.
  • Descriptors leak like memory. Each fd occupies a slot in the process fd table; failing to close() them in long-lived code exhausts the table (EMFILE). Use *_CLOEXEC so they don't leak across exec, and close explicitly on teardown.
  • Signal masks are per-thread. sigprocmask behavior in a multithreaded program is unspecified for the process; use pthread_sigmask, and block the signals in main before creating threads so every thread inherits the mask. If one thread leaves SIGINT unblocked, the signal may be delivered there and never reach your signalfd.
  • eventfd counter overflow. The kernel refuses a write that would push the counter to UINT64_MAX; on a nonblocking fd that write fails with EAGAIN. Don't ignore the write return value.
  • Sharing an fd across threads: two threads reading the same eventfd race over who gets the count. Prefer one owner (the epoll loop) that reads, and other threads only write.
  • Don't read in a signal handler and via signalfd both — pick one delivery mechanism per signal to avoid confusing, order-dependent behavior.

Real-world uses

  • systemd and most modern init/service managers use signalfd + timerfd + epoll as the backbone of their event loop; child-process reaping, timeouts, and SIGTERM shutdown all flow through fds.
  • Event-loop libraries (libuv — Node.js, libevent, glib's main loop) wrap timerfd/eventfd on Linux so timers and cross-thread "async handles" integrate with the same poll as sockets.
  • Databases and servers (Redis, nginx, PostgreSQL background workers, Envoy) use eventfd as the thread-wakeup primitive and timerfd for internal timeouts.
  • Best practice: one thread owns the epoll loop; other threads communicate only by writing an eventfd (or enqueuing + writing an eventfd to signal "queue non-empty"). Handle SIGINT/SIGTERM/SIGHUP via one signalfd for clean shutdown and config reload. Always create fds *_CLOEXEC, always drain on EPOLLIN, and prefer CLOCK_MONOTONIC for timeouts so a wall-clock adjustment can't make a timer fire early or hang.

Practice tasks

  1. eventfd wakeup: Create an eventfd(0, EFD_CLOEXEC | EFD_NONBLOCK). From main, write the value 5, then read it back and print the counter. Confirm a second read returns EAGAIN (counter reset to 0).
  2. One-shot timer: Create a CLOCK_MONOTONIC timerfd, arm it to fire once at 200ms (it_interval = 0), read the expiration count, and print elapsed time to confirm it was ~200ms. Then verify the fd is no longer readable.
  3. Semaphore eventfd: Create an eventfd with EFD_SEMAPHORE, write 3 once, then loop reading — show that you get exactly three successful reads of 1 before EAGAIN. Explain how this differs from the default mode.
  4. Graceful shutdown via signalfd: Block SIGINT and SIGTERM, create a signalfd, put it in an epoll loop with a 500ms periodic timerfd, and cleanly exit (printing which signal arrived) when either signal is received. Verify Ctrl-C no longer kills the process abruptly.
  5. Coalescing under load: Arm a 50ms periodic timerfd in an epoll loop, but inside the loop sleep/busy-wait 200ms whenever the timer fires. Print the uint64_t expiration count each read and show it exceeds 1 — proving expirations coalesced — and confirm your tick total stays correct only when you add exp rather than +1.

Summary

  • eventfd, timerfd, and signalfd turn cross-thread wakeups, timer expirations, and signals into ordinary readable file descriptors, so a single epoll loop can own all of them alongside sockets.
  • Each has a fixed read payload: eventfd/timerfd = a uint64_t (8 bytes), signalfd = a struct signalfd_siginfo. Reading the wrong size fails with EINVAL.
  • eventfd default read returns+zeros the counter (EFD_SEMAPHORE decrements by 1); timerfd read returns how many times it expired (add exp, don't +1); signalfd requires you to block the signals first (sigprocmask) or they take their default action and can kill the process.
  • Always create with *_CLOEXEC, use *_NONBLOCK with epoll and only read after EPOLLIN, drain ready fds (level-triggered epoll spins otherwise), and close() every fd.
  • These patterns power systemd, libuv/libevent, Redis, nginx, and essentially every serious Linux daemon.

Practice with these exercises