328 lines
20 KiB
Markdown
328 lines
20 KiB
Markdown
# Process lifecycle: signals over IPC
|
||
|
||
**Status: increments 1–4 built** (2026-07-12): claim release on death, exit
|
||
reasons, published exit events, and signals + one-shot timers + the service
|
||
harness are all in — the interface below is as-built. The primitives underneath
|
||
predate this design ([process-management.md](process-management.md):
|
||
spawn, the supervision link, kill, child-exit notifications); this document designs
|
||
the layer above them — the standard vocabulary a danos process speaks about its own
|
||
life, and the stable `runtime.process` interface that carries it. Nothing here is
|
||
device- or driver-specific: a driver, the VFS, and a user application all stop,
|
||
reload, and die the same way. The device manager is simply this design's first
|
||
serious customer ([device-manager.md](device-manager.md)).
|
||
|
||
**"POSIX" in this document means the concepts, never the letter of the standard.**
|
||
danos borrows the ideas and the hard-won lessons (what SIGTERM *means*, why SIGPIPE
|
||
was a mistake) without inheriting the mechanism, the API, or the names. The naming
|
||
rule is danos's own and it is strict: plain words that communicate intent
|
||
(`terminate`, `reload`, `exited`) and the IPC vocabulary the system already speaks
|
||
(`bind`, `subscribe`, `publish`, `endpoint`) — never `SIG*`, never a second word for
|
||
a concept that already has one. Literal POSIX arrives later and lives elsewhere: the
|
||
`std.os.danos` seam that makes danos a Zig target, and eventually a **musl-based C
|
||
layer** on the same native surface (see [zig-self-hosting.md](zig-self-hosting.md)) —
|
||
musl's syscall surface retargeted at danos system calls and IPC protocols (files onto
|
||
the VFS protocol, `sigaction`/`wait` onto this lifecycle, sockets onto whatever
|
||
networking becomes). Ported programs see POSIX; the system underneath never does.
|
||
|
||
## Why a standard vocabulary
|
||
|
||
A supervisor can only manage processes it has never heard of if "please exit" means
|
||
the same thing to all of them. That is the one thing POSIX signals got deeply right:
|
||
`SIGTERM` means the same thing to nginx and to a five-line script, which is why
|
||
process supervision on Unix (init systems, container runtimes) is possible at all.
|
||
danos wants that property from day one, because supervision-and-restart is the
|
||
system's core motivation ([resilience.md](resilience.md)).
|
||
|
||
What POSIX got wrong — for a system like this — is the **delivery mechanism**:
|
||
asynchronous control-flow hijack. A Unix handler runs on a stolen stack at an
|
||
arbitrary instruction boundary, which is why the async-signal-safe function list
|
||
exists, why `errno` must be saved, and why the canonical signal bug is a SIGTERM
|
||
handler innocently calling `printf` mid-`malloc`. That entire bug class comes from
|
||
the mechanism, not the vocabulary, and none of it is worth importing.
|
||
|
||
A microkernel already has the right channel: **a signal is a message.** QNX delivers
|
||
POSIX signals over its message passing; seL4 has notification objects; Erlang turned
|
||
"death is a message to whoever linked" into a reliability philosophy. danos has
|
||
already done it once without naming it: a child's death arrives as a notification
|
||
badge on the supervisor's endpoint — the microkernel's SIGCHLD, the IRQ-as-IPC
|
||
pattern reused. Signals are the same pattern reused a third time.
|
||
|
||
## The mechanism
|
||
|
||
- **`signal_bind(endpoint)`** — a process nominates the endpoint its signals arrive
|
||
on, exactly as `irq_bind` nominates where a device's interrupts land. The runtime
|
||
does this at startup for any program that opts in.
|
||
- **`process_signal(id, signal)`** — posts the signal as an asynchronous
|
||
notification to the target's bound endpoint: badge = `notify_badge_bit |
|
||
notify_signal_bit | pending signals`. Non-blocking for the sender, always.
|
||
- **Pending signals coalesce** in a per-process bitmask until the target next waits
|
||
— exactly like interrupt notifications, and exactly POSIX's own semantics for
|
||
non-realtime signals (two pending SIGTERMs are one SIGTERM). The bitmask *is* the
|
||
design: signals carry no payload. Anything with a payload is a protocol message.
|
||
- **Authority**: the supervisor may signal its children — the same link that is
|
||
already the kill authority. A process may signal itself. Anything broader waits
|
||
for transferable process handles.
|
||
- **No binding, no problem**: a process that never calls `signal_bind` is not
|
||
broken — its signals pend unread and only `process_kill` works on it. Simple
|
||
programs stay simple; the vocabulary is opt-in, the kill authority is not.
|
||
|
||
Because delivery is a message into the process's own event loop, there is no
|
||
async-signal-safe list in danos: a handler is ordinary code running at a point the
|
||
process chose. The bug class is gone by construction, not by discipline.
|
||
|
||
## The vocabulary: POSIX.1-1990, sorted honestly
|
||
|
||
The full 1990 set, and what each becomes. Two intrinsically problematic cases get a
|
||
defense below the table.
|
||
|
||
| POSIX.1-1990 | danos disposition | Notes |
|
||
|---|---|---|
|
||
| SIGTERM | signal `terminate` | finish up and exit; the supervisor's polite half |
|
||
| SIGHUP | signal `reload` | re-read configuration / re-scan |
|
||
| SIGINT | signal `interrupt` | interactive interrupt; meaningful once a console can send it, in the vocabulary now so numbering is stable |
|
||
| SIGQUIT | signal `quit` | as SIGINT, without the core-dump baggage |
|
||
| SIGALRM | signal `alarm` | timer expiry as a message; the Unix SIGALRM+`longjmp` timeout hacks are impossible here. In the vocabulary, unbuilt: no consumer yet, and when one appears it is runtime sugar over the existing timer — zero kernel work |
|
||
| SIGUSR1, SIGUSR2 | signals `user_1`, `user_2` | service-defined |
|
||
| SIGCHLD | **already exists** — the exit notification | the badge carries the child id, dodging the classic coalescing bug (Unix code must loop `waitpid`) |
|
||
| SIGKILL | `process_kill` — kernel mechanism | its definition is "cannot be handled"; it was never really a signal |
|
||
| SIGABRT | exit reason `abort` | `abort()` is synchronous self-termination, not an event |
|
||
| SIGSEGV, SIGILL, SIGFPE | exit reasons, **never delivered** | see below |
|
||
| SIGPIPE | **an error return**, not a signal | see below |
|
||
| SIGSTOP, SIGTSTP, SIGTTIN, SIGTTOU, SIGCONT | deferred | job control needs terminals, sessions, and process groups; stop/continue is scheduler territory |
|
||
|
||
**The fault signals (SIGSEGV, SIGILL, SIGFPE) are intrinsically wrong for messages.**
|
||
They are *synchronous* — raised at a specific faulting instruction, not "sometime
|
||
soon". A message cannot be delivered to a process whose next instruction re-faults;
|
||
it never reaches its event loop to read it. POSIX only makes fault handlers "work"
|
||
via the async hijack (run the handler *instead of* the instruction), and even there,
|
||
returning from a SIGSEGV handler without curing the cause is undefined behavior.
|
||
danos's architecture already has the better answer: fault → the kernel kills the
|
||
process ([resilience.md](resilience.md) step 2, built) → the supervisor reads the
|
||
reason → restart. Recovery is restart, not a handler. This is also truer to the 1990
|
||
standard than handling is: the standard's default action for all three was
|
||
"terminate the process".
|
||
|
||
**SIGPIPE deserves special contempt.** Its default kills a process that writes to a
|
||
closed pipe — which is why "the whole server died because one client disconnected"
|
||
is roughly every network daemon's first production bug, and why every mature codebase
|
||
contains the same fix: ignore SIGPIPE, handle the `EPIPE` error return. danos made
|
||
the right choice natively already — a reply owed to a dead peer fails with `-EPEER`.
|
||
Errors from operations are error returns from those operations. The posix layer can
|
||
synthesize SIGPIPE for ported code that expects it.
|
||
|
||
### Statements, not questions
|
||
|
||
A signal and a protocol message both travel over IPC — the difference is the
|
||
**contract**, not the transport. danos IPC has two primitives, both already in
|
||
daily use: the **asynchronous notification** (a badge — bits that coalesce into a
|
||
pending mask; the sender never blocks; no payload, *no reply path*; how IRQs and
|
||
exit events arrive) and the **synchronous call** (a rendezvous — payload both
|
||
ways, the caller waits for the reply; how VFS requests work). A signal is the
|
||
first kind: a *statement*. `terminate` wants no reply — the exit notification is
|
||
its acknowledgement.
|
||
|
||
A health probe is the second kind: a *question*, worthless without its answer —
|
||
and the answer's absence within a deadline is the very thing being measured.
|
||
Asked as a signal it has no reply channel (a coalescing bit can't carry an answer,
|
||
and the authority rule forbids a child signalling its supervisor back); asked as a
|
||
call, the timeout-is-the-diagnosis semantics come free. So there is no `health`
|
||
signal. Liveness is the common **`ping`**: a reserved request every harness-run
|
||
service answers automatically on its main endpoint — still free for the service
|
||
author, still one obvious way — and a supervisor's probe is a `ping` call with a
|
||
deadline.
|
||
|
||
## The two iron rules
|
||
|
||
1. **Cleanup is the kernel's job.** A process can die with no warning — fault,
|
||
kill, power. Correctness must never depend on a `terminate` handler running. On
|
||
any death the kernel releases the address space, IPC handles, IRQ bindings, and
|
||
owed replies (built), and must also release **device, I/O-port, and interrupt
|
||
claims and MSI vectors** (the known gap in
|
||
[process-management.md](process-management.md); increment 1). A signal handler is
|
||
for *graceful* work — flushing, deregistering, saving — never for *necessary*
|
||
work.
|
||
2. **Kill is not a signal, and exit reasons are load-bearing.** The standard stop
|
||
sequence is *terminate → deadline → `process_kill`*; the unhandleable kill stays
|
||
a kernel mechanism. And a supervisor deciding whether to restart must know *how*
|
||
the child died: clean exit (meant to — don't restart), fault (restart with
|
||
backoff), killed (the supervisor did it). The exit notification today carries
|
||
only the id; it grows a reason. Restart policy cannot be written without it.
|
||
|
||
## Who learns of a death
|
||
|
||
A death has three audiences, and conflating them is how systems end up with either
|
||
zombie state or privileged snooping:
|
||
|
||
1. **The supervisor** — gets the exit notification on the endpoint it gave at spawn
|
||
(built), which grows the `ExitReason` (increment 2). The supervisor is the only
|
||
audience that needs the *reason*, because it is the only one deciding whether to
|
||
restart.
|
||
2. **The peer owed a reply** — already built: a client that dies mid-request fails
|
||
the server's reply with `-EPEER`; a server that dies fails its waiting clients
|
||
the same way. This covers the *synchronous* case only.
|
||
3. **The subscribers** — the new piece, and it is the input service's
|
||
publish/subscribe shape ([input.md](input.md)) applied to exits. A stateful
|
||
service accumulates per-client state across many requests: the VFS holds a dead
|
||
client's open file handles, the input service holds its subscriptions, a future
|
||
network stack holds its sockets. None of these are the client's supervisor, and
|
||
none learn anything from a failed reply if the client simply never calls again.
|
||
So the kernel **publishes every exit** to whoever subscribed:
|
||
`process_subscribe(endpoint)` adds a subscriber, and each death posts a
|
||
notification to every subscriber (badge = `notify_exit_bit | process id` — the
|
||
same encoding supervisors already decode, the IRQ-as-IPC pattern once more). The
|
||
subscriber filters for ids it holds state for and releases what the dead client
|
||
held. Correlating is free of bookkeeping: an IPC sender's badge already *is* its
|
||
task id (`runtime.ipc.Received`), so the id a service has been keying client
|
||
state by all along is the id the exit event carries.
|
||
|
||
Subscription, not broadcast-to-everyone: only processes that asked receive
|
||
events, the kernel keeps a bounded subscriber table, and delivery is the same
|
||
non-blocking coalescing notification as everything else — a dying process never
|
||
waits on its mourners. Subscribing is ungated, like `process_enumerate`: what is
|
||
running (and dying) is not a secret between cooperating processes. Subscribers
|
||
do not receive the exit reason — the VFS does not care *why* the client died.
|
||
|
||
This is the service-side mirror of iron rule 1: **a service must never depend on
|
||
its clients cleaning up after themselves.** Handle release on client death is the
|
||
service's job, triggered by the published exit event — never by a courtesy
|
||
"closing now" message that a crashed client will never send.
|
||
|
||
## The stable interface: `runtime.process`
|
||
|
||
`runtime.process` already owns what a process receives at birth (`Init`, the
|
||
argv contract). It grows to own the other end of life.
|
||
|
||
**The runtime is the stable interface; the numbers are not.** danos applications do
|
||
not make system calls — they call the runtime library, and the system-call numbers,
|
||
notification bits, and signal bit positions beneath it are a **private kernel ↔
|
||
runtime contract** that may change at any time (settled 2026-07-12). This is why
|
||
the runtime exists. Today kernel and runtime ship from one tree in one image, so
|
||
"stability" is simply building them together. When driver binaries start shipping
|
||
as separately-versioned applications — the whole point of the restart design — the
|
||
binary's embedded runtime version becomes compatibility metadata (the same idea as
|
||
the protocol version in the device manager's `hello`), and the kernel refuses what
|
||
it cannot serve. Signals therefore need no reserved numbering scheme: the enum
|
||
below is vocabulary, not ABI.
|
||
|
||
```zig
|
||
/// The signal vocabulary. The value is the bit position in the pending mask — a
|
||
/// private kernel/runtime detail, free to change while they ship together.
|
||
pub const Signal = enum(u5) {
|
||
terminate = 0, // SIGTERM: finish up and exit
|
||
reload = 1, // SIGHUP: re-read configuration
|
||
interrupt = 2, // SIGINT
|
||
quit = 3, // SIGQUIT
|
||
alarm = 4, // SIGALRM
|
||
user_1 = 5, // SIGUSR1
|
||
user_2 = 6, // SIGUSR2
|
||
};
|
||
|
||
/// A decoded pending mask: the coalesced set of signals a notification delivered.
|
||
pub const SignalSet = struct {
|
||
pending: u32,
|
||
pub fn has(set: SignalSet, signal: Signal) bool { ... }
|
||
pub fn iterate(set: SignalSet) Iterator { ... }
|
||
};
|
||
|
||
/// Nominate `endpoint` as this process's signal endpoint (signal_bind). The
|
||
/// runtime's service harness calls this; a bare program may call it directly and
|
||
/// fold signals into its own replyWait loop.
|
||
pub fn bindSignals(endpoint: usize) bool { ... }
|
||
|
||
/// Decode a received badge into signals, or null if the badge is not a signal
|
||
/// notification (mirrors ipc.Received.isChildExit).
|
||
pub fn signalsFrom(badge: usize) ?SignalSet { ... }
|
||
|
||
/// Send `signal` to process `id`. Supervisor-gated, like kill; non-blocking.
|
||
pub fn sendSignal(id: u32, signal: Signal) bool { ... }
|
||
|
||
/// The standard stop sequence: terminate, wait up to `deadline_ms` for the exit
|
||
/// notification, then process_kill. The one call a supervisor needs.
|
||
pub fn stop(id: u32, deadline_ms: u64) void { ... }
|
||
|
||
/// Subscribe `endpoint` to published exit events (process_subscribe). Every
|
||
/// process death posts an asynchronous notification: badge = notify_exit_bit |
|
||
/// process id — the same encoding a supervisor's exit notification uses, decoded
|
||
/// by the same ipc.Received helpers. For stateful services: release what the dead
|
||
/// client held (file handles, subscriptions, sockets). Ungated, like
|
||
/// process_enumerate.
|
||
pub fn subscribeExits(endpoint: usize) bool { ... }
|
||
|
||
/// How a process ended — queried after the exit notification (the kernel records
|
||
/// it first, so the two never race). What restart policy reads. (Built in M17.2.)
|
||
pub const ExitReason = enum(u8) {
|
||
exited, // returned from main / clean exit
|
||
aborted, // abort() — deliberate self-termination (SIGABRT's ghost; reserved)
|
||
segmentation_fault, // SIGSEGV's ghost
|
||
illegal_instruction, // SIGILL's ghost
|
||
arithmetic_fault, // SIGFPE's ghost
|
||
protection_fault, // general protection fault
|
||
fault, // any other CPU exception
|
||
killed, // process_kill
|
||
};
|
||
```
|
||
|
||
Two deliberate absences. There is no `mask`/`block` API — a process that is not
|
||
ready for a signal simply has not waited on its endpoint yet; the pending mask *is*
|
||
the blocked set. And there is no per-signal handler registration at this layer —
|
||
dispatch is the process's own `switch` over `SignalSet`, or the service harness's
|
||
callbacks (`on_terminate`, `on_reload`) for programs that want defaults.
|
||
|
||
### The service harness
|
||
|
||
`runtime.service` owns the `replyWait` loop and folds every event source — signals,
|
||
child exits, protocol messages — into callbacks, with the vocabulary's defaults:
|
||
`terminate` returns from the loop (clean exit), the common `ping` is answered automatically,
|
||
`reload` is ignored unless overridden. One loop, no locking, nothing reentrant. A
|
||
service author writes domain logic; the lifecycle contract is satisfied by the
|
||
harness. A process that bypasses the harness and ignores its signals meets the
|
||
deadline-then-kill escalation — you cannot force a process to implement an
|
||
interface, but you can make compliance free and non-compliance fatal.
|
||
|
||
### The musl layer later
|
||
|
||
The POSIX C layer is a **musl port**: musl's arch/syscall layer retargeted so that
|
||
what musl believes are kernel syscalls become danos runtime calls and IPC — `open`
|
||
and `read` onto the VFS protocol, `kill`/`sigaction`/`waitpid` onto this document's
|
||
vocabulary, `exit` onto the runtime's exit path. `sigaction` handlers registered
|
||
through it are invoked by the runtime's loop when the signal message arrives —
|
||
synchronous underneath, async-looking to ported code, delivered at wait boundaries
|
||
the way most Unix programs already experience signals (at syscalls). No stack hijack
|
||
ever happens, `SA_RESTART` semantics come free because nothing was interrupted, and
|
||
SIGPIPE can be synthesized from `-EPEER` for the programs that expect it. C programs
|
||
get POSIX; danos-native programs never pay for it.
|
||
|
||
## Increments
|
||
|
||
1. **Kernel: release device/port/IRQ claims and MSI vectors on death** — the
|
||
cleanup half of iron rule 1, and the prerequisite for any restart story. Test:
|
||
kill a claiming driver, spawn it again, the claim succeeds.
|
||
2. **Exit reason in the death notification** (`ExitReason` above).
|
||
3. **Exit events**: `process_subscribe` in the kernel (bounded subscriber table,
|
||
publishes on every death), `runtime.process.subscribeExits`; the VFS becomes the
|
||
first subscriber — releasing a dead client's handles is its proof test.
|
||
4. **Signals**: `signal_bind` + `process_signal` + the pending mask in the kernel;
|
||
`runtime.process` grows the interface above; the service harness handles
|
||
`terminate` and answers the common `ping`; `stop()` for supervisors.
|
||
|
||
[device-manager.md](device-manager.md) builds directly on all four.
|
||
|
||
## Settled questions (2026-07-12)
|
||
|
||
- **Signal numbering is not ABI**: the runtime is the stable interface; the numbers
|
||
beneath it are a private kernel ↔ runtime contract (see "The stable interface").
|
||
- **Liveness is a `ping` call, not a signal**: signals are statements, questions
|
||
are synchronous calls (see "Statements, not questions"). A service wanting *deep*
|
||
health ("can I reach my hardware?") defines its own protocol message on top.
|
||
- **Process handles: deferred.** Pids + the supervisor gate cover everything
|
||
planned; transferable handles (Fuchsia-style, delegating signalling without
|
||
delegating kill) wait for the capability table to grow types beyond endpoints.
|
||
- **`alarm`: in the vocabulary, unbuilt.** No consumer yet; when one appears it is
|
||
runtime sugar over the existing timer (arm a timer that posts your own signal) —
|
||
zero kernel work, so deferring costs nothing.
|
||
- **Subscription granularity: all exits**, subscriber-side filtering — one
|
||
subscription per service, a bounded kernel table. Per-id subscriptions only if
|
||
event volume ever matters (hundreds of processes, not before).
|
||
- **Client identity across the exit boundary: no convention needed** — an IPC
|
||
sender's badge already is its task id (see "Who learns of a death").
|