docs: full docs-vs-code audit — fix every stale claim across 40 docs

Every doc verified claim-by-claim against the code by parallel audit agents,
then fixed and adversarially re-verified. Two waves of staleness corrected:
the originally audited findings (higher-half boot handoff, kernel VFS
takeover, fault isolation + claim release + driver restart, AML/S5 moving to
ring 3, threading's shipped design, USB+FAT landing) and a second pass of
adjacent claims the verifiers caught (smp.md 'not built yet' intro,
system-requirements' PS/2-only and no-storage claims, halting.md's red-panic
and no-IDT text, testing.md's serial mirroring, router-era vfs-protocol
wording, capsule-first boot loading).

threading.md now documents the shared-fate gap explicitly: the design says a
process dies whole, the kernel today kills only the offending thread.

Also fixes three stale code comments (isr.s exceptionHandler, acpi.zig
sleepValue, build.zig boot-volume) — comments only, no behavior change.
This commit is contained in:
Daniel Samson
2026-07-22 09:09:53 +01:00
parent 52df2ba6f6
commit e854f65623
43 changed files with 821 additions and 546 deletions
+36 -24
View File
@@ -55,10 +55,14 @@ pattern reused. Signals are the same pattern reused a third time.
- **`process_signal(id, signal)`** — posts the signal as an asynchronous
notification to the target's bound endpoint: badge = `notify_badge_bit |
notify_signal_bit | pending signals`. Non-blocking for the sender, always.
- **Pending signals coalesce** in a per-process bitmask until the target next waits
— exactly like interrupt notifications, and exactly POSIX's own semantics for
non-realtime signals (two pending SIGTERMs are one SIGTERM). The bitmask *is* the
design: signals carry no payload. Anything with a payload is a protocol message.
- **Pending signals coalesce** in a per-process bitmask while the target has no
signal endpoint bound, and the whole mask arrives as one notification at bind —
POSIX's own semantics for non-realtime signals (two pending SIGTERMs are one
SIGTERM). Once bound, each `process_signal` flushes the mask straight into the
endpoint's notification ring, so a busy receiver drains separate posts as
separate notifications — harmless, because the badge is a set of bits, never a
count. The bitmask *is* the design: signals carry no payload. Anything with a
payload is a protocol message.
- **Authority**: the supervisor may signal its children — the same link that is
already the kill authority. A process may signal itself. Anything broader waits
for transferable process handles.
@@ -85,7 +89,7 @@ defense below the table.
| SIGUSR1, SIGUSR2 | signals `user_1`, `user_2` | service-defined |
| SIGCHLD | **already exists** — the exit notification | the badge carries the child id, dodging the classic coalescing bug (Unix code must loop `waitpid`) |
| SIGKILL | `process_kill` — kernel mechanism | its definition is "cannot be handled"; it was never really a signal |
| SIGABRT | exit reason `abort` | `abort()` is synchronous self-termination, not an event |
| SIGABRT | exit reason `aborted` | synchronous self-termination is an exit, not an event; recorded for any nonzero exit code |
| SIGSEGV, SIGILL, SIGFPE | exit reasons, **never delivered** | see below |
| SIGPIPE | **an error return**, not a signal | see below |
| SIGSTOP, SIGTSTP, SIGTTIN, SIGTTOU, SIGCONT | deferred | job control needs terminals, sessions, and process groups; stop/continue is scheduler territory |
@@ -135,18 +139,21 @@ deadline.
1. **Cleanup is the kernel's job.** A process can die with no warning — fault,
kill, power. Correctness must never depend on a `terminate` handler running. On
any death the kernel releases the address space, IPC handles, IRQ bindings, and
owed replies (built), and must also release **device, I/O-port, and interrupt
claims and MSI vectors** (the known gap in
[process-management.md](process-management.md); increment 1). A signal handler is
any death the kernel releases the address space, IPC handles, IRQ bindings,
owed replies, and **device, I/O-port, and interrupt claims and MSI vectors** —
the last of these was once the known gap in
[process-management.md](process-management.md), closed by increment 1
(`releaseTaskResourcesLocked`, on every death path). A signal handler is
for *graceful* work — flushing, deregistering, saving — never for *necessary*
work.
2. **Kill is not a signal, and exit reasons are load-bearing.** The standard stop
sequence is *terminate → deadline → `process_kill`*; the unhandleable kill stays
a kernel mechanism. And a supervisor deciding whether to restart must know *how*
the child died: clean exit (meant to — don't restart), fault (restart with
backoff), killed (the supervisor did it). The exit notification today carries
only the id; it grows a reason. Restart policy cannot be written without it.
backoff), killed (the supervisor did it). The exit notification carries only the
id; the reason is recorded before the notification posts and read with the
supervisor-gated `process_exit_reason` query. Restart policy cannot be
written without it.
## Who learns of a death
@@ -154,7 +161,8 @@ A death has three audiences, and conflating them is how systems end up with eith
zombie state or privileged snooping:
1. **The supervisor** — gets the exit notification on the endpoint it gave at spawn
(built), which grows the `ExitReason` (increment 2). The supervisor is the only
(built), then reads the `ExitReason` with the `process_exit_reason` query
(increment 2). The supervisor is the only
audience that needs the *reason*, because it is the only one deciding whether to
restart.
2. **The peer owed a reply** — already built: a client that dies mid-request fails
@@ -162,10 +170,11 @@ zombie state or privileged snooping:
the same way. This covers the *synchronous* case only.
3. **The subscribers** — the new piece, and it is the input service's
publish/subscribe shape ([input.md](input.md)) applied to exits. A stateful
service accumulates per-client state across many requests: the VFS holds a dead
client's open file handles, the input service holds its subscriptions, a future
network stack holds its sockets. None of these are the client's supervisor, and
none learn anything from a failed reply if the client simply never calls again.
service accumulates per-client state across many requests: a filesystem server
(FAT today) holds a dead client's open file handles, the input service holds
its subscriptions, a future network stack holds its sockets. None of these
are the client's supervisor, and none learn anything from a failed reply if
the client simply never calls again.
So the kernel **publishes every exit** to whoever subscribed:
`process_subscribe(endpoint)` adds a subscriber, and each death posts a
notification to every subscriber (badge = `notify_exit_bit | process id` — the
@@ -180,7 +189,8 @@ zombie state or privileged snooping:
non-blocking coalescing notification as everything else — a dying process never
waits on its mourners. Subscribing is ungated, like `process_enumerate`: what is
running (and dying) is not a secret between cooperating processes. Subscribers
do not receive the exit reason — the VFS does not care *why* the client died.
do not receive the exit reason — the filesystem server does not care *why*
the client died.
This is the service-side mirror of iron rule 1: **a service must never depend on
its clients cleaning up after themselves.** Handle release on client death is the
@@ -221,7 +231,6 @@ pub const Signal = enum(u5) {
pub const SignalSet = struct {
pending: u32,
pub fn has(set: SignalSet, signal: Signal) bool { ... }
pub fn iterate(set: SignalSet) Iterator { ... }
};
/// Nominate `endpoint` as this process's signal endpoint (signal_bind). The
@@ -231,14 +240,15 @@ pub fn bindSignals(endpoint: usize) bool { ... }
/// Decode a received badge into signals, or null if the badge is not a signal
/// notification (mirrors ipc.Received.isChildExit).
pub fn signalsFrom(badge: usize) ?SignalSet { ... }
pub fn signalsFrom(badge: u64) ?SignalSet { ... }
/// Send `signal` to process `id`. Supervisor-gated, like kill; non-blocking.
pub fn sendSignal(id: u32, signal: Signal) bool { ... }
/// The standard stop sequence: terminate, wait up to `deadline_ms` for the exit
/// notification, then process_kill. The one call a supervisor needs.
pub fn stop(id: u32, deadline_ms: u64) void { ... }
/// notification on `exit_endpoint` (the endpoint the child was spawned with),
/// then process_kill. The one call a supervisor needs.
pub fn stop(id: u32, deadline_ms: u64, exit_endpoint: usize) void { ... }
/// Subscribe `endpoint` to published exit events (process_subscribe). Every
/// process death posts an asynchronous notification: badge = notify_exit_bit |
@@ -252,7 +262,7 @@ pub fn subscribeExits(endpoint: usize) bool { ... }
/// it first, so the two never race). What restart policy reads. (Built in M17.2.)
pub const ExitReason = enum(u8) {
exited, // returned from main / clean exit
aborted, // abort() — deliberate self-termination (SIGABRT's ghost; reserved)
aborted, // deliberate failure exit — any nonzero exit code (SIGABRT's ghost)
segmentation_fault, // SIGSEGV's ghost
illegal_instruction, // SIGILL's ghost
arithmetic_fault, // SIGFPE's ghost
@@ -299,8 +309,10 @@ get POSIX; danos-native programs never pay for it.
kill a claiming driver, spawn it again, the claim succeeds.
2. **Exit reason in the death notification** (`ExitReason` above).
3. **Exit events**: `process_subscribe` in the kernel (bounded subscriber table,
publishes on every death), `runtime.process.subscribeExits`; the VFS becomes the
first subscriber — releasing a dead client's handles is its proof test.
publishes on every death), `runtime.process.subscribeExits`; the userspace VFS
router was the first subscriber — releasing a dead client's handles was its
proof test — and the FAT server inherited the role when the router moved into
the kernel (clients now hold the filesystem server's node ids directly).
4. **Signals**: `signal_bind` + `process_signal` + the pending mask in the kernel;
`runtime.process` grows the interface above; the service harness handles
`terminate` and answers the common `ping`; `stop()` for supervisors.