From e78810195dbd7c2a244701b2c19d30abb619cddd Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Wed, 22 Jul 2026 10:12:42 +0100 Subject: [PATCH] =?UTF-8?q?docs:=20shared-fate=20plan=20=E2=80=94=20approv?= =?UTF-8?q?ed=20design=20for=20whole-process=20death?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Red-teamed draft: leader id (tgid-style), AddressSpaceRef dying latch, killGroupLocked fan-out, group notification at address-space destruction, shared-memory mapping refs. Leader thread_exit refused with -EPERM (decided at sign-off). Milestones M1-M4. --- docs/shared-fate-plan.md | 291 +++++++++++++++++++++++++++++++++++++++ 1 file changed, 291 insertions(+) create mode 100644 docs/shared-fate-plan.md diff --git a/docs/shared-fate-plan.md b/docs/shared-fate-plan.md new file mode 100644 index 0000000..f03a002 --- /dev/null +++ b/docs/shared-fate-plan.md @@ -0,0 +1,291 @@ +# Shared fate: whole-process death (plan) + +**Status: approved 2026-07-22; leader `thread_exit` → `-EPERM`. Implementation in +progress (M1–M4 below).** + +[threading.md](threading.md) promises that a process dies *whole* — a fault in any +thread, or a kill, takes down every thread. The kernel doesn't do that yet: every +death path (`exit`, `thread_exit`, a ring-3 fault, `process_kill`) tears down +exactly one `Task`, and the address-space refcount keeps the space alive for the +siblings — so a faulting worker orphans its threads, which keep running in the +possibly-corrupted address space (`process.zig` `killCurrentProcess`; +`scheduler.zig` `exitUserLocked`). This plan closes that gap the way Linux, Windows, +and Fuchsia all did: **the process is the unit of fate; only a voluntary +`thread_exit` is per-thread.** The supervisor contract — one exit notification, +then `process_exit_reason` — is deliberately unchanged. + +## The contract + +| Event | Who dies | Reason the supervisor reads (leader's record) | +|---|---|---| +| CPU fault in **any** thread (recoverable vector) | the whole group | the fault class (`segmentation_fault`, …) | +| `process_kill` on **any member id** | the whole group | `killed` | +| `exit(code)` from **any** thread | the whole group | `exited` (0) / `aborted` (≠0) | +| `thread_exit` from a worker | that worker only | — (per-task record: `exited`) | +| `thread_exit` from the **leader** | nobody — refused, `-EPERM` (decided, below) | — | +| NMI, double fault, machine check | the core halts (unchanged) | — | + +`exit` gaining group semantics is the `exit_group` lesson from Linux: the runtime's +main-return path calls `exit`, and a process whose main returned must not leave +workers running. `thread_exit` (what the worker trampoline calls) keeps today's +per-thread behavior, refcount and all. + +**The leader-`thread_exit` rule (decided: refuse).** The syscall is reachable from +the leader even though the runtime never does it. Options weighed: (i) **refuse +with `-EPERM`** — cheapest and honest; the group ends only through +`exit`/fault/kill; (ii) escalate to `exit(0)` — Linux-flavored, but silently turns +a (buggy) library call into process death; (iii) a Linux-style zombie leader whose +slot survives until the group ends — the most faithful, and by far the most +machinery. **(i) chosen at sign-off**; rows and tests below follow it. + +## Group identity: a leader id + +Nothing on `Task` names a process today — threads are tied to their process only by +an equal `address_space`, their name is `"thread"`, and their `supervisor` is +whichever *task* spawned them (possibly another worker), so supervision links form a +chain, not a group. Scanning by `address_space` is also fragile during teardown, +because both death paths zero it. + +So: **add `leader: u32` to `Task`** — the Linux tgid, in danos clothes. +`spawnProcessSupervised` sets `leader = own id`; `spawnThreadSupervised` copies the +*caller's* leader; kernel tasks keep `leader = 0`, which is never followed. The +leader id is exactly the id `system_spawn` returned to the supervisor, so the +outside world already speaks it. Group membership = equal `leader`, where a *live +member* means `state` ∈ {`.ready`, `.blocked`, `.running`} — the same filter +`taskByIdLocked` applies; `.reaping` corpses are excluded. While the field is being +introduced, add `leader` to `ProcessDescriptor` too (the ABI is private, so this is +cheap now and lets `process_enumerate` consumers group threads). + +`process_kill` re-derives its authority through the leader — with the existing +guard order preserved: a kernel task (`address_space == 0`) is `-ESRCH` *before* +any leader resolution (the kernel test asserts exactly that). Then: resolve the +target, follow `target.leader`, require `leader.supervisor == caller`. The kill +capability becomes per-*process*, aimed at any member id, and the odd +today-behavior where a worker can be individually killed by its spawning thread +disappears. (Audited: nothing in-tree kills a worker tid or relies on +thread-supervisor kill semantics.) + +## The group-dying latch + +`AddressSpaceRef` — one per space, refcounted by its member tasks, recycled with a +full struct re-init — gains the group-death state: + +```zig +dying: bool = false, // set by the first trigger; never cleared +group_reason: abi.ExitReason, // what the leader's record will say +exit_endpoint: ?*ipc.Endpoint, // the leader's counted ref, moved here +leader: u32, +``` + +The latch answers three attacks the red team confirmed against a latch-free +design: + +- **The spawn gate.** A member already *inside* `thread_spawn` on another core + when the fan-out runs (it passed the syscall-entry `kill_pending` check, then + spun on the BKL) would otherwise complete the spawn after the fan-out's lock + hold ends — a fresh, uncondemned member that escapes the kill and, worse, holds + a space reference that keeps the group-death hook from ever firing. Fix: + `retainAddressSpace` (equivalently `spawnUserLocked`) **refuses a dying + space**; the in-flight `thread_spawn` fails with `-ESRCH` under the same lock + that would have created the member. +- **Concurrent triggers.** A second member faulting (or exiting) on another core + while the first fan-out runs must not re-run the fan-out, double-bump + `fault_kill_count`, or re-stamp reasons. Every kill path checks the latch first: + already dying → skip straight to `terminateCurrentLocked`, no stamp, no count. + First trigger wins, deterministically. `exit_reason` and `fault_kill_count` + writes move under the BKL as part of this. +- **Notification ownership.** The leader's `exit_endpoint` is a counted + birth-to-death reference dropped at notify time. The stamp pass **moves** that + reference onto the `AddressSpaceRef` and nulls `Task.exit_endpoint` in the same + hold, so the leader's own `releaseTaskResourcesLocked` sees null (no early + notify, no double drop); the group-death hook notifies and drops exactly once. + +## The fan-out: `killGroupLocked` + +One new function in `process.zig`, running under a **single BKL hold** (built from +the `*Locked` primitives — the lock is non-recursive, and `terminateCurrentLocked` +never returns, which forces the shape): + +``` +killGroupLocked(leader: u32, reason: ExitReason, trigger: ?*Task) + 0. Latch: AddressSpaceRef.dying = true, stash {reason, leader, + leader's exit_endpoint (moved)}. + 1. Stamp pass: the LEADER's exit_reason = reason — the leader's + record is the one the supervisor can read, so it carries the + group reason even when the trigger is a worker. The trigger + also keeps `reason` (its own record tells the truth); every + other live member gets .killed. All members get kill_pending. + Stamping precedes any teardown, because recordExitLocked + snapshots the reason first thing. + 2. Reap pass, to fixpoint: reap every member in .ready or .blocked + via reapTaskLocked, re-reading Task.state each iteration — a + member's teardown can wake another member (-EPEER wakes, joiner + wakes), flipping it .blocked → .ready behind the scan cursor. + Terminates in ≤ one pass per member: the scrub calls in + releaseTaskResourcesLocked (abandonSenderLocked, + removeFromWaitQueueLocked, forgetIpcClientLocked, + killOwnedEndpointsLocked) run before destroy, so no wake path + holds a pointer to a reaped member. + 3. Members .running on other cores stay condemned (kill_pending); + a condemned member dies at its next syscall entry, at its own + core's next tick while in user mode, or — once it blocks or is + preempted — at any core's next tick reap. There is no kill IPI. + (The entry check reads kill_pending unlocked; benign on + x86-TSO — a missed read is caught by the next delivery point — + but make the field atomic when touching it.) + 4. If the current task is a member (fault, exit, in-group kill): + terminateCurrentLocked, last, because it switches away and the + reap paths free the kernel stack being stood on. + If the caller is outside the group (supervisor kill): return. +``` + +The invariants this preserves, each load-bearing today: + +- **Only `.ready`/`.blocked` tasks are reaped synchronously.** A member running on + another core can only be condemned — it tears itself down after switching CR3 + off the dying page tables (the stack it stands on is freed later by the reap + list), and its address-space reference protects the page tables its CR3 still + points at. Force-destroying the space under a running sibling is the one + unrecoverable mistake available here. +- **The refcount decides when the space dies.** Reaping N members drops N + references; the last drop — possibly on a condemned sibling's core, a tick + later — destroys the space. No path forces it. +- **`fault_kill_count` bumps once per group**, not per member (`fault-recovery` + asserts `== 1` exactly); the latch is what enforces this under racing faults. +- **Per-tid resource sweeps stay per-tid.** Each member's + `releaseTaskResourcesLocked` releases what *that tid* owns — claims, GSI/MSI + bindings, registered endpoints, handles. That keying is correct under shared + fate (and is today's hazard: a lone worker death already yanks its claims out + from under live siblings). A worker that *does* carry an `exit_endpoint` (the + ABI allows it; the runtime passes `no_cap`) keeps today's per-task posting at + its own teardown — only the leader's notification moves. + +## When is the group dead? The notification + +Today each task posts its own exit notification as the *last* step of its release, +so a supervisor observes a fully-released child. For a group that guarantee must +hold for the **whole group**: if the leader's notification fires while a condemned +sibling still runs on another core, the device manager can respawn the driver into +a claim conflict with a not-yet-dead sibling. + +The clean fix falls out of the refcount: **the group is dead exactly when the +address space is destroyed.** `releaseAddressSpace`'s last-drop path calls a new +`group_exit_hook` (the scheduler already calls up through hooks — +`terminate_current_hook` — precisely to keep this layering), which: + +1. **re-stamps the leader's exit record** with the stashed `group_reason` — the + record is written (again) at group-death time, so "the reason is recorded + before the notification posts" stays true and a supervisor can never be + notified and then read `-ESRCH` because the burst evicted an old record; +2. posts the leader's exit notification (and subscriber broadcast) from the + stashed endpoint, and drops that reference — exactly once. + +Both `releaseAddressSpace` call sites (`exitUserLocked`, `destroyTaskLocked`) run +under the BKL, so the hook does too; its wakes are safe at both (verified). For a +single-threaded process the behavior is *externally indistinguishable* from +today — the order of notify vs. destroy inverts, but both sit inside one lock +hold, so no other core can observe the space destroyed but the notification +unposted, or vice versa. That sentence is the correctness argument; it is also the +first invariant to re-examine if the BKL is ever split, along with +`killGroupLocked`'s single-hold atomicity. (Hand-built spaces that were never +retained take the immediate-destroy path and are out of the hook's scope.) + +Workers' `exit_subscribers` broadcasts still fire per task — the FAT server's +dead-client sweep is keyed by tid and needs those. + +**Signals.** `signal_bind` is per-task and the service harness binds on the main +thread, so signals address the leader in practice; that stays. During a group +death, `process_signal` may return `0` (accepted by a condemned member — never +delivered, every delivery point kills first) or `-ESRCH` (member already reaped); +init's stop sequence already tolerates both, and its timer escalation to +`process_kill` covers the gap. `process_signal` follows `process_kill`'s +leader re-key for consistency. + +## The shared-memory frame hazard + +`dropSharedMemoryReference` frees a region's physical frames when the last *handle* +reference drops, but mappings die only with the address space. If the last handle +lived in a torn-down member while any task still has the region mapped, that task +holds a live mapping onto freed frames — and the red team showed this is **not** +group-specific: a plain `thread_exit` of the handle-holding thread, or a last-ref +drop by a task *outside* the dying group during the condemned window, hits the same +use-after-free. + +So the fix is a property of the **object**, not the dropper: give +`SharedMemoryObject` a per-*mapping* reference — `shared_memory_map` (and create's +self-map) retains; each space's destruction releases. "Last reference" then means +*no handles and no mappings*, both hazard paths collapse into the existing +refcount, and no group-kill special case is needed at all. + +## Deliberately unchanged + +- Worker `thread_exit`: per-thread, full per-tid resource sweep, refcount drop. +- The condemned-but-running window: a member on another core can finish its + in-flight syscall and run user code for up to a tick before dying — identical to + today's single-task `process_kill` semantics ("prompt but asynchronous, like a + Unix signal"). A kill IPI would shrink it; it is not part of this plan. +- `thread_join` returns 0 for a killed thread; joiners inside a dying group are + woken and then reaped like any member. +- The `.reaping` state, reap lists, and stack reaper. + +## Accepted limits (documented, not fixed here) + +- **Notify-ring overflow**: a group death posts one subscriber badge per member + into 8-slot rings; a >8-member group can drop badges. Group size is bounded by + the 48-task table; today's largest production group is 2 (display) and + thread-test already reaches 5. +- **Exit-record ring pressure**: one 64-entry ring, one record per member — made + harmless for the supervisor by the hook's group-death re-stamp. +- **Enumerate shows a partial group** mid-death: reaped members vanish at once, + condemned members linger up to a tick (audited: no in-tree consumer + misbehaves; the `leader` field in `ProcessDescriptor` lets future consumers + group correctly). +- **Pre-existing reap race, not widened**: a task preempted *mid-syscall* is + `.ready` with `in_system_call = true`, and the tick's reap loop will reap it — + an existing hazard the fan-out inherits but must not add new instances of. + Filed to investigate separately. +- **Per-task DMA/shm cursors** (`dma_map_next`, `shared_memory_map_next`): two + sibling threads allocating overlap the same arena — pre-existing thread bug, + adjacent to but not part of this plan (the mmap/MMIO cursors already moved + per-space for exactly this reason). + +## Milestones + +- **M1 — the leader id.** `Task.leader` (kernel tasks: 0, never followed), set on + both spawn paths; `leader` added to `ProcessDescriptor`; `process_kill` and + `process_signal` re-keyed (kernel-task `-ESRCH` guard *before* leader + resolution). No fan-out yet. Existing tests must pass untouched. +- **M2 — the latch + fan-out.** `AddressSpaceRef.dying` + stash; + `retainAddressSpace` refuses dying spaces; `killGroupLocked`; wire the fault + path, `exit`, and `process_kill` into it; leader `thread_exit` → `-EPERM`; + `exit_reason`/`fault_kill_count` writes under the BKL; `kill_pending` made + atomic. Group notification via the `group_exit_hook` re-stamp + post. +- **M3 — shared-memory mapping refs.** `SharedMemoryObject` counts mappings; + space destruction releases them; frames free only at zero handles *and* zero + mappings. +- **M4 — tests + docs.** New `-Dtest-case`s (all `smp: 4` where cross-core + matters), driving `thread-test` with new argv modes: + - `thread-fault-group`: a worker faults; assert both tasks gone from + `enumerate`, `fault_kill_count == 1`, `process_exit_reason(leader) == + segmentation_fault`, address-space and stack-bytes counters return to base. + - `kill-threaded-group`: `process_kill(leader)` with a worker spinning on + another core; assert the worker dies by the deferred path, exactly one exit + badge, delivered only after both members are dead, and + `process_exit_reason(leader) == .killed`. + - `kill-via-worker-tid`: `process_kill(worker)` kills the whole group; + `-EPERM` for a non-supervisor aiming at the worker. + - `racing-triggers`: two members fault/exit simultaneously on different cores; + assert a deterministic leader reason and `fault_kill_count == 1`. + - `exit-group`: a *worker* calls `exit(3)`; group dies, leader reason + `.aborted`. + - `leader-thread-exit`: asserts the chosen rule (`-EPERM`, workers unaffected). + - `thread-exit-solo`: regression — worker `thread_exit` still leaves siblings + running. + - `group-claim-release`: a member claims a device; assert the claim is free and + the leader notification arrives only after every member is dead. + - `shm-mapping-ref`: last handle dropped by a dying thread; sibling's mapping + stays valid until space death (M3 regression). + Then update [threading.md](threading.md) (the shared-fate gap note), + [process-lifecycle.md](process-lifecycle.md), + [process-management.md](process-management.md), and + [ipc.md](ipc.md)/[drivers.md](drivers.md) mentions.