threads(M8): the task reaper — reclaim dead tasks' kernel stacks

A dead task's kernel stack was leaked (no reaper), so every process/thread death
bled kernel memory. Now exit()/exitUserLocked record the dying task in a per-core
reap_after_switch slot and switch away; the task that resumes on that core frees
the stack in switchTo's tail (on its own stack, lock still held so the slot can't
be reused). reapKillPendingLocked drains the slot on the timer tick as a safety
net for the fresh-task case (a fresh task enters via the trampoline, bypassing
switchTo's tail). A task killed while not running is freed directly in
destroyTaskLocked. live_stack_bytes is the observable.

Fixed a migration bug this exposed: the post-switchContext reap read the pc
parameter, but a migrated task carries a stale pc in its saved switchTo frame ->
it freed the wrong core's pending stack (a #GP under SMP). Re-fetch thisCpu()
after the switch.

Gate task-reap PASS (5x isolated, 2x in the 24-case batch); full guardrail 24/24
incl. fault-recovery/supervision/process-kill/smp/affinity; build + host green.
This commit is contained in:
2026-07-20 23:18:30 +01:00
parent 8259678f0a
commit ed3b3f1c45
4 changed files with 124 additions and 14 deletions
+26 -14
View File
@@ -335,23 +335,35 @@ multi-threaded code would hit it. Closed it:
> Reworked it (and the settle loop) to wait on the wall clock instead, so its duration is
> independent of unrelated code changes.
### M8 — The task reaper (cleanup + resilience)
### M8 — The task reaper (cleanup + resilience) ✅
A dead task's **kernel** stack is currently leaked ("no reaper yet") — every process
*and* thread death loses one, so a crash loop bleeds kernel memory. A reaper fixes it and
serves the [resilience](resilience.md) restart goal directly:
A dead task's **kernel** stack was leaked ("no reaper yet") — every process *and* thread
death lost one, so a crash loop bled kernel memory. The reaper fixes it and serves the
[resilience](resilience.md) restart goal directly:
- [ ] A dying task cannot free the kernel stack it runs on, so it hands itself to a
**reap list** and switches away; the kernel stack (and, for a detached thread, its
user stack) is reclaimed from another context — a low-priority reaper step drained
on the scheduler tick and when a core goes idle. Extends the existing
`reap_task_hook`/`destroyTaskLocked` path rather than inventing a parallel one.
- [ ] `-Dtest-case=task-reap`: spawn and exit many threads and processes; assert the
kernel-heap free bytes (a new test observable) return to **baseline** — kernel
stacks reclaimed, no leak — and that the `fault-recovery`/kill paths reclaim too.
- [x] A dying task cannot free the kernel stack it runs on, so `exit()`/`exitUserLocked`
record it in a **per-core `reap_after_switch` slot** and switch away; the task that
resumes on that core frees the stack in `switchTo`'s tail (it's on its own stack, the
big lock is still held so the slot can't have been reused). A **tick-time drain**
(`reapKillPendingLocked`) is the safety net for the case where the next task is
*fresh* (enters via the trampoline, bypassing `switchTo`'s tail). A task killed while
*not* running is freed immediately in `destroyTaskLocked`. A `live_stack_bytes`
counter is the observable. *(Detached-thread user-stack reclaim moves to M9, which
adds the joinable/detached flag.)*
- [x] `-Dtest-case=task-reap` (`smp: 4`): spawn and kill 12 processes; poll the
test-observable `scheduler.liveStackBytes()` until it returns to **baseline** (a
correct reaper gets there in a few ms; a genuine leak times out) — every kernel
stack reclaimed, no leak. Threads exit through the same `exitUserLocked`, so covered.
**Gate:** `task-reap` passes; `fault-recovery`, `supervision`, `process-kill`,
`aspace-refcount` still green.
**Gate (met):** `task-reap` passes (5× isolated + 2× in the full batch); `fault-recovery`,
`supervision`, `process-kill`, `aspace-refcount`, `smp`, `affinity` all still green (24/24
full guardrail); `zig build`/`zig build test` clean.
> **Bug found + fixed here (touches every context switch):** the post-`switchContext` reap
> first read the `pc` **parameter**, but a task that migrated cores carries a *stale* `pc`
> in its saved `switchTo` frame — so it read the wrong core's slot and freed a live stack
> (a #GP under SMP). Fixed to re-fetch `thisCpu()` after the switch (the switch only swaps
> stacks on the current core).
### M9 — Futex-completion join (retire the per-thread endpoint)