threads(M8): the task reaper — reclaim dead tasks' kernel stacks
A dead task's kernel stack was leaked (no reaper), so every process/thread death bled kernel memory. Now exit()/exitUserLocked record the dying task in a per-core reap_after_switch slot and switch away; the task that resumes on that core frees the stack in switchTo's tail (on its own stack, lock still held so the slot can't be reused). reapKillPendingLocked drains the slot on the timer tick as a safety net for the fresh-task case (a fresh task enters via the trampoline, bypassing switchTo's tail). A task killed while not running is freed directly in destroyTaskLocked. live_stack_bytes is the observable. Fixed a migration bug this exposed: the post-switchContext reap read the pc parameter, but a migrated task carries a stale pc in its saved switchTo frame -> it freed the wrong core's pending stack (a #GP under SMP). Re-fetch thisCpu() after the switch. Gate task-reap PASS (5x isolated, 2x in the 24-case batch); full guardrail 24/24 incl. fault-recovery/supervision/process-kill/smp/affinity; build + host green.
This commit is contained in:
+26
-14
@@ -335,23 +335,35 @@ multi-threaded code would hit it. Closed it:
|
||||
> Reworked it (and the settle loop) to wait on the wall clock instead, so its duration is
|
||||
> independent of unrelated code changes.
|
||||
|
||||
### M8 — The task reaper (cleanup + resilience)
|
||||
### M8 — The task reaper (cleanup + resilience) ✅
|
||||
|
||||
A dead task's **kernel** stack is currently leaked ("no reaper yet") — every process
|
||||
*and* thread death loses one, so a crash loop bleeds kernel memory. A reaper fixes it and
|
||||
serves the [resilience](resilience.md) restart goal directly:
|
||||
A dead task's **kernel** stack was leaked ("no reaper yet") — every process *and* thread
|
||||
death lost one, so a crash loop bled kernel memory. The reaper fixes it and serves the
|
||||
[resilience](resilience.md) restart goal directly:
|
||||
|
||||
- [ ] A dying task cannot free the kernel stack it runs on, so it hands itself to a
|
||||
**reap list** and switches away; the kernel stack (and, for a detached thread, its
|
||||
user stack) is reclaimed from another context — a low-priority reaper step drained
|
||||
on the scheduler tick and when a core goes idle. Extends the existing
|
||||
`reap_task_hook`/`destroyTaskLocked` path rather than inventing a parallel one.
|
||||
- [ ] `-Dtest-case=task-reap`: spawn and exit many threads and processes; assert the
|
||||
kernel-heap free bytes (a new test observable) return to **baseline** — kernel
|
||||
stacks reclaimed, no leak — and that the `fault-recovery`/kill paths reclaim too.
|
||||
- [x] A dying task cannot free the kernel stack it runs on, so `exit()`/`exitUserLocked`
|
||||
record it in a **per-core `reap_after_switch` slot** and switch away; the task that
|
||||
resumes on that core frees the stack in `switchTo`'s tail (it's on its own stack, the
|
||||
big lock is still held so the slot can't have been reused). A **tick-time drain**
|
||||
(`reapKillPendingLocked`) is the safety net for the case where the next task is
|
||||
*fresh* (enters via the trampoline, bypassing `switchTo`'s tail). A task killed while
|
||||
*not* running is freed immediately in `destroyTaskLocked`. A `live_stack_bytes`
|
||||
counter is the observable. *(Detached-thread user-stack reclaim moves to M9, which
|
||||
adds the joinable/detached flag.)*
|
||||
- [x] `-Dtest-case=task-reap` (`smp: 4`): spawn and kill 12 processes; poll the
|
||||
test-observable `scheduler.liveStackBytes()` until it returns to **baseline** (a
|
||||
correct reaper gets there in a few ms; a genuine leak times out) — every kernel
|
||||
stack reclaimed, no leak. Threads exit through the same `exitUserLocked`, so covered.
|
||||
|
||||
**Gate:** `task-reap` passes; `fault-recovery`, `supervision`, `process-kill`,
|
||||
`aspace-refcount` still green.
|
||||
**Gate (met):** `task-reap` passes (5× isolated + 2× in the full batch); `fault-recovery`,
|
||||
`supervision`, `process-kill`, `aspace-refcount`, `smp`, `affinity` all still green (24/24
|
||||
full guardrail); `zig build`/`zig build test` clean.
|
||||
|
||||
> **Bug found + fixed here (touches every context switch):** the post-`switchContext` reap
|
||||
> first read the `pc` **parameter**, but a task that migrated cores carries a *stale* `pc`
|
||||
> in its saved `switchTo` frame — so it read the wrong core's slot and freed a live stack
|
||||
> (a #GP under SMP). Fixed to re-fetch `thisCpu()` after the switch (the switch only swaps
|
||||
> stacks on the current core).
|
||||
|
||||
### M9 — Futex-completion join (retire the per-thread endpoint)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user