threads(M8): the task reaper — reclaim dead tasks' kernel stacks
A dead task's kernel stack was leaked (no reaper), so every process/thread death bled kernel memory. Now exit()/exitUserLocked record the dying task in a per-core reap_after_switch slot and switch away; the task that resumes on that core frees the stack in switchTo's tail (on its own stack, lock still held so the slot can't be reused). reapKillPendingLocked drains the slot on the timer tick as a safety net for the fresh-task case (a fresh task enters via the trampoline, bypassing switchTo's tail). A task killed while not running is freed directly in destroyTaskLocked. live_stack_bytes is the observable. Fixed a migration bug this exposed: the post-switchContext reap read the pc parameter, but a migrated task carries a stale pc in its saved switchTo frame -> it freed the wrong core's pending stack (a #GP under SMP). Re-fetch thisCpu() after the switch. Gate task-reap PASS (5x isolated, 2x in the 24-case batch); full guardrail 24/24 incl. fault-recovery/supervision/process-kill/smp/affinity; build + host green.
This commit is contained in:
@@ -153,6 +153,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
|
||||
threadIdTest(boot_information);
|
||||
} else if (eql(case, "thread-alloc")) {
|
||||
threadAllocTest(boot_information);
|
||||
} else if (eql(case, "task-reap")) {
|
||||
taskReapTest(boot_information);
|
||||
} else if (eql(case, "args")) {
|
||||
argsTest(boot_information);
|
||||
} else if (eql(case, "init")) {
|
||||
@@ -1738,6 +1740,42 @@ fn threadAllocTest(boot_information: *const BootInformation) void {
|
||||
result();
|
||||
}
|
||||
|
||||
/// The task reaper (docs/threading-plan.md M8): a dead task's kernel stack used to be
|
||||
/// leaked ("no reaper yet"). Spawn and kill many ring-3 processes and confirm the total
|
||||
/// kernel-stack bytes return to baseline — every stack reclaimed, no leak. (Threads exit
|
||||
/// through the same exitUserLocked path, so this covers them too.)
|
||||
fn taskReapTest(boot_information: *const BootInformation) void {
|
||||
_ = boot_information;
|
||||
log("DANOS-TEST-BEGIN: task-reap\n", .{});
|
||||
const base = scheduler.liveStackBytes();
|
||||
const rounds: u32 = 12;
|
||||
var killed: u32 = 0;
|
||||
var round: u32 = 0;
|
||||
while (round < rounds) : (round += 1) {
|
||||
process.fault_kill_count = 0;
|
||||
const probe = spawnFaultingProcess() orelse break;
|
||||
_ = probe;
|
||||
scheduler.setPriority(1);
|
||||
const deadline = architecture.millis() + 5000;
|
||||
while (process.fault_kill_count < 1 and architecture.millis() < deadline) scheduler.yield();
|
||||
scheduler.setPriority(4);
|
||||
if (process.fault_kill_count >= 1) killed += 1;
|
||||
}
|
||||
// Reaping is asynchronous — a dead task's stack is freed when its core next switches
|
||||
// or ticks. Poll (bounded) until the live bytes return to baseline: a correct reaper
|
||||
// gets there in a few ms; a genuine leak never does and this times out.
|
||||
scheduler.setPriority(1);
|
||||
const settle_deadline = architecture.millis() + 3000;
|
||||
while (scheduler.liveStackBytes() != base and architecture.millis() < settle_deadline) scheduler.yield();
|
||||
scheduler.setPriority(4);
|
||||
|
||||
check("all probes spawned and were killed", killed == rounds);
|
||||
check("kernel stacks reclaimed to baseline (no leak)", scheduler.liveStackBytes() == base);
|
||||
if (killed == rounds and scheduler.liveStackBytes() == base)
|
||||
log("task-reap: kernel stacks reclaimed to baseline ok\n", .{});
|
||||
result();
|
||||
}
|
||||
|
||||
/// The full PID-1 path: the bootloader read /system/services/init off the boot volume and
|
||||
/// handed it over; load it as a user ELF and spawn it as a real ring-3 process
|
||||
/// — the same call the normal boot path makes — then confirm it beats. init
|
||||
|
||||
Reference in New Issue
Block a user