From 568823a4fb52193c498c0a65f1754d23e74910a5 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 11:15:11 +0100 Subject: [PATCH] =?UTF-8?q?docs:=20L1=20was=20wrong=20=E2=80=94=20reclamat?= =?UTF-8?q?ion=20is=20not=20a=20death-sweep=20problem?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The first step of the unattended run was "a dead task's registrations die with its claims". Implementing it would have broken the restart path it was meant to protect. The broker keeps entries on purpose: they describe hardware, which did not go away when a driver died. And device ids must stay stable across a bus restart, because device-manager dedupes re-reports by device_id so a restarted bus does not spawn a second driver instance — stability that comes from the idempotency scan returning the existing id. Removing entries on death would hand a restarted bus fresh ids and duplicate every driver. Everything else a task holds is already reclaimed on every path out: IRQ bindings, IOMMU domains, DMA regions, then its claims. The leak the audit found is real but has two other sources: a device that genuinely goes away has no retirement path, and a bus that enumerates differently on restart strands its old entries. Both belong to the device manager's inventory, and both need the id-stability question settled first — tombstone-and-reuse aliases ids another process still holds, generation tagging changes the id encoding, which is ABI. Recorded as an open question rather than guessed at. --- docs/bounds-track-plan.md | 24 +++++++++++++++++++++++- 1 file changed, 23 insertions(+), 1 deletion(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 8fd5754..de4394c 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -12,7 +12,7 @@ next one starts.* | Step | What | State | |---|---|---| -| L1 | Reclamation: a dead task's registrations die with its claims | not started | +| L1 | Reclamation: a dead task's registrations die with its claims | **stopped — the step was wrong; see open question 4** | | L2 | Bounds build check + allowlist; declare what we have already touched | not started | | L3 | xHCI: slot count from `HCSPARAMS1.MaxSlots`, not 8 | not started | | L4 | USB: configuration descriptor sized by `wTotalLength`, not 512 | not started | @@ -51,6 +51,28 @@ question down instead of inventing an answer. 3. **Which adversarial tests I1–I3 need.** The audit's six real defects were all found by asking what an attacker would do, and the suite had never asked. "Add adversarial cases" is not executable until the attacks are named. +4. **Reclamation is not a death-sweep problem, and L1 as written would have broken the + restart path.** Found on the first attempt at it. The audit is right that `count` + never decreases, but *death is the wrong trigger*: + - The broker keeps entries deliberately: "The devices stay in the table — they + describe hardware, which did not go away — only their ownership clears." A driver + dying does not unplug anything. + - Device ids must stay **stable across a bus restart**, because + `device-manager.driverForDevice` dedupes by `device_id` so that "a re-report after + a bus restart must not spawn a second instance". Stability comes from the + idempotency scan returning the existing id — removing entries on death would give + a restarted bus fresh ids and spawn duplicate driver instances. + - Everything else a task holds *is* already reclaimed on every path out: + `irq.releaseOwner`, `iommu.releaseAllOwnedBy`, `dmaRegistryReleaseOwner`, then the + broker's claims (`process.releaseTaskResourcesLocked`). + + So the real leak has two sources, and neither is death: a device that genuinely + **goes away** (hot-unplug) has no retirement path, and a bus that enumerates + *differently* on restart leaves its stale entries behind forever. Both are the device + manager's inventory problem — phase 3 — and both need the id-stability question + answered first (tombstone-and-reuse aliases stale ids held by another process; + generation-tagged ids change the id encoding, which is ABI). Not an unattended + decision. ### Working rules for the run