diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 8fd5754..de4394c 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -12,7 +12,7 @@ next one starts.* | Step | What | State | |---|---|---| -| L1 | Reclamation: a dead task's registrations die with its claims | not started | +| L1 | Reclamation: a dead task's registrations die with its claims | **stopped — the step was wrong; see open question 4** | | L2 | Bounds build check + allowlist; declare what we have already touched | not started | | L3 | xHCI: slot count from `HCSPARAMS1.MaxSlots`, not 8 | not started | | L4 | USB: configuration descriptor sized by `wTotalLength`, not 512 | not started | @@ -51,6 +51,28 @@ question down instead of inventing an answer. 3. **Which adversarial tests I1–I3 need.** The audit's six real defects were all found by asking what an attacker would do, and the suite had never asked. "Add adversarial cases" is not executable until the attacks are named. +4. **Reclamation is not a death-sweep problem, and L1 as written would have broken the + restart path.** Found on the first attempt at it. The audit is right that `count` + never decreases, but *death is the wrong trigger*: + - The broker keeps entries deliberately: "The devices stay in the table — they + describe hardware, which did not go away — only their ownership clears." A driver + dying does not unplug anything. + - Device ids must stay **stable across a bus restart**, because + `device-manager.driverForDevice` dedupes by `device_id` so that "a re-report after + a bus restart must not spawn a second instance". Stability comes from the + idempotency scan returning the existing id — removing entries on death would give + a restarted bus fresh ids and spawn duplicate driver instances. + - Everything else a task holds *is* already reclaimed on every path out: + `irq.releaseOwner`, `iommu.releaseAllOwnedBy`, `dmaRegistryReleaseOwner`, then the + broker's claims (`process.releaseTaskResourcesLocked`). + + So the real leak has two sources, and neither is death: a device that genuinely + **goes away** (hot-unplug) has no retirement path, and a bus that enumerates + *differently* on restart leaves its stale entries behind forever. Both are the device + manager's inventory problem — phase 3 — and both need the id-stability question + answered first (tombstone-and-reuse aliases stale ids held by another process; + generation-tagged ids change the id encoding, which is ABI). Not an unattended + decision. ### Working rules for the run