docs: L1 was wrong — reclamation is not a death-sweep problem

The first step of the unattended run was "a dead task's registrations die
with its claims". Implementing it would have broken the restart path it was
meant to protect.

The broker keeps entries on purpose: they describe hardware, which did not go
away when a driver died. And device ids must stay stable across a bus
restart, because device-manager dedupes re-reports by device_id so a
restarted bus does not spawn a second driver instance — stability that comes
from the idempotency scan returning the existing id. Removing entries on
death would hand a restarted bus fresh ids and duplicate every driver.

Everything else a task holds is already reclaimed on every path out: IRQ
bindings, IOMMU domains, DMA regions, then its claims.

The leak the audit found is real but has two other sources: a device that
genuinely goes away has no retirement path, and a bus that enumerates
differently on restart strands its old entries. Both belong to the device
manager's inventory, and both need the id-stability question settled first —
tombstone-and-reuse aliases ids another process still holds, generation
tagging changes the id encoding, which is ABI. Recorded as an open question
rather than guessed at.
This commit is contained in:
Daniel Samson
2026-08-08 11:15:11 +01:00
parent 4398eb7cc4
commit 568823a4fb
+23 -1
View File
@@ -12,7 +12,7 @@ next one starts.*
| Step | What | State | | Step | What | State |
|---|---|---| |---|---|---|
| L1 | Reclamation: a dead task's registrations die with its claims | not started | | L1 | Reclamation: a dead task's registrations die with its claims | **stopped — the step was wrong; see open question 4** |
| L2 | Bounds build check + allowlist; declare what we have already touched | not started | | L2 | Bounds build check + allowlist; declare what we have already touched | not started |
| L3 | xHCI: slot count from `HCSPARAMS1.MaxSlots`, not 8 | not started | | L3 | xHCI: slot count from `HCSPARAMS1.MaxSlots`, not 8 | not started |
| L4 | USB: configuration descriptor sized by `wTotalLength`, not 512 | not started | | L4 | USB: configuration descriptor sized by `wTotalLength`, not 512 | not started |
@@ -51,6 +51,28 @@ question down instead of inventing an answer.
3. **Which adversarial tests I1–I3 need.** The audit's six real defects were all found 3. **Which adversarial tests I1–I3 need.** The audit's six real defects were all found
by asking what an attacker would do, and the suite had never asked. "Add adversarial by asking what an attacker would do, and the suite had never asked. "Add adversarial
cases" is not executable until the attacks are named. cases" is not executable until the attacks are named.
4. **Reclamation is not a death-sweep problem, and L1 as written would have broken the
restart path.** Found on the first attempt at it. The audit is right that `count`
never decreases, but *death is the wrong trigger*:
- The broker keeps entries deliberately: "The devices stay in the table — they
describe hardware, which did not go away — only their ownership clears." A driver
dying does not unplug anything.
- Device ids must stay **stable across a bus restart**, because
`device-manager.driverForDevice` dedupes by `device_id` so that "a re-report after
a bus restart must not spawn a second instance". Stability comes from the
idempotency scan returning the existing id — removing entries on death would give
a restarted bus fresh ids and spawn duplicate driver instances.
- Everything else a task holds *is* already reclaimed on every path out:
`irq.releaseOwner`, `iommu.releaseAllOwnedBy`, `dmaRegistryReleaseOwner`, then the
broker's claims (`process.releaseTaskResourcesLocked`).
So the real leak has two sources, and neither is death: a device that genuinely
**goes away** (hot-unplug) has no retirement path, and a bus that enumerates
*differently* on restart leaves its stale entries behind forever. Both are the device
manager's inventory problem — phase 3 — and both need the id-stability question
answered first (tombstone-and-reuse aliases stale ids held by another process;
generation-tagged ids change the id encoding, which is ABI). Not an unattended
decision.
### Working rules for the run ### Working rules for the run