The name list and its predicate existed so drivers could move to delegation one at a time with the suite green throughout. Every driver is delegated now, so the manager simply hands over whatever device a driver was assigned. Deleting it caught a real consequence: crash-test finally got delegated too, and it was still claiming its device — so it got AlreadyClaimed because it already held it, exited, and the restart drill had nothing to restart. Its own comment named what the case was really checking: "the respawn only reaches this line because the kernel released the previous instance's claim at death". That property still holds, by a different mechanism — the device reverts to the manager on death and is handed to the replacement, which is the same guarantee without the race it used to rely on. All four delegation paths verified: the xHCI controller, the PCI bridge, the PS/2 two-node singleton, and virtio-gpu's restart re-attach. Run 3 complete. Suite 118/118.
633 lines
36 KiB
Markdown
633 lines
36 KiB
Markdown
# The bounds track: removing the numbers we invented
|
||
|
||
*Plan, 2026-08-08. Follows [fixed-bounds-audit.md](fixed-bounds-audit.md) (235 ceilings,
|
||
139 on quantities we do not choose) and the AMD Ryzen that found the first one.*
|
||
|
||
---
|
||
|
||
## Live state — the unattended run
|
||
|
||
*This table is the progress view. It is updated at the end of every step, before the
|
||
next one starts.*
|
||
|
||
| Step | What | State |
|
||
|---|---|---|
|
||
| L1 | Reclamation: a dead task's registrations die with its claims | **stopped — the step was wrong; see open question 4** |
|
||
| L2 | Bounds build check + allowlist; declare what we have already touched | **done** — `zig build bounds`, 273 allowlisted, 5 declared |
|
||
| L3 | xHCI: slot count from `HCSPARAMS1.MaxSlots`, not 8 | **done** — QEMU reports 64; the driver tracked 8 |
|
||
| L4 | USB: configuration descriptor sized by `wTotalLength`, not 512 | **done** — QEMU tops out at 211 bytes, so the case catches the class, not the original trigger |
|
||
| L5 | USB: interfaces from the descriptor, and the misattributed-endpoint bug | **done** — fix is by construction; no direct test, see open question 5 |
|
||
| L6 | xHCI: a failed `allocateDevice` stops leaking an enabled slot | **done** — path forced and verified; no regression test, see open question 5 |
|
||
|
||
**Run 1 complete.** L1 stopped (the step was wrong), L2–L6 landed. Suite 115 → 116.
|
||
Allowlist 278 → 269. Two steps ship without a permanent regression test, both because
|
||
QEMU's USB devices are too small to reach the paths — see open question 5, which is the
|
||
audit's own lesson recurring: the test rig is smaller than a real machine.
|
||
|
||
---
|
||
|
||
## Run 2 — device authority: delete the invented ceilings
|
||
|
||
*Design: [device-authority.md](os-development/device-authority.md), which is the **how**
|
||
for the delegation step [device-manager.md](device-driver-development/device-manager.md)
|
||
already settled. Read both before starting; the second is authoritative where they
|
||
differ.*
|
||
|
||
The goal, in the project owner's words: **remove the maximum values we set arbitrarily,
|
||
move the responsibility to the device manager, and keep in the kernel only the parts
|
||
that cannot safely run in user space.**
|
||
|
||
| Step | What | State |
|
||
|---|---|---|
|
||
| D0 | The grant rides `system_spawn` — atomic, so no driver need change to receive one | **done** |
|
||
| D1 | `device_transfer(device_id, task_id)` — the holder gives a device away | **done** — syscall 54; a move, not a copy |
|
||
| D2 | Adversarial case: a process handed nothing is refused | **done** — `device-authority-test`; the claim half joins it at D6 |
|
||
| D3 | The manager claims the seeded devices before any driver is spawned | **merged into D4** |
|
||
| D4 | `usb-xhci-bus` receives its controller | **done** — caught an IOMMU regression I introduced |
|
||
| D5 | `pci-bus`, `virtio-gpu`, `ps2-bus`, discovery receive theirs | **done** — every driver now receives its hardware |
|
||
| D6 | `device_claim` refuses a device the caller was not handed | **blocked on question 10** |
|
||
| D7 | Zero-resource devices stop being kernel objects | **blocked on question 8** |
|
||
| D8 | **`maximum_children_per_parent` deleted** | **done** |
|
||
| D9 | Dynamic table; **`maximum_devices` deleted**; per-registrar allowance | **done** |
|
||
| D10 | Every driver hellos, on its own merits | not started — optional, independent |
|
||
|
||
*Executed in the order D1, D2, D4, D5(part), D9, D0, D5(rest), D8 — the numbering is
|
||
the original plan's, not the sequence. D0 was added mid-run, D3 merged into D4, and D8
|
||
turned out to have been unblocked since D9.*
|
||
|
||
**Run 2 stops, blocked on one question.** Landed: D0, D1, D2, D4, D8, D9, and D5 for
|
||
three of five claimants (`usb-xhci-bus`, `pci-bus`, `virtio-gpu`). Suite 118/118.
|
||
|
||
**Both invented ceilings are gone.** `maximum_devices` and
|
||
`maximum_children_per_parent` no longer exist, and delegation is atomic with the spawn.
|
||
|
||
What remains is not a number: **D6 closes the claiming hole** and is blocked on
|
||
question 9, and D7 moves the inventory and is blocked on question 8.
|
||
|
||
Ordering is load-bearing. D1–D2 built and proved the mechanism with nothing depending on
|
||
it. D4–D5 move each claimant across one at a time, so the suite stays green throughout
|
||
and a regression names the driver that caused it. D6 is the flag day. (The original
|
||
"D7 must precede D9" no longer holds: the per-holder quota bounds zero-resource children
|
||
as well as anything else, so D9 landed without it.)
|
||
|
||
**D3 merged into D4** (found while implementing, recorded rather than worked around).
|
||
The two cannot be separated: the moment the manager claims a device, any driver still
|
||
calling `device_claim` on it is refused `AlreadyClaimed`, so D3 on its own turns the
|
||
suite red — and D3 applied to *nothing* changes no behaviour and cannot be tested.
|
||
They land together, with the manager claiming only for drivers in an explicit
|
||
**delegated set** so every unconverted driver keeps claiming exactly as before.
|
||
`usb-xhci-bus` is the first member, as it was the first driver to conform to `hello`
|
||
(device-manager.md, M18.1). D5 moves the rest in one at a time; the set and the
|
||
`device_claim` path both disappear at D6.
|
||
|
||
### Observations from the run
|
||
|
||
- **D4 introduced an IOMMU regression, caught by converting one driver at a time.**
|
||
`confineDevice` runs inside `systemDeviceClaim`, so a device arriving by *transfer*
|
||
was never confined for its new owner: the driver's DMA rings went unbound (three
|
||
IOMMU+USB cases failed), and two worse consequences were latent — a manager death
|
||
would have torn down a domain a live driver was using, and a driver death would have
|
||
leaked one. `iommu.reassign` moves the confinement with the device, keeping the
|
||
domain and attachment intact so it never translates through nothing. Converting all
|
||
five drivers at once would have produced the same three failures with five suspects.
|
||
- **`usb-hub` failed once in a full run, then passed six times** (four isolated, two
|
||
full). Suspected instance of the known intermittent AP ring-3 fault rather than
|
||
anything in D4 — recorded rather than dismissed, because D4 moved the `hello` earlier
|
||
and so did shift boot timing. Watch it across the remaining steps.
|
||
|
||
---
|
||
|
||
## Run 3 — close the claiming hole
|
||
|
||
*Every driver now receives its hardware. What remains is that `device_claim` still
|
||
works for anyone, so a device nobody holds can still be taken. Six steps, no open
|
||
questions.*
|
||
|
||
### The rule this run implements
|
||
|
||
The kernel records, per device, **who gave it**. From that one field everything follows:
|
||
|
||
- `device_transfer` and the spawn-fused grant set the giver.
|
||
- **`device_claim` refuses a device that has a giver.** A device that was given to
|
||
someone is delegated hardware and must be handed on, never taken.
|
||
- **On death a device reverts to its giver**, if that task is still alive; otherwise its
|
||
claim clears. A grant is a loan, not a gift.
|
||
|
||
Two things fall out rather than being special-cased:
|
||
|
||
- **The framebuffer is untouched.** Nobody delegates it, so it has no giver, so the
|
||
display service can still claim it exactly as today. No exemption in the kernel, no
|
||
mention of display anywhere in the rule.
|
||
- **Restart needs no race.** A dying driver's device returns to the manager, which
|
||
re-delegates it on respawn. Previously the kernel released it to nobody and the
|
||
manager re-claimed first-come, so every restart reopened the hole.
|
||
|
||
**Found while implementing: E2 is the step that closes the hole, not E3.** A delegated
|
||
device is *held*, so an attempt to take it is refused as `AlreadyClaimed` long before
|
||
the giver is consulted — and once a borrower's death returns the device to its lender
|
||
(or clears both when the lender is gone), there is no state where a device is unheld and
|
||
still on loan. E3's check is therefore unreachable today. It stays as one comparison
|
||
that fails closed, guarding any future path that frees a device without clearing its
|
||
giver, and its comment says so rather than implying a protection it is not providing.
|
||
|
||
| Step | What |
|
||
|---|---|
|
||
| E1 | Record a giver per device; `device_transfer` and the spawn grant set it — **done** |
|
||
| E2 | On task death a device reverts to its giver if alive, else its claim clears — **done** |
|
||
| E3 | `device_claim` refuses a device that has a giver — **done**, but unreachable: E2 already closed the window |
|
||
| E4 | The manager claims every resource-bearing device at boot, so nothing is left takeable — **done** (boot snapshot only; see below) |
|
||
| E5 | The attacker fixture gains the claim half it has been waiting for since D2 — **done with E4** |
|
||
| E6 | Delete the delegated-set scaffolding — every driver is delegated now — **done** |
|
||
|
||
Ordering: E1 alone changes no behaviour. E2 must precede E3, or a restart cannot
|
||
re-acquire. E4 must precede E5, or the attacker will find takeable devices and the
|
||
assertion will be wrong about why. E6 is cleanup.
|
||
|
||
**Run 3 complete.** Suite 118/118. Every driver receives its hardware; a grant is a
|
||
loan that returns to its lender when the borrower dies; nothing firmware-discovered is
|
||
left unheld except the framebuffer, which the compositor owns.
|
||
|
||
### E4's scope, stated
|
||
|
||
It covers the **boot snapshot**. A device *reported* later and matched to no driver
|
||
stays claimable — `pci-cap-test` and `iommu-fault-test` both rely on that to reach an
|
||
unmatched NIC. Narrowing it further is a separate change with those fixtures in scope.
|
||
|
||
The real gap it closed was the **HPET**: an MMIO window, an IRQ, no user-space driver,
|
||
and claimable by anyone. Excluded on purpose: the loader's framebuffer (the manager
|
||
starts before display, so taking it would break the boot screen) and anything with no
|
||
resources, which grants nothing.
|
||
|
||
### Not in this run, and not blocking it
|
||
|
||
**Zero-resource devices stay in the kernel.** Moving them out needs an answer to who
|
||
mints their ids, given that the id is the `device_token` on the usb-transfer wire and
|
||
the kernel's idempotency is what keeps it stable across a bus restart. Nothing depends
|
||
on it: both invented ceilings are already gone and this run does not need it. Recorded
|
||
as future work.
|
||
|
||
**Not every driver hellos.** Delivery no longer needs it, so it stands or falls on its
|
||
own merits — uniform liveness and one class of driver. Also future work.
|
||
|
||
---
|
||
|
||
### Question 9 — answered: a singleton gets every device that matched it
|
||
|
||
The 8042 is one controller described by two ACPI nodes, so it cannot be split across
|
||
processes. The manager now hands the single `ps2-bus` instance **every** matching node:
|
||
the first rides the spawn, the rest are transferred to the running instance. Late
|
||
arrival is safe because the ordering is natural rather than lucky — the controller node
|
||
carries the ports and is needed at once, the mouse node not until after identify
|
||
(measured: handed over at 0.336, first touched at 0.456). The count is whatever matched,
|
||
so a machine with no PS/2 ports or one port needs no special case.
|
||
|
||
Discovery turned out not to be special either. The kernel seeds the `acpi-tables` node,
|
||
so it is in the same boot snapshot the manager already scans for the PCI host bridge —
|
||
it is handed over at spawn like everything else.
|
||
|
||
### Open question 10 — the framebuffer is the last thing anyone claims
|
||
|
||
`device_claim` now has exactly two callers: the device manager, which is the acquirer
|
||
and should have it, and `system/services/display/backend.zig`, whose GOP path claims the
|
||
kernel-seeded display node.
|
||
|
||
Closing `device_claim` breaks the compositor's boot floor. Exempting it puts a hole in
|
||
the middle of the authority model, in the one place an exemption is most expensive.
|
||
|
||
The likely answer is neither. **The framebuffer is not a device** — it is where pixels
|
||
go, handed over by the loader, and the kernel wraps it in a `DeviceClass.display`
|
||
descriptor only so `mmio_map` can hand it over write-combining. If that is right, it
|
||
should leave the device table rather than be exempted from its rules, and the compositor
|
||
should receive the pixels some other way. That is a change to the display path, which is
|
||
out of scope for this run.
|
||
|
||
### Settled 2026-08-08: the grant rides `system_spawn` (D0)
|
||
|
||
Questions 6 and 7 both dissolved on inspection — neither `ps2-bus` nor discovery needs
|
||
to start speaking `hello`, and `virtio-gpu` has no standalone path to lose. What remains
|
||
is *where the grant is delivered*, and there are three candidates:
|
||
|
||
| | Race window | Cost |
|
||
|---|---|---|
|
||
| Transfer after spawn | **yes** | none |
|
||
| Every driver hellos | no | `ps2-bus` + discovery gain a handshake |
|
||
| **Grant rides `system_spawn`** | **no — atomic** | one more syscall argument |
|
||
|
||
**Take the third.** The manager cannot transfer before the child exists, so a separate
|
||
transfer always leaves a window in which the child is running and does not yet hold its
|
||
device. It would close on QEMU every time and open occasionally on a machine with
|
||
different timing — the exact failure shape this track exists to delete, and not worth
|
||
introducing while removing the others. Fusing the device into the spawn removes it by
|
||
construction: the child does not exist until it holds the device. No new knowledge in
|
||
the kernel — the same rule, *you may give away what you hold*, made atomic with the call
|
||
that creates the recipient. `system_spawn` uses five of six argument registers, so there
|
||
is room, and `no_device` is already the sentinel for a driver with no assignment.
|
||
|
||
**`hello` for every driver is a good idea on its own merits** — uniform liveness, the
|
||
deadline applied to all rather than some, and the `speaks_protocol` two-class split
|
||
leaving the manager (a wedged `ps2-bus` is invisible to its supervisor today). It is
|
||
D10, kept separate so grant delivery does not force it.
|
||
|
||
### Open questions raised by D5 — resolved except question 8
|
||
|
||
Delegation is delivered in `onHello`. That works for a driver that says hello, and
|
||
**two of the four do not**.
|
||
|
||
6. **`ps2-bus` does not hello**, and that is deliberate: device-manager.md records
|
||
"Legacy drivers (e.g. ps2-bus) are supervised and restarted but **not yet required
|
||
to hello**", and `Driver.speaks_protocol` exists to say so. Either it leaves
|
||
"legacy", or the grant gets a delivery point that is not `hello`.
|
||
|
||
**The `acpi` service is not this case — an earlier version of this question wrongly
|
||
lumped it in.** It is spawned as `addDriver("discovery", no_device, false)`: the
|
||
*discovery service*, one per firmware, with **no device assignment at all**, which
|
||
"finds and claims the acpi-tables (or devicetree-blob) node itself. Not a per-device
|
||
driver." The manager cannot hand it a device because discovery is what produces the
|
||
device tree — there is nothing to match against yet. Its question is not "should it
|
||
hello" but "is the bootstrap exempt from D6, or does the manager claim the
|
||
acpi-tables node and pass it on".
|
||
|
||
**A delivery point that needs no hello already exists in the shape of the code**:
|
||
`process.spawnSupervised` returns the child's pid, so the manager could transfer
|
||
immediately after spawn. If that is acceptable, this question mostly dissolves —
|
||
`ps2-bus` would not need to leave "legacy" and discovery could be handed its node
|
||
too. The cost is ordering: the child may reach for the device before the transfer
|
||
lands, where `hello` guarantees it cannot because the child is the one asking. That
|
||
trade is the actual decision.
|
||
|
||
7. **`virtio-gpu`'s "standalone bring-up" — resolved; there is no such path.** An
|
||
earlier version of this question read the comment "Best-effort: standalone bring-up
|
||
has no manager" as a boot path that delegation would delete. It is not one.
|
||
`virtio-gpu`'s `main` requires `argv[1]` and exits without it, and the only source of
|
||
that argument is the device manager — the chain is ACPI → pci-bus reports the
|
||
function → `devices.csv` matches `1AF4:1050` → the manager spawns the driver with the
|
||
id. There is no way to run it without a manager, so nothing is lost.
|
||
|
||
What the comment is really about is **resilience**: the hello is best-effort so a
|
||
driver whose manager has *died* keeps serving, which device-manager.md states
|
||
("If the manager dies, drivers keep running").
|
||
|
||
The residual is one narrow window: the manager spawns a driver and dies before
|
||
transferring. Today the driver could still claim, because claiming is free-for-all;
|
||
after D6 it exits and the restarted manager respawns it. That is the better
|
||
behaviour — a driver holding hardware nobody assigned it is what D6 exists to stop.
|
||
|
||
Until these are answered, `device_claim` cannot be closed off at D6 for those three, so
|
||
**D6 is blocked on questions 6 and 7**.
|
||
|
||
8. **Removing zero-resource devices from the kernel needs someone else to mint their
|
||
ids, and nothing says who.** A USB interface is registered with
|
||
`resource_count = 0`, and the id `device_register` returns is load-bearing in three
|
||
places: it is the `child_added` packet's target, it is the class driver's `argv[1]`,
|
||
and it is the `device_token` of the **usb-transfer wire protocol** — so the id space
|
||
is visible on the wire, not merely internal. usb-xhci-bus's own comment records the
|
||
fourth constraint: the kernel's idempotency is what makes "the same port and
|
||
interface always map back to the same device id" across a bus restart, which is what
|
||
stops a respawned bus spawning duplicate class drivers.
|
||
|
||
So the mover has to answer: who mints the id, how it stays stable across a *bus*
|
||
restart, how it stays stable across a *manager* restart (open question 3 territory),
|
||
and whether the wire protocol's `device_token` changes meaning. That is a design
|
||
step, not a mechanical one.
|
||
|
||
**D8 is reordered to run after D9**, which is a correction to the original sequencing.
|
||
D8's justification was "the authorisation it stood in for exists" — but with D6 blocked
|
||
it does not, so deleting the shared per-parent cap now would reopen the exhaustion hole
|
||
it was written for. D9's **per-holder quota** closes that hole independently of
|
||
authorisation, and does it better: a rogue exhausts its own allowance rather than the
|
||
table everyone shares. Once the quota exists, the per-parent cap is redundant whether or
|
||
not D6 has landed.
|
||
|
||
D9 also no longer depends on D7. Its original rationale was that zero-resource children
|
||
are the case that sidesteps containment — true, but a per-holder quota bounds them just
|
||
as well as anything else, because it counts entries per holder rather than per parent.
|
||
|
||
### Settled, so the run does not re-litigate them
|
||
|
||
- **The manager claims, it is not granted.** No binary names in the kernel; the rule is
|
||
"you may give away what you hold". The residual — it rests on the manager claiming
|
||
first — is stated in the design and is closed later by the spawn capability
|
||
[drivers.md](device-driver-development/drivers.md) already names as missing.
|
||
- **The framebuffer is not a device.** It is where pixels go, handed over by the loader,
|
||
and the compositor uses it as the boot floor until a real display driver announces
|
||
itself. Nothing in this run touches the display service or its GOP path.
|
||
- **`maximum_endpoints_per_interface` and the wire structs stay.** Widening them is a
|
||
protocol change, out of scope.
|
||
- **A per-holder quota is not a retreat.** Dynamic storage with no bound moves the
|
||
ceiling to the kernel heap, which is shared and fatal rather than partial. A bound
|
||
charged to the task that caused it is isolation, and it is declared through
|
||
[bounds.md](os-development/bounds.md) like anything else.
|
||
|
||
### Working rules
|
||
|
||
As Run 1, unchanged: work in `/Users/danielsamson/Gitea/daniel/danos` on
|
||
`claude/bounds-track`; every step lands with a test that fails before the fix, verified
|
||
by restoring the old behaviour; full suite green before each commit; never two suites at
|
||
once (`pgrep -f qemu_test.py`); 60 GiB free; `git commit -F` with no `Co-Authored-By`;
|
||
update this table before starting the next step. **If a step needs a decision that is
|
||
not written down, stop it, add the question below, and move on** — Run 1's first step
|
||
was wrong and stopping was the right call.
|
||
|
||
**Suite:** 115/115 at the start of the run.
|
||
**Branch:** `claude/bounds-track`.
|
||
|
||
### What this run deliberately does not touch
|
||
|
||
Phases 2 and 3 below — the authorisation gate and moving the inventory to the device
|
||
manager — are **out of scope for unattended work**. They decide whether the OS is
|
||
secure, and they are currently a direction rather than a specification: what a device
|
||
capability *is*, which syscalls change, what replaces `device_claim` for its seven
|
||
callers, how a driver spawned bare behaves. Those want a design session, the way
|
||
`/protocol` had one.
|
||
|
||
Also out of scope: anything touching `maximum_device_resources` (a wire struct, so a
|
||
trust-boundary change, not a resize), and the non-device bounds the audit found in FAT,
|
||
the VFS, logger, init, display and boot.
|
||
|
||
### Open questions this run must not answer on its own
|
||
|
||
Recorded here rather than guessed. If a step runs into one, it stops and writes the
|
||
question down instead of inventing an answer.
|
||
|
||
1. **`device_enumerate` probably narrows rather than retires.** The device manager
|
||
calls it to find `pci_host_bridge` nodes — it cannot ask itself. The likely shape is
|
||
that the kernel keeps the *firmware-discovered roots* (which by principle 5 it holds
|
||
for real reasons, since they come from ACPI rather than a driver's say-so) and
|
||
everything a driver registered lives in the manager. Not decided.
|
||
2. **A device-manager restart has no re-enumerate handshake.** If only the manager
|
||
dies, the buses are alive and never re-send `child_added`, so a restarted manager
|
||
comes back blind. The manager is restartable by design; nothing implements this.
|
||
3. **Which adversarial tests I1–I3 need.** The audit's six real defects were all found
|
||
by asking what an attacker would do, and the suite had never asked. "Add adversarial
|
||
cases" is not executable until the attacks are named.
|
||
4. **Reclamation is not a death-sweep problem, and L1 as written would have broken the
|
||
restart path.** Found on the first attempt at it. The audit is right that `count`
|
||
never decreases, but *death is the wrong trigger*:
|
||
- The broker keeps entries deliberately: "The devices stay in the table — they
|
||
describe hardware, which did not go away — only their ownership clears." A driver
|
||
dying does not unplug anything.
|
||
- Device ids must stay **stable across a bus restart**, because
|
||
`device-manager.driverForDevice` dedupes by `device_id` so that "a re-report after
|
||
a bus restart must not spawn a second instance". Stability comes from the
|
||
idempotency scan returning the existing id — removing entries on death would give
|
||
a restarted bus fresh ids and spawn duplicate driver instances.
|
||
- Everything else a task holds *is* already reclaimed on every path out:
|
||
`irq.releaseOwner`, `iommu.releaseAllOwnedBy`, `dmaRegistryReleaseOwner`, then the
|
||
broker's claims (`process.releaseTaskResourcesLocked`).
|
||
|
||
So the real leak has two sources, and neither is death: a device that genuinely
|
||
**goes away** (hot-unplug) has no retirement path, and a bus that enumerates
|
||
*differently* on restart leaves its stale entries behind forever. Both are the device
|
||
manager's inventory problem — phase 3 — and both need the id-stability question
|
||
answered first (tombstone-and-reuse aliases stale ids held by another process;
|
||
generation-tagged ids change the id encoding, which is ABI). Not an unattended
|
||
decision.
|
||
5. **Driver descriptor parsing cannot be host-tested, so L5's correctness fix ships
|
||
without a direct test.** The endpoint-misattribution bug lives in
|
||
`parseConfiguration`, a pure function over a byte blob — exactly the shape a host
|
||
unit test wants, and `usb-storage/scsi.zig` and `usb-hid/hid-report.zig` already do
|
||
this. But `usb-xhci-library.zig` imports `memory`, `mmio` and `time`, so it cannot
|
||
be a standalone host-test root, and QEMU offers no device that would exercise the
|
||
path anyway: the largest available is `usb-audio,multi=on` at 2 interfaces and 211
|
||
bytes, against a cap of 4.
|
||
|
||
Three ways out, and picking one is a judgement about house style rather than a
|
||
mechanical step: extract the parser to its own file and wire `usb-abi` into a test
|
||
module (build-support currently resolves module names only for `userBinary`);
|
||
extract it and import `usb-abi` by relative path (against the import-by-name
|
||
convention); or accept QEMU-only coverage and say so.
|
||
|
||
Mitigating, and the reason this is recorded rather than blocking: after the fix the
|
||
bug is unreachable **by construction**, not by the added `else`. Interfaces are now
|
||
allocated to exactly the count the descriptor declares, so `interface_count` can
|
||
never reach `interfaces.len` mid-parse. The `else` is belt-and-braces for the
|
||
255-interface clamp. The alternate-setting path that shares it *is* exercised —
|
||
`usb-audio` has alternate settings, and the `usb-large-descriptor` case walks them.
|
||
|
||
### Working rules for the run
|
||
|
||
- Work in `/Users/danielsamson/Gitea/daniel/danos` (not a worktree), on
|
||
`claude/bounds-track`.
|
||
- **Every step lands with a test that fails before the fix**, verified by temporarily
|
||
restoring the old behaviour and watching exactly the intended assertion flip. A test
|
||
that passes both ways is not a test.
|
||
- Full QEMU suite green before each commit. Never run two suites at once — check
|
||
`pgrep -f qemu_test.py` first; a second concurrent run produces false triple faults
|
||
because both share `zig-out`.
|
||
- Check at least 60 GiB free before starting a suite.
|
||
- Commit with `git commit -F <file>`, never `-m` (a backtick in a message is executed
|
||
by the shell and silently eats a word). No `Co-Authored-By` trailers.
|
||
- Update the Live state table **before** starting the next step.
|
||
- If a step needs a decision that is not written down here, stop, add it to the open
|
||
questions above, and move to the next step.
|
||
|
||
---
|
||
|
||
## The principles this is derived from
|
||
|
||
1. **danOS is a microkernel.** Minimise what the kernel is responsible for; move
|
||
responsibility to user space so it can be restarted, or fixed live during
|
||
development, without taking the system down.
|
||
2. **Implement the specifications correctly**, with the limits those specifications
|
||
define — not limits we decide.
|
||
3. **Move as much responsibility as possible to user space** (the device manager).
|
||
4. **What remains in the kernel is minimal.**
|
||
5. **What remains in the kernel is there for security or for a hardware limitation.**
|
||
Nothing else earns a place.
|
||
|
||
Principle 5 is the test every bound is put to. For each one: *is this here because of
|
||
security, or because of a hardware limitation?* If neither, the storage does not belong
|
||
in the kernel and the bound is not a number to be resized — it is a thing to be moved or
|
||
deleted.
|
||
|
||
Applying it to the case that started this:
|
||
|
||
- `maximum_devices = 64` bounds an inventory of hardware. An inventory is neither a
|
||
security control nor a hardware limitation. **The table is in the wrong place**; the
|
||
number is a symptom.
|
||
- `maximum_children_per_parent = 16` exists because `device_claim` is unauthenticated —
|
||
any process can claim any unclaimed device ([devices-broker.zig:164](../system/kernel/devices-broker.zig:164)
|
||
checks only that the device exists and is free). The cap is a crude proxy for an
|
||
authorisation the kernel does not perform. **Fix the authorisation and the cap has
|
||
nothing to defend.**
|
||
- `maximum_domains = 64` bounds IOMMU translation domains. Security — stays in the
|
||
kernel. But VT-d and AMD-Vi both *report* how many domains they support in a
|
||
capability register. Principle 2: read it. We chose 64 without asking.
|
||
|
||
## The security invariants
|
||
|
||
Every phase must leave all five standing. This is the "without punching a hole" half of
|
||
the brief, and each phase below states how it is checked.
|
||
|
||
- **I1 Containment.** A process may map only physical memory inside a resource it was
|
||
granted. A bus may subdivide only what it already holds.
|
||
- **I2 Confinement.** A DMA-capable device is under IOMMU translation before its driver
|
||
can program it, or it is not driven at all.
|
||
- **I3 No self-granted authority.** A process holds what it was handed. It cannot name
|
||
its way into holding more.
|
||
- **I4 Death releases everything.** Every resource a task held is reclaimed when it
|
||
dies, on every path out.
|
||
- **I5 Refusal is attributable.** Every refusal names the rule that refused it.
|
||
|
||
## Phase 0 — Done
|
||
|
||
- **Errno attribution.** One errno space in `system/abi.zig`; `device_register`'s six
|
||
refusals and `device_claim`'s three are distinct codes; call sites name the reason;
|
||
`pci-bus` reconciles found against registered. (I5)
|
||
- **Idempotency ordering.** A re-registration consumes no slot, so a full parent
|
||
re-admits an identical child. A restarted bus is no longer billed for what it
|
||
rediscovers.
|
||
|
||
Suite 114/114.
|
||
|
||
## Phase 1 — Reclamation
|
||
|
||
**Nothing may become dynamic before this.** Today `count` only ever increases and
|
||
`releaseAllOwnedBy` clears a dead driver's *claims* but not its *registrations*. With a
|
||
fixed table that is a slow march to the cap; with dynamic storage it is an unbounded
|
||
leak, and every supervisor restart makes it worse.
|
||
|
||
- Extend the existing death sweep so a task's registrations go with its claims.
|
||
- A registration whose owner is gone is removed; its children are re-parented or removed
|
||
with it (they cannot outlive the authority that published them).
|
||
- Test: register under a claimed parent, kill the owner, assert the entries are gone and
|
||
the ids are not reused while any handle to them lives.
|
||
|
||
Invariant: **I4**.
|
||
|
||
## Phase 2 — Close the authorisation hole
|
||
|
||
The device manager already decides which driver gets which device — it matches against
|
||
`devices.csv` and spawns the driver with the device id as `argv[1]`. Nothing binds that
|
||
decision to the kernel's `claim`. A driver passes an integer; the kernel checks only
|
||
that the device is free.
|
||
|
||
Per principles 3 and 5: **the decision stays in user space; the kernel enforces only
|
||
possession.** The manager hands the driver the device it matched; the kernel's job is
|
||
that a driver holds what it was handed and nothing else.
|
||
|
||
- The manager passes a device to the driver it spawned, over the existing cap-passing
|
||
path. Possession is the authority.
|
||
- `device_claim` stops being a way to *acquire* a device by naming it.
|
||
- Exclusivity stops being a broker refusing a second claimant and becomes the ordinary
|
||
property of a thing only one process was given.
|
||
|
||
**`maximum_children_per_parent` is deleted here**, because after this a bus driver's
|
||
children are the devices it actually enumerated under a bus it was actually given, and
|
||
the rogue-driver-fills-the-table threat the cap was written for no longer exists.
|
||
|
||
Invariants: **I3** (the point of the phase), **I1** (containment is unchanged and still
|
||
checked on every subdivision), **I5**.
|
||
|
||
Acceptance: a driver that names a device it was not given is refused, with its own
|
||
errno. The QEMU suite gains an adversarial case for it — the audit's lesson was that
|
||
"the suite contains no attacker".
|
||
|
||
## Phase 3 — The inventory moves to user space
|
||
|
||
The kernel reads only three things out of a device descriptor: **physical ranges** (to
|
||
check a mapping falls inside one), **interrupt numbers**, and **one PCI BDF** (to key an
|
||
IOMMU domain). Vendor and device ids, class triples, subsystem ids, human-readable
|
||
names, bus numbers and parent links are stored solely so `device_enumerate` can hand
|
||
them back. That is the kernel acting as a distribution mechanism for data it does not
|
||
use — principle 5 excludes it.
|
||
|
||
- **Zero-resource devices leave the kernel entirely.** A USB device addressed through
|
||
its controller conveys no mapping authority; there is nothing for the kernel to
|
||
enforce. It is pure inventory and belongs to the device manager. (This is also the
|
||
case that sidesteps containment, which is why the cap existed.)
|
||
- Identity and topology move to the manager, which already receives them as
|
||
`child_added` reports and already holds the authoritative picture.
|
||
- `device_enumerate` retires; callers ask the manager, whose protocol already reserves
|
||
an `enumerate` verb. Public-ABI change — `docs/os-development/vdso.md` documents it.
|
||
- What the kernel keeps: for each device that carries resources, the ranges, the GSIs,
|
||
the BDF, and the owner.
|
||
|
||
After this, the kernel's table holds only resource-bearing devices, and the remaining
|
||
count is bounded by what the machine physically has rather than by us.
|
||
|
||
Invariants: **I1**, **I2** unchanged — both operate on resources, which do not move.
|
||
**I4** must be re-checked: the manager's table now needs its own reclamation, and it is
|
||
restartable, so it must be able to rebuild from the buses.
|
||
|
||
## Phase 4 — Ask the hardware and the specification
|
||
|
||
Principle 2, applied to every remaining bound. Each of these is a number the machine or
|
||
the standard already states, which we replaced with a guess. Independent of each other;
|
||
can proceed in any order.
|
||
|
||
| Today | Ask instead |
|
||
|---|---|
|
||
| `maximum_domains = 64` (IOMMU) | the VT-d / AMD-Vi capability register reports the domains supported |
|
||
| `max_devices = 8` (xHCI slots) | `HCSPARAMS1.MaxSlots` — the controller says (1–255) |
|
||
| `max_interfaces = 4`, `max_endpoints` | the configuration descriptor says |
|
||
| `blob: [512]u8` (USB config) | the device's `wTotalLength` |
|
||
| `below: [64]Range` (memory map) | UEFI reports the descriptor count |
|
||
| AML blobs capped at 6 | the XSDT's length field gives the entry count |
|
||
| `maximum_cpus = 128` | the MADT entry count |
|
||
| `maximum_gsi = 24` | the I/O APIC's redirection-entry count; and more than one I/O APIC exists |
|
||
| MSI-X vectors | the capability's table-size field (up to 2048) |
|
||
|
||
Several of these are in user space already (xHCI, USB descriptors) and are ordinary
|
||
allocations — principle 1 means those are also the safest to do first, since a mistake
|
||
restarts a driver rather than the machine.
|
||
|
||
Two in this table are **also** correctness fixes the audit found, and should carry their
|
||
regression tests: the xHCI `max_interfaces` path misattributes a fifth interface's
|
||
endpoints to interface 3, and `below: [64]Range` silently turns occupied RAM into a PCI
|
||
aperture — which is an **I1 violation reachable on real hardware**, not merely a lost
|
||
device. That one is the highest-priority item in this phase.
|
||
|
||
## Phase 5 — What legitimately remains
|
||
|
||
After phases 1–4 the survivors should be only:
|
||
|
||
- **Pre-allocator storage**: the PMM's own frame bitmap, the memory map the loader hands
|
||
over, the bootstrap page tables. You cannot allocate the allocator. (Hardware/boot
|
||
limitation — principle 5 admits these.)
|
||
- **Interrupt-context storage**: the IST stack and anything an exception path touches
|
||
without allocating.
|
||
- **Wire structures** whose layout the other side of a trust boundary parses.
|
||
- **Facts that are not ceilings**: a page is 4096 bytes; an ACPI name segment is 4.
|
||
|
||
Each is declared per [bounds.md](os-development/bounds.md) — what it counts, who decides
|
||
its size, what it protects, what happens at the limit, how you find out. And two numbers
|
||
that must agree agree in code, not in a comment:
|
||
|
||
```zig
|
||
comptime {
|
||
if (maximum_domains != devices_broker.maximum_devices)
|
||
@compileError("iommu.confined is indexed by device id; an id past its end is " ++
|
||
"left unconfined while confineDevice still reports success");
|
||
}
|
||
```
|
||
|
||
## The one that must not wait
|
||
|
||
[`iommu.zig:107`](../system/kernel/iommu.zig:107) — `if (device_id >= confined.len) return
|
||
true;` — returns *success* without confining. It is unreachable today only because
|
||
device ids stop at 64. **Phases 3 and 4 both change the device count, and either makes
|
||
it live.** Fix it before them: out of range must refuse, never allow. (I2)
|
||
|
||
This is also the standing rule the audit argues for: at a bound, the safe direction is
|
||
refusal. A ceiling that fails open is not a limit, it is a switch that turns the
|
||
protection off.
|
||
|
||
## How this is verified
|
||
|
||
- The QEMU suite is the arbiter at every step; it is 114 cases and must stay green.
|
||
- Each fix lands with a test that **fails before it** — as the idempotency reorder did,
|
||
where exactly one assertion flipped.
|
||
- Adversarial cases for I1–I3 specifically: the audit's six real defects were all found
|
||
by asking "what would an attacker do", and the suite had never asked.
|
||
- The Ryzen is the acceptance test. It is the machine that found this, and the one that
|
||
proves it fixed.
|
||
|
||
## Sequencing
|
||
|
||
Phase 1 gates everything. Phase 2 gates phase 3 — the inventory cannot move until
|
||
authority is sound, or moving it is the hole. Phase 4 is independent and its user-space
|
||
items are the safest work in the track. Phase 5 is the record of what survived.
|
||
|
||
The IOMMU fail-open is fixed before phase 3 or 4 touches the device count.
|