Third driver converted. It no longer claims the id from argv[1] — the manager holds the device and names it in the call that creates the process, so it is held before the driver's first instruction. display-reattach is the case that matters here: it kills the driver and watches the compositor re-attach to the fresh scanout. It passes, so the restart path survives the fused grant — the manager re-takes the device when the driver dies and hands it to the replacement. ps2-bus and discovery are NOT converted, and the reason is recorded as open question 9 rather than worked around. Both need a device nobody assigned them. ps2-bus ignores its argv[1] entirely: it finds the controller by walking the table for PNP0303, then claims a second device, the PNP0F13 mouse node, which it also finds itself — so it holds two devices and was assigned at most one, while system_spawn carries one. discovery claims the acpi-tables node it locates itself, because it is what produces the device tree and there is nothing to assign at that point. One thing worth checking before designing an answer: devices.csv maps both PS/2 hardware ids to ps2-bus, so the manager may already be spawning two instances where the driver expects one. If so the fix is smaller than it looks. D6 stays blocked — closing device_claim with these two still depending on it would stop the machine booting. Suite 118/118.
548 lines
32 KiB
Markdown
548 lines
32 KiB
Markdown
# The bounds track: removing the numbers we invented
|
||
|
||
*Plan, 2026-08-08. Follows [fixed-bounds-audit.md](fixed-bounds-audit.md) (235 ceilings,
|
||
139 on quantities we do not choose) and the AMD Ryzen that found the first one.*
|
||
|
||
---
|
||
|
||
## Live state — the unattended run
|
||
|
||
*This table is the progress view. It is updated at the end of every step, before the
|
||
next one starts.*
|
||
|
||
| Step | What | State |
|
||
|---|---|---|
|
||
| L1 | Reclamation: a dead task's registrations die with its claims | **stopped — the step was wrong; see open question 4** |
|
||
| L2 | Bounds build check + allowlist; declare what we have already touched | **done** — `zig build bounds`, 273 allowlisted, 5 declared |
|
||
| L3 | xHCI: slot count from `HCSPARAMS1.MaxSlots`, not 8 | **done** — QEMU reports 64; the driver tracked 8 |
|
||
| L4 | USB: configuration descriptor sized by `wTotalLength`, not 512 | **done** — QEMU tops out at 211 bytes, so the case catches the class, not the original trigger |
|
||
| L5 | USB: interfaces from the descriptor, and the misattributed-endpoint bug | **done** — fix is by construction; no direct test, see open question 5 |
|
||
| L6 | xHCI: a failed `allocateDevice` stops leaking an enabled slot | **done** — path forced and verified; no regression test, see open question 5 |
|
||
|
||
**Run 1 complete.** L1 stopped (the step was wrong), L2–L6 landed. Suite 115 → 116.
|
||
Allowlist 278 → 269. Two steps ship without a permanent regression test, both because
|
||
QEMU's USB devices are too small to reach the paths — see open question 5, which is the
|
||
audit's own lesson recurring: the test rig is smaller than a real machine.
|
||
|
||
---
|
||
|
||
## Run 2 — device authority: delete the invented ceilings
|
||
|
||
*Design: [device-authority.md](os-development/device-authority.md), which is the **how**
|
||
for the delegation step [device-manager.md](device-driver-development/device-manager.md)
|
||
already settled. Read both before starting; the second is authoritative where they
|
||
differ.*
|
||
|
||
The goal, in the project owner's words: **remove the maximum values we set arbitrarily,
|
||
move the responsibility to the device manager, and keep in the kernel only the parts
|
||
that cannot safely run in user space.**
|
||
|
||
| Step | What | State |
|
||
|---|---|---|
|
||
| D1 | `device_transfer(device_id, task_id)` — the holder gives a device away | **done** — syscall 54; a move, not a copy |
|
||
| D2 | Adversarial case: a process handed nothing is refused, on a held device and a free one | **done** — `device-authority-test`; the claim half joins it at D6 |
|
||
| D3 | The manager claims the seeded devices at boot, before any driver is spawned | **merged into D4** — see below |
|
||
| D4 | The manager claims + delegates on `hello`; `usb-xhci-bus` is the first driver converted | **done** — caught an IOMMU regression I introduced; see below |
|
||
| D5 | The other four claimants converted: `pci-bus`, `ps2-bus`, `virtio-gpu`, `acpi` | **partial** — `pci-bus` + `virtio-gpu` done; `ps2-bus` and discovery need question 9 |
|
||
| D0 | The grant rides `system_spawn` — atomic, so no driver need change to receive one | **done** — and caught a test-marker bug that made the D2 fixture unfailable |
|
||
| D10 | Every driver hellos, on its own merits (liveness, one class of driver) | not started — optional, independent |
|
||
| D6 | `device_claim` refuses a device the caller was not handed; the hole is closed | not started |
|
||
| D7 | Zero-resource devices stop being kernel objects — inventory moves to the manager | **blocked** — nothing else mints their ids; see question 8 |
|
||
| D8 | **`maximum_children_per_parent` deleted** — the authorisation it stood in for exists | **blocked on D6**, and now ordered after D9 |
|
||
| D9 | The device table becomes dynamic; **`maximum_devices` deleted**; per-holder quota declared | **done** — one of the two invented numbers is gone |
|
||
|
||
**Run 2 resumes at D0.** D1, D2, D4, D5 (`pci-bus` only) and D9 landed; `maximum_devices`
|
||
no longer exists and the suite is 118/118. Questions 6 and 7 dissolved, so the order is
|
||
now **D0 → D5 → D6 → D8**, which deletes `maximum_children_per_parent`. Only D7 is still
|
||
blocked, on question 8, and it is needed for neither ceiling. D10 is optional.
|
||
|
||
Ordering is load-bearing. D1–D2 built and proved the mechanism with nothing depending on
|
||
it. D4–D5 move each claimant across one at a time, so the suite stays green throughout
|
||
and a regression names the driver that caused it. D6 is the flag day. (The original
|
||
"D7 must precede D9" no longer holds: the per-holder quota bounds zero-resource children
|
||
as well as anything else, so D9 landed without it.)
|
||
|
||
**D3 merged into D4** (found while implementing, recorded rather than worked around).
|
||
The two cannot be separated: the moment the manager claims a device, any driver still
|
||
calling `device_claim` on it is refused `AlreadyClaimed`, so D3 on its own turns the
|
||
suite red — and D3 applied to *nothing* changes no behaviour and cannot be tested.
|
||
They land together, with the manager claiming only for drivers in an explicit
|
||
**delegated set** so every unconverted driver keeps claiming exactly as before.
|
||
`usb-xhci-bus` is the first member, as it was the first driver to conform to `hello`
|
||
(device-manager.md, M18.1). D5 moves the rest in one at a time; the set and the
|
||
`device_claim` path both disappear at D6.
|
||
|
||
### Observations from the run
|
||
|
||
- **D4 introduced an IOMMU regression, caught by converting one driver at a time.**
|
||
`confineDevice` runs inside `systemDeviceClaim`, so a device arriving by *transfer*
|
||
was never confined for its new owner: the driver's DMA rings went unbound (three
|
||
IOMMU+USB cases failed), and two worse consequences were latent — a manager death
|
||
would have torn down a domain a live driver was using, and a driver death would have
|
||
leaked one. `iommu.reassign` moves the confinement with the device, keeping the
|
||
domain and attachment intact so it never translates through nothing. Converting all
|
||
five drivers at once would have produced the same three failures with five suspects.
|
||
- **`usb-hub` failed once in a full run, then passed six times** (four isolated, two
|
||
full). Suspected instance of the known intermittent AP ring-3 fault rather than
|
||
anything in D4 — recorded rather than dismissed, because D4 moved the `hello` earlier
|
||
and so did shift boot timing. Watch it across the remaining steps.
|
||
|
||
### Open question 9 — two drivers need a device nobody assigned them
|
||
|
||
D0 made delivery atomic and D5 converted `virtio-gpu` with it. The last two do not fit,
|
||
for the same underlying reason: **they need a device the manager never assigned.**
|
||
|
||
- **`ps2-bus` ignores its `argv[1]` entirely.** It finds the controller by walking the
|
||
table for `PNP0303`, and then claims a *second* device — the `PNP0F13` mouse node —
|
||
which it also finds itself. So it holds two devices and was assigned at most one, and
|
||
`system_spawn` carries one.
|
||
- **discovery** is spawned `addDriver("discovery", no_device, false)` and claims the
|
||
`acpi-tables` node it locates itself, because it is what produces the device tree;
|
||
there is nothing to assign at that point.
|
||
|
||
Three shapes of answer, none written down:
|
||
|
||
1. **The manager assigns every device a driver needs.** For `ps2-bus` it would have to
|
||
understand that one driver serves both `PNP0303` and `PNP0F13` — `devices.csv` maps
|
||
both to it already, so the manager may simply be spawning two instances today where
|
||
the driver expects one. Worth checking before designing.
|
||
2. **A driver asks for a device over the protocol** — a `request(device)` verb the
|
||
manager answers by transferring. General, and it reintroduces a window, though only
|
||
for a device the driver asks for after it is running.
|
||
3. **The bootstrap is exempt**: discovery keeps claiming, and D6 permits a claim of a
|
||
device nobody holds *only* for the discovery role. Narrow and honest, but it is an
|
||
exception in the exact place an exception is most expensive.
|
||
|
||
**D6 stays blocked** until this is answered — closing `device_claim` with these two
|
||
still depending on it would stop the machine booting.
|
||
|
||
### Settled 2026-08-08: the grant rides `system_spawn` (D0)
|
||
|
||
Questions 6 and 7 both dissolved on inspection — neither `ps2-bus` nor discovery needs
|
||
to start speaking `hello`, and `virtio-gpu` has no standalone path to lose. What remains
|
||
is *where the grant is delivered*, and there are three candidates:
|
||
|
||
| | Race window | Cost |
|
||
|---|---|---|
|
||
| Transfer after spawn | **yes** | none |
|
||
| Every driver hellos | no | `ps2-bus` + discovery gain a handshake |
|
||
| **Grant rides `system_spawn`** | **no — atomic** | one more syscall argument |
|
||
|
||
**Take the third.** The manager cannot transfer before the child exists, so a separate
|
||
transfer always leaves a window in which the child is running and does not yet hold its
|
||
device. It would close on QEMU every time and open occasionally on a machine with
|
||
different timing — the exact failure shape this track exists to delete, and not worth
|
||
introducing while removing the others. Fusing the device into the spawn removes it by
|
||
construction: the child does not exist until it holds the device. No new knowledge in
|
||
the kernel — the same rule, *you may give away what you hold*, made atomic with the call
|
||
that creates the recipient. `system_spawn` uses five of six argument registers, so there
|
||
is room, and `no_device` is already the sentinel for a driver with no assignment.
|
||
|
||
**`hello` for every driver is a good idea on its own merits** — uniform liveness, the
|
||
deadline applied to all rather than some, and the `speaks_protocol` two-class split
|
||
leaving the manager (a wedged `ps2-bus` is invisible to its supervisor today). It is
|
||
D10, kept separate so grant delivery does not force it.
|
||
|
||
### Open questions raised by D5 — resolved except question 8
|
||
|
||
Delegation is delivered in `onHello`. That works for a driver that says hello, and
|
||
**two of the four do not**.
|
||
|
||
6. **`ps2-bus` does not hello**, and that is deliberate: device-manager.md records
|
||
"Legacy drivers (e.g. ps2-bus) are supervised and restarted but **not yet required
|
||
to hello**", and `Driver.speaks_protocol` exists to say so. Either it leaves
|
||
"legacy", or the grant gets a delivery point that is not `hello`.
|
||
|
||
**The `acpi` service is not this case — an earlier version of this question wrongly
|
||
lumped it in.** It is spawned as `addDriver("discovery", no_device, false)`: the
|
||
*discovery service*, one per firmware, with **no device assignment at all**, which
|
||
"finds and claims the acpi-tables (or devicetree-blob) node itself. Not a per-device
|
||
driver." The manager cannot hand it a device because discovery is what produces the
|
||
device tree — there is nothing to match against yet. Its question is not "should it
|
||
hello" but "is the bootstrap exempt from D6, or does the manager claim the
|
||
acpi-tables node and pass it on".
|
||
|
||
**A delivery point that needs no hello already exists in the shape of the code**:
|
||
`process.spawnSupervised` returns the child's pid, so the manager could transfer
|
||
immediately after spawn. If that is acceptable, this question mostly dissolves —
|
||
`ps2-bus` would not need to leave "legacy" and discovery could be handed its node
|
||
too. The cost is ordering: the child may reach for the device before the transfer
|
||
lands, where `hello` guarantees it cannot because the child is the one asking. That
|
||
trade is the actual decision.
|
||
|
||
7. **`virtio-gpu`'s "standalone bring-up" — resolved; there is no such path.** An
|
||
earlier version of this question read the comment "Best-effort: standalone bring-up
|
||
has no manager" as a boot path that delegation would delete. It is not one.
|
||
`virtio-gpu`'s `main` requires `argv[1]` and exits without it, and the only source of
|
||
that argument is the device manager — the chain is ACPI → pci-bus reports the
|
||
function → `devices.csv` matches `1AF4:1050` → the manager spawns the driver with the
|
||
id. There is no way to run it without a manager, so nothing is lost.
|
||
|
||
What the comment is really about is **resilience**: the hello is best-effort so a
|
||
driver whose manager has *died* keeps serving, which device-manager.md states
|
||
("If the manager dies, drivers keep running").
|
||
|
||
The residual is one narrow window: the manager spawns a driver and dies before
|
||
transferring. Today the driver could still claim, because claiming is free-for-all;
|
||
after D6 it exits and the restarted manager respawns it. That is the better
|
||
behaviour — a driver holding hardware nobody assigned it is what D6 exists to stop.
|
||
|
||
Until these are answered, `device_claim` cannot be closed off at D6 for those three, so
|
||
**D6 is blocked on questions 6 and 7**.
|
||
|
||
8. **Removing zero-resource devices from the kernel needs someone else to mint their
|
||
ids, and nothing says who.** A USB interface is registered with
|
||
`resource_count = 0`, and the id `device_register` returns is load-bearing in three
|
||
places: it is the `child_added` packet's target, it is the class driver's `argv[1]`,
|
||
and it is the `device_token` of the **usb-transfer wire protocol** — so the id space
|
||
is visible on the wire, not merely internal. usb-xhci-bus's own comment records the
|
||
fourth constraint: the kernel's idempotency is what makes "the same port and
|
||
interface always map back to the same device id" across a bus restart, which is what
|
||
stops a respawned bus spawning duplicate class drivers.
|
||
|
||
So the mover has to answer: who mints the id, how it stays stable across a *bus*
|
||
restart, how it stays stable across a *manager* restart (open question 3 territory),
|
||
and whether the wire protocol's `device_token` changes meaning. That is a design
|
||
step, not a mechanical one.
|
||
|
||
**D8 is reordered to run after D9**, which is a correction to the original sequencing.
|
||
D8's justification was "the authorisation it stood in for exists" — but with D6 blocked
|
||
it does not, so deleting the shared per-parent cap now would reopen the exhaustion hole
|
||
it was written for. D9's **per-holder quota** closes that hole independently of
|
||
authorisation, and does it better: a rogue exhausts its own allowance rather than the
|
||
table everyone shares. Once the quota exists, the per-parent cap is redundant whether or
|
||
not D6 has landed.
|
||
|
||
D9 also no longer depends on D7. Its original rationale was that zero-resource children
|
||
are the case that sidesteps containment — true, but a per-holder quota bounds them just
|
||
as well as anything else, because it counts entries per holder rather than per parent.
|
||
|
||
### Settled, so the run does not re-litigate them
|
||
|
||
- **The manager claims, it is not granted.** No binary names in the kernel; the rule is
|
||
"you may give away what you hold". The residual — it rests on the manager claiming
|
||
first — is stated in the design and is closed later by the spawn capability
|
||
[drivers.md](device-driver-development/drivers.md) already names as missing.
|
||
- **The framebuffer is not a device.** It is where pixels go, handed over by the loader,
|
||
and the compositor uses it as the boot floor until a real display driver announces
|
||
itself. Nothing in this run touches the display service or its GOP path.
|
||
- **`maximum_endpoints_per_interface` and the wire structs stay.** Widening them is a
|
||
protocol change, out of scope.
|
||
- **A per-holder quota is not a retreat.** Dynamic storage with no bound moves the
|
||
ceiling to the kernel heap, which is shared and fatal rather than partial. A bound
|
||
charged to the task that caused it is isolation, and it is declared through
|
||
[bounds.md](os-development/bounds.md) like anything else.
|
||
|
||
### Working rules
|
||
|
||
As Run 1, unchanged: work in `/Users/danielsamson/Gitea/daniel/danos` on
|
||
`claude/bounds-track`; every step lands with a test that fails before the fix, verified
|
||
by restoring the old behaviour; full suite green before each commit; never two suites at
|
||
once (`pgrep -f qemu_test.py`); 60 GiB free; `git commit -F` with no `Co-Authored-By`;
|
||
update this table before starting the next step. **If a step needs a decision that is
|
||
not written down, stop it, add the question below, and move on** — Run 1's first step
|
||
was wrong and stopping was the right call.
|
||
|
||
**Suite:** 115/115 at the start of the run.
|
||
**Branch:** `claude/bounds-track`.
|
||
|
||
### What this run deliberately does not touch
|
||
|
||
Phases 2 and 3 below — the authorisation gate and moving the inventory to the device
|
||
manager — are **out of scope for unattended work**. They decide whether the OS is
|
||
secure, and they are currently a direction rather than a specification: what a device
|
||
capability *is*, which syscalls change, what replaces `device_claim` for its seven
|
||
callers, how a driver spawned bare behaves. Those want a design session, the way
|
||
`/protocol` had one.
|
||
|
||
Also out of scope: anything touching `maximum_device_resources` (a wire struct, so a
|
||
trust-boundary change, not a resize), and the non-device bounds the audit found in FAT,
|
||
the VFS, logger, init, display and boot.
|
||
|
||
### Open questions this run must not answer on its own
|
||
|
||
Recorded here rather than guessed. If a step runs into one, it stops and writes the
|
||
question down instead of inventing an answer.
|
||
|
||
1. **`device_enumerate` probably narrows rather than retires.** The device manager
|
||
calls it to find `pci_host_bridge` nodes — it cannot ask itself. The likely shape is
|
||
that the kernel keeps the *firmware-discovered roots* (which by principle 5 it holds
|
||
for real reasons, since they come from ACPI rather than a driver's say-so) and
|
||
everything a driver registered lives in the manager. Not decided.
|
||
2. **A device-manager restart has no re-enumerate handshake.** If only the manager
|
||
dies, the buses are alive and never re-send `child_added`, so a restarted manager
|
||
comes back blind. The manager is restartable by design; nothing implements this.
|
||
3. **Which adversarial tests I1–I3 need.** The audit's six real defects were all found
|
||
by asking what an attacker would do, and the suite had never asked. "Add adversarial
|
||
cases" is not executable until the attacks are named.
|
||
4. **Reclamation is not a death-sweep problem, and L1 as written would have broken the
|
||
restart path.** Found on the first attempt at it. The audit is right that `count`
|
||
never decreases, but *death is the wrong trigger*:
|
||
- The broker keeps entries deliberately: "The devices stay in the table — they
|
||
describe hardware, which did not go away — only their ownership clears." A driver
|
||
dying does not unplug anything.
|
||
- Device ids must stay **stable across a bus restart**, because
|
||
`device-manager.driverForDevice` dedupes by `device_id` so that "a re-report after
|
||
a bus restart must not spawn a second instance". Stability comes from the
|
||
idempotency scan returning the existing id — removing entries on death would give
|
||
a restarted bus fresh ids and spawn duplicate driver instances.
|
||
- Everything else a task holds *is* already reclaimed on every path out:
|
||
`irq.releaseOwner`, `iommu.releaseAllOwnedBy`, `dmaRegistryReleaseOwner`, then the
|
||
broker's claims (`process.releaseTaskResourcesLocked`).
|
||
|
||
So the real leak has two sources, and neither is death: a device that genuinely
|
||
**goes away** (hot-unplug) has no retirement path, and a bus that enumerates
|
||
*differently* on restart leaves its stale entries behind forever. Both are the device
|
||
manager's inventory problem — phase 3 — and both need the id-stability question
|
||
answered first (tombstone-and-reuse aliases stale ids held by another process;
|
||
generation-tagged ids change the id encoding, which is ABI). Not an unattended
|
||
decision.
|
||
5. **Driver descriptor parsing cannot be host-tested, so L5's correctness fix ships
|
||
without a direct test.** The endpoint-misattribution bug lives in
|
||
`parseConfiguration`, a pure function over a byte blob — exactly the shape a host
|
||
unit test wants, and `usb-storage/scsi.zig` and `usb-hid/hid-report.zig` already do
|
||
this. But `usb-xhci-library.zig` imports `memory`, `mmio` and `time`, so it cannot
|
||
be a standalone host-test root, and QEMU offers no device that would exercise the
|
||
path anyway: the largest available is `usb-audio,multi=on` at 2 interfaces and 211
|
||
bytes, against a cap of 4.
|
||
|
||
Three ways out, and picking one is a judgement about house style rather than a
|
||
mechanical step: extract the parser to its own file and wire `usb-abi` into a test
|
||
module (build-support currently resolves module names only for `userBinary`);
|
||
extract it and import `usb-abi` by relative path (against the import-by-name
|
||
convention); or accept QEMU-only coverage and say so.
|
||
|
||
Mitigating, and the reason this is recorded rather than blocking: after the fix the
|
||
bug is unreachable **by construction**, not by the added `else`. Interfaces are now
|
||
allocated to exactly the count the descriptor declares, so `interface_count` can
|
||
never reach `interfaces.len` mid-parse. The `else` is belt-and-braces for the
|
||
255-interface clamp. The alternate-setting path that shares it *is* exercised —
|
||
`usb-audio` has alternate settings, and the `usb-large-descriptor` case walks them.
|
||
|
||
### Working rules for the run
|
||
|
||
- Work in `/Users/danielsamson/Gitea/daniel/danos` (not a worktree), on
|
||
`claude/bounds-track`.
|
||
- **Every step lands with a test that fails before the fix**, verified by temporarily
|
||
restoring the old behaviour and watching exactly the intended assertion flip. A test
|
||
that passes both ways is not a test.
|
||
- Full QEMU suite green before each commit. Never run two suites at once — check
|
||
`pgrep -f qemu_test.py` first; a second concurrent run produces false triple faults
|
||
because both share `zig-out`.
|
||
- Check at least 60 GiB free before starting a suite.
|
||
- Commit with `git commit -F <file>`, never `-m` (a backtick in a message is executed
|
||
by the shell and silently eats a word). No `Co-Authored-By` trailers.
|
||
- Update the Live state table **before** starting the next step.
|
||
- If a step needs a decision that is not written down here, stop, add it to the open
|
||
questions above, and move to the next step.
|
||
|
||
---
|
||
|
||
## The principles this is derived from
|
||
|
||
1. **danOS is a microkernel.** Minimise what the kernel is responsible for; move
|
||
responsibility to user space so it can be restarted, or fixed live during
|
||
development, without taking the system down.
|
||
2. **Implement the specifications correctly**, with the limits those specifications
|
||
define — not limits we decide.
|
||
3. **Move as much responsibility as possible to user space** (the device manager).
|
||
4. **What remains in the kernel is minimal.**
|
||
5. **What remains in the kernel is there for security or for a hardware limitation.**
|
||
Nothing else earns a place.
|
||
|
||
Principle 5 is the test every bound is put to. For each one: *is this here because of
|
||
security, or because of a hardware limitation?* If neither, the storage does not belong
|
||
in the kernel and the bound is not a number to be resized — it is a thing to be moved or
|
||
deleted.
|
||
|
||
Applying it to the case that started this:
|
||
|
||
- `maximum_devices = 64` bounds an inventory of hardware. An inventory is neither a
|
||
security control nor a hardware limitation. **The table is in the wrong place**; the
|
||
number is a symptom.
|
||
- `maximum_children_per_parent = 16` exists because `device_claim` is unauthenticated —
|
||
any process can claim any unclaimed device ([devices-broker.zig:164](../system/kernel/devices-broker.zig:164)
|
||
checks only that the device exists and is free). The cap is a crude proxy for an
|
||
authorisation the kernel does not perform. **Fix the authorisation and the cap has
|
||
nothing to defend.**
|
||
- `maximum_domains = 64` bounds IOMMU translation domains. Security — stays in the
|
||
kernel. But VT-d and AMD-Vi both *report* how many domains they support in a
|
||
capability register. Principle 2: read it. We chose 64 without asking.
|
||
|
||
## The security invariants
|
||
|
||
Every phase must leave all five standing. This is the "without punching a hole" half of
|
||
the brief, and each phase below states how it is checked.
|
||
|
||
- **I1 Containment.** A process may map only physical memory inside a resource it was
|
||
granted. A bus may subdivide only what it already holds.
|
||
- **I2 Confinement.** A DMA-capable device is under IOMMU translation before its driver
|
||
can program it, or it is not driven at all.
|
||
- **I3 No self-granted authority.** A process holds what it was handed. It cannot name
|
||
its way into holding more.
|
||
- **I4 Death releases everything.** Every resource a task held is reclaimed when it
|
||
dies, on every path out.
|
||
- **I5 Refusal is attributable.** Every refusal names the rule that refused it.
|
||
|
||
## Phase 0 — Done
|
||
|
||
- **Errno attribution.** One errno space in `system/abi.zig`; `device_register`'s six
|
||
refusals and `device_claim`'s three are distinct codes; call sites name the reason;
|
||
`pci-bus` reconciles found against registered. (I5)
|
||
- **Idempotency ordering.** A re-registration consumes no slot, so a full parent
|
||
re-admits an identical child. A restarted bus is no longer billed for what it
|
||
rediscovers.
|
||
|
||
Suite 114/114.
|
||
|
||
## Phase 1 — Reclamation
|
||
|
||
**Nothing may become dynamic before this.** Today `count` only ever increases and
|
||
`releaseAllOwnedBy` clears a dead driver's *claims* but not its *registrations*. With a
|
||
fixed table that is a slow march to the cap; with dynamic storage it is an unbounded
|
||
leak, and every supervisor restart makes it worse.
|
||
|
||
- Extend the existing death sweep so a task's registrations go with its claims.
|
||
- A registration whose owner is gone is removed; its children are re-parented or removed
|
||
with it (they cannot outlive the authority that published them).
|
||
- Test: register under a claimed parent, kill the owner, assert the entries are gone and
|
||
the ids are not reused while any handle to them lives.
|
||
|
||
Invariant: **I4**.
|
||
|
||
## Phase 2 — Close the authorisation hole
|
||
|
||
The device manager already decides which driver gets which device — it matches against
|
||
`devices.csv` and spawns the driver with the device id as `argv[1]`. Nothing binds that
|
||
decision to the kernel's `claim`. A driver passes an integer; the kernel checks only
|
||
that the device is free.
|
||
|
||
Per principles 3 and 5: **the decision stays in user space; the kernel enforces only
|
||
possession.** The manager hands the driver the device it matched; the kernel's job is
|
||
that a driver holds what it was handed and nothing else.
|
||
|
||
- The manager passes a device to the driver it spawned, over the existing cap-passing
|
||
path. Possession is the authority.
|
||
- `device_claim` stops being a way to *acquire* a device by naming it.
|
||
- Exclusivity stops being a broker refusing a second claimant and becomes the ordinary
|
||
property of a thing only one process was given.
|
||
|
||
**`maximum_children_per_parent` is deleted here**, because after this a bus driver's
|
||
children are the devices it actually enumerated under a bus it was actually given, and
|
||
the rogue-driver-fills-the-table threat the cap was written for no longer exists.
|
||
|
||
Invariants: **I3** (the point of the phase), **I1** (containment is unchanged and still
|
||
checked on every subdivision), **I5**.
|
||
|
||
Acceptance: a driver that names a device it was not given is refused, with its own
|
||
errno. The QEMU suite gains an adversarial case for it — the audit's lesson was that
|
||
"the suite contains no attacker".
|
||
|
||
## Phase 3 — The inventory moves to user space
|
||
|
||
The kernel reads only three things out of a device descriptor: **physical ranges** (to
|
||
check a mapping falls inside one), **interrupt numbers**, and **one PCI BDF** (to key an
|
||
IOMMU domain). Vendor and device ids, class triples, subsystem ids, human-readable
|
||
names, bus numbers and parent links are stored solely so `device_enumerate` can hand
|
||
them back. That is the kernel acting as a distribution mechanism for data it does not
|
||
use — principle 5 excludes it.
|
||
|
||
- **Zero-resource devices leave the kernel entirely.** A USB device addressed through
|
||
its controller conveys no mapping authority; there is nothing for the kernel to
|
||
enforce. It is pure inventory and belongs to the device manager. (This is also the
|
||
case that sidesteps containment, which is why the cap existed.)
|
||
- Identity and topology move to the manager, which already receives them as
|
||
`child_added` reports and already holds the authoritative picture.
|
||
- `device_enumerate` retires; callers ask the manager, whose protocol already reserves
|
||
an `enumerate` verb. Public-ABI change — `docs/os-development/vdso.md` documents it.
|
||
- What the kernel keeps: for each device that carries resources, the ranges, the GSIs,
|
||
the BDF, and the owner.
|
||
|
||
After this, the kernel's table holds only resource-bearing devices, and the remaining
|
||
count is bounded by what the machine physically has rather than by us.
|
||
|
||
Invariants: **I1**, **I2** unchanged — both operate on resources, which do not move.
|
||
**I4** must be re-checked: the manager's table now needs its own reclamation, and it is
|
||
restartable, so it must be able to rebuild from the buses.
|
||
|
||
## Phase 4 — Ask the hardware and the specification
|
||
|
||
Principle 2, applied to every remaining bound. Each of these is a number the machine or
|
||
the standard already states, which we replaced with a guess. Independent of each other;
|
||
can proceed in any order.
|
||
|
||
| Today | Ask instead |
|
||
|---|---|
|
||
| `maximum_domains = 64` (IOMMU) | the VT-d / AMD-Vi capability register reports the domains supported |
|
||
| `max_devices = 8` (xHCI slots) | `HCSPARAMS1.MaxSlots` — the controller says (1–255) |
|
||
| `max_interfaces = 4`, `max_endpoints` | the configuration descriptor says |
|
||
| `blob: [512]u8` (USB config) | the device's `wTotalLength` |
|
||
| `below: [64]Range` (memory map) | UEFI reports the descriptor count |
|
||
| AML blobs capped at 6 | the XSDT's length field gives the entry count |
|
||
| `maximum_cpus = 128` | the MADT entry count |
|
||
| `maximum_gsi = 24` | the I/O APIC's redirection-entry count; and more than one I/O APIC exists |
|
||
| MSI-X vectors | the capability's table-size field (up to 2048) |
|
||
|
||
Several of these are in user space already (xHCI, USB descriptors) and are ordinary
|
||
allocations — principle 1 means those are also the safest to do first, since a mistake
|
||
restarts a driver rather than the machine.
|
||
|
||
Two in this table are **also** correctness fixes the audit found, and should carry their
|
||
regression tests: the xHCI `max_interfaces` path misattributes a fifth interface's
|
||
endpoints to interface 3, and `below: [64]Range` silently turns occupied RAM into a PCI
|
||
aperture — which is an **I1 violation reachable on real hardware**, not merely a lost
|
||
device. That one is the highest-priority item in this phase.
|
||
|
||
## Phase 5 — What legitimately remains
|
||
|
||
After phases 1–4 the survivors should be only:
|
||
|
||
- **Pre-allocator storage**: the PMM's own frame bitmap, the memory map the loader hands
|
||
over, the bootstrap page tables. You cannot allocate the allocator. (Hardware/boot
|
||
limitation — principle 5 admits these.)
|
||
- **Interrupt-context storage**: the IST stack and anything an exception path touches
|
||
without allocating.
|
||
- **Wire structures** whose layout the other side of a trust boundary parses.
|
||
- **Facts that are not ceilings**: a page is 4096 bytes; an ACPI name segment is 4.
|
||
|
||
Each is declared per [bounds.md](os-development/bounds.md) — what it counts, who decides
|
||
its size, what it protects, what happens at the limit, how you find out. And two numbers
|
||
that must agree agree in code, not in a comment:
|
||
|
||
```zig
|
||
comptime {
|
||
if (maximum_domains != devices_broker.maximum_devices)
|
||
@compileError("iommu.confined is indexed by device id; an id past its end is " ++
|
||
"left unconfined while confineDevice still reports success");
|
||
}
|
||
```
|
||
|
||
## The one that must not wait
|
||
|
||
[`iommu.zig:107`](../system/kernel/iommu.zig:107) — `if (device_id >= confined.len) return
|
||
true;` — returns *success* without confining. It is unreachable today only because
|
||
device ids stop at 64. **Phases 3 and 4 both change the device count, and either makes
|
||
it live.** Fix it before them: out of range must refuse, never allow. (I2)
|
||
|
||
This is also the standing rule the audit argues for: at a bound, the safe direction is
|
||
refusal. A ceiling that fails open is not a limit, it is a switch that turns the
|
||
protection off.
|
||
|
||
## How this is verified
|
||
|
||
- The QEMU suite is the arbiter at every step; it is 114 cases and must stay green.
|
||
- Each fix lands with a test that **fails before it** — as the idempotency reorder did,
|
||
where exactly one assertion flipped.
|
||
- Adversarial cases for I1–I3 specifically: the audit's six real defects were all found
|
||
by asking "what would an attacker do", and the suite had never asked.
|
||
- The Ryzen is the acceptance test. It is the machine that found this, and the one that
|
||
proves it fixed.
|
||
|
||
## Sequencing
|
||
|
||
Phase 1 gates everything. Phase 2 gates phase 3 — the inventory cannot move until
|
||
authority is sound, or moving it is the hole. Phase 4 is independent and its user-space
|
||
items are the safest work in the track. Phase 5 is the record of what survived.
|
||
|
||
The IOMMU fail-open is fixed before phase 3 or 4 touches the device count.
|