Files
danos/docs/bounds-track-plan.md
T
Daniel Samson 33376f24ab test: the attacker the device suite never had
The audit's sharpest finding was structural, not a bug: a fully green suite
had hidden six real defects because it contains no attacker. Every device
case asserts that a driver handed its own hardware can drive it. None asked
what a process handed NOTHING can do.

device-authority-test is that process. It is spawned with no device and
asserts what it therefore cannot do: it cannot give away a device another
task holds, nor a free one, because the kernel's rule is that you may give
away what you hold and the device's state is irrelevant to a process holding
nothing. Asserted across every device the machine actually has, so it cannot
pass by accident of which one happened to be free at boot — six on QEMU,
none of them its.

A positive control runs first. device_enumerate works from this process, so
the refusals below it are decisions rather than a syscall path that is
simply broken here; without it, "everything failed" would read identically
to "the assertions are meaningless". A nonexistent device is refused as
NoSuchDevice rather than NotHeld, because a refusal that cannot name its own
rule is what cost a debugging session on the Ryzen.

What it deliberately does not assert, and says so in its header:
device_claim is still first-come-first-served at this point in the run. That
is the hole D6 closes, and the claim half of the invariant joins this
fixture then. Asserting it now would be writing a test that documents the
bug.

Verified to discriminate: removing the holder check flips "every transfer by
a non-holder is refused" while the positive control keeps passing.

Suite 117 -> 118.
2026-08-08 17:23:17 +01:00

386 lines
22 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# The bounds track: removing the numbers we invented
*Plan, 2026-08-08. Follows [fixed-bounds-audit.md](fixed-bounds-audit.md) (235 ceilings,
139 on quantities we do not choose) and the AMD Ryzen that found the first one.*
---
## Live state — the unattended run
*This table is the progress view. It is updated at the end of every step, before the
next one starts.*
| Step | What | State |
|---|---|---|
| L1 | Reclamation: a dead task's registrations die with its claims | **stopped — the step was wrong; see open question 4** |
| L2 | Bounds build check + allowlist; declare what we have already touched | **done** — `zig build bounds`, 273 allowlisted, 5 declared |
| L3 | xHCI: slot count from `HCSPARAMS1.MaxSlots`, not 8 | **done** — QEMU reports 64; the driver tracked 8 |
| L4 | USB: configuration descriptor sized by `wTotalLength`, not 512 | **done** — QEMU tops out at 211 bytes, so the case catches the class, not the original trigger |
| L5 | USB: interfaces from the descriptor, and the misattributed-endpoint bug | **done** — fix is by construction; no direct test, see open question 5 |
| L6 | xHCI: a failed `allocateDevice` stops leaking an enabled slot | **done** — path forced and verified; no regression test, see open question 5 |
**Run 1 complete.** L1 stopped (the step was wrong), L2–L6 landed. Suite 115 → 116.
Allowlist 278 → 269. Two steps ship without a permanent regression test, both because
QEMU's USB devices are too small to reach the paths — see open question 5, which is the
audit's own lesson recurring: the test rig is smaller than a real machine.
---
## Run 2 — device authority: delete the invented ceilings
*Design: [device-authority.md](os-development/device-authority.md), which is the **how**
for the delegation step [device-manager.md](device-driver-development/device-manager.md)
already settled. Read both before starting; the second is authoritative where they
differ.*
The goal, in the project owner's words: **remove the maximum values we set arbitrarily,
move the responsibility to the device manager, and keep in the kernel only the parts
that cannot safely run in user space.**
| Step | What | State |
|---|---|---|
| D1 | `device_transfer(device_id, task_id)` — the holder gives a device away | **done** — syscall 54; a move, not a copy |
| D2 | Adversarial case: a process handed nothing is refused, on a held device and a free one | **done** — `device-authority-test`; the claim half joins it at D6 |
| D3 | The manager claims the seeded devices at boot, before any driver is spawned | not started |
| D4 | `usb-xhci-bus` receives its controller in the `hello` reply instead of claiming argv[1] | not started |
| D5 | The other four claimants converted: `pci-bus`, `ps2-bus`, `virtio-gpu`, `acpi` | not started |
| D6 | `device_claim` refuses a device the caller was not handed; the hole is closed | not started |
| D7 | Zero-resource devices stop being kernel objects — inventory moves to the manager | not started |
| D8 | **`maximum_children_per_parent` deleted** — the authorisation it stood in for exists | not started |
| D9 | The device table becomes dynamic; **`maximum_devices` deleted**; per-holder quota declared | not started |
Ordering is load-bearing. D1–D2 build and prove the mechanism with nothing depending on
it. D3–D5 move each claimant across one at a time, so the suite stays green throughout
and a regression names the driver that caused it. D6 is the flag day. D7 must precede
D9, because zero-resource children are the case that sidesteps containment and so the
reason a shared cap was needed at all.
### Settled, so the run does not re-litigate them
- **The manager claims, it is not granted.** No binary names in the kernel; the rule is
"you may give away what you hold". The residual — it rests on the manager claiming
first — is stated in the design and is closed later by the spawn capability
[drivers.md](device-driver-development/drivers.md) already names as missing.
- **The framebuffer is not a device.** It is where pixels go, handed over by the loader,
and the compositor uses it as the boot floor until a real display driver announces
itself. Nothing in this run touches the display service or its GOP path.
- **`maximum_endpoints_per_interface` and the wire structs stay.** Widening them is a
protocol change, out of scope.
- **A per-holder quota is not a retreat.** Dynamic storage with no bound moves the
ceiling to the kernel heap, which is shared and fatal rather than partial. A bound
charged to the task that caused it is isolation, and it is declared through
[bounds.md](os-development/bounds.md) like anything else.
### Working rules
As Run 1, unchanged: work in `/Users/danielsamson/Gitea/daniel/danos` on
`claude/bounds-track`; every step lands with a test that fails before the fix, verified
by restoring the old behaviour; full suite green before each commit; never two suites at
once (`pgrep -f qemu_test.py`); 60 GiB free; `git commit -F` with no `Co-Authored-By`;
update this table before starting the next step. **If a step needs a decision that is
not written down, stop it, add the question below, and move on** — Run 1's first step
was wrong and stopping was the right call.
**Suite:** 115/115 at the start of the run.
**Branch:** `claude/bounds-track`.
### What this run deliberately does not touch
Phases 2 and 3 below — the authorisation gate and moving the inventory to the device
manager — are **out of scope for unattended work**. They decide whether the OS is
secure, and they are currently a direction rather than a specification: what a device
capability *is*, which syscalls change, what replaces `device_claim` for its seven
callers, how a driver spawned bare behaves. Those want a design session, the way
`/protocol` had one.
Also out of scope: anything touching `maximum_device_resources` (a wire struct, so a
trust-boundary change, not a resize), and the non-device bounds the audit found in FAT,
the VFS, logger, init, display and boot.
### Open questions this run must not answer on its own
Recorded here rather than guessed. If a step runs into one, it stops and writes the
question down instead of inventing an answer.
1. **`device_enumerate` probably narrows rather than retires.** The device manager
calls it to find `pci_host_bridge` nodes — it cannot ask itself. The likely shape is
that the kernel keeps the *firmware-discovered roots* (which by principle 5 it holds
for real reasons, since they come from ACPI rather than a driver's say-so) and
everything a driver registered lives in the manager. Not decided.
2. **A device-manager restart has no re-enumerate handshake.** If only the manager
dies, the buses are alive and never re-send `child_added`, so a restarted manager
comes back blind. The manager is restartable by design; nothing implements this.
3. **Which adversarial tests I1–I3 need.** The audit's six real defects were all found
by asking what an attacker would do, and the suite had never asked. "Add adversarial
cases" is not executable until the attacks are named.
4. **Reclamation is not a death-sweep problem, and L1 as written would have broken the
restart path.** Found on the first attempt at it. The audit is right that `count`
never decreases, but *death is the wrong trigger*:
- The broker keeps entries deliberately: "The devices stay in the table — they
describe hardware, which did not go away — only their ownership clears." A driver
dying does not unplug anything.
- Device ids must stay **stable across a bus restart**, because
`device-manager.driverForDevice` dedupes by `device_id` so that "a re-report after
a bus restart must not spawn a second instance". Stability comes from the
idempotency scan returning the existing id — removing entries on death would give
a restarted bus fresh ids and spawn duplicate driver instances.
- Everything else a task holds *is* already reclaimed on every path out:
`irq.releaseOwner`, `iommu.releaseAllOwnedBy`, `dmaRegistryReleaseOwner`, then the
broker's claims (`process.releaseTaskResourcesLocked`).
So the real leak has two sources, and neither is death: a device that genuinely
**goes away** (hot-unplug) has no retirement path, and a bus that enumerates
*differently* on restart leaves its stale entries behind forever. Both are the device
manager's inventory problem — phase 3 — and both need the id-stability question
answered first (tombstone-and-reuse aliases stale ids held by another process;
generation-tagged ids change the id encoding, which is ABI). Not an unattended
decision.
5. **Driver descriptor parsing cannot be host-tested, so L5's correctness fix ships
without a direct test.** The endpoint-misattribution bug lives in
`parseConfiguration`, a pure function over a byte blob — exactly the shape a host
unit test wants, and `usb-storage/scsi.zig` and `usb-hid/hid-report.zig` already do
this. But `usb-xhci-library.zig` imports `memory`, `mmio` and `time`, so it cannot
be a standalone host-test root, and QEMU offers no device that would exercise the
path anyway: the largest available is `usb-audio,multi=on` at 2 interfaces and 211
bytes, against a cap of 4.
Three ways out, and picking one is a judgement about house style rather than a
mechanical step: extract the parser to its own file and wire `usb-abi` into a test
module (build-support currently resolves module names only for `userBinary`);
extract it and import `usb-abi` by relative path (against the import-by-name
convention); or accept QEMU-only coverage and say so.
Mitigating, and the reason this is recorded rather than blocking: after the fix the
bug is unreachable **by construction**, not by the added `else`. Interfaces are now
allocated to exactly the count the descriptor declares, so `interface_count` can
never reach `interfaces.len` mid-parse. The `else` is belt-and-braces for the
255-interface clamp. The alternate-setting path that shares it *is* exercised —
`usb-audio` has alternate settings, and the `usb-large-descriptor` case walks them.
### Working rules for the run
- Work in `/Users/danielsamson/Gitea/daniel/danos` (not a worktree), on
`claude/bounds-track`.
- **Every step lands with a test that fails before the fix**, verified by temporarily
restoring the old behaviour and watching exactly the intended assertion flip. A test
that passes both ways is not a test.
- Full QEMU suite green before each commit. Never run two suites at once — check
`pgrep -f qemu_test.py` first; a second concurrent run produces false triple faults
because both share `zig-out`.
- Check at least 60 GiB free before starting a suite.
- Commit with `git commit -F <file>`, never `-m` (a backtick in a message is executed
by the shell and silently eats a word). No `Co-Authored-By` trailers.
- Update the Live state table **before** starting the next step.
- If a step needs a decision that is not written down here, stop, add it to the open
questions above, and move to the next step.
---
## The principles this is derived from
1. **danOS is a microkernel.** Minimise what the kernel is responsible for; move
responsibility to user space so it can be restarted, or fixed live during
development, without taking the system down.
2. **Implement the specifications correctly**, with the limits those specifications
define — not limits we decide.
3. **Move as much responsibility as possible to user space** (the device manager).
4. **What remains in the kernel is minimal.**
5. **What remains in the kernel is there for security or for a hardware limitation.**
Nothing else earns a place.
Principle 5 is the test every bound is put to. For each one: *is this here because of
security, or because of a hardware limitation?* If neither, the storage does not belong
in the kernel and the bound is not a number to be resized — it is a thing to be moved or
deleted.
Applying it to the case that started this:
- `maximum_devices = 64` bounds an inventory of hardware. An inventory is neither a
security control nor a hardware limitation. **The table is in the wrong place**; the
number is a symptom.
- `maximum_children_per_parent = 16` exists because `device_claim` is unauthenticated —
any process can claim any unclaimed device ([devices-broker.zig:164](../system/kernel/devices-broker.zig:164)
checks only that the device exists and is free). The cap is a crude proxy for an
authorisation the kernel does not perform. **Fix the authorisation and the cap has
nothing to defend.**
- `maximum_domains = 64` bounds IOMMU translation domains. Security — stays in the
kernel. But VT-d and AMD-Vi both *report* how many domains they support in a
capability register. Principle 2: read it. We chose 64 without asking.
## The security invariants
Every phase must leave all five standing. This is the "without punching a hole" half of
the brief, and each phase below states how it is checked.
- **I1 Containment.** A process may map only physical memory inside a resource it was
granted. A bus may subdivide only what it already holds.
- **I2 Confinement.** A DMA-capable device is under IOMMU translation before its driver
can program it, or it is not driven at all.
- **I3 No self-granted authority.** A process holds what it was handed. It cannot name
its way into holding more.
- **I4 Death releases everything.** Every resource a task held is reclaimed when it
dies, on every path out.
- **I5 Refusal is attributable.** Every refusal names the rule that refused it.
## Phase 0 — Done
- **Errno attribution.** One errno space in `system/abi.zig`; `device_register`'s six
refusals and `device_claim`'s three are distinct codes; call sites name the reason;
`pci-bus` reconciles found against registered. (I5)
- **Idempotency ordering.** A re-registration consumes no slot, so a full parent
re-admits an identical child. A restarted bus is no longer billed for what it
rediscovers.
Suite 114/114.
## Phase 1 — Reclamation
**Nothing may become dynamic before this.** Today `count` only ever increases and
`releaseAllOwnedBy` clears a dead driver's *claims* but not its *registrations*. With a
fixed table that is a slow march to the cap; with dynamic storage it is an unbounded
leak, and every supervisor restart makes it worse.
- Extend the existing death sweep so a task's registrations go with its claims.
- A registration whose owner is gone is removed; its children are re-parented or removed
with it (they cannot outlive the authority that published them).
- Test: register under a claimed parent, kill the owner, assert the entries are gone and
the ids are not reused while any handle to them lives.
Invariant: **I4**.
## Phase 2 — Close the authorisation hole
The device manager already decides which driver gets which device — it matches against
`devices.csv` and spawns the driver with the device id as `argv[1]`. Nothing binds that
decision to the kernel's `claim`. A driver passes an integer; the kernel checks only
that the device is free.
Per principles 3 and 5: **the decision stays in user space; the kernel enforces only
possession.** The manager hands the driver the device it matched; the kernel's job is
that a driver holds what it was handed and nothing else.
- The manager passes a device to the driver it spawned, over the existing cap-passing
path. Possession is the authority.
- `device_claim` stops being a way to *acquire* a device by naming it.
- Exclusivity stops being a broker refusing a second claimant and becomes the ordinary
property of a thing only one process was given.
**`maximum_children_per_parent` is deleted here**, because after this a bus driver's
children are the devices it actually enumerated under a bus it was actually given, and
the rogue-driver-fills-the-table threat the cap was written for no longer exists.
Invariants: **I3** (the point of the phase), **I1** (containment is unchanged and still
checked on every subdivision), **I5**.
Acceptance: a driver that names a device it was not given is refused, with its own
errno. The QEMU suite gains an adversarial case for it — the audit's lesson was that
"the suite contains no attacker".
## Phase 3 — The inventory moves to user space
The kernel reads only three things out of a device descriptor: **physical ranges** (to
check a mapping falls inside one), **interrupt numbers**, and **one PCI BDF** (to key an
IOMMU domain). Vendor and device ids, class triples, subsystem ids, human-readable
names, bus numbers and parent links are stored solely so `device_enumerate` can hand
them back. That is the kernel acting as a distribution mechanism for data it does not
use — principle 5 excludes it.
- **Zero-resource devices leave the kernel entirely.** A USB device addressed through
its controller conveys no mapping authority; there is nothing for the kernel to
enforce. It is pure inventory and belongs to the device manager. (This is also the
case that sidesteps containment, which is why the cap existed.)
- Identity and topology move to the manager, which already receives them as
`child_added` reports and already holds the authoritative picture.
- `device_enumerate` retires; callers ask the manager, whose protocol already reserves
an `enumerate` verb. Public-ABI change — `docs/os-development/vdso.md` documents it.
- What the kernel keeps: for each device that carries resources, the ranges, the GSIs,
the BDF, and the owner.
After this, the kernel's table holds only resource-bearing devices, and the remaining
count is bounded by what the machine physically has rather than by us.
Invariants: **I1**, **I2** unchanged — both operate on resources, which do not move.
**I4** must be re-checked: the manager's table now needs its own reclamation, and it is
restartable, so it must be able to rebuild from the buses.
## Phase 4 — Ask the hardware and the specification
Principle 2, applied to every remaining bound. Each of these is a number the machine or
the standard already states, which we replaced with a guess. Independent of each other;
can proceed in any order.
| Today | Ask instead |
|---|---|
| `maximum_domains = 64` (IOMMU) | the VT-d / AMD-Vi capability register reports the domains supported |
| `max_devices = 8` (xHCI slots) | `HCSPARAMS1.MaxSlots` — the controller says (1–255) |
| `max_interfaces = 4`, `max_endpoints` | the configuration descriptor says |
| `blob: [512]u8` (USB config) | the device's `wTotalLength` |
| `below: [64]Range` (memory map) | UEFI reports the descriptor count |
| AML blobs capped at 6 | the XSDT's length field gives the entry count |
| `maximum_cpus = 128` | the MADT entry count |
| `maximum_gsi = 24` | the I/O APIC's redirection-entry count; and more than one I/O APIC exists |
| MSI-X vectors | the capability's table-size field (up to 2048) |
Several of these are in user space already (xHCI, USB descriptors) and are ordinary
allocations — principle 1 means those are also the safest to do first, since a mistake
restarts a driver rather than the machine.
Two in this table are **also** correctness fixes the audit found, and should carry their
regression tests: the xHCI `max_interfaces` path misattributes a fifth interface's
endpoints to interface 3, and `below: [64]Range` silently turns occupied RAM into a PCI
aperture — which is an **I1 violation reachable on real hardware**, not merely a lost
device. That one is the highest-priority item in this phase.
## Phase 5 — What legitimately remains
After phases 1–4 the survivors should be only:
- **Pre-allocator storage**: the PMM's own frame bitmap, the memory map the loader hands
over, the bootstrap page tables. You cannot allocate the allocator. (Hardware/boot
limitation — principle 5 admits these.)
- **Interrupt-context storage**: the IST stack and anything an exception path touches
without allocating.
- **Wire structures** whose layout the other side of a trust boundary parses.
- **Facts that are not ceilings**: a page is 4096 bytes; an ACPI name segment is 4.
Each is declared per [bounds.md](os-development/bounds.md) — what it counts, who decides
its size, what it protects, what happens at the limit, how you find out. And two numbers
that must agree agree in code, not in a comment:
```zig
comptime {
if (maximum_domains != devices_broker.maximum_devices)
@compileError("iommu.confined is indexed by device id; an id past its end is " ++
"left unconfined while confineDevice still reports success");
}
```
## The one that must not wait
[`iommu.zig:107`](../system/kernel/iommu.zig:107) — `if (device_id >= confined.len) return
true;` — returns *success* without confining. It is unreachable today only because
device ids stop at 64. **Phases 3 and 4 both change the device count, and either makes
it live.** Fix it before them: out of range must refuse, never allow. (I2)
This is also the standing rule the audit argues for: at a bound, the safe direction is
refusal. A ceiling that fails open is not a limit, it is a switch that turns the
protection off.
## How this is verified
- The QEMU suite is the arbiter at every step; it is 114 cases and must stay green.
- Each fix lands with a test that **fails before it** — as the idempotency reorder did,
where exactly one assertion flipped.
- Adversarial cases for I1–I3 specifically: the audit's six real defects were all found
by asking "what would an attacker do", and the suite had never asked.
- The Ryzen is the acceptance test. It is the machine that found this, and the one that
proves it fixed.
## Sequencing
Phase 1 gates everything. Phase 2 gates phase 3 — the inventory cannot move until
authority is sound, or moving it is the hole. Phase 4 is independent and its user-space
items are the safest work in the track. Phase 5 is the record of what survived.
The IOMMU fail-open is fixed before phase 3 or 4 touches the device count.