Files
danos/docs/bounds-track-plan.md
T
Daniel Samson 81dd1e9318 pci: the host bridge arrives by delegation too
pci-bus joins usb-xhci-bus in receiving its device from the manager rather
than claiming the id it found in argv[1]. Its hello moves ahead of the ECAM
mapping, since that is where the bridge now arrives, and its hello was
already mandatory so nothing about its failure behaviour changes.

isDelegated compared whole strings, which silently missed this driver: the
manager records the boot-snapshot match as the bare "pci-bus" and a
devices.csv match as the full "/system/drivers/pci-bus". pci-bus was then
neither claiming nor delegated and died on "ECAM mmio_map failed". It now
matches on the last path component. Reintroducing the whole-string compare
breaks usb-xhci-bus instead of pci-bus — the two spellings swap which driver
loses — so usb-hid is the case that catches it, not pci-scan.

pci-scan asserts the delegation on the initial bring-up AND after the
restart drill, with the device id backreferenced so both must name the same
device. That is what proves the manager re-takes a device when its driver
dies and hands it to the replacement, which is the property the whole
supervision design rests on.

The remaining three claimants are NOT converted, and the plan records why
rather than working around it. ps2-bus and the acpi service never hello at
all, which device-manager.md states deliberately ("legacy drivers ... not
yet required to hello"), so delegating to them means either promoting them
out of legacy or giving the grant a delivery point that is not hello.
virtio-gpu hellos best-effort by design — "standalone bring-up has no
manager" — and delegation would make it mandatory. Both are decisions, not
mechanical steps.

Consequence: D6 is blocked, because device_claim cannot be closed off while
three claimants still depend on it. D7-D9 are unaffected — they concern what
the kernel stores and how its table is sized.

Suite 118/118.
2026-08-08 18:25:16 +01:00

25 KiB
Raw Blame History

The bounds track: removing the numbers we invented

Plan, 2026-08-08. Follows fixed-bounds-audit.md (235 ceilings, 139 on quantities we do not choose) and the AMD Ryzen that found the first one.


Live state — the unattended run

This table is the progress view. It is updated at the end of every step, before the next one starts.

Step What State
L1 Reclamation: a dead task's registrations die with its claims stopped — the step was wrong; see open question 4
L2 Bounds build check + allowlist; declare what we have already touched done — zig build bounds, 273 allowlisted, 5 declared
L3 xHCI: slot count from HCSPARAMS1.MaxSlots, not 8 done — QEMU reports 64; the driver tracked 8
L4 USB: configuration descriptor sized by wTotalLength, not 512 done — QEMU tops out at 211 bytes, so the case catches the class, not the original trigger
L5 USB: interfaces from the descriptor, and the misattributed-endpoint bug done — fix is by construction; no direct test, see open question 5
L6 xHCI: a failed allocateDevice stops leaking an enabled slot done — path forced and verified; no regression test, see open question 5

Run 1 complete. L1 stopped (the step was wrong), L2–L6 landed. Suite 115 → 116. Allowlist 278 → 269. Two steps ship without a permanent regression test, both because QEMU's USB devices are too small to reach the paths — see open question 5, which is the audit's own lesson recurring: the test rig is smaller than a real machine.


Run 2 — device authority: delete the invented ceilings

Design: device-authority.md, which is the how for the delegation step device-manager.md already settled. Read both before starting; the second is authoritative where they differ.

The goal, in the project owner's words: remove the maximum values we set arbitrarily, move the responsibility to the device manager, and keep in the kernel only the parts that cannot safely run in user space.

Step What State
D1 device_transfer(device_id, task_id) — the holder gives a device away done — syscall 54; a move, not a copy
D2 Adversarial case: a process handed nothing is refused, on a held device and a free one done — device-authority-test; the claim half joins it at D6
D3 The manager claims the seeded devices at boot, before any driver is spawned merged into D4 — see below
D4 The manager claims + delegates on hello; usb-xhci-bus is the first driver converted done — caught an IOMMU regression I introduced; see below
D5 The other four claimants converted: pci-bus, ps2-bus, virtio-gpu, acpi partial — pci-bus done; the other three need decisions, see below
D6 device_claim refuses a device the caller was not handed; the hole is closed not started
D7 Zero-resource devices stop being kernel objects — inventory moves to the manager not started
D8 maximum_children_per_parent deleted — the authorisation it stood in for exists not started
D9 The device table becomes dynamic; maximum_devices deleted; per-holder quota declared not started

Ordering is load-bearing. D1–D2 build and prove the mechanism with nothing depending on it. D4–D5 move each claimant across one at a time, so the suite stays green throughout and a regression names the driver that caused it. D6 is the flag day. D7 must precede D9, because zero-resource children are the case that sidesteps containment and so the reason a shared cap was needed at all.

D3 merged into D4 (found while implementing, recorded rather than worked around). The two cannot be separated: the moment the manager claims a device, any driver still calling device_claim on it is refused AlreadyClaimed, so D3 on its own turns the suite red — and D3 applied to nothing changes no behaviour and cannot be tested. They land together, with the manager claiming only for drivers in an explicit delegated set so every unconverted driver keeps claiming exactly as before. usb-xhci-bus is the first member, as it was the first driver to conform to hello (device-manager.md, M18.1). D5 moves the rest in one at a time; the set and the device_claim path both disappear at D6.

Observations from the run

  • D4 introduced an IOMMU regression, caught by converting one driver at a time. confineDevice runs inside systemDeviceClaim, so a device arriving by transfer was never confined for its new owner: the driver's DMA rings went unbound (three IOMMU+USB cases failed), and two worse consequences were latent — a manager death would have torn down a domain a live driver was using, and a driver death would have leaked one. iommu.reassign moves the confinement with the device, keeping the domain and attachment intact so it never translates through nothing. Converting all five drivers at once would have produced the same three failures with five suspects.
  • usb-hub failed once in a full run, then passed six times (four isolated, two full). Suspected instance of the known intermittent AP ring-3 fault rather than anything in D4 — recorded rather than dismissed, because D4 moved the hello earlier and so did shift boot timing. Watch it across the remaining steps.

Open questions raised by D5 — three of the four claimants cannot be converted yet

Delegation is delivered in onHello. That works for a driver that says hello, and two of the four do not.

  1. ps2-bus and the acpi service never hello at all, and that is deliberate: device-manager.md records "Legacy drivers (e.g. ps2-bus) are supervised and restarted but not yet required to hello", and the Driver.speaks_protocol flag exists to say so. Delegating to them means either making them speak the protocol — promoting them out of "legacy", which is a change to their documented status — or giving the grant a second delivery point that is not hello. Neither is written down. A second delivery point would also need its own answer to "when", since the whole value of hello here is that it is the moment the driver is known to be alive and is a synchronous point to hand something over.

  2. virtio-gpu hellos, but best-effort by design. Its call is _ = device_manager.hello(.device, device_id); with the comment "Best-effort: standalone bring-up has no manager." Delegation would make the hello mandatory and move it to the front, so the driver could no longer come up without a manager. No test exercises standalone today (the virtio-gpu case boots the manager stack, and devices.csv matches it), so this is a documented intent rather than a live path — but discarding a documented intent is a decision, not a mechanical step.

Until these are answered, device_claim cannot be closed off at D6 for those three, so D6 is blocked on question 6 and 7. D7–D9 are not: they concern what the kernel stores and how the table is sized, and are independent of which drivers have converted.

Settled, so the run does not re-litigate them

  • The manager claims, it is not granted. No binary names in the kernel; the rule is "you may give away what you hold". The residual — it rests on the manager claiming first — is stated in the design and is closed later by the spawn capability drivers.md already names as missing.
  • The framebuffer is not a device. It is where pixels go, handed over by the loader, and the compositor uses it as the boot floor until a real display driver announces itself. Nothing in this run touches the display service or its GOP path.
  • maximum_endpoints_per_interface and the wire structs stay. Widening them is a protocol change, out of scope.
  • A per-holder quota is not a retreat. Dynamic storage with no bound moves the ceiling to the kernel heap, which is shared and fatal rather than partial. A bound charged to the task that caused it is isolation, and it is declared through bounds.md like anything else.

Working rules

As Run 1, unchanged: work in /Users/danielsamson/Gitea/daniel/danos on claude/bounds-track; every step lands with a test that fails before the fix, verified by restoring the old behaviour; full suite green before each commit; never two suites at once (pgrep -f qemu_test.py); 60 GiB free; git commit -F with no Co-Authored-By; update this table before starting the next step. If a step needs a decision that is not written down, stop it, add the question below, and move on — Run 1's first step was wrong and stopping was the right call.

Suite: 115/115 at the start of the run. Branch: claude/bounds-track.

What this run deliberately does not touch

Phases 2 and 3 below — the authorisation gate and moving the inventory to the device manager — are out of scope for unattended work. They decide whether the OS is secure, and they are currently a direction rather than a specification: what a device capability is, which syscalls change, what replaces device_claim for its seven callers, how a driver spawned bare behaves. Those want a design session, the way /protocol had one.

Also out of scope: anything touching maximum_device_resources (a wire struct, so a trust-boundary change, not a resize), and the non-device bounds the audit found in FAT, the VFS, logger, init, display and boot.

Open questions this run must not answer on its own

Recorded here rather than guessed. If a step runs into one, it stops and writes the question down instead of inventing an answer.

  1. device_enumerate probably narrows rather than retires. The device manager calls it to find pci_host_bridge nodes — it cannot ask itself. The likely shape is that the kernel keeps the firmware-discovered roots (which by principle 5 it holds for real reasons, since they come from ACPI rather than a driver's say-so) and everything a driver registered lives in the manager. Not decided.

  2. A device-manager restart has no re-enumerate handshake. If only the manager dies, the buses are alive and never re-send child_added, so a restarted manager comes back blind. The manager is restartable by design; nothing implements this.

  3. Which adversarial tests I1–I3 need. The audit's six real defects were all found by asking what an attacker would do, and the suite had never asked. "Add adversarial cases" is not executable until the attacks are named.

  4. Reclamation is not a death-sweep problem, and L1 as written would have broken the restart path. Found on the first attempt at it. The audit is right that count never decreases, but death is the wrong trigger:

    • The broker keeps entries deliberately: "The devices stay in the table — they describe hardware, which did not go away — only their ownership clears." A driver dying does not unplug anything.
    • Device ids must stay stable across a bus restart, because device-manager.driverForDevice dedupes by device_id so that "a re-report after a bus restart must not spawn a second instance". Stability comes from the idempotency scan returning the existing id — removing entries on death would give a restarted bus fresh ids and spawn duplicate driver instances.
    • Everything else a task holds is already reclaimed on every path out: irq.releaseOwner, iommu.releaseAllOwnedBy, dmaRegistryReleaseOwner, then the broker's claims (process.releaseTaskResourcesLocked).

    So the real leak has two sources, and neither is death: a device that genuinely goes away (hot-unplug) has no retirement path, and a bus that enumerates differently on restart leaves its stale entries behind forever. Both are the device manager's inventory problem — phase 3 — and both need the id-stability question answered first (tombstone-and-reuse aliases stale ids held by another process; generation-tagged ids change the id encoding, which is ABI). Not an unattended decision.

  5. Driver descriptor parsing cannot be host-tested, so L5's correctness fix ships without a direct test. The endpoint-misattribution bug lives in parseConfiguration, a pure function over a byte blob — exactly the shape a host unit test wants, and usb-storage/scsi.zig and usb-hid/hid-report.zig already do this. But usb-xhci-library.zig imports memory, mmio and time, so it cannot be a standalone host-test root, and QEMU offers no device that would exercise the path anyway: the largest available is usb-audio,multi=on at 2 interfaces and 211 bytes, against a cap of 4.

    Three ways out, and picking one is a judgement about house style rather than a mechanical step: extract the parser to its own file and wire usb-abi into a test module (build-support currently resolves module names only for userBinary); extract it and import usb-abi by relative path (against the import-by-name convention); or accept QEMU-only coverage and say so.

    Mitigating, and the reason this is recorded rather than blocking: after the fix the bug is unreachable by construction, not by the added else. Interfaces are now allocated to exactly the count the descriptor declares, so interface_count can never reach interfaces.len mid-parse. The else is belt-and-braces for the 255-interface clamp. The alternate-setting path that shares it is exercised — usb-audio has alternate settings, and the usb-large-descriptor case walks them.

Working rules for the run

  • Work in /Users/danielsamson/Gitea/daniel/danos (not a worktree), on claude/bounds-track.
  • Every step lands with a test that fails before the fix, verified by temporarily restoring the old behaviour and watching exactly the intended assertion flip. A test that passes both ways is not a test.
  • Full QEMU suite green before each commit. Never run two suites at once — check pgrep -f qemu_test.py first; a second concurrent run produces false triple faults because both share zig-out.
  • Check at least 60 GiB free before starting a suite.
  • Commit with git commit -F <file>, never -m (a backtick in a message is executed by the shell and silently eats a word). No Co-Authored-By trailers.
  • Update the Live state table before starting the next step.
  • If a step needs a decision that is not written down here, stop, add it to the open questions above, and move to the next step.

The principles this is derived from

  1. danOS is a microkernel. Minimise what the kernel is responsible for; move responsibility to user space so it can be restarted, or fixed live during development, without taking the system down.
  2. Implement the specifications correctly, with the limits those specifications define — not limits we decide.
  3. Move as much responsibility as possible to user space (the device manager).
  4. What remains in the kernel is minimal.
  5. What remains in the kernel is there for security or for a hardware limitation. Nothing else earns a place.

Principle 5 is the test every bound is put to. For each one: is this here because of security, or because of a hardware limitation? If neither, the storage does not belong in the kernel and the bound is not a number to be resized — it is a thing to be moved or deleted.

Applying it to the case that started this:

  • maximum_devices = 64 bounds an inventory of hardware. An inventory is neither a security control nor a hardware limitation. The table is in the wrong place; the number is a symptom.
  • maximum_children_per_parent = 16 exists because device_claim is unauthenticated — any process can claim any unclaimed device (devices-broker.zig:164 checks only that the device exists and is free). The cap is a crude proxy for an authorisation the kernel does not perform. Fix the authorisation and the cap has nothing to defend.
  • maximum_domains = 64 bounds IOMMU translation domains. Security — stays in the kernel. But VT-d and AMD-Vi both report how many domains they support in a capability register. Principle 2: read it. We chose 64 without asking.

The security invariants

Every phase must leave all five standing. This is the "without punching a hole" half of the brief, and each phase below states how it is checked.

  • I1 Containment. A process may map only physical memory inside a resource it was granted. A bus may subdivide only what it already holds.
  • I2 Confinement. A DMA-capable device is under IOMMU translation before its driver can program it, or it is not driven at all.
  • I3 No self-granted authority. A process holds what it was handed. It cannot name its way into holding more.
  • I4 Death releases everything. Every resource a task held is reclaimed when it dies, on every path out.
  • I5 Refusal is attributable. Every refusal names the rule that refused it.

Phase 0 — Done

  • Errno attribution. One errno space in system/abi.zig; device_register's six refusals and device_claim's three are distinct codes; call sites name the reason; pci-bus reconciles found against registered. (I5)
  • Idempotency ordering. A re-registration consumes no slot, so a full parent re-admits an identical child. A restarted bus is no longer billed for what it rediscovers.

Suite 114/114.

Phase 1 — Reclamation

Nothing may become dynamic before this. Today count only ever increases and releaseAllOwnedBy clears a dead driver's claims but not its registrations. With a fixed table that is a slow march to the cap; with dynamic storage it is an unbounded leak, and every supervisor restart makes it worse.

  • Extend the existing death sweep so a task's registrations go with its claims.
  • A registration whose owner is gone is removed; its children are re-parented or removed with it (they cannot outlive the authority that published them).
  • Test: register under a claimed parent, kill the owner, assert the entries are gone and the ids are not reused while any handle to them lives.

Invariant: I4.

Phase 2 — Close the authorisation hole

The device manager already decides which driver gets which device — it matches against devices.csv and spawns the driver with the device id as argv[1]. Nothing binds that decision to the kernel's claim. A driver passes an integer; the kernel checks only that the device is free.

Per principles 3 and 5: the decision stays in user space; the kernel enforces only possession. The manager hands the driver the device it matched; the kernel's job is that a driver holds what it was handed and nothing else.

  • The manager passes a device to the driver it spawned, over the existing cap-passing path. Possession is the authority.
  • device_claim stops being a way to acquire a device by naming it.
  • Exclusivity stops being a broker refusing a second claimant and becomes the ordinary property of a thing only one process was given.

maximum_children_per_parent is deleted here, because after this a bus driver's children are the devices it actually enumerated under a bus it was actually given, and the rogue-driver-fills-the-table threat the cap was written for no longer exists.

Invariants: I3 (the point of the phase), I1 (containment is unchanged and still checked on every subdivision), I5.

Acceptance: a driver that names a device it was not given is refused, with its own errno. The QEMU suite gains an adversarial case for it — the audit's lesson was that "the suite contains no attacker".

Phase 3 — The inventory moves to user space

The kernel reads only three things out of a device descriptor: physical ranges (to check a mapping falls inside one), interrupt numbers, and one PCI BDF (to key an IOMMU domain). Vendor and device ids, class triples, subsystem ids, human-readable names, bus numbers and parent links are stored solely so device_enumerate can hand them back. That is the kernel acting as a distribution mechanism for data it does not use — principle 5 excludes it.

  • Zero-resource devices leave the kernel entirely. A USB device addressed through its controller conveys no mapping authority; there is nothing for the kernel to enforce. It is pure inventory and belongs to the device manager. (This is also the case that sidesteps containment, which is why the cap existed.)
  • Identity and topology move to the manager, which already receives them as child_added reports and already holds the authoritative picture.
  • device_enumerate retires; callers ask the manager, whose protocol already reserves an enumerate verb. Public-ABI change — docs/os-development/vdso.md documents it.
  • What the kernel keeps: for each device that carries resources, the ranges, the GSIs, the BDF, and the owner.

After this, the kernel's table holds only resource-bearing devices, and the remaining count is bounded by what the machine physically has rather than by us.

Invariants: I1, I2 unchanged — both operate on resources, which do not move. I4 must be re-checked: the manager's table now needs its own reclamation, and it is restartable, so it must be able to rebuild from the buses.

Phase 4 — Ask the hardware and the specification

Principle 2, applied to every remaining bound. Each of these is a number the machine or the standard already states, which we replaced with a guess. Independent of each other; can proceed in any order.

Today Ask instead
maximum_domains = 64 (IOMMU) the VT-d / AMD-Vi capability register reports the domains supported
max_devices = 8 (xHCI slots) HCSPARAMS1.MaxSlots — the controller says (1–255)
max_interfaces = 4, max_endpoints the configuration descriptor says
blob: [512]u8 (USB config) the device's wTotalLength
below: [64]Range (memory map) UEFI reports the descriptor count
AML blobs capped at 6 the XSDT's length field gives the entry count
maximum_cpus = 128 the MADT entry count
maximum_gsi = 24 the I/O APIC's redirection-entry count; and more than one I/O APIC exists
MSI-X vectors the capability's table-size field (up to 2048)

Several of these are in user space already (xHCI, USB descriptors) and are ordinary allocations — principle 1 means those are also the safest to do first, since a mistake restarts a driver rather than the machine.

Two in this table are also correctness fixes the audit found, and should carry their regression tests: the xHCI max_interfaces path misattributes a fifth interface's endpoints to interface 3, and below: [64]Range silently turns occupied RAM into a PCI aperture — which is an I1 violation reachable on real hardware, not merely a lost device. That one is the highest-priority item in this phase.

Phase 5 — What legitimately remains

After phases 1–4 the survivors should be only:

  • Pre-allocator storage: the PMM's own frame bitmap, the memory map the loader hands over, the bootstrap page tables. You cannot allocate the allocator. (Hardware/boot limitation — principle 5 admits these.)
  • Interrupt-context storage: the IST stack and anything an exception path touches without allocating.
  • Wire structures whose layout the other side of a trust boundary parses.
  • Facts that are not ceilings: a page is 4096 bytes; an ACPI name segment is 4.

Each is declared per bounds.md — what it counts, who decides its size, what it protects, what happens at the limit, how you find out. And two numbers that must agree agree in code, not in a comment:

comptime {
    if (maximum_domains != devices_broker.maximum_devices)
        @compileError("iommu.confined is indexed by device id; an id past its end is " ++
            "left unconfined while confineDevice still reports success");
}

The one that must not wait

iommu.zig:107 — if (device_id >= confined.len) return true; — returns success without confining. It is unreachable today only because device ids stop at 64. Phases 3 and 4 both change the device count, and either makes it live. Fix it before them: out of range must refuse, never allow. (I2)

This is also the standing rule the audit argues for: at a bound, the safe direction is refusal. A ceiling that fails open is not a limit, it is a switch that turns the protection off.

How this is verified

  • The QEMU suite is the arbiter at every step; it is 114 cases and must stay green.
  • Each fix lands with a test that fails before it — as the idempotency reorder did, where exactly one assertion flipped.
  • Adversarial cases for I1–I3 specifically: the audit's six real defects were all found by asking "what would an attacker do", and the suite had never asked.
  • The Ryzen is the acceptance test. It is the machine that found this, and the one that proves it fixed.

Sequencing

Phase 1 gates everything. Phase 2 gates phase 3 — the inventory cannot move until authority is sound, or moving it is the hole. Phase 4 is independent and its user-space items are the safest work in the track. Phase 5 is the record of what survived.

The IOMMU fail-open is fixed before phase 3 or 4 touches the device count.