An AMD Ryzen booted to a working compositor with no USB and no storage, and the log said only "register refused". A tree-wide audit of every compile-time ceiling followed: 235 of them, 139 on quantities the machine or a file decides rather than us, 5 documented anywhere, 171 silent when reached. docs/fixed-bounds-audit.md has the inventory. Errno attribution. The errno space was split between the kernel and the envelope, free to drift; it is now one list in system/abi.zig, restated on both sides, with a comptime check in library/device/driver where the two halves are visible. device_register's six refusals and device_claim's three are distinct codes, so a bus driver can say which rule stopped it, and BadParent splits into NoSuchParent and NotYourParent. pci-bus reconciles found against registered instead of counting refused functions as found. Idempotency ordering. The child cap was checked before the identity match, so a restarted bus was refused its own devices — the supervision restart the system leans on ratcheted toward a degraded machine. A re-registration consumes no slot and is now admitted first. IOMMU fail-closed. confineDevice returned success for a device id past the confinement table, leaving the device outside every domain while the caller believed it confined — unreachable only while ids stop at 64, which both the inventory move and a hardware-reported domain count would change. It refuses now, and the coupling to the broker's device cap is a comptime assert rather than a sentence in a comment. PCI apertures. The bridge's MMIO apertures are derived from the holes in the firmware memory map, and the derivation copied sub-4 GiB entries into a fixed [64] array and skipped the rest. A skipped region is not merely lost: the gap finder concludes it is free, so a real machine's 60-200 entry map yields an aperture over live RAM, and containment then admits a child BAR covering kernel memory. Rewritten to walk the map in place, with the hole finder extracted as a pure function and driven by a synthetic 100-entry map in a new test case. Both new tests were verified to fail on the old code. parameters.zig gains the rationale it was missing and loses a stale sentence pointing at the wrong file; vdso.md documents the errno space, including EPEER, which had no written meaning anywhere. docs/os-development/bounds.md is how a ceiling is declared from here. docs/bounds-track-plan.md is the plan to remove the ones we invented. Suite 114 -> 115.
11 KiB
The bounds track: removing the numbers we invented
Plan, 2026-08-08. Follows fixed-bounds-audit.md (235 ceilings, 139 on quantities we do not choose) and the AMD Ryzen that found the first one.
The principles this is derived from
- danOS is a microkernel. Minimise what the kernel is responsible for; move responsibility to user space so it can be restarted, or fixed live during development, without taking the system down.
- Implement the specifications correctly, with the limits those specifications define — not limits we decide.
- Move as much responsibility as possible to user space (the device manager).
- What remains in the kernel is minimal.
- What remains in the kernel is there for security or for a hardware limitation. Nothing else earns a place.
Principle 5 is the test every bound is put to. For each one: is this here because of security, or because of a hardware limitation? If neither, the storage does not belong in the kernel and the bound is not a number to be resized — it is a thing to be moved or deleted.
Applying it to the case that started this:
maximum_devices = 64bounds an inventory of hardware. An inventory is neither a security control nor a hardware limitation. The table is in the wrong place; the number is a symptom.maximum_children_per_parent = 16exists becausedevice_claimis unauthenticated — any process can claim any unclaimed device (devices-broker.zig:164 checks only that the device exists and is free). The cap is a crude proxy for an authorisation the kernel does not perform. Fix the authorisation and the cap has nothing to defend.maximum_domains = 64bounds IOMMU translation domains. Security — stays in the kernel. But VT-d and AMD-Vi both report how many domains they support in a capability register. Principle 2: read it. We chose 64 without asking.
The security invariants
Every phase must leave all five standing. This is the "without punching a hole" half of the brief, and each phase below states how it is checked.
- I1 Containment. A process may map only physical memory inside a resource it was granted. A bus may subdivide only what it already holds.
- I2 Confinement. A DMA-capable device is under IOMMU translation before its driver can program it, or it is not driven at all.
- I3 No self-granted authority. A process holds what it was handed. It cannot name its way into holding more.
- I4 Death releases everything. Every resource a task held is reclaimed when it dies, on every path out.
- I5 Refusal is attributable. Every refusal names the rule that refused it.
Phase 0 — Done
- Errno attribution. One errno space in
system/abi.zig;device_register's six refusals anddevice_claim's three are distinct codes; call sites name the reason;pci-busreconciles found against registered. (I5) - Idempotency ordering. A re-registration consumes no slot, so a full parent re-admits an identical child. A restarted bus is no longer billed for what it rediscovers.
Suite 114/114.
Phase 1 — Reclamation
Nothing may become dynamic before this. Today count only ever increases and
releaseAllOwnedBy clears a dead driver's claims but not its registrations. With a
fixed table that is a slow march to the cap; with dynamic storage it is an unbounded
leak, and every supervisor restart makes it worse.
- Extend the existing death sweep so a task's registrations go with its claims.
- A registration whose owner is gone is removed; its children are re-parented or removed with it (they cannot outlive the authority that published them).
- Test: register under a claimed parent, kill the owner, assert the entries are gone and the ids are not reused while any handle to them lives.
Invariant: I4.
Phase 2 — Close the authorisation hole
The device manager already decides which driver gets which device — it matches against
devices.csv and spawns the driver with the device id as argv[1]. Nothing binds that
decision to the kernel's claim. A driver passes an integer; the kernel checks only
that the device is free.
Per principles 3 and 5: the decision stays in user space; the kernel enforces only possession. The manager hands the driver the device it matched; the kernel's job is that a driver holds what it was handed and nothing else.
- The manager passes a device to the driver it spawned, over the existing cap-passing path. Possession is the authority.
device_claimstops being a way to acquire a device by naming it.- Exclusivity stops being a broker refusing a second claimant and becomes the ordinary property of a thing only one process was given.
maximum_children_per_parent is deleted here, because after this a bus driver's
children are the devices it actually enumerated under a bus it was actually given, and
the rogue-driver-fills-the-table threat the cap was written for no longer exists.
Invariants: I3 (the point of the phase), I1 (containment is unchanged and still checked on every subdivision), I5.
Acceptance: a driver that names a device it was not given is refused, with its own errno. The QEMU suite gains an adversarial case for it — the audit's lesson was that "the suite contains no attacker".
Phase 3 — The inventory moves to user space
The kernel reads only three things out of a device descriptor: physical ranges (to
check a mapping falls inside one), interrupt numbers, and one PCI BDF (to key an
IOMMU domain). Vendor and device ids, class triples, subsystem ids, human-readable
names, bus numbers and parent links are stored solely so device_enumerate can hand
them back. That is the kernel acting as a distribution mechanism for data it does not
use — principle 5 excludes it.
- Zero-resource devices leave the kernel entirely. A USB device addressed through its controller conveys no mapping authority; there is nothing for the kernel to enforce. It is pure inventory and belongs to the device manager. (This is also the case that sidesteps containment, which is why the cap existed.)
- Identity and topology move to the manager, which already receives them as
child_addedreports and already holds the authoritative picture. device_enumerateretires; callers ask the manager, whose protocol already reserves anenumerateverb. Public-ABI change —docs/os-development/vdso.mddocuments it.- What the kernel keeps: for each device that carries resources, the ranges, the GSIs, the BDF, and the owner.
After this, the kernel's table holds only resource-bearing devices, and the remaining count is bounded by what the machine physically has rather than by us.
Invariants: I1, I2 unchanged — both operate on resources, which do not move. I4 must be re-checked: the manager's table now needs its own reclamation, and it is restartable, so it must be able to rebuild from the buses.
Phase 4 — Ask the hardware and the specification
Principle 2, applied to every remaining bound. Each of these is a number the machine or the standard already states, which we replaced with a guess. Independent of each other; can proceed in any order.
| Today | Ask instead |
|---|---|
maximum_domains = 64 (IOMMU) |
the VT-d / AMD-Vi capability register reports the domains supported |
max_devices = 8 (xHCI slots) |
HCSPARAMS1.MaxSlots — the controller says (1–255) |
max_interfaces = 4, max_endpoints |
the configuration descriptor says |
blob: [512]u8 (USB config) |
the device's wTotalLength |
below: [64]Range (memory map) |
UEFI reports the descriptor count |
| AML blobs capped at 6 | the XSDT's length field gives the entry count |
maximum_cpus = 128 |
the MADT entry count |
maximum_gsi = 24 |
the I/O APIC's redirection-entry count; and more than one I/O APIC exists |
| MSI-X vectors | the capability's table-size field (up to 2048) |
Several of these are in user space already (xHCI, USB descriptors) and are ordinary allocations — principle 1 means those are also the safest to do first, since a mistake restarts a driver rather than the machine.
Two in this table are also correctness fixes the audit found, and should carry their
regression tests: the xHCI max_interfaces path misattributes a fifth interface's
endpoints to interface 3, and below: [64]Range silently turns occupied RAM into a PCI
aperture — which is an I1 violation reachable on real hardware, not merely a lost
device. That one is the highest-priority item in this phase.
Phase 5 — What legitimately remains
After phases 1–4 the survivors should be only:
- Pre-allocator storage: the PMM's own frame bitmap, the memory map the loader hands over, the bootstrap page tables. You cannot allocate the allocator. (Hardware/boot limitation — principle 5 admits these.)
- Interrupt-context storage: the IST stack and anything an exception path touches without allocating.
- Wire structures whose layout the other side of a trust boundary parses.
- Facts that are not ceilings: a page is 4096 bytes; an ACPI name segment is 4.
Each is declared per bounds.md — what it counts, who decides its size, what it protects, what happens at the limit, how you find out. And two numbers that must agree agree in code, not in a comment:
comptime {
if (maximum_domains != devices_broker.maximum_devices)
@compileError("iommu.confined is indexed by device id; an id past its end is " ++
"left unconfined while confineDevice still reports success");
}
The one that must not wait
iommu.zig:107 — if (device_id >= confined.len) return true; — returns success without confining. It is unreachable today only because
device ids stop at 64. Phases 3 and 4 both change the device count, and either makes
it live. Fix it before them: out of range must refuse, never allow. (I2)
This is also the standing rule the audit argues for: at a bound, the safe direction is refusal. A ceiling that fails open is not a limit, it is a switch that turns the protection off.
How this is verified
- The QEMU suite is the arbiter at every step; it is 114 cases and must stay green.
- Each fix lands with a test that fails before it — as the idempotency reorder did, where exactly one assertion flipped.
- Adversarial cases for I1–I3 specifically: the audit's six real defects were all found by asking "what would an attacker do", and the suite had never asked.
- The Ryzen is the acceptance test. It is the machine that found this, and the one that proves it fixed.
Sequencing
Phase 1 gates everything. Phase 2 gates phase 3 — the inventory cannot move until authority is sound, or moving it is the hole. Phase 4 is independent and its user-space items are the safest work in the track. Phase 5 is the record of what survived.
The IOMMU fail-open is fixed before phase 3 or 4 touches the device count.