kernel: a refusal names its rule, and two bounds stop failing open

An AMD Ryzen booted to a working compositor with no USB and no storage,
and the log said only "register refused". A tree-wide audit of every
compile-time ceiling followed: 235 of them, 139 on quantities the machine
or a file decides rather than us, 5 documented anywhere, 171 silent when
reached. docs/fixed-bounds-audit.md has the inventory.

Errno attribution. The errno space was split between the kernel and the
envelope, free to drift; it is now one list in system/abi.zig, restated on
both sides, with a comptime check in library/device/driver where the two
halves are visible. device_register's six refusals and device_claim's three
are distinct codes, so a bus driver can say which rule stopped it, and
BadParent splits into NoSuchParent and NotYourParent. pci-bus reconciles
found against registered instead of counting refused functions as found.

Idempotency ordering. The child cap was checked before the identity match,
so a restarted bus was refused its own devices — the supervision restart the
system leans on ratcheted toward a degraded machine. A re-registration
consumes no slot and is now admitted first.

IOMMU fail-closed. confineDevice returned success for a device id past the
confinement table, leaving the device outside every domain while the caller
believed it confined — unreachable only while ids stop at 64, which both the
inventory move and a hardware-reported domain count would change. It refuses
now, and the coupling to the broker's device cap is a comptime assert rather
than a sentence in a comment.

PCI apertures. The bridge's MMIO apertures are derived from the holes in the
firmware memory map, and the derivation copied sub-4 GiB entries into a
fixed [64] array and skipped the rest. A skipped region is not merely lost:
the gap finder concludes it is free, so a real machine's 60-200 entry map
yields an aperture over live RAM, and containment then admits a child BAR
covering kernel memory. Rewritten to walk the map in place, with the hole
finder extracted as a pure function and driven by a synthetic 100-entry map
in a new test case. Both new tests were verified to fail on the old code.

parameters.zig gains the rationale it was missing and loses a stale sentence
pointing at the wrong file; vdso.md documents the errno space, including
EPEER, which had no written meaning anywhere.

docs/os-development/bounds.md is how a ceiling is declared from here.
docs/bounds-track-plan.md is the plan to remove the ones we invented.

Suite 114 -> 115.
This commit is contained in:
Daniel Samson
2026-08-08 11:09:54 +01:00
parent 7db5fba884
commit a86559648e
25 changed files with 1520 additions and 168 deletions
+26 -5
View File
@@ -96,16 +96,37 @@ pub fn init() void {
const Confined = struct { active: bool = false, owner: u32 = 0, bdf: u16 = 0, domain: u16 = invalid_domain };
var confined: [maximum_domains]Confined = .{Confined{}} ** maximum_domains;
// `confined` is indexed by **device id**, so it must cover every id the broker can
// mint. These two numbers agreed only by a sentence in a comment above
// `maximum_domains` — and when they disagreed, `confineDevice` returned success for
// the ids it had no room for, leaving those devices unconfined DMA masters. Coupled
// bounds agree in code, not in prose (docs/os-development/bounds.md).
comptime {
if (maximum_domains < devices_broker.maximum_devices)
@compileError("iommu.confined is indexed by device id but is smaller than the " ++
"broker's device table: ids past its end cannot be confined, and so cannot " ++
"be claimed at all");
}
/// Place a just-claimed PCI function under IOMMU translation on behalf of `owner`: give
/// it a private empty domain, seed it with the device's own firmware reserved region,
/// and attach. Its DMA buffers arrive afterward as explicit grants — the owner's own
/// `dma_alloc`'d regions are bound by the claim path (`mapForDevice`), and cross-process
/// buffers by `dma_bind`. false only if a domain can't be allocated — the caller rolls
/// the claim back (a claim that can't be confined must not stand). No-op success when no
/// IOMMU exists (fail-open).
/// buffers by `dma_bind`. false when the device cannot be confined — the caller rolls
/// the claim back (a claim that can't be confined must not stand).
///
/// **Fail-closed at the table's edge.** There is exactly one deliberate fail-open here:
/// a machine with no IOMMU, which is a fact about the hardware rather than the size of
/// anything. Running out of *room to record* a confinement is not that, and must refuse.
pub fn confineDevice(device_id: u64, bdf: u16, owner: u32) bool {
if (!active) return true;
if (device_id >= confined.len) return true; // unusual id; leave it to fail-open
if (!active) return true; // no IOMMU on this machine — nothing to confine with
// A device id past the end of the record table. This returned `true` — success —
// leaving the device outside every domain while telling the caller it was
// confined, and rolling nothing back. It is unreachable only while device ids stop
// at `confined.len`; moving the inventory out of the kernel and taking the domain
// count from the hardware both change that, and either would have made a silent
// unconfined DMA master out of every device past the 64th.
if (device_id >= confined.len) return false;
const domain = domainCreate(owner, bdf) orelse return false;
// Firmware reserved region for this device, if any (real hardware; QEMU has none).