kernel: a refusal names its rule, and two bounds stop failing open

An AMD Ryzen booted to a working compositor with no USB and no storage,
and the log said only "register refused". A tree-wide audit of every
compile-time ceiling followed: 235 of them, 139 on quantities the machine
or a file decides rather than us, 5 documented anywhere, 171 silent when
reached. docs/fixed-bounds-audit.md has the inventory.

Errno attribution. The errno space was split between the kernel and the
envelope, free to drift; it is now one list in system/abi.zig, restated on
both sides, with a comptime check in library/device/driver where the two
halves are visible. device_register's six refusals and device_claim's three
are distinct codes, so a bus driver can say which rule stopped it, and
BadParent splits into NoSuchParent and NotYourParent. pci-bus reconciles
found against registered instead of counting refused functions as found.

Idempotency ordering. The child cap was checked before the identity match,
so a restarted bus was refused its own devices — the supervision restart the
system leans on ratcheted toward a degraded machine. A re-registration
consumes no slot and is now admitted first.

IOMMU fail-closed. confineDevice returned success for a device id past the
confinement table, leaving the device outside every domain while the caller
believed it confined — unreachable only while ids stop at 64, which both the
inventory move and a hardware-reported domain count would change. It refuses
now, and the coupling to the broker's device cap is a comptime assert rather
than a sentence in a comment.

PCI apertures. The bridge's MMIO apertures are derived from the holes in the
firmware memory map, and the derivation copied sub-4 GiB entries into a
fixed [64] array and skipped the rest. A skipped region is not merely lost:
the gap finder concludes it is free, so a real machine's 60-200 entry map
yields an aperture over live RAM, and containment then admits a child BAR
covering kernel memory. Rewritten to walk the map in place, with the hole
finder extracted as a pure function and driven by a synthetic 100-entry map
in a new test case. Both new tests were verified to fail on the old code.

parameters.zig gains the rationale it was missing and loses a stale sentence
pointing at the wrong file; vdso.md documents the errno space, including
EPEER, which had no written meaning anywhere.

docs/os-development/bounds.md is how a ceiling is declared from here.
docs/bounds-track-plan.md is the plan to remove the ones we invented.

Suite 114 -> 115.
This commit is contained in:
Daniel Samson
2026-08-08 11:09:54 +01:00
parent 7db5fba884
commit a86559648e
25 changed files with 1520 additions and 168 deletions
+38 -28
View File
@@ -402,36 +402,42 @@ fn systemDeviceEnumerate(state: *architecture.CpuState) void {
architecture.setSystemCallResult(state, devices_broker.deviceCount());
}
/// device_claim(id) -> 0/-1: take exclusive ownership of a device for this process.
/// device_claim(id) -> 0/-errno: take exclusive ownership of a device for this process.
/// `-ENODEV` no such id, `-EBUSY` a live task already owns it, `-ECONFINE` the claim
/// could not be placed under IOMMU translation and was rolled back.
fn systemDeviceClaim(state: *architecture.CpuState) void {
const device_id = architecture.systemCallArg(state, 0);
const claim_flags = sync.enter();
defer sync.leave(claim_flags);
if (devices_broker.claim(device_id, scheduler.current().id)) {
// Confine the device's DMA before the driver can program it: a PCI function
// becomes reachable to the IOMMU only once claimed (until now its DMA is
// blocked). A claim that cannot be confined must not stand — roll it back —
// since the whole point is that claiming a DMA device is no longer equivalent
// to ring 0. No-op when no IOMMU exists (fail-open).
if (devices_broker.pciAddressOf(device_id)) |bdf| {
const owner = scheduler.current().id;
if (!iommu.confineDevice(device_id, bdf, owner)) {
_ = devices_broker.unclaim(device_id, owner);
return fail(state);
}
// Bind the buffers this task allocated before claiming the device (a driver
// that dma_alloc'd its rings, then claimed the controller).
dmaBindOwnerRegionsInto(owner, device_id);
devices_broker.claim(device_id, scheduler.current().id) catch |e|
return failErr(state, devices_broker.claimErrnoOf(e));
// Confine the device's DMA before the driver can program it: a PCI function
// becomes reachable to the IOMMU only once claimed (until now its DMA is
// blocked). A claim that cannot be confined must not stand — roll it back —
// since the whole point is that claiming a DMA device is no longer equivalent
// to ring 0. No-op when no IOMMU exists (fail-open).
if (devices_broker.pciAddressOf(device_id)) |bdf| {
const owner = scheduler.current().id;
if (!iommu.confineDevice(device_id, bdf, owner)) {
_ = devices_broker.unclaim(device_id, owner);
// Its own errno: "nobody could confine this" is a different world from
// "someone else already has it", and a driver that cannot tell them apart
// cannot report the one that means the machine's DMA protection ran out.
return failErr(state, ipc.ECONFINE);
}
// A display service just took the framebuffer — quiesce the bootstrap console
// so the kernel and the service don't scribble over each other's pixels. The
// claim releases (and the console resumes) automatically if the service dies;
// see releaseTaskResourcesLocked.
if (devices_broker.displayDevice()) |display_id| {
if (device_id == display_id) console.setSuppressed(true);
}
architecture.setSystemCallResult(state, 0);
} else fail(state);
// Bind the buffers this task allocated before claiming the device (a driver
// that dma_alloc'd its rings, then claimed the controller).
dmaBindOwnerRegionsInto(owner, device_id);
}
// A display service just took the framebuffer — quiesce the bootstrap console
// so the kernel and the service don't scribble over each other's pixels. The
// claim releases (and the console resumes) automatically if the service dies;
// see releaseTaskResourcesLocked.
if (devices_broker.displayDevice()) |display_id| {
if (device_id == display_id) console.setSuppressed(true);
}
architecture.setSystemCallResult(state, 0);
}
/// mmio_map(device_id, resource_index) -> virtual_address: map a claimed device's MMIO window into
@@ -931,17 +937,21 @@ fn systemDeviceRegister(state: *architecture.CpuState) void {
const parent_id = architecture.systemCallArg(state, 0);
const descriptor_ptr = architecture.systemCallArg(state, 1);
const t = scheduler.current();
if (t.address_space == 0) return fail(state);
if (t.address_space == 0) return failErr(state, ipc.EPERM);
var descriptor: device_abi.DeviceDescriptor = undefined;
if (!ipc.copyFromUser(t.address_space, descriptor_ptr, std.mem.asBytes(&descriptor))) return fail(state);
if (!ipc.copyFromUser(t.address_space, descriptor_ptr, std.mem.asBytes(&descriptor))) return failErr(state, ipc.EFAULT);
// Under the big kernel lock: the broker's table is also mutated by the
// death sweep (releaseAllOwnedBy) and read by enumerate on other cores —
// ring-3 registration (M19) made those genuinely concurrent.
const flags = sync.enter();
defer sync.leave(flags);
const id = devices_broker.register(parent_id, t.id, &descriptor) catch return fail(state);
// Each refusal carries its own errno (devices-broker.errnoOf) — the bus driver
// logs which rule stopped it, so "this parent is full" is never again mistaken
// for "the table is full" or "that resource escapes your window".
const id = devices_broker.register(parent_id, t.id, &descriptor) catch |e|
return failErr(state, devices_broker.errnoOf(e));
architecture.setSystemCallResult(state, id);
}