kernel: the device table has no ceiling; a runaway is charged to whoever caused it

maximum_devices = 64 is gone. It was a guess about someone else's computer,
and because it was shared, one driver's enumeration starved every other —
which is how an AMD Ryzen booted with a working display, no USB and no
storage. The table now grows from the kernel heap. It was always built after
heap.init; nothing ever prevented this except it having been written static
first.

What replaces it is an allowance charged to the registrar, so a driver
looping device_register exhausts its own and every other driver carries on.
It is declared as what it is — a runaway detector, NOT a security boundary.
A quota generous enough never to bite a real machine is still generous
enough to be unpleasant, and it is not trying to be the defence; delegation
is. What this catches is a legitimate driver in a loop, early, attributably,
and without collateral. Reaching 4096 is a bug report, not a tuning request.

The initial block is 8, deliberately small. Sizing it for a typical machine
would mean the growth path never ran on the hardware we test on and only
woke up on someone else's larger machine — the exact failure shape this
track exists to stop. At 8 it grows several times every boot; disabling
growth now fails the suite with the HPET not fitting, which is the Ryzen
failure in miniature.

The comptime coupling assert added earlier fired, and was right to. confined
(one slot per device id) and domains (the IOMMU's own translation pool) were
sized by the same constant only because device ids happened to stop at 64
too. Two unrelated quantities: confined now grows with the device table,
while maximum_domains stays as the hardware's number — both VT-d and AMD-Vi
report how many domains they support, and reading it is phase 4. The assert
existed for exactly this and did its job.

Suite 118/118.
This commit is contained in:
Daniel Samson
2026-08-08 18:40:20 +01:00
parent 5492607a61
commit f7151ed577
4 changed files with 142 additions and 41 deletions
+89 -17
View File
@@ -22,16 +22,36 @@ const std = @import("std");
const abi = @import("abi");
const platform = @import("platform");
const device_abi = @import("device-abi");
const heap = @import("heap.zig");
/// bound: device nodes for the whole machine — firmware-discovered plus every child a
/// bus driver registers at runtime
/// decided-by: hardware
/// protects: nothing — this is a sizing guess about someone else's computer, which is
/// why an AMD Ryzen booted with a working display, no USB and no storage
/// at-limit: refuse — ENOSPC from device_register; `dropped` counts discovery losses
/// How many devices a single registrar may put in the table.
///
/// **This is a runaway detector, not a security boundary**, and the difference matters.
/// It cannot stop a malicious driver — a quota generous enough never to bite a real
/// machine is still generous enough to be unpleasant — and it is not trying to. What
/// stops malice is that only a driver the manager handed a device can register children
/// under it (docs/os-development/device-authority.md). What this catches is a
/// *legitimate* driver in a loop, early, and attributably: the driver that did it is
/// refused, named in its own log, and restarted, while every other driver is untouched.
///
/// The shared ceilings it replaces could not do that. `maximum_devices = 64` was a
/// guess about someone else's computer and one driver's enumeration starved every
/// other — which is how an AMD Ryzen came to boot with a working display, no USB and no
/// storage. A per-registrar allowance is the same protection charged to whoever caused
/// it, which is the microkernel property rather than a workaround for it.
///
/// The number: a machine's whole PCI segment tops out at 65536 functions, and the
/// biggest real registrar seen is pci-bus at a few dozen. 4096 is far above anything a
/// real machine produces and around 1.4 MiB of descriptors, well under the kernel heap.
/// **Reaching it is a bug report, not a tuning request** — no correct driver gets near.
///
/// bound: devices one task may register
/// decided-by: ours
/// protects: the kernel heap, against a driver looping device_register
/// at-limit: refuse — ECHILDREN to the registrar; every other driver is unaffected
/// observed-by: the bus driver's line naming the reason (pci-bus reconciles functions
/// found against registered), and kernel.zig:203 for discovery drops
pub const maximum_devices = 64;
/// found against registered)
const maximum_devices_per_registrar = 4096;
/// Cap on children a single parent may have. A zero-resource child (legal — a USB
/// device is addressed through its controller, not by MMIO) sidesteps the containment
@@ -50,18 +70,64 @@ pub const maximum_devices = 64;
/// observed-by: pci-bus logs the reason per refused function, and warns at end of scan
const maximum_children_per_parent = 16;
var devices: [maximum_devices]device_abi.DeviceDescriptor = undefined;
var claimed: [maximum_devices]?u32 = .{null} ** maximum_devices; // owner task id, or null
/// The table, grown on demand from the kernel heap. **There is no ceiling**: how many
/// devices a machine has is the machine's business, and no specification bounds it, so
/// nothing here should. `devices_broker.init` runs after `heap.init` (kernel.zig), so
/// there was never a reason for this to be static beyond it having been written that
/// way first.
var devices: []device_abi.DeviceDescriptor = &.{};
/// Owner task id per device, or null. Parallel to `devices` and grown with it.
var claimed: []?u32 = &.{};
/// The task that called `register` for each device, so the per-registrar allowance can
/// be charged to whoever caused the entry. Firmware-discovered nodes carry `no_registrar`
/// — they are the kernel's own, not anybody's doing.
var registrar: []u32 = &.{};
const no_registrar: u32 = 0;
var count: usize = 0;
/// Grow the three parallel arrays so at least one more device fits. False if the heap
/// cannot satisfy it, which the callers report rather than swallow.
fn reserve() bool {
if (count < devices.len) return true;
// Double from a deliberately SMALL first block. Sizing it for a typical machine
// would mean the growth path never ran on the hardware we test on, and only woke up
// on someone else's larger machine — which is the exact shape of the failure this
// whole track exists to stop. At 8, every boot grows the table several times, so
// the path is exercised constantly and the suite asserts it.
const wanted = if (devices.len == 0) 8 else devices.len * 2;
const allocator = heap.allocator();
const grown_devices = allocator.realloc(devices, wanted) catch return false;
devices = grown_devices;
const grown_claimed = allocator.realloc(claimed, wanted) catch return false;
claimed = grown_claimed;
const grown_registrar = allocator.realloc(registrar, wanted) catch return false;
registrar = grown_registrar;
for (claimed[count..], registrar[count..]) |*slot, *who| {
slot.* = null;
who.* = no_registrar;
}
return true;
}
/// How many devices `task` has registered — the allowance is charged per registrar, so
/// a driver in a loop exhausts its own and no one else's.
fn registeredBy(task: u32) usize {
var n: usize = 0;
for (registrar[0..count]) |who| {
if (who == task) n += 1;
}
return n;
}
/// The id of the seeded framebuffer node (`seedDisplay`), or null when the machine
/// handed over no framebuffer. Lets the process layer recognise the display claim
/// (to quiesce the bootstrap console) without threading the id through every caller.
var display_device: ?u64 = null;
/// Devices discovery found but the table had no room for. Non-zero means the machine
/// is bigger than `maximum_devices` and some hardware is simply invisible to drivers —
/// which would otherwise be an entirely silent failure. Logged at boot.
/// Devices discovery found but could not record. The table grows on demand, so this is
/// no longer "the machine is bigger than our guess" — it means the kernel heap could not
/// satisfy the growth, which would otherwise be an entirely silent failure. Logged at
/// boot.
pub var dropped: usize = 0;
/// Snapshot the device tree into the flat table. Run once, right after discovery.
@@ -69,7 +135,8 @@ pub fn init(device_tree: *const platform.DeviceTree) void {
count = 0;
dropped = 0;
display_device = null;
for (&claimed) |*c| c.* = null;
for (claimed) |*c| c.* = null;
for (registrar) |*r| r.* = no_registrar;
walk(device_tree.root, device_abi.no_parent);
}
@@ -81,7 +148,7 @@ pub fn init(device_tree: *const platform.DeviceTree) void {
/// table is full. Idempotent-ish: only ever call once per boot.
pub fn seedDisplay(base: u64, width: u32, height: u32, pitch: u32, format: u32, refresh_hz: u32) ?u64 {
if (base == 0 or width == 0 or height == 0) return null; // headless
if (count >= maximum_devices) {
if (!reserve()) {
dropped += 1;
return null;
}
@@ -126,7 +193,7 @@ fn walk(node: *platform.Device, parent_id: u64) void {
}
fn record(node: *platform.Device, parent_id: u64) u64 {
if (count >= maximum_devices) {
if (!reserve()) {
dropped += 1;
return device_abi.no_parent; // children of a dropped node become roots, not orphans
}
@@ -431,7 +498,11 @@ pub fn register(parent_id: u64, owner: u32, descriptor: *const device_abi.Device
if (existingChild(parent_id, descriptor)) |existing_id| return existing_id;
if (childCount(parent_id) >= maximum_children_per_parent) return error.TooManyChildren;
if (count >= maximum_devices) return error.NoSpace;
// The allowance is charged to whoever is registering, so a driver in a loop
// exhausts its own and every other driver carries on. There is no machine-wide
// ceiling any more: the table grows.
if (registeredBy(owner) >= maximum_devices_per_registrar) return error.TooManyChildren;
if (!reserve()) return error.NoSpace;
const parent = &devices[@intCast(parent_id)];
for (0..@intCast(descriptor.resource_count)) |i| {
@@ -454,6 +525,7 @@ pub fn register(parent_id: u64, owner: u32, descriptor: *const device_abi.Device
for (0..@intCast(descriptor.resource_count)) |i| d.resources[i] = descriptor.resources[i];
devices[count] = d;
registrar[count] = owner;
count += 1;
return d.id;
}