kernel: the device table has no ceiling; a runaway is charged to whoever caused it

maximum_devices = 64 is gone. It was a guess about someone else's computer,
and because it was shared, one driver's enumeration starved every other —
which is how an AMD Ryzen booted with a working display, no USB and no
storage. The table now grows from the kernel heap. It was always built after
heap.init; nothing ever prevented this except it having been written static
first.

What replaces it is an allowance charged to the registrar, so a driver
looping device_register exhausts its own and every other driver carries on.
It is declared as what it is — a runaway detector, NOT a security boundary.
A quota generous enough never to bite a real machine is still generous
enough to be unpleasant, and it is not trying to be the defence; delegation
is. What this catches is a legitimate driver in a loop, early, attributably,
and without collateral. Reaching 4096 is a bug report, not a tuning request.

The initial block is 8, deliberately small. Sizing it for a typical machine
would mean the growth path never ran on the hardware we test on and only
woke up on someone else's larger machine — the exact failure shape this
track exists to stop. At 8 it grows several times every boot; disabling
growth now fails the suite with the HPET not fitting, which is the Ryzen
failure in miniature.

The comptime coupling assert added earlier fired, and was right to. confined
(one slot per device id) and domains (the IOMMU's own translation pool) were
sized by the same constant only because device ids happened to stop at 64
too. Two unrelated quantities: confined now grows with the device table,
while maximum_domains stays as the hardware's number — both VT-d and AMD-Vi
report how many domains they support, and reading it is phase 4. The assert
existed for exactly this and did its job.

Suite 118/118.
This commit is contained in:
Daniel Samson
2026-08-08 18:40:20 +01:00
parent 5492607a61
commit f7151ed577
4 changed files with 142 additions and 41 deletions
+37 -23
View File
@@ -34,22 +34,24 @@ const platform = @import("platform");
const architecture = @import("architecture");
const devices_broker = @import("devices-broker.zig");
const log = @import("log.zig");
const heap = @import("heap.zig");
const page_size: u64 = abi.page_size;
const page_mask: u64 = page_size - 1;
const huge_page_size: u64 = 2 * 1024 * 1024;
/// One domain per claimed PCI function. Coupled to devices-broker's device cap — and
/// coupled *in code*, by the comptime assert beside `confined` below, because when
/// these two agreed only by this sentence the disagreement failed open.
/// The IOMMU's own translation-domain pool — one per claimed DMA-capable device.
///
/// Both VT-d and AMD-Vi report the number of domains they support in a capability
/// register. We should be reading it rather than choosing 64 — docs/bounds-track-plan.md
/// phase 4.
/// No longer coupled to the device count. It was, by a comment and then by a comptime
/// assert, only because `confined` (one slot per device id) was sized by this same
/// constant; those are two unrelated quantities and making the device table dynamic
/// separated them. This one is genuinely the hardware's: both VT-d and AMD-Vi report
/// how many domains they support in a capability register, so the honest fix is to read
/// it rather than choose 64 — bounds-track-plan.md phase 4.
///
/// bound: IOMMU translation domains, one per claimed DMA-capable device
/// bound: IOMMU translation domains the kernel can hold at once
/// decided-by: hardware
/// protects: the statically sized domain and confinement tables
/// protects: the statically sized domain pool
/// at-limit: refuse — ECONFINE; the claim is rolled back and the device is not driven,
/// because a claim that cannot be confined must not stand
/// observed-by: the claiming driver's own line naming ECONFINE
@@ -107,18 +109,30 @@ pub fn init() void {
/// Per-claimed-device record: its private domain, so a driver's death tears down
/// exactly the domains it held.
const Confined = struct { active: bool = false, owner: u32 = 0, bdf: u16 = 0, domain: u16 = invalid_domain };
var confined: [maximum_domains]Confined = .{Confined{}} ** maximum_domains;
// `confined` is indexed by **device id**, so it must cover every id the broker can
// mint. These two numbers agreed only by a sentence in a comment above
// `maximum_domains` — and when they disagreed, `confineDevice` returned success for
// the ids it had no room for, leaving those devices unconfined DMA masters. Coupled
// bounds agree in code, not in prose (docs/os-development/bounds.md).
comptime {
if (maximum_domains < devices_broker.maximum_devices)
@compileError("iommu.confined is indexed by device id but is smaller than the " ++
"broker's device table: ids past its end cannot be confined, and so cannot " ++
"be claimed at all");
/// Indexed by **device id**, so it must cover every id the broker can mint — and the
/// broker's table has no ceiling any more, so neither can this. It grows on demand.
///
/// This used to be `[maximum_domains]`, sized by the *domain* constant purely because
/// device ids happened to stop at 64 as well. Two unrelated quantities sharing one
/// number: `domains` below is the IOMMU's own translation-domain pool, which the
/// hardware bounds and reports, while this is one slot per device the machine has.
/// A comptime assert held them together while both were fixed; making the device table
/// dynamic is what forced them apart, which is the assert having done its job.
var confined: []Confined = &.{};
/// Grow `confined` to cover `device_id`. False if the heap cannot — and the caller
/// treats that as a refusal to confine, never as permission.
fn reserveConfined(device_id: u64) bool {
if (device_id < confined.len) return true;
if (device_id >= std.math.maxInt(usize) / 2) return false; // absurd id; refuse rather than size to it
var wanted: usize = if (confined.len == 0) 64 else confined.len;
while (wanted <= device_id) wanted *= 2;
const grown = heap.allocator().realloc(confined, wanted) catch return false;
const previous = confined.len;
confined = grown;
for (confined[previous..]) |*record| record.* = .{};
return true;
}
/// Place a just-claimed PCI function under IOMMU translation on behalf of `owner`: give
@@ -139,7 +153,7 @@ pub fn confineDevice(device_id: u64, bdf: u16, owner: u32) bool {
// at `confined.len`; moving the inventory out of the kernel and taking the domain
// count from the hardware both change that, and either would have made a silent
// unconfined DMA master out of every device past the 64th.
if (device_id >= confined.len) return false;
if (!reserveConfined(device_id)) return false;
const domain = domainCreate(owner, bdf) orelse return false;
// Firmware reserved region for this device, if any (real hardware; QEMU has none).
@@ -181,7 +195,7 @@ pub fn unmapForDevice(device_id: u64, physical: u64, len: u64) void {
/// own freshly-`dma_alloc`'d buffer into the devices it drives.
pub fn mapRegionForOwner(owner: u32, physical: u64, len: u64) void {
if (!active) return;
for (&confined) |*c| {
for (confined) |*c| {
if (c.active and c.owner == owner) _ = map(c.domain, physical, len);
}
}
@@ -192,7 +206,7 @@ pub fn mapRegionForOwner(owner: u32, physical: u64, len: u64) void {
/// domain other than its owner's.
pub fn unmapRegionEverywhere(physical: u64, len: u64) void {
if (!active) return;
for (&confined) |*c| {
for (confined) |*c| {
if (c.active) unmap(c.domain, physical, len);
}
}
@@ -222,7 +236,7 @@ pub fn reassign(device_id: u64, owner: u32) void {
pub fn releaseAllOwnedBy(owner: u32) void {
if (!active) return;
for (&confined) |*c| {
for (confined) |*c| {
if (c.active and c.owner == owner) {
detachDevice(c.bdf);
domainDestroy(c.domain);