kernel: a refusal names its rule, and two bounds stop failing open

An AMD Ryzen booted to a working compositor with no USB and no storage,
and the log said only "register refused". A tree-wide audit of every
compile-time ceiling followed: 235 of them, 139 on quantities the machine
or a file decides rather than us, 5 documented anywhere, 171 silent when
reached. docs/fixed-bounds-audit.md has the inventory.

Errno attribution. The errno space was split between the kernel and the
envelope, free to drift; it is now one list in system/abi.zig, restated on
both sides, with a comptime check in library/device/driver where the two
halves are visible. device_register's six refusals and device_claim's three
are distinct codes, so a bus driver can say which rule stopped it, and
BadParent splits into NoSuchParent and NotYourParent. pci-bus reconciles
found against registered instead of counting refused functions as found.

Idempotency ordering. The child cap was checked before the identity match,
so a restarted bus was refused its own devices — the supervision restart the
system leans on ratcheted toward a degraded machine. A re-registration
consumes no slot and is now admitted first.

IOMMU fail-closed. confineDevice returned success for a device id past the
confinement table, leaving the device outside every domain while the caller
believed it confined — unreachable only while ids stop at 64, which both the
inventory move and a hardware-reported domain count would change. It refuses
now, and the coupling to the broker's device cap is a comptime assert rather
than a sentence in a comment.

PCI apertures. The bridge's MMIO apertures are derived from the holes in the
firmware memory map, and the derivation copied sub-4 GiB entries into a
fixed [64] array and skipped the rest. A skipped region is not merely lost:
the gap finder concludes it is free, so a real machine's 60-200 entry map
yields an aperture over live RAM, and containment then admits a child BAR
covering kernel memory. Rewritten to walk the map in place, with the hole
finder extracted as a pure function and driven by a synthetic 100-entry map
in a new test case. Both new tests were verified to fail on the old code.

parameters.zig gains the rationale it was missing and loses a stale sentence
pointing at the wrong file; vdso.md documents the errno space, including
EPEER, which had no written meaning anywhere.

docs/os-development/bounds.md is how a ceiling is declared from here.
docs/bounds-track-plan.md is the plan to remove the ones we invented.

Suite 114 -> 115.
This commit is contained in:
Daniel Samson
2026-08-08 11:09:54 +01:00
parent 7db5fba884
commit a86559648e
25 changed files with 1520 additions and 168 deletions
+40 -2
View File
@@ -43,11 +43,11 @@ pub const SystemCall = enum(u64) {
ipc_call = 9, // ipc_call(h, message, len, reply, cap) -> reply_len: send + block for reply
ipc_reply_wait = 10, // ipc_reply_wait(h, reply, len, receive, cap) -> receive_len (+badge in rdx)
device_enumerate = 11, // device_enumerate(buffer, maximum) -> count: snapshot the device table
device_claim = 12, // device_claim(id) -> ok: take exclusive ownership of a device
device_claim = 12, // device_claim(id) -> 0/-errno: take exclusive ownership of a device (-ENODEV no such id, -EBUSY someone owns it, -ECONFINE the IOMMU would not confine it)
mmio_map = 13, // mmio_map(id, resource_index) -> virtual_address: map a claimed device's MMIO into this address space
irq_bind = 14, // irq_bind(id, resource_index, endpoint): deliver a device IRQ as an IPC notification
irq_ack = 15, // irq_ack(id, resource_index): re-arm a bound IRQ after servicing it
device_register = 16, // device_register(parent_id, descriptor) -> id: publish a child of a device you claimed
device_register = 16, // device_register(parent_id, descriptor) -> id/-errno: publish a child of a device you claimed (-ENOSPC table full, -ECHILDREN parent full, -ERANGE resource escapes the parent, -ENODEV/-EPERM bad parent, -E2BIG too many resources)
system_spawn = 17, // system_spawn(name_ptr, name_len, arguments_ptr, arguments_len, exit_endpoint) -> child process id: start a named initial-ramdisk binary as a new ring-3 process
dma_alloc = 18, // dma_alloc(len, flags) -> virtual_address (rax), physical_address (rdx): contiguous, pinned, uncacheable DMA memory
dma_free = 19, // dma_free(virtual_address, len) -> 0: release a prior dma_alloc
@@ -88,6 +88,44 @@ pub const SystemCall = enum(u64) {
_,
};
/// **The errno space** — the one vocabulary of refusal, returned as `-value` in the
/// system_call result register and echoed by user-space providers in a reply status.
/// Canonical here because it crosses the kernel↔user boundary in both directions:
/// the kernel restates these in system/kernel/ipc-synchronous.zig, and the envelope
/// restates the provider-facing subset in library/protocol/envelope/envelope.zig
/// (which cannot import this module — the `protocol` package deliberately has no
/// dependencies, so a comptime cross-check in library/device/driver/driver.zig and
/// system/kernel/tests.zig holds the two halves together).
///
/// The rule these serve: **a refusal must say which rule refused it.** A caller that
/// gets one number for five different reasons cannot report, retry or route around
/// any of them — see docs/fixed-bounds-audit.md, where a bare -1 turned "this bus is
/// at its child cap" into a machine that booted with no USB and no storage, and cost
/// a debugging session to tell apart from four other causes.
///
/// Values are stable: `failed()` in the runtime treats the top 4096 return values as
/// errors (the Linux convention), so anything here must stay well inside 1..4095.
pub const EBADF: i64 = 1; // bad handle
pub const E2BIG: i64 = 2; // argument exceeds its maximum (a message, a descriptor's resource count)
pub const EFAULT: i64 = 3; // buffer unmapped / outside the user half
pub const ENOENT: i64 = 4; // no such name
pub const ENOSPC: i64 = 5; // a kernel table is full (handles, devices)
pub const ENOMEM: i64 = 6; // out of memory
pub const EPEER: i64 = 7; // peer died before replying (its process exited or was killed)
pub const ESRCH: i64 = 8; // no such process (process_kill of an unknown/dead id)
pub const EPERM: i64 = 9; // not permitted (the caller is not the owner/supervisor)
pub const ENOSYS: i64 = 10; // this protocol has no such operation
pub const EPROTO: i64 = 11; // malformed packet: shorter than the verb it names
pub const EBUSY: i64 = 12; // the thing asked for is held by someone still alive
pub const ENODEV: i64 = 13; // no such device id
pub const ECHILDREN: i64 = 14; // this parent already holds as many children as it can
pub const ERANGE: i64 = 15; // a resource escapes the window it must fall inside
pub const ECONFINE: i64 = 16; // the device could not be placed under IOMMU translation
/// The highest errno defined above. A cheap guard for anyone switching over the
/// space, and the number to bump when adding one.
pub const errno_maximum: i64 = 16;
/// `futex_wait` return codes (in rax).
pub const futex_woken: u64 = 0; // woken by a futex_wake
pub const futex_mismatch: u64 = 1; // *addr != expected on entry; the caller did not block
+20 -6
View File
@@ -44,6 +44,10 @@ var ecam_physical: u64 = 0;
var start_bus: u64 = 0;
var bus_count: u64 = 0;
var manager_handle: ipc.Handle = 0;
/// Functions this scan discovered but could not publish. Counted so the end of the
/// scan can reconcile "found" against "registered" — a scan that silently returns a
/// subset of the machine is the failure this driver is most able to hide.
var refused: u32 = 0;
/// One aligned 32-bit read from a function's configuration space.
fn configRead(bus: u64, dev: u64, function: u64, offset: u64) u32 {
@@ -74,10 +78,10 @@ fn configWrite16(bus: u64, dev: u64, function: u64, offset: u64, value: u16) voi
/// Claim the bridge, map the ECAM, hello the manager, then scan.
fn initialise(endpoint: ipc.Handle) bool {
_ = endpoint;
if (!device.claim(bridge_id)) {
std.log.info("unable to claim bridge device {d}", .{bridge_id});
device.claim(bridge_id) catch |e| {
std.log.info("unable to claim bridge device {d}: {s}", .{ bridge_id, @errorName(e) });
return false;
}
};
const buffer = memory.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = logging.write("/system/drivers/pci-bus: out of memory\n");
return false;
@@ -141,7 +145,13 @@ fn scan() void {
}
}
}
std.log.info("{d} functions found", .{found});
// Reconcile: "found" alone reads as success even when most of the machine was
// refused. If the two disagree, say so at a level that survives a scrollback.
if (refused == 0) {
std.log.info("{d} functions found, all registered", .{found});
} else {
std.log.warn("{d} functions found, {d} REFUSED — {d} registered", .{ found, refused, found - refused });
}
}
/// Register one function under the bridge and report it to the manager. The
@@ -225,8 +235,12 @@ fn registerAndReport(bus: u64, dev: u64, function: u64, class_triple: u32) void
// subsystem are read. Every discovered function is logged, matched or not.
logFunction(bus, dev, function, class_triple, descriptor.vendor, descriptor.device, descriptor.subsystem);
const registered = device.register(bridge_id, &descriptor) orelse {
std.log.info("register refused for {d}:{d}.{d}", .{ bus, dev, function });
// Name the rule that refused. `ParentFull` and `TableFull` are different ceilings
// with different fixes, and `NotContained` is not a ceiling at all — it means the
// BAR escaped the bridge's own window (docs/fixed-bounds-audit.md).
const registered = device.register(bridge_id, &descriptor) catch |e| {
refused += 1;
std.log.warn("register refused for {d}:{d}.{d}: {s}", .{ bus, dev, function, @errorName(e) });
return;
};
// The registered device id is the packet's target — the manager's object
+15 -11
View File
@@ -114,10 +114,10 @@ pub fn main() void {
_ = logging.write("/system/drivers/ps2-bus: found PS/2 controller\n");
_ = logging.write("/system/drivers/ps2-bus: initializing controller\n");
if (!device.claim(controller_device_descriptor.id)) {
_ = logging.write("/system/drivers/ps2-bus: unable to claim controller \n");
device.claim(controller_device_descriptor.id) catch |e| {
std.log.warn("unable to claim controller: {s}", .{@errorName(e)});
return;
}
};
const controller = ps2.Controller.init(controller_device_descriptor) orelse {
_ = logging.write("/system/drivers/ps2-bus: controller is missing its IO ports\n");
@@ -255,14 +255,18 @@ pub fn main() void {
if (port_device_types[@intFromEnum(ps2.Port.two)] != null) {
if (ps2.findMouseDescriptor(buffer)) |descriptor| {
if (findInterruptResourceIndex(descriptor)) |auxiliary_index| {
if (device.claim(descriptor.id) and device.irqBind(descriptor.id, auxiliary_index, endpoint)) {
maybe_auxiliary_interrupt = .{
.device_id = descriptor.id,
.interrupt_index = auxiliary_index,
.gsi = descriptor.resources[auxiliary_index].start,
};
} else {
_ = logging.write("/system/drivers/ps2-bus: auxiliary irq_bind failed\n");
if (device.claim(descriptor.id)) |_| {
if (device.irqBind(descriptor.id, auxiliary_index, endpoint)) {
maybe_auxiliary_interrupt = .{
.device_id = descriptor.id,
.interrupt_index = auxiliary_index,
.gsi = descriptor.resources[auxiliary_index].start,
};
} else {
_ = logging.write("/system/drivers/ps2-bus: auxiliary irq_bind failed\n");
}
} else |e| {
std.log.warn("auxiliary claim failed: {s}", .{@errorName(e)});
}
}
}
+5 -5
View File
@@ -173,10 +173,10 @@ fn initialise(endpoint: ipc.Handle) bool {
if (!channel.bindPatiently("usb-transfer", endpoint))
_ = logging.write("/system/drivers/usb-xhci-bus: /protocol/usb-transfer is another controller's; serving mine unnamed\n");
if (!device.claim(controller_id)) {
std.log.info("unable to claim controller device {d}", .{controller_id});
device.claim(controller_id) catch |e| {
std.log.warn("unable to claim controller device {d}: {s}", .{ controller_id, @errorName(e) });
return false;
}
};
// Fetch our own descriptor back for the controller's resources.
const buffer = memory.allocator().alloc(device.DeviceDescriptor, 64) catch {
@@ -522,8 +522,8 @@ fn reportInterface(manager: ipc.Handle, port: u32, interface: library.InterfaceI
const hid_text = std.fmt.bufPrint(&hid_buffer, "P{d}I{d}", .{ port, interface.number }) catch "";
descriptor.hid_len = hid_text.len;
@memcpy(descriptor.hid[0..hid_text.len], hid_text);
const registered = device.register(controller_id, &descriptor) orelse {
std.log.info("register refused for port {d} interface {d}", .{ port, interface.number });
const registered = device.register(controller_id, &descriptor) catch |e| {
std.log.warn("register refused for port {d} interface {d}: {s}", .{ port, interface.number, @errorName(e) });
return null;
};
+3 -3
View File
@@ -200,10 +200,10 @@ fn testPixel(index: u32) u32 {
fn initialise(endpoint: ipc.Handle) bool {
_ = endpoint;
if (!device.claim(device_id)) {
std.log.info("unable to claim device {d}", .{device_id});
device.claim(device_id) catch |e| {
std.log.info("unable to claim device {d}: {s}", .{ device_id, @errorName(e) });
return false;
}
};
var descriptors: [64]device.DeviceDescriptor = undefined;
const total = device.enumerate(&descriptors);
+82 -40
View File
@@ -615,17 +615,78 @@ var boot_memory_regions: []const boot_handoff.MemoryRegion = &.{};
/// I/O-APIC region, the high one from 4 GiB (or the end of RAM above it) to
/// the 46-bit line. Coarse, mechanical, and AML-free — available at boot no
/// matter what later moved to user space.
pub const AddressRange = struct { base: u64, end: u64 };
/// The largest holes below 4 GiB in a firmware memory map: the ranges the firmware
/// described *nothing* in, which is where a PCI BAR may legitimately live. Fills `out`
/// (keeping the largest, replacing the smallest held so far), ignores holes shorter
/// than `minimum`, and returns how many entries are non-empty; entries the caller
/// reads may be empty and are skipped by `end > base`.
///
/// **It never copies the map, and that is the point.** The firmware chooses how many
/// descriptors its map has — commonly 60–200 on a real machine, 15–25 under OVMF. The
/// previous version copied the sub-4 GiB entries into a fixed `[64]` array and
/// `continue`d past the rest, which did not merely lose them: a region absent from the
/// walk is a region this function concludes is *free*, so a large enough map yields an
/// "aperture" lying over live RAM. `device_register` containment would then admit a
/// child BAR covering kernel memory, and its claimant could `mmio_map` it. A bound
/// whose overflow hands out authority is not a limit; the fix is not a bigger array.
///
/// Pure, allocation-free, and linear-ish in the map (each inner pass consumes at least
/// one region, and it runs once at boot).
pub fn largestHolesBelow4G(
regions: []const boot_handoff.MemoryRegion,
minimum: u64,
out: []AddressRange,
) usize {
const limit: u64 = 1 << 32;
for (out) |*hole| hole.* = .{ .base = 0, .end = 0 };
var cursor: u64 = 0;
while (cursor < limit) {
// Step over every described region covering the cursor. Regions may overlap
// and chain, so repeat until the cursor stops moving.
var moved = true;
while (moved) {
moved = false;
for (regions) |region| {
if (region.base >= limit) continue;
const end = @min(region.base + region.pages * 4096, limit);
if (region.base <= cursor and end > cursor) {
cursor = end;
moved = true;
}
}
}
if (cursor >= limit) break;
// The cursor now sits in a hole; it runs to the next described base, or to
// 4 GiB if nothing is described above it.
var next: u64 = limit;
for (regions) |region| {
if (region.base >= limit) continue;
if (region.base > cursor and region.base < next) next = region.base;
}
if (next - cursor >= minimum and out.len != 0) {
var smallest: usize = 0;
for (out, 0..) |hole, i| {
if (hole.end - hole.base < out[smallest].end - out[smallest].base) smallest = i;
}
if (next - cursor > out[smallest].end - out[smallest].base)
out[smallest] = .{ .base = cursor, .end = next };
}
cursor = next;
}
var found: usize = 0;
for (out) |hole| {
if (hole.end > hole.base) found += 1;
}
return found;
}
fn addBridgeApertures(bridge: *device_model.Device) void {
// Below 4 GiB the described regions are sparse (RAM low, firmware flash
// and tables high), so the holes are the *gaps between* them — a single
// "after the last region" rule dies on OVMF's flash at the very top.
// Sort-merge the described ranges, then keep the three largest gaps
// (resource slots are bounded at 8 per device; ECAM + bus range + 3 + the
// high aperture fits). Above 4 GiB one aperture runs from the end of the
// described space to the 46-bit line.
const Range = struct { base: u64, end: u64 };
var below: [64]Range = undefined;
var below_count: usize = 0;
var high_end: u64 = 1 << 32;
for (boot_memory_regions) |region| {
const end = region.base + region.pages * 4096;
@@ -636,37 +697,18 @@ fn addBridgeApertures(bridge: *device_model.Device) void {
// kernel image, the tables, the ramdisk all live there). Bring-up
// trust: only the bridge's claimant can register into the aperture.
if (region.kind == .usable and end > high_end) high_end = end;
if (region.base >= (1 << 32) or below_count == below.len) continue;
below[below_count] = .{ .base = region.base, .end = @min(end, 1 << 32) };
below_count += 1;
}
// Insertion sort by base (the map is small and this runs once at boot).
for (1..below_count) |i| {
const key = below[i];
var j = i;
while (j > 0 and below[j - 1].base > key.base) : (j -= 1) below[j] = below[j - 1];
below[j] = key;
}
// Walk the sorted ranges, collecting inter-region gaps of at least 1 MiB.
var gaps: [3]Range = .{Range{ .base = 0, .end = 0 }} ** 3;
var cursor: u64 = 0;
var index: usize = 0;
while (index <= below_count) : (index += 1) {
const gap_end = if (index == below_count) (1 << 32) else below[index].base;
if (gap_end > cursor and gap_end - cursor >= (1 << 20)) {
// Keep the three largest, replacing the smallest kept so far.
var smallest: usize = 0;
for (gaps, 0..) |gap, gi| {
if (gap.end - gap.base < gaps[smallest].end - gaps[smallest].base) smallest = gi;
}
if (gap_end - cursor > gaps[smallest].end - gaps[smallest].base) {
gaps[smallest] = .{ .base = cursor, .end = gap_end };
}
}
if (index < below_count and below[index].end > cursor) cursor = below[index].end;
}
for (gaps) |gap| {
if (gap.end > gap.base) _ = bridge.addResource(.memory, gap.base, gap.end - gap.base);
// Three holes below 4 GiB. This three is not a guess about hardware: a
// `DeviceDescriptor` carries `maximum_device_resources` (8) resources, and the
// bridge spends them on ECAM + bus range + these + the high aperture. That cap is
// a wire struct in the kernel↔user ABI, so widening it is a separate change —
// docs/bounds-track-plan.md, phase 4. Keeping *fewer* holes is fail-closed: it
// refuses BARs, it never admits one.
var holes: [3]AddressRange = undefined;
_ = largestHolesBelow4G(boot_memory_regions, 1 << 20, &holes);
for (holes) |hole| {
if (hole.end > hole.base) _ = bridge.addResource(.memory, hole.base, hole.end - hole.base);
}
_ = bridge.addResource(.memory, high_end, (@as(u64, 1) << 46) - high_end);
}
+86 -31
View File
@@ -19,10 +19,11 @@
//! ever subdivide what it was already given.
const std = @import("std");
const abi = @import("abi");
const platform = @import("platform");
const device_abi = @import("device-abi");
const maximum_devices = 64;
pub const maximum_devices = 64;
/// Cap on children a single parent may have. A zero-resource child (legal — a USB
/// device is addressed through its controller, not by MMIO) sidesteps the containment
@@ -157,13 +158,13 @@ pub fn enumerateFrom(start: usize, out: []device_abi.DeviceDescriptor) usize {
return n;
}
/// Take exclusive ownership of device `id` for task `owner`. Fails if the id is
/// out of range or already claimed.
pub fn claim(id: u64, owner: u32) bool {
if (id >= count) return false;
if (claimed[@intCast(id)] != null) return false;
/// Take exclusive ownership of device `id` for task `owner`. The two ways this can
/// fail want different responses from a driver — a stale id means re-enumerate, a
/// live claimant means back off — so they are distinguishable (`ClaimError`).
pub fn claim(id: u64, owner: u32) ClaimError!void {
if (id >= count) return error.NoSuchDevice;
if (claimed[@intCast(id)] != null) return error.AlreadyClaimed;
claimed[@intCast(id)] = owner;
return true;
}
/// The task that owns device `id`, or null.
@@ -274,14 +275,70 @@ fn contains(parent: device_abi.ResourceDescriptor, child: device_abi.ResourceDes
return child.start >= parent.start and child_end <= parent_end;
}
/// Why a `register` was refused. Each variant maps to its own errno (`errnoOf`), so
/// a bus driver's log line can name the rule that stopped it — "this parent is at
/// its child cap" and "the table is full" want different fixes, and telling them
/// apart from a bare -1 cost a debugging session (docs/fixed-bounds-audit.md).
pub const RegisterError = error{
NoSpace, // the device table is full
BadParent, // no such device, or not claimed by this task
TooManyResources,
NoSuchParent, // no device with that id
NotYourParent, // that device exists but this task has not claimed it
TooManyResources, // the descriptor declares more resources than one device may hold
TooManyChildren, // this parent is at maximum_children_per_parent
NotContained, // a child resource escapes its parent's window
};
/// The errno a refused `register` returns to ring 3.
pub fn errnoOf(e: RegisterError) i64 {
return switch (e) {
error.NoSpace => abi.ENOSPC,
error.NoSuchParent => abi.ENODEV,
error.NotYourParent => abi.EPERM,
error.TooManyResources => abi.E2BIG,
error.TooManyChildren => abi.ECHILDREN,
error.NotContained => abi.ERANGE,
};
}
/// Why a `claim` was refused.
pub const ClaimError = error{
NoSuchDevice, // no device with that id
AlreadyClaimed, // a live task already owns it
};
/// The errno a refused `claim` returns to ring 3. (`ECONFINE` — the claim stood but
/// the IOMMU would not confine the device — is raised by the caller in
/// system/kernel/process.zig, which is what rolls the claim back.)
pub fn claimErrnoOf(e: ClaimError) i64 {
return switch (e) {
error.NoSuchDevice => abi.ENODEV,
error.AlreadyClaimed => abi.EBUSY,
};
}
/// The id of a child of `parent_id` already identical to `descriptor`, or null.
/// Exact on class, identity and every resource — anything less would let a bus
/// silently adopt an entry that is not the device it just found. The caller must
/// have bounded `descriptor.resource_count` first.
fn existingChild(parent_id: u64, descriptor: *const device_abi.DeviceDescriptor) ?u64 {
for (devices[0..count]) |*existing| {
if (existing.parent != parent_id) continue;
if (existing.class != descriptor.class) continue;
if (existing.pci_class != descriptor.pci_class) continue;
if (existing.hid_len != descriptor.hid_len) continue;
if (!std.mem.eql(u8, existing.hid[0..@intCast(existing.hid_len)], descriptor.hid[0..@intCast(descriptor.hid_len)])) continue;
if (existing.resource_count != descriptor.resource_count) continue;
var same = true;
for (0..@intCast(descriptor.resource_count)) |i| {
const a = existing.resources[i];
const b = descriptor.resources[i];
if (a.kind != b.kind or a.start != b.start or a.len != b.len) same = false;
}
if (same) return existing.id;
}
return null;
}
/// Number of devices currently recorded with `parent_id` as their parent.
fn childCount(parent_id: u64) usize {
var n: usize = 0;
@@ -299,9 +356,27 @@ fn childCount(parent_id: u64) usize {
/// contained in a parent resource of the same kind. A device with no resources is
/// fine and common: a USB device is addressed through its controller, not by MMIO.
pub fn register(parent_id: u64, owner: u32, descriptor: *const device_abi.DeviceDescriptor) RegisterError!u64 {
const parent_owner = ownerOf(parent_id) orelse return error.BadParent;
if (parent_owner != owner) return error.BadParent;
if (parent_id >= count) return error.NoSuchParent;
const parent_owner = ownerOf(parent_id) orelse return error.NotYourParent;
if (parent_owner != owner) return error.NotYourParent;
// Bounds every `descriptor.resources` read below, including the match scan's.
if (descriptor.resource_count > device_abi.maximum_device_resources) return error.TooManyResources;
// Idempotent on exact match (docs/device-manager.md): a restarted registering
// bus re-registers what it rediscovers, and the table has no unregister — an
// identical child under the same parent returns the existing id instead of
// appending a duplicate.
//
// Checked **before the caps**, because a re-registration consumes no slot.
// Charging it against the child cap refused a restarted bus its own devices the
// second time it started, which turned the supervision restart this system leans
// on into a one-way ratchet toward a degraded machine. The match is exact — class,
// identity, and every resource — so an entry returned this way was contained when
// it was first admitted, and a stored descriptor's resources never change
// afterwards (they are written only by `record`, `seedDisplay` and the append
// below).
if (existingChild(parent_id, descriptor)) |existing_id| return existing_id;
if (childCount(parent_id) >= maximum_children_per_parent) return error.TooManyChildren;
if (count >= maximum_devices) return error.NoSpace;
@@ -315,26 +390,6 @@ pub fn register(parent_id: u64, owner: u32, descriptor: *const device_abi.Device
if (!ok) return error.NotContained;
}
// Idempotent on exact match (docs/device-manager.md): a restarted
// registering bus re-registers what it rediscovers, and the table has no
// unregister — an identical (class, identity, resources) child under the
// same parent returns the existing id instead of appending a duplicate.
for (devices[0..count]) |*existing| {
if (existing.parent != parent_id) continue;
if (existing.class != descriptor.class) continue;
if (existing.pci_class != descriptor.pci_class) continue;
if (existing.hid_len != descriptor.hid_len) continue;
if (!std.mem.eql(u8, existing.hid[0..@intCast(existing.hid_len)], descriptor.hid[0..@intCast(descriptor.hid_len)])) continue;
if (existing.resource_count != descriptor.resource_count) continue;
var same = true;
for (0..@intCast(descriptor.resource_count)) |i| {
const a = existing.resources[i];
const b = descriptor.resources[i];
if (a.kind != b.kind or a.start != b.start or a.len != b.len) same = false;
}
if (same) return existing.id;
}
var d = std.mem.zeroes(device_abi.DeviceDescriptor);
d.id = count;
d.parent = parent_id;
+26 -5
View File
@@ -96,16 +96,37 @@ pub fn init() void {
const Confined = struct { active: bool = false, owner: u32 = 0, bdf: u16 = 0, domain: u16 = invalid_domain };
var confined: [maximum_domains]Confined = .{Confined{}} ** maximum_domains;
// `confined` is indexed by **device id**, so it must cover every id the broker can
// mint. These two numbers agreed only by a sentence in a comment above
// `maximum_domains` — and when they disagreed, `confineDevice` returned success for
// the ids it had no room for, leaving those devices unconfined DMA masters. Coupled
// bounds agree in code, not in prose (docs/os-development/bounds.md).
comptime {
if (maximum_domains < devices_broker.maximum_devices)
@compileError("iommu.confined is indexed by device id but is smaller than the " ++
"broker's device table: ids past its end cannot be confined, and so cannot " ++
"be claimed at all");
}
/// Place a just-claimed PCI function under IOMMU translation on behalf of `owner`: give
/// it a private empty domain, seed it with the device's own firmware reserved region,
/// and attach. Its DMA buffers arrive afterward as explicit grants — the owner's own
/// `dma_alloc`'d regions are bound by the claim path (`mapForDevice`), and cross-process
/// buffers by `dma_bind`. false only if a domain can't be allocated — the caller rolls
/// the claim back (a claim that can't be confined must not stand). No-op success when no
/// IOMMU exists (fail-open).
/// buffers by `dma_bind`. false when the device cannot be confined — the caller rolls
/// the claim back (a claim that can't be confined must not stand).
///
/// **Fail-closed at the table's edge.** There is exactly one deliberate fail-open here:
/// a machine with no IOMMU, which is a fact about the hardware rather than the size of
/// anything. Running out of *room to record* a confinement is not that, and must refuse.
pub fn confineDevice(device_id: u64, bdf: u16, owner: u32) bool {
if (!active) return true;
if (device_id >= confined.len) return true; // unusual id; leave it to fail-open
if (!active) return true; // no IOMMU on this machine — nothing to confine with
// A device id past the end of the record table. This returned `true` — success —
// leaving the device outside every domain while telling the caller it was
// confined, and rolling nothing back. It is unreachable only while device ids stop
// at `confined.len`; moving the inventory out of the kernel and taking the domain
// count from the hardware both change that, and either would have made a silent
// unconfined DMA master out of every device past the 64th.
if (device_id >= confined.len) return false;
const domain = domainCreate(owner, bdf) orelse return false;
// Firmware reserved region for this device, if any (real hardware; QEMU has none).
+16 -9
View File
@@ -42,15 +42,22 @@ pub const MESSAGE_MAXIMUM: usize = 256;
pub const maximum_handles = scheduler.ipc_maximum_handles;
/// Errno-style failures, returned as `-value` in the system_call result register.
pub const EBADF: i64 = 1; // bad handle
pub const E2BIG: i64 = 2; // message exceeds MESSAGE_MAXIMUM
pub const EFAULT: i64 = 3; // buffer unmapped / out of the user half
pub const ENOENT: i64 = 4; // no such name
pub const ENOSPC: i64 = 5; // handle table full
pub const ENOMEM: i64 = 6; // out of memory
pub const EPEER: i64 = 7; // peer died before replying (its process exited or was killed)
pub const ESRCH: i64 = 8; // no such process (process_kill of an unknown/dead id)
pub const EPERM: i64 = 9; // not permitted (process_kill by anyone but the supervisor)
/// Restated from the shared kernel↔user ABI (system/abi.zig), because ring 3 reads
/// the same numbers — the same reason `notify_badge_bit` below is restated. New
/// codes are added *there*, which is where the space is documented.
pub const EBADF = abi.EBADF; // bad handle
pub const E2BIG = abi.E2BIG; // message exceeds MESSAGE_MAXIMUM
pub const EFAULT = abi.EFAULT; // buffer unmapped / out of the user half
pub const ENOENT = abi.ENOENT; // no such name
pub const ENOSPC = abi.ENOSPC; // a kernel table is full
pub const ENOMEM = abi.ENOMEM; // out of memory
pub const EPEER = abi.EPEER; // peer died before replying (its process exited or was killed)
pub const ESRCH = abi.ESRCH; // no such process (process_kill of an unknown/dead id)
pub const EPERM = abi.EPERM; // not permitted (process_kill by anyone but the supervisor)
pub const ENODEV = abi.ENODEV; // no such device id
pub const ECHILDREN = abi.ECHILDREN; // this parent is at its child cap
pub const ERANGE = abi.ERANGE; // a resource escapes its parent's window
pub const ECONFINE = abi.ECONFINE; // the device could not be placed under IOMMU translation
/// A badge with this bit set is an asynchronous notification (e.g. an IRQ), not a
/// message from a client — there is no reply owed. The low bits carry the source
+8
View File
@@ -26,6 +26,14 @@ pub const RegisterAccess = acpi.RegisterAccess;
pub const IsoEntry = acpi.IsoEntry;
pub const Cpu = acpi.Cpu;
/// Where a PCI BAR may legitimately live: the holes in the firmware memory map.
/// Firmware-agnostic — it takes a boot-handoff map, not an ACPI table — and pure, so
/// the kernel self-test can drive it with a synthetic map. That is the only way to
/// check the invariant that matters here: an aperture must never cover memory the
/// firmware described, because containment would then admit a BAR over live RAM.
pub const AddressRange = acpi.AddressRange;
pub const largestHolesBelow4G = acpi.largestHolesBelow4G;
/// The FADT power register map discovery extracted (PM1 control, reset register),
/// for kernel reboot and diagnostics. Sleep-state values are userspace's (S5 is
/// owned by the ring-3 acpi service), so they are not here.
+38 -28
View File
@@ -402,36 +402,42 @@ fn systemDeviceEnumerate(state: *architecture.CpuState) void {
architecture.setSystemCallResult(state, devices_broker.deviceCount());
}
/// device_claim(id) -> 0/-1: take exclusive ownership of a device for this process.
/// device_claim(id) -> 0/-errno: take exclusive ownership of a device for this process.
/// `-ENODEV` no such id, `-EBUSY` a live task already owns it, `-ECONFINE` the claim
/// could not be placed under IOMMU translation and was rolled back.
fn systemDeviceClaim(state: *architecture.CpuState) void {
const device_id = architecture.systemCallArg(state, 0);
const claim_flags = sync.enter();
defer sync.leave(claim_flags);
if (devices_broker.claim(device_id, scheduler.current().id)) {
// Confine the device's DMA before the driver can program it: a PCI function
// becomes reachable to the IOMMU only once claimed (until now its DMA is
// blocked). A claim that cannot be confined must not stand — roll it back —
// since the whole point is that claiming a DMA device is no longer equivalent
// to ring 0. No-op when no IOMMU exists (fail-open).
if (devices_broker.pciAddressOf(device_id)) |bdf| {
const owner = scheduler.current().id;
if (!iommu.confineDevice(device_id, bdf, owner)) {
_ = devices_broker.unclaim(device_id, owner);
return fail(state);
}
// Bind the buffers this task allocated before claiming the device (a driver
// that dma_alloc'd its rings, then claimed the controller).
dmaBindOwnerRegionsInto(owner, device_id);
devices_broker.claim(device_id, scheduler.current().id) catch |e|
return failErr(state, devices_broker.claimErrnoOf(e));
// Confine the device's DMA before the driver can program it: a PCI function
// becomes reachable to the IOMMU only once claimed (until now its DMA is
// blocked). A claim that cannot be confined must not stand — roll it back —
// since the whole point is that claiming a DMA device is no longer equivalent
// to ring 0. No-op when no IOMMU exists (fail-open).
if (devices_broker.pciAddressOf(device_id)) |bdf| {
const owner = scheduler.current().id;
if (!iommu.confineDevice(device_id, bdf, owner)) {
_ = devices_broker.unclaim(device_id, owner);
// Its own errno: "nobody could confine this" is a different world from
// "someone else already has it", and a driver that cannot tell them apart
// cannot report the one that means the machine's DMA protection ran out.
return failErr(state, ipc.ECONFINE);
}
// A display service just took the framebuffer — quiesce the bootstrap console
// so the kernel and the service don't scribble over each other's pixels. The
// claim releases (and the console resumes) automatically if the service dies;
// see releaseTaskResourcesLocked.
if (devices_broker.displayDevice()) |display_id| {
if (device_id == display_id) console.setSuppressed(true);
}
architecture.setSystemCallResult(state, 0);
} else fail(state);
// Bind the buffers this task allocated before claiming the device (a driver
// that dma_alloc'd its rings, then claimed the controller).
dmaBindOwnerRegionsInto(owner, device_id);
}
// A display service just took the framebuffer — quiesce the bootstrap console
// so the kernel and the service don't scribble over each other's pixels. The
// claim releases (and the console resumes) automatically if the service dies;
// see releaseTaskResourcesLocked.
if (devices_broker.displayDevice()) |display_id| {
if (device_id == display_id) console.setSuppressed(true);
}
architecture.setSystemCallResult(state, 0);
}
/// mmio_map(device_id, resource_index) -> virtual_address: map a claimed device's MMIO window into
@@ -931,17 +937,21 @@ fn systemDeviceRegister(state: *architecture.CpuState) void {
const parent_id = architecture.systemCallArg(state, 0);
const descriptor_ptr = architecture.systemCallArg(state, 1);
const t = scheduler.current();
if (t.address_space == 0) return fail(state);
if (t.address_space == 0) return failErr(state, ipc.EPERM);
var descriptor: device_abi.DeviceDescriptor = undefined;
if (!ipc.copyFromUser(t.address_space, descriptor_ptr, std.mem.asBytes(&descriptor))) return fail(state);
if (!ipc.copyFromUser(t.address_space, descriptor_ptr, std.mem.asBytes(&descriptor))) return failErr(state, ipc.EFAULT);
// Under the big kernel lock: the broker's table is also mutated by the
// death sweep (releaseAllOwnedBy) and read by enumerate on other cores —
// ring-3 registration (M19) made those genuinely concurrent.
const flags = sync.enter();
defer sync.leave(flags);
const id = devices_broker.register(parent_id, t.id, &descriptor) catch return fail(state);
// Each refusal carries its own errno (devices-broker.errnoOf) — the bus driver
// logs which rule stopped it, so "this parent is full" is never again mistaken
// for "the table is full" or "that resource escapes your window".
const id = devices_broker.register(parent_id, t.id, &descriptor) catch |e|
return failErr(state, devices_broker.errnoOf(e));
architecture.setSystemCallResult(state, id);
}
+103 -8
View File
@@ -51,6 +51,15 @@ fn check(name: []const u8, ok: bool) void {
}
}
/// `devices_broker.claim` reduced to a bool, for the `check` assertions below. The
/// broker returns `ClaimError` so ring 3 can tell "stale id" from "someone already
/// owns it" (the errno space in system/abi.zig); a test that only asserts the claim
/// succeeded does not care which, and the ones that do match the error directly.
fn claimOk(id: u64, owner: u32) bool {
devices_broker.claim(id, owner) catch return false;
return true;
}
/// Emit the overall result line the harness matches, then the done sentinel.
fn result() void {
log("DANOS-TEST-RESULT: {s} ({d} passed, {d} failed)\n", .{
@@ -250,6 +259,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
irqFreeTest();
} else if (eql(case, "containment")) {
containmentTest();
} else if (eql(case, "apertures")) {
apertureTest();
} else if (eql(case, "device-manager")) {
deviceManagerTest(boot_information);
} else if (eql(case, "protocol-registry")) {
@@ -1475,7 +1486,7 @@ fn ioPortTest() void {
check("discovered the acpi-tables I/O window", true);
const me = scheduler.current();
check("claimed the io_port device", devices_broker.claim(id, me.id));
check("claimed the io_port device", claimOk(id, me.id));
check("an in-range access resolves to port 0x64", process.resolveIoPort(me, id, found_res, 0x64, 1) == 0x64);
check("a 4-byte access at the last port is refused", process.resolveIoPort(me, id, found_res, 0xFFFF, 4) == null);
check("an out-of-range offset is refused", process.resolveIoPort(me, id, found_res, 0x10000, 1) == null);
@@ -2470,8 +2481,8 @@ fn claimReleaseTest(boot_information: *const BootInformation) void {
}
// The broker release in isolation.
check("device 0 claimed by owner 111", devices_broker.claim(0, 111));
check("device 1 claimed by owner 222", devices_broker.claim(1, 222));
check("device 0 claimed by owner 111", claimOk(0, 111));
check("device 1 claimed by owner 222", claimOk(1, 222));
devices_broker.releaseAllOwnedBy(111);
check("owner 111's claim is released", devices_broker.ownerOf(0) == null);
check("owner 222's claim survives", (devices_broker.ownerOf(1) orelse 0) == 222);
@@ -2493,7 +2504,7 @@ fn claimReleaseTest(boot_information: *const BootInformation) void {
};
const child = process.spawnProcessSupervised(image, 4, &.{"/system/services/init"}, me, endpoint) catch 0;
check("supervised child spawned", child != 0);
check("device 0 claimed on the child's behalf", devices_broker.claim(0, child));
check("device 0 claimed on the child's behalf", claimOk(0, child));
check("the kill is accepted", process.killProcess(me, child) == 0);
var badge: u64 = 0;
@@ -2501,7 +2512,7 @@ fn claimReleaseTest(boot_information: *const BootInformation) void {
_ = ipcsync.replyWait(endpoint, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap);
check("the exit notification arrived", badge == abi.notify_badge_bit | abi.notify_exit_bit | child);
check("death released the child's claim", devices_broker.ownerOf(0) == null);
check("the device is claimable again", devices_broker.claim(0, me));
check("the device is claimable again", claimOk(0, me));
devices_broker.releaseAllOwnedBy(me);
result();
}
@@ -3765,8 +3776,8 @@ fn killThreadedGroupTest(boot_information: *const BootInformation) void {
// Claims for BOTH members: group death must release every member's claims
// before the supervisor hears anything — the worker's by the deferred
// (condemned) path.
check("device 0 claimed for the leader", devices_broker.claim(0, child));
check("device 1 claimed for the worker", devices_broker.claim(1, worker));
check("device 0 claimed for the leader", claimOk(0, child));
check("device 1 claimed for the worker", claimOk(1, worker));
scheduler.sleep(100); // let the worker really be running on another core
check("the supervisor's kill is accepted", process.killProcess(me, child) == 0);
const badge = awaitExitBadge(endpoint);
@@ -3954,7 +3965,7 @@ fn containmentTest() void {
result();
return;
}
check("claimed the parent device", devices_broker.claim(parent_id, me));
check("claimed the parent device", claimOk(parent_id, me));
defer devices_broker.releaseAllOwnedBy(me);
// A child whose window lies inside the parent's is accepted.
@@ -3976,10 +3987,94 @@ fn containmentTest() void {
check("re-registering an identical child returns the same id", again != 0 and again == good);
check("re-registering grew nothing", devices_broker.enumerate(&buffer) == before + 1);
// Fill the parent to its child cap with distinct children (same window, different
// identity — the match is on identity, so each is a new device).
var filled: u32 = 0;
var capped = false;
while (filled < 64) : (filled += 1) {
var name: [4]u8 = .{ 'k', 0, 0, 0 };
name[1] = '0' + @as(u8, @intCast(filled / 10));
name[2] = '0' + @as(u8, @intCast(filled % 10));
var extra = childDescriptor(name[0..3], parent_window.start, 0x20);
_ = devices_broker.register(parent_id, me, &extra) catch |err| {
capped = err == error.TooManyChildren;
break;
};
}
check("the parent reaches its child cap (TooManyChildren)", capped);
// The regression this ordering exists for: **a re-registration consumes no slot,
// so a full parent must not refuse one.** A crashed bus driver is restarted by its
// supervisor and re-registers everything it rediscovers; when the cap was checked
// before the identity match, the restart was refused its own devices and the
// machine degraded a little more on every crash.
const readmitted = devices_broker.register(parent_id, me, &fits) catch 0;
check("a full parent still re-admits an identical child", readmitted != 0 and readmitted == good);
// ...and the cap is genuinely still in force for anything new.
var novel = childDescriptor("knew", parent_window.start, 0x20);
const still_capped = if (devices_broker.register(parent_id, me, &novel)) |_| false else |err| err == error.TooManyChildren;
check("a full parent still refuses a new child", still_capped);
result();
}
/// A minimal child descriptor with one memory resource, for the containment test.
/// PCI host-bridge apertures are derived from the *holes* in the firmware memory map,
/// and a registered BAR must fall inside one. So the invariant is not "we find the
/// holes" but "an aperture never covers memory the firmware described" — an aperture
/// over RAM means `device_register` containment admits a child BAR over kernel memory,
/// and its claimant can `mmio_map` it.
///
/// The map's length is the firmware's choice: 60–200 descriptors on a real machine,
/// 15–25 under OVMF, which is why the suite never saw this. The derivation used to
/// copy sub-4 GiB entries into a fixed `[64]` array and skip the rest — and a skipped
/// region is not merely lost, it is one the gap finder concludes is *free*.
fn apertureTest() void {
// 100 described one-page regions, 2 MiB apart: 99 small holes between them, then
// one large hole from the last region up to 4 GiB. Under the old fixed array the
// 36 regions past the 64th vanished, so the "largest hole" ran from ~128 MiB to
// 4 GiB — straight across 36 regions the firmware had described.
const spacing: u64 = 2 << 20;
var regions: [100]boot_handoff.MemoryRegion = undefined;
for (&regions, 0..) |*region, i| {
region.* = .{ .base = @as(u64, i) * spacing, .pages = 1, .kind = .usable };
}
var holes: [3]platform.AddressRange = undefined;
const found = platform.largestHolesBelow4G(&regions, 1 << 20, &holes);
check("apertures were derived from a 100-entry map", found > 0);
var overlaps: usize = 0;
var largest: platform.AddressRange = .{ .base = 0, .end = 0 };
for (holes) |hole| {
if (hole.end <= hole.base) continue;
if (hole.end - hole.base > largest.end - largest.base) largest = hole;
for (regions) |region| {
const region_end = region.base + region.pages * 4096;
if (region.base < hole.end and region_end > hole.base) overlaps += 1;
}
}
check("no aperture overlaps described memory", overlaps == 0);
// The big hole is above the last described region, not across it.
const last_end = (regions.len - 1) * spacing + 4096;
check("the largest aperture starts after the last described region", largest.base == last_end);
check("the largest aperture runs to 4 GiB", largest.end == (1 << 32));
// A map the firmware describes nothing in is one whole hole; a map that describes
// everything has none. Neither may invent an aperture over something described.
var empty: [3]platform.AddressRange = undefined;
check("an empty map yields one hole", platform.largestHolesBelow4G(&.{}, 1 << 20, &empty) == 1);
check("that hole is the whole low space", empty[0].base == 0 and empty[0].end == (1 << 32));
const whole = [_]boot_handoff.MemoryRegion{.{ .base = 0, .pages = (1 << 32) / 4096, .kind = .usable }};
var none: [3]platform.AddressRange = undefined;
check("a fully described map yields no aperture", platform.largestHolesBelow4G(&whole, 1 << 20, &none) == 0);
result();
}
fn childDescriptor(hid: []const u8, start: u64, len: u64) device_abi.DeviceDescriptor {
var child = std.mem.zeroes(device_abi.DeviceDescriptor);
child.class = @intFromEnum(device_abi.DeviceClass.unknown);
+8 -2
View File
@@ -4,8 +4,14 @@
//! hiding the trade-offs. Keeping them here makes them visible at a glance and gives
//! one spot to change them. They're plain `comptime` constants (zero runtime cost);
//! any one can later be promoted to a `-D` build option if a target needs to vary it
//! (see build.zig's `-Dtest-case` for the pattern). This keeps ps2-library.zig to what it
//! actually is — the bootloader↔kernel handoff *contract* — with tunables living here.
//! (see build.zig's `-Dtest-case` for the pattern). This keeps [[boot-handoff]] to what
//! it actually is — the loader↔kernel handoff *contract* — with tunables living here.
//!
//! **Kernel only, and deliberately so.** This file exists because tunables were
//! crowding the loader↔kernel contract they were split out of; it is not a registry for
//! the whole system. A driver's ring size belongs to that driver, a protocol's payload
//! cap to that protocol. How a ceiling is *declared*, wherever it lives, is
//! docs/os-development/bounds.md — a shape, not a shared list.
/// Ceiling on logical CPUs the kernel tracks — the size of the per-CPU bookkeeping
/// arrays (discovery pool, scheduler state, per-core GDT/TSS). Generous headroom:
+5 -5
View File
@@ -119,10 +119,10 @@ pub fn main(init: process.Init) void {
return;
};
node_id = node.id;
if (!device.claim(node_id)) {
_ = logging.write("/system/services/acpi: unable to claim acpi-tables\n");
device.claim(node_id) catch |e| {
std.log.warn("unable to claim acpi-tables: {s}", .{@errorName(e)});
return;
}
};
// Map the node's resources: the AML blobs (bytecode), the FADT (intact
// "FACP" header — decision 3), the io_port grant, and the SCI irq.
@@ -519,8 +519,8 @@ fn registerDevice(node: *aml.Node, hid: [8]u8, interpreter: *aml.Interpreter) vo
@memcpy(descriptor.hid[0..@intCast(hid_len)], hid[0..@intCast(hid_len)]);
applyCrs(&descriptor, node, interpreter);
const id = device.register(node_id, &descriptor) orelse {
std.log.info("register refused for {s}", .{hid[0..@intCast(hid_len)]});
const id = device.register(node_id, &descriptor) catch |e| {
std.log.warn("register refused for {s}: {s}", .{ hid[0..@intCast(hid_len)], @errorName(e) });
return;
};
registered[registered_count] = .{ .hid = hid, .hid_len = @intCast(hid_len), .device_id = id, .resource_count = descriptor.resource_count };
+5 -3
View File
@@ -74,10 +74,12 @@ pub const Gop = struct {
return null;
};
if (!device.claim(found.id)) {
_ = logging.write("display: could not claim the framebuffer\n");
device.claim(found.id) catch |e| {
var line: [96]u8 = undefined;
_ = logging.write(std.fmt.bufPrint(&line, "display: could not claim the framebuffer: {s}\n", .{@errorName(e)}) catch
"display: could not claim the framebuffer\n");
return null;
}
};
// Resource 0 is the framebuffer memory window; the kernel maps it write-combining
// because the resource carries that flag (docs/display-plan.md D1).
const front_base = device.mmioMap(found.id, 0) orelse {