kernel: a refusal names its rule, and two bounds stop failing open
An AMD Ryzen booted to a working compositor with no USB and no storage, and the log said only "register refused". A tree-wide audit of every compile-time ceiling followed: 235 of them, 139 on quantities the machine or a file decides rather than us, 5 documented anywhere, 171 silent when reached. docs/fixed-bounds-audit.md has the inventory. Errno attribution. The errno space was split between the kernel and the envelope, free to drift; it is now one list in system/abi.zig, restated on both sides, with a comptime check in library/device/driver where the two halves are visible. device_register's six refusals and device_claim's three are distinct codes, so a bus driver can say which rule stopped it, and BadParent splits into NoSuchParent and NotYourParent. pci-bus reconciles found against registered instead of counting refused functions as found. Idempotency ordering. The child cap was checked before the identity match, so a restarted bus was refused its own devices — the supervision restart the system leans on ratcheted toward a degraded machine. A re-registration consumes no slot and is now admitted first. IOMMU fail-closed. confineDevice returned success for a device id past the confinement table, leaving the device outside every domain while the caller believed it confined — unreachable only while ids stop at 64, which both the inventory move and a hardware-reported domain count would change. It refuses now, and the coupling to the broker's device cap is a comptime assert rather than a sentence in a comment. PCI apertures. The bridge's MMIO apertures are derived from the holes in the firmware memory map, and the derivation copied sub-4 GiB entries into a fixed [64] array and skipped the rest. A skipped region is not merely lost: the gap finder concludes it is free, so a real machine's 60-200 entry map yields an aperture over live RAM, and containment then admits a child BAR covering kernel memory. Rewritten to walk the map in place, with the hole finder extracted as a pure function and driven by a synthetic 100-entry map in a new test case. Both new tests were verified to fail on the old code. parameters.zig gains the rationale it was missing and loses a stale sentence pointing at the wrong file; vdso.md documents the errno space, including EPEER, which had no written meaning anywhere. docs/os-development/bounds.md is how a ceiling is declared from here. docs/bounds-track-plan.md is the plan to remove the ones we invented. Suite 114 -> 115.
This commit is contained in:
+82
-40
@@ -615,17 +615,78 @@ var boot_memory_regions: []const boot_handoff.MemoryRegion = &.{};
|
||||
/// I/O-APIC region, the high one from 4 GiB (or the end of RAM above it) to
|
||||
/// the 46-bit line. Coarse, mechanical, and AML-free — available at boot no
|
||||
/// matter what later moved to user space.
|
||||
pub const AddressRange = struct { base: u64, end: u64 };
|
||||
|
||||
/// The largest holes below 4 GiB in a firmware memory map: the ranges the firmware
|
||||
/// described *nothing* in, which is where a PCI BAR may legitimately live. Fills `out`
|
||||
/// (keeping the largest, replacing the smallest held so far), ignores holes shorter
|
||||
/// than `minimum`, and returns how many entries are non-empty; entries the caller
|
||||
/// reads may be empty and are skipped by `end > base`.
|
||||
///
|
||||
/// **It never copies the map, and that is the point.** The firmware chooses how many
|
||||
/// descriptors its map has — commonly 60–200 on a real machine, 15–25 under OVMF. The
|
||||
/// previous version copied the sub-4 GiB entries into a fixed `[64]` array and
|
||||
/// `continue`d past the rest, which did not merely lose them: a region absent from the
|
||||
/// walk is a region this function concludes is *free*, so a large enough map yields an
|
||||
/// "aperture" lying over live RAM. `device_register` containment would then admit a
|
||||
/// child BAR covering kernel memory, and its claimant could `mmio_map` it. A bound
|
||||
/// whose overflow hands out authority is not a limit; the fix is not a bigger array.
|
||||
///
|
||||
/// Pure, allocation-free, and linear-ish in the map (each inner pass consumes at least
|
||||
/// one region, and it runs once at boot).
|
||||
pub fn largestHolesBelow4G(
|
||||
regions: []const boot_handoff.MemoryRegion,
|
||||
minimum: u64,
|
||||
out: []AddressRange,
|
||||
) usize {
|
||||
const limit: u64 = 1 << 32;
|
||||
for (out) |*hole| hole.* = .{ .base = 0, .end = 0 };
|
||||
|
||||
var cursor: u64 = 0;
|
||||
while (cursor < limit) {
|
||||
// Step over every described region covering the cursor. Regions may overlap
|
||||
// and chain, so repeat until the cursor stops moving.
|
||||
var moved = true;
|
||||
while (moved) {
|
||||
moved = false;
|
||||
for (regions) |region| {
|
||||
if (region.base >= limit) continue;
|
||||
const end = @min(region.base + region.pages * 4096, limit);
|
||||
if (region.base <= cursor and end > cursor) {
|
||||
cursor = end;
|
||||
moved = true;
|
||||
}
|
||||
}
|
||||
}
|
||||
if (cursor >= limit) break;
|
||||
|
||||
// The cursor now sits in a hole; it runs to the next described base, or to
|
||||
// 4 GiB if nothing is described above it.
|
||||
var next: u64 = limit;
|
||||
for (regions) |region| {
|
||||
if (region.base >= limit) continue;
|
||||
if (region.base > cursor and region.base < next) next = region.base;
|
||||
}
|
||||
|
||||
if (next - cursor >= minimum and out.len != 0) {
|
||||
var smallest: usize = 0;
|
||||
for (out, 0..) |hole, i| {
|
||||
if (hole.end - hole.base < out[smallest].end - out[smallest].base) smallest = i;
|
||||
}
|
||||
if (next - cursor > out[smallest].end - out[smallest].base)
|
||||
out[smallest] = .{ .base = cursor, .end = next };
|
||||
}
|
||||
cursor = next;
|
||||
}
|
||||
|
||||
var found: usize = 0;
|
||||
for (out) |hole| {
|
||||
if (hole.end > hole.base) found += 1;
|
||||
}
|
||||
return found;
|
||||
}
|
||||
|
||||
fn addBridgeApertures(bridge: *device_model.Device) void {
|
||||
// Below 4 GiB the described regions are sparse (RAM low, firmware flash
|
||||
// and tables high), so the holes are the *gaps between* them — a single
|
||||
// "after the last region" rule dies on OVMF's flash at the very top.
|
||||
// Sort-merge the described ranges, then keep the three largest gaps
|
||||
// (resource slots are bounded at 8 per device; ECAM + bus range + 3 + the
|
||||
// high aperture fits). Above 4 GiB one aperture runs from the end of the
|
||||
// described space to the 46-bit line.
|
||||
const Range = struct { base: u64, end: u64 };
|
||||
var below: [64]Range = undefined;
|
||||
var below_count: usize = 0;
|
||||
var high_end: u64 = 1 << 32;
|
||||
for (boot_memory_regions) |region| {
|
||||
const end = region.base + region.pages * 4096;
|
||||
@@ -636,37 +697,18 @@ fn addBridgeApertures(bridge: *device_model.Device) void {
|
||||
// kernel image, the tables, the ramdisk all live there). Bring-up
|
||||
// trust: only the bridge's claimant can register into the aperture.
|
||||
if (region.kind == .usable and end > high_end) high_end = end;
|
||||
if (region.base >= (1 << 32) or below_count == below.len) continue;
|
||||
below[below_count] = .{ .base = region.base, .end = @min(end, 1 << 32) };
|
||||
below_count += 1;
|
||||
}
|
||||
// Insertion sort by base (the map is small and this runs once at boot).
|
||||
for (1..below_count) |i| {
|
||||
const key = below[i];
|
||||
var j = i;
|
||||
while (j > 0 and below[j - 1].base > key.base) : (j -= 1) below[j] = below[j - 1];
|
||||
below[j] = key;
|
||||
}
|
||||
// Walk the sorted ranges, collecting inter-region gaps of at least 1 MiB.
|
||||
var gaps: [3]Range = .{Range{ .base = 0, .end = 0 }} ** 3;
|
||||
var cursor: u64 = 0;
|
||||
var index: usize = 0;
|
||||
while (index <= below_count) : (index += 1) {
|
||||
const gap_end = if (index == below_count) (1 << 32) else below[index].base;
|
||||
if (gap_end > cursor and gap_end - cursor >= (1 << 20)) {
|
||||
// Keep the three largest, replacing the smallest kept so far.
|
||||
var smallest: usize = 0;
|
||||
for (gaps, 0..) |gap, gi| {
|
||||
if (gap.end - gap.base < gaps[smallest].end - gaps[smallest].base) smallest = gi;
|
||||
}
|
||||
if (gap_end - cursor > gaps[smallest].end - gaps[smallest].base) {
|
||||
gaps[smallest] = .{ .base = cursor, .end = gap_end };
|
||||
}
|
||||
}
|
||||
if (index < below_count and below[index].end > cursor) cursor = below[index].end;
|
||||
}
|
||||
for (gaps) |gap| {
|
||||
if (gap.end > gap.base) _ = bridge.addResource(.memory, gap.base, gap.end - gap.base);
|
||||
|
||||
// Three holes below 4 GiB. This three is not a guess about hardware: a
|
||||
// `DeviceDescriptor` carries `maximum_device_resources` (8) resources, and the
|
||||
// bridge spends them on ECAM + bus range + these + the high aperture. That cap is
|
||||
// a wire struct in the kernel↔user ABI, so widening it is a separate change —
|
||||
// docs/bounds-track-plan.md, phase 4. Keeping *fewer* holes is fail-closed: it
|
||||
// refuses BARs, it never admits one.
|
||||
var holes: [3]AddressRange = undefined;
|
||||
_ = largestHolesBelow4G(boot_memory_regions, 1 << 20, &holes);
|
||||
for (holes) |hole| {
|
||||
if (hole.end > hole.base) _ = bridge.addResource(.memory, hole.base, hole.end - hole.base);
|
||||
}
|
||||
_ = bridge.addResource(.memory, high_end, (@as(u64, 1) << 46) - high_end);
|
||||
}
|
||||
|
||||
@@ -19,10 +19,11 @@
|
||||
//! ever subdivide what it was already given.
|
||||
|
||||
const std = @import("std");
|
||||
const abi = @import("abi");
|
||||
const platform = @import("platform");
|
||||
const device_abi = @import("device-abi");
|
||||
|
||||
const maximum_devices = 64;
|
||||
pub const maximum_devices = 64;
|
||||
|
||||
/// Cap on children a single parent may have. A zero-resource child (legal — a USB
|
||||
/// device is addressed through its controller, not by MMIO) sidesteps the containment
|
||||
@@ -157,13 +158,13 @@ pub fn enumerateFrom(start: usize, out: []device_abi.DeviceDescriptor) usize {
|
||||
return n;
|
||||
}
|
||||
|
||||
/// Take exclusive ownership of device `id` for task `owner`. Fails if the id is
|
||||
/// out of range or already claimed.
|
||||
pub fn claim(id: u64, owner: u32) bool {
|
||||
if (id >= count) return false;
|
||||
if (claimed[@intCast(id)] != null) return false;
|
||||
/// Take exclusive ownership of device `id` for task `owner`. The two ways this can
|
||||
/// fail want different responses from a driver — a stale id means re-enumerate, a
|
||||
/// live claimant means back off — so they are distinguishable (`ClaimError`).
|
||||
pub fn claim(id: u64, owner: u32) ClaimError!void {
|
||||
if (id >= count) return error.NoSuchDevice;
|
||||
if (claimed[@intCast(id)] != null) return error.AlreadyClaimed;
|
||||
claimed[@intCast(id)] = owner;
|
||||
return true;
|
||||
}
|
||||
|
||||
/// The task that owns device `id`, or null.
|
||||
@@ -274,14 +275,70 @@ fn contains(parent: device_abi.ResourceDescriptor, child: device_abi.ResourceDes
|
||||
return child.start >= parent.start and child_end <= parent_end;
|
||||
}
|
||||
|
||||
/// Why a `register` was refused. Each variant maps to its own errno (`errnoOf`), so
|
||||
/// a bus driver's log line can name the rule that stopped it — "this parent is at
|
||||
/// its child cap" and "the table is full" want different fixes, and telling them
|
||||
/// apart from a bare -1 cost a debugging session (docs/fixed-bounds-audit.md).
|
||||
pub const RegisterError = error{
|
||||
NoSpace, // the device table is full
|
||||
BadParent, // no such device, or not claimed by this task
|
||||
TooManyResources,
|
||||
NoSuchParent, // no device with that id
|
||||
NotYourParent, // that device exists but this task has not claimed it
|
||||
TooManyResources, // the descriptor declares more resources than one device may hold
|
||||
TooManyChildren, // this parent is at maximum_children_per_parent
|
||||
NotContained, // a child resource escapes its parent's window
|
||||
};
|
||||
|
||||
/// The errno a refused `register` returns to ring 3.
|
||||
pub fn errnoOf(e: RegisterError) i64 {
|
||||
return switch (e) {
|
||||
error.NoSpace => abi.ENOSPC,
|
||||
error.NoSuchParent => abi.ENODEV,
|
||||
error.NotYourParent => abi.EPERM,
|
||||
error.TooManyResources => abi.E2BIG,
|
||||
error.TooManyChildren => abi.ECHILDREN,
|
||||
error.NotContained => abi.ERANGE,
|
||||
};
|
||||
}
|
||||
|
||||
/// Why a `claim` was refused.
|
||||
pub const ClaimError = error{
|
||||
NoSuchDevice, // no device with that id
|
||||
AlreadyClaimed, // a live task already owns it
|
||||
};
|
||||
|
||||
/// The errno a refused `claim` returns to ring 3. (`ECONFINE` — the claim stood but
|
||||
/// the IOMMU would not confine the device — is raised by the caller in
|
||||
/// system/kernel/process.zig, which is what rolls the claim back.)
|
||||
pub fn claimErrnoOf(e: ClaimError) i64 {
|
||||
return switch (e) {
|
||||
error.NoSuchDevice => abi.ENODEV,
|
||||
error.AlreadyClaimed => abi.EBUSY,
|
||||
};
|
||||
}
|
||||
|
||||
/// The id of a child of `parent_id` already identical to `descriptor`, or null.
|
||||
/// Exact on class, identity and every resource — anything less would let a bus
|
||||
/// silently adopt an entry that is not the device it just found. The caller must
|
||||
/// have bounded `descriptor.resource_count` first.
|
||||
fn existingChild(parent_id: u64, descriptor: *const device_abi.DeviceDescriptor) ?u64 {
|
||||
for (devices[0..count]) |*existing| {
|
||||
if (existing.parent != parent_id) continue;
|
||||
if (existing.class != descriptor.class) continue;
|
||||
if (existing.pci_class != descriptor.pci_class) continue;
|
||||
if (existing.hid_len != descriptor.hid_len) continue;
|
||||
if (!std.mem.eql(u8, existing.hid[0..@intCast(existing.hid_len)], descriptor.hid[0..@intCast(descriptor.hid_len)])) continue;
|
||||
if (existing.resource_count != descriptor.resource_count) continue;
|
||||
var same = true;
|
||||
for (0..@intCast(descriptor.resource_count)) |i| {
|
||||
const a = existing.resources[i];
|
||||
const b = descriptor.resources[i];
|
||||
if (a.kind != b.kind or a.start != b.start or a.len != b.len) same = false;
|
||||
}
|
||||
if (same) return existing.id;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/// Number of devices currently recorded with `parent_id` as their parent.
|
||||
fn childCount(parent_id: u64) usize {
|
||||
var n: usize = 0;
|
||||
@@ -299,9 +356,27 @@ fn childCount(parent_id: u64) usize {
|
||||
/// contained in a parent resource of the same kind. A device with no resources is
|
||||
/// fine and common: a USB device is addressed through its controller, not by MMIO.
|
||||
pub fn register(parent_id: u64, owner: u32, descriptor: *const device_abi.DeviceDescriptor) RegisterError!u64 {
|
||||
const parent_owner = ownerOf(parent_id) orelse return error.BadParent;
|
||||
if (parent_owner != owner) return error.BadParent;
|
||||
if (parent_id >= count) return error.NoSuchParent;
|
||||
const parent_owner = ownerOf(parent_id) orelse return error.NotYourParent;
|
||||
if (parent_owner != owner) return error.NotYourParent;
|
||||
// Bounds every `descriptor.resources` read below, including the match scan's.
|
||||
if (descriptor.resource_count > device_abi.maximum_device_resources) return error.TooManyResources;
|
||||
|
||||
// Idempotent on exact match (docs/device-manager.md): a restarted registering
|
||||
// bus re-registers what it rediscovers, and the table has no unregister — an
|
||||
// identical child under the same parent returns the existing id instead of
|
||||
// appending a duplicate.
|
||||
//
|
||||
// Checked **before the caps**, because a re-registration consumes no slot.
|
||||
// Charging it against the child cap refused a restarted bus its own devices the
|
||||
// second time it started, which turned the supervision restart this system leans
|
||||
// on into a one-way ratchet toward a degraded machine. The match is exact — class,
|
||||
// identity, and every resource — so an entry returned this way was contained when
|
||||
// it was first admitted, and a stored descriptor's resources never change
|
||||
// afterwards (they are written only by `record`, `seedDisplay` and the append
|
||||
// below).
|
||||
if (existingChild(parent_id, descriptor)) |existing_id| return existing_id;
|
||||
|
||||
if (childCount(parent_id) >= maximum_children_per_parent) return error.TooManyChildren;
|
||||
if (count >= maximum_devices) return error.NoSpace;
|
||||
|
||||
@@ -315,26 +390,6 @@ pub fn register(parent_id: u64, owner: u32, descriptor: *const device_abi.Device
|
||||
if (!ok) return error.NotContained;
|
||||
}
|
||||
|
||||
// Idempotent on exact match (docs/device-manager.md): a restarted
|
||||
// registering bus re-registers what it rediscovers, and the table has no
|
||||
// unregister — an identical (class, identity, resources) child under the
|
||||
// same parent returns the existing id instead of appending a duplicate.
|
||||
for (devices[0..count]) |*existing| {
|
||||
if (existing.parent != parent_id) continue;
|
||||
if (existing.class != descriptor.class) continue;
|
||||
if (existing.pci_class != descriptor.pci_class) continue;
|
||||
if (existing.hid_len != descriptor.hid_len) continue;
|
||||
if (!std.mem.eql(u8, existing.hid[0..@intCast(existing.hid_len)], descriptor.hid[0..@intCast(descriptor.hid_len)])) continue;
|
||||
if (existing.resource_count != descriptor.resource_count) continue;
|
||||
var same = true;
|
||||
for (0..@intCast(descriptor.resource_count)) |i| {
|
||||
const a = existing.resources[i];
|
||||
const b = descriptor.resources[i];
|
||||
if (a.kind != b.kind or a.start != b.start or a.len != b.len) same = false;
|
||||
}
|
||||
if (same) return existing.id;
|
||||
}
|
||||
|
||||
var d = std.mem.zeroes(device_abi.DeviceDescriptor);
|
||||
d.id = count;
|
||||
d.parent = parent_id;
|
||||
|
||||
+26
-5
@@ -96,16 +96,37 @@ pub fn init() void {
|
||||
const Confined = struct { active: bool = false, owner: u32 = 0, bdf: u16 = 0, domain: u16 = invalid_domain };
|
||||
var confined: [maximum_domains]Confined = .{Confined{}} ** maximum_domains;
|
||||
|
||||
// `confined` is indexed by **device id**, so it must cover every id the broker can
|
||||
// mint. These two numbers agreed only by a sentence in a comment above
|
||||
// `maximum_domains` — and when they disagreed, `confineDevice` returned success for
|
||||
// the ids it had no room for, leaving those devices unconfined DMA masters. Coupled
|
||||
// bounds agree in code, not in prose (docs/os-development/bounds.md).
|
||||
comptime {
|
||||
if (maximum_domains < devices_broker.maximum_devices)
|
||||
@compileError("iommu.confined is indexed by device id but is smaller than the " ++
|
||||
"broker's device table: ids past its end cannot be confined, and so cannot " ++
|
||||
"be claimed at all");
|
||||
}
|
||||
|
||||
/// Place a just-claimed PCI function under IOMMU translation on behalf of `owner`: give
|
||||
/// it a private empty domain, seed it with the device's own firmware reserved region,
|
||||
/// and attach. Its DMA buffers arrive afterward as explicit grants — the owner's own
|
||||
/// `dma_alloc`'d regions are bound by the claim path (`mapForDevice`), and cross-process
|
||||
/// buffers by `dma_bind`. false only if a domain can't be allocated — the caller rolls
|
||||
/// the claim back (a claim that can't be confined must not stand). No-op success when no
|
||||
/// IOMMU exists (fail-open).
|
||||
/// buffers by `dma_bind`. false when the device cannot be confined — the caller rolls
|
||||
/// the claim back (a claim that can't be confined must not stand).
|
||||
///
|
||||
/// **Fail-closed at the table's edge.** There is exactly one deliberate fail-open here:
|
||||
/// a machine with no IOMMU, which is a fact about the hardware rather than the size of
|
||||
/// anything. Running out of *room to record* a confinement is not that, and must refuse.
|
||||
pub fn confineDevice(device_id: u64, bdf: u16, owner: u32) bool {
|
||||
if (!active) return true;
|
||||
if (device_id >= confined.len) return true; // unusual id; leave it to fail-open
|
||||
if (!active) return true; // no IOMMU on this machine — nothing to confine with
|
||||
// A device id past the end of the record table. This returned `true` — success —
|
||||
// leaving the device outside every domain while telling the caller it was
|
||||
// confined, and rolling nothing back. It is unreachable only while device ids stop
|
||||
// at `confined.len`; moving the inventory out of the kernel and taking the domain
|
||||
// count from the hardware both change that, and either would have made a silent
|
||||
// unconfined DMA master out of every device past the 64th.
|
||||
if (device_id >= confined.len) return false;
|
||||
const domain = domainCreate(owner, bdf) orelse return false;
|
||||
|
||||
// Firmware reserved region for this device, if any (real hardware; QEMU has none).
|
||||
|
||||
@@ -42,15 +42,22 @@ pub const MESSAGE_MAXIMUM: usize = 256;
|
||||
pub const maximum_handles = scheduler.ipc_maximum_handles;
|
||||
|
||||
/// Errno-style failures, returned as `-value` in the system_call result register.
|
||||
pub const EBADF: i64 = 1; // bad handle
|
||||
pub const E2BIG: i64 = 2; // message exceeds MESSAGE_MAXIMUM
|
||||
pub const EFAULT: i64 = 3; // buffer unmapped / out of the user half
|
||||
pub const ENOENT: i64 = 4; // no such name
|
||||
pub const ENOSPC: i64 = 5; // handle table full
|
||||
pub const ENOMEM: i64 = 6; // out of memory
|
||||
pub const EPEER: i64 = 7; // peer died before replying (its process exited or was killed)
|
||||
pub const ESRCH: i64 = 8; // no such process (process_kill of an unknown/dead id)
|
||||
pub const EPERM: i64 = 9; // not permitted (process_kill by anyone but the supervisor)
|
||||
/// Restated from the shared kernel↔user ABI (system/abi.zig), because ring 3 reads
|
||||
/// the same numbers — the same reason `notify_badge_bit` below is restated. New
|
||||
/// codes are added *there*, which is where the space is documented.
|
||||
pub const EBADF = abi.EBADF; // bad handle
|
||||
pub const E2BIG = abi.E2BIG; // message exceeds MESSAGE_MAXIMUM
|
||||
pub const EFAULT = abi.EFAULT; // buffer unmapped / out of the user half
|
||||
pub const ENOENT = abi.ENOENT; // no such name
|
||||
pub const ENOSPC = abi.ENOSPC; // a kernel table is full
|
||||
pub const ENOMEM = abi.ENOMEM; // out of memory
|
||||
pub const EPEER = abi.EPEER; // peer died before replying (its process exited or was killed)
|
||||
pub const ESRCH = abi.ESRCH; // no such process (process_kill of an unknown/dead id)
|
||||
pub const EPERM = abi.EPERM; // not permitted (process_kill by anyone but the supervisor)
|
||||
pub const ENODEV = abi.ENODEV; // no such device id
|
||||
pub const ECHILDREN = abi.ECHILDREN; // this parent is at its child cap
|
||||
pub const ERANGE = abi.ERANGE; // a resource escapes its parent's window
|
||||
pub const ECONFINE = abi.ECONFINE; // the device could not be placed under IOMMU translation
|
||||
|
||||
/// A badge with this bit set is an asynchronous notification (e.g. an IRQ), not a
|
||||
/// message from a client — there is no reply owed. The low bits carry the source
|
||||
|
||||
@@ -26,6 +26,14 @@ pub const RegisterAccess = acpi.RegisterAccess;
|
||||
pub const IsoEntry = acpi.IsoEntry;
|
||||
pub const Cpu = acpi.Cpu;
|
||||
|
||||
/// Where a PCI BAR may legitimately live: the holes in the firmware memory map.
|
||||
/// Firmware-agnostic — it takes a boot-handoff map, not an ACPI table — and pure, so
|
||||
/// the kernel self-test can drive it with a synthetic map. That is the only way to
|
||||
/// check the invariant that matters here: an aperture must never cover memory the
|
||||
/// firmware described, because containment would then admit a BAR over live RAM.
|
||||
pub const AddressRange = acpi.AddressRange;
|
||||
pub const largestHolesBelow4G = acpi.largestHolesBelow4G;
|
||||
|
||||
/// The FADT power register map discovery extracted (PM1 control, reset register),
|
||||
/// for kernel reboot and diagnostics. Sleep-state values are userspace's (S5 is
|
||||
/// owned by the ring-3 acpi service), so they are not here.
|
||||
|
||||
+38
-28
@@ -402,36 +402,42 @@ fn systemDeviceEnumerate(state: *architecture.CpuState) void {
|
||||
architecture.setSystemCallResult(state, devices_broker.deviceCount());
|
||||
}
|
||||
|
||||
/// device_claim(id) -> 0/-1: take exclusive ownership of a device for this process.
|
||||
/// device_claim(id) -> 0/-errno: take exclusive ownership of a device for this process.
|
||||
/// `-ENODEV` no such id, `-EBUSY` a live task already owns it, `-ECONFINE` the claim
|
||||
/// could not be placed under IOMMU translation and was rolled back.
|
||||
fn systemDeviceClaim(state: *architecture.CpuState) void {
|
||||
const device_id = architecture.systemCallArg(state, 0);
|
||||
const claim_flags = sync.enter();
|
||||
defer sync.leave(claim_flags);
|
||||
if (devices_broker.claim(device_id, scheduler.current().id)) {
|
||||
// Confine the device's DMA before the driver can program it: a PCI function
|
||||
// becomes reachable to the IOMMU only once claimed (until now its DMA is
|
||||
// blocked). A claim that cannot be confined must not stand — roll it back —
|
||||
// since the whole point is that claiming a DMA device is no longer equivalent
|
||||
// to ring 0. No-op when no IOMMU exists (fail-open).
|
||||
if (devices_broker.pciAddressOf(device_id)) |bdf| {
|
||||
const owner = scheduler.current().id;
|
||||
if (!iommu.confineDevice(device_id, bdf, owner)) {
|
||||
_ = devices_broker.unclaim(device_id, owner);
|
||||
return fail(state);
|
||||
}
|
||||
// Bind the buffers this task allocated before claiming the device (a driver
|
||||
// that dma_alloc'd its rings, then claimed the controller).
|
||||
dmaBindOwnerRegionsInto(owner, device_id);
|
||||
devices_broker.claim(device_id, scheduler.current().id) catch |e|
|
||||
return failErr(state, devices_broker.claimErrnoOf(e));
|
||||
|
||||
// Confine the device's DMA before the driver can program it: a PCI function
|
||||
// becomes reachable to the IOMMU only once claimed (until now its DMA is
|
||||
// blocked). A claim that cannot be confined must not stand — roll it back —
|
||||
// since the whole point is that claiming a DMA device is no longer equivalent
|
||||
// to ring 0. No-op when no IOMMU exists (fail-open).
|
||||
if (devices_broker.pciAddressOf(device_id)) |bdf| {
|
||||
const owner = scheduler.current().id;
|
||||
if (!iommu.confineDevice(device_id, bdf, owner)) {
|
||||
_ = devices_broker.unclaim(device_id, owner);
|
||||
// Its own errno: "nobody could confine this" is a different world from
|
||||
// "someone else already has it", and a driver that cannot tell them apart
|
||||
// cannot report the one that means the machine's DMA protection ran out.
|
||||
return failErr(state, ipc.ECONFINE);
|
||||
}
|
||||
// A display service just took the framebuffer — quiesce the bootstrap console
|
||||
// so the kernel and the service don't scribble over each other's pixels. The
|
||||
// claim releases (and the console resumes) automatically if the service dies;
|
||||
// see releaseTaskResourcesLocked.
|
||||
if (devices_broker.displayDevice()) |display_id| {
|
||||
if (device_id == display_id) console.setSuppressed(true);
|
||||
}
|
||||
architecture.setSystemCallResult(state, 0);
|
||||
} else fail(state);
|
||||
// Bind the buffers this task allocated before claiming the device (a driver
|
||||
// that dma_alloc'd its rings, then claimed the controller).
|
||||
dmaBindOwnerRegionsInto(owner, device_id);
|
||||
}
|
||||
// A display service just took the framebuffer — quiesce the bootstrap console
|
||||
// so the kernel and the service don't scribble over each other's pixels. The
|
||||
// claim releases (and the console resumes) automatically if the service dies;
|
||||
// see releaseTaskResourcesLocked.
|
||||
if (devices_broker.displayDevice()) |display_id| {
|
||||
if (device_id == display_id) console.setSuppressed(true);
|
||||
}
|
||||
architecture.setSystemCallResult(state, 0);
|
||||
}
|
||||
|
||||
/// mmio_map(device_id, resource_index) -> virtual_address: map a claimed device's MMIO window into
|
||||
@@ -931,17 +937,21 @@ fn systemDeviceRegister(state: *architecture.CpuState) void {
|
||||
const parent_id = architecture.systemCallArg(state, 0);
|
||||
const descriptor_ptr = architecture.systemCallArg(state, 1);
|
||||
const t = scheduler.current();
|
||||
if (t.address_space == 0) return fail(state);
|
||||
if (t.address_space == 0) return failErr(state, ipc.EPERM);
|
||||
|
||||
var descriptor: device_abi.DeviceDescriptor = undefined;
|
||||
if (!ipc.copyFromUser(t.address_space, descriptor_ptr, std.mem.asBytes(&descriptor))) return fail(state);
|
||||
if (!ipc.copyFromUser(t.address_space, descriptor_ptr, std.mem.asBytes(&descriptor))) return failErr(state, ipc.EFAULT);
|
||||
|
||||
// Under the big kernel lock: the broker's table is also mutated by the
|
||||
// death sweep (releaseAllOwnedBy) and read by enumerate on other cores —
|
||||
// ring-3 registration (M19) made those genuinely concurrent.
|
||||
const flags = sync.enter();
|
||||
defer sync.leave(flags);
|
||||
const id = devices_broker.register(parent_id, t.id, &descriptor) catch return fail(state);
|
||||
// Each refusal carries its own errno (devices-broker.errnoOf) — the bus driver
|
||||
// logs which rule stopped it, so "this parent is full" is never again mistaken
|
||||
// for "the table is full" or "that resource escapes your window".
|
||||
const id = devices_broker.register(parent_id, t.id, &descriptor) catch |e|
|
||||
return failErr(state, devices_broker.errnoOf(e));
|
||||
architecture.setSystemCallResult(state, id);
|
||||
}
|
||||
|
||||
|
||||
+103
-8
@@ -51,6 +51,15 @@ fn check(name: []const u8, ok: bool) void {
|
||||
}
|
||||
}
|
||||
|
||||
/// `devices_broker.claim` reduced to a bool, for the `check` assertions below. The
|
||||
/// broker returns `ClaimError` so ring 3 can tell "stale id" from "someone already
|
||||
/// owns it" (the errno space in system/abi.zig); a test that only asserts the claim
|
||||
/// succeeded does not care which, and the ones that do match the error directly.
|
||||
fn claimOk(id: u64, owner: u32) bool {
|
||||
devices_broker.claim(id, owner) catch return false;
|
||||
return true;
|
||||
}
|
||||
|
||||
/// Emit the overall result line the harness matches, then the done sentinel.
|
||||
fn result() void {
|
||||
log("DANOS-TEST-RESULT: {s} ({d} passed, {d} failed)\n", .{
|
||||
@@ -250,6 +259,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
|
||||
irqFreeTest();
|
||||
} else if (eql(case, "containment")) {
|
||||
containmentTest();
|
||||
} else if (eql(case, "apertures")) {
|
||||
apertureTest();
|
||||
} else if (eql(case, "device-manager")) {
|
||||
deviceManagerTest(boot_information);
|
||||
} else if (eql(case, "protocol-registry")) {
|
||||
@@ -1475,7 +1486,7 @@ fn ioPortTest() void {
|
||||
check("discovered the acpi-tables I/O window", true);
|
||||
|
||||
const me = scheduler.current();
|
||||
check("claimed the io_port device", devices_broker.claim(id, me.id));
|
||||
check("claimed the io_port device", claimOk(id, me.id));
|
||||
check("an in-range access resolves to port 0x64", process.resolveIoPort(me, id, found_res, 0x64, 1) == 0x64);
|
||||
check("a 4-byte access at the last port is refused", process.resolveIoPort(me, id, found_res, 0xFFFF, 4) == null);
|
||||
check("an out-of-range offset is refused", process.resolveIoPort(me, id, found_res, 0x10000, 1) == null);
|
||||
@@ -2470,8 +2481,8 @@ fn claimReleaseTest(boot_information: *const BootInformation) void {
|
||||
}
|
||||
|
||||
// The broker release in isolation.
|
||||
check("device 0 claimed by owner 111", devices_broker.claim(0, 111));
|
||||
check("device 1 claimed by owner 222", devices_broker.claim(1, 222));
|
||||
check("device 0 claimed by owner 111", claimOk(0, 111));
|
||||
check("device 1 claimed by owner 222", claimOk(1, 222));
|
||||
devices_broker.releaseAllOwnedBy(111);
|
||||
check("owner 111's claim is released", devices_broker.ownerOf(0) == null);
|
||||
check("owner 222's claim survives", (devices_broker.ownerOf(1) orelse 0) == 222);
|
||||
@@ -2493,7 +2504,7 @@ fn claimReleaseTest(boot_information: *const BootInformation) void {
|
||||
};
|
||||
const child = process.spawnProcessSupervised(image, 4, &.{"/system/services/init"}, me, endpoint) catch 0;
|
||||
check("supervised child spawned", child != 0);
|
||||
check("device 0 claimed on the child's behalf", devices_broker.claim(0, child));
|
||||
check("device 0 claimed on the child's behalf", claimOk(0, child));
|
||||
|
||||
check("the kill is accepted", process.killProcess(me, child) == 0);
|
||||
var badge: u64 = 0;
|
||||
@@ -2501,7 +2512,7 @@ fn claimReleaseTest(boot_information: *const BootInformation) void {
|
||||
_ = ipcsync.replyWait(endpoint, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap);
|
||||
check("the exit notification arrived", badge == abi.notify_badge_bit | abi.notify_exit_bit | child);
|
||||
check("death released the child's claim", devices_broker.ownerOf(0) == null);
|
||||
check("the device is claimable again", devices_broker.claim(0, me));
|
||||
check("the device is claimable again", claimOk(0, me));
|
||||
devices_broker.releaseAllOwnedBy(me);
|
||||
result();
|
||||
}
|
||||
@@ -3765,8 +3776,8 @@ fn killThreadedGroupTest(boot_information: *const BootInformation) void {
|
||||
// Claims for BOTH members: group death must release every member's claims
|
||||
// before the supervisor hears anything — the worker's by the deferred
|
||||
// (condemned) path.
|
||||
check("device 0 claimed for the leader", devices_broker.claim(0, child));
|
||||
check("device 1 claimed for the worker", devices_broker.claim(1, worker));
|
||||
check("device 0 claimed for the leader", claimOk(0, child));
|
||||
check("device 1 claimed for the worker", claimOk(1, worker));
|
||||
scheduler.sleep(100); // let the worker really be running on another core
|
||||
check("the supervisor's kill is accepted", process.killProcess(me, child) == 0);
|
||||
const badge = awaitExitBadge(endpoint);
|
||||
@@ -3954,7 +3965,7 @@ fn containmentTest() void {
|
||||
result();
|
||||
return;
|
||||
}
|
||||
check("claimed the parent device", devices_broker.claim(parent_id, me));
|
||||
check("claimed the parent device", claimOk(parent_id, me));
|
||||
defer devices_broker.releaseAllOwnedBy(me);
|
||||
|
||||
// A child whose window lies inside the parent's is accepted.
|
||||
@@ -3976,10 +3987,94 @@ fn containmentTest() void {
|
||||
check("re-registering an identical child returns the same id", again != 0 and again == good);
|
||||
check("re-registering grew nothing", devices_broker.enumerate(&buffer) == before + 1);
|
||||
|
||||
// Fill the parent to its child cap with distinct children (same window, different
|
||||
// identity — the match is on identity, so each is a new device).
|
||||
var filled: u32 = 0;
|
||||
var capped = false;
|
||||
while (filled < 64) : (filled += 1) {
|
||||
var name: [4]u8 = .{ 'k', 0, 0, 0 };
|
||||
name[1] = '0' + @as(u8, @intCast(filled / 10));
|
||||
name[2] = '0' + @as(u8, @intCast(filled % 10));
|
||||
var extra = childDescriptor(name[0..3], parent_window.start, 0x20);
|
||||
_ = devices_broker.register(parent_id, me, &extra) catch |err| {
|
||||
capped = err == error.TooManyChildren;
|
||||
break;
|
||||
};
|
||||
}
|
||||
check("the parent reaches its child cap (TooManyChildren)", capped);
|
||||
|
||||
// The regression this ordering exists for: **a re-registration consumes no slot,
|
||||
// so a full parent must not refuse one.** A crashed bus driver is restarted by its
|
||||
// supervisor and re-registers everything it rediscovers; when the cap was checked
|
||||
// before the identity match, the restart was refused its own devices and the
|
||||
// machine degraded a little more on every crash.
|
||||
const readmitted = devices_broker.register(parent_id, me, &fits) catch 0;
|
||||
check("a full parent still re-admits an identical child", readmitted != 0 and readmitted == good);
|
||||
|
||||
// ...and the cap is genuinely still in force for anything new.
|
||||
var novel = childDescriptor("knew", parent_window.start, 0x20);
|
||||
const still_capped = if (devices_broker.register(parent_id, me, &novel)) |_| false else |err| err == error.TooManyChildren;
|
||||
check("a full parent still refuses a new child", still_capped);
|
||||
|
||||
result();
|
||||
}
|
||||
|
||||
/// A minimal child descriptor with one memory resource, for the containment test.
|
||||
/// PCI host-bridge apertures are derived from the *holes* in the firmware memory map,
|
||||
/// and a registered BAR must fall inside one. So the invariant is not "we find the
|
||||
/// holes" but "an aperture never covers memory the firmware described" — an aperture
|
||||
/// over RAM means `device_register` containment admits a child BAR over kernel memory,
|
||||
/// and its claimant can `mmio_map` it.
|
||||
///
|
||||
/// The map's length is the firmware's choice: 60–200 descriptors on a real machine,
|
||||
/// 15–25 under OVMF, which is why the suite never saw this. The derivation used to
|
||||
/// copy sub-4 GiB entries into a fixed `[64]` array and skip the rest — and a skipped
|
||||
/// region is not merely lost, it is one the gap finder concludes is *free*.
|
||||
fn apertureTest() void {
|
||||
// 100 described one-page regions, 2 MiB apart: 99 small holes between them, then
|
||||
// one large hole from the last region up to 4 GiB. Under the old fixed array the
|
||||
// 36 regions past the 64th vanished, so the "largest hole" ran from ~128 MiB to
|
||||
// 4 GiB — straight across 36 regions the firmware had described.
|
||||
const spacing: u64 = 2 << 20;
|
||||
var regions: [100]boot_handoff.MemoryRegion = undefined;
|
||||
for (®ions, 0..) |*region, i| {
|
||||
region.* = .{ .base = @as(u64, i) * spacing, .pages = 1, .kind = .usable };
|
||||
}
|
||||
|
||||
var holes: [3]platform.AddressRange = undefined;
|
||||
const found = platform.largestHolesBelow4G(®ions, 1 << 20, &holes);
|
||||
check("apertures were derived from a 100-entry map", found > 0);
|
||||
|
||||
var overlaps: usize = 0;
|
||||
var largest: platform.AddressRange = .{ .base = 0, .end = 0 };
|
||||
for (holes) |hole| {
|
||||
if (hole.end <= hole.base) continue;
|
||||
if (hole.end - hole.base > largest.end - largest.base) largest = hole;
|
||||
for (regions) |region| {
|
||||
const region_end = region.base + region.pages * 4096;
|
||||
if (region.base < hole.end and region_end > hole.base) overlaps += 1;
|
||||
}
|
||||
}
|
||||
check("no aperture overlaps described memory", overlaps == 0);
|
||||
|
||||
// The big hole is above the last described region, not across it.
|
||||
const last_end = (regions.len - 1) * spacing + 4096;
|
||||
check("the largest aperture starts after the last described region", largest.base == last_end);
|
||||
check("the largest aperture runs to 4 GiB", largest.end == (1 << 32));
|
||||
|
||||
// A map the firmware describes nothing in is one whole hole; a map that describes
|
||||
// everything has none. Neither may invent an aperture over something described.
|
||||
var empty: [3]platform.AddressRange = undefined;
|
||||
check("an empty map yields one hole", platform.largestHolesBelow4G(&.{}, 1 << 20, &empty) == 1);
|
||||
check("that hole is the whole low space", empty[0].base == 0 and empty[0].end == (1 << 32));
|
||||
|
||||
const whole = [_]boot_handoff.MemoryRegion{.{ .base = 0, .pages = (1 << 32) / 4096, .kind = .usable }};
|
||||
var none: [3]platform.AddressRange = undefined;
|
||||
check("a fully described map yields no aperture", platform.largestHolesBelow4G(&whole, 1 << 20, &none) == 0);
|
||||
|
||||
result();
|
||||
}
|
||||
|
||||
fn childDescriptor(hid: []const u8, start: u64, len: u64) device_abi.DeviceDescriptor {
|
||||
var child = std.mem.zeroes(device_abi.DeviceDescriptor);
|
||||
child.class = @intFromEnum(device_abi.DeviceClass.unknown);
|
||||
|
||||
Reference in New Issue
Block a user