kernel: a refusal names its rule, and two bounds stop failing open
An AMD Ryzen booted to a working compositor with no USB and no storage, and the log said only "register refused". A tree-wide audit of every compile-time ceiling followed: 235 of them, 139 on quantities the machine or a file decides rather than us, 5 documented anywhere, 171 silent when reached. docs/fixed-bounds-audit.md has the inventory. Errno attribution. The errno space was split between the kernel and the envelope, free to drift; it is now one list in system/abi.zig, restated on both sides, with a comptime check in library/device/driver where the two halves are visible. device_register's six refusals and device_claim's three are distinct codes, so a bus driver can say which rule stopped it, and BadParent splits into NoSuchParent and NotYourParent. pci-bus reconciles found against registered instead of counting refused functions as found. Idempotency ordering. The child cap was checked before the identity match, so a restarted bus was refused its own devices — the supervision restart the system leans on ratcheted toward a degraded machine. A re-registration consumes no slot and is now admitted first. IOMMU fail-closed. confineDevice returned success for a device id past the confinement table, leaving the device outside every domain while the caller believed it confined — unreachable only while ids stop at 64, which both the inventory move and a hardware-reported domain count would change. It refuses now, and the coupling to the broker's device cap is a comptime assert rather than a sentence in a comment. PCI apertures. The bridge's MMIO apertures are derived from the holes in the firmware memory map, and the derivation copied sub-4 GiB entries into a fixed [64] array and skipped the rest. A skipped region is not merely lost: the gap finder concludes it is free, so a real machine's 60-200 entry map yields an aperture over live RAM, and containment then admits a child BAR covering kernel memory. Rewritten to walk the map in place, with the hole finder extracted as a pure function and driven by a synthetic 100-entry map in a new test case. Both new tests were verified to fail on the old code. parameters.zig gains the rationale it was missing and loses a stale sentence pointing at the wrong file; vdso.md documents the errno space, including EPEER, which had no written meaning anywhere. docs/os-development/bounds.md is how a ceiling is declared from here. docs/bounds-track-plan.md is the plan to remove the ones we invented. Suite 114 -> 115.
This commit is contained in:
+26
-5
@@ -96,16 +96,37 @@ pub fn init() void {
|
||||
const Confined = struct { active: bool = false, owner: u32 = 0, bdf: u16 = 0, domain: u16 = invalid_domain };
|
||||
var confined: [maximum_domains]Confined = .{Confined{}} ** maximum_domains;
|
||||
|
||||
// `confined` is indexed by **device id**, so it must cover every id the broker can
|
||||
// mint. These two numbers agreed only by a sentence in a comment above
|
||||
// `maximum_domains` — and when they disagreed, `confineDevice` returned success for
|
||||
// the ids it had no room for, leaving those devices unconfined DMA masters. Coupled
|
||||
// bounds agree in code, not in prose (docs/os-development/bounds.md).
|
||||
comptime {
|
||||
if (maximum_domains < devices_broker.maximum_devices)
|
||||
@compileError("iommu.confined is indexed by device id but is smaller than the " ++
|
||||
"broker's device table: ids past its end cannot be confined, and so cannot " ++
|
||||
"be claimed at all");
|
||||
}
|
||||
|
||||
/// Place a just-claimed PCI function under IOMMU translation on behalf of `owner`: give
|
||||
/// it a private empty domain, seed it with the device's own firmware reserved region,
|
||||
/// and attach. Its DMA buffers arrive afterward as explicit grants — the owner's own
|
||||
/// `dma_alloc`'d regions are bound by the claim path (`mapForDevice`), and cross-process
|
||||
/// buffers by `dma_bind`. false only if a domain can't be allocated — the caller rolls
|
||||
/// the claim back (a claim that can't be confined must not stand). No-op success when no
|
||||
/// IOMMU exists (fail-open).
|
||||
/// buffers by `dma_bind`. false when the device cannot be confined — the caller rolls
|
||||
/// the claim back (a claim that can't be confined must not stand).
|
||||
///
|
||||
/// **Fail-closed at the table's edge.** There is exactly one deliberate fail-open here:
|
||||
/// a machine with no IOMMU, which is a fact about the hardware rather than the size of
|
||||
/// anything. Running out of *room to record* a confinement is not that, and must refuse.
|
||||
pub fn confineDevice(device_id: u64, bdf: u16, owner: u32) bool {
|
||||
if (!active) return true;
|
||||
if (device_id >= confined.len) return true; // unusual id; leave it to fail-open
|
||||
if (!active) return true; // no IOMMU on this machine — nothing to confine with
|
||||
// A device id past the end of the record table. This returned `true` — success —
|
||||
// leaving the device outside every domain while telling the caller it was
|
||||
// confined, and rolling nothing back. It is unreachable only while device ids stop
|
||||
// at `confined.len`; moving the inventory out of the kernel and taking the domain
|
||||
// count from the hardware both change that, and either would have made a silent
|
||||
// unconfined DMA master out of every device past the 64th.
|
||||
if (device_id >= confined.len) return false;
|
||||
const domain = domainCreate(owner, bdf) orelse return false;
|
||||
|
||||
// Firmware reserved region for this device, if any (real hardware; QEMU has none).
|
||||
|
||||
Reference in New Issue
Block a user