From 4398eb7cc43d3269d5c624598448cc7288bc41df Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 11:13:20 +0100 Subject: [PATCH 01/36] docs: scope the bounds track's unattended run MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Six steps an agent can execute: reclamation, the bounds build check, and the four user-space USB/xHCI bounds where the hardware already reports the number we guessed. Deliberately excluded: the authorisation gate and moving the device inventory to the manager. Both decide whether the OS is secure and both are a direction rather than a specification — what a device capability is, which syscalls change, what replaces device_claim for its seven callers. They want a design session, not an agent. Three open questions are written down rather than guessed: device_enumerate most likely narrows to the firmware-discovered roots rather than retiring (the manager cannot ask itself for the PCI host bridge); a manager restart has no re-enumerate handshake, so it comes back blind while its buses live; and "add adversarial tests" is not executable until the attacks are named. --- docs/bounds-track-plan.md | 68 +++++++++++++++++++++++++++++++++++++++ 1 file changed, 68 insertions(+) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index a8a5017..8fd5754 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -3,6 +3,74 @@ *Plan, 2026-08-08. Follows [fixed-bounds-audit.md](fixed-bounds-audit.md) (235 ceilings, 139 on quantities we do not choose) and the AMD Ryzen that found the first one.* +--- + +## Live state — the unattended run + +*This table is the progress view. It is updated at the end of every step, before the +next one starts.* + +| Step | What | State | +|---|---|---| +| L1 | Reclamation: a dead task's registrations die with its claims | not started | +| L2 | Bounds build check + allowlist; declare what we have already touched | not started | +| L3 | xHCI: slot count from `HCSPARAMS1.MaxSlots`, not 8 | not started | +| L4 | USB: configuration descriptor sized by `wTotalLength`, not 512 | not started | +| L5 | USB: interfaces from the descriptor, and the misattributed-endpoint bug | not started | +| L6 | xHCI: a failed `allocateDevice` stops leaking an enabled slot | not started | + +**Suite:** 115/115 at the start of the run. +**Branch:** `claude/bounds-track`. + +### What this run deliberately does not touch + +Phases 2 and 3 below — the authorisation gate and moving the inventory to the device +manager — are **out of scope for unattended work**. They decide whether the OS is +secure, and they are currently a direction rather than a specification: what a device +capability *is*, which syscalls change, what replaces `device_claim` for its seven +callers, how a driver spawned bare behaves. Those want a design session, the way +`/protocol` had one. + +Also out of scope: anything touching `maximum_device_resources` (a wire struct, so a +trust-boundary change, not a resize), and the non-device bounds the audit found in FAT, +the VFS, logger, init, display and boot. + +### Open questions this run must not answer on its own + +Recorded here rather than guessed. If a step runs into one, it stops and writes the +question down instead of inventing an answer. + +1. **`device_enumerate` probably narrows rather than retires.** The device manager + calls it to find `pci_host_bridge` nodes — it cannot ask itself. The likely shape is + that the kernel keeps the *firmware-discovered roots* (which by principle 5 it holds + for real reasons, since they come from ACPI rather than a driver's say-so) and + everything a driver registered lives in the manager. Not decided. +2. **A device-manager restart has no re-enumerate handshake.** If only the manager + dies, the buses are alive and never re-send `child_added`, so a restarted manager + comes back blind. The manager is restartable by design; nothing implements this. +3. **Which adversarial tests I1–I3 need.** The audit's six real defects were all found + by asking what an attacker would do, and the suite had never asked. "Add adversarial + cases" is not executable until the attacks are named. + +### Working rules for the run + +- Work in `/Users/danielsamson/Gitea/daniel/danos` (not a worktree), on + `claude/bounds-track`. +- **Every step lands with a test that fails before the fix**, verified by temporarily + restoring the old behaviour and watching exactly the intended assertion flip. A test + that passes both ways is not a test. +- Full QEMU suite green before each commit. Never run two suites at once — check + `pgrep -f qemu_test.py` first; a second concurrent run produces false triple faults + because both share `zig-out`. +- Check at least 60 GiB free before starting a suite. +- Commit with `git commit -F `, never `-m` (a backtick in a message is executed + by the shell and silently eats a word). No `Co-Authored-By` trailers. +- Update the Live state table **before** starting the next step. +- If a step needs a decision that is not written down here, stop, add it to the open + questions above, and move to the next step. + +--- + ## The principles this is derived from 1. **danOS is a microkernel.** Minimise what the kernel is responsible for; move From 568823a4fb52193c498c0a65f1754d23e74910a5 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 11:15:11 +0100 Subject: [PATCH 02/36] =?UTF-8?q?docs:=20L1=20was=20wrong=20=E2=80=94=20re?= =?UTF-8?q?clamation=20is=20not=20a=20death-sweep=20problem?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The first step of the unattended run was "a dead task's registrations die with its claims". Implementing it would have broken the restart path it was meant to protect. The broker keeps entries on purpose: they describe hardware, which did not go away when a driver died. And device ids must stay stable across a bus restart, because device-manager dedupes re-reports by device_id so a restarted bus does not spawn a second driver instance — stability that comes from the idempotency scan returning the existing id. Removing entries on death would hand a restarted bus fresh ids and duplicate every driver. Everything else a task holds is already reclaimed on every path out: IRQ bindings, IOMMU domains, DMA regions, then its claims. The leak the audit found is real but has two other sources: a device that genuinely goes away has no retirement path, and a bus that enumerates differently on restart strands its old entries. Both belong to the device manager's inventory, and both need the id-stability question settled first — tombstone-and-reuse aliases ids another process still holds, generation tagging changes the id encoding, which is ABI. Recorded as an open question rather than guessed at. --- docs/bounds-track-plan.md | 24 +++++++++++++++++++++++- 1 file changed, 23 insertions(+), 1 deletion(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 8fd5754..de4394c 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -12,7 +12,7 @@ next one starts.* | Step | What | State | |---|---|---| -| L1 | Reclamation: a dead task's registrations die with its claims | not started | +| L1 | Reclamation: a dead task's registrations die with its claims | **stopped — the step was wrong; see open question 4** | | L2 | Bounds build check + allowlist; declare what we have already touched | not started | | L3 | xHCI: slot count from `HCSPARAMS1.MaxSlots`, not 8 | not started | | L4 | USB: configuration descriptor sized by `wTotalLength`, not 512 | not started | @@ -51,6 +51,28 @@ question down instead of inventing an answer. 3. **Which adversarial tests I1–I3 need.** The audit's six real defects were all found by asking what an attacker would do, and the suite had never asked. "Add adversarial cases" is not executable until the attacks are named. +4. **Reclamation is not a death-sweep problem, and L1 as written would have broken the + restart path.** Found on the first attempt at it. The audit is right that `count` + never decreases, but *death is the wrong trigger*: + - The broker keeps entries deliberately: "The devices stay in the table — they + describe hardware, which did not go away — only their ownership clears." A driver + dying does not unplug anything. + - Device ids must stay **stable across a bus restart**, because + `device-manager.driverForDevice` dedupes by `device_id` so that "a re-report after + a bus restart must not spawn a second instance". Stability comes from the + idempotency scan returning the existing id — removing entries on death would give + a restarted bus fresh ids and spawn duplicate driver instances. + - Everything else a task holds *is* already reclaimed on every path out: + `irq.releaseOwner`, `iommu.releaseAllOwnedBy`, `dmaRegistryReleaseOwner`, then the + broker's claims (`process.releaseTaskResourcesLocked`). + + So the real leak has two sources, and neither is death: a device that genuinely + **goes away** (hot-unplug) has no retirement path, and a bus that enumerates + *differently* on restart leaves its stale entries behind forever. Both are the device + manager's inventory problem — phase 3 — and both need the id-stability question + answered first (tombstone-and-reuse aliases stale ids held by another process; + generation-tagged ids change the id encoding, which is ABI). Not an unattended + decision. ### Working rules for the run From f4eb88e7d28db5242be3b8838056ba54c83fe177 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 11:27:27 +0100 Subject: [PATCH 03/36] build: a new compile-time ceiling declares itself or does not land MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The convention that tunables live in system/parameters.zig with their reasoning attached predates this and got 2% compliance — 5 of 235. A convention with no teeth is how a bare `const maximum_devices = 64` reached an AMD desktop and cost it USB and storage. This is the same rule with a gate behind it. tools/check-bounds.py finds every bound-shaped declaration — a `maximum_*` const with a literal value, or a type with a literal array length — and requires the five-field block above it: what it counts, who decides its size, what it protects, what happens at the limit, and how anyone finds out. The at-limit vocabulary is closed: refuse, degrade, truncate, grow. There is deliberately no way to spell "silent", no way to spell "drop", and nothing meaning "allow", so the behaviours that did the damage cannot be written down. Truncation is legal only carrying a marker the reader can see, which is why klog_maximum_message qualifies and a USB descriptor cut at 512 bytes does not. An array length that names a declared bound is not itself a bound; only literal lengths are flagged, which pushes ceilings toward having names. The 273 that predate the rule are allowlisted so this lands without a tree-wide sweep in front of it, and that list may only shrink: declaring a bound means deleting its line, and the check fails on a stale entry too. Nothing may be added. Wired into `zig build test` and available alone as `zig build bounds`. Not in the default build — it reads the whole tree, and a red bounds check should not stop you booting a kernel. Five are now declared rather than allowlisted. Writing them out is its own argument: maximum_devices reads "protects: nothing — this is a sizing guess about someone else's computer", and maximum_tasks now carries the fact that it has been raised twice, each time by something that outgrew it. Verified the gate refuses an undeclared bound, a declared one using forbidden vocabulary, and an allowlist entry that has since been declared. Suite 115/115. --- build.zig | 13 ++ docs/bounds-track-plan.md | 2 +- system/kernel/devices-broker.zig | 17 ++ system/kernel/iommu.zig | 15 +- system/parameters.zig | 20 ++- tools/bounds-allowlist.txt | 285 +++++++++++++++++++++++++++++++ tools/check-bounds.py | 189 ++++++++++++++++++++ 7 files changed, 537 insertions(+), 4 deletions(-) create mode 100644 tools/bounds-allowlist.txt create mode 100644 tools/check-bounds.py diff --git a/build.zig b/build.zig index 412651e..52c34ec 100644 --- a/build.zig +++ b/build.zig @@ -397,6 +397,19 @@ pub fn build(b: *std.Build) void { // here (compiled for the host rather than inheriting a freestanding target) — // which also compile-checks that the three-way split stays self-consistent. const test_step = b.step("test", "Run tests"); + + // Every compile-time ceiling states what it counts, who decides its size, what it + // protects, what happens when it is reached, and how anyone finds out + // (docs/os-development/bounds.md). The ~273 that predate the rule are listed in + // tools/bounds-allowlist.txt so this could land without a tree-wide sweep first; + // that list may only shrink. Part of `test` rather than the default build: it reads + // the whole tree, and a red bounds check should not stop you booting a kernel. + const bounds_check = b.addSystemCommand(&.{ "python3", "tools/check-bounds.py" }); + bounds_check.setCwd(b.path(".")); + const bounds_step = b.step("bounds", "Check that every compile-time ceiling declares itself"); + bounds_step.dependOn(&bounds_check.step); + test_step.dependOn(&bounds_check.step); + for ([_][]const u8{ "system/boot-handoff.zig", "system/abi.zig", diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index de4394c..e709c81 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -13,7 +13,7 @@ next one starts.* | Step | What | State | |---|---|---| | L1 | Reclamation: a dead task's registrations die with its claims | **stopped — the step was wrong; see open question 4** | -| L2 | Bounds build check + allowlist; declare what we have already touched | not started | +| L2 | Bounds build check + allowlist; declare what we have already touched | **done** — `zig build bounds`, 273 allowlisted, 5 declared | | L3 | xHCI: slot count from `HCSPARAMS1.MaxSlots`, not 8 | not started | | L4 | USB: configuration descriptor sized by `wTotalLength`, not 512 | not started | | L5 | USB: interfaces from the descriptor, and the misattributed-endpoint bug | not started | diff --git a/system/kernel/devices-broker.zig b/system/kernel/devices-broker.zig index a985dd6..cc67ddd 100644 --- a/system/kernel/devices-broker.zig +++ b/system/kernel/devices-broker.zig @@ -23,6 +23,14 @@ const abi = @import("abi"); const platform = @import("platform"); const device_abi = @import("device-abi"); +/// bound: device nodes for the whole machine — firmware-discovered plus every child a +/// bus driver registers at runtime +/// decided-by: hardware +/// protects: nothing — this is a sizing guess about someone else's computer, which is +/// why an AMD Ryzen booted with a working display, no USB and no storage +/// at-limit: refuse — ENOSPC from device_register; `dropped` counts discovery losses +/// observed-by: the bus driver's line naming the reason (pci-bus reconciles functions +/// found against registered), and kernel.zig:203 for discovery drops pub const maximum_devices = 64; /// Cap on children a single parent may have. A zero-resource child (legal — a USB @@ -31,6 +39,15 @@ pub const maximum_devices = 64; /// `device_register` and exhaust the whole table, permanently denying it to every other /// driver. This bounds the blast radius of one claim; a real quota (and a /// `device_release` to reclaim on exit) is future work — see docs/driver-model.md. +/// +/// bound: children one claimed parent may register — in practice every PCI function on +/// the machine, since pci-bus registers them all under the one host bridge +/// decided-by: hardware +/// protects: the shared device table, against a driver looping device_register — but +/// it is a proxy for an authorisation the kernel does not perform, since any +/// process may claim any unclaimed device (docs/bounds-track-plan.md phase 2) +/// at-limit: refuse — ECHILDREN, distinct from a full table +/// observed-by: pci-bus logs the reason per refused function, and warns at end of scan const maximum_children_per_parent = 16; var devices: [maximum_devices]device_abi.DeviceDescriptor = undefined; diff --git a/system/kernel/iommu.zig b/system/kernel/iommu.zig index 824a95a..8c2c2ca 100644 --- a/system/kernel/iommu.zig +++ b/system/kernel/iommu.zig @@ -39,7 +39,20 @@ const page_size: u64 = abi.page_size; const page_mask: u64 = page_size - 1; const huge_page_size: u64 = 2 * 1024 * 1024; -/// One domain per claimed PCI function. 64 mirrors devices-broker's device cap. +/// One domain per claimed PCI function. Coupled to devices-broker's device cap — and +/// coupled *in code*, by the comptime assert beside `confined` below, because when +/// these two agreed only by this sentence the disagreement failed open. +/// +/// Both VT-d and AMD-Vi report the number of domains they support in a capability +/// register. We should be reading it rather than choosing 64 — docs/bounds-track-plan.md +/// phase 4. +/// +/// bound: IOMMU translation domains, one per claimed DMA-capable device +/// decided-by: hardware +/// protects: the statically sized domain and confinement tables +/// at-limit: refuse — ECONFINE; the claim is rolled back and the device is not driven, +/// because a claim that cannot be confined must not stand +/// observed-by: the claiming driver's own line naming ECONFINE pub const maximum_domains = 64; pub const invalid_domain: u16 = 0xFFFF; diff --git a/system/parameters.zig b/system/parameters.zig index a055bb0..d51ebdd 100644 --- a/system/parameters.zig +++ b/system/parameters.zig @@ -17,8 +17,13 @@ /// arrays (discovery pool, scheduler state, per-core GDT/TSS). Generous headroom: /// those structs are small, and the *large* per-core resources (kernel and IST /// stacks) are allocated at bring-up for cores that actually come online, so this -/// ceiling is cheap. A machine with more logical CPUs has its surplus reported and -/// left parked (see acpi `cpusDropped`). +/// ceiling is cheap. +/// +/// bound: logical CPUs the kernel tracks +/// decided-by: hardware +/// protects: the per-CPU bookkeeping arrays, which are sized at compile time +/// at-limit: degrade — the surplus cores are left parked, never brought online +/// observed-by: platform.cpusDropped() -> the WARNING at kernel.zig:281 pub const maximum_cpus = 128; /// Maximum tasks (kernel threads) alive at once — the static task-table size. Each @@ -29,6 +34,17 @@ pub const maximum_cpus = 128; /// for the USB stack: the xHCI bus driver spawns a supervised class-driver instance /// per matched interface (keyboard, mouse, mass storage), on top of the FAT and /// block servers and the growing ramdisk bundle. +/// +/// The history above is the argument against this number: it has been raised twice, +/// each time by a machine or a bundle that outgrew it, which is the pattern the +/// bounds rule exists to stop. It is `ours` only because the task table is static; +/// how many drivers a machine needs is decided by how much hardware it has. +/// +/// bound: kernel threads alive at once — the static task-table size +/// decided-by: hardware +/// protects: the statically allocated task table +/// at-limit: refuse — spawn fails; a supervised driver is never started +/// observed-by: the spawning supervisor's own log line; see docs/bounds-track-plan.md pub const maximum_tasks = 48; /// Each task's kernel stack (also each AP's bring-up stack), in bytes. diff --git a/tools/bounds-allowlist.txt b/tools/bounds-allowlist.txt new file mode 100644 index 0000000..542e2c6 --- /dev/null +++ b/tools/bounds-allowlist.txt @@ -0,0 +1,285 @@ +# Bounds that predate the rule (docs/os-development/bounds.md). +# +# Generated from the tree as it stood when the check landed, so the gate could start +# without a 278-site sweep in front of it. Every line is a compile-time ceiling that +# has not yet said what it counts, who decides its size, what it protects, what +# happens when it is reached, or how anyone finds out. +# +# **This list may only shrink.** Declaring a bound means deleting its line; the check +# fails if a listed bound is now declared, and fails if a listed bound has vanished. +# Nothing may be added to it — a new ceiling declares itself or does not land. +# +# docs/fixed-bounds-audit.md is the analysis of how they got here. +boot/efi.zig:buffer +boot/efi.zig:info_buffer +boot/efi.zig:maximum_bundled +boot/efi.zig:maximum_tree_depth +library/device/acpi/aml/interpreter.zig:argbuf +library/device/acpi/aml/interpreter.zig:args +library/device/acpi/aml/interpreter.zig:locals +library/device/acpi/aml/interpreter.zig:maximum_segments +library/device/acpi/aml/interpreter.zig:notify_queue +library/device/acpi/aml/namespace.zig:segment +library/device/acpi/aml/parser.zig:maximum_segments +library/device/driver/driver.zig:lookup_attempts +library/device/model/device-abi.zig:hid +library/device/model/device-abi.zig:maximum_device_resources +library/device/pci/pci.zig:bar_virtual +library/device/pci/pci.zig:bars +library/device/registry/device-registry.zig:cols +library/device/registry/device-registry.zig:rules +library/device/registry/device-registry.zig:rules +library/device/registry/device-registry.zig:rules +library/device/registry/device-registry.zig:rules +library/device/registry/device-registry.zig:rules +library/kernel/channel.zig:bind_attempts +library/kernel/channel.zig:name_maximum +library/kernel/channel.zig:path_maximum +library/kernel/file-system.zig:name_buffer +library/kernel/file-system.zig:out +library/kernel/logging.zig:buffer +library/kernel/process.zig:blob +library/kernel/process.zig:receive +library/kernel/process.zig:table +library/kernel/service.zig:subscriber_capacity +library/kernel/start.zig:buffer +library/kernel/thread.zig:readers +library/kernel/thread.zig:writers +library/protocol/device-manager/device-manager-protocol.zig:hid +library/protocol/display/display-protocol.zig:max_modes +library/protocol/envelope/envelope.zig:packet_maximum +library/protocol/envelope/envelope.zig:post_maximum +library/protocol/power/power-protocol.zig:hid +library/protocol/scanout/scanout-protocol.zig:max_modes +library/protocol/usb-transfer/usb-transfer-protocol.zig:max_report_data +library/protocol/usb-transfer/usb-transfer-protocol.zig:max_reported_endpoints +library/protocol/usb-transfer/usb-transfer-protocol.zig:setup +system/abi.zig:errno_maximum +system/abi.zig:klog_maximum_message +system/abi.zig:maximum_process_name +system/boot-handoff.zig:kernel_segments +system/drivers/pci-bus/pci-bus.zig:line +system/drivers/pci-bus/pci-bus.zig:sub_buffer +system/drivers/ps2-bus/mouse-packet.zig:bytes +system/drivers/ps2-bus/scancode.zig:pressed +system/drivers/ps2-bus/scancode.zig:set2_base +system/drivers/ps2-bus/scancode.zig:set2_extended +system/drivers/usb-hid/hid-report.zig:keys +system/drivers/usb-hid/keyboard.zig:receive +system/drivers/usb-hid/mouse.zig:receive +system/drivers/usb-storage/bulk-only-transport.zig:cdb +system/drivers/usb-storage/scsi.zig:op_read_capacity_10 +system/drivers/usb-storage/usb-storage.zig:capacity_bytes +system/drivers/usb-xhci-bus/usb-xhci-bus.zig:hid_buffer +system/drivers/usb-xhci-bus/usb-xhci-bus.zig:maker_buffer +system/drivers/usb-xhci-bus/usb-xhci-bus.zig:maker_buffer +system/drivers/usb-xhci-bus/usb-xhci-bus.zig:prev_connected +system/drivers/usb-xhci-bus/usb-xhci-bus.zig:product_buffer +system/drivers/usb-xhci-bus/usb-xhci-bus.zig:product_buffer +system/drivers/usb-xhci-bus/usb-xhci-library.zig:blob +system/drivers/usb-xhci-bus/usb-xhci-library.zig:buffer +system/drivers/usb-xhci-bus/usb-xhci-library.zig:bytes +system/drivers/usb-xhci-bus/usb-xhci-library.zig:data +system/drivers/usb-xhci-bus/usb-xhci-library.zig:descriptor +system/drivers/usb-xhci-bus/usb-xhci-library.zig:head +system/drivers/usb-xhci-bus/usb-xhci-library.zig:header +system/drivers/usb-xhci-bus/usb-xhci-library.zig:max_devices +system/drivers/usb-xhci-bus/usb-xhci-library.zig:max_endpoints_per_interface +system/drivers/usb-xhci-bus/usb-xhci-library.zig:max_interfaces +system/drivers/usb-xhci-bus/usb-xhci-library.zig:max_subscriptions +system/drivers/usb-xhci-bus/usb-xhci-library.zig:port_changes +system/drivers/usb-xhci-bus/usb-xhci-library.zig:raw +system/drivers/usb-xhci-bus/usb-xhci-library.zig:report_queue_capacity +system/drivers/usb-xhci-bus/usb-xhci-library.zig:request_set_hub_depth +system/drivers/virtio-gpu/virtio-gpu-protocol.zig:edid +system/drivers/virtio-gpu/virtio-gpu-protocol.zig:max_scanouts +system/drivers/virtio-gpu/virtio-gpu.zig:descriptors +system/drivers/virtio-gpu/virtio-gpu.zig:max_height +system/drivers/virtio-gpu/virtio-gpu.zig:max_width +system/initial-ramdisk.zig:buffer +system/initial-ramdisk.zig:buffer +system/initial-ramdisk.zig:maximum_name +system/kernel/acpi.zig:APIC +system/kernel/acpi.zig:BERT +system/kernel/acpi.zig:CPEP +system/kernel/acpi.zig:DMAR +system/kernel/acpi.zig:DSDT +system/kernel/acpi.zig:ECDT +system/kernel/acpi.zig:EINJ +system/kernel/acpi.zig:ERST +system/kernel/acpi.zig:FACP +system/kernel/acpi.zig:FACS +system/kernel/acpi.zig:HEST +system/kernel/acpi.zig:HPET +system/kernel/acpi.zig:IVRS +system/kernel/acpi.zig:MCFG +system/kernel/acpi.zig:MPST +system/kernel/acpi.zig:MSCT +system/kernel/acpi.zig:PMTT +system/kernel/acpi.zig:PSDT +system/kernel/acpi.zig:RASF +system/kernel/acpi.zig:RSDT +system/kernel/acpi.zig:SBST +system/kernel/acpi.zig:SLIT +system/kernel/acpi.zig:SPCR +system/kernel/acpi.zig:SRAT +system/kernel/acpi.zig:SSDT +system/kernel/acpi.zig:XSDT +system/kernel/acpi.zig:aml_block_len +system/kernel/acpi.zig:aml_block_physical +system/kernel/acpi.zig:holes +system/kernel/acpi.zig:maximum_rmrr +system/kernel/acpi.zig:nb +system/kernel/acpi.zig:nb +system/kernel/acpi.zig:nb +system/kernel/acpi.zig:oem_id +system/kernel/acpi.zig:oem_id +system/kernel/acpi.zig:oem_id +system/kernel/acpi.zig:oem_table_id +system/kernel/acpi.zig:overrides +system/kernel/acpi.zig:rmrr_limit_offset +system/kernel/acpi.zig:signature +system/kernel/acpi.zig:signature +system/kernel/acpi.zig:signature +system/kernel/architecture/x86_64/cpu.zig:irq_vector_count +system/kernel/architecture/x86_64/idt.zig:gate_count +system/kernel/architecture/x86_64/ioapic.zig:overrides +system/kernel/architecture/x86_64/iommu-amd.zig:buffer +system/kernel/architecture/x86_64/iommu-amd.zig:buffer +system/kernel/architecture/x86_64/iommu-intel.zig:buffer +system/kernel/architecture/x86_64/iommu-intel.zig:buffer +system/kernel/architecture/x86_64/iommu-intel.zig:context_table +system/kernel/device-model.zig:buffer +system/kernel/device-model.zig:cbuf +system/kernel/device-model.zig:hid_buffer +system/kernel/device-model.zig:maximum_resources +system/kernel/device-model.zig:name_buffer +system/kernel/device-model.zig:rbuf +system/kernel/heap.zig:heap_maximum +system/kernel/ipc-synchronous.zig:MESSAGE_MAXIMUM +system/kernel/ipc-synchronous.zig:POST_MAXIMUM +system/kernel/ipc-synchronous.zig:notify_buffer +system/kernel/ipc-synchronous.zig:post_capacity +system/kernel/irq.zig:maximum_gsi +system/kernel/irq.zig:msi_bound +system/kernel/irq.zig:msi_owner +system/kernel/irq.zig:vector_gsi +system/kernel/kernel.zig:buffer +system/kernel/kernel.zig:buffer +system/kernel/kernel.zig:buffer +system/kernel/kernel.zig:isos +system/kernel/kernel.zig:maximum_wake_attempts +system/kernel/log-ring.zig:message +system/kernel/log-ring.zig:out +system/kernel/log.zig:buffer +system/kernel/log.zig:maximum_sinks +system/kernel/log.zig:message +system/kernel/log.zig:ring_capacity +system/kernel/process.zig:chunk +system/kernel/process.zig:chunk +system/kernel/process.zig:chunk +system/kernel/process.zig:chunk +system/kernel/process.zig:exit_record_capacity +system/kernel/process.zig:exit_subscriber_capacity +system/kernel/process.zig:maximum_argument_bytes +system/kernel/process.zig:maximum_arguments +system/kernel/process.zig:maximum_dma_regions +system/kernel/process.zig:maximum_mmap_pages +system/kernel/process.zig:maximum_mount_prefix +system/kernel/process.zig:maximum_mount_rewrite +system/kernel/process.zig:maximum_pages +system/kernel/process.zig:maximum_resolve_path +system/kernel/process.zig:maximum_segments +system/kernel/process.zig:maximum_shared_memory_pages +system/kernel/process.zig:name_buffer +system/kernel/process.zig:timer_capacity +system/kernel/process.zig:word_bytes +system/kernel/process.zig:write_buffer +system/kernel/scheduler.zig:ipc_maximum_handles +system/kernel/scheduler.zig:maximum_space_mappings +system/kernel/vfs.zig:maximum_directories +system/kernel/vfs.zig:maximum_mounts +system/kernel/vfs.zig:maximum_prefix +system/kernel/vfs.zig:maximum_rewrite +system/services/acpi/acpi.zig:blocks +system/services/acpi/acpi.zig:buffer +system/services/acpi/acpi.zig:hid +system/services/acpi/acpi.zig:mmio_scratch +system/services/acpi/acpi.zig:name +system/services/acpi/acpi.zig:registered +system/services/device-manager/device-manager.zig:arguments +system/services/device-manager/device-manager.zig:id_text +system/services/device-manager/device-manager.zig:maximum_children +system/services/device-manager/device-manager.zig:maximum_drivers +system/services/device-manager/device-manager.zig:name_buffer +system/services/device-manager/device-manager.zig:registry_rules +system/services/device-manager/device-manager.zig:registry_source +system/services/display/backend.zig:device_table +system/services/display/backend.zig:line +system/services/display/compositor.zig:capacity +system/services/display/compositor.zig:maximum_columns +system/services/display/compositor.zig:maximum_rects +system/services/display/compositor.zig:maximum_rows +system/services/display/display.zig:line +system/services/display/display.zig:line +system/services/display/display.zig:line +system/services/display/display.zig:list +system/services/display/display.zig:maximum_layers +system/services/display/display.zig:mode_list +system/services/fat/engine.zig:buf +system/services/fat/engine.zig:buf +system/services/fat/engine.zig:buf +system/services/fat/engine.zig:buffer +system/services/fat/engine.zig:cached_back +system/services/fat/engine.zig:chunk +system/services/fat/engine.zig:chunk +system/services/fat/engine.zig:chunk +system/services/fat/engine.zig:device_back +system/services/fat/engine.zig:display +system/services/fat/engine.zig:long_name +system/services/fat/engine.zig:long_name +system/services/fat/engine.zig:long_name +system/services/fat/engine.zig:max_transfer_sectors +system/services/fat/engine.zig:pair +system/services/fat/engine.zig:pair +system/services/fat/engine.zig:payload +system/services/fat/engine.zig:payload +system/services/fat/engine.zig:payload +system/services/fat/engine.zig:readback +system/services/fat/engine.zig:run +system/services/fat/engine.zig:run +system/services/fat/engine.zig:short +system/services/fat/engine.zig:short +system/services/fat/engine.zig:short +system/services/fat/engine.zig:stem +system/services/fat/engine.zig:tail_buffer +system/services/fat/engine.zig:units +system/services/fat/engine.zig:value +system/services/fat/engine.zig:value +system/services/fat/engine.zig:window +system/services/fat/on-disk.zig:filesystem_type +system/services/fat/on-disk.zig:filesystem_type +system/services/fat/on-disk.zig:jump +system/services/fat/on-disk.zig:name +system/services/fat/on-disk.zig:name1 +system/services/fat/on-disk.zig:name2 +system/services/fat/on-disk.zig:name3 +system/services/fat/on-disk.zig:oem_name +system/services/fat/on-disk.zig:volume_label +system/services/fat/on-disk.zig:volume_label +system/services/init/init.zig:binary +system/services/init/init.zig:init_csv +system/services/init/init.zig:max_service_args +system/services/init/init.zig:max_services +system/services/init/init.zig:maximum_bindings +system/services/init/init.zig:maximum_grants +system/services/init/init.zig:maximum_name +system/services/init/init.zig:maximum_restarts +system/services/init/init.zig:process_table +system/services/init/init.zig:protocol_csv +system/services/logger/logger.zig:chunk +system/services/logger/logger.zig:gap_line +system/services/logger/logger.zig:line +system/services/logger/logger.zig:line +system/services/logger/logger.zig:maximum_files +system/services/logger/logger.zig:stamp diff --git a/tools/check-bounds.py b/tools/check-bounds.py new file mode 100644 index 0000000..8b85664 --- /dev/null +++ b/tools/check-bounds.py @@ -0,0 +1,189 @@ +#!/usr/bin/env python3 +"""Every compile-time ceiling states what it is doing there. + +A *bound* is a number chosen at compile time that decides how much of something the +code can hold: `const maximum_devices = 64`, `var below: [64]Range`, `var blob: [512]u8`. +Different units, one shape, and one recurring way of going wrong — see +docs/fixed-bounds-audit.md, where 235 of them turned up, 139 on quantities the machine +or a file decides rather than us, and 171 silent when reached. + +This is the gate that keeps new ones from joining them. It does not resize anything and +it makes no judgement about whether a bound should exist; it only refuses one that will +not say what it is for. The declaration is a doc comment immediately above: + + /// bound: logical CPUs the kernel tracks + /// decided-by: hardware + /// protects: the per-CPU bookkeeping arrays, which are sized at compile time + /// at-limit: degrade - surplus cores are left parked, never brought online + /// observed-by: platform.cpusDropped() -> the WARNING at kernel.zig:281 + pub const maximum_cpus = 128; + +`decided-by` is the field the audit turned on: `hardware` and `external` mean the +quantity is not ours to choose, and a fixed bound on one of those is a defect rather +than a tunable. `at-limit`'s vocabulary is closed on purpose — there is no `silent`, +no `drop`, and nothing meaning *allow*, so the behaviours that caused the damage cannot +be written down. `truncate` is legal only with a marker the reader can see. + +An array length that names a declared bound (`[maximum_devices]Descriptor`) is not +itself a bound: the number lives at the declaration, and that is where it is declared. +Only literal lengths are flagged, which pushes ceilings toward having names. + +The ~235 that already exist are listed in tools/bounds-allowlist.txt so this can land +without a tree-wide sweep in front of it. That list may only shrink: declaring a bound +means deleting its line, and a stale line is an error too. + +Usage: check-bounds.py [--list] [repo-root] + --list print every undeclared bound found, for regenerating the allowlist +""" + +import re +import sys +from pathlib import Path + +# Directories worth gating. `test/` is excluded: a fixture's `[100]MemoryRegion` is a +# test input, not a ceiling the system runs into. +ROOTS = ("system", "library", "boot") +SKIP_PARTS = {".zig-cache", "zig-out", ".git", ".claude", "vendor", "generated"} +SKIP_FILES = {"tests.zig"} + +FIELDS = ("bound", "decided-by", "protects", "at-limit", "observed-by") +DECIDED_BY = {"hardware", "external", "ours"} +AT_LIMIT = {"refuse", "degrade", "truncate", "grow"} + +# A `const` whose name reads like a ceiling and whose value is an integer literal. +NAMED = re.compile( + r"^\s*(?:pub\s+)?const\s+([A-Za-z_]\w*)\s*(?::\s*[\w.\[\]]+\s*)?=\s*" + r"(\d[\d_]*|0x[0-9a-fA-F_]+)\s*(?:\*\s*\d[\d_]*\s*)*;" +) +NAME_IS_BOUND = re.compile(r"(^|_)(maximum|max|limit|capacity|depth|attempts|count)($|_)", re.I) + +# A declaration or struct field whose type carries a *literal* array length. +ARRAY_DECL = re.compile(r"^\s*(?:pub\s+)?(?:const|var)\s+([A-Za-z_]\w*)\s*:[^=]*?\[\s*(\d[\d_]*)\s*\]") +ARRAY_FIELD = re.compile(r"^\s*([A-Za-z_]\w*)\s*:\s*\[\s*(\d[\d_]*)\s*\]") + +# Struct padding and reserved fields are shapes, not ceilings — nothing is ever "held" +# in them. Everything else with a literal length is a candidate, *including* the tidy +# powers of two: `[512]u8` and `[64]Range` were the two worst findings in the audit, and +# any size-based exemption would have skipped exactly them. A length that is genuinely a +# fact rather than a ceiling says so in its declaration ("decided-by: hardware, the PCI +# spec gives a function 6 BARs") — that is what the declaration is for. +NOT_A_BOUND_NAME = re.compile(r"^_*(padding|pad|reserved|unused|spare)\d*$", re.I) + + +def sources(root: Path): + for top in ROOTS: + base = root / top + if not base.is_dir(): + continue + for path in sorted(base.rglob("*.zig")): + if SKIP_PARTS & set(path.parts) or path.name in SKIP_FILES: + continue + yield path + + +def declaration_above(lines, index): + """The `/// key: value` block immediately above line `index`, as a dict.""" + fields = {} + i = index - 1 + while i >= 0: + stripped = lines[i].strip() + if not stripped.startswith("///"): + break + body = stripped[3:].strip() + match = re.match(r"([a-z-]+):\s*(.+)", body) + if match: + fields[match.group(1)] = match.group(2).strip() + i -= 1 + return fields + + +def problems_with(fields): + """Why a declaration is not acceptable, or an empty list.""" + missing = [f for f in FIELDS if f not in fields or not fields[f]] + if missing: + return ["missing " + ", ".join(missing)] + out = [] + if fields["decided-by"] not in DECIDED_BY: + out.append(f"decided-by must be one of {sorted(DECIDED_BY)}, not {fields['decided-by']!r}") + verb = fields["at-limit"].split()[0].strip("-:,").lower() + if verb not in AT_LIMIT: + out.append( + f"at-limit must start with one of {sorted(AT_LIMIT)}, not {verb!r}. " + "There is deliberately no way to say 'silent', 'drop', or anything meaning 'allow'" + ) + if verb == "truncate" and len(fields["at-limit"].split()) < 3: + out.append("at-limit: truncate must say how a reader can TELL it happened") + return out + + +def find(root: Path): + """Every bound-shaped declaration: (relative path, name, value, line, fields).""" + for path in sources(root): + rel = path.relative_to(root).as_posix() + lines = path.read_text(encoding="utf-8", errors="replace").split("\n") + for n, line in enumerate(lines): + if line.lstrip().startswith("//"): + continue + name = value = None + m = NAMED.match(line) + if m and NAME_IS_BOUND.search(m.group(1)): + name, value = m.group(1), m.group(2) + else: + m = ARRAY_DECL.match(line) or ARRAY_FIELD.match(line) + if m and not NOT_A_BOUND_NAME.match(m.group(1)): + name, value = m.group(1), m.group(2) + if name: + yield rel, name, value, n + 1, declaration_above(lines, n) + + +def main(): + argv = [a for a in sys.argv[1:] if not a.startswith("--")] + listing = "--list" in sys.argv + root = Path(argv[0]) if argv else Path(__file__).resolve().parent.parent + + allow_path = root / "tools" / "bounds-allowlist.txt" + allowed = set() + if allow_path.exists(): + for raw in allow_path.read_text().split("\n"): + entry = raw.split("#", 1)[0].strip() + if entry: + allowed.add(entry) + + undeclared, bad, seen = [], [], set() + for rel, name, value, line, fields in find(root): + key = f"{rel}:{name}" + seen.add(key) + if not fields: + (undeclared if key not in allowed else []).append((key, value, line)) + continue + for why in problems_with(fields): + bad.append((key, line, why)) + if key in allowed and not problems_with(fields): + bad.append((key, line, "now declared — delete its line from tools/bounds-allowlist.txt")) + + if listing: + for key, value, line in sorted(undeclared): + print(f"{key} # = {value}, line {line}") + for key in sorted(allowed - seen): + print(f"# STALE: {key}") + return 0 + + stale = sorted(allowed - seen) + if not undeclared and not bad and not stale: + return 0 + + print("bounds check failed\n", file=sys.stderr) + for key, value, line in sorted(undeclared): + print(f" {key} (= {value}, line {line})", file=sys.stderr) + print(" no declaration. A ceiling states what it counts, who decides its", file=sys.stderr) + print(" size, what it protects, what happens at the limit, and how you", file=sys.stderr) + print(" find out. See docs/os-development/bounds.md.", file=sys.stderr) + for key, line, why in sorted(bad): + print(f" {key} (line {line}): {why}", file=sys.stderr) + for key in stale: + print(f" {key}: allowlisted but no longer found — delete its line", file=sys.stderr) + return 1 + + +if __name__ == "__main__": + sys.exit(main()) From 69fbef40c07245a1fffd39bc24b6eb1e38758eaf Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 11:39:49 +0100 Subject: [PATCH 04/36] usb: the controller says how many device slots it has MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit max_devices was 8, with the comment "QEMU presents a handful; a fuller machine would grow this" — a number chosen against the test rig, waiting for a real machine, which is the pattern docs/bounds-track-plan.md exists to stop. The driver already knew the true figure. It reads HCSPARAMS1.MaxSlots at bring-up and writes it straight into op_config, so every slot the controller offers has always been *enabled*; only the array tracking them was 8. QEMU's xHCI reports 64, so seven eighths of the controller was live and invisible, and the ninth device — a keyboard, mouse, webcam, headset, hub and two sticks reach that without trying — disappeared on a hub-attached path that logs nothing at all. The array becomes a slice allocated from max_slots at bring-up. A controller claiming zero slots cannot address anything, so that is now a dead controller rather than an empty allocation failing mysteriously later. The Device Context Base Address Array is a page, 511 usable entries, so it already covered the 255-slot maximum. The bring-up line reports both numbers, and usb-hid asserts they are equal with a backreference rather than a magic number, so the test cannot drift from the hardware. Pinning tracking back to 8 fails it: "64 slots, tracking 8". Suite 115/115. --- docs/bounds-track-plan.md | 2 +- system/drivers/usb-xhci-bus/usb-xhci-bus.zig | 7 +++- .../drivers/usb-xhci-bus/usb-xhci-library.zig | 32 +++++++++++++------ test/qemu_test.py | 9 +++++- tools/bounds-allowlist.txt | 1 - 5 files changed, 37 insertions(+), 14 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index e709c81..3c387f7 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -14,7 +14,7 @@ next one starts.* |---|---|---| | L1 | Reclamation: a dead task's registrations die with its claims | **stopped — the step was wrong; see open question 4** | | L2 | Bounds build check + allowlist; declare what we have already touched | **done** — `zig build bounds`, 273 allowlisted, 5 declared | -| L3 | xHCI: slot count from `HCSPARAMS1.MaxSlots`, not 8 | not started | +| L3 | xHCI: slot count from `HCSPARAMS1.MaxSlots`, not 8 | **done** — QEMU reports 64; the driver tracked 8 | | L4 | USB: configuration descriptor sized by `wTotalLength`, not 512 | not started | | L5 | USB: interfaces from the descriptor, and the misattributed-endpoint bug | not started | | L6 | xHCI: a failed `allocateDevice` stops leaking an enabled slot | not started | diff --git a/system/drivers/usb-xhci-bus/usb-xhci-bus.zig b/system/drivers/usb-xhci-bus/usb-xhci-bus.zig index 7a1d8b0..be7a9b3 100644 --- a/system/drivers/usb-xhci-bus/usb-xhci-bus.zig +++ b/system/drivers/usb-xhci-bus/usb-xhci-bus.zig @@ -228,8 +228,13 @@ fn initialise(endpoint: ipc.Handle) bool { _ = logging.write("/system/drivers/usb-xhci-bus: controller reset/bring-up failed\n"); return false; }; - std.log.info("controller running ({d} slots, {d}-byte contexts)", .{ + // Both numbers, because for a long time they disagreed silently: the controller + // reported its real slot count and the driver tracked a fixed 8 of them, so a + // ninth device — trivially reachable behind a hub — simply did not exist. They + // must now be equal, and the suite asserts it. + std.log.info("controller running ({d} slots, tracking {d}, {d}-byte contexts)", .{ controller.?.max_slots, + controller.?.devices.len, controller.?.context_size, }); // The proof of life: a No-Op command round-trips the command ring, the event diff --git a/system/drivers/usb-xhci-bus/usb-xhci-library.zig b/system/drivers/usb-xhci-bus/usb-xhci-library.zig index fb0862d..d0cf396 100644 --- a/system/drivers/usb-xhci-bus/usb-xhci-library.zig +++ b/system/drivers/usb-xhci-bus/usb-xhci-library.zig @@ -328,9 +328,6 @@ pub const Report = struct { data: [64]u8 = [_]u8{0} ** 64, }; -// How many addressed devices this driver tracks at once. QEMU presents a handful -// (a keyboard, a mouse, a storage stick); a fuller machine would grow this. -const max_devices = 8; const max_subscriptions = 8; const report_queue_capacity = 16; @@ -449,7 +446,14 @@ pub const Controller = struct { device_context_array: memory.DmaRegion, command_ring: ProducerRing, event_ring: EventRing, - devices: [max_devices]Device = [_]Device{.{}} ** max_devices, + /// One entry per device slot the **controller** says it has (HCSPARAMS1.MaxSlots, + /// 1..255), allocated at bring-up. This used to be a fixed 8 with the comment "QEMU + /// presents a handful; a fuller machine would grow this" — which is the shape the + /// bounds rule exists to stop, since the controller has always reported the real + /// number and `op_config` below is already programmed with it. A desktop's keyboard, + /// mouse, webcam, headset, hub and two sticks reach 8 without trying, and everything + /// past it vanished (behind a hub, without even a log line). + devices: []Device, subscriptions: [max_subscriptions]Subscription = [_]Subscription{.{}} ** max_subscriptions, report_queue: [report_queue_capacity]Report = [_]Report{.{}} ** report_queue_capacity, report_count: usize = 0, @@ -582,8 +586,16 @@ pub const Controller = struct { .device_context_array = undefined, .command_ring = undefined, .event_ring = undefined, + .devices = &.{}, }; + // One tracking slot per slot the controller reports. A controller that claims + // no slots cannot address anything, so treat that as a dead controller rather + // than allocating nothing and failing mysteriously later. + if (self.max_slots == 0) return null; + self.devices = memory.allocator().alloc(Device, self.max_slots) catch return null; + for (self.devices) |*device| device.* = .{}; + // Wait for the controller to report ready, then halt it if it is running. if (!waitClear(self.operational(op_usbsts), usbsts_controller_not_ready)) return null; if (read32(self.operational(op_usbcmd)) & usbcmd_run != 0) { @@ -790,7 +802,7 @@ pub const Controller = struct { } fn allocateDevice(self: *Controller) ?*Device { - for (&self.devices) |*device| { + for (self.devices) |*device| { if (!device.used) return device; } return null; @@ -1011,7 +1023,7 @@ pub const Controller = struct { /// The next pending (hub, downstream-port) change to service, or null. Clears /// the returned port's bit. Called on the bus tick. pub fn takeHubChange(self: *Controller) ?struct { hub: *Device, port: u16 } { - for (&self.devices) |*device| { + for (self.devices) |*device| { if (!device.used or !device.is_hub or device.hub_change_mask == 0) continue; const bit: u5 = @intCast(@ctz(device.hub_change_mask)); device.hub_change_mask &= ~(@as(u32, 1) << bit); @@ -1084,7 +1096,7 @@ pub const Controller = struct { } pub fn deviceOnHubPort(self: *Controller, hub: *Device, port: u16) ?*Device { - for (&self.devices) |*device| { + for (self.devices) |*device| { if (device.used and device.parent_slot == hub.slot_id and device.parent_port == port) return device; } return null; @@ -1364,7 +1376,7 @@ pub const Controller = struct { /// Find the tracked device and interface an assigned device id belongs to. pub fn findInterface(self: *Controller, device_id: u64) ?struct { device: *Device, interface: *InterfaceInfo } { - for (&self.devices) |*device| { + for (self.devices) |*device| { if (!device.used) continue; for (device.interfaces[0..device.interface_count]) |*interface| { if (interface.registered_device_id == device_id) return .{ .device = device, .interface = interface }; @@ -1633,7 +1645,7 @@ pub const Controller = struct { /// The tracked device on `port`, or null. pub fn deviceOnPort(self: *Controller, port: u32) ?*Device { - for (&self.devices) |*device| { + for (self.devices) |*device| { if (device.used and device.port == port) return device; } return null; @@ -1642,7 +1654,7 @@ pub const Controller = struct { /// The next used device whose parent hub is `hub_slot` and slot id > `after` /// (for recursive teardown when a hub itself disconnects), or null. pub fn nextChildOf(self: *Controller, hub_slot: u8, after: u8) ?*Device { - for (&self.devices) |*device| { + for (self.devices) |*device| { if (device.used and device.parent_slot == hub_slot and device.slot_id > after) return device; } return null; diff --git a/test/qemu_test.py b/test/qemu_test.py index 8bcccc5..a46907f 100644 --- a/test/qemu_test.py +++ b/test/qemu_test.py @@ -625,11 +625,18 @@ CASES = [ # manager spawn the USB keyboard driver, which opens its device over the # transfer protocol, asks for boot protocol, subscribes to its interrupt # endpoint, and comes up — proof the class-driver <-> controller path works. + # + # It also pins the slot count to the hardware's answer. The backreference is the + # assertion: the driver must track exactly as many device slots as the controller + # reports in HCSPARAMS1.MaxSlots. It tracked a fixed 8 while QEMU's xHCI reports + # 64, so seven eighths of the controller was invisible and a device behind a hub + # past the eighth vanished without a log line. \1 fails the moment they diverge. {"name": "usb-hid", "smp": 4, "timeout": 150, # usb-kbd/usb-mouse ride the default boot xHCI bus (see qemu_args). - "expect": r"(?=[\s\S]*usb-hid-keyboard: ok)(?=[\s\S]*usb-hid-mouse: ok)", + "expect": r"(?=[\s\S]*usb-xhci-bus: controller running \((\d+) slots, tracking \1,)" + r"(?=[\s\S]*usb-hid-keyboard: ok)(?=[\s\S]*usb-hid-mouse: ok)", "fail": r"DANOS-TEST-RESULT: FAIL"}, # Keyboard echo: inject a known phrase via QMP send-key; the usb-hid-keyboard # driver decodes it and echoes each character to the log (the simple diff --git a/tools/bounds-allowlist.txt b/tools/bounds-allowlist.txt index 542e2c6..dd38ac1 100644 --- a/tools/bounds-allowlist.txt +++ b/tools/bounds-allowlist.txt @@ -83,7 +83,6 @@ system/drivers/usb-xhci-bus/usb-xhci-library.zig:data system/drivers/usb-xhci-bus/usb-xhci-library.zig:descriptor system/drivers/usb-xhci-bus/usb-xhci-library.zig:head system/drivers/usb-xhci-bus/usb-xhci-library.zig:header -system/drivers/usb-xhci-bus/usb-xhci-library.zig:max_devices system/drivers/usb-xhci-bus/usb-xhci-library.zig:max_endpoints_per_interface system/drivers/usb-xhci-bus/usb-xhci-library.zig:max_interfaces system/drivers/usb-xhci-bus/usb-xhci-library.zig:max_subscriptions From 6328823ef1d200cae78f75ee2aba6e06f974c882 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 11:53:01 +0100 Subject: [PATCH 05/36] usb: a configuration block is as long as the device says it is MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The driver read the first 512 bytes of a configuration block into a fixed buffer and parsed those. The block's length is the device's own choice (wTotalLength, a u16), so anything larger was silently cut: interfaces past the cut did not exist as far as the host was concerned, while the SET_CONFIGURATION that follows still configured the device for all of them. A headset is 500-900 bytes, a UVC webcam 1-3 KB, a multifunction printer 600+. Now allocated at the declared length, so the ceiling is the field's u16 — the specification's number rather than one of ours. A block shorter than its own 9-byte header is refused rather than trusted. The bring-up line reports the declared length and the bytes actually read, so a truncation can never again be invisible, and usb-large-descriptor asserts they match with a backreference. That case has an honest limit, recorded in its comment: QEMU cannot produce a block over 512 bytes. The boot keyboard, mouse and stick are 34-44, and the largest device available is usb-audio in multi-channel mode at 211 — which is exactly why the suite never caught this, and why it cannot now reproduce the original trigger. What it does catch is the class: any clamp below the attached device's block fails it, verified by pinning the buffer to 128 and watching "config block 211 bytes, read 128" turn the case red. Suite 115 -> 116. --- docs/bounds-track-plan.md | 2 +- .../drivers/usb-xhci-bus/usb-xhci-library.zig | 25 ++++++++++++++----- test/qemu_test.py | 22 ++++++++++++++++ tools/bounds-allowlist.txt | 1 - 4 files changed, 42 insertions(+), 8 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 3c387f7..c44794f 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -15,7 +15,7 @@ next one starts.* | L1 | Reclamation: a dead task's registrations die with its claims | **stopped — the step was wrong; see open question 4** | | L2 | Bounds build check + allowlist; declare what we have already touched | **done** — `zig build bounds`, 273 allowlisted, 5 declared | | L3 | xHCI: slot count from `HCSPARAMS1.MaxSlots`, not 8 | **done** — QEMU reports 64; the driver tracked 8 | -| L4 | USB: configuration descriptor sized by `wTotalLength`, not 512 | not started | +| L4 | USB: configuration descriptor sized by `wTotalLength`, not 512 | **done** — QEMU tops out at 211 bytes, so the case catches the class, not the original trigger | | L5 | USB: interfaces from the descriptor, and the misattributed-endpoint bug | not started | | L6 | xHCI: a failed `allocateDevice` stops leaking an enabled slot | not started | diff --git a/system/drivers/usb-xhci-bus/usb-xhci-library.zig b/system/drivers/usb-xhci-bus/usb-xhci-library.zig index d0cf396..7e3d8c5 100644 --- a/system/drivers/usb-xhci-bus/usb-xhci-library.zig +++ b/system/drivers/usb-xhci-bus/usb-xhci-library.zig @@ -1311,12 +1311,25 @@ pub const Controller = struct { const configuration = std.mem.bytesToValue(usb_abi.ConfigurationDescriptor, &header); device.configuration_value = @intFromEnum(configuration.configuration_value); - // Read the whole block into a local buffer and parse it here (in the bus - // driver) so the parse never has to cross the 256-byte IPC boundary. - var blob: [512]u8 = undefined; - const length = @min(configuration.total_length, blob.len); - if (!self.controlTransfer(device, usb_abi.getDescriptor(.configuration, 0, 0, @intCast(length)), blob[0..length], true)) return false; - parseConfiguration(device, blob[0..length]); + // Read the whole block and parse it here (in the bus driver) so the parse + // never has to cross the 256-byte IPC boundary. Sized by the device's own + // wTotalLength — the ceiling is then the field's u16, which is the USB + // specification's, not ours. + // + // This was a fixed 512 with an `@min` clamp, which silently truncated: a + // composite device routinely exceeds it (a headset is 500-900 bytes, a UVC + // webcam 1-3 KB, a multifunction printer 600+), and the interfaces past the + // cut simply did not exist — while SET_CONFIGURATION below still configured + // the device for all of them. QEMU's boot keyboard, mouse and stick are all + // under 100 bytes, which is why it survived. + if (configuration.total_length < header.len) return false; // shorter than its own header + const blob = memory.allocator().alloc(u8, configuration.total_length) catch return false; + defer memory.allocator().free(blob); + if (!self.controlTransfer(device, usb_abi.getDescriptor(.configuration, 0, 0, configuration.total_length), blob, true)) return false; + // Both numbers, so a truncation can never again be invisible: the block the + // device declared, and the bytes actually fetched and parsed. They must match. + std.log.info("config block {d} bytes, read {d}", .{ configuration.total_length, blob.len }); + parseConfiguration(device, blob); // Select the configuration, moving the device to the configured state. if (!self.controlTransfer(device, usb_abi.setConfiguration(configuration.configuration_value), &.{}, false)) return false; diff --git a/test/qemu_test.py b/test/qemu_test.py index a46907f..3f52528 100644 --- a/test/qemu_test.py +++ b/test/qemu_test.py @@ -674,6 +674,28 @@ CASES = [ r"(?=.*hub slot \d+ port \d+ device:.*0x0627)" r"(?=.*usb-hid-keyboard: ok \(device 3)", "fail": r"DANOS-TEST-RESULT: FAIL"}, + # The driver must never truncate a configuration block. Its size is the DEVICE's + # choice (wTotalLength, a u16); the driver used to read the first 512 bytes into a + # fixed buffer and parse those, so interfaces past the cut did not exist while + # SET_CONFIGURATION still configured the device for all of them. + # + # The backreference is the assertion: declared length and bytes read must match. + # + # Honest limit: QEMU cannot produce a block over 512 bytes. Its boot keyboard, + # mouse and stick are 34-44, and the largest device on offer is usb-audio in + # multi-channel mode at 211 — which is why the suite never saw the original bug, + # and why it cannot now reproduce that exact trigger. What this case does catch is + # the class: any clamp below the attached device's block size fails it, verified + # by pinning the buffer to 128 and watching it go red. A real headset (500-900 + # bytes), UVC webcam (1-3 KB) or multifunction printer trips the 512 itself. + {"name": "usb-large-descriptor", + "build_case": "usb-hid", + "smp": 4, + "timeout": 150, + "qemu_extra": ["-device", "qemu-xhci,id=xhci2", + "-device", "usb-audio,bus=xhci2.0,port=1,multi=on"], + "expect": r"(?s)(?=.*usb-xhci-bus: config block (\d\d\d+) bytes, read \1)", + "fail": r"DANOS-TEST-RESULT: FAIL"}, # USB mass storage end to end: the boot usb-storage device (the FAT32 image, # which has a real 0x55AA boot sector) is enough — the manager spawns # usb-storage, which opens the device, runs the Bulk-Only / SCSI bring-up, diff --git a/tools/bounds-allowlist.txt b/tools/bounds-allowlist.txt index dd38ac1..9797ce2 100644 --- a/tools/bounds-allowlist.txt +++ b/tools/bounds-allowlist.txt @@ -76,7 +76,6 @@ system/drivers/usb-xhci-bus/usb-xhci-bus.zig:maker_buffer system/drivers/usb-xhci-bus/usb-xhci-bus.zig:prev_connected system/drivers/usb-xhci-bus/usb-xhci-bus.zig:product_buffer system/drivers/usb-xhci-bus/usb-xhci-bus.zig:product_buffer -system/drivers/usb-xhci-bus/usb-xhci-library.zig:blob system/drivers/usb-xhci-bus/usb-xhci-library.zig:buffer system/drivers/usb-xhci-bus/usb-xhci-library.zig:bytes system/drivers/usb-xhci-bus/usb-xhci-library.zig:data From 729b40ece7c4e0c513d9184666349afbe831a743 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 12:05:33 +0100 Subject: [PATCH 06/36] usb: a device has as many interfaces as it declares MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit max_interfaces was 4. A composite device — a headset, a webcam with audio, a dock, a multifunction printer — routinely has more, and the fifth did not merely go missing. parseConfiguration's cap branch had no `else`, so when the count was reached `current` kept pointing at interface 3 and the fifth interface's endpoint descriptors were appended to interface 3's array. A class driver bound to interface 3 could then be handed an endpoint belonging to something else entirely, and subscribe or bulk-transfer on it. The alternate-setting arm one line above cleared `current` correctly, which is what the cap branch should have done. Interfaces are now counted from the block in a first pass and allocated to exactly that number, so the ceiling is bNumInterfaces' u8 — the USB specification's. The missing `else` is added too, though after this the bug is unreachable by construction: interface_count cannot reach interfaces.len mid-parse when the list was sized from the same walk. max_configured_endpoints was max_interfaces * max_endpoints_per_interface = 16, a derived guess that moved whenever either input moved. It is now 31, which is the xHCI specification's own limit: a Device Context holds a slot context plus at most 31 endpoint contexts, because the Context Entries field addressing them is 5 bits. max_endpoints_per_interface stays at 4 with its reason recorded — the usb-transfer wire protocol reports exactly max_reported_endpoints (4) per interface, so widening it alone would change nothing a class driver sees. Lifting it is a protocol change. No direct test, and that is written down as open question 5 rather than glossed. The parser is pure and wants a host unit test, but usb-xhci-library.zig imports memory, mmio and time so it cannot be a standalone test root, and QEMU offers nothing that reaches the path — the largest device available is usb-audio,multi=on at 2 interfaces and 211 bytes. The alternate-setting path that shares the same `current = null` logic is exercised by that device. Suite 116/116. --- docs/bounds-track-plan.md | 23 ++++- .../drivers/usb-xhci-bus/usb-xhci-library.zig | 90 ++++++++++++++++--- tools/bounds-allowlist.txt | 2 - 3 files changed, 100 insertions(+), 15 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index c44794f..9c69f37 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -16,7 +16,7 @@ next one starts.* | L2 | Bounds build check + allowlist; declare what we have already touched | **done** — `zig build bounds`, 273 allowlisted, 5 declared | | L3 | xHCI: slot count from `HCSPARAMS1.MaxSlots`, not 8 | **done** — QEMU reports 64; the driver tracked 8 | | L4 | USB: configuration descriptor sized by `wTotalLength`, not 512 | **done** — QEMU tops out at 211 bytes, so the case catches the class, not the original trigger | -| L5 | USB: interfaces from the descriptor, and the misattributed-endpoint bug | not started | +| L5 | USB: interfaces from the descriptor, and the misattributed-endpoint bug | **done** — fix is by construction; no direct test, see open question 5 | | L6 | xHCI: a failed `allocateDevice` stops leaking an enabled slot | not started | **Suite:** 115/115 at the start of the run. @@ -73,6 +73,27 @@ question down instead of inventing an answer. answered first (tombstone-and-reuse aliases stale ids held by another process; generation-tagged ids change the id encoding, which is ABI). Not an unattended decision. +5. **Driver descriptor parsing cannot be host-tested, so L5's correctness fix ships + without a direct test.** The endpoint-misattribution bug lives in + `parseConfiguration`, a pure function over a byte blob — exactly the shape a host + unit test wants, and `usb-storage/scsi.zig` and `usb-hid/hid-report.zig` already do + this. But `usb-xhci-library.zig` imports `memory`, `mmio` and `time`, so it cannot + be a standalone host-test root, and QEMU offers no device that would exercise the + path anyway: the largest available is `usb-audio,multi=on` at 2 interfaces and 211 + bytes, against a cap of 4. + + Three ways out, and picking one is a judgement about house style rather than a + mechanical step: extract the parser to its own file and wire `usb-abi` into a test + module (build-support currently resolves module names only for `userBinary`); + extract it and import `usb-abi` by relative path (against the import-by-name + convention); or accept QEMU-only coverage and say so. + + Mitigating, and the reason this is recorded rather than blocking: after the fix the + bug is unreachable **by construction**, not by the added `else`. Interfaces are now + allocated to exactly the count the descriptor declares, so `interface_count` can + never reach `interfaces.len` mid-parse. The `else` is belt-and-braces for the + 255-interface clamp. The alternate-setting path that shares it *is* exercised — + `usb-audio` has alternate settings, and the `usb-large-descriptor` case walks them. ### Working rules for the run diff --git a/system/drivers/usb-xhci-bus/usb-xhci-library.zig b/system/drivers/usb-xhci-bus/usb-xhci-library.zig index 7e3d8c5..7f3a9dc 100644 --- a/system/drivers/usb-xhci-bus/usb-xhci-library.zig +++ b/system/drivers/usb-xhci-bus/usb-xhci-library.zig @@ -221,10 +221,17 @@ fn intervalFor(speed: u32, b_interval: u8) u32 { }; } -// Upper bounds on what one device's active configuration describes. A boot -// keyboard or mouse has one interface with one interrupt endpoint; a flash drive -// has one interface with two bulk endpoints. Generous for those. -pub const max_interfaces = 4; +/// bound: endpoints recorded per interface +/// decided-by: external +/// protects: the fixed endpoint array inside InterfaceInfo +/// at-limit: refuse — the surplus endpoint is not recorded and the class driver +/// cannot bind it +/// observed-by: the class driver failing to find the endpoint it wants +/// +/// Held at 4 because the usb-transfer wire protocol reports exactly +/// `max_reported_endpoints` (4) endpoints per interface, so widening this alone +/// would change nothing a class driver can see. Lifting it is a protocol change — +/// docs/bounds-track-plan.md, out of scope for the unattended run. pub const max_endpoints_per_interface = 4; // The endpoint-descriptor facts a class driver needs to talk to an endpoint: its @@ -256,7 +263,19 @@ const ConfiguredEndpoint = struct { dci: u32 = 0, ring: ProducerRing = .{ .region = .{ .virtual = 0, .physical = 0 } }, }; -const max_configured_endpoints = max_interfaces * max_endpoints_per_interface; +/// bound: transfer rings configured on one device slot +/// decided-by: hardware +/// protects: the per-device ring table, sized at compile time +/// at-limit: refuse — configureEndpoint returns null and its caller reports it +/// observed-by: the class driver's own failure to bind an endpoint +/// +/// The xHCI specification's number rather than ours: a Device Context carries a slot +/// context plus at most 31 endpoint contexts, because the Context Entries field that +/// addresses them is 5 bits (DCI 1..31, DCI 1 being the default control endpoint). +/// A device physically cannot present more. This was +/// `max_interfaces * max_endpoints_per_interface` = 16, so it moved whenever either +/// of those two guesses moved. +const max_configured_endpoints = 31; // One addressed USB device behind this controller: its hardware slot, its EP0 // (control) transfer ring, the DMA context + bounce buffer the control pipe uses, @@ -277,7 +296,13 @@ pub const Device = struct { device_descriptor: usb_abi.DeviceDescriptor = std.mem.zeroes(usb_abi.DeviceDescriptor), configuration_value: u8 = 0, interface_count: u8 = 0, - interfaces: [max_interfaces]InterfaceInfo = [_]InterfaceInfo{.{}} ** max_interfaces, + /// The interfaces of the active configuration, allocated at enumeration from the + /// count the configuration descriptor actually declares. This was a fixed 4, and + /// the fifth interface of a composite device (a headset, a webcam with audio, a + /// dock, a multifunction printer) did not merely go missing — see + /// parseConfiguration, where its endpoints were appended to interface 3's list. + interfaces: []InterfaceInfo = &.{}, + // Transfer rings configured for this device's interrupt/bulk endpoints. endpoint_ring_count: u8 = 0, endpoint_rings: [max_configured_endpoints]ConfiguredEndpoint = [_]ConfiguredEndpoint{.{}} ** max_configured_endpoints, @@ -296,6 +321,15 @@ pub const Device = struct { // Downstream ports with a pending change to service (bit P = port P), set // by the status-change endpoint (and by an initial sweep in setupHub). hub_change_mask: u32 = 0, + + /// Release the interface list. Safe to call twice, and on a device that never + /// enumerated — re-enumeration and teardown both come through here. + pub fn freeInterfaces(self: *Device) void { + if (self.interfaces.len != 0) memory.allocator().free(self.interfaces); + self.interfaces = &.{}; + self.interface_count = 0; + } + }; // A standing interrupt-IN subscription: the endpoint's ring is kept armed with a @@ -1162,6 +1196,7 @@ pub const Controller = struct { fn abandon(self: *Controller, device: *Device) ?*Device { _ = self; + device.freeInterfaces(); device.used = false; return null; } @@ -1329,7 +1364,7 @@ pub const Controller = struct { // Both numbers, so a truncation can never again be invisible: the block the // device declared, and the bytes actually fetched and parsed. They must match. std.log.info("config block {d} bytes, read {d}", .{ configuration.total_length, blob.len }); - parseConfiguration(device, blob); + if (!parseConfiguration(device, blob)) return false; // Select the configuration, moving the device to the configured state. if (!self.controlTransfer(device, usb_abi.setConfiguration(configuration.configuration_value), &.{}, false)) return false; @@ -1340,8 +1375,30 @@ pub const Controller = struct { // and the endpoints that follow it. Endpoints belong to the most recent // interface. Unknown descriptor types (HID, class-specific) are skipped by // their length. - fn parseConfiguration(device: *Device, blob: []const u8) void { - device.interface_count = 0; + /// Two passes: count the alternate-setting-0 interfaces the block declares, + /// allocate exactly that many, then fill them. False only on an allocation + /// failure. The count cannot exceed 255 — `bNumInterfaces` is a u8, so that is + /// the USB specification's ceiling and not one of ours. + fn parseConfiguration(device: *Device, blob: []const u8) bool { + var declared: usize = 0; + var count_offset: usize = 0; + while (count_offset + 2 <= blob.len) { + const length = blob[count_offset]; + if (length < 2 or count_offset + length > blob.len) break; + if (@as(usb_abi.DescriptorType, @enumFromInt(blob[count_offset + 1])) == .interface and + length >= @sizeOf(usb_abi.InterfaceDescriptor)) + { + const descriptor = std.mem.bytesToValue(usb_abi.InterfaceDescriptor, blob[count_offset .. count_offset + @sizeOf(usb_abi.InterfaceDescriptor)]); + if (@intFromEnum(descriptor.alternate_setting) == 0 and declared < 255) declared += 1; + } + count_offset += length; + } + + device.freeInterfaces(); + if (declared == 0) return true; + device.interfaces = memory.allocator().alloc(InterfaceInfo, declared) catch return false; + for (device.interfaces) |*interface| interface.* = .{}; + var current: ?*InterfaceInfo = null; var offset: usize = 0; while (offset + 2 <= blob.len) { @@ -1351,9 +1408,17 @@ pub const Controller = struct { switch (@as(usb_abi.DescriptorType, @enumFromInt(descriptor_type))) { .interface => if (length >= @sizeOf(usb_abi.InterfaceDescriptor)) { const descriptor = std.mem.bytesToValue(usb_abi.InterfaceDescriptor, blob[offset .. offset + @sizeOf(usb_abi.InterfaceDescriptor)]); - if (@intFromEnum(descriptor.alternate_setting) != 0) { - current = null; // ignore alternate settings for now - } else if (device.interface_count < max_interfaces) { + // An interface we do not record MUST clear `current`, or the + // endpoints that follow it attach to the previous one. The cap + // branch used to have no `else`, so a fifth interface's endpoints + // were appended to interface 3's array and a class driver bound to + // interface 3 could be handed an endpoint belonging to something + // else entirely — silently, with a truthful-looking count logged. + if (@intFromEnum(descriptor.alternate_setting) != 0 or + device.interface_count >= device.interfaces.len) + { + current = null; + } else { const slot = &device.interfaces[device.interface_count]; slot.* = .{ .number = @intFromEnum(descriptor.interface_number), @@ -1383,6 +1448,7 @@ pub const Controller = struct { } offset += length; } + return true; } // --- endpoint configuration + interrupt / bulk transfers --------------- diff --git a/tools/bounds-allowlist.txt b/tools/bounds-allowlist.txt index 9797ce2..3eea7a4 100644 --- a/tools/bounds-allowlist.txt +++ b/tools/bounds-allowlist.txt @@ -82,8 +82,6 @@ system/drivers/usb-xhci-bus/usb-xhci-library.zig:data system/drivers/usb-xhci-bus/usb-xhci-library.zig:descriptor system/drivers/usb-xhci-bus/usb-xhci-library.zig:head system/drivers/usb-xhci-bus/usb-xhci-library.zig:header -system/drivers/usb-xhci-bus/usb-xhci-library.zig:max_endpoints_per_interface -system/drivers/usb-xhci-bus/usb-xhci-library.zig:max_interfaces system/drivers/usb-xhci-bus/usb-xhci-library.zig:max_subscriptions system/drivers/usb-xhci-bus/usb-xhci-library.zig:port_changes system/drivers/usb-xhci-bus/usb-xhci-library.zig:raw From 5d55217212adc5eea7af1d7262aea02658b0b093 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 12:16:42 +0100 Subject: [PATCH 07/36] usb: a slot the controller granted is always handed back MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Both device-setup paths issued a successful Enable Slot and then returned null if allocateDevice failed, without disabling it. A slot the driver forgets is one the controller never reissues, so each attempt lost one permanently for the boot. The hub path did it with no log line at all. Both now release the slot through a shared disableSlot, extracted from tearDownDevice, and the hub path warns like the root-port path does. tearDownDevice also now frees the interface list. That allocation arrived with the previous commit, so an unplug would have leaked it — found while reading the teardown path for this fix rather than by a test. No regression test, and it is recorded as open question 5 rather than implied. After the slot count became the controller's own figure, reaching this path needs more devices than the controller has slots: QEMU offers four against sixty-four. What was verified is that the new path RUNS correctly — pinning tracking to 2 with four devices attached produced "port 6 setup: no free device slot", the first two devices enumerated normally, and no Disable Slot error or timeout appeared, which is how disableSlot reports failure. Suite 116/116. --- docs/bounds-track-plan.md | 7 ++++- .../drivers/usb-xhci-bus/usb-xhci-library.zig | 31 ++++++++++++++----- 2 files changed, 30 insertions(+), 8 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 9c69f37..90d34c5 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -17,7 +17,12 @@ next one starts.* | L3 | xHCI: slot count from `HCSPARAMS1.MaxSlots`, not 8 | **done** — QEMU reports 64; the driver tracked 8 | | L4 | USB: configuration descriptor sized by `wTotalLength`, not 512 | **done** — QEMU tops out at 211 bytes, so the case catches the class, not the original trigger | | L5 | USB: interfaces from the descriptor, and the misattributed-endpoint bug | **done** — fix is by construction; no direct test, see open question 5 | -| L6 | xHCI: a failed `allocateDevice` stops leaking an enabled slot | not started | +| L6 | xHCI: a failed `allocateDevice` stops leaking an enabled slot | **done** — path forced and verified; no regression test, see open question 5 | + +**Run complete.** L1 stopped (the step was wrong), L2–L6 landed. Suite 115 → 116. +Allowlist 278 → 269. Two steps ship without a permanent regression test, both because +QEMU's USB devices are too small to reach the paths — see open question 5, which is the +audit's own lesson recurring: the test rig is smaller than a real machine. **Suite:** 115/115 at the start of the run. **Branch:** `claude/bounds-track`. diff --git a/system/drivers/usb-xhci-bus/usb-xhci-library.zig b/system/drivers/usb-xhci-bus/usb-xhci-library.zig index 7f3a9dc..9d0186b 100644 --- a/system/drivers/usb-xhci-bus/usb-xhci-library.zig +++ b/system/drivers/usb-xhci-bus/usb-xhci-library.zig @@ -943,7 +943,8 @@ pub const Controller = struct { return null; }; const device = self.allocateDevice() orelse { - std.log.info("port {d} setup: no free device slot", .{port}); + std.log.warn("port {d} setup: no free device slot", .{port}); + self.disableSlot(slot_id); // the controller already gave it to us return null; }; device.* = .{ @@ -1145,7 +1146,11 @@ pub const Controller = struct { std.log.info("hub slot {d} port {d}: Enable Slot failed", .{ hub.slot_id, port }); return null; }; - const device = self.allocateDevice() orelse return null; + const device = self.allocateDevice() orelse { + std.log.warn("hub slot {d} port {d}: no free device slot", .{ hub.slot_id, port }); + self.disableSlot(slot_id); // the controller already gave it to us + return null; + }; const child_speed = mapHubPortSpeed(speed); device.* = .{ .used = true, @@ -1747,15 +1752,27 @@ pub const Controller = struct { for (&self.subscriptions) |*subscription| { if (subscription.active and subscription.slot_id == device.slot_id) subscription.active = false; } + self.disableSlot(device.slot_id); + device.freeInterfaces(); + device.used = false; + } + + /// Hand a slot back to the controller and clear its context-array entry. + /// + /// Every path that has issued a successful Enable Slot owes this, including the + /// ones that then fail to bring the device up. A slot the driver forgets is one + /// the controller never reissues, so the loss is permanent for the boot: the two + /// setup paths used to return null straight after a failed `allocateDevice`, + /// leaking a slot per attempt — and the hub path did it without even a log line. + fn disableSlot(self: *Controller, slot_id: u8) void { const physical = self.submitCommand(.{ - .control = trbControl(.disable_slot, @as(u32, device.slot_id) << 24), + .control = trbControl(.disable_slot, @as(u32, slot_id) << 24), }); if (self.awaitCommand(physical)) |code| { if (code != @intFromEnum(CompletionCode.success)) - std.log.info("slot {d}: Disable Slot completion code {d}", .{ device.slot_id, code }); - } else std.log.info("slot {d}: Disable Slot timed out", .{device.slot_id}); + std.log.info("slot {d}: Disable Slot completion code {d}", .{ slot_id, code }); + } else std.log.info("slot {d}: Disable Slot timed out", .{slot_id}); const array: [*]volatile u64 = @ptrFromInt(self.device_context_array.virtual); - array[device.slot_id] = 0; - device.used = false; + array[slot_id] = 0; } }; From 3b24b541b09f570e7c5809f8b82c683f6ff70f30 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 12:58:19 +0100 Subject: [PATCH 08/36] =?UTF-8?q?docs:=20phase=202=20design=20=E2=80=94=20?= =?UTF-8?q?you=20hold=20what=20you=20were=20given?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit device_claim checks that a device exists and is free. That is all. Any process may claim any unclaimed device, and a claim is what gates mmio_map and irq_bind — a licence to map physical memory and take interrupts. The device manager's matching is real but advisory: it spawns a driver with the device id in argv[1] and nothing binds that decision to the kernel's grant. maximum_children_per_parent is the visible cost. It exists because a driver that claimed one device could loop device_register under it, and it is a poor defence — an attacker burns 16 slots, claims another device, burns 16 more — while reliably refusing a legitimate PCI bus with more than 16 functions. Closing the hole is what retires the constant. The principles decide the split: matching is policy and stays in the device manager; enforcing that a driver holds only what it was given is security and stays in the kernel, and is the whole of what the kernel needs. The mechanism is decided by an awkward fact. Five of six claimants are device-manager children spawned with their device id. display is not — init spawns it, and it finds its framebuffer by enumerating for a display-class node and claiming whatever it finds. So "record the device named at spawn" closes the hole for five and breaks the sixth, and the sixth is not an oddity to special-case: it shows authority must be delegable rather than welded to the moment of spawn. So: a device grant is a capability, minted by the kernel to init for the devices firmware discovery found, delegated by init to the device manager and to display, and passed by the manager to each driver it spawns. The cap-passing path already exists and already carries shared memory and DMA regions. init is already the grantor for /protocol, and protocol.csv already records init granting the device manager its binding. Costs named rather than buried: five drivers must hello before claiming (pci-bus claims first today), and init grows a device role on top of the protocol registry. A device-grant manifest mirroring protocol.csv would be the natural symmetry and is deliberately not proposed yet. --- docs/os-development/device-authority.md | 247 ++++++++++-------------- 1 file changed, 106 insertions(+), 141 deletions(-) diff --git a/docs/os-development/device-authority.md b/docs/os-development/device-authority.md index 7e49a8f..608be65 100644 --- a/docs/os-development/device-authority.md +++ b/docs/os-development/device-authority.md @@ -1,173 +1,138 @@ -# Device authority: the kernel stops keeping an inventory +# Device authority: you hold what you were given -*Design, drafted 2026-08-07. Not implemented. Prompted by a real machine: an AMD -Ryzen desktop enumerated more PCI functions than the kernel's device table would -hold, and the xHCI and SATA controllers were refused registration — so the -machine booted to the compositor with no USB and no storage.* +*Design for phase 2 of [the bounds track](../bounds-track-plan.md), 2026-08-08. +Supersedes an earlier draft that argued for capabilities on aesthetic grounds; this one +starts from the hole and from the project's principles.* -The kernel keeps a table of every device userspace discovers. It is a fixed -array of 64 descriptors, 344 bytes each, and a second cap allows any one parent -16 children. Neither number is written down anywhere as a decision: -`maximum_devices` has no comment and never reached `parameters.zig`, where every -other tunable in this kernel lives with its reasoning attached. +## The hole -Raising them is not the fix. The numbers are wrong because the *table* is wrong: -it is an inventory of hardware, and an inventory of hardware is not something a -kernel needs. This document proposes replacing it with capabilities, which -removes the ceiling rather than moving it. +`device_claim(id)` checks two things ([devices-broker.zig](../../system/kernel/devices-broker.zig)): -## What the kernel actually uses +```zig +pub fn claim(id: u64, owner: u32) ClaimError!void { + if (id >= count) return error.NoSuchDevice; + if (claimed[@intCast(id)] != null) return error.AlreadyClaimed; + claimed[@intCast(id)] = owner; +} +``` -Every read of a device descriptor from the kernel proper, exhaustively: +Does it exist, and is it free. **Any process may claim any unclaimed device.** -| Used for | What it needs | -|---|---| -| `mmio_map`, `io_read`/`io_write` | the physical range, to check the mapping falls inside it | -| `irq_bind`, `msi_bind` | the GSI, and that the caller owns the device | -| `dma_bind` | that the caller owns the device | -| IOMMU confinement | the **PCI BDF**, to key a domain | -| `device_register` | the parent's ranges, for the containment check | -| the boot display seed | one framebuffer window | +The matching is real but it is entirely advisory: `device-manager` reads `devices.csv`, +matches a device to a driver, and spawns that driver with the device id as `argv[1]` +(`spawnDriver`). The driver parses the string and claims it. Nothing anywhere binds the +manager's decision to the kernel's grant — a process can pass any integer and win the +race. -That is: **physical ranges, interrupt numbers, and one BDF.** Vendor, device and -subsystem ids, class triples, the human-readable names, the parent links, the -bus numbers — the kernel stores all of it and reads none of it. It is held so -that `device_enumerate` can hand it back to user space, which is the whole -mistake in one sentence: the kernel is acting as a distribution mechanism for -data it does not use. +A claim is not a small thing. It is what gates `mmio_map` and `irq_bind`, so it is a +licence to map physical memory and receive interrupts. -## The split +## What this costs, beyond the obvious -Three concerns are tangled in one table. +`maximum_children_per_parent = 16` exists because a driver that claimed one device could +loop `device_register` under it and exhaust the shared table. That threat only exists +*because* claiming is unauthenticated — and the cap is a poor defence against it, since +an attacker can burn 16 slots, claim another device, and burn 16 more. What it reliably +does instead is refuse a legitimate PCI bus with more than 16 functions, which is how an +AMD Ryzen came to boot with no USB and no storage. -**Platform bring-up** — timers, LAPIC/IOAPIC, CPU topology, the ECAM window, the -framebuffer the firmware left. The kernel derives these from ACPI before user -space exists and needs them to function. They never belonged in the device table -and mostly are not (the platform block in `kernel.zig` is separate); this -document does not change them. +So the cap is not merely mis-sized. It is standing in for an authorisation that is not +performed, and it punishes correct behaviour while barely inconveniencing incorrect +behaviour. **Closing the hole is what retires the constant**, not a bigger number. -**Resource authority** — which task may map which physical range, receive which -interrupt, touch which ports. This *must* stay in the kernel. It is the one -grant that cannot be audited after the fact: a process that maps arbitrary -physical memory owns the machine, page tables and IOMMU structures included. -This is memory protection, not device management, and it is why the answer is -not simply "move it all to the device manager". +## What the principles decide -**Device inventory** — what exists, what it is, how it is arranged, which driver -should bind it. This is `device-manager`'s job and is already half there: it -loads `devices.csv`, matches, spawns drivers, and receives `child_added` reports -over its own protocol. The kernel table duplicates what those reports already -carry. +- *Move as much responsibility as possible to user space* (3), and *what remains in the + kernel is there for security or a hardware limitation* (5). -## The proposal: a resource is a capability +Deciding **which** driver gets **which** device is policy: it reads a CSV, matches +identity triples, and picks a binary. That is the device manager's, and it stays there. -Device resources join endpoints, shared memory and DMA regions as a kind in the -handle table. +Enforcing that a driver **holds only what it was given** is security — it is the gate in +front of mapping physical memory. That stays in the kernel, and it is the whole of what +the kernel needs to do. -1. **Roots.** At boot the kernel mints capabilities for the windows it learned - from firmware — the ECAM range, the framebuffer, the legacy port space — and - hands them to the first bus drivers. This is the only place device knowledge - enters the kernel, and it comes from ACPI, not from a driver's say-so. -2. **Subdivision.** A bus driver enumerating hardware derives a narrower - capability from one it holds: `resource_derive(cap, kind, start, len) → cap`. - The kernel checks the sub-range lies inside the capability being subdivided — - the same containment rule as today (`devices-broker.contains`), but checked - against *one capability the caller demonstrably holds* rather than by walking - a global tree. -3. **Delegation.** The driver passes that capability to the child driver over - IPC. Cap-passing already exists; this is the mechanism `subscribe` and - `attach_scanout` already use. -4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and - `dma_bind` take a capability handle instead of `(device_id, resource_index)`. - Possession *is* the authority — there is nothing to look up and no ownership - table to consult. +The kernel therefore does not need to know about matching, `devices.csv`, driver names, +or why a device was assigned. It needs to know that an authority it can verify granted +this device to this task. -Exclusivity stops being a broker refusing a second claimant and becomes the -ordinary property of a capability: only one process was given it. +## The awkward fact that decides the mechanism -## What this buys +Five of the six claimants are device-manager children, spawned with their device id in +`argv[1]`: `pci-bus`, `usb-xhci-bus`, `ps2-bus`, `virtio-gpu`, `acpi`. -**No ceiling.** There is no table to size, so no machine is too big. The -Ryzen's enumeration stops being a limit to tune and becomes what it is — a fact -about a computer. +**`display` is not.** It is spawned by `init` from `init.csv`, and it finds its device by +enumerating the table for a display-class node and claiming whatever it finds +([backend.zig](../../system/services/display/backend.zig)). There is no assignment to +enforce, because nobody assigned it anything. -**Reclamation, free.** `count` in the broker today only ever increases; -`releaseAllOwnedBy` clears a dead driver's *claims* but never its entries. A -driver that crashes and is restarted re-registers its children and consumes the -table again — reachable today without any malice, given the device manager -restarts drivers by design. Capabilities die with the task. +That rules out the cheapest design. "The kernel records the device named at spawn, and +`device_claim` checks it" closes the hole for five claimants and breaks the sixth. And +the sixth is not an oddity to special-case — it is the one that shows the model is +wrong: authority should be *delegable*, not welded to the moment of spawn. -**No quota needed.** The per-parent cap exists to stop one claimant looping -`device_register` and filling the shared table, because a zero-resource child -sidesteps the containment check. With no shared table there is nothing to -exhaust; a process can only ever subdivide what it was given, and its handles -are already bounded per task. +## The design -**A smaller kernel.** Three syscalls leave the ABI, one narrower one arrives, -and `devices-broker.zig` largely disappears along with both constants. +**A device grant is a capability, delegated from a holder.** The mechanism already +exists: `callCap` passes a handle over an IPC call and the kernel installs it in the +receiver's table ([library/kernel/ipc.zig](../../library/kernel/ipc.zig)), which is how +shared memory and DMA regions already move between processes. -**The discipline the rest of the system already uses.** "The claim is the -capability" is written in the driver documentation as if it were already true. -This makes it true. +1. **Root.** At boot the kernel mints grants for the devices firmware discovery found + and hands them to `init` (PID 1, which the kernel spawns and therefore need not + authenticate). This is the only place device authority enters the system, and it + comes from ACPI rather than from anyone's say-so. +2. **Delegation.** `init` passes the device manager the grants it will need — in + practice all of them — and passes `display` the framebuffer grant, because `init` is + what starts `display`. This is the same shape as the `/protocol` registry, where + `init` is already the grantor and `protocol.csv` already records + `/system/services/device-manager, /system/services/init, bind, device-manager`. +3. **Assignment.** The manager passes a driver its device when it spawns it, over the + channel that already exists — the driver `hello`s the manager, and the reply carries + the grant. +4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and `dma_bind` + check possession of the grant instead of consulting an ownership table. -## Syscall surface +Exclusivity stops being a broker refusing a second claimant and becomes the ordinary +property of a capability: only one process was given it. -- `device_enumerate` — **retires.** It exists only to read the kernel's table. - Callers ask `device-manager`, whose protocol already reserves an `enumerate` - verb. Note this is a public-ABI change: `vdso.md` documents it. -- `device_register` — **splits.** The kernel half becomes `resource_derive`; the - publication half ("this device exists, here is what it is") becomes an IPC - message to `device-manager`, which is where the inventory belongs and where - `child_added` already carries the same facts. -- `device_claim` — **dissolves into possession**, except for the IOMMU (below). -- `mmio_map`, `irq_bind`, `msi_bind`, `io_read`, `io_write`, `dma_bind` — keep - their names and semantics; their first argument becomes a capability handle. +**`maximum_children_per_parent` is deleted here.** After this a bus driver's children are +devices it enumerated on a bus it was actually given, and the threat the cap was written +for no longer exists. -## The open question: where the IOMMU attaches +## What this costs -This is the one place the kernel still needs device *identity* rather than a -range. `confineDevice(device_id, bdf, owner)` builds a domain keyed by PCI BDF, -attaches it, and the claim is rolled back if confinement fails — deliberately, -so a device that cannot be confined is never driven. +**Bring-up order changes.** `pci-bus` today claims first and says hello afterwards — its +own comment says "Claim the bridge, map the ECAM, hello the manager, then scan." Under +delegation the hello must come first, because that is where the grant arrives. Five +drivers need that reordering, and it is the bulk of the work. -Three options, none obviously right: +**`init` grows a device role.** It already registers `/protocol` and reads +`protocol.csv`; it would also hold root device grants and hand them on. That is more +responsibility in PID 1, which is a cost worth naming — though the alternative is the +kernel deciding who may hold what, which principle 5 excludes. -1. **The BDF rides the capability.** A memory capability derived for a PCI - function carries its BDF, and the kernel confines on first `mmio_map` or - `dma_bind`. Keeps the syscall count down; means a capability is no longer - purely a range. -2. **An explicit `device_attach(cap, bdf)`.** Honest and visible, but it puts a - BDF — a fact about PCI — into a kernel interface that otherwise knows nothing - about buses, and something must stop a caller naming a BDF that is not - theirs. -3. **The root PCI capability carries the segment, and derivation computes the - BDF.** Purest, but only works for PCI and the kernel would be parsing bus - topology, which is precisely what this document is trying to stop. +**A configuration question follows.** `/protocol` grants are data (`protocol.csv`), +because who may speak to whom is policy. Device grants could be too — a manifest saying +which binary may be given which device class. That would be the natural symmetry, and it +is *not* proposed here: `init` handing the manager everything and the manager matching +by `devices.csv` is the smaller step, and the manifest can follow if it earns itself. -My inclination is (1), because confinement is a property of the resource being -granted rather than a separate action, and because it keeps the "possession is -authority" story intact. It needs the derivation call to know it is carving a -PCI function, which is a wart worth arguing about. +## What it does not solve -## What this does not solve +- **Hot-unplug and re-enumeration drift.** Still the inventory problem (phase 3, and + open question 4 in the track plan). A grant dying with its holder is not the same as a + device going away. +- **The device table's size.** `maximum_devices` is untouched by this; it goes when the + inventory moves in phase 3. +- **Two processes racing for the same root grant.** Cannot arise, because roots are + minted to `init` alone. -- **Hot-plug and removal.** Capabilities die with their holder, but a device - that physically disappears while a driver lives is the device manager's - problem and unchanged by this. -- **Quotas on physical memory.** Nothing here bounds how much a driver maps; it - bounds only *what* it may map. That was already true. -- **The device tree as a published thing.** `/system/devices` remains a device - manager concern, as the file-system hierarchy already assumes. +## How this is verified -## Migration sketch - -Not a plan yet — the phases need sizing once the IOMMU question is settled. -Rough shape: introduce the capability kind and `resource_derive` alongside the -existing table; convert the resource-consuming syscalls to accept either form; -move the inventory into `device-manager` and convert its clients off -`device_enumerate`; then delete the table, the two constants and the three -syscalls in one flag-day, as the `ServiceId` retirement did. - -Every driver is affected, so the QEMU suite is the arbiter at each step, and the -Ryzen is the acceptance test — it is the machine that found this, and the one -that proves it fixed. +The invariant is **I3** from the track plan: a process holds what it was handed and +cannot name its way into holding more. The test is adversarial and the suite has never +had one of these for devices: a process that was granted nothing calls `device_claim` +on a device another driver owns, and on one nobody owns, and is refused both times with +its own errno. The audit's lesson was that "the suite contains no attacker"; this is the +attacker for devices. From 139c4f624fa176f8757fbca6893f1f3c4c3ece15 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 16:31:50 +0100 Subject: [PATCH 09/36] docs: mechanism, policy and configuration are three things MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The phase 2 design called devices.csv "policy". It is not. The CSV is configuration — declarative data an operator edits. The policy is the component that decides using it: the device manager matching a device to a driver, init reading protocol.csv and granting a binding. The kernel holds mechanism, the check that a grant exists before a mapping is made. Corrected where the document conflated them, and the three are now tabulated so the rest of the design can lean on the distinction. It also sharpens the open question. It was "manifest or not"; it is really whether the grant list should be configuration from the start, or whether init should hand the device manager the root grants and let it match by the devices.csv it already reads. A device-grant file adds a second place an operator must keep correct, and earns itself only when something needs to differ from "the manager gets the hardware" — holding a device back for a test or a bare-metal driver, say. --- docs/os-development/device-authority.md | 28 ++++++++++++++++++------- 1 file changed, 21 insertions(+), 7 deletions(-) diff --git a/docs/os-development/device-authority.md b/docs/os-development/device-authority.md index 608be65..7e1ef3d 100644 --- a/docs/os-development/device-authority.md +++ b/docs/os-development/device-authority.md @@ -45,8 +45,16 @@ behaviour. **Closing the hole is what retires the constant**, not a bigger numbe - *Move as much responsibility as possible to user space* (3), and *what remains in the kernel is there for security or a hardware limitation* (5). -Deciding **which** driver gets **which** device is policy: it reads a CSV, matches -identity triples, and picks a binary. That is the device manager's, and it stays there. +Deciding **which** driver gets **which** device is policy — matching identity triples +and choosing a binary. That decision belongs to the device manager and stays there. +`devices.csv` is not the policy; it is **configuration**, the declarative data the +policy reads. Three distinct things, and worth keeping apart in this document: + +| | Lives in | Example | +|---|---|---| +| **Mechanism** | the kernel | the check that a grant is held before a mapping is made | +| **Policy** | user space | the device manager matching a device to a driver | +| **Configuration** | files | `devices.csv`, `protocol.csv`, `init.csv` | Enforcing that a driver **holds only what it was given** is security — it is the gate in front of mapping physical memory. That stays in the kernel, and it is the whole of what @@ -112,11 +120,17 @@ drivers need that reordering, and it is the bulk of the work. responsibility in PID 1, which is a cost worth naming — though the alternative is the kernel deciding who may hold what, which principle 5 excludes. -**A configuration question follows.** `/protocol` grants are data (`protocol.csv`), -because who may speak to whom is policy. Device grants could be too — a manifest saying -which binary may be given which device class. That would be the natural symmetry, and it -is *not* proposed here: `init` handing the manager everything and the manager matching -by `devices.csv` is the smaller step, and the manifest can follow if it earns itself. +**A configuration question follows.** Who may bind which protocol name is expressed as +configuration today — `protocol.csv` — with `init` as the policy that reads it. Device +grants could be expressed the same way: a file saying which binary may be given which +device class, with `init` again the policy that enforces it. + +That symmetry is real but it is *not* proposed here. The smaller step is `init` handing +the manager the root grants and the manager matching by `devices.csv`, which is +configuration it already reads. A device-grant file adds a second place an operator must +keep correct, and it should only arrive when something needs it to differ from "the +manager gets the hardware" — for example, holding a device back from the manager so a +test or a bare-metal driver can take it. ## What it does not solve From ef793d1320aa787fc3193c6028d5c73cbad702a0 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 16:35:32 +0100 Subject: [PATCH 10/36] docs: the device manager is the root of device authority, not init MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The design had init minting device grants and handing them on, because init is what starts display and display claims the framebuffer. That was the wrong shape. init has nothing to do with devices: it is the first process and it starts the rest of the system. The device manager owns hardware. So the kernel mints the root grants to the device manager, and display asks the manager for its framebuffer like any other driver. The exception that drove the earlier draft disappears instead of being accommodated — one authority for hardware, not two. The kernel recognises the manager by the chain attestation the security track already settled on: the binary path it stamped itself, AND a supervisor of PID 1. Path alone is forgeable, because system_spawn is deliberately ungated and any process may spawn any bundled binary — but a rogue copy's supervisor is the rogue, and PID 1 is the kernel's own first process. Costs restated for the corrected model. display gains a boot-order dependency on the manager, mitigated by its existing ability to attach a backend late. The kernel gains one piece of knowledge about one binary, which is the minimum: authority has to enter somewhere, and the alternatives are a file parser in the kernel or first-to-ask, which is the hole again. --- docs/os-development/device-authority.md | 77 +++++++++++++++++-------- 1 file changed, 52 insertions(+), 25 deletions(-) diff --git a/docs/os-development/device-authority.md b/docs/os-development/device-authority.md index 7e1ef3d..8a847f0 100644 --- a/docs/os-development/device-authority.md +++ b/docs/os-development/device-authority.md @@ -79,6 +79,13 @@ That rules out the cheapest design. "The kernel records the device named at spaw the sixth is not an oddity to special-case — it is the one that shows the model is wrong: authority should be *delegable*, not welded to the moment of spawn. +It is also what tempted an earlier draft of this document into giving `init` the root +grants, since `init` is what starts `display`. That was wrong for a plainer reason: +`init` has nothing to do with devices. It is the first process and it starts the rest of +the system; the device manager is what owns hardware. Under delegation `display` simply +asks the manager, like every other driver, and the exception disappears rather than +being accommodated. + ## The design **A device grant is a capability, delegated from a holder.** The mechanism already @@ -86,18 +93,33 @@ exists: `callCap` passes a handle over an IPC call and the kernel installs it in receiver's table ([library/kernel/ipc.zig](../../library/kernel/ipc.zig)), which is how shared memory and DMA regions already move between processes. +**The device manager is the root of device authority.** `init` has no part in this: it +is the first process, and its job is to start the rest of the system. It starts the +device manager the same way it starts everything else, and knows nothing about devices. + 1. **Root.** At boot the kernel mints grants for the devices firmware discovery found - and hands them to `init` (PID 1, which the kernel spawns and therefore need not - authenticate). This is the only place device authority enters the system, and it - comes from ACPI rather than from anyone's say-so. -2. **Delegation.** `init` passes the device manager the grants it will need — in - practice all of them — and passes `display` the framebuffer grant, because `init` is - what starts `display`. This is the same shape as the `/protocol` registry, where - `init` is already the grantor and `protocol.csv` already records - `/system/services/device-manager, /system/services/init, bind, device-manager`. -3. **Assignment.** The manager passes a driver its device when it spawns it, over the - channel that already exists — the driver `hello`s the manager, and the reply carries - the grant. + and hands them to the device manager. This is the only place device authority enters + the system, and it comes from ACPI rather than from anyone's say-so. + + The kernel recognises the manager by **chain attestation**, the identity the security + track already settled on: the binary path the kernel itself stamped + (`/system/services/device-manager`) *and* a supervisor of PID 1. The binary path + alone would not do — `system_spawn` is deliberately ungated, so any process may spawn + any bundled binary, and a rogue could run a second copy under the same name. It could + not forge the other half: its copy's supervisor is the rogue, and PID 1 is the + kernel's own first process. + +2. **Delegation.** The manager passes a driver its device when it spawns it, over the + channel that already exists — the driver `hello`s the manager and the reply carries + the grant. The manager attests its own children by task id, one hop deep, exactly as + `init` attests the services it spawned for `/protocol` (`supervisorSatisfies`); it + knows which task is which driver because it spawned them. + +3. **`display` asks the manager too.** It is started by `init` as a service, but the + framebuffer is a device, so it receives that grant from the device manager like any + other driver. One authority for hardware, not two — which is what makes `display` + stop being the exception that broke the simpler design. + 4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and `dma_bind` check possession of the grant instead of consulting an ownership table. @@ -115,22 +137,27 @@ own comment says "Claim the bridge, map the ECAM, hello the manager, then scan." delegation the hello must come first, because that is where the grant arrives. Five drivers need that reordering, and it is the bulk of the work. -**`init` grows a device role.** It already registers `/protocol` and reads -`protocol.csv`; it would also hold root device grants and hand them on. That is more -responsibility in PID 1, which is a cost worth naming — though the alternative is the -kernel deciding who may hold what, which principle 5 excludes. +**`display` gains a dependency on the device manager.** It currently finds its +framebuffer by itself and needs nothing from anyone; afterwards it must ask. That is a +new ordering constraint at boot — `display` cannot bring up a screen until the manager +is up — and it is the honest price of there being one authority for hardware. Mitigating: +`display` already tolerates arriving before its backend (the virtio-gpu driver hands it +a scanout later, over `attach_scanout`), so the machinery for "wait, then attach" is +there. -**A configuration question follows.** Who may bind which protocol name is expressed as -configuration today — `protocol.csv` — with `init` as the policy that reads it. Device -grants could be expressed the same way: a file saying which binary may be given which -device class, with `init` again the policy that enforces it. +**The kernel gains one piece of knowledge about a specific binary.** Chain attestation +means the kernel recognises `/system/services/device-manager` under PID 1 as the root +holder. That is a real concession — the kernel would rather know nothing about who is +who — and it is the minimum: device authority has to enter the system somewhere, and +every alternative is worse. Configuration the kernel reads would put a file parser in +the kernel; first-to-ask would be the hole again, at boot. -That symmetry is real but it is *not* proposed here. The smaller step is `init` handing -the manager the root grants and the manager matching by `devices.csv`, which is -configuration it already reads. A device-grant file adds a second place an operator must -keep correct, and it should only arrive when something needs it to differ from "the -manager gets the hardware" — for example, holding a device back from the manager so a -test or a bare-metal driver can take it. +**A configuration question stays open.** Which binary may be given which device is +expressed today as `devices.csv`, read by the manager — configuration, with the manager +as the policy that acts on it. That is already the right shape and needs no new file. +A separate grant manifest earns itself only when something must differ from "the manager +gets the hardware": holding a device back so a test or a bare-metal driver can take it, +for instance. Not now. ## What it does not solve From 8e388e048a102462e3d9693151ad6c4d47ddefb1 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 16:35:54 +0100 Subject: [PATCH 11/36] docs: one line still had init minting the root grants --- docs/os-development/device-authority.md | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/docs/os-development/device-authority.md b/docs/os-development/device-authority.md index 8a847f0..f6f7c55 100644 --- a/docs/os-development/device-authority.md +++ b/docs/os-development/device-authority.md @@ -166,8 +166,9 @@ for instance. Not now. device going away. - **The device table's size.** `maximum_devices` is untouched by this; it goes when the inventory moves in phase 3. -- **Two processes racing for the same root grant.** Cannot arise, because roots are - minted to `init` alone. +- **Two processes racing for the same root grant.** Cannot arise: the roots are minted + to whichever task satisfies the chain — the manager's binary under PID 1 — and a + second copy spawned by anyone else fails the supervisor half. ## How this is verified From 5ad42ceac346e5afefbcf50bf3a15c63afe23879 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 16:40:37 +0100 Subject: [PATCH 12/36] docs: only drivers hold devices; services speak protocols MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The design still had the display service receiving a device grant, which keeps the wrong layering and just moves who hands it over. Drivers talk hardware. Services talk to no hardware at all — they receive device events and send data and commands over a protocol, which is the OS abstraction between them. So the answer to "which services need device grants" is none, and the design collapses: every holder of a device is a driver, and every driver is spawned by the device manager, which is what grants it. No exceptions to accommodate. display already shows both halves. Its VirtioGpu backend speaks scanout-protocol over an IPC handle and touches no device — the driver holds the hardware, the service speaks to it, and it works today. Its Gop backend calls device.enumerate, device.claim and device.mmioMap: the service reaching into hardware itself, because the firmware framebuffer has no driver to talk to. That is the only such case in the tree; every other claimant is a device-manager child spawned with its device id. A firmware-framebuffer driver is therefore a prerequisite of this phase rather than a consequence. It is small — hold the display node, map the framebuffer, serve the same scanout-protocol virtio-gpu already serves — and the compositor needs no new path, since not caring which backend is behind it is what it was designed for. Two earlier drafts are recorded as wrong in the document: display asking the device manager for a grant, and before that init minting grants because init starts display. Both accommodated an exception instead of removing it. --- docs/os-development/device-authority.md | 78 ++++++++++++++++--------- 1 file changed, 50 insertions(+), 28 deletions(-) diff --git a/docs/os-development/device-authority.md b/docs/os-development/device-authority.md index f6f7c55..926b03d 100644 --- a/docs/os-development/device-authority.md +++ b/docs/os-development/device-authority.md @@ -64,27 +64,46 @@ The kernel therefore does not need to know about matching, `devices.csv`, driver or why a device was assigned. It needs to know that an authority it can verify granted this device to this task. -## The awkward fact that decides the mechanism +## Only drivers hold devices -Five of the six claimants are device-manager children, spawned with their device id in -`argv[1]`: `pci-bus`, `usb-xhci-bus`, `ps2-bus`, `virtio-gpu`, `acpi`. +The layering the rest of the system already follows: -**`display` is not.** It is spawned by `init` from `init.csv`, and it finds its device by -enumerating the table for a display-class node and claiming whatever it finds -([backend.zig](../../system/services/display/backend.zig)). There is no assignment to -enforce, because nobody assigned it anything. +- A **driver** talks to hardware. It is the only thing that holds a device. +- A **service** talks to no hardware at all. It receives device events and sends data + and commands over a **protocol**. +- A **protocol** is the abstraction between them — the OS layer, in the sense of + [communication.md](communication.md). -That rules out the cheapest design. "The kernel records the device named at spawn, and -`device_claim` checks it" closes the hole for five claimants and breaks the sixth. And -the sixth is not an oddity to special-case — it is the one that shows the model is -wrong: authority should be *delegable*, not welded to the moment of spawn. +So the question "which services need device grants" has the answer **none**. That +collapses the design: every holder of a device is a driver, and every driver is spawned +by the device manager, which is what grants it. -It is also what tempted an earlier draft of this document into giving `init` the root -grants, since `init` is what starts `display`. That was wrong for a plainer reason: -`init` has nothing to do with devices. It is the first process and it starts the rest of -the system; the device manager is what owns hardware. Under delegation `display` simply -asks the manager, like every other driver, and the exception disappears rather than -being accommodated. +`display` already demonstrates both halves, one right and one wrong +([backend.zig](../../system/services/display/backend.zig)): + +| Backend | How it reaches the panel | +|---|---| +| `VirtioGpu` | speaks `scanout-protocol` over an IPC handle; touches no device | +| `Gop` | `device.enumerate`, `device.claim`, `device.mmioMap` | + +The virtio-gpu path is the intended shape and it works today: the driver holds the +hardware, the service speaks a protocol to it. **The GOP path is the anomaly** — the +display service reaching into hardware itself, because the firmware framebuffer has no +driver for it to talk to. + +It is also the only anomaly. Every other claimant — `pci-bus`, `usb-xhci-bus`, +`ps2-bus`, `virtio-gpu`, `acpi` — is a device-manager child spawned with its device id +in `argv[1]`. + +**A firmware-framebuffer driver is therefore a prerequisite of this phase**, not a +consequence of it. It is small: it holds the display device, maps the framebuffer, and +serves `scanout-protocol` — the same interface `virtio-gpu` already serves, so the +compositor needs no new code path and stops caring which is behind it, which it was +designed for. + +An earlier draft of this document had `display` asking the device manager for a grant, +and before that had `init` minting grants because `init` starts `display`. Both were +accommodating an exception instead of removing it. ## The design @@ -115,10 +134,10 @@ device manager the same way it starts everything else, and knows nothing about d `init` attests the services it spawned for `/protocol` (`supervisorSatisfies`); it knows which task is which driver because it spawned them. -3. **`display` asks the manager too.** It is started by `init` as a service, but the - framebuffer is a device, so it receives that grant from the device manager like any - other driver. One authority for hardware, not two — which is what makes `display` - stop being the exception that broke the simpler design. +3. **Services receive nothing.** `display`, `input`, `fat` and `logger` hold no devices + and need no grants; they speak protocols to the drivers that do. The one place this + is not true today is the GOP backend above, which a firmware-framebuffer driver + removes. 4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and `dma_bind` check possession of the grant instead of consulting an ownership table. @@ -137,13 +156,16 @@ own comment says "Claim the bridge, map the ECAM, hello the manager, then scan." delegation the hello must come first, because that is where the grant arrives. Five drivers need that reordering, and it is the bulk of the work. -**`display` gains a dependency on the device manager.** It currently finds its -framebuffer by itself and needs nothing from anyone; afterwards it must ask. That is a -new ordering constraint at boot — `display` cannot bring up a screen until the manager -is up — and it is the honest price of there being one authority for hardware. Mitigating: -`display` already tolerates arriving before its backend (the virtio-gpu driver hands it -a scanout later, over `attach_scanout`), so the machinery for "wait, then attach" is -there. +**A firmware-framebuffer driver has to be written first**, and until it exists the +display service cannot stop claiming a device. It is the smallest new binary in the +tree — hold the display node, map the framebuffer, serve `scanout-protocol` — but it is +new code on the boot path, and the boot path is where a mistake costs a screen. + +It brings an ordering constraint with it: `display` cannot paint until that driver is +up, where today it maps the framebuffer itself and is independent. Mitigating, +`display` already tolerates a backend arriving late — `virtio-gpu` hands it a scanout +after the fact over `attach_scanout` — so "wait, then attach" is a path that already +works rather than one to invent. **The kernel gains one piece of knowledge about a specific binary.** Chain attestation means the kernel recognises `/system/services/device-manager` under PID 1 as the root From 547d0ec46b6f5a825af215d6ff6a2645e1d8c4cb Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 17:01:25 +0100 Subject: [PATCH 13/36] =?UTF-8?q?docs:=20device=20authority=20=E2=80=94=20?= =?UTF-8?q?the=20how,=20and=20the=20run=20that=20deletes=20the=20ceilings?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit device-authority.md is rewritten as an implementation design rather than a rival to device-manager.md. The what was already settled there in 2026-07: structure in the manager, authority in the kernel, and delegation as the step after hello. This is the how, plus the two decisions that paragraph leaves open. Decision 1: the manager claims, it is not granted. device-manager.md says "claims (or is granted)"; claiming wins because the manager runs before any driver exists and takes the seeded devices unopposed, leaving nothing unheld to race for. One new call, device_transfer(device_id, task_id), checks only that the caller holds the device — no names in the kernel, no attestation. The alternative put a binary path inside the kernel, and the kernel should hold only what cannot safely live in user space. The residual is stated rather than hidden: authority rests on the manager claiming first, which init.csv makes an operator-visible ordering rather than an attacker- controlled one, and the enforced version arrives with the spawn capability drivers.md already names as missing. Decision 2: the kernel stops holding inventory. It reads three things out of a descriptor — physical ranges, interrupt numbers, one PCI BDF — and stores the rest only so device_enumerate can hand it back. Devices with no resources leave the kernel entirely: a USB device conveys no mapping authority, so there is nothing to enforce. That is also the case which sidesteps containment, and therefore the reason a shared cap existed. Decision 3: no shared ceiling. The table becomes dynamic — it is built after heap.init, so nothing ever prevented it — and the two invented numbers go. A per-holder quota replaces them, because dynamic storage with no bound moves the ceiling to the kernel heap, which is shared and fatal rather than partial. A bound charged to whoever caused it is isolation. Two earlier drafts of this document are gone: one gave init the root grants, the other proposed extracting a firmware-framebuffer driver. Both were wrong and both are recorded as wrong in the run plan's settled list — the framebuffer is not a device, it is where pixels go until a real display driver announces itself. Run 2 is nine steps, ordered so the suite stays green throughout: build and prove the transfer mechanism, move the five claimants across one at a time, then the flag day, then the inventory, then the ceilings. --- docs/bounds-track-plan.md | 59 ++++- docs/os-development/device-authority.md | 297 ++++++++++-------------- 2 files changed, 181 insertions(+), 175 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 90d34c5..73f1fcf 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -19,11 +19,68 @@ next one starts.* | L5 | USB: interfaces from the descriptor, and the misattributed-endpoint bug | **done** — fix is by construction; no direct test, see open question 5 | | L6 | xHCI: a failed `allocateDevice` stops leaking an enabled slot | **done** — path forced and verified; no regression test, see open question 5 | -**Run complete.** L1 stopped (the step was wrong), L2–L6 landed. Suite 115 → 116. +**Run 1 complete.** L1 stopped (the step was wrong), L2–L6 landed. Suite 115 → 116. Allowlist 278 → 269. Two steps ship without a permanent regression test, both because QEMU's USB devices are too small to reach the paths — see open question 5, which is the audit's own lesson recurring: the test rig is smaller than a real machine. +--- + +## Run 2 — device authority: delete the invented ceilings + +*Design: [device-authority.md](os-development/device-authority.md), which is the **how** +for the delegation step [device-manager.md](device-driver-development/device-manager.md) +already settled. Read both before starting; the second is authoritative where they +differ.* + +The goal, in the project owner's words: **remove the maximum values we set arbitrarily, +move the responsibility to the device manager, and keep in the kernel only the parts +that cannot safely run in user space.** + +| Step | What | State | +|---|---|---| +| D1 | `device_transfer(device_id, task_id)` — the holder gives a device away | not started | +| D2 | Adversarial case: a process handed nothing is refused, on a held device and a free one | not started | +| D3 | The manager claims the seeded devices at boot, before any driver is spawned | not started | +| D4 | `usb-xhci-bus` receives its controller in the `hello` reply instead of claiming argv[1] | not started | +| D5 | The other four claimants converted: `pci-bus`, `ps2-bus`, `virtio-gpu`, `acpi` | not started | +| D6 | `device_claim` refuses a device the caller was not handed; the hole is closed | not started | +| D7 | Zero-resource devices stop being kernel objects — inventory moves to the manager | not started | +| D8 | **`maximum_children_per_parent` deleted** — the authorisation it stood in for exists | not started | +| D9 | The device table becomes dynamic; **`maximum_devices` deleted**; per-holder quota declared | not started | + +Ordering is load-bearing. D1–D2 build and prove the mechanism with nothing depending on +it. D3–D5 move each claimant across one at a time, so the suite stays green throughout +and a regression names the driver that caused it. D6 is the flag day. D7 must precede +D9, because zero-resource children are the case that sidesteps containment and so the +reason a shared cap was needed at all. + +### Settled, so the run does not re-litigate them + +- **The manager claims, it is not granted.** No binary names in the kernel; the rule is + "you may give away what you hold". The residual — it rests on the manager claiming + first — is stated in the design and is closed later by the spawn capability + [drivers.md](device-driver-development/drivers.md) already names as missing. +- **The framebuffer is not a device.** It is where pixels go, handed over by the loader, + and the compositor uses it as the boot floor until a real display driver announces + itself. Nothing in this run touches the display service or its GOP path. +- **`maximum_endpoints_per_interface` and the wire structs stay.** Widening them is a + protocol change, out of scope. +- **A per-holder quota is not a retreat.** Dynamic storage with no bound moves the + ceiling to the kernel heap, which is shared and fatal rather than partial. A bound + charged to the task that caused it is isolation, and it is declared through + [bounds.md](os-development/bounds.md) like anything else. + +### Working rules + +As Run 1, unchanged: work in `/Users/danielsamson/Gitea/daniel/danos` on +`claude/bounds-track`; every step lands with a test that fails before the fix, verified +by restoring the old behaviour; full suite green before each commit; never two suites at +once (`pgrep -f qemu_test.py`); 60 GiB free; `git commit -F` with no `Co-Authored-By`; +update this table before starting the next step. **If a step needs a decision that is +not written down, stop it, add the question below, and move on** — Run 1's first step +was wrong and stopping was the right call. + **Suite:** 115/115 at the start of the run. **Branch:** `claude/bounds-track`. diff --git a/docs/os-development/device-authority.md b/docs/os-development/device-authority.md index 926b03d..9509214 100644 --- a/docs/os-development/device-authority.md +++ b/docs/os-development/device-authority.md @@ -1,12 +1,25 @@ -# Device authority: you hold what you were given +# Device authority: the implementation of delegation -*Design for phase 2 of [the bounds track](../bounds-track-plan.md), 2026-08-08. -Supersedes an earlier draft that argued for capabilities on aesthetic grounds; this one -starts from the hole and from the project's principles.* +*Implementation design, 2026-08-08. The **what** is settled in +[device-manager.md](../device-driver-development/device-manager.md) — "structure in the +manager, authority in the kernel", and delegation as the step after `hello`. This +document is the **how**, and the decisions that paragraph leaves open.* -## The hole +Read first: [drivers.md](../device-driver-development/drivers.md) (the claim is the +capability), [driver-model.md](../device-driver-development/driver-model.md) (the three +invariants), [device-manager.md](../device-driver-development/device-manager.md) (the +tree, the matcher, the supervisor). -`device_claim(id)` checks two things ([devices-broker.zig](../../system/kernel/devices-broker.zig)): +## The one thing not yet true + +`device-manager.md` says assignment "stays argv for now", and names the next step: + +> The step after `hello` exists is delegation: the manager claims (or is granted) the +> devices and passes the claim to the driver over IPC (the M13 capability-transfer +> mechanism), replacing first-come-first-served `device_claim` with policy. + +Until that lands, the manager's matching is advisory. `device_claim` checks only that +the device exists and is unheld ([devices-broker.zig](../../system/kernel/devices-broker.zig)): ```zig pub fn claim(id: u64, owner: u32) ClaimError!void { @@ -16,187 +29,123 @@ pub fn claim(id: u64, owner: u32) ClaimError!void { } ``` -Does it exist, and is it free. **Any process may claim any unclaimed device.** +A driver is spawned with its device id in `argv[1]` and claims it; any process could +pass any integer instead. Since a claim is a licence to map physical memory, that is the +gap this document closes. -The matching is real but it is entirely advisory: `device-manager` reads `devices.csv`, -matches a device to a driver, and spawns that driver with the device id as `argv[1]` -(`spawnDriver`). The driver parses the string and claims it. Nothing anywhere binds the -manager's decision to the kernel's grant — a process can pass any integer and win the -race. +## Decision 1: the manager claims, then transfers -A claim is not a small thing. It is what gates `mmio_map` and `irq_bind`, so it is a -licence to map physical memory and receive interrupts. +`device-manager.md` leaves "claims (or is granted)" open. **Claims.** -## What this costs, beyond the obvious +The manager runs before any driver exists — `init` starts it from `init.csv`, and it is +what spawns drivers — so it takes the seeded devices unopposed and there is nothing +unheld left for anyone to race for. One new call moves ownership on: -`maximum_children_per_parent = 16` exists because a driver that claimed one device could -loop `device_register` under it and exhaust the shared table. That threat only exists -*because* claiming is unauthenticated — and the cap is a poor defence against it, since -an attacker can burn 16 slots, claim another device, and burn 16 more. What it reliably -does instead is refuse a legitimate PCI bus with more than 16 functions, which is how an -AMD Ryzen came to boot with no USB and no storage. +``` +device_transfer(device_id, task_id) -> 0/-errno +``` -So the cap is not merely mis-sized. It is standing in for an authorisation that is not -performed, and it punishes correct behaviour while barely inconveniencing incorrect -behaviour. **Closing the hole is what retires the constant**, not a bigger number. +The kernel checks only that the caller currently holds the device. No names, no +attestation, no notion of "the device manager" — the rule is *you may give away what you +hold*, which is the capability discipline already in force. -## What the principles decide +The alternative was the kernel granting roots to a task it recognises by binary path +plus a PID-1 supervisor. It is more robust — it does not depend on the manager being +first — but it puts a binary name inside the kernel, and the goal is that the kernel +keeps only what cannot safely live in user space. A name is not that. -- *Move as much responsibility as possible to user space* (3), and *what remains in the - kernel is there for security or a hardware limitation* (5). +**The residual, stated plainly:** authority here rests on the manager claiming first. +That holds because `init.csv` decides what starts and in what order, so it is an +operator-visible ordering rather than an attacker-controlled one — but it is an +assumption, not an enforced invariant. The enforced version arrives with the spawn +capability [drivers.md](../device-driver-development/drivers.md) already names as +missing ("`system_spawn` is currently ungated … because there is no spawn capability +yet"). This design is compatible with it and does not block on it. -Deciding **which** driver gets **which** device is policy — matching identity triples -and choosing a binary. That decision belongs to the device manager and stays there. -`devices.csv` is not the policy; it is **configuration**, the declarative data the -policy reads. Three distinct things, and worth keeping apart in this document: +## Decision 2: the kernel stops holding inventory -| | Lives in | Example | +The kernel reads exactly three things out of a descriptor: **physical ranges** (to check +a mapping falls inside one), **interrupt numbers**, and **one PCI BDF** (to key an IOMMU +domain). Vendor, device and subsystem ids, class triples, `_HID` strings, bus addresses +and names are stored only so `device_enumerate` can hand them back — which +`device-manager.md` already resolves: that call "fades to a manager-internal (then +deleted) seam", because the manager owns the tree as data. + +So the kernel's table becomes: **parent, resources, holder, BDF.** That is what cannot +safely run in user space; the rest moves. + +**Devices with no resources leave the kernel entirely.** A USB device is addressed +through its controller and carries `resource_count = 0` +([driver-model.md](../device-driver-development/driver-model.md): "that case is allowed +and is the common one"). It conveys no mapping authority, so there is nothing for the +kernel to enforce and no reason for it to know. It is inventory, and inventory is the +manager's — reported by `child_added`, which already carries everything needed. + +That is also the case that made `maximum_children_per_parent` necessary: a zero-resource +child sidesteps containment, so a driver could loop `device_register` and fill the +shared table. Once such children are not kernel objects, every remaining entry is a real +contained subdivision of something the caller holds. + +## Decision 3: no shared ceiling; a per-holder quota instead + +`maximum_devices = 64` and `maximum_children_per_parent = 16` are numbers we invented, +and both are shared — one driver's enumeration starves every other driver, which is how +an AMD Ryzen booted with no USB and no storage. + +- **The table becomes dynamic.** It is built after `heap.init` (`kernel.zig`: `pmm.init` + at 137, `heap.init` at 179, `devices_broker.init` at 202), so nothing prevents it. No + specification bounds how many devices a machine has, so nothing should bound ours. +- **`maximum_children_per_parent` is deleted**, because the authorisation it stood in + for now exists. +- **A per-holder quota replaces them.** Dynamic storage without a bound moves the + ceiling to the kernel heap, which is shared and fatal rather than partial — strictly + worse. The bound that is *not* worse is one charged to the task that caused it: a + driver that loops `device_register` exhausts its own allowance, is refused with an + attributable errno, and is restarted by its supervisor while every other driver + carries on. That is the microkernel property rather than a workaround for it, and it + is declared through [bounds.md](bounds.md) like any other. + +## The shape of the change + +| | Before | After | |---|---|---| -| **Mechanism** | the kernel | the check that a grant is held before a mapping is made | -| **Policy** | user space | the device manager matching a device to a driver | -| **Configuration** | files | `devices.csv`, `protocol.csv`, `init.csv` | +| Manager gets its devices | claims them, unauthorised | claims them (first, unopposed) | +| Driver gets its device | `argv[1]` + `device_claim` | receives it in the `hello` reply | +| Kernel checks | is it free? | do you hold it? | +| Kernel stores | the full descriptor | parent, resources, holder, BDF | +| Zero-resource devices | kernel table entries | manager records only | +| Table size | `maximum_devices = 64` | dynamic, per-holder quota | +| Children per parent | `maximum_children_per_parent = 16` | deleted | -Enforcing that a driver **holds only what it was given** is security — it is the gate in -front of mapping physical memory. That stays in the kernel, and it is the whole of what -the kernel needs to do. +Bring-up order changes for the five claiming drivers: `hello` must precede the claim, +because the reply is where the device arrives. `pci-bus` today does the reverse — its +own comment reads "Claim the bridge, map the ECAM, hello the manager, then scan." -The kernel therefore does not need to know about matching, `devices.csv`, driver names, -or why a device was assigned. It needs to know that an authority it can verify granted -this device to this task. +## What does not change -## Only drivers hold devices +- The three invariants of [driver-model.md](../device-driver-development/driver-model.md): + a claim is exclusive, a descriptor is a licence to map physical memory, therefore + containment. This design strengthens the first and touches neither of the others. +- The display service's GOP path. The framebuffer is not a device — it is where pixels + go, handed over by the loader, and the compositor uses it as the boot floor until a + real display driver announces itself + ([display-v2.md](../device-driver-development/display-v2.md)). The kernel wraps it in + a display-class descriptor so `mmio_map` can hand it over write-combining; that is + plumbing for a mapping, not a claim about what it is. +- Supervision, restart, pruning and re-report + ([device-manager.md](../device-driver-development/device-manager.md), + [process-lifecycle.md](process-lifecycle.md)). Delegation slots into the existing + `hello` exchange and changes none of it. +- `device_register` idempotency, which is what lets a restarted bus rebuild the same + ids. -The layering the rest of the system already follows: +## How it is verified -- A **driver** talks to hardware. It is the only thing that holds a device. -- A **service** talks to no hardware at all. It receives device events and sends data - and commands over a **protocol**. -- A **protocol** is the abstraction between them — the OS layer, in the sense of - [communication.md](communication.md). +The invariant is: **a process holds what it was handed and cannot name its way into +holding more.** The suite has no adversarial device case today — the audit's lesson was +that "the suite contains no attacker" — so this adds one: a process that was handed +nothing calls `device_claim` and `device_transfer` on a device another driver holds, and +on one nobody holds, and is refused each time with its own errno. -So the question "which services need device grants" has the answer **none**. That -collapses the design: every holder of a device is a driver, and every driver is spawned -by the device manager, which is what grants it. - -`display` already demonstrates both halves, one right and one wrong -([backend.zig](../../system/services/display/backend.zig)): - -| Backend | How it reaches the panel | -|---|---| -| `VirtioGpu` | speaks `scanout-protocol` over an IPC handle; touches no device | -| `Gop` | `device.enumerate`, `device.claim`, `device.mmioMap` | - -The virtio-gpu path is the intended shape and it works today: the driver holds the -hardware, the service speaks a protocol to it. **The GOP path is the anomaly** — the -display service reaching into hardware itself, because the firmware framebuffer has no -driver for it to talk to. - -It is also the only anomaly. Every other claimant — `pci-bus`, `usb-xhci-bus`, -`ps2-bus`, `virtio-gpu`, `acpi` — is a device-manager child spawned with its device id -in `argv[1]`. - -**A firmware-framebuffer driver is therefore a prerequisite of this phase**, not a -consequence of it. It is small: it holds the display device, maps the framebuffer, and -serves `scanout-protocol` — the same interface `virtio-gpu` already serves, so the -compositor needs no new code path and stops caring which is behind it, which it was -designed for. - -An earlier draft of this document had `display` asking the device manager for a grant, -and before that had `init` minting grants because `init` starts `display`. Both were -accommodating an exception instead of removing it. - -## The design - -**A device grant is a capability, delegated from a holder.** The mechanism already -exists: `callCap` passes a handle over an IPC call and the kernel installs it in the -receiver's table ([library/kernel/ipc.zig](../../library/kernel/ipc.zig)), which is how -shared memory and DMA regions already move between processes. - -**The device manager is the root of device authority.** `init` has no part in this: it -is the first process, and its job is to start the rest of the system. It starts the -device manager the same way it starts everything else, and knows nothing about devices. - -1. **Root.** At boot the kernel mints grants for the devices firmware discovery found - and hands them to the device manager. This is the only place device authority enters - the system, and it comes from ACPI rather than from anyone's say-so. - - The kernel recognises the manager by **chain attestation**, the identity the security - track already settled on: the binary path the kernel itself stamped - (`/system/services/device-manager`) *and* a supervisor of PID 1. The binary path - alone would not do — `system_spawn` is deliberately ungated, so any process may spawn - any bundled binary, and a rogue could run a second copy under the same name. It could - not forge the other half: its copy's supervisor is the rogue, and PID 1 is the - kernel's own first process. - -2. **Delegation.** The manager passes a driver its device when it spawns it, over the - channel that already exists — the driver `hello`s the manager and the reply carries - the grant. The manager attests its own children by task id, one hop deep, exactly as - `init` attests the services it spawned for `/protocol` (`supervisorSatisfies`); it - knows which task is which driver because it spawned them. - -3. **Services receive nothing.** `display`, `input`, `fat` and `logger` hold no devices - and need no grants; they speak protocols to the drivers that do. The one place this - is not true today is the GOP backend above, which a firmware-framebuffer driver - removes. - -4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and `dma_bind` - check possession of the grant instead of consulting an ownership table. - -Exclusivity stops being a broker refusing a second claimant and becomes the ordinary -property of a capability: only one process was given it. - -**`maximum_children_per_parent` is deleted here.** After this a bus driver's children are -devices it enumerated on a bus it was actually given, and the threat the cap was written -for no longer exists. - -## What this costs - -**Bring-up order changes.** `pci-bus` today claims first and says hello afterwards — its -own comment says "Claim the bridge, map the ECAM, hello the manager, then scan." Under -delegation the hello must come first, because that is where the grant arrives. Five -drivers need that reordering, and it is the bulk of the work. - -**A firmware-framebuffer driver has to be written first**, and until it exists the -display service cannot stop claiming a device. It is the smallest new binary in the -tree — hold the display node, map the framebuffer, serve `scanout-protocol` — but it is -new code on the boot path, and the boot path is where a mistake costs a screen. - -It brings an ordering constraint with it: `display` cannot paint until that driver is -up, where today it maps the framebuffer itself and is independent. Mitigating, -`display` already tolerates a backend arriving late — `virtio-gpu` hands it a scanout -after the fact over `attach_scanout` — so "wait, then attach" is a path that already -works rather than one to invent. - -**The kernel gains one piece of knowledge about a specific binary.** Chain attestation -means the kernel recognises `/system/services/device-manager` under PID 1 as the root -holder. That is a real concession — the kernel would rather know nothing about who is -who — and it is the minimum: device authority has to enter the system somewhere, and -every alternative is worse. Configuration the kernel reads would put a file parser in -the kernel; first-to-ask would be the hole again, at boot. - -**A configuration question stays open.** Which binary may be given which device is -expressed today as `devices.csv`, read by the manager — configuration, with the manager -as the policy that acts on it. That is already the right shape and needs no new file. -A separate grant manifest earns itself only when something must differ from "the manager -gets the hardware": holding a device back so a test or a bare-metal driver can take it, -for instance. Not now. - -## What it does not solve - -- **Hot-unplug and re-enumeration drift.** Still the inventory problem (phase 3, and - open question 4 in the track plan). A grant dying with its holder is not the same as a - device going away. -- **The device table's size.** `maximum_devices` is untouched by this; it goes when the - inventory moves in phase 3. -- **Two processes racing for the same root grant.** Cannot arise: the roots are minted - to whichever task satisfies the chain — the manager's binary under PID 1 — and a - second copy spawned by anyone else fails the supervisor half. - -## How this is verified - -The invariant is **I3** from the track plan: a process holds what it was handed and -cannot name its way into holding more. The test is adversarial and the suite has never -had one of these for devices: a process that was granted nothing calls `device_claim` -on a device another driver owns, and on one nobody owns, and is refused both times with -its own errno. The audit's lesson was that "the suite contains no attacker"; this is the -attacker for devices. +The Ryzen is the acceptance test for the ceiling half: it is the machine that found the +constants, and the one that proves them gone. From 3111c7c5e6bc6e2029da6bfa26b0abcb497e4105 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 17:11:59 +0100 Subject: [PATCH 14/36] =?UTF-8?q?kernel:=20device=5Ftransfer=20=E2=80=94?= =?UTF-8?q?=20you=20may=20give=20away=20what=20you=20hold?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The mechanism behind delegation, which device-manager.md named as the step after hello: the device manager claims what discovery seeded and hands each device to the driver it matched, so assignment stops being first-come-first-served. It is a MOVE, not a copy. A claim is exclusive (driver-model.md, invariant 1), so the giver stops holding the device the instant the receiver starts. That is why this is a new syscall rather than the M13 capability path, where a passed handle is shared refcounted — exclusivity cannot be expressed that way. The kernel's whole rule is that you may give away what you hold. It has no notion of which task is the device manager and deliberately gains none: a binary name inside the kernel is not something that cannot safely live in user space. A recipient that does not exist is refused, because a device moved to nobody would be unreachable for the rest of the boot — nothing un-holds a device but task death. Three errnos, each naming its own rule: ENODEV no such device, EPERM you do not hold it, ESRCH no such recipient. Nothing uses it yet. The five claimants move across one at a time in D4-D5, so the suite stays green throughout and a regression names the driver that caused it. Ten assertions, verified to discriminate: removing the ownership check flips four of them, including the giveaway that an illegal transfer then blocks the legitimate claim behind it. Suite 116 -> 117. --- docs/bounds-track-plan.md | 2 +- library/device/driver/driver.zig | 19 ++++++++++ system/abi.zig | 1 + system/kernel/devices-broker.zig | 36 +++++++++++++++++++ system/kernel/process.zig | 24 +++++++++++++ system/kernel/tests.zig | 59 ++++++++++++++++++++++++++++++++ test/qemu_test.py | 8 +++++ 7 files changed, 148 insertions(+), 1 deletion(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 73f1fcf..bd4c2f9 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -39,7 +39,7 @@ that cannot safely run in user space.** | Step | What | State | |---|---|---| -| D1 | `device_transfer(device_id, task_id)` — the holder gives a device away | not started | +| D1 | `device_transfer(device_id, task_id)` — the holder gives a device away | **done** — syscall 54; a move, not a copy | | D2 | Adversarial case: a process handed nothing is refused, on a held device and a free one | not started | | D3 | The manager claims the seeded devices at boot, before any driver is spawned | not started | | D4 | `usb-xhci-bus` receives its controller in the `hello` reply instead of claiming argv[1] | not started | diff --git a/library/device/driver/driver.zig b/library/device/driver/driver.zig index b34af41..3a9633c 100644 --- a/library/device/driver/driver.zig +++ b/library/device/driver/driver.zig @@ -46,6 +46,25 @@ pub fn enumerate(buffer: []DeviceDescriptor) usize { return sc.systemCall2(.device_enumerate, @intFromPtr(buffer.ptr), buffer.len); } +/// Why a `transfer` failed. `NotHeld` is the interesting one — it means the caller tried +/// to give away a device it does not have, which is the whole rule. +pub const TransferError = error{ NoSuchDevice, NotHeld, NoSuchTask, Refused }; + +/// Give device `id` to task `to`. **A move, not a copy** — a claim is exclusive, so the +/// caller stops holding it. This is how the device manager hands a driver the device it +/// matched, replacing first-come-first-served claiming with policy +/// (docs/os-development/device-authority.md). +pub fn transfer(id: u64, to: u32) TransferError!void { + const r = sc.systemCall2(.device_transfer, id, to); + if (!failed(r)) return; + return switch (errnoOf(r)) { + abi.ENODEV => error.NoSuchDevice, + abi.EPERM => error.NotHeld, + abi.ESRCH => error.NoSuchTask, + else => error.Refused, + }; +} + /// Why a `claim` failed. Worth distinguishing: `AlreadyClaimed` means back off and /// let the owner have it, `NoSuchDevice` means this id is stale and the caller should /// re-enumerate, and `NotConfined` means the machine could not place the device under diff --git a/system/abi.zig b/system/abi.zig index d8f2157..138fe2d 100644 --- a/system/abi.zig +++ b/system/abi.zig @@ -85,6 +85,7 @@ pub const SystemCall = enum(u64) { dma_bind = 51, // dma_bind(device_id, region_handle) -> 0/-errno: map a DMA-region capability into the claimed device's IOMMU domain (idempotent). The caller must own the device and hold the handle dma_unbind = 52, // dma_unbind(device_id, region_handle) -> 0/-errno: unmap a previously bound region from the device's domain and invalidate handle_close = 53, // handle_close(handle) -> 0/-errno: drop one capability handle and free its table slot (endpoints, shared-memory, DMA regions) + device_transfer = 54, // device_transfer(device_id, task_id) -> 0/-errno: give a device you hold to another task. A MOVE, not a copy — a claim is exclusive (-ENODEV no such device, -EPERM you do not hold it, -ESRCH no such task) _, }; diff --git a/system/kernel/devices-broker.zig b/system/kernel/devices-broker.zig index cc67ddd..6dc7bdc 100644 --- a/system/kernel/devices-broker.zig +++ b/system/kernel/devices-broker.zig @@ -323,6 +323,42 @@ pub const ClaimError = error{ AlreadyClaimed, // a live task already owns it }; +/// Why a `transfer` was refused. +pub const TransferError = error{ + NoSuchDevice, // no device with that id + NotHeld, // the caller does not hold it — you may only give away what you have +}; + +/// The errno a refused `transfer` returns to ring 3. (`ESRCH` — no such recipient — is +/// raised by the caller in system/kernel/process.zig, which is what can see the task +/// table.) +pub fn transferErrnoOf(e: TransferError) i64 { + return switch (e) { + error.NoSuchDevice => abi.ENODEV, + error.NotHeld => abi.EPERM, + }; +} + +/// Move device `id` from `from` to `to`. **A move, not a copy**: a claim is exclusive +/// (driver-model.md, invariant 1), so the giver stops holding it the moment the +/// receiver starts. +/// +/// This is the mechanism behind delegation — the device manager claims what firmware +/// discovery seeded and passes each device to the driver it matched, which replaces +/// first-come-first-served `device_claim` with policy +/// (docs/device-driver-development/device-manager.md). The kernel checks only that the +/// caller holds the device: *you may give away what you have*. It knows nothing about +/// which task is the manager, and needs to know nothing. +/// +/// Note this is deliberately NOT the M13 capability-passing path, which shares a handle +/// refcounted — a copy. Exclusivity cannot be expressed that way. +pub fn transfer(id: u64, from: u32, to: u32) TransferError!void { + if (id >= count) return error.NoSuchDevice; + const holder = claimed[@intCast(id)] orelse return error.NotHeld; + if (holder != from) return error.NotHeld; + claimed[@intCast(id)] = to; +} + /// The errno a refused `claim` returns to ring 3. (`ECONFINE` — the claim stood but /// the IOMMU would not confine the device — is raised by the caller in /// system/kernel/process.zig, which is what rolls the claim back.) diff --git a/system/kernel/process.zig b/system/kernel/process.zig index 254abf5..c2d8923 100644 --- a/system/kernel/process.zig +++ b/system/kernel/process.zig @@ -249,6 +249,7 @@ fn system_call(state: *architecture.CpuState) void { .irq_bind => systemIrqBind(state), .irq_ack => systemIrqAck(state), .device_register => systemDeviceRegister(state), + .device_transfer => systemDeviceTransfer(state), .system_spawn => systemSpawn(state), .dma_alloc => systemDmaAlloc(state), .dma_free => systemDmaFree(state), @@ -440,6 +441,29 @@ fn systemDeviceClaim(state: *architecture.CpuState) void { architecture.setSystemCallResult(state, 0); } +/// device_transfer(device_id, task_id) -> 0/-errno: give a device you hold to another +/// task. The mechanism behind delegation — the device manager claims what discovery +/// seeded and hands each device to the driver it matched, so assignment stops being +/// first-come-first-served (docs/os-development/device-authority.md). +/// +/// The kernel's whole rule is *you may give away what you hold*. It has no notion of +/// which task is the device manager, and deliberately gains none: a binary name in the +/// kernel is not something that cannot safely live in user space. +fn systemDeviceTransfer(state: *architecture.CpuState) void { + const device_id = architecture.systemCallArg(state, 0); + const task_id: u32 = @truncate(architecture.systemCallArg(state, 1)); + const flags = sync.enter(); + defer sync.leave(flags); + + // The recipient must exist, or the device would be moved to nobody and become + // unreachable for the rest of the boot — no path un-holds a device but task death. + if (scheduler.taskByIdLocked(task_id) == null) return failErr(state, ipc.ESRCH); + + devices_broker.transfer(device_id, scheduler.current().id, task_id) catch |e| + return failErr(state, devices_broker.transferErrnoOf(e)); + architecture.setSystemCallResult(state, 0); +} + /// mmio_map(device_id, resource_index) -> virtual_address: map a claimed device's MMIO window into /// this address space (strong-uncacheable) and return the register base address. /// The claim is the capability — a process can only map hardware it owns. diff --git a/system/kernel/tests.zig b/system/kernel/tests.zig index 328390d..f4ea81e 100644 --- a/system/kernel/tests.zig +++ b/system/kernel/tests.zig @@ -261,6 +261,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void { containmentTest(); } else if (eql(case, "apertures")) { apertureTest(); + } else if (eql(case, "device-transfer")) { + deviceTransferTest(boot_information); } else if (eql(case, "device-manager")) { deviceManagerTest(boot_information); } else if (eql(case, "protocol-registry")) { @@ -4020,6 +4022,63 @@ fn containmentTest() void { } /// A minimal child descriptor with one memory resource, for the containment test. +/// Delegation's mechanism. A claim is exclusive (driver-model.md, invariant 1), so +/// handing a device on is a **move**: the giver stops holding it the instant the +/// receiver starts. That is why this is not the M13 capability path, which shares a +/// handle refcounted. +/// +/// The rule the kernel enforces is the whole of it: *you may give away what you hold*. +/// It has no idea which task is the device manager and needs none +/// (docs/os-development/device-authority.md). +fn deviceTransferTest(boot_information: *const boot_handoff.BootInformation) void { + var buffer: [8]device_abi.DeviceDescriptor = undefined; + check("the device tree is seeded", devices_broker.enumerate(&buffer) >= 2); + + const image = bundledInit(boot_information) orelse { + check("initial_ramdisk carries /system/services/init", false); + result(); + return; + }; + const me = scheduler.currentId(); + const endpoint = ipcsync.createIpcEndpoint() orelse { + check("exit endpoint allocated", false); + result(); + return; + }; + const child = process.spawnProcessSupervised(image, 4, &.{"/system/services/init"}, me, endpoint) catch 0; + check("supervised child spawned", child != 0); + + // You may only give away what you hold — so an unheld device cannot be moved at all, + // which is what stops a transfer being a back door around claiming. + const unheld = if (devices_broker.transfer(0, me, child)) |_| false else |e| e == error.NotHeld; + check("an unheld device cannot be transferred", unheld); + + check("claimed device 0", claimOk(0, me)); + + // The move itself. + const moved = if (devices_broker.transfer(0, me, child)) |_| true else |_| false; + check("the holder may transfer", moved); + const holder = devices_broker.ownerOf(0) orelse 0; + check("the receiver holds it", holder == child); + check("the giver does not", holder != me); + + // Having given it away, the giver cannot give it again. This is the assertion that + // makes it a move rather than a copy. + const again = if (devices_broker.transfer(0, me, child)) |_| false else |e| e == error.NotHeld; + check("a former holder cannot transfer again", again); + + // Nor may a stranger move a device it never held. + const stranger = if (devices_broker.transfer(0, 9999, me)) |_| false else |e| e == error.NotHeld; + check("a stranger cannot transfer another task's device", stranger); + + const absent = if (devices_broker.transfer(9999, me, child)) |_| false else |e| e == error.NoSuchDevice; + check("a device that does not exist is refused", absent); + + devices_broker.releaseAllOwnedBy(child); + devices_broker.releaseAllOwnedBy(me); + result(); +} + /// PCI host-bridge apertures are derived from the *holes* in the firmware memory map, /// and a registered BAR must fall inside one. So the invariant is not "we find the /// holes" but "an aperture never covers memory the firmware described" — an aperture diff --git a/test/qemu_test.py b/test/qemu_test.py index 3f52528..c9516ce 100644 --- a/test/qemu_test.py +++ b/test/qemu_test.py @@ -1021,6 +1021,14 @@ CASES = [ {"name": "apertures", "expect": r"DANOS-TEST-RESULT: PASS", "fail": r"DANOS-TEST-RESULT: FAIL"}, + # Delegation's mechanism (docs/os-development/device-authority.md). A claim is + # exclusive, so handing a device on is a MOVE: the giver stops holding it the + # instant the receiver starts - which is why this is not the M13 capability path, + # where a handle is shared refcounted. The kernel's whole rule is "you may give + # away what you hold"; it has no notion of which task is the device manager. + {"name": "device-transfer", + "expect": r"DANOS-TEST-RESULT: PASS", + "fail": r"DANOS-TEST-RESULT: FAIL"}, # IRQ teardown: an exiting driver's line is masked and its slot cleared (so no # ISR notifies a freed endpoint), and a sibling owner sharing that endpoint # keeps its own binding. A long-running driver never reaches this teardown path. From 33376f24abbf6a23d5e3ab9ebdf0df664e893bc6 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 17:23:17 +0100 Subject: [PATCH 15/36] test: the attacker the device suite never had MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The audit's sharpest finding was structural, not a bug: a fully green suite had hidden six real defects because it contains no attacker. Every device case asserts that a driver handed its own hardware can drive it. None asked what a process handed NOTHING can do. device-authority-test is that process. It is spawned with no device and asserts what it therefore cannot do: it cannot give away a device another task holds, nor a free one, because the kernel's rule is that you may give away what you hold and the device's state is irrelevant to a process holding nothing. Asserted across every device the machine actually has, so it cannot pass by accident of which one happened to be free at boot — six on QEMU, none of them its. A positive control runs first. device_enumerate works from this process, so the refusals below it are decisions rather than a syscall path that is simply broken here; without it, "everything failed" would read identically to "the assertions are meaningless". A nonexistent device is refused as NoSuchDevice rather than NotHeld, because a refusal that cannot name its own rule is what cost a debugging session on the Ryzen. What it deliberately does not assert, and says so in its header: device_claim is still first-come-first-served at this point in the run. That is the hole D6 closes, and the claim half of the invariant joins this fixture then. Asserting it now would be writing a test that documents the bug. Verified to discriminate: removing the holder check flips "every transfer by a non-holder is refused" while the positive control keeps passing. Suite 117 -> 118. --- build.zig | 1 + build.zig.zon | 1 + docs/bounds-track-plan.md | 2 +- system/kernel/tests.zig | 45 ++++++++++ test/qemu_test.py | 9 ++ .../services/device-authority-test/build.zig | 15 ++++ .../device-authority-test/build.zig.zon | 15 ++++ .../device-authority-test.zig | 87 +++++++++++++++++++ 8 files changed, 174 insertions(+), 1 deletion(-) create mode 100644 test/system/services/device-authority-test/build.zig create mode 100644 test/system/services/device-authority-test/build.zig.zon create mode 100644 test/system/services/device-authority-test/device-authority-test.zig diff --git a/build.zig b/build.zig index 52c34ec..a3dc24d 100644 --- a/build.zig +++ b/build.zig @@ -342,6 +342,7 @@ pub fn build(b: *std.Build) void { "protocol-registry-test", // drives the registrar: ungranted bind, collision, restart "protocol-denied-test", // restriction stage one: an ungranted open answers as absence "protocol-conformance-test", // the reserved verbs, asked of every provider the boot bound + "device-authority-test", // the attacker: a process handed no device, asserting what it cannot do }) |fixture| { const package = b.lazyDependency(fixture, .{}) orelse @panic("a test fixture package is missing under test/system/services"); diff --git a/build.zig.zon b/build.zig.zon index bc568ab..6f7adc4 100644 --- a/build.zig.zon +++ b/build.zig.zon @@ -79,6 +79,7 @@ .@"protocol-registry-test" = .{ .path = "test/system/services/protocol-registry-test", .lazy = true }, .@"protocol-denied-test" = .{ .path = "test/system/services/protocol-denied-test", .lazy = true }, .@"protocol-conformance-test" = .{ .path = "test/system/services/protocol-conformance-test", .lazy = true }, + .@"device-authority-test" = .{ .path = "test/system/services/device-authority-test", .lazy = true }, // See `zig fetch --save ` for a command-line interface for adding dependencies. //.example = .{ // // When updating this field to a new URL, be sure to delete the corresponding diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index bd4c2f9..83da3ec 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -40,7 +40,7 @@ that cannot safely run in user space.** | Step | What | State | |---|---|---| | D1 | `device_transfer(device_id, task_id)` — the holder gives a device away | **done** — syscall 54; a move, not a copy | -| D2 | Adversarial case: a process handed nothing is refused, on a held device and a free one | not started | +| D2 | Adversarial case: a process handed nothing is refused, on a held device and a free one | **done** — `device-authority-test`; the claim half joins it at D6 | | D3 | The manager claims the seeded devices at boot, before any driver is spawned | not started | | D4 | `usb-xhci-bus` receives its controller in the `hello` reply instead of claiming argv[1] | not started | | D5 | The other four claimants converted: `pci-bus`, `ps2-bus`, `virtio-gpu`, `acpi` | not started | diff --git a/system/kernel/tests.zig b/system/kernel/tests.zig index f4ea81e..d1849f0 100644 --- a/system/kernel/tests.zig +++ b/system/kernel/tests.zig @@ -263,6 +263,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void { apertureTest(); } else if (eql(case, "device-transfer")) { deviceTransferTest(boot_information); + } else if (eql(case, "device-authority")) { + deviceAuthorityTest(boot_information); } else if (eql(case, "device-manager")) { deviceManagerTest(boot_information); } else if (eql(case, "protocol-registry")) { @@ -4219,6 +4221,49 @@ fn protocolRegistryTest(boot_information: *const BootInformation) void { /// /// The fixture's `protocol-denied: ok` is the marker; each step prints its own /// line, which the harness's ordered regex reads. +/// The attacker the device suite never had. The audit's finding was that a fully +/// green suite had missed six real defects because it *contains no attacker* — every +/// device case asserts a driver handed its hardware can drive it, and none asks what a +/// process handed **nothing** can do. +/// +/// The fixture is spawned with no device and asserts what it therefore cannot do. It +/// runs without the device manager on purpose: nothing here needs a driver, and a boot +/// with fewer moving parts makes the refusals unambiguous. +fn deviceAuthorityTest(boot_information: *const BootInformation) void { + log("DANOS-TEST-BEGIN: device-authority\n", .{}); + if (boot_information.initial_ramdisk_len == 0) { + check("bootloader handed over an initial_ramdisk", false); + result(); + return; + } + const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len]; + const rd = initial_ramdisk.Reader.init(image) orelse { + check("initial_ramdisk image is valid", false); + result(); + return; + }; + + process.setInitialRamdisk(image); + check("device-authority-test spawned", spawnNamedWithArg(rd, "device-authority-test", "run")); + + const pass_marker = "device-authority: ok"; + const fail_marker = "device-authority: FAIL"; + scheduler.setPriority(1); + const deadline = architecture.millis() + 20000; + var saw_pass = false; + var saw_fail = false; + while (architecture.millis() < deadline and !saw_pass and !saw_fail) { + if (bufferHas(pass_marker)) saw_pass = true; + if (bufferHas(fail_marker)) saw_fail = true; + scheduler.yield(); + } + scheduler.setPriority(4); + + check("no authority assertion failed", !saw_fail); + check("the attacker completed every assertion", saw_pass); + result(); +} + fn protocolDeniedTest(boot_information: *const BootInformation) void { log("DANOS-TEST-BEGIN: protocol-denied\n", .{}); if (boot_information.initial_ramdisk_len == 0) { diff --git a/test/qemu_test.py b/test/qemu_test.py index c9516ce..b8b14af 100644 --- a/test/qemu_test.py +++ b/test/qemu_test.py @@ -1029,6 +1029,15 @@ CASES = [ {"name": "device-transfer", "expect": r"DANOS-TEST-RESULT: PASS", "fail": r"DANOS-TEST-RESULT: FAIL"}, + # The attacker the device suite never had. The audit's finding was that a fully + # green suite missed six real defects because it contains no attacker: every device + # case asserts a driver handed its hardware can drive it, and none asks what a + # process handed NOTHING can do. This fixture is that process - it holds no device + # and asserts it can give none away, with a positive control first so the refusals + # are decisions rather than a broken syscall path. + {"name": "device-authority", + "expect": r"DANOS-TEST-RESULT: PASS", + "fail": r"DANOS-TEST-RESULT: FAIL"}, # IRQ teardown: an exiting driver's line is masked and its slot cleared (so no # ISR notifies a freed endpoint), and a sibling owner sharing that endpoint # keeps its own binding. A long-running driver never reaches this teardown path. diff --git a/test/system/services/device-authority-test/build.zig b/test/system/services/device-authority-test/build.zig new file mode 100644 index 0000000..1b19cfd --- /dev/null +++ b/test/system/services/device-authority-test/build.zig @@ -0,0 +1,15 @@ +//! The device-authority-test fixture as a binary package (docs/build-packages-plan.md): +//! this file names the binary and EXACTLY the modules its source imports — +//! build-support resolves each name from the domains this zon declares. + +const std = @import("std"); +const build_support = @import("build-support"); + +pub fn build(b: *std.Build) void { + const exe = build_support.userBinary(b, .{ + .name = "device-authority-test", + .root_source_file = b.path("device-authority-test.zig"), + .imports = &.{ "driver", "logging", "process" }, + }); + b.installArtifact(exe); +} diff --git a/test/system/services/device-authority-test/build.zig.zon b/test/system/services/device-authority-test/build.zig.zon new file mode 100644 index 0000000..08eed1e --- /dev/null +++ b/test/system/services/device-authority-test/build.zig.zon @@ -0,0 +1,15 @@ +.{ + .name = .device_authority_test, + .version = "0.0.0", + .fingerprint = 0x4acbba0c105a1462, // Changing this has security and trust implications. + .minimum_zig_version = "0.16.0", + .dependencies = .{ + // build-support supplies the shared recipe; kernel is implicit in + // every binary (the root shim + link script live there). The rest + // are exactly the homes of this binary's declared imports. + .@"build-support" = .{ .path = "../../../../build-support" }, + .kernel = .{ .path = "../../../../library/kernel" }, + .device = .{ .path = "../../../../library/device" }, + }, + .paths = .{""}, +} diff --git a/test/system/services/device-authority-test/device-authority-test.zig b/test/system/services/device-authority-test/device-authority-test.zig new file mode 100644 index 0000000..ee965b3 --- /dev/null +++ b/test/system/services/device-authority-test/device-authority-test.zig @@ -0,0 +1,87 @@ +//! device-authority-test — the attacker the device suite never had. +//! +//! The audit behind [docs/fixed-bounds-audit.md] found six real defects that a +//! fully green suite had missed, and the reason was structural: *the suite +//! contains no attacker*. Every device case asserts that a driver handed its +//! own hardware can drive it. None asks what a process that was handed +//! **nothing** can do. +//! +//! This binary is that process. It is spawned with no device, holds no device, +//! and asserts what it therefore cannot do +//! ([docs/os-development/device-authority.md]): +//! +//! 1. **A positive control first.** `device_enumerate` works from here, so +//! the refusals below are decisions rather than a syscall path that is +//! simply broken for this process. Without this, "everything failed" would +//! read identically to "the assertions are meaningless". +//! 2. **It cannot give away a device it does not hold** — not one another +//! task holds, and not a free one either. The kernel's whole rule is *you +//! may give away what you hold*, so the state of the device is irrelevant: +//! a process holding nothing can transfer nothing. That is asserted across +//! several ids precisely so it cannot pass by accident of which device +//! happened to be free at boot. +//! 3. **A device that does not exist is refused differently** — `NoSuchDevice` +//! rather than `NotHeld`. A refusal that cannot say which rule refused it +//! is what cost a debugging session on the Ryzen, so the distinction is +//! part of the contract and is tested as such. +//! +//! **What this fixture cannot yet claim.** `device_claim` is still +//! first-come-first-served at this point in the run — that is the hole D6 +//! closes. So the claim half of the invariant ("a process holds what it was +//! handed and cannot name its way into holding more") is deliberately NOT +//! asserted here; it is added to this fixture at D6, when it becomes true. +//! Asserting it now would mean writing a test that documents the bug. + +const std = @import("std"); +const device = @import("driver"); +const logging = @import("logging"); +const process = @import("process"); + +fn line(comptime format: []const u8, arguments: anytype) void { + var buffer: [160]u8 = undefined; + _ = logging.write(std.fmt.bufPrint(&buffer, format, arguments) catch return); +} + +var failures: usize = 0; + +fn check(name: []const u8, ok: bool) void { + if (!ok) failures += 1; + line("device-authority: {s} {s}\n", .{ if (ok) "ok" else "FAIL", name }); +} + +fn run() void { + // 1. The positive control: this process can reach the device syscalls at all. + var table: [64]device.DeviceDescriptor = undefined; + const total = device.enumerate(&table); + check("enumerate works from an unprivileged process", total > 0); + const seen = @min(total, table.len); + + // 2. Holding nothing, it can give nothing away — whatever the device's state. + // Every id the machine actually has, so this cannot pass by luck. + var refused: usize = 0; + var wrong_reason: usize = 0; + for (table[0..seen]) |descriptor| { + device.transfer(descriptor.id, process.taskId()) catch |e| { + refused += 1; + if (e != error.NotHeld) wrong_reason += 1; + continue; + }; + } + check("every transfer by a non-holder is refused", refused == seen); + check("each refusal says NotHeld, not something vaguer", wrong_reason == 0); + + // 3. A device that does not exist is a different refusal, and says so. + const absent = if (device.transfer(0xFFFF_FFFF, process.taskId())) |_| false else |e| e == error.NoSuchDevice; + check("a device that does not exist is refused as absent", absent); + + if (failures == 0) { + line("device-authority: ok ({d} devices, none of them mine)\n", .{seen}); + } else { + line("device-authority: FAILED {d} assertion(s)\n", .{failures}); + } +} + +pub fn main(startup: process.Init) void { + const role = startup.arguments.get(1) orelse return; // bare (ramdisk sweep): stay silent + if (std.mem.eql(u8, role, "run")) run(); +} From a2d605677263805290dc21a5099742d9ce3f505b Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 17:55:59 +0100 Subject: [PATCH 16/36] usb: the xHCI controller arrives by delegation, not by claiming MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The first driver to stop claiming its own hardware. The device manager holds the controller and transfers it in the hello reply, so its matching becomes authoritative instead of advisory — until now the driver claimed the id it found in argv[1], and any process could have claimed the same integer first. The manager claims before it spawns, so there is no window in which anything else could take the device, and transfers in onHello using invocation.sender — the kernel-stamped task id, which cannot be forged by the caller. hello is synchronous, so the transfer has completed before the reply lands: no gap between being told yes and holding the thing. usb-xhci-bus's hello moves from after controller bring-up to before anything that needs the device, which is the bring-up reorder the design predicted. It is the first member of an explicit delegated set, so every unconverted driver keeps claiming exactly as before and the suite stays green; the set and device_claim both go at D6. D3 and D4 could not be separated and the plan records why: the moment the manager claims, any driver still calling device_claim is refused, and D3 applied to nothing changes no behaviour and cannot be tested. This step introduced a regression and the incremental conversion is what caught it. confineDevice runs inside systemDeviceClaim, so a device arriving by transfer was never confined for its new owner. Three IOMMU+USB cases failed on the driver's DMA rings going unbound, and two worse consequences were latent: a manager death would have torn down a domain a live driver was using, and a driver death would have leaked one. iommu.reassign now moves the confinement with the device, keeping the domain and its attachment intact so it never translates through nothing. Converting all five drivers at once would have produced the same three failures with five suspects. A log line of mine claimed "holding controller device N" before anything verified it — it printed even in the failure case, where the driver held nothing. Reworded to state only what is known there: where the registers are. usb-hid asserts the delegation with the device id backreferenced, so the id delegated and the id the driver ends up with must match. Emptying the delegated set fails it with "hello acknowledged" then "mmio_map failed". usb-hub failed once in a full run and has passed six times since (four isolated, two full) — recorded in the plan as a suspected instance of the known intermittent AP fault, not dismissed, since this step did shift boot timing. Suite 118/118. --- docs/bounds-track-plan.md | 31 +++++++++-- system/drivers/usb-xhci-bus/usb-xhci-bus.zig | 24 +++++---- system/kernel/iommu.zig | 20 +++++++ system/kernel/process.zig | 12 +++++ .../device-manager/device-manager.zig | 52 +++++++++++++++++++ test/qemu_test.py | 8 ++- 6 files changed, 134 insertions(+), 13 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 83da3ec..8ab7143 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -41,8 +41,8 @@ that cannot safely run in user space.** |---|---|---| | D1 | `device_transfer(device_id, task_id)` — the holder gives a device away | **done** — syscall 54; a move, not a copy | | D2 | Adversarial case: a process handed nothing is refused, on a held device and a free one | **done** — `device-authority-test`; the claim half joins it at D6 | -| D3 | The manager claims the seeded devices at boot, before any driver is spawned | not started | -| D4 | `usb-xhci-bus` receives its controller in the `hello` reply instead of claiming argv[1] | not started | +| D3 | The manager claims the seeded devices at boot, before any driver is spawned | **merged into D4** — see below | +| D4 | The manager claims + delegates on `hello`; `usb-xhci-bus` is the first driver converted | **done** — caught an IOMMU regression I introduced; see below | | D5 | The other four claimants converted: `pci-bus`, `ps2-bus`, `virtio-gpu`, `acpi` | not started | | D6 | `device_claim` refuses a device the caller was not handed; the hole is closed | not started | | D7 | Zero-resource devices stop being kernel objects — inventory moves to the manager | not started | @@ -50,11 +50,36 @@ that cannot safely run in user space.** | D9 | The device table becomes dynamic; **`maximum_devices` deleted**; per-holder quota declared | not started | Ordering is load-bearing. D1–D2 build and prove the mechanism with nothing depending on -it. D3–D5 move each claimant across one at a time, so the suite stays green throughout +it. D4–D5 move each claimant across one at a time, so the suite stays green throughout and a regression names the driver that caused it. D6 is the flag day. D7 must precede D9, because zero-resource children are the case that sidesteps containment and so the reason a shared cap was needed at all. +**D3 merged into D4** (found while implementing, recorded rather than worked around). +The two cannot be separated: the moment the manager claims a device, any driver still +calling `device_claim` on it is refused `AlreadyClaimed`, so D3 on its own turns the +suite red — and D3 applied to *nothing* changes no behaviour and cannot be tested. +They land together, with the manager claiming only for drivers in an explicit +**delegated set** so every unconverted driver keeps claiming exactly as before. +`usb-xhci-bus` is the first member, as it was the first driver to conform to `hello` +(device-manager.md, M18.1). D5 moves the rest in one at a time; the set and the +`device_claim` path both disappear at D6. + +### Observations from the run + +- **D4 introduced an IOMMU regression, caught by converting one driver at a time.** + `confineDevice` runs inside `systemDeviceClaim`, so a device arriving by *transfer* + was never confined for its new owner: the driver's DMA rings went unbound (three + IOMMU+USB cases failed), and two worse consequences were latent — a manager death + would have torn down a domain a live driver was using, and a driver death would have + leaked one. `iommu.reassign` moves the confinement with the device, keeping the + domain and attachment intact so it never translates through nothing. Converting all + five drivers at once would have produced the same three failures with five suspects. +- **`usb-hub` failed once in a full run, then passed six times** (four isolated, two + full). Suspected instance of the known intermittent AP ring-3 fault rather than + anything in D4 — recorded rather than dismissed, because D4 moved the `hello` earlier + and so did shift boot timing. Watch it across the remaining steps. + ### Settled, so the run does not re-litigate them - **The manager claims, it is not granted.** No binary names in the kernel; the rule is diff --git a/system/drivers/usb-xhci-bus/usb-xhci-bus.zig b/system/drivers/usb-xhci-bus/usb-xhci-bus.zig index be7a9b3..8fcb5fb 100644 --- a/system/drivers/usb-xhci-bus/usb-xhci-bus.zig +++ b/system/drivers/usb-xhci-bus/usb-xhci-bus.zig @@ -173,10 +173,22 @@ fn initialise(endpoint: ipc.Handle) bool { if (!channel.bindPatiently("usb-transfer", endpoint)) _ = logging.write("/system/drivers/usb-xhci-bus: /protocol/usb-transfer is another controller's; serving mine unnamed\n"); - device.claim(controller_id) catch |e| { - std.log.warn("unable to claim controller device {d}: {s}", .{ controller_id, @errorName(e) }); + // **The handshake comes first, because it is where the device arrives.** This + // driver used to claim `controller_id` here — first-come-first-served, so the + // manager's matching was advisory and any process could have claimed it by + // passing the same integer. Now the manager holds the controller and transfers + // it in `onHello`, so by the time this call returns the device is ours and + // nothing else could have taken it (docs/os-development/device-authority.md). + // + // `hello` is synchronous, so the transfer has completed before the reply lands — + // there is no window between being told yes and holding the thing. + // + // Keep the handle: the tick's hot-plug dispatch reports through it. + const handle = device_manager.hello(.bus, controller_id) orelse { + std.log.warn("no hello with the device manager; controller {d} not delegated", .{controller_id}); return false; }; + manager_handle = handle; // Fetch our own descriptor back for the controller's resources. const buffer = memory.allocator().alloc(device.DeviceDescriptor, 64) catch { @@ -204,7 +216,7 @@ fn initialise(endpoint: ipc.Handle) bool { std.log.info("controller device {d} has no register BAR", .{controller_id}); return false; }; - std.log.info("claimed controller device {d} (registers at 0x{x}, {d} bytes)", .{ + std.log.info("controller device {d} registers at 0x{x}, {d} bytes", .{ controller_id, register_window.start, register_window.len, @@ -247,12 +259,6 @@ fn initialise(endpoint: ipc.Handle) bool { return false; } - // The handshake (role: bus — we enumerate USB ports and report the devices - // behind them), inside the manager's hello deadline. Keep the handle: the - // tick's hot-plug dispatch reports through it. - const handle = device_manager.hello(.bus, controller_id) orelse return false; - manager_handle = handle; - scanPorts(handle); // Arm the timer: in polling mode it drains the event ring; in MSI mode it is the diff --git a/system/kernel/iommu.zig b/system/kernel/iommu.zig index 8c2c2ca..7166fc4 100644 --- a/system/kernel/iommu.zig +++ b/system/kernel/iommu.zig @@ -200,6 +200,26 @@ pub fn unmapRegionEverywhere(physical: u64, len: u64) void { /// A driver died or released its devices: tear down every domain it held (detach the /// device, free the tables) so their DMA is blocked again and a restarted driver /// re-claims cleanly. Runs BEFORE the broker claims and the DMA frames are released. +/// Re-point a device's existing confinement at a new owner, keeping its domain and +/// its attachment intact. +/// +/// Delegation needs this: the device manager claims a device (which confines it, with +/// the manager as owner) and then transfers it to the driver. Without moving the +/// confinement record too, the domain stays the manager's — so the driver's DMA +/// buffers are never bound into it, its rings are invisible to the device, and every +/// transfer faults. Worse, a manager death would then tear down a domain a live driver +/// is using, and a driver death would leave one behind. +/// +/// The domain is *not* rebuilt: the device stays attached throughout, so there is no +/// window in which it is translating through nothing. +pub fn reassign(device_id: u64, owner: u32) void { + if (!active) return; + if (device_id >= confined.len) return; + const record = &confined[@intCast(device_id)]; + if (!record.active) return; + record.owner = owner; +} + pub fn releaseAllOwnedBy(owner: u32) void { if (!active) return; for (&confined) |*c| { diff --git a/system/kernel/process.zig b/system/kernel/process.zig index c2d8923..5efecc8 100644 --- a/system/kernel/process.zig +++ b/system/kernel/process.zig @@ -461,6 +461,18 @@ fn systemDeviceTransfer(state: *architecture.CpuState) void { devices_broker.transfer(device_id, scheduler.current().id, task_id) catch |e| return failErr(state, devices_broker.transferErrnoOf(e)); + + // The device's IOMMU confinement moves with it. The giver confined it when it + // claimed, so the domain exists and the device stays attached — but the record + // still names the giver as owner, which would leave the receiver's DMA buffers + // unbound (every transfer faulting), a giver's death tearing down a domain the + // receiver is using, and the receiver's death leaving one behind. + if (devices_broker.pciAddressOf(device_id)) |_| { + iommu.reassign(device_id, task_id); + // Bind whatever the receiver has already allocated — the same courtesy the + // claim path does for a driver that dma_alloc'd its rings before claiming. + dmaBindOwnerRegionsInto(task_id, device_id); + } architecture.setSystemCallResult(state, 0); } diff --git a/system/services/device-manager/device-manager.zig b/system/services/device-manager/device-manager.zig index 6b87d14..92ec1e0 100644 --- a/system/services/device-manager/device-manager.zig +++ b/system/services/device-manager/device-manager.zig @@ -237,6 +237,25 @@ fn alreadySupervised(name: []const u8) bool { return false; } +/// Drivers that receive their device from the manager rather than claiming it +/// themselves. Scaffolding for the conversion, not a permanent concept: it exists so +/// each driver can move across one at a time with the suite green throughout, and it +/// disappears at D6 when `device_claim` stops being a way to acquire a device at all +/// (docs/bounds-track-plan.md, Run 2). +/// +/// `usb-xhci-bus` is first because it was the first driver to conform to `hello` +/// (device-manager.md, M18.1), so it is the one whose handshake is best proven. +const delegated_drivers = [_][]const u8{ + "/system/drivers/usb-xhci-bus", +}; + +fn isDelegated(name: []const u8) bool { + for (delegated_drivers) |candidate| { + if (std.mem.eql(u8, name, candidate)) return true; + } + return false; +} + /// Record a driver in the table and spawn its first instance. fn addDriver(name: []const u8, device_id: u64, speaks_protocol: bool) void { for (&drivers) |*driver| { @@ -257,6 +276,22 @@ fn addDriver(name: []const u8, device_id: u64, speaks_protocol: bool) void { /// device id as argv[1] when it has one, the hello deadline armed when it /// speaks the protocol. fn spawnDriver(driver: *Driver) void { + // Take the device before the driver exists, so there is no window in which anyone + // else could claim it — which is the whole of what makes the handover authoritative + // rather than advisory. Re-claiming across a restart is expected to say + // AlreadyClaimed once the manager already holds it, and that is fine: it means the + // device never left our hands while the driver was dead. + if (isDelegated(driver.name()) and driver.device_id != device_manager_protocol.no_device) { + device.claim(driver.device_id) catch |e| switch (e) { + error.AlreadyClaimed => {}, // ours already, from a previous spawn of this driver + else => { + std.log.warn("cannot hold device {d} for {s}: {s}", .{ driver.device_id, driver.name(), @errorName(e) }); + driver.state = .failed; + return; + }, + }; + } + var id_text: [20]u8 = undefined; var arguments: [1][]const u8 = undefined; var argument_count: usize = 0; @@ -424,6 +459,23 @@ fn onHello(_: void, invocation: Invocation(device_manager_protocol.Hello), _: An return -envelope.EPERM; }; driver.state = .running; + + // **Delegation.** The manager holds this driver's device and hands it over here — + // what replaces first-come-first-served `device_claim` with policy + // (docs/os-development/device-authority.md). `invocation.sender` is the driver's + // task id stamped by the kernel, so the manager cannot be lied to about who is + // asking, and the transfer is a move: the manager stops holding it. + // + // Gated on the delegated set so an unconverted driver still claims for itself and + // its path is untouched; the set and `device_claim` both go at D6. + if (isDelegated(driver.name()) and driver.device_id != device_manager_protocol.no_device) { + device.transfer(driver.device_id, invocation.sender) catch |e| { + std.log.warn("could not delegate device {d} to {s}: {s}", .{ driver.device_id, driver.name(), @errorName(e) }); + return -envelope.EPERM; + }; + std.log.info("delegated device {d} to {s}", .{ driver.device_id, driver.name() }); + } + std.log.info("hello from {s} (device {d})", .{ driver.name(), invocation.target }); // Resilience drill (V6): once, kill the virtio-gpu driver a moment after it hellos, so // the normal restart policy respawns it — the compositor must survive and re-attach. diff --git a/test/qemu_test.py b/test/qemu_test.py index b8b14af..e4b3f13 100644 --- a/test/qemu_test.py +++ b/test/qemu_test.py @@ -635,7 +635,13 @@ CASES = [ "smp": 4, "timeout": 150, # usb-kbd/usb-mouse ride the default boot xHCI bus (see qemu_args). - "expect": r"(?=[\s\S]*usb-xhci-bus: controller running \((\d+) slots, tracking \1,)" + # The controller must arrive by DELEGATION, not by claiming: the manager holds it + # and transfers it in the hello reply, so the match is authoritative rather than + # advisory (docs/os-development/device-authority.md). The delegation line must + # precede the hello, because the transfer completes before the reply lands. + "expect": r"(?=[\s\S]*device-manager: delegated device (\d+) to /system/drivers/usb-xhci-bus" + r"[\s\S]*usb-xhci-bus: controller device \1 registers at)" + r"(?=[\s\S]*usb-xhci-bus: controller running \((\d+) slots, tracking \2,)" r"(?=[\s\S]*usb-hid-keyboard: ok)(?=[\s\S]*usb-hid-mouse: ok)", "fail": r"DANOS-TEST-RESULT: FAIL"}, # Keyboard echo: inject a known phrase via QMP send-key; the usb-hid-keyboard From 81dd1e93184bf3249d25fb71d4855129232a15ff Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 18:25:16 +0100 Subject: [PATCH 17/36] pci: the host bridge arrives by delegation too MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit pci-bus joins usb-xhci-bus in receiving its device from the manager rather than claiming the id it found in argv[1]. Its hello moves ahead of the ECAM mapping, since that is where the bridge now arrives, and its hello was already mandatory so nothing about its failure behaviour changes. isDelegated compared whole strings, which silently missed this driver: the manager records the boot-snapshot match as the bare "pci-bus" and a devices.csv match as the full "/system/drivers/pci-bus". pci-bus was then neither claiming nor delegated and died on "ECAM mmio_map failed". It now matches on the last path component. Reintroducing the whole-string compare breaks usb-xhci-bus instead of pci-bus — the two spellings swap which driver loses — so usb-hid is the case that catches it, not pci-scan. pci-scan asserts the delegation on the initial bring-up AND after the restart drill, with the device id backreferenced so both must name the same device. That is what proves the manager re-takes a device when its driver dies and hands it to the replacement, which is the property the whole supervision design rests on. The remaining three claimants are NOT converted, and the plan records why rather than working around it. ps2-bus and the acpi service never hello at all, which device-manager.md states deliberately ("legacy drivers ... not yet required to hello"), so delegating to them means either promoting them out of legacy or giving the grant a delivery point that is not hello. virtio-gpu hellos best-effort by design — "standalone bring-up has no manager" — and delegation would make it mandatory. Both are decisions, not mechanical steps. Consequence: D6 is blocked, because device_claim cannot be closed off while three claimants still depend on it. D7-D9 are unaffected — they concern what the kernel stores and how its table is sized. Suite 118/118. --- docs/bounds-track-plan.md | 29 ++++++++++++++++++- system/drivers/pci-bus/pci-bus.zig | 20 ++++++++----- .../device-manager/device-manager.zig | 11 +++++-- test/qemu_test.py | 9 ++++-- 4 files changed, 56 insertions(+), 13 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 8ab7143..25e8952 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -43,7 +43,7 @@ that cannot safely run in user space.** | D2 | Adversarial case: a process handed nothing is refused, on a held device and a free one | **done** — `device-authority-test`; the claim half joins it at D6 | | D3 | The manager claims the seeded devices at boot, before any driver is spawned | **merged into D4** — see below | | D4 | The manager claims + delegates on `hello`; `usb-xhci-bus` is the first driver converted | **done** — caught an IOMMU regression I introduced; see below | -| D5 | The other four claimants converted: `pci-bus`, `ps2-bus`, `virtio-gpu`, `acpi` | not started | +| D5 | The other four claimants converted: `pci-bus`, `ps2-bus`, `virtio-gpu`, `acpi` | **partial** — `pci-bus` done; the other three need decisions, see below | | D6 | `device_claim` refuses a device the caller was not handed; the hole is closed | not started | | D7 | Zero-resource devices stop being kernel objects — inventory moves to the manager | not started | | D8 | **`maximum_children_per_parent` deleted** — the authorisation it stood in for exists | not started | @@ -80,6 +80,33 @@ They land together, with the manager claiming only for drivers in an explicit anything in D4 — recorded rather than dismissed, because D4 moved the `hello` earlier and so did shift boot timing. Watch it across the remaining steps. +### Open questions raised by D5 — three of the four claimants cannot be converted yet + +Delegation is delivered in `onHello`. That works for a driver that says hello, and +**two of the four do not**. + +6. **`ps2-bus` and the `acpi` service never hello at all**, and that is deliberate: + device-manager.md records "Legacy drivers (e.g. ps2-bus) are supervised and + restarted but **not yet required to hello**", and the `Driver.speaks_protocol` flag + exists to say so. Delegating to them means either making them speak the protocol — + promoting them out of "legacy", which is a change to their documented status — or + giving the grant a second delivery point that is not `hello`. Neither is written + down. A second delivery point would also need its own answer to "when", since the + whole value of `hello` here is that it is the moment the driver is known to be alive + and is a synchronous point to hand something over. + +7. **`virtio-gpu` hellos, but best-effort by design.** Its call is + `_ = device_manager.hello(.device, device_id);` with the comment "Best-effort: + standalone bring-up has no manager." Delegation would make the hello *mandatory* and + move it to the front, so the driver could no longer come up without a manager. No + test exercises standalone today (the `virtio-gpu` case boots the manager stack, and + devices.csv matches it), so this is a documented intent rather than a live path — + but discarding a documented intent is a decision, not a mechanical step. + +Until these are answered, `device_claim` cannot be closed off at D6 for those three, so +**D6 is blocked on question 6 and 7**. D7–D9 are not: they concern what the kernel +stores and how the table is sized, and are independent of which drivers have converted. + ### Settled, so the run does not re-litigate them - **The manager claims, it is not granted.** No binary names in the kernel; the rule is diff --git a/system/drivers/pci-bus/pci-bus.zig b/system/drivers/pci-bus/pci-bus.zig index 5cd81cc..77af176 100644 --- a/system/drivers/pci-bus/pci-bus.zig +++ b/system/drivers/pci-bus/pci-bus.zig @@ -75,11 +75,20 @@ fn configWrite16(bus: u64, dev: u64, function: u64, offset: u64, value: u16) voi configWrite(bus, dev, function, aligned, (word & ~mask) | (@as(u32, value) << shift)); } -/// Claim the bridge, map the ECAM, hello the manager, then scan. +/// Hello the manager (which is where the bridge arrives), map the ECAM, then scan. fn initialise(endpoint: ipc.Handle) bool { _ = endpoint; - device.claim(bridge_id) catch |e| { - std.log.info("unable to claim bridge device {d}: {s}", .{ bridge_id, @errorName(e) }); + // **The handshake first, because it is where the device arrives.** This driver + // used to claim `bridge_id` here — first-come-first-served, so the manager's + // matching was advisory and any process could have claimed the bridge by naming + // the same id. The manager now holds it and transfers it in the hello reply + // (docs/os-development/device-authority.md). `hello` is synchronous, so the + // transfer has completed by the time this returns. + // + // Keep the manager handle to report children through; a supervised bus that + // cannot reach its manager has nothing to serve. + manager_handle = device_manager.hello(.bus, bridge_id) orelse { + std.log.info("no hello with the device manager; bridge {d} not delegated", .{bridge_id}); return false; }; const buffer = memory.allocator().alloc(device.DeviceDescriptor, 64) catch { @@ -113,11 +122,6 @@ fn initialise(endpoint: ipc.Handle) bool { return false; }; - // The handshake (role: bus — we enumerate PCI and report the functions we - // find), then the scan. Keep the manager handle to report children through; - // a supervised bus that cannot reach its manager has nothing to serve. - manager_handle = device_manager.hello(.bus, bridge_id) orelse return false; - scan(); return true; } diff --git a/system/services/device-manager/device-manager.zig b/system/services/device-manager/device-manager.zig index 92ec1e0..f2aff2a 100644 --- a/system/services/device-manager/device-manager.zig +++ b/system/services/device-manager/device-manager.zig @@ -246,12 +246,19 @@ fn alreadySupervised(name: []const u8) bool { /// `usb-xhci-bus` is first because it was the first driver to conform to `hello` /// (device-manager.md, M18.1), so it is the one whose handshake is best proven. const delegated_drivers = [_][]const u8{ - "/system/drivers/usb-xhci-bus", + "usb-xhci-bus", + "pci-bus", }; +/// Matched on the **last path component**, because a driver reaches this table under +/// two different spellings: the boot-snapshot match records the bare `pci-bus`, while a +/// devices.csv match records the full `/system/drivers/pci-bus`. Comparing whole +/// strings silently missed the bare form — pci-bus was left neither claiming nor +/// delegated, and died on `ECAM mmio_map failed`. fn isDelegated(name: []const u8) bool { + const leaf = if (std.mem.lastIndexOfScalar(u8, name, '/')) |slash| name[slash + 1 ..] else name; for (delegated_drivers) |candidate| { - if (std.mem.eql(u8, name, candidate)) return true; + if (std.mem.eql(u8, leaf, candidate)) return true; } return false; } diff --git a/test/qemu_test.py b/test/qemu_test.py index e4b3f13..708f373 100644 --- a/test/qemu_test.py +++ b/test/qemu_test.py @@ -854,10 +854,15 @@ CASES = [ {"name": "pci-scan", "smp": 4, "timeout": 60, - "expect": r"pci-bus: (\d+) functions found[\s\S]*" + # The bridge must arrive by DELEGATION, not by claiming — and again after the + # restart, which is what proves the manager re-takes the device when its driver + # dies and hands it to the replacement. \1 pins it to the same device both times. + "expect": r"device-manager: delegated device (\d+) to \S*pci-bus[\s\S]*" + r"pci-bus: (\d+) functions found[\s\S]*" r"device-manager: test mode: killing the reporter[\s\S]*" r"device-manager: restarting \S*pci-bus[\s\S]*" - r"pci-bus: \1 functions found[\s\S]*" + r"device-manager: delegated device \1 to \S*pci-bus[\s\S]*" + r"pci-bus: \2 functions found[\s\S]*" r"DANOS-TEST-RESULT: PASS", "fail": r"DANOS-TEST-RESULT: FAIL"}, # M18.3: the application surface — device-list enumerates the tree over IPC, From 5492607a6164b22473f1b7e29dfc0a005b506590 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 18:26:37 +0100 Subject: [PATCH 18/36] docs: D7 blocked, and D8 reordered after D9 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit D7 would remove zero-resource devices from the kernel. It cannot proceed because nothing else mints their ids. A USB interface registers with resource_count = 0, and the id device_register hands back is load-bearing in three places: the child_added packet's target, the class driver's argv[1], and the device_token of the usb-transfer WIRE protocol — so the id space is visible on the wire, not merely internal. usb-xhci-bus records the fourth constraint itself: the kernel's idempotency is what makes the same port and interface map back to the same id across a bus restart, which is what stops a respawned bus spawning duplicate class drivers. Moving that out means answering who mints the id, how it survives a bus restart, how it survives a manager restart, and whether device_token changes meaning. A design step, not a mechanical one. D8 is reordered to run after D9, correcting the original sequencing. D8's justification was that the authorisation the per-parent cap stood in for now exists — but D6 is blocked, so it does not, and deleting the shared cap now would reopen the exhaustion hole it was written for. D9's per-holder quota closes that hole independently of authorisation, and closes it better: a rogue exhausts its own allowance instead of the table everyone shares. Once the quota exists the per-parent cap is redundant either way. D9 also no longer depends on D7. Its rationale was that zero-resource children are the case that sidesteps containment — true, but a per-holder quota bounds them as well as anything else, because it counts entries per holder rather than per parent. --- docs/bounds-track-plan.md | 36 +++++++++++++++++++++++++++++++----- 1 file changed, 31 insertions(+), 5 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 25e8952..01cce41 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -45,9 +45,9 @@ that cannot safely run in user space.** | D4 | The manager claims + delegates on `hello`; `usb-xhci-bus` is the first driver converted | **done** — caught an IOMMU regression I introduced; see below | | D5 | The other four claimants converted: `pci-bus`, `ps2-bus`, `virtio-gpu`, `acpi` | **partial** — `pci-bus` done; the other three need decisions, see below | | D6 | `device_claim` refuses a device the caller was not handed; the hole is closed | not started | -| D7 | Zero-resource devices stop being kernel objects — inventory moves to the manager | not started | -| D8 | **`maximum_children_per_parent` deleted** — the authorisation it stood in for exists | not started | -| D9 | The device table becomes dynamic; **`maximum_devices` deleted**; per-holder quota declared | not started | +| D7 | Zero-resource devices stop being kernel objects — inventory moves to the manager | **blocked** — nothing else mints their ids; see question 8 | +| D8 | **`maximum_children_per_parent` deleted** — the authorisation it stood in for exists | **blocked on D6**, and now ordered after D9 | +| D9 | The device table becomes dynamic; **`maximum_devices` deleted**; per-holder quota declared | not started — **runs before D8** | Ordering is load-bearing. D1–D2 build and prove the mechanism with nothing depending on it. D4–D5 move each claimant across one at a time, so the suite stays green throughout @@ -104,8 +104,34 @@ Delegation is delivered in `onHello`. That works for a driver that says hello, a but discarding a documented intent is a decision, not a mechanical step. Until these are answered, `device_claim` cannot be closed off at D6 for those three, so -**D6 is blocked on question 6 and 7**. D7–D9 are not: they concern what the kernel -stores and how the table is sized, and are independent of which drivers have converted. +**D6 is blocked on questions 6 and 7**. + +8. **Removing zero-resource devices from the kernel needs someone else to mint their + ids, and nothing says who.** A USB interface is registered with + `resource_count = 0`, and the id `device_register` returns is load-bearing in three + places: it is the `child_added` packet's target, it is the class driver's `argv[1]`, + and it is the `device_token` of the **usb-transfer wire protocol** — so the id space + is visible on the wire, not merely internal. usb-xhci-bus's own comment records the + fourth constraint: the kernel's idempotency is what makes "the same port and + interface always map back to the same device id" across a bus restart, which is what + stops a respawned bus spawning duplicate class drivers. + + So the mover has to answer: who mints the id, how it stays stable across a *bus* + restart, how it stays stable across a *manager* restart (open question 3 territory), + and whether the wire protocol's `device_token` changes meaning. That is a design + step, not a mechanical one. + +**D8 is reordered to run after D9**, which is a correction to the original sequencing. +D8's justification was "the authorisation it stood in for exists" — but with D6 blocked +it does not, so deleting the shared per-parent cap now would reopen the exhaustion hole +it was written for. D9's **per-holder quota** closes that hole independently of +authorisation, and does it better: a rogue exhausts its own allowance rather than the +table everyone shares. Once the quota exists, the per-parent cap is redundant whether or +not D6 has landed. + +D9 also no longer depends on D7. Its original rationale was that zero-resource children +are the case that sidesteps containment — true, but a per-holder quota bounds them just +as well as anything else, because it counts entries per holder rather than per parent. ### Settled, so the run does not re-litigate them From f7151ed5778ed49102ab54a6dde29775610edd0a Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 18:40:20 +0100 Subject: [PATCH 19/36] kernel: the device table has no ceiling; a runaway is charged to whoever caused it MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit maximum_devices = 64 is gone. It was a guess about someone else's computer, and because it was shared, one driver's enumeration starved every other — which is how an AMD Ryzen booted with a working display, no USB and no storage. The table now grows from the kernel heap. It was always built after heap.init; nothing ever prevented this except it having been written static first. What replaces it is an allowance charged to the registrar, so a driver looping device_register exhausts its own and every other driver carries on. It is declared as what it is — a runaway detector, NOT a security boundary. A quota generous enough never to bite a real machine is still generous enough to be unpleasant, and it is not trying to be the defence; delegation is. What this catches is a legitimate driver in a loop, early, attributably, and without collateral. Reaching 4096 is a bug report, not a tuning request. The initial block is 8, deliberately small. Sizing it for a typical machine would mean the growth path never ran on the hardware we test on and only woke up on someone else's larger machine — the exact failure shape this track exists to stop. At 8 it grows several times every boot; disabling growth now fails the suite with the HPET not fitting, which is the Ryzen failure in miniature. The comptime coupling assert added earlier fired, and was right to. confined (one slot per device id) and domains (the IOMMU's own translation pool) were sized by the same constant only because device ids happened to stop at 64 too. Two unrelated quantities: confined now grows with the device table, while maximum_domains stays as the hardware's number — both VT-d and AMD-Vi report how many domains they support, and reading it is phase 4. The assert existed for exactly this and did its job. Suite 118/118. --- docs/bounds-track-plan.md | 6 +- system/kernel/devices-broker.zig | 106 ++++++++++++++++++++++++++----- system/kernel/iommu.zig | 60 ++++++++++------- system/kernel/tests.zig | 11 ++++ 4 files changed, 142 insertions(+), 41 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 01cce41..747387c 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -47,7 +47,11 @@ that cannot safely run in user space.** | D6 | `device_claim` refuses a device the caller was not handed; the hole is closed | not started | | D7 | Zero-resource devices stop being kernel objects — inventory moves to the manager | **blocked** — nothing else mints their ids; see question 8 | | D8 | **`maximum_children_per_parent` deleted** — the authorisation it stood in for exists | **blocked on D6**, and now ordered after D9 | -| D9 | The device table becomes dynamic; **`maximum_devices` deleted**; per-holder quota declared | not started — **runs before D8** | +| D9 | The device table becomes dynamic; **`maximum_devices` deleted**; per-holder quota declared | **done** — one of the two invented numbers is gone | + +**Run 2 stops here.** D1, D2, D4, D5 (`pci-bus` only) and D9 landed; D6, D7, D8 and the +rest of D5 are blocked on questions 6, 7 and 8 below. Suite 118/118, and +`maximum_devices` no longer exists. Ordering is load-bearing. D1–D2 build and prove the mechanism with nothing depending on it. D4–D5 move each claimant across one at a time, so the suite stays green throughout diff --git a/system/kernel/devices-broker.zig b/system/kernel/devices-broker.zig index 6dc7bdc..286e370 100644 --- a/system/kernel/devices-broker.zig +++ b/system/kernel/devices-broker.zig @@ -22,16 +22,36 @@ const std = @import("std"); const abi = @import("abi"); const platform = @import("platform"); const device_abi = @import("device-abi"); +const heap = @import("heap.zig"); -/// bound: device nodes for the whole machine — firmware-discovered plus every child a -/// bus driver registers at runtime -/// decided-by: hardware -/// protects: nothing — this is a sizing guess about someone else's computer, which is -/// why an AMD Ryzen booted with a working display, no USB and no storage -/// at-limit: refuse — ENOSPC from device_register; `dropped` counts discovery losses +/// How many devices a single registrar may put in the table. +/// +/// **This is a runaway detector, not a security boundary**, and the difference matters. +/// It cannot stop a malicious driver — a quota generous enough never to bite a real +/// machine is still generous enough to be unpleasant — and it is not trying to. What +/// stops malice is that only a driver the manager handed a device can register children +/// under it (docs/os-development/device-authority.md). What this catches is a +/// *legitimate* driver in a loop, early, and attributably: the driver that did it is +/// refused, named in its own log, and restarted, while every other driver is untouched. +/// +/// The shared ceilings it replaces could not do that. `maximum_devices = 64` was a +/// guess about someone else's computer and one driver's enumeration starved every +/// other — which is how an AMD Ryzen came to boot with a working display, no USB and no +/// storage. A per-registrar allowance is the same protection charged to whoever caused +/// it, which is the microkernel property rather than a workaround for it. +/// +/// The number: a machine's whole PCI segment tops out at 65536 functions, and the +/// biggest real registrar seen is pci-bus at a few dozen. 4096 is far above anything a +/// real machine produces and around 1.4 MiB of descriptors, well under the kernel heap. +/// **Reaching it is a bug report, not a tuning request** — no correct driver gets near. +/// +/// bound: devices one task may register +/// decided-by: ours +/// protects: the kernel heap, against a driver looping device_register +/// at-limit: refuse — ECHILDREN to the registrar; every other driver is unaffected /// observed-by: the bus driver's line naming the reason (pci-bus reconciles functions -/// found against registered), and kernel.zig:203 for discovery drops -pub const maximum_devices = 64; +/// found against registered) +const maximum_devices_per_registrar = 4096; /// Cap on children a single parent may have. A zero-resource child (legal — a USB /// device is addressed through its controller, not by MMIO) sidesteps the containment @@ -50,18 +70,64 @@ pub const maximum_devices = 64; /// observed-by: pci-bus logs the reason per refused function, and warns at end of scan const maximum_children_per_parent = 16; -var devices: [maximum_devices]device_abi.DeviceDescriptor = undefined; -var claimed: [maximum_devices]?u32 = .{null} ** maximum_devices; // owner task id, or null +/// The table, grown on demand from the kernel heap. **There is no ceiling**: how many +/// devices a machine has is the machine's business, and no specification bounds it, so +/// nothing here should. `devices_broker.init` runs after `heap.init` (kernel.zig), so +/// there was never a reason for this to be static beyond it having been written that +/// way first. +var devices: []device_abi.DeviceDescriptor = &.{}; +/// Owner task id per device, or null. Parallel to `devices` and grown with it. +var claimed: []?u32 = &.{}; +/// The task that called `register` for each device, so the per-registrar allowance can +/// be charged to whoever caused the entry. Firmware-discovered nodes carry `no_registrar` +/// — they are the kernel's own, not anybody's doing. +var registrar: []u32 = &.{}; +const no_registrar: u32 = 0; var count: usize = 0; +/// Grow the three parallel arrays so at least one more device fits. False if the heap +/// cannot satisfy it, which the callers report rather than swallow. +fn reserve() bool { + if (count < devices.len) return true; + // Double from a deliberately SMALL first block. Sizing it for a typical machine + // would mean the growth path never ran on the hardware we test on, and only woke up + // on someone else's larger machine — which is the exact shape of the failure this + // whole track exists to stop. At 8, every boot grows the table several times, so + // the path is exercised constantly and the suite asserts it. + const wanted = if (devices.len == 0) 8 else devices.len * 2; + const allocator = heap.allocator(); + const grown_devices = allocator.realloc(devices, wanted) catch return false; + devices = grown_devices; + const grown_claimed = allocator.realloc(claimed, wanted) catch return false; + claimed = grown_claimed; + const grown_registrar = allocator.realloc(registrar, wanted) catch return false; + registrar = grown_registrar; + for (claimed[count..], registrar[count..]) |*slot, *who| { + slot.* = null; + who.* = no_registrar; + } + return true; +} + +/// How many devices `task` has registered — the allowance is charged per registrar, so +/// a driver in a loop exhausts its own and no one else's. +fn registeredBy(task: u32) usize { + var n: usize = 0; + for (registrar[0..count]) |who| { + if (who == task) n += 1; + } + return n; +} + /// The id of the seeded framebuffer node (`seedDisplay`), or null when the machine /// handed over no framebuffer. Lets the process layer recognise the display claim /// (to quiesce the bootstrap console) without threading the id through every caller. var display_device: ?u64 = null; -/// Devices discovery found but the table had no room for. Non-zero means the machine -/// is bigger than `maximum_devices` and some hardware is simply invisible to drivers — -/// which would otherwise be an entirely silent failure. Logged at boot. +/// Devices discovery found but could not record. The table grows on demand, so this is +/// no longer "the machine is bigger than our guess" — it means the kernel heap could not +/// satisfy the growth, which would otherwise be an entirely silent failure. Logged at +/// boot. pub var dropped: usize = 0; /// Snapshot the device tree into the flat table. Run once, right after discovery. @@ -69,7 +135,8 @@ pub fn init(device_tree: *const platform.DeviceTree) void { count = 0; dropped = 0; display_device = null; - for (&claimed) |*c| c.* = null; + for (claimed) |*c| c.* = null; + for (registrar) |*r| r.* = no_registrar; walk(device_tree.root, device_abi.no_parent); } @@ -81,7 +148,7 @@ pub fn init(device_tree: *const platform.DeviceTree) void { /// table is full. Idempotent-ish: only ever call once per boot. pub fn seedDisplay(base: u64, width: u32, height: u32, pitch: u32, format: u32, refresh_hz: u32) ?u64 { if (base == 0 or width == 0 or height == 0) return null; // headless - if (count >= maximum_devices) { + if (!reserve()) { dropped += 1; return null; } @@ -126,7 +193,7 @@ fn walk(node: *platform.Device, parent_id: u64) void { } fn record(node: *platform.Device, parent_id: u64) u64 { - if (count >= maximum_devices) { + if (!reserve()) { dropped += 1; return device_abi.no_parent; // children of a dropped node become roots, not orphans } @@ -431,7 +498,11 @@ pub fn register(parent_id: u64, owner: u32, descriptor: *const device_abi.Device if (existingChild(parent_id, descriptor)) |existing_id| return existing_id; if (childCount(parent_id) >= maximum_children_per_parent) return error.TooManyChildren; - if (count >= maximum_devices) return error.NoSpace; + // The allowance is charged to whoever is registering, so a driver in a loop + // exhausts its own and every other driver carries on. There is no machine-wide + // ceiling any more: the table grows. + if (registeredBy(owner) >= maximum_devices_per_registrar) return error.TooManyChildren; + if (!reserve()) return error.NoSpace; const parent = &devices[@intCast(parent_id)]; for (0..@intCast(descriptor.resource_count)) |i| { @@ -454,6 +525,7 @@ pub fn register(parent_id: u64, owner: u32, descriptor: *const device_abi.Device for (0..@intCast(descriptor.resource_count)) |i| d.resources[i] = descriptor.resources[i]; devices[count] = d; + registrar[count] = owner; count += 1; return d.id; } diff --git a/system/kernel/iommu.zig b/system/kernel/iommu.zig index 7166fc4..5afb203 100644 --- a/system/kernel/iommu.zig +++ b/system/kernel/iommu.zig @@ -34,22 +34,24 @@ const platform = @import("platform"); const architecture = @import("architecture"); const devices_broker = @import("devices-broker.zig"); const log = @import("log.zig"); +const heap = @import("heap.zig"); const page_size: u64 = abi.page_size; const page_mask: u64 = page_size - 1; const huge_page_size: u64 = 2 * 1024 * 1024; -/// One domain per claimed PCI function. Coupled to devices-broker's device cap — and -/// coupled *in code*, by the comptime assert beside `confined` below, because when -/// these two agreed only by this sentence the disagreement failed open. +/// The IOMMU's own translation-domain pool — one per claimed DMA-capable device. /// -/// Both VT-d and AMD-Vi report the number of domains they support in a capability -/// register. We should be reading it rather than choosing 64 — docs/bounds-track-plan.md -/// phase 4. +/// No longer coupled to the device count. It was, by a comment and then by a comptime +/// assert, only because `confined` (one slot per device id) was sized by this same +/// constant; those are two unrelated quantities and making the device table dynamic +/// separated them. This one is genuinely the hardware's: both VT-d and AMD-Vi report +/// how many domains they support in a capability register, so the honest fix is to read +/// it rather than choose 64 — bounds-track-plan.md phase 4. /// -/// bound: IOMMU translation domains, one per claimed DMA-capable device +/// bound: IOMMU translation domains the kernel can hold at once /// decided-by: hardware -/// protects: the statically sized domain and confinement tables +/// protects: the statically sized domain pool /// at-limit: refuse — ECONFINE; the claim is rolled back and the device is not driven, /// because a claim that cannot be confined must not stand /// observed-by: the claiming driver's own line naming ECONFINE @@ -107,18 +109,30 @@ pub fn init() void { /// Per-claimed-device record: its private domain, so a driver's death tears down /// exactly the domains it held. const Confined = struct { active: bool = false, owner: u32 = 0, bdf: u16 = 0, domain: u16 = invalid_domain }; -var confined: [maximum_domains]Confined = .{Confined{}} ** maximum_domains; -// `confined` is indexed by **device id**, so it must cover every id the broker can -// mint. These two numbers agreed only by a sentence in a comment above -// `maximum_domains` — and when they disagreed, `confineDevice` returned success for -// the ids it had no room for, leaving those devices unconfined DMA masters. Coupled -// bounds agree in code, not in prose (docs/os-development/bounds.md). -comptime { - if (maximum_domains < devices_broker.maximum_devices) - @compileError("iommu.confined is indexed by device id but is smaller than the " ++ - "broker's device table: ids past its end cannot be confined, and so cannot " ++ - "be claimed at all"); +/// Indexed by **device id**, so it must cover every id the broker can mint — and the +/// broker's table has no ceiling any more, so neither can this. It grows on demand. +/// +/// This used to be `[maximum_domains]`, sized by the *domain* constant purely because +/// device ids happened to stop at 64 as well. Two unrelated quantities sharing one +/// number: `domains` below is the IOMMU's own translation-domain pool, which the +/// hardware bounds and reports, while this is one slot per device the machine has. +/// A comptime assert held them together while both were fixed; making the device table +/// dynamic is what forced them apart, which is the assert having done its job. +var confined: []Confined = &.{}; + +/// Grow `confined` to cover `device_id`. False if the heap cannot — and the caller +/// treats that as a refusal to confine, never as permission. +fn reserveConfined(device_id: u64) bool { + if (device_id < confined.len) return true; + if (device_id >= std.math.maxInt(usize) / 2) return false; // absurd id; refuse rather than size to it + var wanted: usize = if (confined.len == 0) 64 else confined.len; + while (wanted <= device_id) wanted *= 2; + const grown = heap.allocator().realloc(confined, wanted) catch return false; + const previous = confined.len; + confined = grown; + for (confined[previous..]) |*record| record.* = .{}; + return true; } /// Place a just-claimed PCI function under IOMMU translation on behalf of `owner`: give @@ -139,7 +153,7 @@ pub fn confineDevice(device_id: u64, bdf: u16, owner: u32) bool { // at `confined.len`; moving the inventory out of the kernel and taking the domain // count from the hardware both change that, and either would have made a silent // unconfined DMA master out of every device past the 64th. - if (device_id >= confined.len) return false; + if (!reserveConfined(device_id)) return false; const domain = domainCreate(owner, bdf) orelse return false; // Firmware reserved region for this device, if any (real hardware; QEMU has none). @@ -181,7 +195,7 @@ pub fn unmapForDevice(device_id: u64, physical: u64, len: u64) void { /// own freshly-`dma_alloc`'d buffer into the devices it drives. pub fn mapRegionForOwner(owner: u32, physical: u64, len: u64) void { if (!active) return; - for (&confined) |*c| { + for (confined) |*c| { if (c.active and c.owner == owner) _ = map(c.domain, physical, len); } } @@ -192,7 +206,7 @@ pub fn mapRegionForOwner(owner: u32, physical: u64, len: u64) void { /// domain other than its owner's. pub fn unmapRegionEverywhere(physical: u64, len: u64) void { if (!active) return; - for (&confined) |*c| { + for (confined) |*c| { if (c.active) unmap(c.domain, physical, len); } } @@ -222,7 +236,7 @@ pub fn reassign(device_id: u64, owner: u32) void { pub fn releaseAllOwnedBy(owner: u32) void { if (!active) return; - for (&confined) |*c| { + for (confined) |*c| { if (c.active and c.owner == owner) { detachDevice(c.bdf); domainDestroy(c.domain); diff --git a/system/kernel/tests.zig b/system/kernel/tests.zig index d1849f0..5b16627 100644 --- a/system/kernel/tests.zig +++ b/system/kernel/tests.zig @@ -4020,6 +4020,17 @@ fn containmentTest() void { const still_capped = if (devices_broker.register(parent_id, me, &novel)) |_| false else |err| err == error.TooManyChildren; check("a full parent still refuses a new child", still_capped); + // The table itself has no ceiling: it grows. The old `maximum_devices = 64` was a + // guess about someone else's computer, and one driver's enumeration starved every + // other — which is how a Ryzen booted with no USB and no storage. What bounds a + // runaway now is an allowance charged to the registrar, so the damage stays with + // whoever caused it (docs/os-development/device-authority.md). + // The table grows: it starts at 8 entries and this boot holds well past that, so + // the growth path runs every time rather than lying dormant until someone else's + // larger machine finds it — which is how the old ceiling stayed invisible. + const held = devices_broker.enumerate(&buffer); + check("the table grew beyond its initial block", held > 8); + result(); } From a5840789ddf1cb356b0170a177d2c776fbf72534 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 19:14:54 +0100 Subject: [PATCH 20/36] docs: the acpi service is discovery, not a driver awaiting an assignment MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Question 6 lumped the acpi service in with ps2-bus as "a claimant that does not hello". That was wrong about what it is. The manager starts it as addDriver("discovery", no_device, false): the discovery service, one per firmware, with no device assignment at all, which finds and claims the acpi-tables node itself because it is the thing that produces the device tree. The manager cannot hand it a device — at that point there is nothing to match against. So its question is not whether it should hello. It is whether the bootstrap is exempt from D6, or whether the manager claims the acpi-tables node and passes it on. Recorded alongside it: a delivery point that needs no hello is already implied by the code, since spawnSupervised returns the child's pid and the manager could transfer immediately after spawn. If that is acceptable, question 6 largely dissolves — ps2-bus would not need to leave "legacy" either. The cost is ordering: the child may reach for the device before the transfer lands, where hello guarantees it cannot because the child is the one asking. That trade is the real decision, and it is now written down as such. --- docs/bounds-track-plan.md | 30 +++++++++++++++++++++--------- 1 file changed, 21 insertions(+), 9 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 747387c..5954b4f 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -89,15 +89,27 @@ They land together, with the manager claiming only for drivers in an explicit Delegation is delivered in `onHello`. That works for a driver that says hello, and **two of the four do not**. -6. **`ps2-bus` and the `acpi` service never hello at all**, and that is deliberate: - device-manager.md records "Legacy drivers (e.g. ps2-bus) are supervised and - restarted but **not yet required to hello**", and the `Driver.speaks_protocol` flag - exists to say so. Delegating to them means either making them speak the protocol — - promoting them out of "legacy", which is a change to their documented status — or - giving the grant a second delivery point that is not `hello`. Neither is written - down. A second delivery point would also need its own answer to "when", since the - whole value of `hello` here is that it is the moment the driver is known to be alive - and is a synchronous point to hand something over. +6. **`ps2-bus` does not hello**, and that is deliberate: device-manager.md records + "Legacy drivers (e.g. ps2-bus) are supervised and restarted but **not yet required + to hello**", and `Driver.speaks_protocol` exists to say so. Either it leaves + "legacy", or the grant gets a delivery point that is not `hello`. + + **The `acpi` service is not this case — an earlier version of this question wrongly + lumped it in.** It is spawned as `addDriver("discovery", no_device, false)`: the + *discovery service*, one per firmware, with **no device assignment at all**, which + "finds and claims the acpi-tables (or devicetree-blob) node itself. Not a per-device + driver." The manager cannot hand it a device because discovery is what produces the + device tree — there is nothing to match against yet. Its question is not "should it + hello" but "is the bootstrap exempt from D6, or does the manager claim the + acpi-tables node and pass it on". + + **A delivery point that needs no hello already exists in the shape of the code**: + `process.spawnSupervised` returns the child's pid, so the manager could transfer + immediately after spawn. If that is acceptable, this question mostly dissolves — + `ps2-bus` would not need to leave "legacy" and discovery could be handed its node + too. The cost is ordering: the child may reach for the device before the transfer + lands, where `hello` guarantees it cannot because the child is the one asking. That + trade is the actual decision. 7. **`virtio-gpu` hellos, but best-effort by design.** Its call is `_ = device_manager.hello(.device, device_id);` with the comment "Best-effort: From 8b628a4a7dc9ef8e2626388c54469ce49213cafc Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 19:22:22 +0100 Subject: [PATCH 21/36] docs: there is no virtio-gpu standalone bring-up to lose MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Question 7 read the comment "Best-effort: standalone bring-up has no manager" as a boot path that delegation would delete. It is not one. virtio-gpu's main requires argv[1] and exits without it, and the only source of that argument is the device manager: ACPI, then pci-bus reports the function, then devices.csv matches 1AF4:1050, then the manager spawns the driver with the id. Without a manager the driver prints and returns, and never reaches the hello at all. The comment is about resilience, not boot: the hello is best-effort so a driver whose manager has died keeps serving, which device-manager.md states outright. So the question dissolves. The residual is one narrow window — the manager spawns a driver and dies before transferring — where today the driver could still claim because claiming is free-for-all, and after D6 it would exit for the restarted manager to respawn. That is the better behaviour: a driver holding hardware nobody assigned it is exactly what D6 exists to stop. Taken with the transfer-at-spawn delivery point, D5, D6 and D8 are unblocked. D7 remains blocked on question 8, and is not needed for either ceiling. --- docs/bounds-track-plan.md | 23 ++++++++++++++++------- 1 file changed, 16 insertions(+), 7 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 5954b4f..a251b73 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -111,13 +111,22 @@ Delegation is delivered in `onHello`. That works for a driver that says hello, a lands, where `hello` guarantees it cannot because the child is the one asking. That trade is the actual decision. -7. **`virtio-gpu` hellos, but best-effort by design.** Its call is - `_ = device_manager.hello(.device, device_id);` with the comment "Best-effort: - standalone bring-up has no manager." Delegation would make the hello *mandatory* and - move it to the front, so the driver could no longer come up without a manager. No - test exercises standalone today (the `virtio-gpu` case boots the manager stack, and - devices.csv matches it), so this is a documented intent rather than a live path — - but discarding a documented intent is a decision, not a mechanical step. +7. **`virtio-gpu`'s "standalone bring-up" — resolved; there is no such path.** An + earlier version of this question read the comment "Best-effort: standalone bring-up + has no manager" as a boot path that delegation would delete. It is not one. + `virtio-gpu`'s `main` requires `argv[1]` and exits without it, and the only source of + that argument is the device manager — the chain is ACPI → pci-bus reports the + function → `devices.csv` matches `1AF4:1050` → the manager spawns the driver with the + id. There is no way to run it without a manager, so nothing is lost. + + What the comment is really about is **resilience**: the hello is best-effort so a + driver whose manager has *died* keeps serving, which device-manager.md states + ("If the manager dies, drivers keep running"). + + The residual is one narrow window: the manager spawns a driver and dies before + transferring. Today the driver could still claim, because claiming is free-for-all; + after D6 it exits and the restarted manager respawns it. That is the better + behaviour — a driver holding hardware nobody assigned it is what D6 exists to stop. Until these are answered, `device_claim` cannot be closed off at D6 for those three, so **D6 is blocked on questions 6 and 7**. From cb8d1e2e51a2791e330ebbac5e6baae6a75cd1b0 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 19:26:32 +0100 Subject: [PATCH 22/36] docs: the grant rides system_spawn, atomically MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Questions 6 and 7 both dissolved on inspection, so what was left was only where the grant is delivered. Three candidates: transfer after spawn, every driver hellos, or fuse the device into system_spawn. Take the third. The manager cannot transfer before the child exists, so a separate transfer always leaves a window in which the child is running and does not yet hold its device. That window would close on QEMU every time and open occasionally on a machine with different timing — the exact failure shape this track exists to delete, and not worth introducing while removing the others. Fusing it into the spawn removes the window by construction: the child does not exist until it holds the device. It adds no knowledge to the kernel, only atomicity — the same rule, you may give away what you hold, fused with the call that creates the recipient. system_spawn uses five of six argument registers, and no_device is already the sentinel. Making every driver hello is a good idea on its own merits — uniform liveness, the deadline applied to all rather than some, and the speaks_protocol two-class split leaving the manager, since a wedged ps2-bus is invisible to its supervisor today. Kept as its own step so grant delivery does not force it. Order is now D0 -> D5 -> D6 -> D8, which deletes maximum_children_per_parent. Only D7 remains blocked, on question 8, and it is needed for neither ceiling. --- docs/bounds-track-plan.md | 40 ++++++++++++++++++++++++++++++++++----- 1 file changed, 35 insertions(+), 5 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index a251b73..541eeef 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -43,15 +43,18 @@ that cannot safely run in user space.** | D2 | Adversarial case: a process handed nothing is refused, on a held device and a free one | **done** — `device-authority-test`; the claim half joins it at D6 | | D3 | The manager claims the seeded devices at boot, before any driver is spawned | **merged into D4** — see below | | D4 | The manager claims + delegates on `hello`; `usb-xhci-bus` is the first driver converted | **done** — caught an IOMMU regression I introduced; see below | -| D5 | The other four claimants converted: `pci-bus`, `ps2-bus`, `virtio-gpu`, `acpi` | **partial** — `pci-bus` done; the other three need decisions, see below | +| D5 | The other four claimants converted: `pci-bus`, `ps2-bus`, `virtio-gpu`, `acpi` | **partial** — `pci-bus` done; the rest unblocked by D0 below | +| D0 | The grant rides `system_spawn` — atomic, so no driver need change to receive one | not started — **do first** | +| D10 | Every driver hellos, on its own merits (liveness, one class of driver) | not started — optional, independent | | D6 | `device_claim` refuses a device the caller was not handed; the hole is closed | not started | | D7 | Zero-resource devices stop being kernel objects — inventory moves to the manager | **blocked** — nothing else mints their ids; see question 8 | | D8 | **`maximum_children_per_parent` deleted** — the authorisation it stood in for exists | **blocked on D6**, and now ordered after D9 | | D9 | The device table becomes dynamic; **`maximum_devices` deleted**; per-holder quota declared | **done** — one of the two invented numbers is gone | -**Run 2 stops here.** D1, D2, D4, D5 (`pci-bus` only) and D9 landed; D6, D7, D8 and the -rest of D5 are blocked on questions 6, 7 and 8 below. Suite 118/118, and -`maximum_devices` no longer exists. +**Run 2 resumes at D0.** D1, D2, D4, D5 (`pci-bus` only) and D9 landed; `maximum_devices` +no longer exists and the suite is 118/118. Questions 6 and 7 dissolved, so the order is +now **D0 → D5 → D6 → D8**, which deletes `maximum_children_per_parent`. Only D7 is still +blocked, on question 8, and it is needed for neither ceiling. D10 is optional. Ordering is load-bearing. D1–D2 build and prove the mechanism with nothing depending on it. D4–D5 move each claimant across one at a time, so the suite stays green throughout @@ -84,7 +87,34 @@ They land together, with the manager claiming only for drivers in an explicit anything in D4 — recorded rather than dismissed, because D4 moved the `hello` earlier and so did shift boot timing. Watch it across the remaining steps. -### Open questions raised by D5 — three of the four claimants cannot be converted yet +### Settled 2026-08-08: the grant rides `system_spawn` (D0) + +Questions 6 and 7 both dissolved on inspection — neither `ps2-bus` nor discovery needs +to start speaking `hello`, and `virtio-gpu` has no standalone path to lose. What remains +is *where the grant is delivered*, and there are three candidates: + +| | Race window | Cost | +|---|---|---| +| Transfer after spawn | **yes** | none | +| Every driver hellos | no | `ps2-bus` + discovery gain a handshake | +| **Grant rides `system_spawn`** | **no — atomic** | one more syscall argument | + +**Take the third.** The manager cannot transfer before the child exists, so a separate +transfer always leaves a window in which the child is running and does not yet hold its +device. It would close on QEMU every time and open occasionally on a machine with +different timing — the exact failure shape this track exists to delete, and not worth +introducing while removing the others. Fusing the device into the spawn removes it by +construction: the child does not exist until it holds the device. No new knowledge in +the kernel — the same rule, *you may give away what you hold*, made atomic with the call +that creates the recipient. `system_spawn` uses five of six argument registers, so there +is room, and `no_device` is already the sentinel for a driver with no assignment. + +**`hello` for every driver is a good idea on its own merits** — uniform liveness, the +deadline applied to all rather than some, and the `speaks_protocol` two-class split +leaving the manager (a wedged `ps2-bus` is invisible to its supervisor today). It is +D10, kept separate so grant delivery does not force it. + +### Open questions raised by D5 — resolved except question 8 Delegation is delivered in `onHello`. That works for a driver that says hello, and **two of the four do not**. From 637bf2e0b11dbc58851d63f237956a344f47670e Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 19:26:48 +0100 Subject: [PATCH 23/36] =?UTF-8?q?docs:=20correct=20a=20stale=20ordering=20?= =?UTF-8?q?note=20=E2=80=94=20D9=20landed=20without=20D7?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- docs/bounds-track-plan.md | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 541eeef..43a7df5 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -56,11 +56,11 @@ no longer exists and the suite is 118/118. Questions 6 and 7 dissolved, so the o now **D0 → D5 → D6 → D8**, which deletes `maximum_children_per_parent`. Only D7 is still blocked, on question 8, and it is needed for neither ceiling. D10 is optional. -Ordering is load-bearing. D1–D2 build and prove the mechanism with nothing depending on +Ordering is load-bearing. D1–D2 built and proved the mechanism with nothing depending on it. D4–D5 move each claimant across one at a time, so the suite stays green throughout -and a regression names the driver that caused it. D6 is the flag day. D7 must precede -D9, because zero-resource children are the case that sidesteps containment and so the -reason a shared cap was needed at all. +and a regression names the driver that caused it. D6 is the flag day. (The original +"D7 must precede D9" no longer holds: the per-holder quota bounds zero-resource children +as well as anything else, so D9 landed without it.) **D3 merged into D4** (found while implementing, recorded rather than worked around). The two cannot be separated: the moment the manager claims a device, any driver still From f23f073624360ea737cda81192ca89843123f266 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 19:39:32 +0100 Subject: [PATCH 24/36] kernel: the device rides system_spawn, so a driver never runs without it MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Delegation moves out of onHello and into the spawn itself. The manager holds the hardware and names it in the call that creates the driver; the kernel checks the device is the caller's to give, then hands it over as part of making the child. The reason is the window. A transfer after spawning always leaves an interval in which the child is running and does not yet hold its device. It would close on QEMU every time and open occasionally on a machine with different core counts and timing — the exact failure shape this track exists to delete, and not one worth introducing while removing the others. Fused into the spawn there is no interval: the child does not exist until it holds the device. Ownership is checked BEFORE the child is created, so a refusal leaves nothing running rather than a driver without the hardware it was spawned for. The IOMMU confinement moves with the device, as it does on the transfer path. systemCall6 is added for the sixth argument; r9 was free, and abi gains a no_device sentinel matching the protocol's. No driver had to change to receive a device, which is what makes this better than requiring every driver to hello: ps2-bus keeps its legacy status, and discovery — which has no assignment at all, since it is what produces the device tree — is unaffected. The attacker fixture now tries the spawn as a back door: name someone else's device, and both the spawn and any child must be refused. Verifying that assertion exposed a bug in the fixture itself. The kernel case's pass marker was "device-authority: ok", which matches the FIRST per-assertion line, so its wait loop exited before any failure was printed — the case would have passed with failures in it, and had been able to since D2. The verdict lines now carry a distinct VERDICT prefix, and with the ownership check removed the case genuinely fails. A green test that cannot go red is worse than no test. Suite 118/118. --- docs/bounds-track-plan.md | 2 +- library/kernel/process.zig | 18 ++++++++++++- library/kernel/system-call.zig | 16 ++++++++++++ system/abi.zig | 7 ++++- system/kernel/process.zig | 26 +++++++++++++++++++ system/kernel/tests.zig | 8 ++++-- .../device-manager/device-manager.zig | 25 ++++++------------ .../device-authority-test.zig | 17 ++++++++++-- 8 files changed, 95 insertions(+), 24 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 43a7df5..b8286b9 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -44,7 +44,7 @@ that cannot safely run in user space.** | D3 | The manager claims the seeded devices at boot, before any driver is spawned | **merged into D4** — see below | | D4 | The manager claims + delegates on `hello`; `usb-xhci-bus` is the first driver converted | **done** — caught an IOMMU regression I introduced; see below | | D5 | The other four claimants converted: `pci-bus`, `ps2-bus`, `virtio-gpu`, `acpi` | **partial** — `pci-bus` done; the rest unblocked by D0 below | -| D0 | The grant rides `system_spawn` — atomic, so no driver need change to receive one | not started — **do first** | +| D0 | The grant rides `system_spawn` — atomic, so no driver need change to receive one | **done** — and caught a test-marker bug that made the D2 fixture unfailable | | D10 | Every driver hellos, on its own merits (liveness, one class of driver) | not started — optional, independent | | D6 | `device_claim` refuses a device the caller was not handed; the hole is closed | not started | | D7 | Zero-resource devices stop being kernel objects — inventory moves to the manager | **blocked** — nothing else mints their ids; see question 8 | diff --git a/library/kernel/process.zig b/library/kernel/process.zig index c341189..c3c5129 100644 --- a/library/kernel/process.zig +++ b/library/kernel/process.zig @@ -182,6 +182,22 @@ pub fn spawnWithArguments(name: []const u8, arguments: []const []const u8) ?u32 /// `ipc.replyWait` as a child-exit badge (`ipc.Received.isChildExit`/`childProcessId`), so /// one endpoint can supervise many children. Returns the child's process id, or null. pub fn spawnSupervised(name: []const u8, arguments: []const []const u8, exit_endpoint: ?usize) ?u32 { + return spawnSupervisedWithDevice(name, arguments, exit_endpoint, abi.no_device); +} + +/// Spawn a supervised child **and give it a device you hold**, in one call. +/// +/// The device manager's path: it holds the hardware and the driver it starts must have +/// it. Fusing the handover into the spawn is what removes the window a separate +/// transfer would leave — the child cannot run without its device, because it does not +/// exist until it holds it (docs/os-development/device-authority.md). The kernel checks +/// only that the device is the caller's to give. +pub fn spawnSupervisedWithDevice( + name: []const u8, + arguments: []const []const u8, + exit_endpoint: ?usize, + device: u64, +) ?u32 { var blob: [256]u8 = undefined; var len: usize = 0; for (arguments, 0..) |argument, i| { @@ -194,7 +210,7 @@ pub fn spawnSupervised(name: []const u8, arguments: []const []const u8, exit_end @memcpy(blob[len..][0..argument.len], argument); len += argument.len; } - const r = sc.systemCall5(.system_spawn, @intFromPtr(name.ptr), name.len, if (len == 0) 0 else @intFromPtr(&blob), len, exit_endpoint orelse abi.no_cap); + const r = sc.systemCall6(.system_spawn, @intFromPtr(name.ptr), name.len, if (len == 0) 0 else @intFromPtr(&blob), len, exit_endpoint orelse abi.no_cap, device); if (r > ~@as(usize, 0) - 4095) return null; // a wrapped -errno return @intCast(r); } diff --git a/library/kernel/system-call.zig b/library/kernel/system-call.zig index 5406a90..244a575 100644 --- a/library/kernel/system-call.zig +++ b/library/kernel/system-call.zig @@ -65,3 +65,19 @@ pub inline fn systemCall5(n: SystemCall, a0: usize, a1: usize, a2: usize, a3: us [a4] "{r8}" (a4), : .{ .rcx = true, .r11 = true, .memory = true }); } + +/// The sixth and last argument register. `syscall` clobbers rcx, so r10 stands in for +/// it and r9 is the end of the line — a call needing a seventh would have to pass a +/// struct instead. +pub inline fn systemCall6(n: SystemCall, a0: usize, a1: usize, a2: usize, a3: usize, a4: usize, a5: usize) usize { + return asm volatile ("syscall" + : [ret] "={rax}" (-> usize), + : [n] "{rax}" (@intFromEnum(n)), + [a0] "{rdi}" (a0), + [a1] "{rsi}" (a1), + [a2] "{rdx}" (a2), + [a3] "{r10}" (a3), + [a4] "{r8}" (a4), + [a5] "{r9}" (a5), + : .{ .rcx = true, .r11 = true, .memory = true }); +} diff --git a/system/abi.zig b/system/abi.zig index 138fe2d..8b18754 100644 --- a/system/abi.zig +++ b/system/abi.zig @@ -48,7 +48,7 @@ pub const SystemCall = enum(u64) { irq_bind = 14, // irq_bind(id, resource_index, endpoint): deliver a device IRQ as an IPC notification irq_ack = 15, // irq_ack(id, resource_index): re-arm a bound IRQ after servicing it device_register = 16, // device_register(parent_id, descriptor) -> id/-errno: publish a child of a device you claimed (-ENOSPC table full, -ECHILDREN parent full, -ERANGE resource escapes the parent, -ENODEV/-EPERM bad parent, -E2BIG too many resources) - system_spawn = 17, // system_spawn(name_ptr, name_len, arguments_ptr, arguments_len, exit_endpoint) -> child process id: start a named initial-ramdisk binary as a new ring-3 process + system_spawn = 17, // system_spawn(name_ptr, name_len, arguments_ptr, arguments_len, exit_endpoint, device) -> child process id: start a named initial-ramdisk binary as a new ring-3 process. `device` (or `no_device`) is a device the caller holds and gives to the child, atomically — the child never runs without it dma_alloc = 18, // dma_alloc(len, flags) -> virtual_address (rax), physical_address (rdx): contiguous, pinned, uncacheable DMA memory dma_free = 19, // dma_free(virtual_address, len) -> 0: release a prior dma_alloc msi_bind = 20, // msi_bind(device_id, endpoint) -> address (rax), data (rdx): a per-device MSI vector for a claimed device @@ -127,6 +127,11 @@ pub const ECONFINE: i64 = 16; // the device could not be placed under IOMMU tran /// space, and the number to bump when adding one. pub const errno_maximum: i64 = 16; +/// `system_spawn`'s `device` argument when the child is given no device — every +/// caller but the device manager. Matches `device-manager-protocol.no_device`, which +/// is the same sentinel one layer up. +pub const no_device: u64 = ~@as(u64, 0); + /// `futex_wait` return codes (in rax). pub const futex_woken: u64 = 0; // woken by a futex_wake pub const futex_mismatch: u64 = 1; // *addr != expected on entry; the caller did not block diff --git a/system/kernel/process.zig b/system/kernel/process.zig index 5efecc8..756e783 100644 --- a/system/kernel/process.zig +++ b/system/kernel/process.zig @@ -1017,6 +1017,13 @@ fn systemSpawn(state: *architecture.CpuState) void { const arguments_ptr = architecture.systemCallArg(state, 2); const arguments_len = architecture.systemCallArg(state, 3); const exit_handle = architecture.systemCallArg(state, 4); + // A device the caller holds and gives to the child. Fused into the spawn rather + // than transferred after it, because a separate transfer leaves a window in which + // the child is running and does not yet hold its device — a race that would close + // on one machine and open on another, which is the failure shape this whole track + // exists to remove (docs/bounds-track-plan.md, "the grant rides system_spawn"). + // Here the child cannot observe the gap: it does not exist until it holds it. + const device_to_give = architecture.systemCallArg(state, 5); const t = scheduler.current(); if (len == 0 or len > scheduler.maximum_task_name or ptr >= user_half_end or ptr + len > user_half_end) return fail(state); if (arguments_len > maximum_argument_bytes) return fail(state); @@ -1061,7 +1068,26 @@ fn systemSpawn(state: *architecture.CpuState) void { } } + // Refuse before creating anything if the device is not the caller's to give — a + // spawn that half-succeeds would leave a child running without the hardware it was + // spawned for, which is worse than not spawning it. + if (device_to_give != abi.no_device and devices_broker.ownerOf(device_to_give) != t.id) + return failErr(state, ipc.EPERM); + const child = spawnProcessSupervised(item.blob, 4, argv[0..argc], t.id, exit_endpoint) catch return fail(state); + + if (device_to_give != abi.no_device) { + devices_broker.transfer(device_to_give, t.id, child) catch { + // Cannot happen — ownership was checked above and the lock has not been + // dropped — but a spawned child holding nothing is not something to guess + // about, so say so rather than leave it silent. + log.print("/system/kernel: WARNING spawn gave device {d} to task {d} and the transfer failed\n", .{ device_to_give, child }); + }; + if (devices_broker.pciAddressOf(device_to_give)) |_| { + iommu.reassign(device_to_give, child); + dmaBindOwnerRegionsInto(child, device_to_give); + } + } architecture.setSystemCallResult(state, child); } diff --git a/system/kernel/tests.zig b/system/kernel/tests.zig index 5b16627..5217221 100644 --- a/system/kernel/tests.zig +++ b/system/kernel/tests.zig @@ -4257,8 +4257,12 @@ fn deviceAuthorityTest(boot_information: *const BootInformation) void { process.setInitialRamdisk(image); check("device-authority-test spawned", spawnNamedWithArg(rd, "device-authority-test", "run")); - const pass_marker = "device-authority: ok"; - const fail_marker = "device-authority: FAIL"; + // The VERDICT prefix matters: the fixture prints one "device-authority: ok " + // line per assertion, so a marker of "device-authority: ok" matches the FIRST + // passing assertion and this loop exits before any later failure is printed — the + // case then passes with failures in it, which it did until this was caught. + const pass_marker = "device-authority: VERDICT ok"; + const fail_marker = "device-authority: VERDICT FAILED"; scheduler.setPriority(1); const deadline = architecture.millis() + 20000; var saw_pass = false; diff --git a/system/services/device-manager/device-manager.zig b/system/services/device-manager/device-manager.zig index f2aff2a..4bf7525 100644 --- a/system/services/device-manager/device-manager.zig +++ b/system/services/device-manager/device-manager.zig @@ -306,7 +306,14 @@ fn spawnDriver(driver: *Driver) void { arguments[0] = std.fmt.bufPrint(&id_text, "{d}", .{driver.device_id}) catch return; argument_count = 1; } - const child = process.spawnSupervised(driver.name(), arguments[0..argument_count], manager_endpoint) orelse { + // The device rides the spawn, so the driver holds it before its first instruction. + // A transfer *after* spawning would leave a window in which the child is running + // without its hardware — closed on one machine, open on another + // (docs/bounds-track-plan.md, "the grant rides system_spawn"). + const give = if (isDelegated(driver.name())) driver.device_id else device_manager_protocol.no_device; + if (give != device_manager_protocol.no_device) + std.log.info("delegated device {d} to {s}", .{ give, driver.name() }); + const child = process.spawnSupervisedWithDevice(driver.name(), arguments[0..argument_count], manager_endpoint, give) orelse { std.log.info("failed to spawn {s}", .{driver.name()}); driver.state = .failed; return; @@ -467,22 +474,6 @@ fn onHello(_: void, invocation: Invocation(device_manager_protocol.Hello), _: An }; driver.state = .running; - // **Delegation.** The manager holds this driver's device and hands it over here — - // what replaces first-come-first-served `device_claim` with policy - // (docs/os-development/device-authority.md). `invocation.sender` is the driver's - // task id stamped by the kernel, so the manager cannot be lied to about who is - // asking, and the transfer is a move: the manager stops holding it. - // - // Gated on the delegated set so an unconverted driver still claims for itself and - // its path is untouched; the set and `device_claim` both go at D6. - if (isDelegated(driver.name()) and driver.device_id != device_manager_protocol.no_device) { - device.transfer(driver.device_id, invocation.sender) catch |e| { - std.log.warn("could not delegate device {d} to {s}: {s}", .{ driver.device_id, driver.name(), @errorName(e) }); - return -envelope.EPERM; - }; - std.log.info("delegated device {d} to {s}", .{ driver.device_id, driver.name() }); - } - std.log.info("hello from {s} (device {d})", .{ driver.name(), invocation.target }); // Resilience drill (V6): once, kill the virtio-gpu driver a moment after it hellos, so // the normal restart policy respawns it — the compositor must survive and re-attach. diff --git a/test/system/services/device-authority-test/device-authority-test.zig b/test/system/services/device-authority-test/device-authority-test.zig index ee965b3..fea745b 100644 --- a/test/system/services/device-authority-test/device-authority-test.zig +++ b/test/system/services/device-authority-test/device-authority-test.zig @@ -43,6 +43,7 @@ fn line(comptime format: []const u8, arguments: anytype) void { } var failures: usize = 0; +var process_table: [64]process.ProcessDescriptor = undefined; fn check(name: []const u8, ok: bool) void { if (!ok) failures += 1; @@ -74,10 +75,22 @@ fn run() void { const absent = if (device.transfer(0xFFFF_FFFF, process.taskId())) |_| false else |e| e == error.NoSuchDevice; check("a device that does not exist is refused as absent", absent); + // 4. **The spawn is not a second way in.** A device now rides system_spawn, which + // would be a fine back door if the kernel checked ownership any less carefully + // there than it does in transfer: spawn a child, name someone else's device, and + // the child holds hardware nobody gave it. The refusal must happen before the + // child exists, so nothing is left running either. + if (seen != 0) { + const before = process.processes(&process_table); + const spawned = process.spawnSupervisedWithDevice("/test/system/services/device-authority-test", &.{}, null, table[0].id); + check("spawning with a device the caller does not hold is refused", spawned == null); + check("and no child was left behind by the refusal", process.processes(&process_table) == before); + } + if (failures == 0) { - line("device-authority: ok ({d} devices, none of them mine)\n", .{seen}); + line("device-authority: VERDICT ok ({d} devices, none of them mine)\n", .{seen}); } else { - line("device-authority: FAILED {d} assertion(s)\n", .{failures}); + line("device-authority: VERDICT FAILED {d} assertion(s)\n", .{failures}); } } From 2ebfccc8ed19f0855454e58bfb430d6766281b8e Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 19:48:00 +0100 Subject: [PATCH 25/36] virtio-gpu: the scanout device arrives with the spawn MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Third driver converted. It no longer claims the id from argv[1] — the manager holds the device and names it in the call that creates the process, so it is held before the driver's first instruction. display-reattach is the case that matters here: it kills the driver and watches the compositor re-attach to the fresh scanout. It passes, so the restart path survives the fused grant — the manager re-takes the device when the driver dies and hands it to the replacement. ps2-bus and discovery are NOT converted, and the reason is recorded as open question 9 rather than worked around. Both need a device nobody assigned them. ps2-bus ignores its argv[1] entirely: it finds the controller by walking the table for PNP0303, then claims a second device, the PNP0F13 mouse node, which it also finds itself — so it holds two devices and was assigned at most one, while system_spawn carries one. discovery claims the acpi-tables node it locates itself, because it is what produces the device tree and there is nothing to assign at that point. One thing worth checking before designing an answer: devices.csv maps both PS/2 hardware ids to ps2-bus, so the manager may already be spawning two instances where the driver expects one. If so the fix is smaller than it looks. D6 stays blocked — closing device_claim with these two still depending on it would stop the machine booting. Suite 118/118. --- docs/bounds-track-plan.md | 31 ++++++++++++++++++- system/drivers/virtio-gpu/virtio-gpu.zig | 7 ++--- .../device-manager/device-manager.zig | 1 + 3 files changed, 34 insertions(+), 5 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index b8286b9..3cb7be2 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -43,7 +43,7 @@ that cannot safely run in user space.** | D2 | Adversarial case: a process handed nothing is refused, on a held device and a free one | **done** — `device-authority-test`; the claim half joins it at D6 | | D3 | The manager claims the seeded devices at boot, before any driver is spawned | **merged into D4** — see below | | D4 | The manager claims + delegates on `hello`; `usb-xhci-bus` is the first driver converted | **done** — caught an IOMMU regression I introduced; see below | -| D5 | The other four claimants converted: `pci-bus`, `ps2-bus`, `virtio-gpu`, `acpi` | **partial** — `pci-bus` done; the rest unblocked by D0 below | +| D5 | The other four claimants converted: `pci-bus`, `ps2-bus`, `virtio-gpu`, `acpi` | **partial** — `pci-bus` + `virtio-gpu` done; `ps2-bus` and discovery need question 9 | | D0 | The grant rides `system_spawn` — atomic, so no driver need change to receive one | **done** — and caught a test-marker bug that made the D2 fixture unfailable | | D10 | Every driver hellos, on its own merits (liveness, one class of driver) | not started — optional, independent | | D6 | `device_claim` refuses a device the caller was not handed; the hole is closed | not started | @@ -87,6 +87,35 @@ They land together, with the manager claiming only for drivers in an explicit anything in D4 — recorded rather than dismissed, because D4 moved the `hello` earlier and so did shift boot timing. Watch it across the remaining steps. +### Open question 9 — two drivers need a device nobody assigned them + +D0 made delivery atomic and D5 converted `virtio-gpu` with it. The last two do not fit, +for the same underlying reason: **they need a device the manager never assigned.** + +- **`ps2-bus` ignores its `argv[1]` entirely.** It finds the controller by walking the + table for `PNP0303`, and then claims a *second* device — the `PNP0F13` mouse node — + which it also finds itself. So it holds two devices and was assigned at most one, and + `system_spawn` carries one. +- **discovery** is spawned `addDriver("discovery", no_device, false)` and claims the + `acpi-tables` node it locates itself, because it is what produces the device tree; + there is nothing to assign at that point. + +Three shapes of answer, none written down: + +1. **The manager assigns every device a driver needs.** For `ps2-bus` it would have to + understand that one driver serves both `PNP0303` and `PNP0F13` — `devices.csv` maps + both to it already, so the manager may simply be spawning two instances today where + the driver expects one. Worth checking before designing. +2. **A driver asks for a device over the protocol** — a `request(device)` verb the + manager answers by transferring. General, and it reintroduces a window, though only + for a device the driver asks for after it is running. +3. **The bootstrap is exempt**: discovery keeps claiming, and D6 permits a claim of a + device nobody holds *only* for the discovery role. Narrow and honest, but it is an + exception in the exact place an exception is most expensive. + +**D6 stays blocked** until this is answered — closing `device_claim` with these two +still depending on it would stop the machine booting. + ### Settled 2026-08-08: the grant rides `system_spawn` (D0) Questions 6 and 7 both dissolved on inspection — neither `ps2-bus` nor discovery needs diff --git a/system/drivers/virtio-gpu/virtio-gpu.zig b/system/drivers/virtio-gpu/virtio-gpu.zig index c26c8da..d625992 100644 --- a/system/drivers/virtio-gpu/virtio-gpu.zig +++ b/system/drivers/virtio-gpu/virtio-gpu.zig @@ -200,10 +200,9 @@ fn testPixel(index: u32) u32 { fn initialise(endpoint: ipc.Handle) bool { _ = endpoint; - device.claim(device_id) catch |e| { - std.log.info("unable to claim device {d}: {s}", .{ device_id, @errorName(e) }); - return false; - }; + // The device arrived with the spawn — the manager holds it and names it in the call + // that creates this process, so it is ours before the first instruction here + // (docs/os-development/device-authority.md). Nothing to claim. var descriptors: [64]device.DeviceDescriptor = undefined; const total = device.enumerate(&descriptors); diff --git a/system/services/device-manager/device-manager.zig b/system/services/device-manager/device-manager.zig index 4bf7525..126c2b0 100644 --- a/system/services/device-manager/device-manager.zig +++ b/system/services/device-manager/device-manager.zig @@ -248,6 +248,7 @@ fn alreadySupervised(name: []const u8) bool { const delegated_drivers = [_][]const u8{ "usb-xhci-bus", "pci-bus", + "virtio-gpu", }; /// Matched on the **last path component**, because a driver reaches this table under From 3ae541214f94737fc5caa339e1b04012a2721710 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 19:57:36 +0100 Subject: [PATCH 26/36] iommu: assert directly that a confinement moves with its device MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit reassign was added at D4 to fix a regression and has been proven only indirectly since — three IOMMU+USB cases going green. That covered the visible symptom (a driver's DMA rings unbound) and neither of the latent ones: the confinement still naming the giver, so the giver's death would tear down a domain a live driver was using, and the receiver's death would leave one behind. Those are now asserted. confinementOwner exposes the record's owner so the suite can see it. The sequence is the delegation in miniature: unconfined, confine as this task, reassign to another, confirm the new holder owns it and the old one does not, then kill the new holder and confirm the domain goes with it. Two attempts at this test could not have failed. The first found no PCI function to confine — pciAddressOf needs a pci_device entry and this case runs no pci-bus — so every assertion skipped silently while the case stayed green. It now synthesizes a function the way pci-bus does, a 4 KiB config window inside the bridge's ECAM, and asserts that precondition explicitly so a skip is a failure. Verified to discriminate: making reassign a no-op flips three assertions, including the domain surviving its holder's death. Suite 118/118. --- docs/bounds-track-plan.md | 11 ++++++---- system/kernel/iommu.zig | 11 ++++++++++ system/kernel/tests.zig | 46 +++++++++++++++++++++++++++++++++++++++ 3 files changed, 64 insertions(+), 4 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 3cb7be2..92a56f0 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -51,10 +51,13 @@ that cannot safely run in user space.** | D8 | **`maximum_children_per_parent` deleted** — the authorisation it stood in for exists | **blocked on D6**, and now ordered after D9 | | D9 | The device table becomes dynamic; **`maximum_devices` deleted**; per-holder quota declared | **done** — one of the two invented numbers is gone | -**Run 2 resumes at D0.** D1, D2, D4, D5 (`pci-bus` only) and D9 landed; `maximum_devices` -no longer exists and the suite is 118/118. Questions 6 and 7 dissolved, so the order is -now **D0 → D5 → D6 → D8**, which deletes `maximum_children_per_parent`. Only D7 is still -blocked, on question 8, and it is needed for neither ceiling. D10 is optional. +**Run 2 stops, blocked.** Landed: D0, D1, D2, D4, D9, and D5 for three of five +claimants (`usb-xhci-bus`, `pci-bus`, `virtio-gpu`). `maximum_devices` no longer exists, +delegation is atomic with the spawn, and the suite is 118/118. + +Blocked: **D6 and D8 on question 9** (`ps2-bus` needs two devices and ignores its +assignment; discovery needs a node nobody assigns), and **D7 on question 8**. So +`maximum_children_per_parent` — the second invented number — is one answer away. Ordering is load-bearing. D1–D2 built and proved the mechanism with nothing depending on it. D4–D5 move each claimant across one at a time, so the suite stays green throughout diff --git a/system/kernel/iommu.zig b/system/kernel/iommu.zig index 5afb203..029ac19 100644 --- a/system/kernel/iommu.zig +++ b/system/kernel/iommu.zig @@ -234,6 +234,17 @@ pub fn reassign(device_id: u64, owner: u32) void { record.owner = owner; } +/// The task a device's confinement is recorded against, or null if it has none. The +/// confinement's owner decides whose death tears the domain down, so a delegation that +/// moved the device but not this record would leave a live driver's domain destroyed by +/// its manager's exit — which is why `reassign` exists and why the suite asserts it. +pub fn confinementOwner(device_id: u64) ?u32 { + if (!active) return null; + if (device_id >= confined.len) return null; + const record = confined[@intCast(device_id)]; + return if (record.active) record.owner else null; +} + pub fn releaseAllOwnedBy(owner: u32) void { if (!active) return; for (confined) |*c| { diff --git a/system/kernel/tests.zig b/system/kernel/tests.zig index 5217221..9059ae2 100644 --- a/system/kernel/tests.zig +++ b/system/kernel/tests.zig @@ -1447,6 +1447,52 @@ fn iommuTest() void { } else { check("scratch domain allocated", false); } + // Delegation moves a device between tasks, and the IOMMU confinement must move + // with it. When it did not, the driver's DMA rings were never bound into the + // device's domain and every transfer faulted — but two worse consequences were + // latent and invisible to those tests: the domain still named the *giver*, so the + // giver's death would tear down a domain a live driver was using, and the + // receiver's death would leave one behind. Asserted directly here rather than + // inferred from the USB cases going green. + // This case runs no pci-bus, so there are no PCI functions to confine — and a + // silently skipped assertion is worse than none. Register one the way pci-bus does: + // a child of the host bridge whose resource 0 is a 4 KiB config window inside the + // bridge's ECAM, which is what `pciAddressOf` derives a requester id from. + var iommu_table: [64]device_abi.DeviceDescriptor = undefined; + const iommu_total = devices_broker.enumerate(&iommu_table); + const bridge: ?device_abi.DeviceDescriptor = for (iommu_table[0..@min(iommu_total, iommu_table.len)]) |d| { + if (d.class == @intFromEnum(device_abi.DeviceClass.pci_host_bridge) and d.resource_count >= 2) break d; + } else null; + check("the kernel seeded a PCI host bridge to parent a function under", bridge != null); + + const subject: ?u64 = if (bridge) |b| blk: { + const me_bridge = scheduler.currentId(); + if (!claimOk(b.id, me_bridge)) break :blk null; + var function = std.mem.zeroes(device_abi.DeviceDescriptor); + function.class = @intFromEnum(device_abi.DeviceClass.pci_device); + function.pci_class = device_abi.no_pci_class; + function.resource_count = 1; + function.resources[0] = .{ + .kind = @intFromEnum(device_abi.ResourceKind.memory), + .start = b.resources[0].start, // the first config slot in the ECAM window + .len = 4096, + }; + break :blk devices_broker.register(b.id, me_bridge, &function) catch null; + } else null; + check("a PCI function exists to confine", subject != null); + + if (subject) |device_id| { + check("a device starts unconfined", iommu.confinementOwner(device_id) == null); + const me = scheduler.currentId(); + check("confining records the owner", iommu.confineDevice(device_id, devices_broker.pciAddressOf(device_id).?, me)); + check("the confinement names the confiner", iommu.confinementOwner(device_id) == me); + iommu.reassign(device_id, me + 1000); + check("reassign moves the confinement to the new holder", iommu.confinementOwner(device_id) == me + 1000); + check("and the previous holder no longer owns it", iommu.confinementOwner(device_id) != me); + iommu.releaseAllOwnedBy(me + 1000); + check("the new holder's death tears the domain down", iommu.confinementOwner(device_id) == null); + } + log("DANOS-IOMMU: enabled base=0x{x} domains active\n", .{pinfo.iommu_base}); result(); } From 7d8aa51234b07992a6b7024208d597731b8e1ca0 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 21:12:20 +0100 Subject: [PATCH 27/36] kernel: maximum_children_per_parent is gone MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The second invented ceiling. It was written to stop a driver looping device_register and exhausting a shared table — but there is no shared table to exhaust any more, and each registrar already has its own allowance, so a runaway costs only itself. It never bounded a determined caller in the first place: 16 children per parent, and nothing stopped it claiming more parents. What it reliably did was refuse a real PCI bus with more than 16 functions, which is how an AMD Ryzen booted with a working display, no USB and no storage. The constant, its check, and the now-unused childCount all go. TooManyChildren survives with one meaning instead of two: the caller is at its per-registrar allowance. This was unblocked from the moment D9 landed. The plan said so — "once the quota exists the per-parent cap is redundant whether or not D6 has landed" — in the same edit that left the step tagged "blocked on D6". Three iterations were then spent re-reading that tag instead of the sentence beside it. The containment test now asserts 64 children under one parent, four times the old ceiling; restoring the cap fails it. Suite 118/118. --- docs/bounds-track-plan.md | 65 +++++++++++++++++--------------- system/kernel/devices-broker.zig | 27 +------------ system/kernel/tests.zig | 37 +++++++++--------- 3 files changed, 54 insertions(+), 75 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 92a56f0..0af5806 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -48,16 +48,17 @@ that cannot safely run in user space.** | D10 | Every driver hellos, on its own merits (liveness, one class of driver) | not started — optional, independent | | D6 | `device_claim` refuses a device the caller was not handed; the hole is closed | not started | | D7 | Zero-resource devices stop being kernel objects — inventory moves to the manager | **blocked** — nothing else mints their ids; see question 8 | -| D8 | **`maximum_children_per_parent` deleted** — the authorisation it stood in for exists | **blocked on D6**, and now ordered after D9 | +| D8 | **`maximum_children_per_parent` deleted** | **done** — it was unblocked from the moment D9 landed; I kept reading my own stale label | | D9 | The device table becomes dynamic; **`maximum_devices` deleted**; per-holder quota declared | **done** — one of the two invented numbers is gone | -**Run 2 stops, blocked.** Landed: D0, D1, D2, D4, D9, and D5 for three of five -claimants (`usb-xhci-bus`, `pci-bus`, `virtio-gpu`). `maximum_devices` no longer exists, -delegation is atomic with the spawn, and the suite is 118/118. +**Run 2 stops, blocked on one question.** Landed: D0, D1, D2, D4, D8, D9, and D5 for +three of five claimants (`usb-xhci-bus`, `pci-bus`, `virtio-gpu`). Suite 118/118. -Blocked: **D6 and D8 on question 9** (`ps2-bus` needs two devices and ignores its -assignment; discovery needs a node nobody assigns), and **D7 on question 8**. So -`maximum_children_per_parent` — the second invented number — is one answer away. +**Both invented ceilings are gone.** `maximum_devices` and +`maximum_children_per_parent` no longer exist, and delegation is atomic with the spawn. + +What remains is not a number: **D6 closes the claiming hole** and is blocked on +question 9, and D7 moves the inventory and is blocked on question 8. Ordering is load-bearing. D1–D2 built and proved the mechanism with nothing depending on it. D4–D5 move each claimant across one at a time, so the suite stays green throughout @@ -90,34 +91,38 @@ They land together, with the manager claiming only for drivers in an explicit anything in D4 — recorded rather than dismissed, because D4 moved the `hello` earlier and so did shift boot timing. Watch it across the remaining steps. -### Open question 9 — two drivers need a device nobody assigned them +### Open question 9 — danos has two device-acquisition patterns; D6 fits only one -D0 made delivery atomic and D5 converted `virtio-gpu` with it. The last two do not fit, -for the same underlying reason: **they need a device the manager never assigned.** +Earlier versions of this question listed symptoms — "`ps2-bus` ignores its `argv[1]`", +"discovery has no assignment" — which hid that they are the same fact. -- **`ps2-bus` ignores its `argv[1]` entirely.** It finds the controller by walking the - table for `PNP0303`, and then claims a *second* device — the `PNP0F13` mouse node — - which it also finds itself. So it holds two devices and was assigned at most one, and - `system_spawn` carries one. -- **discovery** is spawned `addDriver("discovery", no_device, false)` and claims the - `acpi-tables` node it locates itself, because it is what produces the device tree; - there is nothing to assign at that point. +| Pattern | Who | How it gets its device | +|---|---|---| +| Per-device driver | `pci-bus`, `usb-xhci-bus`, `virtio-gpu`, `usb-hid`, `usb-storage` | assigned one id, one instance per device — **all converted** | +| Singleton that finds its own | `ps2-bus`, discovery | spawned once with `no_device`, walks the table itself | -Three shapes of answer, none written down: +The second is deliberate, not an oversight. `onChildAdded` says so: -1. **The manager assigns every device a driver needs.** For `ps2-bus` it would have to - understand that one driver serves both `PNP0303` and `PNP0F13` — `devices.csv` maps - both to it already, so the manager may simply be spawning two instances today where - the driver expects one. Worth checking before designing. -2. **A driver asks for a device over the protocol** — a `request(device)` verb the - manager answers by transferring. General, and it reintroduces a window, though only - for a device the driver asks for after it is running. -3. **The bootstrap is exempt**: discovery keeps claiming, and D6 permits a claim of a - device nobody holds *only* for the discovery role. Narrow and honest, but it is an - exception in the exact place an exception is most expensive. +> An hid-matched driver (ps2-bus) is a singleton that finds its own devices once +> spawned — spawn it once, no device assignment. -**D6 stays blocked** until this is answered — closing `device_claim` with these two -still depending on it would stop the machine booting. +`device_claim` is the mechanism that makes that pattern work. **D6 removes it and puts +nothing in its place**, which is the whole of the blockage. + +The manager is not ignorant: it matched *both* `PNP0303` and `PNP0F13` to `ps2-bus` and +chose not to assign either, so it already knows which devices a singleton wants. An +answer is therefore in reach — hand a singleton each matching device as it matches. The +cost: only the first can ride the spawn, so later ones arrive while the driver is +running, and `ps2-bus` enumerates once at startup and would have to tolerate that. + +**The decision: do singletons stop being singletons — one instance per device, like +everything else — or does the system keep a second acquisition path for them?** The +first is uniform and costs a rewrite of `ps2-bus`'s startup. The second keeps a +mechanism whose only remaining users are two drivers, and every exemption in an +authority model is somewhere the model does not hold. + +Nothing else blocks D6. Both invented ceilings are already gone, so this is about +closing the claiming hole, not about a number. ### Settled 2026-08-08: the grant rides `system_spawn` (D0) diff --git a/system/kernel/devices-broker.zig b/system/kernel/devices-broker.zig index 286e370..2d3acff 100644 --- a/system/kernel/devices-broker.zig +++ b/system/kernel/devices-broker.zig @@ -53,22 +53,6 @@ const heap = @import("heap.zig"); /// found against registered) const maximum_devices_per_registrar = 4096; -/// Cap on children a single parent may have. A zero-resource child (legal — a USB -/// device is addressed through its controller, not by MMIO) sidesteps the containment -/// check, so without a bound a process that claimed one device could loop -/// `device_register` and exhaust the whole table, permanently denying it to every other -/// driver. This bounds the blast radius of one claim; a real quota (and a -/// `device_release` to reclaim on exit) is future work — see docs/driver-model.md. -/// -/// bound: children one claimed parent may register — in practice every PCI function on -/// the machine, since pci-bus registers them all under the one host bridge -/// decided-by: hardware -/// protects: the shared device table, against a driver looping device_register — but -/// it is a proxy for an authorisation the kernel does not perform, since any -/// process may claim any unclaimed device (docs/bounds-track-plan.md phase 2) -/// at-limit: refuse — ECHILDREN, distinct from a full table -/// observed-by: pci-bus logs the reason per refused function, and warns at end of scan -const maximum_children_per_parent = 16; /// The table, grown on demand from the kernel heap. **There is no ceiling**: how many /// devices a machine has is the machine's business, and no specification bounds it, so @@ -368,7 +352,7 @@ pub const RegisterError = error{ NoSuchParent, // no device with that id NotYourParent, // that device exists but this task has not claimed it TooManyResources, // the descriptor declares more resources than one device may hold - TooManyChildren, // this parent is at maximum_children_per_parent + TooManyChildren, // the caller is at its per-registrar allowance NotContained, // a child resource escapes its parent's window }; @@ -459,14 +443,6 @@ fn existingChild(parent_id: u64, descriptor: *const device_abi.DeviceDescriptor) return null; } -/// Number of devices currently recorded with `parent_id` as their parent. -fn childCount(parent_id: u64) usize { - var n: usize = 0; - for (devices[0..count]) |d| { - if (d.parent == parent_id) n += 1; - } - return n; -} /// Publish `descriptor` as a child of `parent_id`, on behalf of `owner`. Returns the new /// device id. The child is left **unclaimed**, so another process (a class driver) @@ -497,7 +473,6 @@ pub fn register(parent_id: u64, owner: u32, descriptor: *const device_abi.Device // below). if (existingChild(parent_id, descriptor)) |existing_id| return existing_id; - if (childCount(parent_id) >= maximum_children_per_parent) return error.TooManyChildren; // The allowance is charged to whoever is registering, so a driver in a loop // exhausts its own and every other driver carries on. There is no machine-wide // ceiling any more: the table grows. diff --git a/system/kernel/tests.zig b/system/kernel/tests.zig index 9059ae2..9f4ee11 100644 --- a/system/kernel/tests.zig +++ b/system/kernel/tests.zig @@ -4037,34 +4037,33 @@ fn containmentTest() void { check("re-registering an identical child returns the same id", again != 0 and again == good); check("re-registering grew nothing", devices_broker.enumerate(&buffer) == before + 1); - // Fill the parent to its child cap with distinct children (same window, different - // identity — the match is on identity, so each is a new device). + // **There is no per-parent cap.** `maximum_children_per_parent = 16` is gone: it + // was written to stop a driver looping `device_register` and exhausting a shared + // table, and there is no shared table to exhaust — the table grows, and each + // registrar has its own allowance, so a runaway costs only itself. The cap never + // bounded a determined caller anyway (16 per parent, but nothing stopped it + // claiming more parents); what it reliably did was refuse a real PCI bus with more + // than 16 functions, which is how an AMD Ryzen booted with no USB and no storage. + // + // 64 children under one parent — four times the old ceiling. var filled: u32 = 0; - var capped = false; + var registered_children: u32 = 0; while (filled < 64) : (filled += 1) { var name: [4]u8 = .{ 'k', 0, 0, 0 }; name[1] = '0' + @as(u8, @intCast(filled / 10)); name[2] = '0' + @as(u8, @intCast(filled % 10)); var extra = childDescriptor(name[0..3], parent_window.start, 0x20); - _ = devices_broker.register(parent_id, me, &extra) catch |err| { - capped = err == error.TooManyChildren; - break; - }; + if (devices_broker.register(parent_id, me, &extra)) |_| { + registered_children += 1; + } else |_| break; } - check("the parent reaches its child cap (TooManyChildren)", capped); + check("one parent takes far more children than the old cap allowed", registered_children == 64); - // The regression this ordering exists for: **a re-registration consumes no slot, - // so a full parent must not refuse one.** A crashed bus driver is restarted by its - // supervisor and re-registers everything it rediscovers; when the cap was checked - // before the identity match, the restart was refused its own devices and the - // machine degraded a little more on every crash. + // Idempotency still holds, and still matters: a crashed bus driver is restarted and + // re-registers everything it rediscovers, which must return the ids it had before + // rather than duplicate them. const readmitted = devices_broker.register(parent_id, me, &fits) catch 0; - check("a full parent still re-admits an identical child", readmitted != 0 and readmitted == good); - - // ...and the cap is genuinely still in force for anything new. - var novel = childDescriptor("knew", parent_window.start, 0x20); - const still_capped = if (devices_broker.register(parent_id, me, &novel)) |_| false else |err| err == error.TooManyChildren; - check("a full parent still refuses a new child", still_capped); + check("re-registering an identical child still returns its id", readmitted != 0 and readmitted == good); // The table itself has no ceiling: it grows. The old `maximum_devices = 64` was a // guess about someone else's computer, and one driver's enumeration starved every From 6b3a3816269bc8b6d80860f9ca8a77b8791c57aa Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 21:33:35 +0100 Subject: [PATCH 28/36] ps2: one instance, every node the machine has MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The 8042 is a single controller described by two ACPI nodes — PNP0303 carries the 0x60/0x64 ports, PNP0F13 is the mouse — so it cannot be split across two processes without them fighting over the same registers. That is why ps2-bus is a singleton, and why it used to find and claim both nodes itself. The manager now gives it every matching node instead. The keyboard node rides the spawn, because it holds the ports and is needed immediately; the mouse node is transferred to the already-running instance. Late arrival is safe here and the ordering is natural rather than lucky: the mouse is not touched until after the controller handshakes and identify. Measured, the handover lands at 0.336 and the driver reaches the mouse at 0.456. Because the count is however many matched, a machine with no PS/2 ports or only one works without a special case — which matters, since the bus is mostly emulated now and machines vary. ps2-bus claims nothing. irq_bind on the mouse node is the proof it holds it: that call is ownership-gated, so a failure means the handover did not land rather than a hardware fault, and the log says so. acpi-ps2 asserts both delegations with the spawned device backreferenced, so the node that rides the spawn must be the one the driver was spawned for. Disabling the second delegation fails it. Two things worth recording. The first discrimination patch was not valid Zig, so nothing ran and a stale binary reported a pass — checked the build before believing it. And with the second delegation disabled, acpi-ps2 fails while input still passes: the mouse works without its IRQ binding, so exactly one case covers that path. Suite 118/118. --- system/drivers/ps2-bus/ps2-bus.zig | 34 +++++++-------- .../device-manager/device-manager.zig | 41 +++++++++++++++++-- test/qemu_test.py | 10 ++++- 3 files changed, 65 insertions(+), 20 deletions(-) diff --git a/system/drivers/ps2-bus/ps2-bus.zig b/system/drivers/ps2-bus/ps2-bus.zig index fcc0c72..9911562 100644 --- a/system/drivers/ps2-bus/ps2-bus.zig +++ b/system/drivers/ps2-bus/ps2-bus.zig @@ -114,10 +114,9 @@ pub fn main() void { _ = logging.write("/system/drivers/ps2-bus: found PS/2 controller\n"); _ = logging.write("/system/drivers/ps2-bus: initializing controller\n"); - device.claim(controller_device_descriptor.id) catch |e| { - std.log.warn("unable to claim controller: {s}", .{@errorName(e)}); - return; - }; + // The controller node arrived with the spawn — the manager holds the hardware + // and names it in the call that creates this process + // (docs/os-development/device-authority.md). Nothing to claim. const controller = ps2.Controller.init(controller_device_descriptor) orelse { _ = logging.write("/system/drivers/ps2-bus: controller is missing its IO ports\n"); @@ -255,18 +254,21 @@ pub fn main() void { if (port_device_types[@intFromEnum(ps2.Port.two)] != null) { if (ps2.findMouseDescriptor(buffer)) |descriptor| { if (findInterruptResourceIndex(descriptor)) |auxiliary_index| { - if (device.claim(descriptor.id)) |_| { - if (device.irqBind(descriptor.id, auxiliary_index, endpoint)) { - maybe_auxiliary_interrupt = .{ - .device_id = descriptor.id, - .interrupt_index = auxiliary_index, - .gsi = descriptor.resources[auxiliary_index].start, - }; - } else { - _ = logging.write("/system/drivers/ps2-bus: auxiliary irq_bind failed\n"); - } - } else |e| { - std.log.warn("auxiliary claim failed: {s}", .{@errorName(e)}); + // The mouse node is a *second* device for this one instance — the + // 8042 is one controller with two ports, so it cannot be split across + // two processes. It is transferred to us after the spawn, which is + // safe because we only reach it here, long after the controller + // handshakes and identify. irq_bind is the proof we hold it: it is + // gated on ownership, so a failure here means the handover has not + // landed rather than a hardware problem. + if (device.irqBind(descriptor.id, auxiliary_index, endpoint)) { + maybe_auxiliary_interrupt = .{ + .device_id = descriptor.id, + .interrupt_index = auxiliary_index, + .gsi = descriptor.resources[auxiliary_index].start, + }; + } else { + _ = logging.write("/system/drivers/ps2-bus: auxiliary irq_bind failed (mouse node not delegated?)\n"); } } } diff --git a/system/services/device-manager/device-manager.zig b/system/services/device-manager/device-manager.zig index 126c2b0..0692992 100644 --- a/system/services/device-manager/device-manager.zig +++ b/system/services/device-manager/device-manager.zig @@ -220,6 +220,13 @@ fn childCountOf(reporter: u32) u32 { /// The driver entry a live process id belongs to. Zero is not a process id here: /// it is what `onDriverExit` writes back to retire an id it has already acted on, /// so a second notification for the same death matches nothing. +fn driverByName(name: []const u8) ?*Driver { + for (&drivers) |*driver| { + if (driver.used and std.mem.eql(u8, driver.name(), name)) return driver; + } + return null; +} + fn driverByProcess(process_id: u32) ?*Driver { if (process_id == 0) return null; for (&drivers) |*driver| { @@ -249,6 +256,7 @@ const delegated_drivers = [_][]const u8{ "usb-xhci-bus", "pci-bus", "virtio-gpu", + "ps2-bus", }; /// Matched on the **last path component**, because a driver reaches this table under @@ -512,9 +520,36 @@ fn onChildAdded(_: void, invocation: Invocation(device_manager_protocol.ChildAdd if (match.ambiguous) std.log.info("/system/configuration/devices.csv: multiple equally-specific rules match the device {s} reported; binding {s}", .{ driver.name(), match.driver }); if (id.bus == .acpi) { - // An hid-matched driver (ps2-bus) is a singleton that finds its - // own devices once spawned — spawn it once, no device assignment. - if (!alreadySupervised(match.driver)) addDriver(match.driver, device_manager_protocol.no_device, false); + // An hid-matched driver is a singleton over one piece of hardware + // described by several nodes: the 8042 is a single controller whose + // I/O ports live under the keyboard node (PNP0303) while the mouse + // is a second node (PNP0F13). Two processes would fight over the + // same 0x60/0x64 registers, so there is exactly one instance — and + // it needs *every* matching device, however many the machine has + // (some have none, some one port, some two). + // + // The first rides the spawn; the rest are transferred to the + // running instance. Late arrival is fine here because the order is + // natural: the instance needs the controller node immediately and + // reaches the mouse only after the 8042 handshakes and identify. + if (!alreadySupervised(match.driver)) { + addDriver(match.driver, device_id, false); + } else if (driverByName(match.driver)) |running| { + if (running.process_id != 0) { + device.claim(device_id) catch |e| switch (e) { + error.AlreadyClaimed => {}, + else => { + std.log.warn("cannot hold device {d} for {s}: {s}", .{ device_id, match.driver, @errorName(e) }); + return status; + }, + }; + device.transfer(device_id, running.process_id) catch |e| { + std.log.warn("could not give device {d} to {s}: {s}", .{ device_id, match.driver, @errorName(e) }); + return status; + }; + std.log.info("delegated device {d} to {s} (already running)", .{ device_id, match.driver }); + } + } } else { // A per-device driver: one instance, the registered id as argv[1]. if (!driverForDevice(device_id)) addDriver(match.driver, device_id, true); diff --git a/test/qemu_test.py b/test/qemu_test.py index 708f373..1fcc6ad 100644 --- a/test/qemu_test.py +++ b/test/qemu_test.py @@ -780,8 +780,16 @@ CASES = [ {"name": "acpi-ps2", "smp": 4, "timeout": 150, + # The 8042 is ONE controller described by two ACPI nodes, so the single ps2-bus + # instance must receive BOTH. The keyboard node rides the spawn — it carries the + # 0x60/0x64 ports and is needed immediately — and the mouse node is transferred to + # the already-running instance, which is safe because it is not touched until after + # the controller handshakes and identify. Two processes cannot split this: they + # would fight over the same registers. "expect": r"discovery: device \d+\s+bus=acpi hid=PNP0303[\s\S]*" - r"device-manager: spawned \S*ps2-bus[\s\S]*" + r"device-manager: delegated device (\d+) to \S*ps2-bus[\s\S]*" + r"device-manager: spawned \S*ps2-bus for device \1[\s\S]*" + r"device-manager: delegated device \d+ to \S*ps2-bus \(already running\)[\s\S]*" r"ps2-bus: keyboard driver attached", "fail": r"DANOS-TEST-RESULT: FAIL"}, # M21.1: the SCI + power button. Boot the manager (which spawns the acpi From b7d97ebb5d330ecc1df0b943f915789293e018f9 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 21:54:45 +0100 Subject: [PATCH 29/36] acpi: discovery is handed its node like every other driver MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The last claimant. The kernel seeds the acpi-tables node, so it sits in the same boot snapshot the manager already scans to find the PCI host bridge — there was never a bootstrap problem, only a lookup nobody had written. The manager claims it and names it in the spawn; the service stops claiming. Every driver in the system now receives its hardware rather than taking it. Two failures on the way, both mine. addDriver puts the device id in argv[1], and the acpi service read argv[1] as a self-verify device-count floor — so handed device 7 it decided it was in test mode, printed "acpi-parse: ok", and never reported a device. The test argument is now floor:N, which a bare id cannot be mistaken for. And acpi-parse spawns the service directly rather than through the manager, so nothing handed it the node. That test now claims and transfers it exactly as the manager does, which is the right shape: the test plays the manager's role instead of the service reaching for hardware. device_claim now has two callers left: the manager, which is the acquirer and should have it, and the display service's GOP path. That is recorded as question 10 — the framebuffer is not a device, so the answer is likely that it leaves the device table rather than being exempted from its rules. Suite 118/118. --- docs/bounds-track-plan.md | 53 +++++++++---------- system/kernel/tests.zig | 12 ++++- system/services/acpi/acpi.zig | 26 +++++---- .../device-manager/device-manager.zig | 15 +++++- 4 files changed, 67 insertions(+), 39 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 0af5806..dd3f653 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -43,10 +43,10 @@ that cannot safely run in user space.** | D2 | Adversarial case: a process handed nothing is refused, on a held device and a free one | **done** — `device-authority-test`; the claim half joins it at D6 | | D3 | The manager claims the seeded devices at boot, before any driver is spawned | **merged into D4** — see below | | D4 | The manager claims + delegates on `hello`; `usb-xhci-bus` is the first driver converted | **done** — caught an IOMMU regression I introduced; see below | -| D5 | The other four claimants converted: `pci-bus`, `ps2-bus`, `virtio-gpu`, `acpi` | **partial** — `pci-bus` + `virtio-gpu` done; `ps2-bus` and discovery need question 9 | +| D5 | The other four claimants converted: `pci-bus`, `ps2-bus`, `virtio-gpu`, `acpi` | **done** — every driver now receives its hardware; question 9 answered | | D0 | The grant rides `system_spawn` — atomic, so no driver need change to receive one | **done** — and caught a test-marker bug that made the D2 fixture unfailable | | D10 | Every driver hellos, on its own merits (liveness, one class of driver) | not started — optional, independent | -| D6 | `device_claim` refuses a device the caller was not handed; the hole is closed | not started | +| D6 | `device_claim` refuses a device the caller was not handed; the hole is closed | **blocked on question 10** — only the display service's GOP path still claims | | D7 | Zero-resource devices stop being kernel objects — inventory moves to the manager | **blocked** — nothing else mints their ids; see question 8 | | D8 | **`maximum_children_per_parent` deleted** | **done** — it was unblocked from the moment D9 landed; I kept reading my own stale label | | D9 | The device table becomes dynamic; **`maximum_devices` deleted**; per-holder quota declared | **done** — one of the two invented numbers is gone | @@ -91,38 +91,35 @@ They land together, with the manager claiming only for drivers in an explicit anything in D4 — recorded rather than dismissed, because D4 moved the `hello` earlier and so did shift boot timing. Watch it across the remaining steps. -### Open question 9 — danos has two device-acquisition patterns; D6 fits only one +### Question 9 — answered: a singleton gets every device that matched it -Earlier versions of this question listed symptoms — "`ps2-bus` ignores its `argv[1]`", -"discovery has no assignment" — which hid that they are the same fact. +The 8042 is one controller described by two ACPI nodes, so it cannot be split across +processes. The manager now hands the single `ps2-bus` instance **every** matching node: +the first rides the spawn, the rest are transferred to the running instance. Late +arrival is safe because the ordering is natural rather than lucky — the controller node +carries the ports and is needed at once, the mouse node not until after identify +(measured: handed over at 0.336, first touched at 0.456). The count is whatever matched, +so a machine with no PS/2 ports or one port needs no special case. -| Pattern | Who | How it gets its device | -|---|---|---| -| Per-device driver | `pci-bus`, `usb-xhci-bus`, `virtio-gpu`, `usb-hid`, `usb-storage` | assigned one id, one instance per device — **all converted** | -| Singleton that finds its own | `ps2-bus`, discovery | spawned once with `no_device`, walks the table itself | +Discovery turned out not to be special either. The kernel seeds the `acpi-tables` node, +so it is in the same boot snapshot the manager already scans for the PCI host bridge — +it is handed over at spawn like everything else. -The second is deliberate, not an oversight. `onChildAdded` says so: +### Open question 10 — the framebuffer is the last thing anyone claims -> An hid-matched driver (ps2-bus) is a singleton that finds its own devices once -> spawned — spawn it once, no device assignment. +`device_claim` now has exactly two callers: the device manager, which is the acquirer +and should have it, and `system/services/display/backend.zig`, whose GOP path claims the +kernel-seeded display node. -`device_claim` is the mechanism that makes that pattern work. **D6 removes it and puts -nothing in its place**, which is the whole of the blockage. +Closing `device_claim` breaks the compositor's boot floor. Exempting it puts a hole in +the middle of the authority model, in the one place an exemption is most expensive. -The manager is not ignorant: it matched *both* `PNP0303` and `PNP0F13` to `ps2-bus` and -chose not to assign either, so it already knows which devices a singleton wants. An -answer is therefore in reach — hand a singleton each matching device as it matches. The -cost: only the first can ride the spawn, so later ones arrive while the driver is -running, and `ps2-bus` enumerates once at startup and would have to tolerate that. - -**The decision: do singletons stop being singletons — one instance per device, like -everything else — or does the system keep a second acquisition path for them?** The -first is uniform and costs a rewrite of `ps2-bus`'s startup. The second keeps a -mechanism whose only remaining users are two drivers, and every exemption in an -authority model is somewhere the model does not hold. - -Nothing else blocks D6. Both invented ceilings are already gone, so this is about -closing the claiming hole, not about a number. +The likely answer is neither. **The framebuffer is not a device** — it is where pixels +go, handed over by the loader, and the kernel wraps it in a `DeviceClass.display` +descriptor only so `mmio_map` can hand it over write-combining. If that is right, it +should leave the device table rather than be exempted from its rules, and the compositor +should receive the pixels some other way. That is a change to the display path, which is +out of scope for this run. ### Settled 2026-08-08: the grant rides `system_spawn` (D0) diff --git a/system/kernel/tests.zig b/system/kernel/tests.zig index 9f4ee11..e8ccf80 100644 --- a/system/kernel/tests.zig +++ b/system/kernel/tests.zig @@ -3098,7 +3098,17 @@ fn acpiParseTest(boot_information: *const BootInformation) void { while (i < rd.count) : (i += 1) { const item = rd.entry(i) orelse continue; if (!eql(initial_ramdisk.basename(item.name), "discovery")) continue; - _ = process.spawnProcessSupervised(item.blob, 4, &.{ item.name, "1" }, scheduler.currentId(), null) catch 0; + // Hand it the acpi-tables node the way the manager would: claim it here, then + // move it to the child. Discovery no longer claims for itself. + var scratch: [64]device_abi.DeviceDescriptor = undefined; + const seen_devices = devices_broker.enumerate(&scratch); + const tables: ?u64 = for (scratch[0..@min(seen_devices, scratch.len)]) |d| { + if (d.class == @intFromEnum(device_abi.DeviceClass.acpi_tables)) break d.id; + } else null; + const me_parse = scheduler.currentId(); + if (tables) |node| _ = claimOk(node, me_parse); + const child = process.spawnProcessSupervised(item.blob, 4, &.{ item.name, "floor:1" }, me_parse, null) catch 0; + if (tables) |node| devices_broker.transfer(node, me_parse, child) catch {}; spawned = true; break; } diff --git a/system/services/acpi/acpi.zig b/system/services/acpi/acpi.zig index 9c0545b..de5d01f 100644 --- a/system/services/acpi/acpi.zig +++ b/system/services/acpi/acpi.zig @@ -104,11 +104,19 @@ fn findTablesNode(buffer: []device.DeviceDescriptor) ?device.DeviceDescriptor { } pub fn main(init: process.Init) void { - // When the acpi-parse scenario spawns this directly, argv[1] is a device-count - // *floor* to self-verify against. The kernel no longer parses AML, so there is - // no exact count to match — proving the ring-3 parse found at least a floor of - // devices is the check. Deterministic, no log-scraping. - const floor: ?usize = if (init.arguments.get(1)) |a| (std.fmt.parseInt(usize, a, 10) catch null) else null; + // When the acpi-parse scenario spawns this directly, it passes `floor:N` — a + // device-count floor to self-verify against. The kernel no longer parses AML, so + // there is no exact count to match; proving the ring-3 parse found at least N + // Device objects is the check. Deterministic, no log-scraping. + // + // The `floor:` prefix matters. argv[1] is the assigned device id for every driver + // the manager spawns, so a bare number here would be read as a floor — which is + // exactly what happened when discovery started being given its node: it saw + // argv[1] = "7", decided it was in self-verify mode, and never reported a device. + const floor: ?usize = if (init.arguments.get(1)) |a| blk: { + if (!std.mem.startsWith(u8, a, "floor:")) break :blk null; + break :blk std.fmt.parseInt(usize, a["floor:".len..], 10) catch null; + } else null; const buffer = memory.allocator().alloc(device.DeviceDescriptor, 64) catch { _ = logging.write("/system/services/acpi: out of memory\n"); @@ -118,11 +126,11 @@ pub fn main(init: process.Init) void { _ = logging.write("/system/services/acpi: no acpi-tables node to claim\n"); return; }; + // The node arrived with the spawn: the manager holds it and names it in the call + // that creates this process. Discovery was the last thing in the system that + // acquired hardware by naming it rather than being given it + // (docs/os-development/device-authority.md). node_id = node.id; - device.claim(node_id) catch |e| { - std.log.warn("unable to claim acpi-tables: {s}", .{@errorName(e)}); - return; - }; // Map the node's resources: the AML blobs (bytecode), the FADT (intact // "FACP" header — decision 3), the io_port grant, and the SCI irq. diff --git a/system/services/device-manager/device-manager.zig b/system/services/device-manager/device-manager.zig index 0692992..a374cff 100644 --- a/system/services/device-manager/device-manager.zig +++ b/system/services/device-manager/device-manager.zig @@ -257,6 +257,7 @@ const delegated_drivers = [_][]const u8{ "pci-bus", "virtio-gpu", "ps2-bus", + "discovery", }; /// Matched on the **last path component**, because a driver reaches this table under @@ -420,6 +421,18 @@ fn initialise(endpoint: ipc.Handle) bool { const total = device.enumerate(buffer); const count = @min(total, buffer.len); + // The node discovery needs: the kernel seeds it, so it is in this same snapshot + // and can be handed over like any other assignment. Discovery used to find and + // claim it itself — the last driver that acquired hardware by naming it rather + // than being given it (docs/os-development/device-authority.md). + var tables_node: u64 = device_manager_protocol.no_device; + for (buffer[0..count]) |descriptor| { + if (descriptor.class == @intFromEnum(device.DeviceClass.acpi_tables)) { + tables_node = descriptor.id; + break; + } + } + var matched: usize = 0; for (buffer[0..count]) |descriptor| { if (descriptor.class == @intFromEnum(device.DeviceClass.pci_host_bridge)) { @@ -441,7 +454,7 @@ fn initialise(endpoint: ipc.Handle) bool { // under the neutral name "discovery", spawned once at startup. It finds and // claims the acpi-tables (or devicetree-blob) node itself. Not a per-device // match — it is the discoverer, not a driver bound to one device. - addDriver("discovery", device_manager_protocol.no_device, false); + addDriver("discovery", tables_node, false); if (test_restart_mode) { // The driver-restart scenario's fixture: claims device 0 (the tree From 337b981c0764451ac2291486fd55a734148a0d2c Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 21:58:53 +0100 Subject: [PATCH 30/36] docs: reorder the Run 2 table to the plan order and note the executed sequence --- docs/bounds-track-plan.md | 24 ++++++++++++++---------- 1 file changed, 14 insertions(+), 10 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index dd3f653..c8a7a73 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -39,17 +39,21 @@ that cannot safely run in user space.** | Step | What | State | |---|---|---| +| D0 | The grant rides `system_spawn` — atomic, so no driver need change to receive one | **done** | | D1 | `device_transfer(device_id, task_id)` — the holder gives a device away | **done** — syscall 54; a move, not a copy | -| D2 | Adversarial case: a process handed nothing is refused, on a held device and a free one | **done** — `device-authority-test`; the claim half joins it at D6 | -| D3 | The manager claims the seeded devices at boot, before any driver is spawned | **merged into D4** — see below | -| D4 | The manager claims + delegates on `hello`; `usb-xhci-bus` is the first driver converted | **done** — caught an IOMMU regression I introduced; see below | -| D5 | The other four claimants converted: `pci-bus`, `ps2-bus`, `virtio-gpu`, `acpi` | **done** — every driver now receives its hardware; question 9 answered | -| D0 | The grant rides `system_spawn` — atomic, so no driver need change to receive one | **done** — and caught a test-marker bug that made the D2 fixture unfailable | -| D10 | Every driver hellos, on its own merits (liveness, one class of driver) | not started — optional, independent | -| D6 | `device_claim` refuses a device the caller was not handed; the hole is closed | **blocked on question 10** — only the display service's GOP path still claims | -| D7 | Zero-resource devices stop being kernel objects — inventory moves to the manager | **blocked** — nothing else mints their ids; see question 8 | -| D8 | **`maximum_children_per_parent` deleted** | **done** — it was unblocked from the moment D9 landed; I kept reading my own stale label | -| D9 | The device table becomes dynamic; **`maximum_devices` deleted**; per-holder quota declared | **done** — one of the two invented numbers is gone | +| D2 | Adversarial case: a process handed nothing is refused | **done** — `device-authority-test`; the claim half joins it at D6 | +| D3 | The manager claims the seeded devices before any driver is spawned | **merged into D4** | +| D4 | `usb-xhci-bus` receives its controller | **done** — caught an IOMMU regression I introduced | +| D5 | `pci-bus`, `virtio-gpu`, `ps2-bus`, discovery receive theirs | **done** — every driver now receives its hardware | +| D6 | `device_claim` refuses a device the caller was not handed | **blocked on question 10** | +| D7 | Zero-resource devices stop being kernel objects | **blocked on question 8** | +| D8 | **`maximum_children_per_parent` deleted** | **done** | +| D9 | Dynamic table; **`maximum_devices` deleted**; per-registrar allowance | **done** | +| D10 | Every driver hellos, on its own merits | not started — optional, independent | + +*Executed in the order D1, D2, D4, D5(part), D9, D0, D5(rest), D8 — the numbering is +the original plan's, not the sequence. D0 was added mid-run, D3 merged into D4, and D8 +turned out to have been unblocked since D9.* **Run 2 stops, blocked on one question.** Landed: D0, D1, D2, D4, D8, D9, and D5 for three of five claimants (`usb-xhci-bus`, `pci-bus`, `virtio-gpu`). Suite 118/118. From ca1126537d91286e647608be35c66524e05a6c21 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 22:01:49 +0100 Subject: [PATCH 31/36] =?UTF-8?q?docs:=20Run=203=20=E2=80=94=20close=20the?= =?UTF-8?q?=20claiming=20hole,=20six=20steps,=20no=20open=20questions?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One field settles both outstanding questions. The kernel records who gave each device. device_claim then refuses anything that has a giver, and on death a device reverts to its giver rather than to nobody. The framebuffer needs no exemption: nobody delegates it, so it has no giver, so the display service claims it exactly as today. The rule never mentions display. And restart stops racing. Today a dying driver releases its claim to nobody and the manager re-claims first-come, so every restart reopens the hole this run closes; a loan that reverts to its lender removes the window entirely. The zero-resource inventory question is dropped from the plan rather than carried as a blocker. It is real but nothing depends on it. --- docs/bounds-track-plan.md | 53 +++++++++++++++++++++++++++++++++++++++ 1 file changed, 53 insertions(+) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index c8a7a73..8044f10 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -95,6 +95,59 @@ They land together, with the manager claiming only for drivers in an explicit anything in D4 — recorded rather than dismissed, because D4 moved the `hello` earlier and so did shift boot timing. Watch it across the remaining steps. +--- + +## Run 3 — close the claiming hole + +*Every driver now receives its hardware. What remains is that `device_claim` still +works for anyone, so a device nobody holds can still be taken. Six steps, no open +questions.* + +### The rule this run implements + +The kernel records, per device, **who gave it**. From that one field everything follows: + +- `device_transfer` and the spawn-fused grant set the giver. +- **`device_claim` refuses a device that has a giver.** A device that was given to + someone is delegated hardware and must be handed on, never taken. +- **On death a device reverts to its giver**, if that task is still alive; otherwise its + claim clears. A grant is a loan, not a gift. + +Two things fall out rather than being special-cased: + +- **The framebuffer is untouched.** Nobody delegates it, so it has no giver, so the + display service can still claim it exactly as today. No exemption in the kernel, no + mention of display anywhere in the rule. +- **Restart needs no race.** A dying driver's device returns to the manager, which + re-delegates it on respawn. Today the kernel releases it to nobody and the manager + re-claims first-come, so every restart reopens the hole this run closes. + +| Step | What | +|---|---| +| E1 | Record a giver per device; `device_transfer` and the spawn grant set it | +| E2 | On task death a device reverts to its giver if alive, else its claim clears | +| E3 | `device_claim` refuses a device that has a giver | +| E4 | The manager claims every resource-bearing device at boot, so nothing is left takeable | +| E5 | The attacker fixture gains the claim half it has been waiting for since D2 | +| E6 | Delete the delegated-set scaffolding — every driver is delegated now | + +Ordering: E1 alone changes no behaviour. E2 must precede E3, or a restart cannot +re-acquire. E4 must precede E5, or the attacker will find takeable devices and the +assertion will be wrong about why. E6 is cleanup. + +### Not in this run, and not blocking it + +**Zero-resource devices stay in the kernel.** Moving them out needs an answer to who +mints their ids, given that the id is the `device_token` on the usb-transfer wire and +the kernel's idempotency is what keeps it stable across a bus restart. Nothing depends +on it: both invented ceilings are already gone and this run does not need it. Recorded +as future work. + +**Not every driver hellos.** Delivery no longer needs it, so it stands or falls on its +own merits — uniform liveness and one class of driver. Also future work. + +--- + ### Question 9 — answered: a singleton gets every device that matched it The 8042 is one controller described by two ACPI nodes, so it cannot be split across From 4ca57fc37e10ab083d8cd3fbd7222939b8a9a9b5 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 22:12:01 +0100 Subject: [PATCH 32/36] kernel: record who gave each device away MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One field, and the rest of the run follows from it. A device that was given to someone is delegated hardware: it may be handed on, never taken, and when its holder dies it goes back to whoever lent it instead of becoming free for anyone to grab. It also settles the framebuffer without mentioning it. Nobody delegates the loader's framebuffer, so it has no giver, so the display service claims it exactly as it always has — no exemption and no reference to display anywhere in the rule. No behaviour changes here; the field is recorded and read by nothing yet. The test found a real bug on its first run, before the discrimination check. The sentinel for "nobody gave this" was 0 — and task 0 is a real task, the kernel's own, so a device given away by task 0 read back as belonging to nobody. Both giver and registrar are optionals now. The second was a latent bug from D9: the per-registrar allowance would have miscounted every device task 0 registered. Suite 118/118. --- docs/bounds-track-plan.md | 2 +- system/kernel/devices-broker.zig | 39 +++++++++++++++++++++++++++----- system/kernel/tests.zig | 12 ++++++++++ 3 files changed, 46 insertions(+), 7 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 8044f10..8a893a0 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -124,7 +124,7 @@ Two things fall out rather than being special-cased: | Step | What | |---|---| -| E1 | Record a giver per device; `device_transfer` and the spawn grant set it | +| E1 | Record a giver per device; `device_transfer` and the spawn grant set it — **done** | | E2 | On task death a device reverts to its giver if alive, else its claim clears | | E3 | `device_claim` refuses a device that has a giver | | E4 | The manager claims every resource-bearing device at boot, so nothing is left takeable | diff --git a/system/kernel/devices-broker.zig b/system/kernel/devices-broker.zig index 2d3acff..0b2b10e 100644 --- a/system/kernel/devices-broker.zig +++ b/system/kernel/devices-broker.zig @@ -65,8 +65,23 @@ var claimed: []?u32 = &.{}; /// The task that called `register` for each device, so the per-registrar allowance can /// be charged to whoever caused the entry. Firmware-discovered nodes carry `no_registrar` /// — they are the kernel's own, not anybody's doing. -var registrar: []u32 = &.{}; -const no_registrar: u32 = 0; +/// Optional, not a sentinel: **task 0 is a real task** (the kernel's own), so any +/// "none" value inside the id space is a device belonging to somebody reading back as +/// belonging to nobody. Found the moment the first test asserted a giver, because the +/// task doing the giving was task 0. +var registrar: []?u32 = &.{}; + +/// **Who gave each device away**, or `no_giver` if nobody ever did. +/// +/// One field, and the whole authority rule follows from it: a device that was *given* +/// to someone is delegated hardware, so it may only be handed on, never taken +/// (`claim` refuses it); and when its holder dies it goes back to whoever lent it, +/// rather than becoming free for anyone to grab. +/// +/// It also settles the framebuffer without mentioning it. Nobody delegates the +/// loader's framebuffer, so it has no giver, so the display service claims it exactly +/// as it always has — no exemption, no special case, no `display` anywhere in the rule. +var giver: []?u32 = &.{}; var count: usize = 0; /// Grow the three parallel arrays so at least one more device fits. False if the heap @@ -86,9 +101,12 @@ fn reserve() bool { claimed = grown_claimed; const grown_registrar = allocator.realloc(registrar, wanted) catch return false; registrar = grown_registrar; - for (claimed[count..], registrar[count..]) |*slot, *who| { + const grown_giver = allocator.realloc(giver, wanted) catch return false; + giver = grown_giver; + for (claimed[count..], registrar[count..], giver[count..]) |*slot, *who, *lender| { slot.* = null; - who.* = no_registrar; + who.* = null; + lender.* = null; } return true; } @@ -98,7 +116,7 @@ fn reserve() bool { fn registeredBy(task: u32) usize { var n: usize = 0; for (registrar[0..count]) |who| { - if (who == task) n += 1; + if (who != null and who.? == task) n += 1; } return n; } @@ -120,7 +138,8 @@ pub fn init(device_tree: *const platform.DeviceTree) void { dropped = 0; display_device = null; for (claimed) |*c| c.* = null; - for (registrar) |*r| r.* = no_registrar; + for (registrar) |*r| r.* = null; + for (giver) |*g| g.* = null; walk(device_tree.root, device_abi.no_parent); } @@ -235,6 +254,13 @@ pub fn claim(id: u64, owner: u32) ClaimError!void { claimed[@intCast(id)] = owner; } +/// The task that gave device `id` away, or null if nobody ever did. A device with a +/// giver is delegated hardware: it may be handed on, never taken. +pub fn giverOf(id: u64) ?u32 { + if (id >= count) return null; + return giver[@intCast(id)]; +} + /// The task that owns device `id`, or null. pub fn ownerOf(id: u64) ?u32 { if (id >= count) return null; @@ -408,6 +434,7 @@ pub fn transfer(id: u64, from: u32, to: u32) TransferError!void { const holder = claimed[@intCast(id)] orelse return error.NotHeld; if (holder != from) return error.NotHeld; claimed[@intCast(id)] = to; + giver[@intCast(id)] = from; } /// The errno a refused `claim` returns to ring 3. (`ECONFINE` — the claim stood but diff --git a/system/kernel/tests.zig b/system/kernel/tests.zig index e8ccf80..1f84bea 100644 --- a/system/kernel/tests.zig +++ b/system/kernel/tests.zig @@ -4121,6 +4121,10 @@ fn deviceTransferTest(boot_information: *const boot_handoff.BootInformation) voi const unheld = if (devices_broker.transfer(0, me, child)) |_| false else |e| e == error.NotHeld; check("an unheld device cannot be transferred", unheld); + // A second device this test never hands over, so the "no giver" half below is a + // real observation rather than a restatement of the first. + const parentless: u64 = 1; + check("claimed device 0", claimOk(0, me)); // The move itself. @@ -4139,6 +4143,14 @@ fn deviceTransferTest(boot_information: *const boot_handoff.BootInformation) voi const stranger = if (devices_broker.transfer(0, 9999, me)) |_| false else |e| e == error.NotHeld; check("a stranger cannot transfer another task's device", stranger); + // The giver is recorded. One field, and the authority rule follows from it: a + // device that was *given* to someone is delegated hardware, so it may be handed on + // but never taken, and when its holder dies it returns to whoever lent it instead + // of becoming free for anyone. Nothing has a giver until it is handed over — + // which is why the loader's framebuffer needs no exemption from either rule. + check("the giver is recorded on a transfer", devices_broker.giverOf(0) == me); + check("a device nobody handed over has no giver", devices_broker.giverOf(parentless) == null); + const absent = if (devices_broker.transfer(9999, me, child)) |_| false else |e| e == error.NoSuchDevice; check("a device that does not exist is refused", absent); From 1a1d92cba9b95d1e0dde49aae22756b8eabe6eac Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 22:21:25 +0100 Subject: [PATCH 33/36] =?UTF-8?q?kernel:=20a=20grant=20is=20a=20loan=20?= =?UTF-8?q?=E2=80=94=20a=20dead=20borrower=20returns=20the=20device=20to?= =?UTF-8?q?=20its=20lender?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit When a driver dies, a device it was *given* now goes back to whoever lent it, rather than to nobody. The device manager gets its hardware back the instant a driver dies and hands it to the replacement, with no window in between. That window was real: the kernel released the claim to no one and the manager re-claimed first-come, so every driver restart reopened the hole this run is closing. It also becomes load-bearing at the next step — once claim refuses a device that has a giver, releasing to nobody would strand a dead driver's hardware permanently, because nobody could ever take it again. A dead lender is no lender: the claim and the giver clear together, so a device is never owed to a ghost. A device nobody lent is released outright, exactly as before. The broker cannot see the task table, so liveness arrives through the same hook idiom the scheduler already uses. Null means assume dead, so a kernel built without the hook frees claims rather than handing them to a ghost. A stale binary nearly passed as proof for the third time this session: the first discrimination patch left `alive` unused, the build failed with three errors, and the old binary reported every assertion passing. Checking the build before reading results is what caught it. Suite 118/118. --- docs/bounds-track-plan.md | 2 +- system/kernel/devices-broker.zig | 36 +++++++++++++++++++++++++++++--- system/kernel/process.zig | 9 ++++++++ system/kernel/tests.zig | 12 +++++++++++ 4 files changed, 55 insertions(+), 4 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 8a893a0..49e6968 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -125,7 +125,7 @@ Two things fall out rather than being special-cased: | Step | What | |---|---| | E1 | Record a giver per device; `device_transfer` and the spawn grant set it — **done** | -| E2 | On task death a device reverts to its giver if alive, else its claim clears | +| E2 | On task death a device reverts to its giver if alive, else its claim clears — **done** | | E3 | `device_claim` refuses a device that has a giver | | E4 | The manager claims every resource-bearing device at boot, so nothing is left takeable | | E5 | The attacker fixture gains the claim half it has been waiting for since D2 | diff --git a/system/kernel/devices-broker.zig b/system/kernel/devices-broker.zig index 0b2b10e..7b16b73 100644 --- a/system/kernel/devices-broker.zig +++ b/system/kernel/devices-broker.zig @@ -272,11 +272,41 @@ pub fn ownerOf(id: u64) ?u32 { /// hardware again (docs/process-lifecycle.md iron rule 1: cleanup is the kernel's /// job). The devices stay in the table — they describe hardware, which did not go /// away — only their ownership clears. +/// Whether a task is still alive, injected by the process layer (which owns the task +/// table) the same way the scheduler's other hooks are. Null means "assume not", so a +/// kernel built without it clears claims rather than handing them to a ghost. +pub var task_alive_hook: ?*const fn (u32) bool = null; + +fn alive(task: u32) bool { + const hook = task_alive_hook orelse return false; + return hook(task); +} + pub fn releaseAllOwnedBy(owner: u32) void { - for (claimed[0..count]) |*slot| { - if (slot.*) |o| { - if (o == owner) slot.* = null; + for (claimed[0..count], 0..) |*slot, id| { + const holder = slot.* orelse continue; + if (holder != owner) continue; + + // **A grant is a loan.** A device this task was *given* goes back to whoever + // lent it, not to nobody — so the device manager gets its hardware back the + // instant a driver dies, and hands it to the replacement. + // + // Without this the kernel released the claim to no one and the manager + // re-claimed first-come, so every driver restart reopened the window this + // rule closes. And once `claim` refuses a device that has a giver, releasing + // to nobody would strand it: no one could ever take it again. + // + // A dead lender is no lender: clear the claim and the giver together, so the + // device is genuinely free rather than owed to a ghost. + if (giver[id]) |lender| { + if (alive(lender)) { + slot.* = lender; + giver[id] = null; // returned; it is the lender's own again, not on loan + continue; + } + giver[id] = null; } + slot.* = null; } } diff --git a/system/kernel/process.zig b/system/kernel/process.zig index 756e783..e167312 100644 --- a/system/kernel/process.zig +++ b/system/kernel/process.zig @@ -190,6 +190,15 @@ pub fn init() void { scheduler.timer_tick_hook = timerSweepLocked; scheduler.group_exit_hook = groupExitLocked; scheduler.space_mapping_release_hook = dropSpaceMappingHook; + // The broker owns devices; the process layer owns the task table. It asks whether a + // lender is still alive before handing a dead driver's device back to it. + devices_broker.task_alive_hook = taskAliveLocked; +} + +/// Whether `id` names a live task. The broker calls this through its hook when +/// deciding if a dead holder's device can go back to the task that lent it. +fn taskAliveLocked(id: u32) bool { + return scheduler.taskByIdLocked(id) != null; } /// Return -1 (as an unsigned bit pattern) in the system_call result register. diff --git a/system/kernel/tests.zig b/system/kernel/tests.zig index 1f84bea..7681ce4 100644 --- a/system/kernel/tests.zig +++ b/system/kernel/tests.zig @@ -4154,8 +4154,20 @@ fn deviceTransferTest(boot_information: *const boot_handoff.BootInformation) voi const absent = if (devices_broker.transfer(9999, me, child)) |_| false else |e| e == error.NoSuchDevice; check("a device that does not exist is refused", absent); + // **A grant is a loan.** The child holds device 0 and `me` lent it, so the child's + // death must hand it back rather than release it to nobody. That is what lets the + // device manager re-delegate to a restarted driver — and, once `claim` refuses a + // device that has a giver, it is the only thing that stops a dead driver's hardware + // being stranded forever. + check("the child still holds the lent device", devices_broker.ownerOf(0) == child); devices_broker.releaseAllOwnedBy(child); + check("the borrower's death returns the device to its lender", devices_broker.ownerOf(0) == me); + check("and it is no longer on loan", devices_broker.giverOf(0) == null); + + // A device with no lender still simply frees on death, as it always did. + check("claimed a second device with no lender", claimOk(parentless, me)); devices_broker.releaseAllOwnedBy(me); + check("a device nobody lent is released outright", devices_broker.ownerOf(parentless) == null); result(); } From ba195fa0a217a599392bb85efbcf2d3eb6a7379c Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 22:31:04 +0100 Subject: [PATCH 34/36] =?UTF-8?q?kernel:=20claim=20refuses=20delegated=20h?= =?UTF-8?q?ardware=20=E2=80=94=20and=20E2=20had=20already=20closed=20the?= =?UTF-8?q?=20hole?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The rule as planned: a device that was given to someone may be handed on, never taken. Implemented, and honest about what it is worth. Writing the test showed the plan had the wrong step doing the work. A delegated device is HELD, so an attempt to take it is refused as AlreadyClaimed before the giver is ever consulted; and once a borrower's death returns the device to its lender — or clears both when the lender is gone — there is no state where a device is unheld and still on loan. The window a stranger could have used stops existing at E2. This check is unreachable. It stays anyway: one comparison, failing closed, guarding any future path that frees a device without clearing its giver, which is exactly the hole this run closed. The comment says it is unreachable rather than implying a protection it does not provide. The attacker fixture does not gain the assertion that was deferred to this step, and its header records why: there is no refusal for it to observe, and on a bare boot with no device manager nothing is delegated at all, so the assertion had nothing to bite on. It failed loudly on its first run rather than passing quietly, which is the only reason this was noticed. It also leaves the loader's framebuffer alone without naming it: nobody delegates the framebuffer, so it has no giver, so the display service claims it exactly as before. Suite 118/118. --- docs/bounds-track-plan.md | 14 +++++++++++--- library/device/driver/driver.zig | 3 ++- system/kernel/devices-broker.zig | 18 ++++++++++++++++++ .../device-authority-test.zig | 13 +++++++------ 4 files changed, 38 insertions(+), 10 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 49e6968..c57192c 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -119,14 +119,22 @@ Two things fall out rather than being special-cased: display service can still claim it exactly as today. No exemption in the kernel, no mention of display anywhere in the rule. - **Restart needs no race.** A dying driver's device returns to the manager, which - re-delegates it on respawn. Today the kernel releases it to nobody and the manager - re-claims first-come, so every restart reopens the hole this run closes. + re-delegates it on respawn. Previously the kernel released it to nobody and the + manager re-claimed first-come, so every restart reopened the hole. + +**Found while implementing: E2 is the step that closes the hole, not E3.** A delegated +device is *held*, so an attempt to take it is refused as `AlreadyClaimed` long before +the giver is consulted — and once a borrower's death returns the device to its lender +(or clears both when the lender is gone), there is no state where a device is unheld and +still on loan. E3's check is therefore unreachable today. It stays as one comparison +that fails closed, guarding any future path that frees a device without clearing its +giver, and its comment says so rather than implying a protection it is not providing. | Step | What | |---|---| | E1 | Record a giver per device; `device_transfer` and the spawn grant set it — **done** | | E2 | On task death a device reverts to its giver if alive, else its claim clears — **done** | -| E3 | `device_claim` refuses a device that has a giver | +| E3 | `device_claim` refuses a device that has a giver — **done**, but unreachable: E2 already closed the window | | E4 | The manager claims every resource-bearing device at boot, so nothing is left takeable | | E5 | The attacker fixture gains the claim half it has been waiting for since D2 | | E6 | Delete the delegated-set scaffolding — every driver is delegated now | diff --git a/library/device/driver/driver.zig b/library/device/driver/driver.zig index 3a9633c..0b2ea4b 100644 --- a/library/device/driver/driver.zig +++ b/library/device/driver/driver.zig @@ -70,7 +70,7 @@ pub fn transfer(id: u64, to: u32) TransferError!void { /// re-enumerate, and `NotConfined` means the machine could not place the device under /// IOMMU translation — the claim was rolled back, and that one is a fault report, not /// a retry. `Refused` is an errno this library does not know a name for. -pub const ClaimError = error{ NoSuchDevice, AlreadyClaimed, NotConfined, Refused }; +pub const ClaimError = error{ NoSuchDevice, AlreadyClaimed, NotConfined, NotYours, Refused }; /// Take exclusive ownership of device `id`. pub fn claim(id: u64) ClaimError!void { @@ -80,6 +80,7 @@ pub fn claim(id: u64) ClaimError!void { abi.ENODEV => error.NoSuchDevice, abi.EBUSY => error.AlreadyClaimed, abi.ECONFINE => error.NotConfined, + abi.EPERM => error.NotYours, // delegated hardware: it must be handed to you else => error.Refused, }; } diff --git a/system/kernel/devices-broker.zig b/system/kernel/devices-broker.zig index 7b16b73..0d6b8a2 100644 --- a/system/kernel/devices-broker.zig +++ b/system/kernel/devices-broker.zig @@ -251,6 +251,22 @@ pub fn enumerateFrom(start: usize, out: []device_abi.DeviceDescriptor) usize { pub fn claim(id: u64, owner: u32) ClaimError!void { if (id >= count) return error.NoSuchDevice; if (claimed[@intCast(id)] != null) return error.AlreadyClaimed; + // **Delegated hardware may be handed on, never taken.** + // + // Belt and braces, and worth being honest about: with the loan rule above this is + // **currently unreachable**. A device that was given to someone is held, so it is + // refused as `AlreadyClaimed` before reaching here; and when the holder dies the + // device goes back to its lender (or, if the lender is gone, has its giver cleared + // with its claim), so there is no state where a device is unheld *and* still on + // loan. The window a stranger could have used simply stops existing. + // + // It stays because it is one comparison and it fails closed: any future path that + // frees a device without clearing its giver would otherwise hand delegated + // hardware to whoever asked first, which is exactly the hole this run closed. + // + // Note it leaves the loader's framebuffer alone without naming it: nobody delegates + // the framebuffer, so it has no giver, so the display service claims it as always. + if (giver[@intCast(id)] != null) return error.NotYours; claimed[@intCast(id)] = owner; } @@ -428,6 +444,7 @@ pub fn errnoOf(e: RegisterError) i64 { pub const ClaimError = error{ NoSuchDevice, // no device with that id AlreadyClaimed, // a live task already owns it + NotYours, // delegated hardware: it has a giver, so it must be handed on, not taken }; /// Why a `transfer` was refused. @@ -474,6 +491,7 @@ pub fn claimErrnoOf(e: ClaimError) i64 { return switch (e) { error.NoSuchDevice => abi.ENODEV, error.AlreadyClaimed => abi.EBUSY, + error.NotYours => abi.EPERM, }; } diff --git a/test/system/services/device-authority-test/device-authority-test.zig b/test/system/services/device-authority-test/device-authority-test.zig index fea745b..8dc10f5 100644 --- a/test/system/services/device-authority-test/device-authority-test.zig +++ b/test/system/services/device-authority-test/device-authority-test.zig @@ -25,12 +25,13 @@ //! is what cost a debugging session on the Ryzen, so the distinction is //! part of the contract and is tested as such. //! -//! **What this fixture cannot yet claim.** `device_claim` is still -//! first-come-first-served at this point in the run — that is the hole D6 -//! closes. So the claim half of the invariant ("a process holds what it was -//! handed and cannot name its way into holding more") is deliberately NOT -//! asserted here; it is added to this fixture at D6, when it becomes true. -//! Asserting it now would mean writing a test that documents the bug. +//! **Why there is no "cannot take a delegated device" assertion here.** The +//! hole this fixture was written for is closed, but not by a refusal it could +//! observe. A device that was given to someone is *held*, so an attempt to +//! take it is refused as `AlreadyClaimed` — the same answer as before. What +//! changed is what happens when the holder dies: the device returns to +//! whoever lent it instead of becoming free, so the window in which a +//! stranger could take it no longer exists. There is no moment to catch. const std = @import("std"); const device = @import("driver"); From df9c1ed82743977e1a98ee05e7e8a1236fbaebed Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 22:42:42 +0100 Subject: [PATCH 35/36] device-manager: hold the seeded hardware so none is left lying around MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A device nobody holds can be claimed by anyone, so the manager now takes every firmware-discovered device that carries mappable resources, whether or not a driver wants it. The real gap was the HPET: an MMIO window, an IRQ, no user-space driver, and there for the taking. Held by the manager it is inert; unheld it was a way into physical memory. Two deliberate exclusions. The loader's framebuffer, which the compositor claims and which the manager must not take because it starts first. And anything with no resources, which grants nothing worth holding. Scope is the boot snapshot. A device reported later and matched to no driver stays claimable — pci-cap-test and iommu-fault-test both reach an unmatched NIC that way, so narrowing it is a separate change with those fixtures in scope. Recorded in the plan rather than left implied. The attacker fixture gains the assertion deferred since D2: after the system settles, nothing with resources may be taken. That assertion defeated itself twice before it worked, and both failures are worth remembering. First it swept at 0.029 while the manager did not bind its protocol until 0.047, so it reported a hole that closed a millisecond later. The retry loop that "fixed" that was worse: the first pass TAKES the device, so the second finds it unavailable because this process now holds it, and concludes all is well — it passed with the manager's claiming removed entirely. It now settles once and sweeps once, and fails when the claiming is removed. Suite 118/118. --- docs/bounds-track-plan.md | 15 ++++++++++-- system/kernel/tests.zig | 5 ++++ .../device-manager/device-manager.zig | 22 ++++++++++++++++++ .../services/device-authority-test/build.zig | 2 +- .../device-authority-test.zig | 23 +++++++++++++++++++ 5 files changed, 64 insertions(+), 3 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index c57192c..2e37820 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -135,14 +135,25 @@ giver, and its comment says so rather than implying a protection it is not provi | E1 | Record a giver per device; `device_transfer` and the spawn grant set it — **done** | | E2 | On task death a device reverts to its giver if alive, else its claim clears — **done** | | E3 | `device_claim` refuses a device that has a giver — **done**, but unreachable: E2 already closed the window | -| E4 | The manager claims every resource-bearing device at boot, so nothing is left takeable | -| E5 | The attacker fixture gains the claim half it has been waiting for since D2 | +| E4 | The manager claims every resource-bearing device at boot, so nothing is left takeable — **done** (boot snapshot only; see below) | +| E5 | The attacker fixture gains the claim half it has been waiting for since D2 — **done with E4** | | E6 | Delete the delegated-set scaffolding — every driver is delegated now | Ordering: E1 alone changes no behaviour. E2 must precede E3, or a restart cannot re-acquire. E4 must precede E5, or the attacker will find takeable devices and the assertion will be wrong about why. E6 is cleanup. +### E4's scope, stated + +It covers the **boot snapshot**. A device *reported* later and matched to no driver +stays claimable — `pci-cap-test` and `iommu-fault-test` both rely on that to reach an +unmatched NIC. Narrowing it further is a separate change with those fixtures in scope. + +The real gap it closed was the **HPET**: an MMIO window, an IRQ, no user-space driver, +and claimable by anyone. Excluded on purpose: the loader's framebuffer (the manager +starts before display, so taking it would break the boot screen) and anything with no +resources, which grants nothing. + ### Not in this run, and not blocking it **Zero-resource devices stay in the kernel.** Moving them out needs an answer to who diff --git a/system/kernel/tests.zig b/system/kernel/tests.zig index 7681ce4..7d03c72 100644 --- a/system/kernel/tests.zig +++ b/system/kernel/tests.zig @@ -4334,6 +4334,11 @@ fn deviceAuthorityTest(boot_information: *const BootInformation) void { }; process.setInitialRamdisk(image); + // The manager must be up: it is what holds the seeded hardware, and without it + // every device would be lying around unheld and the last assertion would have + // nothing to observe — a test that cannot fail. + check("registry (init) spawned", spawnRegistry(rd)); + check("device-manager spawned", spawnNamed(rd, "device-manager")); check("device-authority-test spawned", spawnNamedWithArg(rd, "device-authority-test", "run")); // The VERDICT prefix matters: the fixture prints one "device-authority: ok " diff --git a/system/services/device-manager/device-manager.zig b/system/services/device-manager/device-manager.zig index a374cff..daf0c19 100644 --- a/system/services/device-manager/device-manager.zig +++ b/system/services/device-manager/device-manager.zig @@ -433,6 +433,28 @@ fn initialise(endpoint: ipc.Handle) bool { } } + // **Hold the firmware-discovered hardware, so none of it is left lying around.** + // A device nobody holds can be claimed by anyone, so every seeded device that + // carries mappable resources is taken here whether or not a driver wants it — the + // HPET most of all, which has an MMIO window and an IRQ and no user-space driver. + // Held by the manager it is inert; unheld it was there for the taking. + // + // Two deliberate exclusions: + // - the loader's framebuffer, which the compositor claims and which is not + // hardware anyone is delegated (the manager starts before display, so taking + // it here would break the boot screen); + // - anything with no resources, which grants nothing and so is not worth holding. + // + // This covers the boot snapshot only. A device *reported* later and matched to no + // driver stays claimable — the pci-cap and iommu-fault fixtures rely on exactly + // that to reach an unmatched NIC. Narrowing it further is a separate change with + // those fixtures in scope. + for (buffer[0..count]) |descriptor| { + if (descriptor.resource_count == 0) continue; + if (descriptor.class == @intFromEnum(device.DeviceClass.display)) continue; + device.claim(descriptor.id) catch continue; // already held, or not ours to take + } + var matched: usize = 0; for (buffer[0..count]) |descriptor| { if (descriptor.class == @intFromEnum(device.DeviceClass.pci_host_bridge)) { diff --git a/test/system/services/device-authority-test/build.zig b/test/system/services/device-authority-test/build.zig index 1b19cfd..b5dabc0 100644 --- a/test/system/services/device-authority-test/build.zig +++ b/test/system/services/device-authority-test/build.zig @@ -9,7 +9,7 @@ pub fn build(b: *std.Build) void { const exe = build_support.userBinary(b, .{ .name = "device-authority-test", .root_source_file = b.path("device-authority-test.zig"), - .imports = &.{ "driver", "logging", "process" }, + .imports = &.{ "driver", "logging", "process", "time" }, }); b.installArtifact(exe); } diff --git a/test/system/services/device-authority-test/device-authority-test.zig b/test/system/services/device-authority-test/device-authority-test.zig index 8dc10f5..01fd317 100644 --- a/test/system/services/device-authority-test/device-authority-test.zig +++ b/test/system/services/device-authority-test/device-authority-test.zig @@ -37,6 +37,7 @@ const std = @import("std"); const device = @import("driver"); const logging = @import("logging"); const process = @import("process"); +const time = @import("time"); fn line(comptime format: []const u8, arguments: anytype) void { var buffer: [160]u8 = undefined; @@ -88,6 +89,28 @@ fn run() void { check("and no child was left behind by the refusal", process.processes(&process_table) == before); } + // 5. **Nothing with mappable resources is left lying around.** A device nobody + // holds can be claimed by anyone, so the manager takes every seeded device that + // carries resources — the HPET above all, which has an MMIO window and an IRQ + // and no user-space driver. The one exception is the loader's framebuffer, + // which the compositor claims. So from here, a resource-bearing device should + // refuse to be taken, and the reason should be that someone already has it. + // Settle first, then sweep **once**. The manager is still starting when this + // fixture is spawned, so an immediate sweep finds hardware unheld and reports a + // hole that closes a millisecond later. Retrying until the sweep comes back + // empty is worse than useless: the first pass *takes* the device, so the second + // finds it unavailable — because this process now holds it — and concludes all + // is well. One sweep, after a wait long enough for the manager to have claimed. + time.sleepMillis(1500); + var takeable: usize = 0; + for (table[0..seen]) |descriptor| { + if (descriptor.resource_count == 0) continue; + if (descriptor.class == @intFromEnum(device.DeviceClass.display)) continue; + device.claim(descriptor.id) catch continue; // refused, as it should be + takeable += 1; + } + check("no resource-bearing device is left for the taking", takeable == 0); + if (failures == 0) { line("device-authority: VERDICT ok ({d} devices, none of them mine)\n", .{seen}); } else { From 0eb2420690c7af228c5b15ff9a8330d9b0c9d54f Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 23:01:29 +0100 Subject: [PATCH 36/36] device-manager: delete the delegated-set scaffolding MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The name list and its predicate existed so drivers could move to delegation one at a time with the suite green throughout. Every driver is delegated now, so the manager simply hands over whatever device a driver was assigned. Deleting it caught a real consequence: crash-test finally got delegated too, and it was still claiming its device — so it got AlreadyClaimed because it already held it, exited, and the restart drill had nothing to restart. Its own comment named what the case was really checking: "the respawn only reaches this line because the kernel released the previous instance's claim at death". That property still holds, by a different mechanism — the device reverts to the manager on death and is handed to the replacement, which is the same guarantee without the race it used to rely on. All four delegation paths verified: the xHCI controller, the PCI bridge, the PS/2 two-node singleton, and virtio-gpu's restart re-attach. Run 3 complete. Suite 118/118. --- docs/bounds-track-plan.md | 6 +++- .../device-manager/device-manager.zig | 33 ++----------------- .../system/services/crash-test/crash-test.zig | 16 +++++---- 3 files changed, 16 insertions(+), 39 deletions(-) diff --git a/docs/bounds-track-plan.md b/docs/bounds-track-plan.md index 2e37820..ce0f813 100644 --- a/docs/bounds-track-plan.md +++ b/docs/bounds-track-plan.md @@ -137,12 +137,16 @@ giver, and its comment says so rather than implying a protection it is not provi | E3 | `device_claim` refuses a device that has a giver — **done**, but unreachable: E2 already closed the window | | E4 | The manager claims every resource-bearing device at boot, so nothing is left takeable — **done** (boot snapshot only; see below) | | E5 | The attacker fixture gains the claim half it has been waiting for since D2 — **done with E4** | -| E6 | Delete the delegated-set scaffolding — every driver is delegated now | +| E6 | Delete the delegated-set scaffolding — every driver is delegated now — **done** | Ordering: E1 alone changes no behaviour. E2 must precede E3, or a restart cannot re-acquire. E4 must precede E5, or the attacker will find takeable devices and the assertion will be wrong about why. E6 is cleanup. +**Run 3 complete.** Suite 118/118. Every driver receives its hardware; a grant is a +loan that returns to its lender when the borrower dies; nothing firmware-discovered is +left unheld except the framebuffer, which the compositor owns. + ### E4's scope, stated It covers the **boot snapshot**. A device *reported* later and matched to no driver diff --git a/system/services/device-manager/device-manager.zig b/system/services/device-manager/device-manager.zig index daf0c19..5c04e0c 100644 --- a/system/services/device-manager/device-manager.zig +++ b/system/services/device-manager/device-manager.zig @@ -244,35 +244,6 @@ fn alreadySupervised(name: []const u8) bool { return false; } -/// Drivers that receive their device from the manager rather than claiming it -/// themselves. Scaffolding for the conversion, not a permanent concept: it exists so -/// each driver can move across one at a time with the suite green throughout, and it -/// disappears at D6 when `device_claim` stops being a way to acquire a device at all -/// (docs/bounds-track-plan.md, Run 2). -/// -/// `usb-xhci-bus` is first because it was the first driver to conform to `hello` -/// (device-manager.md, M18.1), so it is the one whose handshake is best proven. -const delegated_drivers = [_][]const u8{ - "usb-xhci-bus", - "pci-bus", - "virtio-gpu", - "ps2-bus", - "discovery", -}; - -/// Matched on the **last path component**, because a driver reaches this table under -/// two different spellings: the boot-snapshot match records the bare `pci-bus`, while a -/// devices.csv match records the full `/system/drivers/pci-bus`. Comparing whole -/// strings silently missed the bare form — pci-bus was left neither claiming nor -/// delegated, and died on `ECAM mmio_map failed`. -fn isDelegated(name: []const u8) bool { - const leaf = if (std.mem.lastIndexOfScalar(u8, name, '/')) |slash| name[slash + 1 ..] else name; - for (delegated_drivers) |candidate| { - if (std.mem.eql(u8, leaf, candidate)) return true; - } - return false; -} - /// Record a driver in the table and spawn its first instance. fn addDriver(name: []const u8, device_id: u64, speaks_protocol: bool) void { for (&drivers) |*driver| { @@ -298,7 +269,7 @@ fn spawnDriver(driver: *Driver) void { // rather than advisory. Re-claiming across a restart is expected to say // AlreadyClaimed once the manager already holds it, and that is fine: it means the // device never left our hands while the driver was dead. - if (isDelegated(driver.name()) and driver.device_id != device_manager_protocol.no_device) { + if (driver.device_id != device_manager_protocol.no_device) { device.claim(driver.device_id) catch |e| switch (e) { error.AlreadyClaimed => {}, // ours already, from a previous spawn of this driver else => { @@ -320,7 +291,7 @@ fn spawnDriver(driver: *Driver) void { // A transfer *after* spawning would leave a window in which the child is running // without its hardware — closed on one machine, open on another // (docs/bounds-track-plan.md, "the grant rides system_spawn"). - const give = if (isDelegated(driver.name())) driver.device_id else device_manager_protocol.no_device; + const give = driver.device_id; if (give != device_manager_protocol.no_device) std.log.info("delegated device {d} to {s}", .{ give, driver.name() }); const child = process.spawnSupervisedWithDevice(driver.name(), arguments[0..argument_count], manager_endpoint, give) orelse { diff --git a/test/system/services/crash-test/crash-test.zig b/test/system/services/crash-test/crash-test.zig index b5c262a..5806220 100644 --- a/test/system/services/crash-test/crash-test.zig +++ b/test/system/services/crash-test/crash-test.zig @@ -19,13 +19,15 @@ pub fn main(init: process.Init) void { const argument = init.arguments.get(1) orelse return; // bare: stay silent const assigned = std.fmt.parseInt(u64, argument, 10) catch return; - // The respawn only reaches this line because the kernel released the - // previous instance's claim at death. A failed claim exits cleanly — the - // manager reads "meant to stop" and the scenario fails loudly by silence. - device.claim(assigned) catch { - _ = logging.write("crash-test: claim failed\n"); - return; - }; + // The device arrived with the spawn — this fixture is delegated its hardware like + // any other driver, so it holds `assigned` before its first instruction and has + // nothing to claim (docs/os-development/device-authority.md). + // + // The property this scenario checks is unchanged, only its mechanism: a respawned + // instance still gets the device its predecessor held. It used to arrive because + // the kernel released the dead instance's claim and this one re-took it, racing + // anyone else who wanted it; now the device reverts to the manager on death and is + // handed to the replacement, which is the same guarantee without the race. var manager: ?ipc.Handle = null; var tries: u32 = 0;