# Device authority: you hold what you were given *Design for phase 2 of [the bounds track](../bounds-track-plan.md), 2026-08-08. Supersedes an earlier draft that argued for capabilities on aesthetic grounds; this one starts from the hole and from the project's principles.* ## The hole `device_claim(id)` checks two things ([devices-broker.zig](../../system/kernel/devices-broker.zig)): ```zig pub fn claim(id: u64, owner: u32) ClaimError!void { if (id >= count) return error.NoSuchDevice; if (claimed[@intCast(id)] != null) return error.AlreadyClaimed; claimed[@intCast(id)] = owner; } ``` Does it exist, and is it free. **Any process may claim any unclaimed device.** The matching is real but it is entirely advisory: `device-manager` reads `devices.csv`, matches a device to a driver, and spawns that driver with the device id as `argv[1]` (`spawnDriver`). The driver parses the string and claims it. Nothing anywhere binds the manager's decision to the kernel's grant — a process can pass any integer and win the race. A claim is not a small thing. It is what gates `mmio_map` and `irq_bind`, so it is a licence to map physical memory and receive interrupts. ## What this costs, beyond the obvious `maximum_children_per_parent = 16` exists because a driver that claimed one device could loop `device_register` under it and exhaust the shared table. That threat only exists *because* claiming is unauthenticated — and the cap is a poor defence against it, since an attacker can burn 16 slots, claim another device, and burn 16 more. What it reliably does instead is refuse a legitimate PCI bus with more than 16 functions, which is how an AMD Ryzen came to boot with no USB and no storage. So the cap is not merely mis-sized. It is standing in for an authorisation that is not performed, and it punishes correct behaviour while barely inconveniencing incorrect behaviour. **Closing the hole is what retires the constant**, not a bigger number. ## What the principles decide - *Move as much responsibility as possible to user space* (3), and *what remains in the kernel is there for security or a hardware limitation* (5). Deciding **which** driver gets **which** device is policy — matching identity triples and choosing a binary. That decision belongs to the device manager and stays there. `devices.csv` is not the policy; it is **configuration**, the declarative data the policy reads. Three distinct things, and worth keeping apart in this document: | | Lives in | Example | |---|---|---| | **Mechanism** | the kernel | the check that a grant is held before a mapping is made | | **Policy** | user space | the device manager matching a device to a driver | | **Configuration** | files | `devices.csv`, `protocol.csv`, `init.csv` | Enforcing that a driver **holds only what it was given** is security — it is the gate in front of mapping physical memory. That stays in the kernel, and it is the whole of what the kernel needs to do. The kernel therefore does not need to know about matching, `devices.csv`, driver names, or why a device was assigned. It needs to know that an authority it can verify granted this device to this task. ## Only drivers hold devices The layering the rest of the system already follows: - A **driver** talks to hardware. It is the only thing that holds a device. - A **service** talks to no hardware at all. It receives device events and sends data and commands over a **protocol**. - A **protocol** is the abstraction between them — the OS layer, in the sense of [communication.md](communication.md). So the question "which services need device grants" has the answer **none**. That collapses the design: every holder of a device is a driver, and every driver is spawned by the device manager, which is what grants it. `display` already demonstrates both halves, one right and one wrong ([backend.zig](../../system/services/display/backend.zig)): | Backend | How it reaches the panel | |---|---| | `VirtioGpu` | speaks `scanout-protocol` over an IPC handle; touches no device | | `Gop` | `device.enumerate`, `device.claim`, `device.mmioMap` | The virtio-gpu path is the intended shape and it works today: the driver holds the hardware, the service speaks a protocol to it. **The GOP path is the anomaly** — the display service reaching into hardware itself, because the firmware framebuffer has no driver for it to talk to. It is also the only anomaly. Every other claimant — `pci-bus`, `usb-xhci-bus`, `ps2-bus`, `virtio-gpu`, `acpi` — is a device-manager child spawned with its device id in `argv[1]`. **A firmware-framebuffer driver is therefore a prerequisite of this phase**, not a consequence of it. It is small: it holds the display device, maps the framebuffer, and serves `scanout-protocol` — the same interface `virtio-gpu` already serves, so the compositor needs no new code path and stops caring which is behind it, which it was designed for. An earlier draft of this document had `display` asking the device manager for a grant, and before that had `init` minting grants because `init` starts `display`. Both were accommodating an exception instead of removing it. ## The design **A device grant is a capability, delegated from a holder.** The mechanism already exists: `callCap` passes a handle over an IPC call and the kernel installs it in the receiver's table ([library/kernel/ipc.zig](../../library/kernel/ipc.zig)), which is how shared memory and DMA regions already move between processes. **The device manager is the root of device authority.** `init` has no part in this: it is the first process, and its job is to start the rest of the system. It starts the device manager the same way it starts everything else, and knows nothing about devices. 1. **Root.** At boot the kernel mints grants for the devices firmware discovery found and hands them to the device manager. This is the only place device authority enters the system, and it comes from ACPI rather than from anyone's say-so. The kernel recognises the manager by **chain attestation**, the identity the security track already settled on: the binary path the kernel itself stamped (`/system/services/device-manager`) *and* a supervisor of PID 1. The binary path alone would not do — `system_spawn` is deliberately ungated, so any process may spawn any bundled binary, and a rogue could run a second copy under the same name. It could not forge the other half: its copy's supervisor is the rogue, and PID 1 is the kernel's own first process. 2. **Delegation.** The manager passes a driver its device when it spawns it, over the channel that already exists — the driver `hello`s the manager and the reply carries the grant. The manager attests its own children by task id, one hop deep, exactly as `init` attests the services it spawned for `/protocol` (`supervisorSatisfies`); it knows which task is which driver because it spawned them. 3. **Services receive nothing.** `display`, `input`, `fat` and `logger` hold no devices and need no grants; they speak protocols to the drivers that do. The one place this is not true today is the GOP backend above, which a firmware-framebuffer driver removes. 4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and `dma_bind` check possession of the grant instead of consulting an ownership table. Exclusivity stops being a broker refusing a second claimant and becomes the ordinary property of a capability: only one process was given it. **`maximum_children_per_parent` is deleted here.** After this a bus driver's children are devices it enumerated on a bus it was actually given, and the threat the cap was written for no longer exists. ## What this costs **Bring-up order changes.** `pci-bus` today claims first and says hello afterwards — its own comment says "Claim the bridge, map the ECAM, hello the manager, then scan." Under delegation the hello must come first, because that is where the grant arrives. Five drivers need that reordering, and it is the bulk of the work. **A firmware-framebuffer driver has to be written first**, and until it exists the display service cannot stop claiming a device. It is the smallest new binary in the tree — hold the display node, map the framebuffer, serve `scanout-protocol` — but it is new code on the boot path, and the boot path is where a mistake costs a screen. It brings an ordering constraint with it: `display` cannot paint until that driver is up, where today it maps the framebuffer itself and is independent. Mitigating, `display` already tolerates a backend arriving late — `virtio-gpu` hands it a scanout after the fact over `attach_scanout` — so "wait, then attach" is a path that already works rather than one to invent. **The kernel gains one piece of knowledge about a specific binary.** Chain attestation means the kernel recognises `/system/services/device-manager` under PID 1 as the root holder. That is a real concession — the kernel would rather know nothing about who is who — and it is the minimum: device authority has to enter the system somewhere, and every alternative is worse. Configuration the kernel reads would put a file parser in the kernel; first-to-ask would be the hole again, at boot. **A configuration question stays open.** Which binary may be given which device is expressed today as `devices.csv`, read by the manager — configuration, with the manager as the policy that acts on it. That is already the right shape and needs no new file. A separate grant manifest earns itself only when something must differ from "the manager gets the hardware": holding a device back so a test or a bare-metal driver can take it, for instance. Not now. ## What it does not solve - **Hot-unplug and re-enumeration drift.** Still the inventory problem (phase 3, and open question 4 in the track plan). A grant dying with its holder is not the same as a device going away. - **The device table's size.** `maximum_devices` is untouched by this; it goes when the inventory moves in phase 3. - **Two processes racing for the same root grant.** Cannot arise: the roots are minted to whichever task satisfies the chain — the manager's binary under PID 1 — and a second copy spawned by anyone else fails the supervisor half. ## How this is verified The invariant is **I3** from the track plan: a process holds what it was handed and cannot name its way into holding more. The test is adversarial and the suite has never had one of these for devices: a process that was granted nothing calls `device_claim` on a device another driver owns, and on one nobody owns, and is refused both times with its own errno. The audit's lesson was that "the suite contains no attacker"; this is the attacker for devices.