# Device authority: the kernel stops keeping an inventory *Design, drafted 2026-08-07. Not implemented. Prompted by a real machine: an AMD Ryzen desktop enumerated more PCI functions than the kernel's device table would hold, and the xHCI and SATA controllers were refused registration — so the machine booted to the compositor with no USB and no storage.* The kernel keeps a table of every device userspace discovers. It is a fixed array of 64 descriptors, 344 bytes each, and a second cap allows any one parent 16 children. Neither number is written down anywhere as a decision: `maximum_devices` has no comment and never reached `parameters.zig`, where every other tunable in this kernel lives with its reasoning attached. Raising them is not the fix. The numbers are wrong because the *table* is wrong: it is an inventory of hardware, and an inventory of hardware is not something a kernel needs. This document proposes replacing it with capabilities, which removes the ceiling rather than moving it. ## What the kernel actually uses Every read of a device descriptor from the kernel proper, exhaustively: | Used for | What it needs | |---|---| | `mmio_map`, `io_read`/`io_write` | the physical range, to check the mapping falls inside it | | `irq_bind`, `msi_bind` | the GSI, and that the caller owns the device | | `dma_bind` | that the caller owns the device | | IOMMU confinement | the **PCI BDF**, to key a domain | | `device_register` | the parent's ranges, for the containment check | | the boot display seed | one framebuffer window | That is: **physical ranges, interrupt numbers, and one BDF.** Vendor, device and subsystem ids, class triples, the human-readable names, the parent links, the bus numbers — the kernel stores all of it and reads none of it. It is held so that `device_enumerate` can hand it back to user space, which is the whole mistake in one sentence: the kernel is acting as a distribution mechanism for data it does not use. ## The split Three concerns are tangled in one table. **Platform bring-up** — timers, LAPIC/IOAPIC, CPU topology, the ECAM window, the framebuffer the firmware left. The kernel derives these from ACPI before user space exists and needs them to function. They never belonged in the device table and mostly are not (the platform block in `kernel.zig` is separate); this document does not change them. **Resource authority** — which task may map which physical range, receive which interrupt, touch which ports. This *must* stay in the kernel. It is the one grant that cannot be audited after the fact: a process that maps arbitrary physical memory owns the machine, page tables and IOMMU structures included. This is memory protection, not device management, and it is why the answer is not simply "move it all to the device manager". **Device inventory** — what exists, what it is, how it is arranged, which driver should bind it. This is `device-manager`'s job and is already half there: it loads `devices.csv`, matches, spawns drivers, and receives `child_added` reports over its own protocol. The kernel table duplicates what those reports already carry. ## The proposal: a resource is a capability Device resources join endpoints, shared memory and DMA regions as a kind in the handle table. 1. **Roots.** At boot the kernel mints capabilities for the windows it learned from firmware — the ECAM range, the framebuffer, the legacy port space — and hands them to the first bus drivers. This is the only place device knowledge enters the kernel, and it comes from ACPI, not from a driver's say-so. 2. **Subdivision.** A bus driver enumerating hardware derives a narrower capability from one it holds: `resource_derive(cap, kind, start, len) → cap`. The kernel checks the sub-range lies inside the capability being subdivided — the same containment rule as today (`devices-broker.contains`), but checked against *one capability the caller demonstrably holds* rather than by walking a global tree. 3. **Delegation.** The driver passes that capability to the child driver over IPC. Cap-passing already exists; this is the mechanism `subscribe` and `attach_scanout` already use. 4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and `dma_bind` take a capability handle instead of `(device_id, resource_index)`. Possession *is* the authority — there is nothing to look up and no ownership table to consult. Exclusivity stops being a broker refusing a second claimant and becomes the ordinary property of a capability: only one process was given it. ## What this buys **No ceiling.** There is no table to size, so no machine is too big. The Ryzen's enumeration stops being a limit to tune and becomes what it is — a fact about a computer. **Reclamation, free.** `count` in the broker today only ever increases; `releaseAllOwnedBy` clears a dead driver's *claims* but never its entries. A driver that crashes and is restarted re-registers its children and consumes the table again — reachable today without any malice, given the device manager restarts drivers by design. Capabilities die with the task. **No quota needed.** The per-parent cap exists to stop one claimant looping `device_register` and filling the shared table, because a zero-resource child sidesteps the containment check. With no shared table there is nothing to exhaust; a process can only ever subdivide what it was given, and its handles are already bounded per task. **A smaller kernel.** Three syscalls leave the ABI, one narrower one arrives, and `devices-broker.zig` largely disappears along with both constants. **The discipline the rest of the system already uses.** "The claim is the capability" is written in the driver documentation as if it were already true. This makes it true. ## Syscall surface - `device_enumerate` — **retires.** It exists only to read the kernel's table. Callers ask `device-manager`, whose protocol already reserves an `enumerate` verb. Note this is a public-ABI change: `vdso.md` documents it. - `device_register` — **splits.** The kernel half becomes `resource_derive`; the publication half ("this device exists, here is what it is") becomes an IPC message to `device-manager`, which is where the inventory belongs and where `child_added` already carries the same facts. - `device_claim` — **dissolves into possession**, except for the IOMMU (below). - `mmio_map`, `irq_bind`, `msi_bind`, `io_read`, `io_write`, `dma_bind` — keep their names and semantics; their first argument becomes a capability handle. ## The open question: where the IOMMU attaches This is the one place the kernel still needs device *identity* rather than a range. `confineDevice(device_id, bdf, owner)` builds a domain keyed by PCI BDF, attaches it, and the claim is rolled back if confinement fails — deliberately, so a device that cannot be confined is never driven. Three options, none obviously right: 1. **The BDF rides the capability.** A memory capability derived for a PCI function carries its BDF, and the kernel confines on first `mmio_map` or `dma_bind`. Keeps the syscall count down; means a capability is no longer purely a range. 2. **An explicit `device_attach(cap, bdf)`.** Honest and visible, but it puts a BDF — a fact about PCI — into a kernel interface that otherwise knows nothing about buses, and something must stop a caller naming a BDF that is not theirs. 3. **The root PCI capability carries the segment, and derivation computes the BDF.** Purest, but only works for PCI and the kernel would be parsing bus topology, which is precisely what this document is trying to stop. My inclination is (1), because confinement is a property of the resource being granted rather than a separate action, and because it keeps the "possession is authority" story intact. It needs the derivation call to know it is carving a PCI function, which is a wart worth arguing about. ## What this does not solve - **Hot-plug and removal.** Capabilities die with their holder, but a device that physically disappears while a driver lives is the device manager's problem and unchanged by this. - **Quotas on physical memory.** Nothing here bounds how much a driver maps; it bounds only *what* it may map. That was already true. - **The device tree as a published thing.** `/system/devices` remains a device manager concern, as the file-system hierarchy already assumes. ## Migration sketch Not a plan yet — the phases need sizing once the IOMMU question is settled. Rough shape: introduce the capability kind and `resource_derive` alongside the existing table; convert the resource-consuming syscalls to accept either form; move the inventory into `device-manager` and convert its clients off `device_enumerate`; then delete the table, the two constants and the three syscalls in one flag-day, as the `ServiceId` retirement did. Every driver is affected, so the QEMU suite is the arbiter at each step, and the Ryzen is the acceptance test — it is the machine that found this, and the one that proves it fixed.