diff --git a/docs/os-development/device-authority.md b/docs/os-development/device-authority.md index 7e49a8f..608be65 100644 --- a/docs/os-development/device-authority.md +++ b/docs/os-development/device-authority.md @@ -1,173 +1,138 @@ -# Device authority: the kernel stops keeping an inventory +# Device authority: you hold what you were given -*Design, drafted 2026-08-07. Not implemented. Prompted by a real machine: an AMD -Ryzen desktop enumerated more PCI functions than the kernel's device table would -hold, and the xHCI and SATA controllers were refused registration — so the -machine booted to the compositor with no USB and no storage.* +*Design for phase 2 of [the bounds track](../bounds-track-plan.md), 2026-08-08. +Supersedes an earlier draft that argued for capabilities on aesthetic grounds; this one +starts from the hole and from the project's principles.* -The kernel keeps a table of every device userspace discovers. It is a fixed -array of 64 descriptors, 344 bytes each, and a second cap allows any one parent -16 children. Neither number is written down anywhere as a decision: -`maximum_devices` has no comment and never reached `parameters.zig`, where every -other tunable in this kernel lives with its reasoning attached. +## The hole -Raising them is not the fix. The numbers are wrong because the *table* is wrong: -it is an inventory of hardware, and an inventory of hardware is not something a -kernel needs. This document proposes replacing it with capabilities, which -removes the ceiling rather than moving it. +`device_claim(id)` checks two things ([devices-broker.zig](../../system/kernel/devices-broker.zig)): -## What the kernel actually uses +```zig +pub fn claim(id: u64, owner: u32) ClaimError!void { + if (id >= count) return error.NoSuchDevice; + if (claimed[@intCast(id)] != null) return error.AlreadyClaimed; + claimed[@intCast(id)] = owner; +} +``` -Every read of a device descriptor from the kernel proper, exhaustively: +Does it exist, and is it free. **Any process may claim any unclaimed device.** -| Used for | What it needs | -|---|---| -| `mmio_map`, `io_read`/`io_write` | the physical range, to check the mapping falls inside it | -| `irq_bind`, `msi_bind` | the GSI, and that the caller owns the device | -| `dma_bind` | that the caller owns the device | -| IOMMU confinement | the **PCI BDF**, to key a domain | -| `device_register` | the parent's ranges, for the containment check | -| the boot display seed | one framebuffer window | +The matching is real but it is entirely advisory: `device-manager` reads `devices.csv`, +matches a device to a driver, and spawns that driver with the device id as `argv[1]` +(`spawnDriver`). The driver parses the string and claims it. Nothing anywhere binds the +manager's decision to the kernel's grant — a process can pass any integer and win the +race. -That is: **physical ranges, interrupt numbers, and one BDF.** Vendor, device and -subsystem ids, class triples, the human-readable names, the parent links, the -bus numbers — the kernel stores all of it and reads none of it. It is held so -that `device_enumerate` can hand it back to user space, which is the whole -mistake in one sentence: the kernel is acting as a distribution mechanism for -data it does not use. +A claim is not a small thing. It is what gates `mmio_map` and `irq_bind`, so it is a +licence to map physical memory and receive interrupts. -## The split +## What this costs, beyond the obvious -Three concerns are tangled in one table. +`maximum_children_per_parent = 16` exists because a driver that claimed one device could +loop `device_register` under it and exhaust the shared table. That threat only exists +*because* claiming is unauthenticated — and the cap is a poor defence against it, since +an attacker can burn 16 slots, claim another device, and burn 16 more. What it reliably +does instead is refuse a legitimate PCI bus with more than 16 functions, which is how an +AMD Ryzen came to boot with no USB and no storage. -**Platform bring-up** — timers, LAPIC/IOAPIC, CPU topology, the ECAM window, the -framebuffer the firmware left. The kernel derives these from ACPI before user -space exists and needs them to function. They never belonged in the device table -and mostly are not (the platform block in `kernel.zig` is separate); this -document does not change them. +So the cap is not merely mis-sized. It is standing in for an authorisation that is not +performed, and it punishes correct behaviour while barely inconveniencing incorrect +behaviour. **Closing the hole is what retires the constant**, not a bigger number. -**Resource authority** — which task may map which physical range, receive which -interrupt, touch which ports. This *must* stay in the kernel. It is the one -grant that cannot be audited after the fact: a process that maps arbitrary -physical memory owns the machine, page tables and IOMMU structures included. -This is memory protection, not device management, and it is why the answer is -not simply "move it all to the device manager". +## What the principles decide -**Device inventory** — what exists, what it is, how it is arranged, which driver -should bind it. This is `device-manager`'s job and is already half there: it -loads `devices.csv`, matches, spawns drivers, and receives `child_added` reports -over its own protocol. The kernel table duplicates what those reports already -carry. +- *Move as much responsibility as possible to user space* (3), and *what remains in the + kernel is there for security or a hardware limitation* (5). -## The proposal: a resource is a capability +Deciding **which** driver gets **which** device is policy: it reads a CSV, matches +identity triples, and picks a binary. That is the device manager's, and it stays there. -Device resources join endpoints, shared memory and DMA regions as a kind in the -handle table. +Enforcing that a driver **holds only what it was given** is security — it is the gate in +front of mapping physical memory. That stays in the kernel, and it is the whole of what +the kernel needs to do. -1. **Roots.** At boot the kernel mints capabilities for the windows it learned - from firmware — the ECAM range, the framebuffer, the legacy port space — and - hands them to the first bus drivers. This is the only place device knowledge - enters the kernel, and it comes from ACPI, not from a driver's say-so. -2. **Subdivision.** A bus driver enumerating hardware derives a narrower - capability from one it holds: `resource_derive(cap, kind, start, len) → cap`. - The kernel checks the sub-range lies inside the capability being subdivided — - the same containment rule as today (`devices-broker.contains`), but checked - against *one capability the caller demonstrably holds* rather than by walking - a global tree. -3. **Delegation.** The driver passes that capability to the child driver over - IPC. Cap-passing already exists; this is the mechanism `subscribe` and - `attach_scanout` already use. -4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and - `dma_bind` take a capability handle instead of `(device_id, resource_index)`. - Possession *is* the authority — there is nothing to look up and no ownership - table to consult. +The kernel therefore does not need to know about matching, `devices.csv`, driver names, +or why a device was assigned. It needs to know that an authority it can verify granted +this device to this task. -Exclusivity stops being a broker refusing a second claimant and becomes the -ordinary property of a capability: only one process was given it. +## The awkward fact that decides the mechanism -## What this buys +Five of the six claimants are device-manager children, spawned with their device id in +`argv[1]`: `pci-bus`, `usb-xhci-bus`, `ps2-bus`, `virtio-gpu`, `acpi`. -**No ceiling.** There is no table to size, so no machine is too big. The -Ryzen's enumeration stops being a limit to tune and becomes what it is — a fact -about a computer. +**`display` is not.** It is spawned by `init` from `init.csv`, and it finds its device by +enumerating the table for a display-class node and claiming whatever it finds +([backend.zig](../../system/services/display/backend.zig)). There is no assignment to +enforce, because nobody assigned it anything. -**Reclamation, free.** `count` in the broker today only ever increases; -`releaseAllOwnedBy` clears a dead driver's *claims* but never its entries. A -driver that crashes and is restarted re-registers its children and consumes the -table again — reachable today without any malice, given the device manager -restarts drivers by design. Capabilities die with the task. +That rules out the cheapest design. "The kernel records the device named at spawn, and +`device_claim` checks it" closes the hole for five claimants and breaks the sixth. And +the sixth is not an oddity to special-case — it is the one that shows the model is +wrong: authority should be *delegable*, not welded to the moment of spawn. -**No quota needed.** The per-parent cap exists to stop one claimant looping -`device_register` and filling the shared table, because a zero-resource child -sidesteps the containment check. With no shared table there is nothing to -exhaust; a process can only ever subdivide what it was given, and its handles -are already bounded per task. +## The design -**A smaller kernel.** Three syscalls leave the ABI, one narrower one arrives, -and `devices-broker.zig` largely disappears along with both constants. +**A device grant is a capability, delegated from a holder.** The mechanism already +exists: `callCap` passes a handle over an IPC call and the kernel installs it in the +receiver's table ([library/kernel/ipc.zig](../../library/kernel/ipc.zig)), which is how +shared memory and DMA regions already move between processes. -**The discipline the rest of the system already uses.** "The claim is the -capability" is written in the driver documentation as if it were already true. -This makes it true. +1. **Root.** At boot the kernel mints grants for the devices firmware discovery found + and hands them to `init` (PID 1, which the kernel spawns and therefore need not + authenticate). This is the only place device authority enters the system, and it + comes from ACPI rather than from anyone's say-so. +2. **Delegation.** `init` passes the device manager the grants it will need — in + practice all of them — and passes `display` the framebuffer grant, because `init` is + what starts `display`. This is the same shape as the `/protocol` registry, where + `init` is already the grantor and `protocol.csv` already records + `/system/services/device-manager, /system/services/init, bind, device-manager`. +3. **Assignment.** The manager passes a driver its device when it spawns it, over the + channel that already exists — the driver `hello`s the manager, and the reply carries + the grant. +4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and `dma_bind` + check possession of the grant instead of consulting an ownership table. -## Syscall surface +Exclusivity stops being a broker refusing a second claimant and becomes the ordinary +property of a capability: only one process was given it. -- `device_enumerate` — **retires.** It exists only to read the kernel's table. - Callers ask `device-manager`, whose protocol already reserves an `enumerate` - verb. Note this is a public-ABI change: `vdso.md` documents it. -- `device_register` — **splits.** The kernel half becomes `resource_derive`; the - publication half ("this device exists, here is what it is") becomes an IPC - message to `device-manager`, which is where the inventory belongs and where - `child_added` already carries the same facts. -- `device_claim` — **dissolves into possession**, except for the IOMMU (below). -- `mmio_map`, `irq_bind`, `msi_bind`, `io_read`, `io_write`, `dma_bind` — keep - their names and semantics; their first argument becomes a capability handle. +**`maximum_children_per_parent` is deleted here.** After this a bus driver's children are +devices it enumerated on a bus it was actually given, and the threat the cap was written +for no longer exists. -## The open question: where the IOMMU attaches +## What this costs -This is the one place the kernel still needs device *identity* rather than a -range. `confineDevice(device_id, bdf, owner)` builds a domain keyed by PCI BDF, -attaches it, and the claim is rolled back if confinement fails — deliberately, -so a device that cannot be confined is never driven. +**Bring-up order changes.** `pci-bus` today claims first and says hello afterwards — its +own comment says "Claim the bridge, map the ECAM, hello the manager, then scan." Under +delegation the hello must come first, because that is where the grant arrives. Five +drivers need that reordering, and it is the bulk of the work. -Three options, none obviously right: +**`init` grows a device role.** It already registers `/protocol` and reads +`protocol.csv`; it would also hold root device grants and hand them on. That is more +responsibility in PID 1, which is a cost worth naming — though the alternative is the +kernel deciding who may hold what, which principle 5 excludes. -1. **The BDF rides the capability.** A memory capability derived for a PCI - function carries its BDF, and the kernel confines on first `mmio_map` or - `dma_bind`. Keeps the syscall count down; means a capability is no longer - purely a range. -2. **An explicit `device_attach(cap, bdf)`.** Honest and visible, but it puts a - BDF — a fact about PCI — into a kernel interface that otherwise knows nothing - about buses, and something must stop a caller naming a BDF that is not - theirs. -3. **The root PCI capability carries the segment, and derivation computes the - BDF.** Purest, but only works for PCI and the kernel would be parsing bus - topology, which is precisely what this document is trying to stop. +**A configuration question follows.** `/protocol` grants are data (`protocol.csv`), +because who may speak to whom is policy. Device grants could be too — a manifest saying +which binary may be given which device class. That would be the natural symmetry, and it +is *not* proposed here: `init` handing the manager everything and the manager matching +by `devices.csv` is the smaller step, and the manifest can follow if it earns itself. -My inclination is (1), because confinement is a property of the resource being -granted rather than a separate action, and because it keeps the "possession is -authority" story intact. It needs the derivation call to know it is carving a -PCI function, which is a wart worth arguing about. +## What it does not solve -## What this does not solve +- **Hot-unplug and re-enumeration drift.** Still the inventory problem (phase 3, and + open question 4 in the track plan). A grant dying with its holder is not the same as a + device going away. +- **The device table's size.** `maximum_devices` is untouched by this; it goes when the + inventory moves in phase 3. +- **Two processes racing for the same root grant.** Cannot arise, because roots are + minted to `init` alone. -- **Hot-plug and removal.** Capabilities die with their holder, but a device - that physically disappears while a driver lives is the device manager's - problem and unchanged by this. -- **Quotas on physical memory.** Nothing here bounds how much a driver maps; it - bounds only *what* it may map. That was already true. -- **The device tree as a published thing.** `/system/devices` remains a device - manager concern, as the file-system hierarchy already assumes. +## How this is verified -## Migration sketch - -Not a plan yet — the phases need sizing once the IOMMU question is settled. -Rough shape: introduce the capability kind and `resource_derive` alongside the -existing table; convert the resource-consuming syscalls to accept either form; -move the inventory into `device-manager` and convert its clients off -`device_enumerate`; then delete the table, the two constants and the three -syscalls in one flag-day, as the `ServiceId` retirement did. - -Every driver is affected, so the QEMU suite is the arbiter at each step, and the -Ryzen is the acceptance test — it is the machine that found this, and the one -that proves it fixed. +The invariant is **I3** from the track plan: a process holds what it was handed and +cannot name its way into holding more. The test is adversarial and the suite has never +had one of these for devices: a process that was granted nothing calls `device_claim` +on a device another driver owns, and on one nobody owns, and is refused both times with +its own errno. The audit's lesson was that "the suite contains no attacker"; this is the +attacker for devices.