From 3b24b541b09f570e7c5809f8b82c683f6ff70f30 Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sat, 8 Aug 2026 12:58:19 +0100 Subject: [PATCH] =?UTF-8?q?docs:=20phase=202=20design=20=E2=80=94=20you=20?= =?UTF-8?q?hold=20what=20you=20were=20given?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit device_claim checks that a device exists and is free. That is all. Any process may claim any unclaimed device, and a claim is what gates mmio_map and irq_bind — a licence to map physical memory and take interrupts. The device manager's matching is real but advisory: it spawns a driver with the device id in argv[1] and nothing binds that decision to the kernel's grant. maximum_children_per_parent is the visible cost. It exists because a driver that claimed one device could loop device_register under it, and it is a poor defence — an attacker burns 16 slots, claims another device, burns 16 more — while reliably refusing a legitimate PCI bus with more than 16 functions. Closing the hole is what retires the constant. The principles decide the split: matching is policy and stays in the device manager; enforcing that a driver holds only what it was given is security and stays in the kernel, and is the whole of what the kernel needs. The mechanism is decided by an awkward fact. Five of six claimants are device-manager children spawned with their device id. display is not — init spawns it, and it finds its framebuffer by enumerating for a display-class node and claiming whatever it finds. So "record the device named at spawn" closes the hole for five and breaks the sixth, and the sixth is not an oddity to special-case: it shows authority must be delegable rather than welded to the moment of spawn. So: a device grant is a capability, minted by the kernel to init for the devices firmware discovery found, delegated by init to the device manager and to display, and passed by the manager to each driver it spawns. The cap-passing path already exists and already carries shared memory and DMA regions. init is already the grantor for /protocol, and protocol.csv already records init granting the device manager its binding. Costs named rather than buried: five drivers must hello before claiming (pci-bus claims first today), and init grows a device role on top of the protocol registry. A device-grant manifest mirroring protocol.csv would be the natural symmetry and is deliberately not proposed yet. --- docs/os-development/device-authority.md | 247 ++++++++++-------------- 1 file changed, 106 insertions(+), 141 deletions(-) diff --git a/docs/os-development/device-authority.md b/docs/os-development/device-authority.md index 7e49a8f..608be65 100644 --- a/docs/os-development/device-authority.md +++ b/docs/os-development/device-authority.md @@ -1,173 +1,138 @@ -# Device authority: the kernel stops keeping an inventory +# Device authority: you hold what you were given -*Design, drafted 2026-08-07. Not implemented. Prompted by a real machine: an AMD -Ryzen desktop enumerated more PCI functions than the kernel's device table would -hold, and the xHCI and SATA controllers were refused registration — so the -machine booted to the compositor with no USB and no storage.* +*Design for phase 2 of [the bounds track](../bounds-track-plan.md), 2026-08-08. +Supersedes an earlier draft that argued for capabilities on aesthetic grounds; this one +starts from the hole and from the project's principles.* -The kernel keeps a table of every device userspace discovers. It is a fixed -array of 64 descriptors, 344 bytes each, and a second cap allows any one parent -16 children. Neither number is written down anywhere as a decision: -`maximum_devices` has no comment and never reached `parameters.zig`, where every -other tunable in this kernel lives with its reasoning attached. +## The hole -Raising them is not the fix. The numbers are wrong because the *table* is wrong: -it is an inventory of hardware, and an inventory of hardware is not something a -kernel needs. This document proposes replacing it with capabilities, which -removes the ceiling rather than moving it. +`device_claim(id)` checks two things ([devices-broker.zig](../../system/kernel/devices-broker.zig)): -## What the kernel actually uses +```zig +pub fn claim(id: u64, owner: u32) ClaimError!void { + if (id >= count) return error.NoSuchDevice; + if (claimed[@intCast(id)] != null) return error.AlreadyClaimed; + claimed[@intCast(id)] = owner; +} +``` -Every read of a device descriptor from the kernel proper, exhaustively: +Does it exist, and is it free. **Any process may claim any unclaimed device.** -| Used for | What it needs | -|---|---| -| `mmio_map`, `io_read`/`io_write` | the physical range, to check the mapping falls inside it | -| `irq_bind`, `msi_bind` | the GSI, and that the caller owns the device | -| `dma_bind` | that the caller owns the device | -| IOMMU confinement | the **PCI BDF**, to key a domain | -| `device_register` | the parent's ranges, for the containment check | -| the boot display seed | one framebuffer window | +The matching is real but it is entirely advisory: `device-manager` reads `devices.csv`, +matches a device to a driver, and spawns that driver with the device id as `argv[1]` +(`spawnDriver`). The driver parses the string and claims it. Nothing anywhere binds the +manager's decision to the kernel's grant — a process can pass any integer and win the +race. -That is: **physical ranges, interrupt numbers, and one BDF.** Vendor, device and -subsystem ids, class triples, the human-readable names, the parent links, the -bus numbers — the kernel stores all of it and reads none of it. It is held so -that `device_enumerate` can hand it back to user space, which is the whole -mistake in one sentence: the kernel is acting as a distribution mechanism for -data it does not use. +A claim is not a small thing. It is what gates `mmio_map` and `irq_bind`, so it is a +licence to map physical memory and receive interrupts. -## The split +## What this costs, beyond the obvious -Three concerns are tangled in one table. +`maximum_children_per_parent = 16` exists because a driver that claimed one device could +loop `device_register` under it and exhaust the shared table. That threat only exists +*because* claiming is unauthenticated — and the cap is a poor defence against it, since +an attacker can burn 16 slots, claim another device, and burn 16 more. What it reliably +does instead is refuse a legitimate PCI bus with more than 16 functions, which is how an +AMD Ryzen came to boot with no USB and no storage. -**Platform bring-up** — timers, LAPIC/IOAPIC, CPU topology, the ECAM window, the -framebuffer the firmware left. The kernel derives these from ACPI before user -space exists and needs them to function. They never belonged in the device table -and mostly are not (the platform block in `kernel.zig` is separate); this -document does not change them. +So the cap is not merely mis-sized. It is standing in for an authorisation that is not +performed, and it punishes correct behaviour while barely inconveniencing incorrect +behaviour. **Closing the hole is what retires the constant**, not a bigger number. -**Resource authority** — which task may map which physical range, receive which -interrupt, touch which ports. This *must* stay in the kernel. It is the one -grant that cannot be audited after the fact: a process that maps arbitrary -physical memory owns the machine, page tables and IOMMU structures included. -This is memory protection, not device management, and it is why the answer is -not simply "move it all to the device manager". +## What the principles decide -**Device inventory** — what exists, what it is, how it is arranged, which driver -should bind it. This is `device-manager`'s job and is already half there: it -loads `devices.csv`, matches, spawns drivers, and receives `child_added` reports -over its own protocol. The kernel table duplicates what those reports already -carry. +- *Move as much responsibility as possible to user space* (3), and *what remains in the + kernel is there for security or a hardware limitation* (5). -## The proposal: a resource is a capability +Deciding **which** driver gets **which** device is policy: it reads a CSV, matches +identity triples, and picks a binary. That is the device manager's, and it stays there. -Device resources join endpoints, shared memory and DMA regions as a kind in the -handle table. +Enforcing that a driver **holds only what it was given** is security — it is the gate in +front of mapping physical memory. That stays in the kernel, and it is the whole of what +the kernel needs to do. -1. **Roots.** At boot the kernel mints capabilities for the windows it learned - from firmware — the ECAM range, the framebuffer, the legacy port space — and - hands them to the first bus drivers. This is the only place device knowledge - enters the kernel, and it comes from ACPI, not from a driver's say-so. -2. **Subdivision.** A bus driver enumerating hardware derives a narrower - capability from one it holds: `resource_derive(cap, kind, start, len) → cap`. - The kernel checks the sub-range lies inside the capability being subdivided — - the same containment rule as today (`devices-broker.contains`), but checked - against *one capability the caller demonstrably holds* rather than by walking - a global tree. -3. **Delegation.** The driver passes that capability to the child driver over - IPC. Cap-passing already exists; this is the mechanism `subscribe` and - `attach_scanout` already use. -4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and - `dma_bind` take a capability handle instead of `(device_id, resource_index)`. - Possession *is* the authority — there is nothing to look up and no ownership - table to consult. +The kernel therefore does not need to know about matching, `devices.csv`, driver names, +or why a device was assigned. It needs to know that an authority it can verify granted +this device to this task. -Exclusivity stops being a broker refusing a second claimant and becomes the -ordinary property of a capability: only one process was given it. +## The awkward fact that decides the mechanism -## What this buys +Five of the six claimants are device-manager children, spawned with their device id in +`argv[1]`: `pci-bus`, `usb-xhci-bus`, `ps2-bus`, `virtio-gpu`, `acpi`. -**No ceiling.** There is no table to size, so no machine is too big. The -Ryzen's enumeration stops being a limit to tune and becomes what it is — a fact -about a computer. +**`display` is not.** It is spawned by `init` from `init.csv`, and it finds its device by +enumerating the table for a display-class node and claiming whatever it finds +([backend.zig](../../system/services/display/backend.zig)). There is no assignment to +enforce, because nobody assigned it anything. -**Reclamation, free.** `count` in the broker today only ever increases; -`releaseAllOwnedBy` clears a dead driver's *claims* but never its entries. A -driver that crashes and is restarted re-registers its children and consumes the -table again — reachable today without any malice, given the device manager -restarts drivers by design. Capabilities die with the task. +That rules out the cheapest design. "The kernel records the device named at spawn, and +`device_claim` checks it" closes the hole for five claimants and breaks the sixth. And +the sixth is not an oddity to special-case — it is the one that shows the model is +wrong: authority should be *delegable*, not welded to the moment of spawn. -**No quota needed.** The per-parent cap exists to stop one claimant looping -`device_register` and filling the shared table, because a zero-resource child -sidesteps the containment check. With no shared table there is nothing to -exhaust; a process can only ever subdivide what it was given, and its handles -are already bounded per task. +## The design -**A smaller kernel.** Three syscalls leave the ABI, one narrower one arrives, -and `devices-broker.zig` largely disappears along with both constants. +**A device grant is a capability, delegated from a holder.** The mechanism already +exists: `callCap` passes a handle over an IPC call and the kernel installs it in the +receiver's table ([library/kernel/ipc.zig](../../library/kernel/ipc.zig)), which is how +shared memory and DMA regions already move between processes. -**The discipline the rest of the system already uses.** "The claim is the -capability" is written in the driver documentation as if it were already true. -This makes it true. +1. **Root.** At boot the kernel mints grants for the devices firmware discovery found + and hands them to `init` (PID 1, which the kernel spawns and therefore need not + authenticate). This is the only place device authority enters the system, and it + comes from ACPI rather than from anyone's say-so. +2. **Delegation.** `init` passes the device manager the grants it will need — in + practice all of them — and passes `display` the framebuffer grant, because `init` is + what starts `display`. This is the same shape as the `/protocol` registry, where + `init` is already the grantor and `protocol.csv` already records + `/system/services/device-manager, /system/services/init, bind, device-manager`. +3. **Assignment.** The manager passes a driver its device when it spawns it, over the + channel that already exists — the driver `hello`s the manager, and the reply carries + the grant. +4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and `dma_bind` + check possession of the grant instead of consulting an ownership table. -## Syscall surface +Exclusivity stops being a broker refusing a second claimant and becomes the ordinary +property of a capability: only one process was given it. -- `device_enumerate` — **retires.** It exists only to read the kernel's table. - Callers ask `device-manager`, whose protocol already reserves an `enumerate` - verb. Note this is a public-ABI change: `vdso.md` documents it. -- `device_register` — **splits.** The kernel half becomes `resource_derive`; the - publication half ("this device exists, here is what it is") becomes an IPC - message to `device-manager`, which is where the inventory belongs and where - `child_added` already carries the same facts. -- `device_claim` — **dissolves into possession**, except for the IOMMU (below). -- `mmio_map`, `irq_bind`, `msi_bind`, `io_read`, `io_write`, `dma_bind` — keep - their names and semantics; their first argument becomes a capability handle. +**`maximum_children_per_parent` is deleted here.** After this a bus driver's children are +devices it enumerated on a bus it was actually given, and the threat the cap was written +for no longer exists. -## The open question: where the IOMMU attaches +## What this costs -This is the one place the kernel still needs device *identity* rather than a -range. `confineDevice(device_id, bdf, owner)` builds a domain keyed by PCI BDF, -attaches it, and the claim is rolled back if confinement fails — deliberately, -so a device that cannot be confined is never driven. +**Bring-up order changes.** `pci-bus` today claims first and says hello afterwards — its +own comment says "Claim the bridge, map the ECAM, hello the manager, then scan." Under +delegation the hello must come first, because that is where the grant arrives. Five +drivers need that reordering, and it is the bulk of the work. -Three options, none obviously right: +**`init` grows a device role.** It already registers `/protocol` and reads +`protocol.csv`; it would also hold root device grants and hand them on. That is more +responsibility in PID 1, which is a cost worth naming — though the alternative is the +kernel deciding who may hold what, which principle 5 excludes. -1. **The BDF rides the capability.** A memory capability derived for a PCI - function carries its BDF, and the kernel confines on first `mmio_map` or - `dma_bind`. Keeps the syscall count down; means a capability is no longer - purely a range. -2. **An explicit `device_attach(cap, bdf)`.** Honest and visible, but it puts a - BDF — a fact about PCI — into a kernel interface that otherwise knows nothing - about buses, and something must stop a caller naming a BDF that is not - theirs. -3. **The root PCI capability carries the segment, and derivation computes the - BDF.** Purest, but only works for PCI and the kernel would be parsing bus - topology, which is precisely what this document is trying to stop. +**A configuration question follows.** `/protocol` grants are data (`protocol.csv`), +because who may speak to whom is policy. Device grants could be too — a manifest saying +which binary may be given which device class. That would be the natural symmetry, and it +is *not* proposed here: `init` handing the manager everything and the manager matching +by `devices.csv` is the smaller step, and the manifest can follow if it earns itself. -My inclination is (1), because confinement is a property of the resource being -granted rather than a separate action, and because it keeps the "possession is -authority" story intact. It needs the derivation call to know it is carving a -PCI function, which is a wart worth arguing about. +## What it does not solve -## What this does not solve +- **Hot-unplug and re-enumeration drift.** Still the inventory problem (phase 3, and + open question 4 in the track plan). A grant dying with its holder is not the same as a + device going away. +- **The device table's size.** `maximum_devices` is untouched by this; it goes when the + inventory moves in phase 3. +- **Two processes racing for the same root grant.** Cannot arise, because roots are + minted to `init` alone. -- **Hot-plug and removal.** Capabilities die with their holder, but a device - that physically disappears while a driver lives is the device manager's - problem and unchanged by this. -- **Quotas on physical memory.** Nothing here bounds how much a driver maps; it - bounds only *what* it may map. That was already true. -- **The device tree as a published thing.** `/system/devices` remains a device - manager concern, as the file-system hierarchy already assumes. +## How this is verified -## Migration sketch - -Not a plan yet — the phases need sizing once the IOMMU question is settled. -Rough shape: introduce the capability kind and `resource_derive` alongside the -existing table; convert the resource-consuming syscalls to accept either form; -move the inventory into `device-manager` and convert its clients off -`device_enumerate`; then delete the table, the two constants and the three -syscalls in one flag-day, as the `ServiceId` retirement did. - -Every driver is affected, so the QEMU suite is the arbiter at each step, and the -Ryzen is the acceptance test — it is the machine that found this, and the one -that proves it fixed. +The invariant is **I3** from the track plan: a process holds what it was handed and +cannot name its way into holding more. The test is adversarial and the suite has never +had one of these for devices: a process that was granted nothing calls `device_claim` +on a device another driver owns, and on one nobody owns, and is refused both times with +its own errno. The audit's lesson was that "the suite contains no attacker"; this is the +attacker for devices.