docs: phase 2 design — you hold what you were given

device_claim checks that a device exists and is free. That is all. Any
process may claim any unclaimed device, and a claim is what gates mmio_map
and irq_bind — a licence to map physical memory and take interrupts. The
device manager's matching is real but advisory: it spawns a driver with the
device id in argv[1] and nothing binds that decision to the kernel's grant.

maximum_children_per_parent is the visible cost. It exists because a driver
that claimed one device could loop device_register under it, and it is a
poor defence — an attacker burns 16 slots, claims another device, burns 16
more — while reliably refusing a legitimate PCI bus with more than 16
functions. Closing the hole is what retires the constant.

The principles decide the split: matching is policy and stays in the device
manager; enforcing that a driver holds only what it was given is security
and stays in the kernel, and is the whole of what the kernel needs.

The mechanism is decided by an awkward fact. Five of six claimants are
device-manager children spawned with their device id. display is not — init
spawns it, and it finds its framebuffer by enumerating for a display-class
node and claiming whatever it finds. So "record the device named at spawn"
closes the hole for five and breaks the sixth, and the sixth is not an
oddity to special-case: it shows authority must be delegable rather than
welded to the moment of spawn.

So: a device grant is a capability, minted by the kernel to init for the
devices firmware discovery found, delegated by init to the device manager
and to display, and passed by the manager to each driver it spawns. The
cap-passing path already exists and already carries shared memory and DMA
regions. init is already the grantor for /protocol, and protocol.csv already
records init granting the device manager its binding.

Costs named rather than buried: five drivers must hello before claiming
(pci-bus claims first today), and init grows a device role on top of the
protocol registry. A device-grant manifest mirroring protocol.csv would be
the natural symmetry and is deliberately not proposed yet.
This commit is contained in:
Daniel Samson
2026-08-08 12:58:19 +01:00
parent 5d55217212
commit 3b24b541b0
+106 -141
View File
@@ -1,173 +1,138 @@
# Device authority: the kernel stops keeping an inventory
# Device authority: you hold what you were given
*Design, drafted 2026-08-07. Not implemented. Prompted by a real machine: an AMD
Ryzen desktop enumerated more PCI functions than the kernel's device table would
hold, and the xHCI and SATA controllers were refused registration — so the
machine booted to the compositor with no USB and no storage.*
*Design for phase 2 of [the bounds track](../bounds-track-plan.md), 2026-08-08.
Supersedes an earlier draft that argued for capabilities on aesthetic grounds; this one
starts from the hole and from the project's principles.*
The kernel keeps a table of every device userspace discovers. It is a fixed
array of 64 descriptors, 344 bytes each, and a second cap allows any one parent
16 children. Neither number is written down anywhere as a decision:
`maximum_devices` has no comment and never reached `parameters.zig`, where every
other tunable in this kernel lives with its reasoning attached.
## The hole
Raising them is not the fix. The numbers are wrong because the *table* is wrong:
it is an inventory of hardware, and an inventory of hardware is not something a
kernel needs. This document proposes replacing it with capabilities, which
removes the ceiling rather than moving it.
`device_claim(id)` checks two things ([devices-broker.zig](../../system/kernel/devices-broker.zig)):
## What the kernel actually uses
```zig
pub fn claim(id: u64, owner: u32) ClaimError!void {
if (id >= count) return error.NoSuchDevice;
if (claimed[@intCast(id)] != null) return error.AlreadyClaimed;
claimed[@intCast(id)] = owner;
}
```
Every read of a device descriptor from the kernel proper, exhaustively:
Does it exist, and is it free. **Any process may claim any unclaimed device.**
| Used for | What it needs |
|---|---|
| `mmio_map`, `io_read`/`io_write` | the physical range, to check the mapping falls inside it |
| `irq_bind`, `msi_bind` | the GSI, and that the caller owns the device |
| `dma_bind` | that the caller owns the device |
| IOMMU confinement | the **PCI BDF**, to key a domain |
| `device_register` | the parent's ranges, for the containment check |
| the boot display seed | one framebuffer window |
The matching is real but it is entirely advisory: `device-manager` reads `devices.csv`,
matches a device to a driver, and spawns that driver with the device id as `argv[1]`
(`spawnDriver`). The driver parses the string and claims it. Nothing anywhere binds the
manager's decision to the kernel's grant — a process can pass any integer and win the
race.
That is: **physical ranges, interrupt numbers, and one BDF.** Vendor, device and
subsystem ids, class triples, the human-readable names, the parent links, the
bus numbers — the kernel stores all of it and reads none of it. It is held so
that `device_enumerate` can hand it back to user space, which is the whole
mistake in one sentence: the kernel is acting as a distribution mechanism for
data it does not use.
A claim is not a small thing. It is what gates `mmio_map` and `irq_bind`, so it is a
licence to map physical memory and receive interrupts.
## The split
## What this costs, beyond the obvious
Three concerns are tangled in one table.
`maximum_children_per_parent = 16` exists because a driver that claimed one device could
loop `device_register` under it and exhaust the shared table. That threat only exists
*because* claiming is unauthenticated — and the cap is a poor defence against it, since
an attacker can burn 16 slots, claim another device, and burn 16 more. What it reliably
does instead is refuse a legitimate PCI bus with more than 16 functions, which is how an
AMD Ryzen came to boot with no USB and no storage.
**Platform bring-up** — timers, LAPIC/IOAPIC, CPU topology, the ECAM window, the
framebuffer the firmware left. The kernel derives these from ACPI before user
space exists and needs them to function. They never belonged in the device table
and mostly are not (the platform block in `kernel.zig` is separate); this
document does not change them.
So the cap is not merely mis-sized. It is standing in for an authorisation that is not
performed, and it punishes correct behaviour while barely inconveniencing incorrect
behaviour. **Closing the hole is what retires the constant**, not a bigger number.
**Resource authority** — which task may map which physical range, receive which
interrupt, touch which ports. This *must* stay in the kernel. It is the one
grant that cannot be audited after the fact: a process that maps arbitrary
physical memory owns the machine, page tables and IOMMU structures included.
This is memory protection, not device management, and it is why the answer is
not simply "move it all to the device manager".
## What the principles decide
**Device inventory** — what exists, what it is, how it is arranged, which driver
should bind it. This is `device-manager`'s job and is already half there: it
loads `devices.csv`, matches, spawns drivers, and receives `child_added` reports
over its own protocol. The kernel table duplicates what those reports already
carry.
- *Move as much responsibility as possible to user space* (3), and *what remains in the
kernel is there for security or a hardware limitation* (5).
## The proposal: a resource is a capability
Deciding **which** driver gets **which** device is policy: it reads a CSV, matches
identity triples, and picks a binary. That is the device manager's, and it stays there.
Device resources join endpoints, shared memory and DMA regions as a kind in the
handle table.
Enforcing that a driver **holds only what it was given** is security — it is the gate in
front of mapping physical memory. That stays in the kernel, and it is the whole of what
the kernel needs to do.
1. **Roots.** At boot the kernel mints capabilities for the windows it learned
from firmware — the ECAM range, the framebuffer, the legacy port space — and
hands them to the first bus drivers. This is the only place device knowledge
enters the kernel, and it comes from ACPI, not from a driver's say-so.
2. **Subdivision.** A bus driver enumerating hardware derives a narrower
capability from one it holds: `resource_derive(cap, kind, start, len) → cap`.
The kernel checks the sub-range lies inside the capability being subdivided —
the same containment rule as today (`devices-broker.contains`), but checked
against *one capability the caller demonstrably holds* rather than by walking
a global tree.
3. **Delegation.** The driver passes that capability to the child driver over
IPC. Cap-passing already exists; this is the mechanism `subscribe` and
`attach_scanout` already use.
4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and
`dma_bind` take a capability handle instead of `(device_id, resource_index)`.
Possession *is* the authority — there is nothing to look up and no ownership
table to consult.
The kernel therefore does not need to know about matching, `devices.csv`, driver names,
or why a device was assigned. It needs to know that an authority it can verify granted
this device to this task.
Exclusivity stops being a broker refusing a second claimant and becomes the
ordinary property of a capability: only one process was given it.
## The awkward fact that decides the mechanism
## What this buys
Five of the six claimants are device-manager children, spawned with their device id in
`argv[1]`: `pci-bus`, `usb-xhci-bus`, `ps2-bus`, `virtio-gpu`, `acpi`.
**No ceiling.** There is no table to size, so no machine is too big. The
Ryzen's enumeration stops being a limit to tune and becomes what it is — a fact
about a computer.
**`display` is not.** It is spawned by `init` from `init.csv`, and it finds its device by
enumerating the table for a display-class node and claiming whatever it finds
([backend.zig](../../system/services/display/backend.zig)). There is no assignment to
enforce, because nobody assigned it anything.
**Reclamation, free.** `count` in the broker today only ever increases;
`releaseAllOwnedBy` clears a dead driver's *claims* but never its entries. A
driver that crashes and is restarted re-registers its children and consumes the
table again — reachable today without any malice, given the device manager
restarts drivers by design. Capabilities die with the task.
That rules out the cheapest design. "The kernel records the device named at spawn, and
`device_claim` checks it" closes the hole for five claimants and breaks the sixth. And
the sixth is not an oddity to special-case — it is the one that shows the model is
wrong: authority should be *delegable*, not welded to the moment of spawn.
**No quota needed.** The per-parent cap exists to stop one claimant looping
`device_register` and filling the shared table, because a zero-resource child
sidesteps the containment check. With no shared table there is nothing to
exhaust; a process can only ever subdivide what it was given, and its handles
are already bounded per task.
## The design
**A smaller kernel.** Three syscalls leave the ABI, one narrower one arrives,
and `devices-broker.zig` largely disappears along with both constants.
**A device grant is a capability, delegated from a holder.** The mechanism already
exists: `callCap` passes a handle over an IPC call and the kernel installs it in the
receiver's table ([library/kernel/ipc.zig](../../library/kernel/ipc.zig)), which is how
shared memory and DMA regions already move between processes.
**The discipline the rest of the system already uses.** "The claim is the
capability" is written in the driver documentation as if it were already true.
This makes it true.
1. **Root.** At boot the kernel mints grants for the devices firmware discovery found
and hands them to `init` (PID 1, which the kernel spawns and therefore need not
authenticate). This is the only place device authority enters the system, and it
comes from ACPI rather than from anyone's say-so.
2. **Delegation.** `init` passes the device manager the grants it will need — in
practice all of them — and passes `display` the framebuffer grant, because `init` is
what starts `display`. This is the same shape as the `/protocol` registry, where
`init` is already the grantor and `protocol.csv` already records
`/system/services/device-manager, /system/services/init, bind, device-manager`.
3. **Assignment.** The manager passes a driver its device when it spawns it, over the
channel that already exists — the driver `hello`s the manager, and the reply carries
the grant.
4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and `dma_bind`
check possession of the grant instead of consulting an ownership table.
## Syscall surface
Exclusivity stops being a broker refusing a second claimant and becomes the ordinary
property of a capability: only one process was given it.
- `device_enumerate` — **retires.** It exists only to read the kernel's table.
Callers ask `device-manager`, whose protocol already reserves an `enumerate`
verb. Note this is a public-ABI change: `vdso.md` documents it.
- `device_register` — **splits.** The kernel half becomes `resource_derive`; the
publication half ("this device exists, here is what it is") becomes an IPC
message to `device-manager`, which is where the inventory belongs and where
`child_added` already carries the same facts.
- `device_claim` — **dissolves into possession**, except for the IOMMU (below).
- `mmio_map`, `irq_bind`, `msi_bind`, `io_read`, `io_write`, `dma_bind` — keep
their names and semantics; their first argument becomes a capability handle.
**`maximum_children_per_parent` is deleted here.** After this a bus driver's children are
devices it enumerated on a bus it was actually given, and the threat the cap was written
for no longer exists.
## The open question: where the IOMMU attaches
## What this costs
This is the one place the kernel still needs device *identity* rather than a
range. `confineDevice(device_id, bdf, owner)` builds a domain keyed by PCI BDF,
attaches it, and the claim is rolled back if confinement fails — deliberately,
so a device that cannot be confined is never driven.
**Bring-up order changes.** `pci-bus` today claims first and says hello afterwards — its
own comment says "Claim the bridge, map the ECAM, hello the manager, then scan." Under
delegation the hello must come first, because that is where the grant arrives. Five
drivers need that reordering, and it is the bulk of the work.
Three options, none obviously right:
**`init` grows a device role.** It already registers `/protocol` and reads
`protocol.csv`; it would also hold root device grants and hand them on. That is more
responsibility in PID 1, which is a cost worth naming — though the alternative is the
kernel deciding who may hold what, which principle 5 excludes.
1. **The BDF rides the capability.** A memory capability derived for a PCI
function carries its BDF, and the kernel confines on first `mmio_map` or
`dma_bind`. Keeps the syscall count down; means a capability is no longer
purely a range.
2. **An explicit `device_attach(cap, bdf)`.** Honest and visible, but it puts a
BDF — a fact about PCI — into a kernel interface that otherwise knows nothing
about buses, and something must stop a caller naming a BDF that is not
theirs.
3. **The root PCI capability carries the segment, and derivation computes the
BDF.** Purest, but only works for PCI and the kernel would be parsing bus
topology, which is precisely what this document is trying to stop.
**A configuration question follows.** `/protocol` grants are data (`protocol.csv`),
because who may speak to whom is policy. Device grants could be too — a manifest saying
which binary may be given which device class. That would be the natural symmetry, and it
is *not* proposed here: `init` handing the manager everything and the manager matching
by `devices.csv` is the smaller step, and the manifest can follow if it earns itself.
My inclination is (1), because confinement is a property of the resource being
granted rather than a separate action, and because it keeps the "possession is
authority" story intact. It needs the derivation call to know it is carving a
PCI function, which is a wart worth arguing about.
## What it does not solve
## What this does not solve
- **Hot-unplug and re-enumeration drift.** Still the inventory problem (phase 3, and
open question 4 in the track plan). A grant dying with its holder is not the same as a
device going away.
- **The device table's size.** `maximum_devices` is untouched by this; it goes when the
inventory moves in phase 3.
- **Two processes racing for the same root grant.** Cannot arise, because roots are
minted to `init` alone.
- **Hot-plug and removal.** Capabilities die with their holder, but a device
that physically disappears while a driver lives is the device manager's
problem and unchanged by this.
- **Quotas on physical memory.** Nothing here bounds how much a driver maps; it
bounds only *what* it may map. That was already true.
- **The device tree as a published thing.** `/system/devices` remains a device
manager concern, as the file-system hierarchy already assumes.
## How this is verified
## Migration sketch
Not a plan yet — the phases need sizing once the IOMMU question is settled.
Rough shape: introduce the capability kind and `resource_derive` alongside the
existing table; convert the resource-consuming syscalls to accept either form;
move the inventory into `device-manager` and convert its clients off
`device_enumerate`; then delete the table, the two constants and the three
syscalls in one flag-day, as the `ServiceId` retirement did.
Every driver is affected, so the QEMU suite is the arbiter at each step, and the
Ryzen is the acceptance test — it is the machine that found this, and the one
that proves it fixed.
The invariant is **I3** from the track plan: a process holds what it was handed and
cannot name its way into holding more. The test is adversarial and the suite has never
had one of these for devices: a process that was granted nothing calls `device_claim`
on a device another driver owns, and on one nobody owns, and is refused both times with
its own errno. The audit's lesson was that "the suite contains no attacker"; this is the
attacker for devices.