docs: phase 2 design — you hold what you were given
device_claim checks that a device exists and is free. That is all. Any process may claim any unclaimed device, and a claim is what gates mmio_map and irq_bind — a licence to map physical memory and take interrupts. The device manager's matching is real but advisory: it spawns a driver with the device id in argv[1] and nothing binds that decision to the kernel's grant. maximum_children_per_parent is the visible cost. It exists because a driver that claimed one device could loop device_register under it, and it is a poor defence — an attacker burns 16 slots, claims another device, burns 16 more — while reliably refusing a legitimate PCI bus with more than 16 functions. Closing the hole is what retires the constant. The principles decide the split: matching is policy and stays in the device manager; enforcing that a driver holds only what it was given is security and stays in the kernel, and is the whole of what the kernel needs. The mechanism is decided by an awkward fact. Five of six claimants are device-manager children spawned with their device id. display is not — init spawns it, and it finds its framebuffer by enumerating for a display-class node and claiming whatever it finds. So "record the device named at spawn" closes the hole for five and breaks the sixth, and the sixth is not an oddity to special-case: it shows authority must be delegable rather than welded to the moment of spawn. So: a device grant is a capability, minted by the kernel to init for the devices firmware discovery found, delegated by init to the device manager and to display, and passed by the manager to each driver it spawns. The cap-passing path already exists and already carries shared memory and DMA regions. init is already the grantor for /protocol, and protocol.csv already records init granting the device manager its binding. Costs named rather than buried: five drivers must hello before claiming (pci-bus claims first today), and init grows a device role on top of the protocol registry. A device-grant manifest mirroring protocol.csv would be the natural symmetry and is deliberately not proposed yet.
This commit is contained in:
@@ -1,173 +1,138 @@
|
|||||||
# Device authority: the kernel stops keeping an inventory
|
# Device authority: you hold what you were given
|
||||||
|
|
||||||
*Design, drafted 2026-08-07. Not implemented. Prompted by a real machine: an AMD
|
*Design for phase 2 of [the bounds track](../bounds-track-plan.md), 2026-08-08.
|
||||||
Ryzen desktop enumerated more PCI functions than the kernel's device table would
|
Supersedes an earlier draft that argued for capabilities on aesthetic grounds; this one
|
||||||
hold, and the xHCI and SATA controllers were refused registration — so the
|
starts from the hole and from the project's principles.*
|
||||||
machine booted to the compositor with no USB and no storage.*
|
|
||||||
|
|
||||||
The kernel keeps a table of every device userspace discovers. It is a fixed
|
## The hole
|
||||||
array of 64 descriptors, 344 bytes each, and a second cap allows any one parent
|
|
||||||
16 children. Neither number is written down anywhere as a decision:
|
|
||||||
`maximum_devices` has no comment and never reached `parameters.zig`, where every
|
|
||||||
other tunable in this kernel lives with its reasoning attached.
|
|
||||||
|
|
||||||
Raising them is not the fix. The numbers are wrong because the *table* is wrong:
|
`device_claim(id)` checks two things ([devices-broker.zig](../../system/kernel/devices-broker.zig)):
|
||||||
it is an inventory of hardware, and an inventory of hardware is not something a
|
|
||||||
kernel needs. This document proposes replacing it with capabilities, which
|
|
||||||
removes the ceiling rather than moving it.
|
|
||||||
|
|
||||||
## What the kernel actually uses
|
```zig
|
||||||
|
pub fn claim(id: u64, owner: u32) ClaimError!void {
|
||||||
|
if (id >= count) return error.NoSuchDevice;
|
||||||
|
if (claimed[@intCast(id)] != null) return error.AlreadyClaimed;
|
||||||
|
claimed[@intCast(id)] = owner;
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
Every read of a device descriptor from the kernel proper, exhaustively:
|
Does it exist, and is it free. **Any process may claim any unclaimed device.**
|
||||||
|
|
||||||
| Used for | What it needs |
|
The matching is real but it is entirely advisory: `device-manager` reads `devices.csv`,
|
||||||
|---|---|
|
matches a device to a driver, and spawns that driver with the device id as `argv[1]`
|
||||||
| `mmio_map`, `io_read`/`io_write` | the physical range, to check the mapping falls inside it |
|
(`spawnDriver`). The driver parses the string and claims it. Nothing anywhere binds the
|
||||||
| `irq_bind`, `msi_bind` | the GSI, and that the caller owns the device |
|
manager's decision to the kernel's grant — a process can pass any integer and win the
|
||||||
| `dma_bind` | that the caller owns the device |
|
race.
|
||||||
| IOMMU confinement | the **PCI BDF**, to key a domain |
|
|
||||||
| `device_register` | the parent's ranges, for the containment check |
|
|
||||||
| the boot display seed | one framebuffer window |
|
|
||||||
|
|
||||||
That is: **physical ranges, interrupt numbers, and one BDF.** Vendor, device and
|
A claim is not a small thing. It is what gates `mmio_map` and `irq_bind`, so it is a
|
||||||
subsystem ids, class triples, the human-readable names, the parent links, the
|
licence to map physical memory and receive interrupts.
|
||||||
bus numbers — the kernel stores all of it and reads none of it. It is held so
|
|
||||||
that `device_enumerate` can hand it back to user space, which is the whole
|
|
||||||
mistake in one sentence: the kernel is acting as a distribution mechanism for
|
|
||||||
data it does not use.
|
|
||||||
|
|
||||||
## The split
|
## What this costs, beyond the obvious
|
||||||
|
|
||||||
Three concerns are tangled in one table.
|
`maximum_children_per_parent = 16` exists because a driver that claimed one device could
|
||||||
|
loop `device_register` under it and exhaust the shared table. That threat only exists
|
||||||
|
*because* claiming is unauthenticated — and the cap is a poor defence against it, since
|
||||||
|
an attacker can burn 16 slots, claim another device, and burn 16 more. What it reliably
|
||||||
|
does instead is refuse a legitimate PCI bus with more than 16 functions, which is how an
|
||||||
|
AMD Ryzen came to boot with no USB and no storage.
|
||||||
|
|
||||||
**Platform bring-up** — timers, LAPIC/IOAPIC, CPU topology, the ECAM window, the
|
So the cap is not merely mis-sized. It is standing in for an authorisation that is not
|
||||||
framebuffer the firmware left. The kernel derives these from ACPI before user
|
performed, and it punishes correct behaviour while barely inconveniencing incorrect
|
||||||
space exists and needs them to function. They never belonged in the device table
|
behaviour. **Closing the hole is what retires the constant**, not a bigger number.
|
||||||
and mostly are not (the platform block in `kernel.zig` is separate); this
|
|
||||||
document does not change them.
|
|
||||||
|
|
||||||
**Resource authority** — which task may map which physical range, receive which
|
## What the principles decide
|
||||||
interrupt, touch which ports. This *must* stay in the kernel. It is the one
|
|
||||||
grant that cannot be audited after the fact: a process that maps arbitrary
|
|
||||||
physical memory owns the machine, page tables and IOMMU structures included.
|
|
||||||
This is memory protection, not device management, and it is why the answer is
|
|
||||||
not simply "move it all to the device manager".
|
|
||||||
|
|
||||||
**Device inventory** — what exists, what it is, how it is arranged, which driver
|
- *Move as much responsibility as possible to user space* (3), and *what remains in the
|
||||||
should bind it. This is `device-manager`'s job and is already half there: it
|
kernel is there for security or a hardware limitation* (5).
|
||||||
loads `devices.csv`, matches, spawns drivers, and receives `child_added` reports
|
|
||||||
over its own protocol. The kernel table duplicates what those reports already
|
|
||||||
carry.
|
|
||||||
|
|
||||||
## The proposal: a resource is a capability
|
Deciding **which** driver gets **which** device is policy: it reads a CSV, matches
|
||||||
|
identity triples, and picks a binary. That is the device manager's, and it stays there.
|
||||||
|
|
||||||
Device resources join endpoints, shared memory and DMA regions as a kind in the
|
Enforcing that a driver **holds only what it was given** is security — it is the gate in
|
||||||
handle table.
|
front of mapping physical memory. That stays in the kernel, and it is the whole of what
|
||||||
|
the kernel needs to do.
|
||||||
|
|
||||||
1. **Roots.** At boot the kernel mints capabilities for the windows it learned
|
The kernel therefore does not need to know about matching, `devices.csv`, driver names,
|
||||||
from firmware — the ECAM range, the framebuffer, the legacy port space — and
|
or why a device was assigned. It needs to know that an authority it can verify granted
|
||||||
hands them to the first bus drivers. This is the only place device knowledge
|
this device to this task.
|
||||||
enters the kernel, and it comes from ACPI, not from a driver's say-so.
|
|
||||||
2. **Subdivision.** A bus driver enumerating hardware derives a narrower
|
|
||||||
capability from one it holds: `resource_derive(cap, kind, start, len) → cap`.
|
|
||||||
The kernel checks the sub-range lies inside the capability being subdivided —
|
|
||||||
the same containment rule as today (`devices-broker.contains`), but checked
|
|
||||||
against *one capability the caller demonstrably holds* rather than by walking
|
|
||||||
a global tree.
|
|
||||||
3. **Delegation.** The driver passes that capability to the child driver over
|
|
||||||
IPC. Cap-passing already exists; this is the mechanism `subscribe` and
|
|
||||||
`attach_scanout` already use.
|
|
||||||
4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and
|
|
||||||
`dma_bind` take a capability handle instead of `(device_id, resource_index)`.
|
|
||||||
Possession *is* the authority — there is nothing to look up and no ownership
|
|
||||||
table to consult.
|
|
||||||
|
|
||||||
Exclusivity stops being a broker refusing a second claimant and becomes the
|
## The awkward fact that decides the mechanism
|
||||||
ordinary property of a capability: only one process was given it.
|
|
||||||
|
|
||||||
## What this buys
|
Five of the six claimants are device-manager children, spawned with their device id in
|
||||||
|
`argv[1]`: `pci-bus`, `usb-xhci-bus`, `ps2-bus`, `virtio-gpu`, `acpi`.
|
||||||
|
|
||||||
**No ceiling.** There is no table to size, so no machine is too big. The
|
**`display` is not.** It is spawned by `init` from `init.csv`, and it finds its device by
|
||||||
Ryzen's enumeration stops being a limit to tune and becomes what it is — a fact
|
enumerating the table for a display-class node and claiming whatever it finds
|
||||||
about a computer.
|
([backend.zig](../../system/services/display/backend.zig)). There is no assignment to
|
||||||
|
enforce, because nobody assigned it anything.
|
||||||
|
|
||||||
**Reclamation, free.** `count` in the broker today only ever increases;
|
That rules out the cheapest design. "The kernel records the device named at spawn, and
|
||||||
`releaseAllOwnedBy` clears a dead driver's *claims* but never its entries. A
|
`device_claim` checks it" closes the hole for five claimants and breaks the sixth. And
|
||||||
driver that crashes and is restarted re-registers its children and consumes the
|
the sixth is not an oddity to special-case — it is the one that shows the model is
|
||||||
table again — reachable today without any malice, given the device manager
|
wrong: authority should be *delegable*, not welded to the moment of spawn.
|
||||||
restarts drivers by design. Capabilities die with the task.
|
|
||||||
|
|
||||||
**No quota needed.** The per-parent cap exists to stop one claimant looping
|
## The design
|
||||||
`device_register` and filling the shared table, because a zero-resource child
|
|
||||||
sidesteps the containment check. With no shared table there is nothing to
|
|
||||||
exhaust; a process can only ever subdivide what it was given, and its handles
|
|
||||||
are already bounded per task.
|
|
||||||
|
|
||||||
**A smaller kernel.** Three syscalls leave the ABI, one narrower one arrives,
|
**A device grant is a capability, delegated from a holder.** The mechanism already
|
||||||
and `devices-broker.zig` largely disappears along with both constants.
|
exists: `callCap` passes a handle over an IPC call and the kernel installs it in the
|
||||||
|
receiver's table ([library/kernel/ipc.zig](../../library/kernel/ipc.zig)), which is how
|
||||||
|
shared memory and DMA regions already move between processes.
|
||||||
|
|
||||||
**The discipline the rest of the system already uses.** "The claim is the
|
1. **Root.** At boot the kernel mints grants for the devices firmware discovery found
|
||||||
capability" is written in the driver documentation as if it were already true.
|
and hands them to `init` (PID 1, which the kernel spawns and therefore need not
|
||||||
This makes it true.
|
authenticate). This is the only place device authority enters the system, and it
|
||||||
|
comes from ACPI rather than from anyone's say-so.
|
||||||
|
2. **Delegation.** `init` passes the device manager the grants it will need — in
|
||||||
|
practice all of them — and passes `display` the framebuffer grant, because `init` is
|
||||||
|
what starts `display`. This is the same shape as the `/protocol` registry, where
|
||||||
|
`init` is already the grantor and `protocol.csv` already records
|
||||||
|
`/system/services/device-manager, /system/services/init, bind, device-manager`.
|
||||||
|
3. **Assignment.** The manager passes a driver its device when it spawns it, over the
|
||||||
|
channel that already exists — the driver `hello`s the manager, and the reply carries
|
||||||
|
the grant.
|
||||||
|
4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and `dma_bind`
|
||||||
|
check possession of the grant instead of consulting an ownership table.
|
||||||
|
|
||||||
## Syscall surface
|
Exclusivity stops being a broker refusing a second claimant and becomes the ordinary
|
||||||
|
property of a capability: only one process was given it.
|
||||||
|
|
||||||
- `device_enumerate` — **retires.** It exists only to read the kernel's table.
|
**`maximum_children_per_parent` is deleted here.** After this a bus driver's children are
|
||||||
Callers ask `device-manager`, whose protocol already reserves an `enumerate`
|
devices it enumerated on a bus it was actually given, and the threat the cap was written
|
||||||
verb. Note this is a public-ABI change: `vdso.md` documents it.
|
for no longer exists.
|
||||||
- `device_register` — **splits.** The kernel half becomes `resource_derive`; the
|
|
||||||
publication half ("this device exists, here is what it is") becomes an IPC
|
|
||||||
message to `device-manager`, which is where the inventory belongs and where
|
|
||||||
`child_added` already carries the same facts.
|
|
||||||
- `device_claim` — **dissolves into possession**, except for the IOMMU (below).
|
|
||||||
- `mmio_map`, `irq_bind`, `msi_bind`, `io_read`, `io_write`, `dma_bind` — keep
|
|
||||||
their names and semantics; their first argument becomes a capability handle.
|
|
||||||
|
|
||||||
## The open question: where the IOMMU attaches
|
## What this costs
|
||||||
|
|
||||||
This is the one place the kernel still needs device *identity* rather than a
|
**Bring-up order changes.** `pci-bus` today claims first and says hello afterwards — its
|
||||||
range. `confineDevice(device_id, bdf, owner)` builds a domain keyed by PCI BDF,
|
own comment says "Claim the bridge, map the ECAM, hello the manager, then scan." Under
|
||||||
attaches it, and the claim is rolled back if confinement fails — deliberately,
|
delegation the hello must come first, because that is where the grant arrives. Five
|
||||||
so a device that cannot be confined is never driven.
|
drivers need that reordering, and it is the bulk of the work.
|
||||||
|
|
||||||
Three options, none obviously right:
|
**`init` grows a device role.** It already registers `/protocol` and reads
|
||||||
|
`protocol.csv`; it would also hold root device grants and hand them on. That is more
|
||||||
|
responsibility in PID 1, which is a cost worth naming — though the alternative is the
|
||||||
|
kernel deciding who may hold what, which principle 5 excludes.
|
||||||
|
|
||||||
1. **The BDF rides the capability.** A memory capability derived for a PCI
|
**A configuration question follows.** `/protocol` grants are data (`protocol.csv`),
|
||||||
function carries its BDF, and the kernel confines on first `mmio_map` or
|
because who may speak to whom is policy. Device grants could be too — a manifest saying
|
||||||
`dma_bind`. Keeps the syscall count down; means a capability is no longer
|
which binary may be given which device class. That would be the natural symmetry, and it
|
||||||
purely a range.
|
is *not* proposed here: `init` handing the manager everything and the manager matching
|
||||||
2. **An explicit `device_attach(cap, bdf)`.** Honest and visible, but it puts a
|
by `devices.csv` is the smaller step, and the manifest can follow if it earns itself.
|
||||||
BDF — a fact about PCI — into a kernel interface that otherwise knows nothing
|
|
||||||
about buses, and something must stop a caller naming a BDF that is not
|
|
||||||
theirs.
|
|
||||||
3. **The root PCI capability carries the segment, and derivation computes the
|
|
||||||
BDF.** Purest, but only works for PCI and the kernel would be parsing bus
|
|
||||||
topology, which is precisely what this document is trying to stop.
|
|
||||||
|
|
||||||
My inclination is (1), because confinement is a property of the resource being
|
## What it does not solve
|
||||||
granted rather than a separate action, and because it keeps the "possession is
|
|
||||||
authority" story intact. It needs the derivation call to know it is carving a
|
|
||||||
PCI function, which is a wart worth arguing about.
|
|
||||||
|
|
||||||
## What this does not solve
|
- **Hot-unplug and re-enumeration drift.** Still the inventory problem (phase 3, and
|
||||||
|
open question 4 in the track plan). A grant dying with its holder is not the same as a
|
||||||
|
device going away.
|
||||||
|
- **The device table's size.** `maximum_devices` is untouched by this; it goes when the
|
||||||
|
inventory moves in phase 3.
|
||||||
|
- **Two processes racing for the same root grant.** Cannot arise, because roots are
|
||||||
|
minted to `init` alone.
|
||||||
|
|
||||||
- **Hot-plug and removal.** Capabilities die with their holder, but a device
|
## How this is verified
|
||||||
that physically disappears while a driver lives is the device manager's
|
|
||||||
problem and unchanged by this.
|
|
||||||
- **Quotas on physical memory.** Nothing here bounds how much a driver maps; it
|
|
||||||
bounds only *what* it may map. That was already true.
|
|
||||||
- **The device tree as a published thing.** `/system/devices` remains a device
|
|
||||||
manager concern, as the file-system hierarchy already assumes.
|
|
||||||
|
|
||||||
## Migration sketch
|
The invariant is **I3** from the track plan: a process holds what it was handed and
|
||||||
|
cannot name its way into holding more. The test is adversarial and the suite has never
|
||||||
Not a plan yet — the phases need sizing once the IOMMU question is settled.
|
had one of these for devices: a process that was granted nothing calls `device_claim`
|
||||||
Rough shape: introduce the capability kind and `resource_derive` alongside the
|
on a device another driver owns, and on one nobody owns, and is refused both times with
|
||||||
existing table; convert the resource-consuming syscalls to accept either form;
|
its own errno. The audit's lesson was that "the suite contains no attacker"; this is the
|
||||||
move the inventory into `device-manager` and convert its clients off
|
attacker for devices.
|
||||||
`device_enumerate`; then delete the table, the two constants and the three
|
|
||||||
syscalls in one flag-day, as the `ServiceId` retirement did.
|
|
||||||
|
|
||||||
Every driver is affected, so the QEMU suite is the arbiter at each step, and the
|
|
||||||
Ryzen is the acceptance test — it is the machine that found this, and the one
|
|
||||||
that proves it fixed.
|
|
||||||
|
|||||||
Reference in New Issue
Block a user