docs: device authority — the how, and the run that deletes the ceilings

device-authority.md is rewritten as an implementation design rather than a
rival to device-manager.md. The what was already settled there in 2026-07:
structure in the manager, authority in the kernel, and delegation as the
step after hello. This is the how, plus the two decisions that paragraph
leaves open.

Decision 1: the manager claims, it is not granted. device-manager.md says
"claims (or is granted)"; claiming wins because the manager runs before any
driver exists and takes the seeded devices unopposed, leaving nothing unheld
to race for. One new call, device_transfer(device_id, task_id), checks only
that the caller holds the device — no names in the kernel, no attestation.
The alternative put a binary path inside the kernel, and the kernel should
hold only what cannot safely live in user space. The residual is stated
rather than hidden: authority rests on the manager claiming first, which
init.csv makes an operator-visible ordering rather than an attacker-
controlled one, and the enforced version arrives with the spawn capability
drivers.md already names as missing.

Decision 2: the kernel stops holding inventory. It reads three things out of
a descriptor — physical ranges, interrupt numbers, one PCI BDF — and stores
the rest only so device_enumerate can hand it back. Devices with no
resources leave the kernel entirely: a USB device conveys no mapping
authority, so there is nothing to enforce. That is also the case which
sidesteps containment, and therefore the reason a shared cap existed.

Decision 3: no shared ceiling. The table becomes dynamic — it is built after
heap.init, so nothing ever prevented it — and the two invented numbers go. A
per-holder quota replaces them, because dynamic storage with no bound moves
the ceiling to the kernel heap, which is shared and fatal rather than
partial. A bound charged to whoever caused it is isolation.

Two earlier drafts of this document are gone: one gave init the root grants,
the other proposed extracting a firmware-framebuffer driver. Both were wrong
and both are recorded as wrong in the run plan's settled list — the
framebuffer is not a device, it is where pixels go until a real display
driver announces itself.

Run 2 is nine steps, ordered so the suite stays green throughout: build and
prove the transfer mechanism, move the five claimants across one at a time,
then the flag day, then the inventory, then the ceilings.
This commit is contained in:
Daniel Samson
2026-08-08 17:01:25 +01:00
parent 5ad42ceac3
commit 547d0ec46b
2 changed files with 181 additions and 175 deletions
+123 -174
View File
@@ -1,12 +1,25 @@
# Device authority: you hold what you were given
# Device authority: the implementation of delegation
*Design for phase 2 of [the bounds track](../bounds-track-plan.md), 2026-08-08.
Supersedes an earlier draft that argued for capabilities on aesthetic grounds; this one
starts from the hole and from the project's principles.*
*Implementation design, 2026-08-08. The **what** is settled in
[device-manager.md](../device-driver-development/device-manager.md) — "structure in the
manager, authority in the kernel", and delegation as the step after `hello`. This
document is the **how**, and the decisions that paragraph leaves open.*
## The hole
Read first: [drivers.md](../device-driver-development/drivers.md) (the claim is the
capability), [driver-model.md](../device-driver-development/driver-model.md) (the three
invariants), [device-manager.md](../device-driver-development/device-manager.md) (the
tree, the matcher, the supervisor).
`device_claim(id)` checks two things ([devices-broker.zig](../../system/kernel/devices-broker.zig)):
## The one thing not yet true
`device-manager.md` says assignment "stays argv for now", and names the next step:
> The step after `hello` exists is delegation: the manager claims (or is granted) the
> devices and passes the claim to the driver over IPC (the M13 capability-transfer
> mechanism), replacing first-come-first-served `device_claim` with policy.
Until that lands, the manager's matching is advisory. `device_claim` checks only that
the device exists and is unheld ([devices-broker.zig](../../system/kernel/devices-broker.zig)):
```zig
pub fn claim(id: u64, owner: u32) ClaimError!void {
@@ -16,187 +29,123 @@ pub fn claim(id: u64, owner: u32) ClaimError!void {
}
```
Does it exist, and is it free. **Any process may claim any unclaimed device.**
A driver is spawned with its device id in `argv[1]` and claims it; any process could
pass any integer instead. Since a claim is a licence to map physical memory, that is the
gap this document closes.
The matching is real but it is entirely advisory: `device-manager` reads `devices.csv`,
matches a device to a driver, and spawns that driver with the device id as `argv[1]`
(`spawnDriver`). The driver parses the string and claims it. Nothing anywhere binds the
manager's decision to the kernel's grant — a process can pass any integer and win the
race.
## Decision 1: the manager claims, then transfers
A claim is not a small thing. It is what gates `mmio_map` and `irq_bind`, so it is a
licence to map physical memory and receive interrupts.
`device-manager.md` leaves "claims (or is granted)" open. **Claims.**
## What this costs, beyond the obvious
The manager runs before any driver exists — `init` starts it from `init.csv`, and it is
what spawns drivers — so it takes the seeded devices unopposed and there is nothing
unheld left for anyone to race for. One new call moves ownership on:
`maximum_children_per_parent = 16` exists because a driver that claimed one device could
loop `device_register` under it and exhaust the shared table. That threat only exists
*because* claiming is unauthenticated — and the cap is a poor defence against it, since
an attacker can burn 16 slots, claim another device, and burn 16 more. What it reliably
does instead is refuse a legitimate PCI bus with more than 16 functions, which is how an
AMD Ryzen came to boot with no USB and no storage.
```
device_transfer(device_id, task_id) -> 0/-errno
```
So the cap is not merely mis-sized. It is standing in for an authorisation that is not
performed, and it punishes correct behaviour while barely inconveniencing incorrect
behaviour. **Closing the hole is what retires the constant**, not a bigger number.
The kernel checks only that the caller currently holds the device. No names, no
attestation, no notion of "the device manager" — the rule is *you may give away what you
hold*, which is the capability discipline already in force.
## What the principles decide
The alternative was the kernel granting roots to a task it recognises by binary path
plus a PID-1 supervisor. It is more robust — it does not depend on the manager being
first — but it puts a binary name inside the kernel, and the goal is that the kernel
keeps only what cannot safely live in user space. A name is not that.
- *Move as much responsibility as possible to user space* (3), and *what remains in the
kernel is there for security or a hardware limitation* (5).
**The residual, stated plainly:** authority here rests on the manager claiming first.
That holds because `init.csv` decides what starts and in what order, so it is an
operator-visible ordering rather than an attacker-controlled one — but it is an
assumption, not an enforced invariant. The enforced version arrives with the spawn
capability [drivers.md](../device-driver-development/drivers.md) already names as
missing ("`system_spawn` is currently ungated … because there is no spawn capability
yet"). This design is compatible with it and does not block on it.
Deciding **which** driver gets **which** device is policy — matching identity triples
and choosing a binary. That decision belongs to the device manager and stays there.
`devices.csv` is not the policy; it is **configuration**, the declarative data the
policy reads. Three distinct things, and worth keeping apart in this document:
## Decision 2: the kernel stops holding inventory
| | Lives in | Example |
The kernel reads exactly three things out of a descriptor: **physical ranges** (to check
a mapping falls inside one), **interrupt numbers**, and **one PCI BDF** (to key an IOMMU
domain). Vendor, device and subsystem ids, class triples, `_HID` strings, bus addresses
and names are stored only so `device_enumerate` can hand them back — which
`device-manager.md` already resolves: that call "fades to a manager-internal (then
deleted) seam", because the manager owns the tree as data.
So the kernel's table becomes: **parent, resources, holder, BDF.** That is what cannot
safely run in user space; the rest moves.
**Devices with no resources leave the kernel entirely.** A USB device is addressed
through its controller and carries `resource_count = 0`
([driver-model.md](../device-driver-development/driver-model.md): "that case is allowed
and is the common one"). It conveys no mapping authority, so there is nothing for the
kernel to enforce and no reason for it to know. It is inventory, and inventory is the
manager's — reported by `child_added`, which already carries everything needed.
That is also the case that made `maximum_children_per_parent` necessary: a zero-resource
child sidesteps containment, so a driver could loop `device_register` and fill the
shared table. Once such children are not kernel objects, every remaining entry is a real
contained subdivision of something the caller holds.
## Decision 3: no shared ceiling; a per-holder quota instead
`maximum_devices = 64` and `maximum_children_per_parent = 16` are numbers we invented,
and both are shared — one driver's enumeration starves every other driver, which is how
an AMD Ryzen booted with no USB and no storage.
- **The table becomes dynamic.** It is built after `heap.init` (`kernel.zig`: `pmm.init`
at 137, `heap.init` at 179, `devices_broker.init` at 202), so nothing prevents it. No
specification bounds how many devices a machine has, so nothing should bound ours.
- **`maximum_children_per_parent` is deleted**, because the authorisation it stood in
for now exists.
- **A per-holder quota replaces them.** Dynamic storage without a bound moves the
ceiling to the kernel heap, which is shared and fatal rather than partial — strictly
worse. The bound that is *not* worse is one charged to the task that caused it: a
driver that loops `device_register` exhausts its own allowance, is refused with an
attributable errno, and is restarted by its supervisor while every other driver
carries on. That is the microkernel property rather than a workaround for it, and it
is declared through [bounds.md](bounds.md) like any other.
## The shape of the change
| | Before | After |
|---|---|---|
| **Mechanism** | the kernel | the check that a grant is held before a mapping is made |
| **Policy** | user space | the device manager matching a device to a driver |
| **Configuration** | files | `devices.csv`, `protocol.csv`, `init.csv` |
| Manager gets its devices | claims them, unauthorised | claims them (first, unopposed) |
| Driver gets its device | `argv[1]` + `device_claim` | receives it in the `hello` reply |
| Kernel checks | is it free? | do you hold it? |
| Kernel stores | the full descriptor | parent, resources, holder, BDF |
| Zero-resource devices | kernel table entries | manager records only |
| Table size | `maximum_devices = 64` | dynamic, per-holder quota |
| Children per parent | `maximum_children_per_parent = 16` | deleted |
Enforcing that a driver **holds only what it was given** is security — it is the gate in
front of mapping physical memory. That stays in the kernel, and it is the whole of what
the kernel needs to do.
Bring-up order changes for the five claiming drivers: `hello` must precede the claim,
because the reply is where the device arrives. `pci-bus` today does the reverse — its
own comment reads "Claim the bridge, map the ECAM, hello the manager, then scan."
The kernel therefore does not need to know about matching, `devices.csv`, driver names,
or why a device was assigned. It needs to know that an authority it can verify granted
this device to this task.
## What does not change
## Only drivers hold devices
- The three invariants of [driver-model.md](../device-driver-development/driver-model.md):
a claim is exclusive, a descriptor is a licence to map physical memory, therefore
containment. This design strengthens the first and touches neither of the others.
- The display service's GOP path. The framebuffer is not a device — it is where pixels
go, handed over by the loader, and the compositor uses it as the boot floor until a
real display driver announces itself
([display-v2.md](../device-driver-development/display-v2.md)). The kernel wraps it in
a display-class descriptor so `mmio_map` can hand it over write-combining; that is
plumbing for a mapping, not a claim about what it is.
- Supervision, restart, pruning and re-report
([device-manager.md](../device-driver-development/device-manager.md),
[process-lifecycle.md](process-lifecycle.md)). Delegation slots into the existing
`hello` exchange and changes none of it.
- `device_register` idempotency, which is what lets a restarted bus rebuild the same
ids.
The layering the rest of the system already follows:
## How it is verified
- A **driver** talks to hardware. It is the only thing that holds a device.
- A **service** talks to no hardware at all. It receives device events and sends data
and commands over a **protocol**.
- A **protocol** is the abstraction between them — the OS layer, in the sense of
[communication.md](communication.md).
The invariant is: **a process holds what it was handed and cannot name its way into
holding more.** The suite has no adversarial device case today — the audit's lesson was
that "the suite contains no attacker" — so this adds one: a process that was handed
nothing calls `device_claim` and `device_transfer` on a device another driver holds, and
on one nobody holds, and is refused each time with its own errno.
So the question "which services need device grants" has the answer **none**. That
collapses the design: every holder of a device is a driver, and every driver is spawned
by the device manager, which is what grants it.
`display` already demonstrates both halves, one right and one wrong
([backend.zig](../../system/services/display/backend.zig)):
| Backend | How it reaches the panel |
|---|---|
| `VirtioGpu` | speaks `scanout-protocol` over an IPC handle; touches no device |
| `Gop` | `device.enumerate`, `device.claim`, `device.mmioMap` |
The virtio-gpu path is the intended shape and it works today: the driver holds the
hardware, the service speaks a protocol to it. **The GOP path is the anomaly** — the
display service reaching into hardware itself, because the firmware framebuffer has no
driver for it to talk to.
It is also the only anomaly. Every other claimant — `pci-bus`, `usb-xhci-bus`,
`ps2-bus`, `virtio-gpu`, `acpi` — is a device-manager child spawned with its device id
in `argv[1]`.
**A firmware-framebuffer driver is therefore a prerequisite of this phase**, not a
consequence of it. It is small: it holds the display device, maps the framebuffer, and
serves `scanout-protocol` — the same interface `virtio-gpu` already serves, so the
compositor needs no new code path and stops caring which is behind it, which it was
designed for.
An earlier draft of this document had `display` asking the device manager for a grant,
and before that had `init` minting grants because `init` starts `display`. Both were
accommodating an exception instead of removing it.
## The design
**A device grant is a capability, delegated from a holder.** The mechanism already
exists: `callCap` passes a handle over an IPC call and the kernel installs it in the
receiver's table ([library/kernel/ipc.zig](../../library/kernel/ipc.zig)), which is how
shared memory and DMA regions already move between processes.
**The device manager is the root of device authority.** `init` has no part in this: it
is the first process, and its job is to start the rest of the system. It starts the
device manager the same way it starts everything else, and knows nothing about devices.
1. **Root.** At boot the kernel mints grants for the devices firmware discovery found
and hands them to the device manager. This is the only place device authority enters
the system, and it comes from ACPI rather than from anyone's say-so.
The kernel recognises the manager by **chain attestation**, the identity the security
track already settled on: the binary path the kernel itself stamped
(`/system/services/device-manager`) *and* a supervisor of PID 1. The binary path
alone would not do — `system_spawn` is deliberately ungated, so any process may spawn
any bundled binary, and a rogue could run a second copy under the same name. It could
not forge the other half: its copy's supervisor is the rogue, and PID 1 is the
kernel's own first process.
2. **Delegation.** The manager passes a driver its device when it spawns it, over the
channel that already exists — the driver `hello`s the manager and the reply carries
the grant. The manager attests its own children by task id, one hop deep, exactly as
`init` attests the services it spawned for `/protocol` (`supervisorSatisfies`); it
knows which task is which driver because it spawned them.
3. **Services receive nothing.** `display`, `input`, `fat` and `logger` hold no devices
and need no grants; they speak protocols to the drivers that do. The one place this
is not true today is the GOP backend above, which a firmware-framebuffer driver
removes.
4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and `dma_bind`
check possession of the grant instead of consulting an ownership table.
Exclusivity stops being a broker refusing a second claimant and becomes the ordinary
property of a capability: only one process was given it.
**`maximum_children_per_parent` is deleted here.** After this a bus driver's children are
devices it enumerated on a bus it was actually given, and the threat the cap was written
for no longer exists.
## What this costs
**Bring-up order changes.** `pci-bus` today claims first and says hello afterwards — its
own comment says "Claim the bridge, map the ECAM, hello the manager, then scan." Under
delegation the hello must come first, because that is where the grant arrives. Five
drivers need that reordering, and it is the bulk of the work.
**A firmware-framebuffer driver has to be written first**, and until it exists the
display service cannot stop claiming a device. It is the smallest new binary in the
tree — hold the display node, map the framebuffer, serve `scanout-protocol` — but it is
new code on the boot path, and the boot path is where a mistake costs a screen.
It brings an ordering constraint with it: `display` cannot paint until that driver is
up, where today it maps the framebuffer itself and is independent. Mitigating,
`display` already tolerates a backend arriving late — `virtio-gpu` hands it a scanout
after the fact over `attach_scanout` — so "wait, then attach" is a path that already
works rather than one to invent.
**The kernel gains one piece of knowledge about a specific binary.** Chain attestation
means the kernel recognises `/system/services/device-manager` under PID 1 as the root
holder. That is a real concession — the kernel would rather know nothing about who is
who — and it is the minimum: device authority has to enter the system somewhere, and
every alternative is worse. Configuration the kernel reads would put a file parser in
the kernel; first-to-ask would be the hole again, at boot.
**A configuration question stays open.** Which binary may be given which device is
expressed today as `devices.csv`, read by the manager — configuration, with the manager
as the policy that acts on it. That is already the right shape and needs no new file.
A separate grant manifest earns itself only when something must differ from "the manager
gets the hardware": holding a device back so a test or a bare-metal driver can take it,
for instance. Not now.
## What it does not solve
- **Hot-unplug and re-enumeration drift.** Still the inventory problem (phase 3, and
open question 4 in the track plan). A grant dying with its holder is not the same as a
device going away.
- **The device table's size.** `maximum_devices` is untouched by this; it goes when the
inventory moves in phase 3.
- **Two processes racing for the same root grant.** Cannot arise: the roots are minted
to whichever task satisfies the chain — the manager's binary under PID 1 — and a
second copy spawned by anyone else fails the supervisor half.
## How this is verified
The invariant is **I3** from the track plan: a process holds what it was handed and
cannot name its way into holding more. The test is adversarial and the suite has never
had one of these for devices: a process that was granted nothing calls `device_claim`
on a device another driver owns, and on one nobody owns, and is refused both times with
its own errno. The audit's lesson was that "the suite contains no attacker"; this is the
attacker for devices.
The Ryzen is the acceptance test for the ceiling half: it is the machine that found the
constants, and the one that proves them gone.