Files
danos/docs/os-development/device-authority.md
T
Daniel Samson 5ad42ceac3 docs: only drivers hold devices; services speak protocols
The design still had the display service receiving a device grant, which
keeps the wrong layering and just moves who hands it over. Drivers talk
hardware. Services talk to no hardware at all — they receive device events
and send data and commands over a protocol, which is the OS abstraction
between them.

So the answer to "which services need device grants" is none, and the design
collapses: every holder of a device is a driver, and every driver is spawned
by the device manager, which is what grants it. No exceptions to
accommodate.

display already shows both halves. Its VirtioGpu backend speaks
scanout-protocol over an IPC handle and touches no device — the driver holds
the hardware, the service speaks to it, and it works today. Its Gop backend
calls device.enumerate, device.claim and device.mmioMap: the service
reaching into hardware itself, because the firmware framebuffer has no
driver to talk to. That is the only such case in the tree; every other
claimant is a device-manager child spawned with its device id.

A firmware-framebuffer driver is therefore a prerequisite of this phase
rather than a consequence. It is small — hold the display node, map the
framebuffer, serve the same scanout-protocol virtio-gpu already serves — and
the compositor needs no new path, since not caring which backend is behind
it is what it was designed for.

Two earlier drafts are recorded as wrong in the document: display asking the
device manager for a grant, and before that init minting grants because init
starts display. Both accommodated an exception instead of removing it.
2026-08-08 16:40:37 +01:00

203 lines
11 KiB
Markdown

# Device authority: you hold what you were given
*Design for phase 2 of [the bounds track](../bounds-track-plan.md), 2026-08-08.
Supersedes an earlier draft that argued for capabilities on aesthetic grounds; this one
starts from the hole and from the project's principles.*
## The hole
`device_claim(id)` checks two things ([devices-broker.zig](../../system/kernel/devices-broker.zig)):
```zig
pub fn claim(id: u64, owner: u32) ClaimError!void {
if (id >= count) return error.NoSuchDevice;
if (claimed[@intCast(id)] != null) return error.AlreadyClaimed;
claimed[@intCast(id)] = owner;
}
```
Does it exist, and is it free. **Any process may claim any unclaimed device.**
The matching is real but it is entirely advisory: `device-manager` reads `devices.csv`,
matches a device to a driver, and spawns that driver with the device id as `argv[1]`
(`spawnDriver`). The driver parses the string and claims it. Nothing anywhere binds the
manager's decision to the kernel's grant — a process can pass any integer and win the
race.
A claim is not a small thing. It is what gates `mmio_map` and `irq_bind`, so it is a
licence to map physical memory and receive interrupts.
## What this costs, beyond the obvious
`maximum_children_per_parent = 16` exists because a driver that claimed one device could
loop `device_register` under it and exhaust the shared table. That threat only exists
*because* claiming is unauthenticated — and the cap is a poor defence against it, since
an attacker can burn 16 slots, claim another device, and burn 16 more. What it reliably
does instead is refuse a legitimate PCI bus with more than 16 functions, which is how an
AMD Ryzen came to boot with no USB and no storage.
So the cap is not merely mis-sized. It is standing in for an authorisation that is not
performed, and it punishes correct behaviour while barely inconveniencing incorrect
behaviour. **Closing the hole is what retires the constant**, not a bigger number.
## What the principles decide
- *Move as much responsibility as possible to user space* (3), and *what remains in the
kernel is there for security or a hardware limitation* (5).
Deciding **which** driver gets **which** device is policy — matching identity triples
and choosing a binary. That decision belongs to the device manager and stays there.
`devices.csv` is not the policy; it is **configuration**, the declarative data the
policy reads. Three distinct things, and worth keeping apart in this document:
| | Lives in | Example |
|---|---|---|
| **Mechanism** | the kernel | the check that a grant is held before a mapping is made |
| **Policy** | user space | the device manager matching a device to a driver |
| **Configuration** | files | `devices.csv`, `protocol.csv`, `init.csv` |
Enforcing that a driver **holds only what it was given** is security — it is the gate in
front of mapping physical memory. That stays in the kernel, and it is the whole of what
the kernel needs to do.
The kernel therefore does not need to know about matching, `devices.csv`, driver names,
or why a device was assigned. It needs to know that an authority it can verify granted
this device to this task.
## Only drivers hold devices
The layering the rest of the system already follows:
- A **driver** talks to hardware. It is the only thing that holds a device.
- A **service** talks to no hardware at all. It receives device events and sends data
and commands over a **protocol**.
- A **protocol** is the abstraction between them — the OS layer, in the sense of
[communication.md](communication.md).
So the question "which services need device grants" has the answer **none**. That
collapses the design: every holder of a device is a driver, and every driver is spawned
by the device manager, which is what grants it.
`display` already demonstrates both halves, one right and one wrong
([backend.zig](../../system/services/display/backend.zig)):
| Backend | How it reaches the panel |
|---|---|
| `VirtioGpu` | speaks `scanout-protocol` over an IPC handle; touches no device |
| `Gop` | `device.enumerate`, `device.claim`, `device.mmioMap` |
The virtio-gpu path is the intended shape and it works today: the driver holds the
hardware, the service speaks a protocol to it. **The GOP path is the anomaly** — the
display service reaching into hardware itself, because the firmware framebuffer has no
driver for it to talk to.
It is also the only anomaly. Every other claimant — `pci-bus`, `usb-xhci-bus`,
`ps2-bus`, `virtio-gpu`, `acpi` — is a device-manager child spawned with its device id
in `argv[1]`.
**A firmware-framebuffer driver is therefore a prerequisite of this phase**, not a
consequence of it. It is small: it holds the display device, maps the framebuffer, and
serves `scanout-protocol` — the same interface `virtio-gpu` already serves, so the
compositor needs no new code path and stops caring which is behind it, which it was
designed for.
An earlier draft of this document had `display` asking the device manager for a grant,
and before that had `init` minting grants because `init` starts `display`. Both were
accommodating an exception instead of removing it.
## The design
**A device grant is a capability, delegated from a holder.** The mechanism already
exists: `callCap` passes a handle over an IPC call and the kernel installs it in the
receiver's table ([library/kernel/ipc.zig](../../library/kernel/ipc.zig)), which is how
shared memory and DMA regions already move between processes.
**The device manager is the root of device authority.** `init` has no part in this: it
is the first process, and its job is to start the rest of the system. It starts the
device manager the same way it starts everything else, and knows nothing about devices.
1. **Root.** At boot the kernel mints grants for the devices firmware discovery found
and hands them to the device manager. This is the only place device authority enters
the system, and it comes from ACPI rather than from anyone's say-so.
The kernel recognises the manager by **chain attestation**, the identity the security
track already settled on: the binary path the kernel itself stamped
(`/system/services/device-manager`) *and* a supervisor of PID 1. The binary path
alone would not do — `system_spawn` is deliberately ungated, so any process may spawn
any bundled binary, and a rogue could run a second copy under the same name. It could
not forge the other half: its copy's supervisor is the rogue, and PID 1 is the
kernel's own first process.
2. **Delegation.** The manager passes a driver its device when it spawns it, over the
channel that already exists — the driver `hello`s the manager and the reply carries
the grant. The manager attests its own children by task id, one hop deep, exactly as
`init` attests the services it spawned for `/protocol` (`supervisorSatisfies`); it
knows which task is which driver because it spawned them.
3. **Services receive nothing.** `display`, `input`, `fat` and `logger` hold no devices
and need no grants; they speak protocols to the drivers that do. The one place this
is not true today is the GOP backend above, which a firmware-framebuffer driver
removes.
4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and `dma_bind`
check possession of the grant instead of consulting an ownership table.
Exclusivity stops being a broker refusing a second claimant and becomes the ordinary
property of a capability: only one process was given it.
**`maximum_children_per_parent` is deleted here.** After this a bus driver's children are
devices it enumerated on a bus it was actually given, and the threat the cap was written
for no longer exists.
## What this costs
**Bring-up order changes.** `pci-bus` today claims first and says hello afterwards — its
own comment says "Claim the bridge, map the ECAM, hello the manager, then scan." Under
delegation the hello must come first, because that is where the grant arrives. Five
drivers need that reordering, and it is the bulk of the work.
**A firmware-framebuffer driver has to be written first**, and until it exists the
display service cannot stop claiming a device. It is the smallest new binary in the
tree — hold the display node, map the framebuffer, serve `scanout-protocol` — but it is
new code on the boot path, and the boot path is where a mistake costs a screen.
It brings an ordering constraint with it: `display` cannot paint until that driver is
up, where today it maps the framebuffer itself and is independent. Mitigating,
`display` already tolerates a backend arriving late — `virtio-gpu` hands it a scanout
after the fact over `attach_scanout` — so "wait, then attach" is a path that already
works rather than one to invent.
**The kernel gains one piece of knowledge about a specific binary.** Chain attestation
means the kernel recognises `/system/services/device-manager` under PID 1 as the root
holder. That is a real concession — the kernel would rather know nothing about who is
who — and it is the minimum: device authority has to enter the system somewhere, and
every alternative is worse. Configuration the kernel reads would put a file parser in
the kernel; first-to-ask would be the hole again, at boot.
**A configuration question stays open.** Which binary may be given which device is
expressed today as `devices.csv`, read by the manager — configuration, with the manager
as the policy that acts on it. That is already the right shape and needs no new file.
A separate grant manifest earns itself only when something must differ from "the manager
gets the hardware": holding a device back so a test or a bare-metal driver can take it,
for instance. Not now.
## What it does not solve
- **Hot-unplug and re-enumeration drift.** Still the inventory problem (phase 3, and
open question 4 in the track plan). A grant dying with its holder is not the same as a
device going away.
- **The device table's size.** `maximum_devices` is untouched by this; it goes when the
inventory moves in phase 3.
- **Two processes racing for the same root grant.** Cannot arise: the roots are minted
to whichever task satisfies the chain — the manager's binary under PID 1 — and a
second copy spawned by anyone else fails the supervisor half.
## How this is verified
The invariant is **I3** from the track plan: a process holds what it was handed and
cannot name its way into holding more. The test is adversarial and the suite has never
had one of these for devices: a process that was granted nothing calls `device_claim`
on a device another driver owns, and on one nobody owns, and is refused both times with
its own errno. The audit's lesson was that "the suite contains no attacker"; this is the
attacker for devices.