docs: only drivers hold devices; services speak protocols

The design still had the display service receiving a device grant, which
keeps the wrong layering and just moves who hands it over. Drivers talk
hardware. Services talk to no hardware at all — they receive device events
and send data and commands over a protocol, which is the OS abstraction
between them.

So the answer to "which services need device grants" is none, and the design
collapses: every holder of a device is a driver, and every driver is spawned
by the device manager, which is what grants it. No exceptions to
accommodate.

display already shows both halves. Its VirtioGpu backend speaks
scanout-protocol over an IPC handle and touches no device — the driver holds
the hardware, the service speaks to it, and it works today. Its Gop backend
calls device.enumerate, device.claim and device.mmioMap: the service
reaching into hardware itself, because the firmware framebuffer has no
driver to talk to. That is the only such case in the tree; every other
claimant is a device-manager child spawned with its device id.

A firmware-framebuffer driver is therefore a prerequisite of this phase
rather than a consequence. It is small — hold the display node, map the
framebuffer, serve the same scanout-protocol virtio-gpu already serves — and
the compositor needs no new path, since not caring which backend is behind
it is what it was designed for.

Two earlier drafts are recorded as wrong in the document: display asking the
device manager for a grant, and before that init minting grants because init
starts display. Both accommodated an exception instead of removing it.
This commit is contained in:
Daniel Samson
2026-08-08 16:40:37 +01:00
parent 8e388e048a
commit 5ad42ceac3
+50 -28
View File
@@ -64,27 +64,46 @@ The kernel therefore does not need to know about matching, `devices.csv`, driver
or why a device was assigned. It needs to know that an authority it can verify granted
this device to this task.
## The awkward fact that decides the mechanism
## Only drivers hold devices
Five of the six claimants are device-manager children, spawned with their device id in
`argv[1]`: `pci-bus`, `usb-xhci-bus`, `ps2-bus`, `virtio-gpu`, `acpi`.
The layering the rest of the system already follows:
**`display` is not.** It is spawned by `init` from `init.csv`, and it finds its device by
enumerating the table for a display-class node and claiming whatever it finds
([backend.zig](../../system/services/display/backend.zig)). There is no assignment to
enforce, because nobody assigned it anything.
- A **driver** talks to hardware. It is the only thing that holds a device.
- A **service** talks to no hardware at all. It receives device events and sends data
and commands over a **protocol**.
- A **protocol** is the abstraction between them — the OS layer, in the sense of
[communication.md](communication.md).
That rules out the cheapest design. "The kernel records the device named at spawn, and
`device_claim` checks it" closes the hole for five claimants and breaks the sixth. And
the sixth is not an oddity to special-case — it is the one that shows the model is
wrong: authority should be *delegable*, not welded to the moment of spawn.
So the question "which services need device grants" has the answer **none**. That
collapses the design: every holder of a device is a driver, and every driver is spawned
by the device manager, which is what grants it.
It is also what tempted an earlier draft of this document into giving `init` the root
grants, since `init` is what starts `display`. That was wrong for a plainer reason:
`init` has nothing to do with devices. It is the first process and it starts the rest of
the system; the device manager is what owns hardware. Under delegation `display` simply
asks the manager, like every other driver, and the exception disappears rather than
being accommodated.
`display` already demonstrates both halves, one right and one wrong
([backend.zig](../../system/services/display/backend.zig)):
| Backend | How it reaches the panel |
|---|---|
| `VirtioGpu` | speaks `scanout-protocol` over an IPC handle; touches no device |
| `Gop` | `device.enumerate`, `device.claim`, `device.mmioMap` |
The virtio-gpu path is the intended shape and it works today: the driver holds the
hardware, the service speaks a protocol to it. **The GOP path is the anomaly** — the
display service reaching into hardware itself, because the firmware framebuffer has no
driver for it to talk to.
It is also the only anomaly. Every other claimant — `pci-bus`, `usb-xhci-bus`,
`ps2-bus`, `virtio-gpu`, `acpi` — is a device-manager child spawned with its device id
in `argv[1]`.
**A firmware-framebuffer driver is therefore a prerequisite of this phase**, not a
consequence of it. It is small: it holds the display device, maps the framebuffer, and
serves `scanout-protocol` — the same interface `virtio-gpu` already serves, so the
compositor needs no new code path and stops caring which is behind it, which it was
designed for.
An earlier draft of this document had `display` asking the device manager for a grant,
and before that had `init` minting grants because `init` starts `display`. Both were
accommodating an exception instead of removing it.
## The design
@@ -115,10 +134,10 @@ device manager the same way it starts everything else, and knows nothing about d
`init` attests the services it spawned for `/protocol` (`supervisorSatisfies`); it
knows which task is which driver because it spawned them.
3. **`display` asks the manager too.** It is started by `init` as a service, but the
framebuffer is a device, so it receives that grant from the device manager like any
other driver. One authority for hardware, not two — which is what makes `display`
stop being the exception that broke the simpler design.
3. **Services receive nothing.** `display`, `input`, `fat` and `logger` hold no devices
and need no grants; they speak protocols to the drivers that do. The one place this
is not true today is the GOP backend above, which a firmware-framebuffer driver
removes.
4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and `dma_bind`
check possession of the grant instead of consulting an ownership table.
@@ -137,13 +156,16 @@ own comment says "Claim the bridge, map the ECAM, hello the manager, then scan."
delegation the hello must come first, because that is where the grant arrives. Five
drivers need that reordering, and it is the bulk of the work.
**`display` gains a dependency on the device manager.** It currently finds its
framebuffer by itself and needs nothing from anyone; afterwards it must ask. That is a
new ordering constraint at boot — `display` cannot bring up a screen until the manager
is up — and it is the honest price of there being one authority for hardware. Mitigating:
`display` already tolerates arriving before its backend (the virtio-gpu driver hands it
a scanout later, over `attach_scanout`), so the machinery for "wait, then attach" is
there.
**A firmware-framebuffer driver has to be written first**, and until it exists the
display service cannot stop claiming a device. It is the smallest new binary in the
tree — hold the display node, map the framebuffer, serve `scanout-protocol` — but it is
new code on the boot path, and the boot path is where a mistake costs a screen.
It brings an ordering constraint with it: `display` cannot paint until that driver is
up, where today it maps the framebuffer itself and is independent. Mitigating,
`display` already tolerates a backend arriving late — `virtio-gpu` hands it a scanout
after the fact over `attach_scanout` — so "wait, then attach" is a path that already
works rather than one to invent.
**The kernel gains one piece of knowledge about a specific binary.** Chain attestation
means the kernel recognises `/system/services/device-manager` under PID 1 as the root