docs: the device manager is the root of device authority, not init

The design had init minting device grants and handing them on, because init
is what starts display and display claims the framebuffer. That was the
wrong shape. init has nothing to do with devices: it is the first process
and it starts the rest of the system. The device manager owns hardware.

So the kernel mints the root grants to the device manager, and display asks
the manager for its framebuffer like any other driver. The exception that
drove the earlier draft disappears instead of being accommodated — one
authority for hardware, not two.

The kernel recognises the manager by the chain attestation the security
track already settled on: the binary path it stamped itself, AND a
supervisor of PID 1. Path alone is forgeable, because system_spawn is
deliberately ungated and any process may spawn any bundled binary — but a
rogue copy's supervisor is the rogue, and PID 1 is the kernel's own first
process.

Costs restated for the corrected model. display gains a boot-order
dependency on the manager, mitigated by its existing ability to attach a
backend late. The kernel gains one piece of knowledge about one binary,
which is the minimum: authority has to enter somewhere, and the
alternatives are a file parser in the kernel or first-to-ask, which is the
hole again.
This commit is contained in:
Daniel Samson
2026-08-08 16:35:32 +01:00
parent 139c4f624f
commit ef793d1320
+52 -25
View File
@@ -79,6 +79,13 @@ That rules out the cheapest design. "The kernel records the device named at spaw
the sixth is not an oddity to special-case — it is the one that shows the model is
wrong: authority should be *delegable*, not welded to the moment of spawn.
It is also what tempted an earlier draft of this document into giving `init` the root
grants, since `init` is what starts `display`. That was wrong for a plainer reason:
`init` has nothing to do with devices. It is the first process and it starts the rest of
the system; the device manager is what owns hardware. Under delegation `display` simply
asks the manager, like every other driver, and the exception disappears rather than
being accommodated.
## The design
**A device grant is a capability, delegated from a holder.** The mechanism already
@@ -86,18 +93,33 @@ exists: `callCap` passes a handle over an IPC call and the kernel installs it in
receiver's table ([library/kernel/ipc.zig](../../library/kernel/ipc.zig)), which is how
shared memory and DMA regions already move between processes.
**The device manager is the root of device authority.** `init` has no part in this: it
is the first process, and its job is to start the rest of the system. It starts the
device manager the same way it starts everything else, and knows nothing about devices.
1. **Root.** At boot the kernel mints grants for the devices firmware discovery found
and hands them to `init` (PID 1, which the kernel spawns and therefore need not
authenticate). This is the only place device authority enters the system, and it
comes from ACPI rather than from anyone's say-so.
2. **Delegation.** `init` passes the device manager the grants it will need — in
practice all of them — and passes `display` the framebuffer grant, because `init` is
what starts `display`. This is the same shape as the `/protocol` registry, where
`init` is already the grantor and `protocol.csv` already records
`/system/services/device-manager, /system/services/init, bind, device-manager`.
3. **Assignment.** The manager passes a driver its device when it spawns it, over the
channel that already exists — the driver `hello`s the manager, and the reply carries
the grant.
and hands them to the device manager. This is the only place device authority enters
the system, and it comes from ACPI rather than from anyone's say-so.
The kernel recognises the manager by **chain attestation**, the identity the security
track already settled on: the binary path the kernel itself stamped
(`/system/services/device-manager`) *and* a supervisor of PID 1. The binary path
alone would not do — `system_spawn` is deliberately ungated, so any process may spawn
any bundled binary, and a rogue could run a second copy under the same name. It could
not forge the other half: its copy's supervisor is the rogue, and PID 1 is the
kernel's own first process.
2. **Delegation.** The manager passes a driver its device when it spawns it, over the
channel that already exists — the driver `hello`s the manager and the reply carries
the grant. The manager attests its own children by task id, one hop deep, exactly as
`init` attests the services it spawned for `/protocol` (`supervisorSatisfies`); it
knows which task is which driver because it spawned them.
3. **`display` asks the manager too.** It is started by `init` as a service, but the
framebuffer is a device, so it receives that grant from the device manager like any
other driver. One authority for hardware, not two — which is what makes `display`
stop being the exception that broke the simpler design.
4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and `dma_bind`
check possession of the grant instead of consulting an ownership table.
@@ -115,22 +137,27 @@ own comment says "Claim the bridge, map the ECAM, hello the manager, then scan."
delegation the hello must come first, because that is where the grant arrives. Five
drivers need that reordering, and it is the bulk of the work.
**`init` grows a device role.** It already registers `/protocol` and reads
`protocol.csv`; it would also hold root device grants and hand them on. That is more
responsibility in PID 1, which is a cost worth naming — though the alternative is the
kernel deciding who may hold what, which principle 5 excludes.
**`display` gains a dependency on the device manager.** It currently finds its
framebuffer by itself and needs nothing from anyone; afterwards it must ask. That is a
new ordering constraint at boot — `display` cannot bring up a screen until the manager
is up — and it is the honest price of there being one authority for hardware. Mitigating:
`display` already tolerates arriving before its backend (the virtio-gpu driver hands it
a scanout later, over `attach_scanout`), so the machinery for "wait, then attach" is
there.
**A configuration question follows.** Who may bind which protocol name is expressed as
configuration today — `protocol.csv` — with `init` as the policy that reads it. Device
grants could be expressed the same way: a file saying which binary may be given which
device class, with `init` again the policy that enforces it.
**The kernel gains one piece of knowledge about a specific binary.** Chain attestation
means the kernel recognises `/system/services/device-manager` under PID 1 as the root
holder. That is a real concession — the kernel would rather know nothing about who is
who — and it is the minimum: device authority has to enter the system somewhere, and
every alternative is worse. Configuration the kernel reads would put a file parser in
the kernel; first-to-ask would be the hole again, at boot.
That symmetry is real but it is *not* proposed here. The smaller step is `init` handing
the manager the root grants and the manager matching by `devices.csv`, which is
configuration it already reads. A device-grant file adds a second place an operator must
keep correct, and it should only arrive when something needs it to differ from "the
manager gets the hardware" — for example, holding a device back from the manager so a
test or a bare-metal driver can take it.
**A configuration question stays open.** Which binary may be given which device is
expressed today as `devices.csv`, read by the manager — configuration, with the manager
as the policy that acts on it. That is already the right shape and needs no new file.
A separate grant manifest earns itself only when something must differ from "the manager
gets the hardware": holding a device back so a test or a bare-metal driver can take it,
for instance. Not now.
## What it does not solve