docs: device authority — the how, and the run that deletes the ceilings
device-authority.md is rewritten as an implementation design rather than a rival to device-manager.md. The what was already settled there in 2026-07: structure in the manager, authority in the kernel, and delegation as the step after hello. This is the how, plus the two decisions that paragraph leaves open. Decision 1: the manager claims, it is not granted. device-manager.md says "claims (or is granted)"; claiming wins because the manager runs before any driver exists and takes the seeded devices unopposed, leaving nothing unheld to race for. One new call, device_transfer(device_id, task_id), checks only that the caller holds the device — no names in the kernel, no attestation. The alternative put a binary path inside the kernel, and the kernel should hold only what cannot safely live in user space. The residual is stated rather than hidden: authority rests on the manager claiming first, which init.csv makes an operator-visible ordering rather than an attacker- controlled one, and the enforced version arrives with the spawn capability drivers.md already names as missing. Decision 2: the kernel stops holding inventory. It reads three things out of a descriptor — physical ranges, interrupt numbers, one PCI BDF — and stores the rest only so device_enumerate can hand it back. Devices with no resources leave the kernel entirely: a USB device conveys no mapping authority, so there is nothing to enforce. That is also the case which sidesteps containment, and therefore the reason a shared cap existed. Decision 3: no shared ceiling. The table becomes dynamic — it is built after heap.init, so nothing ever prevented it — and the two invented numbers go. A per-holder quota replaces them, because dynamic storage with no bound moves the ceiling to the kernel heap, which is shared and fatal rather than partial. A bound charged to whoever caused it is isolation. Two earlier drafts of this document are gone: one gave init the root grants, the other proposed extracting a firmware-framebuffer driver. Both were wrong and both are recorded as wrong in the run plan's settled list — the framebuffer is not a device, it is where pixels go until a real display driver announces itself. Run 2 is nine steps, ordered so the suite stays green throughout: build and prove the transfer mechanism, move the five claimants across one at a time, then the flag day, then the inventory, then the ceilings.
This commit is contained in:
@@ -1,12 +1,25 @@
|
||||
# Device authority: you hold what you were given
|
||||
# Device authority: the implementation of delegation
|
||||
|
||||
*Design for phase 2 of [the bounds track](../bounds-track-plan.md), 2026-08-08.
|
||||
Supersedes an earlier draft that argued for capabilities on aesthetic grounds; this one
|
||||
starts from the hole and from the project's principles.*
|
||||
*Implementation design, 2026-08-08. The **what** is settled in
|
||||
[device-manager.md](../device-driver-development/device-manager.md) — "structure in the
|
||||
manager, authority in the kernel", and delegation as the step after `hello`. This
|
||||
document is the **how**, and the decisions that paragraph leaves open.*
|
||||
|
||||
## The hole
|
||||
Read first: [drivers.md](../device-driver-development/drivers.md) (the claim is the
|
||||
capability), [driver-model.md](../device-driver-development/driver-model.md) (the three
|
||||
invariants), [device-manager.md](../device-driver-development/device-manager.md) (the
|
||||
tree, the matcher, the supervisor).
|
||||
|
||||
`device_claim(id)` checks two things ([devices-broker.zig](../../system/kernel/devices-broker.zig)):
|
||||
## The one thing not yet true
|
||||
|
||||
`device-manager.md` says assignment "stays argv for now", and names the next step:
|
||||
|
||||
> The step after `hello` exists is delegation: the manager claims (or is granted) the
|
||||
> devices and passes the claim to the driver over IPC (the M13 capability-transfer
|
||||
> mechanism), replacing first-come-first-served `device_claim` with policy.
|
||||
|
||||
Until that lands, the manager's matching is advisory. `device_claim` checks only that
|
||||
the device exists and is unheld ([devices-broker.zig](../../system/kernel/devices-broker.zig)):
|
||||
|
||||
```zig
|
||||
pub fn claim(id: u64, owner: u32) ClaimError!void {
|
||||
@@ -16,187 +29,123 @@ pub fn claim(id: u64, owner: u32) ClaimError!void {
|
||||
}
|
||||
```
|
||||
|
||||
Does it exist, and is it free. **Any process may claim any unclaimed device.**
|
||||
A driver is spawned with its device id in `argv[1]` and claims it; any process could
|
||||
pass any integer instead. Since a claim is a licence to map physical memory, that is the
|
||||
gap this document closes.
|
||||
|
||||
The matching is real but it is entirely advisory: `device-manager` reads `devices.csv`,
|
||||
matches a device to a driver, and spawns that driver with the device id as `argv[1]`
|
||||
(`spawnDriver`). The driver parses the string and claims it. Nothing anywhere binds the
|
||||
manager's decision to the kernel's grant — a process can pass any integer and win the
|
||||
race.
|
||||
## Decision 1: the manager claims, then transfers
|
||||
|
||||
A claim is not a small thing. It is what gates `mmio_map` and `irq_bind`, so it is a
|
||||
licence to map physical memory and receive interrupts.
|
||||
`device-manager.md` leaves "claims (or is granted)" open. **Claims.**
|
||||
|
||||
## What this costs, beyond the obvious
|
||||
The manager runs before any driver exists — `init` starts it from `init.csv`, and it is
|
||||
what spawns drivers — so it takes the seeded devices unopposed and there is nothing
|
||||
unheld left for anyone to race for. One new call moves ownership on:
|
||||
|
||||
`maximum_children_per_parent = 16` exists because a driver that claimed one device could
|
||||
loop `device_register` under it and exhaust the shared table. That threat only exists
|
||||
*because* claiming is unauthenticated — and the cap is a poor defence against it, since
|
||||
an attacker can burn 16 slots, claim another device, and burn 16 more. What it reliably
|
||||
does instead is refuse a legitimate PCI bus with more than 16 functions, which is how an
|
||||
AMD Ryzen came to boot with no USB and no storage.
|
||||
```
|
||||
device_transfer(device_id, task_id) -> 0/-errno
|
||||
```
|
||||
|
||||
So the cap is not merely mis-sized. It is standing in for an authorisation that is not
|
||||
performed, and it punishes correct behaviour while barely inconveniencing incorrect
|
||||
behaviour. **Closing the hole is what retires the constant**, not a bigger number.
|
||||
The kernel checks only that the caller currently holds the device. No names, no
|
||||
attestation, no notion of "the device manager" — the rule is *you may give away what you
|
||||
hold*, which is the capability discipline already in force.
|
||||
|
||||
## What the principles decide
|
||||
The alternative was the kernel granting roots to a task it recognises by binary path
|
||||
plus a PID-1 supervisor. It is more robust — it does not depend on the manager being
|
||||
first — but it puts a binary name inside the kernel, and the goal is that the kernel
|
||||
keeps only what cannot safely live in user space. A name is not that.
|
||||
|
||||
- *Move as much responsibility as possible to user space* (3), and *what remains in the
|
||||
kernel is there for security or a hardware limitation* (5).
|
||||
**The residual, stated plainly:** authority here rests on the manager claiming first.
|
||||
That holds because `init.csv` decides what starts and in what order, so it is an
|
||||
operator-visible ordering rather than an attacker-controlled one — but it is an
|
||||
assumption, not an enforced invariant. The enforced version arrives with the spawn
|
||||
capability [drivers.md](../device-driver-development/drivers.md) already names as
|
||||
missing ("`system_spawn` is currently ungated … because there is no spawn capability
|
||||
yet"). This design is compatible with it and does not block on it.
|
||||
|
||||
Deciding **which** driver gets **which** device is policy — matching identity triples
|
||||
and choosing a binary. That decision belongs to the device manager and stays there.
|
||||
`devices.csv` is not the policy; it is **configuration**, the declarative data the
|
||||
policy reads. Three distinct things, and worth keeping apart in this document:
|
||||
## Decision 2: the kernel stops holding inventory
|
||||
|
||||
| | Lives in | Example |
|
||||
The kernel reads exactly three things out of a descriptor: **physical ranges** (to check
|
||||
a mapping falls inside one), **interrupt numbers**, and **one PCI BDF** (to key an IOMMU
|
||||
domain). Vendor, device and subsystem ids, class triples, `_HID` strings, bus addresses
|
||||
and names are stored only so `device_enumerate` can hand them back — which
|
||||
`device-manager.md` already resolves: that call "fades to a manager-internal (then
|
||||
deleted) seam", because the manager owns the tree as data.
|
||||
|
||||
So the kernel's table becomes: **parent, resources, holder, BDF.** That is what cannot
|
||||
safely run in user space; the rest moves.
|
||||
|
||||
**Devices with no resources leave the kernel entirely.** A USB device is addressed
|
||||
through its controller and carries `resource_count = 0`
|
||||
([driver-model.md](../device-driver-development/driver-model.md): "that case is allowed
|
||||
and is the common one"). It conveys no mapping authority, so there is nothing for the
|
||||
kernel to enforce and no reason for it to know. It is inventory, and inventory is the
|
||||
manager's — reported by `child_added`, which already carries everything needed.
|
||||
|
||||
That is also the case that made `maximum_children_per_parent` necessary: a zero-resource
|
||||
child sidesteps containment, so a driver could loop `device_register` and fill the
|
||||
shared table. Once such children are not kernel objects, every remaining entry is a real
|
||||
contained subdivision of something the caller holds.
|
||||
|
||||
## Decision 3: no shared ceiling; a per-holder quota instead
|
||||
|
||||
`maximum_devices = 64` and `maximum_children_per_parent = 16` are numbers we invented,
|
||||
and both are shared — one driver's enumeration starves every other driver, which is how
|
||||
an AMD Ryzen booted with no USB and no storage.
|
||||
|
||||
- **The table becomes dynamic.** It is built after `heap.init` (`kernel.zig`: `pmm.init`
|
||||
at 137, `heap.init` at 179, `devices_broker.init` at 202), so nothing prevents it. No
|
||||
specification bounds how many devices a machine has, so nothing should bound ours.
|
||||
- **`maximum_children_per_parent` is deleted**, because the authorisation it stood in
|
||||
for now exists.
|
||||
- **A per-holder quota replaces them.** Dynamic storage without a bound moves the
|
||||
ceiling to the kernel heap, which is shared and fatal rather than partial — strictly
|
||||
worse. The bound that is *not* worse is one charged to the task that caused it: a
|
||||
driver that loops `device_register` exhausts its own allowance, is refused with an
|
||||
attributable errno, and is restarted by its supervisor while every other driver
|
||||
carries on. That is the microkernel property rather than a workaround for it, and it
|
||||
is declared through [bounds.md](bounds.md) like any other.
|
||||
|
||||
## The shape of the change
|
||||
|
||||
| | Before | After |
|
||||
|---|---|---|
|
||||
| **Mechanism** | the kernel | the check that a grant is held before a mapping is made |
|
||||
| **Policy** | user space | the device manager matching a device to a driver |
|
||||
| **Configuration** | files | `devices.csv`, `protocol.csv`, `init.csv` |
|
||||
| Manager gets its devices | claims them, unauthorised | claims them (first, unopposed) |
|
||||
| Driver gets its device | `argv[1]` + `device_claim` | receives it in the `hello` reply |
|
||||
| Kernel checks | is it free? | do you hold it? |
|
||||
| Kernel stores | the full descriptor | parent, resources, holder, BDF |
|
||||
| Zero-resource devices | kernel table entries | manager records only |
|
||||
| Table size | `maximum_devices = 64` | dynamic, per-holder quota |
|
||||
| Children per parent | `maximum_children_per_parent = 16` | deleted |
|
||||
|
||||
Enforcing that a driver **holds only what it was given** is security — it is the gate in
|
||||
front of mapping physical memory. That stays in the kernel, and it is the whole of what
|
||||
the kernel needs to do.
|
||||
Bring-up order changes for the five claiming drivers: `hello` must precede the claim,
|
||||
because the reply is where the device arrives. `pci-bus` today does the reverse — its
|
||||
own comment reads "Claim the bridge, map the ECAM, hello the manager, then scan."
|
||||
|
||||
The kernel therefore does not need to know about matching, `devices.csv`, driver names,
|
||||
or why a device was assigned. It needs to know that an authority it can verify granted
|
||||
this device to this task.
|
||||
## What does not change
|
||||
|
||||
## Only drivers hold devices
|
||||
- The three invariants of [driver-model.md](../device-driver-development/driver-model.md):
|
||||
a claim is exclusive, a descriptor is a licence to map physical memory, therefore
|
||||
containment. This design strengthens the first and touches neither of the others.
|
||||
- The display service's GOP path. The framebuffer is not a device — it is where pixels
|
||||
go, handed over by the loader, and the compositor uses it as the boot floor until a
|
||||
real display driver announces itself
|
||||
([display-v2.md](../device-driver-development/display-v2.md)). The kernel wraps it in
|
||||
a display-class descriptor so `mmio_map` can hand it over write-combining; that is
|
||||
plumbing for a mapping, not a claim about what it is.
|
||||
- Supervision, restart, pruning and re-report
|
||||
([device-manager.md](../device-driver-development/device-manager.md),
|
||||
[process-lifecycle.md](process-lifecycle.md)). Delegation slots into the existing
|
||||
`hello` exchange and changes none of it.
|
||||
- `device_register` idempotency, which is what lets a restarted bus rebuild the same
|
||||
ids.
|
||||
|
||||
The layering the rest of the system already follows:
|
||||
## How it is verified
|
||||
|
||||
- A **driver** talks to hardware. It is the only thing that holds a device.
|
||||
- A **service** talks to no hardware at all. It receives device events and sends data
|
||||
and commands over a **protocol**.
|
||||
- A **protocol** is the abstraction between them — the OS layer, in the sense of
|
||||
[communication.md](communication.md).
|
||||
The invariant is: **a process holds what it was handed and cannot name its way into
|
||||
holding more.** The suite has no adversarial device case today — the audit's lesson was
|
||||
that "the suite contains no attacker" — so this adds one: a process that was handed
|
||||
nothing calls `device_claim` and `device_transfer` on a device another driver holds, and
|
||||
on one nobody holds, and is refused each time with its own errno.
|
||||
|
||||
So the question "which services need device grants" has the answer **none**. That
|
||||
collapses the design: every holder of a device is a driver, and every driver is spawned
|
||||
by the device manager, which is what grants it.
|
||||
|
||||
`display` already demonstrates both halves, one right and one wrong
|
||||
([backend.zig](../../system/services/display/backend.zig)):
|
||||
|
||||
| Backend | How it reaches the panel |
|
||||
|---|---|
|
||||
| `VirtioGpu` | speaks `scanout-protocol` over an IPC handle; touches no device |
|
||||
| `Gop` | `device.enumerate`, `device.claim`, `device.mmioMap` |
|
||||
|
||||
The virtio-gpu path is the intended shape and it works today: the driver holds the
|
||||
hardware, the service speaks a protocol to it. **The GOP path is the anomaly** — the
|
||||
display service reaching into hardware itself, because the firmware framebuffer has no
|
||||
driver for it to talk to.
|
||||
|
||||
It is also the only anomaly. Every other claimant — `pci-bus`, `usb-xhci-bus`,
|
||||
`ps2-bus`, `virtio-gpu`, `acpi` — is a device-manager child spawned with its device id
|
||||
in `argv[1]`.
|
||||
|
||||
**A firmware-framebuffer driver is therefore a prerequisite of this phase**, not a
|
||||
consequence of it. It is small: it holds the display device, maps the framebuffer, and
|
||||
serves `scanout-protocol` — the same interface `virtio-gpu` already serves, so the
|
||||
compositor needs no new code path and stops caring which is behind it, which it was
|
||||
designed for.
|
||||
|
||||
An earlier draft of this document had `display` asking the device manager for a grant,
|
||||
and before that had `init` minting grants because `init` starts `display`. Both were
|
||||
accommodating an exception instead of removing it.
|
||||
|
||||
## The design
|
||||
|
||||
**A device grant is a capability, delegated from a holder.** The mechanism already
|
||||
exists: `callCap` passes a handle over an IPC call and the kernel installs it in the
|
||||
receiver's table ([library/kernel/ipc.zig](../../library/kernel/ipc.zig)), which is how
|
||||
shared memory and DMA regions already move between processes.
|
||||
|
||||
**The device manager is the root of device authority.** `init` has no part in this: it
|
||||
is the first process, and its job is to start the rest of the system. It starts the
|
||||
device manager the same way it starts everything else, and knows nothing about devices.
|
||||
|
||||
1. **Root.** At boot the kernel mints grants for the devices firmware discovery found
|
||||
and hands them to the device manager. This is the only place device authority enters
|
||||
the system, and it comes from ACPI rather than from anyone's say-so.
|
||||
|
||||
The kernel recognises the manager by **chain attestation**, the identity the security
|
||||
track already settled on: the binary path the kernel itself stamped
|
||||
(`/system/services/device-manager`) *and* a supervisor of PID 1. The binary path
|
||||
alone would not do — `system_spawn` is deliberately ungated, so any process may spawn
|
||||
any bundled binary, and a rogue could run a second copy under the same name. It could
|
||||
not forge the other half: its copy's supervisor is the rogue, and PID 1 is the
|
||||
kernel's own first process.
|
||||
|
||||
2. **Delegation.** The manager passes a driver its device when it spawns it, over the
|
||||
channel that already exists — the driver `hello`s the manager and the reply carries
|
||||
the grant. The manager attests its own children by task id, one hop deep, exactly as
|
||||
`init` attests the services it spawned for `/protocol` (`supervisorSatisfies`); it
|
||||
knows which task is which driver because it spawned them.
|
||||
|
||||
3. **Services receive nothing.** `display`, `input`, `fat` and `logger` hold no devices
|
||||
and need no grants; they speak protocols to the drivers that do. The one place this
|
||||
is not true today is the GOP backend above, which a firmware-framebuffer driver
|
||||
removes.
|
||||
|
||||
4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and `dma_bind`
|
||||
check possession of the grant instead of consulting an ownership table.
|
||||
|
||||
Exclusivity stops being a broker refusing a second claimant and becomes the ordinary
|
||||
property of a capability: only one process was given it.
|
||||
|
||||
**`maximum_children_per_parent` is deleted here.** After this a bus driver's children are
|
||||
devices it enumerated on a bus it was actually given, and the threat the cap was written
|
||||
for no longer exists.
|
||||
|
||||
## What this costs
|
||||
|
||||
**Bring-up order changes.** `pci-bus` today claims first and says hello afterwards — its
|
||||
own comment says "Claim the bridge, map the ECAM, hello the manager, then scan." Under
|
||||
delegation the hello must come first, because that is where the grant arrives. Five
|
||||
drivers need that reordering, and it is the bulk of the work.
|
||||
|
||||
**A firmware-framebuffer driver has to be written first**, and until it exists the
|
||||
display service cannot stop claiming a device. It is the smallest new binary in the
|
||||
tree — hold the display node, map the framebuffer, serve `scanout-protocol` — but it is
|
||||
new code on the boot path, and the boot path is where a mistake costs a screen.
|
||||
|
||||
It brings an ordering constraint with it: `display` cannot paint until that driver is
|
||||
up, where today it maps the framebuffer itself and is independent. Mitigating,
|
||||
`display` already tolerates a backend arriving late — `virtio-gpu` hands it a scanout
|
||||
after the fact over `attach_scanout` — so "wait, then attach" is a path that already
|
||||
works rather than one to invent.
|
||||
|
||||
**The kernel gains one piece of knowledge about a specific binary.** Chain attestation
|
||||
means the kernel recognises `/system/services/device-manager` under PID 1 as the root
|
||||
holder. That is a real concession — the kernel would rather know nothing about who is
|
||||
who — and it is the minimum: device authority has to enter the system somewhere, and
|
||||
every alternative is worse. Configuration the kernel reads would put a file parser in
|
||||
the kernel; first-to-ask would be the hole again, at boot.
|
||||
|
||||
**A configuration question stays open.** Which binary may be given which device is
|
||||
expressed today as `devices.csv`, read by the manager — configuration, with the manager
|
||||
as the policy that acts on it. That is already the right shape and needs no new file.
|
||||
A separate grant manifest earns itself only when something must differ from "the manager
|
||||
gets the hardware": holding a device back so a test or a bare-metal driver can take it,
|
||||
for instance. Not now.
|
||||
|
||||
## What it does not solve
|
||||
|
||||
- **Hot-unplug and re-enumeration drift.** Still the inventory problem (phase 3, and
|
||||
open question 4 in the track plan). A grant dying with its holder is not the same as a
|
||||
device going away.
|
||||
- **The device table's size.** `maximum_devices` is untouched by this; it goes when the
|
||||
inventory moves in phase 3.
|
||||
- **Two processes racing for the same root grant.** Cannot arise: the roots are minted
|
||||
to whichever task satisfies the chain — the manager's binary under PID 1 — and a
|
||||
second copy spawned by anyone else fails the supervisor half.
|
||||
|
||||
## How this is verified
|
||||
|
||||
The invariant is **I3** from the track plan: a process holds what it was handed and
|
||||
cannot name its way into holding more. The test is adversarial and the suite has never
|
||||
had one of these for devices: a process that was granted nothing calls `device_claim`
|
||||
on a device another driver owns, and on one nobody owns, and is refused both times with
|
||||
its own errno. The audit's lesson was that "the suite contains no attacker"; this is the
|
||||
attacker for devices.
|
||||
The Ryzen is the acceptance test for the ceiling half: it is the machine that found the
|
||||
constants, and the one that proves them gone.
|
||||
|
||||
Reference in New Issue
Block a user