docs: device authority — the how, and the run that deletes the ceilings

device-authority.md is rewritten as an implementation design rather than a
rival to device-manager.md. The what was already settled there in 2026-07:
structure in the manager, authority in the kernel, and delegation as the
step after hello. This is the how, plus the two decisions that paragraph
leaves open.

Decision 1: the manager claims, it is not granted. device-manager.md says
"claims (or is granted)"; claiming wins because the manager runs before any
driver exists and takes the seeded devices unopposed, leaving nothing unheld
to race for. One new call, device_transfer(device_id, task_id), checks only
that the caller holds the device — no names in the kernel, no attestation.
The alternative put a binary path inside the kernel, and the kernel should
hold only what cannot safely live in user space. The residual is stated
rather than hidden: authority rests on the manager claiming first, which
init.csv makes an operator-visible ordering rather than an attacker-
controlled one, and the enforced version arrives with the spawn capability
drivers.md already names as missing.

Decision 2: the kernel stops holding inventory. It reads three things out of
a descriptor — physical ranges, interrupt numbers, one PCI BDF — and stores
the rest only so device_enumerate can hand it back. Devices with no
resources leave the kernel entirely: a USB device conveys no mapping
authority, so there is nothing to enforce. That is also the case which
sidesteps containment, and therefore the reason a shared cap existed.

Decision 3: no shared ceiling. The table becomes dynamic — it is built after
heap.init, so nothing ever prevented it — and the two invented numbers go. A
per-holder quota replaces them, because dynamic storage with no bound moves
the ceiling to the kernel heap, which is shared and fatal rather than
partial. A bound charged to whoever caused it is isolation.

Two earlier drafts of this document are gone: one gave init the root grants,
the other proposed extracting a firmware-framebuffer driver. Both were wrong
and both are recorded as wrong in the run plan's settled list — the
framebuffer is not a device, it is where pixels go until a real display
driver announces itself.

Run 2 is nine steps, ordered so the suite stays green throughout: build and
prove the transfer mechanism, move the five claimants across one at a time,
then the flag day, then the inventory, then the ceilings.
This commit is contained in:
Daniel Samson
2026-08-08 17:01:25 +01:00
parent 5ad42ceac3
commit 547d0ec46b
2 changed files with 181 additions and 175 deletions
+58 -1
View File
@@ -19,11 +19,68 @@ next one starts.*
| L5 | USB: interfaces from the descriptor, and the misattributed-endpoint bug | **done** — fix is by construction; no direct test, see open question 5 | | L5 | USB: interfaces from the descriptor, and the misattributed-endpoint bug | **done** — fix is by construction; no direct test, see open question 5 |
| L6 | xHCI: a failed `allocateDevice` stops leaking an enabled slot | **done** — path forced and verified; no regression test, see open question 5 | | L6 | xHCI: a failed `allocateDevice` stops leaking an enabled slot | **done** — path forced and verified; no regression test, see open question 5 |
**Run complete.** L1 stopped (the step was wrong), L2–L6 landed. Suite 115 → 116. **Run 1 complete.** L1 stopped (the step was wrong), L2–L6 landed. Suite 115 → 116.
Allowlist 278 → 269. Two steps ship without a permanent regression test, both because Allowlist 278 → 269. Two steps ship without a permanent regression test, both because
QEMU's USB devices are too small to reach the paths — see open question 5, which is the QEMU's USB devices are too small to reach the paths — see open question 5, which is the
audit's own lesson recurring: the test rig is smaller than a real machine. audit's own lesson recurring: the test rig is smaller than a real machine.
---
## Run 2 — device authority: delete the invented ceilings
*Design: [device-authority.md](os-development/device-authority.md), which is the **how**
for the delegation step [device-manager.md](device-driver-development/device-manager.md)
already settled. Read both before starting; the second is authoritative where they
differ.*
The goal, in the project owner's words: **remove the maximum values we set arbitrarily,
move the responsibility to the device manager, and keep in the kernel only the parts
that cannot safely run in user space.**
| Step | What | State |
|---|---|---|
| D1 | `device_transfer(device_id, task_id)` — the holder gives a device away | not started |
| D2 | Adversarial case: a process handed nothing is refused, on a held device and a free one | not started |
| D3 | The manager claims the seeded devices at boot, before any driver is spawned | not started |
| D4 | `usb-xhci-bus` receives its controller in the `hello` reply instead of claiming argv[1] | not started |
| D5 | The other four claimants converted: `pci-bus`, `ps2-bus`, `virtio-gpu`, `acpi` | not started |
| D6 | `device_claim` refuses a device the caller was not handed; the hole is closed | not started |
| D7 | Zero-resource devices stop being kernel objects — inventory moves to the manager | not started |
| D8 | **`maximum_children_per_parent` deleted** — the authorisation it stood in for exists | not started |
| D9 | The device table becomes dynamic; **`maximum_devices` deleted**; per-holder quota declared | not started |
Ordering is load-bearing. D1–D2 build and prove the mechanism with nothing depending on
it. D3–D5 move each claimant across one at a time, so the suite stays green throughout
and a regression names the driver that caused it. D6 is the flag day. D7 must precede
D9, because zero-resource children are the case that sidesteps containment and so the
reason a shared cap was needed at all.
### Settled, so the run does not re-litigate them
- **The manager claims, it is not granted.** No binary names in the kernel; the rule is
"you may give away what you hold". The residual — it rests on the manager claiming
first — is stated in the design and is closed later by the spawn capability
[drivers.md](device-driver-development/drivers.md) already names as missing.
- **The framebuffer is not a device.** It is where pixels go, handed over by the loader,
and the compositor uses it as the boot floor until a real display driver announces
itself. Nothing in this run touches the display service or its GOP path.
- **`maximum_endpoints_per_interface` and the wire structs stay.** Widening them is a
protocol change, out of scope.
- **A per-holder quota is not a retreat.** Dynamic storage with no bound moves the
ceiling to the kernel heap, which is shared and fatal rather than partial. A bound
charged to the task that caused it is isolation, and it is declared through
[bounds.md](os-development/bounds.md) like anything else.
### Working rules
As Run 1, unchanged: work in `/Users/danielsamson/Gitea/daniel/danos` on
`claude/bounds-track`; every step lands with a test that fails before the fix, verified
by restoring the old behaviour; full suite green before each commit; never two suites at
once (`pgrep -f qemu_test.py`); 60 GiB free; `git commit -F` with no `Co-Authored-By`;
update this table before starting the next step. **If a step needs a decision that is
not written down, stop it, add the question below, and move on** — Run 1's first step
was wrong and stopping was the right call.
**Suite:** 115/115 at the start of the run. **Suite:** 115/115 at the start of the run.
**Branch:** `claude/bounds-track`. **Branch:** `claude/bounds-track`.
+123 -174
View File
@@ -1,12 +1,25 @@
# Device authority: you hold what you were given # Device authority: the implementation of delegation
*Design for phase 2 of [the bounds track](../bounds-track-plan.md), 2026-08-08. *Implementation design, 2026-08-08. The **what** is settled in
Supersedes an earlier draft that argued for capabilities on aesthetic grounds; this one [device-manager.md](../device-driver-development/device-manager.md) — "structure in the
starts from the hole and from the project's principles.* manager, authority in the kernel", and delegation as the step after `hello`. This
document is the **how**, and the decisions that paragraph leaves open.*
## The hole Read first: [drivers.md](../device-driver-development/drivers.md) (the claim is the
capability), [driver-model.md](../device-driver-development/driver-model.md) (the three
invariants), [device-manager.md](../device-driver-development/device-manager.md) (the
tree, the matcher, the supervisor).
`device_claim(id)` checks two things ([devices-broker.zig](../../system/kernel/devices-broker.zig)): ## The one thing not yet true
`device-manager.md` says assignment "stays argv for now", and names the next step:
> The step after `hello` exists is delegation: the manager claims (or is granted) the
> devices and passes the claim to the driver over IPC (the M13 capability-transfer
> mechanism), replacing first-come-first-served `device_claim` with policy.
Until that lands, the manager's matching is advisory. `device_claim` checks only that
the device exists and is unheld ([devices-broker.zig](../../system/kernel/devices-broker.zig)):
```zig ```zig
pub fn claim(id: u64, owner: u32) ClaimError!void { pub fn claim(id: u64, owner: u32) ClaimError!void {
@@ -16,187 +29,123 @@ pub fn claim(id: u64, owner: u32) ClaimError!void {
} }
``` ```
Does it exist, and is it free. **Any process may claim any unclaimed device.** A driver is spawned with its device id in `argv[1]` and claims it; any process could
pass any integer instead. Since a claim is a licence to map physical memory, that is the
gap this document closes.
The matching is real but it is entirely advisory: `device-manager` reads `devices.csv`, ## Decision 1: the manager claims, then transfers
matches a device to a driver, and spawns that driver with the device id as `argv[1]`
(`spawnDriver`). The driver parses the string and claims it. Nothing anywhere binds the
manager's decision to the kernel's grant — a process can pass any integer and win the
race.
A claim is not a small thing. It is what gates `mmio_map` and `irq_bind`, so it is a `device-manager.md` leaves "claims (or is granted)" open. **Claims.**
licence to map physical memory and receive interrupts.
## What this costs, beyond the obvious The manager runs before any driver exists — `init` starts it from `init.csv`, and it is
what spawns drivers — so it takes the seeded devices unopposed and there is nothing
unheld left for anyone to race for. One new call moves ownership on:
`maximum_children_per_parent = 16` exists because a driver that claimed one device could ```
loop `device_register` under it and exhaust the shared table. That threat only exists device_transfer(device_id, task_id) -> 0/-errno
*because* claiming is unauthenticated — and the cap is a poor defence against it, since ```
an attacker can burn 16 slots, claim another device, and burn 16 more. What it reliably
does instead is refuse a legitimate PCI bus with more than 16 functions, which is how an
AMD Ryzen came to boot with no USB and no storage.
So the cap is not merely mis-sized. It is standing in for an authorisation that is not The kernel checks only that the caller currently holds the device. No names, no
performed, and it punishes correct behaviour while barely inconveniencing incorrect attestation, no notion of "the device manager" — the rule is *you may give away what you
behaviour. **Closing the hole is what retires the constant**, not a bigger number. hold*, which is the capability discipline already in force.
## What the principles decide The alternative was the kernel granting roots to a task it recognises by binary path
plus a PID-1 supervisor. It is more robust — it does not depend on the manager being
first — but it puts a binary name inside the kernel, and the goal is that the kernel
keeps only what cannot safely live in user space. A name is not that.
- *Move as much responsibility as possible to user space* (3), and *what remains in the **The residual, stated plainly:** authority here rests on the manager claiming first.
kernel is there for security or a hardware limitation* (5). That holds because `init.csv` decides what starts and in what order, so it is an
operator-visible ordering rather than an attacker-controlled one — but it is an
assumption, not an enforced invariant. The enforced version arrives with the spawn
capability [drivers.md](../device-driver-development/drivers.md) already names as
missing ("`system_spawn` is currently ungated … because there is no spawn capability
yet"). This design is compatible with it and does not block on it.
Deciding **which** driver gets **which** device is policy — matching identity triples ## Decision 2: the kernel stops holding inventory
and choosing a binary. That decision belongs to the device manager and stays there.
`devices.csv` is not the policy; it is **configuration**, the declarative data the
policy reads. Three distinct things, and worth keeping apart in this document:
| | Lives in | Example | The kernel reads exactly three things out of a descriptor: **physical ranges** (to check
a mapping falls inside one), **interrupt numbers**, and **one PCI BDF** (to key an IOMMU
domain). Vendor, device and subsystem ids, class triples, `_HID` strings, bus addresses
and names are stored only so `device_enumerate` can hand them back — which
`device-manager.md` already resolves: that call "fades to a manager-internal (then
deleted) seam", because the manager owns the tree as data.
So the kernel's table becomes: **parent, resources, holder, BDF.** That is what cannot
safely run in user space; the rest moves.
**Devices with no resources leave the kernel entirely.** A USB device is addressed
through its controller and carries `resource_count = 0`
([driver-model.md](../device-driver-development/driver-model.md): "that case is allowed
and is the common one"). It conveys no mapping authority, so there is nothing for the
kernel to enforce and no reason for it to know. It is inventory, and inventory is the
manager's — reported by `child_added`, which already carries everything needed.
That is also the case that made `maximum_children_per_parent` necessary: a zero-resource
child sidesteps containment, so a driver could loop `device_register` and fill the
shared table. Once such children are not kernel objects, every remaining entry is a real
contained subdivision of something the caller holds.
## Decision 3: no shared ceiling; a per-holder quota instead
`maximum_devices = 64` and `maximum_children_per_parent = 16` are numbers we invented,
and both are shared — one driver's enumeration starves every other driver, which is how
an AMD Ryzen booted with no USB and no storage.
- **The table becomes dynamic.** It is built after `heap.init` (`kernel.zig`: `pmm.init`
at 137, `heap.init` at 179, `devices_broker.init` at 202), so nothing prevents it. No
specification bounds how many devices a machine has, so nothing should bound ours.
- **`maximum_children_per_parent` is deleted**, because the authorisation it stood in
for now exists.
- **A per-holder quota replaces them.** Dynamic storage without a bound moves the
ceiling to the kernel heap, which is shared and fatal rather than partial — strictly
worse. The bound that is *not* worse is one charged to the task that caused it: a
driver that loops `device_register` exhausts its own allowance, is refused with an
attributable errno, and is restarted by its supervisor while every other driver
carries on. That is the microkernel property rather than a workaround for it, and it
is declared through [bounds.md](bounds.md) like any other.
## The shape of the change
| | Before | After |
|---|---|---| |---|---|---|
| **Mechanism** | the kernel | the check that a grant is held before a mapping is made | | Manager gets its devices | claims them, unauthorised | claims them (first, unopposed) |
| **Policy** | user space | the device manager matching a device to a driver | | Driver gets its device | `argv[1]` + `device_claim` | receives it in the `hello` reply |
| **Configuration** | files | `devices.csv`, `protocol.csv`, `init.csv` | | Kernel checks | is it free? | do you hold it? |
| Kernel stores | the full descriptor | parent, resources, holder, BDF |
| Zero-resource devices | kernel table entries | manager records only |
| Table size | `maximum_devices = 64` | dynamic, per-holder quota |
| Children per parent | `maximum_children_per_parent = 16` | deleted |
Enforcing that a driver **holds only what it was given** is security — it is the gate in Bring-up order changes for the five claiming drivers: `hello` must precede the claim,
front of mapping physical memory. That stays in the kernel, and it is the whole of what because the reply is where the device arrives. `pci-bus` today does the reverse — its
the kernel needs to do. own comment reads "Claim the bridge, map the ECAM, hello the manager, then scan."
The kernel therefore does not need to know about matching, `devices.csv`, driver names, ## What does not change
or why a device was assigned. It needs to know that an authority it can verify granted
this device to this task.
## Only drivers hold devices - The three invariants of [driver-model.md](../device-driver-development/driver-model.md):
a claim is exclusive, a descriptor is a licence to map physical memory, therefore
containment. This design strengthens the first and touches neither of the others.
- The display service's GOP path. The framebuffer is not a device — it is where pixels
go, handed over by the loader, and the compositor uses it as the boot floor until a
real display driver announces itself
([display-v2.md](../device-driver-development/display-v2.md)). The kernel wraps it in
a display-class descriptor so `mmio_map` can hand it over write-combining; that is
plumbing for a mapping, not a claim about what it is.
- Supervision, restart, pruning and re-report
([device-manager.md](../device-driver-development/device-manager.md),
[process-lifecycle.md](process-lifecycle.md)). Delegation slots into the existing
`hello` exchange and changes none of it.
- `device_register` idempotency, which is what lets a restarted bus rebuild the same
ids.
The layering the rest of the system already follows: ## How it is verified
- A **driver** talks to hardware. It is the only thing that holds a device. The invariant is: **a process holds what it was handed and cannot name its way into
- A **service** talks to no hardware at all. It receives device events and sends data holding more.** The suite has no adversarial device case today — the audit's lesson was
and commands over a **protocol**. that "the suite contains no attacker" — so this adds one: a process that was handed
- A **protocol** is the abstraction between them — the OS layer, in the sense of nothing calls `device_claim` and `device_transfer` on a device another driver holds, and
[communication.md](communication.md). on one nobody holds, and is refused each time with its own errno.
So the question "which services need device grants" has the answer **none**. That The Ryzen is the acceptance test for the ceiling half: it is the machine that found the
collapses the design: every holder of a device is a driver, and every driver is spawned constants, and the one that proves them gone.
by the device manager, which is what grants it.
`display` already demonstrates both halves, one right and one wrong
([backend.zig](../../system/services/display/backend.zig)):
| Backend | How it reaches the panel |
|---|---|
| `VirtioGpu` | speaks `scanout-protocol` over an IPC handle; touches no device |
| `Gop` | `device.enumerate`, `device.claim`, `device.mmioMap` |
The virtio-gpu path is the intended shape and it works today: the driver holds the
hardware, the service speaks a protocol to it. **The GOP path is the anomaly** — the
display service reaching into hardware itself, because the firmware framebuffer has no
driver for it to talk to.
It is also the only anomaly. Every other claimant — `pci-bus`, `usb-xhci-bus`,
`ps2-bus`, `virtio-gpu`, `acpi` — is a device-manager child spawned with its device id
in `argv[1]`.
**A firmware-framebuffer driver is therefore a prerequisite of this phase**, not a
consequence of it. It is small: it holds the display device, maps the framebuffer, and
serves `scanout-protocol` — the same interface `virtio-gpu` already serves, so the
compositor needs no new code path and stops caring which is behind it, which it was
designed for.
An earlier draft of this document had `display` asking the device manager for a grant,
and before that had `init` minting grants because `init` starts `display`. Both were
accommodating an exception instead of removing it.
## The design
**A device grant is a capability, delegated from a holder.** The mechanism already
exists: `callCap` passes a handle over an IPC call and the kernel installs it in the
receiver's table ([library/kernel/ipc.zig](../../library/kernel/ipc.zig)), which is how
shared memory and DMA regions already move between processes.
**The device manager is the root of device authority.** `init` has no part in this: it
is the first process, and its job is to start the rest of the system. It starts the
device manager the same way it starts everything else, and knows nothing about devices.
1. **Root.** At boot the kernel mints grants for the devices firmware discovery found
and hands them to the device manager. This is the only place device authority enters
the system, and it comes from ACPI rather than from anyone's say-so.
The kernel recognises the manager by **chain attestation**, the identity the security
track already settled on: the binary path the kernel itself stamped
(`/system/services/device-manager`) *and* a supervisor of PID 1. The binary path
alone would not do — `system_spawn` is deliberately ungated, so any process may spawn
any bundled binary, and a rogue could run a second copy under the same name. It could
not forge the other half: its copy's supervisor is the rogue, and PID 1 is the
kernel's own first process.
2. **Delegation.** The manager passes a driver its device when it spawns it, over the
channel that already exists — the driver `hello`s the manager and the reply carries
the grant. The manager attests its own children by task id, one hop deep, exactly as
`init` attests the services it spawned for `/protocol` (`supervisorSatisfies`); it
knows which task is which driver because it spawned them.
3. **Services receive nothing.** `display`, `input`, `fat` and `logger` hold no devices
and need no grants; they speak protocols to the drivers that do. The one place this
is not true today is the GOP backend above, which a firmware-framebuffer driver
removes.
4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and `dma_bind`
check possession of the grant instead of consulting an ownership table.
Exclusivity stops being a broker refusing a second claimant and becomes the ordinary
property of a capability: only one process was given it.
**`maximum_children_per_parent` is deleted here.** After this a bus driver's children are
devices it enumerated on a bus it was actually given, and the threat the cap was written
for no longer exists.
## What this costs
**Bring-up order changes.** `pci-bus` today claims first and says hello afterwards — its
own comment says "Claim the bridge, map the ECAM, hello the manager, then scan." Under
delegation the hello must come first, because that is where the grant arrives. Five
drivers need that reordering, and it is the bulk of the work.
**A firmware-framebuffer driver has to be written first**, and until it exists the
display service cannot stop claiming a device. It is the smallest new binary in the
tree — hold the display node, map the framebuffer, serve `scanout-protocol` — but it is
new code on the boot path, and the boot path is where a mistake costs a screen.
It brings an ordering constraint with it: `display` cannot paint until that driver is
up, where today it maps the framebuffer itself and is independent. Mitigating,
`display` already tolerates a backend arriving late — `virtio-gpu` hands it a scanout
after the fact over `attach_scanout` — so "wait, then attach" is a path that already
works rather than one to invent.
**The kernel gains one piece of knowledge about a specific binary.** Chain attestation
means the kernel recognises `/system/services/device-manager` under PID 1 as the root
holder. That is a real concession — the kernel would rather know nothing about who is
who — and it is the minimum: device authority has to enter the system somewhere, and
every alternative is worse. Configuration the kernel reads would put a file parser in
the kernel; first-to-ask would be the hole again, at boot.
**A configuration question stays open.** Which binary may be given which device is
expressed today as `devices.csv`, read by the manager — configuration, with the manager
as the policy that acts on it. That is already the right shape and needs no new file.
A separate grant manifest earns itself only when something must differ from "the manager
gets the hardware": holding a device back so a test or a bare-metal driver can take it,
for instance. Not now.
## What it does not solve
- **Hot-unplug and re-enumeration drift.** Still the inventory problem (phase 3, and
open question 4 in the track plan). A grant dying with its holder is not the same as a
device going away.
- **The device table's size.** `maximum_devices` is untouched by this; it goes when the
inventory moves in phase 3.
- **Two processes racing for the same root grant.** Cannot arise: the roots are minted
to whichever task satisfies the chain — the manager's binary under PID 1 — and a
second copy spawned by anyone else fails the supervisor half.
## How this is verified
The invariant is **I3** from the track plan: a process holds what it was handed and
cannot name its way into holding more. The test is adversarial and the suite has never
had one of these for devices: a process that was granted nothing calls `device_claim`
on a device another driver owns, and on one nobody owns, and is refused both times with
its own errno. The audit's lesson was that "the suite contains no attacker"; this is the
attacker for devices.