Files
danos/docs/os-development/device-authority.md
T

9.7 KiB

Device authority: you hold what you were given

Design for phase 2 of the bounds track, 2026-08-08. Supersedes an earlier draft that argued for capabilities on aesthetic grounds; this one starts from the hole and from the project's principles.

The hole

device_claim(id) checks two things (devices-broker.zig):

pub fn claim(id: u64, owner: u32) ClaimError!void {
    if (id >= count) return error.NoSuchDevice;
    if (claimed[@intCast(id)] != null) return error.AlreadyClaimed;
    claimed[@intCast(id)] = owner;
}

Does it exist, and is it free. Any process may claim any unclaimed device.

The matching is real but it is entirely advisory: device-manager reads devices.csv, matches a device to a driver, and spawns that driver with the device id as argv[1] (spawnDriver). The driver parses the string and claims it. Nothing anywhere binds the manager's decision to the kernel's grant — a process can pass any integer and win the race.

A claim is not a small thing. It is what gates mmio_map and irq_bind, so it is a licence to map physical memory and receive interrupts.

What this costs, beyond the obvious

maximum_children_per_parent = 16 exists because a driver that claimed one device could loop device_register under it and exhaust the shared table. That threat only exists because claiming is unauthenticated — and the cap is a poor defence against it, since an attacker can burn 16 slots, claim another device, and burn 16 more. What it reliably does instead is refuse a legitimate PCI bus with more than 16 functions, which is how an AMD Ryzen came to boot with no USB and no storage.

So the cap is not merely mis-sized. It is standing in for an authorisation that is not performed, and it punishes correct behaviour while barely inconveniencing incorrect behaviour. Closing the hole is what retires the constant, not a bigger number.

What the principles decide

  • Move as much responsibility as possible to user space (3), and what remains in the kernel is there for security or a hardware limitation (5).

Deciding which driver gets which device is policy — matching identity triples and choosing a binary. That decision belongs to the device manager and stays there. devices.csv is not the policy; it is configuration, the declarative data the policy reads. Three distinct things, and worth keeping apart in this document:

Lives in Example
Mechanism the kernel the check that a grant is held before a mapping is made
Policy user space the device manager matching a device to a driver
Configuration files devices.csv, protocol.csv, init.csv

Enforcing that a driver holds only what it was given is security — it is the gate in front of mapping physical memory. That stays in the kernel, and it is the whole of what the kernel needs to do.

The kernel therefore does not need to know about matching, devices.csv, driver names, or why a device was assigned. It needs to know that an authority it can verify granted this device to this task.

The awkward fact that decides the mechanism

Five of the six claimants are device-manager children, spawned with their device id in argv[1]: pci-bus, usb-xhci-bus, ps2-bus, virtio-gpu, acpi.

display is not. It is spawned by init from init.csv, and it finds its device by enumerating the table for a display-class node and claiming whatever it finds (backend.zig). There is no assignment to enforce, because nobody assigned it anything.

That rules out the cheapest design. "The kernel records the device named at spawn, and device_claim checks it" closes the hole for five claimants and breaks the sixth. And the sixth is not an oddity to special-case — it is the one that shows the model is wrong: authority should be delegable, not welded to the moment of spawn.

It is also what tempted an earlier draft of this document into giving init the root grants, since init is what starts display. That was wrong for a plainer reason: init has nothing to do with devices. It is the first process and it starts the rest of the system; the device manager is what owns hardware. Under delegation display simply asks the manager, like every other driver, and the exception disappears rather than being accommodated.

The design

A device grant is a capability, delegated from a holder. The mechanism already exists: callCap passes a handle over an IPC call and the kernel installs it in the receiver's table (library/kernel/ipc.zig), which is how shared memory and DMA regions already move between processes.

The device manager is the root of device authority. init has no part in this: it is the first process, and its job is to start the rest of the system. It starts the device manager the same way it starts everything else, and knows nothing about devices.

  1. Root. At boot the kernel mints grants for the devices firmware discovery found and hands them to the device manager. This is the only place device authority enters the system, and it comes from ACPI rather than from anyone's say-so.

    The kernel recognises the manager by chain attestation, the identity the security track already settled on: the binary path the kernel itself stamped (/system/services/device-manager) and a supervisor of PID 1. The binary path alone would not do — system_spawn is deliberately ungated, so any process may spawn any bundled binary, and a rogue could run a second copy under the same name. It could not forge the other half: its copy's supervisor is the rogue, and PID 1 is the kernel's own first process.

  2. Delegation. The manager passes a driver its device when it spawns it, over the channel that already exists — the driver hellos the manager and the reply carries the grant. The manager attests its own children by task id, one hop deep, exactly as init attests the services it spawned for /protocol (supervisorSatisfies); it knows which task is which driver because it spawned them.

  3. display asks the manager too. It is started by init as a service, but the framebuffer is a device, so it receives that grant from the device manager like any other driver. One authority for hardware, not two — which is what makes display stop being the exception that broke the simpler design.

  4. Use. mmio_map, irq_bind, msi_bind, io_read/io_write and dma_bind check possession of the grant instead of consulting an ownership table.

Exclusivity stops being a broker refusing a second claimant and becomes the ordinary property of a capability: only one process was given it.

maximum_children_per_parent is deleted here. After this a bus driver's children are devices it enumerated on a bus it was actually given, and the threat the cap was written for no longer exists.

What this costs

Bring-up order changes. pci-bus today claims first and says hello afterwards — its own comment says "Claim the bridge, map the ECAM, hello the manager, then scan." Under delegation the hello must come first, because that is where the grant arrives. Five drivers need that reordering, and it is the bulk of the work.

display gains a dependency on the device manager. It currently finds its framebuffer by itself and needs nothing from anyone; afterwards it must ask. That is a new ordering constraint at boot — display cannot bring up a screen until the manager is up — and it is the honest price of there being one authority for hardware. Mitigating: display already tolerates arriving before its backend (the virtio-gpu driver hands it a scanout later, over attach_scanout), so the machinery for "wait, then attach" is there.

The kernel gains one piece of knowledge about a specific binary. Chain attestation means the kernel recognises /system/services/device-manager under PID 1 as the root holder. That is a real concession — the kernel would rather know nothing about who is who — and it is the minimum: device authority has to enter the system somewhere, and every alternative is worse. Configuration the kernel reads would put a file parser in the kernel; first-to-ask would be the hole again, at boot.

A configuration question stays open. Which binary may be given which device is expressed today as devices.csv, read by the manager — configuration, with the manager as the policy that acts on it. That is already the right shape and needs no new file. A separate grant manifest earns itself only when something must differ from "the manager gets the hardware": holding a device back so a test or a bare-metal driver can take it, for instance. Not now.

What it does not solve

  • Hot-unplug and re-enumeration drift. Still the inventory problem (phase 3, and open question 4 in the track plan). A grant dying with its holder is not the same as a device going away.
  • The device table's size. maximum_devices is untouched by this; it goes when the inventory moves in phase 3.
  • Two processes racing for the same root grant. Cannot arise: the roots are minted to whichever task satisfies the chain — the manager's binary under PID 1 — and a second copy spawned by anyone else fails the supervisor half.

How this is verified

The invariant is I3 from the track plan: a process holds what it was handed and cannot name its way into holding more. The test is adversarial and the suite has never had one of these for devices: a process that was granted nothing calls device_claim on a device another driver owns, and on one nobody owns, and is refused both times with its own errno. The audit's lesson was that "the suite contains no attacker"; this is the attacker for devices.