The phase 2 design called devices.csv "policy". It is not. The CSV is configuration — declarative data an operator edits. The policy is the component that decides using it: the device manager matching a device to a driver, init reading protocol.csv and granting a binding. The kernel holds mechanism, the check that a grant exists before a mapping is made. Corrected where the document conflated them, and the three are now tabulated so the rest of the design can lean on the distinction. It also sharpens the open question. It was "manifest or not"; it is really whether the grant list should be configuration from the start, or whether init should hand the device manager the root grants and let it match by the devices.csv it already reads. A device-grant file adds a second place an operator must keep correct, and earns itself only when something needs to differ from "the manager gets the hardware" — holding a device back for a test or a bare-metal driver, say.
8.0 KiB
Device authority: you hold what you were given
Design for phase 2 of the bounds track, 2026-08-08. Supersedes an earlier draft that argued for capabilities on aesthetic grounds; this one starts from the hole and from the project's principles.
The hole
device_claim(id) checks two things (devices-broker.zig):
pub fn claim(id: u64, owner: u32) ClaimError!void {
if (id >= count) return error.NoSuchDevice;
if (claimed[@intCast(id)] != null) return error.AlreadyClaimed;
claimed[@intCast(id)] = owner;
}
Does it exist, and is it free. Any process may claim any unclaimed device.
The matching is real but it is entirely advisory: device-manager reads devices.csv,
matches a device to a driver, and spawns that driver with the device id as argv[1]
(spawnDriver). The driver parses the string and claims it. Nothing anywhere binds the
manager's decision to the kernel's grant — a process can pass any integer and win the
race.
A claim is not a small thing. It is what gates mmio_map and irq_bind, so it is a
licence to map physical memory and receive interrupts.
What this costs, beyond the obvious
maximum_children_per_parent = 16 exists because a driver that claimed one device could
loop device_register under it and exhaust the shared table. That threat only exists
because claiming is unauthenticated — and the cap is a poor defence against it, since
an attacker can burn 16 slots, claim another device, and burn 16 more. What it reliably
does instead is refuse a legitimate PCI bus with more than 16 functions, which is how an
AMD Ryzen came to boot with no USB and no storage.
So the cap is not merely mis-sized. It is standing in for an authorisation that is not performed, and it punishes correct behaviour while barely inconveniencing incorrect behaviour. Closing the hole is what retires the constant, not a bigger number.
What the principles decide
- Move as much responsibility as possible to user space (3), and what remains in the kernel is there for security or a hardware limitation (5).
Deciding which driver gets which device is policy — matching identity triples
and choosing a binary. That decision belongs to the device manager and stays there.
devices.csv is not the policy; it is configuration, the declarative data the
policy reads. Three distinct things, and worth keeping apart in this document:
| Lives in | Example | |
|---|---|---|
| Mechanism | the kernel | the check that a grant is held before a mapping is made |
| Policy | user space | the device manager matching a device to a driver |
| Configuration | files | devices.csv, protocol.csv, init.csv |
Enforcing that a driver holds only what it was given is security — it is the gate in front of mapping physical memory. That stays in the kernel, and it is the whole of what the kernel needs to do.
The kernel therefore does not need to know about matching, devices.csv, driver names,
or why a device was assigned. It needs to know that an authority it can verify granted
this device to this task.
The awkward fact that decides the mechanism
Five of the six claimants are device-manager children, spawned with their device id in
argv[1]: pci-bus, usb-xhci-bus, ps2-bus, virtio-gpu, acpi.
display is not. It is spawned by init from init.csv, and it finds its device by
enumerating the table for a display-class node and claiming whatever it finds
(backend.zig). There is no assignment to
enforce, because nobody assigned it anything.
That rules out the cheapest design. "The kernel records the device named at spawn, and
device_claim checks it" closes the hole for five claimants and breaks the sixth. And
the sixth is not an oddity to special-case — it is the one that shows the model is
wrong: authority should be delegable, not welded to the moment of spawn.
The design
A device grant is a capability, delegated from a holder. The mechanism already
exists: callCap passes a handle over an IPC call and the kernel installs it in the
receiver's table (library/kernel/ipc.zig), which is how
shared memory and DMA regions already move between processes.
- Root. At boot the kernel mints grants for the devices firmware discovery found
and hands them to
init(PID 1, which the kernel spawns and therefore need not authenticate). This is the only place device authority enters the system, and it comes from ACPI rather than from anyone's say-so. - Delegation.
initpasses the device manager the grants it will need — in practice all of them — and passesdisplaythe framebuffer grant, becauseinitis what startsdisplay. This is the same shape as the/protocolregistry, whereinitis already the grantor andprotocol.csvalready records/system/services/device-manager, /system/services/init, bind, device-manager. - Assignment. The manager passes a driver its device when it spawns it, over the
channel that already exists — the driver
hellos the manager, and the reply carries the grant. - Use.
mmio_map,irq_bind,msi_bind,io_read/io_writeanddma_bindcheck possession of the grant instead of consulting an ownership table.
Exclusivity stops being a broker refusing a second claimant and becomes the ordinary property of a capability: only one process was given it.
maximum_children_per_parent is deleted here. After this a bus driver's children are
devices it enumerated on a bus it was actually given, and the threat the cap was written
for no longer exists.
What this costs
Bring-up order changes. pci-bus today claims first and says hello afterwards — its
own comment says "Claim the bridge, map the ECAM, hello the manager, then scan." Under
delegation the hello must come first, because that is where the grant arrives. Five
drivers need that reordering, and it is the bulk of the work.
init grows a device role. It already registers /protocol and reads
protocol.csv; it would also hold root device grants and hand them on. That is more
responsibility in PID 1, which is a cost worth naming — though the alternative is the
kernel deciding who may hold what, which principle 5 excludes.
A configuration question follows. Who may bind which protocol name is expressed as
configuration today — protocol.csv — with init as the policy that reads it. Device
grants could be expressed the same way: a file saying which binary may be given which
device class, with init again the policy that enforces it.
That symmetry is real but it is not proposed here. The smaller step is init handing
the manager the root grants and the manager matching by devices.csv, which is
configuration it already reads. A device-grant file adds a second place an operator must
keep correct, and it should only arrive when something needs it to differ from "the
manager gets the hardware" — for example, holding a device back from the manager so a
test or a bare-metal driver can take it.
What it does not solve
- Hot-unplug and re-enumeration drift. Still the inventory problem (phase 3, and open question 4 in the track plan). A grant dying with its holder is not the same as a device going away.
- The device table's size.
maximum_devicesis untouched by this; it goes when the inventory moves in phase 3. - Two processes racing for the same root grant. Cannot arise, because roots are
minted to
initalone.
How this is verified
The invariant is I3 from the track plan: a process holds what it was handed and
cannot name its way into holding more. The test is adversarial and the suite has never
had one of these for devices: a process that was granted nothing calls device_claim
on a device another driver owns, and on one nobody owns, and is refused both times with
its own errno. The audit's lesson was that "the suite contains no attacker"; this is the
attacker for devices.