Files
danos/docs/os-development/device-authority.md
T
Daniel Samson a86559648e kernel: a refusal names its rule, and two bounds stop failing open
An AMD Ryzen booted to a working compositor with no USB and no storage,
and the log said only "register refused". A tree-wide audit of every
compile-time ceiling followed: 235 of them, 139 on quantities the machine
or a file decides rather than us, 5 documented anywhere, 171 silent when
reached. docs/fixed-bounds-audit.md has the inventory.

Errno attribution. The errno space was split between the kernel and the
envelope, free to drift; it is now one list in system/abi.zig, restated on
both sides, with a comptime check in library/device/driver where the two
halves are visible. device_register's six refusals and device_claim's three
are distinct codes, so a bus driver can say which rule stopped it, and
BadParent splits into NoSuchParent and NotYourParent. pci-bus reconciles
found against registered instead of counting refused functions as found.

Idempotency ordering. The child cap was checked before the identity match,
so a restarted bus was refused its own devices — the supervision restart the
system leans on ratcheted toward a degraded machine. A re-registration
consumes no slot and is now admitted first.

IOMMU fail-closed. confineDevice returned success for a device id past the
confinement table, leaving the device outside every domain while the caller
believed it confined — unreachable only while ids stop at 64, which both the
inventory move and a hardware-reported domain count would change. It refuses
now, and the coupling to the broker's device cap is a comptime assert rather
than a sentence in a comment.

PCI apertures. The bridge's MMIO apertures are derived from the holes in the
firmware memory map, and the derivation copied sub-4 GiB entries into a
fixed [64] array and skipped the rest. A skipped region is not merely lost:
the gap finder concludes it is free, so a real machine's 60-200 entry map
yields an aperture over live RAM, and containment then admits a child BAR
covering kernel memory. Rewritten to walk the map in place, with the hole
finder extracted as a pure function and driven by a synthetic 100-entry map
in a new test case. Both new tests were verified to fail on the old code.

parameters.zig gains the rationale it was missing and loses a stale sentence
pointing at the wrong file; vdso.md documents the errno space, including
EPEER, which had no written meaning anywhere.

docs/os-development/bounds.md is how a ceiling is declared from here.
docs/bounds-track-plan.md is the plan to remove the ones we invented.

Suite 114 -> 115.
2026-08-08 11:09:54 +01:00

174 lines
8.9 KiB
Markdown

# Device authority: the kernel stops keeping an inventory
*Design, drafted 2026-08-07. Not implemented. Prompted by a real machine: an AMD
Ryzen desktop enumerated more PCI functions than the kernel's device table would
hold, and the xHCI and SATA controllers were refused registration — so the
machine booted to the compositor with no USB and no storage.*
The kernel keeps a table of every device userspace discovers. It is a fixed
array of 64 descriptors, 344 bytes each, and a second cap allows any one parent
16 children. Neither number is written down anywhere as a decision:
`maximum_devices` has no comment and never reached `parameters.zig`, where every
other tunable in this kernel lives with its reasoning attached.
Raising them is not the fix. The numbers are wrong because the *table* is wrong:
it is an inventory of hardware, and an inventory of hardware is not something a
kernel needs. This document proposes replacing it with capabilities, which
removes the ceiling rather than moving it.
## What the kernel actually uses
Every read of a device descriptor from the kernel proper, exhaustively:
| Used for | What it needs |
|---|---|
| `mmio_map`, `io_read`/`io_write` | the physical range, to check the mapping falls inside it |
| `irq_bind`, `msi_bind` | the GSI, and that the caller owns the device |
| `dma_bind` | that the caller owns the device |
| IOMMU confinement | the **PCI BDF**, to key a domain |
| `device_register` | the parent's ranges, for the containment check |
| the boot display seed | one framebuffer window |
That is: **physical ranges, interrupt numbers, and one BDF.** Vendor, device and
subsystem ids, class triples, the human-readable names, the parent links, the
bus numbers — the kernel stores all of it and reads none of it. It is held so
that `device_enumerate` can hand it back to user space, which is the whole
mistake in one sentence: the kernel is acting as a distribution mechanism for
data it does not use.
## The split
Three concerns are tangled in one table.
**Platform bring-up** — timers, LAPIC/IOAPIC, CPU topology, the ECAM window, the
framebuffer the firmware left. The kernel derives these from ACPI before user
space exists and needs them to function. They never belonged in the device table
and mostly are not (the platform block in `kernel.zig` is separate); this
document does not change them.
**Resource authority** — which task may map which physical range, receive which
interrupt, touch which ports. This *must* stay in the kernel. It is the one
grant that cannot be audited after the fact: a process that maps arbitrary
physical memory owns the machine, page tables and IOMMU structures included.
This is memory protection, not device management, and it is why the answer is
not simply "move it all to the device manager".
**Device inventory** — what exists, what it is, how it is arranged, which driver
should bind it. This is `device-manager`'s job and is already half there: it
loads `devices.csv`, matches, spawns drivers, and receives `child_added` reports
over its own protocol. The kernel table duplicates what those reports already
carry.
## The proposal: a resource is a capability
Device resources join endpoints, shared memory and DMA regions as a kind in the
handle table.
1. **Roots.** At boot the kernel mints capabilities for the windows it learned
from firmware — the ECAM range, the framebuffer, the legacy port space — and
hands them to the first bus drivers. This is the only place device knowledge
enters the kernel, and it comes from ACPI, not from a driver's say-so.
2. **Subdivision.** A bus driver enumerating hardware derives a narrower
capability from one it holds: `resource_derive(cap, kind, start, len) → cap`.
The kernel checks the sub-range lies inside the capability being subdivided —
the same containment rule as today (`devices-broker.contains`), but checked
against *one capability the caller demonstrably holds* rather than by walking
a global tree.
3. **Delegation.** The driver passes that capability to the child driver over
IPC. Cap-passing already exists; this is the mechanism `subscribe` and
`attach_scanout` already use.
4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and
`dma_bind` take a capability handle instead of `(device_id, resource_index)`.
Possession *is* the authority — there is nothing to look up and no ownership
table to consult.
Exclusivity stops being a broker refusing a second claimant and becomes the
ordinary property of a capability: only one process was given it.
## What this buys
**No ceiling.** There is no table to size, so no machine is too big. The
Ryzen's enumeration stops being a limit to tune and becomes what it is — a fact
about a computer.
**Reclamation, free.** `count` in the broker today only ever increases;
`releaseAllOwnedBy` clears a dead driver's *claims* but never its entries. A
driver that crashes and is restarted re-registers its children and consumes the
table again — reachable today without any malice, given the device manager
restarts drivers by design. Capabilities die with the task.
**No quota needed.** The per-parent cap exists to stop one claimant looping
`device_register` and filling the shared table, because a zero-resource child
sidesteps the containment check. With no shared table there is nothing to
exhaust; a process can only ever subdivide what it was given, and its handles
are already bounded per task.
**A smaller kernel.** Three syscalls leave the ABI, one narrower one arrives,
and `devices-broker.zig` largely disappears along with both constants.
**The discipline the rest of the system already uses.** "The claim is the
capability" is written in the driver documentation as if it were already true.
This makes it true.
## Syscall surface
- `device_enumerate` — **retires.** It exists only to read the kernel's table.
Callers ask `device-manager`, whose protocol already reserves an `enumerate`
verb. Note this is a public-ABI change: `vdso.md` documents it.
- `device_register` — **splits.** The kernel half becomes `resource_derive`; the
publication half ("this device exists, here is what it is") becomes an IPC
message to `device-manager`, which is where the inventory belongs and where
`child_added` already carries the same facts.
- `device_claim` — **dissolves into possession**, except for the IOMMU (below).
- `mmio_map`, `irq_bind`, `msi_bind`, `io_read`, `io_write`, `dma_bind` — keep
their names and semantics; their first argument becomes a capability handle.
## The open question: where the IOMMU attaches
This is the one place the kernel still needs device *identity* rather than a
range. `confineDevice(device_id, bdf, owner)` builds a domain keyed by PCI BDF,
attaches it, and the claim is rolled back if confinement fails — deliberately,
so a device that cannot be confined is never driven.
Three options, none obviously right:
1. **The BDF rides the capability.** A memory capability derived for a PCI
function carries its BDF, and the kernel confines on first `mmio_map` or
`dma_bind`. Keeps the syscall count down; means a capability is no longer
purely a range.
2. **An explicit `device_attach(cap, bdf)`.** Honest and visible, but it puts a
BDF — a fact about PCI — into a kernel interface that otherwise knows nothing
about buses, and something must stop a caller naming a BDF that is not
theirs.
3. **The root PCI capability carries the segment, and derivation computes the
BDF.** Purest, but only works for PCI and the kernel would be parsing bus
topology, which is precisely what this document is trying to stop.
My inclination is (1), because confinement is a property of the resource being
granted rather than a separate action, and because it keeps the "possession is
authority" story intact. It needs the derivation call to know it is carving a
PCI function, which is a wart worth arguing about.
## What this does not solve
- **Hot-plug and removal.** Capabilities die with their holder, but a device
that physically disappears while a driver lives is the device manager's
problem and unchanged by this.
- **Quotas on physical memory.** Nothing here bounds how much a driver maps; it
bounds only *what* it may map. That was already true.
- **The device tree as a published thing.** `/system/devices` remains a device
manager concern, as the file-system hierarchy already assumes.
## Migration sketch
Not a plan yet — the phases need sizing once the IOMMU question is settled.
Rough shape: introduce the capability kind and `resource_derive` alongside the
existing table; convert the resource-consuming syscalls to accept either form;
move the inventory into `device-manager` and convert its clients off
`device_enumerate`; then delete the table, the two constants and the three
syscalls in one flag-day, as the `ServiceId` retirement did.
Every driver is affected, so the QEMU suite is the arbiter at each step, and the
Ryzen is the acceptance test — it is the machine that found this, and the one
that proves it fixed.