An AMD Ryzen booted to a working compositor with no USB and no storage, and the log said only "register refused". A tree-wide audit of every compile-time ceiling followed: 235 of them, 139 on quantities the machine or a file decides rather than us, 5 documented anywhere, 171 silent when reached. docs/fixed-bounds-audit.md has the inventory. Errno attribution. The errno space was split between the kernel and the envelope, free to drift; it is now one list in system/abi.zig, restated on both sides, with a comptime check in library/device/driver where the two halves are visible. device_register's six refusals and device_claim's three are distinct codes, so a bus driver can say which rule stopped it, and BadParent splits into NoSuchParent and NotYourParent. pci-bus reconciles found against registered instead of counting refused functions as found. Idempotency ordering. The child cap was checked before the identity match, so a restarted bus was refused its own devices — the supervision restart the system leans on ratcheted toward a degraded machine. A re-registration consumes no slot and is now admitted first. IOMMU fail-closed. confineDevice returned success for a device id past the confinement table, leaving the device outside every domain while the caller believed it confined — unreachable only while ids stop at 64, which both the inventory move and a hardware-reported domain count would change. It refuses now, and the coupling to the broker's device cap is a comptime assert rather than a sentence in a comment. PCI apertures. The bridge's MMIO apertures are derived from the holes in the firmware memory map, and the derivation copied sub-4 GiB entries into a fixed [64] array and skipped the rest. A skipped region is not merely lost: the gap finder concludes it is free, so a real machine's 60-200 entry map yields an aperture over live RAM, and containment then admits a child BAR covering kernel memory. Rewritten to walk the map in place, with the hole finder extracted as a pure function and driven by a synthetic 100-entry map in a new test case. Both new tests were verified to fail on the old code. parameters.zig gains the rationale it was missing and loses a stale sentence pointing at the wrong file; vdso.md documents the errno space, including EPEER, which had no written meaning anywhere. docs/os-development/bounds.md is how a ceiling is declared from here. docs/bounds-track-plan.md is the plan to remove the ones we invented. Suite 114 -> 115.
174 lines
8.9 KiB
Markdown
174 lines
8.9 KiB
Markdown
# Device authority: the kernel stops keeping an inventory
|
|
|
|
*Design, drafted 2026-08-07. Not implemented. Prompted by a real machine: an AMD
|
|
Ryzen desktop enumerated more PCI functions than the kernel's device table would
|
|
hold, and the xHCI and SATA controllers were refused registration — so the
|
|
machine booted to the compositor with no USB and no storage.*
|
|
|
|
The kernel keeps a table of every device userspace discovers. It is a fixed
|
|
array of 64 descriptors, 344 bytes each, and a second cap allows any one parent
|
|
16 children. Neither number is written down anywhere as a decision:
|
|
`maximum_devices` has no comment and never reached `parameters.zig`, where every
|
|
other tunable in this kernel lives with its reasoning attached.
|
|
|
|
Raising them is not the fix. The numbers are wrong because the *table* is wrong:
|
|
it is an inventory of hardware, and an inventory of hardware is not something a
|
|
kernel needs. This document proposes replacing it with capabilities, which
|
|
removes the ceiling rather than moving it.
|
|
|
|
## What the kernel actually uses
|
|
|
|
Every read of a device descriptor from the kernel proper, exhaustively:
|
|
|
|
| Used for | What it needs |
|
|
|---|---|
|
|
| `mmio_map`, `io_read`/`io_write` | the physical range, to check the mapping falls inside it |
|
|
| `irq_bind`, `msi_bind` | the GSI, and that the caller owns the device |
|
|
| `dma_bind` | that the caller owns the device |
|
|
| IOMMU confinement | the **PCI BDF**, to key a domain |
|
|
| `device_register` | the parent's ranges, for the containment check |
|
|
| the boot display seed | one framebuffer window |
|
|
|
|
That is: **physical ranges, interrupt numbers, and one BDF.** Vendor, device and
|
|
subsystem ids, class triples, the human-readable names, the parent links, the
|
|
bus numbers — the kernel stores all of it and reads none of it. It is held so
|
|
that `device_enumerate` can hand it back to user space, which is the whole
|
|
mistake in one sentence: the kernel is acting as a distribution mechanism for
|
|
data it does not use.
|
|
|
|
## The split
|
|
|
|
Three concerns are tangled in one table.
|
|
|
|
**Platform bring-up** — timers, LAPIC/IOAPIC, CPU topology, the ECAM window, the
|
|
framebuffer the firmware left. The kernel derives these from ACPI before user
|
|
space exists and needs them to function. They never belonged in the device table
|
|
and mostly are not (the platform block in `kernel.zig` is separate); this
|
|
document does not change them.
|
|
|
|
**Resource authority** — which task may map which physical range, receive which
|
|
interrupt, touch which ports. This *must* stay in the kernel. It is the one
|
|
grant that cannot be audited after the fact: a process that maps arbitrary
|
|
physical memory owns the machine, page tables and IOMMU structures included.
|
|
This is memory protection, not device management, and it is why the answer is
|
|
not simply "move it all to the device manager".
|
|
|
|
**Device inventory** — what exists, what it is, how it is arranged, which driver
|
|
should bind it. This is `device-manager`'s job and is already half there: it
|
|
loads `devices.csv`, matches, spawns drivers, and receives `child_added` reports
|
|
over its own protocol. The kernel table duplicates what those reports already
|
|
carry.
|
|
|
|
## The proposal: a resource is a capability
|
|
|
|
Device resources join endpoints, shared memory and DMA regions as a kind in the
|
|
handle table.
|
|
|
|
1. **Roots.** At boot the kernel mints capabilities for the windows it learned
|
|
from firmware — the ECAM range, the framebuffer, the legacy port space — and
|
|
hands them to the first bus drivers. This is the only place device knowledge
|
|
enters the kernel, and it comes from ACPI, not from a driver's say-so.
|
|
2. **Subdivision.** A bus driver enumerating hardware derives a narrower
|
|
capability from one it holds: `resource_derive(cap, kind, start, len) → cap`.
|
|
The kernel checks the sub-range lies inside the capability being subdivided —
|
|
the same containment rule as today (`devices-broker.contains`), but checked
|
|
against *one capability the caller demonstrably holds* rather than by walking
|
|
a global tree.
|
|
3. **Delegation.** The driver passes that capability to the child driver over
|
|
IPC. Cap-passing already exists; this is the mechanism `subscribe` and
|
|
`attach_scanout` already use.
|
|
4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and
|
|
`dma_bind` take a capability handle instead of `(device_id, resource_index)`.
|
|
Possession *is* the authority — there is nothing to look up and no ownership
|
|
table to consult.
|
|
|
|
Exclusivity stops being a broker refusing a second claimant and becomes the
|
|
ordinary property of a capability: only one process was given it.
|
|
|
|
## What this buys
|
|
|
|
**No ceiling.** There is no table to size, so no machine is too big. The
|
|
Ryzen's enumeration stops being a limit to tune and becomes what it is — a fact
|
|
about a computer.
|
|
|
|
**Reclamation, free.** `count` in the broker today only ever increases;
|
|
`releaseAllOwnedBy` clears a dead driver's *claims* but never its entries. A
|
|
driver that crashes and is restarted re-registers its children and consumes the
|
|
table again — reachable today without any malice, given the device manager
|
|
restarts drivers by design. Capabilities die with the task.
|
|
|
|
**No quota needed.** The per-parent cap exists to stop one claimant looping
|
|
`device_register` and filling the shared table, because a zero-resource child
|
|
sidesteps the containment check. With no shared table there is nothing to
|
|
exhaust; a process can only ever subdivide what it was given, and its handles
|
|
are already bounded per task.
|
|
|
|
**A smaller kernel.** Three syscalls leave the ABI, one narrower one arrives,
|
|
and `devices-broker.zig` largely disappears along with both constants.
|
|
|
|
**The discipline the rest of the system already uses.** "The claim is the
|
|
capability" is written in the driver documentation as if it were already true.
|
|
This makes it true.
|
|
|
|
## Syscall surface
|
|
|
|
- `device_enumerate` — **retires.** It exists only to read the kernel's table.
|
|
Callers ask `device-manager`, whose protocol already reserves an `enumerate`
|
|
verb. Note this is a public-ABI change: `vdso.md` documents it.
|
|
- `device_register` — **splits.** The kernel half becomes `resource_derive`; the
|
|
publication half ("this device exists, here is what it is") becomes an IPC
|
|
message to `device-manager`, which is where the inventory belongs and where
|
|
`child_added` already carries the same facts.
|
|
- `device_claim` — **dissolves into possession**, except for the IOMMU (below).
|
|
- `mmio_map`, `irq_bind`, `msi_bind`, `io_read`, `io_write`, `dma_bind` — keep
|
|
their names and semantics; their first argument becomes a capability handle.
|
|
|
|
## The open question: where the IOMMU attaches
|
|
|
|
This is the one place the kernel still needs device *identity* rather than a
|
|
range. `confineDevice(device_id, bdf, owner)` builds a domain keyed by PCI BDF,
|
|
attaches it, and the claim is rolled back if confinement fails — deliberately,
|
|
so a device that cannot be confined is never driven.
|
|
|
|
Three options, none obviously right:
|
|
|
|
1. **The BDF rides the capability.** A memory capability derived for a PCI
|
|
function carries its BDF, and the kernel confines on first `mmio_map` or
|
|
`dma_bind`. Keeps the syscall count down; means a capability is no longer
|
|
purely a range.
|
|
2. **An explicit `device_attach(cap, bdf)`.** Honest and visible, but it puts a
|
|
BDF — a fact about PCI — into a kernel interface that otherwise knows nothing
|
|
about buses, and something must stop a caller naming a BDF that is not
|
|
theirs.
|
|
3. **The root PCI capability carries the segment, and derivation computes the
|
|
BDF.** Purest, but only works for PCI and the kernel would be parsing bus
|
|
topology, which is precisely what this document is trying to stop.
|
|
|
|
My inclination is (1), because confinement is a property of the resource being
|
|
granted rather than a separate action, and because it keeps the "possession is
|
|
authority" story intact. It needs the derivation call to know it is carving a
|
|
PCI function, which is a wart worth arguing about.
|
|
|
|
## What this does not solve
|
|
|
|
- **Hot-plug and removal.** Capabilities die with their holder, but a device
|
|
that physically disappears while a driver lives is the device manager's
|
|
problem and unchanged by this.
|
|
- **Quotas on physical memory.** Nothing here bounds how much a driver maps; it
|
|
bounds only *what* it may map. That was already true.
|
|
- **The device tree as a published thing.** `/system/devices` remains a device
|
|
manager concern, as the file-system hierarchy already assumes.
|
|
|
|
## Migration sketch
|
|
|
|
Not a plan yet — the phases need sizing once the IOMMU question is settled.
|
|
Rough shape: introduce the capability kind and `resource_derive` alongside the
|
|
existing table; convert the resource-consuming syscalls to accept either form;
|
|
move the inventory into `device-manager` and convert its clients off
|
|
`device_enumerate`; then delete the table, the two constants and the three
|
|
syscalls in one flag-day, as the `ServiceId` retirement did.
|
|
|
|
Every driver is affected, so the QEMU suite is the arbiter at each step, and the
|
|
Ryzen is the acceptance test — it is the machine that found this, and the one
|
|
that proves it fixed.
|