kernel: a refusal names its rule, and two bounds stop failing open
An AMD Ryzen booted to a working compositor with no USB and no storage, and the log said only "register refused". A tree-wide audit of every compile-time ceiling followed: 235 of them, 139 on quantities the machine or a file decides rather than us, 5 documented anywhere, 171 silent when reached. docs/fixed-bounds-audit.md has the inventory. Errno attribution. The errno space was split between the kernel and the envelope, free to drift; it is now one list in system/abi.zig, restated on both sides, with a comptime check in library/device/driver where the two halves are visible. device_register's six refusals and device_claim's three are distinct codes, so a bus driver can say which rule stopped it, and BadParent splits into NoSuchParent and NotYourParent. pci-bus reconciles found against registered instead of counting refused functions as found. Idempotency ordering. The child cap was checked before the identity match, so a restarted bus was refused its own devices — the supervision restart the system leans on ratcheted toward a degraded machine. A re-registration consumes no slot and is now admitted first. IOMMU fail-closed. confineDevice returned success for a device id past the confinement table, leaving the device outside every domain while the caller believed it confined — unreachable only while ids stop at 64, which both the inventory move and a hardware-reported domain count would change. It refuses now, and the coupling to the broker's device cap is a comptime assert rather than a sentence in a comment. PCI apertures. The bridge's MMIO apertures are derived from the holes in the firmware memory map, and the derivation copied sub-4 GiB entries into a fixed [64] array and skipped the rest. A skipped region is not merely lost: the gap finder concludes it is free, so a real machine's 60-200 entry map yields an aperture over live RAM, and containment then admits a child BAR covering kernel memory. Rewritten to walk the map in place, with the hole finder extracted as a pure function and driven by a synthetic 100-entry map in a new test case. Both new tests were verified to fail on the old code. parameters.zig gains the rationale it was missing and loses a stale sentence pointing at the wrong file; vdso.md documents the errno space, including EPEER, which had no written meaning anywhere. docs/os-development/bounds.md is how a ceiling is declared from here. docs/bounds-track-plan.md is the plan to remove the ones we invented. Suite 114 -> 115.
This commit is contained in:
@@ -0,0 +1,173 @@
|
||||
# Device authority: the kernel stops keeping an inventory
|
||||
|
||||
*Design, drafted 2026-08-07. Not implemented. Prompted by a real machine: an AMD
|
||||
Ryzen desktop enumerated more PCI functions than the kernel's device table would
|
||||
hold, and the xHCI and SATA controllers were refused registration — so the
|
||||
machine booted to the compositor with no USB and no storage.*
|
||||
|
||||
The kernel keeps a table of every device userspace discovers. It is a fixed
|
||||
array of 64 descriptors, 344 bytes each, and a second cap allows any one parent
|
||||
16 children. Neither number is written down anywhere as a decision:
|
||||
`maximum_devices` has no comment and never reached `parameters.zig`, where every
|
||||
other tunable in this kernel lives with its reasoning attached.
|
||||
|
||||
Raising them is not the fix. The numbers are wrong because the *table* is wrong:
|
||||
it is an inventory of hardware, and an inventory of hardware is not something a
|
||||
kernel needs. This document proposes replacing it with capabilities, which
|
||||
removes the ceiling rather than moving it.
|
||||
|
||||
## What the kernel actually uses
|
||||
|
||||
Every read of a device descriptor from the kernel proper, exhaustively:
|
||||
|
||||
| Used for | What it needs |
|
||||
|---|---|
|
||||
| `mmio_map`, `io_read`/`io_write` | the physical range, to check the mapping falls inside it |
|
||||
| `irq_bind`, `msi_bind` | the GSI, and that the caller owns the device |
|
||||
| `dma_bind` | that the caller owns the device |
|
||||
| IOMMU confinement | the **PCI BDF**, to key a domain |
|
||||
| `device_register` | the parent's ranges, for the containment check |
|
||||
| the boot display seed | one framebuffer window |
|
||||
|
||||
That is: **physical ranges, interrupt numbers, and one BDF.** Vendor, device and
|
||||
subsystem ids, class triples, the human-readable names, the parent links, the
|
||||
bus numbers — the kernel stores all of it and reads none of it. It is held so
|
||||
that `device_enumerate` can hand it back to user space, which is the whole
|
||||
mistake in one sentence: the kernel is acting as a distribution mechanism for
|
||||
data it does not use.
|
||||
|
||||
## The split
|
||||
|
||||
Three concerns are tangled in one table.
|
||||
|
||||
**Platform bring-up** — timers, LAPIC/IOAPIC, CPU topology, the ECAM window, the
|
||||
framebuffer the firmware left. The kernel derives these from ACPI before user
|
||||
space exists and needs them to function. They never belonged in the device table
|
||||
and mostly are not (the platform block in `kernel.zig` is separate); this
|
||||
document does not change them.
|
||||
|
||||
**Resource authority** — which task may map which physical range, receive which
|
||||
interrupt, touch which ports. This *must* stay in the kernel. It is the one
|
||||
grant that cannot be audited after the fact: a process that maps arbitrary
|
||||
physical memory owns the machine, page tables and IOMMU structures included.
|
||||
This is memory protection, not device management, and it is why the answer is
|
||||
not simply "move it all to the device manager".
|
||||
|
||||
**Device inventory** — what exists, what it is, how it is arranged, which driver
|
||||
should bind it. This is `device-manager`'s job and is already half there: it
|
||||
loads `devices.csv`, matches, spawns drivers, and receives `child_added` reports
|
||||
over its own protocol. The kernel table duplicates what those reports already
|
||||
carry.
|
||||
|
||||
## The proposal: a resource is a capability
|
||||
|
||||
Device resources join endpoints, shared memory and DMA regions as a kind in the
|
||||
handle table.
|
||||
|
||||
1. **Roots.** At boot the kernel mints capabilities for the windows it learned
|
||||
from firmware — the ECAM range, the framebuffer, the legacy port space — and
|
||||
hands them to the first bus drivers. This is the only place device knowledge
|
||||
enters the kernel, and it comes from ACPI, not from a driver's say-so.
|
||||
2. **Subdivision.** A bus driver enumerating hardware derives a narrower
|
||||
capability from one it holds: `resource_derive(cap, kind, start, len) → cap`.
|
||||
The kernel checks the sub-range lies inside the capability being subdivided —
|
||||
the same containment rule as today (`devices-broker.contains`), but checked
|
||||
against *one capability the caller demonstrably holds* rather than by walking
|
||||
a global tree.
|
||||
3. **Delegation.** The driver passes that capability to the child driver over
|
||||
IPC. Cap-passing already exists; this is the mechanism `subscribe` and
|
||||
`attach_scanout` already use.
|
||||
4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and
|
||||
`dma_bind` take a capability handle instead of `(device_id, resource_index)`.
|
||||
Possession *is* the authority — there is nothing to look up and no ownership
|
||||
table to consult.
|
||||
|
||||
Exclusivity stops being a broker refusing a second claimant and becomes the
|
||||
ordinary property of a capability: only one process was given it.
|
||||
|
||||
## What this buys
|
||||
|
||||
**No ceiling.** There is no table to size, so no machine is too big. The
|
||||
Ryzen's enumeration stops being a limit to tune and becomes what it is — a fact
|
||||
about a computer.
|
||||
|
||||
**Reclamation, free.** `count` in the broker today only ever increases;
|
||||
`releaseAllOwnedBy` clears a dead driver's *claims* but never its entries. A
|
||||
driver that crashes and is restarted re-registers its children and consumes the
|
||||
table again — reachable today without any malice, given the device manager
|
||||
restarts drivers by design. Capabilities die with the task.
|
||||
|
||||
**No quota needed.** The per-parent cap exists to stop one claimant looping
|
||||
`device_register` and filling the shared table, because a zero-resource child
|
||||
sidesteps the containment check. With no shared table there is nothing to
|
||||
exhaust; a process can only ever subdivide what it was given, and its handles
|
||||
are already bounded per task.
|
||||
|
||||
**A smaller kernel.** Three syscalls leave the ABI, one narrower one arrives,
|
||||
and `devices-broker.zig` largely disappears along with both constants.
|
||||
|
||||
**The discipline the rest of the system already uses.** "The claim is the
|
||||
capability" is written in the driver documentation as if it were already true.
|
||||
This makes it true.
|
||||
|
||||
## Syscall surface
|
||||
|
||||
- `device_enumerate` — **retires.** It exists only to read the kernel's table.
|
||||
Callers ask `device-manager`, whose protocol already reserves an `enumerate`
|
||||
verb. Note this is a public-ABI change: `vdso.md` documents it.
|
||||
- `device_register` — **splits.** The kernel half becomes `resource_derive`; the
|
||||
publication half ("this device exists, here is what it is") becomes an IPC
|
||||
message to `device-manager`, which is where the inventory belongs and where
|
||||
`child_added` already carries the same facts.
|
||||
- `device_claim` — **dissolves into possession**, except for the IOMMU (below).
|
||||
- `mmio_map`, `irq_bind`, `msi_bind`, `io_read`, `io_write`, `dma_bind` — keep
|
||||
their names and semantics; their first argument becomes a capability handle.
|
||||
|
||||
## The open question: where the IOMMU attaches
|
||||
|
||||
This is the one place the kernel still needs device *identity* rather than a
|
||||
range. `confineDevice(device_id, bdf, owner)` builds a domain keyed by PCI BDF,
|
||||
attaches it, and the claim is rolled back if confinement fails — deliberately,
|
||||
so a device that cannot be confined is never driven.
|
||||
|
||||
Three options, none obviously right:
|
||||
|
||||
1. **The BDF rides the capability.** A memory capability derived for a PCI
|
||||
function carries its BDF, and the kernel confines on first `mmio_map` or
|
||||
`dma_bind`. Keeps the syscall count down; means a capability is no longer
|
||||
purely a range.
|
||||
2. **An explicit `device_attach(cap, bdf)`.** Honest and visible, but it puts a
|
||||
BDF — a fact about PCI — into a kernel interface that otherwise knows nothing
|
||||
about buses, and something must stop a caller naming a BDF that is not
|
||||
theirs.
|
||||
3. **The root PCI capability carries the segment, and derivation computes the
|
||||
BDF.** Purest, but only works for PCI and the kernel would be parsing bus
|
||||
topology, which is precisely what this document is trying to stop.
|
||||
|
||||
My inclination is (1), because confinement is a property of the resource being
|
||||
granted rather than a separate action, and because it keeps the "possession is
|
||||
authority" story intact. It needs the derivation call to know it is carving a
|
||||
PCI function, which is a wart worth arguing about.
|
||||
|
||||
## What this does not solve
|
||||
|
||||
- **Hot-plug and removal.** Capabilities die with their holder, but a device
|
||||
that physically disappears while a driver lives is the device manager's
|
||||
problem and unchanged by this.
|
||||
- **Quotas on physical memory.** Nothing here bounds how much a driver maps; it
|
||||
bounds only *what* it may map. That was already true.
|
||||
- **The device tree as a published thing.** `/system/devices` remains a device
|
||||
manager concern, as the file-system hierarchy already assumes.
|
||||
|
||||
## Migration sketch
|
||||
|
||||
Not a plan yet — the phases need sizing once the IOMMU question is settled.
|
||||
Rough shape: introduce the capability kind and `resource_derive` alongside the
|
||||
existing table; convert the resource-consuming syscalls to accept either form;
|
||||
move the inventory into `device-manager` and convert its clients off
|
||||
`device_enumerate`; then delete the table, the two constants and the three
|
||||
syscalls in one flag-day, as the `ServiceId` retirement did.
|
||||
|
||||
Every driver is affected, so the QEMU suite is the arbiter at each step, and the
|
||||
Ryzen is the acceptance test — it is the machine that found this, and the one
|
||||
that proves it fixed.
|
||||
Reference in New Issue
Block a user