kernel: a refusal names its rule, and two bounds stop failing open

An AMD Ryzen booted to a working compositor with no USB and no storage,
and the log said only "register refused". A tree-wide audit of every
compile-time ceiling followed: 235 of them, 139 on quantities the machine
or a file decides rather than us, 5 documented anywhere, 171 silent when
reached. docs/fixed-bounds-audit.md has the inventory.

Errno attribution. The errno space was split between the kernel and the
envelope, free to drift; it is now one list in system/abi.zig, restated on
both sides, with a comptime check in library/device/driver where the two
halves are visible. device_register's six refusals and device_claim's three
are distinct codes, so a bus driver can say which rule stopped it, and
BadParent splits into NoSuchParent and NotYourParent. pci-bus reconciles
found against registered instead of counting refused functions as found.

Idempotency ordering. The child cap was checked before the identity match,
so a restarted bus was refused its own devices — the supervision restart the
system leans on ratcheted toward a degraded machine. A re-registration
consumes no slot and is now admitted first.

IOMMU fail-closed. confineDevice returned success for a device id past the
confinement table, leaving the device outside every domain while the caller
believed it confined — unreachable only while ids stop at 64, which both the
inventory move and a hardware-reported domain count would change. It refuses
now, and the coupling to the broker's device cap is a comptime assert rather
than a sentence in a comment.

PCI apertures. The bridge's MMIO apertures are derived from the holes in the
firmware memory map, and the derivation copied sub-4 GiB entries into a
fixed [64] array and skipped the rest. A skipped region is not merely lost:
the gap finder concludes it is free, so a real machine's 60-200 entry map
yields an aperture over live RAM, and containment then admits a child BAR
covering kernel memory. Rewritten to walk the map in place, with the hole
finder extracted as a pure function and driven by a synthetic 100-entry map
in a new test case. Both new tests were verified to fail on the old code.

parameters.zig gains the rationale it was missing and loses a stale sentence
pointing at the wrong file; vdso.md documents the errno space, including
EPEER, which had no written meaning anywhere.

docs/os-development/bounds.md is how a ceiling is declared from here.
docs/bounds-track-plan.md is the plan to remove the ones we invented.

Suite 114 -> 115.
This commit is contained in:
Daniel Samson
2026-08-08 11:09:54 +01:00
parent 7db5fba884
commit a86559648e
25 changed files with 1520 additions and 168 deletions
+140
View File
@@ -0,0 +1,140 @@
# Bounds: how a ceiling is declared
*Design, 2026-08-08. Follows [fixed-bounds-audit.md](../fixed-bounds-audit.md), which
found 235 compile-time ceilings in this tree: 139 on quantities we do not choose, 5
recorded anywhere with their reasoning, and 171 that pass in silence when reached.*
A bound is a number chosen at compile time that decides how much of something the code
can hold. `const maximum_devices = 64`. `var below: [64]Range`. `var blob: [512]u8`.
Different units — devices, firmware memory-map entries, bytes of a USB descriptor — but
one shape, and one recurring way of going wrong.
## Where a bound lives
**Where the thing it bounds lives.** A driver's transfer-ring size belongs to that
driver; a protocol's payload cap belongs to that protocol; the kernel's task-table size
belongs to the kernel. There is no central list and this document does not propose one.
`system/parameters.zig` is not a counter-example. It is kernel-only, and it exists for a
specific historical reason: tunables had accumulated inside the loader↔kernel handoff
contract, and splitting them out kept that contract to what it actually is. It is a
tidying of one file's contents, not a registry the rest of the system reports to.
This matters for the mechanism below. An earlier draft had every bound declared through
a shared `bounds` module — which would have meant adding a dependency to roughly eight
package manifests, including `library/protocol`, which deliberately depends on nothing.
That is a coupling the problem does not require: a bound is a local fact about local
storage, and the only thing worth sharing is the *shape of the statement*, not a module.
## The declaration
A structured doc comment, immediately above the declaration, in the file that owns it:
```zig
/// bound: logical CPUs the kernel tracks
/// decided-by: hardware
/// protects: the per-CPU bookkeeping arrays, which are sized at compile time
/// at-limit: degrade — surplus cores are left parked, never brought online
/// observed-by: platform.cpusDropped() -> the WARNING at kernel.zig:281
pub const maximum_cpus = 128;
```
Five fields, all mandatory:
| Field | Answers |
|---|---|
| `bound` | what is counted, in plain words |
| `decided-by` | `hardware`, `external`, or `ours` — who chooses how large it gets |
| `protects` | what this ceiling defends against |
| `at-limit` | `refuse` / `degrade` / `truncate` / `grow`, and the detail |
| `observed-by` | how an operator finds out it was reached |
`decided-by` is the classification the audit turned on. `hardware` means the machine
chooses — PCI functions, CPUs, ACPI rows, memory-map entries. `external` means a file,
disk structure or peer chooses. `ours` means we do: a stack size, a tick rate, our own
protocol's payload. A fixed bound on the first two is a defect rather than a tunable.
## What the build step enforces
A step in `build.zig` reads the tree and fails on:
1. **A bound with no declaration.** A fixed-size array or a `maximum_*`/`max_*` constant
with no `bound:` block above it. The 235 that exist today are allowlisted by
file+line+name, so only newly written ones are gated — the rule can land without a
235-site sweep in front of it.
2. **A missing field.** All five or it fails. This alone is the 97 bounds with no
comment at all and the 171 with no observability.
3. **An unspeakable `at-limit`.** The vocabulary is closed. There is no `silent`, no
`drop`, and nothing meaning *allow*. `truncate` is legal only with a marker the
reader can see — `klog_maximum_message` qualifies because the record carries
`klog_flag_truncated`; the USB configuration descriptor cut at 512 bytes does not,
because nothing records that anything was lost.
4. **A stale allowlist entry.** If a listed bound is fixed or deleted, its entry goes
too, so the list can only shrink.
Because this is text and not a Zig type, it also covers `boot/`, which imports almost
nothing, and `tools/*.py`, where the audit found bounds as well. One mechanism, whole
tree, no new dependency edges.
## Coupled bounds
Two numbers that must agree, agreeing in code rather than in a comment — no module
needed, just a `comptime` block where one of them lives:
```zig
comptime {
if (maximum_domains != devices_broker.maximum_devices)
@compileError("iommu.confined is indexed by device id; an id past its end is " ++
"left unconfined while confineDevice still reports success");
}
```
`maximum_domains = 64` and `maximum_devices = 64` agree today only by a sentence in a
comment, and the agreement fails open. This is the clause with a live hole behind it,
and the reason raising `maximum_devices` alone would be a privilege escalation rather
than a fix.
## The worked bad case
`devices_broker.maximum_devices`, which had no comment at all:
```zig
/// bound: device nodes for the whole machine — firmware-discovered plus registered
/// decided-by: hardware
/// protects: nothing; this is a sizing guess about someone else's computer
/// at-limit: refuse — ENOSPC from device_register, dropped++ during discovery
/// observed-by: kernel.zig:203 counts discovery drops only, NOT runtime refusals
const maximum_devices = 64;
```
Writing it out is the argument. `decided-by: hardware` alongside a `protects` that
admits there is no threat describes a bound that should not be fixed at all, and
`observed-by` cannot be filled in honestly. The build step does not reject this — it
makes it impossible to write down without noticing.
## The work-list this produces
`decided-by` is machine-readable, so the sweep is a query: every `hardware` or
`external` bound whose `at-limit` is not `grow`. That is 139 of the 235, and it is the
order the fixing takes — by class, not by our guesses about which machines get run.
Reachability is exactly what a new computer changes; the Ryzen's cap was unreachable
until it wasn't.
## What this does not do
**It is a statement, not a proof.** Nothing checks that the code does what `at-limit`
claims. The enforcement is completeness and vocabulary: you cannot leave the question
unanswered, and you cannot answer it with "silently".
**It carries no occupancy.** Knowing `system/configuration/protocol.csv` sits at 50 of
64 grant rows still needs a counter and somewhere to report it. The declarations make
that cheap to add later; it is not here.
**It resizes nothing.** Declaring `maximum_devices` honestly does not make the Ryzen
work. It makes the next machine's failure legible, and it names the 139 that need real
fixes.
**Enforcement is at build time, not compile time.** A Zig type could have made a missing
field a compile error. That version needed the shared module, and the module was not
worth the coupling — so a missing field is a failed build step instead. In practice both
mean `zig build` stops; the difference is which stage prints the message.
+173
View File
@@ -0,0 +1,173 @@
# Device authority: the kernel stops keeping an inventory
*Design, drafted 2026-08-07. Not implemented. Prompted by a real machine: an AMD
Ryzen desktop enumerated more PCI functions than the kernel's device table would
hold, and the xHCI and SATA controllers were refused registration — so the
machine booted to the compositor with no USB and no storage.*
The kernel keeps a table of every device userspace discovers. It is a fixed
array of 64 descriptors, 344 bytes each, and a second cap allows any one parent
16 children. Neither number is written down anywhere as a decision:
`maximum_devices` has no comment and never reached `parameters.zig`, where every
other tunable in this kernel lives with its reasoning attached.
Raising them is not the fix. The numbers are wrong because the *table* is wrong:
it is an inventory of hardware, and an inventory of hardware is not something a
kernel needs. This document proposes replacing it with capabilities, which
removes the ceiling rather than moving it.
## What the kernel actually uses
Every read of a device descriptor from the kernel proper, exhaustively:
| Used for | What it needs |
|---|---|
| `mmio_map`, `io_read`/`io_write` | the physical range, to check the mapping falls inside it |
| `irq_bind`, `msi_bind` | the GSI, and that the caller owns the device |
| `dma_bind` | that the caller owns the device |
| IOMMU confinement | the **PCI BDF**, to key a domain |
| `device_register` | the parent's ranges, for the containment check |
| the boot display seed | one framebuffer window |
That is: **physical ranges, interrupt numbers, and one BDF.** Vendor, device and
subsystem ids, class triples, the human-readable names, the parent links, the
bus numbers — the kernel stores all of it and reads none of it. It is held so
that `device_enumerate` can hand it back to user space, which is the whole
mistake in one sentence: the kernel is acting as a distribution mechanism for
data it does not use.
## The split
Three concerns are tangled in one table.
**Platform bring-up** — timers, LAPIC/IOAPIC, CPU topology, the ECAM window, the
framebuffer the firmware left. The kernel derives these from ACPI before user
space exists and needs them to function. They never belonged in the device table
and mostly are not (the platform block in `kernel.zig` is separate); this
document does not change them.
**Resource authority** — which task may map which physical range, receive which
interrupt, touch which ports. This *must* stay in the kernel. It is the one
grant that cannot be audited after the fact: a process that maps arbitrary
physical memory owns the machine, page tables and IOMMU structures included.
This is memory protection, not device management, and it is why the answer is
not simply "move it all to the device manager".
**Device inventory** — what exists, what it is, how it is arranged, which driver
should bind it. This is `device-manager`'s job and is already half there: it
loads `devices.csv`, matches, spawns drivers, and receives `child_added` reports
over its own protocol. The kernel table duplicates what those reports already
carry.
## The proposal: a resource is a capability
Device resources join endpoints, shared memory and DMA regions as a kind in the
handle table.
1. **Roots.** At boot the kernel mints capabilities for the windows it learned
from firmware — the ECAM range, the framebuffer, the legacy port space — and
hands them to the first bus drivers. This is the only place device knowledge
enters the kernel, and it comes from ACPI, not from a driver's say-so.
2. **Subdivision.** A bus driver enumerating hardware derives a narrower
capability from one it holds: `resource_derive(cap, kind, start, len) → cap`.
The kernel checks the sub-range lies inside the capability being subdivided —
the same containment rule as today (`devices-broker.contains`), but checked
against *one capability the caller demonstrably holds* rather than by walking
a global tree.
3. **Delegation.** The driver passes that capability to the child driver over
IPC. Cap-passing already exists; this is the mechanism `subscribe` and
`attach_scanout` already use.
4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and
`dma_bind` take a capability handle instead of `(device_id, resource_index)`.
Possession *is* the authority — there is nothing to look up and no ownership
table to consult.
Exclusivity stops being a broker refusing a second claimant and becomes the
ordinary property of a capability: only one process was given it.
## What this buys
**No ceiling.** There is no table to size, so no machine is too big. The
Ryzen's enumeration stops being a limit to tune and becomes what it is — a fact
about a computer.
**Reclamation, free.** `count` in the broker today only ever increases;
`releaseAllOwnedBy` clears a dead driver's *claims* but never its entries. A
driver that crashes and is restarted re-registers its children and consumes the
table again — reachable today without any malice, given the device manager
restarts drivers by design. Capabilities die with the task.
**No quota needed.** The per-parent cap exists to stop one claimant looping
`device_register` and filling the shared table, because a zero-resource child
sidesteps the containment check. With no shared table there is nothing to
exhaust; a process can only ever subdivide what it was given, and its handles
are already bounded per task.
**A smaller kernel.** Three syscalls leave the ABI, one narrower one arrives,
and `devices-broker.zig` largely disappears along with both constants.
**The discipline the rest of the system already uses.** "The claim is the
capability" is written in the driver documentation as if it were already true.
This makes it true.
## Syscall surface
- `device_enumerate` — **retires.** It exists only to read the kernel's table.
Callers ask `device-manager`, whose protocol already reserves an `enumerate`
verb. Note this is a public-ABI change: `vdso.md` documents it.
- `device_register` — **splits.** The kernel half becomes `resource_derive`; the
publication half ("this device exists, here is what it is") becomes an IPC
message to `device-manager`, which is where the inventory belongs and where
`child_added` already carries the same facts.
- `device_claim` — **dissolves into possession**, except for the IOMMU (below).
- `mmio_map`, `irq_bind`, `msi_bind`, `io_read`, `io_write`, `dma_bind` — keep
their names and semantics; their first argument becomes a capability handle.
## The open question: where the IOMMU attaches
This is the one place the kernel still needs device *identity* rather than a
range. `confineDevice(device_id, bdf, owner)` builds a domain keyed by PCI BDF,
attaches it, and the claim is rolled back if confinement fails — deliberately,
so a device that cannot be confined is never driven.
Three options, none obviously right:
1. **The BDF rides the capability.** A memory capability derived for a PCI
function carries its BDF, and the kernel confines on first `mmio_map` or
`dma_bind`. Keeps the syscall count down; means a capability is no longer
purely a range.
2. **An explicit `device_attach(cap, bdf)`.** Honest and visible, but it puts a
BDF — a fact about PCI — into a kernel interface that otherwise knows nothing
about buses, and something must stop a caller naming a BDF that is not
theirs.
3. **The root PCI capability carries the segment, and derivation computes the
BDF.** Purest, but only works for PCI and the kernel would be parsing bus
topology, which is precisely what this document is trying to stop.
My inclination is (1), because confinement is a property of the resource being
granted rather than a separate action, and because it keeps the "possession is
authority" story intact. It needs the derivation call to know it is carving a
PCI function, which is a wart worth arguing about.
## What this does not solve
- **Hot-plug and removal.** Capabilities die with their holder, but a device
that physically disappears while a driver lives is the device manager's
problem and unchanged by this.
- **Quotas on physical memory.** Nothing here bounds how much a driver maps; it
bounds only *what* it may map. That was already true.
- **The device tree as a published thing.** `/system/devices` remains a device
manager concern, as the file-system hierarchy already assumes.
## Migration sketch
Not a plan yet — the phases need sizing once the IOMMU question is settled.
Rough shape: introduce the capability kind and `resource_derive` alongside the
existing table; convert the resource-consuming syscalls to accept either form;
move the inventory into `device-manager` and convert its clients off
`device_enumerate`; then delete the table, the two constants and the three
syscalls in one flag-day, as the `ServiceId` retirement did.
Every driver is affected, so the QEMU suite is the arbiter at each step, and the
Ryzen is the acceptance test — it is the machine that found this, and the one
that proves it fixed.
+51
View File
@@ -150,6 +150,57 @@ they are wire values a Rust program needs verbatim. What stays private in
`abi.zig` is exactly the thing the vDSO exists to hide: the `SystemCall`
numbers and the trap convention.
## Errors: the errno space
A failed call returns `-errno`. The runtime detects failure the way Linux
does — a return value in the top 4096 — so every code stays inside 1..4095.
These are **public**: unlike the call numbers, a caller must be able to read
them verbatim, and they are the same vocabulary whether the number came from
the kernel or from a user-space provider answering over IPC.
They are defined once in `system/abi.zig`. The kernel restates them in
`system/kernel/ipc-synchronous.zig` and the envelope restates the
provider-facing subset in `library/protocol/envelope/envelope.zig` (the
`protocol` package deliberately depends on nothing, so it cannot import
`abi`); a comptime check in `library/device/driver/driver.zig` makes drift a
compile error.
| # | Name | Meaning |
|---|------|---------|
| 1 | `EBADF` | bad handle |
| 2 | `E2BIG` | an argument exceeds its maximum (a message, a descriptor's resource count) |
| 3 | `EFAULT` | buffer unmapped, or outside the user half |
| 4 | `ENOENT` | no such name |
| 5 | `ENOSPC` | a kernel table is full (handles, devices) |
| 6 | `ENOMEM` | out of memory |
| 7 | `EPEER` | the peer died before replying — its process exited or was killed |
| 8 | `ESRCH` | no such process |
| 9 | `EPERM` | not permitted: the caller is not the owner or supervisor |
| 10 | `ENOSYS` | this protocol has no such operation |
| 11 | `EPROTO` | malformed packet: shorter than the verb it names |
| 12 | `EBUSY` | the thing asked for is held by someone still alive |
| 13 | `ENODEV` | no such device id |
| 14 | `ECHILDREN` | this parent already holds as many children as it can |
| 15 | `ERANGE` | a resource escapes the window it must fall inside |
| 16 | `ECONFINE` | the device could not be placed under IOMMU translation |
`EPEER` is the one with no POSIX counterpart and it is worth stating plainly:
synchronous IPC blocks the caller until the server replies, so the caller
needs an answer for "the server died while I was waiting." It is not a
transport error and not a refusal — the request may well have been carried
out — it says only that no reply is coming. A client that treats it as
"retry" can duplicate work; the honest response is to re-resolve the protocol
name, because the provider it held is gone.
**A refusal names the rule that refused it.** This is a rule and not a
courtesy. `device_register` alone can fail six ways, and until each got its
own code a bus driver could only report "refused" — which is how an AMD
desktop came to boot with a working display, no USB and no storage, with
three independent causes indistinguishable in the log. See
[fixed-bounds-audit.md](../fixed-bounds-audit.md). A new failure mode that
does not fit an existing code gets a new one here rather than borrowing the
nearest.
## Enforcement, and an honest threat model
Renumbering only has teeth if the kernel **refuses syscalls that don't come