kernel: a refusal names its rule, and two bounds stop failing open
An AMD Ryzen booted to a working compositor with no USB and no storage, and the log said only "register refused". A tree-wide audit of every compile-time ceiling followed: 235 of them, 139 on quantities the machine or a file decides rather than us, 5 documented anywhere, 171 silent when reached. docs/fixed-bounds-audit.md has the inventory. Errno attribution. The errno space was split between the kernel and the envelope, free to drift; it is now one list in system/abi.zig, restated on both sides, with a comptime check in library/device/driver where the two halves are visible. device_register's six refusals and device_claim's three are distinct codes, so a bus driver can say which rule stopped it, and BadParent splits into NoSuchParent and NotYourParent. pci-bus reconciles found against registered instead of counting refused functions as found. Idempotency ordering. The child cap was checked before the identity match, so a restarted bus was refused its own devices — the supervision restart the system leans on ratcheted toward a degraded machine. A re-registration consumes no slot and is now admitted first. IOMMU fail-closed. confineDevice returned success for a device id past the confinement table, leaving the device outside every domain while the caller believed it confined — unreachable only while ids stop at 64, which both the inventory move and a hardware-reported domain count would change. It refuses now, and the coupling to the broker's device cap is a comptime assert rather than a sentence in a comment. PCI apertures. The bridge's MMIO apertures are derived from the holes in the firmware memory map, and the derivation copied sub-4 GiB entries into a fixed [64] array and skipped the rest. A skipped region is not merely lost: the gap finder concludes it is free, so a real machine's 60-200 entry map yields an aperture over live RAM, and containment then admits a child BAR covering kernel memory. Rewritten to walk the map in place, with the hole finder extracted as a pure function and driven by a synthetic 100-entry map in a new test case. Both new tests were verified to fail on the old code. parameters.zig gains the rationale it was missing and loses a stale sentence pointing at the wrong file; vdso.md documents the errno space, including EPEER, which had no written meaning anywhere. docs/os-development/bounds.md is how a ceiling is declared from here. docs/bounds-track-plan.md is the plan to remove the ones we invented. Suite 114 -> 115.
This commit is contained in:
@@ -0,0 +1,140 @@
|
||||
# Bounds: how a ceiling is declared
|
||||
|
||||
*Design, 2026-08-08. Follows [fixed-bounds-audit.md](../fixed-bounds-audit.md), which
|
||||
found 235 compile-time ceilings in this tree: 139 on quantities we do not choose, 5
|
||||
recorded anywhere with their reasoning, and 171 that pass in silence when reached.*
|
||||
|
||||
A bound is a number chosen at compile time that decides how much of something the code
|
||||
can hold. `const maximum_devices = 64`. `var below: [64]Range`. `var blob: [512]u8`.
|
||||
Different units — devices, firmware memory-map entries, bytes of a USB descriptor — but
|
||||
one shape, and one recurring way of going wrong.
|
||||
|
||||
## Where a bound lives
|
||||
|
||||
**Where the thing it bounds lives.** A driver's transfer-ring size belongs to that
|
||||
driver; a protocol's payload cap belongs to that protocol; the kernel's task-table size
|
||||
belongs to the kernel. There is no central list and this document does not propose one.
|
||||
|
||||
`system/parameters.zig` is not a counter-example. It is kernel-only, and it exists for a
|
||||
specific historical reason: tunables had accumulated inside the loader↔kernel handoff
|
||||
contract, and splitting them out kept that contract to what it actually is. It is a
|
||||
tidying of one file's contents, not a registry the rest of the system reports to.
|
||||
|
||||
This matters for the mechanism below. An earlier draft had every bound declared through
|
||||
a shared `bounds` module — which would have meant adding a dependency to roughly eight
|
||||
package manifests, including `library/protocol`, which deliberately depends on nothing.
|
||||
That is a coupling the problem does not require: a bound is a local fact about local
|
||||
storage, and the only thing worth sharing is the *shape of the statement*, not a module.
|
||||
|
||||
## The declaration
|
||||
|
||||
A structured doc comment, immediately above the declaration, in the file that owns it:
|
||||
|
||||
```zig
|
||||
/// bound: logical CPUs the kernel tracks
|
||||
/// decided-by: hardware
|
||||
/// protects: the per-CPU bookkeeping arrays, which are sized at compile time
|
||||
/// at-limit: degrade — surplus cores are left parked, never brought online
|
||||
/// observed-by: platform.cpusDropped() -> the WARNING at kernel.zig:281
|
||||
pub const maximum_cpus = 128;
|
||||
```
|
||||
|
||||
Five fields, all mandatory:
|
||||
|
||||
| Field | Answers |
|
||||
|---|---|
|
||||
| `bound` | what is counted, in plain words |
|
||||
| `decided-by` | `hardware`, `external`, or `ours` — who chooses how large it gets |
|
||||
| `protects` | what this ceiling defends against |
|
||||
| `at-limit` | `refuse` / `degrade` / `truncate` / `grow`, and the detail |
|
||||
| `observed-by` | how an operator finds out it was reached |
|
||||
|
||||
`decided-by` is the classification the audit turned on. `hardware` means the machine
|
||||
chooses — PCI functions, CPUs, ACPI rows, memory-map entries. `external` means a file,
|
||||
disk structure or peer chooses. `ours` means we do: a stack size, a tick rate, our own
|
||||
protocol's payload. A fixed bound on the first two is a defect rather than a tunable.
|
||||
|
||||
## What the build step enforces
|
||||
|
||||
A step in `build.zig` reads the tree and fails on:
|
||||
|
||||
1. **A bound with no declaration.** A fixed-size array or a `maximum_*`/`max_*` constant
|
||||
with no `bound:` block above it. The 235 that exist today are allowlisted by
|
||||
file+line+name, so only newly written ones are gated — the rule can land without a
|
||||
235-site sweep in front of it.
|
||||
2. **A missing field.** All five or it fails. This alone is the 97 bounds with no
|
||||
comment at all and the 171 with no observability.
|
||||
3. **An unspeakable `at-limit`.** The vocabulary is closed. There is no `silent`, no
|
||||
`drop`, and nothing meaning *allow*. `truncate` is legal only with a marker the
|
||||
reader can see — `klog_maximum_message` qualifies because the record carries
|
||||
`klog_flag_truncated`; the USB configuration descriptor cut at 512 bytes does not,
|
||||
because nothing records that anything was lost.
|
||||
4. **A stale allowlist entry.** If a listed bound is fixed or deleted, its entry goes
|
||||
too, so the list can only shrink.
|
||||
|
||||
Because this is text and not a Zig type, it also covers `boot/`, which imports almost
|
||||
nothing, and `tools/*.py`, where the audit found bounds as well. One mechanism, whole
|
||||
tree, no new dependency edges.
|
||||
|
||||
## Coupled bounds
|
||||
|
||||
Two numbers that must agree, agreeing in code rather than in a comment — no module
|
||||
needed, just a `comptime` block where one of them lives:
|
||||
|
||||
```zig
|
||||
comptime {
|
||||
if (maximum_domains != devices_broker.maximum_devices)
|
||||
@compileError("iommu.confined is indexed by device id; an id past its end is " ++
|
||||
"left unconfined while confineDevice still reports success");
|
||||
}
|
||||
```
|
||||
|
||||
`maximum_domains = 64` and `maximum_devices = 64` agree today only by a sentence in a
|
||||
comment, and the agreement fails open. This is the clause with a live hole behind it,
|
||||
and the reason raising `maximum_devices` alone would be a privilege escalation rather
|
||||
than a fix.
|
||||
|
||||
## The worked bad case
|
||||
|
||||
`devices_broker.maximum_devices`, which had no comment at all:
|
||||
|
||||
```zig
|
||||
/// bound: device nodes for the whole machine — firmware-discovered plus registered
|
||||
/// decided-by: hardware
|
||||
/// protects: nothing; this is a sizing guess about someone else's computer
|
||||
/// at-limit: refuse — ENOSPC from device_register, dropped++ during discovery
|
||||
/// observed-by: kernel.zig:203 counts discovery drops only, NOT runtime refusals
|
||||
const maximum_devices = 64;
|
||||
```
|
||||
|
||||
Writing it out is the argument. `decided-by: hardware` alongside a `protects` that
|
||||
admits there is no threat describes a bound that should not be fixed at all, and
|
||||
`observed-by` cannot be filled in honestly. The build step does not reject this — it
|
||||
makes it impossible to write down without noticing.
|
||||
|
||||
## The work-list this produces
|
||||
|
||||
`decided-by` is machine-readable, so the sweep is a query: every `hardware` or
|
||||
`external` bound whose `at-limit` is not `grow`. That is 139 of the 235, and it is the
|
||||
order the fixing takes — by class, not by our guesses about which machines get run.
|
||||
Reachability is exactly what a new computer changes; the Ryzen's cap was unreachable
|
||||
until it wasn't.
|
||||
|
||||
## What this does not do
|
||||
|
||||
**It is a statement, not a proof.** Nothing checks that the code does what `at-limit`
|
||||
claims. The enforcement is completeness and vocabulary: you cannot leave the question
|
||||
unanswered, and you cannot answer it with "silently".
|
||||
|
||||
**It carries no occupancy.** Knowing `system/configuration/protocol.csv` sits at 50 of
|
||||
64 grant rows still needs a counter and somewhere to report it. The declarations make
|
||||
that cheap to add later; it is not here.
|
||||
|
||||
**It resizes nothing.** Declaring `maximum_devices` honestly does not make the Ryzen
|
||||
work. It makes the next machine's failure legible, and it names the 139 that need real
|
||||
fixes.
|
||||
|
||||
**Enforcement is at build time, not compile time.** A Zig type could have made a missing
|
||||
field a compile error. That version needed the shared module, and the module was not
|
||||
worth the coupling — so a missing field is a failed build step instead. In practice both
|
||||
mean `zig build` stops; the difference is which stage prints the message.
|
||||
@@ -0,0 +1,173 @@
|
||||
# Device authority: the kernel stops keeping an inventory
|
||||
|
||||
*Design, drafted 2026-08-07. Not implemented. Prompted by a real machine: an AMD
|
||||
Ryzen desktop enumerated more PCI functions than the kernel's device table would
|
||||
hold, and the xHCI and SATA controllers were refused registration — so the
|
||||
machine booted to the compositor with no USB and no storage.*
|
||||
|
||||
The kernel keeps a table of every device userspace discovers. It is a fixed
|
||||
array of 64 descriptors, 344 bytes each, and a second cap allows any one parent
|
||||
16 children. Neither number is written down anywhere as a decision:
|
||||
`maximum_devices` has no comment and never reached `parameters.zig`, where every
|
||||
other tunable in this kernel lives with its reasoning attached.
|
||||
|
||||
Raising them is not the fix. The numbers are wrong because the *table* is wrong:
|
||||
it is an inventory of hardware, and an inventory of hardware is not something a
|
||||
kernel needs. This document proposes replacing it with capabilities, which
|
||||
removes the ceiling rather than moving it.
|
||||
|
||||
## What the kernel actually uses
|
||||
|
||||
Every read of a device descriptor from the kernel proper, exhaustively:
|
||||
|
||||
| Used for | What it needs |
|
||||
|---|---|
|
||||
| `mmio_map`, `io_read`/`io_write` | the physical range, to check the mapping falls inside it |
|
||||
| `irq_bind`, `msi_bind` | the GSI, and that the caller owns the device |
|
||||
| `dma_bind` | that the caller owns the device |
|
||||
| IOMMU confinement | the **PCI BDF**, to key a domain |
|
||||
| `device_register` | the parent's ranges, for the containment check |
|
||||
| the boot display seed | one framebuffer window |
|
||||
|
||||
That is: **physical ranges, interrupt numbers, and one BDF.** Vendor, device and
|
||||
subsystem ids, class triples, the human-readable names, the parent links, the
|
||||
bus numbers — the kernel stores all of it and reads none of it. It is held so
|
||||
that `device_enumerate` can hand it back to user space, which is the whole
|
||||
mistake in one sentence: the kernel is acting as a distribution mechanism for
|
||||
data it does not use.
|
||||
|
||||
## The split
|
||||
|
||||
Three concerns are tangled in one table.
|
||||
|
||||
**Platform bring-up** — timers, LAPIC/IOAPIC, CPU topology, the ECAM window, the
|
||||
framebuffer the firmware left. The kernel derives these from ACPI before user
|
||||
space exists and needs them to function. They never belonged in the device table
|
||||
and mostly are not (the platform block in `kernel.zig` is separate); this
|
||||
document does not change them.
|
||||
|
||||
**Resource authority** — which task may map which physical range, receive which
|
||||
interrupt, touch which ports. This *must* stay in the kernel. It is the one
|
||||
grant that cannot be audited after the fact: a process that maps arbitrary
|
||||
physical memory owns the machine, page tables and IOMMU structures included.
|
||||
This is memory protection, not device management, and it is why the answer is
|
||||
not simply "move it all to the device manager".
|
||||
|
||||
**Device inventory** — what exists, what it is, how it is arranged, which driver
|
||||
should bind it. This is `device-manager`'s job and is already half there: it
|
||||
loads `devices.csv`, matches, spawns drivers, and receives `child_added` reports
|
||||
over its own protocol. The kernel table duplicates what those reports already
|
||||
carry.
|
||||
|
||||
## The proposal: a resource is a capability
|
||||
|
||||
Device resources join endpoints, shared memory and DMA regions as a kind in the
|
||||
handle table.
|
||||
|
||||
1. **Roots.** At boot the kernel mints capabilities for the windows it learned
|
||||
from firmware — the ECAM range, the framebuffer, the legacy port space — and
|
||||
hands them to the first bus drivers. This is the only place device knowledge
|
||||
enters the kernel, and it comes from ACPI, not from a driver's say-so.
|
||||
2. **Subdivision.** A bus driver enumerating hardware derives a narrower
|
||||
capability from one it holds: `resource_derive(cap, kind, start, len) → cap`.
|
||||
The kernel checks the sub-range lies inside the capability being subdivided —
|
||||
the same containment rule as today (`devices-broker.contains`), but checked
|
||||
against *one capability the caller demonstrably holds* rather than by walking
|
||||
a global tree.
|
||||
3. **Delegation.** The driver passes that capability to the child driver over
|
||||
IPC. Cap-passing already exists; this is the mechanism `subscribe` and
|
||||
`attach_scanout` already use.
|
||||
4. **Use.** `mmio_map`, `irq_bind`, `msi_bind`, `io_read`/`io_write` and
|
||||
`dma_bind` take a capability handle instead of `(device_id, resource_index)`.
|
||||
Possession *is* the authority — there is nothing to look up and no ownership
|
||||
table to consult.
|
||||
|
||||
Exclusivity stops being a broker refusing a second claimant and becomes the
|
||||
ordinary property of a capability: only one process was given it.
|
||||
|
||||
## What this buys
|
||||
|
||||
**No ceiling.** There is no table to size, so no machine is too big. The
|
||||
Ryzen's enumeration stops being a limit to tune and becomes what it is — a fact
|
||||
about a computer.
|
||||
|
||||
**Reclamation, free.** `count` in the broker today only ever increases;
|
||||
`releaseAllOwnedBy` clears a dead driver's *claims* but never its entries. A
|
||||
driver that crashes and is restarted re-registers its children and consumes the
|
||||
table again — reachable today without any malice, given the device manager
|
||||
restarts drivers by design. Capabilities die with the task.
|
||||
|
||||
**No quota needed.** The per-parent cap exists to stop one claimant looping
|
||||
`device_register` and filling the shared table, because a zero-resource child
|
||||
sidesteps the containment check. With no shared table there is nothing to
|
||||
exhaust; a process can only ever subdivide what it was given, and its handles
|
||||
are already bounded per task.
|
||||
|
||||
**A smaller kernel.** Three syscalls leave the ABI, one narrower one arrives,
|
||||
and `devices-broker.zig` largely disappears along with both constants.
|
||||
|
||||
**The discipline the rest of the system already uses.** "The claim is the
|
||||
capability" is written in the driver documentation as if it were already true.
|
||||
This makes it true.
|
||||
|
||||
## Syscall surface
|
||||
|
||||
- `device_enumerate` — **retires.** It exists only to read the kernel's table.
|
||||
Callers ask `device-manager`, whose protocol already reserves an `enumerate`
|
||||
verb. Note this is a public-ABI change: `vdso.md` documents it.
|
||||
- `device_register` — **splits.** The kernel half becomes `resource_derive`; the
|
||||
publication half ("this device exists, here is what it is") becomes an IPC
|
||||
message to `device-manager`, which is where the inventory belongs and where
|
||||
`child_added` already carries the same facts.
|
||||
- `device_claim` — **dissolves into possession**, except for the IOMMU (below).
|
||||
- `mmio_map`, `irq_bind`, `msi_bind`, `io_read`, `io_write`, `dma_bind` — keep
|
||||
their names and semantics; their first argument becomes a capability handle.
|
||||
|
||||
## The open question: where the IOMMU attaches
|
||||
|
||||
This is the one place the kernel still needs device *identity* rather than a
|
||||
range. `confineDevice(device_id, bdf, owner)` builds a domain keyed by PCI BDF,
|
||||
attaches it, and the claim is rolled back if confinement fails — deliberately,
|
||||
so a device that cannot be confined is never driven.
|
||||
|
||||
Three options, none obviously right:
|
||||
|
||||
1. **The BDF rides the capability.** A memory capability derived for a PCI
|
||||
function carries its BDF, and the kernel confines on first `mmio_map` or
|
||||
`dma_bind`. Keeps the syscall count down; means a capability is no longer
|
||||
purely a range.
|
||||
2. **An explicit `device_attach(cap, bdf)`.** Honest and visible, but it puts a
|
||||
BDF — a fact about PCI — into a kernel interface that otherwise knows nothing
|
||||
about buses, and something must stop a caller naming a BDF that is not
|
||||
theirs.
|
||||
3. **The root PCI capability carries the segment, and derivation computes the
|
||||
BDF.** Purest, but only works for PCI and the kernel would be parsing bus
|
||||
topology, which is precisely what this document is trying to stop.
|
||||
|
||||
My inclination is (1), because confinement is a property of the resource being
|
||||
granted rather than a separate action, and because it keeps the "possession is
|
||||
authority" story intact. It needs the derivation call to know it is carving a
|
||||
PCI function, which is a wart worth arguing about.
|
||||
|
||||
## What this does not solve
|
||||
|
||||
- **Hot-plug and removal.** Capabilities die with their holder, but a device
|
||||
that physically disappears while a driver lives is the device manager's
|
||||
problem and unchanged by this.
|
||||
- **Quotas on physical memory.** Nothing here bounds how much a driver maps; it
|
||||
bounds only *what* it may map. That was already true.
|
||||
- **The device tree as a published thing.** `/system/devices` remains a device
|
||||
manager concern, as the file-system hierarchy already assumes.
|
||||
|
||||
## Migration sketch
|
||||
|
||||
Not a plan yet — the phases need sizing once the IOMMU question is settled.
|
||||
Rough shape: introduce the capability kind and `resource_derive` alongside the
|
||||
existing table; convert the resource-consuming syscalls to accept either form;
|
||||
move the inventory into `device-manager` and convert its clients off
|
||||
`device_enumerate`; then delete the table, the two constants and the three
|
||||
syscalls in one flag-day, as the `ServiceId` retirement did.
|
||||
|
||||
Every driver is affected, so the QEMU suite is the arbiter at each step, and the
|
||||
Ryzen is the acceptance test — it is the machine that found this, and the one
|
||||
that proves it fixed.
|
||||
@@ -150,6 +150,57 @@ they are wire values a Rust program needs verbatim. What stays private in
|
||||
`abi.zig` is exactly the thing the vDSO exists to hide: the `SystemCall`
|
||||
numbers and the trap convention.
|
||||
|
||||
## Errors: the errno space
|
||||
|
||||
A failed call returns `-errno`. The runtime detects failure the way Linux
|
||||
does — a return value in the top 4096 — so every code stays inside 1..4095.
|
||||
These are **public**: unlike the call numbers, a caller must be able to read
|
||||
them verbatim, and they are the same vocabulary whether the number came from
|
||||
the kernel or from a user-space provider answering over IPC.
|
||||
|
||||
They are defined once in `system/abi.zig`. The kernel restates them in
|
||||
`system/kernel/ipc-synchronous.zig` and the envelope restates the
|
||||
provider-facing subset in `library/protocol/envelope/envelope.zig` (the
|
||||
`protocol` package deliberately depends on nothing, so it cannot import
|
||||
`abi`); a comptime check in `library/device/driver/driver.zig` makes drift a
|
||||
compile error.
|
||||
|
||||
| # | Name | Meaning |
|
||||
|---|------|---------|
|
||||
| 1 | `EBADF` | bad handle |
|
||||
| 2 | `E2BIG` | an argument exceeds its maximum (a message, a descriptor's resource count) |
|
||||
| 3 | `EFAULT` | buffer unmapped, or outside the user half |
|
||||
| 4 | `ENOENT` | no such name |
|
||||
| 5 | `ENOSPC` | a kernel table is full (handles, devices) |
|
||||
| 6 | `ENOMEM` | out of memory |
|
||||
| 7 | `EPEER` | the peer died before replying — its process exited or was killed |
|
||||
| 8 | `ESRCH` | no such process |
|
||||
| 9 | `EPERM` | not permitted: the caller is not the owner or supervisor |
|
||||
| 10 | `ENOSYS` | this protocol has no such operation |
|
||||
| 11 | `EPROTO` | malformed packet: shorter than the verb it names |
|
||||
| 12 | `EBUSY` | the thing asked for is held by someone still alive |
|
||||
| 13 | `ENODEV` | no such device id |
|
||||
| 14 | `ECHILDREN` | this parent already holds as many children as it can |
|
||||
| 15 | `ERANGE` | a resource escapes the window it must fall inside |
|
||||
| 16 | `ECONFINE` | the device could not be placed under IOMMU translation |
|
||||
|
||||
`EPEER` is the one with no POSIX counterpart and it is worth stating plainly:
|
||||
synchronous IPC blocks the caller until the server replies, so the caller
|
||||
needs an answer for "the server died while I was waiting." It is not a
|
||||
transport error and not a refusal — the request may well have been carried
|
||||
out — it says only that no reply is coming. A client that treats it as
|
||||
"retry" can duplicate work; the honest response is to re-resolve the protocol
|
||||
name, because the provider it held is gone.
|
||||
|
||||
**A refusal names the rule that refused it.** This is a rule and not a
|
||||
courtesy. `device_register` alone can fail six ways, and until each got its
|
||||
own code a bus driver could only report "refused" — which is how an AMD
|
||||
desktop came to boot with a working display, no USB and no storage, with
|
||||
three independent causes indistinguishable in the log. See
|
||||
[fixed-bounds-audit.md](../fixed-bounds-audit.md). A new failure mode that
|
||||
does not fit an existing code gets a new one here rather than borrowing the
|
||||
nearest.
|
||||
|
||||
## Enforcement, and an honest threat model
|
||||
|
||||
Renumbering only has teeth if the kernel **refuses syscalls that don't come
|
||||
|
||||
Reference in New Issue
Block a user