547d0ec46b6f5a825af215d6ff6a2645e1d8c4cb
9
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
547d0ec46b |
docs: device authority — the how, and the run that deletes the ceilings
device-authority.md is rewritten as an implementation design rather than a rival to device-manager.md. The what was already settled there in 2026-07: structure in the manager, authority in the kernel, and delegation as the step after hello. This is the how, plus the two decisions that paragraph leaves open. Decision 1: the manager claims, it is not granted. device-manager.md says "claims (or is granted)"; claiming wins because the manager runs before any driver exists and takes the seeded devices unopposed, leaving nothing unheld to race for. One new call, device_transfer(device_id, task_id), checks only that the caller holds the device — no names in the kernel, no attestation. The alternative put a binary path inside the kernel, and the kernel should hold only what cannot safely live in user space. The residual is stated rather than hidden: authority rests on the manager claiming first, which init.csv makes an operator-visible ordering rather than an attacker- controlled one, and the enforced version arrives with the spawn capability drivers.md already names as missing. Decision 2: the kernel stops holding inventory. It reads three things out of a descriptor — physical ranges, interrupt numbers, one PCI BDF — and stores the rest only so device_enumerate can hand it back. Devices with no resources leave the kernel entirely: a USB device conveys no mapping authority, so there is nothing to enforce. That is also the case which sidesteps containment, and therefore the reason a shared cap existed. Decision 3: no shared ceiling. The table becomes dynamic — it is built after heap.init, so nothing ever prevented it — and the two invented numbers go. A per-holder quota replaces them, because dynamic storage with no bound moves the ceiling to the kernel heap, which is shared and fatal rather than partial. A bound charged to whoever caused it is isolation. Two earlier drafts of this document are gone: one gave init the root grants, the other proposed extracting a firmware-framebuffer driver. Both were wrong and both are recorded as wrong in the run plan's settled list — the framebuffer is not a device, it is where pixels go until a real display driver announces itself. Run 2 is nine steps, ordered so the suite stays green throughout: build and prove the transfer mechanism, move the five claimants across one at a time, then the flag day, then the inventory, then the ceilings. |
||
|
|
5d55217212 |
usb: a slot the controller granted is always handed back
Both device-setup paths issued a successful Enable Slot and then returned null if allocateDevice failed, without disabling it. A slot the driver forgets is one the controller never reissues, so each attempt lost one permanently for the boot. The hub path did it with no log line at all. Both now release the slot through a shared disableSlot, extracted from tearDownDevice, and the hub path warns like the root-port path does. tearDownDevice also now frees the interface list. That allocation arrived with the previous commit, so an unplug would have leaked it — found while reading the teardown path for this fix rather than by a test. No regression test, and it is recorded as open question 5 rather than implied. After the slot count became the controller's own figure, reaching this path needs more devices than the controller has slots: QEMU offers four against sixty-four. What was verified is that the new path RUNS correctly — pinning tracking to 2 with four devices attached produced "port 6 setup: no free device slot", the first two devices enumerated normally, and no Disable Slot error or timeout appeared, which is how disableSlot reports failure. Suite 116/116. |
||
|
|
729b40ece7 |
usb: a device has as many interfaces as it declares
max_interfaces was 4. A composite device — a headset, a webcam with audio, a dock, a multifunction printer — routinely has more, and the fifth did not merely go missing. parseConfiguration's cap branch had no `else`, so when the count was reached `current` kept pointing at interface 3 and the fifth interface's endpoint descriptors were appended to interface 3's array. A class driver bound to interface 3 could then be handed an endpoint belonging to something else entirely, and subscribe or bulk-transfer on it. The alternate-setting arm one line above cleared `current` correctly, which is what the cap branch should have done. Interfaces are now counted from the block in a first pass and allocated to exactly that number, so the ceiling is bNumInterfaces' u8 — the USB specification's. The missing `else` is added too, though after this the bug is unreachable by construction: interface_count cannot reach interfaces.len mid-parse when the list was sized from the same walk. max_configured_endpoints was max_interfaces * max_endpoints_per_interface = 16, a derived guess that moved whenever either input moved. It is now 31, which is the xHCI specification's own limit: a Device Context holds a slot context plus at most 31 endpoint contexts, because the Context Entries field addressing them is 5 bits. max_endpoints_per_interface stays at 4 with its reason recorded — the usb-transfer wire protocol reports exactly max_reported_endpoints (4) per interface, so widening it alone would change nothing a class driver sees. Lifting it is a protocol change. No direct test, and that is written down as open question 5 rather than glossed. The parser is pure and wants a host unit test, but usb-xhci-library.zig imports memory, mmio and time so it cannot be a standalone test root, and QEMU offers nothing that reaches the path — the largest device available is usb-audio,multi=on at 2 interfaces and 211 bytes. The alternate-setting path that shares the same `current = null` logic is exercised by that device. Suite 116/116. |
||
|
|
6328823ef1 |
usb: a configuration block is as long as the device says it is
The driver read the first 512 bytes of a configuration block into a fixed buffer and parsed those. The block's length is the device's own choice (wTotalLength, a u16), so anything larger was silently cut: interfaces past the cut did not exist as far as the host was concerned, while the SET_CONFIGURATION that follows still configured the device for all of them. A headset is 500-900 bytes, a UVC webcam 1-3 KB, a multifunction printer 600+. Now allocated at the declared length, so the ceiling is the field's u16 — the specification's number rather than one of ours. A block shorter than its own 9-byte header is refused rather than trusted. The bring-up line reports the declared length and the bytes actually read, so a truncation can never again be invisible, and usb-large-descriptor asserts they match with a backreference. That case has an honest limit, recorded in its comment: QEMU cannot produce a block over 512 bytes. The boot keyboard, mouse and stick are 34-44, and the largest device available is usb-audio in multi-channel mode at 211 — which is exactly why the suite never caught this, and why it cannot now reproduce the original trigger. What it does catch is the class: any clamp below the attached device's block fails it, verified by pinning the buffer to 128 and watching "config block 211 bytes, read 128" turn the case red. Suite 115 -> 116. |
||
|
|
69fbef40c0 |
usb: the controller says how many device slots it has
max_devices was 8, with the comment "QEMU presents a handful; a fuller machine would grow this" — a number chosen against the test rig, waiting for a real machine, which is the pattern docs/bounds-track-plan.md exists to stop. The driver already knew the true figure. It reads HCSPARAMS1.MaxSlots at bring-up and writes it straight into op_config, so every slot the controller offers has always been *enabled*; only the array tracking them was 8. QEMU's xHCI reports 64, so seven eighths of the controller was live and invisible, and the ninth device — a keyboard, mouse, webcam, headset, hub and two sticks reach that without trying — disappeared on a hub-attached path that logs nothing at all. The array becomes a slice allocated from max_slots at bring-up. A controller claiming zero slots cannot address anything, so that is now a dead controller rather than an empty allocation failing mysteriously later. The Device Context Base Address Array is a page, 511 usable entries, so it already covered the 255-slot maximum. The bring-up line reports both numbers, and usb-hid asserts they are equal with a backreference rather than a magic number, so the test cannot drift from the hardware. Pinning tracking back to 8 fails it: "64 slots, tracking 8". Suite 115/115. |
||
|
|
f4eb88e7d2 |
build: a new compile-time ceiling declares itself or does not land
The convention that tunables live in system/parameters.zig with their reasoning attached predates this and got 2% compliance — 5 of 235. A convention with no teeth is how a bare `const maximum_devices = 64` reached an AMD desktop and cost it USB and storage. This is the same rule with a gate behind it. tools/check-bounds.py finds every bound-shaped declaration — a `maximum_*` const with a literal value, or a type with a literal array length — and requires the five-field block above it: what it counts, who decides its size, what it protects, what happens at the limit, and how anyone finds out. The at-limit vocabulary is closed: refuse, degrade, truncate, grow. There is deliberately no way to spell "silent", no way to spell "drop", and nothing meaning "allow", so the behaviours that did the damage cannot be written down. Truncation is legal only carrying a marker the reader can see, which is why klog_maximum_message qualifies and a USB descriptor cut at 512 bytes does not. An array length that names a declared bound is not itself a bound; only literal lengths are flagged, which pushes ceilings toward having names. The 273 that predate the rule are allowlisted so this lands without a tree-wide sweep in front of it, and that list may only shrink: declaring a bound means deleting its line, and the check fails on a stale entry too. Nothing may be added. Wired into `zig build test` and available alone as `zig build bounds`. Not in the default build — it reads the whole tree, and a red bounds check should not stop you booting a kernel. Five are now declared rather than allowlisted. Writing them out is its own argument: maximum_devices reads "protects: nothing — this is a sizing guess about someone else's computer", and maximum_tasks now carries the fact that it has been raised twice, each time by something that outgrew it. Verified the gate refuses an undeclared bound, a declared one using forbidden vocabulary, and an allowlist entry that has since been declared. Suite 115/115. |
||
|
|
568823a4fb |
docs: L1 was wrong — reclamation is not a death-sweep problem
The first step of the unattended run was "a dead task's registrations die with its claims". Implementing it would have broken the restart path it was meant to protect. The broker keeps entries on purpose: they describe hardware, which did not go away when a driver died. And device ids must stay stable across a bus restart, because device-manager dedupes re-reports by device_id so a restarted bus does not spawn a second driver instance — stability that comes from the idempotency scan returning the existing id. Removing entries on death would hand a restarted bus fresh ids and duplicate every driver. Everything else a task holds is already reclaimed on every path out: IRQ bindings, IOMMU domains, DMA regions, then its claims. The leak the audit found is real but has two other sources: a device that genuinely goes away has no retirement path, and a bus that enumerates differently on restart strands its old entries. Both belong to the device manager's inventory, and both need the id-stability question settled first — tombstone-and-reuse aliases ids another process still holds, generation tagging changes the id encoding, which is ABI. Recorded as an open question rather than guessed at. |
||
|
|
4398eb7cc4 |
docs: scope the bounds track's unattended run
Six steps an agent can execute: reclamation, the bounds build check, and the four user-space USB/xHCI bounds where the hardware already reports the number we guessed. Deliberately excluded: the authorisation gate and moving the device inventory to the manager. Both decide whether the OS is secure and both are a direction rather than a specification — what a device capability is, which syscalls change, what replaces device_claim for its seven callers. They want a design session, not an agent. Three open questions are written down rather than guessed: device_enumerate most likely narrows to the firmware-discovered roots rather than retiring (the manager cannot ask itself for the PCI host bridge); a manager restart has no re-enumerate handshake, so it comes back blind while its buses live; and "add adversarial tests" is not executable until the attacks are named. |
||
|
|
a86559648e |
kernel: a refusal names its rule, and two bounds stop failing open
An AMD Ryzen booted to a working compositor with no USB and no storage, and the log said only "register refused". A tree-wide audit of every compile-time ceiling followed: 235 of them, 139 on quantities the machine or a file decides rather than us, 5 documented anywhere, 171 silent when reached. docs/fixed-bounds-audit.md has the inventory. Errno attribution. The errno space was split between the kernel and the envelope, free to drift; it is now one list in system/abi.zig, restated on both sides, with a comptime check in library/device/driver where the two halves are visible. device_register's six refusals and device_claim's three are distinct codes, so a bus driver can say which rule stopped it, and BadParent splits into NoSuchParent and NotYourParent. pci-bus reconciles found against registered instead of counting refused functions as found. Idempotency ordering. The child cap was checked before the identity match, so a restarted bus was refused its own devices — the supervision restart the system leans on ratcheted toward a degraded machine. A re-registration consumes no slot and is now admitted first. IOMMU fail-closed. confineDevice returned success for a device id past the confinement table, leaving the device outside every domain while the caller believed it confined — unreachable only while ids stop at 64, which both the inventory move and a hardware-reported domain count would change. It refuses now, and the coupling to the broker's device cap is a comptime assert rather than a sentence in a comment. PCI apertures. The bridge's MMIO apertures are derived from the holes in the firmware memory map, and the derivation copied sub-4 GiB entries into a fixed [64] array and skipped the rest. A skipped region is not merely lost: the gap finder concludes it is free, so a real machine's 60-200 entry map yields an aperture over live RAM, and containment then admits a child BAR covering kernel memory. Rewritten to walk the map in place, with the hole finder extracted as a pure function and driven by a synthetic 100-entry map in a new test case. Both new tests were verified to fail on the old code. parameters.zig gains the rationale it was missing and loses a stale sentence pointing at the wrong file; vdso.md documents the errno space, including EPEER, which had no written meaning anywhere. docs/os-development/bounds.md is how a ceiling is declared from here. docs/bounds-track-plan.md is the plan to remove the ones we invented. Suite 114 -> 115. |