af47d4198946e6b73314461f425608d2944bf044
266
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
af47d41989 |
volume-manager: removal supersedes a pending restart in the poll
Inline V4 review (the boundary-review workflow stalled): the poll ran a due fat-restart before the presence check and returned, so a fat death followed by a device removal would respawn fat against the now-dead channel and churn until the crash cap before the removal was noticed. Reorder: check the specific device's presence first (unmount if gone), and only fire a due restart once the device is confirmed present. Neutral: fat-mount, volume-removal, amd-iommu-usb-storage green. Noted V4 limitations (not fixed here, edge cases outside the user unplug case): a usb-storage DRIVER crash (device stays, driver restarts with a new endpoint) leaves fat holding a dead channel — the device is still present so removal is not detected; fat would need to observe its channel death and exit. Deferred with the medium_changed subscription and multi-volume. |
||
|
|
e3ec9fa668 |
volume-manager: try every mass-storage entry, watch the one we opened
The full suite caught a V4 regression: under AMD-Vi the device-manager tree carries more than one mass-storage-identity entry (a phantom no driver is bound to, which answers a consumer hello with NO channel). V4 split presence from acquisition and picked the FIRST identity match blindly, so it kept helloing the phantom (device 27) and never reached the real storage (device 31). V3's inline loop had skipped no-channel entries with `orelse continue`; the split lost that. Restore it: openAnyStorage tries each matching entry and takes the first whose channel opens, recording its device id. Removal detection then watches THAT specific device id leave the tree (isDevicePresent), not "any mass-storage" — so a phantom that never leaves cannot mask a real removal. Both are bare enumerates; the hello only happens while bringing a volume up. Green: amd-iommu-usb-storage, fat-mount, volume-removal. |
||
|
|
9e67a74232 |
volume-manager: the removal lifecycle — a pulled stick unmounts (V4)
The volume manager stops probing-once and polls storage presence for the life of the boot: findStorageDevice enumerates the device-manager tree (presence only, no consumer-hello, so it is cheap and leaks nothing). The volume is now a field that goes null and back — the whole lifecycle: - storage present + no volume -> open the channel, probe, confine + spawn the filesystem (openStorage is the one consumer-hello, on the insertion edge); - storage gone + have volume -> kill the filesystem (its mounts retire via the kernel dead-backend sweep), close the dead channel, clear the volume; - fat crash -> the same supervised backoff/cap as before, folded into the poll (one timer). This also subsumes the V3-review leak fix (no per-poll consumer-hello) and the no-volume retry (a present-but-unreadable device keeps polling). The user's case — pull the boot stick, plug it back — is a DEVICE unplug (the stick IS the device), so the mass-storage child leaves the device-manager tree and the poll catches it. volume-removal asserts the unmount and discriminates: against the V3 probe-once volume manager the removal is never noticed (0/1). The re-mount on replug is the VM's bringUpVolume firing when the device returns — correct and in place, but not QEMU-testable here: device_add of usb-storage to the boot xHCI controller is not re-presented to the guest (no port-connect on any port), a harness quirk, not a VM issue. On real hardware the bus's per-tick port poll catches a reconnect (H1 proves reconnect on a second controller); bench-verify the full round trip. |
||
|
|
7c6ed2ca09 |
volume-manager: harden the probe and supervision from the V3 review
Five confirmed defects from the boundary review: 1. (security) The VM never checked a partition fit inside the device, so a crafted MBR could hand the driver a range whose base+lba wraps past a u32 — panicking usb-storage in a loop, and at multi-volume overlapping a neighbour. This is the exact invariant the clamp's overflow-safety rests on. partition.firstVolume now skips any entry that runs past the device (host-tested), establishing the invariant where the untrusted bytes are first read. 2. (leak) The probe re-acquired a fresh block channel on every 500 ms retry, leaking a handle each time on a medium-absent device. The channel is now acquired once and kept. 3. (wedge) A failed spawn or defineRange stranded the volume with no retry; both now arm a backoff restart. 4. (loop) fat respawn had no exit-reason gate, no backoff, no crash-loop cap — a faulting filesystem respawned in a zero-delay loop, and a clean exit was resurrected. Supervision now mirrors the device manager: a clean exit is not restarted, a fault backs off, three fast deaths give up. 5. (removable) A device that parsed to no volume was terminal; it now keeps polling so an inserted medium is picked up — the removal-lifecycle trigger. Known limitation (noted, not fixed here): if the VM itself crashes and init restarts it, the orphaned fat keeps serving vfs while the new VM spawns a second fat whose bind is refused — the same "manager restart re-learns the world" gap the device manager also defers. The old fat keeps storage working. Neutral: partition unit tests + fat-mount, volume-probe, block-range, logger all green. |
||
|
|
301bdcaf5b |
volume-manager: the flip — fat is spawned, confined, and handed its channel (V3b)
The load-bearing step. The FAT service stops acquiring its own volume: the volume manager spawns it (per volume), defines its partition range on the storage driver BEFORE it runs, and answers its startup hello with the range-confined block channel over a new volume-manager protocol. fat never finds its storage by name and never sees the whole device — establishment by lineage, one layer up from the driver tree. - New library/protocol/volume-manager: one verb, hello(volume-id) -> the block channel as the reply capability (the P0 reply-cap path). - The volume manager becomes the confinement CONTROLLER: it defines the first range on usb-storage, so no other party can confine a filesystem. It supervises the filesystems it spawns and respawns one on death (the reap- and-rebuild the device manager proved, one layer up). - fat: drops acquireVolume(device-manager); hellos the volume manager for its channel; reads its volume id from argv[1]. main takes process.Init now. - init.csv no longer spawns fat (the volume manager does); protocol.csv rewires fat to be supervised by the volume manager (bind vfs, open volume-manager) and drops fat open device-manager. - The block-range fixture boots registry + device-manager only (not the full tree), so the volume manager is absent and the fixture stays the sole confinement definer — otherwise the volume manager would take the controller first and refuse it. Verified end to end (VM probes -> spawns fat -> confines it -> hands over the channel -> fat mounts) and neutral: 18/18 across the fat family, logging, shutdown, both IOMMU variants, usb restart, vfs, conformance, confinement. |
||
|
|
d56b1b81c0 |
volume-manager: discovery and probe — V3a
The storage layer gains its policy home (storage-architecture.md): a new system/services/volume-manager, spawned by init, that acquires the mass- storage block channel through the device manager (the same lineage a filesystem uses), reads block 0, and parses the first volume out of it. The partition-table walk that lived in the FAT engine moves here, above the driver where it belongs (partition.zig, host-tested: MBR entry, bare-FAT, no-signature). Identity is the MBR disk signature + partition index — the weak rung of the ladder; GPT GUID and FAT serial refine identityOf without changing shape. This increment is discovery + probe + log only, additive: the FAT service still acquires its own volume, so nothing changes for it. Confining each filesystem to its partition and spawning one per volume (the flip) lands next, keeping fat working throughout. Grants + wiring: init.csv spawns it after the device manager; protocol.csv grants bind volume-manager + open device-manager. Verified: volume-probe asserts the parse (bare-FAT volume at lba 0), neutral 10/10 across storage, restart, display, logging, confinement — the volume manager now runs in every boot and disturbs nothing. |
||
|
|
bc67771bfd |
block: reclaim range slots on death, and give confinement one controller
Two defects the V2 boundary review confirmed: 1. The per-badge range table was never reclaimed. usb-storage receives exit notifications (via the subscriber watch), but onNotification handled only the timer and dropped child-exits, so a dead filesystem left its range slot used forever. The medium-removal lifecycle churns filesystems, so after maximum_ranges confine/die cycles define_range would return ENOSPC and no volume could be confined again until reboot. onNotification now frees the dead badge's slot, mirroring fat's open-node sweep. 2. define_range checked only that the CALLER was unconfined, never that it owned the target badge — so any unconfined opener could install a range for another live client and silently redirect its I/O. Confinement now has a single controller: the first unconfined party to define a range (the volume manager, which confines every filesystem before handing it a channel). Only the controller may thereafter; the slot releases on its death so a restarted manager re-takes it. This is the mechanism half; V3 adds the grant half (only the volume manager gets an unconfined channel). Neutral: block-range (the fixture is the sole definer -> controller), fat-mount, usb-report all green. End-to-end exercise of both lands in V3/V4 (the VM+filesystem relationship and the remount churn). |
||
|
|
89d4592777 |
block: close the range-clamp overflow — a confined caller could wrap into the neighbour
The naive bound `lba + count > r.count` wraps for an lba near u64 max: the sum overflows to a small value, sails under the check, and `base + lba` wraps to an absolute block OUTSIDE the range. Calibrated, it is a real confinement escape — a process confined to [1,3) reads absolute block 0 (the boot sector) with lba = maxInt(u64), since base + lba wraps to 0. The bound is rewritten as two subtractions that cannot overflow: lba within the range, and count within what remains. block-range gains a wrap-refused assertion calibrated to be exploitable against the naive form — it FAILS against the old bound (reads block 0) and passes against the fix (verified by reverting the clamp). Caught pre-emptively before the V2 boundary review. |
||
|
|
7af65697cc |
block: medium presence — the medium_changed event and usb-storage as publisher (V2b)
The block protocol gains a pushed medium_changed event (present + a monotonic change counter; presence only, never content). usb-storage becomes a Subscribers provider and runs a slow TEST UNIT READY poll (1 s): success is present, failure absent, and a transition bumps the counter and publishes. This is the second removal trigger — the DEVICE stays while the MEDIUM leaves (card readers, ATAPI trays) — which channel death cannot see (storage-architecture.md, two triggers one lifecycle). The subscriber is the volume manager (V3); until it exists the publish is a no-op fan-out, so this commit is behaviour-neutral, and its end-to-end test (eject -> medium_changed -> unmount/remount) lands in V4 with the real consumer rather than a throwaway subscriber fixture (recorded sequencing). Sense-key inspection to tell medium-absent from other transport errors is a noted refinement; a clean eject reads correctly as not-ready. Neutral: 12/12 across the block-serving surface, restart, confinement, conformance, and logging. |
||
|
|
c37402891a |
test: block-range — the discrimination fixture for range confinement (V2a)
A process acquires a block channel the way a filesystem does (consumer-hello the device manager), confines ITSELF to blocks [1,3), then proves the clamp and the gate: volume-relative LBA 0 maps inside the range and reads; a read reaching past the range is refused; geometry reports the confined size; and a confined caller can no longer call define_range (no widening, no escape). It gates on argv so the ramdisk sweep leaves it silent in other boots, and coexists with fat (ranges are per-badge). Discrimination (verified by reverting usb-storage to pre-clamp f1bdce2~1): the unconfined read still succeeds but define_range returns ENOSYS, so the fixture cannot arm confinement and the case fails — exactly the property the clamp adds. With the clamp: block-range 1/1. |
||
|
|
f1bdce25e0 |
block: per-sender range confinement — V2a mechanism
The block protocol gains define_range (appended, numbers hold): confine the process named by `badge` to blocks [base, base+count). usb-storage keeps a per-badge range table and, in read/write, translates volume-relative LBAs (base added) and refuses any transfer past the volume end. geometry returns the confined size, so a filesystem mounts against what it may actually touch. The security seam (decision 4, settled): the clamp lives at the PROVIDER, so a channel carries exactly the authority it grants — handing a filesystem the whole disk plus a base offset would let it reach the neighbouring partition. The gate: a confined caller may NOT call define_range, so a filesystem cannot widen its own range or confine anyone; only an unconfined party (the volume manager, whole-device) may. The volume manager defines a filesystem's range before handing it the channel, so the ordering holds by construction. Default (no range for a badge) is the whole device — behaviour-neutral for a single-volume boot and what the volume manager itself uses to probe partitions. The range table is declared through bounds.md as a runaway detector (ours, refuse at limit), not a real-partition cap. Neutral: fat-mount, usb-storage, iommu-usb-storage green. The discrimination fixture (a confined process reads past its range and is refused) follows next. |
||
|
|
d63a008148 |
file-system: extract the serving harness from fat — V1
fat was one binary doing four jobs; the three that are not FAT-specific move to library/kernel/file-system-harness, a Server(comptime Engine) generic over the engine type: the badge-scoped open-node table, the nine vfs handlers, the not-mounted politeness, the exit sweep, mount registration, and durable-on- close. A filesystem is now an engine plus a main that hands the harness a mounted volume; a second engine reuses the harness wholesale. Placement note: the plan said library/file-system, but the harness is a specialization of `service` (its sibling) and needs nothing from the device domain, so it lives beside service in library/kernel and stays block-free — durability rides a caller closure (Volume.flush), no backwards kernel->device dependency, no new-domain scaffolding. The engine type is inferred from resolve()'s return, so engine.zig is untouched (its Node stays module-scope). fat keeps only its FAT-specific bring-up (acquireVolume, DMA, engine.mount, the attach/detach round trip) and the three mount prefixes as data. Behavior- neutral: 13/13 across the fat/vfs/logger/IOMMU surface, nothing observable changed. This lands first so every later phase touches the harness once. |
||
|
|
60b41c0e82 |
kernel: mounts have owners — V0 of the volume-manager plan
fs_unmount was gated by nothing but the /protocol carve-out: any process could unmount any prefix — latent with one mount owner, an obvious cross-tenant hole once volumes multiply. Each backend mount now records the mounting task, and the syscall layer enforces two rules that keep the restart story intact: only the owner unmounts (a dead owner's mount is swept lazily by resolution — strangers gain nothing by racing that), and a mount may be REPLACED only by its live owner or after its owner died (the respawned-filesystem path; displacement of a live mount would be worse than unmounting it). Kernel-installed mounts are never displaceable. The vfs-test park role is the discrimination: with the volume provably mounted it attempts the foreign unmount, requires the refusal AND the subtree still resolving, and withholds its "parked" marker otherwise — against the ungated kernel the unmount was ALLOWED and vfs-client-death fails; with the gate, green. (Its verification handle closes immediately: the kernel test string-matches "released 1 handle(s)".) |
||
|
|
a44b397bed |
protocols: attach gets its reverse — block detach, usb-transfer dma_detach
The kernel was always symmetric (dma_bind 51 / dma_unbind 52); the two protocols that forward an attachment up the stack were one-way, so a live client could grant a device reach into its buffer but never revoke it while alive — exactly the one-way lifecycle the storage architecture's enforcement section forbids. Death stays the mechanical backstop; detach is the living process's path. Both verbs are appended, so every existing number holds. The shape mirrors attach precisely: the same region capability rides the cap slot again — the kernel matches the region, so no layer retains anything between the calls (the bus never kept the handle; now it never needs to). fat's bring-up does attach -> detach -> attach, exercising both verbs through the whole chain (fat -> storage -> bus -> kernel) on every boot: a broken detach fails every fat case instead of lying dormant until the first buffer replacement. Honest scope: the round trip proves the plumbing; unbind semantics are the kernel iommu tests' (map/unmap/translationOf); the full composition (detach then DMA faults) is a future iommu-fault extension. |
||
|
|
77e7001878 |
usb: a hub yanked from a root port takes its subtree with it (H2)
The hot-plug matrix's predicted bug, found on first contact: tearDownPort never recursed into a departing hub's children — only tearDownHubDevice (a hub leaving one level down) did. Yank a populated hub from a root port and the downstream slots stayed live against vanished hardware, their class drivers were never reaped, and the replugged hub found its port still occupied, so nothing ever re-enumerated: the subtree was gone for the boot. tearDownPort now recurses children-first, exactly like tearDownHubDevice. The usb-hub-yank case is the discrimination: one device_del removes a hub carrying a keyboard AND a mouse, both drivers must be reaped, and the re-added hub must rebind both — it failed against the unfixed bus and passes with the recursion. |
||
|
|
4a2df587bb |
establishment: unplug reaps like death, so a replug rebinds
The hot-unplug path (onChildRemoved, a report from a live bus) cleared the child but left the bound class driver: a process blocked on reports that will never come, whose stale entry made the matcher's dedupe refuse the respawn when the device was plugged back in — the same wall the restart zombie hit, one path over. Unplug now reaps exactly like reporter death. The usb-hub-unplug case grows the replug: device_del the hub keyboard, then device_add it back (qmp_sequence); the ordered tail — child removed, reaping, delegated, ok — can only be satisfied by the second generation, since every boot keyboard's ok precedes the unplug. Discrimination: without the reap the replug never rebinds and the case times out (verified by stash run). Hub family, restart drill, and the two-controller proof all green (8/8). |
||
|
|
8710944a92 |
establishment: a reporter's death reaps its subtree, and the re-report rebuilds it
P3 of docs/establishment-planes-plan.md — the restart-zombie fix. A class driver cannot observe its provider's death: an HID driver blocks on interrupt reports that will simply never come, and storage answers its callers with refusals forever. Worse, the dead generation's still-used entries made the matcher's dedupe refuse the respawn when the restarted bus re-reported — the subtree was a permanent zombie, which is the exact opposite of the restart-a-driver-live goal the driver model exists for. pruneChildrenOf now reaps: each pruned child's bound driver is killed and its entry cleared (the exit notification finds no entry, so the death is never double-counted; its device returns by the loan rule; its stored endpoint handle is closed). The re-report then spawns a fresh generation whose hellos fetch the successor's channel. The usb-report drill now asserts the subtree WORKS after the restart: the respawned storage opens its device on the NEW bus instance and reads block 0. Discrimination: against the pre-reap manager the drill fails — no reap line, no post-restart respawn (the survivors were zombies), verified by a stash run. Note the scenario boots no input service, so the HID drivers of BOTH generations exit after their input lookup times out — storage is the functional proof. |
||
|
|
1d7850239d |
establishment: block stops being a name, and enumerate learns to page
P2 of docs/establishment-planes-plan.md. usb-storage serves nameless — one process per stick cannot share an exclusive bind, and a second stick used to die silently on -EBUSY before ever helloing. Its one hello now moves both directions at once: the block-serving endpoint up, its controller channel down. fat finds its volume through the manager — a new `.consumer` role asks for the channel of the driver BOUND TO a device (distinct from the device's reporter), found by enumerating the tree for the mass-storage identity. fat stays single-volume; the boot-volume-by-content choice is M21. The conversion immediately caught a live truncation of exactly the audit's shape: ChildEntry grew to 32 bytes, one enumerate reply holds ~7, and a real tree carries a dozen ACPI nodes before the first USB child — the storage entry silently never fit (the protocol comment already said "paging joins the protocol if a tree ever outgrows one packet"). enumerate is now paged: Header.target is the start cursor, a short page is the end; device-list's page-0 read is unchanged. Grant rows move with the code: the block bind and fat's block open die, fat gains open device-manager. Gate: 18 cases green including the registry trio, device-list, and both IOMMU storage variants. |
||
|
|
d603d40b5c |
establishment: usb-transfer stops being a name
P1 of docs/establishment-planes-plan.md. The bus no longer binds
/protocol/usb-transfer — the bind race made whichever instance came second
unreachable, which on a real three-controller Ryzen meant a mouse no class
driver could reach ("could not open device 50"). Each instance hands its
serving endpoint up in the hello that already delegates its controller, and
class drivers receive their OWN controller's channel from helloForChannel —
routed by the manager's lineage, retried while a provider is mid-restart,
on one manager handle so retries never spend handle-table slots.
usb.open(bus, id) now takes the channel it used to look up; the name rows
leave protocol.csv with the code (bind 67, opens 119/120/122); the
conformance fixture's prose stops claiming the bus binds; and the stale
input-client import leaves the bus with the channel one.
Gate: 19 QEMU cases green (usb family, hubs, both IOMMU variants, fat chain,
boot-from-USB, orderly shutdown, conformance).
|
||
|
|
77fe4d220e |
establishment: the mechanics — reply capabilities, helloExchange, lineage routing
P0 of docs/establishment-planes-plan.md; no behavior changes yet, nothing sets the new flag or sends a hello capability. - service.run gains a reply-capability out-slot (replyWithCapability), the registry idiom init already uses, lifted into the harness; null stays the untouched common path. Subscribers gains claimArrival() so a provider handler can keep a turn capability through the same flag the reserved subscribe uses. - Hello wire struct: the padding byte becomes wants_channel — old callers wire-compatibly say 0, and the manager nominates a reply capability ONLY when asked, because a capability sent to a caller that never reads one is a leaked slot in that caller's table. - driver.helloExchange: one handshake can hand a serving endpoint up and receive the device's provider channel down (usb-storage will need both at once). No channel in the reply is retryable, never a verdict. - The manager stores each instance's serving endpoint on its Driver entry, routes consumer hellos by lineage (child -> reporter -> endpoint), replaces on re-hello, and closes the stale handle on death - the 32-slot table is the bound that makes forgetting this boot-fatal. |
||
|
|
5cca580066 |
kernel: AMD-Vi asked for an interrupt where it meant a store
Two command encodings checked against the specification: - COMPLETION_WAIT set bit 1 (I, interrupt) instead of bit 0 (S, store), so the IOMMU was never asked to write the sentinel, the poll always exhausted its spins, and completeAndWait returned without any guarantee the preceding invalidation had executed — no invalidation barrier has ever existed, on QEMU or on silicon. The in-code claim that "QEMU's amd-iommu does not implement the store form" was a misdiagnosis of this bug: with S set, QEMU stores the sentinel fine, and the amd-iommu cases now run without the warn line. On real hardware, which fetches commands asynchronously, the missing barrier was an IOTLB use-after-free window: unmap returned before the invalidation was confirmed and the caller freed the frames. - INVALIDATE_IOMMU_PAGES "invalidate everything" used address bits 51:12 all-ones; the architected encoding is bits 62:12 all-ones (the spec's literal 0x7FFF_FFFF_FFFF_F000). Real silicon is free to misread the non-architected form as a bounded range. |
||
|
|
35f43057f4 |
kernel: the give paths take the lock, and a give confines afresh after a death
Three holes from the real-AMD audit, one shared root: the delegation flag-day added paths that touch the broker table and the IOMMU records without the big kernel lock, and a loan-return rule whose re-delegation skipped confinement. - device_enumerate walked the table with no lock. The table stopped being a static array in the bounds track — reserve() regrows it through realloc on every boot — so an unlocked reader can be mid-copy out of a slice that device_register on another core has already freed and reused, or pair a fresh count with a stale slice. Each chunk is now snapshotted under the lock; the copy to the user stays outside it. - system_spawn's give ran entirely unlocked — the comment claiming "the lock has not been dropped" was false (spawnProcessSupervised takes and releases it internally). The ownership pre-check now only spares creating a doomed child; the give itself re-checks, confines and moves in one lock hold, and a give that fails after the spawn kills the child rather than leaving it running without the hardware it was spawned for. - A re-delegated device after a driver death was never re-confined. Death tears the domain down before the loan returns to the lender, so the next give found no active record, reassign no-op'd, and the respawned driver ran the device with a V=0 device-table entry and no domain — silently unconfined, the exact fail-open the fail-closed claim was built to remove. Both give paths now share one body (giveDeviceLocked): check first, confine afresh when no record is active — refusing with ECONFINE like the claim — and move last, when nothing can fail. The iommu test drives the death-and-respawn sequence directly; with the old reassign-only behaviour its two confinement checks fail, with this change the suite is 118/118. |
||
|
|
0eb2420690 |
device-manager: delete the delegated-set scaffolding
The name list and its predicate existed so drivers could move to delegation one at a time with the suite green throughout. Every driver is delegated now, so the manager simply hands over whatever device a driver was assigned. Deleting it caught a real consequence: crash-test finally got delegated too, and it was still claiming its device — so it got AlreadyClaimed because it already held it, exited, and the restart drill had nothing to restart. Its own comment named what the case was really checking: "the respawn only reaches this line because the kernel released the previous instance's claim at death". That property still holds, by a different mechanism — the device reverts to the manager on death and is handed to the replacement, which is the same guarantee without the race it used to rely on. All four delegation paths verified: the xHCI controller, the PCI bridge, the PS/2 two-node singleton, and virtio-gpu's restart re-attach. Run 3 complete. Suite 118/118. |
||
|
|
df9c1ed827 |
device-manager: hold the seeded hardware so none is left lying around
A device nobody holds can be claimed by anyone, so the manager now takes every firmware-discovered device that carries mappable resources, whether or not a driver wants it. The real gap was the HPET: an MMIO window, an IRQ, no user-space driver, and there for the taking. Held by the manager it is inert; unheld it was a way into physical memory. Two deliberate exclusions. The loader's framebuffer, which the compositor claims and which the manager must not take because it starts first. And anything with no resources, which grants nothing worth holding. Scope is the boot snapshot. A device reported later and matched to no driver stays claimable — pci-cap-test and iommu-fault-test both reach an unmatched NIC that way, so narrowing it is a separate change with those fixtures in scope. Recorded in the plan rather than left implied. The attacker fixture gains the assertion deferred since D2: after the system settles, nothing with resources may be taken. That assertion defeated itself twice before it worked, and both failures are worth remembering. First it swept at 0.029 while the manager did not bind its protocol until 0.047, so it reported a hole that closed a millisecond later. The retry loop that "fixed" that was worse: the first pass TAKES the device, so the second finds it unavailable because this process now holds it, and concludes all is well — it passed with the manager's claiming removed entirely. It now settles once and sweeps once, and fails when the claiming is removed. Suite 118/118. |
||
|
|
ba195fa0a2 |
kernel: claim refuses delegated hardware — and E2 had already closed the hole
The rule as planned: a device that was given to someone may be handed on, never taken. Implemented, and honest about what it is worth. Writing the test showed the plan had the wrong step doing the work. A delegated device is HELD, so an attempt to take it is refused as AlreadyClaimed before the giver is ever consulted; and once a borrower's death returns the device to its lender — or clears both when the lender is gone — there is no state where a device is unheld and still on loan. The window a stranger could have used stops existing at E2. This check is unreachable. It stays anyway: one comparison, failing closed, guarding any future path that frees a device without clearing its giver, which is exactly the hole this run closed. The comment says it is unreachable rather than implying a protection it does not provide. The attacker fixture does not gain the assertion that was deferred to this step, and its header records why: there is no refusal for it to observe, and on a bare boot with no device manager nothing is delegated at all, so the assertion had nothing to bite on. It failed loudly on its first run rather than passing quietly, which is the only reason this was noticed. It also leaves the loader's framebuffer alone without naming it: nobody delegates the framebuffer, so it has no giver, so the display service claims it exactly as before. Suite 118/118. |
||
|
|
1a1d92cba9 |
kernel: a grant is a loan — a dead borrower returns the device to its lender
When a driver dies, a device it was *given* now goes back to whoever lent it, rather than to nobody. The device manager gets its hardware back the instant a driver dies and hands it to the replacement, with no window in between. That window was real: the kernel released the claim to no one and the manager re-claimed first-come, so every driver restart reopened the hole this run is closing. It also becomes load-bearing at the next step — once claim refuses a device that has a giver, releasing to nobody would strand a dead driver's hardware permanently, because nobody could ever take it again. A dead lender is no lender: the claim and the giver clear together, so a device is never owed to a ghost. A device nobody lent is released outright, exactly as before. The broker cannot see the task table, so liveness arrives through the same hook idiom the scheduler already uses. Null means assume dead, so a kernel built without the hook frees claims rather than handing them to a ghost. A stale binary nearly passed as proof for the third time this session: the first discrimination patch left `alive` unused, the build failed with three errors, and the old binary reported every assertion passing. Checking the build before reading results is what caught it. Suite 118/118. |
||
|
|
4ca57fc37e |
kernel: record who gave each device away
One field, and the rest of the run follows from it. A device that was given to someone is delegated hardware: it may be handed on, never taken, and when its holder dies it goes back to whoever lent it instead of becoming free for anyone to grab. It also settles the framebuffer without mentioning it. Nobody delegates the loader's framebuffer, so it has no giver, so the display service claims it exactly as it always has — no exemption and no reference to display anywhere in the rule. No behaviour changes here; the field is recorded and read by nothing yet. The test found a real bug on its first run, before the discrimination check. The sentinel for "nobody gave this" was 0 — and task 0 is a real task, the kernel's own, so a device given away by task 0 read back as belonging to nobody. Both giver and registrar are optionals now. The second was a latent bug from D9: the per-registrar allowance would have miscounted every device task 0 registered. Suite 118/118. |
||
|
|
b7d97ebb5d |
acpi: discovery is handed its node like every other driver
The last claimant. The kernel seeds the acpi-tables node, so it sits in the same boot snapshot the manager already scans to find the PCI host bridge — there was never a bootstrap problem, only a lookup nobody had written. The manager claims it and names it in the spawn; the service stops claiming. Every driver in the system now receives its hardware rather than taking it. Two failures on the way, both mine. addDriver puts the device id in argv[1], and the acpi service read argv[1] as a self-verify device-count floor — so handed device 7 it decided it was in test mode, printed "acpi-parse: ok", and never reported a device. The test argument is now floor:N, which a bare id cannot be mistaken for. And acpi-parse spawns the service directly rather than through the manager, so nothing handed it the node. That test now claims and transfers it exactly as the manager does, which is the right shape: the test plays the manager's role instead of the service reaching for hardware. device_claim now has two callers left: the manager, which is the acquirer and should have it, and the display service's GOP path. That is recorded as question 10 — the framebuffer is not a device, so the answer is likely that it leaves the device table rather than being exempted from its rules. Suite 118/118. |
||
|
|
6b3a381626 |
ps2: one instance, every node the machine has
The 8042 is a single controller described by two ACPI nodes — PNP0303 carries the 0x60/0x64 ports, PNP0F13 is the mouse — so it cannot be split across two processes without them fighting over the same registers. That is why ps2-bus is a singleton, and why it used to find and claim both nodes itself. The manager now gives it every matching node instead. The keyboard node rides the spawn, because it holds the ports and is needed immediately; the mouse node is transferred to the already-running instance. Late arrival is safe here and the ordering is natural rather than lucky: the mouse is not touched until after the controller handshakes and identify. Measured, the handover lands at 0.336 and the driver reaches the mouse at 0.456. Because the count is however many matched, a machine with no PS/2 ports or only one works without a special case — which matters, since the bus is mostly emulated now and machines vary. ps2-bus claims nothing. irq_bind on the mouse node is the proof it holds it: that call is ownership-gated, so a failure means the handover did not land rather than a hardware fault, and the log says so. acpi-ps2 asserts both delegations with the spawned device backreferenced, so the node that rides the spawn must be the one the driver was spawned for. Disabling the second delegation fails it. Two things worth recording. The first discrimination patch was not valid Zig, so nothing ran and a stale binary reported a pass — checked the build before believing it. And with the second delegation disabled, acpi-ps2 fails while input still passes: the mouse works without its IRQ binding, so exactly one case covers that path. Suite 118/118. |
||
|
|
7d8aa51234 |
kernel: maximum_children_per_parent is gone
The second invented ceiling. It was written to stop a driver looping device_register and exhausting a shared table — but there is no shared table to exhaust any more, and each registrar already has its own allowance, so a runaway costs only itself. It never bounded a determined caller in the first place: 16 children per parent, and nothing stopped it claiming more parents. What it reliably did was refuse a real PCI bus with more than 16 functions, which is how an AMD Ryzen booted with a working display, no USB and no storage. The constant, its check, and the now-unused childCount all go. TooManyChildren survives with one meaning instead of two: the caller is at its per-registrar allowance. This was unblocked from the moment D9 landed. The plan said so — "once the quota exists the per-parent cap is redundant whether or not D6 has landed" — in the same edit that left the step tagged "blocked on D6". Three iterations were then spent re-reading that tag instead of the sentence beside it. The containment test now asserts 64 children under one parent, four times the old ceiling; restoring the cap fails it. Suite 118/118. |
||
|
|
3ae541214f |
iommu: assert directly that a confinement moves with its device
reassign was added at D4 to fix a regression and has been proven only indirectly since — three IOMMU+USB cases going green. That covered the visible symptom (a driver's DMA rings unbound) and neither of the latent ones: the confinement still naming the giver, so the giver's death would tear down a domain a live driver was using, and the receiver's death would leave one behind. Those are now asserted. confinementOwner exposes the record's owner so the suite can see it. The sequence is the delegation in miniature: unconfined, confine as this task, reassign to another, confirm the new holder owns it and the old one does not, then kill the new holder and confirm the domain goes with it. Two attempts at this test could not have failed. The first found no PCI function to confine — pciAddressOf needs a pci_device entry and this case runs no pci-bus — so every assertion skipped silently while the case stayed green. It now synthesizes a function the way pci-bus does, a 4 KiB config window inside the bridge's ECAM, and asserts that precondition explicitly so a skip is a failure. Verified to discriminate: making reassign a no-op flips three assertions, including the domain surviving its holder's death. Suite 118/118. |
||
|
|
2ebfccc8ed |
virtio-gpu: the scanout device arrives with the spawn
Third driver converted. It no longer claims the id from argv[1] — the manager holds the device and names it in the call that creates the process, so it is held before the driver's first instruction. display-reattach is the case that matters here: it kills the driver and watches the compositor re-attach to the fresh scanout. It passes, so the restart path survives the fused grant — the manager re-takes the device when the driver dies and hands it to the replacement. ps2-bus and discovery are NOT converted, and the reason is recorded as open question 9 rather than worked around. Both need a device nobody assigned them. ps2-bus ignores its argv[1] entirely: it finds the controller by walking the table for PNP0303, then claims a second device, the PNP0F13 mouse node, which it also finds itself — so it holds two devices and was assigned at most one, while system_spawn carries one. discovery claims the acpi-tables node it locates itself, because it is what produces the device tree and there is nothing to assign at that point. One thing worth checking before designing an answer: devices.csv maps both PS/2 hardware ids to ps2-bus, so the manager may already be spawning two instances where the driver expects one. If so the fix is smaller than it looks. D6 stays blocked — closing device_claim with these two still depending on it would stop the machine booting. Suite 118/118. |
||
|
|
f23f073624 |
kernel: the device rides system_spawn, so a driver never runs without it
Delegation moves out of onHello and into the spawn itself. The manager holds the hardware and names it in the call that creates the driver; the kernel checks the device is the caller's to give, then hands it over as part of making the child. The reason is the window. A transfer after spawning always leaves an interval in which the child is running and does not yet hold its device. It would close on QEMU every time and open occasionally on a machine with different core counts and timing — the exact failure shape this track exists to delete, and not one worth introducing while removing the others. Fused into the spawn there is no interval: the child does not exist until it holds the device. Ownership is checked BEFORE the child is created, so a refusal leaves nothing running rather than a driver without the hardware it was spawned for. The IOMMU confinement moves with the device, as it does on the transfer path. systemCall6 is added for the sixth argument; r9 was free, and abi gains a no_device sentinel matching the protocol's. No driver had to change to receive a device, which is what makes this better than requiring every driver to hello: ps2-bus keeps its legacy status, and discovery — which has no assignment at all, since it is what produces the device tree — is unaffected. The attacker fixture now tries the spawn as a back door: name someone else's device, and both the spawn and any child must be refused. Verifying that assertion exposed a bug in the fixture itself. The kernel case's pass marker was "device-authority: ok", which matches the FIRST per-assertion line, so its wait loop exited before any failure was printed — the case would have passed with failures in it, and had been able to since D2. The verdict lines now carry a distinct VERDICT prefix, and with the ownership check removed the case genuinely fails. A green test that cannot go red is worse than no test. Suite 118/118. |
||
|
|
f7151ed577 |
kernel: the device table has no ceiling; a runaway is charged to whoever caused it
maximum_devices = 64 is gone. It was a guess about someone else's computer, and because it was shared, one driver's enumeration starved every other — which is how an AMD Ryzen booted with a working display, no USB and no storage. The table now grows from the kernel heap. It was always built after heap.init; nothing ever prevented this except it having been written static first. What replaces it is an allowance charged to the registrar, so a driver looping device_register exhausts its own and every other driver carries on. It is declared as what it is — a runaway detector, NOT a security boundary. A quota generous enough never to bite a real machine is still generous enough to be unpleasant, and it is not trying to be the defence; delegation is. What this catches is a legitimate driver in a loop, early, attributably, and without collateral. Reaching 4096 is a bug report, not a tuning request. The initial block is 8, deliberately small. Sizing it for a typical machine would mean the growth path never ran on the hardware we test on and only woke up on someone else's larger machine — the exact failure shape this track exists to stop. At 8 it grows several times every boot; disabling growth now fails the suite with the HPET not fitting, which is the Ryzen failure in miniature. The comptime coupling assert added earlier fired, and was right to. confined (one slot per device id) and domains (the IOMMU's own translation pool) were sized by the same constant only because device ids happened to stop at 64 too. Two unrelated quantities: confined now grows with the device table, while maximum_domains stays as the hardware's number — both VT-d and AMD-Vi report how many domains they support, and reading it is phase 4. The assert existed for exactly this and did its job. Suite 118/118. |
||
|
|
81dd1e9318 |
pci: the host bridge arrives by delegation too
pci-bus joins usb-xhci-bus in receiving its device from the manager rather
than claiming the id it found in argv[1]. Its hello moves ahead of the ECAM
mapping, since that is where the bridge now arrives, and its hello was
already mandatory so nothing about its failure behaviour changes.
isDelegated compared whole strings, which silently missed this driver: the
manager records the boot-snapshot match as the bare "pci-bus" and a
devices.csv match as the full "/system/drivers/pci-bus". pci-bus was then
neither claiming nor delegated and died on "ECAM mmio_map failed". It now
matches on the last path component. Reintroducing the whole-string compare
breaks usb-xhci-bus instead of pci-bus — the two spellings swap which driver
loses — so usb-hid is the case that catches it, not pci-scan.
pci-scan asserts the delegation on the initial bring-up AND after the
restart drill, with the device id backreferenced so both must name the same
device. That is what proves the manager re-takes a device when its driver
dies and hands it to the replacement, which is the property the whole
supervision design rests on.
The remaining three claimants are NOT converted, and the plan records why
rather than working around it. ps2-bus and the acpi service never hello at
all, which device-manager.md states deliberately ("legacy drivers ... not
yet required to hello"), so delegating to them means either promoting them
out of legacy or giving the grant a delivery point that is not hello.
virtio-gpu hellos best-effort by design — "standalone bring-up has no
manager" — and delegation would make it mandatory. Both are decisions, not
mechanical steps.
Consequence: D6 is blocked, because device_claim cannot be closed off while
three claimants still depend on it. D7-D9 are unaffected — they concern what
the kernel stores and how its table is sized.
Suite 118/118.
|
||
|
|
a2d6056772 |
usb: the xHCI controller arrives by delegation, not by claiming
The first driver to stop claiming its own hardware. The device manager holds the controller and transfers it in the hello reply, so its matching becomes authoritative instead of advisory — until now the driver claimed the id it found in argv[1], and any process could have claimed the same integer first. The manager claims before it spawns, so there is no window in which anything else could take the device, and transfers in onHello using invocation.sender — the kernel-stamped task id, which cannot be forged by the caller. hello is synchronous, so the transfer has completed before the reply lands: no gap between being told yes and holding the thing. usb-xhci-bus's hello moves from after controller bring-up to before anything that needs the device, which is the bring-up reorder the design predicted. It is the first member of an explicit delegated set, so every unconverted driver keeps claiming exactly as before and the suite stays green; the set and device_claim both go at D6. D3 and D4 could not be separated and the plan records why: the moment the manager claims, any driver still calling device_claim is refused, and D3 applied to nothing changes no behaviour and cannot be tested. This step introduced a regression and the incremental conversion is what caught it. confineDevice runs inside systemDeviceClaim, so a device arriving by transfer was never confined for its new owner. Three IOMMU+USB cases failed on the driver's DMA rings going unbound, and two worse consequences were latent: a manager death would have torn down a domain a live driver was using, and a driver death would have leaked one. iommu.reassign now moves the confinement with the device, keeping the domain and its attachment intact so it never translates through nothing. Converting all five drivers at once would have produced the same three failures with five suspects. A log line of mine claimed "holding controller device N" before anything verified it — it printed even in the failure case, where the driver held nothing. Reworded to state only what is known there: where the registers are. usb-hid asserts the delegation with the device id backreferenced, so the id delegated and the id the driver ends up with must match. Emptying the delegated set fails it with "hello acknowledged" then "mmio_map failed". usb-hub failed once in a full run and has passed six times since (four isolated, two full) — recorded in the plan as a suspected instance of the known intermittent AP fault, not dismissed, since this step did shift boot timing. Suite 118/118. |
||
|
|
33376f24ab |
test: the attacker the device suite never had
The audit's sharpest finding was structural, not a bug: a fully green suite had hidden six real defects because it contains no attacker. Every device case asserts that a driver handed its own hardware can drive it. None asked what a process handed NOTHING can do. device-authority-test is that process. It is spawned with no device and asserts what it therefore cannot do: it cannot give away a device another task holds, nor a free one, because the kernel's rule is that you may give away what you hold and the device's state is irrelevant to a process holding nothing. Asserted across every device the machine actually has, so it cannot pass by accident of which one happened to be free at boot — six on QEMU, none of them its. A positive control runs first. device_enumerate works from this process, so the refusals below it are decisions rather than a syscall path that is simply broken here; without it, "everything failed" would read identically to "the assertions are meaningless". A nonexistent device is refused as NoSuchDevice rather than NotHeld, because a refusal that cannot name its own rule is what cost a debugging session on the Ryzen. What it deliberately does not assert, and says so in its header: device_claim is still first-come-first-served at this point in the run. That is the hole D6 closes, and the claim half of the invariant joins this fixture then. Asserting it now would be writing a test that documents the bug. Verified to discriminate: removing the holder check flips "every transfer by a non-holder is refused" while the positive control keeps passing. Suite 117 -> 118. |
||
|
|
3111c7c5e6 |
kernel: device_transfer — you may give away what you hold
The mechanism behind delegation, which device-manager.md named as the step after hello: the device manager claims what discovery seeded and hands each device to the driver it matched, so assignment stops being first-come-first-served. It is a MOVE, not a copy. A claim is exclusive (driver-model.md, invariant 1), so the giver stops holding the device the instant the receiver starts. That is why this is a new syscall rather than the M13 capability path, where a passed handle is shared refcounted — exclusivity cannot be expressed that way. The kernel's whole rule is that you may give away what you hold. It has no notion of which task is the device manager and deliberately gains none: a binary name inside the kernel is not something that cannot safely live in user space. A recipient that does not exist is refused, because a device moved to nobody would be unreachable for the rest of the boot — nothing un-holds a device but task death. Three errnos, each naming its own rule: ENODEV no such device, EPERM you do not hold it, ESRCH no such recipient. Nothing uses it yet. The five claimants move across one at a time in D4-D5, so the suite stays green throughout and a regression names the driver that caused it. Ten assertions, verified to discriminate: removing the ownership check flips four of them, including the giveaway that an illegal transfer then blocks the legitimate claim behind it. Suite 116 -> 117. |
||
|
|
5d55217212 |
usb: a slot the controller granted is always handed back
Both device-setup paths issued a successful Enable Slot and then returned null if allocateDevice failed, without disabling it. A slot the driver forgets is one the controller never reissues, so each attempt lost one permanently for the boot. The hub path did it with no log line at all. Both now release the slot through a shared disableSlot, extracted from tearDownDevice, and the hub path warns like the root-port path does. tearDownDevice also now frees the interface list. That allocation arrived with the previous commit, so an unplug would have leaked it — found while reading the teardown path for this fix rather than by a test. No regression test, and it is recorded as open question 5 rather than implied. After the slot count became the controller's own figure, reaching this path needs more devices than the controller has slots: QEMU offers four against sixty-four. What was verified is that the new path RUNS correctly — pinning tracking to 2 with four devices attached produced "port 6 setup: no free device slot", the first two devices enumerated normally, and no Disable Slot error or timeout appeared, which is how disableSlot reports failure. Suite 116/116. |
||
|
|
729b40ece7 |
usb: a device has as many interfaces as it declares
max_interfaces was 4. A composite device — a headset, a webcam with audio, a dock, a multifunction printer — routinely has more, and the fifth did not merely go missing. parseConfiguration's cap branch had no `else`, so when the count was reached `current` kept pointing at interface 3 and the fifth interface's endpoint descriptors were appended to interface 3's array. A class driver bound to interface 3 could then be handed an endpoint belonging to something else entirely, and subscribe or bulk-transfer on it. The alternate-setting arm one line above cleared `current` correctly, which is what the cap branch should have done. Interfaces are now counted from the block in a first pass and allocated to exactly that number, so the ceiling is bNumInterfaces' u8 — the USB specification's. The missing `else` is added too, though after this the bug is unreachable by construction: interface_count cannot reach interfaces.len mid-parse when the list was sized from the same walk. max_configured_endpoints was max_interfaces * max_endpoints_per_interface = 16, a derived guess that moved whenever either input moved. It is now 31, which is the xHCI specification's own limit: a Device Context holds a slot context plus at most 31 endpoint contexts, because the Context Entries field addressing them is 5 bits. max_endpoints_per_interface stays at 4 with its reason recorded — the usb-transfer wire protocol reports exactly max_reported_endpoints (4) per interface, so widening it alone would change nothing a class driver sees. Lifting it is a protocol change. No direct test, and that is written down as open question 5 rather than glossed. The parser is pure and wants a host unit test, but usb-xhci-library.zig imports memory, mmio and time so it cannot be a standalone test root, and QEMU offers nothing that reaches the path — the largest device available is usb-audio,multi=on at 2 interfaces and 211 bytes. The alternate-setting path that shares the same `current = null` logic is exercised by that device. Suite 116/116. |
||
|
|
6328823ef1 |
usb: a configuration block is as long as the device says it is
The driver read the first 512 bytes of a configuration block into a fixed buffer and parsed those. The block's length is the device's own choice (wTotalLength, a u16), so anything larger was silently cut: interfaces past the cut did not exist as far as the host was concerned, while the SET_CONFIGURATION that follows still configured the device for all of them. A headset is 500-900 bytes, a UVC webcam 1-3 KB, a multifunction printer 600+. Now allocated at the declared length, so the ceiling is the field's u16 — the specification's number rather than one of ours. A block shorter than its own 9-byte header is refused rather than trusted. The bring-up line reports the declared length and the bytes actually read, so a truncation can never again be invisible, and usb-large-descriptor asserts they match with a backreference. That case has an honest limit, recorded in its comment: QEMU cannot produce a block over 512 bytes. The boot keyboard, mouse and stick are 34-44, and the largest device available is usb-audio in multi-channel mode at 211 — which is exactly why the suite never caught this, and why it cannot now reproduce the original trigger. What it does catch is the class: any clamp below the attached device's block fails it, verified by pinning the buffer to 128 and watching "config block 211 bytes, read 128" turn the case red. Suite 115 -> 116. |
||
|
|
69fbef40c0 |
usb: the controller says how many device slots it has
max_devices was 8, with the comment "QEMU presents a handful; a fuller machine would grow this" — a number chosen against the test rig, waiting for a real machine, which is the pattern docs/bounds-track-plan.md exists to stop. The driver already knew the true figure. It reads HCSPARAMS1.MaxSlots at bring-up and writes it straight into op_config, so every slot the controller offers has always been *enabled*; only the array tracking them was 8. QEMU's xHCI reports 64, so seven eighths of the controller was live and invisible, and the ninth device — a keyboard, mouse, webcam, headset, hub and two sticks reach that without trying — disappeared on a hub-attached path that logs nothing at all. The array becomes a slice allocated from max_slots at bring-up. A controller claiming zero slots cannot address anything, so that is now a dead controller rather than an empty allocation failing mysteriously later. The Device Context Base Address Array is a page, 511 usable entries, so it already covered the 255-slot maximum. The bring-up line reports both numbers, and usb-hid asserts they are equal with a backreference rather than a magic number, so the test cannot drift from the hardware. Pinning tracking back to 8 fails it: "64 slots, tracking 8". Suite 115/115. |
||
|
|
f4eb88e7d2 |
build: a new compile-time ceiling declares itself or does not land
The convention that tunables live in system/parameters.zig with their reasoning attached predates this and got 2% compliance — 5 of 235. A convention with no teeth is how a bare `const maximum_devices = 64` reached an AMD desktop and cost it USB and storage. This is the same rule with a gate behind it. tools/check-bounds.py finds every bound-shaped declaration — a `maximum_*` const with a literal value, or a type with a literal array length — and requires the five-field block above it: what it counts, who decides its size, what it protects, what happens at the limit, and how anyone finds out. The at-limit vocabulary is closed: refuse, degrade, truncate, grow. There is deliberately no way to spell "silent", no way to spell "drop", and nothing meaning "allow", so the behaviours that did the damage cannot be written down. Truncation is legal only carrying a marker the reader can see, which is why klog_maximum_message qualifies and a USB descriptor cut at 512 bytes does not. An array length that names a declared bound is not itself a bound; only literal lengths are flagged, which pushes ceilings toward having names. The 273 that predate the rule are allowlisted so this lands without a tree-wide sweep in front of it, and that list may only shrink: declaring a bound means deleting its line, and the check fails on a stale entry too. Nothing may be added. Wired into `zig build test` and available alone as `zig build bounds`. Not in the default build — it reads the whole tree, and a red bounds check should not stop you booting a kernel. Five are now declared rather than allowlisted. Writing them out is its own argument: maximum_devices reads "protects: nothing — this is a sizing guess about someone else's computer", and maximum_tasks now carries the fact that it has been raised twice, each time by something that outgrew it. Verified the gate refuses an undeclared bound, a declared one using forbidden vocabulary, and an allowlist entry that has since been declared. Suite 115/115. |
||
|
|
a86559648e |
kernel: a refusal names its rule, and two bounds stop failing open
An AMD Ryzen booted to a working compositor with no USB and no storage, and the log said only "register refused". A tree-wide audit of every compile-time ceiling followed: 235 of them, 139 on quantities the machine or a file decides rather than us, 5 documented anywhere, 171 silent when reached. docs/fixed-bounds-audit.md has the inventory. Errno attribution. The errno space was split between the kernel and the envelope, free to drift; it is now one list in system/abi.zig, restated on both sides, with a comptime check in library/device/driver where the two halves are visible. device_register's six refusals and device_claim's three are distinct codes, so a bus driver can say which rule stopped it, and BadParent splits into NoSuchParent and NotYourParent. pci-bus reconciles found against registered instead of counting refused functions as found. Idempotency ordering. The child cap was checked before the identity match, so a restarted bus was refused its own devices — the supervision restart the system leans on ratcheted toward a degraded machine. A re-registration consumes no slot and is now admitted first. IOMMU fail-closed. confineDevice returned success for a device id past the confinement table, leaving the device outside every domain while the caller believed it confined — unreachable only while ids stop at 64, which both the inventory move and a hardware-reported domain count would change. It refuses now, and the coupling to the broker's device cap is a comptime assert rather than a sentence in a comment. PCI apertures. The bridge's MMIO apertures are derived from the holes in the firmware memory map, and the derivation copied sub-4 GiB entries into a fixed [64] array and skipped the rest. A skipped region is not merely lost: the gap finder concludes it is free, so a real machine's 60-200 entry map yields an aperture over live RAM, and containment then admits a child BAR covering kernel memory. Rewritten to walk the map in place, with the hole finder extracted as a pure function and driven by a synthetic 100-entry map in a new test case. Both new tests were verified to fail on the old code. parameters.zig gains the rationale it was missing and loses a stale sentence pointing at the wrong file; vdso.md documents the errno space, including EPEER, which had no written meaning anywhere. docs/os-development/bounds.md is how a ceiling is declared from here. docs/bounds-track-plan.md is the plan to remove the ones we invented. Suite 114 -> 115. |
||
|
|
7db5fba884 |
kernel: boot on AMD — SYSRET puts the RPL in the STAR base
A Ryzen 3 3200G triple-faulted on its first timer tick after reaching init. Three defects in a chain, each hiding the one beneath it. STAR's SYSRET base was 0x10, so SS came back as base+8 = 0x18 with RPL 0 while CS carried RPL 3. Intel ORs RPL 3 into SS on SYSRET; AMD only does so for CS. Ring 3 ran fine — RPL is not checked on data access — and died the moment an interrupt tried to IRETQ back, where SS.RPL must equal CS.RPL. The base now carries the RPL (0x13), as Linux does. Two fixes below it, both of which made the first one unreadable: scheduler() read IA32_GS_BASE and dereferenced it without testing for zero, so every fault reporter faulted in turn — a panic inside a panic, and the machine reset before printing anything. Cast after the null test, plus a re-entrancy guard in the panic handler. NT is now masked in SFMASK alongside the rest, and isr.s exports isr_return_iretq at the faulting instruction so a frame dump can say which IRETQ died and print the CS/SS it was about to load. That dump is what identified the RPL mismatch. |
||
|
|
5a5146ec13 |
kernel: ring 0 reaches user memory only through the checked copy
SMAP makes the rule the copy layer has followed since it was written into a rule the hardware keeps. A ring-0 read or write of a user page now faults, so any code that reaches for a user pointer directly fails the first time it runs rather than the first time someone attacks it — and the suite becomes the enforcement test, because every case exercises the kernel with the bit on. Nothing had to be fixed to turn it on, which is the retrospective proof that the nine stragglers converted earlier were all of them. The interrupt entry needed one instruction first. Hardware does not clear the alignment-check flag on its way into a handler, and ring 3 sets that flag freely, so a process could have taken an interrupt with SMAP suspended for the duration. The system call path was already covered — its flag mask clears it — but the interrupt path needed a `clac`, which cannot simply be assembled in: it is an invalid instruction on a processor without SMAP, and danos boots on those too. So the entry ships as a three-byte NOP and is patched at boot, through the physmap, because the kernel maps its own text read-only. The ordering that makes that safe is enforced rather than described: the patch sets a flag, and no core will set the SMAP bit until it is true. A translation that fails, or bytes that read back wrong through the address they will actually be fetched from, leave the machine unhardened and saying so — which is the same posture the IOMMU takes, and better than enforcing over an entry path that cannot comply. The patch runs before interrupts are enabled and before any second core exists; a comment says so, because the three bytes pass through an encoding that must never be executed and a future change that moves this later has to deal with that first. Suite 114/114, with a case that reads a user page from ring 0 and requires the fault, and the multi-core case asserting every core that ran work had the bit — the same shape SMEP got, for the same reason: CR4 is per-core, and a hardening is only as wide as its narrowest core. |
||
|
|
cb30faf15f |
kernel: a hostile return address cannot fault the kernel
SYSRETQ with a non-canonical RIP raises a general protection fault in ring 0 — on the kernel stack, an instruction after the swapgs that installed the user's GS base. It is one of the better-known escalation primitives, and ring 3 reaches it without any kernel bug at all: the processor saves the address of the instruction after SYSCALL, so a program whose SYSCALL is the last two bytes of the last canonical page returns to the first non-canonical address. The new test does exactly that. The exit path now sign-extends the return address from bit 47 and compares; if the value changed, it returns through IRETQ instead, which commits the privilege change before fetching the new address, so the fault arrives from ring 3 and the process dies like any other. Four register-only operations and a branch that a correct program can never take — it could not have executed at a non-canonical address in the first place. Bit 47 is the right pivot because danos builds four-level page tables and nothing sets the five-level bit; a future port must move the pivot, and the comment says so. SFMASK grows one bit while we are here. SYSCALL, unlike an interrupt gate, does not clear the nested-task flag, so the kernel had been running every system call with whatever ring 3 last chose — harmless while the only exit was SYSRETQ, and a question worth not having now that one exit is IRETQ. The kernel is never nested; ring 3 still gets its own flag back. Suite 113/113. The new case asserts the refusal counter rather than the dying process: the emulator we test on kills it either way, so only the counter distinguishes a guard that ran from one that did not. |
||
|
|
36a7cc5fe9 |
kernel: ring 0 cannot execute a user page
SMEP turns the classic escalation — divert kernel control flow into a page the attacker wrote — from a silent takeover into an immediate fault with the offending address in the log. The bit is per-core state, so it is set where the syscall MSRs already are: in the per-CPU bring-up both the boot processor and every application processor run on their way in. A core that climbed the trampoline without it would be a hole no boot log would show, which is why the SMP case now reads CR4 on each core it lands on and requires every one of them to be hardened, not just the one that printed the banner. Enabling it that early is only safe because nothing ring 0 executes is mapped for ring 3, and that had to be established rather than assumed: kernel text carries only its ELF flags, the physmap is no-execute, the trampoline page is mapped supervisor and the core running it has not enabled the bit yet, and the boot processor turns it on while still on the loader's tables — which map nothing user-accessible at all. The one indirect call in the kernel takes a kernel address. The CPUID probing that was scattered across the timer code becomes a small shared helper, since the feature question is now asked from two places and each wanted the same maximum-leaf guard. Absence is tolerated and reported, like the IOMMU: danos still boots on a machine without the feature, and says which one it is. The test harness starts asking QEMU for a CPU that has the bit at all — its default model has neither SMEP nor SMAP, so the code would otherwise have been unreachable in every run. No case behaved differently under the richer model. Suite 112/112, with a new case that maps an executable user page, calls into it from the kernel, and requires the fault the CPU is supposed to raise. |
||
|
|
1b1c587c14 |
library: the harness keeps the subscribers, and an id belongs to whoever opened it
Three services had each written the same thing and got it three different ways: input polled the process list to notice a dead subscriber, and only when someone else subscribed; the power service never noticed at all; the device manager noticed drivers but not subscribers. The harness owns the table now, driven by the events a protocol declares — it registers on the reserved verb, frames each event once, posts to everyone interested without waiting on any of them, and reclaims a slot when the kernel says its owner died. Interest masks moved to the envelope, so a subscriber that wants only mice asks the same way everywhere. Two consequences the plan had not foreseen. The device manager now hears a supervised child's death twice, once as its supervisor and once as a subscriber, so restart backoff counted every crash twice and gave up after half as many; it retires the id before counting. And the kernel's published exit table had eight slots for what is now six subscriptions in a plain boot, so it holds sixteen. The other half is a hole the design named early and left standing: a backend handed out a small integer and then honoured it from anyone. A process that guessed a file's node id read another client's file; a display layer had no owner at all, so any client could reconfigure or destroy any layer; a USB device token was never checked against the client that opened it. Each is now bound to the task that opened it, and a wrong owner gets exactly what an unknown id gets — the refusal must not become the oracle the identical answers elsewhere were designed to remove. Closing a file changed with it: it used to succeed unconditionally, which would have told a caller which ids existed. Suite 111/111, with a new case in which one process holds a file and a layer, hands both ids to a second process, and finds them untouched after that process has tried everything with them. |
||
|
|
2719b93530 |
library: the last three protocols speak the envelope
These were the awkward ones. Each began with an operation packed into a single byte — two of them with a version wedged in beside it — so there was no wrapping them: the layouts had to be rebuilt. The device manager's own enumerate and subscribe become the reserved verbs that mean the same thing everywhere, its replies lose three status structs the envelope already carries, and a device id becomes the packet's target. Power drops the version it repeated on every request, because describe is the handshake, and stops claiming a 64-byte ceiling it never needed for calls. USB moves a control transfer's data to the packet tail in both directions, which makes the status length the transferred length and retires a field that had been saying the same thing twice. The danger in this one was not the protocols but their readers. Init recognised a power button by two bytes at the head of a message, the ACPI service dispatched on the first byte, the xHCI driver read its operation with a raw integer load, and the HID drivers reinterpreted a report wholesale — none of which would have failed to compile once the layouts moved. They would simply have stopped: no shutdown on the power button, no reports from the keyboard. Every one of them now reads through the generated types, and the shutdown gate that answers only a subscriber is the same code it was. Two sizes were decided by measuring rather than assuming. The child-added message is both a request and the event broadcast to subscribers, and alignment rounds it to 48 bytes, which puts its packet exactly on the 64-byte push floor — a test pins that, because a field added carelessly would now overflow it. The interrupt report gives up eight bytes of inline room to make space for the header; the two drivers that produce reports send eight and four. Suite 110/110. |