337b981c0764451ac2291486fd55a734148a0d2c
25
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
337b981c07 | docs: reorder the Run 2 table to the plan order and note the executed sequence | ||
|
|
b7d97ebb5d |
acpi: discovery is handed its node like every other driver
The last claimant. The kernel seeds the acpi-tables node, so it sits in the same boot snapshot the manager already scans to find the PCI host bridge — there was never a bootstrap problem, only a lookup nobody had written. The manager claims it and names it in the spawn; the service stops claiming. Every driver in the system now receives its hardware rather than taking it. Two failures on the way, both mine. addDriver puts the device id in argv[1], and the acpi service read argv[1] as a self-verify device-count floor — so handed device 7 it decided it was in test mode, printed "acpi-parse: ok", and never reported a device. The test argument is now floor:N, which a bare id cannot be mistaken for. And acpi-parse spawns the service directly rather than through the manager, so nothing handed it the node. That test now claims and transfers it exactly as the manager does, which is the right shape: the test plays the manager's role instead of the service reaching for hardware. device_claim now has two callers left: the manager, which is the acquirer and should have it, and the display service's GOP path. That is recorded as question 10 — the framebuffer is not a device, so the answer is likely that it leaves the device table rather than being exempted from its rules. Suite 118/118. |
||
|
|
7d8aa51234 |
kernel: maximum_children_per_parent is gone
The second invented ceiling. It was written to stop a driver looping device_register and exhausting a shared table — but there is no shared table to exhaust any more, and each registrar already has its own allowance, so a runaway costs only itself. It never bounded a determined caller in the first place: 16 children per parent, and nothing stopped it claiming more parents. What it reliably did was refuse a real PCI bus with more than 16 functions, which is how an AMD Ryzen booted with a working display, no USB and no storage. The constant, its check, and the now-unused childCount all go. TooManyChildren survives with one meaning instead of two: the caller is at its per-registrar allowance. This was unblocked from the moment D9 landed. The plan said so — "once the quota exists the per-parent cap is redundant whether or not D6 has landed" — in the same edit that left the step tagged "blocked on D6". Three iterations were then spent re-reading that tag instead of the sentence beside it. The containment test now asserts 64 children under one parent, four times the old ceiling; restoring the cap fails it. Suite 118/118. |
||
|
|
3ae541214f |
iommu: assert directly that a confinement moves with its device
reassign was added at D4 to fix a regression and has been proven only indirectly since — three IOMMU+USB cases going green. That covered the visible symptom (a driver's DMA rings unbound) and neither of the latent ones: the confinement still naming the giver, so the giver's death would tear down a domain a live driver was using, and the receiver's death would leave one behind. Those are now asserted. confinementOwner exposes the record's owner so the suite can see it. The sequence is the delegation in miniature: unconfined, confine as this task, reassign to another, confirm the new holder owns it and the old one does not, then kill the new holder and confirm the domain goes with it. Two attempts at this test could not have failed. The first found no PCI function to confine — pciAddressOf needs a pci_device entry and this case runs no pci-bus — so every assertion skipped silently while the case stayed green. It now synthesizes a function the way pci-bus does, a 4 KiB config window inside the bridge's ECAM, and asserts that precondition explicitly so a skip is a failure. Verified to discriminate: making reassign a no-op flips three assertions, including the domain surviving its holder's death. Suite 118/118. |
||
|
|
2ebfccc8ed |
virtio-gpu: the scanout device arrives with the spawn
Third driver converted. It no longer claims the id from argv[1] — the manager holds the device and names it in the call that creates the process, so it is held before the driver's first instruction. display-reattach is the case that matters here: it kills the driver and watches the compositor re-attach to the fresh scanout. It passes, so the restart path survives the fused grant — the manager re-takes the device when the driver dies and hands it to the replacement. ps2-bus and discovery are NOT converted, and the reason is recorded as open question 9 rather than worked around. Both need a device nobody assigned them. ps2-bus ignores its argv[1] entirely: it finds the controller by walking the table for PNP0303, then claims a second device, the PNP0F13 mouse node, which it also finds itself — so it holds two devices and was assigned at most one, while system_spawn carries one. discovery claims the acpi-tables node it locates itself, because it is what produces the device tree and there is nothing to assign at that point. One thing worth checking before designing an answer: devices.csv maps both PS/2 hardware ids to ps2-bus, so the manager may already be spawning two instances where the driver expects one. If so the fix is smaller than it looks. D6 stays blocked — closing device_claim with these two still depending on it would stop the machine booting. Suite 118/118. |
||
|
|
f23f073624 |
kernel: the device rides system_spawn, so a driver never runs without it
Delegation moves out of onHello and into the spawn itself. The manager holds the hardware and names it in the call that creates the driver; the kernel checks the device is the caller's to give, then hands it over as part of making the child. The reason is the window. A transfer after spawning always leaves an interval in which the child is running and does not yet hold its device. It would close on QEMU every time and open occasionally on a machine with different core counts and timing — the exact failure shape this track exists to delete, and not one worth introducing while removing the others. Fused into the spawn there is no interval: the child does not exist until it holds the device. Ownership is checked BEFORE the child is created, so a refusal leaves nothing running rather than a driver without the hardware it was spawned for. The IOMMU confinement moves with the device, as it does on the transfer path. systemCall6 is added for the sixth argument; r9 was free, and abi gains a no_device sentinel matching the protocol's. No driver had to change to receive a device, which is what makes this better than requiring every driver to hello: ps2-bus keeps its legacy status, and discovery — which has no assignment at all, since it is what produces the device tree — is unaffected. The attacker fixture now tries the spawn as a back door: name someone else's device, and both the spawn and any child must be refused. Verifying that assertion exposed a bug in the fixture itself. The kernel case's pass marker was "device-authority: ok", which matches the FIRST per-assertion line, so its wait loop exited before any failure was printed — the case would have passed with failures in it, and had been able to since D2. The verdict lines now carry a distinct VERDICT prefix, and with the ownership check removed the case genuinely fails. A green test that cannot go red is worse than no test. Suite 118/118. |
||
|
|
637bf2e0b1 | docs: correct a stale ordering note — D9 landed without D7 | ||
|
|
cb8d1e2e51 |
docs: the grant rides system_spawn, atomically
Questions 6 and 7 both dissolved on inspection, so what was left was only where the grant is delivered. Three candidates: transfer after spawn, every driver hellos, or fuse the device into system_spawn. Take the third. The manager cannot transfer before the child exists, so a separate transfer always leaves a window in which the child is running and does not yet hold its device. That window would close on QEMU every time and open occasionally on a machine with different timing — the exact failure shape this track exists to delete, and not worth introducing while removing the others. Fusing it into the spawn removes the window by construction: the child does not exist until it holds the device. It adds no knowledge to the kernel, only atomicity — the same rule, you may give away what you hold, fused with the call that creates the recipient. system_spawn uses five of six argument registers, and no_device is already the sentinel. Making every driver hello is a good idea on its own merits — uniform liveness, the deadline applied to all rather than some, and the speaks_protocol two-class split leaving the manager, since a wedged ps2-bus is invisible to its supervisor today. Kept as its own step so grant delivery does not force it. Order is now D0 -> D5 -> D6 -> D8, which deletes maximum_children_per_parent. Only D7 remains blocked, on question 8, and it is needed for neither ceiling. |
||
|
|
8b628a4a7d |
docs: there is no virtio-gpu standalone bring-up to lose
Question 7 read the comment "Best-effort: standalone bring-up has no manager" as a boot path that delegation would delete. It is not one. virtio-gpu's main requires argv[1] and exits without it, and the only source of that argument is the device manager: ACPI, then pci-bus reports the function, then devices.csv matches 1AF4:1050, then the manager spawns the driver with the id. Without a manager the driver prints and returns, and never reaches the hello at all. The comment is about resilience, not boot: the hello is best-effort so a driver whose manager has died keeps serving, which device-manager.md states outright. So the question dissolves. The residual is one narrow window — the manager spawns a driver and dies before transferring — where today the driver could still claim because claiming is free-for-all, and after D6 it would exit for the restarted manager to respawn. That is the better behaviour: a driver holding hardware nobody assigned it is exactly what D6 exists to stop. Taken with the transfer-at-spawn delivery point, D5, D6 and D8 are unblocked. D7 remains blocked on question 8, and is not needed for either ceiling. |
||
|
|
a5840789dd |
docs: the acpi service is discovery, not a driver awaiting an assignment
Question 6 lumped the acpi service in with ps2-bus as "a claimant that does
not hello". That was wrong about what it is. The manager starts it as
addDriver("discovery", no_device, false): the discovery service, one per
firmware, with no device assignment at all, which finds and claims the
acpi-tables node itself because it is the thing that produces the device
tree. The manager cannot hand it a device — at that point there is nothing
to match against.
So its question is not whether it should hello. It is whether the bootstrap
is exempt from D6, or whether the manager claims the acpi-tables node and
passes it on.
Recorded alongside it: a delivery point that needs no hello is already
implied by the code, since spawnSupervised returns the child's pid and the
manager could transfer immediately after spawn. If that is acceptable,
question 6 largely dissolves — ps2-bus would not need to leave "legacy"
either. The cost is ordering: the child may reach for the device before the
transfer lands, where hello guarantees it cannot because the child is the
one asking. That trade is the real decision, and it is now written down as
such.
|
||
|
|
f7151ed577 |
kernel: the device table has no ceiling; a runaway is charged to whoever caused it
maximum_devices = 64 is gone. It was a guess about someone else's computer, and because it was shared, one driver's enumeration starved every other — which is how an AMD Ryzen booted with a working display, no USB and no storage. The table now grows from the kernel heap. It was always built after heap.init; nothing ever prevented this except it having been written static first. What replaces it is an allowance charged to the registrar, so a driver looping device_register exhausts its own and every other driver carries on. It is declared as what it is — a runaway detector, NOT a security boundary. A quota generous enough never to bite a real machine is still generous enough to be unpleasant, and it is not trying to be the defence; delegation is. What this catches is a legitimate driver in a loop, early, attributably, and without collateral. Reaching 4096 is a bug report, not a tuning request. The initial block is 8, deliberately small. Sizing it for a typical machine would mean the growth path never ran on the hardware we test on and only woke up on someone else's larger machine — the exact failure shape this track exists to stop. At 8 it grows several times every boot; disabling growth now fails the suite with the HPET not fitting, which is the Ryzen failure in miniature. The comptime coupling assert added earlier fired, and was right to. confined (one slot per device id) and domains (the IOMMU's own translation pool) were sized by the same constant only because device ids happened to stop at 64 too. Two unrelated quantities: confined now grows with the device table, while maximum_domains stays as the hardware's number — both VT-d and AMD-Vi report how many domains they support, and reading it is phase 4. The assert existed for exactly this and did its job. Suite 118/118. |
||
|
|
5492607a61 |
docs: D7 blocked, and D8 reordered after D9
D7 would remove zero-resource devices from the kernel. It cannot proceed because nothing else mints their ids. A USB interface registers with resource_count = 0, and the id device_register hands back is load-bearing in three places: the child_added packet's target, the class driver's argv[1], and the device_token of the usb-transfer WIRE protocol — so the id space is visible on the wire, not merely internal. usb-xhci-bus records the fourth constraint itself: the kernel's idempotency is what makes the same port and interface map back to the same id across a bus restart, which is what stops a respawned bus spawning duplicate class drivers. Moving that out means answering who mints the id, how it survives a bus restart, how it survives a manager restart, and whether device_token changes meaning. A design step, not a mechanical one. D8 is reordered to run after D9, correcting the original sequencing. D8's justification was that the authorisation the per-parent cap stood in for now exists — but D6 is blocked, so it does not, and deleting the shared cap now would reopen the exhaustion hole it was written for. D9's per-holder quota closes that hole independently of authorisation, and closes it better: a rogue exhausts its own allowance instead of the table everyone shares. Once the quota exists the per-parent cap is redundant either way. D9 also no longer depends on D7. Its rationale was that zero-resource children are the case that sidesteps containment — true, but a per-holder quota bounds them as well as anything else, because it counts entries per holder rather than per parent. |
||
|
|
81dd1e9318 |
pci: the host bridge arrives by delegation too
pci-bus joins usb-xhci-bus in receiving its device from the manager rather
than claiming the id it found in argv[1]. Its hello moves ahead of the ECAM
mapping, since that is where the bridge now arrives, and its hello was
already mandatory so nothing about its failure behaviour changes.
isDelegated compared whole strings, which silently missed this driver: the
manager records the boot-snapshot match as the bare "pci-bus" and a
devices.csv match as the full "/system/drivers/pci-bus". pci-bus was then
neither claiming nor delegated and died on "ECAM mmio_map failed". It now
matches on the last path component. Reintroducing the whole-string compare
breaks usb-xhci-bus instead of pci-bus — the two spellings swap which driver
loses — so usb-hid is the case that catches it, not pci-scan.
pci-scan asserts the delegation on the initial bring-up AND after the
restart drill, with the device id backreferenced so both must name the same
device. That is what proves the manager re-takes a device when its driver
dies and hands it to the replacement, which is the property the whole
supervision design rests on.
The remaining three claimants are NOT converted, and the plan records why
rather than working around it. ps2-bus and the acpi service never hello at
all, which device-manager.md states deliberately ("legacy drivers ... not
yet required to hello"), so delegating to them means either promoting them
out of legacy or giving the grant a delivery point that is not hello.
virtio-gpu hellos best-effort by design — "standalone bring-up has no
manager" — and delegation would make it mandatory. Both are decisions, not
mechanical steps.
Consequence: D6 is blocked, because device_claim cannot be closed off while
three claimants still depend on it. D7-D9 are unaffected — they concern what
the kernel stores and how its table is sized.
Suite 118/118.
|
||
|
|
a2d6056772 |
usb: the xHCI controller arrives by delegation, not by claiming
The first driver to stop claiming its own hardware. The device manager holds the controller and transfers it in the hello reply, so its matching becomes authoritative instead of advisory — until now the driver claimed the id it found in argv[1], and any process could have claimed the same integer first. The manager claims before it spawns, so there is no window in which anything else could take the device, and transfers in onHello using invocation.sender — the kernel-stamped task id, which cannot be forged by the caller. hello is synchronous, so the transfer has completed before the reply lands: no gap between being told yes and holding the thing. usb-xhci-bus's hello moves from after controller bring-up to before anything that needs the device, which is the bring-up reorder the design predicted. It is the first member of an explicit delegated set, so every unconverted driver keeps claiming exactly as before and the suite stays green; the set and device_claim both go at D6. D3 and D4 could not be separated and the plan records why: the moment the manager claims, any driver still calling device_claim is refused, and D3 applied to nothing changes no behaviour and cannot be tested. This step introduced a regression and the incremental conversion is what caught it. confineDevice runs inside systemDeviceClaim, so a device arriving by transfer was never confined for its new owner. Three IOMMU+USB cases failed on the driver's DMA rings going unbound, and two worse consequences were latent: a manager death would have torn down a domain a live driver was using, and a driver death would have leaked one. iommu.reassign now moves the confinement with the device, keeping the domain and its attachment intact so it never translates through nothing. Converting all five drivers at once would have produced the same three failures with five suspects. A log line of mine claimed "holding controller device N" before anything verified it — it printed even in the failure case, where the driver held nothing. Reworded to state only what is known there: where the registers are. usb-hid asserts the delegation with the device id backreferenced, so the id delegated and the id the driver ends up with must match. Emptying the delegated set fails it with "hello acknowledged" then "mmio_map failed". usb-hub failed once in a full run and has passed six times since (four isolated, two full) — recorded in the plan as a suspected instance of the known intermittent AP fault, not dismissed, since this step did shift boot timing. Suite 118/118. |
||
|
|
33376f24ab |
test: the attacker the device suite never had
The audit's sharpest finding was structural, not a bug: a fully green suite had hidden six real defects because it contains no attacker. Every device case asserts that a driver handed its own hardware can drive it. None asked what a process handed NOTHING can do. device-authority-test is that process. It is spawned with no device and asserts what it therefore cannot do: it cannot give away a device another task holds, nor a free one, because the kernel's rule is that you may give away what you hold and the device's state is irrelevant to a process holding nothing. Asserted across every device the machine actually has, so it cannot pass by accident of which one happened to be free at boot — six on QEMU, none of them its. A positive control runs first. device_enumerate works from this process, so the refusals below it are decisions rather than a syscall path that is simply broken here; without it, "everything failed" would read identically to "the assertions are meaningless". A nonexistent device is refused as NoSuchDevice rather than NotHeld, because a refusal that cannot name its own rule is what cost a debugging session on the Ryzen. What it deliberately does not assert, and says so in its header: device_claim is still first-come-first-served at this point in the run. That is the hole D6 closes, and the claim half of the invariant joins this fixture then. Asserting it now would be writing a test that documents the bug. Verified to discriminate: removing the holder check flips "every transfer by a non-holder is refused" while the positive control keeps passing. Suite 117 -> 118. |
||
|
|
3111c7c5e6 |
kernel: device_transfer — you may give away what you hold
The mechanism behind delegation, which device-manager.md named as the step after hello: the device manager claims what discovery seeded and hands each device to the driver it matched, so assignment stops being first-come-first-served. It is a MOVE, not a copy. A claim is exclusive (driver-model.md, invariant 1), so the giver stops holding the device the instant the receiver starts. That is why this is a new syscall rather than the M13 capability path, where a passed handle is shared refcounted — exclusivity cannot be expressed that way. The kernel's whole rule is that you may give away what you hold. It has no notion of which task is the device manager and deliberately gains none: a binary name inside the kernel is not something that cannot safely live in user space. A recipient that does not exist is refused, because a device moved to nobody would be unreachable for the rest of the boot — nothing un-holds a device but task death. Three errnos, each naming its own rule: ENODEV no such device, EPERM you do not hold it, ESRCH no such recipient. Nothing uses it yet. The five claimants move across one at a time in D4-D5, so the suite stays green throughout and a regression names the driver that caused it. Ten assertions, verified to discriminate: removing the ownership check flips four of them, including the giveaway that an illegal transfer then blocks the legitimate claim behind it. Suite 116 -> 117. |
||
|
|
547d0ec46b |
docs: device authority — the how, and the run that deletes the ceilings
device-authority.md is rewritten as an implementation design rather than a rival to device-manager.md. The what was already settled there in 2026-07: structure in the manager, authority in the kernel, and delegation as the step after hello. This is the how, plus the two decisions that paragraph leaves open. Decision 1: the manager claims, it is not granted. device-manager.md says "claims (or is granted)"; claiming wins because the manager runs before any driver exists and takes the seeded devices unopposed, leaving nothing unheld to race for. One new call, device_transfer(device_id, task_id), checks only that the caller holds the device — no names in the kernel, no attestation. The alternative put a binary path inside the kernel, and the kernel should hold only what cannot safely live in user space. The residual is stated rather than hidden: authority rests on the manager claiming first, which init.csv makes an operator-visible ordering rather than an attacker- controlled one, and the enforced version arrives with the spawn capability drivers.md already names as missing. Decision 2: the kernel stops holding inventory. It reads three things out of a descriptor — physical ranges, interrupt numbers, one PCI BDF — and stores the rest only so device_enumerate can hand it back. Devices with no resources leave the kernel entirely: a USB device conveys no mapping authority, so there is nothing to enforce. That is also the case which sidesteps containment, and therefore the reason a shared cap existed. Decision 3: no shared ceiling. The table becomes dynamic — it is built after heap.init, so nothing ever prevented it — and the two invented numbers go. A per-holder quota replaces them, because dynamic storage with no bound moves the ceiling to the kernel heap, which is shared and fatal rather than partial. A bound charged to whoever caused it is isolation. Two earlier drafts of this document are gone: one gave init the root grants, the other proposed extracting a firmware-framebuffer driver. Both were wrong and both are recorded as wrong in the run plan's settled list — the framebuffer is not a device, it is where pixels go until a real display driver announces itself. Run 2 is nine steps, ordered so the suite stays green throughout: build and prove the transfer mechanism, move the five claimants across one at a time, then the flag day, then the inventory, then the ceilings. |
||
|
|
5d55217212 |
usb: a slot the controller granted is always handed back
Both device-setup paths issued a successful Enable Slot and then returned null if allocateDevice failed, without disabling it. A slot the driver forgets is one the controller never reissues, so each attempt lost one permanently for the boot. The hub path did it with no log line at all. Both now release the slot through a shared disableSlot, extracted from tearDownDevice, and the hub path warns like the root-port path does. tearDownDevice also now frees the interface list. That allocation arrived with the previous commit, so an unplug would have leaked it — found while reading the teardown path for this fix rather than by a test. No regression test, and it is recorded as open question 5 rather than implied. After the slot count became the controller's own figure, reaching this path needs more devices than the controller has slots: QEMU offers four against sixty-four. What was verified is that the new path RUNS correctly — pinning tracking to 2 with four devices attached produced "port 6 setup: no free device slot", the first two devices enumerated normally, and no Disable Slot error or timeout appeared, which is how disableSlot reports failure. Suite 116/116. |
||
|
|
729b40ece7 |
usb: a device has as many interfaces as it declares
max_interfaces was 4. A composite device — a headset, a webcam with audio, a dock, a multifunction printer — routinely has more, and the fifth did not merely go missing. parseConfiguration's cap branch had no `else`, so when the count was reached `current` kept pointing at interface 3 and the fifth interface's endpoint descriptors were appended to interface 3's array. A class driver bound to interface 3 could then be handed an endpoint belonging to something else entirely, and subscribe or bulk-transfer on it. The alternate-setting arm one line above cleared `current` correctly, which is what the cap branch should have done. Interfaces are now counted from the block in a first pass and allocated to exactly that number, so the ceiling is bNumInterfaces' u8 — the USB specification's. The missing `else` is added too, though after this the bug is unreachable by construction: interface_count cannot reach interfaces.len mid-parse when the list was sized from the same walk. max_configured_endpoints was max_interfaces * max_endpoints_per_interface = 16, a derived guess that moved whenever either input moved. It is now 31, which is the xHCI specification's own limit: a Device Context holds a slot context plus at most 31 endpoint contexts, because the Context Entries field addressing them is 5 bits. max_endpoints_per_interface stays at 4 with its reason recorded — the usb-transfer wire protocol reports exactly max_reported_endpoints (4) per interface, so widening it alone would change nothing a class driver sees. Lifting it is a protocol change. No direct test, and that is written down as open question 5 rather than glossed. The parser is pure and wants a host unit test, but usb-xhci-library.zig imports memory, mmio and time so it cannot be a standalone test root, and QEMU offers nothing that reaches the path — the largest device available is usb-audio,multi=on at 2 interfaces and 211 bytes. The alternate-setting path that shares the same `current = null` logic is exercised by that device. Suite 116/116. |
||
|
|
6328823ef1 |
usb: a configuration block is as long as the device says it is
The driver read the first 512 bytes of a configuration block into a fixed buffer and parsed those. The block's length is the device's own choice (wTotalLength, a u16), so anything larger was silently cut: interfaces past the cut did not exist as far as the host was concerned, while the SET_CONFIGURATION that follows still configured the device for all of them. A headset is 500-900 bytes, a UVC webcam 1-3 KB, a multifunction printer 600+. Now allocated at the declared length, so the ceiling is the field's u16 — the specification's number rather than one of ours. A block shorter than its own 9-byte header is refused rather than trusted. The bring-up line reports the declared length and the bytes actually read, so a truncation can never again be invisible, and usb-large-descriptor asserts they match with a backreference. That case has an honest limit, recorded in its comment: QEMU cannot produce a block over 512 bytes. The boot keyboard, mouse and stick are 34-44, and the largest device available is usb-audio in multi-channel mode at 211 — which is exactly why the suite never caught this, and why it cannot now reproduce the original trigger. What it does catch is the class: any clamp below the attached device's block fails it, verified by pinning the buffer to 128 and watching "config block 211 bytes, read 128" turn the case red. Suite 115 -> 116. |
||
|
|
69fbef40c0 |
usb: the controller says how many device slots it has
max_devices was 8, with the comment "QEMU presents a handful; a fuller machine would grow this" — a number chosen against the test rig, waiting for a real machine, which is the pattern docs/bounds-track-plan.md exists to stop. The driver already knew the true figure. It reads HCSPARAMS1.MaxSlots at bring-up and writes it straight into op_config, so every slot the controller offers has always been *enabled*; only the array tracking them was 8. QEMU's xHCI reports 64, so seven eighths of the controller was live and invisible, and the ninth device — a keyboard, mouse, webcam, headset, hub and two sticks reach that without trying — disappeared on a hub-attached path that logs nothing at all. The array becomes a slice allocated from max_slots at bring-up. A controller claiming zero slots cannot address anything, so that is now a dead controller rather than an empty allocation failing mysteriously later. The Device Context Base Address Array is a page, 511 usable entries, so it already covered the 255-slot maximum. The bring-up line reports both numbers, and usb-hid asserts they are equal with a backreference rather than a magic number, so the test cannot drift from the hardware. Pinning tracking back to 8 fails it: "64 slots, tracking 8". Suite 115/115. |
||
|
|
f4eb88e7d2 |
build: a new compile-time ceiling declares itself or does not land
The convention that tunables live in system/parameters.zig with their reasoning attached predates this and got 2% compliance — 5 of 235. A convention with no teeth is how a bare `const maximum_devices = 64` reached an AMD desktop and cost it USB and storage. This is the same rule with a gate behind it. tools/check-bounds.py finds every bound-shaped declaration — a `maximum_*` const with a literal value, or a type with a literal array length — and requires the five-field block above it: what it counts, who decides its size, what it protects, what happens at the limit, and how anyone finds out. The at-limit vocabulary is closed: refuse, degrade, truncate, grow. There is deliberately no way to spell "silent", no way to spell "drop", and nothing meaning "allow", so the behaviours that did the damage cannot be written down. Truncation is legal only carrying a marker the reader can see, which is why klog_maximum_message qualifies and a USB descriptor cut at 512 bytes does not. An array length that names a declared bound is not itself a bound; only literal lengths are flagged, which pushes ceilings toward having names. The 273 that predate the rule are allowlisted so this lands without a tree-wide sweep in front of it, and that list may only shrink: declaring a bound means deleting its line, and the check fails on a stale entry too. Nothing may be added. Wired into `zig build test` and available alone as `zig build bounds`. Not in the default build — it reads the whole tree, and a red bounds check should not stop you booting a kernel. Five are now declared rather than allowlisted. Writing them out is its own argument: maximum_devices reads "protects: nothing — this is a sizing guess about someone else's computer", and maximum_tasks now carries the fact that it has been raised twice, each time by something that outgrew it. Verified the gate refuses an undeclared bound, a declared one using forbidden vocabulary, and an allowlist entry that has since been declared. Suite 115/115. |
||
|
|
568823a4fb |
docs: L1 was wrong — reclamation is not a death-sweep problem
The first step of the unattended run was "a dead task's registrations die with its claims". Implementing it would have broken the restart path it was meant to protect. The broker keeps entries on purpose: they describe hardware, which did not go away when a driver died. And device ids must stay stable across a bus restart, because device-manager dedupes re-reports by device_id so a restarted bus does not spawn a second driver instance — stability that comes from the idempotency scan returning the existing id. Removing entries on death would hand a restarted bus fresh ids and duplicate every driver. Everything else a task holds is already reclaimed on every path out: IRQ bindings, IOMMU domains, DMA regions, then its claims. The leak the audit found is real but has two other sources: a device that genuinely goes away has no retirement path, and a bus that enumerates differently on restart strands its old entries. Both belong to the device manager's inventory, and both need the id-stability question settled first — tombstone-and-reuse aliases ids another process still holds, generation tagging changes the id encoding, which is ABI. Recorded as an open question rather than guessed at. |
||
|
|
4398eb7cc4 |
docs: scope the bounds track's unattended run
Six steps an agent can execute: reclamation, the bounds build check, and the four user-space USB/xHCI bounds where the hardware already reports the number we guessed. Deliberately excluded: the authorisation gate and moving the device inventory to the manager. Both decide whether the OS is secure and both are a direction rather than a specification — what a device capability is, which syscalls change, what replaces device_claim for its seven callers. They want a design session, not an agent. Three open questions are written down rather than guessed: device_enumerate most likely narrows to the firmware-discovered roots rather than retiring (the manager cannot ask itself for the PCI host bridge); a manager restart has no re-enumerate handshake, so it comes back blind while its buses live; and "add adversarial tests" is not executable until the attacks are named. |
||
|
|
a86559648e |
kernel: a refusal names its rule, and two bounds stop failing open
An AMD Ryzen booted to a working compositor with no USB and no storage, and the log said only "register refused". A tree-wide audit of every compile-time ceiling followed: 235 of them, 139 on quantities the machine or a file decides rather than us, 5 documented anywhere, 171 silent when reached. docs/fixed-bounds-audit.md has the inventory. Errno attribution. The errno space was split between the kernel and the envelope, free to drift; it is now one list in system/abi.zig, restated on both sides, with a comptime check in library/device/driver where the two halves are visible. device_register's six refusals and device_claim's three are distinct codes, so a bus driver can say which rule stopped it, and BadParent splits into NoSuchParent and NotYourParent. pci-bus reconciles found against registered instead of counting refused functions as found. Idempotency ordering. The child cap was checked before the identity match, so a restarted bus was refused its own devices — the supervision restart the system leans on ratcheted toward a degraded machine. A re-registration consumes no slot and is now admitted first. IOMMU fail-closed. confineDevice returned success for a device id past the confinement table, leaving the device outside every domain while the caller believed it confined — unreachable only while ids stop at 64, which both the inventory move and a hardware-reported domain count would change. It refuses now, and the coupling to the broker's device cap is a comptime assert rather than a sentence in a comment. PCI apertures. The bridge's MMIO apertures are derived from the holes in the firmware memory map, and the derivation copied sub-4 GiB entries into a fixed [64] array and skipped the rest. A skipped region is not merely lost: the gap finder concludes it is free, so a real machine's 60-200 entry map yields an aperture over live RAM, and containment then admits a child BAR covering kernel memory. Rewritten to walk the map in place, with the hole finder extracted as a pure function and driven by a synthetic 100-entry map in a new test case. Both new tests were verified to fail on the old code. parameters.zig gains the rationale it was missing and loses a stale sentence pointing at the wrong file; vdso.md documents the errno space, including EPEER, which had no written meaning anywhere. docs/os-development/bounds.md is how a ceiling is declared from here. docs/bounds-track-plan.md is the plan to remove the ones we invented. Suite 114 -> 115. |