Commit Graph
255 Commits
Author SHA1 Message Date
Daniel Samson 8216be991d docs: correct the storage docs' V0-V4 status — the V5 flip missed several
A V5 close-out audit (docs against the actual code) found the storage docs
overclaiming in both directions: the status header was flipped to "built" but
several body markers were not, and two passages describe mechanisms the code
never implemented. Nine confirmed, adversarially verified against the source:

Underclaims (marked Planned, actually built):
- storage-architecture: driver named sub-ranges / range confinement (V2, tested
  by the block-range case); the pushed medium_changed event (published today
  from a TEST UNIT READY poll); ownership-gated fs_unmount (V0, process.zig
  gates it with EPERM); the filesystem-harness extraction (V1). Scoped the
  remaining *planned* to the genuinely-pending parts (native-signal translation,
  the volume manager consuming medium_changed).

Overclaims (described, never built):
- storage-architecture: the "FAT dirty flag on disk" guarantee — no on-disk
  dirty/clean-shutdown bit exists; only an in-memory device-dirty bool gating a
  device write-cache flush on close. Fixed in all three places.
- storage-architecture: fat "acquires its own volume (first mass-storage child
  by enumeration order)" — the V3b flip removed self-acquisition; fat is handed
  its volume id and channel by the volume manager.
- rationale + plan: the FAT engine's base_lba "deleted rather than moved" — it
  and the engine's MBR walk still exist as now-inert legacy; the authoritative
  walk lives in partition.zig.
- rationale: NVMe namespaces "decision 4 settles as endpoint-per-volume" —
  decision 4 settles the opposite (per-sender confinement, one endpoint);
  endpoint-per-volume is named only as an unbuilt future refactor.
- rationale: the five-rung identity ladder and volumes.csv map stated in flat
  present tense — only rung 4 (MBR signature + index) is built; added the
  build-status hedge and marked each rung.

Docs only; no code or behavior change. Suite unaffected (127/127).
2026-08-09 21:19:29 +01:00
Daniel Samson 5e89b111cf docs: flip the storage-architecture status markers the V0-V4 track made real
The volume manager, per-sender range confinement + gate, medium_changed
event, per-volume spawning + supervision, ownership-gated fs_unmount, and
the removal half of the lifecycle are built. Left honestly pending: the
volumes.csv/filesystems.csv maps, the fuller identity ladder, multi-volume,
the VM consuming medium_changed (removal uses device-presence polling), and
remount-on-replug end-to-end (bench-pending — QEMU can't re-present the
boot-controller device). Full suite 127/127.
2026-08-09 20:20:20 +01:00
Daniel Samson a67a7015bf docs: record the V3c resequencing — removal lifecycle first, multi-volume follows 2026-08-09 18:30:16 +01:00
Daniel Samson 73fbbd3922 docs: record the V2 sequencing — medium_changed emitted in V2b, tested in V4 2026-08-09 17:27:43 +01:00
Daniel Samson b59f981c58 docs: decision 4 settled — a loop with an open decision is not a loop
Per-sender range confinement at the provider, one serving endpoint: the
badge-scoped provider pattern the xHCI bus already uses (the per-client
device-token table), applied to blocks. Endpoint-per-volume would buy the
same enforced property only by inventing a multi-endpoint harness; it stays
available as a future refactor, same wire contract. The plan's flag-for-veto
is gone: nothing in the track waits on a choice.
2026-08-09 16:34:04 +01:00
Daniel Samson 451abba000 docs: the volume-manager plan — V0 unmount ownership through V5 close-out 2026-08-09 16:09:20 +01:00
Daniel Samson 9da63e81e2 docs: the boot volume — identity is what makes yanking it survivable
The sharpest instance of the return story, folded into the architecture: the
boot volume is a recorded content identity (the volume carrying
/system/configuration and /system/logs), the system runs on without it (the
ramdisk is the OS; only log persistence pauses), the kernel ring is the
buffer during absence — bounded, so a wrapped ring is a data loss window
that gets MARKED in the file on resume, never spliced silently — and on
return at any port the same identity remounts the same prefixes and the
logger appends into the same boot-stamp tree. The logger requirement is
named: failed flushes retry on the patient cadence, never abandoned after
the first not_found. The dirty-honesty rule stands; the boot volume gets no
exemption.
2026-08-09 16:01:03 +01:00
Daniel Samson ada251150a docs: volume identity, and volumes.csv as danos's fstab
Decision 8 in the rationale, mirrored into the architecture's volume-manager
section: the mount map keys on CONTENT identity, never port or arrival order
— Linux's /dev/sda1-era fstab broke on every port move until UUID= replaced
it, and danos skips that era. The identity ladder the prober reads off the
medium: GPT partition GUID, filesystem UUID, FAT serial+label, MBR
signature+index, anonymous. Consequences are mechanical: port moves change
nothing (USB, hub, SATA bay, or transport swaps), replug remounts at the
same path, the boot volume is a recorded identity findable anywhere, and
cloned duplicates are a loud policy case instead of silent shadowing.
/volumes/<name> stands as the hierarchy's home for attached media; minting
identifiers (formatting, entropy) stays deliberately out of scope.
2026-08-09 15:59:38 +01:00
Daniel Samson 728b436d0f docs: the storage rationale lives with the architecture it justifies
storage-stack-discussion.md was misplaced at the docs root — that level is
for track plans; this is the file-system domain's design record. Moved to
file-system-development/storage-design-rationale.md, renamed to say what it
is, cross-references updated.
2026-08-09 15:49:12 +01:00
Daniel Samson 092817ba2e docs: transport generality and the media-presence event
The NVMe/SATA assessment folded into the storage discussion: what transfers
untouched (the block contract and everything above it — one driver binary
plus one devices.csv row per transport), NVMe as the shorter stack whose
namespaces are the reserved multi-volume case, AHCI as an open shape choice
(leaning per-port processes, the matrix-proven granularity), and the three
pressure points named honestly: multi-volume is reserved-not-implemented,
synchronous call/reply bottlenecks NVMe until the shm-ring data plane, and
media lifecycle is not device lifecycle.

That last one becomes decision 7 and enters the architecture doc: the
removal path has TWO TRIGGERS, ONE LIFECYCLE — channel death (device leaves)
and a planned pushed medium_changed event (medium leaves, device stays: card
readers and trays, USB ones today), translated by the storage driver from
its transport's native signal, presence never content, consumed by the
volume manager into the same kill-retire-remount path. Without it a swapped
card would be served with the previous card's filesystem state.
2026-08-09 15:36:59 +01:00
Daniel Samson fa54ef6915 docs: the volume lifecycle is enforced, not described
The enforcement section of the storage architecture: the three levers the
device lifecycle already proved, mapped onto volumes — a filesystem can only
be GIVEN its volume (no establishment grants, channel at spawn), the volume
manager supervises with teeth (deadline, kill-on-removal, crash-loop cap),
and the shared harness makes every engine inherit the state machine by
construction (engines never see channels). Kernel backstops: ownership-gated
fs_unmount + the lazy dead-endpoint sweep. Checkable via a lifecycle
conformance drill parameterized over filesystems — supporting a filesystem
MEANS passing it.
2026-08-09 15:13:34 +01:00
Daniel Samson fae616fa5a docs: the storage architecture — layers, boundaries, and who does what when media leaves
The settled shape from the storage-stack discussion, written as the
reference: the data path (vfs -> filesystem service -> block -> driver) as
the application/service/protocol/driver model applied twice; the three kinds
of boundary (protocol between processes, library inside them, control-plane
beside them); the volume manager as the policy home (planned — the FAT
service squats on its duties today, marked as such); adding a filesystem as
engine + shared harness + one configuration row; and the per-layer
responsibility table for removable media — one removal path, kill/retire/
respawn, dirty data lost and SAID to be lost. Indexed from docs/README.md.
2026-08-09 15:08:45 +01:00
Daniel Samson 5a8a2a4d7e docs: the storage-stack discussion — block, volumes, filesystems, against the survey 2026-08-09 14:47:35 +01:00
Daniel Samson d4f8dc51b9 docs: the hot-plug matrix is complete — 124/124 2026-08-09 14:19:39 +01:00
Daniel Samson d565a6b845 docs: the hot-plug matrix plan — unplug anything, replug anywhere 2026-08-09 13:58:29 +01:00
Daniel Samson 3c9f454398 establishment: the two-controller proof, and the docs catch up
P4 of docs/establishment-planes-plan.md. The new usb-two-controllers case is
the Ryzen mouse bug pinned in the suite: a second xHCI controller with its
own keyboard while the boot controller keeps the default one — both must
come up, on different device ids, which requires each class driver to reach
ITS OWN controller. Discrimination: at 72807c2 (name-based establishment)
the case fails — one keyboard is unreachable, exactly the bench failure —
verified against a checkout of that merge; with lineage routing it passes.
The existing second-controller cases could not prove this: the boot bus
always carries a keyboard, so their expects were satisfiable by it.

Docs updated with the code: device-manager.md (hello moves the channels,
paged enumerate, delegation as built, reap-and-rebuild in the restart
sequence), device-authority.md (the fourth as-built decision: driver-layer
channels ride the hello; hello is no longer only the liveness handshake).

Full suite: 118/119, the one failure being iommu-fault's fixed 200 ms
fault window under end-of-run host load — 3/3 green standalone, deflake
flagged separately. usb-two-controllers passed inside the full run.
2026-08-09 12:48:33 +01:00
Daniel Samson 0d5a7394ef docs: establishment planes — the design and the two-seam plan
The namespace holds protocols, never instances; what multiplies is provider
processes, and their establishment routes through the device manager, which
owns the topology. communication.md gets the model; the plan converts the
usb-transfer and block seams in one flag-day, closes the restart-zombie hole
the scoping found, and pins the three-controller mouse bug as a QEMU case.
2026-08-09 11:26:53 +01:00
Daniel Samson 5930c9653c docs: device-authority as built — spawn carries the device, the loan, confinement rebuilt on return 2026-08-09 10:17:46 +01:00
Daniel Samson 0eb2420690 device-manager: delete the delegated-set scaffolding
The name list and its predicate existed so drivers could move to delegation
one at a time with the suite green throughout. Every driver is delegated
now, so the manager simply hands over whatever device a driver was assigned.

Deleting it caught a real consequence: crash-test finally got delegated too,
and it was still claiming its device — so it got AlreadyClaimed because it
already held it, exited, and the restart drill had nothing to restart. Its
own comment named what the case was really checking: "the respawn only
reaches this line because the kernel released the previous instance's claim
at death". That property still holds, by a different mechanism — the device
reverts to the manager on death and is handed to the replacement, which is
the same guarantee without the race it used to rely on.

All four delegation paths verified: the xHCI controller, the PCI bridge, the
PS/2 two-node singleton, and virtio-gpu's restart re-attach.

Run 3 complete. Suite 118/118.
2026-08-08 23:01:29 +01:00
Daniel Samson df9c1ed827 device-manager: hold the seeded hardware so none is left lying around
A device nobody holds can be claimed by anyone, so the manager now takes
every firmware-discovered device that carries mappable resources, whether or
not a driver wants it. The real gap was the HPET: an MMIO window, an IRQ, no
user-space driver, and there for the taking. Held by the manager it is
inert; unheld it was a way into physical memory.

Two deliberate exclusions. The loader's framebuffer, which the compositor
claims and which the manager must not take because it starts first. And
anything with no resources, which grants nothing worth holding.

Scope is the boot snapshot. A device reported later and matched to no driver
stays claimable — pci-cap-test and iommu-fault-test both reach an unmatched
NIC that way, so narrowing it is a separate change with those fixtures in
scope. Recorded in the plan rather than left implied.

The attacker fixture gains the assertion deferred since D2: after the system
settles, nothing with resources may be taken.

That assertion defeated itself twice before it worked, and both failures are
worth remembering. First it swept at 0.029 while the manager did not bind
its protocol until 0.047, so it reported a hole that closed a millisecond
later. The retry loop that "fixed" that was worse: the first pass TAKES the
device, so the second finds it unavailable because this process now holds
it, and concludes all is well — it passed with the manager's claiming
removed entirely. It now settles once and sweeps once, and fails when the
claiming is removed.

Suite 118/118.
2026-08-08 22:42:42 +01:00
Daniel Samson ba195fa0a2 kernel: claim refuses delegated hardware — and E2 had already closed the hole
The rule as planned: a device that was given to someone may be handed on,
never taken. Implemented, and honest about what it is worth.

Writing the test showed the plan had the wrong step doing the work. A
delegated device is HELD, so an attempt to take it is refused as
AlreadyClaimed before the giver is ever consulted; and once a borrower's
death returns the device to its lender — or clears both when the lender is
gone — there is no state where a device is unheld and still on loan. The
window a stranger could have used stops existing at E2. This check is
unreachable.

It stays anyway: one comparison, failing closed, guarding any future path
that frees a device without clearing its giver, which is exactly the hole
this run closed. The comment says it is unreachable rather than implying a
protection it does not provide.

The attacker fixture does not gain the assertion that was deferred to this
step, and its header records why: there is no refusal for it to observe, and
on a bare boot with no device manager nothing is delegated at all, so the
assertion had nothing to bite on. It failed loudly on its first run rather
than passing quietly, which is the only reason this was noticed.

It also leaves the loader's framebuffer alone without naming it: nobody
delegates the framebuffer, so it has no giver, so the display service claims
it exactly as before.

Suite 118/118.
2026-08-08 22:31:04 +01:00
Daniel Samson 1a1d92cba9 kernel: a grant is a loan — a dead borrower returns the device to its lender
When a driver dies, a device it was *given* now goes back to whoever lent
it, rather than to nobody. The device manager gets its hardware back the
instant a driver dies and hands it to the replacement, with no window in
between.

That window was real: the kernel released the claim to no one and the
manager re-claimed first-come, so every driver restart reopened the hole
this run is closing. It also becomes load-bearing at the next step — once
claim refuses a device that has a giver, releasing to nobody would strand a
dead driver's hardware permanently, because nobody could ever take it again.

A dead lender is no lender: the claim and the giver clear together, so a
device is never owed to a ghost. A device nobody lent is released outright,
exactly as before.

The broker cannot see the task table, so liveness arrives through the same
hook idiom the scheduler already uses. Null means assume dead, so a kernel
built without the hook frees claims rather than handing them to a ghost.

A stale binary nearly passed as proof for the third time this session: the
first discrimination patch left `alive` unused, the build failed with three
errors, and the old binary reported every assertion passing. Checking the
build before reading results is what caught it.

Suite 118/118.
2026-08-08 22:21:25 +01:00
Daniel Samson 4ca57fc37e kernel: record who gave each device away
One field, and the rest of the run follows from it. A device that was given
to someone is delegated hardware: it may be handed on, never taken, and when
its holder dies it goes back to whoever lent it instead of becoming free for
anyone to grab.

It also settles the framebuffer without mentioning it. Nobody delegates the
loader's framebuffer, so it has no giver, so the display service claims it
exactly as it always has — no exemption and no reference to display anywhere
in the rule.

No behaviour changes here; the field is recorded and read by nothing yet.

The test found a real bug on its first run, before the discrimination check.
The sentinel for "nobody gave this" was 0 — and task 0 is a real task, the
kernel's own, so a device given away by task 0 read back as belonging to
nobody. Both giver and registrar are optionals now. The second was a latent
bug from D9: the per-registrar allowance would have miscounted every device
task 0 registered.

Suite 118/118.
2026-08-08 22:12:01 +01:00
Daniel Samson ca1126537d docs: Run 3 — close the claiming hole, six steps, no open questions
One field settles both outstanding questions. The kernel records who gave
each device. device_claim then refuses anything that has a giver, and on
death a device reverts to its giver rather than to nobody.

The framebuffer needs no exemption: nobody delegates it, so it has no giver,
so the display service claims it exactly as today. The rule never mentions
display.

And restart stops racing. Today a dying driver releases its claim to nobody
and the manager re-claims first-come, so every restart reopens the hole this
run closes; a loan that reverts to its lender removes the window entirely.

The zero-resource inventory question is dropped from the plan rather than
carried as a blocker. It is real but nothing depends on it.
2026-08-08 22:01:49 +01:00
Daniel Samson 337b981c07 docs: reorder the Run 2 table to the plan order and note the executed sequence 2026-08-08 21:58:53 +01:00
Daniel Samson b7d97ebb5d acpi: discovery is handed its node like every other driver
The last claimant. The kernel seeds the acpi-tables node, so it sits in the
same boot snapshot the manager already scans to find the PCI host bridge —
there was never a bootstrap problem, only a lookup nobody had written. The
manager claims it and names it in the spawn; the service stops claiming.

Every driver in the system now receives its hardware rather than taking it.

Two failures on the way, both mine. addDriver puts the device id in argv[1],
and the acpi service read argv[1] as a self-verify device-count floor — so
handed device 7 it decided it was in test mode, printed "acpi-parse: ok",
and never reported a device. The test argument is now floor:N, which a bare
id cannot be mistaken for.

And acpi-parse spawns the service directly rather than through the manager,
so nothing handed it the node. That test now claims and transfers it exactly
as the manager does, which is the right shape: the test plays the manager's
role instead of the service reaching for hardware.

device_claim now has two callers left: the manager, which is the acquirer
and should have it, and the display service's GOP path. That is recorded as
question 10 — the framebuffer is not a device, so the answer is likely that
it leaves the device table rather than being exempted from its rules.

Suite 118/118.
2026-08-08 21:54:45 +01:00
Daniel Samson 7d8aa51234 kernel: maximum_children_per_parent is gone
The second invented ceiling. It was written to stop a driver looping
device_register and exhausting a shared table — but there is no shared table
to exhaust any more, and each registrar already has its own allowance, so a
runaway costs only itself.

It never bounded a determined caller in the first place: 16 children per
parent, and nothing stopped it claiming more parents. What it reliably did
was refuse a real PCI bus with more than 16 functions, which is how an AMD
Ryzen booted with a working display, no USB and no storage.

The constant, its check, and the now-unused childCount all go. TooManyChildren
survives with one meaning instead of two: the caller is at its per-registrar
allowance.

This was unblocked from the moment D9 landed. The plan said so — "once the
quota exists the per-parent cap is redundant whether or not D6 has landed" —
in the same edit that left the step tagged "blocked on D6". Three iterations
were then spent re-reading that tag instead of the sentence beside it.

The containment test now asserts 64 children under one parent, four times
the old ceiling; restoring the cap fails it.

Suite 118/118.
2026-08-08 21:12:20 +01:00
Daniel Samson 3ae541214f iommu: assert directly that a confinement moves with its device
reassign was added at D4 to fix a regression and has been proven only
indirectly since — three IOMMU+USB cases going green. That covered the
visible symptom (a driver's DMA rings unbound) and neither of the latent
ones: the confinement still naming the giver, so the giver's death would
tear down a domain a live driver was using, and the receiver's death would
leave one behind. Those are now asserted.

confinementOwner exposes the record's owner so the suite can see it. The
sequence is the delegation in miniature: unconfined, confine as this task,
reassign to another, confirm the new holder owns it and the old one does
not, then kill the new holder and confirm the domain goes with it.

Two attempts at this test could not have failed. The first found no PCI
function to confine — pciAddressOf needs a pci_device entry and this case
runs no pci-bus — so every assertion skipped silently while the case stayed
green. It now synthesizes a function the way pci-bus does, a 4 KiB config
window inside the bridge's ECAM, and asserts that precondition explicitly so
a skip is a failure.

Verified to discriminate: making reassign a no-op flips three assertions,
including the domain surviving its holder's death.

Suite 118/118.
2026-08-08 19:57:36 +01:00
Daniel Samson 2ebfccc8ed virtio-gpu: the scanout device arrives with the spawn
Third driver converted. It no longer claims the id from argv[1] — the
manager holds the device and names it in the call that creates the process,
so it is held before the driver's first instruction.

display-reattach is the case that matters here: it kills the driver and
watches the compositor re-attach to the fresh scanout. It passes, so the
restart path survives the fused grant — the manager re-takes the device when
the driver dies and hands it to the replacement.

ps2-bus and discovery are NOT converted, and the reason is recorded as open
question 9 rather than worked around. Both need a device nobody assigned
them. ps2-bus ignores its argv[1] entirely: it finds the controller by
walking the table for PNP0303, then claims a second device, the PNP0F13
mouse node, which it also finds itself — so it holds two devices and was
assigned at most one, while system_spawn carries one. discovery claims the
acpi-tables node it locates itself, because it is what produces the device
tree and there is nothing to assign at that point.

One thing worth checking before designing an answer: devices.csv maps both
PS/2 hardware ids to ps2-bus, so the manager may already be spawning two
instances where the driver expects one. If so the fix is smaller than it
looks.

D6 stays blocked — closing device_claim with these two still depending on it
would stop the machine booting.

Suite 118/118.
2026-08-08 19:48:00 +01:00
Daniel Samson f23f073624 kernel: the device rides system_spawn, so a driver never runs without it
Delegation moves out of onHello and into the spawn itself. The manager holds
the hardware and names it in the call that creates the driver; the kernel
checks the device is the caller's to give, then hands it over as part of
making the child.

The reason is the window. A transfer after spawning always leaves an
interval in which the child is running and does not yet hold its device. It
would close on QEMU every time and open occasionally on a machine with
different core counts and timing — the exact failure shape this track exists
to delete, and not one worth introducing while removing the others. Fused
into the spawn there is no interval: the child does not exist until it holds
the device.

Ownership is checked BEFORE the child is created, so a refusal leaves
nothing running rather than a driver without the hardware it was spawned
for. The IOMMU confinement moves with the device, as it does on the transfer
path. systemCall6 is added for the sixth argument; r9 was free, and abi
gains a no_device sentinel matching the protocol's.

No driver had to change to receive a device, which is what makes this
better than requiring every driver to hello: ps2-bus keeps its legacy
status, and discovery — which has no assignment at all, since it is what
produces the device tree — is unaffected.

The attacker fixture now tries the spawn as a back door: name someone else's
device, and both the spawn and any child must be refused. Verifying that
assertion exposed a bug in the fixture itself. The kernel case's pass marker
was "device-authority: ok", which matches the FIRST per-assertion line, so
its wait loop exited before any failure was printed — the case would have
passed with failures in it, and had been able to since D2. The verdict lines
now carry a distinct VERDICT prefix, and with the ownership check removed
the case genuinely fails. A green test that cannot go red is worse than no
test.

Suite 118/118.
2026-08-08 19:39:32 +01:00
Daniel Samson 637bf2e0b1 docs: correct a stale ordering note — D9 landed without D7 2026-08-08 19:26:48 +01:00
Daniel Samson cb8d1e2e51 docs: the grant rides system_spawn, atomically
Questions 6 and 7 both dissolved on inspection, so what was left was only
where the grant is delivered. Three candidates: transfer after spawn, every
driver hellos, or fuse the device into system_spawn.

Take the third. The manager cannot transfer before the child exists, so a
separate transfer always leaves a window in which the child is running and
does not yet hold its device. That window would close on QEMU every time and
open occasionally on a machine with different timing — the exact failure
shape this track exists to delete, and not worth introducing while removing
the others. Fusing it into the spawn removes the window by construction: the
child does not exist until it holds the device. It adds no knowledge to the
kernel, only atomicity — the same rule, you may give away what you hold,
fused with the call that creates the recipient. system_spawn uses five of
six argument registers, and no_device is already the sentinel.

Making every driver hello is a good idea on its own merits — uniform
liveness, the deadline applied to all rather than some, and the
speaks_protocol two-class split leaving the manager, since a wedged ps2-bus
is invisible to its supervisor today. Kept as its own step so grant delivery
does not force it.

Order is now D0 -> D5 -> D6 -> D8, which deletes maximum_children_per_parent.
Only D7 remains blocked, on question 8, and it is needed for neither ceiling.
2026-08-08 19:26:32 +01:00
Daniel Samson 8b628a4a7d docs: there is no virtio-gpu standalone bring-up to lose
Question 7 read the comment "Best-effort: standalone bring-up has no
manager" as a boot path that delegation would delete. It is not one.
virtio-gpu's main requires argv[1] and exits without it, and the only source
of that argument is the device manager: ACPI, then pci-bus reports the
function, then devices.csv matches 1AF4:1050, then the manager spawns the
driver with the id. Without a manager the driver prints and returns, and
never reaches the hello at all.

The comment is about resilience, not boot: the hello is best-effort so a
driver whose manager has died keeps serving, which device-manager.md states
outright.

So the question dissolves. The residual is one narrow window — the manager
spawns a driver and dies before transferring — where today the driver could
still claim because claiming is free-for-all, and after D6 it would exit for
the restarted manager to respawn. That is the better behaviour: a driver
holding hardware nobody assigned it is exactly what D6 exists to stop.

Taken with the transfer-at-spawn delivery point, D5, D6 and D8 are
unblocked. D7 remains blocked on question 8, and is not needed for either
ceiling.
2026-08-08 19:22:22 +01:00
Daniel Samson a5840789dd docs: the acpi service is discovery, not a driver awaiting an assignment
Question 6 lumped the acpi service in with ps2-bus as "a claimant that does
not hello". That was wrong about what it is. The manager starts it as
addDriver("discovery", no_device, false): the discovery service, one per
firmware, with no device assignment at all, which finds and claims the
acpi-tables node itself because it is the thing that produces the device
tree. The manager cannot hand it a device — at that point there is nothing
to match against.

So its question is not whether it should hello. It is whether the bootstrap
is exempt from D6, or whether the manager claims the acpi-tables node and
passes it on.

Recorded alongside it: a delivery point that needs no hello is already
implied by the code, since spawnSupervised returns the child's pid and the
manager could transfer immediately after spawn. If that is acceptable,
question 6 largely dissolves — ps2-bus would not need to leave "legacy"
either. The cost is ordering: the child may reach for the device before the
transfer lands, where hello guarantees it cannot because the child is the
one asking. That trade is the real decision, and it is now written down as
such.
2026-08-08 19:14:54 +01:00
Daniel Samson f7151ed577 kernel: the device table has no ceiling; a runaway is charged to whoever caused it
maximum_devices = 64 is gone. It was a guess about someone else's computer,
and because it was shared, one driver's enumeration starved every other —
which is how an AMD Ryzen booted with a working display, no USB and no
storage. The table now grows from the kernel heap. It was always built after
heap.init; nothing ever prevented this except it having been written static
first.

What replaces it is an allowance charged to the registrar, so a driver
looping device_register exhausts its own and every other driver carries on.
It is declared as what it is — a runaway detector, NOT a security boundary.
A quota generous enough never to bite a real machine is still generous
enough to be unpleasant, and it is not trying to be the defence; delegation
is. What this catches is a legitimate driver in a loop, early, attributably,
and without collateral. Reaching 4096 is a bug report, not a tuning request.

The initial block is 8, deliberately small. Sizing it for a typical machine
would mean the growth path never ran on the hardware we test on and only
woke up on someone else's larger machine — the exact failure shape this
track exists to stop. At 8 it grows several times every boot; disabling
growth now fails the suite with the HPET not fitting, which is the Ryzen
failure in miniature.

The comptime coupling assert added earlier fired, and was right to. confined
(one slot per device id) and domains (the IOMMU's own translation pool) were
sized by the same constant only because device ids happened to stop at 64
too. Two unrelated quantities: confined now grows with the device table,
while maximum_domains stays as the hardware's number — both VT-d and AMD-Vi
report how many domains they support, and reading it is phase 4. The assert
existed for exactly this and did its job.

Suite 118/118.
2026-08-08 18:40:20 +01:00
Daniel Samson 5492607a61 docs: D7 blocked, and D8 reordered after D9
D7 would remove zero-resource devices from the kernel. It cannot proceed
because nothing else mints their ids. A USB interface registers with
resource_count = 0, and the id device_register hands back is load-bearing in
three places: the child_added packet's target, the class driver's argv[1],
and the device_token of the usb-transfer WIRE protocol — so the id space is
visible on the wire, not merely internal. usb-xhci-bus records the fourth
constraint itself: the kernel's idempotency is what makes the same port and
interface map back to the same id across a bus restart, which is what stops
a respawned bus spawning duplicate class drivers.

Moving that out means answering who mints the id, how it survives a bus
restart, how it survives a manager restart, and whether device_token changes
meaning. A design step, not a mechanical one.

D8 is reordered to run after D9, correcting the original sequencing. D8's
justification was that the authorisation the per-parent cap stood in for now
exists — but D6 is blocked, so it does not, and deleting the shared cap now
would reopen the exhaustion hole it was written for. D9's per-holder quota
closes that hole independently of authorisation, and closes it better: a
rogue exhausts its own allowance instead of the table everyone shares. Once
the quota exists the per-parent cap is redundant either way.

D9 also no longer depends on D7. Its rationale was that zero-resource
children are the case that sidesteps containment — true, but a per-holder
quota bounds them as well as anything else, because it counts entries per
holder rather than per parent.
2026-08-08 18:26:37 +01:00
Daniel Samson 81dd1e9318 pci: the host bridge arrives by delegation too
pci-bus joins usb-xhci-bus in receiving its device from the manager rather
than claiming the id it found in argv[1]. Its hello moves ahead of the ECAM
mapping, since that is where the bridge now arrives, and its hello was
already mandatory so nothing about its failure behaviour changes.

isDelegated compared whole strings, which silently missed this driver: the
manager records the boot-snapshot match as the bare "pci-bus" and a
devices.csv match as the full "/system/drivers/pci-bus". pci-bus was then
neither claiming nor delegated and died on "ECAM mmio_map failed". It now
matches on the last path component. Reintroducing the whole-string compare
breaks usb-xhci-bus instead of pci-bus — the two spellings swap which driver
loses — so usb-hid is the case that catches it, not pci-scan.

pci-scan asserts the delegation on the initial bring-up AND after the
restart drill, with the device id backreferenced so both must name the same
device. That is what proves the manager re-takes a device when its driver
dies and hands it to the replacement, which is the property the whole
supervision design rests on.

The remaining three claimants are NOT converted, and the plan records why
rather than working around it. ps2-bus and the acpi service never hello at
all, which device-manager.md states deliberately ("legacy drivers ... not
yet required to hello"), so delegating to them means either promoting them
out of legacy or giving the grant a delivery point that is not hello.
virtio-gpu hellos best-effort by design — "standalone bring-up has no
manager" — and delegation would make it mandatory. Both are decisions, not
mechanical steps.

Consequence: D6 is blocked, because device_claim cannot be closed off while
three claimants still depend on it. D7-D9 are unaffected — they concern what
the kernel stores and how its table is sized.

Suite 118/118.
2026-08-08 18:25:16 +01:00
Daniel Samson a2d6056772 usb: the xHCI controller arrives by delegation, not by claiming
The first driver to stop claiming its own hardware. The device manager holds
the controller and transfers it in the hello reply, so its matching becomes
authoritative instead of advisory — until now the driver claimed the id it
found in argv[1], and any process could have claimed the same integer first.

The manager claims before it spawns, so there is no window in which anything
else could take the device, and transfers in onHello using invocation.sender
— the kernel-stamped task id, which cannot be forged by the caller. hello is
synchronous, so the transfer has completed before the reply lands: no gap
between being told yes and holding the thing.

usb-xhci-bus's hello moves from after controller bring-up to before anything
that needs the device, which is the bring-up reorder the design predicted.
It is the first member of an explicit delegated set, so every unconverted
driver keeps claiming exactly as before and the suite stays green; the set
and device_claim both go at D6. D3 and D4 could not be separated and the
plan records why: the moment the manager claims, any driver still calling
device_claim is refused, and D3 applied to nothing changes no behaviour and
cannot be tested.

This step introduced a regression and the incremental conversion is what
caught it. confineDevice runs inside systemDeviceClaim, so a device arriving
by transfer was never confined for its new owner. Three IOMMU+USB cases
failed on the driver's DMA rings going unbound, and two worse consequences
were latent: a manager death would have torn down a domain a live driver was
using, and a driver death would have leaked one. iommu.reassign now moves
the confinement with the device, keeping the domain and its attachment
intact so it never translates through nothing. Converting all five drivers
at once would have produced the same three failures with five suspects.

A log line of mine claimed "holding controller device N" before anything
verified it — it printed even in the failure case, where the driver held
nothing. Reworded to state only what is known there: where the registers
are.

usb-hid asserts the delegation with the device id backreferenced, so the id
delegated and the id the driver ends up with must match. Emptying the
delegated set fails it with "hello acknowledged" then "mmio_map failed".

usb-hub failed once in a full run and has passed six times since (four
isolated, two full) — recorded in the plan as a suspected instance of the
known intermittent AP fault, not dismissed, since this step did shift boot
timing.

Suite 118/118.
2026-08-08 17:55:59 +01:00
Daniel Samson 33376f24ab test: the attacker the device suite never had
The audit's sharpest finding was structural, not a bug: a fully green suite
had hidden six real defects because it contains no attacker. Every device
case asserts that a driver handed its own hardware can drive it. None asked
what a process handed NOTHING can do.

device-authority-test is that process. It is spawned with no device and
asserts what it therefore cannot do: it cannot give away a device another
task holds, nor a free one, because the kernel's rule is that you may give
away what you hold and the device's state is irrelevant to a process holding
nothing. Asserted across every device the machine actually has, so it cannot
pass by accident of which one happened to be free at boot — six on QEMU,
none of them its.

A positive control runs first. device_enumerate works from this process, so
the refusals below it are decisions rather than a syscall path that is
simply broken here; without it, "everything failed" would read identically
to "the assertions are meaningless". A nonexistent device is refused as
NoSuchDevice rather than NotHeld, because a refusal that cannot name its own
rule is what cost a debugging session on the Ryzen.

What it deliberately does not assert, and says so in its header:
device_claim is still first-come-first-served at this point in the run. That
is the hole D6 closes, and the claim half of the invariant joins this
fixture then. Asserting it now would be writing a test that documents the
bug.

Verified to discriminate: removing the holder check flips "every transfer by
a non-holder is refused" while the positive control keeps passing.

Suite 117 -> 118.
2026-08-08 17:23:17 +01:00
Daniel Samson 3111c7c5e6 kernel: device_transfer — you may give away what you hold
The mechanism behind delegation, which device-manager.md named as the step
after hello: the device manager claims what discovery seeded and hands each
device to the driver it matched, so assignment stops being
first-come-first-served.

It is a MOVE, not a copy. A claim is exclusive (driver-model.md, invariant
1), so the giver stops holding the device the instant the receiver starts.
That is why this is a new syscall rather than the M13 capability path, where
a passed handle is shared refcounted — exclusivity cannot be expressed that
way.

The kernel's whole rule is that you may give away what you hold. It has no
notion of which task is the device manager and deliberately gains none: a
binary name inside the kernel is not something that cannot safely live in
user space. A recipient that does not exist is refused, because a device
moved to nobody would be unreachable for the rest of the boot — nothing
un-holds a device but task death.

Three errnos, each naming its own rule: ENODEV no such device, EPERM you do
not hold it, ESRCH no such recipient.

Nothing uses it yet. The five claimants move across one at a time in D4-D5,
so the suite stays green throughout and a regression names the driver that
caused it.

Ten assertions, verified to discriminate: removing the ownership check flips
four of them, including the giveaway that an illegal transfer then blocks
the legitimate claim behind it.

Suite 116 -> 117.
2026-08-08 17:11:59 +01:00
Daniel Samson 547d0ec46b docs: device authority — the how, and the run that deletes the ceilings
device-authority.md is rewritten as an implementation design rather than a
rival to device-manager.md. The what was already settled there in 2026-07:
structure in the manager, authority in the kernel, and delegation as the
step after hello. This is the how, plus the two decisions that paragraph
leaves open.

Decision 1: the manager claims, it is not granted. device-manager.md says
"claims (or is granted)"; claiming wins because the manager runs before any
driver exists and takes the seeded devices unopposed, leaving nothing unheld
to race for. One new call, device_transfer(device_id, task_id), checks only
that the caller holds the device — no names in the kernel, no attestation.
The alternative put a binary path inside the kernel, and the kernel should
hold only what cannot safely live in user space. The residual is stated
rather than hidden: authority rests on the manager claiming first, which
init.csv makes an operator-visible ordering rather than an attacker-
controlled one, and the enforced version arrives with the spawn capability
drivers.md already names as missing.

Decision 2: the kernel stops holding inventory. It reads three things out of
a descriptor — physical ranges, interrupt numbers, one PCI BDF — and stores
the rest only so device_enumerate can hand it back. Devices with no
resources leave the kernel entirely: a USB device conveys no mapping
authority, so there is nothing to enforce. That is also the case which
sidesteps containment, and therefore the reason a shared cap existed.

Decision 3: no shared ceiling. The table becomes dynamic — it is built after
heap.init, so nothing ever prevented it — and the two invented numbers go. A
per-holder quota replaces them, because dynamic storage with no bound moves
the ceiling to the kernel heap, which is shared and fatal rather than
partial. A bound charged to whoever caused it is isolation.

Two earlier drafts of this document are gone: one gave init the root grants,
the other proposed extracting a firmware-framebuffer driver. Both were wrong
and both are recorded as wrong in the run plan's settled list — the
framebuffer is not a device, it is where pixels go until a real display
driver announces itself.

Run 2 is nine steps, ordered so the suite stays green throughout: build and
prove the transfer mechanism, move the five claimants across one at a time,
then the flag day, then the inventory, then the ceilings.
2026-08-08 17:01:25 +01:00
Daniel Samson 5ad42ceac3 docs: only drivers hold devices; services speak protocols
The design still had the display service receiving a device grant, which
keeps the wrong layering and just moves who hands it over. Drivers talk
hardware. Services talk to no hardware at all — they receive device events
and send data and commands over a protocol, which is the OS abstraction
between them.

So the answer to "which services need device grants" is none, and the design
collapses: every holder of a device is a driver, and every driver is spawned
by the device manager, which is what grants it. No exceptions to
accommodate.

display already shows both halves. Its VirtioGpu backend speaks
scanout-protocol over an IPC handle and touches no device — the driver holds
the hardware, the service speaks to it, and it works today. Its Gop backend
calls device.enumerate, device.claim and device.mmioMap: the service
reaching into hardware itself, because the firmware framebuffer has no
driver to talk to. That is the only such case in the tree; every other
claimant is a device-manager child spawned with its device id.

A firmware-framebuffer driver is therefore a prerequisite of this phase
rather than a consequence. It is small — hold the display node, map the
framebuffer, serve the same scanout-protocol virtio-gpu already serves — and
the compositor needs no new path, since not caring which backend is behind
it is what it was designed for.

Two earlier drafts are recorded as wrong in the document: display asking the
device manager for a grant, and before that init minting grants because init
starts display. Both accommodated an exception instead of removing it.
2026-08-08 16:40:37 +01:00
Daniel Samson 8e388e048a docs: one line still had init minting the root grants 2026-08-08 16:35:54 +01:00
Daniel Samson ef793d1320 docs: the device manager is the root of device authority, not init
The design had init minting device grants and handing them on, because init
is what starts display and display claims the framebuffer. That was the
wrong shape. init has nothing to do with devices: it is the first process
and it starts the rest of the system. The device manager owns hardware.

So the kernel mints the root grants to the device manager, and display asks
the manager for its framebuffer like any other driver. The exception that
drove the earlier draft disappears instead of being accommodated — one
authority for hardware, not two.

The kernel recognises the manager by the chain attestation the security
track already settled on: the binary path it stamped itself, AND a
supervisor of PID 1. Path alone is forgeable, because system_spawn is
deliberately ungated and any process may spawn any bundled binary — but a
rogue copy's supervisor is the rogue, and PID 1 is the kernel's own first
process.

Costs restated for the corrected model. display gains a boot-order
dependency on the manager, mitigated by its existing ability to attach a
backend late. The kernel gains one piece of knowledge about one binary,
which is the minimum: authority has to enter somewhere, and the
alternatives are a file parser in the kernel or first-to-ask, which is the
hole again.
2026-08-08 16:35:32 +01:00
Daniel Samson 139c4f624f docs: mechanism, policy and configuration are three things
The phase 2 design called devices.csv "policy". It is not. The CSV is
configuration — declarative data an operator edits. The policy is the
component that decides using it: the device manager matching a device to a
driver, init reading protocol.csv and granting a binding. The kernel holds
mechanism, the check that a grant exists before a mapping is made.

Corrected where the document conflated them, and the three are now tabulated
so the rest of the design can lean on the distinction.

It also sharpens the open question. It was "manifest or not"; it is really
whether the grant list should be configuration from the start, or whether
init should hand the device manager the root grants and let it match by the
devices.csv it already reads. A device-grant file adds a second place an
operator must keep correct, and earns itself only when something needs to
differ from "the manager gets the hardware" — holding a device back for a
test or a bare-metal driver, say.
2026-08-08 16:31:50 +01:00
Daniel Samson 3b24b541b0 docs: phase 2 design — you hold what you were given
device_claim checks that a device exists and is free. That is all. Any
process may claim any unclaimed device, and a claim is what gates mmio_map
and irq_bind — a licence to map physical memory and take interrupts. The
device manager's matching is real but advisory: it spawns a driver with the
device id in argv[1] and nothing binds that decision to the kernel's grant.

maximum_children_per_parent is the visible cost. It exists because a driver
that claimed one device could loop device_register under it, and it is a
poor defence — an attacker burns 16 slots, claims another device, burns 16
more — while reliably refusing a legitimate PCI bus with more than 16
functions. Closing the hole is what retires the constant.

The principles decide the split: matching is policy and stays in the device
manager; enforcing that a driver holds only what it was given is security
and stays in the kernel, and is the whole of what the kernel needs.

The mechanism is decided by an awkward fact. Five of six claimants are
device-manager children spawned with their device id. display is not — init
spawns it, and it finds its framebuffer by enumerating for a display-class
node and claiming whatever it finds. So "record the device named at spawn"
closes the hole for five and breaks the sixth, and the sixth is not an
oddity to special-case: it shows authority must be delegable rather than
welded to the moment of spawn.

So: a device grant is a capability, minted by the kernel to init for the
devices firmware discovery found, delegated by init to the device manager
and to display, and passed by the manager to each driver it spawns. The
cap-passing path already exists and already carries shared memory and DMA
regions. init is already the grantor for /protocol, and protocol.csv already
records init granting the device manager its binding.

Costs named rather than buried: five drivers must hello before claiming
(pci-bus claims first today), and init grows a device role on top of the
protocol registry. A device-grant manifest mirroring protocol.csv would be
the natural symmetry and is deliberately not proposed yet.
2026-08-08 12:58:19 +01:00
Daniel Samson 5d55217212 usb: a slot the controller granted is always handed back
Both device-setup paths issued a successful Enable Slot and then returned
null if allocateDevice failed, without disabling it. A slot the driver
forgets is one the controller never reissues, so each attempt lost one
permanently for the boot. The hub path did it with no log line at all.

Both now release the slot through a shared disableSlot, extracted from
tearDownDevice, and the hub path warns like the root-port path does.

tearDownDevice also now frees the interface list. That allocation arrived
with the previous commit, so an unplug would have leaked it — found while
reading the teardown path for this fix rather than by a test.

No regression test, and it is recorded as open question 5 rather than
implied. After the slot count became the controller's own figure, reaching
this path needs more devices than the controller has slots: QEMU offers four
against sixty-four. What was verified is that the new path RUNS correctly —
pinning tracking to 2 with four devices attached produced "port 6 setup: no
free device slot", the first two devices enumerated normally, and no Disable
Slot error or timeout appeared, which is how disableSlot reports failure.

Suite 116/116.
2026-08-08 12:16:42 +01:00
Daniel Samson 729b40ece7 usb: a device has as many interfaces as it declares
max_interfaces was 4. A composite device — a headset, a webcam with audio, a
dock, a multifunction printer — routinely has more, and the fifth did not
merely go missing. parseConfiguration's cap branch had no `else`, so when the
count was reached `current` kept pointing at interface 3 and the fifth
interface's endpoint descriptors were appended to interface 3's array. A
class driver bound to interface 3 could then be handed an endpoint belonging
to something else entirely, and subscribe or bulk-transfer on it. The
alternate-setting arm one line above cleared `current` correctly, which is
what the cap branch should have done.

Interfaces are now counted from the block in a first pass and allocated to
exactly that number, so the ceiling is bNumInterfaces' u8 — the USB
specification's. The missing `else` is added too, though after this the bug
is unreachable by construction: interface_count cannot reach interfaces.len
mid-parse when the list was sized from the same walk.

max_configured_endpoints was max_interfaces * max_endpoints_per_interface =
16, a derived guess that moved whenever either input moved. It is now 31,
which is the xHCI specification's own limit: a Device Context holds a slot
context plus at most 31 endpoint contexts, because the Context Entries field
addressing them is 5 bits.

max_endpoints_per_interface stays at 4 with its reason recorded — the
usb-transfer wire protocol reports exactly max_reported_endpoints (4) per
interface, so widening it alone would change nothing a class driver sees.
Lifting it is a protocol change.

No direct test, and that is written down as open question 5 rather than
glossed. The parser is pure and wants a host unit test, but
usb-xhci-library.zig imports memory, mmio and time so it cannot be a
standalone test root, and QEMU offers nothing that reaches the path — the
largest device available is usb-audio,multi=on at 2 interfaces and 211
bytes. The alternate-setting path that shares the same `current = null`
logic is exercised by that device.

Suite 116/116.
2026-08-08 12:05:33 +01:00
Daniel Samson 6328823ef1 usb: a configuration block is as long as the device says it is
The driver read the first 512 bytes of a configuration block into a fixed
buffer and parsed those. The block's length is the device's own choice
(wTotalLength, a u16), so anything larger was silently cut: interfaces past
the cut did not exist as far as the host was concerned, while the
SET_CONFIGURATION that follows still configured the device for all of them.
A headset is 500-900 bytes, a UVC webcam 1-3 KB, a multifunction printer
600+.

Now allocated at the declared length, so the ceiling is the field's u16 —
the specification's number rather than one of ours. A block shorter than its
own 9-byte header is refused rather than trusted.

The bring-up line reports the declared length and the bytes actually read,
so a truncation can never again be invisible, and usb-large-descriptor
asserts they match with a backreference.

That case has an honest limit, recorded in its comment: QEMU cannot produce
a block over 512 bytes. The boot keyboard, mouse and stick are 34-44, and
the largest device available is usb-audio in multi-channel mode at 211 —
which is exactly why the suite never caught this, and why it cannot now
reproduce the original trigger. What it does catch is the class: any clamp
below the attached device's block fails it, verified by pinning the buffer
to 128 and watching "config block 211 bytes, read 128" turn the case red.

Suite 115 -> 116.
2026-08-08 11:53:01 +01:00
Daniel Samson 69fbef40c0 usb: the controller says how many device slots it has
max_devices was 8, with the comment "QEMU presents a handful; a fuller
machine would grow this" — a number chosen against the test rig, waiting for
a real machine, which is the pattern docs/bounds-track-plan.md exists to
stop.

The driver already knew the true figure. It reads HCSPARAMS1.MaxSlots at
bring-up and writes it straight into op_config, so every slot the controller
offers has always been *enabled*; only the array tracking them was 8. QEMU's
xHCI reports 64, so seven eighths of the controller was live and invisible,
and the ninth device — a keyboard, mouse, webcam, headset, hub and two
sticks reach that without trying — disappeared on a hub-attached path that
logs nothing at all.

The array becomes a slice allocated from max_slots at bring-up. A controller
claiming zero slots cannot address anything, so that is now a dead
controller rather than an empty allocation failing mysteriously later. The
Device Context Base Address Array is a page, 511 usable entries, so it
already covered the 255-slot maximum.

The bring-up line reports both numbers, and usb-hid asserts they are equal
with a backreference rather than a magic number, so the test cannot drift
from the hardware. Pinning tracking back to 8 fails it: "64 slots,
tracking 8".

Suite 115/115.
2026-08-08 11:39:49 +01:00