The storage layer gains its policy home (storage-architecture.md): a new
system/services/volume-manager, spawned by init, that acquires the mass-
storage block channel through the device manager (the same lineage a
filesystem uses), reads block 0, and parses the first volume out of it. The
partition-table walk that lived in the FAT engine moves here, above the
driver where it belongs (partition.zig, host-tested: MBR entry, bare-FAT,
no-signature). Identity is the MBR disk signature + partition index — the
weak rung of the ladder; GPT GUID and FAT serial refine identityOf without
changing shape.
This increment is discovery + probe + log only, additive: the FAT service
still acquires its own volume, so nothing changes for it. Confining each
filesystem to its partition and spawning one per volume (the flip) lands
next, keeping fat working throughout.
Grants + wiring: init.csv spawns it after the device manager; protocol.csv
grants bind volume-manager + open device-manager. Verified: volume-probe
asserts the parse (bare-FAT volume at lba 0), neutral 10/10 across storage,
restart, display, logging, confinement — the volume manager now runs in
every boot and disturbs nothing.
fat was one binary doing four jobs; the three that are not FAT-specific move
to library/kernel/file-system-harness, a Server(comptime Engine) generic over
the engine type: the badge-scoped open-node table, the nine vfs handlers, the
not-mounted politeness, the exit sweep, mount registration, and durable-on-
close. A filesystem is now an engine plus a main that hands the harness a
mounted volume; a second engine reuses the harness wholesale.
Placement note: the plan said library/file-system, but the harness is a
specialization of `service` (its sibling) and needs nothing from the device
domain, so it lives beside service in library/kernel and stays block-free —
durability rides a caller closure (Volume.flush), no backwards kernel->device
dependency, no new-domain scaffolding. The engine type is inferred from
resolve()'s return, so engine.zig is untouched (its Node stays module-scope).
fat keeps only its FAT-specific bring-up (acquireVolume, DMA, engine.mount,
the attach/detach round trip) and the three mount prefixes as data. Behavior-
neutral: 13/13 across the fat/vfs/logger/IOMMU surface, nothing observable
changed. This lands first so every later phase touches the harness once.
The kernel was always symmetric (dma_bind 51 / dma_unbind 52); the two
protocols that forward an attachment up the stack were one-way, so a live
client could grant a device reach into its buffer but never revoke it while
alive — exactly the one-way lifecycle the storage architecture's enforcement
section forbids. Death stays the mechanical backstop; detach is the living
process's path.
Both verbs are appended, so every existing number holds. The shape mirrors
attach precisely: the same region capability rides the cap slot again — the
kernel matches the region, so no layer retains anything between the calls
(the bus never kept the handle; now it never needs to).
fat's bring-up does attach -> detach -> attach, exercising both verbs
through the whole chain (fat -> storage -> bus -> kernel) on every boot: a
broken detach fails every fat case instead of lying dormant until the first
buffer replacement. Honest scope: the round trip proves the plumbing; unbind
semantics are the kernel iommu tests' (map/unmap/translationOf); the full
composition (detach then DMA faults) is a future iommu-fault extension.
The hot-unplug path (onChildRemoved, a report from a live bus) cleared the
child but left the bound class driver: a process blocked on reports that
will never come, whose stale entry made the matcher's dedupe refuse the
respawn when the device was plugged back in — the same wall the restart
zombie hit, one path over. Unplug now reaps exactly like reporter death.
The usb-hub-unplug case grows the replug: device_del the hub keyboard, then
device_add it back (qmp_sequence); the ordered tail — child removed, reaping,
delegated, ok — can only be satisfied by the second generation, since every
boot keyboard's ok precedes the unplug. Discrimination: without the reap the
replug never rebinds and the case times out (verified by stash run). Hub
family, restart drill, and the two-controller proof all green (8/8).
P3 of docs/establishment-planes-plan.md — the restart-zombie fix. A class
driver cannot observe its provider's death: an HID driver blocks on
interrupt reports that will simply never come, and storage answers its
callers with refusals forever. Worse, the dead generation's still-used
entries made the matcher's dedupe refuse the respawn when the restarted bus
re-reported — the subtree was a permanent zombie, which is the exact
opposite of the restart-a-driver-live goal the driver model exists for.
pruneChildrenOf now reaps: each pruned child's bound driver is killed and
its entry cleared (the exit notification finds no entry, so the death is
never double-counted; its device returns by the loan rule; its stored
endpoint handle is closed). The re-report then spawns a fresh generation
whose hellos fetch the successor's channel.
The usb-report drill now asserts the subtree WORKS after the restart: the
respawned storage opens its device on the NEW bus instance and reads block
0. Discrimination: against the pre-reap manager the drill fails — no reap
line, no post-restart respawn (the survivors were zombies), verified by a
stash run. Note the scenario boots no input service, so the HID drivers of
BOTH generations exit after their input lookup times out — storage is the
functional proof.
P2 of docs/establishment-planes-plan.md. usb-storage serves nameless — one
process per stick cannot share an exclusive bind, and a second stick used to
die silently on -EBUSY before ever helloing. Its one hello now moves both
directions at once: the block-serving endpoint up, its controller channel
down. fat finds its volume through the manager — a new `.consumer` role asks
for the channel of the driver BOUND TO a device (distinct from the device's
reporter), found by enumerating the tree for the mass-storage identity.
fat stays single-volume; the boot-volume-by-content choice is M21.
The conversion immediately caught a live truncation of exactly the audit's
shape: ChildEntry grew to 32 bytes, one enumerate reply holds ~7, and a real
tree carries a dozen ACPI nodes before the first USB child — the storage
entry silently never fit (the protocol comment already said "paging joins
the protocol if a tree ever outgrows one packet"). enumerate is now paged:
Header.target is the start cursor, a short page is the end; device-list's
page-0 read is unchanged.
Grant rows move with the code: the block bind and fat's block open die, fat
gains open device-manager. Gate: 18 cases green including the registry
trio, device-list, and both IOMMU storage variants.
P0 of docs/establishment-planes-plan.md; no behavior changes yet, nothing
sets the new flag or sends a hello capability.
- service.run gains a reply-capability out-slot (replyWithCapability), the
registry idiom init already uses, lifted into the harness; null stays the
untouched common path. Subscribers gains claimArrival() so a provider
handler can keep a turn capability through the same flag the reserved
subscribe uses.
- Hello wire struct: the padding byte becomes wants_channel — old callers
wire-compatibly say 0, and the manager nominates a reply capability ONLY
when asked, because a capability sent to a caller that never reads one is
a leaked slot in that caller's table.
- driver.helloExchange: one handshake can hand a serving endpoint up and
receive the device's provider channel down (usb-storage will need both at
once). No channel in the reply is retryable, never a verdict.
- The manager stores each instance's serving endpoint on its Driver entry,
routes consumer hellos by lineage (child -> reporter -> endpoint), replaces
on re-hello, and closes the stale handle on death - the 32-slot table is
the bound that makes forgetting this boot-fatal.
The name list and its predicate existed so drivers could move to delegation
one at a time with the suite green throughout. Every driver is delegated
now, so the manager simply hands over whatever device a driver was assigned.
Deleting it caught a real consequence: crash-test finally got delegated too,
and it was still claiming its device — so it got AlreadyClaimed because it
already held it, exited, and the restart drill had nothing to restart. Its
own comment named what the case was really checking: "the respawn only
reaches this line because the kernel released the previous instance's claim
at death". That property still holds, by a different mechanism — the device
reverts to the manager on death and is handed to the replacement, which is
the same guarantee without the race it used to rely on.
All four delegation paths verified: the xHCI controller, the PCI bridge, the
PS/2 two-node singleton, and virtio-gpu's restart re-attach.
Run 3 complete. Suite 118/118.
A device nobody holds can be claimed by anyone, so the manager now takes
every firmware-discovered device that carries mappable resources, whether or
not a driver wants it. The real gap was the HPET: an MMIO window, an IRQ, no
user-space driver, and there for the taking. Held by the manager it is
inert; unheld it was a way into physical memory.
Two deliberate exclusions. The loader's framebuffer, which the compositor
claims and which the manager must not take because it starts first. And
anything with no resources, which grants nothing worth holding.
Scope is the boot snapshot. A device reported later and matched to no driver
stays claimable — pci-cap-test and iommu-fault-test both reach an unmatched
NIC that way, so narrowing it is a separate change with those fixtures in
scope. Recorded in the plan rather than left implied.
The attacker fixture gains the assertion deferred since D2: after the system
settles, nothing with resources may be taken.
That assertion defeated itself twice before it worked, and both failures are
worth remembering. First it swept at 0.029 while the manager did not bind
its protocol until 0.047, so it reported a hole that closed a millisecond
later. The retry loop that "fixed" that was worse: the first pass TAKES the
device, so the second finds it unavailable because this process now holds
it, and concludes all is well — it passed with the manager's claiming
removed entirely. It now settles once and sweeps once, and fails when the
claiming is removed.
Suite 118/118.
The last claimant. The kernel seeds the acpi-tables node, so it sits in the
same boot snapshot the manager already scans to find the PCI host bridge —
there was never a bootstrap problem, only a lookup nobody had written. The
manager claims it and names it in the spawn; the service stops claiming.
Every driver in the system now receives its hardware rather than taking it.
Two failures on the way, both mine. addDriver puts the device id in argv[1],
and the acpi service read argv[1] as a self-verify device-count floor — so
handed device 7 it decided it was in test mode, printed "acpi-parse: ok",
and never reported a device. The test argument is now floor:N, which a bare
id cannot be mistaken for.
And acpi-parse spawns the service directly rather than through the manager,
so nothing handed it the node. That test now claims and transfers it exactly
as the manager does, which is the right shape: the test plays the manager's
role instead of the service reaching for hardware.
device_claim now has two callers left: the manager, which is the acquirer
and should have it, and the display service's GOP path. That is recorded as
question 10 — the framebuffer is not a device, so the answer is likely that
it leaves the device table rather than being exempted from its rules.
Suite 118/118.
The 8042 is a single controller described by two ACPI nodes — PNP0303
carries the 0x60/0x64 ports, PNP0F13 is the mouse — so it cannot be split
across two processes without them fighting over the same registers. That is
why ps2-bus is a singleton, and why it used to find and claim both nodes
itself.
The manager now gives it every matching node instead. The keyboard node
rides the spawn, because it holds the ports and is needed immediately; the
mouse node is transferred to the already-running instance. Late arrival is
safe here and the ordering is natural rather than lucky: the mouse is not
touched until after the controller handshakes and identify. Measured, the
handover lands at 0.336 and the driver reaches the mouse at 0.456.
Because the count is however many matched, a machine with no PS/2 ports or
only one works without a special case — which matters, since the bus is
mostly emulated now and machines vary.
ps2-bus claims nothing. irq_bind on the mouse node is the proof it holds it:
that call is ownership-gated, so a failure means the handover did not land
rather than a hardware fault, and the log says so.
acpi-ps2 asserts both delegations with the spawned device backreferenced, so
the node that rides the spawn must be the one the driver was spawned for.
Disabling the second delegation fails it.
Two things worth recording. The first discrimination patch was not valid Zig,
so nothing ran and a stale binary reported a pass — checked the build before
believing it. And with the second delegation disabled, acpi-ps2 fails while
input still passes: the mouse works without its IRQ binding, so exactly one
case covers that path.
Suite 118/118.
Third driver converted. It no longer claims the id from argv[1] — the
manager holds the device and names it in the call that creates the process,
so it is held before the driver's first instruction.
display-reattach is the case that matters here: it kills the driver and
watches the compositor re-attach to the fresh scanout. It passes, so the
restart path survives the fused grant — the manager re-takes the device when
the driver dies and hands it to the replacement.
ps2-bus and discovery are NOT converted, and the reason is recorded as open
question 9 rather than worked around. Both need a device nobody assigned
them. ps2-bus ignores its argv[1] entirely: it finds the controller by
walking the table for PNP0303, then claims a second device, the PNP0F13
mouse node, which it also finds itself — so it holds two devices and was
assigned at most one, while system_spawn carries one. discovery claims the
acpi-tables node it locates itself, because it is what produces the device
tree and there is nothing to assign at that point.
One thing worth checking before designing an answer: devices.csv maps both
PS/2 hardware ids to ps2-bus, so the manager may already be spawning two
instances where the driver expects one. If so the fix is smaller than it
looks.
D6 stays blocked — closing device_claim with these two still depending on it
would stop the machine booting.
Suite 118/118.
Delegation moves out of onHello and into the spawn itself. The manager holds
the hardware and names it in the call that creates the driver; the kernel
checks the device is the caller's to give, then hands it over as part of
making the child.
The reason is the window. A transfer after spawning always leaves an
interval in which the child is running and does not yet hold its device. It
would close on QEMU every time and open occasionally on a machine with
different core counts and timing — the exact failure shape this track exists
to delete, and not one worth introducing while removing the others. Fused
into the spawn there is no interval: the child does not exist until it holds
the device.
Ownership is checked BEFORE the child is created, so a refusal leaves
nothing running rather than a driver without the hardware it was spawned
for. The IOMMU confinement moves with the device, as it does on the transfer
path. systemCall6 is added for the sixth argument; r9 was free, and abi
gains a no_device sentinel matching the protocol's.
No driver had to change to receive a device, which is what makes this
better than requiring every driver to hello: ps2-bus keeps its legacy
status, and discovery — which has no assignment at all, since it is what
produces the device tree — is unaffected.
The attacker fixture now tries the spawn as a back door: name someone else's
device, and both the spawn and any child must be refused. Verifying that
assertion exposed a bug in the fixture itself. The kernel case's pass marker
was "device-authority: ok", which matches the FIRST per-assertion line, so
its wait loop exited before any failure was printed — the case would have
passed with failures in it, and had been able to since D2. The verdict lines
now carry a distinct VERDICT prefix, and with the ownership check removed
the case genuinely fails. A green test that cannot go red is worse than no
test.
Suite 118/118.
pci-bus joins usb-xhci-bus in receiving its device from the manager rather
than claiming the id it found in argv[1]. Its hello moves ahead of the ECAM
mapping, since that is where the bridge now arrives, and its hello was
already mandatory so nothing about its failure behaviour changes.
isDelegated compared whole strings, which silently missed this driver: the
manager records the boot-snapshot match as the bare "pci-bus" and a
devices.csv match as the full "/system/drivers/pci-bus". pci-bus was then
neither claiming nor delegated and died on "ECAM mmio_map failed". It now
matches on the last path component. Reintroducing the whole-string compare
breaks usb-xhci-bus instead of pci-bus — the two spellings swap which driver
loses — so usb-hid is the case that catches it, not pci-scan.
pci-scan asserts the delegation on the initial bring-up AND after the
restart drill, with the device id backreferenced so both must name the same
device. That is what proves the manager re-takes a device when its driver
dies and hands it to the replacement, which is the property the whole
supervision design rests on.
The remaining three claimants are NOT converted, and the plan records why
rather than working around it. ps2-bus and the acpi service never hello at
all, which device-manager.md states deliberately ("legacy drivers ... not
yet required to hello"), so delegating to them means either promoting them
out of legacy or giving the grant a delivery point that is not hello.
virtio-gpu hellos best-effort by design — "standalone bring-up has no
manager" — and delegation would make it mandatory. Both are decisions, not
mechanical steps.
Consequence: D6 is blocked, because device_claim cannot be closed off while
three claimants still depend on it. D7-D9 are unaffected — they concern what
the kernel stores and how its table is sized.
Suite 118/118.
The first driver to stop claiming its own hardware. The device manager holds
the controller and transfers it in the hello reply, so its matching becomes
authoritative instead of advisory — until now the driver claimed the id it
found in argv[1], and any process could have claimed the same integer first.
The manager claims before it spawns, so there is no window in which anything
else could take the device, and transfers in onHello using invocation.sender
— the kernel-stamped task id, which cannot be forged by the caller. hello is
synchronous, so the transfer has completed before the reply lands: no gap
between being told yes and holding the thing.
usb-xhci-bus's hello moves from after controller bring-up to before anything
that needs the device, which is the bring-up reorder the design predicted.
It is the first member of an explicit delegated set, so every unconverted
driver keeps claiming exactly as before and the suite stays green; the set
and device_claim both go at D6. D3 and D4 could not be separated and the
plan records why: the moment the manager claims, any driver still calling
device_claim is refused, and D3 applied to nothing changes no behaviour and
cannot be tested.
This step introduced a regression and the incremental conversion is what
caught it. confineDevice runs inside systemDeviceClaim, so a device arriving
by transfer was never confined for its new owner. Three IOMMU+USB cases
failed on the driver's DMA rings going unbound, and two worse consequences
were latent: a manager death would have torn down a domain a live driver was
using, and a driver death would have leaked one. iommu.reassign now moves
the confinement with the device, keeping the domain and its attachment
intact so it never translates through nothing. Converting all five drivers
at once would have produced the same three failures with five suspects.
A log line of mine claimed "holding controller device N" before anything
verified it — it printed even in the failure case, where the driver held
nothing. Reworded to state only what is known there: where the registers
are.
usb-hid asserts the delegation with the device id backreferenced, so the id
delegated and the id the driver ends up with must match. Emptying the
delegated set fails it with "hello acknowledged" then "mmio_map failed".
usb-hub failed once in a full run and has passed six times since (four
isolated, two full) — recorded in the plan as a suspected instance of the
known intermittent AP fault, not dismissed, since this step did shift boot
timing.
Suite 118/118.
An AMD Ryzen booted to a working compositor with no USB and no storage,
and the log said only "register refused". A tree-wide audit of every
compile-time ceiling followed: 235 of them, 139 on quantities the machine
or a file decides rather than us, 5 documented anywhere, 171 silent when
reached. docs/fixed-bounds-audit.md has the inventory.
Errno attribution. The errno space was split between the kernel and the
envelope, free to drift; it is now one list in system/abi.zig, restated on
both sides, with a comptime check in library/device/driver where the two
halves are visible. device_register's six refusals and device_claim's three
are distinct codes, so a bus driver can say which rule stopped it, and
BadParent splits into NoSuchParent and NotYourParent. pci-bus reconciles
found against registered instead of counting refused functions as found.
Idempotency ordering. The child cap was checked before the identity match,
so a restarted bus was refused its own devices — the supervision restart the
system leans on ratcheted toward a degraded machine. A re-registration
consumes no slot and is now admitted first.
IOMMU fail-closed. confineDevice returned success for a device id past the
confinement table, leaving the device outside every domain while the caller
believed it confined — unreachable only while ids stop at 64, which both the
inventory move and a hardware-reported domain count would change. It refuses
now, and the coupling to the broker's device cap is a comptime assert rather
than a sentence in a comment.
PCI apertures. The bridge's MMIO apertures are derived from the holes in the
firmware memory map, and the derivation copied sub-4 GiB entries into a
fixed [64] array and skipped the rest. A skipped region is not merely lost:
the gap finder concludes it is free, so a real machine's 60-200 entry map
yields an aperture over live RAM, and containment then admits a child BAR
covering kernel memory. Rewritten to walk the map in place, with the hole
finder extracted as a pure function and driven by a synthetic 100-entry map
in a new test case. Both new tests were verified to fail on the old code.
parameters.zig gains the rationale it was missing and loses a stale sentence
pointing at the wrong file; vdso.md documents the errno space, including
EPEER, which had no written meaning anywhere.
docs/os-development/bounds.md is how a ceiling is declared from here.
docs/bounds-track-plan.md is the plan to remove the ones we invented.
Suite 114 -> 115.
Three services had each written the same thing and got it three different
ways: input polled the process list to notice a dead subscriber, and only
when someone else subscribed; the power service never noticed at all; the
device manager noticed drivers but not subscribers. The harness owns the
table now, driven by the events a protocol declares — it registers on the
reserved verb, frames each event once, posts to everyone interested without
waiting on any of them, and reclaims a slot when the kernel says its owner
died. Interest masks moved to the envelope, so a subscriber that wants only
mice asks the same way everywhere.
Two consequences the plan had not foreseen. The device manager now hears a
supervised child's death twice, once as its supervisor and once as a
subscriber, so restart backoff counted every crash twice and gave up after
half as many; it retires the id before counting. And the kernel's published
exit table had eight slots for what is now six subscriptions in a plain
boot, so it holds sixteen.
The other half is a hole the design named early and left standing: a
backend handed out a small integer and then honoured it from anyone. A
process that guessed a file's node id read another client's file; a display
layer had no owner at all, so any client could reconfigure or destroy any
layer; a USB device token was never checked against the client that opened
it. Each is now bound to the task that opened it, and a wrong owner gets
exactly what an unknown id gets — the refusal must not become the oracle
the identical answers elsewhere were designed to remove. Closing a file
changed with it: it used to succeed unconditionally, which would have told
a caller which ids existed.
Suite 111/111, with a new case in which one process holds a file and a
layer, hands both ids to a second process, and finds them untouched after
that process has tried everything with them.
These were the awkward ones. Each began with an operation packed into a
single byte — two of them with a version wedged in beside it — so there was
no wrapping them: the layouts had to be rebuilt. The device manager's own
enumerate and subscribe become the reserved verbs that mean the same thing
everywhere, its replies lose three status structs the envelope already
carries, and a device id becomes the packet's target. Power drops the
version it repeated on every request, because describe is the handshake,
and stops claiming a 64-byte ceiling it never needed for calls. USB moves a
control transfer's data to the packet tail in both directions, which makes
the status length the transferred length and retires a field that had been
saying the same thing twice.
The danger in this one was not the protocols but their readers. Init
recognised a power button by two bytes at the head of a message, the ACPI
service dispatched on the first byte, the xHCI driver read its operation
with a raw integer load, and the HID drivers reinterpreted a report
wholesale — none of which would have failed to compile once the layouts
moved. They would simply have stopped: no shutdown on the power button, no
reports from the keyboard. Every one of them now reads through the
generated types, and the shutdown gate that answers only a subscriber is
the same code it was.
Two sizes were decided by measuring rather than assuming. The child-added
message is both a request and the event broadcast to subscribers, and
alignment rounds it to 48 bytes, which puts its packet exactly on the
64-byte push floor — a test pins that, because a field added carelessly
would now overflow it. The interrupt report gives up eight bytes of inline
room to make space for the header; the two drivers that produce reports
send eight and four.
Suite 110/110.
The folded header stops being a rule in a document and becomes the layout
on the wire. Verbs number from sixteen, leaving describe, enumerate,
subscribe and unsubscribe reserved and answered the same way by every
provider — none of them writes a line to do it. What each protocol used to
carry in a field of its own now travels in the header: a vfs node and a
display layer are the packet's target, and a reply opens with a status the
envelope stamps rather than one each protocol spelled for itself.
Display gains the most. One forty-byte request had served eleven verbs, so
attach_scanout smuggled stride through x, refresh through y and format
through colour, and every coordinate crossed as a bitcast. Per-operation
structs end all three: the fields have their own names and their own signs,
and the tile payload grows to 224 bytes because the prefix shrank. Scanout
loses a message maximum of 64 it had no business declaring — it answers
calls, and the floor for a call is 256 — and virtio-gpu stops hard-coding
that number at its harness.
Two changes are semantic rather than notational. A directory now ends at an
entry with no name, because the fixed part of a reply always travels and a
zero-length reply no longer exists to mean anything. And input joins the
service harness, the last loop in the tree that answered no ping and heard
no terminate; its subscriber table, its pruning and its fan-out are the
same code, and a shutdown now asks it to stop instead of killing it.
A new conformance case reads the registry's own listing and asks every
protocol it finds for its name, its version and its verb count, then offers
a verb nobody defines and requires -ENOSYS — the envelope's promise,
checked against providers rather than against itself. What it cannot reach
in that boot it names on the serial line instead of passing quietly.
Suite 110/110.
The registry consults the open rows it has been parsing since P2, so
reaching a contract now takes a grant as well as a binding. A caller
without one is answered exactly as it would be for a name nobody ever
bound: same status, same empty reply, same absent capability, byte for
byte, and no log line on either path — klog_read is ungated, so a line on
one and not the other would be the oracle the design set out to remove.
Refusal and absence being one answer is what lets a supervisor later
narrow, fake or park a child's namespace without the child learning what
it was denied.
The manifest gains a third permission for a shape the plan did not
foresee: attestation is one hop, but the driver tree is three deep — the
PS/2 keyboard and mouse are spawned by ps2-bus, which the device manager
spawned — so no row could name them and PS/2 input would simply stop.
A supervise grant lets a delegate vouch for what its children *reach*,
never for what they claim; the bind path is untouched, and the laundering
deputy is still refused.
The review found the receive side of a rule this track had already
written down. Every process holds a sendable handle to the registrar —
resolve installs one for anyone who asks — and ipc_reply_wait never asked
who owned the endpoint, so a stranger could dequeue there: take the
provider endpoints riding bind requests, and answer other clients' opens
in the registrar's name. Receiving is the owner's privilege, like binding
a signal or a timer; sending remains anyone's.
Suite 109/109.
A protocol is reached by name now, not by a compile-time integer. Init is
PID 1 and already knows which binary it started, so init serves /protocol
as a vfs backend: bind claims a contract with the provider's endpoint
attached, open answers with that endpoint as the reply's capability, and
readdir lists what is bound with the task and binary behind it. The kernel
reserves the prefix — nothing may mount over it, under it, or unmount it —
and ServiceId, ipc_register and ipc_lookup are gone, their syscall numbers
left vacant.
A bind is authorized by who the caller *is*: the kernel-stamped binary
together with the supervising task's identity, matched against
/system/configuration/protocol.csv. Identity, not spelling — spawn is
ungated, so an attacker can run any bundled binary, and a name-only rule
would have let it launder grants through an init of its own making. A name
a live process holds is refused to everyone else; a dead one's is released.
Three review rounds against a hostile ring-3 process found what 108 green
tests could not, because the suite contains no attacker. Publishing init's
supervision endpoint as the registry put PID 1's mailbox in every process's
hands, where two forged bytes reached the shutdown path: privileged traffic
is now believed only from the task that holds the contract it speaks for.
A capability arriving on a request outlived every path that ignored it,
one handle per call until the table was full — in init, and in the harness
ten services share — so the arriving capability is owned by the turn and
released unless a handler says otherwise. And the kernel let anyone holding
an endpoint handle aim signals, timers, exit notices and interrupts at it:
binding now requires having created it.
Suite 108/108. The new protocol-registry case asserts eleven properties,
each one an attack that must fail.
/etc/init.csv and /etc/devices.csv become /system/configuration/*.csv (the
repo's etc/ moves to system/configuration/, mirroring the runtime tree),
/var/log becomes /system/logs, and /mnt/usb becomes /volumes/usb. The
kernel VFS gains a carve-out so FAT may serve exactly /system/configuration
and /system/logs beneath the initrd-backed /system while /system and /test
themselves stay unshadowable; FAT's single /var mount splits into those two
rewritten mounts. The kvfs readdir check learns /system's third child and
the ramdisk spawn sweep skips the configuration tree.
Suite 106/106.
The module-to-domain table in build-support duplicated what each
domain's build.zig already states with its addModule exports. userBinary
now resolves each named import by searching the packages the binary
declared in its own build.zig.zon (b.available_deps), which also makes
the zon the literal include path: an import can only be satisfied by a
domain the binary claims, and naming a module whose domain is missing
fails the build graph with the domain to declare. build-support is down
to the recipe alone. All build variants green; manifest unchanged.
display-client and input-client (files and module names), so a service,
its wire contract, and its client never share a name: `display` the
service, `display-protocol` the contract, `display-client` a program's
view of it. The nine consumers' imports and their packages' declared
lists follow (regenerated from the source scan); build-support's
module_homes table carries the new names.
Every binary's build.zig now names precisely the modules its source
imports (derived by scanning each artifact's sources, transitively
through same-directory files), and its zon carries only the domains
those come from — kernel stays implicit (the root shim + link script
live there). build-support's userBinary resolves each name through one
module-to-domain table (module_homes); Domains/domains()/defaultImports
and the raw recipe entry point are deleted. An undeclared @import is a
compile error (verified: injecting @import("xkeyboard-config") into
logger fails with 'no module named ... available within module
program'), and e.g. xkeyboard-config now appears in exactly two
manifests — the two keyboard drivers. Availability never bloated the
emitted binaries (Zig compiles only what a program imports); this makes
the declared interfaces honest. Production and -Dtest-case manifests
byte-identical; all build variants and standalone package builds
green.
init, fat, display, display-demo, device-manager, input, logger, and
the two discovery fillers (acpi, fdt — each exporting an artifact named
"discovery"; the root -Ddiscovery picks which ships) convert to binary
packages on the pci-bus template. init's serial heartbeat flag rides
the dependency options (the root forwards its -Dserial). fat's and
display's unit tests move into their packages and the root aggregate
delegates to them. Boot-image file list unchanged.
Replaces L2's interim DMA pool (every buffer reachable by every claimed
device) with true per-grant confinement: a device reaches only buffers
whose capability was delegated to its driver.
Kernel:
- DmaRegionObject (handle kind 2): a delegation token naming a dma_alloc'd
region, passable across processes on the IPC cap slot like an endpoint
or shared-memory object. Frames stay owned by the allocating address
space (freed on dma_free/teardown as before); the token carries a `dead`
flag so a stale downstream handle can no longer bind a freed region.
- dma_alloc gains the dma_shareable flag: it returns a capability handle
in r8 and every region is tracked in a registry. A task's own regions
auto-bind into the devices it claims (its rings just work); foreign
buffers are bound explicitly.
- dma_bind / dma_unbind / handle_close syscalls (51-53). dma_bind maps a
held region (or shared-memory) capability into a claimed device's domain;
it is idempotent. handle_close reclaims a table slot (raised 16 -> 32).
- dma_free and task death unmap a region from every domain and invalidate
BEFORE its frames return to the allocator — the stale-IOTLB use-after-
free window, closed structurally.
Protocols (flag-day): block gains attach, usb-transfer gains dma_attach —
each carries a region capability on the cap slot. fat allocates its bounce
buffer shareable and attaches it; usb-storage allocates its transport
buffers shareable, attaches them to the controller, and forwards fat's
capability downstream; usb-xhci-bus binds and closes; virtio-gpu binds its
shared scanout surface. The physical addresses on the wire are unchanged
(identity IOVA), so no register-programming code moved.
Cross-process DMA (fat -> usb-storage -> xHC) now flows only through
delegated capabilities. iommu-usb-storage / iommu-usb-hid / iommu-fault
all green under per-grant enforcement; 104/104 overall (fail-open paths
unchanged).
Each bus's discovery line now prints the device's would-be /etc/devices.csv row
(bus, base, class, prog_if, vendor, device, subsystem / hid) in uppercase hex,
followed by the human-readable names — so a row for a new driver reads straight
off the boot log, for pci, usb, and acpi alike.
- pci-bus: logFunction moved to registerAndReport (where vendor/device/subsystem
are read from config space) and reformatted to columns + names; subsystem
prints '*' when the function has none. Read unclaimed via the bridge ECAM, so
no per-function claim is needed.
- usb-xhci-bus, acpi: the same column framing on their existing discovery lines.
- qemu_test.py: the acpi-ps2 and acpi-report regexes updated to the new ACPI
format (both verified passing in QEMU).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KJqSiLLchDUUCoXn5jsiwd
Replace init's hardcoded boot_services array with an authoritative,
human-readable service list read at boot, mirroring /etc/devices.csv. Each row
is a service binary path followed by its argv; startup order is file order,
shutdown the reverse. There is no hardcoded fallback — a missing file starts
nothing (the no-ramdisk isolation behavior).
- library/csv: shared CSV helpers (comment strip, field iteration) with unit
tests; device-registry is refactored onto them so both /etc/*.csv files parse
through one place.
- init reads /etc/init.csv into fixed-max static tables (the same pattern as the
device registry) and passes each row's argv straight to spawnSupervised. This
also makes boot-time modes (e.g. device-manager test-usb-restart) expressible
as data rather than hardcoded.
- Diagnose mode (-Ddiagnose omits the display stack) becomes build-time file
selection between etc/init.csv and etc/init-diagnose.csv, so init carries no
comptime service logic; the diagnose build option is dropped from init.
- build.zig: csv module wired; /etc/init.csv bundled into the initrd; the
device-registry tests move to a dedicated block since they now import csv.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KJqSiLLchDUUCoXn5jsiwd
Replace the three hardcoded switch tables (pciDriverForIdentity, hidDriverFor,
usbDriverForIdentity) with an authoritative, human-readable device registry the
manager reads at boot. Matching is most-specific-wins across
base/subclass/prog_if/vendor/device/subsystem/hid, so a precise vendor:device
rule and a generic class rule coexist; an unmatched device is logged, never
guessed. This resolves docs' "matching stays code until the third bus".
- ABI: child_added and DeviceDescriptor gain vendor/device/subsystem; child_added
gains a bus discriminator (BusKind) so PCI and USB class triples match against
the right namespace.
- pci-bus reads vendor/device (config 0x00) and subsystem (0x2C, type-0) and
reports them.
- library/device/registry: freestanding CSV parser + matchDriver() with
specificity scoring; 5 unit tests wired into `zig build test`.
- etc/devices.csv bundled into the initrd; the kernel serves /etc directly, so
the manager reads it before any filesystem service is up (fat starts later).
- virtio-gpu: drop the now-redundant post-spawn 1AF4:1050 re-confirm, since the
registry binds this driver by exact identity.
- Remove the orphaned system/drivers/display driver (unreferenced by build or
registry).
- docs: new devices-csv.md; device-manager.md "matching stays code" resolved.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KJqSiLLchDUUCoXn5jsiwd
The 11 QEMU-suite fixtures lived mixed into system/services/ with their
binaries bundled at /system/tests/<name>. Now the repo path is the boot
path, like every real service: test/system/services/<name>. fat-test
moves out of the fat server's directory into its own; display-demo
stays a boot service.
- kernel VFS: setInitialRamdisk derives one read-only initrd mount per
top-level tree named by the ramdisk entry paths (/system, /test),
registers ancestors generically with self-parented roots, and refuses
backend shadowing of any initrd tree
- EFI loader: the fallback walk also enumerates \test (optional — a
volume without fixtures still boots); manifest and capsule unchanged
- path literals: vfs-test self-open + create probe, process-test
process_enumerate matches, the args-echo argv[0] expectation; the
kvfs case now covers the /test root end to end
- docs: DFHS /test rows, tree diagrams, loader prose, and the location
convention gain the third home; fixed the input-source link
100/100 QEMU cases pass.
Every user binary and the two device-logic library modules (pci, usb) now
`@import` the concern modules directly instead of aliasing through `runtime`:
runtime.ipc/process/time/service/input/block/display -> @import("<module>")
runtime.device / runtime.device_manager -> @import("driver")
runtime.fs -> @import("file-system")
runtime.Thread -> @import("thread").Thread
runtime.system.{write,writeRecord,klog*} -> logging.*
runtime.system.{sleep,timerOnce,wallClock,clock} -> time.*
runtime.system.{spawn*,kill,exit,yield,processes,...}-> process.*
runtime.system.{mmap,munmap,PROT_*} -> memory.*
runtime.dma.* / runtime.shared_memory.* / runtime.allocator -> memory.*
Each consumer keeps its own alias name (e.g. `const device = @import("driver")`),
so call sites are unchanged and there are no collisions with local `driver`
variables. build.zig now injects the concern modules into every user binary via
`default_imports`; pci/usb module import lists were updated to match.
The `runtime` and `system` shims remain for one more step (root.zig still uses
runtime); they are deleted in C5. Nothing but root.zig imports `runtime` now.
Verified: zig build, zig build test, and 17 QEMU cases (smoke, device-manager,
logger, fat-mount, fat-mutations, usb-storage, usb-hid, display-native,
virtio-gpu, input, thread-spawn, thread-mutex, process-kill, shared-memory,
driver-restart, acpi-ps2, pci-scan).
Finish the direct-import convention: a protocol is now aliased to its module
name in snake_case in every file that imports one — consumers and the runtime
client wrappers alike. This kills the last of the alias variety (the generic
`protocol`, plus `power` and `transfer`), so `device_manager_protocol.Hello`
means the same thing everywhere and grepping a module name finds all its uses.
Pure rename (identifier uses only; prose references left untouched via a
guarded pattern). 14 files, 299/299 lines.
zig build + test green; 16 QEMU cases pass (smoke, device-manager,
driver-restart, device-list, pci-scan, usb-hid, usb-storage, display-service,
display-native, display-reattach, input, acpi-ps2, vfs, fat-mount,
power-button, orderly-shutdown).
power-protocol -> library/protocol/power/power-protocol.zig. Its two consumers
(init and the acpi discovery service) import it directly; the
runtime.power_protocol re-export and the runtime module import are dropped (no
runtime client speaks it). The now-empty system/services/power/ is removed.
This completes library/protocol/: every driver<->service wire contract lives
there and is imported by module name, no protocol is re-exported through
runtime, and the alias chaos (dp/sp, mixed protocol/<x>_protocol) is resolved.
virtio-gpu-protocol stays a driver-private relative import (a hardware command
set, not a driver<->service seam — like virtio-pci.zig beside it).
zig build + test green; power-button, orderly-shutdown pass.
display-protocol -> library/protocol/display/, scanout-protocol ->
library/protocol/scanout/. Consumers import both directly, killing the cryptic
dp/sp aliases in virtio-gpu (now display_protocol / scanout_protocol) and the
runtime.display_protocol / runtime.scanout_protocol re-exports. runtime.display
(the client) still imports display-protocol by name; scanout has no runtime
client, so its runtime module import is dropped too.
zig build + test green; display-service, virtio-gpu, display-native,
display-reattach pass.
device-manager-protocol -> library/protocol/device-manager/. Its consumers
(pci-bus, usb-xhci-bus, the acpi discovery service, device-manager, device-list,
crash-test, and the not-yet-built intel-integrated display sub-driver) import the
module directly instead of through runtime.device_manager_protocol, which is
deleted. runtime.device_manager (the hello client) already imported the module
by name and is unchanged.
zig build + test green; device-manager, driver-restart, device-list, pci-scan pass.
input-protocol -> library/protocol/input/input-protocol.zig. Its five
consumers (usb-hid keyboard/mouse, ps2 keyboard/mouse, the input service) now
import the module directly and alias it input_protocol, replacing the
runtime.input_protocol re-export (deleted) and the inconsistent local `protocol`
aliases. runtime.input (the client) still imports the module by name.
zig build + test green; input, acpi-ps2, usb-hid pass.
vfs-protocol -> library/protocol/vfs/vfs-protocol.zig. fat, its only consumer,
now imports the module directly (const vfs_protocol = @import("vfs-protocol"))
instead of through runtime.vfs_protocol, and the runtime re-export is deleted.
The runtime's own client (runtime.fs) still imports the module by name.
zig build + test green; vfs, fat-mount pass.
Start the protocol tree: wire protocols move to library/protocol/<name>/,
entry file <name>-protocol.zig, module name unchanged. block and usb-transfer
are pure moves — every consumer already imports them by module name, so only
the b.addModule paths change and the empty system/services/block/ is removed.
block-protocol -> library/protocol/block/block-protocol.zig
usb-transfer-protocol -> library/protocol/usb-transfer/usb-transfer-protocol.zig
zig build + test green.
Eight new QEMU cases (thread-fault-group, kill-threaded-group,
kill-via-worker-tid, racing-triggers, exit-group, leader-thread-exit,
thread-exit-solo, shm-mapping-ref) driving seven new thread-test modes; a
shared checkGroupDead asserts the contract everywhere: one notification,
badged with the leader, reason on the leader's record, no member listed,
claims released first.
The shm-mapping-ref case flushed out the per-task DMA/shared-memory arena
cursor bug directly (a sibling's regions mapped over the worker's), so both
cursors moved to the AddressSpaceRef like the mmap/MMIO cursors before them
(threading M7 pattern). Runtime gains Thread.tryExitCurrent for the leader
-EPERM refusal path. Docs updated: threading.md's shared-fate gap is closed,
process-management.md and process-lifecycle.md describe the leader re-key,
plan status = implemented. Full suite: 100/100.
Every doc verified claim-by-claim against the code by parallel audit agents,
then fixed and adversarially re-verified. Two waves of staleness corrected:
the originally audited findings (higher-half boot handoff, kernel VFS
takeover, fault isolation + claim release + driver restart, AML/S5 moving to
ring 3, threading's shipped design, USB+FAT landing) and a second pass of
adjacent claims the verifiers caught (smp.md 'not built yet' intro,
system-requirements' PS/2-only and no-storage claims, halting.md's red-panic
and no-IDT text, testing.md's serial mirroring, router-era vfs-protocol
wording, capsule-first boot loading).
threading.md now documents the shared-fate gap explicitly: the design says a
process dies whole, the kernel today kills only the offending thread.
Also fixes three stale code comments (isr.s exceptionHandler, acpi.zig
sleepValue, build.zig boot-volume) — comments only, no behavior change.
Directory scans and FAT-chain walks re-read the same sectors constantly:
resolving many paths under one directory (a logging burst opening dozens
of files under /var/log/<stamp>/) re-read that directory's sectors and the
FAT from the device every time. Add a 16-line write-through cache under the
single-sector blockRead/blockWrite path, keyed by filesystem-relative LBA
with round-robin eviction, so repeated metadata reads come from RAM instead
of a USB round trip each.
Disciplines that keep it safe:
- Write-through: every write reaches the device immediately and refreshes
the cache, so it never holds dirty-only data — the crash-safe
data->FAT->directory write order and the flush-on-close are unchanged.
- Bulk file data (the B5a multi-sector run path) bypasses the cache and a
run write invalidates any overlapping cached sector, so metadata that
shares a range can never go stale.
- Cluster-zeroing writes uncached (write-once bulk that would only evict
live metadata), invalidating any stale copy.
A new engine test proves a repeated resolve does zero device reads and that
a write is coherent both in-cache and against a cold-mounted filesystem
(it really reached the device).
Full QEMU suite green (the one intermittent `logger` miss is the documented
pre-existing AP ring-3 fault: 16/16 logger reruns pass on this change and
every other storage case is green; the fat engine is single-threaded and
the cache is bounded global state, so it cannot itself fault intermittently).
The FAT engine read and wrote one sector per device command — one SCSI
READ(10)/WRITE(10) over USB Bulk-Only Transport per 512 bytes, so every
file read, log write, and cluster fill paid a full USB round trip per
sector. The lower stack (runtime.block, the block protocol, usb-storage's
read10/write10 + count*block_size data stage) already carried a
multi-sector count; only the engine's BlockDevice interface and IpcBlock
were single-sector.
Widen BlockDevice to move a run of `count` contiguous sectors per call
(readBlock/writeBlock kept as count=1 wrappers, so metadata call sites —
FAT sectors, directory entries — are untouched). readFile and writeFile
now coalesce the aligned full-sector middle of a transfer into one command
(capped at the cluster boundary and the 4 KiB bounce = 8 sectors), reading
straight into / writing straight from the caller's buffer with no staging
copy. Full-sector writes skip the read-modify-write entirely, since they
overwrite the whole sector. Partial head/tail sectors keep the per-sector
RMW path.
IpcBlock passes count through to the block driver and sizes its bounce from
engine.max_transfer_sectors (already 4 KiB — no new allocation). A new
spc=8 engine test proves runs coalesce (a 21-sector overwrite drops from
>=21 writes to <=5) and that reads/writes round-trip byte-identical,
including an unaligned offset spanning a cluster boundary.
Full QEMU suite 92/92.
B2 — the 1-in-3 boot-time READ CAPACITY failure, root-caused: the xHCI
library's awaitTransfer claimed ANY unclaimed transfer event as its own
completion. An interrupt-endpoint event whose TRB pointer no longer
matched the armed subscription (an error or stale completion from the
keyboard/mouse polling concurrently with storage bring-up) fell through
and was misread as the bulk transfer's completion — desynchronizing the
mass-storage bulk protocol in controller state that SURVIVED driver
restarts, so every retry failed too. Awaited transfers now match the
event's slot id and endpoint DCI; foreign events are dropped and named.
Twelve consecutive runs of the previously-flaky cases pass; the full
suite is green with none of its old intermittents.
B1 — and when storage does fail transiently, the system now heals
instead of giving up forever: a nonzero exit maps to ExitReason.aborted
(a deliberate FAILURE exit — supervisors restart those with backoff,
unlike a clean .exited), usb-storage exits nonzero when a PRESENT
device fails bring-up, and the fat service no longer blocks its harness
polling for a block device and then dies — it serves immediately
(requests fail politely), retries on a 500 ms timer, and mounts
whenever storage appears, including after a driver restart.
Its only content was the redundant announce line — the useful shutdown
marker was written AFTER the final drain, so it reached serial and the
ring but never a file, and a truncated boot could only be inferred from
what was missing. The marker now enters the ring BEFORE the drain, so
the drain carries it into logger.log: a directory whose logger.log ends
with 'shutting down; final flush' is complete through shutdown; one
without it was cut early. The sequence-count epilogue stays serial-only,
after the drain, by design.
The display service claiming the framebuffer suppresses the on-screen
boot transcript — correctly in normal operation, but on a serial-less
machine being debugged, the timeline vanishes just when it matters. A
diagnose image (zig build -Ddiagnose=true) has init skip the display
service and demo: the timestamped transcript stays on screen
indefinitely, and the power button still runs the orderly shutdown (so
the logger's files land when storage works).
First use immediately caught a real bug IN QEMU: an intermittent (~1
in 3) usb-storage READ CAPACITY failure after a ~21 s stall — after
which usb-storage and fat both exit cleanly and nothing retries: one
transient early-boot USB failure leaves the system permanently without
storage (and therefore without logs). That no-retry policy is a prime
suspect for real hardware never mounting /var, and is Track B's first
work item.
Per docs/coding-standards.md (no Unix-abbreviation exception): syscalls
shared_memory_create/map/physical, kernel SharedMemoryObject + handlers,
runtime.shared_memory (library/runtime/shared-memory.zig), the
shared-memory-server/-client test services, the shared-memory QEMU case,
and docs incl. vdso.md's danos_shared_memory_*. 87/87 QEMU tests pass.
runtime.fs now routes every path through fs_resolve: kernel-served
/system nodes are read via fs_node (tokens, no open state); everything
under a userspace mount goes straight to the owning backend's endpoint
with the kernel-rewritten mount-relative path — one syscall of naming,
then the unchanged vfs-protocol rendezvous, public API untouched. mkdir/
unlink/rename resolve-then-forward (rename checks both paths land on
the SAME backend); mount is the fs_mount syscall.
The fat server mounts twice — /mnt/usb from the volume root and /var
from its /var subtree — so the logger now writes the FHS path
/var/log/<boot-stamp>/... and swapping the persistent medium later
touches only fat's two mount calls. With clients holding fat's node ids
directly, fat records each handle's owner, checks it, and sweeps a dead
client's handles via the published exit events (the old router's
pattern, now where the state actually lives).
The userspace vfs server and its router die; ServiceId.vfs=1 stays
reserved-retired; protocol.zig moves to system/vfs-protocol.zig (the
wire contract is backend-only now). vfs-test becomes the ring-3 proof
of the kernel VFS (own-binary ELF magic through /system, read-only
refusals, listing); vfs-client-death becomes the fat sweep test over
the full storage chain, with a ring-scanning check (the last-write
buffer is too racy under a chattering tree).
The logger service drains the tagged ring every 250 ms and demultiplexes
it into one file per process on the flash volume:
/mnt/usb/var/log/2026-07-21T150434Z/system/services/fat.log
[ 0.214] mounted FAT (fat32, 128992 clusters, partition lba 0)
The boot stamp is the wall-clock anchor from klog_status (FAT-safe, no
colons); the kernel's records go to kernel.log; each line carries the
record's monotonic timestamp and level. Storage is best-effort and
late — the ring buffers a whole boot until the mount appears, then the
first drain writes the backlog, storage-stack records included. Lost
records surface as '-- N records lost --' from sequence gaps. Files
close (= fat's device cache flush) after a 2 s quiet period, bounding
data-at-risk without per-record flush thrash. The logger announces
itself once — a periodic status line would feed the stream it drains.
init spawns the logger last, so the reverse-order shutdown stops it
first and its final drain runs over a live storage chain; log-flush and
init's own DANOS.LOG shutdown flush retire (superseded). runtime.fs
makePath treats components at or above a mount point as router names —
create is best-effort per prefix, the final verdict is exists(path).