The adversarial S5 review found a real, unrecoverable defect: onMediumEvent
deduped medium_changed events on a module-global last_medium_change compared by
equality against the event's change_count. But change_count is a PER-DRIVER
counter that restarts at 0 in every usb-storage instance — it is bumped +%=1 and
published only on a real medium transition, so within one instance every count
is unique and monotonic and an equality dedup can never legitimately fire.
The global was carried across a driver restart — S5's OWN crash-rebuild path —
so a fresh instance's first eject (count=1) collided with a stale last==1 and was
dropped. removeDevice never ran; the filesystem kept serving I/O against absent
media forever, and nothing else recovered it: onGeometry answers from a cached
block_count so channelAlive stays true, and isDevicePresent stays true (the
device never left the tree). medium_changed is the sole eject oracle there.
The dedup guarded a re-delivery the driver already makes impossible, and its
only observable effect was this bug. Remove it: react to each present-edge
directly. Both branches are idempotent (a freed device stops matching dev.used)
and the poll reconciles, so acting on every genuine edge is safe — and a fresh
driver instance's counter can no longer be mistaken for the previous one's.
Verified: build + bounds green; volume-removal, volume-medium-change and
volume-driver-restart all pass, so both removal paths survive the change. The
driver-restart-then-eject intersection that triggered the collision cannot be
staged in QEMU — the internal driver-kill cannot be ordered against a QMP eject,
and a usb-storage device_add is not re-presented — so the guarantee rests on the
driver's one-publish-per-transition-with-unique-count contract.
The V4 review's open edge: a storage driver that crashes while its device
stays in the tree left fat wedged on a dead channel — device-presence
polling (a device-manager enumerate) still reported the device present,
so nothing reaped it. pollTick now also probes channelAlive(dev), a
geometry() on the block channel that fails fast on the dead endpoint; a
present device with a dead channel is reaped like a pull, and the adopt
loop re-adopts it on the restarted driver's fresh channel — the rebuild.
The manager's own liveness probe makes fat self-detection unnecessary:
it rebuilds regardless of the wedged filesystem's state.
The drill: the device manager gains a test-storage-restart mode that
kills usb-storage once, ~2s after its hello (post-mount); a new
volume-driver-restart kernel case boots a manual tree with it, and the
QEMU case asserts a SECOND mount of the same id-path after the reap —
the rebuild. A pre-S5 manager, checking only device presence, never
reaps, so the second mount never appears. The manually-spawned volume
manager needed kernel-supervisor protocol grants (bind its name, open
the device manager), as the other manual-tree services already have.
Full suite 133/133.
The block client gains subscribeMedium / unsubscribeMedium /
decodeMediumChanged (the reserved subscribe/unsubscribe verbs plus the
MediumChanged decode), so a consumer never hand-rolls the wire format.
The volume manager subscribes to each device it adopts and consumes the
event through the new on_buffered_message seam — never the protocol
dispatch, whose op numbers collide with the manager's own hello.
A medium leaving while its device stays in the tree (a card reader, an
eject) now runs the SAME kill-retire-remount path as a pulled stick:
absent retires the volume, present re-probes it. That closes the "two
triggers, one lifecycle" the architecture specifies — device-presence
polling alone could never see a medium leave under a present device.
dropDevice unsubscribes before closing so the driver's bounded
subscriber table frees the slot; on a dead channel (a real pull) the
call fails fast, proven by volume-removal still passing.
New volume-medium-change case: eject the medium (not the device) ->
"medium absent" -> the manager unmounts. Fails against a pre-S5 manager
that never subscribed. Suite 132/132 (device-authority is the known
child-cleanup flake, green on rerun).
Lift the one-device/one-volume cap. bringUpVolume now probes the whole
partition table (allVolumes uncapped) and spawns a confined filesystem
per volume; pollTick loops it to adopt every present, not-yet-adopted
device each tick.
The subtlety is adopt-once-and-keep: a device is recorded in the table
the first time it is seen and kept until it leaves the tree, even when it
carries no servable volume or its geometry cannot be read. Dropping an
unservable device would make openAnyStorage hand back the same one every
tick and starve the devices behind it; keeping it lets the scan advance
past it. A genuine removal frees the slot; a re-insert (fresh device id)
is probed anew.
The boot image is a single bare-FAT volume, so the full suite is
unchanged at 128/128.
Replace the single `var volume: ?Volume` and file-global supervision
state with two fixed tables: devices[maximum_devices] owning each adopted
block channel once, and volumes[maximum_volumes] each carrying its own
identity, id, mount prefix, and supervision fields (restarts, spawn_ns,
failed, restart_pending, restart_due_ns). A monotonic next_volume_id
never reuses ids, so a stale hello can't address the wrong child.
Lookups (deviceById, volumeById, volumeByPid, firstUsedVolume) and
claims (claimDevice, claimVolumeIndex) replace the ad-hoc singletons.
pollTick reconciles devices first (removeDevice drops their volumes),
then per-volume restarts, then idle bring-up.
This step stays one-device/one-volume on purpose: bringUpVolume adopts
the first device and caps allVolumes to a single partition, so behavior
is identical and the full suite stays 128/128. Uncapping and adopt-all
land next.
The mechanism the id/label split needs: a `volumes` verb whose reply packs the
mounted volume's {id, mount_path, label} into the tail (VolumeInfo.encode/decode
— three length-prefixed strings). Software keys on the id (the mount path is
/volumes/<id>); a shell or file manager shows the label — the database id/name
split made a query. The VM's onVolumes answers from the mounted volume, empty
reply if none. Two host round-trip tests (encode/decode; too-small buffer and
short-tail rejection). No runtime consumer yet — the first is a userspace shell;
the hello handshake is unaffected (fat-mount/volume-probe green).
The volume manager reads its policy from configuration at boot (loadTables,
mirroring the device-manager registry load): filesystems.csv (content signature
-> service binary) and volumes.csv (optional id -> mount-prefix override), each
held in a static source buffer with declared bounds. On probe it picks the
binary from the volume's signature (unserved + logged if no row matches, like an
unbound device) and composes the mount path — a volumes.csv override, else the
default /volumes/<id> from volume-map.idString — then spawns that binary with
argv {volume-id, mount-prefix}. Behavior-preserving: fat still ignores argv[2..]
and uses its hardcoded mounts, the binary resolves to /system/services/fat, so
the FULL suite stays green (127/127); the flip to argv-driven mounts and the
/volumes/usb -> id-path migration land in step 5.
The identity ladder's flag-day — no behavior change. partition.firstVolume stops
taking one preloaded block-0 slice and takes a SectorReader (a read-one-sector
fn), so it can reach GPT metadata at LBA 1 and each partition's VBR on demand
(the next commits). The u64 identity becomes Identity{rung,key,label}: key is the
id (the mount path derives from it), label is display metadata (empty at rung 4).
Identity equality is id-only (rung+key) — the label never enters it. Only rung-4
(MBR sig+index / bare-FAT index 0) is produced, byte-identical to before; the
four host tests port to a RAM-disk reader, and fat-mount/volume-probe/
volume-removal stay green.
The V4 adversarial review found removeVolume's comment overclaiming: it said
"the kernel sweeps a dead backend's mounts", which reads as an eager death-time
sweep. There is no such sweep. Killing the filesystem marks its backend endpoint
dead (killOwnedEndpointsLocked), and the VFS router retires each mount that
endpoint backed lazily, on the next path resolution under it (resolvePath sees
the dead backend, frees the slot, returns not_found). The functional guarantee
the comment promised — killing the filesystem retires its mounts — holds; only
the described mechanism was wrong. Comment-only; no behavior change.
Inline V4 review (the boundary-review workflow stalled): the poll ran a due
fat-restart before the presence check and returned, so a fat death followed
by a device removal would respawn fat against the now-dead channel and churn
until the crash cap before the removal was noticed. Reorder: check the
specific device's presence first (unmount if gone), and only fire a due
restart once the device is confirmed present. Neutral: fat-mount,
volume-removal, amd-iommu-usb-storage green.
Noted V4 limitations (not fixed here, edge cases outside the user unplug
case): a usb-storage DRIVER crash (device stays, driver restarts with a new
endpoint) leaves fat holding a dead channel — the device is still present so
removal is not detected; fat would need to observe its channel death and
exit. Deferred with the medium_changed subscription and multi-volume.
The full suite caught a V4 regression: under AMD-Vi the device-manager tree
carries more than one mass-storage-identity entry (a phantom no driver is
bound to, which answers a consumer hello with NO channel). V4 split presence
from acquisition and picked the FIRST identity match blindly, so it kept
helloing the phantom (device 27) and never reached the real storage (device
31). V3's inline loop had skipped no-channel entries with `orelse continue`;
the split lost that.
Restore it: openAnyStorage tries each matching entry and takes the first whose
channel opens, recording its device id. Removal detection then watches THAT
specific device id leave the tree (isDevicePresent), not "any mass-storage" —
so a phantom that never leaves cannot mask a real removal. Both are bare
enumerates; the hello only happens while bringing a volume up.
Green: amd-iommu-usb-storage, fat-mount, volume-removal.
The volume manager stops probing-once and polls storage presence for the life
of the boot: findStorageDevice enumerates the device-manager tree (presence
only, no consumer-hello, so it is cheap and leaks nothing). The volume is now
a field that goes null and back — the whole lifecycle:
- storage present + no volume -> open the channel, probe, confine + spawn the
filesystem (openStorage is the one consumer-hello, on the insertion edge);
- storage gone + have volume -> kill the filesystem (its mounts retire via
the kernel dead-backend sweep), close the dead channel, clear the volume;
- fat crash -> the same supervised backoff/cap as before,
folded into the poll (one timer).
This also subsumes the V3-review leak fix (no per-poll consumer-hello) and the
no-volume retry (a present-but-unreadable device keeps polling).
The user's case — pull the boot stick, plug it back — is a DEVICE unplug (the
stick IS the device), so the mass-storage child leaves the device-manager tree
and the poll catches it. volume-removal asserts the unmount and discriminates:
against the V3 probe-once volume manager the removal is never noticed (0/1).
The re-mount on replug is the VM's bringUpVolume firing when the device
returns — correct and in place, but not QEMU-testable here: device_add of
usb-storage to the boot xHCI controller is not re-presented to the guest (no
port-connect on any port), a harness quirk, not a VM issue. On real hardware
the bus's per-tick port poll catches a reconnect (H1 proves reconnect on a
second controller); bench-verify the full round trip.
Five confirmed defects from the boundary review:
1. (security) The VM never checked a partition fit inside the device, so a
crafted MBR could hand the driver a range whose base+lba wraps past a u32
— panicking usb-storage in a loop, and at multi-volume overlapping a
neighbour. This is the exact invariant the clamp's overflow-safety rests
on. partition.firstVolume now skips any entry that runs past the device
(host-tested), establishing the invariant where the untrusted bytes are
first read.
2. (leak) The probe re-acquired a fresh block channel on every 500 ms retry,
leaking a handle each time on a medium-absent device. The channel is now
acquired once and kept.
3. (wedge) A failed spawn or defineRange stranded the volume with no retry;
both now arm a backoff restart.
4. (loop) fat respawn had no exit-reason gate, no backoff, no crash-loop cap
— a faulting filesystem respawned in a zero-delay loop, and a clean exit
was resurrected. Supervision now mirrors the device manager: a clean exit
is not restarted, a fault backs off, three fast deaths give up.
5. (removable) A device that parsed to no volume was terminal; it now keeps
polling so an inserted medium is picked up — the removal-lifecycle trigger.
Known limitation (noted, not fixed here): if the VM itself crashes and init
restarts it, the orphaned fat keeps serving vfs while the new VM spawns a
second fat whose bind is refused — the same "manager restart re-learns the
world" gap the device manager also defers. The old fat keeps storage working.
Neutral: partition unit tests + fat-mount, volume-probe, block-range, logger
all green.
The load-bearing step. The FAT service stops acquiring its own volume: the
volume manager spawns it (per volume), defines its partition range on the
storage driver BEFORE it runs, and answers its startup hello with the
range-confined block channel over a new volume-manager protocol. fat never
finds its storage by name and never sees the whole device — establishment
by lineage, one layer up from the driver tree.
- New library/protocol/volume-manager: one verb, hello(volume-id) -> the
block channel as the reply capability (the P0 reply-cap path).
- The volume manager becomes the confinement CONTROLLER: it defines the first
range on usb-storage, so no other party can confine a filesystem. It
supervises the filesystems it spawns and respawns one on death (the reap-
and-rebuild the device manager proved, one layer up).
- fat: drops acquireVolume(device-manager); hellos the volume manager for its
channel; reads its volume id from argv[1]. main takes process.Init now.
- init.csv no longer spawns fat (the volume manager does); protocol.csv
rewires fat to be supervised by the volume manager (bind vfs, open
volume-manager) and drops fat open device-manager.
- The block-range fixture boots registry + device-manager only (not the full
tree), so the volume manager is absent and the fixture stays the sole
confinement definer — otherwise the volume manager would take the
controller first and refuse it.
Verified end to end (VM probes -> spawns fat -> confines it -> hands over the
channel -> fat mounts) and neutral: 18/18 across the fat family, logging,
shutdown, both IOMMU variants, usb restart, vfs, conformance, confinement.
The storage layer gains its policy home (storage-architecture.md): a new
system/services/volume-manager, spawned by init, that acquires the mass-
storage block channel through the device manager (the same lineage a
filesystem uses), reads block 0, and parses the first volume out of it. The
partition-table walk that lived in the FAT engine moves here, above the
driver where it belongs (partition.zig, host-tested: MBR entry, bare-FAT,
no-signature). Identity is the MBR disk signature + partition index — the
weak rung of the ladder; GPT GUID and FAT serial refine identityOf without
changing shape.
This increment is discovery + probe + log only, additive: the FAT service
still acquires its own volume, so nothing changes for it. Confining each
filesystem to its partition and spawning one per volume (the flip) lands
next, keeping fat working throughout.
Grants + wiring: init.csv spawns it after the device manager; protocol.csv
grants bind volume-manager + open device-manager. Verified: volume-probe
asserts the parse (bare-FAT volume at lba 0), neutral 10/10 across storage,
restart, display, logging, confinement — the volume manager now runs in
every boot and disturbs nothing.