The volume manager, per-sender range confinement + gate, medium_changed event, per-volume spawning + supervision, ownership-gated fs_unmount, and the removal half of the lifecycle are built. Left honestly pending: the volumes.csv/filesystems.csv maps, the fuller identity ladder, multi-volume, the VM consuming medium_changed (removal uses device-presence polling), and remount-on-replug end-to-end (bench-pending — QEMU can't re-present the boot-controller device). Full suite 127/127.
17 KiB
The storage architecture: layers, boundaries, responsibilities
Status: the layered model below is the settled design (storage-design-rationale.md records how it was reached, and volume-manager-plan.md how it was built). Built (the volume-manager track, V0–V4): the data path, the driver range confinement (per-sender clamp + the confinement gate), the
medium_changedpresence event, the volume manager itself — it probes the partition table, confines each filesystem to its partition, spawns one filesystem per volume, and supervises it — and the removal half of the lifecycle (a pulled stick unmounts). Still pending: the fuller identity ladder and thevolumes.csvmount map, multi-volume (one FAT volume today), the volume manager consumingmedium_changed(removal is detected by device-presence polling; the event is published but only a card-reader medium change needs the subscription), and the remount-on-replug end-to-end (the logic is in place; QEMU can't re-present the boot-controller device, so it is bench-verified). A few markers below are left where a duty is still pending.
The model
One pattern, applied twice: an application is a client of a service over a protocol; a service is a client of a driver over a protocol.
application
│ vfs protocol (routed by the kernel mount table)
▼
filesystem service ── one process per volume
│ inside: [file-protocol shell] ── Zig API ── [fs engine]
│ block protocol (channel received at spawn)
▼
storage driver ── one process per device (usb-storage per stick)
│ usb-transfer protocol (channel received via hello, by lineage)
▼
bus driver ── one process per controller (usb-xhci-bus)
│ hardware
Three kinds of boundary, deliberately different:
- Protocol boundaries sit where processes meet (vfs, block, usb-transfer). Each side is independently restartable; channels are established by capability handoff, never by registry names (communication.md "Establishment: two planes"). Steady-state file I/O crosses exactly two — vfs and block.
- Library boundaries sit where concerns meet inside one process. The
filesystem service's engine (on-disk format logic, host-testable, behind
the four-function
BlockDevicevtable) and shell (IPC serving, open-node table, mount registration) talk through a Zig API. No protocol between them: they share fate regardless — a corrupt engine corrupts the answers either way — so a channel there would add a hop per file operation and buy nothing. The seam exists for compile-time testability and reuse; the process is the restart unit. - Control-plane relationships sit beside the data path, never on it. Two supervisors, one per layer: the device manager wires and revives the device layers (bus and storage drivers — devices only); the volume manager (built) wires and revives the volume layer (filesystem services). Neither touches steady-state I/O.
Who does what
Bus driver (usb-xhci-bus, per controller): hardware only. Enumerates, reports children to the device manager, serves the transfer contract to its own children's class drivers, routed by lineage.
Storage driver (usb-storage, per device; later nvme, ahci): hardware →
blocks. Speaks its transport upward from its device; serves the block
protocol (geometry, read, write, flush, attach, detach — and, planned, the
pushed medium_changed presence event, translated from the transport's
native signal). Content-blind, permanently:
it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no
GPT, no filesystem magic, ever. (Planned) it gains one content-blind
mechanism: named sub-ranges — "serve blocks [a, b) as a channel of the
same block contract". It clamps and offsets; it never knows the numbers came
from a partition table. The clamp lives here and nowhere else because a
channel must carry exactly the authority it grants: handing a filesystem the
whole disk plus a polite base offset would let a buggy or compromised
filesystem scribble the neighboring partition — the same authority-overshoot
the device-authority track eliminated for MMIO and DMA.
Volume manager (built; system/services/volume-manager): the policy home
of the volume layer, one service, supervised by init. It watches the device
manager's tree for a storage provider; when one appears it consumer-hellos for
the block channel, reads the partition table and the first blocks itself
(it is the prober), defines the volume's sub-range on the driver, spawns the
matching filesystem service confined to that range, and supervises it (backoff,
crash-loop cap). (Pending): it decides mount placement from volumes.csv and
picks the filesystem binary from filesystems.csv — today it hands every
FAT-shaped volume to the FAT service and the FAT service carries hardcoded mount
prefixes. Those tables are CSV configuration, read by it (the policy), enforced
by nobody else:
filesystems.csv(pending) — content signature → filesystem binary. Adding a filesystem adds a row.volumes.csv(pending) — the mount map, danos's fstab: volume identity → mount prefix, keyed on content identity and never on port, path, or arrival order (the lesson of Linux's/dev/sda1-era fstab, which broke on every port move untilUUID=replaced it). Identity is read off the medium by the prober, strongest first: GPT partition GUID → filesystem UUID → FAT serial + label → MBR signature + partition index → anonymous (generated name, no persistence). Same identity, same mount point: a drive moved to another port — USB to another hub, SATA to another bay, even a stick returning in a different dock — lands exactly where it was. Duplicate identity (cloned sticks, together) is policy: first keeps the name, the second mounts suffixed and is logged loudly. The boot volume is the recorded identity of the volume carrying/system/configurationand/system/logs, findable on any port. Unknown volumes mount under/volumes/<derived name>.
Filesystem service (the FAT service today; one process per volume): the proven unit — block-client + engine + file-protocol provider in one binary. It receives its block channel at spawn; it never discovers devices. It registers its own mounts with the kernel; its write cache lives inside the process, so a write error is observed by the code that owns the volume and surfaces on the owning channel (the anti-fsyncgate rule — never a system-wide dirty pool). (Today, interim:) fat also acquires its own volume (first mass-storage child by enumeration order) and hardcodes its mount prefixes; both migrate to the volume manager.
Kernel (mechanism only): the mount table routes paths to backend
endpoints — resolve and redirect, never data. Remount-replace is the restart
story; a dead backend's slot is swept lazily on the next resolution.
(Planned fix:) fs_unmount gains ownership — only the mounting endpoint's
holder may unmount; today it is ungated, which is safe with one mount owner
and wrong with several.
Adding a filesystem
- Write the engine: pure Zig, no IPC, behind the
BlockDevicevtable (four functions). On-disk layout inalign(1)extern structs. Host-test it withzig build testfixtures — the FAT engine (engine.zig, on-disk.zig) is the template, and its ~500 lines of host tests the standard. - Reuse the shell: the filesystem harness — establishment, the
badge-scoped open-node table, the nine vfs-protocol handlers, mount
registration, the removal path — is shared code, not per-filesystem code.
(Planned: extracted from fat's 434-line shell into a library before the
second engine is written.) An engine plus a
mainwiring it into the harness is a complete filesystem service. - Add the configuration row: one line in
filesystems.csvmapping the on-disk signature to the binary. No other component changes: the volume manager probes, matches, spawns; the vfs protocol is already backend-neutral (a second provider serves it in production today — init's synthetic registry backend). - Prove removal: extend the removable-media suite (below) for the new filesystem — surprise-yank during writes must lose only what was unflushed, said honestly, with the volume consistent enough to remount.
The engine never sees: partitions (it receives a volume-shaped block channel; base offsets are the driver's clamp), device discovery (the channel arrives at spawn), mount policy (prefixes are handed to it), or other volumes (one process, one volume).
Removable media: responsibilities on removal, per layer
The design rule, learned from what Linux cannot do: there is exactly one
surprise-removal path — kill the filesystem process, retire its mounts,
respawn on return. No half-alive states, no remount-ro, no mounts that
error forever (Plan 9's dead-server wart).
The path has two triggers, one lifecycle: the device leaving (the
storage driver dies — channel death, the table below), and the medium
leaving while the device stays (an SD card pulled from its reader, an ATAPI
tray opened — including USB card readers today). The second trigger is a
pushed medium_changed event on the block protocol (planned): the storage
driver translates its transport's native signal (SCSI UNIT ATTENTION, AHCI
PxSSTS, NVMe namespace-change AER) into presence-changed — never content —
and the volume manager runs the same kill-retire path, then re-probes on
medium return exactly as on device return. Without it, a swapped card would
be served with the previous card's filesystem state.
| Layer | Observes | Must do | Guarantees |
|---|---|---|---|
| Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report child_removed per interface |
the device tree is honest within one reconcile tick |
| Device manager | child_removed / reporter death |
prune the child; reap the bound driver (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
| Volume manager (removal built; remount bench-pending) | the storage device leaving the device-manager tree (poll) | kill the filesystem service of that device's volume; its kernel mounts retire | one removal path; mounts never dangle; log persistence stops cleanly |
| Filesystem service | its block channel dies (EPEER) mid-operation, or it is killed by the volume manager |
if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is lost and said to be lost | the FAT dirty flag on disk marks the unclean removal; the process never serves from behind a dead channel |
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated fs_unmount (built, V0) |
resolution under a dead mount is not_found, not a hang |
| Application | not_found / error on paths under the vanished mount |
its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
On return the same table runs upward in reverse: the bus re-enumerates and
re-reports; the device manager respawns the storage driver (built); the volume
manager re-probes — same content, same volume identity — respawns the
filesystem service, and remounts at the same prefix; applications see the
subtree reappear. A different stick in the same port is a different volume
(identity is content, not port) and mounts wherever its identity says —
possibly nowhere but /volumes/….
The boot volume: identity is what makes yanking it survivable
The boot volume is the sharpest instance of the return story, because the
system's own log persistence rides it (/system/logs), and because nothing
else about the running system depends on it at all — every binary and the
boot-time configuration live in the initial ramdisk. Its lifecycle,
end to end:
- At first mount the volume manager records the identity of the volume
carrying
/system/configurationand/system/logs— THE boot volume, from then on a fact about content, not about a port. - While it is absent, the system runs on. The kernel log ring keeps
accumulating — it is the buffer that makes the absence survivable — and
the logger keeps draining it; only persistence pauses. File operations
under the retired mounts fail honestly (
not_found). The ring is bounded, so a long absence overwrites its oldest entries: that window is the data loss, and it must be said — the logger marks the gap in the file when persistence resumes, never splicing the stream silently. - On return — any port, any hub, even a different transport — the
prober reads the same identity, the mount map answers with the same
prefixes, the filesystem service is respawned, and
/system/logsis the same tree it was: the logger resumes appending into the SAME<boot-stamp>directory, per-binary files continuing where they left off (plus the gap marker if the ring wrapped). - What never comes back is the write-back window lost at the yank — the dirty-honesty rule, unchanged; the boot volume gets no exemption.
This is also the requirement that shapes the logger: it treats the log tree
as a volume that comes and goes — failed flushes are retried on the same
patient cadence the fat service already uses for storage that arrives late,
never abandoned after the first not_found.
What is lost on a surprise yank is exactly the write-back window of the filesystem service, no more: the engine owns its cache, so the blast radius of a yank is one volume's unflushed writes, marked by the on-disk dirty flag. A filesystem format with better crash honesty (journaling, copy-on-write — the lesson of QNX's Power-Safe) narrows that window further and slots in as implementation N+1 through the table above, changing nothing else.
How the lifecycle is enforced
Responsibility tables are convention; convention is worthless (the bounds audit's lesson). The volume lifecycle is enforced with the same three levers the device lifecycle already uses, mapped one-to-one:
- A filesystem cannot acquire — it can only be given. Filesystem
binaries hold NO establishment grants: no
open device-manager, no registry name to look up. The only block channel a filesystem process ever has is the one handed to it at spawn by the volume manager. Serving the wrong volume, a second volume, or a self-discovered volume is not forbidden but impossible — the same way a driver cannot claim hardware it was not delegated (device-authority.md). - The volume manager supervises with teeth. Mount-within-deadline or be
stopped (a filesystem wedged on a corrupt volume is killed and crash-loop
capped; the volume is marked bad, not retried forever). On removal the
manager KILLS the filesystem process and retires its mounts — the polite
observe-
EPEER-and-exit path is an optimization; the kill is the guarantee, because a process cannot be trusted to observe its own obsolescence (the reap argument, proven on the device tree). - The shared harness is how every filesystem inherits the lifecycle by construction. The harness — not the engine — owns the state machine: establishment at spawn, mount registration, the dirty-flag set/clear bracket, error-out-and-exit on channel death. The engine sits behind the four-function vtable and never sees a channel; it cannot opt out of the lifecycle for the same reason it cannot find a device. This is why the harness is extracted BEFORE the second engine is written.
Beneath all three, two kernel mechanisms: ownership-gated fs_unmount (only
the mounting endpoint's holder may unmount), and the lazy dead-endpoint mount
sweep as the backstop nothing can disable. And above them, the check: a
lifecycle conformance drill — mount, serve, yank mid-write, verify honest
loss, replug, remount — parameterized over filesystem implementations, so
"danos supports filesystem X" MEANS "X passes the drill through the harness",
exactly as provider conformance means passing the reserved-verb suite.