Flip the storage docs to match S2: filesystems.csv (signature -> binary) and volumes.csv (identity -> optional override) are built; a volume's mount path IS its content id (/volumes/<id>), never a port name; the label is display metadata a `volumes` query returns. fat receives its mount path via argv[2] rather than hardcoding it. Still pending: rung 2 (filesystem UUID, needs a non-FAT engine), multi-volume (fat's boot rewrites stay unconditional until S3), medium_changed consumption, remount bench-verification.
18 KiB
The storage architecture: layers, boundaries, responsibilities
Status: the layered model below is the settled design (storage-design-rationale.md records how it was reached, and volume-manager-plan.md how it was built). Built (V0–V4 + the storage-stack S1/S2): the data path, the driver range confinement (per-sender clamp + the confinement gate), the
medium_changedpresence event, the volume manager itself — it probes the partition table, confines each filesystem to its partition, spawns one filesystem per volume, and supervises it — the removal half of the lifecycle (a pulled stick unmounts), the identity ladder (GPT GUID + name, FAT serial + label, MBR), and the mount map:filesystems.csv(signature → binary) +volumes.csv(identity → optional override), a volume's mount path IS its content id (/volumes/<id>), with the label as display metadata avolumesquery returns. Still pending: thefilesystem UUIDrung (needs a non-FAT engine), multi-volume (one FAT volume today; fat's boot rewrites are unconditional until S3 makes them content-conditional), the volume manager consumingmedium_changed(removal is detected by device-presence polling; the event is published but only a card-reader medium change needs the subscription), and the remount-on-replug end-to-end (the logic is in place; QEMU can't re-present the boot-controller device, so it is bench-verified). A few markers below are left where a duty is still pending.
The model
One pattern, applied twice: an application is a client of a service over a protocol; a service is a client of a driver over a protocol.
application
│ vfs protocol (routed by the kernel mount table)
▼
filesystem service ── one process per volume
│ inside: [file-protocol shell] ── Zig API ── [fs engine]
│ block protocol (channel received at spawn)
▼
storage driver ── one process per device (usb-storage per stick)
│ usb-transfer protocol (channel received via hello, by lineage)
▼
bus driver ── one process per controller (usb-xhci-bus)
│ hardware
Three kinds of boundary, deliberately different:
- Protocol boundaries sit where processes meet (vfs, block, usb-transfer). Each side is independently restartable; channels are established by capability handoff, never by registry names (communication.md "Establishment: two planes"). Steady-state file I/O crosses exactly two — vfs and block.
- Library boundaries sit where concerns meet inside one process. The
filesystem service's engine (on-disk format logic, host-testable, behind
the four-function
BlockDevicevtable) and shell (IPC serving, open-node table, mount registration) talk through a Zig API. No protocol between them: they share fate regardless — a corrupt engine corrupts the answers either way — so a channel there would add a hop per file operation and buy nothing. The seam exists for compile-time testability and reuse; the process is the restart unit. - Control-plane relationships sit beside the data path, never on it. Two supervisors, one per layer: the device manager wires and revives the device layers (bus and storage drivers — devices only); the volume manager (built) wires and revives the volume layer (filesystem services). Neither touches steady-state I/O.
Who does what
Bus driver (usb-xhci-bus, per controller): hardware only. Enumerates, reports children to the device manager, serves the transfer contract to its own children's class drivers, routed by lineage.
Storage driver (usb-storage, per device; later nvme, ahci): hardware →
blocks. Speaks its transport upward from its device; serves the block
protocol (geometry, read, write, flush, attach, detach — and the pushed
medium_changed presence event, published today from a slow TEST UNIT READY
poll; planned: translating it from the transport's native signal instead of
polling). Content-blind, permanently:
it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no
GPT, no filesystem magic, ever. It has one content-blind mechanism (built):
named sub-ranges — "serve blocks [a, b) as a channel of the
same block contract". It clamps and offsets; it never knows the numbers came
from a partition table. The clamp lives here and nowhere else because a
channel must carry exactly the authority it grants: handing a filesystem the
whole disk plus a polite base offset would let a buggy or compromised
filesystem scribble the neighboring partition — the same authority-overshoot
the device-authority track eliminated for MMIO and DMA.
Volume manager (built; system/services/volume-manager): the policy home
of the volume layer, one service, supervised by init. It watches the device
manager's tree for a storage provider; when one appears it consumer-hellos for
the block channel, reads the partition table and the first blocks itself
(it is the prober), defines the volume's sub-range on the driver, spawns the
matching filesystem service confined to that range, and supervises it (backoff,
crash-loop cap). (Built): it picks the filesystem binary from
filesystems.csv by the volume's content signature, and mounts the volume at its
content id (/volumes/<id>) — or a volumes.csv override. The label is display
metadata the volumes query returns, never the path. Those tables are CSV
configuration, read by it (the policy), enforced by nobody else:
filesystems.csv(built) — content signature → filesystem binary. Adding a filesystem adds a row.volumes.csv(built) — the mount map, danos's fstab: an OPTIONAL volume identity → mount prefix override (a volume with no row mounts at its default/volumes/<id>), keyed on content identity and never on port, path, or arrival order (the lesson of Linux's/dev/sda1-era fstab, which broke on every port move untilUUID=replaced it). Identity is read off the medium by the prober, strongest first: GPT partition GUID → filesystem UUID → FAT serial + label → MBR signature + partition index → anonymous (generated name, no persistence). Same identity, same mount point: a drive moved to another port — USB to another hub, SATA to another bay, even a stick returning in a different dock — lands exactly where it was. Duplicate identity (cloned sticks, together) is policy: first keeps the name, the second mounts suffixed and is logged loudly. The boot volume is the recorded identity of the volume carrying/system/configurationand/system/logs, findable on any port. Every volume's default mount is/volumes/<id>— its rendered content identity.
Filesystem service (the FAT service today; one process per volume): the
proven unit — block-client + engine + file-protocol provider in one binary. It
receives its block channel at spawn; it never discovers devices. It registers
its own mounts with the kernel; its write cache lives inside the process, so a
write error is observed by the code that owns the volume and surfaces on the
owning channel (the anti-fsyncgate rule — never a system-wide dirty pool).
(Built:) fat receives its mount path as argv[2] from the volume manager (the
volume's id-path, e.g. /volumes/fat-12345678) and mounts its root there, plus
the two /system hierarchy rewrites it installs in place (unconditional this
increment; S3 makes them content-conditional across volumes). It no longer
self-acquires a volume — the V3b flip made it receive its volume id and block
channel from the volume manager, consistent with "it never discovers devices"
above.
Kernel (mechanism only): the mount table routes paths to backend
endpoints — resolve and redirect, never data. Remount-replace is the restart
story; a dead backend's slot is swept lazily on the next resolution.
fs_unmount is ownership-gated (built, V0) — only the mounting task may
unmount its own prefix; anyone else is refused (EPERM). No dead-owner
exception: a dead owner's slot is swept lazily by resolution, and restart goes
through remount-replace, never through a stranger's unmount.
Adding a filesystem
- Write the engine: pure Zig, no IPC, behind the
BlockDevicevtable (four functions). On-disk layout inalign(1)extern structs. Host-test it withzig build testfixtures — the FAT engine (engine.zig, on-disk.zig) is the template, and its ~500 lines of host tests the standard. - Reuse the shell: the filesystem harness — establishment, the
badge-scoped open-node table, the nine vfs-protocol handlers, mount
registration, the removal path — is shared code, not per-filesystem code.
(Done: extracted from fat's original 434-line shell into
library/kernel/file-system-harness.zig; fat imports it and instantiatesharness.Server(engine.FileSystem).) An engine plus amainwiring it into the harness is a complete filesystem service. - Add the configuration row: one line in
filesystems.csvmapping the on-disk signature to the binary. No other component changes: the volume manager probes, matches, spawns; the vfs protocol is already backend-neutral (a second provider serves it in production today — init's synthetic registry backend). - Prove removal: extend the removable-media suite (below) for the new filesystem — surprise-yank during writes must lose only what was unflushed, said honestly, with the volume consistent enough to remount.
The engine never sees: partitions (it receives a volume-shaped block channel; base offsets are the driver's clamp), device discovery (the channel arrives at spawn), mount policy (prefixes are handed to it), or other volumes (one process, one volume).
Removable media: responsibilities on removal, per layer
The design rule, learned from what Linux cannot do: there is exactly one
surprise-removal path — kill the filesystem process, retire its mounts,
respawn on return. No half-alive states, no remount-ro, no mounts that
error forever (Plan 9's dead-server wart).
The path has two triggers, one lifecycle: the device leaving (the
storage driver dies — channel death, the table below), and the medium
leaving while the device stays (an SD card pulled from its reader, an ATAPI
tray opened — including USB card readers today). The second trigger is the
pushed medium_changed event on the block protocol — published today from a
TEST UNIT READY poll; still planned is the volume manager consuming it
(today removal is driven only by device-presence polling) and translating the
transport's native signal (SCSI UNIT ATTENTION, AHCI PxSSTS, NVMe
namespace-change AER) in place of the poll. On the event the volume manager
runs the same kill-retire path, then re-probes on medium return exactly as on
device return. Without it, a swapped card would be served with the previous
card's filesystem state.
| Layer | Observes | Must do | Guarantees |
|---|---|---|---|
| Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report child_removed per interface |
the device tree is honest within one reconcile tick |
| Device manager | child_removed / reporter death |
prune the child; reap the bound driver (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
| Volume manager (removal built; remount bench-pending) | the storage device leaving the device-manager tree (poll) | kill the filesystem service of that device's volume; its kernel mounts retire | one removal path; mounts never dangle; log persistence stops cleanly |
| Filesystem service | its block channel dies (EPEER) mid-operation, or it is killed by the volume manager |
if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is lost and said to be lost | the unflushed write-back window is dropped on a surprise yank — danos writes no on-disk dirty/clean-shutdown marker today; the process never serves from behind a dead channel |
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated fs_unmount (built, V0) |
resolution under a dead mount is not_found, not a hang |
| Application | not_found / error on paths under the vanished mount |
its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
On return the same table runs upward in reverse: the bus re-enumerates and
re-reports; the device manager respawns the storage driver (built); the volume
manager re-probes — same content, same volume identity — respawns the
filesystem service, and remounts at the same prefix; applications see the
subtree reappear. A different stick in the same port is a different volume
(identity is content, not port) and mounts wherever its identity says —
possibly nowhere but /volumes/….
The boot volume: identity is what makes yanking it survivable
The boot volume is the sharpest instance of the return story, because the
system's own log persistence rides it (/system/logs), and because nothing
else about the running system depends on it at all — every binary and the
boot-time configuration live in the initial ramdisk. Its lifecycle,
end to end:
- At first mount the volume manager records the identity of the volume
carrying
/system/configurationand/system/logs— THE boot volume, from then on a fact about content, not about a port. - While it is absent, the system runs on. The kernel log ring keeps
accumulating — it is the buffer that makes the absence survivable — and
the logger keeps draining it; only persistence pauses. File operations
under the retired mounts fail honestly (
not_found). The ring is bounded, so a long absence overwrites its oldest entries: that window is the data loss, and it must be said — the logger marks the gap in the file when persistence resumes, never splicing the stream silently. - On return — any port, any hub, even a different transport — the
prober reads the same identity, the mount map answers with the same
prefixes, the filesystem service is respawned, and
/system/logsis the same tree it was: the logger resumes appending into the SAME<boot-stamp>directory, per-binary files continuing where they left off (plus the gap marker if the ring wrapped). - What never comes back is the write-back window lost at the yank — the dirty-honesty rule, unchanged; the boot volume gets no exemption.
This is also the requirement that shapes the logger: it treats the log tree
as a volume that comes and goes — failed flushes are retried on the same
patient cadence the fat service already uses for storage that arrives late,
never abandoned after the first not_found.
What is lost on a surprise yank is exactly the write-back window of the filesystem service, no more: the engine owns its cache, so the blast radius of a yank is one volume's unflushed writes, which are simply lost — danos records no on-disk dirty/clean-shutdown bit yet. A filesystem format with better crash honesty (journaling, copy-on-write — the lesson of QNX's Power-Safe) narrows that window further and slots in as implementation N+1 through the table above, changing nothing else.
How the lifecycle is enforced
Responsibility tables are convention; convention is worthless (the bounds audit's lesson). The volume lifecycle is enforced with the same three levers the device lifecycle already uses, mapped one-to-one:
- A filesystem cannot acquire — it can only be given. Filesystem
binaries hold NO establishment grants: no
open device-manager, no registry name to look up. The only block channel a filesystem process ever has is the one handed to it at spawn by the volume manager. Serving the wrong volume, a second volume, or a self-discovered volume is not forbidden but impossible — the same way a driver cannot claim hardware it was not delegated (device-authority.md). - The volume manager supervises with teeth. Mount-within-deadline or be
stopped (a filesystem wedged on a corrupt volume is killed and crash-loop
capped; the volume is marked bad, not retried forever). On removal the
manager KILLS the filesystem process and retires its mounts — the polite
observe-
EPEER-and-exit path is an optimization; the kill is the guarantee, because a process cannot be trusted to observe its own obsolescence (the reap argument, proven on the device tree). - The shared harness is how every filesystem inherits the lifecycle by construction. The harness — not the engine — owns the state machine: establishment at spawn, mount registration, the flush-on-close hook (an in-memory device-dirty check that commits the device write cache on close), error-out-and-exit on channel death. The engine sits behind the four-function vtable and never sees a channel; it cannot opt out of the lifecycle for the same reason it cannot find a device. This is why the harness is extracted BEFORE the second engine is written.
Beneath all three, two kernel mechanisms: ownership-gated fs_unmount (only
the mounting endpoint's holder may unmount), and the lazy dead-endpoint mount
sweep as the backstop nothing can disable. And above them, the check: a
lifecycle conformance drill — mount, serve, yank mid-write, verify honest
loss, replug, remount — parameterized over filesystem implementations, so
"danos supports filesystem X" MEANS "X passes the drill through the harness",
exactly as provider conformance means passing the reserved-verb suite.