# The storage architecture: layers, boundaries, responsibilities > **Status:** the layered model below is the settled design > ([storage-design-rationale.md](storage-design-rationale.md) records how it was > reached, and [volume-manager-plan.md](../volume-manager-plan.md) how it was > built). **Built** (the volume-manager track, V0–V4): the data path, the driver > range confinement (per-sender clamp + the confinement gate), the `medium_changed` > presence event, the volume manager itself — it probes the partition table, > confines each filesystem to its partition, spawns one filesystem per volume, and > supervises it — and the removal half of the lifecycle (a pulled stick unmounts). > **Still pending**: the fuller identity ladder and the `volumes.csv` mount map, > multi-volume (one FAT volume today), the volume manager *consuming* > `medium_changed` (removal is detected by device-presence polling; the event is > published but only a card-reader medium change needs the subscription), and the > remount-on-replug end-to-end (the logic is in place; QEMU can't re-present the > boot-controller device, so it is bench-verified). A few markers below are left > where a duty is still pending. ## The model One pattern, applied twice: an application is a client of a service over a protocol; a service is a client of a driver over a protocol. ``` application │ vfs protocol (routed by the kernel mount table) ▼ filesystem service ── one process per volume │ inside: [file-protocol shell] ── Zig API ── [fs engine] │ block protocol (channel received at spawn) ▼ storage driver ── one process per device (usb-storage per stick) │ usb-transfer protocol (channel received via hello, by lineage) ▼ bus driver ── one process per controller (usb-xhci-bus) │ hardware ``` Three kinds of boundary, deliberately different: - **Protocol boundaries** sit where processes meet (vfs, block, usb-transfer). Each side is independently restartable; channels are established by capability handoff, never by registry names (communication.md "Establishment: two planes"). Steady-state file I/O crosses exactly two — vfs and block. - **Library boundaries** sit where concerns meet inside one process. The filesystem service's engine (on-disk format logic, host-testable, behind the four-function `BlockDevice` vtable) and shell (IPC serving, open-node table, mount registration) talk through a Zig API. No protocol between them: they share fate regardless — a corrupt engine corrupts the answers either way — so a channel there would add a hop per file operation and buy nothing. The seam exists for compile-time testability and reuse; the *process* is the restart unit. - **Control-plane relationships** sit beside the data path, never on it. Two supervisors, one per layer: the **device manager** wires and revives the device layers (bus and storage drivers — devices only); the **volume manager** *(built)* wires and revives the volume layer (filesystem services). Neither touches steady-state I/O. ## Who does what **Bus driver** (usb-xhci-bus, per controller): hardware only. Enumerates, reports children to the device manager, serves the transfer contract to its own children's class drivers, routed by lineage. **Storage driver** (usb-storage, per device; later nvme, ahci): hardware → blocks. Speaks its transport upward from its device; serves the block protocol (geometry, read, write, flush, attach, detach — and the pushed `medium_changed` presence event, published today from a slow TEST UNIT READY poll; *planned*: translating it from the transport's native signal instead of polling). **Content-blind, permanently**: it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no GPT, no filesystem magic, ever. It has one content-blind mechanism *(built)*: **named sub-ranges** — "serve blocks [a, b) as a channel of the same block contract". It clamps and offsets; it never knows the numbers came from a partition table. The clamp lives here and nowhere else because a channel must carry exactly the authority it grants: handing a filesystem the whole disk plus a polite base offset would let a buggy or compromised filesystem scribble the neighboring partition — the same authority-overshoot the device-authority track eliminated for MMIO and DMA. **Volume manager** *(built; `system/services/volume-manager`)*: the policy home of the volume layer, one service, supervised by init. It watches the device manager's tree for a storage provider; when one appears it consumer-hellos for the block channel, reads the partition table and the first blocks itself (**it** is the prober), defines the volume's sub-range on the driver, spawns the matching filesystem service confined to that range, and supervises it (backoff, crash-loop cap). *(Pending)*: it decides mount placement from `volumes.csv` and picks the filesystem binary from `filesystems.csv` — today it hands every FAT-shaped volume to the FAT service and the FAT service carries hardcoded mount prefixes. Those tables are CSV configuration, read by it (the policy), enforced by nobody else: - `filesystems.csv` *(pending)* — content signature → filesystem binary. Adding a filesystem adds a row. - `volumes.csv` *(pending)* — the mount map, danos's fstab: **volume identity → mount prefix**, keyed on content identity and never on port, path, or arrival order (the lesson of Linux's `/dev/sda1`-era fstab, which broke on every port move until `UUID=` replaced it). Identity is read off the medium by the prober, strongest first: GPT partition GUID → filesystem UUID → FAT serial + label → MBR signature + partition index → anonymous (generated name, no persistence). Same identity, same mount point: a drive moved to another port — USB to another hub, SATA to another bay, even a stick returning in a different dock — lands exactly where it was. Duplicate identity (cloned sticks, together) is policy: first keeps the name, the second mounts suffixed and is logged loudly. The boot volume is the recorded identity of the volume carrying `/system/configuration` and `/system/logs`, findable on any port. Unknown volumes mount under `/volumes/`. **Filesystem service** (the FAT service today; one process per volume): the proven unit — block-client + engine + file-protocol provider in one binary. It receives its block channel at spawn; it never discovers devices. It registers its own mounts with the kernel; its write cache lives inside the process, so a write error is observed by the code that owns the volume and surfaces on the owning channel (the anti-fsyncgate rule — never a system-wide dirty pool). *(Today, interim:)* fat still hardcodes its mount prefixes (`/volumes/usb` plus the two boot-volume hierarchy subtrees it rewrites in place); a `volumes.csv` mount map will migrate that to the volume manager. It no longer self-acquires a volume — the V3b flip made it receive its volume id at spawn and its block channel from the volume manager's hello reply, consistent with "it never discovers devices" above. **Kernel** (mechanism only): the mount table routes paths to backend endpoints — resolve and redirect, never data. Remount-replace is the restart story; a dead backend's slot is swept lazily on the next resolution. `fs_unmount` is ownership-gated *(built, V0)* — only the mounting task may unmount its own prefix; anyone else is refused (`EPERM`). No dead-owner exception: a dead owner's slot is swept lazily by resolution, and restart goes through remount-replace, never through a stranger's unmount. ## Adding a filesystem 1. **Write the engine**: pure Zig, no IPC, behind the `BlockDevice` vtable (four functions). On-disk layout in `align(1)` extern structs. Host-test it with `zig build test` fixtures — the FAT engine ([engine.zig](../../system/services/fat/engine.zig), [on-disk.zig](../../system/services/fat/on-disk.zig)) is the template, and its ~500 lines of host tests the standard. 2. **Reuse the shell**: the filesystem harness — establishment, the badge-scoped open-node table, the nine vfs-protocol handlers, mount registration, the removal path — is shared code, not per-filesystem code. *(Done: extracted from fat's original 434-line shell into `library/kernel/file-system-harness.zig`; fat imports it and instantiates `harness.Server(engine.FileSystem)`.)* An engine plus a `main` wiring it into the harness is a complete filesystem service. 3. **Add the configuration row**: one line in `filesystems.csv` mapping the on-disk signature to the binary. No other component changes: the volume manager probes, matches, spawns; the vfs protocol is already backend-neutral (a second provider serves it in production today — init's synthetic registry backend). 4. **Prove removal**: extend the removable-media suite (below) for the new filesystem — surprise-yank during writes must lose only what was unflushed, said honestly, with the volume consistent enough to remount. The engine never sees: partitions (it receives a volume-shaped block channel; base offsets are the driver's clamp), device discovery (the channel arrives at spawn), mount policy (prefixes are handed to it), or other volumes (one process, one volume). ## Removable media: responsibilities on removal, per layer The design rule, learned from what Linux cannot do: there is exactly **one** surprise-removal path — kill the filesystem process, retire its mounts, respawn on return. No half-alive states, no `remount-ro`, no mounts that error forever (Plan 9's dead-server wart). The path has **two triggers, one lifecycle**: the *device* leaving (the storage driver dies — channel death, the table below), and the *medium* leaving while the device stays (an SD card pulled from its reader, an ATAPI tray opened — including USB card readers today). The second trigger is the pushed `medium_changed` event on the block protocol — published today from a TEST UNIT READY poll; still *planned* is the volume manager *consuming* it (today removal is driven only by device-presence polling) and translating the transport's native signal (SCSI UNIT ATTENTION, AHCI PxSSTS, NVMe namespace-change AER) in place of the poll. On the event the volume manager runs the same kill-retire path, then re-probes on medium return exactly as on device return. Without it, a swapped card would be served with the previous card's filesystem state. | Layer | Observes | Must do | Guarantees | |---|---|---|---| | Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick | | Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds | | Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang | | Volume manager *(removal built; remount bench-pending)* | the storage device leaving the device-manager tree (poll) | kill the filesystem service of that device's volume; its kernel mounts retire | one removal path; mounts never dangle; log persistence stops *cleanly* | | Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the unflushed write-back window is dropped on a surprise yank — danos writes no on-disk dirty/clean-shutdown marker today; the process never serves from behind a dead channel | | Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated `fs_unmount` (built, V0) | resolution under a dead mount is `not_found`, not a hang | | Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media | **On return** the same table runs upward in reverse: the bus re-enumerates and re-reports; the device manager respawns the storage driver (built); the volume manager re-probes — same content, same volume identity — respawns the filesystem service, and remounts at the same prefix; applications see the subtree reappear. A different stick in the same port is a *different volume* (identity is content, not port) and mounts wherever its identity says — possibly nowhere but `/volumes/…`. ### The boot volume: identity is what makes yanking it survivable The boot volume is the sharpest instance of the return story, because the system's own log persistence rides it (`/system/logs`), and because nothing else about the running system depends on it at all — every binary and the boot-time configuration live in the initial ramdisk. Its lifecycle, end to end: - **At first mount** the volume manager records the identity of the volume carrying `/system/configuration` and `/system/logs` — THE boot volume, from then on a fact about content, not about a port. - **While it is absent**, the system runs on. The kernel log ring keeps accumulating — it is the buffer that makes the absence survivable — and the logger keeps draining it; only *persistence* pauses. File operations under the retired mounts fail honestly (`not_found`). The ring is bounded, so a long absence overwrites its oldest entries: that window is the data loss, and it must be *said* — the logger marks the gap in the file when persistence resumes, never splicing the stream silently. - **On return — any port, any hub, even a different transport** — the prober reads the same identity, the mount map answers with the same prefixes, the filesystem service is respawned, and `/system/logs` is the same tree it was: the logger resumes appending into the SAME `` directory, per-binary files continuing where they left off (plus the gap marker if the ring wrapped). - **What never comes back** is the write-back window lost at the yank — the dirty-honesty rule, unchanged; the boot volume gets no exemption. This is also the requirement that shapes the logger: it treats the log tree as a volume that comes and goes — failed flushes are retried on the same patient cadence the fat service already uses for storage that arrives late, never abandoned after the first `not_found`. What is lost on a surprise yank is exactly the write-back window of the filesystem service, no more: the engine owns its cache, so the blast radius of a yank is one volume's unflushed writes, which are simply lost — danos records no on-disk dirty/clean-shutdown bit yet. A filesystem format with better crash honesty (journaling, copy-on-write — the lesson of QNX's Power-Safe) narrows that window further and slots in as implementation N+1 through the table above, changing nothing else. ## How the lifecycle is enforced Responsibility tables are convention; convention is worthless (the bounds audit's lesson). The volume lifecycle is enforced with the same three levers the device lifecycle already uses, mapped one-to-one: 1. **A filesystem cannot acquire — it can only be given.** Filesystem binaries hold NO establishment grants: no `open device-manager`, no registry name to look up. The only block channel a filesystem process ever has is the one handed to it at spawn by the volume manager. Serving the wrong volume, a second volume, or a self-discovered volume is not forbidden but *impossible* — the same way a driver cannot claim hardware it was not delegated (device-authority.md). 2. **The volume manager supervises with teeth.** Mount-within-deadline or be stopped (a filesystem wedged on a corrupt volume is killed and crash-loop capped; the volume is marked bad, not retried forever). On removal the manager KILLS the filesystem process and retires its mounts — the polite observe-`EPEER`-and-exit path is an optimization; the kill is the guarantee, because a process cannot be trusted to observe its own obsolescence (the reap argument, proven on the device tree). 3. **The shared harness is how every filesystem inherits the lifecycle by construction.** The harness — not the engine — owns the state machine: establishment at spawn, mount registration, the flush-on-close hook (an in-memory device-dirty check that commits the device write cache on close), error-out-and-exit on channel death. The engine sits behind the four-function vtable and never sees a channel; it cannot opt out of the lifecycle for the same reason it cannot find a device. This is why the harness is extracted BEFORE the second engine is written. Beneath all three, two kernel mechanisms: ownership-gated `fs_unmount` (only the mounting endpoint's holder may unmount), and the lazy dead-endpoint mount sweep as the backstop nothing can disable. And above them, the check: a **lifecycle conformance drill** — mount, serve, yank mid-write, verify honest loss, replug, remount — parameterized over filesystem implementations, so "danos supports filesystem X" MEANS "X passes the drill through the harness", exactly as provider conformance means passing the reserved-verb suite.