# The storage architecture: layers, boundaries, responsibilities > **Status:** the layered model below is the settled design > ([storage-stack-discussion.md](../storage-stack-discussion.md) records how it > was reached and what the surveyed systems taught). The data path — vfs > protocol, kernel mount routing, the FAT service, the block protocol, > usb-storage — is **built**. The volume manager, the driver's range > mechanism, per-volume filesystem spawning, and the removal lifecycle are > **planned**; until they land, the FAT service performs volume-manager duties > itself (marked below). This document is the reference for both states. ## The model One pattern, applied twice: an application is a client of a service over a protocol; a service is a client of a driver over a protocol. ``` application │ vfs protocol (routed by the kernel mount table) ▼ filesystem service ── one process per volume │ inside: [file-protocol shell] ── Zig API ── [fs engine] │ block protocol (channel received at spawn) ▼ storage driver ── one process per device (usb-storage per stick) │ usb-transfer protocol (channel received via hello, by lineage) ▼ bus driver ── one process per controller (usb-xhci-bus) │ hardware ``` Three kinds of boundary, deliberately different: - **Protocol boundaries** sit where processes meet (vfs, block, usb-transfer). Each side is independently restartable; channels are established by capability handoff, never by registry names (communication.md "Establishment: two planes"). Steady-state file I/O crosses exactly two — vfs and block. - **Library boundaries** sit where concerns meet inside one process. The filesystem service's engine (on-disk format logic, host-testable, behind the four-function `BlockDevice` vtable) and shell (IPC serving, open-node table, mount registration) talk through a Zig API. No protocol between them: they share fate regardless — a corrupt engine corrupts the answers either way — so a channel there would add a hop per file operation and buy nothing. The seam exists for compile-time testability and reuse; the *process* is the restart unit. - **Control-plane relationships** sit beside the data path, never on it. Two supervisors, one per layer: the **device manager** wires and revives the device layers (bus and storage drivers — devices only); the **volume manager** *(planned)* wires and revives the volume layer (filesystem services). Neither touches steady-state I/O. ## Who does what **Bus driver** (usb-xhci-bus, per controller): hardware only. Enumerates, reports children to the device manager, serves the transfer contract to its own children's class drivers, routed by lineage. **Storage driver** (usb-storage, per device): hardware → blocks. Speaks SCSI Bulk-Only Transport upward from its device; serves the block protocol (five verbs: geometry, read, write, flush, attach). **Content-blind, permanently**: it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no GPT, no filesystem magic, ever. *(Planned)* it gains one content-blind mechanism: **named sub-ranges** — "serve blocks [a, b) as a channel of the same block contract". It clamps and offsets; it never knows the numbers came from a partition table. The clamp lives here and nowhere else because a channel must carry exactly the authority it grants: handing a filesystem the whole disk plus a polite base offset would let a buggy or compromised filesystem scribble the neighboring partition — the same authority-overshoot the device-authority track eliminated for MMIO and DMA. **Volume manager** *(planned; today the FAT service squats on these duties)*: the policy home of the volume layer, one service, supervised by init. It subscribes to the device manager's child events; when a storage provider appears it consumer-hellos for the block channel, reads the partition table and the first blocks itself (**it** is the prober), consults its configuration, defines sub-ranges on the driver, spawns the matching filesystem service per volume with that volume's channel, supervises it, and decides mount placement. Its tables are CSV configuration, read by it (the policy), enforced by nobody else: - `filesystems.csv` — content signature → filesystem binary. Adding a filesystem adds a row. - the mount map — volume identity → mount prefix. The boot volume is chosen by **content** (the volume carrying `/system/configuration` and `/system/logs`), never by port or arrival order. **Filesystem service** (the FAT service today; one process per volume): the proven unit — block-client + engine + file-protocol provider in one binary. It receives its block channel at spawn; it never discovers devices. It registers its own mounts with the kernel; its write cache lives inside the process, so a write error is observed by the code that owns the volume and surfaces on the owning channel (the anti-fsyncgate rule — never a system-wide dirty pool). *(Today, interim:)* fat also acquires its own volume (first mass-storage child by enumeration order) and hardcodes its mount prefixes; both migrate to the volume manager. **Kernel** (mechanism only): the mount table routes paths to backend endpoints — resolve and redirect, never data. Remount-replace is the restart story; a dead backend's slot is swept lazily on the next resolution. *(Planned fix:)* `fs_unmount` gains ownership — only the mounting endpoint's holder may unmount; today it is ungated, which is safe with one mount owner and wrong with several. ## Adding a filesystem 1. **Write the engine**: pure Zig, no IPC, behind the `BlockDevice` vtable (four functions). On-disk layout in `align(1)` extern structs. Host-test it with `zig build test` fixtures — the FAT engine ([engine.zig](../../system/services/fat/engine.zig), [on-disk.zig](../../system/services/fat/on-disk.zig)) is the template, and its ~500 lines of host tests the standard. 2. **Reuse the shell**: the filesystem harness — establishment, the badge-scoped open-node table, the nine vfs-protocol handlers, mount registration, the removal path — is shared code, not per-filesystem code. *(Planned: extracted from fat's 434-line shell into a library before the second engine is written.)* An engine plus a `main` wiring it into the harness is a complete filesystem service. 3. **Add the configuration row**: one line in `filesystems.csv` mapping the on-disk signature to the binary. No other component changes: the volume manager probes, matches, spawns; the vfs protocol is already backend-neutral (a second provider serves it in production today — init's synthetic registry backend). 4. **Prove removal**: extend the removable-media suite (below) for the new filesystem — surprise-yank during writes must lose only what was unflushed, said honestly, with the volume consistent enough to remount. The engine never sees: partitions (it receives a volume-shaped block channel; base offsets are the driver's clamp), device discovery (the channel arrives at spawn), mount policy (prefixes are handed to it), or other volumes (one process, one volume). ## Removable media: responsibilities on removal, per layer The design rule, learned from what Linux cannot do: there is exactly **one** surprise-removal path — kill the filesystem process, retire its mounts, respawn on return. No half-alive states, no `remount-ro`, no mounts that error forever (Plan 9's dead-server wart). | Layer | Observes | Must do | Guarantees | |---|---|---|---| | Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick | | Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds | | Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang | | Volume manager *(planned)* | the storage provider's channel death / the manager's child events | kill each filesystem service of that device's volumes; retire their kernel mounts; remember the volume identity | one removal path; mounts never dangle; log persistence stops *cleanly* | | Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the FAT dirty flag on disk marks the unclean removal; the process never serves from behind a dead channel | | Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); *(planned)* ownership-gated `fs_unmount` | resolution under a dead mount is `not_found`, not a hang | | Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media | **On return** the same table runs upward in reverse: the bus re-enumerates and re-reports; the device manager respawns the storage driver (built); the volume manager re-probes — same content, same volume identity — respawns the filesystem service, and remounts at the same prefix; applications see the subtree reappear. A different stick in the same port is a *different volume* (identity is content, not port) and mounts wherever its identity says — possibly nowhere but `/volumes/…`. What is lost on a surprise yank is exactly the write-back window of the filesystem service, no more: the engine owns its cache, so the blast radius of a yank is one volume's unflushed writes, marked by the on-disk dirty flag. A filesystem format with better crash honesty (journaling, copy-on-write — the lesson of QNX's Power-Safe) narrows that window further and slots in as implementation N+1 through the table above, changing nothing else. ## How the lifecycle is enforced Responsibility tables are convention; convention is worthless (the bounds audit's lesson). The volume lifecycle is enforced with the same three levers the device lifecycle already uses, mapped one-to-one: 1. **A filesystem cannot acquire — it can only be given.** Filesystem binaries hold NO establishment grants: no `open device-manager`, no registry name to look up. The only block channel a filesystem process ever has is the one handed to it at spawn by the volume manager. Serving the wrong volume, a second volume, or a self-discovered volume is not forbidden but *impossible* — the same way a driver cannot claim hardware it was not delegated (device-authority.md). 2. **The volume manager supervises with teeth.** Mount-within-deadline or be stopped (a filesystem wedged on a corrupt volume is killed and crash-loop capped; the volume is marked bad, not retried forever). On removal the manager KILLS the filesystem process and retires its mounts — the polite observe-`EPEER`-and-exit path is an optimization; the kill is the guarantee, because a process cannot be trusted to observe its own obsolescence (the reap argument, proven on the device tree). 3. **The shared harness is how every filesystem inherits the lifecycle by construction.** The harness — not the engine — owns the state machine: establishment at spawn, mount registration, the dirty-flag set/clear bracket, error-out-and-exit on channel death. The engine sits behind the four-function vtable and never sees a channel; it cannot opt out of the lifecycle for the same reason it cannot find a device. This is why the harness is extracted BEFORE the second engine is written. Beneath all three, two kernel mechanisms: ownership-gated `fs_unmount` (only the mounting endpoint's holder may unmount), and the lazy dead-endpoint mount sweep as the backstop nothing can disable. And above them, the check: a **lifecycle conformance drill** — mount, serve, yank mid-write, verify honest loss, replug, remount — parameterized over filesystem implementations, so "danos supports filesystem X" MEANS "X passes the drill through the harness", exactly as provider conformance means passing the reserved-verb suite.