Files
danos/docs/file-system-development/storage-architecture.md
T
Daniel Samson 6d4992ae02 docs: storage — the mount map is built; the path is the id (S2)
Flip the storage docs to match S2: filesystems.csv (signature -> binary) and
volumes.csv (identity -> optional override) are built; a volume's mount path IS
its content id (/volumes/<id>), never a port name; the label is display metadata
a `volumes` query returns. fat receives its mount path via argv[2] rather than
hardcoding it. Still pending: rung 2 (filesystem UUID, needs a non-FAT engine),
multi-volume (fat's boot rewrites stay unconditional until S3), medium_changed
consumption, remount bench-verification.
2026-08-10 00:49:32 +01:00

18 KiB
Raw Blame History

The storage architecture: layers, boundaries, responsibilities

Status: the layered model below is the settled design (storage-design-rationale.md records how it was reached, and volume-manager-plan.md how it was built). Built (V0–V4 + the storage-stack S1/S2): the data path, the driver range confinement (per-sender clamp + the confinement gate), the medium_changed presence event, the volume manager itself — it probes the partition table, confines each filesystem to its partition, spawns one filesystem per volume, and supervises it — the removal half of the lifecycle (a pulled stick unmounts), the identity ladder (GPT GUID + name, FAT serial + label, MBR), and the mount map: filesystems.csv (signature → binary) + volumes.csv (identity → optional override), a volume's mount path IS its content id (/volumes/<id>), with the label as display metadata a volumes query returns. Still pending: the filesystem UUID rung (needs a non-FAT engine), multi-volume (one FAT volume today; fat's boot rewrites are unconditional until S3 makes them content-conditional), the volume manager consuming medium_changed (removal is detected by device-presence polling; the event is published but only a card-reader medium change needs the subscription), and the remount-on-replug end-to-end (the logic is in place; QEMU can't re-present the boot-controller device, so it is bench-verified). A few markers below are left where a duty is still pending.

The model

One pattern, applied twice: an application is a client of a service over a protocol; a service is a client of a driver over a protocol.

application
    │  vfs protocol            (routed by the kernel mount table)
    ▼
filesystem service  ──  one process per volume
    │                   inside: [file-protocol shell] ── Zig API ── [fs engine]
    │  block protocol          (channel received at spawn)
    ▼
storage driver      ──  one process per device (usb-storage per stick)
    │  usb-transfer protocol   (channel received via hello, by lineage)
    ▼
bus driver          ──  one process per controller (usb-xhci-bus)
    │  hardware

Three kinds of boundary, deliberately different:

  • Protocol boundaries sit where processes meet (vfs, block, usb-transfer). Each side is independently restartable; channels are established by capability handoff, never by registry names (communication.md "Establishment: two planes"). Steady-state file I/O crosses exactly two — vfs and block.
  • Library boundaries sit where concerns meet inside one process. The filesystem service's engine (on-disk format logic, host-testable, behind the four-function BlockDevice vtable) and shell (IPC serving, open-node table, mount registration) talk through a Zig API. No protocol between them: they share fate regardless — a corrupt engine corrupts the answers either way — so a channel there would add a hop per file operation and buy nothing. The seam exists for compile-time testability and reuse; the process is the restart unit.
  • Control-plane relationships sit beside the data path, never on it. Two supervisors, one per layer: the device manager wires and revives the device layers (bus and storage drivers — devices only); the volume manager (built) wires and revives the volume layer (filesystem services). Neither touches steady-state I/O.

Who does what

Bus driver (usb-xhci-bus, per controller): hardware only. Enumerates, reports children to the device manager, serves the transfer contract to its own children's class drivers, routed by lineage.

Storage driver (usb-storage, per device; later nvme, ahci): hardware → blocks. Speaks its transport upward from its device; serves the block protocol (geometry, read, write, flush, attach, detach — and the pushed medium_changed presence event, published today from a slow TEST UNIT READY poll; planned: translating it from the transport's native signal instead of polling). Content-blind, permanently: it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no GPT, no filesystem magic, ever. It has one content-blind mechanism (built): named sub-ranges — "serve blocks [a, b) as a channel of the same block contract". It clamps and offsets; it never knows the numbers came from a partition table. The clamp lives here and nowhere else because a channel must carry exactly the authority it grants: handing a filesystem the whole disk plus a polite base offset would let a buggy or compromised filesystem scribble the neighboring partition — the same authority-overshoot the device-authority track eliminated for MMIO and DMA.

Volume manager (built; system/services/volume-manager): the policy home of the volume layer, one service, supervised by init. It watches the device manager's tree for a storage provider; when one appears it consumer-hellos for the block channel, reads the partition table and the first blocks itself (it is the prober), defines the volume's sub-range on the driver, spawns the matching filesystem service confined to that range, and supervises it (backoff, crash-loop cap). (Built): it picks the filesystem binary from filesystems.csv by the volume's content signature, and mounts the volume at its content id (/volumes/<id>) — or a volumes.csv override. The label is display metadata the volumes query returns, never the path. Those tables are CSV configuration, read by it (the policy), enforced by nobody else:

  • filesystems.csv (built) — content signature → filesystem binary. Adding a filesystem adds a row.
  • volumes.csv (built) — the mount map, danos's fstab: an OPTIONAL volume identity → mount prefix override (a volume with no row mounts at its default /volumes/<id>), keyed on content identity and never on port, path, or arrival order (the lesson of Linux's /dev/sda1-era fstab, which broke on every port move until UUID= replaced it). Identity is read off the medium by the prober, strongest first: GPT partition GUID → filesystem UUID → FAT serial + label → MBR signature + partition index → anonymous (generated name, no persistence). Same identity, same mount point: a drive moved to another port — USB to another hub, SATA to another bay, even a stick returning in a different dock — lands exactly where it was. Duplicate identity (cloned sticks, together) is policy: first keeps the name, the second mounts suffixed and is logged loudly. The boot volume is the recorded identity of the volume carrying /system/configuration and /system/logs, findable on any port. Every volume's default mount is /volumes/<id> — its rendered content identity.

Filesystem service (the FAT service today; one process per volume): the proven unit — block-client + engine + file-protocol provider in one binary. It receives its block channel at spawn; it never discovers devices. It registers its own mounts with the kernel; its write cache lives inside the process, so a write error is observed by the code that owns the volume and surfaces on the owning channel (the anti-fsyncgate rule — never a system-wide dirty pool). (Built:) fat receives its mount path as argv[2] from the volume manager (the volume's id-path, e.g. /volumes/fat-12345678) and mounts its root there, plus the two /system hierarchy rewrites it installs in place (unconditional this increment; S3 makes them content-conditional across volumes). It no longer self-acquires a volume — the V3b flip made it receive its volume id and block channel from the volume manager, consistent with "it never discovers devices" above.

Kernel (mechanism only): the mount table routes paths to backend endpoints — resolve and redirect, never data. Remount-replace is the restart story; a dead backend's slot is swept lazily on the next resolution. fs_unmount is ownership-gated (built, V0) — only the mounting task may unmount its own prefix; anyone else is refused (EPERM). No dead-owner exception: a dead owner's slot is swept lazily by resolution, and restart goes through remount-replace, never through a stranger's unmount.

Adding a filesystem

  1. Write the engine: pure Zig, no IPC, behind the BlockDevice vtable (four functions). On-disk layout in align(1) extern structs. Host-test it with zig build test fixtures — the FAT engine (engine.zig, on-disk.zig) is the template, and its ~500 lines of host tests the standard.
  2. Reuse the shell: the filesystem harness — establishment, the badge-scoped open-node table, the nine vfs-protocol handlers, mount registration, the removal path — is shared code, not per-filesystem code. (Done: extracted from fat's original 434-line shell into library/kernel/file-system-harness.zig; fat imports it and instantiates harness.Server(engine.FileSystem).) An engine plus a main wiring it into the harness is a complete filesystem service.
  3. Add the configuration row: one line in filesystems.csv mapping the on-disk signature to the binary. No other component changes: the volume manager probes, matches, spawns; the vfs protocol is already backend-neutral (a second provider serves it in production today — init's synthetic registry backend).
  4. Prove removal: extend the removable-media suite (below) for the new filesystem — surprise-yank during writes must lose only what was unflushed, said honestly, with the volume consistent enough to remount.

The engine never sees: partitions (it receives a volume-shaped block channel; base offsets are the driver's clamp), device discovery (the channel arrives at spawn), mount policy (prefixes are handed to it), or other volumes (one process, one volume).

Removable media: responsibilities on removal, per layer

The design rule, learned from what Linux cannot do: there is exactly one surprise-removal path — kill the filesystem process, retire its mounts, respawn on return. No half-alive states, no remount-ro, no mounts that error forever (Plan 9's dead-server wart).

The path has two triggers, one lifecycle: the device leaving (the storage driver dies — channel death, the table below), and the medium leaving while the device stays (an SD card pulled from its reader, an ATAPI tray opened — including USB card readers today). The second trigger is the pushed medium_changed event on the block protocol — published today from a TEST UNIT READY poll; still planned is the volume manager consuming it (today removal is driven only by device-presence polling) and translating the transport's native signal (SCSI UNIT ATTENTION, AHCI PxSSTS, NVMe namespace-change AER) in place of the poll. On the event the volume manager runs the same kill-retire path, then re-probes on medium return exactly as on device return. Without it, a swapped card would be served with the previous card's filesystem state.

Layer Observes Must do Guarantees
Bus driver port/hub status change tear down the device's slots (children first, recursively — built, hot-plug matrix), report child_removed per interface the device tree is honest within one reconcile tick
Device manager child_removed / reporter death prune the child; reap the bound driver (built) — the storage driver for that stick dies now, not never no zombie storage processes; re-report rebinds
Storage driver its own death (it IS the removed device's driver) nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death in-flight transfers fail visibly to callers, never hang
Volume manager (removal built; remount bench-pending) the storage device leaving the device-manager tree (poll) kill the filesystem service of that device's volume; its kernel mounts retire one removal path; mounts never dangle; log persistence stops cleanly
Filesystem service its block channel dies (EPEER) mid-operation, or it is killed by the volume manager if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is lost and said to be lost the unflushed write-back window is dropped on a surprise yank — danos writes no on-disk dirty/clean-shutdown marker today; the process never serves from behind a dead channel
Kernel backend endpoint death lazy mount-slot sweep on next resolve (built); ownership-gated fs_unmount (built, V0) resolution under a dead mount is not_found, not a hang
Application not_found / error on paths under the vanished mount its own error handling — the contract is honest absence, identical to the path never existing no operation blocks forever on removed media

On return the same table runs upward in reverse: the bus re-enumerates and re-reports; the device manager respawns the storage driver (built); the volume manager re-probes — same content, same volume identity — respawns the filesystem service, and remounts at the same prefix; applications see the subtree reappear. A different stick in the same port is a different volume (identity is content, not port) and mounts wherever its identity says — possibly nowhere but /volumes/….

The boot volume: identity is what makes yanking it survivable

The boot volume is the sharpest instance of the return story, because the system's own log persistence rides it (/system/logs), and because nothing else about the running system depends on it at all — every binary and the boot-time configuration live in the initial ramdisk. Its lifecycle, end to end:

  • At first mount the volume manager records the identity of the volume carrying /system/configuration and /system/logs — THE boot volume, from then on a fact about content, not about a port.
  • While it is absent, the system runs on. The kernel log ring keeps accumulating — it is the buffer that makes the absence survivable — and the logger keeps draining it; only persistence pauses. File operations under the retired mounts fail honestly (not_found). The ring is bounded, so a long absence overwrites its oldest entries: that window is the data loss, and it must be said — the logger marks the gap in the file when persistence resumes, never splicing the stream silently.
  • On return — any port, any hub, even a different transport — the prober reads the same identity, the mount map answers with the same prefixes, the filesystem service is respawned, and /system/logs is the same tree it was: the logger resumes appending into the SAME <boot-stamp> directory, per-binary files continuing where they left off (plus the gap marker if the ring wrapped).
  • What never comes back is the write-back window lost at the yank — the dirty-honesty rule, unchanged; the boot volume gets no exemption.

This is also the requirement that shapes the logger: it treats the log tree as a volume that comes and goes — failed flushes are retried on the same patient cadence the fat service already uses for storage that arrives late, never abandoned after the first not_found.

What is lost on a surprise yank is exactly the write-back window of the filesystem service, no more: the engine owns its cache, so the blast radius of a yank is one volume's unflushed writes, which are simply lost — danos records no on-disk dirty/clean-shutdown bit yet. A filesystem format with better crash honesty (journaling, copy-on-write — the lesson of QNX's Power-Safe) narrows that window further and slots in as implementation N+1 through the table above, changing nothing else.

How the lifecycle is enforced

Responsibility tables are convention; convention is worthless (the bounds audit's lesson). The volume lifecycle is enforced with the same three levers the device lifecycle already uses, mapped one-to-one:

  1. A filesystem cannot acquire — it can only be given. Filesystem binaries hold NO establishment grants: no open device-manager, no registry name to look up. The only block channel a filesystem process ever has is the one handed to it at spawn by the volume manager. Serving the wrong volume, a second volume, or a self-discovered volume is not forbidden but impossible — the same way a driver cannot claim hardware it was not delegated (device-authority.md).
  2. The volume manager supervises with teeth. Mount-within-deadline or be stopped (a filesystem wedged on a corrupt volume is killed and crash-loop capped; the volume is marked bad, not retried forever). On removal the manager KILLS the filesystem process and retires its mounts — the polite observe-EPEER-and-exit path is an optimization; the kill is the guarantee, because a process cannot be trusted to observe its own obsolescence (the reap argument, proven on the device tree).
  3. The shared harness is how every filesystem inherits the lifecycle by construction. The harness — not the engine — owns the state machine: establishment at spawn, mount registration, the flush-on-close hook (an in-memory device-dirty check that commits the device write cache on close), error-out-and-exit on channel death. The engine sits behind the four-function vtable and never sees a channel; it cannot opt out of the lifecycle for the same reason it cannot find a device. This is why the harness is extracted BEFORE the second engine is written.

Beneath all three, two kernel mechanisms: ownership-gated fs_unmount (only the mounting endpoint's holder may unmount), and the lazy dead-endpoint mount sweep as the backstop nothing can disable. And above them, the check: a lifecycle conformance drill — mount, serve, yank mid-write, verify honest loss, replug, remount — parameterized over filesystem implementations, so "danos supports filesystem X" MEANS "X passes the drill through the harness", exactly as provider conformance means passing the reserved-verb suite.