Files
danos/docs/storage-stack-discussion.md
T

8.6 KiB

The storage stack: block, volumes, filesystems — a discussion

2026-08-09. Design discussion, not a plan. The questions, verbatim: should the block protocol be separate from the VFS? how do channels work with these block devices? how should we wire up different filesystems? Grounded in a survey of how Minix 3, QNX Neutrino, Fuchsia, Redox, Plan 9, and Linux each answered the same questions, and in exactly where our own seams sit today. The layering rule this discussion serves: drivers are the lowest level (hardware only); VFS and the filesystems are higher layers; the protocol layer routes between them.

What the survey says, compressed

Everyone separates block from VFS. Even Plan 9, which unifies the protocol (a disk is a file tree), keeps the layers: kernel sd serves raw ranges, user-space filesystem servers consume them and serve files. The two poles are Minix 3 (every layer a process: one VFS server, one filesystem server per mounted volume, driver processes below — driver crashes proven survivable by fault-injection, tens of thousands of injected faults) and QNX (filesystems as DLLs inside the disk-driver process — fastest path, but one corrupt volume takes every volume behind that controller down). Fuchsia is the modern capability- native reference: a tiny block contract that every layer speaks, control channel separate from a data fast-path with pre-registered buffers, filesystems as separate processes launched by a storage-policy component (fshost) that probes content and hands each filesystem its block channel at startup. The filesystem never discovers devices.

Partition tables are parsed in user space, below the filesystem, above the raw device — never in the kernel, never in the filesystem engine. Fuchsia tried partitions-as-drivers and formally reversed it (RFC-0257). Plan 9 has the cleanest split of all: the kernel offers only a mechanism — "create a named sub-range of this disk" — and a user-space prober parses the table and issues those commands. QNX (io-blk) and Minix (driver-side shared library) agree on placement: the block layer multiplexes one physical block object into several logical block objects of the same contract.

The failure lessons are unanimous. Linux's shared page cache across the filesystem boundary produced fsyncgate (write errors observed by the wrong process, dirty pages marked clean); the lesson is to keep write caching inside the filesystem process, so an error surfaces on the channel that owns the volume. Linux's errors=remount-ro/lazy-unmount half-alive states exist because a monolith cannot kill and respawn a filesystem; Plan 9 leaves dead servers' mounts in the namespace erroring forever. A restartable-parts system should have exactly one surprise-removal path: kill the filesystem process, retire its mounts, respawn on return. QNX's deepest lesson: after years of removable-media pain they concluded detection isn't enough and built a copy-on-write filesystem (Power-Safe) — format-level crash honesty. And QNX's mcd (a standalone policy daemon with declarative insertion/removal rules) plus Fuchsia's fshost agree that media policy is one dedicated component, not code scattered through filesystems.

Where we already stand

Closer than expected. The block protocol is 61 lines, five verbs, and its docblock already reserves Header.target for volumes ("targets = volumes, enumerate lists them; nothing else about the protocol changes"). usb-storage is content-blind — it reads block 0 only as a bring-up self-check and parses nothing. The vfs protocol is already served by a second provider in production (init's registry serves it synthetically), so a second filesystem is not a protocol problem. The FAT service is already internally split: a host-testable engine (1659 + 310 lines) behind a four-function BlockDevice vtable, and a 434-line IPC shell. The kernel mount table already has the right restart semantics (remount-replace; lazy dead-endpoint sweep on resolve).

The misplacements, all small: the MBR walk lives inside the FAT engine (engine.zig:196-212) — every future engine would duplicate it; fat hardcodes its acquisition policy ("first mass-storage child, mount at /volumes/usb") — that is policy in a filesystem binary; fs_unmount is gated by nothing (any process may unmount any prefix — latent now, an obvious cross-tenant hole once mounts multiply); and fat's mounted flag is one-way — no path back after its block channel dies.

The proposed shape

Three layers, matching the stated rule, every boundary a channel:

Block (driver layer, mechanism). usb-storage stays hardware→blocks and content-blind. It gains one mechanism, Plan 9's: named sub-ranges. A define_range-style op creates a logical block object (a partition) clamped and offset onto the disk; sub-ranges are target ids on the same endpoint, or separate endpoints handed out per range — either way they speak the same block contract (Fuchsia's closed-under-layering rule), so a filesystem cannot tell whole-disk from partition. The driver never parses a table; it clamps ranges it is told about. Offsets translate at definition time, so the data path stays one hop (Fuchsia's session-mapping trick, for free).

Volumes (service layer, policy). One new service — the volume manager (fshost/mcd-shaped, init-supervised; the device manager stays devices-only). It subscribes to the device manager's child_added/child_removed, consumer- hellos for each new mass-storage provider's block channel, reads the partition table and the first blocks itself (the prober is policy), consults configuration — filesystems.csv: content signature → filesystem binary; mounts.csv or similar: volume identity → mount prefix — defines sub-ranges on the driver, spawns one filesystem process per volume, hands it its block channel at startup, and supervises it with the reap-and-rebuild idiom the device manager proved: provider dies or medium leaves → kill the filesystem process, retire its mounts; medium returns → re-probe, respawn, remount. The boot volume is chosen by content (which volume carries /system/configuration and /system/logs), closing the two-sticks question honestly.

Filesystems (per volume, one process). The proven unit everywhere from Plan 9's dossrv to Minix to Fuchsia: block-client + engine + file-protocol provider in one binary, one process per volume (9front practice; per-volume fault isolation is what our supervision makes cheap). fat's shell becomes a shared filesystem harness library before a second engine is written; the MBR walk moves out of the engine into the volume manager; write caching stays inside the process (the anti-fsyncgate rule). Each mounts its prefixes into the kernel mount table itself, exactly as today.

Kernel: two small changes only. fs_unmount gains ownership (only the mounting endpoint's holder may unmount — possession-is-capability, consistent with everything else), and the 8-slot mount table gets a declared bound or growth once volumes multiply. The mount table stays the router; per-process namespaces (Plan 9's extra) remain separable future work.

The decisions on the table

  1. Partitions as driver-side sub-ranges (mechanism in usb-storage, parsing in the volume manager) — versus a separate partition process re-serving block (Fuchsia's storage-host). Sub-ranges are ~20 lines of clamp in the driver and keep the data path one hop; a separate process is purer layering at the cost of a hop or session plumbing. Recommendation: sub-ranges.
  2. The volume manager as a new service owning probe, spawn, supervision, and mount policy — with filesystems.csv and the mount map as configuration. Recommendation: yes; it is the missing policy home that fat is currently squatting in.
  3. One filesystem process per volume (fat's binary becomes "the FAT implementation", spawned per FAT volume). Recommendation: yes — it extends recompile-and-restart-live to filesystems and isolates corrupt media.
  4. Sub-range addressing: target ids on the storage endpoint versus one endpoint per volume handed out by the driver. Endpoint-per-volume matches the establishment-plane machinery (a channel per party, caps at establishment) and keeps per-client badge scoping simple. Recommendation: endpoint per volume.
  5. fs_unmount ownership — a defect fix more than a decision.
  6. Later, kept open: the shm-ring data plane (communication.md already names it as the 256-byte ceiling's unlock — Fuchsia's FIFO+VMO is the precedent); format-level crash honesty (a Power-Safe-style journaling or COW filesystem) once danos outgrows FAT; per-process namespaces.