From 5a8a2a4d7e796f283a106cab865dbeb757ca9bfb Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sun, 9 Aug 2026 14:47:35 +0100 Subject: [PATCH] =?UTF-8?q?docs:=20the=20storage-stack=20discussion=20?= =?UTF-8?q?=E2=80=94=20block,=20volumes,=20filesystems,=20against=20the=20?= =?UTF-8?q?survey?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- docs/storage-stack-discussion.md | 138 +++++++++++++++++++++++++++++++ 1 file changed, 138 insertions(+) create mode 100644 docs/storage-stack-discussion.md diff --git a/docs/storage-stack-discussion.md b/docs/storage-stack-discussion.md new file mode 100644 index 0000000..4a3ad64 --- /dev/null +++ b/docs/storage-stack-discussion.md @@ -0,0 +1,138 @@ +# The storage stack: block, volumes, filesystems — a discussion + +*2026-08-09. Design discussion, not a plan. The questions, verbatim: should the +block protocol be separate from the VFS? how do channels work with these block +devices? how should we wire up different filesystems? Grounded in a survey of how +Minix 3, QNX Neutrino, Fuchsia, Redox, Plan 9, and Linux each answered the same +questions, and in exactly where our own seams sit today. The layering rule this +discussion serves: drivers are the lowest level (hardware only); VFS and the +filesystems are higher layers; the protocol layer routes between them.* + +## What the survey says, compressed + +**Everyone separates block from VFS.** Even Plan 9, which unifies the *protocol* +(a disk is a file tree), keeps the layers: kernel `sd` serves raw ranges, +user-space filesystem servers consume them and serve files. The two poles are +Minix 3 (every layer a process: one VFS server, one filesystem server *per +mounted volume*, driver processes below — driver crashes proven survivable by +fault-injection, tens of thousands of injected faults) and QNX (filesystems as +DLLs inside the disk-driver process — fastest path, but one corrupt volume takes +every volume behind that controller down). Fuchsia is the modern capability- +native reference: a tiny block contract that every layer speaks, control channel +separate from a data fast-path with pre-registered buffers, filesystems as +separate processes launched by a storage-policy component (`fshost`) that probes +content and hands each filesystem its block channel at startup. The filesystem +never discovers devices. + +**Partition tables are parsed in user space, below the filesystem, above the +raw device — never in the kernel, never in the filesystem engine.** Fuchsia +tried partitions-as-drivers and formally reversed it (RFC-0257). Plan 9 has the +cleanest split of all: the kernel offers only a *mechanism* — "create a named +sub-range of this disk" — and a user-space prober parses the table and issues +those commands. QNX (`io-blk`) and Minix (driver-side shared library) agree on +placement: the block layer multiplexes one physical block object into several +logical block objects **of the same contract**. + +**The failure lessons are unanimous.** Linux's shared page cache across the +filesystem boundary produced fsyncgate (write errors observed by the wrong +process, dirty pages marked clean); the lesson is to keep write caching *inside* +the filesystem process, so an error surfaces on the channel that owns the +volume. Linux's `errors=remount-ro`/lazy-unmount half-alive states exist because +a monolith cannot kill and respawn a filesystem; Plan 9 leaves dead servers' +mounts in the namespace erroring forever. A restartable-parts system should have +exactly one surprise-removal path: kill the filesystem process, retire its +mounts, respawn on return. QNX's deepest lesson: after years of removable-media +pain they concluded detection isn't enough and built a copy-on-write filesystem +(Power-Safe) — *format-level* crash honesty. And QNX's `mcd` (a standalone +policy daemon with declarative insertion/removal rules) plus Fuchsia's `fshost` +agree that media policy is one dedicated component, not code scattered through +filesystems. + +## Where we already stand + +Closer than expected. The block protocol is 61 lines, five verbs, and its +docblock already reserves `Header.target` for volumes ("targets = volumes, +`enumerate` lists them; nothing else about the protocol changes"). usb-storage +is content-blind — it reads block 0 only as a bring-up self-check and parses +nothing. The vfs protocol is already served by a second provider in production +(init's registry serves it synthetically), so a second filesystem is *not* a +protocol problem. The FAT service is already internally split: a host-testable +engine (1659 + 310 lines) behind a four-function BlockDevice vtable, and a +434-line IPC shell. The kernel mount table already has the right restart +semantics (remount-replace; lazy dead-endpoint sweep on resolve). + +The misplacements, all small: the MBR walk lives *inside the FAT engine* +(`engine.zig:196-212`) — every future engine would duplicate it; fat hardcodes +its acquisition policy ("first mass-storage child, mount at /volumes/usb") — +that is policy in a filesystem binary; `fs_unmount` is gated by nothing (any +process may unmount any prefix — latent now, an obvious cross-tenant hole once +mounts multiply); and fat's `mounted` flag is one-way — no path back after its +block channel dies. + +## The proposed shape + +Three layers, matching the stated rule, every boundary a channel: + +**Block (driver layer, mechanism).** usb-storage stays hardware→blocks and +content-blind. It gains one mechanism, Plan 9's: *named sub-ranges*. A +`define_range`-style op creates a logical block object (a partition) clamped +and offset onto the disk; sub-ranges are `target` ids on the same endpoint, or +separate endpoints handed out per range — either way they speak the **same +block contract** (Fuchsia's closed-under-layering rule), so a filesystem cannot +tell whole-disk from partition. The driver never parses a table; it clamps +ranges it is told about. Offsets translate at definition time, so the data path +stays one hop (Fuchsia's session-mapping trick, for free). + +**Volumes (service layer, policy).** One new service — the *volume manager* +(fshost/mcd-shaped, init-supervised; the device manager stays devices-only). It +subscribes to the device manager's `child_added`/`child_removed`, consumer- +hellos for each new mass-storage provider's block channel, reads the partition +table and the first blocks itself (the prober is policy), consults +configuration — `filesystems.csv`: content signature → filesystem binary; +`mounts.csv` or similar: volume identity → mount prefix — defines sub-ranges on +the driver, spawns **one filesystem process per volume**, hands it its block +channel at startup, and supervises it with the reap-and-rebuild idiom the +device manager proved: provider dies or medium leaves → kill the filesystem +process, retire its mounts; medium returns → re-probe, respawn, remount. The +boot volume is chosen by *content* (which volume carries /system/configuration +and /system/logs), closing the two-sticks question honestly. + +**Filesystems (per volume, one process).** The proven unit everywhere from +Plan 9's `dossrv` to Minix to Fuchsia: block-client + engine + file-protocol +provider in one binary, one process per volume (9front practice; per-volume +fault isolation is what our supervision makes cheap). fat's shell becomes a +shared *filesystem harness* library before a second engine is written; the MBR +walk moves out of the engine into the volume manager; write caching stays +inside the process (the anti-fsyncgate rule). Each mounts its prefixes into the +kernel mount table itself, exactly as today. + +**Kernel: two small changes only.** `fs_unmount` gains ownership (only the +mounting endpoint's holder may unmount — possession-is-capability, consistent +with everything else), and the 8-slot mount table gets a declared bound or +growth once volumes multiply. The mount table stays the router; per-process +namespaces (Plan 9's extra) remain separable future work. + +## The decisions on the table + +1. **Partitions as driver-side sub-ranges** (mechanism in usb-storage, parsing + in the volume manager) — versus a separate partition *process* re-serving + block (Fuchsia's storage-host). Sub-ranges are ~20 lines of clamp in the + driver and keep the data path one hop; a separate process is purer layering + at the cost of a hop or session plumbing. Recommendation: sub-ranges. +2. **The volume manager as a new service** owning probe, spawn, supervision, + and mount policy — with `filesystems.csv` and the mount map as + configuration. Recommendation: yes; it is the missing policy home that fat + is currently squatting in. +3. **One filesystem process per volume** (fat's binary becomes "the FAT + implementation", spawned per FAT volume). Recommendation: yes — it extends + recompile-and-restart-live to filesystems and isolates corrupt media. +4. **Sub-range addressing**: `target` ids on the storage endpoint versus one + endpoint per volume handed out by the driver. Endpoint-per-volume matches + the establishment-plane machinery (a channel per party, caps at + establishment) and keeps per-client badge scoping simple. Recommendation: + endpoint per volume. +5. **`fs_unmount` ownership** — a defect fix more than a decision. +6. **Later, kept open**: the shm-ring data plane (communication.md already + names it as the 256-byte ceiling's unlock — Fuchsia's FIFO+VMO is the + precedent); format-level crash honesty (a Power-Safe-style journaling or COW + filesystem) once danos outgrows FAT; per-process namespaces.