docs: the storage-stack discussion — block, volumes, filesystems, against the survey
This commit is contained in:
@@ -0,0 +1,138 @@
|
||||
# The storage stack: block, volumes, filesystems — a discussion
|
||||
|
||||
*2026-08-09. Design discussion, not a plan. The questions, verbatim: should the
|
||||
block protocol be separate from the VFS? how do channels work with these block
|
||||
devices? how should we wire up different filesystems? Grounded in a survey of how
|
||||
Minix 3, QNX Neutrino, Fuchsia, Redox, Plan 9, and Linux each answered the same
|
||||
questions, and in exactly where our own seams sit today. The layering rule this
|
||||
discussion serves: drivers are the lowest level (hardware only); VFS and the
|
||||
filesystems are higher layers; the protocol layer routes between them.*
|
||||
|
||||
## What the survey says, compressed
|
||||
|
||||
**Everyone separates block from VFS.** Even Plan 9, which unifies the *protocol*
|
||||
(a disk is a file tree), keeps the layers: kernel `sd` serves raw ranges,
|
||||
user-space filesystem servers consume them and serve files. The two poles are
|
||||
Minix 3 (every layer a process: one VFS server, one filesystem server *per
|
||||
mounted volume*, driver processes below — driver crashes proven survivable by
|
||||
fault-injection, tens of thousands of injected faults) and QNX (filesystems as
|
||||
DLLs inside the disk-driver process — fastest path, but one corrupt volume takes
|
||||
every volume behind that controller down). Fuchsia is the modern capability-
|
||||
native reference: a tiny block contract that every layer speaks, control channel
|
||||
separate from a data fast-path with pre-registered buffers, filesystems as
|
||||
separate processes launched by a storage-policy component (`fshost`) that probes
|
||||
content and hands each filesystem its block channel at startup. The filesystem
|
||||
never discovers devices.
|
||||
|
||||
**Partition tables are parsed in user space, below the filesystem, above the
|
||||
raw device — never in the kernel, never in the filesystem engine.** Fuchsia
|
||||
tried partitions-as-drivers and formally reversed it (RFC-0257). Plan 9 has the
|
||||
cleanest split of all: the kernel offers only a *mechanism* — "create a named
|
||||
sub-range of this disk" — and a user-space prober parses the table and issues
|
||||
those commands. QNX (`io-blk`) and Minix (driver-side shared library) agree on
|
||||
placement: the block layer multiplexes one physical block object into several
|
||||
logical block objects **of the same contract**.
|
||||
|
||||
**The failure lessons are unanimous.** Linux's shared page cache across the
|
||||
filesystem boundary produced fsyncgate (write errors observed by the wrong
|
||||
process, dirty pages marked clean); the lesson is to keep write caching *inside*
|
||||
the filesystem process, so an error surfaces on the channel that owns the
|
||||
volume. Linux's `errors=remount-ro`/lazy-unmount half-alive states exist because
|
||||
a monolith cannot kill and respawn a filesystem; Plan 9 leaves dead servers'
|
||||
mounts in the namespace erroring forever. A restartable-parts system should have
|
||||
exactly one surprise-removal path: kill the filesystem process, retire its
|
||||
mounts, respawn on return. QNX's deepest lesson: after years of removable-media
|
||||
pain they concluded detection isn't enough and built a copy-on-write filesystem
|
||||
(Power-Safe) — *format-level* crash honesty. And QNX's `mcd` (a standalone
|
||||
policy daemon with declarative insertion/removal rules) plus Fuchsia's `fshost`
|
||||
agree that media policy is one dedicated component, not code scattered through
|
||||
filesystems.
|
||||
|
||||
## Where we already stand
|
||||
|
||||
Closer than expected. The block protocol is 61 lines, five verbs, and its
|
||||
docblock already reserves `Header.target` for volumes ("targets = volumes,
|
||||
`enumerate` lists them; nothing else about the protocol changes"). usb-storage
|
||||
is content-blind — it reads block 0 only as a bring-up self-check and parses
|
||||
nothing. The vfs protocol is already served by a second provider in production
|
||||
(init's registry serves it synthetically), so a second filesystem is *not* a
|
||||
protocol problem. The FAT service is already internally split: a host-testable
|
||||
engine (1659 + 310 lines) behind a four-function BlockDevice vtable, and a
|
||||
434-line IPC shell. The kernel mount table already has the right restart
|
||||
semantics (remount-replace; lazy dead-endpoint sweep on resolve).
|
||||
|
||||
The misplacements, all small: the MBR walk lives *inside the FAT engine*
|
||||
(`engine.zig:196-212`) — every future engine would duplicate it; fat hardcodes
|
||||
its acquisition policy ("first mass-storage child, mount at /volumes/usb") —
|
||||
that is policy in a filesystem binary; `fs_unmount` is gated by nothing (any
|
||||
process may unmount any prefix — latent now, an obvious cross-tenant hole once
|
||||
mounts multiply); and fat's `mounted` flag is one-way — no path back after its
|
||||
block channel dies.
|
||||
|
||||
## The proposed shape
|
||||
|
||||
Three layers, matching the stated rule, every boundary a channel:
|
||||
|
||||
**Block (driver layer, mechanism).** usb-storage stays hardware→blocks and
|
||||
content-blind. It gains one mechanism, Plan 9's: *named sub-ranges*. A
|
||||
`define_range`-style op creates a logical block object (a partition) clamped
|
||||
and offset onto the disk; sub-ranges are `target` ids on the same endpoint, or
|
||||
separate endpoints handed out per range — either way they speak the **same
|
||||
block contract** (Fuchsia's closed-under-layering rule), so a filesystem cannot
|
||||
tell whole-disk from partition. The driver never parses a table; it clamps
|
||||
ranges it is told about. Offsets translate at definition time, so the data path
|
||||
stays one hop (Fuchsia's session-mapping trick, for free).
|
||||
|
||||
**Volumes (service layer, policy).** One new service — the *volume manager*
|
||||
(fshost/mcd-shaped, init-supervised; the device manager stays devices-only). It
|
||||
subscribes to the device manager's `child_added`/`child_removed`, consumer-
|
||||
hellos for each new mass-storage provider's block channel, reads the partition
|
||||
table and the first blocks itself (the prober is policy), consults
|
||||
configuration — `filesystems.csv`: content signature → filesystem binary;
|
||||
`mounts.csv` or similar: volume identity → mount prefix — defines sub-ranges on
|
||||
the driver, spawns **one filesystem process per volume**, hands it its block
|
||||
channel at startup, and supervises it with the reap-and-rebuild idiom the
|
||||
device manager proved: provider dies or medium leaves → kill the filesystem
|
||||
process, retire its mounts; medium returns → re-probe, respawn, remount. The
|
||||
boot volume is chosen by *content* (which volume carries /system/configuration
|
||||
and /system/logs), closing the two-sticks question honestly.
|
||||
|
||||
**Filesystems (per volume, one process).** The proven unit everywhere from
|
||||
Plan 9's `dossrv` to Minix to Fuchsia: block-client + engine + file-protocol
|
||||
provider in one binary, one process per volume (9front practice; per-volume
|
||||
fault isolation is what our supervision makes cheap). fat's shell becomes a
|
||||
shared *filesystem harness* library before a second engine is written; the MBR
|
||||
walk moves out of the engine into the volume manager; write caching stays
|
||||
inside the process (the anti-fsyncgate rule). Each mounts its prefixes into the
|
||||
kernel mount table itself, exactly as today.
|
||||
|
||||
**Kernel: two small changes only.** `fs_unmount` gains ownership (only the
|
||||
mounting endpoint's holder may unmount — possession-is-capability, consistent
|
||||
with everything else), and the 8-slot mount table gets a declared bound or
|
||||
growth once volumes multiply. The mount table stays the router; per-process
|
||||
namespaces (Plan 9's extra) remain separable future work.
|
||||
|
||||
## The decisions on the table
|
||||
|
||||
1. **Partitions as driver-side sub-ranges** (mechanism in usb-storage, parsing
|
||||
in the volume manager) — versus a separate partition *process* re-serving
|
||||
block (Fuchsia's storage-host). Sub-ranges are ~20 lines of clamp in the
|
||||
driver and keep the data path one hop; a separate process is purer layering
|
||||
at the cost of a hop or session plumbing. Recommendation: sub-ranges.
|
||||
2. **The volume manager as a new service** owning probe, spawn, supervision,
|
||||
and mount policy — with `filesystems.csv` and the mount map as
|
||||
configuration. Recommendation: yes; it is the missing policy home that fat
|
||||
is currently squatting in.
|
||||
3. **One filesystem process per volume** (fat's binary becomes "the FAT
|
||||
implementation", spawned per FAT volume). Recommendation: yes — it extends
|
||||
recompile-and-restart-live to filesystems and isolates corrupt media.
|
||||
4. **Sub-range addressing**: `target` ids on the storage endpoint versus one
|
||||
endpoint per volume handed out by the driver. Endpoint-per-volume matches
|
||||
the establishment-plane machinery (a channel per party, caps at
|
||||
establishment) and keeps per-client badge scoping simple. Recommendation:
|
||||
endpoint per volume.
|
||||
5. **`fs_unmount` ownership** — a defect fix more than a decision.
|
||||
6. **Later, kept open**: the shm-ring data plane (communication.md already
|
||||
names it as the 256-byte ceiling's unlock — Fuchsia's FIFO+VMO is the
|
||||
precedent); format-level crash honesty (a Power-Safe-style journaling or COW
|
||||
filesystem) once danos outgrows FAT; per-process namespaces.
|
||||
Reference in New Issue
Block a user