139 lines
8.6 KiB
Markdown
139 lines
8.6 KiB
Markdown
# The storage stack: block, volumes, filesystems — a discussion
|
|
|
|
*2026-08-09. Design discussion, not a plan. The questions, verbatim: should the
|
|
block protocol be separate from the VFS? how do channels work with these block
|
|
devices? how should we wire up different filesystems? Grounded in a survey of how
|
|
Minix 3, QNX Neutrino, Fuchsia, Redox, Plan 9, and Linux each answered the same
|
|
questions, and in exactly where our own seams sit today. The layering rule this
|
|
discussion serves: drivers are the lowest level (hardware only); VFS and the
|
|
filesystems are higher layers; the protocol layer routes between them.*
|
|
|
|
## What the survey says, compressed
|
|
|
|
**Everyone separates block from VFS.** Even Plan 9, which unifies the *protocol*
|
|
(a disk is a file tree), keeps the layers: kernel `sd` serves raw ranges,
|
|
user-space filesystem servers consume them and serve files. The two poles are
|
|
Minix 3 (every layer a process: one VFS server, one filesystem server *per
|
|
mounted volume*, driver processes below — driver crashes proven survivable by
|
|
fault-injection, tens of thousands of injected faults) and QNX (filesystems as
|
|
DLLs inside the disk-driver process — fastest path, but one corrupt volume takes
|
|
every volume behind that controller down). Fuchsia is the modern capability-
|
|
native reference: a tiny block contract that every layer speaks, control channel
|
|
separate from a data fast-path with pre-registered buffers, filesystems as
|
|
separate processes launched by a storage-policy component (`fshost`) that probes
|
|
content and hands each filesystem its block channel at startup. The filesystem
|
|
never discovers devices.
|
|
|
|
**Partition tables are parsed in user space, below the filesystem, above the
|
|
raw device — never in the kernel, never in the filesystem engine.** Fuchsia
|
|
tried partitions-as-drivers and formally reversed it (RFC-0257). Plan 9 has the
|
|
cleanest split of all: the kernel offers only a *mechanism* — "create a named
|
|
sub-range of this disk" — and a user-space prober parses the table and issues
|
|
those commands. QNX (`io-blk`) and Minix (driver-side shared library) agree on
|
|
placement: the block layer multiplexes one physical block object into several
|
|
logical block objects **of the same contract**.
|
|
|
|
**The failure lessons are unanimous.** Linux's shared page cache across the
|
|
filesystem boundary produced fsyncgate (write errors observed by the wrong
|
|
process, dirty pages marked clean); the lesson is to keep write caching *inside*
|
|
the filesystem process, so an error surfaces on the channel that owns the
|
|
volume. Linux's `errors=remount-ro`/lazy-unmount half-alive states exist because
|
|
a monolith cannot kill and respawn a filesystem; Plan 9 leaves dead servers'
|
|
mounts in the namespace erroring forever. A restartable-parts system should have
|
|
exactly one surprise-removal path: kill the filesystem process, retire its
|
|
mounts, respawn on return. QNX's deepest lesson: after years of removable-media
|
|
pain they concluded detection isn't enough and built a copy-on-write filesystem
|
|
(Power-Safe) — *format-level* crash honesty. And QNX's `mcd` (a standalone
|
|
policy daemon with declarative insertion/removal rules) plus Fuchsia's `fshost`
|
|
agree that media policy is one dedicated component, not code scattered through
|
|
filesystems.
|
|
|
|
## Where we already stand
|
|
|
|
Closer than expected. The block protocol is 61 lines, five verbs, and its
|
|
docblock already reserves `Header.target` for volumes ("targets = volumes,
|
|
`enumerate` lists them; nothing else about the protocol changes"). usb-storage
|
|
is content-blind — it reads block 0 only as a bring-up self-check and parses
|
|
nothing. The vfs protocol is already served by a second provider in production
|
|
(init's registry serves it synthetically), so a second filesystem is *not* a
|
|
protocol problem. The FAT service is already internally split: a host-testable
|
|
engine (1659 + 310 lines) behind a four-function BlockDevice vtable, and a
|
|
434-line IPC shell. The kernel mount table already has the right restart
|
|
semantics (remount-replace; lazy dead-endpoint sweep on resolve).
|
|
|
|
The misplacements, all small: the MBR walk lives *inside the FAT engine*
|
|
(`engine.zig:196-212`) — every future engine would duplicate it; fat hardcodes
|
|
its acquisition policy ("first mass-storage child, mount at /volumes/usb") —
|
|
that is policy in a filesystem binary; `fs_unmount` is gated by nothing (any
|
|
process may unmount any prefix — latent now, an obvious cross-tenant hole once
|
|
mounts multiply); and fat's `mounted` flag is one-way — no path back after its
|
|
block channel dies.
|
|
|
|
## The proposed shape
|
|
|
|
Three layers, matching the stated rule, every boundary a channel:
|
|
|
|
**Block (driver layer, mechanism).** usb-storage stays hardware→blocks and
|
|
content-blind. It gains one mechanism, Plan 9's: *named sub-ranges*. A
|
|
`define_range`-style op creates a logical block object (a partition) clamped
|
|
and offset onto the disk; sub-ranges are `target` ids on the same endpoint, or
|
|
separate endpoints handed out per range — either way they speak the **same
|
|
block contract** (Fuchsia's closed-under-layering rule), so a filesystem cannot
|
|
tell whole-disk from partition. The driver never parses a table; it clamps
|
|
ranges it is told about. Offsets translate at definition time, so the data path
|
|
stays one hop (Fuchsia's session-mapping trick, for free).
|
|
|
|
**Volumes (service layer, policy).** One new service — the *volume manager*
|
|
(fshost/mcd-shaped, init-supervised; the device manager stays devices-only). It
|
|
subscribes to the device manager's `child_added`/`child_removed`, consumer-
|
|
hellos for each new mass-storage provider's block channel, reads the partition
|
|
table and the first blocks itself (the prober is policy), consults
|
|
configuration — `filesystems.csv`: content signature → filesystem binary;
|
|
`mounts.csv` or similar: volume identity → mount prefix — defines sub-ranges on
|
|
the driver, spawns **one filesystem process per volume**, hands it its block
|
|
channel at startup, and supervises it with the reap-and-rebuild idiom the
|
|
device manager proved: provider dies or medium leaves → kill the filesystem
|
|
process, retire its mounts; medium returns → re-probe, respawn, remount. The
|
|
boot volume is chosen by *content* (which volume carries /system/configuration
|
|
and /system/logs), closing the two-sticks question honestly.
|
|
|
|
**Filesystems (per volume, one process).** The proven unit everywhere from
|
|
Plan 9's `dossrv` to Minix to Fuchsia: block-client + engine + file-protocol
|
|
provider in one binary, one process per volume (9front practice; per-volume
|
|
fault isolation is what our supervision makes cheap). fat's shell becomes a
|
|
shared *filesystem harness* library before a second engine is written; the MBR
|
|
walk moves out of the engine into the volume manager; write caching stays
|
|
inside the process (the anti-fsyncgate rule). Each mounts its prefixes into the
|
|
kernel mount table itself, exactly as today.
|
|
|
|
**Kernel: two small changes only.** `fs_unmount` gains ownership (only the
|
|
mounting endpoint's holder may unmount — possession-is-capability, consistent
|
|
with everything else), and the 8-slot mount table gets a declared bound or
|
|
growth once volumes multiply. The mount table stays the router; per-process
|
|
namespaces (Plan 9's extra) remain separable future work.
|
|
|
|
## The decisions on the table
|
|
|
|
1. **Partitions as driver-side sub-ranges** (mechanism in usb-storage, parsing
|
|
in the volume manager) — versus a separate partition *process* re-serving
|
|
block (Fuchsia's storage-host). Sub-ranges are ~20 lines of clamp in the
|
|
driver and keep the data path one hop; a separate process is purer layering
|
|
at the cost of a hop or session plumbing. Recommendation: sub-ranges.
|
|
2. **The volume manager as a new service** owning probe, spawn, supervision,
|
|
and mount policy — with `filesystems.csv` and the mount map as
|
|
configuration. Recommendation: yes; it is the missing policy home that fat
|
|
is currently squatting in.
|
|
3. **One filesystem process per volume** (fat's binary becomes "the FAT
|
|
implementation", spawned per FAT volume). Recommendation: yes — it extends
|
|
recompile-and-restart-live to filesystems and isolates corrupt media.
|
|
4. **Sub-range addressing**: `target` ids on the storage endpoint versus one
|
|
endpoint per volume handed out by the driver. Endpoint-per-volume matches
|
|
the establishment-plane machinery (a channel per party, caps at
|
|
establishment) and keeps per-client badge scoping simple. Recommendation:
|
|
endpoint per volume.
|
|
5. **`fs_unmount` ownership** — a defect fix more than a decision.
|
|
6. **Later, kept open**: the shm-ring data plane (communication.md already
|
|
names it as the 256-byte ceiling's unlock — Fuchsia's FIFO+VMO is the
|
|
precedent); format-level crash honesty (a Power-Safe-style journaling or COW
|
|
filesystem) once danos outgrows FAT; per-process namespaces.
|