docs: the storage rationale lives with the architecture it justifies
storage-stack-discussion.md was misplaced at the docs root — that level is for track plans; this is the file-system domain's design record. Moved to file-system-development/storage-design-rationale.md, renamed to say what it is, cross-references updated.
This commit is contained in:
@@ -1,7 +1,7 @@
|
||||
# The storage architecture: layers, boundaries, responsibilities
|
||||
|
||||
> **Status:** the layered model below is the settled design
|
||||
> ([storage-stack-discussion.md](../storage-stack-discussion.md) records how it
|
||||
> ([storage-design-rationale.md](storage-design-rationale.md) records how it
|
||||
> was reached and what the surveyed systems taught). The data path — vfs
|
||||
> protocol, kernel mount routing, the FAT service, the block protocol,
|
||||
> usb-storage — is **built**. The volume manager, the driver's range
|
||||
|
||||
@@ -0,0 +1,196 @@
|
||||
# The storage design rationale: why the stack is shaped this way
|
||||
|
||||
*2026-08-09. The design record behind
|
||||
[storage-architecture.md](storage-architecture.md): the survey, the
|
||||
trade-offs, and the decisions with their reasons — kept so future changes
|
||||
argue against the evidence rather than rediscovering it. The questions that
|
||||
drove it, verbatim: should the block protocol be separate from the VFS? how
|
||||
do channels work with these block devices? how should we wire up different
|
||||
filesystems? Grounded in how Minix 3, QNX Neutrino, Fuchsia, Redox, Plan 9,
|
||||
and Linux each answered the same questions, and in exactly where our own
|
||||
seams sat. The layering rule it serves: drivers are the lowest level
|
||||
(hardware only); VFS and the filesystems are higher layers; the protocol
|
||||
layer routes between them.*
|
||||
|
||||
## What the survey says, compressed
|
||||
|
||||
**Everyone separates block from VFS.** Even Plan 9, which unifies the *protocol*
|
||||
(a disk is a file tree), keeps the layers: kernel `sd` serves raw ranges,
|
||||
user-space filesystem servers consume them and serve files. The two poles are
|
||||
Minix 3 (every layer a process: one VFS server, one filesystem server *per
|
||||
mounted volume*, driver processes below — driver crashes proven survivable by
|
||||
fault-injection, tens of thousands of injected faults) and QNX (filesystems as
|
||||
DLLs inside the disk-driver process — fastest path, but one corrupt volume takes
|
||||
every volume behind that controller down). Fuchsia is the modern capability-
|
||||
native reference: a tiny block contract that every layer speaks, control channel
|
||||
separate from a data fast-path with pre-registered buffers, filesystems as
|
||||
separate processes launched by a storage-policy component (`fshost`) that probes
|
||||
content and hands each filesystem its block channel at startup. The filesystem
|
||||
never discovers devices.
|
||||
|
||||
**Partition tables are parsed in user space, below the filesystem, above the
|
||||
raw device — never in the kernel, never in the filesystem engine.** Fuchsia
|
||||
tried partitions-as-drivers and formally reversed it (RFC-0257). Plan 9 has the
|
||||
cleanest split of all: the kernel offers only a *mechanism* — "create a named
|
||||
sub-range of this disk" — and a user-space prober parses the table and issues
|
||||
those commands. QNX (`io-blk`) and Minix (driver-side shared library) agree on
|
||||
placement: the block layer multiplexes one physical block object into several
|
||||
logical block objects **of the same contract**.
|
||||
|
||||
**The failure lessons are unanimous.** Linux's shared page cache across the
|
||||
filesystem boundary produced fsyncgate (write errors observed by the wrong
|
||||
process, dirty pages marked clean); the lesson is to keep write caching *inside*
|
||||
the filesystem process, so an error surfaces on the channel that owns the
|
||||
volume. Linux's `errors=remount-ro`/lazy-unmount half-alive states exist because
|
||||
a monolith cannot kill and respawn a filesystem; Plan 9 leaves dead servers'
|
||||
mounts in the namespace erroring forever. A restartable-parts system should have
|
||||
exactly one surprise-removal path: kill the filesystem process, retire its
|
||||
mounts, respawn on return. QNX's deepest lesson: after years of removable-media
|
||||
pain they concluded detection isn't enough and built a copy-on-write filesystem
|
||||
(Power-Safe) — *format-level* crash honesty. And QNX's `mcd` (a standalone
|
||||
policy daemon with declarative insertion/removal rules) plus Fuchsia's `fshost`
|
||||
agree that media policy is one dedicated component, not code scattered through
|
||||
filesystems.
|
||||
|
||||
## Where we already stand
|
||||
|
||||
Closer than expected. The block protocol is 61 lines, five verbs, and its
|
||||
docblock already reserves `Header.target` for volumes ("targets = volumes,
|
||||
`enumerate` lists them; nothing else about the protocol changes"). usb-storage
|
||||
is content-blind — it reads block 0 only as a bring-up self-check and parses
|
||||
nothing. The vfs protocol is already served by a second provider in production
|
||||
(init's registry serves it synthetically), so a second filesystem is *not* a
|
||||
protocol problem. The FAT service is already internally split: a host-testable
|
||||
engine (1659 + 310 lines) behind a four-function BlockDevice vtable, and a
|
||||
434-line IPC shell. The kernel mount table already has the right restart
|
||||
semantics (remount-replace; lazy dead-endpoint sweep on resolve).
|
||||
|
||||
The misplacements, all small: the MBR walk lives *inside the FAT engine*
|
||||
(`engine.zig:196-212`) — every future engine would duplicate it; fat hardcodes
|
||||
its acquisition policy ("first mass-storage child, mount at /volumes/usb") —
|
||||
that is policy in a filesystem binary; `fs_unmount` is gated by nothing (any
|
||||
process may unmount any prefix — latent now, an obvious cross-tenant hole once
|
||||
mounts multiply); and fat's `mounted` flag is one-way — no path back after its
|
||||
block channel dies.
|
||||
|
||||
## The proposed shape
|
||||
|
||||
Three layers, matching the stated rule, every boundary a channel:
|
||||
|
||||
**Block (driver layer, mechanism).** usb-storage stays hardware→blocks and
|
||||
content-blind. It gains one mechanism, Plan 9's: *named sub-ranges*. A
|
||||
`define_range`-style op creates a logical block object (a partition) clamped
|
||||
and offset onto the disk; sub-ranges are `target` ids on the same endpoint, or
|
||||
separate endpoints handed out per range — either way they speak the **same
|
||||
block contract** (Fuchsia's closed-under-layering rule), so a filesystem cannot
|
||||
tell whole-disk from partition. The driver never parses a table; it clamps
|
||||
ranges it is told about. Offsets translate at definition time, so the data path
|
||||
stays one hop (Fuchsia's session-mapping trick, for free).
|
||||
|
||||
**Volumes (service layer, policy).** One new service — the *volume manager*
|
||||
(fshost/mcd-shaped, init-supervised; the device manager stays devices-only). It
|
||||
subscribes to the device manager's `child_added`/`child_removed`, consumer-
|
||||
hellos for each new mass-storage provider's block channel, reads the partition
|
||||
table and the first blocks itself (the prober is policy), consults
|
||||
configuration — `filesystems.csv`: content signature → filesystem binary;
|
||||
`mounts.csv` or similar: volume identity → mount prefix — defines sub-ranges on
|
||||
the driver, spawns **one filesystem process per volume**, hands it its block
|
||||
channel at startup, and supervises it with the reap-and-rebuild idiom the
|
||||
device manager proved: provider dies or medium leaves → kill the filesystem
|
||||
process, retire its mounts; medium returns → re-probe, respawn, remount. The
|
||||
boot volume is chosen by *content* (which volume carries /system/configuration
|
||||
and /system/logs), closing the two-sticks question honestly.
|
||||
|
||||
**Filesystems (per volume, one process).** The proven unit everywhere from
|
||||
Plan 9's `dossrv` to Minix to Fuchsia: block-client + engine + file-protocol
|
||||
provider in one binary, one process per volume (9front practice; per-volume
|
||||
fault isolation is what our supervision makes cheap). fat's shell becomes a
|
||||
shared *filesystem harness* library before a second engine is written; the MBR
|
||||
walk moves out of the engine into the volume manager; write caching stays
|
||||
inside the process (the anti-fsyncgate rule). Each mounts its prefixes into the
|
||||
kernel mount table itself, exactly as today.
|
||||
|
||||
**Kernel: two small changes only.** `fs_unmount` gains ownership (only the
|
||||
mounting endpoint's holder may unmount — possession-is-capability, consistent
|
||||
with everything else), and the 8-slot mount table gets a declared bound or
|
||||
growth once volumes multiply. The mount table stays the router; per-process
|
||||
namespaces (Plan 9's extra) remain separable future work.
|
||||
|
||||
## Transport generality: NVMe and SATA against this design
|
||||
|
||||
The layering was chosen transport-agnostic on purpose (the 9front lesson:
|
||||
filesystem services cannot tell a USB stick from an ATA disk); NVMe and AHCI
|
||||
are the test of that claim, and they fit — with three named pressure points.
|
||||
|
||||
**What transfers untouched:** the block protocol (nothing USB in it), the
|
||||
volume manager, partitions/ranges, filesystems, mounts, and the whole
|
||||
enforcement stack — delegation, IOMMU confinement (these are first-party PCI
|
||||
DMA masters, so confinement applies even more directly than USB),
|
||||
reap-and-rebuild, the hot-plug lifecycle. Integration cost per transport is
|
||||
what the architecture promises: one driver binary plus one `devices.csv` row.
|
||||
|
||||
**NVMe** shortens the stack — no bus/class split; the driver IS the
|
||||
controller driver, one hop fewer than USB. Its structural novelty,
|
||||
**namespaces** (hardware-native multiple volumes behind one controller), is
|
||||
exactly the case the block protocol reserved on day one and decision 4 below
|
||||
settles as endpoint-per-volume. **AHCI** is a shape choice, not a problem: an
|
||||
HBA fronts up to 32 ports plus port multipliers — structurally a bus — so
|
||||
either mirror USB (ahci-bus + a per-port disk driver: maximum restart
|
||||
granularity, the proven shape) or mirror NVMe (one driver per controller, one
|
||||
block channel per port). Leaning per-port processes for consistency with the
|
||||
matrix-proven shape; genuinely open.
|
||||
|
||||
**The pressure points, honestly:**
|
||||
|
||||
1. **Multi-volume providers are reserved, not implemented.** The volume
|
||||
manager flow assumes one provider, one volume; NVMe namespaces make
|
||||
endpoint-per-volume real work with hardware demanding it.
|
||||
2. **The current transport will bottleneck NVMe.** Synchronous call/reply,
|
||||
one operation in flight, one bounce buffer — fine for a USB2 stick,
|
||||
forfeits an NVMe drive's queue depth and per-queue MSI-X. Correctness
|
||||
needs nothing; performance is gated on the shm-ring data plane
|
||||
communication.md already names (Fuchsia's FIFO+VMO is the precedent).
|
||||
NVMe is not worth building before that milestone.
|
||||
3. **Media lifecycle ≠ device lifecycle.** ATAPI trays and SD card readers
|
||||
(USB ones too, today) keep the DEVICE present while the MEDIUM comes and
|
||||
goes — a removal trigger our channel-death lifecycle does not carry. The
|
||||
fix is decision 7 below. (Related small item: fat's bounce sizing
|
||||
hardcodes 512-byte sectors; 4Kn drives need the geometry honored
|
||||
throughout.)
|
||||
|
||||
## The decisions on the table
|
||||
|
||||
1. **Partitions as driver-side sub-ranges** (mechanism in usb-storage, parsing
|
||||
in the volume manager) — versus a separate partition *process* re-serving
|
||||
block (Fuchsia's storage-host). Sub-ranges are ~20 lines of clamp in the
|
||||
driver and keep the data path one hop; a separate process is purer layering
|
||||
at the cost of a hop or session plumbing. Recommendation: sub-ranges.
|
||||
2. **The volume manager as a new service** owning probe, spawn, supervision,
|
||||
and mount policy — with `filesystems.csv` and the mount map as
|
||||
configuration. Recommendation: yes; it is the missing policy home that fat
|
||||
is currently squatting in.
|
||||
3. **One filesystem process per volume** (fat's binary becomes "the FAT
|
||||
implementation", spawned per FAT volume). Recommendation: yes — it extends
|
||||
recompile-and-restart-live to filesystems and isolates corrupt media.
|
||||
4. **Sub-range addressing**: `target` ids on the storage endpoint versus one
|
||||
endpoint per volume handed out by the driver. Endpoint-per-volume matches
|
||||
the establishment-plane machinery (a channel per party, caps at
|
||||
establishment) and keeps per-client badge scoping simple. Recommendation:
|
||||
endpoint per volume.
|
||||
5. **`fs_unmount` ownership** — a defect fix more than a decision.
|
||||
6. **Later, kept open**: the shm-ring data plane (communication.md already
|
||||
names it as the 256-byte ceiling's unlock — Fuchsia's FIFO+VMO is the
|
||||
precedent); format-level crash honesty (a Power-Safe-style journaling or COW
|
||||
filesystem) once danos outgrows FAT; per-process namespaces.
|
||||
7. **The media-presence event** (settled in principle; lands with the volume
|
||||
manager): the block protocol gains a pushed event — `medium_changed`, with
|
||||
present/absent and a change counter — produced by the storage driver from
|
||||
its transport's native signal (SCSI UNIT ATTENTION / TEST UNIT READY for
|
||||
USB and ATAPI, PxSSTS for AHCI, namespace-change AER for NVMe) and
|
||||
consumed by the volume manager, which runs the SAME kill-retire-remount
|
||||
path it runs on channel death — one lifecycle, two triggers. The driver
|
||||
reports presence, never content; a pushed event carries no capability,
|
||||
which the kernel already guarantees. The device staying while its medium
|
||||
leaves is the one removable-media case the channel-death trigger cannot
|
||||
see; without this event a swapped SD card would be served with the old
|
||||
card's filesystem state.
|
||||
Reference in New Issue
Block a user