Files
danos/docs/file-system-development/storage-design-rationale.md
T
Daniel Samson 6d4992ae02 docs: storage — the mount map is built; the path is the id (S2)
Flip the storage docs to match S2: filesystems.csv (signature -> binary) and
volumes.csv (identity -> optional override) are built; a volume's mount path IS
its content id (/volumes/<id>), never a port name; the label is display metadata
a `volumes` query returns. fat receives its mount path via argv[2] rather than
hardcoding it. Still pending: rung 2 (filesystem UUID, needs a non-FAT engine),
multi-volume (fat's boot rewrites stay unconditional until S3), medium_changed
consumption, remount bench-verification.
2026-08-10 00:49:32 +01:00

249 lines
16 KiB
Markdown

# The storage design rationale: why the stack is shaped this way
*2026-08-09. The design record behind
[storage-architecture.md](storage-architecture.md): the survey, the
trade-offs, and the decisions with their reasons — kept so future changes
argue against the evidence rather than rediscovering it. The questions that
drove it, verbatim: should the block protocol be separate from the VFS? how
do channels work with these block devices? how should we wire up different
filesystems? Grounded in how Minix 3, QNX Neutrino, Fuchsia, Redox, Plan 9,
and Linux each answered the same questions, and in exactly where our own
seams sat. The layering rule it serves: drivers are the lowest level
(hardware only); VFS and the filesystems are higher layers; the protocol
layer routes between them.*
## What the survey says, compressed
**Everyone separates block from VFS.** Even Plan 9, which unifies the *protocol*
(a disk is a file tree), keeps the layers: kernel `sd` serves raw ranges,
user-space filesystem servers consume them and serve files. The two poles are
Minix 3 (every layer a process: one VFS server, one filesystem server *per
mounted volume*, driver processes below — driver crashes proven survivable by
fault-injection, tens of thousands of injected faults) and QNX (filesystems as
DLLs inside the disk-driver process — fastest path, but one corrupt volume takes
every volume behind that controller down). Fuchsia is the modern capability-
native reference: a tiny block contract that every layer speaks, control channel
separate from a data fast-path with pre-registered buffers, filesystems as
separate processes launched by a storage-policy component (`fshost`) that probes
content and hands each filesystem its block channel at startup. The filesystem
never discovers devices.
**Partition tables are parsed in user space, below the filesystem, above the
raw device — never in the kernel, never in the filesystem engine.** Fuchsia
tried partitions-as-drivers and formally reversed it (RFC-0257). Plan 9 has the
cleanest split of all: the kernel offers only a *mechanism* — "create a named
sub-range of this disk" — and a user-space prober parses the table and issues
those commands. QNX (`io-blk`) and Minix (driver-side shared library) agree on
placement: the block layer multiplexes one physical block object into several
logical block objects **of the same contract**.
**The failure lessons are unanimous.** Linux's shared page cache across the
filesystem boundary produced fsyncgate (write errors observed by the wrong
process, dirty pages marked clean); the lesson is to keep write caching *inside*
the filesystem process, so an error surfaces on the channel that owns the
volume. Linux's `errors=remount-ro`/lazy-unmount half-alive states exist because
a monolith cannot kill and respawn a filesystem; Plan 9 leaves dead servers'
mounts in the namespace erroring forever. A restartable-parts system should have
exactly one surprise-removal path: kill the filesystem process, retire its
mounts, respawn on return. QNX's deepest lesson: after years of removable-media
pain they concluded detection isn't enough and built a copy-on-write filesystem
(Power-Safe) — *format-level* crash honesty. And QNX's `mcd` (a standalone
policy daemon with declarative insertion/removal rules) plus Fuchsia's `fshost`
agree that media policy is one dedicated component, not code scattered through
filesystems.
## Where we already stand
Closer than expected. The block protocol is 61 lines, five verbs, and its
docblock already reserves `Header.target` for volumes ("targets = volumes,
`enumerate` lists them; nothing else about the protocol changes"). usb-storage
is content-blind — it reads block 0 only as a bring-up self-check and parses
nothing. The vfs protocol is already served by a second provider in production
(init's registry serves it synthetically), so a second filesystem is *not* a
protocol problem. The FAT service is already internally split: a host-testable
engine (1659 + 310 lines) behind a four-function BlockDevice vtable, and a
434-line IPC shell. The kernel mount table already has the right restart
semantics (remount-replace; lazy dead-endpoint sweep on resolve).
The misplacements, all small: the MBR walk lives *inside the FAT engine*
(`engine.zig:196-212`) — every future engine would duplicate it; fat hardcodes
its acquisition policy ("first mass-storage child, mount at /volumes/usb") —
that is policy in a filesystem binary; `fs_unmount` is gated by nothing (any
process may unmount any prefix — latent now, an obvious cross-tenant hole once
mounts multiply); and fat's `mounted` flag is one-way — no path back after its
block channel dies.
## The proposed shape
Three layers, matching the stated rule, every boundary a channel:
**Block (driver layer, mechanism).** usb-storage stays hardware→blocks and
content-blind. It gains one mechanism, Plan 9's: *named sub-ranges*. A
`define_range`-style op creates a logical block object (a partition) clamped
and offset onto the disk; sub-ranges are `target` ids on the same endpoint, or
separate endpoints handed out per range — either way they speak the **same
block contract** (Fuchsia's closed-under-layering rule), so a filesystem cannot
tell whole-disk from partition. The driver never parses a table; it clamps
ranges it is told about. Offsets translate at definition time, so the data path
stays one hop (Fuchsia's session-mapping trick, for free).
**Volumes (service layer, policy).** One new service — the *volume manager*
(fshost/mcd-shaped, init-supervised; the device manager stays devices-only). It
subscribes to the device manager's `child_added`/`child_removed`, consumer-
hellos for each new mass-storage provider's block channel, reads the partition
table and the first blocks itself (the prober is policy), consults
configuration — `filesystems.csv`: content signature → filesystem binary;
`mounts.csv` or similar: volume identity → mount prefix — defines sub-ranges on
the driver, spawns **one filesystem process per volume**, hands it its block
channel at startup, and supervises it with the reap-and-rebuild idiom the
device manager proved: provider dies or medium leaves → kill the filesystem
process, retire its mounts; medium returns → re-probe, respawn, remount. The
boot volume is chosen by *content* (which volume carries /system/configuration
and /system/logs), closing the two-sticks question honestly.
**Filesystems (per volume, one process).** The proven unit everywhere from
Plan 9's `dossrv` to Minix to Fuchsia: block-client + engine + file-protocol
provider in one binary, one process per volume (9front practice; per-volume
fault isolation is what our supervision makes cheap). fat's shell becomes a
shared *filesystem harness* library before a second engine is written; a
partition walk is added in the volume manager (`partition.zig`) — the engine's
own MBR walk currently remains alongside it; write caching stays
inside the process (the anti-fsyncgate rule). Each mounts its prefixes into the
kernel mount table itself, exactly as today.
**Kernel: two small changes only.** `fs_unmount` gains ownership (only the
mounting endpoint's holder may unmount — possession-is-capability, consistent
with everything else), and the 8-slot mount table gets a declared bound or
growth once volumes multiply. The mount table stays the router; per-process
namespaces (Plan 9's extra) remain separable future work.
## Transport generality: NVMe and SATA against this design
The layering was chosen transport-agnostic on purpose (the 9front lesson:
filesystem services cannot tell a USB stick from an ATA disk); NVMe and AHCI
are the test of that claim, and they fit — with three named pressure points.
**What transfers untouched:** the block protocol (nothing USB in it), the
volume manager, partitions/ranges, filesystems, mounts, and the whole
enforcement stack — delegation, IOMMU confinement (these are first-party PCI
DMA masters, so confinement applies even more directly than USB),
reap-and-rebuild, the hot-plug lifecycle. Integration cost per transport is
what the architecture promises: one driver binary plus one `devices.csv` row.
**NVMe** shortens the stack — no bus/class split; the driver IS the
controller driver, one hop fewer than USB. Its structural novelty,
**namespaces** (hardware-native multiple volumes behind one controller), is
exactly the case the block protocol reserved on day one — still reserved, not
implemented (pressure point 1 below): decision 4 settles per-volume addressing
as per-sender confinement on a single endpoint, and names endpoint-per-volume
only as an unbuilt future refactor. **AHCI** is a shape choice, not a problem: an
HBA fronts up to 32 ports plus port multipliers — structurally a bus — so
either mirror USB (ahci-bus + a per-port disk driver: maximum restart
granularity, the proven shape) or mirror NVMe (one driver per controller, one
block channel per port). Leaning per-port processes for consistency with the
matrix-proven shape; genuinely open.
**The pressure points, honestly:**
1. **Multi-volume providers are reserved, not implemented.** The volume
manager flow assumes one provider, one volume; NVMe namespaces make
endpoint-per-volume real work with hardware demanding it.
2. **The current transport will bottleneck NVMe.** Synchronous call/reply,
one operation in flight, one bounce buffer — fine for a USB2 stick,
forfeits an NVMe drive's queue depth and per-queue MSI-X. Correctness
needs nothing; performance is gated on the shm-ring data plane
communication.md already names (Fuchsia's FIFO+VMO is the precedent).
NVMe is not worth building before that milestone.
3. **Media lifecycle ≠ device lifecycle.** ATAPI trays and SD card readers
(USB ones too, today) keep the DEVICE present while the MEDIUM comes and
goes — a removal trigger our channel-death lifecycle does not carry. The
fix is decision 7 below. (Related small item: fat's bounce sizing
hardcodes 512-byte sectors; 4Kn drives need the geometry honored
throughout.)
## The decisions on the table
1. **Partitions as driver-side sub-ranges** (mechanism in usb-storage, parsing
in the volume manager) — versus a separate partition *process* re-serving
block (Fuchsia's storage-host). Sub-ranges are ~20 lines of clamp in the
driver and keep the data path one hop; a separate process is purer layering
at the cost of a hop or session plumbing. Recommendation: sub-ranges.
2. **The volume manager as a new service** owning probe, spawn, supervision,
and mount policy — with `filesystems.csv` and the mount map as
configuration. Recommendation: yes; it is the missing policy home that fat
is currently squatting in.
3. **One filesystem process per volume** (fat's binary becomes "the FAT
implementation", spawned per FAT volume). Recommendation: yes — it extends
recompile-and-restart-live to filesystems and isolates corrupt media.
4. **Sub-range addressing — SETTLED as per-sender confinement at the
provider, one serving endpoint.** The deciding argument is precedent: the
xHCI bus already serves every class driver on one endpoint with authority
scoped by the kernel-stamped badge (the per-client device-token table) —
that IS danos's provider pattern, and per-badge range confinement is the
same pattern applied to blocks. The volume manager sets each filesystem
process's range on the driver; the driver clamps AND translates every
transfer by the sender's range, so filesystems address volume-relative
LBAs from 0. On the confined path the FAT engine's `base_lba` therefore
resolves to 0 on every access; the field and the engine's own MBR walk
remain in `engine.zig` as now-inert legacy code (the authoritative partition
walk lives in the volume manager's `partition.zig`), not yet deleted.
The enforcement point (the clamp at the provider, never in the consumer)
is what carries the security property; endpoint-per-volume would deliver
the same property only by inventing a multi-endpoint harness the pattern
does not need. It stays available as a future refactor if a multi-endpoint
harness ever exists for other reasons; the wire contract is identical
either way.
5. **`fs_unmount` ownership** — a defect fix more than a decision.
6. **Later, kept open**: the shm-ring data plane (communication.md already
names it as the 256-byte ceiling's unlock — Fuchsia's FIFO+VMO is the
precedent); format-level crash honesty (a Power-Safe-style journaling or COW
filesystem) once danos outgrows FAT; per-process namespaces.
7. **The media-presence event** (settled in principle; lands with the volume
manager): the block protocol gains a pushed event — `medium_changed`, with
present/absent and a change counter — produced by the storage driver from
its transport's native signal (SCSI UNIT ATTENTION / TEST UNIT READY for
USB and ATAPI, PxSSTS for AHCI, namespace-change AER for NVMe) and
consumed by the volume manager, which runs the SAME kill-retire-remount
path it runs on channel death — one lifecycle, two triggers. The driver
reports presence, never content; a pushed event carries no capability,
which the kernel already guarantees. The device staying while its medium
leaves is the one removable-media case the channel-death trigger cannot
see; without this event a swapped SD card would be served with the old
card's filesystem state.
8. **Volume identity, and the mount map as danos's fstab** (settled). The
lesson is Linux's own history: fstab keyed on `/dev/sda1` for years and
broke whenever a drive changed ports or enumeration order; `UUID=` entries
exist because device-path identity failed. danos skips that era: the mount
map (`volumes.csv` — configuration, read by the volume manager) keys on
**content identity, never port or discovery order**. Build status: rungs 1,
3, and 4 (GPT partition GUID, FAT serial + label, MBR signature + index) are
implemented (S1); rung 2 waits on a non-FAT engine. The `volumes.csv` map and
the id-derived mount path are built (S2): a volume's mount path IS its content
id (`/volumes/<id>`, e.g. `/volumes/fat-12345678`), or a `volumes.csv`
override; the label is display metadata a `volumes` query returns, never the
path. The ladder the prober reads off the medium, strongest first:
1. GPT partition GUID — 128-bit, unique, stable for the volume's life — **built (S1)**;
2. filesystem UUID (ext-family and most modern formats, in the superblock) *(planned)*;
3. FAT volume serial + label — 32 bits, weak (dd-cloned sticks share it)
but what real sticks carry — **built (S1)**;
4. MBR disk signature + partition index — **built**; a bare FAT with no
table takes index 0 over the whole device;
5. nothing — an anonymous volume: generated mount name, no persistence *(planned)*.
Consequences, each mechanical once identity keys the map: **moving a drive
to a different port changes nothing** — same identity, same mount point,
whether USB port, hub depth, SATA port, or a stick that left as USB and
returned in a SATA dock; **replug remounts at the same path** (a map lookup
once the map exists; today's single volume re-probes and remounts at the
fixed prefix, and remount-on-replug is bench-verified, not QEMU-tested); **the boot volume** is the
recorded identity of the volume carrying `/system/configuration`, findable
on any port; and **duplicate identity is a policy case, not a surprise** —
two cloned sticks at once: first keeps the mapped name, second mounts
suffixed and is logged loudly, never silently shadowed. Unknown identities
mount under a derived name (sanitized label, else generated) at
`/volumes/<name>` — the hierarchy's documented home for attached media,
which stands: `/system` is what danos IS; attached media is what it isn't.
danos never needs to MINT identifiers to detect volumes — detection only
reads — until it grows formatting, which brings the entropy question and
is deliberately out of scope here.