A V5 close-out audit (docs against the actual code) found the storage docs overclaiming in both directions: the status header was flipped to "built" but several body markers were not, and two passages describe mechanisms the code never implemented. Nine confirmed, adversarially verified against the source: Underclaims (marked Planned, actually built): - storage-architecture: driver named sub-ranges / range confinement (V2, tested by the block-range case); the pushed medium_changed event (published today from a TEST UNIT READY poll); ownership-gated fs_unmount (V0, process.zig gates it with EPERM); the filesystem-harness extraction (V1). Scoped the remaining *planned* to the genuinely-pending parts (native-signal translation, the volume manager consuming medium_changed). Overclaims (described, never built): - storage-architecture: the "FAT dirty flag on disk" guarantee — no on-disk dirty/clean-shutdown bit exists; only an in-memory device-dirty bool gating a device write-cache flush on close. Fixed in all three places. - storage-architecture: fat "acquires its own volume (first mass-storage child by enumeration order)" — the V3b flip removed self-acquisition; fat is handed its volume id and channel by the volume manager. - rationale + plan: the FAT engine's base_lba "deleted rather than moved" — it and the engine's MBR walk still exist as now-inert legacy; the authoritative walk lives in partition.zig. - rationale: NVMe namespaces "decision 4 settles as endpoint-per-volume" — decision 4 settles the opposite (per-sender confinement, one endpoint); endpoint-per-volume is named only as an unbuilt future refactor. - rationale: the five-rung identity ladder and volumes.csv map stated in flat present tense — only rung 4 (MBR signature + index) is built; added the build-status hedge and marked each rung. Docs only; no code or behavior change. Suite unaffected (127/127).
248 lines
16 KiB
Markdown
248 lines
16 KiB
Markdown
# The storage design rationale: why the stack is shaped this way
|
|
|
|
*2026-08-09. The design record behind
|
|
[storage-architecture.md](storage-architecture.md): the survey, the
|
|
trade-offs, and the decisions with their reasons — kept so future changes
|
|
argue against the evidence rather than rediscovering it. The questions that
|
|
drove it, verbatim: should the block protocol be separate from the VFS? how
|
|
do channels work with these block devices? how should we wire up different
|
|
filesystems? Grounded in how Minix 3, QNX Neutrino, Fuchsia, Redox, Plan 9,
|
|
and Linux each answered the same questions, and in exactly where our own
|
|
seams sat. The layering rule it serves: drivers are the lowest level
|
|
(hardware only); VFS and the filesystems are higher layers; the protocol
|
|
layer routes between them.*
|
|
|
|
## What the survey says, compressed
|
|
|
|
**Everyone separates block from VFS.** Even Plan 9, which unifies the *protocol*
|
|
(a disk is a file tree), keeps the layers: kernel `sd` serves raw ranges,
|
|
user-space filesystem servers consume them and serve files. The two poles are
|
|
Minix 3 (every layer a process: one VFS server, one filesystem server *per
|
|
mounted volume*, driver processes below — driver crashes proven survivable by
|
|
fault-injection, tens of thousands of injected faults) and QNX (filesystems as
|
|
DLLs inside the disk-driver process — fastest path, but one corrupt volume takes
|
|
every volume behind that controller down). Fuchsia is the modern capability-
|
|
native reference: a tiny block contract that every layer speaks, control channel
|
|
separate from a data fast-path with pre-registered buffers, filesystems as
|
|
separate processes launched by a storage-policy component (`fshost`) that probes
|
|
content and hands each filesystem its block channel at startup. The filesystem
|
|
never discovers devices.
|
|
|
|
**Partition tables are parsed in user space, below the filesystem, above the
|
|
raw device — never in the kernel, never in the filesystem engine.** Fuchsia
|
|
tried partitions-as-drivers and formally reversed it (RFC-0257). Plan 9 has the
|
|
cleanest split of all: the kernel offers only a *mechanism* — "create a named
|
|
sub-range of this disk" — and a user-space prober parses the table and issues
|
|
those commands. QNX (`io-blk`) and Minix (driver-side shared library) agree on
|
|
placement: the block layer multiplexes one physical block object into several
|
|
logical block objects **of the same contract**.
|
|
|
|
**The failure lessons are unanimous.** Linux's shared page cache across the
|
|
filesystem boundary produced fsyncgate (write errors observed by the wrong
|
|
process, dirty pages marked clean); the lesson is to keep write caching *inside*
|
|
the filesystem process, so an error surfaces on the channel that owns the
|
|
volume. Linux's `errors=remount-ro`/lazy-unmount half-alive states exist because
|
|
a monolith cannot kill and respawn a filesystem; Plan 9 leaves dead servers'
|
|
mounts in the namespace erroring forever. A restartable-parts system should have
|
|
exactly one surprise-removal path: kill the filesystem process, retire its
|
|
mounts, respawn on return. QNX's deepest lesson: after years of removable-media
|
|
pain they concluded detection isn't enough and built a copy-on-write filesystem
|
|
(Power-Safe) — *format-level* crash honesty. And QNX's `mcd` (a standalone
|
|
policy daemon with declarative insertion/removal rules) plus Fuchsia's `fshost`
|
|
agree that media policy is one dedicated component, not code scattered through
|
|
filesystems.
|
|
|
|
## Where we already stand
|
|
|
|
Closer than expected. The block protocol is 61 lines, five verbs, and its
|
|
docblock already reserves `Header.target` for volumes ("targets = volumes,
|
|
`enumerate` lists them; nothing else about the protocol changes"). usb-storage
|
|
is content-blind — it reads block 0 only as a bring-up self-check and parses
|
|
nothing. The vfs protocol is already served by a second provider in production
|
|
(init's registry serves it synthetically), so a second filesystem is *not* a
|
|
protocol problem. The FAT service is already internally split: a host-testable
|
|
engine (1659 + 310 lines) behind a four-function BlockDevice vtable, and a
|
|
434-line IPC shell. The kernel mount table already has the right restart
|
|
semantics (remount-replace; lazy dead-endpoint sweep on resolve).
|
|
|
|
The misplacements, all small: the MBR walk lives *inside the FAT engine*
|
|
(`engine.zig:196-212`) — every future engine would duplicate it; fat hardcodes
|
|
its acquisition policy ("first mass-storage child, mount at /volumes/usb") —
|
|
that is policy in a filesystem binary; `fs_unmount` is gated by nothing (any
|
|
process may unmount any prefix — latent now, an obvious cross-tenant hole once
|
|
mounts multiply); and fat's `mounted` flag is one-way — no path back after its
|
|
block channel dies.
|
|
|
|
## The proposed shape
|
|
|
|
Three layers, matching the stated rule, every boundary a channel:
|
|
|
|
**Block (driver layer, mechanism).** usb-storage stays hardware→blocks and
|
|
content-blind. It gains one mechanism, Plan 9's: *named sub-ranges*. A
|
|
`define_range`-style op creates a logical block object (a partition) clamped
|
|
and offset onto the disk; sub-ranges are `target` ids on the same endpoint, or
|
|
separate endpoints handed out per range — either way they speak the **same
|
|
block contract** (Fuchsia's closed-under-layering rule), so a filesystem cannot
|
|
tell whole-disk from partition. The driver never parses a table; it clamps
|
|
ranges it is told about. Offsets translate at definition time, so the data path
|
|
stays one hop (Fuchsia's session-mapping trick, for free).
|
|
|
|
**Volumes (service layer, policy).** One new service — the *volume manager*
|
|
(fshost/mcd-shaped, init-supervised; the device manager stays devices-only). It
|
|
subscribes to the device manager's `child_added`/`child_removed`, consumer-
|
|
hellos for each new mass-storage provider's block channel, reads the partition
|
|
table and the first blocks itself (the prober is policy), consults
|
|
configuration — `filesystems.csv`: content signature → filesystem binary;
|
|
`mounts.csv` or similar: volume identity → mount prefix — defines sub-ranges on
|
|
the driver, spawns **one filesystem process per volume**, hands it its block
|
|
channel at startup, and supervises it with the reap-and-rebuild idiom the
|
|
device manager proved: provider dies or medium leaves → kill the filesystem
|
|
process, retire its mounts; medium returns → re-probe, respawn, remount. The
|
|
boot volume is chosen by *content* (which volume carries /system/configuration
|
|
and /system/logs), closing the two-sticks question honestly.
|
|
|
|
**Filesystems (per volume, one process).** The proven unit everywhere from
|
|
Plan 9's `dossrv` to Minix to Fuchsia: block-client + engine + file-protocol
|
|
provider in one binary, one process per volume (9front practice; per-volume
|
|
fault isolation is what our supervision makes cheap). fat's shell becomes a
|
|
shared *filesystem harness* library before a second engine is written; a
|
|
partition walk is added in the volume manager (`partition.zig`) — the engine's
|
|
own MBR walk currently remains alongside it; write caching stays
|
|
inside the process (the anti-fsyncgate rule). Each mounts its prefixes into the
|
|
kernel mount table itself, exactly as today.
|
|
|
|
**Kernel: two small changes only.** `fs_unmount` gains ownership (only the
|
|
mounting endpoint's holder may unmount — possession-is-capability, consistent
|
|
with everything else), and the 8-slot mount table gets a declared bound or
|
|
growth once volumes multiply. The mount table stays the router; per-process
|
|
namespaces (Plan 9's extra) remain separable future work.
|
|
|
|
## Transport generality: NVMe and SATA against this design
|
|
|
|
The layering was chosen transport-agnostic on purpose (the 9front lesson:
|
|
filesystem services cannot tell a USB stick from an ATA disk); NVMe and AHCI
|
|
are the test of that claim, and they fit — with three named pressure points.
|
|
|
|
**What transfers untouched:** the block protocol (nothing USB in it), the
|
|
volume manager, partitions/ranges, filesystems, mounts, and the whole
|
|
enforcement stack — delegation, IOMMU confinement (these are first-party PCI
|
|
DMA masters, so confinement applies even more directly than USB),
|
|
reap-and-rebuild, the hot-plug lifecycle. Integration cost per transport is
|
|
what the architecture promises: one driver binary plus one `devices.csv` row.
|
|
|
|
**NVMe** shortens the stack — no bus/class split; the driver IS the
|
|
controller driver, one hop fewer than USB. Its structural novelty,
|
|
**namespaces** (hardware-native multiple volumes behind one controller), is
|
|
exactly the case the block protocol reserved on day one — still reserved, not
|
|
implemented (pressure point 1 below): decision 4 settles per-volume addressing
|
|
as per-sender confinement on a single endpoint, and names endpoint-per-volume
|
|
only as an unbuilt future refactor. **AHCI** is a shape choice, not a problem: an
|
|
HBA fronts up to 32 ports plus port multipliers — structurally a bus — so
|
|
either mirror USB (ahci-bus + a per-port disk driver: maximum restart
|
|
granularity, the proven shape) or mirror NVMe (one driver per controller, one
|
|
block channel per port). Leaning per-port processes for consistency with the
|
|
matrix-proven shape; genuinely open.
|
|
|
|
**The pressure points, honestly:**
|
|
|
|
1. **Multi-volume providers are reserved, not implemented.** The volume
|
|
manager flow assumes one provider, one volume; NVMe namespaces make
|
|
endpoint-per-volume real work with hardware demanding it.
|
|
2. **The current transport will bottleneck NVMe.** Synchronous call/reply,
|
|
one operation in flight, one bounce buffer — fine for a USB2 stick,
|
|
forfeits an NVMe drive's queue depth and per-queue MSI-X. Correctness
|
|
needs nothing; performance is gated on the shm-ring data plane
|
|
communication.md already names (Fuchsia's FIFO+VMO is the precedent).
|
|
NVMe is not worth building before that milestone.
|
|
3. **Media lifecycle ≠ device lifecycle.** ATAPI trays and SD card readers
|
|
(USB ones too, today) keep the DEVICE present while the MEDIUM comes and
|
|
goes — a removal trigger our channel-death lifecycle does not carry. The
|
|
fix is decision 7 below. (Related small item: fat's bounce sizing
|
|
hardcodes 512-byte sectors; 4Kn drives need the geometry honored
|
|
throughout.)
|
|
|
|
## The decisions on the table
|
|
|
|
1. **Partitions as driver-side sub-ranges** (mechanism in usb-storage, parsing
|
|
in the volume manager) — versus a separate partition *process* re-serving
|
|
block (Fuchsia's storage-host). Sub-ranges are ~20 lines of clamp in the
|
|
driver and keep the data path one hop; a separate process is purer layering
|
|
at the cost of a hop or session plumbing. Recommendation: sub-ranges.
|
|
2. **The volume manager as a new service** owning probe, spawn, supervision,
|
|
and mount policy — with `filesystems.csv` and the mount map as
|
|
configuration. Recommendation: yes; it is the missing policy home that fat
|
|
is currently squatting in.
|
|
3. **One filesystem process per volume** (fat's binary becomes "the FAT
|
|
implementation", spawned per FAT volume). Recommendation: yes — it extends
|
|
recompile-and-restart-live to filesystems and isolates corrupt media.
|
|
4. **Sub-range addressing — SETTLED as per-sender confinement at the
|
|
provider, one serving endpoint.** The deciding argument is precedent: the
|
|
xHCI bus already serves every class driver on one endpoint with authority
|
|
scoped by the kernel-stamped badge (the per-client device-token table) —
|
|
that IS danos's provider pattern, and per-badge range confinement is the
|
|
same pattern applied to blocks. The volume manager sets each filesystem
|
|
process's range on the driver; the driver clamps AND translates every
|
|
transfer by the sender's range, so filesystems address volume-relative
|
|
LBAs from 0. On the confined path the FAT engine's `base_lba` therefore
|
|
resolves to 0 on every access; the field and the engine's own MBR walk
|
|
remain in `engine.zig` as now-inert legacy code (the authoritative partition
|
|
walk lives in the volume manager's `partition.zig`), not yet deleted.
|
|
The enforcement point (the clamp at the provider, never in the consumer)
|
|
is what carries the security property; endpoint-per-volume would deliver
|
|
the same property only by inventing a multi-endpoint harness the pattern
|
|
does not need. It stays available as a future refactor if a multi-endpoint
|
|
harness ever exists for other reasons; the wire contract is identical
|
|
either way.
|
|
5. **`fs_unmount` ownership** — a defect fix more than a decision.
|
|
6. **Later, kept open**: the shm-ring data plane (communication.md already
|
|
names it as the 256-byte ceiling's unlock — Fuchsia's FIFO+VMO is the
|
|
precedent); format-level crash honesty (a Power-Safe-style journaling or COW
|
|
filesystem) once danos outgrows FAT; per-process namespaces.
|
|
7. **The media-presence event** (settled in principle; lands with the volume
|
|
manager): the block protocol gains a pushed event — `medium_changed`, with
|
|
present/absent and a change counter — produced by the storage driver from
|
|
its transport's native signal (SCSI UNIT ATTENTION / TEST UNIT READY for
|
|
USB and ATAPI, PxSSTS for AHCI, namespace-change AER for NVMe) and
|
|
consumed by the volume manager, which runs the SAME kill-retire-remount
|
|
path it runs on channel death — one lifecycle, two triggers. The driver
|
|
reports presence, never content; a pushed event carries no capability,
|
|
which the kernel already guarantees. The device staying while its medium
|
|
leaves is the one removable-media case the channel-death trigger cannot
|
|
see; without this event a swapped SD card would be served with the old
|
|
card's filesystem state.
|
|
8. **Volume identity, and the mount map as danos's fstab** (settled). The
|
|
lesson is Linux's own history: fstab keyed on `/dev/sda1` for years and
|
|
broke whenever a drive changed ports or enumeration order; `UUID=` entries
|
|
exist because device-path identity failed. danos skips that era: the mount
|
|
map (`volumes.csv` — configuration, read by the volume manager) keys on
|
|
**content identity, never port or discovery order**. Build status: only
|
|
rung 4 (MBR signature + partition index) is implemented today; the fuller
|
|
rungs and the `volumes.csv` map itself land with the identity ladder, so
|
|
today a single volume mounts at the fixed `/volumes/usb` and its recorded
|
|
identity is not yet consulted to pick a path. The target ladder the prober
|
|
reads off the medium, strongest first:
|
|
1. GPT partition GUID — 128-bit, unique, stable for the volume's life *(planned)*;
|
|
2. filesystem UUID (ext-family and most modern formats, in the superblock) *(planned)*;
|
|
3. FAT volume serial + label — 32 bits, weak (dd-cloned sticks share it)
|
|
but what real sticks carry *(planned)*;
|
|
4. MBR disk signature + partition index — **built**; a bare FAT with no
|
|
table takes index 0 over the whole device;
|
|
5. nothing — an anonymous volume: generated mount name, no persistence *(planned)*.
|
|
|
|
Consequences, each mechanical once identity keys the map: **moving a drive
|
|
to a different port changes nothing** — same identity, same mount point,
|
|
whether USB port, hub depth, SATA port, or a stick that left as USB and
|
|
returned in a SATA dock; **replug remounts at the same path** (a map lookup
|
|
once the map exists; today's single volume re-probes and remounts at the
|
|
fixed prefix, and remount-on-replug is bench-verified, not QEMU-tested); **the boot volume** is the
|
|
recorded identity of the volume carrying `/system/configuration`, findable
|
|
on any port; and **duplicate identity is a policy case, not a surprise** —
|
|
two cloned sticks at once: first keeps the mapped name, second mounts
|
|
suffixed and is logged loudly, never silently shadowed. Unknown identities
|
|
mount under a derived name (sanitized label, else generated) at
|
|
`/volumes/<name>` — the hierarchy's documented home for attached media,
|
|
which stands: `/system` is what danos IS; attached media is what it isn't.
|
|
danos never needs to MINT identifiers to detect volumes — detection only
|
|
reads — until it grows formatting, which brings the entropy question and
|
|
is deliberately out of scope here.
|