storage-stack-discussion.md was misplaced at the docs root — that level is for track plans; this is the file-system domain's design record. Moved to file-system-development/storage-design-rationale.md, renamed to say what it is, cross-references updated.
213 lines
13 KiB
Markdown
213 lines
13 KiB
Markdown
# The storage architecture: layers, boundaries, responsibilities
|
|
|
|
> **Status:** the layered model below is the settled design
|
|
> ([storage-design-rationale.md](storage-design-rationale.md) records how it
|
|
> was reached and what the surveyed systems taught). The data path — vfs
|
|
> protocol, kernel mount routing, the FAT service, the block protocol,
|
|
> usb-storage — is **built**. The volume manager, the driver's range
|
|
> mechanism, per-volume filesystem spawning, and the removal lifecycle are
|
|
> **planned**; until they land, the FAT service performs volume-manager duties
|
|
> itself (marked below). This document is the reference for both states.
|
|
|
|
## The model
|
|
|
|
One pattern, applied twice: an application is a client of a service over a
|
|
protocol; a service is a client of a driver over a protocol.
|
|
|
|
```
|
|
application
|
|
│ vfs protocol (routed by the kernel mount table)
|
|
▼
|
|
filesystem service ── one process per volume
|
|
│ inside: [file-protocol shell] ── Zig API ── [fs engine]
|
|
│ block protocol (channel received at spawn)
|
|
▼
|
|
storage driver ── one process per device (usb-storage per stick)
|
|
│ usb-transfer protocol (channel received via hello, by lineage)
|
|
▼
|
|
bus driver ── one process per controller (usb-xhci-bus)
|
|
│ hardware
|
|
```
|
|
|
|
Three kinds of boundary, deliberately different:
|
|
|
|
- **Protocol boundaries** sit where processes meet (vfs, block, usb-transfer).
|
|
Each side is independently restartable; channels are established by
|
|
capability handoff, never by registry names (communication.md
|
|
"Establishment: two planes"). Steady-state file I/O crosses exactly two —
|
|
vfs and block.
|
|
- **Library boundaries** sit where concerns meet inside one process. The
|
|
filesystem service's engine (on-disk format logic, host-testable, behind
|
|
the four-function `BlockDevice` vtable) and shell (IPC serving, open-node
|
|
table, mount registration) talk through a Zig API. No protocol between
|
|
them: they share fate regardless — a corrupt engine corrupts the answers
|
|
either way — so a channel there would add a hop per file operation and buy
|
|
nothing. The seam exists for compile-time testability and reuse; the
|
|
*process* is the restart unit.
|
|
- **Control-plane relationships** sit beside the data path, never on it. Two
|
|
supervisors, one per layer: the **device manager** wires and revives the
|
|
device layers (bus and storage drivers — devices only); the **volume
|
|
manager** *(planned)* wires and revives the volume layer (filesystem
|
|
services). Neither touches steady-state I/O.
|
|
|
|
## Who does what
|
|
|
|
**Bus driver** (usb-xhci-bus, per controller): hardware only. Enumerates,
|
|
reports children to the device manager, serves the transfer contract to its
|
|
own children's class drivers, routed by lineage.
|
|
|
|
**Storage driver** (usb-storage, per device; later nvme, ahci): hardware →
|
|
blocks. Speaks its transport upward from its device; serves the block
|
|
protocol (geometry, read, write, flush, attach, detach — and, *planned*, the
|
|
pushed `medium_changed` presence event, translated from the transport's
|
|
native signal). **Content-blind, permanently**:
|
|
it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no
|
|
GPT, no filesystem magic, ever. *(Planned)* it gains one content-blind
|
|
mechanism: **named sub-ranges** — "serve blocks [a, b) as a channel of the
|
|
same block contract". It clamps and offsets; it never knows the numbers came
|
|
from a partition table. The clamp lives here and nowhere else because a
|
|
channel must carry exactly the authority it grants: handing a filesystem the
|
|
whole disk plus a polite base offset would let a buggy or compromised
|
|
filesystem scribble the neighboring partition — the same authority-overshoot
|
|
the device-authority track eliminated for MMIO and DMA.
|
|
|
|
**Volume manager** *(planned; today the FAT service squats on these duties)*:
|
|
the policy home of the volume layer, one service, supervised by init. It
|
|
subscribes to the device manager's child events; when a storage provider
|
|
appears it consumer-hellos for the block channel, reads the partition table
|
|
and the first blocks itself (**it** is the prober), consults its
|
|
configuration, defines sub-ranges on the driver, spawns the matching
|
|
filesystem service per volume with that volume's channel, supervises it, and
|
|
decides mount placement. Its tables are CSV configuration, read by it (the
|
|
policy), enforced by nobody else:
|
|
|
|
- `filesystems.csv` — content signature → filesystem binary. Adding a
|
|
filesystem adds a row.
|
|
- the mount map — volume identity → mount prefix. The boot volume is chosen
|
|
by **content** (the volume carrying `/system/configuration` and
|
|
`/system/logs`), never by port or arrival order.
|
|
|
|
**Filesystem service** (the FAT service today; one process per volume): the
|
|
proven unit — block-client + engine + file-protocol provider in one binary. It
|
|
receives its block channel at spawn; it never discovers devices. It registers
|
|
its own mounts with the kernel; its write cache lives inside the process, so a
|
|
write error is observed by the code that owns the volume and surfaces on the
|
|
owning channel (the anti-fsyncgate rule — never a system-wide dirty pool).
|
|
*(Today, interim:)* fat also acquires its own volume (first mass-storage child
|
|
by enumeration order) and hardcodes its mount prefixes; both migrate to the
|
|
volume manager.
|
|
|
|
**Kernel** (mechanism only): the mount table routes paths to backend
|
|
endpoints — resolve and redirect, never data. Remount-replace is the restart
|
|
story; a dead backend's slot is swept lazily on the next resolution.
|
|
*(Planned fix:)* `fs_unmount` gains ownership — only the mounting endpoint's
|
|
holder may unmount; today it is ungated, which is safe with one mount owner
|
|
and wrong with several.
|
|
|
|
## Adding a filesystem
|
|
|
|
1. **Write the engine**: pure Zig, no IPC, behind the `BlockDevice` vtable
|
|
(four functions). On-disk layout in `align(1)` extern structs. Host-test it
|
|
with `zig build test` fixtures — the FAT engine
|
|
([engine.zig](../../system/services/fat/engine.zig),
|
|
[on-disk.zig](../../system/services/fat/on-disk.zig)) is the template, and
|
|
its ~500 lines of host tests the standard.
|
|
2. **Reuse the shell**: the filesystem harness — establishment, the
|
|
badge-scoped open-node table, the nine vfs-protocol handlers, mount
|
|
registration, the removal path — is shared code, not per-filesystem code.
|
|
*(Planned: extracted from fat's 434-line shell into a library before the
|
|
second engine is written.)* An engine plus a `main` wiring it into the
|
|
harness is a complete filesystem service.
|
|
3. **Add the configuration row**: one line in `filesystems.csv` mapping the
|
|
on-disk signature to the binary. No other component changes: the volume
|
|
manager probes, matches, spawns; the vfs protocol is already
|
|
backend-neutral (a second provider serves it in production today — init's
|
|
synthetic registry backend).
|
|
4. **Prove removal**: extend the removable-media suite (below) for the new
|
|
filesystem — surprise-yank during writes must lose only what was
|
|
unflushed, said honestly, with the volume consistent enough to remount.
|
|
|
|
The engine never sees: partitions (it receives a volume-shaped block channel;
|
|
base offsets are the driver's clamp), device discovery (the channel arrives at
|
|
spawn), mount policy (prefixes are handed to it), or other volumes (one
|
|
process, one volume).
|
|
|
|
## Removable media: responsibilities on removal, per layer
|
|
|
|
The design rule, learned from what Linux cannot do: there is exactly **one**
|
|
surprise-removal path — kill the filesystem process, retire its mounts,
|
|
respawn on return. No half-alive states, no `remount-ro`, no mounts that
|
|
error forever (Plan 9's dead-server wart).
|
|
|
|
The path has **two triggers, one lifecycle**: the *device* leaving (the
|
|
storage driver dies — channel death, the table below), and the *medium*
|
|
leaving while the device stays (an SD card pulled from its reader, an ATAPI
|
|
tray opened — including USB card readers today). The second trigger is a
|
|
pushed `medium_changed` event on the block protocol *(planned)*: the storage
|
|
driver translates its transport's native signal (SCSI UNIT ATTENTION, AHCI
|
|
PxSSTS, NVMe namespace-change AER) into presence-changed — never content —
|
|
and the volume manager runs the same kill-retire path, then re-probes on
|
|
medium return exactly as on device return. Without it, a swapped card would
|
|
be served with the previous card's filesystem state.
|
|
|
|
| Layer | Observes | Must do | Guarantees |
|
|
|---|---|---|---|
|
|
| Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick |
|
|
| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
|
|
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
|
|
| Volume manager *(planned)* | the storage provider's channel death / the manager's child events | kill each filesystem service of that device's volumes; retire their kernel mounts; remember the volume identity | one removal path; mounts never dangle; log persistence stops *cleanly* |
|
|
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the FAT dirty flag on disk marks the unclean removal; the process never serves from behind a dead channel |
|
|
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); *(planned)* ownership-gated `fs_unmount` | resolution under a dead mount is `not_found`, not a hang |
|
|
| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
|
|
|
|
**On return** the same table runs upward in reverse: the bus re-enumerates and
|
|
re-reports; the device manager respawns the storage driver (built); the volume
|
|
manager re-probes — same content, same volume identity — respawns the
|
|
filesystem service, and remounts at the same prefix; applications see the
|
|
subtree reappear. A different stick in the same port is a *different volume*
|
|
(identity is content, not port) and mounts wherever its identity says —
|
|
possibly nowhere but `/volumes/…`.
|
|
|
|
What is lost on a surprise yank is exactly the write-back window of the
|
|
filesystem service, no more: the engine owns its cache, so the blast radius of
|
|
a yank is one volume's unflushed writes, marked by the on-disk dirty flag. A
|
|
filesystem format with better crash honesty (journaling, copy-on-write — the
|
|
lesson of QNX's Power-Safe) narrows that window further and slots in as
|
|
implementation N+1 through the table above, changing nothing else.
|
|
|
|
## How the lifecycle is enforced
|
|
|
|
Responsibility tables are convention; convention is worthless (the bounds
|
|
audit's lesson). The volume lifecycle is enforced with the same three levers
|
|
the device lifecycle already uses, mapped one-to-one:
|
|
|
|
1. **A filesystem cannot acquire — it can only be given.** Filesystem
|
|
binaries hold NO establishment grants: no `open device-manager`, no
|
|
registry name to look up. The only block channel a filesystem process ever
|
|
has is the one handed to it at spawn by the volume manager. Serving the
|
|
wrong volume, a second volume, or a self-discovered volume is not
|
|
forbidden but *impossible* — the same way a driver cannot claim hardware
|
|
it was not delegated (device-authority.md).
|
|
2. **The volume manager supervises with teeth.** Mount-within-deadline or be
|
|
stopped (a filesystem wedged on a corrupt volume is killed and crash-loop
|
|
capped; the volume is marked bad, not retried forever). On removal the
|
|
manager KILLS the filesystem process and retires its mounts — the polite
|
|
observe-`EPEER`-and-exit path is an optimization; the kill is the
|
|
guarantee, because a process cannot be trusted to observe its own
|
|
obsolescence (the reap argument, proven on the device tree).
|
|
3. **The shared harness is how every filesystem inherits the lifecycle by
|
|
construction.** The harness — not the engine — owns the state machine:
|
|
establishment at spawn, mount registration, the dirty-flag set/clear
|
|
bracket, error-out-and-exit on channel death. The engine sits behind the
|
|
four-function vtable and never sees a channel; it cannot opt out of the
|
|
lifecycle for the same reason it cannot find a device. This is why the
|
|
harness is extracted BEFORE the second engine is written.
|
|
|
|
Beneath all three, two kernel mechanisms: ownership-gated `fs_unmount` (only
|
|
the mounting endpoint's holder may unmount), and the lazy dead-endpoint mount
|
|
sweep as the backstop nothing can disable. And above them, the check: a
|
|
**lifecycle conformance drill** — mount, serve, yank mid-write, verify honest
|
|
loss, replug, remount — parameterized over filesystem implementations, so
|
|
"danos supports filesystem X" MEANS "X passes the drill through the harness",
|
|
exactly as provider conformance means passing the reserved-verb suite.
|