A V5 close-out audit (docs against the actual code) found the storage docs overclaiming in both directions: the status header was flipped to "built" but several body markers were not, and two passages describe mechanisms the code never implemented. Nine confirmed, adversarially verified against the source: Underclaims (marked Planned, actually built): - storage-architecture: driver named sub-ranges / range confinement (V2, tested by the block-range case); the pushed medium_changed event (published today from a TEST UNIT READY poll); ownership-gated fs_unmount (V0, process.zig gates it with EPERM); the filesystem-harness extraction (V1). Scoped the remaining *planned* to the genuinely-pending parts (native-signal translation, the volume manager consuming medium_changed). Overclaims (described, never built): - storage-architecture: the "FAT dirty flag on disk" guarantee — no on-disk dirty/clean-shutdown bit exists; only an in-memory device-dirty bool gating a device write-cache flush on close. Fixed in all three places. - storage-architecture: fat "acquires its own volume (first mass-storage child by enumeration order)" — the V3b flip removed self-acquisition; fat is handed its volume id and channel by the volume manager. - rationale + plan: the FAT engine's base_lba "deleted rather than moved" — it and the engine's MBR walk still exist as now-inert legacy; the authoritative walk lives in partition.zig. - rationale: NVMe namespaces "decision 4 settles as endpoint-per-volume" — decision 4 settles the opposite (per-sender confinement, one endpoint); endpoint-per-volume is named only as an unbuilt future refactor. - rationale: the five-rung identity ladder and volumes.csv map stated in flat present tense — only rung 4 (MBR signature + index) is built; added the build-status hedge and marked each rung. Docs only; no code or behavior change. Suite unaffected (127/127).
275 lines
17 KiB
Markdown
275 lines
17 KiB
Markdown
# The storage architecture: layers, boundaries, responsibilities
|
||
|
||
> **Status:** the layered model below is the settled design
|
||
> ([storage-design-rationale.md](storage-design-rationale.md) records how it was
|
||
> reached, and [volume-manager-plan.md](../volume-manager-plan.md) how it was
|
||
> built). **Built** (the volume-manager track, V0–V4): the data path, the driver
|
||
> range confinement (per-sender clamp + the confinement gate), the `medium_changed`
|
||
> presence event, the volume manager itself — it probes the partition table,
|
||
> confines each filesystem to its partition, spawns one filesystem per volume, and
|
||
> supervises it — and the removal half of the lifecycle (a pulled stick unmounts).
|
||
> **Still pending**: the fuller identity ladder and the `volumes.csv` mount map,
|
||
> multi-volume (one FAT volume today), the volume manager *consuming*
|
||
> `medium_changed` (removal is detected by device-presence polling; the event is
|
||
> published but only a card-reader medium change needs the subscription), and the
|
||
> remount-on-replug end-to-end (the logic is in place; QEMU can't re-present the
|
||
> boot-controller device, so it is bench-verified). A few markers below are left
|
||
> where a duty is still pending.
|
||
|
||
## The model
|
||
|
||
One pattern, applied twice: an application is a client of a service over a
|
||
protocol; a service is a client of a driver over a protocol.
|
||
|
||
```
|
||
application
|
||
│ vfs protocol (routed by the kernel mount table)
|
||
▼
|
||
filesystem service ── one process per volume
|
||
│ inside: [file-protocol shell] ── Zig API ── [fs engine]
|
||
│ block protocol (channel received at spawn)
|
||
▼
|
||
storage driver ── one process per device (usb-storage per stick)
|
||
│ usb-transfer protocol (channel received via hello, by lineage)
|
||
▼
|
||
bus driver ── one process per controller (usb-xhci-bus)
|
||
│ hardware
|
||
```
|
||
|
||
Three kinds of boundary, deliberately different:
|
||
|
||
- **Protocol boundaries** sit where processes meet (vfs, block, usb-transfer).
|
||
Each side is independently restartable; channels are established by
|
||
capability handoff, never by registry names (communication.md
|
||
"Establishment: two planes"). Steady-state file I/O crosses exactly two —
|
||
vfs and block.
|
||
- **Library boundaries** sit where concerns meet inside one process. The
|
||
filesystem service's engine (on-disk format logic, host-testable, behind
|
||
the four-function `BlockDevice` vtable) and shell (IPC serving, open-node
|
||
table, mount registration) talk through a Zig API. No protocol between
|
||
them: they share fate regardless — a corrupt engine corrupts the answers
|
||
either way — so a channel there would add a hop per file operation and buy
|
||
nothing. The seam exists for compile-time testability and reuse; the
|
||
*process* is the restart unit.
|
||
- **Control-plane relationships** sit beside the data path, never on it. Two
|
||
supervisors, one per layer: the **device manager** wires and revives the
|
||
device layers (bus and storage drivers — devices only); the **volume
|
||
manager** *(built)* wires and revives the volume layer (filesystem
|
||
services). Neither touches steady-state I/O.
|
||
|
||
## Who does what
|
||
|
||
**Bus driver** (usb-xhci-bus, per controller): hardware only. Enumerates,
|
||
reports children to the device manager, serves the transfer contract to its
|
||
own children's class drivers, routed by lineage.
|
||
|
||
**Storage driver** (usb-storage, per device; later nvme, ahci): hardware →
|
||
blocks. Speaks its transport upward from its device; serves the block
|
||
protocol (geometry, read, write, flush, attach, detach — and the pushed
|
||
`medium_changed` presence event, published today from a slow TEST UNIT READY
|
||
poll; *planned*: translating it from the transport's native signal instead of
|
||
polling). **Content-blind, permanently**:
|
||
it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no
|
||
GPT, no filesystem magic, ever. It has one content-blind mechanism *(built)*:
|
||
**named sub-ranges** — "serve blocks [a, b) as a channel of the
|
||
same block contract". It clamps and offsets; it never knows the numbers came
|
||
from a partition table. The clamp lives here and nowhere else because a
|
||
channel must carry exactly the authority it grants: handing a filesystem the
|
||
whole disk plus a polite base offset would let a buggy or compromised
|
||
filesystem scribble the neighboring partition — the same authority-overshoot
|
||
the device-authority track eliminated for MMIO and DMA.
|
||
|
||
**Volume manager** *(built; `system/services/volume-manager`)*: the policy home
|
||
of the volume layer, one service, supervised by init. It watches the device
|
||
manager's tree for a storage provider; when one appears it consumer-hellos for
|
||
the block channel, reads the partition table and the first blocks itself
|
||
(**it** is the prober), defines the volume's sub-range on the driver, spawns the
|
||
matching filesystem service confined to that range, and supervises it (backoff,
|
||
crash-loop cap). *(Pending)*: it decides mount placement from `volumes.csv` and
|
||
picks the filesystem binary from `filesystems.csv` — today it hands every
|
||
FAT-shaped volume to the FAT service and the FAT service carries hardcoded mount
|
||
prefixes. Those tables are CSV configuration, read by it (the policy), enforced
|
||
by nobody else:
|
||
|
||
- `filesystems.csv` *(pending)* — content signature → filesystem binary. Adding
|
||
a filesystem adds a row.
|
||
- `volumes.csv` *(pending)* — the mount map, danos's fstab: **volume identity → mount
|
||
prefix**, keyed on content identity and never on port, path, or arrival
|
||
order (the lesson of Linux's `/dev/sda1`-era fstab, which broke on every
|
||
port move until `UUID=` replaced it). Identity is read off the medium by
|
||
the prober, strongest first: GPT partition GUID → filesystem UUID → FAT
|
||
serial + label → MBR signature + partition index → anonymous (generated
|
||
name, no persistence). Same identity, same mount point: a drive moved to
|
||
another port — USB to another hub, SATA to another bay, even a stick
|
||
returning in a different dock — lands exactly where it was. Duplicate
|
||
identity (cloned sticks, together) is policy: first keeps the name, the
|
||
second mounts suffixed and is logged loudly. The boot volume is the
|
||
recorded identity of the volume carrying `/system/configuration` and
|
||
`/system/logs`, findable on any port. Unknown volumes mount under
|
||
`/volumes/<derived name>`.
|
||
|
||
**Filesystem service** (the FAT service today; one process per volume): the
|
||
proven unit — block-client + engine + file-protocol provider in one binary. It
|
||
receives its block channel at spawn; it never discovers devices. It registers
|
||
its own mounts with the kernel; its write cache lives inside the process, so a
|
||
write error is observed by the code that owns the volume and surfaces on the
|
||
owning channel (the anti-fsyncgate rule — never a system-wide dirty pool).
|
||
*(Today, interim:)* fat still hardcodes its mount prefixes (`/volumes/usb` plus
|
||
the two boot-volume hierarchy subtrees it rewrites in place); a `volumes.csv`
|
||
mount map will migrate that to the volume manager. It no longer self-acquires a
|
||
volume — the V3b flip made it receive its volume id at spawn and its block
|
||
channel from the volume manager's hello reply, consistent with "it never
|
||
discovers devices" above.
|
||
|
||
**Kernel** (mechanism only): the mount table routes paths to backend
|
||
endpoints — resolve and redirect, never data. Remount-replace is the restart
|
||
story; a dead backend's slot is swept lazily on the next resolution.
|
||
`fs_unmount` is ownership-gated *(built, V0)* — only the mounting task may
|
||
unmount its own prefix; anyone else is refused (`EPERM`). No dead-owner
|
||
exception: a dead owner's slot is swept lazily by resolution, and restart goes
|
||
through remount-replace, never through a stranger's unmount.
|
||
|
||
## Adding a filesystem
|
||
|
||
1. **Write the engine**: pure Zig, no IPC, behind the `BlockDevice` vtable
|
||
(four functions). On-disk layout in `align(1)` extern structs. Host-test it
|
||
with `zig build test` fixtures — the FAT engine
|
||
([engine.zig](../../system/services/fat/engine.zig),
|
||
[on-disk.zig](../../system/services/fat/on-disk.zig)) is the template, and
|
||
its ~500 lines of host tests the standard.
|
||
2. **Reuse the shell**: the filesystem harness — establishment, the
|
||
badge-scoped open-node table, the nine vfs-protocol handlers, mount
|
||
registration, the removal path — is shared code, not per-filesystem code.
|
||
*(Done: extracted from fat's original 434-line shell into
|
||
`library/kernel/file-system-harness.zig`; fat imports it and instantiates
|
||
`harness.Server(engine.FileSystem)`.)* An engine plus a `main` wiring it into the
|
||
harness is a complete filesystem service.
|
||
3. **Add the configuration row**: one line in `filesystems.csv` mapping the
|
||
on-disk signature to the binary. No other component changes: the volume
|
||
manager probes, matches, spawns; the vfs protocol is already
|
||
backend-neutral (a second provider serves it in production today — init's
|
||
synthetic registry backend).
|
||
4. **Prove removal**: extend the removable-media suite (below) for the new
|
||
filesystem — surprise-yank during writes must lose only what was
|
||
unflushed, said honestly, with the volume consistent enough to remount.
|
||
|
||
The engine never sees: partitions (it receives a volume-shaped block channel;
|
||
base offsets are the driver's clamp), device discovery (the channel arrives at
|
||
spawn), mount policy (prefixes are handed to it), or other volumes (one
|
||
process, one volume).
|
||
|
||
## Removable media: responsibilities on removal, per layer
|
||
|
||
The design rule, learned from what Linux cannot do: there is exactly **one**
|
||
surprise-removal path — kill the filesystem process, retire its mounts,
|
||
respawn on return. No half-alive states, no `remount-ro`, no mounts that
|
||
error forever (Plan 9's dead-server wart).
|
||
|
||
The path has **two triggers, one lifecycle**: the *device* leaving (the
|
||
storage driver dies — channel death, the table below), and the *medium*
|
||
leaving while the device stays (an SD card pulled from its reader, an ATAPI
|
||
tray opened — including USB card readers today). The second trigger is the
|
||
pushed `medium_changed` event on the block protocol — published today from a
|
||
TEST UNIT READY poll; still *planned* is the volume manager *consuming* it
|
||
(today removal is driven only by device-presence polling) and translating the
|
||
transport's native signal (SCSI UNIT ATTENTION, AHCI PxSSTS, NVMe
|
||
namespace-change AER) in place of the poll. On the event the volume manager
|
||
runs the same kill-retire path, then re-probes on medium return exactly as on
|
||
device return. Without it, a swapped card would be served with the previous
|
||
card's filesystem state.
|
||
|
||
| Layer | Observes | Must do | Guarantees |
|
||
|---|---|---|---|
|
||
| Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick |
|
||
| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
|
||
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
|
||
| Volume manager *(removal built; remount bench-pending)* | the storage device leaving the device-manager tree (poll) | kill the filesystem service of that device's volume; its kernel mounts retire | one removal path; mounts never dangle; log persistence stops *cleanly* |
|
||
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the unflushed write-back window is dropped on a surprise yank — danos writes no on-disk dirty/clean-shutdown marker today; the process never serves from behind a dead channel |
|
||
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated `fs_unmount` (built, V0) | resolution under a dead mount is `not_found`, not a hang |
|
||
| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
|
||
|
||
**On return** the same table runs upward in reverse: the bus re-enumerates and
|
||
re-reports; the device manager respawns the storage driver (built); the volume
|
||
manager re-probes — same content, same volume identity — respawns the
|
||
filesystem service, and remounts at the same prefix; applications see the
|
||
subtree reappear. A different stick in the same port is a *different volume*
|
||
(identity is content, not port) and mounts wherever its identity says —
|
||
possibly nowhere but `/volumes/…`.
|
||
|
||
### The boot volume: identity is what makes yanking it survivable
|
||
|
||
The boot volume is the sharpest instance of the return story, because the
|
||
system's own log persistence rides it (`/system/logs`), and because nothing
|
||
else about the running system depends on it at all — every binary and the
|
||
boot-time configuration live in the initial ramdisk. Its lifecycle,
|
||
end to end:
|
||
|
||
- **At first mount** the volume manager records the identity of the volume
|
||
carrying `/system/configuration` and `/system/logs` — THE boot volume,
|
||
from then on a fact about content, not about a port.
|
||
- **While it is absent**, the system runs on. The kernel log ring keeps
|
||
accumulating — it is the buffer that makes the absence survivable — and
|
||
the logger keeps draining it; only *persistence* pauses. File operations
|
||
under the retired mounts fail honestly (`not_found`). The ring is bounded,
|
||
so a long absence overwrites its oldest entries: that window is the data
|
||
loss, and it must be *said* — the logger marks the gap in the file when
|
||
persistence resumes, never splicing the stream silently.
|
||
- **On return — any port, any hub, even a different transport** — the
|
||
prober reads the same identity, the mount map answers with the same
|
||
prefixes, the filesystem service is respawned, and `/system/logs` is the
|
||
same tree it was: the logger resumes appending into the SAME
|
||
`<boot-stamp>` directory, per-binary files continuing where they left
|
||
off (plus the gap marker if the ring wrapped).
|
||
- **What never comes back** is the write-back window lost at the yank —
|
||
the dirty-honesty rule, unchanged; the boot volume gets no exemption.
|
||
|
||
This is also the requirement that shapes the logger: it treats the log tree
|
||
as a volume that comes and goes — failed flushes are retried on the same
|
||
patient cadence the fat service already uses for storage that arrives late,
|
||
never abandoned after the first `not_found`.
|
||
|
||
What is lost on a surprise yank is exactly the write-back window of the
|
||
filesystem service, no more: the engine owns its cache, so the blast radius of
|
||
a yank is one volume's unflushed writes, which are simply lost — danos records
|
||
no on-disk dirty/clean-shutdown bit yet. A
|
||
filesystem format with better crash honesty (journaling, copy-on-write — the
|
||
lesson of QNX's Power-Safe) narrows that window further and slots in as
|
||
implementation N+1 through the table above, changing nothing else.
|
||
|
||
## How the lifecycle is enforced
|
||
|
||
Responsibility tables are convention; convention is worthless (the bounds
|
||
audit's lesson). The volume lifecycle is enforced with the same three levers
|
||
the device lifecycle already uses, mapped one-to-one:
|
||
|
||
1. **A filesystem cannot acquire — it can only be given.** Filesystem
|
||
binaries hold NO establishment grants: no `open device-manager`, no
|
||
registry name to look up. The only block channel a filesystem process ever
|
||
has is the one handed to it at spawn by the volume manager. Serving the
|
||
wrong volume, a second volume, or a self-discovered volume is not
|
||
forbidden but *impossible* — the same way a driver cannot claim hardware
|
||
it was not delegated (device-authority.md).
|
||
2. **The volume manager supervises with teeth.** Mount-within-deadline or be
|
||
stopped (a filesystem wedged on a corrupt volume is killed and crash-loop
|
||
capped; the volume is marked bad, not retried forever). On removal the
|
||
manager KILLS the filesystem process and retires its mounts — the polite
|
||
observe-`EPEER`-and-exit path is an optimization; the kill is the
|
||
guarantee, because a process cannot be trusted to observe its own
|
||
obsolescence (the reap argument, proven on the device tree).
|
||
3. **The shared harness is how every filesystem inherits the lifecycle by
|
||
construction.** The harness — not the engine — owns the state machine:
|
||
establishment at spawn, mount registration, the flush-on-close hook (an
|
||
in-memory device-dirty check that commits the device write cache on close),
|
||
error-out-and-exit on channel death. The engine sits behind the
|
||
four-function vtable and never sees a channel; it cannot opt out of the
|
||
lifecycle for the same reason it cannot find a device. This is why the
|
||
harness is extracted BEFORE the second engine is written.
|
||
|
||
Beneath all three, two kernel mechanisms: ownership-gated `fs_unmount` (only
|
||
the mounting endpoint's holder may unmount), and the lazy dead-endpoint mount
|
||
sweep as the backstop nothing can disable. And above them, the check: a
|
||
**lifecycle conformance drill** — mount, serve, yank mid-write, verify honest
|
||
loss, replug, remount — parameterized over filesystem implementations, so
|
||
"danos supports filesystem X" MEANS "X passes the drill through the harness",
|
||
exactly as provider conformance means passing the reserved-verb suite.
|