Files
Daniel Samson 4037e746aa docs: storage remount-on-replug is bench-pending, not bench-verified
The physical unplug/replug remount was never actually bench-run — the claim was
carried over from S3. QEMU cannot re-present a usb-storage device_add, so this
path is only reachable on real hardware. Correct both docs to say bench-pending,
keeping the honest distinction: the re-adopt+remount CODE PATH is QEMU-proven by
the driver-crash rebuild; only the physical-replug end-to-end awaits a bench pass.
2026-08-10 19:35:04 +01:00

301 lines
20 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# The storage architecture: layers, boundaries, responsibilities
> **Status:** the layered model below is the settled design
> ([storage-design-rationale.md](storage-design-rationale.md) records how it was
> reached, and [volume-manager-plan.md](../volume-manager-plan.md) how it was
> built). **Built** (V0–V4 + the storage-stack S1/S2): the data path, the driver
> range confinement (per-sender clamp + the confinement gate), the `medium_changed`
> presence event, the volume manager itself — it probes the partition table,
> confines each filesystem to its partition, spawns one filesystem per volume, and
> supervises it — the removal half of the lifecycle (a pulled stick unmounts), the
> identity ladder (GPT GUID + name, FAT serial + label, MBR), and the mount map:
> `filesystems.csv` (signature → binary) + `volumes.csv` (identity → optional
> override), a volume's mount path IS its content id (`/volumes/<id>`), with the
> label as display metadata a `volumes` query returns. Multi-volume is **built**:
> the manager adopts every storage device, probes each device's whole partition
> table, and spawns one range-confined FAT per volume — several volumes across
> several devices, or several partitions sharing one device's channel — each at
> its own `/volumes/<id>` path with its own supervision. The boot volume is
> identified by **content** (a volume backs `/system/configuration` + `/system/logs`
> only when it resolves `/system/configuration` on its own media), so it works as
> any partition of any device. exFAT is **built** as a second engine
> (`system/services/exfat`): full read + write, directories, rename, and on-disk
> up-case folding, reusing `library/kernel/file-system-harness` wholesale — the
> reuse claim, proven — and a volume routes to fat or exfat by its VBR, at an
> `exfat-<serial>` id-path. Removal is robust to all three triggers now: a
> pulled device (presence polling), a medium that leaves while its device stays
> (the volume manager CONSUMES `medium_changed`), and a storage driver that
> crashes while its device stays present (a channel-liveness `geometry()` probe
> reaps the volume and rebuilds it on the restarted driver's fresh channel). The
> re-adopt-and-remount path is QEMU-proven by the driver-crash rebuild; a physical
> unplug/replug exercises the same path but is bench-pending (QEMU cannot
> re-present a usb-storage `device_add`). **Still pending**: the `filesystem UUID`
> rung (ext-family superblocks, which need such an engine); and arbitration when
> two volumes both resolve the boot markers (S3 mounts both and logs each claim;
> picking one is deferred). A few
> markers below are left where a duty is still pending.
## The model
One pattern, applied twice: an application is a client of a service over a
protocol; a service is a client of a driver over a protocol.
```
application
│ vfs protocol (routed by the kernel mount table)
▼
filesystem service ── one process per volume
│ inside: [file-protocol shell] ── Zig API ── [fs engine]
│ block protocol (channel received at spawn)
▼
storage driver ── one process per device (usb-storage per stick)
│ usb-transfer protocol (channel received via hello, by lineage)
▼
bus driver ── one process per controller (usb-xhci-bus)
│ hardware
```
Three kinds of boundary, deliberately different:
- **Protocol boundaries** sit where processes meet (vfs, block, usb-transfer).
Each side is independently restartable; channels are established by
capability handoff, never by registry names (communication.md
"Establishment: two planes"). Steady-state file I/O crosses exactly two —
vfs and block.
- **Library boundaries** sit where concerns meet inside one process. The
filesystem service's engine (on-disk format logic, host-testable, behind
the four-function `BlockDevice` vtable) and shell (IPC serving, open-node
table, mount registration) talk through a Zig API. No protocol between
them: they share fate regardless — a corrupt engine corrupts the answers
either way — so a channel there would add a hop per file operation and buy
nothing. The seam exists for compile-time testability and reuse; the
*process* is the restart unit.
- **Control-plane relationships** sit beside the data path, never on it. Two
supervisors, one per layer: the **device manager** wires and revives the
device layers (bus and storage drivers — devices only); the **volume
manager** *(built)* wires and revives the volume layer (filesystem
services). Neither touches steady-state I/O.
## Who does what
**Bus driver** (usb-xhci-bus, per controller): hardware only. Enumerates,
reports children to the device manager, serves the transfer contract to its
own children's class drivers, routed by lineage.
**Storage driver** (usb-storage, per device; later nvme, ahci): hardware →
blocks. Speaks its transport upward from its device; serves the block
protocol (geometry, read, write, flush, attach, detach — and the pushed
`medium_changed` presence event, published today from a slow TEST UNIT READY
poll; *planned*: translating it from the transport's native signal instead of
polling). **Content-blind, permanently**:
it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no
GPT, no filesystem magic, ever. It has one content-blind mechanism *(built)*:
**named sub-ranges** — "serve blocks [a, b) as a channel of the
same block contract". It clamps and offsets; it never knows the numbers came
from a partition table. The clamp lives here and nowhere else because a
channel must carry exactly the authority it grants: handing a filesystem the
whole disk plus a polite base offset would let a buggy or compromised
filesystem scribble the neighboring partition — the same authority-overshoot
the device-authority track eliminated for MMIO and DMA.
**Volume manager** *(built; `system/services/volume-manager`)*: the policy home
of the volume layer, one service, supervised by init. It watches the device
manager's tree for a storage provider; when one appears it consumer-hellos for
the block channel, reads the partition table and the first blocks itself
(**it** is the prober), defines the volume's sub-range on the driver, spawns the
matching filesystem service confined to that range, and supervises it (backoff,
crash-loop cap). *(Built)*: it picks the filesystem binary from
`filesystems.csv` by the volume's content signature, and mounts the volume at its
content id (`/volumes/<id>`) — or a `volumes.csv` override. The label is display
metadata the `volumes` query returns, never the path. Those tables are CSV
configuration, read by it (the policy), enforced by nobody else:
- `filesystems.csv` *(built)* — content signature → filesystem binary. Adding
a filesystem adds a row.
- `volumes.csv` *(built)* — the mount map, danos's fstab: an OPTIONAL **volume
identity → mount prefix** override (a volume with no row mounts at its default
`/volumes/<id>`), keyed on content identity and never on port, path, or arrival
order (the lesson of Linux's `/dev/sda1`-era fstab, which broke on every
port move until `UUID=` replaced it). Identity is read off the medium by
the prober, strongest first: GPT partition GUID → filesystem UUID → FAT
serial + label → MBR signature + partition index → anonymous (generated
name, no persistence). Same identity, same mount point: a drive moved to
another port — USB to another hub, SATA to another bay, even a stick
returning in a different dock — lands exactly where it was. Duplicate
identity (cloned sticks, together) is policy: first keeps the name, the
second mounts suffixed and is logged loudly. The boot volume is the
recorded identity of the volume carrying `/system/configuration` and
`/system/logs`, findable on any port. Every volume's default mount is
`/volumes/<id>` — its rendered content identity.
**Filesystem service** (the FAT service today; one process per volume): the
proven unit — block-client + engine + file-protocol provider in one binary. It
receives its block channel at spawn; it never discovers devices. It registers
its own mounts with the kernel; its write cache lives inside the process, so a
write error is observed by the code that owns the volume and surfaces on the
owning channel (the anti-fsyncgate rule — never a system-wide dirty pool).
*(Built:)* fat receives its mount path as `argv[2]` from the volume manager (the
volume's id-path, e.g. `/volumes/fat-12345678`) and mounts its root there. It
installs the two `/system` hierarchy rewrites (`/system/configuration`,
`/system/logs`) only when it is the boot volume — decided by **content**: it
resolves `/system/configuration` on its own media at mount, so a data volume
mounts at its id-path alone and never shadows the running system. It no longer
self-acquires a volume — the V3b flip made it receive its volume id and block
channel from the volume manager, consistent with "it never discovers devices"
above. Because several volumes now serve at once, no filesystem binds a shared
service name; clients reach each through the kernel mount table (`fs_resolve`
routes by prefix to the backing endpoint).
**Kernel** (mechanism only): the mount table routes paths to backend
endpoints — resolve and redirect, never data. Remount-replace is the restart
story; a dead backend's slot is swept lazily on the next resolution.
`fs_unmount` is ownership-gated *(built, V0)* — only the mounting task may
unmount its own prefix; anyone else is refused (`EPERM`). No dead-owner
exception: a dead owner's slot is swept lazily by resolution, and restart goes
through remount-replace, never through a stranger's unmount.
## Adding a filesystem
1. **Write the engine**: pure Zig, no IPC, behind the `BlockDevice` vtable
(four functions). On-disk layout in `align(1)` extern structs. Host-test it
with `zig build test` fixtures — the FAT engine
([engine.zig](../../system/services/fat/engine.zig),
[on-disk.zig](../../system/services/fat/on-disk.zig)) is the template, and
its ~500 lines of host tests the standard.
2. **Reuse the shell**: the filesystem harness — establishment, the
badge-scoped open-node table, the nine vfs-protocol handlers, mount
registration, the removal path — is shared code, not per-filesystem code.
*(Done: extracted from fat's original 434-line shell into
`library/kernel/file-system-harness.zig`; fat imports it and instantiates
`harness.Server(engine.FileSystem)`.)* An engine plus a `main` wiring it into the
harness is a complete filesystem service.
3. **Add the configuration row**: one line in `filesystems.csv` mapping the
on-disk signature to the binary. No other component changes: the volume
manager probes, matches, spawns; the vfs protocol is already
backend-neutral (a second provider serves it in production today — init's
synthetic registry backend).
4. **Prove removal**: extend the removable-media suite (below) for the new
filesystem — surprise-yank during writes must lose only what was
unflushed, said honestly, with the volume consistent enough to remount.
The engine never sees: partitions (it receives a volume-shaped block channel;
base offsets are the driver's clamp), device discovery (the channel arrives at
spawn), mount policy (prefixes are handed to it), or other volumes (one
process, one volume).
## Removable media: responsibilities on removal, per layer
The design rule, learned from what Linux cannot do: there is exactly **one**
surprise-removal path — kill the filesystem process, retire its mounts,
respawn on return. No half-alive states, no `remount-ro`, no mounts that
error forever (Plan 9's dead-server wart).
The path folds **three triggers into one lifecycle**: the *device* leaving (a
pulled stick — presence polling); the *medium* leaving while the device stays
(an SD card pulled from its reader, an ATAPI tray opened, a USB card reader);
and a storage *driver crashing* while its device stays in the tree. The second
trigger is the pushed `medium_changed` event on the block protocol, published
from a TEST UNIT READY poll — the volume manager now **consumes** it (subscribed
per device), running the same kill-retire path and re-probing on medium return,
so a swapped card is never served with the previous card's filesystem state. The
third is caught by a channel-liveness `geometry()` probe: presence polling alone
sees the device still present, but the channel is dead, so the manager reaps the
volume and rebuilds it on the restarted driver's fresh channel. Still *planned*
is translating the transport's native signal (SCSI UNIT ATTENTION, AHCI PxSSTS,
NVMe namespace-change AER) in place of the presence poll.
| Layer | Observes | Must do | Guarantees |
|---|---|---|---|
| Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick |
| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
| Volume manager *(built)* | a device leaving the tree (poll), a `medium_changed` event, or a dead channel under a still-present device (a crashed driver — `geometry()` liveness probe) | kill that volume's filesystem service (its mounts retire), then re-adopt + remount on return or on the restarted driver's fresh channel | one removal path for all three triggers; mounts never dangle; the manager never serves from behind a dead channel |
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the unflushed write-back window is dropped on a surprise yank — danos writes no on-disk dirty/clean-shutdown marker today; the process never serves from behind a dead channel |
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated `fs_unmount` (built, V0) | resolution under a dead mount is `not_found`, not a hang |
| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
**On return** the same table runs upward in reverse: the bus re-enumerates and
re-reports; the device manager respawns the storage driver (built); the volume
manager re-probes — same content, same volume identity — respawns the
filesystem service, and remounts at the same prefix; applications see the
subtree reappear. A different stick in the same port is a *different volume*
(identity is content, not port) and mounts wherever its identity says —
possibly nowhere but `/volumes/…`.
### The boot volume: identity is what makes yanking it survivable
The boot volume is the sharpest instance of the return story, because the
system's own log persistence rides it (`/system/logs`), and because nothing
else about the running system depends on it at all — every binary and the
boot-time configuration live in the initial ramdisk. Its lifecycle,
end to end:
- **At first mount** the volume manager records the identity of the volume
carrying `/system/configuration` and `/system/logs` — THE boot volume,
from then on a fact about content, not about a port.
- **While it is absent**, the system runs on. The kernel log ring keeps
accumulating — it is the buffer that makes the absence survivable — and
the logger keeps draining it; only *persistence* pauses. File operations
under the retired mounts fail honestly (`not_found`). The ring is bounded,
so a long absence overwrites its oldest entries: that window is the data
loss, and it must be *said* — the logger marks the gap in the file when
persistence resumes, never splicing the stream silently.
- **On return — any port, any hub, even a different transport** — the
prober reads the same identity, the mount map answers with the same
prefixes, the filesystem service is respawned, and `/system/logs` is the
same tree it was: the logger resumes appending into the SAME
`<boot-stamp>` directory, per-binary files continuing where they left
off (plus the gap marker if the ring wrapped).
- **What never comes back** is the write-back window lost at the yank —
the dirty-honesty rule, unchanged; the boot volume gets no exemption.
This is also the requirement that shapes the logger: it treats the log tree
as a volume that comes and goes — failed flushes are retried on the same
patient cadence the fat service already uses for storage that arrives late,
never abandoned after the first `not_found`.
What is lost on a surprise yank is exactly the write-back window of the
filesystem service, no more: the engine owns its cache, so the blast radius of
a yank is one volume's unflushed writes, which are simply lost — danos records
no on-disk dirty/clean-shutdown bit yet. A
filesystem format with better crash honesty (journaling, copy-on-write — the
lesson of QNX's Power-Safe) narrows that window further and slots in as
implementation N+1 through the table above, changing nothing else.
## How the lifecycle is enforced
Responsibility tables are convention; convention is worthless (the bounds
audit's lesson). The volume lifecycle is enforced with the same three levers
the device lifecycle already uses, mapped one-to-one:
1. **A filesystem cannot acquire — it can only be given.** Filesystem
binaries hold NO establishment grants: no `open device-manager`, no
registry name to look up. The only block channel a filesystem process ever
has is the one handed to it at spawn by the volume manager. Serving the
wrong volume, a second volume, or a self-discovered volume is not
forbidden but *impossible* — the same way a driver cannot claim hardware
it was not delegated (device-authority.md).
2. **The volume manager supervises with teeth.** Mount-within-deadline or be
stopped (a filesystem wedged on a corrupt volume is killed and crash-loop
capped; the volume is marked bad, not retried forever). On removal the
manager KILLS the filesystem process and retires its mounts — the polite
observe-`EPEER`-and-exit path is an optimization; the kill is the
guarantee, because a process cannot be trusted to observe its own
obsolescence (the reap argument, proven on the device tree).
3. **The shared harness is how every filesystem inherits the lifecycle by
construction.** The harness — not the engine — owns the state machine:
establishment at spawn, mount registration, the flush-on-close hook (an
in-memory device-dirty check that commits the device write cache on close),
error-out-and-exit on channel death. The engine sits behind the
four-function vtable and never sees a channel; it cannot opt out of the
lifecycle for the same reason it cannot find a device. This is why the
harness is extracted BEFORE the second engine is written.
Beneath all three, two kernel mechanisms: ownership-gated `fs_unmount` (only
the mounting endpoint's holder may unmount), and the lazy dead-endpoint mount
sweep as the backstop nothing can disable. And above them, the check: a
**lifecycle conformance drill** — mount, serve, yank mid-write, verify honest
loss, replug, remount — parameterized over filesystem implementations, so
"danos supports filesystem X" MEANS "X passes the drill through the harness",
exactly as provider conformance means passing the reserved-verb suite.