Flip the storage docs to record exFAT as a built second engine reusing the shared filesystem harness wholesale (the reuse the architecture promised): full read + write, directories, rename, on-disk up-case folding, routed by VBR to fat or exfat at an exfat-<serial> id-path. Note the two surface limits shared by BOTH engines as vfs-layer concerns, not exFAT shortcuts: u32 file offsets (a 4 GiB cap) and ASCII-only names.
297 lines
19 KiB
Markdown
297 lines
19 KiB
Markdown
# The storage architecture: layers, boundaries, responsibilities
|
||
|
||
> **Status:** the layered model below is the settled design
|
||
> ([storage-design-rationale.md](storage-design-rationale.md) records how it was
|
||
> reached, and [volume-manager-plan.md](../volume-manager-plan.md) how it was
|
||
> built). **Built** (V0–V4 + the storage-stack S1/S2): the data path, the driver
|
||
> range confinement (per-sender clamp + the confinement gate), the `medium_changed`
|
||
> presence event, the volume manager itself — it probes the partition table,
|
||
> confines each filesystem to its partition, spawns one filesystem per volume, and
|
||
> supervises it — the removal half of the lifecycle (a pulled stick unmounts), the
|
||
> identity ladder (GPT GUID + name, FAT serial + label, MBR), and the mount map:
|
||
> `filesystems.csv` (signature → binary) + `volumes.csv` (identity → optional
|
||
> override), a volume's mount path IS its content id (`/volumes/<id>`), with the
|
||
> label as display metadata a `volumes` query returns. Multi-volume is **built**:
|
||
> the manager adopts every storage device, probes each device's whole partition
|
||
> table, and spawns one range-confined FAT per volume — several volumes across
|
||
> several devices, or several partitions sharing one device's channel — each at
|
||
> its own `/volumes/<id>` path with its own supervision. The boot volume is
|
||
> identified by **content** (a volume backs `/system/configuration` + `/system/logs`
|
||
> only when it resolves `/system/configuration` on its own media), so it works as
|
||
> any partition of any device. exFAT is **built** as a second engine
|
||
> (`system/services/exfat`): full read + write, directories, rename, and on-disk
|
||
> up-case folding, reusing `library/kernel/file-system-harness` wholesale — the
|
||
> reuse claim, proven — and a volume routes to fat or exfat by its VBR, at an
|
||
> `exfat-<serial>` id-path. **Still pending**: the `filesystem UUID` rung (ext-
|
||
> family superblocks, which need such an engine); the volume manager *consuming*
|
||
> `medium_changed`
|
||
> (removal is detected by device-presence polling; the event is published but only
|
||
> a card-reader medium change needs the subscription); the remount-on-replug
|
||
> end-to-end (the logic is in place; QEMU can't re-present the boot-controller
|
||
> device, so it is bench-verified); and arbitration when two volumes both resolve
|
||
> the boot markers (S3 mounts both and logs each claim; picking one is S4). A few
|
||
> markers below are left where a duty is still pending.
|
||
|
||
## The model
|
||
|
||
One pattern, applied twice: an application is a client of a service over a
|
||
protocol; a service is a client of a driver over a protocol.
|
||
|
||
```
|
||
application
|
||
│ vfs protocol (routed by the kernel mount table)
|
||
▼
|
||
filesystem service ── one process per volume
|
||
│ inside: [file-protocol shell] ── Zig API ── [fs engine]
|
||
│ block protocol (channel received at spawn)
|
||
▼
|
||
storage driver ── one process per device (usb-storage per stick)
|
||
│ usb-transfer protocol (channel received via hello, by lineage)
|
||
▼
|
||
bus driver ── one process per controller (usb-xhci-bus)
|
||
│ hardware
|
||
```
|
||
|
||
Three kinds of boundary, deliberately different:
|
||
|
||
- **Protocol boundaries** sit where processes meet (vfs, block, usb-transfer).
|
||
Each side is independently restartable; channels are established by
|
||
capability handoff, never by registry names (communication.md
|
||
"Establishment: two planes"). Steady-state file I/O crosses exactly two —
|
||
vfs and block.
|
||
- **Library boundaries** sit where concerns meet inside one process. The
|
||
filesystem service's engine (on-disk format logic, host-testable, behind
|
||
the four-function `BlockDevice` vtable) and shell (IPC serving, open-node
|
||
table, mount registration) talk through a Zig API. No protocol between
|
||
them: they share fate regardless — a corrupt engine corrupts the answers
|
||
either way — so a channel there would add a hop per file operation and buy
|
||
nothing. The seam exists for compile-time testability and reuse; the
|
||
*process* is the restart unit.
|
||
- **Control-plane relationships** sit beside the data path, never on it. Two
|
||
supervisors, one per layer: the **device manager** wires and revives the
|
||
device layers (bus and storage drivers — devices only); the **volume
|
||
manager** *(built)* wires and revives the volume layer (filesystem
|
||
services). Neither touches steady-state I/O.
|
||
|
||
## Who does what
|
||
|
||
**Bus driver** (usb-xhci-bus, per controller): hardware only. Enumerates,
|
||
reports children to the device manager, serves the transfer contract to its
|
||
own children's class drivers, routed by lineage.
|
||
|
||
**Storage driver** (usb-storage, per device; later nvme, ahci): hardware →
|
||
blocks. Speaks its transport upward from its device; serves the block
|
||
protocol (geometry, read, write, flush, attach, detach — and the pushed
|
||
`medium_changed` presence event, published today from a slow TEST UNIT READY
|
||
poll; *planned*: translating it from the transport's native signal instead of
|
||
polling). **Content-blind, permanently**:
|
||
it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no
|
||
GPT, no filesystem magic, ever. It has one content-blind mechanism *(built)*:
|
||
**named sub-ranges** — "serve blocks [a, b) as a channel of the
|
||
same block contract". It clamps and offsets; it never knows the numbers came
|
||
from a partition table. The clamp lives here and nowhere else because a
|
||
channel must carry exactly the authority it grants: handing a filesystem the
|
||
whole disk plus a polite base offset would let a buggy or compromised
|
||
filesystem scribble the neighboring partition — the same authority-overshoot
|
||
the device-authority track eliminated for MMIO and DMA.
|
||
|
||
**Volume manager** *(built; `system/services/volume-manager`)*: the policy home
|
||
of the volume layer, one service, supervised by init. It watches the device
|
||
manager's tree for a storage provider; when one appears it consumer-hellos for
|
||
the block channel, reads the partition table and the first blocks itself
|
||
(**it** is the prober), defines the volume's sub-range on the driver, spawns the
|
||
matching filesystem service confined to that range, and supervises it (backoff,
|
||
crash-loop cap). *(Built)*: it picks the filesystem binary from
|
||
`filesystems.csv` by the volume's content signature, and mounts the volume at its
|
||
content id (`/volumes/<id>`) — or a `volumes.csv` override. The label is display
|
||
metadata the `volumes` query returns, never the path. Those tables are CSV
|
||
configuration, read by it (the policy), enforced by nobody else:
|
||
|
||
- `filesystems.csv` *(built)* — content signature → filesystem binary. Adding
|
||
a filesystem adds a row.
|
||
- `volumes.csv` *(built)* — the mount map, danos's fstab: an OPTIONAL **volume
|
||
identity → mount prefix** override (a volume with no row mounts at its default
|
||
`/volumes/<id>`), keyed on content identity and never on port, path, or arrival
|
||
order (the lesson of Linux's `/dev/sda1`-era fstab, which broke on every
|
||
port move until `UUID=` replaced it). Identity is read off the medium by
|
||
the prober, strongest first: GPT partition GUID → filesystem UUID → FAT
|
||
serial + label → MBR signature + partition index → anonymous (generated
|
||
name, no persistence). Same identity, same mount point: a drive moved to
|
||
another port — USB to another hub, SATA to another bay, even a stick
|
||
returning in a different dock — lands exactly where it was. Duplicate
|
||
identity (cloned sticks, together) is policy: first keeps the name, the
|
||
second mounts suffixed and is logged loudly. The boot volume is the
|
||
recorded identity of the volume carrying `/system/configuration` and
|
||
`/system/logs`, findable on any port. Every volume's default mount is
|
||
`/volumes/<id>` — its rendered content identity.
|
||
|
||
**Filesystem service** (the FAT service today; one process per volume): the
|
||
proven unit — block-client + engine + file-protocol provider in one binary. It
|
||
receives its block channel at spawn; it never discovers devices. It registers
|
||
its own mounts with the kernel; its write cache lives inside the process, so a
|
||
write error is observed by the code that owns the volume and surfaces on the
|
||
owning channel (the anti-fsyncgate rule — never a system-wide dirty pool).
|
||
*(Built:)* fat receives its mount path as `argv[2]` from the volume manager (the
|
||
volume's id-path, e.g. `/volumes/fat-12345678`) and mounts its root there. It
|
||
installs the two `/system` hierarchy rewrites (`/system/configuration`,
|
||
`/system/logs`) only when it is the boot volume — decided by **content**: it
|
||
resolves `/system/configuration` on its own media at mount, so a data volume
|
||
mounts at its id-path alone and never shadows the running system. It no longer
|
||
self-acquires a volume — the V3b flip made it receive its volume id and block
|
||
channel from the volume manager, consistent with "it never discovers devices"
|
||
above. Because several volumes now serve at once, no filesystem binds a shared
|
||
service name; clients reach each through the kernel mount table (`fs_resolve`
|
||
routes by prefix to the backing endpoint).
|
||
|
||
**Kernel** (mechanism only): the mount table routes paths to backend
|
||
endpoints — resolve and redirect, never data. Remount-replace is the restart
|
||
story; a dead backend's slot is swept lazily on the next resolution.
|
||
`fs_unmount` is ownership-gated *(built, V0)* — only the mounting task may
|
||
unmount its own prefix; anyone else is refused (`EPERM`). No dead-owner
|
||
exception: a dead owner's slot is swept lazily by resolution, and restart goes
|
||
through remount-replace, never through a stranger's unmount.
|
||
|
||
## Adding a filesystem
|
||
|
||
1. **Write the engine**: pure Zig, no IPC, behind the `BlockDevice` vtable
|
||
(four functions). On-disk layout in `align(1)` extern structs. Host-test it
|
||
with `zig build test` fixtures — the FAT engine
|
||
([engine.zig](../../system/services/fat/engine.zig),
|
||
[on-disk.zig](../../system/services/fat/on-disk.zig)) is the template, and
|
||
its ~500 lines of host tests the standard.
|
||
2. **Reuse the shell**: the filesystem harness — establishment, the
|
||
badge-scoped open-node table, the nine vfs-protocol handlers, mount
|
||
registration, the removal path — is shared code, not per-filesystem code.
|
||
*(Done: extracted from fat's original 434-line shell into
|
||
`library/kernel/file-system-harness.zig`; fat imports it and instantiates
|
||
`harness.Server(engine.FileSystem)`.)* An engine plus a `main` wiring it into the
|
||
harness is a complete filesystem service.
|
||
3. **Add the configuration row**: one line in `filesystems.csv` mapping the
|
||
on-disk signature to the binary. No other component changes: the volume
|
||
manager probes, matches, spawns; the vfs protocol is already
|
||
backend-neutral (a second provider serves it in production today — init's
|
||
synthetic registry backend).
|
||
4. **Prove removal**: extend the removable-media suite (below) for the new
|
||
filesystem — surprise-yank during writes must lose only what was
|
||
unflushed, said honestly, with the volume consistent enough to remount.
|
||
|
||
The engine never sees: partitions (it receives a volume-shaped block channel;
|
||
base offsets are the driver's clamp), device discovery (the channel arrives at
|
||
spawn), mount policy (prefixes are handed to it), or other volumes (one
|
||
process, one volume).
|
||
|
||
## Removable media: responsibilities on removal, per layer
|
||
|
||
The design rule, learned from what Linux cannot do: there is exactly **one**
|
||
surprise-removal path — kill the filesystem process, retire its mounts,
|
||
respawn on return. No half-alive states, no `remount-ro`, no mounts that
|
||
error forever (Plan 9's dead-server wart).
|
||
|
||
The path has **two triggers, one lifecycle**: the *device* leaving (the
|
||
storage driver dies — channel death, the table below), and the *medium*
|
||
leaving while the device stays (an SD card pulled from its reader, an ATAPI
|
||
tray opened — including USB card readers today). The second trigger is the
|
||
pushed `medium_changed` event on the block protocol — published today from a
|
||
TEST UNIT READY poll; still *planned* is the volume manager *consuming* it
|
||
(today removal is driven only by device-presence polling) and translating the
|
||
transport's native signal (SCSI UNIT ATTENTION, AHCI PxSSTS, NVMe
|
||
namespace-change AER) in place of the poll. On the event the volume manager
|
||
runs the same kill-retire path, then re-probes on medium return exactly as on
|
||
device return. Without it, a swapped card would be served with the previous
|
||
card's filesystem state.
|
||
|
||
| Layer | Observes | Must do | Guarantees |
|
||
|---|---|---|---|
|
||
| Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick |
|
||
| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
|
||
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
|
||
| Volume manager *(removal built; remount bench-pending)* | the storage device leaving the device-manager tree (poll) | kill the filesystem service of that device's volume; its kernel mounts retire | one removal path; mounts never dangle; log persistence stops *cleanly* |
|
||
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the unflushed write-back window is dropped on a surprise yank — danos writes no on-disk dirty/clean-shutdown marker today; the process never serves from behind a dead channel |
|
||
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated `fs_unmount` (built, V0) | resolution under a dead mount is `not_found`, not a hang |
|
||
| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
|
||
|
||
**On return** the same table runs upward in reverse: the bus re-enumerates and
|
||
re-reports; the device manager respawns the storage driver (built); the volume
|
||
manager re-probes — same content, same volume identity — respawns the
|
||
filesystem service, and remounts at the same prefix; applications see the
|
||
subtree reappear. A different stick in the same port is a *different volume*
|
||
(identity is content, not port) and mounts wherever its identity says —
|
||
possibly nowhere but `/volumes/…`.
|
||
|
||
### The boot volume: identity is what makes yanking it survivable
|
||
|
||
The boot volume is the sharpest instance of the return story, because the
|
||
system's own log persistence rides it (`/system/logs`), and because nothing
|
||
else about the running system depends on it at all — every binary and the
|
||
boot-time configuration live in the initial ramdisk. Its lifecycle,
|
||
end to end:
|
||
|
||
- **At first mount** the volume manager records the identity of the volume
|
||
carrying `/system/configuration` and `/system/logs` — THE boot volume,
|
||
from then on a fact about content, not about a port.
|
||
- **While it is absent**, the system runs on. The kernel log ring keeps
|
||
accumulating — it is the buffer that makes the absence survivable — and
|
||
the logger keeps draining it; only *persistence* pauses. File operations
|
||
under the retired mounts fail honestly (`not_found`). The ring is bounded,
|
||
so a long absence overwrites its oldest entries: that window is the data
|
||
loss, and it must be *said* — the logger marks the gap in the file when
|
||
persistence resumes, never splicing the stream silently.
|
||
- **On return — any port, any hub, even a different transport** — the
|
||
prober reads the same identity, the mount map answers with the same
|
||
prefixes, the filesystem service is respawned, and `/system/logs` is the
|
||
same tree it was: the logger resumes appending into the SAME
|
||
`<boot-stamp>` directory, per-binary files continuing where they left
|
||
off (plus the gap marker if the ring wrapped).
|
||
- **What never comes back** is the write-back window lost at the yank —
|
||
the dirty-honesty rule, unchanged; the boot volume gets no exemption.
|
||
|
||
This is also the requirement that shapes the logger: it treats the log tree
|
||
as a volume that comes and goes — failed flushes are retried on the same
|
||
patient cadence the fat service already uses for storage that arrives late,
|
||
never abandoned after the first `not_found`.
|
||
|
||
What is lost on a surprise yank is exactly the write-back window of the
|
||
filesystem service, no more: the engine owns its cache, so the blast radius of
|
||
a yank is one volume's unflushed writes, which are simply lost — danos records
|
||
no on-disk dirty/clean-shutdown bit yet. A
|
||
filesystem format with better crash honesty (journaling, copy-on-write — the
|
||
lesson of QNX's Power-Safe) narrows that window further and slots in as
|
||
implementation N+1 through the table above, changing nothing else.
|
||
|
||
## How the lifecycle is enforced
|
||
|
||
Responsibility tables are convention; convention is worthless (the bounds
|
||
audit's lesson). The volume lifecycle is enforced with the same three levers
|
||
the device lifecycle already uses, mapped one-to-one:
|
||
|
||
1. **A filesystem cannot acquire — it can only be given.** Filesystem
|
||
binaries hold NO establishment grants: no `open device-manager`, no
|
||
registry name to look up. The only block channel a filesystem process ever
|
||
has is the one handed to it at spawn by the volume manager. Serving the
|
||
wrong volume, a second volume, or a self-discovered volume is not
|
||
forbidden but *impossible* — the same way a driver cannot claim hardware
|
||
it was not delegated (device-authority.md).
|
||
2. **The volume manager supervises with teeth.** Mount-within-deadline or be
|
||
stopped (a filesystem wedged on a corrupt volume is killed and crash-loop
|
||
capped; the volume is marked bad, not retried forever). On removal the
|
||
manager KILLS the filesystem process and retires its mounts — the polite
|
||
observe-`EPEER`-and-exit path is an optimization; the kill is the
|
||
guarantee, because a process cannot be trusted to observe its own
|
||
obsolescence (the reap argument, proven on the device tree).
|
||
3. **The shared harness is how every filesystem inherits the lifecycle by
|
||
construction.** The harness — not the engine — owns the state machine:
|
||
establishment at spawn, mount registration, the flush-on-close hook (an
|
||
in-memory device-dirty check that commits the device write cache on close),
|
||
error-out-and-exit on channel death. The engine sits behind the
|
||
four-function vtable and never sees a channel; it cannot opt out of the
|
||
lifecycle for the same reason it cannot find a device. This is why the
|
||
harness is extracted BEFORE the second engine is written.
|
||
|
||
Beneath all three, two kernel mechanisms: ownership-gated `fs_unmount` (only
|
||
the mounting endpoint's holder may unmount), and the lazy dead-endpoint mount
|
||
sweep as the backstop nothing can disable. And above them, the check: a
|
||
**lifecycle conformance drill** — mount, serve, yank mid-write, verify honest
|
||
loss, replug, remount — parameterized over filesystem implementations, so
|
||
"danos supports filesystem X" MEANS "X passes the drill through the harness",
|
||
exactly as provider conformance means passing the reserved-verb suite.
|