docs: flip the storage-architecture status markers the V0-V4 track made real

The volume manager, per-sender range confinement + gate, medium_changed
event, per-volume spawning + supervision, ownership-gated fs_unmount, and
the removal half of the lifecycle are built. Left honestly pending: the
volumes.csv/filesystems.csv maps, the fuller identity ladder, multi-volume,
the VM consuming medium_changed (removal uses device-presence polling), and
remount-on-replug end-to-end (bench-pending — QEMU can't re-present the
boot-controller device). Full suite 127/127.
This commit is contained in:
Daniel Samson
2026-08-09 20:20:20 +01:00
parent e3ec9fa668
commit 5e89b111cf
@@ -1,13 +1,20 @@
# The storage architecture: layers, boundaries, responsibilities
> **Status:** the layered model below is the settled design
> ([storage-design-rationale.md](storage-design-rationale.md) records how it
> was reached and what the surveyed systems taught). The data path — vfs
> protocol, kernel mount routing, the FAT service, the block protocol,
> usb-storage — is **built**. The volume manager, the driver's range
> mechanism, per-volume filesystem spawning, and the removal lifecycle are
> **planned**; until they land, the FAT service performs volume-manager duties
> itself (marked below). This document is the reference for both states.
> ([storage-design-rationale.md](storage-design-rationale.md) records how it was
> reached, and [volume-manager-plan.md](../volume-manager-plan.md) how it was
> built). **Built** (the volume-manager track, V0–V4): the data path, the driver
> range confinement (per-sender clamp + the confinement gate), the `medium_changed`
> presence event, the volume manager itself — it probes the partition table,
> confines each filesystem to its partition, spawns one filesystem per volume, and
> supervises it — and the removal half of the lifecycle (a pulled stick unmounts).
> **Still pending**: the fuller identity ladder and the `volumes.csv` mount map,
> multi-volume (one FAT volume today), the volume manager *consuming*
> `medium_changed` (removal is detected by device-presence polling; the event is
> published but only a card-reader medium change needs the subscription), and the
> remount-on-replug end-to-end (the logic is in place; QEMU can't re-present the
> boot-controller device, so it is bench-verified). A few markers below are left
> where a duty is still pending.
## The model
@@ -47,7 +54,7 @@ Three kinds of boundary, deliberately different:
- **Control-plane relationships** sit beside the data path, never on it. Two
supervisors, one per layer: the **device manager** wires and revives the
device layers (bus and storage drivers — devices only); the **volume
manager** *(planned)* wires and revives the volume layer (filesystem
manager** *(built)* wires and revives the volume layer (filesystem
services). Neither touches steady-state I/O.
## Who does what
@@ -71,19 +78,21 @@ whole disk plus a polite base offset would let a buggy or compromised
filesystem scribble the neighboring partition — the same authority-overshoot
the device-authority track eliminated for MMIO and DMA.
**Volume manager** *(planned; today the FAT service squats on these duties)*:
the policy home of the volume layer, one service, supervised by init. It
subscribes to the device manager's child events; when a storage provider
appears it consumer-hellos for the block channel, reads the partition table
and the first blocks itself (**it** is the prober), consults its
configuration, defines sub-ranges on the driver, spawns the matching
filesystem service per volume with that volume's channel, supervises it, and
decides mount placement. Its tables are CSV configuration, read by it (the
policy), enforced by nobody else:
**Volume manager** *(built; `system/services/volume-manager`)*: the policy home
of the volume layer, one service, supervised by init. It watches the device
manager's tree for a storage provider; when one appears it consumer-hellos for
the block channel, reads the partition table and the first blocks itself
(**it** is the prober), defines the volume's sub-range on the driver, spawns the
matching filesystem service confined to that range, and supervises it (backoff,
crash-loop cap). *(Pending)*: it decides mount placement from `volumes.csv` and
picks the filesystem binary from `filesystems.csv` — today it hands every
FAT-shaped volume to the FAT service and the FAT service carries hardcoded mount
prefixes. Those tables are CSV configuration, read by it (the policy), enforced
by nobody else:
- `filesystems.csv` — content signature → filesystem binary. Adding a
filesystem adds a row.
- `volumes.csv` — the mount map, danos's fstab: **volume identity → mount
- `filesystems.csv` *(pending)* — content signature → filesystem binary. Adding
a filesystem adds a row.
- `volumes.csv` *(pending)* — the mount map, danos's fstab: **volume identity → mount
prefix**, keyed on content identity and never on port, path, or arrival
order (the lesson of Linux's `/dev/sda1`-era fstab, which broke on every
port move until `UUID=` replaced it). Identity is read off the medium by
@@ -166,9 +175,9 @@ be served with the previous card's filesystem state.
| Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick |
| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
| Volume manager *(planned)* | the storage provider's channel death / the manager's child events | kill each filesystem service of that device's volumes; retire their kernel mounts; remember the volume identity | one removal path; mounts never dangle; log persistence stops *cleanly* |
| Volume manager *(removal built; remount bench-pending)* | the storage device leaving the device-manager tree (poll) | kill the filesystem service of that device's volume; its kernel mounts retire | one removal path; mounts never dangle; log persistence stops *cleanly* |
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the FAT dirty flag on disk marks the unclean removal; the process never serves from behind a dead channel |
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); *(planned)* ownership-gated `fs_unmount` | resolution under a dead mount is `not_found`, not a hang |
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated `fs_unmount` (built, V0) | resolution under a dead mount is `not_found`, not a hang |
| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
**On return** the same table runs upward in reverse: the bus re-enumerates and