docs: flip the storage-architecture status markers the V0-V4 track made real
The volume manager, per-sender range confinement + gate, medium_changed event, per-volume spawning + supervision, ownership-gated fs_unmount, and the removal half of the lifecycle are built. Left honestly pending: the volumes.csv/filesystems.csv maps, the fuller identity ladder, multi-volume, the VM consuming medium_changed (removal uses device-presence polling), and remount-on-replug end-to-end (bench-pending — QEMU can't re-present the boot-controller device). Full suite 127/127.
This commit is contained in:
@@ -1,13 +1,20 @@
|
||||
# The storage architecture: layers, boundaries, responsibilities
|
||||
|
||||
> **Status:** the layered model below is the settled design
|
||||
> ([storage-design-rationale.md](storage-design-rationale.md) records how it
|
||||
> was reached and what the surveyed systems taught). The data path — vfs
|
||||
> protocol, kernel mount routing, the FAT service, the block protocol,
|
||||
> usb-storage — is **built**. The volume manager, the driver's range
|
||||
> mechanism, per-volume filesystem spawning, and the removal lifecycle are
|
||||
> **planned**; until they land, the FAT service performs volume-manager duties
|
||||
> itself (marked below). This document is the reference for both states.
|
||||
> ([storage-design-rationale.md](storage-design-rationale.md) records how it was
|
||||
> reached, and [volume-manager-plan.md](../volume-manager-plan.md) how it was
|
||||
> built). **Built** (the volume-manager track, V0–V4): the data path, the driver
|
||||
> range confinement (per-sender clamp + the confinement gate), the `medium_changed`
|
||||
> presence event, the volume manager itself — it probes the partition table,
|
||||
> confines each filesystem to its partition, spawns one filesystem per volume, and
|
||||
> supervises it — and the removal half of the lifecycle (a pulled stick unmounts).
|
||||
> **Still pending**: the fuller identity ladder and the `volumes.csv` mount map,
|
||||
> multi-volume (one FAT volume today), the volume manager *consuming*
|
||||
> `medium_changed` (removal is detected by device-presence polling; the event is
|
||||
> published but only a card-reader medium change needs the subscription), and the
|
||||
> remount-on-replug end-to-end (the logic is in place; QEMU can't re-present the
|
||||
> boot-controller device, so it is bench-verified). A few markers below are left
|
||||
> where a duty is still pending.
|
||||
|
||||
## The model
|
||||
|
||||
@@ -47,7 +54,7 @@ Three kinds of boundary, deliberately different:
|
||||
- **Control-plane relationships** sit beside the data path, never on it. Two
|
||||
supervisors, one per layer: the **device manager** wires and revives the
|
||||
device layers (bus and storage drivers — devices only); the **volume
|
||||
manager** *(planned)* wires and revives the volume layer (filesystem
|
||||
manager** *(built)* wires and revives the volume layer (filesystem
|
||||
services). Neither touches steady-state I/O.
|
||||
|
||||
## Who does what
|
||||
@@ -71,19 +78,21 @@ whole disk plus a polite base offset would let a buggy or compromised
|
||||
filesystem scribble the neighboring partition — the same authority-overshoot
|
||||
the device-authority track eliminated for MMIO and DMA.
|
||||
|
||||
**Volume manager** *(planned; today the FAT service squats on these duties)*:
|
||||
the policy home of the volume layer, one service, supervised by init. It
|
||||
subscribes to the device manager's child events; when a storage provider
|
||||
appears it consumer-hellos for the block channel, reads the partition table
|
||||
and the first blocks itself (**it** is the prober), consults its
|
||||
configuration, defines sub-ranges on the driver, spawns the matching
|
||||
filesystem service per volume with that volume's channel, supervises it, and
|
||||
decides mount placement. Its tables are CSV configuration, read by it (the
|
||||
policy), enforced by nobody else:
|
||||
**Volume manager** *(built; `system/services/volume-manager`)*: the policy home
|
||||
of the volume layer, one service, supervised by init. It watches the device
|
||||
manager's tree for a storage provider; when one appears it consumer-hellos for
|
||||
the block channel, reads the partition table and the first blocks itself
|
||||
(**it** is the prober), defines the volume's sub-range on the driver, spawns the
|
||||
matching filesystem service confined to that range, and supervises it (backoff,
|
||||
crash-loop cap). *(Pending)*: it decides mount placement from `volumes.csv` and
|
||||
picks the filesystem binary from `filesystems.csv` — today it hands every
|
||||
FAT-shaped volume to the FAT service and the FAT service carries hardcoded mount
|
||||
prefixes. Those tables are CSV configuration, read by it (the policy), enforced
|
||||
by nobody else:
|
||||
|
||||
- `filesystems.csv` — content signature → filesystem binary. Adding a
|
||||
filesystem adds a row.
|
||||
- `volumes.csv` — the mount map, danos's fstab: **volume identity → mount
|
||||
- `filesystems.csv` *(pending)* — content signature → filesystem binary. Adding
|
||||
a filesystem adds a row.
|
||||
- `volumes.csv` *(pending)* — the mount map, danos's fstab: **volume identity → mount
|
||||
prefix**, keyed on content identity and never on port, path, or arrival
|
||||
order (the lesson of Linux's `/dev/sda1`-era fstab, which broke on every
|
||||
port move until `UUID=` replaced it). Identity is read off the medium by
|
||||
@@ -166,9 +175,9 @@ be served with the previous card's filesystem state.
|
||||
| Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick |
|
||||
| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
|
||||
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
|
||||
| Volume manager *(planned)* | the storage provider's channel death / the manager's child events | kill each filesystem service of that device's volumes; retire their kernel mounts; remember the volume identity | one removal path; mounts never dangle; log persistence stops *cleanly* |
|
||||
| Volume manager *(removal built; remount bench-pending)* | the storage device leaving the device-manager tree (poll) | kill the filesystem service of that device's volume; its kernel mounts retire | one removal path; mounts never dangle; log persistence stops *cleanly* |
|
||||
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the FAT dirty flag on disk marks the unclean removal; the process never serves from behind a dead channel |
|
||||
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); *(planned)* ownership-gated `fs_unmount` | resolution under a dead mount is `not_found`, not a hang |
|
||||
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated `fs_unmount` (built, V0) | resolution under a dead mount is `not_found`, not a hang |
|
||||
| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
|
||||
|
||||
**On return** the same table runs upward in reverse: the bus re-enumerates and
|
||||
|
||||
Reference in New Issue
Block a user