From 5e89b111cf2f9bf6fbaab5023137d8ab9ec7999c Mon Sep 17 00:00:00 2001 From: Daniel Samson <12231216+daniel-samson@users.noreply.github.com> Date: Sun, 9 Aug 2026 20:20:20 +0100 Subject: [PATCH] docs: flip the storage-architecture status markers the V0-V4 track made real MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The volume manager, per-sender range confinement + gate, medium_changed event, per-volume spawning + supervision, ownership-gated fs_unmount, and the removal half of the lifecycle are built. Left honestly pending: the volumes.csv/filesystems.csv maps, the fuller identity ladder, multi-volume, the VM consuming medium_changed (removal uses device-presence polling), and remount-on-replug end-to-end (bench-pending — QEMU can't re-present the boot-controller device). Full suite 127/127. --- .../storage-architecture.md | 53 +++++++++++-------- 1 file changed, 31 insertions(+), 22 deletions(-) diff --git a/docs/file-system-development/storage-architecture.md b/docs/file-system-development/storage-architecture.md index 4f8a30e..f31ae13 100644 --- a/docs/file-system-development/storage-architecture.md +++ b/docs/file-system-development/storage-architecture.md @@ -1,13 +1,20 @@ # The storage architecture: layers, boundaries, responsibilities > **Status:** the layered model below is the settled design -> ([storage-design-rationale.md](storage-design-rationale.md) records how it -> was reached and what the surveyed systems taught). The data path — vfs -> protocol, kernel mount routing, the FAT service, the block protocol, -> usb-storage — is **built**. The volume manager, the driver's range -> mechanism, per-volume filesystem spawning, and the removal lifecycle are -> **planned**; until they land, the FAT service performs volume-manager duties -> itself (marked below). This document is the reference for both states. +> ([storage-design-rationale.md](storage-design-rationale.md) records how it was +> reached, and [volume-manager-plan.md](../volume-manager-plan.md) how it was +> built). **Built** (the volume-manager track, V0–V4): the data path, the driver +> range confinement (per-sender clamp + the confinement gate), the `medium_changed` +> presence event, the volume manager itself — it probes the partition table, +> confines each filesystem to its partition, spawns one filesystem per volume, and +> supervises it — and the removal half of the lifecycle (a pulled stick unmounts). +> **Still pending**: the fuller identity ladder and the `volumes.csv` mount map, +> multi-volume (one FAT volume today), the volume manager *consuming* +> `medium_changed` (removal is detected by device-presence polling; the event is +> published but only a card-reader medium change needs the subscription), and the +> remount-on-replug end-to-end (the logic is in place; QEMU can't re-present the +> boot-controller device, so it is bench-verified). A few markers below are left +> where a duty is still pending. ## The model @@ -47,7 +54,7 @@ Three kinds of boundary, deliberately different: - **Control-plane relationships** sit beside the data path, never on it. Two supervisors, one per layer: the **device manager** wires and revives the device layers (bus and storage drivers — devices only); the **volume - manager** *(planned)* wires and revives the volume layer (filesystem + manager** *(built)* wires and revives the volume layer (filesystem services). Neither touches steady-state I/O. ## Who does what @@ -71,19 +78,21 @@ whole disk plus a polite base offset would let a buggy or compromised filesystem scribble the neighboring partition — the same authority-overshoot the device-authority track eliminated for MMIO and DMA. -**Volume manager** *(planned; today the FAT service squats on these duties)*: -the policy home of the volume layer, one service, supervised by init. It -subscribes to the device manager's child events; when a storage provider -appears it consumer-hellos for the block channel, reads the partition table -and the first blocks itself (**it** is the prober), consults its -configuration, defines sub-ranges on the driver, spawns the matching -filesystem service per volume with that volume's channel, supervises it, and -decides mount placement. Its tables are CSV configuration, read by it (the -policy), enforced by nobody else: +**Volume manager** *(built; `system/services/volume-manager`)*: the policy home +of the volume layer, one service, supervised by init. It watches the device +manager's tree for a storage provider; when one appears it consumer-hellos for +the block channel, reads the partition table and the first blocks itself +(**it** is the prober), defines the volume's sub-range on the driver, spawns the +matching filesystem service confined to that range, and supervises it (backoff, +crash-loop cap). *(Pending)*: it decides mount placement from `volumes.csv` and +picks the filesystem binary from `filesystems.csv` — today it hands every +FAT-shaped volume to the FAT service and the FAT service carries hardcoded mount +prefixes. Those tables are CSV configuration, read by it (the policy), enforced +by nobody else: -- `filesystems.csv` — content signature → filesystem binary. Adding a - filesystem adds a row. -- `volumes.csv` — the mount map, danos's fstab: **volume identity → mount +- `filesystems.csv` *(pending)* — content signature → filesystem binary. Adding + a filesystem adds a row. +- `volumes.csv` *(pending)* — the mount map, danos's fstab: **volume identity → mount prefix**, keyed on content identity and never on port, path, or arrival order (the lesson of Linux's `/dev/sda1`-era fstab, which broke on every port move until `UUID=` replaced it). Identity is read off the medium by @@ -166,9 +175,9 @@ be served with the previous card's filesystem state. | Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick | | Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds | | Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang | -| Volume manager *(planned)* | the storage provider's channel death / the manager's child events | kill each filesystem service of that device's volumes; retire their kernel mounts; remember the volume identity | one removal path; mounts never dangle; log persistence stops *cleanly* | +| Volume manager *(removal built; remount bench-pending)* | the storage device leaving the device-manager tree (poll) | kill the filesystem service of that device's volume; its kernel mounts retire | one removal path; mounts never dangle; log persistence stops *cleanly* | | Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the FAT dirty flag on disk marks the unclean removal; the process never serves from behind a dead channel | -| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); *(planned)* ownership-gated `fs_unmount` | resolution under a dead mount is `not_found`, not a hang | +| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated `fs_unmount` (built, V0) | resolution under a dead mount is `not_found`, not a hang | | Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media | **On return** the same table runs upward in reverse: the bus re-enumerates and