diff --git a/docs/volume-manager-plan.md b/docs/volume-manager-plan.md new file mode 100644 index 0000000..d874279 --- /dev/null +++ b/docs/volume-manager-plan.md @@ -0,0 +1,115 @@ +# The volume manager: the plan + +*2026-08-09. Executes the settled design in +[storage-architecture.md](file-system-development/storage-architecture.md) and +[storage-design-rationale.md](file-system-development/storage-design-rationale.md) +(decisions 1–8). Track discipline as always: one commit per coherent step, suite +green at phase boundaries, every new test shown to fail against the old +behavior, one QEMU suite at a time, work on main.* + +**One refinement of decision 4, flagged for sign-off rather than silently +applied.** The decision said endpoint-per-volume. The service harness serves one +endpoint per process, and kernel-ipc has no wait-on-many; a driver serving N +range endpoints would need threads or a multi-endpoint harness — real machinery, +none of it needed for the security property. The property ("a channel carries +exactly the authority it grants") is delivered instead by **per-sender range +confinement**: every packet already arrives with the kernel-stamped, +unforgeable badge; the storage driver keeps a per-badge range (set by the +volume manager, which spawned the filesystem process and knows its id) and +clamps-and-translates every transfer by the sender's range. A filesystem +process addresses volume-relative LBAs from 0; the provider adds the base — +offset translation at the provider, exactly the Fuchsia session-mapping shape, +and the FAT engine's `base_lba` code is deleted rather than moved. +Endpoint-per-volume can still arrive later with a multi-endpoint harness; the +wire contract does not change either way. **If this refinement is wrong, say so +before V2.** + +## V0 — `fs_unmount` ownership (the defect fix) + +The kernel records the mounting process on each mount slot; `fs_unmount` is +refused for any caller but the owner (the `/protocol` special case stays). +Death cleanup is unaffected (the lazy sweep is not an unmount). Discrimination: +a fixture unmounting a prefix it does not own must be refused — fails against +today's kernel, which lets any process unmount anything. + +## V1 — the filesystem harness (extraction, no behavior change) + +fat's 434-line shell becomes `library/file-system/harness` (name per +convention): establishment, the badge-scoped open-node table, the nine +vfs-protocol handlers, mount registration, bring-up/teardown. fat becomes +engine + on-disk + a thin main wiring the harness. Suite green is the gate; +nothing observable changes. This lands FIRST so every later phase touches the +harness once, not fat and the harness both. + +## V2 — the driver mechanism: ranges, confinement, presence + +- Block protocol additions (appended; numbering holds): `define_range` + (volume-manager-only in practice — see grants), and the pushed + `medium_changed` event (present/absent + change counter). +- usb-storage: per-badge range table (clamp + translate per sender), range + definitions from the volume manager, and presence: a slow idle-time + TEST UNIT READY poll plus sense-key inspection on failed transfers, emitting + `medium_changed` on transitions. QEMU test lever: `eject` / + `blockdev-remove-medium` against a `removable=on` usb-storage device; if + QEMU's model refuses, the fallback drill is device_del/add of the whole + stick (the H-matrix already proves that path) and presence gets its real + test on the bench with a card reader. +- Discrimination: a fixture transferring outside its assigned range must be + refused; fails against a driver without the clamp. + +## V3 — the volume manager service + +New binary `system/services/volume-manager`, spawned by init, serving the +(genuinely singular) name `volume-manager`. Duties, all moved OUT of fat: + +- subscribe to the device manager; consumer-hello each storage provider for + its block channel; +- probe: partition table walk (MBR now, GPT next — the walk LEAVES the FAT + engine) and content identity (the decision-8 ladder: GPT GUID → fs UUID → + FAT serial+label → MBR signature+index → anonymous); +- configuration: `filesystems.csv` (signature → filesystem binary) and + `volumes.csv` (identity → mount prefix; the fstab). Boot volume identity + recorded at first sight of `/system/configuration`; +- define ranges on the driver; spawn one filesystem process per volume + (argv: volume id); answer each filesystem's startup hello with its volume + channel (the reply-capability path, same as the device manager's); +- supervise: hello/mount deadline, crash-loop cap, reap on removal. + +fat sheds `acquireVolume` and its device-manager grant; filesystem binaries +get `open volume-manager` only — a filesystem cannot acquire, only be given. +Grants move with the code in the same commits. + +## V4 — the removal lifecycle, end to end + +Two triggers, one path: storage-channel death and `medium_changed(absent)` +both drive kill-the-filesystem-process + retire-its-mounts; return (device +re-report or `medium_changed(present)`) drives re-probe → respawn → remount +at the identity's prefix. The logger gains resume patience (retry flushes on +the fat cadence, never abandon) and the ring-wrap gap marker. QEMU cases: + +- `volume-replug`: yank the boot stick mid-run, replug, assert remount at the + same prefixes and the logger appending to the SAME boot-stamp tree with the + gap marked. Discrimination: against pre-V4, fat wedges (`mounted` forever) + and no remount happens. +- `volume-two-partitions`: a two-partition FAT image → two volumes, two + filesystem processes, two mounts from one stick; yank once, both die; return + once, both remount. Proves multi-volume and the reap breadth. (New image + fixture beside make-fat-image.py.) +- `volume-clone-policy`: two sticks with identical FAT serials — first keeps + the mapped name, second mounts suffixed, loudly logged. +- The **lifecycle conformance drill**, parameterized by filesystem: mount, + serve, yank mid-write, verify honest loss (dirty flag set, gap said), + replug, remount. FAT is implementation #1; the drill is the definition of + "danos supports filesystem X". + +## V5 — close-out + +Full suite green; the architecture doc's *(planned)* markers flip to built; +`fs_unmount` ownership, ranges, presence, identity, and the volume manager +lose their future-tense; memory updated; Ryzen bench note: pull the stick, +watch the log gap get marked, plug it back anywhere. + +**Not in this track** (recorded so absence is deliberate): GPT parsing beyond +the identity read (row exists in the prober's ladder; full GPT when a GPT +medium matters), AHCI/NVMe drivers (NVMe gated on the shm-ring data plane), +formatting/entropy, per-process namespaces.