docs: the volume-manager plan — V0 unmount ownership through V5 close-out
This commit is contained in:
@@ -0,0 +1,115 @@
|
||||
# The volume manager: the plan
|
||||
|
||||
*2026-08-09. Executes the settled design in
|
||||
[storage-architecture.md](file-system-development/storage-architecture.md) and
|
||||
[storage-design-rationale.md](file-system-development/storage-design-rationale.md)
|
||||
(decisions 1–8). Track discipline as always: one commit per coherent step, suite
|
||||
green at phase boundaries, every new test shown to fail against the old
|
||||
behavior, one QEMU suite at a time, work on main.*
|
||||
|
||||
**One refinement of decision 4, flagged for sign-off rather than silently
|
||||
applied.** The decision said endpoint-per-volume. The service harness serves one
|
||||
endpoint per process, and kernel-ipc has no wait-on-many; a driver serving N
|
||||
range endpoints would need threads or a multi-endpoint harness — real machinery,
|
||||
none of it needed for the security property. The property ("a channel carries
|
||||
exactly the authority it grants") is delivered instead by **per-sender range
|
||||
confinement**: every packet already arrives with the kernel-stamped,
|
||||
unforgeable badge; the storage driver keeps a per-badge range (set by the
|
||||
volume manager, which spawned the filesystem process and knows its id) and
|
||||
clamps-and-translates every transfer by the sender's range. A filesystem
|
||||
process addresses volume-relative LBAs from 0; the provider adds the base —
|
||||
offset translation at the provider, exactly the Fuchsia session-mapping shape,
|
||||
and the FAT engine's `base_lba` code is deleted rather than moved.
|
||||
Endpoint-per-volume can still arrive later with a multi-endpoint harness; the
|
||||
wire contract does not change either way. **If this refinement is wrong, say so
|
||||
before V2.**
|
||||
|
||||
## V0 — `fs_unmount` ownership (the defect fix)
|
||||
|
||||
The kernel records the mounting process on each mount slot; `fs_unmount` is
|
||||
refused for any caller but the owner (the `/protocol` special case stays).
|
||||
Death cleanup is unaffected (the lazy sweep is not an unmount). Discrimination:
|
||||
a fixture unmounting a prefix it does not own must be refused — fails against
|
||||
today's kernel, which lets any process unmount anything.
|
||||
|
||||
## V1 — the filesystem harness (extraction, no behavior change)
|
||||
|
||||
fat's 434-line shell becomes `library/file-system/harness` (name per
|
||||
convention): establishment, the badge-scoped open-node table, the nine
|
||||
vfs-protocol handlers, mount registration, bring-up/teardown. fat becomes
|
||||
engine + on-disk + a thin main wiring the harness. Suite green is the gate;
|
||||
nothing observable changes. This lands FIRST so every later phase touches the
|
||||
harness once, not fat and the harness both.
|
||||
|
||||
## V2 — the driver mechanism: ranges, confinement, presence
|
||||
|
||||
- Block protocol additions (appended; numbering holds): `define_range`
|
||||
(volume-manager-only in practice — see grants), and the pushed
|
||||
`medium_changed` event (present/absent + change counter).
|
||||
- usb-storage: per-badge range table (clamp + translate per sender), range
|
||||
definitions from the volume manager, and presence: a slow idle-time
|
||||
TEST UNIT READY poll plus sense-key inspection on failed transfers, emitting
|
||||
`medium_changed` on transitions. QEMU test lever: `eject` /
|
||||
`blockdev-remove-medium` against a `removable=on` usb-storage device; if
|
||||
QEMU's model refuses, the fallback drill is device_del/add of the whole
|
||||
stick (the H-matrix already proves that path) and presence gets its real
|
||||
test on the bench with a card reader.
|
||||
- Discrimination: a fixture transferring outside its assigned range must be
|
||||
refused; fails against a driver without the clamp.
|
||||
|
||||
## V3 — the volume manager service
|
||||
|
||||
New binary `system/services/volume-manager`, spawned by init, serving the
|
||||
(genuinely singular) name `volume-manager`. Duties, all moved OUT of fat:
|
||||
|
||||
- subscribe to the device manager; consumer-hello each storage provider for
|
||||
its block channel;
|
||||
- probe: partition table walk (MBR now, GPT next — the walk LEAVES the FAT
|
||||
engine) and content identity (the decision-8 ladder: GPT GUID → fs UUID →
|
||||
FAT serial+label → MBR signature+index → anonymous);
|
||||
- configuration: `filesystems.csv` (signature → filesystem binary) and
|
||||
`volumes.csv` (identity → mount prefix; the fstab). Boot volume identity
|
||||
recorded at first sight of `/system/configuration`;
|
||||
- define ranges on the driver; spawn one filesystem process per volume
|
||||
(argv: volume id); answer each filesystem's startup hello with its volume
|
||||
channel (the reply-capability path, same as the device manager's);
|
||||
- supervise: hello/mount deadline, crash-loop cap, reap on removal.
|
||||
|
||||
fat sheds `acquireVolume` and its device-manager grant; filesystem binaries
|
||||
get `open volume-manager` only — a filesystem cannot acquire, only be given.
|
||||
Grants move with the code in the same commits.
|
||||
|
||||
## V4 — the removal lifecycle, end to end
|
||||
|
||||
Two triggers, one path: storage-channel death and `medium_changed(absent)`
|
||||
both drive kill-the-filesystem-process + retire-its-mounts; return (device
|
||||
re-report or `medium_changed(present)`) drives re-probe → respawn → remount
|
||||
at the identity's prefix. The logger gains resume patience (retry flushes on
|
||||
the fat cadence, never abandon) and the ring-wrap gap marker. QEMU cases:
|
||||
|
||||
- `volume-replug`: yank the boot stick mid-run, replug, assert remount at the
|
||||
same prefixes and the logger appending to the SAME boot-stamp tree with the
|
||||
gap marked. Discrimination: against pre-V4, fat wedges (`mounted` forever)
|
||||
and no remount happens.
|
||||
- `volume-two-partitions`: a two-partition FAT image → two volumes, two
|
||||
filesystem processes, two mounts from one stick; yank once, both die; return
|
||||
once, both remount. Proves multi-volume and the reap breadth. (New image
|
||||
fixture beside make-fat-image.py.)
|
||||
- `volume-clone-policy`: two sticks with identical FAT serials — first keeps
|
||||
the mapped name, second mounts suffixed, loudly logged.
|
||||
- The **lifecycle conformance drill**, parameterized by filesystem: mount,
|
||||
serve, yank mid-write, verify honest loss (dirty flag set, gap said),
|
||||
replug, remount. FAT is implementation #1; the drill is the definition of
|
||||
"danos supports filesystem X".
|
||||
|
||||
## V5 — close-out
|
||||
|
||||
Full suite green; the architecture doc's *(planned)* markers flip to built;
|
||||
`fs_unmount` ownership, ranges, presence, identity, and the volume manager
|
||||
lose their future-tense; memory updated; Ryzen bench note: pull the stick,
|
||||
watch the log gap get marked, plug it back anywhere.
|
||||
|
||||
**Not in this track** (recorded so absence is deliberate): GPT parsing beyond
|
||||
the identity read (row exists in the prober's ladder; full GPT when a GPT
|
||||
medium matters), AHCI/NVMe drivers (NVMe gated on the shm-ring data plane),
|
||||
formatting/entropy, per-process namespaces.
|
||||
Reference in New Issue
Block a user