Files
danos/docs/volume-manager-plan.md
T

116 lines
6.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# The volume manager: the plan
*2026-08-09. Executes the settled design in
[storage-architecture.md](file-system-development/storage-architecture.md) and
[storage-design-rationale.md](file-system-development/storage-design-rationale.md)
(decisions 1–8). Track discipline as always: one commit per coherent step, suite
green at phase boundaries, every new test shown to fail against the old
behavior, one QEMU suite at a time, work on main.*
**One refinement of decision 4, flagged for sign-off rather than silently
applied.** The decision said endpoint-per-volume. The service harness serves one
endpoint per process, and kernel-ipc has no wait-on-many; a driver serving N
range endpoints would need threads or a multi-endpoint harness — real machinery,
none of it needed for the security property. The property ("a channel carries
exactly the authority it grants") is delivered instead by **per-sender range
confinement**: every packet already arrives with the kernel-stamped,
unforgeable badge; the storage driver keeps a per-badge range (set by the
volume manager, which spawned the filesystem process and knows its id) and
clamps-and-translates every transfer by the sender's range. A filesystem
process addresses volume-relative LBAs from 0; the provider adds the base —
offset translation at the provider, exactly the Fuchsia session-mapping shape,
and the FAT engine's `base_lba` code is deleted rather than moved.
Endpoint-per-volume can still arrive later with a multi-endpoint harness; the
wire contract does not change either way. **If this refinement is wrong, say so
before V2.**
## V0 — `fs_unmount` ownership (the defect fix)
The kernel records the mounting process on each mount slot; `fs_unmount` is
refused for any caller but the owner (the `/protocol` special case stays).
Death cleanup is unaffected (the lazy sweep is not an unmount). Discrimination:
a fixture unmounting a prefix it does not own must be refused — fails against
today's kernel, which lets any process unmount anything.
## V1 — the filesystem harness (extraction, no behavior change)
fat's 434-line shell becomes `library/file-system/harness` (name per
convention): establishment, the badge-scoped open-node table, the nine
vfs-protocol handlers, mount registration, bring-up/teardown. fat becomes
engine + on-disk + a thin main wiring the harness. Suite green is the gate;
nothing observable changes. This lands FIRST so every later phase touches the
harness once, not fat and the harness both.
## V2 — the driver mechanism: ranges, confinement, presence
- Block protocol additions (appended; numbering holds): `define_range`
(volume-manager-only in practice — see grants), and the pushed
`medium_changed` event (present/absent + change counter).
- usb-storage: per-badge range table (clamp + translate per sender), range
definitions from the volume manager, and presence: a slow idle-time
TEST UNIT READY poll plus sense-key inspection on failed transfers, emitting
`medium_changed` on transitions. QEMU test lever: `eject` /
`blockdev-remove-medium` against a `removable=on` usb-storage device; if
QEMU's model refuses, the fallback drill is device_del/add of the whole
stick (the H-matrix already proves that path) and presence gets its real
test on the bench with a card reader.
- Discrimination: a fixture transferring outside its assigned range must be
refused; fails against a driver without the clamp.
## V3 — the volume manager service
New binary `system/services/volume-manager`, spawned by init, serving the
(genuinely singular) name `volume-manager`. Duties, all moved OUT of fat:
- subscribe to the device manager; consumer-hello each storage provider for
its block channel;
- probe: partition table walk (MBR now, GPT next — the walk LEAVES the FAT
engine) and content identity (the decision-8 ladder: GPT GUID → fs UUID →
FAT serial+label → MBR signature+index → anonymous);
- configuration: `filesystems.csv` (signature → filesystem binary) and
`volumes.csv` (identity → mount prefix; the fstab). Boot volume identity
recorded at first sight of `/system/configuration`;
- define ranges on the driver; spawn one filesystem process per volume
(argv: volume id); answer each filesystem's startup hello with its volume
channel (the reply-capability path, same as the device manager's);
- supervise: hello/mount deadline, crash-loop cap, reap on removal.
fat sheds `acquireVolume` and its device-manager grant; filesystem binaries
get `open volume-manager` only — a filesystem cannot acquire, only be given.
Grants move with the code in the same commits.
## V4 — the removal lifecycle, end to end
Two triggers, one path: storage-channel death and `medium_changed(absent)`
both drive kill-the-filesystem-process + retire-its-mounts; return (device
re-report or `medium_changed(present)`) drives re-probe → respawn → remount
at the identity's prefix. The logger gains resume patience (retry flushes on
the fat cadence, never abandon) and the ring-wrap gap marker. QEMU cases:
- `volume-replug`: yank the boot stick mid-run, replug, assert remount at the
same prefixes and the logger appending to the SAME boot-stamp tree with the
gap marked. Discrimination: against pre-V4, fat wedges (`mounted` forever)
and no remount happens.
- `volume-two-partitions`: a two-partition FAT image → two volumes, two
filesystem processes, two mounts from one stick; yank once, both die; return
once, both remount. Proves multi-volume and the reap breadth. (New image
fixture beside make-fat-image.py.)
- `volume-clone-policy`: two sticks with identical FAT serials — first keeps
the mapped name, second mounts suffixed, loudly logged.
- The **lifecycle conformance drill**, parameterized by filesystem: mount,
serve, yank mid-write, verify honest loss (dirty flag set, gap said),
replug, remount. FAT is implementation #1; the drill is the definition of
"danos supports filesystem X".
## V5 — close-out
Full suite green; the architecture doc's *(planned)* markers flip to built;
`fs_unmount` ownership, ranges, presence, identity, and the volume manager
lose their future-tense; memory updated; Ryzen bench note: pull the stick,
watch the log gap get marked, plug it back anywhere.
**Not in this track** (recorded so absence is deliberate): GPT parsing beyond
the identity read (row exists in the prober's ladder; full GPT when a GPT
medium matters), AHCI/NVMe drivers (NVMe gated on the shm-ring data plane),
formatting/entropy, per-process namespaces.