Files
danos/docs/volume-manager-plan.md
T

125 lines
7.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# The volume manager: the plan
*2026-08-09. Executes the settled design in
[storage-architecture.md](file-system-development/storage-architecture.md) and
[storage-design-rationale.md](file-system-development/storage-design-rationale.md)
(decisions 1–8). Track discipline as always: one commit per coherent step, suite
green at phase boundaries, every new test shown to fail against the old
behavior, one QEMU suite at a time, work on main.*
**No open decisions.** Decision 4 is settled in the rationale as per-sender
range confinement at the provider on one serving endpoint — the badge-scoped
provider pattern the xHCI bus already uses, applied to blocks. The volume
manager sets each filesystem process's range; the driver clamps and translates
every transfer by the sender's kernel-stamped badge; filesystems address
volume-relative LBAs from 0 and the FAT engine's `base_lba` is deleted rather
than moved. Every other decision the phases below execute is recorded in the
rationale (decisions 1–8); nothing in this plan waits on a choice.
## V0 — `fs_unmount` ownership (the defect fix)
The kernel records the mounting process on each mount slot; `fs_unmount` is
refused for any caller but the owner (the `/protocol` special case stays).
Death cleanup is unaffected (the lazy sweep is not an unmount). Discrimination:
a fixture unmounting a prefix it does not own must be refused — fails against
today's kernel, which lets any process unmount anything.
## V1 — the filesystem harness (extraction, no behavior change)
fat's 434-line shell becomes `library/file-system/harness` (name per
convention): establishment, the badge-scoped open-node table, the nine
vfs-protocol handlers, mount registration, bring-up/teardown. fat becomes
engine + on-disk + a thin main wiring the harness. Suite green is the gate;
nothing observable changes. This lands FIRST so every later phase touches the
harness once, not fat and the harness both.
## V2 — the driver mechanism: ranges, confinement, presence
- Block protocol additions (appended; numbering holds): `define_range`
(volume-manager-only in practice — see grants), and the pushed
`medium_changed` event (present/absent + change counter).
- usb-storage: per-badge range table (clamp + translate per sender), range
definitions from the volume manager, and presence: a slow idle-time
TEST UNIT READY poll plus sense-key inspection on failed transfers, emitting
`medium_changed` on transitions. QEMU test lever: `eject` /
`blockdev-remove-medium` against a `removable=on` usb-storage device; if
QEMU's model refuses, the fallback drill is device_del/add of the whole
stick (the H-matrix already proves that path) and presence gets its real
test on the bench with a card reader.
- Discrimination: a fixture transferring outside its assigned range must be
refused; fails against a driver without the clamp.
**Sequencing (as executed).** V2a landed the range clamp + its discrimination
fixture (block-range) — the security mechanism is testable in isolation. V2b
adds `medium_changed` and makes usb-storage a publisher that emits it on
presence transitions, but its END-TO-END test (eject → medium_changed →
unmount/remount) lands in V4 with the real consumer, the volume manager —
rather than a throwaway subscriber fixture V3 would immediately replace. Same
work, no duplicated scaffolding.
## V3 — the volume manager service
New binary `system/services/volume-manager`, spawned by init, serving the
(genuinely singular) name `volume-manager`. Duties, all moved OUT of fat:
- subscribe to the device manager; consumer-hello each storage provider for
its block channel;
- probe: partition table walk (MBR now, GPT next — the walk LEAVES the FAT
engine) and content identity (the decision-8 ladder: GPT GUID → fs UUID →
FAT serial+label → MBR signature+index → anonymous);
- configuration: `filesystems.csv` (signature → filesystem binary) and
`volumes.csv` (identity → mount prefix; the fstab). Boot volume identity
recorded at first sight of `/system/configuration`;
- define ranges on the driver; spawn one filesystem process per volume
(argv: volume id); answer each filesystem's startup hello with its volume
channel (the reply-capability path, same as the device manager's);
- supervise: hello/mount deadline, crash-loop cap, reap on removal.
fat sheds `acquireVolume` and its device-manager grant; filesystem binaries
get `open volume-manager` only — a filesystem cannot acquire, only be given.
Grants move with the code in the same commits.
**Sequencing (as executed).** V3a (discovery+probe) and V3b (the flip: spawn +
confine + hand over the channel, single volume) landed the core. The V3c items —
the `volumes.csv` mount map, the fuller identity ladder (FAT serial, GPT GUID),
and multi-volume spawning — mainly serve the MULTI-volume drills (two-partitions,
clone-policy). The user's goal is the single-volume boot-stick removal lifecycle,
so V4's removal lifecycle runs next on the single volume, and the multi-volume
work + its drills become a documented follow-on (V3c/multi-volume). fat keeps its
hardcoded mount prefixes until the mount map lands.
## V4 — the removal lifecycle, end to end
Two triggers, one path: storage-channel death and `medium_changed(absent)`
both drive kill-the-filesystem-process + retire-its-mounts; return (device
re-report or `medium_changed(present)`) drives re-probe → respawn → remount
at the identity's prefix. The logger gains resume patience (retry flushes on
the fat cadence, never abandon) and the ring-wrap gap marker. QEMU cases:
- `volume-replug`: yank the boot stick mid-run, replug, assert remount at the
same prefixes and the logger appending to the SAME boot-stamp tree with the
gap marked. Discrimination: against pre-V4, fat wedges (`mounted` forever)
and no remount happens.
- `volume-two-partitions`: a two-partition FAT image → two volumes, two
filesystem processes, two mounts from one stick; yank once, both die; return
once, both remount. Proves multi-volume and the reap breadth. (New image
fixture beside make-fat-image.py.)
- `volume-clone-policy`: two sticks with identical FAT serials — first keeps
the mapped name, second mounts suffixed, loudly logged.
- The **lifecycle conformance drill**, parameterized by filesystem: mount,
serve, yank mid-write, verify honest loss (dirty flag set, gap said),
replug, remount. FAT is implementation #1; the drill is the definition of
"danos supports filesystem X".
## V5 — close-out
Full suite green; the architecture doc's *(planned)* markers flip to built;
`fs_unmount` ownership, ranges, presence, identity, and the volume manager
lose their future-tense; memory updated; Ryzen bench note: pull the stick,
watch the log gap get marked, plug it back anywhere.
**Not in this track** (recorded so absence is deliberate): GPT parsing beyond
the identity read (row exists in the prober's ladder; full GPT when a GPT
medium matters), AHCI/NVMe drivers (NVMe gated on the shm-ring data plane),
formatting/entropy, per-process namespaces.