Files
danos/docs/volume-manager-plan.md
Daniel Samson 8216be991d docs: correct the storage docs' V0-V4 status — the V5 flip missed several
A V5 close-out audit (docs against the actual code) found the storage docs
overclaiming in both directions: the status header was flipped to "built" but
several body markers were not, and two passages describe mechanisms the code
never implemented. Nine confirmed, adversarially verified against the source:

Underclaims (marked Planned, actually built):
- storage-architecture: driver named sub-ranges / range confinement (V2, tested
  by the block-range case); the pushed medium_changed event (published today
  from a TEST UNIT READY poll); ownership-gated fs_unmount (V0, process.zig
  gates it with EPERM); the filesystem-harness extraction (V1). Scoped the
  remaining *planned* to the genuinely-pending parts (native-signal translation,
  the volume manager consuming medium_changed).

Overclaims (described, never built):
- storage-architecture: the "FAT dirty flag on disk" guarantee — no on-disk
  dirty/clean-shutdown bit exists; only an in-memory device-dirty bool gating a
  device write-cache flush on close. Fixed in all three places.
- storage-architecture: fat "acquires its own volume (first mass-storage child
  by enumeration order)" — the V3b flip removed self-acquisition; fat is handed
  its volume id and channel by the volume manager.
- rationale + plan: the FAT engine's base_lba "deleted rather than moved" — it
  and the engine's MBR walk still exist as now-inert legacy; the authoritative
  walk lives in partition.zig.
- rationale: NVMe namespaces "decision 4 settles as endpoint-per-volume" —
  decision 4 settles the opposite (per-sender confinement, one endpoint);
  endpoint-per-volume is named only as an unbuilt future refactor.
- rationale: the five-rung identity ladder and volumes.csv map stated in flat
  present tense — only rung 4 (MBR signature + index) is built; added the
  build-status hedge and marked each rung.

Docs only; no code or behavior change. Suite unaffected (127/127).
2026-08-09 21:19:29 +01:00

7.1 KiB
Raw Permalink Blame History

The volume manager: the plan

2026-08-09. Executes the settled design in storage-architecture.md and storage-design-rationale.md (decisions 1–8). Track discipline as always: one commit per coherent step, suite green at phase boundaries, every new test shown to fail against the old behavior, one QEMU suite at a time, work on main.

No open decisions. Decision 4 is settled in the rationale as per-sender range confinement at the provider on one serving endpoint — the badge-scoped provider pattern the xHCI bus already uses, applied to blocks. The volume manager sets each filesystem process's range; the driver clamps and translates every transfer by the sender's kernel-stamped badge; filesystems address volume-relative LBAs from 0 (on the confined path the FAT engine's base_lba resolves to 0; the field and the engine's own MBR walk remain as now-inert legacy, the authoritative walk living in the volume manager's partition.zig). Every other decision the phases below execute is recorded in the rationale (decisions 1–8); nothing in this plan waits on a choice.

V0 — fs_unmount ownership (the defect fix)

The kernel records the mounting process on each mount slot; fs_unmount is refused for any caller but the owner (the /protocol special case stays). Death cleanup is unaffected (the lazy sweep is not an unmount). Discrimination: a fixture unmounting a prefix it does not own must be refused — fails against today's kernel, which lets any process unmount anything.

V1 — the filesystem harness (extraction, no behavior change)

fat's 434-line shell becomes library/file-system/harness (name per convention): establishment, the badge-scoped open-node table, the nine vfs-protocol handlers, mount registration, bring-up/teardown. fat becomes engine + on-disk + a thin main wiring the harness. Suite green is the gate; nothing observable changes. This lands FIRST so every later phase touches the harness once, not fat and the harness both.

V2 — the driver mechanism: ranges, confinement, presence

  • Block protocol additions (appended; numbering holds): define_range (volume-manager-only in practice — see grants), and the pushed medium_changed event (present/absent + change counter).
  • usb-storage: per-badge range table (clamp + translate per sender), range definitions from the volume manager, and presence: a slow idle-time TEST UNIT READY poll plus sense-key inspection on failed transfers, emitting medium_changed on transitions. QEMU test lever: eject / blockdev-remove-medium against a removable=on usb-storage device; if QEMU's model refuses, the fallback drill is device_del/add of the whole stick (the H-matrix already proves that path) and presence gets its real test on the bench with a card reader.
  • Discrimination: a fixture transferring outside its assigned range must be refused; fails against a driver without the clamp.

Sequencing (as executed). V2a landed the range clamp + its discrimination fixture (block-range) — the security mechanism is testable in isolation. V2b adds medium_changed and makes usb-storage a publisher that emits it on presence transitions, but its END-TO-END test (eject → medium_changed → unmount/remount) lands in V4 with the real consumer, the volume manager — rather than a throwaway subscriber fixture V3 would immediately replace. Same work, no duplicated scaffolding.

V3 — the volume manager service

New binary system/services/volume-manager, spawned by init, serving the (genuinely singular) name volume-manager. Duties, all moved OUT of fat:

  • subscribe to the device manager; consumer-hello each storage provider for its block channel;
  • probe: partition table walk (MBR now, GPT next — the walk LEAVES the FAT engine) and content identity (the decision-8 ladder: GPT GUID → fs UUID → FAT serial+label → MBR signature+index → anonymous);
  • configuration: filesystems.csv (signature → filesystem binary) and volumes.csv (identity → mount prefix; the fstab). Boot volume identity recorded at first sight of /system/configuration;
  • define ranges on the driver; spawn one filesystem process per volume (argv: volume id); answer each filesystem's startup hello with its volume channel (the reply-capability path, same as the device manager's);
  • supervise: hello/mount deadline, crash-loop cap, reap on removal.

fat sheds acquireVolume and its device-manager grant; filesystem binaries get open volume-manager only — a filesystem cannot acquire, only be given. Grants move with the code in the same commits.

Sequencing (as executed). V3a (discovery+probe) and V3b (the flip: spawn + confine + hand over the channel, single volume) landed the core. The V3c items — the volumes.csv mount map, the fuller identity ladder (FAT serial, GPT GUID), and multi-volume spawning — mainly serve the MULTI-volume drills (two-partitions, clone-policy). The user's goal is the single-volume boot-stick removal lifecycle, so V4's removal lifecycle runs next on the single volume, and the multi-volume work + its drills become a documented follow-on (V3c/multi-volume). fat keeps its hardcoded mount prefixes until the mount map lands.

V4 — the removal lifecycle, end to end

Two triggers, one path: storage-channel death and medium_changed(absent) both drive kill-the-filesystem-process + retire-its-mounts; return (device re-report or medium_changed(present)) drives re-probe → respawn → remount at the identity's prefix. The logger gains resume patience (retry flushes on the fat cadence, never abandon) and the ring-wrap gap marker. QEMU cases:

  • volume-replug: yank the boot stick mid-run, replug, assert remount at the same prefixes and the logger appending to the SAME boot-stamp tree with the gap marked. Discrimination: against pre-V4, fat wedges (mounted forever) and no remount happens.
  • volume-two-partitions: a two-partition FAT image → two volumes, two filesystem processes, two mounts from one stick; yank once, both die; return once, both remount. Proves multi-volume and the reap breadth. (New image fixture beside make-fat-image.py.)
  • volume-clone-policy: two sticks with identical FAT serials — first keeps the mapped name, second mounts suffixed, loudly logged.
  • The lifecycle conformance drill, parameterized by filesystem: mount, serve, yank mid-write, verify honest loss (dirty flag set, gap said), replug, remount. FAT is implementation #1; the drill is the definition of "danos supports filesystem X".

V5 — close-out

Full suite green; the architecture doc's (planned) markers flip to built; fs_unmount ownership, ranges, presence, identity, and the volume manager lose their future-tense; memory updated; Ryzen bench note: pull the stick, watch the log gap get marked, plug it back anywhere.

Not in this track (recorded so absence is deliberate): GPT parsing beyond the identity read (row exists in the prober's ladder; full GPT when a GPT medium matters), AHCI/NVMe drivers (NVMe gated on the shm-ring data plane), formatting/entropy, per-process namespaces.