# Finishing the storage stack: the S1–S5 plan *2026-08-09. Continues the volume-manager track (V0–V5, on main) from "one FAT volume" to "any filesystem, any number of volumes, identified by content, remounting where they belong, surviving a driver crash." Executes the settled design in [storage-architecture.md](file-system-development/storage-architecture.md) and [storage-design-rationale.md](file-system-development/storage-design-rationale.md). Track discipline as always: work on main; one commit per coherent step with `git commit -F` (no `-m`, no co-author trailer); every new test shown to FAIL against the old behavior; one QEMU suite at a time; CSV is configuration read by the volume manager (the policy), never itself policy; bounds discipline (`tools/check-bounds.py` gate); adversarial boundary review at each phase.* ## Where V0–V4 left it The volume manager probes ONE storage device, parses its MBR (rung-4 identity = `(diskSignature<<8)|index`), spawns ONE FAT service confined to that partition's badge-scoped block range, supervises it, and unmounts it when the device is pulled. `partition.firstVolume` returns the FIRST partition; `openAnyStorage` adopts the FIRST device; `var volume: ?Volume` and `volume_id = 1` are singular; `filesystem_binary` and fat's mount prefixes are hardcoded; `medium_changed` is published by the driver but consumed by no one; remount-on-replug is bench-pending. ## The five phases and how they depend ``` S1 identity ladder ──► S2 mount map ──► S3 multi-volume ──► S4 exFAT (needs S2+S3) S5 removal robustness ── independent; single-volume ── may land any time ``` - **S1** grows the identity read off the medium (GPT GUID, FAT serial+label). It comes FIRST because a volume's mount name is now its identity (below), and the friendly form of that name is the FAT label / GPT name S1 parses. - **S2** moves the last policy out of hardcode into `volumes.csv` + `filesystems.csv`, and names each volume by its identity **id** (a GUID/key), keeping the label as separate, queryable display metadata — no port-name. - **S3** generalizes to N volumes across N devices. - **S4** adds exFAT — a COMPLETE second engine that proves the V1 harness extraction. Needs S2 (to route by signature) and S3 (to run a second volume). - **S5** closes the removal-lifecycle gaps. Independent of the rest; single-volume. Recommended build order is S1 → S2 → S3 → S4 → S5. S5 may be pulled earlier. --- ## S1 — the identity ladder **Goal.** Grow `system/services/volume-manager/partition.zig` from the single rung-4 identity into a ladder that reads the richest available content identity: GPT partition GUID (rung 1, 128-bit), FAT volume serial + label (rung 3), MBR signature + index (rung 4, kept), bare-FAT (kept, enriched to its serial). The `u64` identity becomes a small tagged struct `Identity{ rung, key: u128, label, has_label }` — `key` is the **id** (the path handle), `label` is the **display name** (the GPT 36-char partition name, or the FAT volume label), a separate field per the id/name split S2 relies on. Because GPT metadata is at LBA 1 and the entry array beyond it, and the FAT serial is in each partition's VBR, `firstVolume` stops taking one preloaded block-0 slice and takes a `SectorReader` (context + read-one-sector fn, mirroring the engine's `BlockDevice` vtable) — host-testable against a RAM-disk reader exactly as the four existing `partition.zig` tests are. **Key touchpoints.** `partition.zig` (the `Rung`/`Identity`/`SectorReader` types; `gptFirstVolume`; `fatIdentity`; `firstVolume` control flow — GPT authoritative, else MBR walk skipping type-0xEE, else bare-FAT, each preferring the FAT serial over the disk signature); `volume-manager.zig` (`Volume.identity` type; a `ProbeReader` over the existing 512-byte bounce; the probe log prints `identity.key`); `build.zig` (add `b.dependency("volume-manager", .{})` to the package-test aggregation loop so the host fixtures run under root `zig build test`). Declared bound `gpt_entry_scan_maximum = 128` with the full bounds block; `sector_bytes`/`fat_label_bytes` named consts. **Steps (commits).** (1) The SectorReader/Identity flag-day — pure refactor, no new behavior, all existing tests green. (2) GPT parsing (rung 1) — protective-MBR + `EFI PART` signature + header CRC-32 + per-entry overflow-safe range validation (the confinement-safety invariant the driver's clamp rests on, extended to GPT). (3) FAT serial + label (rung 3), preferred over the disk signature; `Identity.eql`. (4) Test wiring, `check-bounds.py`, full suite, docs, memory. **Discrimination.** Host: a GPT disk yields `rung==.gpt_guid` + the exact GUID key (old code walks the 0xEE protective entry as an ordinary partition); a GPT entry past the device is skipped, an out-of-device-only GPT returns null (the security-boundary guard); an invalid GPT header is not a volume; a bare FAT reports its real serial `0x12345678` not the rung-4 pseudo-signature; an MBR+FAT partition prefers the serial over the disk signature. On-image: the `volume-probe` QEMU regex tightens to `volume 0x0*12345678` — the boot image's real FAT32 serial reaches the running log. **Top risks.** Adversarial GPT input from untrusted media (huge entry counts, bogus offsets, overflowing ranges) — mitigated by header CRC + size bounds + `gpt_entry_scan_maximum` + per-entry overflow-safe validation under boundary review. The `std.hash.crc` symbol in Zig 0.16 is unverified (fallback: a ~15-line reflected CRC-32, poly `0xEDB88320`, used by both parser and fixtures so they never drift onto magic constants). --- ## S2 — the mount map: volumes.csv + filesystems.csv **Goal.** Move the last two pieces of storage policy out of hardcode into configuration read by the volume manager, and split a volume's **id** from its **label** (the database model: the id is the real, stable, unique key software uses; the label is a mutable display name). `filesystems.csv` (content signature → filesystem binary) so the VM picks the binary from the probed signature; `volumes.csv` (id → mount prefix, danos's fstab) as the explicit override for a volume the user wants at a fixed path. **The mount path IS the identity id, never a label or a port/role name.** A volume mounts at `/volumes/` — the GPT partition GUID for a GPT volume, a `fat-` / `mbr--` form otherwise (exact rendering is an impl detail; the point is a stable, unique, content-derived string). Because the path is the id and never the label, two distinct volumes that happen to share a label (`Backup`, `UNTITLED`, unlabeled) get distinct paths automatically and never collide; only genuinely identical *ids* (dd-cloned media) hit the rationale's duplicate-identity policy (first mounts, second suffixed and logged). **The label is display metadata, exposed by a protocol query, not the path.** S1 reads it off the medium (FAT volume label, GPT partition name); a volume-manager `volumes` verb returns `{ id, mount_path, label }` per volume so a future shell/UI can show the friendly name — the id↔name split, like a table's primary key vs its display column. (`volumes.csv` may optionally carry a chosen label override alongside the path override, both keyed on id.) There is no `/volumes/usb` and no `/volumes/boot`; the boot volume is detected by content (it installs the `/system/configuration` + `/system/logs` rewrites) but is named by its id like any other. Parsed with `library/csv` exactly as `device-registry` parses `devices.csv`. The VM hands the binary + volume-id + mount specs to fat at spawn (the argv channel V3b already uses for the volume id); fat retires `fat_mounts` and reads mounts from argv[2..]. **Key touchpoints.** New pure-logic modules `filesystem-map.zig` (parse + `match(signature)`) and `volume-map.zig` (parse + `mountsFor(identity)` + `derivedAnonymous`), mirroring `device-registry`, host-tested; `partition.zig` `Volume` gains a `signature`; `volume-manager.zig` loads both tables in `initialise`, resolves binary + mounts, spawns the chosen binary with the mount argv; `fat.zig` deletes `fat_mounts`, parses argv[2..] into a bounded `MountSpec` array; new `system/configuration/filesystems.csv` + `volumes.csv`; `build.zig` bundles them; `make-fat-image.py` writes a real 4-byte MBR disk signature so the boot volume's identity is a legible non-zero key. **Steps (commits).** (1) partition emits a signature. (2) `filesystem-map` parser + host tests. (3) `volume-map` parser (id→prefix override) + the id-path deriver + the `volumes` label-query verb + host tests. (4) Ship the tables + VM integration **behind fat's still-hardcoded mounts** (behavior-preserving — parsers proven before consumption flips; full suite green). (5) fat consumes argv mounts; the VM mounts the boot volume at its id-path (`/volumes/fat-12345678`, from its serial); **migrate the fixtures + QEMU regexes off `/volumes/usb`** (fat-test, badge-scope-test, vfs-test, and the four cases) in the same commit; the `volume-identity-name` QEMU case (shown failing against HEAD~1); flip the docs' pending markers. **Discrimination.** QEMU `volume-identity-name`: the boot volume mounts at its id-path (`/volumes/fat-12345678`, from its serial), a line the old hardcoded `/volumes/usb` never emits; the `volumes` query returns that id paired with the label `DANOS`. Host: two volumes sharing a label but not an id get distinct id-paths (the collision the label-as-name approach could not resolve stably); a `volumes.csv` override sends a mapped id to its chosen prefix; `match(.fat)` returns the configured binary; fat's argv parser makes installed mounts a function of argv. **Top risks.** Step 5's blast radius — a parser/argv bug, OR the `/volumes/usb`→ `/volumes/DANOS` migration missing a fixture/regex, breaks every fat-dependent case at once; mitigated by landing the VM half behind fat's hardcoded mounts first (step 4) and making the name migration one atomic, complete sweep. The argv blob is 256 bytes (process.zig) — cap emitted mounts and refuse+log on overflow. `/system/logs` is now a `volumes.csv` concern: dropping the boot identity's rewrite rows silently stops log persistence — ship them by default and document the boot-identity contract in the CSV header. --- ## S3 — multi-volume **Goal.** Generalize from `var volume: ?Volume` / `volume_id = 1` / first-device / first-partition to a bounded table of volumes across a bounded table of devices. `partition.allVolumes` returns ALL partitions; the VM adopts EVERY mass-storage provider, probes each device's table, and for each partition spawns one FAT confined to that partition's range (the per-sender clamp is already built), with a distinct `/volumes/` and its OWN backoff/crash-loop state. Removal is per-device. The **boot volume is identified by content** — the FAT process installs the `/system/configuration` + `/system/logs` rewrites only when its own volume resolves `/system/configuration` — so it works as the 2nd partition of the 2nd device just as the 1st of the 1st. (This clarifies the initrd relationship: the kernel already serves `/system/configuration` + binaries read-only from the initrd, which is what lets danos boot with NO volume mounted; a mounted boot volume only adds writable, persistent `/system/configuration` + `/system/logs` that shadow the initrd via longest-prefix match.) **Key touchpoints.** `partition.zig` `firstVolume` → `allVolumes(block0, device_blocks, out) usize` (per-entry overflow-safe skip preserved); `volume-manager.zig` the core refactor — `Volume` absorbs the file-global supervision state as per-volume fields, a `StorageDevice` table owns each adopted device's channel once, `volumes[maximum_volumes]` replaces the singleton, a monotonic `next_volume_id`, `openAllStorage`/`adoptAndProbe`, `gatherPresentStorage` + per-device reconcile in `pollTick`, `onHello`/ `onNotification` keyed across the table; `fat.zig` content-conditional boot mounts; `system/kernel/vfs.zig` raise `maximum_mounts` (8 → 16) with a refreshed bounds annotation; `make-fat-image.py` a partition-table mode; `build/images.zig` the test disk artifacts. **Steps (commits).** (1) `partition.allVolumes` + two-partition host test. (2) Tables, behavior-preserving (still one device / one volume). (3) Multi-device + multi-partition. (4) Per-volume mount naming via argv — each volume by its identity (label→hex, from S2), one per volume, no port-name. (5) fat content-conditional boot mounts. (6) Raise `maximum_mounts`. (7) Partitioned-image tool. (8) `two-volume` QEMU case. (9) `boot-2nd-partition` case. (10) Adversarial review, full suite, docs, memory. **Discrimination.** Host: an MBR with two partitions yields two volumes with distinct identities (old `firstVolume` returns one). QEMU `two-volume`: a two-partition second device yields two mount lines at two base_lbas (old `openAnyStorage` adopts only the first device). `boot-2nd-partition`: the boot volume works as partition 2 (old code confines fat to partition 1, whose `/system` rewrite backs empty space). Per-volume supervision: killing one volume's FAT restarts only that one (old module-scope supervision can't attribute an exit to one of two). **Top risks.** OVMF booting an MBR ESP on partition 2 may be flaky in CI — fallback to a content-detection-ordering assertion + bench-verified boot (the track's existing precedent). N-client range reclamation in usb-storage (`maximum_ranges=64`) must reclaim each of N confined pids' ranges — the V2 mechanism, previously exercised with one live client. Duplicate boot volumes: S3 supports exactly one and must log loudly if a second also resolves the boot markers (arbitration deferred to S4). --- ## S4 — exFAT: the second engine **Goal.** A working exFAT filesystem as `system/services/exfat` that is nothing but an engine + a `main`, reusing `library/kernel/file-system-harness.zig`'s `Server(Engine)` wholesale — the reuse claim the architecture makes, now proven. The harness already owns vfs serving, the badge-scoped open-node table, create-on-open/O_TRUNC, mount registration, the exit sweep, bring-up retry, per-turn time stamping, and durable-on-close. S4 writes only the exFAT-specific bits — but **in full**: complete read AND write, directories, rename, and the real on-disk up-case table for correct case-folding. Not a read-first, minimal- write, or ASCII-only subset. The only limit that survives is the vfs protocol's u32 file-offset surface (a 4 GiB addressable-size cap that applies to FAT too), which is a separate vfs-protocol change, not an exFAT shortcut. **Key touchpoints.** New `system/services/exfat/on-disk.zig` (the Main Boot Sector VBR + the five 32-byte directory-entry types as `align(1)` extern structs; `geometryOf` accepting only `"EXFAT "` + `0xAA55`; `setChecksum`, `nameHash`, and case-folding driven by the volume's **on-disk up-case table**); `engine.zig` (`FileSystem` behind the identical `BlockDevice` vtable with fat's exact method set; **allocation-bitmap** cluster authority — the deepest departure from FAT; read honoring `no_fat_chain` contiguous vs FAT-follow; File+Stream+FileName set assembly with recomputed set checksum); `exfat.zig` (the thin service, a near-clone of fat.zig); build wiring + `service("exfat")`; `tools/make-exfat-image.py` (pure stdlib, correct boot checksum, up-case table — **no committed .img**); the `filesystems.csv` EXFAT row (S2) + the `"EXFAT "` recognizer; a `exfat-test` fixture cloned from fat-test. **Steps (commits).** (1) on-disk.zig byte layout. (2) engine read path. (3) engine write path (bitmap allocate/free, real 32-bit FAT chain with `no_fat_chain=0`, set-checksum recompute). (4) service + build wiring. (5) `make-exfat-image.py` + image assembly. (6) Routing: `filesystems.csv` + recognizer. (7) `exfat-test` fixture + **cross-engine discrimination** host test. (8) In-VM lifecycle drill (second removable device; mount/mutations/removal). (9) Bounds, docs, adversarial review, memory. **Discrimination.** The named one: `fat.mount(exfat_img) == null` (fat reads bytes-per-sector at VBR offset 11 = exFAT's MustBeZero = 0 → reject) AND `exfat.mount(fat_img) == null`, each mounting its own as a control. Host: read across a cluster boundary on both a contiguous and a fragmented file; write across >1 cluster setting the bitmap bits (not the FAT) and a validating set checksum. QEMU: `exfat: mounted /volumes/exfat` + `exfat-test: ok`; `exfat-removal` yanks the exFAT stick mid-write while the FAT boot volume keeps serving. **Top risks.** Allocation authority is the bitmap, not the FAT — allocating without setting the bit silently corrupts free space (highest-attention area). A directory-entry SET can straddle sector/cluster boundaries — scanDirectory, set-checksum, and updateStreamEntry must handle multi-sector sets. vfs offsets are u32 while exFAT DataLength is u64 — clamp and document (as fat does). The in-VM drill needs S3 (a non-boot exFAT volume beside the FAT boot volume); if S4 landed before S3 the discrimination would rest on host tests until multi-volume exists. --- ## S5 — removal robustness **Goal.** Close the three known gaps so every removal trigger is exercised end-to-end. (1) **Consume `medium_changed`** — the VM subscribes to the driver's already-published event so the "device stays, medium leaves" case (a card reader, an ejected removable) runs the same kill-retire-remount path as a pulled stick, closing the second of the "two triggers, one lifecycle" the architecture specifies. (2) **Storage-driver-crash rebuild** — a driver that dies while its device stays present is detected and the volume subtree rebuilt on the restarted driver's fresh channel, instead of leaving fat wedged on a dead channel (the V4 review's open edge). (3) **QEMU-verified remount-on-replug** — the device-return half is proven, not merely asserted-unmount. Single-volume; independent of S1–S4. **Key touchpoints.** `library/kernel/service.zig` an additive, behavior-neutral `on_buffered_message` callback so a buffered-message wake forwards its payload (no existing service sets it); `volume-manager.zig` subscribe on `bringUpVolume` success, `onMediumEvent` with change-count dedupe running a medium-teardown (with `encodeUnsubscribe` before close so the driver's 8-slot table doesn't leak), plus `channelAlive()` (a `geometry()` liveness probe) + `rebuildVolume()` used in `pollTick` and the child-exit path; `fat.zig` re-probe geometry on I/O failure and exit on channel death (device NAK keeps serving); `device-manager.zig` a `test-storage-restart` mode (mirroring `test-scanout-restart`) to kill usb-storage once, post-mount, as the discrimination trigger. **Steps (commits).** (1) **Cheap decisive experiments first** (no commits): QMP- eject the boot medium and confirm `usb-storage: medium absent` fires under QEMU (the whole item-1 chain depends on it); and test whether a boot-controller `device_add` is re-presented (settles whether item 3 extends `volume-removal` or needs a second controller as H1 does). (2) Harness `on_buffered_message` (behavior-neutral). (3) VM consumes `medium_changed`. (4) `volume-medium-change` case (fails pre-step-3). (5) VM driver-crash rebuild. (6) fat observes dead channel and exits. (7) `volume-driver-restart` trigger + case. (8) `volume-replug` (second controller if needed). (9) Docs + the real-hardware bench protocol. (10) Full suite + memory. **Discrimination.** `volume-medium-change`: an eject with the device left in the tree unmounts (old VM never subscribes → the event goes to no one → mount persists). `volume-driver-restart`: killing usb-storage post-mount while its child stays present triggers a rebuild and a SECOND mount + post-kill read (old `pollTick` only checks `isDevicePresent`, still true, and restarts fat against the stale channel → wedge/crash-loop). `volume-replug`: a device return on a second controller drives a remount (the existing case only ever sees the unmount half). **Top risks.** QEMU medium-eject must make TEST UNIT READY report not-ready — step 1(a) validates this before any code. The op-16 overlap (`medium_changed` == `hello` by number) is safe only because async events arrive as `isMessage` notifications and never reach `Serve.dispatch` — the intercept must run in the notification branch and never catch a synchronous hello. fat can't today distinguish EPEER from a device NAK (`CallError` swallows the errno) — the plan uses a geometry re-probe as the liveness oracle, which is correct but indirect. --- ## Decisions (settled) Both flagged decisions are settled: 1. **A volume's path is its id; the label is display metadata (S1/S2).** The mount path is the identity id — the GPT GUID, else a `fat-` / `mbr--` form — a stable, unique, content-derived handle software uses. The label (FAT volume label / GPT partition name) is a mutable display name, NOT in the path; a volume-manager `volumes` verb returns `{ id, mount_path, label }` so a UI can show the friendly name (the database id/name split). Same-label-different-id volumes therefore never collide; only identical ids (dd-clones) hit first-wins-and-log. `volumes.csv` overrides the path (and optionally the label) for a chosen volume, keyed on id. Makes S1 precede S2 and folds the `/volumes/usb` → id-path fixture + regex migration into S2 step 5. 2. **exFAT is implemented in full (S4).** A complete exFAT: full read and write, directories, rename, and the on-disk up-case table for correct case-folding — not a read-first or ASCII-only subset. The one remaining limit is the vfs protocol's u32 file-offset surface, which caps addressable file size at 4 GiB for ALL filesystems (FAT included); widening it to u64 is a separate vfs- protocol change, flagged but out of the exFAT engine's scope. The ~26 smaller design-time questions are settled with the recommended default in the phase text (defer GPT entry-array CRC to correctness-only; GUID key = little-endian u128 pinned now; share the DOS date-time helper into a library module both engines import; a second removable usb-storage device for the exFAT drill; VM-poll `channelAlive()` as the load-bearing crash-detection guarantee).