The phased plan taking the volume-manager track from one FAT volume to any filesystem, N volumes, content identity, remount, and driver-crash survival: S1 identity ladder (GPT GUID + FAT serial), S2 volumes.csv/filesystems.csv mount map, S3 multi-volume, S4 exFAT (the second engine proving the harness reuse), S5 removal robustness (medium_changed consumption + driver-crash rebuild + QEMU-verified remount). Code-grounded: real functions, commit-granular steps, discrimination tests shown to fail against today's behavior, per-phase risks. Design decisions taken with recommended defaults; two forks flagged for review (the /volumes/usb transitional naming and the exFAT write scope). Not started.
330 lines
19 KiB
Markdown
330 lines
19 KiB
Markdown
# Finishing the storage stack: the S1–S5 plan
|
||
|
||
*2026-08-09. Continues the volume-manager track (V0–V5, on main) from "one FAT
|
||
volume" to "any filesystem, any number of volumes, identified by content,
|
||
remounting where they belong, surviving a driver crash." Executes the settled
|
||
design in
|
||
[storage-architecture.md](file-system-development/storage-architecture.md) and
|
||
[storage-design-rationale.md](file-system-development/storage-design-rationale.md).
|
||
Track discipline as always: work on main; one commit per coherent step with
|
||
`git commit -F` (no `-m`, no co-author trailer); every new test shown to FAIL
|
||
against the old behavior; one QEMU suite at a time; CSV is configuration read by
|
||
the volume manager (the policy), never itself policy; bounds discipline
|
||
(`tools/check-bounds.py` gate); adversarial boundary review at each phase.*
|
||
|
||
## Where V0–V4 left it
|
||
|
||
The volume manager probes ONE storage device, parses its MBR (rung-4 identity =
|
||
`(diskSignature<<8)|index`), spawns ONE FAT service confined to that partition's
|
||
badge-scoped block range, supervises it, and unmounts it when the device is
|
||
pulled. `partition.firstVolume` returns the FIRST partition; `openAnyStorage`
|
||
adopts the FIRST device; `var volume: ?Volume` and `volume_id = 1` are singular;
|
||
`filesystem_binary` and fat's mount prefixes are hardcoded; `medium_changed` is
|
||
published by the driver but consumed by no one; remount-on-replug is
|
||
bench-pending.
|
||
|
||
## The five phases and how they depend
|
||
|
||
```
|
||
S1 identity ladder ─────┐
|
||
├──► S2 mount map ──► S3 multi-volume ──► S4 exFAT
|
||
(S2 keys on rung-4, so │ (needs S2+S3)
|
||
is independent of S1) │
|
||
└──► S5 removal robustness (independent; single-volume)
|
||
```
|
||
|
||
- **S1** grows the identity read off the medium (GPT GUID, FAT serial+label).
|
||
- **S2** moves the last policy out of hardcode into `volumes.csv` +
|
||
`filesystems.csv`, keyed on today's rung-4 identity — so it does **not** wait
|
||
on S1; S1 only enriches the identity values the same map consumes.
|
||
- **S3** generalizes to N volumes across N devices.
|
||
- **S4** adds exFAT — the second engine that proves the V1 harness extraction.
|
||
Needs S2 (to route by signature) and S3 (to run a second volume in the VM).
|
||
- **S5** closes the removal-lifecycle gaps. Independent of the rest; single-volume.
|
||
|
||
Recommended build order is S1 → S2 → S3 → S4 → S5. S5 may be pulled earlier if
|
||
removal robustness matters more than multi-volume; S1 and S2 may swap.
|
||
|
||
---
|
||
|
||
## S1 — the identity ladder
|
||
|
||
**Goal.** Grow `system/services/volume-manager/partition.zig` from the single
|
||
rung-4 identity into a ladder that reads the richest available content identity:
|
||
GPT partition GUID (rung 1, 128-bit), FAT volume serial + label (rung 3), MBR
|
||
signature + index (rung 4, kept), bare-FAT (kept, enriched to its serial). The
|
||
`u64` identity becomes a small tagged struct `Identity{ rung, key: u128, label,
|
||
has_label }`. Because GPT metadata is at LBA 1 and the entry array beyond it, and
|
||
the FAT serial is in each partition's VBR, `firstVolume` stops taking one
|
||
preloaded block-0 slice and takes a `SectorReader` (context + read-one-sector
|
||
fn, mirroring the engine's `BlockDevice` vtable) — host-testable against a
|
||
RAM-disk reader exactly as the four existing `partition.zig` tests are.
|
||
|
||
**Key touchpoints.** `partition.zig` (the `Rung`/`Identity`/`SectorReader` types;
|
||
`gptFirstVolume`; `fatIdentity`; `firstVolume` control flow — GPT authoritative,
|
||
else MBR walk skipping type-0xEE, else bare-FAT, each preferring the FAT serial
|
||
over the disk signature); `volume-manager.zig` (`Volume.identity` type; a
|
||
`ProbeReader` over the existing 512-byte bounce; the probe log prints
|
||
`identity.key`); `build.zig` (add `b.dependency("volume-manager", .{})` to the
|
||
package-test aggregation loop so the host fixtures run under root `zig build
|
||
test`). Declared bound `gpt_entry_scan_maximum = 128` with the full bounds block;
|
||
`sector_bytes`/`fat_label_bytes` named consts.
|
||
|
||
**Steps (commits).** (1) The SectorReader/Identity flag-day — pure refactor, no
|
||
new behavior, all existing tests green. (2) GPT parsing (rung 1) — protective-MBR
|
||
+ `EFI PART` signature + header CRC-32 + per-entry overflow-safe range
|
||
validation (the confinement-safety invariant the driver's clamp rests on,
|
||
extended to GPT). (3) FAT serial + label (rung 3), preferred over the disk
|
||
signature; `Identity.eql`. (4) Test wiring, `check-bounds.py`, full suite, docs,
|
||
memory.
|
||
|
||
**Discrimination.** Host: a GPT disk yields `rung==.gpt_guid` + the exact GUID
|
||
key (old code walks the 0xEE protective entry as an ordinary partition); a GPT
|
||
entry past the device is skipped, an out-of-device-only GPT returns null (the
|
||
security-boundary guard); an invalid GPT header is not a volume; a bare FAT
|
||
reports its real serial `0x12345678` not the rung-4 pseudo-signature; an
|
||
MBR+FAT partition prefers the serial over the disk signature. On-image: the
|
||
`volume-probe` QEMU regex tightens to `volume 0x0*12345678` — the boot image's
|
||
real FAT32 serial reaches the running log.
|
||
|
||
**Top risks.** Adversarial GPT input from untrusted media (huge entry counts,
|
||
bogus offsets, overflowing ranges) — mitigated by header CRC + size bounds +
|
||
`gpt_entry_scan_maximum` + per-entry overflow-safe validation under boundary
|
||
review. The `std.hash.crc` symbol in Zig 0.16 is unverified (fallback: a ~15-line
|
||
reflected CRC-32, poly `0xEDB88320`, used by both parser and fixtures so they
|
||
never drift onto magic constants).
|
||
|
||
---
|
||
|
||
## S2 — the mount map: volumes.csv + filesystems.csv
|
||
|
||
**Goal.** Move the last two pieces of storage policy out of hardcode into
|
||
configuration read by the volume manager. `filesystems.csv` (content signature →
|
||
filesystem binary) so the VM picks the binary from the probed signature;
|
||
`volumes.csv` (identity → mount prefix(es), danos's fstab) so the VM decides
|
||
mount placement, with a `/volumes/vol-<hex>` anonymous fallback. Parsed with
|
||
`library/csv` exactly as `device-registry` parses `devices.csv`. The VM hands the
|
||
binary + volume-id + mount specs to fat at spawn (the argv channel V3b already
|
||
uses for the volume id); fat retires `fat_mounts` and reads mounts from
|
||
argv[2..]. Keyed on rung-4 identity, so independent of S1 and forward-compatible.
|
||
|
||
**Key touchpoints.** New pure-logic modules `filesystem-map.zig` (parse +
|
||
`match(signature)`) and `volume-map.zig` (parse + `mountsFor(identity)` +
|
||
`derivedAnonymous`), mirroring `device-registry`, host-tested; `partition.zig`
|
||
`Volume` gains a `signature`; `volume-manager.zig` loads both tables in
|
||
`initialise`, resolves binary + mounts, spawns the chosen binary with the mount
|
||
argv; `fat.zig` deletes `fat_mounts`, parses argv[2..] into a bounded
|
||
`MountSpec` array; new `system/configuration/filesystems.csv` +
|
||
`volumes.csv`; `build.zig` bundles them; `make-fat-image.py` writes a real 4-byte
|
||
MBR disk signature so the boot volume's identity is a legible non-zero key.
|
||
|
||
**Steps (commits).** (1) partition emits a signature. (2) `filesystem-map`
|
||
parser + host tests. (3) `volume-map` parser + host tests. (4) Ship the tables +
|
||
VM integration **behind fat's still-hardcoded mounts** (behavior-preserving —
|
||
the parsers are proven before consumption flips; full suite green). (5) fat
|
||
consumes argv mounts; add the `/volumes/boot` row; the `volume-mapped` QEMU case
|
||
(shown failing against HEAD~1); flip the docs' `volumes.csv`/`filesystems.csv`
|
||
pending markers.
|
||
|
||
**Discrimination.** QEMU `volume-mapped`: the boot volume, mapped in
|
||
`volumes.csv` to `/volumes/boot`, makes fat log `mounted /volumes/boot` — a line
|
||
the old hardcoded `fat_mounts` never emits. Host: `mountsFor(mapped)` returns the
|
||
CSV prefix, `mountsFor(unmapped)` derives `/volumes/vol-<hex>`; `match(.fat)`
|
||
returns the configured binary; fat's argv parser makes installed mounts a
|
||
function of argv.
|
||
|
||
**Top risks.** Step 5's blast radius — a parser/argv bug breaks every
|
||
fat-dependent case at once; mitigated by landing the VM half behind fat's
|
||
hardcoded mounts first (step 4). The argv blob is 256 bytes (process.zig) — cap
|
||
emitted mounts and refuse+log on overflow. `/system/logs` is now a
|
||
`volumes.csv` row: dropping it silently stops log persistence — ship it by
|
||
default and document the boot-identity contract in the CSV header.
|
||
|
||
---
|
||
|
||
## S3 — multi-volume
|
||
|
||
**Goal.** Generalize from `var volume: ?Volume` / `volume_id = 1` / first-device
|
||
/ first-partition to a bounded table of volumes across a bounded table of
|
||
devices. `partition.allVolumes` returns ALL partitions; the VM adopts EVERY
|
||
mass-storage provider, probes each device's table, and for each partition spawns
|
||
one FAT confined to that partition's range (the per-sender clamp is already
|
||
built), with a distinct `/volumes/<name>` and its OWN backoff/crash-loop state.
|
||
Removal is per-device. The **boot volume is identified by content** — the FAT
|
||
process installs the `/system/configuration` + `/system/logs` rewrites only when
|
||
its own volume resolves `/system/configuration` — so it works as the 2nd
|
||
partition of the 2nd device just as the 1st of the 1st. (This clarifies the
|
||
initrd relationship: the kernel already serves `/system/configuration` +
|
||
binaries read-only from the initrd, which is what lets danos boot with NO volume
|
||
mounted; a mounted boot volume only adds writable, persistent
|
||
`/system/configuration` + `/system/logs` that shadow the initrd via
|
||
longest-prefix match.)
|
||
|
||
**Key touchpoints.** `partition.zig` `firstVolume` → `allVolumes(block0,
|
||
device_blocks, out) usize` (per-entry overflow-safe skip preserved);
|
||
`volume-manager.zig` the core refactor — `Volume` absorbs the file-global
|
||
supervision state as per-volume fields, a `StorageDevice` table owns each
|
||
adopted device's channel once, `volumes[maximum_volumes]` replaces the singleton,
|
||
a monotonic `next_volume_id`, `openAllStorage`/`adoptAndProbe`,
|
||
`gatherPresentStorage` + per-device reconcile in `pollTick`, `onHello`/
|
||
`onNotification` keyed across the table; `fat.zig` content-conditional boot
|
||
mounts; `system/kernel/vfs.zig` raise `maximum_mounts` (8 → 16) with a refreshed
|
||
bounds annotation; `make-fat-image.py` a partition-table mode; `build/images.zig`
|
||
the test disk artifacts.
|
||
|
||
**Steps (commits).** (1) `partition.allVolumes` + two-partition host test. (2)
|
||
Tables, behavior-preserving (still one device / one volume). (3) Multi-device +
|
||
multi-partition. (4) Per-volume mount naming via argv (first = `/volumes/usb`,
|
||
rest `/volumes/usb<n>` — bridge until `volumes.csv`). (5) fat content-conditional
|
||
boot mounts. (6) Raise `maximum_mounts`. (7) Partitioned-image tool. (8)
|
||
`two-volume` QEMU case. (9) `boot-2nd-partition` case. (10) Adversarial review,
|
||
full suite, docs, memory.
|
||
|
||
**Discrimination.** Host: an MBR with two partitions yields two volumes with
|
||
distinct identities (old `firstVolume` returns one). QEMU `two-volume`: a
|
||
two-partition second device yields two mount lines at two base_lbas (old
|
||
`openAnyStorage` adopts only the first device). `boot-2nd-partition`: the boot
|
||
volume works as partition 2 (old code confines fat to partition 1, whose
|
||
`/system` rewrite backs empty space). Per-volume supervision: killing one
|
||
volume's FAT restarts only that one (old module-scope supervision can't
|
||
attribute an exit to one of two).
|
||
|
||
**Top risks.** OVMF booting an MBR ESP on partition 2 may be flaky in CI —
|
||
fallback to a content-detection-ordering assertion + bench-verified boot (the
|
||
track's existing precedent). N-client range reclamation in usb-storage
|
||
(`maximum_ranges=64`) must reclaim each of N confined pids' ranges — the V2
|
||
mechanism, previously exercised with one live client. Duplicate boot volumes:
|
||
S3 supports exactly one and must log loudly if a second also resolves the boot
|
||
markers (arbitration deferred to S4).
|
||
|
||
---
|
||
|
||
## S4 — exFAT: the second engine
|
||
|
||
**Goal.** A working exFAT filesystem as `system/services/exfat` that is nothing
|
||
but an engine + a `main`, reusing `library/kernel/file-system-harness.zig`'s
|
||
`Server(Engine)` wholesale — the reuse claim the architecture makes, now proven.
|
||
The harness already owns vfs serving, the badge-scoped open-node table,
|
||
create-on-open/O_TRUNC, mount registration, the exit sweep, bring-up retry,
|
||
per-turn time stamping, and durable-on-close. S4 writes only the exFAT-specific
|
||
bits.
|
||
|
||
**Key touchpoints.** New `system/services/exfat/on-disk.zig` (the Main Boot
|
||
Sector VBR + the five 32-byte directory-entry types as `align(1)` extern structs;
|
||
`geometryOf` accepting only `"EXFAT "` + `0xAA55`; `setChecksum`, `nameHash`,
|
||
`upcaseAscii`); `engine.zig` (`FileSystem` behind the identical `BlockDevice`
|
||
vtable with fat's exact method set; **allocation-bitmap** cluster authority — the
|
||
deepest departure from FAT; read honoring `no_fat_chain` contiguous vs FAT-follow;
|
||
File+Stream+FileName set assembly with recomputed set checksum); `exfat.zig` (the
|
||
thin service, a near-clone of fat.zig); build wiring + `service("exfat")`;
|
||
`tools/make-exfat-image.py` (pure stdlib, correct boot checksum, up-case table —
|
||
**no committed .img**); the `filesystems.csv` EXFAT row (S2) + the `"EXFAT "`
|
||
recognizer; a `exfat-test` fixture cloned from fat-test.
|
||
|
||
**Steps (commits).** (1) on-disk.zig byte layout. (2) engine read path. (3)
|
||
engine write path (bitmap allocate/free, real 32-bit FAT chain with
|
||
`no_fat_chain=0`, set-checksum recompute). (4) service + build wiring. (5)
|
||
`make-exfat-image.py` + image assembly. (6) Routing: `filesystems.csv` +
|
||
recognizer. (7) `exfat-test` fixture + **cross-engine discrimination** host test.
|
||
(8) In-VM lifecycle drill (second removable device; mount/mutations/removal). (9)
|
||
Bounds, docs, adversarial review, memory.
|
||
|
||
**Discrimination.** The named one: `fat.mount(exfat_img) == null` (fat reads
|
||
bytes-per-sector at VBR offset 11 = exFAT's MustBeZero = 0 → reject) AND
|
||
`exfat.mount(fat_img) == null`, each mounting its own as a control. Host: read
|
||
across a cluster boundary on both a contiguous and a fragmented file; write
|
||
across >1 cluster setting the bitmap bits (not the FAT) and a validating set
|
||
checksum. QEMU: `exfat: mounted /volumes/exfat` + `exfat-test: ok`;
|
||
`exfat-removal` yanks the exFAT stick mid-write while the FAT boot volume keeps
|
||
serving.
|
||
|
||
**Top risks.** Allocation authority is the bitmap, not the FAT — allocating
|
||
without setting the bit silently corrupts free space (highest-attention area).
|
||
A directory-entry SET can straddle sector/cluster boundaries — scanDirectory,
|
||
set-checksum, and updateStreamEntry must handle multi-sector sets. vfs offsets
|
||
are u32 while exFAT DataLength is u64 — clamp and document (as fat does). The
|
||
in-VM drill needs S3 (a non-boot exFAT volume beside the FAT boot volume); if S4
|
||
landed before S3 the discrimination would rest on host tests until multi-volume
|
||
exists.
|
||
|
||
---
|
||
|
||
## S5 — removal robustness
|
||
|
||
**Goal.** Close the three known gaps so every removal trigger is exercised
|
||
end-to-end. (1) **Consume `medium_changed`** — the VM subscribes to the driver's
|
||
already-published event so the "device stays, medium leaves" case (a card
|
||
reader, an ejected removable) runs the same kill-retire-remount path as a pulled
|
||
stick, closing the second of the "two triggers, one lifecycle" the architecture
|
||
specifies. (2) **Storage-driver-crash rebuild** — a driver that dies while its
|
||
device stays present is detected and the volume subtree rebuilt on the restarted
|
||
driver's fresh channel, instead of leaving fat wedged on a dead channel (the V4
|
||
review's open edge). (3) **QEMU-verified remount-on-replug** — the device-return
|
||
half is proven, not merely asserted-unmount. Single-volume; independent of S1–S4.
|
||
|
||
**Key touchpoints.** `library/kernel/service.zig` an additive, behavior-neutral
|
||
`on_buffered_message` callback so a buffered-message wake forwards its payload
|
||
(no existing service sets it); `volume-manager.zig` subscribe on `bringUpVolume`
|
||
success, `onMediumEvent` with change-count dedupe running a medium-teardown (with
|
||
`encodeUnsubscribe` before close so the driver's 8-slot table doesn't leak), plus
|
||
`channelAlive()` (a `geometry()` liveness probe) + `rebuildVolume()` used in
|
||
`pollTick` and the child-exit path; `fat.zig` re-probe geometry on I/O failure
|
||
and exit on channel death (device NAK keeps serving); `device-manager.zig` a
|
||
`test-storage-restart` mode (mirroring `test-scanout-restart`) to kill usb-storage
|
||
once, post-mount, as the discrimination trigger.
|
||
|
||
**Steps (commits).** (1) **Cheap decisive experiments first** (no commits): QMP-
|
||
eject the boot medium and confirm `usb-storage: medium absent` fires under QEMU
|
||
(the whole item-1 chain depends on it); and test whether a boot-controller
|
||
`device_add` is re-presented (settles whether item 3 extends `volume-removal` or
|
||
needs a second controller as H1 does). (2) Harness `on_buffered_message`
|
||
(behavior-neutral). (3) VM consumes `medium_changed`. (4) `volume-medium-change`
|
||
case (fails pre-step-3). (5) VM driver-crash rebuild. (6) fat observes dead
|
||
channel and exits. (7) `volume-driver-restart` trigger + case. (8) `volume-replug`
|
||
(second controller if needed). (9) Docs + the real-hardware bench protocol. (10)
|
||
Full suite + memory.
|
||
|
||
**Discrimination.** `volume-medium-change`: an eject with the device left in the
|
||
tree unmounts (old VM never subscribes → the event goes to no one → mount
|
||
persists). `volume-driver-restart`: killing usb-storage post-mount while its
|
||
child stays present triggers a rebuild and a SECOND mount + post-kill read (old
|
||
`pollTick` only checks `isDevicePresent`, still true, and restarts fat against
|
||
the stale channel → wedge/crash-loop). `volume-replug`: a device return on a
|
||
second controller drives a remount (the existing case only ever sees the unmount
|
||
half).
|
||
|
||
**Top risks.** QEMU medium-eject must make TEST UNIT READY report not-ready —
|
||
step 1(a) validates this before any code. The op-16 overlap (`medium_changed` ==
|
||
`hello` by number) is safe only because async events arrive as `isMessage`
|
||
notifications and never reach `Serve.dispatch` — the intercept must run in the
|
||
notification branch and never catch a synchronous hello. fat can't today
|
||
distinguish EPEER from a device NAK (`CallError` swallows the errno) — the plan
|
||
uses a geometry re-probe as the liveness oracle, which is correct but indirect.
|
||
|
||
---
|
||
|
||
## Decisions embedded in this plan (flag if you disagree)
|
||
|
||
Most of the ~28 design-time questions the phase design surfaced are settled with
|
||
the recommended default and noted in the phases above. Two are worth your eye:
|
||
|
||
1. **`/volumes/usb` naming (S2).** The plan **keeps `/volumes/usb`
|
||
transitionally** and adds the mapped `/volumes/boot` — zero fixture churn,
|
||
inherent discrimination. The design-purer alternative is to REPLACE
|
||
`/volumes/usb` (killing the port-name the rationale condemns) and migrate the
|
||
three runtime fixtures + four QEMU regexes in a coupled sweep. Recommendation:
|
||
keep transitionally now, drop `/volumes/usb` as a small follow-up once the map
|
||
is proven.
|
||
|
||
2. **exFAT write scope (S4).** The plan ships **full mutation parity**
|
||
(createFile + writeFile + truncate + removeFile + createDirectory + rename) so
|
||
`exfat-test` reuses fat-test's full matrix. The minimum viable is
|
||
create/write/truncate/remove only. Recommendation: full parity.
|
||
|
||
Everything else (defer GPT entry-array CRC to correctness-only; GUID key =
|
||
little-endian u128 pinned now; share the DOS date-time helper into a library
|
||
module both engines import; a second removable usb-storage device for the exFAT
|
||
drill rather than a boot-disk partition; VM-poll `channelAlive()` as the
|
||
load-bearing crash-detection guarantee with fat's exit as a promptness
|
||
optimization) is taken as the recommended default in the phase text.
|