Files
danos/docs/storage-stack-plan.md
T
Daniel Samson fca41b351e docs: the storage-stack completion plan (S1-S5)
The phased plan taking the volume-manager track from one FAT volume to any
filesystem, N volumes, content identity, remount, and driver-crash survival:
S1 identity ladder (GPT GUID + FAT serial), S2 volumes.csv/filesystems.csv mount
map, S3 multi-volume, S4 exFAT (the second engine proving the harness reuse),
S5 removal robustness (medium_changed consumption + driver-crash rebuild +
QEMU-verified remount). Code-grounded: real functions, commit-granular steps,
discrimination tests shown to fail against today's behavior, per-phase risks.
Design decisions taken with recommended defaults; two forks flagged for review
(the /volumes/usb transitional naming and the exFAT write scope). Not started.
2026-08-09 22:34:00 +01:00

330 lines
19 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Finishing the storage stack: the S1–S5 plan
*2026-08-09. Continues the volume-manager track (V0–V5, on main) from "one FAT
volume" to "any filesystem, any number of volumes, identified by content,
remounting where they belong, surviving a driver crash." Executes the settled
design in
[storage-architecture.md](file-system-development/storage-architecture.md) and
[storage-design-rationale.md](file-system-development/storage-design-rationale.md).
Track discipline as always: work on main; one commit per coherent step with
`git commit -F` (no `-m`, no co-author trailer); every new test shown to FAIL
against the old behavior; one QEMU suite at a time; CSV is configuration read by
the volume manager (the policy), never itself policy; bounds discipline
(`tools/check-bounds.py` gate); adversarial boundary review at each phase.*
## Where V0–V4 left it
The volume manager probes ONE storage device, parses its MBR (rung-4 identity =
`(diskSignature<<8)|index`), spawns ONE FAT service confined to that partition's
badge-scoped block range, supervises it, and unmounts it when the device is
pulled. `partition.firstVolume` returns the FIRST partition; `openAnyStorage`
adopts the FIRST device; `var volume: ?Volume` and `volume_id = 1` are singular;
`filesystem_binary` and fat's mount prefixes are hardcoded; `medium_changed` is
published by the driver but consumed by no one; remount-on-replug is
bench-pending.
## The five phases and how they depend
```
S1 identity ladder ─────┐
├──► S2 mount map ──► S3 multi-volume ──► S4 exFAT
(S2 keys on rung-4, so │ (needs S2+S3)
is independent of S1) │
└──► S5 removal robustness (independent; single-volume)
```
- **S1** grows the identity read off the medium (GPT GUID, FAT serial+label).
- **S2** moves the last policy out of hardcode into `volumes.csv` +
`filesystems.csv`, keyed on today's rung-4 identity — so it does **not** wait
on S1; S1 only enriches the identity values the same map consumes.
- **S3** generalizes to N volumes across N devices.
- **S4** adds exFAT — the second engine that proves the V1 harness extraction.
Needs S2 (to route by signature) and S3 (to run a second volume in the VM).
- **S5** closes the removal-lifecycle gaps. Independent of the rest; single-volume.
Recommended build order is S1 → S2 → S3 → S4 → S5. S5 may be pulled earlier if
removal robustness matters more than multi-volume; S1 and S2 may swap.
---
## S1 — the identity ladder
**Goal.** Grow `system/services/volume-manager/partition.zig` from the single
rung-4 identity into a ladder that reads the richest available content identity:
GPT partition GUID (rung 1, 128-bit), FAT volume serial + label (rung 3), MBR
signature + index (rung 4, kept), bare-FAT (kept, enriched to its serial). The
`u64` identity becomes a small tagged struct `Identity{ rung, key: u128, label,
has_label }`. Because GPT metadata is at LBA 1 and the entry array beyond it, and
the FAT serial is in each partition's VBR, `firstVolume` stops taking one
preloaded block-0 slice and takes a `SectorReader` (context + read-one-sector
fn, mirroring the engine's `BlockDevice` vtable) — host-testable against a
RAM-disk reader exactly as the four existing `partition.zig` tests are.
**Key touchpoints.** `partition.zig` (the `Rung`/`Identity`/`SectorReader` types;
`gptFirstVolume`; `fatIdentity`; `firstVolume` control flow — GPT authoritative,
else MBR walk skipping type-0xEE, else bare-FAT, each preferring the FAT serial
over the disk signature); `volume-manager.zig` (`Volume.identity` type; a
`ProbeReader` over the existing 512-byte bounce; the probe log prints
`identity.key`); `build.zig` (add `b.dependency("volume-manager", .{})` to the
package-test aggregation loop so the host fixtures run under root `zig build
test`). Declared bound `gpt_entry_scan_maximum = 128` with the full bounds block;
`sector_bytes`/`fat_label_bytes` named consts.
**Steps (commits).** (1) The SectorReader/Identity flag-day — pure refactor, no
new behavior, all existing tests green. (2) GPT parsing (rung 1) — protective-MBR
+ `EFI PART` signature + header CRC-32 + per-entry overflow-safe range
validation (the confinement-safety invariant the driver's clamp rests on,
extended to GPT). (3) FAT serial + label (rung 3), preferred over the disk
signature; `Identity.eql`. (4) Test wiring, `check-bounds.py`, full suite, docs,
memory.
**Discrimination.** Host: a GPT disk yields `rung==.gpt_guid` + the exact GUID
key (old code walks the 0xEE protective entry as an ordinary partition); a GPT
entry past the device is skipped, an out-of-device-only GPT returns null (the
security-boundary guard); an invalid GPT header is not a volume; a bare FAT
reports its real serial `0x12345678` not the rung-4 pseudo-signature; an
MBR+FAT partition prefers the serial over the disk signature. On-image: the
`volume-probe` QEMU regex tightens to `volume 0x0*12345678` — the boot image's
real FAT32 serial reaches the running log.
**Top risks.** Adversarial GPT input from untrusted media (huge entry counts,
bogus offsets, overflowing ranges) — mitigated by header CRC + size bounds +
`gpt_entry_scan_maximum` + per-entry overflow-safe validation under boundary
review. The `std.hash.crc` symbol in Zig 0.16 is unverified (fallback: a ~15-line
reflected CRC-32, poly `0xEDB88320`, used by both parser and fixtures so they
never drift onto magic constants).
---
## S2 — the mount map: volumes.csv + filesystems.csv
**Goal.** Move the last two pieces of storage policy out of hardcode into
configuration read by the volume manager. `filesystems.csv` (content signature →
filesystem binary) so the VM picks the binary from the probed signature;
`volumes.csv` (identity → mount prefix(es), danos's fstab) so the VM decides
mount placement, with a `/volumes/vol-<hex>` anonymous fallback. Parsed with
`library/csv` exactly as `device-registry` parses `devices.csv`. The VM hands the
binary + volume-id + mount specs to fat at spawn (the argv channel V3b already
uses for the volume id); fat retires `fat_mounts` and reads mounts from
argv[2..]. Keyed on rung-4 identity, so independent of S1 and forward-compatible.
**Key touchpoints.** New pure-logic modules `filesystem-map.zig` (parse +
`match(signature)`) and `volume-map.zig` (parse + `mountsFor(identity)` +
`derivedAnonymous`), mirroring `device-registry`, host-tested; `partition.zig`
`Volume` gains a `signature`; `volume-manager.zig` loads both tables in
`initialise`, resolves binary + mounts, spawns the chosen binary with the mount
argv; `fat.zig` deletes `fat_mounts`, parses argv[2..] into a bounded
`MountSpec` array; new `system/configuration/filesystems.csv` +
`volumes.csv`; `build.zig` bundles them; `make-fat-image.py` writes a real 4-byte
MBR disk signature so the boot volume's identity is a legible non-zero key.
**Steps (commits).** (1) partition emits a signature. (2) `filesystem-map`
parser + host tests. (3) `volume-map` parser + host tests. (4) Ship the tables +
VM integration **behind fat's still-hardcoded mounts** (behavior-preserving —
the parsers are proven before consumption flips; full suite green). (5) fat
consumes argv mounts; add the `/volumes/boot` row; the `volume-mapped` QEMU case
(shown failing against HEAD~1); flip the docs' `volumes.csv`/`filesystems.csv`
pending markers.
**Discrimination.** QEMU `volume-mapped`: the boot volume, mapped in
`volumes.csv` to `/volumes/boot`, makes fat log `mounted /volumes/boot` — a line
the old hardcoded `fat_mounts` never emits. Host: `mountsFor(mapped)` returns the
CSV prefix, `mountsFor(unmapped)` derives `/volumes/vol-<hex>`; `match(.fat)`
returns the configured binary; fat's argv parser makes installed mounts a
function of argv.
**Top risks.** Step 5's blast radius — a parser/argv bug breaks every
fat-dependent case at once; mitigated by landing the VM half behind fat's
hardcoded mounts first (step 4). The argv blob is 256 bytes (process.zig) — cap
emitted mounts and refuse+log on overflow. `/system/logs` is now a
`volumes.csv` row: dropping it silently stops log persistence — ship it by
default and document the boot-identity contract in the CSV header.
---
## S3 — multi-volume
**Goal.** Generalize from `var volume: ?Volume` / `volume_id = 1` / first-device
/ first-partition to a bounded table of volumes across a bounded table of
devices. `partition.allVolumes` returns ALL partitions; the VM adopts EVERY
mass-storage provider, probes each device's table, and for each partition spawns
one FAT confined to that partition's range (the per-sender clamp is already
built), with a distinct `/volumes/<name>` and its OWN backoff/crash-loop state.
Removal is per-device. The **boot volume is identified by content** — the FAT
process installs the `/system/configuration` + `/system/logs` rewrites only when
its own volume resolves `/system/configuration` — so it works as the 2nd
partition of the 2nd device just as the 1st of the 1st. (This clarifies the
initrd relationship: the kernel already serves `/system/configuration` +
binaries read-only from the initrd, which is what lets danos boot with NO volume
mounted; a mounted boot volume only adds writable, persistent
`/system/configuration` + `/system/logs` that shadow the initrd via
longest-prefix match.)
**Key touchpoints.** `partition.zig` `firstVolume` → `allVolumes(block0,
device_blocks, out) usize` (per-entry overflow-safe skip preserved);
`volume-manager.zig` the core refactor — `Volume` absorbs the file-global
supervision state as per-volume fields, a `StorageDevice` table owns each
adopted device's channel once, `volumes[maximum_volumes]` replaces the singleton,
a monotonic `next_volume_id`, `openAllStorage`/`adoptAndProbe`,
`gatherPresentStorage` + per-device reconcile in `pollTick`, `onHello`/
`onNotification` keyed across the table; `fat.zig` content-conditional boot
mounts; `system/kernel/vfs.zig` raise `maximum_mounts` (8 → 16) with a refreshed
bounds annotation; `make-fat-image.py` a partition-table mode; `build/images.zig`
the test disk artifacts.
**Steps (commits).** (1) `partition.allVolumes` + two-partition host test. (2)
Tables, behavior-preserving (still one device / one volume). (3) Multi-device +
multi-partition. (4) Per-volume mount naming via argv (first = `/volumes/usb`,
rest `/volumes/usb<n>` — bridge until `volumes.csv`). (5) fat content-conditional
boot mounts. (6) Raise `maximum_mounts`. (7) Partitioned-image tool. (8)
`two-volume` QEMU case. (9) `boot-2nd-partition` case. (10) Adversarial review,
full suite, docs, memory.
**Discrimination.** Host: an MBR with two partitions yields two volumes with
distinct identities (old `firstVolume` returns one). QEMU `two-volume`: a
two-partition second device yields two mount lines at two base_lbas (old
`openAnyStorage` adopts only the first device). `boot-2nd-partition`: the boot
volume works as partition 2 (old code confines fat to partition 1, whose
`/system` rewrite backs empty space). Per-volume supervision: killing one
volume's FAT restarts only that one (old module-scope supervision can't
attribute an exit to one of two).
**Top risks.** OVMF booting an MBR ESP on partition 2 may be flaky in CI —
fallback to a content-detection-ordering assertion + bench-verified boot (the
track's existing precedent). N-client range reclamation in usb-storage
(`maximum_ranges=64`) must reclaim each of N confined pids' ranges — the V2
mechanism, previously exercised with one live client. Duplicate boot volumes:
S3 supports exactly one and must log loudly if a second also resolves the boot
markers (arbitration deferred to S4).
---
## S4 — exFAT: the second engine
**Goal.** A working exFAT filesystem as `system/services/exfat` that is nothing
but an engine + a `main`, reusing `library/kernel/file-system-harness.zig`'s
`Server(Engine)` wholesale — the reuse claim the architecture makes, now proven.
The harness already owns vfs serving, the badge-scoped open-node table,
create-on-open/O_TRUNC, mount registration, the exit sweep, bring-up retry,
per-turn time stamping, and durable-on-close. S4 writes only the exFAT-specific
bits.
**Key touchpoints.** New `system/services/exfat/on-disk.zig` (the Main Boot
Sector VBR + the five 32-byte directory-entry types as `align(1)` extern structs;
`geometryOf` accepting only `"EXFAT "` + `0xAA55`; `setChecksum`, `nameHash`,
`upcaseAscii`); `engine.zig` (`FileSystem` behind the identical `BlockDevice`
vtable with fat's exact method set; **allocation-bitmap** cluster authority — the
deepest departure from FAT; read honoring `no_fat_chain` contiguous vs FAT-follow;
File+Stream+FileName set assembly with recomputed set checksum); `exfat.zig` (the
thin service, a near-clone of fat.zig); build wiring + `service("exfat")`;
`tools/make-exfat-image.py` (pure stdlib, correct boot checksum, up-case table —
**no committed .img**); the `filesystems.csv` EXFAT row (S2) + the `"EXFAT "`
recognizer; a `exfat-test` fixture cloned from fat-test.
**Steps (commits).** (1) on-disk.zig byte layout. (2) engine read path. (3)
engine write path (bitmap allocate/free, real 32-bit FAT chain with
`no_fat_chain=0`, set-checksum recompute). (4) service + build wiring. (5)
`make-exfat-image.py` + image assembly. (6) Routing: `filesystems.csv` +
recognizer. (7) `exfat-test` fixture + **cross-engine discrimination** host test.
(8) In-VM lifecycle drill (second removable device; mount/mutations/removal). (9)
Bounds, docs, adversarial review, memory.
**Discrimination.** The named one: `fat.mount(exfat_img) == null` (fat reads
bytes-per-sector at VBR offset 11 = exFAT's MustBeZero = 0 → reject) AND
`exfat.mount(fat_img) == null`, each mounting its own as a control. Host: read
across a cluster boundary on both a contiguous and a fragmented file; write
across >1 cluster setting the bitmap bits (not the FAT) and a validating set
checksum. QEMU: `exfat: mounted /volumes/exfat` + `exfat-test: ok`;
`exfat-removal` yanks the exFAT stick mid-write while the FAT boot volume keeps
serving.
**Top risks.** Allocation authority is the bitmap, not the FAT — allocating
without setting the bit silently corrupts free space (highest-attention area).
A directory-entry SET can straddle sector/cluster boundaries — scanDirectory,
set-checksum, and updateStreamEntry must handle multi-sector sets. vfs offsets
are u32 while exFAT DataLength is u64 — clamp and document (as fat does). The
in-VM drill needs S3 (a non-boot exFAT volume beside the FAT boot volume); if S4
landed before S3 the discrimination would rest on host tests until multi-volume
exists.
---
## S5 — removal robustness
**Goal.** Close the three known gaps so every removal trigger is exercised
end-to-end. (1) **Consume `medium_changed`** — the VM subscribes to the driver's
already-published event so the "device stays, medium leaves" case (a card
reader, an ejected removable) runs the same kill-retire-remount path as a pulled
stick, closing the second of the "two triggers, one lifecycle" the architecture
specifies. (2) **Storage-driver-crash rebuild** — a driver that dies while its
device stays present is detected and the volume subtree rebuilt on the restarted
driver's fresh channel, instead of leaving fat wedged on a dead channel (the V4
review's open edge). (3) **QEMU-verified remount-on-replug** — the device-return
half is proven, not merely asserted-unmount. Single-volume; independent of S1–S4.
**Key touchpoints.** `library/kernel/service.zig` an additive, behavior-neutral
`on_buffered_message` callback so a buffered-message wake forwards its payload
(no existing service sets it); `volume-manager.zig` subscribe on `bringUpVolume`
success, `onMediumEvent` with change-count dedupe running a medium-teardown (with
`encodeUnsubscribe` before close so the driver's 8-slot table doesn't leak), plus
`channelAlive()` (a `geometry()` liveness probe) + `rebuildVolume()` used in
`pollTick` and the child-exit path; `fat.zig` re-probe geometry on I/O failure
and exit on channel death (device NAK keeps serving); `device-manager.zig` a
`test-storage-restart` mode (mirroring `test-scanout-restart`) to kill usb-storage
once, post-mount, as the discrimination trigger.
**Steps (commits).** (1) **Cheap decisive experiments first** (no commits): QMP-
eject the boot medium and confirm `usb-storage: medium absent` fires under QEMU
(the whole item-1 chain depends on it); and test whether a boot-controller
`device_add` is re-presented (settles whether item 3 extends `volume-removal` or
needs a second controller as H1 does). (2) Harness `on_buffered_message`
(behavior-neutral). (3) VM consumes `medium_changed`. (4) `volume-medium-change`
case (fails pre-step-3). (5) VM driver-crash rebuild. (6) fat observes dead
channel and exits. (7) `volume-driver-restart` trigger + case. (8) `volume-replug`
(second controller if needed). (9) Docs + the real-hardware bench protocol. (10)
Full suite + memory.
**Discrimination.** `volume-medium-change`: an eject with the device left in the
tree unmounts (old VM never subscribes → the event goes to no one → mount
persists). `volume-driver-restart`: killing usb-storage post-mount while its
child stays present triggers a rebuild and a SECOND mount + post-kill read (old
`pollTick` only checks `isDevicePresent`, still true, and restarts fat against
the stale channel → wedge/crash-loop). `volume-replug`: a device return on a
second controller drives a remount (the existing case only ever sees the unmount
half).
**Top risks.** QEMU medium-eject must make TEST UNIT READY report not-ready —
step 1(a) validates this before any code. The op-16 overlap (`medium_changed` ==
`hello` by number) is safe only because async events arrive as `isMessage`
notifications and never reach `Serve.dispatch` — the intercept must run in the
notification branch and never catch a synchronous hello. fat can't today
distinguish EPEER from a device NAK (`CallError` swallows the errno) — the plan
uses a geometry re-probe as the liveness oracle, which is correct but indirect.
---
## Decisions embedded in this plan (flag if you disagree)
Most of the ~28 design-time questions the phase design surfaced are settled with
the recommended default and noted in the phases above. Two are worth your eye:
1. **`/volumes/usb` naming (S2).** The plan **keeps `/volumes/usb`
transitionally** and adds the mapped `/volumes/boot` — zero fixture churn,
inherent discrimination. The design-purer alternative is to REPLACE
`/volumes/usb` (killing the port-name the rationale condemns) and migrate the
three runtime fixtures + four QEMU regexes in a coupled sweep. Recommendation:
keep transitionally now, drop `/volumes/usb` as a small follow-up once the map
is proven.
2. **exFAT write scope (S4).** The plan ships **full mutation parity**
(createFile + writeFile + truncate + removeFile + createDirectory + rename) so
`exfat-test` reuses fat-test's full matrix. The minimum viable is
create/write/truncate/remove only. Recommendation: full parity.
Everything else (defer GPT entry-array CRC to correctness-only; GUID key =
little-endian u128 pinned now; share the DOS date-time helper into a library
module both engines import; a second removable usb-storage device for the exFAT
drill rather than a boot-disk partition; VM-poll `channelAlive()` as the
load-bearing crash-detection guarantee with fat's exit as a promptness
optimization) is taken as the recommended default in the phase text.