Two spots still said /volumes/DANOS and "label->hex", contradicting the settled path-is-the-id decision. Corrected to id-path.
22 KiB
Finishing the storage stack: the S1–S5 plan
2026-08-09. Continues the volume-manager track (V0–V5, on main) from "one FAT
volume" to "any filesystem, any number of volumes, identified by content,
remounting where they belong, surviving a driver crash." Executes the settled
design in
storage-architecture.md and
storage-design-rationale.md.
Track discipline as always: work on main; one commit per coherent step with
git commit -F (no -m, no co-author trailer); every new test shown to FAIL
against the old behavior; one QEMU suite at a time; CSV is configuration read by
the volume manager (the policy), never itself policy; bounds discipline
(tools/check-bounds.py gate); adversarial boundary review at each phase.
Where V0–V4 left it
The volume manager probes ONE storage device, parses its MBR (rung-4 identity =
(diskSignature<<8)|index), spawns ONE FAT service confined to that partition's
badge-scoped block range, supervises it, and unmounts it when the device is
pulled. partition.firstVolume returns the FIRST partition; openAnyStorage
adopts the FIRST device; var volume: ?Volume and volume_id = 1 are singular;
filesystem_binary and fat's mount prefixes are hardcoded; medium_changed is
published by the driver but consumed by no one; remount-on-replug is
bench-pending.
The five phases and how they depend
S1 identity ladder ──► S2 mount map ──► S3 multi-volume ──► S4 exFAT
(needs S2+S3)
S5 removal robustness ── independent; single-volume ── may land any time
- S1 grows the identity read off the medium (GPT GUID, FAT serial+label). It comes FIRST because a volume's mount name is now its identity (below), and the friendly form of that name is the FAT label / GPT name S1 parses.
- S2 moves the last policy out of hardcode into
volumes.csv+filesystems.csv, and names each volume by its identity id (a GUID/key), keeping the label as separate, queryable display metadata — no port-name. - S3 generalizes to N volumes across N devices.
- S4 adds exFAT — a COMPLETE second engine that proves the V1 harness extraction. Needs S2 (to route by signature) and S3 (to run a second volume).
- S5 closes the removal-lifecycle gaps. Independent of the rest; single-volume.
Recommended build order is S1 → S2 → S3 → S4 → S5. S5 may be pulled earlier.
S1 — the identity ladder
Goal. Grow system/services/volume-manager/partition.zig from the single
rung-4 identity into a ladder that reads the richest available content identity:
GPT partition GUID (rung 1, 128-bit), FAT volume serial + label (rung 3), MBR
signature + index (rung 4, kept), bare-FAT (kept, enriched to its serial). The
u64 identity becomes a small tagged struct Identity{ rung, key: u128, label, has_label } — key is the id (the path handle), label is the display
name (the GPT 36-char partition name, or the FAT volume label), a separate
field per the id/name split S2 relies on. Because GPT metadata is at LBA 1 and
the entry array beyond it, and
the FAT serial is in each partition's VBR, firstVolume stops taking one
preloaded block-0 slice and takes a SectorReader (context + read-one-sector
fn, mirroring the engine's BlockDevice vtable) — host-testable against a
RAM-disk reader exactly as the four existing partition.zig tests are.
Key touchpoints. partition.zig (the Rung/Identity/SectorReader types;
gptFirstVolume; fatIdentity; firstVolume control flow — GPT authoritative,
else MBR walk skipping type-0xEE, else bare-FAT, each preferring the FAT serial
over the disk signature); volume-manager.zig (Volume.identity type; a
ProbeReader over the existing 512-byte bounce; the probe log prints
identity.key); build.zig (add b.dependency("volume-manager", .{}) to the
package-test aggregation loop so the host fixtures run under root zig build test). Declared bound gpt_entry_scan_maximum = 128 with the full bounds block;
sector_bytes/fat_label_bytes named consts.
Steps (commits). (1) The SectorReader/Identity flag-day — pure refactor, no new behavior, all existing tests green. (2) GPT parsing (rung 1) — protective-MBR
EFI PARTsignature + header CRC-32 + per-entry overflow-safe range validation (the confinement-safety invariant the driver's clamp rests on, extended to GPT). (3) FAT serial + label (rung 3), preferred over the disk signature;Identity.eql. (4) Test wiring,check-bounds.py, full suite, docs, memory.
Discrimination. Host: a GPT disk yields rung==.gpt_guid + the exact GUID
key (old code walks the 0xEE protective entry as an ordinary partition); a GPT
entry past the device is skipped, an out-of-device-only GPT returns null (the
security-boundary guard); an invalid GPT header is not a volume; a bare FAT
reports its real serial 0x12345678 not the rung-4 pseudo-signature; an
MBR+FAT partition prefers the serial over the disk signature. On-image: the
volume-probe QEMU regex tightens to volume 0x0*12345678 — the boot image's
real FAT32 serial reaches the running log.
Top risks. Adversarial GPT input from untrusted media (huge entry counts,
bogus offsets, overflowing ranges) — mitigated by header CRC + size bounds +
gpt_entry_scan_maximum + per-entry overflow-safe validation under boundary
review. The std.hash.crc symbol in Zig 0.16 is unverified (fallback: a ~15-line
reflected CRC-32, poly 0xEDB88320, used by both parser and fixtures so they
never drift onto magic constants).
S2 — the mount map: volumes.csv + filesystems.csv
Goal. Move the last two pieces of storage policy out of hardcode into
configuration read by the volume manager, and split a volume's id from its
label (the database model: the id is the real, stable, unique key software
uses; the label is a mutable display name). filesystems.csv (content signature
→ filesystem binary) so the VM picks the binary from the probed signature;
volumes.csv (id → mount prefix, danos's fstab) as the explicit override for a
volume the user wants at a fixed path.
The mount path IS the identity id, never a label or a port/role name. A
volume mounts at /volumes/<id> — the GPT partition GUID for a GPT volume, a
fat-<serial> / mbr-<sig>-<index> form otherwise (exact rendering is an impl
detail; the point is a stable, unique, content-derived string). Because the path
is the id and never the label, two distinct volumes that happen to share a label
(Backup, UNTITLED, unlabeled) get distinct paths automatically and never
collide; only genuinely identical ids (dd-cloned media) hit the rationale's
duplicate-identity policy (first mounts, second suffixed and logged).
The label is display metadata, exposed by a protocol query, not the path. S1
reads it off the medium (FAT volume label, GPT partition name); a volume-manager
volumes verb returns { id, mount_path, label } per volume so a future
shell/UI can show the friendly name — the id↔name split, like a table's primary
key vs its display column. (volumes.csv may optionally carry a chosen label
override alongside the path override, both keyed on id.) There is no
/volumes/usb and no /volumes/boot; the boot volume is detected by content (it
installs the /system/configuration + /system/logs rewrites) but is named by
its id like any other.
Parsed with library/csv exactly as device-registry parses devices.csv. The
VM hands the binary + volume-id + mount specs to fat at spawn (the argv channel
V3b already uses for the volume id); fat retires fat_mounts and reads mounts
from argv[2..].
Key touchpoints. New pure-logic modules filesystem-map.zig (parse +
match(signature)) and volume-map.zig (parse + mountsFor(identity) +
derivedAnonymous), mirroring device-registry, host-tested; partition.zig
Volume gains a signature; volume-manager.zig loads both tables in
initialise, resolves binary + mounts, spawns the chosen binary with the mount
argv; fat.zig deletes fat_mounts, parses argv[2..] into a bounded
MountSpec array; new system/configuration/filesystems.csv +
volumes.csv; build.zig bundles them; make-fat-image.py writes a real 4-byte
MBR disk signature so the boot volume's identity is a legible non-zero key.
Steps (commits). (1) partition emits a signature. (2) filesystem-map
parser + host tests. (3) volume-map parser (id→prefix override) + the id-path
deriver + the volumes label-query verb + host tests. (4) Ship the tables +
VM integration behind fat's still-hardcoded mounts (behavior-preserving —
parsers proven before consumption flips; full suite green). (5) fat consumes
argv mounts; the VM mounts the boot volume at its id-path (/volumes/fat-12345678,
from its serial); migrate the fixtures + QEMU regexes off /volumes/usb
(fat-test, badge-scope-test, vfs-test, and the four cases) in the same commit;
the volume-identity-name QEMU case (shown failing against HEAD~1); flip the
docs' pending markers.
Discrimination. QEMU volume-identity-name: the boot volume mounts at its
id-path (/volumes/fat-12345678, from its serial), a line the old hardcoded
/volumes/usb never emits; the volumes query returns that id paired with the
label DANOS. Host: two volumes sharing a label but not an id get distinct
id-paths (the collision the label-as-name approach could not resolve stably); a
volumes.csv override sends a mapped id to its chosen prefix; match(.fat)
returns the configured binary; fat's argv parser makes installed mounts a
function of argv.
Top risks. Step 5's blast radius — a parser/argv bug, OR the /volumes/usb→
id-path migration missing a fixture/regex, breaks every fat-dependent case at
once; mitigated by landing the VM half behind fat's hardcoded mounts first (step
4) and making the name migration one atomic, complete sweep. The
argv blob is 256 bytes (process.zig) — cap emitted mounts and refuse+log on
overflow. /system/logs is now a volumes.csv concern: dropping the boot
identity's rewrite rows silently stops log persistence — ship them by default
and document the boot-identity contract in the CSV header.
S3 — multi-volume
Goal. Generalize from var volume: ?Volume / volume_id = 1 / first-device
/ first-partition to a bounded table of volumes across a bounded table of
devices. partition.allVolumes returns ALL partitions; the VM adopts EVERY
mass-storage provider, probes each device's table, and for each partition spawns
one FAT confined to that partition's range (the per-sender clamp is already
built), with a distinct /volumes/<name> and its OWN backoff/crash-loop state.
Removal is per-device. The boot volume is identified by content — the FAT
process installs the /system/configuration + /system/logs rewrites only when
its own volume resolves /system/configuration — so it works as the 2nd
partition of the 2nd device just as the 1st of the 1st. (This clarifies the
initrd relationship: the kernel already serves /system/configuration +
binaries read-only from the initrd, which is what lets danos boot with NO volume
mounted; a mounted boot volume only adds writable, persistent
/system/configuration + /system/logs that shadow the initrd via
longest-prefix match.)
Key touchpoints. partition.zig firstVolume → allVolumes(block0, device_blocks, out) usize (per-entry overflow-safe skip preserved);
volume-manager.zig the core refactor — Volume absorbs the file-global
supervision state as per-volume fields, a StorageDevice table owns each
adopted device's channel once, volumes[maximum_volumes] replaces the singleton,
a monotonic next_volume_id, openAllStorage/adoptAndProbe,
gatherPresentStorage + per-device reconcile in pollTick, onHello/
onNotification keyed across the table; fat.zig content-conditional boot
mounts; system/kernel/vfs.zig raise maximum_mounts (8 → 16) with a refreshed
bounds annotation; make-fat-image.py a partition-table mode; build/images.zig
the test disk artifacts.
Steps (commits). (1) partition.allVolumes + two-partition host test. (2)
Tables, behavior-preserving (still one device / one volume). (3) Multi-device +
multi-partition. (4) Per-volume mount naming via argv — each volume by its
id-path (from S2), one per volume, no port-name. (5) fat content-conditional
boot mounts. (6) Raise maximum_mounts. (7) Partitioned-image tool. (8)
two-volume QEMU case. (9) boot-2nd-partition case. (10) Adversarial review,
full suite, docs, memory.
Discrimination. Host: an MBR with two partitions yields two volumes with
distinct identities (old firstVolume returns one). QEMU two-volume: a
two-partition second device yields two mount lines at two base_lbas (old
openAnyStorage adopts only the first device). boot-2nd-partition: the boot
volume works as partition 2 (old code confines fat to partition 1, whose
/system rewrite backs empty space). Per-volume supervision: killing one
volume's FAT restarts only that one (old module-scope supervision can't
attribute an exit to one of two).
Top risks. OVMF booting an MBR ESP on partition 2 may be flaky in CI —
fallback to a content-detection-ordering assertion + bench-verified boot (the
track's existing precedent). N-client range reclamation in usb-storage
(maximum_ranges=64) must reclaim each of N confined pids' ranges — the V2
mechanism, previously exercised with one live client. Duplicate boot volumes:
S3 supports exactly one and must log loudly if a second also resolves the boot
markers (arbitration deferred to S4).
S4 — exFAT: the second engine
Goal. A working exFAT filesystem as system/services/exfat that is nothing
but an engine + a main, reusing library/kernel/file-system-harness.zig's
Server(Engine) wholesale — the reuse claim the architecture makes, now proven.
The harness already owns vfs serving, the badge-scoped open-node table,
create-on-open/O_TRUNC, mount registration, the exit sweep, bring-up retry,
per-turn time stamping, and durable-on-close. S4 writes only the exFAT-specific
bits — but in full: complete read AND write, directories, rename, and the
real on-disk up-case table for correct case-folding. Not a read-first, minimal-
write, or ASCII-only subset. The only limit that survives is the vfs protocol's
u32 file-offset surface (a 4 GiB addressable-size cap that applies to FAT too),
which is a separate vfs-protocol change, not an exFAT shortcut.
Key touchpoints. New system/services/exfat/on-disk.zig (the Main Boot
Sector VBR + the five 32-byte directory-entry types as align(1) extern structs;
geometryOf accepting only "EXFAT " + 0xAA55; setChecksum, nameHash,
and case-folding driven by the volume's on-disk up-case table); engine.zig (FileSystem behind the identical BlockDevice
vtable with fat's exact method set; allocation-bitmap cluster authority — the
deepest departure from FAT; read honoring no_fat_chain contiguous vs FAT-follow;
File+Stream+FileName set assembly with recomputed set checksum); exfat.zig (the
thin service, a near-clone of fat.zig); build wiring + service("exfat");
tools/make-exfat-image.py (pure stdlib, correct boot checksum, up-case table —
no committed .img); the filesystems.csv EXFAT row (S2) + the "EXFAT "
recognizer; a exfat-test fixture cloned from fat-test.
Steps (commits). (1) on-disk.zig byte layout. (2) engine read path. (3)
engine write path (bitmap allocate/free, real 32-bit FAT chain with
no_fat_chain=0, set-checksum recompute). (4) service + build wiring. (5)
make-exfat-image.py + image assembly. (6) Routing: filesystems.csv +
recognizer. (7) exfat-test fixture + cross-engine discrimination host test.
(8) In-VM lifecycle drill (second removable device; mount/mutations/removal). (9)
Bounds, docs, adversarial review, memory.
Discrimination. The named one: fat.mount(exfat_img) == null (fat reads
bytes-per-sector at VBR offset 11 = exFAT's MustBeZero = 0 → reject) AND
exfat.mount(fat_img) == null, each mounting its own as a control. Host: read
across a cluster boundary on both a contiguous and a fragmented file; write
across >1 cluster setting the bitmap bits (not the FAT) and a validating set
checksum. QEMU: exfat: mounted /volumes/exfat + exfat-test: ok;
exfat-removal yanks the exFAT stick mid-write while the FAT boot volume keeps
serving.
Top risks. Allocation authority is the bitmap, not the FAT — allocating without setting the bit silently corrupts free space (highest-attention area). A directory-entry SET can straddle sector/cluster boundaries — scanDirectory, set-checksum, and updateStreamEntry must handle multi-sector sets. vfs offsets are u32 while exFAT DataLength is u64 — clamp and document (as fat does). The in-VM drill needs S3 (a non-boot exFAT volume beside the FAT boot volume); if S4 landed before S3 the discrimination would rest on host tests until multi-volume exists.
S5 — removal robustness
Goal. Close the three known gaps so every removal trigger is exercised
end-to-end. (1) Consume medium_changed — the VM subscribes to the driver's
already-published event so the "device stays, medium leaves" case (a card
reader, an ejected removable) runs the same kill-retire-remount path as a pulled
stick, closing the second of the "two triggers, one lifecycle" the architecture
specifies. (2) Storage-driver-crash rebuild — a driver that dies while its
device stays present is detected and the volume subtree rebuilt on the restarted
driver's fresh channel, instead of leaving fat wedged on a dead channel (the V4
review's open edge). (3) QEMU-verified remount-on-replug — the device-return
half is proven, not merely asserted-unmount. Single-volume; independent of S1–S4.
Key touchpoints. library/kernel/service.zig an additive, behavior-neutral
on_buffered_message callback so a buffered-message wake forwards its payload
(no existing service sets it); volume-manager.zig subscribe on bringUpVolume
success, onMediumEvent with change-count dedupe running a medium-teardown (with
encodeUnsubscribe before close so the driver's 8-slot table doesn't leak), plus
channelAlive() (a geometry() liveness probe) + rebuildVolume() used in
pollTick and the child-exit path; fat.zig re-probe geometry on I/O failure
and exit on channel death (device NAK keeps serving); device-manager.zig a
test-storage-restart mode (mirroring test-scanout-restart) to kill usb-storage
once, post-mount, as the discrimination trigger.
Steps (commits). (1) Cheap decisive experiments first (no commits): QMP-
eject the boot medium and confirm usb-storage: medium absent fires under QEMU
(the whole item-1 chain depends on it); and test whether a boot-controller
device_add is re-presented (settles whether item 3 extends volume-removal or
needs a second controller as H1 does). (2) Harness on_buffered_message
(behavior-neutral). (3) VM consumes medium_changed. (4) volume-medium-change
case (fails pre-step-3). (5) VM driver-crash rebuild. (6) fat observes dead
channel and exits. (7) volume-driver-restart trigger + case. (8) volume-replug
(second controller if needed). (9) Docs + the real-hardware bench protocol. (10)
Full suite + memory.
Discrimination. volume-medium-change: an eject with the device left in the
tree unmounts (old VM never subscribes → the event goes to no one → mount
persists). volume-driver-restart: killing usb-storage post-mount while its
child stays present triggers a rebuild and a SECOND mount + post-kill read (old
pollTick only checks isDevicePresent, still true, and restarts fat against
the stale channel → wedge/crash-loop). volume-replug: a device return on a
second controller drives a remount (the existing case only ever sees the unmount
half).
Top risks. QEMU medium-eject must make TEST UNIT READY report not-ready —
step 1(a) validates this before any code. The op-16 overlap (medium_changed ==
hello by number) is safe only because async events arrive as isMessage
notifications and never reach Serve.dispatch — the intercept must run in the
notification branch and never catch a synchronous hello. fat can't today
distinguish EPEER from a device NAK (CallError swallows the errno) — the plan
uses a geometry re-probe as the liveness oracle, which is correct but indirect.
Decisions (settled)
Both flagged decisions are settled:
-
A volume's path is its id; the label is display metadata (S1/S2). The mount path is the identity id — the GPT GUID, else a
fat-<serial>/mbr-<sig>-<index>form — a stable, unique, content-derived handle software uses. The label (FAT volume label / GPT partition name) is a mutable display name, NOT in the path; a volume-managervolumesverb returns{ id, mount_path, label }so a UI can show the friendly name (the database id/name split). Same-label-different-id volumes therefore never collide; only identical ids (dd-clones) hit first-wins-and-log.volumes.csvoverrides the path (and optionally the label) for a chosen volume, keyed on id. Makes S1 precede S2 and folds the/volumes/usb→ id-path fixture + regex migration into S2 step 5. -
exFAT is implemented in full (S4). A complete exFAT: full read and write, directories, rename, and the on-disk up-case table for correct case-folding — not a read-first or ASCII-only subset. The one remaining limit is the vfs protocol's u32 file-offset surface, which caps addressable file size at 4 GiB for ALL filesystems (FAT included); widening it to u64 is a separate vfs- protocol change, flagged but out of the exFAT engine's scope.
The ~26 smaller design-time questions are settled with the recommended default
in the phase text (defer GPT entry-array CRC to correctness-only; GUID key =
little-endian u128 pinned now; share the DOS date-time helper into a library
module both engines import; a second removable usb-storage device for the exFAT
drill; VM-poll channelAlive() as the load-bearing crash-detection guarantee).