Files
danos/docs/storage-stack-plan.md
T
Daniel Samson 4f9196c03e docs: storage plan — fix two id-path leftovers (DANOS/label->hex)
Two spots still said /volumes/DANOS and "label->hex", contradicting the settled
path-is-the-id decision. Corrected to id-path.
2026-08-09 22:55:10 +01:00

369 lines
22 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Finishing the storage stack: the S1–S5 plan
*2026-08-09. Continues the volume-manager track (V0–V5, on main) from "one FAT
volume" to "any filesystem, any number of volumes, identified by content,
remounting where they belong, surviving a driver crash." Executes the settled
design in
[storage-architecture.md](file-system-development/storage-architecture.md) and
[storage-design-rationale.md](file-system-development/storage-design-rationale.md).
Track discipline as always: work on main; one commit per coherent step with
`git commit -F` (no `-m`, no co-author trailer); every new test shown to FAIL
against the old behavior; one QEMU suite at a time; CSV is configuration read by
the volume manager (the policy), never itself policy; bounds discipline
(`tools/check-bounds.py` gate); adversarial boundary review at each phase.*
## Where V0–V4 left it
The volume manager probes ONE storage device, parses its MBR (rung-4 identity =
`(diskSignature<<8)|index`), spawns ONE FAT service confined to that partition's
badge-scoped block range, supervises it, and unmounts it when the device is
pulled. `partition.firstVolume` returns the FIRST partition; `openAnyStorage`
adopts the FIRST device; `var volume: ?Volume` and `volume_id = 1` are singular;
`filesystem_binary` and fat's mount prefixes are hardcoded; `medium_changed` is
published by the driver but consumed by no one; remount-on-replug is
bench-pending.
## The five phases and how they depend
```
S1 identity ladder ──► S2 mount map ──► S3 multi-volume ──► S4 exFAT
(needs S2+S3)
S5 removal robustness ── independent; single-volume ── may land any time
```
- **S1** grows the identity read off the medium (GPT GUID, FAT serial+label). It
comes FIRST because a volume's mount name is now its identity (below), and the
friendly form of that name is the FAT label / GPT name S1 parses.
- **S2** moves the last policy out of hardcode into `volumes.csv` +
`filesystems.csv`, and names each volume by its identity **id** (a GUID/key),
keeping the label as separate, queryable display metadata — no port-name.
- **S3** generalizes to N volumes across N devices.
- **S4** adds exFAT — a COMPLETE second engine that proves the V1 harness
extraction. Needs S2 (to route by signature) and S3 (to run a second volume).
- **S5** closes the removal-lifecycle gaps. Independent of the rest; single-volume.
Recommended build order is S1 → S2 → S3 → S4 → S5. S5 may be pulled earlier.
---
## S1 — the identity ladder
**Goal.** Grow `system/services/volume-manager/partition.zig` from the single
rung-4 identity into a ladder that reads the richest available content identity:
GPT partition GUID (rung 1, 128-bit), FAT volume serial + label (rung 3), MBR
signature + index (rung 4, kept), bare-FAT (kept, enriched to its serial). The
`u64` identity becomes a small tagged struct `Identity{ rung, key: u128, label,
has_label }` — `key` is the **id** (the path handle), `label` is the **display
name** (the GPT 36-char partition name, or the FAT volume label), a separate
field per the id/name split S2 relies on. Because GPT metadata is at LBA 1 and
the entry array beyond it, and
the FAT serial is in each partition's VBR, `firstVolume` stops taking one
preloaded block-0 slice and takes a `SectorReader` (context + read-one-sector
fn, mirroring the engine's `BlockDevice` vtable) — host-testable against a
RAM-disk reader exactly as the four existing `partition.zig` tests are.
**Key touchpoints.** `partition.zig` (the `Rung`/`Identity`/`SectorReader` types;
`gptFirstVolume`; `fatIdentity`; `firstVolume` control flow — GPT authoritative,
else MBR walk skipping type-0xEE, else bare-FAT, each preferring the FAT serial
over the disk signature); `volume-manager.zig` (`Volume.identity` type; a
`ProbeReader` over the existing 512-byte bounce; the probe log prints
`identity.key`); `build.zig` (add `b.dependency("volume-manager", .{})` to the
package-test aggregation loop so the host fixtures run under root `zig build
test`). Declared bound `gpt_entry_scan_maximum = 128` with the full bounds block;
`sector_bytes`/`fat_label_bytes` named consts.
**Steps (commits).** (1) The SectorReader/Identity flag-day — pure refactor, no
new behavior, all existing tests green. (2) GPT parsing (rung 1) — protective-MBR
+ `EFI PART` signature + header CRC-32 + per-entry overflow-safe range
validation (the confinement-safety invariant the driver's clamp rests on,
extended to GPT). (3) FAT serial + label (rung 3), preferred over the disk
signature; `Identity.eql`. (4) Test wiring, `check-bounds.py`, full suite, docs,
memory.
**Discrimination.** Host: a GPT disk yields `rung==.gpt_guid` + the exact GUID
key (old code walks the 0xEE protective entry as an ordinary partition); a GPT
entry past the device is skipped, an out-of-device-only GPT returns null (the
security-boundary guard); an invalid GPT header is not a volume; a bare FAT
reports its real serial `0x12345678` not the rung-4 pseudo-signature; an
MBR+FAT partition prefers the serial over the disk signature. On-image: the
`volume-probe` QEMU regex tightens to `volume 0x0*12345678` — the boot image's
real FAT32 serial reaches the running log.
**Top risks.** Adversarial GPT input from untrusted media (huge entry counts,
bogus offsets, overflowing ranges) — mitigated by header CRC + size bounds +
`gpt_entry_scan_maximum` + per-entry overflow-safe validation under boundary
review. The `std.hash.crc` symbol in Zig 0.16 is unverified (fallback: a ~15-line
reflected CRC-32, poly `0xEDB88320`, used by both parser and fixtures so they
never drift onto magic constants).
---
## S2 — the mount map: volumes.csv + filesystems.csv
**Goal.** Move the last two pieces of storage policy out of hardcode into
configuration read by the volume manager, and split a volume's **id** from its
**label** (the database model: the id is the real, stable, unique key software
uses; the label is a mutable display name). `filesystems.csv` (content signature
→ filesystem binary) so the VM picks the binary from the probed signature;
`volumes.csv` (id → mount prefix, danos's fstab) as the explicit override for a
volume the user wants at a fixed path.
**The mount path IS the identity id, never a label or a port/role name.** A
volume mounts at `/volumes/<id>` — the GPT partition GUID for a GPT volume, a
`fat-<serial>` / `mbr-<sig>-<index>` form otherwise (exact rendering is an impl
detail; the point is a stable, unique, content-derived string). Because the path
is the id and never the label, two distinct volumes that happen to share a label
(`Backup`, `UNTITLED`, unlabeled) get distinct paths automatically and never
collide; only genuinely identical *ids* (dd-cloned media) hit the rationale's
duplicate-identity policy (first mounts, second suffixed and logged).
**The label is display metadata, exposed by a protocol query, not the path.** S1
reads it off the medium (FAT volume label, GPT partition name); a volume-manager
`volumes` verb returns `{ id, mount_path, label }` per volume so a future
shell/UI can show the friendly name — the id↔name split, like a table's primary
key vs its display column. (`volumes.csv` may optionally carry a chosen label
override alongside the path override, both keyed on id.) There is no
`/volumes/usb` and no `/volumes/boot`; the boot volume is detected by content (it
installs the `/system/configuration` + `/system/logs` rewrites) but is named by
its id like any other.
Parsed with `library/csv` exactly as `device-registry` parses `devices.csv`. The
VM hands the binary + volume-id + mount specs to fat at spawn (the argv channel
V3b already uses for the volume id); fat retires `fat_mounts` and reads mounts
from argv[2..].
**Key touchpoints.** New pure-logic modules `filesystem-map.zig` (parse +
`match(signature)`) and `volume-map.zig` (parse + `mountsFor(identity)` +
`derivedAnonymous`), mirroring `device-registry`, host-tested; `partition.zig`
`Volume` gains a `signature`; `volume-manager.zig` loads both tables in
`initialise`, resolves binary + mounts, spawns the chosen binary with the mount
argv; `fat.zig` deletes `fat_mounts`, parses argv[2..] into a bounded
`MountSpec` array; new `system/configuration/filesystems.csv` +
`volumes.csv`; `build.zig` bundles them; `make-fat-image.py` writes a real 4-byte
MBR disk signature so the boot volume's identity is a legible non-zero key.
**Steps (commits).** (1) partition emits a signature. (2) `filesystem-map`
parser + host tests. (3) `volume-map` parser (id→prefix override) + the id-path
deriver + the `volumes` label-query verb + host tests. (4) Ship the tables +
VM integration **behind fat's still-hardcoded mounts** (behavior-preserving —
parsers proven before consumption flips; full suite green). (5) fat consumes
argv mounts; the VM mounts the boot volume at its id-path (`/volumes/fat-12345678`,
from its serial); **migrate the fixtures + QEMU regexes off `/volumes/usb`**
(fat-test, badge-scope-test, vfs-test, and the four cases) in the same commit;
the `volume-identity-name` QEMU case (shown failing against HEAD~1); flip the
docs' pending markers.
**Discrimination.** QEMU `volume-identity-name`: the boot volume mounts at its
id-path (`/volumes/fat-12345678`, from its serial), a line the old hardcoded
`/volumes/usb` never emits; the `volumes` query returns that id paired with the
label `DANOS`. Host: two volumes sharing a label but not an id get distinct
id-paths (the collision the label-as-name approach could not resolve stably); a
`volumes.csv` override sends a mapped id to its chosen prefix; `match(.fat)`
returns the configured binary; fat's argv parser makes installed mounts a
function of argv.
**Top risks.** Step 5's blast radius — a parser/argv bug, OR the `/volumes/usb`→
id-path migration missing a fixture/regex, breaks every fat-dependent case at
once; mitigated by landing the VM half behind fat's hardcoded mounts first (step
4) and making the name migration one atomic, complete sweep. The
argv blob is 256 bytes (process.zig) — cap emitted mounts and refuse+log on
overflow. `/system/logs` is now a `volumes.csv` concern: dropping the boot
identity's rewrite rows silently stops log persistence — ship them by default
and document the boot-identity contract in the CSV header.
---
## S3 — multi-volume
**Goal.** Generalize from `var volume: ?Volume` / `volume_id = 1` / first-device
/ first-partition to a bounded table of volumes across a bounded table of
devices. `partition.allVolumes` returns ALL partitions; the VM adopts EVERY
mass-storage provider, probes each device's table, and for each partition spawns
one FAT confined to that partition's range (the per-sender clamp is already
built), with a distinct `/volumes/<name>` and its OWN backoff/crash-loop state.
Removal is per-device. The **boot volume is identified by content** — the FAT
process installs the `/system/configuration` + `/system/logs` rewrites only when
its own volume resolves `/system/configuration` — so it works as the 2nd
partition of the 2nd device just as the 1st of the 1st. (This clarifies the
initrd relationship: the kernel already serves `/system/configuration` +
binaries read-only from the initrd, which is what lets danos boot with NO volume
mounted; a mounted boot volume only adds writable, persistent
`/system/configuration` + `/system/logs` that shadow the initrd via
longest-prefix match.)
**Key touchpoints.** `partition.zig` `firstVolume` → `allVolumes(block0,
device_blocks, out) usize` (per-entry overflow-safe skip preserved);
`volume-manager.zig` the core refactor — `Volume` absorbs the file-global
supervision state as per-volume fields, a `StorageDevice` table owns each
adopted device's channel once, `volumes[maximum_volumes]` replaces the singleton,
a monotonic `next_volume_id`, `openAllStorage`/`adoptAndProbe`,
`gatherPresentStorage` + per-device reconcile in `pollTick`, `onHello`/
`onNotification` keyed across the table; `fat.zig` content-conditional boot
mounts; `system/kernel/vfs.zig` raise `maximum_mounts` (8 → 16) with a refreshed
bounds annotation; `make-fat-image.py` a partition-table mode; `build/images.zig`
the test disk artifacts.
**Steps (commits).** (1) `partition.allVolumes` + two-partition host test. (2)
Tables, behavior-preserving (still one device / one volume). (3) Multi-device +
multi-partition. (4) Per-volume mount naming via argv — each volume by its
id-path (from S2), one per volume, no port-name. (5) fat content-conditional
boot mounts. (6) Raise `maximum_mounts`. (7) Partitioned-image tool. (8)
`two-volume` QEMU case. (9) `boot-2nd-partition` case. (10) Adversarial review,
full suite, docs, memory.
**Discrimination.** Host: an MBR with two partitions yields two volumes with
distinct identities (old `firstVolume` returns one). QEMU `two-volume`: a
two-partition second device yields two mount lines at two base_lbas (old
`openAnyStorage` adopts only the first device). `boot-2nd-partition`: the boot
volume works as partition 2 (old code confines fat to partition 1, whose
`/system` rewrite backs empty space). Per-volume supervision: killing one
volume's FAT restarts only that one (old module-scope supervision can't
attribute an exit to one of two).
**Top risks.** OVMF booting an MBR ESP on partition 2 may be flaky in CI —
fallback to a content-detection-ordering assertion + bench-verified boot (the
track's existing precedent). N-client range reclamation in usb-storage
(`maximum_ranges=64`) must reclaim each of N confined pids' ranges — the V2
mechanism, previously exercised with one live client. Duplicate boot volumes:
S3 supports exactly one and must log loudly if a second also resolves the boot
markers (arbitration deferred to S4).
---
## S4 — exFAT: the second engine
**Goal.** A working exFAT filesystem as `system/services/exfat` that is nothing
but an engine + a `main`, reusing `library/kernel/file-system-harness.zig`'s
`Server(Engine)` wholesale — the reuse claim the architecture makes, now proven.
The harness already owns vfs serving, the badge-scoped open-node table,
create-on-open/O_TRUNC, mount registration, the exit sweep, bring-up retry,
per-turn time stamping, and durable-on-close. S4 writes only the exFAT-specific
bits — but **in full**: complete read AND write, directories, rename, and the
real on-disk up-case table for correct case-folding. Not a read-first, minimal-
write, or ASCII-only subset. The only limit that survives is the vfs protocol's
u32 file-offset surface (a 4 GiB addressable-size cap that applies to FAT too),
which is a separate vfs-protocol change, not an exFAT shortcut.
**Key touchpoints.** New `system/services/exfat/on-disk.zig` (the Main Boot
Sector VBR + the five 32-byte directory-entry types as `align(1)` extern structs;
`geometryOf` accepting only `"EXFAT "` + `0xAA55`; `setChecksum`, `nameHash`,
and case-folding driven by the volume's **on-disk up-case table**); `engine.zig` (`FileSystem` behind the identical `BlockDevice`
vtable with fat's exact method set; **allocation-bitmap** cluster authority — the
deepest departure from FAT; read honoring `no_fat_chain` contiguous vs FAT-follow;
File+Stream+FileName set assembly with recomputed set checksum); `exfat.zig` (the
thin service, a near-clone of fat.zig); build wiring + `service("exfat")`;
`tools/make-exfat-image.py` (pure stdlib, correct boot checksum, up-case table —
**no committed .img**); the `filesystems.csv` EXFAT row (S2) + the `"EXFAT "`
recognizer; a `exfat-test` fixture cloned from fat-test.
**Steps (commits).** (1) on-disk.zig byte layout. (2) engine read path. (3)
engine write path (bitmap allocate/free, real 32-bit FAT chain with
`no_fat_chain=0`, set-checksum recompute). (4) service + build wiring. (5)
`make-exfat-image.py` + image assembly. (6) Routing: `filesystems.csv` +
recognizer. (7) `exfat-test` fixture + **cross-engine discrimination** host test.
(8) In-VM lifecycle drill (second removable device; mount/mutations/removal). (9)
Bounds, docs, adversarial review, memory.
**Discrimination.** The named one: `fat.mount(exfat_img) == null` (fat reads
bytes-per-sector at VBR offset 11 = exFAT's MustBeZero = 0 → reject) AND
`exfat.mount(fat_img) == null`, each mounting its own as a control. Host: read
across a cluster boundary on both a contiguous and a fragmented file; write
across >1 cluster setting the bitmap bits (not the FAT) and a validating set
checksum. QEMU: `exfat: mounted /volumes/exfat` + `exfat-test: ok`;
`exfat-removal` yanks the exFAT stick mid-write while the FAT boot volume keeps
serving.
**Top risks.** Allocation authority is the bitmap, not the FAT — allocating
without setting the bit silently corrupts free space (highest-attention area).
A directory-entry SET can straddle sector/cluster boundaries — scanDirectory,
set-checksum, and updateStreamEntry must handle multi-sector sets. vfs offsets
are u32 while exFAT DataLength is u64 — clamp and document (as fat does). The
in-VM drill needs S3 (a non-boot exFAT volume beside the FAT boot volume); if S4
landed before S3 the discrimination would rest on host tests until multi-volume
exists.
---
## S5 — removal robustness
**Goal.** Close the three known gaps so every removal trigger is exercised
end-to-end. (1) **Consume `medium_changed`** — the VM subscribes to the driver's
already-published event so the "device stays, medium leaves" case (a card
reader, an ejected removable) runs the same kill-retire-remount path as a pulled
stick, closing the second of the "two triggers, one lifecycle" the architecture
specifies. (2) **Storage-driver-crash rebuild** — a driver that dies while its
device stays present is detected and the volume subtree rebuilt on the restarted
driver's fresh channel, instead of leaving fat wedged on a dead channel (the V4
review's open edge). (3) **QEMU-verified remount-on-replug** — the device-return
half is proven, not merely asserted-unmount. Single-volume; independent of S1–S4.
**Key touchpoints.** `library/kernel/service.zig` an additive, behavior-neutral
`on_buffered_message` callback so a buffered-message wake forwards its payload
(no existing service sets it); `volume-manager.zig` subscribe on `bringUpVolume`
success, `onMediumEvent` with change-count dedupe running a medium-teardown (with
`encodeUnsubscribe` before close so the driver's 8-slot table doesn't leak), plus
`channelAlive()` (a `geometry()` liveness probe) + `rebuildVolume()` used in
`pollTick` and the child-exit path; `fat.zig` re-probe geometry on I/O failure
and exit on channel death (device NAK keeps serving); `device-manager.zig` a
`test-storage-restart` mode (mirroring `test-scanout-restart`) to kill usb-storage
once, post-mount, as the discrimination trigger.
**Steps (commits).** (1) **Cheap decisive experiments first** (no commits): QMP-
eject the boot medium and confirm `usb-storage: medium absent` fires under QEMU
(the whole item-1 chain depends on it); and test whether a boot-controller
`device_add` is re-presented (settles whether item 3 extends `volume-removal` or
needs a second controller as H1 does). (2) Harness `on_buffered_message`
(behavior-neutral). (3) VM consumes `medium_changed`. (4) `volume-medium-change`
case (fails pre-step-3). (5) VM driver-crash rebuild. (6) fat observes dead
channel and exits. (7) `volume-driver-restart` trigger + case. (8) `volume-replug`
(second controller if needed). (9) Docs + the real-hardware bench protocol. (10)
Full suite + memory.
**Discrimination.** `volume-medium-change`: an eject with the device left in the
tree unmounts (old VM never subscribes → the event goes to no one → mount
persists). `volume-driver-restart`: killing usb-storage post-mount while its
child stays present triggers a rebuild and a SECOND mount + post-kill read (old
`pollTick` only checks `isDevicePresent`, still true, and restarts fat against
the stale channel → wedge/crash-loop). `volume-replug`: a device return on a
second controller drives a remount (the existing case only ever sees the unmount
half).
**Top risks.** QEMU medium-eject must make TEST UNIT READY report not-ready —
step 1(a) validates this before any code. The op-16 overlap (`medium_changed` ==
`hello` by number) is safe only because async events arrive as `isMessage`
notifications and never reach `Serve.dispatch` — the intercept must run in the
notification branch and never catch a synchronous hello. fat can't today
distinguish EPEER from a device NAK (`CallError` swallows the errno) — the plan
uses a geometry re-probe as the liveness oracle, which is correct but indirect.
---
## Decisions (settled)
Both flagged decisions are settled:
1. **A volume's path is its id; the label is display metadata (S1/S2).** The
mount path is the identity id — the GPT GUID, else a `fat-<serial>` /
`mbr-<sig>-<index>` form — a stable, unique, content-derived handle software
uses. The label (FAT volume label / GPT partition name) is a mutable display
name, NOT in the path; a volume-manager `volumes` verb returns
`{ id, mount_path, label }` so a UI can show the friendly name (the database
id/name split). Same-label-different-id volumes therefore never collide; only
identical ids (dd-clones) hit first-wins-and-log. `volumes.csv` overrides the
path (and optionally the label) for a chosen volume, keyed on id. Makes S1
precede S2 and folds the `/volumes/usb` → id-path fixture + regex migration
into S2 step 5.
2. **exFAT is implemented in full (S4).** A complete exFAT: full read and write,
directories, rename, and the on-disk up-case table for correct case-folding —
not a read-first or ASCII-only subset. The one remaining limit is the vfs
protocol's u32 file-offset surface, which caps addressable file size at 4 GiB
for ALL filesystems (FAT included); widening it to u64 is a separate vfs-
protocol change, flagged but out of the exFAT engine's scope.
The ~26 smaller design-time questions are settled with the recommended default
in the phase text (defer GPT entry-array CRC to correctness-only; GUID key =
little-endian u128 pinned now; share the DOS date-time helper into a library
module both engines import; a second removable usb-storage device for the exFAT
drill; VM-poll `channelAlive()` as the load-bearing crash-detection guarantee).