docs: storage plan — volumes named by identity; exFAT implemented in full

Two decisions settled: (1) a volume's mount name is its own content identity
(FAT label / GPT name, else identity-hex), never a port name (/volumes/usb) or
role name (/volumes/boot); volumes.csv stays as an explicit override. This makes
S1 precede S2 and folds the /volumes/usb fixture+regex migration into S2. (2)
exFAT is a complete implementation — full read+write, directories, rename, and
the on-disk up-case table — not a read-first/ASCII-only subset; the only limit
is the vfs u32 offset surface (a 4 GiB cap on all filesystems), flagged as a
separate vfs change.
This commit is contained in:
Daniel Samson
2026-08-09 22:42:56 +01:00
parent fca41b351e
commit addd264880
+68 -57
View File
@@ -26,24 +26,22 @@ bench-pending.
## The five phases and how they depend
```
S1 identity ladder ─────┐
├──► S2 mount map ──► S3 multi-volume ──► S4 exFAT
(S2 keys on rung-4, so │ (needs S2+S3)
is independent of S1) │
└──► S5 removal robustness (independent; single-volume)
S1 identity ladder ──► S2 mount map ──► S3 multi-volume ──► S4 exFAT
(needs S2+S3)
S5 removal robustness ── independent; single-volume ── may land any time
```
- **S1** grows the identity read off the medium (GPT GUID, FAT serial+label).
- **S1** grows the identity read off the medium (GPT GUID, FAT serial+label). It
comes FIRST because a volume's mount name is now its identity (below), and the
friendly form of that name is the FAT label / GPT name S1 parses.
- **S2** moves the last policy out of hardcode into `volumes.csv` +
`filesystems.csv`, keyed on today's rung-4 identity — so it does **not** wait
on S1; S1 only enriches the identity values the same map consumes.
`filesystems.csv`, and names each volume by its identity — no port-name.
- **S3** generalizes to N volumes across N devices.
- **S4** adds exFAT — the second engine that proves the V1 harness extraction.
Needs S2 (to route by signature) and S3 (to run a second volume in the VM).
- **S4** adds exFAT — a COMPLETE second engine that proves the V1 harness
extraction. Needs S2 (to route by signature) and S3 (to run a second volume).
- **S5** closes the removal-lifecycle gaps. Independent of the rest; single-volume.
Recommended build order is S1 → S2 → S3 → S4 → S5. S5 may be pulled earlier if
removal robustness matters more than multi-volume; S1 and S2 may swap.
Recommended build order is S1 → S2 → S3 → S4 → S5. S5 may be pulled earlier.
---
@@ -101,12 +99,18 @@ never drift onto magic constants).
**Goal.** Move the last two pieces of storage policy out of hardcode into
configuration read by the volume manager. `filesystems.csv` (content signature →
filesystem binary) so the VM picks the binary from the probed signature;
`volumes.csv` (identity → mount prefix(es), danos's fstab) so the VM decides
mount placement, with a `/volumes/vol-<hex>` anonymous fallback. Parsed with
`library/csv` exactly as `device-registry` parses `devices.csv`. The VM hands the
binary + volume-id + mount specs to fat at spawn (the argv channel V3b already
uses for the volume id); fat retires `fat_mounts` and reads mounts from
argv[2..]. Keyed on rung-4 identity, so independent of S1 and forward-compatible.
`volumes.csv` (identity → mount prefix, danos's fstab) as the explicit override
for a volume the user wants at a fixed path. **The default mount name is the
volume's own identity, never a port or role name**: `/volumes/<label>` when the
medium carries a label (the FAT label / GPT partition name S1 reads), else
`/volumes/vol-<hex-key>`; duplicate names follow the rationale's duplicate-
identity policy (first keeps it, second suffixed and logged). There is no
`/volumes/usb` and no `/volumes/boot` — the boot volume is detected by content
(it installs the `/system/configuration` + `/system/logs` rewrites) but is NAMED
by its identity like any other. Parsed with `library/csv` exactly as
`device-registry` parses `devices.csv`. The VM hands the binary + volume-id +
mount specs to fat at spawn (the argv channel V3b already uses for the volume
id); fat retires `fat_mounts` and reads mounts from argv[2..].
**Key touchpoints.** New pure-logic modules `filesystem-map.zig` (parse +
`match(signature)`) and `volume-map.zig` (parse + `mountsFor(identity)` +
@@ -119,26 +123,30 @@ argv; `fat.zig` deletes `fat_mounts`, parses argv[2..] into a bounded
MBR disk signature so the boot volume's identity is a legible non-zero key.
**Steps (commits).** (1) partition emits a signature. (2) `filesystem-map`
parser + host tests. (3) `volume-map` parser + host tests. (4) Ship the tables +
parser + host tests. (3) `volume-map` parser (identity→prefix override) +
identity-derived default naming (label→hex) + host tests. (4) Ship the tables +
VM integration **behind fat's still-hardcoded mounts** (behavior-preserving —
the parsers are proven before consumption flips; full suite green). (5) fat
consumes argv mounts; add the `/volumes/boot` row; the `volume-mapped` QEMU case
(shown failing against HEAD~1); flip the docs' `volumes.csv`/`filesystems.csv`
pending markers.
parsers proven before consumption flips; full suite green). (5) fat consumes
argv mounts; the VM names the boot volume by its identity (`/volumes/DANOS`, from
the FAT label S1 read); **migrate the fixtures + QEMU regexes off `/volumes/usb`**
(fat-test, badge-scope-test, vfs-test, and the four cases) in the same commit;
the `volume-identity-name` QEMU case (shown failing against HEAD~1); flip the
docs' pending markers.
**Discrimination.** QEMU `volume-mapped`: the boot volume, mapped in
`volumes.csv` to `/volumes/boot`, makes fat log `mounted /volumes/boot` — a line
the old hardcoded `fat_mounts` never emits. Host: `mountsFor(mapped)` returns the
CSV prefix, `mountsFor(unmapped)` derives `/volumes/vol-<hex>`; `match(.fat)`
returns the configured binary; fat's argv parser makes installed mounts a
function of argv.
**Discrimination.** QEMU `volume-identity-name`: the boot volume mounts at
`/volumes/DANOS` (its FAT label), a line the old hardcoded `/volumes/usb` never
emits. Host: an unlabeled identity derives `/volumes/vol-<hex>`; a `volumes.csv`
override sends a mapped identity to its chosen prefix; `match(.fat)` returns the
configured binary; fat's argv parser makes installed mounts a function of argv.
**Top risks.** Step 5's blast radius — a parser/argv bug breaks every
fat-dependent case at once; mitigated by landing the VM half behind fat's
hardcoded mounts first (step 4). The argv blob is 256 bytes (process.zig) — cap
emitted mounts and refuse+log on overflow. `/system/logs` is now a
`volumes.csv` row: dropping it silently stops log persistence — ship it by
default and document the boot-identity contract in the CSV header.
**Top risks.** Step 5's blast radius — a parser/argv bug, OR the `/volumes/usb`→
`/volumes/DANOS` migration missing a fixture/regex, breaks every fat-dependent
case at once; mitigated by landing the VM half behind fat's hardcoded mounts
first (step 4) and making the name migration one atomic, complete sweep. The
argv blob is 256 bytes (process.zig) — cap emitted mounts and refuse+log on
overflow. `/system/logs` is now a `volumes.csv` concern: dropping the boot
identity's rewrite rows silently stops log persistence — ship them by default
and document the boot-identity contract in the CSV header.
---
@@ -174,8 +182,8 @@ the test disk artifacts.
**Steps (commits).** (1) `partition.allVolumes` + two-partition host test. (2)
Tables, behavior-preserving (still one device / one volume). (3) Multi-device +
multi-partition. (4) Per-volume mount naming via argv (first = `/volumes/usb`,
rest `/volumes/usb<n>` — bridge until `volumes.csv`). (5) fat content-conditional
multi-partition. (4) Per-volume mount naming via argv — each volume by its
identity (label→hex, from S2), one per volume, no port-name. (5) fat content-conditional
boot mounts. (6) Raise `maximum_mounts`. (7) Partitioned-image tool. (8)
`two-volume` QEMU case. (9) `boot-2nd-partition` case. (10) Adversarial review,
full suite, docs, memory.
@@ -207,12 +215,16 @@ but an engine + a `main`, reusing `library/kernel/file-system-harness.zig`'s
The harness already owns vfs serving, the badge-scoped open-node table,
create-on-open/O_TRUNC, mount registration, the exit sweep, bring-up retry,
per-turn time stamping, and durable-on-close. S4 writes only the exFAT-specific
bits.
bits — but **in full**: complete read AND write, directories, rename, and the
real on-disk up-case table for correct case-folding. Not a read-first, minimal-
write, or ASCII-only subset. The only limit that survives is the vfs protocol's
u32 file-offset surface (a 4 GiB addressable-size cap that applies to FAT too),
which is a separate vfs-protocol change, not an exFAT shortcut.
**Key touchpoints.** New `system/services/exfat/on-disk.zig` (the Main Boot
Sector VBR + the five 32-byte directory-entry types as `align(1)` extern structs;
`geometryOf` accepting only `"EXFAT "` + `0xAA55`; `setChecksum`, `nameHash`,
`upcaseAscii`); `engine.zig` (`FileSystem` behind the identical `BlockDevice`
and case-folding driven by the volume's **on-disk up-case table**); `engine.zig` (`FileSystem` behind the identical `BlockDevice`
vtable with fat's exact method set; **allocation-bitmap** cluster authority — the
deepest departure from FAT; read honoring `no_fat_chain` contiguous vs FAT-follow;
File+Stream+FileName set assembly with recomputed set checksum); `exfat.zig` (the
@@ -303,27 +315,26 @@ uses a geometry re-probe as the liveness oracle, which is correct but indirect.
---
## Decisions embedded in this plan (flag if you disagree)
## Decisions (settled)
Most of the ~28 design-time questions the phase design surfaced are settled with
the recommended default and noted in the phases above. Two are worth your eye:
Both flagged decisions are settled:
1. **`/volumes/usb` naming (S2).** The plan **keeps `/volumes/usb`
transitionally** and adds the mapped `/volumes/boot` — zero fixture churn,
inherent discrimination. The design-purer alternative is to REPLACE
`/volumes/usb` (killing the port-name the rationale condemns) and migrate the
three runtime fixtures + four QEMU regexes in a coupled sweep. Recommendation:
keep transitionally now, drop `/volumes/usb` as a small follow-up once the map
is proven.
1. **Volume names are the identity (S2/S3).** Every volume mounts at
`/volumes/<its-identity>` — the FAT label / GPT name where present, else the
identity-hex — never a port name (`/volumes/usb`) or a role name
(`/volumes/boot`). `volumes.csv` remains the explicit override for a chosen
fixed path. This makes S1 (which reads the label) precede S2, and folds the
`/volumes/usb` → `/volumes/DANOS` fixture + regex migration into S2 step 5.
2. **exFAT write scope (S4).** The plan ships **full mutation parity**
(createFile + writeFile + truncate + removeFile + createDirectory + rename) so
`exfat-test` reuses fat-test's full matrix. The minimum viable is
create/write/truncate/remove only. Recommendation: full parity.
2. **exFAT is implemented in full (S4).** A complete exFAT: full read and write,
directories, rename, and the on-disk up-case table for correct case-folding —
not a read-first or ASCII-only subset. The one remaining limit is the vfs
protocol's u32 file-offset surface, which caps addressable file size at 4 GiB
for ALL filesystems (FAT included); widening it to u64 is a separate vfs-
protocol change, flagged but out of the exFAT engine's scope.
Everything else (defer GPT entry-array CRC to correctness-only; GUID key =
The ~26 smaller design-time questions are settled with the recommended default
in the phase text (defer GPT entry-array CRC to correctness-only; GUID key =
little-endian u128 pinned now; share the DOS date-time helper into a library
module both engines import; a second removable usb-storage device for the exFAT
drill rather than a boot-disk partition; VM-poll `channelAlive()` as the
load-bearing crash-detection guarantee with fat's exit as a promptness
optimization) is taken as the recommended default in the phase text.
drill; VM-poll `channelAlive()` as the load-bearing crash-detection guarantee).