docs: multi-volume is built — storage architecture + rationale (S3)

Flip the storage docs from "one FAT volume today" to the built state:
the manager adopts every device, probes each device's whole partition
table, and spawns one range-confined FAT per volume — several volumes
across several devices, or several partitions on one device's channel.
The boot volume is identified by content; a filesystem installs the
/system rewrites only when it resolves /system/configuration on its own
media. No filesystem binds a shared service name any more — clients route
through the kernel mount table.

Correct two now-stale claims in the rationale, in the honest direction:
"multi-volume providers are reserved, not implemented" becomes built
(multi-namespace-per-provider is the untried NVMe case); and the cloned-
duplicate "second mounts suffixed" was never built — S3 mounts both,
they collide on the shared content id-path (last wins), each boot claim
logged, and distinguishing them is S4 arbitration.
This commit is contained in:
Daniel Samson
2026-08-10 02:35:43 +01:00
parent d59279422e
commit 0b25cd2c94
2 changed files with 38 additions and 22 deletions
@@ -11,15 +11,21 @@
> identity ladder (GPT GUID + name, FAT serial + label, MBR), and the mount map:
> `filesystems.csv` (signature → binary) + `volumes.csv` (identity → optional
> override), a volume's mount path IS its content id (`/volumes/<id>`), with the
> label as display metadata a `volumes` query returns. **Still pending**: the
> `filesystem UUID` rung (needs a non-FAT engine), multi-volume (one FAT volume
> today; fat's boot rewrites are unconditional until S3 makes them
> content-conditional), the volume manager *consuming* `medium_changed` (removal
> is detected by device-presence polling; the event is published but only a
> card-reader medium change needs the subscription), and the remount-on-replug
> label as display metadata a `volumes` query returns. Multi-volume is **built**:
> the manager adopts every storage device, probes each device's whole partition
> table, and spawns one range-confined FAT per volume — several volumes across
> several devices, or several partitions sharing one device's channel — each at
> its own `/volumes/<id>` path with its own supervision. The boot volume is
> identified by **content** (a volume backs `/system/configuration` + `/system/logs`
> only when it resolves `/system/configuration` on its own media), so it works as
> any partition of any device. **Still pending**: the `filesystem UUID` rung and
> a second engine (exFAT, S4); the volume manager *consuming* `medium_changed`
> (removal is detected by device-presence polling; the event is published but only
> a card-reader medium change needs the subscription); the remount-on-replug
> end-to-end (the logic is in place; QEMU can't re-present the boot-controller
> device, so it is bench-verified). A few markers below are left where a duty is
> still pending.
> device, so it is bench-verified); and arbitration when two volumes both resolve
> the boot markers (S3 mounts both and logs each claim; picking one is S4). A few
> markers below are left where a duty is still pending.
## The model
@@ -121,12 +127,16 @@ its own mounts with the kernel; its write cache lives inside the process, so a
write error is observed by the code that owns the volume and surfaces on the
owning channel (the anti-fsyncgate rule — never a system-wide dirty pool).
*(Built:)* fat receives its mount path as `argv[2]` from the volume manager (the
volume's id-path, e.g. `/volumes/fat-12345678`) and mounts its root there, plus
the two `/system` hierarchy rewrites it installs in place (unconditional this
increment; S3 makes them content-conditional across volumes). It no longer
volume's id-path, e.g. `/volumes/fat-12345678`) and mounts its root there. It
installs the two `/system` hierarchy rewrites (`/system/configuration`,
`/system/logs`) only when it is the boot volume — decided by **content**: it
resolves `/system/configuration` on its own media at mount, so a data volume
mounts at its id-path alone and never shadows the running system. It no longer
self-acquires a volume — the V3b flip made it receive its volume id and block
channel from the volume manager, consistent with "it never discovers devices"
above.
above. Because several volumes now serve at once, no filesystem binds a shared
service name; clients reach each through the kernel mount table (`fs_resolve`
routes by prefix to the backing endpoint).
**Kernel** (mechanism only): the mount table routes paths to backend
endpoints — resolve and redirect, never data. Remount-replace is the restart