docs: multi-volume is built — storage architecture + rationale (S3)

Flip the storage docs from "one FAT volume today" to the built state:
the manager adopts every device, probes each device's whole partition
table, and spawns one range-confined FAT per volume — several volumes
across several devices, or several partitions on one device's channel.
The boot volume is identified by content; a filesystem installs the
/system rewrites only when it resolves /system/configuration on its own
media. No filesystem binds a shared service name any more — clients route
through the kernel mount table.

Correct two now-stale claims in the rationale, in the honest direction:
"multi-volume providers are reserved, not implemented" becomes built
(multi-namespace-per-provider is the untried NVMe case); and the cloned-
duplicate "second mounts suffixed" was never built — S3 mounts both,
they collide on the shared content id-path (last wins), each boot claim
logged, and distinguishing them is S4 arbitration.
This commit is contained in:
Daniel Samson
2026-08-10 02:35:43 +01:00
parent d59279422e
commit 0b25cd2c94
2 changed files with 38 additions and 22 deletions
@@ -11,15 +11,21 @@
> identity ladder (GPT GUID + name, FAT serial + label, MBR), and the mount map:
> `filesystems.csv` (signature → binary) + `volumes.csv` (identity → optional
> override), a volume's mount path IS its content id (`/volumes/<id>`), with the
> label as display metadata a `volumes` query returns. **Still pending**: the
> `filesystem UUID` rung (needs a non-FAT engine), multi-volume (one FAT volume
> today; fat's boot rewrites are unconditional until S3 makes them
> content-conditional), the volume manager *consuming* `medium_changed` (removal
> is detected by device-presence polling; the event is published but only a
> card-reader medium change needs the subscription), and the remount-on-replug
> label as display metadata a `volumes` query returns. Multi-volume is **built**:
> the manager adopts every storage device, probes each device's whole partition
> table, and spawns one range-confined FAT per volume — several volumes across
> several devices, or several partitions sharing one device's channel — each at
> its own `/volumes/<id>` path with its own supervision. The boot volume is
> identified by **content** (a volume backs `/system/configuration` + `/system/logs`
> only when it resolves `/system/configuration` on its own media), so it works as
> any partition of any device. **Still pending**: the `filesystem UUID` rung and
> a second engine (exFAT, S4); the volume manager *consuming* `medium_changed`
> (removal is detected by device-presence polling; the event is published but only
> a card-reader medium change needs the subscription); the remount-on-replug
> end-to-end (the logic is in place; QEMU can't re-present the boot-controller
> device, so it is bench-verified). A few markers below are left where a duty is
> still pending.
> device, so it is bench-verified); and arbitration when two volumes both resolve
> the boot markers (S3 mounts both and logs each claim; picking one is S4). A few
> markers below are left where a duty is still pending.
## The model
@@ -121,12 +127,16 @@ its own mounts with the kernel; its write cache lives inside the process, so a
write error is observed by the code that owns the volume and surfaces on the
owning channel (the anti-fsyncgate rule — never a system-wide dirty pool).
*(Built:)* fat receives its mount path as `argv[2]` from the volume manager (the
volume's id-path, e.g. `/volumes/fat-12345678`) and mounts its root there, plus
the two `/system` hierarchy rewrites it installs in place (unconditional this
increment; S3 makes them content-conditional across volumes). It no longer
volume's id-path, e.g. `/volumes/fat-12345678`) and mounts its root there. It
installs the two `/system` hierarchy rewrites (`/system/configuration`,
`/system/logs`) only when it is the boot volume — decided by **content**: it
resolves `/system/configuration` on its own media at mount, so a data volume
mounts at its id-path alone and never shadows the running system. It no longer
self-acquires a volume — the V3b flip made it receive its volume id and block
channel from the volume manager, consistent with "it never discovers devices"
above.
above. Because several volumes now serve at once, no filesystem binds a shared
service name; clients reach each through the kernel mount table (`fs_resolve`
routes by prefix to the backing endpoint).
**Kernel** (mechanism only): the mount table routes paths to backend
endpoints — resolve and redirect, never data. Remount-replace is the restart
@@ -145,9 +145,13 @@ matrix-proven shape; genuinely open.
**The pressure points, honestly:**
1. **Multi-volume providers are reserved, not implemented.** The volume
manager flow assumes one provider, one volume; NVMe namespaces make
endpoint-per-volume real work with hardware demanding it.
1. **Multi-volume is built; multi-namespace-per-provider is untried.** The
volume manager adopts every device and spawns one range-confined FAT per
partition — several volumes across several devices, or several partitions
sharing one device's channel, both proven on USB. What is untried is a single
provider exposing several volumes as *namespaces* (NVMe): the endpoint and
per-badge range machinery generalizes, but no such driver exists yet to
exercise it.
2. **The current transport will bottleneck NVMe.** Synchronous call/reply,
one operation in flight, one bounce buffer — fine for a USB2 stick,
forfeits an NVMe drive's queue depth and per-queue MSI-X. Correctness
@@ -233,13 +237,15 @@ matrix-proven shape; genuinely open.
Consequences, each mechanical once identity keys the map: **moving a drive
to a different port changes nothing** — same identity, same mount point,
whether USB port, hub depth, SATA port, or a stick that left as USB and
returned in a SATA dock; **replug remounts at the same path** (a map lookup
once the map exists; today's single volume re-probes and remounts at the
fixed prefix, and remount-on-replug is bench-verified, not QEMU-tested); **the boot volume** is the
recorded identity of the volume carrying `/system/configuration`, findable
on any port; and **duplicate identity is a policy case, not a surprise** —
two cloned sticks at once: first keeps the mapped name, second mounts
suffixed and is logged loudly, never silently shadowed. Unknown identities
returned in a SATA dock; **replug remounts at the same path** (the id-path is
content-derived, so a volume returns to `/volumes/<id>` wherever it reappears;
remount-on-replug end-to-end is bench-verified, not QEMU-tested, because QEMU
can't re-present the boot-controller device); **the boot volume** is the volume
that resolves `/system/configuration` on its own media, findable on any port or
partition; and **duplicate identity is a known S4 gap** — two cloned sticks
share one content id, so today they collide on `/volumes/<id>` (the kernel
remount-replaces; the last wins) and each boot-volume claim is logged loudly.
Distinguishing them with a suffix is arbitration, deferred to S4. Unknown identities
mount under a derived name (sanitized label, else generated) at
`/volumes/<name>` — the hierarchy's documented home for attached media,
which stands: `/system` is what danos IS; attached media is what it isn't.