docs: multi-volume is built — storage architecture + rationale (S3)

Flip the storage docs from "one FAT volume today" to the built state:
the manager adopts every device, probes each device's whole partition
table, and spawns one range-confined FAT per volume — several volumes
across several devices, or several partitions on one device's channel.
The boot volume is identified by content; a filesystem installs the
/system rewrites only when it resolves /system/configuration on its own
media. No filesystem binds a shared service name any more — clients route
through the kernel mount table.

Correct two now-stale claims in the rationale, in the honest direction:
"multi-volume providers are reserved, not implemented" becomes built
(multi-namespace-per-provider is the untried NVMe case); and the cloned-
duplicate "second mounts suffixed" was never built — S3 mounts both,
they collide on the shared content id-path (last wins), each boot claim
logged, and distinguishing them is S4 arbitration.
This commit is contained in:
Daniel Samson
2026-08-10 02:35:43 +01:00
parent d59279422e
commit 0b25cd2c94
2 changed files with 38 additions and 22 deletions
@@ -11,15 +11,21 @@
> identity ladder (GPT GUID + name, FAT serial + label, MBR), and the mount map: > identity ladder (GPT GUID + name, FAT serial + label, MBR), and the mount map:
> `filesystems.csv` (signature → binary) + `volumes.csv` (identity → optional > `filesystems.csv` (signature → binary) + `volumes.csv` (identity → optional
> override), a volume's mount path IS its content id (`/volumes/<id>`), with the > override), a volume's mount path IS its content id (`/volumes/<id>`), with the
> label as display metadata a `volumes` query returns. **Still pending**: the > label as display metadata a `volumes` query returns. Multi-volume is **built**:
> `filesystem UUID` rung (needs a non-FAT engine), multi-volume (one FAT volume > the manager adopts every storage device, probes each device's whole partition
> today; fat's boot rewrites are unconditional until S3 makes them > table, and spawns one range-confined FAT per volume — several volumes across
> content-conditional), the volume manager *consuming* `medium_changed` (removal > several devices, or several partitions sharing one device's channel — each at
> is detected by device-presence polling; the event is published but only a > its own `/volumes/<id>` path with its own supervision. The boot volume is
> card-reader medium change needs the subscription), and the remount-on-replug > identified by **content** (a volume backs `/system/configuration` + `/system/logs`
> only when it resolves `/system/configuration` on its own media), so it works as
> any partition of any device. **Still pending**: the `filesystem UUID` rung and
> a second engine (exFAT, S4); the volume manager *consuming* `medium_changed`
> (removal is detected by device-presence polling; the event is published but only
> a card-reader medium change needs the subscription); the remount-on-replug
> end-to-end (the logic is in place; QEMU can't re-present the boot-controller > end-to-end (the logic is in place; QEMU can't re-present the boot-controller
> device, so it is bench-verified). A few markers below are left where a duty is > device, so it is bench-verified); and arbitration when two volumes both resolve
> still pending. > the boot markers (S3 mounts both and logs each claim; picking one is S4). A few
> markers below are left where a duty is still pending.
## The model ## The model
@@ -121,12 +127,16 @@ its own mounts with the kernel; its write cache lives inside the process, so a
write error is observed by the code that owns the volume and surfaces on the write error is observed by the code that owns the volume and surfaces on the
owning channel (the anti-fsyncgate rule — never a system-wide dirty pool). owning channel (the anti-fsyncgate rule — never a system-wide dirty pool).
*(Built:)* fat receives its mount path as `argv[2]` from the volume manager (the *(Built:)* fat receives its mount path as `argv[2]` from the volume manager (the
volume's id-path, e.g. `/volumes/fat-12345678`) and mounts its root there, plus volume's id-path, e.g. `/volumes/fat-12345678`) and mounts its root there. It
the two `/system` hierarchy rewrites it installs in place (unconditional this installs the two `/system` hierarchy rewrites (`/system/configuration`,
increment; S3 makes them content-conditional across volumes). It no longer `/system/logs`) only when it is the boot volume — decided by **content**: it
resolves `/system/configuration` on its own media at mount, so a data volume
mounts at its id-path alone and never shadows the running system. It no longer
self-acquires a volume — the V3b flip made it receive its volume id and block self-acquires a volume — the V3b flip made it receive its volume id and block
channel from the volume manager, consistent with "it never discovers devices" channel from the volume manager, consistent with "it never discovers devices"
above. above. Because several volumes now serve at once, no filesystem binds a shared
service name; clients reach each through the kernel mount table (`fs_resolve`
routes by prefix to the backing endpoint).
**Kernel** (mechanism only): the mount table routes paths to backend **Kernel** (mechanism only): the mount table routes paths to backend
endpoints — resolve and redirect, never data. Remount-replace is the restart endpoints — resolve and redirect, never data. Remount-replace is the restart
@@ -145,9 +145,13 @@ matrix-proven shape; genuinely open.
**The pressure points, honestly:** **The pressure points, honestly:**
1. **Multi-volume providers are reserved, not implemented.** The volume 1. **Multi-volume is built; multi-namespace-per-provider is untried.** The
manager flow assumes one provider, one volume; NVMe namespaces make volume manager adopts every device and spawns one range-confined FAT per
endpoint-per-volume real work with hardware demanding it. partition — several volumes across several devices, or several partitions
sharing one device's channel, both proven on USB. What is untried is a single
provider exposing several volumes as *namespaces* (NVMe): the endpoint and
per-badge range machinery generalizes, but no such driver exists yet to
exercise it.
2. **The current transport will bottleneck NVMe.** Synchronous call/reply, 2. **The current transport will bottleneck NVMe.** Synchronous call/reply,
one operation in flight, one bounce buffer — fine for a USB2 stick, one operation in flight, one bounce buffer — fine for a USB2 stick,
forfeits an NVMe drive's queue depth and per-queue MSI-X. Correctness forfeits an NVMe drive's queue depth and per-queue MSI-X. Correctness
@@ -233,13 +237,15 @@ matrix-proven shape; genuinely open.
Consequences, each mechanical once identity keys the map: **moving a drive Consequences, each mechanical once identity keys the map: **moving a drive
to a different port changes nothing** — same identity, same mount point, to a different port changes nothing** — same identity, same mount point,
whether USB port, hub depth, SATA port, or a stick that left as USB and whether USB port, hub depth, SATA port, or a stick that left as USB and
returned in a SATA dock; **replug remounts at the same path** (a map lookup returned in a SATA dock; **replug remounts at the same path** (the id-path is
once the map exists; today's single volume re-probes and remounts at the content-derived, so a volume returns to `/volumes/<id>` wherever it reappears;
fixed prefix, and remount-on-replug is bench-verified, not QEMU-tested); **the boot volume** is the remount-on-replug end-to-end is bench-verified, not QEMU-tested, because QEMU
recorded identity of the volume carrying `/system/configuration`, findable can't re-present the boot-controller device); **the boot volume** is the volume
on any port; and **duplicate identity is a policy case, not a surprise** — that resolves `/system/configuration` on its own media, findable on any port or
two cloned sticks at once: first keeps the mapped name, second mounts partition; and **duplicate identity is a known S4 gap** — two cloned sticks
suffixed and is logged loudly, never silently shadowed. Unknown identities share one content id, so today they collide on `/volumes/<id>` (the kernel
remount-replaces; the last wins) and each boot-volume claim is logged loudly.
Distinguishing them with a suffix is arbitration, deferred to S4. Unknown identities
mount under a derived name (sanitized label, else generated) at mount under a derived name (sanitized label, else generated) at
`/volumes/<name>` — the hierarchy's documented home for attached media, `/volumes/<name>` — the hierarchy's documented home for attached media,
which stands: `/system` is what danos IS; attached media is what it isn't. which stands: `/system` is what danos IS; attached media is what it isn't.