docs: transport generality and the media-presence event

The NVMe/SATA assessment folded into the storage discussion: what transfers
untouched (the block contract and everything above it — one driver binary
plus one devices.csv row per transport), NVMe as the shorter stack whose
namespaces are the reserved multi-volume case, AHCI as an open shape choice
(leaning per-port processes, the matrix-proven granularity), and the three
pressure points named honestly: multi-volume is reserved-not-implemented,
synchronous call/reply bottlenecks NVMe until the shm-ring data plane, and
media lifecycle is not device lifecycle.

That last one becomes decision 7 and enters the architecture doc: the
removal path has TWO TRIGGERS, ONE LIFECYCLE — channel death (device leaves)
and a planned pushed medium_changed event (medium leaves, device stays: card
readers and trays, USB ones today), translated by the storage driver from
its transport's native signal, presence never content, consumed by the
volume manager into the same kill-retire-remount path. Without it a swapped
card would be served with the previous card's filesystem state.
This commit is contained in:
Daniel Samson
2026-08-09 15:36:59 +01:00
parent a44b397bed
commit 092817ba2e
2 changed files with 70 additions and 3 deletions
@@ -56,9 +56,11 @@ Three kinds of boundary, deliberately different:
reports children to the device manager, serves the transfer contract to its reports children to the device manager, serves the transfer contract to its
own children's class drivers, routed by lineage. own children's class drivers, routed by lineage.
**Storage driver** (usb-storage, per device): hardware → blocks. Speaks SCSI **Storage driver** (usb-storage, per device; later nvme, ahci): hardware →
Bulk-Only Transport upward from its device; serves the block protocol (five blocks. Speaks its transport upward from its device; serves the block
verbs: geometry, read, write, flush, attach). **Content-blind, permanently**: protocol (geometry, read, write, flush, attach, detach — and, *planned*, the
pushed `medium_changed` presence event, translated from the transport's
native signal). **Content-blind, permanently**:
it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no
GPT, no filesystem magic, ever. *(Planned)* it gains one content-blind GPT, no filesystem magic, ever. *(Planned)* it gains one content-blind
mechanism: **named sub-ranges** — "serve blocks [a, b) as a channel of the mechanism: **named sub-ranges** — "serve blocks [a, b) as a channel of the
@@ -137,6 +139,17 @@ surprise-removal path — kill the filesystem process, retire its mounts,
respawn on return. No half-alive states, no `remount-ro`, no mounts that respawn on return. No half-alive states, no `remount-ro`, no mounts that
error forever (Plan 9's dead-server wart). error forever (Plan 9's dead-server wart).
The path has **two triggers, one lifecycle**: the *device* leaving (the
storage driver dies — channel death, the table below), and the *medium*
leaving while the device stays (an SD card pulled from its reader, an ATAPI
tray opened — including USB card readers today). The second trigger is a
pushed `medium_changed` event on the block protocol *(planned)*: the storage
driver translates its transport's native signal (SCSI UNIT ATTENTION, AHCI
PxSSTS, NVMe namespace-change AER) into presence-changed — never content —
and the volume manager runs the same kill-retire path, then re-probes on
medium return exactly as on device return. Without it, a swapped card would
be served with the previous card's filesystem state.
| Layer | Observes | Must do | Guarantees | | Layer | Observes | Must do | Guarantees |
|---|---|---|---| |---|---|---|---|
| Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick | | Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick |
+54
View File
@@ -112,6 +112,48 @@ with everything else), and the 8-slot mount table gets a declared bound or
growth once volumes multiply. The mount table stays the router; per-process growth once volumes multiply. The mount table stays the router; per-process
namespaces (Plan 9's extra) remain separable future work. namespaces (Plan 9's extra) remain separable future work.
## Transport generality: NVMe and SATA against this design
The layering was chosen transport-agnostic on purpose (the 9front lesson:
filesystem services cannot tell a USB stick from an ATA disk); NVMe and AHCI
are the test of that claim, and they fit — with three named pressure points.
**What transfers untouched:** the block protocol (nothing USB in it), the
volume manager, partitions/ranges, filesystems, mounts, and the whole
enforcement stack — delegation, IOMMU confinement (these are first-party PCI
DMA masters, so confinement applies even more directly than USB),
reap-and-rebuild, the hot-plug lifecycle. Integration cost per transport is
what the architecture promises: one driver binary plus one `devices.csv` row.
**NVMe** shortens the stack — no bus/class split; the driver IS the
controller driver, one hop fewer than USB. Its structural novelty,
**namespaces** (hardware-native multiple volumes behind one controller), is
exactly the case the block protocol reserved on day one and decision 4 below
settles as endpoint-per-volume. **AHCI** is a shape choice, not a problem: an
HBA fronts up to 32 ports plus port multipliers — structurally a bus — so
either mirror USB (ahci-bus + a per-port disk driver: maximum restart
granularity, the proven shape) or mirror NVMe (one driver per controller, one
block channel per port). Leaning per-port processes for consistency with the
matrix-proven shape; genuinely open.
**The pressure points, honestly:**
1. **Multi-volume providers are reserved, not implemented.** The volume
manager flow assumes one provider, one volume; NVMe namespaces make
endpoint-per-volume real work with hardware demanding it.
2. **The current transport will bottleneck NVMe.** Synchronous call/reply,
one operation in flight, one bounce buffer — fine for a USB2 stick,
forfeits an NVMe drive's queue depth and per-queue MSI-X. Correctness
needs nothing; performance is gated on the shm-ring data plane
communication.md already names (Fuchsia's FIFO+VMO is the precedent).
NVMe is not worth building before that milestone.
3. **Media lifecycle ≠ device lifecycle.** ATAPI trays and SD card readers
(USB ones too, today) keep the DEVICE present while the MEDIUM comes and
goes — a removal trigger our channel-death lifecycle does not carry. The
fix is decision 7 below. (Related small item: fat's bounce sizing
hardcodes 512-byte sectors; 4Kn drives need the geometry honored
throughout.)
## The decisions on the table ## The decisions on the table
1. **Partitions as driver-side sub-ranges** (mechanism in usb-storage, parsing 1. **Partitions as driver-side sub-ranges** (mechanism in usb-storage, parsing
@@ -136,3 +178,15 @@ namespaces (Plan 9's extra) remain separable future work.
names it as the 256-byte ceiling's unlock — Fuchsia's FIFO+VMO is the names it as the 256-byte ceiling's unlock — Fuchsia's FIFO+VMO is the
precedent); format-level crash honesty (a Power-Safe-style journaling or COW precedent); format-level crash honesty (a Power-Safe-style journaling or COW
filesystem) once danos outgrows FAT; per-process namespaces. filesystem) once danos outgrows FAT; per-process namespaces.
7. **The media-presence event** (settled in principle; lands with the volume
manager): the block protocol gains a pushed event — `medium_changed`, with
present/absent and a change counter — produced by the storage driver from
its transport's native signal (SCSI UNIT ATTENTION / TEST UNIT READY for
USB and ATAPI, PxSSTS for AHCI, namespace-change AER for NVMe) and
consumed by the volume manager, which runs the SAME kill-retire-remount
path it runs on channel death — one lifecycle, two triggers. The driver
reports presence, never content; a pushed event carries no capability,
which the kernel already guarantees. The device staying while its medium
leaves is the one removable-media case the channel-death trigger cannot
see; without this event a swapped SD card would be served with the old
card's filesystem state.