docs: correct the storage docs' V0-V4 status — the V5 flip missed several

A V5 close-out audit (docs against the actual code) found the storage docs
overclaiming in both directions: the status header was flipped to "built" but
several body markers were not, and two passages describe mechanisms the code
never implemented. Nine confirmed, adversarially verified against the source:

Underclaims (marked Planned, actually built):
- storage-architecture: driver named sub-ranges / range confinement (V2, tested
  by the block-range case); the pushed medium_changed event (published today
  from a TEST UNIT READY poll); ownership-gated fs_unmount (V0, process.zig
  gates it with EPERM); the filesystem-harness extraction (V1). Scoped the
  remaining *planned* to the genuinely-pending parts (native-signal translation,
  the volume manager consuming medium_changed).

Overclaims (described, never built):
- storage-architecture: the "FAT dirty flag on disk" guarantee — no on-disk
  dirty/clean-shutdown bit exists; only an in-memory device-dirty bool gating a
  device write-cache flush on close. Fixed in all three places.
- storage-architecture: fat "acquires its own volume (first mass-storage child
  by enumeration order)" — the V3b flip removed self-acquisition; fat is handed
  its volume id and channel by the volume manager.
- rationale + plan: the FAT engine's base_lba "deleted rather than moved" — it
  and the engine's MBR walk still exist as now-inert legacy; the authoritative
  walk lives in partition.zig.
- rationale: NVMe namespaces "decision 4 settles as endpoint-per-volume" —
  decision 4 settles the opposite (per-sender confinement, one endpoint);
  endpoint-per-volume is named only as an unbuilt future refactor.
- rationale: the five-rung identity ladder and volumes.csv map stated in flat
  present tense — only rung 4 (MBR signature + index) is built; added the
  build-status hedge and marked each rung.

Docs only; no code or behavior change. Suite unaffected (127/127).
This commit is contained in:
Daniel Samson
2026-08-09 21:19:29 +01:00
parent 68e65803eb
commit 8216be991d
3 changed files with 64 additions and 40 deletions
@@ -65,12 +65,13 @@ own children's class drivers, routed by lineage.
**Storage driver** (usb-storage, per device; later nvme, ahci): hardware → **Storage driver** (usb-storage, per device; later nvme, ahci): hardware →
blocks. Speaks its transport upward from its device; serves the block blocks. Speaks its transport upward from its device; serves the block
protocol (geometry, read, write, flush, attach, detach — and, *planned*, the protocol (geometry, read, write, flush, attach, detach — and the pushed
pushed `medium_changed` presence event, translated from the transport's `medium_changed` presence event, published today from a slow TEST UNIT READY
native signal). **Content-blind, permanently**: poll; *planned*: translating it from the transport's native signal instead of
polling). **Content-blind, permanently**:
it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no
GPT, no filesystem magic, ever. *(Planned)* it gains one content-blind GPT, no filesystem magic, ever. It has one content-blind mechanism *(built)*:
mechanism: **named sub-ranges** — "serve blocks [a, b) as a channel of the **named sub-ranges** — "serve blocks [a, b) as a channel of the
same block contract". It clamps and offsets; it never knows the numbers came same block contract". It clamps and offsets; it never knows the numbers came
from a partition table. The clamp lives here and nowhere else because a from a partition table. The clamp lives here and nowhere else because a
channel must carry exactly the authority it grants: handing a filesystem the channel must carry exactly the authority it grants: handing a filesystem the
@@ -113,16 +114,20 @@ receives its block channel at spawn; it never discovers devices. It registers
its own mounts with the kernel; its write cache lives inside the process, so a its own mounts with the kernel; its write cache lives inside the process, so a
write error is observed by the code that owns the volume and surfaces on the write error is observed by the code that owns the volume and surfaces on the
owning channel (the anti-fsyncgate rule — never a system-wide dirty pool). owning channel (the anti-fsyncgate rule — never a system-wide dirty pool).
*(Today, interim:)* fat also acquires its own volume (first mass-storage child *(Today, interim:)* fat still hardcodes its mount prefixes (`/volumes/usb` plus
by enumeration order) and hardcodes its mount prefixes; both migrate to the the two boot-volume hierarchy subtrees it rewrites in place); a `volumes.csv`
volume manager. mount map will migrate that to the volume manager. It no longer self-acquires a
volume — the V3b flip made it receive its volume id at spawn and its block
channel from the volume manager's hello reply, consistent with "it never
discovers devices" above.
**Kernel** (mechanism only): the mount table routes paths to backend **Kernel** (mechanism only): the mount table routes paths to backend
endpoints — resolve and redirect, never data. Remount-replace is the restart endpoints — resolve and redirect, never data. Remount-replace is the restart
story; a dead backend's slot is swept lazily on the next resolution. story; a dead backend's slot is swept lazily on the next resolution.
*(Planned fix:)* `fs_unmount` gains ownership — only the mounting endpoint's `fs_unmount` is ownership-gated *(built, V0)* — only the mounting task may
holder may unmount; today it is ungated, which is safe with one mount owner unmount its own prefix; anyone else is refused (`EPERM`). No dead-owner
and wrong with several. exception: a dead owner's slot is swept lazily by resolution, and restart goes
through remount-replace, never through a stranger's unmount.
## Adding a filesystem ## Adding a filesystem
@@ -135,8 +140,9 @@ and wrong with several.
2. **Reuse the shell**: the filesystem harness — establishment, the 2. **Reuse the shell**: the filesystem harness — establishment, the
badge-scoped open-node table, the nine vfs-protocol handlers, mount badge-scoped open-node table, the nine vfs-protocol handlers, mount
registration, the removal path — is shared code, not per-filesystem code. registration, the removal path — is shared code, not per-filesystem code.
*(Planned: extracted from fat's 434-line shell into a library before the *(Done: extracted from fat's original 434-line shell into
second engine is written.)* An engine plus a `main` wiring it into the `library/kernel/file-system-harness.zig`; fat imports it and instantiates
`harness.Server(engine.FileSystem)`.)* An engine plus a `main` wiring it into the
harness is a complete filesystem service. harness is a complete filesystem service.
3. **Add the configuration row**: one line in `filesystems.csv` mapping the 3. **Add the configuration row**: one line in `filesystems.csv` mapping the
on-disk signature to the binary. No other component changes: the volume on-disk signature to the binary. No other component changes: the volume
@@ -162,13 +168,15 @@ error forever (Plan 9's dead-server wart).
The path has **two triggers, one lifecycle**: the *device* leaving (the The path has **two triggers, one lifecycle**: the *device* leaving (the
storage driver dies — channel death, the table below), and the *medium* storage driver dies — channel death, the table below), and the *medium*
leaving while the device stays (an SD card pulled from its reader, an ATAPI leaving while the device stays (an SD card pulled from its reader, an ATAPI
tray opened — including USB card readers today). The second trigger is a tray opened — including USB card readers today). The second trigger is the
pushed `medium_changed` event on the block protocol *(planned)*: the storage pushed `medium_changed` event on the block protocol — published today from a
driver translates its transport's native signal (SCSI UNIT ATTENTION, AHCI TEST UNIT READY poll; still *planned* is the volume manager *consuming* it
PxSSTS, NVMe namespace-change AER) into presence-changed — never content — (today removal is driven only by device-presence polling) and translating the
and the volume manager runs the same kill-retire path, then re-probes on transport's native signal (SCSI UNIT ATTENTION, AHCI PxSSTS, NVMe
medium return exactly as on device return. Without it, a swapped card would namespace-change AER) in place of the poll. On the event the volume manager
be served with the previous card's filesystem state. runs the same kill-retire path, then re-probes on medium return exactly as on
device return. Without it, a swapped card would be served with the previous
card's filesystem state.
| Layer | Observes | Must do | Guarantees | | Layer | Observes | Must do | Guarantees |
|---|---|---|---| |---|---|---|---|
@@ -176,7 +184,7 @@ be served with the previous card's filesystem state.
| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds | | Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang | | Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
| Volume manager *(removal built; remount bench-pending)* | the storage device leaving the device-manager tree (poll) | kill the filesystem service of that device's volume; its kernel mounts retire | one removal path; mounts never dangle; log persistence stops *cleanly* | | Volume manager *(removal built; remount bench-pending)* | the storage device leaving the device-manager tree (poll) | kill the filesystem service of that device's volume; its kernel mounts retire | one removal path; mounts never dangle; log persistence stops *cleanly* |
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the FAT dirty flag on disk marks the unclean removal; the process never serves from behind a dead channel | | Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the unflushed write-back window is dropped on a surprise yank — danos writes no on-disk dirty/clean-shutdown marker today; the process never serves from behind a dead channel |
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated `fs_unmount` (built, V0) | resolution under a dead mount is `not_found`, not a hang | | Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated `fs_unmount` (built, V0) | resolution under a dead mount is `not_found`, not a hang |
| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media | | Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
@@ -222,7 +230,8 @@ never abandoned after the first `not_found`.
What is lost on a surprise yank is exactly the write-back window of the What is lost on a surprise yank is exactly the write-back window of the
filesystem service, no more: the engine owns its cache, so the blast radius of filesystem service, no more: the engine owns its cache, so the blast radius of
a yank is one volume's unflushed writes, marked by the on-disk dirty flag. A a yank is one volume's unflushed writes, which are simply lost — danos records
no on-disk dirty/clean-shutdown bit yet. A
filesystem format with better crash honesty (journaling, copy-on-write — the filesystem format with better crash honesty (journaling, copy-on-write — the
lesson of QNX's Power-Safe) narrows that window further and slots in as lesson of QNX's Power-Safe) narrows that window further and slots in as
implementation N+1 through the table above, changing nothing else. implementation N+1 through the table above, changing nothing else.
@@ -249,8 +258,9 @@ the device lifecycle already uses, mapped one-to-one:
obsolescence (the reap argument, proven on the device tree). obsolescence (the reap argument, proven on the device tree).
3. **The shared harness is how every filesystem inherits the lifecycle by 3. **The shared harness is how every filesystem inherits the lifecycle by
construction.** The harness — not the engine — owns the state machine: construction.** The harness — not the engine — owns the state machine:
establishment at spawn, mount registration, the dirty-flag set/clear establishment at spawn, mount registration, the flush-on-close hook (an
bracket, error-out-and-exit on channel death. The engine sits behind the in-memory device-dirty check that commits the device write cache on close),
error-out-and-exit on channel death. The engine sits behind the
four-function vtable and never sees a channel; it cannot opt out of the four-function vtable and never sees a channel; it cannot opt out of the
lifecycle for the same reason it cannot find a device. This is why the lifecycle for the same reason it cannot find a device. This is why the
harness is extracted BEFORE the second engine is written. harness is extracted BEFORE the second engine is written.
@@ -105,8 +105,9 @@ and /system/logs), closing the two-sticks question honestly.
Plan 9's `dossrv` to Minix to Fuchsia: block-client + engine + file-protocol Plan 9's `dossrv` to Minix to Fuchsia: block-client + engine + file-protocol
provider in one binary, one process per volume (9front practice; per-volume provider in one binary, one process per volume (9front practice; per-volume
fault isolation is what our supervision makes cheap). fat's shell becomes a fault isolation is what our supervision makes cheap). fat's shell becomes a
shared *filesystem harness* library before a second engine is written; the MBR shared *filesystem harness* library before a second engine is written; a
walk moves out of the engine into the volume manager; write caching stays partition walk is added in the volume manager (`partition.zig`) — the engine's
own MBR walk currently remains alongside it; write caching stays
inside the process (the anti-fsyncgate rule). Each mounts its prefixes into the inside the process (the anti-fsyncgate rule). Each mounts its prefixes into the
kernel mount table itself, exactly as today. kernel mount table itself, exactly as today.
@@ -132,8 +133,10 @@ what the architecture promises: one driver binary plus one `devices.csv` row.
**NVMe** shortens the stack — no bus/class split; the driver IS the **NVMe** shortens the stack — no bus/class split; the driver IS the
controller driver, one hop fewer than USB. Its structural novelty, controller driver, one hop fewer than USB. Its structural novelty,
**namespaces** (hardware-native multiple volumes behind one controller), is **namespaces** (hardware-native multiple volumes behind one controller), is
exactly the case the block protocol reserved on day one and decision 4 below exactly the case the block protocol reserved on day one — still reserved, not
settles as endpoint-per-volume. **AHCI** is a shape choice, not a problem: an implemented (pressure point 1 below): decision 4 settles per-volume addressing
as per-sender confinement on a single endpoint, and names endpoint-per-volume
only as an unbuilt future refactor. **AHCI** is a shape choice, not a problem: an
HBA fronts up to 32 ports plus port multipliers — structurally a bus — so HBA fronts up to 32 ports plus port multipliers — structurally a bus — so
either mirror USB (ahci-bus + a per-port disk driver: maximum restart either mirror USB (ahci-bus + a per-port disk driver: maximum restart
granularity, the proven shape) or mirror NVMe (one driver per controller, one granularity, the proven shape) or mirror NVMe (one driver per controller, one
@@ -180,7 +183,10 @@ matrix-proven shape; genuinely open.
same pattern applied to blocks. The volume manager sets each filesystem same pattern applied to blocks. The volume manager sets each filesystem
process's range on the driver; the driver clamps AND translates every process's range on the driver; the driver clamps AND translates every
transfer by the sender's range, so filesystems address volume-relative transfer by the sender's range, so filesystems address volume-relative
LBAs from 0 and the FAT engine's `base_lba` is deleted rather than moved. LBAs from 0. On the confined path the FAT engine's `base_lba` therefore
resolves to 0 on every access; the field and the engine's own MBR walk
remain in `engine.zig` as now-inert legacy code (the authoritative partition
walk lives in the volume manager's `partition.zig`), not yet deleted.
The enforcement point (the clamp at the provider, never in the consumer) The enforcement point (the clamp at the provider, never in the consumer)
is what carries the security property; endpoint-per-volume would deliver is what carries the security property; endpoint-per-volume would deliver
the same property only by inventing a multi-endpoint harness the pattern the same property only by inventing a multi-endpoint harness the pattern
@@ -209,20 +215,26 @@ matrix-proven shape; genuinely open.
broke whenever a drive changed ports or enumeration order; `UUID=` entries broke whenever a drive changed ports or enumeration order; `UUID=` entries
exist because device-path identity failed. danos skips that era: the mount exist because device-path identity failed. danos skips that era: the mount
map (`volumes.csv` — configuration, read by the volume manager) keys on map (`volumes.csv` — configuration, read by the volume manager) keys on
**content identity, never port or discovery order**. The prober reads **content identity, never port or discovery order**. Build status: only
identity off the medium, strongest first: rung 4 (MBR signature + partition index) is implemented today; the fuller
1. GPT partition GUID — 128-bit, unique, stable for the volume's life; rungs and the `volumes.csv` map itself land with the identity ladder, so
2. filesystem UUID (ext-family and most modern formats, in the superblock); today a single volume mounts at the fixed `/volumes/usb` and its recorded
identity is not yet consulted to pick a path. The target ladder the prober
reads off the medium, strongest first:
1. GPT partition GUID — 128-bit, unique, stable for the volume's life *(planned)*;
2. filesystem UUID (ext-family and most modern formats, in the superblock) *(planned)*;
3. FAT volume serial + label — 32 bits, weak (dd-cloned sticks share it) 3. FAT volume serial + label — 32 bits, weak (dd-cloned sticks share it)
but what real sticks carry; but what real sticks carry *(planned)*;
4. MBR disk signature + partition index; 4. MBR disk signature + partition index — **built**; a bare FAT with no
5. nothing — an anonymous volume: generated mount name, no persistence. table takes index 0 over the whole device;
5. nothing — an anonymous volume: generated mount name, no persistence *(planned)*.
Consequences, each mechanical once identity keys the map: **moving a drive Consequences, each mechanical once identity keys the map: **moving a drive
to a different port changes nothing** — same identity, same mount point, to a different port changes nothing** — same identity, same mount point,
whether USB port, hub depth, SATA port, or a stick that left as USB and whether USB port, hub depth, SATA port, or a stick that left as USB and
returned in a SATA dock; **replug remounts at the same path** (the returned in a SATA dock; **replug remounts at the same path** (a map lookup
remount-after-return story is a map lookup); **the boot volume** is the once the map exists; today's single volume re-probes and remounts at the
fixed prefix, and remount-on-replug is bench-verified, not QEMU-tested); **the boot volume** is the
recorded identity of the volume carrying `/system/configuration`, findable recorded identity of the volume carrying `/system/configuration`, findable
on any port; and **duplicate identity is a policy case, not a surprise** — on any port; and **duplicate identity is a policy case, not a surprise** —
two cloned sticks at once: first keeps the mapped name, second mounts two cloned sticks at once: first keeps the mapped name, second mounts
+4 -2
View File
@@ -12,8 +12,10 @@ range confinement at the provider on one serving endpoint — the badge-scoped
provider pattern the xHCI bus already uses, applied to blocks. The volume provider pattern the xHCI bus already uses, applied to blocks. The volume
manager sets each filesystem process's range; the driver clamps and translates manager sets each filesystem process's range; the driver clamps and translates
every transfer by the sender's kernel-stamped badge; filesystems address every transfer by the sender's kernel-stamped badge; filesystems address
volume-relative LBAs from 0 and the FAT engine's `base_lba` is deleted rather volume-relative LBAs from 0 (on the confined path the FAT engine's `base_lba`
than moved. Every other decision the phases below execute is recorded in the resolves to 0; the field and the engine's own MBR walk remain as now-inert
legacy, the authoritative walk living in the volume manager's `partition.zig`).
Every other decision the phases below execute is recorded in the
rationale (decisions 1–8); nothing in this plan waits on a choice. rationale (decisions 1–8); nothing in this plan waits on a choice.
## V0 — `fs_unmount` ownership (the defect fix) ## V0 — `fs_unmount` ownership (the defect fix)