docs: correct the storage docs' V0-V4 status — the V5 flip missed several

A V5 close-out audit (docs against the actual code) found the storage docs
overclaiming in both directions: the status header was flipped to "built" but
several body markers were not, and two passages describe mechanisms the code
never implemented. Nine confirmed, adversarially verified against the source:

Underclaims (marked Planned, actually built):
- storage-architecture: driver named sub-ranges / range confinement (V2, tested
  by the block-range case); the pushed medium_changed event (published today
  from a TEST UNIT READY poll); ownership-gated fs_unmount (V0, process.zig
  gates it with EPERM); the filesystem-harness extraction (V1). Scoped the
  remaining *planned* to the genuinely-pending parts (native-signal translation,
  the volume manager consuming medium_changed).

Overclaims (described, never built):
- storage-architecture: the "FAT dirty flag on disk" guarantee — no on-disk
  dirty/clean-shutdown bit exists; only an in-memory device-dirty bool gating a
  device write-cache flush on close. Fixed in all three places.
- storage-architecture: fat "acquires its own volume (first mass-storage child
  by enumeration order)" — the V3b flip removed self-acquisition; fat is handed
  its volume id and channel by the volume manager.
- rationale + plan: the FAT engine's base_lba "deleted rather than moved" — it
  and the engine's MBR walk still exist as now-inert legacy; the authoritative
  walk lives in partition.zig.
- rationale: NVMe namespaces "decision 4 settles as endpoint-per-volume" —
  decision 4 settles the opposite (per-sender confinement, one endpoint);
  endpoint-per-volume is named only as an unbuilt future refactor.
- rationale: the five-rung identity ladder and volumes.csv map stated in flat
  present tense — only rung 4 (MBR signature + index) is built; added the
  build-status hedge and marked each rung.

Docs only; no code or behavior change. Suite unaffected (127/127).
This commit is contained in:
Daniel Samson
2026-08-09 21:19:29 +01:00
parent 68e65803eb
commit 8216be991d
3 changed files with 64 additions and 40 deletions
@@ -65,12 +65,13 @@ own children's class drivers, routed by lineage.
**Storage driver** (usb-storage, per device; later nvme, ahci): hardware →
blocks. Speaks its transport upward from its device; serves the block
protocol (geometry, read, write, flush, attach, detach — and, *planned*, the
pushed `medium_changed` presence event, translated from the transport's
native signal). **Content-blind, permanently**:
protocol (geometry, read, write, flush, attach, detach — and the pushed
`medium_changed` presence event, published today from a slow TEST UNIT READY
poll; *planned*: translating it from the transport's native signal instead of
polling). **Content-blind, permanently**:
it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no
GPT, no filesystem magic, ever. *(Planned)* it gains one content-blind
mechanism: **named sub-ranges** — "serve blocks [a, b) as a channel of the
GPT, no filesystem magic, ever. It has one content-blind mechanism *(built)*:
**named sub-ranges** — "serve blocks [a, b) as a channel of the
same block contract". It clamps and offsets; it never knows the numbers came
from a partition table. The clamp lives here and nowhere else because a
channel must carry exactly the authority it grants: handing a filesystem the
@@ -113,16 +114,20 @@ receives its block channel at spawn; it never discovers devices. It registers
its own mounts with the kernel; its write cache lives inside the process, so a
write error is observed by the code that owns the volume and surfaces on the
owning channel (the anti-fsyncgate rule — never a system-wide dirty pool).
*(Today, interim:)* fat also acquires its own volume (first mass-storage child
by enumeration order) and hardcodes its mount prefixes; both migrate to the
volume manager.
*(Today, interim:)* fat still hardcodes its mount prefixes (`/volumes/usb` plus
the two boot-volume hierarchy subtrees it rewrites in place); a `volumes.csv`
mount map will migrate that to the volume manager. It no longer self-acquires a
volume — the V3b flip made it receive its volume id at spawn and its block
channel from the volume manager's hello reply, consistent with "it never
discovers devices" above.
**Kernel** (mechanism only): the mount table routes paths to backend
endpoints — resolve and redirect, never data. Remount-replace is the restart
story; a dead backend's slot is swept lazily on the next resolution.
*(Planned fix:)* `fs_unmount` gains ownership — only the mounting endpoint's
holder may unmount; today it is ungated, which is safe with one mount owner
and wrong with several.
`fs_unmount` is ownership-gated *(built, V0)* — only the mounting task may
unmount its own prefix; anyone else is refused (`EPERM`). No dead-owner
exception: a dead owner's slot is swept lazily by resolution, and restart goes
through remount-replace, never through a stranger's unmount.
## Adding a filesystem
@@ -135,8 +140,9 @@ and wrong with several.
2. **Reuse the shell**: the filesystem harness — establishment, the
badge-scoped open-node table, the nine vfs-protocol handlers, mount
registration, the removal path — is shared code, not per-filesystem code.
*(Planned: extracted from fat's 434-line shell into a library before the
second engine is written.)* An engine plus a `main` wiring it into the
*(Done: extracted from fat's original 434-line shell into
`library/kernel/file-system-harness.zig`; fat imports it and instantiates
`harness.Server(engine.FileSystem)`.)* An engine plus a `main` wiring it into the
harness is a complete filesystem service.
3. **Add the configuration row**: one line in `filesystems.csv` mapping the
on-disk signature to the binary. No other component changes: the volume
@@ -162,13 +168,15 @@ error forever (Plan 9's dead-server wart).
The path has **two triggers, one lifecycle**: the *device* leaving (the
storage driver dies — channel death, the table below), and the *medium*
leaving while the device stays (an SD card pulled from its reader, an ATAPI
tray opened — including USB card readers today). The second trigger is a
pushed `medium_changed` event on the block protocol *(planned)*: the storage
driver translates its transport's native signal (SCSI UNIT ATTENTION, AHCI
PxSSTS, NVMe namespace-change AER) into presence-changed — never content —
and the volume manager runs the same kill-retire path, then re-probes on
medium return exactly as on device return. Without it, a swapped card would
be served with the previous card's filesystem state.
tray opened — including USB card readers today). The second trigger is the
pushed `medium_changed` event on the block protocol — published today from a
TEST UNIT READY poll; still *planned* is the volume manager *consuming* it
(today removal is driven only by device-presence polling) and translating the
transport's native signal (SCSI UNIT ATTENTION, AHCI PxSSTS, NVMe
namespace-change AER) in place of the poll. On the event the volume manager
runs the same kill-retire path, then re-probes on medium return exactly as on
device return. Without it, a swapped card would be served with the previous
card's filesystem state.
| Layer | Observes | Must do | Guarantees |
|---|---|---|---|
@@ -176,7 +184,7 @@ be served with the previous card's filesystem state.
| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
| Volume manager *(removal built; remount bench-pending)* | the storage device leaving the device-manager tree (poll) | kill the filesystem service of that device's volume; its kernel mounts retire | one removal path; mounts never dangle; log persistence stops *cleanly* |
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the FAT dirty flag on disk marks the unclean removal; the process never serves from behind a dead channel |
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the unflushed write-back window is dropped on a surprise yank — danos writes no on-disk dirty/clean-shutdown marker today; the process never serves from behind a dead channel |
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated `fs_unmount` (built, V0) | resolution under a dead mount is `not_found`, not a hang |
| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
@@ -222,7 +230,8 @@ never abandoned after the first `not_found`.
What is lost on a surprise yank is exactly the write-back window of the
filesystem service, no more: the engine owns its cache, so the blast radius of
a yank is one volume's unflushed writes, marked by the on-disk dirty flag. A
a yank is one volume's unflushed writes, which are simply lost — danos records
no on-disk dirty/clean-shutdown bit yet. A
filesystem format with better crash honesty (journaling, copy-on-write — the
lesson of QNX's Power-Safe) narrows that window further and slots in as
implementation N+1 through the table above, changing nothing else.
@@ -249,8 +258,9 @@ the device lifecycle already uses, mapped one-to-one:
obsolescence (the reap argument, proven on the device tree).
3. **The shared harness is how every filesystem inherits the lifecycle by
construction.** The harness — not the engine — owns the state machine:
establishment at spawn, mount registration, the dirty-flag set/clear
bracket, error-out-and-exit on channel death. The engine sits behind the
establishment at spawn, mount registration, the flush-on-close hook (an
in-memory device-dirty check that commits the device write cache on close),
error-out-and-exit on channel death. The engine sits behind the
four-function vtable and never sees a channel; it cannot opt out of the
lifecycle for the same reason it cannot find a device. This is why the
harness is extracted BEFORE the second engine is written.
@@ -105,8 +105,9 @@ and /system/logs), closing the two-sticks question honestly.
Plan 9's `dossrv` to Minix to Fuchsia: block-client + engine + file-protocol
provider in one binary, one process per volume (9front practice; per-volume
fault isolation is what our supervision makes cheap). fat's shell becomes a
shared *filesystem harness* library before a second engine is written; the MBR
walk moves out of the engine into the volume manager; write caching stays
shared *filesystem harness* library before a second engine is written; a
partition walk is added in the volume manager (`partition.zig`) — the engine's
own MBR walk currently remains alongside it; write caching stays
inside the process (the anti-fsyncgate rule). Each mounts its prefixes into the
kernel mount table itself, exactly as today.
@@ -132,8 +133,10 @@ what the architecture promises: one driver binary plus one `devices.csv` row.
**NVMe** shortens the stack — no bus/class split; the driver IS the
controller driver, one hop fewer than USB. Its structural novelty,
**namespaces** (hardware-native multiple volumes behind one controller), is
exactly the case the block protocol reserved on day one and decision 4 below
settles as endpoint-per-volume. **AHCI** is a shape choice, not a problem: an
exactly the case the block protocol reserved on day one — still reserved, not
implemented (pressure point 1 below): decision 4 settles per-volume addressing
as per-sender confinement on a single endpoint, and names endpoint-per-volume
only as an unbuilt future refactor. **AHCI** is a shape choice, not a problem: an
HBA fronts up to 32 ports plus port multipliers — structurally a bus — so
either mirror USB (ahci-bus + a per-port disk driver: maximum restart
granularity, the proven shape) or mirror NVMe (one driver per controller, one
@@ -180,7 +183,10 @@ matrix-proven shape; genuinely open.
same pattern applied to blocks. The volume manager sets each filesystem
process's range on the driver; the driver clamps AND translates every
transfer by the sender's range, so filesystems address volume-relative
LBAs from 0 and the FAT engine's `base_lba` is deleted rather than moved.
LBAs from 0. On the confined path the FAT engine's `base_lba` therefore
resolves to 0 on every access; the field and the engine's own MBR walk
remain in `engine.zig` as now-inert legacy code (the authoritative partition
walk lives in the volume manager's `partition.zig`), not yet deleted.
The enforcement point (the clamp at the provider, never in the consumer)
is what carries the security property; endpoint-per-volume would deliver
the same property only by inventing a multi-endpoint harness the pattern
@@ -209,20 +215,26 @@ matrix-proven shape; genuinely open.
broke whenever a drive changed ports or enumeration order; `UUID=` entries
exist because device-path identity failed. danos skips that era: the mount
map (`volumes.csv` — configuration, read by the volume manager) keys on
**content identity, never port or discovery order**. The prober reads
identity off the medium, strongest first:
1. GPT partition GUID — 128-bit, unique, stable for the volume's life;
2. filesystem UUID (ext-family and most modern formats, in the superblock);
**content identity, never port or discovery order**. Build status: only
rung 4 (MBR signature + partition index) is implemented today; the fuller
rungs and the `volumes.csv` map itself land with the identity ladder, so
today a single volume mounts at the fixed `/volumes/usb` and its recorded
identity is not yet consulted to pick a path. The target ladder the prober
reads off the medium, strongest first:
1. GPT partition GUID — 128-bit, unique, stable for the volume's life *(planned)*;
2. filesystem UUID (ext-family and most modern formats, in the superblock) *(planned)*;
3. FAT volume serial + label — 32 bits, weak (dd-cloned sticks share it)
but what real sticks carry;
4. MBR disk signature + partition index;
5. nothing — an anonymous volume: generated mount name, no persistence.
but what real sticks carry *(planned)*;
4. MBR disk signature + partition index — **built**; a bare FAT with no
table takes index 0 over the whole device;
5. nothing — an anonymous volume: generated mount name, no persistence *(planned)*.
Consequences, each mechanical once identity keys the map: **moving a drive
to a different port changes nothing** — same identity, same mount point,
whether USB port, hub depth, SATA port, or a stick that left as USB and
returned in a SATA dock; **replug remounts at the same path** (the
remount-after-return story is a map lookup); **the boot volume** is the
returned in a SATA dock; **replug remounts at the same path** (a map lookup
once the map exists; today's single volume re-probes and remounts at the
fixed prefix, and remount-on-replug is bench-verified, not QEMU-tested); **the boot volume** is the
recorded identity of the volume carrying `/system/configuration`, findable
on any port; and **duplicate identity is a policy case, not a surprise** —
two cloned sticks at once: first keeps the mapped name, second mounts
+4 -2
View File
@@ -12,8 +12,10 @@ range confinement at the provider on one serving endpoint — the badge-scoped
provider pattern the xHCI bus already uses, applied to blocks. The volume
manager sets each filesystem process's range; the driver clamps and translates
every transfer by the sender's kernel-stamped badge; filesystems address
volume-relative LBAs from 0 and the FAT engine's `base_lba` is deleted rather
than moved. Every other decision the phases below execute is recorded in the
volume-relative LBAs from 0 (on the confined path the FAT engine's `base_lba`
resolves to 0; the field and the engine's own MBR walk remain as now-inert
legacy, the authoritative walk living in the volume manager's `partition.zig`).
Every other decision the phases below execute is recorded in the
rationale (decisions 1–8); nothing in this plan waits on a choice.
## V0 — `fs_unmount` ownership (the defect fix)