docs: correct the storage docs' V0-V4 status — the V5 flip missed several
A V5 close-out audit (docs against the actual code) found the storage docs overclaiming in both directions: the status header was flipped to "built" but several body markers were not, and two passages describe mechanisms the code never implemented. Nine confirmed, adversarially verified against the source: Underclaims (marked Planned, actually built): - storage-architecture: driver named sub-ranges / range confinement (V2, tested by the block-range case); the pushed medium_changed event (published today from a TEST UNIT READY poll); ownership-gated fs_unmount (V0, process.zig gates it with EPERM); the filesystem-harness extraction (V1). Scoped the remaining *planned* to the genuinely-pending parts (native-signal translation, the volume manager consuming medium_changed). Overclaims (described, never built): - storage-architecture: the "FAT dirty flag on disk" guarantee — no on-disk dirty/clean-shutdown bit exists; only an in-memory device-dirty bool gating a device write-cache flush on close. Fixed in all three places. - storage-architecture: fat "acquires its own volume (first mass-storage child by enumeration order)" — the V3b flip removed self-acquisition; fat is handed its volume id and channel by the volume manager. - rationale + plan: the FAT engine's base_lba "deleted rather than moved" — it and the engine's MBR walk still exist as now-inert legacy; the authoritative walk lives in partition.zig. - rationale: NVMe namespaces "decision 4 settles as endpoint-per-volume" — decision 4 settles the opposite (per-sender confinement, one endpoint); endpoint-per-volume is named only as an unbuilt future refactor. - rationale: the five-rung identity ladder and volumes.csv map stated in flat present tense — only rung 4 (MBR signature + index) is built; added the build-status hedge and marked each rung. Docs only; no code or behavior change. Suite unaffected (127/127).
This commit is contained in:
@@ -65,12 +65,13 @@ own children's class drivers, routed by lineage.
|
|||||||
|
|
||||||
**Storage driver** (usb-storage, per device; later nvme, ahci): hardware →
|
**Storage driver** (usb-storage, per device; later nvme, ahci): hardware →
|
||||||
blocks. Speaks its transport upward from its device; serves the block
|
blocks. Speaks its transport upward from its device; serves the block
|
||||||
protocol (geometry, read, write, flush, attach, detach — and, *planned*, the
|
protocol (geometry, read, write, flush, attach, detach — and the pushed
|
||||||
pushed `medium_changed` presence event, translated from the transport's
|
`medium_changed` presence event, published today from a slow TEST UNIT READY
|
||||||
native signal). **Content-blind, permanently**:
|
poll; *planned*: translating it from the transport's native signal instead of
|
||||||
|
polling). **Content-blind, permanently**:
|
||||||
it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no
|
it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no
|
||||||
GPT, no filesystem magic, ever. *(Planned)* it gains one content-blind
|
GPT, no filesystem magic, ever. It has one content-blind mechanism *(built)*:
|
||||||
mechanism: **named sub-ranges** — "serve blocks [a, b) as a channel of the
|
**named sub-ranges** — "serve blocks [a, b) as a channel of the
|
||||||
same block contract". It clamps and offsets; it never knows the numbers came
|
same block contract". It clamps and offsets; it never knows the numbers came
|
||||||
from a partition table. The clamp lives here and nowhere else because a
|
from a partition table. The clamp lives here and nowhere else because a
|
||||||
channel must carry exactly the authority it grants: handing a filesystem the
|
channel must carry exactly the authority it grants: handing a filesystem the
|
||||||
@@ -113,16 +114,20 @@ receives its block channel at spawn; it never discovers devices. It registers
|
|||||||
its own mounts with the kernel; its write cache lives inside the process, so a
|
its own mounts with the kernel; its write cache lives inside the process, so a
|
||||||
write error is observed by the code that owns the volume and surfaces on the
|
write error is observed by the code that owns the volume and surfaces on the
|
||||||
owning channel (the anti-fsyncgate rule — never a system-wide dirty pool).
|
owning channel (the anti-fsyncgate rule — never a system-wide dirty pool).
|
||||||
*(Today, interim:)* fat also acquires its own volume (first mass-storage child
|
*(Today, interim:)* fat still hardcodes its mount prefixes (`/volumes/usb` plus
|
||||||
by enumeration order) and hardcodes its mount prefixes; both migrate to the
|
the two boot-volume hierarchy subtrees it rewrites in place); a `volumes.csv`
|
||||||
volume manager.
|
mount map will migrate that to the volume manager. It no longer self-acquires a
|
||||||
|
volume — the V3b flip made it receive its volume id at spawn and its block
|
||||||
|
channel from the volume manager's hello reply, consistent with "it never
|
||||||
|
discovers devices" above.
|
||||||
|
|
||||||
**Kernel** (mechanism only): the mount table routes paths to backend
|
**Kernel** (mechanism only): the mount table routes paths to backend
|
||||||
endpoints — resolve and redirect, never data. Remount-replace is the restart
|
endpoints — resolve and redirect, never data. Remount-replace is the restart
|
||||||
story; a dead backend's slot is swept lazily on the next resolution.
|
story; a dead backend's slot is swept lazily on the next resolution.
|
||||||
*(Planned fix:)* `fs_unmount` gains ownership — only the mounting endpoint's
|
`fs_unmount` is ownership-gated *(built, V0)* — only the mounting task may
|
||||||
holder may unmount; today it is ungated, which is safe with one mount owner
|
unmount its own prefix; anyone else is refused (`EPERM`). No dead-owner
|
||||||
and wrong with several.
|
exception: a dead owner's slot is swept lazily by resolution, and restart goes
|
||||||
|
through remount-replace, never through a stranger's unmount.
|
||||||
|
|
||||||
## Adding a filesystem
|
## Adding a filesystem
|
||||||
|
|
||||||
@@ -135,8 +140,9 @@ and wrong with several.
|
|||||||
2. **Reuse the shell**: the filesystem harness — establishment, the
|
2. **Reuse the shell**: the filesystem harness — establishment, the
|
||||||
badge-scoped open-node table, the nine vfs-protocol handlers, mount
|
badge-scoped open-node table, the nine vfs-protocol handlers, mount
|
||||||
registration, the removal path — is shared code, not per-filesystem code.
|
registration, the removal path — is shared code, not per-filesystem code.
|
||||||
*(Planned: extracted from fat's 434-line shell into a library before the
|
*(Done: extracted from fat's original 434-line shell into
|
||||||
second engine is written.)* An engine plus a `main` wiring it into the
|
`library/kernel/file-system-harness.zig`; fat imports it and instantiates
|
||||||
|
`harness.Server(engine.FileSystem)`.)* An engine plus a `main` wiring it into the
|
||||||
harness is a complete filesystem service.
|
harness is a complete filesystem service.
|
||||||
3. **Add the configuration row**: one line in `filesystems.csv` mapping the
|
3. **Add the configuration row**: one line in `filesystems.csv` mapping the
|
||||||
on-disk signature to the binary. No other component changes: the volume
|
on-disk signature to the binary. No other component changes: the volume
|
||||||
@@ -162,13 +168,15 @@ error forever (Plan 9's dead-server wart).
|
|||||||
The path has **two triggers, one lifecycle**: the *device* leaving (the
|
The path has **two triggers, one lifecycle**: the *device* leaving (the
|
||||||
storage driver dies — channel death, the table below), and the *medium*
|
storage driver dies — channel death, the table below), and the *medium*
|
||||||
leaving while the device stays (an SD card pulled from its reader, an ATAPI
|
leaving while the device stays (an SD card pulled from its reader, an ATAPI
|
||||||
tray opened — including USB card readers today). The second trigger is a
|
tray opened — including USB card readers today). The second trigger is the
|
||||||
pushed `medium_changed` event on the block protocol *(planned)*: the storage
|
pushed `medium_changed` event on the block protocol — published today from a
|
||||||
driver translates its transport's native signal (SCSI UNIT ATTENTION, AHCI
|
TEST UNIT READY poll; still *planned* is the volume manager *consuming* it
|
||||||
PxSSTS, NVMe namespace-change AER) into presence-changed — never content —
|
(today removal is driven only by device-presence polling) and translating the
|
||||||
and the volume manager runs the same kill-retire path, then re-probes on
|
transport's native signal (SCSI UNIT ATTENTION, AHCI PxSSTS, NVMe
|
||||||
medium return exactly as on device return. Without it, a swapped card would
|
namespace-change AER) in place of the poll. On the event the volume manager
|
||||||
be served with the previous card's filesystem state.
|
runs the same kill-retire path, then re-probes on medium return exactly as on
|
||||||
|
device return. Without it, a swapped card would be served with the previous
|
||||||
|
card's filesystem state.
|
||||||
|
|
||||||
| Layer | Observes | Must do | Guarantees |
|
| Layer | Observes | Must do | Guarantees |
|
||||||
|---|---|---|---|
|
|---|---|---|---|
|
||||||
@@ -176,7 +184,7 @@ be served with the previous card's filesystem state.
|
|||||||
| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
|
| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
|
||||||
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
|
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
|
||||||
| Volume manager *(removal built; remount bench-pending)* | the storage device leaving the device-manager tree (poll) | kill the filesystem service of that device's volume; its kernel mounts retire | one removal path; mounts never dangle; log persistence stops *cleanly* |
|
| Volume manager *(removal built; remount bench-pending)* | the storage device leaving the device-manager tree (poll) | kill the filesystem service of that device's volume; its kernel mounts retire | one removal path; mounts never dangle; log persistence stops *cleanly* |
|
||||||
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the FAT dirty flag on disk marks the unclean removal; the process never serves from behind a dead channel |
|
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the unflushed write-back window is dropped on a surprise yank — danos writes no on-disk dirty/clean-shutdown marker today; the process never serves from behind a dead channel |
|
||||||
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated `fs_unmount` (built, V0) | resolution under a dead mount is `not_found`, not a hang |
|
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated `fs_unmount` (built, V0) | resolution under a dead mount is `not_found`, not a hang |
|
||||||
| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
|
| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
|
||||||
|
|
||||||
@@ -222,7 +230,8 @@ never abandoned after the first `not_found`.
|
|||||||
|
|
||||||
What is lost on a surprise yank is exactly the write-back window of the
|
What is lost on a surprise yank is exactly the write-back window of the
|
||||||
filesystem service, no more: the engine owns its cache, so the blast radius of
|
filesystem service, no more: the engine owns its cache, so the blast radius of
|
||||||
a yank is one volume's unflushed writes, marked by the on-disk dirty flag. A
|
a yank is one volume's unflushed writes, which are simply lost — danos records
|
||||||
|
no on-disk dirty/clean-shutdown bit yet. A
|
||||||
filesystem format with better crash honesty (journaling, copy-on-write — the
|
filesystem format with better crash honesty (journaling, copy-on-write — the
|
||||||
lesson of QNX's Power-Safe) narrows that window further and slots in as
|
lesson of QNX's Power-Safe) narrows that window further and slots in as
|
||||||
implementation N+1 through the table above, changing nothing else.
|
implementation N+1 through the table above, changing nothing else.
|
||||||
@@ -249,8 +258,9 @@ the device lifecycle already uses, mapped one-to-one:
|
|||||||
obsolescence (the reap argument, proven on the device tree).
|
obsolescence (the reap argument, proven on the device tree).
|
||||||
3. **The shared harness is how every filesystem inherits the lifecycle by
|
3. **The shared harness is how every filesystem inherits the lifecycle by
|
||||||
construction.** The harness — not the engine — owns the state machine:
|
construction.** The harness — not the engine — owns the state machine:
|
||||||
establishment at spawn, mount registration, the dirty-flag set/clear
|
establishment at spawn, mount registration, the flush-on-close hook (an
|
||||||
bracket, error-out-and-exit on channel death. The engine sits behind the
|
in-memory device-dirty check that commits the device write cache on close),
|
||||||
|
error-out-and-exit on channel death. The engine sits behind the
|
||||||
four-function vtable and never sees a channel; it cannot opt out of the
|
four-function vtable and never sees a channel; it cannot opt out of the
|
||||||
lifecycle for the same reason it cannot find a device. This is why the
|
lifecycle for the same reason it cannot find a device. This is why the
|
||||||
harness is extracted BEFORE the second engine is written.
|
harness is extracted BEFORE the second engine is written.
|
||||||
|
|||||||
@@ -105,8 +105,9 @@ and /system/logs), closing the two-sticks question honestly.
|
|||||||
Plan 9's `dossrv` to Minix to Fuchsia: block-client + engine + file-protocol
|
Plan 9's `dossrv` to Minix to Fuchsia: block-client + engine + file-protocol
|
||||||
provider in one binary, one process per volume (9front practice; per-volume
|
provider in one binary, one process per volume (9front practice; per-volume
|
||||||
fault isolation is what our supervision makes cheap). fat's shell becomes a
|
fault isolation is what our supervision makes cheap). fat's shell becomes a
|
||||||
shared *filesystem harness* library before a second engine is written; the MBR
|
shared *filesystem harness* library before a second engine is written; a
|
||||||
walk moves out of the engine into the volume manager; write caching stays
|
partition walk is added in the volume manager (`partition.zig`) — the engine's
|
||||||
|
own MBR walk currently remains alongside it; write caching stays
|
||||||
inside the process (the anti-fsyncgate rule). Each mounts its prefixes into the
|
inside the process (the anti-fsyncgate rule). Each mounts its prefixes into the
|
||||||
kernel mount table itself, exactly as today.
|
kernel mount table itself, exactly as today.
|
||||||
|
|
||||||
@@ -132,8 +133,10 @@ what the architecture promises: one driver binary plus one `devices.csv` row.
|
|||||||
**NVMe** shortens the stack — no bus/class split; the driver IS the
|
**NVMe** shortens the stack — no bus/class split; the driver IS the
|
||||||
controller driver, one hop fewer than USB. Its structural novelty,
|
controller driver, one hop fewer than USB. Its structural novelty,
|
||||||
**namespaces** (hardware-native multiple volumes behind one controller), is
|
**namespaces** (hardware-native multiple volumes behind one controller), is
|
||||||
exactly the case the block protocol reserved on day one and decision 4 below
|
exactly the case the block protocol reserved on day one — still reserved, not
|
||||||
settles as endpoint-per-volume. **AHCI** is a shape choice, not a problem: an
|
implemented (pressure point 1 below): decision 4 settles per-volume addressing
|
||||||
|
as per-sender confinement on a single endpoint, and names endpoint-per-volume
|
||||||
|
only as an unbuilt future refactor. **AHCI** is a shape choice, not a problem: an
|
||||||
HBA fronts up to 32 ports plus port multipliers — structurally a bus — so
|
HBA fronts up to 32 ports plus port multipliers — structurally a bus — so
|
||||||
either mirror USB (ahci-bus + a per-port disk driver: maximum restart
|
either mirror USB (ahci-bus + a per-port disk driver: maximum restart
|
||||||
granularity, the proven shape) or mirror NVMe (one driver per controller, one
|
granularity, the proven shape) or mirror NVMe (one driver per controller, one
|
||||||
@@ -180,7 +183,10 @@ matrix-proven shape; genuinely open.
|
|||||||
same pattern applied to blocks. The volume manager sets each filesystem
|
same pattern applied to blocks. The volume manager sets each filesystem
|
||||||
process's range on the driver; the driver clamps AND translates every
|
process's range on the driver; the driver clamps AND translates every
|
||||||
transfer by the sender's range, so filesystems address volume-relative
|
transfer by the sender's range, so filesystems address volume-relative
|
||||||
LBAs from 0 and the FAT engine's `base_lba` is deleted rather than moved.
|
LBAs from 0. On the confined path the FAT engine's `base_lba` therefore
|
||||||
|
resolves to 0 on every access; the field and the engine's own MBR walk
|
||||||
|
remain in `engine.zig` as now-inert legacy code (the authoritative partition
|
||||||
|
walk lives in the volume manager's `partition.zig`), not yet deleted.
|
||||||
The enforcement point (the clamp at the provider, never in the consumer)
|
The enforcement point (the clamp at the provider, never in the consumer)
|
||||||
is what carries the security property; endpoint-per-volume would deliver
|
is what carries the security property; endpoint-per-volume would deliver
|
||||||
the same property only by inventing a multi-endpoint harness the pattern
|
the same property only by inventing a multi-endpoint harness the pattern
|
||||||
@@ -209,20 +215,26 @@ matrix-proven shape; genuinely open.
|
|||||||
broke whenever a drive changed ports or enumeration order; `UUID=` entries
|
broke whenever a drive changed ports or enumeration order; `UUID=` entries
|
||||||
exist because device-path identity failed. danos skips that era: the mount
|
exist because device-path identity failed. danos skips that era: the mount
|
||||||
map (`volumes.csv` — configuration, read by the volume manager) keys on
|
map (`volumes.csv` — configuration, read by the volume manager) keys on
|
||||||
**content identity, never port or discovery order**. The prober reads
|
**content identity, never port or discovery order**. Build status: only
|
||||||
identity off the medium, strongest first:
|
rung 4 (MBR signature + partition index) is implemented today; the fuller
|
||||||
1. GPT partition GUID — 128-bit, unique, stable for the volume's life;
|
rungs and the `volumes.csv` map itself land with the identity ladder, so
|
||||||
2. filesystem UUID (ext-family and most modern formats, in the superblock);
|
today a single volume mounts at the fixed `/volumes/usb` and its recorded
|
||||||
|
identity is not yet consulted to pick a path. The target ladder the prober
|
||||||
|
reads off the medium, strongest first:
|
||||||
|
1. GPT partition GUID — 128-bit, unique, stable for the volume's life *(planned)*;
|
||||||
|
2. filesystem UUID (ext-family and most modern formats, in the superblock) *(planned)*;
|
||||||
3. FAT volume serial + label — 32 bits, weak (dd-cloned sticks share it)
|
3. FAT volume serial + label — 32 bits, weak (dd-cloned sticks share it)
|
||||||
but what real sticks carry;
|
but what real sticks carry *(planned)*;
|
||||||
4. MBR disk signature + partition index;
|
4. MBR disk signature + partition index — **built**; a bare FAT with no
|
||||||
5. nothing — an anonymous volume: generated mount name, no persistence.
|
table takes index 0 over the whole device;
|
||||||
|
5. nothing — an anonymous volume: generated mount name, no persistence *(planned)*.
|
||||||
|
|
||||||
Consequences, each mechanical once identity keys the map: **moving a drive
|
Consequences, each mechanical once identity keys the map: **moving a drive
|
||||||
to a different port changes nothing** — same identity, same mount point,
|
to a different port changes nothing** — same identity, same mount point,
|
||||||
whether USB port, hub depth, SATA port, or a stick that left as USB and
|
whether USB port, hub depth, SATA port, or a stick that left as USB and
|
||||||
returned in a SATA dock; **replug remounts at the same path** (the
|
returned in a SATA dock; **replug remounts at the same path** (a map lookup
|
||||||
remount-after-return story is a map lookup); **the boot volume** is the
|
once the map exists; today's single volume re-probes and remounts at the
|
||||||
|
fixed prefix, and remount-on-replug is bench-verified, not QEMU-tested); **the boot volume** is the
|
||||||
recorded identity of the volume carrying `/system/configuration`, findable
|
recorded identity of the volume carrying `/system/configuration`, findable
|
||||||
on any port; and **duplicate identity is a policy case, not a surprise** —
|
on any port; and **duplicate identity is a policy case, not a surprise** —
|
||||||
two cloned sticks at once: first keeps the mapped name, second mounts
|
two cloned sticks at once: first keeps the mapped name, second mounts
|
||||||
|
|||||||
@@ -12,8 +12,10 @@ range confinement at the provider on one serving endpoint — the badge-scoped
|
|||||||
provider pattern the xHCI bus already uses, applied to blocks. The volume
|
provider pattern the xHCI bus already uses, applied to blocks. The volume
|
||||||
manager sets each filesystem process's range; the driver clamps and translates
|
manager sets each filesystem process's range; the driver clamps and translates
|
||||||
every transfer by the sender's kernel-stamped badge; filesystems address
|
every transfer by the sender's kernel-stamped badge; filesystems address
|
||||||
volume-relative LBAs from 0 and the FAT engine's `base_lba` is deleted rather
|
volume-relative LBAs from 0 (on the confined path the FAT engine's `base_lba`
|
||||||
than moved. Every other decision the phases below execute is recorded in the
|
resolves to 0; the field and the engine's own MBR walk remain as now-inert
|
||||||
|
legacy, the authoritative walk living in the volume manager's `partition.zig`).
|
||||||
|
Every other decision the phases below execute is recorded in the
|
||||||
rationale (decisions 1–8); nothing in this plan waits on a choice.
|
rationale (decisions 1–8); nothing in this plan waits on a choice.
|
||||||
|
|
||||||
## V0 — `fs_unmount` ownership (the defect fix)
|
## V0 — `fs_unmount` ownership (the defect fix)
|
||||||
|
|||||||
Reference in New Issue
Block a user