diff --git a/docs/file-system-development/storage-architecture.md b/docs/file-system-development/storage-architecture.md index f31ae13..ef5bfb0 100644 --- a/docs/file-system-development/storage-architecture.md +++ b/docs/file-system-development/storage-architecture.md @@ -65,12 +65,13 @@ own children's class drivers, routed by lineage. **Storage driver** (usb-storage, per device; later nvme, ahci): hardware → blocks. Speaks its transport upward from its device; serves the block -protocol (geometry, read, write, flush, attach, detach — and, *planned*, the -pushed `medium_changed` presence event, translated from the transport's -native signal). **Content-blind, permanently**: +protocol (geometry, read, write, flush, attach, detach — and the pushed +`medium_changed` presence event, published today from a slow TEST UNIT READY +poll; *planned*: translating it from the transport's native signal instead of +polling). **Content-blind, permanently**: it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no -GPT, no filesystem magic, ever. *(Planned)* it gains one content-blind -mechanism: **named sub-ranges** — "serve blocks [a, b) as a channel of the +GPT, no filesystem magic, ever. It has one content-blind mechanism *(built)*: +**named sub-ranges** — "serve blocks [a, b) as a channel of the same block contract". It clamps and offsets; it never knows the numbers came from a partition table. The clamp lives here and nowhere else because a channel must carry exactly the authority it grants: handing a filesystem the @@ -113,16 +114,20 @@ receives its block channel at spawn; it never discovers devices. It registers its own mounts with the kernel; its write cache lives inside the process, so a write error is observed by the code that owns the volume and surfaces on the owning channel (the anti-fsyncgate rule — never a system-wide dirty pool). -*(Today, interim:)* fat also acquires its own volume (first mass-storage child -by enumeration order) and hardcodes its mount prefixes; both migrate to the -volume manager. +*(Today, interim:)* fat still hardcodes its mount prefixes (`/volumes/usb` plus +the two boot-volume hierarchy subtrees it rewrites in place); a `volumes.csv` +mount map will migrate that to the volume manager. It no longer self-acquires a +volume — the V3b flip made it receive its volume id at spawn and its block +channel from the volume manager's hello reply, consistent with "it never +discovers devices" above. **Kernel** (mechanism only): the mount table routes paths to backend endpoints — resolve and redirect, never data. Remount-replace is the restart story; a dead backend's slot is swept lazily on the next resolution. -*(Planned fix:)* `fs_unmount` gains ownership — only the mounting endpoint's -holder may unmount; today it is ungated, which is safe with one mount owner -and wrong with several. +`fs_unmount` is ownership-gated *(built, V0)* — only the mounting task may +unmount its own prefix; anyone else is refused (`EPERM`). No dead-owner +exception: a dead owner's slot is swept lazily by resolution, and restart goes +through remount-replace, never through a stranger's unmount. ## Adding a filesystem @@ -135,8 +140,9 @@ and wrong with several. 2. **Reuse the shell**: the filesystem harness — establishment, the badge-scoped open-node table, the nine vfs-protocol handlers, mount registration, the removal path — is shared code, not per-filesystem code. - *(Planned: extracted from fat's 434-line shell into a library before the - second engine is written.)* An engine plus a `main` wiring it into the + *(Done: extracted from fat's original 434-line shell into + `library/kernel/file-system-harness.zig`; fat imports it and instantiates + `harness.Server(engine.FileSystem)`.)* An engine plus a `main` wiring it into the harness is a complete filesystem service. 3. **Add the configuration row**: one line in `filesystems.csv` mapping the on-disk signature to the binary. No other component changes: the volume @@ -162,13 +168,15 @@ error forever (Plan 9's dead-server wart). The path has **two triggers, one lifecycle**: the *device* leaving (the storage driver dies — channel death, the table below), and the *medium* leaving while the device stays (an SD card pulled from its reader, an ATAPI -tray opened — including USB card readers today). The second trigger is a -pushed `medium_changed` event on the block protocol *(planned)*: the storage -driver translates its transport's native signal (SCSI UNIT ATTENTION, AHCI -PxSSTS, NVMe namespace-change AER) into presence-changed — never content — -and the volume manager runs the same kill-retire path, then re-probes on -medium return exactly as on device return. Without it, a swapped card would -be served with the previous card's filesystem state. +tray opened — including USB card readers today). The second trigger is the +pushed `medium_changed` event on the block protocol — published today from a +TEST UNIT READY poll; still *planned* is the volume manager *consuming* it +(today removal is driven only by device-presence polling) and translating the +transport's native signal (SCSI UNIT ATTENTION, AHCI PxSSTS, NVMe +namespace-change AER) in place of the poll. On the event the volume manager +runs the same kill-retire path, then re-probes on medium return exactly as on +device return. Without it, a swapped card would be served with the previous +card's filesystem state. | Layer | Observes | Must do | Guarantees | |---|---|---|---| @@ -176,7 +184,7 @@ be served with the previous card's filesystem state. | Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds | | Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang | | Volume manager *(removal built; remount bench-pending)* | the storage device leaving the device-manager tree (poll) | kill the filesystem service of that device's volume; its kernel mounts retire | one removal path; mounts never dangle; log persistence stops *cleanly* | -| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the FAT dirty flag on disk marks the unclean removal; the process never serves from behind a dead channel | +| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the unflushed write-back window is dropped on a surprise yank — danos writes no on-disk dirty/clean-shutdown marker today; the process never serves from behind a dead channel | | Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated `fs_unmount` (built, V0) | resolution under a dead mount is `not_found`, not a hang | | Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media | @@ -222,7 +230,8 @@ never abandoned after the first `not_found`. What is lost on a surprise yank is exactly the write-back window of the filesystem service, no more: the engine owns its cache, so the blast radius of -a yank is one volume's unflushed writes, marked by the on-disk dirty flag. A +a yank is one volume's unflushed writes, which are simply lost — danos records +no on-disk dirty/clean-shutdown bit yet. A filesystem format with better crash honesty (journaling, copy-on-write — the lesson of QNX's Power-Safe) narrows that window further and slots in as implementation N+1 through the table above, changing nothing else. @@ -249,8 +258,9 @@ the device lifecycle already uses, mapped one-to-one: obsolescence (the reap argument, proven on the device tree). 3. **The shared harness is how every filesystem inherits the lifecycle by construction.** The harness — not the engine — owns the state machine: - establishment at spawn, mount registration, the dirty-flag set/clear - bracket, error-out-and-exit on channel death. The engine sits behind the + establishment at spawn, mount registration, the flush-on-close hook (an + in-memory device-dirty check that commits the device write cache on close), + error-out-and-exit on channel death. The engine sits behind the four-function vtable and never sees a channel; it cannot opt out of the lifecycle for the same reason it cannot find a device. This is why the harness is extracted BEFORE the second engine is written. diff --git a/docs/file-system-development/storage-design-rationale.md b/docs/file-system-development/storage-design-rationale.md index fdb2662..380204e 100644 --- a/docs/file-system-development/storage-design-rationale.md +++ b/docs/file-system-development/storage-design-rationale.md @@ -105,8 +105,9 @@ and /system/logs), closing the two-sticks question honestly. Plan 9's `dossrv` to Minix to Fuchsia: block-client + engine + file-protocol provider in one binary, one process per volume (9front practice; per-volume fault isolation is what our supervision makes cheap). fat's shell becomes a -shared *filesystem harness* library before a second engine is written; the MBR -walk moves out of the engine into the volume manager; write caching stays +shared *filesystem harness* library before a second engine is written; a +partition walk is added in the volume manager (`partition.zig`) — the engine's +own MBR walk currently remains alongside it; write caching stays inside the process (the anti-fsyncgate rule). Each mounts its prefixes into the kernel mount table itself, exactly as today. @@ -132,8 +133,10 @@ what the architecture promises: one driver binary plus one `devices.csv` row. **NVMe** shortens the stack — no bus/class split; the driver IS the controller driver, one hop fewer than USB. Its structural novelty, **namespaces** (hardware-native multiple volumes behind one controller), is -exactly the case the block protocol reserved on day one and decision 4 below -settles as endpoint-per-volume. **AHCI** is a shape choice, not a problem: an +exactly the case the block protocol reserved on day one — still reserved, not +implemented (pressure point 1 below): decision 4 settles per-volume addressing +as per-sender confinement on a single endpoint, and names endpoint-per-volume +only as an unbuilt future refactor. **AHCI** is a shape choice, not a problem: an HBA fronts up to 32 ports plus port multipliers — structurally a bus — so either mirror USB (ahci-bus + a per-port disk driver: maximum restart granularity, the proven shape) or mirror NVMe (one driver per controller, one @@ -180,7 +183,10 @@ matrix-proven shape; genuinely open. same pattern applied to blocks. The volume manager sets each filesystem process's range on the driver; the driver clamps AND translates every transfer by the sender's range, so filesystems address volume-relative - LBAs from 0 and the FAT engine's `base_lba` is deleted rather than moved. + LBAs from 0. On the confined path the FAT engine's `base_lba` therefore + resolves to 0 on every access; the field and the engine's own MBR walk + remain in `engine.zig` as now-inert legacy code (the authoritative partition + walk lives in the volume manager's `partition.zig`), not yet deleted. The enforcement point (the clamp at the provider, never in the consumer) is what carries the security property; endpoint-per-volume would deliver the same property only by inventing a multi-endpoint harness the pattern @@ -209,20 +215,26 @@ matrix-proven shape; genuinely open. broke whenever a drive changed ports or enumeration order; `UUID=` entries exist because device-path identity failed. danos skips that era: the mount map (`volumes.csv` — configuration, read by the volume manager) keys on - **content identity, never port or discovery order**. The prober reads - identity off the medium, strongest first: - 1. GPT partition GUID — 128-bit, unique, stable for the volume's life; - 2. filesystem UUID (ext-family and most modern formats, in the superblock); + **content identity, never port or discovery order**. Build status: only + rung 4 (MBR signature + partition index) is implemented today; the fuller + rungs and the `volumes.csv` map itself land with the identity ladder, so + today a single volume mounts at the fixed `/volumes/usb` and its recorded + identity is not yet consulted to pick a path. The target ladder the prober + reads off the medium, strongest first: + 1. GPT partition GUID — 128-bit, unique, stable for the volume's life *(planned)*; + 2. filesystem UUID (ext-family and most modern formats, in the superblock) *(planned)*; 3. FAT volume serial + label — 32 bits, weak (dd-cloned sticks share it) - but what real sticks carry; - 4. MBR disk signature + partition index; - 5. nothing — an anonymous volume: generated mount name, no persistence. + but what real sticks carry *(planned)*; + 4. MBR disk signature + partition index — **built**; a bare FAT with no + table takes index 0 over the whole device; + 5. nothing — an anonymous volume: generated mount name, no persistence *(planned)*. Consequences, each mechanical once identity keys the map: **moving a drive to a different port changes nothing** — same identity, same mount point, whether USB port, hub depth, SATA port, or a stick that left as USB and - returned in a SATA dock; **replug remounts at the same path** (the - remount-after-return story is a map lookup); **the boot volume** is the + returned in a SATA dock; **replug remounts at the same path** (a map lookup + once the map exists; today's single volume re-probes and remounts at the + fixed prefix, and remount-on-replug is bench-verified, not QEMU-tested); **the boot volume** is the recorded identity of the volume carrying `/system/configuration`, findable on any port; and **duplicate identity is a policy case, not a surprise** — two cloned sticks at once: first keeps the mapped name, second mounts diff --git a/docs/volume-manager-plan.md b/docs/volume-manager-plan.md index cd9d8d0..518b27f 100644 --- a/docs/volume-manager-plan.md +++ b/docs/volume-manager-plan.md @@ -12,8 +12,10 @@ range confinement at the provider on one serving endpoint — the badge-scoped provider pattern the xHCI bus already uses, applied to blocks. The volume manager sets each filesystem process's range; the driver clamps and translates every transfer by the sender's kernel-stamped badge; filesystems address -volume-relative LBAs from 0 and the FAT engine's `base_lba` is deleted rather -than moved. Every other decision the phases below execute is recorded in the +volume-relative LBAs from 0 (on the confined path the FAT engine's `base_lba` +resolves to 0; the field and the engine's own MBR walk remain as now-inert +legacy, the authoritative walk living in the volume manager's `partition.zig`). +Every other decision the phases below execute is recorded in the rationale (decisions 1–8); nothing in this plan waits on a choice. ## V0 — `fs_unmount` ownership (the defect fix)