diff --git a/docs/README.md b/docs/README.md index 5663e48..e79e601 100644 --- a/docs/README.md +++ b/docs/README.md @@ -51,7 +51,11 @@ rather than restate it. Roughly in the order things happen at runtime: 14. **[vfs-protocol.md](file-system-development/vfs-protocol.md) — the VFS wire protocol.** The language-neutral byte-level spec of the file protocol spoken over IPC: request/reply headers, the operation table, mount routing, and the append-only evolution rules — the - first IPC protocol documented as public ABI. + first IPC protocol documented as public ABI. Its architectural frame is + **[storage-architecture.md](file-system-development/storage-architecture.md) — the storage stack**: + the layers from application to hardware, the three kinds of boundary + (protocol, library, control-plane), how to add a filesystem, and the + per-layer responsibilities when removable media is yanked and returned. 15. **[drivers.md](device-driver-development/drivers.md) — writing a driver.** The payoff: a driver is an ordinary ring-3 process that claims a device, maps its registers, and **sleeps until its hardware interrupts it**. The claim is the capability; `irq_ack` is the diff --git a/docs/file-system-development/storage-architecture.md b/docs/file-system-development/storage-architecture.md new file mode 100644 index 0000000..22421da --- /dev/null +++ b/docs/file-system-development/storage-architecture.md @@ -0,0 +1,163 @@ +# The storage architecture: layers, boundaries, responsibilities + +> **Status:** the layered model below is the settled design +> ([storage-stack-discussion.md](../storage-stack-discussion.md) records how it +> was reached and what the surveyed systems taught). The data path — vfs +> protocol, kernel mount routing, the FAT service, the block protocol, +> usb-storage — is **built**. The volume manager, the driver's range +> mechanism, per-volume filesystem spawning, and the removal lifecycle are +> **planned**; until they land, the FAT service performs volume-manager duties +> itself (marked below). This document is the reference for both states. + +## The model + +One pattern, applied twice: an application is a client of a service over a +protocol; a service is a client of a driver over a protocol. + +``` +application + │ vfs protocol (routed by the kernel mount table) + ▼ +filesystem service ── one process per volume + │ inside: [file-protocol shell] ── Zig API ── [fs engine] + │ block protocol (channel received at spawn) + ▼ +storage driver ── one process per device (usb-storage per stick) + │ usb-transfer protocol (channel received via hello, by lineage) + ▼ +bus driver ── one process per controller (usb-xhci-bus) + │ hardware +``` + +Three kinds of boundary, deliberately different: + +- **Protocol boundaries** sit where processes meet (vfs, block, usb-transfer). + Each side is independently restartable; channels are established by + capability handoff, never by registry names (communication.md + "Establishment: two planes"). Steady-state file I/O crosses exactly two — + vfs and block. +- **Library boundaries** sit where concerns meet inside one process. The + filesystem service's engine (on-disk format logic, host-testable, behind + the four-function `BlockDevice` vtable) and shell (IPC serving, open-node + table, mount registration) talk through a Zig API. No protocol between + them: they share fate regardless — a corrupt engine corrupts the answers + either way — so a channel there would add a hop per file operation and buy + nothing. The seam exists for compile-time testability and reuse; the + *process* is the restart unit. +- **Control-plane relationships** sit beside the data path, never on it. Two + supervisors, one per layer: the **device manager** wires and revives the + device layers (bus and storage drivers — devices only); the **volume + manager** *(planned)* wires and revives the volume layer (filesystem + services). Neither touches steady-state I/O. + +## Who does what + +**Bus driver** (usb-xhci-bus, per controller): hardware only. Enumerates, +reports children to the device manager, serves the transfer contract to its +own children's class drivers, routed by lineage. + +**Storage driver** (usb-storage, per device): hardware → blocks. Speaks SCSI +Bulk-Only Transport upward from its device; serves the block protocol (five +verbs: geometry, read, write, flush, attach). **Content-blind, permanently**: +it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no +GPT, no filesystem magic, ever. *(Planned)* it gains one content-blind +mechanism: **named sub-ranges** — "serve blocks [a, b) as a channel of the +same block contract". It clamps and offsets; it never knows the numbers came +from a partition table. The clamp lives here and nowhere else because a +channel must carry exactly the authority it grants: handing a filesystem the +whole disk plus a polite base offset would let a buggy or compromised +filesystem scribble the neighboring partition — the same authority-overshoot +the device-authority track eliminated for MMIO and DMA. + +**Volume manager** *(planned; today the FAT service squats on these duties)*: +the policy home of the volume layer, one service, supervised by init. It +subscribes to the device manager's child events; when a storage provider +appears it consumer-hellos for the block channel, reads the partition table +and the first blocks itself (**it** is the prober), consults its +configuration, defines sub-ranges on the driver, spawns the matching +filesystem service per volume with that volume's channel, supervises it, and +decides mount placement. Its tables are CSV configuration, read by it (the +policy), enforced by nobody else: + +- `filesystems.csv` — content signature → filesystem binary. Adding a + filesystem adds a row. +- the mount map — volume identity → mount prefix. The boot volume is chosen + by **content** (the volume carrying `/system/configuration` and + `/system/logs`), never by port or arrival order. + +**Filesystem service** (the FAT service today; one process per volume): the +proven unit — block-client + engine + file-protocol provider in one binary. It +receives its block channel at spawn; it never discovers devices. It registers +its own mounts with the kernel; its write cache lives inside the process, so a +write error is observed by the code that owns the volume and surfaces on the +owning channel (the anti-fsyncgate rule — never a system-wide dirty pool). +*(Today, interim:)* fat also acquires its own volume (first mass-storage child +by enumeration order) and hardcodes its mount prefixes; both migrate to the +volume manager. + +**Kernel** (mechanism only): the mount table routes paths to backend +endpoints — resolve and redirect, never data. Remount-replace is the restart +story; a dead backend's slot is swept lazily on the next resolution. +*(Planned fix:)* `fs_unmount` gains ownership — only the mounting endpoint's +holder may unmount; today it is ungated, which is safe with one mount owner +and wrong with several. + +## Adding a filesystem + +1. **Write the engine**: pure Zig, no IPC, behind the `BlockDevice` vtable + (four functions). On-disk layout in `align(1)` extern structs. Host-test it + with `zig build test` fixtures — the FAT engine + ([engine.zig](../../system/services/fat/engine.zig), + [on-disk.zig](../../system/services/fat/on-disk.zig)) is the template, and + its ~500 lines of host tests the standard. +2. **Reuse the shell**: the filesystem harness — establishment, the + badge-scoped open-node table, the nine vfs-protocol handlers, mount + registration, the removal path — is shared code, not per-filesystem code. + *(Planned: extracted from fat's 434-line shell into a library before the + second engine is written.)* An engine plus a `main` wiring it into the + harness is a complete filesystem service. +3. **Add the configuration row**: one line in `filesystems.csv` mapping the + on-disk signature to the binary. No other component changes: the volume + manager probes, matches, spawns; the vfs protocol is already + backend-neutral (a second provider serves it in production today — init's + synthetic registry backend). +4. **Prove removal**: extend the removable-media suite (below) for the new + filesystem — surprise-yank during writes must lose only what was + unflushed, said honestly, with the volume consistent enough to remount. + +The engine never sees: partitions (it receives a volume-shaped block channel; +base offsets are the driver's clamp), device discovery (the channel arrives at +spawn), mount policy (prefixes are handed to it), or other volumes (one +process, one volume). + +## Removable media: responsibilities on removal, per layer + +The design rule, learned from what Linux cannot do: there is exactly **one** +surprise-removal path — kill the filesystem process, retire its mounts, +respawn on return. No half-alive states, no `remount-ro`, no mounts that +error forever (Plan 9's dead-server wart). + +| Layer | Observes | Must do | Guarantees | +|---|---|---|---| +| Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick | +| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds | +| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang | +| Volume manager *(planned)* | the storage provider's channel death / the manager's child events | kill each filesystem service of that device's volumes; retire their kernel mounts; remember the volume identity | one removal path; mounts never dangle; log persistence stops *cleanly* | +| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the FAT dirty flag on disk marks the unclean removal; the process never serves from behind a dead channel | +| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); *(planned)* ownership-gated `fs_unmount` | resolution under a dead mount is `not_found`, not a hang | +| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media | + +**On return** the same table runs upward in reverse: the bus re-enumerates and +re-reports; the device manager respawns the storage driver (built); the volume +manager re-probes — same content, same volume identity — respawns the +filesystem service, and remounts at the same prefix; applications see the +subtree reappear. A different stick in the same port is a *different volume* +(identity is content, not port) and mounts wherever its identity says — +possibly nowhere but `/volumes/…`. + +What is lost on a surprise yank is exactly the write-back window of the +filesystem service, no more: the engine owns its cache, so the blast radius of +a yank is one volume's unflushed writes, marked by the on-disk dirty flag. A +filesystem format with better crash honesty (journaling, copy-on-write — the +lesson of QNX's Power-Safe) narrows that window further and slots in as +implementation N+1 through the table above, changing nothing else.