docs: the storage architecture — layers, boundaries, and who does what when media leaves
The settled shape from the storage-stack discussion, written as the reference: the data path (vfs -> filesystem service -> block -> driver) as the application/service/protocol/driver model applied twice; the three kinds of boundary (protocol between processes, library inside them, control-plane beside them); the volume manager as the policy home (planned — the FAT service squats on its duties today, marked as such); adding a filesystem as engine + shared harness + one configuration row; and the per-layer responsibility table for removable media — one removal path, kill/retire/ respawn, dirty data lost and SAID to be lost. Indexed from docs/README.md.
This commit is contained in:
+5
-1
@@ -51,7 +51,11 @@ rather than restate it. Roughly in the order things happen at runtime:
|
||||
14. **[vfs-protocol.md](file-system-development/vfs-protocol.md) — the VFS wire protocol.** The language-neutral
|
||||
byte-level spec of the file protocol spoken over IPC: request/reply headers,
|
||||
the operation table, mount routing, and the append-only evolution rules — the
|
||||
first IPC protocol documented as public ABI.
|
||||
first IPC protocol documented as public ABI. Its architectural frame is
|
||||
**[storage-architecture.md](file-system-development/storage-architecture.md) — the storage stack**:
|
||||
the layers from application to hardware, the three kinds of boundary
|
||||
(protocol, library, control-plane), how to add a filesystem, and the
|
||||
per-layer responsibilities when removable media is yanked and returned.
|
||||
15. **[drivers.md](device-driver-development/drivers.md) — writing a driver.** The payoff: a driver is an
|
||||
ordinary ring-3 process that claims a device, maps its registers, and **sleeps
|
||||
until its hardware interrupts it**. The claim is the capability; `irq_ack` is the
|
||||
|
||||
@@ -0,0 +1,163 @@
|
||||
# The storage architecture: layers, boundaries, responsibilities
|
||||
|
||||
> **Status:** the layered model below is the settled design
|
||||
> ([storage-stack-discussion.md](../storage-stack-discussion.md) records how it
|
||||
> was reached and what the surveyed systems taught). The data path — vfs
|
||||
> protocol, kernel mount routing, the FAT service, the block protocol,
|
||||
> usb-storage — is **built**. The volume manager, the driver's range
|
||||
> mechanism, per-volume filesystem spawning, and the removal lifecycle are
|
||||
> **planned**; until they land, the FAT service performs volume-manager duties
|
||||
> itself (marked below). This document is the reference for both states.
|
||||
|
||||
## The model
|
||||
|
||||
One pattern, applied twice: an application is a client of a service over a
|
||||
protocol; a service is a client of a driver over a protocol.
|
||||
|
||||
```
|
||||
application
|
||||
│ vfs protocol (routed by the kernel mount table)
|
||||
▼
|
||||
filesystem service ── one process per volume
|
||||
│ inside: [file-protocol shell] ── Zig API ── [fs engine]
|
||||
│ block protocol (channel received at spawn)
|
||||
▼
|
||||
storage driver ── one process per device (usb-storage per stick)
|
||||
│ usb-transfer protocol (channel received via hello, by lineage)
|
||||
▼
|
||||
bus driver ── one process per controller (usb-xhci-bus)
|
||||
│ hardware
|
||||
```
|
||||
|
||||
Three kinds of boundary, deliberately different:
|
||||
|
||||
- **Protocol boundaries** sit where processes meet (vfs, block, usb-transfer).
|
||||
Each side is independently restartable; channels are established by
|
||||
capability handoff, never by registry names (communication.md
|
||||
"Establishment: two planes"). Steady-state file I/O crosses exactly two —
|
||||
vfs and block.
|
||||
- **Library boundaries** sit where concerns meet inside one process. The
|
||||
filesystem service's engine (on-disk format logic, host-testable, behind
|
||||
the four-function `BlockDevice` vtable) and shell (IPC serving, open-node
|
||||
table, mount registration) talk through a Zig API. No protocol between
|
||||
them: they share fate regardless — a corrupt engine corrupts the answers
|
||||
either way — so a channel there would add a hop per file operation and buy
|
||||
nothing. The seam exists for compile-time testability and reuse; the
|
||||
*process* is the restart unit.
|
||||
- **Control-plane relationships** sit beside the data path, never on it. Two
|
||||
supervisors, one per layer: the **device manager** wires and revives the
|
||||
device layers (bus and storage drivers — devices only); the **volume
|
||||
manager** *(planned)* wires and revives the volume layer (filesystem
|
||||
services). Neither touches steady-state I/O.
|
||||
|
||||
## Who does what
|
||||
|
||||
**Bus driver** (usb-xhci-bus, per controller): hardware only. Enumerates,
|
||||
reports children to the device manager, serves the transfer contract to its
|
||||
own children's class drivers, routed by lineage.
|
||||
|
||||
**Storage driver** (usb-storage, per device): hardware → blocks. Speaks SCSI
|
||||
Bulk-Only Transport upward from its device; serves the block protocol (five
|
||||
verbs: geometry, read, write, flush, attach). **Content-blind, permanently**:
|
||||
it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no
|
||||
GPT, no filesystem magic, ever. *(Planned)* it gains one content-blind
|
||||
mechanism: **named sub-ranges** — "serve blocks [a, b) as a channel of the
|
||||
same block contract". It clamps and offsets; it never knows the numbers came
|
||||
from a partition table. The clamp lives here and nowhere else because a
|
||||
channel must carry exactly the authority it grants: handing a filesystem the
|
||||
whole disk plus a polite base offset would let a buggy or compromised
|
||||
filesystem scribble the neighboring partition — the same authority-overshoot
|
||||
the device-authority track eliminated for MMIO and DMA.
|
||||
|
||||
**Volume manager** *(planned; today the FAT service squats on these duties)*:
|
||||
the policy home of the volume layer, one service, supervised by init. It
|
||||
subscribes to the device manager's child events; when a storage provider
|
||||
appears it consumer-hellos for the block channel, reads the partition table
|
||||
and the first blocks itself (**it** is the prober), consults its
|
||||
configuration, defines sub-ranges on the driver, spawns the matching
|
||||
filesystem service per volume with that volume's channel, supervises it, and
|
||||
decides mount placement. Its tables are CSV configuration, read by it (the
|
||||
policy), enforced by nobody else:
|
||||
|
||||
- `filesystems.csv` — content signature → filesystem binary. Adding a
|
||||
filesystem adds a row.
|
||||
- the mount map — volume identity → mount prefix. The boot volume is chosen
|
||||
by **content** (the volume carrying `/system/configuration` and
|
||||
`/system/logs`), never by port or arrival order.
|
||||
|
||||
**Filesystem service** (the FAT service today; one process per volume): the
|
||||
proven unit — block-client + engine + file-protocol provider in one binary. It
|
||||
receives its block channel at spawn; it never discovers devices. It registers
|
||||
its own mounts with the kernel; its write cache lives inside the process, so a
|
||||
write error is observed by the code that owns the volume and surfaces on the
|
||||
owning channel (the anti-fsyncgate rule — never a system-wide dirty pool).
|
||||
*(Today, interim:)* fat also acquires its own volume (first mass-storage child
|
||||
by enumeration order) and hardcodes its mount prefixes; both migrate to the
|
||||
volume manager.
|
||||
|
||||
**Kernel** (mechanism only): the mount table routes paths to backend
|
||||
endpoints — resolve and redirect, never data. Remount-replace is the restart
|
||||
story; a dead backend's slot is swept lazily on the next resolution.
|
||||
*(Planned fix:)* `fs_unmount` gains ownership — only the mounting endpoint's
|
||||
holder may unmount; today it is ungated, which is safe with one mount owner
|
||||
and wrong with several.
|
||||
|
||||
## Adding a filesystem
|
||||
|
||||
1. **Write the engine**: pure Zig, no IPC, behind the `BlockDevice` vtable
|
||||
(four functions). On-disk layout in `align(1)` extern structs. Host-test it
|
||||
with `zig build test` fixtures — the FAT engine
|
||||
([engine.zig](../../system/services/fat/engine.zig),
|
||||
[on-disk.zig](../../system/services/fat/on-disk.zig)) is the template, and
|
||||
its ~500 lines of host tests the standard.
|
||||
2. **Reuse the shell**: the filesystem harness — establishment, the
|
||||
badge-scoped open-node table, the nine vfs-protocol handlers, mount
|
||||
registration, the removal path — is shared code, not per-filesystem code.
|
||||
*(Planned: extracted from fat's 434-line shell into a library before the
|
||||
second engine is written.)* An engine plus a `main` wiring it into the
|
||||
harness is a complete filesystem service.
|
||||
3. **Add the configuration row**: one line in `filesystems.csv` mapping the
|
||||
on-disk signature to the binary. No other component changes: the volume
|
||||
manager probes, matches, spawns; the vfs protocol is already
|
||||
backend-neutral (a second provider serves it in production today — init's
|
||||
synthetic registry backend).
|
||||
4. **Prove removal**: extend the removable-media suite (below) for the new
|
||||
filesystem — surprise-yank during writes must lose only what was
|
||||
unflushed, said honestly, with the volume consistent enough to remount.
|
||||
|
||||
The engine never sees: partitions (it receives a volume-shaped block channel;
|
||||
base offsets are the driver's clamp), device discovery (the channel arrives at
|
||||
spawn), mount policy (prefixes are handed to it), or other volumes (one
|
||||
process, one volume).
|
||||
|
||||
## Removable media: responsibilities on removal, per layer
|
||||
|
||||
The design rule, learned from what Linux cannot do: there is exactly **one**
|
||||
surprise-removal path — kill the filesystem process, retire its mounts,
|
||||
respawn on return. No half-alive states, no `remount-ro`, no mounts that
|
||||
error forever (Plan 9's dead-server wart).
|
||||
|
||||
| Layer | Observes | Must do | Guarantees |
|
||||
|---|---|---|---|
|
||||
| Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick |
|
||||
| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
|
||||
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
|
||||
| Volume manager *(planned)* | the storage provider's channel death / the manager's child events | kill each filesystem service of that device's volumes; retire their kernel mounts; remember the volume identity | one removal path; mounts never dangle; log persistence stops *cleanly* |
|
||||
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the FAT dirty flag on disk marks the unclean removal; the process never serves from behind a dead channel |
|
||||
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); *(planned)* ownership-gated `fs_unmount` | resolution under a dead mount is `not_found`, not a hang |
|
||||
| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
|
||||
|
||||
**On return** the same table runs upward in reverse: the bus re-enumerates and
|
||||
re-reports; the device manager respawns the storage driver (built); the volume
|
||||
manager re-probes — same content, same volume identity — respawns the
|
||||
filesystem service, and remounts at the same prefix; applications see the
|
||||
subtree reappear. A different stick in the same port is a *different volume*
|
||||
(identity is content, not port) and mounts wherever its identity says —
|
||||
possibly nowhere but `/volumes/…`.
|
||||
|
||||
What is lost on a surprise yank is exactly the write-back window of the
|
||||
filesystem service, no more: the engine owns its cache, so the blast radius of
|
||||
a yank is one volume's unflushed writes, marked by the on-disk dirty flag. A
|
||||
filesystem format with better crash honesty (journaling, copy-on-write — the
|
||||
lesson of QNX's Power-Safe) narrows that window further and slots in as
|
||||
implementation N+1 through the table above, changing nothing else.
|
||||
Reference in New Issue
Block a user