The settled shape from the storage-stack discussion, written as the reference: the data path (vfs -> filesystem service -> block -> driver) as the application/service/protocol/driver model applied twice; the three kinds of boundary (protocol between processes, library inside them, control-plane beside them); the volume manager as the policy home (planned — the FAT service squats on its duties today, marked as such); adding a filesystem as engine + shared harness + one configuration row; and the per-layer responsibility table for removable media — one removal path, kill/retire/ respawn, dirty data lost and SAID to be lost. Indexed from docs/README.md.
164 lines
10 KiB
Markdown
164 lines
10 KiB
Markdown
# The storage architecture: layers, boundaries, responsibilities
|
|
|
|
> **Status:** the layered model below is the settled design
|
|
> ([storage-stack-discussion.md](../storage-stack-discussion.md) records how it
|
|
> was reached and what the surveyed systems taught). The data path — vfs
|
|
> protocol, kernel mount routing, the FAT service, the block protocol,
|
|
> usb-storage — is **built**. The volume manager, the driver's range
|
|
> mechanism, per-volume filesystem spawning, and the removal lifecycle are
|
|
> **planned**; until they land, the FAT service performs volume-manager duties
|
|
> itself (marked below). This document is the reference for both states.
|
|
|
|
## The model
|
|
|
|
One pattern, applied twice: an application is a client of a service over a
|
|
protocol; a service is a client of a driver over a protocol.
|
|
|
|
```
|
|
application
|
|
│ vfs protocol (routed by the kernel mount table)
|
|
▼
|
|
filesystem service ── one process per volume
|
|
│ inside: [file-protocol shell] ── Zig API ── [fs engine]
|
|
│ block protocol (channel received at spawn)
|
|
▼
|
|
storage driver ── one process per device (usb-storage per stick)
|
|
│ usb-transfer protocol (channel received via hello, by lineage)
|
|
▼
|
|
bus driver ── one process per controller (usb-xhci-bus)
|
|
│ hardware
|
|
```
|
|
|
|
Three kinds of boundary, deliberately different:
|
|
|
|
- **Protocol boundaries** sit where processes meet (vfs, block, usb-transfer).
|
|
Each side is independently restartable; channels are established by
|
|
capability handoff, never by registry names (communication.md
|
|
"Establishment: two planes"). Steady-state file I/O crosses exactly two —
|
|
vfs and block.
|
|
- **Library boundaries** sit where concerns meet inside one process. The
|
|
filesystem service's engine (on-disk format logic, host-testable, behind
|
|
the four-function `BlockDevice` vtable) and shell (IPC serving, open-node
|
|
table, mount registration) talk through a Zig API. No protocol between
|
|
them: they share fate regardless — a corrupt engine corrupts the answers
|
|
either way — so a channel there would add a hop per file operation and buy
|
|
nothing. The seam exists for compile-time testability and reuse; the
|
|
*process* is the restart unit.
|
|
- **Control-plane relationships** sit beside the data path, never on it. Two
|
|
supervisors, one per layer: the **device manager** wires and revives the
|
|
device layers (bus and storage drivers — devices only); the **volume
|
|
manager** *(planned)* wires and revives the volume layer (filesystem
|
|
services). Neither touches steady-state I/O.
|
|
|
|
## Who does what
|
|
|
|
**Bus driver** (usb-xhci-bus, per controller): hardware only. Enumerates,
|
|
reports children to the device manager, serves the transfer contract to its
|
|
own children's class drivers, routed by lineage.
|
|
|
|
**Storage driver** (usb-storage, per device): hardware → blocks. Speaks SCSI
|
|
Bulk-Only Transport upward from its device; serves the block protocol (five
|
|
verbs: geometry, read, write, flush, attach). **Content-blind, permanently**:
|
|
it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no
|
|
GPT, no filesystem magic, ever. *(Planned)* it gains one content-blind
|
|
mechanism: **named sub-ranges** — "serve blocks [a, b) as a channel of the
|
|
same block contract". It clamps and offsets; it never knows the numbers came
|
|
from a partition table. The clamp lives here and nowhere else because a
|
|
channel must carry exactly the authority it grants: handing a filesystem the
|
|
whole disk plus a polite base offset would let a buggy or compromised
|
|
filesystem scribble the neighboring partition — the same authority-overshoot
|
|
the device-authority track eliminated for MMIO and DMA.
|
|
|
|
**Volume manager** *(planned; today the FAT service squats on these duties)*:
|
|
the policy home of the volume layer, one service, supervised by init. It
|
|
subscribes to the device manager's child events; when a storage provider
|
|
appears it consumer-hellos for the block channel, reads the partition table
|
|
and the first blocks itself (**it** is the prober), consults its
|
|
configuration, defines sub-ranges on the driver, spawns the matching
|
|
filesystem service per volume with that volume's channel, supervises it, and
|
|
decides mount placement. Its tables are CSV configuration, read by it (the
|
|
policy), enforced by nobody else:
|
|
|
|
- `filesystems.csv` — content signature → filesystem binary. Adding a
|
|
filesystem adds a row.
|
|
- the mount map — volume identity → mount prefix. The boot volume is chosen
|
|
by **content** (the volume carrying `/system/configuration` and
|
|
`/system/logs`), never by port or arrival order.
|
|
|
|
**Filesystem service** (the FAT service today; one process per volume): the
|
|
proven unit — block-client + engine + file-protocol provider in one binary. It
|
|
receives its block channel at spawn; it never discovers devices. It registers
|
|
its own mounts with the kernel; its write cache lives inside the process, so a
|
|
write error is observed by the code that owns the volume and surfaces on the
|
|
owning channel (the anti-fsyncgate rule — never a system-wide dirty pool).
|
|
*(Today, interim:)* fat also acquires its own volume (first mass-storage child
|
|
by enumeration order) and hardcodes its mount prefixes; both migrate to the
|
|
volume manager.
|
|
|
|
**Kernel** (mechanism only): the mount table routes paths to backend
|
|
endpoints — resolve and redirect, never data. Remount-replace is the restart
|
|
story; a dead backend's slot is swept lazily on the next resolution.
|
|
*(Planned fix:)* `fs_unmount` gains ownership — only the mounting endpoint's
|
|
holder may unmount; today it is ungated, which is safe with one mount owner
|
|
and wrong with several.
|
|
|
|
## Adding a filesystem
|
|
|
|
1. **Write the engine**: pure Zig, no IPC, behind the `BlockDevice` vtable
|
|
(four functions). On-disk layout in `align(1)` extern structs. Host-test it
|
|
with `zig build test` fixtures — the FAT engine
|
|
([engine.zig](../../system/services/fat/engine.zig),
|
|
[on-disk.zig](../../system/services/fat/on-disk.zig)) is the template, and
|
|
its ~500 lines of host tests the standard.
|
|
2. **Reuse the shell**: the filesystem harness — establishment, the
|
|
badge-scoped open-node table, the nine vfs-protocol handlers, mount
|
|
registration, the removal path — is shared code, not per-filesystem code.
|
|
*(Planned: extracted from fat's 434-line shell into a library before the
|
|
second engine is written.)* An engine plus a `main` wiring it into the
|
|
harness is a complete filesystem service.
|
|
3. **Add the configuration row**: one line in `filesystems.csv` mapping the
|
|
on-disk signature to the binary. No other component changes: the volume
|
|
manager probes, matches, spawns; the vfs protocol is already
|
|
backend-neutral (a second provider serves it in production today — init's
|
|
synthetic registry backend).
|
|
4. **Prove removal**: extend the removable-media suite (below) for the new
|
|
filesystem — surprise-yank during writes must lose only what was
|
|
unflushed, said honestly, with the volume consistent enough to remount.
|
|
|
|
The engine never sees: partitions (it receives a volume-shaped block channel;
|
|
base offsets are the driver's clamp), device discovery (the channel arrives at
|
|
spawn), mount policy (prefixes are handed to it), or other volumes (one
|
|
process, one volume).
|
|
|
|
## Removable media: responsibilities on removal, per layer
|
|
|
|
The design rule, learned from what Linux cannot do: there is exactly **one**
|
|
surprise-removal path — kill the filesystem process, retire its mounts,
|
|
respawn on return. No half-alive states, no `remount-ro`, no mounts that
|
|
error forever (Plan 9's dead-server wart).
|
|
|
|
| Layer | Observes | Must do | Guarantees |
|
|
|---|---|---|---|
|
|
| Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick |
|
|
| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
|
|
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
|
|
| Volume manager *(planned)* | the storage provider's channel death / the manager's child events | kill each filesystem service of that device's volumes; retire their kernel mounts; remember the volume identity | one removal path; mounts never dangle; log persistence stops *cleanly* |
|
|
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the FAT dirty flag on disk marks the unclean removal; the process never serves from behind a dead channel |
|
|
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); *(planned)* ownership-gated `fs_unmount` | resolution under a dead mount is `not_found`, not a hang |
|
|
| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
|
|
|
|
**On return** the same table runs upward in reverse: the bus re-enumerates and
|
|
re-reports; the device manager respawns the storage driver (built); the volume
|
|
manager re-probes — same content, same volume identity — respawns the
|
|
filesystem service, and remounts at the same prefix; applications see the
|
|
subtree reappear. A different stick in the same port is a *different volume*
|
|
(identity is content, not port) and mounts wherever its identity says —
|
|
possibly nowhere but `/volumes/…`.
|
|
|
|
What is lost on a surprise yank is exactly the write-back window of the
|
|
filesystem service, no more: the engine owns its cache, so the blast radius of
|
|
a yank is one volume's unflushed writes, marked by the on-disk dirty flag. A
|
|
filesystem format with better crash honesty (journaling, copy-on-write — the
|
|
lesson of QNX's Power-Safe) narrows that window further and slots in as
|
|
implementation N+1 through the table above, changing nothing else.
|