Compare commits
8
Commits
d4f8dc51b9
...
9da63e81e2
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
9da63e81e2 | ||
|
|
ada251150a | ||
|
|
728b436d0f | ||
|
|
092817ba2e | ||
|
|
a44b397bed | ||
|
|
fa54ef6915 | ||
|
|
fae616fa5a | ||
|
|
5a8a2a4d7e |
+5
-1
@@ -51,7 +51,11 @@ rather than restate it. Roughly in the order things happen at runtime:
|
||||
14. **[vfs-protocol.md](file-system-development/vfs-protocol.md) — the VFS wire protocol.** The language-neutral
|
||||
byte-level spec of the file protocol spoken over IPC: request/reply headers,
|
||||
the operation table, mount routing, and the append-only evolution rules — the
|
||||
first IPC protocol documented as public ABI.
|
||||
first IPC protocol documented as public ABI. Its architectural frame is
|
||||
**[storage-architecture.md](file-system-development/storage-architecture.md) — the storage stack**:
|
||||
the layers from application to hardware, the three kinds of boundary
|
||||
(protocol, library, control-plane), how to add a filesystem, and the
|
||||
per-layer responsibilities when removable media is yanked and returned.
|
||||
15. **[drivers.md](device-driver-development/drivers.md) — writing a driver.** The payoff: a driver is an
|
||||
ordinary ring-3 process that claims a device, maps its registers, and **sleeps
|
||||
until its hardware interrupts it**. The claim is the capability; `irq_ack` is the
|
||||
|
||||
@@ -0,0 +1,255 @@
|
||||
# The storage architecture: layers, boundaries, responsibilities
|
||||
|
||||
> **Status:** the layered model below is the settled design
|
||||
> ([storage-design-rationale.md](storage-design-rationale.md) records how it
|
||||
> was reached and what the surveyed systems taught). The data path — vfs
|
||||
> protocol, kernel mount routing, the FAT service, the block protocol,
|
||||
> usb-storage — is **built**. The volume manager, the driver's range
|
||||
> mechanism, per-volume filesystem spawning, and the removal lifecycle are
|
||||
> **planned**; until they land, the FAT service performs volume-manager duties
|
||||
> itself (marked below). This document is the reference for both states.
|
||||
|
||||
## The model
|
||||
|
||||
One pattern, applied twice: an application is a client of a service over a
|
||||
protocol; a service is a client of a driver over a protocol.
|
||||
|
||||
```
|
||||
application
|
||||
│ vfs protocol (routed by the kernel mount table)
|
||||
▼
|
||||
filesystem service ── one process per volume
|
||||
│ inside: [file-protocol shell] ── Zig API ── [fs engine]
|
||||
│ block protocol (channel received at spawn)
|
||||
▼
|
||||
storage driver ── one process per device (usb-storage per stick)
|
||||
│ usb-transfer protocol (channel received via hello, by lineage)
|
||||
▼
|
||||
bus driver ── one process per controller (usb-xhci-bus)
|
||||
│ hardware
|
||||
```
|
||||
|
||||
Three kinds of boundary, deliberately different:
|
||||
|
||||
- **Protocol boundaries** sit where processes meet (vfs, block, usb-transfer).
|
||||
Each side is independently restartable; channels are established by
|
||||
capability handoff, never by registry names (communication.md
|
||||
"Establishment: two planes"). Steady-state file I/O crosses exactly two —
|
||||
vfs and block.
|
||||
- **Library boundaries** sit where concerns meet inside one process. The
|
||||
filesystem service's engine (on-disk format logic, host-testable, behind
|
||||
the four-function `BlockDevice` vtable) and shell (IPC serving, open-node
|
||||
table, mount registration) talk through a Zig API. No protocol between
|
||||
them: they share fate regardless — a corrupt engine corrupts the answers
|
||||
either way — so a channel there would add a hop per file operation and buy
|
||||
nothing. The seam exists for compile-time testability and reuse; the
|
||||
*process* is the restart unit.
|
||||
- **Control-plane relationships** sit beside the data path, never on it. Two
|
||||
supervisors, one per layer: the **device manager** wires and revives the
|
||||
device layers (bus and storage drivers — devices only); the **volume
|
||||
manager** *(planned)* wires and revives the volume layer (filesystem
|
||||
services). Neither touches steady-state I/O.
|
||||
|
||||
## Who does what
|
||||
|
||||
**Bus driver** (usb-xhci-bus, per controller): hardware only. Enumerates,
|
||||
reports children to the device manager, serves the transfer contract to its
|
||||
own children's class drivers, routed by lineage.
|
||||
|
||||
**Storage driver** (usb-storage, per device; later nvme, ahci): hardware →
|
||||
blocks. Speaks its transport upward from its device; serves the block
|
||||
protocol (geometry, read, write, flush, attach, detach — and, *planned*, the
|
||||
pushed `medium_changed` presence event, translated from the transport's
|
||||
native signal). **Content-blind, permanently**:
|
||||
it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no
|
||||
GPT, no filesystem magic, ever. *(Planned)* it gains one content-blind
|
||||
mechanism: **named sub-ranges** — "serve blocks [a, b) as a channel of the
|
||||
same block contract". It clamps and offsets; it never knows the numbers came
|
||||
from a partition table. The clamp lives here and nowhere else because a
|
||||
channel must carry exactly the authority it grants: handing a filesystem the
|
||||
whole disk plus a polite base offset would let a buggy or compromised
|
||||
filesystem scribble the neighboring partition — the same authority-overshoot
|
||||
the device-authority track eliminated for MMIO and DMA.
|
||||
|
||||
**Volume manager** *(planned; today the FAT service squats on these duties)*:
|
||||
the policy home of the volume layer, one service, supervised by init. It
|
||||
subscribes to the device manager's child events; when a storage provider
|
||||
appears it consumer-hellos for the block channel, reads the partition table
|
||||
and the first blocks itself (**it** is the prober), consults its
|
||||
configuration, defines sub-ranges on the driver, spawns the matching
|
||||
filesystem service per volume with that volume's channel, supervises it, and
|
||||
decides mount placement. Its tables are CSV configuration, read by it (the
|
||||
policy), enforced by nobody else:
|
||||
|
||||
- `filesystems.csv` — content signature → filesystem binary. Adding a
|
||||
filesystem adds a row.
|
||||
- `volumes.csv` — the mount map, danos's fstab: **volume identity → mount
|
||||
prefix**, keyed on content identity and never on port, path, or arrival
|
||||
order (the lesson of Linux's `/dev/sda1`-era fstab, which broke on every
|
||||
port move until `UUID=` replaced it). Identity is read off the medium by
|
||||
the prober, strongest first: GPT partition GUID → filesystem UUID → FAT
|
||||
serial + label → MBR signature + partition index → anonymous (generated
|
||||
name, no persistence). Same identity, same mount point: a drive moved to
|
||||
another port — USB to another hub, SATA to another bay, even a stick
|
||||
returning in a different dock — lands exactly where it was. Duplicate
|
||||
identity (cloned sticks, together) is policy: first keeps the name, the
|
||||
second mounts suffixed and is logged loudly. The boot volume is the
|
||||
recorded identity of the volume carrying `/system/configuration` and
|
||||
`/system/logs`, findable on any port. Unknown volumes mount under
|
||||
`/volumes/<derived name>`.
|
||||
|
||||
**Filesystem service** (the FAT service today; one process per volume): the
|
||||
proven unit — block-client + engine + file-protocol provider in one binary. It
|
||||
receives its block channel at spawn; it never discovers devices. It registers
|
||||
its own mounts with the kernel; its write cache lives inside the process, so a
|
||||
write error is observed by the code that owns the volume and surfaces on the
|
||||
owning channel (the anti-fsyncgate rule — never a system-wide dirty pool).
|
||||
*(Today, interim:)* fat also acquires its own volume (first mass-storage child
|
||||
by enumeration order) and hardcodes its mount prefixes; both migrate to the
|
||||
volume manager.
|
||||
|
||||
**Kernel** (mechanism only): the mount table routes paths to backend
|
||||
endpoints — resolve and redirect, never data. Remount-replace is the restart
|
||||
story; a dead backend's slot is swept lazily on the next resolution.
|
||||
*(Planned fix:)* `fs_unmount` gains ownership — only the mounting endpoint's
|
||||
holder may unmount; today it is ungated, which is safe with one mount owner
|
||||
and wrong with several.
|
||||
|
||||
## Adding a filesystem
|
||||
|
||||
1. **Write the engine**: pure Zig, no IPC, behind the `BlockDevice` vtable
|
||||
(four functions). On-disk layout in `align(1)` extern structs. Host-test it
|
||||
with `zig build test` fixtures — the FAT engine
|
||||
([engine.zig](../../system/services/fat/engine.zig),
|
||||
[on-disk.zig](../../system/services/fat/on-disk.zig)) is the template, and
|
||||
its ~500 lines of host tests the standard.
|
||||
2. **Reuse the shell**: the filesystem harness — establishment, the
|
||||
badge-scoped open-node table, the nine vfs-protocol handlers, mount
|
||||
registration, the removal path — is shared code, not per-filesystem code.
|
||||
*(Planned: extracted from fat's 434-line shell into a library before the
|
||||
second engine is written.)* An engine plus a `main` wiring it into the
|
||||
harness is a complete filesystem service.
|
||||
3. **Add the configuration row**: one line in `filesystems.csv` mapping the
|
||||
on-disk signature to the binary. No other component changes: the volume
|
||||
manager probes, matches, spawns; the vfs protocol is already
|
||||
backend-neutral (a second provider serves it in production today — init's
|
||||
synthetic registry backend).
|
||||
4. **Prove removal**: extend the removable-media suite (below) for the new
|
||||
filesystem — surprise-yank during writes must lose only what was
|
||||
unflushed, said honestly, with the volume consistent enough to remount.
|
||||
|
||||
The engine never sees: partitions (it receives a volume-shaped block channel;
|
||||
base offsets are the driver's clamp), device discovery (the channel arrives at
|
||||
spawn), mount policy (prefixes are handed to it), or other volumes (one
|
||||
process, one volume).
|
||||
|
||||
## Removable media: responsibilities on removal, per layer
|
||||
|
||||
The design rule, learned from what Linux cannot do: there is exactly **one**
|
||||
surprise-removal path — kill the filesystem process, retire its mounts,
|
||||
respawn on return. No half-alive states, no `remount-ro`, no mounts that
|
||||
error forever (Plan 9's dead-server wart).
|
||||
|
||||
The path has **two triggers, one lifecycle**: the *device* leaving (the
|
||||
storage driver dies — channel death, the table below), and the *medium*
|
||||
leaving while the device stays (an SD card pulled from its reader, an ATAPI
|
||||
tray opened — including USB card readers today). The second trigger is a
|
||||
pushed `medium_changed` event on the block protocol *(planned)*: the storage
|
||||
driver translates its transport's native signal (SCSI UNIT ATTENTION, AHCI
|
||||
PxSSTS, NVMe namespace-change AER) into presence-changed — never content —
|
||||
and the volume manager runs the same kill-retire path, then re-probes on
|
||||
medium return exactly as on device return. Without it, a swapped card would
|
||||
be served with the previous card's filesystem state.
|
||||
|
||||
| Layer | Observes | Must do | Guarantees |
|
||||
|---|---|---|---|
|
||||
| Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick |
|
||||
| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
|
||||
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
|
||||
| Volume manager *(planned)* | the storage provider's channel death / the manager's child events | kill each filesystem service of that device's volumes; retire their kernel mounts; remember the volume identity | one removal path; mounts never dangle; log persistence stops *cleanly* |
|
||||
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the FAT dirty flag on disk marks the unclean removal; the process never serves from behind a dead channel |
|
||||
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); *(planned)* ownership-gated `fs_unmount` | resolution under a dead mount is `not_found`, not a hang |
|
||||
| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
|
||||
|
||||
**On return** the same table runs upward in reverse: the bus re-enumerates and
|
||||
re-reports; the device manager respawns the storage driver (built); the volume
|
||||
manager re-probes — same content, same volume identity — respawns the
|
||||
filesystem service, and remounts at the same prefix; applications see the
|
||||
subtree reappear. A different stick in the same port is a *different volume*
|
||||
(identity is content, not port) and mounts wherever its identity says —
|
||||
possibly nowhere but `/volumes/…`.
|
||||
|
||||
### The boot volume: identity is what makes yanking it survivable
|
||||
|
||||
The boot volume is the sharpest instance of the return story, because the
|
||||
system's own log persistence rides it (`/system/logs`), and because nothing
|
||||
else about the running system depends on it at all — every binary and the
|
||||
boot-time configuration live in the initial ramdisk. Its lifecycle,
|
||||
end to end:
|
||||
|
||||
- **At first mount** the volume manager records the identity of the volume
|
||||
carrying `/system/configuration` and `/system/logs` — THE boot volume,
|
||||
from then on a fact about content, not about a port.
|
||||
- **While it is absent**, the system runs on. The kernel log ring keeps
|
||||
accumulating — it is the buffer that makes the absence survivable — and
|
||||
the logger keeps draining it; only *persistence* pauses. File operations
|
||||
under the retired mounts fail honestly (`not_found`). The ring is bounded,
|
||||
so a long absence overwrites its oldest entries: that window is the data
|
||||
loss, and it must be *said* — the logger marks the gap in the file when
|
||||
persistence resumes, never splicing the stream silently.
|
||||
- **On return — any port, any hub, even a different transport** — the
|
||||
prober reads the same identity, the mount map answers with the same
|
||||
prefixes, the filesystem service is respawned, and `/system/logs` is the
|
||||
same tree it was: the logger resumes appending into the SAME
|
||||
`<boot-stamp>` directory, per-binary files continuing where they left
|
||||
off (plus the gap marker if the ring wrapped).
|
||||
- **What never comes back** is the write-back window lost at the yank —
|
||||
the dirty-honesty rule, unchanged; the boot volume gets no exemption.
|
||||
|
||||
This is also the requirement that shapes the logger: it treats the log tree
|
||||
as a volume that comes and goes — failed flushes are retried on the same
|
||||
patient cadence the fat service already uses for storage that arrives late,
|
||||
never abandoned after the first `not_found`.
|
||||
|
||||
What is lost on a surprise yank is exactly the write-back window of the
|
||||
filesystem service, no more: the engine owns its cache, so the blast radius of
|
||||
a yank is one volume's unflushed writes, marked by the on-disk dirty flag. A
|
||||
filesystem format with better crash honesty (journaling, copy-on-write — the
|
||||
lesson of QNX's Power-Safe) narrows that window further and slots in as
|
||||
implementation N+1 through the table above, changing nothing else.
|
||||
|
||||
## How the lifecycle is enforced
|
||||
|
||||
Responsibility tables are convention; convention is worthless (the bounds
|
||||
audit's lesson). The volume lifecycle is enforced with the same three levers
|
||||
the device lifecycle already uses, mapped one-to-one:
|
||||
|
||||
1. **A filesystem cannot acquire — it can only be given.** Filesystem
|
||||
binaries hold NO establishment grants: no `open device-manager`, no
|
||||
registry name to look up. The only block channel a filesystem process ever
|
||||
has is the one handed to it at spawn by the volume manager. Serving the
|
||||
wrong volume, a second volume, or a self-discovered volume is not
|
||||
forbidden but *impossible* — the same way a driver cannot claim hardware
|
||||
it was not delegated (device-authority.md).
|
||||
2. **The volume manager supervises with teeth.** Mount-within-deadline or be
|
||||
stopped (a filesystem wedged on a corrupt volume is killed and crash-loop
|
||||
capped; the volume is marked bad, not retried forever). On removal the
|
||||
manager KILLS the filesystem process and retires its mounts — the polite
|
||||
observe-`EPEER`-and-exit path is an optimization; the kill is the
|
||||
guarantee, because a process cannot be trusted to observe its own
|
||||
obsolescence (the reap argument, proven on the device tree).
|
||||
3. **The shared harness is how every filesystem inherits the lifecycle by
|
||||
construction.** The harness — not the engine — owns the state machine:
|
||||
establishment at spawn, mount registration, the dirty-flag set/clear
|
||||
bracket, error-out-and-exit on channel death. The engine sits behind the
|
||||
four-function vtable and never sees a channel; it cannot opt out of the
|
||||
lifecycle for the same reason it cannot find a device. This is why the
|
||||
harness is extracted BEFORE the second engine is written.
|
||||
|
||||
Beneath all three, two kernel mechanisms: ownership-gated `fs_unmount` (only
|
||||
the mounting endpoint's holder may unmount), and the lazy dead-endpoint mount
|
||||
sweep as the backstop nothing can disable. And above them, the check: a
|
||||
**lifecycle conformance drill** — mount, serve, yank mid-write, verify honest
|
||||
loss, replug, remount — parameterized over filesystem implementations, so
|
||||
"danos supports filesystem X" MEANS "X passes the drill through the harness",
|
||||
exactly as provider conformance means passing the reserved-verb suite.
|
||||
@@ -0,0 +1,225 @@
|
||||
# The storage design rationale: why the stack is shaped this way
|
||||
|
||||
*2026-08-09. The design record behind
|
||||
[storage-architecture.md](storage-architecture.md): the survey, the
|
||||
trade-offs, and the decisions with their reasons — kept so future changes
|
||||
argue against the evidence rather than rediscovering it. The questions that
|
||||
drove it, verbatim: should the block protocol be separate from the VFS? how
|
||||
do channels work with these block devices? how should we wire up different
|
||||
filesystems? Grounded in how Minix 3, QNX Neutrino, Fuchsia, Redox, Plan 9,
|
||||
and Linux each answered the same questions, and in exactly where our own
|
||||
seams sat. The layering rule it serves: drivers are the lowest level
|
||||
(hardware only); VFS and the filesystems are higher layers; the protocol
|
||||
layer routes between them.*
|
||||
|
||||
## What the survey says, compressed
|
||||
|
||||
**Everyone separates block from VFS.** Even Plan 9, which unifies the *protocol*
|
||||
(a disk is a file tree), keeps the layers: kernel `sd` serves raw ranges,
|
||||
user-space filesystem servers consume them and serve files. The two poles are
|
||||
Minix 3 (every layer a process: one VFS server, one filesystem server *per
|
||||
mounted volume*, driver processes below — driver crashes proven survivable by
|
||||
fault-injection, tens of thousands of injected faults) and QNX (filesystems as
|
||||
DLLs inside the disk-driver process — fastest path, but one corrupt volume takes
|
||||
every volume behind that controller down). Fuchsia is the modern capability-
|
||||
native reference: a tiny block contract that every layer speaks, control channel
|
||||
separate from a data fast-path with pre-registered buffers, filesystems as
|
||||
separate processes launched by a storage-policy component (`fshost`) that probes
|
||||
content and hands each filesystem its block channel at startup. The filesystem
|
||||
never discovers devices.
|
||||
|
||||
**Partition tables are parsed in user space, below the filesystem, above the
|
||||
raw device — never in the kernel, never in the filesystem engine.** Fuchsia
|
||||
tried partitions-as-drivers and formally reversed it (RFC-0257). Plan 9 has the
|
||||
cleanest split of all: the kernel offers only a *mechanism* — "create a named
|
||||
sub-range of this disk" — and a user-space prober parses the table and issues
|
||||
those commands. QNX (`io-blk`) and Minix (driver-side shared library) agree on
|
||||
placement: the block layer multiplexes one physical block object into several
|
||||
logical block objects **of the same contract**.
|
||||
|
||||
**The failure lessons are unanimous.** Linux's shared page cache across the
|
||||
filesystem boundary produced fsyncgate (write errors observed by the wrong
|
||||
process, dirty pages marked clean); the lesson is to keep write caching *inside*
|
||||
the filesystem process, so an error surfaces on the channel that owns the
|
||||
volume. Linux's `errors=remount-ro`/lazy-unmount half-alive states exist because
|
||||
a monolith cannot kill and respawn a filesystem; Plan 9 leaves dead servers'
|
||||
mounts in the namespace erroring forever. A restartable-parts system should have
|
||||
exactly one surprise-removal path: kill the filesystem process, retire its
|
||||
mounts, respawn on return. QNX's deepest lesson: after years of removable-media
|
||||
pain they concluded detection isn't enough and built a copy-on-write filesystem
|
||||
(Power-Safe) — *format-level* crash honesty. And QNX's `mcd` (a standalone
|
||||
policy daemon with declarative insertion/removal rules) plus Fuchsia's `fshost`
|
||||
agree that media policy is one dedicated component, not code scattered through
|
||||
filesystems.
|
||||
|
||||
## Where we already stand
|
||||
|
||||
Closer than expected. The block protocol is 61 lines, five verbs, and its
|
||||
docblock already reserves `Header.target` for volumes ("targets = volumes,
|
||||
`enumerate` lists them; nothing else about the protocol changes"). usb-storage
|
||||
is content-blind — it reads block 0 only as a bring-up self-check and parses
|
||||
nothing. The vfs protocol is already served by a second provider in production
|
||||
(init's registry serves it synthetically), so a second filesystem is *not* a
|
||||
protocol problem. The FAT service is already internally split: a host-testable
|
||||
engine (1659 + 310 lines) behind a four-function BlockDevice vtable, and a
|
||||
434-line IPC shell. The kernel mount table already has the right restart
|
||||
semantics (remount-replace; lazy dead-endpoint sweep on resolve).
|
||||
|
||||
The misplacements, all small: the MBR walk lives *inside the FAT engine*
|
||||
(`engine.zig:196-212`) — every future engine would duplicate it; fat hardcodes
|
||||
its acquisition policy ("first mass-storage child, mount at /volumes/usb") —
|
||||
that is policy in a filesystem binary; `fs_unmount` is gated by nothing (any
|
||||
process may unmount any prefix — latent now, an obvious cross-tenant hole once
|
||||
mounts multiply); and fat's `mounted` flag is one-way — no path back after its
|
||||
block channel dies.
|
||||
|
||||
## The proposed shape
|
||||
|
||||
Three layers, matching the stated rule, every boundary a channel:
|
||||
|
||||
**Block (driver layer, mechanism).** usb-storage stays hardware→blocks and
|
||||
content-blind. It gains one mechanism, Plan 9's: *named sub-ranges*. A
|
||||
`define_range`-style op creates a logical block object (a partition) clamped
|
||||
and offset onto the disk; sub-ranges are `target` ids on the same endpoint, or
|
||||
separate endpoints handed out per range — either way they speak the **same
|
||||
block contract** (Fuchsia's closed-under-layering rule), so a filesystem cannot
|
||||
tell whole-disk from partition. The driver never parses a table; it clamps
|
||||
ranges it is told about. Offsets translate at definition time, so the data path
|
||||
stays one hop (Fuchsia's session-mapping trick, for free).
|
||||
|
||||
**Volumes (service layer, policy).** One new service — the *volume manager*
|
||||
(fshost/mcd-shaped, init-supervised; the device manager stays devices-only). It
|
||||
subscribes to the device manager's `child_added`/`child_removed`, consumer-
|
||||
hellos for each new mass-storage provider's block channel, reads the partition
|
||||
table and the first blocks itself (the prober is policy), consults
|
||||
configuration — `filesystems.csv`: content signature → filesystem binary;
|
||||
`mounts.csv` or similar: volume identity → mount prefix — defines sub-ranges on
|
||||
the driver, spawns **one filesystem process per volume**, hands it its block
|
||||
channel at startup, and supervises it with the reap-and-rebuild idiom the
|
||||
device manager proved: provider dies or medium leaves → kill the filesystem
|
||||
process, retire its mounts; medium returns → re-probe, respawn, remount. The
|
||||
boot volume is chosen by *content* (which volume carries /system/configuration
|
||||
and /system/logs), closing the two-sticks question honestly.
|
||||
|
||||
**Filesystems (per volume, one process).** The proven unit everywhere from
|
||||
Plan 9's `dossrv` to Minix to Fuchsia: block-client + engine + file-protocol
|
||||
provider in one binary, one process per volume (9front practice; per-volume
|
||||
fault isolation is what our supervision makes cheap). fat's shell becomes a
|
||||
shared *filesystem harness* library before a second engine is written; the MBR
|
||||
walk moves out of the engine into the volume manager; write caching stays
|
||||
inside the process (the anti-fsyncgate rule). Each mounts its prefixes into the
|
||||
kernel mount table itself, exactly as today.
|
||||
|
||||
**Kernel: two small changes only.** `fs_unmount` gains ownership (only the
|
||||
mounting endpoint's holder may unmount — possession-is-capability, consistent
|
||||
with everything else), and the 8-slot mount table gets a declared bound or
|
||||
growth once volumes multiply. The mount table stays the router; per-process
|
||||
namespaces (Plan 9's extra) remain separable future work.
|
||||
|
||||
## Transport generality: NVMe and SATA against this design
|
||||
|
||||
The layering was chosen transport-agnostic on purpose (the 9front lesson:
|
||||
filesystem services cannot tell a USB stick from an ATA disk); NVMe and AHCI
|
||||
are the test of that claim, and they fit — with three named pressure points.
|
||||
|
||||
**What transfers untouched:** the block protocol (nothing USB in it), the
|
||||
volume manager, partitions/ranges, filesystems, mounts, and the whole
|
||||
enforcement stack — delegation, IOMMU confinement (these are first-party PCI
|
||||
DMA masters, so confinement applies even more directly than USB),
|
||||
reap-and-rebuild, the hot-plug lifecycle. Integration cost per transport is
|
||||
what the architecture promises: one driver binary plus one `devices.csv` row.
|
||||
|
||||
**NVMe** shortens the stack — no bus/class split; the driver IS the
|
||||
controller driver, one hop fewer than USB. Its structural novelty,
|
||||
**namespaces** (hardware-native multiple volumes behind one controller), is
|
||||
exactly the case the block protocol reserved on day one and decision 4 below
|
||||
settles as endpoint-per-volume. **AHCI** is a shape choice, not a problem: an
|
||||
HBA fronts up to 32 ports plus port multipliers — structurally a bus — so
|
||||
either mirror USB (ahci-bus + a per-port disk driver: maximum restart
|
||||
granularity, the proven shape) or mirror NVMe (one driver per controller, one
|
||||
block channel per port). Leaning per-port processes for consistency with the
|
||||
matrix-proven shape; genuinely open.
|
||||
|
||||
**The pressure points, honestly:**
|
||||
|
||||
1. **Multi-volume providers are reserved, not implemented.** The volume
|
||||
manager flow assumes one provider, one volume; NVMe namespaces make
|
||||
endpoint-per-volume real work with hardware demanding it.
|
||||
2. **The current transport will bottleneck NVMe.** Synchronous call/reply,
|
||||
one operation in flight, one bounce buffer — fine for a USB2 stick,
|
||||
forfeits an NVMe drive's queue depth and per-queue MSI-X. Correctness
|
||||
needs nothing; performance is gated on the shm-ring data plane
|
||||
communication.md already names (Fuchsia's FIFO+VMO is the precedent).
|
||||
NVMe is not worth building before that milestone.
|
||||
3. **Media lifecycle ≠ device lifecycle.** ATAPI trays and SD card readers
|
||||
(USB ones too, today) keep the DEVICE present while the MEDIUM comes and
|
||||
goes — a removal trigger our channel-death lifecycle does not carry. The
|
||||
fix is decision 7 below. (Related small item: fat's bounce sizing
|
||||
hardcodes 512-byte sectors; 4Kn drives need the geometry honored
|
||||
throughout.)
|
||||
|
||||
## The decisions on the table
|
||||
|
||||
1. **Partitions as driver-side sub-ranges** (mechanism in usb-storage, parsing
|
||||
in the volume manager) — versus a separate partition *process* re-serving
|
||||
block (Fuchsia's storage-host). Sub-ranges are ~20 lines of clamp in the
|
||||
driver and keep the data path one hop; a separate process is purer layering
|
||||
at the cost of a hop or session plumbing. Recommendation: sub-ranges.
|
||||
2. **The volume manager as a new service** owning probe, spawn, supervision,
|
||||
and mount policy — with `filesystems.csv` and the mount map as
|
||||
configuration. Recommendation: yes; it is the missing policy home that fat
|
||||
is currently squatting in.
|
||||
3. **One filesystem process per volume** (fat's binary becomes "the FAT
|
||||
implementation", spawned per FAT volume). Recommendation: yes — it extends
|
||||
recompile-and-restart-live to filesystems and isolates corrupt media.
|
||||
4. **Sub-range addressing**: `target` ids on the storage endpoint versus one
|
||||
endpoint per volume handed out by the driver. Endpoint-per-volume matches
|
||||
the establishment-plane machinery (a channel per party, caps at
|
||||
establishment) and keeps per-client badge scoping simple. Recommendation:
|
||||
endpoint per volume.
|
||||
5. **`fs_unmount` ownership** — a defect fix more than a decision.
|
||||
6. **Later, kept open**: the shm-ring data plane (communication.md already
|
||||
names it as the 256-byte ceiling's unlock — Fuchsia's FIFO+VMO is the
|
||||
precedent); format-level crash honesty (a Power-Safe-style journaling or COW
|
||||
filesystem) once danos outgrows FAT; per-process namespaces.
|
||||
7. **The media-presence event** (settled in principle; lands with the volume
|
||||
manager): the block protocol gains a pushed event — `medium_changed`, with
|
||||
present/absent and a change counter — produced by the storage driver from
|
||||
its transport's native signal (SCSI UNIT ATTENTION / TEST UNIT READY for
|
||||
USB and ATAPI, PxSSTS for AHCI, namespace-change AER for NVMe) and
|
||||
consumed by the volume manager, which runs the SAME kill-retire-remount
|
||||
path it runs on channel death — one lifecycle, two triggers. The driver
|
||||
reports presence, never content; a pushed event carries no capability,
|
||||
which the kernel already guarantees. The device staying while its medium
|
||||
leaves is the one removable-media case the channel-death trigger cannot
|
||||
see; without this event a swapped SD card would be served with the old
|
||||
card's filesystem state.
|
||||
8. **Volume identity, and the mount map as danos's fstab** (settled). The
|
||||
lesson is Linux's own history: fstab keyed on `/dev/sda1` for years and
|
||||
broke whenever a drive changed ports or enumeration order; `UUID=` entries
|
||||
exist because device-path identity failed. danos skips that era: the mount
|
||||
map (`volumes.csv` — configuration, read by the volume manager) keys on
|
||||
**content identity, never port or discovery order**. The prober reads
|
||||
identity off the medium, strongest first:
|
||||
1. GPT partition GUID — 128-bit, unique, stable for the volume's life;
|
||||
2. filesystem UUID (ext-family and most modern formats, in the superblock);
|
||||
3. FAT volume serial + label — 32 bits, weak (dd-cloned sticks share it)
|
||||
but what real sticks carry;
|
||||
4. MBR disk signature + partition index;
|
||||
5. nothing — an anonymous volume: generated mount name, no persistence.
|
||||
|
||||
Consequences, each mechanical once identity keys the map: **moving a drive
|
||||
to a different port changes nothing** — same identity, same mount point,
|
||||
whether USB port, hub depth, SATA port, or a stick that left as USB and
|
||||
returned in a SATA dock; **replug remounts at the same path** (the
|
||||
remount-after-return story is a map lookup); **the boot volume** is the
|
||||
recorded identity of the volume carrying `/system/configuration`, findable
|
||||
on any port; and **duplicate identity is a policy case, not a surprise** —
|
||||
two cloned sticks at once: first keeps the mapped name, second mounts
|
||||
suffixed and is logged loudly, never silently shadowed. Unknown identities
|
||||
mount under a derived name (sanitized label, else generated) at
|
||||
`/volumes/<name>` — the hierarchy's documented home for attached media,
|
||||
which stands: `/system` is what danos IS; attached media is what it isn't.
|
||||
danos never needs to MINT identifiers to detect volumes — detection only
|
||||
reads — until it grows formatting, which brings the entropy question and
|
||||
is deliberately out of scope here.
|
||||
@@ -35,6 +35,14 @@ pub const Device = struct {
|
||||
return self.call(.attach, {}, handle, &reply) != null;
|
||||
}
|
||||
|
||||
/// The reverse of `attach`: the buffer leaves the device's reach. The same
|
||||
/// region capability rides again (the kernel matches the region). Do not name
|
||||
/// the buffer's physical address in `read`/`write` after this.
|
||||
pub fn detach(self: Device, handle: ipc.Handle) bool {
|
||||
var reply: [block_protocol.message_maximum]u8 = undefined;
|
||||
return self.call(.detach, {}, handle, &reply) != null;
|
||||
}
|
||||
|
||||
/// Read `count` blocks starting at `lba` into the DMA buffer at `physical`.
|
||||
pub fn read(self: Device, lba: u64, count: u32, physical: u64) bool {
|
||||
var reply: [block_protocol.message_maximum]u8 = undefined;
|
||||
|
||||
@@ -136,6 +136,15 @@ pub const Device = struct {
|
||||
return self.call(.dma_attach, {}, &.{}, handle, &reply) != null;
|
||||
}
|
||||
|
||||
/// The reverse of `attachDma`: unbind the buffer from the controller's IOMMU
|
||||
/// domain. The same region capability rides again — the kernel matches the
|
||||
/// region, so neither side kept state between the two calls. Do not name the
|
||||
/// buffer's physical address in any transfer after this.
|
||||
pub fn detachDma(self: *Device, handle: ipc.Handle) bool {
|
||||
var reply: [usb_transfer_protocol.message_maximum]u8 = undefined;
|
||||
return self.call(.dma_detach, {}, &.{}, handle, &reply) != null;
|
||||
}
|
||||
|
||||
/// One bulk transfer (IN or OUT per `endpoint_address`'s direction bit) to or
|
||||
/// from the caller's own DMA buffer at `physical`. Returns the bytes moved.
|
||||
pub fn bulk(self: *Device, endpoint_address: u8, physical: u64, length: u32) ?u32 {
|
||||
|
||||
@@ -54,6 +54,12 @@ pub const Protocol = envelope.Define(.{
|
||||
// physical addresses (named in later read/write) are reachable by the
|
||||
// device under an enforcing IOMMU. Call once per buffer before using it.
|
||||
.{ .name = "attach" },
|
||||
// detach(): the reverse — the same region capability rides the cap slot
|
||||
// (the caller still holds its handle; the kernel matches the region) and
|
||||
// the buffer leaves the device's domain. Every grant a live process
|
||||
// makes is revocable by the granter while alive; death remains the
|
||||
// mechanical backstop (storage-architecture.md, the lifecycle rule).
|
||||
.{ .name = "detach" },
|
||||
},
|
||||
});
|
||||
|
||||
|
||||
@@ -156,6 +156,11 @@ pub const Protocol = envelope.Define(.{
|
||||
// target says which caller's device is attaching, so there is nothing
|
||||
// left for a body to carry.
|
||||
.{ .name = "dma_attach" },
|
||||
// dma_detach: the reverse, same shape — the region capability rides the
|
||||
// cap slot again (the kernel matches the region; the provider retains
|
||||
// nothing between the two calls) and the buffer leaves the controller's
|
||||
// domain. Appended, so every existing verb keeps its number.
|
||||
.{ .name = "dma_detach" },
|
||||
},
|
||||
.events = &.{
|
||||
.{ .name = "interrupt_report", .payload = InterruptReport },
|
||||
@@ -188,6 +193,7 @@ test "the verb numbering, and the device token in the header" {
|
||||
try std.testing.expectEqual(@as(u32, 18), @intFromEnum(Operation.interrupt_subscribe));
|
||||
try std.testing.expectEqual(@as(u32, 19), @intFromEnum(Operation.bulk));
|
||||
try std.testing.expectEqual(@as(u32, 20), @intFromEnum(Operation.dma_attach));
|
||||
try std.testing.expectEqual(@as(u32, 21), @intFromEnum(Operation.dma_detach));
|
||||
try std.testing.expectEqual(@as(u32, 16), @intFromEnum(Event.interrupt_report));
|
||||
|
||||
var buffer: [message_maximum]u8 = undefined;
|
||||
|
||||
@@ -204,12 +204,20 @@ fn onAttach(_: void, invocation: Invocation(void), _: Answer(void)) isize {
|
||||
return if (device.attachDma(handle)) 0 else refused;
|
||||
}
|
||||
|
||||
/// The reverse: forward the same region capability so the controller unbinds
|
||||
/// the buffer. As with attach, our copy stays the turn's to close.
|
||||
fn onDetach(_: void, invocation: Invocation(void), _: Answer(void)) isize {
|
||||
const handle = invocation.capability orelse return -envelope.EPROTO;
|
||||
return if (device.detachDma(handle)) 0 else refused;
|
||||
}
|
||||
|
||||
const handlers = Serve.Handlers{
|
||||
.geometry = onGeometry,
|
||||
.read = onRead,
|
||||
.write = onWrite,
|
||||
.flush = onFlush,
|
||||
.attach = onAttach,
|
||||
.detach = onDetach,
|
||||
};
|
||||
|
||||
fn onMessage(message: []const u8, reply: []u8, sender: u32, arrived: *ipc.Arrival) usize {
|
||||
|
||||
@@ -593,6 +593,7 @@ const handlers = Serve.Handlers{
|
||||
.interrupt_subscribe = onInterruptSubscribe,
|
||||
.bulk = onBulk,
|
||||
.dma_attach = onDmaAttach,
|
||||
.dma_detach = onDmaDetach,
|
||||
};
|
||||
|
||||
/// open: the target is the class driver's assigned device id. Resolve it to an
|
||||
@@ -694,6 +695,14 @@ fn onDmaAttach(_: void, invocation: Invocation(void), _: Answer(void)) isize {
|
||||
return if (device.dmaBind(controller_id, handle)) 0 else refused;
|
||||
}
|
||||
|
||||
/// The reverse: the same region capability arrives again and the buffer leaves
|
||||
/// the controller's domain (`dma_unbind` matches the region — nothing was
|
||||
/// retained here between the two calls). The turn closes the arriving copy.
|
||||
fn onDmaDetach(_: void, invocation: Invocation(void), _: Answer(void)) isize {
|
||||
const handle = invocation.capability orelse return -envelope.EPROTO;
|
||||
return if (device.dmaUnbind(controller_id, handle)) 0 else refused;
|
||||
}
|
||||
|
||||
/// A timer tick or an MSI landed: drain the event ring, reconcile ports, and fan out.
|
||||
/// The timer arm re-arms itself (8 ms drain when polling, 250 ms reconcile under MSI);
|
||||
/// the MSI arm clears the interrupter's pending bit FIRST, then drains — so an event
|
||||
|
||||
@@ -196,10 +196,22 @@ fn tryBringUp() void {
|
||||
// enforcing IOMMU. No-op binding otherwise.
|
||||
const bounce = memory.dmaAlloc(engine.max_transfer_sectors * 512, memory.dma_coherent | memory.dma_shareable) orelse return;
|
||||
if (bounce.handle) |handle| {
|
||||
// Attach, detach, and attach again: the round trip exercises BOTH verbs
|
||||
// of the DMA-window lifecycle through the whole chain (fat → storage →
|
||||
// bus → kernel) on every boot, so a broken detach fails every fat case
|
||||
// rather than lying dormant until the first buffer replacement.
|
||||
if (!device.attach(handle)) {
|
||||
_ = logging.write("/system/services/fat: could not attach the DMA bounce buffer\n");
|
||||
return;
|
||||
}
|
||||
if (!device.detach(handle)) {
|
||||
_ = logging.write("/system/services/fat: could not detach the DMA bounce buffer\n");
|
||||
return;
|
||||
}
|
||||
if (!device.attach(handle)) {
|
||||
_ = logging.write("/system/services/fat: could not re-attach the DMA bounce buffer\n");
|
||||
return;
|
||||
}
|
||||
_ = ipc.close(handle); // the binding holds its own reference now
|
||||
}
|
||||
ipc_block = .{ .device = device, .bounce = bounce };
|
||||
|
||||
Reference in New Issue
Block a user