8 Commits
Author SHA1 Message Date
Daniel Samson 9da63e81e2 docs: the boot volume — identity is what makes yanking it survivable
The sharpest instance of the return story, folded into the architecture: the
boot volume is a recorded content identity (the volume carrying
/system/configuration and /system/logs), the system runs on without it (the
ramdisk is the OS; only log persistence pauses), the kernel ring is the
buffer during absence — bounded, so a wrapped ring is a data loss window
that gets MARKED in the file on resume, never spliced silently — and on
return at any port the same identity remounts the same prefixes and the
logger appends into the same boot-stamp tree. The logger requirement is
named: failed flushes retry on the patient cadence, never abandoned after
the first not_found. The dirty-honesty rule stands; the boot volume gets no
exemption.
2026-08-09 16:01:03 +01:00
Daniel Samson ada251150a docs: volume identity, and volumes.csv as danos's fstab
Decision 8 in the rationale, mirrored into the architecture's volume-manager
section: the mount map keys on CONTENT identity, never port or arrival order
— Linux's /dev/sda1-era fstab broke on every port move until UUID= replaced
it, and danos skips that era. The identity ladder the prober reads off the
medium: GPT partition GUID, filesystem UUID, FAT serial+label, MBR
signature+index, anonymous. Consequences are mechanical: port moves change
nothing (USB, hub, SATA bay, or transport swaps), replug remounts at the
same path, the boot volume is a recorded identity findable anywhere, and
cloned duplicates are a loud policy case instead of silent shadowing.
/volumes/<name> stands as the hierarchy's home for attached media; minting
identifiers (formatting, entropy) stays deliberately out of scope.
2026-08-09 15:59:38 +01:00
Daniel Samson 728b436d0f docs: the storage rationale lives with the architecture it justifies
storage-stack-discussion.md was misplaced at the docs root — that level is
for track plans; this is the file-system domain's design record. Moved to
file-system-development/storage-design-rationale.md, renamed to say what it
is, cross-references updated.
2026-08-09 15:49:12 +01:00
Daniel Samson 092817ba2e docs: transport generality and the media-presence event
The NVMe/SATA assessment folded into the storage discussion: what transfers
untouched (the block contract and everything above it — one driver binary
plus one devices.csv row per transport), NVMe as the shorter stack whose
namespaces are the reserved multi-volume case, AHCI as an open shape choice
(leaning per-port processes, the matrix-proven granularity), and the three
pressure points named honestly: multi-volume is reserved-not-implemented,
synchronous call/reply bottlenecks NVMe until the shm-ring data plane, and
media lifecycle is not device lifecycle.

That last one becomes decision 7 and enters the architecture doc: the
removal path has TWO TRIGGERS, ONE LIFECYCLE — channel death (device leaves)
and a planned pushed medium_changed event (medium leaves, device stays: card
readers and trays, USB ones today), translated by the storage driver from
its transport's native signal, presence never content, consumed by the
volume manager into the same kill-retire-remount path. Without it a swapped
card would be served with the previous card's filesystem state.
2026-08-09 15:36:59 +01:00
Daniel Samson a44b397bed protocols: attach gets its reverse — block detach, usb-transfer dma_detach
The kernel was always symmetric (dma_bind 51 / dma_unbind 52); the two
protocols that forward an attachment up the stack were one-way, so a live
client could grant a device reach into its buffer but never revoke it while
alive — exactly the one-way lifecycle the storage architecture's enforcement
section forbids. Death stays the mechanical backstop; detach is the living
process's path.

Both verbs are appended, so every existing number holds. The shape mirrors
attach precisely: the same region capability rides the cap slot again — the
kernel matches the region, so no layer retains anything between the calls
(the bus never kept the handle; now it never needs to).

fat's bring-up does attach -> detach -> attach, exercising both verbs
through the whole chain (fat -> storage -> bus -> kernel) on every boot: a
broken detach fails every fat case instead of lying dormant until the first
buffer replacement. Honest scope: the round trip proves the plumbing; unbind
semantics are the kernel iommu tests' (map/unmap/translationOf); the full
composition (detach then DMA faults) is a future iommu-fault extension.
2026-08-09 15:18:21 +01:00
Daniel Samson fa54ef6915 docs: the volume lifecycle is enforced, not described
The enforcement section of the storage architecture: the three levers the
device lifecycle already proved, mapped onto volumes — a filesystem can only
be GIVEN its volume (no establishment grants, channel at spawn), the volume
manager supervises with teeth (deadline, kill-on-removal, crash-loop cap),
and the shared harness makes every engine inherit the state machine by
construction (engines never see channels). Kernel backstops: ownership-gated
fs_unmount + the lazy dead-endpoint sweep. Checkable via a lifecycle
conformance drill parameterized over filesystems — supporting a filesystem
MEANS passing it.
2026-08-09 15:13:34 +01:00
Daniel Samson fae616fa5a docs: the storage architecture — layers, boundaries, and who does what when media leaves
The settled shape from the storage-stack discussion, written as the
reference: the data path (vfs -> filesystem service -> block -> driver) as
the application/service/protocol/driver model applied twice; the three kinds
of boundary (protocol between processes, library inside them, control-plane
beside them); the volume manager as the policy home (planned — the FAT
service squats on its duties today, marked as such); adding a filesystem as
engine + shared harness + one configuration row; and the per-layer
responsibility table for removable media — one removal path, kill/retire/
respawn, dirty data lost and SAID to be lost. Indexed from docs/README.md.
2026-08-09 15:08:45 +01:00
Daniel Samson 5a8a2a4d7e docs: the storage-stack discussion — block, volumes, filesystems, against the survey 2026-08-09 14:47:35 +01:00
10 changed files with 543 additions and 1 deletions
+5 -1
View File
@@ -51,7 +51,11 @@ rather than restate it. Roughly in the order things happen at runtime:
14. **[vfs-protocol.md](file-system-development/vfs-protocol.md) — the VFS wire protocol.** The language-neutral 14. **[vfs-protocol.md](file-system-development/vfs-protocol.md) — the VFS wire protocol.** The language-neutral
byte-level spec of the file protocol spoken over IPC: request/reply headers, byte-level spec of the file protocol spoken over IPC: request/reply headers,
the operation table, mount routing, and the append-only evolution rules — the the operation table, mount routing, and the append-only evolution rules — the
first IPC protocol documented as public ABI. first IPC protocol documented as public ABI. Its architectural frame is
**[storage-architecture.md](file-system-development/storage-architecture.md) — the storage stack**:
the layers from application to hardware, the three kinds of boundary
(protocol, library, control-plane), how to add a filesystem, and the
per-layer responsibilities when removable media is yanked and returned.
15. **[drivers.md](device-driver-development/drivers.md) — writing a driver.** The payoff: a driver is an 15. **[drivers.md](device-driver-development/drivers.md) — writing a driver.** The payoff: a driver is an
ordinary ring-3 process that claims a device, maps its registers, and **sleeps ordinary ring-3 process that claims a device, maps its registers, and **sleeps
until its hardware interrupts it**. The claim is the capability; `irq_ack` is the until its hardware interrupts it**. The claim is the capability; `irq_ack` is the
@@ -0,0 +1,255 @@
# The storage architecture: layers, boundaries, responsibilities
> **Status:** the layered model below is the settled design
> ([storage-design-rationale.md](storage-design-rationale.md) records how it
> was reached and what the surveyed systems taught). The data path — vfs
> protocol, kernel mount routing, the FAT service, the block protocol,
> usb-storage — is **built**. The volume manager, the driver's range
> mechanism, per-volume filesystem spawning, and the removal lifecycle are
> **planned**; until they land, the FAT service performs volume-manager duties
> itself (marked below). This document is the reference for both states.
## The model
One pattern, applied twice: an application is a client of a service over a
protocol; a service is a client of a driver over a protocol.
```
application
│ vfs protocol (routed by the kernel mount table)
▼
filesystem service ── one process per volume
│ inside: [file-protocol shell] ── Zig API ── [fs engine]
│ block protocol (channel received at spawn)
▼
storage driver ── one process per device (usb-storage per stick)
│ usb-transfer protocol (channel received via hello, by lineage)
▼
bus driver ── one process per controller (usb-xhci-bus)
│ hardware
```
Three kinds of boundary, deliberately different:
- **Protocol boundaries** sit where processes meet (vfs, block, usb-transfer).
Each side is independently restartable; channels are established by
capability handoff, never by registry names (communication.md
"Establishment: two planes"). Steady-state file I/O crosses exactly two —
vfs and block.
- **Library boundaries** sit where concerns meet inside one process. The
filesystem service's engine (on-disk format logic, host-testable, behind
the four-function `BlockDevice` vtable) and shell (IPC serving, open-node
table, mount registration) talk through a Zig API. No protocol between
them: they share fate regardless — a corrupt engine corrupts the answers
either way — so a channel there would add a hop per file operation and buy
nothing. The seam exists for compile-time testability and reuse; the
*process* is the restart unit.
- **Control-plane relationships** sit beside the data path, never on it. Two
supervisors, one per layer: the **device manager** wires and revives the
device layers (bus and storage drivers — devices only); the **volume
manager** *(planned)* wires and revives the volume layer (filesystem
services). Neither touches steady-state I/O.
## Who does what
**Bus driver** (usb-xhci-bus, per controller): hardware only. Enumerates,
reports children to the device manager, serves the transfer contract to its
own children's class drivers, routed by lineage.
**Storage driver** (usb-storage, per device; later nvme, ahci): hardware →
blocks. Speaks its transport upward from its device; serves the block
protocol (geometry, read, write, flush, attach, detach — and, *planned*, the
pushed `medium_changed` presence event, translated from the transport's
native signal). **Content-blind, permanently**:
it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no
GPT, no filesystem magic, ever. *(Planned)* it gains one content-blind
mechanism: **named sub-ranges** — "serve blocks [a, b) as a channel of the
same block contract". It clamps and offsets; it never knows the numbers came
from a partition table. The clamp lives here and nowhere else because a
channel must carry exactly the authority it grants: handing a filesystem the
whole disk plus a polite base offset would let a buggy or compromised
filesystem scribble the neighboring partition — the same authority-overshoot
the device-authority track eliminated for MMIO and DMA.
**Volume manager** *(planned; today the FAT service squats on these duties)*:
the policy home of the volume layer, one service, supervised by init. It
subscribes to the device manager's child events; when a storage provider
appears it consumer-hellos for the block channel, reads the partition table
and the first blocks itself (**it** is the prober), consults its
configuration, defines sub-ranges on the driver, spawns the matching
filesystem service per volume with that volume's channel, supervises it, and
decides mount placement. Its tables are CSV configuration, read by it (the
policy), enforced by nobody else:
- `filesystems.csv` — content signature → filesystem binary. Adding a
filesystem adds a row.
- `volumes.csv` — the mount map, danos's fstab: **volume identity → mount
prefix**, keyed on content identity and never on port, path, or arrival
order (the lesson of Linux's `/dev/sda1`-era fstab, which broke on every
port move until `UUID=` replaced it). Identity is read off the medium by
the prober, strongest first: GPT partition GUID → filesystem UUID → FAT
serial + label → MBR signature + partition index → anonymous (generated
name, no persistence). Same identity, same mount point: a drive moved to
another port — USB to another hub, SATA to another bay, even a stick
returning in a different dock — lands exactly where it was. Duplicate
identity (cloned sticks, together) is policy: first keeps the name, the
second mounts suffixed and is logged loudly. The boot volume is the
recorded identity of the volume carrying `/system/configuration` and
`/system/logs`, findable on any port. Unknown volumes mount under
`/volumes/<derived name>`.
**Filesystem service** (the FAT service today; one process per volume): the
proven unit — block-client + engine + file-protocol provider in one binary. It
receives its block channel at spawn; it never discovers devices. It registers
its own mounts with the kernel; its write cache lives inside the process, so a
write error is observed by the code that owns the volume and surfaces on the
owning channel (the anti-fsyncgate rule — never a system-wide dirty pool).
*(Today, interim:)* fat also acquires its own volume (first mass-storage child
by enumeration order) and hardcodes its mount prefixes; both migrate to the
volume manager.
**Kernel** (mechanism only): the mount table routes paths to backend
endpoints — resolve and redirect, never data. Remount-replace is the restart
story; a dead backend's slot is swept lazily on the next resolution.
*(Planned fix:)* `fs_unmount` gains ownership — only the mounting endpoint's
holder may unmount; today it is ungated, which is safe with one mount owner
and wrong with several.
## Adding a filesystem
1. **Write the engine**: pure Zig, no IPC, behind the `BlockDevice` vtable
(four functions). On-disk layout in `align(1)` extern structs. Host-test it
with `zig build test` fixtures — the FAT engine
([engine.zig](../../system/services/fat/engine.zig),
[on-disk.zig](../../system/services/fat/on-disk.zig)) is the template, and
its ~500 lines of host tests the standard.
2. **Reuse the shell**: the filesystem harness — establishment, the
badge-scoped open-node table, the nine vfs-protocol handlers, mount
registration, the removal path — is shared code, not per-filesystem code.
*(Planned: extracted from fat's 434-line shell into a library before the
second engine is written.)* An engine plus a `main` wiring it into the
harness is a complete filesystem service.
3. **Add the configuration row**: one line in `filesystems.csv` mapping the
on-disk signature to the binary. No other component changes: the volume
manager probes, matches, spawns; the vfs protocol is already
backend-neutral (a second provider serves it in production today — init's
synthetic registry backend).
4. **Prove removal**: extend the removable-media suite (below) for the new
filesystem — surprise-yank during writes must lose only what was
unflushed, said honestly, with the volume consistent enough to remount.
The engine never sees: partitions (it receives a volume-shaped block channel;
base offsets are the driver's clamp), device discovery (the channel arrives at
spawn), mount policy (prefixes are handed to it), or other volumes (one
process, one volume).
## Removable media: responsibilities on removal, per layer
The design rule, learned from what Linux cannot do: there is exactly **one**
surprise-removal path — kill the filesystem process, retire its mounts,
respawn on return. No half-alive states, no `remount-ro`, no mounts that
error forever (Plan 9's dead-server wart).
The path has **two triggers, one lifecycle**: the *device* leaving (the
storage driver dies — channel death, the table below), and the *medium*
leaving while the device stays (an SD card pulled from its reader, an ATAPI
tray opened — including USB card readers today). The second trigger is a
pushed `medium_changed` event on the block protocol *(planned)*: the storage
driver translates its transport's native signal (SCSI UNIT ATTENTION, AHCI
PxSSTS, NVMe namespace-change AER) into presence-changed — never content —
and the volume manager runs the same kill-retire path, then re-probes on
medium return exactly as on device return. Without it, a swapped card would
be served with the previous card's filesystem state.
| Layer | Observes | Must do | Guarantees |
|---|---|---|---|
| Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick |
| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
| Volume manager *(planned)* | the storage provider's channel death / the manager's child events | kill each filesystem service of that device's volumes; retire their kernel mounts; remember the volume identity | one removal path; mounts never dangle; log persistence stops *cleanly* |
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the FAT dirty flag on disk marks the unclean removal; the process never serves from behind a dead channel |
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); *(planned)* ownership-gated `fs_unmount` | resolution under a dead mount is `not_found`, not a hang |
| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
**On return** the same table runs upward in reverse: the bus re-enumerates and
re-reports; the device manager respawns the storage driver (built); the volume
manager re-probes — same content, same volume identity — respawns the
filesystem service, and remounts at the same prefix; applications see the
subtree reappear. A different stick in the same port is a *different volume*
(identity is content, not port) and mounts wherever its identity says —
possibly nowhere but `/volumes/…`.
### The boot volume: identity is what makes yanking it survivable
The boot volume is the sharpest instance of the return story, because the
system's own log persistence rides it (`/system/logs`), and because nothing
else about the running system depends on it at all — every binary and the
boot-time configuration live in the initial ramdisk. Its lifecycle,
end to end:
- **At first mount** the volume manager records the identity of the volume
carrying `/system/configuration` and `/system/logs` — THE boot volume,
from then on a fact about content, not about a port.
- **While it is absent**, the system runs on. The kernel log ring keeps
accumulating — it is the buffer that makes the absence survivable — and
the logger keeps draining it; only *persistence* pauses. File operations
under the retired mounts fail honestly (`not_found`). The ring is bounded,
so a long absence overwrites its oldest entries: that window is the data
loss, and it must be *said* — the logger marks the gap in the file when
persistence resumes, never splicing the stream silently.
- **On return — any port, any hub, even a different transport** — the
prober reads the same identity, the mount map answers with the same
prefixes, the filesystem service is respawned, and `/system/logs` is the
same tree it was: the logger resumes appending into the SAME
`<boot-stamp>` directory, per-binary files continuing where they left
off (plus the gap marker if the ring wrapped).
- **What never comes back** is the write-back window lost at the yank —
the dirty-honesty rule, unchanged; the boot volume gets no exemption.
This is also the requirement that shapes the logger: it treats the log tree
as a volume that comes and goes — failed flushes are retried on the same
patient cadence the fat service already uses for storage that arrives late,
never abandoned after the first `not_found`.
What is lost on a surprise yank is exactly the write-back window of the
filesystem service, no more: the engine owns its cache, so the blast radius of
a yank is one volume's unflushed writes, marked by the on-disk dirty flag. A
filesystem format with better crash honesty (journaling, copy-on-write — the
lesson of QNX's Power-Safe) narrows that window further and slots in as
implementation N+1 through the table above, changing nothing else.
## How the lifecycle is enforced
Responsibility tables are convention; convention is worthless (the bounds
audit's lesson). The volume lifecycle is enforced with the same three levers
the device lifecycle already uses, mapped one-to-one:
1. **A filesystem cannot acquire — it can only be given.** Filesystem
binaries hold NO establishment grants: no `open device-manager`, no
registry name to look up. The only block channel a filesystem process ever
has is the one handed to it at spawn by the volume manager. Serving the
wrong volume, a second volume, or a self-discovered volume is not
forbidden but *impossible* — the same way a driver cannot claim hardware
it was not delegated (device-authority.md).
2. **The volume manager supervises with teeth.** Mount-within-deadline or be
stopped (a filesystem wedged on a corrupt volume is killed and crash-loop
capped; the volume is marked bad, not retried forever). On removal the
manager KILLS the filesystem process and retires its mounts — the polite
observe-`EPEER`-and-exit path is an optimization; the kill is the
guarantee, because a process cannot be trusted to observe its own
obsolescence (the reap argument, proven on the device tree).
3. **The shared harness is how every filesystem inherits the lifecycle by
construction.** The harness — not the engine — owns the state machine:
establishment at spawn, mount registration, the dirty-flag set/clear
bracket, error-out-and-exit on channel death. The engine sits behind the
four-function vtable and never sees a channel; it cannot opt out of the
lifecycle for the same reason it cannot find a device. This is why the
harness is extracted BEFORE the second engine is written.
Beneath all three, two kernel mechanisms: ownership-gated `fs_unmount` (only
the mounting endpoint's holder may unmount), and the lazy dead-endpoint mount
sweep as the backstop nothing can disable. And above them, the check: a
**lifecycle conformance drill** — mount, serve, yank mid-write, verify honest
loss, replug, remount — parameterized over filesystem implementations, so
"danos supports filesystem X" MEANS "X passes the drill through the harness",
exactly as provider conformance means passing the reserved-verb suite.
@@ -0,0 +1,225 @@
# The storage design rationale: why the stack is shaped this way
*2026-08-09. The design record behind
[storage-architecture.md](storage-architecture.md): the survey, the
trade-offs, and the decisions with their reasons — kept so future changes
argue against the evidence rather than rediscovering it. The questions that
drove it, verbatim: should the block protocol be separate from the VFS? how
do channels work with these block devices? how should we wire up different
filesystems? Grounded in how Minix 3, QNX Neutrino, Fuchsia, Redox, Plan 9,
and Linux each answered the same questions, and in exactly where our own
seams sat. The layering rule it serves: drivers are the lowest level
(hardware only); VFS and the filesystems are higher layers; the protocol
layer routes between them.*
## What the survey says, compressed
**Everyone separates block from VFS.** Even Plan 9, which unifies the *protocol*
(a disk is a file tree), keeps the layers: kernel `sd` serves raw ranges,
user-space filesystem servers consume them and serve files. The two poles are
Minix 3 (every layer a process: one VFS server, one filesystem server *per
mounted volume*, driver processes below — driver crashes proven survivable by
fault-injection, tens of thousands of injected faults) and QNX (filesystems as
DLLs inside the disk-driver process — fastest path, but one corrupt volume takes
every volume behind that controller down). Fuchsia is the modern capability-
native reference: a tiny block contract that every layer speaks, control channel
separate from a data fast-path with pre-registered buffers, filesystems as
separate processes launched by a storage-policy component (`fshost`) that probes
content and hands each filesystem its block channel at startup. The filesystem
never discovers devices.
**Partition tables are parsed in user space, below the filesystem, above the
raw device — never in the kernel, never in the filesystem engine.** Fuchsia
tried partitions-as-drivers and formally reversed it (RFC-0257). Plan 9 has the
cleanest split of all: the kernel offers only a *mechanism* — "create a named
sub-range of this disk" — and a user-space prober parses the table and issues
those commands. QNX (`io-blk`) and Minix (driver-side shared library) agree on
placement: the block layer multiplexes one physical block object into several
logical block objects **of the same contract**.
**The failure lessons are unanimous.** Linux's shared page cache across the
filesystem boundary produced fsyncgate (write errors observed by the wrong
process, dirty pages marked clean); the lesson is to keep write caching *inside*
the filesystem process, so an error surfaces on the channel that owns the
volume. Linux's `errors=remount-ro`/lazy-unmount half-alive states exist because
a monolith cannot kill and respawn a filesystem; Plan 9 leaves dead servers'
mounts in the namespace erroring forever. A restartable-parts system should have
exactly one surprise-removal path: kill the filesystem process, retire its
mounts, respawn on return. QNX's deepest lesson: after years of removable-media
pain they concluded detection isn't enough and built a copy-on-write filesystem
(Power-Safe) — *format-level* crash honesty. And QNX's `mcd` (a standalone
policy daemon with declarative insertion/removal rules) plus Fuchsia's `fshost`
agree that media policy is one dedicated component, not code scattered through
filesystems.
## Where we already stand
Closer than expected. The block protocol is 61 lines, five verbs, and its
docblock already reserves `Header.target` for volumes ("targets = volumes,
`enumerate` lists them; nothing else about the protocol changes"). usb-storage
is content-blind — it reads block 0 only as a bring-up self-check and parses
nothing. The vfs protocol is already served by a second provider in production
(init's registry serves it synthetically), so a second filesystem is *not* a
protocol problem. The FAT service is already internally split: a host-testable
engine (1659 + 310 lines) behind a four-function BlockDevice vtable, and a
434-line IPC shell. The kernel mount table already has the right restart
semantics (remount-replace; lazy dead-endpoint sweep on resolve).
The misplacements, all small: the MBR walk lives *inside the FAT engine*
(`engine.zig:196-212`) — every future engine would duplicate it; fat hardcodes
its acquisition policy ("first mass-storage child, mount at /volumes/usb") —
that is policy in a filesystem binary; `fs_unmount` is gated by nothing (any
process may unmount any prefix — latent now, an obvious cross-tenant hole once
mounts multiply); and fat's `mounted` flag is one-way — no path back after its
block channel dies.
## The proposed shape
Three layers, matching the stated rule, every boundary a channel:
**Block (driver layer, mechanism).** usb-storage stays hardware→blocks and
content-blind. It gains one mechanism, Plan 9's: *named sub-ranges*. A
`define_range`-style op creates a logical block object (a partition) clamped
and offset onto the disk; sub-ranges are `target` ids on the same endpoint, or
separate endpoints handed out per range — either way they speak the **same
block contract** (Fuchsia's closed-under-layering rule), so a filesystem cannot
tell whole-disk from partition. The driver never parses a table; it clamps
ranges it is told about. Offsets translate at definition time, so the data path
stays one hop (Fuchsia's session-mapping trick, for free).
**Volumes (service layer, policy).** One new service — the *volume manager*
(fshost/mcd-shaped, init-supervised; the device manager stays devices-only). It
subscribes to the device manager's `child_added`/`child_removed`, consumer-
hellos for each new mass-storage provider's block channel, reads the partition
table and the first blocks itself (the prober is policy), consults
configuration — `filesystems.csv`: content signature → filesystem binary;
`mounts.csv` or similar: volume identity → mount prefix — defines sub-ranges on
the driver, spawns **one filesystem process per volume**, hands it its block
channel at startup, and supervises it with the reap-and-rebuild idiom the
device manager proved: provider dies or medium leaves → kill the filesystem
process, retire its mounts; medium returns → re-probe, respawn, remount. The
boot volume is chosen by *content* (which volume carries /system/configuration
and /system/logs), closing the two-sticks question honestly.
**Filesystems (per volume, one process).** The proven unit everywhere from
Plan 9's `dossrv` to Minix to Fuchsia: block-client + engine + file-protocol
provider in one binary, one process per volume (9front practice; per-volume
fault isolation is what our supervision makes cheap). fat's shell becomes a
shared *filesystem harness* library before a second engine is written; the MBR
walk moves out of the engine into the volume manager; write caching stays
inside the process (the anti-fsyncgate rule). Each mounts its prefixes into the
kernel mount table itself, exactly as today.
**Kernel: two small changes only.** `fs_unmount` gains ownership (only the
mounting endpoint's holder may unmount — possession-is-capability, consistent
with everything else), and the 8-slot mount table gets a declared bound or
growth once volumes multiply. The mount table stays the router; per-process
namespaces (Plan 9's extra) remain separable future work.
## Transport generality: NVMe and SATA against this design
The layering was chosen transport-agnostic on purpose (the 9front lesson:
filesystem services cannot tell a USB stick from an ATA disk); NVMe and AHCI
are the test of that claim, and they fit — with three named pressure points.
**What transfers untouched:** the block protocol (nothing USB in it), the
volume manager, partitions/ranges, filesystems, mounts, and the whole
enforcement stack — delegation, IOMMU confinement (these are first-party PCI
DMA masters, so confinement applies even more directly than USB),
reap-and-rebuild, the hot-plug lifecycle. Integration cost per transport is
what the architecture promises: one driver binary plus one `devices.csv` row.
**NVMe** shortens the stack — no bus/class split; the driver IS the
controller driver, one hop fewer than USB. Its structural novelty,
**namespaces** (hardware-native multiple volumes behind one controller), is
exactly the case the block protocol reserved on day one and decision 4 below
settles as endpoint-per-volume. **AHCI** is a shape choice, not a problem: an
HBA fronts up to 32 ports plus port multipliers — structurally a bus — so
either mirror USB (ahci-bus + a per-port disk driver: maximum restart
granularity, the proven shape) or mirror NVMe (one driver per controller, one
block channel per port). Leaning per-port processes for consistency with the
matrix-proven shape; genuinely open.
**The pressure points, honestly:**
1. **Multi-volume providers are reserved, not implemented.** The volume
manager flow assumes one provider, one volume; NVMe namespaces make
endpoint-per-volume real work with hardware demanding it.
2. **The current transport will bottleneck NVMe.** Synchronous call/reply,
one operation in flight, one bounce buffer — fine for a USB2 stick,
forfeits an NVMe drive's queue depth and per-queue MSI-X. Correctness
needs nothing; performance is gated on the shm-ring data plane
communication.md already names (Fuchsia's FIFO+VMO is the precedent).
NVMe is not worth building before that milestone.
3. **Media lifecycle ≠ device lifecycle.** ATAPI trays and SD card readers
(USB ones too, today) keep the DEVICE present while the MEDIUM comes and
goes — a removal trigger our channel-death lifecycle does not carry. The
fix is decision 7 below. (Related small item: fat's bounce sizing
hardcodes 512-byte sectors; 4Kn drives need the geometry honored
throughout.)
## The decisions on the table
1. **Partitions as driver-side sub-ranges** (mechanism in usb-storage, parsing
in the volume manager) — versus a separate partition *process* re-serving
block (Fuchsia's storage-host). Sub-ranges are ~20 lines of clamp in the
driver and keep the data path one hop; a separate process is purer layering
at the cost of a hop or session plumbing. Recommendation: sub-ranges.
2. **The volume manager as a new service** owning probe, spawn, supervision,
and mount policy — with `filesystems.csv` and the mount map as
configuration. Recommendation: yes; it is the missing policy home that fat
is currently squatting in.
3. **One filesystem process per volume** (fat's binary becomes "the FAT
implementation", spawned per FAT volume). Recommendation: yes — it extends
recompile-and-restart-live to filesystems and isolates corrupt media.
4. **Sub-range addressing**: `target` ids on the storage endpoint versus one
endpoint per volume handed out by the driver. Endpoint-per-volume matches
the establishment-plane machinery (a channel per party, caps at
establishment) and keeps per-client badge scoping simple. Recommendation:
endpoint per volume.
5. **`fs_unmount` ownership** — a defect fix more than a decision.
6. **Later, kept open**: the shm-ring data plane (communication.md already
names it as the 256-byte ceiling's unlock — Fuchsia's FIFO+VMO is the
precedent); format-level crash honesty (a Power-Safe-style journaling or COW
filesystem) once danos outgrows FAT; per-process namespaces.
7. **The media-presence event** (settled in principle; lands with the volume
manager): the block protocol gains a pushed event — `medium_changed`, with
present/absent and a change counter — produced by the storage driver from
its transport's native signal (SCSI UNIT ATTENTION / TEST UNIT READY for
USB and ATAPI, PxSSTS for AHCI, namespace-change AER for NVMe) and
consumed by the volume manager, which runs the SAME kill-retire-remount
path it runs on channel death — one lifecycle, two triggers. The driver
reports presence, never content; a pushed event carries no capability,
which the kernel already guarantees. The device staying while its medium
leaves is the one removable-media case the channel-death trigger cannot
see; without this event a swapped SD card would be served with the old
card's filesystem state.
8. **Volume identity, and the mount map as danos's fstab** (settled). The
lesson is Linux's own history: fstab keyed on `/dev/sda1` for years and
broke whenever a drive changed ports or enumeration order; `UUID=` entries
exist because device-path identity failed. danos skips that era: the mount
map (`volumes.csv` — configuration, read by the volume manager) keys on
**content identity, never port or discovery order**. The prober reads
identity off the medium, strongest first:
1. GPT partition GUID — 128-bit, unique, stable for the volume's life;
2. filesystem UUID (ext-family and most modern formats, in the superblock);
3. FAT volume serial + label — 32 bits, weak (dd-cloned sticks share it)
but what real sticks carry;
4. MBR disk signature + partition index;
5. nothing — an anonymous volume: generated mount name, no persistence.
Consequences, each mechanical once identity keys the map: **moving a drive
to a different port changes nothing** — same identity, same mount point,
whether USB port, hub depth, SATA port, or a stick that left as USB and
returned in a SATA dock; **replug remounts at the same path** (the
remount-after-return story is a map lookup); **the boot volume** is the
recorded identity of the volume carrying `/system/configuration`, findable
on any port; and **duplicate identity is a policy case, not a surprise** —
two cloned sticks at once: first keeps the mapped name, second mounts
suffixed and is logged loudly, never silently shadowed. Unknown identities
mount under a derived name (sanitized label, else generated) at
`/volumes/<name>` — the hierarchy's documented home for attached media,
which stands: `/system` is what danos IS; attached media is what it isn't.
danos never needs to MINT identifiers to detect volumes — detection only
reads — until it grows formatting, which brings the entropy question and
is deliberately out of scope here.
+8
View File
@@ -35,6 +35,14 @@ pub const Device = struct {
return self.call(.attach, {}, handle, &reply) != null; return self.call(.attach, {}, handle, &reply) != null;
} }
/// The reverse of `attach`: the buffer leaves the device's reach. The same
/// region capability rides again (the kernel matches the region). Do not name
/// the buffer's physical address in `read`/`write` after this.
pub fn detach(self: Device, handle: ipc.Handle) bool {
var reply: [block_protocol.message_maximum]u8 = undefined;
return self.call(.detach, {}, handle, &reply) != null;
}
/// Read `count` blocks starting at `lba` into the DMA buffer at `physical`. /// Read `count` blocks starting at `lba` into the DMA buffer at `physical`.
pub fn read(self: Device, lba: u64, count: u32, physical: u64) bool { pub fn read(self: Device, lba: u64, count: u32, physical: u64) bool {
var reply: [block_protocol.message_maximum]u8 = undefined; var reply: [block_protocol.message_maximum]u8 = undefined;
+9
View File
@@ -136,6 +136,15 @@ pub const Device = struct {
return self.call(.dma_attach, {}, &.{}, handle, &reply) != null; return self.call(.dma_attach, {}, &.{}, handle, &reply) != null;
} }
/// The reverse of `attachDma`: unbind the buffer from the controller's IOMMU
/// domain. The same region capability rides again — the kernel matches the
/// region, so neither side kept state between the two calls. Do not name the
/// buffer's physical address in any transfer after this.
pub fn detachDma(self: *Device, handle: ipc.Handle) bool {
var reply: [usb_transfer_protocol.message_maximum]u8 = undefined;
return self.call(.dma_detach, {}, &.{}, handle, &reply) != null;
}
/// One bulk transfer (IN or OUT per `endpoint_address`'s direction bit) to or /// One bulk transfer (IN or OUT per `endpoint_address`'s direction bit) to or
/// from the caller's own DMA buffer at `physical`. Returns the bytes moved. /// from the caller's own DMA buffer at `physical`. Returns the bytes moved.
pub fn bulk(self: *Device, endpoint_address: u8, physical: u64, length: u32) ?u32 { pub fn bulk(self: *Device, endpoint_address: u8, physical: u64, length: u32) ?u32 {
@@ -54,6 +54,12 @@ pub const Protocol = envelope.Define(.{
// physical addresses (named in later read/write) are reachable by the // physical addresses (named in later read/write) are reachable by the
// device under an enforcing IOMMU. Call once per buffer before using it. // device under an enforcing IOMMU. Call once per buffer before using it.
.{ .name = "attach" }, .{ .name = "attach" },
// detach(): the reverse — the same region capability rides the cap slot
// (the caller still holds its handle; the kernel matches the region) and
// the buffer leaves the device's domain. Every grant a live process
// makes is revocable by the granter while alive; death remains the
// mechanical backstop (storage-architecture.md, the lifecycle rule).
.{ .name = "detach" },
}, },
}); });
@@ -156,6 +156,11 @@ pub const Protocol = envelope.Define(.{
// target says which caller's device is attaching, so there is nothing // target says which caller's device is attaching, so there is nothing
// left for a body to carry. // left for a body to carry.
.{ .name = "dma_attach" }, .{ .name = "dma_attach" },
// dma_detach: the reverse, same shape — the region capability rides the
// cap slot again (the kernel matches the region; the provider retains
// nothing between the two calls) and the buffer leaves the controller's
// domain. Appended, so every existing verb keeps its number.
.{ .name = "dma_detach" },
}, },
.events = &.{ .events = &.{
.{ .name = "interrupt_report", .payload = InterruptReport }, .{ .name = "interrupt_report", .payload = InterruptReport },
@@ -188,6 +193,7 @@ test "the verb numbering, and the device token in the header" {
try std.testing.expectEqual(@as(u32, 18), @intFromEnum(Operation.interrupt_subscribe)); try std.testing.expectEqual(@as(u32, 18), @intFromEnum(Operation.interrupt_subscribe));
try std.testing.expectEqual(@as(u32, 19), @intFromEnum(Operation.bulk)); try std.testing.expectEqual(@as(u32, 19), @intFromEnum(Operation.bulk));
try std.testing.expectEqual(@as(u32, 20), @intFromEnum(Operation.dma_attach)); try std.testing.expectEqual(@as(u32, 20), @intFromEnum(Operation.dma_attach));
try std.testing.expectEqual(@as(u32, 21), @intFromEnum(Operation.dma_detach));
try std.testing.expectEqual(@as(u32, 16), @intFromEnum(Event.interrupt_report)); try std.testing.expectEqual(@as(u32, 16), @intFromEnum(Event.interrupt_report));
var buffer: [message_maximum]u8 = undefined; var buffer: [message_maximum]u8 = undefined;
@@ -204,12 +204,20 @@ fn onAttach(_: void, invocation: Invocation(void), _: Answer(void)) isize {
return if (device.attachDma(handle)) 0 else refused; return if (device.attachDma(handle)) 0 else refused;
} }
/// The reverse: forward the same region capability so the controller unbinds
/// the buffer. As with attach, our copy stays the turn's to close.
fn onDetach(_: void, invocation: Invocation(void), _: Answer(void)) isize {
const handle = invocation.capability orelse return -envelope.EPROTO;
return if (device.detachDma(handle)) 0 else refused;
}
const handlers = Serve.Handlers{ const handlers = Serve.Handlers{
.geometry = onGeometry, .geometry = onGeometry,
.read = onRead, .read = onRead,
.write = onWrite, .write = onWrite,
.flush = onFlush, .flush = onFlush,
.attach = onAttach, .attach = onAttach,
.detach = onDetach,
}; };
fn onMessage(message: []const u8, reply: []u8, sender: u32, arrived: *ipc.Arrival) usize { fn onMessage(message: []const u8, reply: []u8, sender: u32, arrived: *ipc.Arrival) usize {
@@ -593,6 +593,7 @@ const handlers = Serve.Handlers{
.interrupt_subscribe = onInterruptSubscribe, .interrupt_subscribe = onInterruptSubscribe,
.bulk = onBulk, .bulk = onBulk,
.dma_attach = onDmaAttach, .dma_attach = onDmaAttach,
.dma_detach = onDmaDetach,
}; };
/// open: the target is the class driver's assigned device id. Resolve it to an /// open: the target is the class driver's assigned device id. Resolve it to an
@@ -694,6 +695,14 @@ fn onDmaAttach(_: void, invocation: Invocation(void), _: Answer(void)) isize {
return if (device.dmaBind(controller_id, handle)) 0 else refused; return if (device.dmaBind(controller_id, handle)) 0 else refused;
} }
/// The reverse: the same region capability arrives again and the buffer leaves
/// the controller's domain (`dma_unbind` matches the region — nothing was
/// retained here between the two calls). The turn closes the arriving copy.
fn onDmaDetach(_: void, invocation: Invocation(void), _: Answer(void)) isize {
const handle = invocation.capability orelse return -envelope.EPROTO;
return if (device.dmaUnbind(controller_id, handle)) 0 else refused;
}
/// A timer tick or an MSI landed: drain the event ring, reconcile ports, and fan out. /// A timer tick or an MSI landed: drain the event ring, reconcile ports, and fan out.
/// The timer arm re-arms itself (8 ms drain when polling, 250 ms reconcile under MSI); /// The timer arm re-arms itself (8 ms drain when polling, 250 ms reconcile under MSI);
/// the MSI arm clears the interrupter's pending bit FIRST, then drains — so an event /// the MSI arm clears the interrupter's pending bit FIRST, then drains — so an event
+12
View File
@@ -196,10 +196,22 @@ fn tryBringUp() void {
// enforcing IOMMU. No-op binding otherwise. // enforcing IOMMU. No-op binding otherwise.
const bounce = memory.dmaAlloc(engine.max_transfer_sectors * 512, memory.dma_coherent | memory.dma_shareable) orelse return; const bounce = memory.dmaAlloc(engine.max_transfer_sectors * 512, memory.dma_coherent | memory.dma_shareable) orelse return;
if (bounce.handle) |handle| { if (bounce.handle) |handle| {
// Attach, detach, and attach again: the round trip exercises BOTH verbs
// of the DMA-window lifecycle through the whole chain (fat → storage →
// bus → kernel) on every boot, so a broken detach fails every fat case
// rather than lying dormant until the first buffer replacement.
if (!device.attach(handle)) { if (!device.attach(handle)) {
_ = logging.write("/system/services/fat: could not attach the DMA bounce buffer\n"); _ = logging.write("/system/services/fat: could not attach the DMA bounce buffer\n");
return; return;
} }
if (!device.detach(handle)) {
_ = logging.write("/system/services/fat: could not detach the DMA bounce buffer\n");
return;
}
if (!device.attach(handle)) {
_ = logging.write("/system/services/fat: could not re-attach the DMA bounce buffer\n");
return;
}
_ = ipc.close(handle); // the binding holds its own reference now _ = ipc.close(handle); // the binding holds its own reference now
} }
ipc_block = .{ .device = device, .bounce = bounce }; ipc_block = .{ .device = device, .bounce = bounce };