Compare commits
8
Commits
d4f8dc51b9
...
9da63e81e2
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
9da63e81e2 | ||
|
|
ada251150a | ||
|
|
728b436d0f | ||
|
|
092817ba2e | ||
|
|
a44b397bed | ||
|
|
fa54ef6915 | ||
|
|
fae616fa5a | ||
|
|
5a8a2a4d7e |
+5
-1
@@ -51,7 +51,11 @@ rather than restate it. Roughly in the order things happen at runtime:
|
|||||||
14. **[vfs-protocol.md](file-system-development/vfs-protocol.md) — the VFS wire protocol.** The language-neutral
|
14. **[vfs-protocol.md](file-system-development/vfs-protocol.md) — the VFS wire protocol.** The language-neutral
|
||||||
byte-level spec of the file protocol spoken over IPC: request/reply headers,
|
byte-level spec of the file protocol spoken over IPC: request/reply headers,
|
||||||
the operation table, mount routing, and the append-only evolution rules — the
|
the operation table, mount routing, and the append-only evolution rules — the
|
||||||
first IPC protocol documented as public ABI.
|
first IPC protocol documented as public ABI. Its architectural frame is
|
||||||
|
**[storage-architecture.md](file-system-development/storage-architecture.md) — the storage stack**:
|
||||||
|
the layers from application to hardware, the three kinds of boundary
|
||||||
|
(protocol, library, control-plane), how to add a filesystem, and the
|
||||||
|
per-layer responsibilities when removable media is yanked and returned.
|
||||||
15. **[drivers.md](device-driver-development/drivers.md) — writing a driver.** The payoff: a driver is an
|
15. **[drivers.md](device-driver-development/drivers.md) — writing a driver.** The payoff: a driver is an
|
||||||
ordinary ring-3 process that claims a device, maps its registers, and **sleeps
|
ordinary ring-3 process that claims a device, maps its registers, and **sleeps
|
||||||
until its hardware interrupts it**. The claim is the capability; `irq_ack` is the
|
until its hardware interrupts it**. The claim is the capability; `irq_ack` is the
|
||||||
|
|||||||
@@ -0,0 +1,255 @@
|
|||||||
|
# The storage architecture: layers, boundaries, responsibilities
|
||||||
|
|
||||||
|
> **Status:** the layered model below is the settled design
|
||||||
|
> ([storage-design-rationale.md](storage-design-rationale.md) records how it
|
||||||
|
> was reached and what the surveyed systems taught). The data path — vfs
|
||||||
|
> protocol, kernel mount routing, the FAT service, the block protocol,
|
||||||
|
> usb-storage — is **built**. The volume manager, the driver's range
|
||||||
|
> mechanism, per-volume filesystem spawning, and the removal lifecycle are
|
||||||
|
> **planned**; until they land, the FAT service performs volume-manager duties
|
||||||
|
> itself (marked below). This document is the reference for both states.
|
||||||
|
|
||||||
|
## The model
|
||||||
|
|
||||||
|
One pattern, applied twice: an application is a client of a service over a
|
||||||
|
protocol; a service is a client of a driver over a protocol.
|
||||||
|
|
||||||
|
```
|
||||||
|
application
|
||||||
|
│ vfs protocol (routed by the kernel mount table)
|
||||||
|
▼
|
||||||
|
filesystem service ── one process per volume
|
||||||
|
│ inside: [file-protocol shell] ── Zig API ── [fs engine]
|
||||||
|
│ block protocol (channel received at spawn)
|
||||||
|
▼
|
||||||
|
storage driver ── one process per device (usb-storage per stick)
|
||||||
|
│ usb-transfer protocol (channel received via hello, by lineage)
|
||||||
|
▼
|
||||||
|
bus driver ── one process per controller (usb-xhci-bus)
|
||||||
|
│ hardware
|
||||||
|
```
|
||||||
|
|
||||||
|
Three kinds of boundary, deliberately different:
|
||||||
|
|
||||||
|
- **Protocol boundaries** sit where processes meet (vfs, block, usb-transfer).
|
||||||
|
Each side is independently restartable; channels are established by
|
||||||
|
capability handoff, never by registry names (communication.md
|
||||||
|
"Establishment: two planes"). Steady-state file I/O crosses exactly two —
|
||||||
|
vfs and block.
|
||||||
|
- **Library boundaries** sit where concerns meet inside one process. The
|
||||||
|
filesystem service's engine (on-disk format logic, host-testable, behind
|
||||||
|
the four-function `BlockDevice` vtable) and shell (IPC serving, open-node
|
||||||
|
table, mount registration) talk through a Zig API. No protocol between
|
||||||
|
them: they share fate regardless — a corrupt engine corrupts the answers
|
||||||
|
either way — so a channel there would add a hop per file operation and buy
|
||||||
|
nothing. The seam exists for compile-time testability and reuse; the
|
||||||
|
*process* is the restart unit.
|
||||||
|
- **Control-plane relationships** sit beside the data path, never on it. Two
|
||||||
|
supervisors, one per layer: the **device manager** wires and revives the
|
||||||
|
device layers (bus and storage drivers — devices only); the **volume
|
||||||
|
manager** *(planned)* wires and revives the volume layer (filesystem
|
||||||
|
services). Neither touches steady-state I/O.
|
||||||
|
|
||||||
|
## Who does what
|
||||||
|
|
||||||
|
**Bus driver** (usb-xhci-bus, per controller): hardware only. Enumerates,
|
||||||
|
reports children to the device manager, serves the transfer contract to its
|
||||||
|
own children's class drivers, routed by lineage.
|
||||||
|
|
||||||
|
**Storage driver** (usb-storage, per device; later nvme, ahci): hardware →
|
||||||
|
blocks. Speaks its transport upward from its device; serves the block
|
||||||
|
protocol (geometry, read, write, flush, attach, detach — and, *planned*, the
|
||||||
|
pushed `medium_changed` presence event, translated from the transport's
|
||||||
|
native signal). **Content-blind, permanently**:
|
||||||
|
it reads block 0 only as a bring-up self-check and parses nothing — no MBR, no
|
||||||
|
GPT, no filesystem magic, ever. *(Planned)* it gains one content-blind
|
||||||
|
mechanism: **named sub-ranges** — "serve blocks [a, b) as a channel of the
|
||||||
|
same block contract". It clamps and offsets; it never knows the numbers came
|
||||||
|
from a partition table. The clamp lives here and nowhere else because a
|
||||||
|
channel must carry exactly the authority it grants: handing a filesystem the
|
||||||
|
whole disk plus a polite base offset would let a buggy or compromised
|
||||||
|
filesystem scribble the neighboring partition — the same authority-overshoot
|
||||||
|
the device-authority track eliminated for MMIO and DMA.
|
||||||
|
|
||||||
|
**Volume manager** *(planned; today the FAT service squats on these duties)*:
|
||||||
|
the policy home of the volume layer, one service, supervised by init. It
|
||||||
|
subscribes to the device manager's child events; when a storage provider
|
||||||
|
appears it consumer-hellos for the block channel, reads the partition table
|
||||||
|
and the first blocks itself (**it** is the prober), consults its
|
||||||
|
configuration, defines sub-ranges on the driver, spawns the matching
|
||||||
|
filesystem service per volume with that volume's channel, supervises it, and
|
||||||
|
decides mount placement. Its tables are CSV configuration, read by it (the
|
||||||
|
policy), enforced by nobody else:
|
||||||
|
|
||||||
|
- `filesystems.csv` — content signature → filesystem binary. Adding a
|
||||||
|
filesystem adds a row.
|
||||||
|
- `volumes.csv` — the mount map, danos's fstab: **volume identity → mount
|
||||||
|
prefix**, keyed on content identity and never on port, path, or arrival
|
||||||
|
order (the lesson of Linux's `/dev/sda1`-era fstab, which broke on every
|
||||||
|
port move until `UUID=` replaced it). Identity is read off the medium by
|
||||||
|
the prober, strongest first: GPT partition GUID → filesystem UUID → FAT
|
||||||
|
serial + label → MBR signature + partition index → anonymous (generated
|
||||||
|
name, no persistence). Same identity, same mount point: a drive moved to
|
||||||
|
another port — USB to another hub, SATA to another bay, even a stick
|
||||||
|
returning in a different dock — lands exactly where it was. Duplicate
|
||||||
|
identity (cloned sticks, together) is policy: first keeps the name, the
|
||||||
|
second mounts suffixed and is logged loudly. The boot volume is the
|
||||||
|
recorded identity of the volume carrying `/system/configuration` and
|
||||||
|
`/system/logs`, findable on any port. Unknown volumes mount under
|
||||||
|
`/volumes/<derived name>`.
|
||||||
|
|
||||||
|
**Filesystem service** (the FAT service today; one process per volume): the
|
||||||
|
proven unit — block-client + engine + file-protocol provider in one binary. It
|
||||||
|
receives its block channel at spawn; it never discovers devices. It registers
|
||||||
|
its own mounts with the kernel; its write cache lives inside the process, so a
|
||||||
|
write error is observed by the code that owns the volume and surfaces on the
|
||||||
|
owning channel (the anti-fsyncgate rule — never a system-wide dirty pool).
|
||||||
|
*(Today, interim:)* fat also acquires its own volume (first mass-storage child
|
||||||
|
by enumeration order) and hardcodes its mount prefixes; both migrate to the
|
||||||
|
volume manager.
|
||||||
|
|
||||||
|
**Kernel** (mechanism only): the mount table routes paths to backend
|
||||||
|
endpoints — resolve and redirect, never data. Remount-replace is the restart
|
||||||
|
story; a dead backend's slot is swept lazily on the next resolution.
|
||||||
|
*(Planned fix:)* `fs_unmount` gains ownership — only the mounting endpoint's
|
||||||
|
holder may unmount; today it is ungated, which is safe with one mount owner
|
||||||
|
and wrong with several.
|
||||||
|
|
||||||
|
## Adding a filesystem
|
||||||
|
|
||||||
|
1. **Write the engine**: pure Zig, no IPC, behind the `BlockDevice` vtable
|
||||||
|
(four functions). On-disk layout in `align(1)` extern structs. Host-test it
|
||||||
|
with `zig build test` fixtures — the FAT engine
|
||||||
|
([engine.zig](../../system/services/fat/engine.zig),
|
||||||
|
[on-disk.zig](../../system/services/fat/on-disk.zig)) is the template, and
|
||||||
|
its ~500 lines of host tests the standard.
|
||||||
|
2. **Reuse the shell**: the filesystem harness — establishment, the
|
||||||
|
badge-scoped open-node table, the nine vfs-protocol handlers, mount
|
||||||
|
registration, the removal path — is shared code, not per-filesystem code.
|
||||||
|
*(Planned: extracted from fat's 434-line shell into a library before the
|
||||||
|
second engine is written.)* An engine plus a `main` wiring it into the
|
||||||
|
harness is a complete filesystem service.
|
||||||
|
3. **Add the configuration row**: one line in `filesystems.csv` mapping the
|
||||||
|
on-disk signature to the binary. No other component changes: the volume
|
||||||
|
manager probes, matches, spawns; the vfs protocol is already
|
||||||
|
backend-neutral (a second provider serves it in production today — init's
|
||||||
|
synthetic registry backend).
|
||||||
|
4. **Prove removal**: extend the removable-media suite (below) for the new
|
||||||
|
filesystem — surprise-yank during writes must lose only what was
|
||||||
|
unflushed, said honestly, with the volume consistent enough to remount.
|
||||||
|
|
||||||
|
The engine never sees: partitions (it receives a volume-shaped block channel;
|
||||||
|
base offsets are the driver's clamp), device discovery (the channel arrives at
|
||||||
|
spawn), mount policy (prefixes are handed to it), or other volumes (one
|
||||||
|
process, one volume).
|
||||||
|
|
||||||
|
## Removable media: responsibilities on removal, per layer
|
||||||
|
|
||||||
|
The design rule, learned from what Linux cannot do: there is exactly **one**
|
||||||
|
surprise-removal path — kill the filesystem process, retire its mounts,
|
||||||
|
respawn on return. No half-alive states, no `remount-ro`, no mounts that
|
||||||
|
error forever (Plan 9's dead-server wart).
|
||||||
|
|
||||||
|
The path has **two triggers, one lifecycle**: the *device* leaving (the
|
||||||
|
storage driver dies — channel death, the table below), and the *medium*
|
||||||
|
leaving while the device stays (an SD card pulled from its reader, an ATAPI
|
||||||
|
tray opened — including USB card readers today). The second trigger is a
|
||||||
|
pushed `medium_changed` event on the block protocol *(planned)*: the storage
|
||||||
|
driver translates its transport's native signal (SCSI UNIT ATTENTION, AHCI
|
||||||
|
PxSSTS, NVMe namespace-change AER) into presence-changed — never content —
|
||||||
|
and the volume manager runs the same kill-retire path, then re-probes on
|
||||||
|
medium return exactly as on device return. Without it, a swapped card would
|
||||||
|
be served with the previous card's filesystem state.
|
||||||
|
|
||||||
|
| Layer | Observes | Must do | Guarantees |
|
||||||
|
|---|---|---|---|
|
||||||
|
| Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick |
|
||||||
|
| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
|
||||||
|
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
|
||||||
|
| Volume manager *(planned)* | the storage provider's channel death / the manager's child events | kill each filesystem service of that device's volumes; retire their kernel mounts; remember the volume identity | one removal path; mounts never dangle; log persistence stops *cleanly* |
|
||||||
|
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the FAT dirty flag on disk marks the unclean removal; the process never serves from behind a dead channel |
|
||||||
|
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); *(planned)* ownership-gated `fs_unmount` | resolution under a dead mount is `not_found`, not a hang |
|
||||||
|
| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
|
||||||
|
|
||||||
|
**On return** the same table runs upward in reverse: the bus re-enumerates and
|
||||||
|
re-reports; the device manager respawns the storage driver (built); the volume
|
||||||
|
manager re-probes — same content, same volume identity — respawns the
|
||||||
|
filesystem service, and remounts at the same prefix; applications see the
|
||||||
|
subtree reappear. A different stick in the same port is a *different volume*
|
||||||
|
(identity is content, not port) and mounts wherever its identity says —
|
||||||
|
possibly nowhere but `/volumes/…`.
|
||||||
|
|
||||||
|
### The boot volume: identity is what makes yanking it survivable
|
||||||
|
|
||||||
|
The boot volume is the sharpest instance of the return story, because the
|
||||||
|
system's own log persistence rides it (`/system/logs`), and because nothing
|
||||||
|
else about the running system depends on it at all — every binary and the
|
||||||
|
boot-time configuration live in the initial ramdisk. Its lifecycle,
|
||||||
|
end to end:
|
||||||
|
|
||||||
|
- **At first mount** the volume manager records the identity of the volume
|
||||||
|
carrying `/system/configuration` and `/system/logs` — THE boot volume,
|
||||||
|
from then on a fact about content, not about a port.
|
||||||
|
- **While it is absent**, the system runs on. The kernel log ring keeps
|
||||||
|
accumulating — it is the buffer that makes the absence survivable — and
|
||||||
|
the logger keeps draining it; only *persistence* pauses. File operations
|
||||||
|
under the retired mounts fail honestly (`not_found`). The ring is bounded,
|
||||||
|
so a long absence overwrites its oldest entries: that window is the data
|
||||||
|
loss, and it must be *said* — the logger marks the gap in the file when
|
||||||
|
persistence resumes, never splicing the stream silently.
|
||||||
|
- **On return — any port, any hub, even a different transport** — the
|
||||||
|
prober reads the same identity, the mount map answers with the same
|
||||||
|
prefixes, the filesystem service is respawned, and `/system/logs` is the
|
||||||
|
same tree it was: the logger resumes appending into the SAME
|
||||||
|
`<boot-stamp>` directory, per-binary files continuing where they left
|
||||||
|
off (plus the gap marker if the ring wrapped).
|
||||||
|
- **What never comes back** is the write-back window lost at the yank —
|
||||||
|
the dirty-honesty rule, unchanged; the boot volume gets no exemption.
|
||||||
|
|
||||||
|
This is also the requirement that shapes the logger: it treats the log tree
|
||||||
|
as a volume that comes and goes — failed flushes are retried on the same
|
||||||
|
patient cadence the fat service already uses for storage that arrives late,
|
||||||
|
never abandoned after the first `not_found`.
|
||||||
|
|
||||||
|
What is lost on a surprise yank is exactly the write-back window of the
|
||||||
|
filesystem service, no more: the engine owns its cache, so the blast radius of
|
||||||
|
a yank is one volume's unflushed writes, marked by the on-disk dirty flag. A
|
||||||
|
filesystem format with better crash honesty (journaling, copy-on-write — the
|
||||||
|
lesson of QNX's Power-Safe) narrows that window further and slots in as
|
||||||
|
implementation N+1 through the table above, changing nothing else.
|
||||||
|
|
||||||
|
## How the lifecycle is enforced
|
||||||
|
|
||||||
|
Responsibility tables are convention; convention is worthless (the bounds
|
||||||
|
audit's lesson). The volume lifecycle is enforced with the same three levers
|
||||||
|
the device lifecycle already uses, mapped one-to-one:
|
||||||
|
|
||||||
|
1. **A filesystem cannot acquire — it can only be given.** Filesystem
|
||||||
|
binaries hold NO establishment grants: no `open device-manager`, no
|
||||||
|
registry name to look up. The only block channel a filesystem process ever
|
||||||
|
has is the one handed to it at spawn by the volume manager. Serving the
|
||||||
|
wrong volume, a second volume, or a self-discovered volume is not
|
||||||
|
forbidden but *impossible* — the same way a driver cannot claim hardware
|
||||||
|
it was not delegated (device-authority.md).
|
||||||
|
2. **The volume manager supervises with teeth.** Mount-within-deadline or be
|
||||||
|
stopped (a filesystem wedged on a corrupt volume is killed and crash-loop
|
||||||
|
capped; the volume is marked bad, not retried forever). On removal the
|
||||||
|
manager KILLS the filesystem process and retires its mounts — the polite
|
||||||
|
observe-`EPEER`-and-exit path is an optimization; the kill is the
|
||||||
|
guarantee, because a process cannot be trusted to observe its own
|
||||||
|
obsolescence (the reap argument, proven on the device tree).
|
||||||
|
3. **The shared harness is how every filesystem inherits the lifecycle by
|
||||||
|
construction.** The harness — not the engine — owns the state machine:
|
||||||
|
establishment at spawn, mount registration, the dirty-flag set/clear
|
||||||
|
bracket, error-out-and-exit on channel death. The engine sits behind the
|
||||||
|
four-function vtable and never sees a channel; it cannot opt out of the
|
||||||
|
lifecycle for the same reason it cannot find a device. This is why the
|
||||||
|
harness is extracted BEFORE the second engine is written.
|
||||||
|
|
||||||
|
Beneath all three, two kernel mechanisms: ownership-gated `fs_unmount` (only
|
||||||
|
the mounting endpoint's holder may unmount), and the lazy dead-endpoint mount
|
||||||
|
sweep as the backstop nothing can disable. And above them, the check: a
|
||||||
|
**lifecycle conformance drill** — mount, serve, yank mid-write, verify honest
|
||||||
|
loss, replug, remount — parameterized over filesystem implementations, so
|
||||||
|
"danos supports filesystem X" MEANS "X passes the drill through the harness",
|
||||||
|
exactly as provider conformance means passing the reserved-verb suite.
|
||||||
@@ -0,0 +1,225 @@
|
|||||||
|
# The storage design rationale: why the stack is shaped this way
|
||||||
|
|
||||||
|
*2026-08-09. The design record behind
|
||||||
|
[storage-architecture.md](storage-architecture.md): the survey, the
|
||||||
|
trade-offs, and the decisions with their reasons — kept so future changes
|
||||||
|
argue against the evidence rather than rediscovering it. The questions that
|
||||||
|
drove it, verbatim: should the block protocol be separate from the VFS? how
|
||||||
|
do channels work with these block devices? how should we wire up different
|
||||||
|
filesystems? Grounded in how Minix 3, QNX Neutrino, Fuchsia, Redox, Plan 9,
|
||||||
|
and Linux each answered the same questions, and in exactly where our own
|
||||||
|
seams sat. The layering rule it serves: drivers are the lowest level
|
||||||
|
(hardware only); VFS and the filesystems are higher layers; the protocol
|
||||||
|
layer routes between them.*
|
||||||
|
|
||||||
|
## What the survey says, compressed
|
||||||
|
|
||||||
|
**Everyone separates block from VFS.** Even Plan 9, which unifies the *protocol*
|
||||||
|
(a disk is a file tree), keeps the layers: kernel `sd` serves raw ranges,
|
||||||
|
user-space filesystem servers consume them and serve files. The two poles are
|
||||||
|
Minix 3 (every layer a process: one VFS server, one filesystem server *per
|
||||||
|
mounted volume*, driver processes below — driver crashes proven survivable by
|
||||||
|
fault-injection, tens of thousands of injected faults) and QNX (filesystems as
|
||||||
|
DLLs inside the disk-driver process — fastest path, but one corrupt volume takes
|
||||||
|
every volume behind that controller down). Fuchsia is the modern capability-
|
||||||
|
native reference: a tiny block contract that every layer speaks, control channel
|
||||||
|
separate from a data fast-path with pre-registered buffers, filesystems as
|
||||||
|
separate processes launched by a storage-policy component (`fshost`) that probes
|
||||||
|
content and hands each filesystem its block channel at startup. The filesystem
|
||||||
|
never discovers devices.
|
||||||
|
|
||||||
|
**Partition tables are parsed in user space, below the filesystem, above the
|
||||||
|
raw device — never in the kernel, never in the filesystem engine.** Fuchsia
|
||||||
|
tried partitions-as-drivers and formally reversed it (RFC-0257). Plan 9 has the
|
||||||
|
cleanest split of all: the kernel offers only a *mechanism* — "create a named
|
||||||
|
sub-range of this disk" — and a user-space prober parses the table and issues
|
||||||
|
those commands. QNX (`io-blk`) and Minix (driver-side shared library) agree on
|
||||||
|
placement: the block layer multiplexes one physical block object into several
|
||||||
|
logical block objects **of the same contract**.
|
||||||
|
|
||||||
|
**The failure lessons are unanimous.** Linux's shared page cache across the
|
||||||
|
filesystem boundary produced fsyncgate (write errors observed by the wrong
|
||||||
|
process, dirty pages marked clean); the lesson is to keep write caching *inside*
|
||||||
|
the filesystem process, so an error surfaces on the channel that owns the
|
||||||
|
volume. Linux's `errors=remount-ro`/lazy-unmount half-alive states exist because
|
||||||
|
a monolith cannot kill and respawn a filesystem; Plan 9 leaves dead servers'
|
||||||
|
mounts in the namespace erroring forever. A restartable-parts system should have
|
||||||
|
exactly one surprise-removal path: kill the filesystem process, retire its
|
||||||
|
mounts, respawn on return. QNX's deepest lesson: after years of removable-media
|
||||||
|
pain they concluded detection isn't enough and built a copy-on-write filesystem
|
||||||
|
(Power-Safe) — *format-level* crash honesty. And QNX's `mcd` (a standalone
|
||||||
|
policy daemon with declarative insertion/removal rules) plus Fuchsia's `fshost`
|
||||||
|
agree that media policy is one dedicated component, not code scattered through
|
||||||
|
filesystems.
|
||||||
|
|
||||||
|
## Where we already stand
|
||||||
|
|
||||||
|
Closer than expected. The block protocol is 61 lines, five verbs, and its
|
||||||
|
docblock already reserves `Header.target` for volumes ("targets = volumes,
|
||||||
|
`enumerate` lists them; nothing else about the protocol changes"). usb-storage
|
||||||
|
is content-blind — it reads block 0 only as a bring-up self-check and parses
|
||||||
|
nothing. The vfs protocol is already served by a second provider in production
|
||||||
|
(init's registry serves it synthetically), so a second filesystem is *not* a
|
||||||
|
protocol problem. The FAT service is already internally split: a host-testable
|
||||||
|
engine (1659 + 310 lines) behind a four-function BlockDevice vtable, and a
|
||||||
|
434-line IPC shell. The kernel mount table already has the right restart
|
||||||
|
semantics (remount-replace; lazy dead-endpoint sweep on resolve).
|
||||||
|
|
||||||
|
The misplacements, all small: the MBR walk lives *inside the FAT engine*
|
||||||
|
(`engine.zig:196-212`) — every future engine would duplicate it; fat hardcodes
|
||||||
|
its acquisition policy ("first mass-storage child, mount at /volumes/usb") —
|
||||||
|
that is policy in a filesystem binary; `fs_unmount` is gated by nothing (any
|
||||||
|
process may unmount any prefix — latent now, an obvious cross-tenant hole once
|
||||||
|
mounts multiply); and fat's `mounted` flag is one-way — no path back after its
|
||||||
|
block channel dies.
|
||||||
|
|
||||||
|
## The proposed shape
|
||||||
|
|
||||||
|
Three layers, matching the stated rule, every boundary a channel:
|
||||||
|
|
||||||
|
**Block (driver layer, mechanism).** usb-storage stays hardware→blocks and
|
||||||
|
content-blind. It gains one mechanism, Plan 9's: *named sub-ranges*. A
|
||||||
|
`define_range`-style op creates a logical block object (a partition) clamped
|
||||||
|
and offset onto the disk; sub-ranges are `target` ids on the same endpoint, or
|
||||||
|
separate endpoints handed out per range — either way they speak the **same
|
||||||
|
block contract** (Fuchsia's closed-under-layering rule), so a filesystem cannot
|
||||||
|
tell whole-disk from partition. The driver never parses a table; it clamps
|
||||||
|
ranges it is told about. Offsets translate at definition time, so the data path
|
||||||
|
stays one hop (Fuchsia's session-mapping trick, for free).
|
||||||
|
|
||||||
|
**Volumes (service layer, policy).** One new service — the *volume manager*
|
||||||
|
(fshost/mcd-shaped, init-supervised; the device manager stays devices-only). It
|
||||||
|
subscribes to the device manager's `child_added`/`child_removed`, consumer-
|
||||||
|
hellos for each new mass-storage provider's block channel, reads the partition
|
||||||
|
table and the first blocks itself (the prober is policy), consults
|
||||||
|
configuration — `filesystems.csv`: content signature → filesystem binary;
|
||||||
|
`mounts.csv` or similar: volume identity → mount prefix — defines sub-ranges on
|
||||||
|
the driver, spawns **one filesystem process per volume**, hands it its block
|
||||||
|
channel at startup, and supervises it with the reap-and-rebuild idiom the
|
||||||
|
device manager proved: provider dies or medium leaves → kill the filesystem
|
||||||
|
process, retire its mounts; medium returns → re-probe, respawn, remount. The
|
||||||
|
boot volume is chosen by *content* (which volume carries /system/configuration
|
||||||
|
and /system/logs), closing the two-sticks question honestly.
|
||||||
|
|
||||||
|
**Filesystems (per volume, one process).** The proven unit everywhere from
|
||||||
|
Plan 9's `dossrv` to Minix to Fuchsia: block-client + engine + file-protocol
|
||||||
|
provider in one binary, one process per volume (9front practice; per-volume
|
||||||
|
fault isolation is what our supervision makes cheap). fat's shell becomes a
|
||||||
|
shared *filesystem harness* library before a second engine is written; the MBR
|
||||||
|
walk moves out of the engine into the volume manager; write caching stays
|
||||||
|
inside the process (the anti-fsyncgate rule). Each mounts its prefixes into the
|
||||||
|
kernel mount table itself, exactly as today.
|
||||||
|
|
||||||
|
**Kernel: two small changes only.** `fs_unmount` gains ownership (only the
|
||||||
|
mounting endpoint's holder may unmount — possession-is-capability, consistent
|
||||||
|
with everything else), and the 8-slot mount table gets a declared bound or
|
||||||
|
growth once volumes multiply. The mount table stays the router; per-process
|
||||||
|
namespaces (Plan 9's extra) remain separable future work.
|
||||||
|
|
||||||
|
## Transport generality: NVMe and SATA against this design
|
||||||
|
|
||||||
|
The layering was chosen transport-agnostic on purpose (the 9front lesson:
|
||||||
|
filesystem services cannot tell a USB stick from an ATA disk); NVMe and AHCI
|
||||||
|
are the test of that claim, and they fit — with three named pressure points.
|
||||||
|
|
||||||
|
**What transfers untouched:** the block protocol (nothing USB in it), the
|
||||||
|
volume manager, partitions/ranges, filesystems, mounts, and the whole
|
||||||
|
enforcement stack — delegation, IOMMU confinement (these are first-party PCI
|
||||||
|
DMA masters, so confinement applies even more directly than USB),
|
||||||
|
reap-and-rebuild, the hot-plug lifecycle. Integration cost per transport is
|
||||||
|
what the architecture promises: one driver binary plus one `devices.csv` row.
|
||||||
|
|
||||||
|
**NVMe** shortens the stack — no bus/class split; the driver IS the
|
||||||
|
controller driver, one hop fewer than USB. Its structural novelty,
|
||||||
|
**namespaces** (hardware-native multiple volumes behind one controller), is
|
||||||
|
exactly the case the block protocol reserved on day one and decision 4 below
|
||||||
|
settles as endpoint-per-volume. **AHCI** is a shape choice, not a problem: an
|
||||||
|
HBA fronts up to 32 ports plus port multipliers — structurally a bus — so
|
||||||
|
either mirror USB (ahci-bus + a per-port disk driver: maximum restart
|
||||||
|
granularity, the proven shape) or mirror NVMe (one driver per controller, one
|
||||||
|
block channel per port). Leaning per-port processes for consistency with the
|
||||||
|
matrix-proven shape; genuinely open.
|
||||||
|
|
||||||
|
**The pressure points, honestly:**
|
||||||
|
|
||||||
|
1. **Multi-volume providers are reserved, not implemented.** The volume
|
||||||
|
manager flow assumes one provider, one volume; NVMe namespaces make
|
||||||
|
endpoint-per-volume real work with hardware demanding it.
|
||||||
|
2. **The current transport will bottleneck NVMe.** Synchronous call/reply,
|
||||||
|
one operation in flight, one bounce buffer — fine for a USB2 stick,
|
||||||
|
forfeits an NVMe drive's queue depth and per-queue MSI-X. Correctness
|
||||||
|
needs nothing; performance is gated on the shm-ring data plane
|
||||||
|
communication.md already names (Fuchsia's FIFO+VMO is the precedent).
|
||||||
|
NVMe is not worth building before that milestone.
|
||||||
|
3. **Media lifecycle ≠ device lifecycle.** ATAPI trays and SD card readers
|
||||||
|
(USB ones too, today) keep the DEVICE present while the MEDIUM comes and
|
||||||
|
goes — a removal trigger our channel-death lifecycle does not carry. The
|
||||||
|
fix is decision 7 below. (Related small item: fat's bounce sizing
|
||||||
|
hardcodes 512-byte sectors; 4Kn drives need the geometry honored
|
||||||
|
throughout.)
|
||||||
|
|
||||||
|
## The decisions on the table
|
||||||
|
|
||||||
|
1. **Partitions as driver-side sub-ranges** (mechanism in usb-storage, parsing
|
||||||
|
in the volume manager) — versus a separate partition *process* re-serving
|
||||||
|
block (Fuchsia's storage-host). Sub-ranges are ~20 lines of clamp in the
|
||||||
|
driver and keep the data path one hop; a separate process is purer layering
|
||||||
|
at the cost of a hop or session plumbing. Recommendation: sub-ranges.
|
||||||
|
2. **The volume manager as a new service** owning probe, spawn, supervision,
|
||||||
|
and mount policy — with `filesystems.csv` and the mount map as
|
||||||
|
configuration. Recommendation: yes; it is the missing policy home that fat
|
||||||
|
is currently squatting in.
|
||||||
|
3. **One filesystem process per volume** (fat's binary becomes "the FAT
|
||||||
|
implementation", spawned per FAT volume). Recommendation: yes — it extends
|
||||||
|
recompile-and-restart-live to filesystems and isolates corrupt media.
|
||||||
|
4. **Sub-range addressing**: `target` ids on the storage endpoint versus one
|
||||||
|
endpoint per volume handed out by the driver. Endpoint-per-volume matches
|
||||||
|
the establishment-plane machinery (a channel per party, caps at
|
||||||
|
establishment) and keeps per-client badge scoping simple. Recommendation:
|
||||||
|
endpoint per volume.
|
||||||
|
5. **`fs_unmount` ownership** — a defect fix more than a decision.
|
||||||
|
6. **Later, kept open**: the shm-ring data plane (communication.md already
|
||||||
|
names it as the 256-byte ceiling's unlock — Fuchsia's FIFO+VMO is the
|
||||||
|
precedent); format-level crash honesty (a Power-Safe-style journaling or COW
|
||||||
|
filesystem) once danos outgrows FAT; per-process namespaces.
|
||||||
|
7. **The media-presence event** (settled in principle; lands with the volume
|
||||||
|
manager): the block protocol gains a pushed event — `medium_changed`, with
|
||||||
|
present/absent and a change counter — produced by the storage driver from
|
||||||
|
its transport's native signal (SCSI UNIT ATTENTION / TEST UNIT READY for
|
||||||
|
USB and ATAPI, PxSSTS for AHCI, namespace-change AER for NVMe) and
|
||||||
|
consumed by the volume manager, which runs the SAME kill-retire-remount
|
||||||
|
path it runs on channel death — one lifecycle, two triggers. The driver
|
||||||
|
reports presence, never content; a pushed event carries no capability,
|
||||||
|
which the kernel already guarantees. The device staying while its medium
|
||||||
|
leaves is the one removable-media case the channel-death trigger cannot
|
||||||
|
see; without this event a swapped SD card would be served with the old
|
||||||
|
card's filesystem state.
|
||||||
|
8. **Volume identity, and the mount map as danos's fstab** (settled). The
|
||||||
|
lesson is Linux's own history: fstab keyed on `/dev/sda1` for years and
|
||||||
|
broke whenever a drive changed ports or enumeration order; `UUID=` entries
|
||||||
|
exist because device-path identity failed. danos skips that era: the mount
|
||||||
|
map (`volumes.csv` — configuration, read by the volume manager) keys on
|
||||||
|
**content identity, never port or discovery order**. The prober reads
|
||||||
|
identity off the medium, strongest first:
|
||||||
|
1. GPT partition GUID — 128-bit, unique, stable for the volume's life;
|
||||||
|
2. filesystem UUID (ext-family and most modern formats, in the superblock);
|
||||||
|
3. FAT volume serial + label — 32 bits, weak (dd-cloned sticks share it)
|
||||||
|
but what real sticks carry;
|
||||||
|
4. MBR disk signature + partition index;
|
||||||
|
5. nothing — an anonymous volume: generated mount name, no persistence.
|
||||||
|
|
||||||
|
Consequences, each mechanical once identity keys the map: **moving a drive
|
||||||
|
to a different port changes nothing** — same identity, same mount point,
|
||||||
|
whether USB port, hub depth, SATA port, or a stick that left as USB and
|
||||||
|
returned in a SATA dock; **replug remounts at the same path** (the
|
||||||
|
remount-after-return story is a map lookup); **the boot volume** is the
|
||||||
|
recorded identity of the volume carrying `/system/configuration`, findable
|
||||||
|
on any port; and **duplicate identity is a policy case, not a surprise** —
|
||||||
|
two cloned sticks at once: first keeps the mapped name, second mounts
|
||||||
|
suffixed and is logged loudly, never silently shadowed. Unknown identities
|
||||||
|
mount under a derived name (sanitized label, else generated) at
|
||||||
|
`/volumes/<name>` — the hierarchy's documented home for attached media,
|
||||||
|
which stands: `/system` is what danos IS; attached media is what it isn't.
|
||||||
|
danos never needs to MINT identifiers to detect volumes — detection only
|
||||||
|
reads — until it grows formatting, which brings the entropy question and
|
||||||
|
is deliberately out of scope here.
|
||||||
@@ -35,6 +35,14 @@ pub const Device = struct {
|
|||||||
return self.call(.attach, {}, handle, &reply) != null;
|
return self.call(.attach, {}, handle, &reply) != null;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// The reverse of `attach`: the buffer leaves the device's reach. The same
|
||||||
|
/// region capability rides again (the kernel matches the region). Do not name
|
||||||
|
/// the buffer's physical address in `read`/`write` after this.
|
||||||
|
pub fn detach(self: Device, handle: ipc.Handle) bool {
|
||||||
|
var reply: [block_protocol.message_maximum]u8 = undefined;
|
||||||
|
return self.call(.detach, {}, handle, &reply) != null;
|
||||||
|
}
|
||||||
|
|
||||||
/// Read `count` blocks starting at `lba` into the DMA buffer at `physical`.
|
/// Read `count` blocks starting at `lba` into the DMA buffer at `physical`.
|
||||||
pub fn read(self: Device, lba: u64, count: u32, physical: u64) bool {
|
pub fn read(self: Device, lba: u64, count: u32, physical: u64) bool {
|
||||||
var reply: [block_protocol.message_maximum]u8 = undefined;
|
var reply: [block_protocol.message_maximum]u8 = undefined;
|
||||||
|
|||||||
@@ -136,6 +136,15 @@ pub const Device = struct {
|
|||||||
return self.call(.dma_attach, {}, &.{}, handle, &reply) != null;
|
return self.call(.dma_attach, {}, &.{}, handle, &reply) != null;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// The reverse of `attachDma`: unbind the buffer from the controller's IOMMU
|
||||||
|
/// domain. The same region capability rides again — the kernel matches the
|
||||||
|
/// region, so neither side kept state between the two calls. Do not name the
|
||||||
|
/// buffer's physical address in any transfer after this.
|
||||||
|
pub fn detachDma(self: *Device, handle: ipc.Handle) bool {
|
||||||
|
var reply: [usb_transfer_protocol.message_maximum]u8 = undefined;
|
||||||
|
return self.call(.dma_detach, {}, &.{}, handle, &reply) != null;
|
||||||
|
}
|
||||||
|
|
||||||
/// One bulk transfer (IN or OUT per `endpoint_address`'s direction bit) to or
|
/// One bulk transfer (IN or OUT per `endpoint_address`'s direction bit) to or
|
||||||
/// from the caller's own DMA buffer at `physical`. Returns the bytes moved.
|
/// from the caller's own DMA buffer at `physical`. Returns the bytes moved.
|
||||||
pub fn bulk(self: *Device, endpoint_address: u8, physical: u64, length: u32) ?u32 {
|
pub fn bulk(self: *Device, endpoint_address: u8, physical: u64, length: u32) ?u32 {
|
||||||
|
|||||||
@@ -54,6 +54,12 @@ pub const Protocol = envelope.Define(.{
|
|||||||
// physical addresses (named in later read/write) are reachable by the
|
// physical addresses (named in later read/write) are reachable by the
|
||||||
// device under an enforcing IOMMU. Call once per buffer before using it.
|
// device under an enforcing IOMMU. Call once per buffer before using it.
|
||||||
.{ .name = "attach" },
|
.{ .name = "attach" },
|
||||||
|
// detach(): the reverse — the same region capability rides the cap slot
|
||||||
|
// (the caller still holds its handle; the kernel matches the region) and
|
||||||
|
// the buffer leaves the device's domain. Every grant a live process
|
||||||
|
// makes is revocable by the granter while alive; death remains the
|
||||||
|
// mechanical backstop (storage-architecture.md, the lifecycle rule).
|
||||||
|
.{ .name = "detach" },
|
||||||
},
|
},
|
||||||
});
|
});
|
||||||
|
|
||||||
|
|||||||
@@ -156,6 +156,11 @@ pub const Protocol = envelope.Define(.{
|
|||||||
// target says which caller's device is attaching, so there is nothing
|
// target says which caller's device is attaching, so there is nothing
|
||||||
// left for a body to carry.
|
// left for a body to carry.
|
||||||
.{ .name = "dma_attach" },
|
.{ .name = "dma_attach" },
|
||||||
|
// dma_detach: the reverse, same shape — the region capability rides the
|
||||||
|
// cap slot again (the kernel matches the region; the provider retains
|
||||||
|
// nothing between the two calls) and the buffer leaves the controller's
|
||||||
|
// domain. Appended, so every existing verb keeps its number.
|
||||||
|
.{ .name = "dma_detach" },
|
||||||
},
|
},
|
||||||
.events = &.{
|
.events = &.{
|
||||||
.{ .name = "interrupt_report", .payload = InterruptReport },
|
.{ .name = "interrupt_report", .payload = InterruptReport },
|
||||||
@@ -188,6 +193,7 @@ test "the verb numbering, and the device token in the header" {
|
|||||||
try std.testing.expectEqual(@as(u32, 18), @intFromEnum(Operation.interrupt_subscribe));
|
try std.testing.expectEqual(@as(u32, 18), @intFromEnum(Operation.interrupt_subscribe));
|
||||||
try std.testing.expectEqual(@as(u32, 19), @intFromEnum(Operation.bulk));
|
try std.testing.expectEqual(@as(u32, 19), @intFromEnum(Operation.bulk));
|
||||||
try std.testing.expectEqual(@as(u32, 20), @intFromEnum(Operation.dma_attach));
|
try std.testing.expectEqual(@as(u32, 20), @intFromEnum(Operation.dma_attach));
|
||||||
|
try std.testing.expectEqual(@as(u32, 21), @intFromEnum(Operation.dma_detach));
|
||||||
try std.testing.expectEqual(@as(u32, 16), @intFromEnum(Event.interrupt_report));
|
try std.testing.expectEqual(@as(u32, 16), @intFromEnum(Event.interrupt_report));
|
||||||
|
|
||||||
var buffer: [message_maximum]u8 = undefined;
|
var buffer: [message_maximum]u8 = undefined;
|
||||||
|
|||||||
@@ -204,12 +204,20 @@ fn onAttach(_: void, invocation: Invocation(void), _: Answer(void)) isize {
|
|||||||
return if (device.attachDma(handle)) 0 else refused;
|
return if (device.attachDma(handle)) 0 else refused;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// The reverse: forward the same region capability so the controller unbinds
|
||||||
|
/// the buffer. As with attach, our copy stays the turn's to close.
|
||||||
|
fn onDetach(_: void, invocation: Invocation(void), _: Answer(void)) isize {
|
||||||
|
const handle = invocation.capability orelse return -envelope.EPROTO;
|
||||||
|
return if (device.detachDma(handle)) 0 else refused;
|
||||||
|
}
|
||||||
|
|
||||||
const handlers = Serve.Handlers{
|
const handlers = Serve.Handlers{
|
||||||
.geometry = onGeometry,
|
.geometry = onGeometry,
|
||||||
.read = onRead,
|
.read = onRead,
|
||||||
.write = onWrite,
|
.write = onWrite,
|
||||||
.flush = onFlush,
|
.flush = onFlush,
|
||||||
.attach = onAttach,
|
.attach = onAttach,
|
||||||
|
.detach = onDetach,
|
||||||
};
|
};
|
||||||
|
|
||||||
fn onMessage(message: []const u8, reply: []u8, sender: u32, arrived: *ipc.Arrival) usize {
|
fn onMessage(message: []const u8, reply: []u8, sender: u32, arrived: *ipc.Arrival) usize {
|
||||||
|
|||||||
@@ -593,6 +593,7 @@ const handlers = Serve.Handlers{
|
|||||||
.interrupt_subscribe = onInterruptSubscribe,
|
.interrupt_subscribe = onInterruptSubscribe,
|
||||||
.bulk = onBulk,
|
.bulk = onBulk,
|
||||||
.dma_attach = onDmaAttach,
|
.dma_attach = onDmaAttach,
|
||||||
|
.dma_detach = onDmaDetach,
|
||||||
};
|
};
|
||||||
|
|
||||||
/// open: the target is the class driver's assigned device id. Resolve it to an
|
/// open: the target is the class driver's assigned device id. Resolve it to an
|
||||||
@@ -694,6 +695,14 @@ fn onDmaAttach(_: void, invocation: Invocation(void), _: Answer(void)) isize {
|
|||||||
return if (device.dmaBind(controller_id, handle)) 0 else refused;
|
return if (device.dmaBind(controller_id, handle)) 0 else refused;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// The reverse: the same region capability arrives again and the buffer leaves
|
||||||
|
/// the controller's domain (`dma_unbind` matches the region — nothing was
|
||||||
|
/// retained here between the two calls). The turn closes the arriving copy.
|
||||||
|
fn onDmaDetach(_: void, invocation: Invocation(void), _: Answer(void)) isize {
|
||||||
|
const handle = invocation.capability orelse return -envelope.EPROTO;
|
||||||
|
return if (device.dmaUnbind(controller_id, handle)) 0 else refused;
|
||||||
|
}
|
||||||
|
|
||||||
/// A timer tick or an MSI landed: drain the event ring, reconcile ports, and fan out.
|
/// A timer tick or an MSI landed: drain the event ring, reconcile ports, and fan out.
|
||||||
/// The timer arm re-arms itself (8 ms drain when polling, 250 ms reconcile under MSI);
|
/// The timer arm re-arms itself (8 ms drain when polling, 250 ms reconcile under MSI);
|
||||||
/// the MSI arm clears the interrupter's pending bit FIRST, then drains — so an event
|
/// the MSI arm clears the interrupter's pending bit FIRST, then drains — so an event
|
||||||
|
|||||||
@@ -196,10 +196,22 @@ fn tryBringUp() void {
|
|||||||
// enforcing IOMMU. No-op binding otherwise.
|
// enforcing IOMMU. No-op binding otherwise.
|
||||||
const bounce = memory.dmaAlloc(engine.max_transfer_sectors * 512, memory.dma_coherent | memory.dma_shareable) orelse return;
|
const bounce = memory.dmaAlloc(engine.max_transfer_sectors * 512, memory.dma_coherent | memory.dma_shareable) orelse return;
|
||||||
if (bounce.handle) |handle| {
|
if (bounce.handle) |handle| {
|
||||||
|
// Attach, detach, and attach again: the round trip exercises BOTH verbs
|
||||||
|
// of the DMA-window lifecycle through the whole chain (fat → storage →
|
||||||
|
// bus → kernel) on every boot, so a broken detach fails every fat case
|
||||||
|
// rather than lying dormant until the first buffer replacement.
|
||||||
if (!device.attach(handle)) {
|
if (!device.attach(handle)) {
|
||||||
_ = logging.write("/system/services/fat: could not attach the DMA bounce buffer\n");
|
_ = logging.write("/system/services/fat: could not attach the DMA bounce buffer\n");
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
if (!device.detach(handle)) {
|
||||||
|
_ = logging.write("/system/services/fat: could not detach the DMA bounce buffer\n");
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (!device.attach(handle)) {
|
||||||
|
_ = logging.write("/system/services/fat: could not re-attach the DMA bounce buffer\n");
|
||||||
|
return;
|
||||||
|
}
|
||||||
_ = ipc.close(handle); // the binding holds its own reference now
|
_ = ipc.close(handle); // the binding holds its own reference now
|
||||||
}
|
}
|
||||||
ipc_block = .{ .device = device, .bounce = bounce };
|
ipc_block = .{ .device = device, .bounce = bounce };
|
||||||
|
|||||||
Reference in New Issue
Block a user