Commit Graph
494 Commits
Author SHA1 Message Date
Daniel Samson 4f9196c03e docs: storage plan — fix two id-path leftovers (DANOS/label->hex)
Two spots still said /volumes/DANOS and "label->hex", contradicting the settled
path-is-the-id decision. Corrected to id-path.
2026-08-09 22:55:10 +01:00
Daniel Samson eff95416d0 docs: storage plan — path is the id, label is queryable display metadata
Refine the naming decision: a volume's mount path IS its identity id (GPT GUID,
else fat-<serial>/mbr-<sig>-<index>) — a stable, unique, content-derived handle
software uses. The label (FAT volume label / GPT partition name) is mutable
display metadata, NOT in the path, exposed by a volume-manager `volumes` verb
returning {id, mount_path, label} — the database id/name split. This dissolves
the label-collision problem: same-label-different-id volumes get distinct paths
automatically; only identical ids (dd-clones) hit first-wins-and-log. S1 carries
both key (id) and label (display); S2 derives the id-path and adds the query.
2026-08-09 22:54:36 +01:00
Daniel Samson addd264880 docs: storage plan — volumes named by identity; exFAT implemented in full
Two decisions settled: (1) a volume's mount name is its own content identity
(FAT label / GPT name, else identity-hex), never a port name (/volumes/usb) or
role name (/volumes/boot); volumes.csv stays as an explicit override. This makes
S1 precede S2 and folds the /volumes/usb fixture+regex migration into S2. (2)
exFAT is a complete implementation — full read+write, directories, rename, and
the on-disk up-case table — not a read-first/ASCII-only subset; the only limit
is the vfs u32 offset surface (a 4 GiB cap on all filesystems), flagged as a
separate vfs change.
2026-08-09 22:42:56 +01:00
Daniel Samson fca41b351e docs: the storage-stack completion plan (S1-S5)
The phased plan taking the volume-manager track from one FAT volume to any
filesystem, N volumes, content identity, remount, and driver-crash survival:
S1 identity ladder (GPT GUID + FAT serial), S2 volumes.csv/filesystems.csv mount
map, S3 multi-volume, S4 exFAT (the second engine proving the harness reuse),
S5 removal robustness (medium_changed consumption + driver-crash rebuild +
QEMU-verified remount). Code-grounded: real functions, commit-granular steps,
discrimination tests shown to fail against today's behavior, per-phase risks.
Design decisions taken with recommended defaults; two forks flagged for review
(the /volumes/usb transitional naming and the exFAT write scope). Not started.
2026-08-09 22:34:00 +01:00
Daniel Samson bf0595763e docs: flip stale status markers across the tracks (audit found 23)
An all-tracks docs-vs-code audit (the same method that caught the storage
drift) found 23 confirmed inaccuracies where a doc's build-status claim no
longer matches the source — status markers that were never flipped after a
track landed, and a few paths left over from completed flag-days. All verified
against the code before editing; docs only, no behavior change.

The systemic ones:
- IOMMU enforcement (driver-model.md, drivers.md): docs said enforcement was
  not built and "device_claim = ring 0" / "memory-safe is not true yet". It is
  built (per-device VT-d/AMD-Vi domains programmed at device_claim, -ECONFINE
  rollback, dma_alloc buffers bound and torn down at death; fail-open only with
  no IOMMU). Restated; M16 marker flipped to done.
- The FHS flag-day paths: /etc/devices.csv -> /system/configuration/devices.csv
  (devices-csv.md, new-driver-checklist.md, device-manager.md), /var/log ->
  /system/logs (logging.md, new-driver-checklist.md), /mnt/usb -> /volumes/usb
  (process-management.md). Following the old paths silently breaks driver match.
- protocol-namespace P4 "remaining" -> landed (only P5 remains); shared-fate
  fan-out "not yet enforced" -> enforced; wall_clock "not built" -> built;
  SMP affinity + fault-recovery "left" -> built; process_enumerate raw-pointer
  trust model -> checked copyToUser/EFAULT; bounds.md maximum_devices static
  hole -> dynamic per-registrar quota; init spawns fat -> volume-manager;
  config "hardcoded, move to /etc" -> already CSV data files; vdso.md three-
  value call; zig-self-hosting library/ layout; python argv "new" -> built.

Found and fixed by a multi-agent audit across 12 doc clusters, each finding
adversarially verified against the source.
2026-08-09 22:04:17 +01:00
Daniel Samson 8216be991d docs: correct the storage docs' V0-V4 status — the V5 flip missed several
A V5 close-out audit (docs against the actual code) found the storage docs
overclaiming in both directions: the status header was flipped to "built" but
several body markers were not, and two passages describe mechanisms the code
never implemented. Nine confirmed, adversarially verified against the source:

Underclaims (marked Planned, actually built):
- storage-architecture: driver named sub-ranges / range confinement (V2, tested
  by the block-range case); the pushed medium_changed event (published today
  from a TEST UNIT READY poll); ownership-gated fs_unmount (V0, process.zig
  gates it with EPERM); the filesystem-harness extraction (V1). Scoped the
  remaining *planned* to the genuinely-pending parts (native-signal translation,
  the volume manager consuming medium_changed).

Overclaims (described, never built):
- storage-architecture: the "FAT dirty flag on disk" guarantee — no on-disk
  dirty/clean-shutdown bit exists; only an in-memory device-dirty bool gating a
  device write-cache flush on close. Fixed in all three places.
- storage-architecture: fat "acquires its own volume (first mass-storage child
  by enumeration order)" — the V3b flip removed self-acquisition; fat is handed
  its volume id and channel by the volume manager.
- rationale + plan: the FAT engine's base_lba "deleted rather than moved" — it
  and the engine's MBR walk still exist as now-inert legacy; the authoritative
  walk lives in partition.zig.
- rationale: NVMe namespaces "decision 4 settles as endpoint-per-volume" —
  decision 4 settles the opposite (per-sender confinement, one endpoint);
  endpoint-per-volume is named only as an unbuilt future refactor.
- rationale: the five-rung identity ladder and volumes.csv map stated in flat
  present tense — only rung 4 (MBR signature + index) is built; added the
  build-status hedge and marked each rung.

Docs only; no code or behavior change. Suite unaffected (127/127).
2026-08-09 21:19:29 +01:00
Daniel Samson 68e65803eb volume-manager: the removal comment says lazy retirement, not an eager sweep
The V4 adversarial review found removeVolume's comment overclaiming: it said
"the kernel sweeps a dead backend's mounts", which reads as an eager death-time
sweep. There is no such sweep. Killing the filesystem marks its backend endpoint
dead (killOwnedEndpointsLocked), and the VFS router retires each mount that
endpoint backed lazily, on the next path resolution under it (resolvePath sees
the dead backend, frees the slot, returns not_found). The functional guarantee
the comment promised — killing the filesystem retires its mounts — holds; only
the described mechanism was wrong. Comment-only; no behavior change.
2026-08-09 20:56:51 +01:00
Daniel Samson af47d41989 volume-manager: removal supersedes a pending restart in the poll
Inline V4 review (the boundary-review workflow stalled): the poll ran a due
fat-restart before the presence check and returned, so a fat death followed
by a device removal would respawn fat against the now-dead channel and churn
until the crash cap before the removal was noticed. Reorder: check the
specific device's presence first (unmount if gone), and only fire a due
restart once the device is confirmed present. Neutral: fat-mount,
volume-removal, amd-iommu-usb-storage green.

Noted V4 limitations (not fixed here, edge cases outside the user unplug
case): a usb-storage DRIVER crash (device stays, driver restarts with a new
endpoint) leaves fat holding a dead channel — the device is still present so
removal is not detected; fat would need to observe its channel death and
exit. Deferred with the medium_changed subscription and multi-volume.
2026-08-09 20:27:58 +01:00
Daniel Samson 5e89b111cf docs: flip the storage-architecture status markers the V0-V4 track made real
The volume manager, per-sender range confinement + gate, medium_changed
event, per-volume spawning + supervision, ownership-gated fs_unmount, and
the removal half of the lifecycle are built. Left honestly pending: the
volumes.csv/filesystems.csv maps, the fuller identity ladder, multi-volume,
the VM consuming medium_changed (removal uses device-presence polling), and
remount-on-replug end-to-end (bench-pending — QEMU can't re-present the
boot-controller device). Full suite 127/127.
2026-08-09 20:20:20 +01:00
Daniel Samson e3ec9fa668 volume-manager: try every mass-storage entry, watch the one we opened
The full suite caught a V4 regression: under AMD-Vi the device-manager tree
carries more than one mass-storage-identity entry (a phantom no driver is
bound to, which answers a consumer hello with NO channel). V4 split presence
from acquisition and picked the FIRST identity match blindly, so it kept
helloing the phantom (device 27) and never reached the real storage (device
31). V3's inline loop had skipped no-channel entries with `orelse continue`;
the split lost that.

Restore it: openAnyStorage tries each matching entry and takes the first whose
channel opens, recording its device id. Removal detection then watches THAT
specific device id leave the tree (isDevicePresent), not "any mass-storage" —
so a phantom that never leaves cannot mask a real removal. Both are bare
enumerates; the hello only happens while bringing a volume up.

Green: amd-iommu-usb-storage, fat-mount, volume-removal.
2026-08-09 20:08:30 +01:00
Daniel Samson 9e67a74232 volume-manager: the removal lifecycle — a pulled stick unmounts (V4)
The volume manager stops probing-once and polls storage presence for the life
of the boot: findStorageDevice enumerates the device-manager tree (presence
only, no consumer-hello, so it is cheap and leaks nothing). The volume is now
a field that goes null and back — the whole lifecycle:

- storage present + no volume  -> open the channel, probe, confine + spawn the
  filesystem (openStorage is the one consumer-hello, on the insertion edge);
- storage gone + have volume    -> kill the filesystem (its mounts retire via
  the kernel dead-backend sweep), close the dead channel, clear the volume;
- fat crash                     -> the same supervised backoff/cap as before,
  folded into the poll (one timer).

This also subsumes the V3-review leak fix (no per-poll consumer-hello) and the
no-volume retry (a present-but-unreadable device keeps polling).

The user's case — pull the boot stick, plug it back — is a DEVICE unplug (the
stick IS the device), so the mass-storage child leaves the device-manager tree
and the poll catches it. volume-removal asserts the unmount and discriminates:
against the V3 probe-once volume manager the removal is never noticed (0/1).

The re-mount on replug is the VM's bringUpVolume firing when the device
returns — correct and in place, but not QEMU-testable here: device_add of
usb-storage to the boot xHCI controller is not re-presented to the guest (no
port-connect on any port), a harness quirk, not a VM issue. On real hardware
the bus's per-tick port poll catches a reconnect (H1 proves reconnect on a
second controller); bench-verify the full round trip.
2026-08-09 19:41:04 +01:00
Daniel Samson b9058fe020 file-system: the mount-failure log has no scratch buffer (orphaned bounds fix)
The mount-failure diagnostic wrote through a fixed [96]u8 bufPrint buffer,
which the bounds gate flags as an undeclared ceiling. Three ring appends
instead — no buffer, no fixed length to justify. This fix was made during
the V2a bounds work but never git-added, so the committed harness still
carried the flagged buffer; committing it now cleans the tip.
2026-08-09 19:40:39 +01:00
Daniel Samson 7c6ed2ca09 volume-manager: harden the probe and supervision from the V3 review
Five confirmed defects from the boundary review:

1. (security) The VM never checked a partition fit inside the device, so a
   crafted MBR could hand the driver a range whose base+lba wraps past a u32
   — panicking usb-storage in a loop, and at multi-volume overlapping a
   neighbour. This is the exact invariant the clamp's overflow-safety rests
   on. partition.firstVolume now skips any entry that runs past the device
   (host-tested), establishing the invariant where the untrusted bytes are
   first read.
2. (leak) The probe re-acquired a fresh block channel on every 500 ms retry,
   leaking a handle each time on a medium-absent device. The channel is now
   acquired once and kept.
3. (wedge) A failed spawn or defineRange stranded the volume with no retry;
   both now arm a backoff restart.
4. (loop) fat respawn had no exit-reason gate, no backoff, no crash-loop cap
   — a faulting filesystem respawned in a zero-delay loop, and a clean exit
   was resurrected. Supervision now mirrors the device manager: a clean exit
   is not restarted, a fault backs off, three fast deaths give up.
5. (removable) A device that parsed to no volume was terminal; it now keeps
   polling so an inserted medium is picked up — the removal-lifecycle trigger.

Known limitation (noted, not fixed here): if the VM itself crashes and init
restarts it, the orphaned fat keeps serving vfs while the new VM spawns a
second fat whose bind is refused — the same "manager restart re-learns the
world" gap the device manager also defers. The old fat keeps storage working.

Neutral: partition unit tests + fat-mount, volume-probe, block-range, logger
all green.
2026-08-09 18:50:01 +01:00
Daniel Samson a67a7015bf docs: record the V3c resequencing — removal lifecycle first, multi-volume follows 2026-08-09 18:30:16 +01:00
Daniel Samson 301bdcaf5b volume-manager: the flip — fat is spawned, confined, and handed its channel (V3b)
The load-bearing step. The FAT service stops acquiring its own volume: the
volume manager spawns it (per volume), defines its partition range on the
storage driver BEFORE it runs, and answers its startup hello with the
range-confined block channel over a new volume-manager protocol. fat never
finds its storage by name and never sees the whole device — establishment
by lineage, one layer up from the driver tree.

- New library/protocol/volume-manager: one verb, hello(volume-id) -> the
  block channel as the reply capability (the P0 reply-cap path).
- The volume manager becomes the confinement CONTROLLER: it defines the first
  range on usb-storage, so no other party can confine a filesystem. It
  supervises the filesystems it spawns and respawns one on death (the reap-
  and-rebuild the device manager proved, one layer up).
- fat: drops acquireVolume(device-manager); hellos the volume manager for its
  channel; reads its volume id from argv[1]. main takes process.Init now.
- init.csv no longer spawns fat (the volume manager does); protocol.csv
  rewires fat to be supervised by the volume manager (bind vfs, open
  volume-manager) and drops fat open device-manager.
- The block-range fixture boots registry + device-manager only (not the full
  tree), so the volume manager is absent and the fixture stays the sole
  confinement definer — otherwise the volume manager would take the
  controller first and refuse it.

Verified end to end (VM probes -> spawns fat -> confines it -> hands over the
channel -> fat mounts) and neutral: 18/18 across the fat family, logging,
shutdown, both IOMMU variants, usb restart, vfs, conformance, confinement.
2026-08-09 18:28:44 +01:00
Daniel Samson d56b1b81c0 volume-manager: discovery and probe — V3a
The storage layer gains its policy home (storage-architecture.md): a new
system/services/volume-manager, spawned by init, that acquires the mass-
storage block channel through the device manager (the same lineage a
filesystem uses), reads block 0, and parses the first volume out of it. The
partition-table walk that lived in the FAT engine moves here, above the
driver where it belongs (partition.zig, host-tested: MBR entry, bare-FAT,
no-signature). Identity is the MBR disk signature + partition index — the
weak rung of the ladder; GPT GUID and FAT serial refine identityOf without
changing shape.

This increment is discovery + probe + log only, additive: the FAT service
still acquires its own volume, so nothing changes for it. Confining each
filesystem to its partition and spawning one per volume (the flip) lands
next, keeping fat working throughout.

Grants + wiring: init.csv spawns it after the device manager; protocol.csv
grants bind volume-manager + open device-manager. Verified: volume-probe
asserts the parse (bare-FAT volume at lba 0), neutral 10/10 across storage,
restart, display, logging, confinement — the volume manager now runs in
every boot and disturbs nothing.
2026-08-09 18:10:43 +01:00
Daniel Samson bc67771bfd block: reclaim range slots on death, and give confinement one controller
Two defects the V2 boundary review confirmed:

1. The per-badge range table was never reclaimed. usb-storage receives exit
   notifications (via the subscriber watch), but onNotification handled only
   the timer and dropped child-exits, so a dead filesystem left its range
   slot used forever. The medium-removal lifecycle churns filesystems, so
   after maximum_ranges confine/die cycles define_range would return ENOSPC
   and no volume could be confined again until reboot. onNotification now
   frees the dead badge's slot, mirroring fat's open-node sweep.

2. define_range checked only that the CALLER was unconfined, never that it
   owned the target badge — so any unconfined opener could install a range
   for another live client and silently redirect its I/O. Confinement now has
   a single controller: the first unconfined party to define a range (the
   volume manager, which confines every filesystem before handing it a
   channel). Only the controller may thereafter; the slot releases on its
   death so a restarted manager re-takes it. This is the mechanism half; V3
   adds the grant half (only the volume manager gets an unconfined channel).

Neutral: block-range (the fixture is the sole definer -> controller),
fat-mount, usb-report all green. End-to-end exercise of both lands in V3/V4
(the VM+filesystem relationship and the remount churn).
2026-08-09 17:58:47 +01:00
Daniel Samson 89d4592777 block: close the range-clamp overflow — a confined caller could wrap into the neighbour
The naive bound `lba + count > r.count` wraps for an lba near u64 max: the
sum overflows to a small value, sails under the check, and `base + lba`
wraps to an absolute block OUTSIDE the range. Calibrated, it is a real
confinement escape — a process confined to [1,3) reads absolute block 0
(the boot sector) with lba = maxInt(u64), since base + lba wraps to 0.

The bound is rewritten as two subtractions that cannot overflow: lba within
the range, and count within what remains. block-range gains a wrap-refused
assertion calibrated to be exploitable against the naive form — it FAILS
against the old bound (reads block 0) and passes against the fix (verified
by reverting the clamp). Caught pre-emptively before the V2 boundary review.
2026-08-09 17:40:45 +01:00
Daniel Samson 7af65697cc block: medium presence — the medium_changed event and usb-storage as publisher (V2b)
The block protocol gains a pushed medium_changed event (present + a monotonic
change counter; presence only, never content). usb-storage becomes a
Subscribers provider and runs a slow TEST UNIT READY poll (1 s): success is
present, failure absent, and a transition bumps the counter and publishes.
This is the second removal trigger — the DEVICE stays while the MEDIUM leaves
(card readers, ATAPI trays) — which channel death cannot see
(storage-architecture.md, two triggers one lifecycle).

The subscriber is the volume manager (V3); until it exists the publish is a
no-op fan-out, so this commit is behaviour-neutral, and its end-to-end test
(eject -> medium_changed -> unmount/remount) lands in V4 with the real
consumer rather than a throwaway subscriber fixture (recorded sequencing).
Sense-key inspection to tell medium-absent from other transport errors is a
noted refinement; a clean eject reads correctly as not-ready.

Neutral: 12/12 across the block-serving surface, restart, confinement,
conformance, and logging.
2026-08-09 17:36:57 +01:00
Daniel Samson 73fbbd3922 docs: record the V2 sequencing — medium_changed emitted in V2b, tested in V4 2026-08-09 17:27:43 +01:00
Daniel Samson c37402891a test: block-range — the discrimination fixture for range confinement (V2a)
A process acquires a block channel the way a filesystem does (consumer-hello
the device manager), confines ITSELF to blocks [1,3), then proves the clamp
and the gate: volume-relative LBA 0 maps inside the range and reads; a read
reaching past the range is refused; geometry reports the confined size; and
a confined caller can no longer call define_range (no widening, no escape).
It gates on argv so the ramdisk sweep leaves it silent in other boots, and
coexists with fat (ranges are per-badge).

Discrimination (verified by reverting usb-storage to pre-clamp f1bdce2~1):
the unconfined read still succeeds but define_range returns ENOSYS, so the
fixture cannot arm confinement and the case fails — exactly the property
the clamp adds. With the clamp: block-range 1/1.
2026-08-09 17:25:43 +01:00
Daniel Samson f1bdce25e0 block: per-sender range confinement — V2a mechanism
The block protocol gains define_range (appended, numbers hold): confine the
process named by `badge` to blocks [base, base+count). usb-storage keeps a
per-badge range table and, in read/write, translates volume-relative LBAs
(base added) and refuses any transfer past the volume end. geometry returns
the confined size, so a filesystem mounts against what it may actually touch.

The security seam (decision 4, settled): the clamp lives at the PROVIDER, so
a channel carries exactly the authority it grants — handing a filesystem the
whole disk plus a base offset would let it reach the neighbouring partition.
The gate: a confined caller may NOT call define_range, so a filesystem cannot
widen its own range or confine anyone; only an unconfined party (the volume
manager, whole-device) may. The volume manager defines a filesystem's range
before handing it the channel, so the ordering holds by construction.

Default (no range for a badge) is the whole device — behaviour-neutral for a
single-volume boot and what the volume manager itself uses to probe
partitions. The range table is declared through bounds.md as a runaway
detector (ours, refuse at limit), not a real-partition cap. Neutral:
fat-mount, usb-storage, iommu-usb-storage green. The discrimination fixture
(a confined process reads past its range and is refused) follows next.
2026-08-09 17:13:26 +01:00
Daniel Samson 7ea54a84e5 file-system: mount-failure diagnostics go to the kernel ring, not std.log
V1 boundary review (4 reviewers converged): the extraction routed mount-
FAILURE logs through std.log where old fat used logging.write. That matters
precisely when the failed mount IS /system/logs — a std.log record would
then have nowhere to land. Restore the direct kernel-ring write for the
failure branch (generic prefix, since the harness is filesystem-agnostic
now); success stays std.log.info as before. Unexercised error path; fat-mount
still green. Everything substantive in the extraction verified neutral.
2026-08-09 17:06:13 +01:00
Daniel Samson d63a008148 file-system: extract the serving harness from fat — V1
fat was one binary doing four jobs; the three that are not FAT-specific move
to library/kernel/file-system-harness, a Server(comptime Engine) generic over
the engine type: the badge-scoped open-node table, the nine vfs handlers, the
not-mounted politeness, the exit sweep, mount registration, and durable-on-
close. A filesystem is now an engine plus a main that hands the harness a
mounted volume; a second engine reuses the harness wholesale.

Placement note: the plan said library/file-system, but the harness is a
specialization of `service` (its sibling) and needs nothing from the device
domain, so it lives beside service in library/kernel and stays block-free —
durability rides a caller closure (Volume.flush), no backwards kernel->device
dependency, no new-domain scaffolding. The engine type is inferred from
resolve()'s return, so engine.zig is untouched (its Node stays module-scope).

fat keeps only its FAT-specific bring-up (acquireVolume, DMA, engine.mount,
the attach/detach round trip) and the three mount prefixes as data. Behavior-
neutral: 13/13 across the fat/vfs/logger/IOMMU surface, nothing observable
changed. This lands first so every later phase touches the harness once.
2026-08-09 16:55:48 +01:00
Daniel Samson b59f981c58 docs: decision 4 settled — a loop with an open decision is not a loop
Per-sender range confinement at the provider, one serving endpoint: the
badge-scoped provider pattern the xHCI bus already uses (the per-client
device-token table), applied to blocks. Endpoint-per-volume would buy the
same enforced property only by inventing a multi-endpoint harness; it stays
available as a future refactor, same wire contract. The plan's flag-for-veto
is gone: nothing in the track waits on a choice.
2026-08-09 16:34:04 +01:00
Daniel Samson 60b41c0e82 kernel: mounts have owners — V0 of the volume-manager plan
fs_unmount was gated by nothing but the /protocol carve-out: any process
could unmount any prefix — latent with one mount owner, an obvious
cross-tenant hole once volumes multiply. Each backend mount now records the
mounting task, and the syscall layer enforces two rules that keep the
restart story intact: only the owner unmounts (a dead owner's mount is
swept lazily by resolution — strangers gain nothing by racing that), and a
mount may be REPLACED only by its live owner or after its owner died (the
respawned-filesystem path; displacement of a live mount would be worse than
unmounting it). Kernel-installed mounts are never displaceable.

The vfs-test park role is the discrimination: with the volume provably
mounted it attempts the foreign unmount, requires the refusal AND the
subtree still resolving, and withholds its "parked" marker otherwise —
against the ungated kernel the unmount was ALLOWED and vfs-client-death
fails; with the gate, green. (Its verification handle closes immediately:
the kernel test string-matches "released 1 handle(s)".)
2026-08-09 16:16:14 +01:00
Daniel Samson 451abba000 docs: the volume-manager plan — V0 unmount ownership through V5 close-out 2026-08-09 16:09:20 +01:00
Daniel Samson 9da63e81e2 docs: the boot volume — identity is what makes yanking it survivable
The sharpest instance of the return story, folded into the architecture: the
boot volume is a recorded content identity (the volume carrying
/system/configuration and /system/logs), the system runs on without it (the
ramdisk is the OS; only log persistence pauses), the kernel ring is the
buffer during absence — bounded, so a wrapped ring is a data loss window
that gets MARKED in the file on resume, never spliced silently — and on
return at any port the same identity remounts the same prefixes and the
logger appends into the same boot-stamp tree. The logger requirement is
named: failed flushes retry on the patient cadence, never abandoned after
the first not_found. The dirty-honesty rule stands; the boot volume gets no
exemption.
2026-08-09 16:01:03 +01:00
Daniel Samson ada251150a docs: volume identity, and volumes.csv as danos's fstab
Decision 8 in the rationale, mirrored into the architecture's volume-manager
section: the mount map keys on CONTENT identity, never port or arrival order
— Linux's /dev/sda1-era fstab broke on every port move until UUID= replaced
it, and danos skips that era. The identity ladder the prober reads off the
medium: GPT partition GUID, filesystem UUID, FAT serial+label, MBR
signature+index, anonymous. Consequences are mechanical: port moves change
nothing (USB, hub, SATA bay, or transport swaps), replug remounts at the
same path, the boot volume is a recorded identity findable anywhere, and
cloned duplicates are a loud policy case instead of silent shadowing.
/volumes/<name> stands as the hierarchy's home for attached media; minting
identifiers (formatting, entropy) stays deliberately out of scope.
2026-08-09 15:59:38 +01:00
Daniel Samson 728b436d0f docs: the storage rationale lives with the architecture it justifies
storage-stack-discussion.md was misplaced at the docs root — that level is
for track plans; this is the file-system domain's design record. Moved to
file-system-development/storage-design-rationale.md, renamed to say what it
is, cross-references updated.
2026-08-09 15:49:12 +01:00
Daniel Samson 092817ba2e docs: transport generality and the media-presence event
The NVMe/SATA assessment folded into the storage discussion: what transfers
untouched (the block contract and everything above it — one driver binary
plus one devices.csv row per transport), NVMe as the shorter stack whose
namespaces are the reserved multi-volume case, AHCI as an open shape choice
(leaning per-port processes, the matrix-proven granularity), and the three
pressure points named honestly: multi-volume is reserved-not-implemented,
synchronous call/reply bottlenecks NVMe until the shm-ring data plane, and
media lifecycle is not device lifecycle.

That last one becomes decision 7 and enters the architecture doc: the
removal path has TWO TRIGGERS, ONE LIFECYCLE — channel death (device leaves)
and a planned pushed medium_changed event (medium leaves, device stays: card
readers and trays, USB ones today), translated by the storage driver from
its transport's native signal, presence never content, consumed by the
volume manager into the same kill-retire-remount path. Without it a swapped
card would be served with the previous card's filesystem state.
2026-08-09 15:36:59 +01:00
Daniel Samson a44b397bed protocols: attach gets its reverse — block detach, usb-transfer dma_detach
The kernel was always symmetric (dma_bind 51 / dma_unbind 52); the two
protocols that forward an attachment up the stack were one-way, so a live
client could grant a device reach into its buffer but never revoke it while
alive — exactly the one-way lifecycle the storage architecture's enforcement
section forbids. Death stays the mechanical backstop; detach is the living
process's path.

Both verbs are appended, so every existing number holds. The shape mirrors
attach precisely: the same region capability rides the cap slot again — the
kernel matches the region, so no layer retains anything between the calls
(the bus never kept the handle; now it never needs to).

fat's bring-up does attach -> detach -> attach, exercising both verbs
through the whole chain (fat -> storage -> bus -> kernel) on every boot: a
broken detach fails every fat case instead of lying dormant until the first
buffer replacement. Honest scope: the round trip proves the plumbing; unbind
semantics are the kernel iommu tests' (map/unmap/translationOf); the full
composition (detach then DMA faults) is a future iommu-fault extension.
2026-08-09 15:18:21 +01:00
Daniel Samson fa54ef6915 docs: the volume lifecycle is enforced, not described
The enforcement section of the storage architecture: the three levers the
device lifecycle already proved, mapped onto volumes — a filesystem can only
be GIVEN its volume (no establishment grants, channel at spawn), the volume
manager supervises with teeth (deadline, kill-on-removal, crash-loop cap),
and the shared harness makes every engine inherit the state machine by
construction (engines never see channels). Kernel backstops: ownership-gated
fs_unmount + the lazy dead-endpoint sweep. Checkable via a lifecycle
conformance drill parameterized over filesystems — supporting a filesystem
MEANS passing it.
2026-08-09 15:13:34 +01:00
Daniel Samson fae616fa5a docs: the storage architecture — layers, boundaries, and who does what when media leaves
The settled shape from the storage-stack discussion, written as the
reference: the data path (vfs -> filesystem service -> block -> driver) as
the application/service/protocol/driver model applied twice; the three kinds
of boundary (protocol between processes, library inside them, control-plane
beside them); the volume manager as the policy home (planned — the FAT
service squats on its duties today, marked as such); adding a filesystem as
engine + shared harness + one configuration row; and the per-layer
responsibility table for removable media — one removal path, kill/retire/
respawn, dirty data lost and SAID to be lost. Indexed from docs/README.md.
2026-08-09 15:08:45 +01:00
Daniel Samson 5a8a2a4d7e docs: the storage-stack discussion — block, volumes, filesystems, against the survey 2026-08-09 14:47:35 +01:00
Daniel Samson d4f8dc51b9 docs: the hot-plug matrix is complete — 124/124 2026-08-09 14:19:39 +01:00
Daniel Samson 36ca3f98b3 test: H4 + H5 — a moved device is a new device, and generations do not leak
The hot-plug matrix's last two cells. H4: unplug from hub port 1.1, replug
on 1.2 — per-port identity means the old child is removed and reaped while
the new port binds a fresh driver; nothing ties a driver to the old port.
H5: three unplug/replug cycles on one port — three reaps, a bind after the
last, proving slots, the bus open table, and the manager's driver entries
are all reusable across generations.
2026-08-09 14:11:05 +01:00
Daniel Samson c3604429d4 test: H3 — a nested hub tree yanked whole rebuilds whole (hot-plug matrix)
hub -> hub -> keyboard, one device_del of the outer hub: the H2 recursion
runs at depth two (the inner hub is torn down as a child, its keyboard
first), the keyboard's driver is reaped, and re-adding all three rebinds.
2026-08-09 14:09:19 +01:00
Daniel Samson 77e7001878 usb: a hub yanked from a root port takes its subtree with it (H2)
The hot-plug matrix's predicted bug, found on first contact: tearDownPort
never recursed into a departing hub's children — only tearDownHubDevice
(a hub leaving one level down) did. Yank a populated hub from a root port
and the downstream slots stayed live against vanished hardware, their class
drivers were never reaped, and the replugged hub found its port still
occupied, so nothing ever re-enumerated: the subtree was gone for the boot.

tearDownPort now recurses children-first, exactly like tearDownHubDevice.
The usb-hub-yank case is the discrimination: one device_del removes a hub
carrying a keyboard AND a mouse, both drivers must be reaped, and the
re-added hub must rebind both — it failed against the unfixed bus and
passes with the recursion.
2026-08-09 14:08:15 +01:00
Daniel Samson a6a3402d92 test: H1 — root-port unplug and replug (hot-plug matrix)
The tearDownPort path, exercised for the first time with a real removal and
return: QEMU raises no root-port change events, so the bus's 250 ms
reconcile tick notices the PORTSC change alone — and it does. Reap, re-add,
rebind, second generation binds on the same port. Discrimination: against
the pre-reap manager (4a2df58~1) the case fails at the dedupe wall.
2026-08-09 14:03:02 +01:00
Daniel Samson d565a6b845 docs: the hot-plug matrix plan — unplug anything, replug anywhere 2026-08-09 13:58:29 +01:00
Daniel Samson 4a2df587bb establishment: unplug reaps like death, so a replug rebinds
The hot-unplug path (onChildRemoved, a report from a live bus) cleared the
child but left the bound class driver: a process blocked on reports that
will never come, whose stale entry made the matcher's dedupe refuse the
respawn when the device was plugged back in — the same wall the restart
zombie hit, one path over. Unplug now reaps exactly like reporter death.

The usb-hub-unplug case grows the replug: device_del the hub keyboard, then
device_add it back (qmp_sequence); the ordered tail — child removed, reaping,
delegated, ok — can only be satisfied by the second generation, since every
boot keyboard's ok precedes the unplug. Discrimination: without the reap the
replug never rebinds and the case times out (verified by stash run). Hub
family, restart drill, and the two-controller proof all green (8/8).
2026-08-09 13:51:07 +01:00
Daniel Samson 3c9f454398 establishment: the two-controller proof, and the docs catch up
P4 of docs/establishment-planes-plan.md. The new usb-two-controllers case is
the Ryzen mouse bug pinned in the suite: a second xHCI controller with its
own keyboard while the boot controller keeps the default one — both must
come up, on different device ids, which requires each class driver to reach
ITS OWN controller. Discrimination: at 72807c2 (name-based establishment)
the case fails — one keyboard is unreachable, exactly the bench failure —
verified against a checkout of that merge; with lineage routing it passes.
The existing second-controller cases could not prove this: the boot bus
always carries a keyboard, so their expects were satisfiable by it.

Docs updated with the code: device-manager.md (hello moves the channels,
paged enumerate, delegation as built, reap-and-rebuild in the restart
sequence), device-authority.md (the fourth as-built decision: driver-layer
channels ride the hello; hello is no longer only the liveness handshake).

Full suite: 118/119, the one failure being iommu-fault's fixed 200 ms
fault window under end-of-run host load — 3/3 green standalone, deflake
flagged separately. usb-two-controllers passed inside the full run.
2026-08-09 12:48:33 +01:00
Daniel Samson 8710944a92 establishment: a reporter's death reaps its subtree, and the re-report rebuilds it
P3 of docs/establishment-planes-plan.md — the restart-zombie fix. A class
driver cannot observe its provider's death: an HID driver blocks on
interrupt reports that will simply never come, and storage answers its
callers with refusals forever. Worse, the dead generation's still-used
entries made the matcher's dedupe refuse the respawn when the restarted bus
re-reported — the subtree was a permanent zombie, which is the exact
opposite of the restart-a-driver-live goal the driver model exists for.

pruneChildrenOf now reaps: each pruned child's bound driver is killed and
its entry cleared (the exit notification finds no entry, so the death is
never double-counted; its device returns by the loan rule; its stored
endpoint handle is closed). The re-report then spawns a fresh generation
whose hellos fetch the successor's channel.

The usb-report drill now asserts the subtree WORKS after the restart: the
respawned storage opens its device on the NEW bus instance and reads block
0. Discrimination: against the pre-reap manager the drill fails — no reap
line, no post-restart respawn (the survivors were zombies), verified by a
stash run. Note the scenario boots no input service, so the HID drivers of
BOTH generations exit after their input lookup times out — storage is the
functional proof.
2026-08-09 12:34:13 +01:00
Daniel Samson 1d7850239d establishment: block stops being a name, and enumerate learns to page
P2 of docs/establishment-planes-plan.md. usb-storage serves nameless — one
process per stick cannot share an exclusive bind, and a second stick used to
die silently on -EBUSY before ever helloing. Its one hello now moves both
directions at once: the block-serving endpoint up, its controller channel
down. fat finds its volume through the manager — a new `.consumer` role asks
for the channel of the driver BOUND TO a device (distinct from the device's
reporter), found by enumerating the tree for the mass-storage identity.
fat stays single-volume; the boot-volume-by-content choice is M21.

The conversion immediately caught a live truncation of exactly the audit's
shape: ChildEntry grew to 32 bytes, one enumerate reply holds ~7, and a real
tree carries a dozen ACPI nodes before the first USB child — the storage
entry silently never fit (the protocol comment already said "paging joins
the protocol if a tree ever outgrows one packet"). enumerate is now paged:
Header.target is the start cursor, a short page is the end; device-list's
page-0 read is unchanged.

Grant rows move with the code: the block bind and fat's block open die, fat
gains open device-manager. Gate: 18 cases green including the registry
trio, device-list, and both IOMMU storage variants.
2026-08-09 12:24:25 +01:00
Daniel Samson d603d40b5c establishment: usb-transfer stops being a name
P1 of docs/establishment-planes-plan.md. The bus no longer binds
/protocol/usb-transfer — the bind race made whichever instance came second
unreachable, which on a real three-controller Ryzen meant a mouse no class
driver could reach ("could not open device 50"). Each instance hands its
serving endpoint up in the hello that already delegates its controller, and
class drivers receive their OWN controller's channel from helloForChannel —
routed by the manager's lineage, retried while a provider is mid-restart,
on one manager handle so retries never spend handle-table slots.

usb.open(bus, id) now takes the channel it used to look up; the name rows
leave protocol.csv with the code (bind 67, opens 119/120/122); the
conformance fixture's prose stops claiming the bus binds; and the stale
input-client import leaves the bus with the channel one.

Gate: 19 QEMU cases green (usb family, hubs, both IOMMU variants, fat chain,
boot-from-USB, orderly shutdown, conformance).
2026-08-09 11:48:31 +01:00
Daniel Samson 77fe4d220e establishment: the mechanics — reply capabilities, helloExchange, lineage routing
P0 of docs/establishment-planes-plan.md; no behavior changes yet, nothing
sets the new flag or sends a hello capability.

- service.run gains a reply-capability out-slot (replyWithCapability), the
  registry idiom init already uses, lifted into the harness; null stays the
  untouched common path. Subscribers gains claimArrival() so a provider
  handler can keep a turn capability through the same flag the reserved
  subscribe uses.
- Hello wire struct: the padding byte becomes wants_channel — old callers
  wire-compatibly say 0, and the manager nominates a reply capability ONLY
  when asked, because a capability sent to a caller that never reads one is
  a leaked slot in that caller's table.
- driver.helloExchange: one handshake can hand a serving endpoint up and
  receive the device's provider channel down (usb-storage will need both at
  once). No channel in the reply is retryable, never a verdict.
- The manager stores each instance's serving endpoint on its Driver entry,
  routes consumer hellos by lineage (child -> reporter -> endpoint), replaces
  on re-hello, and closes the stale handle on death - the 32-slot table is
  the bound that makes forgetting this boot-fatal.
2026-08-09 11:39:29 +01:00
Daniel Samson 0d5a7394ef docs: establishment planes — the design and the two-seam plan
The namespace holds protocols, never instances; what multiplies is provider
processes, and their establishment routes through the device manager, which
owns the topology. communication.md gets the model; the plan converts the
usb-transfer and block seams in one flag-day, closes the restart-zombie hole
the scoping found, and pins the three-controller mouse bug as a QEMU case.
2026-08-09 11:26:53 +01:00
Daniel Samson 5930c9653c docs: device-authority as built — spawn carries the device, the loan, confinement rebuilt on return 2026-08-09 10:17:46 +01:00
Daniel Samson 981ff7cd25 Merge the audit fixes: unlocked give paths, restart re-confinement, AMD-Vi store bit 2026-08-09 10:03:42 +01:00