530 Commits
Author SHA1 Message Date
Daniel Samson 4037e746aa docs: storage remount-on-replug is bench-pending, not bench-verified
The physical unplug/replug remount was never actually bench-run — the claim was
carried over from S3. QEMU cannot re-present a usb-storage device_add, so this
path is only reachable on real hardware. Correct both docs to say bench-pending,
keeping the honest distinction: the re-adopt+remount CODE PATH is QEMU-proven by
the driver-crash rebuild; only the physical-replug end-to-end awaits a bench pass.
2026-08-10 19:35:04 +01:00
Daniel Samson 4dfb5012c0 docs: the removal lifecycle closes — three triggers, one path (S5)
Storage removal is now robust to all three ways a volume can leave, and the docs
say so. storage-architecture.md and storage-design-rationale.md move medium_changed
from "planned to be consumed" to consumed, and record the third trigger:

  - the DEVICE leaving the tree (a pulled stick)          — presence polling
  - the MEDIUM leaving while its device stays (a reader)  — the volume manager
    now consumes the pushed medium_changed event
  - the storage DRIVER crashing while its device stays    — a channel-liveness
    geometry() probe reaps the volume and rebuilds it on the restarted driver's
    fresh channel; presence polling alone cannot see this (the V4 open edge)

The re-adopt-and-remount path is QEMU-proven by the driver-crash rebuild; a
physical unplug/replug is bench-verified (QEMU cannot re-present a usb-storage
device_add). The transport-native eject signal (SCSI UNIT ATTENTION, AHCI
PxSSTS, NVMe namespace-change AER) in place of the TEST UNIT READY poll stays
the documented future refinement.
2026-08-10 06:08:56 +01:00
Daniel Samson 061eb7c004 volume-manager: drop the medium-event dedup — it only ever misfired (S5 review)
The adversarial S5 review found a real, unrecoverable defect: onMediumEvent
deduped medium_changed events on a module-global last_medium_change compared by
equality against the event's change_count. But change_count is a PER-DRIVER
counter that restarts at 0 in every usb-storage instance — it is bumped +%=1 and
published only on a real medium transition, so within one instance every count
is unique and monotonic and an equality dedup can never legitimately fire.

The global was carried across a driver restart — S5's OWN crash-rebuild path —
so a fresh instance's first eject (count=1) collided with a stale last==1 and was
dropped. removeDevice never ran; the filesystem kept serving I/O against absent
media forever, and nothing else recovered it: onGeometry answers from a cached
block_count so channelAlive stays true, and isDevicePresent stays true (the
device never left the tree). medium_changed is the sole eject oracle there.

The dedup guarded a re-delivery the driver already makes impossible, and its
only observable effect was this bug. Remove it: react to each present-edge
directly. Both branches are idempotent (a freed device stops matching dev.used)
and the poll reconciles, so acting on every genuine edge is safe — and a fresh
driver instance's counter can no longer be mistaken for the previous one's.

Verified: build + bounds green; volume-removal, volume-medium-change and
volume-driver-restart all pass, so both removal paths survive the change. The
driver-restart-then-eject intersection that triggered the collision cannot be
staged in QEMU — the internal driver-kill cannot be ordered against a QMP eject,
and a usb-storage device_add is not re-presented — so the guarantee rests on the
driver's one-publish-per-transition-with-unique-count contract.
2026-08-10 05:56:06 +01:00
Daniel Samson 5dc966838a volume-manager: rebuild a volume when its storage driver dies (S5)
The V4 review's open edge: a storage driver that crashes while its device
stays in the tree left fat wedged on a dead channel — device-presence
polling (a device-manager enumerate) still reported the device present,
so nothing reaped it. pollTick now also probes channelAlive(dev), a
geometry() on the block channel that fails fast on the dead endpoint; a
present device with a dead channel is reaped like a pull, and the adopt
loop re-adopts it on the restarted driver's fresh channel — the rebuild.
The manager's own liveness probe makes fat self-detection unnecessary:
it rebuilds regardless of the wedged filesystem's state.

The drill: the device manager gains a test-storage-restart mode that
kills usb-storage once, ~2s after its hello (post-mount); a new
volume-driver-restart kernel case boots a manual tree with it, and the
QEMU case asserts a SECOND mount of the same id-path after the reap —
the rebuild. A pre-S5 manager, checking only device presence, never
reaps, so the second mount never appears. The manually-spawned volume
manager needed kernel-supervisor protocol grants (bind its name, open
the device manager), as the other manual-tree services already have.

Full suite 133/133.
2026-08-10 05:21:48 +01:00
Daniel Samson 700452dc4e volume-manager: consume medium_changed — the second removal trigger (S5)
The block client gains subscribeMedium / unsubscribeMedium /
decodeMediumChanged (the reserved subscribe/unsubscribe verbs plus the
MediumChanged decode), so a consumer never hand-rolls the wire format.
The volume manager subscribes to each device it adopts and consumes the
event through the new on_buffered_message seam — never the protocol
dispatch, whose op numbers collide with the manager's own hello.

A medium leaving while its device stays in the tree (a card reader, an
eject) now runs the SAME kill-retire-remount path as a pulled stick:
absent retires the volume, present re-probes it. That closes the "two
triggers, one lifecycle" the architecture specifies — device-presence
polling alone could never see a medium leave under a present device.
dropDevice unsubscribes before closing so the driver's bounded
subscriber table frees the slot; on a dead channel (a real pull) the
call fails fast, proven by volume-removal still passing.

New volume-medium-change case: eject the medium (not the device) ->
"medium absent" -> the manager unmounts. Fails against a pre-S5 manager
that never subscribed. Suite 132/132 (device-authority is the known
child-cleanup flake, green on rerun).
2026-08-10 04:58:54 +01:00
Daniel Samson 0faa0fd21b service: on_buffered_message for pushed events (S5)
An additive, behavior-neutral callback. A buffered async message
(Received.isMessage — a pushed event from a provider this service
subscribed to) carries a payload in the receive buffer; run() now hands
it to on_buffered_message before falling through to on_notification with
the badge, so a coalesced timer/exit riding the same wake is not lost. A
service that does not set the callback (all of them today) is unchanged:
the isMessage branch is a no-op and on_notification still runs, exactly
as before.

This is the seam the volume manager needs to consume block
medium_changed: the event's reserved op number collides with the volume
manager's own hello, so it must be decoded by hand here, never through
the protocol dispatch. Full suite neutral (the lone device-authority
miss is a known child-cleanup race that passes on rerun).
2026-08-10 04:43:27 +01:00
Daniel Samson c47215821c docs: the second engine is built — exFAT + proven harness reuse (S4)
Flip the storage docs to record exFAT as a built second engine reusing
the shared filesystem harness wholesale (the reuse the architecture
promised): full read + write, directories, rename, on-disk up-case
folding, routed by VBR to fat or exfat at an exfat-<serial> id-path.
Note the two surface limits shared by BOTH engines as vfs-layer concerns,
not exFAT shortcuts: u32 file offsets (a 4 GiB cap) and ASCII-only names.
2026-08-10 04:15:03 +01:00
Daniel Samson 6ddb08091d exfat: adversarial-review fixes — overflow safety, sparse gaps, dir size, big-image bitmap (S4 step 9)
A 7-dimension adversarial review of the engine, tool, and routing found
nine real defects (host tests + the in-VM drill missed them). Fixed:

- geometryOf now rejects a crafted VBR whose cluster shift exceeds the
  exFAT ceiling (bytes+sectors shift > 25) or whose cluster_count exceeds
  the spec max (0xFFFFFFF5) — either would overflow the engine's u32
  cluster-byte / cluster-bounds arithmetic and panic under ReleaseSafe on
  untrusted removable media. validCluster/allocateCluster widened to u64,
  and writeFile's clusters_needed widened, for a >4 GiB file near the u32
  offset boundary.
- writeFile no longer claims valid_data_length = size unconditionally: a
  sparse write past a foreign file's old valid boundary now zero-fills the
  skipped gap on disk, so a read there returns zero, not stale bytes.
- ensureDirCapacity rewrites a grown subdirectory's own DataLength, so a
  spec-compliant reader that bounds a directory by DataLength sees the new
  entries (danos itself bounds by the end marker, but chkdsk / other OSes
  do not).
- make-exfat-image lays the allocation bitmap across as many clusters as
  it needs; a >128 MiB image (whose bitmap exceeds one cluster) was
  self-inconsistent. Verified: the engine mounts+reads both the 48 MiB
  fixture and a 256 MiB image.

Documented (not fixed here — a shared vfs-layer limit, like the u32
offset cap): non-ASCII names fold to '?', the same as the FAT engine.

New host tests pin each fix (crafted-VBR rejection, sparse-gap zero,
subdir-grows-and-records-size). Full suite 131/131, bounds green.
2026-08-10 04:14:54 +01:00
Daniel Samson 77b64229c2 exfat: the in-VM mount + mutation drill (S4 step 8)
exfat-volume attaches a bare exFAT data device (serial e0fa0001, from
make-exfat-image.py) beside the FAT boot volume. The volume manager
content-routes it to the exFAT service — not fat — which mounts it at
its id-path /volumes/exfat-e0fa0001; the exfat-test client then reads the
seeded HELLO.TXT and mutates through the mount (mkdir/write/rename/read/
remove). The second engine reusing the shared harness is now proven end
to end, on a real device, in QEMU.

The drill caught what host tests could not: the exfat service was denied
openEndpoint("volume-manager") — it had no protocol.csv grant — so its
hello never reached the manager and its volume never mounted (a silent
spin, no fault). Added the grant mirroring fat's. A new exfatVolumeTest
kernel case boots the tree and spawns the fixture; the run harness grows
an exfat data-volume flavor. Fails against pre-S4 (no exfat binary, csv
row, VBR recognizer, or grant).
2026-08-10 03:44:29 +01:00
Daniel Samson 90906bcefe exfat: cross-engine discrimination + the exfat-test fixture (S4 step 7)
Each engine now refuses the other's volume at mount(): the exFAT engine
rejects a FAT boot sector (a non-zero byte where exFAT keeps MustBeZero),
and the FAT engine rejects an exFAT one (MustBeZero reads as a zero
bytes-per-sector), each mounting its own as a control. The two engines
can never claim the same medium.

exfat-test is the mount round-trip client, cloned from fat-test: it waits
for /volumes/exfat-e0fa0001, reads the seeded HELLO.TXT, then exercises
mkdir + write + rename + read-back + remove through the VFS and reports
"exfat-test: ok". It ships as a lazy fixture in test builds (the
exfat-volume QEMU case, step 8, drives it). build + zig build test +
bounds green.
2026-08-10 03:32:55 +01:00
Daniel Samson 2c2745e9e5 volume-manager: recognize exFAT and route it by content (S4 step 6)
partition.zig gains an exFAT VBR recognizer: the "EXFAT   " name (where a
FAT BPB keeps its OEM string, so the two never collide) plus the 0x55AA
signature, with VolumeSerialNumber (offset 100) as a new exfat_serial
identity rung. A `recognize` helper tries exFAT, then FAT, then the MBR
disk-signature fallback, and sets each volume's FilesystemKind — so
allVolumes/firstVolume and the GPT path all tag a volume with the engine
its content needs. volume-map renders exfat-<serial> as the id-path, and
filesystems.csv adds the exfat -> /system/services/exfat row: an exFAT
stick now spawns the exFAT service, at its own content id-path.

Host tests: a bare exFAT volume recognized with its serial; a FAT VBR
still recognized as fat (the discrimination); the exfat-<serial> id
render. build + zig build test + bounds green.
2026-08-10 03:28:08 +01:00
Daniel Samson 2a6d604577 tools: make-exfat-image.py — a real exFAT image builder (S4 step 5)
Pure stdlib, no mkfs.exfat: writes a Main Boot Sector + its boot-region
checksum + a backup region, the 32-bit FAT, an allocation bitmap, an
up-case table (with its checksum), and a root directory carrying the
bitmap/up-case/label entries plus a seeded HELLO.TXT file set. --serial
sets the volume serial (the content id-path); --verify re-reads the VBR,
recomputes the boot checksum, and validates each file set's checksum.

Cross-checked against the engine: the Zig exFAT engine mounts a
Python-produced image and reads HELLO.TXT through a case-insensitive
resolve ("exfat hello danos"), proving the tool and engine agree on the
layout — VBR, geometry, bitmap, up-case, and the entry-set checksum. No
image is committed; the drill (step 8) generates one per run.
2026-08-10 03:22:19 +01:00
Daniel Samson e81a4e6f1d exfat: the service + build wiring (S4 step 4)
exfat.zig is a near-clone of fat.zig — the reuse the architecture
promised, now real: same IpcBlock DMA-bounce wrapper, same
acquire-volume-by-hello to the volume manager, same content-conditional
boot rewrites (resolve /system/configuration to decide the system
volume), all differences confined to the engine it wraps. That the
harness's Server(engine.FileSystem) compiles is the proof the exFAT
engine meets the same pub-fn contract as FAT — a comptime check, not a
hope.

The exfat package (build.zig + build.zig.zon) mirrors fat's: it ships in
production_ship, its on-disk + engine host tests run through the root
test aggregate (replacing the temporary standalone entry), and the root
zon declares it. build + zig build test + bounds all green.
2026-08-10 03:18:00 +01:00
Daniel Samson e240341bfb exfat: the engine write path (S4 step 3)
create / write / truncate / remove / mkdir / rename, completing the
engine's pub-fn contract with the shared harness. The allocation BITMAP
is the authority: setAllocated IS the allocation (a set bit), and the
32-bit FAT only records a fragmented chain's order — forgetting the bit
would hand a live cluster out twice, so the write path never touches the
FAT without also owning the bit.

This engine writes FAT-linked (no_fat_chain=0) files: createFile makes
an empty set; writeFile allocates+links+zeroes clusters (so a sparse gap
reads zero and valid_data_length can honestly equal data_length) and
rewrites the Stream entry; a contiguous file opened for growth is first
threaded through the FAT. truncate frees the chain; removeFile clears
each set entry's InUse bit and frees the chain (a non-empty directory is
refused); rename re-homes the same clusters under a new name set.
createDirectory allocates one zeroed cluster — exFAT directories carry no
"." / ".." entries. Every entry-set mutation recomputes the set
checksum. Offsets clamp to the vfs u32 surface.

Ten engine host tests now (read + write across clusters, truncate+reuse,
remove, subdir+inner file, rename); 17 total with on-disk. bounds green.
2026-08-10 03:11:05 +01:00
Daniel Samson 62eb2a748a exfat: the engine read path (S4 step 2)
mount + resolve + list + read over a BlockDevice, host-tested against a
RAM-backed image the tests build with formatExfat. mount reads the VBR,
loads the geometry, and scans the root for the Allocation Bitmap (0x81)
and Up-case Table (0x82); the up-case prefix is decompressed (0xFFFF
identity runs) into a bounded table so names fold correctly.

A directory is read as consecutive 32-byte entries via readChain, so a
File/Stream/Name SET that straddles a sector or cluster boundary
assembles cleanly; each set's checksum is validated before it counts as
a file. A stream's no_fat_chain flag picks contiguous-arithmetic vs
FAT-follow cluster walking. reads honor valid_data_length (allocated-
but-unwritten tail reads as zero) and clamp data_length to the vfs u32
offset surface.

Tests cover mount + up-case fold, listing (skipping the metadata
entries), case-insensitive resolve, and reads across a cluster boundary
on both a contiguous and a FAT-fragmented file. Wired via engine.zig
(which imports on-disk.zig) into the host-test aggregate. 12/12, bounds
green. Write path is step 3.
2026-08-10 03:02:58 +01:00
Daniel Samson 56bd2e7678 exfat: the on-disk layout (S4 step 1)
The pure, host-testable byte layer of the second engine: the Main Boot
Sector (VBR) and the six 32-byte directory-entry types — Allocation
Bitmap, Up-case Table, Volume Label, File, Stream Extension, File Name —
as align(1) extern structs, plus the three exFAT checksums (boot region,
up-case table, directory-entry set), the name hash, and the packed
timestamp <-> Unix-epoch conversion.

geometryOf accepts only "EXFAT   " + 0xAA55 + an all-zero MustBeZero
region; that last guard is the mutual exclusion with FAT — a FAT prober
reads a zero bytes-per-sector there and rejects the volume, and this one
rejects a FAT boot sector for want of the exFAT name. Wire-format widths
are named consts (spec facts, no bare literals) so the bounds gate stays
green; the layout is pinned by @offsetOf/@sizeOf tests.

Wired into the host-test aggregate directly for now; it moves into the
exfat package's own test step when that lands (step 4). 7/7 host tests,
bounds green.
2026-08-10 02:46:23 +01:00
Daniel Samson 0b25cd2c94 docs: multi-volume is built — storage architecture + rationale (S3)
Flip the storage docs from "one FAT volume today" to the built state:
the manager adopts every device, probes each device's whole partition
table, and spawns one range-confined FAT per volume — several volumes
across several devices, or several partitions on one device's channel.
The boot volume is identified by content; a filesystem installs the
/system rewrites only when it resolves /system/configuration on its own
media. No filesystem binds a shared service name any more — clients route
through the kernel mount table.

Correct two now-stale claims in the rationale, in the honest direction:
"multi-volume providers are reserved, not implemented" becomes built
(multi-namespace-per-provider is the untried NVMe case); and the cloned-
duplicate "second mounts suffixed" was never built — S3 mounts both,
they collide on the shared content id-path (last wins), each boot claim
logged, and distinguishing them is S4 arbitration.
2026-08-10 02:35:43 +01:00
Daniel Samson d59279422e test: partitioned-image tool + shared-channel multi-volume proof (S3)
The two-volumes case proves multiple DEVICES; this proves multiple
volumes on ONE device. make-partitioned-image.py lays several FAT32
partitions (each from make-fat-image) behind a classic MBR; the new
partitioned-volume case attaches one such disk (two partitions,
da7a0001 at lba 2048, da7a0002 at lba 83968) as a single usb-storage
device.

partition.allVolumes walks the table and the manager spawns a confined
fat per partition on the SAME block channel, each clamped to its own LBA
range by usb-storage's per-badge range table — so partition B's fat
cannot read partition A's blocks. The case asserts two mount lines at
two distinct non-zero base_lbas on one device; against the pre-uncap
allVolumes (S3 step 2, capped to one partition) only da7a0001 mounts.

No image is committed — the disk is generated per run. Full suite
130/130 (128 + two-volumes + partitioned-volume).
2026-08-10 02:35:32 +01:00
Daniel Samson bf9f8560c6 fat/harness: filesystems coexist without the shared vfs name; two-volume proof (S3)
A second usb-storage device (a generated data volume, serial da7a0001,
an empty FAT with no /system) plugged in beside the boot volume: the
volume manager adopts both devices and spawns a confined fat per volume,
each mounted at its own content id-path.

The test surfaced a real coexistence bug. Every filesystem bound the
single "vfs" contract name under /protocol; the second volume's fat lost
the race, service.run refused-and-exited on the held name, and that
volume never mounted. Clients don't reach filesystems by that name —
fs_resolve routes a path to its backing endpoint through the kernel
mount table by prefix — and nothing consumes "vfs", so the fix is to
bind no shared name: the harness's service_name now defaults to null.
This is the "this fades" the harness comment anticipated for the
volume-manager era; a filesystem's endpoint still serves as its mount
backend without a name.

fat logs "is a data volume" for the non-system branch so the test can
positively assert content-based detection. make-fat-image gains
--serial/--label (default unchanged) so a second image gets a distinct
id-path; the data image is generated per run, never committed. The case
fails against the pre-fix harness (the data volume's fat exits on the
refused bind) — toggle-demonstrated.

Full suite 129/129 (128 + two-volumes); the single-volume path is
unaffected by dropping the vestigial name bind.
2026-08-10 02:10:06 +01:00
Daniel Samson da7dcce64e kernel/vfs: raise the mount ceiling for N volumes; refuse (not drop) a full table (S3)
Multi-volume makes the mount table the bottleneck: each volume installs
one id-path mount and the system volume two FHS rewrites, so at the
volume manager's maximum_volumes (16) the old cap of 8 is far too low.
Raise maximum_mounts to 32 (headroom over the ~20-mount worst case) and
declare its bounds block; drop it from the bounds allowlist.

Fix a latent bug the higher pressure would expose: installMount silently
dropped a mount when the table was full, and mountBackend returned true
anyway — a full table was reported as a successful mount. installMount
now returns whether it placed the mount, and mountBackend propagates a
false so the mounting filesystem's harness logs "could not mount
<prefix>". At-limit is now a refusal that is observed, not a silent
success. (The full-table path has no host unit test: vfs.zig's tests are
not wired into the host aggregate — its import graph reaches the
freestanding kernel — so the correction rests on the propagated return
and the truthful bounds block.)
2026-08-10 01:52:47 +01:00
Daniel Samson 9750db14da fat: install the FHS boot rewrites only on the system volume (S3)
A volume backs /system/configuration and /system/logs only when it
actually carries the /system tree — decided by content (resolve
/system/configuration on its own media at mount), not by spawn order.
The boot volume takes the branch and installs the two rewrites; a data
volume resolves null, mounts only at its id-path, and never shadows the
running system's config or logs with a dead mount.

This retires the "resolve-at-bring-up quirk" the single-volume step
deferred around: that diagnosis was wrong. A boot probe confirmed
resolve() works the instant mount() returns — /system, /system/
configuration, /system/kernel, /system/services all resolve at bring-up
(mount reads LBA 0 through the same block path, so a directory read
cannot fail where the boot-sector read succeeded). No deferral needed.

Behavior-preserving on the single boot volume (it carries the system
tree, so it still installs all three mounts): suite stays 128/128.
2026-08-10 01:37:32 +01:00
Daniel Samson b2a5a0a3c6 volume-manager: adopt every device, a filesystem per partition (S3)
Lift the one-device/one-volume cap. bringUpVolume now probes the whole
partition table (allVolumes uncapped) and spawns a confined filesystem
per volume; pollTick loops it to adopt every present, not-yet-adopted
device each tick.

The subtlety is adopt-once-and-keep: a device is recorded in the table
the first time it is seen and kept until it leaves the tree, even when it
carries no servable volume or its geometry cannot be read. Dropping an
unservable device would make openAnyStorage hand back the same one every
tick and starve the devices behind it; keeping it lets the scan advance
past it. A genuine removal frees the slot; a re-insert (fresh device id)
is probed anew.

The boot image is a single bare-FAT volume, so the full suite is
unchanged at 128/128.
2026-08-10 01:26:47 +01:00
Daniel Samson 7efe7b72d8 volume-manager: N-volume device+volume tables, behavior-preserving (S3)
Replace the single `var volume: ?Volume` and file-global supervision
state with two fixed tables: devices[maximum_devices] owning each adopted
block channel once, and volumes[maximum_volumes] each carrying its own
identity, id, mount prefix, and supervision fields (restarts, spawn_ns,
failed, restart_pending, restart_due_ns). A monotonic next_volume_id
never reuses ids, so a stale hello can't address the wrong child.

Lookups (deviceById, volumeById, volumeByPid, firstUsedVolume) and
claims (claimDevice, claimVolumeIndex) replace the ad-hoc singletons.
pollTick reconciles devices first (removeDevice drops their volumes),
then per-volume restarts, then idle bring-up.

This step stays one-device/one-volume on purpose: bringUpVolume adopts
the first device and caps allVolumes to a single partition, so behavior
is identical and the full suite stays 128/128. Uncapping and adopt-all
land next.
2026-08-10 01:13:52 +01:00
Daniel Samson d4b544d66b volume-manager: partition.allVolumes — every partition, not just the first (S3)
The multi-volume enabler. allVolumes(reader, device_blocks, out) appends every
volume on the device to the caller's buffer and returns the count: GPT
enumerates all valid entries (gptFirstVolume becomes gptAllVolumes), the MBR walk
collects all fitting partitions, and a bare FAT is the single whole-device volume
— each with the same per-entry overflow-safe range validation (the confinement
invariant the driver's clamp rests on) and fatIdentity-over-disk-signature
preference. firstVolume is now the one-element case of allVolumes, so the S1
behavior and its ten tests are unchanged. New host test: a two-partition MBR
yields two volumes with distinct identities (index 0 vs 1); it FAILS when
allVolumes is capped to one (the old firstVolume semantics), passes at 11/11.
2026-08-10 00:57:56 +01:00
Daniel Samson 6d4992ae02 docs: storage — the mount map is built; the path is the id (S2)
Flip the storage docs to match S2: filesystems.csv (signature -> binary) and
volumes.csv (identity -> optional override) are built; a volume's mount path IS
its content id (/volumes/<id>), never a port name; the label is display metadata
a `volumes` query returns. fat receives its mount path via argv[2] rather than
hardcoding it. Still pending: rung 2 (filesystem UUID, needs a non-FAT engine),
multi-volume (fat's boot rewrites stay unconditional until S3), medium_changed
consumption, remount bench-verification.
2026-08-10 00:49:32 +01:00
Daniel Samson df61693065 fat: mount at the id-path from argv; migrate /volumes/usb -> id-path (S2)
The flip that makes the mount path the volume's content id. fat retires its
hardcoded fat_mounts: it reads its mount path from argv[2] (the volume manager
hands it the id-path, e.g. /volumes/fat-12345678, from the FAT serial), mounts
its volume root there, and installs the /system/configuration + /system/logs FHS
rewrites so /system/logs persistence stays decoupled from which volume backs it.
The rewrites are unconditional this increment (the single volume IS the boot
volume); S3 makes them content-conditional across N volumes. Every /volumes/usb
reference migrates to /volumes/fat-12345678 in one commit — the fat-test,
badge-scope-test, and vfs-test fixtures and the four QEMU regexes — plus a new
volume-identity-name case asserting the id-path mount and the /system/logs
rewrite. Discrimination: the regexes now require /volumes/fat-12345678, which the
old hardcoded fat never emitted (it mounted /volumes/usb). Full suite 128/128.
2026-08-10 00:47:19 +01:00
Daniel Samson f1e79d0eeb volume-manager: the volumes query verb — read a volume's id, path, and label (S2)
The mechanism the id/label split needs: a `volumes` verb whose reply packs the
mounted volume's {id, mount_path, label} into the tail (VolumeInfo.encode/decode
— three length-prefixed strings). Software keys on the id (the mount path is
/volumes/<id>); a shell or file manager shows the label — the database id/name
split made a query. The VM's onVolumes answers from the mounted volume, empty
reply if none. Two host round-trip tests (encode/decode; too-small buffer and
short-tail rejection). No runtime consumer yet — the first is a userspace shell;
the hello handshake is unaffected (fat-mount/volume-probe green).
2026-08-10 00:17:22 +01:00
Daniel Samson 167e9c7a9e volume-manager: load the mount map; pick binary by signature, compose the id-path (S2)
The volume manager reads its policy from configuration at boot (loadTables,
mirroring the device-manager registry load): filesystems.csv (content signature
-> service binary) and volumes.csv (optional id -> mount-prefix override), each
held in a static source buffer with declared bounds. On probe it picks the
binary from the volume's signature (unserved + logged if no row matches, like an
unbound device) and composes the mount path — a volumes.csv override, else the
default /volumes/<id> from volume-map.idString — then spawns that binary with
argv {volume-id, mount-prefix}. Behavior-preserving: fat still ignores argv[2..]
and uses its hardcoded mounts, the binary resolves to /system/services/fat, so
the FULL suite stays green (127/127); the flip to argv-driven mounts and the
/volumes/usb -> id-path migration land in step 5.
2026-08-10 00:12:33 +01:00
Daniel Samson 5bfdb75e12 volume-manager: the id-path deriver + volumes.csv override (S2)
NEW volume-map.zig: idString(identity) renders a volume's content identity into
its stable mount id-string — gpt-<32hex>, fat-<8hex>, mbr-<sig>-<index> — the
token whose default mount path is /volumes/<id>, so the path IS the id and never
a port or a label; two volumes that share a label get distinct ids
automatically. parse() reads volumes.csv (id, mount_prefix) into OPTIONAL
overrides; overrideFor returns a pinned prefix or null (the volume takes its
default /volumes/<id>). id_maximum is a declared bound; the fixture sizes are
named. Three host tests (each rung's id token; override hit/miss; malformed rows
counted), wired into the VM package test step with csv. Not yet consumed by the
binary — that lands when the VM loads the tables and composes paths (step 4).
2026-08-09 23:55:08 +01:00
Daniel Samson 3fb8a9b96f volume-manager: content signature + the filesystems.csv map (S2)
partition.Volume gains a FilesystemKind signature (today .fat for every probed
volume; S4 adds a real VBR recognizer for exFAT) — the seam filesystems.csv keys
on to choose a service binary. New filesystem-map.zig parses
`/system/configuration/filesystems.csv` (signature, binary) into rules and
match()es a signature to its binary, mirroring the device registry: a signature
no row matches goes unserved, never guessed; slices point into the source
buffer. Three host tests (fat->binary, the binary is data-driven not hardcoded,
malformed rows counted); the test rule buffer is a named fixture size so the
bounds gate stays quiet. The VM's build gains the csv dependency and wires the
filesystem-map test into its package test step. Not yet consumed by the binary —
that lands when the VM loads the tables (step 4).
2026-08-09 23:50:15 +01:00
Daniel Samson ea8ccf65d0 volume-manager: cover the GPT non-128 entry-size offset path (S1 review)
The S1 adversarial boundary review found the off = (i*entry_size) % 512
arithmetic tested only for 128-byte entries. Add a test with 256-byte entries
and the sole valid entry at index 1 (offset 256), exercising the non-zero-offset
path. No code change — the parser was already correct (off is always a multiple
of entry_size >= 128, so off + 128 <= 512); this closes the coverage gap.
2026-08-09 23:37:52 +01:00
Daniel Samson aca3d5855a volume-manager: S1 close-out — partition fixtures in the root aggregate; docs (S1)
Add volume-manager to build.zig's root package-test loop, so `zig build test`
runs the partition parser's nine host tests and a fixture added there can never
be silently skipped. Flip the identity-ladder build status in
storage-design-rationale.md: rungs 1 (GPT partition GUID) and 3 (FAT serial +
label) are built, joining rung 4; rung 2 (filesystem UUID) waits on a non-FAT
engine; the volumes.csv map and the id-derived mount path land with S2.
2026-08-09 23:26:12 +01:00
Daniel Samson 5637d0e5fc device-manager: the entries-per-reply test asserts 7 (the current shape), not 10
A pre-existing stale assertion, surfaced by wiring volume-manager into the root
test aggregate (its partition fixtures now run under `zig build test`).
entries_per_reply is computed as (packet_maximum 256 - prefix 16) / sizeof
ChildEntry 32 = 7; the test asserted 10, the value the old count-header layout
carried. No behavior change — the enumerate producer and consumers already page
by the real capacity; only the test documented an obsolete number.
2026-08-09 23:26:12 +01:00
Daniel Samson 48ab12e262 volume-manager: the FAT volume serial + label is the identity (rung 3) (S1)
Rung 3, stronger than the MBR disk signature. fatIdentity reads the VBR at the
partition start — 0x55AA plus a 0x28/0x29 extended boot signature; FAT32 iff
fat_size_16 == 0; BS_VolID and BS_VolLab at the FAT12/16 vs FAT32 EBR offsets,
cross-checked against fat/on-disk.zig. firstVolume now prefers it over
mbrIdentity in both the MBR-entry path and the bare-FAT fallback, keeping the
rung-4 id when the VBR is not an extended FAT. The serial becomes the identity
key (the id); the label becomes the display name. Two host tests — a bare FAT32
reports its serial + label; an MBR FAT partition prefers the serial while a
non-FAT partition keeps rung 4 — both FAIL with the preference neutralized (2/9)
and pass with it (9/9). On-image witness: the volume-probe QEMU regex tightens to
the boot image's real serial 0x12345678, which before rung 3 was the ~0x0
pseudo-signature read from VBR offset 440.
2026-08-09 23:22:04 +01:00
Daniel Samson 020e31bc8f volume-manager: GPT parsing — the partition GUID is the id, the name is the label (S1)
Rung 1 of the identity ladder. A protective MBR (a type-0xEE entry) routes
probing to the GPT, authoritatively: gptFirstVolume verifies the LBA-1 header's
'EFI PART' signature and a header CRC-32 (inline reflected poly 0xEDB88320,
shared with the fixtures so parser and tests never drift onto a magic constant),
then walks the entry array — bounded by the declared gpt_entry_scan_maximum —
for the first entry with a non-zero type GUID and an overflow-safe in-device
range. That range check is the confinement-safety guard the driver's clamp
rests on, the invariant firstVolume already enforces for MBR, extended to
untrusted GPT metadata. The unique partition GUID becomes the identity key (the
id / mount-path handle); the 36-char partition name becomes the display label.
Three host tests (GUID-as-id; entry-past-device skipped and an all-out-of-range
table is null; a broken header/CRC is not a volume) — all three FAIL with the
GPT branch neutralized (3/7) and pass with it (7/7). Entry-array CRC deferred
(correctness-only; the range check carries the safety property).
2026-08-09 23:15:23 +01:00
Daniel Samson c81120ef0f volume-manager: partition parser takes a SectorReader; identity is a tagged Identity (S1)
The identity ladder's flag-day — no behavior change. partition.firstVolume stops
taking one preloaded block-0 slice and takes a SectorReader (a read-one-sector
fn), so it can reach GPT metadata at LBA 1 and each partition's VBR on demand
(the next commits). The u64 identity becomes Identity{rung,key,label}: key is the
id (the mount path derives from it), label is display metadata (empty at rung 4).
Identity equality is id-only (rung+key) — the label never enters it. Only rung-4
(MBR sig+index / bare-FAT index 0) is produced, byte-identical to before; the
four host tests port to a RAM-disk reader, and fat-mount/volume-probe/
volume-removal stay green.
2026-08-09 23:06:23 +01:00
Daniel Samson 4f9196c03e docs: storage plan — fix two id-path leftovers (DANOS/label->hex)
Two spots still said /volumes/DANOS and "label->hex", contradicting the settled
path-is-the-id decision. Corrected to id-path.
2026-08-09 22:55:10 +01:00
Daniel Samson eff95416d0 docs: storage plan — path is the id, label is queryable display metadata
Refine the naming decision: a volume's mount path IS its identity id (GPT GUID,
else fat-<serial>/mbr-<sig>-<index>) — a stable, unique, content-derived handle
software uses. The label (FAT volume label / GPT partition name) is mutable
display metadata, NOT in the path, exposed by a volume-manager `volumes` verb
returning {id, mount_path, label} — the database id/name split. This dissolves
the label-collision problem: same-label-different-id volumes get distinct paths
automatically; only identical ids (dd-clones) hit first-wins-and-log. S1 carries
both key (id) and label (display); S2 derives the id-path and adds the query.
2026-08-09 22:54:36 +01:00
Daniel Samson addd264880 docs: storage plan — volumes named by identity; exFAT implemented in full
Two decisions settled: (1) a volume's mount name is its own content identity
(FAT label / GPT name, else identity-hex), never a port name (/volumes/usb) or
role name (/volumes/boot); volumes.csv stays as an explicit override. This makes
S1 precede S2 and folds the /volumes/usb fixture+regex migration into S2. (2)
exFAT is a complete implementation — full read+write, directories, rename, and
the on-disk up-case table — not a read-first/ASCII-only subset; the only limit
is the vfs u32 offset surface (a 4 GiB cap on all filesystems), flagged as a
separate vfs change.
2026-08-09 22:42:56 +01:00
Daniel Samson fca41b351e docs: the storage-stack completion plan (S1-S5)
The phased plan taking the volume-manager track from one FAT volume to any
filesystem, N volumes, content identity, remount, and driver-crash survival:
S1 identity ladder (GPT GUID + FAT serial), S2 volumes.csv/filesystems.csv mount
map, S3 multi-volume, S4 exFAT (the second engine proving the harness reuse),
S5 removal robustness (medium_changed consumption + driver-crash rebuild +
QEMU-verified remount). Code-grounded: real functions, commit-granular steps,
discrimination tests shown to fail against today's behavior, per-phase risks.
Design decisions taken with recommended defaults; two forks flagged for review
(the /volumes/usb transitional naming and the exFAT write scope). Not started.
2026-08-09 22:34:00 +01:00
Daniel Samson bf0595763e docs: flip stale status markers across the tracks (audit found 23)
An all-tracks docs-vs-code audit (the same method that caught the storage
drift) found 23 confirmed inaccuracies where a doc's build-status claim no
longer matches the source — status markers that were never flipped after a
track landed, and a few paths left over from completed flag-days. All verified
against the code before editing; docs only, no behavior change.

The systemic ones:
- IOMMU enforcement (driver-model.md, drivers.md): docs said enforcement was
  not built and "device_claim = ring 0" / "memory-safe is not true yet". It is
  built (per-device VT-d/AMD-Vi domains programmed at device_claim, -ECONFINE
  rollback, dma_alloc buffers bound and torn down at death; fail-open only with
  no IOMMU). Restated; M16 marker flipped to done.
- The FHS flag-day paths: /etc/devices.csv -> /system/configuration/devices.csv
  (devices-csv.md, new-driver-checklist.md, device-manager.md), /var/log ->
  /system/logs (logging.md, new-driver-checklist.md), /mnt/usb -> /volumes/usb
  (process-management.md). Following the old paths silently breaks driver match.
- protocol-namespace P4 "remaining" -> landed (only P5 remains); shared-fate
  fan-out "not yet enforced" -> enforced; wall_clock "not built" -> built;
  SMP affinity + fault-recovery "left" -> built; process_enumerate raw-pointer
  trust model -> checked copyToUser/EFAULT; bounds.md maximum_devices static
  hole -> dynamic per-registrar quota; init spawns fat -> volume-manager;
  config "hardcoded, move to /etc" -> already CSV data files; vdso.md three-
  value call; zig-self-hosting library/ layout; python argv "new" -> built.

Found and fixed by a multi-agent audit across 12 doc clusters, each finding
adversarially verified against the source.
2026-08-09 22:04:17 +01:00
Daniel Samson 8216be991d docs: correct the storage docs' V0-V4 status — the V5 flip missed several
A V5 close-out audit (docs against the actual code) found the storage docs
overclaiming in both directions: the status header was flipped to "built" but
several body markers were not, and two passages describe mechanisms the code
never implemented. Nine confirmed, adversarially verified against the source:

Underclaims (marked Planned, actually built):
- storage-architecture: driver named sub-ranges / range confinement (V2, tested
  by the block-range case); the pushed medium_changed event (published today
  from a TEST UNIT READY poll); ownership-gated fs_unmount (V0, process.zig
  gates it with EPERM); the filesystem-harness extraction (V1). Scoped the
  remaining *planned* to the genuinely-pending parts (native-signal translation,
  the volume manager consuming medium_changed).

Overclaims (described, never built):
- storage-architecture: the "FAT dirty flag on disk" guarantee — no on-disk
  dirty/clean-shutdown bit exists; only an in-memory device-dirty bool gating a
  device write-cache flush on close. Fixed in all three places.
- storage-architecture: fat "acquires its own volume (first mass-storage child
  by enumeration order)" — the V3b flip removed self-acquisition; fat is handed
  its volume id and channel by the volume manager.
- rationale + plan: the FAT engine's base_lba "deleted rather than moved" — it
  and the engine's MBR walk still exist as now-inert legacy; the authoritative
  walk lives in partition.zig.
- rationale: NVMe namespaces "decision 4 settles as endpoint-per-volume" —
  decision 4 settles the opposite (per-sender confinement, one endpoint);
  endpoint-per-volume is named only as an unbuilt future refactor.
- rationale: the five-rung identity ladder and volumes.csv map stated in flat
  present tense — only rung 4 (MBR signature + index) is built; added the
  build-status hedge and marked each rung.

Docs only; no code or behavior change. Suite unaffected (127/127).
2026-08-09 21:19:29 +01:00
Daniel Samson 68e65803eb volume-manager: the removal comment says lazy retirement, not an eager sweep
The V4 adversarial review found removeVolume's comment overclaiming: it said
"the kernel sweeps a dead backend's mounts", which reads as an eager death-time
sweep. There is no such sweep. Killing the filesystem marks its backend endpoint
dead (killOwnedEndpointsLocked), and the VFS router retires each mount that
endpoint backed lazily, on the next path resolution under it (resolvePath sees
the dead backend, frees the slot, returns not_found). The functional guarantee
the comment promised — killing the filesystem retires its mounts — holds; only
the described mechanism was wrong. Comment-only; no behavior change.
2026-08-09 20:56:51 +01:00
Daniel Samson af47d41989 volume-manager: removal supersedes a pending restart in the poll
Inline V4 review (the boundary-review workflow stalled): the poll ran a due
fat-restart before the presence check and returned, so a fat death followed
by a device removal would respawn fat against the now-dead channel and churn
until the crash cap before the removal was noticed. Reorder: check the
specific device's presence first (unmount if gone), and only fire a due
restart once the device is confirmed present. Neutral: fat-mount,
volume-removal, amd-iommu-usb-storage green.

Noted V4 limitations (not fixed here, edge cases outside the user unplug
case): a usb-storage DRIVER crash (device stays, driver restarts with a new
endpoint) leaves fat holding a dead channel — the device is still present so
removal is not detected; fat would need to observe its channel death and
exit. Deferred with the medium_changed subscription and multi-volume.
2026-08-09 20:27:58 +01:00
Daniel Samson 5e89b111cf docs: flip the storage-architecture status markers the V0-V4 track made real
The volume manager, per-sender range confinement + gate, medium_changed
event, per-volume spawning + supervision, ownership-gated fs_unmount, and
the removal half of the lifecycle are built. Left honestly pending: the
volumes.csv/filesystems.csv maps, the fuller identity ladder, multi-volume,
the VM consuming medium_changed (removal uses device-presence polling), and
remount-on-replug end-to-end (bench-pending — QEMU can't re-present the
boot-controller device). Full suite 127/127.
2026-08-09 20:20:20 +01:00
Daniel Samson e3ec9fa668 volume-manager: try every mass-storage entry, watch the one we opened
The full suite caught a V4 regression: under AMD-Vi the device-manager tree
carries more than one mass-storage-identity entry (a phantom no driver is
bound to, which answers a consumer hello with NO channel). V4 split presence
from acquisition and picked the FIRST identity match blindly, so it kept
helloing the phantom (device 27) and never reached the real storage (device
31). V3's inline loop had skipped no-channel entries with `orelse continue`;
the split lost that.

Restore it: openAnyStorage tries each matching entry and takes the first whose
channel opens, recording its device id. Removal detection then watches THAT
specific device id leave the tree (isDevicePresent), not "any mass-storage" —
so a phantom that never leaves cannot mask a real removal. Both are bare
enumerates; the hello only happens while bringing a volume up.

Green: amd-iommu-usb-storage, fat-mount, volume-removal.
2026-08-09 20:08:30 +01:00
Daniel Samson 9e67a74232 volume-manager: the removal lifecycle — a pulled stick unmounts (V4)
The volume manager stops probing-once and polls storage presence for the life
of the boot: findStorageDevice enumerates the device-manager tree (presence
only, no consumer-hello, so it is cheap and leaks nothing). The volume is now
a field that goes null and back — the whole lifecycle:

- storage present + no volume  -> open the channel, probe, confine + spawn the
  filesystem (openStorage is the one consumer-hello, on the insertion edge);
- storage gone + have volume    -> kill the filesystem (its mounts retire via
  the kernel dead-backend sweep), close the dead channel, clear the volume;
- fat crash                     -> the same supervised backoff/cap as before,
  folded into the poll (one timer).

This also subsumes the V3-review leak fix (no per-poll consumer-hello) and the
no-volume retry (a present-but-unreadable device keeps polling).

The user's case — pull the boot stick, plug it back — is a DEVICE unplug (the
stick IS the device), so the mass-storage child leaves the device-manager tree
and the poll catches it. volume-removal asserts the unmount and discriminates:
against the V3 probe-once volume manager the removal is never noticed (0/1).

The re-mount on replug is the VM's bringUpVolume firing when the device
returns — correct and in place, but not QEMU-testable here: device_add of
usb-storage to the boot xHCI controller is not re-presented to the guest (no
port-connect on any port), a harness quirk, not a VM issue. On real hardware
the bus's per-tick port poll catches a reconnect (H1 proves reconnect on a
second controller); bench-verify the full round trip.
2026-08-09 19:41:04 +01:00
Daniel Samson b9058fe020 file-system: the mount-failure log has no scratch buffer (orphaned bounds fix)
The mount-failure diagnostic wrote through a fixed [96]u8 bufPrint buffer,
which the bounds gate flags as an undeclared ceiling. Three ring appends
instead — no buffer, no fixed length to justify. This fix was made during
the V2a bounds work but never git-added, so the committed harness still
carried the flagged buffer; committing it now cleans the tip.
2026-08-09 19:40:39 +01:00
Daniel Samson 7c6ed2ca09 volume-manager: harden the probe and supervision from the V3 review
Five confirmed defects from the boundary review:

1. (security) The VM never checked a partition fit inside the device, so a
   crafted MBR could hand the driver a range whose base+lba wraps past a u32
   — panicking usb-storage in a loop, and at multi-volume overlapping a
   neighbour. This is the exact invariant the clamp's overflow-safety rests
   on. partition.firstVolume now skips any entry that runs past the device
   (host-tested), establishing the invariant where the untrusted bytes are
   first read.
2. (leak) The probe re-acquired a fresh block channel on every 500 ms retry,
   leaking a handle each time on a medium-absent device. The channel is now
   acquired once and kept.
3. (wedge) A failed spawn or defineRange stranded the volume with no retry;
   both now arm a backoff restart.
4. (loop) fat respawn had no exit-reason gate, no backoff, no crash-loop cap
   — a faulting filesystem respawned in a zero-delay loop, and a clean exit
   was resurrected. Supervision now mirrors the device manager: a clean exit
   is not restarted, a fault backs off, three fast deaths give up.
5. (removable) A device that parsed to no volume was terminal; it now keeps
   polling so an inserted medium is picked up — the removal-lifecycle trigger.

Known limitation (noted, not fixed here): if the VM itself crashes and init
restarts it, the orphaned fat keeps serving vfs while the new VM spawns a
second fat whose bind is refused — the same "manager restart re-learns the
world" gap the device manager also defers. The old fat keeps storage working.

Neutral: partition unit tests + fat-mount, volume-probe, block-range, logger
all green.
2026-08-09 18:50:01 +01:00
Daniel Samson a67a7015bf docs: record the V3c resequencing — removal lifecycle first, multi-volume follows 2026-08-09 18:30:16 +01:00