Commit Graph
284 Commits
Author SHA1 Message Date
Daniel Samson 62eb2a748a exfat: the engine read path (S4 step 2)
mount + resolve + list + read over a BlockDevice, host-tested against a
RAM-backed image the tests build with formatExfat. mount reads the VBR,
loads the geometry, and scans the root for the Allocation Bitmap (0x81)
and Up-case Table (0x82); the up-case prefix is decompressed (0xFFFF
identity runs) into a bounded table so names fold correctly.

A directory is read as consecutive 32-byte entries via readChain, so a
File/Stream/Name SET that straddles a sector or cluster boundary
assembles cleanly; each set's checksum is validated before it counts as
a file. A stream's no_fat_chain flag picks contiguous-arithmetic vs
FAT-follow cluster walking. reads honor valid_data_length (allocated-
but-unwritten tail reads as zero) and clamp data_length to the vfs u32
offset surface.

Tests cover mount + up-case fold, listing (skipping the metadata
entries), case-insensitive resolve, and reads across a cluster boundary
on both a contiguous and a FAT-fragmented file. Wired via engine.zig
(which imports on-disk.zig) into the host-test aggregate. 12/12, bounds
green. Write path is step 3.
2026-08-10 03:02:58 +01:00
Daniel Samson 56bd2e7678 exfat: the on-disk layout (S4 step 1)
The pure, host-testable byte layer of the second engine: the Main Boot
Sector (VBR) and the six 32-byte directory-entry types — Allocation
Bitmap, Up-case Table, Volume Label, File, Stream Extension, File Name —
as align(1) extern structs, plus the three exFAT checksums (boot region,
up-case table, directory-entry set), the name hash, and the packed
timestamp <-> Unix-epoch conversion.

geometryOf accepts only "EXFAT   " + 0xAA55 + an all-zero MustBeZero
region; that last guard is the mutual exclusion with FAT — a FAT prober
reads a zero bytes-per-sector there and rejects the volume, and this one
rejects a FAT boot sector for want of the exFAT name. Wire-format widths
are named consts (spec facts, no bare literals) so the bounds gate stays
green; the layout is pinned by @offsetOf/@sizeOf tests.

Wired into the host-test aggregate directly for now; it moves into the
exfat package's own test step when that lands (step 4). 7/7 host tests,
bounds green.
2026-08-10 02:46:23 +01:00
Daniel Samson bf9f8560c6 fat/harness: filesystems coexist without the shared vfs name; two-volume proof (S3)
A second usb-storage device (a generated data volume, serial da7a0001,
an empty FAT with no /system) plugged in beside the boot volume: the
volume manager adopts both devices and spawns a confined fat per volume,
each mounted at its own content id-path.

The test surfaced a real coexistence bug. Every filesystem bound the
single "vfs" contract name under /protocol; the second volume's fat lost
the race, service.run refused-and-exited on the held name, and that
volume never mounted. Clients don't reach filesystems by that name —
fs_resolve routes a path to its backing endpoint through the kernel
mount table by prefix — and nothing consumes "vfs", so the fix is to
bind no shared name: the harness's service_name now defaults to null.
This is the "this fades" the harness comment anticipated for the
volume-manager era; a filesystem's endpoint still serves as its mount
backend without a name.

fat logs "is a data volume" for the non-system branch so the test can
positively assert content-based detection. make-fat-image gains
--serial/--label (default unchanged) so a second image gets a distinct
id-path; the data image is generated per run, never committed. The case
fails against the pre-fix harness (the data volume's fat exits on the
refused bind) — toggle-demonstrated.

Full suite 129/129 (128 + two-volumes); the single-volume path is
unaffected by dropping the vestigial name bind.
2026-08-10 02:10:06 +01:00
Daniel Samson da7dcce64e kernel/vfs: raise the mount ceiling for N volumes; refuse (not drop) a full table (S3)
Multi-volume makes the mount table the bottleneck: each volume installs
one id-path mount and the system volume two FHS rewrites, so at the
volume manager's maximum_volumes (16) the old cap of 8 is far too low.
Raise maximum_mounts to 32 (headroom over the ~20-mount worst case) and
declare its bounds block; drop it from the bounds allowlist.

Fix a latent bug the higher pressure would expose: installMount silently
dropped a mount when the table was full, and mountBackend returned true
anyway — a full table was reported as a successful mount. installMount
now returns whether it placed the mount, and mountBackend propagates a
false so the mounting filesystem's harness logs "could not mount
<prefix>". At-limit is now a refusal that is observed, not a silent
success. (The full-table path has no host unit test: vfs.zig's tests are
not wired into the host aggregate — its import graph reaches the
freestanding kernel — so the correction rests on the propagated return
and the truthful bounds block.)
2026-08-10 01:52:47 +01:00
Daniel Samson 9750db14da fat: install the FHS boot rewrites only on the system volume (S3)
A volume backs /system/configuration and /system/logs only when it
actually carries the /system tree — decided by content (resolve
/system/configuration on its own media at mount), not by spawn order.
The boot volume takes the branch and installs the two rewrites; a data
volume resolves null, mounts only at its id-path, and never shadows the
running system's config or logs with a dead mount.

This retires the "resolve-at-bring-up quirk" the single-volume step
deferred around: that diagnosis was wrong. A boot probe confirmed
resolve() works the instant mount() returns — /system, /system/
configuration, /system/kernel, /system/services all resolve at bring-up
(mount reads LBA 0 through the same block path, so a directory read
cannot fail where the boot-sector read succeeded). No deferral needed.

Behavior-preserving on the single boot volume (it carries the system
tree, so it still installs all three mounts): suite stays 128/128.
2026-08-10 01:37:32 +01:00
Daniel Samson b2a5a0a3c6 volume-manager: adopt every device, a filesystem per partition (S3)
Lift the one-device/one-volume cap. bringUpVolume now probes the whole
partition table (allVolumes uncapped) and spawns a confined filesystem
per volume; pollTick loops it to adopt every present, not-yet-adopted
device each tick.

The subtlety is adopt-once-and-keep: a device is recorded in the table
the first time it is seen and kept until it leaves the tree, even when it
carries no servable volume or its geometry cannot be read. Dropping an
unservable device would make openAnyStorage hand back the same one every
tick and starve the devices behind it; keeping it lets the scan advance
past it. A genuine removal frees the slot; a re-insert (fresh device id)
is probed anew.

The boot image is a single bare-FAT volume, so the full suite is
unchanged at 128/128.
2026-08-10 01:26:47 +01:00
Daniel Samson 7efe7b72d8 volume-manager: N-volume device+volume tables, behavior-preserving (S3)
Replace the single `var volume: ?Volume` and file-global supervision
state with two fixed tables: devices[maximum_devices] owning each adopted
block channel once, and volumes[maximum_volumes] each carrying its own
identity, id, mount prefix, and supervision fields (restarts, spawn_ns,
failed, restart_pending, restart_due_ns). A monotonic next_volume_id
never reuses ids, so a stale hello can't address the wrong child.

Lookups (deviceById, volumeById, volumeByPid, firstUsedVolume) and
claims (claimDevice, claimVolumeIndex) replace the ad-hoc singletons.
pollTick reconciles devices first (removeDevice drops their volumes),
then per-volume restarts, then idle bring-up.

This step stays one-device/one-volume on purpose: bringUpVolume adopts
the first device and caps allVolumes to a single partition, so behavior
is identical and the full suite stays 128/128. Uncapping and adopt-all
land next.
2026-08-10 01:13:52 +01:00
Daniel Samson d4b544d66b volume-manager: partition.allVolumes — every partition, not just the first (S3)
The multi-volume enabler. allVolumes(reader, device_blocks, out) appends every
volume on the device to the caller's buffer and returns the count: GPT
enumerates all valid entries (gptFirstVolume becomes gptAllVolumes), the MBR walk
collects all fitting partitions, and a bare FAT is the single whole-device volume
— each with the same per-entry overflow-safe range validation (the confinement
invariant the driver's clamp rests on) and fatIdentity-over-disk-signature
preference. firstVolume is now the one-element case of allVolumes, so the S1
behavior and its ten tests are unchanged. New host test: a two-partition MBR
yields two volumes with distinct identities (index 0 vs 1); it FAILS when
allVolumes is capped to one (the old firstVolume semantics), passes at 11/11.
2026-08-10 00:57:56 +01:00
Daniel Samson df61693065 fat: mount at the id-path from argv; migrate /volumes/usb -> id-path (S2)
The flip that makes the mount path the volume's content id. fat retires its
hardcoded fat_mounts: it reads its mount path from argv[2] (the volume manager
hands it the id-path, e.g. /volumes/fat-12345678, from the FAT serial), mounts
its volume root there, and installs the /system/configuration + /system/logs FHS
rewrites so /system/logs persistence stays decoupled from which volume backs it.
The rewrites are unconditional this increment (the single volume IS the boot
volume); S3 makes them content-conditional across N volumes. Every /volumes/usb
reference migrates to /volumes/fat-12345678 in one commit — the fat-test,
badge-scope-test, and vfs-test fixtures and the four QEMU regexes — plus a new
volume-identity-name case asserting the id-path mount and the /system/logs
rewrite. Discrimination: the regexes now require /volumes/fat-12345678, which the
old hardcoded fat never emitted (it mounted /volumes/usb). Full suite 128/128.
2026-08-10 00:47:19 +01:00
Daniel Samson f1e79d0eeb volume-manager: the volumes query verb — read a volume's id, path, and label (S2)
The mechanism the id/label split needs: a `volumes` verb whose reply packs the
mounted volume's {id, mount_path, label} into the tail (VolumeInfo.encode/decode
— three length-prefixed strings). Software keys on the id (the mount path is
/volumes/<id>); a shell or file manager shows the label — the database id/name
split made a query. The VM's onVolumes answers from the mounted volume, empty
reply if none. Two host round-trip tests (encode/decode; too-small buffer and
short-tail rejection). No runtime consumer yet — the first is a userspace shell;
the hello handshake is unaffected (fat-mount/volume-probe green).
2026-08-10 00:17:22 +01:00
Daniel Samson 167e9c7a9e volume-manager: load the mount map; pick binary by signature, compose the id-path (S2)
The volume manager reads its policy from configuration at boot (loadTables,
mirroring the device-manager registry load): filesystems.csv (content signature
-> service binary) and volumes.csv (optional id -> mount-prefix override), each
held in a static source buffer with declared bounds. On probe it picks the
binary from the volume's signature (unserved + logged if no row matches, like an
unbound device) and composes the mount path — a volumes.csv override, else the
default /volumes/<id> from volume-map.idString — then spawns that binary with
argv {volume-id, mount-prefix}. Behavior-preserving: fat still ignores argv[2..]
and uses its hardcoded mounts, the binary resolves to /system/services/fat, so
the FULL suite stays green (127/127); the flip to argv-driven mounts and the
/volumes/usb -> id-path migration land in step 5.
2026-08-10 00:12:33 +01:00
Daniel Samson 5bfdb75e12 volume-manager: the id-path deriver + volumes.csv override (S2)
NEW volume-map.zig: idString(identity) renders a volume's content identity into
its stable mount id-string — gpt-<32hex>, fat-<8hex>, mbr-<sig>-<index> — the
token whose default mount path is /volumes/<id>, so the path IS the id and never
a port or a label; two volumes that share a label get distinct ids
automatically. parse() reads volumes.csv (id, mount_prefix) into OPTIONAL
overrides; overrideFor returns a pinned prefix or null (the volume takes its
default /volumes/<id>). id_maximum is a declared bound; the fixture sizes are
named. Three host tests (each rung's id token; override hit/miss; malformed rows
counted), wired into the VM package test step with csv. Not yet consumed by the
binary — that lands when the VM loads the tables and composes paths (step 4).
2026-08-09 23:55:08 +01:00
Daniel Samson 3fb8a9b96f volume-manager: content signature + the filesystems.csv map (S2)
partition.Volume gains a FilesystemKind signature (today .fat for every probed
volume; S4 adds a real VBR recognizer for exFAT) — the seam filesystems.csv keys
on to choose a service binary. New filesystem-map.zig parses
`/system/configuration/filesystems.csv` (signature, binary) into rules and
match()es a signature to its binary, mirroring the device registry: a signature
no row matches goes unserved, never guessed; slices point into the source
buffer. Three host tests (fat->binary, the binary is data-driven not hardcoded,
malformed rows counted); the test rule buffer is a named fixture size so the
bounds gate stays quiet. The VM's build gains the csv dependency and wires the
filesystem-map test into its package test step. Not yet consumed by the binary —
that lands when the VM loads the tables (step 4).
2026-08-09 23:50:15 +01:00
Daniel Samson ea8ccf65d0 volume-manager: cover the GPT non-128 entry-size offset path (S1 review)
The S1 adversarial boundary review found the off = (i*entry_size) % 512
arithmetic tested only for 128-byte entries. Add a test with 256-byte entries
and the sole valid entry at index 1 (offset 256), exercising the non-zero-offset
path. No code change — the parser was already correct (off is always a multiple
of entry_size >= 128, so off + 128 <= 512); this closes the coverage gap.
2026-08-09 23:37:52 +01:00
Daniel Samson 48ab12e262 volume-manager: the FAT volume serial + label is the identity (rung 3) (S1)
Rung 3, stronger than the MBR disk signature. fatIdentity reads the VBR at the
partition start — 0x55AA plus a 0x28/0x29 extended boot signature; FAT32 iff
fat_size_16 == 0; BS_VolID and BS_VolLab at the FAT12/16 vs FAT32 EBR offsets,
cross-checked against fat/on-disk.zig. firstVolume now prefers it over
mbrIdentity in both the MBR-entry path and the bare-FAT fallback, keeping the
rung-4 id when the VBR is not an extended FAT. The serial becomes the identity
key (the id); the label becomes the display name. Two host tests — a bare FAT32
reports its serial + label; an MBR FAT partition prefers the serial while a
non-FAT partition keeps rung 4 — both FAIL with the preference neutralized (2/9)
and pass with it (9/9). On-image witness: the volume-probe QEMU regex tightens to
the boot image's real serial 0x12345678, which before rung 3 was the ~0x0
pseudo-signature read from VBR offset 440.
2026-08-09 23:22:04 +01:00
Daniel Samson 020e31bc8f volume-manager: GPT parsing — the partition GUID is the id, the name is the label (S1)
Rung 1 of the identity ladder. A protective MBR (a type-0xEE entry) routes
probing to the GPT, authoritatively: gptFirstVolume verifies the LBA-1 header's
'EFI PART' signature and a header CRC-32 (inline reflected poly 0xEDB88320,
shared with the fixtures so parser and tests never drift onto a magic constant),
then walks the entry array — bounded by the declared gpt_entry_scan_maximum —
for the first entry with a non-zero type GUID and an overflow-safe in-device
range. That range check is the confinement-safety guard the driver's clamp
rests on, the invariant firstVolume already enforces for MBR, extended to
untrusted GPT metadata. The unique partition GUID becomes the identity key (the
id / mount-path handle); the 36-char partition name becomes the display label.
Three host tests (GUID-as-id; entry-past-device skipped and an all-out-of-range
table is null; a broken header/CRC is not a volume) — all three FAIL with the
GPT branch neutralized (3/7) and pass with it (7/7). Entry-array CRC deferred
(correctness-only; the range check carries the safety property).
2026-08-09 23:15:23 +01:00
Daniel Samson c81120ef0f volume-manager: partition parser takes a SectorReader; identity is a tagged Identity (S1)
The identity ladder's flag-day — no behavior change. partition.firstVolume stops
taking one preloaded block-0 slice and takes a SectorReader (a read-one-sector
fn), so it can reach GPT metadata at LBA 1 and each partition's VBR on demand
(the next commits). The u64 identity becomes Identity{rung,key,label}: key is the
id (the mount path derives from it), label is display metadata (empty at rung 4).
Identity equality is id-only (rung+key) — the label never enters it. Only rung-4
(MBR sig+index / bare-FAT index 0) is produced, byte-identical to before; the
four host tests port to a RAM-disk reader, and fat-mount/volume-probe/
volume-removal stay green.
2026-08-09 23:06:23 +01:00
Daniel Samson 68e65803eb volume-manager: the removal comment says lazy retirement, not an eager sweep
The V4 adversarial review found removeVolume's comment overclaiming: it said
"the kernel sweeps a dead backend's mounts", which reads as an eager death-time
sweep. There is no such sweep. Killing the filesystem marks its backend endpoint
dead (killOwnedEndpointsLocked), and the VFS router retires each mount that
endpoint backed lazily, on the next path resolution under it (resolvePath sees
the dead backend, frees the slot, returns not_found). The functional guarantee
the comment promised — killing the filesystem retires its mounts — holds; only
the described mechanism was wrong. Comment-only; no behavior change.
2026-08-09 20:56:51 +01:00
Daniel Samson af47d41989 volume-manager: removal supersedes a pending restart in the poll
Inline V4 review (the boundary-review workflow stalled): the poll ran a due
fat-restart before the presence check and returned, so a fat death followed
by a device removal would respawn fat against the now-dead channel and churn
until the crash cap before the removal was noticed. Reorder: check the
specific device's presence first (unmount if gone), and only fire a due
restart once the device is confirmed present. Neutral: fat-mount,
volume-removal, amd-iommu-usb-storage green.

Noted V4 limitations (not fixed here, edge cases outside the user unplug
case): a usb-storage DRIVER crash (device stays, driver restarts with a new
endpoint) leaves fat holding a dead channel — the device is still present so
removal is not detected; fat would need to observe its channel death and
exit. Deferred with the medium_changed subscription and multi-volume.
2026-08-09 20:27:58 +01:00
Daniel Samson e3ec9fa668 volume-manager: try every mass-storage entry, watch the one we opened
The full suite caught a V4 regression: under AMD-Vi the device-manager tree
carries more than one mass-storage-identity entry (a phantom no driver is
bound to, which answers a consumer hello with NO channel). V4 split presence
from acquisition and picked the FIRST identity match blindly, so it kept
helloing the phantom (device 27) and never reached the real storage (device
31). V3's inline loop had skipped no-channel entries with `orelse continue`;
the split lost that.

Restore it: openAnyStorage tries each matching entry and takes the first whose
channel opens, recording its device id. Removal detection then watches THAT
specific device id leave the tree (isDevicePresent), not "any mass-storage" —
so a phantom that never leaves cannot mask a real removal. Both are bare
enumerates; the hello only happens while bringing a volume up.

Green: amd-iommu-usb-storage, fat-mount, volume-removal.
2026-08-09 20:08:30 +01:00
Daniel Samson 9e67a74232 volume-manager: the removal lifecycle — a pulled stick unmounts (V4)
The volume manager stops probing-once and polls storage presence for the life
of the boot: findStorageDevice enumerates the device-manager tree (presence
only, no consumer-hello, so it is cheap and leaks nothing). The volume is now
a field that goes null and back — the whole lifecycle:

- storage present + no volume  -> open the channel, probe, confine + spawn the
  filesystem (openStorage is the one consumer-hello, on the insertion edge);
- storage gone + have volume    -> kill the filesystem (its mounts retire via
  the kernel dead-backend sweep), close the dead channel, clear the volume;
- fat crash                     -> the same supervised backoff/cap as before,
  folded into the poll (one timer).

This also subsumes the V3-review leak fix (no per-poll consumer-hello) and the
no-volume retry (a present-but-unreadable device keeps polling).

The user's case — pull the boot stick, plug it back — is a DEVICE unplug (the
stick IS the device), so the mass-storage child leaves the device-manager tree
and the poll catches it. volume-removal asserts the unmount and discriminates:
against the V3 probe-once volume manager the removal is never noticed (0/1).

The re-mount on replug is the VM's bringUpVolume firing when the device
returns — correct and in place, but not QEMU-testable here: device_add of
usb-storage to the boot xHCI controller is not re-presented to the guest (no
port-connect on any port), a harness quirk, not a VM issue. On real hardware
the bus's per-tick port poll catches a reconnect (H1 proves reconnect on a
second controller); bench-verify the full round trip.
2026-08-09 19:41:04 +01:00
Daniel Samson 7c6ed2ca09 volume-manager: harden the probe and supervision from the V3 review
Five confirmed defects from the boundary review:

1. (security) The VM never checked a partition fit inside the device, so a
   crafted MBR could hand the driver a range whose base+lba wraps past a u32
   — panicking usb-storage in a loop, and at multi-volume overlapping a
   neighbour. This is the exact invariant the clamp's overflow-safety rests
   on. partition.firstVolume now skips any entry that runs past the device
   (host-tested), establishing the invariant where the untrusted bytes are
   first read.
2. (leak) The probe re-acquired a fresh block channel on every 500 ms retry,
   leaking a handle each time on a medium-absent device. The channel is now
   acquired once and kept.
3. (wedge) A failed spawn or defineRange stranded the volume with no retry;
   both now arm a backoff restart.
4. (loop) fat respawn had no exit-reason gate, no backoff, no crash-loop cap
   — a faulting filesystem respawned in a zero-delay loop, and a clean exit
   was resurrected. Supervision now mirrors the device manager: a clean exit
   is not restarted, a fault backs off, three fast deaths give up.
5. (removable) A device that parsed to no volume was terminal; it now keeps
   polling so an inserted medium is picked up — the removal-lifecycle trigger.

Known limitation (noted, not fixed here): if the VM itself crashes and init
restarts it, the orphaned fat keeps serving vfs while the new VM spawns a
second fat whose bind is refused — the same "manager restart re-learns the
world" gap the device manager also defers. The old fat keeps storage working.

Neutral: partition unit tests + fat-mount, volume-probe, block-range, logger
all green.
2026-08-09 18:50:01 +01:00
Daniel Samson 301bdcaf5b volume-manager: the flip — fat is spawned, confined, and handed its channel (V3b)
The load-bearing step. The FAT service stops acquiring its own volume: the
volume manager spawns it (per volume), defines its partition range on the
storage driver BEFORE it runs, and answers its startup hello with the
range-confined block channel over a new volume-manager protocol. fat never
finds its storage by name and never sees the whole device — establishment
by lineage, one layer up from the driver tree.

- New library/protocol/volume-manager: one verb, hello(volume-id) -> the
  block channel as the reply capability (the P0 reply-cap path).
- The volume manager becomes the confinement CONTROLLER: it defines the first
  range on usb-storage, so no other party can confine a filesystem. It
  supervises the filesystems it spawns and respawns one on death (the reap-
  and-rebuild the device manager proved, one layer up).
- fat: drops acquireVolume(device-manager); hellos the volume manager for its
  channel; reads its volume id from argv[1]. main takes process.Init now.
- init.csv no longer spawns fat (the volume manager does); protocol.csv
  rewires fat to be supervised by the volume manager (bind vfs, open
  volume-manager) and drops fat open device-manager.
- The block-range fixture boots registry + device-manager only (not the full
  tree), so the volume manager is absent and the fixture stays the sole
  confinement definer — otherwise the volume manager would take the
  controller first and refuse it.

Verified end to end (VM probes -> spawns fat -> confines it -> hands over the
channel -> fat mounts) and neutral: 18/18 across the fat family, logging,
shutdown, both IOMMU variants, usb restart, vfs, conformance, confinement.
2026-08-09 18:28:44 +01:00
Daniel Samson d56b1b81c0 volume-manager: discovery and probe — V3a
The storage layer gains its policy home (storage-architecture.md): a new
system/services/volume-manager, spawned by init, that acquires the mass-
storage block channel through the device manager (the same lineage a
filesystem uses), reads block 0, and parses the first volume out of it. The
partition-table walk that lived in the FAT engine moves here, above the
driver where it belongs (partition.zig, host-tested: MBR entry, bare-FAT,
no-signature). Identity is the MBR disk signature + partition index — the
weak rung of the ladder; GPT GUID and FAT serial refine identityOf without
changing shape.

This increment is discovery + probe + log only, additive: the FAT service
still acquires its own volume, so nothing changes for it. Confining each
filesystem to its partition and spawning one per volume (the flip) lands
next, keeping fat working throughout.

Grants + wiring: init.csv spawns it after the device manager; protocol.csv
grants bind volume-manager + open device-manager. Verified: volume-probe
asserts the parse (bare-FAT volume at lba 0), neutral 10/10 across storage,
restart, display, logging, confinement — the volume manager now runs in
every boot and disturbs nothing.
2026-08-09 18:10:43 +01:00
Daniel Samson bc67771bfd block: reclaim range slots on death, and give confinement one controller
Two defects the V2 boundary review confirmed:

1. The per-badge range table was never reclaimed. usb-storage receives exit
   notifications (via the subscriber watch), but onNotification handled only
   the timer and dropped child-exits, so a dead filesystem left its range
   slot used forever. The medium-removal lifecycle churns filesystems, so
   after maximum_ranges confine/die cycles define_range would return ENOSPC
   and no volume could be confined again until reboot. onNotification now
   frees the dead badge's slot, mirroring fat's open-node sweep.

2. define_range checked only that the CALLER was unconfined, never that it
   owned the target badge — so any unconfined opener could install a range
   for another live client and silently redirect its I/O. Confinement now has
   a single controller: the first unconfined party to define a range (the
   volume manager, which confines every filesystem before handing it a
   channel). Only the controller may thereafter; the slot releases on its
   death so a restarted manager re-takes it. This is the mechanism half; V3
   adds the grant half (only the volume manager gets an unconfined channel).

Neutral: block-range (the fixture is the sole definer -> controller),
fat-mount, usb-report all green. End-to-end exercise of both lands in V3/V4
(the VM+filesystem relationship and the remount churn).
2026-08-09 17:58:47 +01:00
Daniel Samson 89d4592777 block: close the range-clamp overflow — a confined caller could wrap into the neighbour
The naive bound `lba + count > r.count` wraps for an lba near u64 max: the
sum overflows to a small value, sails under the check, and `base + lba`
wraps to an absolute block OUTSIDE the range. Calibrated, it is a real
confinement escape — a process confined to [1,3) reads absolute block 0
(the boot sector) with lba = maxInt(u64), since base + lba wraps to 0.

The bound is rewritten as two subtractions that cannot overflow: lba within
the range, and count within what remains. block-range gains a wrap-refused
assertion calibrated to be exploitable against the naive form — it FAILS
against the old bound (reads block 0) and passes against the fix (verified
by reverting the clamp). Caught pre-emptively before the V2 boundary review.
2026-08-09 17:40:45 +01:00
Daniel Samson 7af65697cc block: medium presence — the medium_changed event and usb-storage as publisher (V2b)
The block protocol gains a pushed medium_changed event (present + a monotonic
change counter; presence only, never content). usb-storage becomes a
Subscribers provider and runs a slow TEST UNIT READY poll (1 s): success is
present, failure absent, and a transition bumps the counter and publishes.
This is the second removal trigger — the DEVICE stays while the MEDIUM leaves
(card readers, ATAPI trays) — which channel death cannot see
(storage-architecture.md, two triggers one lifecycle).

The subscriber is the volume manager (V3); until it exists the publish is a
no-op fan-out, so this commit is behaviour-neutral, and its end-to-end test
(eject -> medium_changed -> unmount/remount) lands in V4 with the real
consumer rather than a throwaway subscriber fixture (recorded sequencing).
Sense-key inspection to tell medium-absent from other transport errors is a
noted refinement; a clean eject reads correctly as not-ready.

Neutral: 12/12 across the block-serving surface, restart, confinement,
conformance, and logging.
2026-08-09 17:36:57 +01:00
Daniel Samson c37402891a test: block-range — the discrimination fixture for range confinement (V2a)
A process acquires a block channel the way a filesystem does (consumer-hello
the device manager), confines ITSELF to blocks [1,3), then proves the clamp
and the gate: volume-relative LBA 0 maps inside the range and reads; a read
reaching past the range is refused; geometry reports the confined size; and
a confined caller can no longer call define_range (no widening, no escape).
It gates on argv so the ramdisk sweep leaves it silent in other boots, and
coexists with fat (ranges are per-badge).

Discrimination (verified by reverting usb-storage to pre-clamp f1bdce2~1):
the unconfined read still succeeds but define_range returns ENOSYS, so the
fixture cannot arm confinement and the case fails — exactly the property
the clamp adds. With the clamp: block-range 1/1.
2026-08-09 17:25:43 +01:00
Daniel Samson f1bdce25e0 block: per-sender range confinement — V2a mechanism
The block protocol gains define_range (appended, numbers hold): confine the
process named by `badge` to blocks [base, base+count). usb-storage keeps a
per-badge range table and, in read/write, translates volume-relative LBAs
(base added) and refuses any transfer past the volume end. geometry returns
the confined size, so a filesystem mounts against what it may actually touch.

The security seam (decision 4, settled): the clamp lives at the PROVIDER, so
a channel carries exactly the authority it grants — handing a filesystem the
whole disk plus a base offset would let it reach the neighbouring partition.
The gate: a confined caller may NOT call define_range, so a filesystem cannot
widen its own range or confine anyone; only an unconfined party (the volume
manager, whole-device) may. The volume manager defines a filesystem's range
before handing it the channel, so the ordering holds by construction.

Default (no range for a badge) is the whole device — behaviour-neutral for a
single-volume boot and what the volume manager itself uses to probe
partitions. The range table is declared through bounds.md as a runaway
detector (ours, refuse at limit), not a real-partition cap. Neutral:
fat-mount, usb-storage, iommu-usb-storage green. The discrimination fixture
(a confined process reads past its range and is refused) follows next.
2026-08-09 17:13:26 +01:00
Daniel Samson d63a008148 file-system: extract the serving harness from fat — V1
fat was one binary doing four jobs; the three that are not FAT-specific move
to library/kernel/file-system-harness, a Server(comptime Engine) generic over
the engine type: the badge-scoped open-node table, the nine vfs handlers, the
not-mounted politeness, the exit sweep, mount registration, and durable-on-
close. A filesystem is now an engine plus a main that hands the harness a
mounted volume; a second engine reuses the harness wholesale.

Placement note: the plan said library/file-system, but the harness is a
specialization of `service` (its sibling) and needs nothing from the device
domain, so it lives beside service in library/kernel and stays block-free —
durability rides a caller closure (Volume.flush), no backwards kernel->device
dependency, no new-domain scaffolding. The engine type is inferred from
resolve()'s return, so engine.zig is untouched (its Node stays module-scope).

fat keeps only its FAT-specific bring-up (acquireVolume, DMA, engine.mount,
the attach/detach round trip) and the three mount prefixes as data. Behavior-
neutral: 13/13 across the fat/vfs/logger/IOMMU surface, nothing observable
changed. This lands first so every later phase touches the harness once.
2026-08-09 16:55:48 +01:00
Daniel Samson 60b41c0e82 kernel: mounts have owners — V0 of the volume-manager plan
fs_unmount was gated by nothing but the /protocol carve-out: any process
could unmount any prefix — latent with one mount owner, an obvious
cross-tenant hole once volumes multiply. Each backend mount now records the
mounting task, and the syscall layer enforces two rules that keep the
restart story intact: only the owner unmounts (a dead owner's mount is
swept lazily by resolution — strangers gain nothing by racing that), and a
mount may be REPLACED only by its live owner or after its owner died (the
respawned-filesystem path; displacement of a live mount would be worse than
unmounting it). Kernel-installed mounts are never displaceable.

The vfs-test park role is the discrimination: with the volume provably
mounted it attempts the foreign unmount, requires the refusal AND the
subtree still resolving, and withholds its "parked" marker otherwise —
against the ungated kernel the unmount was ALLOWED and vfs-client-death
fails; with the gate, green. (Its verification handle closes immediately:
the kernel test string-matches "released 1 handle(s)".)
2026-08-09 16:16:14 +01:00
Daniel Samson a44b397bed protocols: attach gets its reverse — block detach, usb-transfer dma_detach
The kernel was always symmetric (dma_bind 51 / dma_unbind 52); the two
protocols that forward an attachment up the stack were one-way, so a live
client could grant a device reach into its buffer but never revoke it while
alive — exactly the one-way lifecycle the storage architecture's enforcement
section forbids. Death stays the mechanical backstop; detach is the living
process's path.

Both verbs are appended, so every existing number holds. The shape mirrors
attach precisely: the same region capability rides the cap slot again — the
kernel matches the region, so no layer retains anything between the calls
(the bus never kept the handle; now it never needs to).

fat's bring-up does attach -> detach -> attach, exercising both verbs
through the whole chain (fat -> storage -> bus -> kernel) on every boot: a
broken detach fails every fat case instead of lying dormant until the first
buffer replacement. Honest scope: the round trip proves the plumbing; unbind
semantics are the kernel iommu tests' (map/unmap/translationOf); the full
composition (detach then DMA faults) is a future iommu-fault extension.
2026-08-09 15:18:21 +01:00
Daniel Samson 77e7001878 usb: a hub yanked from a root port takes its subtree with it (H2)
The hot-plug matrix's predicted bug, found on first contact: tearDownPort
never recursed into a departing hub's children — only tearDownHubDevice
(a hub leaving one level down) did. Yank a populated hub from a root port
and the downstream slots stayed live against vanished hardware, their class
drivers were never reaped, and the replugged hub found its port still
occupied, so nothing ever re-enumerated: the subtree was gone for the boot.

tearDownPort now recurses children-first, exactly like tearDownHubDevice.
The usb-hub-yank case is the discrimination: one device_del removes a hub
carrying a keyboard AND a mouse, both drivers must be reaped, and the
re-added hub must rebind both — it failed against the unfixed bus and
passes with the recursion.
2026-08-09 14:08:15 +01:00
Daniel Samson 4a2df587bb establishment: unplug reaps like death, so a replug rebinds
The hot-unplug path (onChildRemoved, a report from a live bus) cleared the
child but left the bound class driver: a process blocked on reports that
will never come, whose stale entry made the matcher's dedupe refuse the
respawn when the device was plugged back in — the same wall the restart
zombie hit, one path over. Unplug now reaps exactly like reporter death.

The usb-hub-unplug case grows the replug: device_del the hub keyboard, then
device_add it back (qmp_sequence); the ordered tail — child removed, reaping,
delegated, ok — can only be satisfied by the second generation, since every
boot keyboard's ok precedes the unplug. Discrimination: without the reap the
replug never rebinds and the case times out (verified by stash run). Hub
family, restart drill, and the two-controller proof all green (8/8).
2026-08-09 13:51:07 +01:00
Daniel Samson 8710944a92 establishment: a reporter's death reaps its subtree, and the re-report rebuilds it
P3 of docs/establishment-planes-plan.md — the restart-zombie fix. A class
driver cannot observe its provider's death: an HID driver blocks on
interrupt reports that will simply never come, and storage answers its
callers with refusals forever. Worse, the dead generation's still-used
entries made the matcher's dedupe refuse the respawn when the restarted bus
re-reported — the subtree was a permanent zombie, which is the exact
opposite of the restart-a-driver-live goal the driver model exists for.

pruneChildrenOf now reaps: each pruned child's bound driver is killed and
its entry cleared (the exit notification finds no entry, so the death is
never double-counted; its device returns by the loan rule; its stored
endpoint handle is closed). The re-report then spawns a fresh generation
whose hellos fetch the successor's channel.

The usb-report drill now asserts the subtree WORKS after the restart: the
respawned storage opens its device on the NEW bus instance and reads block
0. Discrimination: against the pre-reap manager the drill fails — no reap
line, no post-restart respawn (the survivors were zombies), verified by a
stash run. Note the scenario boots no input service, so the HID drivers of
BOTH generations exit after their input lookup times out — storage is the
functional proof.
2026-08-09 12:34:13 +01:00
Daniel Samson 1d7850239d establishment: block stops being a name, and enumerate learns to page
P2 of docs/establishment-planes-plan.md. usb-storage serves nameless — one
process per stick cannot share an exclusive bind, and a second stick used to
die silently on -EBUSY before ever helloing. Its one hello now moves both
directions at once: the block-serving endpoint up, its controller channel
down. fat finds its volume through the manager — a new `.consumer` role asks
for the channel of the driver BOUND TO a device (distinct from the device's
reporter), found by enumerating the tree for the mass-storage identity.
fat stays single-volume; the boot-volume-by-content choice is M21.

The conversion immediately caught a live truncation of exactly the audit's
shape: ChildEntry grew to 32 bytes, one enumerate reply holds ~7, and a real
tree carries a dozen ACPI nodes before the first USB child — the storage
entry silently never fit (the protocol comment already said "paging joins
the protocol if a tree ever outgrows one packet"). enumerate is now paged:
Header.target is the start cursor, a short page is the end; device-list's
page-0 read is unchanged.

Grant rows move with the code: the block bind and fat's block open die, fat
gains open device-manager. Gate: 18 cases green including the registry
trio, device-list, and both IOMMU storage variants.
2026-08-09 12:24:25 +01:00
Daniel Samson d603d40b5c establishment: usb-transfer stops being a name
P1 of docs/establishment-planes-plan.md. The bus no longer binds
/protocol/usb-transfer — the bind race made whichever instance came second
unreachable, which on a real three-controller Ryzen meant a mouse no class
driver could reach ("could not open device 50"). Each instance hands its
serving endpoint up in the hello that already delegates its controller, and
class drivers receive their OWN controller's channel from helloForChannel —
routed by the manager's lineage, retried while a provider is mid-restart,
on one manager handle so retries never spend handle-table slots.

usb.open(bus, id) now takes the channel it used to look up; the name rows
leave protocol.csv with the code (bind 67, opens 119/120/122); the
conformance fixture's prose stops claiming the bus binds; and the stale
input-client import leaves the bus with the channel one.

Gate: 19 QEMU cases green (usb family, hubs, both IOMMU variants, fat chain,
boot-from-USB, orderly shutdown, conformance).
2026-08-09 11:48:31 +01:00
Daniel Samson 77fe4d220e establishment: the mechanics — reply capabilities, helloExchange, lineage routing
P0 of docs/establishment-planes-plan.md; no behavior changes yet, nothing
sets the new flag or sends a hello capability.

- service.run gains a reply-capability out-slot (replyWithCapability), the
  registry idiom init already uses, lifted into the harness; null stays the
  untouched common path. Subscribers gains claimArrival() so a provider
  handler can keep a turn capability through the same flag the reserved
  subscribe uses.
- Hello wire struct: the padding byte becomes wants_channel — old callers
  wire-compatibly say 0, and the manager nominates a reply capability ONLY
  when asked, because a capability sent to a caller that never reads one is
  a leaked slot in that caller's table.
- driver.helloExchange: one handshake can hand a serving endpoint up and
  receive the device's provider channel down (usb-storage will need both at
  once). No channel in the reply is retryable, never a verdict.
- The manager stores each instance's serving endpoint on its Driver entry,
  routes consumer hellos by lineage (child -> reporter -> endpoint), replaces
  on re-hello, and closes the stale handle on death - the 32-slot table is
  the bound that makes forgetting this boot-fatal.
2026-08-09 11:39:29 +01:00
Daniel Samson 5cca580066 kernel: AMD-Vi asked for an interrupt where it meant a store
Two command encodings checked against the specification:

- COMPLETION_WAIT set bit 1 (I, interrupt) instead of bit 0 (S, store), so the
  IOMMU was never asked to write the sentinel, the poll always exhausted its
  spins, and completeAndWait returned without any guarantee the preceding
  invalidation had executed — no invalidation barrier has ever existed, on QEMU
  or on silicon. The in-code claim that "QEMU's amd-iommu does not implement
  the store form" was a misdiagnosis of this bug: with S set, QEMU stores the
  sentinel fine, and the amd-iommu cases now run without the warn line. On real
  hardware, which fetches commands asynchronously, the missing barrier was an
  IOTLB use-after-free window: unmap returned before the invalidation was
  confirmed and the caller freed the frames.

- INVALIDATE_IOMMU_PAGES "invalidate everything" used address bits 51:12
  all-ones; the architected encoding is bits 62:12 all-ones (the spec's literal
  0x7FFF_FFFF_FFFF_F000). Real silicon is free to misread the non-architected
  form as a bounded range.
2026-08-09 09:57:55 +01:00
Daniel Samson 35f43057f4 kernel: the give paths take the lock, and a give confines afresh after a death
Three holes from the real-AMD audit, one shared root: the delegation flag-day
added paths that touch the broker table and the IOMMU records without the big
kernel lock, and a loan-return rule whose re-delegation skipped confinement.

- device_enumerate walked the table with no lock. The table stopped being a
  static array in the bounds track — reserve() regrows it through realloc on
  every boot — so an unlocked reader can be mid-copy out of a slice that
  device_register on another core has already freed and reused, or pair a fresh
  count with a stale slice. Each chunk is now snapshotted under the lock; the
  copy to the user stays outside it.

- system_spawn's give ran entirely unlocked — the comment claiming "the lock
  has not been dropped" was false (spawnProcessSupervised takes and releases it
  internally). The ownership pre-check now only spares creating a doomed child;
  the give itself re-checks, confines and moves in one lock hold, and a give
  that fails after the spawn kills the child rather than leaving it running
  without the hardware it was spawned for.

- A re-delegated device after a driver death was never re-confined. Death tears
  the domain down before the loan returns to the lender, so the next give found
  no active record, reassign no-op'd, and the respawned driver ran the device
  with a V=0 device-table entry and no domain — silently unconfined, the exact
  fail-open the fail-closed claim was built to remove. Both give paths now share
  one body (giveDeviceLocked): check first, confine afresh when no record is
  active — refusing with ECONFINE like the claim — and move last, when nothing
  can fail.

The iommu test drives the death-and-respawn sequence directly; with the old
reassign-only behaviour its two confinement checks fail, with this change the
suite is 118/118.
2026-08-09 09:57:55 +01:00
Daniel Samson 0eb2420690 device-manager: delete the delegated-set scaffolding
The name list and its predicate existed so drivers could move to delegation
one at a time with the suite green throughout. Every driver is delegated
now, so the manager simply hands over whatever device a driver was assigned.

Deleting it caught a real consequence: crash-test finally got delegated too,
and it was still claiming its device — so it got AlreadyClaimed because it
already held it, exited, and the restart drill had nothing to restart. Its
own comment named what the case was really checking: "the respawn only
reaches this line because the kernel released the previous instance's claim
at death". That property still holds, by a different mechanism — the device
reverts to the manager on death and is handed to the replacement, which is
the same guarantee without the race it used to rely on.

All four delegation paths verified: the xHCI controller, the PCI bridge, the
PS/2 two-node singleton, and virtio-gpu's restart re-attach.

Run 3 complete. Suite 118/118.
2026-08-08 23:01:29 +01:00
Daniel Samson df9c1ed827 device-manager: hold the seeded hardware so none is left lying around
A device nobody holds can be claimed by anyone, so the manager now takes
every firmware-discovered device that carries mappable resources, whether or
not a driver wants it. The real gap was the HPET: an MMIO window, an IRQ, no
user-space driver, and there for the taking. Held by the manager it is
inert; unheld it was a way into physical memory.

Two deliberate exclusions. The loader's framebuffer, which the compositor
claims and which the manager must not take because it starts first. And
anything with no resources, which grants nothing worth holding.

Scope is the boot snapshot. A device reported later and matched to no driver
stays claimable — pci-cap-test and iommu-fault-test both reach an unmatched
NIC that way, so narrowing it is a separate change with those fixtures in
scope. Recorded in the plan rather than left implied.

The attacker fixture gains the assertion deferred since D2: after the system
settles, nothing with resources may be taken.

That assertion defeated itself twice before it worked, and both failures are
worth remembering. First it swept at 0.029 while the manager did not bind
its protocol until 0.047, so it reported a hole that closed a millisecond
later. The retry loop that "fixed" that was worse: the first pass TAKES the
device, so the second finds it unavailable because this process now holds
it, and concludes all is well — it passed with the manager's claiming
removed entirely. It now settles once and sweeps once, and fails when the
claiming is removed.

Suite 118/118.
2026-08-08 22:42:42 +01:00
Daniel Samson ba195fa0a2 kernel: claim refuses delegated hardware — and E2 had already closed the hole
The rule as planned: a device that was given to someone may be handed on,
never taken. Implemented, and honest about what it is worth.

Writing the test showed the plan had the wrong step doing the work. A
delegated device is HELD, so an attempt to take it is refused as
AlreadyClaimed before the giver is ever consulted; and once a borrower's
death returns the device to its lender — or clears both when the lender is
gone — there is no state where a device is unheld and still on loan. The
window a stranger could have used stops existing at E2. This check is
unreachable.

It stays anyway: one comparison, failing closed, guarding any future path
that frees a device without clearing its giver, which is exactly the hole
this run closed. The comment says it is unreachable rather than implying a
protection it does not provide.

The attacker fixture does not gain the assertion that was deferred to this
step, and its header records why: there is no refusal for it to observe, and
on a bare boot with no device manager nothing is delegated at all, so the
assertion had nothing to bite on. It failed loudly on its first run rather
than passing quietly, which is the only reason this was noticed.

It also leaves the loader's framebuffer alone without naming it: nobody
delegates the framebuffer, so it has no giver, so the display service claims
it exactly as before.

Suite 118/118.
2026-08-08 22:31:04 +01:00
Daniel Samson 1a1d92cba9 kernel: a grant is a loan — a dead borrower returns the device to its lender
When a driver dies, a device it was *given* now goes back to whoever lent
it, rather than to nobody. The device manager gets its hardware back the
instant a driver dies and hands it to the replacement, with no window in
between.

That window was real: the kernel released the claim to no one and the
manager re-claimed first-come, so every driver restart reopened the hole
this run is closing. It also becomes load-bearing at the next step — once
claim refuses a device that has a giver, releasing to nobody would strand a
dead driver's hardware permanently, because nobody could ever take it again.

A dead lender is no lender: the claim and the giver clear together, so a
device is never owed to a ghost. A device nobody lent is released outright,
exactly as before.

The broker cannot see the task table, so liveness arrives through the same
hook idiom the scheduler already uses. Null means assume dead, so a kernel
built without the hook frees claims rather than handing them to a ghost.

A stale binary nearly passed as proof for the third time this session: the
first discrimination patch left `alive` unused, the build failed with three
errors, and the old binary reported every assertion passing. Checking the
build before reading results is what caught it.

Suite 118/118.
2026-08-08 22:21:25 +01:00
Daniel Samson 4ca57fc37e kernel: record who gave each device away
One field, and the rest of the run follows from it. A device that was given
to someone is delegated hardware: it may be handed on, never taken, and when
its holder dies it goes back to whoever lent it instead of becoming free for
anyone to grab.

It also settles the framebuffer without mentioning it. Nobody delegates the
loader's framebuffer, so it has no giver, so the display service claims it
exactly as it always has — no exemption and no reference to display anywhere
in the rule.

No behaviour changes here; the field is recorded and read by nothing yet.

The test found a real bug on its first run, before the discrimination check.
The sentinel for "nobody gave this" was 0 — and task 0 is a real task, the
kernel's own, so a device given away by task 0 read back as belonging to
nobody. Both giver and registrar are optionals now. The second was a latent
bug from D9: the per-registrar allowance would have miscounted every device
task 0 registered.

Suite 118/118.
2026-08-08 22:12:01 +01:00
Daniel Samson b7d97ebb5d acpi: discovery is handed its node like every other driver
The last claimant. The kernel seeds the acpi-tables node, so it sits in the
same boot snapshot the manager already scans to find the PCI host bridge —
there was never a bootstrap problem, only a lookup nobody had written. The
manager claims it and names it in the spawn; the service stops claiming.

Every driver in the system now receives its hardware rather than taking it.

Two failures on the way, both mine. addDriver puts the device id in argv[1],
and the acpi service read argv[1] as a self-verify device-count floor — so
handed device 7 it decided it was in test mode, printed "acpi-parse: ok",
and never reported a device. The test argument is now floor:N, which a bare
id cannot be mistaken for.

And acpi-parse spawns the service directly rather than through the manager,
so nothing handed it the node. That test now claims and transfers it exactly
as the manager does, which is the right shape: the test plays the manager's
role instead of the service reaching for hardware.

device_claim now has two callers left: the manager, which is the acquirer
and should have it, and the display service's GOP path. That is recorded as
question 10 — the framebuffer is not a device, so the answer is likely that
it leaves the device table rather than being exempted from its rules.

Suite 118/118.
2026-08-08 21:54:45 +01:00
Daniel Samson 6b3a381626 ps2: one instance, every node the machine has
The 8042 is a single controller described by two ACPI nodes — PNP0303
carries the 0x60/0x64 ports, PNP0F13 is the mouse — so it cannot be split
across two processes without them fighting over the same registers. That is
why ps2-bus is a singleton, and why it used to find and claim both nodes
itself.

The manager now gives it every matching node instead. The keyboard node
rides the spawn, because it holds the ports and is needed immediately; the
mouse node is transferred to the already-running instance. Late arrival is
safe here and the ordering is natural rather than lucky: the mouse is not
touched until after the controller handshakes and identify. Measured, the
handover lands at 0.336 and the driver reaches the mouse at 0.456.

Because the count is however many matched, a machine with no PS/2 ports or
only one works without a special case — which matters, since the bus is
mostly emulated now and machines vary.

ps2-bus claims nothing. irq_bind on the mouse node is the proof it holds it:
that call is ownership-gated, so a failure means the handover did not land
rather than a hardware fault, and the log says so.

acpi-ps2 asserts both delegations with the spawned device backreferenced, so
the node that rides the spawn must be the one the driver was spawned for.
Disabling the second delegation fails it.

Two things worth recording. The first discrimination patch was not valid Zig,
so nothing ran and a stale binary reported a pass — checked the build before
believing it. And with the second delegation disabled, acpi-ps2 fails while
input still passes: the mouse works without its IRQ binding, so exactly one
case covers that path.

Suite 118/118.
2026-08-08 21:33:35 +01:00
Daniel Samson 7d8aa51234 kernel: maximum_children_per_parent is gone
The second invented ceiling. It was written to stop a driver looping
device_register and exhausting a shared table — but there is no shared table
to exhaust any more, and each registrar already has its own allowance, so a
runaway costs only itself.

It never bounded a determined caller in the first place: 16 children per
parent, and nothing stopped it claiming more parents. What it reliably did
was refuse a real PCI bus with more than 16 functions, which is how an AMD
Ryzen booted with a working display, no USB and no storage.

The constant, its check, and the now-unused childCount all go. TooManyChildren
survives with one meaning instead of two: the caller is at its per-registrar
allowance.

This was unblocked from the moment D9 landed. The plan said so — "once the
quota exists the per-parent cap is redundant whether or not D6 has landed" —
in the same edit that left the step tagged "blocked on D6". Three iterations
were then spent re-reading that tag instead of the sentence beside it.

The containment test now asserts 64 children under one parent, four times
the old ceiling; restoring the cap fails it.

Suite 118/118.
2026-08-08 21:12:20 +01:00
Daniel Samson 3ae541214f iommu: assert directly that a confinement moves with its device
reassign was added at D4 to fix a regression and has been proven only
indirectly since — three IOMMU+USB cases going green. That covered the
visible symptom (a driver's DMA rings unbound) and neither of the latent
ones: the confinement still naming the giver, so the giver's death would
tear down a domain a live driver was using, and the receiver's death would
leave one behind. Those are now asserted.

confinementOwner exposes the record's owner so the suite can see it. The
sequence is the delegation in miniature: unconfined, confine as this task,
reassign to another, confirm the new holder owns it and the old one does
not, then kill the new holder and confirm the domain goes with it.

Two attempts at this test could not have failed. The first found no PCI
function to confine — pciAddressOf needs a pci_device entry and this case
runs no pci-bus — so every assertion skipped silently while the case stayed
green. It now synthesizes a function the way pci-bus does, a 4 KiB config
window inside the bridge's ECAM, and asserts that precondition explicitly so
a skip is a failure.

Verified to discriminate: making reassign a no-op flips three assertions,
including the domain surviving its holder's death.

Suite 118/118.
2026-08-08 19:57:36 +01:00
Daniel Samson 2ebfccc8ed virtio-gpu: the scanout device arrives with the spawn
Third driver converted. It no longer claims the id from argv[1] — the
manager holds the device and names it in the call that creates the process,
so it is held before the driver's first instruction.

display-reattach is the case that matters here: it kills the driver and
watches the compositor re-attach to the fresh scanout. It passes, so the
restart path survives the fused grant — the manager re-takes the device when
the driver dies and hands it to the replacement.

ps2-bus and discovery are NOT converted, and the reason is recorded as open
question 9 rather than worked around. Both need a device nobody assigned
them. ps2-bus ignores its argv[1] entirely: it finds the controller by
walking the table for PNP0303, then claims a second device, the PNP0F13
mouse node, which it also finds itself — so it holds two devices and was
assigned at most one, while system_spawn carries one. discovery claims the
acpi-tables node it locates itself, because it is what produces the device
tree and there is nothing to assign at that point.

One thing worth checking before designing an answer: devices.csv maps both
PS/2 hardware ids to ps2-bus, so the manager may already be spawning two
instances where the driver expects one. If so the fix is smaller than it
looks.

D6 stays blocked — closing device_claim with these two still depending on it
would stop the machine booting.

Suite 118/118.
2026-08-08 19:48:00 +01:00