The namespace holds protocols, never instances; what multiplies is provider processes, and their establishment routes through the device manager, which owns the topology. communication.md gets the model; the plan converts the usb-transfer and block seams in one flag-day, closes the restart-zombie hole the scoping found, and pins the three-controller mouse bug as a QEMU case.
7.6 KiB
Establishment planes: the plan
2026-08-09. The design is settled in
communication.md ("Establishment: two planes, one
namespace"): driver-layer channels stop being registry names and are routed at
establishment by the device manager, which owns the topology. This converts both live
seams — usb-transfer and block — in one flag-day, and fixes the restart hole the
scoping audit found on the way. Found by a real machine: three xHCI controllers, one
name, a mouse nobody could reach.
Non-goals, stated up front. fat stays single-volume (multi-volume mounting, volume
identity/content-probing, and fat's own reconnection after a storage restart are the
M21 remount track). The input service's last-writer-wins button mask across two mice
is a known quirk, not part of this change. Nothing here touches the app-facing names —
input, display, ufs keep their exclusive registry binds.
What the scoping reads established
- Capabilities already travel both directions of a synchronous call, and only
there. Request-direction via
ipc.callCap(r9 →shareCapabilityat server dequeue, a refcounted copy); reply-direction viaipc_reply_wait'ssend_cap(surfaces as the client'sReply.cap) — init's registry is the working in-tree precedent (pending_capability). Pushes structurally cannot carry one. No kernel change is needed. - The one mechanism gap is the service harness:
service.runhardcodes the reply capability to null andAnswerhas no capability slot; the device manager runs on that harness. It needs a reply-cap out-slot following init's pattern, with null staying the untouched default for every other service. - The manager already stores the routing fact:
Child{parent, bus_address, device_id, reporter}— consumer's device id →children[].reporter→driverByProcess(reporter)→ the (new) stored endpoint.Hello.role(.bus vs .device) already exists on the wire to distinguish provider hellos (cap up) from consumer hellos (cap down); the manager currently never reads it. - The restart hole (new finding): today, when a bus instance dies, its class drivers are zombies — HID drivers block forever on interrupt reports that will never come, usb-storage answers fat with errors forever; the manager respawns neither, and the matcher's dedupe actively prevents a respawn on re-report. The user's core goal is restart-a-driver-live; this plan closes the hole.
- Blast radius: every QEMU case boots from USB storage, so this seam failing
fails the whole suite, and the load-bearing serial markers (
fat: mounted /volumes/usb,usb-storage: ready,fat-test: ok) must not change.
P0 — mechanics (library + manager, no behavior change yet)
- Harness reply-capability out-slot. A handler nominates a capability for its
reply (init's
pending_capabilityidiom moved intoservice.run's loop); default null; the existing Arrival take/release ownership discipline is unchanged. Every other service compiles and behaves identically. driver.hellogrows both directions. Provider form: carry the caller's serving endpoint up (ipc.callCap). Consumer form: returnReply.cap. One call can do both (a storage driver hands its block endpoint up and receives its bus channel down in the same hello).- Manager stores and routes.
Drivergains an endpoint field; the cap is claimed atonMessagelevel (thecapability_claimedpattern usb-xhci-bus already uses — the harness's private claim flag composes, verified).onDriverExitmustipc.closethe stored handle (a dead owner's endpoint survives as a refcounted slot; repeated restarts would exhaust the manager's 32-slot table — boot-fatal). Consumer hello with an unknown device id (bus died, not yet re-reported) gets a retryable refusal, never a permanent one.
P1 — the usb-transfer seam
- Bus: delete the
bindPatiently("usb-transfer")block and its serve-unnamed fallback; hello carries the endpoint up. usb.open(device_id)takes the channel from the class driver's hello instead of the name lookup; retry patience stays ≥ today's window (the provider's hello precedes the consumer's spawn in the normal order — the retry covers respawn races).- usb-storage's consumer side uses the same lineage (its own device id).
protocol.csv: delete the bind row (67) and the three open rows (119/120/122) in the same commit as the code — grants and code move together or the old path dies silently first.- Fixtures: conformance table row, the usb-report kernel-side registry spawn.
P2 — the block seam
- usb-storage: drop
.service = "block"(the field is optional; the harness serves nameless — pci-bus and logger already do); itsinitialisekeeps the endpoint it currently discards and hellos it up. Exit-code contract preserved: device absent = clean exit (no restart), device-present failure = exit(1) (restart with backoff) — establishment failure is the latter, never a silent pre-hello death like today's second-stick-EBUSY. - fat:
block.tryOpen's name lookup becomes manager-routed — enumerate, find the block-class provider, hello for its channel. One volume: unambiguous. Two volumes: fat takes the first by device-id order — deterministic within a boot, and no worse than today's bind race; choosing the boot volume by content is deferred to M21 (risk noted:/system/configurationand/system/logsmounts ride this choice). - fat gets a
protocol.csvopen grant fordevice-manager; rows 68/99 die. - Delete
block.open(zero callers, dead code).
P3 — restart integrity (the zombie fix)
- A reporter's death reaps its subtree. When a bus/storage instance dies, the
manager already prunes its children; now it also kills the class drivers spawned
for those children and clears their entries, so the matcher's dedupe stops blocking
the respawn. On re-report (ids are stable — the
P<port>I<iface>registration identity), the matcher spawns fresh class drivers, which hello and receive the successor's channel. Restart a bus driver live: the subtree rebuilds itself. - usb-storage on
-EPEERfrom its bus: exit(1) → supervisor restart with backoff → fresh hello. (fat's own reconnection: M21.) - Extend the existing bus-restart drill (
test-usb-restart, the usb-report case) to assert the subtree works after the restart — a post-restart keyboard event, not just a re-report. Discrimination: today that assertion fails (zombie).
P4 — the multi-instance proof
- New QEMU case: two xHCI controllers, keyboard on one, storage on the other; assert both class drivers come up and function. Discrimination: without the conversion this case fails exactly like the Ryzen mouse (second controller's child unreachable). This is the machine's bug, pinned in the suite forever.
- Full suite green; the four existing second-controller cases (usb-hub and friends) re-checked against their expectations.
- Docs updated with the code: device-manager.md, drivers.md (hello), the usb-transfer and block protocol docblocks, device-authority.md as-built appendix.
Order and verification discipline
P0 → P1 → P2 → P3 → P4, one commit per coherent step, suite green at every phase boundary (P0 and P1 may share a verification run — P0 alone changes no behavior). Every new test must be shown to fail against the old behavior before it counts (the discrimination rule). One QEMU suite at a time. The Ryzen re-test at the end is the acceptance run: mouse moves, and killing one bus instance from a shell one day only blinks the devices on that controller.