establishment: the mechanics — reply capabilities, helloExchange, lineage routing

P0 of docs/establishment-planes-plan.md; no behavior changes yet, nothing
sets the new flag or sends a hello capability.

- service.run gains a reply-capability out-slot (replyWithCapability), the
  registry idiom init already uses, lifted into the harness; null stays the
  untouched common path. Subscribers gains claimArrival() so a provider
  handler can keep a turn capability through the same flag the reserved
  subscribe uses.
- Hello wire struct: the padding byte becomes wants_channel — old callers
  wire-compatibly say 0, and the manager nominates a reply capability ONLY
  when asked, because a capability sent to a caller that never reads one is
  a leaked slot in that caller's table.
- driver.helloExchange: one handshake can hand a serving endpoint up and
  receive the device's provider channel down (usb-storage will need both at
  once). No channel in the reply is retryable, never a verdict.
- The manager stores each instance's serving endpoint on its Driver entry,
  routes consumer hellos by lineage (child -> reporter -> endpoint), replaces
  on re-hello, and closes the stale handle on death - the 32-slot table is
  the bound that makes forgetting this boot-fatal.
This commit is contained in:
Daniel Samson
2026-08-09 11:39:29 +01:00
parent 0d5a7394ef
commit 77fe4d220e
4 changed files with 116 additions and 6 deletions
@@ -140,6 +140,12 @@ const Driver = struct {
spawn_ns: u64 = 0,
hello_deadline_ns: u64 = 0,
restart_due_ns: u64 = 0,
// The serving endpoint this instance handed up in its hello — what consumer
// hellos for its reported children are answered with (establishment by
// lineage, communication.md "Establishment: two planes"). A refcounted
// handle in OUR table: onDriverExit must close it, or every restart leaks a
// slot of the manager's 32 until no capability can arrive at all.
endpoint: ?ipc.Handle = null,
fn name(driver: *const Driver) []const u8 {
return driver.name_buffer[0..driver.name_len];
@@ -217,6 +223,19 @@ fn childCountOf(reporter: u32) u32 {
return n;
}
/// The child a kernel device id belongs to — the lineage lookup: its
/// `.reporter` names the driver instance that provides it, which is how a
/// consumer's hello is routed to the right provider (communication.md
/// "Establishment: two planes"). Null when nothing reported it, or its
/// reporter died and pruned it — a retryable gap, not a verdict.
fn childByDevice(device_id: u64) ?*Child {
if (device_id == device_manager_protocol.no_device) return null;
for (&children) |*child| {
if (child.used and child.device_id == device_id) return child;
}
return null;
}
/// The driver entry a live process id belongs to. Zero is not a process id here:
/// it is what `onDriverExit` writes back to retire an id it has already acted on,
/// so a second notification for the same death matches nothing.
@@ -328,6 +347,13 @@ fn onDriverExit(driver: *Driver) void {
// Without this the backoff would count one death twice and the crash-loop cap
// would fire at half the deaths it names.
driver.process_id = 0;
// The dead instance's serving endpoint is stale the moment it died — the
// kernel marked the endpoint dead, but our refcounted handle would sit in
// the 32-slot table forever. The successor's hello stores a fresh one.
if (driver.endpoint) |stale| {
_ = ipc.close(stale);
driver.endpoint = null;
}
pruneChildrenOf(dead);
const reason = process.exitReason(dead) orelse .fault;
if (reason == .exited) {
@@ -489,6 +515,30 @@ fn onHello(_: void, invocation: Invocation(device_manager_protocol.Hello), _: An
};
driver.state = .running;
// A provider's hello carries its serving endpoint — the channel consumer
// hellos for its reported children are answered with. Stored on the entry
// (claimed from the turn, so the harness's defer leaves it alone); a
// re-hello replaces, closing the old handle rather than leaking the slot.
if (invocation.capability) |serving| {
if (driver.endpoint) |previous| _ = ipc.close(previous);
driver.endpoint = serving;
Serve.claimArrival();
}
// A consumer's hello asks for its device's provider: the child's reporter
// is the routing fact (establishment by lineage). No channel is not a
// refusal — the provider may be mid-restart and its re-report on the way —
// so the hello still acks and the consumer retries. Nothing is nominated
// unless asked: a capability sent to a caller that never reads one is a
// leaked slot in ITS table.
if (invocation.request.wants_channel != 0) {
if (childByDevice(invocation.target)) |child| {
if (driverByProcess(child.reporter)) |provider| {
if (provider.endpoint) |serving| service.replyWithCapability(serving);
}
}
}
std.log.info("hello from {s} (device {d})", .{ driver.name(), invocation.target });
// Resilience drill (V6): once, kill the virtio-gpu driver a moment after it hellos, so
// the normal restart policy respawns it — the compositor must survive and re-attach.