library: the harness keeps the subscribers, and an id belongs to whoever opened it

Three services had each written the same thing and got it three different
ways: input polled the process list to notice a dead subscriber, and only
when someone else subscribed; the power service never noticed at all; the
device manager noticed drivers but not subscribers. The harness owns the
table now, driven by the events a protocol declares — it registers on the
reserved verb, frames each event once, posts to everyone interested without
waiting on any of them, and reclaims a slot when the kernel says its owner
died. Interest masks moved to the envelope, so a subscriber that wants only
mice asks the same way everywhere.

Two consequences the plan had not foreseen. The device manager now hears a
supervised child's death twice, once as its supervisor and once as a
subscriber, so restart backoff counted every crash twice and gave up after
half as many; it retires the id before counting. And the kernel's published
exit table had eight slots for what is now six subscriptions in a plain
boot, so it holds sixteen.

The other half is a hole the design named early and left standing: a
backend handed out a small integer and then honoured it from anyone. A
process that guessed a file's node id read another client's file; a display
layer had no owner at all, so any client could reconfigure or destroy any
layer; a USB device token was never checked against the client that opened
it. Each is now bound to the task that opened it, and a wrong owner gets
exactly what an unknown id gets — the refusal must not become the oracle
the identical answers elsewhere were designed to remove. Closing a file
changed with it: it used to succeed unconditionally, which would have told
a caller which ids existed.

Suite 111/111, with a new case in which one process holds a file and a
layer, hands both ids to a second process, and finds them untouched after
that process has tried everything with them.
This commit is contained in:
Daniel Samson
2026-08-01 09:05:26 +01:00
parent 2719b93530
commit 1b1c587c14
29 changed files with 1072 additions and 335 deletions
+62
View File
@@ -2248,9 +2248,67 @@ fn processKillTest(boot_information: *const BootInformation) void {
}
check("neither victim is listed after its kill", !still_listed);
check("a killed id stays dead (-ESRCH on a second kill)", process.killProcess(me, sleeper) == -ipcsync.ESRCH);
publishedExitChecks(rd, me);
result();
}
/// The mechanism every provider's release-what-a-dead-client-held sweep is built
/// on (docs/process-lifecycle.md, "Who learns of a death"): a **published** exit,
/// fanned out to every subscriber rather than only to the supervisor. The service
/// harness's subscriber sweep, the FAT server's open files, the compositor's
/// layers and the xHCI driver's device tokens all release on exactly this, and
/// several of them are subscribed at once in a normal boot — so what is checked
/// here is the fan-out: three independent subscribers, one death, three
/// notifications carrying the same badge, none of them the supervisor's.
///
/// Run at the end of the process-kill case, because a subscription is for every
/// death from then on and the checks above spawn victims of their own.
fn publishedExitChecks(rd: initial_ramdisk.Reader, me: u32) void {
var subscribers: [3]*ipcsync.Endpoint = undefined;
var subscribed: usize = 0;
while (subscribed < subscribers.len) : (subscribed += 1) {
subscribers[subscribed] = ipcsync.createIpcEndpoint() orelse break;
if (process.subscribeExits(subscribers[subscribed], me) != 0) break;
}
check("three endpoints subscribed to published exits", subscribed == subscribers.len);
if (subscribed != subscribers.len) return;
// A supervised child so the supervisor notification remains distinguishable:
// it lands on `endpoint`, the published ones on the three above.
const supervisor_endpoint = ipcsync.createIpcEndpoint() orelse {
check("supervisor endpoint allocated", false);
return;
};
var child: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(initial_ramdisk.basename(item.name), "args-echo")) continue;
child = process.spawnProcessSupervised(item.blob, 4, &.{ "args-echo", "published-exit" }, me, supervisor_endpoint) catch 0;
break;
}
check("a clean-exit child spawned for the published exit", child != 0);
if (child == 0) return;
var badge: u64 = 0;
var received_cap: u64 = 0;
_ = ipcsync.replyWait(supervisor_endpoint, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap);
check("the supervisor heard the child end", badge == abi.notify_badge_bit | abi.notify_exit_bit | child);
// The publication happens in the same locked section as the supervisor's
// notification and before it, so all three are already queued: a subscriber
// that had not heard would block here and time the harness out rather than
// pass vacuously.
var heard: usize = 0;
for (subscribers) |subscriber| {
badge = 0;
_ = ipcsync.replyWait(subscriber, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap);
if (badge == abi.notify_badge_bit | abi.notify_exit_bit | child) heard += 1;
}
check("every subscriber heard the same death, not just the supervisor", heard == subscribers.len);
}
/// M17.1: a dead process's device claims are released by the reap, so a restarted
/// driver can claim its hardware again (docs/process-lifecycle.md iron rule 1).
/// First the broker release in isolation — two owners, one released, the other's
@@ -2739,6 +2797,10 @@ fn fatMountTest(boot_information: *const BootInformation) void {
const init_ok = if (process.spawnBundled("/system/services/init")) true else |_| false;
check("init spawned (boots the tree, incl. the fat server)", init_ok);
check("fat-test client spawned", spawnNamed(rd, "fat-test"));
// The badge-scoping probe rides the same boot: it needs the fat server for a
// node id and the compositor for a layer id, and init starts both. It spawns
// its own second process — the intruder — with the ids it holds (P4c).
check("badge-scope-test owner spawned", spawnNamed(rd, "badge-scope-test"));
result();
}