library: the harness keeps the subscribers, and an id belongs to whoever opened it

Three services had each written the same thing and got it three different
ways: input polled the process list to notice a dead subscriber, and only
when someone else subscribed; the power service never noticed at all; the
device manager noticed drivers but not subscribers. The harness owns the
table now, driven by the events a protocol declares — it registers on the
reserved verb, frames each event once, posts to everyone interested without
waiting on any of them, and reclaims a slot when the kernel says its owner
died. Interest masks moved to the envelope, so a subscriber that wants only
mice asks the same way everywhere.

Two consequences the plan had not foreseen. The device manager now hears a
supervised child's death twice, once as its supervisor and once as a
subscriber, so restart backoff counted every crash twice and gave up after
half as many; it retires the id before counting. And the kernel's published
exit table had eight slots for what is now six subscriptions in a plain
boot, so it holds sixteen.

The other half is a hole the design named early and left standing: a
backend handed out a small integer and then honoured it from anyone. A
process that guessed a file's node id read another client's file; a display
layer had no owner at all, so any client could reconfigure or destroy any
layer; a USB device token was never checked against the client that opened
it. Each is now bound to the task that opened it, and a wrong owner gets
exactly what an unknown id gets — the refusal must not become the oracle
the identical answers elsewhere were designed to remove. Closing a file
changed with it: it used to succeed unconditionally, which would have told
a caller which ids existed.

Suite 111/111, with a new case in which one process holds a file and a
layer, hands both ids to a second process, and finds them untouched after
that process has tried everything with them.
This commit is contained in:
Daniel Samson
2026-08-01 09:05:26 +01:00
parent 2719b93530
commit 1b1c587c14
29 changed files with 1072 additions and 335 deletions
+23 -4
View File
@@ -1520,7 +1520,15 @@ pub fn exitReasonOf(caller_id: u32, target_id: u32) i64 {
/// subscriptions — that must release what a dead client held and cannot learn it
/// any other way (a client that simply never calls again looks like silence).
/// Bounded like every kernel table; each entry holds its own endpoint reference.
const exit_subscriber_capacity = 8;
///
/// Sixteen, not eight: a subscription is now what *every* provider with
/// per-client state uses to release it — the FAT server's open files, the input,
/// power and device-manager subscriber tables (the service harness subscribes for
/// them), the compositor's layers, and each USB controller driver's device tokens.
/// A single boot already fields six, and a machine with several xHCI controllers
/// fields one per controller, so the old ceiling was within two of a service
/// silently losing its sweep.
const exit_subscriber_capacity = 16;
const ExitSubscriber = struct { endpoint: *ipc.Endpoint, owner: u32 };
var exit_subscribers: [exit_subscriber_capacity]?ExitSubscriber = .{null} ** exit_subscriber_capacity;
@@ -1538,16 +1546,27 @@ fn systemProcessSubscribe(state: *architecture.CpuState) void {
if (t.address_space == 0) return fail(state);
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
if (!ipc.ownedBy(endpoint, t)) return failErr(state, ipc.EPERM);
const result = subscribeExits(endpoint, t.id);
if (result < 0) return failErr(state, @intCast(-result));
architecture.setSystemCallResult(state, 0);
}
/// Take a slot in the published-exit table for `endpoint`, owned by task `owner`.
/// The body of `process_subscribe` minus the authorization, so a kernel test can
/// exercise the fan-out (several subscribers, one death, every one notified) that
/// every provider's release-what-the-dead-client-held sweep is built on. Returns
/// 0, or -ENOSPC.
pub fn subscribeExits(endpoint: *ipc.Endpoint, owner: u32) i64 {
const flags = sync.enter();
defer sync.leave(flags);
for (&exit_subscribers) |*slot| {
if (slot.* == null) {
endpoint.refcount += 1; // the slot's own reference, dropped on unsubscribe-by-death
slot.* = .{ .endpoint = endpoint, .owner = t.id };
return architecture.setSystemCallResult(state, 0);
slot.* = .{ .endpoint = endpoint, .owner = owner };
return 0;
}
}
failErr(state, ipc.ENOSPC);
return -ipc.ENOSPC;
}
/// signal_bind(endpoint): nominate where this process's signals arrive — the