library: the harness keeps the subscribers, and an id belongs to whoever opened it

Three services had each written the same thing and got it three different
ways: input polled the process list to notice a dead subscriber, and only
when someone else subscribed; the power service never noticed at all; the
device manager noticed drivers but not subscribers. The harness owns the
table now, driven by the events a protocol declares — it registers on the
reserved verb, frames each event once, posts to everyone interested without
waiting on any of them, and reclaims a slot when the kernel says its owner
died. Interest masks moved to the envelope, so a subscriber that wants only
mice asks the same way everywhere.

Two consequences the plan had not foreseen. The device manager now hears a
supervised child's death twice, once as its supervisor and once as a
subscriber, so restart backoff counted every crash twice and gave up after
half as many; it retires the id before counting. And the kernel's published
exit table had eight slots for what is now six subscriptions in a plain
boot, so it holds sixteen.

The other half is a hole the design named early and left standing: a
backend handed out a small integer and then honoured it from anyone. A
process that guessed a file's node id read another client's file; a display
layer had no owner at all, so any client could reconfigure or destroy any
layer; a USB device token was never checked against the client that opened
it. Each is now bound to the task that opened it, and a wrong owner gets
exactly what an unknown id gets — the refusal must not become the oracle
the identical answers elsewhere were designed to remove. Closing a file
changed with it: it used to succeed unconditionally, which would have told
a caller which ids existed.

Suite 111/111, with a new case in which one process holds a file and a
layer, hands both ids to a second process, and finds them untouched after
that process has tried everything with them.
This commit is contained in:
Daniel Samson
2026-08-01 09:05:26 +01:00
parent 2719b93530
commit 1b1c587c14
29 changed files with 1072 additions and 335 deletions
+19 -9
View File
@@ -108,15 +108,25 @@ This is the async counterpart of `ipc_call`, and the input service is its first
`publishJoystickEvent`. Publishing is a short synchronous `ipc_call` the service answers at
once; the service's own fan-out is asynchronous, so publishing never blocks on a slow
subscriber.
- The **service** ([input.zig](../../system/services/input/input.zig)) keeps a small subscriber
table (endpoint handle + owning task id + `device_mask`). On `publish` it `ipc_send`s the
event to every subscriber whose mask includes the event's device class. On `subscribe` it
stores the passed capability and mask and, as housekeeping, prunes any slot whose owning
process has exited (checked against `process_enumerate`) — not for correctness (an async
send to an orphaned endpoint is harmless) but to reclaim the slot. The service runs on the
shared harness ([service.zig](../../library/kernel/service.zig)) like every other, so it
answers the universal ping and exits on `terminate`; it was the last hand-rolled receive
loop in the tree, and the last service a shutdown had to kill rather than ask.
- The **service** ([input.zig](../../system/services/input/input.zig)) owns none of that
machinery any more: the subscriber table (endpoint handle + owning task + interest mask),
the reserved `subscribe`/`unsubscribe` verbs, the fan-out, and the dead-subscriber sweep
are the shared harness's (`service.Subscribers` in
[service.zig](../../library/kernel/service.zig)), so every event stream in the system has
identical semantics. What is left in this file is what is actually about input: which
class an event belongs to, and which classes a subscriber asked for. On `publish` it names
the event's class and the harness `ipc_send`s the packet — framed once — to every
subscriber whose mask includes it.
- **A dead subscriber goes away on its exit notification**, not on a poll. The service used
to walk `process_enumerate` on every subscribe and drop slots whose owner had gone; it now
subscribes to the kernel's published exits like the FAT server and the compositor do
([process-lifecycle.md](../os-development/process-lifecycle.md)), which reclaims the slot
*and* closes the endpoint capability in it promptly rather than at the next subscribe.
(The fan-out also drops a subscriber whose `ipc_send` fails, as a backstop for a
notification a full ring dropped.)
- The service runs on the shared harness like every other, so it answers the universal ping
and exits on `terminate`; it was the last hand-rolled receive loop in the tree, and the
last service a shutdown had to kill rather than ask.
Publisher and subscriber must be **separate processes**: a single thread that both
published and serviced its own subscription would deadlock (its `publish` call blocks until