library: the harness keeps the subscribers, and an id belongs to whoever opened it

Three services had each written the same thing and got it three different
ways: input polled the process list to notice a dead subscriber, and only
when someone else subscribed; the power service never noticed at all; the
device manager noticed drivers but not subscribers. The harness owns the
table now, driven by the events a protocol declares — it registers on the
reserved verb, frames each event once, posts to everyone interested without
waiting on any of them, and reclaims a slot when the kernel says its owner
died. Interest masks moved to the envelope, so a subscriber that wants only
mice asks the same way everywhere.

Two consequences the plan had not foreseen. The device manager now hears a
supervised child's death twice, once as its supervisor and once as a
subscriber, so restart backoff counted every crash twice and gave up after
half as many; it retires the id before counting. And the kernel's published
exit table had eight slots for what is now six subscriptions in a plain
boot, so it holds sixteen.

The other half is a hole the design named early and left standing: a
backend handed out a small integer and then honoured it from anyone. A
process that guessed a file's node id read another client's file; a display
layer had no owner at all, so any client could reconfigure or destroy any
layer; a USB device token was never checked against the client that opened
it. Each is now bound to the task that opened it, and a wrong owner gets
exactly what an unknown id gets — the refusal must not become the oracle
the identical answers elsewhere were designed to remove. Closing a file
changed with it: it used to succeed unconditionally, which would have told
a caller which ids existed.

Suite 111/111, with a new case in which one process holds a file and a
layer, hands both ids to a second process, and finds them untouched after
that process has tried everything with them.
This commit is contained in:
Daniel Samson
2026-08-01 09:05:26 +01:00
parent 2719b93530
commit 1b1c587c14
29 changed files with 1072 additions and 335 deletions
@@ -81,6 +81,16 @@ one world.
| app → manager | `subscribe` (reserved verb 2) | receive published add/remove events; the subscriber's endpoint rides as the call's capability |
| manager → app | `child_added` / `child_removed` events | the same two structs, pushed rather than called |
The watcher table behind those last two rows is the **service harness's**
(`service.Subscribers`, shared with input and power), not the manager's: it
answers `subscribe`/`unsubscribe`, frames each event once for the fan-out, and
sweeps a watcher on its exit notification — where the manager previously had no
sweep for watchers at all. Its own supervised-driver exits are a different thing
and unchanged, except that a driver's death now arrives twice (the manager is
both its supervisor and a subscriber to published exits), so the manager retires
a dead driver's process id as it handles the first and the second finds nothing
to act on.
Two of what used to be the manager's own operations are the envelope's **reserved**
verbs, which mean the same thing at every provider in the system, so this protocol
numbers only three of its own (`hello` = 16, `child_added` = 17,
+12
View File
@@ -211,6 +211,18 @@ shell, a terminal, a cursor, and a wallpaper:
| `damage` | mark a region of a layer dirty |
| `present` | request a repaint: composited at the next frame-clock tick |
**A layer belongs to the client that created it.** The id is a slot in a
sixteen-entry table — small, dense, guessable — so every verb above that names one is
answered only for the task whose `create_layer` produced it, and a layer that is
somebody else's is refused exactly as one that never existed (`-ENOENT`), so a client
cannot use the refusal to learn which ids are live
([protocol-namespace.md](../os-development/protocol-namespace.md): handles are scoped
per client, validated against the badge). The compositor's own layers — the cursor
sprite and the startup self-check's pair — are marked service-owned and are created by
direct call rather than over the protocol, so no client can move or destroy the
cursor. A dead client's layers are released on its exit notification, the same sweep
the FAT server runs for open files.
Text is intentionally *not* an operation — a client renders glyphs by blitting tiles
(the [PSF font](../../system/kernel/font.psf) path the console already uses can move into a
client). Keeping the protocol to rectangles and tiles keeps the compositor small and the
+19 -9
View File
@@ -108,15 +108,25 @@ This is the async counterpart of `ipc_call`, and the input service is its first
`publishJoystickEvent`. Publishing is a short synchronous `ipc_call` the service answers at
once; the service's own fan-out is asynchronous, so publishing never blocks on a slow
subscriber.
- The **service** ([input.zig](../../system/services/input/input.zig)) keeps a small subscriber
table (endpoint handle + owning task id + `device_mask`). On `publish` it `ipc_send`s the
event to every subscriber whose mask includes the event's device class. On `subscribe` it
stores the passed capability and mask and, as housekeeping, prunes any slot whose owning
process has exited (checked against `process_enumerate`) — not for correctness (an async
send to an orphaned endpoint is harmless) but to reclaim the slot. The service runs on the
shared harness ([service.zig](../../library/kernel/service.zig)) like every other, so it
answers the universal ping and exits on `terminate`; it was the last hand-rolled receive
loop in the tree, and the last service a shutdown had to kill rather than ask.
- The **service** ([input.zig](../../system/services/input/input.zig)) owns none of that
machinery any more: the subscriber table (endpoint handle + owning task + interest mask),
the reserved `subscribe`/`unsubscribe` verbs, the fan-out, and the dead-subscriber sweep
are the shared harness's (`service.Subscribers` in
[service.zig](../../library/kernel/service.zig)), so every event stream in the system has
identical semantics. What is left in this file is what is actually about input: which
class an event belongs to, and which classes a subscriber asked for. On `publish` it names
the event's class and the harness `ipc_send`s the packet — framed once — to every
subscriber whose mask includes it.
- **A dead subscriber goes away on its exit notification**, not on a poll. The service used
to walk `process_enumerate` on every subscribe and drop slots whose owner had gone; it now
subscribes to the kernel's published exits like the FAT server and the compositor do
([process-lifecycle.md](../os-development/process-lifecycle.md)), which reclaims the slot
*and* closes the endpoint capability in it promptly rather than at the next subscribe.
(The fan-out also drops a subscriber whose `ipc_send` fails, as a backstop for a
notification a full ring dropped.)
- The service runs on the shared harness like every other, so it answers the universal ping
and exits on `terminate`; it was the last hand-rolled receive loop in the tree, and the
last service a shutdown had to kill rather than ask.
Publisher and subscriber must be **separate processes**: a single thread that both
published and serviced its own subscription would deadlock (its `publish` call blocks until
+18 -4
View File
@@ -212,10 +212,24 @@ Open-node ids live in the backend. A client that dies without closing leaks
nothing permanently: the backend (the FAT server) subscribes to the kernel's
published process-exit events (docs/process-lifecycle.md) and releases a dead
client's handles. The kernel VFS root needs no sweep at all — its node tokens
are permanent for a boot and carry no open state. Ids are plain integers, not
capabilities — a backend trusts its callers with each other's ids today, which
is acceptable while every client is part of the system image and worth
revisiting (per-client id namespaces) before third-party binaries arrive.
are permanent for a boot and carry no open state.
Ids are plain integers rather than capabilities, so the backend **scopes them
to the caller's badge**: an open node belongs to the task that opened it, and
every verb that names one — read, write, status, readdir, close — is answered
only for that task. A node id is a small number drawn from a table of
thirty-two, trivially guessable, and until this rule a backend honoured every
client's ids from every other client
(docs/os-development/protocol-namespace.md: *handles must be scoped per
client — validated against the badge, or drawn from a per-client id
namespace*).
The refusal is deliberately **identical to absence**: a node that is somebody
else's answers `-ENOENT`, exactly as one that was never opened, so a prober
learns nothing about which ids are live — the same discipline the protocol
namespace applies to a refused open. The owner is a *task*, because the badge
is: a threaded client uses a node from the thread that opened it, which is
already the granularity of the exit sweep that releases it.
## Evolution rules
+10 -3
View File
@@ -46,8 +46,12 @@ asked once at connect time rather than out of every packet's budget.
Events are published, not polled: like the input service, the service holds
subscriber endpoints as capabilities and `ipc_send`s each event as a buffered
packet, so a slow or dead subscriber can never wedge the source. **The kind is
the packet's operation** — one declared event per named kind, exactly as the
packet, so a slow or dead subscriber can never wedge the source. The table, the
reserved `subscribe`/`unsubscribe` verbs and the fan-out are the **service
harness's** (`service.Subscribers`), shared with input and the device manager, so
the acpi service's own code is the ACPI half only — and a subscriber that dies is
now swept on its exit notification, where before this service had no sweep at
all. **The kind is the packet's operation** — one declared event per named kind, exactly as the
input service delivers one per device class — so a subscriber reads *what
happened* out of the header rather than out of a tag inside the payload. The
vocabulary is hardware-neutral:
@@ -69,7 +73,10 @@ sequence over everything else. The acpi service implements this as a **soft
gate** — it honors `shutdown` only from a process that is a *subscriber*, and
init is the one subscriber. That stands in for "only the system supervisor may
power off" without hard-coding a pid, so it still holds under tests where PID 1
is not init.
is not init. The question is asked of the harness's table now
(`Subscribers.has(sender)`), which is why the harness exposes it: the gate is
unchanged, including the badge being the whole of it — the badge is
kernel-stamped, so nothing inside a packet can claim to be init.
## Orderly shutdown
+13 -3
View File
@@ -188,9 +188,15 @@ zombie state or privileged snooping:
state by all along is the id the exit event carries.
Subscription, not broadcast-to-everyone: only processes that asked receive
events, the kernel keeps a bounded subscriber table, and delivery is the same
non-blocking coalescing notification as everything else — a dying process never
waits on its mourners. Subscribing is ungated, like `process_enumerate`: what is
events, the kernel keeps a bounded subscriber table (sixteen — a normal boot
already fields six, since this is what *every* provider with per-client state
releases on), and delivery is the same non-blocking coalescing notification as
everything else — a dying process never waits on its mourners.
A service does not usually write the sweep itself: the shared service harness
subscribes for it and drops a dead task's event subscriptions
(`service.Subscribers`), and a provider adds its own handler only for state the
harness knows nothing about — open files, layers, device tokens.
Subscribing is ungated, like `process_enumerate`: what is
running (and dying) is not a secret between cooperating processes. Subscribers
do not receive the exit reason — the filesystem server does not care *why*
the client died.
@@ -285,6 +291,10 @@ callbacks (`on_terminate`, `on_reload`) for programs that want defaults.
`service` owns the `replyWait` loop and folds every event source — signals,
child exits, protocol messages — into callbacks, with the vocabulary's defaults:
it also owns the **subscriber side** of any protocol that declares events
(`service.Subscribers`: the table, the reserved `subscribe`/`unsubscribe` verbs,
the fan-out, and the sweep on a subscriber's published exit), so every event
stream in the system behaves identically.
`terminate` returns from the loop (clean exit), the common `ping` is answered automatically,
`reload` is ignored unless overridden. One loop, no locking, nothing reentrant. A
service author writes domain logic; the lifecycle contract is satisfied by the
+45 -7
View File
@@ -36,11 +36,11 @@ plain `main` checkout always tells the truth about where the work is.**
| | |
|---|---|
| Working on | **P4c** — harness subscriber lift + badge-scoped client ids |
| Branch carrying it | `feat/security-group-3` (pushed to origin) |
| On `main` | Phase 0, PM, H1, P1, P2, P3 — groups 1 and 2 merged |
| Awaiting merge | P4a, P4b — land with the group 3 merge |
| Suite | 110 cases, all passing |
| Working on | **H2** — SMEP (group 4) |
| Branch carrying it | `feat/security-group-4` (cut next) |
| On `main` | everything through P4c — groups 1, 2 and 3 merged |
| Awaiting merge | nothing |
| Suite | 111 cases, all passing |
| Last updated | 2026-08-01 |
A checkbox below means the phase met its definition of green and was
@@ -67,8 +67,19 @@ group boundary.
- [x] **merge** group 2 → main, push
- [x] **P4a** — clean protocols rebased onto `Define` (vfs, block, display, scanout, input; display's one overloaded request split per-operation and its field abuse ended, scanout's bogus 64-byte maximum deleted, directory EOF re-spelled as a nameless entry, input moved onto the service harness; new `protocol-conformance` case asks every reachable provider for `describe` and requires `-ENOSYS` for an undefined verb; suite 110/110)
- [x] **P4b** — misfit protocols rebased (device-manager, power, usb-transfer; every leading operation byte folded into the header, and with it the `device_id`/`device_token` that followed it — `Header.target` now carries the device in all three. device-manager's own `enumerate`/`subscribe` became the reserved verbs and its three `{status, reserved}` reply structs the envelope's `Status`; `ChildAdded` is one struct under two numbers, a call and an event, landing exactly on the 64-byte push floor. power's kinds became one declared event each, the input protocol's shape, so init reads *what happened* from the header; usb-transfer's control data stage moved to the packet tail in both directions, which made `Status.len` the transferred length and `actual_length` redundant. The two silent-breakage sites — init's byte-offset power parse and acpi's `message[0]` dispatch — are gone, the shutdown badge gate unchanged; three more rows in the conformance table. Suite 110/110)
- [ ] **P4c** — harness subscriber lift + badge-scoped per-client integers
- [ ] **merge** group 3 → main, push
- [x] **P4c** — harness subscriber lift + badge-scoped per-client integers (the
subscriber table, the reserved subscribe/unsubscribe verbs, the fan-out and the
dead-subscriber sweep are `service.Subscribers` now; input, acpi and
device-manager deleted three hand-rolled variants and their three different
ideas of when a subscriber goes away, standardizing on published exit
notifications — acpi had no sweep at all and input polled the process list on
every subscribe. The three guessable-id namespaces are scoped to the opening
badge: FAT node ids on every verb that names one, xHCI device tokens on open,
control, bulk and interrupt_subscribe, display layers on configure, fill, blit,
damage and destroy — each refusing a wrong owner with the *same* answer as an id
nobody holds. New `badge-scope` case, two processes of one fixture, every
refusal paired with a control; suite 111/111)
- [x] **merge** group 3 → main, push
- [ ] **H2** — SMEP on every core
- [ ] **HS** — SYSRET canonical-RIP guard
- [ ] **H3** — SMAP + boot-patched `clac`; `-cpu max` in the harness; negative tests
@@ -456,6 +467,33 @@ the wire formats will want them:*
node id and asserts refusal; kernel-side unit test for the harness sweep.
Suite 111.
*Landed. Four things the plan had not foreseen:*
- *One sweep idiom means one more kernel subscriber per provider, and the kernel's
published-exit table held **eight**. A normal boot now fields six (fat, input,
power, device-manager, display, and one per xHCI controller), so the table grew
to sixteen. It is not a table anyone notices until a service silently loses its
sweep, which is exactly the failure the old ceiling was two subscriptions away
from.*
- *The device manager hears each of its drivers die **twice** now — it is both the
supervisor its spawn named and, through the harness, a subscriber to published
exits — and the notify ring delivers the two badges separately. Untreated, one
death counted as two: the restart backoff doubled and the crash-loop cap fired
at half the deaths it names. `onDriverExit` therefore retires the dead process
id before it decides anything, and the second notification finds nothing to act
on. (The `driver-restart` and `pci-scan` drills are what would have caught it.)*
- *Refusal-equals-absence has a corollary for the verbs that **release**: FAT's
`close` used to answer 0 for an unknown node, so scoping it had to change that
too — a foreign node and a free one both answer `-ENOENT`, or the pair would
have been an oracle for which ids are live. The same applies to the harness's
`unsubscribe`.*
- *Ownership is per **task**, not per process, because the badge is: the kernel
stamps the sending thread's id, which is already the granularity of the exit
sweep that releases the state (a worker thread's death releases the handles that
worker opened). Nothing in the tree shares an id across its own threads today;
a per-process notion would need the kernel to stamp the leader, and belongs with
P5's spawner-wired namespaces if it is ever wanted.*
## H2 — SMEP
- Generalize the cpuid helper (`apic.zig:351-365`, private, subleaf-0) to