The device manager supervises: hello, backoff, and the crash-loop cap (M18.1)

The manager is now a harness service on the well-known .device_manager
endpoint. Every driver spawns supervised; drivers with an assignment must
hello (device-manager-protocol, versioned) within a deadline enforced by
a timer sweep. Exit reasons drive the restart decision: clean exits stay
down, faults restart with 300/600/1200ms backoff, and three fast deaths
mark a driver failed instead of respawning forever. usb-xhci-bus is the
first conforming driver; the crash-test fixture claims a device, hellos,
and faults on purpose — each respawn re-proving claim release on death
through the manager's own path. maximum_tasks grows 16 -> 32: the
initial-ramdisk sweep (15 binaries at once) was intermittently
overflowing the static pool.
This commit is contained in:
Daniel Samson
2026-07-13 00:19:30 +01:00
parent 36e804b848
commit 3cc1d38dd0
12 changed files with 495 additions and 72 deletions
+7 -1
View File
@@ -45,7 +45,13 @@ only when its definition of green holds.
runtime.service.run with the zero-length ping; VFS converted; `signals`
scenario; suite 51/51)
- [x] **merge** `feat/process-lifecycle` → main, push (merged 2026-07-13)
- [ ] **M18.1** — device-manager protocol: hello + restart policy (branch `feat/device-manager`)
- [x] **M18.1** — device-manager protocol: hello + restart policy
(device-manager-protocol module; the manager as a harness service:
supervised spawns, hello deadline via timer sweep, restart with
300/600/1200ms backoff, exit reasons deciding restart-vs-stopped,
crash-loop cap; usb-xhci-bus first conforming driver; crash-test fixture
re-proving claim release each respawn; `driver-restart` scenario;
maximum_tasks 16→32 — the sweep was overflowing the pool; suite 52/52)
- [ ] **merge** `feat/device-manager` → main, push
- [ ] **M18.2** — xHCI port scan + tree reports (branch `feat/usb-xhci-bus`)
- [ ] **M18.3** — app surface: enumerate/subscribe + device-list