Files
danos/docs/m21-plan.md
Daniel Samson a785efa4a3 Orderly shutdown: init's stop cascade into ring-3 S5 (M21.3)
The capstone. init becomes a real supervisor: it spawns its boot
services supervised against one endpoint that also carries its signals, a
re-arming heartbeat timer, and the power events it subscribes to. On the
power button (or a terminate signal — same path) it logs the shutdown,
runs the M17 stop sequence over its children in reverse spawn order
(vfs last), then asks the power service for S5.

The acpi service honors a shutdown request from a power subscriber — init
is the one subscriber, a soft gate that stands in for 'only the system
supervisor may power off' and, unlike a PID-1 check, survives the test
harness where the kernel's idle tasks take the early ids. The power
service is mechanism (write S5); deciding when to shut down and stopping
everything else first is init's policy — the microkernel split applied to
poweroff.

The orderly-shutdown scenario injects a real QMP power-button event and
watches the whole chain compose: button pressed -> init shutting down ->
entering S5 -> QEMU powers off. That single scenario proves the M17
lifecycle and the M21 event side compose into a clean shutdown. Suite
60/60.
2026-07-13 05:56:58 +01:00

140 lines
8.1 KiB
Markdown

# M21 execution plan: ACPI events + system power
The operational plan for the event side of the acpi service and orderly
shutdown — the capstone [m19-m20-plan.md](m19-m20-plan.md) previewed. Same
rules as its predecessors: one phase at a time, each green before the next;
this file is the build order and the checklist.
**Definition of green, every phase:** `zig build` clean, `zig build test`
clean, `python3 test/qemu_test.py` passes (existing scenarios plus the
phase's new one), and the relevant design doc updated. Commit per green phase
(no co-author trailers). Failing cases preserve their serial logs
(`<case>-failed-serial.log`).
**Workflow:** dedicated worktree; branch `feat/power-events` off `main`;
auto-merge to main when the branch is green; keep the branch; push everything.
## Settled decisions (2026-07-13, approved)
1. **S5 is executed by the acpi service from ring 3.** No new syscall: the
broad port grant (M20 decision 5) already made this physically possible —
the service holds the PM1 control ports in its io grant and derives `_S5`
from its own namespace (`aml.sleepState`). Formalizing it adds no
authority. The kernel keeps `power.zig` for its own test paths and
panic-time use.
2. **The power surface is domain-named** (decision 7 of the last plan): a
`power-protocol` module + `ServiceId.power = 5`, registered by the acpi
service — on ARM, a PSCI/mailbox service registers the same id and
subscribers never know the difference. Messages: `subscribe` (endpoint as
the call's capability, the input/manager pattern), `shutdown` (accepted
only from PID 1 — init), and events published as buffered messages:
`power_button`, `lid`, `ac`, `battery`, generic `notify` with a code.
3. **The service learns event ports from its own FADT copy**: the kernel adds
the FADT as one more memory resource on the acpi-tables node; the service
tells it apart from the AML blobs by signature ("FACP" header — the blob
resources are header-stripped bytecode and start with no signature). The
kernel's own FADT parse is untouched.
4. **The acpi service converts to the harness** (`runtime.service.run`):
protocol messages (subscribe/shutdown), the SCI notification, and the
existing report flow fold into one loop — the shape it was always meant
to have.
5. **GPE/Notify correctness is proven by host unit tests** (synthetic AML
with a Notify inside a method body; aml.zig joins the `zig build test`
loop). The QEMU scenario proves the power button — a *fixed* event,
deterministically injectable via QMP `system_powerdown` — because QEMU
cannot raise GPEs deterministically on this config. Battery/AC/lid and the
embedded controller (`_Qxx`) are interface-complete here and validated on
real hardware (the laptop) later.
## Ground truth the phases build on (verified 2026-07-13)
- `system/devices/power.zig` `shutdown()` is the kernel's S5 write
(SLP_TYP|SLP_EN to PM1a/PM1b control); there is no power syscall.
- init (`system/services/init/init.zig`) spawns vfs/input/device-manager
fire-and-forget — no child ids kept, no signals, no event loop. The whole
stop toolkit exists in `runtime.process` (stop/sendSignal/bindSignals).
- `test/qemu_test.py` has no QMP channel (serial is a one-way file).
- The kernel parses PM1 *control* blocks and SCI_INT from the FADT; the PM1
**event** blocks (offsets 56/60, len at 88) and **GPE0/GPE1** blocks
(offsets 80/84, lens 92/93) are unparsed — the service reads them from its
FADT copy (decision 3).
- The acpi-tables node carries the SCI as its only `len == 1` irq resource
(the broad window is len 256) — that is how the service finds it to
`irqBind`.
- `notify_opcode = 0x86` exists in `system/devices/aml/opcodes.zig` but the
interpreter never handles it — a GPE `_Lxx` body containing Notify fails
evaluation today. Everything else a GPE handler needs (field access,
control flow, method calls) is proven by the ring-3 `_STA`/`_CRS` work.
- The dead-code sweep (spawned task) also edits `system/devices/acpi.zig`;
M21.0 checks whether it landed and rebases before touching that file.
## Status
- [x] **M21.0** — baseline (dead-code sweep confirmed landed on main — no
acpi.zig conflict; `feat/power-events` cut; QMP channel in the harness:
always-on unix socket, client with the capabilities handshake, per-case
`qmp_after` hook, and a hook-must-deliver pass gate that the smoke case
now proves with a harmless query-status; suite 58/58).
- [x] **M21.1** — SCI + the power button (kernel appends the FADT as an
acpi-tables memory resource, tagged by its "FACP" header; `power-protocol`
module + `ServiceId.power = 5`; the acpi service converted to
`runtime.service.run`, registers `.power`, reads PM1 event/control + GPE
ports from its FADT copy, enables ACPI mode if SCI_EN is clear, binds the
SCI (the len-1 irq), sets PWRBTN_EN; the SCI handler clears PM1_STS,
logs `power: button pressed`, publishes `power_button`, acks. Scenario
`power-button` injects a real `system_powerdown` via QMP; initial-ramdisk
timeout 30→60s for the service's added boot work; suite 59/59).
- [x] **M21.2** — Notify + GPE dispatch (interpreter handles `notify_opcode`
into a bounded queue, cleared per-evaluate, drained via
`takeNotifications`; the service walks GPE status/enable bytes, evaluates
`\_GPE._Lxx`/`_Exx` per active bit, maps notified nodes to events
(battery/ac/lid/generic), clears GPE_STS write-1, acks. EC `_Qxx` out.
Host unit test with hand-encoded AML proves the queue; aml.zig joined the
`zig build test` loop. QEMU raises no GPEs — suite is regression net,
59/59).
- [x] **M21.3** — orderly shutdown (init supervises its children on one
endpoint that also carries signals, power events, and a re-arming
heartbeat timer; on `power_button` or a `terminate` signal it logs
`init: shutting down`, runs `stop(child, 2000, endpoint)` in reverse
order, then requests `.power` shutdown; the acpi service honors shutdown
from a subscriber — init is the one subscriber, a soft gate that survives
testing where PID 1 isn't init — and writes SLP_TYP|SLP_EN from ring 3.
`orderly-shutdown` scenario proves button → shutting-down → S5 → QEMU
exit; suite 60/60).
- [ ] **merge** `feat/power-events` → main, push, keep the branch — **loop
ends here**.
---
## Phase notes
**M21.0 QMP:** open the unix socket after Popen, complete the
`qmp_capabilities` handshake, then send the hook's command (for these
scenarios: `{"execute": "system_powerdown"}`). The socket is additive — no
existing case may notice it. Note e3fe3f3 recently reworked how the harness
boots; adapt to its current shape rather than the pre-rework description.
**M21.1 SCI details:** PM1_STS is at the event block base (write-1-to-clear);
PM1_EN at base + block_len/2; PWRBTN bit is 8 in both. If PM1b exists, mirror
reads/writes to both blocks. Enable ACPI mode only when SCI_EN (PM1 control
bit 0) is clear — OVMF boots may already have it set. The publish path reuses
the manager's subscriber table pattern (bounded, drop-on-failed-send).
**M21.2 GPE walk:** GPE0_STS bytes live at the GPE0 block base, GPE0_EN in
the block's upper half; for a set+enabled bit n, the handler method is
`_L%02X` (level) or `_E%02X` (edge) under `\_GPE`. Evaluate, drain the notify
queue, clear the status bit, ack. A missing handler method is clear-and-log,
not an error.
**M21.3 ordering:** init subscribes with retries — the acpi service registers
`.power` well after init starts. The stop sequence runs vfs last (other
services may flush through it). The S5 write mirrors `power.zig`'s
`sleepValue` (SLP_TYP bits [12:10], SLP_EN bit 13); if the write returns, log
`power: S5 write did not take` so the scenario fails loudly instead of
hanging.
**Explicitly out of scope:** the embedded controller and `_Qxx` queries,
battery `_BST`/`_BIF` evaluation beyond the interface stubs, lid/AC on QEMU
(no emulation), reboot over the power protocol, S3 sleep, per-device D-states
(a future lifecycle-vocabulary extension), thermal zones.