Files
danos/docs/m21-plan.md
T
Daniel Samson a785efa4a3 Orderly shutdown: init's stop cascade into ring-3 S5 (M21.3)
The capstone. init becomes a real supervisor: it spawns its boot
services supervised against one endpoint that also carries its signals, a
re-arming heartbeat timer, and the power events it subscribes to. On the
power button (or a terminate signal — same path) it logs the shutdown,
runs the M17 stop sequence over its children in reverse spawn order
(vfs last), then asks the power service for S5.

The acpi service honors a shutdown request from a power subscriber — init
is the one subscriber, a soft gate that stands in for 'only the system
supervisor may power off' and, unlike a PID-1 check, survives the test
harness where the kernel's idle tasks take the early ids. The power
service is mechanism (write S5); deciding when to shut down and stopping
everything else first is init's policy — the microkernel split applied to
poweroff.

The orderly-shutdown scenario injects a real QMP power-button event and
watches the whole chain compose: button pressed -> init shutting down ->
entering S5 -> QEMU powers off. That single scenario proves the M17
lifecycle and the M21 event side compose into a clean shutdown. Suite
60/60.
2026-07-13 05:56:58 +01:00

8.1 KiB

M21 execution plan: ACPI events + system power

The operational plan for the event side of the acpi service and orderly shutdown — the capstone m19-m20-plan.md previewed. Same rules as its predecessors: one phase at a time, each green before the next; this file is the build order and the checklist.

Definition of green, every phase: zig build clean, zig build test clean, python3 test/qemu_test.py passes (existing scenarios plus the phase's new one), and the relevant design doc updated. Commit per green phase (no co-author trailers). Failing cases preserve their serial logs (<case>-failed-serial.log).

Workflow: dedicated worktree; branch feat/power-events off main; auto-merge to main when the branch is green; keep the branch; push everything.

Settled decisions (2026-07-13, approved)

  1. S5 is executed by the acpi service from ring 3. No new syscall: the broad port grant (M20 decision 5) already made this physically possible — the service holds the PM1 control ports in its io grant and derives _S5 from its own namespace (aml.sleepState). Formalizing it adds no authority. The kernel keeps power.zig for its own test paths and panic-time use.
  2. The power surface is domain-named (decision 7 of the last plan): a power-protocol module + ServiceId.power = 5, registered by the acpi service — on ARM, a PSCI/mailbox service registers the same id and subscribers never know the difference. Messages: subscribe (endpoint as the call's capability, the input/manager pattern), shutdown (accepted only from PID 1 — init), and events published as buffered messages: power_button, lid, ac, battery, generic notify with a code.
  3. The service learns event ports from its own FADT copy: the kernel adds the FADT as one more memory resource on the acpi-tables node; the service tells it apart from the AML blobs by signature ("FACP" header — the blob resources are header-stripped bytecode and start with no signature). The kernel's own FADT parse is untouched.
  4. The acpi service converts to the harness (runtime.service.run): protocol messages (subscribe/shutdown), the SCI notification, and the existing report flow fold into one loop — the shape it was always meant to have.
  5. GPE/Notify correctness is proven by host unit tests (synthetic AML with a Notify inside a method body; aml.zig joins the zig build test loop). The QEMU scenario proves the power button — a fixed event, deterministically injectable via QMP system_powerdown — because QEMU cannot raise GPEs deterministically on this config. Battery/AC/lid and the embedded controller (_Qxx) are interface-complete here and validated on real hardware (the laptop) later.

Ground truth the phases build on (verified 2026-07-13)

  • system/devices/power.zig shutdown() is the kernel's S5 write (SLP_TYP|SLP_EN to PM1a/PM1b control); there is no power syscall.
  • init (system/services/init/init.zig) spawns vfs/input/device-manager fire-and-forget — no child ids kept, no signals, no event loop. The whole stop toolkit exists in runtime.process (stop/sendSignal/bindSignals).
  • test/qemu_test.py has no QMP channel (serial is a one-way file).
  • The kernel parses PM1 control blocks and SCI_INT from the FADT; the PM1 event blocks (offsets 56/60, len at 88) and GPE0/GPE1 blocks (offsets 80/84, lens 92/93) are unparsed — the service reads them from its FADT copy (decision 3).
  • The acpi-tables node carries the SCI as its only len == 1 irq resource (the broad window is len 256) — that is how the service finds it to irqBind.
  • notify_opcode = 0x86 exists in system/devices/aml/opcodes.zig but the interpreter never handles it — a GPE _Lxx body containing Notify fails evaluation today. Everything else a GPE handler needs (field access, control flow, method calls) is proven by the ring-3 _STA/_CRS work.
  • The dead-code sweep (spawned task) also edits system/devices/acpi.zig; M21.0 checks whether it landed and rebases before touching that file.

Status

  • M21.0 — baseline (dead-code sweep confirmed landed on main — no acpi.zig conflict; feat/power-events cut; QMP channel in the harness: always-on unix socket, client with the capabilities handshake, per-case qmp_after hook, and a hook-must-deliver pass gate that the smoke case now proves with a harmless query-status; suite 58/58).
  • M21.1 — SCI + the power button (kernel appends the FADT as an acpi-tables memory resource, tagged by its "FACP" header; power-protocol module + ServiceId.power = 5; the acpi service converted to runtime.service.run, registers .power, reads PM1 event/control + GPE ports from its FADT copy, enables ACPI mode if SCI_EN is clear, binds the SCI (the len-1 irq), sets PWRBTN_EN; the SCI handler clears PM1_STS, logs power: button pressed, publishes power_button, acks. Scenario power-button injects a real system_powerdown via QMP; initial-ramdisk timeout 30→60s for the service's added boot work; suite 59/59).
  • M21.2 — Notify + GPE dispatch (interpreter handles notify_opcode into a bounded queue, cleared per-evaluate, drained via takeNotifications; the service walks GPE status/enable bytes, evaluates \_GPE._Lxx/_Exx per active bit, maps notified nodes to events (battery/ac/lid/generic), clears GPE_STS write-1, acks. EC _Qxx out. Host unit test with hand-encoded AML proves the queue; aml.zig joined the zig build test loop. QEMU raises no GPEs — suite is regression net, 59/59).
  • M21.3 — orderly shutdown (init supervises its children on one endpoint that also carries signals, power events, and a re-arming heartbeat timer; on power_button or a terminate signal it logs init: shutting down, runs stop(child, 2000, endpoint) in reverse order, then requests .power shutdown; the acpi service honors shutdown from a subscriber — init is the one subscriber, a soft gate that survives testing where PID 1 isn't init — and writes SLP_TYP|SLP_EN from ring 3. orderly-shutdown scenario proves button → shutting-down → S5 → QEMU exit; suite 60/60).
  • merge feat/power-events → main, push, keep the branch — loop ends here.

Phase notes

M21.0 QMP: open the unix socket after Popen, complete the qmp_capabilities handshake, then send the hook's command (for these scenarios: {"execute": "system_powerdown"}). The socket is additive — no existing case may notice it. Note e3fe3f3 recently reworked how the harness boots; adapt to its current shape rather than the pre-rework description.

M21.1 SCI details: PM1_STS is at the event block base (write-1-to-clear); PM1_EN at base + block_len/2; PWRBTN bit is 8 in both. If PM1b exists, mirror reads/writes to both blocks. Enable ACPI mode only when SCI_EN (PM1 control bit 0) is clear — OVMF boots may already have it set. The publish path reuses the manager's subscriber table pattern (bounded, drop-on-failed-send).

M21.2 GPE walk: GPE0_STS bytes live at the GPE0 block base, GPE0_EN in the block's upper half; for a set+enabled bit n, the handler method is _L%02X (level) or _E%02X (edge) under \_GPE. Evaluate, drain the notify queue, clear the status bit, ack. A missing handler method is clear-and-log, not an error.

M21.3 ordering: init subscribes with retries — the acpi service registers .power well after init starts. The stop sequence runs vfs last (other services may flush through it). The S5 write mirrors power.zig's sleepValue (SLP_TYP bits [12:10], SLP_EN bit 13); if the write returns, log power: S5 write did not take so the scenario fails loudly instead of hanging.

Explicitly out of scope: the embedded controller and _Qxx queries, battery _BST/_BIF evaluation beyond the interface stubs, lid/AC on QEMU (no emulation), reboot over the power protocol, S3 sleep, per-device D-states (a future lifecycle-vocabulary extension), thermal zones.