13 KiB
M19–M20 execution plan: discovery migration
The operational plan for device-manager.md's increment 8: discovery leaves the kernel — a pci-bus driver (M19) and an acpi service (M20), with the kernel's device enumeration retired behind them. Same rules as m17-m18-plan.md: one phase at a time, each green before the next; this file is the build order and the checklist.
Definition of green, every phase: zig build clean, zig build test clean,
python3 test/qemu_test.py passes (existing scenarios plus the phase's new
one), and the relevant design doc updated. Commit per green phase (no co-author
trailers). The full suite is the regression net — the existing
driver-restart / usb-report / device-list / input scenarios must stay
green through the migration, which is the whole point: the system must not be
able to tell who enumerated it.
Workflow: dedicated worktree; branches off main — feat/pci-bus
(M19.0–19.3), feat/acpi-service (M20.1–20.3); auto-merge to main when a
branch is green; keep branches; push everything.
Settled decisions (2026-07-13 — veto before the loop starts)
- What "retiring the kernel scan" means. The kernel keeps, forever, the
parses it needs before user space exists: RSDP/XSDT location, MADT (SMP),
the HPET table (the tick), FADT + the AML
\_S5evaluation (poweroff — the power tests prove it), and MCFG (the host bridge node). What retires is device enumeration: the ECAM function walk (M19.3) and the DSDT/SSDT namespace walk that builds device nodes (M20.3). The AML module stays a shared build module compiled into both the kernel (for\_S5) and the acpi service (for everything else) — same source, two builds, no fork. - Bridge apertures come from the firmware memory map, not AML. Registered PCI functions carry BAR resources, and containment demands the bridge own windows that cover them. The apertures are derived kernel-side from the boot memory map's MMIO holes (regions that are neither RAM nor tables) — mechanical, AML-free, and available at boot regardless of what later moved to user space. (The bridge today carries only ECAM + bus range; this is the prerequisite M19.0 exists for.)
device_registerbecomes idempotent on exact match. A re-registration with identical (parent, class, resources) returns the existing id instead of appending. The kernel table has no unregister, so without this a restarted registering bus would duplicate its children on every respawn — idempotence makes restart-and-re-report safe for every future bus, not just PCI.- The manager matches from reports.
ChildAddedgains adevice_idfield (the kernel-registered id,no_devicefor unregistered leaves like USB ports). After the M19.3 flip, PCI driver matching keys off reported identity (the class triple) instead of the manager's boot-time snapshot — the snapshot match remains only for what the kernel still seeds. One flip phase changes both sides at once so no device is ever matched twice. - The acpi service's authority is one node. The kernel publishes an
acpi-tablesdevice: memory resources covering the table blobs plus a broadio_portresource — the documented trust grant to exactly one process (AML OperationRegions reach EC/PM ports; the claim-gated io_read/io_write calls already exist). The service claims it, maps the tables, and runs the shared AML module in ring 3 behind aHalbacked bymmio_map+io_read/io_write. - Both new processes are protocol drivers under the manager: hello, supervision, restart with backoff — all inherited from M18.1 for free. Registration idempotence (decision 3) is what makes their restarts sound.
- Firmware neutrality is the contract (2026-07-13). The generic layer is
everything at and above the device-manager protocol — descriptors,
containment, reports, matching, supervision — and none of it may become
x86-specific. Discovery is one swappable process per firmware: the acpi
service on x86; an fdt service on the Raspberry Pis (claims a
devicetree-blobnode, reports children from the flattened device tree — pure data, no bytecode, no port grant, strictly simpler than ACPI). The manager owns the tree as data and touches no hardware, ever — AML runs in a crashable, supervised discoverer precisely so a firmware-bytecode fault can never take down the supervisor. Two consequences recorded now:DeviceDescriptor's 8-bytehidcannot hold an FDTcompatiblestring ("brcm,bcm2835-aux-uart") — identity widens before the fdt service exists; and cross-firmware surfaces are named by domain, not firmware (M21 defines a power protocol, not an "ACPI events" protocol — PSCI/mailbox sources feed the same subscribers on ARM). Landed early (2026-07-13): both services exist as placeholders (system/services/acpi, system/services/ fdt) and the build's-Ddiscovery=acpi|fdtoption fills the ramdisk's neutraldiscoveryslot — the manager will spawn "discovery" by that name in M20.3 and never learn which firmware it is on.
Status
- M19.0 — prerequisites (bridge apertures from the memory map's
gaps — the single-hole rule died on OVMF's flash at the top of 4 GiB,
caught by the new every-BAR-contained assert in
discovery; idempotentdevice_registerproven inbus;ChildAdded.device_id; m17-m18-plan.md archived; suite 54/54). - M19.1 — pci-bus driver, scan only (claims the bridge, maps ECAM
through its grant, brute-force walk with the multifunction rule; the
manager matches pci_host_bridge → pci-bus per device with the full
protocol contract;
pci-scanbuilds its expected marker from the kernel's own count — equivalence on the first run; suite 55/55). - M19.2 — register + report (BAR probe mirrored byte-for-byte from the kernel's addBars so dedupe returns the kernel's node ids during coexistence; the bridge gained the io_port aperture I/O BARs need; reports carry the registered device_id; pci-scan drills a forced restart and asserts the PCI node count never grows — plus harness hardening: a failing case now preserves its serial as -failed-serial.log, and the heavy scenarios run at 150s; suite 55/55).
- M19.3 — the flip: kernel
enumeratePci/addBars/PciHeaderall deleted (bridge node stays); manager matches PCI drivers from reported identity, deduped by registered id. Surfaced and fixed a real SMP race the flip created — ring-3 device_register made the broker table concurrent, so mmio_map's lock-free read intermittently tore hpet's resource length (user fault) and overflowedr.len-1into a kernel panic; now the broker read is under the big lock and the arithmetic is guarded, and pci-bus skips size-0 BARs. discovery.md updated; suite 55/55 (driver-restart hammered 6×). - merge
feat/pci-bus→ main, push. - M20.1 — acpi service, parse only (fills the existing placeholder at
system/services/acpi/acpi.zig): kernel publishes
acpi-tables(decision 5) — memory over the table blobs, the broad io_port grant, and the SCI as an irq resource (from the FADT; unused until M21 but free to record now). The service claims it, maps the blobs, runs the shared AML module in ring 3, logs the namespace device count and_HIDs. Scenarioacpi-parse: user-space count equals the kernel walk's count. - M20.2 — register + report: namespace devices with
_HID+_CRSresources registered underacpi-tables(its io_port + the memory-map holes give containment), reported to the manager. Spawn-from-reports for ACPI matches stays off. Scenario: the reported set includes the PS/2 keyboard and mouse nodes with their IRQ resources. - M20.3 — the flip: kernel DSDT device-node building removed (static
tables +
\_S5stay, decision 1); manager matches ACPI-hid drivers (ps2-bus) from reports. Theinputanddevice-managerscenarios are the assertion. discovery.md + acpi.md + device-manager.md updated; device-manager.md increment 8 closed. - merge
feat/acpi-service→ main, push — loop ends here.
Phase notes
M19.0 apertures: the boot memory map already crosses the handoff
([boot-handoff]), but discovery never sees it today — expect a small
pass-through (kernel init hands the map to the platform layer) before the
holes computation, which belongs where the bridge node is built
(parseMcfg). Sanity-check on QEMU q35: the xHCI BAR (0xc0000000-region
values seen in the M18 logs) must land inside a derived aperture, asserted in
the kernel unit test.
M19.1 scanning without owning config access twice: the driver reads config
space through its ECAM mmio_map grant of the bridge window — the same bytes
the kernel walk read. Vendor-id 0xFFFF skip, header-type multifunction rule,
no bridge recursion (matches the kernel's current single-segment walk).
M19.2 BAR sizing: the classic size probe (write all-ones, read mask, restore) is deferred — the BARs' current programmed values and types are enough for containment-checked registration at bring-up; sizing lands with the first driver that needs to move a BAR. Log what is registered so the scenario can assert it.
M19.3 what the manager still seeds from the snapshot: everything the
kernel still enumerates (timers, ACPI nodes until M20.3). The PCI arm of
pciDriverFor switches source; driverFor doesn't move until M20.3.
M20.1 spawn and identity (pre-settled 2026-07-13): the manager spawns
discovery by its neutral ramdisk name at startup, as an ordinary protocol
driver (hello, supervision) — from M20.1 on, on every boot. For reporting ACPI
devices, ChildAdded gains hid: [8]u8 (EISA ids fit; zero = none):
firmware string identity travels beside the numeric identity field until
the FDT-driven widening replaces both (decision 7).
M20.1 Hal in ring 3: mapMmio → device.mmioMap over the claimed
acpi-tables node (plus a table-offset map for blobs); pioRead/pioWrite →
device.ioRead/ioWrite against its io_port resource. The interpreter cannot
tell it moved — that is the assertion of acpi-parse.
M20.2 containment for _CRS: io ports fall inside the node's broad
io_port resource; MMIO windows (HPET, LAPIC ranges some firmwares list) fall
inside the memory-map holes added to the node in M20.1. Anything that doesn't
fit is logged and skipped, loudly — bring-up honesty over silent drops.
M20.3 ps2 ordering: ps2-bus binds nodes the acpi service now reports, so
its spawn moves behind the report (the manager's matching handles this once
the source flips); the input scenario proves the keyboard still types.
Explicitly out of scope: PCI bridge recursion (single segment, flat bus
walk stays); BAR reprogramming/sizing; disk/PCIe hotplug; interrupt routing
changes (_PRT stays wherever it is today); the USB descriptor track;
multi-segment ECAM; per-device power states (D-states, _PSx/_PRx,
suspend/resume — a future lifecycle-vocabulary extension, since "suspend"
has the shape of a signal every driver must answer, and it has no consumer
until laptop sleep); CPU P/C-states.
M21 preview — ACPI events + system power (planned next, not in this loop)
The acpi service grows the event side (settled direction 2026-07-13; detailed phases when M20 lands):
- 21.1 SCI + fixed events: irq_bind the SCI (the resource M20.1 already records), read/clear PM1 status, publish the power-button event to subscribers (the same pub/sub shape the manager uses).
- 21.2 GPE + Notify: Notify dispatch in the shared AML interpreter, GPE
block handling,
Notify(device, code)published per reported node. The acpi service is a bus here: battery (PNP0C0A), AC (ACPI0003), and lid (PNP0C0D) nodes are reported children; small class drivers bind them and speak an evaluate/subscribe protocol to the service — the xHCI split, repeated. The embedded controller (_Qxxqueries) rides this phase; QEMU emulates no battery/EC, so those paths are interface-complete and validated on real hardware (the laptop is the win condition), while the plumbing is proven by the power button. - 21.3 the capstone: QEMU
system_powerdown→ acpi service event → init runs the M17 stop sequence over its children → kernel\_S5— orderly shutdown as the scenario that proves lifecycle + events compose. (The harness grows a QMP poke to inject the event.)