Merge feat/power-events: ACPI events + system power (M21)

The QMP harness channel, the SCI + power button published from ring 3,
Notify/GPE dispatch in the AML interpreter, and orderly shutdown — init's
M17 stop cascade into a ring-3 S5 write. Proven by injecting a real ACPI
power-button event; QEMU powers off through the whole chain.

# Conflicts:
#	system/services/init/init.zig
This commit is contained in:
Daniel Samson
2026-07-13 06:27:18 +01:00
13 changed files with 713 additions and 94 deletions
+7
View File
@@ -226,6 +226,12 @@ pub fn build(b: *std.Build) void {
}); });
runtime_module.addImport("device-manager-protocol", device_manager_protocol_module); runtime_module.addImport("device-manager-protocol", device_manager_protocol_module);
// The power protocol: system power's domain-named surface (docs/m21-plan.md).
const power_protocol_module = b.addModule("power-protocol", .{
.root_source_file = b.path("system/services/power/protocol.zig"),
});
runtime_module.addImport("power-protocol", power_protocol_module);
// Typed volatile MMIO register access + memory-ordering barriers, for drivers on // Typed volatile MMIO register access + memory-ordering barriers, for drivers on
// top of an mmio_map grant. Depends only on `builtin` (arch-conditional barriers); // top of an mmio_map grant. Depends only on `builtin` (arch-conditional barriers);
// no target set, so it inherits each driver's. See library/mmio/mmio.zig. // no target set, so it inherits each driver's. See library/mmio/mmio.zig.
@@ -577,6 +583,7 @@ pub fn build(b: *std.Build) void {
"system/devices/device-abi.zig", "system/devices/device-abi.zig",
"system/devices/pci-class.zig", // class/subclass/prog-IF name decoding "system/devices/pci-class.zig", // class/subclass/prog-IF name decoding
"system/devices/acpi-ids.zig", // _HID name decoding "system/devices/acpi-ids.zig", // _HID name decoding
"system/devices/aml/aml.zig", // AML parse + interpret, incl. Notify dispatch (M21)
"system/devices/usb-abi.zig", // wire sizes + bit packings + set-up packet encodings "system/devices/usb-abi.zig", // wire sizes + bit packings + set-up packet encodings
"system/devices/usb-ids.zig", // class/subclass/protocol code assignments "system/devices/usb-ids.zig", // class/subclass/protocol code assignments
"library/mmio/mmio.zig", // barriers assemble + registers round-trip "library/mmio/mmio.zig", // barriers assemble + registers round-trip
+4 -20
View File
@@ -189,24 +189,8 @@ suspend/resume — a future *lifecycle-vocabulary* extension, since "suspend"
has the shape of a signal every driver must answer, and it has no consumer has the shape of a signal every driver must answer, and it has no consumer
until laptop sleep); CPU P/C-states. until laptop sleep); CPU P/C-states.
## M21 preview — ACPI events + system power (planned next, not in this loop) ## M21 — ACPI events + system power — DONE
The acpi service grows the event side (settled direction 2026-07-13; detailed Built and merged (docs/m21-plan.md, 2026-07-13): the SCI + power button, Notify/GPE
phases when M20 lands): dispatch, and orderly shutdown (init's stop cascade into a ring-3 S5 write).
See that plan for the phase record.
- **21.1 SCI + fixed events**: irq_bind the SCI (the resource M20.1 already
records), read/clear PM1 status, publish the power-button event to
subscribers (the same pub/sub shape the manager uses).
- **21.2 GPE + Notify**: Notify dispatch in the shared AML interpreter, GPE
block handling, `Notify(device, code)` published per reported node. The
acpi service is a **bus** here: battery (PNP0C0A), AC (ACPI0003), and lid
(PNP0C0D) nodes are reported children; small class drivers bind them and
speak an evaluate/subscribe protocol to the service — the xHCI split,
repeated. The embedded controller (`_Qxx` queries) rides this phase;
QEMU emulates no battery/EC, so those paths are interface-complete and
validated on real hardware (the laptop is the win condition), while the
plumbing is proven by the power button.
- **21.3 the capstone**: QEMU `system_powerdown` → acpi service event → init
runs the M17 stop sequence over its children → kernel `\_S5` — orderly
shutdown as the scenario that proves lifecycle + events compose. (The
harness grows a QMP poke to inject the event.)
+31 -27
View File
@@ -70,33 +70,37 @@ auto-merge to main when the branch is green; keep the branch; push everything.
## Status ## Status
- [ ] **M21.0** — baseline: rebase over anything newly merged (the dead-code - [x] **M21.0** — baseline (dead-code sweep confirmed landed on main — no
sweep touches acpi.zig); cut `feat/power-events`; add the QMP channel to acpi.zig conflict; `feat/power-events` cut; QMP channel in the harness:
the harness (`-qmp unix:.../qmp.sock,server,nowait`, a small client with always-on unix socket, client with the capabilities handshake, per-case
the `qmp_capabilities` handshake, a per-case `qmp_after` hook that sends `qmp_after` hook, and a hook-must-deliver pass gate that the smoke case
a command N seconds after boot); existing suite stays green. now proves with a harmless query-status; suite 58/58).
- [ ] **M21.1** — SCI + the power button: kernel appends the FADT as an - [x] **M21.1** — SCI + the power button (kernel appends the FADT as an
acpi-tables memory resource; new `power-protocol` module + acpi-tables memory resource, tagged by its "FACP" header; `power-protocol`
`ServiceId.power`; the acpi service converts to the harness, registers module + `ServiceId.power = 5`; the acpi service converted to
`.power`, parses the event/GPE blocks from its FADT copy, enables ACPI `runtime.service.run`, registers `.power`, reads PM1 event/control + GPE
mode if needed (SMI dance, spin on SCI_EN), binds the SCI, sets ports from its FADT copy, enables ACPI mode if SCI_EN is clear, binds the
PWRBTN_EN; on SCI reads/clears PM1_STS and publishes `power_button` SCI (the len-1 irq), sets PWRBTN_EN; the SCI handler clears PM1_STS,
(log: `power: button pressed`), always irqAck. Scenario `power-button`: logs `power: button pressed`, publishes `power_button`, acks. Scenario
`qmp_after system_powerdown` → expect the log line. `power-button` injects a real `system_powerdown` via QMP; initial-ramdisk
- [ ] **M21.2** — Notify + GPE dispatch: interpreter handles `notify_opcode` timeout 30→60s for the service's added boot work; suite 59/59).
into a bounded queue drained after evaluate(); on GPE status bits the - [x] **M21.2** — Notify + GPE dispatch (interpreter handles `notify_opcode`
service evaluates `\_GPE._Lxx`/`_Exx`, maps notified nodes to events into a bounded queue, cleared per-evaluate, drained via
(PNP0C0A→battery, ACPI0003→ac, PNP0C0D→lid, else generic), clears `takeNotifications`; the service walks GPE status/enable bytes, evaluates
GPE_STS, acks. EC `_Qxx` explicitly out (hardware track). Host unit `\_GPE._Lxx`/`_Exx` per active bit, maps notified nodes to events
tests for Notify in aml.zig; aml.zig joins the `zig build test` loop. (battery/ac/lid/generic), clears GPE_STS write-1, acks. EC `_Qxx` out.
- [ ] **M21.3** — orderly shutdown: init keeps child ids (spawnSupervised + Host unit test with hand-encoded AML proves the queue; aml.zig joined the
exit endpoint), binds signals, subscribes to `.power`; on `power_button` `zig build test` loop. QEMU raises no GPEs — suite is regression net,
logs `init: shutting down`, runs `stop(child, 2000, endpoint)` in 59/59).
reverse spawn order, then sends `shutdown` to `.power`; the acpi service - [x] **M21.3** — orderly shutdown (init supervises its children on one
(sender PID 1 only) logs `power: entering S5` and writes SLP_TYP|SLP_EN endpoint that also carries signals, power events, and a re-arming
from ring 3. Scenario `orderly-shutdown`: boot via init, `qmp_after heartbeat timer; on `power_button` or a `terminate` signal it logs
system_powerdown`, ordered regex button→shutting-down→entering-S5, pass `init: shutting down`, runs `stop(child, 2000, endpoint)` in reverse
on QEMU exit. Docs + memory updated. order, then requests `.power` shutdown; the acpi service honors shutdown
from a subscriber — init is the one subscriber, a soft gate that survives
testing where PID 1 isn't init — and writes SLP_TYP|SLP_EN from ring 3.
`orderly-shutdown` scenario proves button → shutting-down → S5 → QEMU
exit; suite 60/60).
- [ ] **merge** `feat/power-events` → main, push, keep the branch — **loop - [ ] **merge** `feat/power-events` → main, push, keep the branch — **loop
ends here**. ends here**.
+3
View File
@@ -20,6 +20,9 @@ pub const vfs_protocol = @import("vfs-protocol");
/// The device-manager protocol: hello + tree reports (docs/device-manager.md). /// The device-manager protocol: hello + tree reports (docs/device-manager.md).
pub const device_manager_protocol = @import("device-manager-protocol"); pub const device_manager_protocol = @import("device-manager-protocol");
/// The power protocol: events (button, lid, battery) + shutdown (docs/m21-plan.md).
pub const power_protocol = @import("power-protocol");
/// Keyboard-event listening (subscribe/next) and broadcasting (publish), over the input /// Keyboard-event listening (subscribe/next) and broadcasting (publish), over the input
/// service. See library/runtime/input.zig and system/services/input/. /// service. See library/runtime/input.zig and system/services/input/.
pub const input = @import("input.zig"); pub const input = @import("input.zig");
+1
View File
@@ -177,6 +177,7 @@ pub const ServiceId = enum(u32) {
input = 2, input = 2,
ps2_bus = 3, // the 8042 owner; child device drivers attach here for raw bytes ps2_bus = 3, // the 8042 owner; child device drivers attach here for raw bytes
device_manager = 4, // the tree, the matcher, the supervisor (docs/device-manager.md) device_manager = 4, // the tree, the matcher, the supervisor (docs/device-manager.md)
power = 5, // system power: events (button, lid, battery) + shutdown (docs/m21-plan.md; domain-named per decision 7 — the acpi service registers it on x86, a PSCI service will on ARM)
_, _,
}; };
+14
View File
@@ -157,6 +157,13 @@ pub var namespace: ?aml.Namespace = null;
/// Physical address of the DSDT the FADT points at, or 0. /// Physical address of the DSDT the FADT points at, or 0.
pub var dsdt_physical: u64 = 0; pub var dsdt_physical: u64 = 0;
/// The FADT itself (physical + length), published on the acpi-tables node so
/// the ring-3 acpi service can read the PM1 event and GPE blocks it needs for
/// the event side (docs/m21-plan.md decision 3). Distinguished from the AML
/// blob resources by its intact "FACP" header — the blobs are header-stripped.
var fadt_physical: u64 = 0;
var fadt_length: u64 = 0;
// AML blocks (DSDT + any SSDTs) collected during the table walk, as physical // AML blocks (DSDT + any SSDTs) collected during the table walk, as physical
// address + length of each table's post-header bytecode. Scanned after the walk // address + length of each table's post-header bytecode. Scanned after the walk
// for the sleep-state (`_Sx`) packages. // for the sleep-state (`_Sx`) packages.
@@ -373,6 +380,8 @@ pub fn discover(rsdp_physical: u64, memory_regions: []const boot_handoff.MemoryR
// Start clean so a re-run doesn't accumulate stale state. // Start clean so a re-run doesn't accumulate stale state.
power_information = .{}; power_information = .{};
fadt_physical = 0;
fadt_length = 0;
platform_information = .{}; platform_information = .{};
aml_stats = .{}; aml_stats = .{};
namespace = null; namespace = null;
@@ -447,6 +456,9 @@ fn publishAcpiTablesNode(device_tree: *DeviceTree) !void {
// SCI (recorded first, len 1) stays distinct so M21 can pick it out. // SCI (recorded first, len 1) stays distinct so M21 can pick it out.
if (power_information.sci_interrupt != 0) _ = node.addResource(.irq, power_information.sci_interrupt, 1); if (power_information.sci_interrupt != 0) _ = node.addResource(.irq, power_information.sci_interrupt, 1);
_ = node.addResource(.irq, 0, 256); _ = node.addResource(.irq, 0, 256);
// The FADT rides along (M21): the service reads the PM1 event / GPE blocks
// from its own copy, telling it apart from the AML blobs by signature.
if (fadt_physical != 0) _ = node.addResource(.memory, fadt_physical, fadt_length);
} }
/// The number of Device objects in the namespace built during discovery, or 0. /// The number of Device objects in the namespace built during discovery, or 0.
@@ -482,6 +494,8 @@ fn handleTable(device_tree: *DeviceTree, hal: Hal, sdt_physical: u64) !void {
} else if (std.mem.eql(u8, &sig, &HPET)) { } else if (std.mem.eql(u8, &sig, &HPET)) {
try parseHpet(device_tree, hal, header); try parseHpet(device_tree, hal, header);
} else if (std.mem.eql(u8, &sig, &FACP)) { } else if (std.mem.eql(u8, &sig, &FACP)) {
fadt_physical = sdt_physical;
fadt_length = header.length;
parseFadt(header); parseFadt(header);
} else if (std.mem.eql(u8, &sig, &SPCR)) { } else if (std.mem.eql(u8, &sig, &SPCR)) {
parseSpcr(header); parseSpcr(header);
+28
View File
@@ -230,3 +230,31 @@ test "interpreter runs a method with args, arithmetic, and control flow" {
const lo = try interpreter.evaluate(tst, &.{.{ .integer = 2 }}); // 2+5=7 !> 10 -> 0 const lo = try interpreter.evaluate(tst, &.{.{ .integer = 2 }}); // 2+5=7 !> 10 -> 0
try std.testing.expectEqual(@as(u64, 0), try lo.asInteger()); try std.testing.expectEqual(@as(u64, 0), try lo.asInteger());
} }
test "interpreter records Notify(device, code)" {
// Device(DEV_) { Name(_HID, 0x030AD041) } // PNP0A03-ish placeholder
// Method(TST_, 0) { Notify(DEV_, 0x80); Return(Zero) }
// Encoded: a Device holding a Name, then a Method issuing Notify on it.
const blob = [_]u8{
0x5B, 0x82, 0x0F, 0x44, 0x45, 0x56, 0x5F, // Device(DEV_) len=0x0F (pkglen + DEV_ + Name)
0x08, 0x5F, 0x48, 0x49, 0x44, 0x0C, 0x41, 0xD0, 0x0A, 0x03, // Name(_HID, DWord 0x030AD041)
0x14, 0x0F, 0x54, 0x53, 0x54, 0x5F, 0x00, // Method(TST_, 0) len=0x0F (pkglen + TST_ + flags + body)
0x86, 0x44, 0x45, 0x56, 0x5F, 0x0A, 0x80, // Notify(DEV_, 0x80)
0xA4, 0x00, // Return(Zero)
};
var arena = std.heap.ArenaAllocator.init(std.testing.allocator);
defer arena.deinit();
var result = try parse(arena.allocator(), &.{&blob});
const namespace = &result.namespace;
const tst = namespace.resolve(namespace.root, false, 0, &.{.{ 'T', 'S', 'T', '_' }}) orelse return error.NoMethod;
const dev = namespace.resolve(namespace.root, false, 0, &.{.{ 'D', 'E', 'V', '_' }}) orelse return error.NoDevice;
var interpreter = Interpreter.init(namespace, .{ .mapMmio = noMap, .pioRead = noRead, .pioWrite = noWrite }, arena.allocator());
_ = try interpreter.evaluate(tst, &.{});
const events = interpreter.takeNotifications();
try std.testing.expectEqual(@as(usize, 1), events.len);
try std.testing.expectEqual(dev, events[0].node);
try std.testing.expectEqual(@as(u64, 0x80), events[0].code);
}
+41
View File
@@ -141,6 +141,9 @@ const Frame = struct {
/// A CreateField binding: a name that indexes into a buffer object. /// A CreateField binding: a name that indexes into a buffer object.
const BufferField = struct { buffer: *Node, byte_off: usize, bit_width: u32 }; const BufferField = struct { buffer: *Node, byte_off: usize, bit_width: u32 };
/// One Notify(device, code) the interpreter executed.
pub const NotifyEvent = struct { node: *Node, code: u64 };
pub const Interpreter = struct { pub const Interpreter = struct {
namespace: *Namespace, namespace: *Namespace,
hal: Hal, hal: Hal,
@@ -149,6 +152,11 @@ pub const Interpreter = struct {
dynamic_overrides: std.AutoHashMapUnmanaged(*Node, Object) = .{}, dynamic_overrides: std.AutoHashMapUnmanaged(*Node, Object) = .{},
/// CreateField bindings active for the current evaluation. /// CreateField bindings active for the current evaluation.
fields: std.AutoHashMapUnmanaged(*Node, BufferField) = .{}, fields: std.AutoHashMapUnmanaged(*Node, BufferField) = .{},
/// Notify(device, code) operations the last evaluation executed — a GPE or
/// EC handler tells the OS "look at this device" this way. Bounded; the
/// caller drains it with `takeNotifications` after `evaluate` (M21).
notify_queue: [16]NotifyEvent = undefined,
notify_count: usize = 0,
pub fn init(namespace: *Namespace, hal: Hal, arena: std.mem.Allocator) Interpreter { pub fn init(namespace: *Namespace, hal: Hal, arena: std.mem.Allocator) Interpreter {
return .{ .namespace = namespace, .hal = hal, .arena = arena }; return .{ .namespace = namespace, .hal = hal, .arena = arena };
@@ -157,6 +165,7 @@ pub const Interpreter = struct {
/// Evaluate a namespace object: invoke a Method, read a Name's value, or read a /// Evaluate a namespace object: invoke a Method, read a Name's value, or read a
/// Field. Resets per-evaluation runtime state first. /// Field. Resets per-evaluation runtime state first.
pub fn evaluate(self: *Interpreter, node: *Node, args: []const Object) Error!Object { pub fn evaluate(self: *Interpreter, node: *Node, args: []const Object) Error!Object {
self.notify_count = 0;
self.dynamic_overrides.clearRetainingCapacity(); self.dynamic_overrides.clearRetainingCapacity();
self.fields.clearRetainingCapacity(); self.fields.clearRetainingCapacity();
return self.invoke(node, args); return self.invoke(node, args);
@@ -267,6 +276,8 @@ pub const Interpreter = struct {
}, },
opcode.to_buffer_opcode => try self.passThroughUnary(current, frame), opcode.to_buffer_opcode => try self.passThroughUnary(current, frame),
opcode.notify_opcode => try self.notify(current, frame),
opcode.extended_opcode_prefix => try self.ext(current, frame), opcode.extended_opcode_prefix => try self.ext(current, frame),
// CreateXField: source, index, name (bit widths differ by op) // CreateXField: source, index, name (bit widths differ by op)
@@ -542,6 +553,36 @@ pub const Interpreter = struct {
try self.storeInto(current, frame, value); try self.storeInto(current, frame, value);
} }
/// Notify(SuperName, NotifyValue): resolve the named device, evaluate the
/// code, and record the pair for the caller to dispatch. AML control flow
/// continues (Notify returns nothing).
fn notify(self: *Interpreter, current: *Cursor, frame: *Frame) Error!Object {
const lead = current.peek() orelse return error.Truncated;
var target: ?*Node = null;
if (isNameStart(lead)) {
const name_path = try current.nameString();
target = self.namespace.resolve(frame.scope, name_path.rooted, name_path.parents, name_path.slice());
} else {
// A non-name SuperName (Local/Arg holding a reference).
const obj = try self.term(current, frame);
if (obj == .reference) target = obj.reference;
}
const code = try self.evaluateInteger(current, frame);
if (target) |node| {
if (self.notify_count < self.notify_queue.len) {
self.notify_queue[self.notify_count] = .{ .node = node, .code = code };
self.notify_count += 1;
}
}
return .uninitialized;
}
/// The Notify events the last `evaluate` produced. Valid until the next
/// `evaluate` clears the queue.
pub fn takeNotifications(self: *Interpreter) []const NotifyEvent {
return self.notify_queue[0..self.notify_count];
}
fn storeInto(self: *Interpreter, current: *Cursor, frame: *Frame, value: Object) Error!void { fn storeInto(self: *Interpreter, current: *Cursor, frame: *Frame, value: Object) Error!void {
const lead = current.peek() orelse return error.Truncated; const lead = current.peek() orelse return error.Truncated;
if (isNameStart(lead)) { if (isNameStart(lead)) {
+25
View File
@@ -152,6 +152,10 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
acpiReportTest(boot_information); acpiReportTest(boot_information);
} else if (eql(case, "acpi-ps2")) { } else if (eql(case, "acpi-ps2")) {
acpiReportTest(boot_information); // same spawn; the harness regex differs acpiReportTest(boot_information); // same spawn; the harness regex differs
} else if (eql(case, "power-button")) {
acpiReportTest(boot_information); // boot the manager (spawns the acpi service); harness injects the button
} else if (eql(case, "orderly-shutdown")) {
orderlyShutdownTest(boot_information);
} else if (eql(case, "initial-ramdisk")) { } else if (eql(case, "initial-ramdisk")) {
initialRamdiskTest(boot_information); initialRamdiskTest(boot_information);
} else if (eql(case, "vfs")) { } else if (eql(case, "vfs")) {
@@ -1929,6 +1933,27 @@ fn pciScanTest(boot_information: *const BootInformation) void {
result(); result();
} }
/// M21.3 capstone: orderly shutdown. Boot init with the initial-ramdisk
/// published, so init spawns the full service tree (vfs, input, device-manager
/// -> discovery/acpi); the harness injects a real power-button event via QMP;
/// the acpi service publishes it; init runs the stop sequence over its children
/// and asks the power service for S5; the machine powers off (QEMU exits). The
/// kernel test only spawns init — the ordered chain is the harness assertion.
fn orderlyShutdownTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: orderly-shutdown\n", .{});
if (boot_information.init_len == 0 or boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over init and the initial_ramdisk", false);
result();
return;
}
const ramdisk = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
process.setInitialRamdisk(ramdisk);
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len];
const spawned = if (process.spawnProcess(image, 4, &.{"/system/services/init"})) true else |_| false;
check("init spawned as PID root of user space", spawned);
result();
}
/// M20.2: the acpi service registers + reports its _HID devices. Boot normally /// M20.2: the acpi service registers + reports its _HID devices. Boot normally
/// (the manager spawns discovery); the harness's expect regex requires the two /// (the manager spawns discovery); the harness's expect regex requires the two
/// PS/2 nodes among the service's report lines, each with its _CRS resources — /// PS/2 nodes among the service's report lines, each with its _CRS resources —
+323 -32
View File
@@ -4,13 +4,11 @@
//! grant, a broad irq window, the SCI), and runs the **shared AML module** in //! grant, a broad irq window, the SCI), and runs the **shared AML module** in
//! ring 3 — the same parser and interpreter the kernel uses. //! ring 3 — the same parser and interpreter the kernel uses.
//! //!
//! M20.2 (this increment): after parsing, walk the namespace and, for each //! It also owns the **event side** (M21): it registers the domain-named `.power`
//! present Device with a hardware id (`_HID`), evaluate its current resource //! service, binds the SCI (System Control Interrupt), and on a power-button
//! settings (`_CRS`) through a ring-3 `Hal` (port I/O over the claimed node), //! fixed event publishes `power_button` to subscribers — and on init's request
//! register it under the acpi-tables node (its I/O ports and IRQs contained by //! writes S5 to power the machine off. The device discovery (M20) and the event
//! the node's broad grants), and report it to the device manager with its //! handling both run in one `runtime.service.run` loop.
//! EISA-decoded hid as identity. Matching those reports to drivers (ps2-bus)
//! and retiring the kernel's own device build follow in M20.3.
const std = @import("std"); const std = @import("std");
const runtime = @import("runtime"); const runtime = @import("runtime");
@@ -18,6 +16,7 @@ const aml = @import("aml");
const acpi_ids = @import("acpi-ids"); const acpi_ids = @import("acpi-ids");
const device = runtime.device; const device = runtime.device;
const protocol = runtime.device_manager_protocol; const protocol = runtime.device_manager_protocol;
const power = runtime.power_protocol;
/// AML opcode/prefix bytes by name (`zero_opcode`, `byte_prefix`, …) — so the `_HID` /// AML opcode/prefix bytes by name (`zero_opcode`, `byte_prefix`, …) — so the `_HID`
/// integer decode names the opcodes instead of bare 0x0A/0x0B/… (docs/coding-standards.md). /// integer decode names the opcodes instead of bare 0x0A/0x0B/… (docs/coding-standards.md).
const opcodes = aml.opcodes; const opcodes = aml.opcodes;
@@ -31,6 +30,43 @@ fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
// window — the Hal routes every port access through this one claim. // window — the Hal routes every port access through this one claim.
var node_id: u64 = 0; var node_id: u64 = 0;
var io_resource_index: u64 = 0; var io_resource_index: u64 = 0;
// The SCI's irq resource index on the node (the len-1 irq, distinct from the
// broad [0,256) window), for irqBind / irqAck.
var sci_resource_index: u64 = 0;
var has_sci = false;
// PM1 event/control and GPE register ports, read from the FADT copy the kernel
// publishes on the node (M21). Port 0 means absent.
var pm1a_evt: u16 = 0;
var pm1b_evt: u16 = 0;
var pm1_evt_len: u8 = 0;
var pm1a_cnt: u16 = 0;
var pm1b_cnt: u16 = 0;
var gpe0_blk: u16 = 0;
var gpe0_len: u8 = 0;
var gpe1_blk: u16 = 0;
var gpe1_len: u8 = 0;
var smi_cmd: u16 = 0;
var acpi_enable_value: u8 = 0;
var s5_slp_typ_a: u8 = 0;
var s5_slp_typ_b: u8 = 0;
var s5_valid = false;
// PM1 event-register bits (ACPI): PWRBTN in the status/enable word is bit 8;
// the control word's SCI_EN is bit 0; SLP_EN is bit 13.
const pwrbtn_bit: u16 = 1 << 8;
const sci_en_bit: u32 = 1 << 0;
const slp_en: u32 = 1 << 13;
// The `.power` subscribers: endpoints handed over as capabilities, each
// receiving events as buffered messages. Dropped on a failed send. The
// subscriber's task id is kept too — a shutdown request is honored only from a
// subscriber (init subscribes; a stray process does not), the soft gate that
// stands in for "only the system supervisor may power off" without hardcoding
// a pid the kernel's idle tasks would have taken.
const maximum_subscribers = 8;
var subscribers: [maximum_subscribers]?runtime.ipc.Handle = .{null} ** maximum_subscribers;
var subscriber_tasks: [maximum_subscribers]u32 = .{0} ** maximum_subscribers;
// Pass-1 registration record (see main): what pass 2 reports. // Pass-1 registration record (see main): what pass 2 reports.
const Registered = struct { hid: [8]u8 = .{0} ** 8, hid_len: usize = 0, device_id: u64 = 0, resource_count: u64 = 0 }; const Registered = struct { hid: [8]u8 = .{0} ** 8, hid_len: usize = 0, device_id: u64 = 0, resource_count: u64 = 0 };
@@ -85,22 +121,34 @@ pub fn main(init: runtime.process.Init) void {
return; return;
} }
// Map each memory resource (an AML blob) and note the io_port resource. // Map the node's resources: the AML blobs (bytecode), the FADT (intact
// "FACP" header — decision 3), the io_port grant, and the SCI irq.
var blocks: [8][]const u8 = undefined; var blocks: [8][]const u8 = undefined;
var block_count: usize = 0; var block_count: usize = 0;
var found_io = false; var found_io = false;
var fadt: ?[]const u8 = null;
for (node.resources[0..@intCast(node.resource_count)], 0..) |resource, index| { for (node.resources[0..@intCast(node.resource_count)], 0..) |resource, index| {
if (resource.kind == @intFromEnum(device.ResourceKind.io_port) and !found_io) { if (resource.kind == @intFromEnum(device.ResourceKind.io_port) and !found_io) {
io_resource_index = index; io_resource_index = index;
found_io = true; found_io = true;
continue; continue;
} }
if (resource.kind == @intFromEnum(device.ResourceKind.irq) and resource.len == 1) {
sci_resource_index = index;
has_sci = true;
continue;
}
if (resource.kind != @intFromEnum(device.ResourceKind.memory)) continue; if (resource.kind != @intFromEnum(device.ResourceKind.memory)) continue;
const base = device.mmioMap(node_id, index) orelse continue; const base = device.mmioMap(node_id, index) orelse continue;
const pointer: [*]const u8 = @ptrFromInt(base); const pointer: [*]const u8 = @ptrFromInt(base);
blocks[block_count] = pointer[0..@intCast(resource.len)]; const bytes = pointer[0..@intCast(resource.len)];
if (bytes.len >= 4 and std.mem.eql(u8, bytes[0..4], "FACP")) {
fadt = bytes;
continue;
}
if (block_count == blocks.len) continue;
blocks[block_count] = bytes;
block_count += 1; block_count += 1;
if (block_count == blocks.len) break;
} }
if (block_count == 0) { if (block_count == 0) {
_ = runtime.system.write("/system/services/acpi: no AML blobs on the node\n"); _ = runtime.system.write("/system/services/acpi: no AML blobs on the node\n");
@@ -124,29 +172,45 @@ pub fn main(init: runtime.process.Init) void {
while (true) runtime.system.sleep(1000); while (true) runtime.system.sleep(1000);
} }
// Register + report the present _HID devices (M20.2). // Register + report the present _HID devices (M20), then set up the power
var arena = std.heap.ArenaAllocator.init(runtime.allocator()); // event side (M21), then serve — all in one harness loop. The interpreter
var interpreter = aml.Interpreter.init(&namespace, .{ // and namespace outlive this frame (static), so the harness callbacks can
// reach them.
interpreter_arena = std.heap.ArenaAllocator.init(runtime.allocator());
persistent_namespace = namespace;
global_interpreter = aml.Interpreter.init(&persistent_namespace, .{
.mapMmio = halMapMmio, .mapMmio = halMapMmio,
.pioRead = halPioRead, .pioRead = halPioRead,
.pioWrite = halPioWrite, .pioWrite = halPioWrite,
}, arena.allocator()); }, interpreter_arena.allocator());
// Pass 1: register every present _HID device under acpi-tables, remembering readFadt(fadt);
// each (hid, device id). Pass 2: report them all. Registering before any s5_valid = readSleepS5(&persistent_namespace);
// report reaches the manager means a driver it spawns on the first report
// already sees the whole set (no keyboard-before-mouse race for ps2-bus). runtime.service.run(power.message_maximum, .{
.service = .power,
.init = onInit,
.on_message = onMessage,
.on_notification = onNotification,
});
}
// Static so the harness callbacks (which run after main's stack frame is gone)
// can reach the namespace and interpreter.
var persistent_namespace: aml.Namespace = undefined;
var global_interpreter: aml.Interpreter = undefined;
var interpreter_arena: std.heap.ArenaAllocator = undefined;
/// Startup under the harness: register + report the discovered devices to the
/// manager (M20), then enable ACPI mode and arm the power button (M21).
fn onInit(endpoint: runtime.ipc.Handle) bool {
registered_count = 0; registered_count = 0;
walkDevices(namespace.root, &interpreter); walkDevices(persistent_namespace.root, &global_interpreter);
const manager = runtime.ipc.lookup(.device_manager); const manager = runtime.ipc.lookup(.device_manager);
var i: usize = 0; var i: usize = 0;
while (i < registered_count) : (i += 1) { while (i < registered_count) : (i += 1) {
const entry = registered[i]; const entry = registered[i];
// Append the _HID's human-readable name when it is a known standard PnP/ACPI
// id (e.g. PNP0303 -> "PS/2 Keyboard"), so the boot log says what each
// reported device actually is. The description trails the existing fields so
// the acpi-report/acpi-ps2 matchers still see "<hid> (device N, M resources)".
const hid = entry.hid[0..entry.hid_len]; const hid = entry.hid[0..entry.hid_len];
const desc = acpi_ids.description(hid); const desc = acpi_ids.description(hid);
if (desc.len != 0) if (desc.len != 0)
@@ -154,12 +218,7 @@ pub fn main(init: runtime.process.Init) void {
else else
writeLine("/system/services/acpi: reported {s} (device {d}, {d} resources)\n", .{ hid, entry.device_id, entry.resource_count }); writeLine("/system/services/acpi: reported {s} (device {d}, {d} resources)\n", .{ hid, entry.device_id, entry.resource_count });
if (manager) |h| { if (manager) |h| {
var report = protocol.ChildAdded{ var report = protocol.ChildAdded{ .parent = node_id, .bus_address = entry.device_id, .identity = 0, .device_id = entry.device_id };
.parent = node_id,
.bus_address = entry.device_id,
.identity = 0,
.device_id = entry.device_id,
};
@memcpy(report.hid[0..entry.hid_len], entry.hid[0..entry.hid_len]); @memcpy(report.hid[0..entry.hid_len], entry.hid[0..entry.hid_len]);
var reply: [protocol.message_maximum]u8 = undefined; var reply: [protocol.message_maximum]u8 = undefined;
_ = runtime.ipc.call(h, std.mem.asBytes(&report), &reply) catch {}; _ = runtime.ipc.call(h, std.mem.asBytes(&report), &reply) catch {};
@@ -167,9 +226,241 @@ pub fn main(init: runtime.process.Init) void {
} }
writeLine("/system/services/acpi: reported {d} device(s) to the manager\n", .{registered_count}); writeLine("/system/services/acpi: reported {d} device(s) to the manager\n", .{registered_count});
// Stay resident: the claim holds, and the service is here to grow into the armPowerButton(endpoint);
// supervised discoverer (M20.3, then the M21 event side on the SCI). return true;
while (true) runtime.system.sleep(1000); }
// --- power event side (M21) ---------------------------------------------------
/// Read the PM1 event/control and GPE register ports plus the SMI enable pair
/// from the FADT copy on the node. Offsets are from the FADT table start (the
/// SDT header is the first 36 bytes). Prefers the 32-bit port fields; QEMU's
/// FADT populates them.
fn readFadt(fadt: ?[]const u8) void {
const f = fadt orelse {
_ = runtime.system.write("acpi: no FADT on the node — power events off\n");
return;
};
smi_cmd = @truncate(rd32(f, 48));
acpi_enable_value = f[52];
pm1a_evt = @truncate(rd32(f, 56));
pm1b_evt = @truncate(rd32(f, 60));
pm1a_cnt = @truncate(rd32(f, 64));
pm1b_cnt = @truncate(rd32(f, 68));
gpe0_blk = @truncate(rd32(f, 80));
gpe1_blk = @truncate(rd32(f, 84));
pm1_evt_len = if (f.len > 88) f[88] else 4;
gpe0_len = if (f.len > 92) f[92] else 0;
gpe1_len = if (f.len > 93) f[93] else 0;
}
fn readSleepS5(ns: *aml.Namespace) bool {
const st = aml.sleepState(ns, 5) orelse return false;
s5_slp_typ_a = st.slp_typ_a;
s5_slp_typ_b = st.slp_typ_b;
return true;
}
/// Enable ACPI mode if the firmware isn't already in it, then bind the SCI and
/// set PWRBTN_EN so the power button raises an interrupt we can see.
fn armPowerButton(endpoint: runtime.ipc.Handle) void {
if (pm1a_cnt != 0 and (halPioRead(2, pm1a_cnt) & sci_en_bit) == 0 and smi_cmd != 0) {
// Switch to ACPI mode: write ACPI_ENABLE to the SMI command port, then
// spin (bounded) until SCI_EN latches.
halPioWrite(1, smi_cmd, acpi_enable_value);
var tries: u32 = 0;
while (tries < 1000 and (halPioRead(2, pm1a_cnt) & sci_en_bit) == 0) : (tries += 1) {
runtime.system.sleep(1);
}
}
if (!has_sci) {
_ = runtime.system.write("acpi: no SCI resource — power button unavailable\n");
return;
}
if (!device.irqBind(node_id, sci_resource_index, endpoint)) {
_ = runtime.system.write("acpi: SCI irq_bind failed\n");
return;
}
// PWRBTN_EN lives in the PM1 enable register at evt_blk + evt_len/2.
if (pm1a_evt != 0) {
const en_port = pm1a_evt + pm1_evt_len / 2;
halPioWrite(2, en_port, @as(u16, @truncate(halPioRead(2, en_port))) | pwrbtn_bit);
}
if (pm1b_evt != 0) {
const en_port = pm1b_evt + pm1_evt_len / 2;
halPioWrite(2, en_port, @as(u16, @truncate(halPioRead(2, en_port))) | pwrbtn_bit);
}
_ = runtime.system.write("acpi: power button armed\n");
}
/// The SCI fired. Read PM1 status; a set PWRBTN_STS is the power button — clear
/// it (write-1), publish, log. Any other set status is cleared and logged
/// (GPE/Notify dispatch is M21.2). Always re-arm the line.
fn onSci() void {
var handled = false;
inline for (.{ pm1a_evt, pm1b_evt }) |evt_port| {
if (evt_port != 0) {
const sts: u16 = @truncate(halPioRead(2, evt_port));
if (sts & pwrbtn_bit != 0) {
halPioWrite(2, evt_port, pwrbtn_bit); // write-1-to-clear
handled = true;
} else if (sts != 0) {
halPioWrite(2, evt_port, sts); // clear whatever else latched
}
}
}
if (handled) {
_ = runtime.system.write("power: button pressed\n");
publishButton();
}
handleGpe();
_ = device.irqAck(node_id, sci_resource_index);
}
/// General-purpose events: for each set+enabled GPE bit, evaluate its `\_GPE`
/// handler method (`_Lxx` level / `_Exx` edge), drain the Notify queue the
/// method produced, and publish an event per notified device. Then clear the
/// status bit. QEMU raises no GPEs on this config, so this path is exercised by
/// host unit tests (docs/m21-plan.md decision 5); on real hardware it carries
/// battery/AC/lid. The embedded controller's `_Qxx` queries are out of scope.
fn handleGpe() void {
handleGpeBlock(gpe0_blk, gpe0_len, 0);
handleGpeBlock(gpe1_blk, gpe1_len, gpe0_len * 4);
}
fn handleGpeBlock(blk: u16, len: u8, gpe_base: u32) void {
if (blk == 0 or len == 0) return;
const status_bytes = len / 2; // status half, then enable half
var byte_index: u8 = 0;
while (byte_index < status_bytes) : (byte_index += 1) {
const sts: u8 = @truncate(halPioRead(1, blk + byte_index));
const en: u8 = @truncate(halPioRead(1, blk + status_bytes + byte_index));
const active = sts & en;
if (active == 0) continue;
var bit: u3 = 0;
while (true) : (bit += 1) {
if (active & (@as(u8, 1) << bit) != 0) {
dispatchGpe(gpe_base + @as(u32, byte_index) * 8 + bit);
}
if (bit == 7) break;
}
halPioWrite(1, blk + byte_index, active); // write-1-to-clear the serviced bits
}
}
/// Evaluate the `\_GPE._L%02X` or `_E%02X` handler for GPE number `n`, then
/// publish an event for each device it notified.
fn dispatchGpe(n: u32) void {
const gpe_scope = aml.Namespace.resolve(&persistent_namespace, persistent_namespace.root, true, 0, &.{seg4("_GPE")}) orelse return;
var name: [4]u8 = .{ '_', 'L', 0, 0 };
writeHex2(name[2..4], n);
var method = aml.Namespace.childOf(gpe_scope, name);
if (method == null) {
name[1] = 'E';
method = aml.Namespace.childOf(gpe_scope, name);
}
const m = method orelse return; // no handler — the status bit was already cleared
_ = global_interpreter.evaluate(m, &.{}) catch return;
for (global_interpreter.takeNotifications()) |event| publishNotify(event.node, event.code);
}
fn publishNotify(node: *aml.Node, code: u64) void {
// Map the notified device's _HID to a domain event where we recognize it.
var hid: [8]u8 = .{0} ** 8;
if (readHid(node, &global_interpreter)) |h| hid = h;
const which: power.Event = if (std.mem.eql(u8, hid[0..7], "PNP0C0A")) .battery else if (std.mem.eql(u8, hid[0..7], "ACPI0003")) .ac else if (std.mem.eql(u8, hid[0..7], "PNP0C0D")) .lid else .notify;
var event = power.EventMessage{ .event = @intFromEnum(which), .code = @truncate(code) };
event.hid = hid;
writeLine("power: notify {s} code {d}\n", .{ hid[0..7], code });
publishEvent(std.mem.asBytes(&event));
}
/// Two lowercase hex digits of `n` into `out[0..2]`.
fn writeHex2(out: []u8, n: u32) void {
const digits = "0123456789ABCDEF";
out[0] = digits[(n >> 4) & 0xF];
out[1] = digits[n & 0xF];
}
fn publishButton() void {
const event = power.EventMessage{ .event = @intFromEnum(power.Event.power_button) };
publishEvent(std.mem.asBytes(&event));
}
fn publishEvent(bytes: []const u8) void {
for (&subscribers) |*slot| {
if (slot.*) |handle| {
if (!runtime.ipc.send(handle, bytes)) slot.* = null;
}
}
}
fn isSubscriber(task: u32) bool {
for (&subscribers, 0..) |*slot, si| {
if (slot.* != null and subscriber_tasks[si] == task) return true;
}
return false;
}
/// Enter S5 (soft off): write SLP_TYP|SLP_EN to the PM1 control register(s).
/// Mirrors the kernel's power.zig sleepValue. Only reached from a PID-1
/// shutdown request (M21.3).
fn enterS5() void {
if (!s5_valid or pm1a_cnt == 0) {
_ = runtime.system.write("power: S5 unavailable\n");
return;
}
_ = runtime.system.write("power: entering S5\n");
halPioWrite(2, pm1a_cnt, (@as(u32, s5_slp_typ_a & 0x7) << 10) | slp_en);
if (pm1b_cnt != 0) halPioWrite(2, pm1b_cnt, (@as(u32, s5_slp_typ_b & 0x7) << 10) | slp_en);
// If control returns, the write did not take — say so instead of hanging.
runtime.system.sleep(500);
_ = runtime.system.write("power: S5 write did not take\n");
}
// --- harness callbacks --------------------------------------------------------
fn onNotification(badge: u64) void {
// The only notification the service binds is the SCI (an IRQ badge).
_ = badge;
onSci();
}
/// The `.power` protocol: subscribe (endpoint as the call's capability),
/// shutdown (PID 1 only). Device discovery uses a different endpoint (the
/// device manager's), so nothing here handles ChildAdded.
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
if (message.len < 1) return 0;
switch (message[0]) {
@intFromEnum(power.Operation.subscribe) => {
var status: i32 = -1;
if (capability) |handle| {
for (&subscribers, 0..) |*slot, si| {
if (slot.* == null) {
slot.* = handle;
subscriber_tasks[si] = sender;
status = 0;
break;
}
}
}
const r = power.Reply{ .status = status };
@memcpy(reply[0..@sizeOf(power.Reply)], std.mem.asBytes(&r));
return @sizeOf(power.Reply);
},
@intFromEnum(power.Operation.shutdown) => {
// Honored only from a power subscriber — init, which has already run
// the stop sequence over everything else. The power service is
// mechanism (write S5); deciding *when* to shut down and stopping
// the rest of the system first is init's policy.
const allowed = isSubscriber(sender);
const r = power.Reply{ .status = if (allowed) 0 else -1 };
@memcpy(reply[0..@sizeOf(power.Reply)], std.mem.asBytes(&r));
if (allowed) enterS5();
return @sizeOf(power.Reply);
},
else => return 0,
}
} }
/// Depth-first walk: register + report each present device with a _HID, then /// Depth-first walk: register + report each present device with a _HID, then
+102 -15
View File
@@ -1,29 +1,41 @@
//! /system/services/system/services/init: — the first user-space program, PID 1. Built as its own //! /system/services/init — the first user-space program, PID 1. Built as its own
//! freestanding binary (see build.zig), shipped on the boot volume at /system/services/system/services/init:, //! freestanding binary (see build.zig), shipped on the boot volume at /system/services/init,
//! loaded by the bootloader, and started in ring 3 as a scheduled process by the //! loaded by the bootloader, and started in ring 3 as a scheduled process by the
//! kernel (system/kernel/process.zig). It links against the shared user runtime //! kernel (system/kernel/process.zig). It links against the shared user runtime
//! library `runtime` and talks to the kernel only through `runtime`'s system_call wrappers. //! library `runtime` and talks to the kernel only through `runtime`'s system_call wrappers.
//! //!
//! It proves the C-convention heap works, then — as PID 1 — acts as the system's //! It proves the C-convention heap works, then — as PID 1 — acts as the system's
//! **service supervisor**: it spawns the user-space services danos brings up at boot //! **service supervisor**: it spawns the user-space services danos brings up at boot
//! (the VFS server, the device manager), and settles into a heartbeat so it stays //! (the VFS server, the device manager), and settles into an event loop as the root
//! alive as the root of user space. Drivers are *not* its job: the device manager //! of user space. Drivers are *not* its job: the device manager discovers the
//! discovers the hardware and spawns those. This is the service half of the //! hardware and spawns those. This is the service half of the service/driver spawn
//! service/driver spawn split (docs/driver-model.md). //! split (docs/driver-model.md).
//!
//! M21: init also owns **orderly shutdown**. It supervises its children (keeping
//! their ids and an exit endpoint), subscribes to the power service, and on a
//! power-button event runs the stop sequence over its children in reverse order
//! before asking the power service to enter S5 — lifecycle (M17) and events (M21)
//! composing into a clean poweroff.
const std = @import("std");
const runtime = @import("runtime"); const runtime = @import("runtime");
const power = runtime.power_protocol;
/// The system services system/services/init: brings up at boot, in order. This is system/services/init:'s policy — the /// The system services init brings up at boot, in order. This is init's policy — the
/// microkernel keeps such choices in user space, not the kernel. Drivers are absent /// microkernel keeps such choices in user space, not the kernel. Drivers are absent
/// on purpose: the device manager owns those. (A future system/services/init: reads this from a /// on purpose: the device manager owns those. (A future init reads this from a
/// manifest under /system/services instead of a hardcoded list.) /// manifest under /system/services instead of a hardcoded list.)
const boot_services = [_][]const u8{ "vfs", "input", "device-manager" }; const boot_services = [_][]const u8{ "vfs", "input", "device-manager" };
var children: [boot_services.len]u32 = .{0} ** boot_services.len;
var child_count: usize = 0;
var supervision_endpoint: runtime.ipc.Handle = 0;
pub fn main() void { pub fn main() void {
// Prove the heap end to end: allocate through the runtime allocator (which // Prove the heap end to end: allocate through the runtime allocator (which
// mmaps pages from the kernel and carves them with the free list), write into // mmaps pages from the kernel and carves them with the free list), write into
// that heap buffer (exercising the widened debug_write bounds check), and // that heap buffer (exercising the widened debug_write bounds check), and
// free it. A fault here would kill system/services/init: before it heartbeats — so the system/services/init: // free it. A fault here would kill init before it heartbeats — so the init
// test doubles as the heap regression test. (C code links the same heap via // test doubles as the heap regression test. (C code links the same heap via
// the extern malloc/free symbols; Zig code uses this allocator.) // the extern malloc/free symbols; Zig code uses this allocator.)
const gpa = runtime.allocator(); const gpa = runtime.allocator();
@@ -34,19 +46,94 @@ pub fn main() void {
gpa.free(buffer); gpa.free(buffer);
} else |_| {} } else |_| {}
// Bring up the boot services. Best-effort and silent: each service announces its // One endpoint carries everything init waits on: children's exit
// own readiness (`vfs: ready`, ...), and in an isolation test that runs system/services/init: with // notifications (they are spawned supervised against it), init's own
// no system/services/init:ial-ramdisk the spawns simply no-op rather than deranging the heartbeat. // signals, and power events it subscribes to. All arrive in the loop below.
supervision_endpoint = runtime.ipc.createIpcEndpoint() orelse {
_ = runtime.system.write("/system/services/init: no endpoint\n");
return;
};
_ = runtime.process.bindSignals(supervision_endpoint);
// Bring up the boot services, supervised so init can stop them cleanly.
// Best-effort and silent: each service announces its own readiness, and in
// an isolation test with no initial-ramdisk the spawns simply no-op.
for (boot_services) |service| { for (boot_services) |service| {
_ = runtime.system.spawn(service); if (runtime.system.spawnSupervised(service, &.{}, supervision_endpoint)) |id| {
children[child_count] = id;
child_count += 1;
}
} }
// Subscribe to power events (retry: the power service registers well after
// init starts). Best-effort — without it, a `terminate` signal still
// triggers the same shutdown path.
subscribePower();
// A re-arming timer drives the liveness heartbeat: proof PID 1 is alive
// (the init test's marker) while the loop stays free to receive signals,
// power events, and children's exit notifications.
_ = runtime.system.timerOnce(supervision_endpoint, 1000);
var receive: [power.message_maximum]u8 = undefined;
while (true) { while (true) {
_ = runtime.system.write("/system/services/init: heartbeat\n"); const got = runtime.ipc.replyWait(supervision_endpoint, &.{}, &receive, null);
runtime.system.sleep(1000); if (runtime.process.signalsFrom(got.badge)) |signals| {
if (signals.has(.terminate)) shutDown();
continue;
}
if (got.isTimer()) {
_ = runtime.system.write("/system/services/init: heartbeat\n");
_ = runtime.system.timerOnce(supervision_endpoint, 1000);
continue;
}
if (got.isMessage() and got.len >= 2 and receive[0] == @intFromEnum(power.Operation.event)) {
// A power event (the only buffered messages init receives).
if (receive[1] == @intFromEnum(power.Event.power_button)) shutDown();
continue;
}
// Child-exit notifications and anything else: keep waiting.
if (got.isNotification()) continue;
} }
} }
/// Look up the power service and subscribe our endpoint (handed over as the
/// call's capability) so events arrive as buffered messages here.
fn subscribePower() void {
var handle: ?runtime.ipc.Handle = null;
var tries: u32 = 0;
while (handle == null and tries < 200) : (tries += 1) {
handle = runtime.ipc.lookup(.power);
if (handle == null) runtime.system.sleep(20);
}
// A missing power service is not fatal — init proceeds to its heartbeat and
// a `terminate` signal still drives shutdown. Silent so the no-ramdisk init
// test's heartbeat marker is the next line written.
const h = handle orelse return;
const request = power.Subscribe{};
var reply: [power.message_maximum]u8 = undefined;
_ = runtime.ipc.callCap(h, std.mem.asBytes(&request), &reply, supervision_endpoint) catch {};
}
/// The stop sequence: terminate each child in reverse spawn order (vfs last —
/// other services may flush through it), waiting up to a deadline for each to
/// exit before killing it, then ask the power service to enter S5.
fn shutDown() void {
_ = runtime.system.write("/system/services/init: shutting down\n");
var i = child_count;
while (i > 0) {
i -= 1;
if (children[i] != 0) runtime.process.stop(children[i], 2000, supervision_endpoint);
}
if (runtime.ipc.lookup(.power)) |h| {
const request = power.Shutdown{};
var reply: [power.message_maximum]u8 = undefined;
_ = runtime.ipc.call(h, std.mem.asBytes(&request), &reply) catch {};
}
// If S5 did not take, init has nothing left to do but idle.
while (true) runtime.system.sleep(1000);
}
pub const panic = runtime.panic; pub const panic = runtime.panic;
comptime { comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image _ = &runtime.start._start; // pull the runtime entry shim into the image
+68
View File
@@ -0,0 +1,68 @@
//! The power protocol (docs/m21-plan.md): system power's domain-named surface,
//! registered under `ServiceId.power`. On x86 the acpi service serves it; on
//! ARM a PSCI/mailbox service will register the same id — subscribers never
//! learn which firmware they are on (m19-m20-plan.md decision 7). The
//! vfs-protocol pattern: extern-struct messages, a version, reserved fields.
/// The protocol version a client states nowhere yet — reserved for the day a
/// handshake needs it; requests carry it so a mismatch can be refused loudly.
pub const version: u16 = 1;
pub const Operation = enum(u8) {
/// Subscribe to power events: the subscriber's endpoint rides as the
/// call's capability (the input/device-manager pattern); events arrive on
/// it as buffered messages carrying an `EventMessage`.
subscribe = 1,
/// Orderly shutdown's last step: enter S5. Accepted only from PID 1
/// (init) — the process that has already run the stop sequence over
/// everything else.
shutdown = 2,
/// The published event payload (never sent *to* the service).
event = 3,
};
/// What happened. The vocabulary is hardware-neutral: a lid is a lid whether
/// ACPI or a PSCI mailbox reported it.
pub const Event = enum(u8) {
power_button = 1,
lid = 2,
ac = 3,
battery = 4,
/// A device notification that maps to none of the named events — the
/// `code` and `hid` fields say which device and what code.
notify = 5,
};
pub const Subscribe = extern struct {
operation: u8 = @intFromEnum(Operation.subscribe),
reserved0: u8 = 0,
version: u16 = version,
reserved1: u32 = 0,
};
pub const Shutdown = extern struct {
operation: u8 = @intFromEnum(Operation.shutdown),
reserved0: u8 = 0,
version: u16 = version,
reserved1: u32 = 0,
};
/// A published event, as the buffered-message payload subscribers receive.
pub const EventMessage = extern struct {
operation: u8 = @intFromEnum(Operation.event),
/// An Event value.
event: u8,
reserved0: u16 = 0,
/// The device notification code (Notify's second argument), or 0.
code: u32 = 0,
/// The notifying device's hardware id (EISA-decoded), or all zero.
hid: [8]u8 = .{0} ** 8,
};
pub const Reply = extern struct {
status: i32,
reserved: u32 = 0,
};
/// Upper bound on any message in this protocol — sizes endpoint buffers.
pub const message_maximum = 64;
+66
View File
@@ -18,9 +18,11 @@ Usage:
""" """
import argparse import argparse
import json
import os import os
import re import re
import shutil import shutil
import socket
import subprocess import subprocess
import sys import sys
import time import time
@@ -80,7 +82,10 @@ ARCHES = {
# `expect`: a regex that must appear in serial output => pass. # `expect`: a regex that must appear in serial output => pass.
# `fail`: optional regex whose appearance => immediate fail. # `fail`: optional regex whose appearance => immediate fail.
CASES = [ CASES = [
# smoke also proves the QMP channel: the harmless query must be delivered
# (handshake + command) before the case may pass — see run_case.
{"name": "smoke", {"name": "smoke",
"qmp_after": {"delay": 2, "command": "query-status"},
"expect": r"DANOS-TEST-RESULT: PASS", "expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
{"name": "discovery", {"name": "discovery",
@@ -287,6 +292,28 @@ CASES = [
r"device-manager: spawned ps2-bus[\s\S]*" r"device-manager: spawned ps2-bus[\s\S]*"
r"ps2-bus: keyboard driver attached", r"ps2-bus: keyboard driver attached",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# M21.1: the SCI + power button. Boot the manager (which spawns the acpi
# service); ~4s in, QMP system_powerdown raises the ACPI power-button fixed
# event; the service's SCI handler must log the press (docs/m21-plan.md).
{"name": "power-button",
"smp": 4,
"timeout": 60,
"qmp_after": {"delay": 4, "command": "system_powerdown"},
"expect": r"power: button pressed",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M21.3 capstone: orderly shutdown. Boot init (the full tree comes up);
# ~5s in, QMP system_powerdown raises the power button; the acpi service
# publishes it, init stops its children then requests S5, and QEMU exits.
# The ordered regex proves button -> shutting-down -> entering-S5; the case
# passes on QEMU's self-exit through S5 (docs/m21-plan.md).
{"name": "orderly-shutdown",
"smp": 4,
"timeout": 90,
"qmp_after": {"delay": 5, "command": "system_powerdown"},
"expect": r"power: button pressed[\s\S]*"
r"init: shutting down[\s\S]*"
r"power: entering S5",
"fail": r"power: S5 write did not take|DANOS-TEST-RESULT: FAIL"},
# M20.2: the acpi service evaluates _CRS/_STA in ring 3 and registers + # M20.2: the acpi service evaluates _CRS/_STA in ring 3 and registers +
# reports its _HID devices — the two PS/2 nodes must appear with resources # reports its _HID devices — the two PS/2 nodes must appear with resources
# (keyboard: io 0x60/0x64 + IRQ = 3; mouse: IRQ = 1) (docs/m19-m20-plan.md). # (keyboard: io 0x60/0x64 + IRQ = 3; mouse: IRQ = 1) (docs/m19-m20-plan.md).
@@ -337,6 +364,7 @@ CASES = [
# The initial_ramdisk: the loader ferries a bundle of user binaries; the kernel parses # The initial_ramdisk: the loader ferries a bundle of user binaries; the kernel parses
# it and spawns each as a ring-3 process (here the VFS-server stub heartbeats). # it and spawns each as a ring-3 process (here the VFS-server stub heartbeats).
{"name": "initial-ramdisk", {"name": "initial-ramdisk",
"timeout": 60, # the acpi service's boot-time SCI setup can push the marker past 30s under load
"expect": r"DANOS-TEST-RESULT: PASS", "expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# The user-space VFS: a client opens/writes/reads a file through the rt file # The user-space VFS: a client opens/writes/reads a file through the rt file
@@ -418,6 +446,27 @@ def resolve_firmware(arch):
+ "\nInstall OVMF (edk2-ovmf / ovmf) or add its path above.") + "\nInstall OVMF (edk2-ovmf / ovmf) or add its path above.")
def qmp_send(path, command):
"""One QMP command: connect, capabilities handshake, execute. Raises on any
failure — the caller retries until the guest's socket is ready. This is how
a case injects a host-side event (system_powerdown = the ACPI power button)
into the running guest (docs/m21-plan.md)."""
sock = socket.socket(socket.AF_UNIX, socket.SOCK_STREAM)
sock.settimeout(5)
try:
sock.connect(path)
stream = sock.makefile("rw")
stream.readline() # the QMP greeting
stream.write(json.dumps({"execute": "qmp_capabilities"}) + "\n")
stream.flush()
stream.readline() # {"return": {}}
stream.write(json.dumps({"execute": command}) + "\n")
stream.flush()
stream.readline()
finally:
sock.close()
def run_case(arch, case): def run_case(arch, case):
err = build(arch, case["name"]) err = build(arch, case["name"])
if err: if err:
@@ -439,12 +488,27 @@ def run_case(arch, case):
cmd += ["-smp", str(case["smp"])] cmd += ["-smp", str(case["smp"])]
if case.get("qemu_extra"): # extra qemu args, e.g. -device intel-iommu for the IOMMU case if case.get("qemu_extra"): # extra qemu args, e.g. -device intel-iommu for the IOMMU case
cmd += case["qemu_extra"] cmd += case["qemu_extra"]
# A QMP control socket, always present (additive): how a case's `qmp_after`
# hook injects host-side events into the guest mid-run.
qmp_path = os.path.join(WORK, "qmp.sock")
if os.path.exists(qmp_path):
os.remove(qmp_path)
cmd += ["-qmp", f"unix:{qmp_path},server,nowait"]
qmp_after = case.get("qmp_after") # {"delay": seconds, "command": "..."}
qmp_sent = False
started = time.monotonic()
qemu = subprocess.Popen(cmd, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL) qemu = subprocess.Popen(cmd, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
try: try:
timeout = case.get("timeout", TIMEOUT) timeout = case.get("timeout", TIMEOUT)
deadline = time.monotonic() + timeout deadline = time.monotonic() + timeout
while time.monotonic() < deadline: while time.monotonic() < deadline:
time.sleep(0.2) time.sleep(0.2)
if qmp_after and not qmp_sent and time.monotonic() - started >= qmp_after["delay"]:
try:
qmp_send(qmp_path, qmp_after["command"])
qmp_sent = True
except OSError:
pass # socket not up yet; retry next tick
text = "" text = ""
if os.path.exists(serial): if os.path.exists(serial):
with open(serial, "r", errors="replace") as f: with open(serial, "r", errors="replace") as f:
@@ -452,6 +516,8 @@ def run_case(arch, case):
if fail and fail.search(text): if fail and fail.search(text):
return False, "hit failure marker" return False, "hit failure marker"
if expect.search(text): if expect.search(text):
if qmp_after and not qmp_sent:
continue # the hook must deliver before the case may pass
return True, "matched " + repr(case["expect"]) return True, "matched " + repr(case["expect"])
if qemu.poll() is not None: # QEMU exited on its own if qemu.poll() is not None: # QEMU exited on its own
if expect.search(text): if expect.search(text):