Add runtime.time, drop demo drivers, harden TSC timekeeping

Time is a kernel concern in danos: the kernel owns the scheduling timer and
already exposes monotonic time via the clock/sleep/timer_bind syscalls, so a
userspace time service would be a redundant, slower path. This adds the generic
runtime.time module over those syscalls, retires the two demonstration drivers,
reorganizes the milestone docs, and makes the monotonic clock correct on Intel,
AMD, and inside any VM.

runtime.time (library/runtime/time.zig)
- Instant/Duration interface: now, sleep, spin, after, monotonicNanos, available
- a thin layer over system.clock/sleep/timerOnce; unit-tested arithmetic

Remove the demo drivers hpet and bus (a teaching example belongs in the docs,
not shipped in the tree)
- system/drivers/ now holds only real drivers: pci-bus, ps2-bus, usb-xhci-bus
- device-manager end-to-end test repointed to pci-bus (asserts on kernel state:
  the process table and the device tree, not a racy serial marker)
- device_register containment moved to a new in-kernel `containment` test
- the driver-model worked example moved inline into docs/drivers.md

Reorganize milestone docs into topic docs
- m17-m18 / m19-m20 / m21 plans dissolved into process-lifecycle, device-manager,
  discovery, and acpi docs; new docs/power.md and docs/timers.md; ~20 citations
  repointed; plan docs deleted

TSC reliability (apic.zig, smp.zig, cpu.zig, kernel.zig)
- check the invariant-TSC bit (CPUID 0x80000007 EDX[8]) on Intel and AMD
- cross-core "warp" check at SMP bring-up, pairwise BSP<->AP as each core comes up
- fall back to the HPET clocksource when the TSC is not invariant (a bare VM) or
  not synchronized (a warp), switched continuously so time never jumps
- boot log reports the outcome; new tsc-sync test exercises the TSC + warp path

Verified: zig build; zig build test; 60/60 QEMU cases (incl. new containment and
tsc-sync).
This commit is contained in:
Daniel Samson
2026-07-13 11:56:43 +01:00
parent 5b63a841ba
commit 452080e997
36 changed files with 1150 additions and 1267 deletions
+10 -1
View File
@@ -100,7 +100,16 @@ Cutting across all of these:
when to build it, and how to keep it architecture-agnostic.
- **[acpi.md](acpi.md) — finding the ACPI tables.** The concrete x86 locator chain:
how the loader captures the **RSDP**, hands its physical address across in `BootInfo`,
and how the platform derives the **RSDT/XSDT** from it and walks the SDTs.
and how the platform derives the **RSDT/XSDT** from it and walks the SDTs — plus the
live event side (the SCI, the power button, GPE/Notify) the ring-3 acpi service runs.
- **[power.md](power.md) — the power service.** System power as a domain-named
service: button/lid/battery events published to subscribers, and init's orderly
shutdown composing the [lifecycle](process-lifecycle.md) stop sequence with an ACPI
S5 write. Firmware-neutral — a PSCI backend drops in on ARM.
- **[timers.md](timers.md) — timers and time.** The ring-3 surface for reading the
clock and waiting: why `now()` is a syscall rather than a service, and the one-shot
timer notification (`timer_bind`) that gives supervisors a timed wait — built on the
LAPIC heartbeat and calibrated TSC of [device-interrupts.md](device-interrupts.md).
- **[smp.md](smp.md) — multiple cores.** A design/research note on how microkernels
(L4, seL4) handle SMP — big kernel lock vs per-CPU vs multikernel — and how the
right choice depends on whether danos is chasing real-time or resilience.
+55 -1
View File
@@ -107,12 +107,66 @@ firmware-agnostic [device model](discovery.md) gets populated; this note stops a
part that answers "where are the tables?" — everything past the RSDP is just following
more pointers the tables themselves provide.
## ACPI events: the SCI, the power button, and GPEs (M21)
The tables above are static description; ACPI is also a *live* channel. Hardware
raises the **SCI** (System Control Interrupt) — one shared, level-triggered line
whose vector the FADT names — and the OS reads status registers to learn what
happened: a fixed event like the power button, or a **General-Purpose Event**
(GPE) whose handler is an AML method. Since [discovery](discovery.md) moved AML
to ring 3, the event side lives there too, in the same **acpi service** — the
device discoverer and the event source are one process, because both need the
namespace and the port grant.
**The kernel hands the service what it needs and no more.** Reading PM1 event
blocks and GPE blocks requires the FADT, which the kernel already parses for its
own `\_S5` poweroff. Rather than re-parse, the kernel appends the **FADT as one
more memory resource** on the `acpi-tables` node; the service tells it apart
from the AML blob resources by signature — the FADT keeps its intact `"FACP"`
header, while the blob resources are header-stripped bytecode that starts with
no signature. The kernel's own FADT parse is untouched; the service reads the
PM1 *event* blocks (which the kernel never parsed — it only needs PM1 *control*
for `\_S5`) and the GPE0/GPE1 blocks straight from its copy. The **SCI itself**
arrives as the node's one `len == 1` irq resource (distinct from the broad
`[0, 256)` window that covers children's legacy lines), which is how the service
finds the line to `irq_bind`.
With those in hand the service enables ACPI mode (only if `SCI_EN` is clear —
some firmwares boot with it already set), sets `PWRBTN_EN`, and on each SCI:
- **The power button** is a *fixed* event: a set `PWRBTN_STS` bit in PM1 status.
The handler clears it (write-1-to-clear), logs the press, and publishes a
[`power`](power.md) `power_button` event to subscribers.
- **GPEs** are the general path: for each set-and-enabled GPE bit `n`, the
service evaluates its `\_GPE._L%02X` (level) or `_E%02X` (edge) handler
method, drains the **Notify** queue that method produced, maps each notified
device to an event (battery, AC, lid, or a generic `notify` with its code),
and clears the status bit. A missing handler method is clear-and-log, not an
error. Making GPEs work required teaching the interpreter one opcode it never
handled — `Notify` (`0x86`) — which it now folds into a bounded queue drained
per evaluation; everything else a handler needs (field access, control flow,
method calls) was already proven by the ring-3 `_STA`/`_CRS` work.
**How this is tested.** QEMU cannot raise GPEs deterministically on this config,
so GPE/Notify correctness is proven by **host unit tests** — hand-encoded AML
with a `Notify` inside a method body, run under `zig build test`. The QEMU
`power-button` scenario proves the fixed-event path end to end: a QMP
`system_powerdown` injects a real ACPI power-button press, and the service's SCI
handler must log it. Battery/AC/lid and the embedded controller's `_Qxx` queries
are interface-complete but validated on real hardware later.
The service surface these events are *published on* — subscription, the event
vocabulary, and orderly shutdown — is the power service, [power.md](power.md).
## Related
- [efi.md](efi.md) — the loader that captures the RSDP before `ExitBootServices`.
- [memory-map.md](memory-map.md) — the same loader-captures / kernel-consumes seam, and
the ACPI-reclaim memory the RSDP lives in.
- [discovery.md](discovery.md) — the broader (still-evolving) plan for turning these
tables into one neutral device model shared with the ARM device-tree path.
tables into one neutral device model shared with the ARM device-tree path, and how
ACPI enumeration and events moved to the ring-3 acpi service.
- [power.md](power.md) — the domain-named power service the ACPI event side publishes
to (button, lid, battery) and its orderly-shutdown path into S5.
- [arch.md](arch.md) — why the kernel reaches the device code through a `platform`
module and never names ACPI directly.
+2 -2
View File
@@ -96,7 +96,7 @@ That's all — no Unix-abbreviation exception. The source directories are full w
(`system`, `library`, not `src`/`lib`), and there is no daemon `d` suffix: a driver
lives in `system/drivers/` and a service in `system/services/`, so the *location*
already says what it is. Encoding the role in the name too (`busd`, `vfsd`) is
redundant — the program is just `bus`, `vfs`. Don't put in a name what its directory
redundant — the program is just `ps2-bus`, `vfs`. Don't put in a name what its directory
already tells you.
## A note on collisions
@@ -138,7 +138,7 @@ single word or acronym needs no hyphen: `scheduler.zig`, `paging.zig`, `apic.zig
conventions above — `snake_case` — because it's an identifier, not a filename.)
**A sub-project's entry point repeats its directory's name** — `init/init.zig`,
`runtime/runtime.zig`, `hpet/hpet.zig` — and the sub-project is addressed by the
`runtime/runtime.zig`, `ps2-bus/ps2-bus.zig` — and the sub-project is addressed by the
*directory* (`system/services/init`, `library/runtime`), with the repeated leaf
resolving away. See the repository-layout section of [README.md](README.md).
+3 -3
View File
@@ -17,7 +17,7 @@ Most modern Unix and Unix-like operating systems follow the FHS. DanOS has its o
| /srv | Site-specific data served by this system, such as data and scripts for web servers, data offered by FTP servers, and repositories for version control systems |
| /system | DanOS operating system files (similar idea to C:\Windows). A true representation of danos — its layout mirrors the source tree, so `/system` is what danos *is*. |
| /system/devices | danos virtual device tree e.g. similar to /sys on linux but with danos device tree conventions (the structures in the devices module) |
| /system/drivers | driver binaries, one sub-project each (e.g. /system/drivers/hpet) |
| /system/drivers | driver binaries, one sub-project each (e.g. /system/drivers/pci-bus, /system/drivers/ps2-bus) |
| /system/services | system-service binaries — the VFS server, init, and other user-mode servers (e.g. /system/services/vfs, /system/services/init) |
| /system/kernel | the kernel image |
| /tmp | Directory for temporary files (see also /var/tmp). Often not preserved between system reboots and may be severely size-restricted. |
@@ -61,8 +61,8 @@ to the driver in the order written, and a read consumes what is there. Terminals
serial lines, keyboards and mice are all of this shape. These are the natural first
device nodes in danos, because a character driver needs nothing the kernel doesn't
already provide — it claims its device, maps its registers with `mmio_map`, and blocks
on `replyWait` for either an interrupt or a client request. `system/drivers/hpet/hpet.zig` is already
that program, minus the client half.
on `replyWait` for either an interrupt or a client request. `system/drivers/ps2-bus/ps2-bus.zig`
is already that program, minus the file-node client half.
The obstacle was never the file type; it is which hardware a ring-3 driver can reach.
Direct `in`/`out` from user space is still a #GP (no TSS I/O bitmap, IOPL never raised),
+34
View File
@@ -78,6 +78,40 @@ preemption and wakeups (1 ms granularity); the **TSC** is the resolution you rea
time at. Making `sleep` itself sub-millisecond would take a tickless one-shot
timer — a later step.
### Is the TSC trustworthy? Invariant, and synchronized
A cycle counter is only a valid *clock* if two things hold, and danos checks both,
because they decide whether we read time with a cheap `rdtsc` or fall back to the HPET.
**Invariant.** An old TSC counted core clock cycles, so it sped up and slowed down with
frequency scaling — useless as wall time. Modern CPUs (all of danos's targets) provide an
**invariant TSC**: a constant rate across P/C-states that never stops. The guarantee is a
CPUID bit — leaf `0x80000007`, EDX bit 8 — on both Intel *and* AMD. danos reads it in
`calibrate`, and a TSC that doesn't advertise it is not used as the clocksource. AMD is
why this matters in practice: it doesn't populate the Intel leaf `0x15` that enumerates
the TSC *frequency*, so danos already measures AMD's rate against the HPET — but a
measured frequency without the invariance guarantee is not enough.
**Synchronized.** Each core has its own TSC. Even invariant ones can start at different
values (a second socket, some firmware), so a thread migrating from a core reading
`1_000_000` to one reading `999_000` would see time jump *backward*. danos runs a **warp
check** as each application processor comes online (`checkWarpSource`, adapted from
Linux's): the waking core and the BSP hammer a shared "highest seen" TSC under a lock,
and if either ever reads below it, the cores' TSCs are skewed. It's pairwise because APs
come up one at a time ([smp.md](smp.md)).
**The fallback.** When the TSC fails either test — non-invariant (a bare VM such as the
default qemu64), or warped between cores — danos moves the monotonic clock onto the
**HPET** main counter: one fixed-rate counter, so it can neither skew between cores nor
drift with frequency. It costs a memory-mapped read instead of a register read, but it
keeps time *accurate*, which is the whole point. The switch preserves the current value,
so the clock never jumps. The boot log names the outcome:
```
/system/kernel: clocksource tsc (TSC invariant: yes, synchronized: yes) # real Intel/AMD
/system/kernel: clocksource hpet (TSC invariant: no, synchronized: yes) # a bare VM (TCG)
```
## Two kinds of vector, one dispatch
The IDT now installs gates `0-47`: the 32 exceptions plus the device range. Every
+21 -5
View File
@@ -51,7 +51,16 @@ enumeration is a **pci-bus driver**: the manager spawns it against the host brid
like any bus reports children. ACPI becomes an **acpi service** that interprets the
tables and reports the namespace. The manager only orchestrates and merges. Moving
AML interpretation out of ring 0 is its own project on its own track; nothing here
depends on when it lands.
depends on when it lands. (It landed: [discovery.md](discovery.md), M19–M20.)
`device_register` is **idempotent on exact match**: a re-registration with an
identical (parent, class, identity, resources) tuple returns the existing id
instead of appending a duplicate. The kernel table has no unregister, so without
this a restarted registering bus would re-report its children as fresh nodes on
every respawn. Idempotence is what makes restart-and-re-report sound for *every*
reporting bus — pci-bus, the acpi service, a future fdt service — not just one,
and it is why supervision (below) can prune a dead bus's subtree and trust the
restarted instance to rebuild exactly the same ids.
## The protocol
@@ -136,10 +145,17 @@ published exit events, signals + `runtime.process`). On top of those:
the mouse and keyboard QEMU already hangs off it.
7. **App surface**: `enumerate`/`subscribe` over IPC; `device_enumerate` retreats
to a manager-internal seam.
8. **Discovery migration** — DONE (M19–M20, 2026-07-13): pci-bus driver (M19)
then the acpi service (M20) moved enumeration to ring 3; the kernel seeds
only the host bridge and the acpi-tables node. See
[m19-m20-plan.md](m19-m20-plan.md).
8. **Discovery migration** — DONE (M19–M20, 2026-07-13): enumeration moved to
ring 3 as swappable per-firmware discoverers — the pci-bus driver (M19) then
the acpi service (M20), see [discovery.md](discovery.md); the kernel seeds
only the host bridge and the acpi-tables node. Matching moved with it:
`child_added` grew a `device_id` (the kernel-registered id, `no_device` for
unregistered leaves like USB ports) and a firmware `hid`, and the manager now
matches drivers from those **reports** rather than its boot-time snapshot. The
PCI arm flipped in M19.3, the ACPI arm (ps2-bus matched from `_HID`) in M20.3
— each in a single phase so no device is ever matched from both sources at
once. The acpi service reports only the non-PCI `_HID` devices, since pci-bus
already reports PCI functions (M20.2).
## Settled questions (2026-07-12)
+54
View File
@@ -191,3 +191,57 @@ and registers + reports each `_HID` device — the device manager matches driver
(ps2-bus) from those reports. With M19's pci-bus driver, discovery now runs
entirely in user space; the kernel seeds only the host bridge and the
acpi-tables node.
## Discovery is a swappable process per firmware (M19–M20)
Moving PCI and ACPI enumeration out of ring 0 was not just a relocation — it
made discovery **firmware-neutral by construction**, which is the whole reason
to do it before the second architecture rather than after. Everything at and
above the [device-manager](device-manager.md) protocol — descriptors,
containment, reports, matching, supervision — is generic and may never become
x86-specific. Discovery is the single firmware-specific piece, and it is
isolated as **one swappable process per firmware**:
- **x86** boots describe hardware with ACPI, so the discoverer is the **acpi
service** ([acpi.md](acpi.md)): it claims the `acpi-tables` node and runs AML.
- **The Raspberry Pis** hand over a flattened device tree, so the discoverer is
an **fdt service**: it claims a `devicetree-blob` node and walks the tree —
pure data, no bytecode, so it needs neither a port grant nor an interpreter,
strictly simpler than ACPI. (A placeholder until the [aarch64](arm.md)
bring-up fills it in.)
The device manager spawns the discoverer under the **neutral ramdisk name
`discovery`** and never learns which firmware it is on; the build's
`-Ddiscovery=acpi|fdt` option fills that slot (x86 defaults to `acpi`, the
aarch64 target flips the default when it lands). The manager owns the device
tree as *data* and touches no hardware, ever — firmware bytecode runs only
inside the crashable, supervised discoverer, so an AML fault can never take
down the supervisor.
Two consequences of neutrality bind on later work:
- **Cross-firmware surfaces are named by domain, not firmware.** System power is
a [`power`](power.md) protocol, not an "ACPI events" protocol: on x86 the acpi
service registers it, on ARM a PSCI/mailbox service registers the same
`ServiceId.power`, and subscribers never learn the difference.
- **Identity must widen before the fdt service exists.** `DeviceDescriptor`'s
8-byte `hid` holds an EISA id but cannot hold an FDT `compatible` string
(`"brcm,bcm2835-aux-uart"`); the identity field grows before the ARM path can
report a real node.
Two supporting decisions keep the kernel's remaining slice honest:
- **The AML interpreter is a shared build module**, compiled into both the
kernel and the acpi service — one source, two builds, no fork. The kernel
links it for the `\_S5` poweroff evaluation, the service links it for
everything else, and the `acpi-parse` test asserts the two produce the same
device count across the ring-3 move.
- **Bridge apertures come from the firmware memory map, not AML.** Registered
PCI functions carry BAR resources, and `device_register` containment demands
the bridge own windows that cover them. Those apertures are derived
kernel-side from the boot memory map's MMIO holes (regions that are neither
RAM nor tables) — mechanical, AML-free, and available at boot regardless of
what later moved to user space. The acpi service's authority is likewise
exactly one node: the `acpi-tables` node, whose broad io_port grant is the
documented trust boundary for the one process allowed to run firmware
bytecode.
+10 -9
View File
@@ -57,8 +57,9 @@ is not an address window. Discovery is trusted; user space is not.
### What a bus driver looks like
`system/drivers/bus/bus.zig` is the smallest honest one. Its "bus" is the HPET's register block and
its "devices" are the block's comparators:
danos ships no demo bus driver — the real ones are `pci-bus`, `ps2-bus`, and
`usb-xhci-bus`. The smallest *honest* shape, illustrated here with an HPET register block
as the "bus" and its comparators as the "devices", is:
```zig
_ = dev.claim(bus.id); // 1. own the bus
@@ -78,8 +79,8 @@ for (0..n) |i| { // 3. publish each child
Each child is left **unclaimed**, which is the handoff: a comparator driver can now
`device_claim` one and `mmio_map` it, and will see only its own 0x20-byte window. A child
whose window escapes the bus is refused — `bus` asserts that, and the `bus` test
asserts the kernel's table upholds it.
whose window escapes the bus is refused; the in-kernel `containment` test asserts the
kernel's table upholds that ([drivers.md](drivers.md)).
A USB device has *no* resources at all: `resource_count = 0`, because it's addressed
through its controller, not by MMIO. That case is allowed and is the common one.
@@ -143,7 +144,7 @@ If a class driver needs `mmio`, it has become an HCD and should be one.
physically-contiguous, pinned, uncacheable, reclaim-on-teardown buffers with the
physical address exposed (`pmm.allocContiguous`, a DMA arena, `mapUserDmaInto`).
`dma_below_4g` caps the address for legacy engines; `dma_write_combining` is accepted
but falls back to coherent until PAT is programmed. hpet is refactored onto `/lib/mmio`;
but falls back to coherent until PAT is programmed. The bus drivers use `/lib/mmio`;
no DMA driver consumes `dma_alloc` yet.
- **M15** — interrupts for PCI devices, the MSI half. Discovery now gives every PCI
function its 4 KiB ECAM config space as resource 0 (unblocking the capability walk
@@ -302,8 +303,8 @@ rather than an out-struct. The rest of this section is the original design note.
**The blocker, and it's a hard one.** No PCI device can take an interrupt today.
[`addBars`](system/devices/acpi.zig) records `.memory` and `.io_port` BARs and never an
`.irq`; there is no `_PRT` parsing anywhere in the tree. `hpet` only works because the
HPET advertises its own routing options in its own registers — a privilege no ordinary
`.irq`; there is no `_PRT` parsing anywhere in the tree. The HPET is the one exception —
it advertises its own interrupt routing in its own registers, a privilege no ordinary
device has.
**The fix, in two halves.**
@@ -326,7 +327,7 @@ which means **discovery should give each `pci_device` a `.memory` resource for i
4 KiB ECAM slot**. That's a small change to `parseMcfg` and it unblocks the whole
capability walk (MSI, MSI-X, PCIe extended caps) without any new syscall.
Note QEMU's HPET reports `Tn_FSB_INT_DEL_CAP = 0` — no MSI — so `hpet` can never
Note QEMU's HPET reports `Tn_FSB_INT_DEL_CAP = 0` — no MSI — so an HPET timer could never
exercise this path. The first MSI driver will be the first PCI driver.
## M16 — the IOMMU, and the honest caveat ◑ detection done, enforcement pending
@@ -353,7 +354,7 @@ gap should be named rather than implied.
`M13` (capability passing) is independent of `M14`/`M15` and is the cheapest. It
unlocks class drivers, which are the shape with no hardware requirements at all — you
could write a real one against `bus`'s comparators tomorrow.
could write a real one against any device a bus driver publishes tomorrow.
`M14` and `M15` together unlock the first HCD. `M14`'s barrier layer is worth landing
on its own regardless: it's small, obviously correct, and stops every future driver
+40 -29
View File
@@ -22,12 +22,12 @@ say.*
## How a driver gets started: discover, match, spawn
Nothing in the kernel decides that the HPET needs the `hpet` driver — that is policy,
and policy lives in user space. Boot brings user space up as a three-level supervision
hierarchy, each level owning one job:
Nothing in the kernel decides that the PCI host bridge needs the `pci-bus` driver — that
is policy, and policy lives in user space. Boot brings user space up as a three-level
supervision hierarchy, each level owning one job:
```
kernel ──spawns──► init (PID 1) ──spawns──► device-manager ──spawns──► hpet
kernel ──spawns──► init (PID 1) ──spawns──► device-manager ──spawns──► pci-bus
| | |
spawns only init, the service supervisor: the driver supervisor: enumerates
publishes the starts the system /system/devices, matches each device
@@ -184,7 +184,11 @@ Two properties worth knowing:
## A whole driver
`system/drivers/hpet/hpet.zig` is ~150 lines and does all of it. The shape:
A minimal leaf driver is only ~150 lines and does all of it. danos ships **no such
example binary** — the driver model is proven by the real drivers (`pci-bus`, `ps2-bus`,
`usb-xhci-bus`), and a teaching example belongs here, in the docs, rather than as a
compiled program nobody runs. Illustrated with a hypothetical HPET timer driver, the
shape is:
```zig
const hpet = findHpet(buf) orelse return; // device_enumerate, look for
@@ -209,8 +213,8 @@ while (...) {
}
```
The HPET is a good first driver for a reason that isn't obvious. Its *counter* is a
clocksource — the only way to use it is to read it, so it proved `mmio_map` without
The HPET makes a good illustration for a reason that isn't obvious. Its *counter* is a
clocksource — the only way to use it is to read it, so it exercises `mmio_map` without
needing interrupts at all. Its *comparators* are a clockevent, and can be configured
**level-triggered** (`Tn_INT_TYPE_CNF`), which asserts a bit in `GENERAL_INT_STATUS`
that the driver must write-1-to-clear. That's a genuine deassert step, so the full
@@ -252,9 +256,10 @@ bus driver may only ever subdivide what it already owns.
A device with **no resources** is legal and common. A USB device is reached through its
controller, not by MMIO, so it gets `resource_count = 0`.
See [`system/drivers/bus/bus.zig`](../system/drivers/bus/bus.zig) for a complete one, and
[driver-model.md](driver-model.md) for how bus drivers, class drivers and host
controller drivers fit together.
See [`system/drivers/pci-bus/pci-bus.zig`](../system/drivers/pci-bus/pci-bus.zig) for a
real one — it claims a PCI host bridge, maps its ECAM window, and publishes each function
it finds as a child — and [driver-model.md](driver-model.md) for how bus drivers, class
drivers and host controller drivers fit together.
## What the kernel does not do for you
@@ -313,32 +318,38 @@ uncacheable, physical address exposed), and **memory barriers** (`/lib/mmio`'s
## Verifying it
The `hpet` test spawns `hpet` from the initial ramdisk and watches the serial log. The driver
prints `hpet: ok` only after being woken five times, and its loop's only exit is
through `replyWait` returning a notification — it cannot reach that line by polling.
No demo driver ships to prove this end to end; the *real* drivers do, so the tests
target them and the kernel primitives directly:
The last check doesn't trust the driver's self-report at all: the kernel reads the I/O
APIC redirection entry back and asserts the line really is routed to a device vector,
really is level-triggered, and really was left unmasked by the driver's final
`irq_ack`.
- **`device-manager`** — boots only the device manager, which discovers the PCI host
bridge, matches `pci-bus`, and `system_spawn`s it. The test reads kernel state — the
process table and the device tree — to confirm pci-bus came up and registered the
functions it enumerated: the whole discover → match → spawn → driver-up chain.
- **`acpi-ps2`** — a user-space driver (`ps2-bus`) is woken by its device's IRQ,
delivered as an IPC notification, and attaches the keyboard: IRQ-as-IPC, end to end.
- **`pci-scan`** — a user-space driver (`pci-bus`) maps its device's MMIO (the ECAM
window) and walks it: `mmio_map`, end to end.
- **`containment`** — the kernel refuses a `device_register` whose child window escapes
the parent's grant (else it would be a syscall for mapping arbitrary memory), while an
identical re-register stays idempotent. Asserted in-kernel, straight against the broker.
- **`irqfree`** — the teardown path. Binds two owners to one shared endpoint, releases
one, and reads the I/O APIC back: the departing owner's line is masked, the sibling's
is not. That second half is why bindings are keyed on the owning *task* and not on the
endpoint pointer — endpoints are shared, so releasing "everything pointing at this
endpoint" would silently mask a live driver's device.
- **`iopass`** — the `device_grant` teardown rule, so destroying a driver's address
space never returns MMIO frames to the RAM pool.
```
$ python3 test/qemu_test.py hpet irqfree iopass
hpet ... PASS (matched 'DANOS-TEST-RESULT: PASS')
$ python3 test/qemu_test.py device-manager acpi-ps2 pci-scan containment irqfree iopass
device-manager ... PASS (matched 'DANOS-TEST-RESULT: PASS')
acpi-ps2 ... PASS
pci-scan ... PASS (matched 'DANOS-TEST-RESULT: PASS')
containment ... PASS (matched 'DANOS-TEST-RESULT: PASS')
irqfree ... PASS (matched 'DANOS-TEST-RESULT: PASS')
iopass ... PASS (matched 'DANOS-TEST-RESULT: PASS')
```
Two companions cover what `hpet` can't, because it never exits:
- **`irqfree`** — the teardown path. Binds two owners to one shared endpoint, releases
one, and reads the I/O APIC back: the departing owner's line is masked, the sibling's
is not. That second half is why bindings are keyed on the owning *task* and not on
the endpoint pointer — endpoints are shared, so releasing "everything pointing at
this endpoint" would silently mask a live driver's device.
- **`iopass`** — the `device_grant` teardown rule, so destroying a driver's address
space never returns MMIO frames to the RAM pool.
## What's next (not done here)
The big driver-model pieces — capability passing (class drivers), DMA + barriers, MSI,
-210
View File
@@ -1,210 +0,0 @@
# M17–M18 execution plan: process lifecycle + device manager
**Archived — completed 2026-07-13** (every item checked; suite ended 54/54).
Kept as the record of how M17–M18 landed; the successor is
[m19-m20-plan.md](m19-m20-plan.md).
The operational plan for building [process-lifecycle.md](process-lifecycle.md)
(M17) and [device-manager.md](device-manager.md) increments 5–7 (M18). Design is
settled in those documents; this file is the build order — one phase at a time,
each phase green before the next starts. Delete or archive this file when M18
lands.
**Definition of green, every phase:** `zig build` clean, `zig build test` clean,
`python3 test/qemu_test.py` passes (existing scenarios plus the phase's new one),
and the relevant design doc's "known gaps" / status lines updated. Commit per
green phase (no co-author trailers).
**Workflow (settled 2026-07-12):** work happens in a dedicated git worktree, on
feature branches cut from `main` — `feat/process-lifecycle` (M17.1–17.4),
`feat/device-manager` (M18.1), `feat/usb-xhci-bus` (M18.2–18.3). When a branch's
phases are all green it is **auto-merged into `main`**; branches are kept after
merge, not deleted. Merges and branches are pushed to origin. Phase 0 (once):
commit the design docs, merge the outstanding `feat/usb` work into `main`, and
run the existing QEMU suite green before any new work starts.
**Numbering note:** continues the milestone sequence (driver track ended at M16).
## Status
The loop marks a phase `[x]` in the same commit that lands it. A phase is marked
only when its definition of green holds.
- [x] **Phase 0** — baseline: docs committed, feat/usb merged to main, pushed;
`usb-xhci-libary.zig` renamed to `usb-xhci-library.zig`; existing QEMU
suite green from the worktree (48/48, 2026-07-12).
- [x] **M17.1** — kernel releases claims/MSI on death (claims: `releaseAllOwnedBy`
in the reap; MSI was already swept by `irq.releaseOwner`; `claim-release`
test; suite 49/49)
- [x] **M17.2** — exit reasons (`ExitReason` recorded at exit/fault/kill before
the notification; `process_exit_reason` supervisor-gated;
`runtime.process.exitReason`; kernel + ring-3 assertions; suite 49/49)
- [x] **M17.3** — published exit events + VFS subscriber (`process_subscribe`,
bounded ref-counted table, publish on every death;
`runtime.process.subscribeExits`; VFS handles carry owners and are swept on
the owner's death; `vfs-client-death` test; suite 50/50)
- [x] **M17.4** — signals, timer notifications, `runtime.process`, the service
harness (signal_bind/process_signal + coalescing pending mask; timer_bind
on the tick; bindSignals/signalsFrom/sendSignal/stop + timerOnce;
runtime.service.run with the zero-length ping; VFS converted; `signals`
scenario; suite 51/51)
- [x] **merge** `feat/process-lifecycle` → main, push (merged 2026-07-13)
- [x] **M18.1** — device-manager protocol: hello + restart policy
(device-manager-protocol module; the manager as a harness service:
supervised spawns, hello deadline via timer sweep, restart with
300/600/1200ms backoff, exit reasons deciding restart-vs-stopped,
crash-loop cap; usb-xhci-bus first conforming driver; crash-test fixture
re-proving claim release each respawn; `driver-restart` scenario;
maximum_tasks 16→32 — the sweep was overflowing the pool; suite 52/52)
- [x] **merge** `feat/device-manager` → main, push (merged 2026-07-13)
- [x] **M18.2** — xHCI port scan + tree reports (child_added/child_removed in
the protocol; the manager's child mirror with death-pruning; xHCI maps the
register BAR — resource 0 is ECAM — reads CAPLENGTH/HCSPARAMS1, scans
PORTSC, reports connected ports with speed-class identity; `usb-report`
scenario proves report → prune → respawn → re-report; suite 53/53)
- [x] **M18.3** — app surface: enumerate/subscribe over IPC (subscriber
endpoint rides as the call's capability; events are the same structs the
buses send); device-list first client; protocol capped at the kernel's
IPC MESSAGE_MAXIMUM (256); the startUserTask debug print removed — it
sheared concurrent serial lines and was the scenario-flake root cause;
`device-list` scenario; suite 54/54)
- [x] **merge** `feat/usb-xhci-bus` → main, push (merged 2026-07-13) — **plan complete**
---
## M17.1 — the kernel releases a dead process's claims
The cleanup half of iron rule 1; the prerequisite for every restart story.
- `system/kernel/devices-broker.zig`: `releaseAllOwnedBy(owner: u32)` — clear
every `claimed[]` slot holding this task id.
- `system/kernel/process.zig`: call it from the reap path, alongside the existing
IRQ-binding release (the ordering comment there says why IRQs go first — claims
slot in after them, before the exit notification).
- MSI vectors: find where `msi_bind` records per-device vectors (interrupts
module) and release those by owner in the same pass.
- Docs: remove the claims bullet from process-management.md "Known gaps".
**Test:** new QEMU scenario `claim-release` — a test child claims an unclaimed
device, is killed, is respawned, and claims the same device again successfully;
assert both claims in the serial log. Kernel-side unit coverage in
`system/kernel/tests.zig` for `releaseAllOwnedBy` (claim two devices as two owners,
release one owner, verify exactly its claims freed).
## M17.2 — exit reasons
- `system/abi.zig`: `ExitReason` (exited, aborted, segmentation_fault,
illegal_instruction, arithmetic_fault, killed).
- Kernel: record the reason at every death site — clean exit path, each fault
class in `onException`, the kill path. Bounded recent-exits table (ids are never
reused, so a small ring keyed by id is enough).
- New system call `process_exit_reason(id)` — supervisor-gated, like kill; returns
the recorded reason or `-ESRCH` once evicted.
- `library/runtime/process.zig`: `ExitReason` + `exitReason(id: u32)`.
- Docs: remove the no-exit-status bullet from process-management.md.
**Test:** extend the `supervision` scenario — three children: one exits cleanly,
one faults (the fault-recovery pattern), one is killed; the supervisor asserts all
three reasons.
## M17.3 — published exit events
- Kernel: bounded subscriber table (endpoints); new system call
`process_subscribe(endpoint)` (ungated, like `process_enumerate`); every death
posts `notify_exit_bit | id` to each subscriber — the same post the supervisor
path already uses.
- `library/runtime/process.zig`: `subscribeExits(endpoint)`.
- VFS becomes the first subscriber: on an exit event, release every handle keyed
by that task id (badges already are task ids). Log the release.
- Docs: note the convention in ipc.md (exit events reuse the exit-notification
badge encoding).
**Test:** new QEMU scenario `vfs-client-death` — a client opens a file and is
killed without closing; assert the VFS logs the handle release and its open-handle
count returns to baseline.
## M17.4 — signals and the service harness
- Kernel: per-task pending mask + bound endpoint; system calls
`signal_bind(endpoint)` and `process_signal(id, signal)` (supervisor-or-self
gated); delivery posts `notify_signal_bit | pending mask`, coalescing; pending
signals with no bound endpoint pend silently.
- `library/runtime/process.zig`: `Signal`, `SignalSet`, `bindSignals`,
`signalsFrom`, `sendSignal`, `stop(id, deadline_ms)` (terminate → wait for exit
notification → kill). Implement `terminate`, `reload`, `user_1`, `user_2`;
`interrupt`/`quit` are enum members with no sender yet; `alarm` stays unbuilt.
- Kernel: **one-shot timer notifications** — `timer_bind(endpoint, ms)` posts a
notification badge when the deadline lands (IRQ-as-IPC again, on the timer
wheel `sleep` already uses). This is the missing timed-wait primitive:
`replyWait` blocks forever and `sleep` blocks the whole process, but `stop()`'s
escalation, the device manager's `hello` deadline (M18.1), and restart backoff
all need a deadline while staying responsive. It is also the mechanism `alarm`
gets for free later.
- New `library/runtime/service.zig`: the harness — `run(callbacks)` owning the
replyWait loop, folding protocol messages, signals, and child-exit notifications
into `init` / `on_message` / `on_reload` / `on_terminate`; answers the common
`ping` automatically. Define the reserved `ping` request encoding here and
document it in ipc.md (one obvious encoding; smallest that cannot collide with
existing protocols).
- Convert one existing service (input-source or hpet) to the harness as proof it
subtracts code rather than adding it.
**Test:** extend `supervision` — a harness-built child: `sendSignal(reload)`
observed in its log, `ping` answered, `stop()` produces a clean exit with reason
`exited`; a second child that ignores signals (no bind) is killed by `stop()`'s
deadline with reason `killed`.
## M18.1 — device-manager protocol: hello + restart policy
- New `system/services/device-manager/device-manager-protocol.zig` module
(vfs-protocol pattern): `hello { version, role, device_id }`; version constant;
reserved fields.
- Device manager: register the `.device_manager` endpoint; spawn drivers with its
exit endpoint; enforce the hello deadline; restart policy — backoff, crash-loop
cap (three fast deaths → mark failed, log, stop), reasons from M17.2 deciding
restart vs not.
- usb-xhci-bus: adopt the harness + send hello. hpet/ps2-bus follow only if the
conversion is mechanical; otherwise they keep working unconverted (the manager
only enforces hello on drivers spawned with an assignment).
- build.zig: test-loop entry for the protocol module if it grows pure logic.
**Test:** new QEMU scenario `driver-restart` — the xHCI driver takes a test-only
argv flag to fault after hello on its first run; assert: fault, exit reason
recorded, manager respawns with backoff, second run claims the controller
(M17.1) and hellos clean. Assert the crash-loop cap by a driver that always
faults (a tiny test driver, not xhci).
## M18.2 — bus tree reports
- Protocol: `child_added { parent, identity, resources }` / `child_removed { id }`.
- usb-xhci-bus: bring-up to **port scan only** — map the MMIO window (claimed in
M16-era work), controller reset/start per xHCI spec, walk the port registers,
report one `child_added` per connected port with speed + port number as
identity. **No transfer rings, no descriptors** — reading device/interface
descriptors (and therefore USB class triples for matching) is the follow-on USB
track, not this plan.
- Device manager: mirror reports into its tree; prune the subtree (emitting
`child_removed`) when a bus driver dies; assert re-report on restart.
**Test:** QEMU already attaches usb-kbd + usb-mouse on xhci.0 — assert two
`child_added` events reach the manager and appear in its tree dump; kill the
driver, assert two `child_removed` then two fresh `child_added` after respawn.
## M18.3 — the application surface
- Protocol: `enumerate` (tree snapshot) + `subscribe` (published add/remove
events, input-service pattern).
- A small client (`device-list`, the `ps` analog) exercising both; the manager
becomes the one answer to "what devices exist" for user space.
`device_enumerate` stays for drivers/kernel seeding — its retreat is tied to the
discovery migration, out of this plan.
**Test:** QEMU scenario — `device-list` shows the tree including USB children;
during a driver restart the subscribing client logs remove + add events.
---
**Explicitly out of scope** (own tracks, after M18): discovery migration (pci-bus
driver, acpi service, retiring the kernel scan), USB control transfers +
descriptors + class-driver matching, the musl layer, `interrupt`/`quit` senders
(needs a console), job control.
-196
View File
@@ -1,196 +0,0 @@
# M19–M20 execution plan: discovery migration
The operational plan for [device-manager.md](device-manager.md)'s increment 8:
discovery leaves the kernel — a **pci-bus driver** (M19) and an **acpi service**
(M20), with the kernel's device enumeration retired behind them. Same rules as
[m17-m18-plan.md](m17-m18-plan.md): one phase at a time, each green before the
next; this file is the build order and the checklist.
**Definition of green, every phase:** `zig build` clean, `zig build test` clean,
`python3 test/qemu_test.py` passes (existing scenarios plus the phase's new
one), and the relevant design doc updated. Commit per green phase (no co-author
trailers). The full suite is the regression net — the existing
`driver-restart` / `usb-report` / `device-list` / `input` scenarios must stay
green *through* the migration, which is the whole point: the system must not be
able to tell who enumerated it.
**Workflow:** dedicated worktree; branches off `main` — `feat/pci-bus`
(M19.0–19.3), `feat/acpi-service` (M20.1–20.3); auto-merge to main when a
branch is green; keep branches; push everything.
## Settled decisions (2026-07-13 — veto before the loop starts)
1. **What "retiring the kernel scan" means.** The kernel keeps, forever, the
parses it needs before user space exists: RSDP/XSDT location, MADT (SMP),
the HPET table (the tick), FADT + the AML `\_S5` evaluation (poweroff — the
power tests prove it), and MCFG (the host bridge node). What retires is
**device enumeration**: the ECAM function walk (M19.3) and the DSDT/SSDT
namespace walk that builds device nodes (M20.3). The AML module stays a
shared build module compiled into both the kernel (for `\_S5`) and the acpi
service (for everything else) — same source, two builds, no fork.
2. **Bridge apertures come from the firmware memory map, not AML.** Registered
PCI functions carry BAR resources, and containment demands the bridge own
windows that cover them. The apertures are derived kernel-side from the
boot memory map's MMIO holes (regions that are neither RAM nor tables) —
mechanical, AML-free, and available at boot regardless of what later moved
to user space. (The bridge today carries only ECAM + bus range; this is the
prerequisite M19.0 exists for.)
3. **`device_register` becomes idempotent on exact match.** A re-registration
with identical (parent, class, resources) returns the existing id instead
of appending. The kernel table has no unregister, so without this a
restarted registering bus would duplicate its children on every respawn —
idempotence makes restart-and-re-report safe for every future bus, not just
PCI.
4. **The manager matches from reports.** `ChildAdded` gains a `device_id`
field (the kernel-registered id, `no_device` for unregistered leaves like
USB ports). After the M19.3 flip, PCI driver matching keys off reported
identity (the class triple) instead of the manager's boot-time snapshot —
the snapshot match remains only for what the kernel still seeds. One flip
phase changes both sides at once so no device is ever matched twice.
5. **The acpi service's authority is one node.** The kernel publishes an
`acpi-tables` device: memory resources covering the table blobs plus a
broad `io_port` resource — the documented trust grant to exactly one
process (AML OperationRegions reach EC/PM ports; the claim-gated
io_read/io_write calls already exist). The service claims it, maps the
tables, and runs the shared AML module in ring 3 behind a `Hal` backed by
`mmio_map` + `io_read`/`io_write`.
6. **Both new processes are protocol drivers** under the manager: hello,
supervision, restart with backoff — all inherited from M18.1 for free.
Registration idempotence (decision 3) is what makes their restarts sound.
7. **Firmware neutrality is the contract** (2026-07-13). The generic layer is
everything at and above the device-manager protocol — descriptors,
containment, reports, matching, supervision — and none of it may become
x86-specific. Discovery is one swappable process per firmware: the acpi
service on x86; an **fdt service** on the Raspberry Pis (claims a
`devicetree-blob` node, reports children from the flattened device tree —
pure data, no bytecode, no port grant, strictly simpler than ACPI). The
manager owns the tree as *data* and touches no hardware, ever — AML runs in
a crashable, supervised discoverer precisely so a firmware-bytecode fault
can never take down the supervisor. Two consequences recorded now:
`DeviceDescriptor`'s 8-byte `hid` cannot hold an FDT `compatible` string
("brcm,bcm2835-aux-uart") — identity widens before the fdt service exists;
and cross-firmware surfaces are named by **domain, not firmware** (M21
defines a *power* protocol, not an "ACPI events" protocol — PSCI/mailbox
sources feed the same subscribers on ARM). **Landed early (2026-07-13):**
both services exist as placeholders (system/services/acpi, system/services/
fdt) and the build's `-Ddiscovery=acpi|fdt` option fills the ramdisk's
neutral `discovery` slot — the manager will spawn "discovery" by that name
in M20.3 and never learn which firmware it is on.
## Status
- [x] **M19.0** — prerequisites (bridge apertures from the memory map's
*gaps* — the single-hole rule died on OVMF's flash at the top of 4 GiB,
caught by the new every-BAR-contained assert in `discovery`; idempotent
`device_register` proven in `bus`; `ChildAdded.device_id`;
m17-m18-plan.md archived; suite 54/54).
- [x] **M19.1** — pci-bus driver, scan only (claims the bridge, maps ECAM
through its grant, brute-force walk with the multifunction rule; the
manager matches pci_host_bridge → pci-bus per device with the full
protocol contract; `pci-scan` builds its expected marker from the
kernel's own count — equivalence on the first run; suite 55/55).
- [x] **M19.2** — register + report (BAR probe mirrored byte-for-byte from the
kernel's addBars so dedupe returns the kernel's node ids during
coexistence; the bridge gained the io_port aperture I/O BARs need;
reports carry the registered device_id; pci-scan drills a forced restart
and asserts the PCI node count never grows — plus harness hardening: a
failing case now preserves its serial as <case>-failed-serial.log, and
the heavy scenarios run at 150s; suite 55/55).
- [x] **M19.3** — the flip: kernel `enumeratePci`/`addBars`/`PciHeader` all
deleted (bridge node stays); manager matches PCI drivers from reported
identity, deduped by registered id. Surfaced and fixed a real SMP race the
flip created — ring-3 device_register made the broker table concurrent, so
mmio_map's lock-free read intermittently tore hpet's resource length
(user fault) and overflowed `r.len-1` into a kernel panic; now the broker
read is under the big lock and the arithmetic is guarded, and pci-bus
skips size-0 BARs. discovery.md updated; suite 55/55 (driver-restart
hammered 6×).
- [x] **merge** `feat/pci-bus` → main, push (merged 2026-07-13).
- [x] **M20.1** — acpi service, parse only: the AML interpreter is now a build
module compiled into both kernel and service; the kernel publishes the
`acpi-tables` node (AML blobs as memory resources, the broad io_port grant,
the SCI); the service claims it, maps the blobs, runs the shared parser in
ring 3, and self-verifies its Device count against the kernel's (34 = 34,
deterministic via argv, no log-scraping); the manager spawns `discovery`
at startup. Parse-only touches no hardware. Suite 56/56.
- [x] **M20.2** — register + report: the service evaluates `_STA`/`_CRS` in
ring 3 (interpreter Hal = port I/O over the claimed node; a scratch page
backs SystemMemory maps so a stray region can't fault it) and registers +
reports each present `_HID` device under `acpi-tables`. Containment: the
broker's irq check became range-based (len-1 == the old equality) so the
node's broad irq window covers children's legacy lines; io ports fall in
the broad io grant. ChildAdded gained `hid`. Matching stays off. The
`acpi-report` scenario asserts the PS/2 keyboard (3 resources) and mouse
(1 resource) among the reports. Suite 57/57.
- [x] **M20.3** — the flip: the kernel's `wireAcpiDevices` call is gone (the
device-building helpers are retained-but-dead pending a focused sweep,
spawned as a task; static tables + `\_S5` + the acpi-tables node stay).
The manager matches ps2-bus from ACPI `_HID` reports; the service
registers all devices before reporting any (no keyboard-before-mouse
race). The `acpi-ps2` scenario proves report → spawn → ps2-bus attaches
its keyboard; `ioport` retargeted to the acpi-tables I/O window (the
kernel-built PS/2 node is gone). Suite 58/58.
- [x] **merge** `feat/acpi-service` → main, push (merged 2026-07-13) — **discovery migration complete**.
---
## Phase notes
**M19.0 apertures:** the boot memory map already crosses the handoff
([boot-handoff]), but discovery never sees it today — expect a small
pass-through (kernel init hands the map to the platform layer) before the
holes computation, which belongs where the bridge node is built
(`parseMcfg`). Sanity-check on QEMU q35: the xHCI BAR (`0xc0000000`-region
values seen in the M18 logs) must land inside a derived aperture, asserted in
the kernel unit test.
**M19.1 scanning without owning config access twice:** the driver reads config
space through its ECAM mmio_map grant of the *bridge* window — the same bytes
the kernel walk read. Vendor-id `0xFFFF` skip, header-type multifunction rule,
no bridge recursion (matches the kernel's current single-segment walk).
**M19.2 BAR sizing:** the classic size probe (write all-ones, read mask,
restore) is deferred — the BARs' current programmed values and types are
enough for containment-checked registration at bring-up; sizing lands with the
first driver that needs to *move* a BAR. Log what is registered so the
scenario can assert it.
**M19.3 what the manager still seeds from the snapshot:** everything the
kernel still enumerates (timers, ACPI nodes until M20.3). The PCI arm of
`pciDriverFor` switches source; `driverFor` doesn't move until M20.3.
**M20.1 spawn and identity (pre-settled 2026-07-13):** the manager spawns
`discovery` by its neutral ramdisk name at startup, as an ordinary protocol
driver (hello, supervision) — from M20.1 on, on every boot. For reporting ACPI
devices, `ChildAdded` gains `hid: [8]u8` (EISA ids fit; zero = none):
firmware *string* identity travels beside the numeric `identity` field until
the FDT-driven widening replaces both (decision 7).
**M20.1 Hal in ring 3:** `mapMmio` → `device.mmioMap` over the claimed
acpi-tables node (plus a table-offset map for blobs); `pioRead`/`pioWrite` →
`device.ioRead`/`ioWrite` against its io_port resource. The interpreter cannot
tell it moved — that is the assertion of `acpi-parse`.
**M20.2 containment for `_CRS`:** io ports fall inside the node's broad
io_port resource; MMIO windows (HPET, LAPIC ranges some firmwares list) fall
inside the memory-map holes added to the node in M20.1. Anything that doesn't
fit is logged and skipped, loudly — bring-up honesty over silent drops.
**M20.3 ps2 ordering:** ps2-bus binds nodes the acpi service now reports, so
its spawn moves behind the report (the manager's matching handles this once
the source flips); the `input` scenario proves the keyboard still types.
**Explicitly out of scope:** PCI bridge recursion (single segment, flat bus
walk stays); BAR reprogramming/sizing; disk/PCIe hotplug; interrupt routing
changes (`_PRT` stays wherever it is today); the USB descriptor track;
multi-segment ECAM; per-device power states (D-states, `_PSx`/`_PRx`,
suspend/resume — a future *lifecycle-vocabulary* extension, since "suspend"
has the shape of a signal every driver must answer, and it has no consumer
until laptop sleep); CPU P/C-states.
## M21 — ACPI events + system power — DONE
Built and merged (docs/m21-plan.md, 2026-07-13): the SCI + power button, Notify/GPE
dispatch, and orderly shutdown (init's stop cascade into a ring-3 S5 write).
See that plan for the phase record.
-138
View File
@@ -1,138 +0,0 @@
# M21 execution plan: ACPI events + system power
The operational plan for the event side of the acpi service and orderly
shutdown — the capstone [m19-m20-plan.md](m19-m20-plan.md) previewed. Same
rules as its predecessors: one phase at a time, each green before the next;
this file is the build order and the checklist.
**Definition of green, every phase:** `zig build` clean, `zig build test`
clean, `python3 test/qemu_test.py` passes (existing scenarios plus the
phase's new one), and the relevant design doc updated. Commit per green phase
(no co-author trailers). Failing cases preserve their serial logs
(`<case>-failed-serial.log`).
**Workflow:** dedicated worktree; branch `feat/power-events` off `main`;
auto-merge to main when the branch is green; keep the branch; push everything.
## Settled decisions (2026-07-13, approved)
1. **S5 is executed by the acpi service from ring 3.** No new syscall: the
broad port grant (M20 decision 5) already made this physically possible —
the service holds the PM1 control ports in its io grant and derives `_S5`
from its own namespace (`aml.sleepState`). Formalizing it adds no
authority. The kernel keeps `power.zig` for its own test paths and
panic-time use.
2. **The power surface is domain-named** (decision 7 of the last plan): a
`power-protocol` module + `ServiceId.power = 5`, registered by the acpi
service — on ARM, a PSCI/mailbox service registers the same id and
subscribers never know the difference. Messages: `subscribe` (endpoint as
the call's capability, the input/manager pattern), `shutdown` (accepted
only from PID 1 — init), and events published as buffered messages:
`power_button`, `lid`, `ac`, `battery`, generic `notify` with a code.
3. **The service learns event ports from its own FADT copy**: the kernel adds
the FADT as one more memory resource on the acpi-tables node; the service
tells it apart from the AML blobs by signature ("FACP" header — the blob
resources are header-stripped bytecode and start with no signature). The
kernel's own FADT parse is untouched.
4. **The acpi service converts to the harness** (`runtime.service.run`):
protocol messages (subscribe/shutdown), the SCI notification, and the
existing report flow fold into one loop — the shape it was always meant
to have.
5. **GPE/Notify correctness is proven by host unit tests** (synthetic AML
with a Notify inside a method body; aml.zig joins the `zig build test`
loop). The QEMU scenario proves the power button — a *fixed* event,
deterministically injectable via QMP `system_powerdown` — because QEMU
cannot raise GPEs deterministically on this config. Battery/AC/lid and the
embedded controller (`_Qxx`) are interface-complete here and validated on
real hardware (the laptop) later.
## Ground truth the phases build on (verified 2026-07-13)
- `system/devices/power.zig` `shutdown()` is the kernel's S5 write
(SLP_TYP|SLP_EN to PM1a/PM1b control); there is no power syscall.
- init (`system/services/init/init.zig`) spawns vfs/input/device-manager
fire-and-forget — no child ids kept, no signals, no event loop. The whole
stop toolkit exists in `runtime.process` (stop/sendSignal/bindSignals).
- `test/qemu_test.py` has no QMP channel (serial is a one-way file).
- The kernel parses PM1 *control* blocks and SCI_INT from the FADT; the PM1
**event** blocks (offsets 56/60, len at 88) and **GPE0/GPE1** blocks
(offsets 80/84, lens 92/93) are unparsed — the service reads them from its
FADT copy (decision 3).
- The acpi-tables node carries the SCI as its only `len == 1` irq resource
(the broad window is len 256) — that is how the service finds it to
`irqBind`.
- `notify_opcode = 0x86` exists in `system/devices/aml/opcodes.zig` but the
interpreter never handles it — a GPE `_Lxx` body containing Notify fails
evaluation today. Everything else a GPE handler needs (field access,
control flow, method calls) is proven by the ring-3 `_STA`/`_CRS` work.
- The dead-code sweep (spawned task) also edits `system/devices/acpi.zig`;
M21.0 checks whether it landed and rebases before touching that file.
## Status
- [x] **M21.0** — baseline (dead-code sweep confirmed landed on main — no
acpi.zig conflict; `feat/power-events` cut; QMP channel in the harness:
always-on unix socket, client with the capabilities handshake, per-case
`qmp_after` hook, and a hook-must-deliver pass gate that the smoke case
now proves with a harmless query-status; suite 58/58).
- [x] **M21.1** — SCI + the power button (kernel appends the FADT as an
acpi-tables memory resource, tagged by its "FACP" header; `power-protocol`
module + `ServiceId.power = 5`; the acpi service converted to
`runtime.service.run`, registers `.power`, reads PM1 event/control + GPE
ports from its FADT copy, enables ACPI mode if SCI_EN is clear, binds the
SCI (the len-1 irq), sets PWRBTN_EN; the SCI handler clears PM1_STS,
logs `power: button pressed`, publishes `power_button`, acks. Scenario
`power-button` injects a real `system_powerdown` via QMP; initial-ramdisk
timeout 30→60s for the service's added boot work; suite 59/59).
- [x] **M21.2** — Notify + GPE dispatch (interpreter handles `notify_opcode`
into a bounded queue, cleared per-evaluate, drained via
`takeNotifications`; the service walks GPE status/enable bytes, evaluates
`\_GPE._Lxx`/`_Exx` per active bit, maps notified nodes to events
(battery/ac/lid/generic), clears GPE_STS write-1, acks. EC `_Qxx` out.
Host unit test with hand-encoded AML proves the queue; aml.zig joined the
`zig build test` loop. QEMU raises no GPEs — suite is regression net,
59/59).
- [x] **M21.3** — orderly shutdown (init supervises its children on one
endpoint that also carries signals, power events, and a re-arming
heartbeat timer; on `power_button` or a `terminate` signal it logs
`init: shutting down`, runs `stop(child, 2000, endpoint)` in reverse
order, then requests `.power` shutdown; the acpi service honors shutdown
from a subscriber — init is the one subscriber, a soft gate that survives
testing where PID 1 isn't init — and writes SLP_TYP|SLP_EN from ring 3.
`orderly-shutdown` scenario proves button → shutting-down → S5 → QEMU
exit; suite 60/60).
- [x] **merge** `feat/power-events` → main, push, keep the branch (merged 2026-07-13) — **M21 complete**.
---
## Phase notes
**M21.0 QMP:** open the unix socket after Popen, complete the
`qmp_capabilities` handshake, then send the hook's command (for these
scenarios: `{"execute": "system_powerdown"}`). The socket is additive — no
existing case may notice it. Note e3fe3f3 recently reworked how the harness
boots; adapt to its current shape rather than the pre-rework description.
**M21.1 SCI details:** PM1_STS is at the event block base (write-1-to-clear);
PM1_EN at base + block_len/2; PWRBTN bit is 8 in both. If PM1b exists, mirror
reads/writes to both blocks. Enable ACPI mode only when SCI_EN (PM1 control
bit 0) is clear — OVMF boots may already have it set. The publish path reuses
the manager's subscriber table pattern (bounded, drop-on-failed-send).
**M21.2 GPE walk:** GPE0_STS bytes live at the GPE0 block base, GPE0_EN in
the block's upper half; for a set+enabled bit n, the handler method is
`_L%02X` (level) or `_E%02X` (edge) under `\_GPE`. Evaluate, drain the notify
queue, clear the status bit, ack. A missing handler method is clear-and-log,
not an error.
**M21.3 ordering:** init subscribes with retries — the acpi service registers
`.power` well after init starts. The stop sequence runs vfs last (other
services may flush through it). The S5 write mirrors `power.zig`'s
`sleepValue` (SLP_TYP bits [12:10], SLP_EN bit 13); if the write returns, log
`power: S5 write did not take` so the scenario fails loudly instead of
hanging.
**Explicitly out of scope:** the embedded controller and `_Qxx` queries,
battery `_BST`/`_BIF` evaluation beyond the interface stubs, lid/AC on QEMU
(no emulation), reboot over the power protocol, S3 sleep, per-device D-states
(a future lifecycle-vocabulary extension), thermal zones.
+128
View File
@@ -0,0 +1,128 @@
# The power service: events and shutdown
A laptop lid closes, a battery drains, someone presses the power button — and
several parts of the system might care: a session manager dims the screen, a
logger notes it, and ultimately *something* has to turn the machine off. None of
them owns the hardware that reported the event, and the reporter should not know
who is listening. So system power is a **service**: an event source **publishes**
button/lid/battery/AC events, interested processes **subscribe**, and one
privileged caller — init — can ask it to power the machine off. It is the same
publish/subscribe shape as the [input service](input.md), applied to power.
## Why a service, and why it is named for the domain, not the firmware
Where the events come from is firmware-specific — on x86 they ride the ACPI SCI
([acpi.md](acpi.md)); on a Raspberry Pi they would come from PSCI or a mailbox.
What subscribers want is not: *the lid closed* means the same thing regardless of
who noticed. So the surface is **domain-named**. There is a `power-protocol`
module and a well-known `ServiceId.power = 5`; on x86 the **acpi service**
registers it, and on ARM a PSCI/mailbox service will register the *same* id.
Subscribers call `runtime.ipc.lookup(.power)` and never learn which firmware they
are on — the neutrality the whole [discovery](discovery.md) migration exists to
preserve, carried one layer up into a running-system surface.
This is why the protocol is `power`, not "ACPI events": naming a cross-firmware
surface after one firmware would leak x86 into code the ARM port must reuse
unchanged.
## The protocol
The `power-protocol` module ([system/services/power/protocol.zig](../system/services/power/protocol.zig))
follows the vfs-protocol pattern — extern-struct messages, a version, reserved
fields. Three operations:
| Direction | Operation | Purpose |
|---|---|---|
| subscriber → service | `subscribe` | receive published events; the subscriber's endpoint rides as the call's **capability** (the input/device-manager pattern) |
| init → service | `shutdown` | orderly shutdown's last step: enter S5 (soft off) |
| service → subscriber | `event` | a published `EventMessage`, delivered as a buffered message (never sent *to* the service) |
Events are published, not polled: like the input service, the service holds
subscriber endpoints as capabilities and `ipc_send`s each event as a buffered
message, so a slow or dead subscriber can never wedge the source. The event
vocabulary is hardware-neutral:
- `power_button` — the button was pressed (a fixed ACPI event on x86).
- `lid`, `ac`, `battery` — the named GPE-driven events.
- `notify` — a device notification that maps to none of the above; its `code`
(the ACPI `Notify` argument) and the notifying device's `hid` say which device
and what happened.
An `EventMessage` carries the `event` tag plus `code` and an 8-byte `hid`, so a
generic `notify` is fully described without a second round trip.
**`shutdown` is authority, not information.** It is the only operation that
*does* something irreversible, so it is gated: the contract is that only init
(PID 1) may request it, because init is the process that has already run the stop
sequence over everything else. The acpi service implements this as a **soft
gate** — it honors `shutdown` only from a process that is a *subscriber*, and
init is the one subscriber. That stands in for "only the system supervisor may
power off" without hard-coding a pid, so it still holds under tests where PID 1
is not init.
## Orderly shutdown
Powering off cleanly is where the power service, the [process
lifecycle](process-lifecycle.md), and [ACPI events](acpi.md) compose. init
already supervises the services it starts; for shutdown it runs **one event loop
over one endpoint** that carries three things at once: its children's exit
notifications, the lifecycle **signals** it can receive (`terminate`), and the
**power events** it subscribes to — plus a re-arming heartbeat timer proving PID
1 is alive. (init subscribes with retries, because the power service registers
`.power` well after init starts; a missing power service is not fatal — a
`terminate` signal drives the same path.)
On a `power_button` event or a `terminate` signal, init:
1. logs that it is shutting down,
2. runs the standard stop sequence — `runtime.process.stop(child, deadline,
endpoint)` — over its children **in reverse spawn order**, so the VFS stops
last (other services may flush through it), each child getting the
*terminate → deadline → kill* escalation from
[process-lifecycle.md](process-lifecycle.md), and
3. requests `.power` `shutdown`.
The service then enters **S5** (soft off) by writing `SLP_TYP | SLP_EN` to the
PM1 control register(s) from ring 3, mirroring the kernel's own
`system/devices/power.zig` `sleepValue`. If the write returns instead of powering
the machine off, it logs loudly so a test fails rather than hangs.
**No new system call was needed for S5.** The broad io_port grant on the
`acpi-tables` node ([discovery.md](discovery.md)) already put the PM1 control
ports in the acpi service's hands, so writing S5 from ring 3 is something it
could physically already do; formalizing it as a protocol operation added a
contract, not authority. The kernel keeps `power.zig` for its own test paths and
panic-time poweroff, where no user space is available to ask.
## Verifying it
Two QEMU scenarios exercise the path, both injecting a real ACPI power-button
press via QMP `system_powerdown` (there is no other deterministic power event on
this config):
- `power-button` proves the source: the acpi service's SCI handler logs the
press and publishes `power_button` (the ACPI half is in [acpi.md](acpi.md)).
- `orderly-shutdown` proves the whole composition: button → init logs shutting
down → children stopped → the service enters S5 → QEMU exits. The ordered
regex is the proof, and QEMU's self-exit through S5 is the pass.
## Scope
Interface-complete but validated on real hardware (the author's laptop) later,
because QEMU does not emulate them: battery `_BST`/`_BIF` evaluation beyond the
interface stubs, lid and AC events, and the embedded controller's `_Qxx`
queries. Deliberately out of scope for now: reboot over the power protocol, S3
sleep, per-device D-states (a future lifecycle-vocabulary extension, since
"suspend" has the shape of a signal every driver must answer and has no consumer
until laptop sleep), and thermal zones.
## See also
- [acpi.md](acpi.md) — where the events come from on x86: the SCI, the power
button fixed event, and GPE/Notify dispatch in the acpi service.
- [discovery.md](discovery.md) — why the surface is domain-named, and the
firmware neutrality that makes a PSCI backend drop-in on ARM.
- [process-lifecycle.md](process-lifecycle.md) — the stop sequence
(`terminate → deadline → kill`) and signals init composes into shutdown.
- [device-manager.md](device-manager.md) — the supervision model init mirrors for
its own children.
+1 -1
View File
@@ -11,7 +11,7 @@ because the kernel releases a dead process's claims. The `driver-restart` and
`usb-report` scenarios prove kill → release → respawn → re-claim → re-report
end to end. What remains of this document's ladder is scope, not mechanism:
more of the system moved into restartable processes (the discovery migration,
[m19-m20-plan.md](m19-m20-plan.md), is the next rung). This is the property danos is really chasing:
[discovery.md](discovery.md), is the next rung). This is the property danos is really chasing:
**if a part of the OS breaks, isolate it, and re-initialise it — without rebooting.**
A crashed driver gets restarted; a wedged service gets killed and brought back. It's
the reason the [microkernel](vision.md) shape was chosen, and it's a *separate* goal
+117
View File
@@ -0,0 +1,117 @@
# Timers and time
Two different needs hide under the word "timer", and danos keeps them apart:
- **Reading the clock** — *what time is it?* A read of a free-running counter.
- **Waiting** — *wake me in N milliseconds*, or *notify me when a deadline passes.*
Both are answered by the **kernel**, because the kernel already owns a timer: it has
to, to preempt tasks. The LAPIC heartbeat and the calibrated TSC that back all of this
are built in [device-interrupts.md](device-interrupts.md); the scheduler's blocking and
wait queues are in [scheduling.md](scheduling.md). This page is about the surface a
ring-3 program actually uses, and one deliberate absence: **there is no user-space time
service.**
## Why time is a syscall, not a service
The tempting microkernel move is to put a timer *driver* in user space and have
applications ask it for the time over IPC. For a **monotonic clock that is wrong** —
reading `now()` should never cost an IPC round trip. The kernel is already holding the
answer: it computes the current time every time it schedules, from the TSC, in a couple
of instructions. Surfacing that as a system call is pure mechanism; routing it through a
message to another process would be slower *and* redundant, and a device like the HPET
(uncacheable MMIO reads) is a particularly bad thing to read on every `now()`.
This is the same conclusion every serious system reaches: Linux and Zircon read the
counter in the vDSO, L4 exposes a clock field in a shared kernel page, seL4 reads the
cycle counter directly. None of them make a clock read an IPC. danos makes it a syscall.
That "from the TSC" hides a portability question, because the TSC is only a valid clock
when the CPU guarantees it is *invariant* and when every core's TSC is *synchronized*.
danos checks both — the invariant-TSC CPUID bit (`0x80000007` EDX[8], set on Intel and
AMD), and a cross-core "warp" check as the cores come up — and falls back to the HPET
counter when either fails. So `now()` stays accurate on a real Intel box, a real AMD box,
and inside a VM alike; only the source behind it differs. The mechanism is in
[device-interrupts.md](device-interrupts.md).
So the timer hardware lives in the kernel, and there is **no `hpet` driver and no time
server** to consume. (An earlier HPET driver existed only to *demonstrate* the driver
model; that role now lives in [drivers.md](drivers.md), as documentation.) The one place
a user-space time service *is* justified — **wall-clock / calendar time** — is discussed
at the end; it is deliberately not built yet.
## The three system calls
Time and waiting are three entries in the small syscall table ([syscall.md](syscall.md)):
- **`clock` (#23)** → monotonic nanoseconds since boot. It only moves forward. Not
wall-clock: no date, no timezone. Backed by `architecture.nanos()` (TSC, scaled with a
128-bit intermediate so a long uptime can't overflow) — a few nanoseconds of
resolution, and just an `rdtsc` plus a multiply.
- **`sleep` (#3)** → block the caller for N milliseconds. The scheduler records a wake
deadline and the tick sweep wakes it (`scheduler.sleep`).
- **`timer_bind` (#31)** → arm a one-shot timer that, after N milliseconds, posts a
**timer notification** to an IPC endpoint. Unlike `sleep` it does **not** block: a
service can keep answering messages on the same endpoint while a deadline is pending.
This is the timed wait that stop-sequence escalation, hello deadlines, and restart
backoff are built from ([process-lifecycle.md](process-lifecycle.md),
[device-manager.md](device-manager.md)).
The kernel's own scheduling timer (the LAPIC, vector 32) is never exposed to user space;
programs read the TSC through `clock` and get timed wakeups through `sleep`/`timer_bind`,
both riding the scheduler tick.
## `runtime.time` — the generic interface
Applications don't call the syscalls directly; they use `runtime.time`
(`library/runtime/time.zig`), a thin `Instant`/`Duration` layer over them — an ergonomic
front door, not new mechanism.
```zig
const time = @import("runtime").time;
const start = time.now(); // Instant — monotonic
doWork();
const took = start.elapsed(); // Duration
time.sleep(time.Duration.fromMillis(5)); // block ~5 ms
// A deadline delivered as a notification, so a service keeps serving meanwhile:
_ = time.after(endpoint, time.Duration.fromMillis(200));
```
- `Duration` is nanoseconds under the hood, with `fromNanos/fromMicros/fromMillis/
fromSeconds` and `asNanos/asMillis`. `ceilMillis` rounds *up* to the kernel's
millisecond granularity, so a sub-millisecond `sleep` never rounds down to zero and
returns early. All arithmetic saturates rather than wraps.
- `Instant` is a point on the monotonic clock: `since`, `elapsed`, `plus`, `reached` —
built for deadline loops (`while (!deadline.reached()) …`).
- `now()` / `monotonicNanos()` wrap `clock`. `available()` reports whether the clock is
calibrated at all (the kernel returns 0 until the TSC frequency is known, so a caller
that needs real time can treat 0 as "unavailable" rather than assume it advances).
- `sleep(d)` wraps `sleep`; `spin(d)` busy-polls `now()` for the sub-millisecond delays
the millisecond tick can't express; `after(endpoint, d)` wraps `timer_bind`.
The raw wrappers (`system.clock`, `system.sleep`, `system.timerOnce`) stay in
`library/runtime/system.zig`; `runtime.time` is the layer meant for everyday use.
## Wall-clock time (not built)
Everything above is **monotonic**: elapsed time since boot, perfect for timeouts and
measurement, useless for "what is the date?" Calendar time — a real-time clock, time
zones, leap seconds — is genuinely a **user-space** concern, and it *is* the case a time
service is for. It would be backed by an **RTC** driver (the CMOS real-time clock), not
the HPET, and exposed as a `CLOCK_REALTIME`-style service alongside the monotonic
syscall. It is deferred until something needs it; the monotonic clock the kernel already
owns covers every current use.
## Verifying it
`runtime.time`'s `Instant`/`Duration` arithmetic has unit tests that run on the host:
```
$ zig build test # includes library/runtime/time.zig
```
End to end, the proof the clock is real is that it *advances*: read `now()`, `sleep` a
`Duration`, read `now()` again, and the second reading is later — the kernel's timer
driving a ring-3 program with no service in between.