Add runtime.time, drop demo drivers, harden TSC timekeeping
Time is a kernel concern in danos: the kernel owns the scheduling timer and already exposes monotonic time via the clock/sleep/timer_bind syscalls, so a userspace time service would be a redundant, slower path. This adds the generic runtime.time module over those syscalls, retires the two demonstration drivers, reorganizes the milestone docs, and makes the monotonic clock correct on Intel, AMD, and inside any VM. runtime.time (library/runtime/time.zig) - Instant/Duration interface: now, sleep, spin, after, monotonicNanos, available - a thin layer over system.clock/sleep/timerOnce; unit-tested arithmetic Remove the demo drivers hpet and bus (a teaching example belongs in the docs, not shipped in the tree) - system/drivers/ now holds only real drivers: pci-bus, ps2-bus, usb-xhci-bus - device-manager end-to-end test repointed to pci-bus (asserts on kernel state: the process table and the device tree, not a racy serial marker) - device_register containment moved to a new in-kernel `containment` test - the driver-model worked example moved inline into docs/drivers.md Reorganize milestone docs into topic docs - m17-m18 / m19-m20 / m21 plans dissolved into process-lifecycle, device-manager, discovery, and acpi docs; new docs/power.md and docs/timers.md; ~20 citations repointed; plan docs deleted TSC reliability (apic.zig, smp.zig, cpu.zig, kernel.zig) - check the invariant-TSC bit (CPUID 0x80000007 EDX[8]) on Intel and AMD - cross-core "warp" check at SMP bring-up, pairwise BSP<->AP as each core comes up - fall back to the HPET clocksource when the TSC is not invariant (a bare VM) or not synchronized (a warp), switched continuously so time never jumps - boot log reports the outcome; new tsc-sync test exercises the TSC + warp path Verified: zig build; zig build test; 60/60 QEMU cases (incl. new containment and tsc-sync).
This commit is contained in:
+10
-1
@@ -100,7 +100,16 @@ Cutting across all of these:
|
||||
when to build it, and how to keep it architecture-agnostic.
|
||||
- **[acpi.md](acpi.md) — finding the ACPI tables.** The concrete x86 locator chain:
|
||||
how the loader captures the **RSDP**, hands its physical address across in `BootInfo`,
|
||||
and how the platform derives the **RSDT/XSDT** from it and walks the SDTs.
|
||||
and how the platform derives the **RSDT/XSDT** from it and walks the SDTs — plus the
|
||||
live event side (the SCI, the power button, GPE/Notify) the ring-3 acpi service runs.
|
||||
- **[power.md](power.md) — the power service.** System power as a domain-named
|
||||
service: button/lid/battery events published to subscribers, and init's orderly
|
||||
shutdown composing the [lifecycle](process-lifecycle.md) stop sequence with an ACPI
|
||||
S5 write. Firmware-neutral — a PSCI backend drops in on ARM.
|
||||
- **[timers.md](timers.md) — timers and time.** The ring-3 surface for reading the
|
||||
clock and waiting: why `now()` is a syscall rather than a service, and the one-shot
|
||||
timer notification (`timer_bind`) that gives supervisors a timed wait — built on the
|
||||
LAPIC heartbeat and calibrated TSC of [device-interrupts.md](device-interrupts.md).
|
||||
- **[smp.md](smp.md) — multiple cores.** A design/research note on how microkernels
|
||||
(L4, seL4) handle SMP — big kernel lock vs per-CPU vs multikernel — and how the
|
||||
right choice depends on whether danos is chasing real-time or resilience.
|
||||
|
||||
+55
-1
@@ -107,12 +107,66 @@ firmware-agnostic [device model](discovery.md) gets populated; this note stops a
|
||||
part that answers "where are the tables?" — everything past the RSDP is just following
|
||||
more pointers the tables themselves provide.
|
||||
|
||||
## ACPI events: the SCI, the power button, and GPEs (M21)
|
||||
|
||||
The tables above are static description; ACPI is also a *live* channel. Hardware
|
||||
raises the **SCI** (System Control Interrupt) — one shared, level-triggered line
|
||||
whose vector the FADT names — and the OS reads status registers to learn what
|
||||
happened: a fixed event like the power button, or a **General-Purpose Event**
|
||||
(GPE) whose handler is an AML method. Since [discovery](discovery.md) moved AML
|
||||
to ring 3, the event side lives there too, in the same **acpi service** — the
|
||||
device discoverer and the event source are one process, because both need the
|
||||
namespace and the port grant.
|
||||
|
||||
**The kernel hands the service what it needs and no more.** Reading PM1 event
|
||||
blocks and GPE blocks requires the FADT, which the kernel already parses for its
|
||||
own `\_S5` poweroff. Rather than re-parse, the kernel appends the **FADT as one
|
||||
more memory resource** on the `acpi-tables` node; the service tells it apart
|
||||
from the AML blob resources by signature — the FADT keeps its intact `"FACP"`
|
||||
header, while the blob resources are header-stripped bytecode that starts with
|
||||
no signature. The kernel's own FADT parse is untouched; the service reads the
|
||||
PM1 *event* blocks (which the kernel never parsed — it only needs PM1 *control*
|
||||
for `\_S5`) and the GPE0/GPE1 blocks straight from its copy. The **SCI itself**
|
||||
arrives as the node's one `len == 1` irq resource (distinct from the broad
|
||||
`[0, 256)` window that covers children's legacy lines), which is how the service
|
||||
finds the line to `irq_bind`.
|
||||
|
||||
With those in hand the service enables ACPI mode (only if `SCI_EN` is clear —
|
||||
some firmwares boot with it already set), sets `PWRBTN_EN`, and on each SCI:
|
||||
|
||||
- **The power button** is a *fixed* event: a set `PWRBTN_STS` bit in PM1 status.
|
||||
The handler clears it (write-1-to-clear), logs the press, and publishes a
|
||||
[`power`](power.md) `power_button` event to subscribers.
|
||||
- **GPEs** are the general path: for each set-and-enabled GPE bit `n`, the
|
||||
service evaluates its `\_GPE._L%02X` (level) or `_E%02X` (edge) handler
|
||||
method, drains the **Notify** queue that method produced, maps each notified
|
||||
device to an event (battery, AC, lid, or a generic `notify` with its code),
|
||||
and clears the status bit. A missing handler method is clear-and-log, not an
|
||||
error. Making GPEs work required teaching the interpreter one opcode it never
|
||||
handled — `Notify` (`0x86`) — which it now folds into a bounded queue drained
|
||||
per evaluation; everything else a handler needs (field access, control flow,
|
||||
method calls) was already proven by the ring-3 `_STA`/`_CRS` work.
|
||||
|
||||
**How this is tested.** QEMU cannot raise GPEs deterministically on this config,
|
||||
so GPE/Notify correctness is proven by **host unit tests** — hand-encoded AML
|
||||
with a `Notify` inside a method body, run under `zig build test`. The QEMU
|
||||
`power-button` scenario proves the fixed-event path end to end: a QMP
|
||||
`system_powerdown` injects a real ACPI power-button press, and the service's SCI
|
||||
handler must log it. Battery/AC/lid and the embedded controller's `_Qxx` queries
|
||||
are interface-complete but validated on real hardware later.
|
||||
|
||||
The service surface these events are *published on* — subscription, the event
|
||||
vocabulary, and orderly shutdown — is the power service, [power.md](power.md).
|
||||
|
||||
## Related
|
||||
|
||||
- [efi.md](efi.md) — the loader that captures the RSDP before `ExitBootServices`.
|
||||
- [memory-map.md](memory-map.md) — the same loader-captures / kernel-consumes seam, and
|
||||
the ACPI-reclaim memory the RSDP lives in.
|
||||
- [discovery.md](discovery.md) — the broader (still-evolving) plan for turning these
|
||||
tables into one neutral device model shared with the ARM device-tree path.
|
||||
tables into one neutral device model shared with the ARM device-tree path, and how
|
||||
ACPI enumeration and events moved to the ring-3 acpi service.
|
||||
- [power.md](power.md) — the domain-named power service the ACPI event side publishes
|
||||
to (button, lid, battery) and its orderly-shutdown path into S5.
|
||||
- [arch.md](arch.md) — why the kernel reaches the device code through a `platform`
|
||||
module and never names ACPI directly.
|
||||
|
||||
@@ -96,7 +96,7 @@ That's all — no Unix-abbreviation exception. The source directories are full w
|
||||
(`system`, `library`, not `src`/`lib`), and there is no daemon `d` suffix: a driver
|
||||
lives in `system/drivers/` and a service in `system/services/`, so the *location*
|
||||
already says what it is. Encoding the role in the name too (`busd`, `vfsd`) is
|
||||
redundant — the program is just `bus`, `vfs`. Don't put in a name what its directory
|
||||
redundant — the program is just `ps2-bus`, `vfs`. Don't put in a name what its directory
|
||||
already tells you.
|
||||
|
||||
## A note on collisions
|
||||
@@ -138,7 +138,7 @@ single word or acronym needs no hyphen: `scheduler.zig`, `paging.zig`, `apic.zig
|
||||
conventions above — `snake_case` — because it's an identifier, not a filename.)
|
||||
|
||||
**A sub-project's entry point repeats its directory's name** — `init/init.zig`,
|
||||
`runtime/runtime.zig`, `hpet/hpet.zig` — and the sub-project is addressed by the
|
||||
`runtime/runtime.zig`, `ps2-bus/ps2-bus.zig` — and the sub-project is addressed by the
|
||||
*directory* (`system/services/init`, `library/runtime`), with the repeated leaf
|
||||
resolving away. See the repository-layout section of [README.md](README.md).
|
||||
|
||||
|
||||
@@ -17,7 +17,7 @@ Most modern Unix and Unix-like operating systems follow the FHS. DanOS has its o
|
||||
| /srv | Site-specific data served by this system, such as data and scripts for web servers, data offered by FTP servers, and repositories for version control systems |
|
||||
| /system | DanOS operating system files (similar idea to C:\Windows). A true representation of danos — its layout mirrors the source tree, so `/system` is what danos *is*. |
|
||||
| /system/devices | danos virtual device tree e.g. similar to /sys on linux but with danos device tree conventions (the structures in the devices module) |
|
||||
| /system/drivers | driver binaries, one sub-project each (e.g. /system/drivers/hpet) |
|
||||
| /system/drivers | driver binaries, one sub-project each (e.g. /system/drivers/pci-bus, /system/drivers/ps2-bus) |
|
||||
| /system/services | system-service binaries — the VFS server, init, and other user-mode servers (e.g. /system/services/vfs, /system/services/init) |
|
||||
| /system/kernel | the kernel image |
|
||||
| /tmp | Directory for temporary files (see also /var/tmp). Often not preserved between system reboots and may be severely size-restricted. |
|
||||
@@ -61,8 +61,8 @@ to the driver in the order written, and a read consumes what is there. Terminals
|
||||
serial lines, keyboards and mice are all of this shape. These are the natural first
|
||||
device nodes in danos, because a character driver needs nothing the kernel doesn't
|
||||
already provide — it claims its device, maps its registers with `mmio_map`, and blocks
|
||||
on `replyWait` for either an interrupt or a client request. `system/drivers/hpet/hpet.zig` is already
|
||||
that program, minus the client half.
|
||||
on `replyWait` for either an interrupt or a client request. `system/drivers/ps2-bus/ps2-bus.zig`
|
||||
is already that program, minus the file-node client half.
|
||||
|
||||
The obstacle was never the file type; it is which hardware a ring-3 driver can reach.
|
||||
Direct `in`/`out` from user space is still a #GP (no TSS I/O bitmap, IOPL never raised),
|
||||
|
||||
@@ -78,6 +78,40 @@ preemption and wakeups (1 ms granularity); the **TSC** is the resolution you rea
|
||||
time at. Making `sleep` itself sub-millisecond would take a tickless one-shot
|
||||
timer — a later step.
|
||||
|
||||
### Is the TSC trustworthy? Invariant, and synchronized
|
||||
|
||||
A cycle counter is only a valid *clock* if two things hold, and danos checks both,
|
||||
because they decide whether we read time with a cheap `rdtsc` or fall back to the HPET.
|
||||
|
||||
**Invariant.** An old TSC counted core clock cycles, so it sped up and slowed down with
|
||||
frequency scaling — useless as wall time. Modern CPUs (all of danos's targets) provide an
|
||||
**invariant TSC**: a constant rate across P/C-states that never stops. The guarantee is a
|
||||
CPUID bit — leaf `0x80000007`, EDX bit 8 — on both Intel *and* AMD. danos reads it in
|
||||
`calibrate`, and a TSC that doesn't advertise it is not used as the clocksource. AMD is
|
||||
why this matters in practice: it doesn't populate the Intel leaf `0x15` that enumerates
|
||||
the TSC *frequency*, so danos already measures AMD's rate against the HPET — but a
|
||||
measured frequency without the invariance guarantee is not enough.
|
||||
|
||||
**Synchronized.** Each core has its own TSC. Even invariant ones can start at different
|
||||
values (a second socket, some firmware), so a thread migrating from a core reading
|
||||
`1_000_000` to one reading `999_000` would see time jump *backward*. danos runs a **warp
|
||||
check** as each application processor comes online (`checkWarpSource`, adapted from
|
||||
Linux's): the waking core and the BSP hammer a shared "highest seen" TSC under a lock,
|
||||
and if either ever reads below it, the cores' TSCs are skewed. It's pairwise because APs
|
||||
come up one at a time ([smp.md](smp.md)).
|
||||
|
||||
**The fallback.** When the TSC fails either test — non-invariant (a bare VM such as the
|
||||
default qemu64), or warped between cores — danos moves the monotonic clock onto the
|
||||
**HPET** main counter: one fixed-rate counter, so it can neither skew between cores nor
|
||||
drift with frequency. It costs a memory-mapped read instead of a register read, but it
|
||||
keeps time *accurate*, which is the whole point. The switch preserves the current value,
|
||||
so the clock never jumps. The boot log names the outcome:
|
||||
|
||||
```
|
||||
/system/kernel: clocksource tsc (TSC invariant: yes, synchronized: yes) # real Intel/AMD
|
||||
/system/kernel: clocksource hpet (TSC invariant: no, synchronized: yes) # a bare VM (TCG)
|
||||
```
|
||||
|
||||
## Two kinds of vector, one dispatch
|
||||
|
||||
The IDT now installs gates `0-47`: the 32 exceptions plus the device range. Every
|
||||
|
||||
+21
-5
@@ -51,7 +51,16 @@ enumeration is a **pci-bus driver**: the manager spawns it against the host brid
|
||||
like any bus reports children. ACPI becomes an **acpi service** that interprets the
|
||||
tables and reports the namespace. The manager only orchestrates and merges. Moving
|
||||
AML interpretation out of ring 0 is its own project on its own track; nothing here
|
||||
depends on when it lands.
|
||||
depends on when it lands. (It landed: [discovery.md](discovery.md), M19–M20.)
|
||||
|
||||
`device_register` is **idempotent on exact match**: a re-registration with an
|
||||
identical (parent, class, identity, resources) tuple returns the existing id
|
||||
instead of appending a duplicate. The kernel table has no unregister, so without
|
||||
this a restarted registering bus would re-report its children as fresh nodes on
|
||||
every respawn. Idempotence is what makes restart-and-re-report sound for *every*
|
||||
reporting bus — pci-bus, the acpi service, a future fdt service — not just one,
|
||||
and it is why supervision (below) can prune a dead bus's subtree and trust the
|
||||
restarted instance to rebuild exactly the same ids.
|
||||
|
||||
## The protocol
|
||||
|
||||
@@ -136,10 +145,17 @@ published exit events, signals + `runtime.process`). On top of those:
|
||||
the mouse and keyboard QEMU already hangs off it.
|
||||
7. **App surface**: `enumerate`/`subscribe` over IPC; `device_enumerate` retreats
|
||||
to a manager-internal seam.
|
||||
8. **Discovery migration** — DONE (M19–M20, 2026-07-13): pci-bus driver (M19)
|
||||
then the acpi service (M20) moved enumeration to ring 3; the kernel seeds
|
||||
only the host bridge and the acpi-tables node. See
|
||||
[m19-m20-plan.md](m19-m20-plan.md).
|
||||
8. **Discovery migration** — DONE (M19–M20, 2026-07-13): enumeration moved to
|
||||
ring 3 as swappable per-firmware discoverers — the pci-bus driver (M19) then
|
||||
the acpi service (M20), see [discovery.md](discovery.md); the kernel seeds
|
||||
only the host bridge and the acpi-tables node. Matching moved with it:
|
||||
`child_added` grew a `device_id` (the kernel-registered id, `no_device` for
|
||||
unregistered leaves like USB ports) and a firmware `hid`, and the manager now
|
||||
matches drivers from those **reports** rather than its boot-time snapshot. The
|
||||
PCI arm flipped in M19.3, the ACPI arm (ps2-bus matched from `_HID`) in M20.3
|
||||
— each in a single phase so no device is ever matched from both sources at
|
||||
once. The acpi service reports only the non-PCI `_HID` devices, since pci-bus
|
||||
already reports PCI functions (M20.2).
|
||||
|
||||
## Settled questions (2026-07-12)
|
||||
|
||||
|
||||
@@ -191,3 +191,57 @@ and registers + reports each `_HID` device — the device manager matches driver
|
||||
(ps2-bus) from those reports. With M19's pci-bus driver, discovery now runs
|
||||
entirely in user space; the kernel seeds only the host bridge and the
|
||||
acpi-tables node.
|
||||
|
||||
## Discovery is a swappable process per firmware (M19–M20)
|
||||
|
||||
Moving PCI and ACPI enumeration out of ring 0 was not just a relocation — it
|
||||
made discovery **firmware-neutral by construction**, which is the whole reason
|
||||
to do it before the second architecture rather than after. Everything at and
|
||||
above the [device-manager](device-manager.md) protocol — descriptors,
|
||||
containment, reports, matching, supervision — is generic and may never become
|
||||
x86-specific. Discovery is the single firmware-specific piece, and it is
|
||||
isolated as **one swappable process per firmware**:
|
||||
|
||||
- **x86** boots describe hardware with ACPI, so the discoverer is the **acpi
|
||||
service** ([acpi.md](acpi.md)): it claims the `acpi-tables` node and runs AML.
|
||||
- **The Raspberry Pis** hand over a flattened device tree, so the discoverer is
|
||||
an **fdt service**: it claims a `devicetree-blob` node and walks the tree —
|
||||
pure data, no bytecode, so it needs neither a port grant nor an interpreter,
|
||||
strictly simpler than ACPI. (A placeholder until the [aarch64](arm.md)
|
||||
bring-up fills it in.)
|
||||
|
||||
The device manager spawns the discoverer under the **neutral ramdisk name
|
||||
`discovery`** and never learns which firmware it is on; the build's
|
||||
`-Ddiscovery=acpi|fdt` option fills that slot (x86 defaults to `acpi`, the
|
||||
aarch64 target flips the default when it lands). The manager owns the device
|
||||
tree as *data* and touches no hardware, ever — firmware bytecode runs only
|
||||
inside the crashable, supervised discoverer, so an AML fault can never take
|
||||
down the supervisor.
|
||||
|
||||
Two consequences of neutrality bind on later work:
|
||||
|
||||
- **Cross-firmware surfaces are named by domain, not firmware.** System power is
|
||||
a [`power`](power.md) protocol, not an "ACPI events" protocol: on x86 the acpi
|
||||
service registers it, on ARM a PSCI/mailbox service registers the same
|
||||
`ServiceId.power`, and subscribers never learn the difference.
|
||||
- **Identity must widen before the fdt service exists.** `DeviceDescriptor`'s
|
||||
8-byte `hid` holds an EISA id but cannot hold an FDT `compatible` string
|
||||
(`"brcm,bcm2835-aux-uart"`); the identity field grows before the ARM path can
|
||||
report a real node.
|
||||
|
||||
Two supporting decisions keep the kernel's remaining slice honest:
|
||||
|
||||
- **The AML interpreter is a shared build module**, compiled into both the
|
||||
kernel and the acpi service — one source, two builds, no fork. The kernel
|
||||
links it for the `\_S5` poweroff evaluation, the service links it for
|
||||
everything else, and the `acpi-parse` test asserts the two produce the same
|
||||
device count across the ring-3 move.
|
||||
- **Bridge apertures come from the firmware memory map, not AML.** Registered
|
||||
PCI functions carry BAR resources, and `device_register` containment demands
|
||||
the bridge own windows that cover them. Those apertures are derived
|
||||
kernel-side from the boot memory map's MMIO holes (regions that are neither
|
||||
RAM nor tables) — mechanical, AML-free, and available at boot regardless of
|
||||
what later moved to user space. The acpi service's authority is likewise
|
||||
exactly one node: the `acpi-tables` node, whose broad io_port grant is the
|
||||
documented trust boundary for the one process allowed to run firmware
|
||||
bytecode.
|
||||
|
||||
+10
-9
@@ -57,8 +57,9 @@ is not an address window. Discovery is trusted; user space is not.
|
||||
|
||||
### What a bus driver looks like
|
||||
|
||||
`system/drivers/bus/bus.zig` is the smallest honest one. Its "bus" is the HPET's register block and
|
||||
its "devices" are the block's comparators:
|
||||
danos ships no demo bus driver — the real ones are `pci-bus`, `ps2-bus`, and
|
||||
`usb-xhci-bus`. The smallest *honest* shape, illustrated here with an HPET register block
|
||||
as the "bus" and its comparators as the "devices", is:
|
||||
|
||||
```zig
|
||||
_ = dev.claim(bus.id); // 1. own the bus
|
||||
@@ -78,8 +79,8 @@ for (0..n) |i| { // 3. publish each child
|
||||
|
||||
Each child is left **unclaimed**, which is the handoff: a comparator driver can now
|
||||
`device_claim` one and `mmio_map` it, and will see only its own 0x20-byte window. A child
|
||||
whose window escapes the bus is refused — `bus` asserts that, and the `bus` test
|
||||
asserts the kernel's table upholds it.
|
||||
whose window escapes the bus is refused; the in-kernel `containment` test asserts the
|
||||
kernel's table upholds that ([drivers.md](drivers.md)).
|
||||
|
||||
A USB device has *no* resources at all: `resource_count = 0`, because it's addressed
|
||||
through its controller, not by MMIO. That case is allowed and is the common one.
|
||||
@@ -143,7 +144,7 @@ If a class driver needs `mmio`, it has become an HCD and should be one.
|
||||
physically-contiguous, pinned, uncacheable, reclaim-on-teardown buffers with the
|
||||
physical address exposed (`pmm.allocContiguous`, a DMA arena, `mapUserDmaInto`).
|
||||
`dma_below_4g` caps the address for legacy engines; `dma_write_combining` is accepted
|
||||
but falls back to coherent until PAT is programmed. hpet is refactored onto `/lib/mmio`;
|
||||
but falls back to coherent until PAT is programmed. The bus drivers use `/lib/mmio`;
|
||||
no DMA driver consumes `dma_alloc` yet.
|
||||
- **M15** — interrupts for PCI devices, the MSI half. Discovery now gives every PCI
|
||||
function its 4 KiB ECAM config space as resource 0 (unblocking the capability walk
|
||||
@@ -302,8 +303,8 @@ rather than an out-struct. The rest of this section is the original design note.
|
||||
|
||||
**The blocker, and it's a hard one.** No PCI device can take an interrupt today.
|
||||
[`addBars`](system/devices/acpi.zig) records `.memory` and `.io_port` BARs and never an
|
||||
`.irq`; there is no `_PRT` parsing anywhere in the tree. `hpet` only works because the
|
||||
HPET advertises its own routing options in its own registers — a privilege no ordinary
|
||||
`.irq`; there is no `_PRT` parsing anywhere in the tree. The HPET is the one exception —
|
||||
it advertises its own interrupt routing in its own registers, a privilege no ordinary
|
||||
device has.
|
||||
|
||||
**The fix, in two halves.**
|
||||
@@ -326,7 +327,7 @@ which means **discovery should give each `pci_device` a `.memory` resource for i
|
||||
4 KiB ECAM slot**. That's a small change to `parseMcfg` and it unblocks the whole
|
||||
capability walk (MSI, MSI-X, PCIe extended caps) without any new syscall.
|
||||
|
||||
Note QEMU's HPET reports `Tn_FSB_INT_DEL_CAP = 0` — no MSI — so `hpet` can never
|
||||
Note QEMU's HPET reports `Tn_FSB_INT_DEL_CAP = 0` — no MSI — so an HPET timer could never
|
||||
exercise this path. The first MSI driver will be the first PCI driver.
|
||||
|
||||
## M16 — the IOMMU, and the honest caveat ◑ detection done, enforcement pending
|
||||
@@ -353,7 +354,7 @@ gap should be named rather than implied.
|
||||
|
||||
`M13` (capability passing) is independent of `M14`/`M15` and is the cheapest. It
|
||||
unlocks class drivers, which are the shape with no hardware requirements at all — you
|
||||
could write a real one against `bus`'s comparators tomorrow.
|
||||
could write a real one against any device a bus driver publishes tomorrow.
|
||||
|
||||
`M14` and `M15` together unlock the first HCD. `M14`'s barrier layer is worth landing
|
||||
on its own regardless: it's small, obviously correct, and stops every future driver
|
||||
|
||||
+40
-29
@@ -22,12 +22,12 @@ say.*
|
||||
|
||||
## How a driver gets started: discover, match, spawn
|
||||
|
||||
Nothing in the kernel decides that the HPET needs the `hpet` driver — that is policy,
|
||||
and policy lives in user space. Boot brings user space up as a three-level supervision
|
||||
hierarchy, each level owning one job:
|
||||
Nothing in the kernel decides that the PCI host bridge needs the `pci-bus` driver — that
|
||||
is policy, and policy lives in user space. Boot brings user space up as a three-level
|
||||
supervision hierarchy, each level owning one job:
|
||||
|
||||
```
|
||||
kernel ──spawns──► init (PID 1) ──spawns──► device-manager ──spawns──► hpet
|
||||
kernel ──spawns──► init (PID 1) ──spawns──► device-manager ──spawns──► pci-bus
|
||||
| | |
|
||||
spawns only init, the service supervisor: the driver supervisor: enumerates
|
||||
publishes the starts the system /system/devices, matches each device
|
||||
@@ -184,7 +184,11 @@ Two properties worth knowing:
|
||||
|
||||
## A whole driver
|
||||
|
||||
`system/drivers/hpet/hpet.zig` is ~150 lines and does all of it. The shape:
|
||||
A minimal leaf driver is only ~150 lines and does all of it. danos ships **no such
|
||||
example binary** — the driver model is proven by the real drivers (`pci-bus`, `ps2-bus`,
|
||||
`usb-xhci-bus`), and a teaching example belongs here, in the docs, rather than as a
|
||||
compiled program nobody runs. Illustrated with a hypothetical HPET timer driver, the
|
||||
shape is:
|
||||
|
||||
```zig
|
||||
const hpet = findHpet(buf) orelse return; // device_enumerate, look for
|
||||
@@ -209,8 +213,8 @@ while (...) {
|
||||
}
|
||||
```
|
||||
|
||||
The HPET is a good first driver for a reason that isn't obvious. Its *counter* is a
|
||||
clocksource — the only way to use it is to read it, so it proved `mmio_map` without
|
||||
The HPET makes a good illustration for a reason that isn't obvious. Its *counter* is a
|
||||
clocksource — the only way to use it is to read it, so it exercises `mmio_map` without
|
||||
needing interrupts at all. Its *comparators* are a clockevent, and can be configured
|
||||
**level-triggered** (`Tn_INT_TYPE_CNF`), which asserts a bit in `GENERAL_INT_STATUS`
|
||||
that the driver must write-1-to-clear. That's a genuine deassert step, so the full
|
||||
@@ -252,9 +256,10 @@ bus driver may only ever subdivide what it already owns.
|
||||
A device with **no resources** is legal and common. A USB device is reached through its
|
||||
controller, not by MMIO, so it gets `resource_count = 0`.
|
||||
|
||||
See [`system/drivers/bus/bus.zig`](../system/drivers/bus/bus.zig) for a complete one, and
|
||||
[driver-model.md](driver-model.md) for how bus drivers, class drivers and host
|
||||
controller drivers fit together.
|
||||
See [`system/drivers/pci-bus/pci-bus.zig`](../system/drivers/pci-bus/pci-bus.zig) for a
|
||||
real one — it claims a PCI host bridge, maps its ECAM window, and publishes each function
|
||||
it finds as a child — and [driver-model.md](driver-model.md) for how bus drivers, class
|
||||
drivers and host controller drivers fit together.
|
||||
|
||||
## What the kernel does not do for you
|
||||
|
||||
@@ -313,32 +318,38 @@ uncacheable, physical address exposed), and **memory barriers** (`/lib/mmio`'s
|
||||
|
||||
## Verifying it
|
||||
|
||||
The `hpet` test spawns `hpet` from the initial ramdisk and watches the serial log. The driver
|
||||
prints `hpet: ok` only after being woken five times, and its loop's only exit is
|
||||
through `replyWait` returning a notification — it cannot reach that line by polling.
|
||||
No demo driver ships to prove this end to end; the *real* drivers do, so the tests
|
||||
target them and the kernel primitives directly:
|
||||
|
||||
The last check doesn't trust the driver's self-report at all: the kernel reads the I/O
|
||||
APIC redirection entry back and asserts the line really is routed to a device vector,
|
||||
really is level-triggered, and really was left unmasked by the driver's final
|
||||
`irq_ack`.
|
||||
- **`device-manager`** — boots only the device manager, which discovers the PCI host
|
||||
bridge, matches `pci-bus`, and `system_spawn`s it. The test reads kernel state — the
|
||||
process table and the device tree — to confirm pci-bus came up and registered the
|
||||
functions it enumerated: the whole discover → match → spawn → driver-up chain.
|
||||
- **`acpi-ps2`** — a user-space driver (`ps2-bus`) is woken by its device's IRQ,
|
||||
delivered as an IPC notification, and attaches the keyboard: IRQ-as-IPC, end to end.
|
||||
- **`pci-scan`** — a user-space driver (`pci-bus`) maps its device's MMIO (the ECAM
|
||||
window) and walks it: `mmio_map`, end to end.
|
||||
- **`containment`** — the kernel refuses a `device_register` whose child window escapes
|
||||
the parent's grant (else it would be a syscall for mapping arbitrary memory), while an
|
||||
identical re-register stays idempotent. Asserted in-kernel, straight against the broker.
|
||||
- **`irqfree`** — the teardown path. Binds two owners to one shared endpoint, releases
|
||||
one, and reads the I/O APIC back: the departing owner's line is masked, the sibling's
|
||||
is not. That second half is why bindings are keyed on the owning *task* and not on the
|
||||
endpoint pointer — endpoints are shared, so releasing "everything pointing at this
|
||||
endpoint" would silently mask a live driver's device.
|
||||
- **`iopass`** — the `device_grant` teardown rule, so destroying a driver's address
|
||||
space never returns MMIO frames to the RAM pool.
|
||||
|
||||
```
|
||||
$ python3 test/qemu_test.py hpet irqfree iopass
|
||||
hpet ... PASS (matched 'DANOS-TEST-RESULT: PASS')
|
||||
$ python3 test/qemu_test.py device-manager acpi-ps2 pci-scan containment irqfree iopass
|
||||
device-manager ... PASS (matched 'DANOS-TEST-RESULT: PASS')
|
||||
acpi-ps2 ... PASS
|
||||
pci-scan ... PASS (matched 'DANOS-TEST-RESULT: PASS')
|
||||
containment ... PASS (matched 'DANOS-TEST-RESULT: PASS')
|
||||
irqfree ... PASS (matched 'DANOS-TEST-RESULT: PASS')
|
||||
iopass ... PASS (matched 'DANOS-TEST-RESULT: PASS')
|
||||
```
|
||||
|
||||
Two companions cover what `hpet` can't, because it never exits:
|
||||
|
||||
- **`irqfree`** — the teardown path. Binds two owners to one shared endpoint, releases
|
||||
one, and reads the I/O APIC back: the departing owner's line is masked, the sibling's
|
||||
is not. That second half is why bindings are keyed on the owning *task* and not on
|
||||
the endpoint pointer — endpoints are shared, so releasing "everything pointing at
|
||||
this endpoint" would silently mask a live driver's device.
|
||||
- **`iopass`** — the `device_grant` teardown rule, so destroying a driver's address
|
||||
space never returns MMIO frames to the RAM pool.
|
||||
|
||||
## What's next (not done here)
|
||||
|
||||
The big driver-model pieces — capability passing (class drivers), DMA + barriers, MSI,
|
||||
|
||||
@@ -1,210 +0,0 @@
|
||||
# M17–M18 execution plan: process lifecycle + device manager
|
||||
|
||||
**Archived — completed 2026-07-13** (every item checked; suite ended 54/54).
|
||||
Kept as the record of how M17–M18 landed; the successor is
|
||||
[m19-m20-plan.md](m19-m20-plan.md).
|
||||
|
||||
The operational plan for building [process-lifecycle.md](process-lifecycle.md)
|
||||
(M17) and [device-manager.md](device-manager.md) increments 5–7 (M18). Design is
|
||||
settled in those documents; this file is the build order — one phase at a time,
|
||||
each phase green before the next starts. Delete or archive this file when M18
|
||||
lands.
|
||||
|
||||
**Definition of green, every phase:** `zig build` clean, `zig build test` clean,
|
||||
`python3 test/qemu_test.py` passes (existing scenarios plus the phase's new one),
|
||||
and the relevant design doc's "known gaps" / status lines updated. Commit per
|
||||
green phase (no co-author trailers).
|
||||
|
||||
**Workflow (settled 2026-07-12):** work happens in a dedicated git worktree, on
|
||||
feature branches cut from `main` — `feat/process-lifecycle` (M17.1–17.4),
|
||||
`feat/device-manager` (M18.1), `feat/usb-xhci-bus` (M18.2–18.3). When a branch's
|
||||
phases are all green it is **auto-merged into `main`**; branches are kept after
|
||||
merge, not deleted. Merges and branches are pushed to origin. Phase 0 (once):
|
||||
commit the design docs, merge the outstanding `feat/usb` work into `main`, and
|
||||
run the existing QEMU suite green before any new work starts.
|
||||
|
||||
**Numbering note:** continues the milestone sequence (driver track ended at M16).
|
||||
|
||||
## Status
|
||||
|
||||
The loop marks a phase `[x]` in the same commit that lands it. A phase is marked
|
||||
only when its definition of green holds.
|
||||
|
||||
- [x] **Phase 0** — baseline: docs committed, feat/usb merged to main, pushed;
|
||||
`usb-xhci-libary.zig` renamed to `usb-xhci-library.zig`; existing QEMU
|
||||
suite green from the worktree (48/48, 2026-07-12).
|
||||
- [x] **M17.1** — kernel releases claims/MSI on death (claims: `releaseAllOwnedBy`
|
||||
in the reap; MSI was already swept by `irq.releaseOwner`; `claim-release`
|
||||
test; suite 49/49)
|
||||
- [x] **M17.2** — exit reasons (`ExitReason` recorded at exit/fault/kill before
|
||||
the notification; `process_exit_reason` supervisor-gated;
|
||||
`runtime.process.exitReason`; kernel + ring-3 assertions; suite 49/49)
|
||||
- [x] **M17.3** — published exit events + VFS subscriber (`process_subscribe`,
|
||||
bounded ref-counted table, publish on every death;
|
||||
`runtime.process.subscribeExits`; VFS handles carry owners and are swept on
|
||||
the owner's death; `vfs-client-death` test; suite 50/50)
|
||||
- [x] **M17.4** — signals, timer notifications, `runtime.process`, the service
|
||||
harness (signal_bind/process_signal + coalescing pending mask; timer_bind
|
||||
on the tick; bindSignals/signalsFrom/sendSignal/stop + timerOnce;
|
||||
runtime.service.run with the zero-length ping; VFS converted; `signals`
|
||||
scenario; suite 51/51)
|
||||
- [x] **merge** `feat/process-lifecycle` → main, push (merged 2026-07-13)
|
||||
- [x] **M18.1** — device-manager protocol: hello + restart policy
|
||||
(device-manager-protocol module; the manager as a harness service:
|
||||
supervised spawns, hello deadline via timer sweep, restart with
|
||||
300/600/1200ms backoff, exit reasons deciding restart-vs-stopped,
|
||||
crash-loop cap; usb-xhci-bus first conforming driver; crash-test fixture
|
||||
re-proving claim release each respawn; `driver-restart` scenario;
|
||||
maximum_tasks 16→32 — the sweep was overflowing the pool; suite 52/52)
|
||||
- [x] **merge** `feat/device-manager` → main, push (merged 2026-07-13)
|
||||
- [x] **M18.2** — xHCI port scan + tree reports (child_added/child_removed in
|
||||
the protocol; the manager's child mirror with death-pruning; xHCI maps the
|
||||
register BAR — resource 0 is ECAM — reads CAPLENGTH/HCSPARAMS1, scans
|
||||
PORTSC, reports connected ports with speed-class identity; `usb-report`
|
||||
scenario proves report → prune → respawn → re-report; suite 53/53)
|
||||
- [x] **M18.3** — app surface: enumerate/subscribe over IPC (subscriber
|
||||
endpoint rides as the call's capability; events are the same structs the
|
||||
buses send); device-list first client; protocol capped at the kernel's
|
||||
IPC MESSAGE_MAXIMUM (256); the startUserTask debug print removed — it
|
||||
sheared concurrent serial lines and was the scenario-flake root cause;
|
||||
`device-list` scenario; suite 54/54)
|
||||
- [x] **merge** `feat/usb-xhci-bus` → main, push (merged 2026-07-13) — **plan complete**
|
||||
|
||||
---
|
||||
|
||||
## M17.1 — the kernel releases a dead process's claims
|
||||
|
||||
The cleanup half of iron rule 1; the prerequisite for every restart story.
|
||||
|
||||
- `system/kernel/devices-broker.zig`: `releaseAllOwnedBy(owner: u32)` — clear
|
||||
every `claimed[]` slot holding this task id.
|
||||
- `system/kernel/process.zig`: call it from the reap path, alongside the existing
|
||||
IRQ-binding release (the ordering comment there says why IRQs go first — claims
|
||||
slot in after them, before the exit notification).
|
||||
- MSI vectors: find where `msi_bind` records per-device vectors (interrupts
|
||||
module) and release those by owner in the same pass.
|
||||
- Docs: remove the claims bullet from process-management.md "Known gaps".
|
||||
|
||||
**Test:** new QEMU scenario `claim-release` — a test child claims an unclaimed
|
||||
device, is killed, is respawned, and claims the same device again successfully;
|
||||
assert both claims in the serial log. Kernel-side unit coverage in
|
||||
`system/kernel/tests.zig` for `releaseAllOwnedBy` (claim two devices as two owners,
|
||||
release one owner, verify exactly its claims freed).
|
||||
|
||||
## M17.2 — exit reasons
|
||||
|
||||
- `system/abi.zig`: `ExitReason` (exited, aborted, segmentation_fault,
|
||||
illegal_instruction, arithmetic_fault, killed).
|
||||
- Kernel: record the reason at every death site — clean exit path, each fault
|
||||
class in `onException`, the kill path. Bounded recent-exits table (ids are never
|
||||
reused, so a small ring keyed by id is enough).
|
||||
- New system call `process_exit_reason(id)` — supervisor-gated, like kill; returns
|
||||
the recorded reason or `-ESRCH` once evicted.
|
||||
- `library/runtime/process.zig`: `ExitReason` + `exitReason(id: u32)`.
|
||||
- Docs: remove the no-exit-status bullet from process-management.md.
|
||||
|
||||
**Test:** extend the `supervision` scenario — three children: one exits cleanly,
|
||||
one faults (the fault-recovery pattern), one is killed; the supervisor asserts all
|
||||
three reasons.
|
||||
|
||||
## M17.3 — published exit events
|
||||
|
||||
- Kernel: bounded subscriber table (endpoints); new system call
|
||||
`process_subscribe(endpoint)` (ungated, like `process_enumerate`); every death
|
||||
posts `notify_exit_bit | id` to each subscriber — the same post the supervisor
|
||||
path already uses.
|
||||
- `library/runtime/process.zig`: `subscribeExits(endpoint)`.
|
||||
- VFS becomes the first subscriber: on an exit event, release every handle keyed
|
||||
by that task id (badges already are task ids). Log the release.
|
||||
- Docs: note the convention in ipc.md (exit events reuse the exit-notification
|
||||
badge encoding).
|
||||
|
||||
**Test:** new QEMU scenario `vfs-client-death` — a client opens a file and is
|
||||
killed without closing; assert the VFS logs the handle release and its open-handle
|
||||
count returns to baseline.
|
||||
|
||||
## M17.4 — signals and the service harness
|
||||
|
||||
- Kernel: per-task pending mask + bound endpoint; system calls
|
||||
`signal_bind(endpoint)` and `process_signal(id, signal)` (supervisor-or-self
|
||||
gated); delivery posts `notify_signal_bit | pending mask`, coalescing; pending
|
||||
signals with no bound endpoint pend silently.
|
||||
- `library/runtime/process.zig`: `Signal`, `SignalSet`, `bindSignals`,
|
||||
`signalsFrom`, `sendSignal`, `stop(id, deadline_ms)` (terminate → wait for exit
|
||||
notification → kill). Implement `terminate`, `reload`, `user_1`, `user_2`;
|
||||
`interrupt`/`quit` are enum members with no sender yet; `alarm` stays unbuilt.
|
||||
- Kernel: **one-shot timer notifications** — `timer_bind(endpoint, ms)` posts a
|
||||
notification badge when the deadline lands (IRQ-as-IPC again, on the timer
|
||||
wheel `sleep` already uses). This is the missing timed-wait primitive:
|
||||
`replyWait` blocks forever and `sleep` blocks the whole process, but `stop()`'s
|
||||
escalation, the device manager's `hello` deadline (M18.1), and restart backoff
|
||||
all need a deadline while staying responsive. It is also the mechanism `alarm`
|
||||
gets for free later.
|
||||
- New `library/runtime/service.zig`: the harness — `run(callbacks)` owning the
|
||||
replyWait loop, folding protocol messages, signals, and child-exit notifications
|
||||
into `init` / `on_message` / `on_reload` / `on_terminate`; answers the common
|
||||
`ping` automatically. Define the reserved `ping` request encoding here and
|
||||
document it in ipc.md (one obvious encoding; smallest that cannot collide with
|
||||
existing protocols).
|
||||
- Convert one existing service (input-source or hpet) to the harness as proof it
|
||||
subtracts code rather than adding it.
|
||||
|
||||
**Test:** extend `supervision` — a harness-built child: `sendSignal(reload)`
|
||||
observed in its log, `ping` answered, `stop()` produces a clean exit with reason
|
||||
`exited`; a second child that ignores signals (no bind) is killed by `stop()`'s
|
||||
deadline with reason `killed`.
|
||||
|
||||
## M18.1 — device-manager protocol: hello + restart policy
|
||||
|
||||
- New `system/services/device-manager/device-manager-protocol.zig` module
|
||||
(vfs-protocol pattern): `hello { version, role, device_id }`; version constant;
|
||||
reserved fields.
|
||||
- Device manager: register the `.device_manager` endpoint; spawn drivers with its
|
||||
exit endpoint; enforce the hello deadline; restart policy — backoff, crash-loop
|
||||
cap (three fast deaths → mark failed, log, stop), reasons from M17.2 deciding
|
||||
restart vs not.
|
||||
- usb-xhci-bus: adopt the harness + send hello. hpet/ps2-bus follow only if the
|
||||
conversion is mechanical; otherwise they keep working unconverted (the manager
|
||||
only enforces hello on drivers spawned with an assignment).
|
||||
- build.zig: test-loop entry for the protocol module if it grows pure logic.
|
||||
|
||||
**Test:** new QEMU scenario `driver-restart` — the xHCI driver takes a test-only
|
||||
argv flag to fault after hello on its first run; assert: fault, exit reason
|
||||
recorded, manager respawns with backoff, second run claims the controller
|
||||
(M17.1) and hellos clean. Assert the crash-loop cap by a driver that always
|
||||
faults (a tiny test driver, not xhci).
|
||||
|
||||
## M18.2 — bus tree reports
|
||||
|
||||
- Protocol: `child_added { parent, identity, resources }` / `child_removed { id }`.
|
||||
- usb-xhci-bus: bring-up to **port scan only** — map the MMIO window (claimed in
|
||||
M16-era work), controller reset/start per xHCI spec, walk the port registers,
|
||||
report one `child_added` per connected port with speed + port number as
|
||||
identity. **No transfer rings, no descriptors** — reading device/interface
|
||||
descriptors (and therefore USB class triples for matching) is the follow-on USB
|
||||
track, not this plan.
|
||||
- Device manager: mirror reports into its tree; prune the subtree (emitting
|
||||
`child_removed`) when a bus driver dies; assert re-report on restart.
|
||||
|
||||
**Test:** QEMU already attaches usb-kbd + usb-mouse on xhci.0 — assert two
|
||||
`child_added` events reach the manager and appear in its tree dump; kill the
|
||||
driver, assert two `child_removed` then two fresh `child_added` after respawn.
|
||||
|
||||
## M18.3 — the application surface
|
||||
|
||||
- Protocol: `enumerate` (tree snapshot) + `subscribe` (published add/remove
|
||||
events, input-service pattern).
|
||||
- A small client (`device-list`, the `ps` analog) exercising both; the manager
|
||||
becomes the one answer to "what devices exist" for user space.
|
||||
`device_enumerate` stays for drivers/kernel seeding — its retreat is tied to the
|
||||
discovery migration, out of this plan.
|
||||
|
||||
**Test:** QEMU scenario — `device-list` shows the tree including USB children;
|
||||
during a driver restart the subscribing client logs remove + add events.
|
||||
|
||||
---
|
||||
|
||||
**Explicitly out of scope** (own tracks, after M18): discovery migration (pci-bus
|
||||
driver, acpi service, retiring the kernel scan), USB control transfers +
|
||||
descriptors + class-driver matching, the musl layer, `interrupt`/`quit` senders
|
||||
(needs a console), job control.
|
||||
@@ -1,196 +0,0 @@
|
||||
# M19–M20 execution plan: discovery migration
|
||||
|
||||
The operational plan for [device-manager.md](device-manager.md)'s increment 8:
|
||||
discovery leaves the kernel — a **pci-bus driver** (M19) and an **acpi service**
|
||||
(M20), with the kernel's device enumeration retired behind them. Same rules as
|
||||
[m17-m18-plan.md](m17-m18-plan.md): one phase at a time, each green before the
|
||||
next; this file is the build order and the checklist.
|
||||
|
||||
**Definition of green, every phase:** `zig build` clean, `zig build test` clean,
|
||||
`python3 test/qemu_test.py` passes (existing scenarios plus the phase's new
|
||||
one), and the relevant design doc updated. Commit per green phase (no co-author
|
||||
trailers). The full suite is the regression net — the existing
|
||||
`driver-restart` / `usb-report` / `device-list` / `input` scenarios must stay
|
||||
green *through* the migration, which is the whole point: the system must not be
|
||||
able to tell who enumerated it.
|
||||
|
||||
**Workflow:** dedicated worktree; branches off `main` — `feat/pci-bus`
|
||||
(M19.0–19.3), `feat/acpi-service` (M20.1–20.3); auto-merge to main when a
|
||||
branch is green; keep branches; push everything.
|
||||
|
||||
## Settled decisions (2026-07-13 — veto before the loop starts)
|
||||
|
||||
1. **What "retiring the kernel scan" means.** The kernel keeps, forever, the
|
||||
parses it needs before user space exists: RSDP/XSDT location, MADT (SMP),
|
||||
the HPET table (the tick), FADT + the AML `\_S5` evaluation (poweroff — the
|
||||
power tests prove it), and MCFG (the host bridge node). What retires is
|
||||
**device enumeration**: the ECAM function walk (M19.3) and the DSDT/SSDT
|
||||
namespace walk that builds device nodes (M20.3). The AML module stays a
|
||||
shared build module compiled into both the kernel (for `\_S5`) and the acpi
|
||||
service (for everything else) — same source, two builds, no fork.
|
||||
2. **Bridge apertures come from the firmware memory map, not AML.** Registered
|
||||
PCI functions carry BAR resources, and containment demands the bridge own
|
||||
windows that cover them. The apertures are derived kernel-side from the
|
||||
boot memory map's MMIO holes (regions that are neither RAM nor tables) —
|
||||
mechanical, AML-free, and available at boot regardless of what later moved
|
||||
to user space. (The bridge today carries only ECAM + bus range; this is the
|
||||
prerequisite M19.0 exists for.)
|
||||
3. **`device_register` becomes idempotent on exact match.** A re-registration
|
||||
with identical (parent, class, resources) returns the existing id instead
|
||||
of appending. The kernel table has no unregister, so without this a
|
||||
restarted registering bus would duplicate its children on every respawn —
|
||||
idempotence makes restart-and-re-report safe for every future bus, not just
|
||||
PCI.
|
||||
4. **The manager matches from reports.** `ChildAdded` gains a `device_id`
|
||||
field (the kernel-registered id, `no_device` for unregistered leaves like
|
||||
USB ports). After the M19.3 flip, PCI driver matching keys off reported
|
||||
identity (the class triple) instead of the manager's boot-time snapshot —
|
||||
the snapshot match remains only for what the kernel still seeds. One flip
|
||||
phase changes both sides at once so no device is ever matched twice.
|
||||
5. **The acpi service's authority is one node.** The kernel publishes an
|
||||
`acpi-tables` device: memory resources covering the table blobs plus a
|
||||
broad `io_port` resource — the documented trust grant to exactly one
|
||||
process (AML OperationRegions reach EC/PM ports; the claim-gated
|
||||
io_read/io_write calls already exist). The service claims it, maps the
|
||||
tables, and runs the shared AML module in ring 3 behind a `Hal` backed by
|
||||
`mmio_map` + `io_read`/`io_write`.
|
||||
6. **Both new processes are protocol drivers** under the manager: hello,
|
||||
supervision, restart with backoff — all inherited from M18.1 for free.
|
||||
Registration idempotence (decision 3) is what makes their restarts sound.
|
||||
7. **Firmware neutrality is the contract** (2026-07-13). The generic layer is
|
||||
everything at and above the device-manager protocol — descriptors,
|
||||
containment, reports, matching, supervision — and none of it may become
|
||||
x86-specific. Discovery is one swappable process per firmware: the acpi
|
||||
service on x86; an **fdt service** on the Raspberry Pis (claims a
|
||||
`devicetree-blob` node, reports children from the flattened device tree —
|
||||
pure data, no bytecode, no port grant, strictly simpler than ACPI). The
|
||||
manager owns the tree as *data* and touches no hardware, ever — AML runs in
|
||||
a crashable, supervised discoverer precisely so a firmware-bytecode fault
|
||||
can never take down the supervisor. Two consequences recorded now:
|
||||
`DeviceDescriptor`'s 8-byte `hid` cannot hold an FDT `compatible` string
|
||||
("brcm,bcm2835-aux-uart") — identity widens before the fdt service exists;
|
||||
and cross-firmware surfaces are named by **domain, not firmware** (M21
|
||||
defines a *power* protocol, not an "ACPI events" protocol — PSCI/mailbox
|
||||
sources feed the same subscribers on ARM). **Landed early (2026-07-13):**
|
||||
both services exist as placeholders (system/services/acpi, system/services/
|
||||
fdt) and the build's `-Ddiscovery=acpi|fdt` option fills the ramdisk's
|
||||
neutral `discovery` slot — the manager will spawn "discovery" by that name
|
||||
in M20.3 and never learn which firmware it is on.
|
||||
|
||||
## Status
|
||||
|
||||
- [x] **M19.0** — prerequisites (bridge apertures from the memory map's
|
||||
*gaps* — the single-hole rule died on OVMF's flash at the top of 4 GiB,
|
||||
caught by the new every-BAR-contained assert in `discovery`; idempotent
|
||||
`device_register` proven in `bus`; `ChildAdded.device_id`;
|
||||
m17-m18-plan.md archived; suite 54/54).
|
||||
- [x] **M19.1** — pci-bus driver, scan only (claims the bridge, maps ECAM
|
||||
through its grant, brute-force walk with the multifunction rule; the
|
||||
manager matches pci_host_bridge → pci-bus per device with the full
|
||||
protocol contract; `pci-scan` builds its expected marker from the
|
||||
kernel's own count — equivalence on the first run; suite 55/55).
|
||||
- [x] **M19.2** — register + report (BAR probe mirrored byte-for-byte from the
|
||||
kernel's addBars so dedupe returns the kernel's node ids during
|
||||
coexistence; the bridge gained the io_port aperture I/O BARs need;
|
||||
reports carry the registered device_id; pci-scan drills a forced restart
|
||||
and asserts the PCI node count never grows — plus harness hardening: a
|
||||
failing case now preserves its serial as <case>-failed-serial.log, and
|
||||
the heavy scenarios run at 150s; suite 55/55).
|
||||
- [x] **M19.3** — the flip: kernel `enumeratePci`/`addBars`/`PciHeader` all
|
||||
deleted (bridge node stays); manager matches PCI drivers from reported
|
||||
identity, deduped by registered id. Surfaced and fixed a real SMP race the
|
||||
flip created — ring-3 device_register made the broker table concurrent, so
|
||||
mmio_map's lock-free read intermittently tore hpet's resource length
|
||||
(user fault) and overflowed `r.len-1` into a kernel panic; now the broker
|
||||
read is under the big lock and the arithmetic is guarded, and pci-bus
|
||||
skips size-0 BARs. discovery.md updated; suite 55/55 (driver-restart
|
||||
hammered 6×).
|
||||
- [x] **merge** `feat/pci-bus` → main, push (merged 2026-07-13).
|
||||
- [x] **M20.1** — acpi service, parse only: the AML interpreter is now a build
|
||||
module compiled into both kernel and service; the kernel publishes the
|
||||
`acpi-tables` node (AML blobs as memory resources, the broad io_port grant,
|
||||
the SCI); the service claims it, maps the blobs, runs the shared parser in
|
||||
ring 3, and self-verifies its Device count against the kernel's (34 = 34,
|
||||
deterministic via argv, no log-scraping); the manager spawns `discovery`
|
||||
at startup. Parse-only touches no hardware. Suite 56/56.
|
||||
|
||||
- [x] **M20.2** — register + report: the service evaluates `_STA`/`_CRS` in
|
||||
ring 3 (interpreter Hal = port I/O over the claimed node; a scratch page
|
||||
backs SystemMemory maps so a stray region can't fault it) and registers +
|
||||
reports each present `_HID` device under `acpi-tables`. Containment: the
|
||||
broker's irq check became range-based (len-1 == the old equality) so the
|
||||
node's broad irq window covers children's legacy lines; io ports fall in
|
||||
the broad io grant. ChildAdded gained `hid`. Matching stays off. The
|
||||
`acpi-report` scenario asserts the PS/2 keyboard (3 resources) and mouse
|
||||
(1 resource) among the reports. Suite 57/57.
|
||||
- [x] **M20.3** — the flip: the kernel's `wireAcpiDevices` call is gone (the
|
||||
device-building helpers are retained-but-dead pending a focused sweep,
|
||||
spawned as a task; static tables + `\_S5` + the acpi-tables node stay).
|
||||
The manager matches ps2-bus from ACPI `_HID` reports; the service
|
||||
registers all devices before reporting any (no keyboard-before-mouse
|
||||
race). The `acpi-ps2` scenario proves report → spawn → ps2-bus attaches
|
||||
its keyboard; `ioport` retargeted to the acpi-tables I/O window (the
|
||||
kernel-built PS/2 node is gone). Suite 58/58.
|
||||
- [x] **merge** `feat/acpi-service` → main, push (merged 2026-07-13) — **discovery migration complete**.
|
||||
|
||||
---
|
||||
|
||||
## Phase notes
|
||||
|
||||
**M19.0 apertures:** the boot memory map already crosses the handoff
|
||||
([boot-handoff]), but discovery never sees it today — expect a small
|
||||
pass-through (kernel init hands the map to the platform layer) before the
|
||||
holes computation, which belongs where the bridge node is built
|
||||
(`parseMcfg`). Sanity-check on QEMU q35: the xHCI BAR (`0xc0000000`-region
|
||||
values seen in the M18 logs) must land inside a derived aperture, asserted in
|
||||
the kernel unit test.
|
||||
|
||||
**M19.1 scanning without owning config access twice:** the driver reads config
|
||||
space through its ECAM mmio_map grant of the *bridge* window — the same bytes
|
||||
the kernel walk read. Vendor-id `0xFFFF` skip, header-type multifunction rule,
|
||||
no bridge recursion (matches the kernel's current single-segment walk).
|
||||
|
||||
**M19.2 BAR sizing:** the classic size probe (write all-ones, read mask,
|
||||
restore) is deferred — the BARs' current programmed values and types are
|
||||
enough for containment-checked registration at bring-up; sizing lands with the
|
||||
first driver that needs to *move* a BAR. Log what is registered so the
|
||||
scenario can assert it.
|
||||
|
||||
**M19.3 what the manager still seeds from the snapshot:** everything the
|
||||
kernel still enumerates (timers, ACPI nodes until M20.3). The PCI arm of
|
||||
`pciDriverFor` switches source; `driverFor` doesn't move until M20.3.
|
||||
|
||||
**M20.1 spawn and identity (pre-settled 2026-07-13):** the manager spawns
|
||||
`discovery` by its neutral ramdisk name at startup, as an ordinary protocol
|
||||
driver (hello, supervision) — from M20.1 on, on every boot. For reporting ACPI
|
||||
devices, `ChildAdded` gains `hid: [8]u8` (EISA ids fit; zero = none):
|
||||
firmware *string* identity travels beside the numeric `identity` field until
|
||||
the FDT-driven widening replaces both (decision 7).
|
||||
|
||||
**M20.1 Hal in ring 3:** `mapMmio` → `device.mmioMap` over the claimed
|
||||
acpi-tables node (plus a table-offset map for blobs); `pioRead`/`pioWrite` →
|
||||
`device.ioRead`/`ioWrite` against its io_port resource. The interpreter cannot
|
||||
tell it moved — that is the assertion of `acpi-parse`.
|
||||
|
||||
**M20.2 containment for `_CRS`:** io ports fall inside the node's broad
|
||||
io_port resource; MMIO windows (HPET, LAPIC ranges some firmwares list) fall
|
||||
inside the memory-map holes added to the node in M20.1. Anything that doesn't
|
||||
fit is logged and skipped, loudly — bring-up honesty over silent drops.
|
||||
|
||||
**M20.3 ps2 ordering:** ps2-bus binds nodes the acpi service now reports, so
|
||||
its spawn moves behind the report (the manager's matching handles this once
|
||||
the source flips); the `input` scenario proves the keyboard still types.
|
||||
|
||||
**Explicitly out of scope:** PCI bridge recursion (single segment, flat bus
|
||||
walk stays); BAR reprogramming/sizing; disk/PCIe hotplug; interrupt routing
|
||||
changes (`_PRT` stays wherever it is today); the USB descriptor track;
|
||||
multi-segment ECAM; per-device power states (D-states, `_PSx`/`_PRx`,
|
||||
suspend/resume — a future *lifecycle-vocabulary* extension, since "suspend"
|
||||
has the shape of a signal every driver must answer, and it has no consumer
|
||||
until laptop sleep); CPU P/C-states.
|
||||
|
||||
## M21 — ACPI events + system power — DONE
|
||||
|
||||
Built and merged (docs/m21-plan.md, 2026-07-13): the SCI + power button, Notify/GPE
|
||||
dispatch, and orderly shutdown (init's stop cascade into a ring-3 S5 write).
|
||||
See that plan for the phase record.
|
||||
@@ -1,138 +0,0 @@
|
||||
# M21 execution plan: ACPI events + system power
|
||||
|
||||
The operational plan for the event side of the acpi service and orderly
|
||||
shutdown — the capstone [m19-m20-plan.md](m19-m20-plan.md) previewed. Same
|
||||
rules as its predecessors: one phase at a time, each green before the next;
|
||||
this file is the build order and the checklist.
|
||||
|
||||
**Definition of green, every phase:** `zig build` clean, `zig build test`
|
||||
clean, `python3 test/qemu_test.py` passes (existing scenarios plus the
|
||||
phase's new one), and the relevant design doc updated. Commit per green phase
|
||||
(no co-author trailers). Failing cases preserve their serial logs
|
||||
(`<case>-failed-serial.log`).
|
||||
|
||||
**Workflow:** dedicated worktree; branch `feat/power-events` off `main`;
|
||||
auto-merge to main when the branch is green; keep the branch; push everything.
|
||||
|
||||
## Settled decisions (2026-07-13, approved)
|
||||
|
||||
1. **S5 is executed by the acpi service from ring 3.** No new syscall: the
|
||||
broad port grant (M20 decision 5) already made this physically possible —
|
||||
the service holds the PM1 control ports in its io grant and derives `_S5`
|
||||
from its own namespace (`aml.sleepState`). Formalizing it adds no
|
||||
authority. The kernel keeps `power.zig` for its own test paths and
|
||||
panic-time use.
|
||||
2. **The power surface is domain-named** (decision 7 of the last plan): a
|
||||
`power-protocol` module + `ServiceId.power = 5`, registered by the acpi
|
||||
service — on ARM, a PSCI/mailbox service registers the same id and
|
||||
subscribers never know the difference. Messages: `subscribe` (endpoint as
|
||||
the call's capability, the input/manager pattern), `shutdown` (accepted
|
||||
only from PID 1 — init), and events published as buffered messages:
|
||||
`power_button`, `lid`, `ac`, `battery`, generic `notify` with a code.
|
||||
3. **The service learns event ports from its own FADT copy**: the kernel adds
|
||||
the FADT as one more memory resource on the acpi-tables node; the service
|
||||
tells it apart from the AML blobs by signature ("FACP" header — the blob
|
||||
resources are header-stripped bytecode and start with no signature). The
|
||||
kernel's own FADT parse is untouched.
|
||||
4. **The acpi service converts to the harness** (`runtime.service.run`):
|
||||
protocol messages (subscribe/shutdown), the SCI notification, and the
|
||||
existing report flow fold into one loop — the shape it was always meant
|
||||
to have.
|
||||
5. **GPE/Notify correctness is proven by host unit tests** (synthetic AML
|
||||
with a Notify inside a method body; aml.zig joins the `zig build test`
|
||||
loop). The QEMU scenario proves the power button — a *fixed* event,
|
||||
deterministically injectable via QMP `system_powerdown` — because QEMU
|
||||
cannot raise GPEs deterministically on this config. Battery/AC/lid and the
|
||||
embedded controller (`_Qxx`) are interface-complete here and validated on
|
||||
real hardware (the laptop) later.
|
||||
|
||||
## Ground truth the phases build on (verified 2026-07-13)
|
||||
|
||||
- `system/devices/power.zig` `shutdown()` is the kernel's S5 write
|
||||
(SLP_TYP|SLP_EN to PM1a/PM1b control); there is no power syscall.
|
||||
- init (`system/services/init/init.zig`) spawns vfs/input/device-manager
|
||||
fire-and-forget — no child ids kept, no signals, no event loop. The whole
|
||||
stop toolkit exists in `runtime.process` (stop/sendSignal/bindSignals).
|
||||
- `test/qemu_test.py` has no QMP channel (serial is a one-way file).
|
||||
- The kernel parses PM1 *control* blocks and SCI_INT from the FADT; the PM1
|
||||
**event** blocks (offsets 56/60, len at 88) and **GPE0/GPE1** blocks
|
||||
(offsets 80/84, lens 92/93) are unparsed — the service reads them from its
|
||||
FADT copy (decision 3).
|
||||
- The acpi-tables node carries the SCI as its only `len == 1` irq resource
|
||||
(the broad window is len 256) — that is how the service finds it to
|
||||
`irqBind`.
|
||||
- `notify_opcode = 0x86` exists in `system/devices/aml/opcodes.zig` but the
|
||||
interpreter never handles it — a GPE `_Lxx` body containing Notify fails
|
||||
evaluation today. Everything else a GPE handler needs (field access,
|
||||
control flow, method calls) is proven by the ring-3 `_STA`/`_CRS` work.
|
||||
- The dead-code sweep (spawned task) also edits `system/devices/acpi.zig`;
|
||||
M21.0 checks whether it landed and rebases before touching that file.
|
||||
|
||||
## Status
|
||||
|
||||
- [x] **M21.0** — baseline (dead-code sweep confirmed landed on main — no
|
||||
acpi.zig conflict; `feat/power-events` cut; QMP channel in the harness:
|
||||
always-on unix socket, client with the capabilities handshake, per-case
|
||||
`qmp_after` hook, and a hook-must-deliver pass gate that the smoke case
|
||||
now proves with a harmless query-status; suite 58/58).
|
||||
- [x] **M21.1** — SCI + the power button (kernel appends the FADT as an
|
||||
acpi-tables memory resource, tagged by its "FACP" header; `power-protocol`
|
||||
module + `ServiceId.power = 5`; the acpi service converted to
|
||||
`runtime.service.run`, registers `.power`, reads PM1 event/control + GPE
|
||||
ports from its FADT copy, enables ACPI mode if SCI_EN is clear, binds the
|
||||
SCI (the len-1 irq), sets PWRBTN_EN; the SCI handler clears PM1_STS,
|
||||
logs `power: button pressed`, publishes `power_button`, acks. Scenario
|
||||
`power-button` injects a real `system_powerdown` via QMP; initial-ramdisk
|
||||
timeout 30→60s for the service's added boot work; suite 59/59).
|
||||
- [x] **M21.2** — Notify + GPE dispatch (interpreter handles `notify_opcode`
|
||||
into a bounded queue, cleared per-evaluate, drained via
|
||||
`takeNotifications`; the service walks GPE status/enable bytes, evaluates
|
||||
`\_GPE._Lxx`/`_Exx` per active bit, maps notified nodes to events
|
||||
(battery/ac/lid/generic), clears GPE_STS write-1, acks. EC `_Qxx` out.
|
||||
Host unit test with hand-encoded AML proves the queue; aml.zig joined the
|
||||
`zig build test` loop. QEMU raises no GPEs — suite is regression net,
|
||||
59/59).
|
||||
- [x] **M21.3** — orderly shutdown (init supervises its children on one
|
||||
endpoint that also carries signals, power events, and a re-arming
|
||||
heartbeat timer; on `power_button` or a `terminate` signal it logs
|
||||
`init: shutting down`, runs `stop(child, 2000, endpoint)` in reverse
|
||||
order, then requests `.power` shutdown; the acpi service honors shutdown
|
||||
from a subscriber — init is the one subscriber, a soft gate that survives
|
||||
testing where PID 1 isn't init — and writes SLP_TYP|SLP_EN from ring 3.
|
||||
`orderly-shutdown` scenario proves button → shutting-down → S5 → QEMU
|
||||
exit; suite 60/60).
|
||||
- [x] **merge** `feat/power-events` → main, push, keep the branch (merged 2026-07-13) — **M21 complete**.
|
||||
|
||||
---
|
||||
|
||||
## Phase notes
|
||||
|
||||
**M21.0 QMP:** open the unix socket after Popen, complete the
|
||||
`qmp_capabilities` handshake, then send the hook's command (for these
|
||||
scenarios: `{"execute": "system_powerdown"}`). The socket is additive — no
|
||||
existing case may notice it. Note e3fe3f3 recently reworked how the harness
|
||||
boots; adapt to its current shape rather than the pre-rework description.
|
||||
|
||||
**M21.1 SCI details:** PM1_STS is at the event block base (write-1-to-clear);
|
||||
PM1_EN at base + block_len/2; PWRBTN bit is 8 in both. If PM1b exists, mirror
|
||||
reads/writes to both blocks. Enable ACPI mode only when SCI_EN (PM1 control
|
||||
bit 0) is clear — OVMF boots may already have it set. The publish path reuses
|
||||
the manager's subscriber table pattern (bounded, drop-on-failed-send).
|
||||
|
||||
**M21.2 GPE walk:** GPE0_STS bytes live at the GPE0 block base, GPE0_EN in
|
||||
the block's upper half; for a set+enabled bit n, the handler method is
|
||||
`_L%02X` (level) or `_E%02X` (edge) under `\_GPE`. Evaluate, drain the notify
|
||||
queue, clear the status bit, ack. A missing handler method is clear-and-log,
|
||||
not an error.
|
||||
|
||||
**M21.3 ordering:** init subscribes with retries — the acpi service registers
|
||||
`.power` well after init starts. The stop sequence runs vfs last (other
|
||||
services may flush through it). The S5 write mirrors `power.zig`'s
|
||||
`sleepValue` (SLP_TYP bits [12:10], SLP_EN bit 13); if the write returns, log
|
||||
`power: S5 write did not take` so the scenario fails loudly instead of
|
||||
hanging.
|
||||
|
||||
**Explicitly out of scope:** the embedded controller and `_Qxx` queries,
|
||||
battery `_BST`/`_BIF` evaluation beyond the interface stubs, lid/AC on QEMU
|
||||
(no emulation), reboot over the power protocol, S3 sleep, per-device D-states
|
||||
(a future lifecycle-vocabulary extension), thermal zones.
|
||||
+128
@@ -0,0 +1,128 @@
|
||||
# The power service: events and shutdown
|
||||
|
||||
A laptop lid closes, a battery drains, someone presses the power button — and
|
||||
several parts of the system might care: a session manager dims the screen, a
|
||||
logger notes it, and ultimately *something* has to turn the machine off. None of
|
||||
them owns the hardware that reported the event, and the reporter should not know
|
||||
who is listening. So system power is a **service**: an event source **publishes**
|
||||
button/lid/battery/AC events, interested processes **subscribe**, and one
|
||||
privileged caller — init — can ask it to power the machine off. It is the same
|
||||
publish/subscribe shape as the [input service](input.md), applied to power.
|
||||
|
||||
## Why a service, and why it is named for the domain, not the firmware
|
||||
|
||||
Where the events come from is firmware-specific — on x86 they ride the ACPI SCI
|
||||
([acpi.md](acpi.md)); on a Raspberry Pi they would come from PSCI or a mailbox.
|
||||
What subscribers want is not: *the lid closed* means the same thing regardless of
|
||||
who noticed. So the surface is **domain-named**. There is a `power-protocol`
|
||||
module and a well-known `ServiceId.power = 5`; on x86 the **acpi service**
|
||||
registers it, and on ARM a PSCI/mailbox service will register the *same* id.
|
||||
Subscribers call `runtime.ipc.lookup(.power)` and never learn which firmware they
|
||||
are on — the neutrality the whole [discovery](discovery.md) migration exists to
|
||||
preserve, carried one layer up into a running-system surface.
|
||||
|
||||
This is why the protocol is `power`, not "ACPI events": naming a cross-firmware
|
||||
surface after one firmware would leak x86 into code the ARM port must reuse
|
||||
unchanged.
|
||||
|
||||
## The protocol
|
||||
|
||||
The `power-protocol` module ([system/services/power/protocol.zig](../system/services/power/protocol.zig))
|
||||
follows the vfs-protocol pattern — extern-struct messages, a version, reserved
|
||||
fields. Three operations:
|
||||
|
||||
| Direction | Operation | Purpose |
|
||||
|---|---|---|
|
||||
| subscriber → service | `subscribe` | receive published events; the subscriber's endpoint rides as the call's **capability** (the input/device-manager pattern) |
|
||||
| init → service | `shutdown` | orderly shutdown's last step: enter S5 (soft off) |
|
||||
| service → subscriber | `event` | a published `EventMessage`, delivered as a buffered message (never sent *to* the service) |
|
||||
|
||||
Events are published, not polled: like the input service, the service holds
|
||||
subscriber endpoints as capabilities and `ipc_send`s each event as a buffered
|
||||
message, so a slow or dead subscriber can never wedge the source. The event
|
||||
vocabulary is hardware-neutral:
|
||||
|
||||
- `power_button` — the button was pressed (a fixed ACPI event on x86).
|
||||
- `lid`, `ac`, `battery` — the named GPE-driven events.
|
||||
- `notify` — a device notification that maps to none of the above; its `code`
|
||||
(the ACPI `Notify` argument) and the notifying device's `hid` say which device
|
||||
and what happened.
|
||||
|
||||
An `EventMessage` carries the `event` tag plus `code` and an 8-byte `hid`, so a
|
||||
generic `notify` is fully described without a second round trip.
|
||||
|
||||
**`shutdown` is authority, not information.** It is the only operation that
|
||||
*does* something irreversible, so it is gated: the contract is that only init
|
||||
(PID 1) may request it, because init is the process that has already run the stop
|
||||
sequence over everything else. The acpi service implements this as a **soft
|
||||
gate** — it honors `shutdown` only from a process that is a *subscriber*, and
|
||||
init is the one subscriber. That stands in for "only the system supervisor may
|
||||
power off" without hard-coding a pid, so it still holds under tests where PID 1
|
||||
is not init.
|
||||
|
||||
## Orderly shutdown
|
||||
|
||||
Powering off cleanly is where the power service, the [process
|
||||
lifecycle](process-lifecycle.md), and [ACPI events](acpi.md) compose. init
|
||||
already supervises the services it starts; for shutdown it runs **one event loop
|
||||
over one endpoint** that carries three things at once: its children's exit
|
||||
notifications, the lifecycle **signals** it can receive (`terminate`), and the
|
||||
**power events** it subscribes to — plus a re-arming heartbeat timer proving PID
|
||||
1 is alive. (init subscribes with retries, because the power service registers
|
||||
`.power` well after init starts; a missing power service is not fatal — a
|
||||
`terminate` signal drives the same path.)
|
||||
|
||||
On a `power_button` event or a `terminate` signal, init:
|
||||
|
||||
1. logs that it is shutting down,
|
||||
2. runs the standard stop sequence — `runtime.process.stop(child, deadline,
|
||||
endpoint)` — over its children **in reverse spawn order**, so the VFS stops
|
||||
last (other services may flush through it), each child getting the
|
||||
*terminate → deadline → kill* escalation from
|
||||
[process-lifecycle.md](process-lifecycle.md), and
|
||||
3. requests `.power` `shutdown`.
|
||||
|
||||
The service then enters **S5** (soft off) by writing `SLP_TYP | SLP_EN` to the
|
||||
PM1 control register(s) from ring 3, mirroring the kernel's own
|
||||
`system/devices/power.zig` `sleepValue`. If the write returns instead of powering
|
||||
the machine off, it logs loudly so a test fails rather than hangs.
|
||||
|
||||
**No new system call was needed for S5.** The broad io_port grant on the
|
||||
`acpi-tables` node ([discovery.md](discovery.md)) already put the PM1 control
|
||||
ports in the acpi service's hands, so writing S5 from ring 3 is something it
|
||||
could physically already do; formalizing it as a protocol operation added a
|
||||
contract, not authority. The kernel keeps `power.zig` for its own test paths and
|
||||
panic-time poweroff, where no user space is available to ask.
|
||||
|
||||
## Verifying it
|
||||
|
||||
Two QEMU scenarios exercise the path, both injecting a real ACPI power-button
|
||||
press via QMP `system_powerdown` (there is no other deterministic power event on
|
||||
this config):
|
||||
|
||||
- `power-button` proves the source: the acpi service's SCI handler logs the
|
||||
press and publishes `power_button` (the ACPI half is in [acpi.md](acpi.md)).
|
||||
- `orderly-shutdown` proves the whole composition: button → init logs shutting
|
||||
down → children stopped → the service enters S5 → QEMU exits. The ordered
|
||||
regex is the proof, and QEMU's self-exit through S5 is the pass.
|
||||
|
||||
## Scope
|
||||
|
||||
Interface-complete but validated on real hardware (the author's laptop) later,
|
||||
because QEMU does not emulate them: battery `_BST`/`_BIF` evaluation beyond the
|
||||
interface stubs, lid and AC events, and the embedded controller's `_Qxx`
|
||||
queries. Deliberately out of scope for now: reboot over the power protocol, S3
|
||||
sleep, per-device D-states (a future lifecycle-vocabulary extension, since
|
||||
"suspend" has the shape of a signal every driver must answer and has no consumer
|
||||
until laptop sleep), and thermal zones.
|
||||
|
||||
## See also
|
||||
|
||||
- [acpi.md](acpi.md) — where the events come from on x86: the SCI, the power
|
||||
button fixed event, and GPE/Notify dispatch in the acpi service.
|
||||
- [discovery.md](discovery.md) — why the surface is domain-named, and the
|
||||
firmware neutrality that makes a PSCI backend drop-in on ARM.
|
||||
- [process-lifecycle.md](process-lifecycle.md) — the stop sequence
|
||||
(`terminate → deadline → kill`) and signals init composes into shutdown.
|
||||
- [device-manager.md](device-manager.md) — the supervision model init mirrors for
|
||||
its own children.
|
||||
+1
-1
@@ -11,7 +11,7 @@ because the kernel releases a dead process's claims. The `driver-restart` and
|
||||
`usb-report` scenarios prove kill → release → respawn → re-claim → re-report
|
||||
end to end. What remains of this document's ladder is scope, not mechanism:
|
||||
more of the system moved into restartable processes (the discovery migration,
|
||||
[m19-m20-plan.md](m19-m20-plan.md), is the next rung). This is the property danos is really chasing:
|
||||
[discovery.md](discovery.md), is the next rung). This is the property danos is really chasing:
|
||||
**if a part of the OS breaks, isolate it, and re-initialise it — without rebooting.**
|
||||
A crashed driver gets restarted; a wedged service gets killed and brought back. It's
|
||||
the reason the [microkernel](vision.md) shape was chosen, and it's a *separate* goal
|
||||
|
||||
+117
@@ -0,0 +1,117 @@
|
||||
# Timers and time
|
||||
|
||||
Two different needs hide under the word "timer", and danos keeps them apart:
|
||||
|
||||
- **Reading the clock** — *what time is it?* A read of a free-running counter.
|
||||
- **Waiting** — *wake me in N milliseconds*, or *notify me when a deadline passes.*
|
||||
|
||||
Both are answered by the **kernel**, because the kernel already owns a timer: it has
|
||||
to, to preempt tasks. The LAPIC heartbeat and the calibrated TSC that back all of this
|
||||
are built in [device-interrupts.md](device-interrupts.md); the scheduler's blocking and
|
||||
wait queues are in [scheduling.md](scheduling.md). This page is about the surface a
|
||||
ring-3 program actually uses, and one deliberate absence: **there is no user-space time
|
||||
service.**
|
||||
|
||||
## Why time is a syscall, not a service
|
||||
|
||||
The tempting microkernel move is to put a timer *driver* in user space and have
|
||||
applications ask it for the time over IPC. For a **monotonic clock that is wrong** —
|
||||
reading `now()` should never cost an IPC round trip. The kernel is already holding the
|
||||
answer: it computes the current time every time it schedules, from the TSC, in a couple
|
||||
of instructions. Surfacing that as a system call is pure mechanism; routing it through a
|
||||
message to another process would be slower *and* redundant, and a device like the HPET
|
||||
(uncacheable MMIO reads) is a particularly bad thing to read on every `now()`.
|
||||
|
||||
This is the same conclusion every serious system reaches: Linux and Zircon read the
|
||||
counter in the vDSO, L4 exposes a clock field in a shared kernel page, seL4 reads the
|
||||
cycle counter directly. None of them make a clock read an IPC. danos makes it a syscall.
|
||||
|
||||
That "from the TSC" hides a portability question, because the TSC is only a valid clock
|
||||
when the CPU guarantees it is *invariant* and when every core's TSC is *synchronized*.
|
||||
danos checks both — the invariant-TSC CPUID bit (`0x80000007` EDX[8], set on Intel and
|
||||
AMD), and a cross-core "warp" check as the cores come up — and falls back to the HPET
|
||||
counter when either fails. So `now()` stays accurate on a real Intel box, a real AMD box,
|
||||
and inside a VM alike; only the source behind it differs. The mechanism is in
|
||||
[device-interrupts.md](device-interrupts.md).
|
||||
|
||||
So the timer hardware lives in the kernel, and there is **no `hpet` driver and no time
|
||||
server** to consume. (An earlier HPET driver existed only to *demonstrate* the driver
|
||||
model; that role now lives in [drivers.md](drivers.md), as documentation.) The one place
|
||||
a user-space time service *is* justified — **wall-clock / calendar time** — is discussed
|
||||
at the end; it is deliberately not built yet.
|
||||
|
||||
## The three system calls
|
||||
|
||||
Time and waiting are three entries in the small syscall table ([syscall.md](syscall.md)):
|
||||
|
||||
- **`clock` (#23)** → monotonic nanoseconds since boot. It only moves forward. Not
|
||||
wall-clock: no date, no timezone. Backed by `architecture.nanos()` (TSC, scaled with a
|
||||
128-bit intermediate so a long uptime can't overflow) — a few nanoseconds of
|
||||
resolution, and just an `rdtsc` plus a multiply.
|
||||
- **`sleep` (#3)** → block the caller for N milliseconds. The scheduler records a wake
|
||||
deadline and the tick sweep wakes it (`scheduler.sleep`).
|
||||
- **`timer_bind` (#31)** → arm a one-shot timer that, after N milliseconds, posts a
|
||||
**timer notification** to an IPC endpoint. Unlike `sleep` it does **not** block: a
|
||||
service can keep answering messages on the same endpoint while a deadline is pending.
|
||||
This is the timed wait that stop-sequence escalation, hello deadlines, and restart
|
||||
backoff are built from ([process-lifecycle.md](process-lifecycle.md),
|
||||
[device-manager.md](device-manager.md)).
|
||||
|
||||
The kernel's own scheduling timer (the LAPIC, vector 32) is never exposed to user space;
|
||||
programs read the TSC through `clock` and get timed wakeups through `sleep`/`timer_bind`,
|
||||
both riding the scheduler tick.
|
||||
|
||||
## `runtime.time` — the generic interface
|
||||
|
||||
Applications don't call the syscalls directly; they use `runtime.time`
|
||||
(`library/runtime/time.zig`), a thin `Instant`/`Duration` layer over them — an ergonomic
|
||||
front door, not new mechanism.
|
||||
|
||||
```zig
|
||||
const time = @import("runtime").time;
|
||||
|
||||
const start = time.now(); // Instant — monotonic
|
||||
doWork();
|
||||
const took = start.elapsed(); // Duration
|
||||
time.sleep(time.Duration.fromMillis(5)); // block ~5 ms
|
||||
|
||||
// A deadline delivered as a notification, so a service keeps serving meanwhile:
|
||||
_ = time.after(endpoint, time.Duration.fromMillis(200));
|
||||
```
|
||||
|
||||
- `Duration` is nanoseconds under the hood, with `fromNanos/fromMicros/fromMillis/
|
||||
fromSeconds` and `asNanos/asMillis`. `ceilMillis` rounds *up* to the kernel's
|
||||
millisecond granularity, so a sub-millisecond `sleep` never rounds down to zero and
|
||||
returns early. All arithmetic saturates rather than wraps.
|
||||
- `Instant` is a point on the monotonic clock: `since`, `elapsed`, `plus`, `reached` —
|
||||
built for deadline loops (`while (!deadline.reached()) …`).
|
||||
- `now()` / `monotonicNanos()` wrap `clock`. `available()` reports whether the clock is
|
||||
calibrated at all (the kernel returns 0 until the TSC frequency is known, so a caller
|
||||
that needs real time can treat 0 as "unavailable" rather than assume it advances).
|
||||
- `sleep(d)` wraps `sleep`; `spin(d)` busy-polls `now()` for the sub-millisecond delays
|
||||
the millisecond tick can't express; `after(endpoint, d)` wraps `timer_bind`.
|
||||
|
||||
The raw wrappers (`system.clock`, `system.sleep`, `system.timerOnce`) stay in
|
||||
`library/runtime/system.zig`; `runtime.time` is the layer meant for everyday use.
|
||||
|
||||
## Wall-clock time (not built)
|
||||
|
||||
Everything above is **monotonic**: elapsed time since boot, perfect for timeouts and
|
||||
measurement, useless for "what is the date?" Calendar time — a real-time clock, time
|
||||
zones, leap seconds — is genuinely a **user-space** concern, and it *is* the case a time
|
||||
service is for. It would be backed by an **RTC** driver (the CMOS real-time clock), not
|
||||
the HPET, and exposed as a `CLOCK_REALTIME`-style service alongside the monotonic
|
||||
syscall. It is deferred until something needs it; the monotonic clock the kernel already
|
||||
owns covers every current use.
|
||||
|
||||
## Verifying it
|
||||
|
||||
`runtime.time`'s `Instant`/`Duration` arithmetic has unit tests that run on the host:
|
||||
|
||||
```
|
||||
$ zig build test # includes library/runtime/time.zig
|
||||
```
|
||||
|
||||
End to end, the proof the clock is real is that it *advances*: read `now()`, `sleep` a
|
||||
`Duration`, read `now()` again, and the second reading is later — the kernel's timer
|
||||
driving a ring-3 program with no service in between.
|
||||
Reference in New Issue
Block a user