Time is a kernel concern in danos: the kernel owns the scheduling timer and already exposes monotonic time via the clock/sleep/timer_bind syscalls, so a userspace time service would be a redundant, slower path. This adds the generic runtime.time module over those syscalls, retires the two demonstration drivers, reorganizes the milestone docs, and makes the monotonic clock correct on Intel, AMD, and inside any VM. runtime.time (library/runtime/time.zig) - Instant/Duration interface: now, sleep, spin, after, monotonicNanos, available - a thin layer over system.clock/sleep/timerOnce; unit-tested arithmetic Remove the demo drivers hpet and bus (a teaching example belongs in the docs, not shipped in the tree) - system/drivers/ now holds only real drivers: pci-bus, ps2-bus, usb-xhci-bus - device-manager end-to-end test repointed to pci-bus (asserts on kernel state: the process table and the device tree, not a racy serial marker) - device_register containment moved to a new in-kernel `containment` test - the driver-model worked example moved inline into docs/drivers.md Reorganize milestone docs into topic docs - m17-m18 / m19-m20 / m21 plans dissolved into process-lifecycle, device-manager, discovery, and acpi docs; new docs/power.md and docs/timers.md; ~20 citations repointed; plan docs deleted TSC reliability (apic.zig, smp.zig, cpu.zig, kernel.zig) - check the invariant-TSC bit (CPUID 0x80000007 EDX[8]) on Intel and AMD - cross-core "warp" check at SMP bring-up, pairwise BSP<->AP as each core comes up - fall back to the HPET clocksource when the TSC is not invariant (a bare VM) or not synchronized (a warp), switched continuously so time never jumps - boot log reports the outcome; new tsc-sync test exercises the TSC + warp path Verified: zig build; zig build test; 60/60 QEMU cases (incl. new containment and tsc-sync).
248 lines
13 KiB
Markdown
248 lines
13 KiB
Markdown
# Device discovery (ACPI / device tree), the agnostic way
|
||
|
||
"Device discovery" is how the kernel learns **what hardware exists and where** — the
|
||
MMIO addresses, IRQ numbers, CPU count, and interrupt controller it can't just
|
||
assume. On x86 that description comes from **ACPI** tables; on ARM from a **device
|
||
tree** (DTB). This note is a design plan, not built yet: *when* danos should tackle
|
||
it, and *how* to keep it architecture-agnostic — the same discipline the
|
||
[memory map](memory-map.md) and [arch split](arch.md) already follow.
|
||
|
||
## What the kernel assumes today
|
||
|
||
Right now danos discovers almost nothing — it coasts on legacy PC fixtures that are
|
||
guaranteed to exist under QEMU + UEFI:
|
||
|
||
- `arch/x86_64/apic.zig` assumes the **Local APIC** at the default `0xFEE0_0000` and
|
||
calibrates its timer against the **PIT** (the legacy 8254).
|
||
- `arch/x86_64/serial.zig` hardcodes **COM1** at I/O port `0x3F8`.
|
||
- The framebuffer and memory map come from **UEFI** — that *is* discovery, just done
|
||
by the firmware and handed over, not read from ACPI.
|
||
|
||
This works only because PC-compatible hardware promises those legacy pieces exist at
|
||
those addresses. It is a crutch, and it does not travel.
|
||
|
||
## The forcing functions: when to build it
|
||
|
||
Two things drive the need, and they set the timing:
|
||
|
||
1. **The second architecture makes it mandatory.** ARM has *no* legacy fixtures —
|
||
no PIT, no fixed serial port, no standard interrupt controller address. You can't
|
||
find the UART to print a character without reading the device tree. So on x86 we
|
||
can defer discovery a long time (until we want the IOAPIC, PCIe, or SMP), but on
|
||
**aarch64 it's required to boot at all**. The [aarch64 port](arm.md) is what
|
||
forces the issue.
|
||
|
||
2. **Isolated user-space drivers need it.** In the [microkernel vision](vision.md),
|
||
drivers live in user space — but something has to enumerate the hardware and hand
|
||
each driver its MMIO regions and IRQs. That enumeration *is* device discovery. So
|
||
discovery is a prerequisite for real drivers, **not** for user mode itself.
|
||
|
||
The conclusion on timing: **don't build full discovery before user mode.** User mode
|
||
+ address-space isolation needs none of it; the current assumptions are fine there.
|
||
Build the agnostic discovery layer **when the aarch64 port starts** — because that's
|
||
when a second, real implementation makes "agnostic" honest.
|
||
|
||
## Why build it *with* the second arch, not before
|
||
|
||
The same lesson as the arch split: an abstraction with only one implementation
|
||
quietly bends to that implementation. Build "agnostic discovery" x86-only first and
|
||
you'll get an ACPI-shaped interface with device tree bolted on afterward. Build it
|
||
when aarch64 lands and the small, clean DTB parser pulls the abstraction toward the
|
||
right neutral shape, which the x86 side then fills. Design it against two backends or
|
||
it isn't really agnostic.
|
||
|
||
## The one cheap step to take sooner
|
||
|
||
Have the **loader capture the description pointer** into `BootInfo` — a neutral
|
||
handle, no parsing:
|
||
|
||
```zig
|
||
pub const HardwareInfo = extern struct {
|
||
kind: enum(u32) { none, acpi, device_tree },
|
||
addr: u64, // ACPI RSDP, or the DTB blob
|
||
};
|
||
```
|
||
|
||
On x86-UEFI that's the ACPI **RSDP**, read from the UEFI configuration table *before*
|
||
`ExitBootServices` — the same "grab it before exit" pattern as the framebuffer and
|
||
memory map (already flagged in [memory-map.md](memory-map.md)). On ARM it's the DTB
|
||
pointer the firmware passes. This keeps the door open for near-zero cost without
|
||
committing to the parser.
|
||
|
||
## What "agnostic" looks like
|
||
|
||
The happy accident: **the device-tree data model is already a good neutral
|
||
representation.** A DTB is a tree of nodes, each with:
|
||
|
||
- a **`compatible`** string — what the device is,
|
||
- **`reg`** — its MMIO base(s) and size(s),
|
||
- **`interrupts`** — its IRQ number(s),
|
||
- other properties.
|
||
|
||
Even OSes running on ACPI hardware normalize into a unified device model shaped like
|
||
this. So the neutral layer is a **device model** — "here are the devices, each with a
|
||
type, MMIO regions, and IRQs" — fed by two backends behind it:
|
||
|
||
- a **DTB parser** (ARM) — a few hundred lines against a well-specified binary format;
|
||
- **static ACPI + PCI enumeration** (x86).
|
||
|
||
### The caveat that sets the effort: "ACPI" ≠ "AML interpreter"
|
||
|
||
A full ACPI namespace is **AML** (ACPI Machine Language) bytecode, and writing an AML
|
||
interpreter is an enormous undertaking. **We skip it.** Interrupt routing and PCIe
|
||
come entirely from the *static* tables:
|
||
|
||
- **MADT** — the APICs and the **IOAPIC** (interrupt routing),
|
||
- **MCFG** — PCIe configuration space (then walk the PCI bus to enumerate devices),
|
||
- **FADT**, **HPET** — power/reset and a precise timer.
|
||
|
||
Walking PCI config space finds most devices without any AML. So the x86 static-ACPI
|
||
backend is roughly comparable in scope to the DTB parser; it's full AML that's the
|
||
monster, and it isn't on the path.
|
||
|
||
## Where it sits in the seam
|
||
|
||
Discovery layers onto the existing loader↔kernel split cleanly:
|
||
|
||
| Layer | Responsibility | x86 | ARM |
|
||
|-------|----------------|-----|-----|
|
||
| **Capture** (loader) | grab the description pointer | RSDP from UEFI config table | DTB pointer from firmware |
|
||
| **Parse** (kernel, per-mechanism) | pointer → neutral device model | static ACPI + PCI | DTB parser |
|
||
| **Consume** (generic) | use the device model | — same code — | — same code — |
|
||
|
||
Boot-protocol-specific capture, mechanism-specific parse, generic consumption —
|
||
exactly like the memory map, where the loader classifies and the kernel just sees
|
||
neutral regions.
|
||
|
||
## Where it lives in a microkernel
|
||
|
||
For isolated drivers, discovery is not one lump — it splits by *who needs it and
|
||
when*:
|
||
|
||
- **Minimal, in-kernel: the interrupt controller and the timer.** The IOAPIC/GIC and
|
||
the timer are needed *before* user space exists (scheduling and preemption depend
|
||
on them), so the kernel must parse at least these from ACPI/DTB itself. This is also
|
||
the **first real consumer** of discovery — the first genuine reason to read MADT or
|
||
the DTB is "where is the interrupt controller and how do I route an IRQ?" It's the
|
||
point where x86 finally graduates from the PIT/legacy-LAPIC assumptions.
|
||
- **User-space enumeration: a device-manager server.** Everything else — PCI devices,
|
||
peripherals — is parsed (or queried from the kernel's parse) by a privileged
|
||
user-space server that hands each driver process its MMIO regions and IRQ rights
|
||
over [IPC](ipc.md). Combined with **interrupts-as-messages** (an IRQ delivered to a
|
||
driver as a message on a channel — a natural extension of the wait queues and
|
||
channels already built), that's what makes drivers genuinely isolated.
|
||
|
||
So the device-manager server depends on user mode + IPC; only the interrupt-controller
|
||
slice is unavoidably in-kernel.
|
||
|
||
## A nice ARM contrast
|
||
|
||
On ARMv8 the generic timer exposes its frequency directly via the `CNTFRQ` register —
|
||
no calibration needed. That's cleaner than the x86 side, where we measure the LAPIC
|
||
and TSC against the PIT because nothing tells us their frequency (see
|
||
[device-interrupts.md](device-interrupts.md)). Discovery on ARM hands you more for
|
||
free; discovery on x86 is partly about *finding* what ARM just tells you.
|
||
|
||
## Suggested ordering
|
||
|
||
1. **Now (cheap):** plumb the neutral `HardwareInfo` pointer through the loader into
|
||
`BootInfo`. No parser yet.
|
||
2. **Next milestone unchanged:** user mode + address-space isolation — needs no
|
||
discovery.
|
||
3. **With the aarch64 port:** build the agnostic discovery layer, **DTB first** (the
|
||
forcing function), then x86 static-ACPI to fill the same model — starting with the
|
||
**interrupt controller + timer**.
|
||
4. **Then:** a user-space device-manager server + interrupts-as-messages → real
|
||
isolated drivers (keyboard first).
|
||
|
||
## Related
|
||
|
||
- [acpi.md](acpi.md) — the built x86 side of the "capture" step: how the loader grabs
|
||
the RSDP and the platform follows it to the RSDT/XSDT and the SDTs.
|
||
- [arm.md](arm.md) — the aarch64 target that forces genuine discovery (DTB, GIC).
|
||
- [memory-map.md](memory-map.md) — the same loader-captures / kernel-consumes seam,
|
||
and the note about grabbing the RSDP before `ExitBootServices`.
|
||
- [device-interrupts.md](device-interrupts.md) — the LAPIC/timer bring-up that
|
||
discovery will eventually feed (IOAPIC, real IRQ routing).
|
||
- [ipc.md](ipc.md) — the channels that interrupts-as-messages and the device manager
|
||
will ride on.
|
||
- [vision.md](vision.md) — why drivers belong in isolated user space at all.
|
||
|
||
## Update (M19.3, 2026-07-13): PCI enumeration left the kernel
|
||
|
||
The kernel now seeds only the `pci_host_bridge` node (ECAM window, MMIO
|
||
apertures derived from the memory map's holes, bus range, and the 16-bit I/O
|
||
window). The per-function walk moved to the ring-3 `pci-bus` driver
|
||
([device-manager.md](device-manager.md)): it claims the bridge, repeats the
|
||
ECAM scan through its mmio grant, and `device_register`s what it finds, which
|
||
the device manager mirrors and matches. The ACPI namespace walk follows in M20;
|
||
the static tables (MADT, HPET, MCFG, FADT + `\\_S5`) stay kernel-side.
|
||
|
||
## Update (M20.3, 2026-07-13): ACPI enumeration left the kernel too
|
||
|
||
The kernel no longer folds the AML namespace's Device objects into the device
|
||
tree. It still parses the *static* tables (MADT for SMP, HPET for the tick, MCFG
|
||
for the host bridge, FADT) and still builds the AML namespace — but only to read
|
||
the `\\_S5` sleep type for poweroff. Device discovery is the ring-3 **acpi
|
||
service** ([device-manager.md](device-manager.md)): it claims the `acpi-tables`
|
||
node the kernel publishes (the AML blobs, a broad io_port grant, the SCI),
|
||
re-parses the same blobs with the shared AML module, evaluates `_STA`/`_CRS`,
|
||
and registers + reports each `_HID` device — the device manager matches drivers
|
||
(ps2-bus) from those reports. With M19's pci-bus driver, discovery now runs
|
||
entirely in user space; the kernel seeds only the host bridge and the
|
||
acpi-tables node.
|
||
|
||
## Discovery is a swappable process per firmware (M19–M20)
|
||
|
||
Moving PCI and ACPI enumeration out of ring 0 was not just a relocation — it
|
||
made discovery **firmware-neutral by construction**, which is the whole reason
|
||
to do it before the second architecture rather than after. Everything at and
|
||
above the [device-manager](device-manager.md) protocol — descriptors,
|
||
containment, reports, matching, supervision — is generic and may never become
|
||
x86-specific. Discovery is the single firmware-specific piece, and it is
|
||
isolated as **one swappable process per firmware**:
|
||
|
||
- **x86** boots describe hardware with ACPI, so the discoverer is the **acpi
|
||
service** ([acpi.md](acpi.md)): it claims the `acpi-tables` node and runs AML.
|
||
- **The Raspberry Pis** hand over a flattened device tree, so the discoverer is
|
||
an **fdt service**: it claims a `devicetree-blob` node and walks the tree —
|
||
pure data, no bytecode, so it needs neither a port grant nor an interpreter,
|
||
strictly simpler than ACPI. (A placeholder until the [aarch64](arm.md)
|
||
bring-up fills it in.)
|
||
|
||
The device manager spawns the discoverer under the **neutral ramdisk name
|
||
`discovery`** and never learns which firmware it is on; the build's
|
||
`-Ddiscovery=acpi|fdt` option fills that slot (x86 defaults to `acpi`, the
|
||
aarch64 target flips the default when it lands). The manager owns the device
|
||
tree as *data* and touches no hardware, ever — firmware bytecode runs only
|
||
inside the crashable, supervised discoverer, so an AML fault can never take
|
||
down the supervisor.
|
||
|
||
Two consequences of neutrality bind on later work:
|
||
|
||
- **Cross-firmware surfaces are named by domain, not firmware.** System power is
|
||
a [`power`](power.md) protocol, not an "ACPI events" protocol: on x86 the acpi
|
||
service registers it, on ARM a PSCI/mailbox service registers the same
|
||
`ServiceId.power`, and subscribers never learn the difference.
|
||
- **Identity must widen before the fdt service exists.** `DeviceDescriptor`'s
|
||
8-byte `hid` holds an EISA id but cannot hold an FDT `compatible` string
|
||
(`"brcm,bcm2835-aux-uart"`); the identity field grows before the ARM path can
|
||
report a real node.
|
||
|
||
Two supporting decisions keep the kernel's remaining slice honest:
|
||
|
||
- **The AML interpreter is a shared build module**, compiled into both the
|
||
kernel and the acpi service — one source, two builds, no fork. The kernel
|
||
links it for the `\_S5` poweroff evaluation, the service links it for
|
||
everything else, and the `acpi-parse` test asserts the two produce the same
|
||
device count across the ring-3 move.
|
||
- **Bridge apertures come from the firmware memory map, not AML.** Registered
|
||
PCI functions carry BAR resources, and `device_register` containment demands
|
||
the bridge own windows that cover them. Those apertures are derived
|
||
kernel-side from the boot memory map's MMIO holes (regions that are neither
|
||
RAM nor tables) — mechanical, AML-free, and available at boot regardless of
|
||
what later moved to user space. The acpi service's authority is likewise
|
||
exactly one node: the `acpi-tables` node, whose broad io_port grant is the
|
||
documented trust boundary for the one process allowed to run firmware
|
||
bytecode.
|