Author SHA1 Message Date
Daniel Samson 0628944b15 Docs: close the discovery migration (M20.3)
discovery.md records ACPI enumeration leaving the kernel; device-manager.md
increment 8 marked done — enumeration now runs entirely in ring 3.
2026-07-13 03:32:53 +01:00
Daniel Samson e6d0bb7ef0 The flip: ACPI enumeration leaves the kernel (M20.3)
The kernel no longer folds AML Device objects into the device tree — the
ring-3 acpi service is the sole builder of _HID device nodes. The kernel
keeps building the namespace only for the \_S5 sleep type, and still
seeds the static tables (MADT, HPET, MCFG, FADT) and the acpi-tables node.

The device manager matches ps2-bus from the service's _HID reports
(PNP0303 / PNP0F13, singleton-deduped) instead of boot-snapshot nodes;
its dead boot-snapshot ps2 arm is gone. The service registers every
device before reporting any, so a driver the manager spawns on the first
report already sees the full set — no keyboard-before-mouse race. The
acpi-ps2 scenario proves the whole chain: report -> spawn -> ps2-bus
finds the controller and attaches its keyboard, entirely in ring 3. The
ioport test moved to the acpi-tables I/O window, since the kernel-built
PS/2 node it used to scan for no longer exists. The retired
device-building functions in acpi.zig are dead but retained (a botched
mechanical deletion is worse mid-migration than a follow-up sweep, which
is flagged as a task). Suite 58/58.
2026-07-13 03:32:15 +01:00
Daniel Samson 5ca804d827 The acpi service evaluates _CRS/_STA in ring 3 and reports devices (M20.2)
AML method evaluation now runs in userspace touching real hardware: the
service builds an interpreter with a ring-3 Hal (port I/O routed through
its claimed acpi-tables node; a scratch page backs SystemMemory maps so a
stray OperationRegion degrades to zeros instead of faulting a process
that cannot map arbitrary physical memory). It walks the namespace and,
for each present _HID device that is not a PCI root, evaluates _CRS,
registers it under acpi-tables, and reports it with its EISA-decoded hid.

Containment for this needed the broker's irq check to become range-based
— an interrupt line is still indivisible, but a parent may own a range,
so the acpi-tables node's broad irq window contains its children's legacy
lines (a length-1 range is exactly the old equality, so single-irq
parents are unaffected). ChildAdded gained a hid field for firmware
string identity. Matching those reports to drivers stays off until M20.3,
so ps2-bus still comes up via the kernel path — no regression. The
acpi-report scenario proves the PS/2 keyboard (io 0x60/0x64 + IRQ) and
mouse (IRQ) are reported with their resources. Suite 57/57.
2026-07-13 03:19:39 +01:00
Daniel Samson a299363b59 The AML interpreter runs in ring 3: the acpi service parses (M20.1)
The AML module becomes a build module compiled into both the kernel (for
the \_S5 sleep state it still needs) and the new acpi service — one
source, two builds, no fork. The kernel publishes a single acpi-tables
node: the DSDT/SSDT blobs as memory resources, a broad io_port grant (the
honest trust boundary — firmware AML names whatever ports it chose, known
only after parsing), and the SCI for the M21 event track. The acpi
service claims the node, maps each blob through the ordinary mmio grant
(which preserves the sub-page offset onto the bytecode), and runs the
same parser the kernel does. It self-verifies its namespace Device count
against the kernel's — 34 = 34 — deterministically via an argv the
acpi-parse test passes, so no racing the shared serial buffer. Parse-only
touches no hardware; OperationRegion evaluation waits for _CRS/_STA in
M20.2. The manager spawns 'discovery' (the neutral ramdisk name) at
startup. Suite 56/56.
2026-07-13 03:07:11 +01:00
Daniel Samson d8dd62c639 Mark the feat/pci-bus merge done in the M19-M20 plan 2026-07-13 02:55:03 +01:00
Daniel Samson d106b6e8dc Merge feat/pci-bus: PCI enumeration in ring 3 (M19)
The host bridge apertures, idempotent device_register, the pci-bus driver
(scan, register, report), and the flip that retired the kernel's PCI walk
— discovery's first subsystem to leave ring 0.
2026-07-13 02:55:03 +01:00
Daniel Samson af2c766f42 The flip: PCI enumeration leaves the kernel (M19.3)
enumeratePci, addBars, pciConfigurationPtr, and the PciHeader struct are
deleted; the kernel seeds only the host bridge, and the ring-3 pci-bus
driver's reports are the sole source of PCI function nodes. The manager
matches PCI drivers from reported identity, deduped by registered device
id so a bus restart never double-spawns.

The flip did its job by exposing a latent SMP race: ring-3
device_register made the broker table concurrent for the first time, and
mmio_map read it lock-free — under load a torn resource length mapped
hpet's window wrong (its user fault) and underflowed r.len-1 into a
kernel integer-overflow panic. Fixed: the broker read in mmio_map (and
claim) runs under the big kernel lock, the arithmetic rejects
zero-length and wrapping windows cleanly, and pci-bus no longer registers
unimplemented size-0 BARs. driver-restart hammered 6x, suite 55/55.
2026-07-13 02:54:50 +01:00
Daniel Samson d26262bf56 pci-bus registers and reports what it scans (M19.2)
Each function is registered under the bridge with the config-space slice
and BARs sized by the same all-ones probe the kernel uses — byte-for-byte
equal descriptors, so the idempotent register returns the kernel's
existing node ids during coexistence instead of duplicating the tree.
The bridge gained the 16-bit io_port aperture that functions' I/O BARs
need to pass containment. Reports carry the registered device_id, and
the pci-scan scenario drills a forced restart: kill the enumerator after
its reports, watch the respawn re-scan, and assert the broker's PCI node
count never grew. The usb-restart test trigger is pinned to the xHCI
reporter (pci-bus racing it to two reports used to steal the kill).
Harness hardening: failing cases preserve their serial logs; the heavy
scenarios run at 150s.
2026-07-13 02:22:19 +01:00
Daniel Samson 10b89c06ff The PCI scan from ring 3: pci-bus walks the ECAM it mapped (M19.1)
The manager matches the pci_host_bridge node and spawns pci-bus with the
bridge id as its assignment — hello, supervision, restart, all the M18
contract for free. The driver claims the bridge, maps the ECAM window
(resource 0) through the ordinary mmio grant, and repeats the kernel's
brute-force bus/device/function walk from user space. The pci-scan
scenario builds its expected marker from the kernel's own function count,
so the two enumerations must agree exactly — the equivalence that
licenses retiring the kernel walk in M19.3.
2026-07-13 02:05:32 +01:00
Daniel Samson a2a05d0b3d Discovery-migration prerequisites (M19.0)
The host bridge now carries MMIO apertures derived from the boot memory
map's gaps below 4 GiB (largest three, sort-merged; a single after-the-
last-region hole dies on OVMF's flash at the top) plus one aperture above
the described space — so a user-space device_register of PCI functions
with BAR resources can pass containment. The discovery test asserts
every PCI memory resource lies inside a bridge window and names any
escapee. device_register is idempotent on exact (parent, class, identity,
resources) match — a restarted registering bus cannot duplicate its
children; proven directly against the broker in the bus test.
ChildAdded gains device_id so a report can carry the registered kernel
id a matched driver needs as its assignment.
2026-07-13 01:59:32 +01:00
Daniel Samson 75d62660b0 Bring the docs up to the M17-M18 reality; pre-settle M19-M20 ambiguities
resilience.md: steps 1-4 of the ladder are built — supervision, exit
reasons, restart with backoff, crash-loop caps, all proven by scenario;
what remains is scope, not mechanism. README statuses follow. drivers.md
gains the driver-contract section (harness, hello, crash-freely). The
M19-M20 plan pre-settles three things the loop would otherwise have had
to decide alone: the memory-map pass-through for apertures, the manager
spawning 'discovery' from M20.1, and hid[8] riding ChildAdded for ACPI
string identity until the FDT widening.
2026-07-13 01:49:52 +01:00
Daniel Samson a53c2b0193 Placeholder discovery services and the -Ddiscovery build option
system/services/acpi and system/services/fdt exist as documented
placeholders (silent clean-exit mains; the headers say exactly what each
becomes and why). The build's -Ddiscovery=acpi|fdt option fills the
ramdisk's neutral 'discovery' slot — the device manager will spawn
"discovery" by that name in M20.3 and never learn which firmware it is
on (m19-m20-plan.md decision 7). x86 defaults to acpi; the aarch64
target flips the default when it lands.
2026-07-13 01:46:04 +01:00
Daniel Samson bf481c080c Record the firmware-neutrality contract as decision 7
Discovery is one swappable process per firmware (acpi service on x86, an
fdt service on the Pis); everything at and above the device-manager
protocol stays generic. The manager owns the tree as data and never
touches hardware — firmware bytecode runs in a crashable, supervised
discoverer. Flagged now: hid[8] cannot hold an FDT compatible string,
and cross-firmware protocols are named by domain (power, not ACPI).
2026-07-13 01:33:12 +01:00
Daniel Samson 3a78dcab3f Scope ACPI events and system power as M21; record the SCI on acpi-tables
Battery, AC, lid, and the power button ride the acpi service as reported
children with small class drivers — the xHCI split repeated. QEMU can
only prove the power-button path (system_powerdown injects the real fixed
event), so battery/EC are interface-complete and hardware-validated on
the laptop. Per-device power states (D-states, suspend/resume) stay out
of scope: suspend has the shape of a lifecycle signal every driver must
answer, and it has no consumer until laptop sleep.
2026-07-13 01:21:23 +01:00
Daniel Samson 470f93a83d Plan the discovery migration (M19 pci-bus, M20 acpi service) 2026-07-13 01:15:17 +01:00
Daniel Samson 7798706b41 Mark the M17-M18 plan complete 2026-07-13 00:49:04 +01:00
Daniel Samson ad40de03c2 Merge feat/usb-xhci-bus: xHCI port scan, tree reports, and the app surface (M18.2-M18.3) 2026-07-13 00:49:04 +01:00
Daniel Samson d8778b4b70 The application surface: enumerate, subscribe, and device-list (M18.3)
Applications ask the device manager for the tree (enumerate: a header
plus ChildEntry records) and subscribe to published add/remove events by
handing their endpoint over as the call's capability — the input-service
pattern; events are the same ChildAdded/ChildRemoved structs the bus
drivers send, one encoding in both directions. device-list is the first
client: it prints the tree, subscribes, and narrates the events through
a driver restart. The protocol's message maximum is capped at the
kernel's IPC MESSAGE_MAXIMUM (256 bytes, ten entries per reply; paging
joins the protocol when a tree outgrows one message). The startUserTask
debug print is gone: it wrote to serial unserialized against user-space
lines and sheared concurrent log markers in half — the root cause of the
scenario flakes.
2026-07-13 00:49:03 +01:00
Daniel Samson 79d859a111 The xHCI driver scans its root-hub ports and reports the tree (M18.2)
child_added/child_removed join the device-manager protocol. The driver
maps its register BAR (resource 0 is the ECAM config space; the walk
starts at 1), reads CAPLENGTH and HCSPARAMS1, and reads one PORTSC per
port: the connect bit and speed class come straight from hardware, no
rings needed to see the devices. The manager mirrors reported children
keyed by (parent, port), remembers which instance reported each, and
prunes a dead reporter's children before deciding the restart — the
children describe protocol state that died with the process. The
usb-report scenario drives the whole loop: two QEMU devices reported,
reporter killed, children pruned, driver respawned with backoff, and the
new instance re-claims, re-scans, and re-reports.
2026-07-13 00:28:29 +01:00
Daniel Samson 37fb09f75e Mark the feat/device-manager merge done in the M17-M18 plan 2026-07-13 00:19:32 +01:00
Daniel Samson 34ebeb968d Merge feat/device-manager: the supervising device manager (M18.1) 2026-07-13 00:19:32 +01:00
27 changed files with 2138 additions and 217 deletions
+30
View File
@@ -131,6 +131,13 @@ pub fn build(b: *std.Build) void {
});
// ACPI/PnP hardware-ID (_HID) names — the flat analog of pci-class for acpi_device
// nodes. Also shared reference data.
// The AML interpreter, a build module so the ring-3 acpi service can run the
// same parser the kernel does (docs/m19-m20-plan.md decision 1). Pure Zig,
// no kernel imports — one source, two builds.
const aml_module = b.addModule("aml", .{
.root_source_file = b.path("system/devices/aml/aml.zig"),
});
const acpi_ids_module = b.addModule("acpi-ids", .{
.root_source_file = b.path("system/devices/acpi-ids.zig"),
});
@@ -339,9 +346,26 @@ pub fn build(b: *std.Build) void {
const ps2_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-keyboard", "system/drivers/ps2-bus/keyboard.zig");
const ps2_mouse_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-mouse", "system/drivers/ps2-bus/mouse.zig");
const usb_xhci_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-xhci-bus", "system/drivers/usb-xhci-bus/usb-xhci-bus.zig");
const pci_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "pci-bus", "system/drivers/pci-bus/pci-bus.zig");
// A test fixture, not a real driver: hellos to the device manager, then faults —
// what the driver-restart scenario drives the crash-loop cap with.
const crash_test_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "crash-test", "system/services/crash-test/crash-test.zig");
const device_list_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-list", "system/services/device-list/device-list.zig");
// The discovery service: one swappable process per firmware
// (docs/m19-m20-plan.md decision 7), bundled under the neutral ramdisk name
// "discovery" so the device manager never learns which firmware it is on.
// x86 boots describe hardware with ACPI; the Raspberry Pis hand over a
// flattened device tree — the aarch64 target flips the default when it
// lands (docs/arm.md). Both are placeholders until M20.1 (acpi) and the
// ARM bring-up (fdt).
const Discovery = enum { acpi, fdt };
const discovery = b.option(Discovery, "discovery", "Which discovery service fills the ramdisk's 'discovery' slot (default: acpi)") orelse Discovery.acpi;
const discovery_source: []const u8 = switch (discovery) {
.acpi => "system/services/acpi/acpi.zig",
.fdt => "system/services/fdt/fdt.zig",
};
const discovery_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "discovery", discovery_source);
if (discovery == .acpi) discovery_exe.root_module.addImport("aml", aml_module);
const device_manager_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-manager", "system/services/device-manager/device-manager.zig");
// The input service and its exercisers: the fan-out server, a hardware-free synthetic
// source, and a subscriber that doubles as the `input` test's oracle. See docs/input.md.
@@ -373,8 +397,14 @@ pub fn build(b: *std.Build) void {
mk_run.addFileArg(ps2_mouse_exe.getEmittedBin());
mk_run.addArg("usb-xhci-bus");
mk_run.addFileArg(usb_xhci_bus_exe.getEmittedBin());
mk_run.addArg("pci-bus");
mk_run.addFileArg(pci_bus_exe.getEmittedBin());
mk_run.addArg("crash-test");
mk_run.addFileArg(crash_test_exe.getEmittedBin());
mk_run.addArg("device-list");
mk_run.addFileArg(device_list_exe.getEmittedBin());
mk_run.addArg("discovery");
mk_run.addFileArg(discovery_exe.getEmittedBin());
mk_run.addArg("device-manager");
mk_run.addFileArg(device_manager_exe.getEmittedBin());
mk_run.addArg("input");
+8 -6
View File
@@ -45,20 +45,22 @@ rather than restate it. Roughly in the order things happen at runtime:
until its hardware interrupts it**. The claim is the capability; `irq_ack` is the
unmask.
14. **[driver-model.md](driver-model.md) — buses, classes and host controllers.** How
real driver stacks factor into three shapes, how families share code, and the
proposed ABI for the three primitives still missing (capability passing, DMA +
memory barriers, MSI).
real driver stacks factor into three shapes and how families share code. The
three primitives it proposed are long since built (M13 capability passing,
M14 DMA + barriers, M15 MSI), and the driver *contract* on top of them —
hello, supervision, restart — is built too (device-manager.md, M18).
15. **[process-management.md](process-management.md) — process management.** The
microkernel's `ps`/`kill`/SIGCHLD: enumerate as a table snapshot, the
supervision link as the kill authority, and child-exit notifications over the
same endpoints IRQs arrive on.
16. **[process-lifecycle.md](process-lifecycle.md) — the process lifecycle.** Design:
signals over IPC as the one lifecycle vocabulary every process speaks — the
16. **[process-lifecycle.md](process-lifecycle.md) — the process lifecycle.** Built
(M17): signals over IPC as the one lifecycle vocabulary every process speaks — the
POSIX.1-1990 words with message delivery instead of stack hijack, the stable
`runtime.process` interface, exit reasons, published exit events any stateful
service can subscribe to (the VFS releasing dead clients' handles), and the two
iron rules (cleanup is the kernel's job; kill is not a signal).
17. **[device-manager.md](device-manager.md) — the device manager.** Design: the
17. **[device-manager.md](device-manager.md) — the device manager.** Built (M18,
through the app surface): the
tree, the matcher, and the supervisor. Tree structure lives in the manager,
authority stays in the kernel; bus drivers report what they see; drivers are
restarted through the lifecycle vocabulary — the plan that turns
+11 -4
View File
@@ -4,8 +4,13 @@
with its deadline, supervised spawn, restart with backoff, and the crash-loop
cap are in — usb-xhci-bus is the first conforming driver, and the
`driver-restart` scenario proves fault → backoff → re-claim → cap end to end.
Tree reports (M18.2) and the application surface (M18.3) remain design. The
primitives underneath are real ([process-management.md](process-management.md):
Tree reports are built too (M18.2, 2026-07-13): the xHCI driver scans its
root-hub ports and reports each connected device (`child_added`); the manager
mirrors them and prunes a dead reporter's children, and the `usb-report`
scenario proves report → prune → respawn → re-report. The application surface is built (M18.3, 2026-07-13):
`enumerate` and `subscribe` over IPC, with `device-list` as the first client —
the manager is now the one answer to "what devices exist" for applications.
The primitives underneath are real ([process-management.md](process-management.md):
spawn/supervise/kill/exit-notification; [driver-model.md](driver-model.md): the device
table as a capability system; [drivers.md](drivers.md): claim/map/IRQ), and the first
per-device driver spawn works (the device manager matches the xHCI controller by PCI
@@ -131,8 +136,10 @@ published exit events, signals + `runtime.process`). On top of those:
the mouse and keyboard QEMU already hangs off it.
7. **App surface**: `enumerate`/`subscribe` over IPC; `device_enumerate` retreats
to a manager-internal seam.
8. **Discovery migration**: pci-bus driver first, acpi service second, kernel scan
retired last. (AML-in-user-space is its own track.)
8. **Discovery migration** — DONE (M19–M20, 2026-07-13): pci-bus driver (M19)
then the acpi service (M20) moved enumeration to ring 3; the kernel seeds
only the host bridge and the acpi-tables node. See
[m19-m20-plan.md](m19-m20-plan.md).
## Settled questions (2026-07-12)
+24
View File
@@ -167,3 +167,27 @@ free; discovery on x86 is partly about *finding* what ARM just tells you.
- [ipc.md](ipc.md) — the channels that interrupts-as-messages and the device manager
will ride on.
- [vision.md](vision.md) — why drivers belong in isolated user space at all.
## Update (M19.3, 2026-07-13): PCI enumeration left the kernel
The kernel now seeds only the `pci_host_bridge` node (ECAM window, MMIO
apertures derived from the memory map's holes, bus range, and the 16-bit I/O
window). The per-function walk moved to the ring-3 `pci-bus` driver
([device-manager.md](device-manager.md)): it claims the bridge, repeats the
ECAM scan through its mmio grant, and `device_register`s what it finds, which
the device manager mirrors and matches. The ACPI namespace walk follows in M20;
the static tables (MADT, HPET, MCFG, FADT + `\\_S5`) stay kernel-side.
## Update (M20.3, 2026-07-13): ACPI enumeration left the kernel too
The kernel no longer folds the AML namespace's Device objects into the device
tree. It still parses the *static* tables (MADT for SMP, HPET for the tick, MCFG
for the host bridge, FADT) and still builds the AML namespace — but only to read
the `\\_S5` sleep type for poweroff. Device discovery is the ring-3 **acpi
service** ([device-manager.md](device-manager.md)): it claims the `acpi-tables`
node the kernel publishes (the AML blobs, a broad io_port grant, the SCI),
re-parses the same blobs with the shared AML module, evaluates `_STA`/`_CRS`,
and registers + reports each `_HID` device — the device manager matches drivers
(ps2-bus) from those reports. With M19's pci-bus driver, discovery now runs
entirely in user space; the kernel seeds only the host bridge and the
acpi-tables node.
+20
View File
@@ -362,3 +362,23 @@ the first DMA driver to protect and test against) and these smaller items:
- **Interrupt priority / threaded IRQ latency.** `notifyFromIsr` enqueues the woken
driver but doesn't preempt (`wakeLocked` deliberately leaves that to the caller), so
a woken driver waits for the next scheduling point.
## The driver contract (M17–M18)
Claiming and mapping is half of being a danos driver; the other half is the
**lifecycle and protocol contract**, and the runtime makes it nearly free:
- Build on `runtime.service.run` — one replyWait loop folding protocol
requests, signals, and notifications into callbacks. The harness answers the
universal zero-length ping and turns `terminate` into a clean exit for you
([process-lifecycle.md](process-lifecycle.md)).
- A driver spawned with an assignment (its device id as argv[1]) sends the
versioned `hello` to the device manager inside the deadline, and a **bus**
driver reports what it discovers with `child_added`
([device-manager.md](device-manager.md); usb-xhci-bus is the reference
implementation).
- Crash freely — that is the design. The kernel releases your claims, IRQ
bindings, and MSI vectors at death; the manager reads your exit reason,
prunes what you reported, restarts you with backoff, and your fresh instance
re-claims and re-reports. Never depend on your own cleanup running
(iron rule 1).
+17 -4
View File
@@ -1,5 +1,9 @@
# M17–M18 execution plan: process lifecycle + device manager
**Archived — completed 2026-07-13** (every item checked; suite ended 54/54).
Kept as the record of how M17–M18 landed; the successor is
[m19-m20-plan.md](m19-m20-plan.md).
The operational plan for building [process-lifecycle.md](process-lifecycle.md)
(M17) and [device-manager.md](device-manager.md) increments 5–7 (M18). Design is
settled in those documents; this file is the build order — one phase at a time,
@@ -52,10 +56,19 @@ only when its definition of green holds.
crash-loop cap; usb-xhci-bus first conforming driver; crash-test fixture
re-proving claim release each respawn; `driver-restart` scenario;
maximum_tasks 16→32 — the sweep was overflowing the pool; suite 52/52)
- [ ] **merge** `feat/device-manager` → main, push
- [ ] **M18.2** — xHCI port scan + tree reports (branch `feat/usb-xhci-bus`)
- [ ] **M18.3** — app surface: enumerate/subscribe + device-list
- [ ] **merge** `feat/usb-xhci-bus` → main, push — **loop ends here**
- [x] **merge** `feat/device-manager` → main, push (merged 2026-07-13)
- [x] **M18.2** — xHCI port scan + tree reports (child_added/child_removed in
the protocol; the manager's child mirror with death-pruning; xHCI maps the
register BAR — resource 0 is ECAM — reads CAPLENGTH/HCSPARAMS1, scans
PORTSC, reports connected ports with speed-class identity; `usb-report`
scenario proves report → prune → respawn → re-report; suite 53/53)
- [x] **M18.3** — app surface: enumerate/subscribe over IPC (subscriber
endpoint rides as the call's capability; events are the same structs the
buses send); device-list first client; protocol capped at the kernel's
IPC MESSAGE_MAXIMUM (256); the startUserTask debug print removed — it
sheared concurrent serial lines and was the scenario-flake root cause;
`device-list` scenario; suite 54/54)
- [x] **merge** `feat/usb-xhci-bus` → main, push (merged 2026-07-13) — **plan complete**
---
+212
View File
@@ -0,0 +1,212 @@
# M19–M20 execution plan: discovery migration
The operational plan for [device-manager.md](device-manager.md)'s increment 8:
discovery leaves the kernel — a **pci-bus driver** (M19) and an **acpi service**
(M20), with the kernel's device enumeration retired behind them. Same rules as
[m17-m18-plan.md](m17-m18-plan.md): one phase at a time, each green before the
next; this file is the build order and the checklist.
**Definition of green, every phase:** `zig build` clean, `zig build test` clean,
`python3 test/qemu_test.py` passes (existing scenarios plus the phase's new
one), and the relevant design doc updated. Commit per green phase (no co-author
trailers). The full suite is the regression net — the existing
`driver-restart` / `usb-report` / `device-list` / `input` scenarios must stay
green *through* the migration, which is the whole point: the system must not be
able to tell who enumerated it.
**Workflow:** dedicated worktree; branches off `main` — `feat/pci-bus`
(M19.0–19.3), `feat/acpi-service` (M20.1–20.3); auto-merge to main when a
branch is green; keep branches; push everything.
## Settled decisions (2026-07-13 — veto before the loop starts)
1. **What "retiring the kernel scan" means.** The kernel keeps, forever, the
parses it needs before user space exists: RSDP/XSDT location, MADT (SMP),
the HPET table (the tick), FADT + the AML `\_S5` evaluation (poweroff — the
power tests prove it), and MCFG (the host bridge node). What retires is
**device enumeration**: the ECAM function walk (M19.3) and the DSDT/SSDT
namespace walk that builds device nodes (M20.3). The AML module stays a
shared build module compiled into both the kernel (for `\_S5`) and the acpi
service (for everything else) — same source, two builds, no fork.
2. **Bridge apertures come from the firmware memory map, not AML.** Registered
PCI functions carry BAR resources, and containment demands the bridge own
windows that cover them. The apertures are derived kernel-side from the
boot memory map's MMIO holes (regions that are neither RAM nor tables) —
mechanical, AML-free, and available at boot regardless of what later moved
to user space. (The bridge today carries only ECAM + bus range; this is the
prerequisite M19.0 exists for.)
3. **`device_register` becomes idempotent on exact match.** A re-registration
with identical (parent, class, resources) returns the existing id instead
of appending. The kernel table has no unregister, so without this a
restarted registering bus would duplicate its children on every respawn —
idempotence makes restart-and-re-report safe for every future bus, not just
PCI.
4. **The manager matches from reports.** `ChildAdded` gains a `device_id`
field (the kernel-registered id, `no_device` for unregistered leaves like
USB ports). After the M19.3 flip, PCI driver matching keys off reported
identity (the class triple) instead of the manager's boot-time snapshot —
the snapshot match remains only for what the kernel still seeds. One flip
phase changes both sides at once so no device is ever matched twice.
5. **The acpi service's authority is one node.** The kernel publishes an
`acpi-tables` device: memory resources covering the table blobs plus a
broad `io_port` resource — the documented trust grant to exactly one
process (AML OperationRegions reach EC/PM ports; the claim-gated
io_read/io_write calls already exist). The service claims it, maps the
tables, and runs the shared AML module in ring 3 behind a `Hal` backed by
`mmio_map` + `io_read`/`io_write`.
6. **Both new processes are protocol drivers** under the manager: hello,
supervision, restart with backoff — all inherited from M18.1 for free.
Registration idempotence (decision 3) is what makes their restarts sound.
7. **Firmware neutrality is the contract** (2026-07-13). The generic layer is
everything at and above the device-manager protocol — descriptors,
containment, reports, matching, supervision — and none of it may become
x86-specific. Discovery is one swappable process per firmware: the acpi
service on x86; an **fdt service** on the Raspberry Pis (claims a
`devicetree-blob` node, reports children from the flattened device tree —
pure data, no bytecode, no port grant, strictly simpler than ACPI). The
manager owns the tree as *data* and touches no hardware, ever — AML runs in
a crashable, supervised discoverer precisely so a firmware-bytecode fault
can never take down the supervisor. Two consequences recorded now:
`DeviceDescriptor`'s 8-byte `hid` cannot hold an FDT `compatible` string
("brcm,bcm2835-aux-uart") — identity widens before the fdt service exists;
and cross-firmware surfaces are named by **domain, not firmware** (M21
defines a *power* protocol, not an "ACPI events" protocol — PSCI/mailbox
sources feed the same subscribers on ARM). **Landed early (2026-07-13):**
both services exist as placeholders (system/services/acpi, system/services/
fdt) and the build's `-Ddiscovery=acpi|fdt` option fills the ramdisk's
neutral `discovery` slot — the manager will spawn "discovery" by that name
in M20.3 and never learn which firmware it is on.
## Status
- [x] **M19.0** — prerequisites (bridge apertures from the memory map's
*gaps* — the single-hole rule died on OVMF's flash at the top of 4 GiB,
caught by the new every-BAR-contained assert in `discovery`; idempotent
`device_register` proven in `bus`; `ChildAdded.device_id`;
m17-m18-plan.md archived; suite 54/54).
- [x] **M19.1** — pci-bus driver, scan only (claims the bridge, maps ECAM
through its grant, brute-force walk with the multifunction rule; the
manager matches pci_host_bridge → pci-bus per device with the full
protocol contract; `pci-scan` builds its expected marker from the
kernel's own count — equivalence on the first run; suite 55/55).
- [x] **M19.2** — register + report (BAR probe mirrored byte-for-byte from the
kernel's addBars so dedupe returns the kernel's node ids during
coexistence; the bridge gained the io_port aperture I/O BARs need;
reports carry the registered device_id; pci-scan drills a forced restart
and asserts the PCI node count never grows — plus harness hardening: a
failing case now preserves its serial as <case>-failed-serial.log, and
the heavy scenarios run at 150s; suite 55/55).
- [x] **M19.3** — the flip: kernel `enumeratePci`/`addBars`/`PciHeader` all
deleted (bridge node stays); manager matches PCI drivers from reported
identity, deduped by registered id. Surfaced and fixed a real SMP race the
flip created — ring-3 device_register made the broker table concurrent, so
mmio_map's lock-free read intermittently tore hpet's resource length
(user fault) and overflowed `r.len-1` into a kernel panic; now the broker
read is under the big lock and the arithmetic is guarded, and pci-bus
skips size-0 BARs. discovery.md updated; suite 55/55 (driver-restart
hammered 6×).
- [x] **merge** `feat/pci-bus` → main, push (merged 2026-07-13).
- [x] **M20.1** — acpi service, parse only: the AML interpreter is now a build
module compiled into both kernel and service; the kernel publishes the
`acpi-tables` node (AML blobs as memory resources, the broad io_port grant,
the SCI); the service claims it, maps the blobs, runs the shared parser in
ring 3, and self-verifies its Device count against the kernel's (34 = 34,
deterministic via argv, no log-scraping); the manager spawns `discovery`
at startup. Parse-only touches no hardware. Suite 56/56.
- [x] **M20.2** — register + report: the service evaluates `_STA`/`_CRS` in
ring 3 (interpreter Hal = port I/O over the claimed node; a scratch page
backs SystemMemory maps so a stray region can't fault it) and registers +
reports each present `_HID` device under `acpi-tables`. Containment: the
broker's irq check became range-based (len-1 == the old equality) so the
node's broad irq window covers children's legacy lines; io ports fall in
the broad io grant. ChildAdded gained `hid`. Matching stays off. The
`acpi-report` scenario asserts the PS/2 keyboard (3 resources) and mouse
(1 resource) among the reports. Suite 57/57.
- [x] **M20.3** — the flip: the kernel's `wireAcpiDevices` call is gone (the
device-building helpers are retained-but-dead pending a focused sweep,
spawned as a task; static tables + `\_S5` + the acpi-tables node stay).
The manager matches ps2-bus from ACPI `_HID` reports; the service
registers all devices before reporting any (no keyboard-before-mouse
race). The `acpi-ps2` scenario proves report → spawn → ps2-bus attaches
its keyboard; `ioport` retargeted to the acpi-tables I/O window (the
kernel-built PS/2 node is gone). Suite 58/58.
- [ ] **merge** `feat/acpi-service` → main, push — **loop ends here**.
---
## Phase notes
**M19.0 apertures:** the boot memory map already crosses the handoff
([boot-handoff]), but discovery never sees it today — expect a small
pass-through (kernel init hands the map to the platform layer) before the
holes computation, which belongs where the bridge node is built
(`parseMcfg`). Sanity-check on QEMU q35: the xHCI BAR (`0xc0000000`-region
values seen in the M18 logs) must land inside a derived aperture, asserted in
the kernel unit test.
**M19.1 scanning without owning config access twice:** the driver reads config
space through its ECAM mmio_map grant of the *bridge* window — the same bytes
the kernel walk read. Vendor-id `0xFFFF` skip, header-type multifunction rule,
no bridge recursion (matches the kernel's current single-segment walk).
**M19.2 BAR sizing:** the classic size probe (write all-ones, read mask,
restore) is deferred — the BARs' current programmed values and types are
enough for containment-checked registration at bring-up; sizing lands with the
first driver that needs to *move* a BAR. Log what is registered so the
scenario can assert it.
**M19.3 what the manager still seeds from the snapshot:** everything the
kernel still enumerates (timers, ACPI nodes until M20.3). The PCI arm of
`pciDriverFor` switches source; `driverFor` doesn't move until M20.3.
**M20.1 spawn and identity (pre-settled 2026-07-13):** the manager spawns
`discovery` by its neutral ramdisk name at startup, as an ordinary protocol
driver (hello, supervision) — from M20.1 on, on every boot. For reporting ACPI
devices, `ChildAdded` gains `hid: [8]u8` (EISA ids fit; zero = none):
firmware *string* identity travels beside the numeric `identity` field until
the FDT-driven widening replaces both (decision 7).
**M20.1 Hal in ring 3:** `mapMmio` → `device.mmioMap` over the claimed
acpi-tables node (plus a table-offset map for blobs); `pioRead`/`pioWrite` →
`device.ioRead`/`ioWrite` against its io_port resource. The interpreter cannot
tell it moved — that is the assertion of `acpi-parse`.
**M20.2 containment for `_CRS`:** io ports fall inside the node's broad
io_port resource; MMIO windows (HPET, LAPIC ranges some firmwares list) fall
inside the memory-map holes added to the node in M20.1. Anything that doesn't
fit is logged and skipped, loudly — bring-up honesty over silent drops.
**M20.3 ps2 ordering:** ps2-bus binds nodes the acpi service now reports, so
its spawn moves behind the report (the manager's matching handles this once
the source flips); the `input` scenario proves the keyboard still types.
**Explicitly out of scope:** PCI bridge recursion (single segment, flat bus
walk stays); BAR reprogramming/sizing; disk/PCIe hotplug; interrupt routing
changes (`_PRT` stays wherever it is today); the USB descriptor track;
multi-segment ECAM; per-device power states (D-states, `_PSx`/`_PRx`,
suspend/resume — a future *lifecycle-vocabulary* extension, since "suspend"
has the shape of a signal every driver must answer, and it has no consumer
until laptop sleep); CPU P/C-states.
## M21 preview — ACPI events + system power (planned next, not in this loop)
The acpi service grows the event side (settled direction 2026-07-13; detailed
phases when M20 lands):
- **21.1 SCI + fixed events**: irq_bind the SCI (the resource M20.1 already
records), read/clear PM1 status, publish the power-button event to
subscribers (the same pub/sub shape the manager uses).
- **21.2 GPE + Notify**: Notify dispatch in the shared AML interpreter, GPE
block handling, `Notify(device, code)` published per reported node. The
acpi service is a **bus** here: battery (PNP0C0A), AC (ACPI0003), and lid
(PNP0C0D) nodes are reported children; small class drivers bind them and
speak an evaluate/subscribe protocol to the service — the xHCI split,
repeated. The embedded controller (`_Qxx` queries) rides this phase;
QEMU emulates no battery/EC, so those paths are interface-complete and
validated on real hardware (the laptop is the win condition), while the
plumbing is proven by the power button.
- **21.3 the capstone**: QEMU `system_powerdown` → acpi service event → init
runs the M17 stop sequence over its children → kernel `\_S5` — orderly
shutdown as the scenario that proves lifecycle + events compose. (The
harness grows a QMP poke to inject the event.)
+12 -5
View File
@@ -1,10 +1,17 @@
# Resilience: fault isolation and live restart
Steps 1–2 of the ordering below are **built**: user-mode isolation, and fault →
kill the process → keep the core (`onException` in `system/kernel/kernel.zig`; the
`fault-recovery` test proves a crashing ring-3 process dies alone while the system
keeps running). The supervisor notification and restart policy (steps 3+) are
still design. This is the property danos is really chasing:
Steps 1–4 of the ordering below are **built** (M17–M18, 2026-07-13): user-mode
isolation; fault → kill the process → keep the core (`onException`; the
`fault-recovery` test); the supervisor notification **with exit reasons**
([process-lifecycle.md](process-lifecycle.md) — clean exit, fault class, or
killed, recorded before the notice posts); and the **restart policy itself**
([device-manager.md](device-manager.md)): the device manager supervises every
driver, restarts crashes with backoff, caps crash loops, and re-claims work
because the kernel releases a dead process's claims. The `driver-restart` and
`usb-report` scenarios prove kill → release → respawn → re-claim → re-report
end to end. What remains of this document's ladder is scope, not mechanism:
more of the system moved into restartable processes (the discovery migration,
[m19-m20-plan.md](m19-m20-plan.md), is the next rung). This is the property danos is really chasing:
**if a part of the OS breaks, isolate it, and re-initialise it — without rebooting.**
A crashed driver gets restarted; a wedged service gets killed and brought back. It's
the reason the [microkernel](vision.md) shape was chosen, and it's a *separate* goal
+5 -3
View File
@@ -22,8 +22,10 @@ pub const Callbacks = struct {
/// Return false to abort startup (the process exits).
init: ?*const fn (endpoint: ipc.Handle) bool = null,
/// One protocol request from `sender` (a task id): write the reply into
/// `reply`, return its length. The zero-length ping never reaches this.
on_message: *const fn (message: []const u8, reply: []u8, sender: u32) usize,
/// `reply`, return its length. `capability` is the handle the request
/// carried, if any (M13 cap passing — how a subscriber hands over its
/// endpoint). The zero-length ping never reaches this.
on_message: *const fn (message: []const u8, reply: []u8, sender: u32, capability: ?ipc.Handle) usize,
/// A notification that is not a signal — a subscribed exit event, a bound
/// IRQ, a timer landing. The raw badge; decode with the ipc helpers.
on_notification: ?*const fn (badge: u64) void = null,
@@ -76,6 +78,6 @@ pub fn run(comptime maximum_message: usize, callbacks: Callbacks) void {
reply_len = 0; // the universal ping: a zero-length reply, from the harness
continue;
}
reply_len = callbacks.on_message(receive[0..got.len], &reply_buffer, got.senderTaskId());
reply_len = callbacks.on_message(receive[0..got.len], &reply_buffer, got.senderTaskId(), got.cap);
}
}
+126 -134
View File
@@ -41,6 +41,9 @@ pub const RegisterAccess = struct {
/// Everything the power subsystem needs, extracted from the FADT and the AML
/// sleep packages during discovery. Populated by `discover`, read by `power`.
pub const PowerInformation = struct {
/// The System Control Interrupt's GSI (FADT SCI_INT) — the line ACPI events
/// (power button, GPEs) arrive on. Published to the acpi service for M21.
sci_interrupt: u16 = 0,
/// The SMM command port and the value that switches the platform into ACPI mode.
smi_cmd: u16 = 0,
acpi_enable: u8 = 0,
@@ -360,32 +363,14 @@ const Hpet = extern struct {
page_protection: u8,
};
// --- PCI configuration-space header (first 64 bytes, common fields) ---------
const PciHeader = extern struct {
vendor_id: u16 align(1),
device_id: u16 align(1),
command: u16 align(1),
status: u16 align(1),
revision_id: u8,
prog_if: u8,
subclass: u8,
class_code: u8,
cache_line_size: u8,
latency_timer: u8,
/// bit 7 set => multi-function device.
header_type: u8,
bist: u8,
// 0x10 onward (BARs, etc.) depends on header_type; read separately.
};
// --- Entry point ------------------------------------------------------------
/// Discover hardware from the ACPI tables rooted at `rsdp_physical` and populate
/// `device_tree`. `hal` provides MMIO mapping (for PCIe ECAM) and port I/O. Also parses the
/// FADT and the AML sleep-state (`_Sx`) packages into `power_information` for the power service.
pub fn discover(rsdp_physical: u64, device_tree: *DeviceTree, hal: Hal) !void {
pub fn discover(rsdp_physical: u64, memory_regions: []const boot_handoff.MemoryRegion, device_tree: *DeviceTree, hal: Hal) !void {
if (rsdp_physical == 0) return error.NoRsdp;
boot_memory_regions = memory_regions;
// Start clean so a re-run doesn't accumulate stale state.
power_information = .{};
@@ -420,12 +405,57 @@ pub fn discover(rsdp_physical: u64, device_tree: *DeviceTree, hal: Hal) !void {
aml_stats = .{ .nodes = namespace.?.nodeCount(), .consumed = pr.consumed, .total = pr.total };
power_information.s5 = aml.sleepState(&namespace.?, 5);
power_information.s3 = aml.sleepState(&namespace.?, 3);
// Fold the namespace's Device objects into the generic tree.
wireAcpiDevices(device_tree, &namespace.?, hal) catch {};
// The namespace's Device objects are no longer folded into the kernel
// tree (M20.3): the ring-3 acpi service claims the acpi-tables node
// (published below), re-parses the same blobs, and registers + reports
// the _HID devices itself. The kernel keeps the namespace only for the
// \_S5 sleep type above. The device-building helpers below
// (wireAcpiDevices and friends) are retained but unreferenced — a
// focused dead-code sweep follows the migration.
} else |_| {
// AML parse failed (e.g. out of memory); power stays best-effort with
// whatever the FADT alone provided.
}
// Publish the acpi-tables node (docs/m19-m20-plan.md M20): the AML blobs as
// memory resources for the acpi service to map and parse in ring 3, a broad
// io_port grant for the OperationRegion access its interpreter needs, and
// the SCI for the events track (M21). Exactly one node, one trusted
// claimant. Kept even when the kernel-side device building (above) retires
// in M20.3 — the kernel still owns the *static* tables and \_S5.
publishAcpiTablesNode(device_tree) catch {};
}
/// Build the acpi-tables node (see the call site in discover). Best-effort: a
/// failure here leaves the kernel-seeded tree working, only the ring-3 service
/// finds nothing to claim.
fn publishAcpiTablesNode(device_tree: *DeviceTree) !void {
const node = try device_tree.addChild(device_tree.root, .acpi_tables, "acpi-tables");
// One memory resource per AML block — page-aligned base down, length padded
// up to cover the bytecode, so mmio_map hands the service a pointer into it.
var i: usize = 0;
while (i < aml_block_count and i < device_model.maximum_resources - 2) : (i += 1) {
// mmio_map preserves the sub-page offset, so the service maps this and
// gets a pointer straight to the bytecode.
_ = node.addResource(.memory, aml_block_physical[i], aml_block_len[i]);
}
// The broad I/O grant: OperationRegions name whatever ports the firmware
// chose (EC, PM1, GPE, SMBus); which ports cannot be known before the AML
// that names them is parsed, so the grant is the whole space — the honest
// trust boundary of docs/m19-m20-plan.md decision 5.
_ = node.addResource(.io_port, 0, 1 << 16);
// A broad interrupt window: ACPI _CRS names legacy ISA IRQs (the PS/2 lines
// 1 and 12, the RTC, …), and the service registers those devices under this
// node, so it must own a superset. The range [0, 256) covers every GSI; the
// SCI (recorded first, len 1) stays distinct so M21 can pick it out.
if (power_information.sci_interrupt != 0) _ = node.addResource(.irq, power_information.sci_interrupt, 1);
_ = node.addResource(.irq, 0, 256);
}
/// The number of Device objects in the namespace built during discovery, or 0.
pub fn amlDeviceCount() usize {
if (namespace) |*ns| return aml.deviceCount(ns);
return 0;
}
/// Walk the RSDT (Entry = u32) or XSDT (Entry = u64): validate it, then dispatch
@@ -451,7 +481,7 @@ fn handleTable(device_tree: *DeviceTree, hal: Hal, sdt_physical: u64) !void {
if (std.mem.eql(u8, &sig, &APIC)) {
try parseMadt(device_tree, header);
} else if (std.mem.eql(u8, &sig, &MCFG)) {
try parseMcfg(device_tree, hal, header);
try parseMcfg(device_tree, header);
} else if (std.mem.eql(u8, &sig, &HPET)) {
try parseHpet(device_tree, hal, header);
} else if (std.mem.eql(u8, &sig, &FACP)) {
@@ -536,7 +566,7 @@ fn parseMadt(device_tree: *DeviceTree, header: *const SystemDescriptorTableHeade
}
/// MCFG -> a pci_host_bridge per ECAM segment, then a PCI enumeration underneath.
fn parseMcfg(device_tree: *DeviceTree, hal: Hal, header: *const SystemDescriptorTableHeader) !void {
fn parseMcfg(device_tree: *DeviceTree, header: *const SystemDescriptorTableHeader) !void {
const total: usize = header.length;
const base: [*]const u8 = @ptrCast(header);
@@ -551,109 +581,85 @@ fn parseMcfg(device_tree: *DeviceTree, hal: Hal, header: *const SystemDescriptor
// ECAM window: 1 MiB of configuration space per bus.
_ = bridge.addResource(.memory, alloc.base_address, bus_count << 20);
_ = bridge.addResource(.bus_range, alloc.start_bus, bus_count);
addBridgeApertures(bridge);
// The bridge decodes the whole 16-bit I/O space toward its bus — the
// window functions' I/O BARs must register-contain within (M19.2).
_ = bridge.addResource(.io_port, 0, 1 << 16);
try enumeratePci(device_tree, bridge, hal, alloc.*);
// The function walk itself retired to ring 3 (M19.3): the pci-bus
// driver claims this bridge, repeats the scan through its ECAM grant,
// and device_registers what it finds — the kernel seeds only the
// bridge. The scan's equivalence was proven before the hand-off
// (pci-scan), and the walk's history is in git if archaeology calls.
}
}
/// Brute-force scan the ECAM window's bus range for present PCI functions. No
/// bridge recursion yet: on the ECAM path the host bridge decodes every bus in
/// the window, so scanning the declared range finds everything QEMU exposes.
fn enumeratePci(
device_tree: *DeviceTree,
bridge: *device_model.Device,
hal: Hal,
alloc: McfgAllocation,
) !void {
var bus: u16 = alloc.start_bus;
while (bus <= alloc.end_bus) : (bus += 1) {
var device: u8 = 0;
while (device < 32) : (device += 1) {
const h0: *align(1) const PciHeader = @ptrCast(pciConfigurationPtr(alloc, hal, @intCast(bus), device, 0));
if (h0.vendor_id == 0xFFFF) continue; // no function 0 => slot empty
/// The boot memory map, stored at discover() entry for the aperture derivation
/// below (and, in M20, for the acpi-tables node's containment windows).
var boot_memory_regions: []const boot_handoff.MemoryRegion = &.{};
const funcs: u8 = if (h0.header_type & 0x80 != 0) 8 else 1;
var function: u8 = 0;
while (function < funcs) : (function += 1) {
const configuration = pciConfigurationPtr(alloc, hal, @intCast(bus), device, function);
const h: *align(1) const PciHeader = @ptrCast(configuration);
if (h.vendor_id == 0xFFFF) continue;
var nb: [24]u8 = undefined;
const nm = std.fmt.bufPrint(&nb, "{s}:{x:0>2}:{x:0>2}.{d}", .{
bridge.name(), bus, device, function,
}) catch "pcidev";
const node = try device_tree.addChild(bridge, .pci_device, nm);
// Resource 0 is the function's own 4 KiB ECAM configuration space. A
// claimed PCI driver mmio_maps this to reach its command register,
// BARs, and — the point — its capability list (MSI/MSI-X, PCIe
// extended caps), without any new syscall. Physical address per the
// ECAM formula (same as pciConfigurationPtr).
const config_physical = alloc.base_address +
(@as(u64, @as(u8, @intCast(bus)) - alloc.start_bus) << 20) +
(@as(u64, device) << 15) + (@as(u64, function) << 12);
_ = node.addResource(.memory, config_physical, abi.page_size);
node.ids.pci_vendor = h.vendor_id;
node.ids.pci_device = h.device_id;
node.ids.pci_class = (@as(u24, h.class_code) << 16) |
(@as(u24, h.subclass) << 8) | h.prog_if;
node.ids.pci_bdf = (@as(u16, @intCast(bus)) << 8) | (@as(u16, device) << 3) | function;
// BARs only exist in header type 0 (normal devices), not bridges.
if (h.header_type & 0x7F == 0) addBars(node, configuration);
/// The bridge's MMIO apertures, derived from the boot memory map's holes
/// (docs/m19-m20-plan.md decision 2): registered PCI functions carry BAR
/// resources, and `device_register` containment demands the bridge own windows
/// that cover them. Everything the firmware described is "not hole"; the low
/// aperture runs from the end of the described space below 4 GiB up to the
/// I/O-APIC region, the high one from 4 GiB (or the end of RAM above it) to
/// the 46-bit line. Coarse, mechanical, and AML-free — available at boot no
/// matter what later moved to user space.
fn addBridgeApertures(bridge: *device_model.Device) void {
// Below 4 GiB the described regions are sparse (RAM low, firmware flash
// and tables high), so the holes are the *gaps between* them — a single
// "after the last region" rule dies on OVMF's flash at the very top.
// Sort-merge the described ranges, then keep the three largest gaps
// (resource slots are bounded at 8 per device; ECAM + bus range + 3 + the
// high aperture fits). Above 4 GiB one aperture runs from the end of the
// described space to the 46-bit line.
const Range = struct { base: u64, end: u64 };
var below: [64]Range = undefined;
var below_count: usize = 0;
var high_end: u64 = 1 << 32;
for (boot_memory_regions) |region| {
const end = region.base + region.pages * 4096;
// Above 4 GiB only *usable RAM* blocks the aperture: OVMF describes
// its own 64-bit PCI window as a reserved region and then programs
// BARs inside it — honoring reserved there would exclude the very
// space BARs live in. Below 4 GiB every described region blocks (the
// kernel image, the tables, the ramdisk all live there). Bring-up
// trust: only the bridge's claimant can register into the aperture.
if (region.kind == .usable and end > high_end) high_end = end;
if (region.base >= (1 << 32) or below_count == below.len) continue;
below[below_count] = .{ .base = region.base, .end = @min(end, 1 << 32) };
below_count += 1;
}
// Insertion sort by base (the map is small and this runs once at boot).
for (1..below_count) |i| {
const key = below[i];
var j = i;
while (j > 0 and below[j - 1].base > key.base) : (j -= 1) below[j] = below[j - 1];
below[j] = key;
}
// Walk the sorted ranges, collecting inter-region gaps of at least 1 MiB.
var gaps: [3]Range = .{Range{ .base = 0, .end = 0 }} ** 3;
var cursor: u64 = 0;
var index: usize = 0;
while (index <= below_count) : (index += 1) {
const gap_end = if (index == below_count) (1 << 32) else below[index].base;
if (gap_end > cursor and gap_end - cursor >= (1 << 20)) {
// Keep the three largest, replacing the smallest kept so far.
var smallest: usize = 0;
for (gaps, 0..) |gap, gi| {
if (gap.end - gap.base < gaps[smallest].end - gaps[smallest].base) smallest = gi;
}
if (gap_end - cursor > gaps[smallest].end - gaps[smallest].base) {
gaps[smallest] = .{ .base = cursor, .end = gap_end };
}
}
if (index < below_count and below[index].end > cursor) cursor = below[index].end;
}
}
/// Record and size the memory/IO windows named by a device's Base Address
/// Registers. Sizing is the standard probe: disable decode, write all-ones, read
/// back the writable (address) bits, restore. `size = ~mask + 1`.
fn addBars(node: *device_model.Device, configuration: [*]align(1) u8) void {
// Stop the device decoding its BARs while we transiently write all-ones.
const command = rd(u16, configuration, 0x04);
wr(u16, configuration, 0x04, command & ~@as(u16, 0b11));
var i: usize = 0;
while (i < 6) : (i += 1) {
const off = 0x10 + i * 4;
const orig = rd(u32, configuration, off);
if (orig == 0) continue;
if (orig & 1 != 0) {
// I/O-space BAR (16-bit address space on x86).
wr(u32, configuration, off, 0xFFFF_FFFF);
const readback = rd(u32, configuration, off);
wr(u32, configuration, off, orig);
const mask = readback & 0xFFFF_FFFC;
const size: u32 = if (mask == 0) 0 else (~mask +% 1) & 0xFFFF;
_ = node.addResource(.io_port, orig & 0xFFFF_FFFC, size);
} else if ((orig >> 1) & 0x3 == 2) {
// 64-bit memory BAR: this BAR pair spans two configuration slots.
const orig_hi = rd(u32, configuration, off + 4);
wr(u32, configuration, off, 0xFFFF_FFFF);
wr(u32, configuration, off + 4, 0xFFFF_FFFF);
const lo = rd(u32, configuration, off);
const hi = rd(u32, configuration, off + 4);
wr(u32, configuration, off, orig);
wr(u32, configuration, off + 4, orig_hi);
const readback = (@as(u64, hi) << 32) | (lo & 0xFFFF_FFF0);
const size: u64 = if (readback == 0) 0 else ~readback +% 1;
const address = (@as(u64, orig_hi) << 32) | (orig & 0xFFFF_FFF0);
_ = node.addResource(.memory, address, size);
i += 1; // consumed the high half
} else {
// 32-bit memory BAR.
wr(u32, configuration, off, 0xFFFF_FFFF);
const readback = rd(u32, configuration, off);
wr(u32, configuration, off, orig);
const mask = readback & 0xFFFF_FFF0;
const size: u32 = if (mask == 0) 0 else ~mask +% 1;
_ = node.addResource(.memory, orig & 0xFFFF_FFF0, size);
}
for (gaps) |gap| {
if (gap.end > gap.base) _ = bridge.addResource(.memory, gap.base, gap.end - gap.base);
}
wr(u16, configuration, 0x04, command); // restore decode
_ = bridge.addResource(.memory, high_end, (@as(u64, 1) << 46) - high_end);
}
/// HPET -> a timer node with its register block as an MMIO resource, plus the GSI
@@ -712,6 +718,7 @@ const fadt_pm1a_cnt_blk = 64; // u32 (I/O port)
const fadt_pm1b_cnt_blk = 68; // u32 (I/O port)
const fadt_pm_tmr_blk = 76; // u32 (I/O port) — the PM timer counter
const fadt_pm1_cnt_len = 89; // u8 (bytes)
const fadt_sci_int = 46; // u16 (the SCI's GSI)
const fadt_flags = 112; // u32
const fadt_reset_register = 116; // GAS (12 bytes)
const fadt_reset_value = 128; // u8
@@ -729,6 +736,7 @@ fn parseFadt(header: *const SystemDescriptorTableHeader) void {
const len: usize = header.length;
const pi = &power_information;
pi.sci_interrupt = @truncate(fadt(u16, base, len, fadt_sci_int) orelse 0);
pi.smi_cmd = @truncate(fadt(u32, base, len, fadt_smi_cmd) orelse 0);
pi.acpi_enable = fadt(u8, base, len, fadt_acpi_enable) orelse 0;
pi.acpi_disable = fadt(u8, base, len, fadt_acpi_disable) orelse 0;
@@ -1186,16 +1194,6 @@ fn readCntRegister(base: [*]align(1) const u8, len: usize, xoff: usize, legacy_o
/// The mapped configuration space of one PCI function (its 4 KiB ECAM page). Mapped
/// writable so BAR sizing can probe it; reads and writes both go through here.
fn pciConfigurationPtr(alloc: McfgAllocation, hal: Hal, bus: u8, device: u8, function: u8) [*]align(1) u8 {
const physical = alloc.base_address +
(@as(u64, bus - alloc.start_bus) << 20) +
(@as(u64, device) << 15) +
(@as(u64, function) << 12);
// Map the configuration page (writable, for BAR sizing) and use the virtual
// address the HAL hands back.
return @ptrFromInt(hal.mapMmio(physical, abi.page_size, true));
}
/// Read a little-endian integer at `off` from a (possibly unaligned) byte pointer.
/// x86 is little-endian and native, so an unaligned load suffices.
fn rd(comptime T: type, bytes: [*]align(1) const u8, off: usize) T {
@@ -1203,12 +1201,6 @@ fn rd(comptime T: type, bytes: [*]align(1) const u8, off: usize) T {
return p.*;
}
/// Write a little-endian integer at `off` through a (possibly unaligned) pointer.
fn wr(comptime T: type, bytes: [*]align(1) u8, off: usize, value: T) void {
const p: *align(1) T = @ptrCast(bytes + off);
p.* = value;
}
// --- tests ------------------------------------------------------------------
test "eisaIdToStr decodes a packed EISA id" {
+14
View File
@@ -50,6 +50,20 @@ pub fn parse(allocator: std.mem.Allocator, blocks: []const []const u8) !ParseRes
return .{ .namespace = namespace, .consumed = consumed, .total = total };
}
/// Count the Device objects in a parsed namespace — what the acpi service
/// (docs/m19-m20-plan.md M20) reports, and what the kernel's own parse counts
/// so the two can be checked equal across the ring-3 move.
pub fn deviceCount(namespace: *const Namespace) usize {
return countKind(namespace.root, .device);
}
fn countKind(node: *const Node, kind: NodeKind) usize {
var n: usize = if (node.kind == kind) 1 else 0;
var c = node.first_child;
while (c) |child| : (c = child.next_sibling) n += countKind(child, kind);
return n;
}
/// Look up the `\_S{state}` sleep package in a parsed namespace and return its
/// first two integer elements (SLP_TYP for PM1a / PM1b), or null if absent.
pub fn sleepState(namespace: *Namespace, state: u8) ?SleepType {
+5
View File
@@ -28,6 +28,11 @@ pub const DeviceClass = enum(u32) {
/// A device named in the ACPI namespace (from the DSDT/SSDT), carrying a
/// hardware ID (`_HID`) and, where static, current resource settings (`_CRS`).
acpi_device,
/// The ACPI tables themselves, published as one node for the user-space acpi
/// service (docs/m19-m20-plan.md M20): memory resources over the AML blobs,
/// a broad io_port grant for OperationRegion access, and the SCI interrupt.
/// The one node whose claimant is trusted to run firmware bytecode.
acpi_tables,
unknown,
};
+9 -1
View File
@@ -40,6 +40,13 @@ pub fn platformInformation() PlatformInformation {
}
/// AML parse integrity/diagnostics (namespace node count, bytes consumed).
/// The number of Device objects in the kernel's own AML namespace, or 0 if the
/// parse produced none — the `acpi-parse` test compares the ring-3 service's
/// count against this.
pub fn amlDeviceCount() usize {
return acpi.amlDeviceCount();
}
pub fn amlStats() AmlStats {
return acpi.aml_stats;
}
@@ -72,7 +79,8 @@ pub fn discover(
var device_tree = try DeviceTree.init(allocator);
if (boot_information.acpi_rsdp != 0) {
try acpi.discover(boot_information.acpi_rsdp, &device_tree, hal);
const memory_regions = @as([*]const boot_handoff.MemoryRegion, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.memory_map.regions)))[0..boot_information.memory_map.len];
try acpi.discover(boot_information.acpi_rsdp, memory_regions, &device_tree, hal);
} else {
// No ACPI RSDP. A device-tree boot would parse its blob here; today that
// path is a stub, so this reports the machine described itself no way we
+253
View File
@@ -0,0 +1,253 @@
//! /system/drivers/pci-bus — the PCI bus driver: enumeration moved out of ring 0
//! (docs/m19-m20-plan.md, M19). The device manager matches the `pci_host_bridge`
//! node and spawns one instance per bridge, the bridge's device id as argv[1] —
//! the same per-device contract as usb-xhci-bus.
//!
//! M19.1 (this increment): claim the bridge, map its ECAM window (resource 0;
//! the bus range and the MMIO apertures follow it), walk every
//! bus/device/function config header, and log what the walk finds — ending
//! with "pci-bus: N functions found", which the `pci-scan` scenario compares
//! against the kernel's own enumeration. Registration and reports (M19.2), and
//! the kernel walk's retirement (M19.3), build on this proven-equivalent scan.
const std = @import("std");
const runtime = @import("runtime");
const protocol = runtime.device_manager_protocol;
const device = runtime.device;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
var bridge_id: u64 = protocol.no_device;
var ecam_base: usize = 0;
var ecam_physical: u64 = 0;
var start_bus: u64 = 0;
var bus_count: u64 = 0;
var manager_handle: runtime.ipc.Handle = 0;
/// One aligned 32-bit read from a function's configuration space.
fn configRead(bus: u64, dev: u64, function: u64, offset: u64) u32 {
const address = ecam_base + (((bus - start_bus) << 20) | (dev << 15) | (function << 12) | offset);
const register: *volatile u32 = @ptrFromInt(address);
return register.*;
}
fn configWrite(bus: u64, dev: u64, function: u64, offset: u64, value: u32) void {
const address = ecam_base + (((bus - start_bus) << 20) | (dev << 15) | (function << 12) | offset);
const register: *volatile u32 = @ptrFromInt(address);
register.* = value;
}
fn configRead16(bus: u64, dev: u64, function: u64, offset: u64) u16 {
const word = configRead(bus, dev, function, offset & ~@as(u64, 3));
return @truncate(word >> @intCast((offset & 3) * 8));
}
fn configWrite16(bus: u64, dev: u64, function: u64, offset: u64, value: u16) void {
const aligned = offset & ~@as(u64, 3);
const shift: u5 = @intCast((offset & 3) * 8);
const word = configRead(bus, dev, function, aligned);
const mask = @as(u32, 0xFFFF) << shift;
configWrite(bus, dev, function, aligned, (word & ~mask) | (@as(u32, value) << shift));
}
/// Claim the bridge, map the ECAM, hello the manager, then scan.
fn initialise(endpoint: runtime.ipc.Handle) bool {
_ = endpoint;
if (!device.claim(bridge_id)) {
writeLine("pci-bus: unable to claim bridge device {d}\n", .{bridge_id});
return false;
}
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("pci-bus: out of memory\n");
return false;
};
const total = device.enumerate(buffer);
const descriptor = for (buffer[0..@min(total, buffer.len)]) |d| {
if (d.id == bridge_id) break d;
} else {
writeLine("pci-bus: device {d} not in the device tree\n", .{bridge_id});
return false;
};
// Resource 0 is the ECAM window (1 MiB of config space per bus); the bus
// range rides beside it. The MMIO apertures (M19.0) come after both.
if (descriptor.resource_count < 2 or descriptor.resources[0].kind != @intFromEnum(device.ResourceKind.memory)) {
_ = runtime.system.write("pci-bus: bridge has no ECAM window\n");
return false;
}
const bus_range = for (descriptor.resources[0..@intCast(descriptor.resource_count)]) |resource| {
if (resource.kind == @intFromEnum(device.ResourceKind.bus_range)) break resource;
} else {
_ = runtime.system.write("pci-bus: bridge has no bus range\n");
return false;
};
start_bus = bus_range.start;
bus_count = bus_range.len;
ecam_physical = descriptor.resources[0].start;
ecam_base = device.mmioMap(bridge_id, 0) orelse {
_ = runtime.system.write("pci-bus: ECAM mmio_map failed\n");
return false;
};
// The handshake, then the scan (reports join in M19.2).
var manager: ?runtime.ipc.Handle = null;
var tries: u32 = 0;
while (manager == null and tries < 100) : (tries += 1) {
manager = runtime.ipc.lookup(.device_manager);
if (manager == null) runtime.system.sleep(20);
}
const h = manager orelse {
_ = runtime.system.write("pci-bus: no device manager to hello\n");
return false;
};
const hello = protocol.Hello{ .role = @intFromEnum(protocol.Role.bus), .device_id = bridge_id };
var reply: [protocol.message_maximum]u8 = undefined;
const n = runtime.ipc.call(h, std.mem.asBytes(&hello), &reply) catch {
_ = runtime.system.write("pci-bus: hello call failed\n");
return false;
};
if (n < protocol.reply_size or std.mem.bytesToValue(protocol.HelloReply, reply[0..protocol.reply_size]).status != 0) {
_ = runtime.system.write("pci-bus: hello refused\n");
return false;
}
manager_handle = h;
scan();
return true;
}
/// The brute-force walk the kernel does today, from ring 3: every bus in the
/// range, 32 devices, 8 functions; vendor id FFFFh means nothing decodes there,
/// and only multifunction devices get their functions 1..7 probed.
fn scan() void {
var found: u32 = 0;
var bus: u64 = start_bus;
while (bus < start_bus + bus_count) : (bus += 1) {
var dev: u64 = 0;
while (dev < 32) : (dev += 1) {
const first = configRead(bus, dev, 0, 0);
if (first & 0xFFFF == 0xFFFF) continue;
const multifunction = (configRead(bus, dev, 0, 0x0C) >> 16) & 0x80 != 0;
var function: u64 = 0;
while (function < 8) : (function += 1) {
if (function != 0 and !multifunction) break;
const vendor_device = configRead(bus, dev, function, 0);
if (vendor_device & 0xFFFF == 0xFFFF) continue;
const class_revision = configRead(bus, dev, function, 0x08);
found += 1;
writeLine("pci-bus: {d}:{d}.{d} class 0x{x:0>6}\n", .{ bus, dev, function, class_revision >> 8 });
registerAndReport(bus, dev, function, class_revision >> 8);
}
}
}
writeLine("pci-bus: {d} functions found\n", .{found});
}
/// Register one function under the bridge and report it to the manager. The
/// descriptor mirrors the kernel's own recording byte for byte — config slice
/// as resource 0, then the sized BARs — so during coexistence the idempotent
/// device_register (M19.0) returns the kernel's existing node id rather than
/// growing a duplicate, and the report carries the id drivers already use.
fn registerAndReport(bus: u64, dev: u64, function: u64, class_triple: u32) void {
var descriptor = std.mem.zeroes(device.DeviceDescriptor);
descriptor.class = @intFromEnum(device.DeviceClass.pci_device);
descriptor.pci_class = class_triple;
descriptor.resources[0] = .{
.kind = @intFromEnum(device.ResourceKind.memory),
.start = ecam_physical + (((bus - start_bus) << 20) | (dev << 15) | (function << 12)),
.len = 4096,
};
descriptor.resource_count = 1;
// The standard BAR-sizing probe, exactly as the kernel does it: decode off,
// write all-ones, read the writable mask back, restore. Header type 0 only.
const header_type = (configRead(bus, dev, function, 0x0C) >> 16) & 0x7F;
if (header_type == 0) {
const command = configRead16(bus, dev, function, 0x04);
configWrite16(bus, dev, function, 0x04, command & ~@as(u16, 0b11));
var i: u64 = 0;
while (i < 6) : (i += 1) {
if (descriptor.resource_count >= 8) break;
const off = 0x10 + i * 4;
const original = configRead(bus, dev, function, off);
if (original == 0) continue;
const slot: usize = @intCast(descriptor.resource_count);
if (original & 1 != 0) {
configWrite(bus, dev, function, off, 0xFFFF_FFFF);
const readback = configRead(bus, dev, function, off);
configWrite(bus, dev, function, off, original);
const mask = readback & 0xFFFF_FFFC;
const size: u32 = if (mask == 0) 0 else (~mask +% 1) & 0xFFFF;
if (size == 0) continue; // unimplemented BAR — nothing to register
descriptor.resources[slot] = .{ .kind = @intFromEnum(device.ResourceKind.io_port), .start = original & 0xFFFF_FFFC, .len = size };
descriptor.resource_count += 1;
} else if ((original >> 1) & 0x3 == 2) {
const original_high = configRead(bus, dev, function, off + 4);
configWrite(bus, dev, function, off, 0xFFFF_FFFF);
configWrite(bus, dev, function, off + 4, 0xFFFF_FFFF);
const lo = configRead(bus, dev, function, off);
const hi = configRead(bus, dev, function, off + 4);
configWrite(bus, dev, function, off, original);
configWrite(bus, dev, function, off + 4, original_high);
const readback = (@as(u64, hi) << 32) | (lo & 0xFFFF_FFF0);
const size: u64 = if (readback == 0) 0 else ~readback +% 1;
i += 1; // consumed the high half regardless
if (size == 0) continue;
descriptor.resources[slot] = .{ .kind = @intFromEnum(device.ResourceKind.memory), .start = (@as(u64, original_high) << 32) | (original & 0xFFFF_FFF0), .len = size };
descriptor.resource_count += 1;
} else {
configWrite(bus, dev, function, off, 0xFFFF_FFFF);
const readback = configRead(bus, dev, function, off);
configWrite(bus, dev, function, off, original);
const mask = readback & 0xFFFF_FFF0;
const size: u32 = if (mask == 0) 0 else ~mask +% 1;
if (size == 0) continue;
descriptor.resources[slot] = .{ .kind = @intFromEnum(device.ResourceKind.memory), .start = original & 0xFFFF_FFF0, .len = size };
descriptor.resource_count += 1;
}
}
configWrite16(bus, dev, function, 0x04, command);
}
const registered = device.register(bridge_id, &descriptor) orelse {
writeLine("pci-bus: register refused for {d}:{d}.{d}\n", .{ bus, dev, function });
return;
};
const report = protocol.ChildAdded{
.parent = bridge_id,
.bus_address = (bus << 8) | (dev << 3) | function,
.identity = class_triple,
.device_id = registered,
};
var reply: [protocol.message_maximum]u8 = undefined;
_ = runtime.ipc.call(manager_handle, std.mem.asBytes(&report), &reply) catch {
writeLine("pci-bus: child report for {d}:{d}.{d} failed\n", .{ bus, dev, function });
};
}
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = message;
_ = reply;
_ = sender;
_ = capability;
return 0;
}
pub fn main(init: runtime.process.Init) void {
const argument = init.arguments.get(1) orelse return; // bare (ramdisk sweep): stay silent
bridge_id = std.fmt.parseInt(u64, argument, 10) catch {
writeLine("pci-bus: malformed bridge device id '{s}'\n", .{argument});
return;
};
runtime.service.run(protocol.message_maximum, .{
.init = initialise,
.on_message = onMessage,
});
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
+70 -11
View File
@@ -4,11 +4,14 @@
//! argv[1]; this instance claims that device and no other, so multiple
//! instances never fight over hardware.
//!
//! M18.1 (this increment): a harness service and the first conforming driver of
//! the device-manager protocol — claim the controller, `hello` the manager
//! (role, version, assignment) inside its deadline, then serve. Controller
//! bring-up (map the MMIO window, reset, port scan) and tree reports
//! (`child_added` for each connected port) land in M18.2.
//! M18.2 (this increment): after the hello, real hardware — map the xHC's
//! register window (the first memory BAR; resource 0 is the ECAM config
//! space), read the capability registers, and walk the root-hub ports: one
//! `child_added` report to the manager per connected port, carrying the port
//! number and the PORTSC speed class as identity. No transfer rings yet —
//! descriptors and USB class matching are the USB track; the connect bit and
//! speed come straight from PORTSC, which reflects hardware state whether or
//! not the controller is running.
const std = @import("std");
const runtime = @import("runtime");
@@ -47,12 +50,16 @@ fn initialise(endpoint: runtime.ipc.Handle) bool {
return false;
};
// The controller's operational registers live behind BAR0, enumerated as
// the device's first memory resource.
const register_window = for (descriptor.resources[0..@intCast(descriptor.resource_count)]) |resource| {
if (resource.kind == @intFromEnum(device.ResourceKind.memory)) break resource;
// The xHC's registers live behind the first memory BAR. Resource 0 is the
// function's ECAM configuration space (M15), so the walk starts at 1.
var register_index: u64 = 0;
const register_window = for (descriptor.resources[1..@intCast(descriptor.resource_count)], 1..) |resource, index| {
if (resource.kind == @intFromEnum(device.ResourceKind.memory)) {
register_index = index;
break resource;
}
} else {
writeLine("usb-xhci-bus: controller device {d} has no MMIO window\n", .{controller_id});
writeLine("usb-xhci-bus: controller device {d} has no register BAR\n", .{controller_id});
return false;
};
writeLine("usb-xhci-bus: claimed controller device {d} (registers at 0x{x}, {d} bytes)\n", .{
@@ -60,6 +67,10 @@ fn initialise(endpoint: runtime.ipc.Handle) bool {
register_window.start,
register_window.len,
});
register_base = device.mmioMap(controller_id, register_index) orelse {
_ = runtime.system.write("usb-xhci-bus: mmio_map failed\n");
return false;
};
// The handshake: role, protocol version, assignment — inside the manager's
// deadline (the lookup retries cover the manager still registering).
@@ -84,14 +95,62 @@ fn initialise(endpoint: runtime.ipc.Handle) bool {
return false;
}
_ = runtime.system.write("usb-xhci-bus: hello acknowledged\n");
scanPorts(h);
return true;
}
var register_base: usize = 0;
/// One 32-bit volatile register read at `offset` from the mapped window.
fn readRegister(offset: usize) u32 {
const register: *volatile u32 = @ptrFromInt(register_base + offset);
return register.*;
}
/// The root-hub port scan: read the capability registers for the port count
/// and the operational-register offset, then one PORTSC per port. The connect
/// bit (CCS) and the speed field reflect hardware state directly — no
/// controller reset or run needed to *see* the devices; driving them needs the
/// rings (the USB track).
fn scanPorts(manager: runtime.ipc.Handle) void {
// Capability registers: CAPLENGTH is byte 0 of the first dword; HCSPARAMS1
// carries MaxPorts in bits 31:24.
const capability_length = readRegister(0) & 0xFF;
const structural = readRegister(0x04);
const maximum_ports: u32 = structural >> 24;
writeLine("usb-xhci-bus: {d} root-hub ports\n", .{maximum_ports});
// PORTSC registers: operational base + 0x400 + 0x10 per port (1-based).
var port: u32 = 1;
var connected: u32 = 0;
while (port <= maximum_ports) : (port += 1) {
const port_status = readRegister(capability_length + 0x400 + 0x10 * (port - 1));
if (port_status & 1 == 0) continue; // CCS: nothing connected
connected += 1;
const speed = (port_status >> 10) & 0xF; // the PORTSC port-speed class
writeLine("usb-xhci-bus: port {d} connected (speed class {d})\n", .{ port, speed });
const report = protocol.ChildAdded{
.parent = controller_id,
.bus_address = port,
.identity = speed,
};
var reply: [protocol.message_maximum]u8 = undefined;
_ = runtime.ipc.call(manager, std.mem.asBytes(&report), &reply) catch {
writeLine("usb-xhci-bus: child report for port {d} failed\n", .{port});
continue;
};
}
if (connected == 0) _ = runtime.system.write("usb-xhci-bus: no devices connected\n");
}
/// No bus protocol to serve yet — transfer requests arrive with the USB track.
fn onMessage(message: []const u8, reply: []u8, sender: u32) usize {
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = message;
_ = reply;
_ = sender;
_ = capability;
return 0;
}
+30 -1
View File
@@ -131,7 +131,16 @@ pub fn resourceOf(id: u64, index: u64) ?device_abi.ResourceDescriptor {
/// and would otherwise vacuously "fit" anywhere.
fn contains(parent: device_abi.ResourceDescriptor, child: device_abi.ResourceDescriptor) bool {
if (parent.kind != child.kind) return false;
if (child.kind == @intFromEnum(device_abi.ResourceKind.irq)) return parent.start == child.start;
if (child.kind == @intFromEnum(device_abi.ResourceKind.irq)) {
// Range containment: an interrupt line is still indivisible (a child owns
// exactly one GSI), but a parent may own a *range* of lines so a broad
// owner — the acpi-tables node, whose firmware names any legacy IRQ —
// can contain its children's specific lines. A length-1 parent range is
// exactly the old equality rule, so existing single-IRQ parents are
// unaffected.
const span = if (parent.len == 0) 1 else parent.len;
return child.start >= parent.start and child.start < parent.start + span;
}
if (child.len == 0 or parent.len == 0) return false;
// No overflow: a resource that wraps the address space is not containable.
const child_end = std.math.add(u64, child.start, child.len) catch return false;
@@ -180,6 +189,26 @@ pub fn register(parent_id: u64, owner: u32, descriptor: *const device_abi.Device
if (!ok) return error.NotContained;
}
// Idempotent on exact match (docs/m19-m20-plan.md decision 3): a restarted
// registering bus re-registers what it rediscovers, and the table has no
// unregister — an identical (class, identity, resources) child under the
// same parent returns the existing id instead of appending a duplicate.
for (devices[0..count]) |*existing| {
if (existing.parent != parent_id) continue;
if (existing.class != descriptor.class) continue;
if (existing.pci_class != descriptor.pci_class) continue;
if (existing.hid_len != descriptor.hid_len) continue;
if (!std.mem.eql(u8, existing.hid[0..@intCast(existing.hid_len)], descriptor.hid[0..@intCast(descriptor.hid_len)])) continue;
if (existing.resource_count != descriptor.resource_count) continue;
var same = true;
for (0..@intCast(descriptor.resource_count)) |i| {
const a = existing.resources[i];
const b = descriptor.resources[i];
if (a.kind != b.kind or a.start != b.start or a.len != b.len) same = false;
}
if (same) return existing.id;
}
var d = std.mem.zeroes(device_abi.DeviceDescriptor);
d.id = count;
d.parent = parent_id;
+22 -3
View File
@@ -298,6 +298,8 @@ fn systemDeviceEnumerate(state: *architecture.CpuState) void {
/// device_claim(id) -> 0/-1: take exclusive ownership of a device for this process.
fn systemDeviceClaim(state: *architecture.CpuState) void {
const claim_flags = sync.enter();
defer sync.leave(claim_flags);
if (devices_broker.claim(architecture.systemCallArg(state, 0), scheduler.current().id))
architecture.setSystemCallResult(state, 0)
else
@@ -312,10 +314,22 @@ fn systemMmioMap(state: *architecture.CpuState) void {
const resource_index = architecture.systemCallArg(state, 1);
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const owner = devices_broker.ownerOf(device_id) orelse return fail(state);
if (owner != t.id) return fail(state); // not claimed by this process
const r = devices_broker.resourceOf(device_id, resource_index) orelse return fail(state);
// Read the broker table under the lock: ring-3 device_register (M19) now
// mutates it concurrently on other cores, so a lock-free read here could
// see a torn resource (and a torn length used to panic the arithmetic
// below on integer overflow).
const r = blk: {
const flags = sync.enter();
defer sync.leave(flags);
const owner = devices_broker.ownerOf(device_id) orelse return fail(state);
if (owner != t.id) return fail(state); // not claimed by this process
break :blk devices_broker.resourceOf(device_id, resource_index) orelse return fail(state);
};
if (r.kind != @intFromEnum(device_abi.ResourceKind.memory)) return fail(state);
// A zero-length or wrapping window is not mappable — fail cleanly rather
// than underflow `r.len - 1`.
if (r.len == 0) return fail(state);
if (@addWithOverflow(r.start, r.len)[1] != 0) return fail(state);
if (t.device_map_next == 0) t.device_map_next = device_arena_base;
const first = r.start & ~@as(u64, page_size - 1);
@@ -451,6 +465,11 @@ fn systemDeviceRegister(state: *architecture.CpuState) void {
var descriptor: device_abi.DeviceDescriptor = undefined;
if (!ipc.copyFromUser(t.aspace, descriptor_ptr, std.mem.asBytes(&descriptor))) return fail(state);
// Under the big kernel lock: the broker's table is also mutated by the
// death sweep (releaseAllOwnedBy) and read by enumerate on other cores —
// ring-3 registration (M19) made those genuinely concurrent.
const flags = sync.enter();
defer sync.leave(flags);
const id = devices_broker.register(parent_id, t.id, &descriptor) catch return fail(state);
architecture.setSystemCallResult(state, id);
}
+3 -2
View File
@@ -350,8 +350,9 @@ pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, priority: Priority
/// context switch and lock release.
fn startUserTask() void {
const t = current();
var buffer: [96]u8 = undefined;
architecture.serialWrite(std.fmt.bufPrint(&buffer, "DBG startUserTask ip=0x{x} sp=0x{x} aspace=0x{x} kstack=0x{x}\n", .{ t.user_ip, t.user_sp, t.aspace, t.kstack_top }) catch "");
// No serial chatter here: this runs on every spawn, unserialized against
// user-space writes, and its output used to shear concurrent log lines in
// half — the largest source of corrupted markers in the QEMU scenarios.
architecture.jumpToUser(t.user_ip, t.user_sp); // noreturn
}
+350 -16
View File
@@ -140,6 +140,18 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
signalsTest(boot_information);
} else if (eql(case, "driver-restart")) {
driverRestartTest(boot_information);
} else if (eql(case, "usb-report")) {
usbReportTest(boot_information);
} else if (eql(case, "device-list")) {
deviceListTest(boot_information);
} else if (eql(case, "pci-scan")) {
pciScanTest(boot_information);
} else if (eql(case, "acpi-parse")) {
acpiParseTest(boot_information);
} else if (eql(case, "acpi-report")) {
acpiReportTest(boot_information);
} else if (eql(case, "acpi-ps2")) {
acpiReportTest(boot_information); // same spawn; the harness regex differs
} else if (eql(case, "initial-ramdisk")) {
initialRamdiskTest(boot_information);
} else if (eql(case, "vfs")) {
@@ -264,20 +276,56 @@ fn discoveryTest() void {
// M15: every PCI function now carries its own 4 KiB ECAM configuration space as
// resource 0 — the window a driver mmio_maps to walk its capability list (MSI etc).
// M19.3: the kernel seeds only the bridge; functions arrive by the ring-3
// scan (proven equivalent in pci-scan before the walk retired).
var buffer: [64]device_abi.DeviceDescriptor = undefined;
const n = @min(devices_broker.enumerate(&buffer), buffer.len);
var pci_functions: u32 = 0;
var pci_config_ok = true;
var bridges: u32 = 0;
var bridge_shape_ok = false;
for (buffer[0..n]) |d| {
if (d.class != @intFromEnum(device_abi.DeviceClass.pci_host_bridge)) continue;
bridges += 1;
var has_bus_range = false;
var has_io = false;
var memory_windows: u32 = 0;
for (d.resources[0..@intCast(d.resource_count)]) |resource| {
if (resource.kind == @intFromEnum(device_abi.ResourceKind.bus_range)) has_bus_range = true;
if (resource.kind == @intFromEnum(device_abi.ResourceKind.io_port)) has_io = true;
if (resource.kind == @intFromEnum(device_abi.ResourceKind.memory)) memory_windows += 1;
}
// ECAM plus at least one MMIO aperture, the bus range, the I/O window.
if (has_bus_range and has_io and memory_windows >= 2) bridge_shape_ok = true;
}
check("a PCI host bridge was seeded (MCFG)", bridges >= 1);
check("the bridge carries ECAM, apertures, bus range, and the I/O window", bridge_shape_ok);
// M19.0: every PCI memory resource (config slice and BARs alike) must be
// contained in one of its parent bridge's windows — the aperture derivation
// from the memory map is what makes a future user-space device_register of
// these functions pass containment. This is the assert that catches a
// too-coarse hole computation before M19.2 would.
var bars_contained = true;
for (buffer[0..n]) |d| {
if (d.class != @intFromEnum(device_abi.DeviceClass.pci_device)) continue;
pci_functions += 1;
const has_config = d.resource_count >= 1 and
d.resources[0].kind == @intFromEnum(device_abi.ResourceKind.memory) and
d.resources[0].len == abi.page_size;
if (!has_config) pci_config_ok = false;
if (d.parent >= n) {
bars_contained = false;
continue;
}
const bridge = buffer[@intCast(d.parent)];
for (d.resources[0..@intCast(d.resource_count)]) |r| {
if (r.kind != @intFromEnum(device_abi.ResourceKind.memory)) continue;
var inside = false;
for (bridge.resources[0..@intCast(bridge.resource_count)]) |w| {
if (w.kind != @intFromEnum(device_abi.ResourceKind.memory)) continue;
if (r.start >= w.start and r.start + r.len <= w.start + w.len) inside = true;
}
if (!inside) {
bars_contained = false;
log(" escaping BAR: 0x{x}+0x{x} on device {d}\n", .{ r.start, r.len, d.id });
}
}
}
check("PCI functions were enumerated (MCFG/ECAM)", pci_functions >= 1);
check("each PCI function exposes its ECAM config space as resource 0", pci_config_ok);
check("every PCI BAR lies inside a bridge aperture (M19.0)", bars_contained);
result();
}
@@ -1088,12 +1136,18 @@ fn ioPortTest() void {
var buffer: [64]device_abi.DeviceDescriptor = undefined;
const n = @min(devices_broker.enumerate(&buffer), buffer.len);
// Post-M20.3 the PS/2 node is registered at runtime by the ring-3 acpi
// service, so it is absent from this boot snapshot. Exercise the same
// io_port claim/resolve mechanism against the acpi-tables node's broad I/O
// grant — the window that now carries port authority (the service uses it
// for exactly this). The PS/2 status port 0x64 is offset 0x64 within it.
var found_id: ?u64 = null;
var found_res: u64 = 0;
outer: for (buffer[0..n]) |d| {
if (d.class != @intFromEnum(device_abi.DeviceClass.acpi_tables)) continue;
for (0..d.resource_count) |ri| {
const r = d.resources[ri];
if (r.kind == @intFromEnum(device_abi.ResourceKind.io_port) and r.start == 0x64 and r.len >= 1) {
if (r.kind == @intFromEnum(device_abi.ResourceKind.io_port) and r.start == 0 and r.len > 0x64) {
found_id = d.id;
found_res = ri;
break :outer;
@@ -1101,18 +1155,18 @@ fn ioPortTest() void {
}
}
const id = found_id orelse {
check("discovered the PS/2 status port (io_port 0x64)", false);
check("discovered the acpi-tables I/O window", false);
result();
return;
};
check("discovered the PS/2 status port (io_port 0x64)", true);
check("discovered the acpi-tables I/O window", true);
const me = scheduler.current();
check("claimed the io_port device", devices_broker.claim(id, me.id));
check("an in-range access resolves to port 0x64", process.resolveIoPort(me, id, found_res, 0, 1) == 0x64);
check("an over-wide access is refused", process.resolveIoPort(me, id, found_res, 0, 2) == null);
check("an out-of-range offset is refused", process.resolveIoPort(me, id, found_res, 1, 1) == null);
check("an unclaimed device id is refused", process.resolveIoPort(me, 0xDEAD_BEEF, found_res, 0, 1) == null);
check("an in-range access resolves to port 0x64", process.resolveIoPort(me, id, found_res, 0x64, 1) == 0x64);
check("a 4-byte access at the last port is refused", process.resolveIoPort(me, id, found_res, 0xFFFF, 4) == null);
check("an out-of-range offset is refused", process.resolveIoPort(me, id, found_res, 0x10000, 1) == null);
check("an unclaimed device id is refused", process.resolveIoPort(me, 0xDEAD_BEEF, found_res, 0x64, 1) == null);
// The kernel actually issues the `in`. Reaching this line at all proves it didn't
// fault; a width-1 read must return a single byte.
@@ -1694,6 +1748,257 @@ fn driverRestartTest(boot_information: *const BootInformation) void {
result();
}
/// M18.2: bus tree reports, end to end. The manager (test-usb-restart mode)
/// spawns the xHCI driver; the driver maps its BAR, scans the root-hub ports,
/// and reports the two QEMU devices; the manager mirrors them, kills the
/// reporter (the test trigger), prunes both children, restarts the driver with
/// backoff, and the respawned instance re-claims, re-scans, and re-reports.
/// The harness's ordered expect regex is the assertion; this test only
/// orchestrates the spawn.
fn usbReportTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: usb-report\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(image);
var manager: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{ "device-manager", "test-usb-restart" }, scheduler.currentId(), null) catch 0;
break;
}
check("device-manager spawned in test-usb-restart mode", manager != 0);
result();
}
/// M18.3: the application surface. device-list enumerates the manager's tree
/// over IPC, subscribes with its endpoint as a capability, and prints every
/// published event; the manager's delayed test-kill of the reporter produces a
/// removed/added storm the subscriber must observe. The harness's ordered
/// expect regex is the assertion.
fn deviceListTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: device-list\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(image);
var manager: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{ "device-manager", "test-usb-restart" }, scheduler.currentId(), null) catch 0;
break;
}
check("device-manager spawned in test-usb-restart mode", manager != 0);
check("device-list spawned", spawnNamed(rd, "device-list"));
result();
}
/// M19.1: the ring-3 PCI scan agrees with the kernel's. The manager spawns
/// pci-bus for the host bridge; the driver walks the same ECAM window through
/// its mmio_map grant and must find exactly the functions the kernel's own
/// enumeration recorded — the equivalence that licenses retiring the kernel
/// walk in M19.3.
fn pciScanTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: pci-scan\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
// Post-flip (M19.3) ground truth: the kernel no longer enumerates PCI
// functions, so equivalence inverts — the broker's function count after
// the scan must equal what the driver itself reported finding.
var buffer: [64]device_abi.DeviceDescriptor = undefined;
const n = @min(devices_broker.enumerate(&buffer), buffer.len);
var boot_pci: u32 = 0;
for (buffer[0..n]) |d| {
if (d.class == @intFromEnum(device_abi.DeviceClass.pci_device)) boot_pci += 1;
}
check("the kernel seeded no PCI functions (the walk retired)", boot_pci == 0);
process.setInitialRamdisk(image);
process.write_count = 0;
var manager: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{ "device-manager", "test-pci-restart" }, scheduler.currentId(), null) catch 0;
break;
}
check("device-manager spawned (test-pci-restart mode)", manager != 0);
// First scan: wait for the driver's count line and parse the number.
const count_prefix = "pci-bus: ";
const count_suffix = " functions found";
var reported: u32 = 0;
scheduler.setPriority(1);
var deadline = architecture.millis() + 15000;
while (architecture.millis() < deadline and reported == 0) {
if (process.write_len > count_prefix.len + count_suffix.len and eql(process.write_buffer[0..count_prefix.len], count_prefix)) {
const line = process.write_buffer[0..process.write_len];
const digits_end = std.mem.indexOf(u8, line, count_suffix) orelse {
scheduler.yield();
continue;
};
reported = std.fmt.parseInt(u32, line[count_prefix.len..digits_end], 10) catch 0;
}
scheduler.yield();
}
scheduler.setPriority(4);
check("the ring-3 scan reported a function count", reported >= 1);
// Every reported function was registered: the broker holds exactly them.
var registered: [64]device_abi.DeviceDescriptor = undefined;
const r = @min(devices_broker.enumerate(&registered), registered.len);
var registered_pci: u32 = 0;
for (registered[0..r]) |d| {
if (d.class == @intFromEnum(device_abi.DeviceClass.pci_device)) registered_pci += 1;
}
check("the broker holds exactly the reported functions", registered_pci == reported);
const kernel_count = reported; // the no-duplicate check below reuses it
// The restart drill: the manager kills pci-bus after its reports; the
// respawn re-claims, re-scans, and re-registers.
const restart_marker = "device-manager: restarting pci-bus";
scheduler.setPriority(1);
deadline = architecture.millis() + 15000;
var restarted = false;
while (architecture.millis() < deadline and !restarted) {
if (process.write_len >= restart_marker.len and eql(process.write_buffer[0..restart_marker.len], restart_marker)) restarted = true;
scheduler.yield();
}
scheduler.setPriority(4);
check("the manager restarted pci-bus", restarted);
var marker_buffer: [48]u8 = undefined;
const marker = std.fmt.bufPrint(&marker_buffer, "pci-bus: {d} functions found", .{reported}) catch "";
scheduler.setPriority(1);
deadline = architecture.millis() + 15000;
var seen = false;
while (architecture.millis() < deadline and !seen) {
if (process.write_len >= marker.len and eql(process.write_buffer[0..marker.len], marker)) seen = true;
scheduler.yield();
}
scheduler.setPriority(4);
check("the respawned scan reported the same count", seen);
// No duplicates: the registrations deduped against the kernel's own nodes
// on the first pass, and against themselves on the second.
var after: [64]device_abi.DeviceDescriptor = undefined;
const m = @min(devices_broker.enumerate(&after), after.len);
var after_count: u32 = 0;
for (after[0..m]) |d| {
if (d.class == @intFromEnum(device_abi.DeviceClass.pci_device)) after_count += 1;
}
check("no duplicate PCI nodes after register + restart + re-register", after_count == kernel_count);
result();
}
/// M20.2: the acpi service registers + reports its _HID devices. Boot normally
/// (the manager spawns discovery); the harness's expect regex requires the two
/// PS/2 nodes among the service's report lines, each with its _CRS resources —
/// the ring-3 _CRS/_STA evaluation working end to end. The kernel test only
/// starts the manager.
fn acpiReportTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: acpi-report\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(image);
var spawned = false;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue;
_ = process.spawnProcessSupervised(item.blob, 4, &.{"device-manager"}, scheduler.currentId(), null) catch 0;
spawned = true;
break;
}
check("device-manager spawned", spawned);
result();
}
/// M20.1: the ring-3 AML parse agrees with the kernel's. The manager spawns
/// the discovery service (the acpi build variant); it claims the acpi-tables
/// node, maps the blobs, parses them, and logs its Device count — which must
/// equal what the kernel's own parse produced (the equivalence that licenses
/// retiring the kernel's device build in M20.3).
fn acpiParseTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: acpi-parse\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
// The kernel's own count, from the namespace it already built for \_S5.
const kernel_devices = platform.amlDeviceCount();
check("the kernel namespace has devices to compare against", kernel_devices >= 1);
// Spawn the discovery service directly with that count as argv: it parses
// the same blobs in ring 3 and self-verifies, printing "acpi-parse: ok" iff
// the counts match. The harness's expect regex is that marker — deterministic,
// no racing the shared serial buffer.
process.setInitialRamdisk(image);
var count_text: [16]u8 = undefined;
const count_arg = std.fmt.bufPrint(&count_text, "{d}", .{kernel_devices}) catch "0";
var spawned = false;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "discovery")) continue;
_ = process.spawnProcessSupervised(item.blob, 4, &.{ "discovery", count_arg }, scheduler.currentId(), null) catch 0;
spawned = true;
break;
}
check("discovery service spawned", spawned);
result();
}
/// The whole user-side surface at once: spawn process-test's supervisor role,
/// which — entirely from ring 3 — creates an exit endpoint, spawns its two
/// children supervised, sees them in process_enumerate, kills them (one blocked,
@@ -2014,6 +2319,35 @@ fn hpetGsi() ?u32 {
/// land in the device table with the containment invariant intact.
fn busTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: bus\n", .{});
// M19.0: device_register is idempotent on exact match — a restarted
// registering bus must not duplicate its children. Driven directly against
// the broker: claim an unclaimed node, register the same (class, hid,
// resourceless) child twice, expect one id and one table entry.
{
const me = scheduler.currentId();
var probe: [1]device_abi.DeviceDescriptor = undefined;
const total = devices_broker.enumerate(&probe);
check("device tree is seeded for the idempotence check", total >= 1);
if (devices_broker.ownerOf(0) == null) {
check("claimed device 0 for the idempotence check", devices_broker.claim(0, me));
var child = std.mem.zeroes(device_abi.DeviceDescriptor);
child.class = @intFromEnum(device_abi.DeviceClass.unknown);
child.pci_class = device_abi.no_pci_class;
child.hid_len = 4;
child.hid[0..4].* = "idem".*;
const first = devices_broker.register(0, me, &child) catch 0;
check("first register succeeded", first != 0);
const before = devices_broker.enumerate(&probe);
const second = devices_broker.register(0, me, &child) catch 0;
check("re-register returned the same id", second == first);
check("re-register grew nothing", devices_broker.enumerate(&probe) == before);
devices_broker.releaseAllOwnedBy(me);
} else {
check("device 0 unexpectedly claimed before the idempotence check", false);
}
}
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
+364
View File
@@ -0,0 +1,364 @@
//! /system/services/acpi — the ACPI discovery service: the x86 firmware
//! interpreter, moved out of ring 0 (docs/m19-m20-plan.md, M20). Claims the
//! `acpi-tables` node the kernel publishes (the AML blobs, the broad io_port
//! grant, a broad irq window, the SCI), and runs the **shared AML module** in
//! ring 3 — the same parser and interpreter the kernel uses.
//!
//! M20.2 (this increment): after parsing, walk the namespace and, for each
//! present Device with a hardware id (`_HID`), evaluate its current resource
//! settings (`_CRS`) through a ring-3 `Hal` (port I/O over the claimed node),
//! register it under the acpi-tables node (its I/O ports and IRQs contained by
//! the node's broad grants), and report it to the device manager with its
//! EISA-decoded hid as identity. Matching those reports to drivers (ps2-bus)
//! and retiring the kernel's own device build follow in M20.3.
const std = @import("std");
const runtime = @import("runtime");
const aml = @import("aml");
const device = runtime.device;
const protocol = runtime.device_manager_protocol;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
// The claimed acpi-tables node and the resource index of its broad io_port
// window — the Hal routes every port access through this one claim.
var node_id: u64 = 0;
var io_resource_index: u64 = 0;
// Pass-1 registration record (see main): what pass 2 reports.
const Registered = struct { hid: [8]u8 = .{0} ** 8, hid_len: usize = 0, device_id: u64 = 0, resource_count: u64 = 0 };
var registered: [64]Registered = undefined;
var registered_count: usize = 0;
// A scratch page returned for SystemMemory OperationRegion maps: the service
// cannot map arbitrary physical memory from ring 3, so such regions are
// unsupported and degrade to harmless zeros rather than faulting. The M20.2
// targets (ps2, the legacy devices) use SystemIO and static templates.
var mmio_scratch: [4096]u8 align(4096) = .{0} ** 4096;
fn halMapMmio(physical: u64, len: u64, writable: bool) u64 {
_ = physical;
_ = len;
_ = writable;
return @intFromPtr(&mmio_scratch);
}
fn halPioRead(width: u8, port: u16) u32 {
return device.ioRead(node_id, io_resource_index, port, width) orelse 0;
}
fn halPioWrite(width: u8, port: u16, value: u32) void {
_ = device.ioWrite(node_id, io_resource_index, port, width, value);
}
fn findTablesNode(buffer: []device.DeviceDescriptor) ?device.DeviceDescriptor {
const total = device.enumerate(buffer);
for (buffer[0..@min(total, buffer.len)]) |d| {
if (d.class == @intFromEnum(device.DeviceClass.acpi_tables)) return d;
}
return null;
}
pub fn main(init: runtime.process.Init) void {
// When the acpi-parse scenario spawns this directly, argv[1] is the kernel's
// own device count to self-verify against — deterministic, no log-scraping.
const expected: ?usize = if (init.arguments.get(1)) |a| (std.fmt.parseInt(usize, a, 10) catch null) else null;
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("acpi: out of memory\n");
return;
};
const node = findTablesNode(buffer) orelse {
_ = runtime.system.write("acpi: no acpi-tables node to claim\n");
return;
};
node_id = node.id;
if (!device.claim(node_id)) {
_ = runtime.system.write("acpi: unable to claim acpi-tables\n");
return;
}
// Map each memory resource (an AML blob) and note the io_port resource.
var blocks: [8][]const u8 = undefined;
var block_count: usize = 0;
var found_io = false;
for (node.resources[0..@intCast(node.resource_count)], 0..) |resource, index| {
if (resource.kind == @intFromEnum(device.ResourceKind.io_port) and !found_io) {
io_resource_index = index;
found_io = true;
continue;
}
if (resource.kind != @intFromEnum(device.ResourceKind.memory)) continue;
const base = device.mmioMap(node_id, index) orelse continue;
const pointer: [*]const u8 = @ptrFromInt(base);
blocks[block_count] = pointer[0..@intCast(resource.len)];
block_count += 1;
if (block_count == blocks.len) break;
}
if (block_count == 0) {
_ = runtime.system.write("acpi: no AML blobs on the node\n");
return;
}
const result = aml.parse(runtime.allocator(), blocks[0..block_count]) catch {
_ = runtime.system.write("acpi: AML parse failed\n");
return;
};
var namespace = result.namespace;
const devices = aml.deviceCount(&namespace);
writeLine("acpi: parsed {d} AML blob(s), {d} namespace devices\n", .{ block_count, devices });
if (expected) |want| {
if (devices == want) {
_ = runtime.system.write("acpi-parse: ok\n");
} else {
writeLine("acpi-parse: mismatch (ring-3 {d} vs kernel {d})\n", .{ devices, want });
}
// Self-verify mode is standalone (no manager); stop before reporting.
while (true) runtime.system.sleep(1000);
}
// Register + report the present _HID devices (M20.2).
var arena = std.heap.ArenaAllocator.init(runtime.allocator());
var interpreter = aml.Interpreter.init(&namespace, .{
.mapMmio = halMapMmio,
.pioRead = halPioRead,
.pioWrite = halPioWrite,
}, arena.allocator());
// Pass 1: register every present _HID device under acpi-tables, remembering
// each (hid, device id). Pass 2: report them all. Registering before any
// report reaches the manager means a driver it spawns on the first report
// already sees the whole set (no keyboard-before-mouse race for ps2-bus).
registered_count = 0;
walkDevices(namespace.root, &interpreter);
const manager = runtime.ipc.lookup(.device_manager);
var i: usize = 0;
while (i < registered_count) : (i += 1) {
const entry = registered[i];
writeLine("acpi: reported {s} (device {d}, {d} resources)\n", .{ entry.hid[0..entry.hid_len], entry.device_id, entry.resource_count });
if (manager) |h| {
var report = protocol.ChildAdded{
.parent = node_id,
.bus_address = entry.device_id,
.identity = 0,
.device_id = entry.device_id,
};
@memcpy(report.hid[0..entry.hid_len], entry.hid[0..entry.hid_len]);
var reply: [protocol.message_maximum]u8 = undefined;
_ = runtime.ipc.call(h, std.mem.asBytes(&report), &reply) catch {};
}
}
writeLine("acpi: reported {d} device(s) to the manager\n", .{registered_count});
// Stay resident: the claim holds, and the service is here to grow into the
// supervised discoverer (M20.3, then the M21 event side on the SCI).
while (true) runtime.system.sleep(1000);
}
/// Depth-first walk: register + report each present device with a _HID, then
/// descend. Scopes (\_SB, \_GPE …) are descended without producing a node.
fn walkDevices(node: *aml.Node, interpreter: *aml.Interpreter) void {
var child = node.first_child;
while (child) |c| : (child = c.next_sibling) {
if (c.kind != .device) {
walkDevices(c, interpreter);
continue;
}
if (!devicePresent(interpreter, c)) continue; // absent: skip it and its subtree
if (readHid(c, interpreter)) |hid| {
// Skip PCI roots — pci-bus already reports PCI functions; ACPI adds
// only the non-PCI _HID devices (docs/m19-m20-plan.md M20.2).
if (!std.mem.eql(u8, hid[0..7], "PNP0A03") and !std.mem.eql(u8, hid[0..7], "PNP0A08")) {
registerDevice(c, hid, interpreter);
}
}
walkDevices(c, interpreter);
}
}
fn registerDevice(node: *aml.Node, hid: [8]u8, interpreter: *aml.Interpreter) void {
if (registered_count >= registered.len) return;
var descriptor = std.mem.zeroes(device.DeviceDescriptor);
descriptor.class = @intFromEnum(device.DeviceClass.acpi_device);
descriptor.pci_class = device.no_pci_class;
const hid_len: u64 = std.mem.indexOfScalar(u8, &hid, 0) orelse hid.len;
descriptor.hid_len = hid_len;
@memcpy(descriptor.hid[0..@intCast(hid_len)], hid[0..@intCast(hid_len)]);
applyCrs(&descriptor, node, interpreter);
const id = device.register(node_id, &descriptor) orelse {
writeLine("acpi: register refused for {s}\n", .{hid[0..@intCast(hid_len)]});
return;
};
registered[registered_count] = .{ .hid = hid, .hid_len = @intCast(hid_len), .device_id = id, .resource_count = descriptor.resource_count };
registered_count += 1;
}
/// _STA bit 0 (present); absent method or a failed evaluation is treated as
/// present, per the ACPI rules.
fn devicePresent(interpreter: *aml.Interpreter, node: *aml.Node) bool {
const sta = aml.Namespace.childOf(node, seg4("_STA")) orelse return true;
const obj = interpreter.evaluate(sta, &.{}) catch return true;
const status = obj.asInteger() catch return true;
return (status & 0x01) != 0;
}
/// The device's EISA-decoded _HID (e.g. "PNP0303"), or null.
fn readHid(node: *aml.Node, interpreter: *aml.Interpreter) ?[8]u8 {
const hid = aml.Namespace.childOf(node, seg4("_HID")) orelse return null;
var buffer: [8]u8 = .{0} ** 8;
if (hid.kind == .method) {
const obj = interpreter.evaluate(hid, &.{}) catch return null;
switch (obj) {
.integer => |n| {
_ = eisaIdToStr(@truncate(n), &buffer);
return buffer;
},
else => return null,
}
}
if (hid.kind != .name or hid.value.len == 0) return null;
const v = hid.value;
switch (v[0]) {
0x00, 0x01, 0xFF, 0x0A, 0x0B, 0x0C, 0x0E => {
var p: usize = 0;
const n = readIntObj(v, &p) orelse return null;
_ = eisaIdToStr(@truncate(n), &buffer);
return buffer;
},
else => return null,
}
}
// --- _CRS resource-template decode (ported from the kernel's acpi.zig) --------
fn applyCrs(descriptor: *device.DeviceDescriptor, node: *aml.Node, interpreter: *aml.Interpreter) void {
const crs = aml.Namespace.childOf(node, seg4("_CRS")) orelse return;
const obj = interpreter.evaluate(crs, &.{}) catch return;
const bytes = switch (obj) {
.buffer => |b| b,
else => return,
};
var i: usize = 0;
while (i < bytes.len) {
const tag = bytes[i];
if (tag & 0x80 == 0) {
const len: usize = tag & 0x07;
const body = i + 1;
if (body + len > bytes.len) break;
switch ((tag >> 3) & 0x0F) {
0x04 => if (len >= 2) { // IRQ mask
const mask = @as(u16, bytes[body]) | (@as(u16, bytes[body + 1]) << 8);
var b: usize = 0;
while (b < 16) : (b += 1) {
if (mask & (@as(u16, 1) << @intCast(b)) != 0) addResource(descriptor, .irq, b, 1);
}
},
0x08 => if (len >= 7) addResource(descriptor, .io_port, rd16(bytes, body + 1), bytes[body + 6]),
0x09 => if (len >= 3) addResource(descriptor, .io_port, rd16(bytes, body), bytes[body + 2]),
0x0F => break,
else => {},
}
i = body + len;
} else {
if (i + 3 > bytes.len) break;
const len: usize = @intCast(rd16(bytes, i + 1));
const body = i + 3;
if (body + len > bytes.len) break;
switch (tag) {
0x85 => if (len >= 17) addResource(descriptor, .memory, rd32(bytes, body + 1), rd32(bytes, body + 13)),
0x86 => if (len >= 9) addResource(descriptor, .memory, rd32(bytes, body + 1), rd32(bytes, body + 5)),
0x89 => if (len >= 2) {
const count = bytes[body + 1];
var k: usize = 0;
while (k < count and body + 2 + k * 4 + 4 <= body + len) : (k += 1) {
addResource(descriptor, .irq, rd32(bytes, body + 2 + k * 4), 1);
}
},
else => {},
}
i = body + len;
}
}
}
fn addResource(descriptor: *device.DeviceDescriptor, kind: device.ResourceKind, start: u64, len: u64) void {
if (descriptor.resource_count >= descriptor.resources.len) return;
descriptor.resources[@intCast(descriptor.resource_count)] = .{ .kind = @intFromEnum(kind), .start = start, .len = len };
descriptor.resource_count += 1;
}
// --- small helpers ported verbatim from the kernel's acpi.zig ----------------
fn seg4(comptime s: *const [4:0]u8) [4]u8 {
return s[0..4].*;
}
fn hexDigit(n: u8) u8 {
return if (n < 10) '0' + n else 'A' + (n - 10);
}
fn eisaIdToStr(id: u32, buffer: *[8]u8) []const u8 {
const b0: u16 = @intCast(id & 0xFF);
const b1: u16 = @intCast((id >> 8) & 0xFF);
const b2: u8 = @truncate(id >> 16);
const b3: u8 = @truncate(id >> 24);
const mfg = (b0 << 8) | b1;
buffer[0] = '@' + @as(u8, @intCast((mfg >> 10) & 0x1F));
buffer[1] = '@' + @as(u8, @intCast((mfg >> 5) & 0x1F));
buffer[2] = '@' + @as(u8, @intCast(mfg & 0x1F));
buffer[3] = hexDigit((b2 >> 4) & 0xF);
buffer[4] = hexDigit(b2 & 0xF);
buffer[5] = hexDigit((b3 >> 4) & 0xF);
buffer[6] = hexDigit(b3 & 0xF);
buffer[7] = 0;
return buffer[0..7];
}
fn readIntObj(bytes: []const u8, p: *usize) ?u64 {
if (p.* >= bytes.len) return null;
const op = bytes[p.*];
p.* += 1;
switch (op) {
0x00 => return 0,
0x01 => return 1,
0xFF => return 1,
0x0A => {
if (p.* >= bytes.len) return null;
const v = bytes[p.*];
p.* += 1;
return v;
},
0x0B => {
if (p.* + 2 > bytes.len) return null;
const v = rd16(bytes, p.*);
p.* += 2;
return v;
},
0x0C => {
if (p.* + 4 > bytes.len) return null;
const v = rd32(bytes, p.*);
p.* += 4;
return v;
},
else => return null,
}
}
fn rd16(bytes: []const u8, off: usize) u64 {
return @as(u64, bytes[off]) | (@as(u64, bytes[off + 1]) << 8);
}
fn rd32(bytes: []const u8, off: usize) u64 {
return rd16(bytes, off) | (rd16(bytes, off + 2) << 16);
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
@@ -0,0 +1,88 @@
//! device-list — the `ps` analog for the device tree (docs/device-manager.md
//! M18.3): asks the device manager for the tree over IPC, prints it, then
//! subscribes and prints every published add/remove event. The manager is the
//! one answer to "what devices exist" for user space; nothing here touches a
//! device_* system call.
const std = @import("std");
const runtime = @import("runtime");
const protocol = runtime.device_manager_protocol;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [96]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
pub fn main() void {
var manager: ?runtime.ipc.Handle = null;
var tries: u32 = 0;
while (manager == null and tries < 200) : (tries += 1) {
manager = runtime.ipc.lookup(.device_manager);
if (manager == null) runtime.system.sleep(20);
}
const h = manager orelse {
_ = runtime.system.write("device-list: no device manager\n");
return;
};
// The snapshot — polled briefly, because at boot the bus drivers may still
// be scanning: an empty first answer usually just means "too early".
var reply: [protocol.message_maximum]u8 = undefined;
var count: u32 = 0;
var length: usize = 0;
tries = 0;
while (tries < 20) : (tries += 1) {
const request = protocol.Enumerate{};
length = runtime.ipc.call(h, std.mem.asBytes(&request), &reply) catch 0;
if (length >= @sizeOf(protocol.EnumerateReply)) {
count = std.mem.bytesToValue(protocol.EnumerateReply, reply[0..@sizeOf(protocol.EnumerateReply)]).count;
if (count != 0) break;
}
runtime.system.sleep(100);
}
writeLine("device-list: {d} devices\n", .{count});
var offset: usize = @sizeOf(protocol.EnumerateReply);
var index: u32 = 0;
while (index < count and offset + @sizeOf(protocol.ChildEntry) <= length) : (index += 1) {
const entry = std.mem.bytesToValue(protocol.ChildEntry, reply[offset..][0..@sizeOf(protocol.ChildEntry)]);
writeLine("device-list: device {d} port {d} identity {d}\n", .{ entry.parent, entry.bus_address, entry.identity });
offset += @sizeOf(protocol.ChildEntry);
}
// The subscription: our endpoint rides as the call's capability; events
// arrive as buffered messages carrying the same structs the bus sends.
const endpoint = runtime.ipc.createIpcEndpoint() orelse {
_ = runtime.system.write("device-list: no endpoint\n");
return;
};
const subscribe = protocol.Subscribe{};
_ = runtime.ipc.callCap(h, std.mem.asBytes(&subscribe), &reply, endpoint) catch {
_ = runtime.system.write("device-list: subscribe failed\n");
return;
};
_ = runtime.system.write("device-list: subscribed\n");
var receive: [protocol.message_maximum]u8 = undefined;
while (true) {
const got = runtime.ipc.replyWait(endpoint, &.{}, &receive, null);
if (!got.isMessage() or got.len < 1) continue;
switch (receive[0]) {
@intFromEnum(protocol.Operation.child_added) => {
if (got.len < protocol.child_added_size) continue;
const event = std.mem.bytesToValue(protocol.ChildAdded, receive[0..protocol.child_added_size]);
writeLine("device-list: added (device {d} port {d})\n", .{ event.parent, event.bus_address });
},
@intFromEnum(protocol.Operation.child_removed) => {
if (got.len < protocol.child_removed_size) continue;
const event = std.mem.bytesToValue(protocol.ChildRemoved, receive[0..protocol.child_removed_size]);
writeLine("device-list: removed (device {d} port {d})\n", .{ event.parent, event.bus_address });
},
else => {},
}
}
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
@@ -19,10 +19,13 @@ pub const Role = enum(u8) {
device = 2,
};
/// The message kinds. `child_added`/`child_removed` land in M18.2;
/// `enumerate`/`subscribe` in M18.3.
/// The message kinds.
pub const Operation = enum(u8) {
hello = 1,
child_added = 2,
child_removed = 3,
enumerate = 4,
subscribe = 5,
};
/// `Hello.device_id` for a driver that serves no enumerated device (a test
@@ -54,5 +57,92 @@ pub const HelloReply = extern struct {
pub const reply_size = @sizeOf(HelloReply);
/// A bus driver reporting one device it discovered behind its controller
/// (docs/device-manager.md "the tree"). Identity is the bus's native language —
/// for USB a port-speed class; the (class, subclass, protocol) triple joins it
/// once control transfers exist (the USB track). The manager mirrors the child
/// into its tree; when the reporting driver dies, the manager prunes everything
/// it reported (the children describe protocol state that died with it) and the
/// restarted instance rediscovers and re-reports.
pub const ChildAdded = extern struct {
operation: u8 = @intFromEnum(Operation.child_added),
reserved0: u8 = 0,
reserved1: u16 = 0,
reserved2: u32 = 0,
/// The reporting driver's own device (the controller) — the child's parent.
parent: u64,
/// Where on the bus (for USB: the root port number, 1-based).
bus_address: u64,
/// Bus-specific identity (for USB: the PORTSC port-speed class; for PCI:
/// the class triple; for ACPI devices, 0 — identity is the hid below).
identity: u64,
/// The kernel device id this child was `device_register`ed as — what the
/// manager hands a matched driver as its argv assignment — or `no_device`
/// for an unregistered leaf (a USB port before the descriptor track).
device_id: u64 = no_device,
/// The ACPI hardware id (`_HID`), EISA-decoded (e.g. "PNP0303"), for devices
/// discovered by firmware string rather than a numeric bus identity. Empty
/// (all zero) otherwise. Widens for FDT `compatible` strings later.
hid: [8]u8 = .{0} ** 8,
};
pub const child_added_size = @sizeOf(ChildAdded);
/// A bus driver reporting a device gone (hot-unplug). Not yet sent by any
/// driver — the port scan has no unplug interrupt — but the manager handles it;
/// death-pruning covers removal until hotplug lands.
pub const ChildRemoved = extern struct {
operation: u8 = @intFromEnum(Operation.child_removed),
reserved0: u8 = 0,
reserved1: u16 = 0,
reserved2: u32 = 0,
parent: u64,
bus_address: u64,
};
pub const child_removed_size = @sizeOf(ChildRemoved);
/// The manager's answer to a tree report.
pub const ReportReply = extern struct {
status: i32,
reserved: u32 = 0,
};
/// An application asking for the tree (M18.3): the reply is an EnumerateReply
/// header followed by `count` ChildEntry records.
pub const Enumerate = extern struct {
operation: u8 = @intFromEnum(Operation.enumerate),
reserved0: u8 = 0,
reserved1: u16 = 0,
reserved2: u32 = 0,
};
pub const EnumerateReply = extern struct {
status: i32,
/// ChildEntry records following this header.
count: u32,
};
pub const ChildEntry = extern struct {
parent: u64,
bus_address: u64,
identity: u64,
};
/// An application subscribing to published add/remove events (the input-service
/// pattern): the subscriber's endpoint rides as the call's **capability**, and
/// events arrive on it as buffered messages whose payload is the same
/// ChildAdded / ChildRemoved struct the bus drivers send — one encoding, both
/// directions.
pub const Subscribe = extern struct {
operation: u8 = @intFromEnum(Operation.subscribe),
reserved0: u8 = 0,
reserved1: u16 = 0,
reserved2: u32 = 0,
};
/// Upper bound on any message in this protocol — sizes the endpoint buffers.
pub const message_maximum = 64;
/// Capped by the kernel's IPC MESSAGE_MAXIMUM (256): an EnumerateReply carries
/// up to ten ChildEntry records per call, plenty for the mirror's current
/// bounds; paging joins the protocol if a tree ever outgrows one message.
pub const message_maximum = 256;
+258 -20
View File
@@ -34,15 +34,11 @@ fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
/// this comes from a manifest (docs/device-manager.md: the third bus type
/// triggers it); for now a static map. `null` = no driver for this class yet.
fn driverFor(d: device.DeviceDescriptor) ?[]const u8 {
// detect device via DeviceClass
// The HPET timer node is still kernel-seeded (from the HPET table, not AML).
// PS/2 and other _HID devices now arrive as acpi-service reports and match
// in onChildAdded (M20.3), not from this boot snapshot.
if (d.class == @intFromEnum(device.DeviceClass.timer)) return "hpet";
// detect device via hid
const hid = d.hid[0..@intCast(d.hid_len)];
const id = acpi_ids.HardwareId.fromHid(hid) orelse return null;
return switch (id) {
.ps2_keyboard, .ps2_mouse => "ps2-bus",
else => null,
};
return null;
}
/// The PCI class/subclass/prog-IF triple of an xHCI (USB 3) host controller:
@@ -50,17 +46,36 @@ fn driverFor(d: device.DeviceDescriptor) ?[]const u8 {
/// pci-class.zig decodes.
const xhci_pci_class: u64 = 0x0C_03_30;
/// The bus driver that serves a PCI function, or null. A machine can carry
/// several identical controllers — one driver instance per device, the id as
/// argv[1]. These drivers speak the protocol: a hello is expected.
fn pciDriverFor(d: device.DeviceDescriptor) ?[]const u8 {
if (d.class != @intFromEnum(device.DeviceClass.pci_device)) return null;
return switch (d.pci_class) {
/// The driver that serves a *reported* PCI function (M19.3: matching moved
/// from the boot snapshot to the bus reports), or null. A machine can carry
/// several identical controllers — one driver instance per reported device,
/// its registered id as argv[1].
fn pciDriverForIdentity(identity: u64) ?[]const u8 {
return switch (identity) {
xhci_pci_class => "usb-xhci-bus",
else => null,
};
}
/// The driver that serves a *reported* ACPI device by its `_HID` (M20.3:
/// ps2-bus now binds the PS/2 nodes the acpi service reports, not boot-snapshot
/// nodes the kernel used to build). ps2-bus is a singleton that finds both its
/// devices by hid once spawned, so keyboard and mouse map to the same name.
fn hidDriverFor(hid: []const u8) ?[]const u8 {
if (std.mem.eql(u8, hid, "PNP0303")) return "ps2-bus"; // PS/2 keyboard
if (std.mem.eql(u8, hid, "PNP0F13")) return "ps2-bus"; // PS/2 mouse
return null;
}
/// Whether some driver entry already serves registered device `device_id` —
/// a re-report after a bus restart must not spawn a second instance.
fn driverForDevice(device_id: u64) bool {
for (&drivers) |*driver| {
if (driver.used and driver.device_id == device_id) return true;
}
return false;
}
// --- supervision -------------------------------------------------------------
/// How long a protocol driver has to hello after its spawn.
@@ -106,6 +121,88 @@ const maximum_drivers = 16;
var drivers: [maximum_drivers]Driver = .{Driver{}} ** maximum_drivers;
var manager_endpoint: runtime.ipc.Handle = 0;
var test_restart_mode = false;
var test_usb_restart_mode = false;
var test_usb_killed = false;
var test_pci_restart_mode = false;
var test_kill_pid: u32 = 0;
var test_kill_due_ns: u64 = 0;
/// The application subscribers (M18.3, the input-service pattern): endpoints
/// handed over as capabilities, each receiving every child add/remove as a
/// buffered message. A subscriber whose endpoint stops accepting (it died) is
/// dropped on the failed send.
const maximum_subscribers = 8;
var subscribers: [maximum_subscribers]?runtime.ipc.Handle = .{null} ** maximum_subscribers;
/// Publish one event (a ChildAdded or ChildRemoved struct, the same encoding
/// the bus drivers send) to every subscriber.
fn publishEvent(event: []const u8) void {
for (&subscribers) |*slot| {
if (slot.*) |handle| {
if (!runtime.ipc.send(handle, event)) slot.* = null; // dead subscriber
}
}
}
/// The manager's mirror of what bus drivers report (docs/device-manager.md "the
/// tree"): the children, keyed by (parent, bus address), each remembering which
/// driver instance reported it — that is what death-pruning sweeps by.
const Child = struct {
used: bool = false,
parent: u64 = 0,
bus_address: u64 = 0,
identity: u64 = 0,
// The kernel device id (registered by the reporter), or protocol.no_device.
device_id: u64 = 0,
reporter: u32 = 0, // the reporting driver instance's process id
};
const maximum_children = 64; // ACPI adds ~34 device nodes (M20.2), plus PCI + USB
var children: [maximum_children]Child = .{Child{}} ** maximum_children;
/// Record (or refresh) a reported child. Refreshing matters: a restarted bus
/// driver re-reports what it rediscovers, and the same (parent, port) must not
/// duplicate.
fn addChild(parent: u64, bus_address: u64, identity: u64, device_id: u64, reporter: u32) bool {
var free: ?*Child = null;
for (&children) |*child| {
if (child.used and child.parent == parent and child.bus_address == bus_address) {
child.identity = identity;
child.device_id = device_id;
child.reporter = reporter;
return true;
}
if (!child.used and free == null) free = child;
}
const slot = free orelse return false;
slot.* = .{ .used = true, .parent = parent, .bus_address = bus_address, .identity = identity, .device_id = device_id, .reporter = reporter };
return true;
}
/// Prune every child a dead driver instance reported: the children describe
/// protocol state (slots, rings) that died with the process — keeping the nodes
/// would be keeping a lie. The restarted instance rediscovers and re-reports.
/// Watchers hear the honest story: removed now, added again on rediscovery.
fn pruneChildrenOf(reporter: u32) void {
for (&children) |*child| {
if (child.used and child.reporter == reporter) {
writeLine("device-manager: child removed (device {d} port {d})\n", .{ child.parent, child.bus_address });
child.used = false;
const event = protocol.ChildRemoved{ .parent = child.parent, .bus_address = child.bus_address };
publishEvent(std.mem.asBytes(&event));
}
}
}
/// How many children a driver instance has reported (the test-usb-restart
/// trigger counts these).
fn childCountOf(reporter: u32) u32 {
var n: u32 = 0;
for (&children) |*child| {
if (child.used and child.reporter == reporter) n += 1;
}
return n;
}
fn driverByProcess(process_id: u32) ?*Driver {
for (&drivers) |*driver| {
@@ -171,9 +268,11 @@ fn spawnDriver(driver: *Driver) void {
}
}
/// A driver died. The exit reason (M17.2) is the whole decision: a clean exit
/// meant to stop; anything else restarts with backoff until the crash-loop cap.
/// A driver died. Prune what it reported first — then the exit reason (M17.2)
/// is the whole restart decision: a clean exit meant to stop; anything else
/// restarts with backoff until the crash-loop cap.
fn onDriverExit(driver: *Driver) void {
pruneChildrenOf(driver.process_id);
const reason = runtime.process.exitReason(driver.process_id) orelse .fault;
if (reason == .exited) {
driver.state = .stopped;
@@ -201,6 +300,11 @@ fn onDriverExit(driver: *Driver) void {
/// sweep serves every armed deadline.
fn sweepDeadlines() void {
const now = system.clock();
if (test_kill_pid != 0 and now >= test_kill_due_ns) {
writeLine("device-manager: test mode: killing the reporter\n", .{});
_ = system.kill(test_kill_pid);
test_kill_pid = 0;
}
for (&drivers) |*driver| {
if (!driver.used) continue;
switch (driver.state) {
@@ -230,11 +334,15 @@ fn initialise(endpoint: runtime.ipc.Handle) bool {
var matched: usize = 0;
for (buffer[0..count]) |descriptor| {
if (pciDriverFor(descriptor)) |driver_name| {
if (descriptor.class == @intFromEnum(device.DeviceClass.pci_host_bridge)) {
// The PCI bus driver: enumeration in ring 3 (M19), one instance
// per bridge, the bridge id as its assignment.
matched += 1;
addDriver(driver_name, descriptor.id, true);
addDriver("pci-bus", descriptor.id, true);
continue;
}
// PCI functions no longer appear in the boot snapshot (M19.3): the
// pci-bus driver reports them, and onChildAdded matches from reports.
const driver_name = driverFor(descriptor) orelse continue;
matched += 1;
// Skip a singleton that is already alive (the initial-ramdisk sweep test
@@ -245,6 +353,12 @@ fn initialise(endpoint: runtime.ipc.Handle) bool {
}
}
// The discovery service (docs/m19-m20-plan.md M20): one per firmware, packed
// under the neutral name "discovery", spawned once at startup. It finds and
// claims the acpi-tables (or devicetree-blob) node itself. Not a per-device
// match — it is the discoverer, not a driver bound to one device.
addDriver("discovery", protocol.no_device, false);
if (test_restart_mode) {
// The driver-restart scenario's fixture: claims device 0 (the tree
// root, otherwise unclaimed), hellos, then faults — driving backoff,
@@ -260,10 +374,18 @@ fn initialise(endpoint: runtime.ipc.Handle) bool {
return true;
}
fn onMessage(message: []const u8, reply: []u8, sender: u32) usize {
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
if (message.len < 1) return 0;
switch (message[0]) {
@intFromEnum(protocol.Operation.child_added) => return onChildAdded(message, reply, sender),
@intFromEnum(protocol.Operation.child_removed) => return onChildRemoved(message, reply, sender),
@intFromEnum(protocol.Operation.enumerate) => return onEnumerate(reply),
@intFromEnum(protocol.Operation.subscribe) => return onSubscribe(reply, capability),
@intFromEnum(protocol.Operation.hello) => {},
else => return 0,
}
if (message.len < protocol.hello_size) return 0;
const hello = std.mem.bytesToValue(protocol.Hello, message[0..protocol.hello_size]);
if (hello.operation != @intFromEnum(protocol.Operation.hello)) return 0;
var status: i32 = 0;
if (hello.version != protocol.version) {
@@ -281,6 +403,120 @@ fn onMessage(message: []const u8, reply: []u8, sender: u32) usize {
return protocol.reply_size;
}
/// A bus driver reported a discovered device: mirror it, and in
/// test-usb-restart mode kill the reporter once after its second child — the
/// deterministic trigger for prune -> backoff -> respawn -> re-report.
fn onChildAdded(message: []const u8, reply: []u8, sender: u32) usize {
if (message.len < protocol.child_added_size) return 0;
const report = std.mem.bytesToValue(protocol.ChildAdded, message[0..protocol.child_added_size]);
var status: i32 = 0;
if (driverByProcess(sender)) |driver| {
if (!addChild(report.parent, report.bus_address, report.identity, report.device_id, sender)) status = -1;
writeLine("device-manager: child added (device {d} port {d}, identity {d}) by {s}\n", .{ report.parent, report.bus_address, report.identity, driver.name() });
if (status == 0) publishEvent(message[0..protocol.child_added_size]);
// Matching from reports (M19.3): a registered child whose identity
// names a driver gets one, once — re-reports after a bus restart
// dedupe on the registered id, exactly like the registrations do.
if (status == 0 and report.device_id != protocol.no_device) {
if (pciDriverForIdentity(report.identity)) |child_driver| {
if (!driverForDevice(report.device_id)) addDriver(child_driver, report.device_id, true);
}
// ACPI _HID match (M20.3): ps2-bus is a singleton that finds its own
// devices by hid, so spawn it once, without a device assignment.
const hid_len = std.mem.indexOfScalar(u8, &report.hid, 0) orelse report.hid.len;
if (hid_len != 0) {
if (hidDriverFor(report.hid[0..hid_len])) |hid_driver| {
if (!alreadySupervised(hid_driver)) addDriver(hid_driver, protocol.no_device, false);
}
}
}
} else {
status = -1;
}
const report_reply = protocol.ReportReply{ .status = status };
@memcpy(reply[0..@sizeOf(protocol.ReportReply)], std.mem.asBytes(&report_reply));
if (test_pci_restart_mode and !test_usb_killed) {
if (driverByProcess(sender)) |driver| {
if (std.mem.eql(u8, driver.name(), "pci-bus") and childCountOf(sender) >= 3) {
// The pci restart drill: kill the enumerator after it has
// reported; the respawn must re-register without duplicates
// (M19.0 idempotence, proven end to end by pci-scan).
test_usb_killed = true;
test_kill_pid = sender;
test_kill_due_ns = system.clock() + 1_000_000_000;
_ = system.timerOnce(manager_endpoint, 1100);
}
}
}
if (test_usb_restart_mode and !test_usb_killed and childCountOf(sender) >= 2) {
// Only the xHCI reporter is the drill's victim — pci-bus also reports
// now, and whichever finishes second must not trigger the kill.
if (driverByProcess(sender)) |driver| {
if (std.mem.eql(u8, driver.name(), "usb-xhci-bus")) {
// Delayed, not immediate: the device-list scenario's subscriber
// needs a window to enumerate and subscribe before the events.
test_usb_killed = true;
test_kill_pid = sender;
test_kill_due_ns = system.clock() + 2_000_000_000;
_ = system.timerOnce(manager_endpoint, 2100);
}
}
}
return @sizeOf(protocol.ReportReply);
}
/// A bus driver reported a device gone (hot-unplug; no sender exists yet, but
/// the handler is protocol-complete — death-pruning covers removal until then).
fn onChildRemoved(message: []const u8, reply: []u8, sender: u32) usize {
if (message.len < protocol.child_removed_size) return 0;
const report = std.mem.bytesToValue(protocol.ChildRemoved, message[0..protocol.child_removed_size]);
var status: i32 = -1;
for (&children) |*child| {
if (child.used and child.parent == report.parent and child.bus_address == report.bus_address and child.reporter == sender) {
writeLine("device-manager: child removed (device {d} port {d})\n", .{ child.parent, child.bus_address });
child.used = false;
status = 0;
}
}
const report_reply = protocol.ReportReply{ .status = status };
@memcpy(reply[0..@sizeOf(protocol.ReportReply)], std.mem.asBytes(&report_reply));
return @sizeOf(protocol.ReportReply);
}
/// An application asked for the tree: the mirror, as a header plus entries.
fn onEnumerate(reply: []u8) usize {
var count: u32 = 0;
var offset: usize = @sizeOf(protocol.EnumerateReply);
for (&children) |*child| {
if (!child.used) continue;
if (offset + @sizeOf(protocol.ChildEntry) > reply.len) break;
const entry = protocol.ChildEntry{ .parent = child.parent, .bus_address = child.bus_address, .identity = child.identity };
@memcpy(reply[offset..][0..@sizeOf(protocol.ChildEntry)], std.mem.asBytes(&entry));
offset += @sizeOf(protocol.ChildEntry);
count += 1;
}
const header = protocol.EnumerateReply{ .status = 0, .count = count };
@memcpy(reply[0..@sizeOf(protocol.EnumerateReply)], std.mem.asBytes(&header));
return offset;
}
/// An application subscribed: its endpoint arrived as the call's capability.
fn onSubscribe(reply: []u8, capability: ?runtime.ipc.Handle) usize {
var status: i32 = -1;
if (capability) |handle| {
for (&subscribers) |*slot| {
if (slot.* == null) {
slot.* = handle;
status = 0;
break;
}
}
}
const report_reply = protocol.ReportReply{ .status = status };
@memcpy(reply[0..@sizeOf(protocol.ReportReply)], std.mem.asBytes(&report_reply));
return @sizeOf(protocol.ReportReply);
}
fn onNotification(badge: u64) void {
if (badge & runtime.ipc.notify_exit_bit != 0) {
const dead: u32 = @intCast(badge & ~(runtime.ipc.notify_badge_bit | runtime.ipc.notify_exit_bit));
@@ -293,6 +529,8 @@ fn onNotification(badge: u64) void {
pub fn main(init: runtime.process.Init) void {
if (init.arguments.get(1)) |mode| {
test_restart_mode = std.mem.eql(u8, mode, "test-restart");
test_usb_restart_mode = std.mem.eql(u8, mode, "test-usb-restart");
test_pci_restart_mode = std.mem.eql(u8, mode, "test-pci-restart");
}
runtime.service.run(protocol.message_maximum, .{
.service = .device_manager,
+35
View File
@@ -0,0 +1,35 @@
//! /system/services/fdt — the devicetree discovery service: the ARM twin of the
//! acpi service (docs/m19-m20-plan.md decision 7). **Placeholder: not
//! implemented.** It exists so the build's `-Ddiscovery` option has both of its
//! values from day one; the implementation lands with the Raspberry Pi
//! bring-up (docs/arm.md).
//!
//! What it becomes: the per-firmware discoverer for boots that hand over a
//! flattened device tree instead of ACPI tables. It claims the
//! `devicetree-blob` node the kernel publishes (the FDT the loader received),
//! walks the tree — pure data, no bytecode, so unlike the acpi service it
//! needs no port grant and no interpreter — and, like any bus-shaped driver:
//! `device_register`s what it finds (containment against the blob node's
//! recorded apertures), reports each child to the device manager
//! (`child_added`, identity = the node's `compatible` string), and stays
//! resident under the manager's supervision (hello, restart, the usual
//! contract).
//!
//! Known prerequisite recorded in the plan: `DeviceDescriptor`'s 8-byte `hid`
//! cannot hold an FDT `compatible` string ("brcm,bcm2835-aux-uart") — identity
//! widens before this file grows a body.
const runtime = @import("runtime");
pub fn main(init: runtime.process.Init) void {
_ = init;
// Not implemented: exit cleanly and silently (a bare spawn by the
// initial-ramdisk sweep must not derange other tests' markers). The
// supervisor reads a clean exit as "meant to stop" — correct for a
// placeholder.
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
@@ -51,8 +51,9 @@ fn awaitChildExit(endpoint: runtime.ipc.Handle) u32 {
/// The harness-run child of the signals test: echoes requests, logs the two
/// signals it handles. Terminate makes run() return, and returning from main is
/// the clean exit the parent reads as ExitReason.exited.
fn echo(message: []const u8, reply: []u8, sender: u32) usize {
fn echo(message: []const u8, reply: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = sender;
_ = capability;
const n = @min(message.len, reply.len);
@memcpy(reply[0..n], message[0..n]);
return n;
+2 -1
View File
@@ -91,7 +91,8 @@ fn releaseClientHandles(client: u32) void {
}
/// Handle one request from `sender`; write the reply into `out`, return its length.
fn handle(message: []const u8, out: []u8, sender: u32) usize {
fn handle(message: []const u8, out: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = capability;
if (message.len < protocol.request_size) return fail(out);
const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]);
const payload = message[protocol.request_size..];
+75 -2
View File
@@ -168,7 +168,7 @@ CASES = [
# Stress the big kernel lock across cores; heavier, so a longer timeout.
{"name": "smp-stress",
"smp": 4,
"timeout": 90,
"timeout": 150,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Retry: a forced first-wake failure must still bring every core online.
@@ -258,12 +258,79 @@ CASES = [
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M18.2: bus tree reports — the xHCI driver scans its root-hub ports and
# reports both QEMU devices; the manager mirrors, prunes on the reporter's
# death, and the respawned driver re-reports (docs/device-manager.md).
{"name": "usb-report",
"smp": 4,
"timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"device-manager: child added[\s\S]*"
r"device-manager: child added[\s\S]*"
r"device-manager: test mode: killing the reporter[\s\S]*"
r"device-manager: child removed[\s\S]*"
r"device-manager: restarting usb-xhci-bus[\s\S]*"
r"device-manager: child added",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M20.1: the ring-3 AML parse (the acpi service maps the blobs and parses
# them) finds exactly the Device count the kernel's own parse produced.
{"name": "acpi-parse",
"smp": 4,
"timeout": 60,
"expect": r"acpi-parse: ok",
"fail": r"acpi-parse: mismatch|DANOS-TEST-RESULT: FAIL"},
# M20.3: the flip — ps2-bus now comes up from the acpi service's report, not
# a kernel-built node. Ordered: report -> spawn -> the driver attaches its
# keyboard, proving discovery runs entirely in ring 3 (docs/m19-m20-plan.md).
{"name": "acpi-ps2",
"smp": 4,
"timeout": 150,
"expect": r"acpi: reported PNP0303[\s\S]*"
r"device-manager: spawned ps2-bus[\s\S]*"
r"ps2-bus: keyboard driver attached",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M20.2: the acpi service evaluates _CRS/_STA in ring 3 and registers +
# reports its _HID devices — the two PS/2 nodes must appear with resources
# (keyboard: io 0x60/0x64 + IRQ = 3; mouse: IRQ = 1) (docs/m19-m20-plan.md).
{"name": "acpi-report",
"smp": 4,
"timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"acpi: reported PNP0303 \(device \d+, 3 resources\)[\s\S]*"
r"acpi: reported PNP0F13 \(device \d+, 1 resources\)",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M19.1: the ring-3 PCI scan (pci-bus walks the ECAM through its mmio_map
# grant) finds exactly the functions the kernel's own walk recorded.
{"name": "pci-scan",
"smp": 4,
"timeout": 60,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M18.3: the application surface — device-list enumerates the tree over IPC,
# subscribes (endpoint as capability), and observes the removed/added events
# the reporter's test-kill produces (docs/device-manager.md).
{"name": "device-list",
"smp": 4,
"timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"device-list: \d+ devices[\s\S]*"
r"device-list: subscribed[\s\S]*"
r"device-manager: test mode: killing the reporter[\s\S]*"
r"device-list: removed \(device[\s\S]*"
r"device-list: added \(device",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M18.1: the device manager's hello + restart policy — xHCI hellos clean and
# stays; crash-test faults, is restarted with backoff (re-claiming its device
# each time), and hits the crash-loop cap (docs/device-manager.md).
{"name": "driver-restart",
"smp": 4,
"timeout": 90,
"timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
@@ -441,6 +508,12 @@ def main():
for case in selected:
print(f" {case['name']:<12} ... ", end="", flush=True)
ok, detail = run_case(arch, case)
if not ok:
# Keep the evidence: serial.log is otherwise overwritten by the
# next case, and an intermittent failure's log is unrecoverable.
source = os.path.join(WORK, "serial.log")
if os.path.exists(source):
shutil.copy(source, os.path.join(WORK, f"{case['name']}-failed-serial.log"))
print(("PASS" if ok else "FAIL") + f" ({detail})")
if not ok:
failures += 1