64 Commits
Author SHA1 Message Date
Daniel Samson dd22bfbc48 Merge feat/power-events: ACPI events + system power (M21)
The QMP harness channel, the SCI + power button published from ring 3,
Notify/GPE dispatch in the AML interpreter, and orderly shutdown — init's
M17 stop cascade into a ring-3 S5 write. Proven by injecting a real ACPI
power-button event; QEMU powers off through the whole chain.

# Conflicts:
#	system/services/init/init.zig
2026-07-13 06:27:18 +01:00
Daniel Samson 6e60daed6a Fix the drviers typo and make tests robust to source-path debug prefixes
The debug-message refactor prefixed each service/driver line with its
source path (system/drivers/hpet:, ...) for readability, but two things
left main red: a 'drviers' typo in hpet.zig and pci-bus.zig, and five
kernel tests (init, hpet, pci-scan, device-manager, vfs-client-death)
that starts-with-matched the old short markers, which no longer sit at
the front of the prefixed lines.

Fix the typo, and convert the fragile starts-with matchers to substring
matching via a bufferHas helper — 'hpet: ok' now matches inside
'system/drivers/hpet: ok' regardless of prefix. Future-proof against
further prefix changes and harmless for the tests that already passed.
Suite 58/58.
2026-07-13 06:23:21 +01:00
Daniel Samson 446f655c69 Docs: close the M21 track (events + system power)
docs/m19-m20-plan.md's M21 preview now points at the completed plan.
2026-07-13 05:57:17 +01:00
Daniel Samson a785efa4a3 Orderly shutdown: init's stop cascade into ring-3 S5 (M21.3)
The capstone. init becomes a real supervisor: it spawns its boot
services supervised against one endpoint that also carries its signals, a
re-arming heartbeat timer, and the power events it subscribes to. On the
power button (or a terminate signal — same path) it logs the shutdown,
runs the M17 stop sequence over its children in reverse spawn order
(vfs last), then asks the power service for S5.

The acpi service honors a shutdown request from a power subscriber — init
is the one subscriber, a soft gate that stands in for 'only the system
supervisor may power off' and, unlike a PID-1 check, survives the test
harness where the kernel's idle tasks take the early ids. The power
service is mechanism (write S5); deciding when to shut down and stopping
everything else first is init's policy — the microkernel split applied to
poweroff.

The orderly-shutdown scenario injects a real QMP power-button event and
watches the whole chain compose: button pressed -> init shutting down ->
entering S5 -> QEMU powers off. That single scenario proves the M17
lifecycle and the M21 event side compose into a clean shutdown. Suite
60/60.
2026-07-13 05:56:58 +01:00
Daniel Samson 767a2a9a7c Notify dispatch and GPE handlers (M21.2)
The AML interpreter now handles the Notify opcode (0x86, previously
unhandled): it resolves the target device, evaluates the code, and
records the pair in a bounded per-evaluate queue the caller drains with
takeNotifications. A host unit test with hand-encoded AML — a method that
issues Notify(DEV_, 0x80) — proves the device and code come back; aml.zig
joins the zig build test loop so the interpreter is covered on the host.

The acpi service's SCI handler now services general-purpose events too:
for each set-and-enabled GPE bit it evaluates the \_GPE._Lxx (level) or
_Exx (edge) handler method, drains the Notify queue that produced, and
publishes a domain event per notified device — PNP0C0A battery, ACPI0003
AC, PNP0C0D lid, else generic notify — then clears the status bit and
acks. The embedded controller's _Qxx queries are out of scope (hardware
track). QEMU raises no GPEs on this config, so the QEMU suite is the
regression net (the power button still works with GPE servicing in the
path); correctness is the unit test. Suite 59/59.
2026-07-13 05:44:58 +01:00
Daniel Samson dd044fb115 fix / debug 2026-07-13 05:41:57 +01:00
Daniel Samson 3ec14509a0 fix / debug 2026-07-13 05:41:05 +01:00
Daniel Samson 1f2c60b3ec The power button, in ring 3: SCI bound, fixed event published (M21.1)
The kernel publishes the FADT as one more acpi-tables memory resource
(tagged by its intact FACP header — the AML blobs are header-stripped);
the acpi service reads the PM1 event/control and GPE register ports from
that copy, so the kernel's own FADT parse is untouched. A power-protocol
module (ServiceId.power = 5, domain-named so an ARM PSCI service can serve
the same id) carries subscribe / shutdown / events.

The acpi service converts to runtime.service.run — device discovery, the
.power protocol, and the SCI notification all fold into one loop. At
startup it enables ACPI mode if SCI_EN is clear (the SMI dance), binds the
SCI (found as the node's len-1 irq resource, distinct from the broad
window), and sets PWRBTN_EN. On the SCI it reads PM1_STS, clears
PWRBTN_STS write-1, logs the press, publishes power_button to
subscribers, and always acks. The power-button scenario proves it with a
real QMP system_powerdown injected mid-run through the M21.0 channel.
2026-07-13 05:38:03 +01:00
Daniel Samson d71a5f25d3 fix kernel: debug 2026-07-13 05:37:35 +01:00
Daniel Samson 2a0f17ae86 fix kernel: debug 2026-07-13 05:35:29 +01:00
Daniel Samson 9ef61a0844 fix kernel: debug prefix 2026-07-13 05:32:26 +01:00
Daniel Samson 688b9101e8 fix kernel: debug prefix 2026-07-13 05:31:36 +01:00
Daniel Samson 1d7ba814dc fix efi: debug prefix 2026-07-13 05:30:56 +01:00
Daniel Samson 8aba86b4ce fix vfs: debug prefix 2026-07-13 05:24:47 +01:00
Daniel Samson a0c83f4b3f fix input: debug prefix 2026-07-13 05:24:16 +01:00
Daniel Samson 1ea48ed5d6 fix init: debug prefix 2026-07-13 05:23:07 +01:00
Daniel Samson dfc7d6a609 The harness grows a QMP channel (M21.0)
Every case now gets a -qmp unix socket (additive; no case notices). A
minimal client does the capabilities handshake and executes one command;
the per-case qmp_after hook sends it N seconds after boot, retrying until
the guest's socket is up. A case with a hook configured cannot pass until
the hook delivered — and the smoke case now carries a harmless
query-status hook, so the channel is proven end to end on every run.
This is how the power scenarios inject the real ACPI power-button event
(system_powerdown) in M21.1 and M21.3.
2026-07-13 05:22:53 +01:00
Daniel Samson 77a3ccd33d fix hpet: debug prefix 2026-07-13 05:22:27 +01:00
Daniel Samson 07da27dc39 fix pci-bus: debug prefix 2026-07-13 05:21:47 +01:00
Daniel Samson d89657d0a4 fix device-manager: debug prefix 2026-07-13 05:21:08 +01:00
Daniel Samson 849b4b62d4 fix acpi: debug prefix 2026-07-13 05:20:31 +01:00
Daniel Samson 8589bf713b fix usb-xhci-bus debug prefix 2026-07-13 05:18:49 +01:00
Daniel Samson 738f6aa697 Make sort-lines-group-by-start.sh a runnable script
It was a bare awk snippet starting with `|`, meant to be pasted into a
pipeline. Turn it into an executable script that takes the log file as an
argument (tools/sort-lines-group-by-start.sh filename.log) and document its
behaviour and usage in a header comment.
2026-07-13 05:15:56 +01:00
Daniel Samson 01e56e3f36 Plan M21: ACPI events + system power 2026-07-13 05:13:51 +01:00
Daniel Samson d5d15cefcb Decode PCI/ACPI device identities and name their class codes as enums
Two related changes to make device identities legible in the boot log and in
the code that matches on them.

Logging: the pci-bus driver decodes each function's class/subclass/prog-IF
triple to human names (via the existing pci-class module), and the acpi
service appends each _HID's human name (via acpi-ids) to its report line. So
"class 0x01 (Mass Storage Controller) subclass 0x06 (Serial ATA Controller)
progif 0x01 (AHCI 1.0)" reads straight off the log when writing a driver.

Naming: a new coding standard ("Named values, not magic numbers") says a value
with meaning gets a name, prefer an enum for value sets. Applied:
- pci-class is refactored from u8-switch tables into a BaseClass enum plus
  per-class SubClass/ProgIf enums with name() methods (the usb-ids shape). The
  public className/subclassName/progIfName(u8...) API is unchanged, so the
  hardware-byte decoders (pci-bus, the kernel dump) are untouched; output is
  byte-identical.
- the device-manager builds the xHCI class triple from named parts instead of
  a bare 0x0C0330.
- the acpi service's _CRS walk names its resource-descriptor tags as
  SmallResourceType/LargeResourceType enums, and the _HID integer decode uses
  the AML module's existing *_opcode constants (now re-exported from aml.zig)
  rather than bare 0x0A/0xFF/... literals.
2026-07-13 05:05:25 +01:00
Daniel Samson fd96a35eb9 Decode the xHCI port speed in the usb-xhci-bus log
The root-hub scan logged the raw PORTSC port-speed class ("speed class 3").
Decode it to a human name — Low/Full/High/SuperSpeed/SuperSpeedPlus with the
USB generation and line rate — so the boot log says what enumerated on each
port, the USB analog of the pci-bus class line. This is the link speed only;
the device class/subclass/protocol needs descriptor reads (the USB track).
2026-07-13 05:05:13 +01:00
Daniel Samson e3fe3f3f45 Boot zig-out directly in the qemu test harness
The FHS-shaped zig-out IS the boot volume (docs/efi.md), and `zig build
run-x86-64` already presents it to the guest with fat:rw:zig-out. The test
harness instead assembled a separate ESP by copying the boot-critical files
out of zig-out into zig-out/qemu-test/esp — but every (dest, src) pair was
identical, so the copy was pure redundancy.

Drop make_esp and point QEMU straight at zig-out, matching run-x86-64 and the
docs. Removes the now-dead efi_app/kernel/extra arch-config entries.
2026-07-13 05:05:08 +01:00
Daniel Samson 60da667b42 Merge claude/vigilant-swanson-073c72: retire dead kernel AML device-building path (M20.3 cleanup) 2026-07-13 03:54:06 +01:00
Daniel Samson 36145e623b Delete the retired kernel AML device-building path (M20.3 cleanup)
The M20.3 flip moved ACPI namespace enumeration to the ring-3 acpi
service; the kernel now builds the namespace only for the \_S5 sleep
type. That left the kernel's AML-to-device helpers unreferenced.

Remove the dead cluster (wireAcpiDevices, mirrorDevices, applyHid,
setEisaHid, applyCrs, parseResourceTemplate, parseAddressSpace,
devicePresent, matchHostBridge, findPciNode, readAdr, isPciRootNode,
isPciRootHid, PciContext) and every AML-decoding helper it alone used
(eisaIdToStr, seg4, cstr, hexDigit, rd16, rd32, readN, readLE,
readIntObj, packageLength/PkgLen) plus their tests and the now-orphaned
acpi-ids import. The static-table path keeps checksumOk, fadt, readGas,
readCntRegister, and rd. Also tidies two stale comments.
2026-07-13 03:53:04 +01:00
Daniel Samson 565415327d Mark the M19-M20 discovery migration complete 2026-07-13 03:33:00 +01:00
Daniel Samson bf6bdb389d Merge feat/acpi-service: ACPI interpretation in ring 3 (M20)
The AML interpreter as a shared build module, the acpi-tables node, the
acpi service (parse, evaluate _CRS/_STA, register + report), and the flip
that retired the kernel's ACPI device build — discovery's second and final
subsystem to leave ring 0.
2026-07-13 03:32:54 +01:00
Daniel Samson 0628944b15 Docs: close the discovery migration (M20.3)
discovery.md records ACPI enumeration leaving the kernel; device-manager.md
increment 8 marked done — enumeration now runs entirely in ring 3.
2026-07-13 03:32:53 +01:00
Daniel Samson e6d0bb7ef0 The flip: ACPI enumeration leaves the kernel (M20.3)
The kernel no longer folds AML Device objects into the device tree — the
ring-3 acpi service is the sole builder of _HID device nodes. The kernel
keeps building the namespace only for the \_S5 sleep type, and still
seeds the static tables (MADT, HPET, MCFG, FADT) and the acpi-tables node.

The device manager matches ps2-bus from the service's _HID reports
(PNP0303 / PNP0F13, singleton-deduped) instead of boot-snapshot nodes;
its dead boot-snapshot ps2 arm is gone. The service registers every
device before reporting any, so a driver the manager spawns on the first
report already sees the full set — no keyboard-before-mouse race. The
acpi-ps2 scenario proves the whole chain: report -> spawn -> ps2-bus
finds the controller and attaches its keyboard, entirely in ring 3. The
ioport test moved to the acpi-tables I/O window, since the kernel-built
PS/2 node it used to scan for no longer exists. The retired
device-building functions in acpi.zig are dead but retained (a botched
mechanical deletion is worse mid-migration than a follow-up sweep, which
is flagged as a task). Suite 58/58.
2026-07-13 03:32:15 +01:00
Daniel Samson 5ca804d827 The acpi service evaluates _CRS/_STA in ring 3 and reports devices (M20.2)
AML method evaluation now runs in userspace touching real hardware: the
service builds an interpreter with a ring-3 Hal (port I/O routed through
its claimed acpi-tables node; a scratch page backs SystemMemory maps so a
stray OperationRegion degrades to zeros instead of faulting a process
that cannot map arbitrary physical memory). It walks the namespace and,
for each present _HID device that is not a PCI root, evaluates _CRS,
registers it under acpi-tables, and reports it with its EISA-decoded hid.

Containment for this needed the broker's irq check to become range-based
— an interrupt line is still indivisible, but a parent may own a range,
so the acpi-tables node's broad irq window contains its children's legacy
lines (a length-1 range is exactly the old equality, so single-irq
parents are unaffected). ChildAdded gained a hid field for firmware
string identity. Matching those reports to drivers stays off until M20.3,
so ps2-bus still comes up via the kernel path — no regression. The
acpi-report scenario proves the PS/2 keyboard (io 0x60/0x64 + IRQ) and
mouse (IRQ) are reported with their resources. Suite 57/57.
2026-07-13 03:19:39 +01:00
Daniel Samson a299363b59 The AML interpreter runs in ring 3: the acpi service parses (M20.1)
The AML module becomes a build module compiled into both the kernel (for
the \_S5 sleep state it still needs) and the new acpi service — one
source, two builds, no fork. The kernel publishes a single acpi-tables
node: the DSDT/SSDT blobs as memory resources, a broad io_port grant (the
honest trust boundary — firmware AML names whatever ports it chose, known
only after parsing), and the SCI for the M21 event track. The acpi
service claims the node, maps each blob through the ordinary mmio grant
(which preserves the sub-page offset onto the bytecode), and runs the
same parser the kernel does. It self-verifies its namespace Device count
against the kernel's — 34 = 34 — deterministically via an argv the
acpi-parse test passes, so no racing the shared serial buffer. Parse-only
touches no hardware; OperationRegion evaluation waits for _CRS/_STA in
M20.2. The manager spawns 'discovery' (the neutral ramdisk name) at
startup. Suite 56/56.
2026-07-13 03:07:11 +01:00
Daniel Samson d8dd62c639 Mark the feat/pci-bus merge done in the M19-M20 plan 2026-07-13 02:55:03 +01:00
Daniel Samson d106b6e8dc Merge feat/pci-bus: PCI enumeration in ring 3 (M19)
The host bridge apertures, idempotent device_register, the pci-bus driver
(scan, register, report), and the flip that retired the kernel's PCI walk
— discovery's first subsystem to leave ring 0.
2026-07-13 02:55:03 +01:00
Daniel Samson af2c766f42 The flip: PCI enumeration leaves the kernel (M19.3)
enumeratePci, addBars, pciConfigurationPtr, and the PciHeader struct are
deleted; the kernel seeds only the host bridge, and the ring-3 pci-bus
driver's reports are the sole source of PCI function nodes. The manager
matches PCI drivers from reported identity, deduped by registered device
id so a bus restart never double-spawns.

The flip did its job by exposing a latent SMP race: ring-3
device_register made the broker table concurrent for the first time, and
mmio_map read it lock-free — under load a torn resource length mapped
hpet's window wrong (its user fault) and underflowed r.len-1 into a
kernel integer-overflow panic. Fixed: the broker read in mmio_map (and
claim) runs under the big kernel lock, the arithmetic rejects
zero-length and wrapping windows cleanly, and pci-bus no longer registers
unimplemented size-0 BARs. driver-restart hammered 6x, suite 55/55.
2026-07-13 02:54:50 +01:00
Daniel Samson d26262bf56 pci-bus registers and reports what it scans (M19.2)
Each function is registered under the bridge with the config-space slice
and BARs sized by the same all-ones probe the kernel uses — byte-for-byte
equal descriptors, so the idempotent register returns the kernel's
existing node ids during coexistence instead of duplicating the tree.
The bridge gained the 16-bit io_port aperture that functions' I/O BARs
need to pass containment. Reports carry the registered device_id, and
the pci-scan scenario drills a forced restart: kill the enumerator after
its reports, watch the respawn re-scan, and assert the broker's PCI node
count never grew. The usb-restart test trigger is pinned to the xHCI
reporter (pci-bus racing it to two reports used to steal the kill).
Harness hardening: failing cases preserve their serial logs; the heavy
scenarios run at 150s.
2026-07-13 02:22:19 +01:00
Daniel Samson 10b89c06ff The PCI scan from ring 3: pci-bus walks the ECAM it mapped (M19.1)
The manager matches the pci_host_bridge node and spawns pci-bus with the
bridge id as its assignment — hello, supervision, restart, all the M18
contract for free. The driver claims the bridge, maps the ECAM window
(resource 0) through the ordinary mmio grant, and repeats the kernel's
brute-force bus/device/function walk from user space. The pci-scan
scenario builds its expected marker from the kernel's own function count,
so the two enumerations must agree exactly — the equivalence that
licenses retiring the kernel walk in M19.3.
2026-07-13 02:05:32 +01:00
Daniel Samson a2a05d0b3d Discovery-migration prerequisites (M19.0)
The host bridge now carries MMIO apertures derived from the boot memory
map's gaps below 4 GiB (largest three, sort-merged; a single after-the-
last-region hole dies on OVMF's flash at the top) plus one aperture above
the described space — so a user-space device_register of PCI functions
with BAR resources can pass containment. The discovery test asserts
every PCI memory resource lies inside a bridge window and names any
escapee. device_register is idempotent on exact (parent, class, identity,
resources) match — a restarted registering bus cannot duplicate its
children; proven directly against the broker in the bus test.
ChildAdded gains device_id so a report can carry the registered kernel
id a matched driver needs as its assignment.
2026-07-13 01:59:32 +01:00
Daniel Samson 75d62660b0 Bring the docs up to the M17-M18 reality; pre-settle M19-M20 ambiguities
resilience.md: steps 1-4 of the ladder are built — supervision, exit
reasons, restart with backoff, crash-loop caps, all proven by scenario;
what remains is scope, not mechanism. README statuses follow. drivers.md
gains the driver-contract section (harness, hello, crash-freely). The
M19-M20 plan pre-settles three things the loop would otherwise have had
to decide alone: the memory-map pass-through for apertures, the manager
spawning 'discovery' from M20.1, and hid[8] riding ChildAdded for ACPI
string identity until the FDT widening.
2026-07-13 01:49:52 +01:00
Daniel Samson a53c2b0193 Placeholder discovery services and the -Ddiscovery build option
system/services/acpi and system/services/fdt exist as documented
placeholders (silent clean-exit mains; the headers say exactly what each
becomes and why). The build's -Ddiscovery=acpi|fdt option fills the
ramdisk's neutral 'discovery' slot — the device manager will spawn
"discovery" by that name in M20.3 and never learn which firmware it is
on (m19-m20-plan.md decision 7). x86 defaults to acpi; the aarch64
target flips the default when it lands.
2026-07-13 01:46:04 +01:00
Daniel Samson bf481c080c Record the firmware-neutrality contract as decision 7
Discovery is one swappable process per firmware (acpi service on x86, an
fdt service on the Pis); everything at and above the device-manager
protocol stays generic. The manager owns the tree as data and never
touches hardware — firmware bytecode runs in a crashable, supervised
discoverer. Flagged now: hid[8] cannot hold an FDT compatible string,
and cross-firmware protocols are named by domain (power, not ACPI).
2026-07-13 01:33:12 +01:00
Daniel Samson 3a78dcab3f Scope ACPI events and system power as M21; record the SCI on acpi-tables
Battery, AC, lid, and the power button ride the acpi service as reported
children with small class drivers — the xHCI split repeated. QEMU can
only prove the power-button path (system_powerdown injects the real fixed
event), so battery/EC are interface-complete and hardware-validated on
the laptop. Per-device power states (D-states, suspend/resume) stay out
of scope: suspend has the shape of a lifecycle signal every driver must
answer, and it has no consumer until laptop sleep.
2026-07-13 01:21:23 +01:00
Daniel Samson 470f93a83d Plan the discovery migration (M19 pci-bus, M20 acpi service) 2026-07-13 01:15:17 +01:00
Daniel Samson 7798706b41 Mark the M17-M18 plan complete 2026-07-13 00:49:04 +01:00
Daniel Samson ad40de03c2 Merge feat/usb-xhci-bus: xHCI port scan, tree reports, and the app surface (M18.2-M18.3) 2026-07-13 00:49:04 +01:00
Daniel Samson d8778b4b70 The application surface: enumerate, subscribe, and device-list (M18.3)
Applications ask the device manager for the tree (enumerate: a header
plus ChildEntry records) and subscribe to published add/remove events by
handing their endpoint over as the call's capability — the input-service
pattern; events are the same ChildAdded/ChildRemoved structs the bus
drivers send, one encoding in both directions. device-list is the first
client: it prints the tree, subscribes, and narrates the events through
a driver restart. The protocol's message maximum is capped at the
kernel's IPC MESSAGE_MAXIMUM (256 bytes, ten entries per reply; paging
joins the protocol when a tree outgrows one message). The startUserTask
debug print is gone: it wrote to serial unserialized against user-space
lines and sheared concurrent log markers in half — the root cause of the
scenario flakes.
2026-07-13 00:49:03 +01:00
Daniel Samson 79d859a111 The xHCI driver scans its root-hub ports and reports the tree (M18.2)
child_added/child_removed join the device-manager protocol. The driver
maps its register BAR (resource 0 is the ECAM config space; the walk
starts at 1), reads CAPLENGTH and HCSPARAMS1, and reads one PORTSC per
port: the connect bit and speed class come straight from hardware, no
rings needed to see the devices. The manager mirrors reported children
keyed by (parent, port), remembers which instance reported each, and
prunes a dead reporter's children before deciding the restart — the
children describe protocol state that died with the process. The
usb-report scenario drives the whole loop: two QEMU devices reported,
reporter killed, children pruned, driver respawned with backoff, and the
new instance re-claims, re-scans, and re-reports.
2026-07-13 00:28:29 +01:00
Daniel Samson 37fb09f75e Mark the feat/device-manager merge done in the M17-M18 plan 2026-07-13 00:19:32 +01:00
Daniel Samson 34ebeb968d Merge feat/device-manager: the supervising device manager (M18.1) 2026-07-13 00:19:32 +01:00
Daniel Samson 3cc1d38dd0 The device manager supervises: hello, backoff, and the crash-loop cap (M18.1)
The manager is now a harness service on the well-known .device_manager
endpoint. Every driver spawns supervised; drivers with an assignment must
hello (device-manager-protocol, versioned) within a deadline enforced by
a timer sweep. Exit reasons drive the restart decision: clean exits stay
down, faults restart with 300/600/1200ms backoff, and three fast deaths
mark a driver failed instead of respawning forever. usb-xhci-bus is the
first conforming driver; the crash-test fixture claims a device, hellos,
and faults on purpose — each respawn re-proving claim release on death
through the manager's own path. maximum_tasks grows 16 -> 32: the
initial-ramdisk sweep (15 binaries at once) was intermittently
overflowing the static pool.
2026-07-13 00:19:30 +01:00
Daniel Samson 36e804b848 Mark the feat/process-lifecycle merge done in the M17-M18 plan 2026-07-12 23:53:50 +01:00
Daniel Samson be83a42d42 Merge feat/process-lifecycle: the process lifecycle (M17.1-M17.4)
Claim release on death, exit reasons, published exit events with the VFS
as first subscriber, signals over IPC with one-shot timers and the
service harness — docs/process-lifecycle.md increments 1-4, all built.
2026-07-12 23:53:50 +01:00
Daniel Samson 650a1b1595 Signals over IPC, one-shot timers, and the service harness (M17.4)
Signals are statements delivered as coalescing notifications to the
endpoint a process nominates with signal_bind — never a hijacked stack,
never a question (liveness is the zero-length ping the harness answers).
process_signal is supervisor-or-self gated, like kill; unbound targets
accumulate a pending mask delivered on bind. timer_bind is the missing
timed wait: a one-shot deadline landing in the same replyWait as
everything else — what stop(), hello deadlines, and restart backoff are
built from. runtime.service.run folds requests, signals, and
notifications into callbacks; the VFS conversion deletes its hand-rolled
loop and gains the whole lifecycle contract. The signals scenario drives
ping, reload, terminate->exited, the timer, and the deaf-child
deadline->killed path from ring 3. docs/process-lifecycle.md increments
1-4 are now as-built.
2026-07-12 23:53:38 +01:00
Daniel Samson d8c55c6f2f Publish exit events to subscribers; the VFS releases dead clients' handles (M17.3)
process_subscribe adds an endpoint to a bounded, ref-counted subscriber
table; every death posts the same badge encoding a supervisor's exit
notification uses, equally late, so subscribers observe a fully-released
child. A dying subscriber's own subscriptions are removed first — it never
hears about itself. The VFS is the first subscriber: open handles now
record their owner and are swept when the owner dies, because a service
must never depend on clients cleaning up after themselves
(docs/process-lifecycle.md). Proven by the vfs-client-death scenario.
2026-07-12 23:41:44 +01:00
Daniel Samson 2ebfb0c3b0 Record and expose how every process ends (M17.2)
The kernel records an ExitReason at all three death sites — clean exit,
fault (classified by vector), and process_kill — into a bounded ring
before the exit notification posts, so a supervisor's query never races
the notice. process_exit_reason is gated by the same supervisor check as
kill; runtime.process.exitReason is the stable interface. This is the
input restart policy reads (docs/process-lifecycle.md iron rule 2).
2026-07-12 23:34:09 +01:00
Daniel Samson 888eaa74e1 Release a dead process's device claims (M17.1)
Every path out of a process (exit, fault, kill) now releases its device
claims alongside its IRQ and MSI bindings, so a restarted driver can claim
its hardware again — the cleanup half of process-lifecycle.md's iron rule 1.
MSI vectors were already swept by irq.releaseOwner; claims were the gap.
The claim-release test proves kill -> release -> re-claim, plus the broker
release in isolation.
2026-07-12 23:23:49 +01:00
Daniel Samson ed76cbbc79 Mark Phase 0 done: baseline QEMU suite green (48/48) 2026-07-12 23:17:20 +01:00
Daniel Samson 140229b88d Rename usb-xhci-libary.zig to usb-xhci-library.zig (naming typo) 2026-07-12 23:13:33 +01:00
Daniel Samson cb2379fd06 Update README.md 2026-07-12 23:12:39 +01:00
Daniel Samson 1665b239b0 Add the status checklist and workflow to the M17-M18 plan 2026-07-12 23:04:43 +01:00
Daniel Samson 70ed0337f8 Merge feat/usb: USB wire ABI, xHCI detection and spawn, M17-M18 design 2026-07-12 22:56:35 +01:00
54 changed files with 5127 additions and 1030 deletions
+26 -7
View File
@@ -2,12 +2,31 @@
Codename: Shodan Codename: Shodan
Version: 1 Version: 1
A small operating system, written from scratch in Zig — a bootloader (`boot/`) A small resilient operating system, written from scratch in Zig.
and a microkernel (`system/kernel/`), sharing a neutral handoff contract (`system/boot-handoff.zig`).
It boots x86-64 via UEFI, and so far has a framebuffer console, a physical frame ## Zen of DanOS:
allocator, its own paging with W^X permissions, interrupt/exception handling, a
LAPIC timer, a kernel heap, a fixed-priority preemptive scheduler, and in-kernel IPC - Resilient Micro-Kernel Architecture.
channels. See [`docs/`](docs/README.md) for how each piece works. - Every process run in an isolated user space not kernel space.
- Processes cannot take down the entire OS with it when they die or is killed
- Stable public runtime library, private OS ABI.
- Keeps a stable runtime for user space processes between OS versions (great for backwards compatibility)
- Allows the underlying OS to be changed without effecting applications
- Provides a boundary to enable compatibility between OS's e.g. POSIX, MUSL etc
- Drivers are just isolated processes in user space.
- Thin binaries that can be restarted like applications.
- Useful during driver development.
- Drivers can claim MMIO / ports
- Driver resources (e.g. IRQ/Port/MMIO) claims are automatically cleaned up if the driver dies or is killed
- Drivers can also hook into the process lifecyle to clean up or reset hardware
- No legacy to deal with
- Zig code uses a clean coding style (Zen of Zig)
- Favor reading code over writing code.
- No magic numbers.
- No shortend names unless its for ABI compatibility or acronyms
- Inter-Process Communication (IPC)
- Publish and subscribe to Asynchronous Messages
- Talk to services and processes synchronously
## Prerequisites ## Prerequisites
@@ -60,7 +79,7 @@ straight into CI.
## Documentation ## Documentation
Design notes explaining the *why* behind the code live in Design notes explaining *why* behind the code live in
[`docs/`](docs/README.md) — start with [`docs/README.md`](docs/README.md). [`docs/`](docs/README.md) — start with [`docs/README.md`](docs/README.md).
## Logo ## Logo
+6 -6
View File
@@ -29,7 +29,7 @@ pub fn main() uefi.Status {
// report the reason (boot services are still up) and park the machine so the // report the reason (boot services are still up) and park the machine so the
// message stays on screen. // message stays on screen.
boot() catch |err| { boot() catch |err| {
log("\r\ndanos: boot failed: "); log("\r\nEFI: boot failed: ");
logBytes(@errorName(err)); logBytes(@errorName(err));
log("\r\n"); log("\r\n");
while (true) asm volatile ("hlt"); while (true) asm volatile ("hlt");
@@ -65,14 +65,14 @@ fn boot() !noreturn {
// Best effort: a volume without /system/services/init still boots (kernel-only). // Best effort: a volume without /system/services/init still boots (kernel-only).
loadInit(bs, &boot_information) catch |err| { loadInit(bs, &boot_information) catch |err| {
log("danos: no /system/services/init ("); log("EFI: no /system/services/init (");
logBytes(@errorName(err)); logBytes(@errorName(err));
log(") - booting without user space\r\n"); log(") - booting without user space\r\n");
}; };
// Best effort: the initial_ramdisk (VFS server + drivers) is optional too. // Best effort: the initial_ramdisk (VFS server + drivers) is optional too.
loadInitialRamdisk(bs, &boot_information) catch |err| { loadInitialRamdisk(bs, &boot_information) catch |err| {
log("danos: no initial_ramdisk ("); log("EFI: no initial_ramdisk (");
logBytes(@errorName(err)); logBytes(@errorName(err));
log(")\r\n"); log(")\r\n");
}; };
@@ -84,7 +84,7 @@ fn boot() !noreturn {
// the map and exiting would invalidate the map key. // the map and exiting would invalidate the map key.
const cr3 = try buildBootstrapTables(bs, &boot_information); const cr3 = try buildBootstrapTables(bs, &boot_information);
log("danos: kernel loaded, exiting boot services\r\n"); log("EFI: kernel loaded, exiting boot services\r\n");
boot_information.memory_map = try exitBootServices(bs); boot_information.memory_map = try exitBootServices(bs);
// Switch onto our tables and jump to the kernel in one uninterruptible step. // Switch onto our tables and jump to the kernel in one uninterruptible step.
@@ -395,7 +395,7 @@ fn loadInit(bs: *uefi.tables.BootServices, boot_information: *BootInformation) !
const image = try loadFile(bs, init_file_name); const image = try loadFile(bs, init_file_name);
boot_information.init_base = @intFromPtr(image.ptr); boot_information.init_base = @intFromPtr(image.ptr);
boot_information.init_len = image.len; boot_information.init_len = image.len;
log("danos: /system/services/init loaded\r\n"); log("EFI: /system/services/init loaded\r\n");
} }
/// Ferry the initial_ramdisk (the VFS server + drivers) to the kernel, same as init. /// Ferry the initial_ramdisk (the VFS server + drivers) to the kernel, same as init.
@@ -403,7 +403,7 @@ fn loadInitialRamdisk(bs: *uefi.tables.BootServices, boot_information: *BootInfo
const image = try loadFile(bs, initial_ramdisk_file_name); const image = try loadFile(bs, initial_ramdisk_file_name);
boot_information.initial_ramdisk_base = @intFromPtr(image.ptr); boot_information.initial_ramdisk_base = @intFromPtr(image.ptr);
boot_information.initial_ramdisk_len = image.len; boot_information.initial_ramdisk_len = image.len;
log("danos: initial_ramdisk loaded\r\n"); log("EFI: initial_ramdisk loaded\r\n");
} }
/// Validate the ELF, copy every PT_LOAD segment to its physical address, and /// Validate the ELF, copy every PT_LOAD segment to its physical address, and
+54
View File
@@ -131,6 +131,13 @@ pub fn build(b: *std.Build) void {
}); });
// ACPI/PnP hardware-ID (_HID) names — the flat analog of pci-class for acpi_device // ACPI/PnP hardware-ID (_HID) names — the flat analog of pci-class for acpi_device
// nodes. Also shared reference data. // nodes. Also shared reference data.
// The AML interpreter, a build module so the ring-3 acpi service can run the
// same parser the kernel does (docs/m19-m20-plan.md decision 1). Pure Zig,
// no kernel imports — one source, two builds.
const aml_module = b.addModule("aml", .{
.root_source_file = b.path("system/devices/aml/aml.zig"),
});
const acpi_ids_module = b.addModule("acpi-ids", .{ const acpi_ids_module = b.addModule("acpi-ids", .{
.root_source_file = b.path("system/devices/acpi-ids.zig"), .root_source_file = b.path("system/devices/acpi-ids.zig"),
}); });
@@ -212,6 +219,19 @@ pub fn build(b: *std.Build) void {
}, },
}); });
// The device-manager protocol: hello + (M18.2) tree reports, exposed as its
// own module like the other protocol modules. Imported through the runtime.
const device_manager_protocol_module = b.addModule("device-manager-protocol", .{
.root_source_file = b.path("system/services/device-manager/device-manager-protocol.zig"),
});
runtime_module.addImport("device-manager-protocol", device_manager_protocol_module);
// The power protocol: system power's domain-named surface (docs/m21-plan.md).
const power_protocol_module = b.addModule("power-protocol", .{
.root_source_file = b.path("system/services/power/protocol.zig"),
});
runtime_module.addImport("power-protocol", power_protocol_module);
// Typed volatile MMIO register access + memory-ordering barriers, for drivers on // Typed volatile MMIO register access + memory-ordering barriers, for drivers on
// top of an mmio_map grant. Depends only on `builtin` (arch-conditional barriers); // top of an mmio_map grant. Depends only on `builtin` (arch-conditional barriers);
// no target set, so it inherits each driver's. See library/mmio/mmio.zig. // no target set, so it inherits each driver's. See library/mmio/mmio.zig.
@@ -332,7 +352,32 @@ pub fn build(b: *std.Build) void {
const ps2_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-keyboard", "system/drivers/ps2-bus/keyboard.zig"); const ps2_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-keyboard", "system/drivers/ps2-bus/keyboard.zig");
const ps2_mouse_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-mouse", "system/drivers/ps2-bus/mouse.zig"); const ps2_mouse_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-mouse", "system/drivers/ps2-bus/mouse.zig");
const usb_xhci_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-xhci-bus", "system/drivers/usb-xhci-bus/usb-xhci-bus.zig"); const usb_xhci_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-xhci-bus", "system/drivers/usb-xhci-bus/usb-xhci-bus.zig");
const pci_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "pci-bus", "system/drivers/pci-bus/pci-bus.zig");
// The PCI bus driver decodes each function's class triple to human names in its
// boot log (class/subclass/prog-IF), so pull in the shared pci-class reference.
pci_bus_exe.root_module.addImport("pci-class", pci_class_module);
// A test fixture, not a real driver: hellos to the device manager, then faults —
// what the driver-restart scenario drives the crash-loop cap with.
const crash_test_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "crash-test", "system/services/crash-test/crash-test.zig");
const device_list_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-list", "system/services/device-list/device-list.zig");
// The discovery service: one swappable process per firmware
// (docs/m19-m20-plan.md decision 7), bundled under the neutral ramdisk name
// "discovery" so the device manager never learns which firmware it is on.
// x86 boots describe hardware with ACPI; the Raspberry Pis hand over a
// flattened device tree — the aarch64 target flips the default when it
// lands (docs/arm.md). Both are placeholders until M20.1 (acpi) and the
// ARM bring-up (fdt).
const Discovery = enum { acpi, fdt };
const discovery = b.option(Discovery, "discovery", "Which discovery service fills the ramdisk's 'discovery' slot (default: acpi)") orelse Discovery.acpi;
const discovery_source: []const u8 = switch (discovery) {
.acpi => "system/services/acpi/acpi.zig",
.fdt => "system/services/fdt/fdt.zig",
};
const discovery_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "discovery", discovery_source);
if (discovery == .acpi) discovery_exe.root_module.addImport("aml", aml_module);
const device_manager_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-manager", "system/services/device-manager/device-manager.zig"); const device_manager_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-manager", "system/services/device-manager/device-manager.zig");
// Names the xHCI PCI class triple from the shared taxonomy instead of a bare 0x0C0330.
device_manager_exe.root_module.addImport("pci-class", pci_class_module);
// The input service and its exercisers: the fan-out server, a hardware-free synthetic // The input service and its exercisers: the fan-out server, a hardware-free synthetic
// source, and a subscriber that doubles as the `input` test's oracle. See docs/input.md. // source, and a subscriber that doubles as the `input` test's oracle. See docs/input.md.
const input_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input", "system/services/input/input.zig"); const input_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input", "system/services/input/input.zig");
@@ -363,6 +408,14 @@ pub fn build(b: *std.Build) void {
mk_run.addFileArg(ps2_mouse_exe.getEmittedBin()); mk_run.addFileArg(ps2_mouse_exe.getEmittedBin());
mk_run.addArg("usb-xhci-bus"); mk_run.addArg("usb-xhci-bus");
mk_run.addFileArg(usb_xhci_bus_exe.getEmittedBin()); mk_run.addFileArg(usb_xhci_bus_exe.getEmittedBin());
mk_run.addArg("pci-bus");
mk_run.addFileArg(pci_bus_exe.getEmittedBin());
mk_run.addArg("crash-test");
mk_run.addFileArg(crash_test_exe.getEmittedBin());
mk_run.addArg("device-list");
mk_run.addFileArg(device_list_exe.getEmittedBin());
mk_run.addArg("discovery");
mk_run.addFileArg(discovery_exe.getEmittedBin());
mk_run.addArg("device-manager"); mk_run.addArg("device-manager");
mk_run.addFileArg(device_manager_exe.getEmittedBin()); mk_run.addFileArg(device_manager_exe.getEmittedBin());
mk_run.addArg("input"); mk_run.addArg("input");
@@ -530,6 +583,7 @@ pub fn build(b: *std.Build) void {
"system/devices/device-abi.zig", "system/devices/device-abi.zig",
"system/devices/pci-class.zig", // class/subclass/prog-IF name decoding "system/devices/pci-class.zig", // class/subclass/prog-IF name decoding
"system/devices/acpi-ids.zig", // _HID name decoding "system/devices/acpi-ids.zig", // _HID name decoding
"system/devices/aml/aml.zig", // AML parse + interpret, incl. Notify dispatch (M21)
"system/devices/usb-abi.zig", // wire sizes + bit packings + set-up packet encodings "system/devices/usb-abi.zig", // wire sizes + bit packings + set-up packet encodings
"system/devices/usb-ids.zig", // class/subclass/protocol code assignments "system/devices/usb-ids.zig", // class/subclass/protocol code assignments
"library/mmio/mmio.zig", // barriers assemble + registers round-trip "library/mmio/mmio.zig", // barriers assemble + registers round-trip
+8 -6
View File
@@ -45,20 +45,22 @@ rather than restate it. Roughly in the order things happen at runtime:
until its hardware interrupts it**. The claim is the capability; `irq_ack` is the until its hardware interrupts it**. The claim is the capability; `irq_ack` is the
unmask. unmask.
14. **[driver-model.md](driver-model.md) — buses, classes and host controllers.** How 14. **[driver-model.md](driver-model.md) — buses, classes and host controllers.** How
real driver stacks factor into three shapes, how families share code, and the real driver stacks factor into three shapes and how families share code. The
proposed ABI for the three primitives still missing (capability passing, DMA + three primitives it proposed are long since built (M13 capability passing,
memory barriers, MSI). M14 DMA + barriers, M15 MSI), and the driver *contract* on top of them —
hello, supervision, restart — is built too (device-manager.md, M18).
15. **[process-management.md](process-management.md) — process management.** The 15. **[process-management.md](process-management.md) — process management.** The
microkernel's `ps`/`kill`/SIGCHLD: enumerate as a table snapshot, the microkernel's `ps`/`kill`/SIGCHLD: enumerate as a table snapshot, the
supervision link as the kill authority, and child-exit notifications over the supervision link as the kill authority, and child-exit notifications over the
same endpoints IRQs arrive on. same endpoints IRQs arrive on.
16. **[process-lifecycle.md](process-lifecycle.md) — the process lifecycle.** Design: 16. **[process-lifecycle.md](process-lifecycle.md) — the process lifecycle.** Built
signals over IPC as the one lifecycle vocabulary every process speaks — the (M17): signals over IPC as the one lifecycle vocabulary every process speaks — the
POSIX.1-1990 words with message delivery instead of stack hijack, the stable POSIX.1-1990 words with message delivery instead of stack hijack, the stable
`runtime.process` interface, exit reasons, published exit events any stateful `runtime.process` interface, exit reasons, published exit events any stateful
service can subscribe to (the VFS releasing dead clients' handles), and the two service can subscribe to (the VFS releasing dead clients' handles), and the two
iron rules (cleanup is the kernel's job; kill is not a signal). iron rules (cleanup is the kernel's job; kill is not a signal).
17. **[device-manager.md](device-manager.md) — the device manager.** Design: the 17. **[device-manager.md](device-manager.md) — the device manager.** Built (M18,
through the app surface): the
tree, the matcher, and the supervisor. Tree structure lives in the manager, tree, the matcher, and the supervisor. Tree structure lives in the manager,
authority stays in the kernel; bus drivers report what they see; drivers are authority stays in the kernel; bus drivers report what they see; drivers are
restarted through the lifecycle vocabulary — the plan that turns restarted through the lifecycle vocabulary — the plan that turns
+26
View File
@@ -142,6 +142,32 @@ conventions above — `snake_case` — because it's an identifier, not a filenam
*directory* (`system/services/init`, `library/runtime`), with the repeated leaf *directory* (`system/services/init`, `library/runtime`), with the repeated leaf
resolving away. See the repository-layout section of [README.md](README.md). resolving away. See the repository-layout section of [README.md](README.md).
## Named values, not magic numbers
The naming rule has a twin: **a value with meaning gets a name, too.** The same
principle drives both — a reader should never have to leave the code to understand it.
An abbreviated *name* forces a reader to guess; a bare *number* forces them worse, out
to a spec or a header or a comment three files away, to learn what the value even *is*.
If `0x0C` is the PCI serial-bus class, the code says `BaseClass.serial_bus`, not `0x0C`;
if `0x04` is the ACPI IRQ resource descriptor, it says `SmallResourceType.irq`, not
`0x04`. The number is an implementation detail of the name — recorded once, where the
name is defined, and never spelled again at a use site.
**Prefer an `enum`** when the values form a set (device classes, AML opcodes, resource
descriptor types, states): the type then also says *which* set a value belongs to, and
the compiler rejects a value from the wrong one. A lone `pub const` with a descriptive
name suffices for a one-off (`const large_descriptor_bit = 0x80`). Reach for the enum
the moment code elsewhere compares against, packs, or produces the value — a packed PCI
class triple is written from named parts (`.serial_bus`, `.usb`, `.xhci`), never as
`0x0C_03_30` under a comment that decodes the bytes.
The exceptions are the numbers that carry no hidden meaning: `0` and `1` as plain zero
and one, an index step, a field width, a bit shift. `x + 1`, `buffer[0]`, and `<< 8`
need no christening — there is nothing to look up. The test is exactly the naming test:
*would a reader have to look this up to know what it means?* If yes, name it. This is
what `opcodes.zig`'s `*_opcode` constants, `acpi-ids`'s `HardwareId`, and `pci-class`'s
class enums already are — reference data defined once and named everywhere it is used.
## Why acronyms are the line ## Why acronyms are the line
Because an acronym has no letters to restore. `MMIO` doesn't become "memory mapped Because an acronym has no letters to restore. `MMIO` doesn't become "memory mapped
+15 -3
View File
@@ -1,6 +1,16 @@
# The device manager # The device manager
**Status: design.** The primitives this builds on are real ([process-management.md](process-management.md): **Status: the protocol and supervision are built** (M18.1, 2026-07-13): `hello`
with its deadline, supervised spawn, restart with backoff, and the crash-loop
cap are in — usb-xhci-bus is the first conforming driver, and the
`driver-restart` scenario proves fault → backoff → re-claim → cap end to end.
Tree reports are built too (M18.2, 2026-07-13): the xHCI driver scans its
root-hub ports and reports each connected device (`child_added`); the manager
mirrors them and prunes a dead reporter's children, and the `usb-report`
scenario proves report → prune → respawn → re-report. The application surface is built (M18.3, 2026-07-13):
`enumerate` and `subscribe` over IPC, with `device-list` as the first client —
the manager is now the one answer to "what devices exist" for applications.
The primitives underneath are real ([process-management.md](process-management.md):
spawn/supervise/kill/exit-notification; [driver-model.md](driver-model.md): the device spawn/supervise/kill/exit-notification; [driver-model.md](driver-model.md): the device
table as a capability system; [drivers.md](drivers.md): claim/map/IRQ), and the first table as a capability system; [drivers.md](drivers.md): claim/map/IRQ), and the first
per-device driver spawn works (the device manager matches the xHCI controller by PCI per-device driver spawn works (the device manager matches the xHCI controller by PCI
@@ -126,8 +136,10 @@ published exit events, signals + `runtime.process`). On top of those:
the mouse and keyboard QEMU already hangs off it. the mouse and keyboard QEMU already hangs off it.
7. **App surface**: `enumerate`/`subscribe` over IPC; `device_enumerate` retreats 7. **App surface**: `enumerate`/`subscribe` over IPC; `device_enumerate` retreats
to a manager-internal seam. to a manager-internal seam.
8. **Discovery migration**: pci-bus driver first, acpi service second, kernel scan 8. **Discovery migration** — DONE (M19–M20, 2026-07-13): pci-bus driver (M19)
retired last. (AML-in-user-space is its own track.) then the acpi service (M20) moved enumeration to ring 3; the kernel seeds
only the host bridge and the acpi-tables node. See
[m19-m20-plan.md](m19-m20-plan.md).
## Settled questions (2026-07-12) ## Settled questions (2026-07-12)
+24
View File
@@ -167,3 +167,27 @@ free; discovery on x86 is partly about *finding* what ARM just tells you.
- [ipc.md](ipc.md) — the channels that interrupts-as-messages and the device manager - [ipc.md](ipc.md) — the channels that interrupts-as-messages and the device manager
will ride on. will ride on.
- [vision.md](vision.md) — why drivers belong in isolated user space at all. - [vision.md](vision.md) — why drivers belong in isolated user space at all.
## Update (M19.3, 2026-07-13): PCI enumeration left the kernel
The kernel now seeds only the `pci_host_bridge` node (ECAM window, MMIO
apertures derived from the memory map's holes, bus range, and the 16-bit I/O
window). The per-function walk moved to the ring-3 `pci-bus` driver
([device-manager.md](device-manager.md)): it claims the bridge, repeats the
ECAM scan through its mmio grant, and `device_register`s what it finds, which
the device manager mirrors and matches. The ACPI namespace walk follows in M20;
the static tables (MADT, HPET, MCFG, FADT + `\\_S5`) stay kernel-side.
## Update (M20.3, 2026-07-13): ACPI enumeration left the kernel too
The kernel no longer folds the AML namespace's Device objects into the device
tree. It still parses the *static* tables (MADT for SMP, HPET for the tick, MCFG
for the host bridge, FADT) and still builds the AML namespace — but only to read
the `\\_S5` sleep type for poweroff. Device discovery is the ring-3 **acpi
service** ([device-manager.md](device-manager.md)): it claims the `acpi-tables`
node the kernel publishes (the AML blobs, a broad io_port grant, the SCI),
re-parses the same blobs with the shared AML module, evaluates `_STA`/`_CRS`,
and registers + reports each `_HID` device — the device manager matches drivers
(ps2-bus) from those reports. With M19's pci-bus driver, discovery now runs
entirely in user space; the kernel seeds only the host bridge and the
acpi-tables node.
+20
View File
@@ -362,3 +362,23 @@ the first DMA driver to protect and test against) and these smaller items:
- **Interrupt priority / threaded IRQ latency.** `notifyFromIsr` enqueues the woken - **Interrupt priority / threaded IRQ latency.** `notifyFromIsr` enqueues the woken
driver but doesn't preempt (`wakeLocked` deliberately leaves that to the caller), so driver but doesn't preempt (`wakeLocked` deliberately leaves that to the caller), so
a woken driver waits for the next scheduling point. a woken driver waits for the next scheduling point.
## The driver contract (M17–M18)
Claiming and mapping is half of being a danos driver; the other half is the
**lifecycle and protocol contract**, and the runtime makes it nearly free:
- Build on `runtime.service.run` — one replyWait loop folding protocol
requests, signals, and notifications into callbacks. The harness answers the
universal zero-length ping and turns `terminate` into a clean exit for you
([process-lifecycle.md](process-lifecycle.md)).
- A driver spawned with an assignment (its device id as argv[1]) sends the
versioned `hello` to the device manager inside the deadline, and a **bus**
driver reports what it discovers with `child_added`
([device-manager.md](device-manager.md); usb-xhci-bus is the reference
implementation).
- Crash freely — that is the design. The kernel releases your claims, IRQ
bindings, and MSI vectors at death; the manager reads your exit reason,
prunes what you reported, restarts you with backoff, and your fresh instance
re-claims and re-reports. Never depend on your own cleanup running
(iron rule 1).
+19
View File
@@ -103,3 +103,22 @@ This is what makes a user-space driver possible at all, and it's the subject of
the oldest (discrete messages, not a coalescing level like the notification ring). the oldest (discrete messages, not a coalescing level like the notification ring).
- **A bounded reply.** `MSG_MAX` is 256 bytes and the copy runs under the big kernel - **A bounded reply.** `MSG_MAX` is 256 bytes and the copy runs under the big kernel
lock; a bulk transfer wants shared pages, not a copy. lock; a bulk transfer wants shared pages, not a copy.
## Lifecycle conventions over IPC (M17)
Three conventions from [process-lifecycle.md](process-lifecycle.md) ride the
notification mechanism:
- **Signals** arrive as notifications on the endpoint a process nominated with
`signal_bind` (`runtime.process.bindSignals`): badge = the signal bit plus the
coalesced pending mask (`runtime.process.signalsFrom` decodes). Statements,
never questions; no payload, no reply.
- **One-shot timers** (`timer_bind`, `runtime.system.timerOnce`) land as a
timer-bit notification — the timed wait: a service arms a deadline and keeps
serving, instead of blocking in sleep.
- **The universal ping**: a **zero-length request is the liveness probe**,
answered with a zero-length reply by the service harness itself
(`runtime.service.run`). No protocol's requests start at length zero, so the
encoding cannot collide, and a wedged service simply fails to answer — which
is the diagnosis. Deep health ("can I reach my hardware?") stays a per-service
protocol message.
+49
View File
@@ -1,5 +1,9 @@
# M17–M18 execution plan: process lifecycle + device manager # M17–M18 execution plan: process lifecycle + device manager
**Archived — completed 2026-07-13** (every item checked; suite ended 54/54).
Kept as the record of how M17–M18 landed; the successor is
[m19-m20-plan.md](m19-m20-plan.md).
The operational plan for building [process-lifecycle.md](process-lifecycle.md) The operational plan for building [process-lifecycle.md](process-lifecycle.md)
(M17) and [device-manager.md](device-manager.md) increments 5–7 (M18). Design is (M17) and [device-manager.md](device-manager.md) increments 5–7 (M18). Design is
settled in those documents; this file is the build order — one phase at a time, settled in those documents; this file is the build order — one phase at a time,
@@ -21,6 +25,51 @@ run the existing QEMU suite green before any new work starts.
**Numbering note:** continues the milestone sequence (driver track ended at M16). **Numbering note:** continues the milestone sequence (driver track ended at M16).
## Status
The loop marks a phase `[x]` in the same commit that lands it. A phase is marked
only when its definition of green holds.
- [x] **Phase 0** — baseline: docs committed, feat/usb merged to main, pushed;
`usb-xhci-libary.zig` renamed to `usb-xhci-library.zig`; existing QEMU
suite green from the worktree (48/48, 2026-07-12).
- [x] **M17.1** — kernel releases claims/MSI on death (claims: `releaseAllOwnedBy`
in the reap; MSI was already swept by `irq.releaseOwner`; `claim-release`
test; suite 49/49)
- [x] **M17.2** — exit reasons (`ExitReason` recorded at exit/fault/kill before
the notification; `process_exit_reason` supervisor-gated;
`runtime.process.exitReason`; kernel + ring-3 assertions; suite 49/49)
- [x] **M17.3** — published exit events + VFS subscriber (`process_subscribe`,
bounded ref-counted table, publish on every death;
`runtime.process.subscribeExits`; VFS handles carry owners and are swept on
the owner's death; `vfs-client-death` test; suite 50/50)
- [x] **M17.4** — signals, timer notifications, `runtime.process`, the service
harness (signal_bind/process_signal + coalescing pending mask; timer_bind
on the tick; bindSignals/signalsFrom/sendSignal/stop + timerOnce;
runtime.service.run with the zero-length ping; VFS converted; `signals`
scenario; suite 51/51)
- [x] **merge** `feat/process-lifecycle` → main, push (merged 2026-07-13)
- [x] **M18.1** — device-manager protocol: hello + restart policy
(device-manager-protocol module; the manager as a harness service:
supervised spawns, hello deadline via timer sweep, restart with
300/600/1200ms backoff, exit reasons deciding restart-vs-stopped,
crash-loop cap; usb-xhci-bus first conforming driver; crash-test fixture
re-proving claim release each respawn; `driver-restart` scenario;
maximum_tasks 16→32 — the sweep was overflowing the pool; suite 52/52)
- [x] **merge** `feat/device-manager` → main, push (merged 2026-07-13)
- [x] **M18.2** — xHCI port scan + tree reports (child_added/child_removed in
the protocol; the manager's child mirror with death-pruning; xHCI maps the
register BAR — resource 0 is ECAM — reads CAPLENGTH/HCSPARAMS1, scans
PORTSC, reports connected ports with speed-class identity; `usb-report`
scenario proves report → prune → respawn → re-report; suite 53/53)
- [x] **M18.3** — app surface: enumerate/subscribe over IPC (subscriber
endpoint rides as the call's capability; events are the same structs the
buses send); device-list first client; protocol capped at the kernel's
IPC MESSAGE_MAXIMUM (256); the startUserTask debug print removed — it
sheared concurrent serial lines and was the scenario-flake root cause;
`device-list` scenario; suite 54/54)
- [x] **merge** `feat/usb-xhci-bus` → main, push (merged 2026-07-13) — **plan complete**
--- ---
## M17.1 — the kernel releases a dead process's claims ## M17.1 — the kernel releases a dead process's claims
+196
View File
@@ -0,0 +1,196 @@
# M19–M20 execution plan: discovery migration
The operational plan for [device-manager.md](device-manager.md)'s increment 8:
discovery leaves the kernel — a **pci-bus driver** (M19) and an **acpi service**
(M20), with the kernel's device enumeration retired behind them. Same rules as
[m17-m18-plan.md](m17-m18-plan.md): one phase at a time, each green before the
next; this file is the build order and the checklist.
**Definition of green, every phase:** `zig build` clean, `zig build test` clean,
`python3 test/qemu_test.py` passes (existing scenarios plus the phase's new
one), and the relevant design doc updated. Commit per green phase (no co-author
trailers). The full suite is the regression net — the existing
`driver-restart` / `usb-report` / `device-list` / `input` scenarios must stay
green *through* the migration, which is the whole point: the system must not be
able to tell who enumerated it.
**Workflow:** dedicated worktree; branches off `main` — `feat/pci-bus`
(M19.0–19.3), `feat/acpi-service` (M20.1–20.3); auto-merge to main when a
branch is green; keep branches; push everything.
## Settled decisions (2026-07-13 — veto before the loop starts)
1. **What "retiring the kernel scan" means.** The kernel keeps, forever, the
parses it needs before user space exists: RSDP/XSDT location, MADT (SMP),
the HPET table (the tick), FADT + the AML `\_S5` evaluation (poweroff — the
power tests prove it), and MCFG (the host bridge node). What retires is
**device enumeration**: the ECAM function walk (M19.3) and the DSDT/SSDT
namespace walk that builds device nodes (M20.3). The AML module stays a
shared build module compiled into both the kernel (for `\_S5`) and the acpi
service (for everything else) — same source, two builds, no fork.
2. **Bridge apertures come from the firmware memory map, not AML.** Registered
PCI functions carry BAR resources, and containment demands the bridge own
windows that cover them. The apertures are derived kernel-side from the
boot memory map's MMIO holes (regions that are neither RAM nor tables) —
mechanical, AML-free, and available at boot regardless of what later moved
to user space. (The bridge today carries only ECAM + bus range; this is the
prerequisite M19.0 exists for.)
3. **`device_register` becomes idempotent on exact match.** A re-registration
with identical (parent, class, resources) returns the existing id instead
of appending. The kernel table has no unregister, so without this a
restarted registering bus would duplicate its children on every respawn —
idempotence makes restart-and-re-report safe for every future bus, not just
PCI.
4. **The manager matches from reports.** `ChildAdded` gains a `device_id`
field (the kernel-registered id, `no_device` for unregistered leaves like
USB ports). After the M19.3 flip, PCI driver matching keys off reported
identity (the class triple) instead of the manager's boot-time snapshot —
the snapshot match remains only for what the kernel still seeds. One flip
phase changes both sides at once so no device is ever matched twice.
5. **The acpi service's authority is one node.** The kernel publishes an
`acpi-tables` device: memory resources covering the table blobs plus a
broad `io_port` resource — the documented trust grant to exactly one
process (AML OperationRegions reach EC/PM ports; the claim-gated
io_read/io_write calls already exist). The service claims it, maps the
tables, and runs the shared AML module in ring 3 behind a `Hal` backed by
`mmio_map` + `io_read`/`io_write`.
6. **Both new processes are protocol drivers** under the manager: hello,
supervision, restart with backoff — all inherited from M18.1 for free.
Registration idempotence (decision 3) is what makes their restarts sound.
7. **Firmware neutrality is the contract** (2026-07-13). The generic layer is
everything at and above the device-manager protocol — descriptors,
containment, reports, matching, supervision — and none of it may become
x86-specific. Discovery is one swappable process per firmware: the acpi
service on x86; an **fdt service** on the Raspberry Pis (claims a
`devicetree-blob` node, reports children from the flattened device tree —
pure data, no bytecode, no port grant, strictly simpler than ACPI). The
manager owns the tree as *data* and touches no hardware, ever — AML runs in
a crashable, supervised discoverer precisely so a firmware-bytecode fault
can never take down the supervisor. Two consequences recorded now:
`DeviceDescriptor`'s 8-byte `hid` cannot hold an FDT `compatible` string
("brcm,bcm2835-aux-uart") — identity widens before the fdt service exists;
and cross-firmware surfaces are named by **domain, not firmware** (M21
defines a *power* protocol, not an "ACPI events" protocol — PSCI/mailbox
sources feed the same subscribers on ARM). **Landed early (2026-07-13):**
both services exist as placeholders (system/services/acpi, system/services/
fdt) and the build's `-Ddiscovery=acpi|fdt` option fills the ramdisk's
neutral `discovery` slot — the manager will spawn "discovery" by that name
in M20.3 and never learn which firmware it is on.
## Status
- [x] **M19.0** — prerequisites (bridge apertures from the memory map's
*gaps* — the single-hole rule died on OVMF's flash at the top of 4 GiB,
caught by the new every-BAR-contained assert in `discovery`; idempotent
`device_register` proven in `bus`; `ChildAdded.device_id`;
m17-m18-plan.md archived; suite 54/54).
- [x] **M19.1** — pci-bus driver, scan only (claims the bridge, maps ECAM
through its grant, brute-force walk with the multifunction rule; the
manager matches pci_host_bridge → pci-bus per device with the full
protocol contract; `pci-scan` builds its expected marker from the
kernel's own count — equivalence on the first run; suite 55/55).
- [x] **M19.2** — register + report (BAR probe mirrored byte-for-byte from the
kernel's addBars so dedupe returns the kernel's node ids during
coexistence; the bridge gained the io_port aperture I/O BARs need;
reports carry the registered device_id; pci-scan drills a forced restart
and asserts the PCI node count never grows — plus harness hardening: a
failing case now preserves its serial as <case>-failed-serial.log, and
the heavy scenarios run at 150s; suite 55/55).
- [x] **M19.3** — the flip: kernel `enumeratePci`/`addBars`/`PciHeader` all
deleted (bridge node stays); manager matches PCI drivers from reported
identity, deduped by registered id. Surfaced and fixed a real SMP race the
flip created — ring-3 device_register made the broker table concurrent, so
mmio_map's lock-free read intermittently tore hpet's resource length
(user fault) and overflowed `r.len-1` into a kernel panic; now the broker
read is under the big lock and the arithmetic is guarded, and pci-bus
skips size-0 BARs. discovery.md updated; suite 55/55 (driver-restart
hammered 6×).
- [x] **merge** `feat/pci-bus` → main, push (merged 2026-07-13).
- [x] **M20.1** — acpi service, parse only: the AML interpreter is now a build
module compiled into both kernel and service; the kernel publishes the
`acpi-tables` node (AML blobs as memory resources, the broad io_port grant,
the SCI); the service claims it, maps the blobs, runs the shared parser in
ring 3, and self-verifies its Device count against the kernel's (34 = 34,
deterministic via argv, no log-scraping); the manager spawns `discovery`
at startup. Parse-only touches no hardware. Suite 56/56.
- [x] **M20.2** — register + report: the service evaluates `_STA`/`_CRS` in
ring 3 (interpreter Hal = port I/O over the claimed node; a scratch page
backs SystemMemory maps so a stray region can't fault it) and registers +
reports each present `_HID` device under `acpi-tables`. Containment: the
broker's irq check became range-based (len-1 == the old equality) so the
node's broad irq window covers children's legacy lines; io ports fall in
the broad io grant. ChildAdded gained `hid`. Matching stays off. The
`acpi-report` scenario asserts the PS/2 keyboard (3 resources) and mouse
(1 resource) among the reports. Suite 57/57.
- [x] **M20.3** — the flip: the kernel's `wireAcpiDevices` call is gone (the
device-building helpers are retained-but-dead pending a focused sweep,
spawned as a task; static tables + `\_S5` + the acpi-tables node stay).
The manager matches ps2-bus from ACPI `_HID` reports; the service
registers all devices before reporting any (no keyboard-before-mouse
race). The `acpi-ps2` scenario proves report → spawn → ps2-bus attaches
its keyboard; `ioport` retargeted to the acpi-tables I/O window (the
kernel-built PS/2 node is gone). Suite 58/58.
- [x] **merge** `feat/acpi-service` → main, push (merged 2026-07-13) — **discovery migration complete**.
---
## Phase notes
**M19.0 apertures:** the boot memory map already crosses the handoff
([boot-handoff]), but discovery never sees it today — expect a small
pass-through (kernel init hands the map to the platform layer) before the
holes computation, which belongs where the bridge node is built
(`parseMcfg`). Sanity-check on QEMU q35: the xHCI BAR (`0xc0000000`-region
values seen in the M18 logs) must land inside a derived aperture, asserted in
the kernel unit test.
**M19.1 scanning without owning config access twice:** the driver reads config
space through its ECAM mmio_map grant of the *bridge* window — the same bytes
the kernel walk read. Vendor-id `0xFFFF` skip, header-type multifunction rule,
no bridge recursion (matches the kernel's current single-segment walk).
**M19.2 BAR sizing:** the classic size probe (write all-ones, read mask,
restore) is deferred — the BARs' current programmed values and types are
enough for containment-checked registration at bring-up; sizing lands with the
first driver that needs to *move* a BAR. Log what is registered so the
scenario can assert it.
**M19.3 what the manager still seeds from the snapshot:** everything the
kernel still enumerates (timers, ACPI nodes until M20.3). The PCI arm of
`pciDriverFor` switches source; `driverFor` doesn't move until M20.3.
**M20.1 spawn and identity (pre-settled 2026-07-13):** the manager spawns
`discovery` by its neutral ramdisk name at startup, as an ordinary protocol
driver (hello, supervision) — from M20.1 on, on every boot. For reporting ACPI
devices, `ChildAdded` gains `hid: [8]u8` (EISA ids fit; zero = none):
firmware *string* identity travels beside the numeric `identity` field until
the FDT-driven widening replaces both (decision 7).
**M20.1 Hal in ring 3:** `mapMmio` → `device.mmioMap` over the claimed
acpi-tables node (plus a table-offset map for blobs); `pioRead`/`pioWrite` →
`device.ioRead`/`ioWrite` against its io_port resource. The interpreter cannot
tell it moved — that is the assertion of `acpi-parse`.
**M20.2 containment for `_CRS`:** io ports fall inside the node's broad
io_port resource; MMIO windows (HPET, LAPIC ranges some firmwares list) fall
inside the memory-map holes added to the node in M20.1. Anything that doesn't
fit is logged and skipped, loudly — bring-up honesty over silent drops.
**M20.3 ps2 ordering:** ps2-bus binds nodes the acpi service now reports, so
its spawn moves behind the report (the manager's matching handles this once
the source flips); the `input` scenario proves the keyboard still types.
**Explicitly out of scope:** PCI bridge recursion (single segment, flat bus
walk stays); BAR reprogramming/sizing; disk/PCIe hotplug; interrupt routing
changes (`_PRT` stays wherever it is today); the USB descriptor track;
multi-segment ECAM; per-device power states (D-states, `_PSx`/`_PRx`,
suspend/resume — a future *lifecycle-vocabulary* extension, since "suspend"
has the shape of a signal every driver must answer, and it has no consumer
until laptop sleep); CPU P/C-states.
## M21 — ACPI events + system power — DONE
Built and merged (docs/m21-plan.md, 2026-07-13): the SCI + power button, Notify/GPE
dispatch, and orderly shutdown (init's stop cascade into a ring-3 S5 write).
See that plan for the phase record.
+139
View File
@@ -0,0 +1,139 @@
# M21 execution plan: ACPI events + system power
The operational plan for the event side of the acpi service and orderly
shutdown — the capstone [m19-m20-plan.md](m19-m20-plan.md) previewed. Same
rules as its predecessors: one phase at a time, each green before the next;
this file is the build order and the checklist.
**Definition of green, every phase:** `zig build` clean, `zig build test`
clean, `python3 test/qemu_test.py` passes (existing scenarios plus the
phase's new one), and the relevant design doc updated. Commit per green phase
(no co-author trailers). Failing cases preserve their serial logs
(`<case>-failed-serial.log`).
**Workflow:** dedicated worktree; branch `feat/power-events` off `main`;
auto-merge to main when the branch is green; keep the branch; push everything.
## Settled decisions (2026-07-13, approved)
1. **S5 is executed by the acpi service from ring 3.** No new syscall: the
broad port grant (M20 decision 5) already made this physically possible —
the service holds the PM1 control ports in its io grant and derives `_S5`
from its own namespace (`aml.sleepState`). Formalizing it adds no
authority. The kernel keeps `power.zig` for its own test paths and
panic-time use.
2. **The power surface is domain-named** (decision 7 of the last plan): a
`power-protocol` module + `ServiceId.power = 5`, registered by the acpi
service — on ARM, a PSCI/mailbox service registers the same id and
subscribers never know the difference. Messages: `subscribe` (endpoint as
the call's capability, the input/manager pattern), `shutdown` (accepted
only from PID 1 — init), and events published as buffered messages:
`power_button`, `lid`, `ac`, `battery`, generic `notify` with a code.
3. **The service learns event ports from its own FADT copy**: the kernel adds
the FADT as one more memory resource on the acpi-tables node; the service
tells it apart from the AML blobs by signature ("FACP" header — the blob
resources are header-stripped bytecode and start with no signature). The
kernel's own FADT parse is untouched.
4. **The acpi service converts to the harness** (`runtime.service.run`):
protocol messages (subscribe/shutdown), the SCI notification, and the
existing report flow fold into one loop — the shape it was always meant
to have.
5. **GPE/Notify correctness is proven by host unit tests** (synthetic AML
with a Notify inside a method body; aml.zig joins the `zig build test`
loop). The QEMU scenario proves the power button — a *fixed* event,
deterministically injectable via QMP `system_powerdown` — because QEMU
cannot raise GPEs deterministically on this config. Battery/AC/lid and the
embedded controller (`_Qxx`) are interface-complete here and validated on
real hardware (the laptop) later.
## Ground truth the phases build on (verified 2026-07-13)
- `system/devices/power.zig` `shutdown()` is the kernel's S5 write
(SLP_TYP|SLP_EN to PM1a/PM1b control); there is no power syscall.
- init (`system/services/init/init.zig`) spawns vfs/input/device-manager
fire-and-forget — no child ids kept, no signals, no event loop. The whole
stop toolkit exists in `runtime.process` (stop/sendSignal/bindSignals).
- `test/qemu_test.py` has no QMP channel (serial is a one-way file).
- The kernel parses PM1 *control* blocks and SCI_INT from the FADT; the PM1
**event** blocks (offsets 56/60, len at 88) and **GPE0/GPE1** blocks
(offsets 80/84, lens 92/93) are unparsed — the service reads them from its
FADT copy (decision 3).
- The acpi-tables node carries the SCI as its only `len == 1` irq resource
(the broad window is len 256) — that is how the service finds it to
`irqBind`.
- `notify_opcode = 0x86` exists in `system/devices/aml/opcodes.zig` but the
interpreter never handles it — a GPE `_Lxx` body containing Notify fails
evaluation today. Everything else a GPE handler needs (field access,
control flow, method calls) is proven by the ring-3 `_STA`/`_CRS` work.
- The dead-code sweep (spawned task) also edits `system/devices/acpi.zig`;
M21.0 checks whether it landed and rebases before touching that file.
## Status
- [x] **M21.0** — baseline (dead-code sweep confirmed landed on main — no
acpi.zig conflict; `feat/power-events` cut; QMP channel in the harness:
always-on unix socket, client with the capabilities handshake, per-case
`qmp_after` hook, and a hook-must-deliver pass gate that the smoke case
now proves with a harmless query-status; suite 58/58).
- [x] **M21.1** — SCI + the power button (kernel appends the FADT as an
acpi-tables memory resource, tagged by its "FACP" header; `power-protocol`
module + `ServiceId.power = 5`; the acpi service converted to
`runtime.service.run`, registers `.power`, reads PM1 event/control + GPE
ports from its FADT copy, enables ACPI mode if SCI_EN is clear, binds the
SCI (the len-1 irq), sets PWRBTN_EN; the SCI handler clears PM1_STS,
logs `power: button pressed`, publishes `power_button`, acks. Scenario
`power-button` injects a real `system_powerdown` via QMP; initial-ramdisk
timeout 30→60s for the service's added boot work; suite 59/59).
- [x] **M21.2** — Notify + GPE dispatch (interpreter handles `notify_opcode`
into a bounded queue, cleared per-evaluate, drained via
`takeNotifications`; the service walks GPE status/enable bytes, evaluates
`\_GPE._Lxx`/`_Exx` per active bit, maps notified nodes to events
(battery/ac/lid/generic), clears GPE_STS write-1, acks. EC `_Qxx` out.
Host unit test with hand-encoded AML proves the queue; aml.zig joined the
`zig build test` loop. QEMU raises no GPEs — suite is regression net,
59/59).
- [x] **M21.3** — orderly shutdown (init supervises its children on one
endpoint that also carries signals, power events, and a re-arming
heartbeat timer; on `power_button` or a `terminate` signal it logs
`init: shutting down`, runs `stop(child, 2000, endpoint)` in reverse
order, then requests `.power` shutdown; the acpi service honors shutdown
from a subscriber — init is the one subscriber, a soft gate that survives
testing where PID 1 isn't init — and writes SLP_TYP|SLP_EN from ring 3.
`orderly-shutdown` scenario proves button → shutting-down → S5 → QEMU
exit; suite 60/60).
- [ ] **merge** `feat/power-events` → main, push, keep the branch — **loop
ends here**.
---
## Phase notes
**M21.0 QMP:** open the unix socket after Popen, complete the
`qmp_capabilities` handshake, then send the hook's command (for these
scenarios: `{"execute": "system_powerdown"}`). The socket is additive — no
existing case may notice it. Note e3fe3f3 recently reworked how the harness
boots; adapt to its current shape rather than the pre-rework description.
**M21.1 SCI details:** PM1_STS is at the event block base (write-1-to-clear);
PM1_EN at base + block_len/2; PWRBTN bit is 8 in both. If PM1b exists, mirror
reads/writes to both blocks. Enable ACPI mode only when SCI_EN (PM1 control
bit 0) is clear — OVMF boots may already have it set. The publish path reuses
the manager's subscriber table pattern (bounded, drop-on-failed-send).
**M21.2 GPE walk:** GPE0_STS bytes live at the GPE0 block base, GPE0_EN in
the block's upper half; for a set+enabled bit n, the handler method is
`_L%02X` (level) or `_E%02X` (edge) under `\_GPE`. Evaluate, drain the notify
queue, clear the status bit, ack. A missing handler method is clear-and-log,
not an error.
**M21.3 ordering:** init subscribes with retries — the acpi service registers
`.power` well after init starts. The stop sequence runs vfs last (other
services may flush through it). The S5 write mirrors `power.zig`'s
`sleepValue` (SLP_TYP bits [12:10], SLP_EN bit 13); if the write returns, log
`power: S5 write did not take` so the scenario fails loudly instead of
hanging.
**Explicitly out of scope:** the embedded controller and `_Qxx` queries,
battery `_BST`/`_BIF` evaluation beyond the interface stubs, lid/AC on QEMU
(no emulation), reboot over the power protocol, S3 sleep, per-device D-states
(a future lifecycle-vocabulary extension), thermal zones.
+10 -4
View File
@@ -1,6 +1,9 @@
# Process lifecycle: signals over IPC # Process lifecycle: signals over IPC
**Status: design.** The primitives underneath are built ([process-management.md](process-management.md): **Status: increments 1–4 built** (2026-07-12): claim release on death, exit
reasons, published exit events, and signals + one-shot timers + the service
harness are all in — the interface below is as-built. The primitives underneath
predate this design ([process-management.md](process-management.md):
spawn, the supervision link, kill, child-exit notifications); this document designs spawn, the supervision link, kill, child-exit notifications); this document designs
the layer above them — the standard vocabulary a danos process speaks about its own the layer above them — the standard vocabulary a danos process speaks about its own
life, and the stable `runtime.process` interface that carries it. Nothing here is life, and the stable `runtime.process` interface that carries it. Nothing here is
@@ -245,13 +248,16 @@ pub fn stop(id: u32, deadline_ms: u64) void { ... }
/// process_enumerate. /// process_enumerate.
pub fn subscribeExits(endpoint: usize) bool { ... } pub fn subscribeExits(endpoint: usize) bool { ... }
/// How a process ended — from the exit notification. What restart policy reads. /// How a process ended — queried after the exit notification (the kernel records
pub const ExitReason = enum { /// it first, so the two never race). What restart policy reads. (Built in M17.2.)
pub const ExitReason = enum(u8) {
exited, // returned from main / clean exit exited, // returned from main / clean exit
aborted, // abort() — deliberate self-termination (SIGABRT's ghost) aborted, // abort() — deliberate self-termination (SIGABRT's ghost; reserved)
segmentation_fault, // SIGSEGV's ghost segmentation_fault, // SIGSEGV's ghost
illegal_instruction, // SIGILL's ghost illegal_instruction, // SIGILL's ghost
arithmetic_fault, // SIGFPE's ghost arithmetic_fault, // SIGFPE's ghost
protection_fault, // general protection fault
fault, // any other CPU exception
killed, // process_kill killed, // process_kill
}; };
``` ```
+13 -5
View File
@@ -95,11 +95,18 @@ the architecture layer calls up into `tick`.
## Known gaps (bring-up honesty) ## Known gaps (bring-up honesty)
- Device **claims** are not released on death (pre-existing: the fault path has - ~~Device claims are not released on death~~ Closed (M17.1): every path out of a
the same gap) — a killed driver's device stays claimed until reboot. process releases its device claims alongside its IRQ and MSI bindings
(`releaseTaskResourcesLocked`), so a restarted driver can claim its hardware
again — the cleanup half of [process-lifecycle.md](process-lifecycle.md)'s iron
rule 1. The `claim-release` test proves the kill → release → re-claim cycle.
- Kernel stacks of dead tasks are leaked, as on every exit path (no reaper yet). - Kernel stacks of dead tasks are leaked, as on every exit path (no reaper yet).
- There is no exit *status* in the notification, only the id; a supervisor that - ~~There is no exit status in the notification~~ Closed (M17.2): the kernel
needs the code can grow a wait-style call later. records how every process ends — exited, a fault class, or killed — before it
posts the exit notification, and the supervisor reads it with
`process_exit_reason` (`runtime.process.exitReason`). This is the input to
restart policy ([process-lifecycle.md](process-lifecycle.md)); an exit *code*
for the clean case can still ride alongside later.
- Enumerate writes through the caller's raw pointer under the bring-up trust - Enumerate writes through the caller's raw pointer under the bring-up trust
model, like `device_enumerate` (an unmapped page is a self-DoS, not an model, like `device_enumerate` (an unmapped page is a self-DoS, not an
isolation break). isolation break).
@@ -109,4 +116,5 @@ the architecture layer calls up into `tick`.
`process-list` (enumerate), `process-kill` (kernel-level kill paths, refusals, `process-list` (enumerate), `process-kill` (kernel-level kill paths, refusals,
notifications), `supervision` (the whole user-side surface via the process-test notifications), `supervision` (the whole user-side surface via the process-test
service: spawn supervised → enumerate → kill blocked and spinning children → service: spawn supervised → enumerate → kill blocked and spinning children →
notifications → gone). See test/qemu_test.py. notifications → gone), `claim-release` (a killed claim-holder's device is
claimable again). See test/qemu_test.py.
+12 -5
View File
@@ -1,10 +1,17 @@
# Resilience: fault isolation and live restart # Resilience: fault isolation and live restart
Steps 1–2 of the ordering below are **built**: user-mode isolation, and fault → Steps 1–4 of the ordering below are **built** (M17–M18, 2026-07-13): user-mode
kill the process → keep the core (`onException` in `system/kernel/kernel.zig`; the isolation; fault → kill the process → keep the core (`onException`; the
`fault-recovery` test proves a crashing ring-3 process dies alone while the system `fault-recovery` test); the supervisor notification **with exit reasons**
keeps running). The supervisor notification and restart policy (steps 3+) are ([process-lifecycle.md](process-lifecycle.md) — clean exit, fault class, or
still design. This is the property danos is really chasing: killed, recorded before the notice posts); and the **restart policy itself**
([device-manager.md](device-manager.md)): the device manager supervises every
driver, restarts crashes with backoff, caps crash loops, and re-claims work
because the kernel releases a dead process's claims. The `driver-restart` and
`usb-report` scenarios prove kill → release → respawn → re-claim → re-report
end to end. What remains of this document's ladder is scope, not mechanism:
more of the system moved into restartable processes (the discovery migration,
[m19-m20-plan.md](m19-m20-plan.md), is the next rung). This is the property danos is really chasing:
**if a part of the OS breaks, isolate it, and re-initialise it — without rebooting.** **if a part of the OS breaks, isolate it, and re-initialise it — without rebooting.**
A crashed driver gets restarted; a wedged service gets killed and brought back. It's A crashed driver gets restarted; a wedged service gets killed and brought back. It's
the reason the [microkernel](vision.md) shape was chosen, and it's a *separate* goal the reason the [microkernel](vision.md) shape was chosen, and it's a *separate* goal
+20
View File
@@ -94,6 +94,15 @@ pub fn send(h: Handle, message: []const u8) bool {
/// GSI. See `isNotification`. /// GSI. See `isNotification`.
pub const notify_badge_bit: u64 = abi.notify_badge_bit; pub const notify_badge_bit: u64 = abi.notify_badge_bit;
/// Set alongside `notify_badge_bit` when the notification is a **signal** — the
/// lifecycle vocabulary of docs/process-lifecycle.md, delivered to the endpoint
/// nominated with `process.bindSignals`. Decode with `process.signalsFrom`.
pub const notify_signal_bit: u64 = abi.notify_signal_bit;
/// Set alongside `notify_badge_bit` when the notification is a **one-shot timer**
/// landing (`system.timerOnce`).
pub const notify_timer_bit: u64 = abi.notify_timer_bit;
/// Set alongside `notify_badge_bit` when the notification is a **child-exit /// Set alongside `notify_badge_bit` when the notification is a **child-exit
/// notice** — a process this one spawned (with an exit endpoint) has ended — /// notice** — a process this one spawned (with an exit endpoint) has ended —
/// rather than a device interrupt. The low bits carry the child's process id. /// rather than a device interrupt. The low bits carry the child's process id.
@@ -133,6 +142,17 @@ pub const Received = struct {
} }
/// The task id of whoever posted a buffered message, meaningful only when /// The task id of whoever posted a buffered message, meaningful only when
/// Whether this arrival is a signal notification — decode the set with
/// `process.signalsFrom(badge)`.
pub fn isSignal(self: Received) bool {
return self.isNotification() and self.badge & notify_signal_bit != 0;
}
/// Whether this arrival is a one-shot timer landing (`system.timerOnce`).
pub fn isTimer(self: Received) bool {
return self.isNotification() and self.badge & notify_timer_bit != 0;
}
/// `isMessage`. (The badge's low bits, with the three high marker bits masked off.) /// `isMessage`. (The badge's low bits, with the three high marker bits masked off.)
pub fn senderTaskId(self: Received) u32 { pub fn senderTaskId(self: Received) u32 {
return @intCast(self.badge & ~(notify_badge_bit | notify_exit_bit | notify_message_bit)); return @intCast(self.badge & ~(notify_badge_bit | notify_exit_bit | notify_message_bit));
+91 -1
View File
@@ -1,8 +1,15 @@
//! Process-level runtime types: what a user program receives at entry. Mirrors //! Process-level runtime types: what a user program receives at entry (`Init`,
//! the argv contract) and the process end of the lifecycle
//! (docs/process-lifecycle.md) — today the exit reason a supervisor reads to
//! decide restart; signals and the stop sequence land here with M17.4. Mirrors
//! the spirit of `std.process.Init.Minimal` in danos terms — std's `Args` holds //! the spirit of `std.process.Init.Minimal` in danos terms — std's `Args` holds
//! no data on freestanding targets, so the type is danos's own. //! no data on freestanding targets, so the type is danos's own.
const std = @import("std"); const std = @import("std");
const abi = @import("abi");
const sc = @import("system-call.zig");
const ipc = @import("ipc.zig");
const system = @import("system.zig");
/// Everything a program receives at entry. Passed to /// Everything a program receives at entry. Passed to
/// `pub fn main(init: runtime.process.Init)`; programs that need nothing keep /// `pub fn main(init: runtime.process.Init)`; programs that need nothing keep
@@ -45,3 +52,86 @@ pub const Arguments = struct {
} }
}; };
}; };
/// How a process ended — what a supervisor's restart policy reads: a clean exit
/// meant to stop, a fault wants a restart with backoff, killed means the
/// supervisor did it itself (docs/process-lifecycle.md).
pub const ExitReason = abi.ExitReason;
/// How dead child `id` ended. Ask after the exit notification arrives — the
/// kernel records the reason before it posts the notification, so this never
/// races it. Returns null for an id that never lived, is still alive, was
/// evicted from the kernel's bounded record, or is not this process's child
/// (the same authority gate as `kill`).
pub fn exitReason(id: u32) ?ExitReason {
const r = sc.systemCall1(.process_exit_reason, id);
if (r > ~@as(usize, 0) - 4095) return null; // a wrapped -errno
return @enumFromInt(r);
}
/// The signal vocabulary (docs/process-lifecycle.md): POSIX's concepts, danos's
/// names, message delivery. A signal is a one-way coalescing statement — never a
/// question (liveness is the zero-length ping call) and never kill (that is
/// `system.kill`, unhandleable by definition).
pub const Signal = abi.Signal;
/// The coalesced set of signals one notification delivered: two pending
/// terminates arrive as one. Decode a received badge with `signalsFrom`.
pub const SignalSet = struct {
pending: u32,
pub fn has(set: SignalSet, signal: Signal) bool {
return set.pending & (@as(u32, 1) << @intFromEnum(signal)) != 0;
}
};
/// Nominate `endpoint` as this process's signal endpoint. Signals posted while
/// unbound have pended; they are delivered immediately on bind, coalesced.
pub fn bindSignals(endpoint: usize) bool {
return sc.systemCall1(.signal_bind, endpoint) == 0;
}
/// Decode a received badge into the signals it delivered, or null if it is not
/// a signal notification.
pub fn signalsFrom(badge: u64) ?SignalSet {
if (badge & abi.notify_badge_bit == 0 or badge & abi.notify_signal_bit == 0) return null;
return .{ .pending = @truncate(badge & ~(abi.notify_badge_bit | abi.notify_signal_bit)) };
}
/// Post `signal` to child `id` (or to yourself). Supervisor-gated, like kill;
/// non-blocking, always — a statement, not a conversation.
pub fn sendSignal(id: u32, signal: Signal) bool {
return sc.systemCall2(.process_signal, id, @intFromEnum(signal)) == 0;
}
/// The standard stop sequence (docs/process-lifecycle.md): terminate, wait up to
/// `deadline_ms` for the exit notification on `exit_endpoint` (the endpoint the
/// child was spawned with), then kill. Any *other* notifications arriving on
/// that endpoint while stopping are consumed and dropped — a supervisor with
/// concurrent traffic implements the same sequence inside its own event loop
/// (arm `system.timerOnce`, keep serving) instead of calling this.
pub fn stop(id: u32, deadline_ms: u64, exit_endpoint: usize) void {
_ = sendSignal(id, .terminate);
_ = system.timerOnce(exit_endpoint, deadline_ms);
var receive: [8]u8 = undefined;
while (true) {
const got = ipc.replyWait(exit_endpoint, &.{}, &receive, null);
if (got.isChildExit() and got.childProcessId() == id) return;
if (got.isTimer()) break; // the deadline passed first — escalate
}
_ = system.kill(id);
while (true) {
const got = ipc.replyWait(exit_endpoint, &.{}, &receive, null);
if (got.isChildExit() and got.childProcessId() == id) return;
}
}
/// Subscribe `endpoint` to published exit events: every process death posts an
/// asynchronous notification with the same badge encoding as a supervisor's exit
/// notice (decode with `ipc.Received.isChildExit`/`childProcessId`). For stateful
/// services: release what the dead client held — file handles, subscriptions —
/// because a service must never depend on clients cleaning up after themselves
/// (docs/process-lifecycle.md). Ungated, like `system.processes`.
pub fn subscribeExits(endpoint: usize) bool {
return sc.systemCall1(.process_subscribe, endpoint) == 0;
}
+10
View File
@@ -17,6 +17,12 @@ pub const ipc = @import("ipc.zig");
pub const start = @import("start.zig"); pub const start = @import("start.zig");
/// The VFS wire protocol (shared with the VFS server). /// The VFS wire protocol (shared with the VFS server).
pub const vfs_protocol = @import("vfs-protocol"); pub const vfs_protocol = @import("vfs-protocol");
/// The device-manager protocol: hello + tree reports (docs/device-manager.md).
pub const device_manager_protocol = @import("device-manager-protocol");
/// The power protocol: events (button, lid, battery) + shutdown (docs/m21-plan.md).
pub const power_protocol = @import("power-protocol");
/// Keyboard-event listening (subscribe/next) and broadcasting (publish), over the input /// Keyboard-event listening (subscribe/next) and broadcasting (publish), over the input
/// service. See library/runtime/input.zig and system/services/input/. /// service. See library/runtime/input.zig and system/services/input/.
pub const input = @import("input.zig"); pub const input = @import("input.zig");
@@ -35,5 +41,9 @@ pub const panic = start.panic;
/// Process entry types: the `Init` handed to `main`, and its `Arguments`. /// Process entry types: the `Init` handed to `main`, and its `Arguments`.
pub const process = @import("process.zig"); pub const process = @import("process.zig");
/// The service harness: one replyWait loop folding requests, signals, and
/// notifications into callbacks (docs/process-lifecycle.md).
pub const service = @import("service.zig");
/// The heap as a `std.mem.Allocator`, for Zig `std` containers in user code. /// The heap as a `std.mem.Allocator`, for Zig `std` containers in user code.
pub const allocator = heap.allocator; pub const allocator = heap.allocator;
+83
View File
@@ -0,0 +1,83 @@
//! The service harness (docs/process-lifecycle.md): one replyWait loop that
//! folds protocol requests, signals, and subscribed notifications into
//! callbacks — so the lifecycle contract ("answers ping, exits on terminate")
//! is satisfied by construction and a service author writes domain logic only.
//! Nothing is asynchronous inside the process: a callback runs at a point the
//! loop chose, never on a hijacked stack — the whole reason signals are
//! messages.
//!
//! The liveness probe: a **zero-length request is the universal ping**, answered
//! with a zero-length reply by the harness itself. No protocol's requests start
//! at length zero, so the encoding cannot collide, and there is nothing for a
//! service author to implement — a wedged service simply fails to answer, which
//! is the diagnosis (see docs/ipc.md).
const abi = @import("abi");
const ipc = @import("ipc.zig");
const process = @import("process.zig");
pub const Callbacks = struct {
/// Called once with the service's endpoint before the loop starts — the
/// place to subscribe to exit events, bind IRQs, or announce readiness.
/// Return false to abort startup (the process exits).
init: ?*const fn (endpoint: ipc.Handle) bool = null,
/// One protocol request from `sender` (a task id): write the reply into
/// `reply`, return its length. `capability` is the handle the request
/// carried, if any (M13 cap passing — how a subscriber hands over its
/// endpoint). The zero-length ping never reaches this.
on_message: *const fn (message: []const u8, reply: []u8, sender: u32, capability: ?ipc.Handle) usize,
/// A notification that is not a signal — a subscribed exit event, a bound
/// IRQ, a timer landing. The raw badge; decode with the ipc helpers.
on_notification: ?*const fn (badge: u64) void = null,
/// The reload signal. Default: ignored.
on_reload: ?*const fn () void = null,
/// The terminate signal, called before the loop returns. The clean exit is
/// the return itself — never put *necessary* work here (iron rule 1: a kill
/// arrives with no warning; this is for graceful extras only).
on_terminate: ?*const fn () void = null,
/// Publish the endpoint under a well-known service id at startup.
service: ?abi.ServiceId = null,
};
/// Run the service: create and (optionally) register the endpoint, bind signals
/// to it, call `init`, then serve until `terminate` arrives — at which point the
/// loop returns and main's return is the clean exit the supervisor reads as
/// `ExitReason.exited`. `maximum_message` sizes the receive and reply buffers
/// (a service passes its protocol's message maximum).
pub fn run(comptime maximum_message: usize, callbacks: Callbacks) void {
const endpoint = ipc.createIpcEndpoint() orelse return;
if (callbacks.service) |id| {
if (!ipc.register(id, endpoint)) return;
}
_ = process.bindSignals(endpoint);
if (callbacks.init) |initialise| {
if (!initialise(endpoint)) return;
}
var reply_buffer: [maximum_message]u8 = undefined;
var reply_len: usize = 0;
var receive: [maximum_message]u8 = undefined;
while (true) {
const got = ipc.replyWait(endpoint, reply_buffer[0..reply_len], &receive, null);
if (got.isNotification()) {
reply_len = 0; // nothing owed for a notification
if (process.signalsFrom(got.badge)) |signals| {
if (signals.has(.reload)) {
if (callbacks.on_reload) |onReload| onReload();
}
if (signals.has(.terminate)) {
if (callbacks.on_terminate) |onTerminate| onTerminate();
return; // the loop's return IS the clean exit
}
continue;
}
if (callbacks.on_notification) |onNotification| onNotification(got.badge);
continue;
}
if (got.len == 0) {
reply_len = 0; // the universal ping: a zero-length reply, from the harness
continue;
}
reply_len = callbacks.on_message(receive[0..got.len], &reply_buffer, got.senderTaskId(), got.cap);
}
}
+9
View File
@@ -32,6 +32,15 @@ pub fn sleep(ms: usize) void {
_ = sc.systemCall1(.sleep, ms); _ = sc.systemCall1(.sleep, ms);
} }
/// Arm a one-shot timer: after `ms` milliseconds the kernel posts a timer
/// notification (`ipc.Received.isTimer`) to `endpoint`. The timed wait of
/// docs/process-lifecycle.md — a service arms a deadline and keeps serving,
/// instead of blocking in sleep; what stop-sequence escalation, hello deadlines,
/// and restart backoff are built from.
pub fn timerOnce(endpoint: usize, ms: u64) bool {
return sc.systemCall2(.timer_bind, endpoint, ms) == 0;
}
/// Monotonic nanoseconds since boot — a time source for timeouts and short delays. It /// Monotonic nanoseconds since boot — a time source for timeouts and short delays. It
/// only ever moves forward. This is *not* wall-clock time (no date, no timezone — that /// only ever moves forward. This is *not* wall-clock time (no date, no timezone — that
/// is a user-space service layered on top). Deadline pattern for a bounded poll loop: /// is a user-space service layered on top). Deadline pattern for a bounded poll loop:
+53
View File
@@ -53,9 +53,31 @@ pub const SystemCall = enum(u64) {
process_enumerate = 24, // process_enumerate(buffer, maximum) -> total: snapshot the task table process_enumerate = 24, // process_enumerate(buffer, maximum) -> total: snapshot the task table
process_kill = 25, // process_kill(id) -> 0/-errno: end a process this process spawned process_kill = 25, // process_kill(id) -> 0/-errno: end a process this process spawned
ipc_send = 26, // ipc_send(handle, message_ptr, message_len) -> 0/-errno: post a payload to an endpoint's async queue without blocking ipc_send = 26, // ipc_send(handle, message_ptr, message_len) -> 0/-errno: post a payload to an endpoint's async queue without blocking
process_exit_reason = 27, // process_exit_reason(id) -> ExitReason/-errno: how a dead child ended (its supervisor only)
process_subscribe = 28, // process_subscribe(endpoint) -> 0/-errno: subscribe to published exit events — every death posts a notification
signal_bind = 29, // signal_bind(endpoint) -> 0/-errno: nominate the endpoint this process's signals arrive on
process_signal = 30, // process_signal(id, signal) -> 0/-errno: post a signal to a child (or to yourself)
timer_bind = 31, // timer_bind(endpoint, ms) -> 0/-errno: one-shot timer — posts a notification when ms elapse
_, _,
}; };
/// How a process ended — recorded by the kernel at death, queried by the
/// supervisor with `process_exit_reason`, and the input to its restart decision
/// (docs/process-lifecycle.md): a clean exit meant to stop, a fault wants a
/// restart with backoff, killed means the supervisor did it itself. The faults
/// mirror the CPU exceptions a ring-3 process can die of; they are exit reasons,
/// never delivered to the faulting process (recovery is restart, not a handler).
pub const ExitReason = enum(u8) {
exited = 0, // returned from main / called exit
aborted = 1, // deliberate self-termination (reserved: no abort path yet)
segmentation_fault = 2, // page fault
illegal_instruction = 3, // invalid opcode
arithmetic_fault = 4, // divide error, x87 or SIMD fault
protection_fault = 5, // general protection fault
fault = 6, // any other CPU exception
killed = 7, // process_kill
};
/// The x86 MSI message address base (`0xFEE0_0000`): a device raises an MSI by writing /// The x86 MSI message address base (`0xFEE0_0000`): a device raises an MSI by writing
/// `data` to this address, which the Local APIC turns into an interrupt at the vector /// `data` to this address, which the Local APIC turns into an interrupt at the vector
/// in `data`. The kernel returns the concrete (address, data) from `msi_bind`; this is /// in `data`. The kernel returns the concrete (address, data) from `msi_bind`; this is
@@ -93,6 +115,35 @@ pub const notify_exit_bit: u64 = 1 << 62;
/// broadcasts where a rendezvous is the wrong shape (the input service is the first user). /// broadcasts where a rendezvous is the wrong shape (the input service is the first user).
pub const notify_message_bit: u64 = 1 << 61; pub const notify_message_bit: u64 = 1 << 61;
/// Set (alongside `notify_badge_bit`) in the badge of a **signal notification** —
/// the process-lifecycle vocabulary of docs/process-lifecycle.md, delivered to the
/// endpoint the process nominated with `signal_bind`. The low bits carry the
/// coalesced pending mask (bit positions = `Signal` values): signals are
/// statements, not questions, and two pending terminates are one terminate.
pub const notify_signal_bit: u64 = 1 << 60;
/// Set (alongside `notify_badge_bit`) in the badge of a **timer notification** —
/// a one-shot `timer_bind` deadline landing. No payload bits: what to do when the
/// deadline fires is whatever the receiver armed it for (a stop-sequence
/// escalation, a restart backoff, an alarm).
pub const notify_timer_bit: u64 = 1 << 59;
/// The signal vocabulary (docs/process-lifecycle.md): POSIX's concepts, danos's
/// names, message delivery. The value is the bit position in the pending mask — a
/// private kernel/runtime detail, free to change while they ship together. Kill
/// is not here (it is `process_kill`, unhandleable by definition); faults are not
/// here (they are `ExitReason`s — recovery is restart, not a handler); liveness is
/// not here (a question, asked as the zero-length ping call, not a statement).
pub const Signal = enum(u5) {
terminate = 0, // finish up and exit (the polite half of the stop sequence)
reload = 1, // re-read configuration / re-scan
interrupt = 2, // interactive interrupt (no sender until a console exists)
quit = 3, // as interrupt, by convention more final
alarm = 4, // a timer the process armed for itself (unbuilt: no consumer yet)
user_1 = 5, // service-defined
user_2 = 6, // service-defined
};
/// Capacity of `ProcessDescriptor.name` — matches the longest name `system_spawn` /// Capacity of `ProcessDescriptor.name` — matches the longest name `system_spawn`
/// accepts, so a process's recorded name (its argv[0]) is never truncated. /// accepts, so a process's recorded name (its argv[0]) is never truncated.
pub const maximum_process_name = 64; pub const maximum_process_name = 64;
@@ -125,6 +176,8 @@ pub const ServiceId = enum(u32) {
vfs = 1, vfs = 1,
input = 2, input = 2,
ps2_bus = 3, // the 8042 owner; child device drivers attach here for raw bytes ps2_bus = 3, // the 8042 owner; child device drivers attach here for raw bytes
device_manager = 4, // the tree, the matcher, the supervisor (docs/device-manager.md)
power = 5, // system power: events (button, lid, battery) + shutdown (docs/m21-plan.md; domain-named per decision 7 — the acpi service registers it on x86, a PSCI service will on ARM)
_, _,
}; };
+138 -504
View File
@@ -17,7 +17,6 @@
const std = @import("std"); const std = @import("std");
const boot_handoff = @import("boot-handoff"); const boot_handoff = @import("boot-handoff");
const abi = @import("abi"); const abi = @import("abi");
const acpi_ids = @import("acpi-ids");
const parameters = @import("parameters"); const parameters = @import("parameters");
const device_model = @import("device-model.zig"); const device_model = @import("device-model.zig");
const aml = @import("aml/aml.zig"); const aml = @import("aml/aml.zig");
@@ -41,6 +40,9 @@ pub const RegisterAccess = struct {
/// Everything the power subsystem needs, extracted from the FADT and the AML /// Everything the power subsystem needs, extracted from the FADT and the AML
/// sleep packages during discovery. Populated by `discover`, read by `power`. /// sleep packages during discovery. Populated by `discover`, read by `power`.
pub const PowerInformation = struct { pub const PowerInformation = struct {
/// The System Control Interrupt's GSI (FADT SCI_INT) — the line ACPI events
/// (power button, GPEs) arrive on. Published to the acpi service for M21.
sci_interrupt: u16 = 0,
/// The SMM command port and the value that switches the platform into ACPI mode. /// The SMM command port and the value that switches the platform into ACPI mode.
smi_cmd: u16 = 0, smi_cmd: u16 = 0,
acpi_enable: u8 = 0, acpi_enable: u8 = 0,
@@ -155,6 +157,13 @@ pub var namespace: ?aml.Namespace = null;
/// Physical address of the DSDT the FADT points at, or 0. /// Physical address of the DSDT the FADT points at, or 0.
pub var dsdt_physical: u64 = 0; pub var dsdt_physical: u64 = 0;
/// The FADT itself (physical + length), published on the acpi-tables node so
/// the ring-3 acpi service can read the PM1 event and GPE blocks it needs for
/// the event side (docs/m21-plan.md decision 3). Distinguished from the AML
/// blob resources by its intact "FACP" header — the blobs are header-stripped.
var fadt_physical: u64 = 0;
var fadt_length: u64 = 0;
// AML blocks (DSDT + any SSDTs) collected during the table walk, as physical // AML blocks (DSDT + any SSDTs) collected during the table walk, as physical
// address + length of each table's post-header bytecode. Scanned after the walk // address + length of each table's post-header bytecode. Scanned after the walk
// for the sleep-state (`_Sx`) packages. // for the sleep-state (`_Sx`) packages.
@@ -360,35 +369,19 @@ const Hpet = extern struct {
page_protection: u8, page_protection: u8,
}; };
// --- PCI configuration-space header (first 64 bytes, common fields) ---------
const PciHeader = extern struct {
vendor_id: u16 align(1),
device_id: u16 align(1),
command: u16 align(1),
status: u16 align(1),
revision_id: u8,
prog_if: u8,
subclass: u8,
class_code: u8,
cache_line_size: u8,
latency_timer: u8,
/// bit 7 set => multi-function device.
header_type: u8,
bist: u8,
// 0x10 onward (BARs, etc.) depends on header_type; read separately.
};
// --- Entry point ------------------------------------------------------------ // --- Entry point ------------------------------------------------------------
/// Discover hardware from the ACPI tables rooted at `rsdp_physical` and populate /// Discover hardware from the ACPI tables rooted at `rsdp_physical` and populate
/// `device_tree`. `hal` provides MMIO mapping (for PCIe ECAM) and port I/O. Also parses the /// `device_tree`. `hal` provides MMIO mapping (for PCIe ECAM) and port I/O. Also parses the
/// FADT and the AML sleep-state (`_Sx`) packages into `power_information` for the power service. /// FADT and the AML sleep-state (`_Sx`) packages into `power_information` for the power service.
pub fn discover(rsdp_physical: u64, device_tree: *DeviceTree, hal: Hal) !void { pub fn discover(rsdp_physical: u64, memory_regions: []const boot_handoff.MemoryRegion, device_tree: *DeviceTree, hal: Hal) !void {
if (rsdp_physical == 0) return error.NoRsdp; if (rsdp_physical == 0) return error.NoRsdp;
boot_memory_regions = memory_regions;
// Start clean so a re-run doesn't accumulate stale state. // Start clean so a re-run doesn't accumulate stale state.
power_information = .{}; power_information = .{};
fadt_physical = 0;
fadt_length = 0;
platform_information = .{}; platform_information = .{};
aml_stats = .{}; aml_stats = .{};
namespace = null; namespace = null;
@@ -420,12 +413,58 @@ pub fn discover(rsdp_physical: u64, device_tree: *DeviceTree, hal: Hal) !void {
aml_stats = .{ .nodes = namespace.?.nodeCount(), .consumed = pr.consumed, .total = pr.total }; aml_stats = .{ .nodes = namespace.?.nodeCount(), .consumed = pr.consumed, .total = pr.total };
power_information.s5 = aml.sleepState(&namespace.?, 5); power_information.s5 = aml.sleepState(&namespace.?, 5);
power_information.s3 = aml.sleepState(&namespace.?, 3); power_information.s3 = aml.sleepState(&namespace.?, 3);
// Fold the namespace's Device objects into the generic tree. // The namespace's Device objects are no longer folded into the kernel
wireAcpiDevices(device_tree, &namespace.?, hal) catch {}; // tree (M20.3): the ring-3 acpi service claims the acpi-tables node
// (published below), re-parses the same blobs, and registers + reports
// the _HID devices itself. The kernel keeps the namespace only for the
// \_S5 sleep type above.
} else |_| { } else |_| {
// AML parse failed (e.g. out of memory); power stays best-effort with // AML parse failed (e.g. out of memory); power stays best-effort with
// whatever the FADT alone provided. // whatever the FADT alone provided.
} }
// Publish the acpi-tables node (docs/m19-m20-plan.md M20): the AML blobs as
// memory resources for the acpi service to map and parse in ring 3, a broad
// io_port grant for the OperationRegion access its interpreter needs, and
// the SCI for the events track (M21). Exactly one node, one trusted
// claimant. Kept even when the kernel-side device building (above) retires
// in M20.3 — the kernel still owns the *static* tables and \_S5.
publishAcpiTablesNode(device_tree) catch {};
}
/// Build the acpi-tables node (see the call site in discover). Best-effort: a
/// failure here leaves the kernel-seeded tree working, only the ring-3 service
/// finds nothing to claim.
fn publishAcpiTablesNode(device_tree: *DeviceTree) !void {
const node = try device_tree.addChild(device_tree.root, .acpi_tables, "acpi-tables");
// One memory resource per AML block — page-aligned base down, length padded
// up to cover the bytecode, so mmio_map hands the service a pointer into it.
var i: usize = 0;
while (i < aml_block_count and i < device_model.maximum_resources - 2) : (i += 1) {
// mmio_map preserves the sub-page offset, so the service maps this and
// gets a pointer straight to the bytecode.
_ = node.addResource(.memory, aml_block_physical[i], aml_block_len[i]);
}
// The broad I/O grant: OperationRegions name whatever ports the firmware
// chose (EC, PM1, GPE, SMBus); which ports cannot be known before the AML
// that names them is parsed, so the grant is the whole space — the honest
// trust boundary of docs/m19-m20-plan.md decision 5.
_ = node.addResource(.io_port, 0, 1 << 16);
// A broad interrupt window: ACPI _CRS names legacy ISA IRQs (the PS/2 lines
// 1 and 12, the RTC, …), and the service registers those devices under this
// node, so it must own a superset. The range [0, 256) covers every GSI; the
// SCI (recorded first, len 1) stays distinct so M21 can pick it out.
if (power_information.sci_interrupt != 0) _ = node.addResource(.irq, power_information.sci_interrupt, 1);
_ = node.addResource(.irq, 0, 256);
// The FADT rides along (M21): the service reads the PM1 event / GPE blocks
// from its own copy, telling it apart from the AML blobs by signature.
if (fadt_physical != 0) _ = node.addResource(.memory, fadt_physical, fadt_length);
}
/// The number of Device objects in the namespace built during discovery, or 0.
pub fn amlDeviceCount() usize {
if (namespace) |*ns| return aml.deviceCount(ns);
return 0;
} }
/// Walk the RSDT (Entry = u32) or XSDT (Entry = u64): validate it, then dispatch /// Walk the RSDT (Entry = u32) or XSDT (Entry = u64): validate it, then dispatch
@@ -451,10 +490,12 @@ fn handleTable(device_tree: *DeviceTree, hal: Hal, sdt_physical: u64) !void {
if (std.mem.eql(u8, &sig, &APIC)) { if (std.mem.eql(u8, &sig, &APIC)) {
try parseMadt(device_tree, header); try parseMadt(device_tree, header);
} else if (std.mem.eql(u8, &sig, &MCFG)) { } else if (std.mem.eql(u8, &sig, &MCFG)) {
try parseMcfg(device_tree, hal, header); try parseMcfg(device_tree, header);
} else if (std.mem.eql(u8, &sig, &HPET)) { } else if (std.mem.eql(u8, &sig, &HPET)) {
try parseHpet(device_tree, hal, header); try parseHpet(device_tree, hal, header);
} else if (std.mem.eql(u8, &sig, &FACP)) { } else if (std.mem.eql(u8, &sig, &FACP)) {
fadt_physical = sdt_physical;
fadt_length = header.length;
parseFadt(header); parseFadt(header);
} else if (std.mem.eql(u8, &sig, &SPCR)) { } else if (std.mem.eql(u8, &sig, &SPCR)) {
parseSpcr(header); parseSpcr(header);
@@ -536,7 +577,7 @@ fn parseMadt(device_tree: *DeviceTree, header: *const SystemDescriptorTableHeade
} }
/// MCFG -> a pci_host_bridge per ECAM segment, then a PCI enumeration underneath. /// MCFG -> a pci_host_bridge per ECAM segment, then a PCI enumeration underneath.
fn parseMcfg(device_tree: *DeviceTree, hal: Hal, header: *const SystemDescriptorTableHeader) !void { fn parseMcfg(device_tree: *DeviceTree, header: *const SystemDescriptorTableHeader) !void {
const total: usize = header.length; const total: usize = header.length;
const base: [*]const u8 = @ptrCast(header); const base: [*]const u8 = @ptrCast(header);
@@ -551,109 +592,85 @@ fn parseMcfg(device_tree: *DeviceTree, hal: Hal, header: *const SystemDescriptor
// ECAM window: 1 MiB of configuration space per bus. // ECAM window: 1 MiB of configuration space per bus.
_ = bridge.addResource(.memory, alloc.base_address, bus_count << 20); _ = bridge.addResource(.memory, alloc.base_address, bus_count << 20);
_ = bridge.addResource(.bus_range, alloc.start_bus, bus_count); _ = bridge.addResource(.bus_range, alloc.start_bus, bus_count);
addBridgeApertures(bridge);
// The bridge decodes the whole 16-bit I/O space toward its bus — the
// window functions' I/O BARs must register-contain within (M19.2).
_ = bridge.addResource(.io_port, 0, 1 << 16);
try enumeratePci(device_tree, bridge, hal, alloc.*); // The function walk itself retired to ring 3 (M19.3): the pci-bus
// driver claims this bridge, repeats the scan through its ECAM grant,
// and device_registers what it finds — the kernel seeds only the
// bridge. The scan's equivalence was proven before the hand-off
// (pci-scan), and the walk's history is in git if archaeology calls.
} }
} }
/// Brute-force scan the ECAM window's bus range for present PCI functions. No /// The boot memory map, stored at discover() entry for the aperture derivation
/// bridge recursion yet: on the ECAM path the host bridge decodes every bus in /// below (and, in M20, for the acpi-tables node's containment windows).
/// the window, so scanning the declared range finds everything QEMU exposes. var boot_memory_regions: []const boot_handoff.MemoryRegion = &.{};
fn enumeratePci(
device_tree: *DeviceTree,
bridge: *device_model.Device,
hal: Hal,
alloc: McfgAllocation,
) !void {
var bus: u16 = alloc.start_bus;
while (bus <= alloc.end_bus) : (bus += 1) {
var device: u8 = 0;
while (device < 32) : (device += 1) {
const h0: *align(1) const PciHeader = @ptrCast(pciConfigurationPtr(alloc, hal, @intCast(bus), device, 0));
if (h0.vendor_id == 0xFFFF) continue; // no function 0 => slot empty
const funcs: u8 = if (h0.header_type & 0x80 != 0) 8 else 1; /// The bridge's MMIO apertures, derived from the boot memory map's holes
var function: u8 = 0; /// (docs/m19-m20-plan.md decision 2): registered PCI functions carry BAR
while (function < funcs) : (function += 1) { /// resources, and `device_register` containment demands the bridge own windows
const configuration = pciConfigurationPtr(alloc, hal, @intCast(bus), device, function); /// that cover them. Everything the firmware described is "not hole"; the low
const h: *align(1) const PciHeader = @ptrCast(configuration); /// aperture runs from the end of the described space below 4 GiB up to the
if (h.vendor_id == 0xFFFF) continue; /// I/O-APIC region, the high one from 4 GiB (or the end of RAM above it) to
/// the 46-bit line. Coarse, mechanical, and AML-free — available at boot no
var nb: [24]u8 = undefined; /// matter what later moved to user space.
const nm = std.fmt.bufPrint(&nb, "{s}:{x:0>2}:{x:0>2}.{d}", .{ fn addBridgeApertures(bridge: *device_model.Device) void {
bridge.name(), bus, device, function, // Below 4 GiB the described regions are sparse (RAM low, firmware flash
}) catch "pcidev"; // and tables high), so the holes are the *gaps between* them — a single
const node = try device_tree.addChild(bridge, .pci_device, nm); // "after the last region" rule dies on OVMF's flash at the very top.
// Resource 0 is the function's own 4 KiB ECAM configuration space. A // Sort-merge the described ranges, then keep the three largest gaps
// claimed PCI driver mmio_maps this to reach its command register, // (resource slots are bounded at 8 per device; ECAM + bus range + 3 + the
// BARs, and — the point — its capability list (MSI/MSI-X, PCIe // high aperture fits). Above 4 GiB one aperture runs from the end of the
// extended caps), without any new syscall. Physical address per the // described space to the 46-bit line.
// ECAM formula (same as pciConfigurationPtr). const Range = struct { base: u64, end: u64 };
const config_physical = alloc.base_address + var below: [64]Range = undefined;
(@as(u64, @as(u8, @intCast(bus)) - alloc.start_bus) << 20) + var below_count: usize = 0;
(@as(u64, device) << 15) + (@as(u64, function) << 12); var high_end: u64 = 1 << 32;
_ = node.addResource(.memory, config_physical, abi.page_size); for (boot_memory_regions) |region| {
node.ids.pci_vendor = h.vendor_id; const end = region.base + region.pages * 4096;
node.ids.pci_device = h.device_id; // Above 4 GiB only *usable RAM* blocks the aperture: OVMF describes
node.ids.pci_class = (@as(u24, h.class_code) << 16) | // its own 64-bit PCI window as a reserved region and then programs
(@as(u24, h.subclass) << 8) | h.prog_if; // BARs inside it — honoring reserved there would exclude the very
node.ids.pci_bdf = (@as(u16, @intCast(bus)) << 8) | (@as(u16, device) << 3) | function; // space BARs live in. Below 4 GiB every described region blocks (the
// kernel image, the tables, the ramdisk all live there). Bring-up
// BARs only exist in header type 0 (normal devices), not bridges. // trust: only the bridge's claimant can register into the aperture.
if (h.header_type & 0x7F == 0) addBars(node, configuration); if (region.kind == .usable and end > high_end) high_end = end;
if (region.base >= (1 << 32) or below_count == below.len) continue;
below[below_count] = .{ .base = region.base, .end = @min(end, 1 << 32) };
below_count += 1;
}
// Insertion sort by base (the map is small and this runs once at boot).
for (1..below_count) |i| {
const key = below[i];
var j = i;
while (j > 0 and below[j - 1].base > key.base) : (j -= 1) below[j] = below[j - 1];
below[j] = key;
}
// Walk the sorted ranges, collecting inter-region gaps of at least 1 MiB.
var gaps: [3]Range = .{Range{ .base = 0, .end = 0 }} ** 3;
var cursor: u64 = 0;
var index: usize = 0;
while (index <= below_count) : (index += 1) {
const gap_end = if (index == below_count) (1 << 32) else below[index].base;
if (gap_end > cursor and gap_end - cursor >= (1 << 20)) {
// Keep the three largest, replacing the smallest kept so far.
var smallest: usize = 0;
for (gaps, 0..) |gap, gi| {
if (gap.end - gap.base < gaps[smallest].end - gaps[smallest].base) smallest = gi;
}
if (gap_end - cursor > gaps[smallest].end - gaps[smallest].base) {
gaps[smallest] = .{ .base = cursor, .end = gap_end };
} }
} }
if (index < below_count and below[index].end > cursor) cursor = below[index].end;
} }
} for (gaps) |gap| {
if (gap.end > gap.base) _ = bridge.addResource(.memory, gap.base, gap.end - gap.base);
/// Record and size the memory/IO windows named by a device's Base Address
/// Registers. Sizing is the standard probe: disable decode, write all-ones, read
/// back the writable (address) bits, restore. `size = ~mask + 1`.
fn addBars(node: *device_model.Device, configuration: [*]align(1) u8) void {
// Stop the device decoding its BARs while we transiently write all-ones.
const command = rd(u16, configuration, 0x04);
wr(u16, configuration, 0x04, command & ~@as(u16, 0b11));
var i: usize = 0;
while (i < 6) : (i += 1) {
const off = 0x10 + i * 4;
const orig = rd(u32, configuration, off);
if (orig == 0) continue;
if (orig & 1 != 0) {
// I/O-space BAR (16-bit address space on x86).
wr(u32, configuration, off, 0xFFFF_FFFF);
const readback = rd(u32, configuration, off);
wr(u32, configuration, off, orig);
const mask = readback & 0xFFFF_FFFC;
const size: u32 = if (mask == 0) 0 else (~mask +% 1) & 0xFFFF;
_ = node.addResource(.io_port, orig & 0xFFFF_FFFC, size);
} else if ((orig >> 1) & 0x3 == 2) {
// 64-bit memory BAR: this BAR pair spans two configuration slots.
const orig_hi = rd(u32, configuration, off + 4);
wr(u32, configuration, off, 0xFFFF_FFFF);
wr(u32, configuration, off + 4, 0xFFFF_FFFF);
const lo = rd(u32, configuration, off);
const hi = rd(u32, configuration, off + 4);
wr(u32, configuration, off, orig);
wr(u32, configuration, off + 4, orig_hi);
const readback = (@as(u64, hi) << 32) | (lo & 0xFFFF_FFF0);
const size: u64 = if (readback == 0) 0 else ~readback +% 1;
const address = (@as(u64, orig_hi) << 32) | (orig & 0xFFFF_FFF0);
_ = node.addResource(.memory, address, size);
i += 1; // consumed the high half
} else {
// 32-bit memory BAR.
wr(u32, configuration, off, 0xFFFF_FFFF);
const readback = rd(u32, configuration, off);
wr(u32, configuration, off, orig);
const mask = readback & 0xFFFF_FFF0;
const size: u32 = if (mask == 0) 0 else ~mask +% 1;
_ = node.addResource(.memory, orig & 0xFFFF_FFF0, size);
} }
} _ = bridge.addResource(.memory, high_end, (@as(u64, 1) << 46) - high_end);
wr(u16, configuration, 0x04, command); // restore decode
} }
/// HPET -> a timer node with its register block as an MMIO resource, plus the GSI /// HPET -> a timer node with its register block as an MMIO resource, plus the GSI
@@ -712,6 +729,7 @@ const fadt_pm1a_cnt_blk = 64; // u32 (I/O port)
const fadt_pm1b_cnt_blk = 68; // u32 (I/O port) const fadt_pm1b_cnt_blk = 68; // u32 (I/O port)
const fadt_pm_tmr_blk = 76; // u32 (I/O port) — the PM timer counter const fadt_pm_tmr_blk = 76; // u32 (I/O port) — the PM timer counter
const fadt_pm1_cnt_len = 89; // u8 (bytes) const fadt_pm1_cnt_len = 89; // u8 (bytes)
const fadt_sci_int = 46; // u16 (the SCI's GSI)
const fadt_flags = 112; // u32 const fadt_flags = 112; // u32
const fadt_reset_register = 116; // GAS (12 bytes) const fadt_reset_register = 116; // GAS (12 bytes)
const fadt_reset_value = 128; // u8 const fadt_reset_value = 128; // u8
@@ -729,6 +747,7 @@ fn parseFadt(header: *const SystemDescriptorTableHeader) void {
const len: usize = header.length; const len: usize = header.length;
const pi = &power_information; const pi = &power_information;
pi.sci_interrupt = @truncate(fadt(u16, base, len, fadt_sci_int) orelse 0);
pi.smi_cmd = @truncate(fadt(u32, base, len, fadt_smi_cmd) orelse 0); pi.smi_cmd = @truncate(fadt(u32, base, len, fadt_smi_cmd) orelse 0);
pi.acpi_enable = fadt(u8, base, len, fadt_acpi_enable) orelse 0; pi.acpi_enable = fadt(u8, base, len, fadt_acpi_enable) orelse 0;
pi.acpi_disable = fadt(u8, base, len, fadt_acpi_disable) orelse 0; pi.acpi_disable = fadt(u8, base, len, fadt_acpi_disable) orelse 0;
@@ -808,342 +827,6 @@ fn parseDmar(hal: Hal, header: *const SystemDescriptorTableHeader) void {
} }
} }
// --- AML namespace -> generic device tree -----------------------------------
/// The PCI bus context while descending the ACPI namespace: the generic host
/// bridge whose children ACPI address (`_ADR`) devices resolve against, and the bus number.
const PciContext = struct { bridge: *device_model.Device, bus: u8 };
/// Mirror the ACPI namespace's Device objects into the generic tree, *merging*
/// them with the PCI-enumerated nodes: a PCI root bridge (`PNP0A03`/`PNP0A08`)
/// folds onto the existing `pci_host_bridge`, and each addressed (`_ADR`) device folds onto
/// the matching PCI function (annotating it with the ACPI hardware ID (`_HID`) and nesting the
/// ACPI-only children — keyboard, RTC, … — beneath it). Namespace devices with no
/// PCI match land under a synthetic `acpi` node.
fn wireAcpiDevices(device_tree: *DeviceTree, aml_namespace: *aml.Namespace, hal: Hal) !void {
var arena = std.heap.ArenaAllocator.init(device_tree.allocator);
defer arena.deinit();
var interpreter = aml.Interpreter.init(aml_namespace, .{
.mapMmio = hal.mapMmio,
.pioRead = hal.pioRead,
.pioWrite = hal.pioWrite,
}, arena.allocator());
const acpi_root = try device_tree.addChild(device_tree.root, .unknown, "acpi");
try mirrorDevices(device_tree, aml_namespace.root, acpi_root, null, &interpreter);
}
fn mirrorDevices(device_tree: *DeviceTree, node: *aml.Node, parent_device: *device_model.Device, context: ?PciContext, interpreter: *aml.Interpreter) (error{OutOfMemory})!void {
var child = node.first_child;
while (child) |c| : (child = c.next_sibling) {
if (c.kind != .device) {
// A scope — the System Bus (\_SB), General Purpose Events (\_GPE), … —
// descend without adding a node.
try mirrorDevices(device_tree, c, parent_device, context, interpreter);
continue;
}
// Skip devices the firmware reports as not present (via a device-status (`_STA`) method),
// along with their whole subtree — per the ACPI rules.
if (!devicePresent(interpreter, c)) continue;
var mirrored_device: *device_model.Device = undefined;
var child_context = context;
if (isPciRootNode(c)) {
// The PCI root bridge folds onto the generic host bridge.
mirrored_device = matchHostBridge(device_tree) orelse
try device_tree.addChild(parent_device, .acpi_device, &c.segment);
child_context = .{ .bridge = mirrored_device, .bus = 0 };
} else {
// An addressed device folds onto its matching PCI function; anything
// else becomes a fresh node under the current parent.
mirrored_device = pick: {
if (context) |pc| {
if (readAdr(c)) |adr| {
if (findPciNode(pc.bridge, pc.bus, adr)) |pnode| break :pick pnode;
}
}
break :pick try device_tree.addChild(parent_device, .acpi_device, &c.segment);
};
}
applyHid(mirrored_device, c, interpreter);
applyCrs(mirrored_device, c, interpreter);
try mirrorDevices(device_tree, c, mirrored_device, child_context, interpreter);
}
}
/// Evaluate a device's status (`_STA`) to decide if it is present. An absent status
/// (`_STA`) means present by default; an evaluation failure is treated as present too (we'd
/// rather over-report than hide a device we couldn't introspect).
fn devicePresent(interpreter: *aml.Interpreter, node: *aml.Node) bool {
const sta = aml.Namespace.childOf(node, seg4("_STA")) orelse return true;
const obj = interpreter.evaluate(sta, &.{}) catch return true;
const status = obj.asInteger() catch return true;
return (status & 0x01) != 0; // bit 0 = present
}
/// The first PCI host bridge in the generic tree (segment 0).
fn matchHostBridge(device_tree: *DeviceTree) ?*device_model.Device {
var c = device_tree.root.first_child;
while (c) |ch| : (c = ch.next_sibling) {
if (ch.class == .pci_host_bridge) return ch;
}
return null;
}
/// The PCI function node under `bridge` at the address the device's address object
/// (`_ADR`) names (device/function on
/// `bus`), or null.
fn findPciNode(bridge: *device_model.Device, bus: u8, adr: u32) ?*device_model.Device {
const device: u16 = @truncate((adr >> 16) & 0x1F);
const function: u16 = @truncate(adr & 0x7);
const target: u16 = (@as(u16, bus) << 8) | (device << 3) | function;
var c = bridge.first_child;
while (c) |ch| : (c = ch.next_sibling) {
if (ch.ids.pci_bdf) |bdf| {
if (bdf == target) return ch;
}
}
return null;
}
/// A device's address (`_ADR`) — a static integer Name — or null.
fn readAdr(node: *aml.Node) ?u32 {
const n = aml.Namespace.childOf(node, seg4("_ADR")) orelse return null;
if (n.kind != .name) return null;
var p: usize = 0;
return @truncate(readIntObj(n.value, &p) orelse return null);
}
/// Whether a `_HID` string names a PCI(e) host bridge.
fn isPciRootHid(hid: []const u8) bool {
const id = acpi_ids.HardwareId.fromHid(hid) orelse return false;
return id == .pci_bus or id == .pci_express_root_bridge;
}
/// Whether a namespace device is a PCI(e) host bridge. A packed EISA id is decoded
/// to its string form first, so both encodings answer through the one registry.
fn isPciRootNode(node: *aml.Node) bool {
const hid = aml.Namespace.childOf(node, seg4("_HID")) orelse return false;
if (hid.kind != .name or hid.value.len == 0) return false;
switch (hid.value[0]) {
0x00, 0x01, 0xFF, 0x0A, 0x0B, 0x0C, 0x0E => {
var p: usize = 0;
const n = readIntObj(hid.value, &p) orelse return false;
var buffer: [8]u8 = undefined;
return isPciRootHid(eisaIdToStr(@truncate(n), &buffer));
},
0x0D => return isPciRootHid(cstr(hid.value[1..])),
else => return false,
}
}
/// Read a device's hardware ID (`_HID`) into the generic device: an integer decodes as an EISA
/// id ("PNP0A03"), a string is taken verbatim. Handles both the common static
/// Name form and a Method form (evaluated).
fn applyHid(device: *device_model.Device, node: *aml.Node, interpreter: *aml.Interpreter) void {
const hid = aml.Namespace.childOf(node, seg4("_HID")) orelse return;
if (hid.kind == .method) {
const obj = interpreter.evaluate(hid, &.{}) catch return;
switch (obj) {
.integer => |n| setEisaHid(device, @truncate(n)),
.string => |s| device.setHid(s),
else => {},
}
return;
}
if (hid.kind != .name or hid.value.len == 0) return;
const v = hid.value;
switch (v[0]) {
0x00, 0x01, 0xFF, 0x0A, 0x0B, 0x0C, 0x0E => {
var p: usize = 0;
const n = readIntObj(v, &p) orelse return;
setEisaHid(device, @truncate(n));
},
0x0D => device.setHid(cstr(v[1..])), // StringPrefix
else => {},
}
}
fn setEisaHid(device: *device_model.Device, id: u32) void {
device.ids.acpi_hid = id;
var buffer: [8]u8 = undefined;
device.setHid(eisaIdToStr(id, &buffer));
}
/// Parse a device's current resource settings (`_CRS`). The evaluator handles both the static
/// `Buffer` form (a `Name`) and the method form uniformly, yielding the
/// ResourceTemplate bytes we then decode.
fn applyCrs(device: *device_model.Device, node: *aml.Node, interpreter: *aml.Interpreter) void {
const crs = aml.Namespace.childOf(node, seg4("_CRS")) orelse return;
const obj = interpreter.evaluate(crs, &.{}) catch return;
const buffer = switch (obj) {
.buffer => |b| b,
else => return,
};
parseResourceTemplate(device, buffer);
}
/// Walk a ResourceTemplate byte list, adding recognised descriptors as resources.
fn parseResourceTemplate(device: *device_model.Device, bytes: []const u8) void {
var i: usize = 0;
while (i < bytes.len) {
const tag = bytes[i];
if (tag & 0x80 == 0) {
// Small descriptor: length in low 3 bits, type in bits [6:3].
const len: usize = tag & 0x07;
const body = i + 1;
if (body + len > bytes.len) break;
switch ((tag >> 3) & 0x0F) {
0x04 => if (len >= 2) { // IRQ: a 16-bit mask, one resource per set bit
const mask = @as(u16, bytes[body]) | (@as(u16, bytes[body + 1]) << 8);
var b: usize = 0;
while (b < 16) : (b += 1) {
if (mask & (@as(u16, 1) << @intCast(b)) != 0) _ = device.addResource(.irq, b, 1);
}
},
0x08 => if (len >= 7) { // IO port: minimum at +1, length at +6
_ = device.addResource(.io_port, rd16(bytes, body + 1), bytes[body + 6]);
},
0x09 => if (len >= 3) { // Fixed IO: base at +0, length at +2
_ = device.addResource(.io_port, rd16(bytes, body), bytes[body + 2]);
},
0x0F => break, // EndTag
else => {},
}
i = body + len;
} else {
// Large descriptor: 16-bit length follows the tag.
if (i + 3 > bytes.len) break;
const len: usize = @intCast(rd16(bytes, i + 1));
const body = i + 3;
if (body + len > bytes.len) break;
switch (tag) {
0x85 => if (len >= 17) { // Memory32: minimum at +1, length at +13
_ = device.addResource(.memory, rd32(bytes, body + 1), rd32(bytes, body + 13));
},
0x86 => if (len >= 9) { // Memory32Fixed: base at +1, length at +5
_ = device.addResource(.memory, rd32(bytes, body + 1), rd32(bytes, body + 5));
},
0x89 => if (len >= 2) { // Extended IRQ: count at +1, then count u32s
const count = bytes[body + 1];
var k: usize = 0;
while (k < count and body + 2 + k * 4 + 4 <= body + len) : (k += 1) {
_ = device.addResource(.irq, rd32(bytes, body + 2 + k * 4), 1);
}
},
0x87, 0x88, 0x8A => parseAddressSpace(device, tag, bytes[body .. body + len]),
else => {},
}
i = body + len;
}
}
}
/// Word/DWord/QWord address-space descriptors: resource type at [0], then
/// granularity/minimum/maximum/translation/length, each of width `w`.
fn parseAddressSpace(device: *device_model.Device, tag: u8, body: []const u8) void {
const w: usize = switch (tag) {
0x88 => 2, // Word
0x87 => 4, // DWord
else => 8, // QWord (0x8A)
};
if (body.len < 3 + 5 * w) return;
const minimum = readN(body, 3 + w, w);
const length = readN(body, 3 + 4 * w, w);
const kind: device_model.ResourceKind = switch (body[0]) {
0 => .memory,
1 => .io_port,
else => .bus_range,
};
_ = device.addResource(kind, minimum, length);
}
/// Decode a packed EISA id into its 7-char string (e.g. 0x030AD041 -> "PNP0A03").
fn eisaIdToStr(id: u32, buffer: *[8]u8) []const u8 {
const b0: u16 = @intCast(id & 0xFF);
const b1: u16 = @intCast((id >> 8) & 0xFF);
const b2: u8 = @truncate(id >> 16);
const b3: u8 = @truncate(id >> 24);
const mfg = (b0 << 8) | b1;
buffer[0] = '@' + @as(u8, @intCast((mfg >> 10) & 0x1F));
buffer[1] = '@' + @as(u8, @intCast((mfg >> 5) & 0x1F));
buffer[2] = '@' + @as(u8, @intCast(mfg & 0x1F));
buffer[3] = hexDigit((b2 >> 4) & 0xF);
buffer[4] = hexDigit(b2 & 0xF);
buffer[5] = hexDigit((b3 >> 4) & 0xF);
buffer[6] = hexDigit(b3 & 0xF);
return buffer[0..7];
}
fn hexDigit(n: u8) u8 {
return if (n < 10) '0' + n else 'A' + (n - 10);
}
fn seg4(comptime s: *const [4:0]u8) [4]u8 {
return s[0..4].*;
}
fn cstr(bytes: []const u8) []const u8 {
const index = std.mem.indexOfScalar(u8, bytes, 0) orelse bytes.len;
return bytes[0..index];
}
const PkgLen = struct { value: usize, size: usize };
fn packageLength(bytes: []const u8, p: usize) ?PkgLen {
if (p >= bytes.len) return null;
const lead = bytes[p];
const follow: usize = lead >> 6;
if (p + 1 + follow > bytes.len) return null;
if (follow == 0) return .{ .value = lead & 0x3F, .size = 1 };
var value: usize = lead & 0x0F;
var i: usize = 0;
while (i < follow) : (i += 1) value |= @as(usize, bytes[p + 1 + i]) << @intCast(4 + i * 8);
return .{ .value = value, .size = 1 + follow };
}
/// Read an AML integer object at `p`, advancing `p` past it.
fn readIntObj(bytes: []const u8, p: *usize) ?u64 {
if (p.* >= bytes.len) return null;
const opcode = bytes[p.*];
p.* += 1;
return switch (opcode) {
0x00 => 0,
0x01 => 1,
0xFF => 0xFF,
0x0A => readLE(bytes, p, 1),
0x0B => readLE(bytes, p, 2),
0x0C => readLE(bytes, p, 4),
0x0E => readLE(bytes, p, 8),
else => null,
};
}
fn readLE(bytes: []const u8, p: *usize, n: usize) ?u64 {
if (p.* + n > bytes.len) return null;
const v = readN(bytes, p.*, n);
p.* += n;
return v;
}
fn readN(bytes: []const u8, off: usize, n: usize) u64 {
var v: u64 = 0;
var k: usize = 0;
while (k < n and off + k < bytes.len) : (k += 1) v |= @as(u64, bytes[off + k]) << @intCast(k * 8);
return v;
}
fn rd16(bytes: []const u8, off: usize) u64 {
return readN(bytes, off, 2);
}
fn rd32(bytes: []const u8, off: usize) u64 {
return readN(bytes, off, 4);
}
// --- helpers ---------------------------------------------------------------- // --- helpers ----------------------------------------------------------------
/// Sum `len` bytes; an ACPI table/pointer is valid when the low 8 bits are zero. /// Sum `len` bytes; an ACPI table/pointer is valid when the low 8 bits are zero.
@@ -1184,58 +867,9 @@ fn readCntRegister(base: [*]align(1) const u8, len: usize, xoff: usize, legacy_o
return .{ .mmio = false, .address = port, .width = width }; return .{ .mmio = false, .address = port, .width = width };
} }
/// The mapped configuration space of one PCI function (its 4 KiB ECAM page). Mapped
/// writable so BAR sizing can probe it; reads and writes both go through here.
fn pciConfigurationPtr(alloc: McfgAllocation, hal: Hal, bus: u8, device: u8, function: u8) [*]align(1) u8 {
const physical = alloc.base_address +
(@as(u64, bus - alloc.start_bus) << 20) +
(@as(u64, device) << 15) +
(@as(u64, function) << 12);
// Map the configuration page (writable, for BAR sizing) and use the virtual
// address the HAL hands back.
return @ptrFromInt(hal.mapMmio(physical, abi.page_size, true));
}
/// Read a little-endian integer at `off` from a (possibly unaligned) byte pointer. /// Read a little-endian integer at `off` from a (possibly unaligned) byte pointer.
/// x86 is little-endian and native, so an unaligned load suffices. /// x86 is little-endian and native, so an unaligned load suffices.
fn rd(comptime T: type, bytes: [*]align(1) const u8, off: usize) T { fn rd(comptime T: type, bytes: [*]align(1) const u8, off: usize) T {
const p: *align(1) const T = @ptrCast(bytes + off); const p: *align(1) const T = @ptrCast(bytes + off);
return p.*; return p.*;
} }
/// Write a little-endian integer at `off` through a (possibly unaligned) pointer.
fn wr(comptime T: type, bytes: [*]align(1) u8, off: usize, value: T) void {
const p: *align(1) T = @ptrCast(bytes + off);
p.* = value;
}
// --- tests ------------------------------------------------------------------
test "eisaIdToStr decodes a packed EISA id" {
var buffer: [8]u8 = undefined;
// 0x030AD041 is the well-known encoding of "PNP0A03" (PCI root bridge).
try std.testing.expectEqualStrings("PNP0A03", eisaIdToStr(0x030AD041, &buffer));
}
test "parseResourceTemplate extracts IO, IRQ, and fixed memory" {
// ResourceTemplate { IO(minimum 0x60, len 8), IRQ(4), Memory32Fixed(0xFED00000, 0x1000) }
const runtime = [_]u8{
0x47, 0x01, 0x60, 0x00, 0x60, 0x00, 0x01, 0x08, // small IO descriptor
0x22, 0x10, 0x00, // small IRQ descriptor (mask bit 4 -> IRQ 4)
0x86, 0x09, 0x00, 0x01, 0x00, 0x00, 0xD0, 0xFE, 0x00, 0x10, 0x00, 0x00, // Memory32Fixed
0x79, 0x00, // EndTag
};
var device = device_model.Device{};
parseResourceTemplate(&device, &runtime);
try std.testing.expectEqual(@as(u8, 3), device.resource_count);
const rs = device.resources[0..device.resource_count];
try std.testing.expectEqual(device_model.ResourceKind.io_port, rs[0].kind);
try std.testing.expectEqual(@as(u64, 0x60), rs[0].start);
try std.testing.expectEqual(@as(u64, 8), rs[0].len);
try std.testing.expectEqual(device_model.ResourceKind.irq, rs[1].kind);
try std.testing.expectEqual(@as(u64, 4), rs[1].start);
try std.testing.expectEqual(device_model.ResourceKind.memory, rs[2].kind);
try std.testing.expectEqual(@as(u64, 0xFED00000), rs[2].start);
try std.testing.expectEqual(@as(u64, 0x1000), rs[2].len);
}
+47
View File
@@ -12,6 +12,11 @@ const std = @import("std");
const opcode = @import("opcodes.zig"); const opcode = @import("opcodes.zig");
const parser = @import("parser.zig"); const parser = @import("parser.zig");
/// The named AML opcode/prefix bytes (`zero_opcode`, `byte_prefix`, …). Re-exported so
/// callers that decode raw AML bytes — e.g. the acpi service reading a `_HID` integer —
/// name the opcodes instead of writing bare 0x0A/0x0B/… literals (docs/coding-standards.md).
pub const opcodes = @import("opcodes.zig");
pub const Namespace = @import("namespace.zig").Namespace; pub const Namespace = @import("namespace.zig").Namespace;
pub const Node = @import("namespace.zig").Node; pub const Node = @import("namespace.zig").Node;
pub const NodeKind = @import("namespace.zig").NodeKind; pub const NodeKind = @import("namespace.zig").NodeKind;
@@ -50,6 +55,20 @@ pub fn parse(allocator: std.mem.Allocator, blocks: []const []const u8) !ParseRes
return .{ .namespace = namespace, .consumed = consumed, .total = total }; return .{ .namespace = namespace, .consumed = consumed, .total = total };
} }
/// Count the Device objects in a parsed namespace — what the acpi service
/// (docs/m19-m20-plan.md M20) reports, and what the kernel's own parse counts
/// so the two can be checked equal across the ring-3 move.
pub fn deviceCount(namespace: *const Namespace) usize {
return countKind(namespace.root, .device);
}
fn countKind(node: *const Node, kind: NodeKind) usize {
var n: usize = if (node.kind == kind) 1 else 0;
var c = node.first_child;
while (c) |child| : (c = child.next_sibling) n += countKind(child, kind);
return n;
}
/// Look up the `\_S{state}` sleep package in a parsed namespace and return its /// Look up the `\_S{state}` sleep package in a parsed namespace and return its
/// first two integer elements (SLP_TYP for PM1a / PM1b), or null if absent. /// first two integer elements (SLP_TYP for PM1a / PM1b), or null if absent.
pub fn sleepState(namespace: *Namespace, state: u8) ?SleepType { pub fn sleepState(namespace: *Namespace, state: u8) ?SleepType {
@@ -211,3 +230,31 @@ test "interpreter runs a method with args, arithmetic, and control flow" {
const lo = try interpreter.evaluate(tst, &.{.{ .integer = 2 }}); // 2+5=7 !> 10 -> 0 const lo = try interpreter.evaluate(tst, &.{.{ .integer = 2 }}); // 2+5=7 !> 10 -> 0
try std.testing.expectEqual(@as(u64, 0), try lo.asInteger()); try std.testing.expectEqual(@as(u64, 0), try lo.asInteger());
} }
test "interpreter records Notify(device, code)" {
// Device(DEV_) { Name(_HID, 0x030AD041) } // PNP0A03-ish placeholder
// Method(TST_, 0) { Notify(DEV_, 0x80); Return(Zero) }
// Encoded: a Device holding a Name, then a Method issuing Notify on it.
const blob = [_]u8{
0x5B, 0x82, 0x0F, 0x44, 0x45, 0x56, 0x5F, // Device(DEV_) len=0x0F (pkglen + DEV_ + Name)
0x08, 0x5F, 0x48, 0x49, 0x44, 0x0C, 0x41, 0xD0, 0x0A, 0x03, // Name(_HID, DWord 0x030AD041)
0x14, 0x0F, 0x54, 0x53, 0x54, 0x5F, 0x00, // Method(TST_, 0) len=0x0F (pkglen + TST_ + flags + body)
0x86, 0x44, 0x45, 0x56, 0x5F, 0x0A, 0x80, // Notify(DEV_, 0x80)
0xA4, 0x00, // Return(Zero)
};
var arena = std.heap.ArenaAllocator.init(std.testing.allocator);
defer arena.deinit();
var result = try parse(arena.allocator(), &.{&blob});
const namespace = &result.namespace;
const tst = namespace.resolve(namespace.root, false, 0, &.{.{ 'T', 'S', 'T', '_' }}) orelse return error.NoMethod;
const dev = namespace.resolve(namespace.root, false, 0, &.{.{ 'D', 'E', 'V', '_' }}) orelse return error.NoDevice;
var interpreter = Interpreter.init(namespace, .{ .mapMmio = noMap, .pioRead = noRead, .pioWrite = noWrite }, arena.allocator());
_ = try interpreter.evaluate(tst, &.{});
const events = interpreter.takeNotifications();
try std.testing.expectEqual(@as(usize, 1), events.len);
try std.testing.expectEqual(dev, events[0].node);
try std.testing.expectEqual(@as(u64, 0x80), events[0].code);
}
+41
View File
@@ -141,6 +141,9 @@ const Frame = struct {
/// A CreateField binding: a name that indexes into a buffer object. /// A CreateField binding: a name that indexes into a buffer object.
const BufferField = struct { buffer: *Node, byte_off: usize, bit_width: u32 }; const BufferField = struct { buffer: *Node, byte_off: usize, bit_width: u32 };
/// One Notify(device, code) the interpreter executed.
pub const NotifyEvent = struct { node: *Node, code: u64 };
pub const Interpreter = struct { pub const Interpreter = struct {
namespace: *Namespace, namespace: *Namespace,
hal: Hal, hal: Hal,
@@ -149,6 +152,11 @@ pub const Interpreter = struct {
dynamic_overrides: std.AutoHashMapUnmanaged(*Node, Object) = .{}, dynamic_overrides: std.AutoHashMapUnmanaged(*Node, Object) = .{},
/// CreateField bindings active for the current evaluation. /// CreateField bindings active for the current evaluation.
fields: std.AutoHashMapUnmanaged(*Node, BufferField) = .{}, fields: std.AutoHashMapUnmanaged(*Node, BufferField) = .{},
/// Notify(device, code) operations the last evaluation executed — a GPE or
/// EC handler tells the OS "look at this device" this way. Bounded; the
/// caller drains it with `takeNotifications` after `evaluate` (M21).
notify_queue: [16]NotifyEvent = undefined,
notify_count: usize = 0,
pub fn init(namespace: *Namespace, hal: Hal, arena: std.mem.Allocator) Interpreter { pub fn init(namespace: *Namespace, hal: Hal, arena: std.mem.Allocator) Interpreter {
return .{ .namespace = namespace, .hal = hal, .arena = arena }; return .{ .namespace = namespace, .hal = hal, .arena = arena };
@@ -157,6 +165,7 @@ pub const Interpreter = struct {
/// Evaluate a namespace object: invoke a Method, read a Name's value, or read a /// Evaluate a namespace object: invoke a Method, read a Name's value, or read a
/// Field. Resets per-evaluation runtime state first. /// Field. Resets per-evaluation runtime state first.
pub fn evaluate(self: *Interpreter, node: *Node, args: []const Object) Error!Object { pub fn evaluate(self: *Interpreter, node: *Node, args: []const Object) Error!Object {
self.notify_count = 0;
self.dynamic_overrides.clearRetainingCapacity(); self.dynamic_overrides.clearRetainingCapacity();
self.fields.clearRetainingCapacity(); self.fields.clearRetainingCapacity();
return self.invoke(node, args); return self.invoke(node, args);
@@ -267,6 +276,8 @@ pub const Interpreter = struct {
}, },
opcode.to_buffer_opcode => try self.passThroughUnary(current, frame), opcode.to_buffer_opcode => try self.passThroughUnary(current, frame),
opcode.notify_opcode => try self.notify(current, frame),
opcode.extended_opcode_prefix => try self.ext(current, frame), opcode.extended_opcode_prefix => try self.ext(current, frame),
// CreateXField: source, index, name (bit widths differ by op) // CreateXField: source, index, name (bit widths differ by op)
@@ -542,6 +553,36 @@ pub const Interpreter = struct {
try self.storeInto(current, frame, value); try self.storeInto(current, frame, value);
} }
/// Notify(SuperName, NotifyValue): resolve the named device, evaluate the
/// code, and record the pair for the caller to dispatch. AML control flow
/// continues (Notify returns nothing).
fn notify(self: *Interpreter, current: *Cursor, frame: *Frame) Error!Object {
const lead = current.peek() orelse return error.Truncated;
var target: ?*Node = null;
if (isNameStart(lead)) {
const name_path = try current.nameString();
target = self.namespace.resolve(frame.scope, name_path.rooted, name_path.parents, name_path.slice());
} else {
// A non-name SuperName (Local/Arg holding a reference).
const obj = try self.term(current, frame);
if (obj == .reference) target = obj.reference;
}
const code = try self.evaluateInteger(current, frame);
if (target) |node| {
if (self.notify_count < self.notify_queue.len) {
self.notify_queue[self.notify_count] = .{ .node = node, .code = code };
self.notify_count += 1;
}
}
return .uninitialized;
}
/// The Notify events the last `evaluate` produced. Valid until the next
/// `evaluate` clears the queue.
pub fn takeNotifications(self: *Interpreter) []const NotifyEvent {
return self.notify_queue[0..self.notify_count];
}
fn storeInto(self: *Interpreter, current: *Cursor, frame: *Frame, value: Object) Error!void { fn storeInto(self: *Interpreter, current: *Cursor, frame: *Frame, value: Object) Error!void {
const lead = current.peek() orelse return error.Truncated; const lead = current.peek() orelse return error.Truncated;
if (isNameStart(lead)) { if (isNameStart(lead)) {
+5
View File
@@ -28,6 +28,11 @@ pub const DeviceClass = enum(u32) {
/// A device named in the ACPI namespace (from the DSDT/SSDT), carrying a /// A device named in the ACPI namespace (from the DSDT/SSDT), carrying a
/// hardware ID (`_HID`) and, where static, current resource settings (`_CRS`). /// hardware ID (`_HID`) and, where static, current resource settings (`_CRS`).
acpi_device, acpi_device,
/// The ACPI tables themselves, published as one node for the user-space acpi
/// service (docs/m19-m20-plan.md M20): memory resources over the AML blobs,
/// a broad io_port grant for OperationRegion access, and the SCI interrupt.
/// The one node whose claimant is trusted to run firmware bytecode.
acpi_tables,
unknown, unknown,
}; };
+496 -189
View File
@@ -7,6 +7,16 @@
//! apart. Pure reference data (from the PCI spec; see https://wiki.osdev.org/PCI) — no //! apart. Pure reference data (from the PCI spec; see https://wiki.osdev.org/PCI) — no
//! hardware access — so it is shared by kernel discovery (the device-tree dump) and any //! hardware access — so it is shared by kernel discovery (the device-tree dump) and any
//! user-space tool (a future lspci, driver matching). //! user-space tool (a future lspci, driver matching).
//!
//! The taxonomy is named, not numbered (docs/coding-standards.md, "Named values"): the
//! base class is a `BaseClass` enum, and each class with defined subclasses gets a
//! namespace holding its `SubClass` enum (and, where the spec defines them, per-subclass
//! `ProgIf` enums) — the same shape as `usb-ids.zig`. Code that *means* a specific class
//! names it (`BaseClass.serial_bus`, `serial_bus.usb.ProgIf.xhci`) rather than writing a
//! bare 0x0C/0x03/0x30. The `className`/`subclassName`/`progIfName` functions still take
//! the raw bytes a function reports in its header, because that is what hardware hands us.
const std = @import("std");
/// The three bytes of a PCI class code, unpacked from the `0xCCSSPP` value discovery /// The three bytes of a PCI class code, unpacked from the `0xCCSSPP` value discovery
/// records in `Device.ids.pci_class` (CC = base class, SS = subclass, PP = prog-IF). /// records in `Device.ids.pci_class` (CC = base class, SS = subclass, PP = prog-IF).
@@ -22,148 +32,465 @@ pub const ClassCode = struct {
.prog_if = @intCast(packed_code & 0xFF), .prog_if = @intCast(packed_code & 0xFF),
}; };
} }
/// Re-pack the triple into the `0xCCSSPP` form. Lets code name a whole class code
/// from its parts — `pack(.{ .base = @intFromEnum(BaseClass.serial_bus), … })` —
/// instead of writing the literal 0x0C0330.
pub fn pack(self: ClassCode) u24 {
return (@as(u24, self.base) << 16) | (@as(u24, self.subclass) << 8) | self.prog_if;
}
}; };
/// Base class (config byte 0x0B). Non-exhaustive: an unlisted code is a real but
/// unnamed class, decoded as "Unknown" rather than rejected.
pub const BaseClass = enum(u8) {
unclassified = 0x00,
mass_storage = 0x01,
network = 0x02,
display = 0x03,
multimedia = 0x04,
memory = 0x05,
bridge = 0x06,
simple_communication = 0x07,
base_system_peripheral = 0x08,
input_device = 0x09,
docking_station = 0x0A,
processor = 0x0B,
serial_bus = 0x0C,
wireless = 0x0D,
intelligent = 0x0E,
satellite_communication = 0x0F,
encryption = 0x10,
signal_processing = 0x11,
processing_accelerator = 0x12,
non_essential_instrumentation = 0x13,
co_processor = 0x40,
unassigned = 0xFF,
_,
pub fn name(self: BaseClass) []const u8 {
return switch (self) {
.unclassified => "Unclassified",
.mass_storage => "Mass Storage Controller",
.network => "Network Controller",
.display => "Display Controller",
.multimedia => "Multimedia Controller",
.memory => "Memory Controller",
.bridge => "Bridge",
.simple_communication => "Simple Communication Controller",
.base_system_peripheral => "Base System Peripheral",
.input_device => "Input Device Controller",
.docking_station => "Docking Station",
.processor => "Processor",
.serial_bus => "Serial Bus Controller",
.wireless => "Wireless Controller",
.intelligent => "Intelligent Controller",
.satellite_communication => "Satellite Communication Controller",
.encryption => "Encryption Controller",
.signal_processing => "Signal Processing Controller",
.processing_accelerator => "Processing Accelerator",
.non_essential_instrumentation => "Non-Essential Instrumentation",
.co_processor => "Co-Processor",
.unassigned => "Unassigned Class (Vendor specific)",
_ => "Unknown",
};
}
};
// --- Per-class subclass (and prog-IF) taxonomies --------------------------------------
// One namespace per base class that has defined subclasses, named after the class. Each
// holds an exhaustive `SubClass` enum (so an unlisted code decodes to the class default,
// not a wrong name), and, where the spec assigns them, per-subclass `ProgIf` enums.
pub const mass_storage = struct {
pub const SubClass = enum(u8) {
scsi_bus = 0x00,
ide = 0x01,
floppy = 0x02,
ipi_bus = 0x03,
raid = 0x04,
ata = 0x05,
serial_ata = 0x06,
serial_attached_scsi = 0x07,
non_volatile_memory = 0x08,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.scsi_bus => "SCSI Bus Controller",
.ide => "IDE Controller",
.floppy => "Floppy Disk Controller",
.ipi_bus => "IPI Bus Controller",
.raid => "RAID Controller",
.ata => "ATA Controller",
.serial_ata => "Serial ATA Controller",
.serial_attached_scsi => "Serial Attached SCSI Controller",
.non_volatile_memory => "Non-Volatile Memory Controller",
};
}
};
pub const serial_ata = struct {
pub const ProgIf = enum(u8) {
vendor_specific = 0x00,
ahci = 0x01,
serial_storage_bus = 0x02,
pub fn name(self: ProgIf) []const u8 {
return switch (self) {
.vendor_specific => "Vendor Specific Interface",
.ahci => "AHCI 1.0",
.serial_storage_bus => "Serial Storage Bus",
};
}
};
};
pub const non_volatile_memory = struct {
pub const ProgIf = enum(u8) {
nvmhci = 0x01,
nvm_express = 0x02,
pub fn name(self: ProgIf) []const u8 {
return switch (self) {
.nvmhci => "NVMHCI",
.nvm_express => "NVM Express",
};
}
};
};
};
pub const network = struct {
pub const SubClass = enum(u8) {
ethernet = 0x00,
token_ring = 0x01,
fddi = 0x02,
atm = 0x03,
isdn = 0x04,
picmg_multi_computing = 0x06,
infiniband = 0x07,
fabric = 0x08,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.ethernet => "Ethernet Controller",
.token_ring => "Token Ring Controller",
.fddi => "FDDI Controller",
.atm => "ATM Controller",
.isdn => "ISDN Controller",
.picmg_multi_computing => "PICMG 2.14 Multi Computing Controller",
.infiniband => "Infiniband Controller",
.fabric => "Fabric Controller",
};
}
};
};
pub const display = struct {
pub const SubClass = enum(u8) {
vga_compatible = 0x00,
xga = 0x01,
three_dimensional = 0x02,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.vga_compatible => "VGA Compatible Controller",
.xga => "XGA Controller",
.three_dimensional => "3D Controller (Not VGA-Compatible)",
};
}
};
pub const vga_compatible = struct {
pub const ProgIf = enum(u8) {
vga = 0x00,
compatible_8514 = 0x01,
pub fn name(self: ProgIf) []const u8 {
return switch (self) {
.vga => "VGA Controller",
.compatible_8514 => "8514-Compatible Controller",
};
}
};
};
};
pub const multimedia = struct {
pub const SubClass = enum(u8) {
video = 0x00,
audio = 0x01,
telephony = 0x02,
audio_device = 0x03,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.video => "Multimedia Video Controller",
.audio => "Multimedia Audio Controller",
.telephony => "Computer Telephony Device",
.audio_device => "Audio Device",
};
}
};
};
pub const memory = struct {
pub const SubClass = enum(u8) {
ram = 0x00,
flash = 0x01,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.ram => "RAM Controller",
.flash => "Flash Controller",
};
}
};
};
pub const bridge = struct {
pub const SubClass = enum(u8) {
host = 0x00,
isa = 0x01,
eisa = 0x02,
mca = 0x03,
pci_to_pci = 0x04,
pcmcia = 0x05,
nubus = 0x06,
cardbus = 0x07,
raceway = 0x08,
pci_to_pci_semi_transparent = 0x09,
infiniband_to_pci = 0x0A,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.host => "Host Bridge",
.isa => "ISA Bridge",
.eisa => "EISA Bridge",
.mca => "MCA Bridge",
.pci_to_pci => "PCI-to-PCI Bridge",
.pcmcia => "PCMCIA Bridge",
.nubus => "NuBus Bridge",
.cardbus => "CardBus Bridge",
.raceway => "RACEway Bridge",
.pci_to_pci_semi_transparent => "PCI-to-PCI Bridge (Semi-Transparent)",
.infiniband_to_pci => "InfiniBand-to-PCI Host Bridge",
};
}
};
pub const pci_to_pci = struct {
pub const ProgIf = enum(u8) {
normal_decode = 0x00,
subtractive_decode = 0x01,
pub fn name(self: ProgIf) []const u8 {
return switch (self) {
.normal_decode => "Normal Decode",
.subtractive_decode => "Subtractive Decode",
};
}
};
};
};
pub const simple_communication = struct {
pub const SubClass = enum(u8) {
serial = 0x00,
parallel = 0x01,
multiport_serial = 0x02,
modem = 0x03,
gpib = 0x04,
smart_card = 0x05,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.serial => "Serial Controller",
.parallel => "Parallel Controller",
.multiport_serial => "Multiport Serial Controller",
.modem => "Modem",
.gpib => "IEEE 488.1/2 (GPIB) Controller",
.smart_card => "Smart Card Controller",
};
}
};
pub const serial = struct {
pub const ProgIf = enum(u8) {
compatible_8250 = 0x00,
compatible_16450 = 0x01,
compatible_16550 = 0x02,
compatible_16650 = 0x03,
compatible_16750 = 0x04,
compatible_16850 = 0x05,
compatible_16950 = 0x06,
pub fn name(self: ProgIf) []const u8 {
return switch (self) {
.compatible_8250 => "8250-Compatible (Generic XT)",
.compatible_16450 => "16450-Compatible",
.compatible_16550 => "16550-Compatible",
.compatible_16650 => "16650-Compatible",
.compatible_16750 => "16750-Compatible",
.compatible_16850 => "16850-Compatible",
.compatible_16950 => "16950-Compatible",
};
}
};
};
};
pub const base_system_peripheral = struct {
pub const SubClass = enum(u8) {
pic = 0x00,
dma = 0x01,
timer = 0x02,
rtc = 0x03,
pci_hot_plug = 0x04,
sd_host = 0x05,
iommu = 0x06,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.pic => "PIC",
.dma => "DMA Controller",
.timer => "Timer",
.rtc => "RTC Controller",
.pci_hot_plug => "PCI Hot-Plug Controller",
.sd_host => "SD Host Controller",
.iommu => "IOMMU",
};
}
};
};
pub const input_device = struct {
pub const SubClass = enum(u8) {
keyboard = 0x00,
digitizer_pen = 0x01,
mouse = 0x02,
scanner = 0x03,
gameport = 0x04,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.keyboard => "Keyboard Controller",
.digitizer_pen => "Digitizer Pen",
.mouse => "Mouse Controller",
.scanner => "Scanner Controller",
.gameport => "Gameport Controller",
};
}
};
};
pub const serial_bus = struct {
pub const SubClass = enum(u8) {
firewire = 0x00,
access_bus = 0x01,
ssa = 0x02,
usb = 0x03,
fibre_channel = 0x04,
smbus = 0x05,
infiniband = 0x06,
ipmi = 0x07,
sercos = 0x08,
canbus = 0x09,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.firewire => "FireWire (IEEE 1394) Controller",
.access_bus => "ACCESS Bus Controller",
.ssa => "SSA",
.usb => "USB Controller",
.fibre_channel => "Fibre Channel",
.smbus => "SMBus Controller",
.infiniband => "InfiniBand Controller",
.ipmi => "IPMI Interface",
.sercos => "SERCOS Interface (IEC 61491)",
.canbus => "CANbus Controller",
};
}
};
pub const usb = struct {
pub const ProgIf = enum(u8) {
uhci = 0x00,
ohci = 0x10,
ehci = 0x20,
xhci = 0x30,
unspecified = 0x80,
device = 0xFE,
pub fn name(self: ProgIf) []const u8 {
return switch (self) {
.uhci => "UHCI Controller",
.ohci => "OHCI Controller",
.ehci => "EHCI (USB2) Controller",
.xhci => "XHCI (USB3) Controller",
.unspecified => "Unspecified",
.device => "USB Device (not a host controller)",
};
}
};
};
};
pub const wireless = struct {
pub const SubClass = enum(u8) {
irda = 0x00,
consumer_ir = 0x01,
rf = 0x10,
bluetooth = 0x11,
broadband = 0x12,
ethernet_802_1a = 0x20,
ethernet_802_1b = 0x21,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.irda => "iRDA Compatible Controller",
.consumer_ir => "Consumer IR Controller",
.rf => "RF Controller",
.bluetooth => "Bluetooth Controller",
.broadband => "Broadband Controller",
.ethernet_802_1a => "Ethernet Controller (802.1a)",
.ethernet_802_1b => "Ethernet Controller (802.1b)",
};
}
};
};
// --- Raw-byte decoding (what a function reports in its header) -------------------------
/// The name of an exhaustive class-code enum member, or null if `value` is not one — the
/// bridge from a raw config byte to a named taxonomy above.
fn enumName(comptime Enum: type, value: u8) ?[]const u8 {
return (std.enums.fromInt(Enum, value) orelse return null).name();
}
/// Name of the base class (byte 0x0B), e.g. `0x06` -> "Bridge". /// Name of the base class (byte 0x0B), e.g. `0x06` -> "Bridge".
pub fn className(base: u8) []const u8 { pub fn className(base: u8) []const u8 {
return switch (base) { return @as(BaseClass, @enumFromInt(base)).name();
0x00 => "Unclassified",
0x01 => "Mass Storage Controller",
0x02 => "Network Controller",
0x03 => "Display Controller",
0x04 => "Multimedia Controller",
0x05 => "Memory Controller",
0x06 => "Bridge",
0x07 => "Simple Communication Controller",
0x08 => "Base System Peripheral",
0x09 => "Input Device Controller",
0x0A => "Docking Station",
0x0B => "Processor",
0x0C => "Serial Bus Controller",
0x0D => "Wireless Controller",
0x0E => "Intelligent Controller",
0x0F => "Satellite Communication Controller",
0x10 => "Encryption Controller",
0x11 => "Signal Processing Controller",
0x12 => "Processing Accelerator",
0x13 => "Non-Essential Instrumentation",
0x40 => "Co-Processor",
0xFF => "Unassigned Class (Vendor specific)",
else => "Unknown",
};
} }
/// Name of the subclass within its base class, e.g. `(0x06, 0x01)` -> "ISA Bridge". /// Name of the subclass within its base class, e.g. `(0x06, 0x01)` -> "ISA Bridge".
/// Subclass `0x80` is "Other" by PCI convention; anything unlisted is "Unknown". /// Subclass `0x80` is "Other" by PCI convention; anything unlisted is "Unknown".
pub fn subclassName(base: u8, subclass: u8) []const u8 { pub fn subclassName(base: u8, subclass: u8) []const u8 {
return switch (base) { const named: ?[]const u8 = switch (@as(BaseClass, @enumFromInt(base))) {
0x01 => switch (subclass) { .mass_storage => enumName(mass_storage.SubClass, subclass),
0x00 => "SCSI Bus Controller", .network => enumName(network.SubClass, subclass),
0x01 => "IDE Controller", .display => enumName(display.SubClass, subclass),
0x02 => "Floppy Disk Controller", .multimedia => enumName(multimedia.SubClass, subclass),
0x03 => "IPI Bus Controller", .memory => enumName(memory.SubClass, subclass),
0x04 => "RAID Controller", .bridge => enumName(bridge.SubClass, subclass),
0x05 => "ATA Controller", .simple_communication => enumName(simple_communication.SubClass, subclass),
0x06 => "Serial ATA Controller", .base_system_peripheral => enumName(base_system_peripheral.SubClass, subclass),
0x07 => "Serial Attached SCSI Controller", .input_device => enumName(input_device.SubClass, subclass),
0x08 => "Non-Volatile Memory Controller", .serial_bus => enumName(serial_bus.SubClass, subclass),
else => defaultSubclass(subclass), .wireless => enumName(wireless.SubClass, subclass),
}, else => null,
0x02 => switch (subclass) {
0x00 => "Ethernet Controller",
0x01 => "Token Ring Controller",
0x02 => "FDDI Controller",
0x03 => "ATM Controller",
0x04 => "ISDN Controller",
0x06 => "PICMG 2.14 Multi Computing Controller",
0x07 => "Infiniband Controller",
0x08 => "Fabric Controller",
else => defaultSubclass(subclass),
},
0x03 => switch (subclass) {
0x00 => "VGA Compatible Controller",
0x01 => "XGA Controller",
0x02 => "3D Controller (Not VGA-Compatible)",
else => defaultSubclass(subclass),
},
0x04 => switch (subclass) {
0x00 => "Multimedia Video Controller",
0x01 => "Multimedia Audio Controller",
0x02 => "Computer Telephony Device",
0x03 => "Audio Device",
else => defaultSubclass(subclass),
},
0x05 => switch (subclass) {
0x00 => "RAM Controller",
0x01 => "Flash Controller",
else => defaultSubclass(subclass),
},
0x06 => switch (subclass) {
0x00 => "Host Bridge",
0x01 => "ISA Bridge",
0x02 => "EISA Bridge",
0x03 => "MCA Bridge",
0x04 => "PCI-to-PCI Bridge",
0x05 => "PCMCIA Bridge",
0x06 => "NuBus Bridge",
0x07 => "CardBus Bridge",
0x08 => "RACEway Bridge",
0x09 => "PCI-to-PCI Bridge (Semi-Transparent)",
0x0A => "InfiniBand-to-PCI Host Bridge",
else => defaultSubclass(subclass),
},
0x07 => switch (subclass) {
0x00 => "Serial Controller",
0x01 => "Parallel Controller",
0x02 => "Multiport Serial Controller",
0x03 => "Modem",
0x04 => "IEEE 488.1/2 (GPIB) Controller",
0x05 => "Smart Card Controller",
else => defaultSubclass(subclass),
},
0x08 => switch (subclass) {
0x00 => "PIC",
0x01 => "DMA Controller",
0x02 => "Timer",
0x03 => "RTC Controller",
0x04 => "PCI Hot-Plug Controller",
0x05 => "SD Host Controller",
0x06 => "IOMMU",
else => defaultSubclass(subclass),
},
0x09 => switch (subclass) {
0x00 => "Keyboard Controller",
0x01 => "Digitizer Pen",
0x02 => "Mouse Controller",
0x03 => "Scanner Controller",
0x04 => "Gameport Controller",
else => defaultSubclass(subclass),
},
0x0C => switch (subclass) {
0x00 => "FireWire (IEEE 1394) Controller",
0x01 => "ACCESS Bus Controller",
0x02 => "SSA",
0x03 => "USB Controller",
0x04 => "Fibre Channel",
0x05 => "SMBus Controller",
0x06 => "InfiniBand Controller",
0x07 => "IPMI Interface",
0x08 => "SERCOS Interface (IEC 61491)",
0x09 => "CANbus Controller",
else => defaultSubclass(subclass),
},
0x0D => switch (subclass) {
0x00 => "iRDA Compatible Controller",
0x01 => "Consumer IR Controller",
0x10 => "RF Controller",
0x11 => "Bluetooth Controller",
0x12 => "Broadband Controller",
0x20 => "Ethernet Controller (802.1a)",
0x21 => "Ethernet Controller (802.1b)",
else => defaultSubclass(subclass),
},
else => defaultSubclass(subclass),
}; };
return named orelse defaultSubclass(subclass);
} }
fn defaultSubclass(subclass: u8) []const u8 { fn defaultSubclass(subclass: u8) []const u8 {
@@ -175,68 +502,34 @@ fn defaultSubclass(subclass: u8) []const u8 {
/// Returns "" when the prog-IF carries no standard meaning for this class/subclass — /// Returns "" when the prog-IF carries no standard meaning for this class/subclass —
/// callers just print the hex byte in that case. /// callers just print the hex byte in that case.
pub fn progIfName(base: u8, subclass: u8, prog_if: u8) []const u8 { pub fn progIfName(base: u8, subclass: u8, prog_if: u8) []const u8 {
return switch (base) { const named: ?[]const u8 = switch (@as(BaseClass, @enumFromInt(base))) {
0x01 => switch (subclass) { .mass_storage => switch (std.enums.fromInt(mass_storage.SubClass, subclass) orelse return "") {
0x06 => switch (prog_if) { // Serial ATA .serial_ata => enumName(mass_storage.serial_ata.ProgIf, prog_if),
0x00 => "Vendor Specific Interface", .non_volatile_memory => enumName(mass_storage.non_volatile_memory.ProgIf, prog_if),
0x01 => "AHCI 1.0", else => null,
0x02 => "Serial Storage Bus",
else => "",
}, },
0x08 => switch (prog_if) { // Non-Volatile Memory .display => switch (std.enums.fromInt(display.SubClass, subclass) orelse return "") {
0x01 => "NVMHCI", .vga_compatible => enumName(display.vga_compatible.ProgIf, prog_if),
0x02 => "NVM Express", else => null,
else => "",
}, },
else => "", .bridge => switch (std.enums.fromInt(bridge.SubClass, subclass) orelse return "") {
.pci_to_pci => enumName(bridge.pci_to_pci.ProgIf, prog_if),
else => null,
}, },
0x03 => switch (subclass) { .simple_communication => switch (std.enums.fromInt(simple_communication.SubClass, subclass) orelse return "") {
0x00 => switch (prog_if) { // VGA Compatible .serial => enumName(simple_communication.serial.ProgIf, prog_if),
0x00 => "VGA Controller", else => null,
0x01 => "8514-Compatible Controller",
else => "",
}, },
else => "", .serial_bus => switch (std.enums.fromInt(serial_bus.SubClass, subclass) orelse return "") {
.usb => enumName(serial_bus.usb.ProgIf, prog_if),
else => null,
}, },
0x06 => switch (subclass) { else => null,
0x04 => switch (prog_if) { // PCI-to-PCI Bridge
0x00 => "Normal Decode",
0x01 => "Subtractive Decode",
else => "",
},
else => "",
},
0x07 => switch (subclass) {
0x00 => switch (prog_if) { // Serial Controller
0x00 => "8250-Compatible (Generic XT)",
0x01 => "16450-Compatible",
0x02 => "16550-Compatible",
0x03 => "16650-Compatible",
0x04 => "16750-Compatible",
0x05 => "16850-Compatible",
0x06 => "16950-Compatible",
else => "",
},
else => "",
},
0x0C => switch (subclass) {
0x03 => switch (prog_if) { // USB Controller
0x00 => "UHCI Controller",
0x10 => "OHCI Controller",
0x20 => "EHCI (USB2) Controller",
0x30 => "XHCI (USB3) Controller",
0x80 => "Unspecified",
0xFE => "USB Device (not a host controller)",
else => "",
},
else => "",
},
else => "",
}; };
return named orelse "";
} }
test "decodes the common class codes" { test "decodes the common class codes" {
const std = @import("std");
const eq = std.testing.expectEqualStrings; const eq = std.testing.expectEqualStrings;
const isa = ClassCode.unpack(0x06_01_00); const isa = ClassCode.unpack(0x06_01_00);
@@ -251,11 +544,25 @@ test "decodes the common class codes" {
try eq("AHCI 1.0", progIfName(ahci.base, ahci.subclass, ahci.prog_if)); try eq("AHCI 1.0", progIfName(ahci.base, ahci.subclass, ahci.prog_if));
const xhci = ClassCode.unpack(0x0C_03_30); const xhci = ClassCode.unpack(0x0C_03_30);
try eq("Serial Bus Controller", className(xhci.base));
try eq("USB Controller", subclassName(xhci.base, xhci.subclass)); try eq("USB Controller", subclassName(xhci.base, xhci.subclass));
try eq("XHCI (USB3) Controller", progIfName(xhci.base, xhci.subclass, xhci.prog_if)); try eq("XHCI (USB3) Controller", progIfName(xhci.base, xhci.subclass, xhci.prog_if));
}
// Unknowns and the "Other" convention. test "unlisted codes fall back without a wrong name" {
try eq("Other", subclassName(0x02, 0x80)); const eq = std.testing.expectEqualStrings;
try eq("Unknown", subclassName(0x06, 0x7E)); try eq("Unknown", className(0x77)); // no such base class
try eq("", progIfName(0x06, 0x00, 0x00)); // host bridge: prog-IF has no standard name try eq("Other", subclassName(0x01, 0x80)); // 0x80 is the PCI "Other" convention
try eq("Unknown", subclassName(0x01, 0x7A)); // unlisted mass-storage subclass
try eq("", progIfName(0x01, 0x06, 0x7F)); // no standard SATA prog-IF for 0x7F
try eq("", progIfName(0x02, 0x00, 0x00)); // class with no prog-IF taxonomy at all
}
test "named parts pack to the raw triple" {
const xhci = ClassCode{
.base = @intFromEnum(BaseClass.serial_bus),
.subclass = @intFromEnum(serial_bus.SubClass.usb),
.prog_if = @intFromEnum(serial_bus.usb.ProgIf.xhci),
};
try std.testing.expectEqual(@as(u24, 0x0C_03_30), xhci.pack());
} }
+9 -1
View File
@@ -40,6 +40,13 @@ pub fn platformInformation() PlatformInformation {
} }
/// AML parse integrity/diagnostics (namespace node count, bytes consumed). /// AML parse integrity/diagnostics (namespace node count, bytes consumed).
/// The number of Device objects in the kernel's own AML namespace, or 0 if the
/// parse produced none — the `acpi-parse` test compares the ring-3 service's
/// count against this.
pub fn amlDeviceCount() usize {
return acpi.amlDeviceCount();
}
pub fn amlStats() AmlStats { pub fn amlStats() AmlStats {
return acpi.aml_stats; return acpi.aml_stats;
} }
@@ -72,7 +79,8 @@ pub fn discover(
var device_tree = try DeviceTree.init(allocator); var device_tree = try DeviceTree.init(allocator);
if (boot_information.acpi_rsdp != 0) { if (boot_information.acpi_rsdp != 0) {
try acpi.discover(boot_information.acpi_rsdp, &device_tree, hal); const memory_regions = @as([*]const boot_handoff.MemoryRegion, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.memory_map.regions)))[0..boot_information.memory_map.len];
try acpi.discover(boot_information.acpi_rsdp, memory_regions, &device_tree, hal);
} else { } else {
// No ACPI RSDP. A device-tree boot would parse its blob here; today that // No ACPI RSDP. A device-tree boot would parse its blob here; today that
// path is a stub, so this reports the machine described itself no way we // path is a stub, so this reports the machine described itself no way we
+11 -11
View File
@@ -93,21 +93,21 @@ fn findHpet(buffer: []device.DeviceDescriptor) ?Found {
pub fn main() void { pub fn main() void {
// Enumerate into a heap buffer (too big for the one-page user stack). // Enumerate into a heap buffer (too big for the one-page user stack).
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 32) catch { const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 32) catch {
_ = runtime.system.write("hpet: out of memory\n"); _ = runtime.system.write("system/drivers/hpet: out of memory\n");
return; return;
}; };
const hpet = findHpet(buffer) orelse { const hpet = findHpet(buffer) orelse {
_ = runtime.system.write("hpet: no HPET with an IRQ\n"); _ = runtime.system.write("system/drivers/hpet: no HPET with an IRQ\n");
return; return;
}; };
if (!device.claim(hpet.device_id)) { if (!device.claim(hpet.device_id)) {
_ = runtime.system.write("hpet: claim failed\n"); _ = runtime.system.write("system/drivers/hpet: claim failed\n");
return; return;
} }
const base = device.mmioMap(hpet.device_id, hpet.mmio) orelse { const base = device.mmioMap(hpet.device_id, hpet.mmio) orelse {
_ = runtime.system.write("hpet: mmio_map failed\n"); _ = runtime.system.write("system/drivers/hpet: mmio_map failed\n");
return; return;
}; };
@@ -116,7 +116,7 @@ pub fn main() void {
const gsi = hpet.gsi; const gsi = hpet.gsi;
const endpoint = ipc.createIpcEndpoint() orelse { const endpoint = ipc.createIpcEndpoint() orelse {
_ = runtime.system.write("hpet: create_ipc_endpoint failed\n"); _ = runtime.system.write("system/drivers/hpet: create_ipc_endpoint failed\n");
return; return;
}; };
@@ -124,7 +124,7 @@ pub fn main() void {
// Counter period, so we can arm the comparator a fixed wall-clock distance out. // Counter period, so we can arm the comparator a fixed wall-clock distance out.
const femtos_per_tick = rd(base, register_general_cap) >> 32; const femtos_per_tick = rd(base, register_general_cap) >> 32;
if (femtos_per_tick == 0) { if (femtos_per_tick == 0) {
_ = runtime.system.write("hpet: bad HPET period\n"); _ = runtime.system.write("system/drivers/hpet: bad HPET period\n");
return; return;
} }
const ticks_per_ms = 1_000_000_000_000 / femtos_per_tick; const ticks_per_ms = 1_000_000_000_000 / femtos_per_tick;
@@ -147,10 +147,10 @@ pub fn main() void {
wr(base, register_general_configuration, rd(base, register_general_configuration) | configuration_enable); wr(base, register_general_configuration, rd(base, register_general_configuration) | configuration_enable);
if (!device.irqBind(hpet.device_id, hpet.irq, endpoint)) { if (!device.irqBind(hpet.device_id, hpet.irq, endpoint)) {
_ = runtime.system.write("hpet: irq_bind failed\n"); _ = runtime.system.write("system/drivers/hpet: irq_bind failed\n");
return; return;
} }
_ = runtime.system.write("hpet: bound, sleeping until the hardware speaks\n"); _ = runtime.system.write("system/drivers/hpet: bound, sleeping until the hardware speaks\n");
// --- the driver loop ----------------------------------------------------- // --- the driver loop -----------------------------------------------------
// Blocked in replyWait. No polling, no spinning: the next line of this function // Blocked in replyWait. No polling, no spinning: the next line of this function
@@ -178,14 +178,14 @@ pub fn main() void {
wr(base, register_timer0_configuration, rd(base, register_timer0_configuration) & ~tn_int_enb); wr(base, register_timer0_configuration, rd(base, register_timer0_configuration) & ~tn_int_enb);
} }
_ = runtime.system.write("hpet: irq\n"); _ = runtime.system.write("system/drivers/hpet: irq\n");
if (!device.irqAck(hpet.device_id, hpet.irq)) { if (!device.irqAck(hpet.device_id, hpet.irq)) {
_ = runtime.system.write("hpet: irq_ack failed\n"); _ = runtime.system.write("system/drivers/hpet: irq_ack failed\n");
return; return;
} }
} }
_ = runtime.system.write("hpet: ok\n"); _ = runtime.system.write("system/drivers/hpet: ok\n");
while (true) runtime.system.sleep(1000); while (true) runtime.system.sleep(1000);
} }
+270
View File
@@ -0,0 +1,270 @@
//! /system/drivers/pci-bus — the PCI bus driver: enumeration moved out of ring 0
//! (docs/m19-m20-plan.md, M19). The device manager matches the `pci_host_bridge`
//! node and spawns one instance per bridge, the bridge's device id as argv[1] —
//! the same per-device contract as usb-xhci-bus.
//!
//! M19.1 (this increment): claim the bridge, map its ECAM window (resource 0;
//! the bus range and the MMIO apertures follow it), walk every
//! bus/device/function config header, and log what the walk finds — ending
//! with "/system/drivers/pci-bus: N functions found", which the `pci-scan` scenario compares
//! against the kernel's own enumeration. Registration and reports (M19.2), and
//! the kernel walk's retirement (M19.3), build on this proven-equivalent scan.
const std = @import("std");
const runtime = @import("runtime");
const protocol = runtime.device_manager_protocol;
const device = runtime.device;
const pci_class = @import("pci-class");
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
/// Log a discovered function with its (class / subclass / prog-IF) triple decoded
/// to human names — the boot-log breadcrumb that says *what* the hardware is, so
/// "class 0x01 (Mass Storage Controller) subclass 0x06 (Serial ATA Controller)
/// progif 0x01 (AHCI 1.0)" reads straight off the log when writing a new driver.
/// A dedicated wider buffer than `writeLine`'s, since the decoded names are long.
fn logFunction(bus: u64, dev: u64, function: u64, class_triple: u32) void {
const cc = pci_class.ClassCode.unpack(@truncate(class_triple));
const pif = pci_class.progIfName(cc.base, cc.subclass, cc.prog_if);
var line: [200]u8 = undefined;
const text = if (pif.len != 0)
std.fmt.bufPrint(&line, "/system/drivers/pci-bus: {d}:{d}.{d} class 0x{x:0>2} ({s}) subclass 0x{x:0>2} ({s}) progif 0x{x:0>2} ({s})\n", .{ bus, dev, function, cc.base, pci_class.className(cc.base), cc.subclass, pci_class.subclassName(cc.base, cc.subclass), cc.prog_if, pif }) catch return
else
std.fmt.bufPrint(&line, "/system/drivers/pci-bus: {d}:{d}.{d} class 0x{x:0>2} ({s}) subclass 0x{x:0>2} ({s}) progif 0x{x:0>2}\n", .{ bus, dev, function, cc.base, pci_class.className(cc.base), cc.subclass, pci_class.subclassName(cc.base, cc.subclass), cc.prog_if }) catch return;
_ = runtime.system.write(text);
}
var bridge_id: u64 = protocol.no_device;
var ecam_base: usize = 0;
var ecam_physical: u64 = 0;
var start_bus: u64 = 0;
var bus_count: u64 = 0;
var manager_handle: runtime.ipc.Handle = 0;
/// One aligned 32-bit read from a function's configuration space.
fn configRead(bus: u64, dev: u64, function: u64, offset: u64) u32 {
const address = ecam_base + (((bus - start_bus) << 20) | (dev << 15) | (function << 12) | offset);
const register: *volatile u32 = @ptrFromInt(address);
return register.*;
}
fn configWrite(bus: u64, dev: u64, function: u64, offset: u64, value: u32) void {
const address = ecam_base + (((bus - start_bus) << 20) | (dev << 15) | (function << 12) | offset);
const register: *volatile u32 = @ptrFromInt(address);
register.* = value;
}
fn configRead16(bus: u64, dev: u64, function: u64, offset: u64) u16 {
const word = configRead(bus, dev, function, offset & ~@as(u64, 3));
return @truncate(word >> @intCast((offset & 3) * 8));
}
fn configWrite16(bus: u64, dev: u64, function: u64, offset: u64, value: u16) void {
const aligned = offset & ~@as(u64, 3);
const shift: u5 = @intCast((offset & 3) * 8);
const word = configRead(bus, dev, function, aligned);
const mask = @as(u32, 0xFFFF) << shift;
configWrite(bus, dev, function, aligned, (word & ~mask) | (@as(u32, value) << shift));
}
/// Claim the bridge, map the ECAM, hello the manager, then scan.
fn initialise(endpoint: runtime.ipc.Handle) bool {
_ = endpoint;
if (!device.claim(bridge_id)) {
writeLine("/system/drivers/pci-bus: unable to claim bridge device {d}\n", .{bridge_id});
return false;
}
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("/system/drivers/pci-bus: out of memory\n");
return false;
};
const total = device.enumerate(buffer);
const descriptor = for (buffer[0..@min(total, buffer.len)]) |d| {
if (d.id == bridge_id) break d;
} else {
writeLine("/system/drivers/pci-bus: device {d} not in the device tree\n", .{bridge_id});
return false;
};
// Resource 0 is the ECAM window (1 MiB of config space per bus); the bus
// range rides beside it. The MMIO apertures (M19.0) come after both.
if (descriptor.resource_count < 2 or descriptor.resources[0].kind != @intFromEnum(device.ResourceKind.memory)) {
_ = runtime.system.write("/system/drivers/pci-bus: bridge has no ECAM window\n");
return false;
}
const bus_range = for (descriptor.resources[0..@intCast(descriptor.resource_count)]) |resource| {
if (resource.kind == @intFromEnum(device.ResourceKind.bus_range)) break resource;
} else {
_ = runtime.system.write("/system/drivers/pci-bus: bridge has no bus range\n");
return false;
};
start_bus = bus_range.start;
bus_count = bus_range.len;
ecam_physical = descriptor.resources[0].start;
ecam_base = device.mmioMap(bridge_id, 0) orelse {
_ = runtime.system.write("/system/drivers/pci-bus: ECAM mmio_map failed\n");
return false;
};
// The handshake, then the scan (reports join in M19.2).
var manager: ?runtime.ipc.Handle = null;
var tries: u32 = 0;
while (manager == null and tries < 100) : (tries += 1) {
manager = runtime.ipc.lookup(.device_manager);
if (manager == null) runtime.system.sleep(20);
}
const h = manager orelse {
_ = runtime.system.write("/system/drivers/pci-bus: no device manager to hello\n");
return false;
};
const hello = protocol.Hello{ .role = @intFromEnum(protocol.Role.bus), .device_id = bridge_id };
var reply: [protocol.message_maximum]u8 = undefined;
const n = runtime.ipc.call(h, std.mem.asBytes(&hello), &reply) catch {
_ = runtime.system.write("/system/drivers/pci-bus: hello call failed\n");
return false;
};
if (n < protocol.reply_size or std.mem.bytesToValue(protocol.HelloReply, reply[0..protocol.reply_size]).status != 0) {
_ = runtime.system.write("/system/drivers/pci-bus: hello refused\n");
return false;
}
manager_handle = h;
scan();
return true;
}
/// The brute-force walk the kernel does today, from ring 3: every bus in the
/// range, 32 devices, 8 functions; vendor id FFFFh means nothing decodes there,
/// and only multifunction devices get their functions 1..7 probed.
fn scan() void {
var found: u32 = 0;
var bus: u64 = start_bus;
while (bus < start_bus + bus_count) : (bus += 1) {
var dev: u64 = 0;
while (dev < 32) : (dev += 1) {
const first = configRead(bus, dev, 0, 0);
if (first & 0xFFFF == 0xFFFF) continue;
const multifunction = (configRead(bus, dev, 0, 0x0C) >> 16) & 0x80 != 0;
var function: u64 = 0;
while (function < 8) : (function += 1) {
if (function != 0 and !multifunction) break;
const vendor_device = configRead(bus, dev, function, 0);
if (vendor_device & 0xFFFF == 0xFFFF) continue;
const class_revision = configRead(bus, dev, function, 0x08);
found += 1;
logFunction(bus, dev, function, class_revision >> 8);
registerAndReport(bus, dev, function, class_revision >> 8);
}
}
}
writeLine("/system/drivers/pci-bus: {d} functions found\n", .{found});
}
/// Register one function under the bridge and report it to the manager. The
/// descriptor mirrors the kernel's own recording byte for byte — config slice
/// as resource 0, then the sized BARs — so during coexistence the idempotent
/// device_register (M19.0) returns the kernel's existing node id rather than
/// growing a duplicate, and the report carries the id drivers already use.
fn registerAndReport(bus: u64, dev: u64, function: u64, class_triple: u32) void {
var descriptor = std.mem.zeroes(device.DeviceDescriptor);
descriptor.class = @intFromEnum(device.DeviceClass.pci_device);
descriptor.pci_class = class_triple;
descriptor.resources[0] = .{
.kind = @intFromEnum(device.ResourceKind.memory),
.start = ecam_physical + (((bus - start_bus) << 20) | (dev << 15) | (function << 12)),
.len = 4096,
};
descriptor.resource_count = 1;
// The standard BAR-sizing probe, exactly as the kernel does it: decode off,
// write all-ones, read the writable mask back, restore. Header type 0 only.
const header_type = (configRead(bus, dev, function, 0x0C) >> 16) & 0x7F;
if (header_type == 0) {
const command = configRead16(bus, dev, function, 0x04);
configWrite16(bus, dev, function, 0x04, command & ~@as(u16, 0b11));
var i: u64 = 0;
while (i < 6) : (i += 1) {
if (descriptor.resource_count >= 8) break;
const off = 0x10 + i * 4;
const original = configRead(bus, dev, function, off);
if (original == 0) continue;
const slot: usize = @intCast(descriptor.resource_count);
if (original & 1 != 0) {
configWrite(bus, dev, function, off, 0xFFFF_FFFF);
const readback = configRead(bus, dev, function, off);
configWrite(bus, dev, function, off, original);
const mask = readback & 0xFFFF_FFFC;
const size: u32 = if (mask == 0) 0 else (~mask +% 1) & 0xFFFF;
if (size == 0) continue; // unimplemented BAR — nothing to register
descriptor.resources[slot] = .{ .kind = @intFromEnum(device.ResourceKind.io_port), .start = original & 0xFFFF_FFFC, .len = size };
descriptor.resource_count += 1;
} else if ((original >> 1) & 0x3 == 2) {
const original_high = configRead(bus, dev, function, off + 4);
configWrite(bus, dev, function, off, 0xFFFF_FFFF);
configWrite(bus, dev, function, off + 4, 0xFFFF_FFFF);
const lo = configRead(bus, dev, function, off);
const hi = configRead(bus, dev, function, off + 4);
configWrite(bus, dev, function, off, original);
configWrite(bus, dev, function, off + 4, original_high);
const readback = (@as(u64, hi) << 32) | (lo & 0xFFFF_FFF0);
const size: u64 = if (readback == 0) 0 else ~readback +% 1;
i += 1; // consumed the high half regardless
if (size == 0) continue;
descriptor.resources[slot] = .{ .kind = @intFromEnum(device.ResourceKind.memory), .start = (@as(u64, original_high) << 32) | (original & 0xFFFF_FFF0), .len = size };
descriptor.resource_count += 1;
} else {
configWrite(bus, dev, function, off, 0xFFFF_FFFF);
const readback = configRead(bus, dev, function, off);
configWrite(bus, dev, function, off, original);
const mask = readback & 0xFFFF_FFF0;
const size: u32 = if (mask == 0) 0 else ~mask +% 1;
if (size == 0) continue;
descriptor.resources[slot] = .{ .kind = @intFromEnum(device.ResourceKind.memory), .start = original & 0xFFFF_FFF0, .len = size };
descriptor.resource_count += 1;
}
}
configWrite16(bus, dev, function, 0x04, command);
}
const registered = device.register(bridge_id, &descriptor) orelse {
writeLine("/system/drivers/pci-bus: register refused for {d}:{d}.{d}\n", .{ bus, dev, function });
return;
};
const report = protocol.ChildAdded{
.parent = bridge_id,
.bus_address = (bus << 8) | (dev << 3) | function,
.identity = class_triple,
.device_id = registered,
};
var reply: [protocol.message_maximum]u8 = undefined;
_ = runtime.ipc.call(manager_handle, std.mem.asBytes(&report), &reply) catch {
writeLine("/system/drivers/pci-bus: child report for {d}:{d}.{d} failed\n", .{ bus, dev, function });
};
}
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = message;
_ = reply;
_ = sender;
_ = capability;
return 0;
}
pub fn main(init: runtime.process.Init) void {
const argument = init.arguments.get(1) orelse return; // bare (ramdisk sweep): stay silent
bridge_id = std.fmt.parseInt(u64, argument, 10) catch {
writeLine("/system/drivers/pci-bus: malformed bridge device id '{s}'\n", .{argument});
return;
};
runtime.service.run(protocol.message_maximum, .{
.init = initialise,
.on_message = onMessage,
});
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
+11 -11
View File
@@ -72,17 +72,17 @@ fn modifierWord(modifiers: scancode.ModifierSnapshot) u32 {
pub fn main(init: runtime.process.Init) void { pub fn main(init: runtime.process.Init) void {
const hid = init.arguments.get(1).?; const hid = init.arguments.get(1).?;
if (hid.len == 0) { if (hid.len == 0) {
_ = runtime.system.write("system/drivers/ps2-bus/keyboard: no HID argument\n"); _ = runtime.system.write("/system/drivers/ps2-bus/keyboard: no HID argument\n");
return; return;
} }
writeLine("system/drivers/ps2-bus/keyboard: starting for hid {s}\n", .{hid}); writeLine("/system/drivers/ps2-bus/keyboard: starting for hid {s}\n", .{hid});
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch { const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("system/drivers/ps2-bus/keyboard: out of memory\n"); _ = runtime.system.write("/system/drivers/ps2-bus/keyboard: out of memory\n");
return; return;
}; };
if (device.findDeviceDescriptorByHid(buffer, hid) == null) { if (device.findDeviceDescriptorByHid(buffer, hid) == null) {
writeLine("system/drivers/ps2-bus/keyboard: no device for hid {s}\n", .{hid}); writeLine("/system/drivers/ps2-bus/keyboard: no device for hid {s}\n", .{hid});
return; return;
} }
@@ -90,38 +90,38 @@ pub fn main(init: runtime.process.Init) void {
// absent (as today) it defaults to us. // absent (as today) it defaults to us.
const layout_name = init.arguments.get(2) orelse "us"; const layout_name = init.arguments.get(2) orelse "us";
const layout = xkb.byName(layout_name) orelse xkb.us; const layout = xkb.byName(layout_name) orelse xkb.us;
writeLine("system/drivers/ps2-bus/keyboard: layout {s}\n", .{layout.name}); writeLine("/system/drivers/ps2-bus/keyboard: layout {s}\n", .{layout.name});
// Attach to the bus: hand it our endpoint, and it forwards every byte the // Attach to the bus: hand it our endpoint, and it forwards every byte the
// keyboard sends (it owns the controller; we own the decoding). // keyboard sends (it owns the controller; we own the decoding).
const bus = lookupBus() orelse { const bus = lookupBus() orelse {
_ = runtime.system.write("system/drivers/ps2-bus/keyboard: ps2-bus service unavailable\n"); _ = runtime.system.write("/system/drivers/ps2-bus/keyboard: ps2-bus service unavailable\n");
return; return;
}; };
const endpoint = ipc.createIpcEndpoint() orelse { const endpoint = ipc.createIpcEndpoint() orelse {
_ = runtime.system.write("system/drivers/ps2-bus/keyboard: no endpoint\n"); _ = runtime.system.write("/system/drivers/ps2-bus/keyboard: no endpoint\n");
return; return;
}; };
var attach = ps2.AttachRequest{ .device_type = @intFromEnum(ps2.DeviceType.keyboard) }; var attach = ps2.AttachRequest{ .device_type = @intFromEnum(ps2.DeviceType.keyboard) };
var attach_reply: [@sizeOf(ps2.AttachReply)]u8 = undefined; var attach_reply: [@sizeOf(ps2.AttachReply)]u8 = undefined;
const attached = ipc.callCap(bus, std.mem.asBytes(&attach), &attach_reply, endpoint) catch { const attached = ipc.callCap(bus, std.mem.asBytes(&attach), &attach_reply, endpoint) catch {
_ = runtime.system.write("system/drivers/ps2-bus/keyboard: attach call failed\n"); _ = runtime.system.write("/system/drivers/ps2-bus/keyboard: attach call failed\n");
return; return;
}; };
if (attached.len < @sizeOf(ps2.AttachReply) or if (attached.len < @sizeOf(ps2.AttachReply) or
std.mem.bytesToValue(ps2.AttachReply, attach_reply[0..@sizeOf(ps2.AttachReply)]).status != @intFromEnum(ps2.AttachStatus.ok)) std.mem.bytesToValue(ps2.AttachReply, attach_reply[0..@sizeOf(ps2.AttachReply)]).status != @intFromEnum(ps2.AttachStatus.ok))
{ {
_ = runtime.system.write("system/drivers/ps2-bus/keyboard: attach refused\n"); _ = runtime.system.write("/system/drivers/ps2-bus/keyboard: attach refused\n");
return; return;
} }
// Broadcast keyboard events through the input service so programs can listen // Broadcast keyboard events through the input service so programs can listen
// for them (docs/input.md). // for them (docs/input.md).
var source = runtime.input.connectSource() orelse { var source = runtime.input.connectSource() orelse {
_ = runtime.system.write("system/drivers/ps2-bus/keyboard: input service unavailable\n"); _ = runtime.system.write("/system/drivers/ps2-bus/keyboard: input service unavailable\n");
return; return;
}; };
_ = runtime.system.write("system/drivers/ps2-bus/keyboard: ok\n"); _ = runtime.system.write("/system/drivers/ps2-bus/keyboard: ok\n");
var decoder = scancode.Decoder{}; var decoder = scancode.Decoder{};
var state = scancode.KeyboardState{}; var state = scancode.KeyboardState{};
+10 -10
View File
@@ -51,50 +51,50 @@ pub fn main(init: runtime.process.Init) void {
const hid = init.arguments.get(1).?; const hid = init.arguments.get(1).?;
if (hid.len == 0) { if (hid.len == 0) {
_ = runtime.system.write("system/drivers/ps2-bus/mouse: no HID argument\n"); _ = runtime.system.write("/system/drivers/ps2-bus/mouse: no HID argument\n");
return; return;
} }
writeLine("system/drivers/ps2-bus/mouse: starting for hid {s}\n", .{hid}); writeLine("/system/drivers/ps2-bus/mouse: starting for hid {s}\n", .{hid});
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch { const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("system/drivers/ps2-bus/mouse: out of memory\n"); _ = runtime.system.write("/system/drivers/ps2-bus/mouse: out of memory\n");
return; return;
}; };
if (device.findDeviceDescriptorByHid(buffer, hid) == null) { if (device.findDeviceDescriptorByHid(buffer, hid) == null) {
writeLine("system/drivers/ps2-bus/mouse: no device for hid {s}\n", .{hid}); writeLine("/system/drivers/ps2-bus/mouse: no device for hid {s}\n", .{hid});
return; return;
} }
// Attach to the bus: hand it our endpoint, and it forwards every byte the // Attach to the bus: hand it our endpoint, and it forwards every byte the
// mouse sends (it owns the controller; we own the decoding). // mouse sends (it owns the controller; we own the decoding).
const bus = lookupBus() orelse { const bus = lookupBus() orelse {
_ = runtime.system.write("system/drivers/ps2-bus/mouse: ps2-bus service unavailable\n"); _ = runtime.system.write("/system/drivers/ps2-bus/mouse: ps2-bus service unavailable\n");
return; return;
}; };
const endpoint = ipc.createIpcEndpoint() orelse { const endpoint = ipc.createIpcEndpoint() orelse {
_ = runtime.system.write("system/drivers/ps2-bus/mouse: no endpoint\n"); _ = runtime.system.write("/system/drivers/ps2-bus/mouse: no endpoint\n");
return; return;
}; };
var attach = ps2.AttachRequest{ .device_type = @intFromEnum(ps2.DeviceType.mouse) }; var attach = ps2.AttachRequest{ .device_type = @intFromEnum(ps2.DeviceType.mouse) };
var attach_reply: [@sizeOf(ps2.AttachReply)]u8 = undefined; var attach_reply: [@sizeOf(ps2.AttachReply)]u8 = undefined;
const attached = ipc.callCap(bus, std.mem.asBytes(&attach), &attach_reply, endpoint) catch { const attached = ipc.callCap(bus, std.mem.asBytes(&attach), &attach_reply, endpoint) catch {
_ = runtime.system.write("system/drivers/ps2-bus/mouse: attach call failed\n"); _ = runtime.system.write("/system/drivers/ps2-bus/mouse: attach call failed\n");
return; return;
}; };
if (attached.len < @sizeOf(ps2.AttachReply) or if (attached.len < @sizeOf(ps2.AttachReply) or
std.mem.bytesToValue(ps2.AttachReply, attach_reply[0..@sizeOf(ps2.AttachReply)]).status != @intFromEnum(ps2.AttachStatus.ok)) std.mem.bytesToValue(ps2.AttachReply, attach_reply[0..@sizeOf(ps2.AttachReply)]).status != @intFromEnum(ps2.AttachStatus.ok))
{ {
_ = runtime.system.write("system/drivers/ps2-bus/mouse: attach refused\n"); _ = runtime.system.write("/system/drivers/ps2-bus/mouse: attach refused\n");
return; return;
} }
// Broadcast mouse events through the input service so programs can listen // Broadcast mouse events through the input service so programs can listen
// for them (docs/input.md). // for them (docs/input.md).
var source = runtime.input.connectSource() orelse { var source = runtime.input.connectSource() orelse {
_ = runtime.system.write("system/drivers/ps2-bus/mouse: input service unavailable\n"); _ = runtime.system.write("/system/drivers/ps2-bus/mouse: input service unavailable\n");
return; return;
}; };
_ = runtime.system.write("system/drivers/ps2-bus/mouse: ok\n"); _ = runtime.system.write("/system/drivers/ps2-bus/mouse: ok\n");
var assembler = mouse_packet.Assembler{}; var assembler = mouse_packet.Assembler{};
var buttons: u32 = 0; var buttons: u32 = 0;
+32 -32
View File
@@ -30,19 +30,19 @@ fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
/// attaches, or null if nothing was spawned. /// attaches, or null if nothing was spawned.
fn spawnIdentifiedDriver(controller: ps2.Controller, port: ps2.Port) ?ps2.DeviceType { fn spawnIdentifiedDriver(controller: ps2.Controller, port: ps2.Port) ?ps2.DeviceType {
const device_type = controller.identifyDevice(port) orelse { const device_type = controller.identifyDevice(port) orelse {
writeLine("system/drivers/ps2-bus: identify timed out on port {s}\n", .{@tagName(port)}); writeLine("/system/drivers/ps2-bus: identify timed out on port {s}\n", .{@tagName(port)});
return null; return null;
}; };
const driver_name = device_type.driverName() orelse { const driver_name = device_type.driverName() orelse {
writeLine("system/drivers/ps2-bus: unrecognized device on port {s}\n", .{@tagName(port)}); writeLine("/system/drivers/ps2-bus: unrecognized device on port {s}\n", .{@tagName(port)});
return null; return null;
}; };
const hid = device_type.hid() orelse ""; const hid = device_type.hid() orelse "";
if (runtime.system.spawnWithArguments(driver_name, &.{hid}) != null) { if (runtime.system.spawnWithArguments(driver_name, &.{hid}) != null) {
writeLine("system/drivers/ps2-bus: port {s} is a {s}, spawned {s}\n", .{ @tagName(port), hid, driver_name }); writeLine("/system/drivers/ps2-bus: port {s} is a {s}, spawned {s}\n", .{ @tagName(port), hid, driver_name });
return device_type; return device_type;
} }
writeLine("system/drivers/ps2-bus: failed to spawn {s}\n", .{driver_name}); writeLine("/system/drivers/ps2-bus: failed to spawn {s}\n", .{driver_name});
return null; return null;
} }
@@ -83,7 +83,7 @@ fn handleAttach(message: []const u8, got: ipc.Received, out: []u8) usize {
const device_type = maybe_type orelse continue; const device_type = maybe_type orelse continue;
if (@intFromEnum(device_type) != request.device_type) continue; if (@intFromEnum(device_type) != request.device_type) continue;
port_endpoints[port_index] = endpoint; port_endpoints[port_index] = endpoint;
writeLine("system/drivers/ps2-bus: {s} driver attached\n", .{@tagName(device_type)}); writeLine("/system/drivers/ps2-bus: {s} driver attached\n", .{@tagName(device_type)});
return reply.write(out, .ok); return reply.write(out, .ok);
} }
return reply.write(out, .no_such_device); return reply.write(out, .no_such_device);
@@ -91,7 +91,7 @@ fn handleAttach(message: []const u8, got: ipc.Received, out: []u8) usize {
pub fn main() void { pub fn main() void {
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch { const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("system/drivers/ps2-bus: out of memory\n"); _ = runtime.system.write("/system/drivers/ps2-bus: out of memory\n");
return; return;
}; };
@@ -103,16 +103,16 @@ pub fn main() void {
// is on which port is decided later by identify, not by this HID. // is on which port is decided later by identify, not by this HID.
const maybe_controller_device_descriptor = device.findDeviceDescriptorByHid(buffer, acpi_ids.HardwareId.ps2_keyboard.hid()); const maybe_controller_device_descriptor = device.findDeviceDescriptorByHid(buffer, acpi_ids.HardwareId.ps2_keyboard.hid());
if (maybe_controller_device_descriptor) |controller_device_descriptor| { if (maybe_controller_device_descriptor) |controller_device_descriptor| {
_ = runtime.system.write("system/drivers/ps2-bus: found PS/2 controller\n"); _ = runtime.system.write("/system/drivers/ps2-bus: found PS/2 controller\n");
_ = runtime.system.write("system/drivers/ps2-bus: initializing controller\n"); _ = runtime.system.write("/system/drivers/ps2-bus: initializing controller\n");
if (!device.claim(controller_device_descriptor.id)) { if (!device.claim(controller_device_descriptor.id)) {
_ = runtime.system.write("system/drivers/ps2-bus: unable to claim controller \n"); _ = runtime.system.write("/system/drivers/ps2-bus: unable to claim controller \n");
return; return;
} }
const controller = ps2.Controller.init(controller_device_descriptor) orelse { const controller = ps2.Controller.init(controller_device_descriptor) orelse {
_ = runtime.system.write("system/drivers/ps2-bus: controller is missing its IO ports\n"); _ = runtime.system.write("/system/drivers/ps2-bus: controller is missing its IO ports\n");
return; return;
}; };
maybe_controller = controller; maybe_controller = controller;
@@ -123,7 +123,7 @@ pub fn main() void {
controller.flushOutputBuffer(); controller.flushOutputBuffer();
const current = controller.readConfigurationByte() orelse { const current = controller.readConfigurationByte() orelse {
_ = runtime.system.write("system/drivers/ps2-bus: controller configuration timed out\n"); _ = runtime.system.write("/system/drivers/ps2-bus: controller configuration timed out\n");
return; return;
}; };
@@ -132,49 +132,49 @@ pub fn main() void {
ps2.configuration_first_port_translation); ps2.configuration_first_port_translation);
if (controller.writeConfigurationByte(update) == null) { if (controller.writeConfigurationByte(update) == null) {
_ = runtime.system.write("system/drivers/ps2-bus: controller configuration timed out\n"); _ = runtime.system.write("/system/drivers/ps2-bus: controller configuration timed out\n");
return; return;
} }
if (controller.performSelfTest()) |reply| { if (controller.performSelfTest()) |reply| {
if (reply != ps2.response_controller_test_passed) { if (reply != ps2.response_controller_test_passed) {
_ = runtime.system.write("system/drivers/ps2-bus: perform controller self test failed\n"); _ = runtime.system.write("/system/drivers/ps2-bus: perform controller self test failed\n");
return; return;
} }
} else { } else {
_ = runtime.system.write("system/drivers/ps2-bus: controller self test timed out\n"); _ = runtime.system.write("/system/drivers/ps2-bus: controller self test timed out\n");
return; return;
} }
has_two_channels = controller.hasTwoChannels() orelse { has_two_channels = controller.hasTwoChannels() orelse {
_ = runtime.system.write("system/drivers/ps2-bus: controller channels timed out\n"); _ = runtime.system.write("/system/drivers/ps2-bus: controller channels timed out\n");
return; return;
}; };
if (has_two_channels) { if (has_two_channels) {
_ = runtime.system.write("system/drivers/ps2-bus: has two channels\n"); _ = runtime.system.write("/system/drivers/ps2-bus: has two channels\n");
// keep the bus quiet until we have tested the ports and are ready to use them // keep the bus quiet until we have tested the ports and are ready to use them
controller.disablePort(.two); controller.disablePort(.two);
} else { } else {
_ = runtime.system.write("system/drivers/ps2-bus: has one channel\n"); _ = runtime.system.write("/system/drivers/ps2-bus: has one channel\n");
} }
// interface tests: always test port 1, test port 2 only if it exists // interface tests: always test port 1, test port 2 only if it exists
const port_one_works = (controller.testPort(.one) orelse { const port_one_works = (controller.testPort(.one) orelse {
_ = runtime.system.write("system/drivers/ps2-bus: port 1 test timed out\n"); _ = runtime.system.write("/system/drivers/ps2-bus: port 1 test timed out\n");
return; return;
}) == ps2.response_port_test_passed; }) == ps2.response_port_test_passed;
var port_two_works = false; var port_two_works = false;
if (has_two_channels) { if (has_two_channels) {
port_two_works = (controller.testPort(.two) orelse { port_two_works = (controller.testPort(.two) orelse {
_ = runtime.system.write("system/drivers/ps2-bus: port 2 test timed out\n"); _ = runtime.system.write("/system/drivers/ps2-bus: port 2 test timed out\n");
return; return;
}) == ps2.response_port_test_passed; }) == ps2.response_port_test_passed;
} }
if (!port_one_works and !port_two_works) { if (!port_one_works and !port_two_works) {
_ = runtime.system.write("system/drivers/ps2-bus: no usable ports\n"); _ = runtime.system.write("/system/drivers/ps2-bus: no usable ports\n");
return; return;
} }
@@ -188,16 +188,16 @@ pub fn main() void {
// abort bring-up of the other one // abort bring-up of the other one
if (port_one_works) { if (port_one_works) {
if (controller.resetDevice(.one)) |passed| { if (controller.resetDevice(.one)) |passed| {
if (!passed) _ = runtime.system.write("system/drivers/ps2-bus: port 1 device reset failed\n"); if (!passed) _ = runtime.system.write("/system/drivers/ps2-bus: port 1 device reset failed\n");
} else { } else {
_ = runtime.system.write("system/drivers/ps2-bus: port 1 device reset timed out\n"); _ = runtime.system.write("/system/drivers/ps2-bus: port 1 device reset timed out\n");
} }
} }
if (port_two_works) { if (port_two_works) {
if (controller.resetDevice(.two)) |passed| { if (controller.resetDevice(.two)) |passed| {
if (!passed) _ = runtime.system.write("system/drivers/ps2-bus: port 2 device reset failed\n"); if (!passed) _ = runtime.system.write("/system/drivers/ps2-bus: port 2 device reset failed\n");
} else { } else {
_ = runtime.system.write("system/drivers/ps2-bus: port 2 device reset timed out\n"); _ = runtime.system.write("/system/drivers/ps2-bus: port 2 device reset timed out\n");
} }
} }
@@ -207,13 +207,13 @@ pub fn main() void {
if (port_one_works) port_device_types[@intFromEnum(ps2.Port.one)] = spawnIdentifiedDriver(controller, .one); if (port_one_works) port_device_types[@intFromEnum(ps2.Port.one)] = spawnIdentifiedDriver(controller, .one);
if (port_two_works) port_device_types[@intFromEnum(ps2.Port.two)] = spawnIdentifiedDriver(controller, .two); if (port_two_works) port_device_types[@intFromEnum(ps2.Port.two)] = spawnIdentifiedDriver(controller, .two);
} else { } else {
_ = runtime.system.write("system/drivers/ps2-bus: no PS/2 controller found\n"); _ = runtime.system.write("/system/drivers/ps2-bus: no PS/2 controller found\n");
return; return;
} }
const controller = maybe_controller.?; const controller = maybe_controller.?;
const interrupt_index = maybe_interrupt_index orelse { const interrupt_index = maybe_interrupt_index orelse {
_ = runtime.system.write("system/drivers/ps2-bus: controller is missing its IRQ\n"); _ = runtime.system.write("/system/drivers/ps2-bus: controller is missing its IRQ\n");
return; return;
}; };
@@ -221,11 +221,11 @@ pub fn main() void {
// well-known id so the children can find it, the way input subscribers find // well-known id so the children can find it, the way input subscribers find
// the input service. // the input service.
const endpoint = ipc.createIpcEndpoint() orelse { const endpoint = ipc.createIpcEndpoint() orelse {
_ = runtime.system.write("system/drivers/ps2-bus: no endpoint\n"); _ = runtime.system.write("/system/drivers/ps2-bus: no endpoint\n");
return; return;
}; };
if (!ipc.register(.ps2_bus, endpoint)) { if (!ipc.register(.ps2_bus, endpoint)) {
_ = runtime.system.write("system/drivers/ps2-bus: register failed\n"); _ = runtime.system.write("/system/drivers/ps2-bus: register failed\n");
return; return;
} }
@@ -234,7 +234,7 @@ pub fn main() void {
// let the controller raise them — an interrupt with nobody bound is lost. // let the controller raise them — an interrupt with nobody bound is lost.
controller.drainOutputBuffer(); controller.drainOutputBuffer();
if (!device.irqBind(controller.device_id, interrupt_index, endpoint)) { if (!device.irqBind(controller.device_id, interrupt_index, endpoint)) {
_ = runtime.system.write("system/drivers/ps2-bus: irq_bind failed\n"); _ = runtime.system.write("/system/drivers/ps2-bus: irq_bind failed\n");
return; return;
} }
@@ -253,21 +253,21 @@ pub fn main() void {
.gsi = descriptor.resources[auxiliary_index].start, .gsi = descriptor.resources[auxiliary_index].start,
}; };
} else { } else {
_ = runtime.system.write("system/drivers/ps2-bus: auxiliary irq_bind failed\n"); _ = runtime.system.write("/system/drivers/ps2-bus: auxiliary irq_bind failed\n");
} }
} }
} }
} }
var configuration = controller.readConfigurationByte() orelse { var configuration = controller.readConfigurationByte() orelse {
_ = runtime.system.write("system/drivers/ps2-bus: controller configuration timed out\n"); _ = runtime.system.write("/system/drivers/ps2-bus: controller configuration timed out\n");
return; return;
}; };
if (port_device_types[@intFromEnum(ps2.Port.one)] != null) configuration |= ps2.Port.one.interruptBit(); if (port_device_types[@intFromEnum(ps2.Port.one)] != null) configuration |= ps2.Port.one.interruptBit();
if (maybe_auxiliary_interrupt != null) configuration |= ps2.Port.two.interruptBit(); if (maybe_auxiliary_interrupt != null) configuration |= ps2.Port.two.interruptBit();
_ = controller.writeConfigurationByte(configuration); _ = controller.writeConfigurationByte(configuration);
_ = runtime.system.write("system/drivers/ps2-bus: ok\n"); _ = runtime.system.write("/system/drivers/ps2-bus: ok\n");
// The forwarding loop: an IRQ1 notification drains the output buffer, routing // The forwarding loop: an IRQ1 notification drains the output buffer, routing
// each byte to the attached driver of the port it came from; a client message // each byte to the attached driver of the port it came from; a client message
+154 -34
View File
@@ -1,68 +1,188 @@
//! /system/drivers/usb-xhci-bus — the xHCI (USB 3) host-controller bus driver. //! /system/drivers/usb-xhci-bus — the xHCI (USB 3) host-controller bus driver.
//! The device manager spawns **one instance per controller** it discovers (a machine //! The device manager spawns **one instance per controller** it discovers (a
//! can carry several), passing the controller's device-tree id as argv[1]; this //! machine can carry several), passing the controller's device-tree id as
//! instance claims that device and no other, so multiple instances never fight over //! argv[1]; this instance claims that device and no other, so multiple
//! hardware. This increment proves the plumbing: parse the id, claim the controller, //! instances never fight over hardware.
//! and report its MMIO window. The next increments map the registers and bring the //!
//! controller up (reset, rings, port scan), then enumerate the USB devices on the //! M18.2 (this increment): after the hello, real hardware — map the xHC's
//! bus with the usb-abi request builders and publish each with `device_register`. //! register window (the first memory BAR; resource 0 is the ECAM config
//! space), read the capability registers, and walk the root-hub ports: one
//! `child_added` report to the manager per connected port, carrying the port
//! number and the PORTSC speed class as identity. No transfer rings yet —
//! descriptors and USB class matching are the USB track; the connect bit and
//! speed come straight from PORTSC, which reflects hardware state whether or
//! not the controller is running.
const std = @import("std"); const std = @import("std");
const runtime = @import("runtime"); const runtime = @import("runtime");
const protocol = runtime.device_manager_protocol;
const device = runtime.device; const device = runtime.device;
/// Format one whole log line and emit it in a single `debug_write`, so concurrent /// Format one whole log line and emit it in a single `debug_write`, so
/// instances (one per controller) can never interleave mid-line. /// concurrent instances (one per controller) can never interleave mid-line.
fn writeLine(comptime fmt: []const u8, arguments: anytype) void { fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined; var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return); _ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
} }
pub fn main(init: runtime.process.Init) void { var controller_id: u64 = protocol.no_device;
const argument = init.arguments.get(1) orelse {
_ = runtime.system.write("usb-xhci-bus: missing controller device id (argv[1])\n");
return;
};
const controller_id = std.fmt.parseInt(u64, argument, 10) catch {
writeLine("usb-xhci-bus: malformed controller device id '{s}'\n", .{argument});
return;
};
/// Claim the assigned controller, find its register window, and hello the
/// manager. Any failure returns false: the process exits cleanly, which the
/// manager reads as "meant to stop" — a missing assignment is not a crash loop.
fn initialise(endpoint: runtime.ipc.Handle) bool {
_ = endpoint;
if (!device.claim(controller_id)) { if (!device.claim(controller_id)) {
writeLine("usb-xhci-bus: unable to claim controller device {d}\n", .{controller_id}); writeLine("/system/drivers/usb-xhci-bus: unable to claim controller device {d}\n", .{controller_id});
return; return false;
} }
// Fetch our own descriptor back for the controller's resources. // Fetch our own descriptor back for the controller's resources.
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch { const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("usb-xhci-bus: out of memory\n"); _ = runtime.system.write("/system/drivers/usb-xhci-bus: out of memory\n");
return; return false;
}; };
const total = device.enumerate(buffer); const total = device.enumerate(buffer);
const descriptor = for (buffer[0..@min(total, buffer.len)]) |d| { const descriptor = for (buffer[0..@min(total, buffer.len)]) |d| {
if (d.id == controller_id) break d; if (d.id == controller_id) break d;
} else { } else {
writeLine("usb-xhci-bus: device {d} not in the device tree\n", .{controller_id}); writeLine("/system/drivers/usb-xhci-bus: device {d} not in the device tree\n", .{controller_id});
return; return false;
}; };
// The controller's operational registers live behind BAR0, enumerated as the // The xHC's registers live behind the first memory BAR. Resource 0 is the
// device's first memory resource. // function's ECAM configuration space (M15), so the walk starts at 1.
const register_window = for (descriptor.resources[0..@intCast(descriptor.resource_count)]) |resource| { var register_index: u64 = 0;
if (resource.kind == @intFromEnum(device.ResourceKind.memory)) break resource; const register_window = for (descriptor.resources[1..@intCast(descriptor.resource_count)], 1..) |resource, index| {
if (resource.kind == @intFromEnum(device.ResourceKind.memory)) {
register_index = index;
break resource;
}
} else { } else {
writeLine("usb-xhci-bus: controller device {d} has no MMIO window\n", .{controller_id}); writeLine("/system/drivers/usb-xhci-bus: controller device {d} has no register BAR\n", .{controller_id});
return; return false;
}; };
writeLine("usb-xhci-bus: claimed controller device {d} (registers at 0x{x}, {d} bytes)\n", .{ writeLine("/system/drivers/usb-xhci-bus: claimed controller device {d} (registers at 0x{x}, {d} bytes)\n", .{
controller_id, controller_id,
register_window.start, register_window.start,
register_window.len, register_window.len,
}); });
register_base = device.mmioMap(controller_id, register_index) orelse {
_ = runtime.system.write("/system/drivers/usb-xhci-bus: mmio_map failed\n");
return false;
};
// Controller bring-up (map the window, reset, rings, port scan) is the next // The handshake: role, protocol version, assignment — inside the manager's
// increment; stay resident as the bus's supervisor in the meantime. // deadline (the lookup retries cover the manager still registering).
while (true) runtime.system.sleep(1000); var manager: ?runtime.ipc.Handle = null;
var tries: u32 = 0;
while (manager == null and tries < 100) : (tries += 1) {
manager = runtime.ipc.lookup(.device_manager);
if (manager == null) runtime.system.sleep(20);
}
const h = manager orelse {
_ = runtime.system.write("/system/drivers/usb-xhci-bus: no device manager to hello\n");
return false;
};
const hello = protocol.Hello{ .role = @intFromEnum(protocol.Role.bus), .device_id = controller_id };
var reply: [protocol.message_maximum]u8 = undefined;
const n = runtime.ipc.call(h, std.mem.asBytes(&hello), &reply) catch {
_ = runtime.system.write("/system/drivers/usb-xhci-bus: hello call failed\n");
return false;
};
if (n < protocol.reply_size or std.mem.bytesToValue(protocol.HelloReply, reply[0..protocol.reply_size]).status != 0) {
_ = runtime.system.write("/system/drivers/usb-xhci-bus: hello refused\n");
return false;
}
_ = runtime.system.write("/system/drivers/usb-xhci-bus: hello acknowledged\n");
scanPorts(h);
return true;
}
var register_base: usize = 0;
/// One 32-bit volatile register read at `offset` from the mapped window.
fn readRegister(offset: usize) u32 {
const register: *volatile u32 = @ptrFromInt(register_base + offset);
return register.*;
}
/// The xHCI default Protocol Speed IDs (the PORTSC port-speed field, bits 13:10)
/// decoded to human names — the boot-log breadcrumb for what actually enumerated on
/// a port, the USB analog of the pci-bus class-code line. A controller may redefine
/// these through its Supported Protocol capability, but the defaults cover every
/// speed QEMU and real hardware report at this (pre-descriptor) stage.
fn speedName(speed: u32) []const u8 {
return switch (speed) {
1 => "Full-speed (USB 2.0, 12 Mb/s)",
2 => "Low-speed (USB 2.0, 1.5 Mb/s)",
3 => "High-speed (USB 2.0, 480 Mb/s)",
4 => "SuperSpeed (USB 3.0, 5 Gb/s)",
5 => "SuperSpeedPlus (USB 3.1, 10 Gb/s)",
else => "unknown speed",
};
}
/// The root-hub port scan: read the capability registers for the port count
/// and the operational-register offset, then one PORTSC per port. The connect
/// bit (CCS) and the speed field reflect hardware state directly — no
/// controller reset or run needed to *see* the devices; driving them needs the
/// rings (the USB track).
fn scanPorts(manager: runtime.ipc.Handle) void {
// Capability registers: CAPLENGTH is byte 0 of the first dword; HCSPARAMS1
// carries MaxPorts in bits 31:24.
const capability_length = readRegister(0) & 0xFF;
const structural = readRegister(0x04);
const maximum_ports: u32 = structural >> 24;
writeLine("/system/drivers/usb-xhci-bus: {d} root-hub ports\n", .{maximum_ports});
// PORTSC registers: operational base + 0x400 + 0x10 per port (1-based).
var port: u32 = 1;
var connected: u32 = 0;
while (port <= maximum_ports) : (port += 1) {
const port_status = readRegister(capability_length + 0x400 + 0x10 * (port - 1));
if (port_status & 1 == 0) continue; // CCS: nothing connected
connected += 1;
const speed = (port_status >> 10) & 0xF; // the PORTSC port-speed class
writeLine("/system/drivers/usb-xhci-bus: port {d} connected — {s} (speed class {d})\n", .{ port, speedName(speed), speed });
const report = protocol.ChildAdded{
.parent = controller_id,
.bus_address = port,
.identity = speed,
};
var reply: [protocol.message_maximum]u8 = undefined;
_ = runtime.ipc.call(manager, std.mem.asBytes(&report), &reply) catch {
writeLine("/system/drivers/usb-xhci-bus: child report for port {d} failed\n", .{port});
continue;
};
}
if (connected == 0) _ = runtime.system.write("/system/drivers/usb-xhci-bus: no devices connected\n");
}
/// No bus protocol to serve yet — transfer requests arrive with the USB track.
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = message;
_ = reply;
_ = sender;
_ = capability;
return 0;
}
pub fn main(init: runtime.process.Init) void {
const argument = init.arguments.get(1) orelse {
_ = runtime.system.write("/system/drivers/usb-xhci-bus: missing controller device id (argv[1])\n");
return;
};
controller_id = std.fmt.parseInt(u64, argument, 10) catch {
writeLine("/system/drivers/usb-xhci-bus: malformed controller device id '{s}'\n", .{argument});
return;
};
runtime.service.run(protocol.message_maximum, .{
.init = initialise,
.on_message = onMessage,
});
} }
pub const panic = runtime.panic; pub const panic = runtime.panic;
+43 -1
View File
@@ -104,6 +104,19 @@ pub fn ownerOf(id: u64) ?u32 {
return claimed[@intCast(id)]; return claimed[@intCast(id)];
} }
/// Release every claim held by `owner` — called by the process layer on every
/// path out of a process (exit, fault, kill), so a restarted driver can claim its
/// hardware again (docs/process-lifecycle.md iron rule 1: cleanup is the kernel's
/// job). The devices stay in the table — they describe hardware, which did not go
/// away — only their ownership clears.
pub fn releaseAllOwnedBy(owner: u32) void {
for (claimed[0..count]) |*slot| {
if (slot.*) |o| {
if (o == owner) slot.* = null;
}
}
}
/// Resource `index` of device `id`, or null if out of range. /// Resource `index` of device `id`, or null if out of range.
pub fn resourceOf(id: u64, index: u64) ?device_abi.ResourceDescriptor { pub fn resourceOf(id: u64, index: u64) ?device_abi.ResourceDescriptor {
if (id >= count) return null; if (id >= count) return null;
@@ -118,7 +131,16 @@ pub fn resourceOf(id: u64, index: u64) ?device_abi.ResourceDescriptor {
/// and would otherwise vacuously "fit" anywhere. /// and would otherwise vacuously "fit" anywhere.
fn contains(parent: device_abi.ResourceDescriptor, child: device_abi.ResourceDescriptor) bool { fn contains(parent: device_abi.ResourceDescriptor, child: device_abi.ResourceDescriptor) bool {
if (parent.kind != child.kind) return false; if (parent.kind != child.kind) return false;
if (child.kind == @intFromEnum(device_abi.ResourceKind.irq)) return parent.start == child.start; if (child.kind == @intFromEnum(device_abi.ResourceKind.irq)) {
// Range containment: an interrupt line is still indivisible (a child owns
// exactly one GSI), but a parent may own a *range* of lines so a broad
// owner — the acpi-tables node, whose firmware names any legacy IRQ —
// can contain its children's specific lines. A length-1 parent range is
// exactly the old equality rule, so existing single-IRQ parents are
// unaffected.
const span = if (parent.len == 0) 1 else parent.len;
return child.start >= parent.start and child.start < parent.start + span;
}
if (child.len == 0 or parent.len == 0) return false; if (child.len == 0 or parent.len == 0) return false;
// No overflow: a resource that wraps the address space is not containable. // No overflow: a resource that wraps the address space is not containable.
const child_end = std.math.add(u64, child.start, child.len) catch return false; const child_end = std.math.add(u64, child.start, child.len) catch return false;
@@ -167,6 +189,26 @@ pub fn register(parent_id: u64, owner: u32, descriptor: *const device_abi.Device
if (!ok) return error.NotContained; if (!ok) return error.NotContained;
} }
// Idempotent on exact match (docs/m19-m20-plan.md decision 3): a restarted
// registering bus re-registers what it rediscovers, and the table has no
// unregister — an identical (class, identity, resources) child under the
// same parent returns the existing id instead of appending a duplicate.
for (devices[0..count]) |*existing| {
if (existing.parent != parent_id) continue;
if (existing.class != descriptor.class) continue;
if (existing.pci_class != descriptor.pci_class) continue;
if (existing.hid_len != descriptor.hid_len) continue;
if (!std.mem.eql(u8, existing.hid[0..@intCast(existing.hid_len)], descriptor.hid[0..@intCast(descriptor.hid_len)])) continue;
if (existing.resource_count != descriptor.resource_count) continue;
var same = true;
for (0..@intCast(descriptor.resource_count)) |i| {
const a = existing.resources[i];
const b = descriptor.resources[i];
if (a.kind != b.kind or a.start != b.start or a.len != b.len) same = false;
}
if (same) return existing.id;
}
var d = std.mem.zeroes(device_abi.DeviceDescriptor); var d = std.mem.zeroes(device_abi.DeviceDescriptor);
d.id = count; d.id = count;
d.parent = parent_id; d.parent = parent_id;
+40 -27
View File
@@ -77,12 +77,12 @@ fn kmain(boot_information: *const BootInformation) noreturn {
architecture.setFaultHandler(onException); architecture.setFaultHandler(onException);
architecture.init(); architecture.init();
status("danos: initialising kernel...\n"); status("/system/kernel: initialising kernel...\n");
log.write(if (console.present()) log.write(if (console.present())
"danos: framebuffer console online (bootstrap; graphics driver later)\n" "/system/kernel: framebuffer console online (bootstrap; graphics driver later)\n"
else else
"danos: no framebuffer (headless) -> logging to serial/debugcon only\n"); "/system/kernel: no framebuffer (headless) -> logging to serial/debugcon only\n");
log.write("danos: cpu tables online (GDT, IDT, TSS)\n"); log.write("/system/kernel: cpu tables online (GDT, IDT, TSS)\n");
log.print(" resolution : {d}x{d}\n", .{ fb.width, fb.height }); log.print(" resolution : {d}x{d}\n", .{ fb.width, fb.height });
log.print(" pitch : {d} bytes\n", .{fb.pitch}); log.print(" pitch : {d} bytes\n", .{fb.pitch});
log.print(" format : {s}\n", .{@tagName(fb.format)}); log.print(" format : {s}\n", .{@tagName(fb.format)});
@@ -105,7 +105,7 @@ fn kmain(boot_information: *const BootInformation) noreturn {
const total_bytes = total_pages * abi.page_size; const total_bytes = total_pages * abi.page_size;
const gib = 1 << 30; const gib = 1 << 30;
log.write("\ndanos: physical memory\n"); log.write("\n/system/kernel: physical memory\n");
log.print(" total RAM : {d}.{d:0>2} GiB ({d} MiB) - RAM the firmware reported\n", .{ total_bytes / gib, (total_bytes % gib) * 100 / gib, mib(total_pages) }); log.print(" total RAM : {d}.{d:0>2} GiB ({d} MiB) - RAM the firmware reported\n", .{ total_bytes / gib, (total_bytes % gib) * 100 / gib, mib(total_pages) });
log.print(" usable : {d} MiB - free RAM (incl. reclaimed boot-services memory)\n", .{mib(usable_pages)}); log.print(" usable : {d} MiB - free RAM (incl. reclaimed boot-services memory)\n", .{mib(usable_pages)});
log.print(" reserved : {d} MiB - kernel image, boot stack, ACPI, runtime services\n", .{mib(reserved_pages)}); log.print(" reserved : {d} MiB - kernel image, boot stack, ACPI, runtime services\n", .{mib(reserved_pages)});
@@ -119,7 +119,7 @@ fn kmain(boot_information: *const BootInformation) noreturn {
// until SMP bring-up; 0 means none was available (we stay uniprocessor). // until SMP bring-up; 0 means none was available (we stay uniprocessor).
ap_trampoline_page = pmm.allocBelow(0x100000) orelse 0; ap_trampoline_page = pmm.allocBelow(0x100000) orelse 0;
const s1 = pmm.stats(); const s1 = pmm.stats();
log.print("\ndanos: frame allocator online\n", .{}); log.print("\n/system/kernel: frame allocator online\n", .{});
log.print(" free frames: {d} ({d} MiB)\n", .{ s1.free_frames, mib(s1.free_frames) }); log.print(" free frames: {d} ({d} MiB)\n", .{ s1.free_frames, mib(s1.free_frames) });
const f0 = pmm.alloc(); const f0 = pmm.alloc();
const f1 = pmm.alloc(); const f1 = pmm.alloc();
@@ -133,14 +133,14 @@ fn kmain(boot_information: *const BootInformation) noreturn {
// Switch off the firmware's page tables onto our own (with real permissions). // Switch off the firmware's page tables onto our own (with real permissions).
architecture.enablePaging(pmm.alloc, pmm.free, boot_information); architecture.enablePaging(pmm.alloc, pmm.free, boot_information);
log.checkpoint(cp_paging); log.checkpoint(cp_paging);
log.print("\ndanos: paging enabled\n", .{}); log.print("\n/system/kernel: paging enabled\n", .{});
log.print(" page tables: root = 0x{x:0>16}\n", .{architecture.activePageTable()}); log.print(" page tables: root = 0x{x:0>16}\n", .{architecture.activePageTable()});
log.print(" kernel segs: {d} (mapped with W^X permissions)\n", .{boot_information.kernel_segment_count}); log.print(" kernel segs: {d} (mapped with W^X permissions)\n", .{boot_information.kernel_segment_count});
// Bring up the kernel heap (dynamic allocation), built on the VMM. // Bring up the kernel heap (dynamic allocation), built on the VMM.
heap.init(); heap.init();
log.checkpoint(cp_heap); log.checkpoint(cp_heap);
log.write("\ndanos: kernel heap online\n"); log.write("\n/system/kernel: kernel heap online\n");
// Measure the amount of resources the kernel is actually using // Measure the amount of resources the kernel is actually using
const s2 = pmm.stats(); const s2 = pmm.stats();
log.print(" Kernel footprint: {d} KiB\n", .{kib(s1.free_frames - s2.free_frames)}); log.print(" Kernel footprint: {d} KiB\n", .{kib(s1.free_frames - s2.free_frames)});
@@ -156,7 +156,7 @@ fn kmain(boot_information: *const BootInformation) noreturn {
}; };
if (platform.discover(boot_information, heap.allocator(), hal)) |devtree| { if (platform.discover(boot_information, heap.allocator(), hal)) |devtree| {
var device_tree = devtree; var device_tree = devtree;
log.write("\ndanos: device discovery online\n"); log.write("\n/system/kernel: device discovery online\n");
device_tree.dump(log.write); device_tree.dump(log.write);
// Snapshot the device tree for user-space drivers (device_enumerate/claim/ // Snapshot the device tree for user-space drivers (device_enumerate/claim/
@@ -164,7 +164,7 @@ fn kmain(boot_information: *const BootInformation) noreturn {
devices_broker.init(&device_tree); devices_broker.init(&device_tree);
if (devices_broker.dropped > 0) { if (devices_broker.dropped > 0) {
// Otherwise entirely silent: drivers would just never see that hardware. // Otherwise entirely silent: drivers would just never see that hardware.
log.print("danos: WARNING {d} device(s) dropped — table full\n", .{devices_broker.dropped}); log.print("/system/kernel: WARNING {d} device(s) dropped — table full\n", .{devices_broker.dropped});
} }
// Install the device-IRQ trampolines, so a driver's irq_bind has vectors to // Install the device-IRQ trampolines, so a driver's irq_bind has vectors to
@@ -173,7 +173,7 @@ fn kmain(boot_information: *const BootInformation) noreturn {
// Power register map extracted from the FADT + AML, for confidence it parsed. // Power register map extracted from the FADT + AML, for confidence it parsed.
const pw = platform.powerInformation(); const pw = platform.powerInformation();
log.write("danos: power\n"); log.write("/system/kernel: power\n");
log.print(" pm1a_cnt : {s} 0x{x} (width {d})\n", .{ if (pw.pm1a_cnt.mmio) "mmio" else "io", pw.pm1a_cnt.address, pw.pm1a_cnt.width }); log.print(" pm1a_cnt : {s} 0x{x} (width {d})\n", .{ if (pw.pm1a_cnt.mmio) "mmio" else "io", pw.pm1a_cnt.address, pw.pm1a_cnt.width });
if (pw.s5) |s| { if (pw.s5) |s| {
log.print(" S5 slp_typ : a={d} b={d}\n", .{ s.slp_typ_a, s.slp_typ_b }); log.print(" S5 slp_typ : a={d} b={d}\n", .{ s.slp_typ_a, s.slp_typ_b });
@@ -221,7 +221,7 @@ fn kmain(boot_information: *const BootInformation) noreturn {
}); });
if (pinfo.spcr_uart) |u| architecture.serialReconfigure(u.mmio, u.address); if (pinfo.spcr_uart) |u| architecture.serialReconfigure(u.mmio, u.address);
log.write("danos: platform\n"); log.write("/system/kernel: platform\n");
log.print(" 8259 PIC : {s}\n", .{if (pinfo.pic_present) "present" else "absent"}); log.print(" 8259 PIC : {s}\n", .{if (pinfo.pic_present) "present" else "absent"});
log.print(" lapic base : 0x{x}\n", .{pinfo.lapic_base}); log.print(" lapic base : 0x{x}\n", .{pinfo.lapic_base});
log.print(" hpet base : 0x{x}\n", .{hpet_base}); log.print(" hpet base : 0x{x}\n", .{hpet_base});
@@ -237,7 +237,7 @@ fn kmain(boot_information: *const BootInformation) noreturn {
if (platform.cpusDropped() > 0) if (platform.cpusDropped() > 0)
log.print(" cpus : WARNING {d} core(s) beyond pool cap dropped\n", .{platform.cpusDropped()}); log.print(" cpus : WARNING {d} core(s) beyond pool cap dropped\n", .{platform.cpusDropped()});
} else |err| { } else |err| {
log.print("\ndanos: device discovery failed: {s}\n", .{@errorName(err)}); log.print("\n/system/kernel: device discovery failed: {s}\n", .{@errorName(err)});
} }
log.checkpoint(cp_discovery); log.checkpoint(cp_discovery);
@@ -248,14 +248,14 @@ fn kmain(boot_information: *const BootInformation) noreturn {
// Register the current context as the first task before enabling preemption. // Register the current context as the first task before enabling preemption.
scheduler.init(4); scheduler.init(4);
log.checkpoint(cp_scheduler); log.checkpoint(cp_scheduler);
log.write("\ndanos: scheduler online\n"); log.write("\n/system/kernel: scheduler online\n");
// Start the timer and unmask interrupts — the kernel now has a heartbeat, and // Start the timer and unmask interrupts — the kernel now has a heartbeat, and
// the timer preempts among tasks. // the timer preempts among tasks.
architecture.startTimer(); architecture.startTimer();
architecture.enableInterrupts(); architecture.enableInterrupts();
log.checkpoint(cp_timer); log.checkpoint(cp_timer);
log.print("danos: timer online ({d} Hz tick; timer clock {d} MHz, clock {d} MHz; calibrated via {s})\n", .{ architecture.timer_hz, architecture.timerClockHz() / 1_000_000, architecture.clockHz() / 1_000_000, architecture.timerCalibrationSource() }); log.print("/system/kernel: timer online ({d} Hz tick; timer clock {d} MHz, clock {d} MHz; calibrated via {s})\n", .{ architecture.timer_hz, architecture.timerClockHz() / 1_000_000, architecture.clockHz() / 1_000_000, architecture.timerCalibrationSource() });
// Wake the other cores (application processors). A no-op on a single-core // Wake the other cores (application processors). A no-op on a single-core
// machine; on SMP each AP climbs to long mode and reports in (docs/smp.md). // machine; on SMP each AP climbs to long mode and reports in (docs/smp.md).
@@ -269,7 +269,7 @@ fn kmain(boot_information: *const BootInformation) noreturn {
} }
log.checkpoint(cp_running); log.checkpoint(cp_running);
status("kernel initialised.\n"); status("/system/kernel: initialised.\n");
// Publish the initial-ramdisk so user space can `system_spawn` its bundled // Publish the initial-ramdisk so user space can `system_spawn` its bundled
// binaries by name. The kernel no longer launches them itself: init is the // binaries by name. The kernel no longer launches them itself: init is the
@@ -282,10 +282,10 @@ fn kmain(boot_information: *const BootInformation) noreturn {
// manager then discovers the hardware and spawns each driver. init runs on its own // manager then discovers the hardware and spawns each driver. init runs on its own
// address space, preemptively — this boot context becomes the BSP's idle loop. // address space, preemptively — this boot context becomes the BSP's idle loop.
if (boot_information.init_len != 0) { if (boot_information.init_len != 0) {
status("starting /system/services/init...\n"); status("/system/kernel: starting /system/services/init...\n");
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len]; const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len];
process.spawnProcess(image, 4, &.{"/system/services/init"}) catch |err| { process.spawnProcess(image, 4, &.{"/system/services/init"}) catch |err| {
statusPrint("/system/services/init failed to load: {s}\n", .{@errorName(err)}); statusPrint("/system/kernel: /system/services/init failed to load: {s}\n", .{@errorName(err)});
}; };
} else { } else {
status("no /system/services/init on the boot volume.\n"); status("no /system/services/init on the boot volume.\n");
@@ -294,7 +294,7 @@ fn kmain(boot_information: *const BootInformation) noreturn {
// Become the idle task: drop below every real task and halt until an // Become the idle task: drop below every real task and halt until an
// interrupt. The timer keeps preempting into init and any other work. // interrupt. The timer keeps preempting into init and any other work.
scheduler.setPriority(0); scheduler.setPriority(0);
status("\nkernel idle; user space is running.\n"); status("\n/system/kernel: kernel idle; user space is running.\n");
architecture.halt(); architecture.halt();
} }
@@ -321,7 +321,7 @@ fn bringUpSecondaries() void {
// vector addresses it). It's kept for the system's life — armed only during a // vector addresses it). It's kept for the system's life — armed only during a
// wake, inert (zeroed, non-executable) otherwise — so cores can be re-woken later. // wake, inert (zeroed, non-executable) otherwise — so cores can be re-woken later.
if (ap_trampoline_page == 0) { if (ap_trampoline_page == 0) {
log.write("danos: smp: no low page for the AP trampoline; staying uniprocessor\n"); log.write("/system/kernel: smp: no low page for the AP trampoline; staying uniprocessor\n");
return; return;
} }
architecture.setTrampolinePage(ap_trampoline_page); architecture.setTrampolinePage(ap_trampoline_page);
@@ -333,7 +333,7 @@ fn bringUpSecondaries() void {
if (std.mem.eql(u8, tc, "smp-retry")) architecture.testFailNextWakes(1); if (std.mem.eql(u8, tc, "smp-retry")) architecture.testFailNextWakes(1);
} }
log.print("\ndanos: bringing up {d} application processor(s)\n", .{cores.len - 1}); log.print("\n/system/kernel: bringing up {d} application processor(s)\n", .{cores.len - 1});
const maximum_wake_attempts = 3; // a core that misses the first INIT-SIPI-SIPI gets retried const maximum_wake_attempts = 3; // a core that misses the first INIT-SIPI-SIPI gets retried
for (cores[1..], 1..) |core, index| { for (cores[1..], 1..) |core, index| {
const stack = heap.allocator().alloc(u8, parameters.kernel_stack_size) catch { const stack = heap.allocator().alloc(u8, parameters.kernel_stack_size) catch {
@@ -344,7 +344,7 @@ fn bringUpSecondaries() void {
// This core's dedicated fault stack — allocated only now that the core is // This core's dedicated fault stack — allocated only now that the core is
// real, rather than reserved statically for every possible core. // real, rather than reserved statically for every possible core.
const fault_stack = heap.allocator().alloc(u8, architecture.fault_stack_size) catch { const fault_stack = heap.allocator().alloc(u8, architecture.fault_stack_size) catch {
log.print(" cpu apic_id {d}: no fault stack; skipped\n", .{core.apic_id}); log.print("/system/kernel: cpu apic_id {d}: no fault stack; skipped\n", .{core.apic_id});
continue; continue;
}; };
architecture.setFaultStack(index, (@intFromPtr(fault_stack.ptr) + fault_stack.len) & ~@as(usize, 15)); architecture.setFaultStack(index, (@intFromPtr(fault_stack.ptr) + fault_stack.len) & ~@as(usize, 15));
@@ -353,14 +353,14 @@ fn bringUpSecondaries() void {
while (attempt <= maximum_wake_attempts) : (attempt += 1) { while (attempt <= maximum_wake_attempts) : (attempt += 1) {
if (architecture.startSecondary(core.apic_id, stack_top, @intFromPtr(pc), index)) { if (architecture.startSecondary(core.apic_id, stack_top, @intFromPtr(pc), index)) {
pc.online = true; pc.online = true;
log.print(" cpu apic_id {d}: online (attempt {d})\n", .{ core.apic_id, attempt }); log.print("/system/kernel: cpu apic_id {d}: online (attempt {d})\n", .{ core.apic_id, attempt });
break; break;
} }
if (attempt == maximum_wake_attempts) if (attempt == maximum_wake_attempts)
log.print(" cpu apic_id {d}: no response after {d} attempts (parked)\n", .{ core.apic_id, maximum_wake_attempts }); log.print("/system/kernel: cpu apic_id {d}: no response after {d} attempts (parked)\n", .{ core.apic_id, maximum_wake_attempts });
} }
} }
log.print("danos: {d}/{d} cores online\n", .{ scheduler.onlineCount(), cores.len }); log.print("/system/kernel: {d}/{d} cores online\n", .{ scheduler.onlineCount(), cores.len });
} }
/// A user-facing status line: to the diagnostic `log` *and* the on-screen console /// A user-facing status line: to the diagnostic `log` *and* the on-screen console
@@ -412,13 +412,26 @@ fn recoverableFault(vector: u64) bool {
/// plus a POST code and a persistent breadcrumb. (A ring-3 fault on a *borrowed* /// plus a POST code and a persistent breadcrumb. (A ring-3 fault on a *borrowed*
/// kernel thread — process.run, the user-pf isolation probe — also lands here: there /// kernel thread — process.run, the user-pf isolation probe — also lands here: there
/// is no scheduled process to kill.) /// is no scheduled process to kill.)
/// Classify a CPU exception vector as the ExitReason a supervisor reads — the
/// fault classes of docs/process-lifecycle.md. Faults are exit reasons, never
/// signals delivered to the faulting process: recovery is restart, not a handler.
fn exitReasonForVector(vector: u64) abi.ExitReason {
return switch (vector) {
14 => .segmentation_fault, // page fault
6 => .illegal_instruction, // invalid opcode
0, 16, 19 => .arithmetic_fault, // divide error, x87, SIMD
13 => .protection_fault, // general protection
else => .fault,
};
}
fn onException(state: *const architecture.CpuState) noreturn { fn onException(state: *const architecture.CpuState) noreturn {
if (architecture.fromUser(state) and scheduler.currentIsUserProcess() and recoverableFault(state.vector)) { if (architecture.fromUser(state) and scheduler.currentIsUserProcess() and recoverableFault(state.vector)) {
statusPrint("\ndanos: process {d} ({s}) killed by {s} (vector {d}) on core {d}\n", .{ scheduler.currentId(), scheduler.current().name(), architecture.exceptionName(state.vector), state.vector, scheduler.currentCpuIndex() }); statusPrint("\n/system/kernel: process {d} ({s}) killed by {s} (vector {d}) on core {d}\n", .{ scheduler.currentId(), scheduler.current().name(), architecture.exceptionName(state.vector), state.vector, scheduler.currentCpuIndex() });
statusPrint(" error code : 0x{x}\n", .{state.error_code}); statusPrint(" error code : 0x{x}\n", .{state.error_code});
statusPrint(" IP : 0x{x:0>16}\n", .{architecture.instructionPointer(state)}); statusPrint(" IP : 0x{x:0>16}\n", .{architecture.instructionPointer(state)});
if (architecture.faultAddress(state)) |address| statusPrint(" fault addr : 0x{x:0>16}\n", .{address}); if (architecture.faultAddress(state)) |address| statusPrint(" fault addr : 0x{x:0>16}\n", .{address});
process.killCurrentProcess(); // reclaims everything, reschedules; never returns process.killCurrentProcess(exitReasonForVector(state.vector)); // reclaims everything, reschedules; never returns
} }
log.checkpoint(cp_exception); log.checkpoint(cp_exception);
+225 -4
View File
@@ -137,6 +137,7 @@ pub fn init() void {
architecture.setSystemCallHandler(system_call); architecture.setSystemCallHandler(system_call);
scheduler.terminate_current_hook = terminateCurrentLocked; scheduler.terminate_current_hook = terminateCurrentLocked;
scheduler.reap_task_hook = reapTaskLocked; scheduler.reap_task_hook = reapTaskLocked;
scheduler.timer_tick_hook = timerSweepLocked;
} }
/// Return -1 (as an unsigned bit pattern) in the system_call result register. /// Return -1 (as an unsigned bit pattern) in the system_call result register.
@@ -164,6 +165,7 @@ fn system_call(state: *architecture.CpuState) void {
// A scheduled process tears down fully (terminateCurrent); a borrowed // A scheduled process tears down fully (terminateCurrent); a borrowed
// test thread unwinds back to the kernel that entered it. // test thread unwinds back to the kernel that entered it.
if (scheduler.currentIsUserProcess()) { if (scheduler.currentIsUserProcess()) {
scheduler.current().exit_reason = .exited;
terminateCurrent(); terminateCurrent();
} else architecture.userExit(); } else architecture.userExit();
}, },
@@ -199,6 +201,11 @@ fn system_call(state: *architecture.CpuState) void {
.clock => systemClock(state), .clock => systemClock(state),
.process_enumerate => systemProcessEnumerate(state), .process_enumerate => systemProcessEnumerate(state),
.process_kill => systemProcessKill(state), .process_kill => systemProcessKill(state),
.process_exit_reason => systemProcessExitReason(state),
.process_subscribe => systemProcessSubscribe(state),
.signal_bind => systemSignalBind(state),
.process_signal => systemProcessSignal(state),
.timer_bind => systemTimerBind(state),
_ => fail(state), _ => fail(state),
} }
} }
@@ -291,6 +298,8 @@ fn systemDeviceEnumerate(state: *architecture.CpuState) void {
/// device_claim(id) -> 0/-1: take exclusive ownership of a device for this process. /// device_claim(id) -> 0/-1: take exclusive ownership of a device for this process.
fn systemDeviceClaim(state: *architecture.CpuState) void { fn systemDeviceClaim(state: *architecture.CpuState) void {
const claim_flags = sync.enter();
defer sync.leave(claim_flags);
if (devices_broker.claim(architecture.systemCallArg(state, 0), scheduler.current().id)) if (devices_broker.claim(architecture.systemCallArg(state, 0), scheduler.current().id))
architecture.setSystemCallResult(state, 0) architecture.setSystemCallResult(state, 0)
else else
@@ -305,10 +314,22 @@ fn systemMmioMap(state: *architecture.CpuState) void {
const resource_index = architecture.systemCallArg(state, 1); const resource_index = architecture.systemCallArg(state, 1);
const t = scheduler.current(); const t = scheduler.current();
if (t.aspace == 0) return fail(state); if (t.aspace == 0) return fail(state);
// Read the broker table under the lock: ring-3 device_register (M19) now
// mutates it concurrently on other cores, so a lock-free read here could
// see a torn resource (and a torn length used to panic the arithmetic
// below on integer overflow).
const r = blk: {
const flags = sync.enter();
defer sync.leave(flags);
const owner = devices_broker.ownerOf(device_id) orelse return fail(state); const owner = devices_broker.ownerOf(device_id) orelse return fail(state);
if (owner != t.id) return fail(state); // not claimed by this process if (owner != t.id) return fail(state); // not claimed by this process
const r = devices_broker.resourceOf(device_id, resource_index) orelse return fail(state); break :blk devices_broker.resourceOf(device_id, resource_index) orelse return fail(state);
};
if (r.kind != @intFromEnum(device_abi.ResourceKind.memory)) return fail(state); if (r.kind != @intFromEnum(device_abi.ResourceKind.memory)) return fail(state);
// A zero-length or wrapping window is not mappable — fail cleanly rather
// than underflow `r.len - 1`.
if (r.len == 0) return fail(state);
if (@addWithOverflow(r.start, r.len)[1] != 0) return fail(state);
if (t.device_map_next == 0) t.device_map_next = device_arena_base; if (t.device_map_next == 0) t.device_map_next = device_arena_base;
const first = r.start & ~@as(u64, page_size - 1); const first = r.start & ~@as(u64, page_size - 1);
@@ -444,6 +465,11 @@ fn systemDeviceRegister(state: *architecture.CpuState) void {
var descriptor: device_abi.DeviceDescriptor = undefined; var descriptor: device_abi.DeviceDescriptor = undefined;
if (!ipc.copyFromUser(t.aspace, descriptor_ptr, std.mem.asBytes(&descriptor))) return fail(state); if (!ipc.copyFromUser(t.aspace, descriptor_ptr, std.mem.asBytes(&descriptor))) return fail(state);
// Under the big kernel lock: the broker's table is also mutated by the
// death sweep (releaseAllOwnedBy) and read by enumerate on other cores —
// ring-3 registration (M19) made those genuinely concurrent.
const flags = sync.enter();
defer sync.leave(flags);
const id = devices_broker.register(parent_id, t.id, &descriptor) catch return fail(state); const id = devices_broker.register(parent_id, t.id, &descriptor) catch return fail(state);
architecture.setSystemCallResult(state, id); architecture.setSystemCallResult(state, id);
} }
@@ -552,7 +578,11 @@ pub var fault_kill_count: u64 = 0;
/// endpoint reference destroys the Endpoint, and a still-bound GSI would have an /// endpoint reference destroys the Endpoint, and a still-bound GSI would have an
/// ISR call notifyFromIsr on freed memory the next time the device fired. /// ISR call notifyFromIsr on freed memory the next time the device fired.
/// `releaseOwner` also leaves the line masked, so a dead driver's device goes /// `releaseOwner` also leaves the line masked, so a dead driver's device goes
/// quiet rather than storming. /// quiet rather than storming. (It drops MSI vectors by the same owner sweep.)
/// - Device claims are released with the IRQ bindings, so a restarted driver can
/// claim the same hardware again — the cleanup half of process-lifecycle.md's
/// iron rule 1. Claims hold no pointers, so ordering is free; they go here so
/// the exit notification (below, last) observes a fully-released child.
/// - A client this task still owes a reply to (it died between receive and reply) /// - A client this task still owes a reply to (it died between receive and reply)
/// is failed with -EPEER rather than left blocked forever — a dead server must /// is failed with -EPEER rather than left blocked forever — a dead server must
/// not hang its callers. /// not hang its callers.
@@ -565,7 +595,33 @@ pub var fault_kill_count: u64 = 0;
/// reference taken at spawn is dropped with it. /// reference taken at spawn is dropped with it.
/// Precondition: the big kernel lock is held. /// Precondition: the big kernel lock is held.
fn releaseTaskResourcesLocked(t: *scheduler.Task) void { fn releaseTaskResourcesLocked(t: *scheduler.Task) void {
recordExitLocked(t);
irq.releaseOwner(t.id); irq.releaseOwner(t.id);
devices_broker.releaseAllOwnedBy(t.id);
// The dying task's signal endpoint and one-shot timers go with it.
if (t.signal_endpoint) |raw| {
ipc.dropRef(@ptrCast(@alignCast(raw)));
t.signal_endpoint = null;
}
t.pending_signals = 0;
for (&one_shot_timers) |*slot| {
if (slot.*) |timer| {
if (timer.owner == t.id) {
ipc.dropRef(timer.endpoint);
slot.* = null;
}
}
}
// A dead subscriber's own subscriptions go first: it must not hear about
// itself, and the slots' endpoint references drop with it.
for (&exit_subscribers) |*slot| {
if (slot.*) |subscriber| {
if (subscriber.owner == t.id) {
ipc.dropRef(subscriber.endpoint);
slot.* = null;
}
}
}
if (t.ipc_client) |client| { if (t.ipc_client) |client| {
t.ipc_client = null; t.ipc_client = null;
client.ipc_status = -ipc.EPEER; client.ipc_status = -ipc.EPEER;
@@ -575,6 +631,12 @@ fn releaseTaskResourcesLocked(t: *scheduler.Task) void {
scheduler.removeFromWaitQueueLocked(t); scheduler.removeFromWaitQueueLocked(t);
scheduler.forgetIpcClientLocked(t); scheduler.forgetIpcClientLocked(t);
ipc.closeHandles(t); ipc.closeHandles(t);
// Publish the exit to every subscriber (docs/process-lifecycle.md): the same
// badge encoding as the supervisor's notification, and equally late, so a
// subscriber also observes a fully-released child.
for (&exit_subscribers) |*slot| {
if (slot.*) |subscriber| ipc.notifyLocked(subscriber.endpoint, abi.notify_exit_bit | t.id);
}
if (t.exit_endpoint) |raw| { if (t.exit_endpoint) |raw| {
const endpoint: *ipc.Endpoint = @ptrCast(@alignCast(raw)); const endpoint: *ipc.Endpoint = @ptrCast(@alignCast(raw));
t.exit_endpoint = null; t.exit_endpoint = null;
@@ -628,6 +690,7 @@ pub fn killProcess(caller_id: u32, target_id: u32) i64 {
const target = scheduler.taskByIdLocked(target_id) orelse return -ipc.ESRCH; const target = scheduler.taskByIdLocked(target_id) orelse return -ipc.ESRCH;
if (target.aspace == 0) return -ipc.ESRCH; // kernel tasks are not processes if (target.aspace == 0) return -ipc.ESRCH; // kernel tasks are not processes
if (target.supervisor != caller_id) return -ipc.EPERM; if (target.supervisor != caller_id) return -ipc.EPERM;
target.exit_reason = .killed;
if (target.state == .running) { if (target.state == .running) {
target.kill_pending = true; target.kill_pending = true;
} else { } else {
@@ -640,12 +703,170 @@ pub fn killProcess(caller_id: u32, target_id: u32) i64 {
/// The fault is confined to the process — the kernel trapped it on the task's own /// The fault is confined to the process — the kernel trapped it on the task's own
/// kernel stack and is intact — so everything the process held is reclaimed and the /// kernel stack and is intact — so everything the process held is reclaimed and the
/// core reschedules. The system keeps running; only the faulting process dies /// core reschedules. The system keeps running; only the faulting process dies
/// (docs/resilience.md: fault -> kill -> continue). /// (docs/resilience.md: fault -> kill -> continue). `reason` is the fault class
pub fn killCurrentProcess() noreturn { /// (from the vector), recorded for the supervisor's `process_exit_reason`.
pub fn killCurrentProcess(reason: abi.ExitReason) noreturn {
scheduler.current().exit_reason = reason;
fault_kill_count += 1; fault_kill_count += 1;
terminateCurrent(); terminateCurrent();
} }
/// The bounded record of recent deaths, for `process_exit_reason`: ids are never
/// reused, so a ring keyed by id is enough — a record evicted by wraparound reads
/// as -ESRCH, the same as an id that never lived, which a supervisor treats as
/// "too late to ask". Written under the big kernel lock by the reap.
const exit_record_capacity = 64;
const ExitRecord = struct { id: u32 = 0, supervisor: u32 = 0, reason: abi.ExitReason = .exited, valid: bool = false };
var exit_records: [exit_record_capacity]ExitRecord = .{ExitRecord{}} ** exit_record_capacity;
var exit_record_next: usize = 0;
/// Record a dying task's (id, supervisor, reason) — called by the reap before the
/// exit notification is posted, so a supervisor that hears the notification can
/// always still query the reason. Precondition: the big kernel lock is held.
fn recordExitLocked(t: *scheduler.Task) void {
exit_records[exit_record_next] = .{ .id = t.id, .supervisor = t.supervisor, .reason = t.exit_reason, .valid = true };
exit_record_next = (exit_record_next + 1) % exit_record_capacity;
}
/// How dead process `id` ended, for `caller` — the kernel half of the
/// process_exit_reason system call. Returns the ExitReason value, -ESRCH (never
/// lived, still alive, or evicted from the ring), or -EPERM (the caller was not
/// its supervisor — the same authority gate as process_kill).
pub fn exitReasonOf(caller_id: u32, target_id: u32) i64 {
const flags = sync.enter();
defer sync.leave(flags);
for (&exit_records) |*record| {
if (record.valid and record.id == target_id) {
if (record.supervisor != caller_id) return -ipc.EPERM;
return @intFromEnum(record.reason);
}
}
return -ipc.ESRCH;
}
/// The published exit events' subscribers (docs/process-lifecycle.md "Who learns
/// of a death"): stateful services — the VFS's file handles, input's
/// subscriptions — that must release what a dead client held and cannot learn it
/// any other way (a client that simply never calls again looks like silence).
/// Bounded like every kernel table; each entry holds its own endpoint reference.
const exit_subscriber_capacity = 8;
const ExitSubscriber = struct { endpoint: *ipc.Endpoint, owner: u32 };
var exit_subscribers: [exit_subscriber_capacity]?ExitSubscriber = .{null} ** exit_subscriber_capacity;
/// process_subscribe(endpoint): subscribe the caller's endpoint to published exit
/// events. Ungated, like process_enumerate — what is running (and dying) is not a
/// secret between cooperating processes. -ENOSPC when the table is full.
fn systemProcessSubscribe(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
const flags = sync.enter();
defer sync.leave(flags);
for (&exit_subscribers) |*slot| {
if (slot.* == null) {
endpoint.refcount += 1; // the slot's own reference, dropped on unsubscribe-by-death
slot.* = .{ .endpoint = endpoint, .owner = t.id };
return architecture.setSystemCallResult(state, 0);
}
}
failErr(state, ipc.ENOSPC);
}
/// signal_bind(endpoint): nominate where this process's signals arrive — the
/// IRQ-as-IPC pattern a fourth time (docs/process-lifecycle.md). Replacing a
/// binding drops the old reference; signals that pended while unbound are
/// delivered immediately on bind, coalesced into one notification.
fn systemSignalBind(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
const flags = sync.enter();
defer sync.leave(flags);
if (t.signal_endpoint) |raw| ipc.dropRef(@ptrCast(@alignCast(raw)));
endpoint.refcount += 1;
t.signal_endpoint = @ptrCast(endpoint);
if (t.pending_signals != 0) {
ipc.notifyLocked(endpoint, abi.notify_signal_bit | t.pending_signals);
t.pending_signals = 0;
}
architecture.setSystemCallResult(state, 0);
}
/// process_signal(id, signal): post a signal — a one-way, coalescing statement,
/// never a question (docs/process-lifecycle.md). The authority gate is the
/// supervision link, like kill; a process may also signal itself. Unbound
/// targets accumulate the signal in their pending mask.
fn systemProcessSignal(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const id = architecture.systemCallArg(state, 0);
const signal = architecture.systemCallArg(state, 1);
if (id > std.math.maxInt(u32)) return failErr(state, ipc.ESRCH);
if (signal > 31) return failErr(state, ipc.EBADF); // not a Signal bit position
const flags = sync.enter();
defer sync.leave(flags);
const target = scheduler.taskByIdLocked(@intCast(id)) orelse return failErr(state, ipc.ESRCH);
if (target.aspace == 0) return failErr(state, ipc.ESRCH);
if (target.supervisor != t.id and target.id != t.id) return failErr(state, ipc.EPERM);
target.pending_signals |= @as(u32, 1) << @intCast(signal);
if (target.signal_endpoint) |raw| {
const endpoint: *ipc.Endpoint = @ptrCast(@alignCast(raw));
ipc.notifyLocked(endpoint, abi.notify_signal_bit | target.pending_signals);
target.pending_signals = 0;
}
architecture.setSystemCallResult(state, 0);
}
/// The one-shot timers of timer_bind: the missing timed wait. A service arms a
/// deadline and keeps serving; the expiry arrives in the same replyWait as
/// everything else (notify_timer_bit). What stop-sequence escalation, hello
/// deadlines, and restart backoff are built from — and later, `alarm`.
const timer_capacity = 16;
const OneShotTimer = struct { deadline: u64, endpoint: *ipc.Endpoint, owner: u32 };
var one_shot_timers: [timer_capacity]?OneShotTimer = .{null} ** timer_capacity;
/// Sweep expired timers — hung on scheduler.timer_tick_hook, so it runs on every
/// tick with the big kernel lock held, like the sleeper wake it rides beside.
fn timerSweepLocked() void {
const now = architecture.millis();
for (&one_shot_timers) |*slot| {
if (slot.*) |timer| {
if (now >= timer.deadline) {
ipc.notifyLocked(timer.endpoint, abi.notify_timer_bit);
ipc.dropRef(timer.endpoint);
slot.* = null;
}
}
}
}
/// timer_bind(endpoint, ms): arm a one-shot timer. -ENOSPC when the table is full.
fn systemTimerBind(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
const ms = architecture.systemCallArg(state, 1);
const flags = sync.enter();
defer sync.leave(flags);
for (&one_shot_timers) |*slot| {
if (slot.* == null) {
endpoint.refcount += 1;
slot.* = .{ .deadline = architecture.millis() + ms, .endpoint = endpoint, .owner = t.id };
return architecture.setSystemCallResult(state, 0);
}
}
failErr(state, ipc.ENOSPC);
}
fn systemProcessExitReason(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const id = architecture.systemCallArg(state, 0);
if (id > std.math.maxInt(u32)) return failErr(state, ipc.ESRCH);
const r = exitReasonOf(t.id, @intCast(id));
architecture.setSystemCallResult(state, @bitCast(r));
}
/// Resolve `(device_id, resource_index)` to a GSI this process is entitled to bind, or null. /// Resolve `(device_id, resource_index)` to a GSI this process is entitled to bind, or null.
/// The two checks are the whole security story: the device must be *claimed* by the /// The two checks are the whole security story: the device must be *claimed* by the
/// caller, and the resource must be one of that device's `irq` resources as recorded /// caller, and the resource must be one of that device's `irq` resources as recorded
+20 -2
View File
@@ -50,6 +50,17 @@ pub const Task = struct {
// null. Holds its own reference, dropped when the notification is posted. // null. Holds its own reference, dropped when the notification is posted.
// Opaque here for the same reason as `handles` below. // Opaque here for the same reason as `handles` below.
exit_endpoint: ?*anyopaque = null, exit_endpoint: ?*anyopaque = null,
// How this process ended — set by the death paths (exit, fault, kill) just
// before the reap records it for `process_exit_reason`. Meaningless while
// the task lives.
exit_reason: abi.ExitReason = .exited,
// Endpoint this process's signals arrive on (signal_bind), or null — same
// ownership rules as exit_endpoint (holds a reference; opaque here).
signal_endpoint: ?*anyopaque = null,
// Signals posted but not yet delivered: the coalescing pending mask
// (docs/process-lifecycle.md). Bits are abi.Signal values. Signals pend here
// until an endpoint is bound; two pending terminates are one terminate.
pending_signals: u32 = 0,
// Set by process_kill on a task that is running on another core; the kernel // Set by process_kill on a task that is running on another core; the kernel
// finishes the kill at that task's next system call or timer tick. // finishes the kill at that task's next system call or timer tick.
kill_pending: bool = false, kill_pending: bool = false,
@@ -339,8 +350,9 @@ pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, priority: Priority
/// context switch and lock release. /// context switch and lock release.
fn startUserTask() void { fn startUserTask() void {
const t = current(); const t = current();
var buffer: [96]u8 = undefined; // No serial chatter here: this runs on every spawn, unserialized against
architecture.serialWrite(std.fmt.bufPrint(&buffer, "DBG startUserTask ip=0x{x} sp=0x{x} aspace=0x{x} kstack=0x{x}\n", .{ t.user_ip, t.user_sp, t.aspace, t.kstack_top }) catch ""); // user-space writes, and its output used to shear concurrent log lines in
// half — the largest source of corrupted markers in the QEMU scenarios.
architecture.jumpToUser(t.user_ip, t.user_sp); // noreturn architecture.jumpToUser(t.user_ip, t.user_sp); // noreturn
} }
@@ -643,9 +655,15 @@ fn reapKillPendingLocked() void {
/// other critical section, but releases it *without* touching the interrupt flag /// other critical section, but releases it *without* touching the interrupt flag
/// — the handler's `iretq` restores the interrupted context's flags, so /// — the handler's `iretq` restores the interrupted context's flags, so
/// re-enabling here would open a nested-interrupt window before the return. /// re-enabling here would open a nested-interrupt window before the return.
/// Called from the tick with the big kernel lock held — process.zig hangs the
/// one-shot timer sweep here (timer_bind), the same call-up pattern as the
/// teardown hooks below.
pub var timer_tick_hook: ?*const fn () void = null;
pub fn tick() void { pub fn tick() void {
_ = sync.enter(); _ = sync.enter();
wakeExpired(); wakeExpired();
if (timer_tick_hook) |hook| hook();
reapKillPendingLocked(); reapKillPendingLocked();
if (preemption_enabled) schedule(); if (preemption_enabled) schedule();
sync.leaveIsr(); sync.leaveIsr();
+642 -38
View File
@@ -132,6 +132,30 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
processKillTest(boot_information); processKillTest(boot_information);
} else if (eql(case, "supervision")) { } else if (eql(case, "supervision")) {
supervisionTest(boot_information); supervisionTest(boot_information);
} else if (eql(case, "claim-release")) {
claimReleaseTest(boot_information);
} else if (eql(case, "vfs-client-death")) {
vfsClientDeathTest(boot_information);
} else if (eql(case, "signals")) {
signalsTest(boot_information);
} else if (eql(case, "driver-restart")) {
driverRestartTest(boot_information);
} else if (eql(case, "usb-report")) {
usbReportTest(boot_information);
} else if (eql(case, "device-list")) {
deviceListTest(boot_information);
} else if (eql(case, "pci-scan")) {
pciScanTest(boot_information);
} else if (eql(case, "acpi-parse")) {
acpiParseTest(boot_information);
} else if (eql(case, "acpi-report")) {
acpiReportTest(boot_information);
} else if (eql(case, "acpi-ps2")) {
acpiReportTest(boot_information); // same spawn; the harness regex differs
} else if (eql(case, "power-button")) {
acpiReportTest(boot_information); // boot the manager (spawns the acpi service); harness injects the button
} else if (eql(case, "orderly-shutdown")) {
orderlyShutdownTest(boot_information);
} else if (eql(case, "initial-ramdisk")) { } else if (eql(case, "initial-ramdisk")) {
initialRamdiskTest(boot_information); initialRamdiskTest(boot_information);
} else if (eql(case, "vfs")) { } else if (eql(case, "vfs")) {
@@ -187,6 +211,13 @@ fn eql(a: []const u8, b: []const u8) bool {
return std.mem.eql(u8, a, b); return std.mem.eql(u8, a, b);
} }
/// Whether the captured last-write buffer *contains* `needle`. Markers are
/// matched as substrings, not prefixes, so a service's source-path debug prefix
/// (`system/drivers/hpet: ok`) still satisfies a marker like `hpet: ok`.
fn bufferHas(needle: []const u8) bool {
return std.mem.indexOf(u8, process.write_buffer[0..process.write_len], needle) != null;
}
/// Non-destructive checks of the memory map and frame allocator. /// Non-destructive checks of the memory map and frame allocator.
fn smoke(boot_information: *const BootInformation) void { fn smoke(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: smoke\n", .{}); log("DANOS-TEST-BEGIN: smoke\n", .{});
@@ -256,20 +287,56 @@ fn discoveryTest() void {
// M15: every PCI function now carries its own 4 KiB ECAM configuration space as // M15: every PCI function now carries its own 4 KiB ECAM configuration space as
// resource 0 — the window a driver mmio_maps to walk its capability list (MSI etc). // resource 0 — the window a driver mmio_maps to walk its capability list (MSI etc).
// M19.3: the kernel seeds only the bridge; functions arrive by the ring-3
// scan (proven equivalent in pci-scan before the walk retired).
var buffer: [64]device_abi.DeviceDescriptor = undefined; var buffer: [64]device_abi.DeviceDescriptor = undefined;
const n = @min(devices_broker.enumerate(&buffer), buffer.len); const n = @min(devices_broker.enumerate(&buffer), buffer.len);
var pci_functions: u32 = 0; var bridges: u32 = 0;
var pci_config_ok = true; var bridge_shape_ok = false;
for (buffer[0..n]) |d| {
if (d.class != @intFromEnum(device_abi.DeviceClass.pci_host_bridge)) continue;
bridges += 1;
var has_bus_range = false;
var has_io = false;
var memory_windows: u32 = 0;
for (d.resources[0..@intCast(d.resource_count)]) |resource| {
if (resource.kind == @intFromEnum(device_abi.ResourceKind.bus_range)) has_bus_range = true;
if (resource.kind == @intFromEnum(device_abi.ResourceKind.io_port)) has_io = true;
if (resource.kind == @intFromEnum(device_abi.ResourceKind.memory)) memory_windows += 1;
}
// ECAM plus at least one MMIO aperture, the bus range, the I/O window.
if (has_bus_range and has_io and memory_windows >= 2) bridge_shape_ok = true;
}
check("a PCI host bridge was seeded (MCFG)", bridges >= 1);
check("the bridge carries ECAM, apertures, bus range, and the I/O window", bridge_shape_ok);
// M19.0: every PCI memory resource (config slice and BARs alike) must be
// contained in one of its parent bridge's windows — the aperture derivation
// from the memory map is what makes a future user-space device_register of
// these functions pass containment. This is the assert that catches a
// too-coarse hole computation before M19.2 would.
var bars_contained = true;
for (buffer[0..n]) |d| { for (buffer[0..n]) |d| {
if (d.class != @intFromEnum(device_abi.DeviceClass.pci_device)) continue; if (d.class != @intFromEnum(device_abi.DeviceClass.pci_device)) continue;
pci_functions += 1; if (d.parent >= n) {
const has_config = d.resource_count >= 1 and bars_contained = false;
d.resources[0].kind == @intFromEnum(device_abi.ResourceKind.memory) and continue;
d.resources[0].len == abi.page_size;
if (!has_config) pci_config_ok = false;
} }
check("PCI functions were enumerated (MCFG/ECAM)", pci_functions >= 1); const bridge = buffer[@intCast(d.parent)];
check("each PCI function exposes its ECAM config space as resource 0", pci_config_ok); for (d.resources[0..@intCast(d.resource_count)]) |r| {
if (r.kind != @intFromEnum(device_abi.ResourceKind.memory)) continue;
var inside = false;
for (bridge.resources[0..@intCast(bridge.resource_count)]) |w| {
if (w.kind != @intFromEnum(device_abi.ResourceKind.memory)) continue;
if (r.start >= w.start and r.start + r.len <= w.start + w.len) inside = true;
}
if (!inside) {
bars_contained = false;
log(" escaping BAR: 0x{x}+0x{x} on device {d}\n", .{ r.start, r.len, d.id });
}
}
}
check("every PCI BAR lies inside a bridge aperture (M19.0)", bars_contained);
result(); result();
} }
@@ -1080,12 +1147,18 @@ fn ioPortTest() void {
var buffer: [64]device_abi.DeviceDescriptor = undefined; var buffer: [64]device_abi.DeviceDescriptor = undefined;
const n = @min(devices_broker.enumerate(&buffer), buffer.len); const n = @min(devices_broker.enumerate(&buffer), buffer.len);
// Post-M20.3 the PS/2 node is registered at runtime by the ring-3 acpi
// service, so it is absent from this boot snapshot. Exercise the same
// io_port claim/resolve mechanism against the acpi-tables node's broad I/O
// grant — the window that now carries port authority (the service uses it
// for exactly this). The PS/2 status port 0x64 is offset 0x64 within it.
var found_id: ?u64 = null; var found_id: ?u64 = null;
var found_res: u64 = 0; var found_res: u64 = 0;
outer: for (buffer[0..n]) |d| { outer: for (buffer[0..n]) |d| {
if (d.class != @intFromEnum(device_abi.DeviceClass.acpi_tables)) continue;
for (0..d.resource_count) |ri| { for (0..d.resource_count) |ri| {
const r = d.resources[ri]; const r = d.resources[ri];
if (r.kind == @intFromEnum(device_abi.ResourceKind.io_port) and r.start == 0x64 and r.len >= 1) { if (r.kind == @intFromEnum(device_abi.ResourceKind.io_port) and r.start == 0 and r.len > 0x64) {
found_id = d.id; found_id = d.id;
found_res = ri; found_res = ri;
break :outer; break :outer;
@@ -1093,18 +1166,18 @@ fn ioPortTest() void {
} }
} }
const id = found_id orelse { const id = found_id orelse {
check("discovered the PS/2 status port (io_port 0x64)", false); check("discovered the acpi-tables I/O window", false);
result(); result();
return; return;
}; };
check("discovered the PS/2 status port (io_port 0x64)", true); check("discovered the acpi-tables I/O window", true);
const me = scheduler.current(); const me = scheduler.current();
check("claimed the io_port device", devices_broker.claim(id, me.id)); check("claimed the io_port device", devices_broker.claim(id, me.id));
check("an in-range access resolves to port 0x64", process.resolveIoPort(me, id, found_res, 0, 1) == 0x64); check("an in-range access resolves to port 0x64", process.resolveIoPort(me, id, found_res, 0x64, 1) == 0x64);
check("an over-wide access is refused", process.resolveIoPort(me, id, found_res, 0, 2) == null); check("a 4-byte access at the last port is refused", process.resolveIoPort(me, id, found_res, 0xFFFF, 4) == null);
check("an out-of-range offset is refused", process.resolveIoPort(me, id, found_res, 1, 1) == null); check("an out-of-range offset is refused", process.resolveIoPort(me, id, found_res, 0x10000, 1) == null);
check("an unclaimed device id is refused", process.resolveIoPort(me, 0xDEAD_BEEF, found_res, 0, 1) == null); check("an unclaimed device id is refused", process.resolveIoPort(me, 0xDEAD_BEEF, found_res, 0x64, 1) == null);
// The kernel actually issues the `in`. Reaching this line at all proves it didn't // The kernel actually issues the `in`. Reaching this line at all proves it didn't
// fault; a width-1 read must return a single byte. // fault; a width-1 read must return a single byte.
@@ -1207,15 +1280,15 @@ fn userPfTest() void {
/// hand — address space, code page RO+X, stack page RW+NX — because the blob is a /// hand — address space, code page RO+X, stack page RW+NX — because the blob is a
/// raw code fragment, not an ELF `spawnProcess` could load. Returns false if any /// raw code fragment, not an ELF `spawnProcess` could load. Returns false if any
/// allocation fails. /// allocation fails.
fn spawnFaultingProcess() bool { fn spawnFaultingProcess() ?u32 {
const blob = process.pfBlob(); const blob = process.pfBlob();
const flags = sync.enter(); const flags = sync.enter();
defer sync.leave(flags); defer sync.leave(flags);
const aspace = architecture.createAddressSpace() orelse return false; const aspace = architecture.createAddressSpace() orelse return null;
const code_frame = pmm.alloc() orelse { const code_frame = pmm.alloc() orelse {
architecture.destroyAddressSpace(aspace); architecture.destroyAddressSpace(aspace);
return false; return null;
}; };
// Fill through the physmap (the user mapping is read-only); pad with int3 so a // Fill through the physmap (the user mapping is read-only); pad with int3 so a
// stray jump traps instead of sliding. // stray jump traps instead of sliding.
@@ -1226,15 +1299,16 @@ fn spawnFaultingProcess() bool {
const stack_frame = pmm.alloc() orelse { const stack_frame = pmm.alloc() orelse {
architecture.destroyAddressSpace(aspace); // frees code_frame too — it's mapped architecture.destroyAddressSpace(aspace); // frees code_frame too — it's mapped
return false; return null;
}; };
architecture.mapUserPageInto(aspace, process.stack_base_virtual, stack_frame, true, false); // RW + NX architecture.mapUserPageInto(aspace, process.stack_base_virtual, stack_frame, true, false); // RW + NX
if (scheduler.spawnUserLocked(aspace, process.code_virtual, process.stack_base_virtual + abi.page_size, 4, "fault-probe", 0, null) == null) { // Supervised by the calling test task, so exitReasonOf can read the verdict.
const id = scheduler.spawnUserLocked(aspace, process.code_virtual, process.stack_base_virtual + abi.page_size, 4, "fault-probe", scheduler.currentId(), null) orelse {
architecture.destroyAddressSpace(aspace); architecture.destroyAddressSpace(aspace);
return false; return null;
} };
return true; return id;
} }
/// Fault recovery (docs/resilience.md step 2): a scheduled ring-3 process that /// Fault recovery (docs/resilience.md step 2): a scheduled ring-3 process that
@@ -1263,7 +1337,8 @@ fn faultRecoveryTest(boot_information: *const BootInformation) void {
scheduler.setPriority(4); scheduler.setPriority(4);
check("init heartbeat before the fault", process.write_count >= 1); check("init heartbeat before the fault", process.write_count >= 1);
check("faulting process spawned", spawnFaultingProcess()); const probe = spawnFaultingProcess() orelse 0;
check("faulting process spawned", probe != 0);
// The kill: the faulting process #PFs on its first instruction and the kernel // The kill: the faulting process #PFs on its first instruction and the kernel
// reaps it instead of halting. // reaps it instead of halting.
@@ -1272,6 +1347,7 @@ fn faultRecoveryTest(boot_information: *const BootInformation) void {
while (process.fault_kill_count < 1 and architecture.millis() < deadline) scheduler.yield(); while (process.fault_kill_count < 1 and architecture.millis() < deadline) scheduler.yield();
scheduler.setPriority(4); scheduler.setPriority(4);
check("faulting process was killed (not the machine)", process.fault_kill_count == 1); check("faulting process was killed (not the machine)", process.fault_kill_count == 1);
check("the probe's reason reads segmentation_fault", process.exitReasonOf(scheduler.currentId(), probe) == @intFromEnum(abi.ExitReason.segmentation_fault));
// Life after the kill: init must keep beating on the same core. // Life after the kill: init must keep beating on the same core.
const beats_at_kill = process.write_count; const beats_at_kill = process.write_count;
@@ -1311,7 +1387,7 @@ fn initTest(boot_information: *const BootInformation) void {
scheduler.setPriority(4); scheduler.setPriority(4);
const prefix = "init: heartbeat"; const prefix = "init: heartbeat";
const beat_ok = process.write_len >= prefix.len and eql(process.write_buffer[0..prefix.len], prefix); const beat_ok = bufferHas(prefix);
check("init produced repeated heartbeats (>=2)", process.write_count >= 2); check("init produced repeated heartbeats (>=2)", process.write_count >= 2);
check("heartbeat text arrived intact", beat_ok); check("heartbeat text arrived intact", beat_ok);
check("heartbeats came from user mode (CPL 3)", process.write_from_user); check("heartbeats came from user mode (CPL 3)", process.write_from_user);
@@ -1423,6 +1499,12 @@ fn processKillTest(boot_information: *const BootInformation) void {
check("the sleeper's exit notification arrived (length 0)", r == 0); check("the sleeper's exit notification arrived (length 0)", r == 0);
check("its badge carries the exit bit and the child id", badge == abi.notify_badge_bit | abi.notify_exit_bit | sleeper); check("its badge carries the exit bit and the child id", badge == abi.notify_badge_bit | abi.notify_exit_bit | sleeper);
// M17.2: the recorded reason — the notification is the fence, so it is
// already readable, and gated by the same supervisor check as the kill.
check("the sleeper's reason reads killed", process.exitReasonOf(me, sleeper) == @intFromEnum(abi.ExitReason.killed));
check("a non-supervisor may not read the reason (-EPERM)", process.exitReasonOf(me + 12345, sleeper) == -ipcsync.EPERM);
check("an unknown id has no reason (-ESRCH)", process.exitReasonOf(me, 0xFFFF_FF00) == -ipcsync.ESRCH);
const beats_at_kill = process.write_count; const beats_at_kill = process.write_count;
scheduler.sleep(1500); // more than one heartbeat period scheduler.sleep(1500); // more than one heartbeat period
check("the heartbeat stopped with the kill", process.write_count == beats_at_kill); check("the heartbeat stopped with the kill", process.write_count == beats_at_kill);
@@ -1443,6 +1525,21 @@ fn processKillTest(boot_information: *const BootInformation) void {
check("the spinner's exit notification arrived (length 0)", r == 0); check("the spinner's exit notification arrived (length 0)", r == 0);
check("its badge carries the exit bit and the child id", badge == abi.notify_badge_bit | abi.notify_exit_bit | spinner); check("its badge carries the exit bit and the child id", badge == abi.notify_badge_bit | abi.notify_exit_bit | spinner);
// M17.2: a child that ends on its own must read exited, not killed —
// args-echo with arguments echoes once and returns from main.
var clean: u32 = 0;
i = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "args-echo")) continue;
clean = process.spawnProcessSupervised(item.blob, 4, &.{ "args-echo", "clean-exit" }, me, endpoint) catch 0;
break;
}
check("args-echo spawned as the clean-exit child", clean != 0);
r = ipcsync.replyWait(endpoint, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap);
check("the clean child's exit notification arrived", badge == abi.notify_badge_bit | abi.notify_exit_bit | clean);
check("the clean child's reason reads exited", process.exitReasonOf(me, clean) == @intFromEnum(abi.ExitReason.exited));
var table: [32]abi.ProcessDescriptor = undefined; var table: [32]abi.ProcessDescriptor = undefined;
const total = scheduler.enumerate(&table); const total = scheduler.enumerate(&table);
var still_listed = false; var still_listed = false;
@@ -1454,6 +1551,484 @@ fn processKillTest(boot_information: *const BootInformation) void {
result(); result();
} }
/// M17.1: a dead process's device claims are released by the reap, so a restarted
/// driver can claim its hardware again (docs/process-lifecycle.md iron rule 1).
/// First the broker release in isolation — two owners, one released, the other's
/// claim must survive. Then the death-path wiring with a real child: the claim is
/// made on the child's behalf (the broker is kernel-callable), the child is
/// killed, and once the exit notification arrives — posted last, after release —
/// the device must be unclaimed and claimable again.
fn claimReleaseTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: claim-release\n", .{});
var buffer: [2]device_abi.DeviceDescriptor = undefined;
const total = devices_broker.enumerate(&buffer);
check("the device tree is seeded (>= 2 devices)", total >= 2);
if (total < 2) {
result();
return;
}
// The broker release in isolation.
check("device 0 claimed by owner 111", devices_broker.claim(0, 111));
check("device 1 claimed by owner 222", devices_broker.claim(1, 222));
devices_broker.releaseAllOwnedBy(111);
check("owner 111's claim is released", devices_broker.ownerOf(0) == null);
check("owner 222's claim survives", (devices_broker.ownerOf(1) orelse 0) == 222);
devices_broker.releaseAllOwnedBy(222);
check("cleanup released owner 222", devices_broker.ownerOf(1) == null);
// The death-path wiring: a real process dies holding a claim.
check("bootloader handed over /system/services/init", boot_information.init_len != 0);
if (boot_information.init_len == 0) {
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len];
const me = scheduler.currentId();
const endpoint = ipcsync.createIpcEndpoint() orelse {
check("exit endpoint allocated", false);
result();
return;
};
const child = process.spawnProcessSupervised(image, 4, &.{"/system/services/init"}, me, endpoint) catch 0;
check("supervised child spawned", child != 0);
check("device 0 claimed on the child's behalf", devices_broker.claim(0, child));
check("the kill is accepted", process.killProcess(me, child) == 0);
var badge: u64 = 0;
var received_cap: u64 = 0;
_ = ipcsync.replyWait(endpoint, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap);
check("the exit notification arrived", badge == abi.notify_badge_bit | abi.notify_exit_bit | child);
check("death released the child's claim", devices_broker.ownerOf(0) == null);
check("the device is claimable again", devices_broker.claim(0, me));
devices_broker.releaseAllOwnedBy(me);
result();
}
/// M17.3: the published exit events, proven by their first subscriber. The VFS
/// subscribes at startup; a client opens a file and parks holding the handle;
/// the kill posts the exit event to the VFS's endpoint; the VFS releases the
/// dead client's handle and says so — the service-side mirror of iron rule 1
/// (a service must never depend on clients cleaning up after themselves).
fn vfsClientDeathTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: vfs-client-death\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.write_count = 0;
check("vfs spawned", spawnNamed(rd, "vfs"));
const me = scheduler.currentId();
const endpoint = ipcsync.createIpcEndpoint() orelse {
check("exit endpoint allocated", false);
result();
return;
};
var client: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "vfs-test")) continue;
client = process.spawnProcessSupervised(item.blob, 4, &.{ "vfs-test", "park" }, me, endpoint) catch 0;
break;
}
check("parked client spawned (supervised)", client != 0);
// Its heartbeat is the fence: once it beats, the handle is open.
const parked = "vfstest: parked";
scheduler.setPriority(1);
var deadline = architecture.millis() + 10000;
while (architecture.millis() < deadline) {
if (bufferHas(parked)) break;
scheduler.yield();
}
scheduler.setPriority(4);
check("client parked holding an open handle", bufferHas(parked));
check("the kill is accepted", process.killProcess(me, client) == 0);
var badge: u64 = 0;
var received_cap: u64 = 0;
_ = ipcsync.replyWait(endpoint, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap);
check("the exit notification arrived", badge == abi.notify_badge_bit | abi.notify_exit_bit | client);
// The VFS heard the same published event; its release line is the proof.
const released = "vfs: released 1 handle(s) for dead client";
scheduler.setPriority(1);
deadline = architecture.millis() + 10000;
while (architecture.millis() < deadline) {
if (bufferHas(released)) break;
scheduler.yield();
}
scheduler.setPriority(4);
check("the VFS released the dead client's handle", bufferHas(released));
result();
}
/// M17.4 from ring 3: process-test's signal-run role drives the whole lifecycle
/// surface — the zero-length ping (answered by the harness), signals as
/// statements (reload logged, terminate = clean exit), the one-shot timer, and
/// both endings of the stop sequence (polite -> exited, deaf -> killed at the
/// deadline). Its "process-test: signals ok" is the pass marker.
fn signalsTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: signals\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(image); // the parent system_spawns its children by name
process.write_count = 0;
var runner: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "process-test")) continue;
runner = process.spawnProcessSupervised(item.blob, 4, &.{ "process-test", "signal-run" }, scheduler.currentId(), null) catch 0;
break;
}
check("signal-run parent spawned", runner != 0);
const pass_marker = "process-test: signals ok";
const fail_marker = "process-test: FAIL";
scheduler.setPriority(1);
const deadline = architecture.millis() + 15000;
var saw_pass = false;
var saw_fail = false;
while (architecture.millis() < deadline and !saw_pass and !saw_fail) {
if (bufferHas(pass_marker)) saw_pass = true;
if (bufferHas(fail_marker)) saw_fail = true;
scheduler.yield();
}
scheduler.setPriority(4);
check("the signal-run parent reported ok", saw_pass and !saw_fail);
result();
}
/// M18.1: the device manager's restart machinery, end to end. In test-restart
/// mode the manager also supervises crash-test: a fixture that claims device 0,
/// hellos, and faults. The scenario asserts three markers in order — the real
/// xHCI driver hellos clean and stays; crash-test is restarted with backoff
/// (each respawn re-claiming the device the dead instance held, M17.1 through
/// the manager's path); the crash loop caps and the manager gives up.
fn driverRestartTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: driver-restart\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(image); // the manager system_spawns drivers by name
process.write_count = 0;
var manager: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{ "device-manager", "test-restart" }, scheduler.currentId(), null) catch 0;
break;
}
check("device-manager spawned in test-restart mode", manager != 0);
// The assertions live in the harness: its expect regex requires, in order,
// the xHCI hello ack, a crash-test restart, and the crash-loop cap — read
// from the whole serial capture, immune to the transient-line races a
// write_buffer poll would have here (many processes log concurrently).
result();
}
/// M18.2: bus tree reports, end to end. The manager (test-usb-restart mode)
/// spawns the xHCI driver; the driver maps its BAR, scans the root-hub ports,
/// and reports the two QEMU devices; the manager mirrors them, kills the
/// reporter (the test trigger), prunes both children, restarts the driver with
/// backoff, and the respawned instance re-claims, re-scans, and re-reports.
/// The harness's ordered expect regex is the assertion; this test only
/// orchestrates the spawn.
fn usbReportTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: usb-report\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(image);
var manager: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{ "device-manager", "test-usb-restart" }, scheduler.currentId(), null) catch 0;
break;
}
check("device-manager spawned in test-usb-restart mode", manager != 0);
result();
}
/// M18.3: the application surface. device-list enumerates the manager's tree
/// over IPC, subscribes with its endpoint as a capability, and prints every
/// published event; the manager's delayed test-kill of the reporter produces a
/// removed/added storm the subscriber must observe. The harness's ordered
/// expect regex is the assertion.
fn deviceListTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: device-list\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(image);
var manager: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{ "device-manager", "test-usb-restart" }, scheduler.currentId(), null) catch 0;
break;
}
check("device-manager spawned in test-usb-restart mode", manager != 0);
check("device-list spawned", spawnNamed(rd, "device-list"));
result();
}
/// M19.1: the ring-3 PCI scan agrees with the kernel's. The manager spawns
/// pci-bus for the host bridge; the driver walks the same ECAM window through
/// its mmio_map grant and must find exactly the functions the kernel's own
/// enumeration recorded — the equivalence that licenses retiring the kernel
/// walk in M19.3.
fn pciScanTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: pci-scan\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
// Post-flip (M19.3) ground truth: the kernel no longer enumerates PCI
// functions, so equivalence inverts — the broker's function count after
// the scan must equal what the driver itself reported finding.
var buffer: [64]device_abi.DeviceDescriptor = undefined;
const n = @min(devices_broker.enumerate(&buffer), buffer.len);
var boot_pci: u32 = 0;
for (buffer[0..n]) |d| {
if (d.class == @intFromEnum(device_abi.DeviceClass.pci_device)) boot_pci += 1;
}
check("the kernel seeded no PCI functions (the walk retired)", boot_pci == 0);
process.setInitialRamdisk(image);
process.write_count = 0;
var manager: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{ "device-manager", "test-pci-restart" }, scheduler.currentId(), null) catch 0;
break;
}
check("device-manager spawned (test-pci-restart mode)", manager != 0);
// First scan: wait for the driver's count line and parse the number.
const count_prefix = "pci-bus: ";
const count_suffix = " functions found";
var reported: u32 = 0;
scheduler.setPriority(1);
var deadline = architecture.millis() + 15000;
while (architecture.millis() < deadline and reported == 0) {
const line = process.write_buffer[0..process.write_len];
if (std.mem.indexOf(u8, line, count_prefix)) |start| {
if (std.mem.indexOf(u8, line, count_suffix)) |digits_end| {
reported = std.fmt.parseInt(u32, line[start + count_prefix.len .. digits_end], 10) catch 0;
}
}
scheduler.yield();
}
scheduler.setPriority(4);
check("the ring-3 scan reported a function count", reported >= 1);
// Every reported function was registered: the broker holds exactly them.
var registered: [64]device_abi.DeviceDescriptor = undefined;
const r = @min(devices_broker.enumerate(&registered), registered.len);
var registered_pci: u32 = 0;
for (registered[0..r]) |d| {
if (d.class == @intFromEnum(device_abi.DeviceClass.pci_device)) registered_pci += 1;
}
check("the broker holds exactly the reported functions", registered_pci == reported);
const kernel_count = reported; // the no-duplicate check below reuses it
// The restart drill: the manager kills pci-bus after its reports; the
// respawn re-claims, re-scans, and re-registers.
const restart_marker = "device-manager: restarting pci-bus";
scheduler.setPriority(1);
deadline = architecture.millis() + 15000;
var restarted = false;
while (architecture.millis() < deadline and !restarted) {
if (bufferHas(restart_marker)) restarted = true;
scheduler.yield();
}
scheduler.setPriority(4);
check("the manager restarted pci-bus", restarted);
var marker_buffer: [48]u8 = undefined;
const marker = std.fmt.bufPrint(&marker_buffer, "pci-bus: {d} functions found", .{reported}) catch "";
scheduler.setPriority(1);
deadline = architecture.millis() + 15000;
var seen = false;
while (architecture.millis() < deadline and !seen) {
if (bufferHas(marker)) seen = true;
scheduler.yield();
}
scheduler.setPriority(4);
check("the respawned scan reported the same count", seen);
// No duplicates: the registrations deduped against the kernel's own nodes
// on the first pass, and against themselves on the second.
var after: [64]device_abi.DeviceDescriptor = undefined;
const m = @min(devices_broker.enumerate(&after), after.len);
var after_count: u32 = 0;
for (after[0..m]) |d| {
if (d.class == @intFromEnum(device_abi.DeviceClass.pci_device)) after_count += 1;
}
check("no duplicate PCI nodes after register + restart + re-register", after_count == kernel_count);
result();
}
/// M21.3 capstone: orderly shutdown. Boot init with the initial-ramdisk
/// published, so init spawns the full service tree (vfs, input, device-manager
/// -> discovery/acpi); the harness injects a real power-button event via QMP;
/// the acpi service publishes it; init runs the stop sequence over its children
/// and asks the power service for S5; the machine powers off (QEMU exits). The
/// kernel test only spawns init — the ordered chain is the harness assertion.
fn orderlyShutdownTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: orderly-shutdown\n", .{});
if (boot_information.init_len == 0 or boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over init and the initial_ramdisk", false);
result();
return;
}
const ramdisk = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
process.setInitialRamdisk(ramdisk);
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len];
const spawned = if (process.spawnProcess(image, 4, &.{"/system/services/init"})) true else |_| false;
check("init spawned as PID root of user space", spawned);
result();
}
/// M20.2: the acpi service registers + reports its _HID devices. Boot normally
/// (the manager spawns discovery); the harness's expect regex requires the two
/// PS/2 nodes among the service's report lines, each with its _CRS resources —
/// the ring-3 _CRS/_STA evaluation working end to end. The kernel test only
/// starts the manager.
fn acpiReportTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: acpi-report\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(image);
var spawned = false;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue;
_ = process.spawnProcessSupervised(item.blob, 4, &.{"device-manager"}, scheduler.currentId(), null) catch 0;
spawned = true;
break;
}
check("device-manager spawned", spawned);
result();
}
/// M20.1: the ring-3 AML parse agrees with the kernel's. The manager spawns
/// the discovery service (the acpi build variant); it claims the acpi-tables
/// node, maps the blobs, parses them, and logs its Device count — which must
/// equal what the kernel's own parse produced (the equivalence that licenses
/// retiring the kernel's device build in M20.3).
fn acpiParseTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: acpi-parse\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
// The kernel's own count, from the namespace it already built for \_S5.
const kernel_devices = platform.amlDeviceCount();
check("the kernel namespace has devices to compare against", kernel_devices >= 1);
// Spawn the discovery service directly with that count as argv: it parses
// the same blobs in ring 3 and self-verifies, printing "acpi-parse: ok" iff
// the counts match. The harness's expect regex is that marker — deterministic,
// no racing the shared serial buffer.
process.setInitialRamdisk(image);
var count_text: [16]u8 = undefined;
const count_arg = std.fmt.bufPrint(&count_text, "{d}", .{kernel_devices}) catch "0";
var spawned = false;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "discovery")) continue;
_ = process.spawnProcessSupervised(item.blob, 4, &.{ "discovery", count_arg }, scheduler.currentId(), null) catch 0;
spawned = true;
break;
}
check("discovery service spawned", spawned);
result();
}
/// The whole user-side surface at once: spawn process-test's supervisor role, /// The whole user-side surface at once: spawn process-test's supervisor role,
/// which — entirely from ring 3 — creates an exit endpoint, spawns its two /// which — entirely from ring 3 — creates an exit endpoint, spawns its two
/// children supervised, sees them in process_enumerate, kills them (one blocked, /// children supervised, sees them in process_enumerate, kills them (one blocked,
@@ -1490,12 +2065,12 @@ fn supervisionTest(boot_information: *const BootInformation) void {
scheduler.setPriority(1); scheduler.setPriority(1);
const deadline = architecture.millis() + 10000; const deadline = architecture.millis() + 10000;
while (architecture.millis() < deadline) { while (architecture.millis() < deadline) {
if (process.write_len >= marker.len and eql(process.write_buffer[0..marker.len], marker)) break; if (bufferHas(marker)) break;
scheduler.yield(); scheduler.yield();
} }
scheduler.setPriority(4); scheduler.setPriority(4);
const ok = process.write_len >= marker.len and eql(process.write_buffer[0..marker.len], marker); const ok = bufferHas(marker);
if (!ok and process.write_len > 0) log("DANOS-SUPERVISION: got \"{s}\"\n", .{process.write_buffer[0..process.write_len]}); if (!ok and process.write_len > 0) log("DANOS-SUPERVISION: got \"{s}\"\n", .{process.write_buffer[0..process.write_len]});
check("the supervisor completed every step (spawn/list/kill/notify)", ok); check("the supervisor completed every step (spawn/list/kill/notify)", ok);
check("it ran in user mode (CPL 3)", process.write_from_user); check("it ran in user mode (CPL 3)", process.write_from_user);
@@ -1575,12 +2150,12 @@ fn vfsTest(boot_information: *const BootInformation) void {
scheduler.setPriority(1); scheduler.setPriority(1);
const deadline = architecture.millis() + 10000; const deadline = architecture.millis() + 10000;
while (architecture.millis() < deadline) { while (architecture.millis() < deadline) {
if (process.write_len >= prefix.len and eql(process.write_buffer[0..prefix.len], prefix) and process.write_count >= 2) break; if (bufferHas(prefix) and process.write_count >= 2) break;
scheduler.yield(); scheduler.yield();
} }
scheduler.setPriority(4); scheduler.setPriority(4);
const ok = process.write_len >= prefix.len and eql(process.write_buffer[0..prefix.len], prefix); const ok = bufferHas(prefix);
check("client completed the VFS round trip (open/write/read matched)", ok); check("client completed the VFS round trip (open/write/read matched)", ok);
check("the round trip ran repeatedly (server stays up)", process.write_count >= 2); check("the round trip ran repeatedly (server stays up)", process.write_count >= 2);
check("client syscalls came from user mode (CPL 3)", process.write_from_user); check("client syscalls came from user mode (CPL 3)", process.write_from_user);
@@ -1620,12 +2195,12 @@ fn inputTest(boot_information: *const BootInformation) void {
scheduler.setPriority(1); scheduler.setPriority(1);
const deadline = architecture.millis() + 12000; const deadline = architecture.millis() + 12000;
while (architecture.millis() < deadline) { while (architecture.millis() < deadline) {
if (process.write_len >= prefix.len and eql(process.write_buffer[0..prefix.len], prefix) and process.write_count >= 2) break; if (bufferHas(prefix) and process.write_count >= 2) break;
scheduler.yield(); scheduler.yield();
} }
scheduler.setPriority(4); scheduler.setPriority(4);
const ok = process.write_len >= prefix.len and eql(process.write_buffer[0..prefix.len], prefix); const ok = bufferHas(prefix);
check("a subscriber received a broadcast key event over IPC (source -> service -> subscriber)", ok); check("a subscriber received a broadcast key event over IPC (source -> service -> subscriber)", ok);
check("events kept flowing (service + async send stay up)", process.write_count >= 2); check("events kept flowing (service + async send stay up)", process.write_count >= 2);
check("client syscalls came from user mode (CPL 3)", process.write_from_user); check("client syscalls came from user mode (CPL 3)", process.write_from_user);
@@ -1717,12 +2292,12 @@ fn hpetTest(boot_information: *const BootInformation) void {
scheduler.setPriority(1); scheduler.setPriority(1);
const deadline = architecture.millis() + 10000; const deadline = architecture.millis() + 10000;
while (architecture.millis() < deadline) { while (architecture.millis() < deadline) {
if (process.write_len >= prefix.len and eql(process.write_buffer[0..prefix.len], prefix) and process.write_count >= 2) break; if (bufferHas(prefix) and process.write_count >= 2) break;
scheduler.yield(); scheduler.yield();
} }
scheduler.setPriority(4); scheduler.setPriority(4);
const ok = process.write_len >= prefix.len and eql(process.write_buffer[0..prefix.len], prefix); const ok = bufferHas(prefix);
check("user driver mapped HPET MMIO and was woken by its interrupt", ok); check("user driver mapped HPET MMIO and was woken by its interrupt", ok);
check("driver syscalls came from user mode (CPL 3)", process.write_from_user); check("driver syscalls came from user mode (CPL 3)", process.write_from_user);
check("kernel routed and re-armed the HPET's line at the I/O APIC", hpetRouteOk()); check("kernel routed and re-armed the HPET's line at the I/O APIC", hpetRouteOk());
@@ -1774,6 +2349,35 @@ fn hpetGsi() ?u32 {
/// land in the device table with the containment invariant intact. /// land in the device table with the containment invariant intact.
fn busTest(boot_information: *const BootInformation) void { fn busTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: bus\n", .{}); log("DANOS-TEST-BEGIN: bus\n", .{});
// M19.0: device_register is idempotent on exact match — a restarted
// registering bus must not duplicate its children. Driven directly against
// the broker: claim an unclaimed node, register the same (class, hid,
// resourceless) child twice, expect one id and one table entry.
{
const me = scheduler.currentId();
var probe: [1]device_abi.DeviceDescriptor = undefined;
const total = devices_broker.enumerate(&probe);
check("device tree is seeded for the idempotence check", total >= 1);
if (devices_broker.ownerOf(0) == null) {
check("claimed device 0 for the idempotence check", devices_broker.claim(0, me));
var child = std.mem.zeroes(device_abi.DeviceDescriptor);
child.class = @intFromEnum(device_abi.DeviceClass.unknown);
child.pci_class = device_abi.no_pci_class;
child.hid_len = 4;
child.hid[0..4].* = "idem".*;
const first = devices_broker.register(0, me, &child) catch 0;
check("first register succeeded", first != 0);
const before = devices_broker.enumerate(&probe);
const second = devices_broker.register(0, me, &child) catch 0;
check("re-register returned the same id", second == first);
check("re-register grew nothing", devices_broker.enumerate(&probe) == before);
devices_broker.releaseAllOwnedBy(me);
} else {
check("device 0 unexpectedly claimed before the idempotence check", false);
}
}
if (boot_information.initial_ramdisk_len == 0) { if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false); check("bootloader handed over an initial_ramdisk", false);
result(); result();
@@ -1794,12 +2398,12 @@ fn busTest(boot_information: *const BootInformation) void {
scheduler.setPriority(1); scheduler.setPriority(1);
const deadline = architecture.millis() + 10000; const deadline = architecture.millis() + 10000;
while (architecture.millis() < deadline) { while (architecture.millis() < deadline) {
if (process.write_len >= prefix.len and eql(process.write_buffer[0..prefix.len], prefix)) break; if (bufferHas(prefix)) break;
scheduler.yield(); scheduler.yield();
} }
scheduler.setPriority(4); scheduler.setPriority(4);
const ok = process.write_len >= prefix.len and eql(process.write_buffer[0..prefix.len], prefix); const ok = bufferHas(prefix);
check("bus driver published children and the kernel refused an out-of-window one", ok); check("bus driver published children and the kernel refused an out-of-window one", ok);
check("driver syscalls came from user mode (CPL 3)", process.write_from_user); check("driver syscalls came from user mode (CPL 3)", process.write_from_user);
check("every registered child is contained in its parent", childrenContained()); check("every registered child is contained in its parent", childrenContained());
@@ -1843,12 +2447,12 @@ fn deviceManagerTest(boot_information: *const BootInformation) void {
scheduler.setPriority(1); scheduler.setPriority(1);
const deadline = architecture.millis() + 10000; const deadline = architecture.millis() + 10000;
while (architecture.millis() < deadline) { while (architecture.millis() < deadline) {
if (process.write_len >= prefix.len and eql(process.write_buffer[0..prefix.len], prefix)) break; if (bufferHas(prefix)) break;
scheduler.yield(); scheduler.yield();
} }
scheduler.setPriority(4); scheduler.setPriority(4);
const ok = process.write_len >= prefix.len and eql(process.write_buffer[0..prefix.len], prefix); const ok = bufferHas(prefix);
check("device manager matched the timer and system_spawn'd hpet, which came up", ok); check("device manager matched the timer and system_spawn'd hpet, which came up", ok);
check("its syscalls came from user mode (CPL 3)", process.write_from_user); check("its syscalls came from user mode (CPL 3)", process.write_from_user);
result(); result();
+5 -2
View File
@@ -16,8 +16,11 @@
pub const maximum_cpus = 128; pub const maximum_cpus = 128;
/// Maximum tasks (kernel threads) alive at once — the static task-table size. Each /// Maximum tasks (kernel threads) alive at once — the static task-table size. Each
/// online core consumes one slot for its idle task, plus task 0 on the BSP. /// online core consumes one slot for its idle task, plus task 0 on the BSP. Sized
pub const maximum_tasks = 16; /// for the initial-ramdisk sweep (15 bundled binaries spawned at once) plus the
/// device manager's supervised children with room to grow — at 16 the sweep
/// started failing spawns once the bundle passed a dozen binaries.
pub const maximum_tasks = 32;
/// Each task's kernel stack (also each AP's bring-up stack), in bytes. /// Each task's kernel stack (also each AP's bring-up stack), in bytes.
pub const kernel_stack_size = 16 * 1024; pub const kernel_stack_size = 16 * 1024;
+699
View File
@@ -0,0 +1,699 @@
//! /system/services/acpi — the ACPI discovery service: the x86 firmware
//! interpreter, moved out of ring 0 (docs/m19-m20-plan.md, M20). Claims the
//! `acpi-tables` node the kernel publishes (the AML blobs, the broad io_port
//! grant, a broad irq window, the SCI), and runs the **shared AML module** in
//! ring 3 — the same parser and interpreter the kernel uses.
//!
//! It also owns the **event side** (M21): it registers the domain-named `.power`
//! service, binds the SCI (System Control Interrupt), and on a power-button
//! fixed event publishes `power_button` to subscribers — and on init's request
//! writes S5 to power the machine off. The device discovery (M20) and the event
//! handling both run in one `runtime.service.run` loop.
const std = @import("std");
const runtime = @import("runtime");
const aml = @import("aml");
const acpi_ids = @import("acpi-ids");
const device = runtime.device;
const protocol = runtime.device_manager_protocol;
const power = runtime.power_protocol;
/// AML opcode/prefix bytes by name (`zero_opcode`, `byte_prefix`, …) — so the `_HID`
/// integer decode names the opcodes instead of bare 0x0A/0x0B/… (docs/coding-standards.md).
const opcodes = aml.opcodes;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
// The claimed acpi-tables node and the resource index of its broad io_port
// window — the Hal routes every port access through this one claim.
var node_id: u64 = 0;
var io_resource_index: u64 = 0;
// The SCI's irq resource index on the node (the len-1 irq, distinct from the
// broad [0,256) window), for irqBind / irqAck.
var sci_resource_index: u64 = 0;
var has_sci = false;
// PM1 event/control and GPE register ports, read from the FADT copy the kernel
// publishes on the node (M21). Port 0 means absent.
var pm1a_evt: u16 = 0;
var pm1b_evt: u16 = 0;
var pm1_evt_len: u8 = 0;
var pm1a_cnt: u16 = 0;
var pm1b_cnt: u16 = 0;
var gpe0_blk: u16 = 0;
var gpe0_len: u8 = 0;
var gpe1_blk: u16 = 0;
var gpe1_len: u8 = 0;
var smi_cmd: u16 = 0;
var acpi_enable_value: u8 = 0;
var s5_slp_typ_a: u8 = 0;
var s5_slp_typ_b: u8 = 0;
var s5_valid = false;
// PM1 event-register bits (ACPI): PWRBTN in the status/enable word is bit 8;
// the control word's SCI_EN is bit 0; SLP_EN is bit 13.
const pwrbtn_bit: u16 = 1 << 8;
const sci_en_bit: u32 = 1 << 0;
const slp_en: u32 = 1 << 13;
// The `.power` subscribers: endpoints handed over as capabilities, each
// receiving events as buffered messages. Dropped on a failed send. The
// subscriber's task id is kept too — a shutdown request is honored only from a
// subscriber (init subscribes; a stray process does not), the soft gate that
// stands in for "only the system supervisor may power off" without hardcoding
// a pid the kernel's idle tasks would have taken.
const maximum_subscribers = 8;
var subscribers: [maximum_subscribers]?runtime.ipc.Handle = .{null} ** maximum_subscribers;
var subscriber_tasks: [maximum_subscribers]u32 = .{0} ** maximum_subscribers;
// Pass-1 registration record (see main): what pass 2 reports.
const Registered = struct { hid: [8]u8 = .{0} ** 8, hid_len: usize = 0, device_id: u64 = 0, resource_count: u64 = 0 };
var registered: [64]Registered = undefined;
var registered_count: usize = 0;
// A scratch page returned for SystemMemory OperationRegion maps: the service
// cannot map arbitrary physical memory from ring 3, so such regions are
// unsupported and degrade to harmless zeros rather than faulting. The M20.2
// targets (ps2, the legacy devices) use SystemIO and static templates.
var mmio_scratch: [4096]u8 align(4096) = .{0} ** 4096;
fn halMapMmio(physical: u64, len: u64, writable: bool) u64 {
_ = physical;
_ = len;
_ = writable;
return @intFromPtr(&mmio_scratch);
}
fn halPioRead(width: u8, port: u16) u32 {
return device.ioRead(node_id, io_resource_index, port, width) orelse 0;
}
fn halPioWrite(width: u8, port: u16, value: u32) void {
_ = device.ioWrite(node_id, io_resource_index, port, width, value);
}
fn findTablesNode(buffer: []device.DeviceDescriptor) ?device.DeviceDescriptor {
const total = device.enumerate(buffer);
for (buffer[0..@min(total, buffer.len)]) |d| {
if (d.class == @intFromEnum(device.DeviceClass.acpi_tables)) return d;
}
return null;
}
pub fn main(init: runtime.process.Init) void {
// When the acpi-parse scenario spawns this directly, argv[1] is the kernel's
// own device count to self-verify against — deterministic, no log-scraping.
const expected: ?usize = if (init.arguments.get(1)) |a| (std.fmt.parseInt(usize, a, 10) catch null) else null;
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("/system/services/acpi: out of memory\n");
return;
};
const node = findTablesNode(buffer) orelse {
_ = runtime.system.write("/system/services/acpi: no acpi-tables node to claim\n");
return;
};
node_id = node.id;
if (!device.claim(node_id)) {
_ = runtime.system.write("/system/services/acpi: unable to claim acpi-tables\n");
return;
}
// Map the node's resources: the AML blobs (bytecode), the FADT (intact
// "FACP" header — decision 3), the io_port grant, and the SCI irq.
var blocks: [8][]const u8 = undefined;
var block_count: usize = 0;
var found_io = false;
var fadt: ?[]const u8 = null;
for (node.resources[0..@intCast(node.resource_count)], 0..) |resource, index| {
if (resource.kind == @intFromEnum(device.ResourceKind.io_port) and !found_io) {
io_resource_index = index;
found_io = true;
continue;
}
if (resource.kind == @intFromEnum(device.ResourceKind.irq) and resource.len == 1) {
sci_resource_index = index;
has_sci = true;
continue;
}
if (resource.kind != @intFromEnum(device.ResourceKind.memory)) continue;
const base = device.mmioMap(node_id, index) orelse continue;
const pointer: [*]const u8 = @ptrFromInt(base);
const bytes = pointer[0..@intCast(resource.len)];
if (bytes.len >= 4 and std.mem.eql(u8, bytes[0..4], "FACP")) {
fadt = bytes;
continue;
}
if (block_count == blocks.len) continue;
blocks[block_count] = bytes;
block_count += 1;
}
if (block_count == 0) {
_ = runtime.system.write("/system/services/acpi: no AML blobs on the node\n");
return;
}
const result = aml.parse(runtime.allocator(), blocks[0..block_count]) catch {
_ = runtime.system.write("/system/services/acpi: AML parse failed\n");
return;
};
var namespace = result.namespace;
const devices = aml.deviceCount(&namespace);
writeLine("/system/services/acpi: parsed {d} AML blob(s), {d} namespace devices\n", .{ block_count, devices });
if (expected) |want| {
if (devices == want) {
_ = runtime.system.write("acpi-parse: ok\n");
} else {
writeLine("acpi-parse: mismatch (ring-3 {d} vs kernel {d})\n", .{ devices, want });
}
// Self-verify mode is standalone (no manager); stop before reporting.
while (true) runtime.system.sleep(1000);
}
// Register + report the present _HID devices (M20), then set up the power
// event side (M21), then serve — all in one harness loop. The interpreter
// and namespace outlive this frame (static), so the harness callbacks can
// reach them.
interpreter_arena = std.heap.ArenaAllocator.init(runtime.allocator());
persistent_namespace = namespace;
global_interpreter = aml.Interpreter.init(&persistent_namespace, .{
.mapMmio = halMapMmio,
.pioRead = halPioRead,
.pioWrite = halPioWrite,
}, interpreter_arena.allocator());
readFadt(fadt);
s5_valid = readSleepS5(&persistent_namespace);
runtime.service.run(power.message_maximum, .{
.service = .power,
.init = onInit,
.on_message = onMessage,
.on_notification = onNotification,
});
}
// Static so the harness callbacks (which run after main's stack frame is gone)
// can reach the namespace and interpreter.
var persistent_namespace: aml.Namespace = undefined;
var global_interpreter: aml.Interpreter = undefined;
var interpreter_arena: std.heap.ArenaAllocator = undefined;
/// Startup under the harness: register + report the discovered devices to the
/// manager (M20), then enable ACPI mode and arm the power button (M21).
fn onInit(endpoint: runtime.ipc.Handle) bool {
registered_count = 0;
walkDevices(persistent_namespace.root, &global_interpreter);
const manager = runtime.ipc.lookup(.device_manager);
var i: usize = 0;
while (i < registered_count) : (i += 1) {
const entry = registered[i];
const hid = entry.hid[0..entry.hid_len];
const desc = acpi_ids.description(hid);
if (desc.len != 0)
writeLine("/system/services/acpi: reported {s} (device {d}, {d} resources) — {s}\n", .{ hid, entry.device_id, entry.resource_count, desc })
else
writeLine("/system/services/acpi: reported {s} (device {d}, {d} resources)\n", .{ hid, entry.device_id, entry.resource_count });
if (manager) |h| {
var report = protocol.ChildAdded{ .parent = node_id, .bus_address = entry.device_id, .identity = 0, .device_id = entry.device_id };
@memcpy(report.hid[0..entry.hid_len], entry.hid[0..entry.hid_len]);
var reply: [protocol.message_maximum]u8 = undefined;
_ = runtime.ipc.call(h, std.mem.asBytes(&report), &reply) catch {};
}
}
writeLine("/system/services/acpi: reported {d} device(s) to the manager\n", .{registered_count});
armPowerButton(endpoint);
return true;
}
// --- power event side (M21) ---------------------------------------------------
/// Read the PM1 event/control and GPE register ports plus the SMI enable pair
/// from the FADT copy on the node. Offsets are from the FADT table start (the
/// SDT header is the first 36 bytes). Prefers the 32-bit port fields; QEMU's
/// FADT populates them.
fn readFadt(fadt: ?[]const u8) void {
const f = fadt orelse {
_ = runtime.system.write("acpi: no FADT on the node — power events off\n");
return;
};
smi_cmd = @truncate(rd32(f, 48));
acpi_enable_value = f[52];
pm1a_evt = @truncate(rd32(f, 56));
pm1b_evt = @truncate(rd32(f, 60));
pm1a_cnt = @truncate(rd32(f, 64));
pm1b_cnt = @truncate(rd32(f, 68));
gpe0_blk = @truncate(rd32(f, 80));
gpe1_blk = @truncate(rd32(f, 84));
pm1_evt_len = if (f.len > 88) f[88] else 4;
gpe0_len = if (f.len > 92) f[92] else 0;
gpe1_len = if (f.len > 93) f[93] else 0;
}
fn readSleepS5(ns: *aml.Namespace) bool {
const st = aml.sleepState(ns, 5) orelse return false;
s5_slp_typ_a = st.slp_typ_a;
s5_slp_typ_b = st.slp_typ_b;
return true;
}
/// Enable ACPI mode if the firmware isn't already in it, then bind the SCI and
/// set PWRBTN_EN so the power button raises an interrupt we can see.
fn armPowerButton(endpoint: runtime.ipc.Handle) void {
if (pm1a_cnt != 0 and (halPioRead(2, pm1a_cnt) & sci_en_bit) == 0 and smi_cmd != 0) {
// Switch to ACPI mode: write ACPI_ENABLE to the SMI command port, then
// spin (bounded) until SCI_EN latches.
halPioWrite(1, smi_cmd, acpi_enable_value);
var tries: u32 = 0;
while (tries < 1000 and (halPioRead(2, pm1a_cnt) & sci_en_bit) == 0) : (tries += 1) {
runtime.system.sleep(1);
}
}
if (!has_sci) {
_ = runtime.system.write("acpi: no SCI resource — power button unavailable\n");
return;
}
if (!device.irqBind(node_id, sci_resource_index, endpoint)) {
_ = runtime.system.write("acpi: SCI irq_bind failed\n");
return;
}
// PWRBTN_EN lives in the PM1 enable register at evt_blk + evt_len/2.
if (pm1a_evt != 0) {
const en_port = pm1a_evt + pm1_evt_len / 2;
halPioWrite(2, en_port, @as(u16, @truncate(halPioRead(2, en_port))) | pwrbtn_bit);
}
if (pm1b_evt != 0) {
const en_port = pm1b_evt + pm1_evt_len / 2;
halPioWrite(2, en_port, @as(u16, @truncate(halPioRead(2, en_port))) | pwrbtn_bit);
}
_ = runtime.system.write("acpi: power button armed\n");
}
/// The SCI fired. Read PM1 status; a set PWRBTN_STS is the power button — clear
/// it (write-1), publish, log. Any other set status is cleared and logged
/// (GPE/Notify dispatch is M21.2). Always re-arm the line.
fn onSci() void {
var handled = false;
inline for (.{ pm1a_evt, pm1b_evt }) |evt_port| {
if (evt_port != 0) {
const sts: u16 = @truncate(halPioRead(2, evt_port));
if (sts & pwrbtn_bit != 0) {
halPioWrite(2, evt_port, pwrbtn_bit); // write-1-to-clear
handled = true;
} else if (sts != 0) {
halPioWrite(2, evt_port, sts); // clear whatever else latched
}
}
}
if (handled) {
_ = runtime.system.write("power: button pressed\n");
publishButton();
}
handleGpe();
_ = device.irqAck(node_id, sci_resource_index);
}
/// General-purpose events: for each set+enabled GPE bit, evaluate its `\_GPE`
/// handler method (`_Lxx` level / `_Exx` edge), drain the Notify queue the
/// method produced, and publish an event per notified device. Then clear the
/// status bit. QEMU raises no GPEs on this config, so this path is exercised by
/// host unit tests (docs/m21-plan.md decision 5); on real hardware it carries
/// battery/AC/lid. The embedded controller's `_Qxx` queries are out of scope.
fn handleGpe() void {
handleGpeBlock(gpe0_blk, gpe0_len, 0);
handleGpeBlock(gpe1_blk, gpe1_len, gpe0_len * 4);
}
fn handleGpeBlock(blk: u16, len: u8, gpe_base: u32) void {
if (blk == 0 or len == 0) return;
const status_bytes = len / 2; // status half, then enable half
var byte_index: u8 = 0;
while (byte_index < status_bytes) : (byte_index += 1) {
const sts: u8 = @truncate(halPioRead(1, blk + byte_index));
const en: u8 = @truncate(halPioRead(1, blk + status_bytes + byte_index));
const active = sts & en;
if (active == 0) continue;
var bit: u3 = 0;
while (true) : (bit += 1) {
if (active & (@as(u8, 1) << bit) != 0) {
dispatchGpe(gpe_base + @as(u32, byte_index) * 8 + bit);
}
if (bit == 7) break;
}
halPioWrite(1, blk + byte_index, active); // write-1-to-clear the serviced bits
}
}
/// Evaluate the `\_GPE._L%02X` or `_E%02X` handler for GPE number `n`, then
/// publish an event for each device it notified.
fn dispatchGpe(n: u32) void {
const gpe_scope = aml.Namespace.resolve(&persistent_namespace, persistent_namespace.root, true, 0, &.{seg4("_GPE")}) orelse return;
var name: [4]u8 = .{ '_', 'L', 0, 0 };
writeHex2(name[2..4], n);
var method = aml.Namespace.childOf(gpe_scope, name);
if (method == null) {
name[1] = 'E';
method = aml.Namespace.childOf(gpe_scope, name);
}
const m = method orelse return; // no handler — the status bit was already cleared
_ = global_interpreter.evaluate(m, &.{}) catch return;
for (global_interpreter.takeNotifications()) |event| publishNotify(event.node, event.code);
}
fn publishNotify(node: *aml.Node, code: u64) void {
// Map the notified device's _HID to a domain event where we recognize it.
var hid: [8]u8 = .{0} ** 8;
if (readHid(node, &global_interpreter)) |h| hid = h;
const which: power.Event = if (std.mem.eql(u8, hid[0..7], "PNP0C0A")) .battery else if (std.mem.eql(u8, hid[0..7], "ACPI0003")) .ac else if (std.mem.eql(u8, hid[0..7], "PNP0C0D")) .lid else .notify;
var event = power.EventMessage{ .event = @intFromEnum(which), .code = @truncate(code) };
event.hid = hid;
writeLine("power: notify {s} code {d}\n", .{ hid[0..7], code });
publishEvent(std.mem.asBytes(&event));
}
/// Two lowercase hex digits of `n` into `out[0..2]`.
fn writeHex2(out: []u8, n: u32) void {
const digits = "0123456789ABCDEF";
out[0] = digits[(n >> 4) & 0xF];
out[1] = digits[n & 0xF];
}
fn publishButton() void {
const event = power.EventMessage{ .event = @intFromEnum(power.Event.power_button) };
publishEvent(std.mem.asBytes(&event));
}
fn publishEvent(bytes: []const u8) void {
for (&subscribers) |*slot| {
if (slot.*) |handle| {
if (!runtime.ipc.send(handle, bytes)) slot.* = null;
}
}
}
fn isSubscriber(task: u32) bool {
for (&subscribers, 0..) |*slot, si| {
if (slot.* != null and subscriber_tasks[si] == task) return true;
}
return false;
}
/// Enter S5 (soft off): write SLP_TYP|SLP_EN to the PM1 control register(s).
/// Mirrors the kernel's power.zig sleepValue. Only reached from a PID-1
/// shutdown request (M21.3).
fn enterS5() void {
if (!s5_valid or pm1a_cnt == 0) {
_ = runtime.system.write("power: S5 unavailable\n");
return;
}
_ = runtime.system.write("power: entering S5\n");
halPioWrite(2, pm1a_cnt, (@as(u32, s5_slp_typ_a & 0x7) << 10) | slp_en);
if (pm1b_cnt != 0) halPioWrite(2, pm1b_cnt, (@as(u32, s5_slp_typ_b & 0x7) << 10) | slp_en);
// If control returns, the write did not take — say so instead of hanging.
runtime.system.sleep(500);
_ = runtime.system.write("power: S5 write did not take\n");
}
// --- harness callbacks --------------------------------------------------------
fn onNotification(badge: u64) void {
// The only notification the service binds is the SCI (an IRQ badge).
_ = badge;
onSci();
}
/// The `.power` protocol: subscribe (endpoint as the call's capability),
/// shutdown (PID 1 only). Device discovery uses a different endpoint (the
/// device manager's), so nothing here handles ChildAdded.
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
if (message.len < 1) return 0;
switch (message[0]) {
@intFromEnum(power.Operation.subscribe) => {
var status: i32 = -1;
if (capability) |handle| {
for (&subscribers, 0..) |*slot, si| {
if (slot.* == null) {
slot.* = handle;
subscriber_tasks[si] = sender;
status = 0;
break;
}
}
}
const r = power.Reply{ .status = status };
@memcpy(reply[0..@sizeOf(power.Reply)], std.mem.asBytes(&r));
return @sizeOf(power.Reply);
},
@intFromEnum(power.Operation.shutdown) => {
// Honored only from a power subscriber — init, which has already run
// the stop sequence over everything else. The power service is
// mechanism (write S5); deciding *when* to shut down and stopping
// the rest of the system first is init's policy.
const allowed = isSubscriber(sender);
const r = power.Reply{ .status = if (allowed) 0 else -1 };
@memcpy(reply[0..@sizeOf(power.Reply)], std.mem.asBytes(&r));
if (allowed) enterS5();
return @sizeOf(power.Reply);
},
else => return 0,
}
}
/// Depth-first walk: register + report each present device with a _HID, then
/// descend. Scopes (\_SB, \_GPE …) are descended without producing a node.
fn walkDevices(node: *aml.Node, interpreter: *aml.Interpreter) void {
var child = node.first_child;
while (child) |c| : (child = c.next_sibling) {
if (c.kind != .device) {
walkDevices(c, interpreter);
continue;
}
if (!devicePresent(interpreter, c)) continue; // absent: skip it and its subtree
if (readHid(c, interpreter)) |hid| {
// Skip PCI roots — pci-bus already reports PCI functions; ACPI adds
// only the non-PCI _HID devices (docs/m19-m20-plan.md M20.2). The two
// roots are named through the shared registry, not bare _HID strings.
const id = acpi_ids.HardwareId.fromHid(hid[0..7]);
if (id != .pci_bus and id != .pci_express_root_bridge) {
registerDevice(c, hid, interpreter);
}
}
walkDevices(c, interpreter);
}
}
fn registerDevice(node: *aml.Node, hid: [8]u8, interpreter: *aml.Interpreter) void {
if (registered_count >= registered.len) return;
var descriptor = std.mem.zeroes(device.DeviceDescriptor);
descriptor.class = @intFromEnum(device.DeviceClass.acpi_device);
descriptor.pci_class = device.no_pci_class;
const hid_len: u64 = std.mem.indexOfScalar(u8, &hid, 0) orelse hid.len;
descriptor.hid_len = hid_len;
@memcpy(descriptor.hid[0..@intCast(hid_len)], hid[0..@intCast(hid_len)]);
applyCrs(&descriptor, node, interpreter);
const id = device.register(node_id, &descriptor) orelse {
writeLine("/system/services/acpi: register refused for {s}\n", .{hid[0..@intCast(hid_len)]});
return;
};
registered[registered_count] = .{ .hid = hid, .hid_len = @intCast(hid_len), .device_id = id, .resource_count = descriptor.resource_count };
registered_count += 1;
}
/// _STA bit 0 (present); absent method or a failed evaluation is treated as
/// present, per the ACPI rules.
fn devicePresent(interpreter: *aml.Interpreter, node: *aml.Node) bool {
const sta = aml.Namespace.childOf(node, seg4("_STA")) orelse return true;
const obj = interpreter.evaluate(sta, &.{}) catch return true;
const status = obj.asInteger() catch return true;
return (status & 0x01) != 0;
}
/// The device's EISA-decoded _HID (e.g. "PNP0303"), or null.
fn readHid(node: *aml.Node, interpreter: *aml.Interpreter) ?[8]u8 {
const hid = aml.Namespace.childOf(node, seg4("_HID")) orelse return null;
var buffer: [8]u8 = .{0} ** 8;
if (hid.kind == .method) {
const obj = interpreter.evaluate(hid, &.{}) catch return null;
switch (obj) {
.integer => |n| {
_ = eisaIdToStr(@truncate(n), &buffer);
return buffer;
},
else => return null,
}
}
if (hid.kind != .name or hid.value.len == 0) return null;
const v = hid.value;
switch (v[0]) {
// A static _HID names an integer EISA id: Zero/One/Ones or a Byte/Word/DWord/
// QWord integer prefix. Anything else is not an integer we can EISA-decode.
opcodes.zero_opcode, opcodes.one_opcode, opcodes.ones_opcode, opcodes.byte_prefix, opcodes.word_prefix, opcodes.dword_prefix, opcodes.qword_prefix => {
var p: usize = 0;
const n = readIntObj(v, &p) orelse return null;
_ = eisaIdToStr(@truncate(n), &buffer);
return buffer;
},
else => return null,
}
}
// --- _CRS resource-template decode (ported from the kernel's acpi.zig) --------
/// A resource template is a byte list of descriptors. Each starts with a tag byte whose
/// high bit picks the encoding: a *small* descriptor carries its type in bits [6:3] and
/// its length in bits [2:0]; a *large* descriptor is the whole tag byte, followed by a
/// 16-bit length. These are the descriptor types danos decodes into resources — named so
/// the walk below reads by descriptor, not by 0x04/0x85/… (docs/coding-standards.md).
const large_descriptor_bit: u8 = 0x80; // set in a tag byte => large descriptor
const small_length_mask: u8 = 0x07; // low 3 bits of a small tag = body length
const small_type_shift: u3 = 3; // small type sits in bits [6:3]
/// Small resource descriptor types (tag bits [6:3]). Non-exhaustive: an unhandled type
/// is skipped by its length, not misread.
const SmallResourceType = enum(u8) {
irq = 0x04,
io_port = 0x08,
fixed_io_port = 0x09,
end_tag = 0x0F,
_,
};
/// Large resource descriptor types (the whole tag byte). Non-exhaustive for the same reason.
const LargeResourceType = enum(u8) {
memory32 = 0x85,
memory32_fixed = 0x86,
extended_irq = 0x89,
_,
};
fn applyCrs(descriptor: *device.DeviceDescriptor, node: *aml.Node, interpreter: *aml.Interpreter) void {
const crs = aml.Namespace.childOf(node, seg4("_CRS")) orelse return;
const obj = interpreter.evaluate(crs, &.{}) catch return;
const bytes = switch (obj) {
.buffer => |b| b,
else => return,
};
var i: usize = 0;
while (i < bytes.len) {
const tag = bytes[i];
if (tag & large_descriptor_bit == 0) {
const len: usize = tag & small_length_mask;
const body = i + 1;
if (body + len > bytes.len) break;
switch (@as(SmallResourceType, @enumFromInt((tag >> small_type_shift) & 0x0F))) {
.irq => if (len >= 2) { // IRQ mask
const mask = @as(u16, bytes[body]) | (@as(u16, bytes[body + 1]) << 8);
var b: usize = 0;
while (b < 16) : (b += 1) {
if (mask & (@as(u16, 1) << @intCast(b)) != 0) addResource(descriptor, .irq, b, 1);
}
},
.io_port => if (len >= 7) addResource(descriptor, .io_port, rd16(bytes, body + 1), bytes[body + 6]),
.fixed_io_port => if (len >= 3) addResource(descriptor, .io_port, rd16(bytes, body), bytes[body + 2]),
.end_tag => break,
else => {},
}
i = body + len;
} else {
if (i + 3 > bytes.len) break;
const len: usize = @intCast(rd16(bytes, i + 1));
const body = i + 3;
if (body + len > bytes.len) break;
switch (@as(LargeResourceType, @enumFromInt(tag))) {
.memory32 => if (len >= 17) addResource(descriptor, .memory, rd32(bytes, body + 1), rd32(bytes, body + 13)),
.memory32_fixed => if (len >= 9) addResource(descriptor, .memory, rd32(bytes, body + 1), rd32(bytes, body + 5)),
.extended_irq => if (len >= 2) {
const count = bytes[body + 1];
var k: usize = 0;
while (k < count and body + 2 + k * 4 + 4 <= body + len) : (k += 1) {
addResource(descriptor, .irq, rd32(bytes, body + 2 + k * 4), 1);
}
},
else => {},
}
i = body + len;
}
}
}
fn addResource(descriptor: *device.DeviceDescriptor, kind: device.ResourceKind, start: u64, len: u64) void {
if (descriptor.resource_count >= descriptor.resources.len) return;
descriptor.resources[@intCast(descriptor.resource_count)] = .{ .kind = @intFromEnum(kind), .start = start, .len = len };
descriptor.resource_count += 1;
}
// --- small helpers ported verbatim from the kernel's acpi.zig ----------------
fn seg4(comptime s: *const [4:0]u8) [4]u8 {
return s[0..4].*;
}
fn hexDigit(n: u8) u8 {
return if (n < 10) '0' + n else 'A' + (n - 10);
}
fn eisaIdToStr(id: u32, buffer: *[8]u8) []const u8 {
const b0: u16 = @intCast(id & 0xFF);
const b1: u16 = @intCast((id >> 8) & 0xFF);
const b2: u8 = @truncate(id >> 16);
const b3: u8 = @truncate(id >> 24);
const mfg = (b0 << 8) | b1;
buffer[0] = '@' + @as(u8, @intCast((mfg >> 10) & 0x1F));
buffer[1] = '@' + @as(u8, @intCast((mfg >> 5) & 0x1F));
buffer[2] = '@' + @as(u8, @intCast(mfg & 0x1F));
buffer[3] = hexDigit((b2 >> 4) & 0xF);
buffer[4] = hexDigit(b2 & 0xF);
buffer[5] = hexDigit((b3 >> 4) & 0xF);
buffer[6] = hexDigit(b3 & 0xF);
buffer[7] = 0;
return buffer[0..7];
}
fn readIntObj(bytes: []const u8, p: *usize) ?u64 {
if (p.* >= bytes.len) return null;
const op = bytes[p.*];
p.* += 1;
switch (op) {
opcodes.zero_opcode => return 0,
opcodes.one_opcode => return 1,
opcodes.ones_opcode => return 1,
opcodes.byte_prefix => {
if (p.* >= bytes.len) return null;
const v = bytes[p.*];
p.* += 1;
return v;
},
opcodes.word_prefix => {
if (p.* + 2 > bytes.len) return null;
const v = rd16(bytes, p.*);
p.* += 2;
return v;
},
opcodes.dword_prefix => {
if (p.* + 4 > bytes.len) return null;
const v = rd32(bytes, p.*);
p.* += 4;
return v;
},
else => return null,
}
}
fn rd16(bytes: []const u8, off: usize) u64 {
return @as(u64, bytes[off]) | (@as(u64, bytes[off + 1]) << 8);
}
fn rd32(bytes: []const u8, off: usize) u64 {
return rd16(bytes, off) | (rd16(bytes, off + 2) << 16);
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
+44
View File
@@ -0,0 +1,44 @@
//! crash-test — a test fixture, not a driver: claims the device it is assigned,
//! hellos the device manager, announces itself, then faults on purpose. The
//! driver-restart scenario drives the manager's whole restart machinery with
//! it: fault → exit reason → backoff → respawn → the **same claim succeeding
//! again** (claim release on death, M17.1, through the manager's path) → the
//! crash-loop cap. Spawned bare (the initial-ramdisk sweep starts every bundled
//! binary), it exits silently so it cannot derange other tests.
const std = @import("std");
const runtime = @import("runtime");
const protocol = runtime.device_manager_protocol;
pub fn main(init: runtime.process.Init) void {
const argument = init.arguments.get(1) orelse return; // bare: stay silent
const assigned = std.fmt.parseInt(u64, argument, 10) catch return;
// The respawn only reaches this line because the kernel released the
// previous instance's claim at death. A failed claim exits cleanly — the
// manager reads "meant to stop" and the scenario fails loudly by silence.
if (!runtime.device.claim(assigned)) {
_ = runtime.system.write("crash-test: claim failed\n");
return;
}
var manager: ?runtime.ipc.Handle = null;
var tries: u32 = 0;
while (manager == null and tries < 100) : (tries += 1) {
manager = runtime.ipc.lookup(.device_manager);
if (manager == null) runtime.system.sleep(20);
}
const h = manager orelse return;
const hello = protocol.Hello{ .role = @intFromEnum(protocol.Role.device), .device_id = assigned };
var reply: [protocol.message_maximum]u8 = undefined;
_ = runtime.ipc.call(h, std.mem.asBytes(&hello), &reply) catch return;
_ = runtime.system.write("crash-test: faulting now\n");
const poison: *volatile u32 = @ptrFromInt(0xdead0000);
poison.* = 1; // the restart machinery's fuel: a real segmentation fault
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
@@ -0,0 +1,88 @@
//! device-list — the `ps` analog for the device tree (docs/device-manager.md
//! M18.3): asks the device manager for the tree over IPC, prints it, then
//! subscribes and prints every published add/remove event. The manager is the
//! one answer to "what devices exist" for user space; nothing here touches a
//! device_* system call.
const std = @import("std");
const runtime = @import("runtime");
const protocol = runtime.device_manager_protocol;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [96]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
pub fn main() void {
var manager: ?runtime.ipc.Handle = null;
var tries: u32 = 0;
while (manager == null and tries < 200) : (tries += 1) {
manager = runtime.ipc.lookup(.device_manager);
if (manager == null) runtime.system.sleep(20);
}
const h = manager orelse {
_ = runtime.system.write("device-list: no device manager\n");
return;
};
// The snapshot — polled briefly, because at boot the bus drivers may still
// be scanning: an empty first answer usually just means "too early".
var reply: [protocol.message_maximum]u8 = undefined;
var count: u32 = 0;
var length: usize = 0;
tries = 0;
while (tries < 20) : (tries += 1) {
const request = protocol.Enumerate{};
length = runtime.ipc.call(h, std.mem.asBytes(&request), &reply) catch 0;
if (length >= @sizeOf(protocol.EnumerateReply)) {
count = std.mem.bytesToValue(protocol.EnumerateReply, reply[0..@sizeOf(protocol.EnumerateReply)]).count;
if (count != 0) break;
}
runtime.system.sleep(100);
}
writeLine("device-list: {d} devices\n", .{count});
var offset: usize = @sizeOf(protocol.EnumerateReply);
var index: u32 = 0;
while (index < count and offset + @sizeOf(protocol.ChildEntry) <= length) : (index += 1) {
const entry = std.mem.bytesToValue(protocol.ChildEntry, reply[offset..][0..@sizeOf(protocol.ChildEntry)]);
writeLine("device-list: device {d} port {d} identity {d}\n", .{ entry.parent, entry.bus_address, entry.identity });
offset += @sizeOf(protocol.ChildEntry);
}
// The subscription: our endpoint rides as the call's capability; events
// arrive as buffered messages carrying the same structs the bus sends.
const endpoint = runtime.ipc.createIpcEndpoint() orelse {
_ = runtime.system.write("device-list: no endpoint\n");
return;
};
const subscribe = protocol.Subscribe{};
_ = runtime.ipc.callCap(h, std.mem.asBytes(&subscribe), &reply, endpoint) catch {
_ = runtime.system.write("device-list: subscribe failed\n");
return;
};
_ = runtime.system.write("device-list: subscribed\n");
var receive: [protocol.message_maximum]u8 = undefined;
while (true) {
const got = runtime.ipc.replyWait(endpoint, &.{}, &receive, null);
if (!got.isMessage() or got.len < 1) continue;
switch (receive[0]) {
@intFromEnum(protocol.Operation.child_added) => {
if (got.len < protocol.child_added_size) continue;
const event = std.mem.bytesToValue(protocol.ChildAdded, receive[0..protocol.child_added_size]);
writeLine("device-list: added (device {d} port {d})\n", .{ event.parent, event.bus_address });
},
@intFromEnum(protocol.Operation.child_removed) => {
if (got.len < protocol.child_removed_size) continue;
const event = std.mem.bytesToValue(protocol.ChildRemoved, receive[0..protocol.child_removed_size]);
writeLine("device-list: removed (device {d} port {d})\n", .{ event.parent, event.bus_address });
},
else => {},
}
}
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
@@ -0,0 +1,148 @@
//! The device-manager protocol (docs/device-manager.md): what drivers and
//! applications say to the device manager over its well-known endpoint. The
//! vfs-protocol pattern — extern-struct messages, a version in the handshake,
//! reserved fields — so both sides depend on the contract by name. Deliberately
//! contains nothing lifecycle-shaped: stopping, liveness (the zero-length ping),
//! and exit reasons are the universal vocabulary of
//! docs/process-lifecycle.md, not this protocol.
/// The protocol version a driver states in its hello. A manager that cannot
/// serve a driver's version refuses the hello, and the mismatch is loud at
/// startup instead of quiet corruption later.
pub const version: u16 = 1;
/// What kind of driver is talking (docs/driver-model.md's shapes).
pub const Role = enum(u8) {
/// Owns a controller and reports the devices behind it (`child_added`).
bus = 1,
/// Serves one device, reached through a bus's transfer protocol.
device = 2,
};
/// The message kinds.
pub const Operation = enum(u8) {
hello = 1,
child_added = 2,
child_removed = 3,
enumerate = 4,
subscribe = 5,
};
/// `Hello.device_id` for a driver that serves no enumerated device (a test
/// fixture, a synthetic source).
pub const no_device: u64 = ~@as(u64, 0);
/// The handshake, sent once by every driver the manager spawns — the manager's
/// one self-enforced deadline: spawned and silent past it means wrong binary,
/// wrong version, or wedged before main, and the stop sequence follows.
pub const Hello = extern struct {
operation: u8 = @intFromEnum(Operation.hello),
/// A Role value.
role: u8,
/// The protocol version this driver was built against (`version`).
version: u16 = version,
reserved: u32 = 0,
/// The device this driver was assigned (its argv[1]), or `no_device`.
device_id: u64,
};
pub const hello_size = @sizeOf(Hello);
/// The manager's answer to a hello. Nonzero status = refused (version mismatch,
/// unknown sender); a refused driver should exit cleanly.
pub const HelloReply = extern struct {
status: i32,
reserved: u32 = 0,
};
pub const reply_size = @sizeOf(HelloReply);
/// A bus driver reporting one device it discovered behind its controller
/// (docs/device-manager.md "the tree"). Identity is the bus's native language —
/// for USB a port-speed class; the (class, subclass, protocol) triple joins it
/// once control transfers exist (the USB track). The manager mirrors the child
/// into its tree; when the reporting driver dies, the manager prunes everything
/// it reported (the children describe protocol state that died with it) and the
/// restarted instance rediscovers and re-reports.
pub const ChildAdded = extern struct {
operation: u8 = @intFromEnum(Operation.child_added),
reserved0: u8 = 0,
reserved1: u16 = 0,
reserved2: u32 = 0,
/// The reporting driver's own device (the controller) — the child's parent.
parent: u64,
/// Where on the bus (for USB: the root port number, 1-based).
bus_address: u64,
/// Bus-specific identity (for USB: the PORTSC port-speed class; for PCI:
/// the class triple; for ACPI devices, 0 — identity is the hid below).
identity: u64,
/// The kernel device id this child was `device_register`ed as — what the
/// manager hands a matched driver as its argv assignment — or `no_device`
/// for an unregistered leaf (a USB port before the descriptor track).
device_id: u64 = no_device,
/// The ACPI hardware id (`_HID`), EISA-decoded (e.g. "PNP0303"), for devices
/// discovered by firmware string rather than a numeric bus identity. Empty
/// (all zero) otherwise. Widens for FDT `compatible` strings later.
hid: [8]u8 = .{0} ** 8,
};
pub const child_added_size = @sizeOf(ChildAdded);
/// A bus driver reporting a device gone (hot-unplug). Not yet sent by any
/// driver — the port scan has no unplug interrupt — but the manager handles it;
/// death-pruning covers removal until hotplug lands.
pub const ChildRemoved = extern struct {
operation: u8 = @intFromEnum(Operation.child_removed),
reserved0: u8 = 0,
reserved1: u16 = 0,
reserved2: u32 = 0,
parent: u64,
bus_address: u64,
};
pub const child_removed_size = @sizeOf(ChildRemoved);
/// The manager's answer to a tree report.
pub const ReportReply = extern struct {
status: i32,
reserved: u32 = 0,
};
/// An application asking for the tree (M18.3): the reply is an EnumerateReply
/// header followed by `count` ChildEntry records.
pub const Enumerate = extern struct {
operation: u8 = @intFromEnum(Operation.enumerate),
reserved0: u8 = 0,
reserved1: u16 = 0,
reserved2: u32 = 0,
};
pub const EnumerateReply = extern struct {
status: i32,
/// ChildEntry records following this header.
count: u32,
};
pub const ChildEntry = extern struct {
parent: u64,
bus_address: u64,
identity: u64,
};
/// An application subscribing to published add/remove events (the input-service
/// pattern): the subscriber's endpoint rides as the call's **capability**, and
/// events arrive on it as buffered messages whose payload is the same
/// ChildAdded / ChildRemoved struct the bus drivers send — one encoding, both
/// directions.
pub const Subscribe = extern struct {
operation: u8 = @intFromEnum(Operation.subscribe),
reserved0: u8 = 0,
reserved1: u16 = 0,
reserved2: u32 = 0,
};
/// Upper bound on any message in this protocol — sizes the endpoint buffers.
/// Capped by the kernel's IPC MESSAGE_MAXIMUM (256): an EnumerateReply carries
/// up to ten ChildEntry records per call, plenty for the mirror's current
/// bounds; paging joins the protocol if a tree ever outgrows one message.
pub const message_maximum = 256;
+496 -60
View File
@@ -1,20 +1,25 @@
//! /system/services/device-manager — the ring-3 process that turns the device //! /system/services/device-manager — the ring-3 process that turns the device
//! tree into a running system. The kernel enumerates the hardware and enforces the //! tree into a running system: **the matcher and the supervisor**
//! claim capability (mechanism); this decides *which driver serves which device* //! (docs/device-manager.md). The kernel enumerates the hardware and enforces the
//! and, eventually, spawns it (policy). Keeping that split in user space is the //! claim capability (mechanism); this decides which driver serves which device,
//! whole point of the microkernel: the manager is an ordinary, restartable process //! spawns it, and keeps it alive (policy). Keeping that split in user space is
//! with no special privilege — it uses the same `device_*` system calls any process //! the whole point of the microkernel: the manager is an ordinary, restartable
//! could ([drivers.md](../../../docs/drivers.md), [driver-model.md]). //! process with no special privilege.
//! //!
//! Increment 2 (this file): enumerate /system/devices, *match* each device to a //! M18.1 (this increment): the manager is a harness service on the well-known
//! driver, and *spawn* it with `system_spawn` — the kernel loads the named binary //! `.device_manager` endpoint. Every driver is spawned **supervised** — exit
//! from the initial-ramdisk as a fresh ring-3 process. On QEMU this discovers the //! notifications land in the same loop as protocol messages. Drivers with an
//! HPET, decides `hpet` serves it, and brings that driver all the way up. (The //! assignment must `hello` within a deadline or be stopped; a driver that dies
//! kernel still auto-spawns the whole initial-ramdisk at boot; increment 3 removes //! is restarted with backoff, and a crash loop (three fast deaths) marks it
//! that redundancy so the manager is the sole owner of driver spawning.) //! failed instead of respawning forever. Exit reasons (M17.2) drive the
//! decision: a clean exit meant to stop; only faults and missed deadlines
//! restart. Tree reports (`child_added`) land in M18.2.
const std = @import("std"); const std = @import("std");
const runtime = @import("runtime"); const runtime = @import("runtime");
const acpi_ids = @import("acpi-ids"); const acpi_ids = @import("acpi-ids");
const pci_class = @import("pci-class");
const protocol = runtime.device_manager_protocol;
const device = runtime.device; const device = runtime.device;
const system = runtime.system; const system = runtime.system;
@@ -27,86 +32,517 @@ fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
} }
/// The driver that serves each device — the policy table. In a fuller system /// The driver that serves each device — the policy table. In a fuller system
/// this comes from the drivers describing what they bind (or a manifest under /// this comes from a manifest (docs/device-manager.md: the third bus type
/// /system/drivers); for now it is a small static map, which is enough to prove the /// triggers it); for now a static map. `null` = no driver for this class yet.
/// manager reads the tree and decides. `null` = no driver for this class yet.
fn driverFor(d: device.DeviceDescriptor) ?[]const u8 { fn driverFor(d: device.DeviceDescriptor) ?[]const u8 {
// detect device via DeviceClass // The HPET timer node is still kernel-seeded (from the HPET table, not AML).
// PS/2 and other _HID devices now arrive as acpi-service reports and match
// in onChildAdded (M20.3), not from this boot snapshot.
if (d.class == @intFromEnum(device.DeviceClass.timer)) return "hpet"; if (d.class == @intFromEnum(device.DeviceClass.timer)) return "hpet";
// detect device via hid return null;
const hid = d.hid[0..@intCast(d.hid_len)];
const id = acpi_ids.HardwareId.fromHid(hid) orelse return null;
return switch (id) {
.ps2_keyboard, .ps2_mouse => "ps2-bus",
else => null,
};
} }
/// The PCI class/subclass/prog-IF triple of an xHCI (USB 3) host controller: /// The PCI class/subclass/prog-IF triple of an xHCI (USB 3) host controller —
/// Serial Bus Controller (0x0C) / USB Controller (0x03) / XHCI (0x30) — the names /// Serial Bus Controller / USB Controller / XHCI — named from pci-class.zig rather
/// pci-class.zig decodes. /// than written as the bare 0x0C0330 (docs/coding-standards.md, "Named values").
const xhci_pci_class: u64 = 0x0C_03_30; const xhci_pci_class: u64 = pci_class.ClassCode.pack(.{
.base = @intFromEnum(pci_class.BaseClass.serial_bus),
.subclass = @intFromEnum(pci_class.serial_bus.SubClass.usb),
.prog_if = @intFromEnum(pci_class.serial_bus.usb.ProgIf.xhci),
});
/// The bus driver that serves a PCI function, or null. Unlike the singleton drivers /// The driver that serves a *reported* PCI function (M19.3: matching moved
/// in `driverFor`, a machine can carry several identical controllers — so the caller /// from the boot snapshot to the bus reports), or null. A machine can carry
/// spawns one driver instance *per device*, passing the device id as argv[1] for the /// several identical controllers — one driver instance per reported device,
/// instance to claim. /// its registered id as argv[1].
fn pciDriverFor(d: device.DeviceDescriptor) ?[]const u8 { fn pciDriverForIdentity(identity: u64) ?[]const u8 {
if (d.class != @intFromEnum(device.DeviceClass.pci_device)) return null; return switch (identity) {
return switch (d.pci_class) {
xhci_pci_class => "usb-xhci-bus", xhci_pci_class => "usb-xhci-bus",
else => null, else => null,
}; };
} }
/// Spawn one instance of `driver_name` to serve the specific device `id` — the id /// The driver that serves a *reported* ACPI device by its `_HID` (M20.3:
/// arrives as argv[1]. No isProcessRunning gate here: the name alone cannot tell two /// ps2-bus now binds the PS/2 nodes the acpi service reports, not boot-snapshot
/// instances apart, and this manager is the sole spawner of drivers. /// nodes the kernel used to build). ps2-bus is a singleton that finds both its
fn spawnForDevice(driver_name: []const u8, id: u64) void { /// devices by hid once spawned, so keyboard and mouse map to the same name.
var text: [20]u8 = undefined; fn hidDriverFor(hid: []const u8) ?[]const u8 {
const id_text = std.fmt.bufPrint(&text, "{d}", .{id}) catch return; if (std.mem.eql(u8, hid, "PNP0303")) return "ps2-bus"; // PS/2 keyboard
if (system.spawnWithArguments(driver_name, &.{id_text}) != null) { if (std.mem.eql(u8, hid, "PNP0F13")) return "ps2-bus"; // PS/2 mouse
writeLine("device-manager: spawned {s} for device {d}\n", .{ driver_name, id }); return null;
} else { }
writeLine("device-manager: failed to spawn {s} for device {d}\n", .{ driver_name, id });
/// Whether some driver entry already serves registered device `device_id` —
/// a re-report after a bus restart must not spawn a second instance.
fn driverForDevice(device_id: u64) bool {
for (&drivers) |*driver| {
if (driver.used and driver.device_id == device_id) return true;
}
return false;
}
// --- supervision -------------------------------------------------------------
/// How long a protocol driver has to hello after its spawn.
const hello_deadline_ms: u64 = 3000;
/// Deaths faster than this count toward the crash loop; slower ones reset it.
const fast_death_ns: u64 = 2_000_000_000;
/// Consecutive fast deaths before the manager gives up on a driver.
const crash_loop_cap: u32 = 3;
/// Restart backoff: base << (restarts - 1), so 300 ms, 600 ms, 1200 ms.
const backoff_base_ms: u64 = 300;
const DriverState = enum {
awaiting_hello, // spawned; the deadline is armed (protocol drivers only)
running,
restarting, // dead; respawn due at restart_due_ns
stopped, // exited cleanly — it meant to; not restarted
failed, // crash loop, or unspawnable; the manager gave up
};
const Driver = struct {
used: bool = false,
name_buffer: [24]u8 = undefined,
name_len: usize = 0,
// The assigned device id (becomes argv[1]), or protocol.no_device.
device_id: u64 = protocol.no_device,
// Whether this driver speaks the protocol (hello expected, deadline
// enforced). Legacy drivers (hpet, ps2-bus) are supervised and restarted
// but not yet required to hello.
speaks_protocol: bool = false,
process_id: u32 = 0,
state: DriverState = .running,
restarts: u32 = 0,
spawn_ns: u64 = 0,
hello_deadline_ns: u64 = 0,
restart_due_ns: u64 = 0,
fn name(driver: *const Driver) []const u8 {
return driver.name_buffer[0..driver.name_len];
}
};
const maximum_drivers = 16;
var drivers: [maximum_drivers]Driver = .{Driver{}} ** maximum_drivers;
var manager_endpoint: runtime.ipc.Handle = 0;
var test_restart_mode = false;
var test_usb_restart_mode = false;
var test_usb_killed = false;
var test_pci_restart_mode = false;
var test_kill_pid: u32 = 0;
var test_kill_due_ns: u64 = 0;
/// The application subscribers (M18.3, the input-service pattern): endpoints
/// handed over as capabilities, each receiving every child add/remove as a
/// buffered message. A subscriber whose endpoint stops accepting (it died) is
/// dropped on the failed send.
const maximum_subscribers = 8;
var subscribers: [maximum_subscribers]?runtime.ipc.Handle = .{null} ** maximum_subscribers;
/// Publish one event (a ChildAdded or ChildRemoved struct, the same encoding
/// the bus drivers send) to every subscriber.
fn publishEvent(event: []const u8) void {
for (&subscribers) |*slot| {
if (slot.*) |handle| {
if (!runtime.ipc.send(handle, event)) slot.* = null; // dead subscriber
}
} }
} }
pub fn main() void { /// The manager's mirror of what bus drivers report (docs/device-manager.md "the
/// tree"): the children, keyed by (parent, bus address), each remembering which
/// driver instance reported it — that is what death-pruning sweeps by.
const Child = struct {
used: bool = false,
parent: u64 = 0,
bus_address: u64 = 0,
identity: u64 = 0,
// The kernel device id (registered by the reporter), or protocol.no_device.
device_id: u64 = 0,
reporter: u32 = 0, // the reporting driver instance's process id
};
const maximum_children = 64; // ACPI adds ~34 device nodes (M20.2), plus PCI + USB
var children: [maximum_children]Child = .{Child{}} ** maximum_children;
/// Record (or refresh) a reported child. Refreshing matters: a restarted bus
/// driver re-reports what it rediscovers, and the same (parent, port) must not
/// duplicate.
fn addChild(parent: u64, bus_address: u64, identity: u64, device_id: u64, reporter: u32) bool {
var free: ?*Child = null;
for (&children) |*child| {
if (child.used and child.parent == parent and child.bus_address == bus_address) {
child.identity = identity;
child.device_id = device_id;
child.reporter = reporter;
return true;
}
if (!child.used and free == null) free = child;
}
const slot = free orelse return false;
slot.* = .{ .used = true, .parent = parent, .bus_address = bus_address, .identity = identity, .device_id = device_id, .reporter = reporter };
return true;
}
/// Prune every child a dead driver instance reported: the children describe
/// protocol state (slots, rings) that died with the process — keeping the nodes
/// would be keeping a lie. The restarted instance rediscovers and re-reports.
/// Watchers hear the honest story: removed now, added again on rediscovery.
fn pruneChildrenOf(reporter: u32) void {
for (&children) |*child| {
if (child.used and child.reporter == reporter) {
writeLine("/system/services/device-manager: child removed (device {d} port {d})\n", .{ child.parent, child.bus_address });
child.used = false;
const event = protocol.ChildRemoved{ .parent = child.parent, .bus_address = child.bus_address };
publishEvent(std.mem.asBytes(&event));
}
}
}
/// How many children a driver instance has reported (the test-usb-restart
/// trigger counts these).
fn childCountOf(reporter: u32) u32 {
var n: u32 = 0;
for (&children) |*child| {
if (child.used and child.reporter == reporter) n += 1;
}
return n;
}
fn driverByProcess(process_id: u32) ?*Driver {
for (&drivers) |*driver| {
if (driver.used and driver.process_id == process_id) return driver;
}
return null;
}
/// Whether a singleton driver is already in the table (two ACPI nodes can both
/// map to ps2-bus; one instance serves both).
fn alreadySupervised(name: []const u8) bool {
for (&drivers) |*driver| {
if (driver.used and std.mem.eql(u8, driver.name(), name)) return true;
}
return false;
}
/// Record a driver in the table and spawn its first instance.
fn addDriver(name: []const u8, device_id: u64, speaks_protocol: bool) void {
for (&drivers) |*driver| {
if (driver.used) continue;
const n = @min(name.len, driver.name_buffer.len);
@memcpy(driver.name_buffer[0..n], name[0..n]);
driver.name_len = n;
driver.device_id = device_id;
driver.speaks_protocol = speaks_protocol;
driver.used = true;
spawnDriver(driver);
return;
}
writeLine("/system/services/device-manager: driver table full; cannot supervise {s}\n", .{name});
}
/// (Re)spawn a driver instance: supervised on the manager's own endpoint, the
/// device id as argv[1] when it has one, the hello deadline armed when it
/// speaks the protocol.
fn spawnDriver(driver: *Driver) void {
var id_text: [20]u8 = undefined;
var arguments: [1][]const u8 = undefined;
var argument_count: usize = 0;
if (driver.device_id != protocol.no_device) {
arguments[0] = std.fmt.bufPrint(&id_text, "{d}", .{driver.device_id}) catch return;
argument_count = 1;
}
const child = system.spawnSupervised(driver.name(), arguments[0..argument_count], manager_endpoint) orelse {
writeLine("/system/services/device-manager: failed to spawn {s}\n", .{driver.name()});
driver.state = .failed;
return;
};
driver.process_id = child;
driver.spawn_ns = system.clock();
if (driver.speaks_protocol) {
driver.state = .awaiting_hello;
driver.hello_deadline_ns = driver.spawn_ns + hello_deadline_ms * 1_000_000;
_ = system.timerOnce(manager_endpoint, hello_deadline_ms + 100);
} else {
driver.state = .running;
}
if (driver.device_id != protocol.no_device) {
writeLine("/system/services/device-manager: spawned {s} for device {d}\n", .{ driver.name(), driver.device_id });
} else {
writeLine("/system/services/device-manager: spawned {s}\n", .{driver.name()});
}
}
/// A driver died. Prune what it reported first — then the exit reason (M17.2)
/// is the whole restart decision: a clean exit meant to stop; anything else
/// restarts with backoff until the crash-loop cap.
fn onDriverExit(driver: *Driver) void {
pruneChildrenOf(driver.process_id);
const reason = runtime.process.exitReason(driver.process_id) orelse .fault;
if (reason == .exited) {
driver.state = .stopped;
writeLine("/system/services/device-manager: {s} exited cleanly; not restarting\n", .{driver.name()});
return;
}
const now = system.clock();
const alive_ns = now - driver.spawn_ns;
driver.restarts = if (alive_ns < fast_death_ns) driver.restarts + 1 else 1;
if (driver.restarts >= crash_loop_cap) {
driver.state = .failed;
writeLine("/system/services/device-manager: {s} is failing repeatedly (crash loop); giving up\n", .{driver.name()});
return;
}
const delay_ms = backoff_base_ms << @intCast(driver.restarts - 1);
driver.state = .restarting;
driver.restart_due_ns = now + delay_ms * 1_000_000;
writeLine("/system/services/device-manager: restarting {s} in {d} ms (died: {s})\n", .{ driver.name(), delay_ms, @tagName(reason) });
_ = system.timerOnce(manager_endpoint, delay_ms + 50);
}
/// A timer landed: sweep every deadline. Overdue hellos are killed (the exit
/// notification then routes through the normal restart policy); due restarts
/// respawn. Timers carry no id on purpose — the table is the state, and one
/// sweep serves every armed deadline.
fn sweepDeadlines() void {
const now = system.clock();
if (test_kill_pid != 0 and now >= test_kill_due_ns) {
writeLine("/system/services/device-manager: test mode: killing the reporter\n", .{});
_ = system.kill(test_kill_pid);
test_kill_pid = 0;
}
for (&drivers) |*driver| {
if (!driver.used) continue;
switch (driver.state) {
.awaiting_hello => if (now >= driver.hello_deadline_ns) {
writeLine("/system/services/device-manager: {s} missed its hello deadline\n", .{driver.name()});
_ = system.kill(driver.process_id);
// The exit notification finishes the job via onDriverExit.
},
.restarting => if (now >= driver.restart_due_ns) spawnDriver(driver),
else => {},
}
}
}
// --- the harness callbacks -----------------------------------------------------
fn initialise(endpoint: runtime.ipc.Handle) bool {
manager_endpoint = endpoint;
// Enumerate into a heap buffer (too big for the one-page user stack). // Enumerate into a heap buffer (too big for the one-page user stack).
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch { const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("device-manager: out of memory\n"); _ = runtime.system.write("/system/services/device-manager: out of memory\n");
return; return false;
}; };
const total = device.enumerate(buffer); const total = device.enumerate(buffer);
const count = @min(total, buffer.len); const count = @min(total, buffer.len);
var matched: usize = 0; var matched: usize = 0;
for (buffer[0..count]) |descriptor| { for (buffer[0..count]) |descriptor| {
if (pciDriverFor(descriptor)) |driver_name| { if (descriptor.class == @intFromEnum(device.DeviceClass.pci_host_bridge)) {
// The PCI bus driver: enumeration in ring 3 (M19), one instance
// per bridge, the bridge id as its assignment.
matched += 1; matched += 1;
spawnForDevice(driver_name, descriptor.id); addDriver("pci-bus", descriptor.id, true);
continue; continue;
} }
// PCI functions no longer appear in the boot snapshot (M19.3): the
// pci-bus driver reports them, and onChildAdded matches from reports.
const driver_name = driverFor(descriptor) orelse continue; const driver_name = driverFor(descriptor) orelse continue;
matched += 1; matched += 1;
if (!system.isProcessRunning(driver_name)) { // Skip a singleton that is already alive (the initial-ramdisk sweep test
if (runtime.system.spawn(driver_name) != null) { // starts every bundled binary bare, this manager included) — spawning a
writeLine("device-manager: spawned {s}\n", .{driver_name}); // second instance would only lose the claim race and churn the log.
} else { if (!alreadySupervised(driver_name) and !system.isProcessRunning(driver_name)) {
writeLine("device-manager: failed to spawn {s}\n", .{driver_name}); addDriver(driver_name, protocol.no_device, false);
}
} else {
writeLine("device-manager: already spawned {s}\n", .{driver_name});
} }
} }
// The discovery service (docs/m19-m20-plan.md M20): one per firmware, packed
// under the neutral name "discovery", spawned once at startup. It finds and
// claims the acpi-tables (or devicetree-blob) node itself. Not a per-device
// match — it is the discoverer, not a driver bound to one device.
addDriver("discovery", protocol.no_device, false);
if (test_restart_mode) {
// The driver-restart scenario's fixture: claims device 0 (the tree
// root, otherwise unclaimed), hellos, then faults — driving backoff,
// re-claim-after-death, and the crash-loop cap deterministically.
addDriver("crash-test", 0, true);
}
if (matched == 0) { if (matched == 0) {
_ = runtime.system.write("device-manager: no matchable devices\n"); _ = runtime.system.write("/system/services/device-manager: no matchable devices\n");
} else {
_ = runtime.system.write("/system/services/device-manager: ok\n");
}
return true;
}
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
if (message.len < 1) return 0;
switch (message[0]) {
@intFromEnum(protocol.Operation.child_added) => return onChildAdded(message, reply, sender),
@intFromEnum(protocol.Operation.child_removed) => return onChildRemoved(message, reply, sender),
@intFromEnum(protocol.Operation.enumerate) => return onEnumerate(reply),
@intFromEnum(protocol.Operation.subscribe) => return onSubscribe(reply, capability),
@intFromEnum(protocol.Operation.hello) => {},
else => return 0,
}
if (message.len < protocol.hello_size) return 0;
const hello = std.mem.bytesToValue(protocol.Hello, message[0..protocol.hello_size]);
var status: i32 = 0;
if (hello.version != protocol.version) {
status = -1;
writeLine("/system/services/device-manager: refused hello (version {d}) from process {d}\n", .{ hello.version, sender });
} else if (driverByProcess(sender)) |driver| {
driver.state = .running;
writeLine("/system/services/device-manager: hello from {s} (device {d})\n", .{ driver.name(), hello.device_id });
} else {
status = -1;
writeLine("/system/services/device-manager: hello from unknown process {d}\n", .{sender});
}
const hello_reply = protocol.HelloReply{ .status = status };
@memcpy(reply[0..protocol.reply_size], std.mem.asBytes(&hello_reply));
return protocol.reply_size;
}
/// A bus driver reported a discovered device: mirror it, and in
/// test-usb-restart mode kill the reporter once after its second child — the
/// deterministic trigger for prune -> backoff -> respawn -> re-report.
fn onChildAdded(message: []const u8, reply: []u8, sender: u32) usize {
if (message.len < protocol.child_added_size) return 0;
const report = std.mem.bytesToValue(protocol.ChildAdded, message[0..protocol.child_added_size]);
var status: i32 = 0;
if (driverByProcess(sender)) |driver| {
if (!addChild(report.parent, report.bus_address, report.identity, report.device_id, sender)) status = -1;
writeLine("/system/services/device-manager: child added (device {d} port {d}, identity {d}) by {s}\n", .{ report.parent, report.bus_address, report.identity, driver.name() });
if (status == 0) publishEvent(message[0..protocol.child_added_size]);
// Matching from reports (M19.3): a registered child whose identity
// names a driver gets one, once — re-reports after a bus restart
// dedupe on the registered id, exactly like the registrations do.
if (status == 0 and report.device_id != protocol.no_device) {
if (pciDriverForIdentity(report.identity)) |child_driver| {
if (!driverForDevice(report.device_id)) addDriver(child_driver, report.device_id, true);
}
// ACPI _HID match (M20.3): ps2-bus is a singleton that finds its own
// devices by hid, so spawn it once, without a device assignment.
const hid_len = std.mem.indexOfScalar(u8, &report.hid, 0) orelse report.hid.len;
if (hid_len != 0) {
if (hidDriverFor(report.hid[0..hid_len])) |hid_driver| {
if (!alreadySupervised(hid_driver)) addDriver(hid_driver, protocol.no_device, false);
}
}
}
} else {
status = -1;
}
const report_reply = protocol.ReportReply{ .status = status };
@memcpy(reply[0..@sizeOf(protocol.ReportReply)], std.mem.asBytes(&report_reply));
if (test_pci_restart_mode and !test_usb_killed) {
if (driverByProcess(sender)) |driver| {
if (std.mem.eql(u8, driver.name(), "pci-bus") and childCountOf(sender) >= 3) {
// The pci restart drill: kill the enumerator after it has
// reported; the respawn must re-register without duplicates
// (M19.0 idempotence, proven end to end by pci-scan).
test_usb_killed = true;
test_kill_pid = sender;
test_kill_due_ns = system.clock() + 1_000_000_000;
_ = system.timerOnce(manager_endpoint, 1100);
}
}
}
if (test_usb_restart_mode and !test_usb_killed and childCountOf(sender) >= 2) {
// Only the xHCI reporter is the drill's victim — pci-bus also reports
// now, and whichever finishes second must not trigger the kill.
if (driverByProcess(sender)) |driver| {
if (std.mem.eql(u8, driver.name(), "usb-xhci-bus")) {
// Delayed, not immediate: the device-list scenario's subscriber
// needs a window to enumerate and subscribe before the events.
test_usb_killed = true;
test_kill_pid = sender;
test_kill_due_ns = system.clock() + 2_000_000_000;
_ = system.timerOnce(manager_endpoint, 2100);
}
}
}
return @sizeOf(protocol.ReportReply);
}
/// A bus driver reported a device gone (hot-unplug; no sender exists yet, but
/// the handler is protocol-complete — death-pruning covers removal until then).
fn onChildRemoved(message: []const u8, reply: []u8, sender: u32) usize {
if (message.len < protocol.child_removed_size) return 0;
const report = std.mem.bytesToValue(protocol.ChildRemoved, message[0..protocol.child_removed_size]);
var status: i32 = -1;
for (&children) |*child| {
if (child.used and child.parent == report.parent and child.bus_address == report.bus_address and child.reporter == sender) {
writeLine("/system/services/device-manager: child removed (device {d} port {d})\n", .{ child.parent, child.bus_address });
child.used = false;
status = 0;
}
}
const report_reply = protocol.ReportReply{ .status = status };
@memcpy(reply[0..@sizeOf(protocol.ReportReply)], std.mem.asBytes(&report_reply));
return @sizeOf(protocol.ReportReply);
}
/// An application asked for the tree: the mirror, as a header plus entries.
fn onEnumerate(reply: []u8) usize {
var count: u32 = 0;
var offset: usize = @sizeOf(protocol.EnumerateReply);
for (&children) |*child| {
if (!child.used) continue;
if (offset + @sizeOf(protocol.ChildEntry) > reply.len) break;
const entry = protocol.ChildEntry{ .parent = child.parent, .bus_address = child.bus_address, .identity = child.identity };
@memcpy(reply[offset..][0..@sizeOf(protocol.ChildEntry)], std.mem.asBytes(&entry));
offset += @sizeOf(protocol.ChildEntry);
count += 1;
}
const header = protocol.EnumerateReply{ .status = 0, .count = count };
@memcpy(reply[0..@sizeOf(protocol.EnumerateReply)], std.mem.asBytes(&header));
return offset;
}
/// An application subscribed: its endpoint arrived as the call's capability.
fn onSubscribe(reply: []u8, capability: ?runtime.ipc.Handle) usize {
var status: i32 = -1;
if (capability) |handle| {
for (&subscribers) |*slot| {
if (slot.* == null) {
slot.* = handle;
status = 0;
break;
}
}
}
const report_reply = protocol.ReportReply{ .status = status };
@memcpy(reply[0..@sizeOf(protocol.ReportReply)], std.mem.asBytes(&report_reply));
return @sizeOf(protocol.ReportReply);
}
fn onNotification(badge: u64) void {
if (badge & runtime.ipc.notify_exit_bit != 0) {
const dead: u32 = @intCast(badge & ~(runtime.ipc.notify_badge_bit | runtime.ipc.notify_exit_bit));
if (driverByProcess(dead)) |driver| onDriverExit(driver);
return; return;
} }
_ = runtime.system.write("device-manager: ok\n"); if (badge & runtime.ipc.notify_timer_bit != 0) sweepDeadlines();
while (true) runtime.system.sleep(1000); }
pub fn main(init: runtime.process.Init) void {
if (init.arguments.get(1)) |mode| {
test_restart_mode = std.mem.eql(u8, mode, "test-restart");
test_usb_restart_mode = std.mem.eql(u8, mode, "test-usb-restart");
test_pci_restart_mode = std.mem.eql(u8, mode, "test-pci-restart");
}
runtime.service.run(protocol.message_maximum, .{
.service = .device_manager,
.init = initialise,
.on_message = onMessage,
.on_notification = onNotification,
});
} }
pub const panic = runtime.panic; pub const panic = runtime.panic;
+35
View File
@@ -0,0 +1,35 @@
//! /system/services/fdt — the devicetree discovery service: the ARM twin of the
//! acpi service (docs/m19-m20-plan.md decision 7). **Placeholder: not
//! implemented.** It exists so the build's `-Ddiscovery` option has both of its
//! values from day one; the implementation lands with the Raspberry Pi
//! bring-up (docs/arm.md).
//!
//! What it becomes: the per-firmware discoverer for boots that hand over a
//! flattened device tree instead of ACPI tables. It claims the
//! `devicetree-blob` node the kernel publishes (the FDT the loader received),
//! walks the tree — pure data, no bytecode, so unlike the acpi service it
//! needs no port grant and no interpreter — and, like any bus-shaped driver:
//! `device_register`s what it finds (containment against the blob node's
//! recorded apertures), reports each child to the device manager
//! (`child_added`, identity = the node's `compatible` string), and stays
//! resident under the manager's supervision (hello, restart, the usual
//! contract).
//!
//! Known prerequisite recorded in the plan: `DeviceDescriptor`'s 8-byte `hid`
//! cannot hold an FDT `compatible` string ("brcm,bcm2835-aux-uart") — identity
//! widens before this file grows a body.
const runtime = @import("runtime");
pub fn main(init: runtime.process.Init) void {
_ = init;
// Not implemented: exit cleanly and silently (a bare spawn by the
// initial-ramdisk sweep must not derange other tests' markers). The
// supervisor reads a clean exit as "meant to stop" — correct for a
// placeholder.
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
+98 -11
View File
@@ -6,12 +6,20 @@
//! //!
//! It proves the C-convention heap works, then — as PID 1 — acts as the system's //! It proves the C-convention heap works, then — as PID 1 — acts as the system's
//! **service supervisor**: it spawns the user-space services danos brings up at boot //! **service supervisor**: it spawns the user-space services danos brings up at boot
//! (the VFS server, the device manager), and settles into a heartbeat so it stays //! (the VFS server, the device manager), and settles into an event loop as the root
//! alive as the root of user space. Drivers are *not* its job: the device manager //! of user space. Drivers are *not* its job: the device manager discovers the
//! discovers the hardware and spawns those. This is the service half of the //! hardware and spawns those. This is the service half of the service/driver spawn
//! service/driver spawn split (docs/driver-model.md). //! split (docs/driver-model.md).
//!
//! M21: init also owns **orderly shutdown**. It supervises its children (keeping
//! their ids and an exit endpoint), subscribes to the power service, and on a
//! power-button event runs the stop sequence over its children in reverse order
//! before asking the power service to enter S5 — lifecycle (M17) and events (M21)
//! composing into a clean poweroff.
const std = @import("std");
const runtime = @import("runtime"); const runtime = @import("runtime");
const power = runtime.power_protocol;
/// The system services init brings up at boot, in order. This is init's policy — the /// The system services init brings up at boot, in order. This is init's policy — the
/// microkernel keeps such choices in user space, not the kernel. Drivers are absent /// microkernel keeps such choices in user space, not the kernel. Drivers are absent
@@ -19,6 +27,10 @@ const runtime = @import("runtime");
/// manifest under /system/services instead of a hardcoded list.) /// manifest under /system/services instead of a hardcoded list.)
const boot_services = [_][]const u8{ "vfs", "input", "device-manager" }; const boot_services = [_][]const u8{ "vfs", "input", "device-manager" };
var children: [boot_services.len]u32 = .{0} ** boot_services.len;
var child_count: usize = 0;
var supervision_endpoint: runtime.ipc.Handle = 0;
pub fn main() void { pub fn main() void {
// Prove the heap end to end: allocate through the runtime allocator (which // Prove the heap end to end: allocate through the runtime allocator (which
// mmaps pages from the kernel and carves them with the free list), write into // mmaps pages from the kernel and carves them with the free list), write into
@@ -28,23 +40,98 @@ pub fn main() void {
// the extern malloc/free symbols; Zig code uses this allocator.) // the extern malloc/free symbols; Zig code uses this allocator.)
const gpa = runtime.allocator(); const gpa = runtime.allocator();
if (gpa.alloc(u8, 64)) |buffer| { if (gpa.alloc(u8, 64)) |buffer| {
const message = "init: heap ok\n"; const message = "/system/services/init: heap ok\n";
@memcpy(buffer[0..message.len], message); @memcpy(buffer[0..message.len], message);
_ = runtime.system.write(buffer[0..message.len]); _ = runtime.system.write(buffer[0..message.len]);
gpa.free(buffer); gpa.free(buffer);
} else |_| {} } else |_| {}
// Bring up the boot services. Best-effort and silent: each service announces its // One endpoint carries everything init waits on: children's exit
// own readiness (`vfs: ready`, ...), and in an isolation test that runs init with // notifications (they are spawned supervised against it), init's own
// no initial-ramdisk the spawns simply no-op rather than deranging the heartbeat. // signals, and power events it subscribes to. All arrive in the loop below.
supervision_endpoint = runtime.ipc.createIpcEndpoint() orelse {
_ = runtime.system.write("/system/services/init: no endpoint\n");
return;
};
_ = runtime.process.bindSignals(supervision_endpoint);
// Bring up the boot services, supervised so init can stop them cleanly.
// Best-effort and silent: each service announces its own readiness, and in
// an isolation test with no initial-ramdisk the spawns simply no-op.
for (boot_services) |service| { for (boot_services) |service| {
_ = runtime.system.spawn(service); if (runtime.system.spawnSupervised(service, &.{}, supervision_endpoint)) |id| {
children[child_count] = id;
child_count += 1;
}
} }
// Subscribe to power events (retry: the power service registers well after
// init starts). Best-effort — without it, a `terminate` signal still
// triggers the same shutdown path.
subscribePower();
// A re-arming timer drives the liveness heartbeat: proof PID 1 is alive
// (the init test's marker) while the loop stays free to receive signals,
// power events, and children's exit notifications.
_ = runtime.system.timerOnce(supervision_endpoint, 1000);
var receive: [power.message_maximum]u8 = undefined;
while (true) { while (true) {
_ = runtime.system.write("init: heartbeat\n"); const got = runtime.ipc.replyWait(supervision_endpoint, &.{}, &receive, null);
runtime.system.sleep(1000); if (runtime.process.signalsFrom(got.badge)) |signals| {
if (signals.has(.terminate)) shutDown();
continue;
} }
if (got.isTimer()) {
_ = runtime.system.write("/system/services/init: heartbeat\n");
_ = runtime.system.timerOnce(supervision_endpoint, 1000);
continue;
}
if (got.isMessage() and got.len >= 2 and receive[0] == @intFromEnum(power.Operation.event)) {
// A power event (the only buffered messages init receives).
if (receive[1] == @intFromEnum(power.Event.power_button)) shutDown();
continue;
}
// Child-exit notifications and anything else: keep waiting.
if (got.isNotification()) continue;
}
}
/// Look up the power service and subscribe our endpoint (handed over as the
/// call's capability) so events arrive as buffered messages here.
fn subscribePower() void {
var handle: ?runtime.ipc.Handle = null;
var tries: u32 = 0;
while (handle == null and tries < 200) : (tries += 1) {
handle = runtime.ipc.lookup(.power);
if (handle == null) runtime.system.sleep(20);
}
// A missing power service is not fatal — init proceeds to its heartbeat and
// a `terminate` signal still drives shutdown. Silent so the no-ramdisk init
// test's heartbeat marker is the next line written.
const h = handle orelse return;
const request = power.Subscribe{};
var reply: [power.message_maximum]u8 = undefined;
_ = runtime.ipc.callCap(h, std.mem.asBytes(&request), &reply, supervision_endpoint) catch {};
}
/// The stop sequence: terminate each child in reverse spawn order (vfs last —
/// other services may flush through it), waiting up to a deadline for each to
/// exit before killing it, then ask the power service to enter S5.
fn shutDown() void {
_ = runtime.system.write("/system/services/init: shutting down\n");
var i = child_count;
while (i > 0) {
i -= 1;
if (children[i] != 0) runtime.process.stop(children[i], 2000, supervision_endpoint);
}
if (runtime.ipc.lookup(.power)) |h| {
const request = power.Shutdown{};
var reply: [power.message_maximum]u8 = undefined;
_ = runtime.ipc.call(h, std.mem.asBytes(&request), &reply) catch {};
}
// If S5 did not take, init has nothing left to do but idle.
while (true) runtime.system.sleep(1000);
} }
pub const panic = runtime.panic; pub const panic = runtime.panic;
+3 -3
View File
@@ -115,14 +115,14 @@ fn handle(message: []const u8, got: ipc.Received, out: []u8) usize {
pub fn main() void { pub fn main() void {
const endpoint = ipc.createIpcEndpoint() orelse { const endpoint = ipc.createIpcEndpoint() orelse {
_ = system.write("input: no endpoint\n"); _ = system.write("/system/services/input: no endpoint\n");
return; return;
}; };
if (!ipc.register(.input, endpoint)) { if (!ipc.register(.input, endpoint)) {
_ = system.write("input: register failed\n"); _ = system.write("/system/services/input: register failed\n");
return; return;
} }
_ = system.write("input: ready\n"); _ = system.write("/system/services/input: ready\n");
var reply_buffer: [protocol.reply_size]u8 = undefined; var reply_buffer: [protocol.reply_size]u8 = undefined;
var reply_len: usize = 0; var reply_len: usize = 0;
+68
View File
@@ -0,0 +1,68 @@
//! The power protocol (docs/m21-plan.md): system power's domain-named surface,
//! registered under `ServiceId.power`. On x86 the acpi service serves it; on
//! ARM a PSCI/mailbox service will register the same id — subscribers never
//! learn which firmware they are on (m19-m20-plan.md decision 7). The
//! vfs-protocol pattern: extern-struct messages, a version, reserved fields.
/// The protocol version a client states nowhere yet — reserved for the day a
/// handshake needs it; requests carry it so a mismatch can be refused loudly.
pub const version: u16 = 1;
pub const Operation = enum(u8) {
/// Subscribe to power events: the subscriber's endpoint rides as the
/// call's capability (the input/device-manager pattern); events arrive on
/// it as buffered messages carrying an `EventMessage`.
subscribe = 1,
/// Orderly shutdown's last step: enter S5. Accepted only from PID 1
/// (init) — the process that has already run the stop sequence over
/// everything else.
shutdown = 2,
/// The published event payload (never sent *to* the service).
event = 3,
};
/// What happened. The vocabulary is hardware-neutral: a lid is a lid whether
/// ACPI or a PSCI mailbox reported it.
pub const Event = enum(u8) {
power_button = 1,
lid = 2,
ac = 3,
battery = 4,
/// A device notification that maps to none of the named events — the
/// `code` and `hid` fields say which device and what code.
notify = 5,
};
pub const Subscribe = extern struct {
operation: u8 = @intFromEnum(Operation.subscribe),
reserved0: u8 = 0,
version: u16 = version,
reserved1: u32 = 0,
};
pub const Shutdown = extern struct {
operation: u8 = @intFromEnum(Operation.shutdown),
reserved0: u8 = 0,
version: u16 = version,
reserved1: u32 = 0,
};
/// A published event, as the buffered-message payload subscribers receive.
pub const EventMessage = extern struct {
operation: u8 = @intFromEnum(Operation.event),
/// An Event value.
event: u8,
reserved0: u16 = 0,
/// The device notification code (Notify's second argument), or 0.
code: u32 = 0,
/// The notifying device's hardware id (EISA-decoded), or all zero.
hid: [8]u8 = .{0} ** 8,
};
pub const Reply = extern struct {
status: i32,
reserved: u32 = 0,
};
/// Upper bound on any message in this protocol — sizes endpoint buffers.
pub const message_maximum = 64;
@@ -48,11 +48,92 @@ fn awaitChildExit(endpoint: runtime.ipc.Handle) u32 {
return received.childProcessId(); return received.childProcessId();
} }
/// The harness-run child of the signals test: echoes requests, logs the two
/// signals it handles. Terminate makes run() return, and returning from main is
/// the clean exit the parent reads as ExitReason.exited.
fn echo(message: []const u8, reply: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = sender;
_ = capability;
const n = @min(message.len, reply.len);
@memcpy(reply[0..n], message[0..n]);
return n;
}
fn onReload() void {
_ = runtime.system.write("process-test: reloaded\n");
}
fn onTerminate() void {
_ = runtime.system.write("process-test: terminating\n");
}
/// The parent of the signals test: drives ping, echo, reload, the one-shot
/// timer, and both endings of the stop sequence (polite -> exited; deaf ->
/// killed at the deadline). Prints "process-test: signals ok" as the marker.
fn signalRun() void {
const endpoint = runtime.ipc.createIpcEndpoint() orelse fail("create exit endpoint");
const child = runtime.system.spawnSupervised("process-test", &.{"service"}, endpoint) orelse fail("spawn service child");
// Reach the child's endpoint through the registry (retry: it may not be up).
var service_handle: ?runtime.ipc.Handle = null;
var tries: u32 = 0;
while (service_handle == null and tries < 200) : (tries += 1) {
service_handle = runtime.ipc.lookup(.input);
if (service_handle == null) runtime.system.sleep(20);
}
const h = service_handle orelse fail("service child never registered");
// The universal ping: a zero-length call answered zero-length by the harness.
var reply: [16]u8 = undefined;
const pong = runtime.ipc.call(h, &.{}, &reply) catch fail("ping call failed");
if (pong != 0) fail("ping reply not empty");
// An ordinary request still reaches on_message.
const n = runtime.ipc.call(h, "echo!", &reply) catch fail("echo call failed");
if (n != 5 or !std.mem.eql(u8, reply[0..5], "echo!")) fail("echo mismatch");
// reload: a statement — the child logs it; the kernel test reads the serial.
if (!runtime.process.sendSignal(child, .reload)) fail("send reload");
runtime.system.sleep(200);
// The one-shot timer: armed on our endpoint, lands as isTimer.
if (!runtime.system.timerOnce(endpoint, 100)) fail("arm timer");
var scratch: [8]u8 = undefined;
const landing = runtime.ipc.replyWait(endpoint, scratch[0..0], &scratch, null);
if (!landing.isTimer()) fail("expected the timer landing");
// The stop sequence, polite path: terminate, clean exit inside the deadline.
runtime.process.stop(child, 2000, endpoint);
if ((runtime.process.exitReason(child) orelse .killed) != .exited) fail("service child reason not exited");
// The deaf child: binds nothing, hears nothing — the deadline kills it.
const deaf = runtime.system.spawnSupervised("process-test", &.{"sleeper"}, endpoint) orelse fail("spawn deaf child");
runtime.system.sleep(50); // let it reach its sleep
runtime.process.stop(deaf, 300, endpoint);
if ((runtime.process.exitReason(deaf) orelse .exited) != .killed) fail("deaf child reason not killed");
_ = runtime.system.write("process-test: signals ok\n");
}
pub fn main(init: runtime.process.Init) void { pub fn main(init: runtime.process.Init) void {
const role = init.arguments.get(1) orelse return; // spawned bare (ramdisk sweep): stay silent const role = init.arguments.get(1) orelse return; // spawned bare (ramdisk sweep): stay silent
if (std.mem.eql(u8, role, "sleeper")) { if (std.mem.eql(u8, role, "sleeper")) {
while (true) runtime.system.sleep(500); while (true) runtime.system.sleep(500);
} }
if (std.mem.eql(u8, role, "service")) {
// Borrowed well-known id: the input service is not part of this scenario.
runtime.service.run(64, .{
.service = .input,
.on_message = echo,
.on_reload = onReload,
.on_terminate = onTerminate,
});
return; // terminate arrived; returning is the clean exit
}
if (std.mem.eql(u8, role, "signal-run")) {
signalRun();
return;
}
if (std.mem.eql(u8, role, "spinner")) { if (std.mem.eql(u8, role, "spinner")) {
var beat: u64 = 0; var beat: u64 = 0;
const touch: *volatile u64 = &beat; const touch: *volatile u64 = &beat;
@@ -89,6 +170,12 @@ pub fn main(init: runtime.process.Init) void {
if (listed(sleeper, "process-test")) fail("sleeper still listed after kill"); if (listed(sleeper, "process-test")) fail("sleeper still listed after kill");
if (listed(spinner, "process-test")) fail("spinner still listed after kill"); if (listed(spinner, "process-test")) fail("spinner still listed after kill");
// M17.2: both children were killed by us, and the reason says so — the whole
// restart-policy input, read through the runtime like a real supervisor would.
if ((runtime.process.exitReason(sleeper) orelse .exited) != .killed) fail("sleeper reason not killed");
if ((runtime.process.exitReason(spinner) orelse .exited) != .killed) fail("spinner reason not killed");
if (runtime.process.exitReason(0xFFFF_FFF0) != null) fail("unknown id had a reason");
_ = runtime.system.write("process-test: ok\n"); _ = runtime.system.write("process-test: ok\n");
} }
+21 -1
View File
@@ -6,10 +6,30 @@
const std = @import("std"); const std = @import("std");
const runtime = @import("runtime"); const runtime = @import("runtime");
pub fn main() void { pub fn main(init: runtime.process.Init) void {
const u = @import("posix").unistd; const u = @import("posix").unistd;
const payload = "hello-vfs"; const payload = "hello-vfs";
// The "park" role (the vfs-client-death test): open a file, then hold the
// handle forever without closing — the kill and the VFS's release-on-death
// are the point.
if (init.arguments.count > 1) {
var fd: i32 = -1;
var tries: u32 = 0;
while (fd < 0 and tries < 200) : (tries += 1) {
fd = u.open("parked", u.O_CREAT);
if (fd < 0) runtime.system.sleep(20);
}
if (fd < 0) {
_ = runtime.system.write("vfstest: park open failed\n");
return;
}
while (true) {
_ = runtime.system.write("vfstest: parked\n");
runtime.system.sleep(500);
}
}
// The VFS server may not have registered yet — retry open until it's up. // The VFS server may not have registered yet — retry open until it's up.
var fd: i32 = -1; var fd: i32 = -1;
var tries: u32 = 0; var tries: u32 = 0;
+54 -19
View File
@@ -23,6 +23,10 @@ const Node = struct {
const OpenFile = struct { const OpenFile = struct {
used: bool = false, used: bool = false,
node: usize = 0, node: usize = 0,
// The client (task id — an IPC badge is one) that opened this handle. What
// release-on-death sweeps by: a service must never depend on its clients
// cleaning up after themselves (docs/process-lifecycle.md).
owner: u32 = 0,
}; };
var nodes = [_]Node{.{}} ** 8; var nodes = [_]Node{.{}} ** 8;
@@ -65,8 +69,30 @@ fn fail(out: []u8) usize {
return writeReply(out, .{ .status = -1 }, &.{}); return writeReply(out, .{ .status = -1 }, &.{});
} }
/// Handle one request; write the reply into `out`, return its length. /// Format one whole log line and emit it in a single `debug_write`, so lines from
fn handle(message: []const u8, out: []u8) usize { /// concurrent processes can never land in the middle of it.
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [96]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
/// Release every open handle `client` held — called on that client's published
/// exit event. The nodes (the files) stay: ramfs contents outlive their writers,
/// only the dead client's handles go.
fn releaseClientHandles(client: u32) void {
var released: u32 = 0;
for (&opens) |*o| {
if (o.used and o.owner == client) {
o.used = false;
released += 1;
}
}
if (released != 0) writeLine("/system/services/vfs: released {d} handle(s) for dead client {d}\n", .{ released, client });
}
/// Handle one request from `sender`; write the reply into `out`, return its length.
fn handle(message: []const u8, out: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = capability;
if (message.len < protocol.request_size) return fail(out); if (message.len < protocol.request_size) return fail(out);
const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]); const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]);
const payload = message[protocol.request_size..]; const payload = message[protocol.request_size..];
@@ -77,7 +103,7 @@ fn handle(message: []const u8, out: []u8) usize {
const ni = findNode(name) orelse createNode(name) orelse return fail(out); const ni = findNode(name) orelse createNode(name) orelse return fail(out);
for (&opens, 0..) |*o, i| { for (&opens, 0..) |*o, i| {
if (!o.used) { if (!o.used) {
o.* = .{ .used = true, .node = ni }; o.* = .{ .used = true, .node = ni, .owner = sender };
return writeReply(out, .{ .status = 0, .node = i }, &.{}); return writeReply(out, .{ .status = 0, .node = i }, &.{});
} }
} }
@@ -113,27 +139,36 @@ fn handle(message: []const u8, out: []u8) usize {
} }
} }
pub fn main() void { /// Startup, under the harness: subscribe to the published exit events — when a
const endpoint = runtime.ipc.createIpcEndpoint() orelse { /// client dies holding open handles, the exit notification is how the VFS learns
_ = runtime.system.write("vfs: no endpoint\n"); /// to release them (docs/process-lifecycle.md).
return; fn initialise(endpoint: runtime.ipc.Handle) bool {
}; if (!runtime.process.subscribeExits(endpoint)) {
if (!runtime.ipc.register(.vfs, endpoint)) { _ = runtime.system.write("/system/services/vfs: exit subscription failed\n");
_ = runtime.system.write("vfs: register failed\n");
return;
} }
_ = runtime.system.write("vfs: ready\n"); _ = runtime.system.write("/system/services/vfs: ready\n");
return true;
}
var reply_buffer: [protocol.message_maximum]u8 = undefined; /// A non-signal notification: the only kind the VFS subscribes to is exit events.
var reply_len: usize = 0; fn onNotification(badge: u64) void {
var receive: [protocol.message_maximum]u8 = undefined; if (badge & runtime.ipc.notify_exit_bit != 0) {
while (true) { releaseClientHandles(@intCast(badge & ~(runtime.ipc.notify_badge_bit | runtime.ipc.notify_exit_bit)));
const got = runtime.ipc.replyWait(endpoint, reply_buffer[0..reply_len], &receive, null);
// Ignore notifications (none expected here); handle a request.
reply_len = handle(receive[0..got.len], &reply_buffer);
} }
} }
pub fn main() void {
// The harness owns the loop: requests dispatch to handle(), exit events to
// onNotification(), ping and terminate are answered for free — this service
// gained the whole lifecycle contract by deleting its hand-rolled loop.
runtime.service.run(protocol.message_maximum, .{
.service = .vfs,
.init = initialise,
.on_message = handle,
.on_notification = onNotification,
});
}
pub const panic = runtime.panic; pub const panic = runtime.panic;
comptime { comptime {
_ = &runtime.start._start; _ = &runtime.start._start;
+180 -31
View File
@@ -18,9 +18,11 @@ Usage:
""" """
import argparse import argparse
import json
import os import os
import re import re
import shutil import shutil
import socket
import subprocess import subprocess
import sys import sys
import time import time
@@ -55,20 +57,16 @@ ARCHES = {
"/opt/homebrew/share/qemu/edk2-i386-vars.fd", # macOS Homebrew (Apple Silicon) "/opt/homebrew/share/qemu/edk2-i386-vars.fd", # macOS Homebrew (Apple Silicon)
"/usr/local/share/qemu/edk2-i386-vars.fd", # macOS Homebrew (Intel) "/usr/local/share/qemu/edk2-i386-vars.fd", # macOS Homebrew (Intel)
], ],
# zig-out is a FHS-shaped image and the boot volume; the harness copies the # zig-out is itself the FHS-shaped boot volume (docs/efi.md): the build
# boot-critical files from their FHS paths into a fresh ESP with the same # installs BOOTX64.efi, the kernel, init, and the initial-ramdisk at their
# layout. (dest in ESP, source path under zig-out) — identical here. # boot paths. The harness presents zig-out to the guest directly — exactly
"efi_app": ("EFI/BOOT/BOOTX64.efi", "EFI/BOOT/BOOTX64.efi"), # as `zig build run-x86-64` does — so there is no separate ESP to assemble.
"kernel": ("system/kernel", "system/kernel"),
# The init user program and the initial-ramdisk (VFS server + drivers).
"extra": [("system/services/init", "system/services/init"),
("boot/initial-ramdisk.img", "boot/initial-ramdisk.img")],
# Built as a function so we can splice in per-run paths. # Built as a function so we can splice in per-run paths.
"qemu_args": lambda a, esp, vars_fd, serial: [ "qemu_args": lambda a, boot_volume, vars_fd, serial: [
"-machine", "q35", "-m", "128M", "-machine", "q35", "-m", "128M",
"-drive", f"if=pflash,format=raw,readonly=on,file={a['ovmf_code']}", "-drive", f"if=pflash,format=raw,readonly=on,file={a['ovmf_code']}",
"-drive", f"if=pflash,format=raw,file={vars_fd}", "-drive", f"if=pflash,format=raw,file={vars_fd}",
"-drive", f"format=raw,file=fat:rw:{esp}", "-drive", f"format=raw,file=fat:rw:{boot_volume}",
"-net", "none", "-net", "none",
"-vga", "none", "-device", "VGA,edid=on,xres=1280,yres=720", "-vga", "none", "-device", "VGA,edid=on,xres=1280,yres=720",
"-display", "none", "-display", "none",
@@ -84,7 +82,10 @@ ARCHES = {
# `expect`: a regex that must appear in serial output => pass. # `expect`: a regex that must appear in serial output => pass.
# `fail`: optional regex whose appearance => immediate fail. # `fail`: optional regex whose appearance => immediate fail.
CASES = [ CASES = [
# smoke also proves the QMP channel: the harmless query must be delivered
# (handshake + command) before the case may pass — see run_case.
{"name": "smoke", {"name": "smoke",
"qmp_after": {"delay": 2, "command": "query-status"},
"expect": r"DANOS-TEST-RESULT: PASS", "expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
{"name": "discovery", {"name": "discovery",
@@ -168,7 +169,7 @@ CASES = [
# Stress the big kernel lock across cores; heavier, so a longer timeout. # Stress the big kernel lock across cores; heavier, so a longer timeout.
{"name": "smp-stress", {"name": "smp-stress",
"smp": 4, "smp": 4,
"timeout": 90, "timeout": 150,
"expect": r"DANOS-TEST-RESULT: PASS", "expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# Retry: a forced first-wake failure must still bring every core online. # Retry: a forced first-wake failure must still bring every core online.
@@ -240,9 +241,130 @@ CASES = [
"smp": 4, "smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS", "expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# M17.1: a dead process's device claims are released by the reap — kill a child
# holding a claim, the device must be claimable again (process-lifecycle.md).
{"name": "claim-release",
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M17.3: published exit events — the VFS subscribes, a client dies holding an
# open handle, and the VFS releases it (process-lifecycle.md "Who learns of a death").
{"name": "vfs-client-death",
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M17.4: signals over IPC — ping, reload, terminate (clean exit), the one-shot
# timer, and the stop sequence's two endings, all driven from ring 3.
{"name": "signals",
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M18.2: bus tree reports — the xHCI driver scans its root-hub ports and
# reports both QEMU devices; the manager mirrors, prunes on the reporter's
# death, and the respawned driver re-reports (docs/device-manager.md).
{"name": "usb-report",
"smp": 4,
"timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"device-manager: child added[\s\S]*"
r"device-manager: child added[\s\S]*"
r"device-manager: test mode: killing the reporter[\s\S]*"
r"device-manager: child removed[\s\S]*"
r"device-manager: restarting usb-xhci-bus[\s\S]*"
r"device-manager: child added",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M20.1: the ring-3 AML parse (the acpi service maps the blobs and parses
# them) finds exactly the Device count the kernel's own parse produced.
{"name": "acpi-parse",
"smp": 4,
"timeout": 60,
"expect": r"acpi-parse: ok",
"fail": r"acpi-parse: mismatch|DANOS-TEST-RESULT: FAIL"},
# M20.3: the flip — ps2-bus now comes up from the acpi service's report, not
# a kernel-built node. Ordered: report -> spawn -> the driver attaches its
# keyboard, proving discovery runs entirely in ring 3 (docs/m19-m20-plan.md).
{"name": "acpi-ps2",
"smp": 4,
"timeout": 150,
"expect": r"acpi: reported PNP0303[\s\S]*"
r"device-manager: spawned ps2-bus[\s\S]*"
r"ps2-bus: keyboard driver attached",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M21.1: the SCI + power button. Boot the manager (which spawns the acpi
# service); ~4s in, QMP system_powerdown raises the ACPI power-button fixed
# event; the service's SCI handler must log the press (docs/m21-plan.md).
{"name": "power-button",
"smp": 4,
"timeout": 60,
"qmp_after": {"delay": 4, "command": "system_powerdown"},
"expect": r"power: button pressed",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M21.3 capstone: orderly shutdown. Boot init (the full tree comes up);
# ~5s in, QMP system_powerdown raises the power button; the acpi service
# publishes it, init stops its children then requests S5, and QEMU exits.
# The ordered regex proves button -> shutting-down -> entering-S5; the case
# passes on QEMU's self-exit through S5 (docs/m21-plan.md).
{"name": "orderly-shutdown",
"smp": 4,
"timeout": 90,
"qmp_after": {"delay": 5, "command": "system_powerdown"},
"expect": r"power: button pressed[\s\S]*"
r"init: shutting down[\s\S]*"
r"power: entering S5",
"fail": r"power: S5 write did not take|DANOS-TEST-RESULT: FAIL"},
# M20.2: the acpi service evaluates _CRS/_STA in ring 3 and registers +
# reports its _HID devices — the two PS/2 nodes must appear with resources
# (keyboard: io 0x60/0x64 + IRQ = 3; mouse: IRQ = 1) (docs/m19-m20-plan.md).
{"name": "acpi-report",
"smp": 4,
"timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"acpi: reported PNP0303 \(device \d+, 3 resources\)[\s\S]*"
r"acpi: reported PNP0F13 \(device \d+, 1 resources\)",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M19.1: the ring-3 PCI scan (pci-bus walks the ECAM through its mmio_map
# grant) finds exactly the functions the kernel's own walk recorded.
{"name": "pci-scan",
"smp": 4,
"timeout": 60,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M18.3: the application surface — device-list enumerates the tree over IPC,
# subscribes (endpoint as capability), and observes the removed/added events
# the reporter's test-kill produces (docs/device-manager.md).
{"name": "device-list",
"smp": 4,
"timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"device-list: \d+ devices[\s\S]*"
r"device-list: subscribed[\s\S]*"
r"device-manager: test mode: killing the reporter[\s\S]*"
r"device-list: removed \(device[\s\S]*"
r"device-list: added \(device",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M18.1: the device manager's hello + restart policy — xHCI hellos clean and
# stays; crash-test faults, is restarted with backoff (re-claiming its device
# each time), and hits the crash-loop cap (docs/device-manager.md).
{"name": "driver-restart",
"smp": 4,
"timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"usb-xhci-bus: hello acknowledged[\s\S]*"
r"device-manager: restarting crash-test[\s\S]*"
r"device-manager: crash-test is failing repeatedly",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# The initial_ramdisk: the loader ferries a bundle of user binaries; the kernel parses # The initial_ramdisk: the loader ferries a bundle of user binaries; the kernel parses
# it and spawns each as a ring-3 process (here the VFS-server stub heartbeats). # it and spawns each as a ring-3 process (here the VFS-server stub heartbeats).
{"name": "initial-ramdisk", {"name": "initial-ramdisk",
"timeout": 60, # the acpi service's boot-time SCI setup can push the marker past 30s under load
"expect": r"DANOS-TEST-RESULT: PASS", "expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# The user-space VFS: a client opens/writes/reads a file through the rt file # The user-space VFS: a client opens/writes/reads a file through the rt file
@@ -306,24 +428,6 @@ def build(arch, case):
return None return None
def make_esp(arch):
"""Assemble a fresh EFI System Partition from the freshly built binaries."""
esp = os.path.join(WORK, "esp")
if os.path.exists(esp):
shutil.rmtree(esp)
efi_dest, efi_src = arch["efi_app"]
kern_dest, kern_src = arch["kernel"]
fhs = os.path.join(REPO, "zig-out") # zig-out is the FHS image
os.makedirs(os.path.join(esp, os.path.dirname(efi_dest)), exist_ok=True)
os.makedirs(os.path.join(esp, os.path.dirname(kern_dest)), exist_ok=True)
shutil.copy(os.path.join(fhs, efi_src), os.path.join(esp, efi_dest))
shutil.copy(os.path.join(fhs, kern_src), os.path.join(esp, kern_dest))
for dest, src in arch.get("extra", []):
os.makedirs(os.path.join(esp, os.path.dirname(dest)), exist_ok=True)
shutil.copy(os.path.join(fhs, src), os.path.join(esp, dest))
return esp
def resolve_firmware(arch): def resolve_firmware(arch):
"""Collapse the ovmf_code/ovmf_vars candidate lists to the first path that """Collapse the ovmf_code/ovmf_vars candidate lists to the first path that
exists on this machine. Mutates `arch` in place; idempotent (a resolved exists on this machine. Mutates `arch` in place; idempotent (a resolved
@@ -342,12 +446,34 @@ def resolve_firmware(arch):
+ "\nInstall OVMF (edk2-ovmf / ovmf) or add its path above.") + "\nInstall OVMF (edk2-ovmf / ovmf) or add its path above.")
def qmp_send(path, command):
"""One QMP command: connect, capabilities handshake, execute. Raises on any
failure — the caller retries until the guest's socket is ready. This is how
a case injects a host-side event (system_powerdown = the ACPI power button)
into the running guest (docs/m21-plan.md)."""
sock = socket.socket(socket.AF_UNIX, socket.SOCK_STREAM)
sock.settimeout(5)
try:
sock.connect(path)
stream = sock.makefile("rw")
stream.readline() # the QMP greeting
stream.write(json.dumps({"execute": "qmp_capabilities"}) + "\n")
stream.flush()
stream.readline() # {"return": {}}
stream.write(json.dumps({"execute": command}) + "\n")
stream.flush()
stream.readline()
finally:
sock.close()
def run_case(arch, case): def run_case(arch, case):
err = build(arch, case["name"]) err = build(arch, case["name"])
if err: if err:
return False, "build failed:\n" + err return False, "build failed:\n" + err
esp = make_esp(arch) # zig-out is the FHS boot volume; hand it to the guest as-is (see qemu_args).
boot_volume = os.path.join(REPO, "zig-out")
vars_fd = os.path.join(WORK, "vars.fd") vars_fd = os.path.join(WORK, "vars.fd")
shutil.copy(arch["ovmf_vars"], vars_fd) shutil.copy(arch["ovmf_vars"], vars_fd)
serial = os.path.join(WORK, "serial.log") serial = os.path.join(WORK, "serial.log")
@@ -357,17 +483,32 @@ def run_case(arch, case):
expect = re.compile(case["expect"]) expect = re.compile(case["expect"])
fail = re.compile(case["fail"]) if case.get("fail") else None fail = re.compile(case["fail"]) if case.get("fail") else None
cmd = [arch["qemu"]] + arch["qemu_args"](arch, esp, vars_fd, serial) cmd = [arch["qemu"]] + arch["qemu_args"](arch, boot_volume, vars_fd, serial)
if case.get("smp"): # some cases need more than one core (e.g. parallelism) if case.get("smp"): # some cases need more than one core (e.g. parallelism)
cmd += ["-smp", str(case["smp"])] cmd += ["-smp", str(case["smp"])]
if case.get("qemu_extra"): # extra qemu args, e.g. -device intel-iommu for the IOMMU case if case.get("qemu_extra"): # extra qemu args, e.g. -device intel-iommu for the IOMMU case
cmd += case["qemu_extra"] cmd += case["qemu_extra"]
# A QMP control socket, always present (additive): how a case's `qmp_after`
# hook injects host-side events into the guest mid-run.
qmp_path = os.path.join(WORK, "qmp.sock")
if os.path.exists(qmp_path):
os.remove(qmp_path)
cmd += ["-qmp", f"unix:{qmp_path},server,nowait"]
qmp_after = case.get("qmp_after") # {"delay": seconds, "command": "..."}
qmp_sent = False
started = time.monotonic()
qemu = subprocess.Popen(cmd, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL) qemu = subprocess.Popen(cmd, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
try: try:
timeout = case.get("timeout", TIMEOUT) timeout = case.get("timeout", TIMEOUT)
deadline = time.monotonic() + timeout deadline = time.monotonic() + timeout
while time.monotonic() < deadline: while time.monotonic() < deadline:
time.sleep(0.2) time.sleep(0.2)
if qmp_after and not qmp_sent and time.monotonic() - started >= qmp_after["delay"]:
try:
qmp_send(qmp_path, qmp_after["command"])
qmp_sent = True
except OSError:
pass # socket not up yet; retry next tick
text = "" text = ""
if os.path.exists(serial): if os.path.exists(serial):
with open(serial, "r", errors="replace") as f: with open(serial, "r", errors="replace") as f:
@@ -375,6 +516,8 @@ def run_case(arch, case):
if fail and fail.search(text): if fail and fail.search(text):
return False, "hit failure marker" return False, "hit failure marker"
if expect.search(text): if expect.search(text):
if qmp_after and not qmp_sent:
continue # the hook must deliver before the case may pass
return True, "matched " + repr(case["expect"]) return True, "matched " + repr(case["expect"])
if qemu.poll() is not None: # QEMU exited on its own if qemu.poll() is not None: # QEMU exited on its own
if expect.search(text): if expect.search(text):
@@ -410,6 +553,12 @@ def main():
for case in selected: for case in selected:
print(f" {case['name']:<12} ... ", end="", flush=True) print(f" {case['name']:<12} ... ", end="", flush=True)
ok, detail = run_case(arch, case) ok, detail = run_case(arch, case)
if not ok:
# Keep the evidence: serial.log is otherwise overwritten by the
# next case, and an intermittent failure's log is unrecoverable.
source = os.path.join(WORK, "serial.log")
if os.path.exists(source):
shutil.copy(source, os.path.join(WORK, f"{case['name']}-failed-serial.log"))
print(("PASS" if ok else "FAIL") + f" ({detail})") print(("PASS" if ok else "FAIL") + f" ({detail})")
if not ok: if not ok:
failures += 1 failures += 1
+22
View File
@@ -0,0 +1,22 @@
#!/usr/bin/env bash
#
# sort-lines-group-by-start — cluster lines that share their first
# whitespace-separated field ($1). Keys appear in first-seen order, and lines
# within a key keep their original order. It groups; it does NOT sort.
#
# Pass the log file as an argument; result is written to stdout:
#
# tools/sort-lines-group-by-start.sh filename.log
#
# Useful for a serial/boot log where several sources interleave and each line is
# prefixed with its source (the first field): this pulls every source's lines
# back together, in the order the sources first appeared, without reordering
# within a source.
#
# input output
# pci-bus: scan start pci-bus: scan start
# acpi: reported PNP0303 pci-bus: 5 functions
# pci-bus: 5 functions acpi: reported PNP0303
# acpi: reported PNP0501 acpi: reported PNP0501
awk '{lines[$1] = lines[$1] ? lines[$1] ORS $0 : $0; if (!seen[$1]++) order[++count] = $1} END {for (i=1; i<=count; i++) print lines[order[i]]}' "$@"