Author SHA1 Message Date
Daniel Samson 446f655c69 Docs: close the M21 track (events + system power)
docs/m19-m20-plan.md's M21 preview now points at the completed plan.
2026-07-13 05:57:17 +01:00
Daniel Samson a785efa4a3 Orderly shutdown: init's stop cascade into ring-3 S5 (M21.3)
The capstone. init becomes a real supervisor: it spawns its boot
services supervised against one endpoint that also carries its signals, a
re-arming heartbeat timer, and the power events it subscribes to. On the
power button (or a terminate signal — same path) it logs the shutdown,
runs the M17 stop sequence over its children in reverse spawn order
(vfs last), then asks the power service for S5.

The acpi service honors a shutdown request from a power subscriber — init
is the one subscriber, a soft gate that stands in for 'only the system
supervisor may power off' and, unlike a PID-1 check, survives the test
harness where the kernel's idle tasks take the early ids. The power
service is mechanism (write S5); deciding when to shut down and stopping
everything else first is init's policy — the microkernel split applied to
poweroff.

The orderly-shutdown scenario injects a real QMP power-button event and
watches the whole chain compose: button pressed -> init shutting down ->
entering S5 -> QEMU powers off. That single scenario proves the M17
lifecycle and the M21 event side compose into a clean shutdown. Suite
60/60.
2026-07-13 05:56:58 +01:00
Daniel Samson 767a2a9a7c Notify dispatch and GPE handlers (M21.2)
The AML interpreter now handles the Notify opcode (0x86, previously
unhandled): it resolves the target device, evaluates the code, and
records the pair in a bounded per-evaluate queue the caller drains with
takeNotifications. A host unit test with hand-encoded AML — a method that
issues Notify(DEV_, 0x80) — proves the device and code come back; aml.zig
joins the zig build test loop so the interpreter is covered on the host.

The acpi service's SCI handler now services general-purpose events too:
for each set-and-enabled GPE bit it evaluates the \_GPE._Lxx (level) or
_Exx (edge) handler method, drains the Notify queue that produced, and
publishes a domain event per notified device — PNP0C0A battery, ACPI0003
AC, PNP0C0D lid, else generic notify — then clears the status bit and
acks. The embedded controller's _Qxx queries are out of scope (hardware
track). QEMU raises no GPEs on this config, so the QEMU suite is the
regression net (the power button still works with GPE servicing in the
path); correctness is the unit test. Suite 59/59.
2026-07-13 05:44:58 +01:00
Daniel Samson 1f2c60b3ec The power button, in ring 3: SCI bound, fixed event published (M21.1)
The kernel publishes the FADT as one more acpi-tables memory resource
(tagged by its intact FACP header — the AML blobs are header-stripped);
the acpi service reads the PM1 event/control and GPE register ports from
that copy, so the kernel's own FADT parse is untouched. A power-protocol
module (ServiceId.power = 5, domain-named so an ARM PSCI service can serve
the same id) carries subscribe / shutdown / events.

The acpi service converts to runtime.service.run — device discovery, the
.power protocol, and the SCI notification all fold into one loop. At
startup it enables ACPI mode if SCI_EN is clear (the SMI dance), binds the
SCI (found as the node's len-1 irq resource, distinct from the broad
window), and sets PWRBTN_EN. On the SCI it reads PM1_STS, clears
PWRBTN_STS write-1, logs the press, publishes power_button to
subscribers, and always acks. The power-button scenario proves it with a
real QMP system_powerdown injected mid-run through the M21.0 channel.
2026-07-13 05:38:03 +01:00
Daniel Samson dfc7d6a609 The harness grows a QMP channel (M21.0)
Every case now gets a -qmp unix socket (additive; no case notices). A
minimal client does the capabilities handshake and executes one command;
the per-case qmp_after hook sends it N seconds after boot, retrying until
the guest's socket is up. A case with a hook configured cannot pass until
the hook delivered — and the smoke case now carries a harmless
query-status hook, so the channel is proven end to end on every run.
This is how the power scenarios inject the real ACPI power-button event
(system_powerdown) in M21.1 and M21.3.
2026-07-13 05:22:53 +01:00
Daniel Samson 738f6aa697 Make sort-lines-group-by-start.sh a runnable script
It was a bare awk snippet starting with `|`, meant to be pasted into a
pipeline. Turn it into an executable script that takes the log file as an
argument (tools/sort-lines-group-by-start.sh filename.log) and document its
behaviour and usage in a header comment.
2026-07-13 05:15:56 +01:00
Daniel Samson 01e56e3f36 Plan M21: ACPI events + system power 2026-07-13 05:13:51 +01:00
Daniel Samson d5d15cefcb Decode PCI/ACPI device identities and name their class codes as enums
Two related changes to make device identities legible in the boot log and in
the code that matches on them.

Logging: the pci-bus driver decodes each function's class/subclass/prog-IF
triple to human names (via the existing pci-class module), and the acpi
service appends each _HID's human name (via acpi-ids) to its report line. So
"class 0x01 (Mass Storage Controller) subclass 0x06 (Serial ATA Controller)
progif 0x01 (AHCI 1.0)" reads straight off the log when writing a driver.

Naming: a new coding standard ("Named values, not magic numbers") says a value
with meaning gets a name, prefer an enum for value sets. Applied:
- pci-class is refactored from u8-switch tables into a BaseClass enum plus
  per-class SubClass/ProgIf enums with name() methods (the usb-ids shape). The
  public className/subclassName/progIfName(u8...) API is unchanged, so the
  hardware-byte decoders (pci-bus, the kernel dump) are untouched; output is
  byte-identical.
- the device-manager builds the xHCI class triple from named parts instead of
  a bare 0x0C0330.
- the acpi service's _CRS walk names its resource-descriptor tags as
  SmallResourceType/LargeResourceType enums, and the _HID integer decode uses
  the AML module's existing *_opcode constants (now re-exported from aml.zig)
  rather than bare 0x0A/0xFF/... literals.
2026-07-13 05:05:25 +01:00
Daniel Samson fd96a35eb9 Decode the xHCI port speed in the usb-xhci-bus log
The root-hub scan logged the raw PORTSC port-speed class ("speed class 3").
Decode it to a human name — Low/Full/High/SuperSpeed/SuperSpeedPlus with the
USB generation and line rate — so the boot log says what enumerated on each
port, the USB analog of the pci-bus class line. This is the link speed only;
the device class/subclass/protocol needs descriptor reads (the USB track).
2026-07-13 05:05:13 +01:00
Daniel Samson e3fe3f3f45 Boot zig-out directly in the qemu test harness
The FHS-shaped zig-out IS the boot volume (docs/efi.md), and `zig build
run-x86-64` already presents it to the guest with fat:rw:zig-out. The test
harness instead assembled a separate ESP by copying the boot-critical files
out of zig-out into zig-out/qemu-test/esp — but every (dest, src) pair was
identical, so the copy was pure redundancy.

Drop make_esp and point QEMU straight at zig-out, matching run-x86-64 and the
docs. Removes the now-dead efi_app/kernel/extra arch-config entries.
2026-07-13 05:05:08 +01:00
Daniel Samson 60da667b42 Merge claude/vigilant-swanson-073c72: retire dead kernel AML device-building path (M20.3 cleanup) 2026-07-13 03:54:06 +01:00
Daniel Samson 36145e623b Delete the retired kernel AML device-building path (M20.3 cleanup)
The M20.3 flip moved ACPI namespace enumeration to the ring-3 acpi
service; the kernel now builds the namespace only for the \_S5 sleep
type. That left the kernel's AML-to-device helpers unreferenced.

Remove the dead cluster (wireAcpiDevices, mirrorDevices, applyHid,
setEisaHid, applyCrs, parseResourceTemplate, parseAddressSpace,
devicePresent, matchHostBridge, findPciNode, readAdr, isPciRootNode,
isPciRootHid, PciContext) and every AML-decoding helper it alone used
(eisaIdToStr, seg4, cstr, hexDigit, rd16, rd32, readN, readLE,
readIntObj, packageLength/PkgLen) plus their tests and the now-orphaned
acpi-ids import. The static-table path keeps checksumOk, fadt, readGas,
readCntRegister, and rd. Also tidies two stale comments.
2026-07-13 03:53:04 +01:00
Daniel Samson 565415327d Mark the M19-M20 discovery migration complete 2026-07-13 03:33:00 +01:00
Daniel Samson bf6bdb389d Merge feat/acpi-service: ACPI interpretation in ring 3 (M20)
The AML interpreter as a shared build module, the acpi-tables node, the
acpi service (parse, evaluate _CRS/_STA, register + report), and the flip
that retired the kernel's ACPI device build — discovery's second and final
subsystem to leave ring 0.
2026-07-13 03:32:54 +01:00
Daniel Samson 0628944b15 Docs: close the discovery migration (M20.3)
discovery.md records ACPI enumeration leaving the kernel; device-manager.md
increment 8 marked done — enumeration now runs entirely in ring 3.
2026-07-13 03:32:53 +01:00
Daniel Samson e6d0bb7ef0 The flip: ACPI enumeration leaves the kernel (M20.3)
The kernel no longer folds AML Device objects into the device tree — the
ring-3 acpi service is the sole builder of _HID device nodes. The kernel
keeps building the namespace only for the \_S5 sleep type, and still
seeds the static tables (MADT, HPET, MCFG, FADT) and the acpi-tables node.

The device manager matches ps2-bus from the service's _HID reports
(PNP0303 / PNP0F13, singleton-deduped) instead of boot-snapshot nodes;
its dead boot-snapshot ps2 arm is gone. The service registers every
device before reporting any, so a driver the manager spawns on the first
report already sees the full set — no keyboard-before-mouse race. The
acpi-ps2 scenario proves the whole chain: report -> spawn -> ps2-bus
finds the controller and attaches its keyboard, entirely in ring 3. The
ioport test moved to the acpi-tables I/O window, since the kernel-built
PS/2 node it used to scan for no longer exists. The retired
device-building functions in acpi.zig are dead but retained (a botched
mechanical deletion is worse mid-migration than a follow-up sweep, which
is flagged as a task). Suite 58/58.
2026-07-13 03:32:15 +01:00
Daniel Samson 5ca804d827 The acpi service evaluates _CRS/_STA in ring 3 and reports devices (M20.2)
AML method evaluation now runs in userspace touching real hardware: the
service builds an interpreter with a ring-3 Hal (port I/O routed through
its claimed acpi-tables node; a scratch page backs SystemMemory maps so a
stray OperationRegion degrades to zeros instead of faulting a process
that cannot map arbitrary physical memory). It walks the namespace and,
for each present _HID device that is not a PCI root, evaluates _CRS,
registers it under acpi-tables, and reports it with its EISA-decoded hid.

Containment for this needed the broker's irq check to become range-based
— an interrupt line is still indivisible, but a parent may own a range,
so the acpi-tables node's broad irq window contains its children's legacy
lines (a length-1 range is exactly the old equality, so single-irq
parents are unaffected). ChildAdded gained a hid field for firmware
string identity. Matching those reports to drivers stays off until M20.3,
so ps2-bus still comes up via the kernel path — no regression. The
acpi-report scenario proves the PS/2 keyboard (io 0x60/0x64 + IRQ) and
mouse (IRQ) are reported with their resources. Suite 57/57.
2026-07-13 03:19:39 +01:00
Daniel Samson a299363b59 The AML interpreter runs in ring 3: the acpi service parses (M20.1)
The AML module becomes a build module compiled into both the kernel (for
the \_S5 sleep state it still needs) and the new acpi service — one
source, two builds, no fork. The kernel publishes a single acpi-tables
node: the DSDT/SSDT blobs as memory resources, a broad io_port grant (the
honest trust boundary — firmware AML names whatever ports it chose, known
only after parsing), and the SCI for the M21 event track. The acpi
service claims the node, maps each blob through the ordinary mmio grant
(which preserves the sub-page offset onto the bytecode), and runs the
same parser the kernel does. It self-verifies its namespace Device count
against the kernel's — 34 = 34 — deterministically via an argv the
acpi-parse test passes, so no racing the shared serial buffer. Parse-only
touches no hardware; OperationRegion evaluation waits for _CRS/_STA in
M20.2. The manager spawns 'discovery' (the neutral ramdisk name) at
startup. Suite 56/56.
2026-07-13 03:07:11 +01:00
Daniel Samson d8dd62c639 Mark the feat/pci-bus merge done in the M19-M20 plan 2026-07-13 02:55:03 +01:00
Daniel Samson d106b6e8dc Merge feat/pci-bus: PCI enumeration in ring 3 (M19)
The host bridge apertures, idempotent device_register, the pci-bus driver
(scan, register, report), and the flip that retired the kernel's PCI walk
— discovery's first subsystem to leave ring 0.
2026-07-13 02:55:03 +01:00
Daniel Samson af2c766f42 The flip: PCI enumeration leaves the kernel (M19.3)
enumeratePci, addBars, pciConfigurationPtr, and the PciHeader struct are
deleted; the kernel seeds only the host bridge, and the ring-3 pci-bus
driver's reports are the sole source of PCI function nodes. The manager
matches PCI drivers from reported identity, deduped by registered device
id so a bus restart never double-spawns.

The flip did its job by exposing a latent SMP race: ring-3
device_register made the broker table concurrent for the first time, and
mmio_map read it lock-free — under load a torn resource length mapped
hpet's window wrong (its user fault) and underflowed r.len-1 into a
kernel integer-overflow panic. Fixed: the broker read in mmio_map (and
claim) runs under the big kernel lock, the arithmetic rejects
zero-length and wrapping windows cleanly, and pci-bus no longer registers
unimplemented size-0 BARs. driver-restart hammered 6x, suite 55/55.
2026-07-13 02:54:50 +01:00
Daniel Samson d26262bf56 pci-bus registers and reports what it scans (M19.2)
Each function is registered under the bridge with the config-space slice
and BARs sized by the same all-ones probe the kernel uses — byte-for-byte
equal descriptors, so the idempotent register returns the kernel's
existing node ids during coexistence instead of duplicating the tree.
The bridge gained the 16-bit io_port aperture that functions' I/O BARs
need to pass containment. Reports carry the registered device_id, and
the pci-scan scenario drills a forced restart: kill the enumerator after
its reports, watch the respawn re-scan, and assert the broker's PCI node
count never grew. The usb-restart test trigger is pinned to the xHCI
reporter (pci-bus racing it to two reports used to steal the kill).
Harness hardening: failing cases preserve their serial logs; the heavy
scenarios run at 150s.
2026-07-13 02:22:19 +01:00
Daniel Samson 10b89c06ff The PCI scan from ring 3: pci-bus walks the ECAM it mapped (M19.1)
The manager matches the pci_host_bridge node and spawns pci-bus with the
bridge id as its assignment — hello, supervision, restart, all the M18
contract for free. The driver claims the bridge, maps the ECAM window
(resource 0) through the ordinary mmio grant, and repeats the kernel's
brute-force bus/device/function walk from user space. The pci-scan
scenario builds its expected marker from the kernel's own function count,
so the two enumerations must agree exactly — the equivalence that
licenses retiring the kernel walk in M19.3.
2026-07-13 02:05:32 +01:00
Daniel Samson a2a05d0b3d Discovery-migration prerequisites (M19.0)
The host bridge now carries MMIO apertures derived from the boot memory
map's gaps below 4 GiB (largest three, sort-merged; a single after-the-
last-region hole dies on OVMF's flash at the top) plus one aperture above
the described space — so a user-space device_register of PCI functions
with BAR resources can pass containment. The discovery test asserts
every PCI memory resource lies inside a bridge window and names any
escapee. device_register is idempotent on exact (parent, class, identity,
resources) match — a restarted registering bus cannot duplicate its
children; proven directly against the broker in the bus test.
ChildAdded gains device_id so a report can carry the registered kernel
id a matched driver needs as its assignment.
2026-07-13 01:59:32 +01:00
Daniel Samson 75d62660b0 Bring the docs up to the M17-M18 reality; pre-settle M19-M20 ambiguities
resilience.md: steps 1-4 of the ladder are built — supervision, exit
reasons, restart with backoff, crash-loop caps, all proven by scenario;
what remains is scope, not mechanism. README statuses follow. drivers.md
gains the driver-contract section (harness, hello, crash-freely). The
M19-M20 plan pre-settles three things the loop would otherwise have had
to decide alone: the memory-map pass-through for apertures, the manager
spawning 'discovery' from M20.1, and hid[8] riding ChildAdded for ACPI
string identity until the FDT widening.
2026-07-13 01:49:52 +01:00
Daniel Samson a53c2b0193 Placeholder discovery services and the -Ddiscovery build option
system/services/acpi and system/services/fdt exist as documented
placeholders (silent clean-exit mains; the headers say exactly what each
becomes and why). The build's -Ddiscovery=acpi|fdt option fills the
ramdisk's neutral 'discovery' slot — the device manager will spawn
"discovery" by that name in M20.3 and never learn which firmware it is
on (m19-m20-plan.md decision 7). x86 defaults to acpi; the aarch64
target flips the default when it lands.
2026-07-13 01:46:04 +01:00
Daniel Samson bf481c080c Record the firmware-neutrality contract as decision 7
Discovery is one swappable process per firmware (acpi service on x86, an
fdt service on the Pis); everything at and above the device-manager
protocol stays generic. The manager owns the tree as data and never
touches hardware — firmware bytecode runs in a crashable, supervised
discoverer. Flagged now: hid[8] cannot hold an FDT compatible string,
and cross-firmware protocols are named by domain (power, not ACPI).
2026-07-13 01:33:12 +01:00
Daniel Samson 3a78dcab3f Scope ACPI events and system power as M21; record the SCI on acpi-tables
Battery, AC, lid, and the power button ride the acpi service as reported
children with small class drivers — the xHCI split repeated. QEMU can
only prove the power-button path (system_powerdown injects the real fixed
event), so battery/EC are interface-complete and hardware-validated on
the laptop. Per-device power states (D-states, suspend/resume) stay out
of scope: suspend has the shape of a lifecycle signal every driver must
answer, and it has no consumer until laptop sleep.
2026-07-13 01:21:23 +01:00
Daniel Samson 470f93a83d Plan the discovery migration (M19 pci-bus, M20 acpi service) 2026-07-13 01:15:17 +01:00
Daniel Samson 7798706b41 Mark the M17-M18 plan complete 2026-07-13 00:49:04 +01:00
Daniel Samson ad40de03c2 Merge feat/usb-xhci-bus: xHCI port scan, tree reports, and the app surface (M18.2-M18.3) 2026-07-13 00:49:04 +01:00
Daniel Samson d8778b4b70 The application surface: enumerate, subscribe, and device-list (M18.3)
Applications ask the device manager for the tree (enumerate: a header
plus ChildEntry records) and subscribe to published add/remove events by
handing their endpoint over as the call's capability — the input-service
pattern; events are the same ChildAdded/ChildRemoved structs the bus
drivers send, one encoding in both directions. device-list is the first
client: it prints the tree, subscribes, and narrates the events through
a driver restart. The protocol's message maximum is capped at the
kernel's IPC MESSAGE_MAXIMUM (256 bytes, ten entries per reply; paging
joins the protocol when a tree outgrows one message). The startUserTask
debug print is gone: it wrote to serial unserialized against user-space
lines and sheared concurrent log markers in half — the root cause of the
scenario flakes.
2026-07-13 00:49:03 +01:00
Daniel Samson 79d859a111 The xHCI driver scans its root-hub ports and reports the tree (M18.2)
child_added/child_removed join the device-manager protocol. The driver
maps its register BAR (resource 0 is the ECAM config space; the walk
starts at 1), reads CAPLENGTH and HCSPARAMS1, and reads one PORTSC per
port: the connect bit and speed class come straight from hardware, no
rings needed to see the devices. The manager mirrors reported children
keyed by (parent, port), remembers which instance reported each, and
prunes a dead reporter's children before deciding the restart — the
children describe protocol state that died with the process. The
usb-report scenario drives the whole loop: two QEMU devices reported,
reporter killed, children pruned, driver respawned with backoff, and the
new instance re-claims, re-scans, and re-reports.
2026-07-13 00:28:29 +01:00
Daniel Samson 37fb09f75e Mark the feat/device-manager merge done in the M17-M18 plan 2026-07-13 00:19:32 +01:00
Daniel Samson 34ebeb968d Merge feat/device-manager: the supervising device manager (M18.1) 2026-07-13 00:19:32 +01:00
Daniel Samson 3cc1d38dd0 The device manager supervises: hello, backoff, and the crash-loop cap (M18.1)
The manager is now a harness service on the well-known .device_manager
endpoint. Every driver spawns supervised; drivers with an assignment must
hello (device-manager-protocol, versioned) within a deadline enforced by
a timer sweep. Exit reasons drive the restart decision: clean exits stay
down, faults restart with 300/600/1200ms backoff, and three fast deaths
mark a driver failed instead of respawning forever. usb-xhci-bus is the
first conforming driver; the crash-test fixture claims a device, hellos,
and faults on purpose — each respawn re-proving claim release on death
through the manager's own path. maximum_tasks grows 16 -> 32: the
initial-ramdisk sweep (15 binaries at once) was intermittently
overflowing the static pool.
2026-07-13 00:19:30 +01:00
Daniel Samson 36e804b848 Mark the feat/process-lifecycle merge done in the M17-M18 plan 2026-07-12 23:53:50 +01:00
Daniel Samson be83a42d42 Merge feat/process-lifecycle: the process lifecycle (M17.1-M17.4)
Claim release on death, exit reasons, published exit events with the VFS
as first subscriber, signals over IPC with one-shot timers and the
service harness — docs/process-lifecycle.md increments 1-4, all built.
2026-07-12 23:53:50 +01:00
Daniel Samson 650a1b1595 Signals over IPC, one-shot timers, and the service harness (M17.4)
Signals are statements delivered as coalescing notifications to the
endpoint a process nominates with signal_bind — never a hijacked stack,
never a question (liveness is the zero-length ping the harness answers).
process_signal is supervisor-or-self gated, like kill; unbound targets
accumulate a pending mask delivered on bind. timer_bind is the missing
timed wait: a one-shot deadline landing in the same replyWait as
everything else — what stop(), hello deadlines, and restart backoff are
built from. runtime.service.run folds requests, signals, and
notifications into callbacks; the VFS conversion deletes its hand-rolled
loop and gains the whole lifecycle contract. The signals scenario drives
ping, reload, terminate->exited, the timer, and the deaf-child
deadline->killed path from ring 3. docs/process-lifecycle.md increments
1-4 are now as-built.
2026-07-12 23:53:38 +01:00
Daniel Samson d8c55c6f2f Publish exit events to subscribers; the VFS releases dead clients' handles (M17.3)
process_subscribe adds an endpoint to a bounded, ref-counted subscriber
table; every death posts the same badge encoding a supervisor's exit
notification uses, equally late, so subscribers observe a fully-released
child. A dying subscriber's own subscriptions are removed first — it never
hears about itself. The VFS is the first subscriber: open handles now
record their owner and are swept when the owner dies, because a service
must never depend on clients cleaning up after themselves
(docs/process-lifecycle.md). Proven by the vfs-client-death scenario.
2026-07-12 23:41:44 +01:00
Daniel Samson 2ebfb0c3b0 Record and expose how every process ends (M17.2)
The kernel records an ExitReason at all three death sites — clean exit,
fault (classified by vector), and process_kill — into a bounded ring
before the exit notification posts, so a supervisor's query never races
the notice. process_exit_reason is gated by the same supervisor check as
kill; runtime.process.exitReason is the stable interface. This is the
input restart policy reads (docs/process-lifecycle.md iron rule 2).
2026-07-12 23:34:09 +01:00
Daniel Samson 888eaa74e1 Release a dead process's device claims (M17.1)
Every path out of a process (exit, fault, kill) now releases its device
claims alongside its IRQ and MSI bindings, so a restarted driver can claim
its hardware again — the cleanup half of process-lifecycle.md's iron rule 1.
MSI vectors were already swept by irq.releaseOwner; claims were the gap.
The claim-release test proves kill -> release -> re-claim, plus the broker
release in isolation.
2026-07-12 23:23:49 +01:00
Daniel Samson ed76cbbc79 Mark Phase 0 done: baseline QEMU suite green (48/48) 2026-07-12 23:17:20 +01:00
Daniel Samson 140229b88d Rename usb-xhci-libary.zig to usb-xhci-library.zig (naming typo) 2026-07-12 23:13:33 +01:00
Daniel Samson cb2379fd06 Update README.md 2026-07-12 23:12:39 +01:00
Daniel Samson 1665b239b0 Add the status checklist and workflow to the M17-M18 plan 2026-07-12 23:04:43 +01:00
Daniel Samson 70ed0337f8 Merge feat/usb: USB wire ABI, xHCI detection and spawn, M17-M18 design 2026-07-12 22:56:35 +01:00
Daniel Samson 116b8f6c41 Design the process lifecycle and the device manager (M17-M18)
Signals over IPC (POSIX concepts, message delivery), published exit events,
the stable runtime.process interface, and the device manager as tree +
matcher + supervisor. All open questions settled; docs/m17-m18-plan.md is
the phase-by-phase execution plan.
2026-07-12 22:56:34 +01:00
Daniel Samson 77901bbba6 WIP: USB 2026-07-12 22:24:47 +01:00
Daniel Samson 78582d24d2 code lint 2026-07-12 19:36:10 +01:00
Daniel Samson 1cdffe21b1 fixing comments 2026-07-12 16:19:09 +01:00
Daniel Samson 4df90bc212 add tools/rewrap-comments.py 2026-07-12 16:18:56 +01:00
Daniel Samson 713e77354b gitattributes 2026-07-12 16:09:18 +01:00
Daniel Samson abb7b1b634 editorconfig 2026-07-12 16:09:11 +01:00
Daniel Samson f5f0e15769 zig fmt 2026-07-12 16:04:58 +01:00
Daniel Samson 8652b4a724 add qemu-xhci with usb-mouse and usb-kbd to build.zig 2026-07-12 01:31:38 +01:00
Daniel Samson e8233127c7 Install ps2-bus under its own name instead of clobbering bus
The ps2-bus executable was built with the artifact name "bus" — a
copy-paste from the generic bus driver's line above it. The initial
ramdisk was unaffected (the packer pairs names with binaries
explicitly), so the driver ran at boot; but the FHS install uses the
artifact's own name, so both drivers landed on
zig-out/system/drivers/bus, one overwriting the other, and
zig-out/system/drivers/ps2-bus never existed.
2026-07-11 23:41:08 +01:00
Daniel Samson 88e92254e9 Name ACPI hardware IDs instead of magic _HID strings
Turn acpi-ids.zig's flat name table into a HardwareId enum modeled on
ps2-library's Port: one entry() switch holds the registry (variant ->
_HID string + human-readable name), with hid(), description(), and
fromHid() methods. The free description(hid) lookup the kernel's
device-tree dump uses survives, implemented over the enum, and a new
test round-trips every variant through fromHid.

Callers now name the device instead of quoting its id:

- ps2-library's DeviceType.hid() and ps2-bus's descriptor lookups use
  HardwareId.ps2_keyboard / .ps2_mouse.
- device-manager's driverFor parses the HID once with fromHid and
  switches on named values.
- acpi.zig's isPciRootNode carried the same ids twice, as strings and
  as packed-EISA integers (0x030AD041/0x080AD041); both branches now
  decode to the string form and answer through one isPciRootHid helper
  using .pci_bus / .pci_express_root_bridge.
- build.zig threads the acpi-ids module (previously kernel-only) into
  every user binary, like xkeyboard-config.
2026-07-11 23:23:16 +01:00
Daniel Samson 5725d35e5b Name the attach reply statuses instead of magic numbers
Add an AttachStatus enum to ps2-library.zig for AttachReply.status,
distinguishing the three failure causes handleAttach previously
collapsed into a bare -1: invalid_request (message too short),
missing_endpoint (no capability passed), and no_such_device (no port
identified the requested device type). The keyboard and mouse drivers
check against AttachStatus.ok rather than a literal 0.
2026-07-11 23:11:45 +01:00
Daniel Samson 8c95525793 Wire real PS/2 mouse packets through to input events
Replace the mouse driver's synthetic stream with the real path, the
way the keyboard was wired:

- The auxiliary port's IRQ12 is enumerated on the mouse's own ACPI
  node (PNP0F13), and the kernel only lets a device's claimer bind or
  ack its IRQs — so ps2-bus now claims that node alongside the
  controller whenever port 2 carries a device, binds IRQ12 to its one
  endpoint, and re-arms whichever line the notification's badge names.
  The forwarding loop already routed auxiliary bytes by status bit 5.
- mouse-packet.zig (new, pure, host-tested): three-byte stream-mode
  packet assembly — bit-3 sync with resynchronization, ACK/BAT bytes
  dropped at packet start, nine-bit two's-complement movement,
  overflow packets discarded, and PS/2 positive-Y-up converted to the
  screen convention (positive down).
- mouse.zig mirrors the keyboard driver: no hardware claim, attaches
  to the bus as its mouse, and publishes button_down/button_up per
  changed button plus motion events with the pressed-button mask.

Verified end to end in QEMU via monitor mouse_move/mouse_button:
motion round-trips in screen coordinates, buttons transition with the
right mask, and keyboard events keep flowing alongside. The follow-up
is the IntelliMouse magic-knock for a scroll wheel (four-byte packets)
and scroll events.
2026-07-11 23:09:04 +01:00
Daniel Samson 5bba5d3363 Wire real PS/2 scancodes through to input events and characters
Replace the keyboard driver's synthetic stream with the real path:

- ps2-bus binds IRQ1 (interrupt bits set only after the bind), drains
  port 0x60 on each interrupt, and forwards every byte to the attached
  child driver over async ipc_send, routed by the status register's
  auxiliary-output bit. Children attach via the new well-known ps2_bus
  service, handing over their endpoint as a capability.
- scancode.zig (new, pure, host-tested): scancode set 2 -> USB HID
  usage decoding (F0/E0/E1 prefix state machine) plus keyboard state —
  pressed-key bitmap, typematic-repeat classification, modifier and
  caps-lock tracking.
- keyboard.zig decodes the forwarded stream and publishes real
  key_down/key_press/key_up events, filling key_press characters via
  xkeyboard-config (layout from argv[2], default us) and synthesizing
  ASCII control characters for Enter/Tab/Backspace/Escape.
- protocol.zig names the full HID usage set in Keycode; build.zig
  threads the xkeyboard-config module into user binaries.

Verified end to end in QEMU via monitor sendkey: shift, caps lock,
and control-character synthesis all decode correctly. The mouse
driver still publishes its synthetic stream; attaching it to the
bus the same way is the follow-up.
2026-07-11 22:48:59 +01:00
Daniel Samson 80b72db676 Removing DAN-INIT 2026-07-11 21:43:49 +01:00
Daniel Samson aa0c97353a fixing arguments 2026-07-11 21:39:27 +01:00
daniel dd93204b44 Merge pull request 'claude/input-module-keyboard-events-379361' (#6) from claude/input-module-keyboard-events-379361 into main
Reviewed-on: #6
2026-07-11 20:28:46 +00:00
daniel c7b17aaa0e Merge branch 'main' into claude/input-module-keyboard-events-379361 2026-07-11 20:28:21 +00:00
Daniel Samson d7a154a596 Add xkeyboard-config: X11 keyboard layouts compiled to Zig
Turn a keycode + modifiers into a character. The input module delivers HID
usage keycodes but nothing mapped them to characters; rather than hand-maintain
layout tables, vendor the X11 xkeyboard-config database and compile it to native
Zig at build time (no X11 runtime), the way make-initial-ramdisk.py packs the
ramdisk.

- tools/make-xkeyboard-config.py: `fetch` downloads the pinned xkeyboard-config
  release (2.44, sha256-verified), resolves the include graph for the configured
  layouts, and vendors only the reached symbols files + keysymdef.h + COPYING +
  PROVENANCE into library/xkeyboard-config/vendor/. `generate` parses that
  (keycodes via a HID->xkb-name table, symbols with include/augment/override and
  per-key type, keysymdef for keysym->Unicode) and emits generated/layouts.zig
  deterministically.
- library/xkeyboard-config/xkeyboard-config.zig: the API over the generated data
  — map(layout, hid_usage, mods) -> { keysym, character }, byName, and the
  level-selection semantics (the generated tables stay pure data). Host tests
  assert US letters/digits with Shift/Caps, GB £ vs US # on Shift+3, and French
  AZERTY q-where-US-has-a — the end-to-end proof of the parse->emit->lookup path.
- build.zig: `xkeyboard-config` + `layouts` modules, the test wired into
  `zig build test`, and a `zig build gen-xkeyboard-config` convenience step.
- Layouts: us, gb, de, fr, es, dvorak. Scope (documented): group 1, no dead-key
  composition, curated key types. Standalone library; wiring it into the input
  path to fill KeyEvent.character is a documented follow-up.

zig build test green (incl. the new keymap tests); regeneration is byte-identical;
full QEMU suite 48/48 (unaffected — no kernel/runtime/service change).
2026-07-11 15:57:09 +01:00
Daniel Samson 1bf91115dd Generalize input module to mouse and joystick/gamepad events
Extend the input service beyond the keyboard so mouse and joystick/gamepad
drivers can broadcast too, with per-device publish and subscribe methods.

- protocol: KeyEvent joins MouseEvent (motion/buttons/scroll) and
  JoystickEvent (axes/buttons), all carried in a common InputEvent envelope
  tagged with a DeviceKind. A subscribe request carries a device_mask, so a
  subscriber names the classes it wants and the service routes each event only
  to interested subscribers (a mouse-only listener never wakes for keystrokes).
- runtime: per-device publish methods (publishKeyboardEvent/publishMouseEvent/
  publishJoystickEvent) and subscribe helpers (subscribeKeyboard/Mouse/Joystick,
  each typed, plus subscribe(mask)/subscribeAll returning the tagged envelope).
- service: subscriber table gains a device_mask; broadcast routes by the
  event's device class.
- mouse driver now publishes (synthetic) mouse events like the keyboard driver;
  input-source cycles all three classes; input-test subscribes to all and only
  emits its "ok" marker once it has received one of each class — so the passing
  test proves per-device routing, not just delivery. Real HID decoding stays a
  follow-up.

No kernel changes: ipc_send is generic and the 36-byte InputEvent fits its
64-byte payload. Full QEMU suite 48/48; serial log confirms keyboard, mouse,
and joystick all reach one subscription.
2026-07-11 15:21:09 +01:00
Daniel Samson 65244e3103 Add input module: broadcast keyboard events over IPC
Programs can now subscribe to keyboard events (key_down/key_up/key_press)
and drivers can broadcast them, through a new user-space input service.

The delivery model is forced by danos IPC: a synchronous rendezvous holds
one pending reply, so a server cannot park N subscribers blocked in a
"wait for next event" call — delivery must be push. But a synchronous push
has no timeout and the kernel never wakes a sender parked on a dead peer's
endpoint, so one dying subscriber would hang all input. So this lands the
roadmap's planned asynchronous buffered send and builds the service on it:

- ipc_send (syscall 26): non-blocking post to an endpoint's bounded payload
  ring, delivered through reply_wait as a buffered message (notify_message_bit).
  A full ring drops the oldest. It can never hang on a dead/slow peer.
- input-protocol + runtime.input helpers (subscribe/next, connectSource/
  publish) — the first real consumer of M13 capability passing: a subscriber
  hands the service its own endpoint as a capability.
- input service (fan-out via ipc_send, dead-subscriber pruning), a synthetic
  input-source, and input-test; the ps2-bus keyboard driver publishes to it.
  Real IRQ1 scancode decoding (which must live in the bus, the PNP0303 owner)
  is a documented follow-up; the source is synthetic for now.
- build/init wiring, an `input` QEMU case, and docs/input.md.

Full QEMU suite 48/48, including the new input case and every IPC/endpoint
regression (ipc, ipc-call, ipc-cap, vfs, hpet, bus, irqfree).
2026-07-11 15:03:24 +01:00
Daniel Samson 75ccfff171 Pass arguments to main via runtime.process.Init, dispatched on signature 2026-07-11 14:22:48 +01:00
Daniel Samson 2a583d55a8 Finishing PS/2 bus driver 2026-07-11 14:12:33 +01:00
Daniel Samson d218d93f79 Add process management: enumerate, supervisor-gated kill, exit notifications
process_enumerate snapshots the task table (the device_enumerate shape, so
ps is a user program); system_spawn returns the child id, records the caller
as supervisor, and takes an exit endpoint; process_kill is allowed only for
the supervisor. Every death — exit, fault, or kill — posts a child-exit badge
to that endpoint (the IRQ-as-IPC pattern as SIGCHLD). A target caught off-CPU
is reaped in place; a running one is condemned and finished at its next
system call or tick, guarded so teardown never lands mid-kernel-operation.
Tested by process-list, process-kill, and supervision (a ring-3 supervisor
exercising the whole surface); design notes in docs/process-management.md.
2026-07-11 09:32:25 +01:00
Daniel Samson a5fe63c1dd Pass argv to processes on a SysV entry stack; grow the user stack to 32 KiB
Processes now start with C-compatible arguments: the kernel builds the
System V AMD64 entry block (argc, argv, empty envp, auxiliary vector)
at the top of the stack, argv[0] is the path or initial-ramdisk name
the process was spawned as, and system_spawn carries an optional
NUL-separated blob that becomes argv[1..]. The runtime parses the block
(runtime.argumentCount/argument) and its spawn wrappers pass arguments
through. The name is also recorded on the task, so a fault report says
which binary died, not just its id.

The user stack grows from one page to eight (32 KiB,
parameters.user_stack_pages), with the page below left unmapped as a
guard so an overflow faults into a clean process kill rather than
corrupting the image. Task.name_buffer is zero-initialised, not
undefined: an undefined default is materialised as a 0xAA fill that
moved the static task pool out of .bss and made the whole kernel ~7x
slower under QEMU TCG (caught by the affinity test).

Proven end to end by the new args test: args-echo respawns itself with
arguments via the syscall blob, burns more stack than one page could
hold, and echoes its argv intact. Full suite: 44/44.
2026-07-11 08:33:12 +01:00
Daniel Samson 6b3ae0c997 Kill a faulting user process instead of halting the machine
A CPU exception raised in ring 3 by a scheduled process now kills that
process - IRQ bindings, IPC handles, and address space reclaimed, a
client it owed a reply to failed with the new -EPEER instead of hung -
and the core reschedules (docs/resilience.md step 2). Kernel-mode
faults, NMI, double fault, and machine check stay terminal, as does the
borrowed-thread isolation probe. Proven by the new fault-recovery QEMU
test: init keeps heartbeating after a process page-faults to death.
2026-07-11 04:57:02 +01:00
98 changed files with 22648 additions and 1028 deletions
+16
View File
@@ -0,0 +1,16 @@
# EditorConfig: https://editorconfig.org/
# Follows the Zig style guide: https://ziglang.org/documentation/0.16.0/#Style-Guide
root = true
[*]
charset = utf-8
end_of_line = lf
indent_style = space
indent_size = 4
trim_trailing_whitespace = true
insert_final_newline = true
[*.zig]
# "Line length: aim for 100; use common sense."
max_line_length = 100
+1
View File
@@ -0,0 +1 @@
*.zig text eol=lf
+26 -7
View File
@@ -2,12 +2,31 @@
Codename: Shodan
Version: 1
A small operating system, written from scratch in Zig — a bootloader (`boot/`)
and a microkernel (`system/kernel/`), sharing a neutral handoff contract (`system/boot-handoff.zig`).
It boots x86-64 via UEFI, and so far has a framebuffer console, a physical frame
allocator, its own paging with W^X permissions, interrupt/exception handling, a
LAPIC timer, a kernel heap, a fixed-priority preemptive scheduler, and in-kernel IPC
channels. See [`docs/`](docs/README.md) for how each piece works.
A small resilient operating system, written from scratch in Zig.
## Zen of DanOS:
- Resilient Micro-Kernel Architecture.
- Every process run in an isolated user space not kernel space.
- Processes cannot take down the entire OS with it when they die or is killed
- Stable public runtime library, private OS ABI.
- Keeps a stable runtime for user space processes between OS versions (great for backwards compatibility)
- Allows the underlying OS to be changed without effecting applications
- Provides a boundary to enable compatibility between OS's e.g. POSIX, MUSL etc
- Drivers are just isolated processes in user space.
- Thin binaries that can be restarted like applications.
- Useful during driver development.
- Drivers can claim MMIO / ports
- Driver resources (e.g. IRQ/Port/MMIO) claims are automatically cleaned up if the driver dies or is killed
- Drivers can also hook into the process lifecyle to clean up or reset hardware
- No legacy to deal with
- Zig code uses a clean coding style (Zen of Zig)
- Favor reading code over writing code.
- No magic numbers.
- No shortend names unless its for ABI compatibility or acronyms
- Inter-Process Communication (IPC)
- Publish and subscribe to Asynchronous Messages
- Talk to services and processes synchronously
## Prerequisites
@@ -60,7 +79,7 @@ straight into CI.
## Documentation
Design notes explaining the *why* behind the code live in
Design notes explaining *why* behind the code live in
[`docs/`](docs/README.md) — start with [`docs/README.md`](docs/README.md).
## Logo
+162 -6
View File
@@ -60,6 +60,8 @@ fn addUserBinary(
runtime_module: *std.Build.Module,
posix_module: *std.Build.Module,
mmio_module: *std.Build.Module,
xkeyboard_config_module: *std.Build.Module,
acpi_ids_module: *std.Build.Module,
name: []const u8,
root: []const u8,
) *std.Build.Step.Compile {
@@ -81,6 +83,12 @@ fn addUserBinary(
.{ .name = "posix", .module = posix_module },
// Typed volatile MMIO + memory barriers, for drivers. See library/mmio/.
.{ .name = "mmio", .module = mmio_module },
// Keyboard layouts (keycode + modifiers -> keysym/character), available
// to any program that wants it. See library/xkeyboard-config/.
.{ .name = "xkeyboard-config", .module = xkeyboard_config_module },
// ACPI/PnP hardware-ID registry, so drivers name devices
// (HardwareId.ps2_keyboard) instead of magic "_HID" strings.
.{ .name = "acpi-ids", .module = acpi_ids_module },
},
}),
});
@@ -123,6 +131,13 @@ pub fn build(b: *std.Build) void {
});
// ACPI/PnP hardware-ID (_HID) names — the flat analog of pci-class for acpi_device
// nodes. Also shared reference data.
// The AML interpreter, a build module so the ring-3 acpi service can run the
// same parser the kernel does (docs/m19-m20-plan.md decision 1). Pure Zig,
// no kernel imports — one source, two builds.
const aml_module = b.addModule("aml", .{
.root_source_file = b.path("system/devices/aml/aml.zig"),
});
const acpi_ids_module = b.addModule("acpi-ids", .{
.root_source_file = b.path("system/devices/acpi-ids.zig"),
});
@@ -178,6 +193,13 @@ pub fn build(b: *std.Build) void {
.root_source_file = b.path("system/services/vfs/protocol.zig"),
});
// The input wire protocol: the input service's public interface, exposed as its own
// module the same way vfs-protocol is. Shared by the input service, the runtime's
// `input` helper (subscribe/publish), and every source and subscriber.
const input_protocol_module = b.addModule("input-protocol", .{
.root_source_file = b.path("system/services/input/protocol.zig"),
});
// The danos-native user-space runtime: system_call wrappers, the C-convention
// heap, IPC helpers, the process start shim, device access. This is the stable
// application ABI; POSIX compatibility is a separate library on top (see below).
@@ -193,9 +215,23 @@ pub fn build(b: *std.Build) void {
.{ .name = "abi", .module = abi_module },
.{ .name = "device-abi", .module = device_abi_module },
.{ .name = "vfs-protocol", .module = vfs_protocol_module },
.{ .name = "input-protocol", .module = input_protocol_module },
},
});
// The device-manager protocol: hello + (M18.2) tree reports, exposed as its
// own module like the other protocol modules. Imported through the runtime.
const device_manager_protocol_module = b.addModule("device-manager-protocol", .{
.root_source_file = b.path("system/services/device-manager/device-manager-protocol.zig"),
});
runtime_module.addImport("device-manager-protocol", device_manager_protocol_module);
// The power protocol: system power's domain-named surface (docs/m21-plan.md).
const power_protocol_module = b.addModule("power-protocol", .{
.root_source_file = b.path("system/services/power/protocol.zig"),
});
runtime_module.addImport("power-protocol", power_protocol_module);
// Typed volatile MMIO register access + memory-ordering barriers, for drivers on
// top of an mmio_map grant. Depends only on `builtin` (arch-conditional barriers);
// no target set, so it inherits each driver's. See library/mmio/mmio.zig.
@@ -203,6 +239,20 @@ pub fn build(b: *std.Build) void {
.root_source_file = b.path("library/mmio/mmio.zig"),
});
// Keyboard layouts compiled from the X11 xkeyboard-config database into native Zig
// (keycode + modifiers -> keysym/character). The `layouts` tables are generated by
// tools/make-xkeyboard-config.py; `xkeyboard-config` is the hand-written API over them.
// No target set, so each inherits its importer's. See library/xkeyboard-config/.
const xkb_layouts_module = b.addModule("layouts", .{
.root_source_file = b.path("library/xkeyboard-config/generated/layouts.zig"),
});
const xkeyboard_config_module = b.addModule("xkeyboard-config", .{
.root_source_file = b.path("library/xkeyboard-config/xkeyboard-config.zig"),
.imports = &.{
.{ .name = "layouts", .module = xkb_layouts_module },
},
});
// The POSIX / C compatibility layer, a separate library layered strictly over the
// runtime (it calls the runtime's IPC/heap, never system calls directly). This is
// the one place POSIX/C spellings are allowed verbatim — see docs/coding-standards.md
@@ -286,7 +336,7 @@ pub fn build(b: *std.Build) void {
// Built by the shared user-binary recipe (see addUserBinary): freestanding,
// linked into the kernel's user region against the `runtime` runtime library, and
// started in ring 3 by the kernel's user-ELF loader.
const init_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, "init", "system/services/init/init.zig");
const init_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "init", "system/services/init/init.zig");
const init_install = b.addInstallArtifact(init_exe, .{ .dest_dir = .{ .override = .{ .custom = "system/services" } } });
b.getInstallStep().dependOn(&init_install.step);
@@ -294,11 +344,47 @@ pub fn build(b: *std.Build) void {
// Each is built by the same user-binary recipe, then packed into one image by
// the host-side make-initial-ramdisk tool. The bootloader ferries the image to the kernel,
// which unpacks it and spawns each program (system/initial-ramdisk.zig).
const vfs_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, "vfs", "system/services/vfs/vfs.zig");
const vfstest_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, "vfs-test", "system/services/vfs/vfs-test.zig");
const hpet_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, "hpet", "system/drivers/hpet/hpet.zig");
const bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, "bus", "system/drivers/bus/bus.zig");
const device_manager_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, "device-manager", "system/services/device-manager/device-manager.zig");
const vfs_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "vfs", "system/services/vfs/vfs.zig");
const vfstest_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "vfs-test", "system/services/vfs/vfs-test.zig");
const hpet_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "hpet", "system/drivers/hpet/hpet.zig");
const bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "bus", "system/drivers/bus/bus.zig");
const ps2_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-bus", "system/drivers/ps2-bus/ps2-bus.zig");
const ps2_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-keyboard", "system/drivers/ps2-bus/keyboard.zig");
const ps2_mouse_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-mouse", "system/drivers/ps2-bus/mouse.zig");
const usb_xhci_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-xhci-bus", "system/drivers/usb-xhci-bus/usb-xhci-bus.zig");
const pci_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "pci-bus", "system/drivers/pci-bus/pci-bus.zig");
// The PCI bus driver decodes each function's class triple to human names in its
// boot log (class/subclass/prog-IF), so pull in the shared pci-class reference.
pci_bus_exe.root_module.addImport("pci-class", pci_class_module);
// A test fixture, not a real driver: hellos to the device manager, then faults —
// what the driver-restart scenario drives the crash-loop cap with.
const crash_test_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "crash-test", "system/services/crash-test/crash-test.zig");
const device_list_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-list", "system/services/device-list/device-list.zig");
// The discovery service: one swappable process per firmware
// (docs/m19-m20-plan.md decision 7), bundled under the neutral ramdisk name
// "discovery" so the device manager never learns which firmware it is on.
// x86 boots describe hardware with ACPI; the Raspberry Pis hand over a
// flattened device tree — the aarch64 target flips the default when it
// lands (docs/arm.md). Both are placeholders until M20.1 (acpi) and the
// ARM bring-up (fdt).
const Discovery = enum { acpi, fdt };
const discovery = b.option(Discovery, "discovery", "Which discovery service fills the ramdisk's 'discovery' slot (default: acpi)") orelse Discovery.acpi;
const discovery_source: []const u8 = switch (discovery) {
.acpi => "system/services/acpi/acpi.zig",
.fdt => "system/services/fdt/fdt.zig",
};
const discovery_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "discovery", discovery_source);
if (discovery == .acpi) discovery_exe.root_module.addImport("aml", aml_module);
const device_manager_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-manager", "system/services/device-manager/device-manager.zig");
// Names the xHCI PCI class triple from the shared taxonomy instead of a bare 0x0C0330.
device_manager_exe.root_module.addImport("pci-class", pci_class_module);
// The input service and its exercisers: the fan-out server, a hardware-free synthetic
// source, and a subscriber that doubles as the `input` test's oracle. See docs/input.md.
const input_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input", "system/services/input/input.zig");
const input_source_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input-source", "system/services/input-source/input-source.zig");
const input_test_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input-test", "system/services/input-test/input-test.zig");
const args_echo_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "args-echo", "system/services/args-echo/args-echo.zig");
const process_test_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "process-test", "system/services/process-test/process-test.zig");
// Pack the user binaries into the initial_ramdisk image with the host-side Python tool
// (the container format is trivial, and Python sidesteps std API churn). Args:
@@ -314,16 +400,47 @@ pub fn build(b: *std.Build) void {
mk_run.addFileArg(hpet_exe.getEmittedBin());
mk_run.addArg("bus");
mk_run.addFileArg(bus_exe.getEmittedBin());
mk_run.addArg("ps2-bus");
mk_run.addFileArg(ps2_bus_exe.getEmittedBin());
mk_run.addArg("ps2-keyboard");
mk_run.addFileArg(ps2_keyboard_exe.getEmittedBin());
mk_run.addArg("ps2-mouse");
mk_run.addFileArg(ps2_mouse_exe.getEmittedBin());
mk_run.addArg("usb-xhci-bus");
mk_run.addFileArg(usb_xhci_bus_exe.getEmittedBin());
mk_run.addArg("pci-bus");
mk_run.addFileArg(pci_bus_exe.getEmittedBin());
mk_run.addArg("crash-test");
mk_run.addFileArg(crash_test_exe.getEmittedBin());
mk_run.addArg("device-list");
mk_run.addFileArg(device_list_exe.getEmittedBin());
mk_run.addArg("discovery");
mk_run.addFileArg(discovery_exe.getEmittedBin());
mk_run.addArg("device-manager");
mk_run.addFileArg(device_manager_exe.getEmittedBin());
mk_run.addArg("input");
mk_run.addFileArg(input_exe.getEmittedBin());
mk_run.addArg("input-source");
mk_run.addFileArg(input_source_exe.getEmittedBin());
mk_run.addArg("input-test");
mk_run.addFileArg(input_test_exe.getEmittedBin());
mk_run.addArg("args-echo");
mk_run.addFileArg(args_echo_exe.getEmittedBin());
mk_run.addArg("process-test");
mk_run.addFileArg(process_test_exe.getEmittedBin());
// Also install the packed binaries to their FHS homes, so zig-out is a true image
// of the filesystem — even though at boot they arrive inside the initial-ramdisk.
for ([_]struct { *std.Build.Step.Compile, []const u8 }{
.{ vfs_exe, "system/services" },
.{ device_manager_exe, "system/services" },
.{ input_exe, "system/services" },
.{ hpet_exe, "system/drivers" },
.{ bus_exe, "system/drivers" },
.{ ps2_bus_exe, "system/drivers" },
.{ ps2_keyboard_exe, "system/drivers" },
.{ ps2_mouse_exe, "system/drivers" },
.{ usb_xhci_bus_exe, "system/drivers" },
}) |entry| {
const step = b.addInstallArtifact(entry[0], .{ .dest_dir = .{ .override = .{ .custom = entry[1] } } });
b.getInstallStep().dependOn(&step.step);
@@ -395,6 +512,19 @@ pub fn build(b: *std.Build) void {
const run_efi = b.addSystemCommand(&.{
"qemu-system-x86_64",
"-device",
"qemu-xhci,id=xhci",
"-device",
"usb-mouse,bus=xhci.0",
"-device",
"usb-kbd,bus=xhci.0",
// "-usb",
// "-device",
// "usb-ehci,id=ehci",
// "-device",
// "usb-tablet,bus=usb-bus.0",
// "-device",
// "usb-mouse,bus=ehci.0",
"-machine",
"q35",
"-m",
@@ -453,7 +583,12 @@ pub fn build(b: *std.Build) void {
"system/devices/device-abi.zig",
"system/devices/pci-class.zig", // class/subclass/prog-IF name decoding
"system/devices/acpi-ids.zig", // _HID name decoding
"system/devices/aml/aml.zig", // AML parse + interpret, incl. Notify dispatch (M21)
"system/devices/usb-abi.zig", // wire sizes + bit packings + set-up packet encodings
"system/devices/usb-ids.zig", // class/subclass/protocol code assignments
"library/mmio/mmio.zig", // barriers assemble + registers round-trip
"system/drivers/ps2-bus/scancode.zig", // set-2 decode + keyboard state machine
"system/drivers/ps2-bus/mouse-packet.zig", // 3-byte mouse packet assembly
}) |root| {
const mod_tests = b.addTest(.{
.root_module = b.createModule(.{
@@ -464,4 +599,25 @@ pub fn build(b: *std.Build) void {
});
test_step.dependOn(&b.addRunArtifact(mod_tests).step);
}
// The xkeyboard-config keymap tests need its generated `layouts` import wired, so they
// don't fit the plain loop above. Its keycode->character assertions are the end-to-end
// proof that the xkb-data -> generator -> Zig-lookup pipeline is correct.
const xkb_tests = b.addTest(.{
.root_module = b.createModule(.{
.root_source_file = b.path("library/xkeyboard-config/xkeyboard-config.zig"),
.target = target,
.optimize = optimize,
.imports = &.{
.{ .name = "layouts", .module = xkb_layouts_module },
},
}),
});
test_step.dependOn(&b.addRunArtifact(xkb_tests).step);
// Convenience: `zig build gen-xkeyboard-config` regenerates the layout tables from the
// vendored data (offline). `fetch` (the network step) stays a manual script run.
const gen_xkb = b.addSystemCommand(&.{ "python3", "tools/make-xkeyboard-config.py", "generate" });
const gen_xkb_step = b.step("gen-xkeyboard-config", "Regenerate library/xkeyboard-config/generated from the vendored data");
gen_xkb_step.dependOn(&gen_xkb.step);
}
+25 -4
View File
@@ -45,10 +45,31 @@ rather than restate it. Roughly in the order things happen at runtime:
until its hardware interrupts it**. The claim is the capability; `irq_ack` is the
unmask.
14. **[driver-model.md](driver-model.md) — buses, classes and host controllers.** How
real driver stacks factor into three shapes, how families share code, and the
proposed ABI for the three primitives still missing (capability passing, DMA +
memory barriers, MSI).
15. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
real driver stacks factor into three shapes and how families share code. The
three primitives it proposed are long since built (M13 capability passing,
M14 DMA + barriers, M15 MSI), and the driver *contract* on top of them —
hello, supervision, restart — is built too (device-manager.md, M18).
15. **[process-management.md](process-management.md) — process management.** The
microkernel's `ps`/`kill`/SIGCHLD: enumerate as a table snapshot, the
supervision link as the kill authority, and child-exit notifications over the
same endpoints IRQs arrive on.
16. **[process-lifecycle.md](process-lifecycle.md) — the process lifecycle.** Built
(M17): signals over IPC as the one lifecycle vocabulary every process speaks — the
POSIX.1-1990 words with message delivery instead of stack hijack, the stable
`runtime.process` interface, exit reasons, published exit events any stateful
service can subscribe to (the VFS releasing dead clients' handles), and the two
iron rules (cleanup is the kernel's job; kill is not a signal).
17. **[device-manager.md](device-manager.md) — the device manager.** Built (M18,
through the app surface): the
tree, the matcher, and the supervisor. Tree structure lives in the manager,
authority stays in the kernel; bus drivers report what they see; drivers are
restarted through the lifecycle vocabulary — the plan that turns
[resilience.md](resilience.md)'s restart goal into increments.
18. **[input.md](input.md) — the input module.** Broadcasting input events (keyboard,
mouse, joystick): why a synchronous rendezvous can't fan out to many listeners, the
asynchronous `ipc_send` primitive built to fix it, and the per-device subscribe/publish
service layered on top.
19. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
Start with the north star:
+42
View File
@@ -142,6 +142,32 @@ conventions above — `snake_case` — because it's an identifier, not a filenam
*directory* (`system/services/init`, `library/runtime`), with the repeated leaf
resolving away. See the repository-layout section of [README.md](README.md).
## Named values, not magic numbers
The naming rule has a twin: **a value with meaning gets a name, too.** The same
principle drives both — a reader should never have to leave the code to understand it.
An abbreviated *name* forces a reader to guess; a bare *number* forces them worse, out
to a spec or a header or a comment three files away, to learn what the value even *is*.
If `0x0C` is the PCI serial-bus class, the code says `BaseClass.serial_bus`, not `0x0C`;
if `0x04` is the ACPI IRQ resource descriptor, it says `SmallResourceType.irq`, not
`0x04`. The number is an implementation detail of the name — recorded once, where the
name is defined, and never spelled again at a use site.
**Prefer an `enum`** when the values form a set (device classes, AML opcodes, resource
descriptor types, states): the type then also says *which* set a value belongs to, and
the compiler rejects a value from the wrong one. A lone `pub const` with a descriptive
name suffices for a one-off (`const large_descriptor_bit = 0x80`). Reach for the enum
the moment code elsewhere compares against, packs, or produces the value — a packed PCI
class triple is written from named parts (`.serial_bus`, `.usb`, `.xhci`), never as
`0x0C_03_30` under a comment that decodes the bytes.
The exceptions are the numbers that carry no hidden meaning: `0` and `1` as plain zero
and one, an index step, a field width, a bit shift. `x + 1`, `buffer[0]`, and `<< 8`
need no christening — there is nothing to look up. The test is exactly the naming test:
*would a reader have to look this up to know what it means?* If yes, name it. This is
what `opcodes.zig`'s `*_opcode` constants, `acpi-ids`'s `HardwareId`, and `pci-class`'s
class enums already are — reference data defined once and named everywhere it is used.
## Why acronyms are the line
Because an acronym has no letters to restore. `MMIO` doesn't become "memory mapped
@@ -150,3 +176,19 @@ input output" in code — that expansion is what the acronym *is for*. But `msg`
test for "is this an abbreviation I must expand" is simply: *is there a longer word this
is a clipped form of?* If yes, write the word. If it's an initialism standing in for a
phrase, leave it.
## Zen of Zig
* Communicate intent precisely.
* Edge cases matter.
* Favor reading code over writing code.
* Only one obvious way to do things.
* Runtime crashes are better than bugs.
* Compile errors are better than runtime crashes.
* Incremental improvements.
* Avoid local maximums.
* Reduce the amount one must remember.
* Focus on code rather than style.
* Resource allocation may fail; resource deallocation must succeed.
* Memory is a resource.
* Together we serve the users.
+156
View File
@@ -0,0 +1,156 @@
# The device manager
**Status: the protocol and supervision are built** (M18.1, 2026-07-13): `hello`
with its deadline, supervised spawn, restart with backoff, and the crash-loop
cap are in — usb-xhci-bus is the first conforming driver, and the
`driver-restart` scenario proves fault → backoff → re-claim → cap end to end.
Tree reports are built too (M18.2, 2026-07-13): the xHCI driver scans its
root-hub ports and reports each connected device (`child_added`); the manager
mirrors them and prunes a dead reporter's children, and the `usb-report`
scenario proves report → prune → respawn → re-report. The application surface is built (M18.3, 2026-07-13):
`enumerate` and `subscribe` over IPC, with `device-list` as the first client —
the manager is now the one answer to "what devices exist" for applications.
The primitives underneath are real ([process-management.md](process-management.md):
spawn/supervise/kill/exit-notification; [driver-model.md](driver-model.md): the device
table as a capability system; [drivers.md](drivers.md): claim/map/IRQ), and the first
per-device driver spawn works (the device manager matches the xHCI controller by PCI
class and spawns `usb-xhci-bus` with the device id as argv[1]). This document designs
the rest: the device manager as **the tree, the matcher, and the supervisor** — the
policy process that turns [resilience.md](resilience.md)'s restart goal into practice
for drivers.
How processes stop, reload, and report their deaths is deliberately **not** in this
document: that is the universal lifecycle every danos process speaks —
[process-lifecycle.md](process-lifecycle.md), signals over IPC and the stable
`runtime.process` interface. The device manager is that design's first serious
customer, not its owner. Its own protocol contains nothing lifecycle-shaped; a
driver is stopped, health-checked, and buried exactly like any other process.
## The tree: structure in the manager, authority in the kernel
The device tree is two things fused: *information* (what exists, how it nests) and
*authority* (a descriptor is a licence to map physical memory). They separate:
- The **kernel keeps the capability system** — device, I/O-port, and interrupt
claims, resource containment on `device_register`, the
`mmio_map`/`irq_bind`/`msi_bind` gates — and **cleans all of it up when a process
dies** (settled; it is increment 1 of
[process-lifecycle.md](process-lifecycle.md)). The three invariants in
[driver-model.md](driver-model.md) stay exactly where they are. A device manager
that could mint MMIO mappings by its own say-so would be a second kernel, and a
buggy one would un-earn everything the microkernel bought.
- The **device manager owns the tree as data** — identity, topology, naming, driver
matching, hotplug events, and being the one process everything else asks about
devices. Firmware discovery seeds it (today via the kernel's snapshot); **bus
drivers grow it** by reporting what they see; applications query and watch it.
`device_enumerate` fades to a manager-internal (then deleted) seam.
Long-term, discovery itself leaves the kernel — but not *into* the manager. PCI
enumeration is a **pci-bus driver**: the manager spawns it against the host bridge
(already a device with the ECAM window as a resource), it scans, it reports functions
like any bus reports children. ACPI becomes an **acpi service** that interprets the
tables and reports the namespace. The manager only orchestrates and merges. Moving
AML interpretation out of ring 0 is its own project on its own track; nothing here
depends on when it lands.
## The protocol
A `device-manager-protocol` module (the vfs-protocol pattern): extern-struct
messages, a version in the handshake, reserved fields everywhere. The manager is a
well-known endpoint (`ipc.register(.device_manager)`); the badge tells it who is
talking; the same endpoint receives its children's exit notifications — one loop,
one world.
| Direction | Message | Purpose |
|---|---|---|
| driver → manager | `hello { version, role, device_id }` | confirms the argv assignment, starts the deadline clock |
| bus → manager | `child_added { parent, identity, resources }` | one node the bus discovered |
| bus → manager | `child_removed { id }` | unplug, or the bus lost it |
| app → manager | `enumerate` | snapshot of the tree (read-only) |
| app → manager | `subscribe` | receive published add/remove events |
`hello` is the one deadline the manager enforces itself: spawned and silent past the
deadline means wrong binary, wrong protocol version, or wedged before main — apply
the stop sequence and the restart policy. Everything else lifecycle-shaped
(terminate, the common `ping` liveness call, exit reasons) arrives through
[process-lifecycle.md](process-lifecycle.md)'s vocabulary, not this protocol.
Assignment stays argv (`usb-xhci-bus <device id>`) for now — simple, and it works.
The step after `hello` exists is delegation: the manager claims (or is granted) the
devices and passes the claim to the driver over IPC (the M13 capability-transfer
mechanism), replacing first-come-first-served `device_claim` with policy. Identity in
`child_added` is per-bus: PCI children carry the class triple (`pci_class`, as the
xHCI match already uses); USB children carry the (class, subclass, protocol) triple
from usb-ids.zig — each bus's native language, decoded by the shared ids modules.
## Supervision and restart
Every driver is spawned with the manager's exit endpoint (`spawnSupervised` — built).
On a death notification:
1. **Read the reason** ([process-lifecycle.md](process-lifecycle.md) increment 2).
Clean exit → it meant to; don't restart. Fault or missed `hello` deadline →
restart with **backoff**, and a crash-loop cap (three fast deaths → mark failed,
stop respawning, log loudly; a later `reload` to the manager can retry).
2. **Prune the subtree** the dead bus driver reported. Its children describe
protocol state (xHCI slot ids, transfer rings) that died with the process;
keeping the nodes would be keeping a lie. Watchers receive `child_removed` — the
input service losing, then regaining, a keyboard is the *honest* description of
what happened. The restarted instance rediscovers and re-reports.
3. **The claim is already free** because the kernel released it at death — the
restarted instance claims the same controller and comes up.
Who supervises the supervisor: **init** (PID 1), which already supervises the
services it starts. If the manager dies, drivers keep running (they hold their
claims; the kernel doesn't care who their supervisor was — though their exit
notifications now dangle harmlessly). The restarted manager re-learns the world:
kernel snapshot, then a re-`hello` round — drivers answer a broadcast or are stopped
and respawned. Full state handoff is deliberately not attempted.
## Thin drivers, class protocols
The [driver-model.md](driver-model.md) three-shape split, restated as processes:
- A **bus driver** (usb-xhci-bus) owns its controller — claim, MMIO, IRQ/MSI, DMA
rings — and offers a *transfer* protocol ("submit a control transfer to device N",
built from the usb-abi request constructors) plus tree reports to the manager.
- A **class driver** (usb-hid, usb-storage) owns nothing: it is matched to a reported
child by its identity triple, speaks the bus's transfer protocol downward and its
service's protocol upward — HID reports to the input service, blocks to the block
service. It works unchanged over any controller.
- **Services** (input, display, block) aggregate class drivers and face applications.
Each arrow is a protocol module. The manager routes none of the data plane — it
introduces the parties (matching), supervises them (lifecycle), and gets out of the
way.
## Increments
Increments 1–4 are the lifecycle prerequisites and live in
[process-lifecycle.md](process-lifecycle.md) (claim cleanup on death, exit reasons,
published exit events, signals + `runtime.process`). On top of those:
5. **device-manager-protocol**: `hello`, supervised spawn with restart policy;
usb-xhci-bus becomes the first conforming driver.
6. **Tree reports**: `child_added`/`child_removed`; the manager mirrors; xHCI reports
the mouse and keyboard QEMU already hangs off it.
7. **App surface**: `enumerate`/`subscribe` over IPC; `device_enumerate` retreats
to a manager-internal seam.
8. **Discovery migration** — DONE (M19–M20, 2026-07-13): pci-bus driver (M19)
then the acpi service (M20) moved enumeration to ring 3; the kernel seeds
only the host bridge and the acpi-tables node. See
[m19-m20-plan.md](m19-m20-plan.md).
## Settled questions (2026-07-12)
- **Stateful buses**: pruning the subtree on bus-driver death is right for USB. A
future storage bus with in-flight writes wants drain-before-terminate — which is
exactly the `deadline_ms` parameter `stop()` already has; a per-driver deadline
is one value in the manager's policy table when such a bus arrives. No design
change.
- **Manager death**: drivers survive the manager; the restarted manager re-learns
the world (above). Checkpointing driver state with the manager is deferred until
something demonstrates the need.
- **Matching stays code until the third bus.** `driverFor`/`pciDriverFor` are
honest at two bus types; the third triggers the manifest (a driver declares what
it binds: a PCI class triple, a USB class triple, an ACPI `_HID`).
+24
View File
@@ -167,3 +167,27 @@ free; discovery on x86 is partly about *finding* what ARM just tells you.
- [ipc.md](ipc.md) — the channels that interrupts-as-messages and the device manager
will ride on.
- [vision.md](vision.md) — why drivers belong in isolated user space at all.
## Update (M19.3, 2026-07-13): PCI enumeration left the kernel
The kernel now seeds only the `pci_host_bridge` node (ECAM window, MMIO
apertures derived from the memory map's holes, bus range, and the 16-bit I/O
window). The per-function walk moved to the ring-3 `pci-bus` driver
([device-manager.md](device-manager.md)): it claims the bridge, repeats the
ECAM scan through its mmio grant, and `device_register`s what it finds, which
the device manager mirrors and matches. The ACPI namespace walk follows in M20;
the static tables (MADT, HPET, MCFG, FADT + `\\_S5`) stay kernel-side.
## Update (M20.3, 2026-07-13): ACPI enumeration left the kernel too
The kernel no longer folds the AML namespace's Device objects into the device
tree. It still parses the *static* tables (MADT for SMP, HPET for the tick, MCFG
for the host bridge, FADT) and still builds the AML namespace — but only to read
the `\\_S5` sleep type for poweroff. Device discovery is the ring-3 **acpi
service** ([device-manager.md](device-manager.md)): it claims the `acpi-tables`
node the kernel publishes (the AML blobs, a broad io_port grant, the SCI),
re-parses the same blobs with the shared AML module, evaluates `_STA`/`_CRS`,
and registers + reports each `_HID` device — the device manager matches drivers
(ps2-bus) from those reports. With M19's pci-bus driver, discovery now runs
entirely in user space; the kernel seeds only the host bridge and the
acpi-tables node.
+5 -2
View File
@@ -164,8 +164,11 @@ If a class driver needs `mmio`, it has become an HCD and should be one.
DMA is still unprotected** (the caveat below). Enforcement lands with the first DMA
driver, which is what there is to protect and test against. Proven in the `iommu` test,
booted with an emulated `intel-iommu`.
- **`system_spawn`** — a user-space supervisor starts a driver: `system_spawn(name)`
loads a binary bundled in the initial-ramdisk as a fresh ring-3 process. This is what
- **`system_spawn`** — a user-space supervisor starts a driver:
`system_spawn(name, arguments)` loads a binary bundled in the initial-ramdisk as a
fresh ring-3 process; `name` becomes the child's argv[0] and the optional
NUL-separated `arguments` blob its argv[1..], delivered on a SysV entry stack
([sysv.md](sysv.md)). This is what
turned the device manager from "log the match" into "run the driver": the kernel now
spawns only `init`, `init` spawns the services, and the **device-manager** discovers
the hardware and spawns each driver ([drivers.md](drivers.md)). Ungated for now — a
+25 -3
View File
@@ -37,8 +37,9 @@ kernel ──spawns──► init (PID 1) ──spawns──► device-manag
```
The kernel launches exactly one process — `init` — and hands it nothing but the raw
ability to start more (`system_spawn(name)`, which loads a binary bundled in the
initial-ramdisk as a fresh ring-3 process). Everything else is a user-space decision:
ability to start more (`system_spawn(name, arguments)`, which loads a binary bundled
in the initial-ramdisk as a fresh ring-3 process — `name` becoming its argv[0],
the optional arguments its argv[1..], on a SysV entry stack, see sysv.md). Everything else is a user-space decision:
- **init** ([system/services/init](system/services/init/init.zig)) is the **service
supervisor**. It spawns the system services danos brings up at boot — today `vfs` and
@@ -52,7 +53,8 @@ initial-ramdisk as a fresh ring-3 process). Everything else is a user-space deci
is a table (`driverFor`): today a static `timer → hpet` map; a fuller system reads
what each driver *binds* (a manifest under `/system/drivers`, or the driver
describing its own match).
3. **Spawn** — `system_spawn(driver_name)` starts the matched driver, which then claims
3. **Spawn** — `system_spawn(driver_name, arguments)` starts the matched driver (the
arguments can carry *which* device it matched), which then claims
its device and runs the event loop below.
So "how is a driver discovered and configured" has two halves: **discovery** is the
@@ -360,3 +362,23 @@ the first DMA driver to protect and test against) and these smaller items:
- **Interrupt priority / threaded IRQ latency.** `notifyFromIsr` enqueues the woken
driver but doesn't preempt (`wakeLocked` deliberately leaves that to the caller), so
a woken driver waits for the next scheduling point.
## The driver contract (M17–M18)
Claiming and mapping is half of being a danos driver; the other half is the
**lifecycle and protocol contract**, and the runtime makes it nearly free:
- Build on `runtime.service.run` — one replyWait loop folding protocol
requests, signals, and notifications into callbacks. The harness answers the
universal zero-length ping and turns `terminate` into a clean exit for you
([process-lifecycle.md](process-lifecycle.md)).
- A driver spawned with an assignment (its device id as argv[1]) sends the
versioned `hello` to the device manager inside the deadline, and a **bus**
driver reports what it discovers with `child_added`
([device-manager.md](device-manager.md); usb-xhci-bus is the reference
implementation).
- Crash freely — that is the design. The kernel releases your claims, IRQ
bindings, and MSI vectors at death; the manager reads your exit reason,
prunes what you reported, restarts you with backoff, and your fresh instance
re-claims and re-reports. Never depend on your own cleanup running
(iron rule 1).
+161
View File
@@ -0,0 +1,161 @@
# The input module: broadcasting input events
A keyboard driver has one keystroke and *many* programs that might want it — a shell, a
window server, a logger. None of them owns the hardware, and the driver should not know
who is listening. So between the drivers and the listeners sits the **input service**
(`system/services/input/`): drivers **publish** events to it, programs **subscribe**, and
it fans each event out to every interested subscriber. It is an ordinary ring-3 process
reached over IPC, like the [VFS server](../system/services/vfs/vfs.zig) — no kernel knows
what a key is.
## One service, several device classes
The service carries three device classes today — **keyboard**, **mouse**, and
**joystick/gamepad** — and is built to take more
([protocol.zig](../system/services/input/protocol.zig)). Each class has its own typed
event:
- `KeyEvent` — `key_down`/`key_up` (physical make/break) and `key_press` (a character was
produced, carrying the Unicode scalar); plus a layout-independent `keycode` and a
`modifiers` bitmask.
- `MouseEvent` — relative `motion` (`dx`/`dy`), `button_down`/`button_up`, and `scroll`.
- `JoystickEvent` — `axis` moves (a signed value on a `control` index) and
`button_down`/`button_up`.
All three travel in one **`InputEvent` envelope** tagged with a `DeviceKind`, so the
fan-out is a single code path and a subscriber can take a mix of classes on one stream.
Decode an envelope with `asKeyboard()` / `asMouse()` / `asJoystick()` (each returns null
unless the tag matches). A subscriber names the classes it wants with a **`device_mask`**,
and the service routes each event only to subscribers whose mask includes its class — so a
mouse-only listener never wakes for keystrokes.
## Why this needed a new kernel primitive
The interesting part is delivery, and it runs straight into the shape of danos IPC.
[ipc.md](ipc.md) describes a **synchronous rendezvous**: a server holds exactly one
pending reply (`Task.ipc_client`) and *must* answer it on its next `replyWait`. Two
consequences decide the whole design:
1. **You cannot block N subscribers waiting for "the next event".** A server can hold only
one caller at a time, so the natural "subscriber calls `next_event()` and blocks" API
is impossible for more than one subscriber. Delivery therefore has to be **push** — the
service reaching out to subscribers — not pull.
2. **A synchronous push can hang the whole service.** If the service delivered with
`ipc_call`, it would block until each subscriber replied. `ipc_call` has no timeout, and
the kernel does **not** wake a caller parked on a *dead* peer's endpoint (it only fails a
peer that was mid-reply — see [process.zig](../system/kernel/process.zig)
`releaseTaskResourcesLocked`). One subscriber that exits mid-delivery would wedge input
for everyone. That is the opposite of the resilience the microkernel is for.
The fix is the asynchronous send that [ipc.md](ipc.md) had already earmarked as future
work ("asynchronous / buffered send … for notifications between servers"):
```
ipc_send(handle, message_ptr, message_len) -> 0 / -errno
```
`ipc_send` copies a small payload into the endpoint's **bounded queue** and wakes a
receiver, then returns immediately — it never blocks and so can never hang on a dead or
slow subscriber. The receiver picks it up through the same `replyWait` it already runs:
the wake arrives as a **buffered message** — `notify_badge_bit | notify_message_bit` set in
the badge (distinguishing it from a bare IRQ/child-exit notification), the sender's task id
in the low bits, and the payload in the receive buffer, with no reply owed. The queue holds
16 messages per endpoint; a full queue **drops the oldest**, because a buffered message is
discrete data, not a coalescing "level" like an interrupt. See
[ipc-synchronous.zig](../system/kernel/ipc-synchronous.zig) (`sendLocked`, `popPost`, and
the `replyWait` receive loop).
This is the async counterpart of `ipc_call`, and the input service is its first consumer.
## How the pieces fit
```
keyboard/mouse driver, input-source input service subscriber(s)
----------------------------------- ------------- -------------
connectSource(); loop: replyWait: subscribeKeyboard()/…All:
publishKeyboardEvent(k) ─ ipc_call ─▶ publish → broadcast: createIpcEndpoint()
publishMouseEvent(m) for each sub whose callCap(subscribe,
publishJoystickEvent(j) mask matches event.device: send_cap = ep,
ipc_send(sub_ep) ──────▶ device_mask)
reply ok loop: next()
subscribe → store {ep cap, └─ replyWait(ep)
task id, device_mask} → InputEvent
```
- A **subscriber** calls `input.subscribe(mask)` — or a typed helper: `subscribeKeyboard()`,
`subscribeMouse()`, `subscribeJoystick()` (one class, `next()` returns the decoded event),
or `subscribeAll()` (every class, `next()` returns a tagged `InputEvent`)
([library/runtime/input.zig](../library/runtime/input.zig)). It creates its own endpoint
and hands it to the service as a **capability** (M13 capability passing — the input
service is that feature's first real user), along with its `device_mask`. Then it loops on
`next()`, a `replyWait` on that endpoint returning each pushed event.
- A **source** (a keyboard, mouse, or joystick driver) calls `input.connectSource()` and the
method for its class: `publishKeyboardEvent`, `publishMouseEvent`, or
`publishJoystickEvent`. Publishing is a short synchronous `ipc_call` the service answers at
once; the service's own fan-out is asynchronous, so publishing never blocks on a slow
subscriber.
- The **service** ([input.zig](../system/services/input/input.zig)) keeps a small subscriber
table (endpoint handle + owning task id + `device_mask`). On `publish` it `ipc_send`s the
event to every subscriber whose mask includes the event's device class. On `subscribe` it
stores the passed capability and mask and, as housekeeping, prunes any slot whose owning
process has exited (checked against `process_enumerate`) — not for correctness (an async
send to an orphaned endpoint is harmless) but to reclaim the slot.
Publisher and subscriber must be **separate processes**: a single thread that both
published and serviced its own subscription would deadlock (its `publish` call blocks until
the service delivers to its endpoint, which only the same thread could receive).
## Status and follow-ups
- **The keyboard is real.** The `ps2-bus` driver owns PNP0303, which carries *both* the
0x60/0x64 ports and IRQ1, so reading the hardware lives in the bus, not in
[keyboard.zig](../system/drivers/ps2-bus/keyboard.zig): the bus binds IRQ1 and, on each
interrupt, drains port 0x60, routing every byte by the status register's
auxiliary-output bit to whichever child driver **attached** for that device (an
`AttachRequest` to the well-known `ps2_bus` service, carrying the child's endpoint as a
capability; the bytes then arrive as asynchronous `ForwardedByte` messages, so the IRQ
path never blocks on a child). The keyboard driver decodes the stream — scancode **set 2**,
what the keyboard sends with the 8042's legacy translation off, decoded by
[scancode.zig](../system/drivers/ps2-bus/scancode.zig) into USB HID usage keycodes with
make/break, typematic-repeat, and modifier tracking (host-tested under `zig build test`) —
and publishes real `key_down`/`key_press`/`key_up` events.
- **Keycode → character** is wired in: the keyboard driver fills a `key_press` event's
`character` through [`library/xkeyboard-config`](../library/xkeyboard-config/README.md)
(`xkb.map(layout, keycode, mods)` → keysym + Unicode character), synthesizing the ASCII
control characters for Enter/Tab/Backspace/Escape, whose keysyms map to no Unicode. The
layout defaults to `us`; the bus can pass another as the driver's argv[2] — the seam for
a future settings source.
- **The mouse is real too.** IRQ12 is enumerated on the auxiliary device's own ACPI node
(PNP0F13), so the bus claims that node alongside the controller and routes both IRQs to
its one endpoint, acking whichever line the notification's badge names.
[mouse.zig](../system/drivers/ps2-bus/mouse.zig) attaches the way the keyboard does and
assembles the forwarded bytes with
[mouse-packet.zig](../system/drivers/ps2-bus/mouse-packet.zig) (three-byte stream-mode
packets: sync/overflow handling, nine-bit movement, screen-convention `dy` — host-tested
under `zig build test`) into `button_down`/`button_up` transitions and `motion` events.
**Follow-up:** the IntelliMouse magic-knock for a scroll wheel (four-byte packets) and
`scroll` events. The hardware-free `input-source` still rotates through all three classes
synthetically (including a joystick, which has no driver yet) via the
`input.synthetic*Event` helpers.
- **Drop-oldest under overflow** is a defined loss; the 16-slot ring absorbs normal bursts.
Real backpressure/flow-control is future work.
- **`publish` is unauthenticated** — any process may publish, consistent with the current
bring-up trust model (see [driver-model.md](driver-model.md)). A source capability is
future work.
## Verifying it
The `input` case (`python3 test/qemu_test.py input`, in
[tests.zig](../system/kernel/tests.zig) `inputTest`) boots the real kernel and spawns the
service, the synthetic source (which cycles keyboard, mouse, and joystick events), and a
subscriber that took all three classes. It passes only when the subscriber heartbeats
`input-test: ok` — proof that an event travelled source → service → subscriber over IPC,
exercising `ipc_send`, capability-passing subscription, and per-device routing. Each
serial line names the class received, so the log shows all three arriving on one stream.
## See also
- [ipc.md](ipc.md) — the synchronous rendezvous and the notification path `ipc_send` extends.
- [syscall.md](syscall.md) — the system-call surface, including `ipc_send`.
- [driver-model.md](driver-model.md) — class drivers, capability passing (M13), the trust model.
+15 -4
View File
@@ -80,11 +80,22 @@ inline). `build.zig` adds `isr.s` to the arch module.
## Reporting a fault
`isr_common` calls `exceptionHandler`, which forwards to a swappable `on_fault`
hook. The generic kernel installs a reporter (`onException` in `main.zig`) that
prints, in red, the exception name and vector, the error code, the faulting RIP
hook. The generic kernel installs a reporter (`onException` in `kernel.zig`) that
prints the exception name and vector, the error code, the faulting RIP
and RSP, and — for a page fault (#PF, vector 14) — the faulting address from
**CR2**. Then it halts. There's no fault *recovery* yet, so every exception is
terminal; the point is that it's now **visible** instead of a silent reset.
**CR2**. What happens next depends on where the fault came from:
- **User mode (CPL 3): kill the process, keep the machine.** The kernel is intact
(the CPU trapped onto the task's kernel stack), so the faulting process is
killed — address space, IRQ bindings, and IPC handles reclaimed; a client it
owed a reply to is failed with `-EPEER` — and the core reschedules. A crashing
driver takes itself down, never the OS. This is fault recovery step 2 of
[resilience.md](resilience.md). NMI, double fault, and machine check are
excluded: they report machine trouble regardless of what was running.
- **Kernel mode: halt this core.** The trusted base itself is broken, so there is
nothing safe to kill; the fault is still *contained* to the core (an
application-processor fault leaves the rest of the system running), and the
report makes it **visible** instead of a silent reset.
The hook is set before `arch.init()` in `kmain`, so a fault during setup is still
caught.
+25 -1
View File
@@ -95,6 +95,30 @@ This is what makes a user-space driver possible at all, and it's the subject of
every capability is either well-known (the registry) or inherited — there's no way
to delegate one.
- **Asynchronous / buffered send** for the cases where a rendezvous is the wrong
shape (logging, notifications between servers).
shape (logging, notifications between servers). *Landed as `ipc_send`* — a
non-blocking post to an endpoint's bounded payload queue, delivered through
`reply_wait` as a buffered message (badge bit `notify_message_bit`). Built for, and
first used by, the [input service](input.md)'s keyboard-event broadcast, where a
synchronous push would let one dead subscriber hang the fan-out. A full queue drops
the oldest (discrete messages, not a coalescing level like the notification ring).
- **A bounded reply.** `MSG_MAX` is 256 bytes and the copy runs under the big kernel
lock; a bulk transfer wants shared pages, not a copy.
## Lifecycle conventions over IPC (M17)
Three conventions from [process-lifecycle.md](process-lifecycle.md) ride the
notification mechanism:
- **Signals** arrive as notifications on the endpoint a process nominated with
`signal_bind` (`runtime.process.bindSignals`): badge = the signal bit plus the
coalesced pending mask (`runtime.process.signalsFrom` decodes). Statements,
never questions; no payload, no reply.
- **One-shot timers** (`timer_bind`, `runtime.system.timerOnce`) land as a
timer-bit notification — the timed wait: a service arms a deadline and keeps
serving, instead of blocking in sleep.
- **The universal ping**: a **zero-length request is the liveness probe**,
answered with a zero-length reply by the service harness itself
(`runtime.service.run`). No protocol's requests start at length zero, so the
encoding cannot collide, and a wedged service simply fails to answer — which
is the diagnosis. Deep health ("can I reach my hardware?") stays a per-service
protocol message.
+210
View File
@@ -0,0 +1,210 @@
# M17–M18 execution plan: process lifecycle + device manager
**Archived — completed 2026-07-13** (every item checked; suite ended 54/54).
Kept as the record of how M17–M18 landed; the successor is
[m19-m20-plan.md](m19-m20-plan.md).
The operational plan for building [process-lifecycle.md](process-lifecycle.md)
(M17) and [device-manager.md](device-manager.md) increments 5–7 (M18). Design is
settled in those documents; this file is the build order — one phase at a time,
each phase green before the next starts. Delete or archive this file when M18
lands.
**Definition of green, every phase:** `zig build` clean, `zig build test` clean,
`python3 test/qemu_test.py` passes (existing scenarios plus the phase's new one),
and the relevant design doc's "known gaps" / status lines updated. Commit per
green phase (no co-author trailers).
**Workflow (settled 2026-07-12):** work happens in a dedicated git worktree, on
feature branches cut from `main` — `feat/process-lifecycle` (M17.1–17.4),
`feat/device-manager` (M18.1), `feat/usb-xhci-bus` (M18.2–18.3). When a branch's
phases are all green it is **auto-merged into `main`**; branches are kept after
merge, not deleted. Merges and branches are pushed to origin. Phase 0 (once):
commit the design docs, merge the outstanding `feat/usb` work into `main`, and
run the existing QEMU suite green before any new work starts.
**Numbering note:** continues the milestone sequence (driver track ended at M16).
## Status
The loop marks a phase `[x]` in the same commit that lands it. A phase is marked
only when its definition of green holds.
- [x] **Phase 0** — baseline: docs committed, feat/usb merged to main, pushed;
`usb-xhci-libary.zig` renamed to `usb-xhci-library.zig`; existing QEMU
suite green from the worktree (48/48, 2026-07-12).
- [x] **M17.1** — kernel releases claims/MSI on death (claims: `releaseAllOwnedBy`
in the reap; MSI was already swept by `irq.releaseOwner`; `claim-release`
test; suite 49/49)
- [x] **M17.2** — exit reasons (`ExitReason` recorded at exit/fault/kill before
the notification; `process_exit_reason` supervisor-gated;
`runtime.process.exitReason`; kernel + ring-3 assertions; suite 49/49)
- [x] **M17.3** — published exit events + VFS subscriber (`process_subscribe`,
bounded ref-counted table, publish on every death;
`runtime.process.subscribeExits`; VFS handles carry owners and are swept on
the owner's death; `vfs-client-death` test; suite 50/50)
- [x] **M17.4** — signals, timer notifications, `runtime.process`, the service
harness (signal_bind/process_signal + coalescing pending mask; timer_bind
on the tick; bindSignals/signalsFrom/sendSignal/stop + timerOnce;
runtime.service.run with the zero-length ping; VFS converted; `signals`
scenario; suite 51/51)
- [x] **merge** `feat/process-lifecycle` → main, push (merged 2026-07-13)
- [x] **M18.1** — device-manager protocol: hello + restart policy
(device-manager-protocol module; the manager as a harness service:
supervised spawns, hello deadline via timer sweep, restart with
300/600/1200ms backoff, exit reasons deciding restart-vs-stopped,
crash-loop cap; usb-xhci-bus first conforming driver; crash-test fixture
re-proving claim release each respawn; `driver-restart` scenario;
maximum_tasks 16→32 — the sweep was overflowing the pool; suite 52/52)
- [x] **merge** `feat/device-manager` → main, push (merged 2026-07-13)
- [x] **M18.2** — xHCI port scan + tree reports (child_added/child_removed in
the protocol; the manager's child mirror with death-pruning; xHCI maps the
register BAR — resource 0 is ECAM — reads CAPLENGTH/HCSPARAMS1, scans
PORTSC, reports connected ports with speed-class identity; `usb-report`
scenario proves report → prune → respawn → re-report; suite 53/53)
- [x] **M18.3** — app surface: enumerate/subscribe over IPC (subscriber
endpoint rides as the call's capability; events are the same structs the
buses send); device-list first client; protocol capped at the kernel's
IPC MESSAGE_MAXIMUM (256); the startUserTask debug print removed — it
sheared concurrent serial lines and was the scenario-flake root cause;
`device-list` scenario; suite 54/54)
- [x] **merge** `feat/usb-xhci-bus` → main, push (merged 2026-07-13) — **plan complete**
---
## M17.1 — the kernel releases a dead process's claims
The cleanup half of iron rule 1; the prerequisite for every restart story.
- `system/kernel/devices-broker.zig`: `releaseAllOwnedBy(owner: u32)` — clear
every `claimed[]` slot holding this task id.
- `system/kernel/process.zig`: call it from the reap path, alongside the existing
IRQ-binding release (the ordering comment there says why IRQs go first — claims
slot in after them, before the exit notification).
- MSI vectors: find where `msi_bind` records per-device vectors (interrupts
module) and release those by owner in the same pass.
- Docs: remove the claims bullet from process-management.md "Known gaps".
**Test:** new QEMU scenario `claim-release` — a test child claims an unclaimed
device, is killed, is respawned, and claims the same device again successfully;
assert both claims in the serial log. Kernel-side unit coverage in
`system/kernel/tests.zig` for `releaseAllOwnedBy` (claim two devices as two owners,
release one owner, verify exactly its claims freed).
## M17.2 — exit reasons
- `system/abi.zig`: `ExitReason` (exited, aborted, segmentation_fault,
illegal_instruction, arithmetic_fault, killed).
- Kernel: record the reason at every death site — clean exit path, each fault
class in `onException`, the kill path. Bounded recent-exits table (ids are never
reused, so a small ring keyed by id is enough).
- New system call `process_exit_reason(id)` — supervisor-gated, like kill; returns
the recorded reason or `-ESRCH` once evicted.
- `library/runtime/process.zig`: `ExitReason` + `exitReason(id: u32)`.
- Docs: remove the no-exit-status bullet from process-management.md.
**Test:** extend the `supervision` scenario — three children: one exits cleanly,
one faults (the fault-recovery pattern), one is killed; the supervisor asserts all
three reasons.
## M17.3 — published exit events
- Kernel: bounded subscriber table (endpoints); new system call
`process_subscribe(endpoint)` (ungated, like `process_enumerate`); every death
posts `notify_exit_bit | id` to each subscriber — the same post the supervisor
path already uses.
- `library/runtime/process.zig`: `subscribeExits(endpoint)`.
- VFS becomes the first subscriber: on an exit event, release every handle keyed
by that task id (badges already are task ids). Log the release.
- Docs: note the convention in ipc.md (exit events reuse the exit-notification
badge encoding).
**Test:** new QEMU scenario `vfs-client-death` — a client opens a file and is
killed without closing; assert the VFS logs the handle release and its open-handle
count returns to baseline.
## M17.4 — signals and the service harness
- Kernel: per-task pending mask + bound endpoint; system calls
`signal_bind(endpoint)` and `process_signal(id, signal)` (supervisor-or-self
gated); delivery posts `notify_signal_bit | pending mask`, coalescing; pending
signals with no bound endpoint pend silently.
- `library/runtime/process.zig`: `Signal`, `SignalSet`, `bindSignals`,
`signalsFrom`, `sendSignal`, `stop(id, deadline_ms)` (terminate → wait for exit
notification → kill). Implement `terminate`, `reload`, `user_1`, `user_2`;
`interrupt`/`quit` are enum members with no sender yet; `alarm` stays unbuilt.
- Kernel: **one-shot timer notifications** — `timer_bind(endpoint, ms)` posts a
notification badge when the deadline lands (IRQ-as-IPC again, on the timer
wheel `sleep` already uses). This is the missing timed-wait primitive:
`replyWait` blocks forever and `sleep` blocks the whole process, but `stop()`'s
escalation, the device manager's `hello` deadline (M18.1), and restart backoff
all need a deadline while staying responsive. It is also the mechanism `alarm`
gets for free later.
- New `library/runtime/service.zig`: the harness — `run(callbacks)` owning the
replyWait loop, folding protocol messages, signals, and child-exit notifications
into `init` / `on_message` / `on_reload` / `on_terminate`; answers the common
`ping` automatically. Define the reserved `ping` request encoding here and
document it in ipc.md (one obvious encoding; smallest that cannot collide with
existing protocols).
- Convert one existing service (input-source or hpet) to the harness as proof it
subtracts code rather than adding it.
**Test:** extend `supervision` — a harness-built child: `sendSignal(reload)`
observed in its log, `ping` answered, `stop()` produces a clean exit with reason
`exited`; a second child that ignores signals (no bind) is killed by `stop()`'s
deadline with reason `killed`.
## M18.1 — device-manager protocol: hello + restart policy
- New `system/services/device-manager/device-manager-protocol.zig` module
(vfs-protocol pattern): `hello { version, role, device_id }`; version constant;
reserved fields.
- Device manager: register the `.device_manager` endpoint; spawn drivers with its
exit endpoint; enforce the hello deadline; restart policy — backoff, crash-loop
cap (three fast deaths → mark failed, log, stop), reasons from M17.2 deciding
restart vs not.
- usb-xhci-bus: adopt the harness + send hello. hpet/ps2-bus follow only if the
conversion is mechanical; otherwise they keep working unconverted (the manager
only enforces hello on drivers spawned with an assignment).
- build.zig: test-loop entry for the protocol module if it grows pure logic.
**Test:** new QEMU scenario `driver-restart` — the xHCI driver takes a test-only
argv flag to fault after hello on its first run; assert: fault, exit reason
recorded, manager respawns with backoff, second run claims the controller
(M17.1) and hellos clean. Assert the crash-loop cap by a driver that always
faults (a tiny test driver, not xhci).
## M18.2 — bus tree reports
- Protocol: `child_added { parent, identity, resources }` / `child_removed { id }`.
- usb-xhci-bus: bring-up to **port scan only** — map the MMIO window (claimed in
M16-era work), controller reset/start per xHCI spec, walk the port registers,
report one `child_added` per connected port with speed + port number as
identity. **No transfer rings, no descriptors** — reading device/interface
descriptors (and therefore USB class triples for matching) is the follow-on USB
track, not this plan.
- Device manager: mirror reports into its tree; prune the subtree (emitting
`child_removed`) when a bus driver dies; assert re-report on restart.
**Test:** QEMU already attaches usb-kbd + usb-mouse on xhci.0 — assert two
`child_added` events reach the manager and appear in its tree dump; kill the
driver, assert two `child_removed` then two fresh `child_added` after respawn.
## M18.3 — the application surface
- Protocol: `enumerate` (tree snapshot) + `subscribe` (published add/remove
events, input-service pattern).
- A small client (`device-list`, the `ps` analog) exercising both; the manager
becomes the one answer to "what devices exist" for user space.
`device_enumerate` stays for drivers/kernel seeding — its retreat is tied to the
discovery migration, out of this plan.
**Test:** QEMU scenario — `device-list` shows the tree including USB children;
during a driver restart the subscribing client logs remove + add events.
---
**Explicitly out of scope** (own tracks, after M18): discovery migration (pci-bus
driver, acpi service, retiring the kernel scan), USB control transfers +
descriptors + class-driver matching, the musl layer, `interrupt`/`quit` senders
(needs a console), job control.
+196
View File
@@ -0,0 +1,196 @@
# M19–M20 execution plan: discovery migration
The operational plan for [device-manager.md](device-manager.md)'s increment 8:
discovery leaves the kernel — a **pci-bus driver** (M19) and an **acpi service**
(M20), with the kernel's device enumeration retired behind them. Same rules as
[m17-m18-plan.md](m17-m18-plan.md): one phase at a time, each green before the
next; this file is the build order and the checklist.
**Definition of green, every phase:** `zig build` clean, `zig build test` clean,
`python3 test/qemu_test.py` passes (existing scenarios plus the phase's new
one), and the relevant design doc updated. Commit per green phase (no co-author
trailers). The full suite is the regression net — the existing
`driver-restart` / `usb-report` / `device-list` / `input` scenarios must stay
green *through* the migration, which is the whole point: the system must not be
able to tell who enumerated it.
**Workflow:** dedicated worktree; branches off `main` — `feat/pci-bus`
(M19.0–19.3), `feat/acpi-service` (M20.1–20.3); auto-merge to main when a
branch is green; keep branches; push everything.
## Settled decisions (2026-07-13 — veto before the loop starts)
1. **What "retiring the kernel scan" means.** The kernel keeps, forever, the
parses it needs before user space exists: RSDP/XSDT location, MADT (SMP),
the HPET table (the tick), FADT + the AML `\_S5` evaluation (poweroff — the
power tests prove it), and MCFG (the host bridge node). What retires is
**device enumeration**: the ECAM function walk (M19.3) and the DSDT/SSDT
namespace walk that builds device nodes (M20.3). The AML module stays a
shared build module compiled into both the kernel (for `\_S5`) and the acpi
service (for everything else) — same source, two builds, no fork.
2. **Bridge apertures come from the firmware memory map, not AML.** Registered
PCI functions carry BAR resources, and containment demands the bridge own
windows that cover them. The apertures are derived kernel-side from the
boot memory map's MMIO holes (regions that are neither RAM nor tables) —
mechanical, AML-free, and available at boot regardless of what later moved
to user space. (The bridge today carries only ECAM + bus range; this is the
prerequisite M19.0 exists for.)
3. **`device_register` becomes idempotent on exact match.** A re-registration
with identical (parent, class, resources) returns the existing id instead
of appending. The kernel table has no unregister, so without this a
restarted registering bus would duplicate its children on every respawn —
idempotence makes restart-and-re-report safe for every future bus, not just
PCI.
4. **The manager matches from reports.** `ChildAdded` gains a `device_id`
field (the kernel-registered id, `no_device` for unregistered leaves like
USB ports). After the M19.3 flip, PCI driver matching keys off reported
identity (the class triple) instead of the manager's boot-time snapshot —
the snapshot match remains only for what the kernel still seeds. One flip
phase changes both sides at once so no device is ever matched twice.
5. **The acpi service's authority is one node.** The kernel publishes an
`acpi-tables` device: memory resources covering the table blobs plus a
broad `io_port` resource — the documented trust grant to exactly one
process (AML OperationRegions reach EC/PM ports; the claim-gated
io_read/io_write calls already exist). The service claims it, maps the
tables, and runs the shared AML module in ring 3 behind a `Hal` backed by
`mmio_map` + `io_read`/`io_write`.
6. **Both new processes are protocol drivers** under the manager: hello,
supervision, restart with backoff — all inherited from M18.1 for free.
Registration idempotence (decision 3) is what makes their restarts sound.
7. **Firmware neutrality is the contract** (2026-07-13). The generic layer is
everything at and above the device-manager protocol — descriptors,
containment, reports, matching, supervision — and none of it may become
x86-specific. Discovery is one swappable process per firmware: the acpi
service on x86; an **fdt service** on the Raspberry Pis (claims a
`devicetree-blob` node, reports children from the flattened device tree —
pure data, no bytecode, no port grant, strictly simpler than ACPI). The
manager owns the tree as *data* and touches no hardware, ever — AML runs in
a crashable, supervised discoverer precisely so a firmware-bytecode fault
can never take down the supervisor. Two consequences recorded now:
`DeviceDescriptor`'s 8-byte `hid` cannot hold an FDT `compatible` string
("brcm,bcm2835-aux-uart") — identity widens before the fdt service exists;
and cross-firmware surfaces are named by **domain, not firmware** (M21
defines a *power* protocol, not an "ACPI events" protocol — PSCI/mailbox
sources feed the same subscribers on ARM). **Landed early (2026-07-13):**
both services exist as placeholders (system/services/acpi, system/services/
fdt) and the build's `-Ddiscovery=acpi|fdt` option fills the ramdisk's
neutral `discovery` slot — the manager will spawn "discovery" by that name
in M20.3 and never learn which firmware it is on.
## Status
- [x] **M19.0** — prerequisites (bridge apertures from the memory map's
*gaps* — the single-hole rule died on OVMF's flash at the top of 4 GiB,
caught by the new every-BAR-contained assert in `discovery`; idempotent
`device_register` proven in `bus`; `ChildAdded.device_id`;
m17-m18-plan.md archived; suite 54/54).
- [x] **M19.1** — pci-bus driver, scan only (claims the bridge, maps ECAM
through its grant, brute-force walk with the multifunction rule; the
manager matches pci_host_bridge → pci-bus per device with the full
protocol contract; `pci-scan` builds its expected marker from the
kernel's own count — equivalence on the first run; suite 55/55).
- [x] **M19.2** — register + report (BAR probe mirrored byte-for-byte from the
kernel's addBars so dedupe returns the kernel's node ids during
coexistence; the bridge gained the io_port aperture I/O BARs need;
reports carry the registered device_id; pci-scan drills a forced restart
and asserts the PCI node count never grows — plus harness hardening: a
failing case now preserves its serial as <case>-failed-serial.log, and
the heavy scenarios run at 150s; suite 55/55).
- [x] **M19.3** — the flip: kernel `enumeratePci`/`addBars`/`PciHeader` all
deleted (bridge node stays); manager matches PCI drivers from reported
identity, deduped by registered id. Surfaced and fixed a real SMP race the
flip created — ring-3 device_register made the broker table concurrent, so
mmio_map's lock-free read intermittently tore hpet's resource length
(user fault) and overflowed `r.len-1` into a kernel panic; now the broker
read is under the big lock and the arithmetic is guarded, and pci-bus
skips size-0 BARs. discovery.md updated; suite 55/55 (driver-restart
hammered 6×).
- [x] **merge** `feat/pci-bus` → main, push (merged 2026-07-13).
- [x] **M20.1** — acpi service, parse only: the AML interpreter is now a build
module compiled into both kernel and service; the kernel publishes the
`acpi-tables` node (AML blobs as memory resources, the broad io_port grant,
the SCI); the service claims it, maps the blobs, runs the shared parser in
ring 3, and self-verifies its Device count against the kernel's (34 = 34,
deterministic via argv, no log-scraping); the manager spawns `discovery`
at startup. Parse-only touches no hardware. Suite 56/56.
- [x] **M20.2** — register + report: the service evaluates `_STA`/`_CRS` in
ring 3 (interpreter Hal = port I/O over the claimed node; a scratch page
backs SystemMemory maps so a stray region can't fault it) and registers +
reports each present `_HID` device under `acpi-tables`. Containment: the
broker's irq check became range-based (len-1 == the old equality) so the
node's broad irq window covers children's legacy lines; io ports fall in
the broad io grant. ChildAdded gained `hid`. Matching stays off. The
`acpi-report` scenario asserts the PS/2 keyboard (3 resources) and mouse
(1 resource) among the reports. Suite 57/57.
- [x] **M20.3** — the flip: the kernel's `wireAcpiDevices` call is gone (the
device-building helpers are retained-but-dead pending a focused sweep,
spawned as a task; static tables + `\_S5` + the acpi-tables node stay).
The manager matches ps2-bus from ACPI `_HID` reports; the service
registers all devices before reporting any (no keyboard-before-mouse
race). The `acpi-ps2` scenario proves report → spawn → ps2-bus attaches
its keyboard; `ioport` retargeted to the acpi-tables I/O window (the
kernel-built PS/2 node is gone). Suite 58/58.
- [x] **merge** `feat/acpi-service` → main, push (merged 2026-07-13) — **discovery migration complete**.
---
## Phase notes
**M19.0 apertures:** the boot memory map already crosses the handoff
([boot-handoff]), but discovery never sees it today — expect a small
pass-through (kernel init hands the map to the platform layer) before the
holes computation, which belongs where the bridge node is built
(`parseMcfg`). Sanity-check on QEMU q35: the xHCI BAR (`0xc0000000`-region
values seen in the M18 logs) must land inside a derived aperture, asserted in
the kernel unit test.
**M19.1 scanning without owning config access twice:** the driver reads config
space through its ECAM mmio_map grant of the *bridge* window — the same bytes
the kernel walk read. Vendor-id `0xFFFF` skip, header-type multifunction rule,
no bridge recursion (matches the kernel's current single-segment walk).
**M19.2 BAR sizing:** the classic size probe (write all-ones, read mask,
restore) is deferred — the BARs' current programmed values and types are
enough for containment-checked registration at bring-up; sizing lands with the
first driver that needs to *move* a BAR. Log what is registered so the
scenario can assert it.
**M19.3 what the manager still seeds from the snapshot:** everything the
kernel still enumerates (timers, ACPI nodes until M20.3). The PCI arm of
`pciDriverFor` switches source; `driverFor` doesn't move until M20.3.
**M20.1 spawn and identity (pre-settled 2026-07-13):** the manager spawns
`discovery` by its neutral ramdisk name at startup, as an ordinary protocol
driver (hello, supervision) — from M20.1 on, on every boot. For reporting ACPI
devices, `ChildAdded` gains `hid: [8]u8` (EISA ids fit; zero = none):
firmware *string* identity travels beside the numeric `identity` field until
the FDT-driven widening replaces both (decision 7).
**M20.1 Hal in ring 3:** `mapMmio` → `device.mmioMap` over the claimed
acpi-tables node (plus a table-offset map for blobs); `pioRead`/`pioWrite` →
`device.ioRead`/`ioWrite` against its io_port resource. The interpreter cannot
tell it moved — that is the assertion of `acpi-parse`.
**M20.2 containment for `_CRS`:** io ports fall inside the node's broad
io_port resource; MMIO windows (HPET, LAPIC ranges some firmwares list) fall
inside the memory-map holes added to the node in M20.1. Anything that doesn't
fit is logged and skipped, loudly — bring-up honesty over silent drops.
**M20.3 ps2 ordering:** ps2-bus binds nodes the acpi service now reports, so
its spawn moves behind the report (the manager's matching handles this once
the source flips); the `input` scenario proves the keyboard still types.
**Explicitly out of scope:** PCI bridge recursion (single segment, flat bus
walk stays); BAR reprogramming/sizing; disk/PCIe hotplug; interrupt routing
changes (`_PRT` stays wherever it is today); the USB descriptor track;
multi-segment ECAM; per-device power states (D-states, `_PSx`/`_PRx`,
suspend/resume — a future *lifecycle-vocabulary* extension, since "suspend"
has the shape of a signal every driver must answer, and it has no consumer
until laptop sleep); CPU P/C-states.
## M21 — ACPI events + system power — DONE
Built and merged (docs/m21-plan.md, 2026-07-13): the SCI + power button, Notify/GPE
dispatch, and orderly shutdown (init's stop cascade into a ring-3 S5 write).
See that plan for the phase record.
+139
View File
@@ -0,0 +1,139 @@
# M21 execution plan: ACPI events + system power
The operational plan for the event side of the acpi service and orderly
shutdown — the capstone [m19-m20-plan.md](m19-m20-plan.md) previewed. Same
rules as its predecessors: one phase at a time, each green before the next;
this file is the build order and the checklist.
**Definition of green, every phase:** `zig build` clean, `zig build test`
clean, `python3 test/qemu_test.py` passes (existing scenarios plus the
phase's new one), and the relevant design doc updated. Commit per green phase
(no co-author trailers). Failing cases preserve their serial logs
(`<case>-failed-serial.log`).
**Workflow:** dedicated worktree; branch `feat/power-events` off `main`;
auto-merge to main when the branch is green; keep the branch; push everything.
## Settled decisions (2026-07-13, approved)
1. **S5 is executed by the acpi service from ring 3.** No new syscall: the
broad port grant (M20 decision 5) already made this physically possible —
the service holds the PM1 control ports in its io grant and derives `_S5`
from its own namespace (`aml.sleepState`). Formalizing it adds no
authority. The kernel keeps `power.zig` for its own test paths and
panic-time use.
2. **The power surface is domain-named** (decision 7 of the last plan): a
`power-protocol` module + `ServiceId.power = 5`, registered by the acpi
service — on ARM, a PSCI/mailbox service registers the same id and
subscribers never know the difference. Messages: `subscribe` (endpoint as
the call's capability, the input/manager pattern), `shutdown` (accepted
only from PID 1 — init), and events published as buffered messages:
`power_button`, `lid`, `ac`, `battery`, generic `notify` with a code.
3. **The service learns event ports from its own FADT copy**: the kernel adds
the FADT as one more memory resource on the acpi-tables node; the service
tells it apart from the AML blobs by signature ("FACP" header — the blob
resources are header-stripped bytecode and start with no signature). The
kernel's own FADT parse is untouched.
4. **The acpi service converts to the harness** (`runtime.service.run`):
protocol messages (subscribe/shutdown), the SCI notification, and the
existing report flow fold into one loop — the shape it was always meant
to have.
5. **GPE/Notify correctness is proven by host unit tests** (synthetic AML
with a Notify inside a method body; aml.zig joins the `zig build test`
loop). The QEMU scenario proves the power button — a *fixed* event,
deterministically injectable via QMP `system_powerdown` — because QEMU
cannot raise GPEs deterministically on this config. Battery/AC/lid and the
embedded controller (`_Qxx`) are interface-complete here and validated on
real hardware (the laptop) later.
## Ground truth the phases build on (verified 2026-07-13)
- `system/devices/power.zig` `shutdown()` is the kernel's S5 write
(SLP_TYP|SLP_EN to PM1a/PM1b control); there is no power syscall.
- init (`system/services/init/init.zig`) spawns vfs/input/device-manager
fire-and-forget — no child ids kept, no signals, no event loop. The whole
stop toolkit exists in `runtime.process` (stop/sendSignal/bindSignals).
- `test/qemu_test.py` has no QMP channel (serial is a one-way file).
- The kernel parses PM1 *control* blocks and SCI_INT from the FADT; the PM1
**event** blocks (offsets 56/60, len at 88) and **GPE0/GPE1** blocks
(offsets 80/84, lens 92/93) are unparsed — the service reads them from its
FADT copy (decision 3).
- The acpi-tables node carries the SCI as its only `len == 1` irq resource
(the broad window is len 256) — that is how the service finds it to
`irqBind`.
- `notify_opcode = 0x86` exists in `system/devices/aml/opcodes.zig` but the
interpreter never handles it — a GPE `_Lxx` body containing Notify fails
evaluation today. Everything else a GPE handler needs (field access,
control flow, method calls) is proven by the ring-3 `_STA`/`_CRS` work.
- The dead-code sweep (spawned task) also edits `system/devices/acpi.zig`;
M21.0 checks whether it landed and rebases before touching that file.
## Status
- [x] **M21.0** — baseline (dead-code sweep confirmed landed on main — no
acpi.zig conflict; `feat/power-events` cut; QMP channel in the harness:
always-on unix socket, client with the capabilities handshake, per-case
`qmp_after` hook, and a hook-must-deliver pass gate that the smoke case
now proves with a harmless query-status; suite 58/58).
- [x] **M21.1** — SCI + the power button (kernel appends the FADT as an
acpi-tables memory resource, tagged by its "FACP" header; `power-protocol`
module + `ServiceId.power = 5`; the acpi service converted to
`runtime.service.run`, registers `.power`, reads PM1 event/control + GPE
ports from its FADT copy, enables ACPI mode if SCI_EN is clear, binds the
SCI (the len-1 irq), sets PWRBTN_EN; the SCI handler clears PM1_STS,
logs `power: button pressed`, publishes `power_button`, acks. Scenario
`power-button` injects a real `system_powerdown` via QMP; initial-ramdisk
timeout 30→60s for the service's added boot work; suite 59/59).
- [x] **M21.2** — Notify + GPE dispatch (interpreter handles `notify_opcode`
into a bounded queue, cleared per-evaluate, drained via
`takeNotifications`; the service walks GPE status/enable bytes, evaluates
`\_GPE._Lxx`/`_Exx` per active bit, maps notified nodes to events
(battery/ac/lid/generic), clears GPE_STS write-1, acks. EC `_Qxx` out.
Host unit test with hand-encoded AML proves the queue; aml.zig joined the
`zig build test` loop. QEMU raises no GPEs — suite is regression net,
59/59).
- [x] **M21.3** — orderly shutdown (init supervises its children on one
endpoint that also carries signals, power events, and a re-arming
heartbeat timer; on `power_button` or a `terminate` signal it logs
`init: shutting down`, runs `stop(child, 2000, endpoint)` in reverse
order, then requests `.power` shutdown; the acpi service honors shutdown
from a subscriber — init is the one subscriber, a soft gate that survives
testing where PID 1 isn't init — and writes SLP_TYP|SLP_EN from ring 3.
`orderly-shutdown` scenario proves button → shutting-down → S5 → QEMU
exit; suite 60/60).
- [ ] **merge** `feat/power-events` → main, push, keep the branch — **loop
ends here**.
---
## Phase notes
**M21.0 QMP:** open the unix socket after Popen, complete the
`qmp_capabilities` handshake, then send the hook's command (for these
scenarios: `{"execute": "system_powerdown"}`). The socket is additive — no
existing case may notice it. Note e3fe3f3 recently reworked how the harness
boots; adapt to its current shape rather than the pre-rework description.
**M21.1 SCI details:** PM1_STS is at the event block base (write-1-to-clear);
PM1_EN at base + block_len/2; PWRBTN bit is 8 in both. If PM1b exists, mirror
reads/writes to both blocks. Enable ACPI mode only when SCI_EN (PM1 control
bit 0) is clear — OVMF boots may already have it set. The publish path reuses
the manager's subscriber table pattern (bounded, drop-on-failed-send).
**M21.2 GPE walk:** GPE0_STS bytes live at the GPE0 block base, GPE0_EN in
the block's upper half; for a set+enabled bit n, the handler method is
`_L%02X` (level) or `_E%02X` (edge) under `\_GPE`. Evaluate, drain the notify
queue, clear the status bit, ack. A missing handler method is clear-and-log,
not an error.
**M21.3 ordering:** init subscribes with retries — the acpi service registers
`.power` well after init starts. The stop sequence runs vfs last (other
services may flush through it). The S5 write mirrors `power.zig`'s
`sleepValue` (SLP_TYP bits [12:10], SLP_EN bit 13); if the write returns, log
`power: S5 write did not take` so the scenario fails loudly instead of
hanging.
**Explicitly out of scope:** the embedded controller and `_Qxx` queries,
battery `_BST`/`_BIF` evaluation beyond the interface stubs, lid/AC on QEMU
(no emulation), reboot over the power protocol, S3 sleep, per-device D-states
(a future lifecycle-vocabulary extension), thermal zones.
+327
View File
@@ -0,0 +1,327 @@
# Process lifecycle: signals over IPC
**Status: increments 1–4 built** (2026-07-12): claim release on death, exit
reasons, published exit events, and signals + one-shot timers + the service
harness are all in — the interface below is as-built. The primitives underneath
predate this design ([process-management.md](process-management.md):
spawn, the supervision link, kill, child-exit notifications); this document designs
the layer above them — the standard vocabulary a danos process speaks about its own
life, and the stable `runtime.process` interface that carries it. Nothing here is
device- or driver-specific: a driver, the VFS, and a user application all stop,
reload, and die the same way. The device manager is simply this design's first
serious customer ([device-manager.md](device-manager.md)).
**"POSIX" in this document means the concepts, never the letter of the standard.**
danos borrows the ideas and the hard-won lessons (what SIGTERM *means*, why SIGPIPE
was a mistake) without inheriting the mechanism, the API, or the names. The naming
rule is danos's own and it is strict: plain words that communicate intent
(`terminate`, `reload`, `exited`) and the IPC vocabulary the system already speaks
(`bind`, `subscribe`, `publish`, `endpoint`) — never `SIG*`, never a second word for
a concept that already has one. Literal POSIX arrives later and lives elsewhere: a
**musl-based C layer** (growing out of library/posix) that wires C programs to the
danos runtime — musl's syscall surface retargeted at danos system calls and IPC
protocols (files onto the VFS protocol, `sigaction`/`wait` onto this lifecycle,
sockets onto whatever networking becomes). Ported programs see POSIX; the system
underneath never does.
## Why a standard vocabulary
A supervisor can only manage processes it has never heard of if "please exit" means
the same thing to all of them. That is the one thing POSIX signals got deeply right:
`SIGTERM` means the same thing to nginx and to a five-line script, which is why
process supervision on Unix (init systems, container runtimes) is possible at all.
danos wants that property from day one, because supervision-and-restart is the
system's core motivation ([resilience.md](resilience.md)).
What POSIX got wrong — for a system like this — is the **delivery mechanism**:
asynchronous control-flow hijack. A Unix handler runs on a stolen stack at an
arbitrary instruction boundary, which is why the async-signal-safe function list
exists, why `errno` must be saved, and why the canonical signal bug is a SIGTERM
handler innocently calling `printf` mid-`malloc`. That entire bug class comes from
the mechanism, not the vocabulary, and none of it is worth importing.
A microkernel already has the right channel: **a signal is a message.** QNX delivers
POSIX signals over its message passing; seL4 has notification objects; Erlang turned
"death is a message to whoever linked" into a reliability philosophy. danos has
already done it once without naming it: a child's death arrives as a notification
badge on the supervisor's endpoint — the microkernel's SIGCHLD, the IRQ-as-IPC
pattern reused. Signals are the same pattern reused a third time.
## The mechanism
- **`signal_bind(endpoint)`** — a process nominates the endpoint its signals arrive
on, exactly as `irq_bind` nominates where a device's interrupts land. The runtime
does this at startup for any program that opts in.
- **`process_signal(id, signal)`** — posts the signal as an asynchronous
notification to the target's bound endpoint: badge = `notify_badge_bit |
notify_signal_bit | pending signals`. Non-blocking for the sender, always.
- **Pending signals coalesce** in a per-process bitmask until the target next waits
— exactly like interrupt notifications, and exactly POSIX's own semantics for
non-realtime signals (two pending SIGTERMs are one SIGTERM). The bitmask *is* the
design: signals carry no payload. Anything with a payload is a protocol message.
- **Authority**: the supervisor may signal its children — the same link that is
already the kill authority. A process may signal itself. Anything broader waits
for transferable process handles.
- **No binding, no problem**: a process that never calls `signal_bind` is not
broken — its signals pend unread and only `process_kill` works on it. Simple
programs stay simple; the vocabulary is opt-in, the kill authority is not.
Because delivery is a message into the process's own event loop, there is no
async-signal-safe list in danos: a handler is ordinary code running at a point the
process chose. The bug class is gone by construction, not by discipline.
## The vocabulary: POSIX.1-1990, sorted honestly
The full 1990 set, and what each becomes. Two intrinsically problematic cases get a
defense below the table.
| POSIX.1-1990 | danos disposition | Notes |
|---|---|---|
| SIGTERM | signal `terminate` | finish up and exit; the supervisor's polite half |
| SIGHUP | signal `reload` | re-read configuration / re-scan |
| SIGINT | signal `interrupt` | interactive interrupt; meaningful once a console can send it, in the vocabulary now so numbering is stable |
| SIGQUIT | signal `quit` | as SIGINT, without the core-dump baggage |
| SIGALRM | signal `alarm` | timer expiry as a message; the Unix SIGALRM+`longjmp` timeout hacks are impossible here. In the vocabulary, unbuilt: no consumer yet, and when one appears it is runtime sugar over the existing timer — zero kernel work |
| SIGUSR1, SIGUSR2 | signals `user_1`, `user_2` | service-defined |
| SIGCHLD | **already exists** — the exit notification | the badge carries the child id, dodging the classic coalescing bug (Unix code must loop `waitpid`) |
| SIGKILL | `process_kill` — kernel mechanism | its definition is "cannot be handled"; it was never really a signal |
| SIGABRT | exit reason `abort` | `abort()` is synchronous self-termination, not an event |
| SIGSEGV, SIGILL, SIGFPE | exit reasons, **never delivered** | see below |
| SIGPIPE | **an error return**, not a signal | see below |
| SIGSTOP, SIGTSTP, SIGTTIN, SIGTTOU, SIGCONT | deferred | job control needs terminals, sessions, and process groups; stop/continue is scheduler territory |
**The fault signals (SIGSEGV, SIGILL, SIGFPE) are intrinsically wrong for messages.**
They are *synchronous* — raised at a specific faulting instruction, not "sometime
soon". A message cannot be delivered to a process whose next instruction re-faults;
it never reaches its event loop to read it. POSIX only makes fault handlers "work"
via the async hijack (run the handler *instead of* the instruction), and even there,
returning from a SIGSEGV handler without curing the cause is undefined behavior.
danos's architecture already has the better answer: fault → the kernel kills the
process ([resilience.md](resilience.md) step 2, built) → the supervisor reads the
reason → restart. Recovery is restart, not a handler. This is also truer to the 1990
standard than handling is: the standard's default action for all three was
"terminate the process".
**SIGPIPE deserves special contempt.** Its default kills a process that writes to a
closed pipe — which is why "the whole server died because one client disconnected"
is roughly every network daemon's first production bug, and why every mature codebase
contains the same fix: ignore SIGPIPE, handle the `EPIPE` error return. danos made
the right choice natively already — a reply owed to a dead peer fails with `-EPEER`.
Errors from operations are error returns from those operations. The posix layer can
synthesize SIGPIPE for ported code that expects it.
### Statements, not questions
A signal and a protocol message both travel over IPC — the difference is the
**contract**, not the transport. danos IPC has two primitives, both already in
daily use: the **asynchronous notification** (a badge — bits that coalesce into a
pending mask; the sender never blocks; no payload, *no reply path*; how IRQs and
exit events arrive) and the **synchronous call** (a rendezvous — payload both
ways, the caller waits for the reply; how VFS requests work). A signal is the
first kind: a *statement*. `terminate` wants no reply — the exit notification is
its acknowledgement.
A health probe is the second kind: a *question*, worthless without its answer —
and the answer's absence within a deadline is the very thing being measured.
Asked as a signal it has no reply channel (a coalescing bit can't carry an answer,
and the authority rule forbids a child signalling its supervisor back); asked as a
call, the timeout-is-the-diagnosis semantics come free. So there is no `health`
signal. Liveness is the common **`ping`**: a reserved request every harness-run
service answers automatically on its main endpoint — still free for the service
author, still one obvious way — and a supervisor's probe is a `ping` call with a
deadline.
## The two iron rules
1. **Cleanup is the kernel's job.** A process can die with no warning — fault,
kill, power. Correctness must never depend on a `terminate` handler running. On
any death the kernel releases the address space, IPC handles, IRQ bindings, and
owed replies (built), and must also release **device, I/O-port, and interrupt
claims and MSI vectors** (the known gap in
[process-management.md](process-management.md); increment 1). A signal handler is
for *graceful* work — flushing, deregistering, saving — never for *necessary*
work.
2. **Kill is not a signal, and exit reasons are load-bearing.** The standard stop
sequence is *terminate → deadline → `process_kill`*; the unhandleable kill stays
a kernel mechanism. And a supervisor deciding whether to restart must know *how*
the child died: clean exit (meant to — don't restart), fault (restart with
backoff), killed (the supervisor did it). The exit notification today carries
only the id; it grows a reason. Restart policy cannot be written without it.
## Who learns of a death
A death has three audiences, and conflating them is how systems end up with either
zombie state or privileged snooping:
1. **The supervisor** — gets the exit notification on the endpoint it gave at spawn
(built), which grows the `ExitReason` (increment 2). The supervisor is the only
audience that needs the *reason*, because it is the only one deciding whether to
restart.
2. **The peer owed a reply** — already built: a client that dies mid-request fails
the server's reply with `-EPEER`; a server that dies fails its waiting clients
the same way. This covers the *synchronous* case only.
3. **The subscribers** — the new piece, and it is the input service's
publish/subscribe shape ([input.md](input.md)) applied to exits. A stateful
service accumulates per-client state across many requests: the VFS holds a dead
client's open file handles, the input service holds its subscriptions, a future
network stack holds its sockets. None of these are the client's supervisor, and
none learn anything from a failed reply if the client simply never calls again.
So the kernel **publishes every exit** to whoever subscribed:
`process_subscribe(endpoint)` adds a subscriber, and each death posts a
notification to every subscriber (badge = `notify_exit_bit | process id` — the
same encoding supervisors already decode, the IRQ-as-IPC pattern once more). The
subscriber filters for ids it holds state for and releases what the dead client
held. Correlating is free of bookkeeping: an IPC sender's badge already *is* its
task id (`runtime.ipc.Received`), so the id a service has been keying client
state by all along is the id the exit event carries.
Subscription, not broadcast-to-everyone: only processes that asked receive
events, the kernel keeps a bounded subscriber table, and delivery is the same
non-blocking coalescing notification as everything else — a dying process never
waits on its mourners. Subscribing is ungated, like `process_enumerate`: what is
running (and dying) is not a secret between cooperating processes. Subscribers
do not receive the exit reason — the VFS does not care *why* the client died.
This is the service-side mirror of iron rule 1: **a service must never depend on
its clients cleaning up after themselves.** Handle release on client death is the
service's job, triggered by the published exit event — never by a courtesy
"closing now" message that a crashed client will never send.
## The stable interface: `runtime.process`
`runtime.process` already owns what a process receives at birth (`Init`, the
argv contract). It grows to own the other end of life.
**The runtime is the stable interface; the numbers are not.** danos applications do
not make system calls — they call the runtime library, and the system-call numbers,
notification bits, and signal bit positions beneath it are a **private kernel ↔
runtime contract** that may change at any time (settled 2026-07-12). This is why
the runtime exists. Today kernel and runtime ship from one tree in one image, so
"stability" is simply building them together. When driver binaries start shipping
as separately-versioned applications — the whole point of the restart design — the
binary's embedded runtime version becomes compatibility metadata (the same idea as
the protocol version in the device manager's `hello`), and the kernel refuses what
it cannot serve. Signals therefore need no reserved numbering scheme: the enum
below is vocabulary, not ABI.
```zig
/// The signal vocabulary. The value is the bit position in the pending mask — a
/// private kernel/runtime detail, free to change while they ship together.
pub const Signal = enum(u5) {
terminate = 0, // SIGTERM: finish up and exit
reload = 1, // SIGHUP: re-read configuration
interrupt = 2, // SIGINT
quit = 3, // SIGQUIT
alarm = 4, // SIGALRM
user_1 = 5, // SIGUSR1
user_2 = 6, // SIGUSR2
};
/// A decoded pending mask: the coalesced set of signals a notification delivered.
pub const SignalSet = struct {
pending: u32,
pub fn has(set: SignalSet, signal: Signal) bool { ... }
pub fn iterate(set: SignalSet) Iterator { ... }
};
/// Nominate `endpoint` as this process's signal endpoint (signal_bind). The
/// runtime's service harness calls this; a bare program may call it directly and
/// fold signals into its own replyWait loop.
pub fn bindSignals(endpoint: usize) bool { ... }
/// Decode a received badge into signals, or null if the badge is not a signal
/// notification (mirrors ipc.Received.isChildExit).
pub fn signalsFrom(badge: usize) ?SignalSet { ... }
/// Send `signal` to process `id`. Supervisor-gated, like kill; non-blocking.
pub fn sendSignal(id: u32, signal: Signal) bool { ... }
/// The standard stop sequence: terminate, wait up to `deadline_ms` for the exit
/// notification, then process_kill. The one call a supervisor needs.
pub fn stop(id: u32, deadline_ms: u64) void { ... }
/// Subscribe `endpoint` to published exit events (process_subscribe). Every
/// process death posts an asynchronous notification: badge = notify_exit_bit |
/// process id — the same encoding a supervisor's exit notification uses, decoded
/// by the same ipc.Received helpers. For stateful services: release what the dead
/// client held (file handles, subscriptions, sockets). Ungated, like
/// process_enumerate.
pub fn subscribeExits(endpoint: usize) bool { ... }
/// How a process ended — queried after the exit notification (the kernel records
/// it first, so the two never race). What restart policy reads. (Built in M17.2.)
pub const ExitReason = enum(u8) {
exited, // returned from main / clean exit
aborted, // abort() — deliberate self-termination (SIGABRT's ghost; reserved)
segmentation_fault, // SIGSEGV's ghost
illegal_instruction, // SIGILL's ghost
arithmetic_fault, // SIGFPE's ghost
protection_fault, // general protection fault
fault, // any other CPU exception
killed, // process_kill
};
```
Two deliberate absences. There is no `mask`/`block` API — a process that is not
ready for a signal simply has not waited on its endpoint yet; the pending mask *is*
the blocked set. And there is no per-signal handler registration at this layer —
dispatch is the process's own `switch` over `SignalSet`, or the service harness's
callbacks (`on_terminate`, `on_reload`) for programs that want defaults.
### The service harness
`runtime.service` owns the `replyWait` loop and folds every event source — signals,
child exits, protocol messages — into callbacks, with the vocabulary's defaults:
`terminate` returns from the loop (clean exit), the common `ping` is answered automatically,
`reload` is ignored unless overridden. One loop, no locking, nothing reentrant. A
service author writes domain logic; the lifecycle contract is satisfied by the
harness. A process that bypasses the harness and ignores its signals meets the
deadline-then-kill escalation — you cannot force a process to implement an
interface, but you can make compliance free and non-compliance fatal.
### The musl layer later
The POSIX C layer is a **musl port**: musl's arch/syscall layer retargeted so that
what musl believes are kernel syscalls become danos runtime calls and IPC — `open`
and `read` onto the VFS protocol, `kill`/`sigaction`/`waitpid` onto this document's
vocabulary, `exit` onto the runtime's exit path. `sigaction` handlers registered
through it are invoked by the runtime's loop when the signal message arrives —
synchronous underneath, async-looking to ported code, delivered at wait boundaries
the way most Unix programs already experience signals (at syscalls). No stack hijack
ever happens, `SA_RESTART` semantics come free because nothing was interrupted, and
SIGPIPE can be synthesized from `-EPEER` for the programs that expect it. C programs
get POSIX; danos-native programs never pay for it.
## Increments
1. **Kernel: release device/port/IRQ claims and MSI vectors on death** — the
cleanup half of iron rule 1, and the prerequisite for any restart story. Test:
kill a claiming driver, spawn it again, the claim succeeds.
2. **Exit reason in the death notification** (`ExitReason` above).
3. **Exit events**: `process_subscribe` in the kernel (bounded subscriber table,
publishes on every death), `runtime.process.subscribeExits`; the VFS becomes the
first subscriber — releasing a dead client's handles is its proof test.
4. **Signals**: `signal_bind` + `process_signal` + the pending mask in the kernel;
`runtime.process` grows the interface above; the service harness handles
`terminate` and answers the common `ping`; `stop()` for supervisors.
[device-manager.md](device-manager.md) builds directly on all four.
## Settled questions (2026-07-12)
- **Signal numbering is not ABI**: the runtime is the stable interface; the numbers
beneath it are a private kernel ↔ runtime contract (see "The stable interface").
- **Liveness is a `ping` call, not a signal**: signals are statements, questions
are synchronous calls (see "Statements, not questions"). A service wanting *deep*
health ("can I reach my hardware?") defines its own protocol message on top.
- **Process handles: deferred.** Pids + the supervisor gate cover everything
planned; transferable handles (Fuchsia-style, delegating signalling without
delegating kill) wait for the capability table to grow types beyond endpoints.
- **`alarm`: in the vocabulary, unbuilt.** No consumer yet; when one appears it is
runtime sugar over the existing timer (arm a timer that posts your own signal) —
zero kernel work, so deferring costs nothing.
- **Subscription granularity: all exits**, subscriber-side filtering — one
subscription per service, a bounded kernel table. Per-id subscriptions only if
event volume ever matters (hundreds of processes, not before).
- **Client identity across the exit boundary: no convention needed** — an IPC
sender's badge already is its task id (see "Who learns of a death").
+120
View File
@@ -0,0 +1,120 @@
# Process Management
How danos lists, supervises, and kills processes — the microkernel answer to
`ps`, `kill`, and `SIGCHLD`/`wait`.
## Why system calls, not `/proc`
Unix systems sit on a spectrum. Classic BSD/macOS list processes through
syscalls (`sysctl(KERN_PROC)`) and kill through `kill(2)`; Linux renders the
process table as `/proc` for *reading* but still kills through a syscall; Plan 9
made the file tree the whole interface (`echo kill > /proc/n/ctl`). Microkernels
mostly abandon ambient PIDs: Minix and QNX route everything through a user-space
process-manager server, and Fuchsia/seL4 control processes only through handles.
danos rules out `/proc` **as the primitive**: here a `/proc` would be served by
the VFS server — a user process — which would put the VFS in the path of process
control. If the VFS (or anything under it) hangs, nothing could be listed or
killed, *including the hung VFS*. The control plane for processes must not
depend on a process. So the primitives are kernel system calls; a read-only
`/proc` rendering can be layered on later, and a POSIX-style process-manager
server can be built *from* these primitives when one is needed.
## The three primitives
### `process_enumerate(buffer, maximum) -> total`
A snapshot of the task table into a caller buffer of `abi.ProcessDescriptor`
(id, supervisor, state, priority, name) — the exact shape of
`device_enumerate`, so `ps` is a user program over a snapshot, not a kernel
service. The total may exceed what fit; call again with a larger buffer. Kernel
tasks are included with an empty name — an honest listing shows the idle tasks
too. Ungated and read-only: what is running is not a secret between cooperating
bring-up processes.
### `system_spawn(..., exit_endpoint) -> child id`, and the supervision link
`system_spawn` records the caller as the child's **supervisor** and returns the
child's process id (ids are monotonic, never reused — a stale id can only miss).
That link is the kill authority: it answers "who may kill process 7?" without
inventing users or permissions, the same way a device *claim* is the capability
for `mmio_map`. It composes with the supervision hierarchy the device manager
already forms: init supervises the services it starts, the device manager
supervises the drivers it matches. (A transferable process *handle* — Fuchsia
style — can replace the id once the handle table grows types beyond endpoints.)
`exit_endpoint` (a handle, or `abi.no_cap`) is the supervisor's death-watch: when
the child ends — clean exit, CPU fault, or `process_kill` — the kernel posts an
asynchronous notification to that endpoint, exactly like a bound IRQ. The badge
carries `abi.notify_badge_bit | abi.notify_exit_bit | child_id`, so one endpoint
supervises many children and can even share with IRQ notifications. This is the
microkernel's SIGCHLD: no new mechanism, just the IRQ-as-IPC pattern reused, and
a supervisor's event loop (`ipc.replyWait`) already knows how to receive it. The
child holds a reference to the endpoint from birth, so the notification cannot
dangle even if the supervisor dies first.
### `process_kill(id) -> 0 / -ESRCH / -EPERM`
Only the supervisor may kill; kernel tasks are not killable processes. Like a
signal, delivery is prompt but asynchronous — 0 means the kill is accepted and
irrevocable; the exit notification confirms completion.
## How a kill lands (the kernel mechanics)
Everything below runs under the big kernel lock, where task states cannot move.
- **Target ready or blocked** (not on any core): reaped on the killer's own
call. The reap releases what death always releases (IRQ bindings first, then
a client the target still owed a reply to is failed with `-EPEER`, IPC handles
closed, the exit notification posted last) — plus the unlinking only a
*remote* death needs: out of the ready queue, out of an endpoint's sender FIFO
(`Task.ipc_wait_endpoint`), out of a receive wait queue (`Task.wait_queue`),
and out of any server's owed-reply slot, so nothing ever dequeues a dangling
pointer. Destroying the address space is safe because no core can have it
loaded: every switch away from a task loads the next task's tables.
- **Target running on another core**: it cannot be torn down mid-instruction,
so it is condemned (`Task.kill_pending`) and dies at whichever comes first:
- its next **system_call entry** — checked before dispatch, so a condemned
process cannot spawn, claim, or message anything on its way out;
- its core's next **timer tick** — but only when the task is not inside one
of its own system calls (`Task.in_system_call`): the tick may have
interrupted kernel code mid-operation, where teardown would leak whatever
the operation held. User-mode execution is always a safe kill point. The
tick-time terminate abandons the interrupt frame exactly like the fault
path (the LAPIC is acknowledged before the tick hook runs);
- any core's tick finding it **blocked or ready** (it entered a syscall and
parked after being condemned) — reaped by the same remote-reap path.
A pure user-mode spin loop that never makes a system call therefore dies
within one tick; nothing a process does can outrun the kill.
The scheduler stays below the process layer: finishing a kill (IRQ bindings,
handles, the notification) is called *up* through two hooks process.zig
registers at boot (`terminate_current_hook`, `reap_task_hook`), mirroring how
the architecture layer calls up into `tick`.
## Known gaps (bring-up honesty)
- ~~Device claims are not released on death~~ Closed (M17.1): every path out of a
process releases its device claims alongside its IRQ and MSI bindings
(`releaseTaskResourcesLocked`), so a restarted driver can claim its hardware
again — the cleanup half of [process-lifecycle.md](process-lifecycle.md)'s iron
rule 1. The `claim-release` test proves the kill → release → re-claim cycle.
- Kernel stacks of dead tasks are leaked, as on every exit path (no reaper yet).
- ~~There is no exit status in the notification~~ Closed (M17.2): the kernel
records how every process ends — exited, a fault class, or killed — before it
posts the exit notification, and the supervisor reads it with
`process_exit_reason` (`runtime.process.exitReason`). This is the input to
restart policy ([process-lifecycle.md](process-lifecycle.md)); an exit *code*
for the clean case can still ride alongside later.
- Enumerate writes through the caller's raw pointer under the bring-up trust
model, like `device_enumerate` (an unmapped page is a self-DoS, not an
isolation break).
## Tests
`process-list` (enumerate), `process-kill` (kernel-level kill paths, refusals,
notifications), `supervision` (the whole user-side surface via the process-test
service: spawn supervised → enumerate → kill blocked and spinning children →
notifications → gone), `claim-release` (a killed claim-holder's device is
claimable again). See test/qemu_test.py.
+16 -3
View File
@@ -1,6 +1,17 @@
# Resilience: fault isolation and live restart
A design/research note, not built yet. This is the property danos is really chasing:
Steps 1–4 of the ordering below are **built** (M17–M18, 2026-07-13): user-mode
isolation; fault → kill the process → keep the core (`onException`; the
`fault-recovery` test); the supervisor notification **with exit reasons**
([process-lifecycle.md](process-lifecycle.md) — clean exit, fault class, or
killed, recorded before the notice posts); and the **restart policy itself**
([device-manager.md](device-manager.md)): the device manager supervises every
driver, restarts crashes with backoff, caps crash loops, and re-claims work
because the kernel releases a dead process's claims. The `driver-restart` and
`usb-report` scenarios prove kill → release → respawn → re-claim → re-report
end to end. What remains of this document's ladder is scope, not mechanism:
more of the system moved into restartable processes (the discovery migration,
[m19-m20-plan.md](m19-m20-plan.md), is the next rung). This is the property danos is really chasing:
**if a part of the OS breaks, isolate it, and re-initialise it — without rebooting.**
A crashed driver gets restarted; a wedged service gets killed and brought back. It's
the reason the [microkernel](vision.md) shape was chosen, and it's a *separate* goal
@@ -111,9 +122,11 @@ Honest boundaries:
## Suggested ordering
1. **User mode + address-space isolation** — the shared prerequisite (also on the
path for everything else).
path for everything else). **Done.**
2. **Kernel: fault → kill process → notify.** Turn today's "halt on fault" into
"confine to the process and report it."
"confine to the process and report it." **Done** (the kill and reclaim; the
supervisor notification waits for step 3's supervisor). A killed server's
pending client is unblocked with `-EPEER` rather than hung.
3. **A minimal supervisor server** that can (re)start a process.
4. **Resource cleanup on death** — reclaim memory/MMIO/IPC/IRQ, via caps or a grant
table.
+2
View File
@@ -47,6 +47,8 @@ Everything else---including`read()`,`write()`,`malloc()`, and`fork()`---will run
- **What it does:**Used strictly by your background user-space servers (like your disk driver or filesystem). It sends a reply to the last client that called it, and immediately puts the server to sleep until the next request arrives.[[1](https://news.ycombinator.com/item?id=33078441)]
3. **`Yield()`/`Thread_Ctrl()`**
- **What it does:**Allows a thread to voluntarily give up its CPU time slice, or allows a root task to spawn/kill threads.
4. **`ipc_send(endpoint, message_buffer)`(Asynchronous Send)**
- **What it does:**Posts a small payload to an endpoint's bounded queue and returns *without* blocking — no rendezvous, no reply. The receiver picks it up through the same `IPC_ReplyWait`, as a buffered message. It is the async counterpart of `IPC_Call`, for one-to-many broadcasts where a synchronous rendezvous would let one dead or slow receiver hang the sender. The [input service](input.md) — keyboard-event fan-out — is its first user. A full queue drops the oldest message (a buffered message is discrete data, unlike a coalescing interrupt notification).
* * * * *
+32
View File
@@ -65,6 +65,38 @@ function-pointer type and the kernel's `_start` both carry
whole reason `kernel_abi` lives in the shared contract — see [efi.md](efi.md) for
the handoff it governs.
## The process-entry stack (argc/argv)
The SysV ABI also fixes what a *fresh process* finds on its stack — and danos
follows it, so its own runtime and any future C libc read arguments the same way.
At the first user instruction, `rsp` is 16-byte aligned and points at (addresses
growing upward):
```
rsp → argc u64
argv[0] … argv[argc-1] pointers into the strings area below
NULL argv terminator
NULL envp terminator (no environment yet)
{AT_PAGESZ, page size} auxiliary vector
{AT_NULL, 0} auxiliary-vector terminator
argv string bytes NUL-terminated
───────────────────────── stack top (stack_top_virtual)
```
The kernel builds this block at the top of the process's stack — 8 pages (32 KiB,
`parameters.user_stack_pages`) mapped RW+NX below a fixed top, with the page below
them left unmapped as a **guard**, so a stack overflow faults (killing only that
process) instead of silently corrupting the image
(`buildEntryStack` in `system/kernel/process.zig`); `argv[0]` is always the path
or initial-ramdisk name the process was spawned as, and `system_spawn`'s optional
argument blob becomes `argv[1..]`. The runtime's `_start`
(`library/runtime/start.zig`) hands the block to `rt_start`, which builds a
`runtime.process.Init` from it and passes that to the program's `main`
(`pub fn main(init: runtime.process.Init)`; a parameterless `main()` is also
accepted). A C runtime's `crt0` would walk
the identical layout unmodified — that's the compatibility being bought. The
`args` test proves the round trip.
## Where else it surfaces
- **The red zone → `red_zone = false`.** `build.zig` disables the red zone for the
+21
View File
@@ -3,6 +3,7 @@
//! ownership of its hardware; the claim is the capability the kernel checks before
//! mapping registers or routing an IRQ.
const std = @import("std");
const abi = @import("abi");
const device_abi = @import("device-abi");
const sc = @import("system-call.zig");
@@ -36,6 +37,10 @@ pub fn mmioMap(device_id: u64, resource_index: u64) ?usize {
/// `DeviceDescriptor.parent` for a device with no parent.
pub const no_parent = device_abi.no_parent;
/// `DeviceDescriptor.pci_class` for a device that is not a PCI function. Set this on
/// descriptors passed to `register` unless the child really is one.
pub const no_pci_class = device_abi.no_pci_class;
/// Publish `descriptor` as a child of `parent_id`, which this process must have claimed.
/// Returns the new device id. The child is left unclaimed, so whichever driver owns
/// that class of device can `claim` it — that is how a bus hands off a device.
@@ -108,3 +113,19 @@ pub fn ioRead(device_id: u64, resource_index: u64, offset: u64, width: u8) ?u32
pub fn ioWrite(device_id: u64, resource_index: u64, offset: u64, width: u8, value: u32) bool {
return !failed(sc.systemCall5(.io_write, device_id, resource_index, offset, width, value));
}
/// Find DeviceDescription by hid
///
/// Utility function for driver development
pub fn findDeviceDescriptorByHid(buffer: []DeviceDescriptor, hid_needle: []const u8) ?DeviceDescriptor {
const total = enumerate(buffer);
const n = @min(total, buffer.len);
for (@as([]DeviceDescriptor, buffer[0..n])) |d| {
const hid_haystack = d.hid[0..@intCast(d.hid_len)];
if (std.mem.eql(u8, hid_haystack, hid_needle)) {
return d;
}
}
return null;
}
+220
View File
@@ -0,0 +1,220 @@
//! User-space input helpers: the client and publisher sides of the input service, so a
//! program listening for input events — or a driver broadcasting them — doesn't hand-roll
//! the IPC. Layered over `ipc` (endpoints, capability passing, `send`) and the shared
//! `input-protocol` wire format, the way `device.zig` layers over the raw `device_*` calls.
//! See system/services/input/input.zig.
//!
//! The service carries several device classes (keyboard, mouse, joystick/gamepad). A
//! **source** publishes its class with the matching method:
//! var source = input.connectSource() orelse return;
//! _ = source.publishKeyboardEvent(.{ .kind = ..., .keycode = ..., ... });
//! _ = source.publishMouseEvent(.{ ... });
//! _ = source.publishJoystickEvent(.{ ... });
//!
//! A **subscriber** either takes one class with a typed helper —
//! var keys = input.subscribeKeyboard() orelse return;
//! while (true) { const key = keys.next() orelse continue; ... }
//! — or takes several at once and inspects the tagged envelope:
//! var listener = input.subscribeAll() orelse return;
//! while (true) {
//! const event = listener.next() orelse continue;
//! if (event.asKeyboard()) |k| { ... } else if (event.asMouse()) |m| { ... }
//! }
const std = @import("std");
const abi = @import("abi");
const ipc = @import("ipc.zig");
const system = @import("system.zig");
const protocol = @import("input-protocol");
pub const DeviceKind = protocol.DeviceKind;
pub const InputEvent = protocol.InputEvent;
pub const KeyEvent = protocol.KeyEvent;
pub const MouseEvent = protocol.MouseEvent;
pub const JoystickEvent = protocol.JoystickEvent;
pub const EventKind = protocol.EventKind;
pub const MouseEventKind = protocol.MouseEventKind;
pub const JoystickEventKind = protocol.JoystickEventKind;
pub const Keycode = protocol.Keycode;
/// Interest masks re-exported so a caller can `subscribe(input.device_keyboard |
/// input.device_mouse)`.
pub const device_keyboard = protocol.device_keyboard;
pub const device_mouse = protocol.device_mouse;
pub const device_joystick = protocol.device_joystick;
pub const device_all = protocol.device_all;
/// Look up the input service, retrying while it is still coming up. Both a subscriber and
/// a source race the service's registration at boot, so both wait for it here rather than
/// failing. Returns the service endpoint handle, or null if it never appears.
fn lookupService() ?ipc.Handle {
var attempts: usize = 0;
while (attempts < 100) : (attempts += 1) {
if (ipc.lookup(.input)) |handle| return handle;
system.sleep(50);
}
return null;
}
// --- subscribing ------------------------------------------------------------
/// A subscription to the input service: our own endpoint, which the service pushes events
/// to. `next` returns each event as a tagged `InputEvent`; use `asKeyboard`/`asMouse`/
/// `asJoystick` to decode. Created with `subscribe`/`subscribeAll`; for a single device
/// class prefer the typed helpers (`subscribeKeyboard`, ...), which return decoded events.
pub const Subscriber = struct {
/// The endpoint the service delivers events to (created and owned by us; its handle
/// was handed to the service as a capability at subscribe time).
endpoint: ipc.Handle,
receive: [protocol.event_size]u8 = undefined,
/// Block until the next event is pushed, and return it. Events arrive as asynchronous
/// buffered messages (`ipc_send` from the service), so nothing is owed in reply — the
/// empty reply this issues is a harmless no-op. Returns null for any non-event wake-up
/// (there should be none), so callers can loop.
pub fn next(self: *Subscriber) ?InputEvent {
const got = ipc.replyWait(self.endpoint, &.{}, &self.receive, null);
if (!got.isMessage() or got.len < protocol.event_size) return null;
return std.mem.bytesToValue(InputEvent, self.receive[0..protocol.event_size]);
}
};
/// Subscribe to the input classes named in `device_mask` (an OR of `device_*`, or
/// `device_all`). Creates an endpoint for the service to push to and hands it over as a
/// capability. Returns a `Subscriber` to loop `next` on, or null on failure.
pub fn subscribe(device_mask: u32) ?Subscriber {
const service = lookupService() orelse return null;
const endpoint = ipc.createIpcEndpoint() orelse return null;
var request = protocol.Request{ .operation = @intFromEnum(protocol.Operation.subscribe), .device_mask = device_mask };
var reply: [protocol.reply_size]u8 = undefined;
const result = ipc.callCap(service, std.mem.asBytes(&request), &reply, endpoint) catch return null;
if (result.len < protocol.reply_size) return null;
if (std.mem.bytesToValue(protocol.Reply, reply[0..protocol.reply_size]).status != 0) return null;
return .{ .endpoint = endpoint };
}
/// Subscribe to every input class (keyboard, mouse, joystick) on one stream.
pub fn subscribeAll() ?Subscriber {
return subscribe(device_all);
}
/// A subscriber filtered to keyboard events, whose `next` returns a decoded `KeyEvent`.
pub const KeyboardSubscriber = struct {
inner: Subscriber,
pub fn next(self: *KeyboardSubscriber) ?KeyEvent {
return (self.inner.next() orelse return null).asKeyboard();
}
};
/// A subscriber filtered to mouse events, whose `next` returns a decoded `MouseEvent`.
pub const MouseSubscriber = struct {
inner: Subscriber,
pub fn next(self: *MouseSubscriber) ?MouseEvent {
return (self.inner.next() orelse return null).asMouse();
}
};
/// A subscriber filtered to joystick/gamepad events, whose `next` returns a decoded
/// `JoystickEvent`.
pub const JoystickSubscriber = struct {
inner: Subscriber,
pub fn next(self: *JoystickSubscriber) ?JoystickEvent {
return (self.inner.next() orelse return null).asJoystick();
}
};
/// Subscribe to keyboard events only; `next` returns decoded `KeyEvent`s.
pub fn subscribeKeyboard() ?KeyboardSubscriber {
return .{ .inner = subscribe(device_keyboard) orelse return null };
}
/// Subscribe to mouse events only; `next` returns decoded `MouseEvent`s.
pub fn subscribeMouse() ?MouseSubscriber {
return .{ .inner = subscribe(device_mouse) orelse return null };
}
/// Subscribe to joystick/gamepad events only; `next` returns decoded `JoystickEvent`s.
pub fn subscribeJoystick() ?JoystickSubscriber {
return .{ .inner = subscribe(device_joystick) orelse return null };
}
// --- publishing -------------------------------------------------------------
/// A connection to the input service for a source (a keyboard/mouse/joystick driver) that
/// publishes events. Each `publish*Event` is a short synchronous call the service answers
/// at once; its own fan-out to subscribers is asynchronous, so publishing never blocks on
/// a slow subscriber.
pub const Publisher = struct {
service: ipc.Handle,
fn publish(self: Publisher, event: InputEvent) bool {
var request = protocol.Request{ .operation = @intFromEnum(protocol.Operation.publish), .event = event };
var reply: [protocol.reply_size]u8 = undefined;
const len = ipc.call(self.service, std.mem.asBytes(&request), &reply) catch return false;
if (len < protocol.reply_size) return false;
return std.mem.bytesToValue(protocol.Reply, reply[0..protocol.reply_size]).status == 0;
}
/// Broadcast a keyboard event to every subscriber that took keyboard events.
pub fn publishKeyboardEvent(self: Publisher, event: KeyEvent) bool {
return self.publish(InputEvent.fromKeyboard(event));
}
/// Broadcast a mouse event to every subscriber that took mouse events.
pub fn publishMouseEvent(self: Publisher, event: MouseEvent) bool {
return self.publish(InputEvent.fromMouse(event));
}
/// Broadcast a joystick/gamepad event to every subscriber that took joystick events.
pub fn publishJoystickEvent(self: Publisher, event: JoystickEvent) bool {
return self.publish(InputEvent.fromJoystick(event));
}
};
/// Connect to the input service as an event source, waiting for it to come up. Returns a
/// `Publisher`, or null if the service never registered.
pub fn connectSource() ?Publisher {
return .{ .service = lookupService() orelse return null };
}
// --- synthetic scaffolding --------------------------------------------------
/// Synthetic key events, shared by the demo source and the keyboard driver's placeholder
/// stream while real scancode decoding is still a follow-up. `step` rolls through A..E,
/// emitting for each key a `key_down`, then a `key_press` carrying the character, then a
/// `key_up`. Scaffolding, not wire protocol — hence it lives with the helpers.
pub fn syntheticKeyEvent(step: usize) KeyEvent {
const Key = struct { code: Keycode, character: u32 };
const keys = [_]Key{
.{ .code = .a, .character = 'A' },
.{ .code = .b, .character = 'B' },
.{ .code = .c, .character = 'C' },
.{ .code = .d, .character = 'D' },
.{ .code = .e, .character = 'E' },
};
const key = keys[(step / 3) % keys.len];
return switch (step % 3) {
0 => .{ .kind = @intFromEnum(EventKind.key_down), .keycode = @intFromEnum(key.code), .character = 0, .modifiers = 0 },
1 => .{ .kind = @intFromEnum(EventKind.key_press), .keycode = @intFromEnum(key.code), .character = key.character, .modifiers = 0 },
else => .{ .kind = @intFromEnum(EventKind.key_up), .keycode = @intFromEnum(key.code), .character = 0, .modifiers = 0 },
};
}
/// Synthetic mouse events (placeholder until real PS/2 packet decoding). `step` alternates
/// a small diagonal motion with a left-button click.
pub fn syntheticMouseEvent(step: usize) MouseEvent {
return switch (step % 3) {
0 => .{ .kind = @intFromEnum(MouseEventKind.motion), .button = 0, .dx = 1, .dy = 1, .scroll_x = 0, .scroll_y = 0, .buttons = 0 },
1 => .{ .kind = @intFromEnum(MouseEventKind.button_down), .button = protocol.mouse_button_left, .dx = 0, .dy = 0, .scroll_x = 0, .scroll_y = 0, .buttons = protocol.mouse_button_left },
else => .{ .kind = @intFromEnum(MouseEventKind.button_up), .button = protocol.mouse_button_left, .dx = 0, .dy = 0, .scroll_x = 0, .scroll_y = 0, .buttons = 0 },
};
}
/// Synthetic joystick/gamepad events (placeholder until a real controller driver). `step`
/// sweeps axis 0 and toggles button 0.
pub fn syntheticJoystickEvent(step: usize) JoystickEvent {
return switch (step % 3) {
0 => .{ .kind = @intFromEnum(JoystickEventKind.axis), .control = 0, .value = 16384, .buttons = 0 },
1 => .{ .kind = @intFromEnum(JoystickEventKind.button_down), .control = 0, .value = 0, .buttons = 1 },
else => .{ .kind = @intFromEnum(JoystickEventKind.button_up), .control = 0, .value = 0, .buttons = 0 },
};
}
+70 -3
View File
@@ -79,11 +79,41 @@ pub fn call(h: Handle, message: []const u8, reply: []u8) CallError!usize {
return (try callCap(h, message, reply, null)).len;
}
/// Post `message` to endpoint `h`'s asynchronous queue and return immediately — no
/// rendezvous, no reply, no blocking. The receiver picks it up through `replyWait` as a
/// buffered message (`Received.isMessage`). Unlike `call`, this **cannot hang on a dead
/// or slow peer**, which is why a broadcaster (the input service) delivers events this
/// way. The payload must fit an endpoint slot (64 bytes); a full queue drops the oldest
/// message. Returns false on failure (bad handle, oversized payload, bad buffer).
pub fn send(h: Handle, message: []const u8) bool {
return !failed(sc.systemCall3(.ipc_send, h, @intFromPtr(message.ptr), message.len));
}
/// Set in `Received.badge` when what arrived is an asynchronous notification — a
/// bound device interrupt — rather than a client's message. The low bits carry the
/// GSI. See `isNotification`.
pub const notify_badge_bit: u64 = abi.notify_badge_bit;
/// Set alongside `notify_badge_bit` when the notification is a **signal** — the
/// lifecycle vocabulary of docs/process-lifecycle.md, delivered to the endpoint
/// nominated with `process.bindSignals`. Decode with `process.signalsFrom`.
pub const notify_signal_bit: u64 = abi.notify_signal_bit;
/// Set alongside `notify_badge_bit` when the notification is a **one-shot timer**
/// landing (`system.timerOnce`).
pub const notify_timer_bit: u64 = abi.notify_timer_bit;
/// Set alongside `notify_badge_bit` when the notification is a **child-exit
/// notice** — a process this one spawned (with an exit endpoint) has ended —
/// rather than a device interrupt. The low bits carry the child's process id.
pub const notify_exit_bit: u64 = abi.notify_exit_bit;
/// Set alongside `notify_badge_bit` when the wake-up is a **buffered message** — a payload
/// posted with `send` (`ipc_send`) — rather than a bare device interrupt or child-exit
/// notice. The payload is in the `replyWait` receive buffer (`Received.len` bytes); the
/// low bits of the badge carry the sender's task id. See `Received.isMessage`.
pub const notify_message_bit: u64 = abi.notify_message_bit;
/// The result of a `replyWait`: the request length, the sender's badge (a task id, or
/// an IRQ notification if the high bit is set), and any capability the request carried.
pub const Received = struct {
@@ -91,16 +121,53 @@ pub const Received = struct {
badge: u64,
cap: ?Handle,
/// True if this wake-up was a device interrupt, not a client request. A driver's
/// event loop branches on this; there is no reply owed on the notification path.
/// True if this wake-up was an asynchronous notification (a device interrupt
/// or a child-exit notice), not a client request. An event loop branches on
/// this; there is no reply owed on the notification path.
pub fn isNotification(self: Received) bool {
return self.badge & notify_badge_bit != 0;
}
/// The interrupt source (a GSI), meaningful only when `isNotification`.
/// True if this wake-up tells of a supervised child's end — the notification
/// requested by passing an exit endpoint to `system.spawnSupervised`.
pub fn isChildExit(self: Received) bool {
return self.isNotification() and self.badge & notify_exit_bit != 0;
}
/// True if this wake-up is a **buffered message** posted with `send` (`ipc_send`):
/// there is a payload in the receive buffer (`self.len` bytes) and no reply is owed.
/// The subscriber side of a broadcast branches on this.
pub fn isMessage(self: Received) bool {
return self.isNotification() and self.badge & notify_message_bit != 0;
}
/// The task id of whoever posted a buffered message, meaningful only when
/// Whether this arrival is a signal notification — decode the set with
/// `process.signalsFrom(badge)`.
pub fn isSignal(self: Received) bool {
return self.isNotification() and self.badge & notify_signal_bit != 0;
}
/// Whether this arrival is a one-shot timer landing (`system.timerOnce`).
pub fn isTimer(self: Received) bool {
return self.isNotification() and self.badge & notify_timer_bit != 0;
}
/// `isMessage`. (The badge's low bits, with the three high marker bits masked off.)
pub fn senderTaskId(self: Received) u32 {
return @intCast(self.badge & ~(notify_badge_bit | notify_exit_bit | notify_message_bit));
}
/// The interrupt source (a GSI), meaningful only when `isNotification` and
/// not `isChildExit`.
pub fn source(self: Received) u64 {
return self.badge & ~notify_badge_bit;
}
/// The ended child's process id, meaningful only when `isChildExit`.
pub fn childProcessId(self: Received) u32 {
return @intCast(self.badge & ~(notify_badge_bit | notify_exit_bit));
}
};
/// Server side of IPC_ReplyWait: deliver `reply` to the client last received (if any,
+137
View File
@@ -0,0 +1,137 @@
//! Process-level runtime types: what a user program receives at entry (`Init`,
//! the argv contract) and the process end of the lifecycle
//! (docs/process-lifecycle.md) — today the exit reason a supervisor reads to
//! decide restart; signals and the stop sequence land here with M17.4. Mirrors
//! the spirit of `std.process.Init.Minimal` in danos terms — std's `Args` holds
//! no data on freestanding targets, so the type is danos's own.
const std = @import("std");
const abi = @import("abi");
const sc = @import("system-call.zig");
const ipc = @import("ipc.zig");
const system = @import("system.zig");
/// Everything a program receives at entry. Passed to
/// `pub fn main(init: runtime.process.Init)`; programs that need nothing keep
/// `pub fn main() void`. An `environment` field is added here once the kernel
/// passes a non-empty envp (today it is always empty — see docs/sysv.md).
pub const Init = struct {
arguments: Arguments,
};
/// The process arguments (argc/argv), parsed from the kernel-built System V
/// entry block. The bytes live in the entry block at the top of the stack page,
/// NUL-terminated, valid for the process's lifetime.
pub const Arguments = struct {
/// argc — at least 1: argument 0 is the path or name this binary was
/// spawned as.
count: usize,
/// The argv pointers in the entry block (NULL-terminated after `count`
/// entries).
vector: [*]const [*:0]const u8,
/// Argument `index` (0 = the program's own path/name), or null if out of
/// range.
pub fn get(arguments: Arguments, index: usize) ?[:0]const u8 {
if (index >= arguments.count) return null;
return std.mem.span(arguments.vector[index]);
}
pub fn iterate(arguments: Arguments) Iterator {
return .{ .arguments = arguments };
}
pub const Iterator = struct {
arguments: Arguments,
index: usize = 0,
pub fn next(iterator: *Iterator) ?[:0]const u8 {
const argument = iterator.arguments.get(iterator.index) orelse return null;
iterator.index += 1;
return argument;
}
};
};
/// How a process ended — what a supervisor's restart policy reads: a clean exit
/// meant to stop, a fault wants a restart with backoff, killed means the
/// supervisor did it itself (docs/process-lifecycle.md).
pub const ExitReason = abi.ExitReason;
/// How dead child `id` ended. Ask after the exit notification arrives — the
/// kernel records the reason before it posts the notification, so this never
/// races it. Returns null for an id that never lived, is still alive, was
/// evicted from the kernel's bounded record, or is not this process's child
/// (the same authority gate as `kill`).
pub fn exitReason(id: u32) ?ExitReason {
const r = sc.systemCall1(.process_exit_reason, id);
if (r > ~@as(usize, 0) - 4095) return null; // a wrapped -errno
return @enumFromInt(r);
}
/// The signal vocabulary (docs/process-lifecycle.md): POSIX's concepts, danos's
/// names, message delivery. A signal is a one-way coalescing statement — never a
/// question (liveness is the zero-length ping call) and never kill (that is
/// `system.kill`, unhandleable by definition).
pub const Signal = abi.Signal;
/// The coalesced set of signals one notification delivered: two pending
/// terminates arrive as one. Decode a received badge with `signalsFrom`.
pub const SignalSet = struct {
pending: u32,
pub fn has(set: SignalSet, signal: Signal) bool {
return set.pending & (@as(u32, 1) << @intFromEnum(signal)) != 0;
}
};
/// Nominate `endpoint` as this process's signal endpoint. Signals posted while
/// unbound have pended; they are delivered immediately on bind, coalesced.
pub fn bindSignals(endpoint: usize) bool {
return sc.systemCall1(.signal_bind, endpoint) == 0;
}
/// Decode a received badge into the signals it delivered, or null if it is not
/// a signal notification.
pub fn signalsFrom(badge: u64) ?SignalSet {
if (badge & abi.notify_badge_bit == 0 or badge & abi.notify_signal_bit == 0) return null;
return .{ .pending = @truncate(badge & ~(abi.notify_badge_bit | abi.notify_signal_bit)) };
}
/// Post `signal` to child `id` (or to yourself). Supervisor-gated, like kill;
/// non-blocking, always — a statement, not a conversation.
pub fn sendSignal(id: u32, signal: Signal) bool {
return sc.systemCall2(.process_signal, id, @intFromEnum(signal)) == 0;
}
/// The standard stop sequence (docs/process-lifecycle.md): terminate, wait up to
/// `deadline_ms` for the exit notification on `exit_endpoint` (the endpoint the
/// child was spawned with), then kill. Any *other* notifications arriving on
/// that endpoint while stopping are consumed and dropped — a supervisor with
/// concurrent traffic implements the same sequence inside its own event loop
/// (arm `system.timerOnce`, keep serving) instead of calling this.
pub fn stop(id: u32, deadline_ms: u64, exit_endpoint: usize) void {
_ = sendSignal(id, .terminate);
_ = system.timerOnce(exit_endpoint, deadline_ms);
var receive: [8]u8 = undefined;
while (true) {
const got = ipc.replyWait(exit_endpoint, &.{}, &receive, null);
if (got.isChildExit() and got.childProcessId() == id) return;
if (got.isTimer()) break; // the deadline passed first — escalate
}
_ = system.kill(id);
while (true) {
const got = ipc.replyWait(exit_endpoint, &.{}, &receive, null);
if (got.isChildExit() and got.childProcessId() == id) return;
}
}
/// Subscribe `endpoint` to published exit events: every process death posts an
/// asynchronous notification with the same badge encoding as a supervisor's exit
/// notice (decode with `ipc.Received.isChildExit`/`childProcessId`). For stateful
/// services: release what the dead client held — file handles, subscriptions —
/// because a service must never depend on clients cleaning up after themselves
/// (docs/process-lifecycle.md). Ungated, like `system.processes`.
pub fn subscribeExits(endpoint: usize) bool {
return sc.systemCall1(.process_subscribe, endpoint) == 0;
}
+20 -1
View File
@@ -8,7 +8,8 @@
//! const runtime = @import("runtime");
//! pub const panic = runtime.panic;
//! comptime { _ = &runtime.start._start; } // pull the entry shim in
//! and a `pub fn main() void`.
//! and a `pub fn main() void` or `pub fn main(init: runtime.process.Init) void`
//! (arguments arrive via `init`).
pub const system = @import("system.zig");
pub const heap = @import("heap.zig");
@@ -16,6 +17,17 @@ pub const ipc = @import("ipc.zig");
pub const start = @import("start.zig");
/// The VFS wire protocol (shared with the VFS server).
pub const vfs_protocol = @import("vfs-protocol");
/// The device-manager protocol: hello + tree reports (docs/device-manager.md).
pub const device_manager_protocol = @import("device-manager-protocol");
/// The power protocol: events (button, lid, battery) + shutdown (docs/m21-plan.md).
pub const power_protocol = @import("power-protocol");
/// Keyboard-event listening (subscribe/next) and broadcasting (publish), over the input
/// service. See library/runtime/input.zig and system/services/input/.
pub const input = @import("input.zig");
/// The input wire protocol (shared with the input service and its clients).
pub const input_protocol = @import("input-protocol");
/// POSIX-style file API: open/read/write/lseek/stat/close.
/// C stdio: fopen/fread/fwrite/fseek/ftell/fclose over unistd.
/// Device access for drivers: enumerate/claim/mmioMap.
@@ -26,5 +38,12 @@ pub const dma = @import("dma.zig");
/// Re-exported so a user binary can `pub const panic = runtime.panic;`.
pub const panic = start.panic;
/// Process entry types: the `Init` handed to `main`, and its `Arguments`.
pub const process = @import("process.zig");
/// The service harness: one replyWait loop folding requests, signals, and
/// notifications into callbacks (docs/process-lifecycle.md).
pub const service = @import("service.zig");
/// The heap as a `std.mem.Allocator`, for Zig `std` containers in user code.
pub const allocator = heap.allocator;
+83
View File
@@ -0,0 +1,83 @@
//! The service harness (docs/process-lifecycle.md): one replyWait loop that
//! folds protocol requests, signals, and subscribed notifications into
//! callbacks — so the lifecycle contract ("answers ping, exits on terminate")
//! is satisfied by construction and a service author writes domain logic only.
//! Nothing is asynchronous inside the process: a callback runs at a point the
//! loop chose, never on a hijacked stack — the whole reason signals are
//! messages.
//!
//! The liveness probe: a **zero-length request is the universal ping**, answered
//! with a zero-length reply by the harness itself. No protocol's requests start
//! at length zero, so the encoding cannot collide, and there is nothing for a
//! service author to implement — a wedged service simply fails to answer, which
//! is the diagnosis (see docs/ipc.md).
const abi = @import("abi");
const ipc = @import("ipc.zig");
const process = @import("process.zig");
pub const Callbacks = struct {
/// Called once with the service's endpoint before the loop starts — the
/// place to subscribe to exit events, bind IRQs, or announce readiness.
/// Return false to abort startup (the process exits).
init: ?*const fn (endpoint: ipc.Handle) bool = null,
/// One protocol request from `sender` (a task id): write the reply into
/// `reply`, return its length. `capability` is the handle the request
/// carried, if any (M13 cap passing — how a subscriber hands over its
/// endpoint). The zero-length ping never reaches this.
on_message: *const fn (message: []const u8, reply: []u8, sender: u32, capability: ?ipc.Handle) usize,
/// A notification that is not a signal — a subscribed exit event, a bound
/// IRQ, a timer landing. The raw badge; decode with the ipc helpers.
on_notification: ?*const fn (badge: u64) void = null,
/// The reload signal. Default: ignored.
on_reload: ?*const fn () void = null,
/// The terminate signal, called before the loop returns. The clean exit is
/// the return itself — never put *necessary* work here (iron rule 1: a kill
/// arrives with no warning; this is for graceful extras only).
on_terminate: ?*const fn () void = null,
/// Publish the endpoint under a well-known service id at startup.
service: ?abi.ServiceId = null,
};
/// Run the service: create and (optionally) register the endpoint, bind signals
/// to it, call `init`, then serve until `terminate` arrives — at which point the
/// loop returns and main's return is the clean exit the supervisor reads as
/// `ExitReason.exited`. `maximum_message` sizes the receive and reply buffers
/// (a service passes its protocol's message maximum).
pub fn run(comptime maximum_message: usize, callbacks: Callbacks) void {
const endpoint = ipc.createIpcEndpoint() orelse return;
if (callbacks.service) |id| {
if (!ipc.register(id, endpoint)) return;
}
_ = process.bindSignals(endpoint);
if (callbacks.init) |initialise| {
if (!initialise(endpoint)) return;
}
var reply_buffer: [maximum_message]u8 = undefined;
var reply_len: usize = 0;
var receive: [maximum_message]u8 = undefined;
while (true) {
const got = ipc.replyWait(endpoint, reply_buffer[0..reply_len], &receive, null);
if (got.isNotification()) {
reply_len = 0; // nothing owed for a notification
if (process.signalsFrom(got.badge)) |signals| {
if (signals.has(.reload)) {
if (callbacks.on_reload) |onReload| onReload();
}
if (signals.has(.terminate)) {
if (callbacks.on_terminate) |onTerminate| onTerminate();
return; // the loop's return IS the clean exit
}
continue;
}
if (callbacks.on_notification) |onNotification| onNotification(got.badge);
continue;
}
if (got.len == 0) {
reply_len = 0; // the universal ping: a zero-length reply, from the harness
continue;
}
reply_len = callbacks.on_message(receive[0..got.len], &reply_buffer, got.senderTaskId(), got.cap);
}
}
+63 -9
View File
@@ -4,24 +4,78 @@
const std = @import("std");
const system = @import("system.zig");
const process = @import("process.zig");
/// The kernel enters at `_start` with rsp 16-aligned, but a SystemV function expects
/// rsp ≡ 8 (mod 16) on entry (as if reached by `call`). The `call` below pushes
/// the 8-byte return address, satisfying the ABI before any Zig frame runs; the
/// `ud2` is a safety net if `rt_start` ever returns.
/// The kernel enters at `_start` with rsp 16-aligned, pointing at the System V
/// process-entry block it built: argc, argv pointers, NULL, envp terminator, the
/// auxiliary vector, then the strings (see system/kernel/process.zig,
/// `buildEntryStack`). Capture that address in rdi — the first SysV argument —
/// before `call` disturbs the stack; the call's pushed return address also puts
/// rsp ≡ 8 (mod 16), satisfying the ABI before any Zig frame runs. The `ud2` is a
/// safety net if `rt_start` ever returns.
pub export fn _start() callconv(.naked) noreturn {
asm volatile (
\\mov %%rsp, %%rdi
\\call rt_start
\\ud2
);
}
/// The first Zig frame. The heap is lazy (first alloc grows it), so there is no
/// runtime init to order here — just hand control to the program's `main`.
export fn rt_start() callconv(.c) noreturn {
/// The first Zig frame, entered with `stack` pointing at the kernel-built entry
/// block. Build the `process.Init` from it and dispatch to the program's `main`,
/// whose signature is inspected at comptime. The heap is lazy (first alloc grows
/// it), so there is no other runtime init to order here.
export fn rt_start(stack: [*]const u64) callconv(.c) noreturn {
const init: process.Init = .{ .arguments = .{
.count = stack[0],
.vector = @ptrCast(stack + 1),
} };
system.exit(callMain(init));
}
/// Comptime-dispatch on root.main's signature, in the spirit of std's start.zig:
/// zero parameters or one `process.Init`; returns void, noreturn, u8, !void, or !u8.
fn callMain(init: process.Init) u8 {
const root = @import("root"); // the user binary's root source file
root.main();
system.exit(0);
const main_information = @typeInfo(@TypeOf(root.main)).@"fn";
const call_arguments = switch (main_information.params.len) {
0 => .{},
1 => arguments: {
const Parameter = main_information.params[0].type orelse
@compileError("main's parameter must be runtime.process.Init (not anytype)");
if (Parameter != process.Init)
@compileError("main's parameter must be runtime.process.Init, found " ++ @typeName(Parameter));
break :arguments .{init};
},
else => @compileError("main takes no parameters or a single runtime.process.Init"),
};
const ReturnType = main_information.return_type.?;
switch (@typeInfo(ReturnType)) {
.noreturn => @call(.auto, root.main, call_arguments),
.void => {
@call(.auto, root.main, call_arguments);
return 0;
},
.int => {
if (ReturnType != u8)
@compileError("main's integer return type must be u8, found " ++ @typeName(ReturnType));
return @call(.auto, root.main, call_arguments);
},
.error_union => {
const payload = @call(.auto, root.main, call_arguments) catch |err| {
var buffer: [128]u8 = undefined;
const line = std.fmt.bufPrint(&buffer, "main returned error: {s}\n", .{@errorName(err)}) catch "main returned an error\n";
_ = system.write(line);
return 1; // distinct from panic's 127
};
if (@TypeOf(payload) == void) return 0;
if (@TypeOf(payload) == u8) return payload;
@compileError("main's error-union payload must be void or u8, found " ++ @typeName(@TypeOf(payload)));
},
else => @compileError("main must return void, noreturn, u8, !void, or !u8, found " ++ @typeName(ReturnType)),
}
}
/// No runtime to unwind into — report a panic as a nonzero exit code.
+20 -5
View File
@@ -19,34 +19,49 @@ pub inline fn systemCall0(n: SystemCall) usize {
pub inline fn systemCall1(n: SystemCall, a0: usize) usize {
return asm volatile ("syscall"
: [ret] "={rax}" (-> usize),
: [n] "{rax}" (@intFromEnum(n)), [a0] "{rdi}" (a0),
: [n] "{rax}" (@intFromEnum(n)),
[a0] "{rdi}" (a0),
: .{ .rcx = true, .r11 = true, .memory = true });
}
pub inline fn systemCall2(n: SystemCall, a0: usize, a1: usize) usize {
return asm volatile ("syscall"
: [ret] "={rax}" (-> usize),
: [n] "{rax}" (@intFromEnum(n)), [a0] "{rdi}" (a0), [a1] "{rsi}" (a1),
: [n] "{rax}" (@intFromEnum(n)),
[a0] "{rdi}" (a0),
[a1] "{rsi}" (a1),
: .{ .rcx = true, .r11 = true, .memory = true });
}
pub inline fn systemCall3(n: SystemCall, a0: usize, a1: usize, a2: usize) usize {
return asm volatile ("syscall"
: [ret] "={rax}" (-> usize),
: [n] "{rax}" (@intFromEnum(n)), [a0] "{rdi}" (a0), [a1] "{rsi}" (a1), [a2] "{rdx}" (a2),
: [n] "{rax}" (@intFromEnum(n)),
[a0] "{rdi}" (a0),
[a1] "{rsi}" (a1),
[a2] "{rdx}" (a2),
: .{ .rcx = true, .r11 = true, .memory = true });
}
pub inline fn systemCall4(n: SystemCall, a0: usize, a1: usize, a2: usize, a3: usize) usize {
return asm volatile ("syscall"
: [ret] "={rax}" (-> usize),
: [n] "{rax}" (@intFromEnum(n)), [a0] "{rdi}" (a0), [a1] "{rsi}" (a1), [a2] "{rdx}" (a2), [a3] "{r10}" (a3),
: [n] "{rax}" (@intFromEnum(n)),
[a0] "{rdi}" (a0),
[a1] "{rsi}" (a1),
[a2] "{rdx}" (a2),
[a3] "{r10}" (a3),
: .{ .rcx = true, .r11 = true, .memory = true });
}
pub inline fn systemCall5(n: SystemCall, a0: usize, a1: usize, a2: usize, a3: usize, a4: usize) usize {
return asm volatile ("syscall"
: [ret] "={rax}" (-> usize),
: [n] "{rax}" (@intFromEnum(n)), [a0] "{rdi}" (a0), [a1] "{rsi}" (a1), [a2] "{rdx}" (a2), [a3] "{r10}" (a3), [a4] "{r8}" (a4),
: [n] "{rax}" (@intFromEnum(n)),
[a0] "{rdi}" (a0),
[a1] "{rsi}" (a1),
[a2] "{rdx}" (a2),
[a3] "{r10}" (a3),
[a4] "{r8}" (a4),
: .{ .rcx = true, .r11 = true, .memory = true });
}
+82 -5
View File
@@ -2,6 +2,7 @@
//! stubs, one per kernel call. Numbers come from `abi.SystemCall`, the single
//! source of truth shared with the kernel dispatcher.
const std = @import("std");
const abi = @import("abi");
const sc = @import("system-call.zig");
@@ -11,6 +12,10 @@ pub const PROT_READ: usize = abi.prot_read;
pub const PROT_WRITE: usize = abi.prot_write;
pub const PROT_EXEC: usize = abi.prot_exec;
/// One `processes` entry — re-exported from the shared ABI so a user program can
/// declare its snapshot buffer without importing `abi` itself.
pub const ProcessDescriptor = abi.ProcessDescriptor;
/// Give up the rest of this quantum.
pub fn yield() void {
_ = sc.systemCall0(.yield);
@@ -27,6 +32,15 @@ pub fn sleep(ms: usize) void {
_ = sc.systemCall1(.sleep, ms);
}
/// Arm a one-shot timer: after `ms` milliseconds the kernel posts a timer
/// notification (`ipc.Received.isTimer`) to `endpoint`. The timed wait of
/// docs/process-lifecycle.md — a service arms a deadline and keeps serving,
/// instead of blocking in sleep; what stop-sequence escalation, hello deadlines,
/// and restart backoff are built from.
pub fn timerOnce(endpoint: usize, ms: u64) bool {
return sc.systemCall2(.timer_bind, endpoint, ms) == 0;
}
/// Monotonic nanoseconds since boot — a time source for timeouts and short delays. It
/// only ever moves forward. This is *not* wall-clock time (no date, no timezone — that
/// is a user-space service layered on top). Deadline pattern for a bounded poll loop:
@@ -44,11 +58,74 @@ pub fn exit(code: usize) noreturn {
}
/// Start the binary bundled in the initial-ramdisk under `name` as a new ring-3
/// process, returning true on success. This is how a supervisor (the device manager)
/// launches a driver it matched — danos-native, not POSIX (a spawn/exec family comes
/// with the process work later).
pub fn spawn(name: []const u8) bool {
return sc.systemCall2(.system_spawn, @intFromPtr(name.ptr), name.len) == 0;
/// process, returning the child's process id (or null on failure). The child's
/// argv[0] is `name`, and the caller becomes its **supervisor** — the only process
/// allowed to `kill` it. This is how a supervisor (the device manager) launches a
/// driver it matched — danos-native, not POSIX (a spawn/exec family comes with the
/// POSIX layer later).
pub fn spawn(name: []const u8) ?u32 {
return spawnSupervised(name, &.{}, null);
}
/// Like `spawn`, but hands the child command-line arguments: they arrive as
/// argv[1..] on its System V entry stack (argv[0] is still `name`).
pub fn spawnWithArguments(name: []const u8, arguments: []const []const u8) ?u32 {
return spawnSupervised(name, arguments, null);
}
/// The full spawn: command-line arguments for the child, and an optional endpoint
/// (a handle from `ipc.createIpcEndpoint`) the kernel notifies when the child ends
/// — any way it ends: clean exit, fault, or `kill`. The notification arrives via
/// `ipc.replyWait` as a badge with the child-exit bit set and the child's id in
/// the low bits (`ipc.Received.isChildExit`/`childProcessId`), so one endpoint can
/// supervise many children. Arguments are marshalled to the kernel as one
/// NUL-separated blob; the combined arguments must fit `blob` (the kernel caps the
/// blob at 256 bytes and argc at 8 anyway). Returns the child's process id, or
/// null on failure.
pub fn spawnSupervised(name: []const u8, arguments: []const []const u8, exit_endpoint: ?usize) ?u32 {
var blob: [256]u8 = undefined;
var len: usize = 0;
for (arguments, 0..) |argument, i| {
if (i != 0) {
if (len >= blob.len) return null;
blob[len] = 0;
len += 1;
}
if (len + argument.len > blob.len) return null;
@memcpy(blob[len..][0..argument.len], argument);
len += argument.len;
}
const r = sc.systemCall5(.system_spawn, @intFromPtr(name.ptr), name.len, if (len == 0) 0 else @intFromPtr(&blob), len, exit_endpoint orelse abi.no_cap);
if (r > ~@as(usize, 0) - 4095) return null; // a wrapped -errno
return @intCast(r);
}
/// Snapshot the process table into `out` (up to its length) and return the total
/// number of live processes — which may exceed `out.len`; call again with a larger
/// buffer for the full listing. Kernel tasks are included, with an empty name.
/// The primitive `ps` is built on.
pub fn processes(out: []abi.ProcessDescriptor) usize {
return sc.systemCall2(.process_enumerate, @intFromPtr(out.ptr), out.len);
}
/// Whether a process spawned under `name` (its argv[0]) is currently alive.
pub fn isProcessRunning(name: []const u8) bool {
var table: [32]ProcessDescriptor = undefined;
const total = processes(&table);
for (table[0..@min(total, table.len)]) |descriptor| {
if (std.mem.eql(u8, descriptor.name[0..descriptor.name_length], name)) return true;
}
return false;
}
/// End process `id`. Only its supervisor — the process that spawned it — may;
/// anyone else gets false, as does a stale or unknown id (ids are never reused).
/// Delivery is prompt but asynchronous, like a signal: a target caught running on
/// another core dies at its next system call or timer tick. True means the kill
/// is accepted and irrevocable; the exit notification (if an endpoint was given
/// at spawn) confirms completion.
pub fn kill(id: u32) bool {
return sc.systemCall1(.process_kill, id) == 0;
}
/// Grant `len` bytes (rounded up to whole pages) of fresh, zeroed, writable
+65
View File
@@ -0,0 +1,65 @@
# xkeyboard-config — X11 keyboard layouts, compiled to Zig
This module turns a physical key (a **USB HID usage**, as the [input module](../../docs/input.md)
delivers in `KeyEvent.keycode`) plus a modifier state into a **keysym** and, when the key
produces one, a **character** (a Unicode scalar). It is what lets a `keycode` become a
`character` — a keymap — without danos shipping an X11 runtime.
The layout data comes from the X11 [xkeyboard-config](https://gitlab.freedesktop.org/xkeyboard-config/xkeyboard-config)
database, but it is **compiled to native Zig at build time** rather than parsed at runtime.
`tools/make-xkeyboard-config.py` reads the vendored xkb data and emits pure-data tables into
`generated/layouts.zig`; `xkeyboard-config.zig` is the hand-written API over them. This is
the same build-time-codegen pattern as `tools/make-initial-ramdisk.py`.
## Using it
```zig
const xkb = @import("xkeyboard-config");
const m = xkb.map(xkb.us, key_event.keycode, .{ .shift = shift_held, .caps_lock = caps });
if (m.character) |ch| { /* a printable Unicode scalar */ }
// m.keysym is always set (e.g. an X11 keysym for Return / F1 / a dead key).
const layout = xkb.byName("gb") orelse xkb.us; // choose a layout by name
for (xkb.all) |l| { /* enumerate available layouts */ }
```
`Modifiers` carries `shift`, `caps_lock`, `level3` (AltGr), and `control`. `map` selects the
level from the key's XKB *type* (the generated data) and those modifiers (the policy, in
`xkeyboard-config.zig`), so data and semantics stay separable.
Layouts: **us, gb, de, fr, es, dvorak**.
## Regenerating
```sh
python3 tools/make-xkeyboard-config.py fetch # network: download + vendor the data subset
python3 tools/make-xkeyboard-config.py generate # offline: emit generated/layouts.zig
# or, from the build:
zig build gen-xkeyboard-config
```
- **`fetch`** downloads the pinned xkeyboard-config release (version + sha256 in the script),
resolves the `include` graph for the configured layouts, and vendors *only* the symbols
files actually reached (plus `keysymdef.h`, `COPYING`, and `PROVENANCE.md`) into `vendor/`.
Run it when bumping the version or adding a layout.
- **`generate`** is deterministic and offline — same vendored input produces byte-identical
output. To add a layout, extend `TARGETS` (and `HID_TO_NAME` if a new physical key is
involved), then re-run `fetch` (to vendor any new includes) and `generate`.
## Scope
A pragmatic subset, enough for real Latin-script typing:
- **Group 1 only** — no multi-layout group switching.
- **No dead-key / compose composition** — a dead key returns its keysym with no `character`
(composing `´` + `e` → `é` is a higher layer's job).
- **Curated key types** — the common XKB types (one/two-level, alphabetic, four-level, …);
unmapped keys and unknown types fall back to level-by-shift.
- **6 layouts** — extend via `TARGETS` as above.
## Licensing
xkeyboard-config and `keysymdef.h` (xorgproto) are MIT/X11 licensed. The vendored data
subset carries the upstream `vendor/COPYING`, and `vendor/PROVENANCE.md` records the exact
version, source URL, and sha256. The generated tables are a derived work under the same terms.
File diff suppressed because it is too large Load Diff
+190
View File
@@ -0,0 +1,190 @@
Copyright 1996 by Joseph Moss
Copyright (C) 2002-2007 Free Software Foundation, Inc.
Copyright (C) Dmitry Golubev <lastguru@mail.ru>, 2003-2004
Copyright (C) 2004, Gregory Mokhin <mokhin@bog.msu.ru>
Copyright (C) 2006 Erdal Ronahî
Permission to use, copy, modify, distribute, and sell this software and its
documentation for any purpose is hereby granted without fee, provided that
the above copyright notice appear in all copies and that both that
copyright notice and this permission notice appear in supporting
documentation, and that the name of the copyright holder(s) not be used in
advertising or publicity pertaining to distribution of the software without
specific, written prior permission. The copyright holder(s) makes no
representations about the suitability of this software for any purpose. It
is provided "as is" without express or implied warranty.
THE COPYRIGHT HOLDER(S) DISCLAIMS ALL WARRANTIES WITH REGARD TO THIS SOFTWARE,
INCLUDING ALL IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS, IN NO
EVENT SHALL THE COPYRIGHT HOLDER(S) BE LIABLE FOR ANY SPECIAL, INDIRECT OR
CONSEQUENTIAL DAMAGES OR ANY DAMAGES WHATSOEVER RESULTING FROM LOSS OF USE,
DATA OR PROFITS, WHETHER IN AN ACTION OF CONTRACT, NEGLIGENCE OR OTHER
TORTIOUS ACTION, ARISING OUT OF OR IN CONNECTION WITH THE USE OR
PERFORMANCE OF THIS SOFTWARE.
Copyright (c) 1996 Digital Equipment Corporation
Permission is hereby granted, free of charge, to any person obtaining
a copy of this software and associated documentation files (the
"Software"), to deal in the Software without restriction, including
without limitation the rights to use, copy, modify, merge, publish,
distribute, sublicense, and sell copies of the Software, and to
permit persons to whom the Software is furnished to do so, subject to
the following conditions:
The above copyright notice and this permission notice shall be included
in all copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS
OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF
MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.
IN NO EVENT SHALL DIGITAL EQUIPMENT CORPORATION BE LIABLE FOR ANY CLAIM,
DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR
OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR
THE USE OR OTHER DEALINGS IN THE SOFTWARE.
Except as contained in this notice, the name of the Digital Equipment
Corporation shall not be used in advertising or otherwise to promote
the sale, use or other dealings in this Software without prior written
authorization from Digital Equipment Corporation.
Copyright 1996, 1998 The Open Group
Permission to use, copy, modify, distribute, and sell this software and its
documentation for any purpose is hereby granted without fee, provided that
the above copyright notice appear in all copies and that both that
copyright notice and this permission notice appear in supporting
documentation.
The above copyright notice and this permission notice shall be
included in all copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND,
EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF
MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.
IN NO EVENT SHALL THE OPEN GROUP BE LIABLE FOR ANY CLAIM, DAMAGES OR
OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE,
ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR
OTHER DEALINGS IN THE SOFTWARE.
Except as contained in this notice, the name of The Open Group shall
not be used in advertising or otherwise to promote the sale, use or
other dealings in this Software without prior written authorization
from The Open Group.
Copyright 2004-2005 Sun Microsystems, Inc. All rights reserved.
Permission is hereby granted, free of charge, to any person obtaining a
copy of this software and associated documentation files (the "Software"),
to deal in the Software without restriction, including without limitation
the rights to use, copy, modify, merge, publish, distribute, sublicense,
and/or sell copies of the Software, and to permit persons to whom the
Software is furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice (including the next
paragraph) shall be included in all copies or substantial portions of the
Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL
THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING
FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER
DEALINGS IN THE SOFTWARE.
Copyright (c) 1996 by Silicon Graphics Computer Systems, Inc.
Permission to use, copy, modify, and distribute this
software and its documentation for any purpose and without
fee is hereby granted, provided that the above copyright
notice appear in all copies and that both that copyright
notice and this permission notice appear in supporting
documentation, and that the name of Silicon Graphics not be
used in advertising or publicity pertaining to distribution
of the software without specific prior written permission.
Silicon Graphics makes no representation about the suitability
of this software for any purpose. It is provided "as is"
without any express or implied warranty.
SILICON GRAPHICS DISCLAIMS ALL WARRANTIES WITH REGARD TO THIS
SOFTWARE, INCLUDING ALL IMPLIED WARRANTIES OF MERCHANTABILITY
AND FITNESS FOR A PARTICULAR PURPOSE. IN NO EVENT SHALL SILICON
GRAPHICS BE LIABLE FOR ANY SPECIAL, INDIRECT OR CONSEQUENTIAL
DAMAGES OR ANY DAMAGES WHATSOEVER RESULTING FROM LOSS OF USE,
DATA OR PROFITS, WHETHER IN AN ACTION OF CONTRACT, NEGLIGENCE
OR OTHER TORTIOUS ACTION, ARISING OUT OF OR IN CONNECTION WITH
THE USE OR PERFORMANCE OF THIS SOFTWARE.
Copyright (c) 1996 X Consortium
Permission is hereby granted, free of charge, to any person obtaining
a copy of this software and associated documentation files (the
"Software"), to deal in the Software without restriction, including
without limitation the rights to use, copy, modify, merge, publish,
distribute, sublicense, and/or sell copies of the Software, and to
permit persons to whom the Software is furnished to do so, subject to
the following conditions:
The above copyright notice and this permission notice shall be
included in all copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND,
EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF
MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.
IN NO EVENT SHALL THE X CONSORTIUM BE LIABLE FOR ANY CLAIM, DAMAGES OR
OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE,
ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR
OTHER DEALINGS IN THE SOFTWARE.
Except as contained in this notice, the name of the X Consortium shall
not be used in advertising or otherwise to promote the sale, use or
other dealings in this Software without prior written authorization
from the X Consortium.
Copyright (C) 2004, 2006 Ævar Arnfjörð Bjarmason <avarab@gmail.com>
Permission to use, copy, modify, distribute, and sell this software and its
documentation for any purpose is hereby granted without fee, provided that
the above copyright notice appear in all copies and that both that
copyright notice and this permission notice appear in supporting
documentation.
The above copyright notice and this permission notice shall be
included in all copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND,
EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF
MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.
IN NO EVENT SHALL THE OPEN GROUP BE LIABLE FOR ANY CLAIM, DAMAGES OR
OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE,
ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR
OTHER DEALINGS IN THE SOFTWARE.
Except as contained in this notice, the name of a copyright holder shall
not be used in advertising or otherwise to promote the sale, use or
other dealings in this Software without prior written authorization of
the copyright holder.
Copyright (C) 1999, 2000 by Anton Zinoviev <anton@lml.bas.bg>
This software may be used, modified, copied, distributed, and sold,
in both source and binary form provided that the above copyright
and these terms are retained. Under no circumstances is the author
responsible for the proper functioning of this software, nor does
the author assume any responsibility for damages incurred with its
use.
Permission is granted to anyone to use, distribute and modify
this file in any way, provided that the above copyright notice
is left intact and the author of the modification summarizes
the changes in this header.
This file is distributed without any expressed or implied warranty.
+21
View File
@@ -0,0 +1,21 @@
# Vendored xkeyboard-config subset
- **Package**: xkeyboard-config 2.44
- **Source**: https://gitlab.freedesktop.org/xkeyboard-config/xkeyboard-config/-/archive/xkeyboard-config-2.44/xkeyboard-config-2.44.tar.gz
- **sha256**: `35e34edeaf4e8da8d0696ff6b241ee11ddb1b8c6730bac7252d4d0a88ea5f05b`
- **keysymdef.h**: xorgproto, copied from `/opt/homebrew/include/X11/keysymdef.h`
- **License**: MIT/X11 (see COPYING)
Only the symbols files reachable from the generated layouts (tools/make-xkeyboard-config.py `TARGETS`) are vendored; regenerate with
`python3 tools/make-xkeyboard-config.py fetch` then `... generate`.
Vendored symbols files:
- `symbols/de`
- `symbols/es`
- `symbols/fr`
- `symbols/gb`
- `symbols/kpdl`
- `symbols/latin`
- `symbols/level3`
- `symbols/us`
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+250
View File
@@ -0,0 +1,250 @@
// Keyboard layouts for Spain.
// Modified for a real Spanish keyboard by Jon Tombs.
default partial alphanumeric_keys
xkb_symbols "basic" {
include "latin(type4)"
name[Group1]="Spanish";
key <TLDE> { [ masculine, ordfeminine, backslash, backslash ] };
key <AE01> { [ 1, exclam, bar, exclamdown ] };
key <AE03> { [ 3, periodcentered, numbersign, sterling ] };
key <AE04> { [ 4, dollar, asciitilde, dollar ] };
key <AE11> { [apostrophe, question, backslash, questiondown ] };
key <AE12> { [exclamdown, questiondown, dead_cedilla, dead_ogonek] };
key <AD11> { [dead_grave, dead_circumflex, bracketleft, dead_abovering ] };
key <AD12> { [ plus, asterisk, bracketright, dead_macron ] };
key <AC10> { [ ntilde, Ntilde, dead_tilde, dead_doubleacute ] };
key <AC11> { [dead_acute, dead_diaeresis, braceleft, dead_caron ] };
key <BKSL> { [ ccedilla, Ccedilla, braceright, dead_breve ] };
include "level3(ralt_switch)"
};
partial alphanumeric_keys
xkb_symbols "winkeys" {
include "es(basic)"
name[Group1]="Spanish (Windows)";
include "eurosign(5)"
};
partial alphanumeric_keys
xkb_symbols "nodeadkeys" {
include "es(basic)"
name[Group1]="Spanish (no dead keys)";
key <AE12> { [exclamdown, questiondown, cedilla, ogonek ] };
key <AD11> { [ grave, asciicircum, bracketleft, degree ] };
key <AD12> { [ plus, asterisk, bracketright, macron ] };
key <AC07> { [ j, J, ezh, EZH ] };
key <AC10> { [ ntilde, Ntilde, asciitilde, doubleacute ] };
key <AC11> { [ acute, diaeresis, braceleft, caron ] };
key <BKSL> { [ ccedilla, Ccedilla, braceright, breve ] };
key <AB10> { [ minus, underscore, ellipsis, abovedot ] };
};
// Spanish Dvorak mapping (note R-H exchange)
partial alphanumeric_keys
xkb_symbols "dvorak" {
name[Group1]="Spanish (Dvorak)";
key <TLDE> {[ masculine, ordfeminine, backslash, degree ]};
key <AE01> {[ 1, exclam, bar, onesuperior ]};
key <AE02> {[ 2, quotedbl, at, twosuperior ]};
key <AE03> {[ 3, periodcentered, numbersign, threesuperior ]};
key <AE04> {[ 4, dollar, asciitilde, onequarter ]};
key <AE05> {[ 5, percent, brokenbar, fiveeighths ]};
key <AE06> {[ 6, ampersand, notsign, threequarters ]};
key <AE07> {[ 7, slash, onehalf, seveneighths ]};
key <AE08> {[ 8, parenleft, oneeighth, threeeighths ]};
key <AE09> {[ 9, parenright, asciicircum ]};
key <AE10> {[ 0, equal, grave, dead_doubleacute ]};
key <AE11> {[ apostrophe, question, dead_macron, dead_ogonek ]};
key <AE12> {[ exclamdown, questiondown, dead_breve, dead_abovedot ]};
key <AD01> {[ period, colon, less, guillemotleft ]};
key <AD02> {[ comma, semicolon, greater, guillemotright ]};
key <AD03> {[ ntilde, Ntilde, lstroke, Lstroke ]};
key <AD04> {[ p, P, paragraph ]};
key <AD05> {[ y, Y, yen ]};
key <AD06> {[ f, F, tslash, Tslash ]};
key <AD07> {[ g, G, dstroke, Dstroke ]};
key <AD08> {[ c, C, cent, copyright ]};
key <AD09> {[ h, H, hstroke, Hstroke ]};
key <AD10> {[ l, L, sterling ]};
key <AD11> {[ dead_grave, dead_circumflex, bracketleft, dead_caron ]};
key <AD12> {[ plus, asterisk, bracketright, plusminus ]};
key <AC01> {[ a, A, ae, AE ]};
key <AC02> {[ o, O, oslash, Oslash ]};
key <AC03> {[ e, E, EuroSign ]};
key <AC04> {[ u, U, aring, Aring ]};
key <AC05> {[ i, I, oe, OE ]};
key <AC06> {[ d, D, eth, ETH ]};
key <AC07> {[ r, R, registered, trademark ]};
key <AC08> {[ t, T, thorn, THORN ]};
key <AC09> {[ n, N, eng, ENG ]};
key <AC10> {[ s, S, ssharp, section ]};
key <AC11> {[ dead_acute, dead_diaeresis, braceleft, dead_tilde ]};
key <BKSL> {[ ccedilla, Ccedilla, braceright, dead_cedilla ]};
key <LSGT> {[ less, greater, guillemotleft, guillemotright ]};
key <AB01> {[ minus, underscore, hyphen, macron ]};
key <AB02> {[ q, Q, currency ]};
key <AB03> {[ j, J ]};
key <AB04> {[ k, K, kra ]};
key <AB05> {[ x, X, multiply, division ]};
key <AB06> {[ b, B ]};
key <AB07> {[ m, M, mu ]};
key <AB08> {[ w, W ]};
key <AB09> {[ v, V ]};
key <AB10> {[ z, Z ]};
include "level3(ralt_switch)"
};
partial alphanumeric_keys
xkb_symbols "cat" {
include "es(basic)"
name[Group1]="Catalan (Spain, with middle-dot L)";
key <AC09> { [ l, L, 0x1000140, 0x100013F ] };
};
partial alphanumeric_keys
xkb_symbols "ast" {
include "es(basic)"
name[Group1]="Asturian (Spain, with bottom-dot H and L)";
key <AC06> { [ h, H, 0x1001E25, 0x1001E24 ] };
key <AC09> { [ l, L, 0x1001E37, 0x1001E36 ] };
};
partial alphanumeric_keys
xkb_symbols "olpc" {
// #HW-SPECIFIC
// http://wiki.laptop.org/go/OLPC_Spanish_Keyboard
include "us(basic)"
name[Group1]="Spanish";
key <AE00> { [ masculine, ordfeminine ] };
key <AE01> { [ 1, exclam, bar ] };
key <AE02> { [ 2, quotedbl, at ] };
key <AE03> { [ 3, dead_grave, numbersign, grave ] };
key <AE05> { [ 5, percent, asciicircum, dead_circumflex ] };
key <AE06> { [ 6, ampersand, notsign ] };
key <AE07> { [ 7, slash, backslash ] };
key <AE08> { [ 8, parenleft ] };
key <AE09> { [ 9, parenright ] };
key <AE10> { [ 0, equal ] };
key <AE11> { [ apostrophe, question ] };
key <AE12> { [ exclamdown, questiondown ] };
key <AD03> { [ e, E, EuroSign ] };
key <AD11> { [ dead_acute, dead_diaeresis, acute, dead_abovering ] };
key <AD12> { [ bracketleft, braceleft ] };
key <AC10> { [ ntilde, Ntilde ] };
key <AC11> { [ plus, asterisk, dead_tilde ] };
key <AC12> { [ bracketright, braceright, section ] };
key <AB08> { [ comma, semicolon ] };
key <AB09> { [ period, colon ] };
key <AB10> { [ minus, underscore ] };
key <I219> { [ less, greater, ISO_Next_Group ] };
include "level3(ralt_switch)"
};
partial alphanumeric_keys
xkb_symbols "olpcm" {
// #HW-SPECIFIC
// Mechanical (non-membrane) OLPC Spanish keyboard layout.
// See: http://wiki.laptop.org/go/OLPC_Spanish_Non-membrane_Keyboard
include "us(basic)"
name[Group1]="Spanish";
key <AE00> { [ questiondown, exclamdown, backslash ] };
key <AE01> { [ 1, exclam, bar ] };
key <AE02> { [ 2, quotedbl, at ] };
key <AE03> { [ 3, dead_grave, numbersign, grave ] };
key <AE04> { [ 4, dollar, asciitilde, dead_tilde ] };
key <AE05> { [ 5, percent, asciicircum, dead_circumflex ] };
key <AE06> { [ 6, ampersand, notsign ] };
key <AE07> { [ 7, slash, backslash ] }; // no '\' label on olpcm, leave for compatibility
key <AE08> { [ 8, parenleft, masculine ] };
key <AE09> { [ 9, parenright, ordfeminine ] };
key <AE10> { [ 0, equal ] };
key <AE11> { [ apostrophe, question ] };
key <AD03> { [ e, E, EuroSign ] };
key <AD11> { [ dead_acute, dead_diaeresis, dead_abovering, acute ] };
key <AD12> { [ plus, asterisk ] };
key <AC10> { [ ntilde, Ntilde ] };
// no AC11 or AC12 on olpcm
key <AB08> { [ comma, semicolon ] };
key <AB09> { [ period, colon ] };
key <AB10> { [ minus, underscore ] };
key <AA02> { [ less, greater ] };
key <AA06> { [ bracketleft, braceleft, ccedilla, Ccedilla ] };
key <AA07> { [ bracketright, braceright ] };
include "level3(ralt_switch)"
};
partial alphanumeric_keys
xkb_symbols "deadtilde" {
include "es(basic)"
name[Group1]="Spanish (dead tilde)";
key <AE04> { [ 4, dollar, dead_tilde, dollar ] };
key <AC10> { [ ntilde, Ntilde, asciitilde, dead_doubleacute ] };
};
partial alphanumeric_keys
xkb_symbols "olpc2" {
// #HW-SPECIFIC
// Modified variant of US International layout, specifically for Peru
// Contact: Sayamindu Dasgupta <sayamindu@laptop.org>
include "us(olpc)"
name[Group1]="Spanish";
key <AE03> { [ 3, numbersign, dead_grave, dead_grave] }; // combining grave
key <I236> { [ XF86Start ] };
include "level3(ralt_switch)"
};
// EXTRAS:
partial alphanumeric_keys
xkb_symbols "sun_type6" {
include "sun_vndr/es(sun_type6)"
};
File diff suppressed because it is too large Load Diff
+249
View File
@@ -0,0 +1,249 @@
// Keyboard layouts for Great Britain.
default partial alphanumeric_keys
xkb_symbols "basic" {
// The basic UK layout, also known as the IBM 166 layout,
// but with the useless brokenbar pushed two levels up.
include "latin"
name[Group1]="English (UK)";
key <TLDE> { [ grave, notsign, bar, bar ] };
key <AE02> { [ 2, quotedbl, twosuperior, oneeighth ] };
key <AE03> { [ 3, sterling, threesuperior, sterling ] };
key <AE04> { [ 4, dollar, EuroSign, onequarter ] };
key <AC11> { [apostrophe, at, dead_circumflex, dead_caron] };
key <BKSL> { [numbersign, asciitilde, dead_grave, dead_breve ] };
key <LSGT> { [ backslash, bar, bar, brokenbar ] };
include "level3(ralt_switch)"
};
partial alphanumeric_keys
xkb_symbols "intl" {
// A UK layout but with five accents made into dead keys:
// grave, diaeresis, circumflex, acute, and tilde.
// By Phil Jones <philjones1 at blueyonder.co.uk>.
include "latin"
name[Group1]="English (UK, intl., with dead keys)";
key <TLDE> { [ dead_grave, notsign, bar, bar ] };
key <AE02> { [ 2, dead_diaeresis, twosuperior, onehalf ] };
key <AE03> { [ 3, sterling, threesuperior, onethird ] };
key <AE04> { [ 4, dollar, EuroSign, onequarter ] };
key <AE06> { [ 6, dead_circumflex, threequarters, onesixth ] };
key <AC11> { [ dead_acute, at, apostrophe, bar ] };
key <BKSL> { [ numbersign, dead_tilde, bar, bar ] };
key <LSGT> { [ backslash, bar, bar, bar ] };
key <AB08> { [ comma, less, ccedilla, Ccedilla ] };
include "level3(ralt_switch)"
};
partial alphanumeric_keys
xkb_symbols "extd" {
// Clone of the Microsoft "United Kingdom Extended" layout, which
// includes dead keys for: grave; diaeresis; circumflex; tilde; and
// accute. It also enables direct access to accute characters using
// the Multi_key (Alt Gr).
//
// Taken from...
// "Windows Keyboard Layouts"
// https://docs.microsoft.com/en-gb/globalization/windows-keyboard-layouts#U
//
// -- Jonathan Miles <jon@cybah.co.uk>
include "latin"
name[Group1]="English (UK, extended, Windows)";
key <TLDE> { [ dead_grave, notsign, brokenbar, NoSymbol ] };
key <AE02> { [ 2, quotedbl, dead_diaeresis, onehalf ] };
key <AE03> { [ 3, sterling, threesuperior, onethird ] };
key <AE04> { [ 4, dollar, EuroSign, onequarter ] };
key <AE06> { [ 6, asciicircum, dead_circumflex, NoSymbol ] };
key <AD02> { [ w, W, wacute, Wacute ] };
key <AD03> { [ e, E, eacute, Eacute ] };
key <AD06> { [ y, Y, yacute, Yacute ] };
key <AD07> { [ u, U, uacute, Uacute ] };
key <AD08> { [ i, I, iacute, Iacute ] };
key <AD09> { [ o, O, oacute, Oacute ] };
key <AD12> { [ bracketright, braceright, NoSymbol, bar ] };
key <AC01> { [ a, A, aacute, Aacute ] };
key <AC11> { [ apostrophe, at, dead_acute, grave ] };
key <BKSL> { [ numbersign, asciitilde, dead_tilde, backslash ] };
key <LSGT> { [ backslash, bar, NoSymbol, NoSymbol ] };
key <AB03> { [ c, C, ccedilla, Ccedilla ] };
include "level3(ralt_switch)"
};
// Describe the differences between the US Colemak layout
// and a UK variant. By Andy Buckley (andy@insectnation.org)
partial alphanumeric_keys
xkb_symbols "colemak" {
include "us(colemak)"
name[Group1]="English (UK, Colemak)";
key <TLDE> { [ grave, notsign, bar, asciitilde ] };
key <AE02> { [ 2, quotedbl, twosuperior, oneeighth ] };
key <AE03> { [ 3, sterling, threesuperior, sterling ] };
key <AE04> { [ 4, dollar, EuroSign, onequarter ] };
key <AC11> { [apostrophe, at, dead_circumflex, dead_caron] };
key <BKSL> { [numbersign, asciitilde, dead_grave, dead_breve ] };
key <LSGT> { [ backslash, bar, asciitilde, brokenbar ] };
};
// Colemak-DH (ISO) layout, UK Variant, https://colemakmods.github.io/mod-dh/
partial alphanumeric_keys
xkb_symbols "colemak_dh" {
include "us(colemak_dh)"
name[Group1]="English (UK, Colemak-DH)";
key <TLDE> { [ grave, notsign, bar, asciitilde ] };
key <AE02> { [ 2, quotedbl, twosuperior, oneeighth ] };
key <AE03> { [ 3, sterling, threesuperior, sterling ] };
key <AE04> { [ 4, dollar, EuroSign, onequarter ] };
key <AC11> { [apostrophe, at, dead_circumflex, dead_caron] };
key <BKSL> { [numbersign, asciitilde, dead_grave, dead_breve ] };
key <AB05> { [ backslash, bar, asciitilde, brokenbar ] };
};
// Dvorak (UK) keymap (by odaen) allowing the usage of
// the £ and ? key and swapping the @ and " keys.
partial alphanumeric_keys
xkb_symbols "dvorak" {
include "us(dvorak-alt-intl)"
name[Group1]="English (UK, Dvorak)";
key <TLDE> { [ grave, notsign, bar, bar ] };
key <AE02> { [ 2, quotedbl, twosuperior, NoSymbol ] };
key <AE03> { [ 3, sterling, threesuperior, NoSymbol ] };
key <AD01> { [ apostrophe, at ] };
key <BKSL> { [ numbersign, asciitilde ] };
key <LSGT> { [ backslash, bar ] };
};
// Dvorak letter positions, but punctuation all in the normal UK positions.
partial alphanumeric_keys
xkb_symbols "dvorakukp" {
include "gb(dvorak)"
name[Group1]="English (UK, Dvorak, with UK punctuation)";
key <AE11> { [ minus, underscore ] };
key <AE12> { [ equal, plus ] };
key <AD11> { [ bracketleft, braceleft ] };
key <AD12> { [ bracketright, braceright ] };
key <AD01> { [ slash, question ] };
key <AC11> { [apostrophe, at, dead_circumflex, dead_caron] };
};
partial alphanumeric_keys
xkb_symbols "mac" {
include "latin"
name[Group1]= "English (UK, Macintosh)";
key <TLDE> { [ section, plusminus ] };
key <AE02> { [ 2, at, EuroSign ] };
key <AE03> { [ 3, sterling, numbersign ] };
key <LSGT> { [ grave, asciitilde ] };
include "level3(ralt_switch)"
include "level3(enter_switch)"
};
partial alphanumeric_keys
xkb_symbols "mac_intl" {
include "latin"
name[Group1]="English (UK, Macintosh, intl.)";
key <TLDE> { [ section, plusminus, notsign, notsign ] }; //dead_grave
key <AE02> { [ 2, at, EuroSign, onehalf ] };
key <AE03> { [ 3, sterling, twosuperior, onethird ] };
key <AE04> { [ 4, dollar, threesuperior, onequarter ] };
key <AE06> { [ 6, dead_circumflex, NoSymbol, onesixth ] };
key <AD09> { [ o, O, oe, OE ] };
key <AC11> { [ dead_acute, dead_diaeresis, dead_diaeresis, bar ] }; //dead_doubleacute
key <BKSL> { [ backslash, bar, numbersign, bar ] };
key <LSGT> { [ dead_grave, dead_tilde, brokenbar, bar ] };
include "level3(ralt_switch)"
};
partial alphanumeric_keys
xkb_symbols "pl" {
// Polish accented letters on upper levels of corresponding base letters.
// Idea from Wawrzyniec Niewodniczański, adapted by Aleksander Kowalski.
include "gb(basic)"
name[Group1]="Polish (British keyboard)";
key <AD03> { [ e, E, eogonek, Eogonek ] };
key <AD09> { [ o, O, oacute, Oacute ] };
key <AC01> { [ a, A, aogonek, Aogonek ] };
key <AC02> { [ s, S, sacute, Sacute ] };
key <AB01> { [ z, Z, zabovedot, Zabovedot ] };
key <AB02> { [ x, X, zacute, Zacute ] };
key <AB03> { [ c, C, cacute, Cacute ] };
key <AB06> { [ n, N, nacute, Nacute ] };
};
partial alphanumeric_keys
xkb_symbols "gla" {
// Grave-accented letters on the upper levels of the relevant vowels.
include "gb(basic)"
name[Group1]="Scottish Gaelic";
key <AD03> { [ e, E, egrave, Egrave ] };
key <AD07> { [ u, U, ugrave, Ugrave ] };
key <AD08> { [ i, I, igrave, Igrave ] };
key <AD09> { [ o, O, ograve, Ograve ] };
key <AC01> { [ a, A, agrave, Agrave ] };
};
// EXTRAS:
partial alphanumeric_keys
xkb_symbols "sun_type6" {
include "sun_vndr/gb(sun_type6)"
};
+102
View File
@@ -0,0 +1,102 @@
// The <KPDL> key is a mess.
// It was probably originally meant to be a decimal separator.
// Except since it was declared by USA people it didn't use the original
// SI separator "," but a "." (since then the USA managed to f-up the SI
// by making "." an accepted alternative, but standards still use "," as
// default)
// As a result users of SI-abiding countries expect either a "." or a ","
// or a "decimal_separator" which may or may not be translated in one of the
// above depending on applications.
// It's not possible to define a default per-country since user expectations
// depend on the conflicting choices of their most-used applications,
// operating system, etc. Therefore it needs to be a configuration setting
// Copyright © 2007 Nicolas Mailhot <nicolas.mailhot @ laposte.net>
// Legacy <KPDL> #1
// This assumes KP_Decimal will be translated in a dot
partial keypad_keys
xkb_symbols "dot" {
key.type[Group1]="KEYPAD" ;
key <KPDL> { [ KP_Delete, KP_Decimal ] }; // <delete> <separator>
};
// Legacy <KPDL> #2
// This assumes KP_Separator will be translated in a comma
partial keypad_keys
xkb_symbols "comma" {
key.type[Group1]="KEYPAD" ;
key <KPDL> { [ KP_Delete, KP_Separator ] }; // <delete> <separator>
};
// Period <KPDL>, usual keyboard serigraphy in most countries
partial keypad_keys
xkb_symbols "dotoss" {
key.type[Group1]="FOUR_LEVEL_MIXED_KEYPAD" ;
key <KPDL> { [ KP_Delete, period, comma, 0x100202F ] }; // <delete> . , ⍽ (narrow no-break space)
};
// Period <KPDL>, usual keyboard serigraphy in most countries, latin-9 restriction
partial keypad_keys
xkb_symbols "dotoss_latin9" {
key.type[Group1]="FOUR_LEVEL_MIXED_KEYPAD" ;
key <KPDL> { [ KP_Delete, period, comma, nobreakspace ] }; // <delete> . , ⍽ (no-break space)
};
// Comma <KPDL>, what most non anglo-saxon people consider the real separator
partial keypad_keys
xkb_symbols "commaoss" {
key.type[Group1]="FOUR_LEVEL_MIXED_KEYPAD" ;
key <KPDL> { [ KP_Delete, comma, period, 0x100202F ] }; // <delete> , . ⍽ (narrow no-break space)
};
// Momayyez <KPDL>: Bahrain, Iran, Iraq, Kuwait, Oman, Qatar, Saudi Arabia, Syria, UAE
partial keypad_keys
xkb_symbols "momayyezoss" {
key.type[Group1]="FOUR_LEVEL_MIXED_KEYPAD" ;
key <KPDL> { [ KP_Delete, 0x100066B, comma, 0x100202F ] }; // <delete> ? , ⍽ (narrow no-break space)
};
// Abstracted <KPDL>, pray everything will work out (it usually does not)
partial keypad_keys
xkb_symbols "kposs" {
key.type[Group1]="FOUR_LEVEL_MIXED_KEYPAD" ;
key <KPDL> { [ KP_Delete, KP_Decimal, KP_Separator, 0x100202F ] }; // <delete> ? ? ⍽ (narrow no-break space)
};
// Spreadsheets may be configured to use the dot as decimal
// punctuation, comma as a thousands separator and then semi-colon as
// the list separator. Of these, dot and semi-colon is most important
// when entering data by the keyboard; the comma can then be inferred
// and added to the presentation afterwards. Using semi-colon as a
// general separator may in fact be preferred to avoid ambiguities
// in data files. Most times a decimal separator is hard-coded, it
// seems to be period, probably since this is the syntax used in
// (most) programming languages.
partial keypad_keys
xkb_symbols "semi" {
key.type[Group1]="FOUR_LEVEL_MIXED_KEYPAD" ;
key <KPDL> { [ NoSymbol, NoSymbol, semicolon ] };
};
+255
View File
@@ -0,0 +1,255 @@
// Common Latin alphabet layout
default partial
xkb_symbols "basic" {
key <AE01> { [ 1, exclam, onesuperior, exclamdown ] };
key <AE02> { [ 2, at, twosuperior, oneeighth ] };
key <AE03> { [ 3, numbersign, threesuperior, sterling ] };
key <AE04> { [ 4, dollar, onequarter, dollar ] };
key <AE05> { [ 5, percent, onehalf, threeeighths ] };
key <AE06> { [ 6, asciicircum, threequarters, fiveeighths ] };
key <AE07> { [ 7, ampersand, braceleft, seveneighths ] };
key <AE08> { [ 8, asterisk, bracketleft, trademark ] };
key <AE09> { [ 9, parenleft, bracketright, plusminus ] };
key <AE10> { [ 0, parenright, braceright, degree ] };
key <AE11> { [ minus, underscore, backslash, questiondown ] };
key <AE12> { [ equal, plus, dead_cedilla, dead_ogonek ] };
key <AD01> { [ q, Q, at, Greek_OMEGA ] };
key <AD02> { [ w, W, U017F, section ] };
key <AD03> { [ e, E, e, E ] };
key <AD04> { [ r, R, paragraph, registered ] };
key <AD05> { [ t, T, tslash, Tslash ] };
key <AD06> { [ y, Y, leftarrow, yen ] };
key <AD07> { [ u, U, downarrow, uparrow ] };
key <AD08> { [ i, I, rightarrow, idotless ] };
key <AD09> { [ o, O, oslash, Oslash ] };
key <AD10> { [ p, P, thorn, THORN ] };
key <AD11> { [bracketleft, braceleft, dead_diaeresis, dead_abovering ] };
key <AD12> { [bracketright, braceright, dead_tilde, dead_macron ] };
key <AC01> { [ a, A, ae, AE ] };
key <AC02> { [ s, S, ssharp, U1E9E ] };
key <AC03> { [ d, D, eth, ETH ] };
key <AC04> { [ f, F, dstroke, ordfeminine ] };
key <AC05> { [ g, G, eng, ENG ] };
key <AC06> { [ h, H, hstroke, Hstroke ] };
key <AC07> { [ j, J, dead_hook, dead_horn ] };
key <AC08> { [ k, K, kra, ampersand ] };
key <AC09> { [ l, L, lstroke, Lstroke ] };
key <AC10> { [ semicolon, colon, dead_acute, dead_doubleacute ] };
key <AC11> { [apostrophe, quotedbl, dead_circumflex, dead_caron ] };
key <TLDE> { [ grave, asciitilde, notsign, notsign ] };
key <BKSL> { [ backslash, bar, dead_grave, dead_breve ] };
key <AB01> { [ z, Z, guillemotleft, less ] };
key <AB02> { [ x, X, guillemotright, greater ] };
key <AB03> { [ c, C, cent, copyright ] };
key <AB04> { [ v, V, doublelowquotemark, singlelowquotemark ] };
key <AB05> { [ b, B, leftdoublequotemark, leftsinglequotemark ] };
key <AB06> { [ n, N, rightdoublequotemark, rightsinglequotemark ] };
key <AB07> { [ m, M, mu, masculine ] };
key <AB08> { [ comma, less, U2022, multiply ] }; // bullet
key <AB09> { [ period, greater, periodcentered, division ] };
key <AB10> { [ slash, question, dead_belowdot, dead_abovedot ] };
};
// Northern Europe ( Danish, Finnish, Norwegian, Swedish) common layout
partial
xkb_symbols "type2" {
include "latin"
key <AE01> { [ 1, exclam, exclamdown, onesuperior ] };
key <AE02> { [ 2, quotedbl, at, twosuperior ] };
key <AE03> { [ 3, numbersign, sterling, threesuperior] };
key <AE04> { [ 4, currency, dollar, onequarter ] };
key <AE05> { [ 5, percent, onehalf, cent ] };
key <AE06> { [ 6, ampersand, yen, fiveeighths ] };
key <AE07> { [ 7, slash, braceleft, division ] };
key <AE08> { [ 8, parenleft, bracketleft, guillemotleft] };
key <AE09> { [ 9, parenright, bracketright, guillemotright] };
key <AE10> { [ 0, equal, braceright, degree ] };
key <AD03> { [ e, E, EuroSign, cent ] };
key <AD04> { [ r, R, registered, registered ] };
key <AD05> { [ t, T, thorn, THORN ] };
key <AD09> { [ o, O, oe, OE ] };
key <AD11> { [ aring, Aring, dead_diaeresis, dead_abovering ] };
key <AD12> { [dead_diaeresis, dead_circumflex, dead_tilde, dead_caron ] };
key <AC01> { [ a, A, ordfeminine, masculine ] };
key <AB03> { [ c, C, copyright, copyright ] };
key <AB08> { [ comma, semicolon, dead_cedilla, dead_ogonek ] };
key <AB09> { [ period, colon, periodcentered, dead_abovedot ] };
key <AB10> { [ minus, underscore, dead_belowdot, dead_abovedot ] };
};
// Slavic Latin ( Albanian, Croatian, Polish, Slovene, Yugoslav)
// common layout
partial
xkb_symbols "type3" {
include "latin"
key <AD01> { [ q, Q, backslash, Greek_OMEGA ] };
key <AD02> { [ w, W, bar, section ] };
key <AD06> { [ z, Z, leftarrow, yen ] };
key <AC04> { [ f, F, bracketleft, ordfeminine ] };
key <AC05> { [ g, G, bracketright, ENG ] };
key <AC08> { [ k, K, lstroke, ampersand ] };
key <AB01> { [ y, Y, guillemotleft, less ] };
key <AB04> { [ v, V, at, grave ] };
key <AB05> { [ b, B, braceleft, apostrophe ] };
key <AB06> { [ n, N, braceright, acute ] };
key <AB07> { [ m, M, section, masculine ] };
key <AB08> { [ comma, semicolon, less, multiply ] };
key <AB09> { [ period, colon, greater, division ] };
};
// Another common Latin layout
// (German, Estonian, Spanish, Icelandic, Italian, Latin American, Portuguese)
partial
xkb_symbols "type4" {
include "latin"
key <AE02> { [ 2, quotedbl, at, oneeighth ] };
key <AE06> { [ 6, ampersand, notsign, fiveeighths ] };
key <AE07> { [ 7, slash, braceleft, seveneighths ] };
key <AE08> { [ 8, parenleft, bracketleft, trademark ] };
key <AE09> { [ 9, parenright, bracketright, plusminus ] };
key <AE10> { [ 0, equal, braceright, degree ] };
key <AD03> { [ e, E, EuroSign, cent ] };
key <AB08> { [ comma, semicolon, U2022, multiply ] }; // bullet
key <AB09> { [ period, colon, periodcentered, division ] };
key <AB10> { [ minus, underscore, dead_belowdot, dead_abovedot ] };
};
partial
xkb_symbols "nodeadkeys" {
key <AE12> { [ equal, plus, cedilla, ogonek ] };
key <AD11> { [bracketleft, braceleft, diaeresis, degree ] };
key <AD12> { [bracketright, braceright, asciitilde, macron ] };
key <AC07> { [ j, J, ezh, EZH ] };
key <AC10> { [ semicolon, colon, acute, doubleacute ] };
key <AC11> { [apostrophe, quotedbl, asciicircum, caron ] };
key <BKSL> { [ backslash, bar, grave, breve ] };
key <AB10> { [ slash, question, ellipsis, abovedot ] };
};
partial
xkb_symbols "type2_nodeadkeys" {
include "latin(nodeadkeys)"
key <AD11> { [ aring, Aring, diaeresis, degree ] };
key <AD12> { [ diaeresis, asciicircum, asciitilde, caron ] };
key <AB08> { [ comma, semicolon, cedilla, ogonek ] };
key <AB09> { [ period, colon, periodcentered, abovedot ] };
key <AB10> { [ minus, underscore, ellipsis, abovedot ] };
};
partial
xkb_symbols "type3_nodeadkeys" {
include "latin(nodeadkeys)"
};
partial
xkb_symbols "type4_nodeadkeys" {
include "latin(nodeadkeys)"
key <AB10> { [ minus, underscore, ellipsis, abovedot ] };
};
// Added 2008.03.05 by Marcin Woliński
// See http://marcinwolinski.pl/keyboard/ for a description.
// Used by pl(intl)
//
// ┌─────┐
// │ 2 4 │ 2 = Shift, 4 = Level3 + Shift
// │ 1 3 │ 1 = Normal, 3 = Level3
// └─────┘
// ┌─────┬─────┬─────┬─────┬─────┬─────┬─────┬─────┬─────┬─────┬─────┬─────┬─────┲━━━━━━━━━┓
// │ ~ ~ │ ! ' │ @ " │ # ˝ │ $ ¸ │ % ˇ │ ^ ^ │ & ˘ │ * ̇ │ ( ̣ │ ) ° │ _ ¯ │ + ˛ ┃ ⌫ Back- ┃
// │ ` ` │ 1 ¡ │ 2 © │ 3 • │ 4 § │ 5 € │ 6 ¢ │ 7 − │ 8 × │ 9 ÷ │ 0 ° │ - – │ = — ┃ space ┃
// ┢━━━━━┷━┱───┴─┬───┴─┬───┴─┬───┴─┬───┴─┬───┴─┬───┴─┬───┴─┬───┴─┬───┴─┬───┴─┬───┺━┳━━━━━━━┫
// ┃ ┃ Q │ W │ E │ R │ T │ Y │ U │ I │ O │ P │ { « │ } » ┃ Enter ┃
// ┃Tab ↹ ┃ q │ w │ e │ r │ t │ y │ u │ i │ o │ p │ [ ‹ │ ] › ┃ ⏎ ┃
// ┣━━━━━━━┻┱────┴┬────┴┬────┴┬────┴┬────┴┬────┴┬────┴┬────┴┬────┴┬────┴┬────┴┬────┺┓ ┃
// ┃ ┃ A │ S │ D │ F │ G │ H │ J │ K │ L │ : “ │ " ” │ | ¶ ┃ ┃
// ┃Caps ⇬ ┃ a │ s │ d │ f │ g │ h │ j │ k │ l │ ; ‘ │ ' ’ │ \ ┃ ┃
// ┣━━━━━━━━┹────┬┴────┬┴────┬┴────┬┴────┬┴────┬┴────┬┴────┬┴────┬┴────┬┴────┲┷━━━━━┻━━━━━━┫
// ┃ │ Z │ X │ C │ V │ B │ N │ M │ < „ │ > · │ ? ¿ ┃ ┃
// ┃Shift ⇧ │ z │ x │ c │ v │ b │ n │ m │ , ‚ │ . … │ / ⁄ ┃Shift ⇧ ┃
// ┣━━━━━━━┳━━━━━┷━┳━━━┷━━━┱─┴─────┴─────┴─────┴─────┴─────┴───┲━┷━━━━━╈━━━━━┻━┳━━━━━━━┳━━━┛
// ┃ ┃ ┃ ┃ ␣ ⍽ ┃ ┃ ┃ ┃
// ┃Ctrl ┃Meta ┃Alt ┃ ␣ Space ⍽ ┃AltGr ⇮┃Menu ┃Ctrl ┃
// ┗━━━━━━━┻━━━━━━━┻━━━━━━━┹───────────────────────────────────┺━━━━━━━┻━━━━━━━┻━━━━━━━┛
partial
xkb_symbols "intl" {
key <TLDE> { [ grave, asciitilde, dead_grave, dead_tilde ] };
key <AE01> { [ 1, exclam, exclamdown, dead_acute ] };
key <AE02> { [ 2, at, copyright, dead_diaeresis ] };
key <AE03> { [ 3, numbersign, U2022, dead_doubleacute ] }; // U+2022 is bullet (the name bullet does not work)
key <AE04> { [ 4, dollar, section, dead_cedilla ] };
key <AE05> { [ 5, percent, EuroSign, dead_caron ] };
key <AE06> { [ 6, asciicircum, cent, dead_circumflex ] };
key <AE07> { [ 7, ampersand, U2212, dead_breve ] }; // U+2212 is MINUS SIGN
key <AE08> { [ 8, asterisk, multiply, dead_abovedot ] };
key <AE09> { [ 9, parenleft, division, dead_belowdot ] };
key <AE10> { [ 0, parenright, degree, dead_abovering ] };
key <AE11> { [ minus, underscore, endash, dead_macron ] };
key <AE12> { [ equal, plus, emdash, dead_ogonek ] };
key <AD01> { [ q, Q ] };
key <AD02> { [ w, W ] };
key <AD03> { [ e, E ] };
key <AD04> { [ r, R ] };
key <AD05> { [ t, T ] };
key <AD06> { [ y, Y ] };
key <AD07> { [ u, U ] };
key <AD08> { [ i, I ] };
key <AD09> { [ o, O ] };
key <AD10> { [ p, P ] };
key <AD11> { [bracketleft, braceleft, U2039, guillemotleft ] };
key <AD12> { [bracketright, braceright, U203A, guillemotright ] };
key <AC01> { [ a, A ] };
key <AC02> { [ s, S ] };
key <AC03> { [ d, D ] };
key <AC04> { [ f, F ] };
key <AC05> { [ g, G ] };
key <AC06> { [ h, H ] };
key <AC07> { [ j, J ] };
key <AC08> { [ k, K ] };
key <AC09> { [ l, L ] };
key <AC10> { [ semicolon, colon, leftsinglequotemark, leftdoublequotemark ] };
key <AC11> { [apostrophe, quotedbl, rightsinglequotemark, rightdoublequotemark ] };
key <BKSL> { [ backslash, bar, NoSymbol, paragraph ] };
key <AB01> { [ z, Z ] };
key <AB02> { [ x, X ] };
key <AB03> { [ c, C ] };
key <AB04> { [ v, V ] };
key <AB05> { [ b, B ] };
key <AB06> { [ n, N ] };
key <AB07> { [ m, M ] };
key <AB08> { [ comma, less, singlelowquotemark, doublelowquotemark ] };
key <AB09> { [ period, greater, ellipsis, periodcentered ] };
key <AB10> { [ slash, question, U2044, questiondown ] }; // U+2044 is FRACTION SLASH
};
+156
View File
@@ -0,0 +1,156 @@
// These variants assign ISO_Level3_Shift to various keys
// so that levels 3 and 4 can be reached.
// The default behaviour:
// the right Alt key (AltGr) chooses the third symbol engraved on a key.
default partial modifier_keys
xkb_symbols "ralt_switch" {
key <RALT> {[ ISO_Level3_Shift ], type[group1]="ONE_LEVEL" };
};
// The right Alt key never chooses the third level.
// This option attempts to undo the effect of a layout's inclusion of
// 'ralt_switch'. You may want to also select another level3 option
// to map the level3 shift to some other key.
partial modifier_keys
xkb_symbols "ralt_alt" {
key <RALT> {[ Alt_R, Meta_R ], type[group1]="TWO_LEVEL" };
modifier_map Mod1 { <RALT> };
};
// The right Alt key (while pressed) chooses the third shift level,
// and Compose is mapped to its second level.
partial modifier_keys
xkb_symbols "ralt_switch_multikey" {
key <RALT> {[ ISO_Level3_Shift, Multi_key ], type[group1]="TWO_LEVEL" };
};
// Either Alt key (while pressed) chooses the third shift level.
// (To be used mostly to imitate Mac OS functionality.)
partial modifier_keys
xkb_symbols "alt_switch" {
include "level3(lalt_switch)"
include "level3(ralt_switch)"
};
// The left Alt key (while pressed) chooses the third shift level.
partial modifier_keys
xkb_symbols "lalt_switch" {
key <LALT> {[ ISO_Level3_Shift ], type[group1]="ONE_LEVEL" };
};
// The right Ctrl key (while pressed) chooses the third shift level.
partial modifier_keys
xkb_symbols "switch" {
key <RCTL> {[ ISO_Level3_Shift ], type[group1]="ONE_LEVEL" };
};
// The Menu key (while pressed) chooses the third shift level.
partial modifier_keys
xkb_symbols "menu_switch" {
key <MENU> {[ ISO_Level3_Shift ], type[group1]="ONE_LEVEL" };
};
// Either Win key (while pressed) chooses the third shift level.
partial modifier_keys
xkb_symbols "win_switch" {
include "level3(lwin_switch)"
include "level3(rwin_switch)"
};
// The left Win key (while pressed) chooses the third shift level.
partial modifier_keys
xkb_symbols "lwin_switch" {
key <LWIN> {[ ISO_Level3_Shift ], type[group1]="ONE_LEVEL" };
};
// The right Win key (while pressed) chooses the third shift level.
partial modifier_keys
xkb_symbols "rwin_switch" {
key <RWIN> {[ ISO_Level3_Shift ], type[group1]="ONE_LEVEL" };
};
// The Enter key on the kepypad (while pressed) chooses the third shift level.
// (This is especially useful for Mac laptops which miss the right Alt key.)
partial modifier_keys
xkb_symbols "enter_switch" {
key <KPEN> {[ ISO_Level3_Shift ], type[group1]="ONE_LEVEL" };
};
// The CapsLock key (while pressed) chooses the third shift level.
partial modifier_keys
xkb_symbols "caps_switch" {
key <CAPS> {[ ISO_Level3_Shift ], type[group1]="ONE_LEVEL" };
};
// The CapsLock key (while pressed) chooses the third shift level and
// Ctrl + CapsLock has the original CapsLock function.
// The 2023 DIN standard for German keyboards recommends it as an option:
// - https://de.wikipedia.org/wiki/E1_(Tastaturbelegung)#Feststelltaste/Umschaltsperre
// - https://en.wikipedia.org/wiki/Caps_Lock#Abolition
partial modifier_keys
xkb_symbols "caps_switch_capslock_with_ctrl" {
virtual_modifiers LevelThree;
key <CAPS> {
type[Group1] = "PC_CONTROL_LEVEL2",
symbols[Group1] = [ ISO_Level3_Shift, Caps_Lock ],
// Explicit actions are preferred over modMap None/Mod5 { Caps_Lock }
// because they have no side effect
actions[Group1] = [ SetMods(modifiers = LevelThree), LockMods(modifiers = Lock) ]
};
};
// The Backslash key (while pressed) chooses the third shift level.
partial modifier_keys
xkb_symbols "bksl_switch" {
key <BKSL> {[ ISO_Level3_Shift ], type[group1]="ONE_LEVEL" };
};
// The AC11 key (while pressed) chooses the third shift level.
partial modifier_keys
xkb_symbols "ac11_switch" {
key <AC11> {[ ISO_Level3_Shift ], type[Group1]="ONE_LEVEL" };
};
// The Less/Greater key (while pressed) chooses the third shift level.
partial modifier_keys
xkb_symbols "lsgt_switch" {
key <LSGT> {[ ISO_Level3_Shift ], type[group1]="ONE_LEVEL" };
};
// The CapsLock key (while pressed) chooses the third shift level,
// and latches when pressed together with another third-level chooser.
partial modifier_keys
xkb_symbols "caps_switch_latch" {
key <CAPS> {[ ISO_Level3_Shift, ISO_Level3_Shift, ISO_Level3_Latch ],
type[group1]="THREE_LEVEL" };
};
// The Backslash key (while pressed) chooses the third shift level,
// and latches when pressed together with another third-level chooser.
partial modifier_keys
xkb_symbols "bksl_switch_latch" {
key <BKSL> {[ ISO_Level3_Shift, ISO_Level3_Shift, ISO_Level3_Latch ],
type[group1]="THREE_LEVEL" };
};
// The Less/Greater key (while pressed) chooses the third shift level,
// and latches when pressed together with another third-level chooser.
partial modifier_keys
xkb_symbols "lsgt_switch_latch" {
key <LSGT> {[ ISO_Level3_Shift, ISO_Level3_Shift, ISO_Level3_Latch ],
type[group1]="THREE_LEVEL" };
};
// Top-row digit key 4 chooses third shift level when pressed alone.
partial modifier_keys
xkb_symbols "4_switch_isolated" {
override key <AE04> {[ ISO_Level3_Shift ]};
};
// Top-row digit key 9 chooses third shift level when pressed alone.
partial modifier_keys
xkb_symbols "9_switch_isolated" {
override key <AE09> {[ ISO_Level3_Shift ]};
};
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,153 @@
//! xkeyboard-config — keyboard layouts, compiled from the X11 xkeyboard-config database
//! into native Zig. It turns a physical key (a USB HID usage, as the input module delivers)
//! plus a modifier state into a **keysym** and, when the key produces one, a **character**
//! (a Unicode scalar). This is the piece that lets a `KeyEvent.keycode` become a
//! `KeyEvent.character`, without shipping an X11 runtime.
//!
//! The layout tables in `generated/layouts.zig` are produced by
//! `tools/make-xkeyboard-config.py` (see ./README.md to regenerate). Those tables are
//! deliberately pure data — each key carries its up-to-four levels and an XKB *type*. The
//! type -> level selection semantics (which modifier picks which level) live here, so the
//! data and the policy are separable.
//!
//! Scope (documented in README.md): group 1 only, no dead-key/compose composition (a dead
//! key returns its keysym with no character), and a curated set of key types. Layouts:
//! us, gb, de, fr, es, dvorak.
//!
//! Upstream xkeyboard-config and keysymdef.h are MIT/X11 licensed; see vendor/COPYING and
//! vendor/PROVENANCE.md.
const std = @import("std");
const generated = @import("layouts");
pub const Level = generated.Level;
pub const KeyType = generated.KeyType;
pub const Key = generated.Key;
pub const Layout = generated.Layout;
/// The generated layouts, by name — as pointers, so they share identity with `all` and
/// `byName` (and match the `*const Layout` that `map` takes).
pub const us: *const Layout = &generated.us;
pub const gb: *const Layout = &generated.gb;
pub const de: *const Layout = &generated.de;
pub const fr: *const Layout = &generated.fr;
pub const es: *const Layout = &generated.es;
pub const dvorak: *const Layout = &generated.dvorak;
/// Every generated layout, for enumeration (e.g. a settings UI).
pub const all = generated.all;
/// The modifier state that selects a key's level. `level3` is AltGr (ISO Level3 Shift);
/// `control` is accepted for completeness but does not affect level selection here.
pub const Modifiers = struct {
shift: bool = false,
caps_lock: bool = false,
level3: bool = false,
control: bool = false,
};
/// The result of a lookup: the X11 `keysym`, and the `character` it produces (a Unicode
/// scalar) when it is a printable key — null for keys that produce none (Return, F1, a
/// bare dead key, an unmapped key).
pub const Mapping = struct {
keysym: u32,
character: ?u21,
};
/// Which level (0..3) a key of `kind` selects under `mods`. XKB's canonical semantics:
/// Shift picks the odd level, AltGr (level3) adds 2, and Caps acts like Shift for the
/// alphabetic types. See the XKB "key types" — this covers the ones the vendored layouts
/// use; anything else falls back to shift-or-not.
fn selectLevel(kind: KeyType, mods: Modifiers) usize {
const shift_or_caps = mods.shift != mods.caps_lock; // XOR: Caps behaves like Shift
const low: usize = if (mods.shift) 1 else 0;
const high: usize = if (mods.level3) 2 else 0;
return switch (kind) {
.one_level => 0,
.two_level, .keypad, .other => low,
.alphabetic => if (shift_or_caps) 1 else 0,
.four_level => low + high,
.four_level_alphabetic => (if (shift_or_caps) @as(usize, 1) else 0) + high,
// Caps affects only the base pair, not the AltGr pair.
.four_level_semialphabetic => if (mods.level3) 2 + low else (if (shift_or_caps) @as(usize, 1) else 0),
};
}
/// Map a physical key (`hid_usage`, a USB HID keyboard-page usage) under `mods` on
/// `layout` to its keysym and character. Falls back gracefully when the selected level is
/// undefined for the key: it drops the AltGr component, then the shift component, so a key
/// with only a base/shift pair still yields something sensible under AltGr.
pub fn map(layout: *const Layout, hid_usage: u8, mods: Modifiers) Mapping {
const key = &layout.keys[hid_usage];
var level = selectLevel(key.kind, mods);
// Fall back to a defined level: full -> without AltGr -> base.
if (key.levels[level].keysym == 0 and key.levels[level].unicode == 0) {
const candidates = [_]usize{ level & 1, 0 };
for (candidates) |candidate| {
if (key.levels[candidate].keysym != 0 or key.levels[candidate].unicode != 0) {
level = candidate;
break;
}
}
}
const chosen = key.levels[level];
return .{
.keysym = chosen.keysym,
.character = if (chosen.unicode != 0) @intCast(chosen.unicode) else null,
};
}
/// Look up a layout by its name (`"us"`, `"gb"`, ...), or null if unknown.
pub fn byName(name: []const u8) ?*const Layout {
for (all) |layout| {
if (std.mem.eql(u8, layout.name, name)) return layout;
}
return null;
}
// --- tests (host-run via `zig build test`) ---------------------------------
const testing = std.testing;
// USB HID usages used in the tests (keyboard page 0x07).
const hid_a: u8 = 0x04;
const hid_1: u8 = 0x1e;
const hid_3: u8 = 0x20;
test "us: letters obey shift and caps" {
try testing.expectEqual(@as(?u21, 'a'), map(us, hid_a, .{}).character);
try testing.expectEqual(@as(?u21, 'A'), map(us, hid_a, .{ .shift = true }).character);
try testing.expectEqual(@as(?u21, 'A'), map(us, hid_a, .{ .caps_lock = true }).character);
// Shift + Caps cancels for an alphabetic key.
try testing.expectEqual(@as(?u21, 'a'), map(us, hid_a, .{ .shift = true, .caps_lock = true }).character);
}
test "us: digits and their shifted symbols" {
try testing.expectEqual(@as(?u21, '1'), map(us, hid_1, .{}).character);
try testing.expectEqual(@as(?u21, '!'), map(us, hid_1, .{ .shift = true }).character);
try testing.expectEqual(@as(?u21, '3'), map(us, hid_3, .{}).character);
try testing.expectEqual(@as(?u21, '#'), map(us, hid_3, .{ .shift = true }).character);
// A digit is not alphabetic: Caps alone must not shift it.
try testing.expectEqual(@as(?u21, '3'), map(us, hid_3, .{ .caps_lock = true }).character);
}
test "layouts differ: GB pound vs US hash on shift+3" {
try testing.expectEqual(@as(?u21, '#'), map(us, hid_3, .{ .shift = true }).character);
try testing.expectEqual(@as(?u21, '£'), map(gb, hid_3, .{ .shift = true }).character);
}
test "french azerty places q where us has a" {
try testing.expectEqual(@as(?u21, 'q'), map(fr, hid_a, .{}).character);
try testing.expectEqual(@as(?u21, 'Q'), map(fr, hid_a, .{ .shift = true }).character);
}
test "byName resolves and rejects" {
try testing.expect(byName("us") == us);
try testing.expect(byName("gb") == gb);
try testing.expect(byName("nonsense") == null);
}
test "unmapped key yields no character" {
// HID 0x00 is not a key; every level is empty.
try testing.expectEqual(@as(?u21, null), map(us, 0x00, .{}).character);
}
+105 -5
View File
@@ -43,16 +43,41 @@ pub const SystemCall = enum(u64) {
irq_bind = 14, // irq_bind(id, resource_index, endpoint): deliver a device IRQ as an IPC notification
irq_ack = 15, // irq_ack(id, resource_index): re-arm a bound IRQ after servicing it
device_register = 16, // device_register(parent_id, descriptor) -> id: publish a child of a device you claimed
system_spawn = 17, // system_spawn(name_ptr, name_len) -> 0: start a named initial-ramdisk binary as a new ring-3 process
system_spawn = 17, // system_spawn(name_ptr, name_len, arguments_ptr, arguments_len, exit_endpoint) -> child process id: start a named initial-ramdisk binary as a new ring-3 process
dma_alloc = 18, // dma_alloc(len, flags) -> vaddr (rax), paddr (rdx): contiguous, pinned, uncacheable DMA memory
dma_free = 19, // dma_free(vaddr, len) -> 0: release a prior dma_alloc
msi_bind = 20, // msi_bind(device_id, endpoint) -> address (rax), data (rdx): a per-device MSI vector for a claimed device
io_read = 21, // io_read(device_id, resource_index, offset, width) -> value: read a port in a claimed device's io_port resource
io_write = 22, // io_write(device_id, resource_index, offset, width, value) -> 0: write a port in a claimed device's io_port resource
clock = 23, // clock() -> nanoseconds since boot: a monotonic time source (for timeouts/delays)
process_enumerate = 24, // process_enumerate(buffer, maximum) -> total: snapshot the task table
process_kill = 25, // process_kill(id) -> 0/-errno: end a process this process spawned
ipc_send = 26, // ipc_send(handle, message_ptr, message_len) -> 0/-errno: post a payload to an endpoint's async queue without blocking
process_exit_reason = 27, // process_exit_reason(id) -> ExitReason/-errno: how a dead child ended (its supervisor only)
process_subscribe = 28, // process_subscribe(endpoint) -> 0/-errno: subscribe to published exit events — every death posts a notification
signal_bind = 29, // signal_bind(endpoint) -> 0/-errno: nominate the endpoint this process's signals arrive on
process_signal = 30, // process_signal(id, signal) -> 0/-errno: post a signal to a child (or to yourself)
timer_bind = 31, // timer_bind(endpoint, ms) -> 0/-errno: one-shot timer — posts a notification when ms elapse
_,
};
/// How a process ended — recorded by the kernel at death, queried by the
/// supervisor with `process_exit_reason`, and the input to its restart decision
/// (docs/process-lifecycle.md): a clean exit meant to stop, a fault wants a
/// restart with backoff, killed means the supervisor did it itself. The faults
/// mirror the CPU exceptions a ring-3 process can die of; they are exit reasons,
/// never delivered to the faulting process (recovery is restart, not a handler).
pub const ExitReason = enum(u8) {
exited = 0, // returned from main / called exit
aborted = 1, // deliberate self-termination (reserved: no abort path yet)
segmentation_fault = 2, // page fault
illegal_instruction = 3, // invalid opcode
arithmetic_fault = 4, // divide error, x87 or SIMD fault
protection_fault = 5, // general protection fault
fault = 6, // any other CPU exception
killed = 7, // process_kill
};
/// The x86 MSI message address base (`0xFEE0_0000`): a device raises an MSI by writing
/// `data` to this address, which the Local APIC turns into an interrupt at the vector
/// in `data`. The kernel returns the concrete (address, data) from `msi_bind`; this is
@@ -67,17 +92,92 @@ pub const dma_write_combining: u64 = 2; // write-combining (framebuffers); needs
pub const dma_below_4g: u64 = 4; // physical address must fit 32 bits (legacy DMA engines)
/// Set in the badge returned by `ipc_reply_wait` when what arrived is an
/// **asynchronous notification** (today: a device interrupt bound with `irq_bind`)
/// rather than a message from a client. There is no payload and no reply owed; the
/// low bits carry the source, a GSI. Shared so the kernel's ISR and the driver's
/// event loop can't disagree about which bit means "the hardware spoke".
/// **asynchronous notification** (a device interrupt bound with `irq_bind`, or a
/// child-exit notice — see `notify_exit_bit`) rather than a message from a client.
/// There is no payload and no reply owed; the low bits carry the source. Shared so
/// the kernel's ISR and the driver's event loop can't disagree about which bit
/// means "the hardware spoke".
pub const notify_badge_bit: u64 = 1 << 63;
/// Set (alongside `notify_badge_bit`) in the badge of a **child-exit notification**:
/// posted to the endpoint a supervisor passed to `system_spawn` when that child ends
/// — by clean exit, by a fault, or by `process_kill`. The low bits carry the child's
/// process id, so one endpoint can supervise many children (and even share with IRQ
/// notifications, which never set this bit). The microkernel's SIGCHLD.
pub const notify_exit_bit: u64 = 1 << 62;
/// Set (alongside `notify_badge_bit`) in the badge of a **buffered message** — a payload
/// posted to an endpoint's async queue by `ipc_send`, delivered through `ipc_reply_wait`
/// like a notification (no reply owed) but carrying bytes in the receive buffer, not just
/// a badge. This is what distinguishes a payload-bearing async message from a bare IRQ /
/// child-exit notification (which sets neither this nor `notify_exit_bit`). The low bits
/// carry the sender's task id. The async counterpart of the synchronous `ipc_call`, for
/// broadcasts where a rendezvous is the wrong shape (the input service is the first user).
pub const notify_message_bit: u64 = 1 << 61;
/// Set (alongside `notify_badge_bit`) in the badge of a **signal notification** —
/// the process-lifecycle vocabulary of docs/process-lifecycle.md, delivered to the
/// endpoint the process nominated with `signal_bind`. The low bits carry the
/// coalesced pending mask (bit positions = `Signal` values): signals are
/// statements, not questions, and two pending terminates are one terminate.
pub const notify_signal_bit: u64 = 1 << 60;
/// Set (alongside `notify_badge_bit`) in the badge of a **timer notification** —
/// a one-shot `timer_bind` deadline landing. No payload bits: what to do when the
/// deadline fires is whatever the receiver armed it for (a stop-sequence
/// escalation, a restart backoff, an alarm).
pub const notify_timer_bit: u64 = 1 << 59;
/// The signal vocabulary (docs/process-lifecycle.md): POSIX's concepts, danos's
/// names, message delivery. The value is the bit position in the pending mask — a
/// private kernel/runtime detail, free to change while they ship together. Kill
/// is not here (it is `process_kill`, unhandleable by definition); faults are not
/// here (they are `ExitReason`s — recovery is restart, not a handler); liveness is
/// not here (a question, asked as the zero-length ping call, not a statement).
pub const Signal = enum(u5) {
terminate = 0, // finish up and exit (the polite half of the stop sequence)
reload = 1, // re-read configuration / re-scan
interrupt = 2, // interactive interrupt (no sender until a console exists)
quit = 3, // as interrupt, by convention more final
alarm = 4, // a timer the process armed for itself (unbuilt: no consumer yet)
user_1 = 5, // service-defined
user_2 = 6, // service-defined
};
/// Capacity of `ProcessDescriptor.name` — matches the longest name `system_spawn`
/// accepts, so a process's recorded name (its argv[0]) is never truncated.
pub const maximum_process_name = 64;
/// What a process is doing right now, as reported by `process_enumerate`. Crosses
/// the system_call boundary as `ProcessDescriptor.state`.
pub const ProcessState = enum(u32) {
ready = 0, // runnable, waiting for a core
running = 1, // executing on a core right now
blocked = 2, // waiting (sleeping, or blocked in IPC)
};
/// One `process_enumerate` entry — the kernel's view of a live task, kernel tasks
/// included (they carry an empty name and id 0 is the boot task). Fixed layout
/// (extern) because it crosses the kernel↔user boundary by memory copy, like
/// `DeviceDescriptor` in the device ABI.
pub const ProcessDescriptor = extern struct {
id: u32, // kernel-assigned process id; never reused (monotonic)
supervisor: u32, // id of the process that spawned it (0 = the kernel)
state: u32, // a ProcessState value
priority: u32,
name_length: u32,
name: [maximum_process_name]u8, // argv[0] at spawn; empty for kernel tasks
};
/// Well-known IPC service ids for the bootstrap name registry (create_ipc_endpoint +
/// ipc_register/ipc_lookup). Small integers, so no string interning is needed
/// during bring-up. The VFS server registers under `vfs`; clients look it up.
pub const ServiceId = enum(u32) {
vfs = 1,
input = 2,
ps2_bus = 3, // the 8042 owner; child device drivers attach here for raw bytes
device_manager = 4, // the tree, the matcher, the supervisor (docs/device-manager.md)
power = 5, // system power: events (button, lid, battery) + shutdown (docs/m21-plan.md; domain-named per decision 7 — the acpi service registers it on x86, a PSCI service will on ARM)
_,
};
+4 -4
View File
@@ -32,10 +32,10 @@ pub const PixelFormat = enum(u32) {
/// Output Protocol (a headless server, say). The kernel must treat on-screen
/// output as optional and never assume a framebuffer exists.
pub const Framebuffer = extern struct {
base: usize, // the memory address where pixel data starts (0 = none)
width: u32, // visible pixels per row (e.g. 1920)
height: u32, // visible rows (e.g. 1080)
pitch: u32, // bytes from the start of one row to the start of the next
base: usize, // the memory address where pixel data starts (0 = none)
width: u32, // visible pixels per row (e.g. 1920)
height: u32, // visible rows (e.g. 1080)
pitch: u32, // bytes from the start of one row to the start of the next
format: PixelFormat,
/// Whether a usable framebuffer was handed over.
+116 -46
View File
@@ -1,60 +1,120 @@
//! ACPI / PnP hardware-ID (`_HID`) names: the flat analog of pci-class.zig for
//! `acpi_device` nodes. Unlike PCI, ACPI has no class/subclass/prog-IF taxonomy — a
//! device's identity *is* its `_HID` string (`PNP0303` simply means "PS/2 keyboard"),
//! so this is a plain id -> description registry rather than a hierarchical decoder.
//! so this is a plain id <-> name registry rather than a hierarchical decoder.
//! The well-known PnP/ACPI IDs; vendor-specific ids (e.g. `QEMU0002`, `INTC1234`) have
//! no standard name and return "". Pure reference data, so it is shared by kernel
//! discovery (the device-tree dump) and any user-space tool.
//! no standard name and decode to nothing. Pure reference data, so it is shared by
//! kernel discovery (the device-tree dump) and any user-space driver or tool.
//!
//! Code that means a specific device names the `HardwareId` variant instead of its
//! `_HID` string — `HardwareId.ps2_keyboard.hid()` reads without a registry lookup,
//! where a bare `"PNP0303"` does not.
const std = @import("std");
const Entry = struct { hid: []const u8, name: []const u8 };
/// The common standard PnP/ACPI hardware IDs, as named values. Prefix ranges hint at
/// the grouping (PNP03xx keyboards, PNP0Fxx pointing devices, PNP0Cxx ACPI
/// power/thermal, PNP0Axx buses), but there is no formal hierarchy — hence a flat
/// enum over a flat registry.
pub const HardwareId = enum {
programmable_interrupt_controller,
system_timer,
high_precision_event_timer,
dma_controller,
ps2_keyboard,
parallel_port,
ecp_parallel_port,
serial_port,
floppy_disk_controller,
system_speaker,
pci_bus,
generic_container,
/// The second id the ACPI spec assigns the same "Generic Container Device" name.
generic_container_extended,
pci_express_root_bridge,
real_time_clock,
system_board,
motherboard_reserved_resources,
math_coprocessor,
acpi_system_board,
embedded_controller,
control_method_battery,
fan,
power_button,
lid,
sleep_button,
pci_interrupt_link,
microsoft_ps2_mouse,
ps2_mouse,
ac_adapter,
processor_device,
processor_aggregator,
processor_container,
/// The common standard PnP/ACPI hardware IDs. Prefix ranges hint at the grouping
/// (PNP03xx keyboards, PNP0Fxx pointing devices, PNP0Cxx ACPI power/thermal,
/// PNP0Axx buses), but there is no formal hierarchy — hence a flat table.
const table = [_]Entry{
.{ .hid = "PNP0000", .name = "Programmable Interrupt Controller (PIC)" },
.{ .hid = "PNP0100", .name = "System Timer (PIT)" },
.{ .hid = "PNP0103", .name = "High Precision Event Timer (HPET)" },
.{ .hid = "PNP0200", .name = "DMA Controller" },
.{ .hid = "PNP0303", .name = "PS/2 Keyboard" },
.{ .hid = "PNP0400", .name = "Standard LPT Parallel Port" },
.{ .hid = "PNP0401", .name = "ECP Parallel Port" },
.{ .hid = "PNP0501", .name = "16550A-compatible Serial Port" },
.{ .hid = "PNP0700", .name = "PC Floppy Disk Controller" },
.{ .hid = "PNP0800", .name = "System Speaker" },
.{ .hid = "PNP0A03", .name = "PCI Bus" },
.{ .hid = "PNP0A05", .name = "Generic Container Device" },
.{ .hid = "PNP0A06", .name = "Generic Container Device" },
.{ .hid = "PNP0A08", .name = "PCI Express Root Bridge" },
.{ .hid = "PNP0B00", .name = "Real-Time Clock (RTC)" },
.{ .hid = "PNP0C01", .name = "System Board" },
.{ .hid = "PNP0C02", .name = "Motherboard Reserved Resources" },
.{ .hid = "PNP0C04", .name = "Math Coprocessor" },
.{ .hid = "PNP0C08", .name = "ACPI System Board" },
.{ .hid = "PNP0C09", .name = "ACPI Embedded Controller" },
.{ .hid = "PNP0C0A", .name = "ACPI Control Method Battery" },
.{ .hid = "PNP0C0B", .name = "ACPI Fan" },
.{ .hid = "PNP0C0C", .name = "ACPI Power Button" },
.{ .hid = "PNP0C0D", .name = "ACPI Lid" },
.{ .hid = "PNP0C0E", .name = "ACPI Sleep Button" },
.{ .hid = "PNP0C0F", .name = "PCI Interrupt Link Device" },
.{ .hid = "PNP0F03", .name = "Microsoft PS/2 Mouse" },
.{ .hid = "PNP0F13", .name = "PS/2 Mouse" },
.{ .hid = "ACPI0003", .name = "AC Adapter" },
.{ .hid = "ACPI0007", .name = "Processor Device" },
.{ .hid = "ACPI000C", .name = "Processor Aggregator" },
.{ .hid = "ACPI0010", .name = "Processor Container" },
const Entry = struct { hid: []const u8, name: []const u8 };
/// The registry row for this id: its `_HID` string and human-readable name.
fn entry(self: HardwareId) Entry {
return switch (self) {
.programmable_interrupt_controller => .{ .hid = "PNP0000", .name = "Programmable Interrupt Controller (PIC)" },
.system_timer => .{ .hid = "PNP0100", .name = "System Timer (PIT)" },
.high_precision_event_timer => .{ .hid = "PNP0103", .name = "High Precision Event Timer (HPET)" },
.dma_controller => .{ .hid = "PNP0200", .name = "DMA Controller" },
.ps2_keyboard => .{ .hid = "PNP0303", .name = "PS/2 Keyboard" },
.parallel_port => .{ .hid = "PNP0400", .name = "Standard LPT Parallel Port" },
.ecp_parallel_port => .{ .hid = "PNP0401", .name = "ECP Parallel Port" },
.serial_port => .{ .hid = "PNP0501", .name = "16550A-compatible Serial Port" },
.floppy_disk_controller => .{ .hid = "PNP0700", .name = "PC Floppy Disk Controller" },
.system_speaker => .{ .hid = "PNP0800", .name = "System Speaker" },
.pci_bus => .{ .hid = "PNP0A03", .name = "PCI Bus" },
.generic_container => .{ .hid = "PNP0A05", .name = "Generic Container Device" },
.generic_container_extended => .{ .hid = "PNP0A06", .name = "Generic Container Device" },
.pci_express_root_bridge => .{ .hid = "PNP0A08", .name = "PCI Express Root Bridge" },
.real_time_clock => .{ .hid = "PNP0B00", .name = "Real-Time Clock (RTC)" },
.system_board => .{ .hid = "PNP0C01", .name = "System Board" },
.motherboard_reserved_resources => .{ .hid = "PNP0C02", .name = "Motherboard Reserved Resources" },
.math_coprocessor => .{ .hid = "PNP0C04", .name = "Math Coprocessor" },
.acpi_system_board => .{ .hid = "PNP0C08", .name = "ACPI System Board" },
.embedded_controller => .{ .hid = "PNP0C09", .name = "ACPI Embedded Controller" },
.control_method_battery => .{ .hid = "PNP0C0A", .name = "ACPI Control Method Battery" },
.fan => .{ .hid = "PNP0C0B", .name = "ACPI Fan" },
.power_button => .{ .hid = "PNP0C0C", .name = "ACPI Power Button" },
.lid => .{ .hid = "PNP0C0D", .name = "ACPI Lid" },
.sleep_button => .{ .hid = "PNP0C0E", .name = "ACPI Sleep Button" },
.pci_interrupt_link => .{ .hid = "PNP0C0F", .name = "PCI Interrupt Link Device" },
.microsoft_ps2_mouse => .{ .hid = "PNP0F03", .name = "Microsoft PS/2 Mouse" },
.ps2_mouse => .{ .hid = "PNP0F13", .name = "PS/2 Mouse" },
.ac_adapter => .{ .hid = "ACPI0003", .name = "AC Adapter" },
.processor_device => .{ .hid = "ACPI0007", .name = "Processor Device" },
.processor_aggregator => .{ .hid = "ACPI000C", .name = "Processor Aggregator" },
.processor_container => .{ .hid = "ACPI0010", .name = "Processor Container" },
};
}
/// This id's `_HID` string (e.g. `.ps2_keyboard` -> "PNP0303").
pub fn hid(self: HardwareId) []const u8 {
return self.entry().hid;
}
/// This id's human-readable name (e.g. `.ps2_keyboard` -> "PS/2 Keyboard").
pub fn description(self: HardwareId) []const u8 {
return self.entry().name;
}
/// The named value for a `_HID` string, or null if it is not a known standard
/// id (vendor-specific ids are not in the registry).
pub fn fromHid(hid_string: []const u8) ?HardwareId {
for (std.enums.values(HardwareId)) |id| {
if (std.mem.eql(u8, id.hid(), hid_string)) return id;
}
return null;
}
};
/// The human-readable name for a `_HID`, or "" if it is not a known standard id
/// (vendor-specific ids have no registry name — callers just print the raw HID).
/// The human-readable name for a `_HID` string, or "" if it is not a known standard
/// id (vendor-specific ids have no registry name — callers just print the raw HID).
pub fn description(hid: []const u8) []const u8 {
for (table) |entry| {
if (std.mem.eql(u8, entry.hid, hid)) return entry.name;
}
return "";
return (HardwareId.fromHid(hid) orelse return "").description();
}
test "decodes standard PnP/ACPI ids and leaves the rest alone" {
@@ -66,3 +126,13 @@ test "decodes standard PnP/ACPI ids and leaves the rest alone" {
try eq("", description("QEMU0002")); // vendor-specific: no standard name
try eq("", description("")); // no HID at all
}
test "named values round-trip through their _HID strings" {
const testing = std.testing;
try testing.expectEqualStrings("PNP0303", HardwareId.ps2_keyboard.hid());
try testing.expectEqual(@as(?HardwareId, .ps2_mouse), HardwareId.fromHid("PNP0F13"));
try testing.expectEqual(@as(?HardwareId, null), HardwareId.fromHid("QEMU0002"));
for (std.enums.values(HardwareId)) |id| {
try testing.expectEqual(@as(?HardwareId, id), HardwareId.fromHid(id.hid()));
}
}
+142 -500
View File
@@ -40,6 +40,9 @@ pub const RegisterAccess = struct {
/// Everything the power subsystem needs, extracted from the FADT and the AML
/// sleep packages during discovery. Populated by `discover`, read by `power`.
pub const PowerInformation = struct {
/// The System Control Interrupt's GSI (FADT SCI_INT) — the line ACPI events
/// (power button, GPEs) arrive on. Published to the acpi service for M21.
sci_interrupt: u16 = 0,
/// The SMM command port and the value that switches the platform into ACPI mode.
smi_cmd: u16 = 0,
acpi_enable: u8 = 0,
@@ -51,7 +54,8 @@ pub const PowerInformation = struct {
reset: RegisterAccess = .{},
reset_value: u8 = 0,
reset_supported: bool = false,
/// SLP_TYP values for S5 (soft off) and S3 (suspend), from the AML sleep-state (`_Sx`) packages.
/// SLP_TYP values for S5 (soft off) and S3 (suspend), from the AML sleep-state (`_Sx`)
/// packages.
s5: ?aml.SleepType = null,
s3: ?aml.SleepType = null,
};
@@ -153,6 +157,13 @@ pub var namespace: ?aml.Namespace = null;
/// Physical address of the DSDT the FADT points at, or 0.
pub var dsdt_physical: u64 = 0;
/// The FADT itself (physical + length), published on the acpi-tables node so
/// the ring-3 acpi service can read the PM1 event and GPE blocks it needs for
/// the event side (docs/m21-plan.md decision 3). Distinguished from the AML
/// blob resources by its intact "FACP" header — the blobs are header-stripped.
var fadt_physical: u64 = 0;
var fadt_length: u64 = 0;
// AML blocks (DSDT + any SSDTs) collected during the table walk, as physical
// address + length of each table's post-header bytecode. Scanned after the walk
// for the sleep-state (`_Sx`) packages.
@@ -197,7 +208,8 @@ const ExtendedSystemDescriptorPointer = extern struct {
root_system_description_table_address: u32 align(1),
/// The size of the RSDP.
length: u32 align(1),
/// A 64-bit physical address pointing to the XSDT. If the revision is at least 2, the XSDT should be used regardless of architecture, as the RSDT was deprecated.
/// A 64-bit physical address pointing to the XSDT. If the revision is at least 2, the XSDT
/// should be used regardless of architecture, as the RSDT was deprecated.
extended_system_descriptor_table_address: u64 align(1),
/// A checksum used for the entire table.
extended_checksum: u8,
@@ -357,35 +369,19 @@ const Hpet = extern struct {
page_protection: u8,
};
// --- PCI configuration-space header (first 64 bytes, common fields) ---------
const PciHeader = extern struct {
vendor_id: u16 align(1),
device_id: u16 align(1),
command: u16 align(1),
status: u16 align(1),
revision_id: u8,
prog_if: u8,
subclass: u8,
class_code: u8,
cache_line_size: u8,
latency_timer: u8,
/// bit 7 set => multi-function device.
header_type: u8,
bist: u8,
// 0x10 onward (BARs, etc.) depends on header_type; read separately.
};
// --- Entry point ------------------------------------------------------------
/// Discover hardware from the ACPI tables rooted at `rsdp_physical` and populate
/// `device_tree`. `hal` provides MMIO mapping (for PCIe ECAM) and port I/O. Also parses the
/// FADT and the AML sleep-state (`_Sx`) packages into `power_information` for the power service.
pub fn discover(rsdp_physical: u64, device_tree: *DeviceTree, hal: Hal) !void {
pub fn discover(rsdp_physical: u64, memory_regions: []const boot_handoff.MemoryRegion, device_tree: *DeviceTree, hal: Hal) !void {
if (rsdp_physical == 0) return error.NoRsdp;
boot_memory_regions = memory_regions;
// Start clean so a re-run doesn't accumulate stale state.
power_information = .{};
fadt_physical = 0;
fadt_length = 0;
platform_information = .{};
aml_stats = .{};
namespace = null;
@@ -417,12 +413,58 @@ pub fn discover(rsdp_physical: u64, device_tree: *DeviceTree, hal: Hal) !void {
aml_stats = .{ .nodes = namespace.?.nodeCount(), .consumed = pr.consumed, .total = pr.total };
power_information.s5 = aml.sleepState(&namespace.?, 5);
power_information.s3 = aml.sleepState(&namespace.?, 3);
// Fold the namespace's Device objects into the generic tree.
wireAcpiDevices(device_tree, &namespace.?, hal) catch {};
// The namespace's Device objects are no longer folded into the kernel
// tree (M20.3): the ring-3 acpi service claims the acpi-tables node
// (published below), re-parses the same blobs, and registers + reports
// the _HID devices itself. The kernel keeps the namespace only for the
// \_S5 sleep type above.
} else |_| {
// AML parse failed (e.g. out of memory); power stays best-effort with
// whatever the FADT alone provided.
}
// Publish the acpi-tables node (docs/m19-m20-plan.md M20): the AML blobs as
// memory resources for the acpi service to map and parse in ring 3, a broad
// io_port grant for the OperationRegion access its interpreter needs, and
// the SCI for the events track (M21). Exactly one node, one trusted
// claimant. Kept even when the kernel-side device building (above) retires
// in M20.3 — the kernel still owns the *static* tables and \_S5.
publishAcpiTablesNode(device_tree) catch {};
}
/// Build the acpi-tables node (see the call site in discover). Best-effort: a
/// failure here leaves the kernel-seeded tree working, only the ring-3 service
/// finds nothing to claim.
fn publishAcpiTablesNode(device_tree: *DeviceTree) !void {
const node = try device_tree.addChild(device_tree.root, .acpi_tables, "acpi-tables");
// One memory resource per AML block — page-aligned base down, length padded
// up to cover the bytecode, so mmio_map hands the service a pointer into it.
var i: usize = 0;
while (i < aml_block_count and i < device_model.maximum_resources - 2) : (i += 1) {
// mmio_map preserves the sub-page offset, so the service maps this and
// gets a pointer straight to the bytecode.
_ = node.addResource(.memory, aml_block_physical[i], aml_block_len[i]);
}
// The broad I/O grant: OperationRegions name whatever ports the firmware
// chose (EC, PM1, GPE, SMBus); which ports cannot be known before the AML
// that names them is parsed, so the grant is the whole space — the honest
// trust boundary of docs/m19-m20-plan.md decision 5.
_ = node.addResource(.io_port, 0, 1 << 16);
// A broad interrupt window: ACPI _CRS names legacy ISA IRQs (the PS/2 lines
// 1 and 12, the RTC, …), and the service registers those devices under this
// node, so it must own a superset. The range [0, 256) covers every GSI; the
// SCI (recorded first, len 1) stays distinct so M21 can pick it out.
if (power_information.sci_interrupt != 0) _ = node.addResource(.irq, power_information.sci_interrupt, 1);
_ = node.addResource(.irq, 0, 256);
// The FADT rides along (M21): the service reads the PM1 event / GPE blocks
// from its own copy, telling it apart from the AML blobs by signature.
if (fadt_physical != 0) _ = node.addResource(.memory, fadt_physical, fadt_length);
}
/// The number of Device objects in the namespace built during discovery, or 0.
pub fn amlDeviceCount() usize {
if (namespace) |*ns| return aml.deviceCount(ns);
return 0;
}
/// Walk the RSDT (Entry = u32) or XSDT (Entry = u64): validate it, then dispatch
@@ -448,10 +490,12 @@ fn handleTable(device_tree: *DeviceTree, hal: Hal, sdt_physical: u64) !void {
if (std.mem.eql(u8, &sig, &APIC)) {
try parseMadt(device_tree, header);
} else if (std.mem.eql(u8, &sig, &MCFG)) {
try parseMcfg(device_tree, hal, header);
try parseMcfg(device_tree, header);
} else if (std.mem.eql(u8, &sig, &HPET)) {
try parseHpet(device_tree, hal, header);
} else if (std.mem.eql(u8, &sig, &FACP)) {
fadt_physical = sdt_physical;
fadt_length = header.length;
parseFadt(header);
} else if (std.mem.eql(u8, &sig, &SPCR)) {
parseSpcr(header);
@@ -533,7 +577,7 @@ fn parseMadt(device_tree: *DeviceTree, header: *const SystemDescriptorTableHeade
}
/// MCFG -> a pci_host_bridge per ECAM segment, then a PCI enumeration underneath.
fn parseMcfg(device_tree: *DeviceTree, hal: Hal, header: *const SystemDescriptorTableHeader) !void {
fn parseMcfg(device_tree: *DeviceTree, header: *const SystemDescriptorTableHeader) !void {
const total: usize = header.length;
const base: [*]const u8 = @ptrCast(header);
@@ -548,109 +592,85 @@ fn parseMcfg(device_tree: *DeviceTree, hal: Hal, header: *const SystemDescriptor
// ECAM window: 1 MiB of configuration space per bus.
_ = bridge.addResource(.memory, alloc.base_address, bus_count << 20);
_ = bridge.addResource(.bus_range, alloc.start_bus, bus_count);
addBridgeApertures(bridge);
// The bridge decodes the whole 16-bit I/O space toward its bus — the
// window functions' I/O BARs must register-contain within (M19.2).
_ = bridge.addResource(.io_port, 0, 1 << 16);
try enumeratePci(device_tree, bridge, hal, alloc.*);
// The function walk itself retired to ring 3 (M19.3): the pci-bus
// driver claims this bridge, repeats the scan through its ECAM grant,
// and device_registers what it finds — the kernel seeds only the
// bridge. The scan's equivalence was proven before the hand-off
// (pci-scan), and the walk's history is in git if archaeology calls.
}
}
/// Brute-force scan the ECAM window's bus range for present PCI functions. No
/// bridge recursion yet: on the ECAM path the host bridge decodes every bus in
/// the window, so scanning the declared range finds everything QEMU exposes.
fn enumeratePci(
device_tree: *DeviceTree,
bridge: *device_model.Device,
hal: Hal,
alloc: McfgAllocation,
) !void {
var bus: u16 = alloc.start_bus;
while (bus <= alloc.end_bus) : (bus += 1) {
var device: u8 = 0;
while (device < 32) : (device += 1) {
const h0: *align(1) const PciHeader = @ptrCast(pciConfigurationPtr(alloc, hal, @intCast(bus), device, 0));
if (h0.vendor_id == 0xFFFF) continue; // no function 0 => slot empty
/// The boot memory map, stored at discover() entry for the aperture derivation
/// below (and, in M20, for the acpi-tables node's containment windows).
var boot_memory_regions: []const boot_handoff.MemoryRegion = &.{};
const funcs: u8 = if (h0.header_type & 0x80 != 0) 8 else 1;
var function: u8 = 0;
while (function < funcs) : (function += 1) {
const configuration = pciConfigurationPtr(alloc, hal, @intCast(bus), device, function);
const h: *align(1) const PciHeader = @ptrCast(configuration);
if (h.vendor_id == 0xFFFF) continue;
var nb: [24]u8 = undefined;
const nm = std.fmt.bufPrint(&nb, "{s}:{x:0>2}:{x:0>2}.{d}", .{
bridge.name(), bus, device, function,
}) catch "pcidev";
const node = try device_tree.addChild(bridge, .pci_device, nm);
// Resource 0 is the function's own 4 KiB ECAM configuration space. A
// claimed PCI driver mmio_maps this to reach its command register,
// BARs, and — the point — its capability list (MSI/MSI-X, PCIe
// extended caps), without any new syscall. Physical address per the
// ECAM formula (same as pciConfigurationPtr).
const config_physical = alloc.base_address +
(@as(u64, @as(u8, @intCast(bus)) - alloc.start_bus) << 20) +
(@as(u64, device) << 15) + (@as(u64, function) << 12);
_ = node.addResource(.memory, config_physical, abi.page_size);
node.ids.pci_vendor = h.vendor_id;
node.ids.pci_device = h.device_id;
node.ids.pci_class = (@as(u24, h.class_code) << 16) |
(@as(u24, h.subclass) << 8) | h.prog_if;
node.ids.pci_bdf = (@as(u16, @intCast(bus)) << 8) | (@as(u16, device) << 3) | function;
// BARs only exist in header type 0 (normal devices), not bridges.
if (h.header_type & 0x7F == 0) addBars(node, configuration);
/// The bridge's MMIO apertures, derived from the boot memory map's holes
/// (docs/m19-m20-plan.md decision 2): registered PCI functions carry BAR
/// resources, and `device_register` containment demands the bridge own windows
/// that cover them. Everything the firmware described is "not hole"; the low
/// aperture runs from the end of the described space below 4 GiB up to the
/// I/O-APIC region, the high one from 4 GiB (or the end of RAM above it) to
/// the 46-bit line. Coarse, mechanical, and AML-free — available at boot no
/// matter what later moved to user space.
fn addBridgeApertures(bridge: *device_model.Device) void {
// Below 4 GiB the described regions are sparse (RAM low, firmware flash
// and tables high), so the holes are the *gaps between* them — a single
// "after the last region" rule dies on OVMF's flash at the very top.
// Sort-merge the described ranges, then keep the three largest gaps
// (resource slots are bounded at 8 per device; ECAM + bus range + 3 + the
// high aperture fits). Above 4 GiB one aperture runs from the end of the
// described space to the 46-bit line.
const Range = struct { base: u64, end: u64 };
var below: [64]Range = undefined;
var below_count: usize = 0;
var high_end: u64 = 1 << 32;
for (boot_memory_regions) |region| {
const end = region.base + region.pages * 4096;
// Above 4 GiB only *usable RAM* blocks the aperture: OVMF describes
// its own 64-bit PCI window as a reserved region and then programs
// BARs inside it — honoring reserved there would exclude the very
// space BARs live in. Below 4 GiB every described region blocks (the
// kernel image, the tables, the ramdisk all live there). Bring-up
// trust: only the bridge's claimant can register into the aperture.
if (region.kind == .usable and end > high_end) high_end = end;
if (region.base >= (1 << 32) or below_count == below.len) continue;
below[below_count] = .{ .base = region.base, .end = @min(end, 1 << 32) };
below_count += 1;
}
// Insertion sort by base (the map is small and this runs once at boot).
for (1..below_count) |i| {
const key = below[i];
var j = i;
while (j > 0 and below[j - 1].base > key.base) : (j -= 1) below[j] = below[j - 1];
below[j] = key;
}
// Walk the sorted ranges, collecting inter-region gaps of at least 1 MiB.
var gaps: [3]Range = .{Range{ .base = 0, .end = 0 }} ** 3;
var cursor: u64 = 0;
var index: usize = 0;
while (index <= below_count) : (index += 1) {
const gap_end = if (index == below_count) (1 << 32) else below[index].base;
if (gap_end > cursor and gap_end - cursor >= (1 << 20)) {
// Keep the three largest, replacing the smallest kept so far.
var smallest: usize = 0;
for (gaps, 0..) |gap, gi| {
if (gap.end - gap.base < gaps[smallest].end - gaps[smallest].base) smallest = gi;
}
if (gap_end - cursor > gaps[smallest].end - gaps[smallest].base) {
gaps[smallest] = .{ .base = cursor, .end = gap_end };
}
}
if (index < below_count and below[index].end > cursor) cursor = below[index].end;
}
}
/// Record and size the memory/IO windows named by a device's Base Address
/// Registers. Sizing is the standard probe: disable decode, write all-ones, read
/// back the writable (address) bits, restore. `size = ~mask + 1`.
fn addBars(node: *device_model.Device, configuration: [*]align(1) u8) void {
// Stop the device decoding its BARs while we transiently write all-ones.
const command = rd(u16, configuration, 0x04);
wr(u16, configuration, 0x04, command & ~@as(u16, 0b11));
var i: usize = 0;
while (i < 6) : (i += 1) {
const off = 0x10 + i * 4;
const orig = rd(u32, configuration, off);
if (orig == 0) continue;
if (orig & 1 != 0) {
// I/O-space BAR (16-bit address space on x86).
wr(u32, configuration, off, 0xFFFF_FFFF);
const readback = rd(u32, configuration, off);
wr(u32, configuration, off, orig);
const mask = readback & 0xFFFF_FFFC;
const size: u32 = if (mask == 0) 0 else (~mask +% 1) & 0xFFFF;
_ = node.addResource(.io_port, orig & 0xFFFF_FFFC, size);
} else if ((orig >> 1) & 0x3 == 2) {
// 64-bit memory BAR: this BAR pair spans two configuration slots.
const orig_hi = rd(u32, configuration, off + 4);
wr(u32, configuration, off, 0xFFFF_FFFF);
wr(u32, configuration, off + 4, 0xFFFF_FFFF);
const lo = rd(u32, configuration, off);
const hi = rd(u32, configuration, off + 4);
wr(u32, configuration, off, orig);
wr(u32, configuration, off + 4, orig_hi);
const readback = (@as(u64, hi) << 32) | (lo & 0xFFFF_FFF0);
const size: u64 = if (readback == 0) 0 else ~readback +% 1;
const address = (@as(u64, orig_hi) << 32) | (orig & 0xFFFF_FFF0);
_ = node.addResource(.memory, address, size);
i += 1; // consumed the high half
} else {
// 32-bit memory BAR.
wr(u32, configuration, off, 0xFFFF_FFFF);
const readback = rd(u32, configuration, off);
wr(u32, configuration, off, orig);
const mask = readback & 0xFFFF_FFF0;
const size: u32 = if (mask == 0) 0 else ~mask +% 1;
_ = node.addResource(.memory, orig & 0xFFFF_FFF0, size);
}
for (gaps) |gap| {
if (gap.end > gap.base) _ = bridge.addResource(.memory, gap.base, gap.end - gap.base);
}
wr(u16, configuration, 0x04, command); // restore decode
_ = bridge.addResource(.memory, high_end, (@as(u64, 1) << 46) - high_end);
}
/// HPET -> a timer node with its register block as an MMIO resource, plus the GSI
@@ -709,6 +729,7 @@ const fadt_pm1a_cnt_blk = 64; // u32 (I/O port)
const fadt_pm1b_cnt_blk = 68; // u32 (I/O port)
const fadt_pm_tmr_blk = 76; // u32 (I/O port) — the PM timer counter
const fadt_pm1_cnt_len = 89; // u8 (bytes)
const fadt_sci_int = 46; // u16 (the SCI's GSI)
const fadt_flags = 112; // u32
const fadt_reset_register = 116; // GAS (12 bytes)
const fadt_reset_value = 128; // u8
@@ -726,6 +747,7 @@ fn parseFadt(header: *const SystemDescriptorTableHeader) void {
const len: usize = header.length;
const pi = &power_information;
pi.sci_interrupt = @truncate(fadt(u16, base, len, fadt_sci_int) orelse 0);
pi.smi_cmd = @truncate(fadt(u32, base, len, fadt_smi_cmd) orelse 0);
pi.acpi_enable = fadt(u8, base, len, fadt_acpi_enable) orelse 0;
pi.acpi_disable = fadt(u8, base, len, fadt_acpi_disable) orelse 0;
@@ -805,337 +827,6 @@ fn parseDmar(hal: Hal, header: *const SystemDescriptorTableHeader) void {
}
}
// --- AML namespace -> generic device tree -----------------------------------
/// The PCI bus context while descending the ACPI namespace: the generic host
/// bridge whose children ACPI address (`_ADR`) devices resolve against, and the bus number.
const PciContext = struct { bridge: *device_model.Device, bus: u8 };
/// Mirror the ACPI namespace's Device objects into the generic tree, *merging*
/// them with the PCI-enumerated nodes: a PCI root bridge (`PNP0A03`/`PNP0A08`)
/// folds onto the existing `pci_host_bridge`, and each addressed (`_ADR`) device folds onto
/// the matching PCI function (annotating it with the ACPI hardware ID (`_HID`) and nesting the
/// ACPI-only children — keyboard, RTC, … — beneath it). Namespace devices with no
/// PCI match land under a synthetic `acpi` node.
fn wireAcpiDevices(device_tree: *DeviceTree, aml_namespace: *aml.Namespace, hal: Hal) !void {
var arena = std.heap.ArenaAllocator.init(device_tree.allocator);
defer arena.deinit();
var interpreter = aml.Interpreter.init(aml_namespace, .{
.mapMmio = hal.mapMmio,
.pioRead = hal.pioRead,
.pioWrite = hal.pioWrite,
}, arena.allocator());
const acpi_root = try device_tree.addChild(device_tree.root, .unknown, "acpi");
try mirrorDevices(device_tree, aml_namespace.root, acpi_root, null, &interpreter);
}
fn mirrorDevices(device_tree: *DeviceTree, node: *aml.Node, parent_device: *device_model.Device, context: ?PciContext, interpreter: *aml.Interpreter) (error{OutOfMemory})!void {
var child = node.first_child;
while (child) |c| : (child = c.next_sibling) {
if (c.kind != .device) {
// A scope — the System Bus (\_SB), General Purpose Events (\_GPE), … —
// descend without adding a node.
try mirrorDevices(device_tree, c, parent_device, context, interpreter);
continue;
}
// Skip devices the firmware reports as not present (via a device-status (`_STA`) method),
// along with their whole subtree — per the ACPI rules.
if (!devicePresent(interpreter, c)) continue;
var mirrored_device: *device_model.Device = undefined;
var child_context = context;
if (isPciRootNode(c)) {
// The PCI root bridge folds onto the generic host bridge.
mirrored_device = matchHostBridge(device_tree) orelse
try device_tree.addChild(parent_device, .acpi_device, &c.segment);
child_context = .{ .bridge = mirrored_device, .bus = 0 };
} else {
// An addressed device folds onto its matching PCI function; anything
// else becomes a fresh node under the current parent.
mirrored_device = pick: {
if (context) |pc| {
if (readAdr(c)) |adr| {
if (findPciNode(pc.bridge, pc.bus, adr)) |pnode| break :pick pnode;
}
}
break :pick try device_tree.addChild(parent_device, .acpi_device, &c.segment);
};
}
applyHid(mirrored_device, c, interpreter);
applyCrs(mirrored_device, c, interpreter);
try mirrorDevices(device_tree, c, mirrored_device, child_context, interpreter);
}
}
/// Evaluate a device's status (`_STA`) to decide if it is present. An absent status
/// (`_STA`) means present by default; an evaluation failure is treated as present too (we'd
/// rather over-report than hide a device we couldn't introspect).
fn devicePresent(interpreter: *aml.Interpreter, node: *aml.Node) bool {
const sta = aml.Namespace.childOf(node, seg4("_STA")) orelse return true;
const obj = interpreter.evaluate(sta, &.{}) catch return true;
const status = obj.asInteger() catch return true;
return (status & 0x01) != 0; // bit 0 = present
}
/// The first PCI host bridge in the generic tree (segment 0).
fn matchHostBridge(device_tree: *DeviceTree) ?*device_model.Device {
var c = device_tree.root.first_child;
while (c) |ch| : (c = ch.next_sibling) {
if (ch.class == .pci_host_bridge) return ch;
}
return null;
}
/// The PCI function node under `bridge` at the address the device's address object
/// (`_ADR`) names (device/function on
/// `bus`), or null.
fn findPciNode(bridge: *device_model.Device, bus: u8, adr: u32) ?*device_model.Device {
const device: u16 = @truncate((adr >> 16) & 0x1F);
const function: u16 = @truncate(adr & 0x7);
const target: u16 = (@as(u16, bus) << 8) | (device << 3) | function;
var c = bridge.first_child;
while (c) |ch| : (c = ch.next_sibling) {
if (ch.ids.pci_bdf) |bdf| {
if (bdf == target) return ch;
}
}
return null;
}
/// A device's address (`_ADR`) — a static integer Name — or null.
fn readAdr(node: *aml.Node) ?u32 {
const n = aml.Namespace.childOf(node, seg4("_ADR")) orelse return null;
if (n.kind != .name) return null;
var p: usize = 0;
return @truncate(readIntObj(n.value, &p) orelse return null);
}
/// Whether a namespace device is a PCI(e) host bridge (`PNP0A03` / `PNP0A08`).
fn isPciRootNode(node: *aml.Node) bool {
const hid = aml.Namespace.childOf(node, seg4("_HID")) orelse return false;
if (hid.kind != .name or hid.value.len == 0) return false;
switch (hid.value[0]) {
0x00, 0x01, 0xFF, 0x0A, 0x0B, 0x0C, 0x0E => {
var p: usize = 0;
const n = readIntObj(hid.value, &p) orelse return false;
return n == 0x030AD041 or n == 0x080AD041; // PNP0A03 / PNP0A08
},
0x0D => {
const s = cstr(hid.value[1..]);
return std.mem.eql(u8, s, "PNP0A03") or std.mem.eql(u8, s, "PNP0A08");
},
else => return false,
}
}
/// Read a device's hardware ID (`_HID`) into the generic device: an integer decodes as an EISA
/// id ("PNP0A03"), a string is taken verbatim. Handles both the common static
/// Name form and a Method form (evaluated).
fn applyHid(device: *device_model.Device, node: *aml.Node, interpreter: *aml.Interpreter) void {
const hid = aml.Namespace.childOf(node, seg4("_HID")) orelse return;
if (hid.kind == .method) {
const obj = interpreter.evaluate(hid, &.{}) catch return;
switch (obj) {
.integer => |n| setEisaHid(device, @truncate(n)),
.string => |s| device.setHid(s),
else => {},
}
return;
}
if (hid.kind != .name or hid.value.len == 0) return;
const v = hid.value;
switch (v[0]) {
0x00, 0x01, 0xFF, 0x0A, 0x0B, 0x0C, 0x0E => {
var p: usize = 0;
const n = readIntObj(v, &p) orelse return;
setEisaHid(device, @truncate(n));
},
0x0D => device.setHid(cstr(v[1..])), // StringPrefix
else => {},
}
}
fn setEisaHid(device: *device_model.Device, id: u32) void {
device.ids.acpi_hid = id;
var buffer: [8]u8 = undefined;
device.setHid(eisaIdToStr(id, &buffer));
}
/// Parse a device's current resource settings (`_CRS`). The evaluator handles both the static
/// `Buffer` form (a `Name`) and the method form uniformly, yielding the
/// ResourceTemplate bytes we then decode.
fn applyCrs(device: *device_model.Device, node: *aml.Node, interpreter: *aml.Interpreter) void {
const crs = aml.Namespace.childOf(node, seg4("_CRS")) orelse return;
const obj = interpreter.evaluate(crs, &.{}) catch return;
const buffer = switch (obj) {
.buffer => |b| b,
else => return,
};
parseResourceTemplate(device, buffer);
}
/// Walk a ResourceTemplate byte list, adding recognised descriptors as resources.
fn parseResourceTemplate(device: *device_model.Device, bytes: []const u8) void {
var i: usize = 0;
while (i < bytes.len) {
const tag = bytes[i];
if (tag & 0x80 == 0) {
// Small descriptor: length in low 3 bits, type in bits [6:3].
const len: usize = tag & 0x07;
const body = i + 1;
if (body + len > bytes.len) break;
switch ((tag >> 3) & 0x0F) {
0x04 => if (len >= 2) { // IRQ: a 16-bit mask, one resource per set bit
const mask = @as(u16, bytes[body]) | (@as(u16, bytes[body + 1]) << 8);
var b: usize = 0;
while (b < 16) : (b += 1) {
if (mask & (@as(u16, 1) << @intCast(b)) != 0) _ = device.addResource(.irq, b, 1);
}
},
0x08 => if (len >= 7) { // IO port: minimum at +1, length at +6
_ = device.addResource(.io_port, rd16(bytes, body + 1), bytes[body + 6]);
},
0x09 => if (len >= 3) { // Fixed IO: base at +0, length at +2
_ = device.addResource(.io_port, rd16(bytes, body), bytes[body + 2]);
},
0x0F => break, // EndTag
else => {},
}
i = body + len;
} else {
// Large descriptor: 16-bit length follows the tag.
if (i + 3 > bytes.len) break;
const len: usize = @intCast(rd16(bytes, i + 1));
const body = i + 3;
if (body + len > bytes.len) break;
switch (tag) {
0x85 => if (len >= 17) { // Memory32: minimum at +1, length at +13
_ = device.addResource(.memory, rd32(bytes, body + 1), rd32(bytes, body + 13));
},
0x86 => if (len >= 9) { // Memory32Fixed: base at +1, length at +5
_ = device.addResource(.memory, rd32(bytes, body + 1), rd32(bytes, body + 5));
},
0x89 => if (len >= 2) { // Extended IRQ: count at +1, then count u32s
const count = bytes[body + 1];
var k: usize = 0;
while (k < count and body + 2 + k * 4 + 4 <= body + len) : (k += 1) {
_ = device.addResource(.irq, rd32(bytes, body + 2 + k * 4), 1);
}
},
0x87, 0x88, 0x8A => parseAddressSpace(device, tag, bytes[body .. body + len]),
else => {},
}
i = body + len;
}
}
}
/// Word/DWord/QWord address-space descriptors: resource type at [0], then
/// granularity/minimum/maximum/translation/length, each of width `w`.
fn parseAddressSpace(device: *device_model.Device, tag: u8, body: []const u8) void {
const w: usize = switch (tag) {
0x88 => 2, // Word
0x87 => 4, // DWord
else => 8, // QWord (0x8A)
};
if (body.len < 3 + 5 * w) return;
const minimum = readN(body, 3 + w, w);
const length = readN(body, 3 + 4 * w, w);
const kind: device_model.ResourceKind = switch (body[0]) {
0 => .memory,
1 => .io_port,
else => .bus_range,
};
_ = device.addResource(kind, minimum, length);
}
/// Decode a packed EISA id into its 7-char string (e.g. 0x030AD041 -> "PNP0A03").
fn eisaIdToStr(id: u32, buffer: *[8]u8) []const u8 {
const b0: u16 = @intCast(id & 0xFF);
const b1: u16 = @intCast((id >> 8) & 0xFF);
const b2: u8 = @truncate(id >> 16);
const b3: u8 = @truncate(id >> 24);
const mfg = (b0 << 8) | b1;
buffer[0] = '@' + @as(u8, @intCast((mfg >> 10) & 0x1F));
buffer[1] = '@' + @as(u8, @intCast((mfg >> 5) & 0x1F));
buffer[2] = '@' + @as(u8, @intCast(mfg & 0x1F));
buffer[3] = hexDigit((b2 >> 4) & 0xF);
buffer[4] = hexDigit(b2 & 0xF);
buffer[5] = hexDigit((b3 >> 4) & 0xF);
buffer[6] = hexDigit(b3 & 0xF);
return buffer[0..7];
}
fn hexDigit(n: u8) u8 {
return if (n < 10) '0' + n else 'A' + (n - 10);
}
fn seg4(comptime s: *const [4:0]u8) [4]u8 {
return s[0..4].*;
}
fn cstr(bytes: []const u8) []const u8 {
const index = std.mem.indexOfScalar(u8, bytes, 0) orelse bytes.len;
return bytes[0..index];
}
const PkgLen = struct { value: usize, size: usize };
fn packageLength(bytes: []const u8, p: usize) ?PkgLen {
if (p >= bytes.len) return null;
const lead = bytes[p];
const follow: usize = lead >> 6;
if (p + 1 + follow > bytes.len) return null;
if (follow == 0) return .{ .value = lead & 0x3F, .size = 1 };
var value: usize = lead & 0x0F;
var i: usize = 0;
while (i < follow) : (i += 1) value |= @as(usize, bytes[p + 1 + i]) << @intCast(4 + i * 8);
return .{ .value = value, .size = 1 + follow };
}
/// Read an AML integer object at `p`, advancing `p` past it.
fn readIntObj(bytes: []const u8, p: *usize) ?u64 {
if (p.* >= bytes.len) return null;
const opcode = bytes[p.*];
p.* += 1;
return switch (opcode) {
0x00 => 0,
0x01 => 1,
0xFF => 0xFF,
0x0A => readLE(bytes, p, 1),
0x0B => readLE(bytes, p, 2),
0x0C => readLE(bytes, p, 4),
0x0E => readLE(bytes, p, 8),
else => null,
};
}
fn readLE(bytes: []const u8, p: *usize, n: usize) ?u64 {
if (p.* + n > bytes.len) return null;
const v = readN(bytes, p.*, n);
p.* += n;
return v;
}
fn readN(bytes: []const u8, off: usize, n: usize) u64 {
var v: u64 = 0;
var k: usize = 0;
while (k < n and off + k < bytes.len) : (k += 1) v |= @as(u64, bytes[off + k]) << @intCast(k * 8);
return v;
}
fn rd16(bytes: []const u8, off: usize) u64 {
return readN(bytes, off, 2);
}
fn rd32(bytes: []const u8, off: usize) u64 {
return readN(bytes, off, 4);
}
// --- helpers ----------------------------------------------------------------
/// Sum `len` bytes; an ACPI table/pointer is valid when the low 8 bits are zero.
@@ -1176,58 +867,9 @@ fn readCntRegister(base: [*]align(1) const u8, len: usize, xoff: usize, legacy_o
return .{ .mmio = false, .address = port, .width = width };
}
/// The mapped configuration space of one PCI function (its 4 KiB ECAM page). Mapped
/// writable so BAR sizing can probe it; reads and writes both go through here.
fn pciConfigurationPtr(alloc: McfgAllocation, hal: Hal, bus: u8, device: u8, function: u8) [*]align(1) u8 {
const physical = alloc.base_address +
(@as(u64, bus - alloc.start_bus) << 20) +
(@as(u64, device) << 15) +
(@as(u64, function) << 12);
// Map the configuration page (writable, for BAR sizing) and use the virtual
// address the HAL hands back.
return @ptrFromInt(hal.mapMmio(physical, abi.page_size, true));
}
/// Read a little-endian integer at `off` from a (possibly unaligned) byte pointer.
/// x86 is little-endian and native, so an unaligned load suffices.
fn rd(comptime T: type, bytes: [*]align(1) const u8, off: usize) T {
const p: *align(1) const T = @ptrCast(bytes + off);
return p.*;
}
/// Write a little-endian integer at `off` through a (possibly unaligned) pointer.
fn wr(comptime T: type, bytes: [*]align(1) u8, off: usize, value: T) void {
const p: *align(1) T = @ptrCast(bytes + off);
p.* = value;
}
// --- tests ------------------------------------------------------------------
test "eisaIdToStr decodes a packed EISA id" {
var buffer: [8]u8 = undefined;
// 0x030AD041 is the well-known encoding of "PNP0A03" (PCI root bridge).
try std.testing.expectEqualStrings("PNP0A03", eisaIdToStr(0x030AD041, &buffer));
}
test "parseResourceTemplate extracts IO, IRQ, and fixed memory" {
// ResourceTemplate { IO(minimum 0x60, len 8), IRQ(4), Memory32Fixed(0xFED00000, 0x1000) }
const runtime = [_]u8{
0x47, 0x01, 0x60, 0x00, 0x60, 0x00, 0x01, 0x08, // small IO descriptor
0x22, 0x10, 0x00, // small IRQ descriptor (mask bit 4 -> IRQ 4)
0x86, 0x09, 0x00, 0x01, 0x00, 0x00, 0xD0, 0xFE, 0x00, 0x10, 0x00, 0x00, // Memory32Fixed
0x79, 0x00, // EndTag
};
var device = device_model.Device{};
parseResourceTemplate(&device, &runtime);
try std.testing.expectEqual(@as(u8, 3), device.resource_count);
const rs = device.resources[0..device.resource_count];
try std.testing.expectEqual(device_model.ResourceKind.io_port, rs[0].kind);
try std.testing.expectEqual(@as(u64, 0x60), rs[0].start);
try std.testing.expectEqual(@as(u64, 8), rs[0].len);
try std.testing.expectEqual(device_model.ResourceKind.irq, rs[1].kind);
try std.testing.expectEqual(@as(u64, 4), rs[1].start);
try std.testing.expectEqual(device_model.ResourceKind.memory, rs[2].kind);
try std.testing.expectEqual(@as(u64, 0xFED00000), rs[2].start);
try std.testing.expectEqual(@as(u64, 0x1000), rs[2].len);
}
+57 -5
View File
@@ -12,6 +12,11 @@ const std = @import("std");
const opcode = @import("opcodes.zig");
const parser = @import("parser.zig");
/// The named AML opcode/prefix bytes (`zero_opcode`, `byte_prefix`, …). Re-exported so
/// callers that decode raw AML bytes — e.g. the acpi service reading a `_HID` integer —
/// name the opcodes instead of writing bare 0x0A/0x0B/… literals (docs/coding-standards.md).
pub const opcodes = @import("opcodes.zig");
pub const Namespace = @import("namespace.zig").Namespace;
pub const Node = @import("namespace.zig").Node;
pub const NodeKind = @import("namespace.zig").NodeKind;
@@ -50,6 +55,20 @@ pub fn parse(allocator: std.mem.Allocator, blocks: []const []const u8) !ParseRes
return .{ .namespace = namespace, .consumed = consumed, .total = total };
}
/// Count the Device objects in a parsed namespace — what the acpi service
/// (docs/m19-m20-plan.md M20) reports, and what the kernel's own parse counts
/// so the two can be checked equal across the ring-3 move.
pub fn deviceCount(namespace: *const Namespace) usize {
return countKind(namespace.root, .device);
}
fn countKind(node: *const Node, kind: NodeKind) usize {
var n: usize = if (node.kind == kind) 1 else 0;
var c = node.first_child;
while (c) |child| : (c = child.next_sibling) n += countKind(child, kind);
return n;
}
/// Look up the `\_S{state}` sleep package in a parsed namespace and return its
/// first two integer elements (SLP_TYP for PM1a / PM1b), or null if absent.
pub fn sleepState(namespace: *Namespace, state: u8) ?SleepType {
@@ -126,17 +145,22 @@ test "parses a nested namespace and finds the sleep package" {
// Scope(\_SB) packagelen=0x27
0x10, 0x27, 0x5C, 0x5F, 0x53, 0x42, 0x5F,
// Device(PCI0) packagelen=0x1F
0x5B, 0x82, 0x1F, 0x50, 0x43, 0x49, 0x30,
0x5B, 0x82, 0x1F, 0x50, 0x43,
0x49, 0x30,
// Name(_HID, 0x11)
0x08, 0x5F, 0x48, 0x49, 0x44, 0x0A, 0x11,
// Method(MTHD, flags=1) empty, packagelen=0x06
0x14, 0x06, 0x4D, 0x54, 0x48, 0x44, 0x01,
0x14, 0x06, 0x4D,
0x54, 0x48, 0x44, 0x01,
// Method(CALL, flags=0) { MTHD(Zero) }, packagelen=0x0B
0x14, 0x0B, 0x43, 0x41, 0x4C, 0x4C, 0x00, 0x4D, 0x54, 0x48, 0x44, 0x00,
0x14, 0x0B, 0x43, 0x41, 0x4C, 0x4C, 0x00, 0x4D,
0x54, 0x48, 0x44, 0x00,
// OperationRegion(DBG0, SystemIO, Word 0x0402, Byte 1)
0x5B, 0x80, 0x44, 0x42, 0x47, 0x30, 0x01, 0x0B, 0x02, 0x04, 0x0A, 0x01,
0x5B, 0x80, 0x44, 0x42, 0x47, 0x30, 0x01, 0x0B,
0x02, 0x04, 0x0A, 0x01,
// Field(DBG0, flags=1) { DBGB, 8 }, packagelen=0x0B
0x5B, 0x81, 0x0B, 0x44, 0x42, 0x47, 0x30, 0x01, 0x44, 0x42, 0x47, 0x42, 0x08,
0x5B, 0x81, 0x0B, 0x44, 0x42, 0x47, 0x30, 0x01,
0x44, 0x42, 0x47, 0x42, 0x08,
};
var arena = std.heap.ArenaAllocator.init(std.testing.allocator);
@@ -206,3 +230,31 @@ test "interpreter runs a method with args, arithmetic, and control flow" {
const lo = try interpreter.evaluate(tst, &.{.{ .integer = 2 }}); // 2+5=7 !> 10 -> 0
try std.testing.expectEqual(@as(u64, 0), try lo.asInteger());
}
test "interpreter records Notify(device, code)" {
// Device(DEV_) { Name(_HID, 0x030AD041) } // PNP0A03-ish placeholder
// Method(TST_, 0) { Notify(DEV_, 0x80); Return(Zero) }
// Encoded: a Device holding a Name, then a Method issuing Notify on it.
const blob = [_]u8{
0x5B, 0x82, 0x0F, 0x44, 0x45, 0x56, 0x5F, // Device(DEV_) len=0x0F (pkglen + DEV_ + Name)
0x08, 0x5F, 0x48, 0x49, 0x44, 0x0C, 0x41, 0xD0, 0x0A, 0x03, // Name(_HID, DWord 0x030AD041)
0x14, 0x0F, 0x54, 0x53, 0x54, 0x5F, 0x00, // Method(TST_, 0) len=0x0F (pkglen + TST_ + flags + body)
0x86, 0x44, 0x45, 0x56, 0x5F, 0x0A, 0x80, // Notify(DEV_, 0x80)
0xA4, 0x00, // Return(Zero)
};
var arena = std.heap.ArenaAllocator.init(std.testing.allocator);
defer arena.deinit();
var result = try parse(arena.allocator(), &.{&blob});
const namespace = &result.namespace;
const tst = namespace.resolve(namespace.root, false, 0, &.{.{ 'T', 'S', 'T', '_' }}) orelse return error.NoMethod;
const dev = namespace.resolve(namespace.root, false, 0, &.{.{ 'D', 'E', 'V', '_' }}) orelse return error.NoDevice;
var interpreter = Interpreter.init(namespace, .{ .mapMmio = noMap, .pioRead = noRead, .pioWrite = noWrite }, arena.allocator());
_ = try interpreter.evaluate(tst, &.{});
const events = interpreter.takeNotifications();
try std.testing.expectEqual(@as(usize, 1), events.len);
try std.testing.expectEqual(dev, events[0].node);
try std.testing.expectEqual(@as(u64, 0x80), events[0].code);
}
+41
View File
@@ -141,6 +141,9 @@ const Frame = struct {
/// A CreateField binding: a name that indexes into a buffer object.
const BufferField = struct { buffer: *Node, byte_off: usize, bit_width: u32 };
/// One Notify(device, code) the interpreter executed.
pub const NotifyEvent = struct { node: *Node, code: u64 };
pub const Interpreter = struct {
namespace: *Namespace,
hal: Hal,
@@ -149,6 +152,11 @@ pub const Interpreter = struct {
dynamic_overrides: std.AutoHashMapUnmanaged(*Node, Object) = .{},
/// CreateField bindings active for the current evaluation.
fields: std.AutoHashMapUnmanaged(*Node, BufferField) = .{},
/// Notify(device, code) operations the last evaluation executed — a GPE or
/// EC handler tells the OS "look at this device" this way. Bounded; the
/// caller drains it with `takeNotifications` after `evaluate` (M21).
notify_queue: [16]NotifyEvent = undefined,
notify_count: usize = 0,
pub fn init(namespace: *Namespace, hal: Hal, arena: std.mem.Allocator) Interpreter {
return .{ .namespace = namespace, .hal = hal, .arena = arena };
@@ -157,6 +165,7 @@ pub const Interpreter = struct {
/// Evaluate a namespace object: invoke a Method, read a Name's value, or read a
/// Field. Resets per-evaluation runtime state first.
pub fn evaluate(self: *Interpreter, node: *Node, args: []const Object) Error!Object {
self.notify_count = 0;
self.dynamic_overrides.clearRetainingCapacity();
self.fields.clearRetainingCapacity();
return self.invoke(node, args);
@@ -267,6 +276,8 @@ pub const Interpreter = struct {
},
opcode.to_buffer_opcode => try self.passThroughUnary(current, frame),
opcode.notify_opcode => try self.notify(current, frame),
opcode.extended_opcode_prefix => try self.ext(current, frame),
// CreateXField: source, index, name (bit widths differ by op)
@@ -542,6 +553,36 @@ pub const Interpreter = struct {
try self.storeInto(current, frame, value);
}
/// Notify(SuperName, NotifyValue): resolve the named device, evaluate the
/// code, and record the pair for the caller to dispatch. AML control flow
/// continues (Notify returns nothing).
fn notify(self: *Interpreter, current: *Cursor, frame: *Frame) Error!Object {
const lead = current.peek() orelse return error.Truncated;
var target: ?*Node = null;
if (isNameStart(lead)) {
const name_path = try current.nameString();
target = self.namespace.resolve(frame.scope, name_path.rooted, name_path.parents, name_path.slice());
} else {
// A non-name SuperName (Local/Arg holding a reference).
const obj = try self.term(current, frame);
if (obj == .reference) target = obj.reference;
}
const code = try self.evaluateInteger(current, frame);
if (target) |node| {
if (self.notify_count < self.notify_queue.len) {
self.notify_queue[self.notify_count] = .{ .node = node, .code = code };
self.notify_count += 1;
}
}
return .uninitialized;
}
/// The Notify events the last `evaluate` produced. Valid until the next
/// `evaluate` clears the queue.
pub fn takeNotifications(self: *Interpreter) []const NotifyEvent {
return self.notify_queue[0..self.notify_count];
}
fn storeInto(self: *Interpreter, current: *Cursor, frame: *Frame, value: Object) Error!void {
const lead = current.peek() orelse return error.Truncated;
if (isNameStart(lead)) {
+14
View File
@@ -28,6 +28,11 @@ pub const DeviceClass = enum(u32) {
/// A device named in the ACPI namespace (from the DSDT/SSDT), carrying a
/// hardware ID (`_HID`) and, where static, current resource settings (`_CRS`).
acpi_device,
/// The ACPI tables themselves, published as one node for the user-space acpi
/// service (docs/m19-m20-plan.md M20): memory resources over the AML blobs,
/// a broad io_port grant for OperationRegion access, and the SCI interrupt.
/// The one node whose claimant is trusted to run firmware bytecode.
acpi_tables,
unknown,
};
@@ -56,6 +61,10 @@ pub const maximum_device_resources = 8;
/// `DeviceDescriptor.parent` for a device with no parent — a root of the device tree.
pub const no_parent: u64 = ~@as(u64, 0);
/// `DeviceDescriptor.pci_class` for a device that is not a PCI function. (Zero would be
/// ambiguous: 0x000000 is a real class code, "unclassified device".)
pub const no_pci_class: u64 = ~@as(u64, 0);
/// A device, as snapshotted for user space by `device_enumerate`. A driver scans
/// these to find the hardware it owns, claims it, and maps its MMIO.
///
@@ -69,6 +78,11 @@ pub const DeviceDescriptor = extern struct {
id: u64,
parent: u64, // a device id, or `no_parent`
class: u64, // a DeviceClass value
// The PCI class/subclass/prog-IF triple packed as 0xCCSSPP when this device is a PCI
// function, or `no_pci_class` otherwise. This is how a manager tells *what* a
// `pci_device` is (an xHCI controller, an AHCI controller) — decode the triple into
// names with the pci-class module.
pci_class: u64,
hid_len: u64,
resource_count: u64,
hid: [8]u8,
+2 -3
View File
@@ -3,7 +3,7 @@
//! Discovery backends (ACPI today, device-tree later) translate their native
//! hardware description into this one shape, so the rest of the kernel walks a
//! plain `Device` tree without knowing which firmware described the machine —
//! the same discipline `root.zig`'s `MemoryKind` applies to memory and `architecture`
//! the same discipline `ps2-library.zig`'s `MemoryKind` applies to memory and `architecture`
//! applies to the CPU.
//!
//! This is deliberately minimal: enough to *describe* what was discovered (a
@@ -198,8 +198,7 @@ fn dumpNode(device: *const Device, depth: usize, emit: *const fn ([]const u8) vo
std.fmt.bufPrint(buffer[indent..], "{s} [{s}] hid={s} ({s})\n", .{ device.name(), @tagName(device.class), device.hid(), desc }) catch return
else
std.fmt.bufPrint(buffer[indent..], "{s} [{s}] hid={s}\n", .{ device.name(), @tagName(device.class), device.hid() }) catch return;
} else
std.fmt.bufPrint(buffer[indent..], "{s} [{s}]\n", .{ device.name(), @tagName(device.class) }) catch return;
} else std.fmt.bufPrint(buffer[indent..], "{s} [{s}]\n", .{ device.name(), @tagName(device.class) }) catch return;
emit(buffer[0 .. indent + body.len]);
// For a PCI function, decode its class code — the (class / subclass / prog-IF)
+496 -189
View File
@@ -7,6 +7,16 @@
//! apart. Pure reference data (from the PCI spec; see https://wiki.osdev.org/PCI) — no
//! hardware access — so it is shared by kernel discovery (the device-tree dump) and any
//! user-space tool (a future lspci, driver matching).
//!
//! The taxonomy is named, not numbered (docs/coding-standards.md, "Named values"): the
//! base class is a `BaseClass` enum, and each class with defined subclasses gets a
//! namespace holding its `SubClass` enum (and, where the spec defines them, per-subclass
//! `ProgIf` enums) — the same shape as `usb-ids.zig`. Code that *means* a specific class
//! names it (`BaseClass.serial_bus`, `serial_bus.usb.ProgIf.xhci`) rather than writing a
//! bare 0x0C/0x03/0x30. The `className`/`subclassName`/`progIfName` functions still take
//! the raw bytes a function reports in its header, because that is what hardware hands us.
const std = @import("std");
/// The three bytes of a PCI class code, unpacked from the `0xCCSSPP` value discovery
/// records in `Device.ids.pci_class` (CC = base class, SS = subclass, PP = prog-IF).
@@ -22,148 +32,465 @@ pub const ClassCode = struct {
.prog_if = @intCast(packed_code & 0xFF),
};
}
/// Re-pack the triple into the `0xCCSSPP` form. Lets code name a whole class code
/// from its parts — `pack(.{ .base = @intFromEnum(BaseClass.serial_bus), … })` —
/// instead of writing the literal 0x0C0330.
pub fn pack(self: ClassCode) u24 {
return (@as(u24, self.base) << 16) | (@as(u24, self.subclass) << 8) | self.prog_if;
}
};
/// Base class (config byte 0x0B). Non-exhaustive: an unlisted code is a real but
/// unnamed class, decoded as "Unknown" rather than rejected.
pub const BaseClass = enum(u8) {
unclassified = 0x00,
mass_storage = 0x01,
network = 0x02,
display = 0x03,
multimedia = 0x04,
memory = 0x05,
bridge = 0x06,
simple_communication = 0x07,
base_system_peripheral = 0x08,
input_device = 0x09,
docking_station = 0x0A,
processor = 0x0B,
serial_bus = 0x0C,
wireless = 0x0D,
intelligent = 0x0E,
satellite_communication = 0x0F,
encryption = 0x10,
signal_processing = 0x11,
processing_accelerator = 0x12,
non_essential_instrumentation = 0x13,
co_processor = 0x40,
unassigned = 0xFF,
_,
pub fn name(self: BaseClass) []const u8 {
return switch (self) {
.unclassified => "Unclassified",
.mass_storage => "Mass Storage Controller",
.network => "Network Controller",
.display => "Display Controller",
.multimedia => "Multimedia Controller",
.memory => "Memory Controller",
.bridge => "Bridge",
.simple_communication => "Simple Communication Controller",
.base_system_peripheral => "Base System Peripheral",
.input_device => "Input Device Controller",
.docking_station => "Docking Station",
.processor => "Processor",
.serial_bus => "Serial Bus Controller",
.wireless => "Wireless Controller",
.intelligent => "Intelligent Controller",
.satellite_communication => "Satellite Communication Controller",
.encryption => "Encryption Controller",
.signal_processing => "Signal Processing Controller",
.processing_accelerator => "Processing Accelerator",
.non_essential_instrumentation => "Non-Essential Instrumentation",
.co_processor => "Co-Processor",
.unassigned => "Unassigned Class (Vendor specific)",
_ => "Unknown",
};
}
};
// --- Per-class subclass (and prog-IF) taxonomies --------------------------------------
// One namespace per base class that has defined subclasses, named after the class. Each
// holds an exhaustive `SubClass` enum (so an unlisted code decodes to the class default,
// not a wrong name), and, where the spec assigns them, per-subclass `ProgIf` enums.
pub const mass_storage = struct {
pub const SubClass = enum(u8) {
scsi_bus = 0x00,
ide = 0x01,
floppy = 0x02,
ipi_bus = 0x03,
raid = 0x04,
ata = 0x05,
serial_ata = 0x06,
serial_attached_scsi = 0x07,
non_volatile_memory = 0x08,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.scsi_bus => "SCSI Bus Controller",
.ide => "IDE Controller",
.floppy => "Floppy Disk Controller",
.ipi_bus => "IPI Bus Controller",
.raid => "RAID Controller",
.ata => "ATA Controller",
.serial_ata => "Serial ATA Controller",
.serial_attached_scsi => "Serial Attached SCSI Controller",
.non_volatile_memory => "Non-Volatile Memory Controller",
};
}
};
pub const serial_ata = struct {
pub const ProgIf = enum(u8) {
vendor_specific = 0x00,
ahci = 0x01,
serial_storage_bus = 0x02,
pub fn name(self: ProgIf) []const u8 {
return switch (self) {
.vendor_specific => "Vendor Specific Interface",
.ahci => "AHCI 1.0",
.serial_storage_bus => "Serial Storage Bus",
};
}
};
};
pub const non_volatile_memory = struct {
pub const ProgIf = enum(u8) {
nvmhci = 0x01,
nvm_express = 0x02,
pub fn name(self: ProgIf) []const u8 {
return switch (self) {
.nvmhci => "NVMHCI",
.nvm_express => "NVM Express",
};
}
};
};
};
pub const network = struct {
pub const SubClass = enum(u8) {
ethernet = 0x00,
token_ring = 0x01,
fddi = 0x02,
atm = 0x03,
isdn = 0x04,
picmg_multi_computing = 0x06,
infiniband = 0x07,
fabric = 0x08,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.ethernet => "Ethernet Controller",
.token_ring => "Token Ring Controller",
.fddi => "FDDI Controller",
.atm => "ATM Controller",
.isdn => "ISDN Controller",
.picmg_multi_computing => "PICMG 2.14 Multi Computing Controller",
.infiniband => "Infiniband Controller",
.fabric => "Fabric Controller",
};
}
};
};
pub const display = struct {
pub const SubClass = enum(u8) {
vga_compatible = 0x00,
xga = 0x01,
three_dimensional = 0x02,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.vga_compatible => "VGA Compatible Controller",
.xga => "XGA Controller",
.three_dimensional => "3D Controller (Not VGA-Compatible)",
};
}
};
pub const vga_compatible = struct {
pub const ProgIf = enum(u8) {
vga = 0x00,
compatible_8514 = 0x01,
pub fn name(self: ProgIf) []const u8 {
return switch (self) {
.vga => "VGA Controller",
.compatible_8514 => "8514-Compatible Controller",
};
}
};
};
};
pub const multimedia = struct {
pub const SubClass = enum(u8) {
video = 0x00,
audio = 0x01,
telephony = 0x02,
audio_device = 0x03,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.video => "Multimedia Video Controller",
.audio => "Multimedia Audio Controller",
.telephony => "Computer Telephony Device",
.audio_device => "Audio Device",
};
}
};
};
pub const memory = struct {
pub const SubClass = enum(u8) {
ram = 0x00,
flash = 0x01,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.ram => "RAM Controller",
.flash => "Flash Controller",
};
}
};
};
pub const bridge = struct {
pub const SubClass = enum(u8) {
host = 0x00,
isa = 0x01,
eisa = 0x02,
mca = 0x03,
pci_to_pci = 0x04,
pcmcia = 0x05,
nubus = 0x06,
cardbus = 0x07,
raceway = 0x08,
pci_to_pci_semi_transparent = 0x09,
infiniband_to_pci = 0x0A,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.host => "Host Bridge",
.isa => "ISA Bridge",
.eisa => "EISA Bridge",
.mca => "MCA Bridge",
.pci_to_pci => "PCI-to-PCI Bridge",
.pcmcia => "PCMCIA Bridge",
.nubus => "NuBus Bridge",
.cardbus => "CardBus Bridge",
.raceway => "RACEway Bridge",
.pci_to_pci_semi_transparent => "PCI-to-PCI Bridge (Semi-Transparent)",
.infiniband_to_pci => "InfiniBand-to-PCI Host Bridge",
};
}
};
pub const pci_to_pci = struct {
pub const ProgIf = enum(u8) {
normal_decode = 0x00,
subtractive_decode = 0x01,
pub fn name(self: ProgIf) []const u8 {
return switch (self) {
.normal_decode => "Normal Decode",
.subtractive_decode => "Subtractive Decode",
};
}
};
};
};
pub const simple_communication = struct {
pub const SubClass = enum(u8) {
serial = 0x00,
parallel = 0x01,
multiport_serial = 0x02,
modem = 0x03,
gpib = 0x04,
smart_card = 0x05,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.serial => "Serial Controller",
.parallel => "Parallel Controller",
.multiport_serial => "Multiport Serial Controller",
.modem => "Modem",
.gpib => "IEEE 488.1/2 (GPIB) Controller",
.smart_card => "Smart Card Controller",
};
}
};
pub const serial = struct {
pub const ProgIf = enum(u8) {
compatible_8250 = 0x00,
compatible_16450 = 0x01,
compatible_16550 = 0x02,
compatible_16650 = 0x03,
compatible_16750 = 0x04,
compatible_16850 = 0x05,
compatible_16950 = 0x06,
pub fn name(self: ProgIf) []const u8 {
return switch (self) {
.compatible_8250 => "8250-Compatible (Generic XT)",
.compatible_16450 => "16450-Compatible",
.compatible_16550 => "16550-Compatible",
.compatible_16650 => "16650-Compatible",
.compatible_16750 => "16750-Compatible",
.compatible_16850 => "16850-Compatible",
.compatible_16950 => "16950-Compatible",
};
}
};
};
};
pub const base_system_peripheral = struct {
pub const SubClass = enum(u8) {
pic = 0x00,
dma = 0x01,
timer = 0x02,
rtc = 0x03,
pci_hot_plug = 0x04,
sd_host = 0x05,
iommu = 0x06,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.pic => "PIC",
.dma => "DMA Controller",
.timer => "Timer",
.rtc => "RTC Controller",
.pci_hot_plug => "PCI Hot-Plug Controller",
.sd_host => "SD Host Controller",
.iommu => "IOMMU",
};
}
};
};
pub const input_device = struct {
pub const SubClass = enum(u8) {
keyboard = 0x00,
digitizer_pen = 0x01,
mouse = 0x02,
scanner = 0x03,
gameport = 0x04,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.keyboard => "Keyboard Controller",
.digitizer_pen => "Digitizer Pen",
.mouse => "Mouse Controller",
.scanner => "Scanner Controller",
.gameport => "Gameport Controller",
};
}
};
};
pub const serial_bus = struct {
pub const SubClass = enum(u8) {
firewire = 0x00,
access_bus = 0x01,
ssa = 0x02,
usb = 0x03,
fibre_channel = 0x04,
smbus = 0x05,
infiniband = 0x06,
ipmi = 0x07,
sercos = 0x08,
canbus = 0x09,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.firewire => "FireWire (IEEE 1394) Controller",
.access_bus => "ACCESS Bus Controller",
.ssa => "SSA",
.usb => "USB Controller",
.fibre_channel => "Fibre Channel",
.smbus => "SMBus Controller",
.infiniband => "InfiniBand Controller",
.ipmi => "IPMI Interface",
.sercos => "SERCOS Interface (IEC 61491)",
.canbus => "CANbus Controller",
};
}
};
pub const usb = struct {
pub const ProgIf = enum(u8) {
uhci = 0x00,
ohci = 0x10,
ehci = 0x20,
xhci = 0x30,
unspecified = 0x80,
device = 0xFE,
pub fn name(self: ProgIf) []const u8 {
return switch (self) {
.uhci => "UHCI Controller",
.ohci => "OHCI Controller",
.ehci => "EHCI (USB2) Controller",
.xhci => "XHCI (USB3) Controller",
.unspecified => "Unspecified",
.device => "USB Device (not a host controller)",
};
}
};
};
};
pub const wireless = struct {
pub const SubClass = enum(u8) {
irda = 0x00,
consumer_ir = 0x01,
rf = 0x10,
bluetooth = 0x11,
broadband = 0x12,
ethernet_802_1a = 0x20,
ethernet_802_1b = 0x21,
pub fn name(self: SubClass) []const u8 {
return switch (self) {
.irda => "iRDA Compatible Controller",
.consumer_ir => "Consumer IR Controller",
.rf => "RF Controller",
.bluetooth => "Bluetooth Controller",
.broadband => "Broadband Controller",
.ethernet_802_1a => "Ethernet Controller (802.1a)",
.ethernet_802_1b => "Ethernet Controller (802.1b)",
};
}
};
};
// --- Raw-byte decoding (what a function reports in its header) -------------------------
/// The name of an exhaustive class-code enum member, or null if `value` is not one — the
/// bridge from a raw config byte to a named taxonomy above.
fn enumName(comptime Enum: type, value: u8) ?[]const u8 {
return (std.enums.fromInt(Enum, value) orelse return null).name();
}
/// Name of the base class (byte 0x0B), e.g. `0x06` -> "Bridge".
pub fn className(base: u8) []const u8 {
return switch (base) {
0x00 => "Unclassified",
0x01 => "Mass Storage Controller",
0x02 => "Network Controller",
0x03 => "Display Controller",
0x04 => "Multimedia Controller",
0x05 => "Memory Controller",
0x06 => "Bridge",
0x07 => "Simple Communication Controller",
0x08 => "Base System Peripheral",
0x09 => "Input Device Controller",
0x0A => "Docking Station",
0x0B => "Processor",
0x0C => "Serial Bus Controller",
0x0D => "Wireless Controller",
0x0E => "Intelligent Controller",
0x0F => "Satellite Communication Controller",
0x10 => "Encryption Controller",
0x11 => "Signal Processing Controller",
0x12 => "Processing Accelerator",
0x13 => "Non-Essential Instrumentation",
0x40 => "Co-Processor",
0xFF => "Unassigned Class (Vendor specific)",
else => "Unknown",
};
return @as(BaseClass, @enumFromInt(base)).name();
}
/// Name of the subclass within its base class, e.g. `(0x06, 0x01)` -> "ISA Bridge".
/// Subclass `0x80` is "Other" by PCI convention; anything unlisted is "Unknown".
pub fn subclassName(base: u8, subclass: u8) []const u8 {
return switch (base) {
0x01 => switch (subclass) {
0x00 => "SCSI Bus Controller",
0x01 => "IDE Controller",
0x02 => "Floppy Disk Controller",
0x03 => "IPI Bus Controller",
0x04 => "RAID Controller",
0x05 => "ATA Controller",
0x06 => "Serial ATA Controller",
0x07 => "Serial Attached SCSI Controller",
0x08 => "Non-Volatile Memory Controller",
else => defaultSubclass(subclass),
},
0x02 => switch (subclass) {
0x00 => "Ethernet Controller",
0x01 => "Token Ring Controller",
0x02 => "FDDI Controller",
0x03 => "ATM Controller",
0x04 => "ISDN Controller",
0x06 => "PICMG 2.14 Multi Computing Controller",
0x07 => "Infiniband Controller",
0x08 => "Fabric Controller",
else => defaultSubclass(subclass),
},
0x03 => switch (subclass) {
0x00 => "VGA Compatible Controller",
0x01 => "XGA Controller",
0x02 => "3D Controller (Not VGA-Compatible)",
else => defaultSubclass(subclass),
},
0x04 => switch (subclass) {
0x00 => "Multimedia Video Controller",
0x01 => "Multimedia Audio Controller",
0x02 => "Computer Telephony Device",
0x03 => "Audio Device",
else => defaultSubclass(subclass),
},
0x05 => switch (subclass) {
0x00 => "RAM Controller",
0x01 => "Flash Controller",
else => defaultSubclass(subclass),
},
0x06 => switch (subclass) {
0x00 => "Host Bridge",
0x01 => "ISA Bridge",
0x02 => "EISA Bridge",
0x03 => "MCA Bridge",
0x04 => "PCI-to-PCI Bridge",
0x05 => "PCMCIA Bridge",
0x06 => "NuBus Bridge",
0x07 => "CardBus Bridge",
0x08 => "RACEway Bridge",
0x09 => "PCI-to-PCI Bridge (Semi-Transparent)",
0x0A => "InfiniBand-to-PCI Host Bridge",
else => defaultSubclass(subclass),
},
0x07 => switch (subclass) {
0x00 => "Serial Controller",
0x01 => "Parallel Controller",
0x02 => "Multiport Serial Controller",
0x03 => "Modem",
0x04 => "IEEE 488.1/2 (GPIB) Controller",
0x05 => "Smart Card Controller",
else => defaultSubclass(subclass),
},
0x08 => switch (subclass) {
0x00 => "PIC",
0x01 => "DMA Controller",
0x02 => "Timer",
0x03 => "RTC Controller",
0x04 => "PCI Hot-Plug Controller",
0x05 => "SD Host Controller",
0x06 => "IOMMU",
else => defaultSubclass(subclass),
},
0x09 => switch (subclass) {
0x00 => "Keyboard Controller",
0x01 => "Digitizer Pen",
0x02 => "Mouse Controller",
0x03 => "Scanner Controller",
0x04 => "Gameport Controller",
else => defaultSubclass(subclass),
},
0x0C => switch (subclass) {
0x00 => "FireWire (IEEE 1394) Controller",
0x01 => "ACCESS Bus Controller",
0x02 => "SSA",
0x03 => "USB Controller",
0x04 => "Fibre Channel",
0x05 => "SMBus Controller",
0x06 => "InfiniBand Controller",
0x07 => "IPMI Interface",
0x08 => "SERCOS Interface (IEC 61491)",
0x09 => "CANbus Controller",
else => defaultSubclass(subclass),
},
0x0D => switch (subclass) {
0x00 => "iRDA Compatible Controller",
0x01 => "Consumer IR Controller",
0x10 => "RF Controller",
0x11 => "Bluetooth Controller",
0x12 => "Broadband Controller",
0x20 => "Ethernet Controller (802.1a)",
0x21 => "Ethernet Controller (802.1b)",
else => defaultSubclass(subclass),
},
else => defaultSubclass(subclass),
const named: ?[]const u8 = switch (@as(BaseClass, @enumFromInt(base))) {
.mass_storage => enumName(mass_storage.SubClass, subclass),
.network => enumName(network.SubClass, subclass),
.display => enumName(display.SubClass, subclass),
.multimedia => enumName(multimedia.SubClass, subclass),
.memory => enumName(memory.SubClass, subclass),
.bridge => enumName(bridge.SubClass, subclass),
.simple_communication => enumName(simple_communication.SubClass, subclass),
.base_system_peripheral => enumName(base_system_peripheral.SubClass, subclass),
.input_device => enumName(input_device.SubClass, subclass),
.serial_bus => enumName(serial_bus.SubClass, subclass),
.wireless => enumName(wireless.SubClass, subclass),
else => null,
};
return named orelse defaultSubclass(subclass);
}
fn defaultSubclass(subclass: u8) []const u8 {
@@ -175,68 +502,34 @@ fn defaultSubclass(subclass: u8) []const u8 {
/// Returns "" when the prog-IF carries no standard meaning for this class/subclass —
/// callers just print the hex byte in that case.
pub fn progIfName(base: u8, subclass: u8, prog_if: u8) []const u8 {
return switch (base) {
0x01 => switch (subclass) {
0x06 => switch (prog_if) { // Serial ATA
0x00 => "Vendor Specific Interface",
0x01 => "AHCI 1.0",
0x02 => "Serial Storage Bus",
else => "",
},
0x08 => switch (prog_if) { // Non-Volatile Memory
0x01 => "NVMHCI",
0x02 => "NVM Express",
else => "",
},
else => "",
const named: ?[]const u8 = switch (@as(BaseClass, @enumFromInt(base))) {
.mass_storage => switch (std.enums.fromInt(mass_storage.SubClass, subclass) orelse return "") {
.serial_ata => enumName(mass_storage.serial_ata.ProgIf, prog_if),
.non_volatile_memory => enumName(mass_storage.non_volatile_memory.ProgIf, prog_if),
else => null,
},
0x03 => switch (subclass) {
0x00 => switch (prog_if) { // VGA Compatible
0x00 => "VGA Controller",
0x01 => "8514-Compatible Controller",
else => "",
},
else => "",
.display => switch (std.enums.fromInt(display.SubClass, subclass) orelse return "") {
.vga_compatible => enumName(display.vga_compatible.ProgIf, prog_if),
else => null,
},
0x06 => switch (subclass) {
0x04 => switch (prog_if) { // PCI-to-PCI Bridge
0x00 => "Normal Decode",
0x01 => "Subtractive Decode",
else => "",
},
else => "",
.bridge => switch (std.enums.fromInt(bridge.SubClass, subclass) orelse return "") {
.pci_to_pci => enumName(bridge.pci_to_pci.ProgIf, prog_if),
else => null,
},
0x07 => switch (subclass) {
0x00 => switch (prog_if) { // Serial Controller
0x00 => "8250-Compatible (Generic XT)",
0x01 => "16450-Compatible",
0x02 => "16550-Compatible",
0x03 => "16650-Compatible",
0x04 => "16750-Compatible",
0x05 => "16850-Compatible",
0x06 => "16950-Compatible",
else => "",
},
else => "",
.simple_communication => switch (std.enums.fromInt(simple_communication.SubClass, subclass) orelse return "") {
.serial => enumName(simple_communication.serial.ProgIf, prog_if),
else => null,
},
0x0C => switch (subclass) {
0x03 => switch (prog_if) { // USB Controller
0x00 => "UHCI Controller",
0x10 => "OHCI Controller",
0x20 => "EHCI (USB2) Controller",
0x30 => "XHCI (USB3) Controller",
0x80 => "Unspecified",
0xFE => "USB Device (not a host controller)",
else => "",
},
else => "",
.serial_bus => switch (std.enums.fromInt(serial_bus.SubClass, subclass) orelse return "") {
.usb => enumName(serial_bus.usb.ProgIf, prog_if),
else => null,
},
else => "",
else => null,
};
return named orelse "";
}
test "decodes the common class codes" {
const std = @import("std");
const eq = std.testing.expectEqualStrings;
const isa = ClassCode.unpack(0x06_01_00);
@@ -251,11 +544,25 @@ test "decodes the common class codes" {
try eq("AHCI 1.0", progIfName(ahci.base, ahci.subclass, ahci.prog_if));
const xhci = ClassCode.unpack(0x0C_03_30);
try eq("Serial Bus Controller", className(xhci.base));
try eq("USB Controller", subclassName(xhci.base, xhci.subclass));
try eq("XHCI (USB3) Controller", progIfName(xhci.base, xhci.subclass, xhci.prog_if));
}
// Unknowns and the "Other" convention.
try eq("Other", subclassName(0x02, 0x80));
try eq("Unknown", subclassName(0x06, 0x7E));
try eq("", progIfName(0x06, 0x00, 0x00)); // host bridge: prog-IF has no standard name
test "unlisted codes fall back without a wrong name" {
const eq = std.testing.expectEqualStrings;
try eq("Unknown", className(0x77)); // no such base class
try eq("Other", subclassName(0x01, 0x80)); // 0x80 is the PCI "Other" convention
try eq("Unknown", subclassName(0x01, 0x7A)); // unlisted mass-storage subclass
try eq("", progIfName(0x01, 0x06, 0x7F)); // no standard SATA prog-IF for 0x7F
try eq("", progIfName(0x02, 0x00, 0x00)); // class with no prog-IF taxonomy at all
}
test "named parts pack to the raw triple" {
const xhci = ClassCode{
.base = @intFromEnum(BaseClass.serial_bus),
.subclass = @intFromEnum(serial_bus.SubClass.usb),
.prog_if = @intFromEnum(serial_bus.usb.ProgIf.xhci),
};
try std.testing.expectEqual(@as(u24, 0x0C_03_30), xhci.pack());
}
+9 -1
View File
@@ -40,6 +40,13 @@ pub fn platformInformation() PlatformInformation {
}
/// AML parse integrity/diagnostics (namespace node count, bytes consumed).
/// The number of Device objects in the kernel's own AML namespace, or 0 if the
/// parse produced none — the `acpi-parse` test compares the ring-3 service's
/// count against this.
pub fn amlDeviceCount() usize {
return acpi.amlDeviceCount();
}
pub fn amlStats() AmlStats {
return acpi.aml_stats;
}
@@ -72,7 +79,8 @@ pub fn discover(
var device_tree = try DeviceTree.init(allocator);
if (boot_information.acpi_rsdp != 0) {
try acpi.discover(boot_information.acpi_rsdp, &device_tree, hal);
const memory_regions = @as([*]const boot_handoff.MemoryRegion, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.memory_map.regions)))[0..boot_information.memory_map.len];
try acpi.discover(boot_information.acpi_rsdp, memory_regions, &device_tree, hal);
} else {
// No ACPI RSDP. A device-tree boot would parse its blob here; today that
// path is a stub, so this reports the machine described itself no way we
+857
View File
@@ -0,0 +1,857 @@
//! USB device-framework wire ABI: the set-up packets, standard requests, and standard
//! descriptors every USB device speaks over its default control pipe, as defined by chapter 9
//! of the USB 2.0 specification (see https://wiki.osdev.org/Universal_Serial_Bus). Pure data
//! definitions — no hardware access — shared by the host-controller bus drivers (which build
//! the requests) and anything that parses what devices return (device naming, driver
//! matching, configuration). The structs mirror the wire byte-for-byte: multi-byte fields are
//! little-endian and align(1), so a descriptor can be bit-cast straight out of a transfer
//! buffer at any offset, and bitmap bytes are packed structs so no caller ever needs a magic
//! mask. Class, subclass, and protocol code tables live in usb-ids.zig.
const DeviceState = enum(u8) {
// Immediately after the USB device is attached to the USB system, it is in this state.
// The USB specifications do not define the state of a USB device that is detached from
// a USB system.
attached,
// A device is in this state after it has both been attached to the bus, and the VBUS line is
// applied to the device (the host controller drives the VBUS at +5V, however this is only
// particularly important for hardware developers). In this state, the device must not respond
// to any bus transactions. The USB specification recognizes three potential scenarios with
// respect to how a device draws power:
// - Self-Powered Devices draw power from an external power source (e.g, a USB printer plugs
// into the wall as well as a USB port). Although the device may be considered
// technically "powered" even before attachment to the USB, it is still only considered
// powered after the VBUS line is applied to the device.
// - Bus-Powered Devices draw power solely from the USB up to 100mA.
// - Self- or Bus-Powered Devices may draw power from either the bus or an external power
// source, depending on the configuration. These devices may change power source at any
// time. If a device is currently self-powered and requires more than 100mA of power, but
// switches to being bus-powered, then the device must return to the Address state.
powered,
// A device in the powered state enters the default state after receiving a bus reset. In this
// state, the device is addressable at the default, reserved address of 0. At this point, the
// device is operating at the correct speed. The host is expected to allow 10 milliseconds
// before expecting the device to respond to data transfers after reset.
default,
// A device enters this state after the host assigns it an address via the default control pipe,
// which is always accessible whether the device's address has been set or not.
address,
// A device is in this state after the host examines its possible configurations and selects
// one. All endpoint's data toggle bits are initialized to zero when a device enters this state.
configured,
// When no traffic is observed on the bus for a period of 1 millisecond, a USB device enters
// this state, characterized by its low power consumption. The device's address and
// configuration settings are maintained while suspended. A device exits the suspended state as
// soon as it begins seeing bus activity again. The host is expected to allow 10 milliseconds
// before expecting the device to respond to data transfers after resume.
suspended,
};
const RequestCode = enum(u8) {
get_status = 0,
clear_feature = 1,
set_feature = 3,
set_address = 5,
get_descriptor = 6,
set_descriptor = 7,
get_configuration = 8,
set_configuration = 9,
get_interface = 10,
set_interface = 11,
sync_frame = 12,
};
// Direction of an endpoint, from the host's point of view
const EndpointDirection = enum(u1) {
out = 0,
in = 1,
};
// Identifier newtypes: distinct wire-sized types for values that identify something on the
// device rather than count something. Each is a non-exhaustive enum whose values originate
// in the descriptors below and flow, still typed, into the standard request constructors —
// so an interface number can never be passed where a configuration value is expected.
// The bus address of a device, assigned by the host with SET_ADDRESS. Addresses are 7 bits
// wide.
const DeviceAddress = enum(u7) {
// The default address every device answers at after a reset, until SET_ADDRESS
// completes
default = 0,
_,
};
// Identifies a configuration; from ConfigurationDescriptor.configuration_value.
const ConfigurationValue = enum(u8) {
// Not configured: returned by GET_CONFIGURATION while the device is in the address
// state, and passed to SET_CONFIGURATION to return a configured device to the address
// state
none = 0,
_,
};
// Identifies an interface within a configuration; from
// InterfaceDescriptor.interface_number.
const InterfaceNumber = enum(u8) { _ };
// Selects between the alternate settings of one interface; from
// InterfaceDescriptor.alternate_setting.
const AlternateSetting = enum(u8) {
// The default setting of an interface
default = 0,
_,
};
// The number of an endpoint within a device, 4 bits wide. The direction bit carried
// alongside it tells the two endpoints sharing a number apart.
const EndpointNumber = enum(u4) {
// Endpoint zero: the default control pipe every device provides
default_control = 0,
_,
};
// Index of a STRING descriptor, stored in descriptors that reference a string and passed to
// GET_DESCRIPTOR to read it.
const StringIndex = enum(u8) {
// The device has no string descriptor for this field
none = 0,
_,
};
// Characteristics of a device request (the bmRequestType field of a set-up packet). Fields are
// declared least-significant first: recipient occupies bits 4...0, kind bits 6...5, and
// direction bit 7.
const RequestType = packed struct(u8) {
// The recipient of the request (values 4...31 are reserved)
recipient: Recipient,
// The type of the request
kind: Kind,
// Data transfer direction. The value of this bit is ignored when length is zero.
direction: Direction,
const Recipient = enum(u5) {
device = 0,
interface = 1,
endpoint = 2,
other = 3,
};
const Kind = enum(u2) {
standard = 0,
class = 1,
vendor = 2,
reserved = 3,
};
const Direction = enum(u1) {
host_to_device = 0,
device_to_host = 1,
};
};
const Request = extern struct {
// Characteristics of the request
request_type: RequestType,
// Specific request
request_code: RequestCode,
// Word-sized field that may (or may not) serve as a parameter to the request, depending
// on the specific request. For GET_DESCRIPTOR and SET_DESCRIPTOR, bit-cast a
// DescriptorValue into this field.
value: u16 align(1),
// Word-sized field that may (or may not) serve as a parameter to the request, depending
// on the specific request. Typically this field holds an index or an offset value. When
// request_type specifies an endpoint or an interface as the recipient, bit-cast an
// EndpointIndex or an InterfaceIndex into this field.
index: u16 align(1),
// Number of bytes to transfer if there is a DATA stage.
// - If this field is non-zero, and request_type indicates a transfer from
// device-to-host, then the device must never return more than length bytes of data.
// However, a device may return less.
// - If this field is non-zero, and request_type indicates a transfer from
// host-to-device, then the host must send exactly length bytes of data. If the host
// sends more than length bytes, the behavior of the device is undefined.
length: u16 align(1),
// The format of the index field when request_type specifies an endpoint as the
// recipient. The host should always set the direction bit to zero (but the device
// should accept either value) when the endpoint is part of a control pipe.
const EndpointIndex = packed struct(u16) {
// Endpoint number
number: EndpointNumber,
// Reserved (reset to zero)
reserved: u3 = 0,
// Selects the OUT or the IN endpoint with the specified endpoint number
direction: EndpointDirection,
// Reserved (reset to zero)
reserved_high: u8 = 0,
};
// The format of the index field when request_type specifies an interface as the
// recipient.
const InterfaceIndex = packed struct(u16) {
// Interface number
number: u8,
// Reserved (reset to zero)
reserved: u8 = 0,
};
// The format of the value field of GET_DESCRIPTOR and SET_DESCRIPTOR requests: the
// descriptor type in the high byte, and the descriptor index in the low byte. The index
// is used to select a specific descriptor (only for CONFIGURATION and STRING
// descriptors) when several descriptors of that type are implemented by a device.
const DescriptorValue = packed struct(u16) {
// Descriptor index
index: u8 = 0,
// Descriptor type
kind: DescriptorType,
};
};
// Feature selectors, used as the value field of CLEAR_FEATURE and SET_FEATURE requests. The
// comment on each value notes the recipient the selector applies to.
const FeatureSelector = enum(u16) {
// Halts an endpoint (recipient: endpoint)
endpoint_halt = 0,
// Enables or disables the device's remote wakeup capability (recipient: device)
device_remote_wakeup = 1,
// Puts a hi-speed device into a test mode, selected by a TestMode value in the high
// byte of the index field (recipient: device)
test_mode = 2,
};
// Test mode selectors, passed in the high byte of the index field of a SET_FEATURE request
// with the test_mode feature selector. Values 06h...3Fh are reserved for standard test
// selectors and C0h...FFh for vendor-specific test modes; all other unlisted values are
// reserved.
const TestMode = enum(u8) {
test_j = 0x01,
test_k = 0x02,
test_se0_nak = 0x03,
test_packet = 0x04,
test_force_enable = 0x05,
_,
};
// The two bytes returned by a GET_STATUS request directed at a device. Fields are declared
// least-significant first.
const DeviceStatus = packed struct(u16) {
// Whether the device is currently self-powered (as opposed to bus-powered). This bit
// cannot be changed with the SET_FEATURE or CLEAR_FEATURE requests.
self_powered: bool,
// Whether the device is currently enabled to request remote wakeup. Changed with the
// SET_FEATURE and CLEAR_FEATURE requests using the device_remote_wakeup feature
// selector.
remote_wakeup: bool,
// Reserved (reset to zero)
reserved: u14,
};
// The two bytes returned by a GET_STATUS request directed at an endpoint. (A GET_STATUS
// request directed at an interface returns two bytes that are entirely reserved.)
const EndpointStatus = packed struct(u16) {
// Whether the endpoint is currently halted. Set with the SET_FEATURE request using the
// endpoint_halt feature selector, and cleared with CLEAR_FEATURE.
halted: bool,
// Reserved (reset to zero)
reserved: u15,
};
// A target for the standard requests that may be directed at the device, an interface, or
// an endpoint.
const Target = union(enum) {
device,
interface: InterfaceNumber,
endpoint: Request.EndpointIndex,
fn recipient(target: Target) RequestType.Recipient {
return switch (target) {
.device => .device,
.interface => .interface,
.endpoint => .endpoint,
};
}
fn index(target: Target) u16 {
return switch (target) {
.device => 0,
.interface => |number| @intFromEnum(number),
.endpoint => |endpoint| @bitCast(endpoint),
};
}
};
// Constructors for the standard device requests, one per RequestCode. Each returns a
// ready-to-send set-up packet with the request_type, value, index, and length fields the
// specification prescribes for that request.
// Reads the status of the given target: bit-cast the two bytes the device returns into a
// DeviceStatus or an EndpointStatus. (The two bytes returned for an interface are entirely
// reserved.)
fn getStatus(target: Target) Request {
return .{
.request_type = .{
.recipient = target.recipient(),
.kind = .standard,
.direction = .device_to_host,
},
.request_code = .get_status,
.value = 0,
.index = target.index(),
.length = 2,
};
}
// Clears or disables the given feature. A device cannot be taken out of a test mode with
// this request; test_mode is only cleared by cycling power.
fn clearFeature(feature: FeatureSelector, target: Target) Request {
return .{
.request_type = .{
.recipient = target.recipient(),
.kind = .standard,
.direction = .host_to_device,
},
.request_code = .clear_feature,
.value = @intFromEnum(feature),
.index = target.index(),
.length = 0,
};
}
// Sets or enables the given feature. For the test_mode feature selector, use setTestMode
// instead: the test selector rides in the high byte of the index field.
fn setFeature(feature: FeatureSelector, target: Target) Request {
return .{
.request_type = .{
.recipient = target.recipient(),
.kind = .standard,
.direction = .host_to_device,
},
.request_code = .set_feature,
.value = @intFromEnum(feature),
.index = target.index(),
.length = 0,
};
}
// Puts a hi-speed device into the given test mode: a SET_FEATURE request with the test_mode
// feature selector and the test selector in the high byte of the index field.
fn setTestMode(mode: TestMode) Request {
return .{
.request_type = .{
.recipient = .device,
.kind = .standard,
.direction = .host_to_device,
},
.request_code = .set_feature,
.value = @intFromEnum(FeatureSelector.test_mode),
.index = @as(u16, @intFromEnum(mode)) << 8,
.length = 0,
};
}
// Assigns the device its bus address, moving it from the default state to the address
// state. The device does not answer at the new address until the status stage of this
// request completes.
fn setAddress(address: DeviceAddress) Request {
return .{
.request_type = .{
.recipient = .device,
.kind = .standard,
.direction = .host_to_device,
},
.request_code = .set_address,
.value = @intFromEnum(address),
.index = 0,
.length = 0,
};
}
// Reads a descriptor from the device.
// - descriptor_index selects among descriptors of the same type, and is only used for
// configuration and string descriptors.
// - language_id selects the language of a string descriptor, and is zero otherwise.
// - length is the number of bytes to read; a device never returns more than length bytes,
// but may return less if the descriptor is shorter.
fn getDescriptor(kind: DescriptorType, descriptor_index: u8, language_id: u16, length: u16) Request {
return .{
.request_type = .{
.recipient = .device,
.kind = .standard,
.direction = .device_to_host,
},
.request_code = .get_descriptor,
.value = @bitCast(Request.DescriptorValue{ .index = descriptor_index, .kind = kind }),
.index = language_id,
.length = length,
};
}
// Updates an existing descriptor or adds a new one (optional; many devices do not support
// this request). The parameters mirror getDescriptor; the descriptor itself is sent in the
// DATA stage.
fn setDescriptor(kind: DescriptorType, descriptor_index: u8, language_id: u16, length: u16) Request {
return .{
.request_type = .{
.recipient = .device,
.kind = .standard,
.direction = .host_to_device,
},
.request_code = .set_descriptor,
.value = @bitCast(Request.DescriptorValue{ .index = descriptor_index, .kind = kind }),
.index = language_id,
.length = length,
};
}
// Reads the currently active configuration: @enumFromInt the byte the device returns into a
// ConfigurationValue, which is none while the device is not configured.
fn getConfiguration() Request {
return .{
.request_type = .{
.recipient = .device,
.kind = .standard,
.direction = .device_to_host,
},
.request_code = .get_configuration,
.value = 0,
.index = 0,
.length = 1,
};
}
// Selects the configuration with the given configuration_value (from
// ConfigurationDescriptor.configuration_value), moving the device from the address state to
// the configured state. Selecting none returns the device to the address state.
fn setConfiguration(configuration_value: ConfigurationValue) Request {
return .{
.request_type = .{
.recipient = .device,
.kind = .standard,
.direction = .host_to_device,
},
.request_code = .set_configuration,
.value = @intFromEnum(configuration_value),
.index = 0,
.length = 0,
};
}
// Reads the alternate setting currently selected for the given interface: @enumFromInt the
// byte the device returns into an AlternateSetting.
fn getInterface(interface: InterfaceNumber) Request {
return .{
.request_type = .{
.recipient = .interface,
.kind = .standard,
.direction = .device_to_host,
},
.request_code = .get_interface,
.value = 0,
.index = @intFromEnum(interface),
.length = 1,
};
}
// Selects an alternate setting (from InterfaceDescriptor.alternate_setting) for the given
// interface.
fn setInterface(interface: InterfaceNumber, alternate_setting: AlternateSetting) Request {
return .{
.request_type = .{
.recipient = .interface,
.kind = .standard,
.direction = .host_to_device,
},
.request_code = .set_interface,
.value = @intFromEnum(alternate_setting),
.index = @intFromEnum(interface),
.length = 0,
};
}
// Reads the two-byte number of the frame in which the given isochronous endpoint's
// repeating pattern of transfers begins.
fn syncFrame(endpoint: Request.EndpointIndex) Request {
return .{
.request_type = .{
.recipient = .endpoint,
.kind = .standard,
.direction = .device_to_host,
},
.request_code = .sync_frame,
.value = 0,
.index = @bitCast(endpoint),
.length = 2,
};
}
const DescriptorType = enum(u8) {
device = 1,
configuration = 2,
string = 3,
interface = 4,
endpoint = 5,
device_qualifier = 6,
other_speed_configuration = 7,
interface_power = 8,
_,
};
const DeviceDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// DEVICE Descriptor Type
descriptor_type: DescriptorType,
// USB Specification Release Number in Binary-Coded Decimal (i.e, 2.10 is expressed as 210h).
// Identifies the release of the USB Specification with with the device and its
// descriptors are compliant.
bcd_usb: u16 align(1),
// Class code (assigned by the USB-IF)
// - This field is reset to zero if each interface within a configuration specifies its own
// class information and the various interfaces operate independently.
// - A value of FFh in this field indicates the device class is vendor-specific.
device_class: u8,
// Subclass Code (assigned by the USB-IF)
// - The subclass code of a device is qualified by the class code of that device.
// - If device_class is reset to zero, then this field must also be reset to zero.
// - When device_class is not set to FFh, then all values for this field are reserved for
// assignment by the USB-IF.
device_subclass: u8,
// Protocol code (assigned by the USB-IF)
// - The protocol code of a device is qualified by both the class and subclass codes of
// that device.
// - A value of 00h in this field means that the device may specify class-specific
// protocols on an interface basis, though this is not a requirement.
// - If this field is set to FFh, then the device uses a vendor-specific protocol.
device_protocol: u8,
// Maximum packet size for endpoint zero (8, 16, 32, or 64 are the only valid options)
max_packet_size_0: u8,
// Vendor ID (assigned by the USB-IF)
vendor_id: u16 align(1),
// Product ID (assigned by the USB-IF)
product_id: u16 align(1),
// Device release number in binary-coded decimal
bcd_device: u16 align(1),
// Index of STRING descriptor describing manufacturer
manufacturer_index: StringIndex,
// Index of STRING descriptor describing product
product_index: StringIndex,
// Index of STRING descriptor describing the device's serial number
serial_number_index: StringIndex,
// Number of possible configurations
configuration_count: u8,
};
const DeviceQualifierDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// DEVICE_QUALIFIER Descriptor Type
descriptor_type: DescriptorType,
// USB Specification Release Number in Binary-Coded Decimal (i.e, 2.00 is expressed as 200h).
// Identifies the release of the USB Specification with with the device and its
// descriptors are compliant. This field must be at least 0200h.
bcd_usb: u16 align(1),
// Class code (assigned by the USB-IF)
device_class: u8,
// Subclass Code (assigned by the USB-IF)
device_subclass: u8,
// Protocol code (assigned by the USB-IF)
device_protocol: u8,
// Maximum packet size for endpoint zero (8, 16, 32, or 64 are the only valid options)
max_packet_size_0: u8,
// Number of possible configurations
configuration_count: u8,
// Reserved for future uses, must be zero.
reserved: u8,
};
const ConfigurationDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// CONFIGURATION Descriptor Type
descriptor_type: DescriptorType,
// The total combined length in bytes of all the descriptors returned with the request for
// this CONFIGURATION descriptor (including CONFIGURATION, INTERFACE, ENDPOINT, class- and
// vendor-specific descriptors).
total_length: u16 align(1),
// Number of interfaces supported by this configuration
interface_count: u8,
// Value which when used as an argument in the SET_CONFIGURATION request, causes the device
// to assume the configuration described by this descriptor.
configuration_value: ConfigurationValue,
// Index of STRING descriptor describing this configuration.
configuration_index: StringIndex,
// Configuration Characteristics
attributes: Attributes,
// Maximum power consumption of this device from the bus when fully operational and using
// this configuration. Expressed in units of 2mA (i.e., a value of 50 in this field
// indicates 100mA).
// - A device reports with the attributes field whether the configuration is bus- or
// self-powered, but the device status (retrieved with a GET_STATUS request) reports
// whether the device is currently self-powered.
// - If a device is disconnected from an external power source, it may not draw more
// power from the bus than specified in this field.
max_power: u8,
// Configuration characteristics. Fields are declared least-significant first.
const Attributes = packed struct(u8) {
// Reserved, reset to zero (D4...0)
reserved: u5,
// Whether Remote Wakeup is supported by this configuration (D5)
remote_wakeup: bool,
// Self-Powered (D6)
// - false: Device runs on power supplied by the bus
// - true: Device provides a local power source; if max_power is non-zero, the
// device also may use bus power.
self_powered: bool,
// Reserved, must be set to one for historical reasons (D7)
reserved_one: u1,
};
};
// This descriptor describes the configuration of a high-speed device if it were operating at
// its alternative speed. The structure of the OTHER_SPEED_CONFIGURATION is identical to that
// of the CONFIGURATION descriptor; the only difference is that the descriptor_type field
// reflects that the descriptor is an OTHER_SPEED_CONFIGURATION descriptor.
const OtherSpeedConfigurationDescriptor = ConfigurationDescriptor;
const InterfaceDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// INTERFACE Descriptor Type
descriptor_type: DescriptorType,
// Number of this interface. Zero-based value which identifies the index of this interface
// in the array of interfaces supported within a configuration.
interface_number: InterfaceNumber,
// Value used to select the alternate settings described by this INTERFACE descriptor for
// the interface with the interface_number in the previous field. This value is zero if
// this descriptor describes the default settings for a particular interface.
alternate_setting: AlternateSetting,
// Number of endpoints used by this interface, not including endpoint zero.
endpoint_count: u8,
// Class code (assigned by the USB-IF)
// - A value of zero here is reserved for future standardization.
// - If this value is FFh, the interface class is vendor-specific.
// - All other values are reserved for assignment by the USB-IF.
interface_class: u8,
// Subclass code (assigned by the USB-IF)
// - The subclass code in this field is qualified by the value of the interface_class
// field.
// - If interface_class is reset to zero, then this field must also be reset to zero.
// - If interface_class is not set to the value of FFh, then all values of this field are
// reserved for assignment by the USB-IF.
interface_subclass: u8,
// Protocol code (assigned by the USB-IF)
// - The protocol code in this field is qualified by the values of the interface_class
// and interface_subclass fields.
// - If an interface supports class-specific requests, then this field identifies the
// protocols that the device uses as defined by the specifications of the device class.
// - If this field is reset to zero, then the device does not use a class-specific
// protocol on this interface.
// - If this field is set to FFh, then the device uses a vendor-specific protocol on
// this interface.
interface_protocol: u8,
// Index of STRING descriptor describing this interface
interface_index: StringIndex,
};
const EndpointDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// ENDPOINT Descriptor Type
descriptor_type: DescriptorType,
// The address of the endpoint on the USB device described by this descriptor
endpoint_address: Address,
// The endpoint's attributes
attributes: Attributes,
// Maximum packet size that this endpoint is capable of sending or receiving. For
// isochronous endpoints, this value is used to reserve bus time; the pipe, however, may
// not always use all of the reserved bus time.
max_packet_size: MaxPacketSize align(1),
// Interval for polling a device during a data transfer, expressed in units of microframes
// for high-speed devices, and frames for low- and full-speed devices. The exact meaning of
// the value in this field depends on the endpoint type and the operating speed of the
// device:
// - Full- and High-speed isochronous endpoints, and high-speed interrupt endpoints:
// This field must be in the range from 1 to 16, and is used to calculate the period
// as 2^(interval - 1). That is, a value of 4 calculates to 2^(4 - 1) = 2^3 = 8.
// - Full- and Low-speed interrupt endpoints: This field must be in the range from
// 1 to 255.
// - High-speed bulk and control OUT endpoints: This field must be in the range from
// 0 to 255, and specifies the maximum NAK rate of the endpoint. A value of zero
// indicates that the endpoint never NAKs; other values indicate at most 1 NAK each
// interval number of microframes.
interval: u8,
// The address of an endpoint. Fields are declared least-significant first.
const Address = packed struct(u8) {
// Endpoint Number (D3...0)
number: EndpointNumber,
// Reserved, reset to zero (D6...4)
reserved: u3,
// Direction, ignored for control endpoints (D7)
direction: EndpointDirection,
};
// An endpoint's attributes. Fields are declared least-significant first.
const Attributes = packed struct(u8) {
// Transfer Type (D1...0)
transfer_type: TransferType,
// Synchronization Type; isochronous endpoints only, reserved and reset to zero for
// other endpoint types (D3...2)
synchronization: Synchronization,
// Usage Type; isochronous endpoints only, reserved and reset to zero for other
// endpoints (D5...4)
usage: Usage,
// Reserved, reset to zero (D7...6)
reserved: u2,
};
const TransferType = enum(u2) {
control = 0,
isochronous = 1,
bulk = 2,
interrupt = 3,
};
const Synchronization = enum(u2) {
none = 0,
asynchronous = 1,
adaptive = 2,
synchronous = 3,
};
const Usage = enum(u2) {
data = 0,
feedback = 1,
implicit_feedback_data = 2,
_,
};
// The maximum packet size of an endpoint. Fields are declared least-significant first.
const MaxPacketSize = packed struct(u16) {
// Maximum packet size in bytes (bits 10...0)
size: u11,
// Number of additional transaction opportunities per microframe, for high-speed
// isochronous and interrupt endpoints; reserved and reset to zero for other
// endpoints (bits 12...11)
additional_transactions: AdditionalTransactions,
// Reserved, must be reset to zero (bits 15...13)
reserved: u3,
};
const AdditionalTransactions = enum(u2) {
// None (1 transaction per microframe)
none = 0,
// 1 additional (2 transactions per microframe)
one = 1,
// 2 additional (3 transactions per microframe)
two = 2,
_,
};
};
// A STRING descriptor at index zero returns the list of LANGID codes supported by the
// device; all other indices return a Unicode string. Both forms start with this two-byte
// header, followed by the variable-length payload:
// - index 0: an array of two-byte LANGID codes (wLangID[0] through wLangID[x])
// - other indices: a Unicode string of N bytes
const StringDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// STRING Descriptor Type
descriptor_type: DescriptorType,
};
const std = @import("std");
test "wire sizes and offsets match the specification" {
const expectEqual = std.testing.expectEqual;
try expectEqual(8, @sizeOf(Request));
try expectEqual(18, @sizeOf(DeviceDescriptor));
try expectEqual(10, @sizeOf(DeviceQualifierDescriptor));
try expectEqual(9, @sizeOf(ConfigurationDescriptor));
try expectEqual(9, @sizeOf(InterfaceDescriptor));
try expectEqual(7, @sizeOf(EndpointDescriptor));
try expectEqual(2, @sizeOf(StringDescriptor));
try expectEqual(2, @offsetOf(DeviceDescriptor, "bcd_usb"));
try expectEqual(8, @offsetOf(DeviceDescriptor, "vendor_id"));
try expectEqual(17, @offsetOf(DeviceDescriptor, "configuration_count"));
try expectEqual(2, @offsetOf(ConfigurationDescriptor, "total_length"));
try expectEqual(4, @offsetOf(EndpointDescriptor, "max_packet_size"));
}
test "bitmap packings match the specification" {
const expectEqual = std.testing.expectEqual;
const expect = std.testing.expect;
// bmRequestType for GET_DESCRIPTOR: device-to-host | standard | device = 80h
const request_type = RequestType{
.recipient = .device,
.kind = .standard,
.direction = .device_to_host,
};
try expectEqual(0x80, @as(u8, @bitCast(request_type)));
// wValue for GET_DESCRIPTOR(CONFIGURATION, index 0) = 0200h
const descriptor_value = Request.DescriptorValue{ .kind = .configuration };
try expectEqual(0x0200, @as(u16, @bitCast(descriptor_value)));
// wIndex for the IN endpoint 1 = 0081h
const endpoint_index = Request.EndpointIndex{ .number = @enumFromInt(1), .direction = .in };
try expectEqual(0x0081, @as(u16, @bitCast(endpoint_index)));
// Endpoint address 81h = IN endpoint 1
const address: EndpointDescriptor.Address = @bitCast(@as(u8, 0x81));
try expectEqual(1, @intFromEnum(address.number));
try expectEqual(.in, address.direction);
// Endpoint attributes 03h = interrupt transfer
const attributes: EndpointDescriptor.Attributes = @bitCast(@as(u8, 0x03));
try expectEqual(.interrupt, attributes.transfer_type);
// wMaxPacketSize 0008h = 8 bytes, no additional transactions
const max_packet_size: EndpointDescriptor.MaxPacketSize = @bitCast(@as(u16, 0x0008));
try expectEqual(8, max_packet_size.size);
try expectEqual(.none, max_packet_size.additional_transactions);
// Configuration attributes C0h = self-powered, with the historical D7 bit set
const configuration_attributes: ConfigurationDescriptor.Attributes = @bitCast(@as(u8, 0xC0));
try expect(configuration_attributes.self_powered);
try expect(!configuration_attributes.remote_wakeup);
try expectEqual(1, configuration_attributes.reserved_one);
// GET_STATUS words: device 0001h = self-powered; endpoint 0001h = halted
const device_status: DeviceStatus = @bitCast(@as(u16, 0x0001));
try expect(device_status.self_powered and !device_status.remote_wakeup);
const endpoint_status: EndpointStatus = @bitCast(@as(u16, 0x0001));
try expect(endpoint_status.halted);
// DescriptorType is non-exhaustive: class-specific values (HID = 21h) pass through
const hid_type: DescriptorType = @enumFromInt(0x21);
try expectEqual(0x21, @intFromEnum(hid_type));
try expect(hid_type != .device);
}
fn expectRequestBytes(request: Request, expected: [8]u8) !void {
try std.testing.expectEqualSlices(u8, &expected, std.mem.asBytes(&request));
}
test "standard request constructors encode the specification's set-up packets" {
try expectRequestBytes(getStatus(.device), .{ 0x80, 0, 0, 0, 0, 0, 2, 0 });
try expectRequestBytes(getStatus(.{ .interface = @enumFromInt(3) }), .{ 0x81, 0, 0, 0, 3, 0, 2, 0 });
try expectRequestBytes(getStatus(.{ .endpoint = .{ .number = @enumFromInt(2), .direction = .in } }), .{ 0x82, 0, 0, 0, 0x82, 0, 2, 0 });
try expectRequestBytes(clearFeature(.endpoint_halt, .{ .endpoint = .{ .number = @enumFromInt(1), .direction = .out } }), .{ 0x02, 1, 0, 0, 0x01, 0, 0, 0 });
try expectRequestBytes(setFeature(.device_remote_wakeup, .device), .{ 0x00, 3, 1, 0, 0, 0, 0, 0 });
try expectRequestBytes(setTestMode(.test_packet), .{ 0x00, 3, 2, 0, 0, 0x04, 0, 0 });
try expectRequestBytes(setAddress(@enumFromInt(5)), .{ 0x00, 5, 5, 0, 0, 0, 0, 0 });
try expectRequestBytes(getDescriptor(.device, 0, 0, 18), .{ 0x80, 6, 0, 1, 0, 0, 18, 0 });
try expectRequestBytes(getDescriptor(.string, 2, 0x0409, 255), .{ 0x80, 6, 2, 3, 0x09, 0x04, 255, 0 });
try expectRequestBytes(setDescriptor(.string, 2, 0x0409, 16), .{ 0x00, 7, 2, 3, 0x09, 0x04, 16, 0 });
try expectRequestBytes(getConfiguration(), .{ 0x80, 8, 0, 0, 0, 0, 1, 0 });
try expectRequestBytes(setConfiguration(@enumFromInt(1)), .{ 0x00, 9, 1, 0, 0, 0, 0, 0 });
try expectRequestBytes(getInterface(@enumFromInt(2)), .{ 0x81, 10, 0, 0, 2, 0, 1, 0 });
try expectRequestBytes(setInterface(@enumFromInt(2), @enumFromInt(1)), .{ 0x01, 11, 1, 0, 2, 0, 0, 0 });
try expectRequestBytes(syncFrame(.{ .number = @enumFromInt(3), .direction = .in }), .{ 0x82, 12, 0, 0, 0x83, 0, 2, 0 });
}
+264
View File
@@ -0,0 +1,264 @@
//! USB class-code decoding: turn the (class, subclass, protocol) triple a USB device or
//! interface reports in its descriptors into typed values. The device descriptor carries one
//! triple for the whole device, and each interface descriptor carries its own; a class code
//! of zero at the device level defers entirely to the interfaces. Subclass and protocol
//! codes are qualified by the class code — the same value means different things under
//! different classes — so there is no single SubClass or Protocol enum: each class with
//! spec-defined codes gets its own namespace below. Pure reference data (from the USB-IF
//! defined class codes; see https://www.usb.org/defined-class-codes) — no hardware access —
//! so it is shared by kernel discovery and any user-space tool (device naming, driver
//! matching).
// Base class codes (assigned by the USB-IF). The comment on each value notes where the code
// may legally appear: in the device descriptor, in interface descriptors, or both.
const Class = enum(u8) {
// Use class information in the interface descriptors (device descriptor only). Each
// interface within a configuration specifies its own class information and the various
// interfaces operate independently.
per_interface = 0x00,
// Audio: speakers, microphones, sound cards (interface)
audio = 0x01,
// Communications and CDC control: modems, network adapters (both)
communications = 0x02,
// Human Interface Device: keyboards, mice, game controllers (interface)
hid = 0x03,
// Physical: force-feedback devices (interface)
physical = 0x05,
// Image: still-imaging cameras, scanners (interface)
image = 0x06,
// Printer (interface)
printer = 0x07,
// Mass storage: flash drives, external disks, card readers (interface)
mass_storage = 0x08,
// Hub (device descriptor only)
hub = 0x09,
// CDC-Data: the data interfaces paired with a communications control interface
// (interface)
cdc_data = 0x0A,
// Smart card readers (interface)
smart_card = 0x0B,
// Content security (interface)
content_security = 0x0D,
// Video: webcams (interface)
video = 0x0E,
// Personal healthcare devices (interface)
personal_healthcare = 0x0F,
// Audio/Video devices (interface)
audio_video = 0x10,
// Billboard: describes alternate modes a USB Type-C device supports (device descriptor
// only)
billboard = 0x11,
// USB Type-C bridge (interface)
type_c_bridge = 0x12,
// USB Bulk Display Protocol devices (interface)
bulk_display = 0x13,
// MCTP over USB protocol endpoint devices (interface)
mctp = 0x14,
// I3C devices (interface)
i3c = 0x3C,
// Diagnostic devices (both)
diagnostic = 0xDC,
// Wireless controllers: Bluetooth adapters (interface)
wireless_controller = 0xE0,
// Miscellaneous (both)
miscellaneous = 0xEF,
// Application-specific: firmware upgrade, IrDA bridges, test and measurement
// (interface)
application_specific = 0xFE,
// Vendor-specific (both)
vendor_specific = 0xFF,
_,
};
// Subclass and protocol codes qualified by Class.hub. Hubs have no subclass codes; the
// protocol distinguishes the hub's transaction-translator arrangement.
const hub = struct {
const Protocol = enum(u8) {
// Full-speed hub
full_speed = 0x00,
// Hi-speed hub with a single transaction translator
hi_speed_single_tt = 0x01,
// Hi-speed hub with multiple transaction translators
hi_speed_multi_tt = 0x02,
// SuperSpeed hub (USB 3)
super_speed = 0x03,
_,
};
};
// Subclass and protocol codes qualified by Class.hid.
const hid = struct {
const SubClass = enum(u8) {
// No subclass
none = 0x00,
// Boot interface: the device also supports the simplified boot protocol, usable by
// firmware before a full HID report-descriptor parser is available
boot = 0x01,
_,
};
// Only meaningful when the subclass is boot
const Protocol = enum(u8) {
none = 0x00,
keyboard = 0x01,
mouse = 0x02,
_,
};
};
// Subclass and protocol codes qualified by Class.mass_storage. The subclass identifies the
// command set the device understands; the protocol identifies the transport used to carry
// commands, data, and status over the bus.
const mass_storage = struct {
const SubClass = enum(u8) {
// SCSI command set not reported; de facto, treat as scsi
not_reported = 0x00,
// Reduced Block Commands: typically flash devices
rbc = 0x01,
// MMC-5 (ATAPI): CD and DVD drives
atapi = 0x02,
// QIC-157 tape drives (obsolete)
qic_157 = 0x03,
// UFI: floppy disk drives
ufi = 0x04,
// SFF-8070i (obsolete)
sff_8070i = 0x05,
// Transparent SCSI command set: the common case for flash drives and disks
scsi = 0x06,
// LSD FS: negotiated access to large storage devices
lsd_fs = 0x07,
// IEEE 1667
ieee_1667 = 0x08,
// Vendor-specific
vendor_specific = 0xFF,
_,
};
const Protocol = enum(u8) {
// Control/Bulk/Interrupt with command completion interrupt
cbi_completion_interrupt = 0x00,
// Control/Bulk/Interrupt without command completion interrupt
cbi = 0x01,
// Bulk-only transport: the common case for flash drives and disks
bulk_only = 0x50,
// USB attached SCSI
uas = 0x62,
// Vendor-specific
vendor_specific = 0xFF,
_,
};
};
// Subclass and protocol codes qualified by Class.communications (CDC). The protocol codes
// are model-specific; the useful invariant is the subclass, which selects the control model
// the interface implements.
const communications = struct {
const SubClass = enum(u8) {
// Direct line control model
direct_line = 0x01,
// Abstract control model: USB modems and serial adapters
abstract_control = 0x02,
// Telephone control model
telephone = 0x03,
// Multi-channel control model
multi_channel = 0x04,
// CAPI control model
capi = 0x05,
// Ethernet networking control model
ethernet = 0x06,
// ATM networking control model
atm = 0x07,
// Wireless handset control model
wireless_handset = 0x08,
// Device management
device_management = 0x09,
// Mobile direct line model
mobile_direct_line = 0x0A,
// OBEX
obex = 0x0B,
// Ethernet emulation model
ethernet_emulation = 0x0C,
// Network control model
network_control = 0x0D,
_,
};
};
// Subclass and protocol codes qualified by Class.wireless_controller.
const wireless_controller = struct {
const SubClass = enum(u8) {
// Radio frequency controllers
radio_frequency = 0x01,
_,
};
// Only meaningful when the subclass is radio_frequency
const Protocol = enum(u8) {
// Bluetooth programming interface
bluetooth = 0x01,
// Ultra-wideband radio control
ultra_wideband = 0x02,
// Remote NDIS
remote_ndis = 0x03,
// Bluetooth AMP controller
bluetooth_amp = 0x04,
_,
};
};
// Subclass and protocol codes qualified by Class.miscellaneous.
const miscellaneous = struct {
const SubClass = enum(u8) {
// Common class
common = 0x02,
_,
};
// Only meaningful when the subclass is common
const Protocol = enum(u8) {
// Interface association descriptor: at the device level, announces that the
// configuration groups interfaces into functions with IADs
interface_association = 0x01,
_,
};
};
// Subclass and protocol codes qualified by Class.application_specific.
const application_specific = struct {
const SubClass = enum(u8) {
// Device firmware upgrade
firmware_upgrade = 0x01,
// IrDA bridge
irda_bridge = 0x02,
// Test and measurement
test_and_measurement = 0x03,
_,
};
};
test "class codes match the USB-IF assignments" {
const std = @import("std");
const expectEqual = std.testing.expectEqual;
try expectEqual(0x03, @intFromEnum(Class.hid));
try expectEqual(0x09, @intFromEnum(Class.hub));
try expectEqual(0xFF, @intFromEnum(Class.vendor_specific));
// A typical flash drive: mass storage, transparent SCSI, bulk-only transport.
try expectEqual(0x06, @intFromEnum(mass_storage.SubClass.scsi));
try expectEqual(0x50, @intFromEnum(mass_storage.Protocol.bulk_only));
// A boot keyboard: HID, boot subclass, keyboard protocol.
try expectEqual(0x01, @intFromEnum(hid.SubClass.boot));
try expectEqual(0x01, @intFromEnum(hid.Protocol.keyboard));
// Class codes are non-exhaustive: unlisted values pass through undamaged.
const unknown: Class = @enumFromInt(0x42);
try expectEqual(0x42, @intFromEnum(unknown));
_ = hub.Protocol.hi_speed_multi_tt;
_ = communications.SubClass.abstract_control;
_ = wireless_controller.Protocol.bluetooth;
_ = miscellaneous.Protocol.interface_association;
_ = application_specific.SubClass.firmware_upgrade;
}
+1
View File
@@ -100,6 +100,7 @@ pub fn main() void {
while (n < n_children) : (n += 1) {
var child = std.mem.zeroes(device.DeviceDescriptor);
child.class = @intFromEnum(device.DeviceClass.timer);
child.pci_class = device.no_pci_class;
child.hid_len = 6;
child.hid[0..6].* = "hpet-t".*;
child.resource_count = 1;
+270
View File
@@ -0,0 +1,270 @@
//! /system/drivers/pci-bus — the PCI bus driver: enumeration moved out of ring 0
//! (docs/m19-m20-plan.md, M19). The device manager matches the `pci_host_bridge`
//! node and spawns one instance per bridge, the bridge's device id as argv[1] —
//! the same per-device contract as usb-xhci-bus.
//!
//! M19.1 (this increment): claim the bridge, map its ECAM window (resource 0;
//! the bus range and the MMIO apertures follow it), walk every
//! bus/device/function config header, and log what the walk finds — ending
//! with "pci-bus: N functions found", which the `pci-scan` scenario compares
//! against the kernel's own enumeration. Registration and reports (M19.2), and
//! the kernel walk's retirement (M19.3), build on this proven-equivalent scan.
const std = @import("std");
const runtime = @import("runtime");
const protocol = runtime.device_manager_protocol;
const device = runtime.device;
const pci_class = @import("pci-class");
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
/// Log a discovered function with its (class / subclass / prog-IF) triple decoded
/// to human names — the boot-log breadcrumb that says *what* the hardware is, so
/// "class 0x01 (Mass Storage Controller) subclass 0x06 (Serial ATA Controller)
/// progif 0x01 (AHCI 1.0)" reads straight off the log when writing a new driver.
/// A dedicated wider buffer than `writeLine`'s, since the decoded names are long.
fn logFunction(bus: u64, dev: u64, function: u64, class_triple: u32) void {
const cc = pci_class.ClassCode.unpack(@truncate(class_triple));
const pif = pci_class.progIfName(cc.base, cc.subclass, cc.prog_if);
var line: [200]u8 = undefined;
const text = if (pif.len != 0)
std.fmt.bufPrint(&line, "pci-bus: {d}:{d}.{d} class 0x{x:0>2} ({s}) subclass 0x{x:0>2} ({s}) progif 0x{x:0>2} ({s})\n", .{ bus, dev, function, cc.base, pci_class.className(cc.base), cc.subclass, pci_class.subclassName(cc.base, cc.subclass), cc.prog_if, pif }) catch return
else
std.fmt.bufPrint(&line, "pci-bus: {d}:{d}.{d} class 0x{x:0>2} ({s}) subclass 0x{x:0>2} ({s}) progif 0x{x:0>2}\n", .{ bus, dev, function, cc.base, pci_class.className(cc.base), cc.subclass, pci_class.subclassName(cc.base, cc.subclass), cc.prog_if }) catch return;
_ = runtime.system.write(text);
}
var bridge_id: u64 = protocol.no_device;
var ecam_base: usize = 0;
var ecam_physical: u64 = 0;
var start_bus: u64 = 0;
var bus_count: u64 = 0;
var manager_handle: runtime.ipc.Handle = 0;
/// One aligned 32-bit read from a function's configuration space.
fn configRead(bus: u64, dev: u64, function: u64, offset: u64) u32 {
const address = ecam_base + (((bus - start_bus) << 20) | (dev << 15) | (function << 12) | offset);
const register: *volatile u32 = @ptrFromInt(address);
return register.*;
}
fn configWrite(bus: u64, dev: u64, function: u64, offset: u64, value: u32) void {
const address = ecam_base + (((bus - start_bus) << 20) | (dev << 15) | (function << 12) | offset);
const register: *volatile u32 = @ptrFromInt(address);
register.* = value;
}
fn configRead16(bus: u64, dev: u64, function: u64, offset: u64) u16 {
const word = configRead(bus, dev, function, offset & ~@as(u64, 3));
return @truncate(word >> @intCast((offset & 3) * 8));
}
fn configWrite16(bus: u64, dev: u64, function: u64, offset: u64, value: u16) void {
const aligned = offset & ~@as(u64, 3);
const shift: u5 = @intCast((offset & 3) * 8);
const word = configRead(bus, dev, function, aligned);
const mask = @as(u32, 0xFFFF) << shift;
configWrite(bus, dev, function, aligned, (word & ~mask) | (@as(u32, value) << shift));
}
/// Claim the bridge, map the ECAM, hello the manager, then scan.
fn initialise(endpoint: runtime.ipc.Handle) bool {
_ = endpoint;
if (!device.claim(bridge_id)) {
writeLine("pci-bus: unable to claim bridge device {d}\n", .{bridge_id});
return false;
}
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("pci-bus: out of memory\n");
return false;
};
const total = device.enumerate(buffer);
const descriptor = for (buffer[0..@min(total, buffer.len)]) |d| {
if (d.id == bridge_id) break d;
} else {
writeLine("pci-bus: device {d} not in the device tree\n", .{bridge_id});
return false;
};
// Resource 0 is the ECAM window (1 MiB of config space per bus); the bus
// range rides beside it. The MMIO apertures (M19.0) come after both.
if (descriptor.resource_count < 2 or descriptor.resources[0].kind != @intFromEnum(device.ResourceKind.memory)) {
_ = runtime.system.write("pci-bus: bridge has no ECAM window\n");
return false;
}
const bus_range = for (descriptor.resources[0..@intCast(descriptor.resource_count)]) |resource| {
if (resource.kind == @intFromEnum(device.ResourceKind.bus_range)) break resource;
} else {
_ = runtime.system.write("pci-bus: bridge has no bus range\n");
return false;
};
start_bus = bus_range.start;
bus_count = bus_range.len;
ecam_physical = descriptor.resources[0].start;
ecam_base = device.mmioMap(bridge_id, 0) orelse {
_ = runtime.system.write("pci-bus: ECAM mmio_map failed\n");
return false;
};
// The handshake, then the scan (reports join in M19.2).
var manager: ?runtime.ipc.Handle = null;
var tries: u32 = 0;
while (manager == null and tries < 100) : (tries += 1) {
manager = runtime.ipc.lookup(.device_manager);
if (manager == null) runtime.system.sleep(20);
}
const h = manager orelse {
_ = runtime.system.write("pci-bus: no device manager to hello\n");
return false;
};
const hello = protocol.Hello{ .role = @intFromEnum(protocol.Role.bus), .device_id = bridge_id };
var reply: [protocol.message_maximum]u8 = undefined;
const n = runtime.ipc.call(h, std.mem.asBytes(&hello), &reply) catch {
_ = runtime.system.write("pci-bus: hello call failed\n");
return false;
};
if (n < protocol.reply_size or std.mem.bytesToValue(protocol.HelloReply, reply[0..protocol.reply_size]).status != 0) {
_ = runtime.system.write("pci-bus: hello refused\n");
return false;
}
manager_handle = h;
scan();
return true;
}
/// The brute-force walk the kernel does today, from ring 3: every bus in the
/// range, 32 devices, 8 functions; vendor id FFFFh means nothing decodes there,
/// and only multifunction devices get their functions 1..7 probed.
fn scan() void {
var found: u32 = 0;
var bus: u64 = start_bus;
while (bus < start_bus + bus_count) : (bus += 1) {
var dev: u64 = 0;
while (dev < 32) : (dev += 1) {
const first = configRead(bus, dev, 0, 0);
if (first & 0xFFFF == 0xFFFF) continue;
const multifunction = (configRead(bus, dev, 0, 0x0C) >> 16) & 0x80 != 0;
var function: u64 = 0;
while (function < 8) : (function += 1) {
if (function != 0 and !multifunction) break;
const vendor_device = configRead(bus, dev, function, 0);
if (vendor_device & 0xFFFF == 0xFFFF) continue;
const class_revision = configRead(bus, dev, function, 0x08);
found += 1;
logFunction(bus, dev, function, class_revision >> 8);
registerAndReport(bus, dev, function, class_revision >> 8);
}
}
}
writeLine("pci-bus: {d} functions found\n", .{found});
}
/// Register one function under the bridge and report it to the manager. The
/// descriptor mirrors the kernel's own recording byte for byte — config slice
/// as resource 0, then the sized BARs — so during coexistence the idempotent
/// device_register (M19.0) returns the kernel's existing node id rather than
/// growing a duplicate, and the report carries the id drivers already use.
fn registerAndReport(bus: u64, dev: u64, function: u64, class_triple: u32) void {
var descriptor = std.mem.zeroes(device.DeviceDescriptor);
descriptor.class = @intFromEnum(device.DeviceClass.pci_device);
descriptor.pci_class = class_triple;
descriptor.resources[0] = .{
.kind = @intFromEnum(device.ResourceKind.memory),
.start = ecam_physical + (((bus - start_bus) << 20) | (dev << 15) | (function << 12)),
.len = 4096,
};
descriptor.resource_count = 1;
// The standard BAR-sizing probe, exactly as the kernel does it: decode off,
// write all-ones, read the writable mask back, restore. Header type 0 only.
const header_type = (configRead(bus, dev, function, 0x0C) >> 16) & 0x7F;
if (header_type == 0) {
const command = configRead16(bus, dev, function, 0x04);
configWrite16(bus, dev, function, 0x04, command & ~@as(u16, 0b11));
var i: u64 = 0;
while (i < 6) : (i += 1) {
if (descriptor.resource_count >= 8) break;
const off = 0x10 + i * 4;
const original = configRead(bus, dev, function, off);
if (original == 0) continue;
const slot: usize = @intCast(descriptor.resource_count);
if (original & 1 != 0) {
configWrite(bus, dev, function, off, 0xFFFF_FFFF);
const readback = configRead(bus, dev, function, off);
configWrite(bus, dev, function, off, original);
const mask = readback & 0xFFFF_FFFC;
const size: u32 = if (mask == 0) 0 else (~mask +% 1) & 0xFFFF;
if (size == 0) continue; // unimplemented BAR — nothing to register
descriptor.resources[slot] = .{ .kind = @intFromEnum(device.ResourceKind.io_port), .start = original & 0xFFFF_FFFC, .len = size };
descriptor.resource_count += 1;
} else if ((original >> 1) & 0x3 == 2) {
const original_high = configRead(bus, dev, function, off + 4);
configWrite(bus, dev, function, off, 0xFFFF_FFFF);
configWrite(bus, dev, function, off + 4, 0xFFFF_FFFF);
const lo = configRead(bus, dev, function, off);
const hi = configRead(bus, dev, function, off + 4);
configWrite(bus, dev, function, off, original);
configWrite(bus, dev, function, off + 4, original_high);
const readback = (@as(u64, hi) << 32) | (lo & 0xFFFF_FFF0);
const size: u64 = if (readback == 0) 0 else ~readback +% 1;
i += 1; // consumed the high half regardless
if (size == 0) continue;
descriptor.resources[slot] = .{ .kind = @intFromEnum(device.ResourceKind.memory), .start = (@as(u64, original_high) << 32) | (original & 0xFFFF_FFF0), .len = size };
descriptor.resource_count += 1;
} else {
configWrite(bus, dev, function, off, 0xFFFF_FFFF);
const readback = configRead(bus, dev, function, off);
configWrite(bus, dev, function, off, original);
const mask = readback & 0xFFFF_FFF0;
const size: u32 = if (mask == 0) 0 else ~mask +% 1;
if (size == 0) continue;
descriptor.resources[slot] = .{ .kind = @intFromEnum(device.ResourceKind.memory), .start = original & 0xFFFF_FFF0, .len = size };
descriptor.resource_count += 1;
}
}
configWrite16(bus, dev, function, 0x04, command);
}
const registered = device.register(bridge_id, &descriptor) orelse {
writeLine("pci-bus: register refused for {d}:{d}.{d}\n", .{ bus, dev, function });
return;
};
const report = protocol.ChildAdded{
.parent = bridge_id,
.bus_address = (bus << 8) | (dev << 3) | function,
.identity = class_triple,
.device_id = registered,
};
var reply: [protocol.message_maximum]u8 = undefined;
_ = runtime.ipc.call(manager_handle, std.mem.asBytes(&report), &reply) catch {
writeLine("pci-bus: child report for {d}:{d}.{d} failed\n", .{ bus, dev, function });
};
}
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = message;
_ = reply;
_ = sender;
_ = capability;
return 0;
}
pub fn main(init: runtime.process.Init) void {
const argument = init.arguments.get(1) orelse return; // bare (ramdisk sweep): stay silent
bridge_id = std.fmt.parseInt(u64, argument, 10) catch {
writeLine("pci-bus: malformed bridge device id '{s}'\n", .{argument});
return;
};
runtime.service.run(protocol.message_maximum, .{
.init = initialise,
.on_message = onMessage,
});
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
+184
View File
@@ -0,0 +1,184 @@
//! PS/2 Keyboard Driver
//!
//! Spawned by the ps2-bus driver once the controller is initialized and the port
//! has passed its interface test and device reset. The bus driver hands us our
//! device HID as argv[1] and, optionally, a layout name (`"us"`, `"gb"`, ...) as
//! argv[2].
//!
//! The 8042's ports (0x60/0x64) and IRQ1 live on the PNP0303 node, which the
//! ps2-bus driver exclusively owns — so this driver never touches the hardware.
//! Instead it **attaches** to the bus (handing over its endpoint as a capability)
//! and receives every scancode byte as a forwarded asynchronous message. Each byte
//! feeds the set-2 decoder; a decoded key becomes input-protocol events:
//!
//! scancode byte -> HID usage keycode -> key_down / key_up
//! -> xkeyboard-config -> character -> key_press
const std = @import("std");
const runtime = @import("runtime");
const xkb = @import("xkeyboard-config");
const ps2 = @import("ps2-library.zig");
const scancode = @import("scancode.zig");
const device = runtime.device;
const ipc = runtime.ipc;
const protocol = runtime.input_protocol;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
/// Look up the ps2-bus service, retrying while the bus (which spawned us before
/// registering) is still coming up.
fn lookupBus() ?ipc.Handle {
var attempts: usize = 0;
while (attempts < 100) : (attempts += 1) {
if (ipc.lookup(.ps2_bus)) |handle| return handle;
runtime.system.sleep(50);
}
return null;
}
/// The character a pressed key produces under `modifiers`, or 0 for none. The
/// layout lookup answers for printable keys; the keys whose keysym has no Unicode
/// mapping but that every consumer still expects as a character (Enter, Tab,
/// Backspace, Escape) are given their ASCII control characters here.
fn characterFor(layout: *const xkb.Layout, usage: u8, modifiers: scancode.ModifierSnapshot) u32 {
const mapping = xkb.map(layout, usage, .{
.shift = modifiers.shift,
.caps_lock = modifiers.caps_lock,
.level3 = modifiers.right_alt,
.control = modifiers.control,
});
if (mapping.character) |character| return character;
return switch (@as(protocol.Keycode, @enumFromInt(usage))) {
.enter, .keypad_enter => '\n',
.tab => '\t',
.backspace => 0x08,
.escape => 0x1B,
else => 0,
};
}
/// The input protocol's modifier word for a snapshot.
fn modifierWord(modifiers: scancode.ModifierSnapshot) u32 {
var word: u32 = 0;
if (modifiers.shift) word |= protocol.modifier_shift;
if (modifiers.control) word |= protocol.modifier_control;
if (modifiers.alt) word |= protocol.modifier_alt;
return word;
}
pub fn main(init: runtime.process.Init) void {
const hid = init.arguments.get(1).?;
if (hid.len == 0) {
_ = runtime.system.write("/system/drivers/ps2-bus/keyboard: no HID argument\n");
return;
}
writeLine("/system/drivers/ps2-bus/keyboard: starting for hid {s}\n", .{hid});
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("/system/drivers/ps2-bus/keyboard: out of memory\n");
return;
};
if (device.findDeviceDescriptorByHid(buffer, hid) == null) {
writeLine("/system/drivers/ps2-bus/keyboard: no device for hid {s}\n", .{hid});
return;
}
// The layout is a spawn argument so a later settings source can choose it;
// absent (as today) it defaults to us.
const layout_name = init.arguments.get(2) orelse "us";
const layout = xkb.byName(layout_name) orelse xkb.us;
writeLine("/system/drivers/ps2-bus/keyboard: layout {s}\n", .{layout.name});
// Attach to the bus: hand it our endpoint, and it forwards every byte the
// keyboard sends (it owns the controller; we own the decoding).
const bus = lookupBus() orelse {
_ = runtime.system.write("/system/drivers/ps2-bus/keyboard: ps2-bus service unavailable\n");
return;
};
const endpoint = ipc.createIpcEndpoint() orelse {
_ = runtime.system.write("/system/drivers/ps2-bus/keyboard: no endpoint\n");
return;
};
var attach = ps2.AttachRequest{ .device_type = @intFromEnum(ps2.DeviceType.keyboard) };
var attach_reply: [@sizeOf(ps2.AttachReply)]u8 = undefined;
const attached = ipc.callCap(bus, std.mem.asBytes(&attach), &attach_reply, endpoint) catch {
_ = runtime.system.write("/system/drivers/ps2-bus/keyboard: attach call failed\n");
return;
};
if (attached.len < @sizeOf(ps2.AttachReply) or
std.mem.bytesToValue(ps2.AttachReply, attach_reply[0..@sizeOf(ps2.AttachReply)]).status != @intFromEnum(ps2.AttachStatus.ok))
{
_ = runtime.system.write("/system/drivers/ps2-bus/keyboard: attach refused\n");
return;
}
// Broadcast keyboard events through the input service so programs can listen
// for them (docs/input.md).
var source = runtime.input.connectSource() orelse {
_ = runtime.system.write("/system/drivers/ps2-bus/keyboard: input service unavailable\n");
return;
};
_ = runtime.system.write("/system/drivers/ps2-bus/keyboard: ok\n");
var decoder = scancode.Decoder{};
var state = scancode.KeyboardState{};
var receive: [@sizeOf(ps2.ForwardedByte)]u8 = undefined;
while (true) {
const got = ipc.replyWait(endpoint, &.{}, &receive, null);
if (!got.isMessage() or got.len < @sizeOf(ps2.ForwardedByte)) continue;
const forwarded = std.mem.bytesToValue(ps2.ForwardedByte, receive[0..@sizeOf(ps2.ForwardedByte)]);
const key = decoder.feed(@intCast(forwarded.byte & 0xFF)) orelse continue;
const transition = state.apply(key);
const modifiers = modifierWord(transition.modifiers);
switch (transition.action) {
.pressed => {
_ = source.publishKeyboardEvent(.{
.kind = @intFromEnum(protocol.EventKind.key_down),
.keycode = key.usage,
.character = 0,
.modifiers = modifiers,
});
const character = characterFor(layout, key.usage, transition.modifiers);
if (character != 0) {
_ = source.publishKeyboardEvent(.{
.kind = @intFromEnum(protocol.EventKind.key_press),
.keycode = key.usage,
.character = character,
.modifiers = modifiers,
});
}
},
// Typematic repeat: the key did not physically go down again, so no
// key_down — but it keeps producing its character.
.repeated => {
const character = characterFor(layout, key.usage, transition.modifiers);
if (character != 0) {
_ = source.publishKeyboardEvent(.{
.kind = @intFromEnum(protocol.EventKind.key_press),
.keycode = key.usage,
.character = character,
.modifiers = modifiers,
});
}
},
.released => {
_ = source.publishKeyboardEvent(.{
.kind = @intFromEnum(protocol.EventKind.key_up),
.keycode = key.usage,
.character = 0,
.modifiers = modifiers,
});
},
}
}
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
+143
View File
@@ -0,0 +1,143 @@
//! PS/2 mouse packet assembly — the byte stream a streaming mouse sends, turned
//! into decoded movement/button reports.
//!
//! A standard PS/2 mouse in stream mode sends three-byte packets:
//!
//! byte 0: | Y ovf | X ovf | Y sign | X sign | 1 | middle | right | left |
//! byte 1: X movement (low eight bits; the sign bit lives in byte 0)
//! byte 2: Y movement (likewise)
//!
//! Movement is nine-bit two's complement, PS/2 convention: positive X right,
//! positive Y **up**. The decoded packet converts Y to the screen convention
//! (positive down), matching what every consumer of relative motion expects.
//! Bit 3 of byte 0 is always set — the resynchronization anchor: a byte at
//! packet start with bit 3 clear cannot be a packet header and is dropped.
//!
//! Everything here is pure (no imports beyond `std`, no IO), so it is
//! host-testable: the tests at the bottom run under `zig build test`.
const std = @import("std");
/// One decoded movement/button report, in screen convention (positive dy down).
pub const Packet = struct {
left: bool,
right: bool,
middle: bool,
dx: i16,
dy: i16,
};
const header_always_set: u8 = 1 << 3;
const header_left: u8 = 1 << 0;
const header_right: u8 = 1 << 1;
const header_middle: u8 = 1 << 2;
const header_x_sign: u8 = 1 << 4;
const header_y_sign: u8 = 1 << 5;
const header_x_overflow: u8 = 1 << 6;
const header_y_overflow: u8 = 1 << 7;
/// Device protocol bytes that can reach the packet stream around bring-up (the
/// acknowledge to enable-reporting, a reset's self-test result). Both have bit 3
/// set, so the header check alone cannot reject them; they are recognized only
/// at packet start, where a real header cannot be one of them in practice.
const response_acknowledge: u8 = 0xFA;
const response_self_test_passed: u8 = 0xAA;
/// Accumulates the byte stream into `Packet`s. Feed it every byte the mouse
/// sends; the third byte of each well-formed packet returns one.
pub const Assembler = struct {
bytes: [3]u8 = undefined,
count: u8 = 0,
pub fn feed(self: *Assembler, byte: u8) ?Packet {
if (self.count == 0) {
// Resynchronize: a packet must start with a plausible header.
if (byte & header_always_set == 0) return null;
if (byte == response_acknowledge or byte == response_self_test_passed) return null;
}
self.bytes[self.count] = byte;
self.count += 1;
if (self.count < 3) return null;
self.count = 0;
const header = self.bytes[0];
// An overflowed count is garbage by definition; discard the packet.
if (header & (header_x_overflow | header_y_overflow) != 0) return null;
return .{
.left = header & header_left != 0,
.right = header & header_right != 0,
.middle = header & header_middle != 0,
.dx = movement(self.bytes[1], header & header_x_sign != 0),
// PS/2 positive Y is up; screen positive Y is down.
.dy = -movement(self.bytes[2], header & header_y_sign != 0),
};
}
/// Nine-bit two's complement: the eight movement bits plus the header's sign.
fn movement(low: u8, negative: bool) i16 {
const value: i16 = low;
return if (negative) value - 256 else value;
}
};
// --- tests (host-run via `zig build test`) ------------------------------------
const testing = std.testing;
fn feedAll(assembler: *Assembler, bytes: []const u8) ?Packet {
var result: ?Packet = null;
for (bytes) |byte| {
if (assembler.feed(byte)) |packet| result = packet;
}
return result;
}
test "plain motion decodes with screen-convention y" {
var assembler = Assembler{};
const packet = feedAll(&assembler, &.{ 0x08, 5, 3 }).?;
try testing.expectEqual(@as(i16, 5), packet.dx);
try testing.expectEqual(@as(i16, -3), packet.dy); // PS/2 up 3 -> screen -3
try testing.expect(!packet.left and !packet.right and !packet.middle);
}
test "negative movement sign-extends through the header bits" {
var assembler = Assembler{};
// X sign and Y sign set: dx = 0xFB - 256 = -5, dy raw = 0xFE - 256 = -2 -> screen +2.
const packet = feedAll(&assembler, &.{ 0x08 | 0x10 | 0x20, 0xFB, 0xFE }).?;
try testing.expectEqual(@as(i16, -5), packet.dx);
try testing.expectEqual(@as(i16, 2), packet.dy);
}
test "buttons decode from the header" {
var assembler = Assembler{};
const packet = feedAll(&assembler, &.{ 0x08 | 0x01 | 0x02, 0, 0 }).?;
try testing.expect(packet.left);
try testing.expect(packet.right);
try testing.expect(!packet.middle);
}
test "a byte with bit 3 clear at packet start is dropped" {
var assembler = Assembler{};
// The stray 0x02 cannot be a header; the following packet still decodes.
try testing.expectEqual(@as(?Packet, null), assembler.feed(0x02));
const packet = feedAll(&assembler, &.{ 0x09, 1, 0 }).?;
try testing.expect(packet.left);
try testing.expectEqual(@as(i16, 1), packet.dx);
}
test "protocol bytes at packet start are dropped" {
var assembler = Assembler{};
try testing.expectEqual(@as(?Packet, null), assembler.feed(0xFA)); // enable-reporting ACK
try testing.expectEqual(@as(?Packet, null), assembler.feed(0xAA)); // self-test passed
const packet = feedAll(&assembler, &.{ 0x08, 7, 0 }).?;
try testing.expectEqual(@as(i16, 7), packet.dx);
}
test "an overflowed packet is discarded whole" {
var assembler = Assembler{};
try testing.expectEqual(@as(?Packet, null), feedAll(&assembler, &.{ 0x08 | 0x40, 0xFF, 0xFF }));
// The assembler is back at packet start.
const packet = feedAll(&assembler, &.{ 0x08, 1, 1 }).?;
try testing.expectEqual(@as(i16, 1), packet.dx);
}
+144
View File
@@ -0,0 +1,144 @@
//! PS/2 Mouse Driver
//!
//! Spawned by the ps2-bus driver once the controller is initialized and the port
//! has passed its interface test and device reset. The bus driver hands us our
//! device HID as argv[1].
//!
//! Like the keyboard, this driver never touches the hardware: the 8042's ports
//! and both port IRQs are owned by the ps2-bus driver (the auxiliary port's
//! IRQ12 lives on the PNP0F13 node, which the bus claims alongside the
//! controller). The driver **attaches** to the bus and receives every byte the
//! mouse sends as a forwarded asynchronous message. The bytes assemble into
//! three-byte packets, and each packet becomes input-protocol events:
//!
//! packet -> button transitions -> button_down / button_up
//! -> movement -> motion (dx/dy, screen convention)
const std = @import("std");
const runtime = @import("runtime");
const ps2 = @import("ps2-library.zig");
const mouse_packet = @import("mouse-packet.zig");
const device = runtime.device;
const ipc = runtime.ipc;
const protocol = runtime.input_protocol;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
/// Look up the ps2-bus service, retrying while the bus (which spawned us before
/// registering) is still coming up.
fn lookupBus() ?ipc.Handle {
var attempts: usize = 0;
while (attempts < 100) : (attempts += 1) {
if (ipc.lookup(.ps2_bus)) |handle| return handle;
runtime.system.sleep(50);
}
return null;
}
/// The protocol's pressed-button bitmask for a packet.
fn buttonMask(packet: mouse_packet.Packet) u32 {
var mask: u32 = 0;
if (packet.left) mask |= protocol.mouse_button_left;
if (packet.right) mask |= protocol.mouse_button_right;
if (packet.middle) mask |= protocol.mouse_button_middle;
return mask;
}
pub fn main(init: runtime.process.Init) void {
const hid = init.arguments.get(1).?;
if (hid.len == 0) {
_ = runtime.system.write("/system/drivers/ps2-bus/mouse: no HID argument\n");
return;
}
writeLine("/system/drivers/ps2-bus/mouse: starting for hid {s}\n", .{hid});
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("/system/drivers/ps2-bus/mouse: out of memory\n");
return;
};
if (device.findDeviceDescriptorByHid(buffer, hid) == null) {
writeLine("/system/drivers/ps2-bus/mouse: no device for hid {s}\n", .{hid});
return;
}
// Attach to the bus: hand it our endpoint, and it forwards every byte the
// mouse sends (it owns the controller; we own the decoding).
const bus = lookupBus() orelse {
_ = runtime.system.write("/system/drivers/ps2-bus/mouse: ps2-bus service unavailable\n");
return;
};
const endpoint = ipc.createIpcEndpoint() orelse {
_ = runtime.system.write("/system/drivers/ps2-bus/mouse: no endpoint\n");
return;
};
var attach = ps2.AttachRequest{ .device_type = @intFromEnum(ps2.DeviceType.mouse) };
var attach_reply: [@sizeOf(ps2.AttachReply)]u8 = undefined;
const attached = ipc.callCap(bus, std.mem.asBytes(&attach), &attach_reply, endpoint) catch {
_ = runtime.system.write("/system/drivers/ps2-bus/mouse: attach call failed\n");
return;
};
if (attached.len < @sizeOf(ps2.AttachReply) or
std.mem.bytesToValue(ps2.AttachReply, attach_reply[0..@sizeOf(ps2.AttachReply)]).status != @intFromEnum(ps2.AttachStatus.ok))
{
_ = runtime.system.write("/system/drivers/ps2-bus/mouse: attach refused\n");
return;
}
// Broadcast mouse events through the input service so programs can listen
// for them (docs/input.md).
var source = runtime.input.connectSource() orelse {
_ = runtime.system.write("/system/drivers/ps2-bus/mouse: input service unavailable\n");
return;
};
_ = runtime.system.write("/system/drivers/ps2-bus/mouse: ok\n");
var assembler = mouse_packet.Assembler{};
var buttons: u32 = 0;
var receive: [@sizeOf(ps2.ForwardedByte)]u8 = undefined;
while (true) {
const got = ipc.replyWait(endpoint, &.{}, &receive, null);
if (!got.isMessage() or got.len < @sizeOf(ps2.ForwardedByte)) continue;
const forwarded = std.mem.bytesToValue(ps2.ForwardedByte, receive[0..@sizeOf(ps2.ForwardedByte)]);
const packet = assembler.feed(@intCast(forwarded.byte & 0xFF)) orelse continue;
const new_buttons = buttonMask(packet);
// A button transition per changed button, carrying the new whole mask.
const changed = buttons ^ new_buttons;
for ([_]u32{ protocol.mouse_button_left, protocol.mouse_button_right, protocol.mouse_button_middle }) |button| {
if (changed & button == 0) continue;
const kind: protocol.MouseEventKind = if (new_buttons & button != 0) .button_down else .button_up;
_ = source.publishMouseEvent(.{
.kind = @intFromEnum(kind),
.button = button,
.dx = 0,
.dy = 0,
.scroll_x = 0,
.scroll_y = 0,
.buttons = new_buttons,
});
}
buttons = new_buttons;
if (packet.dx != 0 or packet.dy != 0) {
_ = source.publishMouseEvent(.{
.kind = @intFromEnum(protocol.MouseEventKind.motion),
.button = 0,
.dx = packet.dx,
.dy = packet.dy,
.scroll_x = 0,
.scroll_y = 0,
.buttons = new_buttons,
});
}
}
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
+315
View File
@@ -0,0 +1,315 @@
//! The PS/2 Controller is located on the mainboard.
//! In the early days the controller was a single chip (Intel 8042).
//! As of today it is part of the Advanced Integrated Peripheral.
//!
//! It shows up in the device discovery as:
//! KBD_ [acpi_device] hid=PNP0303 (PS/2 Keyboard)
//! - io_port 0x60 len 0x1
//! - io_port 0x64 len 0x1
//! - irq 0x1 len 0x1
//! MOU_ [acpi_device] hid=PNP0F13 (PS/2 Mouse)
//! - irq 0xc len 0x1
const std = @import("std");
const runtime = @import("runtime");
const acpi_ids = @import("acpi-ids");
const ps2 = @import("ps2-library.zig");
const device = runtime.device;
const ipc = runtime.ipc;
/// Format one whole log line and emit it in a single `debug_write`, so output
/// from the child drivers (which run concurrently) can never interleave with it.
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
/// Ask the device on `port` what it is, then spawn the matching driver from the
/// initial-ramdisk, handing it the device's HID as argv[1]. The driver is chosen
/// from what the device reports, not from the port number. Returns the identified
/// type so the forwarding loop can route that port's bytes to the driver once it
/// attaches, or null if nothing was spawned.
fn spawnIdentifiedDriver(controller: ps2.Controller, port: ps2.Port) ?ps2.DeviceType {
const device_type = controller.identifyDevice(port) orelse {
writeLine("/system/drivers/ps2-bus: identify timed out on port {s}\n", .{@tagName(port)});
return null;
};
const driver_name = device_type.driverName() orelse {
writeLine("/system/drivers/ps2-bus: unrecognized device on port {s}\n", .{@tagName(port)});
return null;
};
const hid = device_type.hid() orelse "";
if (runtime.system.spawnWithArguments(driver_name, &.{hid}) != null) {
writeLine("/system/drivers/ps2-bus: port {s} is a {s}, spawned {s}\n", .{ @tagName(port), hid, driver_name });
return device_type;
}
writeLine("/system/drivers/ps2-bus: failed to spawn {s}\n", .{driver_name});
return null;
}
/// Resource index of the controller's IRQ (IRQ1) on the PNP0303 descriptor, found
/// the way the ports are found in `Controller.init`.
fn findInterruptResourceIndex(descriptor: device.DeviceDescriptor) ?u64 {
for (0..descriptor.resource_count) |index| {
if (descriptor.resources[index].kind == @intFromEnum(device.ResourceKind.irq)) return index;
}
return null;
}
/// Forwarding endpoints of the attached child drivers, indexed by `ps2.Port`.
/// Written when a child's `AttachRequest` arrives, read on every forwarded byte.
var port_endpoints = [_]?ipc.Handle{ null, null };
/// Which device type each port identified as, so an attaching child (which knows
/// its type, not its port) can be matched to the right port's byte stream.
var port_device_types = [_]?ps2.DeviceType{ null, null };
/// Handle a child driver's `AttachRequest`: record the endpoint capability it
/// passed as the forwarding target for the port whose device matches its type.
/// Writes an `AttachReply` into `out` and returns its length.
fn handleAttach(message: []const u8, got: ipc.Received, out: []u8) usize {
const reply = struct {
fn write(buffer: []u8, status: ps2.AttachStatus) usize {
const header = ps2.AttachReply{ .status = @intFromEnum(status) };
@memcpy(buffer[0..@sizeOf(ps2.AttachReply)], std.mem.asBytes(&header));
return @sizeOf(ps2.AttachReply);
}
};
if (message.len < @sizeOf(ps2.AttachRequest)) return reply.write(out, .invalid_request);
const request = std.mem.bytesToValue(ps2.AttachRequest, message[0..@sizeOf(ps2.AttachRequest)]);
const endpoint = got.cap orelse return reply.write(out, .missing_endpoint);
for (&port_device_types, 0..) |maybe_type, port_index| {
const device_type = maybe_type orelse continue;
if (@intFromEnum(device_type) != request.device_type) continue;
port_endpoints[port_index] = endpoint;
writeLine("/system/drivers/ps2-bus: {s} driver attached\n", .{@tagName(device_type)});
return reply.write(out, .ok);
}
return reply.write(out, .no_such_device);
}
pub fn main() void {
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("/system/drivers/ps2-bus: out of memory\n");
return;
};
var has_two_channels = false;
var maybe_controller: ?ps2.Controller = null;
var maybe_interrupt_index: ?u64 = null;
// The 8042's IO ports (0x60/0x64) are enumerated under the keyboard ACPI node
// (PNP0303), so we init the controller from that descriptor — but which device
// is on which port is decided later by identify, not by this HID.
const maybe_controller_device_descriptor = device.findDeviceDescriptorByHid(buffer, acpi_ids.HardwareId.ps2_keyboard.hid());
if (maybe_controller_device_descriptor) |controller_device_descriptor| {
_ = runtime.system.write("/system/drivers/ps2-bus: found PS/2 controller\n");
_ = runtime.system.write("/system/drivers/ps2-bus: initializing controller\n");
if (!device.claim(controller_device_descriptor.id)) {
_ = runtime.system.write("/system/drivers/ps2-bus: unable to claim controller \n");
return;
}
const controller = ps2.Controller.init(controller_device_descriptor) orelse {
_ = runtime.system.write("/system/drivers/ps2-bus: controller is missing its IO ports\n");
return;
};
maybe_controller = controller;
maybe_interrupt_index = findInterruptResourceIndex(controller_device_descriptor);
controller.disablePort(.one);
controller.disablePort(.two);
controller.flushOutputBuffer();
const current = controller.readConfigurationByte() orelse {
_ = runtime.system.write("/system/drivers/ps2-bus: controller configuration timed out\n");
return;
};
const update = current & ~(ps2.configuration_first_port_interrupt |
ps2.configuration_second_port_interrupt |
ps2.configuration_first_port_translation);
if (controller.writeConfigurationByte(update) == null) {
_ = runtime.system.write("/system/drivers/ps2-bus: controller configuration timed out\n");
return;
}
if (controller.performSelfTest()) |reply| {
if (reply != ps2.response_controller_test_passed) {
_ = runtime.system.write("/system/drivers/ps2-bus: perform controller self test failed\n");
return;
}
} else {
_ = runtime.system.write("/system/drivers/ps2-bus: controller self test timed out\n");
return;
}
has_two_channels = controller.hasTwoChannels() orelse {
_ = runtime.system.write("/system/drivers/ps2-bus: controller channels timed out\n");
return;
};
if (has_two_channels) {
_ = runtime.system.write("/system/drivers/ps2-bus: has two channels\n");
// keep the bus quiet until we have tested the ports and are ready to use them
controller.disablePort(.two);
} else {
_ = runtime.system.write("/system/drivers/ps2-bus: has one channel\n");
}
// interface tests: always test port 1, test port 2 only if it exists
const port_one_works = (controller.testPort(.one) orelse {
_ = runtime.system.write("/system/drivers/ps2-bus: port 1 test timed out\n");
return;
}) == ps2.response_port_test_passed;
var port_two_works = false;
if (has_two_channels) {
port_two_works = (controller.testPort(.two) orelse {
_ = runtime.system.write("/system/drivers/ps2-bus: port 2 test timed out\n");
return;
}) == ps2.response_port_test_passed;
}
if (!port_one_works and !port_two_works) {
_ = runtime.system.write("/system/drivers/ps2-bus: no usable ports\n");
return;
}
// Enable the working ports. Their interrupts stay off until IRQ1 is bound
// below — reset and identify use polled reads, which must never race the
// interrupt-driven drain loop for bytes.
controller.enablePort(.one);
if (port_two_works) controller.enablePort(.two);
// reset each working device; a failing device is logged but does not
// abort bring-up of the other one
if (port_one_works) {
if (controller.resetDevice(.one)) |passed| {
if (!passed) _ = runtime.system.write("/system/drivers/ps2-bus: port 1 device reset failed\n");
} else {
_ = runtime.system.write("/system/drivers/ps2-bus: port 1 device reset timed out\n");
}
}
if (port_two_works) {
if (controller.resetDevice(.two)) |passed| {
if (!passed) _ = runtime.system.write("/system/drivers/ps2-bus: port 2 device reset failed\n");
} else {
_ = runtime.system.write("/system/drivers/ps2-bus: port 2 device reset timed out\n");
}
}
// Identify the device on each working port and hand it off to the driver
// that matches what it reported — a port is not assumed to be a keyboard
// or a mouse by its number.
if (port_one_works) port_device_types[@intFromEnum(ps2.Port.one)] = spawnIdentifiedDriver(controller, .one);
if (port_two_works) port_device_types[@intFromEnum(ps2.Port.two)] = spawnIdentifiedDriver(controller, .two);
} else {
_ = runtime.system.write("/system/drivers/ps2-bus: no PS/2 controller found\n");
return;
}
const controller = maybe_controller.?;
const interrupt_index = maybe_interrupt_index orelse {
_ = runtime.system.write("/system/drivers/ps2-bus: controller is missing its IRQ\n");
return;
};
// The endpoint the child drivers attach to and IRQ1 wakes. Registered under a
// well-known id so the children can find it, the way input subscribers find
// the input service.
const endpoint = ipc.createIpcEndpoint() orelse {
_ = runtime.system.write("/system/drivers/ps2-bus: no endpoint\n");
return;
};
if (!ipc.register(.ps2_bus, endpoint)) {
_ = runtime.system.write("/system/drivers/ps2-bus: register failed\n");
return;
}
// From here on, only the interrupt path reads the data port. Drop anything a
// device sent between enable-scanning and now, bind the IRQs, and only then
// let the controller raise them — an interrupt with nobody bound is lost.
controller.drainOutputBuffer();
if (!device.irqBind(controller.device_id, interrupt_index, endpoint)) {
_ = runtime.system.write("/system/drivers/ps2-bus: irq_bind failed\n");
return;
}
// Port 2's interrupt (IRQ12) is enumerated on the auxiliary device's own ACPI
// node (PNP0F13), not on the controller's — so if port 2 carries a device,
// claim that node too and route its IRQ to the same endpoint. The IRQ belongs
// to the *port*, whatever device identify found on it.
var maybe_auxiliary_interrupt: ?struct { device_id: u64, interrupt_index: u64, gsi: u64 } = null;
if (port_device_types[@intFromEnum(ps2.Port.two)] != null) {
if (device.findDeviceDescriptorByHid(buffer, acpi_ids.HardwareId.ps2_mouse.hid())) |descriptor| {
if (findInterruptResourceIndex(descriptor)) |auxiliary_index| {
if (device.claim(descriptor.id) and device.irqBind(descriptor.id, auxiliary_index, endpoint)) {
maybe_auxiliary_interrupt = .{
.device_id = descriptor.id,
.interrupt_index = auxiliary_index,
.gsi = descriptor.resources[auxiliary_index].start,
};
} else {
_ = runtime.system.write("/system/drivers/ps2-bus: auxiliary irq_bind failed\n");
}
}
}
}
var configuration = controller.readConfigurationByte() orelse {
_ = runtime.system.write("/system/drivers/ps2-bus: controller configuration timed out\n");
return;
};
if (port_device_types[@intFromEnum(ps2.Port.one)] != null) configuration |= ps2.Port.one.interruptBit();
if (maybe_auxiliary_interrupt != null) configuration |= ps2.Port.two.interruptBit();
_ = controller.writeConfigurationByte(configuration);
_ = runtime.system.write("/system/drivers/ps2-bus: ok\n");
// The forwarding loop: an IRQ1 notification drains the output buffer, routing
// each byte to the attached driver of the port it came from; a client message
// is a child driver's AttachRequest.
var reply_buffer: [@sizeOf(ps2.AttachReply)]u8 = undefined;
var reply_len: usize = 0;
var receive: [@sizeOf(ps2.AttachRequest)]u8 = undefined;
while (true) {
const got = ipc.replyWait(endpoint, reply_buffer[0..reply_len], &receive, null);
if (got.isNotification()) {
reply_len = 0;
if (got.isMessage() or got.isChildExit()) continue; // nothing sends us these
while (true) {
const current_status = ps2.status(controller.device_id, controller.status_index);
if (current_status & ps2.status_output_buffer_full == 0) break;
const byte = device.ioRead(controller.device_id, controller.data_index, 0, 1) orelse break;
const port: ps2.Port = if (current_status & ps2.status_auxiliary_output != 0) .two else .one;
if (port_endpoints[@intFromEnum(port)]) |child| {
const forwarded = ps2.ForwardedByte{ .port = @intFromEnum(port), .byte = byte };
_ = ipc.send(child, std.mem.asBytes(&forwarded));
}
// An unattached port's byte is dropped — e.g. a keystroke before
// the keyboard driver has attached.
}
// Re-arm the line that woke us: the notification badge carries the
// GSI, and IRQ1 and IRQ12 are acked through different device claims.
if (maybe_auxiliary_interrupt) |auxiliary| {
if (got.source() == auxiliary.gsi) {
_ = device.irqAck(auxiliary.device_id, auxiliary.interrupt_index);
} else {
_ = device.irqAck(controller.device_id, interrupt_index);
}
} else {
_ = device.irqAck(controller.device_id, interrupt_index);
}
continue;
}
reply_len = handleAttach(receive[0..got.len], got, &reply_buffer);
}
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
+485
View File
@@ -0,0 +1,485 @@
//! shared definitions between the different PS/2 drivers
const std = @import("std");
const runtime = @import("runtime");
const acpi_ids = @import("acpi-ids");
const device = runtime.device;
const system = runtime.system;
/// PS-2 io ports:
/// The PS/2 Controller itself uses 2 IO ports (usually, IO ports 0x60 and 0x64). Like many IO
/// ports, reads and writes may access different internal registers.
///
/// Historical note: The PC-XT PPI had used port 0x61 to reset the keyboard interrupt request
/// signal (among other unrelated functions). Port 0x61 has no keyboard related functions on AT and
/// PS/2 compatibles.
///
/// The Data Port (typically IO Port 0x60) is used for reading data that was received from a PS/2
/// device or from the PS/2 controller itself and writing data to a PS/2 device or to the PS/2
/// controller itself.
// Access type: Read/Write
pub const dataPort = 0x60;
// Access type: Read
pub const statusRegisterPort = 0x64;
// Access type: Write
pub const CommandRegisterPort = 0x64;
/// How long to poll the status register before giving up. PS/2 controller
/// responses normally arrive within a few milliseconds.
pub const default_wait_timeout_nanoseconds: u64 = 10_000_000; // 10 ms
/// A PS/2 device reset (0xFF) runs the device's self-test (BAT), whose reply can
/// take far longer than an ordinary controller response.
pub const device_reset_timeout_nanoseconds: u64 = 750_000_000; // 750 ms
/// PS/2 controller commands, written to the command register (port 0x64).
pub const cmd_read_configuration_byte: u8 = 0x20; // read controller configuration byte (internal RAM byte 0)
pub const cmd_write_configuration_byte: u8 = 0x60; // write controller configuration byte (internal RAM byte 0)
pub const cmd_disable_second_port: u8 = 0xA7; // disable second PS/2 port (dual-channel controllers only)
pub const cmd_enable_second_port: u8 = 0xA8; // enable second PS/2 port (dual-channel controllers only)
pub const cmd_test_second_port: u8 = 0xA9; // test second PS/2 port
pub const cmd_test_controller: u8 = 0xAA; // controller self-test
pub const cmd_test_first_port: u8 = 0xAB; // test first PS/2 port
pub const cmd_diagnostic_dump: u8 = 0xAC; // read all bytes of internal RAM
pub const cmd_disable_first_port: u8 = 0xAD; // disable first PS/2 port
pub const cmd_enable_first_port: u8 = 0xAE; // enable first PS/2 port
pub const cmd_read_controller_input_port: u8 = 0xC0; // read controller input port
pub const cmd_read_controller_output_port: u8 = 0xD0; // read controller output port
pub const cmd_write_controller_output_port: u8 = 0xD1; // write next data byte to the controller output port
pub const cmd_write_first_port_output: u8 = 0xD2; // write next data byte to the first port output buffer
pub const cmd_write_second_port_output: u8 = 0xD3; // write next data byte to the second port output buffer
pub const cmd_write_second_port_input: u8 = 0xD4; // write next data byte to the second port input buffer (to the mouse)
pub const cmd_pulse_system_reset: u8 = 0xFE; // pulse output line 0 low: resets the CPU
/// PS/2 status register bits (read from the status port, 0x64). Bit 4 is
/// chipset-specific and intentionally omitted.
pub const status_output_buffer_full: u8 = 1 << 0; // 1 = a byte is waiting to be read from the data port
pub const status_input_buffer_full: u8 = 1 << 1; // 1 = the controller has not yet consumed the last write
pub const status_system_flag: u8 = 1 << 2; // set once the controller passes POST
pub const status_command_or_data: u8 = 1 << 3; // 1 = last write was a command, 0 = data
/// Chipset-specific in the original spec, universal in practice on dual-channel
/// controllers: set = the waiting byte came from the second port (the mouse).
pub const status_auxiliary_output: u8 = 1 << 5;
pub const status_timeout_error: u8 = 1 << 6; // 1 = time-out error
pub const status_parity_error: u8 = 1 << 7; // 1 = parity error
/// Controller configuration byte bits (internal RAM byte 0; read/written via 0x20/0x60).
pub const configuration_first_port_interrupt: u8 = 1 << 0; // 1 = first port IRQ (IRQ1) enabled
pub const configuration_second_port_interrupt: u8 = 1 << 1; // 1 = second port IRQ (IRQ12) enabled
pub const configuration_system_flag: u8 = 1 << 2; // 1 = system passed POST
pub const configuration_first_port_clock_disabled: u8 = 1 << 4; // 1 = first port clock disabled
pub const configuration_second_port_clock_disabled: u8 = 1 << 5; // 1 = second port clock disabled
pub const configuration_first_port_translation: u8 = 1 << 6; // 1 = first port scancode translation enabled
/// Controller output port bits (read/written via 0xD0/0xD1).
pub const output_port_system_reset: u8 = 1 << 0; // WARNING: keep this 1; writing 0 can lock the machine
pub const output_port_a20_gate: u8 = 1 << 1; // A20 gate
pub const output_port_second_port_clock: u8 = 1 << 2; // dual-channel controllers only
pub const output_port_second_port_data: u8 = 1 << 3; // dual-channel controllers only
pub const output_port_first_port_output_full: u8 = 1 << 4; // output buffer full from first port (IRQ1)
pub const output_port_second_port_output_full: u8 = 1 << 5; // output buffer full from second port (IRQ12)
pub const output_port_first_port_clock: u8 = 1 << 6; // first port clock
pub const output_port_first_port_data: u8 = 1 << 7; // first port data
/// Controller self-test (0xAA) result codes.
pub const response_controller_test_passed: u8 = 0x55;
pub const response_controller_test_failed: u8 = 0xFC;
/// Port test (0xAB / 0xA9) result codes.
pub const response_port_test_passed: u8 = 0x00;
pub const response_port_test_clock_stuck_low: u8 = 0x01;
pub const response_port_test_clock_stuck_high: u8 = 0x02;
pub const response_port_test_data_stuck_low: u8 = 0x03;
pub const response_port_test_data_stuck_high: u8 = 0x04;
/// PS/2 device commands, written to the data port (0x60) to reach the attached device.
pub const device_cmd_identify: u8 = 0xF2; // identify device
pub const device_cmd_enable_scanning: u8 = 0xF4;
pub const device_cmd_disable_scanning: u8 = 0xF5;
pub const device_cmd_reset: u8 = 0xFF; // reset and run the device self-test (BAT)
/// PS/2 device response bytes, read from the data port (0x60).
pub const device_response_self_test_passed: u8 = 0xAA; // BAT succeeded after a reset
pub const device_response_echo: u8 = 0xEE;
pub const device_response_acknowledge: u8 = 0xFA; // ACK
pub const device_response_self_test_failed_1: u8 = 0xFC; // BAT failure
pub const device_response_self_test_failed_2: u8 = 0xFD; // BAT failure
pub const device_response_resend: u8 = 0xFE; // ask the host to resend the last byte
/// PS/2 device identify (0xF2) reply bytes. A keyboard returns a two-byte id
/// beginning with 0xAB; a mouse returns a single-byte id (0x00/0x03/0x04); an
/// ancient AT keyboard returns nothing at all.
pub const identify_keyboard_mf2: u8 = 0xAB; // first byte of a MF2 keyboard id (a subtype byte follows)
pub const identify_mouse_standard: u8 = 0x00;
pub const identify_mouse_scroll: u8 = 0x03; // mouse with scroll wheel
pub const identify_mouse_five_button: u8 = 0x04; // 5-button mouse
fn waitReadable(id: u64, cmd_index: u64, wait_timeout_nanoseconds: u64) bool {
const deadline = system.clock() + wait_timeout_nanoseconds;
while (system.clock() < deadline) {
if (status(id, cmd_index) & status_output_buffer_full != 0) return true; // OBF set -> data ready
}
return false;
}
fn waitWritable(id: u64, cmd_index: u64, wait_timeout_nanoseconds: u64) bool {
const deadline = system.clock() + wait_timeout_nanoseconds;
while (system.clock() < deadline) {
if (status(id, cmd_index) & status_input_buffer_full == 0) return true; // IBF clear -> ok to write
}
return false; // timed out
}
pub fn status(id: u64, cmd_index: u64) u8 {
return @intCast(device.ioRead(id, cmd_index, 0, 1) orelse 0);
}
pub fn sendCommand(id: u64, cmd_index: u64, byte: u8, timeout_nanoseconds: u64) bool {
// wait IBF clear
if (!waitWritable(id, cmd_index, timeout_nanoseconds)) return false;
return device.ioWrite(id, cmd_index, 0, 1, byte);
}
pub fn readData(id: u64, status_index: u64, data_index: u64, timeout_nanoseconds: u64) ?u8 {
// OBF lives in the status register (0x64); wait for it there, then read the data port (0x60)
if (!waitReadable(id, status_index, timeout_nanoseconds)) return null;
return @intCast(device.ioRead(id, data_index, 0, 1) orelse 0);
}
pub fn writeData(id: u64, status_index: u64, data_index: u64, byte: u8, timeout_nanoseconds: u64) bool {
// IBF lives in the status register (0x64); wait for it to clear there, then write the data port
// (0x60)
if (!waitWritable(id, status_index, timeout_nanoseconds)) return false;
return device.ioWrite(id, data_index, 0, 1, byte);
}
pub const Port = enum(u2) {
one,
two,
/// Command register byte that disables this port.
fn disableCommand(self: Port) u8 {
return switch (self) {
.one => cmd_disable_first_port,
.two => cmd_disable_second_port,
};
}
/// Command register byte that enables this port (and its clock).
fn enableCommand(self: Port) u8 {
return switch (self) {
.one => cmd_enable_first_port,
.two => cmd_enable_second_port,
};
}
/// Command register byte that runs this port's interface test.
fn testCommand(self: Port) u8 {
return switch (self) {
.one => cmd_test_first_port,
.two => cmd_test_second_port,
};
}
/// Configuration-byte bit that, when set, disables this port's clock.
pub fn clockDisabledBit(self: Port) u8 {
return switch (self) {
.one => configuration_first_port_clock_disabled,
.two => configuration_second_port_clock_disabled,
};
}
/// Configuration-byte bit that, when set, enables this port's interrupt.
pub fn interruptBit(self: Port) u8 {
return switch (self) {
.one => configuration_first_port_interrupt,
.two => configuration_second_port_interrupt,
};
}
/// Command register byte that writes the next data byte into this port's
/// output buffer (makes a byte appear as if it came from the device).
pub fn writeOutputBufferCommand(self: Port) u8 {
return switch (self) {
.one => cmd_write_first_port_output,
.two => cmd_write_second_port_output,
};
}
/// Controller command that must prefix a byte destined for this port's
/// device. Port 1 is the default target of the data port, so it needs no
/// prefix (null); port 2 requires the "write second port input" command.
pub fn deviceInputCommand(self: Port) ?u8 {
return switch (self) {
.one => null,
.two => cmd_write_second_port_input,
};
}
/// Controller output-port bit driving this port's clock line.
pub fn outputPortClockBit(self: Port) u8 {
return switch (self) {
.one => output_port_first_port_clock,
.two => output_port_second_port_clock,
};
}
/// Controller output-port bit driving this port's data line.
pub fn outputPortDataBit(self: Port) u8 {
return switch (self) {
.one => output_port_first_port_data,
.two => output_port_second_port_data,
};
}
/// Controller output-port bit set when this port's output buffer is full
/// (wired to the port's IRQ line).
pub fn outputPortBufferFullBit(self: Port) u8 {
return switch (self) {
.one => output_port_first_port_output_full,
.two => output_port_second_port_output_full,
};
}
};
/// The kind of device attached to a port, as reported by the device itself in
/// response to the identify command — not assumed from the port number. Fixed
/// `u32` values because the type also travels in an `AttachRequest`.
pub const DeviceType = enum(u32) {
keyboard = 0,
mouse = 1,
unknown = 2,
/// Initial-ramdisk name of the driver that serves this device type, or null
/// if we could not classify it.
pub fn driverName(self: DeviceType) ?[]const u8 {
return switch (self) {
.keyboard => "ps2-keyboard",
.mouse => "ps2-mouse",
.unknown => null,
};
}
/// Canonical ACPI HID for this device type, handed to the spawned driver as
/// its command-line argument, or null if we could not classify it.
pub fn hid(self: DeviceType) ?[]const u8 {
return switch (self) {
.keyboard => acpi_ids.HardwareId.ps2_keyboard.hid(),
.mouse => acpi_ids.HardwareId.ps2_mouse.hid(),
.unknown => null,
};
}
};
// --- the bus <-> child-driver forwarding protocol -----------------------------
//
// The 8042's ports and IRQ1 live on the PNP0303 node that only the ps2-bus driver
// claims, so the child device drivers (ps2-keyboard, ps2-mouse) cannot read port
// 0x60 themselves. Instead each child **attaches**: it calls the bus's well-known
// `ps2_bus` endpoint with an `AttachRequest`, handing over its own endpoint as the
// call's capability. From then on the bus forwards every byte the device sends as
// a `ForwardedByte` via the asynchronous `ipc.send` — the IRQ path in the bus can
// never block on a slow child, and the child never touches the controller.
/// A child driver registering for its device's bytes. `device_type` is a
/// `DeviceType` value; the child's receive endpoint travels as the call's
/// capability (`send_cap`).
pub const AttachRequest = extern struct {
device_type: u32,
};
/// How the bus answered an `AttachRequest` (`AttachReply.status`).
pub const AttachStatus = enum(i32) {
ok = 0,
/// The request was malformed (too short to be an `AttachRequest`).
invalid_request = -1,
/// The call carried no endpoint capability to forward to.
missing_endpoint = -2,
/// No port identified a device of the requested type.
no_such_device = -3,
};
/// Reply to an `AttachRequest`. `status` is an `AttachStatus` value.
pub const AttachReply = extern struct {
status: i32,
_padding: u32 = 0,
};
/// One raw byte read from the data port, forwarded to the attached child whose
/// port it came from (routed by the status register's auxiliary-output bit).
pub const ForwardedByte = extern struct {
/// The `Port` the byte came from, as `@intFromEnum`.
port: u32,
byte: u32,
};
/// A single PS/2 (8042) controller. Construct one with `Controller.init` and
/// drive the controller through its methods; there is only ever one 8042 per
/// machine, but holding the resolved resource indices in an instance keeps the
/// call sites free of global state.
pub const Controller = struct {
device_id: u64,
/// Resource index of the command/status port (0x64).
status_index: u64,
/// Resource index of the data port (0x60).
data_index: u64,
/// Resolve the controller's IO-port resource indices from its device
/// descriptor. Returns null if either the data or command/status port is
/// missing from the descriptor.
pub fn init(device_descriptor: device.DeviceDescriptor) ?Controller {
var data_index: ?u64 = null;
var status_index: ?u64 = null;
for (device_descriptor.resources, 0..device_descriptor.resource_count) |resource, resource_index| {
if (resource.kind != @intFromEnum(device.ResourceKind.io_port)) continue;
if (resource.start == dataPort) {
data_index = @intCast(resource_index);
} else if (resource.start == statusRegisterPort) {
status_index = @intCast(resource_index);
}
}
return .{
.device_id = device_descriptor.id,
.data_index = data_index orelse return null,
.status_index = status_index orelse return null,
};
}
pub fn disablePort(self: Controller, port: Port) void {
// port enable/disable are controller commands and go to the command register (0x64)
_ = sendCommand(self.device_id, self.status_index, port.disableCommand(), default_wait_timeout_nanoseconds);
}
pub fn enablePort(self: Controller, port: Port) void {
// enabling a port also starts its clock
_ = sendCommand(self.device_id, self.status_index, port.enableCommand(), default_wait_timeout_nanoseconds);
}
/// Run a port's interface test. Returns the controller's reply — compare it
/// to `response_port_test_passed` (0x00) — or null on timeout.
pub fn testPort(self: Controller, port: Port) ?u8 {
if (!sendCommand(self.device_id, self.status_index, port.testCommand(), default_wait_timeout_nanoseconds)) return null;
return readData(self.device_id, self.status_index, self.data_index, default_wait_timeout_nanoseconds);
}
pub fn flushOutputBuffer(self: Controller) void {
// flush any stale byte the controller buffered
_ = device.ioRead(self.device_id, self.data_index, 0, 1);
}
pub fn readConfigurationByte(self: Controller) ?u8 {
// ask the controller to place its configuration byte in the output buffer, then read it
if (!sendCommand(self.device_id, self.status_index, cmd_read_configuration_byte, default_wait_timeout_nanoseconds)) return null;
return readData(self.device_id, self.status_index, self.data_index, default_wait_timeout_nanoseconds);
}
pub fn writeConfigurationByte(self: Controller, update_byte: u8) ?u8 {
// command 0x60 makes the controller store the next data-port byte as its configuration byte
if (!sendCommand(self.device_id, self.status_index, cmd_write_configuration_byte, default_wait_timeout_nanoseconds)) return null;
if (!writeData(self.device_id, self.status_index, self.data_index, update_byte, default_wait_timeout_nanoseconds)) return null;
return update_byte;
}
/// Run the controller self-test. Returns the reply — compare it to
/// `response_controller_test_passed` (0x55) — or null on timeout.
pub fn performSelfTest(self: Controller) ?u8 {
if (!sendCommand(self.device_id, self.status_index, cmd_test_controller, default_wait_timeout_nanoseconds)) return null;
return readData(self.device_id, self.status_index, self.data_index, default_wait_timeout_nanoseconds);
}
/// Detect whether this is a dual-channel controller by temporarily enabling
/// port 2 and checking whether its clock turned on. Note: this leaves port 2
/// enabled; the caller should disable it again to keep the bus quiet until
/// device bring-up.
pub fn hasTwoChannels(self: Controller) ?bool {
self.enablePort(.two);
const configuration = self.readConfigurationByte() orelse return null;
return (configuration & Port.two.clockDisabledBit()) == 0;
}
/// Reset the device attached to `port` (device command 0xFF) and wait for
/// its power-on self-test (BAT) result. Returns true if the device both
/// acknowledged and passed, false if it reported a self-test failure, or
/// null on timeout. The BAT reply can be slow, so the response reads use
/// `device_reset_timeout_nanoseconds`.
pub fn resetDevice(self: Controller, port: Port) ?bool {
// A byte destined for port 2 must be prefixed with the "write to second
// port input buffer" controller command (0xD4); port 1 is the default.
if (port.deviceInputCommand()) |prefix| {
if (!sendCommand(self.device_id, self.status_index, prefix, default_wait_timeout_nanoseconds)) return null;
}
if (!writeData(self.device_id, self.status_index, self.data_index, device_cmd_reset, default_wait_timeout_nanoseconds)) return null;
// A successful reset yields both an ACK (0xFA) and a self-test-passed
// byte (0xAA). Their order is not guaranteed, so accept either ordering.
var saw_acknowledge = false;
var saw_self_test_passed = false;
var reads: u8 = 0;
while (reads < 2) : (reads += 1) {
const reply = readData(self.device_id, self.status_index, self.data_index, device_reset_timeout_nanoseconds) orelse return null;
switch (reply) {
device_response_acknowledge => saw_acknowledge = true,
device_response_self_test_passed => saw_self_test_passed = true,
device_response_self_test_failed_1, device_response_self_test_failed_2 => return false,
else => {},
}
}
return saw_acknowledge and saw_self_test_passed;
}
/// Send one command byte to the device on `port` (applying the port-2 prefix
/// as needed) and consume its acknowledgement. Returns true on ACK (0xFA),
/// false on any other reply, or null on timeout.
pub fn sendToDevice(self: Controller, port: Port, byte: u8) ?bool {
if (port.deviceInputCommand()) |prefix| {
if (!sendCommand(self.device_id, self.status_index, prefix, default_wait_timeout_nanoseconds)) return null;
}
if (!writeData(self.device_id, self.status_index, self.data_index, byte, default_wait_timeout_nanoseconds)) return null;
const reply = readData(self.device_id, self.status_index, self.data_index, default_wait_timeout_nanoseconds) orelse return null;
return reply == device_response_acknowledge;
}
/// Discard any bytes sitting in the output buffer (for example the device-id
/// byte a mouse emits after a reset) so they cannot be mistaken for the reply
/// to a subsequent command.
pub fn drainOutputBuffer(self: Controller) void {
var guard: u8 = 0;
while (guard < 16) : (guard += 1) {
if (status(self.device_id, self.status_index) & status_output_buffer_full == 0) return;
_ = device.ioRead(self.device_id, self.data_index, 0, 1);
}
}
/// Ask the device on `port` what it is (command 0xF2) and classify the reply.
/// Scanning is disabled around the query so a streaming device cannot inject
/// data bytes that look like the identifier. Returns the device type, or null
/// if the identify command itself timed out.
pub fn identifyDevice(self: Controller, port: Port) ?DeviceType {
// Clear any leftover bytes (e.g. a post-reset mouse id) before we start.
self.drainOutputBuffer();
// Stop the device reporting so its data can't be mistaken for the reply.
if (self.sendToDevice(port, device_cmd_disable_scanning) == null) return null;
if (self.sendToDevice(port, device_cmd_identify) == null) return null;
// After the ACK, the device sends 0, 1, or 2 identifier bytes.
const first = readData(self.device_id, self.status_index, self.data_index, default_wait_timeout_nanoseconds);
const device_type: DeviceType = if (first) |id| switch (id) {
identify_keyboard_mf2 => blk: {
// A MF2 keyboard sends a second subtype byte; consume and ignore it.
_ = readData(self.device_id, self.status_index, self.data_index, default_wait_timeout_nanoseconds);
break :blk .keyboard;
},
identify_mouse_standard, identify_mouse_scroll, identify_mouse_five_button => .mouse,
else => .unknown,
} else
// No identifier bytes at all is a legacy AT keyboard.
.keyboard;
// Resume scanning so the device works once its driver takes over.
_ = self.sendToDevice(port, device_cmd_enable_scanning);
return device_type;
}
};
+389
View File
@@ -0,0 +1,389 @@
//! PS/2 scancode set 2 → USB HID usage decoding, plus the keyboard state a driver
//! needs on top of it (pressed keys, modifier tracking, caps-lock toggle).
//!
//! Set 2 is what a keyboard sends when the 8042's legacy set-1 translation is off —
//! which is how ps2-bus.zig deliberately configures the controller. A key's **make**
//! code is one byte (two with an `E0` prefix for the "extended" keys added after the
//! original AT layout); its **break** code is the same code behind an `F0` prefix.
//! Pause alone is an eight-byte `E1` sequence with no break.
//!
//! The output vocabulary is USB HID keyboard-page usages (a=4, enter=40, ...), the
//! same numbering the input protocol's `Keycode` and the xkeyboard-config layout
//! tables use — so a decoded usage indexes a layout directly.
//!
//! Everything here is pure (no imports beyond `std`, no IO), so it is host-testable:
//! the tests at the bottom run under `zig build test`.
const std = @import("std");
// --- USB HID usages the state machine itself needs to recognize --------------
pub const usage_caps_lock: u8 = 0x39;
pub const usage_left_control: u8 = 0xE0;
pub const usage_left_shift: u8 = 0xE1;
pub const usage_left_alt: u8 = 0xE2;
pub const usage_right_control: u8 = 0xE4;
pub const usage_right_shift: u8 = 0xE5;
pub const usage_right_alt: u8 = 0xE6; // AltGr — selects XKB level 3
// --- scancode set 2 → HID usage tables ---------------------------------------
/// Single-byte (non-`E0`) make codes. Zero means "no key" — protocol bytes (ACK,
/// BAT results) and reserved codes land there and decode to nothing.
pub const set2_base: [256]u8 = blk: {
var table = [_]u8{0} ** 256;
// function row
table[0x01] = 0x42; // F9
table[0x03] = 0x3E; // F5
table[0x04] = 0x3C; // F3
table[0x05] = 0x3A; // F1
table[0x06] = 0x3B; // F2
table[0x07] = 0x45; // F12
table[0x09] = 0x43; // F10
table[0x0A] = 0x41; // F8
table[0x0B] = 0x3F; // F6
table[0x0C] = 0x3D; // F4
table[0x78] = 0x44; // F11
table[0x83] = 0x40; // F7
// letters
table[0x1C] = 0x04; // A
table[0x32] = 0x05; // B
table[0x21] = 0x06; // C
table[0x23] = 0x07; // D
table[0x24] = 0x08; // E
table[0x2B] = 0x09; // F
table[0x34] = 0x0A; // G
table[0x33] = 0x0B; // H
table[0x43] = 0x0C; // I
table[0x3B] = 0x0D; // J
table[0x42] = 0x0E; // K
table[0x4B] = 0x0F; // L
table[0x3A] = 0x10; // M
table[0x31] = 0x11; // N
table[0x44] = 0x12; // O
table[0x4D] = 0x13; // P
table[0x15] = 0x14; // Q
table[0x2D] = 0x15; // R
table[0x1B] = 0x16; // S
table[0x2C] = 0x17; // T
table[0x3C] = 0x18; // U
table[0x2A] = 0x19; // V
table[0x1D] = 0x1A; // W
table[0x22] = 0x1B; // X
table[0x35] = 0x1C; // Y
table[0x1A] = 0x1D; // Z
// digit row
table[0x16] = 0x1E; // 1
table[0x1E] = 0x1F; // 2
table[0x26] = 0x20; // 3
table[0x25] = 0x21; // 4
table[0x2E] = 0x22; // 5
table[0x36] = 0x23; // 6
table[0x3D] = 0x24; // 7
table[0x3E] = 0x25; // 8
table[0x46] = 0x26; // 9
table[0x45] = 0x27; // 0
// control and whitespace
table[0x5A] = 0x28; // Enter
table[0x76] = 0x29; // Escape
table[0x66] = 0x2A; // Backspace
table[0x0D] = 0x2B; // Tab
table[0x29] = 0x2C; // Space
// punctuation
table[0x4E] = 0x2D; // - _
table[0x55] = 0x2E; // = +
table[0x54] = 0x2F; // [ {
table[0x5B] = 0x30; // ] }
table[0x5D] = 0x31; // \ | (non-US hash on ISO boards, same position)
table[0x4C] = 0x33; // ; :
table[0x52] = 0x34; // ' "
table[0x0E] = 0x35; // ` ~
table[0x41] = 0x36; // , <
table[0x49] = 0x37; // . >
table[0x4A] = 0x38; // / ?
table[0x61] = 0x64; // non-US backslash (the extra ISO key between shift and Z)
// locks
table[0x58] = usage_caps_lock;
table[0x77] = 0x53; // Num Lock
table[0x7E] = 0x47; // Scroll Lock
// keypad
table[0x7C] = 0x55; // keypad *
table[0x7B] = 0x56; // keypad -
table[0x79] = 0x57; // keypad +
table[0x69] = 0x59; // keypad 1
table[0x72] = 0x5A; // keypad 2
table[0x7A] = 0x5B; // keypad 3
table[0x6B] = 0x5C; // keypad 4
table[0x73] = 0x5D; // keypad 5
table[0x74] = 0x5E; // keypad 6
table[0x6C] = 0x5F; // keypad 7
table[0x75] = 0x60; // keypad 8
table[0x7D] = 0x61; // keypad 9
table[0x70] = 0x62; // keypad 0
table[0x71] = 0x63; // keypad .
// modifiers
table[0x14] = usage_left_control;
table[0x12] = usage_left_shift;
table[0x11] = usage_left_alt;
table[0x59] = usage_right_shift;
break :blk table;
};
/// `E0`-prefixed make codes. `E0 12` is the "fake shift" the keyboard wraps around
/// Print Screen and navigation keys when a real shift is involved; it maps to zero
/// here, so it decodes to nothing and only the real key comes through.
pub const set2_extended: [256]u8 = blk: {
var table = [_]u8{0} ** 256;
table[0x11] = usage_right_alt;
table[0x14] = usage_right_control;
table[0x1F] = 0xE3; // left GUI
table[0x27] = 0xE7; // right GUI
table[0x2F] = 0x65; // application (menu)
table[0x7C] = 0x46; // Print Screen (arrives as E0 12 E0 7C; the E0 12 decodes to nothing)
table[0x4A] = 0x54; // keypad /
table[0x5A] = 0x58; // keypad Enter
table[0x70] = 0x49; // Insert
table[0x6C] = 0x4A; // Home
table[0x7D] = 0x4B; // Page Up
table[0x71] = 0x4C; // Delete
table[0x69] = 0x4D; // End
table[0x7A] = 0x4E; // Page Down
table[0x74] = 0x4F; // right arrow
table[0x6B] = 0x50; // left arrow
table[0x72] = 0x51; // down arrow
table[0x75] = 0x52; // up arrow
break :blk table;
};
// --- the byte-stream decoder --------------------------------------------------
/// One decoded key transition: which key (as a USB HID usage) and whether this is
/// a make (press or typematic repeat) or a break (release).
pub const DecodedKey = struct {
usage: u8,
make: bool,
};
/// Turns the raw set-2 byte stream into `DecodedKey`s. Feed it every byte the
/// keyboard sends; most bytes complete a key and return one, prefix bytes return
/// null and arm the state machine for the next byte.
pub const Decoder = struct {
const State = enum {
idle,
extended, // saw E0
break_prefix, // saw F0
extended_break, // saw E0 F0
pause_skip, // inside the 8-byte E1 Pause sequence
};
state: State = .idle,
/// Bytes still to swallow in `pause_skip`.
skip: u8 = 0,
/// The whole Pause make sequence is `E1 14 77 E1 F0 14 F0 77` — seven bytes
/// after the leading `E1`, and no break sequence ever follows.
const pause_bytes_after_e1: u8 = 7;
pub fn feed(self: *Decoder, byte: u8) ?DecodedKey {
switch (self.state) {
.idle => switch (byte) {
0xE0 => self.state = .extended,
0xF0 => self.state = .break_prefix,
0xE1 => {
self.state = .pause_skip;
self.skip = pause_bytes_after_e1;
},
// Anything else is a make code — or a protocol byte (0xFA ACK,
// 0xAA BAT-passed, 0xEE echo, ...), which the tables map to zero.
else => return decoded(set2_base[byte], true),
},
.extended => switch (byte) {
0xF0 => self.state = .extended_break,
else => {
self.state = .idle;
return decoded(set2_extended[byte], true);
},
},
.break_prefix => {
self.state = .idle;
return decoded(set2_base[byte], false);
},
.extended_break => {
self.state = .idle;
return decoded(set2_extended[byte], false);
},
.pause_skip => {
self.skip -= 1;
if (self.skip == 0) self.state = .idle;
},
}
return null;
}
fn decoded(usage: u8, make: bool) ?DecodedKey {
if (usage == 0) return null; // unmapped or a protocol byte
return .{ .usage = usage, .make = make };
}
};
// --- driver-side keyboard state -----------------------------------------------
/// What a key transition did, plus the modifier state to stamp on the resulting
/// events (snapshotted after the transition was applied).
pub const Transition = struct {
pub const Action = enum {
pressed, // physical make of a key that was up
repeated, // typematic make of a key already down — no new key_down
released, // physical break
};
action: Action,
modifiers: ModifierSnapshot,
};
/// The modifier state at one instant, in both vocabularies a driver needs: the
/// input protocol's coarse bits (shift/control/alt) and the level-selection
/// inputs xkeyboard-config takes (shift, caps_lock, AltGr as level3).
pub const ModifierSnapshot = struct {
shift: bool, // either shift held
control: bool, // either control held
alt: bool, // either alt held (including AltGr)
right_alt: bool, // AltGr specifically — the XKB level-3 selector
caps_lock: bool, // the toggle, not the key
};
/// Tracks which keys are physically down and the caps-lock toggle, and classifies
/// each decoded transition. Pure state — no IO — so repeat detection and modifier
/// snapshots are host-testable.
pub const KeyboardState = struct {
/// One bit per HID usage: set while the key is physically down.
pressed: [32]u8 = [_]u8{0} ** 32,
caps_lock: bool = false,
pub fn apply(self: *KeyboardState, key: DecodedKey) Transition {
const already_down = self.isPressed(key.usage);
if (key.make) {
if (!already_down) {
self.setPressed(key.usage, true);
if (key.usage == usage_caps_lock) self.caps_lock = !self.caps_lock;
}
return .{
.action = if (already_down) .repeated else .pressed,
.modifiers = self.snapshot(),
};
}
self.setPressed(key.usage, false);
return .{ .action = .released, .modifiers = self.snapshot() };
}
pub fn isPressed(self: *const KeyboardState, usage: u8) bool {
return self.pressed[usage / 8] & (@as(u8, 1) << @intCast(usage % 8)) != 0;
}
fn setPressed(self: *KeyboardState, usage: u8, down: bool) void {
const bit = @as(u8, 1) << @intCast(usage % 8);
if (down) {
self.pressed[usage / 8] |= bit;
} else {
self.pressed[usage / 8] &= ~bit;
}
}
fn snapshot(self: *const KeyboardState) ModifierSnapshot {
const right_alt = self.isPressed(usage_right_alt);
return .{
.shift = self.isPressed(usage_left_shift) or self.isPressed(usage_right_shift),
.control = self.isPressed(usage_left_control) or self.isPressed(usage_right_control),
.alt = self.isPressed(usage_left_alt) or right_alt,
.right_alt = right_alt,
.caps_lock = self.caps_lock,
};
}
};
// --- tests (host-run via `zig build test`) ------------------------------------
const testing = std.testing;
/// Feed `bytes` and return the single DecodedKey they should produce (fails the
/// test if they produce none or more than one).
fn feedOne(decoder: *Decoder, bytes: []const u8) !DecodedKey {
var result: ?DecodedKey = null;
for (bytes) |byte| {
if (decoder.feed(byte)) |key| {
try testing.expect(result == null);
result = key;
}
}
return result orelse error.TestExpectedResult;
}
fn feedNone(decoder: *Decoder, bytes: []const u8) !void {
for (bytes) |byte| try testing.expectEqual(@as(?DecodedKey, null), decoder.feed(byte));
}
test "base make and break: A" {
var decoder = Decoder{};
try testing.expectEqual(DecodedKey{ .usage = 0x04, .make = true }, try feedOne(&decoder, &.{0x1C}));
try testing.expectEqual(DecodedKey{ .usage = 0x04, .make = false }, try feedOne(&decoder, &.{ 0xF0, 0x1C }));
}
test "extended make and break: right arrow" {
var decoder = Decoder{};
try testing.expectEqual(DecodedKey{ .usage = 0x4F, .make = true }, try feedOne(&decoder, &.{ 0xE0, 0x74 }));
try testing.expectEqual(DecodedKey{ .usage = 0x4F, .make = false }, try feedOne(&decoder, &.{ 0xE0, 0xF0, 0x74 }));
}
test "pause: the E1 sequence is consumed silently" {
var decoder = Decoder{};
try feedNone(&decoder, &.{ 0xE1, 0x14, 0x77, 0xE1, 0xF0, 0x14, 0xF0, 0x77 });
// The decoder is back in idle: an ordinary key still decodes.
try testing.expectEqual(DecodedKey{ .usage = 0x04, .make = true }, try feedOne(&decoder, &.{0x1C}));
}
test "protocol bytes decode to nothing" {
var decoder = Decoder{};
try feedNone(&decoder, &.{ 0xFA, 0xAA, 0xEE }); // ACK, BAT-passed, echo
}
test "print screen: the fake-shift E0 12 decodes to nothing" {
var decoder = Decoder{};
try feedNone(&decoder, &.{ 0xE0, 0x12 });
try testing.expectEqual(DecodedKey{ .usage = 0x46, .make = true }, try feedOne(&decoder, &.{ 0xE0, 0x7C }));
}
test "typematic repeat is classified, not re-pressed" {
var state = KeyboardState{};
const a = DecodedKey{ .usage = 0x04, .make = true };
try testing.expectEqual(Transition.Action.pressed, state.apply(a).action);
try testing.expectEqual(Transition.Action.repeated, state.apply(a).action);
try testing.expectEqual(Transition.Action.repeated, state.apply(a).action);
try testing.expectEqual(Transition.Action.released, state.apply(.{ .usage = 0x04, .make = false }).action);
try testing.expectEqual(Transition.Action.pressed, state.apply(a).action);
}
test "shift held shows in the snapshot of other keys" {
var state = KeyboardState{};
_ = state.apply(.{ .usage = usage_left_shift, .make = true });
const transition = state.apply(.{ .usage = 0x04, .make = true });
try testing.expect(transition.modifiers.shift);
try testing.expect(!transition.modifiers.control);
_ = state.apply(.{ .usage = usage_left_shift, .make = false });
_ = state.apply(.{ .usage = 0x04, .make = false });
try testing.expect(!state.apply(.{ .usage = 0x04, .make = true }).modifiers.shift);
}
test "right alt reports both alt and the level-3 selector" {
var state = KeyboardState{};
_ = state.apply(.{ .usage = usage_right_alt, .make = true });
const transition = state.apply(.{ .usage = 0x04, .make = true });
try testing.expect(transition.modifiers.alt);
try testing.expect(transition.modifiers.right_alt);
}
test "caps lock toggles on make, not on repeat or break" {
var state = KeyboardState{};
try testing.expect(state.apply(.{ .usage = usage_caps_lock, .make = true }).modifiers.caps_lock);
try testing.expect(state.apply(.{ .usage = usage_caps_lock, .make = true }).modifiers.caps_lock); // repeat
try testing.expect(state.apply(.{ .usage = usage_caps_lock, .make = false }).modifiers.caps_lock);
try testing.expect(!state.apply(.{ .usage = usage_caps_lock, .make = true }).modifiers.caps_lock); // second press: off
}
@@ -0,0 +1,191 @@
//! /system/drivers/usb-xhci-bus — the xHCI (USB 3) host-controller bus driver.
//! The device manager spawns **one instance per controller** it discovers (a
//! machine can carry several), passing the controller's device-tree id as
//! argv[1]; this instance claims that device and no other, so multiple
//! instances never fight over hardware.
//!
//! M18.2 (this increment): after the hello, real hardware — map the xHC's
//! register window (the first memory BAR; resource 0 is the ECAM config
//! space), read the capability registers, and walk the root-hub ports: one
//! `child_added` report to the manager per connected port, carrying the port
//! number and the PORTSC speed class as identity. No transfer rings yet —
//! descriptors and USB class matching are the USB track; the connect bit and
//! speed come straight from PORTSC, which reflects hardware state whether or
//! not the controller is running.
const std = @import("std");
const runtime = @import("runtime");
const protocol = runtime.device_manager_protocol;
const device = runtime.device;
/// Format one whole log line and emit it in a single `debug_write`, so
/// concurrent instances (one per controller) can never interleave mid-line.
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
var controller_id: u64 = protocol.no_device;
/// Claim the assigned controller, find its register window, and hello the
/// manager. Any failure returns false: the process exits cleanly, which the
/// manager reads as "meant to stop" — a missing assignment is not a crash loop.
fn initialise(endpoint: runtime.ipc.Handle) bool {
_ = endpoint;
if (!device.claim(controller_id)) {
writeLine("usb-xhci-bus: unable to claim controller device {d}\n", .{controller_id});
return false;
}
// Fetch our own descriptor back for the controller's resources.
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("usb-xhci-bus: out of memory\n");
return false;
};
const total = device.enumerate(buffer);
const descriptor = for (buffer[0..@min(total, buffer.len)]) |d| {
if (d.id == controller_id) break d;
} else {
writeLine("usb-xhci-bus: device {d} not in the device tree\n", .{controller_id});
return false;
};
// The xHC's registers live behind the first memory BAR. Resource 0 is the
// function's ECAM configuration space (M15), so the walk starts at 1.
var register_index: u64 = 0;
const register_window = for (descriptor.resources[1..@intCast(descriptor.resource_count)], 1..) |resource, index| {
if (resource.kind == @intFromEnum(device.ResourceKind.memory)) {
register_index = index;
break resource;
}
} else {
writeLine("usb-xhci-bus: controller device {d} has no register BAR\n", .{controller_id});
return false;
};
writeLine("usb-xhci-bus: claimed controller device {d} (registers at 0x{x}, {d} bytes)\n", .{
controller_id,
register_window.start,
register_window.len,
});
register_base = device.mmioMap(controller_id, register_index) orelse {
_ = runtime.system.write("usb-xhci-bus: mmio_map failed\n");
return false;
};
// The handshake: role, protocol version, assignment — inside the manager's
// deadline (the lookup retries cover the manager still registering).
var manager: ?runtime.ipc.Handle = null;
var tries: u32 = 0;
while (manager == null and tries < 100) : (tries += 1) {
manager = runtime.ipc.lookup(.device_manager);
if (manager == null) runtime.system.sleep(20);
}
const h = manager orelse {
_ = runtime.system.write("usb-xhci-bus: no device manager to hello\n");
return false;
};
const hello = protocol.Hello{ .role = @intFromEnum(protocol.Role.bus), .device_id = controller_id };
var reply: [protocol.message_maximum]u8 = undefined;
const n = runtime.ipc.call(h, std.mem.asBytes(&hello), &reply) catch {
_ = runtime.system.write("usb-xhci-bus: hello call failed\n");
return false;
};
if (n < protocol.reply_size or std.mem.bytesToValue(protocol.HelloReply, reply[0..protocol.reply_size]).status != 0) {
_ = runtime.system.write("usb-xhci-bus: hello refused\n");
return false;
}
_ = runtime.system.write("usb-xhci-bus: hello acknowledged\n");
scanPorts(h);
return true;
}
var register_base: usize = 0;
/// One 32-bit volatile register read at `offset` from the mapped window.
fn readRegister(offset: usize) u32 {
const register: *volatile u32 = @ptrFromInt(register_base + offset);
return register.*;
}
/// The xHCI default Protocol Speed IDs (the PORTSC port-speed field, bits 13:10)
/// decoded to human names — the boot-log breadcrumb for what actually enumerated on
/// a port, the USB analog of the pci-bus class-code line. A controller may redefine
/// these through its Supported Protocol capability, but the defaults cover every
/// speed QEMU and real hardware report at this (pre-descriptor) stage.
fn speedName(speed: u32) []const u8 {
return switch (speed) {
1 => "Full-speed (USB 2.0, 12 Mb/s)",
2 => "Low-speed (USB 2.0, 1.5 Mb/s)",
3 => "High-speed (USB 2.0, 480 Mb/s)",
4 => "SuperSpeed (USB 3.0, 5 Gb/s)",
5 => "SuperSpeedPlus (USB 3.1, 10 Gb/s)",
else => "unknown speed",
};
}
/// The root-hub port scan: read the capability registers for the port count
/// and the operational-register offset, then one PORTSC per port. The connect
/// bit (CCS) and the speed field reflect hardware state directly — no
/// controller reset or run needed to *see* the devices; driving them needs the
/// rings (the USB track).
fn scanPorts(manager: runtime.ipc.Handle) void {
// Capability registers: CAPLENGTH is byte 0 of the first dword; HCSPARAMS1
// carries MaxPorts in bits 31:24.
const capability_length = readRegister(0) & 0xFF;
const structural = readRegister(0x04);
const maximum_ports: u32 = structural >> 24;
writeLine("usb-xhci-bus: {d} root-hub ports\n", .{maximum_ports});
// PORTSC registers: operational base + 0x400 + 0x10 per port (1-based).
var port: u32 = 1;
var connected: u32 = 0;
while (port <= maximum_ports) : (port += 1) {
const port_status = readRegister(capability_length + 0x400 + 0x10 * (port - 1));
if (port_status & 1 == 0) continue; // CCS: nothing connected
connected += 1;
const speed = (port_status >> 10) & 0xF; // the PORTSC port-speed class
writeLine("usb-xhci-bus: port {d} connected — {s} (speed class {d})\n", .{ port, speedName(speed), speed });
const report = protocol.ChildAdded{
.parent = controller_id,
.bus_address = port,
.identity = speed,
};
var reply: [protocol.message_maximum]u8 = undefined;
_ = runtime.ipc.call(manager, std.mem.asBytes(&report), &reply) catch {
writeLine("usb-xhci-bus: child report for port {d} failed\n", .{port});
continue;
};
}
if (connected == 0) _ = runtime.system.write("usb-xhci-bus: no devices connected\n");
}
/// No bus protocol to serve yet — transfer requests arrive with the USB track.
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = message;
_ = reply;
_ = sender;
_ = capability;
return 0;
}
pub fn main(init: runtime.process.Init) void {
const argument = init.arguments.get(1) orelse {
_ = runtime.system.write("usb-xhci-bus: missing controller device id (argv[1])\n");
return;
};
controller_id = std.fmt.parseInt(u64, argument, 10) catch {
writeLine("usb-xhci-bus: malformed controller device id '{s}'\n", .{argument});
return;
};
runtime.service.run(protocol.message_maximum, .{
.init = initialise,
.on_message = onMessage,
});
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
+1 -2
View File
@@ -490,8 +490,7 @@ pub fn saveInterrupts() u64 {
\\cli
: [f] "=r" (flags),
:
: .{ .memory = true }
);
: .{ .memory = true });
return flags;
}
+24 -21
View File
@@ -3,9 +3,11 @@
//! triple-faults and silently resets the machine. With it, the CPU vectors into
//! our stubs, which capture the register state and hand it to a dispatcher.
//!
//! Vectors split in two: 0-31 are CPU exceptions (terminal — reported and
//! halted); 32+ are device interrupts (a registered handler runs, the APIC is
//! acknowledged, and we return to the interrupted code).
//! Vectors split in two: 0-31 are CPU exceptions, handed to the `on_fault` hook
//! and never returned from (the kernel's handler kills a faulting user process
//! and reschedules, or halts the core for a kernel-mode fault); 32+ are device
//! interrupts (a registered handler runs, the APIC is acknowledged, and we
//! return to the interrupted code).
const gdt = @import("gdt.zig");
const tss = @import("tss.zig");
@@ -76,22 +78,22 @@ fn defaultFault(_: *const CpuState) noreturn {
/// Names for the 32 defined exception vectors, for readable output.
const names = [_][]const u8{
"divide error", "debug",
"NMI", "breakpoint",
"overflow", "bound range exceeded",
"invalid opcode", "device not available",
"double fault", "coprocessor segment overrun",
"invalid TSS", "segment not present",
"stack-segment fault", "general protection fault",
"page fault", "reserved (15)",
"x87 floating-point", "alignment check",
"machine check", "SIMD floating-point",
"virtualization", "control protection",
"reserved (22)", "reserved (23)",
"reserved (24)", "reserved (25)",
"reserved (26)", "reserved (27)",
"hypervisor injection", "VMM communication",
"security exception", "reserved (31)",
"divide error", "debug",
"NMI", "breakpoint",
"overflow", "bound range exceeded",
"invalid opcode", "device not available",
"double fault", "coprocessor segment overrun",
"invalid TSS", "segment not present",
"stack-segment fault", "general protection fault",
"page fault", "reserved (15)",
"x87 floating-point", "alignment check",
"machine check", "SIMD floating-point",
"virtualization", "control protection",
"reserved (22)", "reserved (23)",
"reserved (24)", "reserved (25)",
"reserved (26)", "reserved (27)",
"hypervisor injection", "VMM communication",
"security exception", "reserved (31)",
};
pub fn vectorName(vector: u64) []const u8 {
@@ -162,8 +164,9 @@ pub fn loadOnThisCpu() void {
}
/// Called by isr_common (isr.s) with a pointer to the trap frame. Exported so the
/// assembly stubs can `call` it by name. Exceptions are terminal; device
/// interrupts run their handler, get acknowledged, and return.
/// assembly stubs can `call` it by name. Exceptions never return here (on_fault
/// kills the faulting process or halts the core); device interrupts run their
/// handler, get acknowledged, and return.
export fn interruptDispatch(state: *CpuState) callconv(.c) void {
if (state.vector < 32) {
on_fault(state); // CPU exception — never returns
+2 -4
View File
@@ -166,8 +166,7 @@ pub fn init(allocFrame: *const fn () ?u64, freeFrame: *const fn (u64) void, boot
asm volatile ("mov %[pml4], %%cr3"
:
: [pml4] "r" (pml4),
: .{ .memory = true }
);
: .{ .memory = true });
on_own_tables = true; // now on the kernel's physmap (covers all RAM)
init_done = true; // the kernel half is fixed from here
}
@@ -409,6 +408,5 @@ fn invalidate(virtual: u64) void {
\\invlpg (%%rax)
:
: [v] "r" (virtual),
: .{ .rax = true, .memory = true }
);
: .{ .rax = true, .memory = true });
}
-2
View File
@@ -154,5 +154,3 @@ pub const Console = struct {
while (x < self.fb.width) : (x += 1) destination[x] = source[x];
}
};
+45 -1
View File
@@ -66,6 +66,7 @@ fn record(node: *platform.Device, parent_id: u64) u64 {
d.id = count;
d.parent = parent_id;
d.class = @intFromEnum(node.class);
d.pci_class = if (node.ids.pci_class) |code| code else device_abi.no_pci_class;
const h = node.hid();
d.hid_len = @min(h.len, d.hid.len);
@memcpy(d.hid[0..d.hid_len], h[0..d.hid_len]);
@@ -103,6 +104,19 @@ pub fn ownerOf(id: u64) ?u32 {
return claimed[@intCast(id)];
}
/// Release every claim held by `owner` — called by the process layer on every
/// path out of a process (exit, fault, kill), so a restarted driver can claim its
/// hardware again (docs/process-lifecycle.md iron rule 1: cleanup is the kernel's
/// job). The devices stay in the table — they describe hardware, which did not go
/// away — only their ownership clears.
pub fn releaseAllOwnedBy(owner: u32) void {
for (claimed[0..count]) |*slot| {
if (slot.*) |o| {
if (o == owner) slot.* = null;
}
}
}
/// Resource `index` of device `id`, or null if out of range.
pub fn resourceOf(id: u64, index: u64) ?device_abi.ResourceDescriptor {
if (id >= count) return null;
@@ -117,7 +131,16 @@ pub fn resourceOf(id: u64, index: u64) ?device_abi.ResourceDescriptor {
/// and would otherwise vacuously "fit" anywhere.
fn contains(parent: device_abi.ResourceDescriptor, child: device_abi.ResourceDescriptor) bool {
if (parent.kind != child.kind) return false;
if (child.kind == @intFromEnum(device_abi.ResourceKind.irq)) return parent.start == child.start;
if (child.kind == @intFromEnum(device_abi.ResourceKind.irq)) {
// Range containment: an interrupt line is still indivisible (a child owns
// exactly one GSI), but a parent may own a *range* of lines so a broad
// owner — the acpi-tables node, whose firmware names any legacy IRQ —
// can contain its children's specific lines. A length-1 parent range is
// exactly the old equality rule, so existing single-IRQ parents are
// unaffected.
const span = if (parent.len == 0) 1 else parent.len;
return child.start >= parent.start and child.start < parent.start + span;
}
if (child.len == 0 or parent.len == 0) return false;
// No overflow: a resource that wraps the address space is not containable.
const child_end = std.math.add(u64, child.start, child.len) catch return false;
@@ -166,10 +189,31 @@ pub fn register(parent_id: u64, owner: u32, descriptor: *const device_abi.Device
if (!ok) return error.NotContained;
}
// Idempotent on exact match (docs/m19-m20-plan.md decision 3): a restarted
// registering bus re-registers what it rediscovers, and the table has no
// unregister — an identical (class, identity, resources) child under the
// same parent returns the existing id instead of appending a duplicate.
for (devices[0..count]) |*existing| {
if (existing.parent != parent_id) continue;
if (existing.class != descriptor.class) continue;
if (existing.pci_class != descriptor.pci_class) continue;
if (existing.hid_len != descriptor.hid_len) continue;
if (!std.mem.eql(u8, existing.hid[0..@intCast(existing.hid_len)], descriptor.hid[0..@intCast(descriptor.hid_len)])) continue;
if (existing.resource_count != descriptor.resource_count) continue;
var same = true;
for (0..@intCast(descriptor.resource_count)) |i| {
const a = existing.resources[i];
const b = descriptor.resources[i];
if (a.kind != b.kind or a.start != b.start or a.len != b.len) same = false;
}
if (same) return existing.id;
}
var d = std.mem.zeroes(device_abi.DeviceDescriptor);
d.id = count;
d.parent = parent_id;
d.class = descriptor.class;
d.pci_class = descriptor.pci_class;
d.hid_len = @min(descriptor.hid_len, d.hid.len);
@memcpy(d.hid[0..@intCast(d.hid_len)], descriptor.hid[0..@intCast(d.hid_len)]);
d.resource_count = descriptor.resource_count;
+107 -1
View File
@@ -46,6 +46,9 @@ pub const EFAULT: i64 = 3; // buffer unmapped / out of the user half
pub const ENOENT: i64 = 4; // no such registered service
pub const ENOSPC: i64 = 5; // handle table or registry full
pub const ENOMEM: i64 = 6; // out of memory
pub const EPEER: i64 = 7; // peer died before replying (its process exited or was killed)
pub const ESRCH: i64 = 8; // no such process (process_kill of an unknown/dead id)
pub const EPERM: i64 = 9; // not permitted (process_kill by anyone but the supervisor)
/// A badge with this bit set is an asynchronous notification (e.g. an IRQ), not a
/// message from a client — there is no reply owed. The low bits carry the source
@@ -54,6 +57,29 @@ pub const ENOMEM: i64 = 6; // out of memory
/// shared kernel↔user ABI (system/abi.zig), because ring 3 has to test the same bit.
pub const notify_badge_bit: u64 = abi.notify_badge_bit;
/// Set (with `notify_badge_bit`) when a `replyWait` wake carries a buffered payload
/// posted by `send` (`ipc_send`), rather than a bare IRQ/exit notification. Shared with
/// ring 3 through the ABI so the receiver can tell "a message arrived" from "the hardware
/// spoke".
pub const notify_message_bit: u64 = abi.notify_message_bit;
/// Largest payload a single `send` (`ipc_send`) may post. Kept small — the payload rides
/// inline in every `Endpoint`, and the async path is for events (a `KeyEvent` is 16
/// bytes), not bulk transfer, which is what `call` and future shared pages are for.
pub const POST_MAXIMUM: usize = 64;
/// Depth of an endpoint's async payload ring. Absorbs a burst while a receiver is briefly
/// busy; a full ring drops the *oldest* message (see `send`).
const post_capacity: usize = 16;
/// One buffered message: a length-prefixed payload plus the sender's task id (delivered
/// in the low bits of the receiver's badge).
const PostSlot = struct {
length: u16 = 0,
sender_id: u64 = 0,
bytes: [POST_MAXIMUM]u8 = undefined,
};
/// End of the user (low) canonical half — user buffers must lie below it.
const user_half_end: u64 = 0x0000_8000_0000_0000;
@@ -71,6 +97,12 @@ pub const Endpoint = struct {
notify_buffer: [8]u64 = undefined,
notify_head: u8 = 0,
notify_tail: u8 = 0,
// Pending buffered messages (payloads posted by `send`), a small FIFO ring. Unlike
// notifications — which are a level and coalesce — these are discrete messages, so a
// full ring drops the oldest rather than merging.
post_buffer: [post_capacity]PostSlot = undefined,
post_head: u16 = 0,
post_tail: u16 = 0,
};
pub fn createIpcEndpoint() ?*Endpoint {
@@ -92,6 +124,7 @@ pub fn dropRef(endpoint: *Endpoint) void {
// --- sender FIFO (endpoint-local, via Task.next) ----------------------------
fn enqueueSender(endpoint: *Endpoint, t: *Task) void {
t.ipc_wait_endpoint = @ptrCast(endpoint); // so a kill can unlink a parked caller
t.next = null;
if (endpoint.sender_tail) |tail| tail.next = t else endpoint.sender_head = t;
endpoint.sender_tail = t;
@@ -101,10 +134,34 @@ fn dequeueSender(endpoint: *Endpoint) ?*Task {
const t = endpoint.sender_head orelse return null;
endpoint.sender_head = t.next;
if (endpoint.sender_head == null) endpoint.sender_tail = null;
t.ipc_wait_endpoint = null;
t.next = null;
return t;
}
/// Unlink `t` from the sender FIFO it queues in, if any — the kill path for a
/// client parked in `call` that no server has received yet. Without this, a dead
/// caller would later be dequeued as a dangling pointer. The endpoint is still
/// alive here: `t`'s own handle table holds a reference until closeHandles runs
/// (which the kill path does *after* this). Precondition: the big kernel lock is
/// held.
pub fn abandonSenderLocked(t: *Task) void {
const endpoint: *Endpoint = @ptrCast(@alignCast(t.ipc_wait_endpoint orelse return));
t.ipc_wait_endpoint = null;
var previous: ?*Task = null;
var node = endpoint.sender_head;
while (node) |n| : ({
previous = n;
node = n.next;
}) {
if (n != t) continue;
if (previous) |p| p.next = t.next else endpoint.sender_head = t.next;
if (endpoint.sender_tail == t) endpoint.sender_tail = previous;
t.next = null;
return;
}
}
// --- cross-address-space copy ----------------------------------------------
/// Copy `len` bytes from `source_va` in address space `source_as` to `destination_va` in
@@ -237,12 +294,23 @@ pub fn replyWait(endpoint: *Endpoint, reply_ptr: u64, reply_len: u64, receive_pt
scheduler.readyLocked(client); // its `call` now returns
}
// (2) Receive the next request (or notification), blocking until one is ready.
// (2) Receive the next request (or notification / buffered message), blocking until
// one is ready. Bare notifications (IRQ/exit) come first — they're latency-sensitive
// and carry no payload — then buffered messages, then synchronous client requests.
while (true) {
if (popNotify(endpoint)) |badge| {
out_badge.* = badge | notify_badge_bit;
return 0; // notification: no payload, no reply owed, no cap
}
if (popPost(endpoint)) |slot| {
const n = @min(@as(usize, slot.length), receive_cap);
// Copy from the kernel-resident ring slot (source aspace 0) into the receiver.
if (!copyAcross(0, @intFromPtr(&slot.bytes), me.aspace, receive_ptr, n)) {
continue; // bad receive buffer: drop this message, keep serving
}
out_badge.* = slot.sender_id | notify_badge_bit | notify_message_bit;
return @intCast(n); // async message: payload delivered, no reply owed, no cap
}
if (dequeueSender(endpoint)) |caller| {
const n = @min(caller.ipc_send_len, receive_cap);
if (!copyAcross(caller.aspace, caller.ipc_send_ptr, me.aspace, receive_ptr, n)) {
@@ -278,6 +346,44 @@ fn popNotify(endpoint: *Endpoint) ?u64 {
return badge;
}
/// Take the oldest buffered message from the post ring, or null if empty. Returns a
/// pointer into the endpoint's own storage — valid until the next `send`/`popPost` under
/// the same lock region, which is all the copy-out in `replyWait` needs.
fn popPost(endpoint: *Endpoint) ?*const PostSlot {
if (endpoint.post_head == endpoint.post_tail) return null;
const slot = &endpoint.post_buffer[endpoint.post_head % post_capacity];
endpoint.post_head +%= 1;
return slot;
}
/// Client-free side of async IPC (`ipc_send`): copy `[source_va, len)` from address space
/// `source_as` into `endpoint`'s post ring and wake a waiting receiver — **without
/// blocking the sender** and with no reply owed. `sender_id` rides along, delivered in the
/// low bits of the receiver's badge. Returns 0, or a negative errno (`-E2BIG` if the
/// payload exceeds `POST_MAXIMUM`, `-EFAULT` if the source buffer is unmapped / out of the
/// user half). A full ring drops the *oldest* message (advancing `post_head`), because a
/// buffered message is discrete, not a level: keeping the newest keeps input responsive.
/// Precondition: the big kernel lock is held.
pub fn sendLocked(endpoint: *Endpoint, source_as: u64, source_va: u64, len: u64, sender_id: u64) i64 {
if (len > POST_MAXIMUM) return -E2BIG;
// Drop the oldest if the ring is full, so this newest message always lands.
if (endpoint.post_tail -% endpoint.post_head >= post_capacity) endpoint.post_head +%= 1;
const slot = &endpoint.post_buffer[endpoint.post_tail % post_capacity];
if (!copyFromUser(source_as, source_va, slot.bytes[0..@intCast(len)])) return -EFAULT;
slot.length = @intCast(len);
slot.sender_id = sender_id;
endpoint.post_tail +%= 1;
scheduler.wakeLocked(&endpoint.receive_wait_queue);
return 0;
}
/// `sendLocked` wrapped in its own critical section, for the `ipc_send` syscall path.
pub fn send(endpoint: *Endpoint, source_as: u64, source_va: u64, len: u64, sender_id: u64) i64 {
const flags = sync.enter();
defer sync.leave(flags);
return sendLocked(endpoint, source_as, source_va, len, sender_id);
}
/// Post an asynchronous notification carrying `badge` to `endpoint` and wake a waiting
/// receiver. Precondition: the big kernel lock is held.
///
+50 -8
View File
@@ -87,7 +87,7 @@ fn kmain(boot_information: *const BootInformation) noreturn {
log.print(" pitch : {d} bytes\n", .{fb.pitch});
log.print(" format : {s}\n", .{@tagName(fb.format)});
log.print(" framebuffer: 0x{x:0>16}\n", .{fb.base});
log.print (" footprint : {d} MiB\n", .{(fb.pitch * fb.height) / (1024 * 1024)});
log.print(" footprint : {d} MiB\n", .{(fb.pitch * fb.height) / (1024 * 1024)});
// Summarise the physical memory the loader handed us. The array is danos's
// own MemoryRegion, so this is a plain slice — no firmware layout in sight.
@@ -284,7 +284,7 @@ fn kmain(boot_information: *const BootInformation) noreturn {
if (boot_information.init_len != 0) {
status("starting /system/services/init...\n");
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len];
process.spawnProcess(image, 4) catch |err| {
process.spawnProcess(image, 4, &.{"/system/services/init"}) catch |err| {
statusPrint("/system/services/init failed to load: {s}\n", .{@errorName(err)});
};
} else {
@@ -385,13 +385,55 @@ fn kib(frames: u64) u64 {
return frames * abi.page_size / (1024);
}
/// Report a CPU exception and halt **this core**. There's no fault recovery yet, so
/// the faulting core is terminal — but the fault is *contained* to it: on an
/// application processor only that core stops, and the rest of the system keeps
/// running (full recovery — kill the task, keep the core — is the resilience track,
/// see docs/resilience.md). The report names the core so an AP fault is attributed,
/// and goes to every output sink plus a POST code and a persistent breadcrumb.
/// Whether a ring-3 exception is attributable to the process that raised it — and
/// therefore recoverable by killing that process. NMI (2), double fault (8), and
/// machine check (18) report machine or kernel trouble even when they arrive with a
/// user CS (an NMI interrupts whatever happens to be running), so they stay terminal.
fn recoverableFault(vector: u64) bool {
return switch (vector) {
2, 8, 18 => false,
else => true,
};
}
/// Report a CPU exception. Two outcomes (docs/resilience.md):
///
/// **A fault taken in user mode kills the faulting process, not the machine.** The
/// kernel is intact — the CPU trapped onto the task's kernel stack — so the process
/// is killed, everything it held (address space, IRQ bindings, IPC handles, device
/// grants' frames) is reclaimed, and the core reschedules. A crashing driver takes
/// itself down, never the OS.
///
/// **Everything else halts this core.** A kernel-mode fault means the trusted base
/// itself is broken — there is nothing safe to kill — and NMI/#DF/#MC report machine
/// trouble regardless of CS (`recoverableFault`). Even then the fault is *contained*:
/// on an application processor only that core stops and the rest keep running. The
/// report names the core so an AP fault is attributed, and goes to every output sink
/// plus a POST code and a persistent breadcrumb. (A ring-3 fault on a *borrowed*
/// kernel thread — process.run, the user-pf isolation probe — also lands here: there
/// is no scheduled process to kill.)
/// Classify a CPU exception vector as the ExitReason a supervisor reads — the
/// fault classes of docs/process-lifecycle.md. Faults are exit reasons, never
/// signals delivered to the faulting process: recovery is restart, not a handler.
fn exitReasonForVector(vector: u64) abi.ExitReason {
return switch (vector) {
14 => .segmentation_fault, // page fault
6 => .illegal_instruction, // invalid opcode
0, 16, 19 => .arithmetic_fault, // divide error, x87, SIMD
13 => .protection_fault, // general protection
else => .fault,
};
}
fn onException(state: *const architecture.CpuState) noreturn {
if (architecture.fromUser(state) and scheduler.currentIsUserProcess() and recoverableFault(state.vector)) {
statusPrint("\ndanos: process {d} ({s}) killed by {s} (vector {d}) on core {d}\n", .{ scheduler.currentId(), scheduler.current().name(), architecture.exceptionName(state.vector), state.vector, scheduler.currentCpuIndex() });
statusPrint(" error code : 0x{x}\n", .{state.error_code});
statusPrint(" IP : 0x{x:0>16}\n", .{architecture.instructionPointer(state)});
if (architecture.faultAddress(state)) |address| statusPrint(" fault addr : 0x{x:0>16}\n", .{address});
process.killCurrentProcess(exitReasonForVector(state.vector)); // reclaims everything, reschedules; never returns
}
log.checkpoint(cp_exception);
const core = scheduler.currentCpuIndex();
// A fault is user-facing enough to paint on screen too (via statusPrint), on
+579 -45
View File
@@ -24,6 +24,7 @@ const elf = std.elf;
const boot_handoff = @import("boot-handoff");
const abi = @import("abi");
const device_abi = @import("device-abi");
const parameters = @import("parameters");
const architecture = @import("architecture");
const pmm = @import("pmm.zig");
const scheduler = @import("scheduler.zig");
@@ -40,9 +41,18 @@ const SystemCall = abi.SystemCall;
/// User virtual addresses. PML4 index 224 — a user-exclusive region, far from
/// the identity map (low indices) and the vmm test address (index 128), so
/// setting the U/S bit on its intermediate tables widens no kernel mapping.
/// An ELF image may occupy [code_virtual, stack_virtual); the stack page sits above.
/// An ELF image may occupy [code_virtual, stack_virtual); the stack sits above.
pub const code_virtual: u64 = 0x0000_7000_0000_0000;
pub const stack_virtual: u64 = 0x0000_7000_0020_0000;
/// The stack region, above the image. The page at `stack_virtual` is **never
/// mapped** — it is the guard page: a process that overflows its stack walks into
/// it and faults (killing only that process) rather than silently corrupting the
/// top of its own image. The stack proper is `parameters.user_stack_pages` pages
/// at [stack_base_virtual, stack_top_virtual), RW + NX, with the System V entry
/// block (argc/argv) at the very top.
pub const stack_virtual: u64 = 0x0000_7000_0020_0000; // guard page (unmapped)
pub const stack_base_virtual: u64 = stack_virtual + page_size;
pub const stack_top_virtual: u64 = stack_base_virtual + parameters.user_stack_pages * page_size;
/// The mmap grant arena: where `mmap` hands out fresh user pages, above the image
/// and stack but still inside PML4[224] (so no kernel mapping is widened). Each
@@ -73,6 +83,19 @@ pub const dma_arena_end: u64 = dma_arena_base + (256 << 20); // 256 MiB per proc
/// chunks, so this bound is generous; it also caps the frame scratch array below.
const maximum_mmap_pages = 256;
/// Ceiling on a process's argv entries, including argv[0]. Arguments are spawn
/// parameters ("you are the driver for device 12"), not bulk data — IPC carries
/// that — so the bound is small and everything fits the single stack page.
pub const maximum_arguments = 8;
/// Ceiling on the `system_spawn` extra-arguments blob (argv[1..], NUL-separated).
pub const maximum_argument_bytes = 256;
/// Auxiliary-vector entry types (System V AMD64 process entry). Only what the
/// kernel emits today; a C runtime scans the vector until the null terminator.
const auxiliary_vector_null: u64 = 0; // AT_NULL — end of the vector
const auxiliary_vector_page_size: u64 = 6; // AT_PAGESZ
// The hand-assembled user program blob (isr.s, .rodata) — the isolation probe.
const pf_start = @extern([*]const u8, .{ .name = "user_pf_start" });
const pf_end = @extern([*]const u8, .{ .name = "user_pf_end" });
@@ -106,9 +129,15 @@ pub fn setInitialRamdisk(image: []const u8) void {
/// written back into the trap frame, since the entry paths restore user registers
/// from it. One handler serves both the system_call/sysret and int-0x80 entry paths.
///
/// Install it once at boot (before any user code runs) via `init`.
/// Install it once at boot (before any user code runs) via `init`. Also registers
/// the scheduler's kill hooks: the scheduler sits below this layer, so finishing a
/// deferred process_kill (IRQ bindings, IPC handles, the exit notification) is
/// called back up into here from the tick (see scheduler.reapKillPendingLocked).
pub fn init() void {
architecture.setSystemCallHandler(system_call);
scheduler.terminate_current_hook = terminateCurrentLocked;
scheduler.reap_task_hook = reapTaskLocked;
scheduler.timer_tick_hook = timerSweepLocked;
}
/// Return -1 (as an unsigned bit pattern) in the system_call result register.
@@ -117,21 +146,27 @@ fn fail(state: *architecture.CpuState) void {
}
fn system_call(state: *architecture.CpuState) void {
const t = scheduler.current();
const user = t.aspace != 0;
if (user) {
// A condemned process (process_kill caught it running) dies at its next
// kernel entry — before it can spawn, claim, or message anything else.
if (t.kill_pending) terminateCurrent();
// Mark the span of this call so the timer tick never tears the task down
// in the middle of a kernel operation (scheduler.reapKillPendingLocked).
t.in_system_call = true;
}
defer if (user) {
t.in_system_call = false;
};
switch (@as(SystemCall, @enumFromInt(architecture.systemCallNumber(state)))) {
.exit => {
exit_code = architecture.systemCallArg(state, 0);
// A scheduled process drops its endpoint references, frees its address
// space, and reschedules; a borrowed test thread unwinds back to the
// kernel that entered it.
// A scheduled process tears down fully (terminateCurrent); a borrowed
// test thread unwinds back to the kernel that entered it.
if (scheduler.currentIsUserProcess()) {
// Unbind before closeHandles: dropping the last reference destroys the
// Endpoint, and a still-bound GSI would have an ISR call
// notifyFromIsr on freed memory the next time the device fired.
// unbindAll also leaves the line masked, so a dead driver's device
// goes quiet rather than storming.
releaseIrqs(scheduler.current());
ipc.closeHandles(scheduler.current());
scheduler.exitUser();
scheduler.current().exit_reason = .exited;
terminateCurrent();
} else architecture.userExit();
},
.yield => {
@@ -150,6 +185,7 @@ fn system_call(state: *architecture.CpuState) void {
.ipc_lookup => systemIpcLookup(state),
.ipc_call => systemIpcCall(state),
.ipc_reply_wait => systemIpcReplyWait(state),
.ipc_send => systemIpcSend(state),
.device_enumerate => systemDeviceEnumerate(state),
.device_claim => systemDeviceClaim(state),
.mmio_map => systemMmioMap(state),
@@ -163,6 +199,13 @@ fn system_call(state: *architecture.CpuState) void {
.io_read => systemIoRead(state),
.io_write => systemIoWrite(state),
.clock => systemClock(state),
.process_enumerate => systemProcessEnumerate(state),
.process_kill => systemProcessKill(state),
.process_exit_reason => systemProcessExitReason(state),
.process_subscribe => systemProcessSubscribe(state),
.signal_bind => systemSignalBind(state),
.process_signal => systemProcessSignal(state),
.timer_bind => systemTimerBind(state),
_ => fail(state),
}
}
@@ -228,6 +271,18 @@ fn systemIpcReplyWait(state: *architecture.CpuState) void {
architecture.setSystemCallResult3(state, received_cap);
}
/// ipc_send(handle, message_ptr, message_len) -> 0/-errno: post a payload to an
/// endpoint's async queue and wake a receiver, without blocking the caller. The async
/// counterpart of ipc_call — for broadcasts (the input service) where a rendezvous would
/// let one dead subscriber hang the sender. Delivered through ipc_reply_wait as a
/// buffered message (badge carries notify_message_bit and the caller's task id).
fn systemIpcSend(state: *architecture.CpuState) void {
const me = scheduler.current();
const endpoint = ipc.resolveHandle(me, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
const r = ipc.send(endpoint, me.aspace, architecture.systemCallArg(state, 1), architecture.systemCallArg(state, 2), me.id);
architecture.setSystemCallResult(state, @bitCast(r));
}
/// device_enumerate(buffer, maximum) -> total: snapshot the device table into the caller's
/// buffer (up to `maximum` entries), returning the total device count.
fn systemDeviceEnumerate(state: *architecture.CpuState) void {
@@ -243,6 +298,8 @@ fn systemDeviceEnumerate(state: *architecture.CpuState) void {
/// device_claim(id) -> 0/-1: take exclusive ownership of a device for this process.
fn systemDeviceClaim(state: *architecture.CpuState) void {
const claim_flags = sync.enter();
defer sync.leave(claim_flags);
if (devices_broker.claim(architecture.systemCallArg(state, 0), scheduler.current().id))
architecture.setSystemCallResult(state, 0)
else
@@ -257,10 +314,22 @@ fn systemMmioMap(state: *architecture.CpuState) void {
const resource_index = architecture.systemCallArg(state, 1);
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const owner = devices_broker.ownerOf(device_id) orelse return fail(state);
if (owner != t.id) return fail(state); // not claimed by this process
const r = devices_broker.resourceOf(device_id, resource_index) orelse return fail(state);
// Read the broker table under the lock: ring-3 device_register (M19) now
// mutates it concurrently on other cores, so a lock-free read here could
// see a torn resource (and a torn length used to panic the arithmetic
// below on integer overflow).
const r = blk: {
const flags = sync.enter();
defer sync.leave(flags);
const owner = devices_broker.ownerOf(device_id) orelse return fail(state);
if (owner != t.id) return fail(state); // not claimed by this process
break :blk devices_broker.resourceOf(device_id, resource_index) orelse return fail(state);
};
if (r.kind != @intFromEnum(device_abi.ResourceKind.memory)) return fail(state);
// A zero-length or wrapping window is not mappable — fail cleanly rather
// than underflow `r.len - 1`.
if (r.len == 0) return fail(state);
if (@addWithOverflow(r.start, r.len)[1] != 0) return fail(state);
if (t.device_map_next == 0) t.device_map_next = device_arena_base;
const first = r.start & ~@as(u64, page_size - 1);
@@ -396,44 +465,406 @@ fn systemDeviceRegister(state: *architecture.CpuState) void {
var descriptor: device_abi.DeviceDescriptor = undefined;
if (!ipc.copyFromUser(t.aspace, descriptor_ptr, std.mem.asBytes(&descriptor))) return fail(state);
// Under the big kernel lock: the broker's table is also mutated by the
// death sweep (releaseAllOwnedBy) and read by enumerate on other cores —
// ring-3 registration (M19) made those genuinely concurrent.
const flags = sync.enter();
defer sync.leave(flags);
const id = devices_broker.register(parent_id, t.id, &descriptor) catch return fail(state);
architecture.setSystemCallResult(state, id);
}
/// system_spawn(name_ptr, name_len) -> 0 on success, -1 on failure. Load the binary
/// bundled in the initial-ramdisk under `name` as a fresh ring-3 process. This is the
/// mechanism a user-space supervisor (the device manager) uses to start a driver it
/// matched: discovery and policy stay in user space, the kernel only spawns.
/// system_spawn(name_ptr, name_len, arguments_ptr, arguments_len, exit_endpoint)
/// -> the child's process id on success, -1 on failure. Load the binary bundled in
/// the initial-ramdisk under `name` as a fresh ring-3 process. `name` becomes the
/// child's argv[0] (and its task name, so a fault report can say which binary
/// died); `arguments` is an optional NUL-separated blob that becomes argv[1..] —
/// how a supervisor parameterises what it starts ("you are the driver for device
/// 12"). 0/0 means no extra arguments. This is the mechanism a user-space
/// supervisor (the device manager) uses to start a driver it matched: discovery
/// and policy stay in user space, the kernel only spawns.
///
/// Ungated for now — any process may spawn any bundled binary. A capability (only a
/// supervisor holds the right to spawn) belongs here once the model grows one; see
/// docs/driver-model.md. The name is bounds-checked into the user half exactly like
/// `debug_write`, and an unknown name or a load failure returns -1.
/// The caller is recorded as the child's **supervisor** — the sole holder of the
/// right to `process_kill` it (docs/process-management.md). `exit_endpoint` (a
/// handle, or `abi.no_cap` for none) names an endpoint of the caller's to notify
/// when the child ends, any way it ends — the IRQ-as-IPC pattern reused as the
/// microkernel's SIGCHLD.
///
/// Spawning itself is still ungated — any process may spawn any bundled binary; a
/// spawn capability belongs here once the model grows one (docs/driver-model.md).
/// Both buffers are bounds-checked into the user half exactly like `debug_write`,
/// and an unknown name or a load failure returns -1.
fn systemSpawn(state: *architecture.CpuState) void {
const ptr = architecture.systemCallArg(state, 0);
const len = architecture.systemCallArg(state, 1);
if (len == 0 or len > 64 or ptr >= user_half_end or ptr + len > user_half_end) return fail(state);
const arguments_ptr = architecture.systemCallArg(state, 2);
const arguments_len = architecture.systemCallArg(state, 3);
const exit_handle = architecture.systemCallArg(state, 4);
const t = scheduler.current();
if (len == 0 or len > scheduler.maximum_task_name or ptr >= user_half_end or ptr + len > user_half_end) return fail(state);
if (arguments_len > maximum_argument_bytes) return fail(state);
if (arguments_len != 0 and (arguments_ptr >= user_half_end or arguments_ptr + arguments_len > user_half_end)) return fail(state);
const exit_endpoint: ?*ipc.Endpoint = if (exit_handle == abi.no_cap)
null
else
ipc.resolveHandle(t, exit_handle) orelse return failErr(state, ipc.EBADF);
const image = ramdisk_image orelse return fail(state);
const rd = initial_ramdisk.Reader.init(image) orelse return fail(state);
const name = @as([*]const u8, @ptrFromInt(ptr))[0..len];
var argv: [maximum_arguments][]const u8 = undefined;
argv[0] = name;
var argc: usize = 1;
if (arguments_len != 0) {
const blob = @as([*]const u8, @ptrFromInt(arguments_ptr))[0..arguments_len];
var pieces = std.mem.tokenizeScalar(u8, blob, 0);
while (pieces.next()) |piece| {
if (argc == maximum_arguments) return fail(state);
argv[argc] = piece;
argc += 1;
}
}
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!std.mem.eql(u8, item.name, name)) continue;
spawnProcess(item.blob, 4) catch return fail(state);
architecture.setSystemCallResult(state, 0);
const child = spawnProcessSupervised(item.blob, 4, argv[0..argc], t.id, exit_endpoint) catch return fail(state);
architecture.setSystemCallResult(state, child);
return;
}
fail(state); // no bundled binary by that name
}
/// Drop every IRQ binding `t` made. Called on exit, before the handle table is closed
/// (which is what frees the endpoints an ISR would otherwise notify into).
fn releaseIrqs(t: *scheduler.Task) void {
/// process_enumerate(buffer, maximum) -> total: snapshot the task table into the
/// caller's buffer (up to `maximum` `abi.ProcessDescriptor` entries), returning
/// the total live-task count — the exact shape of `device_enumerate`, so a `ps`
/// is a user program over a snapshot, not a kernel service. Read-only and
/// ungated: what is running is not a secret between cooperating bring-up
/// processes.
fn systemProcessEnumerate(state: *architecture.CpuState) void {
const buffer_ptr = architecture.systemCallArg(state, 0);
const maximum = architecture.systemCallArg(state, 1);
const t = scheduler.current();
if (t.aspace == 0 or buffer_ptr >= user_half_end) return fail(state);
const sz = @sizeOf(abi.ProcessDescriptor);
const cap = @min(maximum, (user_half_end - buffer_ptr) / sz); // clamp to the user half
const out: [*]abi.ProcessDescriptor = @ptrFromInt(buffer_ptr);
architecture.setSystemCallResult(state, scheduler.enumerate(out[0..@intCast(cap)]));
}
/// process_kill(id) -> 0 / -ESRCH / -EPERM: end the process `id`. Only its
/// supervisor — the process that spawned it — may do so; the supervision link is
/// the kill capability, so no user/permission model is needed and a stray id
/// cannot be a weapon (ids are never reused, so a stale one just misses).
fn systemProcessKill(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const id = architecture.systemCallArg(state, 0);
if (id > std.math.maxInt(u32)) return failErr(state, ipc.ESRCH);
const r = killProcess(t.id, @intCast(id));
architecture.setSystemCallResult(state, @bitCast(r));
}
/// Processes killed by a CPU fault rather than a clean exit. Evidence for the
/// fault-recovery test, and a health signal a supervisor can consult later.
pub var fault_kill_count: u64 = 0;
/// Release everything a dying task holds and tell its supervisor — the shared
/// half of every path out of a process: clean exit, fault kill, and process_kill
/// (both the immediate reap and the deferred tick-time terminate). The order
/// matters:
/// - IRQ bindings are dropped before the handle table closes: dropping the last
/// endpoint reference destroys the Endpoint, and a still-bound GSI would have an
/// ISR call notifyFromIsr on freed memory the next time the device fired.
/// `releaseOwner` also leaves the line masked, so a dead driver's device goes
/// quiet rather than storming. (It drops MSI vectors by the same owner sweep.)
/// - Device claims are released with the IRQ bindings, so a restarted driver can
/// claim the same hardware again — the cleanup half of process-lifecycle.md's
/// iron rule 1. Claims hold no pointers, so ordering is free; they go here so
/// the exit notification (below, last) observes a fully-released child.
/// - A client this task still owes a reply to (it died between receive and reply)
/// is failed with -EPEER rather than left blocked forever — a dead server must
/// not hang its callers.
/// - The task is unlinked from wherever IPC parked it (an endpoint's sender FIFO,
/// a receive wait queue, or a server's owed-reply slot) *before* the handles
/// close, so nothing ever dequeues a dangling pointer. These are no-ops for a
/// running task ending itself; they matter when process_kill reaps a blocked one.
/// - The exit notification is posted last, once the process can no longer act, so
/// a supervisor that receives it observes a fully-released child. The endpoint
/// reference taken at spawn is dropped with it.
/// Precondition: the big kernel lock is held.
fn releaseTaskResourcesLocked(t: *scheduler.Task) void {
recordExitLocked(t);
irq.releaseOwner(t.id);
devices_broker.releaseAllOwnedBy(t.id);
// The dying task's signal endpoint and one-shot timers go with it.
if (t.signal_endpoint) |raw| {
ipc.dropRef(@ptrCast(@alignCast(raw)));
t.signal_endpoint = null;
}
t.pending_signals = 0;
for (&one_shot_timers) |*slot| {
if (slot.*) |timer| {
if (timer.owner == t.id) {
ipc.dropRef(timer.endpoint);
slot.* = null;
}
}
}
// A dead subscriber's own subscriptions go first: it must not hear about
// itself, and the slots' endpoint references drop with it.
for (&exit_subscribers) |*slot| {
if (slot.*) |subscriber| {
if (subscriber.owner == t.id) {
ipc.dropRef(subscriber.endpoint);
slot.* = null;
}
}
}
if (t.ipc_client) |client| {
t.ipc_client = null;
client.ipc_status = -ipc.EPEER;
scheduler.readyLocked(client); // its blocked `call` now returns the error
}
ipc.abandonSenderLocked(t);
scheduler.removeFromWaitQueueLocked(t);
scheduler.forgetIpcClientLocked(t);
ipc.closeHandles(t);
// Publish the exit to every subscriber (docs/process-lifecycle.md): the same
// badge encoding as the supervisor's notification, and equally late, so a
// subscriber also observes a fully-released child.
for (&exit_subscribers) |*slot| {
if (slot.*) |subscriber| ipc.notifyLocked(subscriber.endpoint, abi.notify_exit_bit | t.id);
}
if (t.exit_endpoint) |raw| {
const endpoint: *ipc.Endpoint = @ptrCast(@alignCast(raw));
t.exit_endpoint = null;
ipc.notifyLocked(endpoint, abi.notify_exit_bit | t.id);
ipc.dropRef(endpoint);
}
}
/// Tear down the current user process and reschedule; never returns. Shared by the
/// exit system call and the fault path (`killCurrentProcess`). See
/// `releaseTaskResourcesLocked` for what is released, and in what order.
pub fn terminateCurrent() noreturn {
_ = sync.enter(); // handed off through the exit switch, released by the resumed task
terminateCurrentLocked();
}
/// The body of `terminateCurrent` for a caller that already holds the big kernel
/// lock — the scheduler's tick calls this (via `terminate_current_hook`) to finish
/// a deferred process_kill on its own core's current task. Never returns; the
/// tick's abandoned interrupt frame is fine (the LAPIC was acknowledged before the
/// tick hook ran), exactly as on the fault path.
fn terminateCurrentLocked() noreturn {
releaseTaskResourcesLocked(scheduler.current());
scheduler.exitUserLocked();
}
/// Reap a condemned task that is NOT running on any core (ready or blocked — and
/// it cannot start running: state changes need the lock we hold). The other half
/// of a deferred process_kill, called by the scheduler's tick (via
/// `reap_task_hook`) and directly by `killProcess` for targets caught off-CPU.
/// Precondition: the big kernel lock is held.
fn reapTaskLocked(t: *scheduler.Task) void {
releaseTaskResourcesLocked(t);
scheduler.removeFromReadyQueueLocked(t); // no-op unless it was ready in a queue
scheduler.destroyTaskLocked(t);
}
/// Kill process `target_id` on behalf of `caller_id` — the kernel half of the
/// process_kill system call. Returns 0, -ESRCH (no such live process — kernel
/// tasks are not killable processes and stale ids miss, since ids are never
/// reused), or -EPERM (the caller is not the target's supervisor).
///
/// A target that is ready or blocked is reaped on the spot. One that is running
/// on another core cannot be torn down mid-instruction, so it is condemned
/// (`kill_pending`) and dies at its next system_call entry, block, or timer tick
/// — like a Unix signal, delivery is prompt but asynchronous. Either way the
/// call returns 0: the kill is accepted and irrevocable.
pub fn killProcess(caller_id: u32, target_id: u32) i64 {
const flags = sync.enter();
defer sync.leave(flags);
irq.releaseOwner(t.id);
const target = scheduler.taskByIdLocked(target_id) orelse return -ipc.ESRCH;
if (target.aspace == 0) return -ipc.ESRCH; // kernel tasks are not processes
if (target.supervisor != caller_id) return -ipc.EPERM;
target.exit_reason = .killed;
if (target.state == .running) {
target.kill_pending = true;
} else {
reapTaskLocked(target);
}
return 0;
}
/// Kill the current user process in response to a CPU fault it raised in ring 3.
/// The fault is confined to the process — the kernel trapped it on the task's own
/// kernel stack and is intact — so everything the process held is reclaimed and the
/// core reschedules. The system keeps running; only the faulting process dies
/// (docs/resilience.md: fault -> kill -> continue). `reason` is the fault class
/// (from the vector), recorded for the supervisor's `process_exit_reason`.
pub fn killCurrentProcess(reason: abi.ExitReason) noreturn {
scheduler.current().exit_reason = reason;
fault_kill_count += 1;
terminateCurrent();
}
/// The bounded record of recent deaths, for `process_exit_reason`: ids are never
/// reused, so a ring keyed by id is enough — a record evicted by wraparound reads
/// as -ESRCH, the same as an id that never lived, which a supervisor treats as
/// "too late to ask". Written under the big kernel lock by the reap.
const exit_record_capacity = 64;
const ExitRecord = struct { id: u32 = 0, supervisor: u32 = 0, reason: abi.ExitReason = .exited, valid: bool = false };
var exit_records: [exit_record_capacity]ExitRecord = .{ExitRecord{}} ** exit_record_capacity;
var exit_record_next: usize = 0;
/// Record a dying task's (id, supervisor, reason) — called by the reap before the
/// exit notification is posted, so a supervisor that hears the notification can
/// always still query the reason. Precondition: the big kernel lock is held.
fn recordExitLocked(t: *scheduler.Task) void {
exit_records[exit_record_next] = .{ .id = t.id, .supervisor = t.supervisor, .reason = t.exit_reason, .valid = true };
exit_record_next = (exit_record_next + 1) % exit_record_capacity;
}
/// How dead process `id` ended, for `caller` — the kernel half of the
/// process_exit_reason system call. Returns the ExitReason value, -ESRCH (never
/// lived, still alive, or evicted from the ring), or -EPERM (the caller was not
/// its supervisor — the same authority gate as process_kill).
pub fn exitReasonOf(caller_id: u32, target_id: u32) i64 {
const flags = sync.enter();
defer sync.leave(flags);
for (&exit_records) |*record| {
if (record.valid and record.id == target_id) {
if (record.supervisor != caller_id) return -ipc.EPERM;
return @intFromEnum(record.reason);
}
}
return -ipc.ESRCH;
}
/// The published exit events' subscribers (docs/process-lifecycle.md "Who learns
/// of a death"): stateful services — the VFS's file handles, input's
/// subscriptions — that must release what a dead client held and cannot learn it
/// any other way (a client that simply never calls again looks like silence).
/// Bounded like every kernel table; each entry holds its own endpoint reference.
const exit_subscriber_capacity = 8;
const ExitSubscriber = struct { endpoint: *ipc.Endpoint, owner: u32 };
var exit_subscribers: [exit_subscriber_capacity]?ExitSubscriber = .{null} ** exit_subscriber_capacity;
/// process_subscribe(endpoint): subscribe the caller's endpoint to published exit
/// events. Ungated, like process_enumerate — what is running (and dying) is not a
/// secret between cooperating processes. -ENOSPC when the table is full.
fn systemProcessSubscribe(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
const flags = sync.enter();
defer sync.leave(flags);
for (&exit_subscribers) |*slot| {
if (slot.* == null) {
endpoint.refcount += 1; // the slot's own reference, dropped on unsubscribe-by-death
slot.* = .{ .endpoint = endpoint, .owner = t.id };
return architecture.setSystemCallResult(state, 0);
}
}
failErr(state, ipc.ENOSPC);
}
/// signal_bind(endpoint): nominate where this process's signals arrive — the
/// IRQ-as-IPC pattern a fourth time (docs/process-lifecycle.md). Replacing a
/// binding drops the old reference; signals that pended while unbound are
/// delivered immediately on bind, coalesced into one notification.
fn systemSignalBind(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
const flags = sync.enter();
defer sync.leave(flags);
if (t.signal_endpoint) |raw| ipc.dropRef(@ptrCast(@alignCast(raw)));
endpoint.refcount += 1;
t.signal_endpoint = @ptrCast(endpoint);
if (t.pending_signals != 0) {
ipc.notifyLocked(endpoint, abi.notify_signal_bit | t.pending_signals);
t.pending_signals = 0;
}
architecture.setSystemCallResult(state, 0);
}
/// process_signal(id, signal): post a signal — a one-way, coalescing statement,
/// never a question (docs/process-lifecycle.md). The authority gate is the
/// supervision link, like kill; a process may also signal itself. Unbound
/// targets accumulate the signal in their pending mask.
fn systemProcessSignal(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const id = architecture.systemCallArg(state, 0);
const signal = architecture.systemCallArg(state, 1);
if (id > std.math.maxInt(u32)) return failErr(state, ipc.ESRCH);
if (signal > 31) return failErr(state, ipc.EBADF); // not a Signal bit position
const flags = sync.enter();
defer sync.leave(flags);
const target = scheduler.taskByIdLocked(@intCast(id)) orelse return failErr(state, ipc.ESRCH);
if (target.aspace == 0) return failErr(state, ipc.ESRCH);
if (target.supervisor != t.id and target.id != t.id) return failErr(state, ipc.EPERM);
target.pending_signals |= @as(u32, 1) << @intCast(signal);
if (target.signal_endpoint) |raw| {
const endpoint: *ipc.Endpoint = @ptrCast(@alignCast(raw));
ipc.notifyLocked(endpoint, abi.notify_signal_bit | target.pending_signals);
target.pending_signals = 0;
}
architecture.setSystemCallResult(state, 0);
}
/// The one-shot timers of timer_bind: the missing timed wait. A service arms a
/// deadline and keeps serving; the expiry arrives in the same replyWait as
/// everything else (notify_timer_bit). What stop-sequence escalation, hello
/// deadlines, and restart backoff are built from — and later, `alarm`.
const timer_capacity = 16;
const OneShotTimer = struct { deadline: u64, endpoint: *ipc.Endpoint, owner: u32 };
var one_shot_timers: [timer_capacity]?OneShotTimer = .{null} ** timer_capacity;
/// Sweep expired timers — hung on scheduler.timer_tick_hook, so it runs on every
/// tick with the big kernel lock held, like the sleeper wake it rides beside.
fn timerSweepLocked() void {
const now = architecture.millis();
for (&one_shot_timers) |*slot| {
if (slot.*) |timer| {
if (now >= timer.deadline) {
ipc.notifyLocked(timer.endpoint, abi.notify_timer_bit);
ipc.dropRef(timer.endpoint);
slot.* = null;
}
}
}
}
/// timer_bind(endpoint, ms): arm a one-shot timer. -ENOSPC when the table is full.
fn systemTimerBind(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
const ms = architecture.systemCallArg(state, 1);
const flags = sync.enter();
defer sync.leave(flags);
for (&one_shot_timers) |*slot| {
if (slot.* == null) {
endpoint.refcount += 1;
slot.* = .{ .deadline = architecture.millis() + ms, .endpoint = endpoint, .owner = t.id };
return architecture.setSystemCallResult(state, 0);
}
}
failErr(state, ipc.ENOSPC);
}
fn systemProcessExitReason(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const id = architecture.systemCallArg(state, 0);
if (id > std.math.maxInt(u32)) return failErr(state, ipc.ESRCH);
const r = exitReasonOf(t.id, @intCast(id));
architecture.setSystemCallResult(state, @bitCast(r));
}
/// Resolve `(device_id, resource_index)` to a GSI this process is entitled to bind, or null.
@@ -503,6 +934,11 @@ fn systemIrqAck(state: *architecture.CpuState) void {
if (irq.ack(gsi)) architecture.setSystemCallResult(state, 0) else fail(state);
}
/// Whether the debug_write stream sits at the start of a line — the last emitted
/// byte was a newline (true at boot: nothing emitted yet). Guarded by the kernel
/// lock in `systemDebugWrite`, like the stream it describes.
var write_at_line_start: bool = true;
/// debug_write(ptr, len): copy bytes from user memory into the kernel log.
/// A bring-up diagnostic — real output goes through the VFS/console later.
///
@@ -512,17 +948,26 @@ fn systemIrqAck(state: *architecture.CpuState) void {
/// Known gap (fine for trusted user code): a pointer into an *unmapped* hole in
/// the user half passes the check and the read #PFs -> on_fault halts — a
/// self-DoS, not an isolation break. Fault-recovering copy-in is a later item.
///
/// The emit runs under the kernel lock, so a message is atomic on the wire — two
/// processes writing from different cores can interleave *messages*, never bytes.
/// The "DANOS-INIT: " marker is emitted only at the start of a line (not per
/// call), so a process may assemble a line from several writes without the marker
/// (or, with the lock, another byte) landing in the middle. Cleanly-terminated
/// lines from concurrent writers stay whole either way.
fn systemDebugWrite(state: *architecture.CpuState) void {
const ptr = architecture.systemCallArg(state, 0);
const len = architecture.systemCallArg(state, 1);
if (len <= write_buffer.len and ptr < user_half_end and ptr + len <= user_half_end) {
const source: [*]const u8 = @ptrFromInt(ptr);
const flags = sync.enter();
defer sync.leave(flags);
@memcpy(write_buffer[0..len], source[0..len]); // keep the latest message
write_len = len;
write_from_user = architecture.fromUser(state);
write_count += 1;
log.write("DANOS-INIT: ");
log.write(source[0..len]);
if (len != 0) write_at_line_start = source[len - 1] == '\n';
architecture.setSystemCallResult(state, len);
} else {
fail(state);
@@ -617,15 +1062,15 @@ pub fn run(blob: []const u8) RunError!void {
@memset(code[blob.len..page_size], 0xCC);
architecture.mapUserPage(code_virtual, code_frame, false, true); // RO + X
architecture.mapUserPage(stack_virtual, stack_frame, true, false); // RW + NX
architecture.mapUserPage(stack_base_virtual, stack_frame, true, false); // RW + NX (one page; the probe barely stacks)
resetRecords();
architecture.enterUser(scheduler.currentCpuIndex(), code_virtual, stack_virtual + page_size);
architecture.enterUser(scheduler.currentCpuIndex(), code_virtual, stack_base_virtual + page_size);
// Back via the exit system_call; the interrupt gate left IF clear.
architecture.enableInterrupts();
architecture.unmapPage(code_virtual);
architecture.unmapPage(stack_virtual);
architecture.unmapPage(stack_base_virtual);
pmm.free(code_frame);
pmm.free(stack_frame);
}
@@ -636,6 +1081,7 @@ pub const InitError = error{
BadElf, // malformed/inapplicable image (magic, class, machine, type, bounds)
BadSegment, // PT_LOAD unaligned, out of the user region, W&X, or overlapping
BadEntry, // e_entry not inside an executable segment
BadArguments, // no argv[0], too many entries, or too many bytes for the entry stack
ProgramTooBig, // more pages than the loader's budget
OutOfMemory,
};
@@ -736,13 +1182,84 @@ fn loadPageInto(aspace: u64, image: []const u8, seg: Segment, page_index: u64) I
architecture.mapUserPageInto(aspace, seg.vaddr + page_off, frame, seg.writable, seg.executable);
}
/// Build the System V AMD64 process-entry block at the top of a process's stack
/// page and return the initial user stack pointer. At the first user instruction,
/// rsp is 16-byte aligned and points at (addresses growing upward):
///
/// argc, argv[0..argc-1], NULL, NULL (empty envp), auxiliary vector, strings
///
/// — the layout every C runtime's startup code walks, so danos's own runtime and a
/// future libc port read arguments identically (docs/sysv.md). `page` is the kernel
/// (physmap) view of the stack's **top** frame and `page_user_base` that frame's
/// user address (stack_top_virtual - page_size); the pointers written into it are
/// user addresses inside that page. The caller has validated the sizes
/// (`entryStackBytes`), so this cannot overrun.
fn buildEntryStack(page: [*]u8, page_user_base: u64, argv: []const []const u8) u64 {
// The strings live at the very top of the page, packed from the end downward.
var string_offset: usize = page_size;
var pointers: [maximum_arguments]u64 = undefined;
var i: usize = argv.len;
while (i > 0) {
i -= 1;
string_offset -= argv[i].len + 1;
@memcpy(page[string_offset..][0..argv[i].len], argv[i]);
page[string_offset + argv[i].len] = 0; // NUL-terminated, as C expects
pointers[i] = page_user_base + string_offset;
}
// The vector sits below the strings: argc, the argv pointers, the argv
// terminator, an empty envp (terminator only), then the auxiliary vector.
const word_count = 1 + argv.len + 1 + 1 + 4;
const vector_offset = (string_offset - word_count * 8) & ~@as(usize, 15); // entry rsp % 16 == 0
const words: [*]u64 = @ptrCast(@alignCast(page + vector_offset));
var w: usize = 0;
words[w] = argv.len; // argc
w += 1;
for (pointers[0..argv.len]) |pointer| {
words[w] = pointer;
w += 1;
}
words[w] = 0; // argv terminator
words[w + 1] = 0; // envp: no environment yet, just the terminator
words[w + 2] = auxiliary_vector_page_size;
words[w + 3] = page_size;
words[w + 4] = auxiliary_vector_null; // end of the auxiliary vector
words[w + 5] = 0;
return page_user_base + vector_offset;
}
/// Bytes the entry block for `argv` occupies at the top of the stack page:
/// strings (each NUL-terminated), vector words, and the alignment slack.
fn entryStackBytes(argv: []const []const u8) usize {
var string_bytes: usize = 0;
for (argv) |argument| string_bytes += argument.len + 1;
return string_bytes + (1 + argv.len + 1 + 1 + 4) * 8 + 16;
}
/// Load a user ELF image into a fresh address space and spawn it as a scheduled
/// ring-3 process at `priority`. Returns immediately — the process runs
/// preemptively on its own page tables alongside everything else, and its exit
/// is handled by the system_call layer. The whole build (address space + ELF load +
/// task) runs under the kernel lock so it appears atomically and can't race
/// pmm/heap on another core.
pub fn spawnProcess(image: []const u8, priority: u3) InitError!void {
/// ring-3 process at `priority`, entered with `argv` on its stack per the System V
/// convention (`buildEntryStack`). `argv[0]` is required — it names the process:
/// the path or initial-ramdisk name it was spawned as. It is also recorded on the
/// task, so a fault report can say *which* binary died, not just its id.
/// The kernel-internal spawn (init at boot, tests): supervisor 0, no exit
/// notification. `spawnProcessSupervised` is the full form.
pub fn spawnProcess(image: []const u8, priority: u3, argv: []const []const u8) InitError!void {
_ = try spawnProcessSupervised(image, priority, argv, 0, null);
}
/// `spawnProcess`, recording `supervisor` (the id of the process that asked — the
/// kill authority) and, if given, `exit_endpoint` to notify when the child ends
/// (a reference is taken here and dropped when the notification posts).
/// Returns the child's process id.
/// Returns immediately — the process runs preemptively on its own page tables
/// alongside everything else, and its exit is handled by the system_call layer.
/// The whole build (address space + ELF load + task) runs under the kernel lock so
/// it appears atomically and can't race pmm/heap on another core.
pub fn spawnProcessSupervised(image: []const u8, priority: u3, argv: []const []const u8, supervisor: u32, exit_endpoint: ?*ipc.Endpoint) InitError!u32 {
if (argv.len == 0 or argv.len > maximum_arguments) return error.BadArguments;
// The entry block must leave most of the page as actual stack.
if (entryStackBytes(argv) > page_size / 2) return error.BadArguments;
var segs: [maximum_segments]Segment = undefined;
const parsed = try parseSegments(image, &segs);
@@ -755,11 +1272,28 @@ pub fn spawnProcess(image: []const u8, priority: u3) InitError!void {
for (segs[0..parsed.count]) |seg| {
for (0..seg.pages()) |i| try loadPageInto(aspace, image, seg, i);
}
const stack_frame = pmm.alloc() orelse return error.OutOfMemory;
architecture.mapUserPageInto(aspace, stack_virtual, stack_frame, true, false); // RW + NX
if (!scheduler.spawnUserLocked(aspace, parsed.entry, stack_virtual + page_size, priority))
// The stack: `user_stack_pages` zeroed pages below stack_top_virtual, RW + NX.
// The page below them (`stack_virtual`) stays unmapped as the overflow guard.
// The entry block goes at the top of the highest page.
var user_sp: u64 = 0;
for (0..parameters.user_stack_pages) |i| {
const stack_frame = pmm.alloc() orelse return error.OutOfMemory;
const stack_page: [*]u8 = @ptrFromInt(boot_handoff.physicalToVirtual(stack_frame));
@memset(stack_page[0..page_size], 0); // no stale frame contents leak into user space
const page_virtual = stack_base_virtual + i * page_size;
if (i == parameters.user_stack_pages - 1)
user_sp = buildEntryStack(stack_page, page_virtual, argv);
architecture.mapUserPageInto(aspace, page_virtual, stack_frame, true, false); // RW + NX
}
const child = scheduler.spawnUserLocked(aspace, parsed.entry, user_sp, priority, argv[0], supervisor, if (exit_endpoint) |endpoint| @ptrCast(endpoint) else null) orelse
return error.OutOfMemory;
// The child holds a reference to its exit endpoint from birth to death. Taken
// only now, after nothing can fail; the lock is still held, so the child
// cannot run (let alone die) before the reference exists.
if (exit_endpoint) |endpoint| endpoint.refcount += 1;
return child;
}
/// clock() -> nanoseconds since boot: a monotonic time source. The kernel already owns
+246 -13
View File
@@ -18,6 +18,7 @@
//! shared queues.
const std = @import("std");
const abi = @import("abi");
const parameters = @import("parameters");
const architecture = @import("architecture");
const heap = @import("heap.zig");
@@ -41,6 +42,37 @@ pub const Task = struct {
kstack_top: usize = 0, // top of `stack` (== TSS.rsp0 for a user task); 0 = none
wake_at: u64 = 0, // uptime (ms) to wake a sleeping task; 0 = not sleeping
affinity: ?u32 = null, // null = runs on any core; else the index of its pinned core
// --- process management (process.zig) ---
// Id of the process that spawned this one (0 = the kernel). The supervision
// link is the kill authority: only the supervisor may process_kill a child.
supervisor: u32 = 0,
// Endpoint to notify when this process ends (any way: exit, fault, kill), or
// null. Holds its own reference, dropped when the notification is posted.
// Opaque here for the same reason as `handles` below.
exit_endpoint: ?*anyopaque = null,
// How this process ended — set by the death paths (exit, fault, kill) just
// before the reap records it for `process_exit_reason`. Meaningless while
// the task lives.
exit_reason: abi.ExitReason = .exited,
// Endpoint this process's signals arrive on (signal_bind), or null — same
// ownership rules as exit_endpoint (holds a reference; opaque here).
signal_endpoint: ?*anyopaque = null,
// Signals posted but not yet delivered: the coalescing pending mask
// (docs/process-lifecycle.md). Bits are abi.Signal values. Signals pend here
// until an endpoint is bound; two pending terminates are one terminate.
pending_signals: u32 = 0,
// Set by process_kill on a task that is running on another core; the kernel
// finishes the kill at that task's next system call or timer tick.
kill_pending: bool = false,
// True while this task executes its own system call — the timer tick must not
// tear a task down in the middle of a kernel operation, only while it runs
// user code (or sits at a block point, where teardown is safe).
in_system_call: bool = false,
// Where this task is parked while blocked, so a kill can unlink it: the
// WaitQueue it waits on (maintained by waitLocked/wakeLocked), or the endpoint
// whose sender FIFO it queues in (maintained by the IPC layer; opaque here).
wait_queue: ?*WaitQueue = null,
ipc_wait_endpoint: ?*anyopaque = null,
// Physical root of this task's address space, or 0 for a kernel task (which
// runs on the shared kernel page tables). A user task carries its own.
aspace: u64 = 0,
@@ -70,8 +102,25 @@ pub const Task = struct {
ipc_send_cap: u64 = ~@as(u64, 0), // handle to transfer with this message (abi.no_cap = none)
ipc_received_cap: u64 = ~@as(u64, 0), // client: handle the reply's transferred cap landed at (abi.no_cap = none)
next: ?*Task = null, // ready-queue link (also the endpoint sender-FIFO link)
// The process's name — argv[0] as it was spawned (a boot-volume path for init,
// an initial-ramdisk name for everything else); empty for kernel tasks. Fixed
// storage, so the fault path can name the dead without touching the heap.
// Zero-initialised (not `undefined`): an undefined default is materialised as
// a 0xAA fill, which would move the whole static task pool out of .bss.
name_buffer: [maximum_task_name]u8 = .{0} ** maximum_task_name,
name_length: u8 = 0,
/// The task's name (argv[0] at spawn), or empty for a kernel task.
pub fn name(self: *const Task) []const u8 {
return self.name_buffer[0..self.name_length];
}
};
/// Capacity of `Task.name_buffer` — matches the longest name `system_spawn`
/// accepts, so a spawned name is never truncated. Shared with the ABI's
/// ProcessDescriptor, so `enumerate` copies names without clipping.
pub const maximum_task_name = abi.maximum_process_name;
/// Size of each task's IPC handle table. Kept here (not in ipc_sync.zig) because
/// it dimensions a field of `Task`; ipc_sync.zig re-exports it.
pub const ipc_maximum_handles = 16;
@@ -258,14 +307,19 @@ pub fn spawnOn(entry: *const fn () void, priority: Priority, cpu: u32) bool {
}
/// Spawn a **user** task: a task with its own address space (`aspace`) that starts
/// in user mode at `entry` on `user_sp`. It gets a fresh kernel stack for
/// syscalls/interrupts, and its first switch-in lands in `user_task_trampoline`.
/// Returns false (creating nothing) if the table is full or out of memory.
/// in user mode at `entry` on `user_sp`, recorded under `name` (its argv[0]).
/// `supervisor` is the id of the spawning process (0 = the kernel) — the kill
/// authority — and `exit_endpoint` (an *ipc.Endpoint whose reference the caller
/// has already taken, or null) is notified when this process ends.
/// It gets a fresh kernel stack for syscalls/interrupts, and its first switch-in
/// lands in `user_task_trampoline`.
/// Returns the new process id, or null (creating nothing) if the table is full or
/// out of memory.
/// **Caller must hold the kernel lock** (the loader that builds `aspace` holds it
/// across the whole spawn, so the address space and the task appear atomically).
pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, priority: Priority) bool {
const t = freeSlot() orelse return false;
const stack = heap.allocator().alloc(u8, stack_size) catch return false;
pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, priority: Priority, task_name: []const u8, supervisor: u32, exit_endpoint: ?*anyopaque) ?u32 {
const t = freeSlot() orelse return null;
const stack = heap.allocator().alloc(u8, stack_size) catch return null;
t.* = .{
.id = next_id,
.state = .ready,
@@ -274,7 +328,12 @@ pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, priority: Priority
.aspace = aspace,
.user_ip = entry,
.user_sp = user_sp,
.supervisor = supervisor,
.exit_endpoint = exit_endpoint,
};
const name_length = @min(task_name.len, maximum_task_name);
@memcpy(t.name_buffer[0..name_length], task_name[0..name_length]);
t.name_length = @intCast(name_length);
next_id += 1;
const top = @intFromPtr(stack.ptr) + stack.len;
t.kstack_top = top;
@@ -282,7 +341,7 @@ pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, priority: Priority
// the user entry/stack from the Task itself).
t.sp = architecture.initTaskStack(top, @intFromPtr(&startUserTask));
enqueue(t);
return true;
return t.id;
}
/// The first thing a fresh user task runs (in ring 0, via task_trampoline). It
@@ -291,8 +350,9 @@ pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, priority: Priority
/// context switch and lock release.
fn startUserTask() void {
const t = current();
var buffer: [96]u8 = undefined;
architecture.serialWrite(std.fmt.bufPrint(&buffer, "DBG startUserTask ip=0x{x} sp=0x{x} aspace=0x{x} kstack=0x{x}\n", .{ t.user_ip, t.user_sp, t.aspace, t.kstack_top }) catch "");
// No serial chatter here: this runs on every spawn, unserialized against
// user-space writes, and its output used to shear concurrent log lines in
// half — the largest source of corrupted markers in the QEMU scenarios.
architecture.jumpToUser(t.user_ip, t.user_sp); // noreturn
}
@@ -394,6 +454,7 @@ pub const WaitQueue = struct {
pub fn waitLocked(wait_queue: *WaitQueue) void {
const t = current();
t.state = .blocked;
t.wait_queue = wait_queue; // so a kill can unlink a parked waiter
t.next = wait_queue.head;
wait_queue.head = t;
schedule();
@@ -418,10 +479,78 @@ pub fn wakeLocked(wait_queue: *WaitQueue) void {
}
const t = best orelse return;
if (best_previous) |p| p.next = t.next else wait_queue.head = t.next;
t.wait_queue = null;
t.state = .ready;
enqueue(t);
}
/// Unlink `t` from the wait queue it is parked on, if any (the kill path — a
/// killed waiter must not be woken later as a dangling pointer). Precondition:
/// the big kernel lock is held.
pub fn removeFromWaitQueueLocked(t: *Task) void {
const wait_queue = t.wait_queue orelse return;
t.wait_queue = null;
var previous: ?*Task = null;
var node = wait_queue.head;
while (node) |n| : ({
previous = n;
node = n.next;
}) {
if (n != t) continue;
if (previous) |p| p.next = t.next else wait_queue.head = t.next;
t.next = null;
return;
}
}
/// Unlink `t` from the ready queue it sits in (global, or its affinity core's
/// pinned queue) — the kill path for a task that is runnable but not running.
/// Precondition: the big kernel lock is held.
pub fn removeFromReadyQueueLocked(t: *Task) void {
if (t.affinity) |cpu| {
const pc = &cpus[cpu];
removeFrom(&pc.pinned_head, &pc.pinned_tail, &pc.pinned_bitmap, t);
} else {
removeFrom(&ready_head, &ready_tail, &ready_bitmap, t);
}
}
fn removeFrom(head: *[number_priorities]?*Task, tail: *[number_priorities]?*Task, bitmap: *u8, t: *Task) void {
const level: usize = t.priority;
var previous: ?*Task = null;
var node = head[level];
while (node) |n| : ({
previous = n;
node = n.next;
}) {
if (n != t) continue;
if (previous) |p| p.next = t.next else head[level] = t.next;
if (tail[level] == t) tail[level] = previous;
if (head[level] == null) bitmap.* &= ~(@as(u8, 1) << @intCast(level));
t.next = null;
return;
}
}
/// Find a live task by process id, or null. Ids are monotonic and never reused,
/// so a stale id misses cleanly rather than naming a recycled slot.
/// Precondition: the big kernel lock is held.
pub fn taskByIdLocked(id: u32) ?*Task {
for (&tasks) |*t| {
if (t.state != .free and t.id == id) return t;
}
return null;
}
/// Make every server that still holds `t` as the client it owes a reply to forget
/// it — the reply of a dead client is dropped, not delivered into freed state.
/// Precondition: the big kernel lock is held.
pub fn forgetIpcClientLocked(t: *Task) void {
for (&tasks) |*other| {
if (other.state != .free and other.ipc_client == t) other.ipc_client = null;
}
}
/// Block the current task and switch away, without putting it on any wait queue —
/// the caller has already linked it wherever it belongs (e.g. an endpoint's sender
/// FIFO). Precondition: the big kernel lock is held; still held on return (when the
@@ -481,14 +610,61 @@ fn wakeExpired() void {
}
}
// Process-teardown hooks, registered by process.zig at init — the scheduler sits
// below the process layer, so finishing a kill (IRQ bindings, IPC handles, exit
// notification) is called *up* through these, mirroring how the architecture
// layer calls up into `tick`.
//
// `terminate_current_hook` ends the task running on THIS core (lock held, never
// returns — it switches away like `exitUserLocked`). `reap_task_hook` tears down
// a task that is NOT running on any core (lock held).
pub var terminate_current_hook: ?*const fn () noreturn = null;
pub var reap_task_hook: ?*const fn (*Task) void = null;
/// Finish any pending kills this core can see (the deferred half of process_kill;
/// the immediate half runs in the killer's own call). Precondition: the big kernel
/// lock is held, from `tick`.
///
/// - This core's *current* task, if condemned, is terminated here — but only when
/// it is not inside one of its own system calls (`in_system_call`): the tick may
/// have interrupted kernel code mid-operation, where teardown would leak or
/// corrupt what that operation holds. User-mode execution (and the system_call
/// entry/exit stubs, which hold nothing) are safe termination points. A task
/// that *is* mid-call dies at its next block, tick, or system_call entry instead.
/// The hook never returns; abandoning the interrupt frame is fine — the LAPIC
/// was acknowledged before the tick hook ran (see apic.timerTick), exactly as on
/// the fault-kill path.
/// - Condemned tasks that are ready or blocked are not running anywhere (state
/// changes need the lock we hold), so they are reaped in place.
fn reapKillPendingLocked() void {
const pc = thisCpu();
const cur = pc.current;
if (cur.kill_pending and cur.aspace != 0 and !cur.in_system_call) {
if (terminate_current_hook) |hook| hook(); // noreturn
}
if (reap_task_hook) |hook| {
for (&tasks) |*t| {
if (!t.kill_pending) continue;
if (t.state == .ready or t.state == .blocked) hook(t);
}
}
}
/// Called from the timer interrupt (interrupts already disabled): wake due
/// sleepers, then preempt. Takes the kernel lock like any other critical section,
/// but releases it *without* touching the interrupt flag — the handler's `iretq`
/// restores the interrupted context's flags, so re-enabling here would open a
/// nested-interrupt window before the return.
/// sleepers, finish pending kills, then preempt. Takes the kernel lock like any
/// other critical section, but releases it *without* touching the interrupt flag
/// — the handler's `iretq` restores the interrupted context's flags, so
/// re-enabling here would open a nested-interrupt window before the return.
/// Called from the tick with the big kernel lock held — process.zig hangs the
/// one-shot timer sweep here (timer_bind), the same call-up pattern as the
/// teardown hooks below.
pub var timer_tick_hook: ?*const fn () void = null;
pub fn tick() void {
_ = sync.enter();
wakeExpired();
if (timer_tick_hook) |hook| hook();
reapKillPendingLocked();
if (preemption_enabled) schedule();
sync.leaveIsr();
}
@@ -520,6 +696,13 @@ pub fn exit() noreturn {
/// itself is leaked, as in `exit` (no reaper yet). Never returns.
pub fn exitUser() noreturn {
_ = sync.enter();
exitUserLocked();
}
/// The body of `exitUser` for callers that already hold the big kernel lock (the
/// tick-time terminate path, which enters with the lock held). The lock is handed
/// off through the switch and released by the task that resumes. Never returns.
pub fn exitUserLocked() noreturn {
const pc = thisCpu();
const dying = pc.current;
const as = dying.aspace;
@@ -531,6 +714,8 @@ pub fn exitUser() noreturn {
}
dying.state = .free;
dying.aspace = 0;
dying.kill_pending = false;
dying.in_system_call = false;
const next = dequeueHighest(pc) orelse @panic("sched: no task left to run");
next.state = .running;
pc.current = next;
@@ -539,6 +724,54 @@ pub fn exitUser() noreturn {
unreachable;
}
/// Free a task that is NOT running on any core (it is ready or blocked, and the
/// caller — the kill path — has already unlinked it from every queue and released
/// what it held). Destroys its address space: safe here because no core can have
/// it loaded (every switch away from a task loads the next task's tables, and the
/// task isn't running). The kernel stack is leaked, as in `exitUser` (no reaper
/// yet). Precondition: the big kernel lock is held.
pub fn destroyTaskLocked(t: *Task) void {
if (t.aspace != 0) architecture.destroyAddressSpace(t.aspace);
t.aspace = 0;
t.kill_pending = false;
t.in_system_call = false;
t.wake_at = 0;
t.state = .free;
}
/// Snapshot the task table into `out` (up to its length), returning the total
/// number of live tasks — the kernel half of `process_enumerate`, mirroring
/// devices_broker.enumerate. Kernel tasks are included (empty name, supervisor 0):
/// an honest `ps` shows the idle tasks too. `out` may be user memory: the caller's
/// address space is loaded during its system call, and the same bring-up trust
/// applies as for device_enumerate (an unmapped user page faults the kernel).
pub fn enumerate(out: []abi.ProcessDescriptor) u64 {
const flags = sync.enter();
defer sync.leave(flags);
var total: u64 = 0;
for (&tasks) |*t| {
if (t.state == .free) continue;
if (total < out.len) {
const d = &out[total];
d.* = .{
.id = t.id,
.supervisor = t.supervisor,
.state = @intFromEnum(@as(abi.ProcessState, switch (t.state) {
.ready => .ready,
.running => .running,
.blocked => .blocked,
.free => unreachable,
})),
.priority = t.priority,
.name_length = t.name_length,
.name = t.name_buffer,
};
}
total += 1;
}
return total;
}
/// Whether the running task is a user process (has its own address space).
pub fn currentIsUserProcess() bool {
return current().aspace != 0;
+985 -26
View File
File diff suppressed because it is too large Load Diff
+12 -3
View File
@@ -4,7 +4,7 @@
//! hiding the trade-offs. Keeping them here makes them visible at a glance and gives
//! one spot to change them. They're plain `comptime` constants (zero runtime cost);
//! any one can later be promoted to a `-D` build option if a target needs to vary it
//! (see build.zig's `-Dtest-case` for the pattern). This keeps root.zig to what it
//! (see build.zig's `-Dtest-case` for the pattern). This keeps ps2-library.zig to what it
//! actually is — the bootloader↔kernel handoff *contract* — with tunables living here.
/// Ceiling on logical CPUs the kernel tracks — the size of the per-CPU bookkeeping
@@ -16,12 +16,21 @@
pub const maximum_cpus = 128;
/// Maximum tasks (kernel threads) alive at once — the static task-table size. Each
/// online core consumes one slot for its idle task, plus task 0 on the BSP.
pub const maximum_tasks = 16;
/// online core consumes one slot for its idle task, plus task 0 on the BSP. Sized
/// for the initial-ramdisk sweep (15 bundled binaries spawned at once) plus the
/// device manager's supervised children with room to grow — at 16 the sweep
/// started failing spawns once the bundle passed a dozen binaries.
pub const maximum_tasks = 32;
/// Each task's kernel stack (also each AP's bring-up stack), in bytes.
pub const kernel_stack_size = 16 * 1024;
/// Each user process's stack, in pages (32 KiB). Mapped just below a fixed top;
/// the System V entry block (argc/argv) occupies the top of the highest page, and
/// the page below the mapping is left unmapped as a guard, so an overflow faults
/// (killing only that process) instead of silently corrupting the image.
pub const user_stack_pages = 8;
/// Each core's IST (double-fault) stack, in bytes. The BSP's is static; an AP's is
/// heap-allocated at bring-up.
pub const ist_stack_size = 16 * 1024;
+699
View File
@@ -0,0 +1,699 @@
//! /system/services/acpi — the ACPI discovery service: the x86 firmware
//! interpreter, moved out of ring 0 (docs/m19-m20-plan.md, M20). Claims the
//! `acpi-tables` node the kernel publishes (the AML blobs, the broad io_port
//! grant, a broad irq window, the SCI), and runs the **shared AML module** in
//! ring 3 — the same parser and interpreter the kernel uses.
//!
//! It also owns the **event side** (M21): it registers the domain-named `.power`
//! service, binds the SCI (System Control Interrupt), and on a power-button
//! fixed event publishes `power_button` to subscribers — and on init's request
//! writes S5 to power the machine off. The device discovery (M20) and the event
//! handling both run in one `runtime.service.run` loop.
const std = @import("std");
const runtime = @import("runtime");
const aml = @import("aml");
const acpi_ids = @import("acpi-ids");
const device = runtime.device;
const protocol = runtime.device_manager_protocol;
const power = runtime.power_protocol;
/// AML opcode/prefix bytes by name (`zero_opcode`, `byte_prefix`, …) — so the `_HID`
/// integer decode names the opcodes instead of bare 0x0A/0x0B/… (docs/coding-standards.md).
const opcodes = aml.opcodes;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
// The claimed acpi-tables node and the resource index of its broad io_port
// window — the Hal routes every port access through this one claim.
var node_id: u64 = 0;
var io_resource_index: u64 = 0;
// The SCI's irq resource index on the node (the len-1 irq, distinct from the
// broad [0,256) window), for irqBind / irqAck.
var sci_resource_index: u64 = 0;
var has_sci = false;
// PM1 event/control and GPE register ports, read from the FADT copy the kernel
// publishes on the node (M21). Port 0 means absent.
var pm1a_evt: u16 = 0;
var pm1b_evt: u16 = 0;
var pm1_evt_len: u8 = 0;
var pm1a_cnt: u16 = 0;
var pm1b_cnt: u16 = 0;
var gpe0_blk: u16 = 0;
var gpe0_len: u8 = 0;
var gpe1_blk: u16 = 0;
var gpe1_len: u8 = 0;
var smi_cmd: u16 = 0;
var acpi_enable_value: u8 = 0;
var s5_slp_typ_a: u8 = 0;
var s5_slp_typ_b: u8 = 0;
var s5_valid = false;
// PM1 event-register bits (ACPI): PWRBTN in the status/enable word is bit 8;
// the control word's SCI_EN is bit 0; SLP_EN is bit 13.
const pwrbtn_bit: u16 = 1 << 8;
const sci_en_bit: u32 = 1 << 0;
const slp_en: u32 = 1 << 13;
// The `.power` subscribers: endpoints handed over as capabilities, each
// receiving events as buffered messages. Dropped on a failed send. The
// subscriber's task id is kept too — a shutdown request is honored only from a
// subscriber (init subscribes; a stray process does not), the soft gate that
// stands in for "only the system supervisor may power off" without hardcoding
// a pid the kernel's idle tasks would have taken.
const maximum_subscribers = 8;
var subscribers: [maximum_subscribers]?runtime.ipc.Handle = .{null} ** maximum_subscribers;
var subscriber_tasks: [maximum_subscribers]u32 = .{0} ** maximum_subscribers;
// Pass-1 registration record (see main): what pass 2 reports.
const Registered = struct { hid: [8]u8 = .{0} ** 8, hid_len: usize = 0, device_id: u64 = 0, resource_count: u64 = 0 };
var registered: [64]Registered = undefined;
var registered_count: usize = 0;
// A scratch page returned for SystemMemory OperationRegion maps: the service
// cannot map arbitrary physical memory from ring 3, so such regions are
// unsupported and degrade to harmless zeros rather than faulting. The M20.2
// targets (ps2, the legacy devices) use SystemIO and static templates.
var mmio_scratch: [4096]u8 align(4096) = .{0} ** 4096;
fn halMapMmio(physical: u64, len: u64, writable: bool) u64 {
_ = physical;
_ = len;
_ = writable;
return @intFromPtr(&mmio_scratch);
}
fn halPioRead(width: u8, port: u16) u32 {
return device.ioRead(node_id, io_resource_index, port, width) orelse 0;
}
fn halPioWrite(width: u8, port: u16, value: u32) void {
_ = device.ioWrite(node_id, io_resource_index, port, width, value);
}
fn findTablesNode(buffer: []device.DeviceDescriptor) ?device.DeviceDescriptor {
const total = device.enumerate(buffer);
for (buffer[0..@min(total, buffer.len)]) |d| {
if (d.class == @intFromEnum(device.DeviceClass.acpi_tables)) return d;
}
return null;
}
pub fn main(init: runtime.process.Init) void {
// When the acpi-parse scenario spawns this directly, argv[1] is the kernel's
// own device count to self-verify against — deterministic, no log-scraping.
const expected: ?usize = if (init.arguments.get(1)) |a| (std.fmt.parseInt(usize, a, 10) catch null) else null;
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("acpi: out of memory\n");
return;
};
const node = findTablesNode(buffer) orelse {
_ = runtime.system.write("acpi: no acpi-tables node to claim\n");
return;
};
node_id = node.id;
if (!device.claim(node_id)) {
_ = runtime.system.write("acpi: unable to claim acpi-tables\n");
return;
}
// Map the node's resources: the AML blobs (bytecode), the FADT (intact
// "FACP" header — decision 3), the io_port grant, and the SCI irq.
var blocks: [8][]const u8 = undefined;
var block_count: usize = 0;
var found_io = false;
var fadt: ?[]const u8 = null;
for (node.resources[0..@intCast(node.resource_count)], 0..) |resource, index| {
if (resource.kind == @intFromEnum(device.ResourceKind.io_port) and !found_io) {
io_resource_index = index;
found_io = true;
continue;
}
if (resource.kind == @intFromEnum(device.ResourceKind.irq) and resource.len == 1) {
sci_resource_index = index;
has_sci = true;
continue;
}
if (resource.kind != @intFromEnum(device.ResourceKind.memory)) continue;
const base = device.mmioMap(node_id, index) orelse continue;
const pointer: [*]const u8 = @ptrFromInt(base);
const bytes = pointer[0..@intCast(resource.len)];
if (bytes.len >= 4 and std.mem.eql(u8, bytes[0..4], "FACP")) {
fadt = bytes;
continue;
}
if (block_count == blocks.len) continue;
blocks[block_count] = bytes;
block_count += 1;
}
if (block_count == 0) {
_ = runtime.system.write("acpi: no AML blobs on the node\n");
return;
}
const result = aml.parse(runtime.allocator(), blocks[0..block_count]) catch {
_ = runtime.system.write("acpi: AML parse failed\n");
return;
};
var namespace = result.namespace;
const devices = aml.deviceCount(&namespace);
writeLine("acpi: parsed {d} AML blob(s), {d} namespace devices\n", .{ block_count, devices });
if (expected) |want| {
if (devices == want) {
_ = runtime.system.write("acpi-parse: ok\n");
} else {
writeLine("acpi-parse: mismatch (ring-3 {d} vs kernel {d})\n", .{ devices, want });
}
// Self-verify mode is standalone (no manager); stop before reporting.
while (true) runtime.system.sleep(1000);
}
// Register + report the present _HID devices (M20), then set up the power
// event side (M21), then serve — all in one harness loop. The interpreter
// and namespace outlive this frame (static), so the harness callbacks can
// reach them.
interpreter_arena = std.heap.ArenaAllocator.init(runtime.allocator());
persistent_namespace = namespace;
global_interpreter = aml.Interpreter.init(&persistent_namespace, .{
.mapMmio = halMapMmio,
.pioRead = halPioRead,
.pioWrite = halPioWrite,
}, interpreter_arena.allocator());
readFadt(fadt);
s5_valid = readSleepS5(&persistent_namespace);
runtime.service.run(power.message_maximum, .{
.service = .power,
.init = onInit,
.on_message = onMessage,
.on_notification = onNotification,
});
}
// Static so the harness callbacks (which run after main's stack frame is gone)
// can reach the namespace and interpreter.
var persistent_namespace: aml.Namespace = undefined;
var global_interpreter: aml.Interpreter = undefined;
var interpreter_arena: std.heap.ArenaAllocator = undefined;
/// Startup under the harness: register + report the discovered devices to the
/// manager (M20), then enable ACPI mode and arm the power button (M21).
fn onInit(endpoint: runtime.ipc.Handle) bool {
registered_count = 0;
walkDevices(persistent_namespace.root, &global_interpreter);
const manager = runtime.ipc.lookup(.device_manager);
var i: usize = 0;
while (i < registered_count) : (i += 1) {
const entry = registered[i];
const hid = entry.hid[0..entry.hid_len];
const desc = acpi_ids.description(hid);
if (desc.len != 0)
writeLine("acpi: reported {s} (device {d}, {d} resources) — {s}\n", .{ hid, entry.device_id, entry.resource_count, desc })
else
writeLine("acpi: reported {s} (device {d}, {d} resources)\n", .{ hid, entry.device_id, entry.resource_count });
if (manager) |h| {
var report = protocol.ChildAdded{ .parent = node_id, .bus_address = entry.device_id, .identity = 0, .device_id = entry.device_id };
@memcpy(report.hid[0..entry.hid_len], entry.hid[0..entry.hid_len]);
var reply: [protocol.message_maximum]u8 = undefined;
_ = runtime.ipc.call(h, std.mem.asBytes(&report), &reply) catch {};
}
}
writeLine("acpi: reported {d} device(s) to the manager\n", .{registered_count});
armPowerButton(endpoint);
return true;
}
// --- power event side (M21) ---------------------------------------------------
/// Read the PM1 event/control and GPE register ports plus the SMI enable pair
/// from the FADT copy on the node. Offsets are from the FADT table start (the
/// SDT header is the first 36 bytes). Prefers the 32-bit port fields; QEMU's
/// FADT populates them.
fn readFadt(fadt: ?[]const u8) void {
const f = fadt orelse {
_ = runtime.system.write("acpi: no FADT on the node — power events off\n");
return;
};
smi_cmd = @truncate(rd32(f, 48));
acpi_enable_value = f[52];
pm1a_evt = @truncate(rd32(f, 56));
pm1b_evt = @truncate(rd32(f, 60));
pm1a_cnt = @truncate(rd32(f, 64));
pm1b_cnt = @truncate(rd32(f, 68));
gpe0_blk = @truncate(rd32(f, 80));
gpe1_blk = @truncate(rd32(f, 84));
pm1_evt_len = if (f.len > 88) f[88] else 4;
gpe0_len = if (f.len > 92) f[92] else 0;
gpe1_len = if (f.len > 93) f[93] else 0;
}
fn readSleepS5(ns: *aml.Namespace) bool {
const st = aml.sleepState(ns, 5) orelse return false;
s5_slp_typ_a = st.slp_typ_a;
s5_slp_typ_b = st.slp_typ_b;
return true;
}
/// Enable ACPI mode if the firmware isn't already in it, then bind the SCI and
/// set PWRBTN_EN so the power button raises an interrupt we can see.
fn armPowerButton(endpoint: runtime.ipc.Handle) void {
if (pm1a_cnt != 0 and (halPioRead(2, pm1a_cnt) & sci_en_bit) == 0 and smi_cmd != 0) {
// Switch to ACPI mode: write ACPI_ENABLE to the SMI command port, then
// spin (bounded) until SCI_EN latches.
halPioWrite(1, smi_cmd, acpi_enable_value);
var tries: u32 = 0;
while (tries < 1000 and (halPioRead(2, pm1a_cnt) & sci_en_bit) == 0) : (tries += 1) {
runtime.system.sleep(1);
}
}
if (!has_sci) {
_ = runtime.system.write("acpi: no SCI resource — power button unavailable\n");
return;
}
if (!device.irqBind(node_id, sci_resource_index, endpoint)) {
_ = runtime.system.write("acpi: SCI irq_bind failed\n");
return;
}
// PWRBTN_EN lives in the PM1 enable register at evt_blk + evt_len/2.
if (pm1a_evt != 0) {
const en_port = pm1a_evt + pm1_evt_len / 2;
halPioWrite(2, en_port, @as(u16, @truncate(halPioRead(2, en_port))) | pwrbtn_bit);
}
if (pm1b_evt != 0) {
const en_port = pm1b_evt + pm1_evt_len / 2;
halPioWrite(2, en_port, @as(u16, @truncate(halPioRead(2, en_port))) | pwrbtn_bit);
}
_ = runtime.system.write("acpi: power button armed\n");
}
/// The SCI fired. Read PM1 status; a set PWRBTN_STS is the power button — clear
/// it (write-1), publish, log. Any other set status is cleared and logged
/// (GPE/Notify dispatch is M21.2). Always re-arm the line.
fn onSci() void {
var handled = false;
inline for (.{ pm1a_evt, pm1b_evt }) |evt_port| {
if (evt_port != 0) {
const sts: u16 = @truncate(halPioRead(2, evt_port));
if (sts & pwrbtn_bit != 0) {
halPioWrite(2, evt_port, pwrbtn_bit); // write-1-to-clear
handled = true;
} else if (sts != 0) {
halPioWrite(2, evt_port, sts); // clear whatever else latched
}
}
}
if (handled) {
_ = runtime.system.write("power: button pressed\n");
publishButton();
}
handleGpe();
_ = device.irqAck(node_id, sci_resource_index);
}
/// General-purpose events: for each set+enabled GPE bit, evaluate its `\_GPE`
/// handler method (`_Lxx` level / `_Exx` edge), drain the Notify queue the
/// method produced, and publish an event per notified device. Then clear the
/// status bit. QEMU raises no GPEs on this config, so this path is exercised by
/// host unit tests (docs/m21-plan.md decision 5); on real hardware it carries
/// battery/AC/lid. The embedded controller's `_Qxx` queries are out of scope.
fn handleGpe() void {
handleGpeBlock(gpe0_blk, gpe0_len, 0);
handleGpeBlock(gpe1_blk, gpe1_len, gpe0_len * 4);
}
fn handleGpeBlock(blk: u16, len: u8, gpe_base: u32) void {
if (blk == 0 or len == 0) return;
const status_bytes = len / 2; // status half, then enable half
var byte_index: u8 = 0;
while (byte_index < status_bytes) : (byte_index += 1) {
const sts: u8 = @truncate(halPioRead(1, blk + byte_index));
const en: u8 = @truncate(halPioRead(1, blk + status_bytes + byte_index));
const active = sts & en;
if (active == 0) continue;
var bit: u3 = 0;
while (true) : (bit += 1) {
if (active & (@as(u8, 1) << bit) != 0) {
dispatchGpe(gpe_base + @as(u32, byte_index) * 8 + bit);
}
if (bit == 7) break;
}
halPioWrite(1, blk + byte_index, active); // write-1-to-clear the serviced bits
}
}
/// Evaluate the `\_GPE._L%02X` or `_E%02X` handler for GPE number `n`, then
/// publish an event for each device it notified.
fn dispatchGpe(n: u32) void {
const gpe_scope = aml.Namespace.resolve(&persistent_namespace, persistent_namespace.root, true, 0, &.{seg4("_GPE")}) orelse return;
var name: [4]u8 = .{ '_', 'L', 0, 0 };
writeHex2(name[2..4], n);
var method = aml.Namespace.childOf(gpe_scope, name);
if (method == null) {
name[1] = 'E';
method = aml.Namespace.childOf(gpe_scope, name);
}
const m = method orelse return; // no handler — the status bit was already cleared
_ = global_interpreter.evaluate(m, &.{}) catch return;
for (global_interpreter.takeNotifications()) |event| publishNotify(event.node, event.code);
}
fn publishNotify(node: *aml.Node, code: u64) void {
// Map the notified device's _HID to a domain event where we recognize it.
var hid: [8]u8 = .{0} ** 8;
if (readHid(node, &global_interpreter)) |h| hid = h;
const which: power.Event = if (std.mem.eql(u8, hid[0..7], "PNP0C0A")) .battery else if (std.mem.eql(u8, hid[0..7], "ACPI0003")) .ac else if (std.mem.eql(u8, hid[0..7], "PNP0C0D")) .lid else .notify;
var event = power.EventMessage{ .event = @intFromEnum(which), .code = @truncate(code) };
event.hid = hid;
writeLine("power: notify {s} code {d}\n", .{ hid[0..7], code });
publishEvent(std.mem.asBytes(&event));
}
/// Two lowercase hex digits of `n` into `out[0..2]`.
fn writeHex2(out: []u8, n: u32) void {
const digits = "0123456789ABCDEF";
out[0] = digits[(n >> 4) & 0xF];
out[1] = digits[n & 0xF];
}
fn publishButton() void {
const event = power.EventMessage{ .event = @intFromEnum(power.Event.power_button) };
publishEvent(std.mem.asBytes(&event));
}
fn publishEvent(bytes: []const u8) void {
for (&subscribers) |*slot| {
if (slot.*) |handle| {
if (!runtime.ipc.send(handle, bytes)) slot.* = null;
}
}
}
fn isSubscriber(task: u32) bool {
for (&subscribers, 0..) |*slot, si| {
if (slot.* != null and subscriber_tasks[si] == task) return true;
}
return false;
}
/// Enter S5 (soft off): write SLP_TYP|SLP_EN to the PM1 control register(s).
/// Mirrors the kernel's power.zig sleepValue. Only reached from a PID-1
/// shutdown request (M21.3).
fn enterS5() void {
if (!s5_valid or pm1a_cnt == 0) {
_ = runtime.system.write("power: S5 unavailable\n");
return;
}
_ = runtime.system.write("power: entering S5\n");
halPioWrite(2, pm1a_cnt, (@as(u32, s5_slp_typ_a & 0x7) << 10) | slp_en);
if (pm1b_cnt != 0) halPioWrite(2, pm1b_cnt, (@as(u32, s5_slp_typ_b & 0x7) << 10) | slp_en);
// If control returns, the write did not take — say so instead of hanging.
runtime.system.sleep(500);
_ = runtime.system.write("power: S5 write did not take\n");
}
// --- harness callbacks --------------------------------------------------------
fn onNotification(badge: u64) void {
// The only notification the service binds is the SCI (an IRQ badge).
_ = badge;
onSci();
}
/// The `.power` protocol: subscribe (endpoint as the call's capability),
/// shutdown (PID 1 only). Device discovery uses a different endpoint (the
/// device manager's), so nothing here handles ChildAdded.
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
if (message.len < 1) return 0;
switch (message[0]) {
@intFromEnum(power.Operation.subscribe) => {
var status: i32 = -1;
if (capability) |handle| {
for (&subscribers, 0..) |*slot, si| {
if (slot.* == null) {
slot.* = handle;
subscriber_tasks[si] = sender;
status = 0;
break;
}
}
}
const r = power.Reply{ .status = status };
@memcpy(reply[0..@sizeOf(power.Reply)], std.mem.asBytes(&r));
return @sizeOf(power.Reply);
},
@intFromEnum(power.Operation.shutdown) => {
// Honored only from a power subscriber — init, which has already run
// the stop sequence over everything else. The power service is
// mechanism (write S5); deciding *when* to shut down and stopping
// the rest of the system first is init's policy.
const allowed = isSubscriber(sender);
const r = power.Reply{ .status = if (allowed) 0 else -1 };
@memcpy(reply[0..@sizeOf(power.Reply)], std.mem.asBytes(&r));
if (allowed) enterS5();
return @sizeOf(power.Reply);
},
else => return 0,
}
}
/// Depth-first walk: register + report each present device with a _HID, then
/// descend. Scopes (\_SB, \_GPE …) are descended without producing a node.
fn walkDevices(node: *aml.Node, interpreter: *aml.Interpreter) void {
var child = node.first_child;
while (child) |c| : (child = c.next_sibling) {
if (c.kind != .device) {
walkDevices(c, interpreter);
continue;
}
if (!devicePresent(interpreter, c)) continue; // absent: skip it and its subtree
if (readHid(c, interpreter)) |hid| {
// Skip PCI roots — pci-bus already reports PCI functions; ACPI adds
// only the non-PCI _HID devices (docs/m19-m20-plan.md M20.2). The two
// roots are named through the shared registry, not bare _HID strings.
const id = acpi_ids.HardwareId.fromHid(hid[0..7]);
if (id != .pci_bus and id != .pci_express_root_bridge) {
registerDevice(c, hid, interpreter);
}
}
walkDevices(c, interpreter);
}
}
fn registerDevice(node: *aml.Node, hid: [8]u8, interpreter: *aml.Interpreter) void {
if (registered_count >= registered.len) return;
var descriptor = std.mem.zeroes(device.DeviceDescriptor);
descriptor.class = @intFromEnum(device.DeviceClass.acpi_device);
descriptor.pci_class = device.no_pci_class;
const hid_len: u64 = std.mem.indexOfScalar(u8, &hid, 0) orelse hid.len;
descriptor.hid_len = hid_len;
@memcpy(descriptor.hid[0..@intCast(hid_len)], hid[0..@intCast(hid_len)]);
applyCrs(&descriptor, node, interpreter);
const id = device.register(node_id, &descriptor) orelse {
writeLine("acpi: register refused for {s}\n", .{hid[0..@intCast(hid_len)]});
return;
};
registered[registered_count] = .{ .hid = hid, .hid_len = @intCast(hid_len), .device_id = id, .resource_count = descriptor.resource_count };
registered_count += 1;
}
/// _STA bit 0 (present); absent method or a failed evaluation is treated as
/// present, per the ACPI rules.
fn devicePresent(interpreter: *aml.Interpreter, node: *aml.Node) bool {
const sta = aml.Namespace.childOf(node, seg4("_STA")) orelse return true;
const obj = interpreter.evaluate(sta, &.{}) catch return true;
const status = obj.asInteger() catch return true;
return (status & 0x01) != 0;
}
/// The device's EISA-decoded _HID (e.g. "PNP0303"), or null.
fn readHid(node: *aml.Node, interpreter: *aml.Interpreter) ?[8]u8 {
const hid = aml.Namespace.childOf(node, seg4("_HID")) orelse return null;
var buffer: [8]u8 = .{0} ** 8;
if (hid.kind == .method) {
const obj = interpreter.evaluate(hid, &.{}) catch return null;
switch (obj) {
.integer => |n| {
_ = eisaIdToStr(@truncate(n), &buffer);
return buffer;
},
else => return null,
}
}
if (hid.kind != .name or hid.value.len == 0) return null;
const v = hid.value;
switch (v[0]) {
// A static _HID names an integer EISA id: Zero/One/Ones or a Byte/Word/DWord/
// QWord integer prefix. Anything else is not an integer we can EISA-decode.
opcodes.zero_opcode, opcodes.one_opcode, opcodes.ones_opcode, opcodes.byte_prefix, opcodes.word_prefix, opcodes.dword_prefix, opcodes.qword_prefix => {
var p: usize = 0;
const n = readIntObj(v, &p) orelse return null;
_ = eisaIdToStr(@truncate(n), &buffer);
return buffer;
},
else => return null,
}
}
// --- _CRS resource-template decode (ported from the kernel's acpi.zig) --------
/// A resource template is a byte list of descriptors. Each starts with a tag byte whose
/// high bit picks the encoding: a *small* descriptor carries its type in bits [6:3] and
/// its length in bits [2:0]; a *large* descriptor is the whole tag byte, followed by a
/// 16-bit length. These are the descriptor types danos decodes into resources — named so
/// the walk below reads by descriptor, not by 0x04/0x85/… (docs/coding-standards.md).
const large_descriptor_bit: u8 = 0x80; // set in a tag byte => large descriptor
const small_length_mask: u8 = 0x07; // low 3 bits of a small tag = body length
const small_type_shift: u3 = 3; // small type sits in bits [6:3]
/// Small resource descriptor types (tag bits [6:3]). Non-exhaustive: an unhandled type
/// is skipped by its length, not misread.
const SmallResourceType = enum(u8) {
irq = 0x04,
io_port = 0x08,
fixed_io_port = 0x09,
end_tag = 0x0F,
_,
};
/// Large resource descriptor types (the whole tag byte). Non-exhaustive for the same reason.
const LargeResourceType = enum(u8) {
memory32 = 0x85,
memory32_fixed = 0x86,
extended_irq = 0x89,
_,
};
fn applyCrs(descriptor: *device.DeviceDescriptor, node: *aml.Node, interpreter: *aml.Interpreter) void {
const crs = aml.Namespace.childOf(node, seg4("_CRS")) orelse return;
const obj = interpreter.evaluate(crs, &.{}) catch return;
const bytes = switch (obj) {
.buffer => |b| b,
else => return,
};
var i: usize = 0;
while (i < bytes.len) {
const tag = bytes[i];
if (tag & large_descriptor_bit == 0) {
const len: usize = tag & small_length_mask;
const body = i + 1;
if (body + len > bytes.len) break;
switch (@as(SmallResourceType, @enumFromInt((tag >> small_type_shift) & 0x0F))) {
.irq => if (len >= 2) { // IRQ mask
const mask = @as(u16, bytes[body]) | (@as(u16, bytes[body + 1]) << 8);
var b: usize = 0;
while (b < 16) : (b += 1) {
if (mask & (@as(u16, 1) << @intCast(b)) != 0) addResource(descriptor, .irq, b, 1);
}
},
.io_port => if (len >= 7) addResource(descriptor, .io_port, rd16(bytes, body + 1), bytes[body + 6]),
.fixed_io_port => if (len >= 3) addResource(descriptor, .io_port, rd16(bytes, body), bytes[body + 2]),
.end_tag => break,
else => {},
}
i = body + len;
} else {
if (i + 3 > bytes.len) break;
const len: usize = @intCast(rd16(bytes, i + 1));
const body = i + 3;
if (body + len > bytes.len) break;
switch (@as(LargeResourceType, @enumFromInt(tag))) {
.memory32 => if (len >= 17) addResource(descriptor, .memory, rd32(bytes, body + 1), rd32(bytes, body + 13)),
.memory32_fixed => if (len >= 9) addResource(descriptor, .memory, rd32(bytes, body + 1), rd32(bytes, body + 5)),
.extended_irq => if (len >= 2) {
const count = bytes[body + 1];
var k: usize = 0;
while (k < count and body + 2 + k * 4 + 4 <= body + len) : (k += 1) {
addResource(descriptor, .irq, rd32(bytes, body + 2 + k * 4), 1);
}
},
else => {},
}
i = body + len;
}
}
}
fn addResource(descriptor: *device.DeviceDescriptor, kind: device.ResourceKind, start: u64, len: u64) void {
if (descriptor.resource_count >= descriptor.resources.len) return;
descriptor.resources[@intCast(descriptor.resource_count)] = .{ .kind = @intFromEnum(kind), .start = start, .len = len };
descriptor.resource_count += 1;
}
// --- small helpers ported verbatim from the kernel's acpi.zig ----------------
fn seg4(comptime s: *const [4:0]u8) [4]u8 {
return s[0..4].*;
}
fn hexDigit(n: u8) u8 {
return if (n < 10) '0' + n else 'A' + (n - 10);
}
fn eisaIdToStr(id: u32, buffer: *[8]u8) []const u8 {
const b0: u16 = @intCast(id & 0xFF);
const b1: u16 = @intCast((id >> 8) & 0xFF);
const b2: u8 = @truncate(id >> 16);
const b3: u8 = @truncate(id >> 24);
const mfg = (b0 << 8) | b1;
buffer[0] = '@' + @as(u8, @intCast((mfg >> 10) & 0x1F));
buffer[1] = '@' + @as(u8, @intCast((mfg >> 5) & 0x1F));
buffer[2] = '@' + @as(u8, @intCast(mfg & 0x1F));
buffer[3] = hexDigit((b2 >> 4) & 0xF);
buffer[4] = hexDigit(b2 & 0xF);
buffer[5] = hexDigit((b3 >> 4) & 0xF);
buffer[6] = hexDigit(b3 & 0xF);
buffer[7] = 0;
return buffer[0..7];
}
fn readIntObj(bytes: []const u8, p: *usize) ?u64 {
if (p.* >= bytes.len) return null;
const op = bytes[p.*];
p.* += 1;
switch (op) {
opcodes.zero_opcode => return 0,
opcodes.one_opcode => return 1,
opcodes.ones_opcode => return 1,
opcodes.byte_prefix => {
if (p.* >= bytes.len) return null;
const v = bytes[p.*];
p.* += 1;
return v;
},
opcodes.word_prefix => {
if (p.* + 2 > bytes.len) return null;
const v = rd16(bytes, p.*);
p.* += 2;
return v;
},
opcodes.dword_prefix => {
if (p.* + 4 > bytes.len) return null;
const v = rd32(bytes, p.*);
p.* += 4;
return v;
},
else => return null,
}
}
fn rd16(bytes: []const u8, off: usize) u64 {
return @as(u64, bytes[off]) | (@as(u64, bytes[off + 1]) << 8);
}
fn rd32(bytes: []const u8, off: usize) u64 {
return rd16(bytes, off) | (rd16(bytes, off + 2) << 16);
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
+55
View File
@@ -0,0 +1,55 @@
//! args-echo — a test fixture for process arguments (bundled in the
//! initial-ramdisk, spawned only by the `args` test case). Run with no arguments,
//! it respawns itself *with* some via `spawnWithArguments` — exercising the
//! system_spawn argument blob. Run with arguments, it burns more stack than one
//! page could hold (proving the multi-page stack: on a single-page stack the
//! recursion would hit the guard and the process would be killed before echoing),
//! then echoes its whole argv in one `debug_write` the kernel test asserts on —
//! proving the kernel-built System V entry stack (argc, argv pointers,
//! NUL-terminated strings) and the runtime's parsing of it, end to end.
const runtime = @import("runtime");
/// Recurse with a real frame each level: `depth` levels of ~0.5 KiB, touched
/// through a volatile pointer so no optimiser can flatten the frames away.
fn burnStack(depth: usize) u8 {
var frame: [512]u8 = undefined;
const touch: *volatile [512]u8 = &frame;
touch[0] = @truncate(depth);
touch[511] = touch[0];
if (depth == 0) return touch[511];
return touch[0] +% burnStack(depth - 1);
}
pub fn main(init: runtime.process.Init) void {
if (init.arguments.count <= 1) {
// First instance: spawn the second with real arguments, then exit.
_ = runtime.system.spawnWithArguments("args-echo", &.{ "alpha", "beta-42" });
return;
}
// ~16 x 0.5 KiB frames: comfortably past one page, well inside the 32 KiB stack.
_ = burnStack(16);
// Second instance: echo "args: <argv0> <argv1> ..." for the test to match.
var buffer: [128]u8 = undefined;
const prefix = "args:";
@memcpy(buffer[0..prefix.len], prefix);
var len: usize = prefix.len;
var iterator = init.arguments.iterate();
while (iterator.next()) |argument| {
if (len + 1 + argument.len + 1 > buffer.len) break;
buffer[len] = ' ';
len += 1;
@memcpy(buffer[len..][0..argument.len], argument);
len += argument.len;
}
buffer[len] = '\n';
len += 1;
_ = runtime.system.write(buffer[0..len]);
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
+44
View File
@@ -0,0 +1,44 @@
//! crash-test — a test fixture, not a driver: claims the device it is assigned,
//! hellos the device manager, announces itself, then faults on purpose. The
//! driver-restart scenario drives the manager's whole restart machinery with
//! it: fault → exit reason → backoff → respawn → the **same claim succeeding
//! again** (claim release on death, M17.1, through the manager's path) → the
//! crash-loop cap. Spawned bare (the initial-ramdisk sweep starts every bundled
//! binary), it exits silently so it cannot derange other tests.
const std = @import("std");
const runtime = @import("runtime");
const protocol = runtime.device_manager_protocol;
pub fn main(init: runtime.process.Init) void {
const argument = init.arguments.get(1) orelse return; // bare: stay silent
const assigned = std.fmt.parseInt(u64, argument, 10) catch return;
// The respawn only reaches this line because the kernel released the
// previous instance's claim at death. A failed claim exits cleanly — the
// manager reads "meant to stop" and the scenario fails loudly by silence.
if (!runtime.device.claim(assigned)) {
_ = runtime.system.write("crash-test: claim failed\n");
return;
}
var manager: ?runtime.ipc.Handle = null;
var tries: u32 = 0;
while (manager == null and tries < 100) : (tries += 1) {
manager = runtime.ipc.lookup(.device_manager);
if (manager == null) runtime.system.sleep(20);
}
const h = manager orelse return;
const hello = protocol.Hello{ .role = @intFromEnum(protocol.Role.device), .device_id = assigned };
var reply: [protocol.message_maximum]u8 = undefined;
_ = runtime.ipc.call(h, std.mem.asBytes(&hello), &reply) catch return;
_ = runtime.system.write("crash-test: faulting now\n");
const poison: *volatile u32 = @ptrFromInt(0xdead0000);
poison.* = 1; // the restart machinery's fuel: a real segmentation fault
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
@@ -0,0 +1,88 @@
//! device-list — the `ps` analog for the device tree (docs/device-manager.md
//! M18.3): asks the device manager for the tree over IPC, prints it, then
//! subscribes and prints every published add/remove event. The manager is the
//! one answer to "what devices exist" for user space; nothing here touches a
//! device_* system call.
const std = @import("std");
const runtime = @import("runtime");
const protocol = runtime.device_manager_protocol;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [96]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
pub fn main() void {
var manager: ?runtime.ipc.Handle = null;
var tries: u32 = 0;
while (manager == null and tries < 200) : (tries += 1) {
manager = runtime.ipc.lookup(.device_manager);
if (manager == null) runtime.system.sleep(20);
}
const h = manager orelse {
_ = runtime.system.write("device-list: no device manager\n");
return;
};
// The snapshot — polled briefly, because at boot the bus drivers may still
// be scanning: an empty first answer usually just means "too early".
var reply: [protocol.message_maximum]u8 = undefined;
var count: u32 = 0;
var length: usize = 0;
tries = 0;
while (tries < 20) : (tries += 1) {
const request = protocol.Enumerate{};
length = runtime.ipc.call(h, std.mem.asBytes(&request), &reply) catch 0;
if (length >= @sizeOf(protocol.EnumerateReply)) {
count = std.mem.bytesToValue(protocol.EnumerateReply, reply[0..@sizeOf(protocol.EnumerateReply)]).count;
if (count != 0) break;
}
runtime.system.sleep(100);
}
writeLine("device-list: {d} devices\n", .{count});
var offset: usize = @sizeOf(protocol.EnumerateReply);
var index: u32 = 0;
while (index < count and offset + @sizeOf(protocol.ChildEntry) <= length) : (index += 1) {
const entry = std.mem.bytesToValue(protocol.ChildEntry, reply[offset..][0..@sizeOf(protocol.ChildEntry)]);
writeLine("device-list: device {d} port {d} identity {d}\n", .{ entry.parent, entry.bus_address, entry.identity });
offset += @sizeOf(protocol.ChildEntry);
}
// The subscription: our endpoint rides as the call's capability; events
// arrive as buffered messages carrying the same structs the bus sends.
const endpoint = runtime.ipc.createIpcEndpoint() orelse {
_ = runtime.system.write("device-list: no endpoint\n");
return;
};
const subscribe = protocol.Subscribe{};
_ = runtime.ipc.callCap(h, std.mem.asBytes(&subscribe), &reply, endpoint) catch {
_ = runtime.system.write("device-list: subscribe failed\n");
return;
};
_ = runtime.system.write("device-list: subscribed\n");
var receive: [protocol.message_maximum]u8 = undefined;
while (true) {
const got = runtime.ipc.replyWait(endpoint, &.{}, &receive, null);
if (!got.isMessage() or got.len < 1) continue;
switch (receive[0]) {
@intFromEnum(protocol.Operation.child_added) => {
if (got.len < protocol.child_added_size) continue;
const event = std.mem.bytesToValue(protocol.ChildAdded, receive[0..protocol.child_added_size]);
writeLine("device-list: added (device {d} port {d})\n", .{ event.parent, event.bus_address });
},
@intFromEnum(protocol.Operation.child_removed) => {
if (got.len < protocol.child_removed_size) continue;
const event = std.mem.bytesToValue(protocol.ChildRemoved, receive[0..protocol.child_removed_size]);
writeLine("device-list: removed (device {d} port {d})\n", .{ event.parent, event.bus_address });
},
else => {},
}
}
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
@@ -0,0 +1,148 @@
//! The device-manager protocol (docs/device-manager.md): what drivers and
//! applications say to the device manager over its well-known endpoint. The
//! vfs-protocol pattern — extern-struct messages, a version in the handshake,
//! reserved fields — so both sides depend on the contract by name. Deliberately
//! contains nothing lifecycle-shaped: stopping, liveness (the zero-length ping),
//! and exit reasons are the universal vocabulary of
//! docs/process-lifecycle.md, not this protocol.
/// The protocol version a driver states in its hello. A manager that cannot
/// serve a driver's version refuses the hello, and the mismatch is loud at
/// startup instead of quiet corruption later.
pub const version: u16 = 1;
/// What kind of driver is talking (docs/driver-model.md's shapes).
pub const Role = enum(u8) {
/// Owns a controller and reports the devices behind it (`child_added`).
bus = 1,
/// Serves one device, reached through a bus's transfer protocol.
device = 2,
};
/// The message kinds.
pub const Operation = enum(u8) {
hello = 1,
child_added = 2,
child_removed = 3,
enumerate = 4,
subscribe = 5,
};
/// `Hello.device_id` for a driver that serves no enumerated device (a test
/// fixture, a synthetic source).
pub const no_device: u64 = ~@as(u64, 0);
/// The handshake, sent once by every driver the manager spawns — the manager's
/// one self-enforced deadline: spawned and silent past it means wrong binary,
/// wrong version, or wedged before main, and the stop sequence follows.
pub const Hello = extern struct {
operation: u8 = @intFromEnum(Operation.hello),
/// A Role value.
role: u8,
/// The protocol version this driver was built against (`version`).
version: u16 = version,
reserved: u32 = 0,
/// The device this driver was assigned (its argv[1]), or `no_device`.
device_id: u64,
};
pub const hello_size = @sizeOf(Hello);
/// The manager's answer to a hello. Nonzero status = refused (version mismatch,
/// unknown sender); a refused driver should exit cleanly.
pub const HelloReply = extern struct {
status: i32,
reserved: u32 = 0,
};
pub const reply_size = @sizeOf(HelloReply);
/// A bus driver reporting one device it discovered behind its controller
/// (docs/device-manager.md "the tree"). Identity is the bus's native language —
/// for USB a port-speed class; the (class, subclass, protocol) triple joins it
/// once control transfers exist (the USB track). The manager mirrors the child
/// into its tree; when the reporting driver dies, the manager prunes everything
/// it reported (the children describe protocol state that died with it) and the
/// restarted instance rediscovers and re-reports.
pub const ChildAdded = extern struct {
operation: u8 = @intFromEnum(Operation.child_added),
reserved0: u8 = 0,
reserved1: u16 = 0,
reserved2: u32 = 0,
/// The reporting driver's own device (the controller) — the child's parent.
parent: u64,
/// Where on the bus (for USB: the root port number, 1-based).
bus_address: u64,
/// Bus-specific identity (for USB: the PORTSC port-speed class; for PCI:
/// the class triple; for ACPI devices, 0 — identity is the hid below).
identity: u64,
/// The kernel device id this child was `device_register`ed as — what the
/// manager hands a matched driver as its argv assignment — or `no_device`
/// for an unregistered leaf (a USB port before the descriptor track).
device_id: u64 = no_device,
/// The ACPI hardware id (`_HID`), EISA-decoded (e.g. "PNP0303"), for devices
/// discovered by firmware string rather than a numeric bus identity. Empty
/// (all zero) otherwise. Widens for FDT `compatible` strings later.
hid: [8]u8 = .{0} ** 8,
};
pub const child_added_size = @sizeOf(ChildAdded);
/// A bus driver reporting a device gone (hot-unplug). Not yet sent by any
/// driver — the port scan has no unplug interrupt — but the manager handles it;
/// death-pruning covers removal until hotplug lands.
pub const ChildRemoved = extern struct {
operation: u8 = @intFromEnum(Operation.child_removed),
reserved0: u8 = 0,
reserved1: u16 = 0,
reserved2: u32 = 0,
parent: u64,
bus_address: u64,
};
pub const child_removed_size = @sizeOf(ChildRemoved);
/// The manager's answer to a tree report.
pub const ReportReply = extern struct {
status: i32,
reserved: u32 = 0,
};
/// An application asking for the tree (M18.3): the reply is an EnumerateReply
/// header followed by `count` ChildEntry records.
pub const Enumerate = extern struct {
operation: u8 = @intFromEnum(Operation.enumerate),
reserved0: u8 = 0,
reserved1: u16 = 0,
reserved2: u32 = 0,
};
pub const EnumerateReply = extern struct {
status: i32,
/// ChildEntry records following this header.
count: u32,
};
pub const ChildEntry = extern struct {
parent: u64,
bus_address: u64,
identity: u64,
};
/// An application subscribing to published add/remove events (the input-service
/// pattern): the subscriber's endpoint rides as the call's **capability**, and
/// events arrive on it as buffered messages whose payload is the same
/// ChildAdded / ChildRemoved struct the bus drivers send — one encoding, both
/// directions.
pub const Subscribe = extern struct {
operation: u8 = @intFromEnum(Operation.subscribe),
reserved0: u8 = 0,
reserved1: u16 = 0,
reserved2: u32 = 0,
};
/// Upper bound on any message in this protocol — sizes the endpoint buffers.
/// Capped by the kernel's IPC MESSAGE_MAXIMUM (256): an EnumerateReply carries
/// up to ten ChildEntry records per call, plenty for the mirror's current
/// bounds; paging joins the protocol if a tree ever outgrows one message.
pub const message_maximum = 256;
+520 -32
View File
@@ -1,60 +1,548 @@
//! /system/services/device-manager — the ring-3 process that turns the device
//! tree into a running system. The kernel enumerates the hardware and enforces the
//! claim capability (mechanism); this decides *which driver serves which device*
//! and, eventually, spawns it (policy). Keeping that split in user space is the
//! whole point of the microkernel: the manager is an ordinary, restartable process
//! with no special privilege — it uses the same `device_*` system calls any process
//! could ([drivers.md](../../../docs/drivers.md), [driver-model.md]).
//! tree into a running system: **the matcher and the supervisor**
//! (docs/device-manager.md). The kernel enumerates the hardware and enforces the
//! claim capability (mechanism); this decides which driver serves which device,
//! spawns it, and keeps it alive (policy). Keeping that split in user space is
//! the whole point of the microkernel: the manager is an ordinary, restartable
//! process with no special privilege.
//!
//! Increment 2 (this file): enumerate /system/devices, *match* each device to a
//! driver, and *spawn* it with `system_spawn` — the kernel loads the named binary
//! from the initial-ramdisk as a fresh ring-3 process. On QEMU this discovers the
//! HPET, decides `hpet` serves it, and brings that driver all the way up. (The
//! kernel still auto-spawns the whole initial-ramdisk at boot; increment 3 removes
//! that redundancy so the manager is the sole owner of driver spawning.)
//! M18.1 (this increment): the manager is a harness service on the well-known
//! `.device_manager` endpoint. Every driver is spawned **supervised** — exit
//! notifications land in the same loop as protocol messages. Drivers with an
//! assignment must `hello` within a deadline or be stopped; a driver that dies
//! is restarted with backoff, and a crash loop (three fast deaths) marks it
//! failed instead of respawning forever. Exit reasons (M17.2) drive the
//! decision: a clean exit meant to stop; only faults and missed deadlines
//! restart. Tree reports (`child_added`) land in M18.2.
const std = @import("std");
const runtime = @import("runtime");
const acpi_ids = @import("acpi-ids");
const pci_class = @import("pci-class");
const protocol = runtime.device_manager_protocol;
const device = runtime.device;
const system = runtime.system;
/// The driver that serves each device class — the policy table. In a fuller system
/// this comes from the drivers describing what they bind (or a manifest under
/// /system/drivers); for now it is a small static map, which is enough to prove the
/// manager reads the tree and decides. `null` = no driver for this class yet.
fn driverFor(class: u64) ?[]const u8 {
if (class == @intFromEnum(device.DeviceClass.timer)) return "hpet"; // the HPET
/// Format one whole log line and emit it in a single `debug_write`, so output
/// from the drivers this manager starts (which run concurrently) can never land
/// in the middle of it.
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
/// The driver that serves each device — the policy table. In a fuller system
/// this comes from a manifest (docs/device-manager.md: the third bus type
/// triggers it); for now a static map. `null` = no driver for this class yet.
fn driverFor(d: device.DeviceDescriptor) ?[]const u8 {
// The HPET timer node is still kernel-seeded (from the HPET table, not AML).
// PS/2 and other _HID devices now arrive as acpi-service reports and match
// in onChildAdded (M20.3), not from this boot snapshot.
if (d.class == @intFromEnum(device.DeviceClass.timer)) return "hpet";
return null;
}
pub fn main() void {
/// The PCI class/subclass/prog-IF triple of an xHCI (USB 3) host controller —
/// Serial Bus Controller / USB Controller / XHCI — named from pci-class.zig rather
/// than written as the bare 0x0C0330 (docs/coding-standards.md, "Named values").
const xhci_pci_class: u64 = pci_class.ClassCode.pack(.{
.base = @intFromEnum(pci_class.BaseClass.serial_bus),
.subclass = @intFromEnum(pci_class.serial_bus.SubClass.usb),
.prog_if = @intFromEnum(pci_class.serial_bus.usb.ProgIf.xhci),
});
/// The driver that serves a *reported* PCI function (M19.3: matching moved
/// from the boot snapshot to the bus reports), or null. A machine can carry
/// several identical controllers — one driver instance per reported device,
/// its registered id as argv[1].
fn pciDriverForIdentity(identity: u64) ?[]const u8 {
return switch (identity) {
xhci_pci_class => "usb-xhci-bus",
else => null,
};
}
/// The driver that serves a *reported* ACPI device by its `_HID` (M20.3:
/// ps2-bus now binds the PS/2 nodes the acpi service reports, not boot-snapshot
/// nodes the kernel used to build). ps2-bus is a singleton that finds both its
/// devices by hid once spawned, so keyboard and mouse map to the same name.
fn hidDriverFor(hid: []const u8) ?[]const u8 {
if (std.mem.eql(u8, hid, "PNP0303")) return "ps2-bus"; // PS/2 keyboard
if (std.mem.eql(u8, hid, "PNP0F13")) return "ps2-bus"; // PS/2 mouse
return null;
}
/// Whether some driver entry already serves registered device `device_id` —
/// a re-report after a bus restart must not spawn a second instance.
fn driverForDevice(device_id: u64) bool {
for (&drivers) |*driver| {
if (driver.used and driver.device_id == device_id) return true;
}
return false;
}
// --- supervision -------------------------------------------------------------
/// How long a protocol driver has to hello after its spawn.
const hello_deadline_ms: u64 = 3000;
/// Deaths faster than this count toward the crash loop; slower ones reset it.
const fast_death_ns: u64 = 2_000_000_000;
/// Consecutive fast deaths before the manager gives up on a driver.
const crash_loop_cap: u32 = 3;
/// Restart backoff: base << (restarts - 1), so 300 ms, 600 ms, 1200 ms.
const backoff_base_ms: u64 = 300;
const DriverState = enum {
awaiting_hello, // spawned; the deadline is armed (protocol drivers only)
running,
restarting, // dead; respawn due at restart_due_ns
stopped, // exited cleanly — it meant to; not restarted
failed, // crash loop, or unspawnable; the manager gave up
};
const Driver = struct {
used: bool = false,
name_buffer: [24]u8 = undefined,
name_len: usize = 0,
// The assigned device id (becomes argv[1]), or protocol.no_device.
device_id: u64 = protocol.no_device,
// Whether this driver speaks the protocol (hello expected, deadline
// enforced). Legacy drivers (hpet, ps2-bus) are supervised and restarted
// but not yet required to hello.
speaks_protocol: bool = false,
process_id: u32 = 0,
state: DriverState = .running,
restarts: u32 = 0,
spawn_ns: u64 = 0,
hello_deadline_ns: u64 = 0,
restart_due_ns: u64 = 0,
fn name(driver: *const Driver) []const u8 {
return driver.name_buffer[0..driver.name_len];
}
};
const maximum_drivers = 16;
var drivers: [maximum_drivers]Driver = .{Driver{}} ** maximum_drivers;
var manager_endpoint: runtime.ipc.Handle = 0;
var test_restart_mode = false;
var test_usb_restart_mode = false;
var test_usb_killed = false;
var test_pci_restart_mode = false;
var test_kill_pid: u32 = 0;
var test_kill_due_ns: u64 = 0;
/// The application subscribers (M18.3, the input-service pattern): endpoints
/// handed over as capabilities, each receiving every child add/remove as a
/// buffered message. A subscriber whose endpoint stops accepting (it died) is
/// dropped on the failed send.
const maximum_subscribers = 8;
var subscribers: [maximum_subscribers]?runtime.ipc.Handle = .{null} ** maximum_subscribers;
/// Publish one event (a ChildAdded or ChildRemoved struct, the same encoding
/// the bus drivers send) to every subscriber.
fn publishEvent(event: []const u8) void {
for (&subscribers) |*slot| {
if (slot.*) |handle| {
if (!runtime.ipc.send(handle, event)) slot.* = null; // dead subscriber
}
}
}
/// The manager's mirror of what bus drivers report (docs/device-manager.md "the
/// tree"): the children, keyed by (parent, bus address), each remembering which
/// driver instance reported it — that is what death-pruning sweeps by.
const Child = struct {
used: bool = false,
parent: u64 = 0,
bus_address: u64 = 0,
identity: u64 = 0,
// The kernel device id (registered by the reporter), or protocol.no_device.
device_id: u64 = 0,
reporter: u32 = 0, // the reporting driver instance's process id
};
const maximum_children = 64; // ACPI adds ~34 device nodes (M20.2), plus PCI + USB
var children: [maximum_children]Child = .{Child{}} ** maximum_children;
/// Record (or refresh) a reported child. Refreshing matters: a restarted bus
/// driver re-reports what it rediscovers, and the same (parent, port) must not
/// duplicate.
fn addChild(parent: u64, bus_address: u64, identity: u64, device_id: u64, reporter: u32) bool {
var free: ?*Child = null;
for (&children) |*child| {
if (child.used and child.parent == parent and child.bus_address == bus_address) {
child.identity = identity;
child.device_id = device_id;
child.reporter = reporter;
return true;
}
if (!child.used and free == null) free = child;
}
const slot = free orelse return false;
slot.* = .{ .used = true, .parent = parent, .bus_address = bus_address, .identity = identity, .device_id = device_id, .reporter = reporter };
return true;
}
/// Prune every child a dead driver instance reported: the children describe
/// protocol state (slots, rings) that died with the process — keeping the nodes
/// would be keeping a lie. The restarted instance rediscovers and re-reports.
/// Watchers hear the honest story: removed now, added again on rediscovery.
fn pruneChildrenOf(reporter: u32) void {
for (&children) |*child| {
if (child.used and child.reporter == reporter) {
writeLine("device-manager: child removed (device {d} port {d})\n", .{ child.parent, child.bus_address });
child.used = false;
const event = protocol.ChildRemoved{ .parent = child.parent, .bus_address = child.bus_address };
publishEvent(std.mem.asBytes(&event));
}
}
}
/// How many children a driver instance has reported (the test-usb-restart
/// trigger counts these).
fn childCountOf(reporter: u32) u32 {
var n: u32 = 0;
for (&children) |*child| {
if (child.used and child.reporter == reporter) n += 1;
}
return n;
}
fn driverByProcess(process_id: u32) ?*Driver {
for (&drivers) |*driver| {
if (driver.used and driver.process_id == process_id) return driver;
}
return null;
}
/// Whether a singleton driver is already in the table (two ACPI nodes can both
/// map to ps2-bus; one instance serves both).
fn alreadySupervised(name: []const u8) bool {
for (&drivers) |*driver| {
if (driver.used and std.mem.eql(u8, driver.name(), name)) return true;
}
return false;
}
/// Record a driver in the table and spawn its first instance.
fn addDriver(name: []const u8, device_id: u64, speaks_protocol: bool) void {
for (&drivers) |*driver| {
if (driver.used) continue;
const n = @min(name.len, driver.name_buffer.len);
@memcpy(driver.name_buffer[0..n], name[0..n]);
driver.name_len = n;
driver.device_id = device_id;
driver.speaks_protocol = speaks_protocol;
driver.used = true;
spawnDriver(driver);
return;
}
writeLine("device-manager: driver table full; cannot supervise {s}\n", .{name});
}
/// (Re)spawn a driver instance: supervised on the manager's own endpoint, the
/// device id as argv[1] when it has one, the hello deadline armed when it
/// speaks the protocol.
fn spawnDriver(driver: *Driver) void {
var id_text: [20]u8 = undefined;
var arguments: [1][]const u8 = undefined;
var argument_count: usize = 0;
if (driver.device_id != protocol.no_device) {
arguments[0] = std.fmt.bufPrint(&id_text, "{d}", .{driver.device_id}) catch return;
argument_count = 1;
}
const child = system.spawnSupervised(driver.name(), arguments[0..argument_count], manager_endpoint) orelse {
writeLine("device-manager: failed to spawn {s}\n", .{driver.name()});
driver.state = .failed;
return;
};
driver.process_id = child;
driver.spawn_ns = system.clock();
if (driver.speaks_protocol) {
driver.state = .awaiting_hello;
driver.hello_deadline_ns = driver.spawn_ns + hello_deadline_ms * 1_000_000;
_ = system.timerOnce(manager_endpoint, hello_deadline_ms + 100);
} else {
driver.state = .running;
}
if (driver.device_id != protocol.no_device) {
writeLine("device-manager: spawned {s} for device {d}\n", .{ driver.name(), driver.device_id });
} else {
writeLine("device-manager: spawned {s}\n", .{driver.name()});
}
}
/// A driver died. Prune what it reported first — then the exit reason (M17.2)
/// is the whole restart decision: a clean exit meant to stop; anything else
/// restarts with backoff until the crash-loop cap.
fn onDriverExit(driver: *Driver) void {
pruneChildrenOf(driver.process_id);
const reason = runtime.process.exitReason(driver.process_id) orelse .fault;
if (reason == .exited) {
driver.state = .stopped;
writeLine("device-manager: {s} exited cleanly; not restarting\n", .{driver.name()});
return;
}
const now = system.clock();
const alive_ns = now - driver.spawn_ns;
driver.restarts = if (alive_ns < fast_death_ns) driver.restarts + 1 else 1;
if (driver.restarts >= crash_loop_cap) {
driver.state = .failed;
writeLine("device-manager: {s} is failing repeatedly (crash loop); giving up\n", .{driver.name()});
return;
}
const delay_ms = backoff_base_ms << @intCast(driver.restarts - 1);
driver.state = .restarting;
driver.restart_due_ns = now + delay_ms * 1_000_000;
writeLine("device-manager: restarting {s} in {d} ms (died: {s})\n", .{ driver.name(), delay_ms, @tagName(reason) });
_ = system.timerOnce(manager_endpoint, delay_ms + 50);
}
/// A timer landed: sweep every deadline. Overdue hellos are killed (the exit
/// notification then routes through the normal restart policy); due restarts
/// respawn. Timers carry no id on purpose — the table is the state, and one
/// sweep serves every armed deadline.
fn sweepDeadlines() void {
const now = system.clock();
if (test_kill_pid != 0 and now >= test_kill_due_ns) {
writeLine("device-manager: test mode: killing the reporter\n", .{});
_ = system.kill(test_kill_pid);
test_kill_pid = 0;
}
for (&drivers) |*driver| {
if (!driver.used) continue;
switch (driver.state) {
.awaiting_hello => if (now >= driver.hello_deadline_ns) {
writeLine("device-manager: {s} missed its hello deadline\n", .{driver.name()});
_ = system.kill(driver.process_id);
// The exit notification finishes the job via onDriverExit.
},
.restarting => if (now >= driver.restart_due_ns) spawnDriver(driver),
else => {},
}
}
}
// --- the harness callbacks -----------------------------------------------------
fn initialise(endpoint: runtime.ipc.Handle) bool {
manager_endpoint = endpoint;
// Enumerate into a heap buffer (too big for the one-page user stack).
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("device-manager: out of memory\n");
return;
return false;
};
const total = device.enumerate(buffer);
const count = @min(total, buffer.len);
var matched: usize = 0;
for (buffer[0..count]) |descriptor| {
const driver_name = driverFor(descriptor.class) orelse continue;
matched += 1;
if (runtime.system.spawn(driver_name)) {
_ = runtime.system.write("device-manager: spawned ");
_ = runtime.system.write(driver_name);
_ = runtime.system.write("\n");
} else {
_ = runtime.system.write("device-manager: failed to spawn ");
_ = runtime.system.write(driver_name);
_ = runtime.system.write("\n");
if (descriptor.class == @intFromEnum(device.DeviceClass.pci_host_bridge)) {
// The PCI bus driver: enumeration in ring 3 (M19), one instance
// per bridge, the bridge id as its assignment.
matched += 1;
addDriver("pci-bus", descriptor.id, true);
continue;
}
// PCI functions no longer appear in the boot snapshot (M19.3): the
// pci-bus driver reports them, and onChildAdded matches from reports.
const driver_name = driverFor(descriptor) orelse continue;
matched += 1;
// Skip a singleton that is already alive (the initial-ramdisk sweep test
// starts every bundled binary bare, this manager included) — spawning a
// second instance would only lose the claim race and churn the log.
if (!alreadySupervised(driver_name) and !system.isProcessRunning(driver_name)) {
addDriver(driver_name, protocol.no_device, false);
}
}
// The discovery service (docs/m19-m20-plan.md M20): one per firmware, packed
// under the neutral name "discovery", spawned once at startup. It finds and
// claims the acpi-tables (or devicetree-blob) node itself. Not a per-device
// match — it is the discoverer, not a driver bound to one device.
addDriver("discovery", protocol.no_device, false);
if (test_restart_mode) {
// The driver-restart scenario's fixture: claims device 0 (the tree
// root, otherwise unclaimed), hellos, then faults — driving backoff,
// re-claim-after-death, and the crash-loop cap deterministically.
addDriver("crash-test", 0, true);
}
if (matched == 0) {
_ = runtime.system.write("device-manager: no matchable devices\n");
} else {
_ = runtime.system.write("device-manager: ok\n");
}
return true;
}
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
if (message.len < 1) return 0;
switch (message[0]) {
@intFromEnum(protocol.Operation.child_added) => return onChildAdded(message, reply, sender),
@intFromEnum(protocol.Operation.child_removed) => return onChildRemoved(message, reply, sender),
@intFromEnum(protocol.Operation.enumerate) => return onEnumerate(reply),
@intFromEnum(protocol.Operation.subscribe) => return onSubscribe(reply, capability),
@intFromEnum(protocol.Operation.hello) => {},
else => return 0,
}
if (message.len < protocol.hello_size) return 0;
const hello = std.mem.bytesToValue(protocol.Hello, message[0..protocol.hello_size]);
var status: i32 = 0;
if (hello.version != protocol.version) {
status = -1;
writeLine("device-manager: refused hello (version {d}) from process {d}\n", .{ hello.version, sender });
} else if (driverByProcess(sender)) |driver| {
driver.state = .running;
writeLine("device-manager: hello from {s} (device {d})\n", .{ driver.name(), hello.device_id });
} else {
status = -1;
writeLine("device-manager: hello from unknown process {d}\n", .{sender});
}
const hello_reply = protocol.HelloReply{ .status = status };
@memcpy(reply[0..protocol.reply_size], std.mem.asBytes(&hello_reply));
return protocol.reply_size;
}
/// A bus driver reported a discovered device: mirror it, and in
/// test-usb-restart mode kill the reporter once after its second child — the
/// deterministic trigger for prune -> backoff -> respawn -> re-report.
fn onChildAdded(message: []const u8, reply: []u8, sender: u32) usize {
if (message.len < protocol.child_added_size) return 0;
const report = std.mem.bytesToValue(protocol.ChildAdded, message[0..protocol.child_added_size]);
var status: i32 = 0;
if (driverByProcess(sender)) |driver| {
if (!addChild(report.parent, report.bus_address, report.identity, report.device_id, sender)) status = -1;
writeLine("device-manager: child added (device {d} port {d}, identity {d}) by {s}\n", .{ report.parent, report.bus_address, report.identity, driver.name() });
if (status == 0) publishEvent(message[0..protocol.child_added_size]);
// Matching from reports (M19.3): a registered child whose identity
// names a driver gets one, once — re-reports after a bus restart
// dedupe on the registered id, exactly like the registrations do.
if (status == 0 and report.device_id != protocol.no_device) {
if (pciDriverForIdentity(report.identity)) |child_driver| {
if (!driverForDevice(report.device_id)) addDriver(child_driver, report.device_id, true);
}
// ACPI _HID match (M20.3): ps2-bus is a singleton that finds its own
// devices by hid, so spawn it once, without a device assignment.
const hid_len = std.mem.indexOfScalar(u8, &report.hid, 0) orelse report.hid.len;
if (hid_len != 0) {
if (hidDriverFor(report.hid[0..hid_len])) |hid_driver| {
if (!alreadySupervised(hid_driver)) addDriver(hid_driver, protocol.no_device, false);
}
}
}
} else {
status = -1;
}
const report_reply = protocol.ReportReply{ .status = status };
@memcpy(reply[0..@sizeOf(protocol.ReportReply)], std.mem.asBytes(&report_reply));
if (test_pci_restart_mode and !test_usb_killed) {
if (driverByProcess(sender)) |driver| {
if (std.mem.eql(u8, driver.name(), "pci-bus") and childCountOf(sender) >= 3) {
// The pci restart drill: kill the enumerator after it has
// reported; the respawn must re-register without duplicates
// (M19.0 idempotence, proven end to end by pci-scan).
test_usb_killed = true;
test_kill_pid = sender;
test_kill_due_ns = system.clock() + 1_000_000_000;
_ = system.timerOnce(manager_endpoint, 1100);
}
}
}
if (test_usb_restart_mode and !test_usb_killed and childCountOf(sender) >= 2) {
// Only the xHCI reporter is the drill's victim — pci-bus also reports
// now, and whichever finishes second must not trigger the kill.
if (driverByProcess(sender)) |driver| {
if (std.mem.eql(u8, driver.name(), "usb-xhci-bus")) {
// Delayed, not immediate: the device-list scenario's subscriber
// needs a window to enumerate and subscribe before the events.
test_usb_killed = true;
test_kill_pid = sender;
test_kill_due_ns = system.clock() + 2_000_000_000;
_ = system.timerOnce(manager_endpoint, 2100);
}
}
}
return @sizeOf(protocol.ReportReply);
}
/// A bus driver reported a device gone (hot-unplug; no sender exists yet, but
/// the handler is protocol-complete — death-pruning covers removal until then).
fn onChildRemoved(message: []const u8, reply: []u8, sender: u32) usize {
if (message.len < protocol.child_removed_size) return 0;
const report = std.mem.bytesToValue(protocol.ChildRemoved, message[0..protocol.child_removed_size]);
var status: i32 = -1;
for (&children) |*child| {
if (child.used and child.parent == report.parent and child.bus_address == report.bus_address and child.reporter == sender) {
writeLine("device-manager: child removed (device {d} port {d})\n", .{ child.parent, child.bus_address });
child.used = false;
status = 0;
}
}
const report_reply = protocol.ReportReply{ .status = status };
@memcpy(reply[0..@sizeOf(protocol.ReportReply)], std.mem.asBytes(&report_reply));
return @sizeOf(protocol.ReportReply);
}
/// An application asked for the tree: the mirror, as a header plus entries.
fn onEnumerate(reply: []u8) usize {
var count: u32 = 0;
var offset: usize = @sizeOf(protocol.EnumerateReply);
for (&children) |*child| {
if (!child.used) continue;
if (offset + @sizeOf(protocol.ChildEntry) > reply.len) break;
const entry = protocol.ChildEntry{ .parent = child.parent, .bus_address = child.bus_address, .identity = child.identity };
@memcpy(reply[offset..][0..@sizeOf(protocol.ChildEntry)], std.mem.asBytes(&entry));
offset += @sizeOf(protocol.ChildEntry);
count += 1;
}
const header = protocol.EnumerateReply{ .status = 0, .count = count };
@memcpy(reply[0..@sizeOf(protocol.EnumerateReply)], std.mem.asBytes(&header));
return offset;
}
/// An application subscribed: its endpoint arrived as the call's capability.
fn onSubscribe(reply: []u8, capability: ?runtime.ipc.Handle) usize {
var status: i32 = -1;
if (capability) |handle| {
for (&subscribers) |*slot| {
if (slot.* == null) {
slot.* = handle;
status = 0;
break;
}
}
}
const report_reply = protocol.ReportReply{ .status = status };
@memcpy(reply[0..@sizeOf(protocol.ReportReply)], std.mem.asBytes(&report_reply));
return @sizeOf(protocol.ReportReply);
}
fn onNotification(badge: u64) void {
if (badge & runtime.ipc.notify_exit_bit != 0) {
const dead: u32 = @intCast(badge & ~(runtime.ipc.notify_badge_bit | runtime.ipc.notify_exit_bit));
if (driverByProcess(dead)) |driver| onDriverExit(driver);
return;
}
_ = runtime.system.write("device-manager: ok\n");
while (true) runtime.system.sleep(1000);
if (badge & runtime.ipc.notify_timer_bit != 0) sweepDeadlines();
}
pub fn main(init: runtime.process.Init) void {
if (init.arguments.get(1)) |mode| {
test_restart_mode = std.mem.eql(u8, mode, "test-restart");
test_usb_restart_mode = std.mem.eql(u8, mode, "test-usb-restart");
test_pci_restart_mode = std.mem.eql(u8, mode, "test-pci-restart");
}
runtime.service.run(protocol.message_maximum, .{
.service = .device_manager,
.init = initialise,
.on_message = onMessage,
.on_notification = onNotification,
});
}
pub const panic = runtime.panic;
+35
View File
@@ -0,0 +1,35 @@
//! /system/services/fdt — the devicetree discovery service: the ARM twin of the
//! acpi service (docs/m19-m20-plan.md decision 7). **Placeholder: not
//! implemented.** It exists so the build's `-Ddiscovery` option has both of its
//! values from day one; the implementation lands with the Raspberry Pi
//! bring-up (docs/arm.md).
//!
//! What it becomes: the per-firmware discoverer for boots that hand over a
//! flattened device tree instead of ACPI tables. It claims the
//! `devicetree-blob` node the kernel publishes (the FDT the loader received),
//! walks the tree — pure data, no bytecode, so unlike the acpi service it
//! needs no port grant and no interpreter — and, like any bus-shaped driver:
//! `device_register`s what it finds (containment against the blob node's
//! recorded apertures), reports each child to the device manager
//! (`child_added`, identity = the node's `compatible` string), and stays
//! resident under the manager's supervision (hello, restart, the usual
//! contract).
//!
//! Known prerequisite recorded in the plan: `DeviceDescriptor`'s 8-byte `hid`
//! cannot hold an FDT `compatible` string ("brcm,bcm2835-aux-uart") — identity
//! widens before this file grows a body.
const runtime = @import("runtime");
pub fn main(init: runtime.process.Init) void {
_ = init;
// Not implemented: exit cleanly and silently (a bare spawn by the
// initial-ramdisk sweep must not derange other tests' markers). The
// supervisor reads a clean exit as "meant to stop" — correct for a
// placeholder.
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
+98 -11
View File
@@ -6,18 +6,30 @@
//!
//! It proves the C-convention heap works, then — as PID 1 — acts as the system's
//! **service supervisor**: it spawns the user-space services danos brings up at boot
//! (the VFS server, the device manager), and settles into a heartbeat so it stays
//! alive as the root of user space. Drivers are *not* its job: the device manager
//! discovers the hardware and spawns those. This is the service half of the
//! service/driver spawn split (docs/driver-model.md).
//! (the VFS server, the device manager), and settles into an event loop as the root
//! of user space. Drivers are *not* its job: the device manager discovers the
//! hardware and spawns those. This is the service half of the service/driver spawn
//! split (docs/driver-model.md).
//!
//! M21: init also owns **orderly shutdown**. It supervises its children (keeping
//! their ids and an exit endpoint), subscribes to the power service, and on a
//! power-button event runs the stop sequence over its children in reverse order
//! before asking the power service to enter S5 — lifecycle (M17) and events (M21)
//! composing into a clean poweroff.
const std = @import("std");
const runtime = @import("runtime");
const power = runtime.power_protocol;
/// The system services init brings up at boot, in order. This is init's policy — the
/// microkernel keeps such choices in user space, not the kernel. Drivers are absent
/// on purpose: the device manager owns those. (A future init reads this from a
/// manifest under /system/services instead of a hardcoded list.)
const boot_services = [_][]const u8{ "vfs", "device-manager" };
const boot_services = [_][]const u8{ "vfs", "input", "device-manager" };
var children: [boot_services.len]u32 = .{0} ** boot_services.len;
var child_count: usize = 0;
var supervision_endpoint: runtime.ipc.Handle = 0;
pub fn main() void {
// Prove the heap end to end: allocate through the runtime allocator (which
@@ -34,19 +46,94 @@ pub fn main() void {
gpa.free(buffer);
} else |_| {}
// Bring up the boot services. Best-effort and silent: each service announces its
// own readiness (`vfs: ready`, ...), and in an isolation test that runs init with
// no initial-ramdisk the spawns simply no-op rather than deranging the heartbeat.
// One endpoint carries everything init waits on: children's exit
// notifications (they are spawned supervised against it), init's own
// signals, and power events it subscribes to. All arrive in the loop below.
supervision_endpoint = runtime.ipc.createIpcEndpoint() orelse {
_ = runtime.system.write("init: no endpoint\n");
return;
};
_ = runtime.process.bindSignals(supervision_endpoint);
// Bring up the boot services, supervised so init can stop them cleanly.
// Best-effort and silent: each service announces its own readiness, and in
// an isolation test with no initial-ramdisk the spawns simply no-op.
for (boot_services) |service| {
_ = runtime.system.spawn(service);
if (runtime.system.spawnSupervised(service, &.{}, supervision_endpoint)) |id| {
children[child_count] = id;
child_count += 1;
}
}
// Subscribe to power events (retry: the power service registers well after
// init starts). Best-effort — without it, a `terminate` signal still
// triggers the same shutdown path.
subscribePower();
// A re-arming timer drives the liveness heartbeat: proof PID 1 is alive
// (the init test's marker) while the loop stays free to receive signals,
// power events, and children's exit notifications.
_ = runtime.system.timerOnce(supervision_endpoint, 1000);
var receive: [power.message_maximum]u8 = undefined;
while (true) {
_ = runtime.system.write("init: heartbeat\n");
runtime.system.sleep(1000);
const got = runtime.ipc.replyWait(supervision_endpoint, &.{}, &receive, null);
if (runtime.process.signalsFrom(got.badge)) |signals| {
if (signals.has(.terminate)) shutDown();
continue;
}
if (got.isTimer()) {
_ = runtime.system.write("init: heartbeat\n");
_ = runtime.system.timerOnce(supervision_endpoint, 1000);
continue;
}
if (got.isMessage() and got.len >= 2 and receive[0] == @intFromEnum(power.Operation.event)) {
// A power event (the only buffered messages init receives).
if (receive[1] == @intFromEnum(power.Event.power_button)) shutDown();
continue;
}
// Child-exit notifications and anything else: keep waiting.
if (got.isNotification()) continue;
}
}
/// Look up the power service and subscribe our endpoint (handed over as the
/// call's capability) so events arrive as buffered messages here.
fn subscribePower() void {
var handle: ?runtime.ipc.Handle = null;
var tries: u32 = 0;
while (handle == null and tries < 200) : (tries += 1) {
handle = runtime.ipc.lookup(.power);
if (handle == null) runtime.system.sleep(20);
}
// A missing power service is not fatal — init proceeds to its heartbeat and
// a `terminate` signal still drives shutdown. Silent so the no-ramdisk init
// test's heartbeat marker is the next line written.
const h = handle orelse return;
const request = power.Subscribe{};
var reply: [power.message_maximum]u8 = undefined;
_ = runtime.ipc.callCap(h, std.mem.asBytes(&request), &reply, supervision_endpoint) catch {};
}
/// The stop sequence: terminate each child in reverse spawn order (vfs last —
/// other services may flush through it), waiting up to a deadline for each to
/// exit before killing it, then ask the power service to enter S5.
fn shutDown() void {
_ = runtime.system.write("init: shutting down\n");
var i = child_count;
while (i > 0) {
i -= 1;
if (children[i] != 0) runtime.process.stop(children[i], 2000, supervision_endpoint);
}
if (runtime.ipc.lookup(.power)) |h| {
const request = power.Shutdown{};
var reply: [power.message_maximum]u8 = undefined;
_ = runtime.ipc.call(h, std.mem.asBytes(&request), &reply) catch {};
}
// If S5 did not take, init has nothing left to do but idle.
while (true) runtime.system.sleep(1000);
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
@@ -0,0 +1,40 @@
//! system/services/input-source — a hardware-free synthetic input source, used to exercise
//! the input service end to end without a real PS/2 controller (the `input` test case, and
//! any bring-up where there is no hardware). It stands in for a driver: it connects to the
//! input service and publishes a rolling stream that cycles through all device classes —
//! keyboard, mouse, and joystick/gamepad — which the service routes to interested
//! subscribers.
//!
//! It stays silent after startup (no per-event logging) so it can share the boot serial
//! transcript with a subscriber whose output is the test's success marker. The real
//! keyboard and mouse drivers publish their own synthetic streams today; swapping in
//! decoded hardware is a follow-up (see docs/input.md).
const runtime = @import("runtime");
const input = runtime.input;
const system = runtime.system;
pub fn main() void {
var source = input.connectSource() orelse {
_ = system.write("input-source: input service unavailable\n");
return;
};
_ = system.write("input-source: publishing synthetic input events\n");
var step: usize = 0;
while (true) : (step +%= 1) {
// Rotate across the device classes so every publish path (and the service's
// per-device routing) is exercised.
switch (step % 3) {
0 => _ = source.publishKeyboardEvent(input.syntheticKeyEvent(step)),
1 => _ = source.publishMouseEvent(input.syntheticMouseEvent(step)),
else => _ = source.publishJoystickEvent(input.syntheticJoystickEvent(step)),
}
system.sleep(200);
}
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
+56
View File
@@ -0,0 +1,56 @@
//! system/services/input-test — the input service's client and test oracle, the input
//! counterpart of vfs-test. It subscribes to *all* device classes and loops receiving the
//! events a source broadcasts, logging each with its class. It emits the success marker
//! `"input-test: ok"` only **after it has received at least one of each class** (keyboard,
//! mouse, and joystick), then heartbeats it. So the in-kernel `input` test case seeing that
//! marker proves not just that IPC delivery works but that the service *routed* all three
//! device classes to one subscription — source -> service -> subscriber, per device.
const std = @import("std");
const runtime = @import("runtime");
const input = runtime.input;
const system = runtime.system;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
pub fn main() void {
var listener = input.subscribeAll() orelse {
_ = system.write("input-test: could not subscribe\n");
return;
};
_ = system.write("input-test: subscribed\n");
var seen_keyboard = false;
var seen_mouse = false;
var seen_joystick = false;
while (true) {
const event = listener.next() orelse continue;
// Decode the class-specific payload from the tagged envelope and note the class.
if (event.asKeyboard()) |key| {
seen_keyboard = true;
writeLine("input-test: got keyboard code={d} char={d}\n", .{ key.keycode, key.character });
} else if (event.asMouse()) |mouse| {
seen_mouse = true;
writeLine("input-test: got mouse dx={d} dy={d} buttons={d}\n", .{ mouse.dx, mouse.dy, mouse.buttons });
} else if (event.asJoystick()) |joystick| {
seen_joystick = true;
writeLine("input-test: got joystick control={d} value={d}\n", .{ joystick.control, joystick.value });
} else {
writeLine("input-test: got device={d}\n", .{event.device});
}
// The success marker: only once every class has been routed here does this appear,
// and then it heartbeats. Seeing "input-test: ok" proves per-device fan-out works.
if (seen_keyboard and seen_mouse and seen_joystick) {
_ = system.write("input-test: ok all classes received (keyboard, mouse, joystick)\n");
}
}
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
+145
View File
@@ -0,0 +1,145 @@
//! system/services/input — the user-space input service. Shipped in the initial_ramdisk,
//! spawned as a ring-3 process, and published under the well-known `input` service id. It
//! is the fan-out point between **sources** (keyboard, mouse, and joystick/gamepad drivers)
//! and **subscribers** (any program that wants input): a source `publish`es an
//! `InputEvent`, and the service pushes it to every subscriber whose interest mask includes
//! that event's device class (keyboard / mouse / joystick).
//!
//! The delivery discipline is the whole design (see docs/input.md). The kernel's IPC is a
//! synchronous rendezvous: a server holds one pending reply, so it cannot park N
//! subscribers blocked in a "wait for next event" call. Broadcasting therefore has to be
//! *push* — the service delivering to subscribers. But a synchronous push (`ipc_call`)
//! would let one dead or wedged subscriber hang the whole broadcast, since the kernel
//! never wakes a sender parked on a dead peer's endpoint. So delivery uses the
//! asynchronous `ipc.send`: it posts the event to each subscriber's endpoint queue and
//! returns at once, and can never block on a subscriber. That primitive exists for exactly
//! this ([ipc.md](../../../docs/ipc.md), "asynchronous / buffered send").
//!
//! A subscriber registers by handing the service its own endpoint as a capability (M13
//! capability passing — this service is its first real consumer). The service keeps that
//! handle and `ipc.send`s each event to it.
const std = @import("std");
const runtime = @import("runtime");
const protocol = runtime.input_protocol;
const ipc = runtime.ipc;
const system = runtime.system;
/// One registered subscriber: the endpoint we push events to (a capability it handed us at
/// subscribe time) and the task id that owns it (the subscribe call's badge), so a slot
/// left behind by a subscriber that exited can be reclaimed.
const Subscriber = struct {
used: bool = false,
endpoint: ipc.Handle = 0,
task_id: u32 = 0,
/// Which device classes this subscriber wants (an OR of protocol.device_*). An event
/// is delivered only if its device's bit is set here.
device_mask: u32 = 0,
};
var subscribers = [_]Subscriber{.{}} ** 8;
/// Drop any subscriber whose owning process is no longer alive, so its slot (and the
/// endpoint reference it holds) can be reused. Cheap and only run on subscribe — the async
/// `send` to a dead subscriber's orphaned endpoint is harmless (it just fills a queue no
/// one drains), so this is housekeeping, not correctness.
fn pruneDeadSubscribers() void {
var table: [32]system.ProcessDescriptor = undefined;
const total = system.processes(&table);
const count = @min(total, table.len);
for (&subscribers) |*sub| {
if (!sub.used) continue;
var alive = false;
for (table[0..count]) |descriptor| {
if (descriptor.id == sub.task_id) {
alive = true;
break;
}
}
if (!alive) sub.* = .{};
}
}
/// Register `endpoint` (owned by task `task_id`) to receive the device classes in
/// `device_mask`. Returns false if the subscriber table is full.
fn addSubscriber(endpoint: ipc.Handle, task_id: u32, device_mask: u32) bool {
for (&subscribers) |*sub| {
if (!sub.used) {
sub.* = .{ .used = true, .endpoint = endpoint, .task_id = task_id, .device_mask = device_mask };
return true;
}
}
return false;
}
/// Push `event` to every subscriber whose interest mask includes its device class.
/// `ipc.send` never blocks, so a slow or dead subscriber cannot stall delivery to others.
fn broadcast(event: protocol.InputEvent) void {
const bytes = std.mem.asBytes(&event);
const bit = protocol.deviceBit(event.device);
for (&subscribers) |*sub| {
if (sub.used and sub.device_mask & bit != 0) _ = ipc.send(sub.endpoint, bytes);
}
}
/// Handle one request. `got` carries the sender badge (a task id) and, for subscribe, the
/// subscriber's endpoint capability in `got.cap`. Writes a `Reply` into `out` and returns
/// its length.
fn handle(message: []const u8, got: ipc.Received, out: []u8) usize {
const reply = struct {
fn write(buffer: []u8, status: i32) usize {
const header = protocol.Reply{ .status = status };
@memcpy(buffer[0..protocol.reply_size], std.mem.asBytes(&header));
return protocol.reply_size;
}
};
if (message.len < protocol.request_size) return reply.write(out, -1);
const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]);
switch (@as(protocol.Operation, @enumFromInt(request.operation))) {
.subscribe => {
const endpoint = got.cap orelse return reply.write(out, -1); // no endpoint passed
// A zero mask means "everything" (a subscriber that named no class still wants input).
const mask = if (request.device_mask == 0) protocol.device_all else request.device_mask;
pruneDeadSubscribers();
if (!addSubscriber(endpoint, @intCast(got.badge), mask)) return reply.write(out, -1); // table full
return reply.write(out, 0);
},
.publish => {
broadcast(request.event);
return reply.write(out, 0);
},
}
}
pub fn main() void {
const endpoint = ipc.createIpcEndpoint() orelse {
_ = system.write("input: no endpoint\n");
return;
};
if (!ipc.register(.input, endpoint)) {
_ = system.write("input: register failed\n");
return;
}
_ = system.write("input: ready\n");
var reply_buffer: [protocol.reply_size]u8 = undefined;
var reply_len: usize = 0;
var receive: [protocol.request_size]u8 = undefined;
while (true) {
const got = ipc.replyWait(endpoint, reply_buffer[0..reply_len], &receive, null);
// Only synchronous client requests (subscribe/publish) arrive here; nothing sends
// this service asynchronous messages, so a notification wake would be spurious.
if (got.isNotification()) {
reply_len = 0;
continue;
}
reply_len = handle(receive[0..got.len], got, &reply_buffer);
}
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
+319
View File
@@ -0,0 +1,319 @@
//! The input wire protocol — the message format spoken between the user-space input
//! service ([input.zig](input.zig)) and the two kinds of process that reach it: a
//! **source** (a keyboard, mouse, or joystick/gamepad driver) that publishes events, and a
//! **subscriber** (any program) that subscribes and is then pushed each event.
//!
//! The service handles several device classes over one endpoint. Each class has its own
//! typed event (`KeyEvent`, `MouseEvent`, `JoystickEvent`); they all travel in a common
//! `InputEvent` envelope tagged with a `DeviceKind`, so the fan-out path is one code path
//! and a subscriber can take a mix of devices on a single stream. A subscriber declares
//! which classes it wants with a `device_mask`, and the service routes accordingly.
//!
//! Two message shapes ride over the endpoint, tagged by `Operation`, like the
//! [VFS protocol](../vfs/protocol.zig):
//!
//! - **subscribe / publish**: a synchronous `ipc_call` carrying a `Request`. `subscribe`
//! hands the service the subscriber's own endpoint as a capability (`send_cap`) and a
//! `device_mask`; `publish` carries an `InputEvent`. The reply is a `Reply`.
//! - **delivery**: the service pushes each `InputEvent` to every interested subscriber with
//! the asynchronous `ipc_send` — no reply owed, and a dead subscriber can never stall the
//! broadcast. Received in the subscriber's buffer with `Received.isMessage()` set.
//!
//! This is a danos-native contract, shared by the input service, the `runtime.input`
//! client helpers, and every source/subscriber. Everything fits one IPC message.
const std = @import("std");
/// The classes of input device the service fans out. Each names a typed event and a bit in
/// the subscription mask.
pub const DeviceKind = enum(u32) {
keyboard = 0,
mouse = 1,
joystick = 2, // joysticks and gamepads/controllers
};
/// Subscription-interest bits (`Request.device_mask`) — which device classes a subscriber
/// wants. OR them together, or use `device_all`.
pub const device_keyboard: u32 = 1 << 0;
pub const device_mouse: u32 = 1 << 1;
pub const device_joystick: u32 = 1 << 2;
pub const device_all: u32 = device_keyboard | device_mouse | device_joystick;
/// The subscription bit for a `DeviceKind` value (as it appears in `InputEvent.device`).
/// An unknown device maps to 0, so it matches no subscriber.
pub fn deviceBit(device: u32) u32 {
return switch (device) {
@intFromEnum(DeviceKind.keyboard) => device_keyboard,
@intFromEnum(DeviceKind.mouse) => device_mouse,
@intFromEnum(DeviceKind.joystick) => device_joystick,
else => 0,
};
}
// --- keyboard ---------------------------------------------------------------
/// What happened to a key. `key_down`/`key_up` are the physical make/break; `key_press`
/// is the higher-level "a character was produced", carrying it in `KeyEvent.character`.
pub const EventKind = enum(u32) {
key_down = 0,
key_up = 1,
key_press = 2,
};
/// One keyboard event. `keycode` names the physical key (layout-independent); `character`
/// is the Unicode scalar for `key_press` (else 0); `modifiers` is an OR of `modifier_*`.
pub const KeyEvent = extern struct {
kind: u32, // an EventKind
keycode: u32, // a Keycode
character: u32, // Unicode scalar for key_press, else 0
modifiers: u32, // OR of modifier_*
};
pub const modifier_shift: u32 = 1 << 0;
pub const modifier_control: u32 = 1 << 1;
pub const modifier_alt: u32 = 1 << 2;
/// The danos-native keycode namespace: USB HID keyboard-page usages (page 0x07), the
/// numbering the PS/2 scancode decoder emits and the xkeyboard-config layout tables are
/// indexed by. Non-exhaustive, so an unnamed usage still travels as a valid value.
pub const Keycode = enum(u32) {
unknown = 0,
// letters
a = 4,
b = 5,
c = 6,
d = 7,
e = 8,
f = 9,
g = 10,
h = 11,
i = 12,
j = 13,
k = 14,
l = 15,
m = 16,
n = 17,
o = 18,
p = 19,
q = 20,
r = 21,
s = 22,
t = 23,
u = 24,
v = 25,
w = 26,
x = 27,
y = 28,
z = 29,
// digit row
one = 30,
two = 31,
three = 32,
four = 33,
five = 34,
six = 35,
seven = 36,
eight = 37,
nine = 38,
zero = 39,
// control and whitespace
enter = 40,
escape = 41,
backspace = 42,
tab = 43,
spacebar = 44,
// punctuation
minus = 45,
equal = 46,
left_bracket = 47,
right_bracket = 48,
backslash = 49,
non_us_hash = 50,
semicolon = 51,
apostrophe = 52,
grave = 53,
comma = 54,
period = 55,
slash = 56,
caps_lock = 57,
// function row
f1 = 58,
f2 = 59,
f3 = 60,
f4 = 61,
f5 = 62,
f6 = 63,
f7 = 64,
f8 = 65,
f9 = 66,
f10 = 67,
f11 = 68,
f12 = 69,
print_screen = 70,
scroll_lock = 71,
pause = 72,
// navigation
insert = 73,
home = 74,
page_up = 75,
delete = 76,
end = 77,
page_down = 78,
right_arrow = 79,
left_arrow = 80,
down_arrow = 81,
up_arrow = 82,
// keypad
num_lock = 83,
keypad_slash = 84,
keypad_asterisk = 85,
keypad_minus = 86,
keypad_plus = 87,
keypad_enter = 88,
keypad_one = 89,
keypad_two = 90,
keypad_three = 91,
keypad_four = 92,
keypad_five = 93,
keypad_six = 94,
keypad_seven = 95,
keypad_eight = 96,
keypad_nine = 97,
keypad_zero = 98,
keypad_period = 99,
non_us_backslash = 100,
application = 101,
// modifiers
left_control = 224,
left_shift = 225,
left_alt = 226,
left_gui = 227,
right_control = 228,
right_shift = 229,
right_alt = 230,
right_gui = 231,
_,
};
// --- mouse ------------------------------------------------------------------
/// What a mouse event reports. `motion` carries relative `dx`/`dy`; `button_down`/`up`
/// name a button in `button`; `scroll` carries `scroll_x`/`scroll_y`.
pub const MouseEventKind = enum(u32) {
motion = 0,
button_down = 1,
button_up = 2,
scroll = 3,
};
pub const mouse_button_left: u32 = 1 << 0;
pub const mouse_button_right: u32 = 1 << 1;
pub const mouse_button_middle: u32 = 1 << 2;
/// One mouse event. Relative motion (`dx`/`dy`) and wheel (`scroll_*`) are signed;
/// `buttons` is the current pressed-button bitmask (`mouse_button_*`).
pub const MouseEvent = extern struct {
kind: u32, // a MouseEventKind
button: u32, // the mouse_button_* bit for button_down/up, else 0
dx: i32, // relative X motion (.motion)
dy: i32, // relative Y motion (.motion)
scroll_x: i32, // horizontal wheel (.scroll)
scroll_y: i32, // vertical wheel (.scroll)
buttons: u32, // current pressed-button bitmask
};
// --- joystick / gamepad -----------------------------------------------------
/// What a joystick/gamepad event reports. `axis` carries a signed `value` on axis
/// `control`; `button_down`/`up` name a button index in `control`.
pub const JoystickEventKind = enum(u32) {
axis = 0,
button_down = 1,
button_up = 2,
};
/// One joystick/gamepad event. `control` is the axis index (`.axis`) or button index
/// (button events); `value` is the axis position (signed, e.g. -32768..32767) for `.axis`;
/// `buttons` is the current pressed-button bitmask.
pub const JoystickEvent = extern struct {
kind: u32, // a JoystickEventKind
control: u32, // axis index (.axis) or button index (button events)
value: i32, // axis value for .axis, else 0
buttons: u32, // current pressed-button bitmask
};
// --- the common envelope ----------------------------------------------------
/// The largest per-device event, so `InputEvent` can hold any of them inline.
pub const max_event_size: usize = @max(@sizeOf(KeyEvent), @max(@sizeOf(MouseEvent), @sizeOf(JoystickEvent)));
/// The tagged envelope broadcast to subscribers: a `DeviceKind` plus the raw bytes of the
/// matching per-device event. Decode it with `asKeyboard`/`asMouse`/`asJoystick` (each
/// returns null unless `device` matches), or build one with the `from*` constructors.
pub const InputEvent = extern struct {
device: u32, // a DeviceKind
_padding: u32 = 0,
data: [max_event_size]u8 = [_]u8{0} ** max_event_size,
pub fn asKeyboard(self: InputEvent) ?KeyEvent {
if (self.device != @intFromEnum(DeviceKind.keyboard)) return null;
return std.mem.bytesToValue(KeyEvent, self.data[0..@sizeOf(KeyEvent)]);
}
pub fn asMouse(self: InputEvent) ?MouseEvent {
if (self.device != @intFromEnum(DeviceKind.mouse)) return null;
return std.mem.bytesToValue(MouseEvent, self.data[0..@sizeOf(MouseEvent)]);
}
pub fn asJoystick(self: InputEvent) ?JoystickEvent {
if (self.device != @intFromEnum(DeviceKind.joystick)) return null;
return std.mem.bytesToValue(JoystickEvent, self.data[0..@sizeOf(JoystickEvent)]);
}
pub fn fromKeyboard(event: KeyEvent) InputEvent {
return pack(.keyboard, std.mem.asBytes(&event));
}
pub fn fromMouse(event: MouseEvent) InputEvent {
return pack(.mouse, std.mem.asBytes(&event));
}
pub fn fromJoystick(event: JoystickEvent) InputEvent {
return pack(.joystick, std.mem.asBytes(&event));
}
fn pack(device: DeviceKind, bytes: []const u8) InputEvent {
var self = InputEvent{ .device = @intFromEnum(device) };
@memcpy(self.data[0..bytes.len], bytes);
return self;
}
};
// --- request / reply --------------------------------------------------------
/// Which side of a request this is.
pub const Operation = enum(u32) {
subscribe = 0, // register the caller's endpoint (send_cap) for the classes in device_mask
publish = 1, // a source submits `event` to broadcast to interested subscribers
};
/// Request header. For `subscribe`, `device_mask` is the OR of `device_*` bits the caller
/// wants (0 means all) and the caller's receive endpoint travels as the call's capability;
/// `event` is ignored. For `publish`, `event` is the event to broadcast.
pub const Request = extern struct {
operation: u32, // an Operation
device_mask: u32 = 0, // subscribe: interested device classes (0 => all)
event: InputEvent = .{ .device = 0 },
};
/// Reply header. `status` is 0 on success or a negative errno.
pub const Reply = extern struct {
status: i32,
_padding: u32 = 0,
};
pub const request_size: usize = @sizeOf(Request);
pub const reply_size: usize = @sizeOf(Reply);
pub const event_size: usize = @sizeOf(InputEvent);
comptime {
// The delivery path posts a bare InputEvent through ipc_send, so it must fit an
// endpoint's async payload slot (POST_MAXIMUM is 64).
if (event_size > 64) @compileError("InputEvent must fit the ipc_send payload (POST_MAXIMUM)");
}
+68
View File
@@ -0,0 +1,68 @@
//! The power protocol (docs/m21-plan.md): system power's domain-named surface,
//! registered under `ServiceId.power`. On x86 the acpi service serves it; on
//! ARM a PSCI/mailbox service will register the same id — subscribers never
//! learn which firmware they are on (m19-m20-plan.md decision 7). The
//! vfs-protocol pattern: extern-struct messages, a version, reserved fields.
/// The protocol version a client states nowhere yet — reserved for the day a
/// handshake needs it; requests carry it so a mismatch can be refused loudly.
pub const version: u16 = 1;
pub const Operation = enum(u8) {
/// Subscribe to power events: the subscriber's endpoint rides as the
/// call's capability (the input/device-manager pattern); events arrive on
/// it as buffered messages carrying an `EventMessage`.
subscribe = 1,
/// Orderly shutdown's last step: enter S5. Accepted only from PID 1
/// (init) — the process that has already run the stop sequence over
/// everything else.
shutdown = 2,
/// The published event payload (never sent *to* the service).
event = 3,
};
/// What happened. The vocabulary is hardware-neutral: a lid is a lid whether
/// ACPI or a PSCI mailbox reported it.
pub const Event = enum(u8) {
power_button = 1,
lid = 2,
ac = 3,
battery = 4,
/// A device notification that maps to none of the named events — the
/// `code` and `hid` fields say which device and what code.
notify = 5,
};
pub const Subscribe = extern struct {
operation: u8 = @intFromEnum(Operation.subscribe),
reserved0: u8 = 0,
version: u16 = version,
reserved1: u32 = 0,
};
pub const Shutdown = extern struct {
operation: u8 = @intFromEnum(Operation.shutdown),
reserved0: u8 = 0,
version: u16 = version,
reserved1: u32 = 0,
};
/// A published event, as the buffered-message payload subscribers receive.
pub const EventMessage = extern struct {
operation: u8 = @intFromEnum(Operation.event),
/// An Event value.
event: u8,
reserved0: u16 = 0,
/// The device notification code (Notify's second argument), or 0.
code: u32 = 0,
/// The notifying device's hardware id (EISA-decoded), or all zero.
hid: [8]u8 = .{0} ** 8,
};
pub const Reply = extern struct {
status: i32,
reserved: u32 = 0,
};
/// Upper bound on any message in this protocol — sizes endpoint buffers.
pub const message_maximum = 64;
@@ -0,0 +1,185 @@
//! process-test — a test fixture for process management (bundled in the
//! initial-ramdisk, driven by the `supervision` test case). One binary, three
//! roles picked by argv, so the whole user-side surface is exercised end to end:
//!
//! - `process-test run` — the supervisor: spawns the two children below with an
//! exit-notification endpoint, sees them in `process_enumerate`, kills them,
//! receives both exit notifications, and confirms they are gone. Prints
//! "process-test: ok" for the kernel test to match, or a FAIL line naming the
//! step that broke.
//! - `process-test sleeper` — a child that blocks in `sleep` forever: its kill
//! exercises the immediate reap of a blocked task.
//! - `process-test spinner` — a child that spins in user mode making no system
//! calls: its kill exercises the deferred path (kill_pending, finished by the
//! timer tick).
//!
//! Spawned with no arguments (the initial-ramdisk sweep test starts every bundled
//! binary bare), it exits silently so it cannot derange other tests' output.
const std = @import("std");
const runtime = @import("runtime");
fn fail(step: []const u8) noreturn {
_ = runtime.system.write("process-test: FAIL ");
_ = runtime.system.write(step);
_ = runtime.system.write("\n");
runtime.system.exit(1);
}
/// Whether process `id` appears in a fresh `process_enumerate` snapshot, named
/// `name` (an id present under the wrong name is a table mix-up, not a pass).
fn listed(id: u32, name: []const u8) bool {
var table: [32]runtime.system.ProcessDescriptor = undefined;
const total = runtime.system.processes(&table);
for (table[0..@min(total, table.len)]) |descriptor| {
if (descriptor.id != id) continue;
return std.mem.eql(u8, descriptor.name[0..descriptor.name_length], name);
}
return false;
}
/// Block on the exit endpoint until a child-exit notification arrives; returns
/// the ended child's id. A wrong wake-up (there should be none — nothing else
/// knows this endpoint) fails the test rather than looping forever.
fn awaitChildExit(endpoint: runtime.ipc.Handle) u32 {
var scratch: [8]u8 = undefined;
const received = runtime.ipc.replyWait(endpoint, scratch[0..0], &scratch, null);
if (!received.isChildExit()) fail("expected a child-exit notification");
return received.childProcessId();
}
/// The harness-run child of the signals test: echoes requests, logs the two
/// signals it handles. Terminate makes run() return, and returning from main is
/// the clean exit the parent reads as ExitReason.exited.
fn echo(message: []const u8, reply: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = sender;
_ = capability;
const n = @min(message.len, reply.len);
@memcpy(reply[0..n], message[0..n]);
return n;
}
fn onReload() void {
_ = runtime.system.write("process-test: reloaded\n");
}
fn onTerminate() void {
_ = runtime.system.write("process-test: terminating\n");
}
/// The parent of the signals test: drives ping, echo, reload, the one-shot
/// timer, and both endings of the stop sequence (polite -> exited; deaf ->
/// killed at the deadline). Prints "process-test: signals ok" as the marker.
fn signalRun() void {
const endpoint = runtime.ipc.createIpcEndpoint() orelse fail("create exit endpoint");
const child = runtime.system.spawnSupervised("process-test", &.{"service"}, endpoint) orelse fail("spawn service child");
// Reach the child's endpoint through the registry (retry: it may not be up).
var service_handle: ?runtime.ipc.Handle = null;
var tries: u32 = 0;
while (service_handle == null and tries < 200) : (tries += 1) {
service_handle = runtime.ipc.lookup(.input);
if (service_handle == null) runtime.system.sleep(20);
}
const h = service_handle orelse fail("service child never registered");
// The universal ping: a zero-length call answered zero-length by the harness.
var reply: [16]u8 = undefined;
const pong = runtime.ipc.call(h, &.{}, &reply) catch fail("ping call failed");
if (pong != 0) fail("ping reply not empty");
// An ordinary request still reaches on_message.
const n = runtime.ipc.call(h, "echo!", &reply) catch fail("echo call failed");
if (n != 5 or !std.mem.eql(u8, reply[0..5], "echo!")) fail("echo mismatch");
// reload: a statement — the child logs it; the kernel test reads the serial.
if (!runtime.process.sendSignal(child, .reload)) fail("send reload");
runtime.system.sleep(200);
// The one-shot timer: armed on our endpoint, lands as isTimer.
if (!runtime.system.timerOnce(endpoint, 100)) fail("arm timer");
var scratch: [8]u8 = undefined;
const landing = runtime.ipc.replyWait(endpoint, scratch[0..0], &scratch, null);
if (!landing.isTimer()) fail("expected the timer landing");
// The stop sequence, polite path: terminate, clean exit inside the deadline.
runtime.process.stop(child, 2000, endpoint);
if ((runtime.process.exitReason(child) orelse .killed) != .exited) fail("service child reason not exited");
// The deaf child: binds nothing, hears nothing — the deadline kills it.
const deaf = runtime.system.spawnSupervised("process-test", &.{"sleeper"}, endpoint) orelse fail("spawn deaf child");
runtime.system.sleep(50); // let it reach its sleep
runtime.process.stop(deaf, 300, endpoint);
if ((runtime.process.exitReason(deaf) orelse .exited) != .killed) fail("deaf child reason not killed");
_ = runtime.system.write("process-test: signals ok\n");
}
pub fn main(init: runtime.process.Init) void {
const role = init.arguments.get(1) orelse return; // spawned bare (ramdisk sweep): stay silent
if (std.mem.eql(u8, role, "sleeper")) {
while (true) runtime.system.sleep(500);
}
if (std.mem.eql(u8, role, "service")) {
// Borrowed well-known id: the input service is not part of this scenario.
runtime.service.run(64, .{
.service = .input,
.on_message = echo,
.on_reload = onReload,
.on_terminate = onTerminate,
});
return; // terminate arrived; returning is the clean exit
}
if (std.mem.eql(u8, role, "signal-run")) {
signalRun();
return;
}
if (std.mem.eql(u8, role, "spinner")) {
var beat: u64 = 0;
const touch: *volatile u64 = &beat;
while (true) touch.* +%= 1; // user mode only — no system calls to die at
}
// The supervisor ("run").
const endpoint = runtime.ipc.createIpcEndpoint() orelse fail("create exit endpoint");
const sleeper = runtime.system.spawnSupervised("process-test", &.{"sleeper"}, endpoint) orelse fail("spawn sleeper");
const spinner = runtime.system.spawnSupervised("process-test", &.{"spinner"}, endpoint) orelse fail("spawn spinner");
runtime.system.sleep(100); // let the sleeper block and the spinner get a core
if (!listed(sleeper, "process-test")) fail("sleeper not in process_enumerate");
if (!listed(spinner, "process-test")) fail("spinner not in process_enumerate");
// Kills that must be refused: a kernel task (id 0), and an id that was never
// issued — both -ESRCH. (-EPERM needs a second supervisor; the kernel-level
// `process-kill` test covers it.)
if (runtime.system.kill(0)) fail("killing a kernel task was allowed");
if (runtime.system.kill(0xFFFF_FFF0)) fail("killing an unknown id was allowed");
// The blocked child: usually reaped on the spot (it sits in `sleep`). The
// notification is the fence — after it, the child is certainly gone, so the
// second kill must miss (its id is never reused).
if (!runtime.system.kill(sleeper)) fail("kill sleeper");
if (awaitChildExit(endpoint) != sleeper) fail("sleeper exit notification");
if (runtime.system.kill(sleeper)) fail("double kill was allowed");
// The running child: the deferred path — condemned now, dead by the next tick.
if (!runtime.system.kill(spinner)) fail("kill spinner");
if (awaitChildExit(endpoint) != spinner) fail("spinner exit notification");
if (listed(sleeper, "process-test")) fail("sleeper still listed after kill");
if (listed(spinner, "process-test")) fail("spinner still listed after kill");
// M17.2: both children were killed by us, and the reason says so — the whole
// restart-policy input, read through the runtime like a real supervisor would.
if ((runtime.process.exitReason(sleeper) orelse .exited) != .killed) fail("sleeper reason not killed");
if ((runtime.process.exitReason(spinner) orelse .exited) != .killed) fail("spinner reason not killed");
if (runtime.process.exitReason(0xFFFF_FFF0) != null) fail("unknown id had a reason");
_ = runtime.system.write("process-test: ok\n");
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
+2 -2
View File
@@ -8,8 +8,8 @@
//! spellings (`stat`, `O_CREAT`, ...) live only in the POSIX layer
//! (library/posix/unistd.zig), which translates to these.
//!
//! This is user-space only — the kernel knows nothing of files or paths; it only
//! moves the bytes. Shared by library/posix/unistd.zig (client) and system/services/vfs/vfs.zig (server).
//! This is user-space only — the kernel knows nothing of files or paths; it only moves the bytes.
//! Shared by library/posix/unistd.zig (client) and system/services/vfs/vfs.zig (server).
pub const Operation = enum(u32) {
open, // open(path) -> node id
+21 -1
View File
@@ -6,10 +6,30 @@
const std = @import("std");
const runtime = @import("runtime");
pub fn main() void {
pub fn main(init: runtime.process.Init) void {
const u = @import("posix").unistd;
const payload = "hello-vfs";
// The "park" role (the vfs-client-death test): open a file, then hold the
// handle forever without closing — the kill and the VFS's release-on-death
// are the point.
if (init.arguments.count > 1) {
var fd: i32 = -1;
var tries: u32 = 0;
while (fd < 0 and tries < 200) : (tries += 1) {
fd = u.open("parked", u.O_CREAT);
if (fd < 0) runtime.system.sleep(20);
}
if (fd < 0) {
_ = runtime.system.write("vfstest: park open failed\n");
return;
}
while (true) {
_ = runtime.system.write("vfstest: parked\n");
runtime.system.sleep(500);
}
}
// The VFS server may not have registered yet — retry open until it's up.
var fd: i32 = -1;
var tries: u32 = 0;
+53 -18
View File
@@ -23,6 +23,10 @@ const Node = struct {
const OpenFile = struct {
used: bool = false,
node: usize = 0,
// The client (task id — an IPC badge is one) that opened this handle. What
// release-on-death sweeps by: a service must never depend on its clients
// cleaning up after themselves (docs/process-lifecycle.md).
owner: u32 = 0,
};
var nodes = [_]Node{.{}} ** 8;
@@ -65,8 +69,30 @@ fn fail(out: []u8) usize {
return writeReply(out, .{ .status = -1 }, &.{});
}
/// Handle one request; write the reply into `out`, return its length.
fn handle(message: []const u8, out: []u8) usize {
/// Format one whole log line and emit it in a single `debug_write`, so lines from
/// concurrent processes can never land in the middle of it.
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [96]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
/// Release every open handle `client` held — called on that client's published
/// exit event. The nodes (the files) stay: ramfs contents outlive their writers,
/// only the dead client's handles go.
fn releaseClientHandles(client: u32) void {
var released: u32 = 0;
for (&opens) |*o| {
if (o.used and o.owner == client) {
o.used = false;
released += 1;
}
}
if (released != 0) writeLine("vfs: released {d} handle(s) for dead client {d}\n", .{ released, client });
}
/// Handle one request from `sender`; write the reply into `out`, return its length.
fn handle(message: []const u8, out: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = capability;
if (message.len < protocol.request_size) return fail(out);
const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]);
const payload = message[protocol.request_size..];
@@ -77,7 +103,7 @@ fn handle(message: []const u8, out: []u8) usize {
const ni = findNode(name) orelse createNode(name) orelse return fail(out);
for (&opens, 0..) |*o, i| {
if (!o.used) {
o.* = .{ .used = true, .node = ni };
o.* = .{ .used = true, .node = ni, .owner = sender };
return writeReply(out, .{ .status = 0, .node = i }, &.{});
}
}
@@ -113,27 +139,36 @@ fn handle(message: []const u8, out: []u8) usize {
}
}
pub fn main() void {
const endpoint = runtime.ipc.createIpcEndpoint() orelse {
_ = runtime.system.write("vfs: no endpoint\n");
return;
};
if (!runtime.ipc.register(.vfs, endpoint)) {
_ = runtime.system.write("vfs: register failed\n");
return;
/// Startup, under the harness: subscribe to the published exit events — when a
/// client dies holding open handles, the exit notification is how the VFS learns
/// to release them (docs/process-lifecycle.md).
fn initialise(endpoint: runtime.ipc.Handle) bool {
if (!runtime.process.subscribeExits(endpoint)) {
_ = runtime.system.write("vfs: exit subscription failed\n");
}
_ = runtime.system.write("vfs: ready\n");
return true;
}
var reply_buffer: [protocol.message_maximum]u8 = undefined;
var reply_len: usize = 0;
var receive: [protocol.message_maximum]u8 = undefined;
while (true) {
const got = runtime.ipc.replyWait(endpoint, reply_buffer[0..reply_len], &receive, null);
// Ignore notifications (none expected here); handle a request.
reply_len = handle(receive[0..got.len], &reply_buffer);
/// A non-signal notification: the only kind the VFS subscribes to is exit events.
fn onNotification(badge: u64) void {
if (badge & runtime.ipc.notify_exit_bit != 0) {
releaseClientHandles(@intCast(badge & ~(runtime.ipc.notify_badge_bit | runtime.ipc.notify_exit_bit)));
}
}
pub fn main() void {
// The harness owns the loop: requests dispatch to handle(), exit events to
// onNotification(), ping and terminate are answered for free — this service
// gained the whole lifecycle contract by deleting its hand-rolled loop.
runtime.service.run(protocol.message_maximum, .{
.service = .vfs,
.init = initialise,
.on_message = handle,
.on_notification = onNotification,
});
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
+215 -31
View File
@@ -18,9 +18,11 @@ Usage:
"""
import argparse
import json
import os
import re
import shutil
import socket
import subprocess
import sys
import time
@@ -55,20 +57,16 @@ ARCHES = {
"/opt/homebrew/share/qemu/edk2-i386-vars.fd", # macOS Homebrew (Apple Silicon)
"/usr/local/share/qemu/edk2-i386-vars.fd", # macOS Homebrew (Intel)
],
# zig-out is a FHS-shaped image and the boot volume; the harness copies the
# boot-critical files from their FHS paths into a fresh ESP with the same
# layout. (dest in ESP, source path under zig-out) — identical here.
"efi_app": ("EFI/BOOT/BOOTX64.efi", "EFI/BOOT/BOOTX64.efi"),
"kernel": ("system/kernel", "system/kernel"),
# The init user program and the initial-ramdisk (VFS server + drivers).
"extra": [("system/services/init", "system/services/init"),
("boot/initial-ramdisk.img", "boot/initial-ramdisk.img")],
# zig-out is itself the FHS-shaped boot volume (docs/efi.md): the build
# installs BOOTX64.efi, the kernel, init, and the initial-ramdisk at their
# boot paths. The harness presents zig-out to the guest directly — exactly
# as `zig build run-x86-64` does — so there is no separate ESP to assemble.
# Built as a function so we can splice in per-run paths.
"qemu_args": lambda a, esp, vars_fd, serial: [
"qemu_args": lambda a, boot_volume, vars_fd, serial: [
"-machine", "q35", "-m", "128M",
"-drive", f"if=pflash,format=raw,readonly=on,file={a['ovmf_code']}",
"-drive", f"if=pflash,format=raw,file={vars_fd}",
"-drive", f"format=raw,file=fat:rw:{esp}",
"-drive", f"format=raw,file=fat:rw:{boot_volume}",
"-net", "none",
"-vga", "none", "-device", "VGA,edid=on,xres=1280,yres=720",
"-display", "none",
@@ -84,7 +82,10 @@ ARCHES = {
# `expect`: a regex that must appear in serial output => pass.
# `fail`: optional regex whose appearance => immediate fail.
CASES = [
# smoke also proves the QMP channel: the harmless query must be delivered
# (handshake + command) before the case may pass — see run_case.
{"name": "smoke",
"qmp_after": {"delay": 2, "command": "query-status"},
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
{"name": "discovery",
@@ -168,7 +169,7 @@ CASES = [
# Stress the big kernel lock across cores; heavier, so a longer timeout.
{"name": "smp-stress",
"smp": 4,
"timeout": 90,
"timeout": 150,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Retry: a forced first-wake failure must still bring every core online.
@@ -199,6 +200,17 @@ CASES = [
{"name": "user-pf",
"expect": r"page fault \(vector 14\)[\s\S]*error code : 0x5[\s\S]*IP\s*: 0x00007000000000",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Fault recovery: a scheduled ring-3 process that faults is killed — resources
# reclaimed, core kept — while init's heartbeat proves the OS survived.
{"name": "fault-recovery",
"timeout": 60,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Process arguments: argv arrives on the SysV entry stack (argv[0] = the spawned
# name, argv[1..] = the system_spawn argument blob) and echoes back intact.
{"name": "args",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# The real user binary: the bootloader ships /system/services/init off the ESP, the
# kernel loads the ELF and runs it in ring 3, and it writes + exits cleanly.
{"name": "init",
@@ -211,9 +223,148 @@ CASES = [
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# process_enumerate: a task-table snapshot lists spawned processes by name and
# id alongside the kernel tasks, and a too-small buffer still reports the total.
{"name": "process-list",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# process_kill + the exit notification: only the supervisor may kill; a blocked
# victim is reaped in place and a spinning one dies by the deferred (tick) path;
# each death posts one exit badge to the endpoint given at spawn.
{"name": "process-kill",
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# The user-side whole: process-test supervises two children from ring 3 —
# spawn with an exit endpoint, enumerate, kill, notification, gone.
{"name": "supervision",
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M17.1: a dead process's device claims are released by the reap — kill a child
# holding a claim, the device must be claimable again (process-lifecycle.md).
{"name": "claim-release",
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M17.3: published exit events — the VFS subscribes, a client dies holding an
# open handle, and the VFS releases it (process-lifecycle.md "Who learns of a death").
{"name": "vfs-client-death",
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M17.4: signals over IPC — ping, reload, terminate (clean exit), the one-shot
# timer, and the stop sequence's two endings, all driven from ring 3.
{"name": "signals",
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M18.2: bus tree reports — the xHCI driver scans its root-hub ports and
# reports both QEMU devices; the manager mirrors, prunes on the reporter's
# death, and the respawned driver re-reports (docs/device-manager.md).
{"name": "usb-report",
"smp": 4,
"timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"device-manager: child added[\s\S]*"
r"device-manager: child added[\s\S]*"
r"device-manager: test mode: killing the reporter[\s\S]*"
r"device-manager: child removed[\s\S]*"
r"device-manager: restarting usb-xhci-bus[\s\S]*"
r"device-manager: child added",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M20.1: the ring-3 AML parse (the acpi service maps the blobs and parses
# them) finds exactly the Device count the kernel's own parse produced.
{"name": "acpi-parse",
"smp": 4,
"timeout": 60,
"expect": r"acpi-parse: ok",
"fail": r"acpi-parse: mismatch|DANOS-TEST-RESULT: FAIL"},
# M20.3: the flip — ps2-bus now comes up from the acpi service's report, not
# a kernel-built node. Ordered: report -> spawn -> the driver attaches its
# keyboard, proving discovery runs entirely in ring 3 (docs/m19-m20-plan.md).
{"name": "acpi-ps2",
"smp": 4,
"timeout": 150,
"expect": r"acpi: reported PNP0303[\s\S]*"
r"device-manager: spawned ps2-bus[\s\S]*"
r"ps2-bus: keyboard driver attached",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M21.1: the SCI + power button. Boot the manager (which spawns the acpi
# service); ~4s in, QMP system_powerdown raises the ACPI power-button fixed
# event; the service's SCI handler must log the press (docs/m21-plan.md).
{"name": "power-button",
"smp": 4,
"timeout": 60,
"qmp_after": {"delay": 4, "command": "system_powerdown"},
"expect": r"power: button pressed",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M21.3 capstone: orderly shutdown. Boot init (the full tree comes up);
# ~5s in, QMP system_powerdown raises the power button; the acpi service
# publishes it, init stops its children then requests S5, and QEMU exits.
# The ordered regex proves button -> shutting-down -> entering-S5; the case
# passes on QEMU's self-exit through S5 (docs/m21-plan.md).
{"name": "orderly-shutdown",
"smp": 4,
"timeout": 90,
"qmp_after": {"delay": 5, "command": "system_powerdown"},
"expect": r"power: button pressed[\s\S]*"
r"init: shutting down[\s\S]*"
r"power: entering S5",
"fail": r"power: S5 write did not take|DANOS-TEST-RESULT: FAIL"},
# M20.2: the acpi service evaluates _CRS/_STA in ring 3 and registers +
# reports its _HID devices — the two PS/2 nodes must appear with resources
# (keyboard: io 0x60/0x64 + IRQ = 3; mouse: IRQ = 1) (docs/m19-m20-plan.md).
{"name": "acpi-report",
"smp": 4,
"timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"acpi: reported PNP0303 \(device \d+, 3 resources\)[\s\S]*"
r"acpi: reported PNP0F13 \(device \d+, 1 resources\)",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M19.1: the ring-3 PCI scan (pci-bus walks the ECAM through its mmio_map
# grant) finds exactly the functions the kernel's own walk recorded.
{"name": "pci-scan",
"smp": 4,
"timeout": 60,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M18.3: the application surface — device-list enumerates the tree over IPC,
# subscribes (endpoint as capability), and observes the removed/added events
# the reporter's test-kill produces (docs/device-manager.md).
{"name": "device-list",
"smp": 4,
"timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"device-list: \d+ devices[\s\S]*"
r"device-list: subscribed[\s\S]*"
r"device-manager: test mode: killing the reporter[\s\S]*"
r"device-list: removed \(device[\s\S]*"
r"device-list: added \(device",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M18.1: the device manager's hello + restart policy — xHCI hellos clean and
# stays; crash-test faults, is restarted with backoff (re-claiming its device
# each time), and hits the crash-loop cap (docs/device-manager.md).
{"name": "driver-restart",
"smp": 4,
"timeout": 150,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"usb-xhci-bus: hello acknowledged[\s\S]*"
r"device-manager: restarting crash-test[\s\S]*"
r"device-manager: crash-test is failing repeatedly",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# The initial_ramdisk: the loader ferries a bundle of user binaries; the kernel parses
# it and spawns each as a ring-3 process (here the VFS-server stub heartbeats).
{"name": "initial-ramdisk",
"timeout": 60, # the acpi service's boot-time SCI setup can push the marker past 30s under load
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# The user-space VFS: a client opens/writes/reads a file through the rt file
@@ -221,6 +372,12 @@ CASES = [
{"name": "vfs",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# The input service: a synthetic keyboard source publishes events, the service
# broadcasts them over the async ipc_send primitive, and a subscriber (which joined by
# passing its endpoint as a capability) receives them — source -> service -> subscriber.
{"name": "input",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# IO passthrough + IRQ-as-IPC: a user-space HPET driver maps device MMIO into
# its own address space, binds the device's interrupt to an IPC endpoint, and
# is woken by the hardware five times while blocked (never polling).
@@ -271,24 +428,6 @@ def build(arch, case):
return None
def make_esp(arch):
"""Assemble a fresh EFI System Partition from the freshly built binaries."""
esp = os.path.join(WORK, "esp")
if os.path.exists(esp):
shutil.rmtree(esp)
efi_dest, efi_src = arch["efi_app"]
kern_dest, kern_src = arch["kernel"]
fhs = os.path.join(REPO, "zig-out") # zig-out is the FHS image
os.makedirs(os.path.join(esp, os.path.dirname(efi_dest)), exist_ok=True)
os.makedirs(os.path.join(esp, os.path.dirname(kern_dest)), exist_ok=True)
shutil.copy(os.path.join(fhs, efi_src), os.path.join(esp, efi_dest))
shutil.copy(os.path.join(fhs, kern_src), os.path.join(esp, kern_dest))
for dest, src in arch.get("extra", []):
os.makedirs(os.path.join(esp, os.path.dirname(dest)), exist_ok=True)
shutil.copy(os.path.join(fhs, src), os.path.join(esp, dest))
return esp
def resolve_firmware(arch):
"""Collapse the ovmf_code/ovmf_vars candidate lists to the first path that
exists on this machine. Mutates `arch` in place; idempotent (a resolved
@@ -307,12 +446,34 @@ def resolve_firmware(arch):
+ "\nInstall OVMF (edk2-ovmf / ovmf) or add its path above.")
def qmp_send(path, command):
"""One QMP command: connect, capabilities handshake, execute. Raises on any
failure — the caller retries until the guest's socket is ready. This is how
a case injects a host-side event (system_powerdown = the ACPI power button)
into the running guest (docs/m21-plan.md)."""
sock = socket.socket(socket.AF_UNIX, socket.SOCK_STREAM)
sock.settimeout(5)
try:
sock.connect(path)
stream = sock.makefile("rw")
stream.readline() # the QMP greeting
stream.write(json.dumps({"execute": "qmp_capabilities"}) + "\n")
stream.flush()
stream.readline() # {"return": {}}
stream.write(json.dumps({"execute": command}) + "\n")
stream.flush()
stream.readline()
finally:
sock.close()
def run_case(arch, case):
err = build(arch, case["name"])
if err:
return False, "build failed:\n" + err
esp = make_esp(arch)
# zig-out is the FHS boot volume; hand it to the guest as-is (see qemu_args).
boot_volume = os.path.join(REPO, "zig-out")
vars_fd = os.path.join(WORK, "vars.fd")
shutil.copy(arch["ovmf_vars"], vars_fd)
serial = os.path.join(WORK, "serial.log")
@@ -322,17 +483,32 @@ def run_case(arch, case):
expect = re.compile(case["expect"])
fail = re.compile(case["fail"]) if case.get("fail") else None
cmd = [arch["qemu"]] + arch["qemu_args"](arch, esp, vars_fd, serial)
cmd = [arch["qemu"]] + arch["qemu_args"](arch, boot_volume, vars_fd, serial)
if case.get("smp"): # some cases need more than one core (e.g. parallelism)
cmd += ["-smp", str(case["smp"])]
if case.get("qemu_extra"): # extra qemu args, e.g. -device intel-iommu for the IOMMU case
cmd += case["qemu_extra"]
# A QMP control socket, always present (additive): how a case's `qmp_after`
# hook injects host-side events into the guest mid-run.
qmp_path = os.path.join(WORK, "qmp.sock")
if os.path.exists(qmp_path):
os.remove(qmp_path)
cmd += ["-qmp", f"unix:{qmp_path},server,nowait"]
qmp_after = case.get("qmp_after") # {"delay": seconds, "command": "..."}
qmp_sent = False
started = time.monotonic()
qemu = subprocess.Popen(cmd, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
try:
timeout = case.get("timeout", TIMEOUT)
deadline = time.monotonic() + timeout
while time.monotonic() < deadline:
time.sleep(0.2)
if qmp_after and not qmp_sent and time.monotonic() - started >= qmp_after["delay"]:
try:
qmp_send(qmp_path, qmp_after["command"])
qmp_sent = True
except OSError:
pass # socket not up yet; retry next tick
text = ""
if os.path.exists(serial):
with open(serial, "r", errors="replace") as f:
@@ -340,6 +516,8 @@ def run_case(arch, case):
if fail and fail.search(text):
return False, "hit failure marker"
if expect.search(text):
if qmp_after and not qmp_sent:
continue # the hook must deliver before the case may pass
return True, "matched " + repr(case["expect"])
if qemu.poll() is not None: # QEMU exited on its own
if expect.search(text):
@@ -375,6 +553,12 @@ def main():
for case in selected:
print(f" {case['name']:<12} ... ", end="", flush=True)
ok, detail = run_case(arch, case)
if not ok:
# Keep the evidence: serial.log is otherwise overwritten by the
# next case, and an intermittent failure's log is unrecoverable.
source = os.path.join(WORK, "serial.log")
if os.path.exists(source):
shutil.copy(source, os.path.join(WORK, f"{case['name']}-failed-serial.log"))
print(("PASS" if ok else "FAIL") + f" ({detail})")
if not ok:
failures += 1
+468
View File
@@ -0,0 +1,468 @@
#!/usr/bin/env python3
"""Vendor xkeyboard-config and compile it to a native Zig keymap library.
danos does not ship an X11 runtime, but it wants X11's keyboard *data*: the layout
tables (US, UK, German, ...) that turn a physical key + modifiers into a character.
So we compile xkeyboard-config down to Zig at build time, the same way the initial
ramdisk is packed by a host-side Python tool.
Two subcommands:
fetch Download the pinned xkeyboard-config release, verify its sha256, and copy
the transitively-needed data (the symbols files for the configured layouts
plus everything they `include`) into library/xkeyboard-config/vendor/,
alongside a vendored keysymdef.h, the upstream COPYING, and a PROVENANCE.md.
This is the one step that needs the network; run it when bumping the version.
generate Parse the vendored data and emit library/xkeyboard-config/generated/layouts.zig
(deterministic, offline). Run whenever the layout list or emitter changes.
The pipeline per key: our events carry a USB HID usage; HID_TO_NAME maps it to an XKB
key name (<AC01>); the layout's symbols give the keysyms per level; keysymdef.h resolves
each keysym name to its value and Unicode scalar. Level *selection* (which modifier picks
which level) is deliberately left to the Zig side (xkeyboard-config.zig) — we only emit the
per-key type and its up-to-4 keysyms here.
"""
import argparse
import hashlib
import io
import os
import re
import sys
import tarfile
import urllib.request
# --- pinned upstream -------------------------------------------------------
VERSION = "2.44"
TARBALL_URL = (
"https://gitlab.freedesktop.org/xkeyboard-config/xkeyboard-config/-/archive/"
f"xkeyboard-config-{VERSION}/xkeyboard-config-{VERSION}.tar.gz"
)
TARBALL_SHA256 = "35e34edeaf4e8da8d0696ff6b241ee11ddb1b8c6730bac7252d4d0a88ea5f05b"
# keysymdef.h is xorgproto, not xkeyboard-config; fetch copies it from the build host.
KEYSYMDEF_CANDIDATES = [
"/opt/homebrew/include/X11/keysymdef.h",
"/usr/include/X11/keysymdef.h",
"/usr/X11/include/X11/keysymdef.h",
"/usr/X11R6/include/X11/keysymdef.h",
]
# The layouts we generate: (zig name, symbols file, section or None=default).
TARGETS = [
("us", "us", None),
("gb", "gb", None),
("de", "de", None),
("fr", "fr", None),
("es", "es", None),
("dvorak", "us", "dvorak"),
]
# USB HID usage (keyboard page 0x07) -> XKB key name. The physical keys we can turn into
# characters; positions are the ANSI standard shared by HID usages and XKB names.
HID_TO_NAME = {
0x04: "AC01", 0x05: "AB05", 0x06: "AB03", 0x07: "AC03", 0x08: "AD03",
0x09: "AC04", 0x0A: "AC05", 0x0B: "AC06", 0x0C: "AD08", 0x0D: "AC07",
0x0E: "AC08", 0x0F: "AC09", 0x10: "AB07", 0x11: "AB06", 0x12: "AD09",
0x13: "AD10", 0x14: "AD01", 0x15: "AD04", 0x16: "AC02", 0x17: "AD05",
0x18: "AD07", 0x19: "AB04", 0x1A: "AD02", 0x1B: "AB02", 0x1C: "AD06",
0x1D: "AB01",
0x1E: "AE01", 0x1F: "AE02", 0x20: "AE03", 0x21: "AE04", 0x22: "AE05",
0x23: "AE06", 0x24: "AE07", 0x25: "AE08", 0x26: "AE09", 0x27: "AE10",
0x2C: "SPCE",
0x2D: "AE11", 0x2E: "AE12", 0x2F: "AD11", 0x30: "AD12", 0x31: "BKSL",
0x33: "AC10", 0x34: "AC11", 0x35: "TLDE",
0x36: "AB08", 0x37: "AB09", 0x38: "AB10",
0x64: "LSGT", # the extra key on ISO keyboards (102nd key)
}
# XKB type name -> the KeyType enum tag emitted for the Zig side.
TYPE_TO_TAG = {
"ONE_LEVEL": "one_level",
"TWO_LEVEL": "two_level",
"ALPHABETIC": "alphabetic",
"FOUR_LEVEL": "four_level",
"FOUR_LEVEL_ALPHABETIC": "four_level_alphabetic",
"FOUR_LEVEL_SEMIALPHABETIC": "four_level_semialphabetic",
"FOUR_LEVEL_MIXED_KEYPAD": "keypad",
"FOUR_LEVEL_KEYPAD": "keypad",
"KEYPAD": "keypad",
"PC_SYSRQ": "other",
}
REPO = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
LIB = os.path.join(REPO, "library", "xkeyboard-config")
VENDOR = os.path.join(LIB, "vendor")
# --- keysymdef.h: keysym name -> (value, unicode scalar or None) ------------
def parse_keysymdef(path):
table = {}
line_re = re.compile(r"#define\s+XK_(\w+)\s+0x([0-9a-fA-F]+)\s*(?:/\*\s*U\+([0-9A-Fa-f]+)|/\*<U\+([0-9A-Fa-f]+))?")
with open(path, encoding="latin-1") as f:
for line in f:
m = line_re.match(line)
if not m:
continue
name = m.group(1)
value = int(m.group(2), 16)
uni = m.group(3) or m.group(4)
unicode_scalar = int(uni, 16) if uni else None
# First definition wins (keysymdef lists the canonical one first).
table.setdefault(name, (value, unicode_scalar))
return table
def resolve_keysym(name, keysymdef):
"""A keysym token from a symbols file -> (keysym value, unicode scalar or 0)."""
if name in ("NoSymbol", "VoidSymbol", ""):
return (0, 0)
# Unicode escape forms: U00E9 / u00E9.
m = re.fullmatch(r"[Uu]([0-9A-Fa-f]{4,6})", name)
if m:
cp = int(m.group(1), 16)
return (0x01000000 + cp, cp)
# Raw hex keysym value: 0x1000161 (a Unicode keysym) or a plain keysym number.
m = re.fullmatch(r"0x([0-9A-Fa-f]+)", name)
if m:
value = int(m.group(1), 16)
if 0x01000100 <= value <= 0x0110FFFF:
return (value, value - 0x01000000)
if 0x20 <= value <= 0x7E or 0xA0 <= value <= 0xFF:
return (value, value)
return (value, 0)
entry = keysymdef.get(name)
if entry is None:
return (0, 0) # unknown named keysym (dead_*, ISO_*, rare) -> no character
value, unicode_scalar = entry
return (value, unicode_scalar or 0)
# --- symbols parsing --------------------------------------------------------
def strip_comments(text):
text = re.sub(r"/\*.*?\*/", "", text, flags=re.S)
text = re.sub(r"//[^\n]*", "", text)
return text
class KeyDef:
__slots__ = ("levels", "type_name")
def __init__(self, levels, type_name):
self.levels = levels # list of keysym-name strings (group 1)
self.type_name = type_name # explicit XKB type name, or None
class Symbols:
"""Loads and resolves xkb_symbols sections from a directory of symbols files."""
def __init__(self, symbols_dir):
self.dir = symbols_dir
self._files = {} # filename -> {section_name: body, "__default__": name}
self.visited_files = set()
def _load(self, fname):
if fname in self._files:
return self._files[fname]
path = os.path.join(self.dir, fname)
self.visited_files.add(fname)
with open(path, encoding="latin-1") as f:
text = strip_comments(f.read())
sections = {}
default = None
# Find each `[flags] xkb_symbols "NAME" {` and brace-match its body.
for m in re.finditer(r'(\w[\w\s]*?)?\bxkb_symbols\s+"([^"]+)"\s*\{', text):
flags = m.group(1) or ""
name = m.group(2)
body, _ = self._brace_match(text, m.end() - 1)
sections[name] = body
if "default" in flags.split() and default is None:
default = name
if default is None and sections:
default = next(iter(sections))
sections["__default__"] = default
self._files[fname] = sections
return sections
@staticmethod
def _brace_match(text, open_index):
depth = 0
for i in range(open_index, len(text)):
c = text[i]
if c == "{":
depth += 1
elif c == "}":
depth -= 1
if depth == 0:
return text[open_index + 1:i], i
raise ValueError("unbalanced braces")
def resolve(self, fname, section=None, stack=()):
"""Merged {key name -> KeyDef} for (fname, section), following includes."""
sections = self._load(fname)
if section is None:
section = sections["__default__"]
key = (fname, section)
if key in stack:
return {}
body = sections.get(section)
if body is None:
return {}
keys = {}
default_type = None
for stmt in self._statements(body):
kind = stmt[0]
if kind == "include":
augment, inc_file, inc_section = stmt[1], stmt[2], stmt[3]
inc = self.resolve(inc_file, inc_section, stack + (key,))
for n, kd in inc.items():
if augment and n in keys:
continue
keys[n] = kd
elif kind == "default_type":
default_type = stmt[1]
elif kind == "key":
augment, name, kd = stmt[1], stmt[2], stmt[3]
if augment and name in keys:
continue
keys[name] = kd
if default_type is not None:
for kd in keys.values():
if kd.type_name is None:
kd.type_name = default_type
return keys
def _statements(self, body):
"""Yield ('include', augment, file, section) / ('default_type', name) /
('key', augment, name, KeyDef) in source order."""
# Recognise the three statement heads and step through the body in order.
head = re.compile(
r'(?P<inc>(?:(?P<augi>augment|override|replace)\s+)?include\s+"(?P<incarg>[^"]+)")'
r'|(?P<dtype>key\.type(?:\[[^\]]*\])?\s*=\s*"(?P<dtypearg>[^"]+)")'
r'|(?P<key>(?:(?P<augk>augment|override|replace)\s+)?key\s+<(?P<kname>[^>]+)>\s*\{)'
)
i = 0
while i < len(body):
m = head.search(body, i)
if not m:
break
if m.group("inc"):
inc_file, inc_section = self._parse_include(m.group("incarg"))
yield ("include", m.group("augi") == "augment", inc_file, inc_section)
i = m.end()
elif m.group("dtype"):
yield ("default_type", m.group("dtypearg"))
i = m.end()
else: # a key definition; brace-match its body
kbody, end = self._brace_match(body, m.end() - 1)
kd = self._parse_key_body(kbody)
if kd is not None:
yield ("key", m.group("augk") == "augment", m.group("kname"), kd)
i = end + 1
@staticmethod
def _parse_include(arg):
m = re.fullmatch(r"([^()]+)(?:\(([^)]+)\))?", arg.strip())
return (m.group(1), m.group(2))
@staticmethod
def _parse_key_body(kbody):
type_name = None
mt = re.search(r'type(?:\[[^\]]*\])?\s*=\s*"([^"]+)"', kbody)
if mt:
type_name = mt.group(1)
groups = re.findall(r"\[([^\]]*)\]", kbody)
if not groups:
return None
levels = [tok.strip() for tok in groups[0].split(",")]
levels = [t for t in levels if t != ""]
if not levels:
return None
return KeyDef(levels, type_name)
# --- type inference + emission ---------------------------------------------
def is_case_pair(a, b):
return len(a) == 1 and len(b) == 1 and a.isalpha() and a.islower() and b == a.upper()
def key_type_tag(kd):
if kd.type_name is not None:
return TYPE_TO_TAG.get(kd.type_name, "other")
n = len(kd.levels)
if n <= 1:
return "one_level"
if n == 2:
return "alphabetic" if is_case_pair(kd.levels[0], kd.levels[1]) else "two_level"
return "four_level_alphabetic" if is_case_pair(kd.levels[0], kd.levels[1]) else "four_level"
def build_layout(symbols, keysymdef, fname, section):
keys = symbols.resolve(fname, section)
table = [] # 256 entries: (tag, [(keysym, unicode) x4])
for hid in range(256):
name = HID_TO_NAME.get(hid)
kd = keys.get(name) if name else None
if kd is None:
table.append(("one_level", [(0, 0)] * 4))
continue
levels = [resolve_keysym(kd.levels[i], keysymdef) if i < len(kd.levels) else (0, 0)
for i in range(4)]
table.append((key_type_tag(kd), levels))
return table
def emit_zig(layouts):
out = io.StringIO()
out.write("// GENERATED by tools/make-xkeyboard-config.py from xkeyboard-config "
f"{VERSION}. Do not edit.\n")
out.write("// Keyboard layout tables compiled from the X11 xkeyboard-config database\n")
out.write("// (MIT/X11 licensed; see ../vendor/COPYING and ../vendor/PROVENANCE.md).\n\n")
out.write("pub const Level = struct { keysym: u32 = 0, unicode: u21 = 0 };\n\n")
out.write("pub const KeyType = enum {\n")
out.write(" one_level,\n two_level,\n alphabetic,\n four_level,\n"
" four_level_alphabetic,\n four_level_semialphabetic,\n keypad,\n other,\n};\n\n")
out.write("pub const Key = struct { kind: KeyType = .one_level, levels: [4]Level = [_]Level{.{}} ** 4 };\n\n")
out.write("pub const Layout = struct { name: []const u8, keys: [256]Key };\n\n")
names = []
for zig_name, fname, section in TARGETS:
table = layouts[zig_name]
names.append(zig_name)
out.write(f"pub const {zig_name}: Layout = .{{\n")
out.write(f' .name = "{zig_name}",\n')
out.write(" .keys = .{\n")
for hid, (tag, levels) in enumerate(table):
if tag == "one_level" and all(k == 0 and u == 0 for k, u in levels):
out.write(" .{},\n")
continue
parts = ", ".join(f".{{ .keysym = {k}, .unicode = {u} }}" for k, u in levels)
out.write(f" .{{ .kind = .{tag}, .levels = .{{ {parts} }} }},\n")
out.write(" },\n};\n\n")
out.write("pub const all = [_]*const Layout{ " + ", ".join("&" + n for n in names) + " };\n")
return out.getvalue()
# --- fetch ------------------------------------------------------------------
def collect_needed(symbols):
for _, fname, section in TARGETS:
symbols.resolve(fname, section)
return set(symbols.visited_files)
def find_keysymdef():
env = os.environ.get("KEYSYMDEF")
if env and os.path.isfile(env):
return env
for p in KEYSYMDEF_CANDIDATES:
if os.path.isfile(p):
return p
sys.exit("error: keysymdef.h not found; install xorgproto or set KEYSYMDEF=/path/to/keysymdef.h")
def cmd_fetch(_args):
local = os.environ.get("XKB_TARBALL")
if local:
data = open(local, "rb").read()
else:
print(f"downloading {TARBALL_URL}")
data = urllib.request.urlopen(TARBALL_URL).read()
digest = hashlib.sha256(data).hexdigest()
if digest != TARBALL_SHA256:
sys.exit(f"error: sha256 mismatch\n expected {TARBALL_SHA256}\n got {digest}")
tar = tarfile.open(fileobj=io.BytesIO(data), mode="r:gz")
members = tar.getnames()
root = members[0].split("/")[0]
# Extract symbols/ + COPYING to a temp view, resolve includes, keep only what's needed.
tmp = os.path.join(VENDOR, ".upstream")
for m in tar.getmembers():
if m.name.startswith(f"{root}/symbols/") or m.name == f"{root}/COPYING":
m.name = m.name[len(root) + 1:]
tar.extract(m, tmp)
tar.close()
needed = collect_needed(Symbols(os.path.join(tmp, "symbols")))
os.makedirs(os.path.join(VENDOR, "symbols"), exist_ok=True)
for fname in sorted(needed):
src = os.path.join(tmp, "symbols", fname)
dst = os.path.join(VENDOR, "symbols", fname)
os.makedirs(os.path.dirname(dst), exist_ok=True)
with open(src, "rb") as s, open(dst, "wb") as d:
d.write(s.read())
_copy(os.path.join(tmp, "COPYING"), os.path.join(VENDOR, "COPYING"))
keysymdef = find_keysymdef()
_copy(keysymdef, os.path.join(VENDOR, "keysymdef.h"))
_rmtree(tmp)
with open(os.path.join(VENDOR, "PROVENANCE.md"), "w") as f:
f.write("# Vendored xkeyboard-config subset\n\n")
f.write(f"- **Package**: xkeyboard-config {VERSION}\n")
f.write(f"- **Source**: {TARBALL_URL}\n")
f.write(f"- **sha256**: `{TARBALL_SHA256}`\n")
f.write(f"- **keysymdef.h**: xorgproto, copied from `{keysymdef}`\n")
f.write("- **License**: MIT/X11 (see COPYING)\n\n")
f.write("Only the symbols files reachable from the generated layouts "
"(tools/make-xkeyboard-config.py `TARGETS`) are vendored; regenerate with\n"
"`python3 tools/make-xkeyboard-config.py fetch` then `... generate`.\n\n")
f.write("Vendored symbols files:\n\n")
for fname in sorted(needed):
f.write(f"- `symbols/{fname}`\n")
print(f"vendored {len(needed)} symbols files + keysymdef.h + COPYING into {VENDOR}")
def cmd_generate(args):
symbols_dir = args.symbols or os.path.join(VENDOR, "symbols")
keysymdef_path = args.keysymdef or os.path.join(VENDOR, "keysymdef.h")
keysymdef = parse_keysymdef(keysymdef_path)
symbols = Symbols(symbols_dir)
layouts = {zig_name: build_layout(symbols, keysymdef, fname, section)
for zig_name, fname, section in TARGETS}
out_dir = os.path.join(LIB, "generated")
os.makedirs(out_dir, exist_ok=True)
out_path = os.path.join(out_dir, "layouts.zig")
with open(out_path, "w") as f:
f.write(emit_zig(layouts))
print(f"wrote {out_path} ({len(TARGETS)} layouts)")
def _copy(src, dst):
os.makedirs(os.path.dirname(dst), exist_ok=True)
with open(src, "rb") as s, open(dst, "wb") as d:
d.write(s.read())
def _rmtree(path):
for root, dirs, files in os.walk(path, topdown=False):
for name in files:
os.remove(os.path.join(root, name))
for name in dirs:
os.rmdir(os.path.join(root, name))
if os.path.isdir(path):
os.rmdir(path)
def main():
parser = argparse.ArgumentParser(description=__doc__)
sub = parser.add_subparsers(dest="command", required=True)
sub.add_parser("fetch", help="download + vendor the needed xkeyboard-config data")
g = sub.add_parser("generate", help="emit generated/layouts.zig from the vendored data")
g.add_argument("--symbols", help="override the vendored symbols/ dir (for development)")
g.add_argument("--keysymdef", help="override the vendored keysymdef.h (for development)")
args = parser.parse_args()
{"fetch": cmd_fetch, "generate": cmd_generate}[args.command](args)
if __name__ == "__main__":
main()
+152
View File
@@ -0,0 +1,152 @@
#!/usr/bin/env python3
"""Rewrap over-long Zig comment lines at a column limit (default 100).
Rules:
- Only comment-only lines are touched; trailing comments after code are left alone.
- Consecutive comment lines with the same indentation and marker (`//`, `///`, `//!`)
form a block. Blank comment lines separate paragraphs within a block.
- A line whose text starts with `- ` begins a bullet paragraph; its continuation
lines are the ones indented to align under the bullet's text (bullet lead + 2).
- Plain paragraphs join consecutive lines with the same text indentation;
wrapped lines align where the first line's text begins.
- A paragraph is rewrapped only if at least one of its lines exceeds the limit,
so deliberate short line breaks elsewhere are preserved.
"""
import argparse
import difflib
import re
import subprocess
import sys
import textwrap
LIMIT = 100
COMMENT_RE = re.compile(r"^(\s*)(//[/!]?)(?:\s(.*))?$")
BULLET_RE = re.compile(r"^(\s*)- (.*)$")
def split_paragraphs(texts):
"""texts: list of comment text (None for a bare marker line).
Returns paragraphs: dicts with lead/hang/bullet/texts/lines(indices)."""
paragraphs = []
current = None
for index, text in enumerate(texts):
if text is None or text.strip() == "":
paragraphs.append({"literal": True, "lines": [index]})
current = None
continue
lead = len(text) - len(text.lstrip(" "))
bullet = BULLET_RE.match(text)
if bullet:
current = {
"lead": len(bullet.group(1)),
"hang": len(bullet.group(1)) + 2,
"bullet": True,
"texts": [bullet.group(2)],
"lines": [index],
}
paragraphs.append(current)
elif current is not None and lead == current["hang"]:
current["texts"].append(text)
current["lines"].append(index)
else:
current = {
"lead": lead,
"hang": lead,
"bullet": False,
"texts": [text],
"lines": [index],
}
paragraphs.append(current)
return paragraphs
def wrap_paragraph(paragraph, prefix):
joined = re.sub(r"\s+", " ", " ".join(t.strip() for t in paragraph["texts"]))
if paragraph["bullet"]:
initial = prefix + " " * paragraph["lead"] + "- "
else:
initial = prefix + " " * paragraph["lead"]
subsequent = prefix + " " * paragraph["hang"]
return textwrap.wrap(
joined,
width=LIMIT,
initial_indent=initial,
subsequent_indent=subsequent,
break_long_words=False,
break_on_hyphens=False,
)
def rewrap_block(original_lines, indent, marker, texts):
prefix = indent + marker + " "
output = []
for paragraph in split_paragraphs(texts):
block_originals = [original_lines[i] for i in paragraph["lines"]]
if paragraph.get("literal") or all(len(l) <= LIMIT for l in block_originals):
output.extend(block_originals)
else:
output.extend(wrap_paragraph(paragraph, prefix))
return output
def process(source):
lines = source.split("\n")
result = []
i = 0
while i < len(lines):
match = COMMENT_RE.match(lines[i])
if not match:
result.append(lines[i])
i += 1
continue
indent, marker = match.group(1), match.group(2)
block_lines, texts = [], []
while i < len(lines):
m = COMMENT_RE.match(lines[i])
if not m or m.group(1) != indent or m.group(2) != marker:
break
block_lines.append(lines[i])
texts.append(m.group(3))
i += 1
result.extend(rewrap_block(block_lines, indent, marker, texts))
return "\n".join(result)
def main():
parser = argparse.ArgumentParser()
parser.add_argument("--write", action="store_true", help="apply changes (default: diff only)")
parser.add_argument("files", nargs="*", help="files to process (default: git ls-files '*.zig')")
args = parser.parse_args()
files = args.files or subprocess.run(
["git", "ls-files", "*.zig"], capture_output=True, text=True, check=True
).stdout.split()
changed = 0
for path in files:
with open(path, encoding="utf-8") as f:
source = f.read()
rewrapped = process(source)
if rewrapped == source:
continue
changed += 1
if args.write:
with open(path, "w", encoding="utf-8") as f:
f.write(rewrapped)
print(f"rewrapped {path}")
else:
sys.stdout.writelines(
difflib.unified_diff(
source.splitlines(keepends=True),
rewrapped.splitlines(keepends=True),
fromfile=path,
tofile=path,
)
)
if not changed:
print("no comments over the limit")
if __name__ == "__main__":
main()
+22
View File
@@ -0,0 +1,22 @@
#!/usr/bin/env bash
#
# sort-lines-group-by-start — cluster lines that share their first
# whitespace-separated field ($1). Keys appear in first-seen order, and lines
# within a key keep their original order. It groups; it does NOT sort.
#
# Pass the log file as an argument; result is written to stdout:
#
# tools/sort-lines-group-by-start.sh filename.log
#
# Useful for a serial/boot log where several sources interleave and each line is
# prefixed with its source (the first field): this pulls every source's lines
# back together, in the order the sources first appeared, without reordering
# within a source.
#
# input output
# pci-bus: scan start pci-bus: scan start
# acpi: reported PNP0303 pci-bus: 5 functions
# pci-bus: 5 functions acpi: reported PNP0303
# acpi: reported PNP0501 acpi: reported PNP0501
awk '{lines[$1] = lines[$1] ? lines[$1] ORS $0 : $0; if (!seen[$1]++) order[++count] = $1} END {for (i=1; i<=count; i++) print lines[order[i]]}' "$@"