36 Commits
Author SHA1 Message Date
Daniel Samson 116b8f6c41 Design the process lifecycle and the device manager (M17-M18)
Signals over IPC (POSIX concepts, message delivery), published exit events,
the stable runtime.process interface, and the device manager as tree +
matcher + supervisor. All open questions settled; docs/m17-m18-plan.md is
the phase-by-phase execution plan.
2026-07-12 22:56:34 +01:00
Daniel Samson 77901bbba6 WIP: USB 2026-07-12 22:24:47 +01:00
Daniel Samson 78582d24d2 code lint 2026-07-12 19:36:10 +01:00
Daniel Samson 1cdffe21b1 fixing comments 2026-07-12 16:19:09 +01:00
Daniel Samson 4df90bc212 add tools/rewrap-comments.py 2026-07-12 16:18:56 +01:00
Daniel Samson 713e77354b gitattributes 2026-07-12 16:09:18 +01:00
Daniel Samson abb7b1b634 editorconfig 2026-07-12 16:09:11 +01:00
Daniel Samson f5f0e15769 zig fmt 2026-07-12 16:04:58 +01:00
Daniel Samson 8652b4a724 add qemu-xhci with usb-mouse and usb-kbd to build.zig 2026-07-12 01:31:38 +01:00
Daniel Samson e8233127c7 Install ps2-bus under its own name instead of clobbering bus
The ps2-bus executable was built with the artifact name "bus" — a
copy-paste from the generic bus driver's line above it. The initial
ramdisk was unaffected (the packer pairs names with binaries
explicitly), so the driver ran at boot; but the FHS install uses the
artifact's own name, so both drivers landed on
zig-out/system/drivers/bus, one overwriting the other, and
zig-out/system/drivers/ps2-bus never existed.
2026-07-11 23:41:08 +01:00
Daniel Samson 88e92254e9 Name ACPI hardware IDs instead of magic _HID strings
Turn acpi-ids.zig's flat name table into a HardwareId enum modeled on
ps2-library's Port: one entry() switch holds the registry (variant ->
_HID string + human-readable name), with hid(), description(), and
fromHid() methods. The free description(hid) lookup the kernel's
device-tree dump uses survives, implemented over the enum, and a new
test round-trips every variant through fromHid.

Callers now name the device instead of quoting its id:

- ps2-library's DeviceType.hid() and ps2-bus's descriptor lookups use
  HardwareId.ps2_keyboard / .ps2_mouse.
- device-manager's driverFor parses the HID once with fromHid and
  switches on named values.
- acpi.zig's isPciRootNode carried the same ids twice, as strings and
  as packed-EISA integers (0x030AD041/0x080AD041); both branches now
  decode to the string form and answer through one isPciRootHid helper
  using .pci_bus / .pci_express_root_bridge.
- build.zig threads the acpi-ids module (previously kernel-only) into
  every user binary, like xkeyboard-config.
2026-07-11 23:23:16 +01:00
Daniel Samson 5725d35e5b Name the attach reply statuses instead of magic numbers
Add an AttachStatus enum to ps2-library.zig for AttachReply.status,
distinguishing the three failure causes handleAttach previously
collapsed into a bare -1: invalid_request (message too short),
missing_endpoint (no capability passed), and no_such_device (no port
identified the requested device type). The keyboard and mouse drivers
check against AttachStatus.ok rather than a literal 0.
2026-07-11 23:11:45 +01:00
Daniel Samson 8c95525793 Wire real PS/2 mouse packets through to input events
Replace the mouse driver's synthetic stream with the real path, the
way the keyboard was wired:

- The auxiliary port's IRQ12 is enumerated on the mouse's own ACPI
  node (PNP0F13), and the kernel only lets a device's claimer bind or
  ack its IRQs — so ps2-bus now claims that node alongside the
  controller whenever port 2 carries a device, binds IRQ12 to its one
  endpoint, and re-arms whichever line the notification's badge names.
  The forwarding loop already routed auxiliary bytes by status bit 5.
- mouse-packet.zig (new, pure, host-tested): three-byte stream-mode
  packet assembly — bit-3 sync with resynchronization, ACK/BAT bytes
  dropped at packet start, nine-bit two's-complement movement,
  overflow packets discarded, and PS/2 positive-Y-up converted to the
  screen convention (positive down).
- mouse.zig mirrors the keyboard driver: no hardware claim, attaches
  to the bus as its mouse, and publishes button_down/button_up per
  changed button plus motion events with the pressed-button mask.

Verified end to end in QEMU via monitor mouse_move/mouse_button:
motion round-trips in screen coordinates, buttons transition with the
right mask, and keyboard events keep flowing alongside. The follow-up
is the IntelliMouse magic-knock for a scroll wheel (four-byte packets)
and scroll events.
2026-07-11 23:09:04 +01:00
Daniel Samson 5bba5d3363 Wire real PS/2 scancodes through to input events and characters
Replace the keyboard driver's synthetic stream with the real path:

- ps2-bus binds IRQ1 (interrupt bits set only after the bind), drains
  port 0x60 on each interrupt, and forwards every byte to the attached
  child driver over async ipc_send, routed by the status register's
  auxiliary-output bit. Children attach via the new well-known ps2_bus
  service, handing over their endpoint as a capability.
- scancode.zig (new, pure, host-tested): scancode set 2 -> USB HID
  usage decoding (F0/E0/E1 prefix state machine) plus keyboard state —
  pressed-key bitmap, typematic-repeat classification, modifier and
  caps-lock tracking.
- keyboard.zig decodes the forwarded stream and publishes real
  key_down/key_press/key_up events, filling key_press characters via
  xkeyboard-config (layout from argv[2], default us) and synthesizing
  ASCII control characters for Enter/Tab/Backspace/Escape.
- protocol.zig names the full HID usage set in Keycode; build.zig
  threads the xkeyboard-config module into user binaries.

Verified end to end in QEMU via monitor sendkey: shift, caps lock,
and control-character synthesis all decode correctly. The mouse
driver still publishes its synthetic stream; attaching it to the
bus the same way is the follow-up.
2026-07-11 22:48:59 +01:00
Daniel Samson 80b72db676 Removing DAN-INIT 2026-07-11 21:43:49 +01:00
Daniel Samson aa0c97353a fixing arguments 2026-07-11 21:39:27 +01:00
daniel dd93204b44 Merge pull request 'claude/input-module-keyboard-events-379361' (#6) from claude/input-module-keyboard-events-379361 into main
Reviewed-on: #6
2026-07-11 20:28:46 +00:00
daniel c7b17aaa0e Merge branch 'main' into claude/input-module-keyboard-events-379361 2026-07-11 20:28:21 +00:00
Daniel Samson d7a154a596 Add xkeyboard-config: X11 keyboard layouts compiled to Zig
Turn a keycode + modifiers into a character. The input module delivers HID
usage keycodes but nothing mapped them to characters; rather than hand-maintain
layout tables, vendor the X11 xkeyboard-config database and compile it to native
Zig at build time (no X11 runtime), the way make-initial-ramdisk.py packs the
ramdisk.

- tools/make-xkeyboard-config.py: `fetch` downloads the pinned xkeyboard-config
  release (2.44, sha256-verified), resolves the include graph for the configured
  layouts, and vendors only the reached symbols files + keysymdef.h + COPYING +
  PROVENANCE into library/xkeyboard-config/vendor/. `generate` parses that
  (keycodes via a HID->xkb-name table, symbols with include/augment/override and
  per-key type, keysymdef for keysym->Unicode) and emits generated/layouts.zig
  deterministically.
- library/xkeyboard-config/xkeyboard-config.zig: the API over the generated data
  — map(layout, hid_usage, mods) -> { keysym, character }, byName, and the
  level-selection semantics (the generated tables stay pure data). Host tests
  assert US letters/digits with Shift/Caps, GB £ vs US # on Shift+3, and French
  AZERTY q-where-US-has-a — the end-to-end proof of the parse->emit->lookup path.
- build.zig: `xkeyboard-config` + `layouts` modules, the test wired into
  `zig build test`, and a `zig build gen-xkeyboard-config` convenience step.
- Layouts: us, gb, de, fr, es, dvorak. Scope (documented): group 1, no dead-key
  composition, curated key types. Standalone library; wiring it into the input
  path to fill KeyEvent.character is a documented follow-up.

zig build test green (incl. the new keymap tests); regeneration is byte-identical;
full QEMU suite 48/48 (unaffected — no kernel/runtime/service change).
2026-07-11 15:57:09 +01:00
Daniel Samson 1bf91115dd Generalize input module to mouse and joystick/gamepad events
Extend the input service beyond the keyboard so mouse and joystick/gamepad
drivers can broadcast too, with per-device publish and subscribe methods.

- protocol: KeyEvent joins MouseEvent (motion/buttons/scroll) and
  JoystickEvent (axes/buttons), all carried in a common InputEvent envelope
  tagged with a DeviceKind. A subscribe request carries a device_mask, so a
  subscriber names the classes it wants and the service routes each event only
  to interested subscribers (a mouse-only listener never wakes for keystrokes).
- runtime: per-device publish methods (publishKeyboardEvent/publishMouseEvent/
  publishJoystickEvent) and subscribe helpers (subscribeKeyboard/Mouse/Joystick,
  each typed, plus subscribe(mask)/subscribeAll returning the tagged envelope).
- service: subscriber table gains a device_mask; broadcast routes by the
  event's device class.
- mouse driver now publishes (synthetic) mouse events like the keyboard driver;
  input-source cycles all three classes; input-test subscribes to all and only
  emits its "ok" marker once it has received one of each class — so the passing
  test proves per-device routing, not just delivery. Real HID decoding stays a
  follow-up.

No kernel changes: ipc_send is generic and the 36-byte InputEvent fits its
64-byte payload. Full QEMU suite 48/48; serial log confirms keyboard, mouse,
and joystick all reach one subscription.
2026-07-11 15:21:09 +01:00
Daniel Samson 65244e3103 Add input module: broadcast keyboard events over IPC
Programs can now subscribe to keyboard events (key_down/key_up/key_press)
and drivers can broadcast them, through a new user-space input service.

The delivery model is forced by danos IPC: a synchronous rendezvous holds
one pending reply, so a server cannot park N subscribers blocked in a
"wait for next event" call — delivery must be push. But a synchronous push
has no timeout and the kernel never wakes a sender parked on a dead peer's
endpoint, so one dying subscriber would hang all input. So this lands the
roadmap's planned asynchronous buffered send and builds the service on it:

- ipc_send (syscall 26): non-blocking post to an endpoint's bounded payload
  ring, delivered through reply_wait as a buffered message (notify_message_bit).
  A full ring drops the oldest. It can never hang on a dead/slow peer.
- input-protocol + runtime.input helpers (subscribe/next, connectSource/
  publish) — the first real consumer of M13 capability passing: a subscriber
  hands the service its own endpoint as a capability.
- input service (fan-out via ipc_send, dead-subscriber pruning), a synthetic
  input-source, and input-test; the ps2-bus keyboard driver publishes to it.
  Real IRQ1 scancode decoding (which must live in the bus, the PNP0303 owner)
  is a documented follow-up; the source is synthetic for now.
- build/init wiring, an `input` QEMU case, and docs/input.md.

Full QEMU suite 48/48, including the new input case and every IPC/endpoint
regression (ipc, ipc-call, ipc-cap, vfs, hpet, bus, irqfree).
2026-07-11 15:03:24 +01:00
Daniel Samson 75ccfff171 Pass arguments to main via runtime.process.Init, dispatched on signature 2026-07-11 14:22:48 +01:00
Daniel Samson 2a583d55a8 Finishing PS/2 bus driver 2026-07-11 14:12:33 +01:00
Daniel Samson d218d93f79 Add process management: enumerate, supervisor-gated kill, exit notifications
process_enumerate snapshots the task table (the device_enumerate shape, so
ps is a user program); system_spawn returns the child id, records the caller
as supervisor, and takes an exit endpoint; process_kill is allowed only for
the supervisor. Every death — exit, fault, or kill — posts a child-exit badge
to that endpoint (the IRQ-as-IPC pattern as SIGCHLD). A target caught off-CPU
is reaped in place; a running one is condemned and finished at its next
system call or tick, guarded so teardown never lands mid-kernel-operation.
Tested by process-list, process-kill, and supervision (a ring-3 supervisor
exercising the whole surface); design notes in docs/process-management.md.
2026-07-11 09:32:25 +01:00
Daniel Samson a5fe63c1dd Pass argv to processes on a SysV entry stack; grow the user stack to 32 KiB
Processes now start with C-compatible arguments: the kernel builds the
System V AMD64 entry block (argc, argv, empty envp, auxiliary vector)
at the top of the stack, argv[0] is the path or initial-ramdisk name
the process was spawned as, and system_spawn carries an optional
NUL-separated blob that becomes argv[1..]. The runtime parses the block
(runtime.argumentCount/argument) and its spawn wrappers pass arguments
through. The name is also recorded on the task, so a fault report says
which binary died, not just its id.

The user stack grows from one page to eight (32 KiB,
parameters.user_stack_pages), with the page below left unmapped as a
guard so an overflow faults into a clean process kill rather than
corrupting the image. Task.name_buffer is zero-initialised, not
undefined: an undefined default is materialised as a 0xAA fill that
moved the static task pool out of .bss and made the whole kernel ~7x
slower under QEMU TCG (caught by the affinity test).

Proven end to end by the new args test: args-echo respawns itself with
arguments via the syscall blob, burns more stack than one page could
hold, and echoes its argv intact. Full suite: 44/44.
2026-07-11 08:33:12 +01:00
Daniel Samson 6b3ae0c997 Kill a faulting user process instead of halting the machine
A CPU exception raised in ring 3 by a scheduled process now kills that
process - IRQ bindings, IPC handles, and address space reclaimed, a
client it owed a reply to failed with the new -EPEER instead of hung -
and the core reschedules (docs/resilience.md step 2). Kernel-mode
faults, NMI, double fault, and machine check stay terminal, as does the
borrowed-thread isolation probe. Proven by the new fault-recovery QEMU
test: init keeps heartbeating after a process page-faults to death.
2026-07-11 04:57:02 +01:00
Daniel Samson 59104dd988 refactor 2026-07-11 04:32:17 +01:00
Daniel Samson e499f500c3 adding clock system call 2026-07-11 03:43:33 +01:00
Daniel Samson 2ab0d129a2 Decode ACPI _HID names in device discovery
The flat analog of pci-class for acpi_device nodes. ACPI has no class/subclass/prog-IF
taxonomy — a device's identity is its _HID string itself (PNP0303 *is* "PS/2
keyboard") — so this is a plain id -> name registry, not a hierarchical decoder.

New system/devices/acpi-ids.zig (module `acpi-ids`): the common standard PnP/ACPI
hardware IDs; vendor-specific ids (QEMU0002, etc.) have no registry name and print the
raw HID. Shared reference data like pci-class. The dump now names each _HID:

    KBD_ [acpi_device] hid=PNP0303 (PS/2 Keyboard)
    COM1 [acpi_device] hid=PNP0501 (16550A-compatible Serial Port)
    RTC_ [acpi_device] hid=PNP0B00 (Real-Time Clock (RTC))
    LNKA [acpi_device] hid=PNP0C0F (PCI Interrupt Link Device)
    FWCF [acpi_device] hid=QEMU0002          (vendor-specific: raw HID)

Host test covers known ids and the unknown/empty fallthrough. Suite 41/41 plus host
tests.
2026-07-10 21:44:44 +01:00
Daniel Samson d702d2e9ae Decode PCI class codes in device discovery
`pci_device` alone says nothing — an ISA bridge, an AHCI controller, and an xHCI USB
controller are all just `pci_device` by DeviceClass. The identity lives in the 24-bit
class code (base class / subclass / prog-IF) that discovery already recorded in
ids.pci_class but the dump threw away.

New system/devices/pci-class.zig (module `pci-class`): pure reference data from the PCI
spec (per the OSDev PCI table) decoding the triple into names — className, subclassName,
progIfName, plus ClassCode.unpack. No hardware access, so it's shared by kernel
discovery and any future user-space PCI tool. device_tree.dump now prints, under each
PCI function, its class/subclass/prog-IF as both hex and name:

    pci0:00:1f.0 [pci_device]
      class 0x06 (Bridge)  subclass 0x01 (ISA Bridge)  progif 0x00
    pci0:00:1f.2 [pci_device]
      class 0x01 (Mass Storage Controller)  subclass 0x06 (Serial ATA Controller)  progif 0x01 (AHCI 1.0)

Host test covers the decoder (bridge/AHCI/xHCI, the "Other" 0x80 convention, unknowns).
Suite 41/41 plus host tests.
2026-07-10 21:35:20 +01:00
Daniel Samson 3ec2d1828a docs: abi is the kernel↔runtime contract, not spoken by applications
Correct the abi.zig header (and the matching build.zig / README lines): the syscall ABI
is the private contract between the kernel and the runtime library, not something every
user program speaks. danos applications call the `runtime` (the stable, danos-native
ABI); the runtime is the only thing that issues system calls, and POSIX layers over the
runtime — the same split as libSystem on macOS or win32 over the NT syscalls. The call
numbers here are an implementation detail the runtime hides and may renumber, not a
public interface.
2026-07-10 20:38:41 +01:00
Daniel Samson 21657943a8 Port I/O grants (io_read / io_write) + refresh stale driver docs
Ring 3 still has no direct in/out (no TSS I/O bitmap, IOPL never raised — a #GP), but a
driver no longer needs it: io_read(device_id, resource_index, offset, width) and
io_write(..., value) grant port access the same way mmio_map grants memory. The claim
plus the device's discovered io_port resource are the capability — resolveIoPort checks
the device is claimed by the caller, the resource is io_port, and [offset, offset+width)
stays inside it, then issues the in/out via architecture.pioRead/pioWrite. So a PS/2 or
16550 driver is now writable; the low-rate legacy hardware that needs port I/O is fine
with a syscall per access. io_port resources were recorded by discovery and ignored —
now they're used. Runtime: device.ioRead/ioWrite.

New `ioport` test claims QEMU's PS/2 controller (io_port 0x64, discovered via ACPI) and
checks the gate admits an in-range access, refuses over-wide / out-of-range / unclaimed,
and that the kernel actually reads the status port (0x1c). Suite 41/41 plus host tests.

Docs refreshed for the whole M13-M16 + port-I/O reality: drivers.md "Limits" no longer
lists port I/O, DMA memory, or barriers as missing (they exist) and its "what's next"
reflects that; the FSH doc's character-device and block-device sections are corrected
("cannot host a block driver at all" is no longer true — writable now, not yet memory-
safe pending IOMMU enforcement); device-interrupts.md unblocks the keyboard and narrows
the interrupt gap to MSI-X.
2026-07-10 20:36:14 +01:00
Daniel Samson e612d948d2 M16: IOMMU detection (DMAR parsing)
Detect the IOMMU: discovery now parses the ACPI DMAR table, finds the first VT-d
DMA-remapping unit (DRHD), maps its register block, and records its version and
capabilities (iommu_present/base/version/capabilities in the platform info). On QEMU's
emulated intel-iommu this reads back a real unit (base 0xfed90000, version 1.0).

This is detection only, and deliberately so. A full VT-d bring-up — per-device
translation domains that confine a driver's DMA to the buffers it dma_alloc'd — is the
real device-side safety guarantee, but it cannot be verified without a DMA-capable
device driver (none exist yet) and QEMU's intel-iommu to fault against. Writing that
enforcement now would be a large body of unverifiable page-table code; it belongs with
the first DMA driver, which is both the natural order and the only way to test it. Until
then the caveat stands in full: device_claim on a DMA-capable device is still equivalent
to granting ring 0. The docs say so plainly.

New `iommu` test (harness boots it with -device intel-iommu via a new per-case qemu_extra
hook) confirms the DMAR is parsed and the unit's registers read. Suite 40/40 plus host
tests.
2026-07-10 20:07:01 +01:00
Daniel Samson 4ef21fa083 M15: interrupts for PCI devices (ECAM config space + MSI)
Two parts, both blockers for real PCI drivers.

ECAM config space per function: enumeratePci now gives every pci_device its own 4 KiB
configuration window as resource 0. A claimed PCI driver mmio_maps that to reach its
command register, BARs, and — the point — its capability list (MSI/MSI-X, PCIe extended
caps), with no new syscall. Verified in the discovery test (QEMU q35's functions each
carry it).

MSI: msi_bind(device_id, endpoint) -> address (rax), data (rdx) allocates a per-device
edge-triggered vector, binds it to the endpoint, and returns the (address, data) the
driver programs into its own MSI capability. Unlike irq_bind there's no GSI, no I/O APIC
entry, no sharing, and no ack cycle — dispatch recognises an MSI vector (vector_gsi ==
none, msi_bound set), EOIs, and notifies. Owner-keyed release drops the binding on exit.
Legacy INTx (_PRT parsing + shared lines) is deliberately skipped; MSI is the real
answer.

QEMU's HPET has no MSI, so the new `msi` test proves the vector-routing path with a
self-IPI (new apic.selfIpi) standing in for the device's MSI write: bind a vector,
fire it, the bound endpoint is notified. The msi_bind syscall wraps irq.msiBind with
the claim check and lands its first real use with the first PCI driver. Suite 39/39
plus host tests.
2026-07-10 19:59:09 +01:00
Daniel Samson 125a3b4993 M14b: DMA memory (dma_alloc / dma_free)
An HCD programs a bus-master engine: it needs a descriptor ring that is physically
contiguous, at a physical address it knows, uncacheable, and pinned. mmap gives none
of those. Add dma_alloc(len, flags) -> vaddr (rax), paddr (rdx) and dma_free(vaddr,
len): grant contiguous, zeroed, pinned, strong-uncacheable memory in a per-process DMA
arena (PML4[228]) and hand back both addresses.

Pieces: pmm.allocContiguous(count, max_phys) finds a run of contiguous free frames
below a cap (dma_below_4g for 32-bit engines); mapUserDmaInto maps them uncacheable
(PCD|PWT) but WITHOUT device_grant, so unlike an MMIO grant these frames are real RAM
and freeSubtree returns them on teardown — a driver that dies leaks nothing. dma_free
is bounded to the DMA arena so it can never unmap the caller's stack/heap/MMIO.
dma_write_combining is accepted but falls back to coherent (WC needs PAT programming).

Runtime: runtime.dma.alloc/free (a two-return-value stub, like replyWait). New `dma`
kernel test drives the mechanism directly — contiguity, the below-4G cap, coherent
mapping, and reclaim-on-teardown (no leak). The thin syscall wrappers follow the tested
mmap/mmio_map shape and land their first real use with the first DMA driver. Suite
38/38 plus host tests.
2026-07-10 19:43:59 +01:00
Daniel Samson e7c7e7b94c M14a: memory-ordering / MMIO layer (library/mmio)
The tree had zero memory barriers — correct-by-accident on x86 (TSO + strong-
uncacheable MMIO), but a landmine for the first DMA driver and for ARM, which is the
win condition. Add /lib/mmio: typed volatile register access (read/write) plus mb /
rmb / wmb, lowered per-architecture (mfence/lfence/sfence on x86_64, dsb sy/ld/st on
aarch64) so the ordering rules are a named primitive, not scattered `asm volatile`.
`volatile` is not a barrier — it says nothing about ordinary stores (a DMA descriptor
in WB RAM) relative to a volatile doorbell write; wmb() between them is the fix.

Prove it on the one existing caller: hpet now does its register access through
mmio.read/write. It needs no barriers itself (pure MMIO, no DMA, UC grant on x86) —
the point is the typed, arch-portable access every driver should use; the barriers are
there for the DMA drivers to come.

New `mmio` module injected into addUserBinary; host test asserts the barriers assemble
and a register round-trips. Suite 37/37 plus host tests.
2026-07-10 19:32:57 +01:00
90 changed files with 18991 additions and 287 deletions
+16
View File
@@ -0,0 +1,16 @@
# EditorConfig: https://editorconfig.org/
# Follows the Zig style guide: https://ziglang.org/documentation/0.16.0/#Style-Guide
root = true
[*]
charset = utf-8
end_of_line = lf
indent_style = space
indent_size = 4
trim_trailing_whitespace = true
insert_final_newline = true
[*.zig]
# "Line length: aim for 100; use common sense."
max_line_length = 100
+1
View File
@@ -0,0 +1 @@
*.zig text eol=lf
+134 -7
View File
@@ -59,6 +59,9 @@ fn addUserBinary(
target: std.Build.ResolvedTarget,
runtime_module: *std.Build.Module,
posix_module: *std.Build.Module,
mmio_module: *std.Build.Module,
xkeyboard_config_module: *std.Build.Module,
acpi_ids_module: *std.Build.Module,
name: []const u8,
root: []const u8,
) *std.Build.Step.Compile {
@@ -78,6 +81,14 @@ fn addUserBinary(
// POSIX/C compatibility layer, available to any program that wants it
// (danos-native code uses `runtime` directly). See library/posix/.
.{ .name = "posix", .module = posix_module },
// Typed volatile MMIO + memory barriers, for drivers. See library/mmio/.
.{ .name = "mmio", .module = mmio_module },
// Keyboard layouts (keycode + modifiers -> keysym/character), available
// to any program that wants it. See library/xkeyboard-config/.
.{ .name = "xkeyboard-config", .module = xkeyboard_config_module },
// ACPI/PnP hardware-ID registry, so drivers name devices
// (HardwareId.ps2_keyboard) instead of magic "_HID" strings.
.{ .name = "acpi-ids", .module = acpi_ids_module },
},
}),
});
@@ -99,7 +110,7 @@ pub fn build(b: *std.Build) void {
// declares which one it speaks (no target is set, so each inherits the target of
// whichever binary imports it). See docs/coding-standards.md.
// boot-handoff : loader <-> kernel (BootInformation, framebuffer, VM layout)
// abi : kernel <-> user, core (SystemCall, mmap prot flags, page_size)
// abi : kernel <-> runtime, core (SystemCall, mmap prot flags, page_size)
// device-abi : kernel <-> user, devices (DeviceDescriptor, DeviceClass, ...)
const boot_handoff_module = b.addModule("boot-handoff", .{
.root_source_file = b.path("system/boot-handoff.zig"),
@@ -113,6 +124,16 @@ pub fn build(b: *std.Build) void {
const device_abi_module = b.addModule("device-abi", .{
.root_source_file = b.path("system/devices/device-abi.zig"),
});
// PCI class-code decoding (class/subclass/prog-IF -> names). Pure reference data,
// shared by kernel discovery (the device-tree dump) and any user-space PCI tool.
const pci_class_module = b.addModule("pci-class", .{
.root_source_file = b.path("system/devices/pci-class.zig"),
});
// ACPI/PnP hardware-ID (_HID) names — the flat analog of pci-class for acpi_device
// nodes. Also shared reference data.
const acpi_ids_module = b.addModule("acpi-ids", .{
.root_source_file = b.path("system/devices/acpi-ids.zig"),
});
// Kernel tunables (maximum_cpus, stack sizes, tick rate). A dependency-free module of
// compile-time constants, imported wherever a knob is read; keeps the trade-offs
@@ -150,6 +171,8 @@ pub fn build(b: *std.Build) void {
.{ .name = "boot-handoff", .module = boot_handoff_module }, // BootInformation (carries the ACPI RSDP), physicalToVirtual
.{ .name = "abi", .module = abi_module }, // acpi.zig works in page_size units
.{ .name = "device-abi", .module = device_abi_module }, // device-model's DeviceClass/ResourceKind live here
.{ .name = "pci-class", .module = pci_class_module }, // decode PCI class codes in the device dump
.{ .name = "acpi-ids", .module = acpi_ids_module }, // decode ACPI _HID names in the device dump
.{ .name = "parameters", .module = parameters_module }, // maximum_cpus (the discovery pool)
},
});
@@ -163,6 +186,13 @@ pub fn build(b: *std.Build) void {
.root_source_file = b.path("system/services/vfs/protocol.zig"),
});
// The input wire protocol: the input service's public interface, exposed as its own
// module the same way vfs-protocol is. Shared by the input service, the runtime's
// `input` helper (subscribe/publish), and every source and subscriber.
const input_protocol_module = b.addModule("input-protocol", .{
.root_source_file = b.path("system/services/input/protocol.zig"),
});
// The danos-native user-space runtime: system_call wrappers, the C-convention
// heap, IPC helpers, the process start shim, device access. This is the stable
// application ABI; POSIX compatibility is a separate library on top (see below).
@@ -178,6 +208,28 @@ pub fn build(b: *std.Build) void {
.{ .name = "abi", .module = abi_module },
.{ .name = "device-abi", .module = device_abi_module },
.{ .name = "vfs-protocol", .module = vfs_protocol_module },
.{ .name = "input-protocol", .module = input_protocol_module },
},
});
// Typed volatile MMIO register access + memory-ordering barriers, for drivers on
// top of an mmio_map grant. Depends only on `builtin` (arch-conditional barriers);
// no target set, so it inherits each driver's. See library/mmio/mmio.zig.
const mmio_module = b.addModule("mmio", .{
.root_source_file = b.path("library/mmio/mmio.zig"),
});
// Keyboard layouts compiled from the X11 xkeyboard-config database into native Zig
// (keycode + modifiers -> keysym/character). The `layouts` tables are generated by
// tools/make-xkeyboard-config.py; `xkeyboard-config` is the hand-written API over them.
// No target set, so each inherits its importer's. See library/xkeyboard-config/.
const xkb_layouts_module = b.addModule("layouts", .{
.root_source_file = b.path("library/xkeyboard-config/generated/layouts.zig"),
});
const xkeyboard_config_module = b.addModule("xkeyboard-config", .{
.root_source_file = b.path("library/xkeyboard-config/xkeyboard-config.zig"),
.imports = &.{
.{ .name = "layouts", .module = xkb_layouts_module },
},
});
@@ -264,7 +316,7 @@ pub fn build(b: *std.Build) void {
// Built by the shared user-binary recipe (see addUserBinary): freestanding,
// linked into the kernel's user region against the `runtime` runtime library, and
// started in ring 3 by the kernel's user-ELF loader.
const init_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, "init", "system/services/init/init.zig");
const init_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "init", "system/services/init/init.zig");
const init_install = b.addInstallArtifact(init_exe, .{ .dest_dir = .{ .override = .{ .custom = "system/services" } } });
b.getInstallStep().dependOn(&init_install.step);
@@ -272,11 +324,22 @@ pub fn build(b: *std.Build) void {
// Each is built by the same user-binary recipe, then packed into one image by
// the host-side make-initial-ramdisk tool. The bootloader ferries the image to the kernel,
// which unpacks it and spawns each program (system/initial-ramdisk.zig).
const vfs_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, "vfs", "system/services/vfs/vfs.zig");
const vfstest_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, "vfs-test", "system/services/vfs/vfs-test.zig");
const hpet_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, "hpet", "system/drivers/hpet/hpet.zig");
const bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, "bus", "system/drivers/bus/bus.zig");
const device_manager_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, "device-manager", "system/services/device-manager/device-manager.zig");
const vfs_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "vfs", "system/services/vfs/vfs.zig");
const vfstest_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "vfs-test", "system/services/vfs/vfs-test.zig");
const hpet_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "hpet", "system/drivers/hpet/hpet.zig");
const bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "bus", "system/drivers/bus/bus.zig");
const ps2_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-bus", "system/drivers/ps2-bus/ps2-bus.zig");
const ps2_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-keyboard", "system/drivers/ps2-bus/keyboard.zig");
const ps2_mouse_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-mouse", "system/drivers/ps2-bus/mouse.zig");
const usb_xhci_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-xhci-bus", "system/drivers/usb-xhci-bus/usb-xhci-bus.zig");
const device_manager_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-manager", "system/services/device-manager/device-manager.zig");
// The input service and its exercisers: the fan-out server, a hardware-free synthetic
// source, and a subscriber that doubles as the `input` test's oracle. See docs/input.md.
const input_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input", "system/services/input/input.zig");
const input_source_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input-source", "system/services/input-source/input-source.zig");
const input_test_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "input-test", "system/services/input-test/input-test.zig");
const args_echo_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "args-echo", "system/services/args-echo/args-echo.zig");
const process_test_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "process-test", "system/services/process-test/process-test.zig");
// Pack the user binaries into the initial_ramdisk image with the host-side Python tool
// (the container format is trivial, and Python sidesteps std API churn). Args:
@@ -292,16 +355,39 @@ pub fn build(b: *std.Build) void {
mk_run.addFileArg(hpet_exe.getEmittedBin());
mk_run.addArg("bus");
mk_run.addFileArg(bus_exe.getEmittedBin());
mk_run.addArg("ps2-bus");
mk_run.addFileArg(ps2_bus_exe.getEmittedBin());
mk_run.addArg("ps2-keyboard");
mk_run.addFileArg(ps2_keyboard_exe.getEmittedBin());
mk_run.addArg("ps2-mouse");
mk_run.addFileArg(ps2_mouse_exe.getEmittedBin());
mk_run.addArg("usb-xhci-bus");
mk_run.addFileArg(usb_xhci_bus_exe.getEmittedBin());
mk_run.addArg("device-manager");
mk_run.addFileArg(device_manager_exe.getEmittedBin());
mk_run.addArg("input");
mk_run.addFileArg(input_exe.getEmittedBin());
mk_run.addArg("input-source");
mk_run.addFileArg(input_source_exe.getEmittedBin());
mk_run.addArg("input-test");
mk_run.addFileArg(input_test_exe.getEmittedBin());
mk_run.addArg("args-echo");
mk_run.addFileArg(args_echo_exe.getEmittedBin());
mk_run.addArg("process-test");
mk_run.addFileArg(process_test_exe.getEmittedBin());
// Also install the packed binaries to their FHS homes, so zig-out is a true image
// of the filesystem — even though at boot they arrive inside the initial-ramdisk.
for ([_]struct { *std.Build.Step.Compile, []const u8 }{
.{ vfs_exe, "system/services" },
.{ device_manager_exe, "system/services" },
.{ input_exe, "system/services" },
.{ hpet_exe, "system/drivers" },
.{ bus_exe, "system/drivers" },
.{ ps2_bus_exe, "system/drivers" },
.{ ps2_keyboard_exe, "system/drivers" },
.{ ps2_mouse_exe, "system/drivers" },
.{ usb_xhci_bus_exe, "system/drivers" },
}) |entry| {
const step = b.addInstallArtifact(entry[0], .{ .dest_dir = .{ .override = .{ .custom = entry[1] } } });
b.getInstallStep().dependOn(&step.step);
@@ -373,6 +459,19 @@ pub fn build(b: *std.Build) void {
const run_efi = b.addSystemCommand(&.{
"qemu-system-x86_64",
"-device",
"qemu-xhci,id=xhci",
"-device",
"usb-mouse,bus=xhci.0",
"-device",
"usb-kbd,bus=xhci.0",
// "-usb",
// "-device",
// "usb-ehci,id=ehci",
// "-device",
// "usb-tablet,bus=usb-bus.0",
// "-device",
// "usb-mouse,bus=ehci.0",
"-machine",
"q35",
"-m",
@@ -429,6 +528,13 @@ pub fn build(b: *std.Build) void {
"system/boot-handoff.zig",
"system/abi.zig",
"system/devices/device-abi.zig",
"system/devices/pci-class.zig", // class/subclass/prog-IF name decoding
"system/devices/acpi-ids.zig", // _HID name decoding
"system/devices/usb-abi.zig", // wire sizes + bit packings + set-up packet encodings
"system/devices/usb-ids.zig", // class/subclass/protocol code assignments
"library/mmio/mmio.zig", // barriers assemble + registers round-trip
"system/drivers/ps2-bus/scancode.zig", // set-2 decode + keyboard state machine
"system/drivers/ps2-bus/mouse-packet.zig", // 3-byte mouse packet assembly
}) |root| {
const mod_tests = b.addTest(.{
.root_module = b.createModule(.{
@@ -439,4 +545,25 @@ pub fn build(b: *std.Build) void {
});
test_step.dependOn(&b.addRunArtifact(mod_tests).step);
}
// The xkeyboard-config keymap tests need its generated `layouts` import wired, so they
// don't fit the plain loop above. Its keycode->character assertions are the end-to-end
// proof that the xkb-data -> generator -> Zig-lookup pipeline is correct.
const xkb_tests = b.addTest(.{
.root_module = b.createModule(.{
.root_source_file = b.path("library/xkeyboard-config/xkeyboard-config.zig"),
.target = target,
.optimize = optimize,
.imports = &.{
.{ .name = "layouts", .module = xkb_layouts_module },
},
}),
});
test_step.dependOn(&b.addRunArtifact(xkb_tests).step);
// Convenience: `zig build gen-xkeyboard-config` regenerates the layout tables from the
// vendored data (offline). `fetch` (the network step) stays a manual script run.
const gen_xkb = b.addSystemCommand(&.{ "python3", "tools/make-xkeyboard-config.py", "generate" });
const gen_xkb_step = b.step("gen-xkeyboard-config", "Regenerate library/xkeyboard-config/generated from the vendored data");
gen_xkb_step.dependOn(&gen_xkb.step);
}
+22 -3
View File
@@ -48,7 +48,26 @@ rather than restate it. Roughly in the order things happen at runtime:
real driver stacks factor into three shapes, how families share code, and the
proposed ABI for the three primitives still missing (capability passing, DMA +
memory barriers, MSI).
15. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
15. **[process-management.md](process-management.md) — process management.** The
microkernel's `ps`/`kill`/SIGCHLD: enumerate as a table snapshot, the
supervision link as the kill authority, and child-exit notifications over the
same endpoints IRQs arrive on.
16. **[process-lifecycle.md](process-lifecycle.md) — the process lifecycle.** Design:
signals over IPC as the one lifecycle vocabulary every process speaks — the
POSIX.1-1990 words with message delivery instead of stack hijack, the stable
`runtime.process` interface, exit reasons, published exit events any stateful
service can subscribe to (the VFS releasing dead clients' handles), and the two
iron rules (cleanup is the kernel's job; kill is not a signal).
17. **[device-manager.md](device-manager.md) — the device manager.** Design: the
tree, the matcher, and the supervisor. Tree structure lives in the manager,
authority stays in the kernel; bus drivers report what they see; drivers are
restarted through the lifecycle vocabulary — the plan that turns
[resilience.md](resilience.md)'s restart goal into increments.
18. **[input.md](input.md) — the input module.** Broadcasting input events (keyboard,
mouse, joystick): why a synchronous rendezvous can't fan out to many listeners, the
asynchronous `ipc_send` primitive built to fix it, and the per-device subscribe/publish
service layered on top.
19. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
Start with the north star:
@@ -147,7 +166,7 @@ A sub-project's extra files are reached through the module, never as separate pa
```
system/ → /system danos's own internals (the self-representation)
boot-handoff.zig the loader↔kernel contract (the `boot-handoff` module)
abi.zig the core kernel↔user ABI (the `abi` module)
abi.zig the private kernel↔runtime syscall ABI (the `abi` module)
parameters.zig initial-ramdisk.zig shared contracts
kernel/ IPC, memory, scheduling, the private syscall dispatch
architecture/x86_64/ the `architecture` module (never named by generic code)
@@ -180,7 +199,7 @@ appears in the private-ABI path.
| Boot methods (one per way of booting the kernel) | `boot/` — `efi.zig` (UEFI) → `BOOTX64.efi` |
| Kernel entry, panic, bring-up | `system/kernel/kernel.zig` |
| Loader↔kernel handoff (`BootInfo`, `Framebuffer`, `MemoryMap`, VM layout) | `system/boot-handoff.zig` |
| Core kernel↔user ABI (`SystemCall`, mmap prot flags, `page_size`) | `system/abi.zig` |
| Private kernel↔runtime syscall ABI (`SystemCall`, mmap prot flags, `page_size`) — the runtime speaks it, not apps | `system/abi.zig` |
| Device wire types (`DeviceDescriptor`, `DeviceClass`, …) | `system/devices/device-abi.zig` |
| Physical frame allocator | `system/kernel/pmm.zig` |
| Kernel heap (`std.mem.Allocator`) | `system/kernel/heap.zig` |
+16
View File
@@ -150,3 +150,19 @@ input output" in code — that expansion is what the acronym *is for*. But `msg`
test for "is this an abbreviation I must expand" is simply: *is there a longer word this
is a clipped form of?* If yes, write the word. If it's an initialism standing in for a
phrase, leave it.
## Zen of Zig
* Communicate intent precisely.
* Edge cases matter.
* Favor reading code over writing code.
* Only one obvious way to do things.
* Runtime crashes are better than bugs.
* Compile errors are better than runtime crashes.
* Incremental improvements.
* Avoid local maximums.
* Reduce the amount one must remember.
* Focus on code rather than style.
* Resource allocation may fail; resource deallocation must succeed.
* Memory is a resource.
* Together we serve the users.
+25 -21
View File
@@ -64,14 +64,15 @@ already provide — it claims its device, maps its registers with `mmio_map`, an
on `replyWait` for either an interrupt or a client request. `system/drivers/hpet/hpet.zig` is already
that program, minus the client half.
The obstacle is not the file type, it is which hardware a ring-3 driver can actually
drive. Port I/O is unavailable to user space — the TSS I/O permission bitmap is absent
and IOPL is never raised — so `in`/`out` from a driver is a #GP. That excludes the
16550 UART at `0x3F8` and PS/2 at `0x60`/`0x64`, which is to say it excludes the
obvious implementations of `/dev/tty`, `/dev/ttyS0` and a keyboard node. Until either
port I/O grants or a memory-mapped UART exist, serial output stays a kernel service
reached through the `write` system call rather than a file. A memory-mapped device such
as the framebuffer has no such problem and is the more likely first real entry here.
The obstacle was never the file type; it is which hardware a ring-3 driver can reach.
Direct `in`/`out` from user space is still a #GP (no TSS I/O bitmap, IOPL never raised),
but a driver no longer needs it: **`io_read`/`io_write`** grant port access the same way
`mmio_map` grants memory — gated by `device_claim` and the device's discovered `io_port`
resource. So the 16550 UART at `0x3F8` and the PS/2 controller at `0x60`/`0x64` (and thus
`/dev/ttyS0` and a keyboard node) are now writable as ordinary ring-3 drivers; the
low-rate legacy hardware that needs port I/O is fine with a syscall per access. A
memory-mapped device such as the framebuffer, needing no port I/O at all, remains the
easiest first entry.
### Block devices
@@ -79,20 +80,23 @@ A block device is addressed in fixed-size blocks and, unlike a character device,
layer above is free to buffer, reorder, coalesce and retry requests against it. Disks
and other persistent storage are the whole population of this class.
**danos cannot host a block driver at all today,** and the reason is worth stating
plainly because it is not a matter of unwritten code. Every storage controller worth
naming is a bus master: it is programmed by handing it the physical address of a
descriptor ring and left to read and write memory on its own. A ring-3 driver cannot
build such a ring, because `mmap` returns writeback-cached, physically discontiguous
pages and never discloses their physical address. Nor should it be allowed to: a device
programmed with an arbitrary physical address writes to arbitrary physical memory, and
page tables do not sit between a device and RAM — an IOMMU does. Granting a DMA-capable
device to a driver process, with no IOMMU programmed, is equivalent to granting ring 0,
which would forfeit the isolation that motivates user-space drivers in the first place.
A block driver is now **writable, but not yet memory-safe.** Every storage controller
worth naming is a bus master: it is programmed by handing it the physical address of a
descriptor ring and left to read and write memory on its own. That ring is exactly what
**`dma_alloc`** now provides — physically contiguous, pinned, uncacheable, with its
physical address disclosed — and **`/lib/mmio`**'s barriers order the descriptor writes
against the doorbell, and **`msi_bind`** delivers completions. So an AHCI or NVMe driver
can be written today (the M14/M15 work in [driver-model.md](driver-model.md); the earlier
"cannot host a block driver at all" is no longer true).
Block devices therefore wait on DMA-capable memory, memory barriers, and VT-d/DMAR —
the M14–M16 work in [driver-model.md](driver-model.md). A ramdisk over the initial ramdisk is
the one block-shaped thing implementable now, and it needs no driver process.
What is *not* yet true is that it is safe. A device programmed with an arbitrary physical
address writes to arbitrary physical memory, and page tables do not sit between a device
and RAM — an IOMMU does. The IOMMU is now *detected* (M16), but no translation domains
are programmed, so granting a DMA-capable device to a driver process is still equivalent
to granting ring 0. Until per-device domains confine a driver's DMA to the buffers it
`dma_alloc`'d, a block driver works but forfeits the isolation that motivates user-space
drivers — enforcement is the next step, and lands with that first driver. A ramdisk over
the initial ramdisk remains the one block-shaped thing that needs no driver process at all.
### Pseudo-devices
+6 -5
View File
@@ -156,10 +156,11 @@ spinning in unrelated code — is the whole mechanism working end to end.
## What's next (not done here)
- **The keyboard**: the PS/2 controller is port-mapped (`0x60`/`0x64`), and ring 3
has no port I/O yet, so the first *input* device is blocked on either an I/O
permission bitmap or `io_in`/`io_out` syscalls ([drivers.md](drivers.md)).
- **MSI/MSI-X**: per-device vectors, edge-triggered and unshared, which retire the
I/O APIC's mask/ack cycle and its 24-GSI ceiling.
- **The keyboard**: the PS/2 controller is port-mapped (`0x60`/`0x64`), and port I/O is
now available to ring 3 via the claim-gated `io_read`/`io_write` syscalls
([drivers.md](drivers.md)) — so the first *input* device is unblocked; it just needs
writing (claim the controller, `irq_bind` GSI 1, read scancodes from `0x60`).
- **MSI-X**: `msi_bind` gives one per-device edge-triggered vector (M15); MSI-X's
multi-vector table (many queues per device, e.g. NVMe) is the remaining extension.
- **The LAPIC's own page** is still mapped writeback-cacheable like the rest of the
identity map. QEMU tolerates it; real hardware wants it uncacheable.
+144
View File
@@ -0,0 +1,144 @@
# The device manager
**Status: design.** The primitives this builds on are real ([process-management.md](process-management.md):
spawn/supervise/kill/exit-notification; [driver-model.md](driver-model.md): the device
table as a capability system; [drivers.md](drivers.md): claim/map/IRQ), and the first
per-device driver spawn works (the device manager matches the xHCI controller by PCI
class and spawns `usb-xhci-bus` with the device id as argv[1]). This document designs
the rest: the device manager as **the tree, the matcher, and the supervisor** — the
policy process that turns [resilience.md](resilience.md)'s restart goal into practice
for drivers.
How processes stop, reload, and report their deaths is deliberately **not** in this
document: that is the universal lifecycle every danos process speaks —
[process-lifecycle.md](process-lifecycle.md), signals over IPC and the stable
`runtime.process` interface. The device manager is that design's first serious
customer, not its owner. Its own protocol contains nothing lifecycle-shaped; a
driver is stopped, health-checked, and buried exactly like any other process.
## The tree: structure in the manager, authority in the kernel
The device tree is two things fused: *information* (what exists, how it nests) and
*authority* (a descriptor is a licence to map physical memory). They separate:
- The **kernel keeps the capability system** — device, I/O-port, and interrupt
claims, resource containment on `device_register`, the
`mmio_map`/`irq_bind`/`msi_bind` gates — and **cleans all of it up when a process
dies** (settled; it is increment 1 of
[process-lifecycle.md](process-lifecycle.md)). The three invariants in
[driver-model.md](driver-model.md) stay exactly where they are. A device manager
that could mint MMIO mappings by its own say-so would be a second kernel, and a
buggy one would un-earn everything the microkernel bought.
- The **device manager owns the tree as data** — identity, topology, naming, driver
matching, hotplug events, and being the one process everything else asks about
devices. Firmware discovery seeds it (today via the kernel's snapshot); **bus
drivers grow it** by reporting what they see; applications query and watch it.
`device_enumerate` fades to a manager-internal (then deleted) seam.
Long-term, discovery itself leaves the kernel — but not *into* the manager. PCI
enumeration is a **pci-bus driver**: the manager spawns it against the host bridge
(already a device with the ECAM window as a resource), it scans, it reports functions
like any bus reports children. ACPI becomes an **acpi service** that interprets the
tables and reports the namespace. The manager only orchestrates and merges. Moving
AML interpretation out of ring 0 is its own project on its own track; nothing here
depends on when it lands.
## The protocol
A `device-manager-protocol` module (the vfs-protocol pattern): extern-struct
messages, a version in the handshake, reserved fields everywhere. The manager is a
well-known endpoint (`ipc.register(.device_manager)`); the badge tells it who is
talking; the same endpoint receives its children's exit notifications — one loop,
one world.
| Direction | Message | Purpose |
|---|---|---|
| driver → manager | `hello { version, role, device_id }` | confirms the argv assignment, starts the deadline clock |
| bus → manager | `child_added { parent, identity, resources }` | one node the bus discovered |
| bus → manager | `child_removed { id }` | unplug, or the bus lost it |
| app → manager | `enumerate` | snapshot of the tree (read-only) |
| app → manager | `subscribe` | receive published add/remove events |
`hello` is the one deadline the manager enforces itself: spawned and silent past the
deadline means wrong binary, wrong protocol version, or wedged before main — apply
the stop sequence and the restart policy. Everything else lifecycle-shaped
(terminate, the common `ping` liveness call, exit reasons) arrives through
[process-lifecycle.md](process-lifecycle.md)'s vocabulary, not this protocol.
Assignment stays argv (`usb-xhci-bus <device id>`) for now — simple, and it works.
The step after `hello` exists is delegation: the manager claims (or is granted) the
devices and passes the claim to the driver over IPC (the M13 capability-transfer
mechanism), replacing first-come-first-served `device_claim` with policy. Identity in
`child_added` is per-bus: PCI children carry the class triple (`pci_class`, as the
xHCI match already uses); USB children carry the (class, subclass, protocol) triple
from usb-ids.zig — each bus's native language, decoded by the shared ids modules.
## Supervision and restart
Every driver is spawned with the manager's exit endpoint (`spawnSupervised` — built).
On a death notification:
1. **Read the reason** ([process-lifecycle.md](process-lifecycle.md) increment 2).
Clean exit → it meant to; don't restart. Fault or missed `hello` deadline →
restart with **backoff**, and a crash-loop cap (three fast deaths → mark failed,
stop respawning, log loudly; a later `reload` to the manager can retry).
2. **Prune the subtree** the dead bus driver reported. Its children describe
protocol state (xHCI slot ids, transfer rings) that died with the process;
keeping the nodes would be keeping a lie. Watchers receive `child_removed` — the
input service losing, then regaining, a keyboard is the *honest* description of
what happened. The restarted instance rediscovers and re-reports.
3. **The claim is already free** because the kernel released it at death — the
restarted instance claims the same controller and comes up.
Who supervises the supervisor: **init** (PID 1), which already supervises the
services it starts. If the manager dies, drivers keep running (they hold their
claims; the kernel doesn't care who their supervisor was — though their exit
notifications now dangle harmlessly). The restarted manager re-learns the world:
kernel snapshot, then a re-`hello` round — drivers answer a broadcast or are stopped
and respawned. Full state handoff is deliberately not attempted.
## Thin drivers, class protocols
The [driver-model.md](driver-model.md) three-shape split, restated as processes:
- A **bus driver** (usb-xhci-bus) owns its controller — claim, MMIO, IRQ/MSI, DMA
rings — and offers a *transfer* protocol ("submit a control transfer to device N",
built from the usb-abi request constructors) plus tree reports to the manager.
- A **class driver** (usb-hid, usb-storage) owns nothing: it is matched to a reported
child by its identity triple, speaks the bus's transfer protocol downward and its
service's protocol upward — HID reports to the input service, blocks to the block
service. It works unchanged over any controller.
- **Services** (input, display, block) aggregate class drivers and face applications.
Each arrow is a protocol module. The manager routes none of the data plane — it
introduces the parties (matching), supervises them (lifecycle), and gets out of the
way.
## Increments
Increments 1–4 are the lifecycle prerequisites and live in
[process-lifecycle.md](process-lifecycle.md) (claim cleanup on death, exit reasons,
published exit events, signals + `runtime.process`). On top of those:
5. **device-manager-protocol**: `hello`, supervised spawn with restart policy;
usb-xhci-bus becomes the first conforming driver.
6. **Tree reports**: `child_added`/`child_removed`; the manager mirrors; xHCI reports
the mouse and keyboard QEMU already hangs off it.
7. **App surface**: `enumerate`/`subscribe` over IPC; `device_enumerate` retreats
to a manager-internal seam.
8. **Discovery migration**: pci-bus driver first, acpi service second, kernel scan
retired last. (AML-in-user-space is its own track.)
## Settled questions (2026-07-12)
- **Stateful buses**: pruning the subtree on bus-driver death is right for USB. A
future storage bus with in-flight writes wants drain-before-terminate — which is
exactly the `deadline_ms` parameter `stop()` already has; a per-driver deadline
is one value in the manager's policy table when such a bus arrives. No design
change.
- **Manager death**: drivers survive the manager; the restarted manager re-learns
the world (above). Checkpointing driver state with the manager is deferred until
something demonstrates the need.
- **Matching stays code until the third bus.** `driverFor`/`pciDriverFor` are
honest at two bus types; the third triggers the manifest (a driver declares what
it binds: a PCI class triple, a USB class triple, an ACPI `_HID`).
+51 -5
View File
@@ -138,8 +138,37 @@ If a class driver needs `mmio`, it has become an HCD and should be one.
table fails `-ENOSPC` and does not half-deliver. This is the "open" primitive — a bus
driver mints a per-device endpoint and hands it to a class driver. The runtime exposes
`callCap` and `replyWait(..., send_cap)`; no class driver consumes it yet.
- **`system_spawn`** — a user-space supervisor starts a driver: `system_spawn(name)`
loads a binary bundled in the initial-ramdisk as a fresh ring-3 process. This is what
- **M14** — DMA memory + the memory-ordering layer. `/lib/mmio` gives drivers typed
volatile access and `mb`/`rmb`/`wmb` (per-arch); `dma_alloc`/`dma_free` grant
physically-contiguous, pinned, uncacheable, reclaim-on-teardown buffers with the
physical address exposed (`pmm.allocContiguous`, a DMA arena, `mapUserDmaInto`).
`dma_below_4g` caps the address for legacy engines; `dma_write_combining` is accepted
but falls back to coherent until PAT is programmed. hpet is refactored onto `/lib/mmio`;
no DMA driver consumes `dma_alloc` yet.
- **M15** — interrupts for PCI devices, the MSI half. Discovery now gives every PCI
function its 4 KiB ECAM config space as resource 0 (unblocking the capability walk
with no new syscall), and `msi_bind(device_id, endpoint) -> address, data` allocates a
per-device edge-triggered vector, delivered as an IPC notification with no mask and no
ack cycle. Legacy INTx (`_PRT` parsing + shared lines) is deliberately skipped — MSI
is the real answer. QEMU's HPET has no MSI, so delivery is proven with a self-IPI; the
first PCI driver is the first real consumer.
- **Port I/O** — `io_read`/`io_write(device_id, resource_index, offset, width[, value])`:
a claimed device's `io_port` resource lets a driver read/write its ports, gated exactly
like `mmio_map` gates memory (direct ring-3 `in`/`out` stays a #GP). This is what makes
a PS/2 or 16550 driver possible; the low-rate legacy hardware that needs it is fine with
a syscall per access. `io_port` resources were recorded by discovery and ignored — now
they're used.
- **M16 (detection)** — the IOMMU is now *found*: discovery parses the ACPI DMAR table,
maps the first VT-d unit, and reads its version + capabilities (`iommu_present` in the
platform info). This is detection only — **no translation domains are programmed, so
DMA is still unprotected** (the caveat below). Enforcement lands with the first DMA
driver, which is what there is to protect and test against. Proven in the `iommu` test,
booted with an emulated `intel-iommu`.
- **`system_spawn`** — a user-space supervisor starts a driver:
`system_spawn(name, arguments)` loads a binary bundled in the initial-ramdisk as a
fresh ring-3 process; `name` becomes the child's argv[0] and the optional
NUL-separated `arguments` blob its argv[1..], delivered on a SysV entry stack
([sysv.md](sysv.md)). This is what
turned the device manager from "log the match" into "run the driver": the kernel now
spawns only `init`, `init` spawns the services, and the **device-manager** discovers
the hardware and spawns each driver ([drivers.md](drivers.md)). Ungated for now — a
@@ -194,7 +223,12 @@ const dev_ep = ipc.callCap(h, // ... mint a per-device endpoint,
// now dev_ep is a private channel to that one device
```
## M14 — DMA memory and the memory-ordering contract, for HCDs
## M14 — DMA memory and the memory-ordering contract, for HCDs ✅ done
*Implemented: `/lib/mmio` (typed volatile access + `mb`/`rmb`/`wmb`, per-arch) and
`dma_alloc`/`dma_free` (contiguous, pinned, uncacheable, reclaim-on-teardown, physical
address exposed). `dma_write_combining` still falls back to coherent — real WC needs
PAT, a small follow-up. The rest of this section is the original design note.*
**The blocker.** An HCD is a DMA-engine programmer. It needs a descriptor ring the
device can read, which means memory that is (a) physically contiguous, (b) at a
@@ -259,7 +293,12 @@ condition. Build the abstraction while there is one caller to fix.
(Zig note: `@fence` was **removed in 0.16**. Use `@atomicRmw(..., .seq_cst)` for a full
barrier, or per-arch inline asm — which is what `library/mmio.zig` should hide.)
## M15 — interrupts for PCI devices
## M15 — interrupts for PCI devices ✅ done (MSI)
*Implemented the MSI half: ECAM config space per PCI function (resource 0) and
`msi_bind` (per-device edge-triggered vector, delivered as a notification). Legacy INTx
`_PRT` parsing is skipped on purpose. `msi_bind` returns (address, data) as two values
rather than an out-struct. The rest of this section is the original design note.*
**The blocker, and it's a hard one.** No PCI device can take an interrupt today.
[`addBars`](system/devices/acpi.zig) records `.memory` and `.io_port` BARs and never an
@@ -290,7 +329,14 @@ capability walk (MSI, MSI-X, PCIe extended caps) without any new syscall.
Note QEMU's HPET reports `Tn_FSB_INT_DEL_CAP = 0` — no MSI — so `hpet` can never
exercise this path. The first MSI driver will be the first PCI driver.
## M16 — the IOMMU, and the honest caveat
## M16 — the IOMMU, and the honest caveat ◑ detection done, enforcement pending
*The IOMMU is now detected (DMAR parsed, VT-d unit mapped and read — see the `iommu`
test), but **enforcement is not built**: no translation domains are programmed, so the
caveat below still holds in full. Detection can't be taken further usefully until there
is a DMA driver to protect and QEMU's `intel-iommu` to test the protection against —
building the per-device domains alongside that first driver is both the natural order
and the only way to verify them. The rest of this section is the original caveat.*
Everything above is capability-gated at the *CPU*. None of it is gated at the *device*.
A driver that can program a bus-mastering engine can make that device write to any
+24 -26
View File
@@ -37,8 +37,9 @@ kernel ──spawns──► init (PID 1) ──spawns──► device-manag
```
The kernel launches exactly one process — `init` — and hands it nothing but the raw
ability to start more (`system_spawn(name)`, which loads a binary bundled in the
initial-ramdisk as a fresh ring-3 process). Everything else is a user-space decision:
ability to start more (`system_spawn(name, arguments)`, which loads a binary bundled
in the initial-ramdisk as a fresh ring-3 process — `name` becoming its argv[0],
the optional arguments its argv[1..], on a SysV entry stack, see sysv.md). Everything else is a user-space decision:
- **init** ([system/services/init](system/services/init/init.zig)) is the **service
supervisor**. It spawns the system services danos brings up at boot — today `vfs` and
@@ -52,7 +53,8 @@ initial-ramdisk as a fresh ring-3 process). Everything else is a user-space deci
is a table (`driverFor`): today a static `timer → hpet` map; a fuller system reads
what each driver *binds* (a manifest under `/system/drivers`, or the driver
describing its own match).
3. **Spawn** — `system_spawn(driver_name)` starts the matched driver, which then claims
3. **Spawn** — `system_spawn(driver_name, arguments)` starts the matched driver (the
arguments can carry *which* device it matched), which then claims
its device and runs the event loop below.
So "how is a driver discovered and configured" has two halves: **discovery** is the
@@ -189,7 +191,7 @@ const hpet = findHpet(buf) orelse return; // device_enumerate, look for
// class=timer with memory + irq
_ = dev.claim(hpet.dev_id); // the capability
const base = dev.mmioMap(hpet.dev_id, hpet.mmio).?;
const endpoint = ipc.createEndpoint().?;
const endpoint = ipc.createIpcEndpoint().?;
// program the hardware over the mapping we were just handed
reg(base, 0x100).* = level | int_enb | (hpet.gsi << 9); // timer 0 config
@@ -266,26 +268,24 @@ controller drivers fit together.
Worth knowing before you write the second driver:
- **Ring 3 has no port I/O.** The TSS I/O permission bitmap is absent
(`tss.zig`: `iomap_base = @sizeOf(Tss)`), and IOPL is never raised, so `in`/`out`
from a driver is a #GP. That rules out a user-space 16550 UART (`0x3F8`), PS/2
(`0x60`/`0x64`), and legacy PCI config (`0xCF8`/`0xCFC`). Everything must be MMIO.
`io_port` resources are recorded by discovery and then ignored.
Several things this list used to warn about are now available (see
[driver-model.md](driver-model.md)): **port I/O** (`io_read`/`io_write`, claim-gated by
the device's `io_port` resource — direct ring-3 `in`/`out` is still a #GP, so a PS/2 or
16550 driver goes through these), **DMA memory** (`dma_alloc`: contiguous, pinned,
uncacheable, physical address exposed), and **memory barriers** (`/lib/mmio`'s
`mb`/`rmb`/`wmb`). What remains:
- **Page granularity.** `mmio_map` rounds to 4 KiB. Two devices sharing a page means
granting one grants the other. A `device_register`ed child's *resource* can be narrower
than a page, but its *mapping* can't.
- **No DMA memory.** `mmap` gives you writeback-cached, non-contiguous pages and never
tells you their physical address, so you cannot build a descriptor ring. Any driver
for a bus-mastering device is blocked on this.
- **No memory barriers.** There are none in the tree, and `volatile` is not one — it
won't stop the compiler sinking an ordinary store (your DMA descriptor) past a
volatile MMIO store (your doorbell). On x86 you mostly get away with it; on ARM you
will not. See [driver-model.md](driver-model.md#m14).
- **DMA is not contained.** A driver that can program a bus-mastering device can make
that device write to *any* physical address — page tables don't sit between a device
and RAM; an IOMMU does. Until VT-d/DMAR is programmed, `device_claim` on a DMA-capable
device is effectively equivalent to granting ring 0. This is the largest gap between
the design's promise and what it delivers.
and RAM; an IOMMU does. The IOMMU is now *detected* (M16), but no translation domains
are programmed, so `device_claim` on a DMA-capable device is still effectively
equivalent to granting ring 0. This is the largest gap between the design's promise and
what it delivers; enforcement lands with the first DMA driver.
- **No `dev_release`.** A claim is never dropped (only IRQ/MSI bindings are, on exit), so
a device stays owned for the life of its driver — which blocks restart.
- **One endpoint per GSI**, so shared legacy PCI INTx lines can't be split between two
drivers. MSI/MSI-X — one vector per device, edge-triggered, unshared — is the real
answer, and QEMU's HPET doesn't offer it (`Tn_FSB_INT_DEL_CAP = 0`).
@@ -341,14 +341,12 @@ Two companions cover what `hpet` can't, because it never exits:
## What's next (not done here)
The big ones — capability passing (class drivers), DMA + barriers and MSI (host
controller drivers), and the IOMMU — have proposed signatures in
[driver-model.md](driver-model.md). Smaller items:
The big driver-model pieces — capability passing (class drivers), DMA + barriers, MSI,
and IOMMU detection — are **now done** ([driver-model.md](driver-model.md), M13–M16), as
is **port I/O** (`io_read`/`io_write`, the claim-gated syscalls that make a PS/2 or 16550
driver possible). What's left is IOMMU *enforcement* (per-device domains — it waits on
the first DMA driver to protect and test against) and these smaller items:
- **Port I/O grants**, so a PS/2 or 16550 driver is possible: either a per-device TSS
I/O permission bitmap swapped on context switch, or `io_in`/`io_out` syscalls gated
by the same claim. The legacy devices that need it are all low-rate, so the syscall
is likely fast enough.
- **Releasing a claim.** There is no `dev_release`, and `devices_broker` never drops a claim on
exit — only IRQ bindings are released. A dead driver's device stays owned forever,
which blocks restart.
+161
View File
@@ -0,0 +1,161 @@
# The input module: broadcasting input events
A keyboard driver has one keystroke and *many* programs that might want it — a shell, a
window server, a logger. None of them owns the hardware, and the driver should not know
who is listening. So between the drivers and the listeners sits the **input service**
(`system/services/input/`): drivers **publish** events to it, programs **subscribe**, and
it fans each event out to every interested subscriber. It is an ordinary ring-3 process
reached over IPC, like the [VFS server](../system/services/vfs/vfs.zig) — no kernel knows
what a key is.
## One service, several device classes
The service carries three device classes today — **keyboard**, **mouse**, and
**joystick/gamepad** — and is built to take more
([protocol.zig](../system/services/input/protocol.zig)). Each class has its own typed
event:
- `KeyEvent` — `key_down`/`key_up` (physical make/break) and `key_press` (a character was
produced, carrying the Unicode scalar); plus a layout-independent `keycode` and a
`modifiers` bitmask.
- `MouseEvent` — relative `motion` (`dx`/`dy`), `button_down`/`button_up`, and `scroll`.
- `JoystickEvent` — `axis` moves (a signed value on a `control` index) and
`button_down`/`button_up`.
All three travel in one **`InputEvent` envelope** tagged with a `DeviceKind`, so the
fan-out is a single code path and a subscriber can take a mix of classes on one stream.
Decode an envelope with `asKeyboard()` / `asMouse()` / `asJoystick()` (each returns null
unless the tag matches). A subscriber names the classes it wants with a **`device_mask`**,
and the service routes each event only to subscribers whose mask includes its class — so a
mouse-only listener never wakes for keystrokes.
## Why this needed a new kernel primitive
The interesting part is delivery, and it runs straight into the shape of danos IPC.
[ipc.md](ipc.md) describes a **synchronous rendezvous**: a server holds exactly one
pending reply (`Task.ipc_client`) and *must* answer it on its next `replyWait`. Two
consequences decide the whole design:
1. **You cannot block N subscribers waiting for "the next event".** A server can hold only
one caller at a time, so the natural "subscriber calls `next_event()` and blocks" API
is impossible for more than one subscriber. Delivery therefore has to be **push** — the
service reaching out to subscribers — not pull.
2. **A synchronous push can hang the whole service.** If the service delivered with
`ipc_call`, it would block until each subscriber replied. `ipc_call` has no timeout, and
the kernel does **not** wake a caller parked on a *dead* peer's endpoint (it only fails a
peer that was mid-reply — see [process.zig](../system/kernel/process.zig)
`releaseTaskResourcesLocked`). One subscriber that exits mid-delivery would wedge input
for everyone. That is the opposite of the resilience the microkernel is for.
The fix is the asynchronous send that [ipc.md](ipc.md) had already earmarked as future
work ("asynchronous / buffered send … for notifications between servers"):
```
ipc_send(handle, message_ptr, message_len) -> 0 / -errno
```
`ipc_send` copies a small payload into the endpoint's **bounded queue** and wakes a
receiver, then returns immediately — it never blocks and so can never hang on a dead or
slow subscriber. The receiver picks it up through the same `replyWait` it already runs:
the wake arrives as a **buffered message** — `notify_badge_bit | notify_message_bit` set in
the badge (distinguishing it from a bare IRQ/child-exit notification), the sender's task id
in the low bits, and the payload in the receive buffer, with no reply owed. The queue holds
16 messages per endpoint; a full queue **drops the oldest**, because a buffered message is
discrete data, not a coalescing "level" like an interrupt. See
[ipc-synchronous.zig](../system/kernel/ipc-synchronous.zig) (`sendLocked`, `popPost`, and
the `replyWait` receive loop).
This is the async counterpart of `ipc_call`, and the input service is its first consumer.
## How the pieces fit
```
keyboard/mouse driver, input-source input service subscriber(s)
----------------------------------- ------------- -------------
connectSource(); loop: replyWait: subscribeKeyboard()/…All:
publishKeyboardEvent(k) ─ ipc_call ─▶ publish → broadcast: createIpcEndpoint()
publishMouseEvent(m) for each sub whose callCap(subscribe,
publishJoystickEvent(j) mask matches event.device: send_cap = ep,
ipc_send(sub_ep) ──────▶ device_mask)
reply ok loop: next()
subscribe → store {ep cap, └─ replyWait(ep)
task id, device_mask} → InputEvent
```
- A **subscriber** calls `input.subscribe(mask)` — or a typed helper: `subscribeKeyboard()`,
`subscribeMouse()`, `subscribeJoystick()` (one class, `next()` returns the decoded event),
or `subscribeAll()` (every class, `next()` returns a tagged `InputEvent`)
([library/runtime/input.zig](../library/runtime/input.zig)). It creates its own endpoint
and hands it to the service as a **capability** (M13 capability passing — the input
service is that feature's first real user), along with its `device_mask`. Then it loops on
`next()`, a `replyWait` on that endpoint returning each pushed event.
- A **source** (a keyboard, mouse, or joystick driver) calls `input.connectSource()` and the
method for its class: `publishKeyboardEvent`, `publishMouseEvent`, or
`publishJoystickEvent`. Publishing is a short synchronous `ipc_call` the service answers at
once; the service's own fan-out is asynchronous, so publishing never blocks on a slow
subscriber.
- The **service** ([input.zig](../system/services/input/input.zig)) keeps a small subscriber
table (endpoint handle + owning task id + `device_mask`). On `publish` it `ipc_send`s the
event to every subscriber whose mask includes the event's device class. On `subscribe` it
stores the passed capability and mask and, as housekeeping, prunes any slot whose owning
process has exited (checked against `process_enumerate`) — not for correctness (an async
send to an orphaned endpoint is harmless) but to reclaim the slot.
Publisher and subscriber must be **separate processes**: a single thread that both
published and serviced its own subscription would deadlock (its `publish` call blocks until
the service delivers to its endpoint, which only the same thread could receive).
## Status and follow-ups
- **The keyboard is real.** The `ps2-bus` driver owns PNP0303, which carries *both* the
0x60/0x64 ports and IRQ1, so reading the hardware lives in the bus, not in
[keyboard.zig](../system/drivers/ps2-bus/keyboard.zig): the bus binds IRQ1 and, on each
interrupt, drains port 0x60, routing every byte by the status register's
auxiliary-output bit to whichever child driver **attached** for that device (an
`AttachRequest` to the well-known `ps2_bus` service, carrying the child's endpoint as a
capability; the bytes then arrive as asynchronous `ForwardedByte` messages, so the IRQ
path never blocks on a child). The keyboard driver decodes the stream — scancode **set 2**,
what the keyboard sends with the 8042's legacy translation off, decoded by
[scancode.zig](../system/drivers/ps2-bus/scancode.zig) into USB HID usage keycodes with
make/break, typematic-repeat, and modifier tracking (host-tested under `zig build test`) —
and publishes real `key_down`/`key_press`/`key_up` events.
- **Keycode → character** is wired in: the keyboard driver fills a `key_press` event's
`character` through [`library/xkeyboard-config`](../library/xkeyboard-config/README.md)
(`xkb.map(layout, keycode, mods)` → keysym + Unicode character), synthesizing the ASCII
control characters for Enter/Tab/Backspace/Escape, whose keysyms map to no Unicode. The
layout defaults to `us`; the bus can pass another as the driver's argv[2] — the seam for
a future settings source.
- **The mouse is real too.** IRQ12 is enumerated on the auxiliary device's own ACPI node
(PNP0F13), so the bus claims that node alongside the controller and routes both IRQs to
its one endpoint, acking whichever line the notification's badge names.
[mouse.zig](../system/drivers/ps2-bus/mouse.zig) attaches the way the keyboard does and
assembles the forwarded bytes with
[mouse-packet.zig](../system/drivers/ps2-bus/mouse-packet.zig) (three-byte stream-mode
packets: sync/overflow handling, nine-bit movement, screen-convention `dy` — host-tested
under `zig build test`) into `button_down`/`button_up` transitions and `motion` events.
**Follow-up:** the IntelliMouse magic-knock for a scroll wheel (four-byte packets) and
`scroll` events. The hardware-free `input-source` still rotates through all three classes
synthetically (including a joystick, which has no driver yet) via the
`input.synthetic*Event` helpers.
- **Drop-oldest under overflow** is a defined loss; the 16-slot ring absorbs normal bursts.
Real backpressure/flow-control is future work.
- **`publish` is unauthenticated** — any process may publish, consistent with the current
bring-up trust model (see [driver-model.md](driver-model.md)). A source capability is
future work.
## Verifying it
The `input` case (`python3 test/qemu_test.py input`, in
[tests.zig](../system/kernel/tests.zig) `inputTest`) boots the real kernel and spawns the
service, the synthetic source (which cycles keyboard, mouse, and joystick events), and a
subscriber that took all three classes. It passes only when the subscriber heartbeats
`input-test: ok` — proof that an event travelled source → service → subscriber over IPC,
exercising `ipc_send`, capability-passing subscription, and per-device routing. Each
serial line names the class received, so the log shows all three arriving on one stream.
## See also
- [ipc.md](ipc.md) — the synchronous rendezvous and the notification path `ipc_send` extends.
- [syscall.md](syscall.md) — the system-call surface, including `ipc_send`.
- [driver-model.md](driver-model.md) — class drivers, capability passing (M13), the trust model.
+15 -4
View File
@@ -80,11 +80,22 @@ inline). `build.zig` adds `isr.s` to the arch module.
## Reporting a fault
`isr_common` calls `exceptionHandler`, which forwards to a swappable `on_fault`
hook. The generic kernel installs a reporter (`onException` in `main.zig`) that
prints, in red, the exception name and vector, the error code, the faulting RIP
hook. The generic kernel installs a reporter (`onException` in `kernel.zig`) that
prints the exception name and vector, the error code, the faulting RIP
and RSP, and — for a page fault (#PF, vector 14) — the faulting address from
**CR2**. Then it halts. There's no fault *recovery* yet, so every exception is
terminal; the point is that it's now **visible** instead of a silent reset.
**CR2**. What happens next depends on where the fault came from:
- **User mode (CPL 3): kill the process, keep the machine.** The kernel is intact
(the CPU trapped onto the task's kernel stack), so the faulting process is
killed — address space, IRQ bindings, and IPC handles reclaimed; a client it
owed a reply to is failed with `-EPEER` — and the core reschedules. A crashing
driver takes itself down, never the OS. This is fault recovery step 2 of
[resilience.md](resilience.md). NMI, double fault, and machine check are
excluded: they report machine trouble regardless of what was running.
- **Kernel mode: halt this core.** The trusted base itself is broken, so there is
nothing safe to kill; the fault is still *contained* to the core (an
application-processor fault leaves the rest of the system running), and the
report makes it **visible** instead of a silent reset.
The hook is set before `arch.init()` in `kmain`, so a fault during setup is still
caught.
+6 -1
View File
@@ -95,6 +95,11 @@ This is what makes a user-space driver possible at all, and it's the subject of
every capability is either well-known (the registry) or inherited — there's no way
to delegate one.
- **Asynchronous / buffered send** for the cases where a rendezvous is the wrong
shape (logging, notifications between servers).
shape (logging, notifications between servers). *Landed as `ipc_send`* — a
non-blocking post to an endpoint's bounded payload queue, delivered through
`reply_wait` as a buffered message (badge bit `notify_message_bit`). Built for, and
first used by, the [input service](input.md)'s keyboard-event broadcast, where a
synchronous push would let one dead subscriber hang the fan-out. A full queue drops
the oldest (discrete messages, not a coalescing level like the notification ring).
- **A bounded reply.** `MSG_MAX` is 256 bytes and the copy runs under the big kernel
lock; a bulk transfer wants shared pages, not a copy.
+161
View File
@@ -0,0 +1,161 @@
# M17–M18 execution plan: process lifecycle + device manager
The operational plan for building [process-lifecycle.md](process-lifecycle.md)
(M17) and [device-manager.md](device-manager.md) increments 5–7 (M18). Design is
settled in those documents; this file is the build order — one phase at a time,
each phase green before the next starts. Delete or archive this file when M18
lands.
**Definition of green, every phase:** `zig build` clean, `zig build test` clean,
`python3 test/qemu_test.py` passes (existing scenarios plus the phase's new one),
and the relevant design doc's "known gaps" / status lines updated. Commit per
green phase (no co-author trailers).
**Workflow (settled 2026-07-12):** work happens in a dedicated git worktree, on
feature branches cut from `main` — `feat/process-lifecycle` (M17.1–17.4),
`feat/device-manager` (M18.1), `feat/usb-xhci-bus` (M18.2–18.3). When a branch's
phases are all green it is **auto-merged into `main`**; branches are kept after
merge, not deleted. Merges and branches are pushed to origin. Phase 0 (once):
commit the design docs, merge the outstanding `feat/usb` work into `main`, and
run the existing QEMU suite green before any new work starts.
**Numbering note:** continues the milestone sequence (driver track ended at M16).
---
## M17.1 — the kernel releases a dead process's claims
The cleanup half of iron rule 1; the prerequisite for every restart story.
- `system/kernel/devices-broker.zig`: `releaseAllOwnedBy(owner: u32)` — clear
every `claimed[]` slot holding this task id.
- `system/kernel/process.zig`: call it from the reap path, alongside the existing
IRQ-binding release (the ordering comment there says why IRQs go first — claims
slot in after them, before the exit notification).
- MSI vectors: find where `msi_bind` records per-device vectors (interrupts
module) and release those by owner in the same pass.
- Docs: remove the claims bullet from process-management.md "Known gaps".
**Test:** new QEMU scenario `claim-release` — a test child claims an unclaimed
device, is killed, is respawned, and claims the same device again successfully;
assert both claims in the serial log. Kernel-side unit coverage in
`system/kernel/tests.zig` for `releaseAllOwnedBy` (claim two devices as two owners,
release one owner, verify exactly its claims freed).
## M17.2 — exit reasons
- `system/abi.zig`: `ExitReason` (exited, aborted, segmentation_fault,
illegal_instruction, arithmetic_fault, killed).
- Kernel: record the reason at every death site — clean exit path, each fault
class in `onException`, the kill path. Bounded recent-exits table (ids are never
reused, so a small ring keyed by id is enough).
- New system call `process_exit_reason(id)` — supervisor-gated, like kill; returns
the recorded reason or `-ESRCH` once evicted.
- `library/runtime/process.zig`: `ExitReason` + `exitReason(id: u32)`.
- Docs: remove the no-exit-status bullet from process-management.md.
**Test:** extend the `supervision` scenario — three children: one exits cleanly,
one faults (the fault-recovery pattern), one is killed; the supervisor asserts all
three reasons.
## M17.3 — published exit events
- Kernel: bounded subscriber table (endpoints); new system call
`process_subscribe(endpoint)` (ungated, like `process_enumerate`); every death
posts `notify_exit_bit | id` to each subscriber — the same post the supervisor
path already uses.
- `library/runtime/process.zig`: `subscribeExits(endpoint)`.
- VFS becomes the first subscriber: on an exit event, release every handle keyed
by that task id (badges already are task ids). Log the release.
- Docs: note the convention in ipc.md (exit events reuse the exit-notification
badge encoding).
**Test:** new QEMU scenario `vfs-client-death` — a client opens a file and is
killed without closing; assert the VFS logs the handle release and its open-handle
count returns to baseline.
## M17.4 — signals and the service harness
- Kernel: per-task pending mask + bound endpoint; system calls
`signal_bind(endpoint)` and `process_signal(id, signal)` (supervisor-or-self
gated); delivery posts `notify_signal_bit | pending mask`, coalescing; pending
signals with no bound endpoint pend silently.
- `library/runtime/process.zig`: `Signal`, `SignalSet`, `bindSignals`,
`signalsFrom`, `sendSignal`, `stop(id, deadline_ms)` (terminate → wait for exit
notification → kill). Implement `terminate`, `reload`, `user_1`, `user_2`;
`interrupt`/`quit` are enum members with no sender yet; `alarm` stays unbuilt.
- Kernel: **one-shot timer notifications** — `timer_bind(endpoint, ms)` posts a
notification badge when the deadline lands (IRQ-as-IPC again, on the timer
wheel `sleep` already uses). This is the missing timed-wait primitive:
`replyWait` blocks forever and `sleep` blocks the whole process, but `stop()`'s
escalation, the device manager's `hello` deadline (M18.1), and restart backoff
all need a deadline while staying responsive. It is also the mechanism `alarm`
gets for free later.
- New `library/runtime/service.zig`: the harness — `run(callbacks)` owning the
replyWait loop, folding protocol messages, signals, and child-exit notifications
into `init` / `on_message` / `on_reload` / `on_terminate`; answers the common
`ping` automatically. Define the reserved `ping` request encoding here and
document it in ipc.md (one obvious encoding; smallest that cannot collide with
existing protocols).
- Convert one existing service (input-source or hpet) to the harness as proof it
subtracts code rather than adding it.
**Test:** extend `supervision` — a harness-built child: `sendSignal(reload)`
observed in its log, `ping` answered, `stop()` produces a clean exit with reason
`exited`; a second child that ignores signals (no bind) is killed by `stop()`'s
deadline with reason `killed`.
## M18.1 — device-manager protocol: hello + restart policy
- New `system/services/device-manager/device-manager-protocol.zig` module
(vfs-protocol pattern): `hello { version, role, device_id }`; version constant;
reserved fields.
- Device manager: register the `.device_manager` endpoint; spawn drivers with its
exit endpoint; enforce the hello deadline; restart policy — backoff, crash-loop
cap (three fast deaths → mark failed, log, stop), reasons from M17.2 deciding
restart vs not.
- usb-xhci-bus: adopt the harness + send hello. hpet/ps2-bus follow only if the
conversion is mechanical; otherwise they keep working unconverted (the manager
only enforces hello on drivers spawned with an assignment).
- build.zig: test-loop entry for the protocol module if it grows pure logic.
**Test:** new QEMU scenario `driver-restart` — the xHCI driver takes a test-only
argv flag to fault after hello on its first run; assert: fault, exit reason
recorded, manager respawns with backoff, second run claims the controller
(M17.1) and hellos clean. Assert the crash-loop cap by a driver that always
faults (a tiny test driver, not xhci).
## M18.2 — bus tree reports
- Protocol: `child_added { parent, identity, resources }` / `child_removed { id }`.
- usb-xhci-bus: bring-up to **port scan only** — map the MMIO window (claimed in
M16-era work), controller reset/start per xHCI spec, walk the port registers,
report one `child_added` per connected port with speed + port number as
identity. **No transfer rings, no descriptors** — reading device/interface
descriptors (and therefore USB class triples for matching) is the follow-on USB
track, not this plan.
- Device manager: mirror reports into its tree; prune the subtree (emitting
`child_removed`) when a bus driver dies; assert re-report on restart.
**Test:** QEMU already attaches usb-kbd + usb-mouse on xhci.0 — assert two
`child_added` events reach the manager and appear in its tree dump; kill the
driver, assert two `child_removed` then two fresh `child_added` after respawn.
## M18.3 — the application surface
- Protocol: `enumerate` (tree snapshot) + `subscribe` (published add/remove
events, input-service pattern).
- A small client (`device-list`, the `ps` analog) exercising both; the manager
becomes the one answer to "what devices exist" for user space.
`device_enumerate` stays for drivers/kernel seeding — its retreat is tied to the
discovery migration, out of this plan.
**Test:** QEMU scenario — `device-list` shows the tree including USB children;
during a driver restart the subscribing client logs remove + add events.
---
**Explicitly out of scope** (own tracks, after M18): discovery migration (pci-bus
driver, acpi service, retiring the kernel scan), USB control transfers +
descriptors + class-driver matching, the musl layer, `interrupt`/`quit` senders
(needs a console), job control.
+321
View File
@@ -0,0 +1,321 @@
# Process lifecycle: signals over IPC
**Status: design.** The primitives underneath are built ([process-management.md](process-management.md):
spawn, the supervision link, kill, child-exit notifications); this document designs
the layer above them — the standard vocabulary a danos process speaks about its own
life, and the stable `runtime.process` interface that carries it. Nothing here is
device- or driver-specific: a driver, the VFS, and a user application all stop,
reload, and die the same way. The device manager is simply this design's first
serious customer ([device-manager.md](device-manager.md)).
**"POSIX" in this document means the concepts, never the letter of the standard.**
danos borrows the ideas and the hard-won lessons (what SIGTERM *means*, why SIGPIPE
was a mistake) without inheriting the mechanism, the API, or the names. The naming
rule is danos's own and it is strict: plain words that communicate intent
(`terminate`, `reload`, `exited`) and the IPC vocabulary the system already speaks
(`bind`, `subscribe`, `publish`, `endpoint`) — never `SIG*`, never a second word for
a concept that already has one. Literal POSIX arrives later and lives elsewhere: a
**musl-based C layer** (growing out of library/posix) that wires C programs to the
danos runtime — musl's syscall surface retargeted at danos system calls and IPC
protocols (files onto the VFS protocol, `sigaction`/`wait` onto this lifecycle,
sockets onto whatever networking becomes). Ported programs see POSIX; the system
underneath never does.
## Why a standard vocabulary
A supervisor can only manage processes it has never heard of if "please exit" means
the same thing to all of them. That is the one thing POSIX signals got deeply right:
`SIGTERM` means the same thing to nginx and to a five-line script, which is why
process supervision on Unix (init systems, container runtimes) is possible at all.
danos wants that property from day one, because supervision-and-restart is the
system's core motivation ([resilience.md](resilience.md)).
What POSIX got wrong — for a system like this — is the **delivery mechanism**:
asynchronous control-flow hijack. A Unix handler runs on a stolen stack at an
arbitrary instruction boundary, which is why the async-signal-safe function list
exists, why `errno` must be saved, and why the canonical signal bug is a SIGTERM
handler innocently calling `printf` mid-`malloc`. That entire bug class comes from
the mechanism, not the vocabulary, and none of it is worth importing.
A microkernel already has the right channel: **a signal is a message.** QNX delivers
POSIX signals over its message passing; seL4 has notification objects; Erlang turned
"death is a message to whoever linked" into a reliability philosophy. danos has
already done it once without naming it: a child's death arrives as a notification
badge on the supervisor's endpoint — the microkernel's SIGCHLD, the IRQ-as-IPC
pattern reused. Signals are the same pattern reused a third time.
## The mechanism
- **`signal_bind(endpoint)`** — a process nominates the endpoint its signals arrive
on, exactly as `irq_bind` nominates where a device's interrupts land. The runtime
does this at startup for any program that opts in.
- **`process_signal(id, signal)`** — posts the signal as an asynchronous
notification to the target's bound endpoint: badge = `notify_badge_bit |
notify_signal_bit | pending signals`. Non-blocking for the sender, always.
- **Pending signals coalesce** in a per-process bitmask until the target next waits
— exactly like interrupt notifications, and exactly POSIX's own semantics for
non-realtime signals (two pending SIGTERMs are one SIGTERM). The bitmask *is* the
design: signals carry no payload. Anything with a payload is a protocol message.
- **Authority**: the supervisor may signal its children — the same link that is
already the kill authority. A process may signal itself. Anything broader waits
for transferable process handles.
- **No binding, no problem**: a process that never calls `signal_bind` is not
broken — its signals pend unread and only `process_kill` works on it. Simple
programs stay simple; the vocabulary is opt-in, the kill authority is not.
Because delivery is a message into the process's own event loop, there is no
async-signal-safe list in danos: a handler is ordinary code running at a point the
process chose. The bug class is gone by construction, not by discipline.
## The vocabulary: POSIX.1-1990, sorted honestly
The full 1990 set, and what each becomes. Two intrinsically problematic cases get a
defense below the table.
| POSIX.1-1990 | danos disposition | Notes |
|---|---|---|
| SIGTERM | signal `terminate` | finish up and exit; the supervisor's polite half |
| SIGHUP | signal `reload` | re-read configuration / re-scan |
| SIGINT | signal `interrupt` | interactive interrupt; meaningful once a console can send it, in the vocabulary now so numbering is stable |
| SIGQUIT | signal `quit` | as SIGINT, without the core-dump baggage |
| SIGALRM | signal `alarm` | timer expiry as a message; the Unix SIGALRM+`longjmp` timeout hacks are impossible here. In the vocabulary, unbuilt: no consumer yet, and when one appears it is runtime sugar over the existing timer — zero kernel work |
| SIGUSR1, SIGUSR2 | signals `user_1`, `user_2` | service-defined |
| SIGCHLD | **already exists** — the exit notification | the badge carries the child id, dodging the classic coalescing bug (Unix code must loop `waitpid`) |
| SIGKILL | `process_kill` — kernel mechanism | its definition is "cannot be handled"; it was never really a signal |
| SIGABRT | exit reason `abort` | `abort()` is synchronous self-termination, not an event |
| SIGSEGV, SIGILL, SIGFPE | exit reasons, **never delivered** | see below |
| SIGPIPE | **an error return**, not a signal | see below |
| SIGSTOP, SIGTSTP, SIGTTIN, SIGTTOU, SIGCONT | deferred | job control needs terminals, sessions, and process groups; stop/continue is scheduler territory |
**The fault signals (SIGSEGV, SIGILL, SIGFPE) are intrinsically wrong for messages.**
They are *synchronous* — raised at a specific faulting instruction, not "sometime
soon". A message cannot be delivered to a process whose next instruction re-faults;
it never reaches its event loop to read it. POSIX only makes fault handlers "work"
via the async hijack (run the handler *instead of* the instruction), and even there,
returning from a SIGSEGV handler without curing the cause is undefined behavior.
danos's architecture already has the better answer: fault → the kernel kills the
process ([resilience.md](resilience.md) step 2, built) → the supervisor reads the
reason → restart. Recovery is restart, not a handler. This is also truer to the 1990
standard than handling is: the standard's default action for all three was
"terminate the process".
**SIGPIPE deserves special contempt.** Its default kills a process that writes to a
closed pipe — which is why "the whole server died because one client disconnected"
is roughly every network daemon's first production bug, and why every mature codebase
contains the same fix: ignore SIGPIPE, handle the `EPIPE` error return. danos made
the right choice natively already — a reply owed to a dead peer fails with `-EPEER`.
Errors from operations are error returns from those operations. The posix layer can
synthesize SIGPIPE for ported code that expects it.
### Statements, not questions
A signal and a protocol message both travel over IPC — the difference is the
**contract**, not the transport. danos IPC has two primitives, both already in
daily use: the **asynchronous notification** (a badge — bits that coalesce into a
pending mask; the sender never blocks; no payload, *no reply path*; how IRQs and
exit events arrive) and the **synchronous call** (a rendezvous — payload both
ways, the caller waits for the reply; how VFS requests work). A signal is the
first kind: a *statement*. `terminate` wants no reply — the exit notification is
its acknowledgement.
A health probe is the second kind: a *question*, worthless without its answer —
and the answer's absence within a deadline is the very thing being measured.
Asked as a signal it has no reply channel (a coalescing bit can't carry an answer,
and the authority rule forbids a child signalling its supervisor back); asked as a
call, the timeout-is-the-diagnosis semantics come free. So there is no `health`
signal. Liveness is the common **`ping`**: a reserved request every harness-run
service answers automatically on its main endpoint — still free for the service
author, still one obvious way — and a supervisor's probe is a `ping` call with a
deadline.
## The two iron rules
1. **Cleanup is the kernel's job.** A process can die with no warning — fault,
kill, power. Correctness must never depend on a `terminate` handler running. On
any death the kernel releases the address space, IPC handles, IRQ bindings, and
owed replies (built), and must also release **device, I/O-port, and interrupt
claims and MSI vectors** (the known gap in
[process-management.md](process-management.md); increment 1). A signal handler is
for *graceful* work — flushing, deregistering, saving — never for *necessary*
work.
2. **Kill is not a signal, and exit reasons are load-bearing.** The standard stop
sequence is *terminate → deadline → `process_kill`*; the unhandleable kill stays
a kernel mechanism. And a supervisor deciding whether to restart must know *how*
the child died: clean exit (meant to — don't restart), fault (restart with
backoff), killed (the supervisor did it). The exit notification today carries
only the id; it grows a reason. Restart policy cannot be written without it.
## Who learns of a death
A death has three audiences, and conflating them is how systems end up with either
zombie state or privileged snooping:
1. **The supervisor** — gets the exit notification on the endpoint it gave at spawn
(built), which grows the `ExitReason` (increment 2). The supervisor is the only
audience that needs the *reason*, because it is the only one deciding whether to
restart.
2. **The peer owed a reply** — already built: a client that dies mid-request fails
the server's reply with `-EPEER`; a server that dies fails its waiting clients
the same way. This covers the *synchronous* case only.
3. **The subscribers** — the new piece, and it is the input service's
publish/subscribe shape ([input.md](input.md)) applied to exits. A stateful
service accumulates per-client state across many requests: the VFS holds a dead
client's open file handles, the input service holds its subscriptions, a future
network stack holds its sockets. None of these are the client's supervisor, and
none learn anything from a failed reply if the client simply never calls again.
So the kernel **publishes every exit** to whoever subscribed:
`process_subscribe(endpoint)` adds a subscriber, and each death posts a
notification to every subscriber (badge = `notify_exit_bit | process id` — the
same encoding supervisors already decode, the IRQ-as-IPC pattern once more). The
subscriber filters for ids it holds state for and releases what the dead client
held. Correlating is free of bookkeeping: an IPC sender's badge already *is* its
task id (`runtime.ipc.Received`), so the id a service has been keying client
state by all along is the id the exit event carries.
Subscription, not broadcast-to-everyone: only processes that asked receive
events, the kernel keeps a bounded subscriber table, and delivery is the same
non-blocking coalescing notification as everything else — a dying process never
waits on its mourners. Subscribing is ungated, like `process_enumerate`: what is
running (and dying) is not a secret between cooperating processes. Subscribers
do not receive the exit reason — the VFS does not care *why* the client died.
This is the service-side mirror of iron rule 1: **a service must never depend on
its clients cleaning up after themselves.** Handle release on client death is the
service's job, triggered by the published exit event — never by a courtesy
"closing now" message that a crashed client will never send.
## The stable interface: `runtime.process`
`runtime.process` already owns what a process receives at birth (`Init`, the
argv contract). It grows to own the other end of life.
**The runtime is the stable interface; the numbers are not.** danos applications do
not make system calls — they call the runtime library, and the system-call numbers,
notification bits, and signal bit positions beneath it are a **private kernel ↔
runtime contract** that may change at any time (settled 2026-07-12). This is why
the runtime exists. Today kernel and runtime ship from one tree in one image, so
"stability" is simply building them together. When driver binaries start shipping
as separately-versioned applications — the whole point of the restart design — the
binary's embedded runtime version becomes compatibility metadata (the same idea as
the protocol version in the device manager's `hello`), and the kernel refuses what
it cannot serve. Signals therefore need no reserved numbering scheme: the enum
below is vocabulary, not ABI.
```zig
/// The signal vocabulary. The value is the bit position in the pending mask — a
/// private kernel/runtime detail, free to change while they ship together.
pub const Signal = enum(u5) {
terminate = 0, // SIGTERM: finish up and exit
reload = 1, // SIGHUP: re-read configuration
interrupt = 2, // SIGINT
quit = 3, // SIGQUIT
alarm = 4, // SIGALRM
user_1 = 5, // SIGUSR1
user_2 = 6, // SIGUSR2
};
/// A decoded pending mask: the coalesced set of signals a notification delivered.
pub const SignalSet = struct {
pending: u32,
pub fn has(set: SignalSet, signal: Signal) bool { ... }
pub fn iterate(set: SignalSet) Iterator { ... }
};
/// Nominate `endpoint` as this process's signal endpoint (signal_bind). The
/// runtime's service harness calls this; a bare program may call it directly and
/// fold signals into its own replyWait loop.
pub fn bindSignals(endpoint: usize) bool { ... }
/// Decode a received badge into signals, or null if the badge is not a signal
/// notification (mirrors ipc.Received.isChildExit).
pub fn signalsFrom(badge: usize) ?SignalSet { ... }
/// Send `signal` to process `id`. Supervisor-gated, like kill; non-blocking.
pub fn sendSignal(id: u32, signal: Signal) bool { ... }
/// The standard stop sequence: terminate, wait up to `deadline_ms` for the exit
/// notification, then process_kill. The one call a supervisor needs.
pub fn stop(id: u32, deadline_ms: u64) void { ... }
/// Subscribe `endpoint` to published exit events (process_subscribe). Every
/// process death posts an asynchronous notification: badge = notify_exit_bit |
/// process id — the same encoding a supervisor's exit notification uses, decoded
/// by the same ipc.Received helpers. For stateful services: release what the dead
/// client held (file handles, subscriptions, sockets). Ungated, like
/// process_enumerate.
pub fn subscribeExits(endpoint: usize) bool { ... }
/// How a process ended — from the exit notification. What restart policy reads.
pub const ExitReason = enum {
exited, // returned from main / clean exit
aborted, // abort() — deliberate self-termination (SIGABRT's ghost)
segmentation_fault, // SIGSEGV's ghost
illegal_instruction, // SIGILL's ghost
arithmetic_fault, // SIGFPE's ghost
killed, // process_kill
};
```
Two deliberate absences. There is no `mask`/`block` API — a process that is not
ready for a signal simply has not waited on its endpoint yet; the pending mask *is*
the blocked set. And there is no per-signal handler registration at this layer —
dispatch is the process's own `switch` over `SignalSet`, or the service harness's
callbacks (`on_terminate`, `on_reload`) for programs that want defaults.
### The service harness
`runtime.service` owns the `replyWait` loop and folds every event source — signals,
child exits, protocol messages — into callbacks, with the vocabulary's defaults:
`terminate` returns from the loop (clean exit), the common `ping` is answered automatically,
`reload` is ignored unless overridden. One loop, no locking, nothing reentrant. A
service author writes domain logic; the lifecycle contract is satisfied by the
harness. A process that bypasses the harness and ignores its signals meets the
deadline-then-kill escalation — you cannot force a process to implement an
interface, but you can make compliance free and non-compliance fatal.
### The musl layer later
The POSIX C layer is a **musl port**: musl's arch/syscall layer retargeted so that
what musl believes are kernel syscalls become danos runtime calls and IPC — `open`
and `read` onto the VFS protocol, `kill`/`sigaction`/`waitpid` onto this document's
vocabulary, `exit` onto the runtime's exit path. `sigaction` handlers registered
through it are invoked by the runtime's loop when the signal message arrives —
synchronous underneath, async-looking to ported code, delivered at wait boundaries
the way most Unix programs already experience signals (at syscalls). No stack hijack
ever happens, `SA_RESTART` semantics come free because nothing was interrupted, and
SIGPIPE can be synthesized from `-EPEER` for the programs that expect it. C programs
get POSIX; danos-native programs never pay for it.
## Increments
1. **Kernel: release device/port/IRQ claims and MSI vectors on death** — the
cleanup half of iron rule 1, and the prerequisite for any restart story. Test:
kill a claiming driver, spawn it again, the claim succeeds.
2. **Exit reason in the death notification** (`ExitReason` above).
3. **Exit events**: `process_subscribe` in the kernel (bounded subscriber table,
publishes on every death), `runtime.process.subscribeExits`; the VFS becomes the
first subscriber — releasing a dead client's handles is its proof test.
4. **Signals**: `signal_bind` + `process_signal` + the pending mask in the kernel;
`runtime.process` grows the interface above; the service harness handles
`terminate` and answers the common `ping`; `stop()` for supervisors.
[device-manager.md](device-manager.md) builds directly on all four.
## Settled questions (2026-07-12)
- **Signal numbering is not ABI**: the runtime is the stable interface; the numbers
beneath it are a private kernel ↔ runtime contract (see "The stable interface").
- **Liveness is a `ping` call, not a signal**: signals are statements, questions
are synchronous calls (see "Statements, not questions"). A service wanting *deep*
health ("can I reach my hardware?") defines its own protocol message on top.
- **Process handles: deferred.** Pids + the supervisor gate cover everything
planned; transferable handles (Fuchsia-style, delegating signalling without
delegating kill) wait for the capability table to grow types beyond endpoints.
- **`alarm`: in the vocabulary, unbuilt.** No consumer yet; when one appears it is
runtime sugar over the existing timer (arm a timer that posts your own signal) —
zero kernel work, so deferring costs nothing.
- **Subscription granularity: all exits**, subscriber-side filtering — one
subscription per service, a bounded kernel table. Per-id subscriptions only if
event volume ever matters (hundreds of processes, not before).
- **Client identity across the exit boundary: no convention needed** — an IPC
sender's badge already is its task id (see "Who learns of a death").
+112
View File
@@ -0,0 +1,112 @@
# Process Management
How danos lists, supervises, and kills processes — the microkernel answer to
`ps`, `kill`, and `SIGCHLD`/`wait`.
## Why system calls, not `/proc`
Unix systems sit on a spectrum. Classic BSD/macOS list processes through
syscalls (`sysctl(KERN_PROC)`) and kill through `kill(2)`; Linux renders the
process table as `/proc` for *reading* but still kills through a syscall; Plan 9
made the file tree the whole interface (`echo kill > /proc/n/ctl`). Microkernels
mostly abandon ambient PIDs: Minix and QNX route everything through a user-space
process-manager server, and Fuchsia/seL4 control processes only through handles.
danos rules out `/proc` **as the primitive**: here a `/proc` would be served by
the VFS server — a user process — which would put the VFS in the path of process
control. If the VFS (or anything under it) hangs, nothing could be listed or
killed, *including the hung VFS*. The control plane for processes must not
depend on a process. So the primitives are kernel system calls; a read-only
`/proc` rendering can be layered on later, and a POSIX-style process-manager
server can be built *from* these primitives when one is needed.
## The three primitives
### `process_enumerate(buffer, maximum) -> total`
A snapshot of the task table into a caller buffer of `abi.ProcessDescriptor`
(id, supervisor, state, priority, name) — the exact shape of
`device_enumerate`, so `ps` is a user program over a snapshot, not a kernel
service. The total may exceed what fit; call again with a larger buffer. Kernel
tasks are included with an empty name — an honest listing shows the idle tasks
too. Ungated and read-only: what is running is not a secret between cooperating
bring-up processes.
### `system_spawn(..., exit_endpoint) -> child id`, and the supervision link
`system_spawn` records the caller as the child's **supervisor** and returns the
child's process id (ids are monotonic, never reused — a stale id can only miss).
That link is the kill authority: it answers "who may kill process 7?" without
inventing users or permissions, the same way a device *claim* is the capability
for `mmio_map`. It composes with the supervision hierarchy the device manager
already forms: init supervises the services it starts, the device manager
supervises the drivers it matches. (A transferable process *handle* — Fuchsia
style — can replace the id once the handle table grows types beyond endpoints.)
`exit_endpoint` (a handle, or `abi.no_cap`) is the supervisor's death-watch: when
the child ends — clean exit, CPU fault, or `process_kill` — the kernel posts an
asynchronous notification to that endpoint, exactly like a bound IRQ. The badge
carries `abi.notify_badge_bit | abi.notify_exit_bit | child_id`, so one endpoint
supervises many children and can even share with IRQ notifications. This is the
microkernel's SIGCHLD: no new mechanism, just the IRQ-as-IPC pattern reused, and
a supervisor's event loop (`ipc.replyWait`) already knows how to receive it. The
child holds a reference to the endpoint from birth, so the notification cannot
dangle even if the supervisor dies first.
### `process_kill(id) -> 0 / -ESRCH / -EPERM`
Only the supervisor may kill; kernel tasks are not killable processes. Like a
signal, delivery is prompt but asynchronous — 0 means the kill is accepted and
irrevocable; the exit notification confirms completion.
## How a kill lands (the kernel mechanics)
Everything below runs under the big kernel lock, where task states cannot move.
- **Target ready or blocked** (not on any core): reaped on the killer's own
call. The reap releases what death always releases (IRQ bindings first, then
a client the target still owed a reply to is failed with `-EPEER`, IPC handles
closed, the exit notification posted last) — plus the unlinking only a
*remote* death needs: out of the ready queue, out of an endpoint's sender FIFO
(`Task.ipc_wait_endpoint`), out of a receive wait queue (`Task.wait_queue`),
and out of any server's owed-reply slot, so nothing ever dequeues a dangling
pointer. Destroying the address space is safe because no core can have it
loaded: every switch away from a task loads the next task's tables.
- **Target running on another core**: it cannot be torn down mid-instruction,
so it is condemned (`Task.kill_pending`) and dies at whichever comes first:
- its next **system_call entry** — checked before dispatch, so a condemned
process cannot spawn, claim, or message anything on its way out;
- its core's next **timer tick** — but only when the task is not inside one
of its own system calls (`Task.in_system_call`): the tick may have
interrupted kernel code mid-operation, where teardown would leak whatever
the operation held. User-mode execution is always a safe kill point. The
tick-time terminate abandons the interrupt frame exactly like the fault
path (the LAPIC is acknowledged before the tick hook runs);
- any core's tick finding it **blocked or ready** (it entered a syscall and
parked after being condemned) — reaped by the same remote-reap path.
A pure user-mode spin loop that never makes a system call therefore dies
within one tick; nothing a process does can outrun the kill.
The scheduler stays below the process layer: finishing a kill (IRQ bindings,
handles, the notification) is called *up* through two hooks process.zig
registers at boot (`terminate_current_hook`, `reap_task_hook`), mirroring how
the architecture layer calls up into `tick`.
## Known gaps (bring-up honesty)
- Device **claims** are not released on death (pre-existing: the fault path has
the same gap) — a killed driver's device stays claimed until reboot.
- Kernel stacks of dead tasks are leaked, as on every exit path (no reaper yet).
- There is no exit *status* in the notification, only the id; a supervisor that
needs the code can grow a wait-style call later.
- Enumerate writes through the caller's raw pointer under the bring-up trust
model, like `device_enumerate` (an unmapped page is a self-DoS, not an
isolation break).
## Tests
`process-list` (enumerate), `process-kill` (kernel-level kill paths, refusals,
notifications), `supervision` (the whole user-side surface via the process-test
service: spawn supervised → enumerate → kill blocked and spinning children →
notifications → gone). See test/qemu_test.py.
+9 -3
View File
@@ -1,6 +1,10 @@
# Resilience: fault isolation and live restart
A design/research note, not built yet. This is the property danos is really chasing:
Steps 1–2 of the ordering below are **built**: user-mode isolation, and fault →
kill the process → keep the core (`onException` in `system/kernel/kernel.zig`; the
`fault-recovery` test proves a crashing ring-3 process dies alone while the system
keeps running). The supervisor notification and restart policy (steps 3+) are
still design. This is the property danos is really chasing:
**if a part of the OS breaks, isolate it, and re-initialise it — without rebooting.**
A crashed driver gets restarted; a wedged service gets killed and brought back. It's
the reason the [microkernel](vision.md) shape was chosen, and it's a *separate* goal
@@ -111,9 +115,11 @@ Honest boundaries:
## Suggested ordering
1. **User mode + address-space isolation** — the shared prerequisite (also on the
path for everything else).
path for everything else). **Done.**
2. **Kernel: fault → kill process → notify.** Turn today's "halt on fault" into
"confine to the process and report it."
"confine to the process and report it." **Done** (the kill and reclaim; the
supervisor notification waits for step 3's supervisor). A killed server's
pending client is unblocked with `-EPEER` rather than hung.
3. **A minimal supervisor server** that can (re)start a process.
4. **Resource cleanup on death** — reclaim memory/MMIO/IPC/IRQ, via caps or a grant
table.
+2
View File
@@ -47,6 +47,8 @@ Everything else---including`read()`,`write()`,`malloc()`, and`fork()`---will run
- **What it does:**Used strictly by your background user-space servers (like your disk driver or filesystem). It sends a reply to the last client that called it, and immediately puts the server to sleep until the next request arrives.[[1](https://news.ycombinator.com/item?id=33078441)]
3. **`Yield()`/`Thread_Ctrl()`**
- **What it does:**Allows a thread to voluntarily give up its CPU time slice, or allows a root task to spawn/kill threads.
4. **`ipc_send(endpoint, message_buffer)`(Asynchronous Send)**
- **What it does:**Posts a small payload to an endpoint's bounded queue and returns *without* blocking — no rendezvous, no reply. The receiver picks it up through the same `IPC_ReplyWait`, as a buffered message. It is the async counterpart of `IPC_Call`, for one-to-many broadcasts where a synchronous rendezvous would let one dead or slow receiver hang the sender. The [input service](input.md) — keyboard-event fan-out — is its first user. A full queue drops the oldest message (a buffered message is discrete data, unlike a coalescing interrupt notification).
* * * * *
+32
View File
@@ -65,6 +65,38 @@ function-pointer type and the kernel's `_start` both carry
whole reason `kernel_abi` lives in the shared contract — see [efi.md](efi.md) for
the handoff it governs.
## The process-entry stack (argc/argv)
The SysV ABI also fixes what a *fresh process* finds on its stack — and danos
follows it, so its own runtime and any future C libc read arguments the same way.
At the first user instruction, `rsp` is 16-byte aligned and points at (addresses
growing upward):
```
rsp → argc u64
argv[0] … argv[argc-1] pointers into the strings area below
NULL argv terminator
NULL envp terminator (no environment yet)
{AT_PAGESZ, page size} auxiliary vector
{AT_NULL, 0} auxiliary-vector terminator
argv string bytes NUL-terminated
───────────────────────── stack top (stack_top_virtual)
```
The kernel builds this block at the top of the process's stack — 8 pages (32 KiB,
`parameters.user_stack_pages`) mapped RW+NX below a fixed top, with the page below
them left unmapped as a **guard**, so a stack overflow faults (killing only that
process) instead of silently corrupting the image
(`buildEntryStack` in `system/kernel/process.zig`); `argv[0]` is always the path
or initial-ramdisk name the process was spawned as, and `system_spawn`'s optional
argument blob becomes `argv[1..]`. The runtime's `_start`
(`library/runtime/start.zig`) hands the block to `rt_start`, which builds a
`runtime.process.Init` from it and passes that to the program's `main`
(`pub fn main(init: runtime.process.Init)`; a parameterless `main()` is also
accepted). A C runtime's `crt0` would walk
the identical layout unmodified — that's the compatibility being bought. The
`args` test proves the round trip.
## Where else it surfaces
- **The red zone → `red_zone = false`.** `build.zig` disables the red zone for the
+80
View File
@@ -0,0 +1,80 @@
//! /lib/mmio — typed volatile MMIO register access, plus the memory-ordering
//! barriers a device driver needs. Used by drivers on top of an `mmio_map` grant.
//!
//! **`volatile` is not a barrier.** In Zig it means only: don't elide this access, and
//! don't reorder it against *other volatile* accesses. It says nothing about ordinary
//! stores — the DMA descriptor you just filled in write-back RAM — which the compiler
//! (and, on weakly-ordered hardware, the CPU) may freely move past a volatile MMIO
//! write. The canonical bug:
//!
//! ring[i] = descriptor; // ordinary store to WB RAM
//! doorbell.* = i; // volatile store to UC MMIO
//! // nothing orders these; the device can read a stale descriptor
//!
//! Put a `wmb()` between them. The barriers lower per-architecture — which is the whole
//! reason they are a named primitive and not scattered `asm volatile`:
//!
//! x86_64 aarch64
//! mb() mfence dsb sy
//! rmb() lfence dsb ld
//! wmb() sfence dsb st
//!
//! x86 is forgiving (TSO + strong-uncacheable MMIO), so a compiler barrier usually
//! suffices; ARM is not, and ARM is the win condition (docs/vision.md) — so the
//! abstraction exists now, while there is one caller (hpet) to get right. See
//! docs/driver-model.md (M14) for the full ordering contract.
const builtin = @import("builtin");
/// Read a register of type `T` at absolute virtual address `addr` — a location inside
/// a device's `mmio_map` grant. `volatile`: never elided, never reordered against
/// another volatile access.
pub inline fn read(comptime T: type, addr: usize) T {
return @as(*const volatile T, @ptrFromInt(addr)).*;
}
/// Write `value` of type `T` to the register at absolute virtual address `addr`.
pub inline fn write(comptime T: type, addr: usize, value: T) void {
@as(*volatile T, @ptrFromInt(addr)).* = value;
}
/// Full barrier: all loads and stores before it are globally visible before any after
/// it. Use when an MMIO write must complete before a following read.
pub inline fn mb() void {
switch (builtin.target.cpu.arch) {
.x86_64 => asm volatile ("mfence" ::: .{ .memory = true }),
.aarch64 => asm volatile ("dsb sy" ::: .{ .memory = true }),
else => @compileError("mmio.mb: unsupported architecture"),
}
}
/// Read barrier: loads before it complete before loads after it. Use after an IRQ
/// wake, before reading what the device wrote to shared memory.
pub inline fn rmb() void {
switch (builtin.target.cpu.arch) {
.x86_64 => asm volatile ("lfence" ::: .{ .memory = true }),
.aarch64 => asm volatile ("dsb ld" ::: .{ .memory = true }),
else => @compileError("mmio.rmb: unsupported architecture"),
}
}
/// Write barrier: stores before it become visible before stores after it. Use between
/// filling a DMA descriptor in RAM and ringing the device's doorbell.
pub inline fn wmb() void {
switch (builtin.target.cpu.arch) {
.x86_64 => asm volatile ("sfence" ::: .{ .memory = true }),
.aarch64 => asm volatile ("dsb st" ::: .{ .memory = true }),
else => @compileError("mmio.wmb: unsupported architecture"),
}
}
test "barriers emit and registers round-trip through a RAM cell" {
// The barriers must at least assemble for the host arch; ordering can't be unit
// tested, but a missing/mistyped mnemonic is caught here.
wmb();
rmb();
mb();
var cell: u64 = 0;
write(u64, @intFromPtr(&cell), 0xDEAD_BEEF);
try @import("std").testing.expectEqual(@as(u64, 0xDEAD_BEEF), read(u64, @intFromPtr(&cell)));
}
+64
View File
@@ -3,6 +3,8 @@
//! ownership of its hardware; the claim is the capability the kernel checks before
//! mapping registers or routing an IRQ.
const std = @import("std");
const abi = @import("abi");
const device_abi = @import("device-abi");
const sc = @import("system-call.zig");
@@ -35,6 +37,10 @@ pub fn mmioMap(device_id: u64, resource_index: u64) ?usize {
/// `DeviceDescriptor.parent` for a device with no parent.
pub const no_parent = device_abi.no_parent;
/// `DeviceDescriptor.pci_class` for a device that is not a PCI function. Set this on
/// descriptors passed to `register` unless the child really is one.
pub const no_pci_class = device_abi.no_pci_class;
/// Publish `descriptor` as a child of `parent_id`, which this process must have claimed.
/// Returns the new device id. The child is left unclaimed, so whichever driver owns
/// that class of device can `claim` it — that is how a bus hands off a device.
@@ -65,3 +71,61 @@ pub fn irqBind(device_id: u64, resource_index: u64, endpoint: usize) bool {
pub fn irqAck(device_id: u64, resource_index: u64) bool {
return !failed(sc.systemCall2(.irq_ack, device_id, resource_index));
}
/// The Message-Signalled Interrupt address/data a driver programs into its device's
/// MSI capability. The device raises the interrupt by writing `data` to `address`.
pub const Msi = struct { address: u64, data: u32 };
/// Set up MSI for a claimed device: the kernel allocates a per-device edge-triggered
/// vector, binds it to `endpoint` (delivered like `irqBind`, but with no mask and no
/// `irqAck` cycle), and returns the (address, data) to write into the device's MSI
/// capability — found by mmio_mapping the device's ECAM config space (resource 0) and
/// walking its capability list. Returns null on failure. Two return values (address in
/// rax, data in rdx), so a hand-written stub.
pub fn msiBind(device_id: u64, endpoint: usize) ?Msi {
var rax: usize = undefined;
var rdx: usize = undefined;
asm volatile ("syscall"
: [rax] "={rax}" (rax),
[rdx] "={rdx}" (rdx),
: [n] "{rax}" (@intFromEnum(abi.SystemCall.msi_bind)),
[a0] "{rdi}" (device_id),
[a1] "{rsi}" (endpoint),
: .{ .rcx = true, .r11 = true, .memory = true });
if (failed(rax)) return null;
return .{ .address = rax, .data = @intCast(rdx) };
}
/// Read `width` bytes (1, 2, or 4) from a port in a claimed device's `io_port`
/// resource, at byte `offset` within it. Ring 3 has no direct `in`/`out`, so a legacy
/// driver (PS/2, 16550 UART) reaches its ports through this claim-gated call — each
/// access is a syscall, which is fine for the low-rate hardware that needs it. Returns
/// null if the capability check fails (device not claimed, wrong resource, out of
/// range). A device that decodes no data returns all-ones, which is a valid value, not
/// a failure.
pub fn ioRead(device_id: u64, resource_index: u64, offset: u64, width: u8) ?u32 {
const r = sc.systemCall4(.io_read, device_id, resource_index, offset, width);
return if (failed(r)) null else @intCast(r);
}
/// Write `value` (its low `width` bytes, 1/2/4) to a port in a claimed device's
/// `io_port` resource, at byte `offset`. Same capability gate as `ioRead`.
pub fn ioWrite(device_id: u64, resource_index: u64, offset: u64, width: u8, value: u32) bool {
return !failed(sc.systemCall5(.io_write, device_id, resource_index, offset, width, value));
}
/// Find DeviceDescription by hid
///
/// Utility function for driver development
pub fn findDeviceDescriptorByHid(buffer: []DeviceDescriptor, hid_needle: []const u8) ?DeviceDescriptor {
const total = enumerate(buffer);
const n = @min(total, buffer.len);
for (@as([]DeviceDescriptor, buffer[0..n])) |d| {
const hid_haystack = d.hid[0..@intCast(d.hid_len)];
if (std.mem.eql(u8, hid_haystack, hid_needle)) {
return d;
}
}
return null;
}
+48
View File
@@ -0,0 +1,48 @@
//! User-space DMA memory: `dma_alloc` / `dma_free`. A driver that programs a
//! bus-mastering engine needs a descriptor ring the device can read — memory that is
//! physically contiguous, at a physical address the driver knows, uncacheable, and
//! pinned. `mmap` gives none of those; this does. Pair it with the barriers in
//! `/lib/mmio` (fill the ring, `wmb()`, ring the doorbell). See docs/driver-model.md.
const abi = @import("abi");
const sc = @import("system-call.zig");
/// Allocation flags. `coherent` (uncacheable) is the portable default; the rest are
/// opt-in for specific hardware — see `abi`.
pub const coherent: usize = abi.dma_coherent;
pub const write_combining: usize = abi.dma_write_combining;
pub const below_4g: usize = abi.dma_below_4g;
/// A DMA allocation: the `virtual` address the CPU touches, and the `physical` address
/// to program into the device's descriptor-ring / base registers.
pub const Region = struct {
virtual: usize,
physical: usize,
};
inline fn failed(r: usize) bool {
return r > ~@as(usize, 0) - 4095;
}
/// Allocate `len` bytes of DMA-capable memory with `flags` (e.g. `coherent`, or
/// `coherent | below_4g`). Returns the virtual/physical pair, or null on failure. Two
/// return values — the virtual address in rax, the physical address in rdx — so it
/// needs a hand-written stub.
pub fn alloc(len: usize, flags: usize) ?Region {
var rax: usize = undefined;
var rdx: usize = undefined; // out: physical address
asm volatile ("syscall"
: [rax] "={rax}" (rax),
[rdx] "={rdx}" (rdx),
: [n] "{rax}" (@intFromEnum(abi.SystemCall.dma_alloc)),
[a0] "{rdi}" (len),
[a1] "{rsi}" (flags),
: .{ .rcx = true, .r11 = true, .memory = true });
if (failed(rax)) return null;
return .{ .virtual = rax, .physical = rdx };
}
/// Release a region from a prior `alloc` (`virtual` and the same `len`).
pub fn free(virtual: usize, len: usize) void {
_ = sc.systemCall2(.dma_free, virtual, len);
}
+220
View File
@@ -0,0 +1,220 @@
//! User-space input helpers: the client and publisher sides of the input service, so a
//! program listening for input events — or a driver broadcasting them — doesn't hand-roll
//! the IPC. Layered over `ipc` (endpoints, capability passing, `send`) and the shared
//! `input-protocol` wire format, the way `device.zig` layers over the raw `device_*` calls.
//! See system/services/input/input.zig.
//!
//! The service carries several device classes (keyboard, mouse, joystick/gamepad). A
//! **source** publishes its class with the matching method:
//! var source = input.connectSource() orelse return;
//! _ = source.publishKeyboardEvent(.{ .kind = ..., .keycode = ..., ... });
//! _ = source.publishMouseEvent(.{ ... });
//! _ = source.publishJoystickEvent(.{ ... });
//!
//! A **subscriber** either takes one class with a typed helper —
//! var keys = input.subscribeKeyboard() orelse return;
//! while (true) { const key = keys.next() orelse continue; ... }
//! — or takes several at once and inspects the tagged envelope:
//! var listener = input.subscribeAll() orelse return;
//! while (true) {
//! const event = listener.next() orelse continue;
//! if (event.asKeyboard()) |k| { ... } else if (event.asMouse()) |m| { ... }
//! }
const std = @import("std");
const abi = @import("abi");
const ipc = @import("ipc.zig");
const system = @import("system.zig");
const protocol = @import("input-protocol");
pub const DeviceKind = protocol.DeviceKind;
pub const InputEvent = protocol.InputEvent;
pub const KeyEvent = protocol.KeyEvent;
pub const MouseEvent = protocol.MouseEvent;
pub const JoystickEvent = protocol.JoystickEvent;
pub const EventKind = protocol.EventKind;
pub const MouseEventKind = protocol.MouseEventKind;
pub const JoystickEventKind = protocol.JoystickEventKind;
pub const Keycode = protocol.Keycode;
/// Interest masks re-exported so a caller can `subscribe(input.device_keyboard |
/// input.device_mouse)`.
pub const device_keyboard = protocol.device_keyboard;
pub const device_mouse = protocol.device_mouse;
pub const device_joystick = protocol.device_joystick;
pub const device_all = protocol.device_all;
/// Look up the input service, retrying while it is still coming up. Both a subscriber and
/// a source race the service's registration at boot, so both wait for it here rather than
/// failing. Returns the service endpoint handle, or null if it never appears.
fn lookupService() ?ipc.Handle {
var attempts: usize = 0;
while (attempts < 100) : (attempts += 1) {
if (ipc.lookup(.input)) |handle| return handle;
system.sleep(50);
}
return null;
}
// --- subscribing ------------------------------------------------------------
/// A subscription to the input service: our own endpoint, which the service pushes events
/// to. `next` returns each event as a tagged `InputEvent`; use `asKeyboard`/`asMouse`/
/// `asJoystick` to decode. Created with `subscribe`/`subscribeAll`; for a single device
/// class prefer the typed helpers (`subscribeKeyboard`, ...), which return decoded events.
pub const Subscriber = struct {
/// The endpoint the service delivers events to (created and owned by us; its handle
/// was handed to the service as a capability at subscribe time).
endpoint: ipc.Handle,
receive: [protocol.event_size]u8 = undefined,
/// Block until the next event is pushed, and return it. Events arrive as asynchronous
/// buffered messages (`ipc_send` from the service), so nothing is owed in reply — the
/// empty reply this issues is a harmless no-op. Returns null for any non-event wake-up
/// (there should be none), so callers can loop.
pub fn next(self: *Subscriber) ?InputEvent {
const got = ipc.replyWait(self.endpoint, &.{}, &self.receive, null);
if (!got.isMessage() or got.len < protocol.event_size) return null;
return std.mem.bytesToValue(InputEvent, self.receive[0..protocol.event_size]);
}
};
/// Subscribe to the input classes named in `device_mask` (an OR of `device_*`, or
/// `device_all`). Creates an endpoint for the service to push to and hands it over as a
/// capability. Returns a `Subscriber` to loop `next` on, or null on failure.
pub fn subscribe(device_mask: u32) ?Subscriber {
const service = lookupService() orelse return null;
const endpoint = ipc.createIpcEndpoint() orelse return null;
var request = protocol.Request{ .operation = @intFromEnum(protocol.Operation.subscribe), .device_mask = device_mask };
var reply: [protocol.reply_size]u8 = undefined;
const result = ipc.callCap(service, std.mem.asBytes(&request), &reply, endpoint) catch return null;
if (result.len < protocol.reply_size) return null;
if (std.mem.bytesToValue(protocol.Reply, reply[0..protocol.reply_size]).status != 0) return null;
return .{ .endpoint = endpoint };
}
/// Subscribe to every input class (keyboard, mouse, joystick) on one stream.
pub fn subscribeAll() ?Subscriber {
return subscribe(device_all);
}
/// A subscriber filtered to keyboard events, whose `next` returns a decoded `KeyEvent`.
pub const KeyboardSubscriber = struct {
inner: Subscriber,
pub fn next(self: *KeyboardSubscriber) ?KeyEvent {
return (self.inner.next() orelse return null).asKeyboard();
}
};
/// A subscriber filtered to mouse events, whose `next` returns a decoded `MouseEvent`.
pub const MouseSubscriber = struct {
inner: Subscriber,
pub fn next(self: *MouseSubscriber) ?MouseEvent {
return (self.inner.next() orelse return null).asMouse();
}
};
/// A subscriber filtered to joystick/gamepad events, whose `next` returns a decoded
/// `JoystickEvent`.
pub const JoystickSubscriber = struct {
inner: Subscriber,
pub fn next(self: *JoystickSubscriber) ?JoystickEvent {
return (self.inner.next() orelse return null).asJoystick();
}
};
/// Subscribe to keyboard events only; `next` returns decoded `KeyEvent`s.
pub fn subscribeKeyboard() ?KeyboardSubscriber {
return .{ .inner = subscribe(device_keyboard) orelse return null };
}
/// Subscribe to mouse events only; `next` returns decoded `MouseEvent`s.
pub fn subscribeMouse() ?MouseSubscriber {
return .{ .inner = subscribe(device_mouse) orelse return null };
}
/// Subscribe to joystick/gamepad events only; `next` returns decoded `JoystickEvent`s.
pub fn subscribeJoystick() ?JoystickSubscriber {
return .{ .inner = subscribe(device_joystick) orelse return null };
}
// --- publishing -------------------------------------------------------------
/// A connection to the input service for a source (a keyboard/mouse/joystick driver) that
/// publishes events. Each `publish*Event` is a short synchronous call the service answers
/// at once; its own fan-out to subscribers is asynchronous, so publishing never blocks on
/// a slow subscriber.
pub const Publisher = struct {
service: ipc.Handle,
fn publish(self: Publisher, event: InputEvent) bool {
var request = protocol.Request{ .operation = @intFromEnum(protocol.Operation.publish), .event = event };
var reply: [protocol.reply_size]u8 = undefined;
const len = ipc.call(self.service, std.mem.asBytes(&request), &reply) catch return false;
if (len < protocol.reply_size) return false;
return std.mem.bytesToValue(protocol.Reply, reply[0..protocol.reply_size]).status == 0;
}
/// Broadcast a keyboard event to every subscriber that took keyboard events.
pub fn publishKeyboardEvent(self: Publisher, event: KeyEvent) bool {
return self.publish(InputEvent.fromKeyboard(event));
}
/// Broadcast a mouse event to every subscriber that took mouse events.
pub fn publishMouseEvent(self: Publisher, event: MouseEvent) bool {
return self.publish(InputEvent.fromMouse(event));
}
/// Broadcast a joystick/gamepad event to every subscriber that took joystick events.
pub fn publishJoystickEvent(self: Publisher, event: JoystickEvent) bool {
return self.publish(InputEvent.fromJoystick(event));
}
};
/// Connect to the input service as an event source, waiting for it to come up. Returns a
/// `Publisher`, or null if the service never registered.
pub fn connectSource() ?Publisher {
return .{ .service = lookupService() orelse return null };
}
// --- synthetic scaffolding --------------------------------------------------
/// Synthetic key events, shared by the demo source and the keyboard driver's placeholder
/// stream while real scancode decoding is still a follow-up. `step` rolls through A..E,
/// emitting for each key a `key_down`, then a `key_press` carrying the character, then a
/// `key_up`. Scaffolding, not wire protocol — hence it lives with the helpers.
pub fn syntheticKeyEvent(step: usize) KeyEvent {
const Key = struct { code: Keycode, character: u32 };
const keys = [_]Key{
.{ .code = .a, .character = 'A' },
.{ .code = .b, .character = 'B' },
.{ .code = .c, .character = 'C' },
.{ .code = .d, .character = 'D' },
.{ .code = .e, .character = 'E' },
};
const key = keys[(step / 3) % keys.len];
return switch (step % 3) {
0 => .{ .kind = @intFromEnum(EventKind.key_down), .keycode = @intFromEnum(key.code), .character = 0, .modifiers = 0 },
1 => .{ .kind = @intFromEnum(EventKind.key_press), .keycode = @intFromEnum(key.code), .character = key.character, .modifiers = 0 },
else => .{ .kind = @intFromEnum(EventKind.key_up), .keycode = @intFromEnum(key.code), .character = 0, .modifiers = 0 },
};
}
/// Synthetic mouse events (placeholder until real PS/2 packet decoding). `step` alternates
/// a small diagonal motion with a left-button click.
pub fn syntheticMouseEvent(step: usize) MouseEvent {
return switch (step % 3) {
0 => .{ .kind = @intFromEnum(MouseEventKind.motion), .button = 0, .dx = 1, .dy = 1, .scroll_x = 0, .scroll_y = 0, .buttons = 0 },
1 => .{ .kind = @intFromEnum(MouseEventKind.button_down), .button = protocol.mouse_button_left, .dx = 0, .dy = 0, .scroll_x = 0, .scroll_y = 0, .buttons = protocol.mouse_button_left },
else => .{ .kind = @intFromEnum(MouseEventKind.button_up), .button = protocol.mouse_button_left, .dx = 0, .dy = 0, .scroll_x = 0, .scroll_y = 0, .buttons = 0 },
};
}
/// Synthetic joystick/gamepad events (placeholder until a real controller driver). `step`
/// sweeps axis 0 and toggles button 0.
pub fn syntheticJoystickEvent(step: usize) JoystickEvent {
return switch (step % 3) {
0 => .{ .kind = @intFromEnum(JoystickEventKind.axis), .control = 0, .value = 16384, .buttons = 0 },
1 => .{ .kind = @intFromEnum(JoystickEventKind.button_down), .control = 0, .value = 0, .buttons = 1 },
else => .{ .kind = @intFromEnum(JoystickEventKind.button_up), .control = 0, .value = 0, .buttons = 0 },
};
}
+52 -5
View File
@@ -24,8 +24,8 @@ inline fn failed(r: usize) bool {
}
/// Create a new endpoint owned by this process; returns its handle.
pub fn createEndpoint() ?Handle {
const r = sc.systemCall0(.create_endpoint);
pub fn createIpcEndpoint() ?Handle {
const r = sc.systemCall0(.create_ipc_endpoint);
return if (failed(r)) null else r;
}
@@ -79,11 +79,32 @@ pub fn call(h: Handle, message: []const u8, reply: []u8) CallError!usize {
return (try callCap(h, message, reply, null)).len;
}
/// Post `message` to endpoint `h`'s asynchronous queue and return immediately — no
/// rendezvous, no reply, no blocking. The receiver picks it up through `replyWait` as a
/// buffered message (`Received.isMessage`). Unlike `call`, this **cannot hang on a dead
/// or slow peer**, which is why a broadcaster (the input service) delivers events this
/// way. The payload must fit an endpoint slot (64 bytes); a full queue drops the oldest
/// message. Returns false on failure (bad handle, oversized payload, bad buffer).
pub fn send(h: Handle, message: []const u8) bool {
return !failed(sc.systemCall3(.ipc_send, h, @intFromPtr(message.ptr), message.len));
}
/// Set in `Received.badge` when what arrived is an asynchronous notification — a
/// bound device interrupt — rather than a client's message. The low bits carry the
/// GSI. See `isNotification`.
pub const notify_badge_bit: u64 = abi.notify_badge_bit;
/// Set alongside `notify_badge_bit` when the notification is a **child-exit
/// notice** — a process this one spawned (with an exit endpoint) has ended —
/// rather than a device interrupt. The low bits carry the child's process id.
pub const notify_exit_bit: u64 = abi.notify_exit_bit;
/// Set alongside `notify_badge_bit` when the wake-up is a **buffered message** — a payload
/// posted with `send` (`ipc_send`) — rather than a bare device interrupt or child-exit
/// notice. The payload is in the `replyWait` receive buffer (`Received.len` bytes); the
/// low bits of the badge carry the sender's task id. See `Received.isMessage`.
pub const notify_message_bit: u64 = abi.notify_message_bit;
/// The result of a `replyWait`: the request length, the sender's badge (a task id, or
/// an IRQ notification if the high bit is set), and any capability the request carried.
pub const Received = struct {
@@ -91,16 +112,42 @@ pub const Received = struct {
badge: u64,
cap: ?Handle,
/// True if this wake-up was a device interrupt, not a client request. A driver's
/// event loop branches on this; there is no reply owed on the notification path.
/// True if this wake-up was an asynchronous notification (a device interrupt
/// or a child-exit notice), not a client request. An event loop branches on
/// this; there is no reply owed on the notification path.
pub fn isNotification(self: Received) bool {
return self.badge & notify_badge_bit != 0;
}
/// The interrupt source (a GSI), meaningful only when `isNotification`.
/// True if this wake-up tells of a supervised child's end — the notification
/// requested by passing an exit endpoint to `system.spawnSupervised`.
pub fn isChildExit(self: Received) bool {
return self.isNotification() and self.badge & notify_exit_bit != 0;
}
/// True if this wake-up is a **buffered message** posted with `send` (`ipc_send`):
/// there is a payload in the receive buffer (`self.len` bytes) and no reply is owed.
/// The subscriber side of a broadcast branches on this.
pub fn isMessage(self: Received) bool {
return self.isNotification() and self.badge & notify_message_bit != 0;
}
/// The task id of whoever posted a buffered message, meaningful only when
/// `isMessage`. (The badge's low bits, with the three high marker bits masked off.)
pub fn senderTaskId(self: Received) u32 {
return @intCast(self.badge & ~(notify_badge_bit | notify_exit_bit | notify_message_bit));
}
/// The interrupt source (a GSI), meaningful only when `isNotification` and
/// not `isChildExit`.
pub fn source(self: Received) u64 {
return self.badge & ~notify_badge_bit;
}
/// The ended child's process id, meaningful only when `isChildExit`.
pub fn childProcessId(self: Received) u32 {
return @intCast(self.badge & ~(notify_badge_bit | notify_exit_bit));
}
};
/// Server side of IPC_ReplyWait: deliver `reply` to the client last received (if any,
+47
View File
@@ -0,0 +1,47 @@
//! Process-level runtime types: what a user program receives at entry. Mirrors
//! the spirit of `std.process.Init.Minimal` in danos terms — std's `Args` holds
//! no data on freestanding targets, so the type is danos's own.
const std = @import("std");
/// Everything a program receives at entry. Passed to
/// `pub fn main(init: runtime.process.Init)`; programs that need nothing keep
/// `pub fn main() void`. An `environment` field is added here once the kernel
/// passes a non-empty envp (today it is always empty — see docs/sysv.md).
pub const Init = struct {
arguments: Arguments,
};
/// The process arguments (argc/argv), parsed from the kernel-built System V
/// entry block. The bytes live in the entry block at the top of the stack page,
/// NUL-terminated, valid for the process's lifetime.
pub const Arguments = struct {
/// argc — at least 1: argument 0 is the path or name this binary was
/// spawned as.
count: usize,
/// The argv pointers in the entry block (NULL-terminated after `count`
/// entries).
vector: [*]const [*:0]const u8,
/// Argument `index` (0 = the program's own path/name), or null if out of
/// range.
pub fn get(arguments: Arguments, index: usize) ?[:0]const u8 {
if (index >= arguments.count) return null;
return std.mem.span(arguments.vector[index]);
}
pub fn iterate(arguments: Arguments) Iterator {
return .{ .arguments = arguments };
}
pub const Iterator = struct {
arguments: Arguments,
index: usize = 0,
pub fn next(iterator: *Iterator) ?[:0]const u8 {
const argument = iterator.arguments.get(iterator.index) orelse return null;
iterator.index += 1;
return argument;
}
};
};
+12 -1
View File
@@ -8,7 +8,8 @@
//! const runtime = @import("runtime");
//! pub const panic = runtime.panic;
//! comptime { _ = &runtime.start._start; } // pull the entry shim in
//! and a `pub fn main() void`.
//! and a `pub fn main() void` or `pub fn main(init: runtime.process.Init) void`
//! (arguments arrive via `init`).
pub const system = @import("system.zig");
pub const heap = @import("heap.zig");
@@ -16,13 +17,23 @@ pub const ipc = @import("ipc.zig");
pub const start = @import("start.zig");
/// The VFS wire protocol (shared with the VFS server).
pub const vfs_protocol = @import("vfs-protocol");
/// Keyboard-event listening (subscribe/next) and broadcasting (publish), over the input
/// service. See library/runtime/input.zig and system/services/input/.
pub const input = @import("input.zig");
/// The input wire protocol (shared with the input service and its clients).
pub const input_protocol = @import("input-protocol");
/// POSIX-style file API: open/read/write/lseek/stat/close.
/// C stdio: fopen/fread/fwrite/fseek/ftell/fclose over unistd.
/// Device access for drivers: enumerate/claim/mmioMap.
pub const device = @import("device.zig");
/// DMA-capable memory for drivers: contiguous, pinned, uncacheable buffers.
pub const dma = @import("dma.zig");
/// Re-exported so a user binary can `pub const panic = runtime.panic;`.
pub const panic = start.panic;
/// Process entry types: the `Init` handed to `main`, and its `Arguments`.
pub const process = @import("process.zig");
/// The heap as a `std.mem.Allocator`, for Zig `std` containers in user code.
pub const allocator = heap.allocator;
+63 -9
View File
@@ -4,24 +4,78 @@
const std = @import("std");
const system = @import("system.zig");
const process = @import("process.zig");
/// The kernel enters at `_start` with rsp 16-aligned, but a SystemV function expects
/// rsp ≡ 8 (mod 16) on entry (as if reached by `call`). The `call` below pushes
/// the 8-byte return address, satisfying the ABI before any Zig frame runs; the
/// `ud2` is a safety net if `rt_start` ever returns.
/// The kernel enters at `_start` with rsp 16-aligned, pointing at the System V
/// process-entry block it built: argc, argv pointers, NULL, envp terminator, the
/// auxiliary vector, then the strings (see system/kernel/process.zig,
/// `buildEntryStack`). Capture that address in rdi — the first SysV argument —
/// before `call` disturbs the stack; the call's pushed return address also puts
/// rsp ≡ 8 (mod 16), satisfying the ABI before any Zig frame runs. The `ud2` is a
/// safety net if `rt_start` ever returns.
pub export fn _start() callconv(.naked) noreturn {
asm volatile (
\\mov %%rsp, %%rdi
\\call rt_start
\\ud2
);
}
/// The first Zig frame. The heap is lazy (first alloc grows it), so there is no
/// runtime init to order here — just hand control to the program's `main`.
export fn rt_start() callconv(.c) noreturn {
/// The first Zig frame, entered with `stack` pointing at the kernel-built entry
/// block. Build the `process.Init` from it and dispatch to the program's `main`,
/// whose signature is inspected at comptime. The heap is lazy (first alloc grows
/// it), so there is no other runtime init to order here.
export fn rt_start(stack: [*]const u64) callconv(.c) noreturn {
const init: process.Init = .{ .arguments = .{
.count = stack[0],
.vector = @ptrCast(stack + 1),
} };
system.exit(callMain(init));
}
/// Comptime-dispatch on root.main's signature, in the spirit of std's start.zig:
/// zero parameters or one `process.Init`; returns void, noreturn, u8, !void, or !u8.
fn callMain(init: process.Init) u8 {
const root = @import("root"); // the user binary's root source file
root.main();
system.exit(0);
const main_information = @typeInfo(@TypeOf(root.main)).@"fn";
const call_arguments = switch (main_information.params.len) {
0 => .{},
1 => arguments: {
const Parameter = main_information.params[0].type orelse
@compileError("main's parameter must be runtime.process.Init (not anytype)");
if (Parameter != process.Init)
@compileError("main's parameter must be runtime.process.Init, found " ++ @typeName(Parameter));
break :arguments .{init};
},
else => @compileError("main takes no parameters or a single runtime.process.Init"),
};
const ReturnType = main_information.return_type.?;
switch (@typeInfo(ReturnType)) {
.noreturn => @call(.auto, root.main, call_arguments),
.void => {
@call(.auto, root.main, call_arguments);
return 0;
},
.int => {
if (ReturnType != u8)
@compileError("main's integer return type must be u8, found " ++ @typeName(ReturnType));
return @call(.auto, root.main, call_arguments);
},
.error_union => {
const payload = @call(.auto, root.main, call_arguments) catch |err| {
var buffer: [128]u8 = undefined;
const line = std.fmt.bufPrint(&buffer, "main returned error: {s}\n", .{@errorName(err)}) catch "main returned an error\n";
_ = system.write(line);
return 1; // distinct from panic's 127
};
if (@TypeOf(payload) == void) return 0;
if (@TypeOf(payload) == u8) return payload;
@compileError("main's error-union payload must be void or u8, found " ++ @typeName(@TypeOf(payload)));
},
else => @compileError("main must return void, noreturn, u8, !void, or !u8, found " ++ @typeName(ReturnType)),
}
}
/// No runtime to unwind into — report a panic as a nonzero exit code.
+20 -5
View File
@@ -19,34 +19,49 @@ pub inline fn systemCall0(n: SystemCall) usize {
pub inline fn systemCall1(n: SystemCall, a0: usize) usize {
return asm volatile ("syscall"
: [ret] "={rax}" (-> usize),
: [n] "{rax}" (@intFromEnum(n)), [a0] "{rdi}" (a0),
: [n] "{rax}" (@intFromEnum(n)),
[a0] "{rdi}" (a0),
: .{ .rcx = true, .r11 = true, .memory = true });
}
pub inline fn systemCall2(n: SystemCall, a0: usize, a1: usize) usize {
return asm volatile ("syscall"
: [ret] "={rax}" (-> usize),
: [n] "{rax}" (@intFromEnum(n)), [a0] "{rdi}" (a0), [a1] "{rsi}" (a1),
: [n] "{rax}" (@intFromEnum(n)),
[a0] "{rdi}" (a0),
[a1] "{rsi}" (a1),
: .{ .rcx = true, .r11 = true, .memory = true });
}
pub inline fn systemCall3(n: SystemCall, a0: usize, a1: usize, a2: usize) usize {
return asm volatile ("syscall"
: [ret] "={rax}" (-> usize),
: [n] "{rax}" (@intFromEnum(n)), [a0] "{rdi}" (a0), [a1] "{rsi}" (a1), [a2] "{rdx}" (a2),
: [n] "{rax}" (@intFromEnum(n)),
[a0] "{rdi}" (a0),
[a1] "{rsi}" (a1),
[a2] "{rdx}" (a2),
: .{ .rcx = true, .r11 = true, .memory = true });
}
pub inline fn systemCall4(n: SystemCall, a0: usize, a1: usize, a2: usize, a3: usize) usize {
return asm volatile ("syscall"
: [ret] "={rax}" (-> usize),
: [n] "{rax}" (@intFromEnum(n)), [a0] "{rdi}" (a0), [a1] "{rsi}" (a1), [a2] "{rdx}" (a2), [a3] "{r10}" (a3),
: [n] "{rax}" (@intFromEnum(n)),
[a0] "{rdi}" (a0),
[a1] "{rsi}" (a1),
[a2] "{rdx}" (a2),
[a3] "{r10}" (a3),
: .{ .rcx = true, .r11 = true, .memory = true });
}
pub inline fn systemCall5(n: SystemCall, a0: usize, a1: usize, a2: usize, a3: usize, a4: usize) usize {
return asm volatile ("syscall"
: [ret] "={rax}" (-> usize),
: [n] "{rax}" (@intFromEnum(n)), [a0] "{rdi}" (a0), [a1] "{rsi}" (a1), [a2] "{rdx}" (a2), [a3] "{r10}" (a3), [a4] "{r8}" (a4),
: [n] "{rax}" (@intFromEnum(n)),
[a0] "{rdi}" (a0),
[a1] "{rsi}" (a1),
[a2] "{rdx}" (a2),
[a3] "{r10}" (a3),
[a4] "{r8}" (a4),
: .{ .rcx = true, .r11 = true, .memory = true });
}
+83 -5
View File
@@ -2,6 +2,7 @@
//! stubs, one per kernel call. Numbers come from `abi.SystemCall`, the single
//! source of truth shared with the kernel dispatcher.
const std = @import("std");
const abi = @import("abi");
const sc = @import("system-call.zig");
@@ -11,6 +12,10 @@ pub const PROT_READ: usize = abi.prot_read;
pub const PROT_WRITE: usize = abi.prot_write;
pub const PROT_EXEC: usize = abi.prot_exec;
/// One `processes` entry — re-exported from the shared ABI so a user program can
/// declare its snapshot buffer without importing `abi` itself.
pub const ProcessDescriptor = abi.ProcessDescriptor;
/// Give up the rest of this quantum.
pub fn yield() void {
_ = sc.systemCall0(.yield);
@@ -27,6 +32,16 @@ pub fn sleep(ms: usize) void {
_ = sc.systemCall1(.sleep, ms);
}
/// Monotonic nanoseconds since boot — a time source for timeouts and short delays. It
/// only ever moves forward. This is *not* wall-clock time (no date, no timezone — that
/// is a user-space service layered on top). Deadline pattern for a bounded poll loop:
///
/// const deadline = clock() + timeout_ns;
/// while (clock() < deadline) { ... }
pub fn clock() u64 {
return @intCast(sc.systemCall0(.clock));
}
/// End the process. Never returns.
pub fn exit(code: usize) noreturn {
_ = sc.systemCall1(.exit, code);
@@ -34,11 +49,74 @@ pub fn exit(code: usize) noreturn {
}
/// Start the binary bundled in the initial-ramdisk under `name` as a new ring-3
/// process, returning true on success. This is how a supervisor (the device manager)
/// launches a driver it matched — danos-native, not POSIX (a spawn/exec family comes
/// with the process work later).
pub fn spawn(name: []const u8) bool {
return sc.systemCall2(.system_spawn, @intFromPtr(name.ptr), name.len) == 0;
/// process, returning the child's process id (or null on failure). The child's
/// argv[0] is `name`, and the caller becomes its **supervisor** — the only process
/// allowed to `kill` it. This is how a supervisor (the device manager) launches a
/// driver it matched — danos-native, not POSIX (a spawn/exec family comes with the
/// POSIX layer later).
pub fn spawn(name: []const u8) ?u32 {
return spawnSupervised(name, &.{}, null);
}
/// Like `spawn`, but hands the child command-line arguments: they arrive as
/// argv[1..] on its System V entry stack (argv[0] is still `name`).
pub fn spawnWithArguments(name: []const u8, arguments: []const []const u8) ?u32 {
return spawnSupervised(name, arguments, null);
}
/// The full spawn: command-line arguments for the child, and an optional endpoint
/// (a handle from `ipc.createIpcEndpoint`) the kernel notifies when the child ends
/// — any way it ends: clean exit, fault, or `kill`. The notification arrives via
/// `ipc.replyWait` as a badge with the child-exit bit set and the child's id in
/// the low bits (`ipc.Received.isChildExit`/`childProcessId`), so one endpoint can
/// supervise many children. Arguments are marshalled to the kernel as one
/// NUL-separated blob; the combined arguments must fit `blob` (the kernel caps the
/// blob at 256 bytes and argc at 8 anyway). Returns the child's process id, or
/// null on failure.
pub fn spawnSupervised(name: []const u8, arguments: []const []const u8, exit_endpoint: ?usize) ?u32 {
var blob: [256]u8 = undefined;
var len: usize = 0;
for (arguments, 0..) |argument, i| {
if (i != 0) {
if (len >= blob.len) return null;
blob[len] = 0;
len += 1;
}
if (len + argument.len > blob.len) return null;
@memcpy(blob[len..][0..argument.len], argument);
len += argument.len;
}
const r = sc.systemCall5(.system_spawn, @intFromPtr(name.ptr), name.len, if (len == 0) 0 else @intFromPtr(&blob), len, exit_endpoint orelse abi.no_cap);
if (r > ~@as(usize, 0) - 4095) return null; // a wrapped -errno
return @intCast(r);
}
/// Snapshot the process table into `out` (up to its length) and return the total
/// number of live processes — which may exceed `out.len`; call again with a larger
/// buffer for the full listing. Kernel tasks are included, with an empty name.
/// The primitive `ps` is built on.
pub fn processes(out: []abi.ProcessDescriptor) usize {
return sc.systemCall2(.process_enumerate, @intFromPtr(out.ptr), out.len);
}
/// Whether a process spawned under `name` (its argv[0]) is currently alive.
pub fn isProcessRunning(name: []const u8) bool {
var table: [32]ProcessDescriptor = undefined;
const total = processes(&table);
for (table[0..@min(total, table.len)]) |descriptor| {
if (std.mem.eql(u8, descriptor.name[0..descriptor.name_length], name)) return true;
}
return false;
}
/// End process `id`. Only its supervisor — the process that spawned it — may;
/// anyone else gets false, as does a stale or unknown id (ids are never reused).
/// Delivery is prompt but asynchronous, like a signal: a target caught running on
/// another core dies at its next system call or timer tick. True means the kill
/// is accepted and irrevocable; the exit notification (if an endpoint was given
/// at spawn) confirms completion.
pub fn kill(id: u32) bool {
return sc.systemCall1(.process_kill, id) == 0;
}
/// Grant `len` bytes (rounded up to whole pages) of fresh, zeroed, writable
+65
View File
@@ -0,0 +1,65 @@
# xkeyboard-config — X11 keyboard layouts, compiled to Zig
This module turns a physical key (a **USB HID usage**, as the [input module](../../docs/input.md)
delivers in `KeyEvent.keycode`) plus a modifier state into a **keysym** and, when the key
produces one, a **character** (a Unicode scalar). It is what lets a `keycode` become a
`character` — a keymap — without danos shipping an X11 runtime.
The layout data comes from the X11 [xkeyboard-config](https://gitlab.freedesktop.org/xkeyboard-config/xkeyboard-config)
database, but it is **compiled to native Zig at build time** rather than parsed at runtime.
`tools/make-xkeyboard-config.py` reads the vendored xkb data and emits pure-data tables into
`generated/layouts.zig`; `xkeyboard-config.zig` is the hand-written API over them. This is
the same build-time-codegen pattern as `tools/make-initial-ramdisk.py`.
## Using it
```zig
const xkb = @import("xkeyboard-config");
const m = xkb.map(xkb.us, key_event.keycode, .{ .shift = shift_held, .caps_lock = caps });
if (m.character) |ch| { /* a printable Unicode scalar */ }
// m.keysym is always set (e.g. an X11 keysym for Return / F1 / a dead key).
const layout = xkb.byName("gb") orelse xkb.us; // choose a layout by name
for (xkb.all) |l| { /* enumerate available layouts */ }
```
`Modifiers` carries `shift`, `caps_lock`, `level3` (AltGr), and `control`. `map` selects the
level from the key's XKB *type* (the generated data) and those modifiers (the policy, in
`xkeyboard-config.zig`), so data and semantics stay separable.
Layouts: **us, gb, de, fr, es, dvorak**.
## Regenerating
```sh
python3 tools/make-xkeyboard-config.py fetch # network: download + vendor the data subset
python3 tools/make-xkeyboard-config.py generate # offline: emit generated/layouts.zig
# or, from the build:
zig build gen-xkeyboard-config
```
- **`fetch`** downloads the pinned xkeyboard-config release (version + sha256 in the script),
resolves the `include` graph for the configured layouts, and vendors *only* the symbols
files actually reached (plus `keysymdef.h`, `COPYING`, and `PROVENANCE.md`) into `vendor/`.
Run it when bumping the version or adding a layout.
- **`generate`** is deterministic and offline — same vendored input produces byte-identical
output. To add a layout, extend `TARGETS` (and `HID_TO_NAME` if a new physical key is
involved), then re-run `fetch` (to vendor any new includes) and `generate`.
## Scope
A pragmatic subset, enough for real Latin-script typing:
- **Group 1 only** — no multi-layout group switching.
- **No dead-key / compose composition** — a dead key returns its keysym with no `character`
(composing `´` + `e` → `é` is a higher layer's job).
- **Curated key types** — the common XKB types (one/two-level, alphabetic, four-level, …);
unmapped keys and unknown types fall back to level-by-shift.
- **6 layouts** — extend via `TARGETS` as above.
## Licensing
xkeyboard-config and `keysymdef.h` (xorgproto) are MIT/X11 licensed. The vendored data
subset carries the upstream `vendor/COPYING`, and `vendor/PROVENANCE.md` records the exact
version, source URL, and sha256. The generated tables are a derived work under the same terms.
File diff suppressed because it is too large Load Diff
+190
View File
@@ -0,0 +1,190 @@
Copyright 1996 by Joseph Moss
Copyright (C) 2002-2007 Free Software Foundation, Inc.
Copyright (C) Dmitry Golubev <lastguru@mail.ru>, 2003-2004
Copyright (C) 2004, Gregory Mokhin <mokhin@bog.msu.ru>
Copyright (C) 2006 Erdal Ronahî
Permission to use, copy, modify, distribute, and sell this software and its
documentation for any purpose is hereby granted without fee, provided that
the above copyright notice appear in all copies and that both that
copyright notice and this permission notice appear in supporting
documentation, and that the name of the copyright holder(s) not be used in
advertising or publicity pertaining to distribution of the software without
specific, written prior permission. The copyright holder(s) makes no
representations about the suitability of this software for any purpose. It
is provided "as is" without express or implied warranty.
THE COPYRIGHT HOLDER(S) DISCLAIMS ALL WARRANTIES WITH REGARD TO THIS SOFTWARE,
INCLUDING ALL IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS, IN NO
EVENT SHALL THE COPYRIGHT HOLDER(S) BE LIABLE FOR ANY SPECIAL, INDIRECT OR
CONSEQUENTIAL DAMAGES OR ANY DAMAGES WHATSOEVER RESULTING FROM LOSS OF USE,
DATA OR PROFITS, WHETHER IN AN ACTION OF CONTRACT, NEGLIGENCE OR OTHER
TORTIOUS ACTION, ARISING OUT OF OR IN CONNECTION WITH THE USE OR
PERFORMANCE OF THIS SOFTWARE.
Copyright (c) 1996 Digital Equipment Corporation
Permission is hereby granted, free of charge, to any person obtaining
a copy of this software and associated documentation files (the
"Software"), to deal in the Software without restriction, including
without limitation the rights to use, copy, modify, merge, publish,
distribute, sublicense, and sell copies of the Software, and to
permit persons to whom the Software is furnished to do so, subject to
the following conditions:
The above copyright notice and this permission notice shall be included
in all copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS
OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF
MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.
IN NO EVENT SHALL DIGITAL EQUIPMENT CORPORATION BE LIABLE FOR ANY CLAIM,
DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR
OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR
THE USE OR OTHER DEALINGS IN THE SOFTWARE.
Except as contained in this notice, the name of the Digital Equipment
Corporation shall not be used in advertising or otherwise to promote
the sale, use or other dealings in this Software without prior written
authorization from Digital Equipment Corporation.
Copyright 1996, 1998 The Open Group
Permission to use, copy, modify, distribute, and sell this software and its
documentation for any purpose is hereby granted without fee, provided that
the above copyright notice appear in all copies and that both that
copyright notice and this permission notice appear in supporting
documentation.
The above copyright notice and this permission notice shall be
included in all copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND,
EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF
MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.
IN NO EVENT SHALL THE OPEN GROUP BE LIABLE FOR ANY CLAIM, DAMAGES OR
OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE,
ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR
OTHER DEALINGS IN THE SOFTWARE.
Except as contained in this notice, the name of The Open Group shall
not be used in advertising or otherwise to promote the sale, use or
other dealings in this Software without prior written authorization
from The Open Group.
Copyright 2004-2005 Sun Microsystems, Inc. All rights reserved.
Permission is hereby granted, free of charge, to any person obtaining a
copy of this software and associated documentation files (the "Software"),
to deal in the Software without restriction, including without limitation
the rights to use, copy, modify, merge, publish, distribute, sublicense,
and/or sell copies of the Software, and to permit persons to whom the
Software is furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice (including the next
paragraph) shall be included in all copies or substantial portions of the
Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL
THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING
FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER
DEALINGS IN THE SOFTWARE.
Copyright (c) 1996 by Silicon Graphics Computer Systems, Inc.
Permission to use, copy, modify, and distribute this
software and its documentation for any purpose and without
fee is hereby granted, provided that the above copyright
notice appear in all copies and that both that copyright
notice and this permission notice appear in supporting
documentation, and that the name of Silicon Graphics not be
used in advertising or publicity pertaining to distribution
of the software without specific prior written permission.
Silicon Graphics makes no representation about the suitability
of this software for any purpose. It is provided "as is"
without any express or implied warranty.
SILICON GRAPHICS DISCLAIMS ALL WARRANTIES WITH REGARD TO THIS
SOFTWARE, INCLUDING ALL IMPLIED WARRANTIES OF MERCHANTABILITY
AND FITNESS FOR A PARTICULAR PURPOSE. IN NO EVENT SHALL SILICON
GRAPHICS BE LIABLE FOR ANY SPECIAL, INDIRECT OR CONSEQUENTIAL
DAMAGES OR ANY DAMAGES WHATSOEVER RESULTING FROM LOSS OF USE,
DATA OR PROFITS, WHETHER IN AN ACTION OF CONTRACT, NEGLIGENCE
OR OTHER TORTIOUS ACTION, ARISING OUT OF OR IN CONNECTION WITH
THE USE OR PERFORMANCE OF THIS SOFTWARE.
Copyright (c) 1996 X Consortium
Permission is hereby granted, free of charge, to any person obtaining
a copy of this software and associated documentation files (the
"Software"), to deal in the Software without restriction, including
without limitation the rights to use, copy, modify, merge, publish,
distribute, sublicense, and/or sell copies of the Software, and to
permit persons to whom the Software is furnished to do so, subject to
the following conditions:
The above copyright notice and this permission notice shall be
included in all copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND,
EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF
MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.
IN NO EVENT SHALL THE X CONSORTIUM BE LIABLE FOR ANY CLAIM, DAMAGES OR
OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE,
ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR
OTHER DEALINGS IN THE SOFTWARE.
Except as contained in this notice, the name of the X Consortium shall
not be used in advertising or otherwise to promote the sale, use or
other dealings in this Software without prior written authorization
from the X Consortium.
Copyright (C) 2004, 2006 Ævar Arnfjörð Bjarmason <avarab@gmail.com>
Permission to use, copy, modify, distribute, and sell this software and its
documentation for any purpose is hereby granted without fee, provided that
the above copyright notice appear in all copies and that both that
copyright notice and this permission notice appear in supporting
documentation.
The above copyright notice and this permission notice shall be
included in all copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND,
EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF
MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.
IN NO EVENT SHALL THE OPEN GROUP BE LIABLE FOR ANY CLAIM, DAMAGES OR
OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE,
ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR
OTHER DEALINGS IN THE SOFTWARE.
Except as contained in this notice, the name of a copyright holder shall
not be used in advertising or otherwise to promote the sale, use or
other dealings in this Software without prior written authorization of
the copyright holder.
Copyright (C) 1999, 2000 by Anton Zinoviev <anton@lml.bas.bg>
This software may be used, modified, copied, distributed, and sold,
in both source and binary form provided that the above copyright
and these terms are retained. Under no circumstances is the author
responsible for the proper functioning of this software, nor does
the author assume any responsibility for damages incurred with its
use.
Permission is granted to anyone to use, distribute and modify
this file in any way, provided that the above copyright notice
is left intact and the author of the modification summarizes
the changes in this header.
This file is distributed without any expressed or implied warranty.
+21
View File
@@ -0,0 +1,21 @@
# Vendored xkeyboard-config subset
- **Package**: xkeyboard-config 2.44
- **Source**: https://gitlab.freedesktop.org/xkeyboard-config/xkeyboard-config/-/archive/xkeyboard-config-2.44/xkeyboard-config-2.44.tar.gz
- **sha256**: `35e34edeaf4e8da8d0696ff6b241ee11ddb1b8c6730bac7252d4d0a88ea5f05b`
- **keysymdef.h**: xorgproto, copied from `/opt/homebrew/include/X11/keysymdef.h`
- **License**: MIT/X11 (see COPYING)
Only the symbols files reachable from the generated layouts (tools/make-xkeyboard-config.py `TARGETS`) are vendored; regenerate with
`python3 tools/make-xkeyboard-config.py fetch` then `... generate`.
Vendored symbols files:
- `symbols/de`
- `symbols/es`
- `symbols/fr`
- `symbols/gb`
- `symbols/kpdl`
- `symbols/latin`
- `symbols/level3`
- `symbols/us`
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+250
View File
@@ -0,0 +1,250 @@
// Keyboard layouts for Spain.
// Modified for a real Spanish keyboard by Jon Tombs.
default partial alphanumeric_keys
xkb_symbols "basic" {
include "latin(type4)"
name[Group1]="Spanish";
key <TLDE> { [ masculine, ordfeminine, backslash, backslash ] };
key <AE01> { [ 1, exclam, bar, exclamdown ] };
key <AE03> { [ 3, periodcentered, numbersign, sterling ] };
key <AE04> { [ 4, dollar, asciitilde, dollar ] };
key <AE11> { [apostrophe, question, backslash, questiondown ] };
key <AE12> { [exclamdown, questiondown, dead_cedilla, dead_ogonek] };
key <AD11> { [dead_grave, dead_circumflex, bracketleft, dead_abovering ] };
key <AD12> { [ plus, asterisk, bracketright, dead_macron ] };
key <AC10> { [ ntilde, Ntilde, dead_tilde, dead_doubleacute ] };
key <AC11> { [dead_acute, dead_diaeresis, braceleft, dead_caron ] };
key <BKSL> { [ ccedilla, Ccedilla, braceright, dead_breve ] };
include "level3(ralt_switch)"
};
partial alphanumeric_keys
xkb_symbols "winkeys" {
include "es(basic)"
name[Group1]="Spanish (Windows)";
include "eurosign(5)"
};
partial alphanumeric_keys
xkb_symbols "nodeadkeys" {
include "es(basic)"
name[Group1]="Spanish (no dead keys)";
key <AE12> { [exclamdown, questiondown, cedilla, ogonek ] };
key <AD11> { [ grave, asciicircum, bracketleft, degree ] };
key <AD12> { [ plus, asterisk, bracketright, macron ] };
key <AC07> { [ j, J, ezh, EZH ] };
key <AC10> { [ ntilde, Ntilde, asciitilde, doubleacute ] };
key <AC11> { [ acute, diaeresis, braceleft, caron ] };
key <BKSL> { [ ccedilla, Ccedilla, braceright, breve ] };
key <AB10> { [ minus, underscore, ellipsis, abovedot ] };
};
// Spanish Dvorak mapping (note R-H exchange)
partial alphanumeric_keys
xkb_symbols "dvorak" {
name[Group1]="Spanish (Dvorak)";
key <TLDE> {[ masculine, ordfeminine, backslash, degree ]};
key <AE01> {[ 1, exclam, bar, onesuperior ]};
key <AE02> {[ 2, quotedbl, at, twosuperior ]};
key <AE03> {[ 3, periodcentered, numbersign, threesuperior ]};
key <AE04> {[ 4, dollar, asciitilde, onequarter ]};
key <AE05> {[ 5, percent, brokenbar, fiveeighths ]};
key <AE06> {[ 6, ampersand, notsign, threequarters ]};
key <AE07> {[ 7, slash, onehalf, seveneighths ]};
key <AE08> {[ 8, parenleft, oneeighth, threeeighths ]};
key <AE09> {[ 9, parenright, asciicircum ]};
key <AE10> {[ 0, equal, grave, dead_doubleacute ]};
key <AE11> {[ apostrophe, question, dead_macron, dead_ogonek ]};
key <AE12> {[ exclamdown, questiondown, dead_breve, dead_abovedot ]};
key <AD01> {[ period, colon, less, guillemotleft ]};
key <AD02> {[ comma, semicolon, greater, guillemotright ]};
key <AD03> {[ ntilde, Ntilde, lstroke, Lstroke ]};
key <AD04> {[ p, P, paragraph ]};
key <AD05> {[ y, Y, yen ]};
key <AD06> {[ f, F, tslash, Tslash ]};
key <AD07> {[ g, G, dstroke, Dstroke ]};
key <AD08> {[ c, C, cent, copyright ]};
key <AD09> {[ h, H, hstroke, Hstroke ]};
key <AD10> {[ l, L, sterling ]};
key <AD11> {[ dead_grave, dead_circumflex, bracketleft, dead_caron ]};
key <AD12> {[ plus, asterisk, bracketright, plusminus ]};
key <AC01> {[ a, A, ae, AE ]};
key <AC02> {[ o, O, oslash, Oslash ]};
key <AC03> {[ e, E, EuroSign ]};
key <AC04> {[ u, U, aring, Aring ]};
key <AC05> {[ i, I, oe, OE ]};
key <AC06> {[ d, D, eth, ETH ]};
key <AC07> {[ r, R, registered, trademark ]};
key <AC08> {[ t, T, thorn, THORN ]};
key <AC09> {[ n, N, eng, ENG ]};
key <AC10> {[ s, S, ssharp, section ]};
key <AC11> {[ dead_acute, dead_diaeresis, braceleft, dead_tilde ]};
key <BKSL> {[ ccedilla, Ccedilla, braceright, dead_cedilla ]};
key <LSGT> {[ less, greater, guillemotleft, guillemotright ]};
key <AB01> {[ minus, underscore, hyphen, macron ]};
key <AB02> {[ q, Q, currency ]};
key <AB03> {[ j, J ]};
key <AB04> {[ k, K, kra ]};
key <AB05> {[ x, X, multiply, division ]};
key <AB06> {[ b, B ]};
key <AB07> {[ m, M, mu ]};
key <AB08> {[ w, W ]};
key <AB09> {[ v, V ]};
key <AB10> {[ z, Z ]};
include "level3(ralt_switch)"
};
partial alphanumeric_keys
xkb_symbols "cat" {
include "es(basic)"
name[Group1]="Catalan (Spain, with middle-dot L)";
key <AC09> { [ l, L, 0x1000140, 0x100013F ] };
};
partial alphanumeric_keys
xkb_symbols "ast" {
include "es(basic)"
name[Group1]="Asturian (Spain, with bottom-dot H and L)";
key <AC06> { [ h, H, 0x1001E25, 0x1001E24 ] };
key <AC09> { [ l, L, 0x1001E37, 0x1001E36 ] };
};
partial alphanumeric_keys
xkb_symbols "olpc" {
// #HW-SPECIFIC
// http://wiki.laptop.org/go/OLPC_Spanish_Keyboard
include "us(basic)"
name[Group1]="Spanish";
key <AE00> { [ masculine, ordfeminine ] };
key <AE01> { [ 1, exclam, bar ] };
key <AE02> { [ 2, quotedbl, at ] };
key <AE03> { [ 3, dead_grave, numbersign, grave ] };
key <AE05> { [ 5, percent, asciicircum, dead_circumflex ] };
key <AE06> { [ 6, ampersand, notsign ] };
key <AE07> { [ 7, slash, backslash ] };
key <AE08> { [ 8, parenleft ] };
key <AE09> { [ 9, parenright ] };
key <AE10> { [ 0, equal ] };
key <AE11> { [ apostrophe, question ] };
key <AE12> { [ exclamdown, questiondown ] };
key <AD03> { [ e, E, EuroSign ] };
key <AD11> { [ dead_acute, dead_diaeresis, acute, dead_abovering ] };
key <AD12> { [ bracketleft, braceleft ] };
key <AC10> { [ ntilde, Ntilde ] };
key <AC11> { [ plus, asterisk, dead_tilde ] };
key <AC12> { [ bracketright, braceright, section ] };
key <AB08> { [ comma, semicolon ] };
key <AB09> { [ period, colon ] };
key <AB10> { [ minus, underscore ] };
key <I219> { [ less, greater, ISO_Next_Group ] };
include "level3(ralt_switch)"
};
partial alphanumeric_keys
xkb_symbols "olpcm" {
// #HW-SPECIFIC
// Mechanical (non-membrane) OLPC Spanish keyboard layout.
// See: http://wiki.laptop.org/go/OLPC_Spanish_Non-membrane_Keyboard
include "us(basic)"
name[Group1]="Spanish";
key <AE00> { [ questiondown, exclamdown, backslash ] };
key <AE01> { [ 1, exclam, bar ] };
key <AE02> { [ 2, quotedbl, at ] };
key <AE03> { [ 3, dead_grave, numbersign, grave ] };
key <AE04> { [ 4, dollar, asciitilde, dead_tilde ] };
key <AE05> { [ 5, percent, asciicircum, dead_circumflex ] };
key <AE06> { [ 6, ampersand, notsign ] };
key <AE07> { [ 7, slash, backslash ] }; // no '\' label on olpcm, leave for compatibility
key <AE08> { [ 8, parenleft, masculine ] };
key <AE09> { [ 9, parenright, ordfeminine ] };
key <AE10> { [ 0, equal ] };
key <AE11> { [ apostrophe, question ] };
key <AD03> { [ e, E, EuroSign ] };
key <AD11> { [ dead_acute, dead_diaeresis, dead_abovering, acute ] };
key <AD12> { [ plus, asterisk ] };
key <AC10> { [ ntilde, Ntilde ] };
// no AC11 or AC12 on olpcm
key <AB08> { [ comma, semicolon ] };
key <AB09> { [ period, colon ] };
key <AB10> { [ minus, underscore ] };
key <AA02> { [ less, greater ] };
key <AA06> { [ bracketleft, braceleft, ccedilla, Ccedilla ] };
key <AA07> { [ bracketright, braceright ] };
include "level3(ralt_switch)"
};
partial alphanumeric_keys
xkb_symbols "deadtilde" {
include "es(basic)"
name[Group1]="Spanish (dead tilde)";
key <AE04> { [ 4, dollar, dead_tilde, dollar ] };
key <AC10> { [ ntilde, Ntilde, asciitilde, dead_doubleacute ] };
};
partial alphanumeric_keys
xkb_symbols "olpc2" {
// #HW-SPECIFIC
// Modified variant of US International layout, specifically for Peru
// Contact: Sayamindu Dasgupta <sayamindu@laptop.org>
include "us(olpc)"
name[Group1]="Spanish";
key <AE03> { [ 3, numbersign, dead_grave, dead_grave] }; // combining grave
key <I236> { [ XF86Start ] };
include "level3(ralt_switch)"
};
// EXTRAS:
partial alphanumeric_keys
xkb_symbols "sun_type6" {
include "sun_vndr/es(sun_type6)"
};
File diff suppressed because it is too large Load Diff
+249
View File
@@ -0,0 +1,249 @@
// Keyboard layouts for Great Britain.
default partial alphanumeric_keys
xkb_symbols "basic" {
// The basic UK layout, also known as the IBM 166 layout,
// but with the useless brokenbar pushed two levels up.
include "latin"
name[Group1]="English (UK)";
key <TLDE> { [ grave, notsign, bar, bar ] };
key <AE02> { [ 2, quotedbl, twosuperior, oneeighth ] };
key <AE03> { [ 3, sterling, threesuperior, sterling ] };
key <AE04> { [ 4, dollar, EuroSign, onequarter ] };
key <AC11> { [apostrophe, at, dead_circumflex, dead_caron] };
key <BKSL> { [numbersign, asciitilde, dead_grave, dead_breve ] };
key <LSGT> { [ backslash, bar, bar, brokenbar ] };
include "level3(ralt_switch)"
};
partial alphanumeric_keys
xkb_symbols "intl" {
// A UK layout but with five accents made into dead keys:
// grave, diaeresis, circumflex, acute, and tilde.
// By Phil Jones <philjones1 at blueyonder.co.uk>.
include "latin"
name[Group1]="English (UK, intl., with dead keys)";
key <TLDE> { [ dead_grave, notsign, bar, bar ] };
key <AE02> { [ 2, dead_diaeresis, twosuperior, onehalf ] };
key <AE03> { [ 3, sterling, threesuperior, onethird ] };
key <AE04> { [ 4, dollar, EuroSign, onequarter ] };
key <AE06> { [ 6, dead_circumflex, threequarters, onesixth ] };
key <AC11> { [ dead_acute, at, apostrophe, bar ] };
key <BKSL> { [ numbersign, dead_tilde, bar, bar ] };
key <LSGT> { [ backslash, bar, bar, bar ] };
key <AB08> { [ comma, less, ccedilla, Ccedilla ] };
include "level3(ralt_switch)"
};
partial alphanumeric_keys
xkb_symbols "extd" {
// Clone of the Microsoft "United Kingdom Extended" layout, which
// includes dead keys for: grave; diaeresis; circumflex; tilde; and
// accute. It also enables direct access to accute characters using
// the Multi_key (Alt Gr).
//
// Taken from...
// "Windows Keyboard Layouts"
// https://docs.microsoft.com/en-gb/globalization/windows-keyboard-layouts#U
//
// -- Jonathan Miles <jon@cybah.co.uk>
include "latin"
name[Group1]="English (UK, extended, Windows)";
key <TLDE> { [ dead_grave, notsign, brokenbar, NoSymbol ] };
key <AE02> { [ 2, quotedbl, dead_diaeresis, onehalf ] };
key <AE03> { [ 3, sterling, threesuperior, onethird ] };
key <AE04> { [ 4, dollar, EuroSign, onequarter ] };
key <AE06> { [ 6, asciicircum, dead_circumflex, NoSymbol ] };
key <AD02> { [ w, W, wacute, Wacute ] };
key <AD03> { [ e, E, eacute, Eacute ] };
key <AD06> { [ y, Y, yacute, Yacute ] };
key <AD07> { [ u, U, uacute, Uacute ] };
key <AD08> { [ i, I, iacute, Iacute ] };
key <AD09> { [ o, O, oacute, Oacute ] };
key <AD12> { [ bracketright, braceright, NoSymbol, bar ] };
key <AC01> { [ a, A, aacute, Aacute ] };
key <AC11> { [ apostrophe, at, dead_acute, grave ] };
key <BKSL> { [ numbersign, asciitilde, dead_tilde, backslash ] };
key <LSGT> { [ backslash, bar, NoSymbol, NoSymbol ] };
key <AB03> { [ c, C, ccedilla, Ccedilla ] };
include "level3(ralt_switch)"
};
// Describe the differences between the US Colemak layout
// and a UK variant. By Andy Buckley (andy@insectnation.org)
partial alphanumeric_keys
xkb_symbols "colemak" {
include "us(colemak)"
name[Group1]="English (UK, Colemak)";
key <TLDE> { [ grave, notsign, bar, asciitilde ] };
key <AE02> { [ 2, quotedbl, twosuperior, oneeighth ] };
key <AE03> { [ 3, sterling, threesuperior, sterling ] };
key <AE04> { [ 4, dollar, EuroSign, onequarter ] };
key <AC11> { [apostrophe, at, dead_circumflex, dead_caron] };
key <BKSL> { [numbersign, asciitilde, dead_grave, dead_breve ] };
key <LSGT> { [ backslash, bar, asciitilde, brokenbar ] };
};
// Colemak-DH (ISO) layout, UK Variant, https://colemakmods.github.io/mod-dh/
partial alphanumeric_keys
xkb_symbols "colemak_dh" {
include "us(colemak_dh)"
name[Group1]="English (UK, Colemak-DH)";
key <TLDE> { [ grave, notsign, bar, asciitilde ] };
key <AE02> { [ 2, quotedbl, twosuperior, oneeighth ] };
key <AE03> { [ 3, sterling, threesuperior, sterling ] };
key <AE04> { [ 4, dollar, EuroSign, onequarter ] };
key <AC11> { [apostrophe, at, dead_circumflex, dead_caron] };
key <BKSL> { [numbersign, asciitilde, dead_grave, dead_breve ] };
key <AB05> { [ backslash, bar, asciitilde, brokenbar ] };
};
// Dvorak (UK) keymap (by odaen) allowing the usage of
// the £ and ? key and swapping the @ and " keys.
partial alphanumeric_keys
xkb_symbols "dvorak" {
include "us(dvorak-alt-intl)"
name[Group1]="English (UK, Dvorak)";
key <TLDE> { [ grave, notsign, bar, bar ] };
key <AE02> { [ 2, quotedbl, twosuperior, NoSymbol ] };
key <AE03> { [ 3, sterling, threesuperior, NoSymbol ] };
key <AD01> { [ apostrophe, at ] };
key <BKSL> { [ numbersign, asciitilde ] };
key <LSGT> { [ backslash, bar ] };
};
// Dvorak letter positions, but punctuation all in the normal UK positions.
partial alphanumeric_keys
xkb_symbols "dvorakukp" {
include "gb(dvorak)"
name[Group1]="English (UK, Dvorak, with UK punctuation)";
key <AE11> { [ minus, underscore ] };
key <AE12> { [ equal, plus ] };
key <AD11> { [ bracketleft, braceleft ] };
key <AD12> { [ bracketright, braceright ] };
key <AD01> { [ slash, question ] };
key <AC11> { [apostrophe, at, dead_circumflex, dead_caron] };
};
partial alphanumeric_keys
xkb_symbols "mac" {
include "latin"
name[Group1]= "English (UK, Macintosh)";
key <TLDE> { [ section, plusminus ] };
key <AE02> { [ 2, at, EuroSign ] };
key <AE03> { [ 3, sterling, numbersign ] };
key <LSGT> { [ grave, asciitilde ] };
include "level3(ralt_switch)"
include "level3(enter_switch)"
};
partial alphanumeric_keys
xkb_symbols "mac_intl" {
include "latin"
name[Group1]="English (UK, Macintosh, intl.)";
key <TLDE> { [ section, plusminus, notsign, notsign ] }; //dead_grave
key <AE02> { [ 2, at, EuroSign, onehalf ] };
key <AE03> { [ 3, sterling, twosuperior, onethird ] };
key <AE04> { [ 4, dollar, threesuperior, onequarter ] };
key <AE06> { [ 6, dead_circumflex, NoSymbol, onesixth ] };
key <AD09> { [ o, O, oe, OE ] };
key <AC11> { [ dead_acute, dead_diaeresis, dead_diaeresis, bar ] }; //dead_doubleacute
key <BKSL> { [ backslash, bar, numbersign, bar ] };
key <LSGT> { [ dead_grave, dead_tilde, brokenbar, bar ] };
include "level3(ralt_switch)"
};
partial alphanumeric_keys
xkb_symbols "pl" {
// Polish accented letters on upper levels of corresponding base letters.
// Idea from Wawrzyniec Niewodniczański, adapted by Aleksander Kowalski.
include "gb(basic)"
name[Group1]="Polish (British keyboard)";
key <AD03> { [ e, E, eogonek, Eogonek ] };
key <AD09> { [ o, O, oacute, Oacute ] };
key <AC01> { [ a, A, aogonek, Aogonek ] };
key <AC02> { [ s, S, sacute, Sacute ] };
key <AB01> { [ z, Z, zabovedot, Zabovedot ] };
key <AB02> { [ x, X, zacute, Zacute ] };
key <AB03> { [ c, C, cacute, Cacute ] };
key <AB06> { [ n, N, nacute, Nacute ] };
};
partial alphanumeric_keys
xkb_symbols "gla" {
// Grave-accented letters on the upper levels of the relevant vowels.
include "gb(basic)"
name[Group1]="Scottish Gaelic";
key <AD03> { [ e, E, egrave, Egrave ] };
key <AD07> { [ u, U, ugrave, Ugrave ] };
key <AD08> { [ i, I, igrave, Igrave ] };
key <AD09> { [ o, O, ograve, Ograve ] };
key <AC01> { [ a, A, agrave, Agrave ] };
};
// EXTRAS:
partial alphanumeric_keys
xkb_symbols "sun_type6" {
include "sun_vndr/gb(sun_type6)"
};
+102
View File
@@ -0,0 +1,102 @@
// The <KPDL> key is a mess.
// It was probably originally meant to be a decimal separator.
// Except since it was declared by USA people it didn't use the original
// SI separator "," but a "." (since then the USA managed to f-up the SI
// by making "." an accepted alternative, but standards still use "," as
// default)
// As a result users of SI-abiding countries expect either a "." or a ","
// or a "decimal_separator" which may or may not be translated in one of the
// above depending on applications.
// It's not possible to define a default per-country since user expectations
// depend on the conflicting choices of their most-used applications,
// operating system, etc. Therefore it needs to be a configuration setting
// Copyright © 2007 Nicolas Mailhot <nicolas.mailhot @ laposte.net>
// Legacy <KPDL> #1
// This assumes KP_Decimal will be translated in a dot
partial keypad_keys
xkb_symbols "dot" {
key.type[Group1]="KEYPAD" ;
key <KPDL> { [ KP_Delete, KP_Decimal ] }; // <delete> <separator>
};
// Legacy <KPDL> #2
// This assumes KP_Separator will be translated in a comma
partial keypad_keys
xkb_symbols "comma" {
key.type[Group1]="KEYPAD" ;
key <KPDL> { [ KP_Delete, KP_Separator ] }; // <delete> <separator>
};
// Period <KPDL>, usual keyboard serigraphy in most countries
partial keypad_keys
xkb_symbols "dotoss" {
key.type[Group1]="FOUR_LEVEL_MIXED_KEYPAD" ;
key <KPDL> { [ KP_Delete, period, comma, 0x100202F ] }; // <delete> . , ⍽ (narrow no-break space)
};
// Period <KPDL>, usual keyboard serigraphy in most countries, latin-9 restriction
partial keypad_keys
xkb_symbols "dotoss_latin9" {
key.type[Group1]="FOUR_LEVEL_MIXED_KEYPAD" ;
key <KPDL> { [ KP_Delete, period, comma, nobreakspace ] }; // <delete> . , ⍽ (no-break space)
};
// Comma <KPDL>, what most non anglo-saxon people consider the real separator
partial keypad_keys
xkb_symbols "commaoss" {
key.type[Group1]="FOUR_LEVEL_MIXED_KEYPAD" ;
key <KPDL> { [ KP_Delete, comma, period, 0x100202F ] }; // <delete> , . ⍽ (narrow no-break space)
};
// Momayyez <KPDL>: Bahrain, Iran, Iraq, Kuwait, Oman, Qatar, Saudi Arabia, Syria, UAE
partial keypad_keys
xkb_symbols "momayyezoss" {
key.type[Group1]="FOUR_LEVEL_MIXED_KEYPAD" ;
key <KPDL> { [ KP_Delete, 0x100066B, comma, 0x100202F ] }; // <delete> ? , ⍽ (narrow no-break space)
};
// Abstracted <KPDL>, pray everything will work out (it usually does not)
partial keypad_keys
xkb_symbols "kposs" {
key.type[Group1]="FOUR_LEVEL_MIXED_KEYPAD" ;
key <KPDL> { [ KP_Delete, KP_Decimal, KP_Separator, 0x100202F ] }; // <delete> ? ? ⍽ (narrow no-break space)
};
// Spreadsheets may be configured to use the dot as decimal
// punctuation, comma as a thousands separator and then semi-colon as
// the list separator. Of these, dot and semi-colon is most important
// when entering data by the keyboard; the comma can then be inferred
// and added to the presentation afterwards. Using semi-colon as a
// general separator may in fact be preferred to avoid ambiguities
// in data files. Most times a decimal separator is hard-coded, it
// seems to be period, probably since this is the syntax used in
// (most) programming languages.
partial keypad_keys
xkb_symbols "semi" {
key.type[Group1]="FOUR_LEVEL_MIXED_KEYPAD" ;
key <KPDL> { [ NoSymbol, NoSymbol, semicolon ] };
};
+255
View File
@@ -0,0 +1,255 @@
// Common Latin alphabet layout
default partial
xkb_symbols "basic" {
key <AE01> { [ 1, exclam, onesuperior, exclamdown ] };
key <AE02> { [ 2, at, twosuperior, oneeighth ] };
key <AE03> { [ 3, numbersign, threesuperior, sterling ] };
key <AE04> { [ 4, dollar, onequarter, dollar ] };
key <AE05> { [ 5, percent, onehalf, threeeighths ] };
key <AE06> { [ 6, asciicircum, threequarters, fiveeighths ] };
key <AE07> { [ 7, ampersand, braceleft, seveneighths ] };
key <AE08> { [ 8, asterisk, bracketleft, trademark ] };
key <AE09> { [ 9, parenleft, bracketright, plusminus ] };
key <AE10> { [ 0, parenright, braceright, degree ] };
key <AE11> { [ minus, underscore, backslash, questiondown ] };
key <AE12> { [ equal, plus, dead_cedilla, dead_ogonek ] };
key <AD01> { [ q, Q, at, Greek_OMEGA ] };
key <AD02> { [ w, W, U017F, section ] };
key <AD03> { [ e, E, e, E ] };
key <AD04> { [ r, R, paragraph, registered ] };
key <AD05> { [ t, T, tslash, Tslash ] };
key <AD06> { [ y, Y, leftarrow, yen ] };
key <AD07> { [ u, U, downarrow, uparrow ] };
key <AD08> { [ i, I, rightarrow, idotless ] };
key <AD09> { [ o, O, oslash, Oslash ] };
key <AD10> { [ p, P, thorn, THORN ] };
key <AD11> { [bracketleft, braceleft, dead_diaeresis, dead_abovering ] };
key <AD12> { [bracketright, braceright, dead_tilde, dead_macron ] };
key <AC01> { [ a, A, ae, AE ] };
key <AC02> { [ s, S, ssharp, U1E9E ] };
key <AC03> { [ d, D, eth, ETH ] };
key <AC04> { [ f, F, dstroke, ordfeminine ] };
key <AC05> { [ g, G, eng, ENG ] };
key <AC06> { [ h, H, hstroke, Hstroke ] };
key <AC07> { [ j, J, dead_hook, dead_horn ] };
key <AC08> { [ k, K, kra, ampersand ] };
key <AC09> { [ l, L, lstroke, Lstroke ] };
key <AC10> { [ semicolon, colon, dead_acute, dead_doubleacute ] };
key <AC11> { [apostrophe, quotedbl, dead_circumflex, dead_caron ] };
key <TLDE> { [ grave, asciitilde, notsign, notsign ] };
key <BKSL> { [ backslash, bar, dead_grave, dead_breve ] };
key <AB01> { [ z, Z, guillemotleft, less ] };
key <AB02> { [ x, X, guillemotright, greater ] };
key <AB03> { [ c, C, cent, copyright ] };
key <AB04> { [ v, V, doublelowquotemark, singlelowquotemark ] };
key <AB05> { [ b, B, leftdoublequotemark, leftsinglequotemark ] };
key <AB06> { [ n, N, rightdoublequotemark, rightsinglequotemark ] };
key <AB07> { [ m, M, mu, masculine ] };
key <AB08> { [ comma, less, U2022, multiply ] }; // bullet
key <AB09> { [ period, greater, periodcentered, division ] };
key <AB10> { [ slash, question, dead_belowdot, dead_abovedot ] };
};
// Northern Europe ( Danish, Finnish, Norwegian, Swedish) common layout
partial
xkb_symbols "type2" {
include "latin"
key <AE01> { [ 1, exclam, exclamdown, onesuperior ] };
key <AE02> { [ 2, quotedbl, at, twosuperior ] };
key <AE03> { [ 3, numbersign, sterling, threesuperior] };
key <AE04> { [ 4, currency, dollar, onequarter ] };
key <AE05> { [ 5, percent, onehalf, cent ] };
key <AE06> { [ 6, ampersand, yen, fiveeighths ] };
key <AE07> { [ 7, slash, braceleft, division ] };
key <AE08> { [ 8, parenleft, bracketleft, guillemotleft] };
key <AE09> { [ 9, parenright, bracketright, guillemotright] };
key <AE10> { [ 0, equal, braceright, degree ] };
key <AD03> { [ e, E, EuroSign, cent ] };
key <AD04> { [ r, R, registered, registered ] };
key <AD05> { [ t, T, thorn, THORN ] };
key <AD09> { [ o, O, oe, OE ] };
key <AD11> { [ aring, Aring, dead_diaeresis, dead_abovering ] };
key <AD12> { [dead_diaeresis, dead_circumflex, dead_tilde, dead_caron ] };
key <AC01> { [ a, A, ordfeminine, masculine ] };
key <AB03> { [ c, C, copyright, copyright ] };
key <AB08> { [ comma, semicolon, dead_cedilla, dead_ogonek ] };
key <AB09> { [ period, colon, periodcentered, dead_abovedot ] };
key <AB10> { [ minus, underscore, dead_belowdot, dead_abovedot ] };
};
// Slavic Latin ( Albanian, Croatian, Polish, Slovene, Yugoslav)
// common layout
partial
xkb_symbols "type3" {
include "latin"
key <AD01> { [ q, Q, backslash, Greek_OMEGA ] };
key <AD02> { [ w, W, bar, section ] };
key <AD06> { [ z, Z, leftarrow, yen ] };
key <AC04> { [ f, F, bracketleft, ordfeminine ] };
key <AC05> { [ g, G, bracketright, ENG ] };
key <AC08> { [ k, K, lstroke, ampersand ] };
key <AB01> { [ y, Y, guillemotleft, less ] };
key <AB04> { [ v, V, at, grave ] };
key <AB05> { [ b, B, braceleft, apostrophe ] };
key <AB06> { [ n, N, braceright, acute ] };
key <AB07> { [ m, M, section, masculine ] };
key <AB08> { [ comma, semicolon, less, multiply ] };
key <AB09> { [ period, colon, greater, division ] };
};
// Another common Latin layout
// (German, Estonian, Spanish, Icelandic, Italian, Latin American, Portuguese)
partial
xkb_symbols "type4" {
include "latin"
key <AE02> { [ 2, quotedbl, at, oneeighth ] };
key <AE06> { [ 6, ampersand, notsign, fiveeighths ] };
key <AE07> { [ 7, slash, braceleft, seveneighths ] };
key <AE08> { [ 8, parenleft, bracketleft, trademark ] };
key <AE09> { [ 9, parenright, bracketright, plusminus ] };
key <AE10> { [ 0, equal, braceright, degree ] };
key <AD03> { [ e, E, EuroSign, cent ] };
key <AB08> { [ comma, semicolon, U2022, multiply ] }; // bullet
key <AB09> { [ period, colon, periodcentered, division ] };
key <AB10> { [ minus, underscore, dead_belowdot, dead_abovedot ] };
};
partial
xkb_symbols "nodeadkeys" {
key <AE12> { [ equal, plus, cedilla, ogonek ] };
key <AD11> { [bracketleft, braceleft, diaeresis, degree ] };
key <AD12> { [bracketright, braceright, asciitilde, macron ] };
key <AC07> { [ j, J, ezh, EZH ] };
key <AC10> { [ semicolon, colon, acute, doubleacute ] };
key <AC11> { [apostrophe, quotedbl, asciicircum, caron ] };
key <BKSL> { [ backslash, bar, grave, breve ] };
key <AB10> { [ slash, question, ellipsis, abovedot ] };
};
partial
xkb_symbols "type2_nodeadkeys" {
include "latin(nodeadkeys)"
key <AD11> { [ aring, Aring, diaeresis, degree ] };
key <AD12> { [ diaeresis, asciicircum, asciitilde, caron ] };
key <AB08> { [ comma, semicolon, cedilla, ogonek ] };
key <AB09> { [ period, colon, periodcentered, abovedot ] };
key <AB10> { [ minus, underscore, ellipsis, abovedot ] };
};
partial
xkb_symbols "type3_nodeadkeys" {
include "latin(nodeadkeys)"
};
partial
xkb_symbols "type4_nodeadkeys" {
include "latin(nodeadkeys)"
key <AB10> { [ minus, underscore, ellipsis, abovedot ] };
};
// Added 2008.03.05 by Marcin Woliński
// See http://marcinwolinski.pl/keyboard/ for a description.
// Used by pl(intl)
//
// ┌─────┐
// │ 2 4 │ 2 = Shift, 4 = Level3 + Shift
// │ 1 3 │ 1 = Normal, 3 = Level3
// └─────┘
// ┌─────┬─────┬─────┬─────┬─────┬─────┬─────┬─────┬─────┬─────┬─────┬─────┬─────┲━━━━━━━━━┓
// │ ~ ~ │ ! ' │ @ " │ # ˝ │ $ ¸ │ % ˇ │ ^ ^ │ & ˘ │ * ̇ │ ( ̣ │ ) ° │ _ ¯ │ + ˛ ┃ ⌫ Back- ┃
// │ ` ` │ 1 ¡ │ 2 © │ 3 • │ 4 § │ 5 € │ 6 ¢ │ 7 − │ 8 × │ 9 ÷ │ 0 ° │ - – │ = — ┃ space ┃
// ┢━━━━━┷━┱───┴─┬───┴─┬───┴─┬───┴─┬───┴─┬───┴─┬───┴─┬───┴─┬───┴─┬───┴─┬───┴─┬───┺━┳━━━━━━━┫
// ┃ ┃ Q │ W │ E │ R │ T │ Y │ U │ I │ O │ P │ { « │ } » ┃ Enter ┃
// ┃Tab ↹ ┃ q │ w │ e │ r │ t │ y │ u │ i │ o │ p │ [ ‹ │ ] › ┃ ⏎ ┃
// ┣━━━━━━━┻┱────┴┬────┴┬────┴┬────┴┬────┴┬────┴┬────┴┬────┴┬────┴┬────┴┬────┴┬────┺┓ ┃
// ┃ ┃ A │ S │ D │ F │ G │ H │ J │ K │ L │ : “ │ " ” │ | ¶ ┃ ┃
// ┃Caps ⇬ ┃ a │ s │ d │ f │ g │ h │ j │ k │ l │ ; ‘ │ ' ’ │ \ ┃ ┃
// ┣━━━━━━━━┹────┬┴────┬┴────┬┴────┬┴────┬┴────┬┴────┬┴────┬┴────┬┴────┬┴────┲┷━━━━━┻━━━━━━┫
// ┃ │ Z │ X │ C │ V │ B │ N │ M │ < „ │ > · │ ? ¿ ┃ ┃
// ┃Shift ⇧ │ z │ x │ c │ v │ b │ n │ m │ , ‚ │ . … │ / ⁄ ┃Shift ⇧ ┃
// ┣━━━━━━━┳━━━━━┷━┳━━━┷━━━┱─┴─────┴─────┴─────┴─────┴─────┴───┲━┷━━━━━╈━━━━━┻━┳━━━━━━━┳━━━┛
// ┃ ┃ ┃ ┃ ␣ ⍽ ┃ ┃ ┃ ┃
// ┃Ctrl ┃Meta ┃Alt ┃ ␣ Space ⍽ ┃AltGr ⇮┃Menu ┃Ctrl ┃
// ┗━━━━━━━┻━━━━━━━┻━━━━━━━┹───────────────────────────────────┺━━━━━━━┻━━━━━━━┻━━━━━━━┛
partial
xkb_symbols "intl" {
key <TLDE> { [ grave, asciitilde, dead_grave, dead_tilde ] };
key <AE01> { [ 1, exclam, exclamdown, dead_acute ] };
key <AE02> { [ 2, at, copyright, dead_diaeresis ] };
key <AE03> { [ 3, numbersign, U2022, dead_doubleacute ] }; // U+2022 is bullet (the name bullet does not work)
key <AE04> { [ 4, dollar, section, dead_cedilla ] };
key <AE05> { [ 5, percent, EuroSign, dead_caron ] };
key <AE06> { [ 6, asciicircum, cent, dead_circumflex ] };
key <AE07> { [ 7, ampersand, U2212, dead_breve ] }; // U+2212 is MINUS SIGN
key <AE08> { [ 8, asterisk, multiply, dead_abovedot ] };
key <AE09> { [ 9, parenleft, division, dead_belowdot ] };
key <AE10> { [ 0, parenright, degree, dead_abovering ] };
key <AE11> { [ minus, underscore, endash, dead_macron ] };
key <AE12> { [ equal, plus, emdash, dead_ogonek ] };
key <AD01> { [ q, Q ] };
key <AD02> { [ w, W ] };
key <AD03> { [ e, E ] };
key <AD04> { [ r, R ] };
key <AD05> { [ t, T ] };
key <AD06> { [ y, Y ] };
key <AD07> { [ u, U ] };
key <AD08> { [ i, I ] };
key <AD09> { [ o, O ] };
key <AD10> { [ p, P ] };
key <AD11> { [bracketleft, braceleft, U2039, guillemotleft ] };
key <AD12> { [bracketright, braceright, U203A, guillemotright ] };
key <AC01> { [ a, A ] };
key <AC02> { [ s, S ] };
key <AC03> { [ d, D ] };
key <AC04> { [ f, F ] };
key <AC05> { [ g, G ] };
key <AC06> { [ h, H ] };
key <AC07> { [ j, J ] };
key <AC08> { [ k, K ] };
key <AC09> { [ l, L ] };
key <AC10> { [ semicolon, colon, leftsinglequotemark, leftdoublequotemark ] };
key <AC11> { [apostrophe, quotedbl, rightsinglequotemark, rightdoublequotemark ] };
key <BKSL> { [ backslash, bar, NoSymbol, paragraph ] };
key <AB01> { [ z, Z ] };
key <AB02> { [ x, X ] };
key <AB03> { [ c, C ] };
key <AB04> { [ v, V ] };
key <AB05> { [ b, B ] };
key <AB06> { [ n, N ] };
key <AB07> { [ m, M ] };
key <AB08> { [ comma, less, singlelowquotemark, doublelowquotemark ] };
key <AB09> { [ period, greater, ellipsis, periodcentered ] };
key <AB10> { [ slash, question, U2044, questiondown ] }; // U+2044 is FRACTION SLASH
};
+156
View File
@@ -0,0 +1,156 @@
// These variants assign ISO_Level3_Shift to various keys
// so that levels 3 and 4 can be reached.
// The default behaviour:
// the right Alt key (AltGr) chooses the third symbol engraved on a key.
default partial modifier_keys
xkb_symbols "ralt_switch" {
key <RALT> {[ ISO_Level3_Shift ], type[group1]="ONE_LEVEL" };
};
// The right Alt key never chooses the third level.
// This option attempts to undo the effect of a layout's inclusion of
// 'ralt_switch'. You may want to also select another level3 option
// to map the level3 shift to some other key.
partial modifier_keys
xkb_symbols "ralt_alt" {
key <RALT> {[ Alt_R, Meta_R ], type[group1]="TWO_LEVEL" };
modifier_map Mod1 { <RALT> };
};
// The right Alt key (while pressed) chooses the third shift level,
// and Compose is mapped to its second level.
partial modifier_keys
xkb_symbols "ralt_switch_multikey" {
key <RALT> {[ ISO_Level3_Shift, Multi_key ], type[group1]="TWO_LEVEL" };
};
// Either Alt key (while pressed) chooses the third shift level.
// (To be used mostly to imitate Mac OS functionality.)
partial modifier_keys
xkb_symbols "alt_switch" {
include "level3(lalt_switch)"
include "level3(ralt_switch)"
};
// The left Alt key (while pressed) chooses the third shift level.
partial modifier_keys
xkb_symbols "lalt_switch" {
key <LALT> {[ ISO_Level3_Shift ], type[group1]="ONE_LEVEL" };
};
// The right Ctrl key (while pressed) chooses the third shift level.
partial modifier_keys
xkb_symbols "switch" {
key <RCTL> {[ ISO_Level3_Shift ], type[group1]="ONE_LEVEL" };
};
// The Menu key (while pressed) chooses the third shift level.
partial modifier_keys
xkb_symbols "menu_switch" {
key <MENU> {[ ISO_Level3_Shift ], type[group1]="ONE_LEVEL" };
};
// Either Win key (while pressed) chooses the third shift level.
partial modifier_keys
xkb_symbols "win_switch" {
include "level3(lwin_switch)"
include "level3(rwin_switch)"
};
// The left Win key (while pressed) chooses the third shift level.
partial modifier_keys
xkb_symbols "lwin_switch" {
key <LWIN> {[ ISO_Level3_Shift ], type[group1]="ONE_LEVEL" };
};
// The right Win key (while pressed) chooses the third shift level.
partial modifier_keys
xkb_symbols "rwin_switch" {
key <RWIN> {[ ISO_Level3_Shift ], type[group1]="ONE_LEVEL" };
};
// The Enter key on the kepypad (while pressed) chooses the third shift level.
// (This is especially useful for Mac laptops which miss the right Alt key.)
partial modifier_keys
xkb_symbols "enter_switch" {
key <KPEN> {[ ISO_Level3_Shift ], type[group1]="ONE_LEVEL" };
};
// The CapsLock key (while pressed) chooses the third shift level.
partial modifier_keys
xkb_symbols "caps_switch" {
key <CAPS> {[ ISO_Level3_Shift ], type[group1]="ONE_LEVEL" };
};
// The CapsLock key (while pressed) chooses the third shift level and
// Ctrl + CapsLock has the original CapsLock function.
// The 2023 DIN standard for German keyboards recommends it as an option:
// - https://de.wikipedia.org/wiki/E1_(Tastaturbelegung)#Feststelltaste/Umschaltsperre
// - https://en.wikipedia.org/wiki/Caps_Lock#Abolition
partial modifier_keys
xkb_symbols "caps_switch_capslock_with_ctrl" {
virtual_modifiers LevelThree;
key <CAPS> {
type[Group1] = "PC_CONTROL_LEVEL2",
symbols[Group1] = [ ISO_Level3_Shift, Caps_Lock ],
// Explicit actions are preferred over modMap None/Mod5 { Caps_Lock }
// because they have no side effect
actions[Group1] = [ SetMods(modifiers = LevelThree), LockMods(modifiers = Lock) ]
};
};
// The Backslash key (while pressed) chooses the third shift level.
partial modifier_keys
xkb_symbols "bksl_switch" {
key <BKSL> {[ ISO_Level3_Shift ], type[group1]="ONE_LEVEL" };
};
// The AC11 key (while pressed) chooses the third shift level.
partial modifier_keys
xkb_symbols "ac11_switch" {
key <AC11> {[ ISO_Level3_Shift ], type[Group1]="ONE_LEVEL" };
};
// The Less/Greater key (while pressed) chooses the third shift level.
partial modifier_keys
xkb_symbols "lsgt_switch" {
key <LSGT> {[ ISO_Level3_Shift ], type[group1]="ONE_LEVEL" };
};
// The CapsLock key (while pressed) chooses the third shift level,
// and latches when pressed together with another third-level chooser.
partial modifier_keys
xkb_symbols "caps_switch_latch" {
key <CAPS> {[ ISO_Level3_Shift, ISO_Level3_Shift, ISO_Level3_Latch ],
type[group1]="THREE_LEVEL" };
};
// The Backslash key (while pressed) chooses the third shift level,
// and latches when pressed together with another third-level chooser.
partial modifier_keys
xkb_symbols "bksl_switch_latch" {
key <BKSL> {[ ISO_Level3_Shift, ISO_Level3_Shift, ISO_Level3_Latch ],
type[group1]="THREE_LEVEL" };
};
// The Less/Greater key (while pressed) chooses the third shift level,
// and latches when pressed together with another third-level chooser.
partial modifier_keys
xkb_symbols "lsgt_switch_latch" {
key <LSGT> {[ ISO_Level3_Shift, ISO_Level3_Shift, ISO_Level3_Latch ],
type[group1]="THREE_LEVEL" };
};
// Top-row digit key 4 chooses third shift level when pressed alone.
partial modifier_keys
xkb_symbols "4_switch_isolated" {
override key <AE04> {[ ISO_Level3_Shift ]};
};
// Top-row digit key 9 chooses third shift level when pressed alone.
partial modifier_keys
xkb_symbols "9_switch_isolated" {
override key <AE09> {[ ISO_Level3_Shift ]};
};
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,153 @@
//! xkeyboard-config — keyboard layouts, compiled from the X11 xkeyboard-config database
//! into native Zig. It turns a physical key (a USB HID usage, as the input module delivers)
//! plus a modifier state into a **keysym** and, when the key produces one, a **character**
//! (a Unicode scalar). This is the piece that lets a `KeyEvent.keycode` become a
//! `KeyEvent.character`, without shipping an X11 runtime.
//!
//! The layout tables in `generated/layouts.zig` are produced by
//! `tools/make-xkeyboard-config.py` (see ./README.md to regenerate). Those tables are
//! deliberately pure data — each key carries its up-to-four levels and an XKB *type*. The
//! type -> level selection semantics (which modifier picks which level) live here, so the
//! data and the policy are separable.
//!
//! Scope (documented in README.md): group 1 only, no dead-key/compose composition (a dead
//! key returns its keysym with no character), and a curated set of key types. Layouts:
//! us, gb, de, fr, es, dvorak.
//!
//! Upstream xkeyboard-config and keysymdef.h are MIT/X11 licensed; see vendor/COPYING and
//! vendor/PROVENANCE.md.
const std = @import("std");
const generated = @import("layouts");
pub const Level = generated.Level;
pub const KeyType = generated.KeyType;
pub const Key = generated.Key;
pub const Layout = generated.Layout;
/// The generated layouts, by name — as pointers, so they share identity with `all` and
/// `byName` (and match the `*const Layout` that `map` takes).
pub const us: *const Layout = &generated.us;
pub const gb: *const Layout = &generated.gb;
pub const de: *const Layout = &generated.de;
pub const fr: *const Layout = &generated.fr;
pub const es: *const Layout = &generated.es;
pub const dvorak: *const Layout = &generated.dvorak;
/// Every generated layout, for enumeration (e.g. a settings UI).
pub const all = generated.all;
/// The modifier state that selects a key's level. `level3` is AltGr (ISO Level3 Shift);
/// `control` is accepted for completeness but does not affect level selection here.
pub const Modifiers = struct {
shift: bool = false,
caps_lock: bool = false,
level3: bool = false,
control: bool = false,
};
/// The result of a lookup: the X11 `keysym`, and the `character` it produces (a Unicode
/// scalar) when it is a printable key — null for keys that produce none (Return, F1, a
/// bare dead key, an unmapped key).
pub const Mapping = struct {
keysym: u32,
character: ?u21,
};
/// Which level (0..3) a key of `kind` selects under `mods`. XKB's canonical semantics:
/// Shift picks the odd level, AltGr (level3) adds 2, and Caps acts like Shift for the
/// alphabetic types. See the XKB "key types" — this covers the ones the vendored layouts
/// use; anything else falls back to shift-or-not.
fn selectLevel(kind: KeyType, mods: Modifiers) usize {
const shift_or_caps = mods.shift != mods.caps_lock; // XOR: Caps behaves like Shift
const low: usize = if (mods.shift) 1 else 0;
const high: usize = if (mods.level3) 2 else 0;
return switch (kind) {
.one_level => 0,
.two_level, .keypad, .other => low,
.alphabetic => if (shift_or_caps) 1 else 0,
.four_level => low + high,
.four_level_alphabetic => (if (shift_or_caps) @as(usize, 1) else 0) + high,
// Caps affects only the base pair, not the AltGr pair.
.four_level_semialphabetic => if (mods.level3) 2 + low else (if (shift_or_caps) @as(usize, 1) else 0),
};
}
/// Map a physical key (`hid_usage`, a USB HID keyboard-page usage) under `mods` on
/// `layout` to its keysym and character. Falls back gracefully when the selected level is
/// undefined for the key: it drops the AltGr component, then the shift component, so a key
/// with only a base/shift pair still yields something sensible under AltGr.
pub fn map(layout: *const Layout, hid_usage: u8, mods: Modifiers) Mapping {
const key = &layout.keys[hid_usage];
var level = selectLevel(key.kind, mods);
// Fall back to a defined level: full -> without AltGr -> base.
if (key.levels[level].keysym == 0 and key.levels[level].unicode == 0) {
const candidates = [_]usize{ level & 1, 0 };
for (candidates) |candidate| {
if (key.levels[candidate].keysym != 0 or key.levels[candidate].unicode != 0) {
level = candidate;
break;
}
}
}
const chosen = key.levels[level];
return .{
.keysym = chosen.keysym,
.character = if (chosen.unicode != 0) @intCast(chosen.unicode) else null,
};
}
/// Look up a layout by its name (`"us"`, `"gb"`, ...), or null if unknown.
pub fn byName(name: []const u8) ?*const Layout {
for (all) |layout| {
if (std.mem.eql(u8, layout.name, name)) return layout;
}
return null;
}
// --- tests (host-run via `zig build test`) ---------------------------------
const testing = std.testing;
// USB HID usages used in the tests (keyboard page 0x07).
const hid_a: u8 = 0x04;
const hid_1: u8 = 0x1e;
const hid_3: u8 = 0x20;
test "us: letters obey shift and caps" {
try testing.expectEqual(@as(?u21, 'a'), map(us, hid_a, .{}).character);
try testing.expectEqual(@as(?u21, 'A'), map(us, hid_a, .{ .shift = true }).character);
try testing.expectEqual(@as(?u21, 'A'), map(us, hid_a, .{ .caps_lock = true }).character);
// Shift + Caps cancels for an alphabetic key.
try testing.expectEqual(@as(?u21, 'a'), map(us, hid_a, .{ .shift = true, .caps_lock = true }).character);
}
test "us: digits and their shifted symbols" {
try testing.expectEqual(@as(?u21, '1'), map(us, hid_1, .{}).character);
try testing.expectEqual(@as(?u21, '!'), map(us, hid_1, .{ .shift = true }).character);
try testing.expectEqual(@as(?u21, '3'), map(us, hid_3, .{}).character);
try testing.expectEqual(@as(?u21, '#'), map(us, hid_3, .{ .shift = true }).character);
// A digit is not alphabetic: Caps alone must not shift it.
try testing.expectEqual(@as(?u21, '3'), map(us, hid_3, .{ .caps_lock = true }).character);
}
test "layouts differ: GB pound vs US hash on shift+3" {
try testing.expectEqual(@as(?u21, '#'), map(us, hid_3, .{ .shift = true }).character);
try testing.expectEqual(@as(?u21, '£'), map(gb, hid_3, .{ .shift = true }).character);
}
test "french azerty places q where us has a" {
try testing.expectEqual(@as(?u21, 'q'), map(fr, hid_a, .{}).character);
try testing.expectEqual(@as(?u21, 'Q'), map(fr, hid_a, .{ .shift = true }).character);
}
test "byName resolves and rejects" {
try testing.expect(byName("us") == us);
try testing.expect(byName("gb") == gb);
try testing.expect(byName("nonsense") == null);
}
test "unmapped key yields no character" {
// HID 0x00 is not a key; every level is empty.
try testing.expectEqual(@as(?u21, null), map(us, 0x00, .{}).character);
}
+88 -15
View File
@@ -1,16 +1,23 @@
//! The **kernel ↔ user** ABI: the core contract every user program speaks to the
//! kernel — the system_call numbers, `mmap` protection flags, the page size those
//! calls work in, and the IPC name-registry ids and notification bit. Shared by the
//! kernel dispatcher (system/kernel/process.zig) and the user runtime library
//! (library/runtime/), so the two can never drift.
//! The **private kernel ↔ runtime** ABI: the raw system_call contract — the call
//! numbers, `mmap` protection flags, the page size those calls work in, and the IPC
//! name-registry ids and notification bit. Shared by the kernel dispatcher
//! (system/kernel/process.zig) and the user-space runtime library (library/runtime/),
//! so the two can never drift.
//!
//! This is the *core* ABI; the device half — `DeviceDescriptor` and friends, which
//! **Application code does not speak this.** danos programs call the `runtime` library —
//! the stable, danos-native ABI — and the runtime is the one thing that issues the
//! actual system calls (POSIX code layers over the runtime, never on this directly). It
//! is the same split as libSystem on macOS or win32 over the NT syscalls: the numbers
//! here are an implementation detail the runtime hides and may renumber, not a public
//! interface. See docs/coding-standards.md and library/runtime/.
//!
//! This is the *core* contract; the device half — `DeviceDescriptor` and friends, which
//! also cross this boundary — lives with the device sub-project as [[device-abi]]
//! (system/devices/device-abi.zig). The loader↔kernel handoff is [[boot-handoff]].
/// Page size every `mmap`/`munmap` grant and the boot memory map are measured in.
/// 4 KiB on every architecture danos targets so far. Part of the ABI because user
/// code aligns to it (grants are page-granular) and the kernel guarantees it.
/// 4 KiB on every architecture danos targets so far. Part of the ABI because the
/// runtime aligns to it (grants are page-granular) and the kernel guarantees it.
pub const page_size = 4096;
/// The kernel system_call numbers — the single source of truth shared by the kernel
@@ -25,7 +32,7 @@ pub const SystemCall = enum(u64) {
sleep = 3, // sleep(ms): block the caller for ms milliseconds
mmap = 4, // mmap(len, prot) -> base: grant zeroed, page-aligned user pages
munmap = 5, // munmap(base, len): release pages from a prior mmap
create_endpoint = 6, // create_endpoint() -> handle: a new IPC endpoint
create_ipc_endpoint = 6, // create_ipc_endpoint() -> handle: a new IPC endpoint
ipc_register = 7, // ipc_register(service_id, handle): publish an endpoint by well-known id
ipc_lookup = 8, // ipc_lookup(service_id) -> handle: find a published endpoint
ipc_call = 9, // ipc_call(h, message, len, reply, cap) -> reply_len: send + block for reply
@@ -36,22 +43,88 @@ pub const SystemCall = enum(u64) {
irq_bind = 14, // irq_bind(id, resource_index, endpoint): deliver a device IRQ as an IPC notification
irq_ack = 15, // irq_ack(id, resource_index): re-arm a bound IRQ after servicing it
device_register = 16, // device_register(parent_id, descriptor) -> id: publish a child of a device you claimed
system_spawn = 17, // system_spawn(name_ptr, name_len) -> 0: start a named initial-ramdisk binary as a new ring-3 process
system_spawn = 17, // system_spawn(name_ptr, name_len, arguments_ptr, arguments_len, exit_endpoint) -> child process id: start a named initial-ramdisk binary as a new ring-3 process
dma_alloc = 18, // dma_alloc(len, flags) -> vaddr (rax), paddr (rdx): contiguous, pinned, uncacheable DMA memory
dma_free = 19, // dma_free(vaddr, len) -> 0: release a prior dma_alloc
msi_bind = 20, // msi_bind(device_id, endpoint) -> address (rax), data (rdx): a per-device MSI vector for a claimed device
io_read = 21, // io_read(device_id, resource_index, offset, width) -> value: read a port in a claimed device's io_port resource
io_write = 22, // io_write(device_id, resource_index, offset, width, value) -> 0: write a port in a claimed device's io_port resource
clock = 23, // clock() -> nanoseconds since boot: a monotonic time source (for timeouts/delays)
process_enumerate = 24, // process_enumerate(buffer, maximum) -> total: snapshot the task table
process_kill = 25, // process_kill(id) -> 0/-errno: end a process this process spawned
ipc_send = 26, // ipc_send(handle, message_ptr, message_len) -> 0/-errno: post a payload to an endpoint's async queue without blocking
_,
};
/// The x86 MSI message address base (`0xFEE0_0000`): a device raises an MSI by writing
/// `data` to this address, which the Local APIC turns into an interrupt at the vector
/// in `data`. The kernel returns the concrete (address, data) from `msi_bind`; this is
/// the fixed prefix, exposed so a driver's config-space programming reads clearly.
pub const msi_address_base: u64 = 0xFEE0_0000;
/// `dma_alloc` flags. `coherent` (uncacheable) is the portable default; the others are
/// opt-in for specific hardware. `write_combining` needs PAT programming (not yet — it
/// currently falls back to coherent); see docs/driver-model.md (M14).
pub const dma_coherent: u64 = 1; // strong-uncacheable — the default, the only portable one
pub const dma_write_combining: u64 = 2; // write-combining (framebuffers); needs PAT
pub const dma_below_4g: u64 = 4; // physical address must fit 32 bits (legacy DMA engines)
/// Set in the badge returned by `ipc_reply_wait` when what arrived is an
/// **asynchronous notification** (today: a device interrupt bound with `irq_bind`)
/// rather than a message from a client. There is no payload and no reply owed; the
/// low bits carry the source, a GSI. Shared so the kernel's ISR and the driver's
/// event loop can't disagree about which bit means "the hardware spoke".
/// **asynchronous notification** (a device interrupt bound with `irq_bind`, or a
/// child-exit notice — see `notify_exit_bit`) rather than a message from a client.
/// There is no payload and no reply owed; the low bits carry the source. Shared so
/// the kernel's ISR and the driver's event loop can't disagree about which bit
/// means "the hardware spoke".
pub const notify_badge_bit: u64 = 1 << 63;
/// Well-known IPC service ids for the bootstrap name registry (create_endpoint +
/// Set (alongside `notify_badge_bit`) in the badge of a **child-exit notification**:
/// posted to the endpoint a supervisor passed to `system_spawn` when that child ends
/// — by clean exit, by a fault, or by `process_kill`. The low bits carry the child's
/// process id, so one endpoint can supervise many children (and even share with IRQ
/// notifications, which never set this bit). The microkernel's SIGCHLD.
pub const notify_exit_bit: u64 = 1 << 62;
/// Set (alongside `notify_badge_bit`) in the badge of a **buffered message** — a payload
/// posted to an endpoint's async queue by `ipc_send`, delivered through `ipc_reply_wait`
/// like a notification (no reply owed) but carrying bytes in the receive buffer, not just
/// a badge. This is what distinguishes a payload-bearing async message from a bare IRQ /
/// child-exit notification (which sets neither this nor `notify_exit_bit`). The low bits
/// carry the sender's task id. The async counterpart of the synchronous `ipc_call`, for
/// broadcasts where a rendezvous is the wrong shape (the input service is the first user).
pub const notify_message_bit: u64 = 1 << 61;
/// Capacity of `ProcessDescriptor.name` — matches the longest name `system_spawn`
/// accepts, so a process's recorded name (its argv[0]) is never truncated.
pub const maximum_process_name = 64;
/// What a process is doing right now, as reported by `process_enumerate`. Crosses
/// the system_call boundary as `ProcessDescriptor.state`.
pub const ProcessState = enum(u32) {
ready = 0, // runnable, waiting for a core
running = 1, // executing on a core right now
blocked = 2, // waiting (sleeping, or blocked in IPC)
};
/// One `process_enumerate` entry — the kernel's view of a live task, kernel tasks
/// included (they carry an empty name and id 0 is the boot task). Fixed layout
/// (extern) because it crosses the kernel↔user boundary by memory copy, like
/// `DeviceDescriptor` in the device ABI.
pub const ProcessDescriptor = extern struct {
id: u32, // kernel-assigned process id; never reused (monotonic)
supervisor: u32, // id of the process that spawned it (0 = the kernel)
state: u32, // a ProcessState value
priority: u32,
name_length: u32,
name: [maximum_process_name]u8, // argv[0] at spawn; empty for kernel tasks
};
/// Well-known IPC service ids for the bootstrap name registry (create_ipc_endpoint +
/// ipc_register/ipc_lookup). Small integers, so no string interning is needed
/// during bring-up. The VFS server registers under `vfs`; clients look it up.
pub const ServiceId = enum(u32) {
vfs = 1,
input = 2,
ps2_bus = 3, // the 8042 owner; child device drivers attach here for raw bytes
_,
};
+4 -4
View File
@@ -32,10 +32,10 @@ pub const PixelFormat = enum(u32) {
/// Output Protocol (a headless server, say). The kernel must treat on-screen
/// output as optional and never assume a framebuffer exists.
pub const Framebuffer = extern struct {
base: usize, // the memory address where pixel data starts (0 = none)
width: u32, // visible pixels per row (e.g. 1920)
height: u32, // visible rows (e.g. 1080)
pitch: u32, // bytes from the start of one row to the start of the next
base: usize, // the memory address where pixel data starts (0 = none)
width: u32, // visible pixels per row (e.g. 1920)
height: u32, // visible rows (e.g. 1080)
pitch: u32, // bytes from the start of one row to the start of the next
format: PixelFormat,
/// Whether a usable framebuffer was handed over.
+138
View File
@@ -0,0 +1,138 @@
//! ACPI / PnP hardware-ID (`_HID`) names: the flat analog of pci-class.zig for
//! `acpi_device` nodes. Unlike PCI, ACPI has no class/subclass/prog-IF taxonomy — a
//! device's identity *is* its `_HID` string (`PNP0303` simply means "PS/2 keyboard"),
//! so this is a plain id <-> name registry rather than a hierarchical decoder.
//! The well-known PnP/ACPI IDs; vendor-specific ids (e.g. `QEMU0002`, `INTC1234`) have
//! no standard name and decode to nothing. Pure reference data, so it is shared by
//! kernel discovery (the device-tree dump) and any user-space driver or tool.
//!
//! Code that means a specific device names the `HardwareId` variant instead of its
//! `_HID` string — `HardwareId.ps2_keyboard.hid()` reads without a registry lookup,
//! where a bare `"PNP0303"` does not.
const std = @import("std");
/// The common standard PnP/ACPI hardware IDs, as named values. Prefix ranges hint at
/// the grouping (PNP03xx keyboards, PNP0Fxx pointing devices, PNP0Cxx ACPI
/// power/thermal, PNP0Axx buses), but there is no formal hierarchy — hence a flat
/// enum over a flat registry.
pub const HardwareId = enum {
programmable_interrupt_controller,
system_timer,
high_precision_event_timer,
dma_controller,
ps2_keyboard,
parallel_port,
ecp_parallel_port,
serial_port,
floppy_disk_controller,
system_speaker,
pci_bus,
generic_container,
/// The second id the ACPI spec assigns the same "Generic Container Device" name.
generic_container_extended,
pci_express_root_bridge,
real_time_clock,
system_board,
motherboard_reserved_resources,
math_coprocessor,
acpi_system_board,
embedded_controller,
control_method_battery,
fan,
power_button,
lid,
sleep_button,
pci_interrupt_link,
microsoft_ps2_mouse,
ps2_mouse,
ac_adapter,
processor_device,
processor_aggregator,
processor_container,
const Entry = struct { hid: []const u8, name: []const u8 };
/// The registry row for this id: its `_HID` string and human-readable name.
fn entry(self: HardwareId) Entry {
return switch (self) {
.programmable_interrupt_controller => .{ .hid = "PNP0000", .name = "Programmable Interrupt Controller (PIC)" },
.system_timer => .{ .hid = "PNP0100", .name = "System Timer (PIT)" },
.high_precision_event_timer => .{ .hid = "PNP0103", .name = "High Precision Event Timer (HPET)" },
.dma_controller => .{ .hid = "PNP0200", .name = "DMA Controller" },
.ps2_keyboard => .{ .hid = "PNP0303", .name = "PS/2 Keyboard" },
.parallel_port => .{ .hid = "PNP0400", .name = "Standard LPT Parallel Port" },
.ecp_parallel_port => .{ .hid = "PNP0401", .name = "ECP Parallel Port" },
.serial_port => .{ .hid = "PNP0501", .name = "16550A-compatible Serial Port" },
.floppy_disk_controller => .{ .hid = "PNP0700", .name = "PC Floppy Disk Controller" },
.system_speaker => .{ .hid = "PNP0800", .name = "System Speaker" },
.pci_bus => .{ .hid = "PNP0A03", .name = "PCI Bus" },
.generic_container => .{ .hid = "PNP0A05", .name = "Generic Container Device" },
.generic_container_extended => .{ .hid = "PNP0A06", .name = "Generic Container Device" },
.pci_express_root_bridge => .{ .hid = "PNP0A08", .name = "PCI Express Root Bridge" },
.real_time_clock => .{ .hid = "PNP0B00", .name = "Real-Time Clock (RTC)" },
.system_board => .{ .hid = "PNP0C01", .name = "System Board" },
.motherboard_reserved_resources => .{ .hid = "PNP0C02", .name = "Motherboard Reserved Resources" },
.math_coprocessor => .{ .hid = "PNP0C04", .name = "Math Coprocessor" },
.acpi_system_board => .{ .hid = "PNP0C08", .name = "ACPI System Board" },
.embedded_controller => .{ .hid = "PNP0C09", .name = "ACPI Embedded Controller" },
.control_method_battery => .{ .hid = "PNP0C0A", .name = "ACPI Control Method Battery" },
.fan => .{ .hid = "PNP0C0B", .name = "ACPI Fan" },
.power_button => .{ .hid = "PNP0C0C", .name = "ACPI Power Button" },
.lid => .{ .hid = "PNP0C0D", .name = "ACPI Lid" },
.sleep_button => .{ .hid = "PNP0C0E", .name = "ACPI Sleep Button" },
.pci_interrupt_link => .{ .hid = "PNP0C0F", .name = "PCI Interrupt Link Device" },
.microsoft_ps2_mouse => .{ .hid = "PNP0F03", .name = "Microsoft PS/2 Mouse" },
.ps2_mouse => .{ .hid = "PNP0F13", .name = "PS/2 Mouse" },
.ac_adapter => .{ .hid = "ACPI0003", .name = "AC Adapter" },
.processor_device => .{ .hid = "ACPI0007", .name = "Processor Device" },
.processor_aggregator => .{ .hid = "ACPI000C", .name = "Processor Aggregator" },
.processor_container => .{ .hid = "ACPI0010", .name = "Processor Container" },
};
}
/// This id's `_HID` string (e.g. `.ps2_keyboard` -> "PNP0303").
pub fn hid(self: HardwareId) []const u8 {
return self.entry().hid;
}
/// This id's human-readable name (e.g. `.ps2_keyboard` -> "PS/2 Keyboard").
pub fn description(self: HardwareId) []const u8 {
return self.entry().name;
}
/// The named value for a `_HID` string, or null if it is not a known standard
/// id (vendor-specific ids are not in the registry).
pub fn fromHid(hid_string: []const u8) ?HardwareId {
for (std.enums.values(HardwareId)) |id| {
if (std.mem.eql(u8, id.hid(), hid_string)) return id;
}
return null;
}
};
/// The human-readable name for a `_HID` string, or "" if it is not a known standard
/// id (vendor-specific ids have no registry name — callers just print the raw HID).
pub fn description(hid: []const u8) []const u8 {
return (HardwareId.fromHid(hid) orelse return "").description();
}
test "decodes standard PnP/ACPI ids and leaves the rest alone" {
const eq = std.testing.expectEqualStrings;
try eq("PS/2 Keyboard", description("PNP0303"));
try eq("PS/2 Mouse", description("PNP0F13"));
try eq("PCI Express Root Bridge", description("PNP0A08"));
try eq("Real-Time Clock (RTC)", description("PNP0B00"));
try eq("", description("QEMU0002")); // vendor-specific: no standard name
try eq("", description("")); // no HID at all
}
test "named values round-trip through their _HID strings" {
const testing = std.testing;
try testing.expectEqualStrings("PNP0303", HardwareId.ps2_keyboard.hid());
try testing.expectEqual(@as(?HardwareId, .ps2_mouse), HardwareId.fromHid("PNP0F13"));
try testing.expectEqual(@as(?HardwareId, null), HardwareId.fromHid("QEMU0002"));
for (std.enums.values(HardwareId)) |id| {
try testing.expectEqual(@as(?HardwareId, id), HardwareId.fromHid(id.hid()));
}
}
+80 -8
View File
@@ -17,6 +17,7 @@
const std = @import("std");
const boot_handoff = @import("boot-handoff");
const abi = @import("abi");
const acpi_ids = @import("acpi-ids");
const parameters = @import("parameters");
const device_model = @import("device-model.zig");
const aml = @import("aml/aml.zig");
@@ -51,7 +52,8 @@ pub const PowerInformation = struct {
reset: RegisterAccess = .{},
reset_value: u8 = 0,
reset_supported: bool = false,
/// SLP_TYP values for S5 (soft off) and S3 (suspend), from the AML sleep-state (`_Sx`) packages.
/// SLP_TYP values for S5 (soft off) and S3 (suspend), from the AML sleep-state (`_Sx`)
/// packages.
s5: ?aml.SleepType = null,
s3: ?aml.SleepType = null,
};
@@ -89,6 +91,20 @@ pub const PlatformInformation = struct {
/// ISA-IRQ-to-GSI remappings from the MADT (for future IOAPIC routing).
overrides: [16]IsoEntry = undefined,
override_count: usize = 0,
/// Whether an IOMMU (VT-d DMA-remapping unit) was found in the ACPI DMAR table.
/// When false, `device_claim` on a DMA-capable device is equivalent to granting
/// ring 0 — a device can DMA to any physical address (docs/driver-model.md M16).
/// Detection is the first step; per-device domain enforcement lands with the first
/// DMA driver.
iommu_present: bool = false,
/// MMIO base of the first DMA-remapping hardware unit (DMAR DRHD), when present.
iommu_base: u64 = 0,
/// The unit's Version register (offset 0x00) — its low byte is major.minor;
/// reading it back nonzero confirms a real, mappable VT-d unit.
iommu_version: u32 = 0,
/// The unit's Capability register (offset 0x08): supported address widths, number
/// of domains, etc. Recorded now; consumed when enforcement is built.
iommu_capabilities: u64 = 0,
};
/// Filled in by `discover`; the architecture layer reads it during bring-up.
@@ -183,7 +199,8 @@ const ExtendedSystemDescriptorPointer = extern struct {
root_system_description_table_address: u32 align(1),
/// The size of the RSDP.
length: u32 align(1),
/// A 64-bit physical address pointing to the XSDT. If the revision is at least 2, the XSDT should be used regardless of architecture, as the RSDT was deprecated.
/// A 64-bit physical address pointing to the XSDT. If the revision is at least 2, the XSDT
/// should be used regardless of architecture, as the RSDT was deprecated.
extended_system_descriptor_table_address: u64 align(1),
/// A checksum used for the entire table.
extended_checksum: u8,
@@ -233,6 +250,7 @@ const SLIT: [4]u8 = "SLIT".*;
/// System Resource Affinity Table (SRAT)
const SRAT: [4]u8 = "SRAT".*;
/// Secondary System Description Table (SSDT)
const DMAR: [4]u8 = "DMAR".*;
const SSDT: [4]u8 = "SSDT".*;
/// Serial Port Console Redirection table (SPCR) — the firmware's console UART.
const SPCR: [4]u8 = "SPCR".*;
@@ -440,6 +458,8 @@ fn handleTable(device_tree: *DeviceTree, hal: Hal, sdt_physical: u64) !void {
parseFadt(header);
} else if (std.mem.eql(u8, &sig, &SPCR)) {
parseSpcr(header);
} else if (std.mem.eql(u8, &sig, &DMAR)) {
parseDmar(hal, header);
} else if (std.mem.eql(u8, &sig, &SSDT)) {
// Secondary namespace bytecode — collect for the sleep-state (`_Sx`) scan.
addAmlBlock(sdt_physical);
@@ -564,6 +584,15 @@ fn enumeratePci(
bridge.name(), bus, device, function,
}) catch "pcidev";
const node = try device_tree.addChild(bridge, .pci_device, nm);
// Resource 0 is the function's own 4 KiB ECAM configuration space. A
// claimed PCI driver mmio_maps this to reach its command register,
// BARs, and — the point — its capability list (MSI/MSI-X, PCIe
// extended caps), without any new syscall. Physical address per the
// ECAM formula (same as pciConfigurationPtr).
const config_physical = alloc.base_address +
(@as(u64, @as(u8, @intCast(bus)) - alloc.start_bus) << 20) +
(@as(u64, device) << 15) + (@as(u64, function) << 12);
_ = node.addResource(.memory, config_physical, abi.page_size);
node.ids.pci_vendor = h.vendor_id;
node.ids.pci_device = h.device_id;
node.ids.pci_class = (@as(u24, h.class_code) << 16) |
@@ -741,6 +770,44 @@ fn parseSpcr(header: *const SystemDescriptorTableHeader) void {
platform_information.spcr_kind = fadt(u8, base, len, spcr_interface_type) orelse 0;
}
// DMAR remapping-structure layout (Intel VT-d spec §8): the DMAR-specific header is 12
// bytes (host-address-width, flags, 10 reserved), then a list of {type u16, length u16}
// structures. Type 0 is a DRHD (DMA Remapping Hardware Unit Definition), whose 64-bit
// register base sits at offset 8 within it.
const dmar_structures_offset = 48; // 36-byte ACPI header + 12-byte DMAR header
const dmar_type_drhd: u16 = 0;
const drhd_register_base_offset = 8;
/// DMAR -> detect the IOMMU. Find the first DMA-remapping hardware unit, map its
/// register block, and record its version and capabilities. This is *detection only*:
/// it tells the system an IOMMU exists (so `device_claim` on a DMA device could one day
/// be gated by a per-device translation domain), but no domains are programmed yet —
/// enforcement is built with the first DMA driver, which is what there is to protect and
/// test against. See docs/driver-model.md (M16), the honest caveat.
fn parseDmar(hal: Hal, header: *const SystemDescriptorTableHeader) void {
const base: [*]align(1) const u8 = @ptrCast(header);
const total: usize = header.length;
var off: usize = dmar_structures_offset;
while (off + 4 <= total) {
const kind = fadt(u16, base, total, off) orelse break;
const length = fadt(u16, base, total, off + 2) orelse break;
if (length < 4 or off + length > total) break; // malformed; stop rather than loop
if (kind == dmar_type_drhd) {
const register_base = fadt(u64, base, total, off + drhd_register_base_offset) orelse 0;
if (register_base != 0) {
const regs = hal.mapMmio(register_base, abi.page_size, true);
platform_information.iommu_present = true;
platform_information.iommu_base = register_base;
platform_information.iommu_version = @as(*const volatile u32, @ptrFromInt(regs + 0x00)).*;
platform_information.iommu_capabilities = @as(*const volatile u64, @ptrFromInt(regs + 0x08)).*;
return; // first unit is enough for detection; multi-unit is future
}
}
off += length;
}
}
// --- AML namespace -> generic device tree -----------------------------------
/// The PCI bus context while descending the ACPI namespace: the generic host
@@ -850,7 +917,14 @@ fn readAdr(node: *aml.Node) ?u32 {
return @truncate(readIntObj(n.value, &p) orelse return null);
}
/// Whether a namespace device is a PCI(e) host bridge (`PNP0A03` / `PNP0A08`).
/// Whether a `_HID` string names a PCI(e) host bridge.
fn isPciRootHid(hid: []const u8) bool {
const id = acpi_ids.HardwareId.fromHid(hid) orelse return false;
return id == .pci_bus or id == .pci_express_root_bridge;
}
/// Whether a namespace device is a PCI(e) host bridge. A packed EISA id is decoded
/// to its string form first, so both encodings answer through the one registry.
fn isPciRootNode(node: *aml.Node) bool {
const hid = aml.Namespace.childOf(node, seg4("_HID")) orelse return false;
if (hid.kind != .name or hid.value.len == 0) return false;
@@ -858,12 +932,10 @@ fn isPciRootNode(node: *aml.Node) bool {
0x00, 0x01, 0xFF, 0x0A, 0x0B, 0x0C, 0x0E => {
var p: usize = 0;
const n = readIntObj(hid.value, &p) orelse return false;
return n == 0x030AD041 or n == 0x080AD041; // PNP0A03 / PNP0A08
},
0x0D => {
const s = cstr(hid.value[1..]);
return std.mem.eql(u8, s, "PNP0A03") or std.mem.eql(u8, s, "PNP0A08");
var buffer: [8]u8 = undefined;
return isPciRootHid(eisaIdToStr(@truncate(n), &buffer));
},
0x0D => return isPciRootHid(cstr(hid.value[1..])),
else => return false,
}
}
+10 -5
View File
@@ -126,17 +126,22 @@ test "parses a nested namespace and finds the sleep package" {
// Scope(\_SB) packagelen=0x27
0x10, 0x27, 0x5C, 0x5F, 0x53, 0x42, 0x5F,
// Device(PCI0) packagelen=0x1F
0x5B, 0x82, 0x1F, 0x50, 0x43, 0x49, 0x30,
0x5B, 0x82, 0x1F, 0x50, 0x43,
0x49, 0x30,
// Name(_HID, 0x11)
0x08, 0x5F, 0x48, 0x49, 0x44, 0x0A, 0x11,
// Method(MTHD, flags=1) empty, packagelen=0x06
0x14, 0x06, 0x4D, 0x54, 0x48, 0x44, 0x01,
0x14, 0x06, 0x4D,
0x54, 0x48, 0x44, 0x01,
// Method(CALL, flags=0) { MTHD(Zero) }, packagelen=0x0B
0x14, 0x0B, 0x43, 0x41, 0x4C, 0x4C, 0x00, 0x4D, 0x54, 0x48, 0x44, 0x00,
0x14, 0x0B, 0x43, 0x41, 0x4C, 0x4C, 0x00, 0x4D,
0x54, 0x48, 0x44, 0x00,
// OperationRegion(DBG0, SystemIO, Word 0x0402, Byte 1)
0x5B, 0x80, 0x44, 0x42, 0x47, 0x30, 0x01, 0x0B, 0x02, 0x04, 0x0A, 0x01,
0x5B, 0x80, 0x44, 0x42, 0x47, 0x30, 0x01, 0x0B,
0x02, 0x04, 0x0A, 0x01,
// Field(DBG0, flags=1) { DBGB, 8 }, packagelen=0x0B
0x5B, 0x81, 0x0B, 0x44, 0x42, 0x47, 0x30, 0x01, 0x44, 0x42, 0x47, 0x42, 0x08,
0x5B, 0x81, 0x0B, 0x44, 0x42, 0x47, 0x30, 0x01,
0x44, 0x42, 0x47, 0x42, 0x08,
};
var arena = std.heap.ArenaAllocator.init(std.testing.allocator);
+9
View File
@@ -56,6 +56,10 @@ pub const maximum_device_resources = 8;
/// `DeviceDescriptor.parent` for a device with no parent — a root of the device tree.
pub const no_parent: u64 = ~@as(u64, 0);
/// `DeviceDescriptor.pci_class` for a device that is not a PCI function. (Zero would be
/// ambiguous: 0x000000 is a real class code, "unclassified device".)
pub const no_pci_class: u64 = ~@as(u64, 0);
/// A device, as snapshotted for user space by `device_enumerate`. A driver scans
/// these to find the hardware it owns, claims it, and maps its MMIO.
///
@@ -69,6 +73,11 @@ pub const DeviceDescriptor = extern struct {
id: u64,
parent: u64, // a device id, or `no_parent`
class: u64, // a DeviceClass value
// The PCI class/subclass/prog-IF triple packed as 0xCCSSPP when this device is a PCI
// function, or `no_pci_class` otherwise. This is how a manager tells *what* a
// `pci_device` is (an xHCI controller, an AHCI controller) — decode the triple into
// names with the pci-class module.
pci_class: u64,
hid_len: u64,
resource_count: u64,
hid: [8]u8,
+26 -5
View File
@@ -3,7 +3,7 @@
//! Discovery backends (ACPI today, device-tree later) translate their native
//! hardware description into this one shape, so the rest of the kernel walks a
//! plain `Device` tree without knowing which firmware described the machine —
//! the same discipline `root.zig`'s `MemoryKind` applies to memory and `architecture`
//! the same discipline `ps2-library.zig`'s `MemoryKind` applies to memory and `architecture`
//! applies to the CPU.
//!
//! This is deliberately minimal: enough to *describe* what was discovered (a
@@ -13,6 +13,8 @@
const std = @import("std");
const device_abi = @import("device-abi");
const pci_class = @import("pci-class");
const acpi_ids = @import("acpi-ids");
/// The hardware primitives a discovery backend needs but can't express portably.
/// The kernel injects an implementation (the architecture VMM + port I/O), so the device
@@ -189,12 +191,31 @@ fn dumpNode(device: *const Device, depth: usize, emit: *const fn ([]const u8) vo
var buffer: [200]u8 = undefined;
@memset(buffer[0..indent], ' ');
const body = if (device.hid_len != 0)
std.fmt.bufPrint(buffer[indent..], "{s} [{s}] hid={s}\n", .{ device.name(), @tagName(device.class), device.hid() }) catch return
else
std.fmt.bufPrint(buffer[indent..], "{s} [{s}]\n", .{ device.name(), @tagName(device.class) }) catch return;
const body = if (device.hid_len != 0) blk: {
// Decode the _HID to a human name when it's a known standard PnP/ACPI id.
const desc = acpi_ids.description(device.hid());
break :blk if (desc.len != 0)
std.fmt.bufPrint(buffer[indent..], "{s} [{s}] hid={s} ({s})\n", .{ device.name(), @tagName(device.class), device.hid(), desc }) catch return
else
std.fmt.bufPrint(buffer[indent..], "{s} [{s}] hid={s}\n", .{ device.name(), @tagName(device.class), device.hid() }) catch return;
} else std.fmt.bufPrint(buffer[indent..], "{s} [{s}]\n", .{ device.name(), @tagName(device.class) }) catch return;
emit(buffer[0 .. indent + body.len]);
// For a PCI function, decode its class code — the (class / subclass / prog-IF)
// triple that says what it actually is, which the coarse `DeviceClass` can't.
if (device.ids.pci_class) |packed_code| {
const cc = pci_class.ClassCode.unpack(packed_code);
var cbuf: [200]u8 = undefined;
const pad = @min(indent + 2, 42);
@memset(cbuf[0..pad], ' ');
const pif = pci_class.progIfName(cc.base, cc.subclass, cc.prog_if);
const cline = if (pif.len != 0)
std.fmt.bufPrint(cbuf[pad..], "class 0x{x:0>2} ({s}) subclass 0x{x:0>2} ({s}) progif 0x{x:0>2} ({s})\n", .{ cc.base, pci_class.className(cc.base), cc.subclass, pci_class.subclassName(cc.base, cc.subclass), cc.prog_if, pif }) catch return
else
std.fmt.bufPrint(cbuf[pad..], "class 0x{x:0>2} ({s}) subclass 0x{x:0>2} ({s}) progif 0x{x:0>2}\n", .{ cc.base, pci_class.className(cc.base), cc.subclass, pci_class.subclassName(cc.base, cc.subclass), cc.prog_if }) catch return;
emit(cbuf[0 .. pad + cline.len]);
}
for (device.resources[0..device.resource_count]) |r| {
var rbuf: [200]u8 = undefined;
const pad = @min(indent + 2, 42);
+261
View File
@@ -0,0 +1,261 @@
//! PCI class-code decoding: turn the (class, subclass, prog-IF) triple a PCI function
//! reports in its configuration header into human-readable names. Every PCI function
//! carries a 24-bit class code — base class (config byte 0x0B), subclass (0x0A), and
//! programming interface (0x09) — that says *what it is* far more precisely than
//! danos's coarse `DeviceClass`: an ISA bridge, a SATA/AHCI controller, and an xHCI USB
//! controller are all just `pci_device` by class, and only this triple tells them
//! apart. Pure reference data (from the PCI spec; see https://wiki.osdev.org/PCI) — no
//! hardware access — so it is shared by kernel discovery (the device-tree dump) and any
//! user-space tool (a future lspci, driver matching).
/// The three bytes of a PCI class code, unpacked from the `0xCCSSPP` value discovery
/// records in `Device.ids.pci_class` (CC = base class, SS = subclass, PP = prog-IF).
pub const ClassCode = struct {
base: u8, // class code (config offset 0x0B)
subclass: u8, // subclass (0x0A)
prog_if: u8, // programming interface (0x09)
pub fn unpack(packed_code: u24) ClassCode {
return .{
.base = @intCast((packed_code >> 16) & 0xFF),
.subclass = @intCast((packed_code >> 8) & 0xFF),
.prog_if = @intCast(packed_code & 0xFF),
};
}
};
/// Name of the base class (byte 0x0B), e.g. `0x06` -> "Bridge".
pub fn className(base: u8) []const u8 {
return switch (base) {
0x00 => "Unclassified",
0x01 => "Mass Storage Controller",
0x02 => "Network Controller",
0x03 => "Display Controller",
0x04 => "Multimedia Controller",
0x05 => "Memory Controller",
0x06 => "Bridge",
0x07 => "Simple Communication Controller",
0x08 => "Base System Peripheral",
0x09 => "Input Device Controller",
0x0A => "Docking Station",
0x0B => "Processor",
0x0C => "Serial Bus Controller",
0x0D => "Wireless Controller",
0x0E => "Intelligent Controller",
0x0F => "Satellite Communication Controller",
0x10 => "Encryption Controller",
0x11 => "Signal Processing Controller",
0x12 => "Processing Accelerator",
0x13 => "Non-Essential Instrumentation",
0x40 => "Co-Processor",
0xFF => "Unassigned Class (Vendor specific)",
else => "Unknown",
};
}
/// Name of the subclass within its base class, e.g. `(0x06, 0x01)` -> "ISA Bridge".
/// Subclass `0x80` is "Other" by PCI convention; anything unlisted is "Unknown".
pub fn subclassName(base: u8, subclass: u8) []const u8 {
return switch (base) {
0x01 => switch (subclass) {
0x00 => "SCSI Bus Controller",
0x01 => "IDE Controller",
0x02 => "Floppy Disk Controller",
0x03 => "IPI Bus Controller",
0x04 => "RAID Controller",
0x05 => "ATA Controller",
0x06 => "Serial ATA Controller",
0x07 => "Serial Attached SCSI Controller",
0x08 => "Non-Volatile Memory Controller",
else => defaultSubclass(subclass),
},
0x02 => switch (subclass) {
0x00 => "Ethernet Controller",
0x01 => "Token Ring Controller",
0x02 => "FDDI Controller",
0x03 => "ATM Controller",
0x04 => "ISDN Controller",
0x06 => "PICMG 2.14 Multi Computing Controller",
0x07 => "Infiniband Controller",
0x08 => "Fabric Controller",
else => defaultSubclass(subclass),
},
0x03 => switch (subclass) {
0x00 => "VGA Compatible Controller",
0x01 => "XGA Controller",
0x02 => "3D Controller (Not VGA-Compatible)",
else => defaultSubclass(subclass),
},
0x04 => switch (subclass) {
0x00 => "Multimedia Video Controller",
0x01 => "Multimedia Audio Controller",
0x02 => "Computer Telephony Device",
0x03 => "Audio Device",
else => defaultSubclass(subclass),
},
0x05 => switch (subclass) {
0x00 => "RAM Controller",
0x01 => "Flash Controller",
else => defaultSubclass(subclass),
},
0x06 => switch (subclass) {
0x00 => "Host Bridge",
0x01 => "ISA Bridge",
0x02 => "EISA Bridge",
0x03 => "MCA Bridge",
0x04 => "PCI-to-PCI Bridge",
0x05 => "PCMCIA Bridge",
0x06 => "NuBus Bridge",
0x07 => "CardBus Bridge",
0x08 => "RACEway Bridge",
0x09 => "PCI-to-PCI Bridge (Semi-Transparent)",
0x0A => "InfiniBand-to-PCI Host Bridge",
else => defaultSubclass(subclass),
},
0x07 => switch (subclass) {
0x00 => "Serial Controller",
0x01 => "Parallel Controller",
0x02 => "Multiport Serial Controller",
0x03 => "Modem",
0x04 => "IEEE 488.1/2 (GPIB) Controller",
0x05 => "Smart Card Controller",
else => defaultSubclass(subclass),
},
0x08 => switch (subclass) {
0x00 => "PIC",
0x01 => "DMA Controller",
0x02 => "Timer",
0x03 => "RTC Controller",
0x04 => "PCI Hot-Plug Controller",
0x05 => "SD Host Controller",
0x06 => "IOMMU",
else => defaultSubclass(subclass),
},
0x09 => switch (subclass) {
0x00 => "Keyboard Controller",
0x01 => "Digitizer Pen",
0x02 => "Mouse Controller",
0x03 => "Scanner Controller",
0x04 => "Gameport Controller",
else => defaultSubclass(subclass),
},
0x0C => switch (subclass) {
0x00 => "FireWire (IEEE 1394) Controller",
0x01 => "ACCESS Bus Controller",
0x02 => "SSA",
0x03 => "USB Controller",
0x04 => "Fibre Channel",
0x05 => "SMBus Controller",
0x06 => "InfiniBand Controller",
0x07 => "IPMI Interface",
0x08 => "SERCOS Interface (IEC 61491)",
0x09 => "CANbus Controller",
else => defaultSubclass(subclass),
},
0x0D => switch (subclass) {
0x00 => "iRDA Compatible Controller",
0x01 => "Consumer IR Controller",
0x10 => "RF Controller",
0x11 => "Bluetooth Controller",
0x12 => "Broadband Controller",
0x20 => "Ethernet Controller (802.1a)",
0x21 => "Ethernet Controller (802.1b)",
else => defaultSubclass(subclass),
},
else => defaultSubclass(subclass),
};
}
fn defaultSubclass(subclass: u8) []const u8 {
return if (subclass == 0x80) "Other" else "Unknown";
}
/// Name of the programming interface, for the subclasses that define standard ones
/// (IDE modes, SATA/AHCI, NVMe, PCI-bridge decode, UART generation, USB host type).
/// Returns "" when the prog-IF carries no standard meaning for this class/subclass —
/// callers just print the hex byte in that case.
pub fn progIfName(base: u8, subclass: u8, prog_if: u8) []const u8 {
return switch (base) {
0x01 => switch (subclass) {
0x06 => switch (prog_if) { // Serial ATA
0x00 => "Vendor Specific Interface",
0x01 => "AHCI 1.0",
0x02 => "Serial Storage Bus",
else => "",
},
0x08 => switch (prog_if) { // Non-Volatile Memory
0x01 => "NVMHCI",
0x02 => "NVM Express",
else => "",
},
else => "",
},
0x03 => switch (subclass) {
0x00 => switch (prog_if) { // VGA Compatible
0x00 => "VGA Controller",
0x01 => "8514-Compatible Controller",
else => "",
},
else => "",
},
0x06 => switch (subclass) {
0x04 => switch (prog_if) { // PCI-to-PCI Bridge
0x00 => "Normal Decode",
0x01 => "Subtractive Decode",
else => "",
},
else => "",
},
0x07 => switch (subclass) {
0x00 => switch (prog_if) { // Serial Controller
0x00 => "8250-Compatible (Generic XT)",
0x01 => "16450-Compatible",
0x02 => "16550-Compatible",
0x03 => "16650-Compatible",
0x04 => "16750-Compatible",
0x05 => "16850-Compatible",
0x06 => "16950-Compatible",
else => "",
},
else => "",
},
0x0C => switch (subclass) {
0x03 => switch (prog_if) { // USB Controller
0x00 => "UHCI Controller",
0x10 => "OHCI Controller",
0x20 => "EHCI (USB2) Controller",
0x30 => "XHCI (USB3) Controller",
0x80 => "Unspecified",
0xFE => "USB Device (not a host controller)",
else => "",
},
else => "",
},
else => "",
};
}
test "decodes the common class codes" {
const std = @import("std");
const eq = std.testing.expectEqualStrings;
const isa = ClassCode.unpack(0x06_01_00);
try std.testing.expectEqual(@as(u8, 0x06), isa.base);
try std.testing.expectEqual(@as(u8, 0x01), isa.subclass);
try eq("Bridge", className(isa.base));
try eq("ISA Bridge", subclassName(isa.base, isa.subclass));
const ahci = ClassCode.unpack(0x01_06_01);
try eq("Mass Storage Controller", className(ahci.base));
try eq("Serial ATA Controller", subclassName(ahci.base, ahci.subclass));
try eq("AHCI 1.0", progIfName(ahci.base, ahci.subclass, ahci.prog_if));
const xhci = ClassCode.unpack(0x0C_03_30);
try eq("USB Controller", subclassName(xhci.base, xhci.subclass));
try eq("XHCI (USB3) Controller", progIfName(xhci.base, xhci.subclass, xhci.prog_if));
// Unknowns and the "Other" convention.
try eq("Other", subclassName(0x02, 0x80));
try eq("Unknown", subclassName(0x06, 0x7E));
try eq("", progIfName(0x06, 0x00, 0x00)); // host bridge: prog-IF has no standard name
}
+857
View File
@@ -0,0 +1,857 @@
//! USB device-framework wire ABI: the set-up packets, standard requests, and standard
//! descriptors every USB device speaks over its default control pipe, as defined by chapter 9
//! of the USB 2.0 specification (see https://wiki.osdev.org/Universal_Serial_Bus). Pure data
//! definitions — no hardware access — shared by the host-controller bus drivers (which build
//! the requests) and anything that parses what devices return (device naming, driver
//! matching, configuration). The structs mirror the wire byte-for-byte: multi-byte fields are
//! little-endian and align(1), so a descriptor can be bit-cast straight out of a transfer
//! buffer at any offset, and bitmap bytes are packed structs so no caller ever needs a magic
//! mask. Class, subclass, and protocol code tables live in usb-ids.zig.
const DeviceState = enum(u8) {
// Immediately after the USB device is attached to the USB system, it is in this state.
// The USB specifications do not define the state of a USB device that is detached from
// a USB system.
attached,
// A device is in this state after it has both been attached to the bus, and the VBUS line is
// applied to the device (the host controller drives the VBUS at +5V, however this is only
// particularly important for hardware developers). In this state, the device must not respond
// to any bus transactions. The USB specification recognizes three potential scenarios with
// respect to how a device draws power:
// - Self-Powered Devices draw power from an external power source (e.g, a USB printer plugs
// into the wall as well as a USB port). Although the device may be considered
// technically "powered" even before attachment to the USB, it is still only considered
// powered after the VBUS line is applied to the device.
// - Bus-Powered Devices draw power solely from the USB up to 100mA.
// - Self- or Bus-Powered Devices may draw power from either the bus or an external power
// source, depending on the configuration. These devices may change power source at any
// time. If a device is currently self-powered and requires more than 100mA of power, but
// switches to being bus-powered, then the device must return to the Address state.
powered,
// A device in the powered state enters the default state after receiving a bus reset. In this
// state, the device is addressable at the default, reserved address of 0. At this point, the
// device is operating at the correct speed. The host is expected to allow 10 milliseconds
// before expecting the device to respond to data transfers after reset.
default,
// A device enters this state after the host assigns it an address via the default control pipe,
// which is always accessible whether the device's address has been set or not.
address,
// A device is in this state after the host examines its possible configurations and selects
// one. All endpoint's data toggle bits are initialized to zero when a device enters this state.
configured,
// When no traffic is observed on the bus for a period of 1 millisecond, a USB device enters
// this state, characterized by its low power consumption. The device's address and
// configuration settings are maintained while suspended. A device exits the suspended state as
// soon as it begins seeing bus activity again. The host is expected to allow 10 milliseconds
// before expecting the device to respond to data transfers after resume.
suspended,
};
const RequestCode = enum(u8) {
get_status = 0,
clear_feature = 1,
set_feature = 3,
set_address = 5,
get_descriptor = 6,
set_descriptor = 7,
get_configuration = 8,
set_configuration = 9,
get_interface = 10,
set_interface = 11,
sync_frame = 12,
};
// Direction of an endpoint, from the host's point of view
const EndpointDirection = enum(u1) {
out = 0,
in = 1,
};
// Identifier newtypes: distinct wire-sized types for values that identify something on the
// device rather than count something. Each is a non-exhaustive enum whose values originate
// in the descriptors below and flow, still typed, into the standard request constructors —
// so an interface number can never be passed where a configuration value is expected.
// The bus address of a device, assigned by the host with SET_ADDRESS. Addresses are 7 bits
// wide.
const DeviceAddress = enum(u7) {
// The default address every device answers at after a reset, until SET_ADDRESS
// completes
default = 0,
_,
};
// Identifies a configuration; from ConfigurationDescriptor.configuration_value.
const ConfigurationValue = enum(u8) {
// Not configured: returned by GET_CONFIGURATION while the device is in the address
// state, and passed to SET_CONFIGURATION to return a configured device to the address
// state
none = 0,
_,
};
// Identifies an interface within a configuration; from
// InterfaceDescriptor.interface_number.
const InterfaceNumber = enum(u8) { _ };
// Selects between the alternate settings of one interface; from
// InterfaceDescriptor.alternate_setting.
const AlternateSetting = enum(u8) {
// The default setting of an interface
default = 0,
_,
};
// The number of an endpoint within a device, 4 bits wide. The direction bit carried
// alongside it tells the two endpoints sharing a number apart.
const EndpointNumber = enum(u4) {
// Endpoint zero: the default control pipe every device provides
default_control = 0,
_,
};
// Index of a STRING descriptor, stored in descriptors that reference a string and passed to
// GET_DESCRIPTOR to read it.
const StringIndex = enum(u8) {
// The device has no string descriptor for this field
none = 0,
_,
};
// Characteristics of a device request (the bmRequestType field of a set-up packet). Fields are
// declared least-significant first: recipient occupies bits 4...0, kind bits 6...5, and
// direction bit 7.
const RequestType = packed struct(u8) {
// The recipient of the request (values 4...31 are reserved)
recipient: Recipient,
// The type of the request
kind: Kind,
// Data transfer direction. The value of this bit is ignored when length is zero.
direction: Direction,
const Recipient = enum(u5) {
device = 0,
interface = 1,
endpoint = 2,
other = 3,
};
const Kind = enum(u2) {
standard = 0,
class = 1,
vendor = 2,
reserved = 3,
};
const Direction = enum(u1) {
host_to_device = 0,
device_to_host = 1,
};
};
const Request = extern struct {
// Characteristics of the request
request_type: RequestType,
// Specific request
request_code: RequestCode,
// Word-sized field that may (or may not) serve as a parameter to the request, depending
// on the specific request. For GET_DESCRIPTOR and SET_DESCRIPTOR, bit-cast a
// DescriptorValue into this field.
value: u16 align(1),
// Word-sized field that may (or may not) serve as a parameter to the request, depending
// on the specific request. Typically this field holds an index or an offset value. When
// request_type specifies an endpoint or an interface as the recipient, bit-cast an
// EndpointIndex or an InterfaceIndex into this field.
index: u16 align(1),
// Number of bytes to transfer if there is a DATA stage.
// - If this field is non-zero, and request_type indicates a transfer from
// device-to-host, then the device must never return more than length bytes of data.
// However, a device may return less.
// - If this field is non-zero, and request_type indicates a transfer from
// host-to-device, then the host must send exactly length bytes of data. If the host
// sends more than length bytes, the behavior of the device is undefined.
length: u16 align(1),
// The format of the index field when request_type specifies an endpoint as the
// recipient. The host should always set the direction bit to zero (but the device
// should accept either value) when the endpoint is part of a control pipe.
const EndpointIndex = packed struct(u16) {
// Endpoint number
number: EndpointNumber,
// Reserved (reset to zero)
reserved: u3 = 0,
// Selects the OUT or the IN endpoint with the specified endpoint number
direction: EndpointDirection,
// Reserved (reset to zero)
reserved_high: u8 = 0,
};
// The format of the index field when request_type specifies an interface as the
// recipient.
const InterfaceIndex = packed struct(u16) {
// Interface number
number: u8,
// Reserved (reset to zero)
reserved: u8 = 0,
};
// The format of the value field of GET_DESCRIPTOR and SET_DESCRIPTOR requests: the
// descriptor type in the high byte, and the descriptor index in the low byte. The index
// is used to select a specific descriptor (only for CONFIGURATION and STRING
// descriptors) when several descriptors of that type are implemented by a device.
const DescriptorValue = packed struct(u16) {
// Descriptor index
index: u8 = 0,
// Descriptor type
kind: DescriptorType,
};
};
// Feature selectors, used as the value field of CLEAR_FEATURE and SET_FEATURE requests. The
// comment on each value notes the recipient the selector applies to.
const FeatureSelector = enum(u16) {
// Halts an endpoint (recipient: endpoint)
endpoint_halt = 0,
// Enables or disables the device's remote wakeup capability (recipient: device)
device_remote_wakeup = 1,
// Puts a hi-speed device into a test mode, selected by a TestMode value in the high
// byte of the index field (recipient: device)
test_mode = 2,
};
// Test mode selectors, passed in the high byte of the index field of a SET_FEATURE request
// with the test_mode feature selector. Values 06h...3Fh are reserved for standard test
// selectors and C0h...FFh for vendor-specific test modes; all other unlisted values are
// reserved.
const TestMode = enum(u8) {
test_j = 0x01,
test_k = 0x02,
test_se0_nak = 0x03,
test_packet = 0x04,
test_force_enable = 0x05,
_,
};
// The two bytes returned by a GET_STATUS request directed at a device. Fields are declared
// least-significant first.
const DeviceStatus = packed struct(u16) {
// Whether the device is currently self-powered (as opposed to bus-powered). This bit
// cannot be changed with the SET_FEATURE or CLEAR_FEATURE requests.
self_powered: bool,
// Whether the device is currently enabled to request remote wakeup. Changed with the
// SET_FEATURE and CLEAR_FEATURE requests using the device_remote_wakeup feature
// selector.
remote_wakeup: bool,
// Reserved (reset to zero)
reserved: u14,
};
// The two bytes returned by a GET_STATUS request directed at an endpoint. (A GET_STATUS
// request directed at an interface returns two bytes that are entirely reserved.)
const EndpointStatus = packed struct(u16) {
// Whether the endpoint is currently halted. Set with the SET_FEATURE request using the
// endpoint_halt feature selector, and cleared with CLEAR_FEATURE.
halted: bool,
// Reserved (reset to zero)
reserved: u15,
};
// A target for the standard requests that may be directed at the device, an interface, or
// an endpoint.
const Target = union(enum) {
device,
interface: InterfaceNumber,
endpoint: Request.EndpointIndex,
fn recipient(target: Target) RequestType.Recipient {
return switch (target) {
.device => .device,
.interface => .interface,
.endpoint => .endpoint,
};
}
fn index(target: Target) u16 {
return switch (target) {
.device => 0,
.interface => |number| @intFromEnum(number),
.endpoint => |endpoint| @bitCast(endpoint),
};
}
};
// Constructors for the standard device requests, one per RequestCode. Each returns a
// ready-to-send set-up packet with the request_type, value, index, and length fields the
// specification prescribes for that request.
// Reads the status of the given target: bit-cast the two bytes the device returns into a
// DeviceStatus or an EndpointStatus. (The two bytes returned for an interface are entirely
// reserved.)
fn getStatus(target: Target) Request {
return .{
.request_type = .{
.recipient = target.recipient(),
.kind = .standard,
.direction = .device_to_host,
},
.request_code = .get_status,
.value = 0,
.index = target.index(),
.length = 2,
};
}
// Clears or disables the given feature. A device cannot be taken out of a test mode with
// this request; test_mode is only cleared by cycling power.
fn clearFeature(feature: FeatureSelector, target: Target) Request {
return .{
.request_type = .{
.recipient = target.recipient(),
.kind = .standard,
.direction = .host_to_device,
},
.request_code = .clear_feature,
.value = @intFromEnum(feature),
.index = target.index(),
.length = 0,
};
}
// Sets or enables the given feature. For the test_mode feature selector, use setTestMode
// instead: the test selector rides in the high byte of the index field.
fn setFeature(feature: FeatureSelector, target: Target) Request {
return .{
.request_type = .{
.recipient = target.recipient(),
.kind = .standard,
.direction = .host_to_device,
},
.request_code = .set_feature,
.value = @intFromEnum(feature),
.index = target.index(),
.length = 0,
};
}
// Puts a hi-speed device into the given test mode: a SET_FEATURE request with the test_mode
// feature selector and the test selector in the high byte of the index field.
fn setTestMode(mode: TestMode) Request {
return .{
.request_type = .{
.recipient = .device,
.kind = .standard,
.direction = .host_to_device,
},
.request_code = .set_feature,
.value = @intFromEnum(FeatureSelector.test_mode),
.index = @as(u16, @intFromEnum(mode)) << 8,
.length = 0,
};
}
// Assigns the device its bus address, moving it from the default state to the address
// state. The device does not answer at the new address until the status stage of this
// request completes.
fn setAddress(address: DeviceAddress) Request {
return .{
.request_type = .{
.recipient = .device,
.kind = .standard,
.direction = .host_to_device,
},
.request_code = .set_address,
.value = @intFromEnum(address),
.index = 0,
.length = 0,
};
}
// Reads a descriptor from the device.
// - descriptor_index selects among descriptors of the same type, and is only used for
// configuration and string descriptors.
// - language_id selects the language of a string descriptor, and is zero otherwise.
// - length is the number of bytes to read; a device never returns more than length bytes,
// but may return less if the descriptor is shorter.
fn getDescriptor(kind: DescriptorType, descriptor_index: u8, language_id: u16, length: u16) Request {
return .{
.request_type = .{
.recipient = .device,
.kind = .standard,
.direction = .device_to_host,
},
.request_code = .get_descriptor,
.value = @bitCast(Request.DescriptorValue{ .index = descriptor_index, .kind = kind }),
.index = language_id,
.length = length,
};
}
// Updates an existing descriptor or adds a new one (optional; many devices do not support
// this request). The parameters mirror getDescriptor; the descriptor itself is sent in the
// DATA stage.
fn setDescriptor(kind: DescriptorType, descriptor_index: u8, language_id: u16, length: u16) Request {
return .{
.request_type = .{
.recipient = .device,
.kind = .standard,
.direction = .host_to_device,
},
.request_code = .set_descriptor,
.value = @bitCast(Request.DescriptorValue{ .index = descriptor_index, .kind = kind }),
.index = language_id,
.length = length,
};
}
// Reads the currently active configuration: @enumFromInt the byte the device returns into a
// ConfigurationValue, which is none while the device is not configured.
fn getConfiguration() Request {
return .{
.request_type = .{
.recipient = .device,
.kind = .standard,
.direction = .device_to_host,
},
.request_code = .get_configuration,
.value = 0,
.index = 0,
.length = 1,
};
}
// Selects the configuration with the given configuration_value (from
// ConfigurationDescriptor.configuration_value), moving the device from the address state to
// the configured state. Selecting none returns the device to the address state.
fn setConfiguration(configuration_value: ConfigurationValue) Request {
return .{
.request_type = .{
.recipient = .device,
.kind = .standard,
.direction = .host_to_device,
},
.request_code = .set_configuration,
.value = @intFromEnum(configuration_value),
.index = 0,
.length = 0,
};
}
// Reads the alternate setting currently selected for the given interface: @enumFromInt the
// byte the device returns into an AlternateSetting.
fn getInterface(interface: InterfaceNumber) Request {
return .{
.request_type = .{
.recipient = .interface,
.kind = .standard,
.direction = .device_to_host,
},
.request_code = .get_interface,
.value = 0,
.index = @intFromEnum(interface),
.length = 1,
};
}
// Selects an alternate setting (from InterfaceDescriptor.alternate_setting) for the given
// interface.
fn setInterface(interface: InterfaceNumber, alternate_setting: AlternateSetting) Request {
return .{
.request_type = .{
.recipient = .interface,
.kind = .standard,
.direction = .host_to_device,
},
.request_code = .set_interface,
.value = @intFromEnum(alternate_setting),
.index = @intFromEnum(interface),
.length = 0,
};
}
// Reads the two-byte number of the frame in which the given isochronous endpoint's
// repeating pattern of transfers begins.
fn syncFrame(endpoint: Request.EndpointIndex) Request {
return .{
.request_type = .{
.recipient = .endpoint,
.kind = .standard,
.direction = .device_to_host,
},
.request_code = .sync_frame,
.value = 0,
.index = @bitCast(endpoint),
.length = 2,
};
}
const DescriptorType = enum(u8) {
device = 1,
configuration = 2,
string = 3,
interface = 4,
endpoint = 5,
device_qualifier = 6,
other_speed_configuration = 7,
interface_power = 8,
_,
};
const DeviceDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// DEVICE Descriptor Type
descriptor_type: DescriptorType,
// USB Specification Release Number in Binary-Coded Decimal (i.e, 2.10 is expressed as 210h).
// Identifies the release of the USB Specification with with the device and its
// descriptors are compliant.
bcd_usb: u16 align(1),
// Class code (assigned by the USB-IF)
// - This field is reset to zero if each interface within a configuration specifies its own
// class information and the various interfaces operate independently.
// - A value of FFh in this field indicates the device class is vendor-specific.
device_class: u8,
// Subclass Code (assigned by the USB-IF)
// - The subclass code of a device is qualified by the class code of that device.
// - If device_class is reset to zero, then this field must also be reset to zero.
// - When device_class is not set to FFh, then all values for this field are reserved for
// assignment by the USB-IF.
device_subclass: u8,
// Protocol code (assigned by the USB-IF)
// - The protocol code of a device is qualified by both the class and subclass codes of
// that device.
// - A value of 00h in this field means that the device may specify class-specific
// protocols on an interface basis, though this is not a requirement.
// - If this field is set to FFh, then the device uses a vendor-specific protocol.
device_protocol: u8,
// Maximum packet size for endpoint zero (8, 16, 32, or 64 are the only valid options)
max_packet_size_0: u8,
// Vendor ID (assigned by the USB-IF)
vendor_id: u16 align(1),
// Product ID (assigned by the USB-IF)
product_id: u16 align(1),
// Device release number in binary-coded decimal
bcd_device: u16 align(1),
// Index of STRING descriptor describing manufacturer
manufacturer_index: StringIndex,
// Index of STRING descriptor describing product
product_index: StringIndex,
// Index of STRING descriptor describing the device's serial number
serial_number_index: StringIndex,
// Number of possible configurations
configuration_count: u8,
};
const DeviceQualifierDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// DEVICE_QUALIFIER Descriptor Type
descriptor_type: DescriptorType,
// USB Specification Release Number in Binary-Coded Decimal (i.e, 2.00 is expressed as 200h).
// Identifies the release of the USB Specification with with the device and its
// descriptors are compliant. This field must be at least 0200h.
bcd_usb: u16 align(1),
// Class code (assigned by the USB-IF)
device_class: u8,
// Subclass Code (assigned by the USB-IF)
device_subclass: u8,
// Protocol code (assigned by the USB-IF)
device_protocol: u8,
// Maximum packet size for endpoint zero (8, 16, 32, or 64 are the only valid options)
max_packet_size_0: u8,
// Number of possible configurations
configuration_count: u8,
// Reserved for future uses, must be zero.
reserved: u8,
};
const ConfigurationDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// CONFIGURATION Descriptor Type
descriptor_type: DescriptorType,
// The total combined length in bytes of all the descriptors returned with the request for
// this CONFIGURATION descriptor (including CONFIGURATION, INTERFACE, ENDPOINT, class- and
// vendor-specific descriptors).
total_length: u16 align(1),
// Number of interfaces supported by this configuration
interface_count: u8,
// Value which when used as an argument in the SET_CONFIGURATION request, causes the device
// to assume the configuration described by this descriptor.
configuration_value: ConfigurationValue,
// Index of STRING descriptor describing this configuration.
configuration_index: StringIndex,
// Configuration Characteristics
attributes: Attributes,
// Maximum power consumption of this device from the bus when fully operational and using
// this configuration. Expressed in units of 2mA (i.e., a value of 50 in this field
// indicates 100mA).
// - A device reports with the attributes field whether the configuration is bus- or
// self-powered, but the device status (retrieved with a GET_STATUS request) reports
// whether the device is currently self-powered.
// - If a device is disconnected from an external power source, it may not draw more
// power from the bus than specified in this field.
max_power: u8,
// Configuration characteristics. Fields are declared least-significant first.
const Attributes = packed struct(u8) {
// Reserved, reset to zero (D4...0)
reserved: u5,
// Whether Remote Wakeup is supported by this configuration (D5)
remote_wakeup: bool,
// Self-Powered (D6)
// - false: Device runs on power supplied by the bus
// - true: Device provides a local power source; if max_power is non-zero, the
// device also may use bus power.
self_powered: bool,
// Reserved, must be set to one for historical reasons (D7)
reserved_one: u1,
};
};
// This descriptor describes the configuration of a high-speed device if it were operating at
// its alternative speed. The structure of the OTHER_SPEED_CONFIGURATION is identical to that
// of the CONFIGURATION descriptor; the only difference is that the descriptor_type field
// reflects that the descriptor is an OTHER_SPEED_CONFIGURATION descriptor.
const OtherSpeedConfigurationDescriptor = ConfigurationDescriptor;
const InterfaceDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// INTERFACE Descriptor Type
descriptor_type: DescriptorType,
// Number of this interface. Zero-based value which identifies the index of this interface
// in the array of interfaces supported within a configuration.
interface_number: InterfaceNumber,
// Value used to select the alternate settings described by this INTERFACE descriptor for
// the interface with the interface_number in the previous field. This value is zero if
// this descriptor describes the default settings for a particular interface.
alternate_setting: AlternateSetting,
// Number of endpoints used by this interface, not including endpoint zero.
endpoint_count: u8,
// Class code (assigned by the USB-IF)
// - A value of zero here is reserved for future standardization.
// - If this value is FFh, the interface class is vendor-specific.
// - All other values are reserved for assignment by the USB-IF.
interface_class: u8,
// Subclass code (assigned by the USB-IF)
// - The subclass code in this field is qualified by the value of the interface_class
// field.
// - If interface_class is reset to zero, then this field must also be reset to zero.
// - If interface_class is not set to the value of FFh, then all values of this field are
// reserved for assignment by the USB-IF.
interface_subclass: u8,
// Protocol code (assigned by the USB-IF)
// - The protocol code in this field is qualified by the values of the interface_class
// and interface_subclass fields.
// - If an interface supports class-specific requests, then this field identifies the
// protocols that the device uses as defined by the specifications of the device class.
// - If this field is reset to zero, then the device does not use a class-specific
// protocol on this interface.
// - If this field is set to FFh, then the device uses a vendor-specific protocol on
// this interface.
interface_protocol: u8,
// Index of STRING descriptor describing this interface
interface_index: StringIndex,
};
const EndpointDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// ENDPOINT Descriptor Type
descriptor_type: DescriptorType,
// The address of the endpoint on the USB device described by this descriptor
endpoint_address: Address,
// The endpoint's attributes
attributes: Attributes,
// Maximum packet size that this endpoint is capable of sending or receiving. For
// isochronous endpoints, this value is used to reserve bus time; the pipe, however, may
// not always use all of the reserved bus time.
max_packet_size: MaxPacketSize align(1),
// Interval for polling a device during a data transfer, expressed in units of microframes
// for high-speed devices, and frames for low- and full-speed devices. The exact meaning of
// the value in this field depends on the endpoint type and the operating speed of the
// device:
// - Full- and High-speed isochronous endpoints, and high-speed interrupt endpoints:
// This field must be in the range from 1 to 16, and is used to calculate the period
// as 2^(interval - 1). That is, a value of 4 calculates to 2^(4 - 1) = 2^3 = 8.
// - Full- and Low-speed interrupt endpoints: This field must be in the range from
// 1 to 255.
// - High-speed bulk and control OUT endpoints: This field must be in the range from
// 0 to 255, and specifies the maximum NAK rate of the endpoint. A value of zero
// indicates that the endpoint never NAKs; other values indicate at most 1 NAK each
// interval number of microframes.
interval: u8,
// The address of an endpoint. Fields are declared least-significant first.
const Address = packed struct(u8) {
// Endpoint Number (D3...0)
number: EndpointNumber,
// Reserved, reset to zero (D6...4)
reserved: u3,
// Direction, ignored for control endpoints (D7)
direction: EndpointDirection,
};
// An endpoint's attributes. Fields are declared least-significant first.
const Attributes = packed struct(u8) {
// Transfer Type (D1...0)
transfer_type: TransferType,
// Synchronization Type; isochronous endpoints only, reserved and reset to zero for
// other endpoint types (D3...2)
synchronization: Synchronization,
// Usage Type; isochronous endpoints only, reserved and reset to zero for other
// endpoints (D5...4)
usage: Usage,
// Reserved, reset to zero (D7...6)
reserved: u2,
};
const TransferType = enum(u2) {
control = 0,
isochronous = 1,
bulk = 2,
interrupt = 3,
};
const Synchronization = enum(u2) {
none = 0,
asynchronous = 1,
adaptive = 2,
synchronous = 3,
};
const Usage = enum(u2) {
data = 0,
feedback = 1,
implicit_feedback_data = 2,
_,
};
// The maximum packet size of an endpoint. Fields are declared least-significant first.
const MaxPacketSize = packed struct(u16) {
// Maximum packet size in bytes (bits 10...0)
size: u11,
// Number of additional transaction opportunities per microframe, for high-speed
// isochronous and interrupt endpoints; reserved and reset to zero for other
// endpoints (bits 12...11)
additional_transactions: AdditionalTransactions,
// Reserved, must be reset to zero (bits 15...13)
reserved: u3,
};
const AdditionalTransactions = enum(u2) {
// None (1 transaction per microframe)
none = 0,
// 1 additional (2 transactions per microframe)
one = 1,
// 2 additional (3 transactions per microframe)
two = 2,
_,
};
};
// A STRING descriptor at index zero returns the list of LANGID codes supported by the
// device; all other indices return a Unicode string. Both forms start with this two-byte
// header, followed by the variable-length payload:
// - index 0: an array of two-byte LANGID codes (wLangID[0] through wLangID[x])
// - other indices: a Unicode string of N bytes
const StringDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// STRING Descriptor Type
descriptor_type: DescriptorType,
};
const std = @import("std");
test "wire sizes and offsets match the specification" {
const expectEqual = std.testing.expectEqual;
try expectEqual(8, @sizeOf(Request));
try expectEqual(18, @sizeOf(DeviceDescriptor));
try expectEqual(10, @sizeOf(DeviceQualifierDescriptor));
try expectEqual(9, @sizeOf(ConfigurationDescriptor));
try expectEqual(9, @sizeOf(InterfaceDescriptor));
try expectEqual(7, @sizeOf(EndpointDescriptor));
try expectEqual(2, @sizeOf(StringDescriptor));
try expectEqual(2, @offsetOf(DeviceDescriptor, "bcd_usb"));
try expectEqual(8, @offsetOf(DeviceDescriptor, "vendor_id"));
try expectEqual(17, @offsetOf(DeviceDescriptor, "configuration_count"));
try expectEqual(2, @offsetOf(ConfigurationDescriptor, "total_length"));
try expectEqual(4, @offsetOf(EndpointDescriptor, "max_packet_size"));
}
test "bitmap packings match the specification" {
const expectEqual = std.testing.expectEqual;
const expect = std.testing.expect;
// bmRequestType for GET_DESCRIPTOR: device-to-host | standard | device = 80h
const request_type = RequestType{
.recipient = .device,
.kind = .standard,
.direction = .device_to_host,
};
try expectEqual(0x80, @as(u8, @bitCast(request_type)));
// wValue for GET_DESCRIPTOR(CONFIGURATION, index 0) = 0200h
const descriptor_value = Request.DescriptorValue{ .kind = .configuration };
try expectEqual(0x0200, @as(u16, @bitCast(descriptor_value)));
// wIndex for the IN endpoint 1 = 0081h
const endpoint_index = Request.EndpointIndex{ .number = @enumFromInt(1), .direction = .in };
try expectEqual(0x0081, @as(u16, @bitCast(endpoint_index)));
// Endpoint address 81h = IN endpoint 1
const address: EndpointDescriptor.Address = @bitCast(@as(u8, 0x81));
try expectEqual(1, @intFromEnum(address.number));
try expectEqual(.in, address.direction);
// Endpoint attributes 03h = interrupt transfer
const attributes: EndpointDescriptor.Attributes = @bitCast(@as(u8, 0x03));
try expectEqual(.interrupt, attributes.transfer_type);
// wMaxPacketSize 0008h = 8 bytes, no additional transactions
const max_packet_size: EndpointDescriptor.MaxPacketSize = @bitCast(@as(u16, 0x0008));
try expectEqual(8, max_packet_size.size);
try expectEqual(.none, max_packet_size.additional_transactions);
// Configuration attributes C0h = self-powered, with the historical D7 bit set
const configuration_attributes: ConfigurationDescriptor.Attributes = @bitCast(@as(u8, 0xC0));
try expect(configuration_attributes.self_powered);
try expect(!configuration_attributes.remote_wakeup);
try expectEqual(1, configuration_attributes.reserved_one);
// GET_STATUS words: device 0001h = self-powered; endpoint 0001h = halted
const device_status: DeviceStatus = @bitCast(@as(u16, 0x0001));
try expect(device_status.self_powered and !device_status.remote_wakeup);
const endpoint_status: EndpointStatus = @bitCast(@as(u16, 0x0001));
try expect(endpoint_status.halted);
// DescriptorType is non-exhaustive: class-specific values (HID = 21h) pass through
const hid_type: DescriptorType = @enumFromInt(0x21);
try expectEqual(0x21, @intFromEnum(hid_type));
try expect(hid_type != .device);
}
fn expectRequestBytes(request: Request, expected: [8]u8) !void {
try std.testing.expectEqualSlices(u8, &expected, std.mem.asBytes(&request));
}
test "standard request constructors encode the specification's set-up packets" {
try expectRequestBytes(getStatus(.device), .{ 0x80, 0, 0, 0, 0, 0, 2, 0 });
try expectRequestBytes(getStatus(.{ .interface = @enumFromInt(3) }), .{ 0x81, 0, 0, 0, 3, 0, 2, 0 });
try expectRequestBytes(getStatus(.{ .endpoint = .{ .number = @enumFromInt(2), .direction = .in } }), .{ 0x82, 0, 0, 0, 0x82, 0, 2, 0 });
try expectRequestBytes(clearFeature(.endpoint_halt, .{ .endpoint = .{ .number = @enumFromInt(1), .direction = .out } }), .{ 0x02, 1, 0, 0, 0x01, 0, 0, 0 });
try expectRequestBytes(setFeature(.device_remote_wakeup, .device), .{ 0x00, 3, 1, 0, 0, 0, 0, 0 });
try expectRequestBytes(setTestMode(.test_packet), .{ 0x00, 3, 2, 0, 0, 0x04, 0, 0 });
try expectRequestBytes(setAddress(@enumFromInt(5)), .{ 0x00, 5, 5, 0, 0, 0, 0, 0 });
try expectRequestBytes(getDescriptor(.device, 0, 0, 18), .{ 0x80, 6, 0, 1, 0, 0, 18, 0 });
try expectRequestBytes(getDescriptor(.string, 2, 0x0409, 255), .{ 0x80, 6, 2, 3, 0x09, 0x04, 255, 0 });
try expectRequestBytes(setDescriptor(.string, 2, 0x0409, 16), .{ 0x00, 7, 2, 3, 0x09, 0x04, 16, 0 });
try expectRequestBytes(getConfiguration(), .{ 0x80, 8, 0, 0, 0, 0, 1, 0 });
try expectRequestBytes(setConfiguration(@enumFromInt(1)), .{ 0x00, 9, 1, 0, 0, 0, 0, 0 });
try expectRequestBytes(getInterface(@enumFromInt(2)), .{ 0x81, 10, 0, 0, 2, 0, 1, 0 });
try expectRequestBytes(setInterface(@enumFromInt(2), @enumFromInt(1)), .{ 0x01, 11, 1, 0, 2, 0, 0, 0 });
try expectRequestBytes(syncFrame(.{ .number = @enumFromInt(3), .direction = .in }), .{ 0x82, 12, 0, 0, 0x83, 0, 2, 0 });
}
+264
View File
@@ -0,0 +1,264 @@
//! USB class-code decoding: turn the (class, subclass, protocol) triple a USB device or
//! interface reports in its descriptors into typed values. The device descriptor carries one
//! triple for the whole device, and each interface descriptor carries its own; a class code
//! of zero at the device level defers entirely to the interfaces. Subclass and protocol
//! codes are qualified by the class code — the same value means different things under
//! different classes — so there is no single SubClass or Protocol enum: each class with
//! spec-defined codes gets its own namespace below. Pure reference data (from the USB-IF
//! defined class codes; see https://www.usb.org/defined-class-codes) — no hardware access —
//! so it is shared by kernel discovery and any user-space tool (device naming, driver
//! matching).
// Base class codes (assigned by the USB-IF). The comment on each value notes where the code
// may legally appear: in the device descriptor, in interface descriptors, or both.
const Class = enum(u8) {
// Use class information in the interface descriptors (device descriptor only). Each
// interface within a configuration specifies its own class information and the various
// interfaces operate independently.
per_interface = 0x00,
// Audio: speakers, microphones, sound cards (interface)
audio = 0x01,
// Communications and CDC control: modems, network adapters (both)
communications = 0x02,
// Human Interface Device: keyboards, mice, game controllers (interface)
hid = 0x03,
// Physical: force-feedback devices (interface)
physical = 0x05,
// Image: still-imaging cameras, scanners (interface)
image = 0x06,
// Printer (interface)
printer = 0x07,
// Mass storage: flash drives, external disks, card readers (interface)
mass_storage = 0x08,
// Hub (device descriptor only)
hub = 0x09,
// CDC-Data: the data interfaces paired with a communications control interface
// (interface)
cdc_data = 0x0A,
// Smart card readers (interface)
smart_card = 0x0B,
// Content security (interface)
content_security = 0x0D,
// Video: webcams (interface)
video = 0x0E,
// Personal healthcare devices (interface)
personal_healthcare = 0x0F,
// Audio/Video devices (interface)
audio_video = 0x10,
// Billboard: describes alternate modes a USB Type-C device supports (device descriptor
// only)
billboard = 0x11,
// USB Type-C bridge (interface)
type_c_bridge = 0x12,
// USB Bulk Display Protocol devices (interface)
bulk_display = 0x13,
// MCTP over USB protocol endpoint devices (interface)
mctp = 0x14,
// I3C devices (interface)
i3c = 0x3C,
// Diagnostic devices (both)
diagnostic = 0xDC,
// Wireless controllers: Bluetooth adapters (interface)
wireless_controller = 0xE0,
// Miscellaneous (both)
miscellaneous = 0xEF,
// Application-specific: firmware upgrade, IrDA bridges, test and measurement
// (interface)
application_specific = 0xFE,
// Vendor-specific (both)
vendor_specific = 0xFF,
_,
};
// Subclass and protocol codes qualified by Class.hub. Hubs have no subclass codes; the
// protocol distinguishes the hub's transaction-translator arrangement.
const hub = struct {
const Protocol = enum(u8) {
// Full-speed hub
full_speed = 0x00,
// Hi-speed hub with a single transaction translator
hi_speed_single_tt = 0x01,
// Hi-speed hub with multiple transaction translators
hi_speed_multi_tt = 0x02,
// SuperSpeed hub (USB 3)
super_speed = 0x03,
_,
};
};
// Subclass and protocol codes qualified by Class.hid.
const hid = struct {
const SubClass = enum(u8) {
// No subclass
none = 0x00,
// Boot interface: the device also supports the simplified boot protocol, usable by
// firmware before a full HID report-descriptor parser is available
boot = 0x01,
_,
};
// Only meaningful when the subclass is boot
const Protocol = enum(u8) {
none = 0x00,
keyboard = 0x01,
mouse = 0x02,
_,
};
};
// Subclass and protocol codes qualified by Class.mass_storage. The subclass identifies the
// command set the device understands; the protocol identifies the transport used to carry
// commands, data, and status over the bus.
const mass_storage = struct {
const SubClass = enum(u8) {
// SCSI command set not reported; de facto, treat as scsi
not_reported = 0x00,
// Reduced Block Commands: typically flash devices
rbc = 0x01,
// MMC-5 (ATAPI): CD and DVD drives
atapi = 0x02,
// QIC-157 tape drives (obsolete)
qic_157 = 0x03,
// UFI: floppy disk drives
ufi = 0x04,
// SFF-8070i (obsolete)
sff_8070i = 0x05,
// Transparent SCSI command set: the common case for flash drives and disks
scsi = 0x06,
// LSD FS: negotiated access to large storage devices
lsd_fs = 0x07,
// IEEE 1667
ieee_1667 = 0x08,
// Vendor-specific
vendor_specific = 0xFF,
_,
};
const Protocol = enum(u8) {
// Control/Bulk/Interrupt with command completion interrupt
cbi_completion_interrupt = 0x00,
// Control/Bulk/Interrupt without command completion interrupt
cbi = 0x01,
// Bulk-only transport: the common case for flash drives and disks
bulk_only = 0x50,
// USB attached SCSI
uas = 0x62,
// Vendor-specific
vendor_specific = 0xFF,
_,
};
};
// Subclass and protocol codes qualified by Class.communications (CDC). The protocol codes
// are model-specific; the useful invariant is the subclass, which selects the control model
// the interface implements.
const communications = struct {
const SubClass = enum(u8) {
// Direct line control model
direct_line = 0x01,
// Abstract control model: USB modems and serial adapters
abstract_control = 0x02,
// Telephone control model
telephone = 0x03,
// Multi-channel control model
multi_channel = 0x04,
// CAPI control model
capi = 0x05,
// Ethernet networking control model
ethernet = 0x06,
// ATM networking control model
atm = 0x07,
// Wireless handset control model
wireless_handset = 0x08,
// Device management
device_management = 0x09,
// Mobile direct line model
mobile_direct_line = 0x0A,
// OBEX
obex = 0x0B,
// Ethernet emulation model
ethernet_emulation = 0x0C,
// Network control model
network_control = 0x0D,
_,
};
};
// Subclass and protocol codes qualified by Class.wireless_controller.
const wireless_controller = struct {
const SubClass = enum(u8) {
// Radio frequency controllers
radio_frequency = 0x01,
_,
};
// Only meaningful when the subclass is radio_frequency
const Protocol = enum(u8) {
// Bluetooth programming interface
bluetooth = 0x01,
// Ultra-wideband radio control
ultra_wideband = 0x02,
// Remote NDIS
remote_ndis = 0x03,
// Bluetooth AMP controller
bluetooth_amp = 0x04,
_,
};
};
// Subclass and protocol codes qualified by Class.miscellaneous.
const miscellaneous = struct {
const SubClass = enum(u8) {
// Common class
common = 0x02,
_,
};
// Only meaningful when the subclass is common
const Protocol = enum(u8) {
// Interface association descriptor: at the device level, announces that the
// configuration groups interfaces into functions with IADs
interface_association = 0x01,
_,
};
};
// Subclass and protocol codes qualified by Class.application_specific.
const application_specific = struct {
const SubClass = enum(u8) {
// Device firmware upgrade
firmware_upgrade = 0x01,
// IrDA bridge
irda_bridge = 0x02,
// Test and measurement
test_and_measurement = 0x03,
_,
};
};
test "class codes match the USB-IF assignments" {
const std = @import("std");
const expectEqual = std.testing.expectEqual;
try expectEqual(0x03, @intFromEnum(Class.hid));
try expectEqual(0x09, @intFromEnum(Class.hub));
try expectEqual(0xFF, @intFromEnum(Class.vendor_specific));
// A typical flash drive: mass storage, transparent SCSI, bulk-only transport.
try expectEqual(0x06, @intFromEnum(mass_storage.SubClass.scsi));
try expectEqual(0x50, @intFromEnum(mass_storage.Protocol.bulk_only));
// A boot keyboard: HID, boot subclass, keyboard protocol.
try expectEqual(0x01, @intFromEnum(hid.SubClass.boot));
try expectEqual(0x01, @intFromEnum(hid.Protocol.keyboard));
// Class codes are non-exhaustive: unlisted values pass through undamaged.
const unknown: Class = @enumFromInt(0x42);
try expectEqual(0x42, @intFromEnum(unknown));
_ = hub.Protocol.hi_speed_multi_tt;
_ = communications.SubClass.abstract_control;
_ = wireless_controller.Protocol.bluetooth;
_ = miscellaneous.Protocol.interface_association;
_ = application_specific.SubClass.firmware_upgrade;
}
+1
View File
@@ -100,6 +100,7 @@ pub fn main() void {
while (n < n_children) : (n += 1) {
var child = std.mem.zeroes(device.DeviceDescriptor);
child.class = @intFromEnum(device.DeviceClass.timer);
child.pci_class = device.no_pci_class;
child.hid_len = 6;
child.hid[0..6].* = "hpet-t".*;
child.resource_count = 1;
+25 -17
View File
@@ -29,6 +29,7 @@
//! 0x108 TIMER0_COMPARATOR
const runtime = @import("runtime");
const mmio = @import("mmio");
const device = runtime.device;
const ipc = runtime.ipc;
@@ -50,8 +51,15 @@ const tn_route_mask: u64 = 0x1F << tn_route_shift;
/// Interrupts to observe before declaring victory.
const target_ticks = 5;
fn register(base: usize, off: usize) *volatile u64 {
return @ptrFromInt(base + off);
/// Read/write a 64-bit HPET register through the typed volatile MMIO layer (/lib/mmio).
/// The HPET is pure MMIO with no DMA, and on x86 its grant is strong-uncacheable (so
/// UC writes are already ordered) — no barriers are needed here; the point is the
/// typed, arch-portable access every driver should use.
inline fn rd(base: usize, off: usize) u64 {
return mmio.read(u64, base + off);
}
inline fn wr(base: usize, off: usize, value: u64) void {
mmio.write(u64, base + off, value);
}
/// A timer-class device exposing both an MMIO window and an IRQ: its id, the two
@@ -66,16 +74,16 @@ fn findHpet(buffer: []device.DeviceDescriptor) ?Found {
// Skip comparator children a bus driver may have published below the block
// (see system/drivers/bus/bus.zig) — we want the register block itself.
if (d.parent != device.no_parent) continue;
var mmio: ?u64 = null;
var mmio_index: ?u64 = null;
var irq: ?u64 = null;
for (0..d.resource_count) |j| {
switch (d.resources[j].kind) {
@intFromEnum(device.ResourceKind.memory) => mmio = mmio orelse j,
@intFromEnum(device.ResourceKind.memory) => mmio_index = mmio_index orelse j,
@intFromEnum(device.ResourceKind.irq) => irq = irq orelse j,
else => {},
}
}
if (mmio) |m| if (irq) |i| {
if (mmio_index) |m| if (irq) |i| {
return .{ .device_id = d.id, .mmio = m, .irq = i, .gsi = d.resources[i].start };
};
}
@@ -107,14 +115,14 @@ pub fn main() void {
// to raise exactly this line — the kernel will only bind the one it recorded.
const gsi = hpet.gsi;
const endpoint = ipc.createEndpoint() orelse {
_ = runtime.system.write("hpet: create_endpoint failed\n");
const endpoint = ipc.createIpcEndpoint() orelse {
_ = runtime.system.write("hpet: create_ipc_endpoint failed\n");
return;
};
// --- program the hardware ------------------------------------------------
// Counter period, so we can arm the comparator a fixed wall-clock distance out.
const femtos_per_tick = register(base, register_general_cap).* >> 32;
const femtos_per_tick = rd(base, register_general_cap) >> 32;
if (femtos_per_tick == 0) {
_ = runtime.system.write("hpet: bad HPET period\n");
return;
@@ -122,21 +130,21 @@ pub fn main() void {
const ticks_per_ms = 1_000_000_000_000 / femtos_per_tick;
// Stop the counter and take the legacy route off while we reconfigure.
register(base, register_general_configuration).* &= ~(configuration_enable | configuration_leg_rt);
wr(base, register_general_configuration, rd(base, register_general_configuration) & ~(configuration_enable | configuration_leg_rt));
// Timer 0: one-shot, level-triggered, routed to our GSI, interrupt enabled.
// One-shot (not periodic) sidesteps the HPET's Tn_value_SET accumulator quirk —
// we simply re-arm from the driver on each interrupt, which is what a tickless
// timer driver does anyway.
var t0 = register(base, register_timer0_configuration).*;
var t0 = rd(base, register_timer0_configuration);
t0 &= ~(tn_route_mask | tn_type_periodic);
t0 |= tn_int_type_level | tn_int_enb | (gsi << tn_route_shift);
register(base, register_timer0_configuration).* = t0;
wr(base, register_timer0_configuration, t0);
// Clear any stale assertion, then arm ~100 ms out and start the counter.
register(base, register_int_status).* = 1;
register(base, register_timer0_comparator).* = register(base, register_main_counter).* + ticks_per_ms * 100;
register(base, register_general_configuration).* |= configuration_enable;
wr(base, register_int_status, 1);
wr(base, register_timer0_comparator, rd(base, register_main_counter) + ticks_per_ms * 100);
wr(base, register_general_configuration, rd(base, register_general_configuration) | configuration_enable);
if (!device.irqBind(hpet.device_id, hpet.irq, endpoint)) {
_ = runtime.system.write("hpet: irq_bind failed\n");
@@ -157,17 +165,17 @@ pub fn main() void {
// Quiet the device: write 1 to timer 0's status bit. Until this lands, the
// line is still asserted and unmasking would refire immediately.
register(base, register_int_status).* = 1;
wr(base, register_int_status, 1);
count += 1;
if (count < target_ticks) {
register(base, register_timer0_comparator).* = register(base, register_main_counter).* + ticks_per_ms * 100;
wr(base, register_timer0_comparator, rd(base, register_main_counter) + ticks_per_ms * 100);
} else {
// Last one: stop the source rather than re-arming, so the line is left
// both quiet *and* unmasked by the ack below. Re-arming here would leave
// a pending interrupt that nobody is waiting for, and the ISR would mask
// the line again a moment later.
register(base, register_timer0_configuration).* &= ~tn_int_enb;
wr(base, register_timer0_configuration, rd(base, register_timer0_configuration) & ~tn_int_enb);
}
_ = runtime.system.write("hpet: irq\n");
+184
View File
@@ -0,0 +1,184 @@
//! PS/2 Keyboard Driver
//!
//! Spawned by the ps2-bus driver once the controller is initialized and the port
//! has passed its interface test and device reset. The bus driver hands us our
//! device HID as argv[1] and, optionally, a layout name (`"us"`, `"gb"`, ...) as
//! argv[2].
//!
//! The 8042's ports (0x60/0x64) and IRQ1 live on the PNP0303 node, which the
//! ps2-bus driver exclusively owns — so this driver never touches the hardware.
//! Instead it **attaches** to the bus (handing over its endpoint as a capability)
//! and receives every scancode byte as a forwarded asynchronous message. Each byte
//! feeds the set-2 decoder; a decoded key becomes input-protocol events:
//!
//! scancode byte -> HID usage keycode -> key_down / key_up
//! -> xkeyboard-config -> character -> key_press
const std = @import("std");
const runtime = @import("runtime");
const xkb = @import("xkeyboard-config");
const ps2 = @import("ps2-library.zig");
const scancode = @import("scancode.zig");
const device = runtime.device;
const ipc = runtime.ipc;
const protocol = runtime.input_protocol;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
/// Look up the ps2-bus service, retrying while the bus (which spawned us before
/// registering) is still coming up.
fn lookupBus() ?ipc.Handle {
var attempts: usize = 0;
while (attempts < 100) : (attempts += 1) {
if (ipc.lookup(.ps2_bus)) |handle| return handle;
runtime.system.sleep(50);
}
return null;
}
/// The character a pressed key produces under `modifiers`, or 0 for none. The
/// layout lookup answers for printable keys; the keys whose keysym has no Unicode
/// mapping but that every consumer still expects as a character (Enter, Tab,
/// Backspace, Escape) are given their ASCII control characters here.
fn characterFor(layout: *const xkb.Layout, usage: u8, modifiers: scancode.ModifierSnapshot) u32 {
const mapping = xkb.map(layout, usage, .{
.shift = modifiers.shift,
.caps_lock = modifiers.caps_lock,
.level3 = modifiers.right_alt,
.control = modifiers.control,
});
if (mapping.character) |character| return character;
return switch (@as(protocol.Keycode, @enumFromInt(usage))) {
.enter, .keypad_enter => '\n',
.tab => '\t',
.backspace => 0x08,
.escape => 0x1B,
else => 0,
};
}
/// The input protocol's modifier word for a snapshot.
fn modifierWord(modifiers: scancode.ModifierSnapshot) u32 {
var word: u32 = 0;
if (modifiers.shift) word |= protocol.modifier_shift;
if (modifiers.control) word |= protocol.modifier_control;
if (modifiers.alt) word |= protocol.modifier_alt;
return word;
}
pub fn main(init: runtime.process.Init) void {
const hid = init.arguments.get(1).?;
if (hid.len == 0) {
_ = runtime.system.write("system/drivers/ps2-bus/keyboard: no HID argument\n");
return;
}
writeLine("system/drivers/ps2-bus/keyboard: starting for hid {s}\n", .{hid});
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("system/drivers/ps2-bus/keyboard: out of memory\n");
return;
};
if (device.findDeviceDescriptorByHid(buffer, hid) == null) {
writeLine("system/drivers/ps2-bus/keyboard: no device for hid {s}\n", .{hid});
return;
}
// The layout is a spawn argument so a later settings source can choose it;
// absent (as today) it defaults to us.
const layout_name = init.arguments.get(2) orelse "us";
const layout = xkb.byName(layout_name) orelse xkb.us;
writeLine("system/drivers/ps2-bus/keyboard: layout {s}\n", .{layout.name});
// Attach to the bus: hand it our endpoint, and it forwards every byte the
// keyboard sends (it owns the controller; we own the decoding).
const bus = lookupBus() orelse {
_ = runtime.system.write("system/drivers/ps2-bus/keyboard: ps2-bus service unavailable\n");
return;
};
const endpoint = ipc.createIpcEndpoint() orelse {
_ = runtime.system.write("system/drivers/ps2-bus/keyboard: no endpoint\n");
return;
};
var attach = ps2.AttachRequest{ .device_type = @intFromEnum(ps2.DeviceType.keyboard) };
var attach_reply: [@sizeOf(ps2.AttachReply)]u8 = undefined;
const attached = ipc.callCap(bus, std.mem.asBytes(&attach), &attach_reply, endpoint) catch {
_ = runtime.system.write("system/drivers/ps2-bus/keyboard: attach call failed\n");
return;
};
if (attached.len < @sizeOf(ps2.AttachReply) or
std.mem.bytesToValue(ps2.AttachReply, attach_reply[0..@sizeOf(ps2.AttachReply)]).status != @intFromEnum(ps2.AttachStatus.ok))
{
_ = runtime.system.write("system/drivers/ps2-bus/keyboard: attach refused\n");
return;
}
// Broadcast keyboard events through the input service so programs can listen
// for them (docs/input.md).
var source = runtime.input.connectSource() orelse {
_ = runtime.system.write("system/drivers/ps2-bus/keyboard: input service unavailable\n");
return;
};
_ = runtime.system.write("system/drivers/ps2-bus/keyboard: ok\n");
var decoder = scancode.Decoder{};
var state = scancode.KeyboardState{};
var receive: [@sizeOf(ps2.ForwardedByte)]u8 = undefined;
while (true) {
const got = ipc.replyWait(endpoint, &.{}, &receive, null);
if (!got.isMessage() or got.len < @sizeOf(ps2.ForwardedByte)) continue;
const forwarded = std.mem.bytesToValue(ps2.ForwardedByte, receive[0..@sizeOf(ps2.ForwardedByte)]);
const key = decoder.feed(@intCast(forwarded.byte & 0xFF)) orelse continue;
const transition = state.apply(key);
const modifiers = modifierWord(transition.modifiers);
switch (transition.action) {
.pressed => {
_ = source.publishKeyboardEvent(.{
.kind = @intFromEnum(protocol.EventKind.key_down),
.keycode = key.usage,
.character = 0,
.modifiers = modifiers,
});
const character = characterFor(layout, key.usage, transition.modifiers);
if (character != 0) {
_ = source.publishKeyboardEvent(.{
.kind = @intFromEnum(protocol.EventKind.key_press),
.keycode = key.usage,
.character = character,
.modifiers = modifiers,
});
}
},
// Typematic repeat: the key did not physically go down again, so no
// key_down — but it keeps producing its character.
.repeated => {
const character = characterFor(layout, key.usage, transition.modifiers);
if (character != 0) {
_ = source.publishKeyboardEvent(.{
.kind = @intFromEnum(protocol.EventKind.key_press),
.keycode = key.usage,
.character = character,
.modifiers = modifiers,
});
}
},
.released => {
_ = source.publishKeyboardEvent(.{
.kind = @intFromEnum(protocol.EventKind.key_up),
.keycode = key.usage,
.character = 0,
.modifiers = modifiers,
});
},
}
}
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
+143
View File
@@ -0,0 +1,143 @@
//! PS/2 mouse packet assembly — the byte stream a streaming mouse sends, turned
//! into decoded movement/button reports.
//!
//! A standard PS/2 mouse in stream mode sends three-byte packets:
//!
//! byte 0: | Y ovf | X ovf | Y sign | X sign | 1 | middle | right | left |
//! byte 1: X movement (low eight bits; the sign bit lives in byte 0)
//! byte 2: Y movement (likewise)
//!
//! Movement is nine-bit two's complement, PS/2 convention: positive X right,
//! positive Y **up**. The decoded packet converts Y to the screen convention
//! (positive down), matching what every consumer of relative motion expects.
//! Bit 3 of byte 0 is always set — the resynchronization anchor: a byte at
//! packet start with bit 3 clear cannot be a packet header and is dropped.
//!
//! Everything here is pure (no imports beyond `std`, no IO), so it is
//! host-testable: the tests at the bottom run under `zig build test`.
const std = @import("std");
/// One decoded movement/button report, in screen convention (positive dy down).
pub const Packet = struct {
left: bool,
right: bool,
middle: bool,
dx: i16,
dy: i16,
};
const header_always_set: u8 = 1 << 3;
const header_left: u8 = 1 << 0;
const header_right: u8 = 1 << 1;
const header_middle: u8 = 1 << 2;
const header_x_sign: u8 = 1 << 4;
const header_y_sign: u8 = 1 << 5;
const header_x_overflow: u8 = 1 << 6;
const header_y_overflow: u8 = 1 << 7;
/// Device protocol bytes that can reach the packet stream around bring-up (the
/// acknowledge to enable-reporting, a reset's self-test result). Both have bit 3
/// set, so the header check alone cannot reject them; they are recognized only
/// at packet start, where a real header cannot be one of them in practice.
const response_acknowledge: u8 = 0xFA;
const response_self_test_passed: u8 = 0xAA;
/// Accumulates the byte stream into `Packet`s. Feed it every byte the mouse
/// sends; the third byte of each well-formed packet returns one.
pub const Assembler = struct {
bytes: [3]u8 = undefined,
count: u8 = 0,
pub fn feed(self: *Assembler, byte: u8) ?Packet {
if (self.count == 0) {
// Resynchronize: a packet must start with a plausible header.
if (byte & header_always_set == 0) return null;
if (byte == response_acknowledge or byte == response_self_test_passed) return null;
}
self.bytes[self.count] = byte;
self.count += 1;
if (self.count < 3) return null;
self.count = 0;
const header = self.bytes[0];
// An overflowed count is garbage by definition; discard the packet.
if (header & (header_x_overflow | header_y_overflow) != 0) return null;
return .{
.left = header & header_left != 0,
.right = header & header_right != 0,
.middle = header & header_middle != 0,
.dx = movement(self.bytes[1], header & header_x_sign != 0),
// PS/2 positive Y is up; screen positive Y is down.
.dy = -movement(self.bytes[2], header & header_y_sign != 0),
};
}
/// Nine-bit two's complement: the eight movement bits plus the header's sign.
fn movement(low: u8, negative: bool) i16 {
const value: i16 = low;
return if (negative) value - 256 else value;
}
};
// --- tests (host-run via `zig build test`) ------------------------------------
const testing = std.testing;
fn feedAll(assembler: *Assembler, bytes: []const u8) ?Packet {
var result: ?Packet = null;
for (bytes) |byte| {
if (assembler.feed(byte)) |packet| result = packet;
}
return result;
}
test "plain motion decodes with screen-convention y" {
var assembler = Assembler{};
const packet = feedAll(&assembler, &.{ 0x08, 5, 3 }).?;
try testing.expectEqual(@as(i16, 5), packet.dx);
try testing.expectEqual(@as(i16, -3), packet.dy); // PS/2 up 3 -> screen -3
try testing.expect(!packet.left and !packet.right and !packet.middle);
}
test "negative movement sign-extends through the header bits" {
var assembler = Assembler{};
// X sign and Y sign set: dx = 0xFB - 256 = -5, dy raw = 0xFE - 256 = -2 -> screen +2.
const packet = feedAll(&assembler, &.{ 0x08 | 0x10 | 0x20, 0xFB, 0xFE }).?;
try testing.expectEqual(@as(i16, -5), packet.dx);
try testing.expectEqual(@as(i16, 2), packet.dy);
}
test "buttons decode from the header" {
var assembler = Assembler{};
const packet = feedAll(&assembler, &.{ 0x08 | 0x01 | 0x02, 0, 0 }).?;
try testing.expect(packet.left);
try testing.expect(packet.right);
try testing.expect(!packet.middle);
}
test "a byte with bit 3 clear at packet start is dropped" {
var assembler = Assembler{};
// The stray 0x02 cannot be a header; the following packet still decodes.
try testing.expectEqual(@as(?Packet, null), assembler.feed(0x02));
const packet = feedAll(&assembler, &.{ 0x09, 1, 0 }).?;
try testing.expect(packet.left);
try testing.expectEqual(@as(i16, 1), packet.dx);
}
test "protocol bytes at packet start are dropped" {
var assembler = Assembler{};
try testing.expectEqual(@as(?Packet, null), assembler.feed(0xFA)); // enable-reporting ACK
try testing.expectEqual(@as(?Packet, null), assembler.feed(0xAA)); // self-test passed
const packet = feedAll(&assembler, &.{ 0x08, 7, 0 }).?;
try testing.expectEqual(@as(i16, 7), packet.dx);
}
test "an overflowed packet is discarded whole" {
var assembler = Assembler{};
try testing.expectEqual(@as(?Packet, null), feedAll(&assembler, &.{ 0x08 | 0x40, 0xFF, 0xFF }));
// The assembler is back at packet start.
const packet = feedAll(&assembler, &.{ 0x08, 1, 1 }).?;
try testing.expectEqual(@as(i16, 1), packet.dx);
}
+144
View File
@@ -0,0 +1,144 @@
//! PS/2 Mouse Driver
//!
//! Spawned by the ps2-bus driver once the controller is initialized and the port
//! has passed its interface test and device reset. The bus driver hands us our
//! device HID as argv[1].
//!
//! Like the keyboard, this driver never touches the hardware: the 8042's ports
//! and both port IRQs are owned by the ps2-bus driver (the auxiliary port's
//! IRQ12 lives on the PNP0F13 node, which the bus claims alongside the
//! controller). The driver **attaches** to the bus and receives every byte the
//! mouse sends as a forwarded asynchronous message. The bytes assemble into
//! three-byte packets, and each packet becomes input-protocol events:
//!
//! packet -> button transitions -> button_down / button_up
//! -> movement -> motion (dx/dy, screen convention)
const std = @import("std");
const runtime = @import("runtime");
const ps2 = @import("ps2-library.zig");
const mouse_packet = @import("mouse-packet.zig");
const device = runtime.device;
const ipc = runtime.ipc;
const protocol = runtime.input_protocol;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
/// Look up the ps2-bus service, retrying while the bus (which spawned us before
/// registering) is still coming up.
fn lookupBus() ?ipc.Handle {
var attempts: usize = 0;
while (attempts < 100) : (attempts += 1) {
if (ipc.lookup(.ps2_bus)) |handle| return handle;
runtime.system.sleep(50);
}
return null;
}
/// The protocol's pressed-button bitmask for a packet.
fn buttonMask(packet: mouse_packet.Packet) u32 {
var mask: u32 = 0;
if (packet.left) mask |= protocol.mouse_button_left;
if (packet.right) mask |= protocol.mouse_button_right;
if (packet.middle) mask |= protocol.mouse_button_middle;
return mask;
}
pub fn main(init: runtime.process.Init) void {
const hid = init.arguments.get(1).?;
if (hid.len == 0) {
_ = runtime.system.write("system/drivers/ps2-bus/mouse: no HID argument\n");
return;
}
writeLine("system/drivers/ps2-bus/mouse: starting for hid {s}\n", .{hid});
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("system/drivers/ps2-bus/mouse: out of memory\n");
return;
};
if (device.findDeviceDescriptorByHid(buffer, hid) == null) {
writeLine("system/drivers/ps2-bus/mouse: no device for hid {s}\n", .{hid});
return;
}
// Attach to the bus: hand it our endpoint, and it forwards every byte the
// mouse sends (it owns the controller; we own the decoding).
const bus = lookupBus() orelse {
_ = runtime.system.write("system/drivers/ps2-bus/mouse: ps2-bus service unavailable\n");
return;
};
const endpoint = ipc.createIpcEndpoint() orelse {
_ = runtime.system.write("system/drivers/ps2-bus/mouse: no endpoint\n");
return;
};
var attach = ps2.AttachRequest{ .device_type = @intFromEnum(ps2.DeviceType.mouse) };
var attach_reply: [@sizeOf(ps2.AttachReply)]u8 = undefined;
const attached = ipc.callCap(bus, std.mem.asBytes(&attach), &attach_reply, endpoint) catch {
_ = runtime.system.write("system/drivers/ps2-bus/mouse: attach call failed\n");
return;
};
if (attached.len < @sizeOf(ps2.AttachReply) or
std.mem.bytesToValue(ps2.AttachReply, attach_reply[0..@sizeOf(ps2.AttachReply)]).status != @intFromEnum(ps2.AttachStatus.ok))
{
_ = runtime.system.write("system/drivers/ps2-bus/mouse: attach refused\n");
return;
}
// Broadcast mouse events through the input service so programs can listen
// for them (docs/input.md).
var source = runtime.input.connectSource() orelse {
_ = runtime.system.write("system/drivers/ps2-bus/mouse: input service unavailable\n");
return;
};
_ = runtime.system.write("system/drivers/ps2-bus/mouse: ok\n");
var assembler = mouse_packet.Assembler{};
var buttons: u32 = 0;
var receive: [@sizeOf(ps2.ForwardedByte)]u8 = undefined;
while (true) {
const got = ipc.replyWait(endpoint, &.{}, &receive, null);
if (!got.isMessage() or got.len < @sizeOf(ps2.ForwardedByte)) continue;
const forwarded = std.mem.bytesToValue(ps2.ForwardedByte, receive[0..@sizeOf(ps2.ForwardedByte)]);
const packet = assembler.feed(@intCast(forwarded.byte & 0xFF)) orelse continue;
const new_buttons = buttonMask(packet);
// A button transition per changed button, carrying the new whole mask.
const changed = buttons ^ new_buttons;
for ([_]u32{ protocol.mouse_button_left, protocol.mouse_button_right, protocol.mouse_button_middle }) |button| {
if (changed & button == 0) continue;
const kind: protocol.MouseEventKind = if (new_buttons & button != 0) .button_down else .button_up;
_ = source.publishMouseEvent(.{
.kind = @intFromEnum(kind),
.button = button,
.dx = 0,
.dy = 0,
.scroll_x = 0,
.scroll_y = 0,
.buttons = new_buttons,
});
}
buttons = new_buttons;
if (packet.dx != 0 or packet.dy != 0) {
_ = source.publishMouseEvent(.{
.kind = @intFromEnum(protocol.MouseEventKind.motion),
.button = 0,
.dx = packet.dx,
.dy = packet.dy,
.scroll_x = 0,
.scroll_y = 0,
.buttons = new_buttons,
});
}
}
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
+315
View File
@@ -0,0 +1,315 @@
//! The PS/2 Controller is located on the mainboard.
//! In the early days the controller was a single chip (Intel 8042).
//! As of today it is part of the Advanced Integrated Peripheral.
//!
//! It shows up in the device discovery as:
//! KBD_ [acpi_device] hid=PNP0303 (PS/2 Keyboard)
//! - io_port 0x60 len 0x1
//! - io_port 0x64 len 0x1
//! - irq 0x1 len 0x1
//! MOU_ [acpi_device] hid=PNP0F13 (PS/2 Mouse)
//! - irq 0xc len 0x1
const std = @import("std");
const runtime = @import("runtime");
const acpi_ids = @import("acpi-ids");
const ps2 = @import("ps2-library.zig");
const device = runtime.device;
const ipc = runtime.ipc;
/// Format one whole log line and emit it in a single `debug_write`, so output
/// from the child drivers (which run concurrently) can never interleave with it.
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
/// Ask the device on `port` what it is, then spawn the matching driver from the
/// initial-ramdisk, handing it the device's HID as argv[1]. The driver is chosen
/// from what the device reports, not from the port number. Returns the identified
/// type so the forwarding loop can route that port's bytes to the driver once it
/// attaches, or null if nothing was spawned.
fn spawnIdentifiedDriver(controller: ps2.Controller, port: ps2.Port) ?ps2.DeviceType {
const device_type = controller.identifyDevice(port) orelse {
writeLine("system/drivers/ps2-bus: identify timed out on port {s}\n", .{@tagName(port)});
return null;
};
const driver_name = device_type.driverName() orelse {
writeLine("system/drivers/ps2-bus: unrecognized device on port {s}\n", .{@tagName(port)});
return null;
};
const hid = device_type.hid() orelse "";
if (runtime.system.spawnWithArguments(driver_name, &.{hid}) != null) {
writeLine("system/drivers/ps2-bus: port {s} is a {s}, spawned {s}\n", .{ @tagName(port), hid, driver_name });
return device_type;
}
writeLine("system/drivers/ps2-bus: failed to spawn {s}\n", .{driver_name});
return null;
}
/// Resource index of the controller's IRQ (IRQ1) on the PNP0303 descriptor, found
/// the way the ports are found in `Controller.init`.
fn findInterruptResourceIndex(descriptor: device.DeviceDescriptor) ?u64 {
for (0..descriptor.resource_count) |index| {
if (descriptor.resources[index].kind == @intFromEnum(device.ResourceKind.irq)) return index;
}
return null;
}
/// Forwarding endpoints of the attached child drivers, indexed by `ps2.Port`.
/// Written when a child's `AttachRequest` arrives, read on every forwarded byte.
var port_endpoints = [_]?ipc.Handle{ null, null };
/// Which device type each port identified as, so an attaching child (which knows
/// its type, not its port) can be matched to the right port's byte stream.
var port_device_types = [_]?ps2.DeviceType{ null, null };
/// Handle a child driver's `AttachRequest`: record the endpoint capability it
/// passed as the forwarding target for the port whose device matches its type.
/// Writes an `AttachReply` into `out` and returns its length.
fn handleAttach(message: []const u8, got: ipc.Received, out: []u8) usize {
const reply = struct {
fn write(buffer: []u8, status: ps2.AttachStatus) usize {
const header = ps2.AttachReply{ .status = @intFromEnum(status) };
@memcpy(buffer[0..@sizeOf(ps2.AttachReply)], std.mem.asBytes(&header));
return @sizeOf(ps2.AttachReply);
}
};
if (message.len < @sizeOf(ps2.AttachRequest)) return reply.write(out, .invalid_request);
const request = std.mem.bytesToValue(ps2.AttachRequest, message[0..@sizeOf(ps2.AttachRequest)]);
const endpoint = got.cap orelse return reply.write(out, .missing_endpoint);
for (&port_device_types, 0..) |maybe_type, port_index| {
const device_type = maybe_type orelse continue;
if (@intFromEnum(device_type) != request.device_type) continue;
port_endpoints[port_index] = endpoint;
writeLine("system/drivers/ps2-bus: {s} driver attached\n", .{@tagName(device_type)});
return reply.write(out, .ok);
}
return reply.write(out, .no_such_device);
}
pub fn main() void {
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("system/drivers/ps2-bus: out of memory\n");
return;
};
var has_two_channels = false;
var maybe_controller: ?ps2.Controller = null;
var maybe_interrupt_index: ?u64 = null;
// The 8042's IO ports (0x60/0x64) are enumerated under the keyboard ACPI node
// (PNP0303), so we init the controller from that descriptor — but which device
// is on which port is decided later by identify, not by this HID.
const maybe_controller_device_descriptor = device.findDeviceDescriptorByHid(buffer, acpi_ids.HardwareId.ps2_keyboard.hid());
if (maybe_controller_device_descriptor) |controller_device_descriptor| {
_ = runtime.system.write("system/drivers/ps2-bus: found PS/2 controller\n");
_ = runtime.system.write("system/drivers/ps2-bus: initializing controller\n");
if (!device.claim(controller_device_descriptor.id)) {
_ = runtime.system.write("system/drivers/ps2-bus: unable to claim controller \n");
return;
}
const controller = ps2.Controller.init(controller_device_descriptor) orelse {
_ = runtime.system.write("system/drivers/ps2-bus: controller is missing its IO ports\n");
return;
};
maybe_controller = controller;
maybe_interrupt_index = findInterruptResourceIndex(controller_device_descriptor);
controller.disablePort(.one);
controller.disablePort(.two);
controller.flushOutputBuffer();
const current = controller.readConfigurationByte() orelse {
_ = runtime.system.write("system/drivers/ps2-bus: controller configuration timed out\n");
return;
};
const update = current & ~(ps2.configuration_first_port_interrupt |
ps2.configuration_second_port_interrupt |
ps2.configuration_first_port_translation);
if (controller.writeConfigurationByte(update) == null) {
_ = runtime.system.write("system/drivers/ps2-bus: controller configuration timed out\n");
return;
}
if (controller.performSelfTest()) |reply| {
if (reply != ps2.response_controller_test_passed) {
_ = runtime.system.write("system/drivers/ps2-bus: perform controller self test failed\n");
return;
}
} else {
_ = runtime.system.write("system/drivers/ps2-bus: controller self test timed out\n");
return;
}
has_two_channels = controller.hasTwoChannels() orelse {
_ = runtime.system.write("system/drivers/ps2-bus: controller channels timed out\n");
return;
};
if (has_two_channels) {
_ = runtime.system.write("system/drivers/ps2-bus: has two channels\n");
// keep the bus quiet until we have tested the ports and are ready to use them
controller.disablePort(.two);
} else {
_ = runtime.system.write("system/drivers/ps2-bus: has one channel\n");
}
// interface tests: always test port 1, test port 2 only if it exists
const port_one_works = (controller.testPort(.one) orelse {
_ = runtime.system.write("system/drivers/ps2-bus: port 1 test timed out\n");
return;
}) == ps2.response_port_test_passed;
var port_two_works = false;
if (has_two_channels) {
port_two_works = (controller.testPort(.two) orelse {
_ = runtime.system.write("system/drivers/ps2-bus: port 2 test timed out\n");
return;
}) == ps2.response_port_test_passed;
}
if (!port_one_works and !port_two_works) {
_ = runtime.system.write("system/drivers/ps2-bus: no usable ports\n");
return;
}
// Enable the working ports. Their interrupts stay off until IRQ1 is bound
// below — reset and identify use polled reads, which must never race the
// interrupt-driven drain loop for bytes.
controller.enablePort(.one);
if (port_two_works) controller.enablePort(.two);
// reset each working device; a failing device is logged but does not
// abort bring-up of the other one
if (port_one_works) {
if (controller.resetDevice(.one)) |passed| {
if (!passed) _ = runtime.system.write("system/drivers/ps2-bus: port 1 device reset failed\n");
} else {
_ = runtime.system.write("system/drivers/ps2-bus: port 1 device reset timed out\n");
}
}
if (port_two_works) {
if (controller.resetDevice(.two)) |passed| {
if (!passed) _ = runtime.system.write("system/drivers/ps2-bus: port 2 device reset failed\n");
} else {
_ = runtime.system.write("system/drivers/ps2-bus: port 2 device reset timed out\n");
}
}
// Identify the device on each working port and hand it off to the driver
// that matches what it reported — a port is not assumed to be a keyboard
// or a mouse by its number.
if (port_one_works) port_device_types[@intFromEnum(ps2.Port.one)] = spawnIdentifiedDriver(controller, .one);
if (port_two_works) port_device_types[@intFromEnum(ps2.Port.two)] = spawnIdentifiedDriver(controller, .two);
} else {
_ = runtime.system.write("system/drivers/ps2-bus: no PS/2 controller found\n");
return;
}
const controller = maybe_controller.?;
const interrupt_index = maybe_interrupt_index orelse {
_ = runtime.system.write("system/drivers/ps2-bus: controller is missing its IRQ\n");
return;
};
// The endpoint the child drivers attach to and IRQ1 wakes. Registered under a
// well-known id so the children can find it, the way input subscribers find
// the input service.
const endpoint = ipc.createIpcEndpoint() orelse {
_ = runtime.system.write("system/drivers/ps2-bus: no endpoint\n");
return;
};
if (!ipc.register(.ps2_bus, endpoint)) {
_ = runtime.system.write("system/drivers/ps2-bus: register failed\n");
return;
}
// From here on, only the interrupt path reads the data port. Drop anything a
// device sent between enable-scanning and now, bind the IRQs, and only then
// let the controller raise them — an interrupt with nobody bound is lost.
controller.drainOutputBuffer();
if (!device.irqBind(controller.device_id, interrupt_index, endpoint)) {
_ = runtime.system.write("system/drivers/ps2-bus: irq_bind failed\n");
return;
}
// Port 2's interrupt (IRQ12) is enumerated on the auxiliary device's own ACPI
// node (PNP0F13), not on the controller's — so if port 2 carries a device,
// claim that node too and route its IRQ to the same endpoint. The IRQ belongs
// to the *port*, whatever device identify found on it.
var maybe_auxiliary_interrupt: ?struct { device_id: u64, interrupt_index: u64, gsi: u64 } = null;
if (port_device_types[@intFromEnum(ps2.Port.two)] != null) {
if (device.findDeviceDescriptorByHid(buffer, acpi_ids.HardwareId.ps2_mouse.hid())) |descriptor| {
if (findInterruptResourceIndex(descriptor)) |auxiliary_index| {
if (device.claim(descriptor.id) and device.irqBind(descriptor.id, auxiliary_index, endpoint)) {
maybe_auxiliary_interrupt = .{
.device_id = descriptor.id,
.interrupt_index = auxiliary_index,
.gsi = descriptor.resources[auxiliary_index].start,
};
} else {
_ = runtime.system.write("system/drivers/ps2-bus: auxiliary irq_bind failed\n");
}
}
}
}
var configuration = controller.readConfigurationByte() orelse {
_ = runtime.system.write("system/drivers/ps2-bus: controller configuration timed out\n");
return;
};
if (port_device_types[@intFromEnum(ps2.Port.one)] != null) configuration |= ps2.Port.one.interruptBit();
if (maybe_auxiliary_interrupt != null) configuration |= ps2.Port.two.interruptBit();
_ = controller.writeConfigurationByte(configuration);
_ = runtime.system.write("system/drivers/ps2-bus: ok\n");
// The forwarding loop: an IRQ1 notification drains the output buffer, routing
// each byte to the attached driver of the port it came from; a client message
// is a child driver's AttachRequest.
var reply_buffer: [@sizeOf(ps2.AttachReply)]u8 = undefined;
var reply_len: usize = 0;
var receive: [@sizeOf(ps2.AttachRequest)]u8 = undefined;
while (true) {
const got = ipc.replyWait(endpoint, reply_buffer[0..reply_len], &receive, null);
if (got.isNotification()) {
reply_len = 0;
if (got.isMessage() or got.isChildExit()) continue; // nothing sends us these
while (true) {
const current_status = ps2.status(controller.device_id, controller.status_index);
if (current_status & ps2.status_output_buffer_full == 0) break;
const byte = device.ioRead(controller.device_id, controller.data_index, 0, 1) orelse break;
const port: ps2.Port = if (current_status & ps2.status_auxiliary_output != 0) .two else .one;
if (port_endpoints[@intFromEnum(port)]) |child| {
const forwarded = ps2.ForwardedByte{ .port = @intFromEnum(port), .byte = byte };
_ = ipc.send(child, std.mem.asBytes(&forwarded));
}
// An unattached port's byte is dropped — e.g. a keystroke before
// the keyboard driver has attached.
}
// Re-arm the line that woke us: the notification badge carries the
// GSI, and IRQ1 and IRQ12 are acked through different device claims.
if (maybe_auxiliary_interrupt) |auxiliary| {
if (got.source() == auxiliary.gsi) {
_ = device.irqAck(auxiliary.device_id, auxiliary.interrupt_index);
} else {
_ = device.irqAck(controller.device_id, interrupt_index);
}
} else {
_ = device.irqAck(controller.device_id, interrupt_index);
}
continue;
}
reply_len = handleAttach(receive[0..got.len], got, &reply_buffer);
}
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
+485
View File
@@ -0,0 +1,485 @@
//! shared definitions between the different PS/2 drivers
const std = @import("std");
const runtime = @import("runtime");
const acpi_ids = @import("acpi-ids");
const device = runtime.device;
const system = runtime.system;
/// PS-2 io ports:
/// The PS/2 Controller itself uses 2 IO ports (usually, IO ports 0x60 and 0x64). Like many IO
/// ports, reads and writes may access different internal registers.
///
/// Historical note: The PC-XT PPI had used port 0x61 to reset the keyboard interrupt request
/// signal (among other unrelated functions). Port 0x61 has no keyboard related functions on AT and
/// PS/2 compatibles.
///
/// The Data Port (typically IO Port 0x60) is used for reading data that was received from a PS/2
/// device or from the PS/2 controller itself and writing data to a PS/2 device or to the PS/2
/// controller itself.
// Access type: Read/Write
pub const dataPort = 0x60;
// Access type: Read
pub const statusRegisterPort = 0x64;
// Access type: Write
pub const CommandRegisterPort = 0x64;
/// How long to poll the status register before giving up. PS/2 controller
/// responses normally arrive within a few milliseconds.
pub const default_wait_timeout_nanoseconds: u64 = 10_000_000; // 10 ms
/// A PS/2 device reset (0xFF) runs the device's self-test (BAT), whose reply can
/// take far longer than an ordinary controller response.
pub const device_reset_timeout_nanoseconds: u64 = 750_000_000; // 750 ms
/// PS/2 controller commands, written to the command register (port 0x64).
pub const cmd_read_configuration_byte: u8 = 0x20; // read controller configuration byte (internal RAM byte 0)
pub const cmd_write_configuration_byte: u8 = 0x60; // write controller configuration byte (internal RAM byte 0)
pub const cmd_disable_second_port: u8 = 0xA7; // disable second PS/2 port (dual-channel controllers only)
pub const cmd_enable_second_port: u8 = 0xA8; // enable second PS/2 port (dual-channel controllers only)
pub const cmd_test_second_port: u8 = 0xA9; // test second PS/2 port
pub const cmd_test_controller: u8 = 0xAA; // controller self-test
pub const cmd_test_first_port: u8 = 0xAB; // test first PS/2 port
pub const cmd_diagnostic_dump: u8 = 0xAC; // read all bytes of internal RAM
pub const cmd_disable_first_port: u8 = 0xAD; // disable first PS/2 port
pub const cmd_enable_first_port: u8 = 0xAE; // enable first PS/2 port
pub const cmd_read_controller_input_port: u8 = 0xC0; // read controller input port
pub const cmd_read_controller_output_port: u8 = 0xD0; // read controller output port
pub const cmd_write_controller_output_port: u8 = 0xD1; // write next data byte to the controller output port
pub const cmd_write_first_port_output: u8 = 0xD2; // write next data byte to the first port output buffer
pub const cmd_write_second_port_output: u8 = 0xD3; // write next data byte to the second port output buffer
pub const cmd_write_second_port_input: u8 = 0xD4; // write next data byte to the second port input buffer (to the mouse)
pub const cmd_pulse_system_reset: u8 = 0xFE; // pulse output line 0 low: resets the CPU
/// PS/2 status register bits (read from the status port, 0x64). Bit 4 is
/// chipset-specific and intentionally omitted.
pub const status_output_buffer_full: u8 = 1 << 0; // 1 = a byte is waiting to be read from the data port
pub const status_input_buffer_full: u8 = 1 << 1; // 1 = the controller has not yet consumed the last write
pub const status_system_flag: u8 = 1 << 2; // set once the controller passes POST
pub const status_command_or_data: u8 = 1 << 3; // 1 = last write was a command, 0 = data
/// Chipset-specific in the original spec, universal in practice on dual-channel
/// controllers: set = the waiting byte came from the second port (the mouse).
pub const status_auxiliary_output: u8 = 1 << 5;
pub const status_timeout_error: u8 = 1 << 6; // 1 = time-out error
pub const status_parity_error: u8 = 1 << 7; // 1 = parity error
/// Controller configuration byte bits (internal RAM byte 0; read/written via 0x20/0x60).
pub const configuration_first_port_interrupt: u8 = 1 << 0; // 1 = first port IRQ (IRQ1) enabled
pub const configuration_second_port_interrupt: u8 = 1 << 1; // 1 = second port IRQ (IRQ12) enabled
pub const configuration_system_flag: u8 = 1 << 2; // 1 = system passed POST
pub const configuration_first_port_clock_disabled: u8 = 1 << 4; // 1 = first port clock disabled
pub const configuration_second_port_clock_disabled: u8 = 1 << 5; // 1 = second port clock disabled
pub const configuration_first_port_translation: u8 = 1 << 6; // 1 = first port scancode translation enabled
/// Controller output port bits (read/written via 0xD0/0xD1).
pub const output_port_system_reset: u8 = 1 << 0; // WARNING: keep this 1; writing 0 can lock the machine
pub const output_port_a20_gate: u8 = 1 << 1; // A20 gate
pub const output_port_second_port_clock: u8 = 1 << 2; // dual-channel controllers only
pub const output_port_second_port_data: u8 = 1 << 3; // dual-channel controllers only
pub const output_port_first_port_output_full: u8 = 1 << 4; // output buffer full from first port (IRQ1)
pub const output_port_second_port_output_full: u8 = 1 << 5; // output buffer full from second port (IRQ12)
pub const output_port_first_port_clock: u8 = 1 << 6; // first port clock
pub const output_port_first_port_data: u8 = 1 << 7; // first port data
/// Controller self-test (0xAA) result codes.
pub const response_controller_test_passed: u8 = 0x55;
pub const response_controller_test_failed: u8 = 0xFC;
/// Port test (0xAB / 0xA9) result codes.
pub const response_port_test_passed: u8 = 0x00;
pub const response_port_test_clock_stuck_low: u8 = 0x01;
pub const response_port_test_clock_stuck_high: u8 = 0x02;
pub const response_port_test_data_stuck_low: u8 = 0x03;
pub const response_port_test_data_stuck_high: u8 = 0x04;
/// PS/2 device commands, written to the data port (0x60) to reach the attached device.
pub const device_cmd_identify: u8 = 0xF2; // identify device
pub const device_cmd_enable_scanning: u8 = 0xF4;
pub const device_cmd_disable_scanning: u8 = 0xF5;
pub const device_cmd_reset: u8 = 0xFF; // reset and run the device self-test (BAT)
/// PS/2 device response bytes, read from the data port (0x60).
pub const device_response_self_test_passed: u8 = 0xAA; // BAT succeeded after a reset
pub const device_response_echo: u8 = 0xEE;
pub const device_response_acknowledge: u8 = 0xFA; // ACK
pub const device_response_self_test_failed_1: u8 = 0xFC; // BAT failure
pub const device_response_self_test_failed_2: u8 = 0xFD; // BAT failure
pub const device_response_resend: u8 = 0xFE; // ask the host to resend the last byte
/// PS/2 device identify (0xF2) reply bytes. A keyboard returns a two-byte id
/// beginning with 0xAB; a mouse returns a single-byte id (0x00/0x03/0x04); an
/// ancient AT keyboard returns nothing at all.
pub const identify_keyboard_mf2: u8 = 0xAB; // first byte of a MF2 keyboard id (a subtype byte follows)
pub const identify_mouse_standard: u8 = 0x00;
pub const identify_mouse_scroll: u8 = 0x03; // mouse with scroll wheel
pub const identify_mouse_five_button: u8 = 0x04; // 5-button mouse
fn waitReadable(id: u64, cmd_index: u64, wait_timeout_nanoseconds: u64) bool {
const deadline = system.clock() + wait_timeout_nanoseconds;
while (system.clock() < deadline) {
if (status(id, cmd_index) & status_output_buffer_full != 0) return true; // OBF set -> data ready
}
return false;
}
fn waitWritable(id: u64, cmd_index: u64, wait_timeout_nanoseconds: u64) bool {
const deadline = system.clock() + wait_timeout_nanoseconds;
while (system.clock() < deadline) {
if (status(id, cmd_index) & status_input_buffer_full == 0) return true; // IBF clear -> ok to write
}
return false; // timed out
}
pub fn status(id: u64, cmd_index: u64) u8 {
return @intCast(device.ioRead(id, cmd_index, 0, 1) orelse 0);
}
pub fn sendCommand(id: u64, cmd_index: u64, byte: u8, timeout_nanoseconds: u64) bool {
// wait IBF clear
if (!waitWritable(id, cmd_index, timeout_nanoseconds)) return false;
return device.ioWrite(id, cmd_index, 0, 1, byte);
}
pub fn readData(id: u64, status_index: u64, data_index: u64, timeout_nanoseconds: u64) ?u8 {
// OBF lives in the status register (0x64); wait for it there, then read the data port (0x60)
if (!waitReadable(id, status_index, timeout_nanoseconds)) return null;
return @intCast(device.ioRead(id, data_index, 0, 1) orelse 0);
}
pub fn writeData(id: u64, status_index: u64, data_index: u64, byte: u8, timeout_nanoseconds: u64) bool {
// IBF lives in the status register (0x64); wait for it to clear there, then write the data port
// (0x60)
if (!waitWritable(id, status_index, timeout_nanoseconds)) return false;
return device.ioWrite(id, data_index, 0, 1, byte);
}
pub const Port = enum(u2) {
one,
two,
/// Command register byte that disables this port.
fn disableCommand(self: Port) u8 {
return switch (self) {
.one => cmd_disable_first_port,
.two => cmd_disable_second_port,
};
}
/// Command register byte that enables this port (and its clock).
fn enableCommand(self: Port) u8 {
return switch (self) {
.one => cmd_enable_first_port,
.two => cmd_enable_second_port,
};
}
/// Command register byte that runs this port's interface test.
fn testCommand(self: Port) u8 {
return switch (self) {
.one => cmd_test_first_port,
.two => cmd_test_second_port,
};
}
/// Configuration-byte bit that, when set, disables this port's clock.
pub fn clockDisabledBit(self: Port) u8 {
return switch (self) {
.one => configuration_first_port_clock_disabled,
.two => configuration_second_port_clock_disabled,
};
}
/// Configuration-byte bit that, when set, enables this port's interrupt.
pub fn interruptBit(self: Port) u8 {
return switch (self) {
.one => configuration_first_port_interrupt,
.two => configuration_second_port_interrupt,
};
}
/// Command register byte that writes the next data byte into this port's
/// output buffer (makes a byte appear as if it came from the device).
pub fn writeOutputBufferCommand(self: Port) u8 {
return switch (self) {
.one => cmd_write_first_port_output,
.two => cmd_write_second_port_output,
};
}
/// Controller command that must prefix a byte destined for this port's
/// device. Port 1 is the default target of the data port, so it needs no
/// prefix (null); port 2 requires the "write second port input" command.
pub fn deviceInputCommand(self: Port) ?u8 {
return switch (self) {
.one => null,
.two => cmd_write_second_port_input,
};
}
/// Controller output-port bit driving this port's clock line.
pub fn outputPortClockBit(self: Port) u8 {
return switch (self) {
.one => output_port_first_port_clock,
.two => output_port_second_port_clock,
};
}
/// Controller output-port bit driving this port's data line.
pub fn outputPortDataBit(self: Port) u8 {
return switch (self) {
.one => output_port_first_port_data,
.two => output_port_second_port_data,
};
}
/// Controller output-port bit set when this port's output buffer is full
/// (wired to the port's IRQ line).
pub fn outputPortBufferFullBit(self: Port) u8 {
return switch (self) {
.one => output_port_first_port_output_full,
.two => output_port_second_port_output_full,
};
}
};
/// The kind of device attached to a port, as reported by the device itself in
/// response to the identify command — not assumed from the port number. Fixed
/// `u32` values because the type also travels in an `AttachRequest`.
pub const DeviceType = enum(u32) {
keyboard = 0,
mouse = 1,
unknown = 2,
/// Initial-ramdisk name of the driver that serves this device type, or null
/// if we could not classify it.
pub fn driverName(self: DeviceType) ?[]const u8 {
return switch (self) {
.keyboard => "ps2-keyboard",
.mouse => "ps2-mouse",
.unknown => null,
};
}
/// Canonical ACPI HID for this device type, handed to the spawned driver as
/// its command-line argument, or null if we could not classify it.
pub fn hid(self: DeviceType) ?[]const u8 {
return switch (self) {
.keyboard => acpi_ids.HardwareId.ps2_keyboard.hid(),
.mouse => acpi_ids.HardwareId.ps2_mouse.hid(),
.unknown => null,
};
}
};
// --- the bus <-> child-driver forwarding protocol -----------------------------
//
// The 8042's ports and IRQ1 live on the PNP0303 node that only the ps2-bus driver
// claims, so the child device drivers (ps2-keyboard, ps2-mouse) cannot read port
// 0x60 themselves. Instead each child **attaches**: it calls the bus's well-known
// `ps2_bus` endpoint with an `AttachRequest`, handing over its own endpoint as the
// call's capability. From then on the bus forwards every byte the device sends as
// a `ForwardedByte` via the asynchronous `ipc.send` — the IRQ path in the bus can
// never block on a slow child, and the child never touches the controller.
/// A child driver registering for its device's bytes. `device_type` is a
/// `DeviceType` value; the child's receive endpoint travels as the call's
/// capability (`send_cap`).
pub const AttachRequest = extern struct {
device_type: u32,
};
/// How the bus answered an `AttachRequest` (`AttachReply.status`).
pub const AttachStatus = enum(i32) {
ok = 0,
/// The request was malformed (too short to be an `AttachRequest`).
invalid_request = -1,
/// The call carried no endpoint capability to forward to.
missing_endpoint = -2,
/// No port identified a device of the requested type.
no_such_device = -3,
};
/// Reply to an `AttachRequest`. `status` is an `AttachStatus` value.
pub const AttachReply = extern struct {
status: i32,
_padding: u32 = 0,
};
/// One raw byte read from the data port, forwarded to the attached child whose
/// port it came from (routed by the status register's auxiliary-output bit).
pub const ForwardedByte = extern struct {
/// The `Port` the byte came from, as `@intFromEnum`.
port: u32,
byte: u32,
};
/// A single PS/2 (8042) controller. Construct one with `Controller.init` and
/// drive the controller through its methods; there is only ever one 8042 per
/// machine, but holding the resolved resource indices in an instance keeps the
/// call sites free of global state.
pub const Controller = struct {
device_id: u64,
/// Resource index of the command/status port (0x64).
status_index: u64,
/// Resource index of the data port (0x60).
data_index: u64,
/// Resolve the controller's IO-port resource indices from its device
/// descriptor. Returns null if either the data or command/status port is
/// missing from the descriptor.
pub fn init(device_descriptor: device.DeviceDescriptor) ?Controller {
var data_index: ?u64 = null;
var status_index: ?u64 = null;
for (device_descriptor.resources, 0..device_descriptor.resource_count) |resource, resource_index| {
if (resource.kind != @intFromEnum(device.ResourceKind.io_port)) continue;
if (resource.start == dataPort) {
data_index = @intCast(resource_index);
} else if (resource.start == statusRegisterPort) {
status_index = @intCast(resource_index);
}
}
return .{
.device_id = device_descriptor.id,
.data_index = data_index orelse return null,
.status_index = status_index orelse return null,
};
}
pub fn disablePort(self: Controller, port: Port) void {
// port enable/disable are controller commands and go to the command register (0x64)
_ = sendCommand(self.device_id, self.status_index, port.disableCommand(), default_wait_timeout_nanoseconds);
}
pub fn enablePort(self: Controller, port: Port) void {
// enabling a port also starts its clock
_ = sendCommand(self.device_id, self.status_index, port.enableCommand(), default_wait_timeout_nanoseconds);
}
/// Run a port's interface test. Returns the controller's reply — compare it
/// to `response_port_test_passed` (0x00) — or null on timeout.
pub fn testPort(self: Controller, port: Port) ?u8 {
if (!sendCommand(self.device_id, self.status_index, port.testCommand(), default_wait_timeout_nanoseconds)) return null;
return readData(self.device_id, self.status_index, self.data_index, default_wait_timeout_nanoseconds);
}
pub fn flushOutputBuffer(self: Controller) void {
// flush any stale byte the controller buffered
_ = device.ioRead(self.device_id, self.data_index, 0, 1);
}
pub fn readConfigurationByte(self: Controller) ?u8 {
// ask the controller to place its configuration byte in the output buffer, then read it
if (!sendCommand(self.device_id, self.status_index, cmd_read_configuration_byte, default_wait_timeout_nanoseconds)) return null;
return readData(self.device_id, self.status_index, self.data_index, default_wait_timeout_nanoseconds);
}
pub fn writeConfigurationByte(self: Controller, update_byte: u8) ?u8 {
// command 0x60 makes the controller store the next data-port byte as its configuration byte
if (!sendCommand(self.device_id, self.status_index, cmd_write_configuration_byte, default_wait_timeout_nanoseconds)) return null;
if (!writeData(self.device_id, self.status_index, self.data_index, update_byte, default_wait_timeout_nanoseconds)) return null;
return update_byte;
}
/// Run the controller self-test. Returns the reply — compare it to
/// `response_controller_test_passed` (0x55) — or null on timeout.
pub fn performSelfTest(self: Controller) ?u8 {
if (!sendCommand(self.device_id, self.status_index, cmd_test_controller, default_wait_timeout_nanoseconds)) return null;
return readData(self.device_id, self.status_index, self.data_index, default_wait_timeout_nanoseconds);
}
/// Detect whether this is a dual-channel controller by temporarily enabling
/// port 2 and checking whether its clock turned on. Note: this leaves port 2
/// enabled; the caller should disable it again to keep the bus quiet until
/// device bring-up.
pub fn hasTwoChannels(self: Controller) ?bool {
self.enablePort(.two);
const configuration = self.readConfigurationByte() orelse return null;
return (configuration & Port.two.clockDisabledBit()) == 0;
}
/// Reset the device attached to `port` (device command 0xFF) and wait for
/// its power-on self-test (BAT) result. Returns true if the device both
/// acknowledged and passed, false if it reported a self-test failure, or
/// null on timeout. The BAT reply can be slow, so the response reads use
/// `device_reset_timeout_nanoseconds`.
pub fn resetDevice(self: Controller, port: Port) ?bool {
// A byte destined for port 2 must be prefixed with the "write to second
// port input buffer" controller command (0xD4); port 1 is the default.
if (port.deviceInputCommand()) |prefix| {
if (!sendCommand(self.device_id, self.status_index, prefix, default_wait_timeout_nanoseconds)) return null;
}
if (!writeData(self.device_id, self.status_index, self.data_index, device_cmd_reset, default_wait_timeout_nanoseconds)) return null;
// A successful reset yields both an ACK (0xFA) and a self-test-passed
// byte (0xAA). Their order is not guaranteed, so accept either ordering.
var saw_acknowledge = false;
var saw_self_test_passed = false;
var reads: u8 = 0;
while (reads < 2) : (reads += 1) {
const reply = readData(self.device_id, self.status_index, self.data_index, device_reset_timeout_nanoseconds) orelse return null;
switch (reply) {
device_response_acknowledge => saw_acknowledge = true,
device_response_self_test_passed => saw_self_test_passed = true,
device_response_self_test_failed_1, device_response_self_test_failed_2 => return false,
else => {},
}
}
return saw_acknowledge and saw_self_test_passed;
}
/// Send one command byte to the device on `port` (applying the port-2 prefix
/// as needed) and consume its acknowledgement. Returns true on ACK (0xFA),
/// false on any other reply, or null on timeout.
pub fn sendToDevice(self: Controller, port: Port, byte: u8) ?bool {
if (port.deviceInputCommand()) |prefix| {
if (!sendCommand(self.device_id, self.status_index, prefix, default_wait_timeout_nanoseconds)) return null;
}
if (!writeData(self.device_id, self.status_index, self.data_index, byte, default_wait_timeout_nanoseconds)) return null;
const reply = readData(self.device_id, self.status_index, self.data_index, default_wait_timeout_nanoseconds) orelse return null;
return reply == device_response_acknowledge;
}
/// Discard any bytes sitting in the output buffer (for example the device-id
/// byte a mouse emits after a reset) so they cannot be mistaken for the reply
/// to a subsequent command.
pub fn drainOutputBuffer(self: Controller) void {
var guard: u8 = 0;
while (guard < 16) : (guard += 1) {
if (status(self.device_id, self.status_index) & status_output_buffer_full == 0) return;
_ = device.ioRead(self.device_id, self.data_index, 0, 1);
}
}
/// Ask the device on `port` what it is (command 0xF2) and classify the reply.
/// Scanning is disabled around the query so a streaming device cannot inject
/// data bytes that look like the identifier. Returns the device type, or null
/// if the identify command itself timed out.
pub fn identifyDevice(self: Controller, port: Port) ?DeviceType {
// Clear any leftover bytes (e.g. a post-reset mouse id) before we start.
self.drainOutputBuffer();
// Stop the device reporting so its data can't be mistaken for the reply.
if (self.sendToDevice(port, device_cmd_disable_scanning) == null) return null;
if (self.sendToDevice(port, device_cmd_identify) == null) return null;
// After the ACK, the device sends 0, 1, or 2 identifier bytes.
const first = readData(self.device_id, self.status_index, self.data_index, default_wait_timeout_nanoseconds);
const device_type: DeviceType = if (first) |id| switch (id) {
identify_keyboard_mf2 => blk: {
// A MF2 keyboard sends a second subtype byte; consume and ignore it.
_ = readData(self.device_id, self.status_index, self.data_index, default_wait_timeout_nanoseconds);
break :blk .keyboard;
},
identify_mouse_standard, identify_mouse_scroll, identify_mouse_five_button => .mouse,
else => .unknown,
} else
// No identifier bytes at all is a legacy AT keyboard.
.keyboard;
// Resume scanning so the device works once its driver takes over.
_ = self.sendToDevice(port, device_cmd_enable_scanning);
return device_type;
}
};
+389
View File
@@ -0,0 +1,389 @@
//! PS/2 scancode set 2 → USB HID usage decoding, plus the keyboard state a driver
//! needs on top of it (pressed keys, modifier tracking, caps-lock toggle).
//!
//! Set 2 is what a keyboard sends when the 8042's legacy set-1 translation is off —
//! which is how ps2-bus.zig deliberately configures the controller. A key's **make**
//! code is one byte (two with an `E0` prefix for the "extended" keys added after the
//! original AT layout); its **break** code is the same code behind an `F0` prefix.
//! Pause alone is an eight-byte `E1` sequence with no break.
//!
//! The output vocabulary is USB HID keyboard-page usages (a=4, enter=40, ...), the
//! same numbering the input protocol's `Keycode` and the xkeyboard-config layout
//! tables use — so a decoded usage indexes a layout directly.
//!
//! Everything here is pure (no imports beyond `std`, no IO), so it is host-testable:
//! the tests at the bottom run under `zig build test`.
const std = @import("std");
// --- USB HID usages the state machine itself needs to recognize --------------
pub const usage_caps_lock: u8 = 0x39;
pub const usage_left_control: u8 = 0xE0;
pub const usage_left_shift: u8 = 0xE1;
pub const usage_left_alt: u8 = 0xE2;
pub const usage_right_control: u8 = 0xE4;
pub const usage_right_shift: u8 = 0xE5;
pub const usage_right_alt: u8 = 0xE6; // AltGr — selects XKB level 3
// --- scancode set 2 → HID usage tables ---------------------------------------
/// Single-byte (non-`E0`) make codes. Zero means "no key" — protocol bytes (ACK,
/// BAT results) and reserved codes land there and decode to nothing.
pub const set2_base: [256]u8 = blk: {
var table = [_]u8{0} ** 256;
// function row
table[0x01] = 0x42; // F9
table[0x03] = 0x3E; // F5
table[0x04] = 0x3C; // F3
table[0x05] = 0x3A; // F1
table[0x06] = 0x3B; // F2
table[0x07] = 0x45; // F12
table[0x09] = 0x43; // F10
table[0x0A] = 0x41; // F8
table[0x0B] = 0x3F; // F6
table[0x0C] = 0x3D; // F4
table[0x78] = 0x44; // F11
table[0x83] = 0x40; // F7
// letters
table[0x1C] = 0x04; // A
table[0x32] = 0x05; // B
table[0x21] = 0x06; // C
table[0x23] = 0x07; // D
table[0x24] = 0x08; // E
table[0x2B] = 0x09; // F
table[0x34] = 0x0A; // G
table[0x33] = 0x0B; // H
table[0x43] = 0x0C; // I
table[0x3B] = 0x0D; // J
table[0x42] = 0x0E; // K
table[0x4B] = 0x0F; // L
table[0x3A] = 0x10; // M
table[0x31] = 0x11; // N
table[0x44] = 0x12; // O
table[0x4D] = 0x13; // P
table[0x15] = 0x14; // Q
table[0x2D] = 0x15; // R
table[0x1B] = 0x16; // S
table[0x2C] = 0x17; // T
table[0x3C] = 0x18; // U
table[0x2A] = 0x19; // V
table[0x1D] = 0x1A; // W
table[0x22] = 0x1B; // X
table[0x35] = 0x1C; // Y
table[0x1A] = 0x1D; // Z
// digit row
table[0x16] = 0x1E; // 1
table[0x1E] = 0x1F; // 2
table[0x26] = 0x20; // 3
table[0x25] = 0x21; // 4
table[0x2E] = 0x22; // 5
table[0x36] = 0x23; // 6
table[0x3D] = 0x24; // 7
table[0x3E] = 0x25; // 8
table[0x46] = 0x26; // 9
table[0x45] = 0x27; // 0
// control and whitespace
table[0x5A] = 0x28; // Enter
table[0x76] = 0x29; // Escape
table[0x66] = 0x2A; // Backspace
table[0x0D] = 0x2B; // Tab
table[0x29] = 0x2C; // Space
// punctuation
table[0x4E] = 0x2D; // - _
table[0x55] = 0x2E; // = +
table[0x54] = 0x2F; // [ {
table[0x5B] = 0x30; // ] }
table[0x5D] = 0x31; // \ | (non-US hash on ISO boards, same position)
table[0x4C] = 0x33; // ; :
table[0x52] = 0x34; // ' "
table[0x0E] = 0x35; // ` ~
table[0x41] = 0x36; // , <
table[0x49] = 0x37; // . >
table[0x4A] = 0x38; // / ?
table[0x61] = 0x64; // non-US backslash (the extra ISO key between shift and Z)
// locks
table[0x58] = usage_caps_lock;
table[0x77] = 0x53; // Num Lock
table[0x7E] = 0x47; // Scroll Lock
// keypad
table[0x7C] = 0x55; // keypad *
table[0x7B] = 0x56; // keypad -
table[0x79] = 0x57; // keypad +
table[0x69] = 0x59; // keypad 1
table[0x72] = 0x5A; // keypad 2
table[0x7A] = 0x5B; // keypad 3
table[0x6B] = 0x5C; // keypad 4
table[0x73] = 0x5D; // keypad 5
table[0x74] = 0x5E; // keypad 6
table[0x6C] = 0x5F; // keypad 7
table[0x75] = 0x60; // keypad 8
table[0x7D] = 0x61; // keypad 9
table[0x70] = 0x62; // keypad 0
table[0x71] = 0x63; // keypad .
// modifiers
table[0x14] = usage_left_control;
table[0x12] = usage_left_shift;
table[0x11] = usage_left_alt;
table[0x59] = usage_right_shift;
break :blk table;
};
/// `E0`-prefixed make codes. `E0 12` is the "fake shift" the keyboard wraps around
/// Print Screen and navigation keys when a real shift is involved; it maps to zero
/// here, so it decodes to nothing and only the real key comes through.
pub const set2_extended: [256]u8 = blk: {
var table = [_]u8{0} ** 256;
table[0x11] = usage_right_alt;
table[0x14] = usage_right_control;
table[0x1F] = 0xE3; // left GUI
table[0x27] = 0xE7; // right GUI
table[0x2F] = 0x65; // application (menu)
table[0x7C] = 0x46; // Print Screen (arrives as E0 12 E0 7C; the E0 12 decodes to nothing)
table[0x4A] = 0x54; // keypad /
table[0x5A] = 0x58; // keypad Enter
table[0x70] = 0x49; // Insert
table[0x6C] = 0x4A; // Home
table[0x7D] = 0x4B; // Page Up
table[0x71] = 0x4C; // Delete
table[0x69] = 0x4D; // End
table[0x7A] = 0x4E; // Page Down
table[0x74] = 0x4F; // right arrow
table[0x6B] = 0x50; // left arrow
table[0x72] = 0x51; // down arrow
table[0x75] = 0x52; // up arrow
break :blk table;
};
// --- the byte-stream decoder --------------------------------------------------
/// One decoded key transition: which key (as a USB HID usage) and whether this is
/// a make (press or typematic repeat) or a break (release).
pub const DecodedKey = struct {
usage: u8,
make: bool,
};
/// Turns the raw set-2 byte stream into `DecodedKey`s. Feed it every byte the
/// keyboard sends; most bytes complete a key and return one, prefix bytes return
/// null and arm the state machine for the next byte.
pub const Decoder = struct {
const State = enum {
idle,
extended, // saw E0
break_prefix, // saw F0
extended_break, // saw E0 F0
pause_skip, // inside the 8-byte E1 Pause sequence
};
state: State = .idle,
/// Bytes still to swallow in `pause_skip`.
skip: u8 = 0,
/// The whole Pause make sequence is `E1 14 77 E1 F0 14 F0 77` — seven bytes
/// after the leading `E1`, and no break sequence ever follows.
const pause_bytes_after_e1: u8 = 7;
pub fn feed(self: *Decoder, byte: u8) ?DecodedKey {
switch (self.state) {
.idle => switch (byte) {
0xE0 => self.state = .extended,
0xF0 => self.state = .break_prefix,
0xE1 => {
self.state = .pause_skip;
self.skip = pause_bytes_after_e1;
},
// Anything else is a make code — or a protocol byte (0xFA ACK,
// 0xAA BAT-passed, 0xEE echo, ...), which the tables map to zero.
else => return decoded(set2_base[byte], true),
},
.extended => switch (byte) {
0xF0 => self.state = .extended_break,
else => {
self.state = .idle;
return decoded(set2_extended[byte], true);
},
},
.break_prefix => {
self.state = .idle;
return decoded(set2_base[byte], false);
},
.extended_break => {
self.state = .idle;
return decoded(set2_extended[byte], false);
},
.pause_skip => {
self.skip -= 1;
if (self.skip == 0) self.state = .idle;
},
}
return null;
}
fn decoded(usage: u8, make: bool) ?DecodedKey {
if (usage == 0) return null; // unmapped or a protocol byte
return .{ .usage = usage, .make = make };
}
};
// --- driver-side keyboard state -----------------------------------------------
/// What a key transition did, plus the modifier state to stamp on the resulting
/// events (snapshotted after the transition was applied).
pub const Transition = struct {
pub const Action = enum {
pressed, // physical make of a key that was up
repeated, // typematic make of a key already down — no new key_down
released, // physical break
};
action: Action,
modifiers: ModifierSnapshot,
};
/// The modifier state at one instant, in both vocabularies a driver needs: the
/// input protocol's coarse bits (shift/control/alt) and the level-selection
/// inputs xkeyboard-config takes (shift, caps_lock, AltGr as level3).
pub const ModifierSnapshot = struct {
shift: bool, // either shift held
control: bool, // either control held
alt: bool, // either alt held (including AltGr)
right_alt: bool, // AltGr specifically — the XKB level-3 selector
caps_lock: bool, // the toggle, not the key
};
/// Tracks which keys are physically down and the caps-lock toggle, and classifies
/// each decoded transition. Pure state — no IO — so repeat detection and modifier
/// snapshots are host-testable.
pub const KeyboardState = struct {
/// One bit per HID usage: set while the key is physically down.
pressed: [32]u8 = [_]u8{0} ** 32,
caps_lock: bool = false,
pub fn apply(self: *KeyboardState, key: DecodedKey) Transition {
const already_down = self.isPressed(key.usage);
if (key.make) {
if (!already_down) {
self.setPressed(key.usage, true);
if (key.usage == usage_caps_lock) self.caps_lock = !self.caps_lock;
}
return .{
.action = if (already_down) .repeated else .pressed,
.modifiers = self.snapshot(),
};
}
self.setPressed(key.usage, false);
return .{ .action = .released, .modifiers = self.snapshot() };
}
pub fn isPressed(self: *const KeyboardState, usage: u8) bool {
return self.pressed[usage / 8] & (@as(u8, 1) << @intCast(usage % 8)) != 0;
}
fn setPressed(self: *KeyboardState, usage: u8, down: bool) void {
const bit = @as(u8, 1) << @intCast(usage % 8);
if (down) {
self.pressed[usage / 8] |= bit;
} else {
self.pressed[usage / 8] &= ~bit;
}
}
fn snapshot(self: *const KeyboardState) ModifierSnapshot {
const right_alt = self.isPressed(usage_right_alt);
return .{
.shift = self.isPressed(usage_left_shift) or self.isPressed(usage_right_shift),
.control = self.isPressed(usage_left_control) or self.isPressed(usage_right_control),
.alt = self.isPressed(usage_left_alt) or right_alt,
.right_alt = right_alt,
.caps_lock = self.caps_lock,
};
}
};
// --- tests (host-run via `zig build test`) ------------------------------------
const testing = std.testing;
/// Feed `bytes` and return the single DecodedKey they should produce (fails the
/// test if they produce none or more than one).
fn feedOne(decoder: *Decoder, bytes: []const u8) !DecodedKey {
var result: ?DecodedKey = null;
for (bytes) |byte| {
if (decoder.feed(byte)) |key| {
try testing.expect(result == null);
result = key;
}
}
return result orelse error.TestExpectedResult;
}
fn feedNone(decoder: *Decoder, bytes: []const u8) !void {
for (bytes) |byte| try testing.expectEqual(@as(?DecodedKey, null), decoder.feed(byte));
}
test "base make and break: A" {
var decoder = Decoder{};
try testing.expectEqual(DecodedKey{ .usage = 0x04, .make = true }, try feedOne(&decoder, &.{0x1C}));
try testing.expectEqual(DecodedKey{ .usage = 0x04, .make = false }, try feedOne(&decoder, &.{ 0xF0, 0x1C }));
}
test "extended make and break: right arrow" {
var decoder = Decoder{};
try testing.expectEqual(DecodedKey{ .usage = 0x4F, .make = true }, try feedOne(&decoder, &.{ 0xE0, 0x74 }));
try testing.expectEqual(DecodedKey{ .usage = 0x4F, .make = false }, try feedOne(&decoder, &.{ 0xE0, 0xF0, 0x74 }));
}
test "pause: the E1 sequence is consumed silently" {
var decoder = Decoder{};
try feedNone(&decoder, &.{ 0xE1, 0x14, 0x77, 0xE1, 0xF0, 0x14, 0xF0, 0x77 });
// The decoder is back in idle: an ordinary key still decodes.
try testing.expectEqual(DecodedKey{ .usage = 0x04, .make = true }, try feedOne(&decoder, &.{0x1C}));
}
test "protocol bytes decode to nothing" {
var decoder = Decoder{};
try feedNone(&decoder, &.{ 0xFA, 0xAA, 0xEE }); // ACK, BAT-passed, echo
}
test "print screen: the fake-shift E0 12 decodes to nothing" {
var decoder = Decoder{};
try feedNone(&decoder, &.{ 0xE0, 0x12 });
try testing.expectEqual(DecodedKey{ .usage = 0x46, .make = true }, try feedOne(&decoder, &.{ 0xE0, 0x7C }));
}
test "typematic repeat is classified, not re-pressed" {
var state = KeyboardState{};
const a = DecodedKey{ .usage = 0x04, .make = true };
try testing.expectEqual(Transition.Action.pressed, state.apply(a).action);
try testing.expectEqual(Transition.Action.repeated, state.apply(a).action);
try testing.expectEqual(Transition.Action.repeated, state.apply(a).action);
try testing.expectEqual(Transition.Action.released, state.apply(.{ .usage = 0x04, .make = false }).action);
try testing.expectEqual(Transition.Action.pressed, state.apply(a).action);
}
test "shift held shows in the snapshot of other keys" {
var state = KeyboardState{};
_ = state.apply(.{ .usage = usage_left_shift, .make = true });
const transition = state.apply(.{ .usage = 0x04, .make = true });
try testing.expect(transition.modifiers.shift);
try testing.expect(!transition.modifiers.control);
_ = state.apply(.{ .usage = usage_left_shift, .make = false });
_ = state.apply(.{ .usage = 0x04, .make = false });
try testing.expect(!state.apply(.{ .usage = 0x04, .make = true }).modifiers.shift);
}
test "right alt reports both alt and the level-3 selector" {
var state = KeyboardState{};
_ = state.apply(.{ .usage = usage_right_alt, .make = true });
const transition = state.apply(.{ .usage = 0x04, .make = true });
try testing.expect(transition.modifiers.alt);
try testing.expect(transition.modifiers.right_alt);
}
test "caps lock toggles on make, not on repeat or break" {
var state = KeyboardState{};
try testing.expect(state.apply(.{ .usage = usage_caps_lock, .make = true }).modifiers.caps_lock);
try testing.expect(state.apply(.{ .usage = usage_caps_lock, .make = true }).modifiers.caps_lock); // repeat
try testing.expect(state.apply(.{ .usage = usage_caps_lock, .make = false }).modifiers.caps_lock);
try testing.expect(!state.apply(.{ .usage = usage_caps_lock, .make = true }).modifiers.caps_lock); // second press: off
}
@@ -0,0 +1,71 @@
//! /system/drivers/usb-xhci-bus — the xHCI (USB 3) host-controller bus driver.
//! The device manager spawns **one instance per controller** it discovers (a machine
//! can carry several), passing the controller's device-tree id as argv[1]; this
//! instance claims that device and no other, so multiple instances never fight over
//! hardware. This increment proves the plumbing: parse the id, claim the controller,
//! and report its MMIO window. The next increments map the registers and bring the
//! controller up (reset, rings, port scan), then enumerate the USB devices on the
//! bus with the usb-abi request builders and publish each with `device_register`.
const std = @import("std");
const runtime = @import("runtime");
const device = runtime.device;
/// Format one whole log line and emit it in a single `debug_write`, so concurrent
/// instances (one per controller) can never interleave mid-line.
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
pub fn main(init: runtime.process.Init) void {
const argument = init.arguments.get(1) orelse {
_ = runtime.system.write("usb-xhci-bus: missing controller device id (argv[1])\n");
return;
};
const controller_id = std.fmt.parseInt(u64, argument, 10) catch {
writeLine("usb-xhci-bus: malformed controller device id '{s}'\n", .{argument});
return;
};
if (!device.claim(controller_id)) {
writeLine("usb-xhci-bus: unable to claim controller device {d}\n", .{controller_id});
return;
}
// Fetch our own descriptor back for the controller's resources.
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("usb-xhci-bus: out of memory\n");
return;
};
const total = device.enumerate(buffer);
const descriptor = for (buffer[0..@min(total, buffer.len)]) |d| {
if (d.id == controller_id) break d;
} else {
writeLine("usb-xhci-bus: device {d} not in the device tree\n", .{controller_id});
return;
};
// The controller's operational registers live behind BAR0, enumerated as the
// device's first memory resource.
const register_window = for (descriptor.resources[0..@intCast(descriptor.resource_count)]) |resource| {
if (resource.kind == @intFromEnum(device.ResourceKind.memory)) break resource;
} else {
writeLine("usb-xhci-bus: controller device {d} has no MMIO window\n", .{controller_id});
return;
};
writeLine("usb-xhci-bus: claimed controller device {d} (registers at 0x{x}, {d} bytes)\n", .{
controller_id,
register_window.start,
register_window.len,
});
// Controller bring-up (map the window, reset, rings, port scan) is the next
// increment; stay resident as the bus's supervisor in the meantime.
while (true) runtime.system.sleep(1000);
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
@@ -160,6 +160,15 @@ fn waitIcrIdle() void {
while (read(register_icr_low) & icr_delivery_pending != 0) {}
}
/// Send this core a fixed interrupt at `vector` (the "self" destination shorthand). A
/// device raises an MSI by writing its (address, data) to the LAPIC; with no such
/// device on QEMU's HPET, a self-IPI is the stand-in that lets the MSI vector-routing
/// path be tested end to end. Shorthand self (bits 19:18 = 01) | assert (bit 14).
pub fn selfIpi(vector: u8) void {
write(register_icr_low, 0x4_4000 | @as(u32, vector));
waitIcrIdle();
}
/// The calibration window: we time everything against a 10 ms reference interval.
const calib_ms = 10;
+13 -2
View File
@@ -168,6 +168,12 @@ pub fn mapUserDeviceInto(root: u64, virtual: u64, physical: u64, len: u64) void
paging.mapUserDeviceInto(root, virtual, physical, len);
}
/// Map coherent DMA RAM into address space `root`: strong-uncacheable, RW+NX, but
/// reclaimed on teardown (real RAM, not MMIO). For dma_alloc.
pub fn mapUserDmaInto(root: u64, virtual: u64, physical: u64, len: u64) void {
paging.mapUserDmaInto(root, virtual, physical, len);
}
/// Map a page into the kernel address space (non-executable). For the heap, etc.
pub fn mapPage(virtual: u64, physical: u64, writable: bool) void {
paging.map(virtual, physical, writable);
@@ -418,6 +424,12 @@ pub fn irqEoi() void {
apic.eoi();
}
/// Send this core a fixed interrupt at `vector`. Stands in for a device's MSI write
/// so the MSI vector-routing path can be exercised without MSI-capable hardware.
pub fn selfIpi(vector: u8) void {
apic.selfIpi(vector);
}
/// Enable the Local APIC, calibrate its timer against the best available reference
/// (see apic.calibrate — no longer the PIT by default), and start it firing at
/// `timer_hz` — the kernel's real-time heartbeat. Interrupts still have to be
@@ -478,8 +490,7 @@ pub fn saveInterrupts() u64 {
\\cli
: [f] "=r" (flags),
:
: .{ .memory = true }
);
: .{ .memory = true });
return flags;
}
+24 -21
View File
@@ -3,9 +3,11 @@
//! triple-faults and silently resets the machine. With it, the CPU vectors into
//! our stubs, which capture the register state and hand it to a dispatcher.
//!
//! Vectors split in two: 0-31 are CPU exceptions (terminal — reported and
//! halted); 32+ are device interrupts (a registered handler runs, the APIC is
//! acknowledged, and we return to the interrupted code).
//! Vectors split in two: 0-31 are CPU exceptions, handed to the `on_fault` hook
//! and never returned from (the kernel's handler kills a faulting user process
//! and reschedules, or halts the core for a kernel-mode fault); 32+ are device
//! interrupts (a registered handler runs, the APIC is acknowledged, and we
//! return to the interrupted code).
const gdt = @import("gdt.zig");
const tss = @import("tss.zig");
@@ -76,22 +78,22 @@ fn defaultFault(_: *const CpuState) noreturn {
/// Names for the 32 defined exception vectors, for readable output.
const names = [_][]const u8{
"divide error", "debug",
"NMI", "breakpoint",
"overflow", "bound range exceeded",
"invalid opcode", "device not available",
"double fault", "coprocessor segment overrun",
"invalid TSS", "segment not present",
"stack-segment fault", "general protection fault",
"page fault", "reserved (15)",
"x87 floating-point", "alignment check",
"machine check", "SIMD floating-point",
"virtualization", "control protection",
"reserved (22)", "reserved (23)",
"reserved (24)", "reserved (25)",
"reserved (26)", "reserved (27)",
"hypervisor injection", "VMM communication",
"security exception", "reserved (31)",
"divide error", "debug",
"NMI", "breakpoint",
"overflow", "bound range exceeded",
"invalid opcode", "device not available",
"double fault", "coprocessor segment overrun",
"invalid TSS", "segment not present",
"stack-segment fault", "general protection fault",
"page fault", "reserved (15)",
"x87 floating-point", "alignment check",
"machine check", "SIMD floating-point",
"virtualization", "control protection",
"reserved (22)", "reserved (23)",
"reserved (24)", "reserved (25)",
"reserved (26)", "reserved (27)",
"hypervisor injection", "VMM communication",
"security exception", "reserved (31)",
};
pub fn vectorName(vector: u64) []const u8 {
@@ -162,8 +164,9 @@ pub fn loadOnThisCpu() void {
}
/// Called by isr_common (isr.s) with a pointer to the trap frame. Exported so the
/// assembly stubs can `call` it by name. Exceptions are terminal; device
/// interrupts run their handler, get acknowledged, and return.
/// assembly stubs can `call` it by name. Exceptions never return here (on_fault
/// kills the faulting process or halts the core); device interrupts run their
/// handler, get acknowledged, and return.
export fn interruptDispatch(state: *CpuState) callconv(.c) void {
if (state.vector < 32) {
on_fault(state); // CPU exception — never returns
+26 -4
View File
@@ -166,8 +166,7 @@ pub fn init(allocFrame: *const fn () ?u64, freeFrame: *const fn (u64) void, boot
asm volatile ("mov %[pml4], %%cr3"
:
: [pml4] "r" (pml4),
: .{ .memory = true }
);
: .{ .memory = true });
on_own_tables = true; // now on the kernel's physmap (covers all RAM)
init_done = true; // the kernel half is fixed from here
}
@@ -273,6 +272,30 @@ pub fn mapUserDeviceInto(pml4: u64, virtual: u64, physical: u64, len: u64) void
}
}
/// Map `[physical, physical+len)` into the user half rooted at `pml4` as **coherent
/// DMA memory**: strong-uncacheable (PCD|PWT — a device reads/writes this RAM without
/// snooping the CPU caches) but, unlike `mapUserDeviceInto`, **without** `device_grant`
/// — because these frames are real RAM from `pmm.allocContiguous`, so teardown
/// (`freeSubtree`) must return them to the allocator like any other user page. RW + NX.
/// The caller aligns `virtual`/`physical` and places `virtual` in the DMA arena.
pub fn mapUserDmaInto(pml4: u64, virtual: u64, physical: u64, len: u64) void {
const flags: u64 = present | user | writable | no_execute | pcd | pwt;
const first = physical & ~@as(u64, page_size - 1);
const last = (physical + (if (len == 0) 1 else len) - 1) & ~@as(u64, page_size - 1);
var off: u64 = 0;
while (first + off <= last) : (off += page_size) {
const v = virtual + off;
const pml4e = &tableAt(pml4)[(v >> 39) & 0x1FF];
const pdpt = descendUser(pml4e);
const pdpte = &tableAt(pdpt)[(v >> 30) & 0x1FF];
const pd = descendUser(pdpte);
const pde = &tableAt(pd)[(v >> 21) & 0x1FF];
const pt = descendUser(pde);
tableAt(pt)[(v >> 12) & 0x1FF] = ((first + off) & address_mask) | flags;
invalidate(v);
}
}
/// Create a new address space: a fresh PML4 with an empty user half and the
/// kernel's higher half shared in (copying PML4[256..512), whose entries point
/// at the kernel's PDPTs — pre-created at init and never restaled, so growth in
@@ -385,6 +408,5 @@ fn invalidate(virtual: u64) void {
\\invlpg (%%rax)
:
: [v] "r" (virtual),
: .{ .rax = true, .memory = true }
);
: .{ .rax = true, .memory = true });
}
-2
View File
@@ -154,5 +154,3 @@ pub const Console = struct {
while (x < self.fb.width) : (x += 1) destination[x] = source[x];
}
};
+2
View File
@@ -66,6 +66,7 @@ fn record(node: *platform.Device, parent_id: u64) u64 {
d.id = count;
d.parent = parent_id;
d.class = @intFromEnum(node.class);
d.pci_class = if (node.ids.pci_class) |code| code else device_abi.no_pci_class;
const h = node.hid();
d.hid_len = @min(h.len, d.hid.len);
@memcpy(d.hid[0..d.hid_len], h[0..d.hid_len]);
@@ -170,6 +171,7 @@ pub fn register(parent_id: u64, owner: u32, descriptor: *const device_abi.Device
d.id = count;
d.parent = parent_id;
d.class = descriptor.class;
d.pci_class = descriptor.pci_class;
d.hid_len = @min(descriptor.hid_len, d.hid.len);
@memcpy(d.hid[0..@intCast(d.hid_len)], descriptor.hid[0..@intCast(d.hid_len)]);
d.resource_count = descriptor.resource_count;
+108 -2
View File
@@ -46,6 +46,9 @@ pub const EFAULT: i64 = 3; // buffer unmapped / out of the user half
pub const ENOENT: i64 = 4; // no such registered service
pub const ENOSPC: i64 = 5; // handle table or registry full
pub const ENOMEM: i64 = 6; // out of memory
pub const EPEER: i64 = 7; // peer died before replying (its process exited or was killed)
pub const ESRCH: i64 = 8; // no such process (process_kill of an unknown/dead id)
pub const EPERM: i64 = 9; // not permitted (process_kill by anyone but the supervisor)
/// A badge with this bit set is an asynchronous notification (e.g. an IRQ), not a
/// message from a client — there is no reply owed. The low bits carry the source
@@ -54,6 +57,29 @@ pub const ENOMEM: i64 = 6; // out of memory
/// shared kernel↔user ABI (system/abi.zig), because ring 3 has to test the same bit.
pub const notify_badge_bit: u64 = abi.notify_badge_bit;
/// Set (with `notify_badge_bit`) when a `replyWait` wake carries a buffered payload
/// posted by `send` (`ipc_send`), rather than a bare IRQ/exit notification. Shared with
/// ring 3 through the ABI so the receiver can tell "a message arrived" from "the hardware
/// spoke".
pub const notify_message_bit: u64 = abi.notify_message_bit;
/// Largest payload a single `send` (`ipc_send`) may post. Kept small — the payload rides
/// inline in every `Endpoint`, and the async path is for events (a `KeyEvent` is 16
/// bytes), not bulk transfer, which is what `call` and future shared pages are for.
pub const POST_MAXIMUM: usize = 64;
/// Depth of an endpoint's async payload ring. Absorbs a burst while a receiver is briefly
/// busy; a full ring drops the *oldest* message (see `send`).
const post_capacity: usize = 16;
/// One buffered message: a length-prefixed payload plus the sender's task id (delivered
/// in the low bits of the receiver's badge).
const PostSlot = struct {
length: u16 = 0,
sender_id: u64 = 0,
bytes: [POST_MAXIMUM]u8 = undefined,
};
/// End of the user (low) canonical half — user buffers must lie below it.
const user_half_end: u64 = 0x0000_8000_0000_0000;
@@ -71,9 +97,15 @@ pub const Endpoint = struct {
notify_buffer: [8]u64 = undefined,
notify_head: u8 = 0,
notify_tail: u8 = 0,
// Pending buffered messages (payloads posted by `send`), a small FIFO ring. Unlike
// notifications — which are a level and coalesce — these are discrete messages, so a
// full ring drops the oldest rather than merging.
post_buffer: [post_capacity]PostSlot = undefined,
post_head: u16 = 0,
post_tail: u16 = 0,
};
pub fn createEndpoint() ?*Endpoint {
pub fn createIpcEndpoint() ?*Endpoint {
const endpoint = heap.allocator().create(Endpoint) catch return null;
endpoint.* = .{};
return endpoint;
@@ -92,6 +124,7 @@ pub fn dropRef(endpoint: *Endpoint) void {
// --- sender FIFO (endpoint-local, via Task.next) ----------------------------
fn enqueueSender(endpoint: *Endpoint, t: *Task) void {
t.ipc_wait_endpoint = @ptrCast(endpoint); // so a kill can unlink a parked caller
t.next = null;
if (endpoint.sender_tail) |tail| tail.next = t else endpoint.sender_head = t;
endpoint.sender_tail = t;
@@ -101,10 +134,34 @@ fn dequeueSender(endpoint: *Endpoint) ?*Task {
const t = endpoint.sender_head orelse return null;
endpoint.sender_head = t.next;
if (endpoint.sender_head == null) endpoint.sender_tail = null;
t.ipc_wait_endpoint = null;
t.next = null;
return t;
}
/// Unlink `t` from the sender FIFO it queues in, if any — the kill path for a
/// client parked in `call` that no server has received yet. Without this, a dead
/// caller would later be dequeued as a dangling pointer. The endpoint is still
/// alive here: `t`'s own handle table holds a reference until closeHandles runs
/// (which the kill path does *after* this). Precondition: the big kernel lock is
/// held.
pub fn abandonSenderLocked(t: *Task) void {
const endpoint: *Endpoint = @ptrCast(@alignCast(t.ipc_wait_endpoint orelse return));
t.ipc_wait_endpoint = null;
var previous: ?*Task = null;
var node = endpoint.sender_head;
while (node) |n| : ({
previous = n;
node = n.next;
}) {
if (n != t) continue;
if (previous) |p| p.next = t.next else endpoint.sender_head = t.next;
if (endpoint.sender_tail == t) endpoint.sender_tail = previous;
t.next = null;
return;
}
}
// --- cross-address-space copy ----------------------------------------------
/// Copy `len` bytes from `source_va` in address space `source_as` to `destination_va` in
@@ -237,12 +294,23 @@ pub fn replyWait(endpoint: *Endpoint, reply_ptr: u64, reply_len: u64, receive_pt
scheduler.readyLocked(client); // its `call` now returns
}
// (2) Receive the next request (or notification), blocking until one is ready.
// (2) Receive the next request (or notification / buffered message), blocking until
// one is ready. Bare notifications (IRQ/exit) come first — they're latency-sensitive
// and carry no payload — then buffered messages, then synchronous client requests.
while (true) {
if (popNotify(endpoint)) |badge| {
out_badge.* = badge | notify_badge_bit;
return 0; // notification: no payload, no reply owed, no cap
}
if (popPost(endpoint)) |slot| {
const n = @min(@as(usize, slot.length), receive_cap);
// Copy from the kernel-resident ring slot (source aspace 0) into the receiver.
if (!copyAcross(0, @intFromPtr(&slot.bytes), me.aspace, receive_ptr, n)) {
continue; // bad receive buffer: drop this message, keep serving
}
out_badge.* = slot.sender_id | notify_badge_bit | notify_message_bit;
return @intCast(n); // async message: payload delivered, no reply owed, no cap
}
if (dequeueSender(endpoint)) |caller| {
const n = @min(caller.ipc_send_len, receive_cap);
if (!copyAcross(caller.aspace, caller.ipc_send_ptr, me.aspace, receive_ptr, n)) {
@@ -278,6 +346,44 @@ fn popNotify(endpoint: *Endpoint) ?u64 {
return badge;
}
/// Take the oldest buffered message from the post ring, or null if empty. Returns a
/// pointer into the endpoint's own storage — valid until the next `send`/`popPost` under
/// the same lock region, which is all the copy-out in `replyWait` needs.
fn popPost(endpoint: *Endpoint) ?*const PostSlot {
if (endpoint.post_head == endpoint.post_tail) return null;
const slot = &endpoint.post_buffer[endpoint.post_head % post_capacity];
endpoint.post_head +%= 1;
return slot;
}
/// Client-free side of async IPC (`ipc_send`): copy `[source_va, len)` from address space
/// `source_as` into `endpoint`'s post ring and wake a waiting receiver — **without
/// blocking the sender** and with no reply owed. `sender_id` rides along, delivered in the
/// low bits of the receiver's badge. Returns 0, or a negative errno (`-E2BIG` if the
/// payload exceeds `POST_MAXIMUM`, `-EFAULT` if the source buffer is unmapped / out of the
/// user half). A full ring drops the *oldest* message (advancing `post_head`), because a
/// buffered message is discrete, not a level: keeping the newest keeps input responsive.
/// Precondition: the big kernel lock is held.
pub fn sendLocked(endpoint: *Endpoint, source_as: u64, source_va: u64, len: u64, sender_id: u64) i64 {
if (len > POST_MAXIMUM) return -E2BIG;
// Drop the oldest if the ring is full, so this newest message always lands.
if (endpoint.post_tail -% endpoint.post_head >= post_capacity) endpoint.post_head +%= 1;
const slot = &endpoint.post_buffer[endpoint.post_tail % post_capacity];
if (!copyFromUser(source_as, source_va, slot.bytes[0..@intCast(len)])) return -EFAULT;
slot.length = @intCast(len);
slot.sender_id = sender_id;
endpoint.post_tail +%= 1;
scheduler.wakeLocked(&endpoint.receive_wait_queue);
return 0;
}
/// `sendLocked` wrapped in its own critical section, for the `ipc_send` syscall path.
pub fn send(endpoint: *Endpoint, source_as: u64, source_va: u64, len: u64, sender_id: u64) i64 {
const flags = sync.enter();
defer sync.leave(flags);
return sendLocked(endpoint, source_as, source_va, len, sender_id);
}
/// Post an asynchronous notification carrying `badge` to `endpoint` and wake a waiting
/// receiver. Precondition: the big kernel lock is held.
///
+41 -4
View File
@@ -67,6 +67,13 @@ var vector_gsi: [256]u32 = .{no_gsi} ** 256;
/// Vector currently assigned to each GSI (0 = none), so a rebind reuses it.
var gsi_vector: [maximum_gsi]u8 = .{0} ** maximum_gsi;
/// MSI bindings, keyed by **vector** (not GSI): an MSI has no I/O APIC redirection
/// entry, it is a per-device edge-triggered vector the device raises by writing the
/// (address, data) pair `msi_bind` hands back. Read from the ISR and written from
/// syscalls, always under the big kernel lock, like `bound`.
var msi_bound: [256]?*ipc_sync.Endpoint = .{null} ** 256;
var msi_owner: [256]u32 = .{0} ** 256;
/// Set once the trampolines are installed.
var installed = false;
@@ -75,9 +82,15 @@ var installed = false;
fn dispatch(vector: u8) void {
const gsi = vector_gsi[vector];
if (gsi == no_gsi) {
// Nothing is routed here. Acknowledge so the LAPIC doesn't wedge on an
// in-service bit that never clears, but touch no redirection entry.
// No GSI is routed here. It may be an MSI vector (edge-triggered, unshared, no
// I/O APIC entry): just EOI and notify — there is no line to mask and no ack
// cycle. Or it's spurious, and the EOI alone keeps the LAPIC from wedging on an
// in-service bit. Same lock discipline as below: msi_bound[vector] is read and
// used under the acquisition `msiRelease` writes it under.
_ = sync.enter();
defer sync.leaveIsr();
architecture.irqEoi();
if (msi_bound[vector]) |endpoint| ipc_sync.notifyLocked(endpoint, vector);
return;
}
@@ -126,9 +139,9 @@ pub fn init() void {
fn allocVector() ?u8 {
var v: u8 = architecture.irq_vector_base;
while (v < architecture.irq_vector_base + architecture.irq_vector_count) : (v += 1) {
var used = false;
var used = msi_bound[v] != null; // taken by an MSI binding
for (gsi_vector) |gv| {
if (gv == v) used = true;
if (gv == v) used = true; // taken by a GSI binding
}
if (!used) return v;
}
@@ -162,6 +175,21 @@ pub fn bind(gsi: u32, endpoint: *ipc_sync.Endpoint, owner: u32) BindError!void {
architecture.irqUnmask(gsi);
}
pub const MsiError = error{NoVector};
/// Allocate an interrupt vector and bind it to `endpoint` for task `owner`, returning
/// the vector. The `msi_bind` mechanism: MSI is edge-triggered and unshared, so there
/// is no GSI, no I/O APIC redirection entry, no mask, and no ack cycle — the device
/// raises the interrupt by *writing* the (address, data) pair the syscall derives from
/// this vector, and `dispatch` just EOIs and notifies. Caller holds the big kernel
/// lock and has verified `owner` claimed the device. See docs/driver-model.md (M15).
pub fn msiBind(endpoint: *ipc_sync.Endpoint, owner: u32) MsiError!u8 {
const vector = allocVector() orelse return error.NoVector;
msi_bound[vector] = endpoint;
msi_owner[vector] = owner;
return vector;
}
/// Re-arm `gsi` after the driver has quieted the device. Caller holds the big lock
/// and has verified ownership.
pub fn ack(gsi: u32) bool {
@@ -187,4 +215,13 @@ pub fn releaseOwner(owner: u32) void {
gsi_vector[gsi] = 0;
}
}
// MSI bindings carry no I/O APIC entry to mask — just drop them so a stale vector
// stops notifying a freed endpoint. A late edge on a released vector is spurious
// (msi_bound is null) and `dispatch` EOIs it harmlessly.
for (&msi_bound, 0..) |*slot, vector| {
if (slot.* != null and msi_owner[vector] == owner) {
slot.* = null;
msi_owner[vector] = 0;
}
}
}
+37 -8
View File
@@ -87,7 +87,7 @@ fn kmain(boot_information: *const BootInformation) noreturn {
log.print(" pitch : {d} bytes\n", .{fb.pitch});
log.print(" format : {s}\n", .{@tagName(fb.format)});
log.print(" framebuffer: 0x{x:0>16}\n", .{fb.base});
log.print (" footprint : {d} MiB\n", .{(fb.pitch * fb.height) / (1024 * 1024)});
log.print(" footprint : {d} MiB\n", .{(fb.pitch * fb.height) / (1024 * 1024)});
// Summarise the physical memory the loader handed us. The array is danos's
// own MemoryRegion, so this is a plain slice — no firmware layout in sight.
@@ -284,7 +284,7 @@ fn kmain(boot_information: *const BootInformation) noreturn {
if (boot_information.init_len != 0) {
status("starting /system/services/init...\n");
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len];
process.spawnProcess(image, 4) catch |err| {
process.spawnProcess(image, 4, &.{"/system/services/init"}) catch |err| {
statusPrint("/system/services/init failed to load: {s}\n", .{@errorName(err)});
};
} else {
@@ -385,13 +385,42 @@ fn kib(frames: u64) u64 {
return frames * abi.page_size / (1024);
}
/// Report a CPU exception and halt **this core**. There's no fault recovery yet, so
/// the faulting core is terminal — but the fault is *contained* to it: on an
/// application processor only that core stops, and the rest of the system keeps
/// running (full recovery — kill the task, keep the core — is the resilience track,
/// see docs/resilience.md). The report names the core so an AP fault is attributed,
/// and goes to every output sink plus a POST code and a persistent breadcrumb.
/// Whether a ring-3 exception is attributable to the process that raised it — and
/// therefore recoverable by killing that process. NMI (2), double fault (8), and
/// machine check (18) report machine or kernel trouble even when they arrive with a
/// user CS (an NMI interrupts whatever happens to be running), so they stay terminal.
fn recoverableFault(vector: u64) bool {
return switch (vector) {
2, 8, 18 => false,
else => true,
};
}
/// Report a CPU exception. Two outcomes (docs/resilience.md):
///
/// **A fault taken in user mode kills the faulting process, not the machine.** The
/// kernel is intact — the CPU trapped onto the task's kernel stack — so the process
/// is killed, everything it held (address space, IRQ bindings, IPC handles, device
/// grants' frames) is reclaimed, and the core reschedules. A crashing driver takes
/// itself down, never the OS.
///
/// **Everything else halts this core.** A kernel-mode fault means the trusted base
/// itself is broken — there is nothing safe to kill — and NMI/#DF/#MC report machine
/// trouble regardless of CS (`recoverableFault`). Even then the fault is *contained*:
/// on an application processor only that core stops and the rest keep running. The
/// report names the core so an AP fault is attributed, and goes to every output sink
/// plus a POST code and a persistent breadcrumb. (A ring-3 fault on a *borrowed*
/// kernel thread — process.run, the user-pf isolation probe — also lands here: there
/// is no scheduled process to kill.)
fn onException(state: *const architecture.CpuState) noreturn {
if (architecture.fromUser(state) and scheduler.currentIsUserProcess() and recoverableFault(state.vector)) {
statusPrint("\ndanos: process {d} ({s}) killed by {s} (vector {d}) on core {d}\n", .{ scheduler.currentId(), scheduler.current().name(), architecture.exceptionName(state.vector), state.vector, scheduler.currentCpuIndex() });
statusPrint(" error code : 0x{x}\n", .{state.error_code});
statusPrint(" IP : 0x{x:0>16}\n", .{architecture.instructionPointer(state)});
if (architecture.faultAddress(state)) |address| statusPrint(" fault addr : 0x{x:0>16}\n", .{address});
process.killCurrentProcess(); // reclaims everything, reschedules; never returns
}
log.checkpoint(cp_exception);
const core = scheduler.currentCpuIndex();
// A fault is user-facing enough to paint on screen too (via statusPrint), on
+28
View File
@@ -173,3 +173,31 @@ pub fn free(address: u64) void {
used_frames -= 1;
if (f < next_hint) next_hint = f;
}
/// Allocate `count` physically **contiguous** frames whose highest byte is below
/// `max_phys` (pass `~0` for no limit; use a real limit for DMA engines with 32-bit
/// addressing). Returns the physical base, or null if no free run of that size fits.
/// A DMA descriptor ring needs contiguity, a known physical address, and pinning —
/// none of which one-frame `alloc` gives. Linear scan for a run of clear bits: fine
/// for the small rings DMA needs; a buddy allocator is a later optimisation. The
/// frames are freed individually with `free`, so there is no bespoke free path.
pub fn allocContiguous(count: usize, max_phys: u64) ?u64 {
if (count == 0) return null;
const limit: usize = @intCast(@min(@as(u64, total_frames), max_phys / page_size));
var start: usize = 1; // frame 0 stays reserved as the "none" address
while (start + count <= limit) {
if (isUsed(start)) {
start += 1;
continue;
}
var run: usize = 0;
while (run < count and !isUsed(start + run)) : (run += 1) {}
if (run == count) {
for (0..count) |i| setUsed(start + i);
used_frames += count;
return @as(u64, start) * page_size;
}
start += run + 1; // the frame at start+run is used; skip past it
}
return null; // no contiguous run of `count` frames below max_phys
}
+505 -46
View File
@@ -24,6 +24,7 @@ const elf = std.elf;
const boot_handoff = @import("boot-handoff");
const abi = @import("abi");
const device_abi = @import("device-abi");
const parameters = @import("parameters");
const architecture = @import("architecture");
const pmm = @import("pmm.zig");
const scheduler = @import("scheduler.zig");
@@ -40,9 +41,18 @@ const SystemCall = abi.SystemCall;
/// User virtual addresses. PML4 index 224 — a user-exclusive region, far from
/// the identity map (low indices) and the vmm test address (index 128), so
/// setting the U/S bit on its intermediate tables widens no kernel mapping.
/// An ELF image may occupy [code_virtual, stack_virtual); the stack page sits above.
/// An ELF image may occupy [code_virtual, stack_virtual); the stack sits above.
pub const code_virtual: u64 = 0x0000_7000_0000_0000;
pub const stack_virtual: u64 = 0x0000_7000_0020_0000;
/// The stack region, above the image. The page at `stack_virtual` is **never
/// mapped** — it is the guard page: a process that overflows its stack walks into
/// it and faults (killing only that process) rather than silently corrupting the
/// top of its own image. The stack proper is `parameters.user_stack_pages` pages
/// at [stack_base_virtual, stack_top_virtual), RW + NX, with the System V entry
/// block (argc/argv) at the very top.
pub const stack_virtual: u64 = 0x0000_7000_0020_0000; // guard page (unmapped)
pub const stack_base_virtual: u64 = stack_virtual + page_size;
pub const stack_top_virtual: u64 = stack_base_virtual + parameters.user_stack_pages * page_size;
/// The mmap grant arena: where `mmap` hands out fresh user pages, above the image
/// and stack but still inside PML4[224] (so no kernel mapping is widened). Each
@@ -62,10 +72,30 @@ pub const user_half_end: u64 = 0x0000_8000_0000_0000;
pub const device_arena_base: u64 = 0x0000_7100_0000_0000;
pub const device_arena_end: u64 = device_arena_base + (4 << 30);
/// The DMA arena: where `dma_alloc` places coherent DMA buffers, in PML4[228] — a
/// user-exclusive region distinct from the MMIO arena. Unlike MMIO grants these back
/// real RAM (contiguous frames), so they are reclaimed on teardown. Per-process cursor
/// in `Task.dma_map_next`.
pub const dma_arena_base: u64 = 0x0000_7200_0000_0000;
pub const dma_arena_end: u64 = dma_arena_base + (256 << 20); // 256 MiB per process
/// Largest single `mmap` grant, in pages (1 MiB). The user heap grows in small
/// chunks, so this bound is generous; it also caps the frame scratch array below.
const maximum_mmap_pages = 256;
/// Ceiling on a process's argv entries, including argv[0]. Arguments are spawn
/// parameters ("you are the driver for device 12"), not bulk data — IPC carries
/// that — so the bound is small and everything fits the single stack page.
pub const maximum_arguments = 8;
/// Ceiling on the `system_spawn` extra-arguments blob (argv[1..], NUL-separated).
pub const maximum_argument_bytes = 256;
/// Auxiliary-vector entry types (System V AMD64 process entry). Only what the
/// kernel emits today; a C runtime scans the vector until the null terminator.
const auxiliary_vector_null: u64 = 0; // AT_NULL — end of the vector
const auxiliary_vector_page_size: u64 = 6; // AT_PAGESZ
// The hand-assembled user program blob (isr.s, .rodata) — the isolation probe.
const pf_start = @extern([*]const u8, .{ .name = "user_pf_start" });
const pf_end = @extern([*]const u8, .{ .name = "user_pf_end" });
@@ -99,9 +129,14 @@ pub fn setInitialRamdisk(image: []const u8) void {
/// written back into the trap frame, since the entry paths restore user registers
/// from it. One handler serves both the system_call/sysret and int-0x80 entry paths.
///
/// Install it once at boot (before any user code runs) via `init`.
/// Install it once at boot (before any user code runs) via `init`. Also registers
/// the scheduler's kill hooks: the scheduler sits below this layer, so finishing a
/// deferred process_kill (IRQ bindings, IPC handles, the exit notification) is
/// called back up into here from the tick (see scheduler.reapKillPendingLocked).
pub fn init() void {
architecture.setSystemCallHandler(system_call);
scheduler.terminate_current_hook = terminateCurrentLocked;
scheduler.reap_task_hook = reapTaskLocked;
}
/// Return -1 (as an unsigned bit pattern) in the system_call result register.
@@ -110,21 +145,26 @@ fn fail(state: *architecture.CpuState) void {
}
fn system_call(state: *architecture.CpuState) void {
const t = scheduler.current();
const user = t.aspace != 0;
if (user) {
// A condemned process (process_kill caught it running) dies at its next
// kernel entry — before it can spawn, claim, or message anything else.
if (t.kill_pending) terminateCurrent();
// Mark the span of this call so the timer tick never tears the task down
// in the middle of a kernel operation (scheduler.reapKillPendingLocked).
t.in_system_call = true;
}
defer if (user) {
t.in_system_call = false;
};
switch (@as(SystemCall, @enumFromInt(architecture.systemCallNumber(state)))) {
.exit => {
exit_code = architecture.systemCallArg(state, 0);
// A scheduled process drops its endpoint references, frees its address
// space, and reschedules; a borrowed test thread unwinds back to the
// kernel that entered it.
// A scheduled process tears down fully (terminateCurrent); a borrowed
// test thread unwinds back to the kernel that entered it.
if (scheduler.currentIsUserProcess()) {
// Unbind before closeHandles: dropping the last reference destroys the
// Endpoint, and a still-bound GSI would have an ISR call
// notifyFromIsr on freed memory the next time the device fired.
// unbindAll also leaves the line masked, so a dead driver's device
// goes quiet rather than storming.
releaseIrqs(scheduler.current());
ipc.closeHandles(scheduler.current());
scheduler.exitUser();
terminateCurrent();
} else architecture.userExit();
},
.yield => {
@@ -138,11 +178,12 @@ fn system_call(state: *architecture.CpuState) void {
.debug_write => systemDebugWrite(state),
.mmap => systemMmap(state),
.munmap => systemMunmap(state),
.create_endpoint => systemCreateEndpoint(state),
.create_ipc_endpoint => systemCreateIpcEndpoint(state),
.ipc_register => systemIpcRegister(state),
.ipc_lookup => systemIpcLookup(state),
.ipc_call => systemIpcCall(state),
.ipc_reply_wait => systemIpcReplyWait(state),
.ipc_send => systemIpcSend(state),
.device_enumerate => systemDeviceEnumerate(state),
.device_claim => systemDeviceClaim(state),
.mmio_map => systemMmioMap(state),
@@ -150,6 +191,14 @@ fn system_call(state: *architecture.CpuState) void {
.irq_ack => systemIrqAck(state),
.device_register => systemDeviceRegister(state),
.system_spawn => systemSpawn(state),
.dma_alloc => systemDmaAlloc(state),
.dma_free => systemDmaFree(state),
.msi_bind => systemMsiBind(state),
.io_read => systemIoRead(state),
.io_write => systemIoWrite(state),
.clock => systemClock(state),
.process_enumerate => systemProcessEnumerate(state),
.process_kill => systemProcessKill(state),
_ => fail(state),
}
}
@@ -159,10 +208,10 @@ fn failErr(state: *architecture.CpuState, errno: i64) void {
architecture.setSystemCallResult(state, @bitCast(-errno));
}
/// create_endpoint() -> handle: allocate an endpoint and install it in the
/// create_ipc_endpoint() -> handle: allocate an endpoint and install it in the
/// caller's handle table.
fn systemCreateEndpoint(state: *architecture.CpuState) void {
const endpoint = ipc.createEndpoint() orelse return failErr(state, ipc.ENOMEM);
fn systemCreateIpcEndpoint(state: *architecture.CpuState) void {
const endpoint = ipc.createIpcEndpoint() orelse return failErr(state, ipc.ENOMEM);
const h = ipc.installHandle(scheduler.current(), endpoint);
if (h < 0) {
ipc.dropRef(endpoint);
@@ -215,6 +264,18 @@ fn systemIpcReplyWait(state: *architecture.CpuState) void {
architecture.setSystemCallResult3(state, received_cap);
}
/// ipc_send(handle, message_ptr, message_len) -> 0/-errno: post a payload to an
/// endpoint's async queue and wake a receiver, without blocking the caller. The async
/// counterpart of ipc_call — for broadcasts (the input service) where a rendezvous would
/// let one dead subscriber hang the sender. Delivered through ipc_reply_wait as a
/// buffered message (badge carries notify_message_bit and the caller's task id).
fn systemIpcSend(state: *architecture.CpuState) void {
const me = scheduler.current();
const endpoint = ipc.resolveHandle(me, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
const r = ipc.send(endpoint, me.aspace, architecture.systemCallArg(state, 1), architecture.systemCallArg(state, 2), me.id);
architecture.setSystemCallResult(state, @bitCast(r));
}
/// device_enumerate(buffer, maximum) -> total: snapshot the device table into the caller's
/// buffer (up to `maximum` entries), returning the total device count.
fn systemDeviceEnumerate(state: *architecture.CpuState) void {
@@ -261,6 +322,107 @@ fn systemMmioMap(state: *architecture.CpuState) void {
architecture.setSystemCallResult(state, base_v + (r.start & (page_size - 1))); // register base
}
/// Resolve a port-I/O access against the caller's claims. The device must be claimed by
/// `t`, `resource_index` must name one of its `io_port` resources, and the access
/// `[offset, offset+width)` must fall wholly inside it. Returns the absolute 16-bit
/// port, or null if the capability check fails. The claim plus the discovered `io_port`
/// resource are the capability — exactly like `mmio_map` for memory, so a driver can
/// only touch the ports its device actually owns, never a raw `in`/`out` to anywhere.
pub fn resolveIoPort(t: *scheduler.Task, device_id: u64, resource_index: u64, offset: u64, width: u64) ?u16 {
if (width != 1 and width != 2 and width != 4) return null;
const owner = devices_broker.ownerOf(device_id) orelse return null;
if (owner != t.id) return null; // not claimed by this process
const r = devices_broker.resourceOf(device_id, resource_index) orelse return null;
if (r.kind != @intFromEnum(device_abi.ResourceKind.io_port)) return null;
if (offset + width > r.len) return null; // access escapes the claimed port range
const port = r.start + offset;
if (port + width > 0x1_0000) return null; // I/O ports are 16-bit
return @intCast(port);
}
/// io_read(device_id, resource_index, offset, width) -> value: read `width` bytes (1/2/4)
/// from a port in a claimed device's `io_port` resource. Ring 3 has no direct `in`/`out`
/// (no TSS I/O bitmap, IOPL never raised), so a legacy driver (PS/2, 16550 UART) reaches
/// its ports through this claim-gated call — low-rate hardware, so a syscall per access
/// is fine. See docs/drivers.md.
fn systemIoRead(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const width = architecture.systemCallArg(state, 3);
const port = resolveIoPort(t, architecture.systemCallArg(state, 0), architecture.systemCallArg(state, 1), architecture.systemCallArg(state, 2), width) orelse return fail(state);
architecture.setSystemCallResult(state, architecture.pioRead(@intCast(width), port));
}
/// io_write(device_id, resource_index, offset, width, value) -> 0: write `value` (low
/// `width` bytes) to a port in a claimed device's `io_port` resource. Same capability
/// gate as `io_read`.
fn systemIoWrite(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const width = architecture.systemCallArg(state, 3);
const port = resolveIoPort(t, architecture.systemCallArg(state, 0), architecture.systemCallArg(state, 1), architecture.systemCallArg(state, 2), width) orelse return fail(state);
architecture.pioWrite(@intCast(width), port, @intCast(architecture.systemCallArg(state, 4)));
architecture.setSystemCallResult(state, 0);
}
/// dma_alloc(len, flags) -> vaddr (rax), paddr (rdx): grant `len` bytes (rounded up to
/// whole pages) of DMA-capable memory — physically contiguous, zeroed, pinned, and
/// strong-uncacheable (coherent) — mapping it into the caller's DMA arena and handing
/// back both the virtual address to touch and the physical address to program into the
/// device. This is what `sysMmap` can't do: mmap frames are scattered, cacheable, and
/// their physical address is never disclosed. `dma_below_4g` caps the physical address
/// for legacy engines; `dma_write_combining` is accepted but falls back to coherent
/// until PAT is programmed. See docs/driver-model.md (M14).
fn systemDmaAlloc(state: *architecture.CpuState) void {
const len = architecture.systemCallArg(state, 0);
const flags = architecture.systemCallArg(state, 1);
const t = scheduler.current();
if (t.aspace == 0 or len == 0) return fail(state);
const pages: usize = @intCast((len + page_size - 1) / page_size);
const max_phys: u64 = if (flags & abi.dma_below_4g != 0) (@as(u64, 4) << 30) else ~@as(u64, 0);
const phys = pmm.allocContiguous(pages, max_phys) orelse return fail(state);
if (t.dma_map_next == 0) t.dma_map_next = dma_arena_base;
const base_v = t.dma_map_next;
if (base_v + pages * page_size > dma_arena_end) {
for (0..pages) |i| pmm.free(phys + i * page_size); // arena exhausted; give the frames back
return fail(state);
}
// Zero through the physmap (the frames aren't mapped in the caller yet), then map.
const kernel_view: [*]u8 = @ptrFromInt(boot_handoff.physicalToVirtual(phys));
@memset(kernel_view[0 .. pages * page_size], 0);
architecture.mapUserDmaInto(t.aspace, base_v, phys, pages * page_size);
t.dma_map_next = base_v + pages * page_size;
architecture.setSystemCallResult(state, base_v); // virtual address for the CPU
architecture.setSystemCallResult2(state, phys); // physical address for the device
}
/// dma_free(vaddr, len) -> 0: release a prior `dma_alloc`. Bounded to the DMA arena so
/// it can never unmap-and-free the caller's stack, heap, or an MMIO grant; only pages
/// actually mapped are freed (an unmapped hole is skipped). Teardown also reclaims any
/// DMA pages left mapped at exit (they carry no `device_grant`, so `freeSubtree` frees
/// them as ordinary RAM), so a driver that just dies leaks nothing.
fn systemDmaFree(state: *architecture.CpuState) void {
const base_v = architecture.systemCallArg(state, 0);
const len = architecture.systemCallArg(state, 1);
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const pages: usize = @intCast((len + page_size - 1) / page_size);
if (base_v < dma_arena_base or base_v + pages * page_size > dma_arena_end) return fail(state);
for (0..pages) |i| {
const va = base_v + i * page_size;
if (architecture.translate(t.aspace, va)) |phys| {
architecture.unmapUserPageInto(t.aspace, va);
pmm.free(phys);
}
}
architecture.setSystemCallResult(state, 0);
}
/// device_register(parent_id, descriptor_ptr) -> id: publish a child device below a device
/// this process has claimed. The bus-driver primitive: a process that owns a bus
/// enumerates it and hands each device it finds to the table, where a class driver
@@ -286,40 +448,202 @@ fn systemDeviceRegister(state: *architecture.CpuState) void {
architecture.setSystemCallResult(state, id);
}
/// system_spawn(name_ptr, name_len) -> 0 on success, -1 on failure. Load the binary
/// bundled in the initial-ramdisk under `name` as a fresh ring-3 process. This is the
/// mechanism a user-space supervisor (the device manager) uses to start a driver it
/// matched: discovery and policy stay in user space, the kernel only spawns.
/// system_spawn(name_ptr, name_len, arguments_ptr, arguments_len, exit_endpoint)
/// -> the child's process id on success, -1 on failure. Load the binary bundled in
/// the initial-ramdisk under `name` as a fresh ring-3 process. `name` becomes the
/// child's argv[0] (and its task name, so a fault report can say which binary
/// died); `arguments` is an optional NUL-separated blob that becomes argv[1..] —
/// how a supervisor parameterises what it starts ("you are the driver for device
/// 12"). 0/0 means no extra arguments. This is the mechanism a user-space
/// supervisor (the device manager) uses to start a driver it matched: discovery
/// and policy stay in user space, the kernel only spawns.
///
/// Ungated for now — any process may spawn any bundled binary. A capability (only a
/// supervisor holds the right to spawn) belongs here once the model grows one; see
/// docs/driver-model.md. The name is bounds-checked into the user half exactly like
/// `debug_write`, and an unknown name or a load failure returns -1.
/// The caller is recorded as the child's **supervisor** — the sole holder of the
/// right to `process_kill` it (docs/process-management.md). `exit_endpoint` (a
/// handle, or `abi.no_cap` for none) names an endpoint of the caller's to notify
/// when the child ends, any way it ends — the IRQ-as-IPC pattern reused as the
/// microkernel's SIGCHLD.
///
/// Spawning itself is still ungated — any process may spawn any bundled binary; a
/// spawn capability belongs here once the model grows one (docs/driver-model.md).
/// Both buffers are bounds-checked into the user half exactly like `debug_write`,
/// and an unknown name or a load failure returns -1.
fn systemSpawn(state: *architecture.CpuState) void {
const ptr = architecture.systemCallArg(state, 0);
const len = architecture.systemCallArg(state, 1);
if (len == 0 or len > 64 or ptr >= user_half_end or ptr + len > user_half_end) return fail(state);
const arguments_ptr = architecture.systemCallArg(state, 2);
const arguments_len = architecture.systemCallArg(state, 3);
const exit_handle = architecture.systemCallArg(state, 4);
const t = scheduler.current();
if (len == 0 or len > scheduler.maximum_task_name or ptr >= user_half_end or ptr + len > user_half_end) return fail(state);
if (arguments_len > maximum_argument_bytes) return fail(state);
if (arguments_len != 0 and (arguments_ptr >= user_half_end or arguments_ptr + arguments_len > user_half_end)) return fail(state);
const exit_endpoint: ?*ipc.Endpoint = if (exit_handle == abi.no_cap)
null
else
ipc.resolveHandle(t, exit_handle) orelse return failErr(state, ipc.EBADF);
const image = ramdisk_image orelse return fail(state);
const rd = initial_ramdisk.Reader.init(image) orelse return fail(state);
const name = @as([*]const u8, @ptrFromInt(ptr))[0..len];
var argv: [maximum_arguments][]const u8 = undefined;
argv[0] = name;
var argc: usize = 1;
if (arguments_len != 0) {
const blob = @as([*]const u8, @ptrFromInt(arguments_ptr))[0..arguments_len];
var pieces = std.mem.tokenizeScalar(u8, blob, 0);
while (pieces.next()) |piece| {
if (argc == maximum_arguments) return fail(state);
argv[argc] = piece;
argc += 1;
}
}
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!std.mem.eql(u8, item.name, name)) continue;
spawnProcess(item.blob, 4) catch return fail(state);
architecture.setSystemCallResult(state, 0);
const child = spawnProcessSupervised(item.blob, 4, argv[0..argc], t.id, exit_endpoint) catch return fail(state);
architecture.setSystemCallResult(state, child);
return;
}
fail(state); // no bundled binary by that name
}
/// Drop every IRQ binding `t` made. Called on exit, before the handle table is closed
/// (which is what frees the endpoints an ISR would otherwise notify into).
fn releaseIrqs(t: *scheduler.Task) void {
/// process_enumerate(buffer, maximum) -> total: snapshot the task table into the
/// caller's buffer (up to `maximum` `abi.ProcessDescriptor` entries), returning
/// the total live-task count — the exact shape of `device_enumerate`, so a `ps`
/// is a user program over a snapshot, not a kernel service. Read-only and
/// ungated: what is running is not a secret between cooperating bring-up
/// processes.
fn systemProcessEnumerate(state: *architecture.CpuState) void {
const buffer_ptr = architecture.systemCallArg(state, 0);
const maximum = architecture.systemCallArg(state, 1);
const t = scheduler.current();
if (t.aspace == 0 or buffer_ptr >= user_half_end) return fail(state);
const sz = @sizeOf(abi.ProcessDescriptor);
const cap = @min(maximum, (user_half_end - buffer_ptr) / sz); // clamp to the user half
const out: [*]abi.ProcessDescriptor = @ptrFromInt(buffer_ptr);
architecture.setSystemCallResult(state, scheduler.enumerate(out[0..@intCast(cap)]));
}
/// process_kill(id) -> 0 / -ESRCH / -EPERM: end the process `id`. Only its
/// supervisor — the process that spawned it — may do so; the supervision link is
/// the kill capability, so no user/permission model is needed and a stray id
/// cannot be a weapon (ids are never reused, so a stale one just misses).
fn systemProcessKill(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const id = architecture.systemCallArg(state, 0);
if (id > std.math.maxInt(u32)) return failErr(state, ipc.ESRCH);
const r = killProcess(t.id, @intCast(id));
architecture.setSystemCallResult(state, @bitCast(r));
}
/// Processes killed by a CPU fault rather than a clean exit. Evidence for the
/// fault-recovery test, and a health signal a supervisor can consult later.
pub var fault_kill_count: u64 = 0;
/// Release everything a dying task holds and tell its supervisor — the shared
/// half of every path out of a process: clean exit, fault kill, and process_kill
/// (both the immediate reap and the deferred tick-time terminate). The order
/// matters:
/// - IRQ bindings are dropped before the handle table closes: dropping the last
/// endpoint reference destroys the Endpoint, and a still-bound GSI would have an
/// ISR call notifyFromIsr on freed memory the next time the device fired.
/// `releaseOwner` also leaves the line masked, so a dead driver's device goes
/// quiet rather than storming.
/// - A client this task still owes a reply to (it died between receive and reply)
/// is failed with -EPEER rather than left blocked forever — a dead server must
/// not hang its callers.
/// - The task is unlinked from wherever IPC parked it (an endpoint's sender FIFO,
/// a receive wait queue, or a server's owed-reply slot) *before* the handles
/// close, so nothing ever dequeues a dangling pointer. These are no-ops for a
/// running task ending itself; they matter when process_kill reaps a blocked one.
/// - The exit notification is posted last, once the process can no longer act, so
/// a supervisor that receives it observes a fully-released child. The endpoint
/// reference taken at spawn is dropped with it.
/// Precondition: the big kernel lock is held.
fn releaseTaskResourcesLocked(t: *scheduler.Task) void {
irq.releaseOwner(t.id);
if (t.ipc_client) |client| {
t.ipc_client = null;
client.ipc_status = -ipc.EPEER;
scheduler.readyLocked(client); // its blocked `call` now returns the error
}
ipc.abandonSenderLocked(t);
scheduler.removeFromWaitQueueLocked(t);
scheduler.forgetIpcClientLocked(t);
ipc.closeHandles(t);
if (t.exit_endpoint) |raw| {
const endpoint: *ipc.Endpoint = @ptrCast(@alignCast(raw));
t.exit_endpoint = null;
ipc.notifyLocked(endpoint, abi.notify_exit_bit | t.id);
ipc.dropRef(endpoint);
}
}
/// Tear down the current user process and reschedule; never returns. Shared by the
/// exit system call and the fault path (`killCurrentProcess`). See
/// `releaseTaskResourcesLocked` for what is released, and in what order.
pub fn terminateCurrent() noreturn {
_ = sync.enter(); // handed off through the exit switch, released by the resumed task
terminateCurrentLocked();
}
/// The body of `terminateCurrent` for a caller that already holds the big kernel
/// lock — the scheduler's tick calls this (via `terminate_current_hook`) to finish
/// a deferred process_kill on its own core's current task. Never returns; the
/// tick's abandoned interrupt frame is fine (the LAPIC was acknowledged before the
/// tick hook ran), exactly as on the fault path.
fn terminateCurrentLocked() noreturn {
releaseTaskResourcesLocked(scheduler.current());
scheduler.exitUserLocked();
}
/// Reap a condemned task that is NOT running on any core (ready or blocked — and
/// it cannot start running: state changes need the lock we hold). The other half
/// of a deferred process_kill, called by the scheduler's tick (via
/// `reap_task_hook`) and directly by `killProcess` for targets caught off-CPU.
/// Precondition: the big kernel lock is held.
fn reapTaskLocked(t: *scheduler.Task) void {
releaseTaskResourcesLocked(t);
scheduler.removeFromReadyQueueLocked(t); // no-op unless it was ready in a queue
scheduler.destroyTaskLocked(t);
}
/// Kill process `target_id` on behalf of `caller_id` — the kernel half of the
/// process_kill system call. Returns 0, -ESRCH (no such live process — kernel
/// tasks are not killable processes and stale ids miss, since ids are never
/// reused), or -EPERM (the caller is not the target's supervisor).
///
/// A target that is ready or blocked is reaped on the spot. One that is running
/// on another core cannot be torn down mid-instruction, so it is condemned
/// (`kill_pending`) and dies at its next system_call entry, block, or timer tick
/// — like a Unix signal, delivery is prompt but asynchronous. Either way the
/// call returns 0: the kill is accepted and irrevocable.
pub fn killProcess(caller_id: u32, target_id: u32) i64 {
const flags = sync.enter();
defer sync.leave(flags);
irq.releaseOwner(t.id);
const target = scheduler.taskByIdLocked(target_id) orelse return -ipc.ESRCH;
if (target.aspace == 0) return -ipc.ESRCH; // kernel tasks are not processes
if (target.supervisor != caller_id) return -ipc.EPERM;
if (target.state == .running) {
target.kill_pending = true;
} else {
reapTaskLocked(target);
}
return 0;
}
/// Kill the current user process in response to a CPU fault it raised in ring 3.
/// The fault is confined to the process — the kernel trapped it on the task's own
/// kernel stack and is intact — so everything the process held is reclaimed and the
/// core reschedules. The system keeps running; only the faulting process dies
/// (docs/resilience.md: fault -> kill -> continue).
pub fn killCurrentProcess() noreturn {
fault_kill_count += 1;
terminateCurrent();
}
/// Resolve `(device_id, resource_index)` to a GSI this process is entitled to bind, or null.
@@ -352,6 +676,29 @@ fn systemIrqBind(state: *architecture.CpuState) void {
architecture.setSystemCallResult(state, 0);
}
/// msi_bind(device_id, endpoint_handle) -> address (rax), data (rdx): set up
/// Message-Signalled Interrupts for a claimed device. The kernel allocates a vector,
/// binds it to `endpoint`, and hands back the (address, data) the driver programs into
/// its own MSI capability (found via its ECAM config space, resource 0). Unlike
/// `irq_bind` there is no GSI, no sharing, and no ack — MSI is edge-triggered. The
/// claim is the capability. See docs/driver-model.md (M15).
fn systemMsiBind(state: *architecture.CpuState) void {
const device_id = architecture.systemCallArg(state, 0);
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const owner = devices_broker.ownerOf(device_id) orelse return fail(state);
if (owner != t.id) return fail(state); // not claimed by this process
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 1)) orelse return failErr(state, ipc.EBADF);
const flags = sync.enter();
defer sync.leave(flags);
const vector = irq.msiBind(endpoint, t.id) catch return fail(state);
// x86 MSI: the address routes to the BSP's LAPIC (physical destination, fixed
// delivery — dest field 0); the data carries the vector.
architecture.setSystemCallResult(state, abi.msi_address_base);
architecture.setSystemCallResult2(state, vector);
}
/// irq_ack(device_id, resource_index) -> 0/-1: re-arm a bound IRQ. The ISR left the line
/// masked (it could not quiet the device — that's this driver's job), so nothing
/// more arrives until the driver says it has serviced the hardware.
@@ -366,6 +713,11 @@ fn systemIrqAck(state: *architecture.CpuState) void {
if (irq.ack(gsi)) architecture.setSystemCallResult(state, 0) else fail(state);
}
/// Whether the debug_write stream sits at the start of a line — the last emitted
/// byte was a newline (true at boot: nothing emitted yet). Guarded by the kernel
/// lock in `systemDebugWrite`, like the stream it describes.
var write_at_line_start: bool = true;
/// debug_write(ptr, len): copy bytes from user memory into the kernel log.
/// A bring-up diagnostic — real output goes through the VFS/console later.
///
@@ -375,17 +727,26 @@ fn systemIrqAck(state: *architecture.CpuState) void {
/// Known gap (fine for trusted user code): a pointer into an *unmapped* hole in
/// the user half passes the check and the read #PFs -> on_fault halts — a
/// self-DoS, not an isolation break. Fault-recovering copy-in is a later item.
///
/// The emit runs under the kernel lock, so a message is atomic on the wire — two
/// processes writing from different cores can interleave *messages*, never bytes.
/// The "DANOS-INIT: " marker is emitted only at the start of a line (not per
/// call), so a process may assemble a line from several writes without the marker
/// (or, with the lock, another byte) landing in the middle. Cleanly-terminated
/// lines from concurrent writers stay whole either way.
fn systemDebugWrite(state: *architecture.CpuState) void {
const ptr = architecture.systemCallArg(state, 0);
const len = architecture.systemCallArg(state, 1);
if (len <= write_buffer.len and ptr < user_half_end and ptr + len <= user_half_end) {
const source: [*]const u8 = @ptrFromInt(ptr);
const flags = sync.enter();
defer sync.leave(flags);
@memcpy(write_buffer[0..len], source[0..len]); // keep the latest message
write_len = len;
write_from_user = architecture.fromUser(state);
write_count += 1;
log.write("DANOS-INIT: ");
log.write(source[0..len]);
if (len != 0) write_at_line_start = source[len - 1] == '\n';
architecture.setSystemCallResult(state, len);
} else {
fail(state);
@@ -480,15 +841,15 @@ pub fn run(blob: []const u8) RunError!void {
@memset(code[blob.len..page_size], 0xCC);
architecture.mapUserPage(code_virtual, code_frame, false, true); // RO + X
architecture.mapUserPage(stack_virtual, stack_frame, true, false); // RW + NX
architecture.mapUserPage(stack_base_virtual, stack_frame, true, false); // RW + NX (one page; the probe barely stacks)
resetRecords();
architecture.enterUser(scheduler.currentCpuIndex(), code_virtual, stack_virtual + page_size);
architecture.enterUser(scheduler.currentCpuIndex(), code_virtual, stack_base_virtual + page_size);
// Back via the exit system_call; the interrupt gate left IF clear.
architecture.enableInterrupts();
architecture.unmapPage(code_virtual);
architecture.unmapPage(stack_virtual);
architecture.unmapPage(stack_base_virtual);
pmm.free(code_frame);
pmm.free(stack_frame);
}
@@ -499,6 +860,7 @@ pub const InitError = error{
BadElf, // malformed/inapplicable image (magic, class, machine, type, bounds)
BadSegment, // PT_LOAD unaligned, out of the user region, W&X, or overlapping
BadEntry, // e_entry not inside an executable segment
BadArguments, // no argv[0], too many entries, or too many bytes for the entry stack
ProgramTooBig, // more pages than the loader's budget
OutOfMemory,
};
@@ -599,13 +961,84 @@ fn loadPageInto(aspace: u64, image: []const u8, seg: Segment, page_index: u64) I
architecture.mapUserPageInto(aspace, seg.vaddr + page_off, frame, seg.writable, seg.executable);
}
/// Build the System V AMD64 process-entry block at the top of a process's stack
/// page and return the initial user stack pointer. At the first user instruction,
/// rsp is 16-byte aligned and points at (addresses growing upward):
///
/// argc, argv[0..argc-1], NULL, NULL (empty envp), auxiliary vector, strings
///
/// — the layout every C runtime's startup code walks, so danos's own runtime and a
/// future libc port read arguments identically (docs/sysv.md). `page` is the kernel
/// (physmap) view of the stack's **top** frame and `page_user_base` that frame's
/// user address (stack_top_virtual - page_size); the pointers written into it are
/// user addresses inside that page. The caller has validated the sizes
/// (`entryStackBytes`), so this cannot overrun.
fn buildEntryStack(page: [*]u8, page_user_base: u64, argv: []const []const u8) u64 {
// The strings live at the very top of the page, packed from the end downward.
var string_offset: usize = page_size;
var pointers: [maximum_arguments]u64 = undefined;
var i: usize = argv.len;
while (i > 0) {
i -= 1;
string_offset -= argv[i].len + 1;
@memcpy(page[string_offset..][0..argv[i].len], argv[i]);
page[string_offset + argv[i].len] = 0; // NUL-terminated, as C expects
pointers[i] = page_user_base + string_offset;
}
// The vector sits below the strings: argc, the argv pointers, the argv
// terminator, an empty envp (terminator only), then the auxiliary vector.
const word_count = 1 + argv.len + 1 + 1 + 4;
const vector_offset = (string_offset - word_count * 8) & ~@as(usize, 15); // entry rsp % 16 == 0
const words: [*]u64 = @ptrCast(@alignCast(page + vector_offset));
var w: usize = 0;
words[w] = argv.len; // argc
w += 1;
for (pointers[0..argv.len]) |pointer| {
words[w] = pointer;
w += 1;
}
words[w] = 0; // argv terminator
words[w + 1] = 0; // envp: no environment yet, just the terminator
words[w + 2] = auxiliary_vector_page_size;
words[w + 3] = page_size;
words[w + 4] = auxiliary_vector_null; // end of the auxiliary vector
words[w + 5] = 0;
return page_user_base + vector_offset;
}
/// Bytes the entry block for `argv` occupies at the top of the stack page:
/// strings (each NUL-terminated), vector words, and the alignment slack.
fn entryStackBytes(argv: []const []const u8) usize {
var string_bytes: usize = 0;
for (argv) |argument| string_bytes += argument.len + 1;
return string_bytes + (1 + argv.len + 1 + 1 + 4) * 8 + 16;
}
/// Load a user ELF image into a fresh address space and spawn it as a scheduled
/// ring-3 process at `priority`. Returns immediately — the process runs
/// preemptively on its own page tables alongside everything else, and its exit
/// is handled by the system_call layer. The whole build (address space + ELF load +
/// task) runs under the kernel lock so it appears atomically and can't race
/// pmm/heap on another core.
pub fn spawnProcess(image: []const u8, priority: u3) InitError!void {
/// ring-3 process at `priority`, entered with `argv` on its stack per the System V
/// convention (`buildEntryStack`). `argv[0]` is required — it names the process:
/// the path or initial-ramdisk name it was spawned as. It is also recorded on the
/// task, so a fault report can say *which* binary died, not just its id.
/// The kernel-internal spawn (init at boot, tests): supervisor 0, no exit
/// notification. `spawnProcessSupervised` is the full form.
pub fn spawnProcess(image: []const u8, priority: u3, argv: []const []const u8) InitError!void {
_ = try spawnProcessSupervised(image, priority, argv, 0, null);
}
/// `spawnProcess`, recording `supervisor` (the id of the process that asked — the
/// kill authority) and, if given, `exit_endpoint` to notify when the child ends
/// (a reference is taken here and dropped when the notification posts).
/// Returns the child's process id.
/// Returns immediately — the process runs preemptively on its own page tables
/// alongside everything else, and its exit is handled by the system_call layer.
/// The whole build (address space + ELF load + task) runs under the kernel lock so
/// it appears atomically and can't race pmm/heap on another core.
pub fn spawnProcessSupervised(image: []const u8, priority: u3, argv: []const []const u8, supervisor: u32, exit_endpoint: ?*ipc.Endpoint) InitError!u32 {
if (argv.len == 0 or argv.len > maximum_arguments) return error.BadArguments;
// The entry block must leave most of the page as actual stack.
if (entryStackBytes(argv) > page_size / 2) return error.BadArguments;
var segs: [maximum_segments]Segment = undefined;
const parsed = try parseSegments(image, &segs);
@@ -618,9 +1051,35 @@ pub fn spawnProcess(image: []const u8, priority: u3) InitError!void {
for (segs[0..parsed.count]) |seg| {
for (0..seg.pages()) |i| try loadPageInto(aspace, image, seg, i);
}
const stack_frame = pmm.alloc() orelse return error.OutOfMemory;
architecture.mapUserPageInto(aspace, stack_virtual, stack_frame, true, false); // RW + NX
if (!scheduler.spawnUserLocked(aspace, parsed.entry, stack_virtual + page_size, priority))
// The stack: `user_stack_pages` zeroed pages below stack_top_virtual, RW + NX.
// The page below them (`stack_virtual`) stays unmapped as the overflow guard.
// The entry block goes at the top of the highest page.
var user_sp: u64 = 0;
for (0..parameters.user_stack_pages) |i| {
const stack_frame = pmm.alloc() orelse return error.OutOfMemory;
const stack_page: [*]u8 = @ptrFromInt(boot_handoff.physicalToVirtual(stack_frame));
@memset(stack_page[0..page_size], 0); // no stale frame contents leak into user space
const page_virtual = stack_base_virtual + i * page_size;
if (i == parameters.user_stack_pages - 1)
user_sp = buildEntryStack(stack_page, page_virtual, argv);
architecture.mapUserPageInto(aspace, page_virtual, stack_frame, true, false); // RW + NX
}
const child = scheduler.spawnUserLocked(aspace, parsed.entry, user_sp, priority, argv[0], supervisor, if (exit_endpoint) |endpoint| @ptrCast(endpoint) else null) orelse
return error.OutOfMemory;
// The child holds a reference to its exit endpoint from birth to death. Taken
// only now, after nothing can fail; the lock is still held, so the child
// cannot run (let alone die) before the reference exists.
if (exit_endpoint) |endpoint| endpoint.refcount += 1;
return child;
}
/// clock() -> nanoseconds since boot: a monotonic time source. The kernel already owns
/// the scheduling timer and computes this for preemption, so surfacing it is pure
/// mechanism — no policy (wall-clock time, calendars, timezones are a user-space
/// service layered on top). It lets a driver bound a poll loop by real time instead of
/// a spin count, and time short delays.
fn systemClock(state: *architecture.CpuState) void {
architecture.setSystemCallResult(state, architecture.nanos());
}
+227 -11
View File
@@ -18,6 +18,7 @@
//! shared queues.
const std = @import("std");
const abi = @import("abi");
const parameters = @import("parameters");
const architecture = @import("architecture");
const heap = @import("heap.zig");
@@ -41,6 +42,26 @@ pub const Task = struct {
kstack_top: usize = 0, // top of `stack` (== TSS.rsp0 for a user task); 0 = none
wake_at: u64 = 0, // uptime (ms) to wake a sleeping task; 0 = not sleeping
affinity: ?u32 = null, // null = runs on any core; else the index of its pinned core
// --- process management (process.zig) ---
// Id of the process that spawned this one (0 = the kernel). The supervision
// link is the kill authority: only the supervisor may process_kill a child.
supervisor: u32 = 0,
// Endpoint to notify when this process ends (any way: exit, fault, kill), or
// null. Holds its own reference, dropped when the notification is posted.
// Opaque here for the same reason as `handles` below.
exit_endpoint: ?*anyopaque = null,
// Set by process_kill on a task that is running on another core; the kernel
// finishes the kill at that task's next system call or timer tick.
kill_pending: bool = false,
// True while this task executes its own system call — the timer tick must not
// tear a task down in the middle of a kernel operation, only while it runs
// user code (or sits at a block point, where teardown is safe).
in_system_call: bool = false,
// Where this task is parked while blocked, so a kill can unlink it: the
// WaitQueue it waits on (maintained by waitLocked/wakeLocked), or the endpoint
// whose sender FIFO it queues in (maintained by the IPC layer; opaque here).
wait_queue: ?*WaitQueue = null,
ipc_wait_endpoint: ?*anyopaque = null,
// Physical root of this task's address space, or 0 for a kernel task (which
// runs on the shared kernel page tables). A user task carries its own.
aspace: u64 = 0,
@@ -66,11 +87,29 @@ pub const Task = struct {
ipc_reply_ptr: u64 = 0, // client: reply buffer (vaddr)
ipc_reply_cap: u64 = 0,
ipc_status: i64 = 0, // client: reply length / -errno, written by the replier
dma_map_next: u64 = 0, // bump pointer into this task's DMA arena (0 = unseeded)
ipc_send_cap: u64 = ~@as(u64, 0), // handle to transfer with this message (abi.no_cap = none)
ipc_received_cap: u64 = ~@as(u64, 0), // client: handle the reply's transferred cap landed at (abi.no_cap = none)
next: ?*Task = null, // ready-queue link (also the endpoint sender-FIFO link)
// The process's name — argv[0] as it was spawned (a boot-volume path for init,
// an initial-ramdisk name for everything else); empty for kernel tasks. Fixed
// storage, so the fault path can name the dead without touching the heap.
// Zero-initialised (not `undefined`): an undefined default is materialised as
// a 0xAA fill, which would move the whole static task pool out of .bss.
name_buffer: [maximum_task_name]u8 = .{0} ** maximum_task_name,
name_length: u8 = 0,
/// The task's name (argv[0] at spawn), or empty for a kernel task.
pub fn name(self: *const Task) []const u8 {
return self.name_buffer[0..self.name_length];
}
};
/// Capacity of `Task.name_buffer` — matches the longest name `system_spawn`
/// accepts, so a spawned name is never truncated. Shared with the ABI's
/// ProcessDescriptor, so `enumerate` copies names without clipping.
pub const maximum_task_name = abi.maximum_process_name;
/// Size of each task's IPC handle table. Kept here (not in ipc_sync.zig) because
/// it dimensions a field of `Task`; ipc_sync.zig re-exports it.
pub const ipc_maximum_handles = 16;
@@ -257,14 +296,19 @@ pub fn spawnOn(entry: *const fn () void, priority: Priority, cpu: u32) bool {
}
/// Spawn a **user** task: a task with its own address space (`aspace`) that starts
/// in user mode at `entry` on `user_sp`. It gets a fresh kernel stack for
/// syscalls/interrupts, and its first switch-in lands in `user_task_trampoline`.
/// Returns false (creating nothing) if the table is full or out of memory.
/// in user mode at `entry` on `user_sp`, recorded under `name` (its argv[0]).
/// `supervisor` is the id of the spawning process (0 = the kernel) — the kill
/// authority — and `exit_endpoint` (an *ipc.Endpoint whose reference the caller
/// has already taken, or null) is notified when this process ends.
/// It gets a fresh kernel stack for syscalls/interrupts, and its first switch-in
/// lands in `user_task_trampoline`.
/// Returns the new process id, or null (creating nothing) if the table is full or
/// out of memory.
/// **Caller must hold the kernel lock** (the loader that builds `aspace` holds it
/// across the whole spawn, so the address space and the task appear atomically).
pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, priority: Priority) bool {
const t = freeSlot() orelse return false;
const stack = heap.allocator().alloc(u8, stack_size) catch return false;
pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, priority: Priority, task_name: []const u8, supervisor: u32, exit_endpoint: ?*anyopaque) ?u32 {
const t = freeSlot() orelse return null;
const stack = heap.allocator().alloc(u8, stack_size) catch return null;
t.* = .{
.id = next_id,
.state = .ready,
@@ -273,7 +317,12 @@ pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, priority: Priority
.aspace = aspace,
.user_ip = entry,
.user_sp = user_sp,
.supervisor = supervisor,
.exit_endpoint = exit_endpoint,
};
const name_length = @min(task_name.len, maximum_task_name);
@memcpy(t.name_buffer[0..name_length], task_name[0..name_length]);
t.name_length = @intCast(name_length);
next_id += 1;
const top = @intFromPtr(stack.ptr) + stack.len;
t.kstack_top = top;
@@ -281,7 +330,7 @@ pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, priority: Priority
// the user entry/stack from the Task itself).
t.sp = architecture.initTaskStack(top, @intFromPtr(&startUserTask));
enqueue(t);
return true;
return t.id;
}
/// The first thing a fresh user task runs (in ring 0, via task_trampoline). It
@@ -393,6 +442,7 @@ pub const WaitQueue = struct {
pub fn waitLocked(wait_queue: *WaitQueue) void {
const t = current();
t.state = .blocked;
t.wait_queue = wait_queue; // so a kill can unlink a parked waiter
t.next = wait_queue.head;
wait_queue.head = t;
schedule();
@@ -417,10 +467,78 @@ pub fn wakeLocked(wait_queue: *WaitQueue) void {
}
const t = best orelse return;
if (best_previous) |p| p.next = t.next else wait_queue.head = t.next;
t.wait_queue = null;
t.state = .ready;
enqueue(t);
}
/// Unlink `t` from the wait queue it is parked on, if any (the kill path — a
/// killed waiter must not be woken later as a dangling pointer). Precondition:
/// the big kernel lock is held.
pub fn removeFromWaitQueueLocked(t: *Task) void {
const wait_queue = t.wait_queue orelse return;
t.wait_queue = null;
var previous: ?*Task = null;
var node = wait_queue.head;
while (node) |n| : ({
previous = n;
node = n.next;
}) {
if (n != t) continue;
if (previous) |p| p.next = t.next else wait_queue.head = t.next;
t.next = null;
return;
}
}
/// Unlink `t` from the ready queue it sits in (global, or its affinity core's
/// pinned queue) — the kill path for a task that is runnable but not running.
/// Precondition: the big kernel lock is held.
pub fn removeFromReadyQueueLocked(t: *Task) void {
if (t.affinity) |cpu| {
const pc = &cpus[cpu];
removeFrom(&pc.pinned_head, &pc.pinned_tail, &pc.pinned_bitmap, t);
} else {
removeFrom(&ready_head, &ready_tail, &ready_bitmap, t);
}
}
fn removeFrom(head: *[number_priorities]?*Task, tail: *[number_priorities]?*Task, bitmap: *u8, t: *Task) void {
const level: usize = t.priority;
var previous: ?*Task = null;
var node = head[level];
while (node) |n| : ({
previous = n;
node = n.next;
}) {
if (n != t) continue;
if (previous) |p| p.next = t.next else head[level] = t.next;
if (tail[level] == t) tail[level] = previous;
if (head[level] == null) bitmap.* &= ~(@as(u8, 1) << @intCast(level));
t.next = null;
return;
}
}
/// Find a live task by process id, or null. Ids are monotonic and never reused,
/// so a stale id misses cleanly rather than naming a recycled slot.
/// Precondition: the big kernel lock is held.
pub fn taskByIdLocked(id: u32) ?*Task {
for (&tasks) |*t| {
if (t.state != .free and t.id == id) return t;
}
return null;
}
/// Make every server that still holds `t` as the client it owes a reply to forget
/// it — the reply of a dead client is dropped, not delivered into freed state.
/// Precondition: the big kernel lock is held.
pub fn forgetIpcClientLocked(t: *Task) void {
for (&tasks) |*other| {
if (other.state != .free and other.ipc_client == t) other.ipc_client = null;
}
}
/// Block the current task and switch away, without putting it on any wait queue —
/// the caller has already linked it wherever it belongs (e.g. an endpoint's sender
/// FIFO). Precondition: the big kernel lock is held; still held on return (when the
@@ -480,14 +598,55 @@ fn wakeExpired() void {
}
}
// Process-teardown hooks, registered by process.zig at init — the scheduler sits
// below the process layer, so finishing a kill (IRQ bindings, IPC handles, exit
// notification) is called *up* through these, mirroring how the architecture
// layer calls up into `tick`.
//
// `terminate_current_hook` ends the task running on THIS core (lock held, never
// returns — it switches away like `exitUserLocked`). `reap_task_hook` tears down
// a task that is NOT running on any core (lock held).
pub var terminate_current_hook: ?*const fn () noreturn = null;
pub var reap_task_hook: ?*const fn (*Task) void = null;
/// Finish any pending kills this core can see (the deferred half of process_kill;
/// the immediate half runs in the killer's own call). Precondition: the big kernel
/// lock is held, from `tick`.
///
/// - This core's *current* task, if condemned, is terminated here — but only when
/// it is not inside one of its own system calls (`in_system_call`): the tick may
/// have interrupted kernel code mid-operation, where teardown would leak or
/// corrupt what that operation holds. User-mode execution (and the system_call
/// entry/exit stubs, which hold nothing) are safe termination points. A task
/// that *is* mid-call dies at its next block, tick, or system_call entry instead.
/// The hook never returns; abandoning the interrupt frame is fine — the LAPIC
/// was acknowledged before the tick hook ran (see apic.timerTick), exactly as on
/// the fault-kill path.
/// - Condemned tasks that are ready or blocked are not running anywhere (state
/// changes need the lock we hold), so they are reaped in place.
fn reapKillPendingLocked() void {
const pc = thisCpu();
const cur = pc.current;
if (cur.kill_pending and cur.aspace != 0 and !cur.in_system_call) {
if (terminate_current_hook) |hook| hook(); // noreturn
}
if (reap_task_hook) |hook| {
for (&tasks) |*t| {
if (!t.kill_pending) continue;
if (t.state == .ready or t.state == .blocked) hook(t);
}
}
}
/// Called from the timer interrupt (interrupts already disabled): wake due
/// sleepers, then preempt. Takes the kernel lock like any other critical section,
/// but releases it *without* touching the interrupt flag — the handler's `iretq`
/// restores the interrupted context's flags, so re-enabling here would open a
/// nested-interrupt window before the return.
/// sleepers, finish pending kills, then preempt. Takes the kernel lock like any
/// other critical section, but releases it *without* touching the interrupt flag
/// — the handler's `iretq` restores the interrupted context's flags, so
/// re-enabling here would open a nested-interrupt window before the return.
pub fn tick() void {
_ = sync.enter();
wakeExpired();
reapKillPendingLocked();
if (preemption_enabled) schedule();
sync.leaveIsr();
}
@@ -519,6 +678,13 @@ pub fn exit() noreturn {
/// itself is leaked, as in `exit` (no reaper yet). Never returns.
pub fn exitUser() noreturn {
_ = sync.enter();
exitUserLocked();
}
/// The body of `exitUser` for callers that already hold the big kernel lock (the
/// tick-time terminate path, which enters with the lock held). The lock is handed
/// off through the switch and released by the task that resumes. Never returns.
pub fn exitUserLocked() noreturn {
const pc = thisCpu();
const dying = pc.current;
const as = dying.aspace;
@@ -530,6 +696,8 @@ pub fn exitUser() noreturn {
}
dying.state = .free;
dying.aspace = 0;
dying.kill_pending = false;
dying.in_system_call = false;
const next = dequeueHighest(pc) orelse @panic("sched: no task left to run");
next.state = .running;
pc.current = next;
@@ -538,6 +706,54 @@ pub fn exitUser() noreturn {
unreachable;
}
/// Free a task that is NOT running on any core (it is ready or blocked, and the
/// caller — the kill path — has already unlinked it from every queue and released
/// what it held). Destroys its address space: safe here because no core can have
/// it loaded (every switch away from a task loads the next task's tables, and the
/// task isn't running). The kernel stack is leaked, as in `exitUser` (no reaper
/// yet). Precondition: the big kernel lock is held.
pub fn destroyTaskLocked(t: *Task) void {
if (t.aspace != 0) architecture.destroyAddressSpace(t.aspace);
t.aspace = 0;
t.kill_pending = false;
t.in_system_call = false;
t.wake_at = 0;
t.state = .free;
}
/// Snapshot the task table into `out` (up to its length), returning the total
/// number of live tasks — the kernel half of `process_enumerate`, mirroring
/// devices_broker.enumerate. Kernel tasks are included (empty name, supervisor 0):
/// an honest `ps` shows the idle tasks too. `out` may be user memory: the caller's
/// address space is loaded during its system call, and the same bring-up trust
/// applies as for device_enumerate (an unmapped user page faults the kernel).
pub fn enumerate(out: []abi.ProcessDescriptor) u64 {
const flags = sync.enter();
defer sync.leave(flags);
var total: u64 = 0;
for (&tasks) |*t| {
if (t.state == .free) continue;
if (total < out.len) {
const d = &out[total];
d.* = .{
.id = t.id,
.supervisor = t.supervisor,
.state = @intFromEnum(@as(abi.ProcessState, switch (t.state) {
.ready => .ready,
.running => .running,
.blocked => .blocked,
.free => unreachable,
})),
.priority = t.priority,
.name_length = t.name_length,
.name = t.name_buffer,
};
}
total += 1;
}
return total;
}
/// Whether the running task is a user process (has its own address space).
pub fn currentIsUserProcess() bool {
return current().aspace != 0;
+566 -15
View File
@@ -84,6 +84,16 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
ipcCallTest();
} else if (eql(case, "ipc-cap")) {
capabilityTest();
} else if (eql(case, "dma")) {
dmaTest();
} else if (eql(case, "msi")) {
msiTest();
} else if (eql(case, "iommu")) {
iommuTest();
} else if (eql(case, "ioport")) {
ioPortTest();
} else if (eql(case, "clock")) {
clockTest();
} else if (eql(case, "smp")) {
smpTest();
} else if (eql(case, "affinity")) {
@@ -108,14 +118,26 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
userMemTest();
} else if (eql(case, "user-pf")) {
userPfTest();
} else if (eql(case, "fault-recovery")) {
faultRecoveryTest(boot_information);
} else if (eql(case, "args")) {
argsTest(boot_information);
} else if (eql(case, "init")) {
initTest(boot_information);
} else if (eql(case, "process")) {
processTest(boot_information);
} else if (eql(case, "process-list")) {
processListTest(boot_information);
} else if (eql(case, "process-kill")) {
processKillTest(boot_information);
} else if (eql(case, "supervision")) {
supervisionTest(boot_information);
} else if (eql(case, "initial-ramdisk")) {
initialRamdiskTest(boot_information);
} else if (eql(case, "vfs")) {
vfsTest(boot_information);
} else if (eql(case, "input")) {
inputTest(boot_information);
} else if (eql(case, "hpet")) {
hpetTest(boot_information);
} else if (eql(case, "iopass")) {
@@ -232,6 +254,23 @@ fn discoveryTest() void {
check("AML parsed completely (consumed == total)", am.total > 0 and am.consumed == am.total);
check("at least one CPU enumerated (MADT)", platform.cpus().len >= 1);
// M15: every PCI function now carries its own 4 KiB ECAM configuration space as
// resource 0 — the window a driver mmio_maps to walk its capability list (MSI etc).
var buffer: [64]device_abi.DeviceDescriptor = undefined;
const n = @min(devices_broker.enumerate(&buffer), buffer.len);
var pci_functions: u32 = 0;
var pci_config_ok = true;
for (buffer[0..n]) |d| {
if (d.class != @intFromEnum(device_abi.DeviceClass.pci_device)) continue;
pci_functions += 1;
const has_config = d.resource_count >= 1 and
d.resources[0].kind == @intFromEnum(device_abi.ResourceKind.memory) and
d.resources[0].len == abi.page_size;
if (!has_config) pci_config_ok = false;
}
check("PCI functions were enumerated (MCFG/ECAM)", pci_functions >= 1);
check("each PCI function exposes its ECAM config space as resource 0", pci_config_ok);
result();
}
@@ -845,7 +884,7 @@ fn ipcClient() void {
/// prove the block/wake and reply-routing paths. (Kernel tasks, so no user ELF.)
fn ipcCallTest() void {
log("DANOS-TEST-BEGIN: ipc-call\n", .{});
ipc_endpoint = ipcsync.createEndpoint().?;
ipc_endpoint = ipcsync.createIpcEndpoint().?;
ipc_replies_ok = false;
ipc_done = false;
scheduler.spawn(ipcServer, 5); // above this task, so the workers run
@@ -873,7 +912,7 @@ var cap_done: bool = false;
/// verify it, then reply handing back a capability of its own.
fn capServer() void {
const me = scheduler.current();
cap_ep_y = ipcsync.createEndpoint().?;
cap_ep_y = ipcsync.createIpcEndpoint().?;
const h_y = ipcsync.installHandle(me, cap_ep_y); // the handle to send back in the reply
var reply_buffer: [8]u8 = .{0} ** 8;
@@ -896,7 +935,7 @@ fn capServer() void {
/// Client half of "open": mint a capability, send it in a call, receive one back.
fn capClient() void {
const me = scheduler.current();
cap_ep_x = ipcsync.createEndpoint().?;
cap_ep_x = ipcsync.createIpcEndpoint().?;
const h_x = ipcsync.installHandle(me, cap_ep_x);
var message: [8]u8 = .{0} ** 8;
@@ -918,7 +957,7 @@ fn capClient() void {
/// equal) and was *shared*, not moved (refcount bumped to 2).
fn capabilityTest() void {
log("DANOS-TEST-BEGIN: ipc-cap\n", .{});
cap_endpoint = ipcsync.createEndpoint().?;
cap_endpoint = ipcsync.createIpcEndpoint().?;
cap_server_got_x = false;
cap_client_got_y = false;
cap_done = false;
@@ -935,6 +974,170 @@ fn capabilityTest() void {
result();
}
/// DMA memory (M14): the properties a bus-mastering driver needs — physically
/// contiguous, a known physical address, correct cacheability, pinned, and reclaimed
/// on teardown. Exercises the kernel mechanism directly (`pmm.allocContiguous` +
/// `mapUserDmaInto`); the `dma_alloc`/`dma_free` syscalls are thin wrappers over it,
/// following the tested `mmap`/`mmio_map` shape, and land their first real use with the
/// first DMA driver.
fn dmaTest() void {
log("DANOS-TEST-BEGIN: dma\n", .{});
const base_free = pmm.stats().free_frames;
// A contiguous run: aligned, and it consumed exactly that many frames.
const frames = 4;
const phys = pmm.allocContiguous(frames, ~@as(u64, 0)) orelse {
check("allocContiguous(4) succeeded", false);
result();
return;
};
check("contiguous run is page-aligned", phys % abi.page_size == 0);
check("contiguous run consumed 4 frames", pmm.stats().free_frames == base_free - frames);
// The below-4G cap is honoured (legacy 32-bit DMA engines).
const low = pmm.allocContiguous(2, @as(u64, 4) << 30) orelse 0;
check("below-4G run stays under 4 GiB", low != 0 and low + 2 * abi.page_size <= (@as(u64, 4) << 30));
// Map the run into a fresh address space as coherent DMA and translate each page
// back: the same physical run, in order — proving contiguity and the mapping.
const aspace = architecture.createAddressSpace().?;
architecture.mapUserDmaInto(aspace, process.dma_arena_base, phys, frames * abi.page_size);
var mapped_ok = true;
for (0..frames) |i| {
const va = process.dma_arena_base + i * abi.page_size;
const got = architecture.translate(aspace, va) orelse {
mapped_ok = false;
break;
};
if (got != phys + i * abi.page_size) mapped_ok = false;
}
check("DMA pages translate to the contiguous physical run", mapped_ok);
// Teardown must reclaim the DMA RAM (the leaves carry no device_grant, so
// freeSubtree frees them as ordinary frames) — a driver that just dies leaks none.
architecture.destroyAddressSpace(aspace);
for (0..2) |i| pmm.free(low + i * abi.page_size);
check("no frames leaked after DMA teardown", pmm.stats().free_frames == base_free);
result();
}
/// MSI (M15): a per-device, edge-triggered interrupt vector delivered as an IPC
/// notification. QEMU's HPET has no MSI, so this exercises the vector-routing path with
/// a self-IPI standing in for the device's MSI memory write — proving the kernel
/// allocates a vector, `dispatch` recognises it as MSI (EOI + notify, no mask cycle),
/// and the bound endpoint is notified. The `msi_bind` syscall wraps `irq.msiBind` with
/// the device-claim check and lands its first real use with the first PCI driver.
fn msiTest() void {
log("DANOS-TEST-BEGIN: msi\n", .{});
const endpoint = ipcsync.createIpcEndpoint().?;
const flags = sync.enter();
const vector = irq.msiBind(endpoint, 0) catch {
sync.leave(flags);
check("msiBind allocated a vector", false);
result();
return;
};
sync.leave(flags);
check("msiBind allocated a vector in the device window", vector >= architecture.irq_vector_base and
vector < architecture.irq_vector_base + architecture.irq_vector_count);
// Fire the vector — the stand-in for the device writing its MSI (address, data).
const tail: *volatile u8 = &endpoint.notify_tail;
const before = tail.*;
architecture.selfIpi(vector);
var spins: u64 = 0;
while (tail.* == before and spins < 100_000_000) : (spins += 1) scheduler.yield();
check("self-IPI at the MSI vector notified the bound endpoint", tail.* != before);
result();
}
/// IOMMU (M16): with an emulated VT-d unit present (the harness boots this case with
/// `-device intel-iommu`), danos must find it in the ACPI DMAR table, map its register
/// block, and read back a real version. This is *detection*, the honest first step —
/// no translation domains are programmed yet, so DMA is still unprotected; enforcement
/// lands with the first DMA driver (docs/driver-model.md M16).
fn iommuTest() void {
log("DANOS-TEST-BEGIN: iommu\n", .{});
const pinfo = platform.platformInformation();
check("IOMMU found in the DMAR table", pinfo.iommu_present);
check("VT-d unit has a register base", pinfo.iommu_base != 0);
check("VT-d version register reads back nonzero (real, mappable unit)", pinfo.iommu_version != 0);
log("DANOS-IOMMU: base=0x{x} version=0x{x} capabilities=0x{x}\n", .{ pinfo.iommu_base, pinfo.iommu_version, pinfo.iommu_capabilities });
result();
}
/// Port I/O grants: ring 3 has no `in`/`out`, so a legacy driver reaches its ports
/// through `io_read`/`io_write`, gated by `device_claim` and the device's `io_port`
/// resource exactly like `mmio_map` gates memory. Target the PS/2 controller's status
/// port (0x64) — discovered on every PC and side-effect-free to read. Proves the
/// capability gate (`resolveIoPort` admits an in-range access, refuses out-of-range,
/// over-wide, and unclaimed) and that the kernel actually performs the `in`. The
/// `io_read`/`io_write` syscalls wrap this with the same ring-3 dispatch every device
/// driver already uses.
fn ioPortTest() void {
log("DANOS-TEST-BEGIN: ioport\n", .{});
var buffer: [64]device_abi.DeviceDescriptor = undefined;
const n = @min(devices_broker.enumerate(&buffer), buffer.len);
var found_id: ?u64 = null;
var found_res: u64 = 0;
outer: for (buffer[0..n]) |d| {
for (0..d.resource_count) |ri| {
const r = d.resources[ri];
if (r.kind == @intFromEnum(device_abi.ResourceKind.io_port) and r.start == 0x64 and r.len >= 1) {
found_id = d.id;
found_res = ri;
break :outer;
}
}
}
const id = found_id orelse {
check("discovered the PS/2 status port (io_port 0x64)", false);
result();
return;
};
check("discovered the PS/2 status port (io_port 0x64)", true);
const me = scheduler.current();
check("claimed the io_port device", devices_broker.claim(id, me.id));
check("an in-range access resolves to port 0x64", process.resolveIoPort(me, id, found_res, 0, 1) == 0x64);
check("an over-wide access is refused", process.resolveIoPort(me, id, found_res, 0, 2) == null);
check("an out-of-range offset is refused", process.resolveIoPort(me, id, found_res, 1, 1) == null);
check("an unclaimed device id is refused", process.resolveIoPort(me, 0xDEAD_BEEF, found_res, 0, 1) == null);
// The kernel actually issues the `in`. Reaching this line at all proves it didn't
// fault; a width-1 read must return a single byte.
const status = architecture.pioRead(1, 0x64);
check("reading the PS/2 status port returned a byte", status <= 0xFF);
log("DANOS-IOPORT: PS/2 status = 0x{x}\n", .{status});
result();
}
/// The monotonic clock (the source `clock()` surfaces to user space). It must be
/// calibrated, move forward over a spin, and never run backwards — the guarantees a
/// driver's deadline timeout depends on. The syscall is a thin wrapper over this same
/// `architecture.nanos()`.
fn clockTest() void {
log("DANOS-TEST-BEGIN: clock\n", .{});
const t0 = architecture.nanos();
check("monotonic clock is calibrated (nonzero)", t0 != 0);
var last = t0;
var monotonic = true;
var advanced = false;
var i: u32 = 0;
while (i < 1_000_000) : (i += 1) {
const t = architecture.nanos();
if (t < last) monotonic = false;
if (t > t0) advanced = true;
last = t;
}
check("clock advanced over the spin", advanced);
check("clock never ran backwards (monotonic)", monotonic);
result();
}
var proc_worker_run: bool = true;
var proc_worker_ran: bool = false;
@@ -969,8 +1172,8 @@ fn processTest(boot_information: *const BootInformation) void {
scheduler.spawn(procWorker, 4); // kernel task at the processes' priority
var spawned: u32 = 0;
if (process.spawnProcess(image, 4)) spawned += 1 else |_| {}
if (process.spawnProcess(image, 4)) spawned += 1 else |_| {}
if (process.spawnProcess(image, 4, &.{"/system/services/init"})) spawned += 1 else |_| {}
if (process.spawnProcess(image, 4, &.{"/system/services/init"})) spawned += 1 else |_| {}
// Wait (real time) for several heartbeats across the two processes. Each
// process sleeps ~1 s between beats, so a few seconds yields several.
@@ -999,6 +1202,87 @@ fn userPfTest() void {
log("DANOS-TEST-RESULT: FAIL (user read of kernel memory did not fault)\n", .{});
}
/// Spawn a real scheduled ring-3 process whose body is the user-pf blob (a read of
/// a kernel-only page, then a spin), so its first instruction raises #PF. Built by
/// hand — address space, code page RO+X, stack page RW+NX — because the blob is a
/// raw code fragment, not an ELF `spawnProcess` could load. Returns false if any
/// allocation fails.
fn spawnFaultingProcess() bool {
const blob = process.pfBlob();
const flags = sync.enter();
defer sync.leave(flags);
const aspace = architecture.createAddressSpace() orelse return false;
const code_frame = pmm.alloc() orelse {
architecture.destroyAddressSpace(aspace);
return false;
};
// Fill through the physmap (the user mapping is read-only); pad with int3 so a
// stray jump traps instead of sliding.
const code: [*]u8 = @ptrFromInt(boot_handoff.physicalToVirtual(code_frame));
@memset(code[0..abi.page_size], 0xCC);
@memcpy(code[0..blob.len], blob);
architecture.mapUserPageInto(aspace, process.code_virtual, code_frame, false, true); // RO + X
const stack_frame = pmm.alloc() orelse {
architecture.destroyAddressSpace(aspace); // frees code_frame too — it's mapped
return false;
};
architecture.mapUserPageInto(aspace, process.stack_base_virtual, stack_frame, true, false); // RW + NX
if (scheduler.spawnUserLocked(aspace, process.code_virtual, process.stack_base_virtual + abi.page_size, 4, "fault-probe", 0, null) == null) {
architecture.destroyAddressSpace(aspace);
return false;
}
return true;
}
/// Fault recovery (docs/resilience.md step 2): a scheduled ring-3 process that
/// faults must be killed — counted, resources reclaimed — while the rest of the
/// system keeps running. init heartbeats before and after the kill are the proof
/// the OS survived; the old behaviour (halt the core) would freeze the beat and
/// time the harness out.
fn faultRecoveryTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: fault-recovery\n", .{});
check("bootloader handed over /system/services/init", boot_information.init_len != 0);
if (boot_information.init_len == 0) {
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len];
process.write_count = 0;
process.fault_kill_count = 0;
const spawned = if (process.spawnProcess(image, 4, &.{"/system/services/init"})) true else |_| false;
check("init spawned as the surviving process", spawned);
// A first heartbeat proves init runs before the fault.
scheduler.setPriority(1); // drop below the processes so they get the core
var deadline = architecture.millis() + 8000;
while (process.write_count < 1 and architecture.millis() < deadline) scheduler.yield();
scheduler.setPriority(4);
check("init heartbeat before the fault", process.write_count >= 1);
check("faulting process spawned", spawnFaultingProcess());
// The kill: the faulting process #PFs on its first instruction and the kernel
// reaps it instead of halting.
scheduler.setPriority(1);
deadline = architecture.millis() + 5000;
while (process.fault_kill_count < 1 and architecture.millis() < deadline) scheduler.yield();
scheduler.setPriority(4);
check("faulting process was killed (not the machine)", process.fault_kill_count == 1);
// Life after the kill: init must keep beating on the same core.
const beats_at_kill = process.write_count;
scheduler.setPriority(1);
deadline = architecture.millis() + 8000;
while (process.write_count < beats_at_kill + 2 and architecture.millis() < deadline) scheduler.yield();
scheduler.setPriority(4);
check("init kept heartbeating after the kill", process.write_count >= beats_at_kill + 2);
result();
}
/// The full PID-1 path: the bootloader read /system/services/init off the boot volume and
/// handed it over; load it as a user ELF and spawn it as a real ring-3 process
/// — the same call the normal boot path makes — then confirm it beats. init
@@ -1013,7 +1297,7 @@ fn initTest(boot_information: *const BootInformation) void {
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len];
process.write_count = 0;
const spawned = if (process.spawnProcess(image, 4)) true else |err| blk: {
const spawned = if (process.spawnProcess(image, 4, &.{"/system/services/init"})) true else |err| blk: {
log("DANOS-INIT-ERR: {s}\n", .{@errorName(err)});
break :blk false;
};
@@ -1035,6 +1319,189 @@ fn initTest(boot_information: *const BootInformation) void {
result();
}
/// process_enumerate's kernel half: spawn two init processes next to the kernel
/// tasks and snapshot the table. The snapshot must list both by name with distinct,
/// kernel-supervised ids, include the kernel tasks (id 0, empty name), and report
/// the same total through a too-small buffer (the truncation contract: the caller
/// learns how big a buffer to bring).
fn processListTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: process-list\n", .{});
check("bootloader handed over /system/services/init", boot_information.init_len != 0);
if (boot_information.init_len == 0) {
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len];
var spawned: u32 = 0;
if (process.spawnProcess(image, 4, &.{"/system/services/init"})) spawned += 1 else |_| {}
if (process.spawnProcess(image, 4, &.{"/system/services/init"})) spawned += 1 else |_| {}
check("two init processes spawned", spawned == 2);
var table: [32]abi.ProcessDescriptor = undefined;
const total = scheduler.enumerate(&table);
check("enumerate counts the boot task and both processes (>=3)", total >= 3);
var inits: u32 = 0;
var init_ids: [2]u32 = .{ 0, 0 };
var kernel_task_seen = false;
var states_sane = true;
for (table[0..@min(total, table.len)]) |descriptor| {
if (descriptor.state > @intFromEnum(abi.ProcessState.blocked)) states_sane = false;
if (descriptor.name_length == 0) kernel_task_seen = true;
if (eql(descriptor.name[0..descriptor.name_length], "/system/services/init")) {
if (inits < 2) init_ids[inits] = descriptor.id;
inits += 1;
check("init entry is kernel-supervised (supervisor 0)", descriptor.supervisor == 0);
}
}
check("both init processes listed by name", inits == 2);
check("listed processes carry distinct ids", init_ids[0] != init_ids[1]);
check("kernel tasks are listed too (empty name)", kernel_task_seen);
check("every state is a ProcessState value", states_sane);
var one: [1]abi.ProcessDescriptor = undefined;
check("a too-small buffer still learns the true total", scheduler.enumerate(&one) == total);
result();
}
/// process_kill + the exit notification, kernel half. Two victims, two paths:
/// - init, which heartbeats and sleeps: caught blocked, reaped on the killer's
/// own call — and its heartbeat must stop.
/// - process-test's spinner role (from the initial ramdisk), which loops in user
/// mode making no system calls: with more cores it is caught running, taking
/// the deferred path (kill_pending, finished by the victim core's next tick).
/// Each death must post one exit notification badge (exit bit + the child's id)
/// on the endpoint given at spawn; wrong-supervisor and unknown-id kills must be
/// refused. The waits block in replyWait, so a lost notification times the
/// harness out rather than passing vacuously.
fn processKillTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: process-kill\n", .{});
check("bootloader handed over /system/services/init", boot_information.init_len != 0);
if (boot_information.init_len == 0 or boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", boot_information.initial_ramdisk_len != 0);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len];
const ramdisk = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(ramdisk) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
const me = scheduler.currentId();
const endpoint = ipcsync.createIpcEndpoint() orelse {
check("exit endpoint allocated", false);
result();
return;
};
process.write_count = 0;
const sleeper = process.spawnProcessSupervised(image, 4, &.{"/system/services/init"}, me, endpoint) catch 0;
check("init spawned as the supervised sleeper victim", sleeper != 0);
// Let it reach its heartbeat loop (write, then a 1 s sleep) so the kill most
// likely catches it blocked.
scheduler.setPriority(1);
const deadline = architecture.millis() + 8000;
while (process.write_count < 1 and architecture.millis() < deadline) scheduler.yield();
scheduler.setPriority(4);
check("victim heartbeat before the kill", process.write_count >= 1);
// Kills that must be refused, before the one that must not be.
check("a non-supervisor may not kill (-EPERM)", process.killProcess(me + 12345, sleeper) == -ipcsync.EPERM);
check("an unknown id misses (-ESRCH)", process.killProcess(me, 0xFFFF_FF00) == -ipcsync.ESRCH);
check("a kernel task is not a killable process (-ESRCH)", process.killProcess(me, 0) == -ipcsync.ESRCH);
check("the supervisor's kill is accepted", process.killProcess(me, sleeper) == 0);
var badge: u64 = 0;
var received_cap: u64 = 0;
var r = ipcsync.replyWait(endpoint, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap);
check("the sleeper's exit notification arrived (length 0)", r == 0);
check("its badge carries the exit bit and the child id", badge == abi.notify_badge_bit | abi.notify_exit_bit | sleeper);
const beats_at_kill = process.write_count;
scheduler.sleep(1500); // more than one heartbeat period
check("the heartbeat stopped with the kill", process.write_count == beats_at_kill);
// The spinner: no system calls, so only the tick can deliver a deferred kill.
var spinner: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "process-test")) continue;
spinner = process.spawnProcessSupervised(item.blob, 4, &.{ "process-test", "spinner" }, me, endpoint) catch 0;
break;
}
check("process-test spawned as the supervised spinner victim", spinner != 0);
scheduler.sleep(100); // give another core a chance to be running it
check("the spinner's kill is accepted", process.killProcess(me, spinner) == 0);
r = ipcsync.replyWait(endpoint, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap);
check("the spinner's exit notification arrived (length 0)", r == 0);
check("its badge carries the exit bit and the child id", badge == abi.notify_badge_bit | abi.notify_exit_bit | spinner);
var table: [32]abi.ProcessDescriptor = undefined;
const total = scheduler.enumerate(&table);
var still_listed = false;
for (table[0..@min(total, table.len)]) |descriptor| {
if (descriptor.id == sleeper or descriptor.id == spinner) still_listed = true;
}
check("neither victim is listed after its kill", !still_listed);
check("a killed id stays dead (-ESRCH on a second kill)", process.killProcess(me, sleeper) == -ipcsync.ESRCH);
result();
}
/// The whole user-side surface at once: spawn process-test's supervisor role,
/// which — entirely from ring 3 — creates an exit endpoint, spawns its two
/// children supervised, sees them in process_enumerate, kills them (one blocked,
/// one spinning), collects both exit notifications, and confirms they are gone.
/// Its "process-test: ok" is the pass marker; any FAIL line is specific.
fn supervisionTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: supervision\n", .{});
check("bootloader handed over an initial_ramdisk", boot_information.initial_ramdisk_len != 0);
if (boot_information.initial_ramdisk_len == 0) {
result();
return;
}
const ramdisk = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(ramdisk) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(ramdisk); // the supervisor system_spawns its children by name
process.write_count = 0;
process.write_from_user = false;
var started = false;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "process-test")) continue;
started = if (process.spawnProcess(item.blob, 4, &.{ "process-test", "run" })) true else |_| false;
break;
}
check("process-test spawned as the user-space supervisor", started);
const marker = "process-test: ok";
scheduler.setPriority(1);
const deadline = architecture.millis() + 10000;
while (architecture.millis() < deadline) {
if (process.write_len >= marker.len and eql(process.write_buffer[0..marker.len], marker)) break;
scheduler.yield();
}
scheduler.setPriority(4);
const ok = process.write_len >= marker.len and eql(process.write_buffer[0..marker.len], marker);
if (!ok and process.write_len > 0) log("DANOS-SUPERVISION: got \"{s}\"\n", .{process.write_buffer[0..process.write_len]});
check("the supervisor completed every step (spawn/list/kill/notify)", ok);
check("it ran in user mode (CPL 3)", process.write_from_user);
result();
}
/// The initial_ramdisk path: the bootloader handed over an image bundling extra user
/// binaries; parse it, spawn every program, and confirm one (the vfs stub)
/// reaches ring 3 and heartbeats — proving the whole ferry-parse-spawn pipeline.
@@ -1060,7 +1527,7 @@ fn initialRamdiskTest(boot_information: *const BootInformation) void {
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (process.spawnProcess(item.blob, 4)) spawned += 1 else |err| {
if (process.spawnProcess(item.blob, 4, &.{item.name})) spawned += 1 else |err| {
log("DANOS-INITRD-ERR: {s}: {s}\n", .{ item.name, @errorName(err) });
}
}
@@ -1120,6 +1587,89 @@ fn vfsTest(boot_information: *const BootInformation) void {
result();
}
/// The full input path: spawn the input service, a synthetic source, and a subscriber from
/// the initial_ramdisk. The source publishes keyboard, mouse, and joystick events in turn;
/// the service routes them (with the asynchronous ipc_send) to the subscriber, which took
/// all three classes and — only once it has received one — heartbeats "input-test: ok".
/// Seeing that marker proves an event travelled source -> service -> subscriber over IPC,
/// exercising the async buffered-send primitive, capability-passing subscription, and
/// per-device routing. The source and service stay silent after startup so the subscriber's
/// line is the one left in the shared evidence buffer.
fn inputTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: input\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.write_count = 0;
process.write_from_user = false;
_ = spawnNamed(rd, "input"); // the fan-out service
_ = spawnNamed(rd, "input-source"); // a synthetic keyboard publishing events
_ = spawnNamed(rd, "input-test"); // the subscriber whose "ok" line is the marker
// Wait for the subscriber's success heartbeat (it beats once per received event).
const prefix = "input-test: ok";
scheduler.setPriority(1);
const deadline = architecture.millis() + 12000;
while (architecture.millis() < deadline) {
if (process.write_len >= prefix.len and eql(process.write_buffer[0..prefix.len], prefix) and process.write_count >= 2) break;
scheduler.yield();
}
scheduler.setPriority(4);
const ok = process.write_len >= prefix.len and eql(process.write_buffer[0..prefix.len], prefix);
check("a subscriber received a broadcast key event over IPC (source -> service -> subscriber)", ok);
check("events kept flowing (service + async send stay up)", process.write_count >= 2);
check("client syscalls came from user mode (CPL 3)", process.write_from_user);
result();
}
/// Process arguments, end to end: spawn args-echo bare (its argv[0] is the
/// initial-ramdisk name). Instance 1 sees argc == 1 and respawns itself through
/// `system_spawn` with the extra arguments "alpha beta-42" — the syscall argument
/// blob. Instance 2 parses the kernel-built System V entry stack via the runtime
/// and echoes its whole argv in one write, which must arrive exactly as sent.
fn argsTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: args\n", .{});
check("bootloader handed over an initial_ramdisk", boot_information.initial_ramdisk_len != 0);
if (boot_information.initial_ramdisk_len == 0) {
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(image); // args-echo respawns itself through system_spawn
process.write_count = 0;
process.write_from_user = false;
check("args-echo spawned from the initial_ramdisk", spawnNamed(rd, "args-echo"));
// Wait for the *second* instance's echo (the first writes nothing).
scheduler.setPriority(1);
const deadline = architecture.millis() + 8000;
while (process.write_count < 1 and architecture.millis() < deadline) scheduler.yield();
scheduler.setPriority(4);
const expected = "args: args-echo alpha beta-42\n";
const echoed = process.write_len == expected.len and eql(process.write_buffer[0..process.write_len], expected);
if (!echoed and process.write_len > 0) log("DANOS-ARGS: got \"{s}\"\n", .{process.write_buffer[0..process.write_len]});
check("argv arrived intact (argv[0] = name, argv[1..] = spawn arguments)", echoed);
check("echo came from user mode (CPL 3)", process.write_from_user);
result();
}
/// Spawn the initial_ramdisk binary named `name` as a ring-3 process. Returns false if it
/// isn't in the image or fails to load.
fn spawnNamed(rd: initial_ramdisk.Reader, name: []const u8) bool {
@@ -1127,7 +1677,7 @@ fn spawnNamed(rd: initial_ramdisk.Reader, name: []const u8) bool {
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (eql(item.name, name)) {
return if (process.spawnProcess(item.blob, 4)) true else |_| false;
return if (process.spawnProcess(item.blob, 4, &.{item.name})) true else |_| false;
}
}
return false;
@@ -1376,7 +1926,7 @@ fn irqFreeTest() void {
return;
};
const endpoint = ipcsync.createEndpoint() orelse {
const endpoint = ipcsync.createIpcEndpoint() orelse {
check("allocated an endpoint", false);
result();
return;
@@ -1473,7 +2023,10 @@ fn faultNull() void {
// access). `allowzero` skips the same null check on the cast. The write then
// hits the unmapped page 0 and takes a real hardware #PF.
var address: u64 = 0;
address = asm ("" : [ret] "=r" (-> u64) : [in] "0" (address));
address = asm (""
: [ret] "=r" (-> u64),
: [in] "0" (address),
);
const p: *allowzero volatile u64 = @ptrFromInt(address);
p.* = 1;
}
@@ -1499,8 +2052,7 @@ fn faultDoubleFault() void {
\\ud2
:
: [sp] "r" (bad_sp),
: .{ .memory = true }
);
: .{ .memory = true });
bad_sp += 0;
}
@@ -1519,8 +2071,7 @@ fn apDoubleFaultTask() void {
\\ud2
:
: [sp] "r" (bad_sp),
: .{ .memory = true }
);
: .{ .memory = true });
bad_sp += 0;
}
+7 -1
View File
@@ -4,7 +4,7 @@
//! hiding the trade-offs. Keeping them here makes them visible at a glance and gives
//! one spot to change them. They're plain `comptime` constants (zero runtime cost);
//! any one can later be promoted to a `-D` build option if a target needs to vary it
//! (see build.zig's `-Dtest-case` for the pattern). This keeps root.zig to what it
//! (see build.zig's `-Dtest-case` for the pattern). This keeps ps2-library.zig to what it
//! actually is — the bootloader↔kernel handoff *contract* — with tunables living here.
/// Ceiling on logical CPUs the kernel tracks — the size of the per-CPU bookkeeping
@@ -22,6 +22,12 @@ pub const maximum_tasks = 16;
/// Each task's kernel stack (also each AP's bring-up stack), in bytes.
pub const kernel_stack_size = 16 * 1024;
/// Each user process's stack, in pages (32 KiB). Mapped just below a fixed top;
/// the System V entry block (argc/argv) occupies the top of the highest page, and
/// the page below the mapping is left unmapped as a guard, so an overflow faults
/// (killing only that process) instead of silently corrupting the image.
pub const user_stack_pages = 8;
/// Each core's IST (double-fault) stack, in bytes. The BSP's is static; an AP's is
/// heap-allocated at bring-up.
pub const ist_stack_size = 16 * 1024;
+55
View File
@@ -0,0 +1,55 @@
//! args-echo — a test fixture for process arguments (bundled in the
//! initial-ramdisk, spawned only by the `args` test case). Run with no arguments,
//! it respawns itself *with* some via `spawnWithArguments` — exercising the
//! system_spawn argument blob. Run with arguments, it burns more stack than one
//! page could hold (proving the multi-page stack: on a single-page stack the
//! recursion would hit the guard and the process would be killed before echoing),
//! then echoes its whole argv in one `debug_write` the kernel test asserts on —
//! proving the kernel-built System V entry stack (argc, argv pointers,
//! NUL-terminated strings) and the runtime's parsing of it, end to end.
const runtime = @import("runtime");
/// Recurse with a real frame each level: `depth` levels of ~0.5 KiB, touched
/// through a volatile pointer so no optimiser can flatten the frames away.
fn burnStack(depth: usize) u8 {
var frame: [512]u8 = undefined;
const touch: *volatile [512]u8 = &frame;
touch[0] = @truncate(depth);
touch[511] = touch[0];
if (depth == 0) return touch[511];
return touch[0] +% burnStack(depth - 1);
}
pub fn main(init: runtime.process.Init) void {
if (init.arguments.count <= 1) {
// First instance: spawn the second with real arguments, then exit.
_ = runtime.system.spawnWithArguments("args-echo", &.{ "alpha", "beta-42" });
return;
}
// ~16 x 0.5 KiB frames: comfortably past one page, well inside the 32 KiB stack.
_ = burnStack(16);
// Second instance: echo "args: <argv0> <argv1> ..." for the test to match.
var buffer: [128]u8 = undefined;
const prefix = "args:";
@memcpy(buffer[0..prefix.len], prefix);
var len: usize = prefix.len;
var iterator = init.arguments.iterate();
while (iterator.next()) |argument| {
if (len + 1 + argument.len + 1 > buffer.len) break;
buffer[len] = ' ';
len += 1;
@memcpy(buffer[len..][0..argument.len], argument);
len += argument.len;
}
buffer[len] = '\n';
len += 1;
_ = runtime.system.write(buffer[0..len]);
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
@@ -12,17 +12,64 @@
//! HPET, decides `hpet` serves it, and brings that driver all the way up. (The
//! kernel still auto-spawns the whole initial-ramdisk at boot; increment 3 removes
//! that redundancy so the manager is the sole owner of driver spawning.)
const std = @import("std");
const runtime = @import("runtime");
const acpi_ids = @import("acpi-ids");
const device = runtime.device;
const system = runtime.system;
/// The driver that serves each device class — the policy table. In a fuller system
/// Format one whole log line and emit it in a single `debug_write`, so output
/// from the drivers this manager starts (which run concurrently) can never land
/// in the middle of it.
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
/// The driver that serves each device — the policy table. In a fuller system
/// this comes from the drivers describing what they bind (or a manifest under
/// /system/drivers); for now it is a small static map, which is enough to prove the
/// manager reads the tree and decides. `null` = no driver for this class yet.
fn driverFor(class: u64) ?[]const u8 {
if (class == @intFromEnum(device.DeviceClass.timer)) return "hpet"; // the HPET
return null;
fn driverFor(d: device.DeviceDescriptor) ?[]const u8 {
// detect device via DeviceClass
if (d.class == @intFromEnum(device.DeviceClass.timer)) return "hpet";
// detect device via hid
const hid = d.hid[0..@intCast(d.hid_len)];
const id = acpi_ids.HardwareId.fromHid(hid) orelse return null;
return switch (id) {
.ps2_keyboard, .ps2_mouse => "ps2-bus",
else => null,
};
}
/// The PCI class/subclass/prog-IF triple of an xHCI (USB 3) host controller:
/// Serial Bus Controller (0x0C) / USB Controller (0x03) / XHCI (0x30) — the names
/// pci-class.zig decodes.
const xhci_pci_class: u64 = 0x0C_03_30;
/// The bus driver that serves a PCI function, or null. Unlike the singleton drivers
/// in `driverFor`, a machine can carry several identical controllers — so the caller
/// spawns one driver instance *per device*, passing the device id as argv[1] for the
/// instance to claim.
fn pciDriverFor(d: device.DeviceDescriptor) ?[]const u8 {
if (d.class != @intFromEnum(device.DeviceClass.pci_device)) return null;
return switch (d.pci_class) {
xhci_pci_class => "usb-xhci-bus",
else => null,
};
}
/// Spawn one instance of `driver_name` to serve the specific device `id` — the id
/// arrives as argv[1]. No isProcessRunning gate here: the name alone cannot tell two
/// instances apart, and this manager is the sole spawner of drivers.
fn spawnForDevice(driver_name: []const u8, id: u64) void {
var text: [20]u8 = undefined;
const id_text = std.fmt.bufPrint(&text, "{d}", .{id}) catch return;
if (system.spawnWithArguments(driver_name, &.{id_text}) != null) {
writeLine("device-manager: spawned {s} for device {d}\n", .{ driver_name, id });
} else {
writeLine("device-manager: failed to spawn {s} for device {d}\n", .{ driver_name, id });
}
}
pub fn main() void {
@@ -36,16 +83,21 @@ pub fn main() void {
var matched: usize = 0;
for (buffer[0..count]) |descriptor| {
const driver_name = driverFor(descriptor.class) orelse continue;
if (pciDriverFor(descriptor)) |driver_name| {
matched += 1;
spawnForDevice(driver_name, descriptor.id);
continue;
}
const driver_name = driverFor(descriptor) orelse continue;
matched += 1;
if (runtime.system.spawn(driver_name)) {
_ = runtime.system.write("device-manager: spawned ");
_ = runtime.system.write(driver_name);
_ = runtime.system.write("\n");
if (!system.isProcessRunning(driver_name)) {
if (runtime.system.spawn(driver_name) != null) {
writeLine("device-manager: spawned {s}\n", .{driver_name});
} else {
writeLine("device-manager: failed to spawn {s}\n", .{driver_name});
}
} else {
_ = runtime.system.write("device-manager: failed to spawn ");
_ = runtime.system.write(driver_name);
_ = runtime.system.write("\n");
writeLine("device-manager: already spawned {s}\n", .{driver_name});
}
}
+1 -1
View File
@@ -17,7 +17,7 @@ const runtime = @import("runtime");
/// microkernel keeps such choices in user space, not the kernel. Drivers are absent
/// on purpose: the device manager owns those. (A future init reads this from a
/// manifest under /system/services instead of a hardcoded list.)
const boot_services = [_][]const u8{ "vfs", "device-manager" };
const boot_services = [_][]const u8{ "vfs", "input", "device-manager" };
pub fn main() void {
// Prove the heap end to end: allocate through the runtime allocator (which
@@ -0,0 +1,40 @@
//! system/services/input-source — a hardware-free synthetic input source, used to exercise
//! the input service end to end without a real PS/2 controller (the `input` test case, and
//! any bring-up where there is no hardware). It stands in for a driver: it connects to the
//! input service and publishes a rolling stream that cycles through all device classes —
//! keyboard, mouse, and joystick/gamepad — which the service routes to interested
//! subscribers.
//!
//! It stays silent after startup (no per-event logging) so it can share the boot serial
//! transcript with a subscriber whose output is the test's success marker. The real
//! keyboard and mouse drivers publish their own synthetic streams today; swapping in
//! decoded hardware is a follow-up (see docs/input.md).
const runtime = @import("runtime");
const input = runtime.input;
const system = runtime.system;
pub fn main() void {
var source = input.connectSource() orelse {
_ = system.write("input-source: input service unavailable\n");
return;
};
_ = system.write("input-source: publishing synthetic input events\n");
var step: usize = 0;
while (true) : (step +%= 1) {
// Rotate across the device classes so every publish path (and the service's
// per-device routing) is exercised.
switch (step % 3) {
0 => _ = source.publishKeyboardEvent(input.syntheticKeyEvent(step)),
1 => _ = source.publishMouseEvent(input.syntheticMouseEvent(step)),
else => _ = source.publishJoystickEvent(input.syntheticJoystickEvent(step)),
}
system.sleep(200);
}
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
+56
View File
@@ -0,0 +1,56 @@
//! system/services/input-test — the input service's client and test oracle, the input
//! counterpart of vfs-test. It subscribes to *all* device classes and loops receiving the
//! events a source broadcasts, logging each with its class. It emits the success marker
//! `"input-test: ok"` only **after it has received at least one of each class** (keyboard,
//! mouse, and joystick), then heartbeats it. So the in-kernel `input` test case seeing that
//! marker proves not just that IPC delivery works but that the service *routed* all three
//! device classes to one subscription — source -> service -> subscriber, per device.
const std = @import("std");
const runtime = @import("runtime");
const input = runtime.input;
const system = runtime.system;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
pub fn main() void {
var listener = input.subscribeAll() orelse {
_ = system.write("input-test: could not subscribe\n");
return;
};
_ = system.write("input-test: subscribed\n");
var seen_keyboard = false;
var seen_mouse = false;
var seen_joystick = false;
while (true) {
const event = listener.next() orelse continue;
// Decode the class-specific payload from the tagged envelope and note the class.
if (event.asKeyboard()) |key| {
seen_keyboard = true;
writeLine("input-test: got keyboard code={d} char={d}\n", .{ key.keycode, key.character });
} else if (event.asMouse()) |mouse| {
seen_mouse = true;
writeLine("input-test: got mouse dx={d} dy={d} buttons={d}\n", .{ mouse.dx, mouse.dy, mouse.buttons });
} else if (event.asJoystick()) |joystick| {
seen_joystick = true;
writeLine("input-test: got joystick control={d} value={d}\n", .{ joystick.control, joystick.value });
} else {
writeLine("input-test: got device={d}\n", .{event.device});
}
// The success marker: only once every class has been routed here does this appear,
// and then it heartbeats. Seeing "input-test: ok" proves per-device fan-out works.
if (seen_keyboard and seen_mouse and seen_joystick) {
_ = system.write("input-test: ok all classes received (keyboard, mouse, joystick)\n");
}
}
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
+145
View File
@@ -0,0 +1,145 @@
//! system/services/input — the user-space input service. Shipped in the initial_ramdisk,
//! spawned as a ring-3 process, and published under the well-known `input` service id. It
//! is the fan-out point between **sources** (keyboard, mouse, and joystick/gamepad drivers)
//! and **subscribers** (any program that wants input): a source `publish`es an
//! `InputEvent`, and the service pushes it to every subscriber whose interest mask includes
//! that event's device class (keyboard / mouse / joystick).
//!
//! The delivery discipline is the whole design (see docs/input.md). The kernel's IPC is a
//! synchronous rendezvous: a server holds one pending reply, so it cannot park N
//! subscribers blocked in a "wait for next event" call. Broadcasting therefore has to be
//! *push* — the service delivering to subscribers. But a synchronous push (`ipc_call`)
//! would let one dead or wedged subscriber hang the whole broadcast, since the kernel
//! never wakes a sender parked on a dead peer's endpoint. So delivery uses the
//! asynchronous `ipc.send`: it posts the event to each subscriber's endpoint queue and
//! returns at once, and can never block on a subscriber. That primitive exists for exactly
//! this ([ipc.md](../../../docs/ipc.md), "asynchronous / buffered send").
//!
//! A subscriber registers by handing the service its own endpoint as a capability (M13
//! capability passing — this service is its first real consumer). The service keeps that
//! handle and `ipc.send`s each event to it.
const std = @import("std");
const runtime = @import("runtime");
const protocol = runtime.input_protocol;
const ipc = runtime.ipc;
const system = runtime.system;
/// One registered subscriber: the endpoint we push events to (a capability it handed us at
/// subscribe time) and the task id that owns it (the subscribe call's badge), so a slot
/// left behind by a subscriber that exited can be reclaimed.
const Subscriber = struct {
used: bool = false,
endpoint: ipc.Handle = 0,
task_id: u32 = 0,
/// Which device classes this subscriber wants (an OR of protocol.device_*). An event
/// is delivered only if its device's bit is set here.
device_mask: u32 = 0,
};
var subscribers = [_]Subscriber{.{}} ** 8;
/// Drop any subscriber whose owning process is no longer alive, so its slot (and the
/// endpoint reference it holds) can be reused. Cheap and only run on subscribe — the async
/// `send` to a dead subscriber's orphaned endpoint is harmless (it just fills a queue no
/// one drains), so this is housekeeping, not correctness.
fn pruneDeadSubscribers() void {
var table: [32]system.ProcessDescriptor = undefined;
const total = system.processes(&table);
const count = @min(total, table.len);
for (&subscribers) |*sub| {
if (!sub.used) continue;
var alive = false;
for (table[0..count]) |descriptor| {
if (descriptor.id == sub.task_id) {
alive = true;
break;
}
}
if (!alive) sub.* = .{};
}
}
/// Register `endpoint` (owned by task `task_id`) to receive the device classes in
/// `device_mask`. Returns false if the subscriber table is full.
fn addSubscriber(endpoint: ipc.Handle, task_id: u32, device_mask: u32) bool {
for (&subscribers) |*sub| {
if (!sub.used) {
sub.* = .{ .used = true, .endpoint = endpoint, .task_id = task_id, .device_mask = device_mask };
return true;
}
}
return false;
}
/// Push `event` to every subscriber whose interest mask includes its device class.
/// `ipc.send` never blocks, so a slow or dead subscriber cannot stall delivery to others.
fn broadcast(event: protocol.InputEvent) void {
const bytes = std.mem.asBytes(&event);
const bit = protocol.deviceBit(event.device);
for (&subscribers) |*sub| {
if (sub.used and sub.device_mask & bit != 0) _ = ipc.send(sub.endpoint, bytes);
}
}
/// Handle one request. `got` carries the sender badge (a task id) and, for subscribe, the
/// subscriber's endpoint capability in `got.cap`. Writes a `Reply` into `out` and returns
/// its length.
fn handle(message: []const u8, got: ipc.Received, out: []u8) usize {
const reply = struct {
fn write(buffer: []u8, status: i32) usize {
const header = protocol.Reply{ .status = status };
@memcpy(buffer[0..protocol.reply_size], std.mem.asBytes(&header));
return protocol.reply_size;
}
};
if (message.len < protocol.request_size) return reply.write(out, -1);
const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]);
switch (@as(protocol.Operation, @enumFromInt(request.operation))) {
.subscribe => {
const endpoint = got.cap orelse return reply.write(out, -1); // no endpoint passed
// A zero mask means "everything" (a subscriber that named no class still wants input).
const mask = if (request.device_mask == 0) protocol.device_all else request.device_mask;
pruneDeadSubscribers();
if (!addSubscriber(endpoint, @intCast(got.badge), mask)) return reply.write(out, -1); // table full
return reply.write(out, 0);
},
.publish => {
broadcast(request.event);
return reply.write(out, 0);
},
}
}
pub fn main() void {
const endpoint = ipc.createIpcEndpoint() orelse {
_ = system.write("input: no endpoint\n");
return;
};
if (!ipc.register(.input, endpoint)) {
_ = system.write("input: register failed\n");
return;
}
_ = system.write("input: ready\n");
var reply_buffer: [protocol.reply_size]u8 = undefined;
var reply_len: usize = 0;
var receive: [protocol.request_size]u8 = undefined;
while (true) {
const got = ipc.replyWait(endpoint, reply_buffer[0..reply_len], &receive, null);
// Only synchronous client requests (subscribe/publish) arrive here; nothing sends
// this service asynchronous messages, so a notification wake would be spurious.
if (got.isNotification()) {
reply_len = 0;
continue;
}
reply_len = handle(receive[0..got.len], got, &reply_buffer);
}
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start;
}
+319
View File
@@ -0,0 +1,319 @@
//! The input wire protocol — the message format spoken between the user-space input
//! service ([input.zig](input.zig)) and the two kinds of process that reach it: a
//! **source** (a keyboard, mouse, or joystick/gamepad driver) that publishes events, and a
//! **subscriber** (any program) that subscribes and is then pushed each event.
//!
//! The service handles several device classes over one endpoint. Each class has its own
//! typed event (`KeyEvent`, `MouseEvent`, `JoystickEvent`); they all travel in a common
//! `InputEvent` envelope tagged with a `DeviceKind`, so the fan-out path is one code path
//! and a subscriber can take a mix of devices on a single stream. A subscriber declares
//! which classes it wants with a `device_mask`, and the service routes accordingly.
//!
//! Two message shapes ride over the endpoint, tagged by `Operation`, like the
//! [VFS protocol](../vfs/protocol.zig):
//!
//! - **subscribe / publish**: a synchronous `ipc_call` carrying a `Request`. `subscribe`
//! hands the service the subscriber's own endpoint as a capability (`send_cap`) and a
//! `device_mask`; `publish` carries an `InputEvent`. The reply is a `Reply`.
//! - **delivery**: the service pushes each `InputEvent` to every interested subscriber with
//! the asynchronous `ipc_send` — no reply owed, and a dead subscriber can never stall the
//! broadcast. Received in the subscriber's buffer with `Received.isMessage()` set.
//!
//! This is a danos-native contract, shared by the input service, the `runtime.input`
//! client helpers, and every source/subscriber. Everything fits one IPC message.
const std = @import("std");
/// The classes of input device the service fans out. Each names a typed event and a bit in
/// the subscription mask.
pub const DeviceKind = enum(u32) {
keyboard = 0,
mouse = 1,
joystick = 2, // joysticks and gamepads/controllers
};
/// Subscription-interest bits (`Request.device_mask`) — which device classes a subscriber
/// wants. OR them together, or use `device_all`.
pub const device_keyboard: u32 = 1 << 0;
pub const device_mouse: u32 = 1 << 1;
pub const device_joystick: u32 = 1 << 2;
pub const device_all: u32 = device_keyboard | device_mouse | device_joystick;
/// The subscription bit for a `DeviceKind` value (as it appears in `InputEvent.device`).
/// An unknown device maps to 0, so it matches no subscriber.
pub fn deviceBit(device: u32) u32 {
return switch (device) {
@intFromEnum(DeviceKind.keyboard) => device_keyboard,
@intFromEnum(DeviceKind.mouse) => device_mouse,
@intFromEnum(DeviceKind.joystick) => device_joystick,
else => 0,
};
}
// --- keyboard ---------------------------------------------------------------
/// What happened to a key. `key_down`/`key_up` are the physical make/break; `key_press`
/// is the higher-level "a character was produced", carrying it in `KeyEvent.character`.
pub const EventKind = enum(u32) {
key_down = 0,
key_up = 1,
key_press = 2,
};
/// One keyboard event. `keycode` names the physical key (layout-independent); `character`
/// is the Unicode scalar for `key_press` (else 0); `modifiers` is an OR of `modifier_*`.
pub const KeyEvent = extern struct {
kind: u32, // an EventKind
keycode: u32, // a Keycode
character: u32, // Unicode scalar for key_press, else 0
modifiers: u32, // OR of modifier_*
};
pub const modifier_shift: u32 = 1 << 0;
pub const modifier_control: u32 = 1 << 1;
pub const modifier_alt: u32 = 1 << 2;
/// The danos-native keycode namespace: USB HID keyboard-page usages (page 0x07), the
/// numbering the PS/2 scancode decoder emits and the xkeyboard-config layout tables are
/// indexed by. Non-exhaustive, so an unnamed usage still travels as a valid value.
pub const Keycode = enum(u32) {
unknown = 0,
// letters
a = 4,
b = 5,
c = 6,
d = 7,
e = 8,
f = 9,
g = 10,
h = 11,
i = 12,
j = 13,
k = 14,
l = 15,
m = 16,
n = 17,
o = 18,
p = 19,
q = 20,
r = 21,
s = 22,
t = 23,
u = 24,
v = 25,
w = 26,
x = 27,
y = 28,
z = 29,
// digit row
one = 30,
two = 31,
three = 32,
four = 33,
five = 34,
six = 35,
seven = 36,
eight = 37,
nine = 38,
zero = 39,
// control and whitespace
enter = 40,
escape = 41,
backspace = 42,
tab = 43,
spacebar = 44,
// punctuation
minus = 45,
equal = 46,
left_bracket = 47,
right_bracket = 48,
backslash = 49,
non_us_hash = 50,
semicolon = 51,
apostrophe = 52,
grave = 53,
comma = 54,
period = 55,
slash = 56,
caps_lock = 57,
// function row
f1 = 58,
f2 = 59,
f3 = 60,
f4 = 61,
f5 = 62,
f6 = 63,
f7 = 64,
f8 = 65,
f9 = 66,
f10 = 67,
f11 = 68,
f12 = 69,
print_screen = 70,
scroll_lock = 71,
pause = 72,
// navigation
insert = 73,
home = 74,
page_up = 75,
delete = 76,
end = 77,
page_down = 78,
right_arrow = 79,
left_arrow = 80,
down_arrow = 81,
up_arrow = 82,
// keypad
num_lock = 83,
keypad_slash = 84,
keypad_asterisk = 85,
keypad_minus = 86,
keypad_plus = 87,
keypad_enter = 88,
keypad_one = 89,
keypad_two = 90,
keypad_three = 91,
keypad_four = 92,
keypad_five = 93,
keypad_six = 94,
keypad_seven = 95,
keypad_eight = 96,
keypad_nine = 97,
keypad_zero = 98,
keypad_period = 99,
non_us_backslash = 100,
application = 101,
// modifiers
left_control = 224,
left_shift = 225,
left_alt = 226,
left_gui = 227,
right_control = 228,
right_shift = 229,
right_alt = 230,
right_gui = 231,
_,
};
// --- mouse ------------------------------------------------------------------
/// What a mouse event reports. `motion` carries relative `dx`/`dy`; `button_down`/`up`
/// name a button in `button`; `scroll` carries `scroll_x`/`scroll_y`.
pub const MouseEventKind = enum(u32) {
motion = 0,
button_down = 1,
button_up = 2,
scroll = 3,
};
pub const mouse_button_left: u32 = 1 << 0;
pub const mouse_button_right: u32 = 1 << 1;
pub const mouse_button_middle: u32 = 1 << 2;
/// One mouse event. Relative motion (`dx`/`dy`) and wheel (`scroll_*`) are signed;
/// `buttons` is the current pressed-button bitmask (`mouse_button_*`).
pub const MouseEvent = extern struct {
kind: u32, // a MouseEventKind
button: u32, // the mouse_button_* bit for button_down/up, else 0
dx: i32, // relative X motion (.motion)
dy: i32, // relative Y motion (.motion)
scroll_x: i32, // horizontal wheel (.scroll)
scroll_y: i32, // vertical wheel (.scroll)
buttons: u32, // current pressed-button bitmask
};
// --- joystick / gamepad -----------------------------------------------------
/// What a joystick/gamepad event reports. `axis` carries a signed `value` on axis
/// `control`; `button_down`/`up` name a button index in `control`.
pub const JoystickEventKind = enum(u32) {
axis = 0,
button_down = 1,
button_up = 2,
};
/// One joystick/gamepad event. `control` is the axis index (`.axis`) or button index
/// (button events); `value` is the axis position (signed, e.g. -32768..32767) for `.axis`;
/// `buttons` is the current pressed-button bitmask.
pub const JoystickEvent = extern struct {
kind: u32, // a JoystickEventKind
control: u32, // axis index (.axis) or button index (button events)
value: i32, // axis value for .axis, else 0
buttons: u32, // current pressed-button bitmask
};
// --- the common envelope ----------------------------------------------------
/// The largest per-device event, so `InputEvent` can hold any of them inline.
pub const max_event_size: usize = @max(@sizeOf(KeyEvent), @max(@sizeOf(MouseEvent), @sizeOf(JoystickEvent)));
/// The tagged envelope broadcast to subscribers: a `DeviceKind` plus the raw bytes of the
/// matching per-device event. Decode it with `asKeyboard`/`asMouse`/`asJoystick` (each
/// returns null unless `device` matches), or build one with the `from*` constructors.
pub const InputEvent = extern struct {
device: u32, // a DeviceKind
_padding: u32 = 0,
data: [max_event_size]u8 = [_]u8{0} ** max_event_size,
pub fn asKeyboard(self: InputEvent) ?KeyEvent {
if (self.device != @intFromEnum(DeviceKind.keyboard)) return null;
return std.mem.bytesToValue(KeyEvent, self.data[0..@sizeOf(KeyEvent)]);
}
pub fn asMouse(self: InputEvent) ?MouseEvent {
if (self.device != @intFromEnum(DeviceKind.mouse)) return null;
return std.mem.bytesToValue(MouseEvent, self.data[0..@sizeOf(MouseEvent)]);
}
pub fn asJoystick(self: InputEvent) ?JoystickEvent {
if (self.device != @intFromEnum(DeviceKind.joystick)) return null;
return std.mem.bytesToValue(JoystickEvent, self.data[0..@sizeOf(JoystickEvent)]);
}
pub fn fromKeyboard(event: KeyEvent) InputEvent {
return pack(.keyboard, std.mem.asBytes(&event));
}
pub fn fromMouse(event: MouseEvent) InputEvent {
return pack(.mouse, std.mem.asBytes(&event));
}
pub fn fromJoystick(event: JoystickEvent) InputEvent {
return pack(.joystick, std.mem.asBytes(&event));
}
fn pack(device: DeviceKind, bytes: []const u8) InputEvent {
var self = InputEvent{ .device = @intFromEnum(device) };
@memcpy(self.data[0..bytes.len], bytes);
return self;
}
};
// --- request / reply --------------------------------------------------------
/// Which side of a request this is.
pub const Operation = enum(u32) {
subscribe = 0, // register the caller's endpoint (send_cap) for the classes in device_mask
publish = 1, // a source submits `event` to broadcast to interested subscribers
};
/// Request header. For `subscribe`, `device_mask` is the OR of `device_*` bits the caller
/// wants (0 means all) and the caller's receive endpoint travels as the call's capability;
/// `event` is ignored. For `publish`, `event` is the event to broadcast.
pub const Request = extern struct {
operation: u32, // an Operation
device_mask: u32 = 0, // subscribe: interested device classes (0 => all)
event: InputEvent = .{ .device = 0 },
};
/// Reply header. `status` is 0 on success or a negative errno.
pub const Reply = extern struct {
status: i32,
_padding: u32 = 0,
};
pub const request_size: usize = @sizeOf(Request);
pub const reply_size: usize = @sizeOf(Reply);
pub const event_size: usize = @sizeOf(InputEvent);
comptime {
// The delivery path posts a bare InputEvent through ipc_send, so it must fit an
// endpoint's async payload slot (POST_MAXIMUM is 64).
if (event_size > 64) @compileError("InputEvent must fit the ipc_send payload (POST_MAXIMUM)");
}
@@ -0,0 +1,98 @@
//! process-test — a test fixture for process management (bundled in the
//! initial-ramdisk, driven by the `supervision` test case). One binary, three
//! roles picked by argv, so the whole user-side surface is exercised end to end:
//!
//! - `process-test run` — the supervisor: spawns the two children below with an
//! exit-notification endpoint, sees them in `process_enumerate`, kills them,
//! receives both exit notifications, and confirms they are gone. Prints
//! "process-test: ok" for the kernel test to match, or a FAIL line naming the
//! step that broke.
//! - `process-test sleeper` — a child that blocks in `sleep` forever: its kill
//! exercises the immediate reap of a blocked task.
//! - `process-test spinner` — a child that spins in user mode making no system
//! calls: its kill exercises the deferred path (kill_pending, finished by the
//! timer tick).
//!
//! Spawned with no arguments (the initial-ramdisk sweep test starts every bundled
//! binary bare), it exits silently so it cannot derange other tests' output.
const std = @import("std");
const runtime = @import("runtime");
fn fail(step: []const u8) noreturn {
_ = runtime.system.write("process-test: FAIL ");
_ = runtime.system.write(step);
_ = runtime.system.write("\n");
runtime.system.exit(1);
}
/// Whether process `id` appears in a fresh `process_enumerate` snapshot, named
/// `name` (an id present under the wrong name is a table mix-up, not a pass).
fn listed(id: u32, name: []const u8) bool {
var table: [32]runtime.system.ProcessDescriptor = undefined;
const total = runtime.system.processes(&table);
for (table[0..@min(total, table.len)]) |descriptor| {
if (descriptor.id != id) continue;
return std.mem.eql(u8, descriptor.name[0..descriptor.name_length], name);
}
return false;
}
/// Block on the exit endpoint until a child-exit notification arrives; returns
/// the ended child's id. A wrong wake-up (there should be none — nothing else
/// knows this endpoint) fails the test rather than looping forever.
fn awaitChildExit(endpoint: runtime.ipc.Handle) u32 {
var scratch: [8]u8 = undefined;
const received = runtime.ipc.replyWait(endpoint, scratch[0..0], &scratch, null);
if (!received.isChildExit()) fail("expected a child-exit notification");
return received.childProcessId();
}
pub fn main(init: runtime.process.Init) void {
const role = init.arguments.get(1) orelse return; // spawned bare (ramdisk sweep): stay silent
if (std.mem.eql(u8, role, "sleeper")) {
while (true) runtime.system.sleep(500);
}
if (std.mem.eql(u8, role, "spinner")) {
var beat: u64 = 0;
const touch: *volatile u64 = &beat;
while (true) touch.* +%= 1; // user mode only — no system calls to die at
}
// The supervisor ("run").
const endpoint = runtime.ipc.createIpcEndpoint() orelse fail("create exit endpoint");
const sleeper = runtime.system.spawnSupervised("process-test", &.{"sleeper"}, endpoint) orelse fail("spawn sleeper");
const spinner = runtime.system.spawnSupervised("process-test", &.{"spinner"}, endpoint) orelse fail("spawn spinner");
runtime.system.sleep(100); // let the sleeper block and the spinner get a core
if (!listed(sleeper, "process-test")) fail("sleeper not in process_enumerate");
if (!listed(spinner, "process-test")) fail("spinner not in process_enumerate");
// Kills that must be refused: a kernel task (id 0), and an id that was never
// issued — both -ESRCH. (-EPERM needs a second supervisor; the kernel-level
// `process-kill` test covers it.)
if (runtime.system.kill(0)) fail("killing a kernel task was allowed");
if (runtime.system.kill(0xFFFF_FFF0)) fail("killing an unknown id was allowed");
// The blocked child: usually reaped on the spot (it sits in `sleep`). The
// notification is the fence — after it, the child is certainly gone, so the
// second kill must miss (its id is never reused).
if (!runtime.system.kill(sleeper)) fail("kill sleeper");
if (awaitChildExit(endpoint) != sleeper) fail("sleeper exit notification");
if (runtime.system.kill(sleeper)) fail("double kill was allowed");
// The running child: the deferred path — condemned now, dead by the next tick.
if (!runtime.system.kill(spinner)) fail("kill spinner");
if (awaitChildExit(endpoint) != spinner) fail("spinner exit notification");
if (listed(sleeper, "process-test")) fail("sleeper still listed after kill");
if (listed(spinner, "process-test")) fail("spinner still listed after kill");
_ = runtime.system.write("process-test: ok\n");
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
+2 -2
View File
@@ -8,8 +8,8 @@
//! spellings (`stat`, `O_CREAT`, ...) live only in the POSIX layer
//! (library/posix/unistd.zig), which translates to these.
//!
//! This is user-space only — the kernel knows nothing of files or paths; it only
//! moves the bytes. Shared by library/posix/unistd.zig (client) and system/services/vfs/vfs.zig (server).
//! This is user-space only — the kernel knows nothing of files or paths; it only moves the bytes.
//! Shared by library/posix/unistd.zig (client) and system/services/vfs/vfs.zig (server).
pub const Operation = enum(u32) {
open, // open(path) -> node id
+1 -1
View File
@@ -114,7 +114,7 @@ fn handle(message: []const u8, out: []u8) usize {
}
pub fn main() void {
const endpoint = runtime.ipc.createEndpoint() orelse {
const endpoint = runtime.ipc.createIpcEndpoint() orelse {
_ = runtime.system.write("vfs: no endpoint\n");
return;
};
+62
View File
@@ -130,6 +130,31 @@ CASES = [
{"name": "ipc-cap",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# DMA memory (M14): contiguous frame allocation, below-4G cap, coherent mapping,
# and reclaim on teardown.
{"name": "dma",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# MSI (M15): allocate a per-device vector and deliver it as a notification (a
# self-IPI stands in for the device's MSI write, since the HPET has no MSI).
{"name": "msi",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# IOMMU (M16): boot with an emulated VT-d unit and confirm danos parses the DMAR
# table and reads the unit's registers. Detection only — enforcement is future.
{"name": "iommu",
"qemu_extra": ["-device", "intel-iommu,intremap=off"],
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Port I/O grants: a claimed device's io_port resource lets a driver read/write its
# ports (PS/2 status 0x64), gated by the claim; out-of-range/unclaimed is refused.
{"name": "ioport",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Monotonic clock (clock() syscall source): calibrated, advancing, never backwards.
{"name": "clock",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Parallelism: needs more than one core, so this case boots with -smp 4.
{"name": "smp",
"smp": 4,
@@ -174,6 +199,17 @@ CASES = [
{"name": "user-pf",
"expect": r"page fault \(vector 14\)[\s\S]*error code : 0x5[\s\S]*IP\s*: 0x00007000000000",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Fault recovery: a scheduled ring-3 process that faults is killed — resources
# reclaimed, core kept — while init's heartbeat proves the OS survived.
{"name": "fault-recovery",
"timeout": 60,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Process arguments: argv arrives on the SysV entry stack (argv[0] = the spawned
# name, argv[1..] = the system_spawn argument blob) and echoes back intact.
{"name": "args",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# The real user binary: the bootloader ships /system/services/init off the ESP, the
# kernel loads the ELF and runs it in ring 3, and it writes + exits cleanly.
{"name": "init",
@@ -186,6 +222,24 @@ CASES = [
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# process_enumerate: a task-table snapshot lists spawned processes by name and
# id alongside the kernel tasks, and a too-small buffer still reports the total.
{"name": "process-list",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# process_kill + the exit notification: only the supervisor may kill; a blocked
# victim is reaped in place and a spinning one dies by the deferred (tick) path;
# each death posts one exit badge to the endpoint given at spawn.
{"name": "process-kill",
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# The user-side whole: process-test supervises two children from ring 3 —
# spawn with an exit endpoint, enumerate, kill, notification, gone.
{"name": "supervision",
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# The initial_ramdisk: the loader ferries a bundle of user binaries; the kernel parses
# it and spawns each as a ring-3 process (here the VFS-server stub heartbeats).
{"name": "initial-ramdisk",
@@ -196,6 +250,12 @@ CASES = [
{"name": "vfs",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# The input service: a synthetic keyboard source publishes events, the service
# broadcasts them over the async ipc_send primitive, and a subscriber (which joined by
# passing its endpoint as a capability) receives them — source -> service -> subscriber.
{"name": "input",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# IO passthrough + IRQ-as-IPC: a user-space HPET driver maps device MMIO into
# its own address space, binds the device's interrupt to an IPC endpoint, and
# is woken by the hardware five times while blocked (never polling).
@@ -300,6 +360,8 @@ def run_case(arch, case):
cmd = [arch["qemu"]] + arch["qemu_args"](arch, esp, vars_fd, serial)
if case.get("smp"): # some cases need more than one core (e.g. parallelism)
cmd += ["-smp", str(case["smp"])]
if case.get("qemu_extra"): # extra qemu args, e.g. -device intel-iommu for the IOMMU case
cmd += case["qemu_extra"]
qemu = subprocess.Popen(cmd, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
try:
timeout = case.get("timeout", TIMEOUT)
+468
View File
@@ -0,0 +1,468 @@
#!/usr/bin/env python3
"""Vendor xkeyboard-config and compile it to a native Zig keymap library.
danos does not ship an X11 runtime, but it wants X11's keyboard *data*: the layout
tables (US, UK, German, ...) that turn a physical key + modifiers into a character.
So we compile xkeyboard-config down to Zig at build time, the same way the initial
ramdisk is packed by a host-side Python tool.
Two subcommands:
fetch Download the pinned xkeyboard-config release, verify its sha256, and copy
the transitively-needed data (the symbols files for the configured layouts
plus everything they `include`) into library/xkeyboard-config/vendor/,
alongside a vendored keysymdef.h, the upstream COPYING, and a PROVENANCE.md.
This is the one step that needs the network; run it when bumping the version.
generate Parse the vendored data and emit library/xkeyboard-config/generated/layouts.zig
(deterministic, offline). Run whenever the layout list or emitter changes.
The pipeline per key: our events carry a USB HID usage; HID_TO_NAME maps it to an XKB
key name (<AC01>); the layout's symbols give the keysyms per level; keysymdef.h resolves
each keysym name to its value and Unicode scalar. Level *selection* (which modifier picks
which level) is deliberately left to the Zig side (xkeyboard-config.zig) — we only emit the
per-key type and its up-to-4 keysyms here.
"""
import argparse
import hashlib
import io
import os
import re
import sys
import tarfile
import urllib.request
# --- pinned upstream -------------------------------------------------------
VERSION = "2.44"
TARBALL_URL = (
"https://gitlab.freedesktop.org/xkeyboard-config/xkeyboard-config/-/archive/"
f"xkeyboard-config-{VERSION}/xkeyboard-config-{VERSION}.tar.gz"
)
TARBALL_SHA256 = "35e34edeaf4e8da8d0696ff6b241ee11ddb1b8c6730bac7252d4d0a88ea5f05b"
# keysymdef.h is xorgproto, not xkeyboard-config; fetch copies it from the build host.
KEYSYMDEF_CANDIDATES = [
"/opt/homebrew/include/X11/keysymdef.h",
"/usr/include/X11/keysymdef.h",
"/usr/X11/include/X11/keysymdef.h",
"/usr/X11R6/include/X11/keysymdef.h",
]
# The layouts we generate: (zig name, symbols file, section or None=default).
TARGETS = [
("us", "us", None),
("gb", "gb", None),
("de", "de", None),
("fr", "fr", None),
("es", "es", None),
("dvorak", "us", "dvorak"),
]
# USB HID usage (keyboard page 0x07) -> XKB key name. The physical keys we can turn into
# characters; positions are the ANSI standard shared by HID usages and XKB names.
HID_TO_NAME = {
0x04: "AC01", 0x05: "AB05", 0x06: "AB03", 0x07: "AC03", 0x08: "AD03",
0x09: "AC04", 0x0A: "AC05", 0x0B: "AC06", 0x0C: "AD08", 0x0D: "AC07",
0x0E: "AC08", 0x0F: "AC09", 0x10: "AB07", 0x11: "AB06", 0x12: "AD09",
0x13: "AD10", 0x14: "AD01", 0x15: "AD04", 0x16: "AC02", 0x17: "AD05",
0x18: "AD07", 0x19: "AB04", 0x1A: "AD02", 0x1B: "AB02", 0x1C: "AD06",
0x1D: "AB01",
0x1E: "AE01", 0x1F: "AE02", 0x20: "AE03", 0x21: "AE04", 0x22: "AE05",
0x23: "AE06", 0x24: "AE07", 0x25: "AE08", 0x26: "AE09", 0x27: "AE10",
0x2C: "SPCE",
0x2D: "AE11", 0x2E: "AE12", 0x2F: "AD11", 0x30: "AD12", 0x31: "BKSL",
0x33: "AC10", 0x34: "AC11", 0x35: "TLDE",
0x36: "AB08", 0x37: "AB09", 0x38: "AB10",
0x64: "LSGT", # the extra key on ISO keyboards (102nd key)
}
# XKB type name -> the KeyType enum tag emitted for the Zig side.
TYPE_TO_TAG = {
"ONE_LEVEL": "one_level",
"TWO_LEVEL": "two_level",
"ALPHABETIC": "alphabetic",
"FOUR_LEVEL": "four_level",
"FOUR_LEVEL_ALPHABETIC": "four_level_alphabetic",
"FOUR_LEVEL_SEMIALPHABETIC": "four_level_semialphabetic",
"FOUR_LEVEL_MIXED_KEYPAD": "keypad",
"FOUR_LEVEL_KEYPAD": "keypad",
"KEYPAD": "keypad",
"PC_SYSRQ": "other",
}
REPO = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
LIB = os.path.join(REPO, "library", "xkeyboard-config")
VENDOR = os.path.join(LIB, "vendor")
# --- keysymdef.h: keysym name -> (value, unicode scalar or None) ------------
def parse_keysymdef(path):
table = {}
line_re = re.compile(r"#define\s+XK_(\w+)\s+0x([0-9a-fA-F]+)\s*(?:/\*\s*U\+([0-9A-Fa-f]+)|/\*<U\+([0-9A-Fa-f]+))?")
with open(path, encoding="latin-1") as f:
for line in f:
m = line_re.match(line)
if not m:
continue
name = m.group(1)
value = int(m.group(2), 16)
uni = m.group(3) or m.group(4)
unicode_scalar = int(uni, 16) if uni else None
# First definition wins (keysymdef lists the canonical one first).
table.setdefault(name, (value, unicode_scalar))
return table
def resolve_keysym(name, keysymdef):
"""A keysym token from a symbols file -> (keysym value, unicode scalar or 0)."""
if name in ("NoSymbol", "VoidSymbol", ""):
return (0, 0)
# Unicode escape forms: U00E9 / u00E9.
m = re.fullmatch(r"[Uu]([0-9A-Fa-f]{4,6})", name)
if m:
cp = int(m.group(1), 16)
return (0x01000000 + cp, cp)
# Raw hex keysym value: 0x1000161 (a Unicode keysym) or a plain keysym number.
m = re.fullmatch(r"0x([0-9A-Fa-f]+)", name)
if m:
value = int(m.group(1), 16)
if 0x01000100 <= value <= 0x0110FFFF:
return (value, value - 0x01000000)
if 0x20 <= value <= 0x7E or 0xA0 <= value <= 0xFF:
return (value, value)
return (value, 0)
entry = keysymdef.get(name)
if entry is None:
return (0, 0) # unknown named keysym (dead_*, ISO_*, rare) -> no character
value, unicode_scalar = entry
return (value, unicode_scalar or 0)
# --- symbols parsing --------------------------------------------------------
def strip_comments(text):
text = re.sub(r"/\*.*?\*/", "", text, flags=re.S)
text = re.sub(r"//[^\n]*", "", text)
return text
class KeyDef:
__slots__ = ("levels", "type_name")
def __init__(self, levels, type_name):
self.levels = levels # list of keysym-name strings (group 1)
self.type_name = type_name # explicit XKB type name, or None
class Symbols:
"""Loads and resolves xkb_symbols sections from a directory of symbols files."""
def __init__(self, symbols_dir):
self.dir = symbols_dir
self._files = {} # filename -> {section_name: body, "__default__": name}
self.visited_files = set()
def _load(self, fname):
if fname in self._files:
return self._files[fname]
path = os.path.join(self.dir, fname)
self.visited_files.add(fname)
with open(path, encoding="latin-1") as f:
text = strip_comments(f.read())
sections = {}
default = None
# Find each `[flags] xkb_symbols "NAME" {` and brace-match its body.
for m in re.finditer(r'(\w[\w\s]*?)?\bxkb_symbols\s+"([^"]+)"\s*\{', text):
flags = m.group(1) or ""
name = m.group(2)
body, _ = self._brace_match(text, m.end() - 1)
sections[name] = body
if "default" in flags.split() and default is None:
default = name
if default is None and sections:
default = next(iter(sections))
sections["__default__"] = default
self._files[fname] = sections
return sections
@staticmethod
def _brace_match(text, open_index):
depth = 0
for i in range(open_index, len(text)):
c = text[i]
if c == "{":
depth += 1
elif c == "}":
depth -= 1
if depth == 0:
return text[open_index + 1:i], i
raise ValueError("unbalanced braces")
def resolve(self, fname, section=None, stack=()):
"""Merged {key name -> KeyDef} for (fname, section), following includes."""
sections = self._load(fname)
if section is None:
section = sections["__default__"]
key = (fname, section)
if key in stack:
return {}
body = sections.get(section)
if body is None:
return {}
keys = {}
default_type = None
for stmt in self._statements(body):
kind = stmt[0]
if kind == "include":
augment, inc_file, inc_section = stmt[1], stmt[2], stmt[3]
inc = self.resolve(inc_file, inc_section, stack + (key,))
for n, kd in inc.items():
if augment and n in keys:
continue
keys[n] = kd
elif kind == "default_type":
default_type = stmt[1]
elif kind == "key":
augment, name, kd = stmt[1], stmt[2], stmt[3]
if augment and name in keys:
continue
keys[name] = kd
if default_type is not None:
for kd in keys.values():
if kd.type_name is None:
kd.type_name = default_type
return keys
def _statements(self, body):
"""Yield ('include', augment, file, section) / ('default_type', name) /
('key', augment, name, KeyDef) in source order."""
# Recognise the three statement heads and step through the body in order.
head = re.compile(
r'(?P<inc>(?:(?P<augi>augment|override|replace)\s+)?include\s+"(?P<incarg>[^"]+)")'
r'|(?P<dtype>key\.type(?:\[[^\]]*\])?\s*=\s*"(?P<dtypearg>[^"]+)")'
r'|(?P<key>(?:(?P<augk>augment|override|replace)\s+)?key\s+<(?P<kname>[^>]+)>\s*\{)'
)
i = 0
while i < len(body):
m = head.search(body, i)
if not m:
break
if m.group("inc"):
inc_file, inc_section = self._parse_include(m.group("incarg"))
yield ("include", m.group("augi") == "augment", inc_file, inc_section)
i = m.end()
elif m.group("dtype"):
yield ("default_type", m.group("dtypearg"))
i = m.end()
else: # a key definition; brace-match its body
kbody, end = self._brace_match(body, m.end() - 1)
kd = self._parse_key_body(kbody)
if kd is not None:
yield ("key", m.group("augk") == "augment", m.group("kname"), kd)
i = end + 1
@staticmethod
def _parse_include(arg):
m = re.fullmatch(r"([^()]+)(?:\(([^)]+)\))?", arg.strip())
return (m.group(1), m.group(2))
@staticmethod
def _parse_key_body(kbody):
type_name = None
mt = re.search(r'type(?:\[[^\]]*\])?\s*=\s*"([^"]+)"', kbody)
if mt:
type_name = mt.group(1)
groups = re.findall(r"\[([^\]]*)\]", kbody)
if not groups:
return None
levels = [tok.strip() for tok in groups[0].split(",")]
levels = [t for t in levels if t != ""]
if not levels:
return None
return KeyDef(levels, type_name)
# --- type inference + emission ---------------------------------------------
def is_case_pair(a, b):
return len(a) == 1 and len(b) == 1 and a.isalpha() and a.islower() and b == a.upper()
def key_type_tag(kd):
if kd.type_name is not None:
return TYPE_TO_TAG.get(kd.type_name, "other")
n = len(kd.levels)
if n <= 1:
return "one_level"
if n == 2:
return "alphabetic" if is_case_pair(kd.levels[0], kd.levels[1]) else "two_level"
return "four_level_alphabetic" if is_case_pair(kd.levels[0], kd.levels[1]) else "four_level"
def build_layout(symbols, keysymdef, fname, section):
keys = symbols.resolve(fname, section)
table = [] # 256 entries: (tag, [(keysym, unicode) x4])
for hid in range(256):
name = HID_TO_NAME.get(hid)
kd = keys.get(name) if name else None
if kd is None:
table.append(("one_level", [(0, 0)] * 4))
continue
levels = [resolve_keysym(kd.levels[i], keysymdef) if i < len(kd.levels) else (0, 0)
for i in range(4)]
table.append((key_type_tag(kd), levels))
return table
def emit_zig(layouts):
out = io.StringIO()
out.write("// GENERATED by tools/make-xkeyboard-config.py from xkeyboard-config "
f"{VERSION}. Do not edit.\n")
out.write("// Keyboard layout tables compiled from the X11 xkeyboard-config database\n")
out.write("// (MIT/X11 licensed; see ../vendor/COPYING and ../vendor/PROVENANCE.md).\n\n")
out.write("pub const Level = struct { keysym: u32 = 0, unicode: u21 = 0 };\n\n")
out.write("pub const KeyType = enum {\n")
out.write(" one_level,\n two_level,\n alphabetic,\n four_level,\n"
" four_level_alphabetic,\n four_level_semialphabetic,\n keypad,\n other,\n};\n\n")
out.write("pub const Key = struct { kind: KeyType = .one_level, levels: [4]Level = [_]Level{.{}} ** 4 };\n\n")
out.write("pub const Layout = struct { name: []const u8, keys: [256]Key };\n\n")
names = []
for zig_name, fname, section in TARGETS:
table = layouts[zig_name]
names.append(zig_name)
out.write(f"pub const {zig_name}: Layout = .{{\n")
out.write(f' .name = "{zig_name}",\n')
out.write(" .keys = .{\n")
for hid, (tag, levels) in enumerate(table):
if tag == "one_level" and all(k == 0 and u == 0 for k, u in levels):
out.write(" .{},\n")
continue
parts = ", ".join(f".{{ .keysym = {k}, .unicode = {u} }}" for k, u in levels)
out.write(f" .{{ .kind = .{tag}, .levels = .{{ {parts} }} }},\n")
out.write(" },\n};\n\n")
out.write("pub const all = [_]*const Layout{ " + ", ".join("&" + n for n in names) + " };\n")
return out.getvalue()
# --- fetch ------------------------------------------------------------------
def collect_needed(symbols):
for _, fname, section in TARGETS:
symbols.resolve(fname, section)
return set(symbols.visited_files)
def find_keysymdef():
env = os.environ.get("KEYSYMDEF")
if env and os.path.isfile(env):
return env
for p in KEYSYMDEF_CANDIDATES:
if os.path.isfile(p):
return p
sys.exit("error: keysymdef.h not found; install xorgproto or set KEYSYMDEF=/path/to/keysymdef.h")
def cmd_fetch(_args):
local = os.environ.get("XKB_TARBALL")
if local:
data = open(local, "rb").read()
else:
print(f"downloading {TARBALL_URL}")
data = urllib.request.urlopen(TARBALL_URL).read()
digest = hashlib.sha256(data).hexdigest()
if digest != TARBALL_SHA256:
sys.exit(f"error: sha256 mismatch\n expected {TARBALL_SHA256}\n got {digest}")
tar = tarfile.open(fileobj=io.BytesIO(data), mode="r:gz")
members = tar.getnames()
root = members[0].split("/")[0]
# Extract symbols/ + COPYING to a temp view, resolve includes, keep only what's needed.
tmp = os.path.join(VENDOR, ".upstream")
for m in tar.getmembers():
if m.name.startswith(f"{root}/symbols/") or m.name == f"{root}/COPYING":
m.name = m.name[len(root) + 1:]
tar.extract(m, tmp)
tar.close()
needed = collect_needed(Symbols(os.path.join(tmp, "symbols")))
os.makedirs(os.path.join(VENDOR, "symbols"), exist_ok=True)
for fname in sorted(needed):
src = os.path.join(tmp, "symbols", fname)
dst = os.path.join(VENDOR, "symbols", fname)
os.makedirs(os.path.dirname(dst), exist_ok=True)
with open(src, "rb") as s, open(dst, "wb") as d:
d.write(s.read())
_copy(os.path.join(tmp, "COPYING"), os.path.join(VENDOR, "COPYING"))
keysymdef = find_keysymdef()
_copy(keysymdef, os.path.join(VENDOR, "keysymdef.h"))
_rmtree(tmp)
with open(os.path.join(VENDOR, "PROVENANCE.md"), "w") as f:
f.write("# Vendored xkeyboard-config subset\n\n")
f.write(f"- **Package**: xkeyboard-config {VERSION}\n")
f.write(f"- **Source**: {TARBALL_URL}\n")
f.write(f"- **sha256**: `{TARBALL_SHA256}`\n")
f.write(f"- **keysymdef.h**: xorgproto, copied from `{keysymdef}`\n")
f.write("- **License**: MIT/X11 (see COPYING)\n\n")
f.write("Only the symbols files reachable from the generated layouts "
"(tools/make-xkeyboard-config.py `TARGETS`) are vendored; regenerate with\n"
"`python3 tools/make-xkeyboard-config.py fetch` then `... generate`.\n\n")
f.write("Vendored symbols files:\n\n")
for fname in sorted(needed):
f.write(f"- `symbols/{fname}`\n")
print(f"vendored {len(needed)} symbols files + keysymdef.h + COPYING into {VENDOR}")
def cmd_generate(args):
symbols_dir = args.symbols or os.path.join(VENDOR, "symbols")
keysymdef_path = args.keysymdef or os.path.join(VENDOR, "keysymdef.h")
keysymdef = parse_keysymdef(keysymdef_path)
symbols = Symbols(symbols_dir)
layouts = {zig_name: build_layout(symbols, keysymdef, fname, section)
for zig_name, fname, section in TARGETS}
out_dir = os.path.join(LIB, "generated")
os.makedirs(out_dir, exist_ok=True)
out_path = os.path.join(out_dir, "layouts.zig")
with open(out_path, "w") as f:
f.write(emit_zig(layouts))
print(f"wrote {out_path} ({len(TARGETS)} layouts)")
def _copy(src, dst):
os.makedirs(os.path.dirname(dst), exist_ok=True)
with open(src, "rb") as s, open(dst, "wb") as d:
d.write(s.read())
def _rmtree(path):
for root, dirs, files in os.walk(path, topdown=False):
for name in files:
os.remove(os.path.join(root, name))
for name in dirs:
os.rmdir(os.path.join(root, name))
if os.path.isdir(path):
os.rmdir(path)
def main():
parser = argparse.ArgumentParser(description=__doc__)
sub = parser.add_subparsers(dest="command", required=True)
sub.add_parser("fetch", help="download + vendor the needed xkeyboard-config data")
g = sub.add_parser("generate", help="emit generated/layouts.zig from the vendored data")
g.add_argument("--symbols", help="override the vendored symbols/ dir (for development)")
g.add_argument("--keysymdef", help="override the vendored keysymdef.h (for development)")
args = parser.parse_args()
{"fetch": cmd_fetch, "generate": cmd_generate}[args.command](args)
if __name__ == "__main__":
main()
+152
View File
@@ -0,0 +1,152 @@
#!/usr/bin/env python3
"""Rewrap over-long Zig comment lines at a column limit (default 100).
Rules:
- Only comment-only lines are touched; trailing comments after code are left alone.
- Consecutive comment lines with the same indentation and marker (`//`, `///`, `//!`)
form a block. Blank comment lines separate paragraphs within a block.
- A line whose text starts with `- ` begins a bullet paragraph; its continuation
lines are the ones indented to align under the bullet's text (bullet lead + 2).
- Plain paragraphs join consecutive lines with the same text indentation;
wrapped lines align where the first line's text begins.
- A paragraph is rewrapped only if at least one of its lines exceeds the limit,
so deliberate short line breaks elsewhere are preserved.
"""
import argparse
import difflib
import re
import subprocess
import sys
import textwrap
LIMIT = 100
COMMENT_RE = re.compile(r"^(\s*)(//[/!]?)(?:\s(.*))?$")
BULLET_RE = re.compile(r"^(\s*)- (.*)$")
def split_paragraphs(texts):
"""texts: list of comment text (None for a bare marker line).
Returns paragraphs: dicts with lead/hang/bullet/texts/lines(indices)."""
paragraphs = []
current = None
for index, text in enumerate(texts):
if text is None or text.strip() == "":
paragraphs.append({"literal": True, "lines": [index]})
current = None
continue
lead = len(text) - len(text.lstrip(" "))
bullet = BULLET_RE.match(text)
if bullet:
current = {
"lead": len(bullet.group(1)),
"hang": len(bullet.group(1)) + 2,
"bullet": True,
"texts": [bullet.group(2)],
"lines": [index],
}
paragraphs.append(current)
elif current is not None and lead == current["hang"]:
current["texts"].append(text)
current["lines"].append(index)
else:
current = {
"lead": lead,
"hang": lead,
"bullet": False,
"texts": [text],
"lines": [index],
}
paragraphs.append(current)
return paragraphs
def wrap_paragraph(paragraph, prefix):
joined = re.sub(r"\s+", " ", " ".join(t.strip() for t in paragraph["texts"]))
if paragraph["bullet"]:
initial = prefix + " " * paragraph["lead"] + "- "
else:
initial = prefix + " " * paragraph["lead"]
subsequent = prefix + " " * paragraph["hang"]
return textwrap.wrap(
joined,
width=LIMIT,
initial_indent=initial,
subsequent_indent=subsequent,
break_long_words=False,
break_on_hyphens=False,
)
def rewrap_block(original_lines, indent, marker, texts):
prefix = indent + marker + " "
output = []
for paragraph in split_paragraphs(texts):
block_originals = [original_lines[i] for i in paragraph["lines"]]
if paragraph.get("literal") or all(len(l) <= LIMIT for l in block_originals):
output.extend(block_originals)
else:
output.extend(wrap_paragraph(paragraph, prefix))
return output
def process(source):
lines = source.split("\n")
result = []
i = 0
while i < len(lines):
match = COMMENT_RE.match(lines[i])
if not match:
result.append(lines[i])
i += 1
continue
indent, marker = match.group(1), match.group(2)
block_lines, texts = [], []
while i < len(lines):
m = COMMENT_RE.match(lines[i])
if not m or m.group(1) != indent or m.group(2) != marker:
break
block_lines.append(lines[i])
texts.append(m.group(3))
i += 1
result.extend(rewrap_block(block_lines, indent, marker, texts))
return "\n".join(result)
def main():
parser = argparse.ArgumentParser()
parser.add_argument("--write", action="store_true", help="apply changes (default: diff only)")
parser.add_argument("files", nargs="*", help="files to process (default: git ls-files '*.zig')")
args = parser.parse_args()
files = args.files or subprocess.run(
["git", "ls-files", "*.zig"], capture_output=True, text=True, check=True
).stdout.split()
changed = 0
for path in files:
with open(path, encoding="utf-8") as f:
source = f.read()
rewrapped = process(source)
if rewrapped == source:
continue
changed += 1
if args.write:
with open(path, "w", encoding="utf-8") as f:
f.write(rewrapped)
print(f"rewrapped {path}")
else:
sys.stdout.writelines(
difflib.unified_diff(
source.splitlines(keepends=True),
rewrapped.splitlines(keepends=True),
fromfile=path,
tofile=path,
)
)
if not changed:
print("no comments over the limit")
if __name__ == "__main__":
main()