Author SHA1 Message Date
Daniel Samson 650a1b1595 Signals over IPC, one-shot timers, and the service harness (M17.4)
Signals are statements delivered as coalescing notifications to the
endpoint a process nominates with signal_bind — never a hijacked stack,
never a question (liveness is the zero-length ping the harness answers).
process_signal is supervisor-or-self gated, like kill; unbound targets
accumulate a pending mask delivered on bind. timer_bind is the missing
timed wait: a one-shot deadline landing in the same replyWait as
everything else — what stop(), hello deadlines, and restart backoff are
built from. runtime.service.run folds requests, signals, and
notifications into callbacks; the VFS conversion deletes its hand-rolled
loop and gains the whole lifecycle contract. The signals scenario drives
ping, reload, terminate->exited, the timer, and the deaf-child
deadline->killed path from ring 3. docs/process-lifecycle.md increments
1-4 are now as-built.
2026-07-12 23:53:38 +01:00
Daniel Samson d8c55c6f2f Publish exit events to subscribers; the VFS releases dead clients' handles (M17.3)
process_subscribe adds an endpoint to a bounded, ref-counted subscriber
table; every death posts the same badge encoding a supervisor's exit
notification uses, equally late, so subscribers observe a fully-released
child. A dying subscriber's own subscriptions are removed first — it never
hears about itself. The VFS is the first subscriber: open handles now
record their owner and are swept when the owner dies, because a service
must never depend on clients cleaning up after themselves
(docs/process-lifecycle.md). Proven by the vfs-client-death scenario.
2026-07-12 23:41:44 +01:00
Daniel Samson 2ebfb0c3b0 Record and expose how every process ends (M17.2)
The kernel records an ExitReason at all three death sites — clean exit,
fault (classified by vector), and process_kill — into a bounded ring
before the exit notification posts, so a supervisor's query never races
the notice. process_exit_reason is gated by the same supervisor check as
kill; runtime.process.exitReason is the stable interface. This is the
input restart policy reads (docs/process-lifecycle.md iron rule 2).
2026-07-12 23:34:09 +01:00
Daniel Samson 888eaa74e1 Release a dead process's device claims (M17.1)
Every path out of a process (exit, fault, kill) now releases its device
claims alongside its IRQ and MSI bindings, so a restarted driver can claim
its hardware again — the cleanup half of process-lifecycle.md's iron rule 1.
MSI vectors were already swept by irq.releaseOwner; claims were the gap.
The claim-release test proves kill -> release -> re-claim, plus the broker
release in isolation.
2026-07-12 23:23:49 +01:00
Daniel Samson ed76cbbc79 Mark Phase 0 done: baseline QEMU suite green (48/48) 2026-07-12 23:17:20 +01:00
Daniel Samson 140229b88d Rename usb-xhci-libary.zig to usb-xhci-library.zig (naming typo) 2026-07-12 23:13:33 +01:00
Daniel Samson cb2379fd06 Update README.md 2026-07-12 23:12:39 +01:00
Daniel Samson 1665b239b0 Add the status checklist and workflow to the M17-M18 plan 2026-07-12 23:04:43 +01:00
Daniel Samson 70ed0337f8 Merge feat/usb: USB wire ABI, xHCI detection and spawn, M17-M18 design 2026-07-12 22:56:35 +01:00
Daniel Samson 116b8f6c41 Design the process lifecycle and the device manager (M17-M18)
Signals over IPC (POSIX concepts, message delivery), published exit events,
the stable runtime.process interface, and the device manager as tree +
matcher + supervisor. All open questions settled; docs/m17-m18-plan.md is
the phase-by-phase execution plan.
2026-07-12 22:56:34 +01:00
Daniel Samson 77901bbba6 WIP: USB 2026-07-12 22:24:47 +01:00
Daniel Samson 78582d24d2 code lint 2026-07-12 19:36:10 +01:00
Daniel Samson 1cdffe21b1 fixing comments 2026-07-12 16:19:09 +01:00
Daniel Samson 4df90bc212 add tools/rewrap-comments.py 2026-07-12 16:18:56 +01:00
Daniel Samson 713e77354b gitattributes 2026-07-12 16:09:18 +01:00
Daniel Samson abb7b1b634 editorconfig 2026-07-12 16:09:11 +01:00
Daniel Samson f5f0e15769 zig fmt 2026-07-12 16:04:58 +01:00
Daniel Samson 8652b4a724 add qemu-xhci with usb-mouse and usb-kbd to build.zig 2026-07-12 01:31:38 +01:00
Daniel Samson e8233127c7 Install ps2-bus under its own name instead of clobbering bus
The ps2-bus executable was built with the artifact name "bus" — a
copy-paste from the generic bus driver's line above it. The initial
ramdisk was unaffected (the packer pairs names with binaries
explicitly), so the driver ran at boot; but the FHS install uses the
artifact's own name, so both drivers landed on
zig-out/system/drivers/bus, one overwriting the other, and
zig-out/system/drivers/ps2-bus never existed.
2026-07-11 23:41:08 +01:00
48 changed files with 3185 additions and 143 deletions
+16
View File
@@ -0,0 +1,16 @@
# EditorConfig: https://editorconfig.org/
# Follows the Zig style guide: https://ziglang.org/documentation/0.16.0/#Style-Guide
root = true
[*]
charset = utf-8
end_of_line = lf
indent_style = space
indent_size = 4
trim_trailing_whitespace = true
insert_final_newline = true
[*.zig]
# "Line length: aim for 100; use common sense."
max_line_length = 100
+1
View File
@@ -0,0 +1 @@
*.zig text eol=lf
+26 -7
View File
@@ -2,12 +2,31 @@
Codename: Shodan Codename: Shodan
Version: 1 Version: 1
A small operating system, written from scratch in Zig — a bootloader (`boot/`) A small resilient operating system, written from scratch in Zig.
and a microkernel (`system/kernel/`), sharing a neutral handoff contract (`system/boot-handoff.zig`).
It boots x86-64 via UEFI, and so far has a framebuffer console, a physical frame ## Zen of DanOS:
allocator, its own paging with W^X permissions, interrupt/exception handling, a
LAPIC timer, a kernel heap, a fixed-priority preemptive scheduler, and in-kernel IPC - Resilient Micro-Kernel Architecture.
channels. See [`docs/`](docs/README.md) for how each piece works. - Every process run in an isolated user space not kernel space.
- Processes cannot take down the entire OS with it when they die or is killed
- Stable public runtime library, private OS ABI.
- Keeps a stable runtime for user space processes between OS versions (great for backwards compatibility)
- Allows the underlying OS to be changed without effecting applications
- Provides a boundary to enable compatibility between OS's e.g. POSIX, MUSL etc
- Drivers are just isolated processes in user space.
- Thin binaries that can be restarted like applications.
- Useful during driver development.
- Drivers can claim MMIO / ports
- Driver resources (e.g. IRQ/Port/MMIO) claims are automatically cleaned up if the driver dies or is killed
- Drivers can also hook into the process lifecyle to clean up or reset hardware
- No legacy to deal with
- Zig code uses a clean coding style (Zen of Zig)
- Favor reading code over writing code.
- No magic numbers.
- No shortend names unless its for ABI compatibility or acronyms
- Inter-Process Communication (IPC)
- Publish and subscribe to Asynchronous Messages
- Talk to services and processes synchronously
## Prerequisites ## Prerequisites
@@ -60,7 +79,7 @@ straight into CI.
## Documentation ## Documentation
Design notes explaining the *why* behind the code live in Design notes explaining *why* behind the code live in
[`docs/`](docs/README.md) — start with [`docs/README.md`](docs/README.md). [`docs/`](docs/README.md) — start with [`docs/README.md`](docs/README.md).
## Logo ## Logo
+20 -1
View File
@@ -328,9 +328,10 @@ pub fn build(b: *std.Build) void {
const vfstest_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "vfs-test", "system/services/vfs/vfs-test.zig"); const vfstest_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "vfs-test", "system/services/vfs/vfs-test.zig");
const hpet_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "hpet", "system/drivers/hpet/hpet.zig"); const hpet_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "hpet", "system/drivers/hpet/hpet.zig");
const bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "bus", "system/drivers/bus/bus.zig"); const bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "bus", "system/drivers/bus/bus.zig");
const ps2_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "bus", "system/drivers/ps2-bus/ps2-bus.zig"); const ps2_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-bus", "system/drivers/ps2-bus/ps2-bus.zig");
const ps2_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-keyboard", "system/drivers/ps2-bus/keyboard.zig"); const ps2_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-keyboard", "system/drivers/ps2-bus/keyboard.zig");
const ps2_mouse_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-mouse", "system/drivers/ps2-bus/mouse.zig"); const ps2_mouse_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-mouse", "system/drivers/ps2-bus/mouse.zig");
const usb_xhci_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-xhci-bus", "system/drivers/usb-xhci-bus/usb-xhci-bus.zig");
const device_manager_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-manager", "system/services/device-manager/device-manager.zig"); const device_manager_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-manager", "system/services/device-manager/device-manager.zig");
// The input service and its exercisers: the fan-out server, a hardware-free synthetic // The input service and its exercisers: the fan-out server, a hardware-free synthetic
// source, and a subscriber that doubles as the `input` test's oracle. See docs/input.md. // source, and a subscriber that doubles as the `input` test's oracle. See docs/input.md.
@@ -360,6 +361,8 @@ pub fn build(b: *std.Build) void {
mk_run.addFileArg(ps2_keyboard_exe.getEmittedBin()); mk_run.addFileArg(ps2_keyboard_exe.getEmittedBin());
mk_run.addArg("ps2-mouse"); mk_run.addArg("ps2-mouse");
mk_run.addFileArg(ps2_mouse_exe.getEmittedBin()); mk_run.addFileArg(ps2_mouse_exe.getEmittedBin());
mk_run.addArg("usb-xhci-bus");
mk_run.addFileArg(usb_xhci_bus_exe.getEmittedBin());
mk_run.addArg("device-manager"); mk_run.addArg("device-manager");
mk_run.addFileArg(device_manager_exe.getEmittedBin()); mk_run.addFileArg(device_manager_exe.getEmittedBin());
mk_run.addArg("input"); mk_run.addArg("input");
@@ -384,6 +387,7 @@ pub fn build(b: *std.Build) void {
.{ ps2_bus_exe, "system/drivers" }, .{ ps2_bus_exe, "system/drivers" },
.{ ps2_keyboard_exe, "system/drivers" }, .{ ps2_keyboard_exe, "system/drivers" },
.{ ps2_mouse_exe, "system/drivers" }, .{ ps2_mouse_exe, "system/drivers" },
.{ usb_xhci_bus_exe, "system/drivers" },
}) |entry| { }) |entry| {
const step = b.addInstallArtifact(entry[0], .{ .dest_dir = .{ .override = .{ .custom = entry[1] } } }); const step = b.addInstallArtifact(entry[0], .{ .dest_dir = .{ .override = .{ .custom = entry[1] } } });
b.getInstallStep().dependOn(&step.step); b.getInstallStep().dependOn(&step.step);
@@ -455,6 +459,19 @@ pub fn build(b: *std.Build) void {
const run_efi = b.addSystemCommand(&.{ const run_efi = b.addSystemCommand(&.{
"qemu-system-x86_64", "qemu-system-x86_64",
"-device",
"qemu-xhci,id=xhci",
"-device",
"usb-mouse,bus=xhci.0",
"-device",
"usb-kbd,bus=xhci.0",
// "-usb",
// "-device",
// "usb-ehci,id=ehci",
// "-device",
// "usb-tablet,bus=usb-bus.0",
// "-device",
// "usb-mouse,bus=ehci.0",
"-machine", "-machine",
"q35", "q35",
"-m", "-m",
@@ -513,6 +530,8 @@ pub fn build(b: *std.Build) void {
"system/devices/device-abi.zig", "system/devices/device-abi.zig",
"system/devices/pci-class.zig", // class/subclass/prog-IF name decoding "system/devices/pci-class.zig", // class/subclass/prog-IF name decoding
"system/devices/acpi-ids.zig", // _HID name decoding "system/devices/acpi-ids.zig", // _HID name decoding
"system/devices/usb-abi.zig", // wire sizes + bit packings + set-up packet encodings
"system/devices/usb-ids.zig", // class/subclass/protocol code assignments
"library/mmio/mmio.zig", // barriers assemble + registers round-trip "library/mmio/mmio.zig", // barriers assemble + registers round-trip
"system/drivers/ps2-bus/scancode.zig", // set-2 decode + keyboard state machine "system/drivers/ps2-bus/scancode.zig", // set-2 decode + keyboard state machine
"system/drivers/ps2-bus/mouse-packet.zig", // 3-byte mouse packet assembly "system/drivers/ps2-bus/mouse-packet.zig", // 3-byte mouse packet assembly
+13 -2
View File
@@ -52,11 +52,22 @@ rather than restate it. Roughly in the order things happen at runtime:
microkernel's `ps`/`kill`/SIGCHLD: enumerate as a table snapshot, the microkernel's `ps`/`kill`/SIGCHLD: enumerate as a table snapshot, the
supervision link as the kill authority, and child-exit notifications over the supervision link as the kill authority, and child-exit notifications over the
same endpoints IRQs arrive on. same endpoints IRQs arrive on.
16. **[input.md](input.md) — the input module.** Broadcasting input events (keyboard, 16. **[process-lifecycle.md](process-lifecycle.md) — the process lifecycle.** Design:
signals over IPC as the one lifecycle vocabulary every process speaks — the
POSIX.1-1990 words with message delivery instead of stack hijack, the stable
`runtime.process` interface, exit reasons, published exit events any stateful
service can subscribe to (the VFS releasing dead clients' handles), and the two
iron rules (cleanup is the kernel's job; kill is not a signal).
17. **[device-manager.md](device-manager.md) — the device manager.** Design: the
tree, the matcher, and the supervisor. Tree structure lives in the manager,
authority stays in the kernel; bus drivers report what they see; drivers are
restarted through the lifecycle vocabulary — the plan that turns
[resilience.md](resilience.md)'s restart goal into increments.
18. **[input.md](input.md) — the input module.** Broadcasting input events (keyboard,
mouse, joystick): why a synchronous rendezvous can't fan out to many listeners, the mouse, joystick): why a synchronous rendezvous can't fan out to many listeners, the
asynchronous `ipc_send` primitive built to fix it, and the per-device subscribe/publish asynchronous `ipc_send` primitive built to fix it, and the per-device subscribe/publish
service layered on top. service layered on top.
17. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and 19. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
how `while (true) hlt` parks the CPU safely once there's nothing left to do. how `while (true) hlt` parks the CPU safely once there's nothing left to do.
Start with the north star: Start with the north star:
+16
View File
@@ -150,3 +150,19 @@ input output" in code — that expansion is what the acronym *is for*. But `msg`
test for "is this an abbreviation I must expand" is simply: *is there a longer word this test for "is this an abbreviation I must expand" is simply: *is there a longer word this
is a clipped form of?* If yes, write the word. If it's an initialism standing in for a is a clipped form of?* If yes, write the word. If it's an initialism standing in for a
phrase, leave it. phrase, leave it.
## Zen of Zig
* Communicate intent precisely.
* Edge cases matter.
* Favor reading code over writing code.
* Only one obvious way to do things.
* Runtime crashes are better than bugs.
* Compile errors are better than runtime crashes.
* Incremental improvements.
* Avoid local maximums.
* Reduce the amount one must remember.
* Focus on code rather than style.
* Resource allocation may fail; resource deallocation must succeed.
* Memory is a resource.
* Together we serve the users.
+144
View File
@@ -0,0 +1,144 @@
# The device manager
**Status: design.** The primitives this builds on are real ([process-management.md](process-management.md):
spawn/supervise/kill/exit-notification; [driver-model.md](driver-model.md): the device
table as a capability system; [drivers.md](drivers.md): claim/map/IRQ), and the first
per-device driver spawn works (the device manager matches the xHCI controller by PCI
class and spawns `usb-xhci-bus` with the device id as argv[1]). This document designs
the rest: the device manager as **the tree, the matcher, and the supervisor** — the
policy process that turns [resilience.md](resilience.md)'s restart goal into practice
for drivers.
How processes stop, reload, and report their deaths is deliberately **not** in this
document: that is the universal lifecycle every danos process speaks —
[process-lifecycle.md](process-lifecycle.md), signals over IPC and the stable
`runtime.process` interface. The device manager is that design's first serious
customer, not its owner. Its own protocol contains nothing lifecycle-shaped; a
driver is stopped, health-checked, and buried exactly like any other process.
## The tree: structure in the manager, authority in the kernel
The device tree is two things fused: *information* (what exists, how it nests) and
*authority* (a descriptor is a licence to map physical memory). They separate:
- The **kernel keeps the capability system** — device, I/O-port, and interrupt
claims, resource containment on `device_register`, the
`mmio_map`/`irq_bind`/`msi_bind` gates — and **cleans all of it up when a process
dies** (settled; it is increment 1 of
[process-lifecycle.md](process-lifecycle.md)). The three invariants in
[driver-model.md](driver-model.md) stay exactly where they are. A device manager
that could mint MMIO mappings by its own say-so would be a second kernel, and a
buggy one would un-earn everything the microkernel bought.
- The **device manager owns the tree as data** — identity, topology, naming, driver
matching, hotplug events, and being the one process everything else asks about
devices. Firmware discovery seeds it (today via the kernel's snapshot); **bus
drivers grow it** by reporting what they see; applications query and watch it.
`device_enumerate` fades to a manager-internal (then deleted) seam.
Long-term, discovery itself leaves the kernel — but not *into* the manager. PCI
enumeration is a **pci-bus driver**: the manager spawns it against the host bridge
(already a device with the ECAM window as a resource), it scans, it reports functions
like any bus reports children. ACPI becomes an **acpi service** that interprets the
tables and reports the namespace. The manager only orchestrates and merges. Moving
AML interpretation out of ring 0 is its own project on its own track; nothing here
depends on when it lands.
## The protocol
A `device-manager-protocol` module (the vfs-protocol pattern): extern-struct
messages, a version in the handshake, reserved fields everywhere. The manager is a
well-known endpoint (`ipc.register(.device_manager)`); the badge tells it who is
talking; the same endpoint receives its children's exit notifications — one loop,
one world.
| Direction | Message | Purpose |
|---|---|---|
| driver → manager | `hello { version, role, device_id }` | confirms the argv assignment, starts the deadline clock |
| bus → manager | `child_added { parent, identity, resources }` | one node the bus discovered |
| bus → manager | `child_removed { id }` | unplug, or the bus lost it |
| app → manager | `enumerate` | snapshot of the tree (read-only) |
| app → manager | `subscribe` | receive published add/remove events |
`hello` is the one deadline the manager enforces itself: spawned and silent past the
deadline means wrong binary, wrong protocol version, or wedged before main — apply
the stop sequence and the restart policy. Everything else lifecycle-shaped
(terminate, the common `ping` liveness call, exit reasons) arrives through
[process-lifecycle.md](process-lifecycle.md)'s vocabulary, not this protocol.
Assignment stays argv (`usb-xhci-bus <device id>`) for now — simple, and it works.
The step after `hello` exists is delegation: the manager claims (or is granted) the
devices and passes the claim to the driver over IPC (the M13 capability-transfer
mechanism), replacing first-come-first-served `device_claim` with policy. Identity in
`child_added` is per-bus: PCI children carry the class triple (`pci_class`, as the
xHCI match already uses); USB children carry the (class, subclass, protocol) triple
from usb-ids.zig — each bus's native language, decoded by the shared ids modules.
## Supervision and restart
Every driver is spawned with the manager's exit endpoint (`spawnSupervised` — built).
On a death notification:
1. **Read the reason** ([process-lifecycle.md](process-lifecycle.md) increment 2).
Clean exit → it meant to; don't restart. Fault or missed `hello` deadline →
restart with **backoff**, and a crash-loop cap (three fast deaths → mark failed,
stop respawning, log loudly; a later `reload` to the manager can retry).
2. **Prune the subtree** the dead bus driver reported. Its children describe
protocol state (xHCI slot ids, transfer rings) that died with the process;
keeping the nodes would be keeping a lie. Watchers receive `child_removed` — the
input service losing, then regaining, a keyboard is the *honest* description of
what happened. The restarted instance rediscovers and re-reports.
3. **The claim is already free** because the kernel released it at death — the
restarted instance claims the same controller and comes up.
Who supervises the supervisor: **init** (PID 1), which already supervises the
services it starts. If the manager dies, drivers keep running (they hold their
claims; the kernel doesn't care who their supervisor was — though their exit
notifications now dangle harmlessly). The restarted manager re-learns the world:
kernel snapshot, then a re-`hello` round — drivers answer a broadcast or are stopped
and respawned. Full state handoff is deliberately not attempted.
## Thin drivers, class protocols
The [driver-model.md](driver-model.md) three-shape split, restated as processes:
- A **bus driver** (usb-xhci-bus) owns its controller — claim, MMIO, IRQ/MSI, DMA
rings — and offers a *transfer* protocol ("submit a control transfer to device N",
built from the usb-abi request constructors) plus tree reports to the manager.
- A **class driver** (usb-hid, usb-storage) owns nothing: it is matched to a reported
child by its identity triple, speaks the bus's transfer protocol downward and its
service's protocol upward — HID reports to the input service, blocks to the block
service. It works unchanged over any controller.
- **Services** (input, display, block) aggregate class drivers and face applications.
Each arrow is a protocol module. The manager routes none of the data plane — it
introduces the parties (matching), supervises them (lifecycle), and gets out of the
way.
## Increments
Increments 1–4 are the lifecycle prerequisites and live in
[process-lifecycle.md](process-lifecycle.md) (claim cleanup on death, exit reasons,
published exit events, signals + `runtime.process`). On top of those:
5. **device-manager-protocol**: `hello`, supervised spawn with restart policy;
usb-xhci-bus becomes the first conforming driver.
6. **Tree reports**: `child_added`/`child_removed`; the manager mirrors; xHCI reports
the mouse and keyboard QEMU already hangs off it.
7. **App surface**: `enumerate`/`subscribe` over IPC; `device_enumerate` retreats
to a manager-internal seam.
8. **Discovery migration**: pci-bus driver first, acpi service second, kernel scan
retired last. (AML-in-user-space is its own track.)
## Settled questions (2026-07-12)
- **Stateful buses**: pruning the subtree on bus-driver death is right for USB. A
future storage bus with in-flight writes wants drain-before-terminate — which is
exactly the `deadline_ms` parameter `stop()` already has; a per-driver deadline
is one value in the manager's policy table when such a bus arrives. No design
change.
- **Manager death**: drivers survive the manager; the restarted manager re-learns
the world (above). Checkpointing driver state with the manager is deferred until
something demonstrates the need.
- **Matching stays code until the third bus.** `driverFor`/`pciDriverFor` are
honest at two bus types; the third triggers the manifest (a driver declares what
it binds: a PCI class triple, a USB class triple, an ACPI `_HID`).
+19
View File
@@ -103,3 +103,22 @@ This is what makes a user-space driver possible at all, and it's the subject of
the oldest (discrete messages, not a coalescing level like the notification ring). the oldest (discrete messages, not a coalescing level like the notification ring).
- **A bounded reply.** `MSG_MAX` is 256 bytes and the copy runs under the big kernel - **A bounded reply.** `MSG_MAX` is 256 bytes and the copy runs under the big kernel
lock; a bulk transfer wants shared pages, not a copy. lock; a bulk transfer wants shared pages, not a copy.
## Lifecycle conventions over IPC (M17)
Three conventions from [process-lifecycle.md](process-lifecycle.md) ride the
notification mechanism:
- **Signals** arrive as notifications on the endpoint a process nominated with
`signal_bind` (`runtime.process.bindSignals`): badge = the signal bit plus the
coalesced pending mask (`runtime.process.signalsFrom` decodes). Statements,
never questions; no payload, no reply.
- **One-shot timers** (`timer_bind`, `runtime.system.timerOnce`) land as a
timer-bit notification — the timed wait: a service arms a deadline and keeps
serving, instead of blocking in sleep.
- **The universal ping**: a **zero-length request is the liveness probe**,
answered with a zero-length reply by the service harness itself
(`runtime.service.run`). No protocol's requests start at length zero, so the
encoding cannot collide, and a wedged service simply fails to answer — which
is the diagnosis. Deep health ("can I reach my hardware?") stays a per-service
protocol message.
+191
View File
@@ -0,0 +1,191 @@
# M17–M18 execution plan: process lifecycle + device manager
The operational plan for building [process-lifecycle.md](process-lifecycle.md)
(M17) and [device-manager.md](device-manager.md) increments 5–7 (M18). Design is
settled in those documents; this file is the build order — one phase at a time,
each phase green before the next starts. Delete or archive this file when M18
lands.
**Definition of green, every phase:** `zig build` clean, `zig build test` clean,
`python3 test/qemu_test.py` passes (existing scenarios plus the phase's new one),
and the relevant design doc's "known gaps" / status lines updated. Commit per
green phase (no co-author trailers).
**Workflow (settled 2026-07-12):** work happens in a dedicated git worktree, on
feature branches cut from `main` — `feat/process-lifecycle` (M17.1–17.4),
`feat/device-manager` (M18.1), `feat/usb-xhci-bus` (M18.2–18.3). When a branch's
phases are all green it is **auto-merged into `main`**; branches are kept after
merge, not deleted. Merges and branches are pushed to origin. Phase 0 (once):
commit the design docs, merge the outstanding `feat/usb` work into `main`, and
run the existing QEMU suite green before any new work starts.
**Numbering note:** continues the milestone sequence (driver track ended at M16).
## Status
The loop marks a phase `[x]` in the same commit that lands it. A phase is marked
only when its definition of green holds.
- [x] **Phase 0** — baseline: docs committed, feat/usb merged to main, pushed;
`usb-xhci-libary.zig` renamed to `usb-xhci-library.zig`; existing QEMU
suite green from the worktree (48/48, 2026-07-12).
- [x] **M17.1** — kernel releases claims/MSI on death (claims: `releaseAllOwnedBy`
in the reap; MSI was already swept by `irq.releaseOwner`; `claim-release`
test; suite 49/49)
- [x] **M17.2** — exit reasons (`ExitReason` recorded at exit/fault/kill before
the notification; `process_exit_reason` supervisor-gated;
`runtime.process.exitReason`; kernel + ring-3 assertions; suite 49/49)
- [x] **M17.3** — published exit events + VFS subscriber (`process_subscribe`,
bounded ref-counted table, publish on every death;
`runtime.process.subscribeExits`; VFS handles carry owners and are swept on
the owner's death; `vfs-client-death` test; suite 50/50)
- [x] **M17.4** — signals, timer notifications, `runtime.process`, the service
harness (signal_bind/process_signal + coalescing pending mask; timer_bind
on the tick; bindSignals/signalsFrom/sendSignal/stop + timerOnce;
runtime.service.run with the zero-length ping; VFS converted; `signals`
scenario; suite 51/51)
- [ ] **merge** `feat/process-lifecycle` → main, push
- [ ] **M18.1** — device-manager protocol: hello + restart policy (branch `feat/device-manager`)
- [ ] **merge** `feat/device-manager` → main, push
- [ ] **M18.2** — xHCI port scan + tree reports (branch `feat/usb-xhci-bus`)
- [ ] **M18.3** — app surface: enumerate/subscribe + device-list
- [ ] **merge** `feat/usb-xhci-bus` → main, push — **loop ends here**
---
## M17.1 — the kernel releases a dead process's claims
The cleanup half of iron rule 1; the prerequisite for every restart story.
- `system/kernel/devices-broker.zig`: `releaseAllOwnedBy(owner: u32)` — clear
every `claimed[]` slot holding this task id.
- `system/kernel/process.zig`: call it from the reap path, alongside the existing
IRQ-binding release (the ordering comment there says why IRQs go first — claims
slot in after them, before the exit notification).
- MSI vectors: find where `msi_bind` records per-device vectors (interrupts
module) and release those by owner in the same pass.
- Docs: remove the claims bullet from process-management.md "Known gaps".
**Test:** new QEMU scenario `claim-release` — a test child claims an unclaimed
device, is killed, is respawned, and claims the same device again successfully;
assert both claims in the serial log. Kernel-side unit coverage in
`system/kernel/tests.zig` for `releaseAllOwnedBy` (claim two devices as two owners,
release one owner, verify exactly its claims freed).
## M17.2 — exit reasons
- `system/abi.zig`: `ExitReason` (exited, aborted, segmentation_fault,
illegal_instruction, arithmetic_fault, killed).
- Kernel: record the reason at every death site — clean exit path, each fault
class in `onException`, the kill path. Bounded recent-exits table (ids are never
reused, so a small ring keyed by id is enough).
- New system call `process_exit_reason(id)` — supervisor-gated, like kill; returns
the recorded reason or `-ESRCH` once evicted.
- `library/runtime/process.zig`: `ExitReason` + `exitReason(id: u32)`.
- Docs: remove the no-exit-status bullet from process-management.md.
**Test:** extend the `supervision` scenario — three children: one exits cleanly,
one faults (the fault-recovery pattern), one is killed; the supervisor asserts all
three reasons.
## M17.3 — published exit events
- Kernel: bounded subscriber table (endpoints); new system call
`process_subscribe(endpoint)` (ungated, like `process_enumerate`); every death
posts `notify_exit_bit | id` to each subscriber — the same post the supervisor
path already uses.
- `library/runtime/process.zig`: `subscribeExits(endpoint)`.
- VFS becomes the first subscriber: on an exit event, release every handle keyed
by that task id (badges already are task ids). Log the release.
- Docs: note the convention in ipc.md (exit events reuse the exit-notification
badge encoding).
**Test:** new QEMU scenario `vfs-client-death` — a client opens a file and is
killed without closing; assert the VFS logs the handle release and its open-handle
count returns to baseline.
## M17.4 — signals and the service harness
- Kernel: per-task pending mask + bound endpoint; system calls
`signal_bind(endpoint)` and `process_signal(id, signal)` (supervisor-or-self
gated); delivery posts `notify_signal_bit | pending mask`, coalescing; pending
signals with no bound endpoint pend silently.
- `library/runtime/process.zig`: `Signal`, `SignalSet`, `bindSignals`,
`signalsFrom`, `sendSignal`, `stop(id, deadline_ms)` (terminate → wait for exit
notification → kill). Implement `terminate`, `reload`, `user_1`, `user_2`;
`interrupt`/`quit` are enum members with no sender yet; `alarm` stays unbuilt.
- Kernel: **one-shot timer notifications** — `timer_bind(endpoint, ms)` posts a
notification badge when the deadline lands (IRQ-as-IPC again, on the timer
wheel `sleep` already uses). This is the missing timed-wait primitive:
`replyWait` blocks forever and `sleep` blocks the whole process, but `stop()`'s
escalation, the device manager's `hello` deadline (M18.1), and restart backoff
all need a deadline while staying responsive. It is also the mechanism `alarm`
gets for free later.
- New `library/runtime/service.zig`: the harness — `run(callbacks)` owning the
replyWait loop, folding protocol messages, signals, and child-exit notifications
into `init` / `on_message` / `on_reload` / `on_terminate`; answers the common
`ping` automatically. Define the reserved `ping` request encoding here and
document it in ipc.md (one obvious encoding; smallest that cannot collide with
existing protocols).
- Convert one existing service (input-source or hpet) to the harness as proof it
subtracts code rather than adding it.
**Test:** extend `supervision` — a harness-built child: `sendSignal(reload)`
observed in its log, `ping` answered, `stop()` produces a clean exit with reason
`exited`; a second child that ignores signals (no bind) is killed by `stop()`'s
deadline with reason `killed`.
## M18.1 — device-manager protocol: hello + restart policy
- New `system/services/device-manager/device-manager-protocol.zig` module
(vfs-protocol pattern): `hello { version, role, device_id }`; version constant;
reserved fields.
- Device manager: register the `.device_manager` endpoint; spawn drivers with its
exit endpoint; enforce the hello deadline; restart policy — backoff, crash-loop
cap (three fast deaths → mark failed, log, stop), reasons from M17.2 deciding
restart vs not.
- usb-xhci-bus: adopt the harness + send hello. hpet/ps2-bus follow only if the
conversion is mechanical; otherwise they keep working unconverted (the manager
only enforces hello on drivers spawned with an assignment).
- build.zig: test-loop entry for the protocol module if it grows pure logic.
**Test:** new QEMU scenario `driver-restart` — the xHCI driver takes a test-only
argv flag to fault after hello on its first run; assert: fault, exit reason
recorded, manager respawns with backoff, second run claims the controller
(M17.1) and hellos clean. Assert the crash-loop cap by a driver that always
faults (a tiny test driver, not xhci).
## M18.2 — bus tree reports
- Protocol: `child_added { parent, identity, resources }` / `child_removed { id }`.
- usb-xhci-bus: bring-up to **port scan only** — map the MMIO window (claimed in
M16-era work), controller reset/start per xHCI spec, walk the port registers,
report one `child_added` per connected port with speed + port number as
identity. **No transfer rings, no descriptors** — reading device/interface
descriptors (and therefore USB class triples for matching) is the follow-on USB
track, not this plan.
- Device manager: mirror reports into its tree; prune the subtree (emitting
`child_removed`) when a bus driver dies; assert re-report on restart.
**Test:** QEMU already attaches usb-kbd + usb-mouse on xhci.0 — assert two
`child_added` events reach the manager and appear in its tree dump; kill the
driver, assert two `child_removed` then two fresh `child_added` after respawn.
## M18.3 — the application surface
- Protocol: `enumerate` (tree snapshot) + `subscribe` (published add/remove
events, input-service pattern).
- A small client (`device-list`, the `ps` analog) exercising both; the manager
becomes the one answer to "what devices exist" for user space.
`device_enumerate` stays for drivers/kernel seeding — its retreat is tied to the
discovery migration, out of this plan.
**Test:** QEMU scenario — `device-list` shows the tree including USB children;
during a driver restart the subscribing client logs remove + add events.
---
**Explicitly out of scope** (own tracks, after M18): discovery migration (pci-bus
driver, acpi service, retiring the kernel scan), USB control transfers +
descriptors + class-driver matching, the musl layer, `interrupt`/`quit` senders
(needs a console), job control.
+327
View File
@@ -0,0 +1,327 @@
# Process lifecycle: signals over IPC
**Status: increments 1–4 built** (2026-07-12): claim release on death, exit
reasons, published exit events, and signals + one-shot timers + the service
harness are all in — the interface below is as-built. The primitives underneath
predate this design ([process-management.md](process-management.md):
spawn, the supervision link, kill, child-exit notifications); this document designs
the layer above them — the standard vocabulary a danos process speaks about its own
life, and the stable `runtime.process` interface that carries it. Nothing here is
device- or driver-specific: a driver, the VFS, and a user application all stop,
reload, and die the same way. The device manager is simply this design's first
serious customer ([device-manager.md](device-manager.md)).
**"POSIX" in this document means the concepts, never the letter of the standard.**
danos borrows the ideas and the hard-won lessons (what SIGTERM *means*, why SIGPIPE
was a mistake) without inheriting the mechanism, the API, or the names. The naming
rule is danos's own and it is strict: plain words that communicate intent
(`terminate`, `reload`, `exited`) and the IPC vocabulary the system already speaks
(`bind`, `subscribe`, `publish`, `endpoint`) — never `SIG*`, never a second word for
a concept that already has one. Literal POSIX arrives later and lives elsewhere: a
**musl-based C layer** (growing out of library/posix) that wires C programs to the
danos runtime — musl's syscall surface retargeted at danos system calls and IPC
protocols (files onto the VFS protocol, `sigaction`/`wait` onto this lifecycle,
sockets onto whatever networking becomes). Ported programs see POSIX; the system
underneath never does.
## Why a standard vocabulary
A supervisor can only manage processes it has never heard of if "please exit" means
the same thing to all of them. That is the one thing POSIX signals got deeply right:
`SIGTERM` means the same thing to nginx and to a five-line script, which is why
process supervision on Unix (init systems, container runtimes) is possible at all.
danos wants that property from day one, because supervision-and-restart is the
system's core motivation ([resilience.md](resilience.md)).
What POSIX got wrong — for a system like this — is the **delivery mechanism**:
asynchronous control-flow hijack. A Unix handler runs on a stolen stack at an
arbitrary instruction boundary, which is why the async-signal-safe function list
exists, why `errno` must be saved, and why the canonical signal bug is a SIGTERM
handler innocently calling `printf` mid-`malloc`. That entire bug class comes from
the mechanism, not the vocabulary, and none of it is worth importing.
A microkernel already has the right channel: **a signal is a message.** QNX delivers
POSIX signals over its message passing; seL4 has notification objects; Erlang turned
"death is a message to whoever linked" into a reliability philosophy. danos has
already done it once without naming it: a child's death arrives as a notification
badge on the supervisor's endpoint — the microkernel's SIGCHLD, the IRQ-as-IPC
pattern reused. Signals are the same pattern reused a third time.
## The mechanism
- **`signal_bind(endpoint)`** — a process nominates the endpoint its signals arrive
on, exactly as `irq_bind` nominates where a device's interrupts land. The runtime
does this at startup for any program that opts in.
- **`process_signal(id, signal)`** — posts the signal as an asynchronous
notification to the target's bound endpoint: badge = `notify_badge_bit |
notify_signal_bit | pending signals`. Non-blocking for the sender, always.
- **Pending signals coalesce** in a per-process bitmask until the target next waits
— exactly like interrupt notifications, and exactly POSIX's own semantics for
non-realtime signals (two pending SIGTERMs are one SIGTERM). The bitmask *is* the
design: signals carry no payload. Anything with a payload is a protocol message.
- **Authority**: the supervisor may signal its children — the same link that is
already the kill authority. A process may signal itself. Anything broader waits
for transferable process handles.
- **No binding, no problem**: a process that never calls `signal_bind` is not
broken — its signals pend unread and only `process_kill` works on it. Simple
programs stay simple; the vocabulary is opt-in, the kill authority is not.
Because delivery is a message into the process's own event loop, there is no
async-signal-safe list in danos: a handler is ordinary code running at a point the
process chose. The bug class is gone by construction, not by discipline.
## The vocabulary: POSIX.1-1990, sorted honestly
The full 1990 set, and what each becomes. Two intrinsically problematic cases get a
defense below the table.
| POSIX.1-1990 | danos disposition | Notes |
|---|---|---|
| SIGTERM | signal `terminate` | finish up and exit; the supervisor's polite half |
| SIGHUP | signal `reload` | re-read configuration / re-scan |
| SIGINT | signal `interrupt` | interactive interrupt; meaningful once a console can send it, in the vocabulary now so numbering is stable |
| SIGQUIT | signal `quit` | as SIGINT, without the core-dump baggage |
| SIGALRM | signal `alarm` | timer expiry as a message; the Unix SIGALRM+`longjmp` timeout hacks are impossible here. In the vocabulary, unbuilt: no consumer yet, and when one appears it is runtime sugar over the existing timer — zero kernel work |
| SIGUSR1, SIGUSR2 | signals `user_1`, `user_2` | service-defined |
| SIGCHLD | **already exists** — the exit notification | the badge carries the child id, dodging the classic coalescing bug (Unix code must loop `waitpid`) |
| SIGKILL | `process_kill` — kernel mechanism | its definition is "cannot be handled"; it was never really a signal |
| SIGABRT | exit reason `abort` | `abort()` is synchronous self-termination, not an event |
| SIGSEGV, SIGILL, SIGFPE | exit reasons, **never delivered** | see below |
| SIGPIPE | **an error return**, not a signal | see below |
| SIGSTOP, SIGTSTP, SIGTTIN, SIGTTOU, SIGCONT | deferred | job control needs terminals, sessions, and process groups; stop/continue is scheduler territory |
**The fault signals (SIGSEGV, SIGILL, SIGFPE) are intrinsically wrong for messages.**
They are *synchronous* — raised at a specific faulting instruction, not "sometime
soon". A message cannot be delivered to a process whose next instruction re-faults;
it never reaches its event loop to read it. POSIX only makes fault handlers "work"
via the async hijack (run the handler *instead of* the instruction), and even there,
returning from a SIGSEGV handler without curing the cause is undefined behavior.
danos's architecture already has the better answer: fault → the kernel kills the
process ([resilience.md](resilience.md) step 2, built) → the supervisor reads the
reason → restart. Recovery is restart, not a handler. This is also truer to the 1990
standard than handling is: the standard's default action for all three was
"terminate the process".
**SIGPIPE deserves special contempt.** Its default kills a process that writes to a
closed pipe — which is why "the whole server died because one client disconnected"
is roughly every network daemon's first production bug, and why every mature codebase
contains the same fix: ignore SIGPIPE, handle the `EPIPE` error return. danos made
the right choice natively already — a reply owed to a dead peer fails with `-EPEER`.
Errors from operations are error returns from those operations. The posix layer can
synthesize SIGPIPE for ported code that expects it.
### Statements, not questions
A signal and a protocol message both travel over IPC — the difference is the
**contract**, not the transport. danos IPC has two primitives, both already in
daily use: the **asynchronous notification** (a badge — bits that coalesce into a
pending mask; the sender never blocks; no payload, *no reply path*; how IRQs and
exit events arrive) and the **synchronous call** (a rendezvous — payload both
ways, the caller waits for the reply; how VFS requests work). A signal is the
first kind: a *statement*. `terminate` wants no reply — the exit notification is
its acknowledgement.
A health probe is the second kind: a *question*, worthless without its answer —
and the answer's absence within a deadline is the very thing being measured.
Asked as a signal it has no reply channel (a coalescing bit can't carry an answer,
and the authority rule forbids a child signalling its supervisor back); asked as a
call, the timeout-is-the-diagnosis semantics come free. So there is no `health`
signal. Liveness is the common **`ping`**: a reserved request every harness-run
service answers automatically on its main endpoint — still free for the service
author, still one obvious way — and a supervisor's probe is a `ping` call with a
deadline.
## The two iron rules
1. **Cleanup is the kernel's job.** A process can die with no warning — fault,
kill, power. Correctness must never depend on a `terminate` handler running. On
any death the kernel releases the address space, IPC handles, IRQ bindings, and
owed replies (built), and must also release **device, I/O-port, and interrupt
claims and MSI vectors** (the known gap in
[process-management.md](process-management.md); increment 1). A signal handler is
for *graceful* work — flushing, deregistering, saving — never for *necessary*
work.
2. **Kill is not a signal, and exit reasons are load-bearing.** The standard stop
sequence is *terminate → deadline → `process_kill`*; the unhandleable kill stays
a kernel mechanism. And a supervisor deciding whether to restart must know *how*
the child died: clean exit (meant to — don't restart), fault (restart with
backoff), killed (the supervisor did it). The exit notification today carries
only the id; it grows a reason. Restart policy cannot be written without it.
## Who learns of a death
A death has three audiences, and conflating them is how systems end up with either
zombie state or privileged snooping:
1. **The supervisor** — gets the exit notification on the endpoint it gave at spawn
(built), which grows the `ExitReason` (increment 2). The supervisor is the only
audience that needs the *reason*, because it is the only one deciding whether to
restart.
2. **The peer owed a reply** — already built: a client that dies mid-request fails
the server's reply with `-EPEER`; a server that dies fails its waiting clients
the same way. This covers the *synchronous* case only.
3. **The subscribers** — the new piece, and it is the input service's
publish/subscribe shape ([input.md](input.md)) applied to exits. A stateful
service accumulates per-client state across many requests: the VFS holds a dead
client's open file handles, the input service holds its subscriptions, a future
network stack holds its sockets. None of these are the client's supervisor, and
none learn anything from a failed reply if the client simply never calls again.
So the kernel **publishes every exit** to whoever subscribed:
`process_subscribe(endpoint)` adds a subscriber, and each death posts a
notification to every subscriber (badge = `notify_exit_bit | process id` — the
same encoding supervisors already decode, the IRQ-as-IPC pattern once more). The
subscriber filters for ids it holds state for and releases what the dead client
held. Correlating is free of bookkeeping: an IPC sender's badge already *is* its
task id (`runtime.ipc.Received`), so the id a service has been keying client
state by all along is the id the exit event carries.
Subscription, not broadcast-to-everyone: only processes that asked receive
events, the kernel keeps a bounded subscriber table, and delivery is the same
non-blocking coalescing notification as everything else — a dying process never
waits on its mourners. Subscribing is ungated, like `process_enumerate`: what is
running (and dying) is not a secret between cooperating processes. Subscribers
do not receive the exit reason — the VFS does not care *why* the client died.
This is the service-side mirror of iron rule 1: **a service must never depend on
its clients cleaning up after themselves.** Handle release on client death is the
service's job, triggered by the published exit event — never by a courtesy
"closing now" message that a crashed client will never send.
## The stable interface: `runtime.process`
`runtime.process` already owns what a process receives at birth (`Init`, the
argv contract). It grows to own the other end of life.
**The runtime is the stable interface; the numbers are not.** danos applications do
not make system calls — they call the runtime library, and the system-call numbers,
notification bits, and signal bit positions beneath it are a **private kernel ↔
runtime contract** that may change at any time (settled 2026-07-12). This is why
the runtime exists. Today kernel and runtime ship from one tree in one image, so
"stability" is simply building them together. When driver binaries start shipping
as separately-versioned applications — the whole point of the restart design — the
binary's embedded runtime version becomes compatibility metadata (the same idea as
the protocol version in the device manager's `hello`), and the kernel refuses what
it cannot serve. Signals therefore need no reserved numbering scheme: the enum
below is vocabulary, not ABI.
```zig
/// The signal vocabulary. The value is the bit position in the pending mask — a
/// private kernel/runtime detail, free to change while they ship together.
pub const Signal = enum(u5) {
terminate = 0, // SIGTERM: finish up and exit
reload = 1, // SIGHUP: re-read configuration
interrupt = 2, // SIGINT
quit = 3, // SIGQUIT
alarm = 4, // SIGALRM
user_1 = 5, // SIGUSR1
user_2 = 6, // SIGUSR2
};
/// A decoded pending mask: the coalesced set of signals a notification delivered.
pub const SignalSet = struct {
pending: u32,
pub fn has(set: SignalSet, signal: Signal) bool { ... }
pub fn iterate(set: SignalSet) Iterator { ... }
};
/// Nominate `endpoint` as this process's signal endpoint (signal_bind). The
/// runtime's service harness calls this; a bare program may call it directly and
/// fold signals into its own replyWait loop.
pub fn bindSignals(endpoint: usize) bool { ... }
/// Decode a received badge into signals, or null if the badge is not a signal
/// notification (mirrors ipc.Received.isChildExit).
pub fn signalsFrom(badge: usize) ?SignalSet { ... }
/// Send `signal` to process `id`. Supervisor-gated, like kill; non-blocking.
pub fn sendSignal(id: u32, signal: Signal) bool { ... }
/// The standard stop sequence: terminate, wait up to `deadline_ms` for the exit
/// notification, then process_kill. The one call a supervisor needs.
pub fn stop(id: u32, deadline_ms: u64) void { ... }
/// Subscribe `endpoint` to published exit events (process_subscribe). Every
/// process death posts an asynchronous notification: badge = notify_exit_bit |
/// process id — the same encoding a supervisor's exit notification uses, decoded
/// by the same ipc.Received helpers. For stateful services: release what the dead
/// client held (file handles, subscriptions, sockets). Ungated, like
/// process_enumerate.
pub fn subscribeExits(endpoint: usize) bool { ... }
/// How a process ended — queried after the exit notification (the kernel records
/// it first, so the two never race). What restart policy reads. (Built in M17.2.)
pub const ExitReason = enum(u8) {
exited, // returned from main / clean exit
aborted, // abort() — deliberate self-termination (SIGABRT's ghost; reserved)
segmentation_fault, // SIGSEGV's ghost
illegal_instruction, // SIGILL's ghost
arithmetic_fault, // SIGFPE's ghost
protection_fault, // general protection fault
fault, // any other CPU exception
killed, // process_kill
};
```
Two deliberate absences. There is no `mask`/`block` API — a process that is not
ready for a signal simply has not waited on its endpoint yet; the pending mask *is*
the blocked set. And there is no per-signal handler registration at this layer —
dispatch is the process's own `switch` over `SignalSet`, or the service harness's
callbacks (`on_terminate`, `on_reload`) for programs that want defaults.
### The service harness
`runtime.service` owns the `replyWait` loop and folds every event source — signals,
child exits, protocol messages — into callbacks, with the vocabulary's defaults:
`terminate` returns from the loop (clean exit), the common `ping` is answered automatically,
`reload` is ignored unless overridden. One loop, no locking, nothing reentrant. A
service author writes domain logic; the lifecycle contract is satisfied by the
harness. A process that bypasses the harness and ignores its signals meets the
deadline-then-kill escalation — you cannot force a process to implement an
interface, but you can make compliance free and non-compliance fatal.
### The musl layer later
The POSIX C layer is a **musl port**: musl's arch/syscall layer retargeted so that
what musl believes are kernel syscalls become danos runtime calls and IPC — `open`
and `read` onto the VFS protocol, `kill`/`sigaction`/`waitpid` onto this document's
vocabulary, `exit` onto the runtime's exit path. `sigaction` handlers registered
through it are invoked by the runtime's loop when the signal message arrives —
synchronous underneath, async-looking to ported code, delivered at wait boundaries
the way most Unix programs already experience signals (at syscalls). No stack hijack
ever happens, `SA_RESTART` semantics come free because nothing was interrupted, and
SIGPIPE can be synthesized from `-EPEER` for the programs that expect it. C programs
get POSIX; danos-native programs never pay for it.
## Increments
1. **Kernel: release device/port/IRQ claims and MSI vectors on death** — the
cleanup half of iron rule 1, and the prerequisite for any restart story. Test:
kill a claiming driver, spawn it again, the claim succeeds.
2. **Exit reason in the death notification** (`ExitReason` above).
3. **Exit events**: `process_subscribe` in the kernel (bounded subscriber table,
publishes on every death), `runtime.process.subscribeExits`; the VFS becomes the
first subscriber — releasing a dead client's handles is its proof test.
4. **Signals**: `signal_bind` + `process_signal` + the pending mask in the kernel;
`runtime.process` grows the interface above; the service harness handles
`terminate` and answers the common `ping`; `stop()` for supervisors.
[device-manager.md](device-manager.md) builds directly on all four.
## Settled questions (2026-07-12)
- **Signal numbering is not ABI**: the runtime is the stable interface; the numbers
beneath it are a private kernel ↔ runtime contract (see "The stable interface").
- **Liveness is a `ping` call, not a signal**: signals are statements, questions
are synchronous calls (see "Statements, not questions"). A service wanting *deep*
health ("can I reach my hardware?") defines its own protocol message on top.
- **Process handles: deferred.** Pids + the supervisor gate cover everything
planned; transferable handles (Fuchsia-style, delegating signalling without
delegating kill) wait for the capability table to grow types beyond endpoints.
- **`alarm`: in the vocabulary, unbuilt.** No consumer yet; when one appears it is
runtime sugar over the existing timer (arm a timer that posts your own signal) —
zero kernel work, so deferring costs nothing.
- **Subscription granularity: all exits**, subscriber-side filtering — one
subscription per service, a bounded kernel table. Per-id subscriptions only if
event volume ever matters (hundreds of processes, not before).
- **Client identity across the exit boundary: no convention needed** — an IPC
sender's badge already is its task id (see "Who learns of a death").
+13 -5
View File
@@ -95,11 +95,18 @@ the architecture layer calls up into `tick`.
## Known gaps (bring-up honesty) ## Known gaps (bring-up honesty)
- Device **claims** are not released on death (pre-existing: the fault path has - ~~Device claims are not released on death~~ Closed (M17.1): every path out of a
the same gap) — a killed driver's device stays claimed until reboot. process releases its device claims alongside its IRQ and MSI bindings
(`releaseTaskResourcesLocked`), so a restarted driver can claim its hardware
again — the cleanup half of [process-lifecycle.md](process-lifecycle.md)'s iron
rule 1. The `claim-release` test proves the kill → release → re-claim cycle.
- Kernel stacks of dead tasks are leaked, as on every exit path (no reaper yet). - Kernel stacks of dead tasks are leaked, as on every exit path (no reaper yet).
- There is no exit *status* in the notification, only the id; a supervisor that - ~~There is no exit status in the notification~~ Closed (M17.2): the kernel
needs the code can grow a wait-style call later. records how every process ends — exited, a fault class, or killed — before it
posts the exit notification, and the supervisor reads it with
`process_exit_reason` (`runtime.process.exitReason`). This is the input to
restart policy ([process-lifecycle.md](process-lifecycle.md)); an exit *code*
for the clean case can still ride alongside later.
- Enumerate writes through the caller's raw pointer under the bring-up trust - Enumerate writes through the caller's raw pointer under the bring-up trust
model, like `device_enumerate` (an unmapped page is a self-DoS, not an model, like `device_enumerate` (an unmapped page is a self-DoS, not an
isolation break). isolation break).
@@ -109,4 +116,5 @@ the architecture layer calls up into `tick`.
`process-list` (enumerate), `process-kill` (kernel-level kill paths, refusals, `process-list` (enumerate), `process-kill` (kernel-level kill paths, refusals,
notifications), `supervision` (the whole user-side surface via the process-test notifications), `supervision` (the whole user-side surface via the process-test
service: spawn supervised → enumerate → kill blocked and spinning children → service: spawn supervised → enumerate → kill blocked and spinning children →
notifications → gone). See test/qemu_test.py. notifications → gone), `claim-release` (a killed claim-holder's device is
claimable again). See test/qemu_test.py.
+4
View File
@@ -37,6 +37,10 @@ pub fn mmioMap(device_id: u64, resource_index: u64) ?usize {
/// `DeviceDescriptor.parent` for a device with no parent. /// `DeviceDescriptor.parent` for a device with no parent.
pub const no_parent = device_abi.no_parent; pub const no_parent = device_abi.no_parent;
/// `DeviceDescriptor.pci_class` for a device that is not a PCI function. Set this on
/// descriptors passed to `register` unless the child really is one.
pub const no_pci_class = device_abi.no_pci_class;
/// Publish `descriptor` as a child of `parent_id`, which this process must have claimed. /// Publish `descriptor` as a child of `parent_id`, which this process must have claimed.
/// Returns the new device id. The child is left unclaimed, so whichever driver owns /// Returns the new device id. The child is left unclaimed, so whichever driver owns
/// that class of device can `claim` it — that is how a bus hands off a device. /// that class of device can `claim` it — that is how a bus hands off a device.
+2 -1
View File
@@ -37,7 +37,8 @@ pub const MouseEventKind = protocol.MouseEventKind;
pub const JoystickEventKind = protocol.JoystickEventKind; pub const JoystickEventKind = protocol.JoystickEventKind;
pub const Keycode = protocol.Keycode; pub const Keycode = protocol.Keycode;
/// Interest masks re-exported so a caller can `subscribe(input.device_keyboard | input.device_mouse)`. /// Interest masks re-exported so a caller can `subscribe(input.device_keyboard |
/// input.device_mouse)`.
pub const device_keyboard = protocol.device_keyboard; pub const device_keyboard = protocol.device_keyboard;
pub const device_mouse = protocol.device_mouse; pub const device_mouse = protocol.device_mouse;
pub const device_joystick = protocol.device_joystick; pub const device_joystick = protocol.device_joystick;
+20
View File
@@ -94,6 +94,15 @@ pub fn send(h: Handle, message: []const u8) bool {
/// GSI. See `isNotification`. /// GSI. See `isNotification`.
pub const notify_badge_bit: u64 = abi.notify_badge_bit; pub const notify_badge_bit: u64 = abi.notify_badge_bit;
/// Set alongside `notify_badge_bit` when the notification is a **signal** — the
/// lifecycle vocabulary of docs/process-lifecycle.md, delivered to the endpoint
/// nominated with `process.bindSignals`. Decode with `process.signalsFrom`.
pub const notify_signal_bit: u64 = abi.notify_signal_bit;
/// Set alongside `notify_badge_bit` when the notification is a **one-shot timer**
/// landing (`system.timerOnce`).
pub const notify_timer_bit: u64 = abi.notify_timer_bit;
/// Set alongside `notify_badge_bit` when the notification is a **child-exit /// Set alongside `notify_badge_bit` when the notification is a **child-exit
/// notice** — a process this one spawned (with an exit endpoint) has ended — /// notice** — a process this one spawned (with an exit endpoint) has ended —
/// rather than a device interrupt. The low bits carry the child's process id. /// rather than a device interrupt. The low bits carry the child's process id.
@@ -133,6 +142,17 @@ pub const Received = struct {
} }
/// The task id of whoever posted a buffered message, meaningful only when /// The task id of whoever posted a buffered message, meaningful only when
/// Whether this arrival is a signal notification — decode the set with
/// `process.signalsFrom(badge)`.
pub fn isSignal(self: Received) bool {
return self.isNotification() and self.badge & notify_signal_bit != 0;
}
/// Whether this arrival is a one-shot timer landing (`system.timerOnce`).
pub fn isTimer(self: Received) bool {
return self.isNotification() and self.badge & notify_timer_bit != 0;
}
/// `isMessage`. (The badge's low bits, with the three high marker bits masked off.) /// `isMessage`. (The badge's low bits, with the three high marker bits masked off.)
pub fn senderTaskId(self: Received) u32 { pub fn senderTaskId(self: Received) u32 {
return @intCast(self.badge & ~(notify_badge_bit | notify_exit_bit | notify_message_bit)); return @intCast(self.badge & ~(notify_badge_bit | notify_exit_bit | notify_message_bit));
+91 -1
View File
@@ -1,8 +1,15 @@
//! Process-level runtime types: what a user program receives at entry. Mirrors //! Process-level runtime types: what a user program receives at entry (`Init`,
//! the argv contract) and the process end of the lifecycle
//! (docs/process-lifecycle.md) — today the exit reason a supervisor reads to
//! decide restart; signals and the stop sequence land here with M17.4. Mirrors
//! the spirit of `std.process.Init.Minimal` in danos terms — std's `Args` holds //! the spirit of `std.process.Init.Minimal` in danos terms — std's `Args` holds
//! no data on freestanding targets, so the type is danos's own. //! no data on freestanding targets, so the type is danos's own.
const std = @import("std"); const std = @import("std");
const abi = @import("abi");
const sc = @import("system-call.zig");
const ipc = @import("ipc.zig");
const system = @import("system.zig");
/// Everything a program receives at entry. Passed to /// Everything a program receives at entry. Passed to
/// `pub fn main(init: runtime.process.Init)`; programs that need nothing keep /// `pub fn main(init: runtime.process.Init)`; programs that need nothing keep
@@ -45,3 +52,86 @@ pub const Arguments = struct {
} }
}; };
}; };
/// How a process ended — what a supervisor's restart policy reads: a clean exit
/// meant to stop, a fault wants a restart with backoff, killed means the
/// supervisor did it itself (docs/process-lifecycle.md).
pub const ExitReason = abi.ExitReason;
/// How dead child `id` ended. Ask after the exit notification arrives — the
/// kernel records the reason before it posts the notification, so this never
/// races it. Returns null for an id that never lived, is still alive, was
/// evicted from the kernel's bounded record, or is not this process's child
/// (the same authority gate as `kill`).
pub fn exitReason(id: u32) ?ExitReason {
const r = sc.systemCall1(.process_exit_reason, id);
if (r > ~@as(usize, 0) - 4095) return null; // a wrapped -errno
return @enumFromInt(r);
}
/// The signal vocabulary (docs/process-lifecycle.md): POSIX's concepts, danos's
/// names, message delivery. A signal is a one-way coalescing statement — never a
/// question (liveness is the zero-length ping call) and never kill (that is
/// `system.kill`, unhandleable by definition).
pub const Signal = abi.Signal;
/// The coalesced set of signals one notification delivered: two pending
/// terminates arrive as one. Decode a received badge with `signalsFrom`.
pub const SignalSet = struct {
pending: u32,
pub fn has(set: SignalSet, signal: Signal) bool {
return set.pending & (@as(u32, 1) << @intFromEnum(signal)) != 0;
}
};
/// Nominate `endpoint` as this process's signal endpoint. Signals posted while
/// unbound have pended; they are delivered immediately on bind, coalesced.
pub fn bindSignals(endpoint: usize) bool {
return sc.systemCall1(.signal_bind, endpoint) == 0;
}
/// Decode a received badge into the signals it delivered, or null if it is not
/// a signal notification.
pub fn signalsFrom(badge: u64) ?SignalSet {
if (badge & abi.notify_badge_bit == 0 or badge & abi.notify_signal_bit == 0) return null;
return .{ .pending = @truncate(badge & ~(abi.notify_badge_bit | abi.notify_signal_bit)) };
}
/// Post `signal` to child `id` (or to yourself). Supervisor-gated, like kill;
/// non-blocking, always — a statement, not a conversation.
pub fn sendSignal(id: u32, signal: Signal) bool {
return sc.systemCall2(.process_signal, id, @intFromEnum(signal)) == 0;
}
/// The standard stop sequence (docs/process-lifecycle.md): terminate, wait up to
/// `deadline_ms` for the exit notification on `exit_endpoint` (the endpoint the
/// child was spawned with), then kill. Any *other* notifications arriving on
/// that endpoint while stopping are consumed and dropped — a supervisor with
/// concurrent traffic implements the same sequence inside its own event loop
/// (arm `system.timerOnce`, keep serving) instead of calling this.
pub fn stop(id: u32, deadline_ms: u64, exit_endpoint: usize) void {
_ = sendSignal(id, .terminate);
_ = system.timerOnce(exit_endpoint, deadline_ms);
var receive: [8]u8 = undefined;
while (true) {
const got = ipc.replyWait(exit_endpoint, &.{}, &receive, null);
if (got.isChildExit() and got.childProcessId() == id) return;
if (got.isTimer()) break; // the deadline passed first — escalate
}
_ = system.kill(id);
while (true) {
const got = ipc.replyWait(exit_endpoint, &.{}, &receive, null);
if (got.isChildExit() and got.childProcessId() == id) return;
}
}
/// Subscribe `endpoint` to published exit events: every process death posts an
/// asynchronous notification with the same badge encoding as a supervisor's exit
/// notice (decode with `ipc.Received.isChildExit`/`childProcessId`). For stateful
/// services: release what the dead client held — file handles, subscriptions —
/// because a service must never depend on clients cleaning up after themselves
/// (docs/process-lifecycle.md). Ungated, like `system.processes`.
pub fn subscribeExits(endpoint: usize) bool {
return sc.systemCall1(.process_subscribe, endpoint) == 0;
}
+4
View File
@@ -35,5 +35,9 @@ pub const panic = start.panic;
/// Process entry types: the `Init` handed to `main`, and its `Arguments`. /// Process entry types: the `Init` handed to `main`, and its `Arguments`.
pub const process = @import("process.zig"); pub const process = @import("process.zig");
/// The service harness: one replyWait loop folding requests, signals, and
/// notifications into callbacks (docs/process-lifecycle.md).
pub const service = @import("service.zig");
/// The heap as a `std.mem.Allocator`, for Zig `std` containers in user code. /// The heap as a `std.mem.Allocator`, for Zig `std` containers in user code.
pub const allocator = heap.allocator; pub const allocator = heap.allocator;
+81
View File
@@ -0,0 +1,81 @@
//! The service harness (docs/process-lifecycle.md): one replyWait loop that
//! folds protocol requests, signals, and subscribed notifications into
//! callbacks — so the lifecycle contract ("answers ping, exits on terminate")
//! is satisfied by construction and a service author writes domain logic only.
//! Nothing is asynchronous inside the process: a callback runs at a point the
//! loop chose, never on a hijacked stack — the whole reason signals are
//! messages.
//!
//! The liveness probe: a **zero-length request is the universal ping**, answered
//! with a zero-length reply by the harness itself. No protocol's requests start
//! at length zero, so the encoding cannot collide, and there is nothing for a
//! service author to implement — a wedged service simply fails to answer, which
//! is the diagnosis (see docs/ipc.md).
const abi = @import("abi");
const ipc = @import("ipc.zig");
const process = @import("process.zig");
pub const Callbacks = struct {
/// Called once with the service's endpoint before the loop starts — the
/// place to subscribe to exit events, bind IRQs, or announce readiness.
/// Return false to abort startup (the process exits).
init: ?*const fn (endpoint: ipc.Handle) bool = null,
/// One protocol request from `sender` (a task id): write the reply into
/// `reply`, return its length. The zero-length ping never reaches this.
on_message: *const fn (message: []const u8, reply: []u8, sender: u32) usize,
/// A notification that is not a signal — a subscribed exit event, a bound
/// IRQ, a timer landing. The raw badge; decode with the ipc helpers.
on_notification: ?*const fn (badge: u64) void = null,
/// The reload signal. Default: ignored.
on_reload: ?*const fn () void = null,
/// The terminate signal, called before the loop returns. The clean exit is
/// the return itself — never put *necessary* work here (iron rule 1: a kill
/// arrives with no warning; this is for graceful extras only).
on_terminate: ?*const fn () void = null,
/// Publish the endpoint under a well-known service id at startup.
service: ?abi.ServiceId = null,
};
/// Run the service: create and (optionally) register the endpoint, bind signals
/// to it, call `init`, then serve until `terminate` arrives — at which point the
/// loop returns and main's return is the clean exit the supervisor reads as
/// `ExitReason.exited`. `maximum_message` sizes the receive and reply buffers
/// (a service passes its protocol's message maximum).
pub fn run(comptime maximum_message: usize, callbacks: Callbacks) void {
const endpoint = ipc.createIpcEndpoint() orelse return;
if (callbacks.service) |id| {
if (!ipc.register(id, endpoint)) return;
}
_ = process.bindSignals(endpoint);
if (callbacks.init) |initialise| {
if (!initialise(endpoint)) return;
}
var reply_buffer: [maximum_message]u8 = undefined;
var reply_len: usize = 0;
var receive: [maximum_message]u8 = undefined;
while (true) {
const got = ipc.replyWait(endpoint, reply_buffer[0..reply_len], &receive, null);
if (got.isNotification()) {
reply_len = 0; // nothing owed for a notification
if (process.signalsFrom(got.badge)) |signals| {
if (signals.has(.reload)) {
if (callbacks.on_reload) |onReload| onReload();
}
if (signals.has(.terminate)) {
if (callbacks.on_terminate) |onTerminate| onTerminate();
return; // the loop's return IS the clean exit
}
continue;
}
if (callbacks.on_notification) |onNotification| onNotification(got.badge);
continue;
}
if (got.len == 0) {
reply_len = 0; // the universal ping: a zero-length reply, from the harness
continue;
}
reply_len = callbacks.on_message(receive[0..got.len], &reply_buffer, got.senderTaskId());
}
}
+20 -5
View File
@@ -19,34 +19,49 @@ pub inline fn systemCall0(n: SystemCall) usize {
pub inline fn systemCall1(n: SystemCall, a0: usize) usize { pub inline fn systemCall1(n: SystemCall, a0: usize) usize {
return asm volatile ("syscall" return asm volatile ("syscall"
: [ret] "={rax}" (-> usize), : [ret] "={rax}" (-> usize),
: [n] "{rax}" (@intFromEnum(n)), [a0] "{rdi}" (a0), : [n] "{rax}" (@intFromEnum(n)),
[a0] "{rdi}" (a0),
: .{ .rcx = true, .r11 = true, .memory = true }); : .{ .rcx = true, .r11 = true, .memory = true });
} }
pub inline fn systemCall2(n: SystemCall, a0: usize, a1: usize) usize { pub inline fn systemCall2(n: SystemCall, a0: usize, a1: usize) usize {
return asm volatile ("syscall" return asm volatile ("syscall"
: [ret] "={rax}" (-> usize), : [ret] "={rax}" (-> usize),
: [n] "{rax}" (@intFromEnum(n)), [a0] "{rdi}" (a0), [a1] "{rsi}" (a1), : [n] "{rax}" (@intFromEnum(n)),
[a0] "{rdi}" (a0),
[a1] "{rsi}" (a1),
: .{ .rcx = true, .r11 = true, .memory = true }); : .{ .rcx = true, .r11 = true, .memory = true });
} }
pub inline fn systemCall3(n: SystemCall, a0: usize, a1: usize, a2: usize) usize { pub inline fn systemCall3(n: SystemCall, a0: usize, a1: usize, a2: usize) usize {
return asm volatile ("syscall" return asm volatile ("syscall"
: [ret] "={rax}" (-> usize), : [ret] "={rax}" (-> usize),
: [n] "{rax}" (@intFromEnum(n)), [a0] "{rdi}" (a0), [a1] "{rsi}" (a1), [a2] "{rdx}" (a2), : [n] "{rax}" (@intFromEnum(n)),
[a0] "{rdi}" (a0),
[a1] "{rsi}" (a1),
[a2] "{rdx}" (a2),
: .{ .rcx = true, .r11 = true, .memory = true }); : .{ .rcx = true, .r11 = true, .memory = true });
} }
pub inline fn systemCall4(n: SystemCall, a0: usize, a1: usize, a2: usize, a3: usize) usize { pub inline fn systemCall4(n: SystemCall, a0: usize, a1: usize, a2: usize, a3: usize) usize {
return asm volatile ("syscall" return asm volatile ("syscall"
: [ret] "={rax}" (-> usize), : [ret] "={rax}" (-> usize),
: [n] "{rax}" (@intFromEnum(n)), [a0] "{rdi}" (a0), [a1] "{rsi}" (a1), [a2] "{rdx}" (a2), [a3] "{r10}" (a3), : [n] "{rax}" (@intFromEnum(n)),
[a0] "{rdi}" (a0),
[a1] "{rsi}" (a1),
[a2] "{rdx}" (a2),
[a3] "{r10}" (a3),
: .{ .rcx = true, .r11 = true, .memory = true }); : .{ .rcx = true, .r11 = true, .memory = true });
} }
pub inline fn systemCall5(n: SystemCall, a0: usize, a1: usize, a2: usize, a3: usize, a4: usize) usize { pub inline fn systemCall5(n: SystemCall, a0: usize, a1: usize, a2: usize, a3: usize, a4: usize) usize {
return asm volatile ("syscall" return asm volatile ("syscall"
: [ret] "={rax}" (-> usize), : [ret] "={rax}" (-> usize),
: [n] "{rax}" (@intFromEnum(n)), [a0] "{rdi}" (a0), [a1] "{rsi}" (a1), [a2] "{rdx}" (a2), [a3] "{r10}" (a3), [a4] "{r8}" (a4), : [n] "{rax}" (@intFromEnum(n)),
[a0] "{rdi}" (a0),
[a1] "{rsi}" (a1),
[a2] "{rdx}" (a2),
[a3] "{r10}" (a3),
[a4] "{r8}" (a4),
: .{ .rcx = true, .r11 = true, .memory = true }); : .{ .rcx = true, .r11 = true, .memory = true });
} }
+9
View File
@@ -32,6 +32,15 @@ pub fn sleep(ms: usize) void {
_ = sc.systemCall1(.sleep, ms); _ = sc.systemCall1(.sleep, ms);
} }
/// Arm a one-shot timer: after `ms` milliseconds the kernel posts a timer
/// notification (`ipc.Received.isTimer`) to `endpoint`. The timed wait of
/// docs/process-lifecycle.md — a service arms a deadline and keeps serving,
/// instead of blocking in sleep; what stop-sequence escalation, hello deadlines,
/// and restart backoff are built from.
pub fn timerOnce(endpoint: usize, ms: u64) bool {
return sc.systemCall2(.timer_bind, endpoint, ms) == 0;
}
/// Monotonic nanoseconds since boot — a time source for timeouts and short delays. It /// Monotonic nanoseconds since boot — a time source for timeouts and short delays. It
/// only ever moves forward. This is *not* wall-clock time (no date, no timezone — that /// only ever moves forward. This is *not* wall-clock time (no date, no timezone — that
/// is a user-space service layered on top). Deadline pattern for a bounded poll loop: /// is a user-space service layered on top). Deadline pattern for a bounded poll loop:
+51
View File
@@ -53,9 +53,31 @@ pub const SystemCall = enum(u64) {
process_enumerate = 24, // process_enumerate(buffer, maximum) -> total: snapshot the task table process_enumerate = 24, // process_enumerate(buffer, maximum) -> total: snapshot the task table
process_kill = 25, // process_kill(id) -> 0/-errno: end a process this process spawned process_kill = 25, // process_kill(id) -> 0/-errno: end a process this process spawned
ipc_send = 26, // ipc_send(handle, message_ptr, message_len) -> 0/-errno: post a payload to an endpoint's async queue without blocking ipc_send = 26, // ipc_send(handle, message_ptr, message_len) -> 0/-errno: post a payload to an endpoint's async queue without blocking
process_exit_reason = 27, // process_exit_reason(id) -> ExitReason/-errno: how a dead child ended (its supervisor only)
process_subscribe = 28, // process_subscribe(endpoint) -> 0/-errno: subscribe to published exit events — every death posts a notification
signal_bind = 29, // signal_bind(endpoint) -> 0/-errno: nominate the endpoint this process's signals arrive on
process_signal = 30, // process_signal(id, signal) -> 0/-errno: post a signal to a child (or to yourself)
timer_bind = 31, // timer_bind(endpoint, ms) -> 0/-errno: one-shot timer — posts a notification when ms elapse
_, _,
}; };
/// How a process ended — recorded by the kernel at death, queried by the
/// supervisor with `process_exit_reason`, and the input to its restart decision
/// (docs/process-lifecycle.md): a clean exit meant to stop, a fault wants a
/// restart with backoff, killed means the supervisor did it itself. The faults
/// mirror the CPU exceptions a ring-3 process can die of; they are exit reasons,
/// never delivered to the faulting process (recovery is restart, not a handler).
pub const ExitReason = enum(u8) {
exited = 0, // returned from main / called exit
aborted = 1, // deliberate self-termination (reserved: no abort path yet)
segmentation_fault = 2, // page fault
illegal_instruction = 3, // invalid opcode
arithmetic_fault = 4, // divide error, x87 or SIMD fault
protection_fault = 5, // general protection fault
fault = 6, // any other CPU exception
killed = 7, // process_kill
};
/// The x86 MSI message address base (`0xFEE0_0000`): a device raises an MSI by writing /// The x86 MSI message address base (`0xFEE0_0000`): a device raises an MSI by writing
/// `data` to this address, which the Local APIC turns into an interrupt at the vector /// `data` to this address, which the Local APIC turns into an interrupt at the vector
/// in `data`. The kernel returns the concrete (address, data) from `msi_bind`; this is /// in `data`. The kernel returns the concrete (address, data) from `msi_bind`; this is
@@ -93,6 +115,35 @@ pub const notify_exit_bit: u64 = 1 << 62;
/// broadcasts where a rendezvous is the wrong shape (the input service is the first user). /// broadcasts where a rendezvous is the wrong shape (the input service is the first user).
pub const notify_message_bit: u64 = 1 << 61; pub const notify_message_bit: u64 = 1 << 61;
/// Set (alongside `notify_badge_bit`) in the badge of a **signal notification** —
/// the process-lifecycle vocabulary of docs/process-lifecycle.md, delivered to the
/// endpoint the process nominated with `signal_bind`. The low bits carry the
/// coalesced pending mask (bit positions = `Signal` values): signals are
/// statements, not questions, and two pending terminates are one terminate.
pub const notify_signal_bit: u64 = 1 << 60;
/// Set (alongside `notify_badge_bit`) in the badge of a **timer notification** —
/// a one-shot `timer_bind` deadline landing. No payload bits: what to do when the
/// deadline fires is whatever the receiver armed it for (a stop-sequence
/// escalation, a restart backoff, an alarm).
pub const notify_timer_bit: u64 = 1 << 59;
/// The signal vocabulary (docs/process-lifecycle.md): POSIX's concepts, danos's
/// names, message delivery. The value is the bit position in the pending mask — a
/// private kernel/runtime detail, free to change while they ship together. Kill
/// is not here (it is `process_kill`, unhandleable by definition); faults are not
/// here (they are `ExitReason`s — recovery is restart, not a handler); liveness is
/// not here (a question, asked as the zero-length ping call, not a statement).
pub const Signal = enum(u5) {
terminate = 0, // finish up and exit (the polite half of the stop sequence)
reload = 1, // re-read configuration / re-scan
interrupt = 2, // interactive interrupt (no sender until a console exists)
quit = 3, // as interrupt, by convention more final
alarm = 4, // a timer the process armed for itself (unbuilt: no consumer yet)
user_1 = 5, // service-defined
user_2 = 6, // service-defined
};
/// Capacity of `ProcessDescriptor.name` — matches the longest name `system_spawn` /// Capacity of `ProcessDescriptor.name` — matches the longest name `system_spawn`
/// accepts, so a process's recorded name (its argv[0]) is never truncated. /// accepts, so a process's recorded name (its argv[0]) is never truncated.
pub const maximum_process_name = 64; pub const maximum_process_name = 64;
+4 -2
View File
@@ -52,7 +52,8 @@ pub const PowerInformation = struct {
reset: RegisterAccess = .{}, reset: RegisterAccess = .{},
reset_value: u8 = 0, reset_value: u8 = 0,
reset_supported: bool = false, reset_supported: bool = false,
/// SLP_TYP values for S5 (soft off) and S3 (suspend), from the AML sleep-state (`_Sx`) packages. /// SLP_TYP values for S5 (soft off) and S3 (suspend), from the AML sleep-state (`_Sx`)
/// packages.
s5: ?aml.SleepType = null, s5: ?aml.SleepType = null,
s3: ?aml.SleepType = null, s3: ?aml.SleepType = null,
}; };
@@ -198,7 +199,8 @@ const ExtendedSystemDescriptorPointer = extern struct {
root_system_description_table_address: u32 align(1), root_system_description_table_address: u32 align(1),
/// The size of the RSDP. /// The size of the RSDP.
length: u32 align(1), length: u32 align(1),
/// A 64-bit physical address pointing to the XSDT. If the revision is at least 2, the XSDT should be used regardless of architecture, as the RSDT was deprecated. /// A 64-bit physical address pointing to the XSDT. If the revision is at least 2, the XSDT
/// should be used regardless of architecture, as the RSDT was deprecated.
extended_system_descriptor_table_address: u64 align(1), extended_system_descriptor_table_address: u64 align(1),
/// A checksum used for the entire table. /// A checksum used for the entire table.
extended_checksum: u8, extended_checksum: u8,
+10 -5
View File
@@ -126,17 +126,22 @@ test "parses a nested namespace and finds the sleep package" {
// Scope(\_SB) packagelen=0x27 // Scope(\_SB) packagelen=0x27
0x10, 0x27, 0x5C, 0x5F, 0x53, 0x42, 0x5F, 0x10, 0x27, 0x5C, 0x5F, 0x53, 0x42, 0x5F,
// Device(PCI0) packagelen=0x1F // Device(PCI0) packagelen=0x1F
0x5B, 0x82, 0x1F, 0x50, 0x43, 0x49, 0x30, 0x5B, 0x82, 0x1F, 0x50, 0x43,
0x49, 0x30,
// Name(_HID, 0x11) // Name(_HID, 0x11)
0x08, 0x5F, 0x48, 0x49, 0x44, 0x0A, 0x11, 0x08, 0x5F, 0x48, 0x49, 0x44, 0x0A, 0x11,
// Method(MTHD, flags=1) empty, packagelen=0x06 // Method(MTHD, flags=1) empty, packagelen=0x06
0x14, 0x06, 0x4D, 0x54, 0x48, 0x44, 0x01, 0x14, 0x06, 0x4D,
0x54, 0x48, 0x44, 0x01,
// Method(CALL, flags=0) { MTHD(Zero) }, packagelen=0x0B // Method(CALL, flags=0) { MTHD(Zero) }, packagelen=0x0B
0x14, 0x0B, 0x43, 0x41, 0x4C, 0x4C, 0x00, 0x4D, 0x54, 0x48, 0x44, 0x00, 0x14, 0x0B, 0x43, 0x41, 0x4C, 0x4C, 0x00, 0x4D,
0x54, 0x48, 0x44, 0x00,
// OperationRegion(DBG0, SystemIO, Word 0x0402, Byte 1) // OperationRegion(DBG0, SystemIO, Word 0x0402, Byte 1)
0x5B, 0x80, 0x44, 0x42, 0x47, 0x30, 0x01, 0x0B, 0x02, 0x04, 0x0A, 0x01, 0x5B, 0x80, 0x44, 0x42, 0x47, 0x30, 0x01, 0x0B,
0x02, 0x04, 0x0A, 0x01,
// Field(DBG0, flags=1) { DBGB, 8 }, packagelen=0x0B // Field(DBG0, flags=1) { DBGB, 8 }, packagelen=0x0B
0x5B, 0x81, 0x0B, 0x44, 0x42, 0x47, 0x30, 0x01, 0x44, 0x42, 0x47, 0x42, 0x08, 0x5B, 0x81, 0x0B, 0x44, 0x42, 0x47, 0x30, 0x01,
0x44, 0x42, 0x47, 0x42, 0x08,
}; };
var arena = std.heap.ArenaAllocator.init(std.testing.allocator); var arena = std.heap.ArenaAllocator.init(std.testing.allocator);
+9
View File
@@ -56,6 +56,10 @@ pub const maximum_device_resources = 8;
/// `DeviceDescriptor.parent` for a device with no parent — a root of the device tree. /// `DeviceDescriptor.parent` for a device with no parent — a root of the device tree.
pub const no_parent: u64 = ~@as(u64, 0); pub const no_parent: u64 = ~@as(u64, 0);
/// `DeviceDescriptor.pci_class` for a device that is not a PCI function. (Zero would be
/// ambiguous: 0x000000 is a real class code, "unclassified device".)
pub const no_pci_class: u64 = ~@as(u64, 0);
/// A device, as snapshotted for user space by `device_enumerate`. A driver scans /// A device, as snapshotted for user space by `device_enumerate`. A driver scans
/// these to find the hardware it owns, claims it, and maps its MMIO. /// these to find the hardware it owns, claims it, and maps its MMIO.
/// ///
@@ -69,6 +73,11 @@ pub const DeviceDescriptor = extern struct {
id: u64, id: u64,
parent: u64, // a device id, or `no_parent` parent: u64, // a device id, or `no_parent`
class: u64, // a DeviceClass value class: u64, // a DeviceClass value
// The PCI class/subclass/prog-IF triple packed as 0xCCSSPP when this device is a PCI
// function, or `no_pci_class` otherwise. This is how a manager tells *what* a
// `pci_device` is (an xHCI controller, an AHCI controller) — decode the triple into
// names with the pci-class module.
pci_class: u64,
hid_len: u64, hid_len: u64,
resource_count: u64, resource_count: u64,
hid: [8]u8, hid: [8]u8,
+1 -2
View File
@@ -198,8 +198,7 @@ fn dumpNode(device: *const Device, depth: usize, emit: *const fn ([]const u8) vo
std.fmt.bufPrint(buffer[indent..], "{s} [{s}] hid={s} ({s})\n", .{ device.name(), @tagName(device.class), device.hid(), desc }) catch return std.fmt.bufPrint(buffer[indent..], "{s} [{s}] hid={s} ({s})\n", .{ device.name(), @tagName(device.class), device.hid(), desc }) catch return
else else
std.fmt.bufPrint(buffer[indent..], "{s} [{s}] hid={s}\n", .{ device.name(), @tagName(device.class), device.hid() }) catch return; std.fmt.bufPrint(buffer[indent..], "{s} [{s}] hid={s}\n", .{ device.name(), @tagName(device.class), device.hid() }) catch return;
} else } else std.fmt.bufPrint(buffer[indent..], "{s} [{s}]\n", .{ device.name(), @tagName(device.class) }) catch return;
std.fmt.bufPrint(buffer[indent..], "{s} [{s}]\n", .{ device.name(), @tagName(device.class) }) catch return;
emit(buffer[0 .. indent + body.len]); emit(buffer[0 .. indent + body.len]);
// For a PCI function, decode its class code — the (class / subclass / prog-IF) // For a PCI function, decode its class code — the (class / subclass / prog-IF)
+857
View File
@@ -0,0 +1,857 @@
//! USB device-framework wire ABI: the set-up packets, standard requests, and standard
//! descriptors every USB device speaks over its default control pipe, as defined by chapter 9
//! of the USB 2.0 specification (see https://wiki.osdev.org/Universal_Serial_Bus). Pure data
//! definitions — no hardware access — shared by the host-controller bus drivers (which build
//! the requests) and anything that parses what devices return (device naming, driver
//! matching, configuration). The structs mirror the wire byte-for-byte: multi-byte fields are
//! little-endian and align(1), so a descriptor can be bit-cast straight out of a transfer
//! buffer at any offset, and bitmap bytes are packed structs so no caller ever needs a magic
//! mask. Class, subclass, and protocol code tables live in usb-ids.zig.
const DeviceState = enum(u8) {
// Immediately after the USB device is attached to the USB system, it is in this state.
// The USB specifications do not define the state of a USB device that is detached from
// a USB system.
attached,
// A device is in this state after it has both been attached to the bus, and the VBUS line is
// applied to the device (the host controller drives the VBUS at +5V, however this is only
// particularly important for hardware developers). In this state, the device must not respond
// to any bus transactions. The USB specification recognizes three potential scenarios with
// respect to how a device draws power:
// - Self-Powered Devices draw power from an external power source (e.g, a USB printer plugs
// into the wall as well as a USB port). Although the device may be considered
// technically "powered" even before attachment to the USB, it is still only considered
// powered after the VBUS line is applied to the device.
// - Bus-Powered Devices draw power solely from the USB up to 100mA.
// - Self- or Bus-Powered Devices may draw power from either the bus or an external power
// source, depending on the configuration. These devices may change power source at any
// time. If a device is currently self-powered and requires more than 100mA of power, but
// switches to being bus-powered, then the device must return to the Address state.
powered,
// A device in the powered state enters the default state after receiving a bus reset. In this
// state, the device is addressable at the default, reserved address of 0. At this point, the
// device is operating at the correct speed. The host is expected to allow 10 milliseconds
// before expecting the device to respond to data transfers after reset.
default,
// A device enters this state after the host assigns it an address via the default control pipe,
// which is always accessible whether the device's address has been set or not.
address,
// A device is in this state after the host examines its possible configurations and selects
// one. All endpoint's data toggle bits are initialized to zero when a device enters this state.
configured,
// When no traffic is observed on the bus for a period of 1 millisecond, a USB device enters
// this state, characterized by its low power consumption. The device's address and
// configuration settings are maintained while suspended. A device exits the suspended state as
// soon as it begins seeing bus activity again. The host is expected to allow 10 milliseconds
// before expecting the device to respond to data transfers after resume.
suspended,
};
const RequestCode = enum(u8) {
get_status = 0,
clear_feature = 1,
set_feature = 3,
set_address = 5,
get_descriptor = 6,
set_descriptor = 7,
get_configuration = 8,
set_configuration = 9,
get_interface = 10,
set_interface = 11,
sync_frame = 12,
};
// Direction of an endpoint, from the host's point of view
const EndpointDirection = enum(u1) {
out = 0,
in = 1,
};
// Identifier newtypes: distinct wire-sized types for values that identify something on the
// device rather than count something. Each is a non-exhaustive enum whose values originate
// in the descriptors below and flow, still typed, into the standard request constructors —
// so an interface number can never be passed where a configuration value is expected.
// The bus address of a device, assigned by the host with SET_ADDRESS. Addresses are 7 bits
// wide.
const DeviceAddress = enum(u7) {
// The default address every device answers at after a reset, until SET_ADDRESS
// completes
default = 0,
_,
};
// Identifies a configuration; from ConfigurationDescriptor.configuration_value.
const ConfigurationValue = enum(u8) {
// Not configured: returned by GET_CONFIGURATION while the device is in the address
// state, and passed to SET_CONFIGURATION to return a configured device to the address
// state
none = 0,
_,
};
// Identifies an interface within a configuration; from
// InterfaceDescriptor.interface_number.
const InterfaceNumber = enum(u8) { _ };
// Selects between the alternate settings of one interface; from
// InterfaceDescriptor.alternate_setting.
const AlternateSetting = enum(u8) {
// The default setting of an interface
default = 0,
_,
};
// The number of an endpoint within a device, 4 bits wide. The direction bit carried
// alongside it tells the two endpoints sharing a number apart.
const EndpointNumber = enum(u4) {
// Endpoint zero: the default control pipe every device provides
default_control = 0,
_,
};
// Index of a STRING descriptor, stored in descriptors that reference a string and passed to
// GET_DESCRIPTOR to read it.
const StringIndex = enum(u8) {
// The device has no string descriptor for this field
none = 0,
_,
};
// Characteristics of a device request (the bmRequestType field of a set-up packet). Fields are
// declared least-significant first: recipient occupies bits 4...0, kind bits 6...5, and
// direction bit 7.
const RequestType = packed struct(u8) {
// The recipient of the request (values 4...31 are reserved)
recipient: Recipient,
// The type of the request
kind: Kind,
// Data transfer direction. The value of this bit is ignored when length is zero.
direction: Direction,
const Recipient = enum(u5) {
device = 0,
interface = 1,
endpoint = 2,
other = 3,
};
const Kind = enum(u2) {
standard = 0,
class = 1,
vendor = 2,
reserved = 3,
};
const Direction = enum(u1) {
host_to_device = 0,
device_to_host = 1,
};
};
const Request = extern struct {
// Characteristics of the request
request_type: RequestType,
// Specific request
request_code: RequestCode,
// Word-sized field that may (or may not) serve as a parameter to the request, depending
// on the specific request. For GET_DESCRIPTOR and SET_DESCRIPTOR, bit-cast a
// DescriptorValue into this field.
value: u16 align(1),
// Word-sized field that may (or may not) serve as a parameter to the request, depending
// on the specific request. Typically this field holds an index or an offset value. When
// request_type specifies an endpoint or an interface as the recipient, bit-cast an
// EndpointIndex or an InterfaceIndex into this field.
index: u16 align(1),
// Number of bytes to transfer if there is a DATA stage.
// - If this field is non-zero, and request_type indicates a transfer from
// device-to-host, then the device must never return more than length bytes of data.
// However, a device may return less.
// - If this field is non-zero, and request_type indicates a transfer from
// host-to-device, then the host must send exactly length bytes of data. If the host
// sends more than length bytes, the behavior of the device is undefined.
length: u16 align(1),
// The format of the index field when request_type specifies an endpoint as the
// recipient. The host should always set the direction bit to zero (but the device
// should accept either value) when the endpoint is part of a control pipe.
const EndpointIndex = packed struct(u16) {
// Endpoint number
number: EndpointNumber,
// Reserved (reset to zero)
reserved: u3 = 0,
// Selects the OUT or the IN endpoint with the specified endpoint number
direction: EndpointDirection,
// Reserved (reset to zero)
reserved_high: u8 = 0,
};
// The format of the index field when request_type specifies an interface as the
// recipient.
const InterfaceIndex = packed struct(u16) {
// Interface number
number: u8,
// Reserved (reset to zero)
reserved: u8 = 0,
};
// The format of the value field of GET_DESCRIPTOR and SET_DESCRIPTOR requests: the
// descriptor type in the high byte, and the descriptor index in the low byte. The index
// is used to select a specific descriptor (only for CONFIGURATION and STRING
// descriptors) when several descriptors of that type are implemented by a device.
const DescriptorValue = packed struct(u16) {
// Descriptor index
index: u8 = 0,
// Descriptor type
kind: DescriptorType,
};
};
// Feature selectors, used as the value field of CLEAR_FEATURE and SET_FEATURE requests. The
// comment on each value notes the recipient the selector applies to.
const FeatureSelector = enum(u16) {
// Halts an endpoint (recipient: endpoint)
endpoint_halt = 0,
// Enables or disables the device's remote wakeup capability (recipient: device)
device_remote_wakeup = 1,
// Puts a hi-speed device into a test mode, selected by a TestMode value in the high
// byte of the index field (recipient: device)
test_mode = 2,
};
// Test mode selectors, passed in the high byte of the index field of a SET_FEATURE request
// with the test_mode feature selector. Values 06h...3Fh are reserved for standard test
// selectors and C0h...FFh for vendor-specific test modes; all other unlisted values are
// reserved.
const TestMode = enum(u8) {
test_j = 0x01,
test_k = 0x02,
test_se0_nak = 0x03,
test_packet = 0x04,
test_force_enable = 0x05,
_,
};
// The two bytes returned by a GET_STATUS request directed at a device. Fields are declared
// least-significant first.
const DeviceStatus = packed struct(u16) {
// Whether the device is currently self-powered (as opposed to bus-powered). This bit
// cannot be changed with the SET_FEATURE or CLEAR_FEATURE requests.
self_powered: bool,
// Whether the device is currently enabled to request remote wakeup. Changed with the
// SET_FEATURE and CLEAR_FEATURE requests using the device_remote_wakeup feature
// selector.
remote_wakeup: bool,
// Reserved (reset to zero)
reserved: u14,
};
// The two bytes returned by a GET_STATUS request directed at an endpoint. (A GET_STATUS
// request directed at an interface returns two bytes that are entirely reserved.)
const EndpointStatus = packed struct(u16) {
// Whether the endpoint is currently halted. Set with the SET_FEATURE request using the
// endpoint_halt feature selector, and cleared with CLEAR_FEATURE.
halted: bool,
// Reserved (reset to zero)
reserved: u15,
};
// A target for the standard requests that may be directed at the device, an interface, or
// an endpoint.
const Target = union(enum) {
device,
interface: InterfaceNumber,
endpoint: Request.EndpointIndex,
fn recipient(target: Target) RequestType.Recipient {
return switch (target) {
.device => .device,
.interface => .interface,
.endpoint => .endpoint,
};
}
fn index(target: Target) u16 {
return switch (target) {
.device => 0,
.interface => |number| @intFromEnum(number),
.endpoint => |endpoint| @bitCast(endpoint),
};
}
};
// Constructors for the standard device requests, one per RequestCode. Each returns a
// ready-to-send set-up packet with the request_type, value, index, and length fields the
// specification prescribes for that request.
// Reads the status of the given target: bit-cast the two bytes the device returns into a
// DeviceStatus or an EndpointStatus. (The two bytes returned for an interface are entirely
// reserved.)
fn getStatus(target: Target) Request {
return .{
.request_type = .{
.recipient = target.recipient(),
.kind = .standard,
.direction = .device_to_host,
},
.request_code = .get_status,
.value = 0,
.index = target.index(),
.length = 2,
};
}
// Clears or disables the given feature. A device cannot be taken out of a test mode with
// this request; test_mode is only cleared by cycling power.
fn clearFeature(feature: FeatureSelector, target: Target) Request {
return .{
.request_type = .{
.recipient = target.recipient(),
.kind = .standard,
.direction = .host_to_device,
},
.request_code = .clear_feature,
.value = @intFromEnum(feature),
.index = target.index(),
.length = 0,
};
}
// Sets or enables the given feature. For the test_mode feature selector, use setTestMode
// instead: the test selector rides in the high byte of the index field.
fn setFeature(feature: FeatureSelector, target: Target) Request {
return .{
.request_type = .{
.recipient = target.recipient(),
.kind = .standard,
.direction = .host_to_device,
},
.request_code = .set_feature,
.value = @intFromEnum(feature),
.index = target.index(),
.length = 0,
};
}
// Puts a hi-speed device into the given test mode: a SET_FEATURE request with the test_mode
// feature selector and the test selector in the high byte of the index field.
fn setTestMode(mode: TestMode) Request {
return .{
.request_type = .{
.recipient = .device,
.kind = .standard,
.direction = .host_to_device,
},
.request_code = .set_feature,
.value = @intFromEnum(FeatureSelector.test_mode),
.index = @as(u16, @intFromEnum(mode)) << 8,
.length = 0,
};
}
// Assigns the device its bus address, moving it from the default state to the address
// state. The device does not answer at the new address until the status stage of this
// request completes.
fn setAddress(address: DeviceAddress) Request {
return .{
.request_type = .{
.recipient = .device,
.kind = .standard,
.direction = .host_to_device,
},
.request_code = .set_address,
.value = @intFromEnum(address),
.index = 0,
.length = 0,
};
}
// Reads a descriptor from the device.
// - descriptor_index selects among descriptors of the same type, and is only used for
// configuration and string descriptors.
// - language_id selects the language of a string descriptor, and is zero otherwise.
// - length is the number of bytes to read; a device never returns more than length bytes,
// but may return less if the descriptor is shorter.
fn getDescriptor(kind: DescriptorType, descriptor_index: u8, language_id: u16, length: u16) Request {
return .{
.request_type = .{
.recipient = .device,
.kind = .standard,
.direction = .device_to_host,
},
.request_code = .get_descriptor,
.value = @bitCast(Request.DescriptorValue{ .index = descriptor_index, .kind = kind }),
.index = language_id,
.length = length,
};
}
// Updates an existing descriptor or adds a new one (optional; many devices do not support
// this request). The parameters mirror getDescriptor; the descriptor itself is sent in the
// DATA stage.
fn setDescriptor(kind: DescriptorType, descriptor_index: u8, language_id: u16, length: u16) Request {
return .{
.request_type = .{
.recipient = .device,
.kind = .standard,
.direction = .host_to_device,
},
.request_code = .set_descriptor,
.value = @bitCast(Request.DescriptorValue{ .index = descriptor_index, .kind = kind }),
.index = language_id,
.length = length,
};
}
// Reads the currently active configuration: @enumFromInt the byte the device returns into a
// ConfigurationValue, which is none while the device is not configured.
fn getConfiguration() Request {
return .{
.request_type = .{
.recipient = .device,
.kind = .standard,
.direction = .device_to_host,
},
.request_code = .get_configuration,
.value = 0,
.index = 0,
.length = 1,
};
}
// Selects the configuration with the given configuration_value (from
// ConfigurationDescriptor.configuration_value), moving the device from the address state to
// the configured state. Selecting none returns the device to the address state.
fn setConfiguration(configuration_value: ConfigurationValue) Request {
return .{
.request_type = .{
.recipient = .device,
.kind = .standard,
.direction = .host_to_device,
},
.request_code = .set_configuration,
.value = @intFromEnum(configuration_value),
.index = 0,
.length = 0,
};
}
// Reads the alternate setting currently selected for the given interface: @enumFromInt the
// byte the device returns into an AlternateSetting.
fn getInterface(interface: InterfaceNumber) Request {
return .{
.request_type = .{
.recipient = .interface,
.kind = .standard,
.direction = .device_to_host,
},
.request_code = .get_interface,
.value = 0,
.index = @intFromEnum(interface),
.length = 1,
};
}
// Selects an alternate setting (from InterfaceDescriptor.alternate_setting) for the given
// interface.
fn setInterface(interface: InterfaceNumber, alternate_setting: AlternateSetting) Request {
return .{
.request_type = .{
.recipient = .interface,
.kind = .standard,
.direction = .host_to_device,
},
.request_code = .set_interface,
.value = @intFromEnum(alternate_setting),
.index = @intFromEnum(interface),
.length = 0,
};
}
// Reads the two-byte number of the frame in which the given isochronous endpoint's
// repeating pattern of transfers begins.
fn syncFrame(endpoint: Request.EndpointIndex) Request {
return .{
.request_type = .{
.recipient = .endpoint,
.kind = .standard,
.direction = .device_to_host,
},
.request_code = .sync_frame,
.value = 0,
.index = @bitCast(endpoint),
.length = 2,
};
}
const DescriptorType = enum(u8) {
device = 1,
configuration = 2,
string = 3,
interface = 4,
endpoint = 5,
device_qualifier = 6,
other_speed_configuration = 7,
interface_power = 8,
_,
};
const DeviceDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// DEVICE Descriptor Type
descriptor_type: DescriptorType,
// USB Specification Release Number in Binary-Coded Decimal (i.e, 2.10 is expressed as 210h).
// Identifies the release of the USB Specification with with the device and its
// descriptors are compliant.
bcd_usb: u16 align(1),
// Class code (assigned by the USB-IF)
// - This field is reset to zero if each interface within a configuration specifies its own
// class information and the various interfaces operate independently.
// - A value of FFh in this field indicates the device class is vendor-specific.
device_class: u8,
// Subclass Code (assigned by the USB-IF)
// - The subclass code of a device is qualified by the class code of that device.
// - If device_class is reset to zero, then this field must also be reset to zero.
// - When device_class is not set to FFh, then all values for this field are reserved for
// assignment by the USB-IF.
device_subclass: u8,
// Protocol code (assigned by the USB-IF)
// - The protocol code of a device is qualified by both the class and subclass codes of
// that device.
// - A value of 00h in this field means that the device may specify class-specific
// protocols on an interface basis, though this is not a requirement.
// - If this field is set to FFh, then the device uses a vendor-specific protocol.
device_protocol: u8,
// Maximum packet size for endpoint zero (8, 16, 32, or 64 are the only valid options)
max_packet_size_0: u8,
// Vendor ID (assigned by the USB-IF)
vendor_id: u16 align(1),
// Product ID (assigned by the USB-IF)
product_id: u16 align(1),
// Device release number in binary-coded decimal
bcd_device: u16 align(1),
// Index of STRING descriptor describing manufacturer
manufacturer_index: StringIndex,
// Index of STRING descriptor describing product
product_index: StringIndex,
// Index of STRING descriptor describing the device's serial number
serial_number_index: StringIndex,
// Number of possible configurations
configuration_count: u8,
};
const DeviceQualifierDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// DEVICE_QUALIFIER Descriptor Type
descriptor_type: DescriptorType,
// USB Specification Release Number in Binary-Coded Decimal (i.e, 2.00 is expressed as 200h).
// Identifies the release of the USB Specification with with the device and its
// descriptors are compliant. This field must be at least 0200h.
bcd_usb: u16 align(1),
// Class code (assigned by the USB-IF)
device_class: u8,
// Subclass Code (assigned by the USB-IF)
device_subclass: u8,
// Protocol code (assigned by the USB-IF)
device_protocol: u8,
// Maximum packet size for endpoint zero (8, 16, 32, or 64 are the only valid options)
max_packet_size_0: u8,
// Number of possible configurations
configuration_count: u8,
// Reserved for future uses, must be zero.
reserved: u8,
};
const ConfigurationDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// CONFIGURATION Descriptor Type
descriptor_type: DescriptorType,
// The total combined length in bytes of all the descriptors returned with the request for
// this CONFIGURATION descriptor (including CONFIGURATION, INTERFACE, ENDPOINT, class- and
// vendor-specific descriptors).
total_length: u16 align(1),
// Number of interfaces supported by this configuration
interface_count: u8,
// Value which when used as an argument in the SET_CONFIGURATION request, causes the device
// to assume the configuration described by this descriptor.
configuration_value: ConfigurationValue,
// Index of STRING descriptor describing this configuration.
configuration_index: StringIndex,
// Configuration Characteristics
attributes: Attributes,
// Maximum power consumption of this device from the bus when fully operational and using
// this configuration. Expressed in units of 2mA (i.e., a value of 50 in this field
// indicates 100mA).
// - A device reports with the attributes field whether the configuration is bus- or
// self-powered, but the device status (retrieved with a GET_STATUS request) reports
// whether the device is currently self-powered.
// - If a device is disconnected from an external power source, it may not draw more
// power from the bus than specified in this field.
max_power: u8,
// Configuration characteristics. Fields are declared least-significant first.
const Attributes = packed struct(u8) {
// Reserved, reset to zero (D4...0)
reserved: u5,
// Whether Remote Wakeup is supported by this configuration (D5)
remote_wakeup: bool,
// Self-Powered (D6)
// - false: Device runs on power supplied by the bus
// - true: Device provides a local power source; if max_power is non-zero, the
// device also may use bus power.
self_powered: bool,
// Reserved, must be set to one for historical reasons (D7)
reserved_one: u1,
};
};
// This descriptor describes the configuration of a high-speed device if it were operating at
// its alternative speed. The structure of the OTHER_SPEED_CONFIGURATION is identical to that
// of the CONFIGURATION descriptor; the only difference is that the descriptor_type field
// reflects that the descriptor is an OTHER_SPEED_CONFIGURATION descriptor.
const OtherSpeedConfigurationDescriptor = ConfigurationDescriptor;
const InterfaceDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// INTERFACE Descriptor Type
descriptor_type: DescriptorType,
// Number of this interface. Zero-based value which identifies the index of this interface
// in the array of interfaces supported within a configuration.
interface_number: InterfaceNumber,
// Value used to select the alternate settings described by this INTERFACE descriptor for
// the interface with the interface_number in the previous field. This value is zero if
// this descriptor describes the default settings for a particular interface.
alternate_setting: AlternateSetting,
// Number of endpoints used by this interface, not including endpoint zero.
endpoint_count: u8,
// Class code (assigned by the USB-IF)
// - A value of zero here is reserved for future standardization.
// - If this value is FFh, the interface class is vendor-specific.
// - All other values are reserved for assignment by the USB-IF.
interface_class: u8,
// Subclass code (assigned by the USB-IF)
// - The subclass code in this field is qualified by the value of the interface_class
// field.
// - If interface_class is reset to zero, then this field must also be reset to zero.
// - If interface_class is not set to the value of FFh, then all values of this field are
// reserved for assignment by the USB-IF.
interface_subclass: u8,
// Protocol code (assigned by the USB-IF)
// - The protocol code in this field is qualified by the values of the interface_class
// and interface_subclass fields.
// - If an interface supports class-specific requests, then this field identifies the
// protocols that the device uses as defined by the specifications of the device class.
// - If this field is reset to zero, then the device does not use a class-specific
// protocol on this interface.
// - If this field is set to FFh, then the device uses a vendor-specific protocol on
// this interface.
interface_protocol: u8,
// Index of STRING descriptor describing this interface
interface_index: StringIndex,
};
const EndpointDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// ENDPOINT Descriptor Type
descriptor_type: DescriptorType,
// The address of the endpoint on the USB device described by this descriptor
endpoint_address: Address,
// The endpoint's attributes
attributes: Attributes,
// Maximum packet size that this endpoint is capable of sending or receiving. For
// isochronous endpoints, this value is used to reserve bus time; the pipe, however, may
// not always use all of the reserved bus time.
max_packet_size: MaxPacketSize align(1),
// Interval for polling a device during a data transfer, expressed in units of microframes
// for high-speed devices, and frames for low- and full-speed devices. The exact meaning of
// the value in this field depends on the endpoint type and the operating speed of the
// device:
// - Full- and High-speed isochronous endpoints, and high-speed interrupt endpoints:
// This field must be in the range from 1 to 16, and is used to calculate the period
// as 2^(interval - 1). That is, a value of 4 calculates to 2^(4 - 1) = 2^3 = 8.
// - Full- and Low-speed interrupt endpoints: This field must be in the range from
// 1 to 255.
// - High-speed bulk and control OUT endpoints: This field must be in the range from
// 0 to 255, and specifies the maximum NAK rate of the endpoint. A value of zero
// indicates that the endpoint never NAKs; other values indicate at most 1 NAK each
// interval number of microframes.
interval: u8,
// The address of an endpoint. Fields are declared least-significant first.
const Address = packed struct(u8) {
// Endpoint Number (D3...0)
number: EndpointNumber,
// Reserved, reset to zero (D6...4)
reserved: u3,
// Direction, ignored for control endpoints (D7)
direction: EndpointDirection,
};
// An endpoint's attributes. Fields are declared least-significant first.
const Attributes = packed struct(u8) {
// Transfer Type (D1...0)
transfer_type: TransferType,
// Synchronization Type; isochronous endpoints only, reserved and reset to zero for
// other endpoint types (D3...2)
synchronization: Synchronization,
// Usage Type; isochronous endpoints only, reserved and reset to zero for other
// endpoints (D5...4)
usage: Usage,
// Reserved, reset to zero (D7...6)
reserved: u2,
};
const TransferType = enum(u2) {
control = 0,
isochronous = 1,
bulk = 2,
interrupt = 3,
};
const Synchronization = enum(u2) {
none = 0,
asynchronous = 1,
adaptive = 2,
synchronous = 3,
};
const Usage = enum(u2) {
data = 0,
feedback = 1,
implicit_feedback_data = 2,
_,
};
// The maximum packet size of an endpoint. Fields are declared least-significant first.
const MaxPacketSize = packed struct(u16) {
// Maximum packet size in bytes (bits 10...0)
size: u11,
// Number of additional transaction opportunities per microframe, for high-speed
// isochronous and interrupt endpoints; reserved and reset to zero for other
// endpoints (bits 12...11)
additional_transactions: AdditionalTransactions,
// Reserved, must be reset to zero (bits 15...13)
reserved: u3,
};
const AdditionalTransactions = enum(u2) {
// None (1 transaction per microframe)
none = 0,
// 1 additional (2 transactions per microframe)
one = 1,
// 2 additional (3 transactions per microframe)
two = 2,
_,
};
};
// A STRING descriptor at index zero returns the list of LANGID codes supported by the
// device; all other indices return a Unicode string. Both forms start with this two-byte
// header, followed by the variable-length payload:
// - index 0: an array of two-byte LANGID codes (wLangID[0] through wLangID[x])
// - other indices: a Unicode string of N bytes
const StringDescriptor = extern struct {
// Size of this descriptor in bytes
length: u8,
// STRING Descriptor Type
descriptor_type: DescriptorType,
};
const std = @import("std");
test "wire sizes and offsets match the specification" {
const expectEqual = std.testing.expectEqual;
try expectEqual(8, @sizeOf(Request));
try expectEqual(18, @sizeOf(DeviceDescriptor));
try expectEqual(10, @sizeOf(DeviceQualifierDescriptor));
try expectEqual(9, @sizeOf(ConfigurationDescriptor));
try expectEqual(9, @sizeOf(InterfaceDescriptor));
try expectEqual(7, @sizeOf(EndpointDescriptor));
try expectEqual(2, @sizeOf(StringDescriptor));
try expectEqual(2, @offsetOf(DeviceDescriptor, "bcd_usb"));
try expectEqual(8, @offsetOf(DeviceDescriptor, "vendor_id"));
try expectEqual(17, @offsetOf(DeviceDescriptor, "configuration_count"));
try expectEqual(2, @offsetOf(ConfigurationDescriptor, "total_length"));
try expectEqual(4, @offsetOf(EndpointDescriptor, "max_packet_size"));
}
test "bitmap packings match the specification" {
const expectEqual = std.testing.expectEqual;
const expect = std.testing.expect;
// bmRequestType for GET_DESCRIPTOR: device-to-host | standard | device = 80h
const request_type = RequestType{
.recipient = .device,
.kind = .standard,
.direction = .device_to_host,
};
try expectEqual(0x80, @as(u8, @bitCast(request_type)));
// wValue for GET_DESCRIPTOR(CONFIGURATION, index 0) = 0200h
const descriptor_value = Request.DescriptorValue{ .kind = .configuration };
try expectEqual(0x0200, @as(u16, @bitCast(descriptor_value)));
// wIndex for the IN endpoint 1 = 0081h
const endpoint_index = Request.EndpointIndex{ .number = @enumFromInt(1), .direction = .in };
try expectEqual(0x0081, @as(u16, @bitCast(endpoint_index)));
// Endpoint address 81h = IN endpoint 1
const address: EndpointDescriptor.Address = @bitCast(@as(u8, 0x81));
try expectEqual(1, @intFromEnum(address.number));
try expectEqual(.in, address.direction);
// Endpoint attributes 03h = interrupt transfer
const attributes: EndpointDescriptor.Attributes = @bitCast(@as(u8, 0x03));
try expectEqual(.interrupt, attributes.transfer_type);
// wMaxPacketSize 0008h = 8 bytes, no additional transactions
const max_packet_size: EndpointDescriptor.MaxPacketSize = @bitCast(@as(u16, 0x0008));
try expectEqual(8, max_packet_size.size);
try expectEqual(.none, max_packet_size.additional_transactions);
// Configuration attributes C0h = self-powered, with the historical D7 bit set
const configuration_attributes: ConfigurationDescriptor.Attributes = @bitCast(@as(u8, 0xC0));
try expect(configuration_attributes.self_powered);
try expect(!configuration_attributes.remote_wakeup);
try expectEqual(1, configuration_attributes.reserved_one);
// GET_STATUS words: device 0001h = self-powered; endpoint 0001h = halted
const device_status: DeviceStatus = @bitCast(@as(u16, 0x0001));
try expect(device_status.self_powered and !device_status.remote_wakeup);
const endpoint_status: EndpointStatus = @bitCast(@as(u16, 0x0001));
try expect(endpoint_status.halted);
// DescriptorType is non-exhaustive: class-specific values (HID = 21h) pass through
const hid_type: DescriptorType = @enumFromInt(0x21);
try expectEqual(0x21, @intFromEnum(hid_type));
try expect(hid_type != .device);
}
fn expectRequestBytes(request: Request, expected: [8]u8) !void {
try std.testing.expectEqualSlices(u8, &expected, std.mem.asBytes(&request));
}
test "standard request constructors encode the specification's set-up packets" {
try expectRequestBytes(getStatus(.device), .{ 0x80, 0, 0, 0, 0, 0, 2, 0 });
try expectRequestBytes(getStatus(.{ .interface = @enumFromInt(3) }), .{ 0x81, 0, 0, 0, 3, 0, 2, 0 });
try expectRequestBytes(getStatus(.{ .endpoint = .{ .number = @enumFromInt(2), .direction = .in } }), .{ 0x82, 0, 0, 0, 0x82, 0, 2, 0 });
try expectRequestBytes(clearFeature(.endpoint_halt, .{ .endpoint = .{ .number = @enumFromInt(1), .direction = .out } }), .{ 0x02, 1, 0, 0, 0x01, 0, 0, 0 });
try expectRequestBytes(setFeature(.device_remote_wakeup, .device), .{ 0x00, 3, 1, 0, 0, 0, 0, 0 });
try expectRequestBytes(setTestMode(.test_packet), .{ 0x00, 3, 2, 0, 0, 0x04, 0, 0 });
try expectRequestBytes(setAddress(@enumFromInt(5)), .{ 0x00, 5, 5, 0, 0, 0, 0, 0 });
try expectRequestBytes(getDescriptor(.device, 0, 0, 18), .{ 0x80, 6, 0, 1, 0, 0, 18, 0 });
try expectRequestBytes(getDescriptor(.string, 2, 0x0409, 255), .{ 0x80, 6, 2, 3, 0x09, 0x04, 255, 0 });
try expectRequestBytes(setDescriptor(.string, 2, 0x0409, 16), .{ 0x00, 7, 2, 3, 0x09, 0x04, 16, 0 });
try expectRequestBytes(getConfiguration(), .{ 0x80, 8, 0, 0, 0, 0, 1, 0 });
try expectRequestBytes(setConfiguration(@enumFromInt(1)), .{ 0x00, 9, 1, 0, 0, 0, 0, 0 });
try expectRequestBytes(getInterface(@enumFromInt(2)), .{ 0x81, 10, 0, 0, 2, 0, 1, 0 });
try expectRequestBytes(setInterface(@enumFromInt(2), @enumFromInt(1)), .{ 0x01, 11, 1, 0, 2, 0, 0, 0 });
try expectRequestBytes(syncFrame(.{ .number = @enumFromInt(3), .direction = .in }), .{ 0x82, 12, 0, 0, 0x83, 0, 2, 0 });
}
+264
View File
@@ -0,0 +1,264 @@
//! USB class-code decoding: turn the (class, subclass, protocol) triple a USB device or
//! interface reports in its descriptors into typed values. The device descriptor carries one
//! triple for the whole device, and each interface descriptor carries its own; a class code
//! of zero at the device level defers entirely to the interfaces. Subclass and protocol
//! codes are qualified by the class code — the same value means different things under
//! different classes — so there is no single SubClass or Protocol enum: each class with
//! spec-defined codes gets its own namespace below. Pure reference data (from the USB-IF
//! defined class codes; see https://www.usb.org/defined-class-codes) — no hardware access —
//! so it is shared by kernel discovery and any user-space tool (device naming, driver
//! matching).
// Base class codes (assigned by the USB-IF). The comment on each value notes where the code
// may legally appear: in the device descriptor, in interface descriptors, or both.
const Class = enum(u8) {
// Use class information in the interface descriptors (device descriptor only). Each
// interface within a configuration specifies its own class information and the various
// interfaces operate independently.
per_interface = 0x00,
// Audio: speakers, microphones, sound cards (interface)
audio = 0x01,
// Communications and CDC control: modems, network adapters (both)
communications = 0x02,
// Human Interface Device: keyboards, mice, game controllers (interface)
hid = 0x03,
// Physical: force-feedback devices (interface)
physical = 0x05,
// Image: still-imaging cameras, scanners (interface)
image = 0x06,
// Printer (interface)
printer = 0x07,
// Mass storage: flash drives, external disks, card readers (interface)
mass_storage = 0x08,
// Hub (device descriptor only)
hub = 0x09,
// CDC-Data: the data interfaces paired with a communications control interface
// (interface)
cdc_data = 0x0A,
// Smart card readers (interface)
smart_card = 0x0B,
// Content security (interface)
content_security = 0x0D,
// Video: webcams (interface)
video = 0x0E,
// Personal healthcare devices (interface)
personal_healthcare = 0x0F,
// Audio/Video devices (interface)
audio_video = 0x10,
// Billboard: describes alternate modes a USB Type-C device supports (device descriptor
// only)
billboard = 0x11,
// USB Type-C bridge (interface)
type_c_bridge = 0x12,
// USB Bulk Display Protocol devices (interface)
bulk_display = 0x13,
// MCTP over USB protocol endpoint devices (interface)
mctp = 0x14,
// I3C devices (interface)
i3c = 0x3C,
// Diagnostic devices (both)
diagnostic = 0xDC,
// Wireless controllers: Bluetooth adapters (interface)
wireless_controller = 0xE0,
// Miscellaneous (both)
miscellaneous = 0xEF,
// Application-specific: firmware upgrade, IrDA bridges, test and measurement
// (interface)
application_specific = 0xFE,
// Vendor-specific (both)
vendor_specific = 0xFF,
_,
};
// Subclass and protocol codes qualified by Class.hub. Hubs have no subclass codes; the
// protocol distinguishes the hub's transaction-translator arrangement.
const hub = struct {
const Protocol = enum(u8) {
// Full-speed hub
full_speed = 0x00,
// Hi-speed hub with a single transaction translator
hi_speed_single_tt = 0x01,
// Hi-speed hub with multiple transaction translators
hi_speed_multi_tt = 0x02,
// SuperSpeed hub (USB 3)
super_speed = 0x03,
_,
};
};
// Subclass and protocol codes qualified by Class.hid.
const hid = struct {
const SubClass = enum(u8) {
// No subclass
none = 0x00,
// Boot interface: the device also supports the simplified boot protocol, usable by
// firmware before a full HID report-descriptor parser is available
boot = 0x01,
_,
};
// Only meaningful when the subclass is boot
const Protocol = enum(u8) {
none = 0x00,
keyboard = 0x01,
mouse = 0x02,
_,
};
};
// Subclass and protocol codes qualified by Class.mass_storage. The subclass identifies the
// command set the device understands; the protocol identifies the transport used to carry
// commands, data, and status over the bus.
const mass_storage = struct {
const SubClass = enum(u8) {
// SCSI command set not reported; de facto, treat as scsi
not_reported = 0x00,
// Reduced Block Commands: typically flash devices
rbc = 0x01,
// MMC-5 (ATAPI): CD and DVD drives
atapi = 0x02,
// QIC-157 tape drives (obsolete)
qic_157 = 0x03,
// UFI: floppy disk drives
ufi = 0x04,
// SFF-8070i (obsolete)
sff_8070i = 0x05,
// Transparent SCSI command set: the common case for flash drives and disks
scsi = 0x06,
// LSD FS: negotiated access to large storage devices
lsd_fs = 0x07,
// IEEE 1667
ieee_1667 = 0x08,
// Vendor-specific
vendor_specific = 0xFF,
_,
};
const Protocol = enum(u8) {
// Control/Bulk/Interrupt with command completion interrupt
cbi_completion_interrupt = 0x00,
// Control/Bulk/Interrupt without command completion interrupt
cbi = 0x01,
// Bulk-only transport: the common case for flash drives and disks
bulk_only = 0x50,
// USB attached SCSI
uas = 0x62,
// Vendor-specific
vendor_specific = 0xFF,
_,
};
};
// Subclass and protocol codes qualified by Class.communications (CDC). The protocol codes
// are model-specific; the useful invariant is the subclass, which selects the control model
// the interface implements.
const communications = struct {
const SubClass = enum(u8) {
// Direct line control model
direct_line = 0x01,
// Abstract control model: USB modems and serial adapters
abstract_control = 0x02,
// Telephone control model
telephone = 0x03,
// Multi-channel control model
multi_channel = 0x04,
// CAPI control model
capi = 0x05,
// Ethernet networking control model
ethernet = 0x06,
// ATM networking control model
atm = 0x07,
// Wireless handset control model
wireless_handset = 0x08,
// Device management
device_management = 0x09,
// Mobile direct line model
mobile_direct_line = 0x0A,
// OBEX
obex = 0x0B,
// Ethernet emulation model
ethernet_emulation = 0x0C,
// Network control model
network_control = 0x0D,
_,
};
};
// Subclass and protocol codes qualified by Class.wireless_controller.
const wireless_controller = struct {
const SubClass = enum(u8) {
// Radio frequency controllers
radio_frequency = 0x01,
_,
};
// Only meaningful when the subclass is radio_frequency
const Protocol = enum(u8) {
// Bluetooth programming interface
bluetooth = 0x01,
// Ultra-wideband radio control
ultra_wideband = 0x02,
// Remote NDIS
remote_ndis = 0x03,
// Bluetooth AMP controller
bluetooth_amp = 0x04,
_,
};
};
// Subclass and protocol codes qualified by Class.miscellaneous.
const miscellaneous = struct {
const SubClass = enum(u8) {
// Common class
common = 0x02,
_,
};
// Only meaningful when the subclass is common
const Protocol = enum(u8) {
// Interface association descriptor: at the device level, announces that the
// configuration groups interfaces into functions with IADs
interface_association = 0x01,
_,
};
};
// Subclass and protocol codes qualified by Class.application_specific.
const application_specific = struct {
const SubClass = enum(u8) {
// Device firmware upgrade
firmware_upgrade = 0x01,
// IrDA bridge
irda_bridge = 0x02,
// Test and measurement
test_and_measurement = 0x03,
_,
};
};
test "class codes match the USB-IF assignments" {
const std = @import("std");
const expectEqual = std.testing.expectEqual;
try expectEqual(0x03, @intFromEnum(Class.hid));
try expectEqual(0x09, @intFromEnum(Class.hub));
try expectEqual(0xFF, @intFromEnum(Class.vendor_specific));
// A typical flash drive: mass storage, transparent SCSI, bulk-only transport.
try expectEqual(0x06, @intFromEnum(mass_storage.SubClass.scsi));
try expectEqual(0x50, @intFromEnum(mass_storage.Protocol.bulk_only));
// A boot keyboard: HID, boot subclass, keyboard protocol.
try expectEqual(0x01, @intFromEnum(hid.SubClass.boot));
try expectEqual(0x01, @intFromEnum(hid.Protocol.keyboard));
// Class codes are non-exhaustive: unlisted values pass through undamaged.
const unknown: Class = @enumFromInt(0x42);
try expectEqual(0x42, @intFromEnum(unknown));
_ = hub.Protocol.hi_speed_multi_tt;
_ = communications.SubClass.abstract_control;
_ = wireless_controller.Protocol.bluetooth;
_ = miscellaneous.Protocol.interface_association;
_ = application_specific.SubClass.firmware_upgrade;
}
+1
View File
@@ -100,6 +100,7 @@ pub fn main() void {
while (n < n_children) : (n += 1) { while (n < n_children) : (n += 1) {
var child = std.mem.zeroes(device.DeviceDescriptor); var child = std.mem.zeroes(device.DeviceDescriptor);
child.class = @intFromEnum(device.DeviceClass.timer); child.class = @intFromEnum(device.DeviceClass.timer);
child.pci_class = device.no_pci_class;
child.hid_len = 6; child.hid_len = 6;
child.hid[0..6].* = "hpet-t".*; child.hid[0..6].* = "hpet-t".*;
child.resource_count = 1; child.resource_count = 1;
+16 -16
View File
@@ -118,8 +118,8 @@ pub fn main() void {
maybe_controller = controller; maybe_controller = controller;
maybe_interrupt_index = findInterruptResourceIndex(controller_device_descriptor); maybe_interrupt_index = findInterruptResourceIndex(controller_device_descriptor);
controller.disablePort(.One); controller.disablePort(.one);
controller.disablePort(.Two); controller.disablePort(.two);
controller.flushOutputBuffer(); controller.flushOutputBuffer();
const current = controller.readConfigurationByte() orelse { const current = controller.readConfigurationByte() orelse {
@@ -136,7 +136,7 @@ pub fn main() void {
return; return;
} }
if (controller.performSelfTest()) | reply | { if (controller.performSelfTest()) |reply| {
if (reply != ps2.response_controller_test_passed) { if (reply != ps2.response_controller_test_passed) {
_ = runtime.system.write("system/drivers/ps2-bus: perform controller self test failed\n"); _ = runtime.system.write("system/drivers/ps2-bus: perform controller self test failed\n");
return; return;
@@ -154,20 +154,20 @@ pub fn main() void {
if (has_two_channels) { if (has_two_channels) {
_ = runtime.system.write("system/drivers/ps2-bus: has two channels\n"); _ = runtime.system.write("system/drivers/ps2-bus: has two channels\n");
// keep the bus quiet until we have tested the ports and are ready to use them // keep the bus quiet until we have tested the ports and are ready to use them
controller.disablePort(.Two); controller.disablePort(.two);
} else { } else {
_ = runtime.system.write("system/drivers/ps2-bus: has one channel\n"); _ = runtime.system.write("system/drivers/ps2-bus: has one channel\n");
} }
// interface tests: always test port 1, test port 2 only if it exists // interface tests: always test port 1, test port 2 only if it exists
const port_one_works = (controller.testPort(.One) orelse { const port_one_works = (controller.testPort(.one) orelse {
_ = runtime.system.write("system/drivers/ps2-bus: port 1 test timed out\n"); _ = runtime.system.write("system/drivers/ps2-bus: port 1 test timed out\n");
return; return;
}) == ps2.response_port_test_passed; }) == ps2.response_port_test_passed;
var port_two_works = false; var port_two_works = false;
if (has_two_channels) { if (has_two_channels) {
port_two_works = (controller.testPort(.Two) orelse { port_two_works = (controller.testPort(.two) orelse {
_ = runtime.system.write("system/drivers/ps2-bus: port 2 test timed out\n"); _ = runtime.system.write("system/drivers/ps2-bus: port 2 test timed out\n");
return; return;
}) == ps2.response_port_test_passed; }) == ps2.response_port_test_passed;
@@ -181,20 +181,20 @@ pub fn main() void {
// Enable the working ports. Their interrupts stay off until IRQ1 is bound // Enable the working ports. Their interrupts stay off until IRQ1 is bound
// below — reset and identify use polled reads, which must never race the // below — reset and identify use polled reads, which must never race the
// interrupt-driven drain loop for bytes. // interrupt-driven drain loop for bytes.
controller.enablePort(.One); controller.enablePort(.one);
if (port_two_works) controller.enablePort(.Two); if (port_two_works) controller.enablePort(.two);
// reset each working device; a failing device is logged but does not // reset each working device; a failing device is logged but does not
// abort bring-up of the other one // abort bring-up of the other one
if (port_one_works) { if (port_one_works) {
if (controller.resetDevice(.One)) |passed| { if (controller.resetDevice(.one)) |passed| {
if (!passed) _ = runtime.system.write("system/drivers/ps2-bus: port 1 device reset failed\n"); if (!passed) _ = runtime.system.write("system/drivers/ps2-bus: port 1 device reset failed\n");
} else { } else {
_ = runtime.system.write("system/drivers/ps2-bus: port 1 device reset timed out\n"); _ = runtime.system.write("system/drivers/ps2-bus: port 1 device reset timed out\n");
} }
} }
if (port_two_works) { if (port_two_works) {
if (controller.resetDevice(.Two)) |passed| { if (controller.resetDevice(.two)) |passed| {
if (!passed) _ = runtime.system.write("system/drivers/ps2-bus: port 2 device reset failed\n"); if (!passed) _ = runtime.system.write("system/drivers/ps2-bus: port 2 device reset failed\n");
} else { } else {
_ = runtime.system.write("system/drivers/ps2-bus: port 2 device reset timed out\n"); _ = runtime.system.write("system/drivers/ps2-bus: port 2 device reset timed out\n");
@@ -204,8 +204,8 @@ pub fn main() void {
// Identify the device on each working port and hand it off to the driver // Identify the device on each working port and hand it off to the driver
// that matches what it reported — a port is not assumed to be a keyboard // that matches what it reported — a port is not assumed to be a keyboard
// or a mouse by its number. // or a mouse by its number.
if (port_one_works) port_device_types[@intFromEnum(ps2.Port.One)] = spawnIdentifiedDriver(controller, .One); if (port_one_works) port_device_types[@intFromEnum(ps2.Port.one)] = spawnIdentifiedDriver(controller, .one);
if (port_two_works) port_device_types[@intFromEnum(ps2.Port.Two)] = spawnIdentifiedDriver(controller, .Two); if (port_two_works) port_device_types[@intFromEnum(ps2.Port.two)] = spawnIdentifiedDriver(controller, .two);
} else { } else {
_ = runtime.system.write("system/drivers/ps2-bus: no PS/2 controller found\n"); _ = runtime.system.write("system/drivers/ps2-bus: no PS/2 controller found\n");
return; return;
@@ -243,7 +243,7 @@ pub fn main() void {
// claim that node too and route its IRQ to the same endpoint. The IRQ belongs // claim that node too and route its IRQ to the same endpoint. The IRQ belongs
// to the *port*, whatever device identify found on it. // to the *port*, whatever device identify found on it.
var maybe_auxiliary_interrupt: ?struct { device_id: u64, interrupt_index: u64, gsi: u64 } = null; var maybe_auxiliary_interrupt: ?struct { device_id: u64, interrupt_index: u64, gsi: u64 } = null;
if (port_device_types[@intFromEnum(ps2.Port.Two)] != null) { if (port_device_types[@intFromEnum(ps2.Port.two)] != null) {
if (device.findDeviceDescriptorByHid(buffer, acpi_ids.HardwareId.ps2_mouse.hid())) |descriptor| { if (device.findDeviceDescriptorByHid(buffer, acpi_ids.HardwareId.ps2_mouse.hid())) |descriptor| {
if (findInterruptResourceIndex(descriptor)) |auxiliary_index| { if (findInterruptResourceIndex(descriptor)) |auxiliary_index| {
if (device.claim(descriptor.id) and device.irqBind(descriptor.id, auxiliary_index, endpoint)) { if (device.claim(descriptor.id) and device.irqBind(descriptor.id, auxiliary_index, endpoint)) {
@@ -263,8 +263,8 @@ pub fn main() void {
_ = runtime.system.write("system/drivers/ps2-bus: controller configuration timed out\n"); _ = runtime.system.write("system/drivers/ps2-bus: controller configuration timed out\n");
return; return;
}; };
if (port_device_types[@intFromEnum(ps2.Port.One)] != null) configuration |= ps2.Port.One.interruptBit(); if (port_device_types[@intFromEnum(ps2.Port.one)] != null) configuration |= ps2.Port.one.interruptBit();
if (maybe_auxiliary_interrupt != null) configuration |= ps2.Port.Two.interruptBit(); if (maybe_auxiliary_interrupt != null) configuration |= ps2.Port.two.interruptBit();
_ = controller.writeConfigurationByte(configuration); _ = controller.writeConfigurationByte(configuration);
_ = runtime.system.write("system/drivers/ps2-bus: ok\n"); _ = runtime.system.write("system/drivers/ps2-bus: ok\n");
@@ -284,7 +284,7 @@ pub fn main() void {
const current_status = ps2.status(controller.device_id, controller.status_index); const current_status = ps2.status(controller.device_id, controller.status_index);
if (current_status & ps2.status_output_buffer_full == 0) break; if (current_status & ps2.status_output_buffer_full == 0) break;
const byte = device.ioRead(controller.device_id, controller.data_index, 0, 1) orelse break; const byte = device.ioRead(controller.device_id, controller.data_index, 0, 1) orelse break;
const port: ps2.Port = if (current_status & ps2.status_auxiliary_output != 0) .Two else .One; const port: ps2.Port = if (current_status & ps2.status_auxiliary_output != 0) .two else .one;
if (port_endpoints[@intFromEnum(port)]) |child| { if (port_endpoints[@intFromEnum(port)]) |child| {
const forwarded = ps2.ForwardedByte{ .port = @intFromEnum(port), .byte = byte }; const forwarded = ps2.ForwardedByte{ .port = @intFromEnum(port), .byte = byte };
_ = ipc.send(child, std.mem.asBytes(&forwarded)); _ = ipc.send(child, std.mem.asBytes(&forwarded));
+26 -25
View File
@@ -146,52 +146,53 @@ pub fn readData(id: u64, status_index: u64, data_index: u64, timeout_nanoseconds
} }
pub fn writeData(id: u64, status_index: u64, data_index: u64, byte: u8, timeout_nanoseconds: u64) bool { pub fn writeData(id: u64, status_index: u64, data_index: u64, byte: u8, timeout_nanoseconds: u64) bool {
// IBF lives in the status register (0x64); wait for it to clear there, then write the data port (0x60) // IBF lives in the status register (0x64); wait for it to clear there, then write the data port
// (0x60)
if (!waitWritable(id, status_index, timeout_nanoseconds)) return false; if (!waitWritable(id, status_index, timeout_nanoseconds)) return false;
return device.ioWrite(id, data_index, 0, 1, byte); return device.ioWrite(id, data_index, 0, 1, byte);
} }
pub const Port = enum(u2) { pub const Port = enum(u2) {
One, one,
Two, two,
/// Command register byte that disables this port. /// Command register byte that disables this port.
fn disableCommand(self: Port) u8 { fn disableCommand(self: Port) u8 {
return switch (self) { return switch (self) {
.One => cmd_disable_first_port, .one => cmd_disable_first_port,
.Two => cmd_disable_second_port, .two => cmd_disable_second_port,
}; };
} }
/// Command register byte that enables this port (and its clock). /// Command register byte that enables this port (and its clock).
fn enableCommand(self: Port) u8 { fn enableCommand(self: Port) u8 {
return switch (self) { return switch (self) {
.One => cmd_enable_first_port, .one => cmd_enable_first_port,
.Two => cmd_enable_second_port, .two => cmd_enable_second_port,
}; };
} }
/// Command register byte that runs this port's interface test. /// Command register byte that runs this port's interface test.
fn testCommand(self: Port) u8 { fn testCommand(self: Port) u8 {
return switch (self) { return switch (self) {
.One => cmd_test_first_port, .one => cmd_test_first_port,
.Two => cmd_test_second_port, .two => cmd_test_second_port,
}; };
} }
/// Configuration-byte bit that, when set, disables this port's clock. /// Configuration-byte bit that, when set, disables this port's clock.
pub fn clockDisabledBit(self: Port) u8 { pub fn clockDisabledBit(self: Port) u8 {
return switch (self) { return switch (self) {
.One => configuration_first_port_clock_disabled, .one => configuration_first_port_clock_disabled,
.Two => configuration_second_port_clock_disabled, .two => configuration_second_port_clock_disabled,
}; };
} }
/// Configuration-byte bit that, when set, enables this port's interrupt. /// Configuration-byte bit that, when set, enables this port's interrupt.
pub fn interruptBit(self: Port) u8 { pub fn interruptBit(self: Port) u8 {
return switch (self) { return switch (self) {
.One => configuration_first_port_interrupt, .one => configuration_first_port_interrupt,
.Two => configuration_second_port_interrupt, .two => configuration_second_port_interrupt,
}; };
} }
@@ -199,8 +200,8 @@ pub const Port = enum(u2) {
/// output buffer (makes a byte appear as if it came from the device). /// output buffer (makes a byte appear as if it came from the device).
pub fn writeOutputBufferCommand(self: Port) u8 { pub fn writeOutputBufferCommand(self: Port) u8 {
return switch (self) { return switch (self) {
.One => cmd_write_first_port_output, .one => cmd_write_first_port_output,
.Two => cmd_write_second_port_output, .two => cmd_write_second_port_output,
}; };
} }
@@ -209,24 +210,24 @@ pub const Port = enum(u2) {
/// prefix (null); port 2 requires the "write second port input" command. /// prefix (null); port 2 requires the "write second port input" command.
pub fn deviceInputCommand(self: Port) ?u8 { pub fn deviceInputCommand(self: Port) ?u8 {
return switch (self) { return switch (self) {
.One => null, .one => null,
.Two => cmd_write_second_port_input, .two => cmd_write_second_port_input,
}; };
} }
/// Controller output-port bit driving this port's clock line. /// Controller output-port bit driving this port's clock line.
pub fn outputPortClockBit(self: Port) u8 { pub fn outputPortClockBit(self: Port) u8 {
return switch (self) { return switch (self) {
.One => output_port_first_port_clock, .one => output_port_first_port_clock,
.Two => output_port_second_port_clock, .two => output_port_second_port_clock,
}; };
} }
/// Controller output-port bit driving this port's data line. /// Controller output-port bit driving this port's data line.
pub fn outputPortDataBit(self: Port) u8 { pub fn outputPortDataBit(self: Port) u8 {
return switch (self) { return switch (self) {
.One => output_port_first_port_data, .one => output_port_first_port_data,
.Two => output_port_second_port_data, .two => output_port_second_port_data,
}; };
} }
@@ -234,8 +235,8 @@ pub const Port = enum(u2) {
/// (wired to the port's IRQ line). /// (wired to the port's IRQ line).
pub fn outputPortBufferFullBit(self: Port) u8 { pub fn outputPortBufferFullBit(self: Port) u8 {
return switch (self) { return switch (self) {
.One => output_port_first_port_output_full, .one => output_port_first_port_output_full,
.Two => output_port_second_port_output_full, .two => output_port_second_port_output_full,
}; };
} }
}; };
@@ -392,9 +393,9 @@ pub const Controller = struct {
/// enabled; the caller should disable it again to keep the bus quiet until /// enabled; the caller should disable it again to keep the bus quiet until
/// device bring-up. /// device bring-up.
pub fn hasTwoChannels(self: Controller) ?bool { pub fn hasTwoChannels(self: Controller) ?bool {
self.enablePort(.Two); self.enablePort(.two);
const configuration = self.readConfigurationByte() orelse return null; const configuration = self.readConfigurationByte() orelse return null;
return (configuration & Port.Two.clockDisabledBit()) == 0; return (configuration & Port.two.clockDisabledBit()) == 0;
} }
/// Reset the device attached to `port` (device command 0xFF) and wait for /// Reset the device attached to `port` (device command 0xFF) and wait for
@@ -0,0 +1,71 @@
//! /system/drivers/usb-xhci-bus — the xHCI (USB 3) host-controller bus driver.
//! The device manager spawns **one instance per controller** it discovers (a machine
//! can carry several), passing the controller's device-tree id as argv[1]; this
//! instance claims that device and no other, so multiple instances never fight over
//! hardware. This increment proves the plumbing: parse the id, claim the controller,
//! and report its MMIO window. The next increments map the registers and bring the
//! controller up (reset, rings, port scan), then enumerate the USB devices on the
//! bus with the usb-abi request builders and publish each with `device_register`.
const std = @import("std");
const runtime = @import("runtime");
const device = runtime.device;
/// Format one whole log line and emit it in a single `debug_write`, so concurrent
/// instances (one per controller) can never interleave mid-line.
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
pub fn main(init: runtime.process.Init) void {
const argument = init.arguments.get(1) orelse {
_ = runtime.system.write("usb-xhci-bus: missing controller device id (argv[1])\n");
return;
};
const controller_id = std.fmt.parseInt(u64, argument, 10) catch {
writeLine("usb-xhci-bus: malformed controller device id '{s}'\n", .{argument});
return;
};
if (!device.claim(controller_id)) {
writeLine("usb-xhci-bus: unable to claim controller device {d}\n", .{controller_id});
return;
}
// Fetch our own descriptor back for the controller's resources.
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("usb-xhci-bus: out of memory\n");
return;
};
const total = device.enumerate(buffer);
const descriptor = for (buffer[0..@min(total, buffer.len)]) |d| {
if (d.id == controller_id) break d;
} else {
writeLine("usb-xhci-bus: device {d} not in the device tree\n", .{controller_id});
return;
};
// The controller's operational registers live behind BAR0, enumerated as the
// device's first memory resource.
const register_window = for (descriptor.resources[0..@intCast(descriptor.resource_count)]) |resource| {
if (resource.kind == @intFromEnum(device.ResourceKind.memory)) break resource;
} else {
writeLine("usb-xhci-bus: controller device {d} has no MMIO window\n", .{controller_id});
return;
};
writeLine("usb-xhci-bus: claimed controller device {d} (registers at 0x{x}, {d} bytes)\n", .{
controller_id,
register_window.start,
register_window.len,
});
// Controller bring-up (map the window, reset, rings, port scan) is the next
// increment; stay resident as the bus's supervisor in the meantime.
while (true) runtime.system.sleep(1000);
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
+1 -2
View File
@@ -490,8 +490,7 @@ pub fn saveInterrupts() u64 {
\\cli \\cli
: [f] "=r" (flags), : [f] "=r" (flags),
: :
: .{ .memory = true } : .{ .memory = true });
);
return flags; return flags;
} }
+2 -4
View File
@@ -166,8 +166,7 @@ pub fn init(allocFrame: *const fn () ?u64, freeFrame: *const fn (u64) void, boot
asm volatile ("mov %[pml4], %%cr3" asm volatile ("mov %[pml4], %%cr3"
: :
: [pml4] "r" (pml4), : [pml4] "r" (pml4),
: .{ .memory = true } : .{ .memory = true });
);
on_own_tables = true; // now on the kernel's physmap (covers all RAM) on_own_tables = true; // now on the kernel's physmap (covers all RAM)
init_done = true; // the kernel half is fixed from here init_done = true; // the kernel half is fixed from here
} }
@@ -409,6 +408,5 @@ fn invalidate(virtual: u64) void {
\\invlpg (%%rax) \\invlpg (%%rax)
: :
: [v] "r" (virtual), : [v] "r" (virtual),
: .{ .rax = true, .memory = true } : .{ .rax = true, .memory = true });
);
} }
-2
View File
@@ -154,5 +154,3 @@ pub const Console = struct {
while (x < self.fb.width) : (x += 1) destination[x] = source[x]; while (x < self.fb.width) : (x += 1) destination[x] = source[x];
} }
}; };
+15
View File
@@ -66,6 +66,7 @@ fn record(node: *platform.Device, parent_id: u64) u64 {
d.id = count; d.id = count;
d.parent = parent_id; d.parent = parent_id;
d.class = @intFromEnum(node.class); d.class = @intFromEnum(node.class);
d.pci_class = if (node.ids.pci_class) |code| code else device_abi.no_pci_class;
const h = node.hid(); const h = node.hid();
d.hid_len = @min(h.len, d.hid.len); d.hid_len = @min(h.len, d.hid.len);
@memcpy(d.hid[0..d.hid_len], h[0..d.hid_len]); @memcpy(d.hid[0..d.hid_len], h[0..d.hid_len]);
@@ -103,6 +104,19 @@ pub fn ownerOf(id: u64) ?u32 {
return claimed[@intCast(id)]; return claimed[@intCast(id)];
} }
/// Release every claim held by `owner` — called by the process layer on every
/// path out of a process (exit, fault, kill), so a restarted driver can claim its
/// hardware again (docs/process-lifecycle.md iron rule 1: cleanup is the kernel's
/// job). The devices stay in the table — they describe hardware, which did not go
/// away — only their ownership clears.
pub fn releaseAllOwnedBy(owner: u32) void {
for (claimed[0..count]) |*slot| {
if (slot.*) |o| {
if (o == owner) slot.* = null;
}
}
}
/// Resource `index` of device `id`, or null if out of range. /// Resource `index` of device `id`, or null if out of range.
pub fn resourceOf(id: u64, index: u64) ?device_abi.ResourceDescriptor { pub fn resourceOf(id: u64, index: u64) ?device_abi.ResourceDescriptor {
if (id >= count) return null; if (id >= count) return null;
@@ -170,6 +184,7 @@ pub fn register(parent_id: u64, owner: u32, descriptor: *const device_abi.Device
d.id = count; d.id = count;
d.parent = parent_id; d.parent = parent_id;
d.class = descriptor.class; d.class = descriptor.class;
d.pci_class = descriptor.pci_class;
d.hid_len = @min(descriptor.hid_len, d.hid.len); d.hid_len = @min(descriptor.hid_len, d.hid.len);
@memcpy(d.hid[0..@intCast(d.hid_len)], descriptor.hid[0..@intCast(d.hid_len)]); @memcpy(d.hid[0..@intCast(d.hid_len)], descriptor.hid[0..@intCast(d.hid_len)]);
d.resource_count = descriptor.resource_count; d.resource_count = descriptor.resource_count;
+15 -2
View File
@@ -87,7 +87,7 @@ fn kmain(boot_information: *const BootInformation) noreturn {
log.print(" pitch : {d} bytes\n", .{fb.pitch}); log.print(" pitch : {d} bytes\n", .{fb.pitch});
log.print(" format : {s}\n", .{@tagName(fb.format)}); log.print(" format : {s}\n", .{@tagName(fb.format)});
log.print(" framebuffer: 0x{x:0>16}\n", .{fb.base}); log.print(" framebuffer: 0x{x:0>16}\n", .{fb.base});
log.print (" footprint : {d} MiB\n", .{(fb.pitch * fb.height) / (1024 * 1024)}); log.print(" footprint : {d} MiB\n", .{(fb.pitch * fb.height) / (1024 * 1024)});
// Summarise the physical memory the loader handed us. The array is danos's // Summarise the physical memory the loader handed us. The array is danos's
// own MemoryRegion, so this is a plain slice — no firmware layout in sight. // own MemoryRegion, so this is a plain slice — no firmware layout in sight.
@@ -412,13 +412,26 @@ fn recoverableFault(vector: u64) bool {
/// plus a POST code and a persistent breadcrumb. (A ring-3 fault on a *borrowed* /// plus a POST code and a persistent breadcrumb. (A ring-3 fault on a *borrowed*
/// kernel thread — process.run, the user-pf isolation probe — also lands here: there /// kernel thread — process.run, the user-pf isolation probe — also lands here: there
/// is no scheduled process to kill.) /// is no scheduled process to kill.)
/// Classify a CPU exception vector as the ExitReason a supervisor reads — the
/// fault classes of docs/process-lifecycle.md. Faults are exit reasons, never
/// signals delivered to the faulting process: recovery is restart, not a handler.
fn exitReasonForVector(vector: u64) abi.ExitReason {
return switch (vector) {
14 => .segmentation_fault, // page fault
6 => .illegal_instruction, // invalid opcode
0, 16, 19 => .arithmetic_fault, // divide error, x87, SIMD
13 => .protection_fault, // general protection
else => .fault,
};
}
fn onException(state: *const architecture.CpuState) noreturn { fn onException(state: *const architecture.CpuState) noreturn {
if (architecture.fromUser(state) and scheduler.currentIsUserProcess() and recoverableFault(state.vector)) { if (architecture.fromUser(state) and scheduler.currentIsUserProcess() and recoverableFault(state.vector)) {
statusPrint("\ndanos: process {d} ({s}) killed by {s} (vector {d}) on core {d}\n", .{ scheduler.currentId(), scheduler.current().name(), architecture.exceptionName(state.vector), state.vector, scheduler.currentCpuIndex() }); statusPrint("\ndanos: process {d} ({s}) killed by {s} (vector {d}) on core {d}\n", .{ scheduler.currentId(), scheduler.current().name(), architecture.exceptionName(state.vector), state.vector, scheduler.currentCpuIndex() });
statusPrint(" error code : 0x{x}\n", .{state.error_code}); statusPrint(" error code : 0x{x}\n", .{state.error_code});
statusPrint(" IP : 0x{x:0>16}\n", .{architecture.instructionPointer(state)}); statusPrint(" IP : 0x{x:0>16}\n", .{architecture.instructionPointer(state)});
if (architecture.faultAddress(state)) |address| statusPrint(" fault addr : 0x{x:0>16}\n", .{address}); if (architecture.faultAddress(state)) |address| statusPrint(" fault addr : 0x{x:0>16}\n", .{address});
process.killCurrentProcess(); // reclaims everything, reschedules; never returns process.killCurrentProcess(exitReasonForVector(state.vector)); // reclaims everything, reschedules; never returns
} }
log.checkpoint(cp_exception); log.checkpoint(cp_exception);
+205 -3
View File
@@ -137,6 +137,7 @@ pub fn init() void {
architecture.setSystemCallHandler(system_call); architecture.setSystemCallHandler(system_call);
scheduler.terminate_current_hook = terminateCurrentLocked; scheduler.terminate_current_hook = terminateCurrentLocked;
scheduler.reap_task_hook = reapTaskLocked; scheduler.reap_task_hook = reapTaskLocked;
scheduler.timer_tick_hook = timerSweepLocked;
} }
/// Return -1 (as an unsigned bit pattern) in the system_call result register. /// Return -1 (as an unsigned bit pattern) in the system_call result register.
@@ -164,6 +165,7 @@ fn system_call(state: *architecture.CpuState) void {
// A scheduled process tears down fully (terminateCurrent); a borrowed // A scheduled process tears down fully (terminateCurrent); a borrowed
// test thread unwinds back to the kernel that entered it. // test thread unwinds back to the kernel that entered it.
if (scheduler.currentIsUserProcess()) { if (scheduler.currentIsUserProcess()) {
scheduler.current().exit_reason = .exited;
terminateCurrent(); terminateCurrent();
} else architecture.userExit(); } else architecture.userExit();
}, },
@@ -199,6 +201,11 @@ fn system_call(state: *architecture.CpuState) void {
.clock => systemClock(state), .clock => systemClock(state),
.process_enumerate => systemProcessEnumerate(state), .process_enumerate => systemProcessEnumerate(state),
.process_kill => systemProcessKill(state), .process_kill => systemProcessKill(state),
.process_exit_reason => systemProcessExitReason(state),
.process_subscribe => systemProcessSubscribe(state),
.signal_bind => systemSignalBind(state),
.process_signal => systemProcessSignal(state),
.timer_bind => systemTimerBind(state),
_ => fail(state), _ => fail(state),
} }
} }
@@ -552,7 +559,11 @@ pub var fault_kill_count: u64 = 0;
/// endpoint reference destroys the Endpoint, and a still-bound GSI would have an /// endpoint reference destroys the Endpoint, and a still-bound GSI would have an
/// ISR call notifyFromIsr on freed memory the next time the device fired. /// ISR call notifyFromIsr on freed memory the next time the device fired.
/// `releaseOwner` also leaves the line masked, so a dead driver's device goes /// `releaseOwner` also leaves the line masked, so a dead driver's device goes
/// quiet rather than storming. /// quiet rather than storming. (It drops MSI vectors by the same owner sweep.)
/// - Device claims are released with the IRQ bindings, so a restarted driver can
/// claim the same hardware again — the cleanup half of process-lifecycle.md's
/// iron rule 1. Claims hold no pointers, so ordering is free; they go here so
/// the exit notification (below, last) observes a fully-released child.
/// - A client this task still owes a reply to (it died between receive and reply) /// - A client this task still owes a reply to (it died between receive and reply)
/// is failed with -EPEER rather than left blocked forever — a dead server must /// is failed with -EPEER rather than left blocked forever — a dead server must
/// not hang its callers. /// not hang its callers.
@@ -565,7 +576,33 @@ pub var fault_kill_count: u64 = 0;
/// reference taken at spawn is dropped with it. /// reference taken at spawn is dropped with it.
/// Precondition: the big kernel lock is held. /// Precondition: the big kernel lock is held.
fn releaseTaskResourcesLocked(t: *scheduler.Task) void { fn releaseTaskResourcesLocked(t: *scheduler.Task) void {
recordExitLocked(t);
irq.releaseOwner(t.id); irq.releaseOwner(t.id);
devices_broker.releaseAllOwnedBy(t.id);
// The dying task's signal endpoint and one-shot timers go with it.
if (t.signal_endpoint) |raw| {
ipc.dropRef(@ptrCast(@alignCast(raw)));
t.signal_endpoint = null;
}
t.pending_signals = 0;
for (&one_shot_timers) |*slot| {
if (slot.*) |timer| {
if (timer.owner == t.id) {
ipc.dropRef(timer.endpoint);
slot.* = null;
}
}
}
// A dead subscriber's own subscriptions go first: it must not hear about
// itself, and the slots' endpoint references drop with it.
for (&exit_subscribers) |*slot| {
if (slot.*) |subscriber| {
if (subscriber.owner == t.id) {
ipc.dropRef(subscriber.endpoint);
slot.* = null;
}
}
}
if (t.ipc_client) |client| { if (t.ipc_client) |client| {
t.ipc_client = null; t.ipc_client = null;
client.ipc_status = -ipc.EPEER; client.ipc_status = -ipc.EPEER;
@@ -575,6 +612,12 @@ fn releaseTaskResourcesLocked(t: *scheduler.Task) void {
scheduler.removeFromWaitQueueLocked(t); scheduler.removeFromWaitQueueLocked(t);
scheduler.forgetIpcClientLocked(t); scheduler.forgetIpcClientLocked(t);
ipc.closeHandles(t); ipc.closeHandles(t);
// Publish the exit to every subscriber (docs/process-lifecycle.md): the same
// badge encoding as the supervisor's notification, and equally late, so a
// subscriber also observes a fully-released child.
for (&exit_subscribers) |*slot| {
if (slot.*) |subscriber| ipc.notifyLocked(subscriber.endpoint, abi.notify_exit_bit | t.id);
}
if (t.exit_endpoint) |raw| { if (t.exit_endpoint) |raw| {
const endpoint: *ipc.Endpoint = @ptrCast(@alignCast(raw)); const endpoint: *ipc.Endpoint = @ptrCast(@alignCast(raw));
t.exit_endpoint = null; t.exit_endpoint = null;
@@ -628,6 +671,7 @@ pub fn killProcess(caller_id: u32, target_id: u32) i64 {
const target = scheduler.taskByIdLocked(target_id) orelse return -ipc.ESRCH; const target = scheduler.taskByIdLocked(target_id) orelse return -ipc.ESRCH;
if (target.aspace == 0) return -ipc.ESRCH; // kernel tasks are not processes if (target.aspace == 0) return -ipc.ESRCH; // kernel tasks are not processes
if (target.supervisor != caller_id) return -ipc.EPERM; if (target.supervisor != caller_id) return -ipc.EPERM;
target.exit_reason = .killed;
if (target.state == .running) { if (target.state == .running) {
target.kill_pending = true; target.kill_pending = true;
} else { } else {
@@ -640,12 +684,170 @@ pub fn killProcess(caller_id: u32, target_id: u32) i64 {
/// The fault is confined to the process — the kernel trapped it on the task's own /// The fault is confined to the process — the kernel trapped it on the task's own
/// kernel stack and is intact — so everything the process held is reclaimed and the /// kernel stack and is intact — so everything the process held is reclaimed and the
/// core reschedules. The system keeps running; only the faulting process dies /// core reschedules. The system keeps running; only the faulting process dies
/// (docs/resilience.md: fault -> kill -> continue). /// (docs/resilience.md: fault -> kill -> continue). `reason` is the fault class
pub fn killCurrentProcess() noreturn { /// (from the vector), recorded for the supervisor's `process_exit_reason`.
pub fn killCurrentProcess(reason: abi.ExitReason) noreturn {
scheduler.current().exit_reason = reason;
fault_kill_count += 1; fault_kill_count += 1;
terminateCurrent(); terminateCurrent();
} }
/// The bounded record of recent deaths, for `process_exit_reason`: ids are never
/// reused, so a ring keyed by id is enough — a record evicted by wraparound reads
/// as -ESRCH, the same as an id that never lived, which a supervisor treats as
/// "too late to ask". Written under the big kernel lock by the reap.
const exit_record_capacity = 64;
const ExitRecord = struct { id: u32 = 0, supervisor: u32 = 0, reason: abi.ExitReason = .exited, valid: bool = false };
var exit_records: [exit_record_capacity]ExitRecord = .{ExitRecord{}} ** exit_record_capacity;
var exit_record_next: usize = 0;
/// Record a dying task's (id, supervisor, reason) — called by the reap before the
/// exit notification is posted, so a supervisor that hears the notification can
/// always still query the reason. Precondition: the big kernel lock is held.
fn recordExitLocked(t: *scheduler.Task) void {
exit_records[exit_record_next] = .{ .id = t.id, .supervisor = t.supervisor, .reason = t.exit_reason, .valid = true };
exit_record_next = (exit_record_next + 1) % exit_record_capacity;
}
/// How dead process `id` ended, for `caller` — the kernel half of the
/// process_exit_reason system call. Returns the ExitReason value, -ESRCH (never
/// lived, still alive, or evicted from the ring), or -EPERM (the caller was not
/// its supervisor — the same authority gate as process_kill).
pub fn exitReasonOf(caller_id: u32, target_id: u32) i64 {
const flags = sync.enter();
defer sync.leave(flags);
for (&exit_records) |*record| {
if (record.valid and record.id == target_id) {
if (record.supervisor != caller_id) return -ipc.EPERM;
return @intFromEnum(record.reason);
}
}
return -ipc.ESRCH;
}
/// The published exit events' subscribers (docs/process-lifecycle.md "Who learns
/// of a death"): stateful services — the VFS's file handles, input's
/// subscriptions — that must release what a dead client held and cannot learn it
/// any other way (a client that simply never calls again looks like silence).
/// Bounded like every kernel table; each entry holds its own endpoint reference.
const exit_subscriber_capacity = 8;
const ExitSubscriber = struct { endpoint: *ipc.Endpoint, owner: u32 };
var exit_subscribers: [exit_subscriber_capacity]?ExitSubscriber = .{null} ** exit_subscriber_capacity;
/// process_subscribe(endpoint): subscribe the caller's endpoint to published exit
/// events. Ungated, like process_enumerate — what is running (and dying) is not a
/// secret between cooperating processes. -ENOSPC when the table is full.
fn systemProcessSubscribe(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
const flags = sync.enter();
defer sync.leave(flags);
for (&exit_subscribers) |*slot| {
if (slot.* == null) {
endpoint.refcount += 1; // the slot's own reference, dropped on unsubscribe-by-death
slot.* = .{ .endpoint = endpoint, .owner = t.id };
return architecture.setSystemCallResult(state, 0);
}
}
failErr(state, ipc.ENOSPC);
}
/// signal_bind(endpoint): nominate where this process's signals arrive — the
/// IRQ-as-IPC pattern a fourth time (docs/process-lifecycle.md). Replacing a
/// binding drops the old reference; signals that pended while unbound are
/// delivered immediately on bind, coalesced into one notification.
fn systemSignalBind(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
const flags = sync.enter();
defer sync.leave(flags);
if (t.signal_endpoint) |raw| ipc.dropRef(@ptrCast(@alignCast(raw)));
endpoint.refcount += 1;
t.signal_endpoint = @ptrCast(endpoint);
if (t.pending_signals != 0) {
ipc.notifyLocked(endpoint, abi.notify_signal_bit | t.pending_signals);
t.pending_signals = 0;
}
architecture.setSystemCallResult(state, 0);
}
/// process_signal(id, signal): post a signal — a one-way, coalescing statement,
/// never a question (docs/process-lifecycle.md). The authority gate is the
/// supervision link, like kill; a process may also signal itself. Unbound
/// targets accumulate the signal in their pending mask.
fn systemProcessSignal(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const id = architecture.systemCallArg(state, 0);
const signal = architecture.systemCallArg(state, 1);
if (id > std.math.maxInt(u32)) return failErr(state, ipc.ESRCH);
if (signal > 31) return failErr(state, ipc.EBADF); // not a Signal bit position
const flags = sync.enter();
defer sync.leave(flags);
const target = scheduler.taskByIdLocked(@intCast(id)) orelse return failErr(state, ipc.ESRCH);
if (target.aspace == 0) return failErr(state, ipc.ESRCH);
if (target.supervisor != t.id and target.id != t.id) return failErr(state, ipc.EPERM);
target.pending_signals |= @as(u32, 1) << @intCast(signal);
if (target.signal_endpoint) |raw| {
const endpoint: *ipc.Endpoint = @ptrCast(@alignCast(raw));
ipc.notifyLocked(endpoint, abi.notify_signal_bit | target.pending_signals);
target.pending_signals = 0;
}
architecture.setSystemCallResult(state, 0);
}
/// The one-shot timers of timer_bind: the missing timed wait. A service arms a
/// deadline and keeps serving; the expiry arrives in the same replyWait as
/// everything else (notify_timer_bit). What stop-sequence escalation, hello
/// deadlines, and restart backoff are built from — and later, `alarm`.
const timer_capacity = 16;
const OneShotTimer = struct { deadline: u64, endpoint: *ipc.Endpoint, owner: u32 };
var one_shot_timers: [timer_capacity]?OneShotTimer = .{null} ** timer_capacity;
/// Sweep expired timers — hung on scheduler.timer_tick_hook, so it runs on every
/// tick with the big kernel lock held, like the sleeper wake it rides beside.
fn timerSweepLocked() void {
const now = architecture.millis();
for (&one_shot_timers) |*slot| {
if (slot.*) |timer| {
if (now >= timer.deadline) {
ipc.notifyLocked(timer.endpoint, abi.notify_timer_bit);
ipc.dropRef(timer.endpoint);
slot.* = null;
}
}
}
}
/// timer_bind(endpoint, ms): arm a one-shot timer. -ENOSPC when the table is full.
fn systemTimerBind(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
const ms = architecture.systemCallArg(state, 1);
const flags = sync.enter();
defer sync.leave(flags);
for (&one_shot_timers) |*slot| {
if (slot.* == null) {
endpoint.refcount += 1;
slot.* = .{ .deadline = architecture.millis() + ms, .endpoint = endpoint, .owner = t.id };
return architecture.setSystemCallResult(state, 0);
}
}
failErr(state, ipc.ENOSPC);
}
fn systemProcessExitReason(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const id = architecture.systemCallArg(state, 0);
if (id > std.math.maxInt(u32)) return failErr(state, ipc.ESRCH);
const r = exitReasonOf(t.id, @intCast(id));
architecture.setSystemCallResult(state, @bitCast(r));
}
/// Resolve `(device_id, resource_index)` to a GSI this process is entitled to bind, or null. /// Resolve `(device_id, resource_index)` to a GSI this process is entitled to bind, or null.
/// The two checks are the whole security story: the device must be *claimed* by the /// The two checks are the whole security story: the device must be *claimed* by the
/// caller, and the resource must be one of that device's `irq` resources as recorded /// caller, and the resource must be one of that device's `irq` resources as recorded
+17
View File
@@ -50,6 +50,17 @@ pub const Task = struct {
// null. Holds its own reference, dropped when the notification is posted. // null. Holds its own reference, dropped when the notification is posted.
// Opaque here for the same reason as `handles` below. // Opaque here for the same reason as `handles` below.
exit_endpoint: ?*anyopaque = null, exit_endpoint: ?*anyopaque = null,
// How this process ended — set by the death paths (exit, fault, kill) just
// before the reap records it for `process_exit_reason`. Meaningless while
// the task lives.
exit_reason: abi.ExitReason = .exited,
// Endpoint this process's signals arrive on (signal_bind), or null — same
// ownership rules as exit_endpoint (holds a reference; opaque here).
signal_endpoint: ?*anyopaque = null,
// Signals posted but not yet delivered: the coalescing pending mask
// (docs/process-lifecycle.md). Bits are abi.Signal values. Signals pend here
// until an endpoint is bound; two pending terminates are one terminate.
pending_signals: u32 = 0,
// Set by process_kill on a task that is running on another core; the kernel // Set by process_kill on a task that is running on another core; the kernel
// finishes the kill at that task's next system call or timer tick. // finishes the kill at that task's next system call or timer tick.
kill_pending: bool = false, kill_pending: bool = false,
@@ -643,9 +654,15 @@ fn reapKillPendingLocked() void {
/// other critical section, but releases it *without* touching the interrupt flag /// other critical section, but releases it *without* touching the interrupt flag
/// — the handler's `iretq` restores the interrupted context's flags, so /// — the handler's `iretq` restores the interrupted context's flags, so
/// re-enabling here would open a nested-interrupt window before the return. /// re-enabling here would open a nested-interrupt window before the return.
/// Called from the tick with the big kernel lock held — process.zig hangs the
/// one-shot timer sweep here (timer_bind), the same call-up pattern as the
/// teardown hooks below.
pub var timer_tick_hook: ?*const fn () void = null;
pub fn tick() void { pub fn tick() void {
_ = sync.enter(); _ = sync.enter();
wakeExpired(); wakeExpired();
if (timer_tick_hook) |hook| hook();
reapKillPendingLocked(); reapKillPendingLocked();
if (preemption_enabled) schedule(); if (preemption_enabled) schedule();
sync.leaveIsr(); sync.leaveIsr();
+215 -14
View File
@@ -132,6 +132,12 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
processKillTest(boot_information); processKillTest(boot_information);
} else if (eql(case, "supervision")) { } else if (eql(case, "supervision")) {
supervisionTest(boot_information); supervisionTest(boot_information);
} else if (eql(case, "claim-release")) {
claimReleaseTest(boot_information);
} else if (eql(case, "vfs-client-death")) {
vfsClientDeathTest(boot_information);
} else if (eql(case, "signals")) {
signalsTest(boot_information);
} else if (eql(case, "initial-ramdisk")) { } else if (eql(case, "initial-ramdisk")) {
initialRamdiskTest(boot_information); initialRamdiskTest(boot_information);
} else if (eql(case, "vfs")) { } else if (eql(case, "vfs")) {
@@ -1207,15 +1213,15 @@ fn userPfTest() void {
/// hand — address space, code page RO+X, stack page RW+NX — because the blob is a /// hand — address space, code page RO+X, stack page RW+NX — because the blob is a
/// raw code fragment, not an ELF `spawnProcess` could load. Returns false if any /// raw code fragment, not an ELF `spawnProcess` could load. Returns false if any
/// allocation fails. /// allocation fails.
fn spawnFaultingProcess() bool { fn spawnFaultingProcess() ?u32 {
const blob = process.pfBlob(); const blob = process.pfBlob();
const flags = sync.enter(); const flags = sync.enter();
defer sync.leave(flags); defer sync.leave(flags);
const aspace = architecture.createAddressSpace() orelse return false; const aspace = architecture.createAddressSpace() orelse return null;
const code_frame = pmm.alloc() orelse { const code_frame = pmm.alloc() orelse {
architecture.destroyAddressSpace(aspace); architecture.destroyAddressSpace(aspace);
return false; return null;
}; };
// Fill through the physmap (the user mapping is read-only); pad with int3 so a // Fill through the physmap (the user mapping is read-only); pad with int3 so a
// stray jump traps instead of sliding. // stray jump traps instead of sliding.
@@ -1226,15 +1232,16 @@ fn spawnFaultingProcess() bool {
const stack_frame = pmm.alloc() orelse { const stack_frame = pmm.alloc() orelse {
architecture.destroyAddressSpace(aspace); // frees code_frame too — it's mapped architecture.destroyAddressSpace(aspace); // frees code_frame too — it's mapped
return false; return null;
}; };
architecture.mapUserPageInto(aspace, process.stack_base_virtual, stack_frame, true, false); // RW + NX architecture.mapUserPageInto(aspace, process.stack_base_virtual, stack_frame, true, false); // RW + NX
if (scheduler.spawnUserLocked(aspace, process.code_virtual, process.stack_base_virtual + abi.page_size, 4, "fault-probe", 0, null) == null) { // Supervised by the calling test task, so exitReasonOf can read the verdict.
const id = scheduler.spawnUserLocked(aspace, process.code_virtual, process.stack_base_virtual + abi.page_size, 4, "fault-probe", scheduler.currentId(), null) orelse {
architecture.destroyAddressSpace(aspace); architecture.destroyAddressSpace(aspace);
return false; return null;
} };
return true; return id;
} }
/// Fault recovery (docs/resilience.md step 2): a scheduled ring-3 process that /// Fault recovery (docs/resilience.md step 2): a scheduled ring-3 process that
@@ -1263,7 +1270,8 @@ fn faultRecoveryTest(boot_information: *const BootInformation) void {
scheduler.setPriority(4); scheduler.setPriority(4);
check("init heartbeat before the fault", process.write_count >= 1); check("init heartbeat before the fault", process.write_count >= 1);
check("faulting process spawned", spawnFaultingProcess()); const probe = spawnFaultingProcess() orelse 0;
check("faulting process spawned", probe != 0);
// The kill: the faulting process #PFs on its first instruction and the kernel // The kill: the faulting process #PFs on its first instruction and the kernel
// reaps it instead of halting. // reaps it instead of halting.
@@ -1272,6 +1280,7 @@ fn faultRecoveryTest(boot_information: *const BootInformation) void {
while (process.fault_kill_count < 1 and architecture.millis() < deadline) scheduler.yield(); while (process.fault_kill_count < 1 and architecture.millis() < deadline) scheduler.yield();
scheduler.setPriority(4); scheduler.setPriority(4);
check("faulting process was killed (not the machine)", process.fault_kill_count == 1); check("faulting process was killed (not the machine)", process.fault_kill_count == 1);
check("the probe's reason reads segmentation_fault", process.exitReasonOf(scheduler.currentId(), probe) == @intFromEnum(abi.ExitReason.segmentation_fault));
// Life after the kill: init must keep beating on the same core. // Life after the kill: init must keep beating on the same core.
const beats_at_kill = process.write_count; const beats_at_kill = process.write_count;
@@ -1423,6 +1432,12 @@ fn processKillTest(boot_information: *const BootInformation) void {
check("the sleeper's exit notification arrived (length 0)", r == 0); check("the sleeper's exit notification arrived (length 0)", r == 0);
check("its badge carries the exit bit and the child id", badge == abi.notify_badge_bit | abi.notify_exit_bit | sleeper); check("its badge carries the exit bit and the child id", badge == abi.notify_badge_bit | abi.notify_exit_bit | sleeper);
// M17.2: the recorded reason — the notification is the fence, so it is
// already readable, and gated by the same supervisor check as the kill.
check("the sleeper's reason reads killed", process.exitReasonOf(me, sleeper) == @intFromEnum(abi.ExitReason.killed));
check("a non-supervisor may not read the reason (-EPERM)", process.exitReasonOf(me + 12345, sleeper) == -ipcsync.EPERM);
check("an unknown id has no reason (-ESRCH)", process.exitReasonOf(me, 0xFFFF_FF00) == -ipcsync.ESRCH);
const beats_at_kill = process.write_count; const beats_at_kill = process.write_count;
scheduler.sleep(1500); // more than one heartbeat period scheduler.sleep(1500); // more than one heartbeat period
check("the heartbeat stopped with the kill", process.write_count == beats_at_kill); check("the heartbeat stopped with the kill", process.write_count == beats_at_kill);
@@ -1443,6 +1458,21 @@ fn processKillTest(boot_information: *const BootInformation) void {
check("the spinner's exit notification arrived (length 0)", r == 0); check("the spinner's exit notification arrived (length 0)", r == 0);
check("its badge carries the exit bit and the child id", badge == abi.notify_badge_bit | abi.notify_exit_bit | spinner); check("its badge carries the exit bit and the child id", badge == abi.notify_badge_bit | abi.notify_exit_bit | spinner);
// M17.2: a child that ends on its own must read exited, not killed —
// args-echo with arguments echoes once and returns from main.
var clean: u32 = 0;
i = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "args-echo")) continue;
clean = process.spawnProcessSupervised(item.blob, 4, &.{ "args-echo", "clean-exit" }, me, endpoint) catch 0;
break;
}
check("args-echo spawned as the clean-exit child", clean != 0);
r = ipcsync.replyWait(endpoint, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap);
check("the clean child's exit notification arrived", badge == abi.notify_badge_bit | abi.notify_exit_bit | clean);
check("the clean child's reason reads exited", process.exitReasonOf(me, clean) == @intFromEnum(abi.ExitReason.exited));
var table: [32]abi.ProcessDescriptor = undefined; var table: [32]abi.ProcessDescriptor = undefined;
const total = scheduler.enumerate(&table); const total = scheduler.enumerate(&table);
var still_listed = false; var still_listed = false;
@@ -1454,6 +1484,176 @@ fn processKillTest(boot_information: *const BootInformation) void {
result(); result();
} }
/// M17.1: a dead process's device claims are released by the reap, so a restarted
/// driver can claim its hardware again (docs/process-lifecycle.md iron rule 1).
/// First the broker release in isolation — two owners, one released, the other's
/// claim must survive. Then the death-path wiring with a real child: the claim is
/// made on the child's behalf (the broker is kernel-callable), the child is
/// killed, and once the exit notification arrives — posted last, after release —
/// the device must be unclaimed and claimable again.
fn claimReleaseTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: claim-release\n", .{});
var buffer: [2]device_abi.DeviceDescriptor = undefined;
const total = devices_broker.enumerate(&buffer);
check("the device tree is seeded (>= 2 devices)", total >= 2);
if (total < 2) {
result();
return;
}
// The broker release in isolation.
check("device 0 claimed by owner 111", devices_broker.claim(0, 111));
check("device 1 claimed by owner 222", devices_broker.claim(1, 222));
devices_broker.releaseAllOwnedBy(111);
check("owner 111's claim is released", devices_broker.ownerOf(0) == null);
check("owner 222's claim survives", (devices_broker.ownerOf(1) orelse 0) == 222);
devices_broker.releaseAllOwnedBy(222);
check("cleanup released owner 222", devices_broker.ownerOf(1) == null);
// The death-path wiring: a real process dies holding a claim.
check("bootloader handed over /system/services/init", boot_information.init_len != 0);
if (boot_information.init_len == 0) {
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len];
const me = scheduler.currentId();
const endpoint = ipcsync.createIpcEndpoint() orelse {
check("exit endpoint allocated", false);
result();
return;
};
const child = process.spawnProcessSupervised(image, 4, &.{"/system/services/init"}, me, endpoint) catch 0;
check("supervised child spawned", child != 0);
check("device 0 claimed on the child's behalf", devices_broker.claim(0, child));
check("the kill is accepted", process.killProcess(me, child) == 0);
var badge: u64 = 0;
var received_cap: u64 = 0;
_ = ipcsync.replyWait(endpoint, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap);
check("the exit notification arrived", badge == abi.notify_badge_bit | abi.notify_exit_bit | child);
check("death released the child's claim", devices_broker.ownerOf(0) == null);
check("the device is claimable again", devices_broker.claim(0, me));
devices_broker.releaseAllOwnedBy(me);
result();
}
/// M17.3: the published exit events, proven by their first subscriber. The VFS
/// subscribes at startup; a client opens a file and parks holding the handle;
/// the kill posts the exit event to the VFS's endpoint; the VFS releases the
/// dead client's handle and says so — the service-side mirror of iron rule 1
/// (a service must never depend on clients cleaning up after themselves).
fn vfsClientDeathTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: vfs-client-death\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.write_count = 0;
check("vfs spawned", spawnNamed(rd, "vfs"));
const me = scheduler.currentId();
const endpoint = ipcsync.createIpcEndpoint() orelse {
check("exit endpoint allocated", false);
result();
return;
};
var client: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "vfs-test")) continue;
client = process.spawnProcessSupervised(item.blob, 4, &.{ "vfs-test", "park" }, me, endpoint) catch 0;
break;
}
check("parked client spawned (supervised)", client != 0);
// Its heartbeat is the fence: once it beats, the handle is open.
const parked = "vfstest: parked";
scheduler.setPriority(1);
var deadline = architecture.millis() + 10000;
while (architecture.millis() < deadline) {
if (process.write_len >= parked.len and eql(process.write_buffer[0..parked.len], parked)) break;
scheduler.yield();
}
scheduler.setPriority(4);
check("client parked holding an open handle", process.write_len >= parked.len and eql(process.write_buffer[0..parked.len], parked));
check("the kill is accepted", process.killProcess(me, client) == 0);
var badge: u64 = 0;
var received_cap: u64 = 0;
_ = ipcsync.replyWait(endpoint, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap);
check("the exit notification arrived", badge == abi.notify_badge_bit | abi.notify_exit_bit | client);
// The VFS heard the same published event; its release line is the proof.
const released = "vfs: released 1 handle(s) for dead client";
scheduler.setPriority(1);
deadline = architecture.millis() + 10000;
while (architecture.millis() < deadline) {
if (process.write_len >= released.len and eql(process.write_buffer[0..released.len], released)) break;
scheduler.yield();
}
scheduler.setPriority(4);
check("the VFS released the dead client's handle", process.write_len >= released.len and eql(process.write_buffer[0..released.len], released));
result();
}
/// M17.4 from ring 3: process-test's signal-run role drives the whole lifecycle
/// surface — the zero-length ping (answered by the harness), signals as
/// statements (reload logged, terminate = clean exit), the one-shot timer, and
/// both endings of the stop sequence (polite -> exited, deaf -> killed at the
/// deadline). Its "process-test: signals ok" is the pass marker.
fn signalsTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: signals\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(image); // the parent system_spawns its children by name
process.write_count = 0;
var runner: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "process-test")) continue;
runner = process.spawnProcessSupervised(item.blob, 4, &.{ "process-test", "signal-run" }, scheduler.currentId(), null) catch 0;
break;
}
check("signal-run parent spawned", runner != 0);
const pass_marker = "process-test: signals ok";
const fail_marker = "process-test: FAIL";
scheduler.setPriority(1);
const deadline = architecture.millis() + 15000;
var saw_pass = false;
var saw_fail = false;
while (architecture.millis() < deadline and !saw_pass and !saw_fail) {
if (process.write_len >= pass_marker.len and eql(process.write_buffer[0..pass_marker.len], pass_marker)) saw_pass = true;
if (process.write_len >= fail_marker.len and eql(process.write_buffer[0..fail_marker.len], fail_marker)) saw_fail = true;
scheduler.yield();
}
scheduler.setPriority(4);
check("the signal-run parent reported ok", saw_pass and !saw_fail);
result();
}
/// The whole user-side surface at once: spawn process-test's supervisor role, /// The whole user-side surface at once: spawn process-test's supervisor role,
/// which — entirely from ring 3 — creates an exit endpoint, spawns its two /// which — entirely from ring 3 — creates an exit endpoint, spawns its two
/// children supervised, sees them in process_enumerate, kills them (one blocked, /// children supervised, sees them in process_enumerate, kills them (one blocked,
@@ -2023,7 +2223,10 @@ fn faultNull() void {
// access). `allowzero` skips the same null check on the cast. The write then // access). `allowzero` skips the same null check on the cast. The write then
// hits the unmapped page 0 and takes a real hardware #PF. // hits the unmapped page 0 and takes a real hardware #PF.
var address: u64 = 0; var address: u64 = 0;
address = asm ("" : [ret] "=r" (-> u64) : [in] "0" (address)); address = asm (""
: [ret] "=r" (-> u64),
: [in] "0" (address),
);
const p: *allowzero volatile u64 = @ptrFromInt(address); const p: *allowzero volatile u64 = @ptrFromInt(address);
p.* = 1; p.* = 1;
} }
@@ -2049,8 +2252,7 @@ fn faultDoubleFault() void {
\\ud2 \\ud2
: :
: [sp] "r" (bad_sp), : [sp] "r" (bad_sp),
: .{ .memory = true } : .{ .memory = true });
);
bad_sp += 0; bad_sp += 0;
} }
@@ -2069,8 +2271,7 @@ fn apDoubleFaultTask() void {
\\ud2 \\ud2
: :
: [sp] "r" (bad_sp), : [sp] "r" (bad_sp),
: .{ .memory = true } : .{ .memory = true });
);
bad_sp += 0; bad_sp += 0;
} }
@@ -42,6 +42,35 @@ fn driverFor(d: device.DeviceDescriptor) ?[]const u8 {
}; };
} }
/// The PCI class/subclass/prog-IF triple of an xHCI (USB 3) host controller:
/// Serial Bus Controller (0x0C) / USB Controller (0x03) / XHCI (0x30) — the names
/// pci-class.zig decodes.
const xhci_pci_class: u64 = 0x0C_03_30;
/// The bus driver that serves a PCI function, or null. Unlike the singleton drivers
/// in `driverFor`, a machine can carry several identical controllers — so the caller
/// spawns one driver instance *per device*, passing the device id as argv[1] for the
/// instance to claim.
fn pciDriverFor(d: device.DeviceDescriptor) ?[]const u8 {
if (d.class != @intFromEnum(device.DeviceClass.pci_device)) return null;
return switch (d.pci_class) {
xhci_pci_class => "usb-xhci-bus",
else => null,
};
}
/// Spawn one instance of `driver_name` to serve the specific device `id` — the id
/// arrives as argv[1]. No isProcessRunning gate here: the name alone cannot tell two
/// instances apart, and this manager is the sole spawner of drivers.
fn spawnForDevice(driver_name: []const u8, id: u64) void {
var text: [20]u8 = undefined;
const id_text = std.fmt.bufPrint(&text, "{d}", .{id}) catch return;
if (system.spawnWithArguments(driver_name, &.{id_text}) != null) {
writeLine("device-manager: spawned {s} for device {d}\n", .{ driver_name, id });
} else {
writeLine("device-manager: failed to spawn {s} for device {d}\n", .{ driver_name, id });
}
}
pub fn main() void { pub fn main() void {
// Enumerate into a heap buffer (too big for the one-page user stack). // Enumerate into a heap buffer (too big for the one-page user stack).
@@ -54,6 +83,11 @@ pub fn main() void {
var matched: usize = 0; var matched: usize = 0;
for (buffer[0..count]) |descriptor| { for (buffer[0..count]) |descriptor| {
if (pciDriverFor(descriptor)) |driver_name| {
matched += 1;
spawnForDevice(driver_name, descriptor.id);
continue;
}
const driver_name = driverFor(descriptor) orelse continue; const driver_name = driverFor(descriptor) orelse continue;
matched += 1; matched += 1;
if (!system.isProcessRunning(driver_name)) { if (!system.isProcessRunning(driver_name)) {
@@ -48,11 +48,91 @@ fn awaitChildExit(endpoint: runtime.ipc.Handle) u32 {
return received.childProcessId(); return received.childProcessId();
} }
/// The harness-run child of the signals test: echoes requests, logs the two
/// signals it handles. Terminate makes run() return, and returning from main is
/// the clean exit the parent reads as ExitReason.exited.
fn echo(message: []const u8, reply: []u8, sender: u32) usize {
_ = sender;
const n = @min(message.len, reply.len);
@memcpy(reply[0..n], message[0..n]);
return n;
}
fn onReload() void {
_ = runtime.system.write("process-test: reloaded\n");
}
fn onTerminate() void {
_ = runtime.system.write("process-test: terminating\n");
}
/// The parent of the signals test: drives ping, echo, reload, the one-shot
/// timer, and both endings of the stop sequence (polite -> exited; deaf ->
/// killed at the deadline). Prints "process-test: signals ok" as the marker.
fn signalRun() void {
const endpoint = runtime.ipc.createIpcEndpoint() orelse fail("create exit endpoint");
const child = runtime.system.spawnSupervised("process-test", &.{"service"}, endpoint) orelse fail("spawn service child");
// Reach the child's endpoint through the registry (retry: it may not be up).
var service_handle: ?runtime.ipc.Handle = null;
var tries: u32 = 0;
while (service_handle == null and tries < 200) : (tries += 1) {
service_handle = runtime.ipc.lookup(.input);
if (service_handle == null) runtime.system.sleep(20);
}
const h = service_handle orelse fail("service child never registered");
// The universal ping: a zero-length call answered zero-length by the harness.
var reply: [16]u8 = undefined;
const pong = runtime.ipc.call(h, &.{}, &reply) catch fail("ping call failed");
if (pong != 0) fail("ping reply not empty");
// An ordinary request still reaches on_message.
const n = runtime.ipc.call(h, "echo!", &reply) catch fail("echo call failed");
if (n != 5 or !std.mem.eql(u8, reply[0..5], "echo!")) fail("echo mismatch");
// reload: a statement — the child logs it; the kernel test reads the serial.
if (!runtime.process.sendSignal(child, .reload)) fail("send reload");
runtime.system.sleep(200);
// The one-shot timer: armed on our endpoint, lands as isTimer.
if (!runtime.system.timerOnce(endpoint, 100)) fail("arm timer");
var scratch: [8]u8 = undefined;
const landing = runtime.ipc.replyWait(endpoint, scratch[0..0], &scratch, null);
if (!landing.isTimer()) fail("expected the timer landing");
// The stop sequence, polite path: terminate, clean exit inside the deadline.
runtime.process.stop(child, 2000, endpoint);
if ((runtime.process.exitReason(child) orelse .killed) != .exited) fail("service child reason not exited");
// The deaf child: binds nothing, hears nothing — the deadline kills it.
const deaf = runtime.system.spawnSupervised("process-test", &.{"sleeper"}, endpoint) orelse fail("spawn deaf child");
runtime.system.sleep(50); // let it reach its sleep
runtime.process.stop(deaf, 300, endpoint);
if ((runtime.process.exitReason(deaf) orelse .exited) != .killed) fail("deaf child reason not killed");
_ = runtime.system.write("process-test: signals ok\n");
}
pub fn main(init: runtime.process.Init) void { pub fn main(init: runtime.process.Init) void {
const role = init.arguments.get(1) orelse return; // spawned bare (ramdisk sweep): stay silent const role = init.arguments.get(1) orelse return; // spawned bare (ramdisk sweep): stay silent
if (std.mem.eql(u8, role, "sleeper")) { if (std.mem.eql(u8, role, "sleeper")) {
while (true) runtime.system.sleep(500); while (true) runtime.system.sleep(500);
} }
if (std.mem.eql(u8, role, "service")) {
// Borrowed well-known id: the input service is not part of this scenario.
runtime.service.run(64, .{
.service = .input,
.on_message = echo,
.on_reload = onReload,
.on_terminate = onTerminate,
});
return; // terminate arrived; returning is the clean exit
}
if (std.mem.eql(u8, role, "signal-run")) {
signalRun();
return;
}
if (std.mem.eql(u8, role, "spinner")) { if (std.mem.eql(u8, role, "spinner")) {
var beat: u64 = 0; var beat: u64 = 0;
const touch: *volatile u64 = &beat; const touch: *volatile u64 = &beat;
@@ -89,6 +169,12 @@ pub fn main(init: runtime.process.Init) void {
if (listed(sleeper, "process-test")) fail("sleeper still listed after kill"); if (listed(sleeper, "process-test")) fail("sleeper still listed after kill");
if (listed(spinner, "process-test")) fail("spinner still listed after kill"); if (listed(spinner, "process-test")) fail("spinner still listed after kill");
// M17.2: both children were killed by us, and the reason says so — the whole
// restart-policy input, read through the runtime like a real supervisor would.
if ((runtime.process.exitReason(sleeper) orelse .exited) != .killed) fail("sleeper reason not killed");
if ((runtime.process.exitReason(spinner) orelse .exited) != .killed) fail("spinner reason not killed");
if (runtime.process.exitReason(0xFFFF_FFF0) != null) fail("unknown id had a reason");
_ = runtime.system.write("process-test: ok\n"); _ = runtime.system.write("process-test: ok\n");
} }
+2 -2
View File
@@ -8,8 +8,8 @@
//! spellings (`stat`, `O_CREAT`, ...) live only in the POSIX layer //! spellings (`stat`, `O_CREAT`, ...) live only in the POSIX layer
//! (library/posix/unistd.zig), which translates to these. //! (library/posix/unistd.zig), which translates to these.
//! //!
//! This is user-space only — the kernel knows nothing of files or paths; it only //! This is user-space only — the kernel knows nothing of files or paths; it only moves the bytes.
//! moves the bytes. Shared by library/posix/unistd.zig (client) and system/services/vfs/vfs.zig (server). //! Shared by library/posix/unistd.zig (client) and system/services/vfs/vfs.zig (server).
pub const Operation = enum(u32) { pub const Operation = enum(u32) {
open, // open(path) -> node id open, // open(path) -> node id
+21 -1
View File
@@ -6,10 +6,30 @@
const std = @import("std"); const std = @import("std");
const runtime = @import("runtime"); const runtime = @import("runtime");
pub fn main() void { pub fn main(init: runtime.process.Init) void {
const u = @import("posix").unistd; const u = @import("posix").unistd;
const payload = "hello-vfs"; const payload = "hello-vfs";
// The "park" role (the vfs-client-death test): open a file, then hold the
// handle forever without closing — the kill and the VFS's release-on-death
// are the point.
if (init.arguments.count > 1) {
var fd: i32 = -1;
var tries: u32 = 0;
while (fd < 0 and tries < 200) : (tries += 1) {
fd = u.open("parked", u.O_CREAT);
if (fd < 0) runtime.system.sleep(20);
}
if (fd < 0) {
_ = runtime.system.write("vfstest: park open failed\n");
return;
}
while (true) {
_ = runtime.system.write("vfstest: parked\n");
runtime.system.sleep(500);
}
}
// The VFS server may not have registered yet — retry open until it's up. // The VFS server may not have registered yet — retry open until it's up.
var fd: i32 = -1; var fd: i32 = -1;
var tries: u32 = 0; var tries: u32 = 0;
+52 -18
View File
@@ -23,6 +23,10 @@ const Node = struct {
const OpenFile = struct { const OpenFile = struct {
used: bool = false, used: bool = false,
node: usize = 0, node: usize = 0,
// The client (task id — an IPC badge is one) that opened this handle. What
// release-on-death sweeps by: a service must never depend on its clients
// cleaning up after themselves (docs/process-lifecycle.md).
owner: u32 = 0,
}; };
var nodes = [_]Node{.{}} ** 8; var nodes = [_]Node{.{}} ** 8;
@@ -65,8 +69,29 @@ fn fail(out: []u8) usize {
return writeReply(out, .{ .status = -1 }, &.{}); return writeReply(out, .{ .status = -1 }, &.{});
} }
/// Handle one request; write the reply into `out`, return its length. /// Format one whole log line and emit it in a single `debug_write`, so lines from
fn handle(message: []const u8, out: []u8) usize { /// concurrent processes can never land in the middle of it.
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [96]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
/// Release every open handle `client` held — called on that client's published
/// exit event. The nodes (the files) stay: ramfs contents outlive their writers,
/// only the dead client's handles go.
fn releaseClientHandles(client: u32) void {
var released: u32 = 0;
for (&opens) |*o| {
if (o.used and o.owner == client) {
o.used = false;
released += 1;
}
}
if (released != 0) writeLine("vfs: released {d} handle(s) for dead client {d}\n", .{ released, client });
}
/// Handle one request from `sender`; write the reply into `out`, return its length.
fn handle(message: []const u8, out: []u8, sender: u32) usize {
if (message.len < protocol.request_size) return fail(out); if (message.len < protocol.request_size) return fail(out);
const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]); const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]);
const payload = message[protocol.request_size..]; const payload = message[protocol.request_size..];
@@ -77,7 +102,7 @@ fn handle(message: []const u8, out: []u8) usize {
const ni = findNode(name) orelse createNode(name) orelse return fail(out); const ni = findNode(name) orelse createNode(name) orelse return fail(out);
for (&opens, 0..) |*o, i| { for (&opens, 0..) |*o, i| {
if (!o.used) { if (!o.used) {
o.* = .{ .used = true, .node = ni }; o.* = .{ .used = true, .node = ni, .owner = sender };
return writeReply(out, .{ .status = 0, .node = i }, &.{}); return writeReply(out, .{ .status = 0, .node = i }, &.{});
} }
} }
@@ -113,27 +138,36 @@ fn handle(message: []const u8, out: []u8) usize {
} }
} }
pub fn main() void { /// Startup, under the harness: subscribe to the published exit events — when a
const endpoint = runtime.ipc.createIpcEndpoint() orelse { /// client dies holding open handles, the exit notification is how the VFS learns
_ = runtime.system.write("vfs: no endpoint\n"); /// to release them (docs/process-lifecycle.md).
return; fn initialise(endpoint: runtime.ipc.Handle) bool {
}; if (!runtime.process.subscribeExits(endpoint)) {
if (!runtime.ipc.register(.vfs, endpoint)) { _ = runtime.system.write("vfs: exit subscription failed\n");
_ = runtime.system.write("vfs: register failed\n");
return;
} }
_ = runtime.system.write("vfs: ready\n"); _ = runtime.system.write("vfs: ready\n");
return true;
}
var reply_buffer: [protocol.message_maximum]u8 = undefined; /// A non-signal notification: the only kind the VFS subscribes to is exit events.
var reply_len: usize = 0; fn onNotification(badge: u64) void {
var receive: [protocol.message_maximum]u8 = undefined; if (badge & runtime.ipc.notify_exit_bit != 0) {
while (true) { releaseClientHandles(@intCast(badge & ~(runtime.ipc.notify_badge_bit | runtime.ipc.notify_exit_bit)));
const got = runtime.ipc.replyWait(endpoint, reply_buffer[0..reply_len], &receive, null);
// Ignore notifications (none expected here); handle a request.
reply_len = handle(receive[0..got.len], &reply_buffer);
} }
} }
pub fn main() void {
// The harness owns the loop: requests dispatch to handle(), exit events to
// onNotification(), ping and terminate are answered for free — this service
// gained the whole lifecycle contract by deleting its hand-rolled loop.
runtime.service.run(protocol.message_maximum, .{
.service = .vfs,
.init = initialise,
.on_message = handle,
.on_notification = onNotification,
});
}
pub const panic = runtime.panic; pub const panic = runtime.panic;
comptime { comptime {
_ = &runtime.start._start; _ = &runtime.start._start;
+18
View File
@@ -240,6 +240,24 @@ CASES = [
"smp": 4, "smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS", "expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# M17.1: a dead process's device claims are released by the reap — kill a child
# holding a claim, the device must be claimable again (process-lifecycle.md).
{"name": "claim-release",
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M17.3: published exit events — the VFS subscribes, a client dies holding an
# open handle, and the VFS releases it (process-lifecycle.md "Who learns of a death").
{"name": "vfs-client-death",
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M17.4: signals over IPC — ping, reload, terminate (clean exit), the one-shot
# timer, and the stop sequence's two endings, all driven from ring 3.
{"name": "signals",
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# The initial_ramdisk: the loader ferries a bundle of user binaries; the kernel parses # The initial_ramdisk: the loader ferries a bundle of user binaries; the kernel parses
# it and spawns each as a ring-3 process (here the VFS-server stub heartbeats). # it and spawns each as a ring-3 process (here the VFS-server stub heartbeats).
{"name": "initial-ramdisk", {"name": "initial-ramdisk",
+152
View File
@@ -0,0 +1,152 @@
#!/usr/bin/env python3
"""Rewrap over-long Zig comment lines at a column limit (default 100).
Rules:
- Only comment-only lines are touched; trailing comments after code are left alone.
- Consecutive comment lines with the same indentation and marker (`//`, `///`, `//!`)
form a block. Blank comment lines separate paragraphs within a block.
- A line whose text starts with `- ` begins a bullet paragraph; its continuation
lines are the ones indented to align under the bullet's text (bullet lead + 2).
- Plain paragraphs join consecutive lines with the same text indentation;
wrapped lines align where the first line's text begins.
- A paragraph is rewrapped only if at least one of its lines exceeds the limit,
so deliberate short line breaks elsewhere are preserved.
"""
import argparse
import difflib
import re
import subprocess
import sys
import textwrap
LIMIT = 100
COMMENT_RE = re.compile(r"^(\s*)(//[/!]?)(?:\s(.*))?$")
BULLET_RE = re.compile(r"^(\s*)- (.*)$")
def split_paragraphs(texts):
"""texts: list of comment text (None for a bare marker line).
Returns paragraphs: dicts with lead/hang/bullet/texts/lines(indices)."""
paragraphs = []
current = None
for index, text in enumerate(texts):
if text is None or text.strip() == "":
paragraphs.append({"literal": True, "lines": [index]})
current = None
continue
lead = len(text) - len(text.lstrip(" "))
bullet = BULLET_RE.match(text)
if bullet:
current = {
"lead": len(bullet.group(1)),
"hang": len(bullet.group(1)) + 2,
"bullet": True,
"texts": [bullet.group(2)],
"lines": [index],
}
paragraphs.append(current)
elif current is not None and lead == current["hang"]:
current["texts"].append(text)
current["lines"].append(index)
else:
current = {
"lead": lead,
"hang": lead,
"bullet": False,
"texts": [text],
"lines": [index],
}
paragraphs.append(current)
return paragraphs
def wrap_paragraph(paragraph, prefix):
joined = re.sub(r"\s+", " ", " ".join(t.strip() for t in paragraph["texts"]))
if paragraph["bullet"]:
initial = prefix + " " * paragraph["lead"] + "- "
else:
initial = prefix + " " * paragraph["lead"]
subsequent = prefix + " " * paragraph["hang"]
return textwrap.wrap(
joined,
width=LIMIT,
initial_indent=initial,
subsequent_indent=subsequent,
break_long_words=False,
break_on_hyphens=False,
)
def rewrap_block(original_lines, indent, marker, texts):
prefix = indent + marker + " "
output = []
for paragraph in split_paragraphs(texts):
block_originals = [original_lines[i] for i in paragraph["lines"]]
if paragraph.get("literal") or all(len(l) <= LIMIT for l in block_originals):
output.extend(block_originals)
else:
output.extend(wrap_paragraph(paragraph, prefix))
return output
def process(source):
lines = source.split("\n")
result = []
i = 0
while i < len(lines):
match = COMMENT_RE.match(lines[i])
if not match:
result.append(lines[i])
i += 1
continue
indent, marker = match.group(1), match.group(2)
block_lines, texts = [], []
while i < len(lines):
m = COMMENT_RE.match(lines[i])
if not m or m.group(1) != indent or m.group(2) != marker:
break
block_lines.append(lines[i])
texts.append(m.group(3))
i += 1
result.extend(rewrap_block(block_lines, indent, marker, texts))
return "\n".join(result)
def main():
parser = argparse.ArgumentParser()
parser.add_argument("--write", action="store_true", help="apply changes (default: diff only)")
parser.add_argument("files", nargs="*", help="files to process (default: git ls-files '*.zig')")
args = parser.parse_args()
files = args.files or subprocess.run(
["git", "ls-files", "*.zig"], capture_output=True, text=True, check=True
).stdout.split()
changed = 0
for path in files:
with open(path, encoding="utf-8") as f:
source = f.read()
rewrapped = process(source)
if rewrapped == source:
continue
changed += 1
if args.write:
with open(path, "w", encoding="utf-8") as f:
f.write(rewrapped)
print(f"rewrapped {path}")
else:
sys.stdout.writelines(
difflib.unified_diff(
source.splitlines(keepends=True),
rewrapped.splitlines(keepends=True),
fromfile=path,
tofile=path,
)
)
if not changed:
print("no comments over the limit")
if __name__ == "__main__":
main()