Author SHA1 Message Date
Daniel Samson a53c2b0193 Placeholder discovery services and the -Ddiscovery build option
system/services/acpi and system/services/fdt exist as documented
placeholders (silent clean-exit mains; the headers say exactly what each
becomes and why). The build's -Ddiscovery=acpi|fdt option fills the
ramdisk's neutral 'discovery' slot — the device manager will spawn
"discovery" by that name in M20.3 and never learn which firmware it is
on (m19-m20-plan.md decision 7). x86 defaults to acpi; the aarch64
target flips the default when it lands.
2026-07-13 01:46:04 +01:00
Daniel Samson bf481c080c Record the firmware-neutrality contract as decision 7
Discovery is one swappable process per firmware (acpi service on x86, an
fdt service on the Pis); everything at and above the device-manager
protocol stays generic. The manager owns the tree as data and never
touches hardware — firmware bytecode runs in a crashable, supervised
discoverer. Flagged now: hid[8] cannot hold an FDT compatible string,
and cross-firmware protocols are named by domain (power, not ACPI).
2026-07-13 01:33:12 +01:00
Daniel Samson 3a78dcab3f Scope ACPI events and system power as M21; record the SCI on acpi-tables
Battery, AC, lid, and the power button ride the acpi service as reported
children with small class drivers — the xHCI split repeated. QEMU can
only prove the power-button path (system_powerdown injects the real fixed
event), so battery/EC are interface-complete and hardware-validated on
the laptop. Per-device power states (D-states, suspend/resume) stay out
of scope: suspend has the shape of a lifecycle signal every driver must
answer, and it has no consumer until laptop sleep.
2026-07-13 01:21:23 +01:00
Daniel Samson 470f93a83d Plan the discovery migration (M19 pci-bus, M20 acpi service) 2026-07-13 01:15:17 +01:00
Daniel Samson 7798706b41 Mark the M17-M18 plan complete 2026-07-13 00:49:04 +01:00
Daniel Samson ad40de03c2 Merge feat/usb-xhci-bus: xHCI port scan, tree reports, and the app surface (M18.2-M18.3) 2026-07-13 00:49:04 +01:00
Daniel Samson d8778b4b70 The application surface: enumerate, subscribe, and device-list (M18.3)
Applications ask the device manager for the tree (enumerate: a header
plus ChildEntry records) and subscribe to published add/remove events by
handing their endpoint over as the call's capability — the input-service
pattern; events are the same ChildAdded/ChildRemoved structs the bus
drivers send, one encoding in both directions. device-list is the first
client: it prints the tree, subscribes, and narrates the events through
a driver restart. The protocol's message maximum is capped at the
kernel's IPC MESSAGE_MAXIMUM (256 bytes, ten entries per reply; paging
joins the protocol when a tree outgrows one message). The startUserTask
debug print is gone: it wrote to serial unserialized against user-space
lines and sheared concurrent log markers in half — the root cause of the
scenario flakes.
2026-07-13 00:49:03 +01:00
Daniel Samson 79d859a111 The xHCI driver scans its root-hub ports and reports the tree (M18.2)
child_added/child_removed join the device-manager protocol. The driver
maps its register BAR (resource 0 is the ECAM config space; the walk
starts at 1), reads CAPLENGTH and HCSPARAMS1, and reads one PORTSC per
port: the connect bit and speed class come straight from hardware, no
rings needed to see the devices. The manager mirrors reported children
keyed by (parent, port), remembers which instance reported each, and
prunes a dead reporter's children before deciding the restart — the
children describe protocol state that died with the process. The
usb-report scenario drives the whole loop: two QEMU devices reported,
reporter killed, children pruned, driver respawned with backoff, and the
new instance re-claims, re-scans, and re-reports.
2026-07-13 00:28:29 +01:00
Daniel Samson 37fb09f75e Mark the feat/device-manager merge done in the M17-M18 plan 2026-07-13 00:19:32 +01:00
Daniel Samson 34ebeb968d Merge feat/device-manager: the supervising device manager (M18.1) 2026-07-13 00:19:32 +01:00
Daniel Samson 3cc1d38dd0 The device manager supervises: hello, backoff, and the crash-loop cap (M18.1)
The manager is now a harness service on the well-known .device_manager
endpoint. Every driver spawns supervised; drivers with an assignment must
hello (device-manager-protocol, versioned) within a deadline enforced by
a timer sweep. Exit reasons drive the restart decision: clean exits stay
down, faults restart with 300/600/1200ms backoff, and three fast deaths
mark a driver failed instead of respawning forever. usb-xhci-bus is the
first conforming driver; the crash-test fixture claims a device, hellos,
and faults on purpose — each respawn re-proving claim release on death
through the manager's own path. maximum_tasks grows 16 -> 32: the
initial-ramdisk sweep (15 binaries at once) was intermittently
overflowing the static pool.
2026-07-13 00:19:30 +01:00
Daniel Samson 36e804b848 Mark the feat/process-lifecycle merge done in the M17-M18 plan 2026-07-12 23:53:50 +01:00
Daniel Samson be83a42d42 Merge feat/process-lifecycle: the process lifecycle (M17.1-M17.4)
Claim release on death, exit reasons, published exit events with the VFS
as first subscriber, signals over IPC with one-shot timers and the
service harness — docs/process-lifecycle.md increments 1-4, all built.
2026-07-12 23:53:50 +01:00
Daniel Samson 650a1b1595 Signals over IPC, one-shot timers, and the service harness (M17.4)
Signals are statements delivered as coalescing notifications to the
endpoint a process nominates with signal_bind — never a hijacked stack,
never a question (liveness is the zero-length ping the harness answers).
process_signal is supervisor-or-self gated, like kill; unbound targets
accumulate a pending mask delivered on bind. timer_bind is the missing
timed wait: a one-shot deadline landing in the same replyWait as
everything else — what stop(), hello deadlines, and restart backoff are
built from. runtime.service.run folds requests, signals, and
notifications into callbacks; the VFS conversion deletes its hand-rolled
loop and gains the whole lifecycle contract. The signals scenario drives
ping, reload, terminate->exited, the timer, and the deaf-child
deadline->killed path from ring 3. docs/process-lifecycle.md increments
1-4 are now as-built.
2026-07-12 23:53:38 +01:00
Daniel Samson d8c55c6f2f Publish exit events to subscribers; the VFS releases dead clients' handles (M17.3)
process_subscribe adds an endpoint to a bounded, ref-counted subscriber
table; every death posts the same badge encoding a supervisor's exit
notification uses, equally late, so subscribers observe a fully-released
child. A dying subscriber's own subscriptions are removed first — it never
hears about itself. The VFS is the first subscriber: open handles now
record their owner and are swept when the owner dies, because a service
must never depend on clients cleaning up after themselves
(docs/process-lifecycle.md). Proven by the vfs-client-death scenario.
2026-07-12 23:41:44 +01:00
Daniel Samson 2ebfb0c3b0 Record and expose how every process ends (M17.2)
The kernel records an ExitReason at all three death sites — clean exit,
fault (classified by vector), and process_kill — into a bounded ring
before the exit notification posts, so a supervisor's query never races
the notice. process_exit_reason is gated by the same supervisor check as
kill; runtime.process.exitReason is the stable interface. This is the
input restart policy reads (docs/process-lifecycle.md iron rule 2).
2026-07-12 23:34:09 +01:00
Daniel Samson 888eaa74e1 Release a dead process's device claims (M17.1)
Every path out of a process (exit, fault, kill) now releases its device
claims alongside its IRQ and MSI bindings, so a restarted driver can claim
its hardware again — the cleanup half of process-lifecycle.md's iron rule 1.
MSI vectors were already swept by irq.releaseOwner; claims were the gap.
The claim-release test proves kill -> release -> re-claim, plus the broker
release in isolation.
2026-07-12 23:23:49 +01:00
Daniel Samson ed76cbbc79 Mark Phase 0 done: baseline QEMU suite green (48/48) 2026-07-12 23:17:20 +01:00
Daniel Samson 140229b88d Rename usb-xhci-libary.zig to usb-xhci-library.zig (naming typo) 2026-07-12 23:13:33 +01:00
Daniel Samson cb2379fd06 Update README.md 2026-07-12 23:12:39 +01:00
Daniel Samson 1665b239b0 Add the status checklist and workflow to the M17-M18 plan 2026-07-12 23:04:43 +01:00
Daniel Samson 70ed0337f8 Merge feat/usb: USB wire ABI, xHCI detection and spawn, M17-M18 design 2026-07-12 22:56:35 +01:00
32 changed files with 2296 additions and 141 deletions
+26 -7
View File
@@ -2,12 +2,31 @@
Codename: Shodan Codename: Shodan
Version: 1 Version: 1
A small operating system, written from scratch in Zig — a bootloader (`boot/`) A small resilient operating system, written from scratch in Zig.
and a microkernel (`system/kernel/`), sharing a neutral handoff contract (`system/boot-handoff.zig`).
It boots x86-64 via UEFI, and so far has a framebuffer console, a physical frame ## Zen of DanOS:
allocator, its own paging with W^X permissions, interrupt/exception handling, a
LAPIC timer, a kernel heap, a fixed-priority preemptive scheduler, and in-kernel IPC - Resilient Micro-Kernel Architecture.
channels. See [`docs/`](docs/README.md) for how each piece works. - Every process run in an isolated user space not kernel space.
- Processes cannot take down the entire OS with it when they die or is killed
- Stable public runtime library, private OS ABI.
- Keeps a stable runtime for user space processes between OS versions (great for backwards compatibility)
- Allows the underlying OS to be changed without effecting applications
- Provides a boundary to enable compatibility between OS's e.g. POSIX, MUSL etc
- Drivers are just isolated processes in user space.
- Thin binaries that can be restarted like applications.
- Useful during driver development.
- Drivers can claim MMIO / ports
- Driver resources (e.g. IRQ/Port/MMIO) claims are automatically cleaned up if the driver dies or is killed
- Drivers can also hook into the process lifecyle to clean up or reset hardware
- No legacy to deal with
- Zig code uses a clean coding style (Zen of Zig)
- Favor reading code over writing code.
- No magic numbers.
- No shortend names unless its for ABI compatibility or acronyms
- Inter-Process Communication (IPC)
- Publish and subscribe to Asynchronous Messages
- Talk to services and processes synchronously
## Prerequisites ## Prerequisites
@@ -60,7 +79,7 @@ straight into CI.
## Documentation ## Documentation
Design notes explaining the *why* behind the code live in Design notes explaining *why* behind the code live in
[`docs/`](docs/README.md) — start with [`docs/README.md`](docs/README.md). [`docs/`](docs/README.md) — start with [`docs/README.md`](docs/README.md).
## Logo ## Logo
+31
View File
@@ -212,6 +212,13 @@ pub fn build(b: *std.Build) void {
}, },
}); });
// The device-manager protocol: hello + (M18.2) tree reports, exposed as its
// own module like the other protocol modules. Imported through the runtime.
const device_manager_protocol_module = b.addModule("device-manager-protocol", .{
.root_source_file = b.path("system/services/device-manager/device-manager-protocol.zig"),
});
runtime_module.addImport("device-manager-protocol", device_manager_protocol_module);
// Typed volatile MMIO register access + memory-ordering barriers, for drivers on // Typed volatile MMIO register access + memory-ordering barriers, for drivers on
// top of an mmio_map grant. Depends only on `builtin` (arch-conditional barriers); // top of an mmio_map grant. Depends only on `builtin` (arch-conditional barriers);
// no target set, so it inherits each driver's. See library/mmio/mmio.zig. // no target set, so it inherits each driver's. See library/mmio/mmio.zig.
@@ -332,6 +339,24 @@ pub fn build(b: *std.Build) void {
const ps2_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-keyboard", "system/drivers/ps2-bus/keyboard.zig"); const ps2_keyboard_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-keyboard", "system/drivers/ps2-bus/keyboard.zig");
const ps2_mouse_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-mouse", "system/drivers/ps2-bus/mouse.zig"); const ps2_mouse_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "ps2-mouse", "system/drivers/ps2-bus/mouse.zig");
const usb_xhci_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-xhci-bus", "system/drivers/usb-xhci-bus/usb-xhci-bus.zig"); const usb_xhci_bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "usb-xhci-bus", "system/drivers/usb-xhci-bus/usb-xhci-bus.zig");
// A test fixture, not a real driver: hellos to the device manager, then faults —
// what the driver-restart scenario drives the crash-loop cap with.
const crash_test_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "crash-test", "system/services/crash-test/crash-test.zig");
const device_list_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-list", "system/services/device-list/device-list.zig");
// The discovery service: one swappable process per firmware
// (docs/m19-m20-plan.md decision 7), bundled under the neutral ramdisk name
// "discovery" so the device manager never learns which firmware it is on.
// x86 boots describe hardware with ACPI; the Raspberry Pis hand over a
// flattened device tree — the aarch64 target flips the default when it
// lands (docs/arm.md). Both are placeholders until M20.1 (acpi) and the
// ARM bring-up (fdt).
const Discovery = enum { acpi, fdt };
const discovery = b.option(Discovery, "discovery", "Which discovery service fills the ramdisk's 'discovery' slot (default: acpi)") orelse Discovery.acpi;
const discovery_source: []const u8 = switch (discovery) {
.acpi => "system/services/acpi/acpi.zig",
.fdt => "system/services/fdt/fdt.zig",
};
const discovery_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "discovery", discovery_source);
const device_manager_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-manager", "system/services/device-manager/device-manager.zig"); const device_manager_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "device-manager", "system/services/device-manager/device-manager.zig");
// The input service and its exercisers: the fan-out server, a hardware-free synthetic // The input service and its exercisers: the fan-out server, a hardware-free synthetic
// source, and a subscriber that doubles as the `input` test's oracle. See docs/input.md. // source, and a subscriber that doubles as the `input` test's oracle. See docs/input.md.
@@ -363,6 +388,12 @@ pub fn build(b: *std.Build) void {
mk_run.addFileArg(ps2_mouse_exe.getEmittedBin()); mk_run.addFileArg(ps2_mouse_exe.getEmittedBin());
mk_run.addArg("usb-xhci-bus"); mk_run.addArg("usb-xhci-bus");
mk_run.addFileArg(usb_xhci_bus_exe.getEmittedBin()); mk_run.addFileArg(usb_xhci_bus_exe.getEmittedBin());
mk_run.addArg("crash-test");
mk_run.addFileArg(crash_test_exe.getEmittedBin());
mk_run.addArg("device-list");
mk_run.addFileArg(device_list_exe.getEmittedBin());
mk_run.addArg("discovery");
mk_run.addFileArg(discovery_exe.getEmittedBin());
mk_run.addArg("device-manager"); mk_run.addArg("device-manager");
mk_run.addFileArg(device_manager_exe.getEmittedBin()); mk_run.addFileArg(device_manager_exe.getEmittedBin());
mk_run.addArg("input"); mk_run.addArg("input");
+11 -1
View File
@@ -1,6 +1,16 @@
# The device manager # The device manager
**Status: design.** The primitives this builds on are real ([process-management.md](process-management.md): **Status: the protocol and supervision are built** (M18.1, 2026-07-13): `hello`
with its deadline, supervised spawn, restart with backoff, and the crash-loop
cap are in — usb-xhci-bus is the first conforming driver, and the
`driver-restart` scenario proves fault → backoff → re-claim → cap end to end.
Tree reports are built too (M18.2, 2026-07-13): the xHCI driver scans its
root-hub ports and reports each connected device (`child_added`); the manager
mirrors them and prunes a dead reporter's children, and the `usb-report`
scenario proves report → prune → respawn → re-report. The application surface is built (M18.3, 2026-07-13):
`enumerate` and `subscribe` over IPC, with `device-list` as the first client —
the manager is now the one answer to "what devices exist" for applications.
The primitives underneath are real ([process-management.md](process-management.md):
spawn/supervise/kill/exit-notification; [driver-model.md](driver-model.md): the device spawn/supervise/kill/exit-notification; [driver-model.md](driver-model.md): the device
table as a capability system; [drivers.md](drivers.md): claim/map/IRQ), and the first table as a capability system; [drivers.md](drivers.md): claim/map/IRQ), and the first
per-device driver spawn works (the device manager matches the xHCI controller by PCI per-device driver spawn works (the device manager matches the xHCI controller by PCI
+19
View File
@@ -103,3 +103,22 @@ This is what makes a user-space driver possible at all, and it's the subject of
the oldest (discrete messages, not a coalescing level like the notification ring). the oldest (discrete messages, not a coalescing level like the notification ring).
- **A bounded reply.** `MSG_MAX` is 256 bytes and the copy runs under the big kernel - **A bounded reply.** `MSG_MAX` is 256 bytes and the copy runs under the big kernel
lock; a bulk transfer wants shared pages, not a copy. lock; a bulk transfer wants shared pages, not a copy.
## Lifecycle conventions over IPC (M17)
Three conventions from [process-lifecycle.md](process-lifecycle.md) ride the
notification mechanism:
- **Signals** arrive as notifications on the endpoint a process nominated with
`signal_bind` (`runtime.process.bindSignals`): badge = the signal bit plus the
coalesced pending mask (`runtime.process.signalsFrom` decodes). Statements,
never questions; no payload, no reply.
- **One-shot timers** (`timer_bind`, `runtime.system.timerOnce`) land as a
timer-bit notification — the timed wait: a service arms a deadline and keeps
serving, instead of blocking in sleep.
- **The universal ping**: a **zero-length request is the liveness probe**,
answered with a zero-length reply by the service harness itself
(`runtime.service.run`). No protocol's requests start at length zero, so the
encoding cannot collide, and a wedged service simply fails to answer — which
is the diagnosis. Deep health ("can I reach my hardware?") stays a per-service
protocol message.
+45
View File
@@ -21,6 +21,51 @@ run the existing QEMU suite green before any new work starts.
**Numbering note:** continues the milestone sequence (driver track ended at M16). **Numbering note:** continues the milestone sequence (driver track ended at M16).
## Status
The loop marks a phase `[x]` in the same commit that lands it. A phase is marked
only when its definition of green holds.
- [x] **Phase 0** — baseline: docs committed, feat/usb merged to main, pushed;
`usb-xhci-libary.zig` renamed to `usb-xhci-library.zig`; existing QEMU
suite green from the worktree (48/48, 2026-07-12).
- [x] **M17.1** — kernel releases claims/MSI on death (claims: `releaseAllOwnedBy`
in the reap; MSI was already swept by `irq.releaseOwner`; `claim-release`
test; suite 49/49)
- [x] **M17.2** — exit reasons (`ExitReason` recorded at exit/fault/kill before
the notification; `process_exit_reason` supervisor-gated;
`runtime.process.exitReason`; kernel + ring-3 assertions; suite 49/49)
- [x] **M17.3** — published exit events + VFS subscriber (`process_subscribe`,
bounded ref-counted table, publish on every death;
`runtime.process.subscribeExits`; VFS handles carry owners and are swept on
the owner's death; `vfs-client-death` test; suite 50/50)
- [x] **M17.4** — signals, timer notifications, `runtime.process`, the service
harness (signal_bind/process_signal + coalescing pending mask; timer_bind
on the tick; bindSignals/signalsFrom/sendSignal/stop + timerOnce;
runtime.service.run with the zero-length ping; VFS converted; `signals`
scenario; suite 51/51)
- [x] **merge** `feat/process-lifecycle` → main, push (merged 2026-07-13)
- [x] **M18.1** — device-manager protocol: hello + restart policy
(device-manager-protocol module; the manager as a harness service:
supervised spawns, hello deadline via timer sweep, restart with
300/600/1200ms backoff, exit reasons deciding restart-vs-stopped,
crash-loop cap; usb-xhci-bus first conforming driver; crash-test fixture
re-proving claim release each respawn; `driver-restart` scenario;
maximum_tasks 16→32 — the sweep was overflowing the pool; suite 52/52)
- [x] **merge** `feat/device-manager` → main, push (merged 2026-07-13)
- [x] **M18.2** — xHCI port scan + tree reports (child_added/child_removed in
the protocol; the manager's child mirror with death-pruning; xHCI maps the
register BAR — resource 0 is ECAM — reads CAPLENGTH/HCSPARAMS1, scans
PORTSC, reports connected ports with speed-class identity; `usb-report`
scenario proves report → prune → respawn → re-report; suite 53/53)
- [x] **M18.3** — app surface: enumerate/subscribe over IPC (subscriber
endpoint rides as the call's capability; events are the same structs the
buses send); device-list first client; protocol capped at the kernel's
IPC MESSAGE_MAXIMUM (256); the startUserTask debug print removed — it
sheared concurrent serial lines and was the scenario-flake root cause;
`device-list` scenario; suite 54/54)
- [x] **merge** `feat/usb-xhci-bus` → main, push (merged 2026-07-13) — **plan complete**
--- ---
## M17.1 — the kernel releases a dead process's claims ## M17.1 — the kernel releases a dead process's claims
+187
View File
@@ -0,0 +1,187 @@
# M19–M20 execution plan: discovery migration
The operational plan for [device-manager.md](device-manager.md)'s increment 8:
discovery leaves the kernel — a **pci-bus driver** (M19) and an **acpi service**
(M20), with the kernel's device enumeration retired behind them. Same rules as
[m17-m18-plan.md](m17-m18-plan.md): one phase at a time, each green before the
next; this file is the build order and the checklist.
**Definition of green, every phase:** `zig build` clean, `zig build test` clean,
`python3 test/qemu_test.py` passes (existing scenarios plus the phase's new
one), and the relevant design doc updated. Commit per green phase (no co-author
trailers). The full suite is the regression net — the existing
`driver-restart` / `usb-report` / `device-list` / `input` scenarios must stay
green *through* the migration, which is the whole point: the system must not be
able to tell who enumerated it.
**Workflow:** dedicated worktree; branches off `main` — `feat/pci-bus`
(M19.0–19.3), `feat/acpi-service` (M20.1–20.3); auto-merge to main when a
branch is green; keep branches; push everything.
## Settled decisions (2026-07-13 — veto before the loop starts)
1. **What "retiring the kernel scan" means.** The kernel keeps, forever, the
parses it needs before user space exists: RSDP/XSDT location, MADT (SMP),
the HPET table (the tick), FADT + the AML `\_S5` evaluation (poweroff — the
power tests prove it), and MCFG (the host bridge node). What retires is
**device enumeration**: the ECAM function walk (M19.3) and the DSDT/SSDT
namespace walk that builds device nodes (M20.3). The AML module stays a
shared build module compiled into both the kernel (for `\_S5`) and the acpi
service (for everything else) — same source, two builds, no fork.
2. **Bridge apertures come from the firmware memory map, not AML.** Registered
PCI functions carry BAR resources, and containment demands the bridge own
windows that cover them. The apertures are derived kernel-side from the
boot memory map's MMIO holes (regions that are neither RAM nor tables) —
mechanical, AML-free, and available at boot regardless of what later moved
to user space. (The bridge today carries only ECAM + bus range; this is the
prerequisite M19.0 exists for.)
3. **`device_register` becomes idempotent on exact match.** A re-registration
with identical (parent, class, resources) returns the existing id instead
of appending. The kernel table has no unregister, so without this a
restarted registering bus would duplicate its children on every respawn —
idempotence makes restart-and-re-report safe for every future bus, not just
PCI.
4. **The manager matches from reports.** `ChildAdded` gains a `device_id`
field (the kernel-registered id, `no_device` for unregistered leaves like
USB ports). After the M19.3 flip, PCI driver matching keys off reported
identity (the class triple) instead of the manager's boot-time snapshot —
the snapshot match remains only for what the kernel still seeds. One flip
phase changes both sides at once so no device is ever matched twice.
5. **The acpi service's authority is one node.** The kernel publishes an
`acpi-tables` device: memory resources covering the table blobs plus a
broad `io_port` resource — the documented trust grant to exactly one
process (AML OperationRegions reach EC/PM ports; the claim-gated
io_read/io_write calls already exist). The service claims it, maps the
tables, and runs the shared AML module in ring 3 behind a `Hal` backed by
`mmio_map` + `io_read`/`io_write`.
6. **Both new processes are protocol drivers** under the manager: hello,
supervision, restart with backoff — all inherited from M18.1 for free.
Registration idempotence (decision 3) is what makes their restarts sound.
7. **Firmware neutrality is the contract** (2026-07-13). The generic layer is
everything at and above the device-manager protocol — descriptors,
containment, reports, matching, supervision — and none of it may become
x86-specific. Discovery is one swappable process per firmware: the acpi
service on x86; an **fdt service** on the Raspberry Pis (claims a
`devicetree-blob` node, reports children from the flattened device tree —
pure data, no bytecode, no port grant, strictly simpler than ACPI). The
manager owns the tree as *data* and touches no hardware, ever — AML runs in
a crashable, supervised discoverer precisely so a firmware-bytecode fault
can never take down the supervisor. Two consequences recorded now:
`DeviceDescriptor`'s 8-byte `hid` cannot hold an FDT `compatible` string
("brcm,bcm2835-aux-uart") — identity widens before the fdt service exists;
and cross-firmware surfaces are named by **domain, not firmware** (M21
defines a *power* protocol, not an "ACPI events" protocol — PSCI/mailbox
sources feed the same subscribers on ARM). **Landed early (2026-07-13):**
both services exist as placeholders (system/services/acpi, system/services/
fdt) and the build's `-Ddiscovery=acpi|fdt` option fills the ramdisk's
neutral `discovery` slot — the manager will spawn "discovery" by that name
in M20.3 and never learn which firmware it is on.
## Status
- [ ] **M19.0** — prerequisites on `feat/pci-bus`: bridge MMIO apertures from
the memory-map holes; `device_register` idempotence (+ kernel unit
checks); `ChildAdded.device_id`; archive note on m17-m18-plan.md.
- [ ] **M19.1** — pci-bus driver, scan only: claim the host bridge, map the
ECAM window, walk bus/device/function headers, log what it finds.
Scenario `pci-scan`: the kernel test compares the driver's reported count
against the broker table's `pci_device` count — equivalence, per class.
- [ ] **M19.2** — register + report: each function registered under the bridge
(config-space slice + BARs, `pci_class` in the descriptor), reported with
`child_added { device_id, identity = class triple }`. Manager mirrors;
spawn-from-reports stays **off**. Scenario extends `pci-scan`:
registered ids resolve, no duplicates after a forced driver restart
(idempotence proven end to end).
- [ ] **M19.3** — the flip: kernel `enumeratePci` call removed (bridge node
stays); manager matches PCI drivers from reports. One commit. The
existing xHCI scenarios (`driver-restart`, `usb-report`, `device-list`)
are the assertion — xhci must come up spawned off a pci-bus report, and
the suite must not be able to tell the difference. discovery.md updated.
- [ ] **merge** `feat/pci-bus` → main, push.
- [ ] **M20.1** — acpi service, parse only (fills the existing placeholder at
system/services/acpi/acpi.zig): kernel publishes `acpi-tables`
(decision 5) — memory over the table blobs, the broad io_port grant, and
**the SCI as an irq resource** (from the FADT; unused until M21 but free
to record now). The service claims it, maps the blobs, runs the shared
AML module in ring 3, logs the namespace device count and `_HID`s.
Scenario `acpi-parse`: user-space count equals the kernel walk's count.
- [ ] **M20.2** — register + report: namespace devices with `_HID` + `_CRS`
resources registered under `acpi-tables` (its io_port + the memory-map
holes give containment), reported to the manager. Spawn-from-reports for
ACPI matches stays off. Scenario: the reported set includes the PS/2
keyboard and mouse nodes with their IRQ resources.
- [ ] **M20.3** — the flip: kernel DSDT device-node building removed (static
tables + `\_S5` stay, decision 1); manager matches ACPI-hid drivers
(ps2-bus) from reports. The `input` and `device-manager` scenarios are
the assertion. discovery.md + acpi.md + device-manager.md updated;
device-manager.md increment 8 closed.
- [ ] **merge** `feat/acpi-service` → main, push — **loop ends here**.
---
## Phase notes
**M19.0 apertures:** the boot memory map already crosses the handoff
([boot-handoff]); the holes computation belongs where the bridge node is built
(`parseMcfg`). Sanity-check on QEMU q35: the xHCI BAR (`0xc0000000`-region
values seen in the M18 logs) must land inside a derived aperture, asserted in
the kernel unit test.
**M19.1 scanning without owning config access twice:** the driver reads config
space through its ECAM mmio_map grant of the *bridge* window — the same bytes
the kernel walk read. Vendor-id `0xFFFF` skip, header-type multifunction rule,
no bridge recursion (matches the kernel's current single-segment walk).
**M19.2 BAR sizing:** the classic size probe (write all-ones, read mask,
restore) is deferred — the BARs' current programmed values and types are
enough for containment-checked registration at bring-up; sizing lands with the
first driver that needs to *move* a BAR. Log what is registered so the
scenario can assert it.
**M19.3 what the manager still seeds from the snapshot:** everything the
kernel still enumerates (timers, ACPI nodes until M20.3). The PCI arm of
`pciDriverFor` switches source; `driverFor` doesn't move until M20.3.
**M20.1 Hal in ring 3:** `mapMmio` → `device.mmioMap` over the claimed
acpi-tables node (plus a table-offset map for blobs); `pioRead`/`pioWrite` →
`device.ioRead`/`ioWrite` against its io_port resource. The interpreter cannot
tell it moved — that is the assertion of `acpi-parse`.
**M20.2 containment for `_CRS`:** io ports fall inside the node's broad
io_port resource; MMIO windows (HPET, LAPIC ranges some firmwares list) fall
inside the memory-map holes added to the node in M20.1. Anything that doesn't
fit is logged and skipped, loudly — bring-up honesty over silent drops.
**M20.3 ps2 ordering:** ps2-bus binds nodes the acpi service now reports, so
its spawn moves behind the report (the manager's matching handles this once
the source flips); the `input` scenario proves the keyboard still types.
**Explicitly out of scope:** PCI bridge recursion (single segment, flat bus
walk stays); BAR reprogramming/sizing; disk/PCIe hotplug; interrupt routing
changes (`_PRT` stays wherever it is today); the USB descriptor track;
multi-segment ECAM; per-device power states (D-states, `_PSx`/`_PRx`,
suspend/resume — a future *lifecycle-vocabulary* extension, since "suspend"
has the shape of a signal every driver must answer, and it has no consumer
until laptop sleep); CPU P/C-states.
## M21 preview — ACPI events + system power (planned next, not in this loop)
The acpi service grows the event side (settled direction 2026-07-13; detailed
phases when M20 lands):
- **21.1 SCI + fixed events**: irq_bind the SCI (the resource M20.1 already
records), read/clear PM1 status, publish the power-button event to
subscribers (the same pub/sub shape the manager uses).
- **21.2 GPE + Notify**: Notify dispatch in the shared AML interpreter, GPE
block handling, `Notify(device, code)` published per reported node. The
acpi service is a **bus** here: battery (PNP0C0A), AC (ACPI0003), and lid
(PNP0C0D) nodes are reported children; small class drivers bind them and
speak an evaluate/subscribe protocol to the service — the xHCI split,
repeated. The embedded controller (`_Qxx` queries) rides this phase;
QEMU emulates no battery/EC, so those paths are interface-complete and
validated on real hardware (the laptop is the win condition), while the
plumbing is proven by the power button.
- **21.3 the capstone**: QEMU `system_powerdown` → acpi service event → init
runs the M17 stop sequence over its children → kernel `\_S5` — orderly
shutdown as the scenario that proves lifecycle + events compose. (The
harness grows a QMP poke to inject the event.)
+10 -4
View File
@@ -1,6 +1,9 @@
# Process lifecycle: signals over IPC # Process lifecycle: signals over IPC
**Status: design.** The primitives underneath are built ([process-management.md](process-management.md): **Status: increments 1–4 built** (2026-07-12): claim release on death, exit
reasons, published exit events, and signals + one-shot timers + the service
harness are all in — the interface below is as-built. The primitives underneath
predate this design ([process-management.md](process-management.md):
spawn, the supervision link, kill, child-exit notifications); this document designs spawn, the supervision link, kill, child-exit notifications); this document designs
the layer above them — the standard vocabulary a danos process speaks about its own the layer above them — the standard vocabulary a danos process speaks about its own
life, and the stable `runtime.process` interface that carries it. Nothing here is life, and the stable `runtime.process` interface that carries it. Nothing here is
@@ -245,13 +248,16 @@ pub fn stop(id: u32, deadline_ms: u64) void { ... }
/// process_enumerate. /// process_enumerate.
pub fn subscribeExits(endpoint: usize) bool { ... } pub fn subscribeExits(endpoint: usize) bool { ... }
/// How a process ended — from the exit notification. What restart policy reads. /// How a process ended — queried after the exit notification (the kernel records
pub const ExitReason = enum { /// it first, so the two never race). What restart policy reads. (Built in M17.2.)
pub const ExitReason = enum(u8) {
exited, // returned from main / clean exit exited, // returned from main / clean exit
aborted, // abort() — deliberate self-termination (SIGABRT's ghost) aborted, // abort() — deliberate self-termination (SIGABRT's ghost; reserved)
segmentation_fault, // SIGSEGV's ghost segmentation_fault, // SIGSEGV's ghost
illegal_instruction, // SIGILL's ghost illegal_instruction, // SIGILL's ghost
arithmetic_fault, // SIGFPE's ghost arithmetic_fault, // SIGFPE's ghost
protection_fault, // general protection fault
fault, // any other CPU exception
killed, // process_kill killed, // process_kill
}; };
``` ```
+13 -5
View File
@@ -95,11 +95,18 @@ the architecture layer calls up into `tick`.
## Known gaps (bring-up honesty) ## Known gaps (bring-up honesty)
- Device **claims** are not released on death (pre-existing: the fault path has - ~~Device claims are not released on death~~ Closed (M17.1): every path out of a
the same gap) — a killed driver's device stays claimed until reboot. process releases its device claims alongside its IRQ and MSI bindings
(`releaseTaskResourcesLocked`), so a restarted driver can claim its hardware
again — the cleanup half of [process-lifecycle.md](process-lifecycle.md)'s iron
rule 1. The `claim-release` test proves the kill → release → re-claim cycle.
- Kernel stacks of dead tasks are leaked, as on every exit path (no reaper yet). - Kernel stacks of dead tasks are leaked, as on every exit path (no reaper yet).
- There is no exit *status* in the notification, only the id; a supervisor that - ~~There is no exit status in the notification~~ Closed (M17.2): the kernel
needs the code can grow a wait-style call later. records how every process ends — exited, a fault class, or killed — before it
posts the exit notification, and the supervisor reads it with
`process_exit_reason` (`runtime.process.exitReason`). This is the input to
restart policy ([process-lifecycle.md](process-lifecycle.md)); an exit *code*
for the clean case can still ride alongside later.
- Enumerate writes through the caller's raw pointer under the bring-up trust - Enumerate writes through the caller's raw pointer under the bring-up trust
model, like `device_enumerate` (an unmapped page is a self-DoS, not an model, like `device_enumerate` (an unmapped page is a self-DoS, not an
isolation break). isolation break).
@@ -109,4 +116,5 @@ the architecture layer calls up into `tick`.
`process-list` (enumerate), `process-kill` (kernel-level kill paths, refusals, `process-list` (enumerate), `process-kill` (kernel-level kill paths, refusals,
notifications), `supervision` (the whole user-side surface via the process-test notifications), `supervision` (the whole user-side surface via the process-test
service: spawn supervised → enumerate → kill blocked and spinning children → service: spawn supervised → enumerate → kill blocked and spinning children →
notifications → gone). See test/qemu_test.py. notifications → gone), `claim-release` (a killed claim-holder's device is
claimable again). See test/qemu_test.py.
+20
View File
@@ -94,6 +94,15 @@ pub fn send(h: Handle, message: []const u8) bool {
/// GSI. See `isNotification`. /// GSI. See `isNotification`.
pub const notify_badge_bit: u64 = abi.notify_badge_bit; pub const notify_badge_bit: u64 = abi.notify_badge_bit;
/// Set alongside `notify_badge_bit` when the notification is a **signal** — the
/// lifecycle vocabulary of docs/process-lifecycle.md, delivered to the endpoint
/// nominated with `process.bindSignals`. Decode with `process.signalsFrom`.
pub const notify_signal_bit: u64 = abi.notify_signal_bit;
/// Set alongside `notify_badge_bit` when the notification is a **one-shot timer**
/// landing (`system.timerOnce`).
pub const notify_timer_bit: u64 = abi.notify_timer_bit;
/// Set alongside `notify_badge_bit` when the notification is a **child-exit /// Set alongside `notify_badge_bit` when the notification is a **child-exit
/// notice** — a process this one spawned (with an exit endpoint) has ended — /// notice** — a process this one spawned (with an exit endpoint) has ended —
/// rather than a device interrupt. The low bits carry the child's process id. /// rather than a device interrupt. The low bits carry the child's process id.
@@ -133,6 +142,17 @@ pub const Received = struct {
} }
/// The task id of whoever posted a buffered message, meaningful only when /// The task id of whoever posted a buffered message, meaningful only when
/// Whether this arrival is a signal notification — decode the set with
/// `process.signalsFrom(badge)`.
pub fn isSignal(self: Received) bool {
return self.isNotification() and self.badge & notify_signal_bit != 0;
}
/// Whether this arrival is a one-shot timer landing (`system.timerOnce`).
pub fn isTimer(self: Received) bool {
return self.isNotification() and self.badge & notify_timer_bit != 0;
}
/// `isMessage`. (The badge's low bits, with the three high marker bits masked off.) /// `isMessage`. (The badge's low bits, with the three high marker bits masked off.)
pub fn senderTaskId(self: Received) u32 { pub fn senderTaskId(self: Received) u32 {
return @intCast(self.badge & ~(notify_badge_bit | notify_exit_bit | notify_message_bit)); return @intCast(self.badge & ~(notify_badge_bit | notify_exit_bit | notify_message_bit));
+91 -1
View File
@@ -1,8 +1,15 @@
//! Process-level runtime types: what a user program receives at entry. Mirrors //! Process-level runtime types: what a user program receives at entry (`Init`,
//! the argv contract) and the process end of the lifecycle
//! (docs/process-lifecycle.md) — today the exit reason a supervisor reads to
//! decide restart; signals and the stop sequence land here with M17.4. Mirrors
//! the spirit of `std.process.Init.Minimal` in danos terms — std's `Args` holds //! the spirit of `std.process.Init.Minimal` in danos terms — std's `Args` holds
//! no data on freestanding targets, so the type is danos's own. //! no data on freestanding targets, so the type is danos's own.
const std = @import("std"); const std = @import("std");
const abi = @import("abi");
const sc = @import("system-call.zig");
const ipc = @import("ipc.zig");
const system = @import("system.zig");
/// Everything a program receives at entry. Passed to /// Everything a program receives at entry. Passed to
/// `pub fn main(init: runtime.process.Init)`; programs that need nothing keep /// `pub fn main(init: runtime.process.Init)`; programs that need nothing keep
@@ -45,3 +52,86 @@ pub const Arguments = struct {
} }
}; };
}; };
/// How a process ended — what a supervisor's restart policy reads: a clean exit
/// meant to stop, a fault wants a restart with backoff, killed means the
/// supervisor did it itself (docs/process-lifecycle.md).
pub const ExitReason = abi.ExitReason;
/// How dead child `id` ended. Ask after the exit notification arrives — the
/// kernel records the reason before it posts the notification, so this never
/// races it. Returns null for an id that never lived, is still alive, was
/// evicted from the kernel's bounded record, or is not this process's child
/// (the same authority gate as `kill`).
pub fn exitReason(id: u32) ?ExitReason {
const r = sc.systemCall1(.process_exit_reason, id);
if (r > ~@as(usize, 0) - 4095) return null; // a wrapped -errno
return @enumFromInt(r);
}
/// The signal vocabulary (docs/process-lifecycle.md): POSIX's concepts, danos's
/// names, message delivery. A signal is a one-way coalescing statement — never a
/// question (liveness is the zero-length ping call) and never kill (that is
/// `system.kill`, unhandleable by definition).
pub const Signal = abi.Signal;
/// The coalesced set of signals one notification delivered: two pending
/// terminates arrive as one. Decode a received badge with `signalsFrom`.
pub const SignalSet = struct {
pending: u32,
pub fn has(set: SignalSet, signal: Signal) bool {
return set.pending & (@as(u32, 1) << @intFromEnum(signal)) != 0;
}
};
/// Nominate `endpoint` as this process's signal endpoint. Signals posted while
/// unbound have pended; they are delivered immediately on bind, coalesced.
pub fn bindSignals(endpoint: usize) bool {
return sc.systemCall1(.signal_bind, endpoint) == 0;
}
/// Decode a received badge into the signals it delivered, or null if it is not
/// a signal notification.
pub fn signalsFrom(badge: u64) ?SignalSet {
if (badge & abi.notify_badge_bit == 0 or badge & abi.notify_signal_bit == 0) return null;
return .{ .pending = @truncate(badge & ~(abi.notify_badge_bit | abi.notify_signal_bit)) };
}
/// Post `signal` to child `id` (or to yourself). Supervisor-gated, like kill;
/// non-blocking, always — a statement, not a conversation.
pub fn sendSignal(id: u32, signal: Signal) bool {
return sc.systemCall2(.process_signal, id, @intFromEnum(signal)) == 0;
}
/// The standard stop sequence (docs/process-lifecycle.md): terminate, wait up to
/// `deadline_ms` for the exit notification on `exit_endpoint` (the endpoint the
/// child was spawned with), then kill. Any *other* notifications arriving on
/// that endpoint while stopping are consumed and dropped — a supervisor with
/// concurrent traffic implements the same sequence inside its own event loop
/// (arm `system.timerOnce`, keep serving) instead of calling this.
pub fn stop(id: u32, deadline_ms: u64, exit_endpoint: usize) void {
_ = sendSignal(id, .terminate);
_ = system.timerOnce(exit_endpoint, deadline_ms);
var receive: [8]u8 = undefined;
while (true) {
const got = ipc.replyWait(exit_endpoint, &.{}, &receive, null);
if (got.isChildExit() and got.childProcessId() == id) return;
if (got.isTimer()) break; // the deadline passed first — escalate
}
_ = system.kill(id);
while (true) {
const got = ipc.replyWait(exit_endpoint, &.{}, &receive, null);
if (got.isChildExit() and got.childProcessId() == id) return;
}
}
/// Subscribe `endpoint` to published exit events: every process death posts an
/// asynchronous notification with the same badge encoding as a supervisor's exit
/// notice (decode with `ipc.Received.isChildExit`/`childProcessId`). For stateful
/// services: release what the dead client held — file handles, subscriptions —
/// because a service must never depend on clients cleaning up after themselves
/// (docs/process-lifecycle.md). Ungated, like `system.processes`.
pub fn subscribeExits(endpoint: usize) bool {
return sc.systemCall1(.process_subscribe, endpoint) == 0;
}
+7
View File
@@ -17,6 +17,9 @@ pub const ipc = @import("ipc.zig");
pub const start = @import("start.zig"); pub const start = @import("start.zig");
/// The VFS wire protocol (shared with the VFS server). /// The VFS wire protocol (shared with the VFS server).
pub const vfs_protocol = @import("vfs-protocol"); pub const vfs_protocol = @import("vfs-protocol");
/// The device-manager protocol: hello + tree reports (docs/device-manager.md).
pub const device_manager_protocol = @import("device-manager-protocol");
/// Keyboard-event listening (subscribe/next) and broadcasting (publish), over the input /// Keyboard-event listening (subscribe/next) and broadcasting (publish), over the input
/// service. See library/runtime/input.zig and system/services/input/. /// service. See library/runtime/input.zig and system/services/input/.
pub const input = @import("input.zig"); pub const input = @import("input.zig");
@@ -35,5 +38,9 @@ pub const panic = start.panic;
/// Process entry types: the `Init` handed to `main`, and its `Arguments`. /// Process entry types: the `Init` handed to `main`, and its `Arguments`.
pub const process = @import("process.zig"); pub const process = @import("process.zig");
/// The service harness: one replyWait loop folding requests, signals, and
/// notifications into callbacks (docs/process-lifecycle.md).
pub const service = @import("service.zig");
/// The heap as a `std.mem.Allocator`, for Zig `std` containers in user code. /// The heap as a `std.mem.Allocator`, for Zig `std` containers in user code.
pub const allocator = heap.allocator; pub const allocator = heap.allocator;
+83
View File
@@ -0,0 +1,83 @@
//! The service harness (docs/process-lifecycle.md): one replyWait loop that
//! folds protocol requests, signals, and subscribed notifications into
//! callbacks — so the lifecycle contract ("answers ping, exits on terminate")
//! is satisfied by construction and a service author writes domain logic only.
//! Nothing is asynchronous inside the process: a callback runs at a point the
//! loop chose, never on a hijacked stack — the whole reason signals are
//! messages.
//!
//! The liveness probe: a **zero-length request is the universal ping**, answered
//! with a zero-length reply by the harness itself. No protocol's requests start
//! at length zero, so the encoding cannot collide, and there is nothing for a
//! service author to implement — a wedged service simply fails to answer, which
//! is the diagnosis (see docs/ipc.md).
const abi = @import("abi");
const ipc = @import("ipc.zig");
const process = @import("process.zig");
pub const Callbacks = struct {
/// Called once with the service's endpoint before the loop starts — the
/// place to subscribe to exit events, bind IRQs, or announce readiness.
/// Return false to abort startup (the process exits).
init: ?*const fn (endpoint: ipc.Handle) bool = null,
/// One protocol request from `sender` (a task id): write the reply into
/// `reply`, return its length. `capability` is the handle the request
/// carried, if any (M13 cap passing — how a subscriber hands over its
/// endpoint). The zero-length ping never reaches this.
on_message: *const fn (message: []const u8, reply: []u8, sender: u32, capability: ?ipc.Handle) usize,
/// A notification that is not a signal — a subscribed exit event, a bound
/// IRQ, a timer landing. The raw badge; decode with the ipc helpers.
on_notification: ?*const fn (badge: u64) void = null,
/// The reload signal. Default: ignored.
on_reload: ?*const fn () void = null,
/// The terminate signal, called before the loop returns. The clean exit is
/// the return itself — never put *necessary* work here (iron rule 1: a kill
/// arrives with no warning; this is for graceful extras only).
on_terminate: ?*const fn () void = null,
/// Publish the endpoint under a well-known service id at startup.
service: ?abi.ServiceId = null,
};
/// Run the service: create and (optionally) register the endpoint, bind signals
/// to it, call `init`, then serve until `terminate` arrives — at which point the
/// loop returns and main's return is the clean exit the supervisor reads as
/// `ExitReason.exited`. `maximum_message` sizes the receive and reply buffers
/// (a service passes its protocol's message maximum).
pub fn run(comptime maximum_message: usize, callbacks: Callbacks) void {
const endpoint = ipc.createIpcEndpoint() orelse return;
if (callbacks.service) |id| {
if (!ipc.register(id, endpoint)) return;
}
_ = process.bindSignals(endpoint);
if (callbacks.init) |initialise| {
if (!initialise(endpoint)) return;
}
var reply_buffer: [maximum_message]u8 = undefined;
var reply_len: usize = 0;
var receive: [maximum_message]u8 = undefined;
while (true) {
const got = ipc.replyWait(endpoint, reply_buffer[0..reply_len], &receive, null);
if (got.isNotification()) {
reply_len = 0; // nothing owed for a notification
if (process.signalsFrom(got.badge)) |signals| {
if (signals.has(.reload)) {
if (callbacks.on_reload) |onReload| onReload();
}
if (signals.has(.terminate)) {
if (callbacks.on_terminate) |onTerminate| onTerminate();
return; // the loop's return IS the clean exit
}
continue;
}
if (callbacks.on_notification) |onNotification| onNotification(got.badge);
continue;
}
if (got.len == 0) {
reply_len = 0; // the universal ping: a zero-length reply, from the harness
continue;
}
reply_len = callbacks.on_message(receive[0..got.len], &reply_buffer, got.senderTaskId(), got.cap);
}
}
+9
View File
@@ -32,6 +32,15 @@ pub fn sleep(ms: usize) void {
_ = sc.systemCall1(.sleep, ms); _ = sc.systemCall1(.sleep, ms);
} }
/// Arm a one-shot timer: after `ms` milliseconds the kernel posts a timer
/// notification (`ipc.Received.isTimer`) to `endpoint`. The timed wait of
/// docs/process-lifecycle.md — a service arms a deadline and keeps serving,
/// instead of blocking in sleep; what stop-sequence escalation, hello deadlines,
/// and restart backoff are built from.
pub fn timerOnce(endpoint: usize, ms: u64) bool {
return sc.systemCall2(.timer_bind, endpoint, ms) == 0;
}
/// Monotonic nanoseconds since boot — a time source for timeouts and short delays. It /// Monotonic nanoseconds since boot — a time source for timeouts and short delays. It
/// only ever moves forward. This is *not* wall-clock time (no date, no timezone — that /// only ever moves forward. This is *not* wall-clock time (no date, no timezone — that
/// is a user-space service layered on top). Deadline pattern for a bounded poll loop: /// is a user-space service layered on top). Deadline pattern for a bounded poll loop:
+52
View File
@@ -53,9 +53,31 @@ pub const SystemCall = enum(u64) {
process_enumerate = 24, // process_enumerate(buffer, maximum) -> total: snapshot the task table process_enumerate = 24, // process_enumerate(buffer, maximum) -> total: snapshot the task table
process_kill = 25, // process_kill(id) -> 0/-errno: end a process this process spawned process_kill = 25, // process_kill(id) -> 0/-errno: end a process this process spawned
ipc_send = 26, // ipc_send(handle, message_ptr, message_len) -> 0/-errno: post a payload to an endpoint's async queue without blocking ipc_send = 26, // ipc_send(handle, message_ptr, message_len) -> 0/-errno: post a payload to an endpoint's async queue without blocking
process_exit_reason = 27, // process_exit_reason(id) -> ExitReason/-errno: how a dead child ended (its supervisor only)
process_subscribe = 28, // process_subscribe(endpoint) -> 0/-errno: subscribe to published exit events — every death posts a notification
signal_bind = 29, // signal_bind(endpoint) -> 0/-errno: nominate the endpoint this process's signals arrive on
process_signal = 30, // process_signal(id, signal) -> 0/-errno: post a signal to a child (or to yourself)
timer_bind = 31, // timer_bind(endpoint, ms) -> 0/-errno: one-shot timer — posts a notification when ms elapse
_, _,
}; };
/// How a process ended — recorded by the kernel at death, queried by the
/// supervisor with `process_exit_reason`, and the input to its restart decision
/// (docs/process-lifecycle.md): a clean exit meant to stop, a fault wants a
/// restart with backoff, killed means the supervisor did it itself. The faults
/// mirror the CPU exceptions a ring-3 process can die of; they are exit reasons,
/// never delivered to the faulting process (recovery is restart, not a handler).
pub const ExitReason = enum(u8) {
exited = 0, // returned from main / called exit
aborted = 1, // deliberate self-termination (reserved: no abort path yet)
segmentation_fault = 2, // page fault
illegal_instruction = 3, // invalid opcode
arithmetic_fault = 4, // divide error, x87 or SIMD fault
protection_fault = 5, // general protection fault
fault = 6, // any other CPU exception
killed = 7, // process_kill
};
/// The x86 MSI message address base (`0xFEE0_0000`): a device raises an MSI by writing /// The x86 MSI message address base (`0xFEE0_0000`): a device raises an MSI by writing
/// `data` to this address, which the Local APIC turns into an interrupt at the vector /// `data` to this address, which the Local APIC turns into an interrupt at the vector
/// in `data`. The kernel returns the concrete (address, data) from `msi_bind`; this is /// in `data`. The kernel returns the concrete (address, data) from `msi_bind`; this is
@@ -93,6 +115,35 @@ pub const notify_exit_bit: u64 = 1 << 62;
/// broadcasts where a rendezvous is the wrong shape (the input service is the first user). /// broadcasts where a rendezvous is the wrong shape (the input service is the first user).
pub const notify_message_bit: u64 = 1 << 61; pub const notify_message_bit: u64 = 1 << 61;
/// Set (alongside `notify_badge_bit`) in the badge of a **signal notification** —
/// the process-lifecycle vocabulary of docs/process-lifecycle.md, delivered to the
/// endpoint the process nominated with `signal_bind`. The low bits carry the
/// coalesced pending mask (bit positions = `Signal` values): signals are
/// statements, not questions, and two pending terminates are one terminate.
pub const notify_signal_bit: u64 = 1 << 60;
/// Set (alongside `notify_badge_bit`) in the badge of a **timer notification** —
/// a one-shot `timer_bind` deadline landing. No payload bits: what to do when the
/// deadline fires is whatever the receiver armed it for (a stop-sequence
/// escalation, a restart backoff, an alarm).
pub const notify_timer_bit: u64 = 1 << 59;
/// The signal vocabulary (docs/process-lifecycle.md): POSIX's concepts, danos's
/// names, message delivery. The value is the bit position in the pending mask — a
/// private kernel/runtime detail, free to change while they ship together. Kill
/// is not here (it is `process_kill`, unhandleable by definition); faults are not
/// here (they are `ExitReason`s — recovery is restart, not a handler); liveness is
/// not here (a question, asked as the zero-length ping call, not a statement).
pub const Signal = enum(u5) {
terminate = 0, // finish up and exit (the polite half of the stop sequence)
reload = 1, // re-read configuration / re-scan
interrupt = 2, // interactive interrupt (no sender until a console exists)
quit = 3, // as interrupt, by convention more final
alarm = 4, // a timer the process armed for itself (unbuilt: no consumer yet)
user_1 = 5, // service-defined
user_2 = 6, // service-defined
};
/// Capacity of `ProcessDescriptor.name` — matches the longest name `system_spawn` /// Capacity of `ProcessDescriptor.name` — matches the longest name `system_spawn`
/// accepts, so a process's recorded name (its argv[0]) is never truncated. /// accepts, so a process's recorded name (its argv[0]) is never truncated.
pub const maximum_process_name = 64; pub const maximum_process_name = 64;
@@ -125,6 +176,7 @@ pub const ServiceId = enum(u32) {
vfs = 1, vfs = 1,
input = 2, input = 2,
ps2_bus = 3, // the 8042 owner; child device drivers attach here for raw bytes ps2_bus = 3, // the 8042 owner; child device drivers attach here for raw bytes
device_manager = 4, // the tree, the matcher, the supervisor (docs/device-manager.md)
_, _,
}; };
+149 -45
View File
@@ -1,68 +1,172 @@
//! /system/drivers/usb-xhci-bus — the xHCI (USB 3) host-controller bus driver. //! /system/drivers/usb-xhci-bus — the xHCI (USB 3) host-controller bus driver.
//! The device manager spawns **one instance per controller** it discovers (a machine //! The device manager spawns **one instance per controller** it discovers (a
//! can carry several), passing the controller's device-tree id as argv[1]; this //! machine can carry several), passing the controller's device-tree id as
//! instance claims that device and no other, so multiple instances never fight over //! argv[1]; this instance claims that device and no other, so multiple
//! hardware. This increment proves the plumbing: parse the id, claim the controller, //! instances never fight over hardware.
//! and report its MMIO window. The next increments map the registers and bring the //!
//! controller up (reset, rings, port scan), then enumerate the USB devices on the //! M18.2 (this increment): after the hello, real hardware — map the xHC's
//! bus with the usb-abi request builders and publish each with `device_register`. //! register window (the first memory BAR; resource 0 is the ECAM config
//! space), read the capability registers, and walk the root-hub ports: one
//! `child_added` report to the manager per connected port, carrying the port
//! number and the PORTSC speed class as identity. No transfer rings yet —
//! descriptors and USB class matching are the USB track; the connect bit and
//! speed come straight from PORTSC, which reflects hardware state whether or
//! not the controller is running.
const std = @import("std"); const std = @import("std");
const runtime = @import("runtime"); const runtime = @import("runtime");
const protocol = runtime.device_manager_protocol;
const device = runtime.device; const device = runtime.device;
/// Format one whole log line and emit it in a single `debug_write`, so concurrent /// Format one whole log line and emit it in a single `debug_write`, so
/// instances (one per controller) can never interleave mid-line. /// concurrent instances (one per controller) can never interleave mid-line.
fn writeLine(comptime fmt: []const u8, arguments: anytype) void { fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined; var line: [128]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return); _ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
} }
var controller_id: u64 = protocol.no_device;
/// Claim the assigned controller, find its register window, and hello the
/// manager. Any failure returns false: the process exits cleanly, which the
/// manager reads as "meant to stop" — a missing assignment is not a crash loop.
fn initialise(endpoint: runtime.ipc.Handle) bool {
_ = endpoint;
if (!device.claim(controller_id)) {
writeLine("usb-xhci-bus: unable to claim controller device {d}\n", .{controller_id});
return false;
}
// Fetch our own descriptor back for the controller's resources.
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("usb-xhci-bus: out of memory\n");
return false;
};
const total = device.enumerate(buffer);
const descriptor = for (buffer[0..@min(total, buffer.len)]) |d| {
if (d.id == controller_id) break d;
} else {
writeLine("usb-xhci-bus: device {d} not in the device tree\n", .{controller_id});
return false;
};
// The xHC's registers live behind the first memory BAR. Resource 0 is the
// function's ECAM configuration space (M15), so the walk starts at 1.
var register_index: u64 = 0;
const register_window = for (descriptor.resources[1..@intCast(descriptor.resource_count)], 1..) |resource, index| {
if (resource.kind == @intFromEnum(device.ResourceKind.memory)) {
register_index = index;
break resource;
}
} else {
writeLine("usb-xhci-bus: controller device {d} has no register BAR\n", .{controller_id});
return false;
};
writeLine("usb-xhci-bus: claimed controller device {d} (registers at 0x{x}, {d} bytes)\n", .{
controller_id,
register_window.start,
register_window.len,
});
register_base = device.mmioMap(controller_id, register_index) orelse {
_ = runtime.system.write("usb-xhci-bus: mmio_map failed\n");
return false;
};
// The handshake: role, protocol version, assignment — inside the manager's
// deadline (the lookup retries cover the manager still registering).
var manager: ?runtime.ipc.Handle = null;
var tries: u32 = 0;
while (manager == null and tries < 100) : (tries += 1) {
manager = runtime.ipc.lookup(.device_manager);
if (manager == null) runtime.system.sleep(20);
}
const h = manager orelse {
_ = runtime.system.write("usb-xhci-bus: no device manager to hello\n");
return false;
};
const hello = protocol.Hello{ .role = @intFromEnum(protocol.Role.bus), .device_id = controller_id };
var reply: [protocol.message_maximum]u8 = undefined;
const n = runtime.ipc.call(h, std.mem.asBytes(&hello), &reply) catch {
_ = runtime.system.write("usb-xhci-bus: hello call failed\n");
return false;
};
if (n < protocol.reply_size or std.mem.bytesToValue(protocol.HelloReply, reply[0..protocol.reply_size]).status != 0) {
_ = runtime.system.write("usb-xhci-bus: hello refused\n");
return false;
}
_ = runtime.system.write("usb-xhci-bus: hello acknowledged\n");
scanPorts(h);
return true;
}
var register_base: usize = 0;
/// One 32-bit volatile register read at `offset` from the mapped window.
fn readRegister(offset: usize) u32 {
const register: *volatile u32 = @ptrFromInt(register_base + offset);
return register.*;
}
/// The root-hub port scan: read the capability registers for the port count
/// and the operational-register offset, then one PORTSC per port. The connect
/// bit (CCS) and the speed field reflect hardware state directly — no
/// controller reset or run needed to *see* the devices; driving them needs the
/// rings (the USB track).
fn scanPorts(manager: runtime.ipc.Handle) void {
// Capability registers: CAPLENGTH is byte 0 of the first dword; HCSPARAMS1
// carries MaxPorts in bits 31:24.
const capability_length = readRegister(0) & 0xFF;
const structural = readRegister(0x04);
const maximum_ports: u32 = structural >> 24;
writeLine("usb-xhci-bus: {d} root-hub ports\n", .{maximum_ports});
// PORTSC registers: operational base + 0x400 + 0x10 per port (1-based).
var port: u32 = 1;
var connected: u32 = 0;
while (port <= maximum_ports) : (port += 1) {
const port_status = readRegister(capability_length + 0x400 + 0x10 * (port - 1));
if (port_status & 1 == 0) continue; // CCS: nothing connected
connected += 1;
const speed = (port_status >> 10) & 0xF; // the PORTSC port-speed class
writeLine("usb-xhci-bus: port {d} connected (speed class {d})\n", .{ port, speed });
const report = protocol.ChildAdded{
.parent = controller_id,
.bus_address = port,
.identity = speed,
};
var reply: [protocol.message_maximum]u8 = undefined;
_ = runtime.ipc.call(manager, std.mem.asBytes(&report), &reply) catch {
writeLine("usb-xhci-bus: child report for port {d} failed\n", .{port});
continue;
};
}
if (connected == 0) _ = runtime.system.write("usb-xhci-bus: no devices connected\n");
}
/// No bus protocol to serve yet — transfer requests arrive with the USB track.
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = message;
_ = reply;
_ = sender;
_ = capability;
return 0;
}
pub fn main(init: runtime.process.Init) void { pub fn main(init: runtime.process.Init) void {
const argument = init.arguments.get(1) orelse { const argument = init.arguments.get(1) orelse {
_ = runtime.system.write("usb-xhci-bus: missing controller device id (argv[1])\n"); _ = runtime.system.write("usb-xhci-bus: missing controller device id (argv[1])\n");
return; return;
}; };
const controller_id = std.fmt.parseInt(u64, argument, 10) catch { controller_id = std.fmt.parseInt(u64, argument, 10) catch {
writeLine("usb-xhci-bus: malformed controller device id '{s}'\n", .{argument}); writeLine("usb-xhci-bus: malformed controller device id '{s}'\n", .{argument});
return; return;
}; };
runtime.service.run(protocol.message_maximum, .{
if (!device.claim(controller_id)) { .init = initialise,
writeLine("usb-xhci-bus: unable to claim controller device {d}\n", .{controller_id}); .on_message = onMessage,
return;
}
// Fetch our own descriptor back for the controller's resources.
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("usb-xhci-bus: out of memory\n");
return;
};
const total = device.enumerate(buffer);
const descriptor = for (buffer[0..@min(total, buffer.len)]) |d| {
if (d.id == controller_id) break d;
} else {
writeLine("usb-xhci-bus: device {d} not in the device tree\n", .{controller_id});
return;
};
// The controller's operational registers live behind BAR0, enumerated as the
// device's first memory resource.
const register_window = for (descriptor.resources[0..@intCast(descriptor.resource_count)]) |resource| {
if (resource.kind == @intFromEnum(device.ResourceKind.memory)) break resource;
} else {
writeLine("usb-xhci-bus: controller device {d} has no MMIO window\n", .{controller_id});
return;
};
writeLine("usb-xhci-bus: claimed controller device {d} (registers at 0x{x}, {d} bytes)\n", .{
controller_id,
register_window.start,
register_window.len,
}); });
// Controller bring-up (map the window, reset, rings, port scan) is the next
// increment; stay resident as the bus's supervisor in the meantime.
while (true) runtime.system.sleep(1000);
} }
pub const panic = runtime.panic; pub const panic = runtime.panic;
+13
View File
@@ -104,6 +104,19 @@ pub fn ownerOf(id: u64) ?u32 {
return claimed[@intCast(id)]; return claimed[@intCast(id)];
} }
/// Release every claim held by `owner` — called by the process layer on every
/// path out of a process (exit, fault, kill), so a restarted driver can claim its
/// hardware again (docs/process-lifecycle.md iron rule 1: cleanup is the kernel's
/// job). The devices stay in the table — they describe hardware, which did not go
/// away — only their ownership clears.
pub fn releaseAllOwnedBy(owner: u32) void {
for (claimed[0..count]) |*slot| {
if (slot.*) |o| {
if (o == owner) slot.* = null;
}
}
}
/// Resource `index` of device `id`, or null if out of range. /// Resource `index` of device `id`, or null if out of range.
pub fn resourceOf(id: u64, index: u64) ?device_abi.ResourceDescriptor { pub fn resourceOf(id: u64, index: u64) ?device_abi.ResourceDescriptor {
if (id >= count) return null; if (id >= count) return null;
+14 -1
View File
@@ -412,13 +412,26 @@ fn recoverableFault(vector: u64) bool {
/// plus a POST code and a persistent breadcrumb. (A ring-3 fault on a *borrowed* /// plus a POST code and a persistent breadcrumb. (A ring-3 fault on a *borrowed*
/// kernel thread — process.run, the user-pf isolation probe — also lands here: there /// kernel thread — process.run, the user-pf isolation probe — also lands here: there
/// is no scheduled process to kill.) /// is no scheduled process to kill.)
/// Classify a CPU exception vector as the ExitReason a supervisor reads — the
/// fault classes of docs/process-lifecycle.md. Faults are exit reasons, never
/// signals delivered to the faulting process: recovery is restart, not a handler.
fn exitReasonForVector(vector: u64) abi.ExitReason {
return switch (vector) {
14 => .segmentation_fault, // page fault
6 => .illegal_instruction, // invalid opcode
0, 16, 19 => .arithmetic_fault, // divide error, x87, SIMD
13 => .protection_fault, // general protection
else => .fault,
};
}
fn onException(state: *const architecture.CpuState) noreturn { fn onException(state: *const architecture.CpuState) noreturn {
if (architecture.fromUser(state) and scheduler.currentIsUserProcess() and recoverableFault(state.vector)) { if (architecture.fromUser(state) and scheduler.currentIsUserProcess() and recoverableFault(state.vector)) {
statusPrint("\ndanos: process {d} ({s}) killed by {s} (vector {d}) on core {d}\n", .{ scheduler.currentId(), scheduler.current().name(), architecture.exceptionName(state.vector), state.vector, scheduler.currentCpuIndex() }); statusPrint("\ndanos: process {d} ({s}) killed by {s} (vector {d}) on core {d}\n", .{ scheduler.currentId(), scheduler.current().name(), architecture.exceptionName(state.vector), state.vector, scheduler.currentCpuIndex() });
statusPrint(" error code : 0x{x}\n", .{state.error_code}); statusPrint(" error code : 0x{x}\n", .{state.error_code});
statusPrint(" IP : 0x{x:0>16}\n", .{architecture.instructionPointer(state)}); statusPrint(" IP : 0x{x:0>16}\n", .{architecture.instructionPointer(state)});
if (architecture.faultAddress(state)) |address| statusPrint(" fault addr : 0x{x:0>16}\n", .{address}); if (architecture.faultAddress(state)) |address| statusPrint(" fault addr : 0x{x:0>16}\n", .{address});
process.killCurrentProcess(); // reclaims everything, reschedules; never returns process.killCurrentProcess(exitReasonForVector(state.vector)); // reclaims everything, reschedules; never returns
} }
log.checkpoint(cp_exception); log.checkpoint(cp_exception);
+205 -3
View File
@@ -137,6 +137,7 @@ pub fn init() void {
architecture.setSystemCallHandler(system_call); architecture.setSystemCallHandler(system_call);
scheduler.terminate_current_hook = terminateCurrentLocked; scheduler.terminate_current_hook = terminateCurrentLocked;
scheduler.reap_task_hook = reapTaskLocked; scheduler.reap_task_hook = reapTaskLocked;
scheduler.timer_tick_hook = timerSweepLocked;
} }
/// Return -1 (as an unsigned bit pattern) in the system_call result register. /// Return -1 (as an unsigned bit pattern) in the system_call result register.
@@ -164,6 +165,7 @@ fn system_call(state: *architecture.CpuState) void {
// A scheduled process tears down fully (terminateCurrent); a borrowed // A scheduled process tears down fully (terminateCurrent); a borrowed
// test thread unwinds back to the kernel that entered it. // test thread unwinds back to the kernel that entered it.
if (scheduler.currentIsUserProcess()) { if (scheduler.currentIsUserProcess()) {
scheduler.current().exit_reason = .exited;
terminateCurrent(); terminateCurrent();
} else architecture.userExit(); } else architecture.userExit();
}, },
@@ -199,6 +201,11 @@ fn system_call(state: *architecture.CpuState) void {
.clock => systemClock(state), .clock => systemClock(state),
.process_enumerate => systemProcessEnumerate(state), .process_enumerate => systemProcessEnumerate(state),
.process_kill => systemProcessKill(state), .process_kill => systemProcessKill(state),
.process_exit_reason => systemProcessExitReason(state),
.process_subscribe => systemProcessSubscribe(state),
.signal_bind => systemSignalBind(state),
.process_signal => systemProcessSignal(state),
.timer_bind => systemTimerBind(state),
_ => fail(state), _ => fail(state),
} }
} }
@@ -552,7 +559,11 @@ pub var fault_kill_count: u64 = 0;
/// endpoint reference destroys the Endpoint, and a still-bound GSI would have an /// endpoint reference destroys the Endpoint, and a still-bound GSI would have an
/// ISR call notifyFromIsr on freed memory the next time the device fired. /// ISR call notifyFromIsr on freed memory the next time the device fired.
/// `releaseOwner` also leaves the line masked, so a dead driver's device goes /// `releaseOwner` also leaves the line masked, so a dead driver's device goes
/// quiet rather than storming. /// quiet rather than storming. (It drops MSI vectors by the same owner sweep.)
/// - Device claims are released with the IRQ bindings, so a restarted driver can
/// claim the same hardware again — the cleanup half of process-lifecycle.md's
/// iron rule 1. Claims hold no pointers, so ordering is free; they go here so
/// the exit notification (below, last) observes a fully-released child.
/// - A client this task still owes a reply to (it died between receive and reply) /// - A client this task still owes a reply to (it died between receive and reply)
/// is failed with -EPEER rather than left blocked forever — a dead server must /// is failed with -EPEER rather than left blocked forever — a dead server must
/// not hang its callers. /// not hang its callers.
@@ -565,7 +576,33 @@ pub var fault_kill_count: u64 = 0;
/// reference taken at spawn is dropped with it. /// reference taken at spawn is dropped with it.
/// Precondition: the big kernel lock is held. /// Precondition: the big kernel lock is held.
fn releaseTaskResourcesLocked(t: *scheduler.Task) void { fn releaseTaskResourcesLocked(t: *scheduler.Task) void {
recordExitLocked(t);
irq.releaseOwner(t.id); irq.releaseOwner(t.id);
devices_broker.releaseAllOwnedBy(t.id);
// The dying task's signal endpoint and one-shot timers go with it.
if (t.signal_endpoint) |raw| {
ipc.dropRef(@ptrCast(@alignCast(raw)));
t.signal_endpoint = null;
}
t.pending_signals = 0;
for (&one_shot_timers) |*slot| {
if (slot.*) |timer| {
if (timer.owner == t.id) {
ipc.dropRef(timer.endpoint);
slot.* = null;
}
}
}
// A dead subscriber's own subscriptions go first: it must not hear about
// itself, and the slots' endpoint references drop with it.
for (&exit_subscribers) |*slot| {
if (slot.*) |subscriber| {
if (subscriber.owner == t.id) {
ipc.dropRef(subscriber.endpoint);
slot.* = null;
}
}
}
if (t.ipc_client) |client| { if (t.ipc_client) |client| {
t.ipc_client = null; t.ipc_client = null;
client.ipc_status = -ipc.EPEER; client.ipc_status = -ipc.EPEER;
@@ -575,6 +612,12 @@ fn releaseTaskResourcesLocked(t: *scheduler.Task) void {
scheduler.removeFromWaitQueueLocked(t); scheduler.removeFromWaitQueueLocked(t);
scheduler.forgetIpcClientLocked(t); scheduler.forgetIpcClientLocked(t);
ipc.closeHandles(t); ipc.closeHandles(t);
// Publish the exit to every subscriber (docs/process-lifecycle.md): the same
// badge encoding as the supervisor's notification, and equally late, so a
// subscriber also observes a fully-released child.
for (&exit_subscribers) |*slot| {
if (slot.*) |subscriber| ipc.notifyLocked(subscriber.endpoint, abi.notify_exit_bit | t.id);
}
if (t.exit_endpoint) |raw| { if (t.exit_endpoint) |raw| {
const endpoint: *ipc.Endpoint = @ptrCast(@alignCast(raw)); const endpoint: *ipc.Endpoint = @ptrCast(@alignCast(raw));
t.exit_endpoint = null; t.exit_endpoint = null;
@@ -628,6 +671,7 @@ pub fn killProcess(caller_id: u32, target_id: u32) i64 {
const target = scheduler.taskByIdLocked(target_id) orelse return -ipc.ESRCH; const target = scheduler.taskByIdLocked(target_id) orelse return -ipc.ESRCH;
if (target.aspace == 0) return -ipc.ESRCH; // kernel tasks are not processes if (target.aspace == 0) return -ipc.ESRCH; // kernel tasks are not processes
if (target.supervisor != caller_id) return -ipc.EPERM; if (target.supervisor != caller_id) return -ipc.EPERM;
target.exit_reason = .killed;
if (target.state == .running) { if (target.state == .running) {
target.kill_pending = true; target.kill_pending = true;
} else { } else {
@@ -640,12 +684,170 @@ pub fn killProcess(caller_id: u32, target_id: u32) i64 {
/// The fault is confined to the process — the kernel trapped it on the task's own /// The fault is confined to the process — the kernel trapped it on the task's own
/// kernel stack and is intact — so everything the process held is reclaimed and the /// kernel stack and is intact — so everything the process held is reclaimed and the
/// core reschedules. The system keeps running; only the faulting process dies /// core reschedules. The system keeps running; only the faulting process dies
/// (docs/resilience.md: fault -> kill -> continue). /// (docs/resilience.md: fault -> kill -> continue). `reason` is the fault class
pub fn killCurrentProcess() noreturn { /// (from the vector), recorded for the supervisor's `process_exit_reason`.
pub fn killCurrentProcess(reason: abi.ExitReason) noreturn {
scheduler.current().exit_reason = reason;
fault_kill_count += 1; fault_kill_count += 1;
terminateCurrent(); terminateCurrent();
} }
/// The bounded record of recent deaths, for `process_exit_reason`: ids are never
/// reused, so a ring keyed by id is enough — a record evicted by wraparound reads
/// as -ESRCH, the same as an id that never lived, which a supervisor treats as
/// "too late to ask". Written under the big kernel lock by the reap.
const exit_record_capacity = 64;
const ExitRecord = struct { id: u32 = 0, supervisor: u32 = 0, reason: abi.ExitReason = .exited, valid: bool = false };
var exit_records: [exit_record_capacity]ExitRecord = .{ExitRecord{}} ** exit_record_capacity;
var exit_record_next: usize = 0;
/// Record a dying task's (id, supervisor, reason) — called by the reap before the
/// exit notification is posted, so a supervisor that hears the notification can
/// always still query the reason. Precondition: the big kernel lock is held.
fn recordExitLocked(t: *scheduler.Task) void {
exit_records[exit_record_next] = .{ .id = t.id, .supervisor = t.supervisor, .reason = t.exit_reason, .valid = true };
exit_record_next = (exit_record_next + 1) % exit_record_capacity;
}
/// How dead process `id` ended, for `caller` — the kernel half of the
/// process_exit_reason system call. Returns the ExitReason value, -ESRCH (never
/// lived, still alive, or evicted from the ring), or -EPERM (the caller was not
/// its supervisor — the same authority gate as process_kill).
pub fn exitReasonOf(caller_id: u32, target_id: u32) i64 {
const flags = sync.enter();
defer sync.leave(flags);
for (&exit_records) |*record| {
if (record.valid and record.id == target_id) {
if (record.supervisor != caller_id) return -ipc.EPERM;
return @intFromEnum(record.reason);
}
}
return -ipc.ESRCH;
}
/// The published exit events' subscribers (docs/process-lifecycle.md "Who learns
/// of a death"): stateful services — the VFS's file handles, input's
/// subscriptions — that must release what a dead client held and cannot learn it
/// any other way (a client that simply never calls again looks like silence).
/// Bounded like every kernel table; each entry holds its own endpoint reference.
const exit_subscriber_capacity = 8;
const ExitSubscriber = struct { endpoint: *ipc.Endpoint, owner: u32 };
var exit_subscribers: [exit_subscriber_capacity]?ExitSubscriber = .{null} ** exit_subscriber_capacity;
/// process_subscribe(endpoint): subscribe the caller's endpoint to published exit
/// events. Ungated, like process_enumerate — what is running (and dying) is not a
/// secret between cooperating processes. -ENOSPC when the table is full.
fn systemProcessSubscribe(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
const flags = sync.enter();
defer sync.leave(flags);
for (&exit_subscribers) |*slot| {
if (slot.* == null) {
endpoint.refcount += 1; // the slot's own reference, dropped on unsubscribe-by-death
slot.* = .{ .endpoint = endpoint, .owner = t.id };
return architecture.setSystemCallResult(state, 0);
}
}
failErr(state, ipc.ENOSPC);
}
/// signal_bind(endpoint): nominate where this process's signals arrive — the
/// IRQ-as-IPC pattern a fourth time (docs/process-lifecycle.md). Replacing a
/// binding drops the old reference; signals that pended while unbound are
/// delivered immediately on bind, coalesced into one notification.
fn systemSignalBind(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
const flags = sync.enter();
defer sync.leave(flags);
if (t.signal_endpoint) |raw| ipc.dropRef(@ptrCast(@alignCast(raw)));
endpoint.refcount += 1;
t.signal_endpoint = @ptrCast(endpoint);
if (t.pending_signals != 0) {
ipc.notifyLocked(endpoint, abi.notify_signal_bit | t.pending_signals);
t.pending_signals = 0;
}
architecture.setSystemCallResult(state, 0);
}
/// process_signal(id, signal): post a signal — a one-way, coalescing statement,
/// never a question (docs/process-lifecycle.md). The authority gate is the
/// supervision link, like kill; a process may also signal itself. Unbound
/// targets accumulate the signal in their pending mask.
fn systemProcessSignal(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const id = architecture.systemCallArg(state, 0);
const signal = architecture.systemCallArg(state, 1);
if (id > std.math.maxInt(u32)) return failErr(state, ipc.ESRCH);
if (signal > 31) return failErr(state, ipc.EBADF); // not a Signal bit position
const flags = sync.enter();
defer sync.leave(flags);
const target = scheduler.taskByIdLocked(@intCast(id)) orelse return failErr(state, ipc.ESRCH);
if (target.aspace == 0) return failErr(state, ipc.ESRCH);
if (target.supervisor != t.id and target.id != t.id) return failErr(state, ipc.EPERM);
target.pending_signals |= @as(u32, 1) << @intCast(signal);
if (target.signal_endpoint) |raw| {
const endpoint: *ipc.Endpoint = @ptrCast(@alignCast(raw));
ipc.notifyLocked(endpoint, abi.notify_signal_bit | target.pending_signals);
target.pending_signals = 0;
}
architecture.setSystemCallResult(state, 0);
}
/// The one-shot timers of timer_bind: the missing timed wait. A service arms a
/// deadline and keeps serving; the expiry arrives in the same replyWait as
/// everything else (notify_timer_bit). What stop-sequence escalation, hello
/// deadlines, and restart backoff are built from — and later, `alarm`.
const timer_capacity = 16;
const OneShotTimer = struct { deadline: u64, endpoint: *ipc.Endpoint, owner: u32 };
var one_shot_timers: [timer_capacity]?OneShotTimer = .{null} ** timer_capacity;
/// Sweep expired timers — hung on scheduler.timer_tick_hook, so it runs on every
/// tick with the big kernel lock held, like the sleeper wake it rides beside.
fn timerSweepLocked() void {
const now = architecture.millis();
for (&one_shot_timers) |*slot| {
if (slot.*) |timer| {
if (now >= timer.deadline) {
ipc.notifyLocked(timer.endpoint, abi.notify_timer_bit);
ipc.dropRef(timer.endpoint);
slot.* = null;
}
}
}
}
/// timer_bind(endpoint, ms): arm a one-shot timer. -ENOSPC when the table is full.
fn systemTimerBind(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const endpoint = ipc.resolveHandle(t, architecture.systemCallArg(state, 0)) orelse return failErr(state, ipc.EBADF);
const ms = architecture.systemCallArg(state, 1);
const flags = sync.enter();
defer sync.leave(flags);
for (&one_shot_timers) |*slot| {
if (slot.* == null) {
endpoint.refcount += 1;
slot.* = .{ .deadline = architecture.millis() + ms, .endpoint = endpoint, .owner = t.id };
return architecture.setSystemCallResult(state, 0);
}
}
failErr(state, ipc.ENOSPC);
}
fn systemProcessExitReason(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const id = architecture.systemCallArg(state, 0);
if (id > std.math.maxInt(u32)) return failErr(state, ipc.ESRCH);
const r = exitReasonOf(t.id, @intCast(id));
architecture.setSystemCallResult(state, @bitCast(r));
}
/// Resolve `(device_id, resource_index)` to a GSI this process is entitled to bind, or null. /// Resolve `(device_id, resource_index)` to a GSI this process is entitled to bind, or null.
/// The two checks are the whole security story: the device must be *claimed* by the /// The two checks are the whole security story: the device must be *claimed* by the
/// caller, and the resource must be one of that device's `irq` resources as recorded /// caller, and the resource must be one of that device's `irq` resources as recorded
+20 -2
View File
@@ -50,6 +50,17 @@ pub const Task = struct {
// null. Holds its own reference, dropped when the notification is posted. // null. Holds its own reference, dropped when the notification is posted.
// Opaque here for the same reason as `handles` below. // Opaque here for the same reason as `handles` below.
exit_endpoint: ?*anyopaque = null, exit_endpoint: ?*anyopaque = null,
// How this process ended — set by the death paths (exit, fault, kill) just
// before the reap records it for `process_exit_reason`. Meaningless while
// the task lives.
exit_reason: abi.ExitReason = .exited,
// Endpoint this process's signals arrive on (signal_bind), or null — same
// ownership rules as exit_endpoint (holds a reference; opaque here).
signal_endpoint: ?*anyopaque = null,
// Signals posted but not yet delivered: the coalescing pending mask
// (docs/process-lifecycle.md). Bits are abi.Signal values. Signals pend here
// until an endpoint is bound; two pending terminates are one terminate.
pending_signals: u32 = 0,
// Set by process_kill on a task that is running on another core; the kernel // Set by process_kill on a task that is running on another core; the kernel
// finishes the kill at that task's next system call or timer tick. // finishes the kill at that task's next system call or timer tick.
kill_pending: bool = false, kill_pending: bool = false,
@@ -339,8 +350,9 @@ pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, priority: Priority
/// context switch and lock release. /// context switch and lock release.
fn startUserTask() void { fn startUserTask() void {
const t = current(); const t = current();
var buffer: [96]u8 = undefined; // No serial chatter here: this runs on every spawn, unserialized against
architecture.serialWrite(std.fmt.bufPrint(&buffer, "DBG startUserTask ip=0x{x} sp=0x{x} aspace=0x{x} kstack=0x{x}\n", .{ t.user_ip, t.user_sp, t.aspace, t.kstack_top }) catch ""); // user-space writes, and its output used to shear concurrent log lines in
// half — the largest source of corrupted markers in the QEMU scenarios.
architecture.jumpToUser(t.user_ip, t.user_sp); // noreturn architecture.jumpToUser(t.user_ip, t.user_sp); // noreturn
} }
@@ -643,9 +655,15 @@ fn reapKillPendingLocked() void {
/// other critical section, but releases it *without* touching the interrupt flag /// other critical section, but releases it *without* touching the interrupt flag
/// — the handler's `iretq` restores the interrupted context's flags, so /// — the handler's `iretq` restores the interrupted context's flags, so
/// re-enabling here would open a nested-interrupt window before the return. /// re-enabling here would open a nested-interrupt window before the return.
/// Called from the tick with the big kernel lock held — process.zig hangs the
/// one-shot timer sweep here (timer_bind), the same call-up pattern as the
/// teardown hooks below.
pub var timer_tick_hook: ?*const fn () void = null;
pub fn tick() void { pub fn tick() void {
_ = sync.enter(); _ = sync.enter();
wakeExpired(); wakeExpired();
if (timer_tick_hook) |hook| hook();
reapKillPendingLocked(); reapKillPendingLocked();
if (preemption_enabled) schedule(); if (preemption_enabled) schedule();
sync.leaveIsr(); sync.leaveIsr();
+320 -9
View File
@@ -132,6 +132,18 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
processKillTest(boot_information); processKillTest(boot_information);
} else if (eql(case, "supervision")) { } else if (eql(case, "supervision")) {
supervisionTest(boot_information); supervisionTest(boot_information);
} else if (eql(case, "claim-release")) {
claimReleaseTest(boot_information);
} else if (eql(case, "vfs-client-death")) {
vfsClientDeathTest(boot_information);
} else if (eql(case, "signals")) {
signalsTest(boot_information);
} else if (eql(case, "driver-restart")) {
driverRestartTest(boot_information);
} else if (eql(case, "usb-report")) {
usbReportTest(boot_information);
} else if (eql(case, "device-list")) {
deviceListTest(boot_information);
} else if (eql(case, "initial-ramdisk")) { } else if (eql(case, "initial-ramdisk")) {
initialRamdiskTest(boot_information); initialRamdiskTest(boot_information);
} else if (eql(case, "vfs")) { } else if (eql(case, "vfs")) {
@@ -1207,15 +1219,15 @@ fn userPfTest() void {
/// hand — address space, code page RO+X, stack page RW+NX — because the blob is a /// hand — address space, code page RO+X, stack page RW+NX — because the blob is a
/// raw code fragment, not an ELF `spawnProcess` could load. Returns false if any /// raw code fragment, not an ELF `spawnProcess` could load. Returns false if any
/// allocation fails. /// allocation fails.
fn spawnFaultingProcess() bool { fn spawnFaultingProcess() ?u32 {
const blob = process.pfBlob(); const blob = process.pfBlob();
const flags = sync.enter(); const flags = sync.enter();
defer sync.leave(flags); defer sync.leave(flags);
const aspace = architecture.createAddressSpace() orelse return false; const aspace = architecture.createAddressSpace() orelse return null;
const code_frame = pmm.alloc() orelse { const code_frame = pmm.alloc() orelse {
architecture.destroyAddressSpace(aspace); architecture.destroyAddressSpace(aspace);
return false; return null;
}; };
// Fill through the physmap (the user mapping is read-only); pad with int3 so a // Fill through the physmap (the user mapping is read-only); pad with int3 so a
// stray jump traps instead of sliding. // stray jump traps instead of sliding.
@@ -1226,15 +1238,16 @@ fn spawnFaultingProcess() bool {
const stack_frame = pmm.alloc() orelse { const stack_frame = pmm.alloc() orelse {
architecture.destroyAddressSpace(aspace); // frees code_frame too — it's mapped architecture.destroyAddressSpace(aspace); // frees code_frame too — it's mapped
return false; return null;
}; };
architecture.mapUserPageInto(aspace, process.stack_base_virtual, stack_frame, true, false); // RW + NX architecture.mapUserPageInto(aspace, process.stack_base_virtual, stack_frame, true, false); // RW + NX
if (scheduler.spawnUserLocked(aspace, process.code_virtual, process.stack_base_virtual + abi.page_size, 4, "fault-probe", 0, null) == null) { // Supervised by the calling test task, so exitReasonOf can read the verdict.
const id = scheduler.spawnUserLocked(aspace, process.code_virtual, process.stack_base_virtual + abi.page_size, 4, "fault-probe", scheduler.currentId(), null) orelse {
architecture.destroyAddressSpace(aspace); architecture.destroyAddressSpace(aspace);
return false; return null;
} };
return true; return id;
} }
/// Fault recovery (docs/resilience.md step 2): a scheduled ring-3 process that /// Fault recovery (docs/resilience.md step 2): a scheduled ring-3 process that
@@ -1263,7 +1276,8 @@ fn faultRecoveryTest(boot_information: *const BootInformation) void {
scheduler.setPriority(4); scheduler.setPriority(4);
check("init heartbeat before the fault", process.write_count >= 1); check("init heartbeat before the fault", process.write_count >= 1);
check("faulting process spawned", spawnFaultingProcess()); const probe = spawnFaultingProcess() orelse 0;
check("faulting process spawned", probe != 0);
// The kill: the faulting process #PFs on its first instruction and the kernel // The kill: the faulting process #PFs on its first instruction and the kernel
// reaps it instead of halting. // reaps it instead of halting.
@@ -1272,6 +1286,7 @@ fn faultRecoveryTest(boot_information: *const BootInformation) void {
while (process.fault_kill_count < 1 and architecture.millis() < deadline) scheduler.yield(); while (process.fault_kill_count < 1 and architecture.millis() < deadline) scheduler.yield();
scheduler.setPriority(4); scheduler.setPriority(4);
check("faulting process was killed (not the machine)", process.fault_kill_count == 1); check("faulting process was killed (not the machine)", process.fault_kill_count == 1);
check("the probe's reason reads segmentation_fault", process.exitReasonOf(scheduler.currentId(), probe) == @intFromEnum(abi.ExitReason.segmentation_fault));
// Life after the kill: init must keep beating on the same core. // Life after the kill: init must keep beating on the same core.
const beats_at_kill = process.write_count; const beats_at_kill = process.write_count;
@@ -1423,6 +1438,12 @@ fn processKillTest(boot_information: *const BootInformation) void {
check("the sleeper's exit notification arrived (length 0)", r == 0); check("the sleeper's exit notification arrived (length 0)", r == 0);
check("its badge carries the exit bit and the child id", badge == abi.notify_badge_bit | abi.notify_exit_bit | sleeper); check("its badge carries the exit bit and the child id", badge == abi.notify_badge_bit | abi.notify_exit_bit | sleeper);
// M17.2: the recorded reason — the notification is the fence, so it is
// already readable, and gated by the same supervisor check as the kill.
check("the sleeper's reason reads killed", process.exitReasonOf(me, sleeper) == @intFromEnum(abi.ExitReason.killed));
check("a non-supervisor may not read the reason (-EPERM)", process.exitReasonOf(me + 12345, sleeper) == -ipcsync.EPERM);
check("an unknown id has no reason (-ESRCH)", process.exitReasonOf(me, 0xFFFF_FF00) == -ipcsync.ESRCH);
const beats_at_kill = process.write_count; const beats_at_kill = process.write_count;
scheduler.sleep(1500); // more than one heartbeat period scheduler.sleep(1500); // more than one heartbeat period
check("the heartbeat stopped with the kill", process.write_count == beats_at_kill); check("the heartbeat stopped with the kill", process.write_count == beats_at_kill);
@@ -1443,6 +1464,21 @@ fn processKillTest(boot_information: *const BootInformation) void {
check("the spinner's exit notification arrived (length 0)", r == 0); check("the spinner's exit notification arrived (length 0)", r == 0);
check("its badge carries the exit bit and the child id", badge == abi.notify_badge_bit | abi.notify_exit_bit | spinner); check("its badge carries the exit bit and the child id", badge == abi.notify_badge_bit | abi.notify_exit_bit | spinner);
// M17.2: a child that ends on its own must read exited, not killed —
// args-echo with arguments echoes once and returns from main.
var clean: u32 = 0;
i = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "args-echo")) continue;
clean = process.spawnProcessSupervised(item.blob, 4, &.{ "args-echo", "clean-exit" }, me, endpoint) catch 0;
break;
}
check("args-echo spawned as the clean-exit child", clean != 0);
r = ipcsync.replyWait(endpoint, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap);
check("the clean child's exit notification arrived", badge == abi.notify_badge_bit | abi.notify_exit_bit | clean);
check("the clean child's reason reads exited", process.exitReasonOf(me, clean) == @intFromEnum(abi.ExitReason.exited));
var table: [32]abi.ProcessDescriptor = undefined; var table: [32]abi.ProcessDescriptor = undefined;
const total = scheduler.enumerate(&table); const total = scheduler.enumerate(&table);
var still_listed = false; var still_listed = false;
@@ -1454,6 +1490,281 @@ fn processKillTest(boot_information: *const BootInformation) void {
result(); result();
} }
/// M17.1: a dead process's device claims are released by the reap, so a restarted
/// driver can claim its hardware again (docs/process-lifecycle.md iron rule 1).
/// First the broker release in isolation — two owners, one released, the other's
/// claim must survive. Then the death-path wiring with a real child: the claim is
/// made on the child's behalf (the broker is kernel-callable), the child is
/// killed, and once the exit notification arrives — posted last, after release —
/// the device must be unclaimed and claimable again.
fn claimReleaseTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: claim-release\n", .{});
var buffer: [2]device_abi.DeviceDescriptor = undefined;
const total = devices_broker.enumerate(&buffer);
check("the device tree is seeded (>= 2 devices)", total >= 2);
if (total < 2) {
result();
return;
}
// The broker release in isolation.
check("device 0 claimed by owner 111", devices_broker.claim(0, 111));
check("device 1 claimed by owner 222", devices_broker.claim(1, 222));
devices_broker.releaseAllOwnedBy(111);
check("owner 111's claim is released", devices_broker.ownerOf(0) == null);
check("owner 222's claim survives", (devices_broker.ownerOf(1) orelse 0) == 222);
devices_broker.releaseAllOwnedBy(222);
check("cleanup released owner 222", devices_broker.ownerOf(1) == null);
// The death-path wiring: a real process dies holding a claim.
check("bootloader handed over /system/services/init", boot_information.init_len != 0);
if (boot_information.init_len == 0) {
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len];
const me = scheduler.currentId();
const endpoint = ipcsync.createIpcEndpoint() orelse {
check("exit endpoint allocated", false);
result();
return;
};
const child = process.spawnProcessSupervised(image, 4, &.{"/system/services/init"}, me, endpoint) catch 0;
check("supervised child spawned", child != 0);
check("device 0 claimed on the child's behalf", devices_broker.claim(0, child));
check("the kill is accepted", process.killProcess(me, child) == 0);
var badge: u64 = 0;
var received_cap: u64 = 0;
_ = ipcsync.replyWait(endpoint, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap);
check("the exit notification arrived", badge == abi.notify_badge_bit | abi.notify_exit_bit | child);
check("death released the child's claim", devices_broker.ownerOf(0) == null);
check("the device is claimable again", devices_broker.claim(0, me));
devices_broker.releaseAllOwnedBy(me);
result();
}
/// M17.3: the published exit events, proven by their first subscriber. The VFS
/// subscribes at startup; a client opens a file and parks holding the handle;
/// the kill posts the exit event to the VFS's endpoint; the VFS releases the
/// dead client's handle and says so — the service-side mirror of iron rule 1
/// (a service must never depend on clients cleaning up after themselves).
fn vfsClientDeathTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: vfs-client-death\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.write_count = 0;
check("vfs spawned", spawnNamed(rd, "vfs"));
const me = scheduler.currentId();
const endpoint = ipcsync.createIpcEndpoint() orelse {
check("exit endpoint allocated", false);
result();
return;
};
var client: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "vfs-test")) continue;
client = process.spawnProcessSupervised(item.blob, 4, &.{ "vfs-test", "park" }, me, endpoint) catch 0;
break;
}
check("parked client spawned (supervised)", client != 0);
// Its heartbeat is the fence: once it beats, the handle is open.
const parked = "vfstest: parked";
scheduler.setPriority(1);
var deadline = architecture.millis() + 10000;
while (architecture.millis() < deadline) {
if (process.write_len >= parked.len and eql(process.write_buffer[0..parked.len], parked)) break;
scheduler.yield();
}
scheduler.setPriority(4);
check("client parked holding an open handle", process.write_len >= parked.len and eql(process.write_buffer[0..parked.len], parked));
check("the kill is accepted", process.killProcess(me, client) == 0);
var badge: u64 = 0;
var received_cap: u64 = 0;
_ = ipcsync.replyWait(endpoint, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap);
check("the exit notification arrived", badge == abi.notify_badge_bit | abi.notify_exit_bit | client);
// The VFS heard the same published event; its release line is the proof.
const released = "vfs: released 1 handle(s) for dead client";
scheduler.setPriority(1);
deadline = architecture.millis() + 10000;
while (architecture.millis() < deadline) {
if (process.write_len >= released.len and eql(process.write_buffer[0..released.len], released)) break;
scheduler.yield();
}
scheduler.setPriority(4);
check("the VFS released the dead client's handle", process.write_len >= released.len and eql(process.write_buffer[0..released.len], released));
result();
}
/// M17.4 from ring 3: process-test's signal-run role drives the whole lifecycle
/// surface — the zero-length ping (answered by the harness), signals as
/// statements (reload logged, terminate = clean exit), the one-shot timer, and
/// both endings of the stop sequence (polite -> exited, deaf -> killed at the
/// deadline). Its "process-test: signals ok" is the pass marker.
fn signalsTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: signals\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(image); // the parent system_spawns its children by name
process.write_count = 0;
var runner: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "process-test")) continue;
runner = process.spawnProcessSupervised(item.blob, 4, &.{ "process-test", "signal-run" }, scheduler.currentId(), null) catch 0;
break;
}
check("signal-run parent spawned", runner != 0);
const pass_marker = "process-test: signals ok";
const fail_marker = "process-test: FAIL";
scheduler.setPriority(1);
const deadline = architecture.millis() + 15000;
var saw_pass = false;
var saw_fail = false;
while (architecture.millis() < deadline and !saw_pass and !saw_fail) {
if (process.write_len >= pass_marker.len and eql(process.write_buffer[0..pass_marker.len], pass_marker)) saw_pass = true;
if (process.write_len >= fail_marker.len and eql(process.write_buffer[0..fail_marker.len], fail_marker)) saw_fail = true;
scheduler.yield();
}
scheduler.setPriority(4);
check("the signal-run parent reported ok", saw_pass and !saw_fail);
result();
}
/// M18.1: the device manager's restart machinery, end to end. In test-restart
/// mode the manager also supervises crash-test: a fixture that claims device 0,
/// hellos, and faults. The scenario asserts three markers in order — the real
/// xHCI driver hellos clean and stays; crash-test is restarted with backoff
/// (each respawn re-claiming the device the dead instance held, M17.1 through
/// the manager's path); the crash loop caps and the manager gives up.
fn driverRestartTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: driver-restart\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(image); // the manager system_spawns drivers by name
process.write_count = 0;
var manager: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{ "device-manager", "test-restart" }, scheduler.currentId(), null) catch 0;
break;
}
check("device-manager spawned in test-restart mode", manager != 0);
// The assertions live in the harness: its expect regex requires, in order,
// the xHCI hello ack, a crash-test restart, and the crash-loop cap — read
// from the whole serial capture, immune to the transient-line races a
// write_buffer poll would have here (many processes log concurrently).
result();
}
/// M18.2: bus tree reports, end to end. The manager (test-usb-restart mode)
/// spawns the xHCI driver; the driver maps its BAR, scans the root-hub ports,
/// and reports the two QEMU devices; the manager mirrors them, kills the
/// reporter (the test trigger), prunes both children, restarts the driver with
/// backoff, and the respawned instance re-claims, re-scans, and re-reports.
/// The harness's ordered expect regex is the assertion; this test only
/// orchestrates the spawn.
fn usbReportTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: usb-report\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(image);
var manager: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{ "device-manager", "test-usb-restart" }, scheduler.currentId(), null) catch 0;
break;
}
check("device-manager spawned in test-usb-restart mode", manager != 0);
result();
}
/// M18.3: the application surface. device-list enumerates the manager's tree
/// over IPC, subscribes with its endpoint as a capability, and prints every
/// published event; the manager's delayed test-kill of the reporter produces a
/// removed/added storm the subscriber must observe. The harness's ordered
/// expect regex is the assertion.
fn deviceListTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: device-list\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(image);
var manager: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{ "device-manager", "test-usb-restart" }, scheduler.currentId(), null) catch 0;
break;
}
check("device-manager spawned in test-usb-restart mode", manager != 0);
check("device-list spawned", spawnNamed(rd, "device-list"));
result();
}
/// The whole user-side surface at once: spawn process-test's supervisor role, /// The whole user-side surface at once: spawn process-test's supervisor role,
/// which — entirely from ring 3 — creates an exit endpoint, spawns its two /// which — entirely from ring 3 — creates an exit endpoint, spawns its two
/// children supervised, sees them in process_enumerate, kills them (one blocked, /// children supervised, sees them in process_enumerate, kills them (one blocked,
+5 -2
View File
@@ -16,8 +16,11 @@
pub const maximum_cpus = 128; pub const maximum_cpus = 128;
/// Maximum tasks (kernel threads) alive at once — the static task-table size. Each /// Maximum tasks (kernel threads) alive at once — the static task-table size. Each
/// online core consumes one slot for its idle task, plus task 0 on the BSP. /// online core consumes one slot for its idle task, plus task 0 on the BSP. Sized
pub const maximum_tasks = 16; /// for the initial-ramdisk sweep (15 bundled binaries spawned at once) plus the
/// device manager's supervised children with room to grow — at 16 the sweep
/// started failing spawns once the bundle passed a dozen binaries.
pub const maximum_tasks = 32;
/// Each task's kernel stack (also each AP's bring-up stack), in bytes. /// Each task's kernel stack (also each AP's bring-up stack), in bytes.
pub const kernel_stack_size = 16 * 1024; pub const kernel_stack_size = 16 * 1024;
+30
View File
@@ -0,0 +1,30 @@
//! /system/services/acpi — the ACPI discovery service: the x86 firmware
//! interpreter, moved out of ring 0 (docs/m19-m20-plan.md, M20). **Placeholder:
//! not implemented until M20.1** — it exists so the build's `-Ddiscovery`
//! option has both of its values and the ramdisk's neutral `discovery` slot is
//! wired before the implementation lands.
//!
//! What it becomes (the plan's decisions 5 and 7): claim the `acpi-tables`
//! node the kernel publishes (table blobs + the broad io_port grant + the SCI),
//! map the tables, and run the **shared AML module** in ring 3 behind a `Hal`
//! backed by `mmio_map` and `io_read`/`io_write` — the interpreter cannot tell
//! it moved. Then the bus-driver shape: `device_register` the namespace
//! devices (`_HID`, `_CRS` resources, containment against the node's
//! apertures), report each to the device manager, stay resident under its
//! supervision. M21 grows the event side on the same claim: the SCI, PM1 fixed
//! events, GPEs, Notify — published through the domain-named power protocol,
//! never an "ACPI events" protocol.
const runtime = @import("runtime");
pub fn main() void {
// Not implemented: exit cleanly and silently (a bare spawn by the
// initial-ramdisk sweep must not derange other tests' markers). The
// supervisor reads a clean exit as "meant to stop" — correct for a
// placeholder.
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
+44
View File
@@ -0,0 +1,44 @@
//! crash-test — a test fixture, not a driver: claims the device it is assigned,
//! hellos the device manager, announces itself, then faults on purpose. The
//! driver-restart scenario drives the manager's whole restart machinery with
//! it: fault → exit reason → backoff → respawn → the **same claim succeeding
//! again** (claim release on death, M17.1, through the manager's path) → the
//! crash-loop cap. Spawned bare (the initial-ramdisk sweep starts every bundled
//! binary), it exits silently so it cannot derange other tests.
const std = @import("std");
const runtime = @import("runtime");
const protocol = runtime.device_manager_protocol;
pub fn main(init: runtime.process.Init) void {
const argument = init.arguments.get(1) orelse return; // bare: stay silent
const assigned = std.fmt.parseInt(u64, argument, 10) catch return;
// The respawn only reaches this line because the kernel released the
// previous instance's claim at death. A failed claim exits cleanly — the
// manager reads "meant to stop" and the scenario fails loudly by silence.
if (!runtime.device.claim(assigned)) {
_ = runtime.system.write("crash-test: claim failed\n");
return;
}
var manager: ?runtime.ipc.Handle = null;
var tries: u32 = 0;
while (manager == null and tries < 100) : (tries += 1) {
manager = runtime.ipc.lookup(.device_manager);
if (manager == null) runtime.system.sleep(20);
}
const h = manager orelse return;
const hello = protocol.Hello{ .role = @intFromEnum(protocol.Role.device), .device_id = assigned };
var reply: [protocol.message_maximum]u8 = undefined;
_ = runtime.ipc.call(h, std.mem.asBytes(&hello), &reply) catch return;
_ = runtime.system.write("crash-test: faulting now\n");
const poison: *volatile u32 = @ptrFromInt(0xdead0000);
poison.* = 1; // the restart machinery's fuel: a real segmentation fault
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
@@ -0,0 +1,88 @@
//! device-list — the `ps` analog for the device tree (docs/device-manager.md
//! M18.3): asks the device manager for the tree over IPC, prints it, then
//! subscribes and prints every published add/remove event. The manager is the
//! one answer to "what devices exist" for user space; nothing here touches a
//! device_* system call.
const std = @import("std");
const runtime = @import("runtime");
const protocol = runtime.device_manager_protocol;
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [96]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
pub fn main() void {
var manager: ?runtime.ipc.Handle = null;
var tries: u32 = 0;
while (manager == null and tries < 200) : (tries += 1) {
manager = runtime.ipc.lookup(.device_manager);
if (manager == null) runtime.system.sleep(20);
}
const h = manager orelse {
_ = runtime.system.write("device-list: no device manager\n");
return;
};
// The snapshot — polled briefly, because at boot the bus drivers may still
// be scanning: an empty first answer usually just means "too early".
var reply: [protocol.message_maximum]u8 = undefined;
var count: u32 = 0;
var length: usize = 0;
tries = 0;
while (tries < 20) : (tries += 1) {
const request = protocol.Enumerate{};
length = runtime.ipc.call(h, std.mem.asBytes(&request), &reply) catch 0;
if (length >= @sizeOf(protocol.EnumerateReply)) {
count = std.mem.bytesToValue(protocol.EnumerateReply, reply[0..@sizeOf(protocol.EnumerateReply)]).count;
if (count != 0) break;
}
runtime.system.sleep(100);
}
writeLine("device-list: {d} devices\n", .{count});
var offset: usize = @sizeOf(protocol.EnumerateReply);
var index: u32 = 0;
while (index < count and offset + @sizeOf(protocol.ChildEntry) <= length) : (index += 1) {
const entry = std.mem.bytesToValue(protocol.ChildEntry, reply[offset..][0..@sizeOf(protocol.ChildEntry)]);
writeLine("device-list: device {d} port {d} identity {d}\n", .{ entry.parent, entry.bus_address, entry.identity });
offset += @sizeOf(protocol.ChildEntry);
}
// The subscription: our endpoint rides as the call's capability; events
// arrive as buffered messages carrying the same structs the bus sends.
const endpoint = runtime.ipc.createIpcEndpoint() orelse {
_ = runtime.system.write("device-list: no endpoint\n");
return;
};
const subscribe = protocol.Subscribe{};
_ = runtime.ipc.callCap(h, std.mem.asBytes(&subscribe), &reply, endpoint) catch {
_ = runtime.system.write("device-list: subscribe failed\n");
return;
};
_ = runtime.system.write("device-list: subscribed\n");
var receive: [protocol.message_maximum]u8 = undefined;
while (true) {
const got = runtime.ipc.replyWait(endpoint, &.{}, &receive, null);
if (!got.isMessage() or got.len < 1) continue;
switch (receive[0]) {
@intFromEnum(protocol.Operation.child_added) => {
if (got.len < protocol.child_added_size) continue;
const event = std.mem.bytesToValue(protocol.ChildAdded, receive[0..protocol.child_added_size]);
writeLine("device-list: added (device {d} port {d})\n", .{ event.parent, event.bus_address });
},
@intFromEnum(protocol.Operation.child_removed) => {
if (got.len < protocol.child_removed_size) continue;
const event = std.mem.bytesToValue(protocol.ChildRemoved, receive[0..protocol.child_removed_size]);
writeLine("device-list: removed (device {d} port {d})\n", .{ event.parent, event.bus_address });
},
else => {},
}
}
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
@@ -0,0 +1,139 @@
//! The device-manager protocol (docs/device-manager.md): what drivers and
//! applications say to the device manager over its well-known endpoint. The
//! vfs-protocol pattern — extern-struct messages, a version in the handshake,
//! reserved fields — so both sides depend on the contract by name. Deliberately
//! contains nothing lifecycle-shaped: stopping, liveness (the zero-length ping),
//! and exit reasons are the universal vocabulary of
//! docs/process-lifecycle.md, not this protocol.
/// The protocol version a driver states in its hello. A manager that cannot
/// serve a driver's version refuses the hello, and the mismatch is loud at
/// startup instead of quiet corruption later.
pub const version: u16 = 1;
/// What kind of driver is talking (docs/driver-model.md's shapes).
pub const Role = enum(u8) {
/// Owns a controller and reports the devices behind it (`child_added`).
bus = 1,
/// Serves one device, reached through a bus's transfer protocol.
device = 2,
};
/// The message kinds.
pub const Operation = enum(u8) {
hello = 1,
child_added = 2,
child_removed = 3,
enumerate = 4,
subscribe = 5,
};
/// `Hello.device_id` for a driver that serves no enumerated device (a test
/// fixture, a synthetic source).
pub const no_device: u64 = ~@as(u64, 0);
/// The handshake, sent once by every driver the manager spawns — the manager's
/// one self-enforced deadline: spawned and silent past it means wrong binary,
/// wrong version, or wedged before main, and the stop sequence follows.
pub const Hello = extern struct {
operation: u8 = @intFromEnum(Operation.hello),
/// A Role value.
role: u8,
/// The protocol version this driver was built against (`version`).
version: u16 = version,
reserved: u32 = 0,
/// The device this driver was assigned (its argv[1]), or `no_device`.
device_id: u64,
};
pub const hello_size = @sizeOf(Hello);
/// The manager's answer to a hello. Nonzero status = refused (version mismatch,
/// unknown sender); a refused driver should exit cleanly.
pub const HelloReply = extern struct {
status: i32,
reserved: u32 = 0,
};
pub const reply_size = @sizeOf(HelloReply);
/// A bus driver reporting one device it discovered behind its controller
/// (docs/device-manager.md "the tree"). Identity is the bus's native language —
/// for USB a port-speed class; the (class, subclass, protocol) triple joins it
/// once control transfers exist (the USB track). The manager mirrors the child
/// into its tree; when the reporting driver dies, the manager prunes everything
/// it reported (the children describe protocol state that died with it) and the
/// restarted instance rediscovers and re-reports.
pub const ChildAdded = extern struct {
operation: u8 = @intFromEnum(Operation.child_added),
reserved0: u8 = 0,
reserved1: u16 = 0,
reserved2: u32 = 0,
/// The reporting driver's own device (the controller) — the child's parent.
parent: u64,
/// Where on the bus (for USB: the root port number, 1-based).
bus_address: u64,
/// Bus-specific identity (for USB: the PORTSC port-speed class).
identity: u64,
};
pub const child_added_size = @sizeOf(ChildAdded);
/// A bus driver reporting a device gone (hot-unplug). Not yet sent by any
/// driver — the port scan has no unplug interrupt — but the manager handles it;
/// death-pruning covers removal until hotplug lands.
pub const ChildRemoved = extern struct {
operation: u8 = @intFromEnum(Operation.child_removed),
reserved0: u8 = 0,
reserved1: u16 = 0,
reserved2: u32 = 0,
parent: u64,
bus_address: u64,
};
pub const child_removed_size = @sizeOf(ChildRemoved);
/// The manager's answer to a tree report.
pub const ReportReply = extern struct {
status: i32,
reserved: u32 = 0,
};
/// An application asking for the tree (M18.3): the reply is an EnumerateReply
/// header followed by `count` ChildEntry records.
pub const Enumerate = extern struct {
operation: u8 = @intFromEnum(Operation.enumerate),
reserved0: u8 = 0,
reserved1: u16 = 0,
reserved2: u32 = 0,
};
pub const EnumerateReply = extern struct {
status: i32,
/// ChildEntry records following this header.
count: u32,
};
pub const ChildEntry = extern struct {
parent: u64,
bus_address: u64,
identity: u64,
};
/// An application subscribing to published add/remove events (the input-service
/// pattern): the subscriber's endpoint rides as the call's **capability**, and
/// events arrive on it as buffered messages whose payload is the same
/// ChildAdded / ChildRemoved struct the bus drivers send — one encoding, both
/// directions.
pub const Subscribe = extern struct {
operation: u8 = @intFromEnum(Operation.subscribe),
reserved0: u8 = 0,
reserved1: u16 = 0,
reserved2: u32 = 0,
};
/// Upper bound on any message in this protocol — sizes the endpoint buffers.
/// Capped by the kernel's IPC MESSAGE_MAXIMUM (256): an EnumerateReply carries
/// up to ten ChildEntry records per call, plenty for the mirror's current
/// bounds; paging joins the protocol if a tree ever outgrows one message.
pub const message_maximum = 256;
+407 -41
View File
@@ -1,20 +1,24 @@
//! /system/services/device-manager — the ring-3 process that turns the device //! /system/services/device-manager — the ring-3 process that turns the device
//! tree into a running system. The kernel enumerates the hardware and enforces the //! tree into a running system: **the matcher and the supervisor**
//! claim capability (mechanism); this decides *which driver serves which device* //! (docs/device-manager.md). The kernel enumerates the hardware and enforces the
//! and, eventually, spawns it (policy). Keeping that split in user space is the //! claim capability (mechanism); this decides which driver serves which device,
//! whole point of the microkernel: the manager is an ordinary, restartable process //! spawns it, and keeps it alive (policy). Keeping that split in user space is
//! with no special privilege — it uses the same `device_*` system calls any process //! the whole point of the microkernel: the manager is an ordinary, restartable
//! could ([drivers.md](../../../docs/drivers.md), [driver-model.md]). //! process with no special privilege.
//! //!
//! Increment 2 (this file): enumerate /system/devices, *match* each device to a //! M18.1 (this increment): the manager is a harness service on the well-known
//! driver, and *spawn* it with `system_spawn` — the kernel loads the named binary //! `.device_manager` endpoint. Every driver is spawned **supervised** — exit
//! from the initial-ramdisk as a fresh ring-3 process. On QEMU this discovers the //! notifications land in the same loop as protocol messages. Drivers with an
//! HPET, decides `hpet` serves it, and brings that driver all the way up. (The //! assignment must `hello` within a deadline or be stopped; a driver that dies
//! kernel still auto-spawns the whole initial-ramdisk at boot; increment 3 removes //! is restarted with backoff, and a crash loop (three fast deaths) marks it
//! that redundancy so the manager is the sole owner of driver spawning.) //! failed instead of respawning forever. Exit reasons (M17.2) drive the
//! decision: a clean exit meant to stop; only faults and missed deadlines
//! restart. Tree reports (`child_added`) land in M18.2.
const std = @import("std"); const std = @import("std");
const runtime = @import("runtime"); const runtime = @import("runtime");
const acpi_ids = @import("acpi-ids"); const acpi_ids = @import("acpi-ids");
const protocol = runtime.device_manager_protocol;
const device = runtime.device; const device = runtime.device;
const system = runtime.system; const system = runtime.system;
@@ -27,9 +31,8 @@ fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
} }
/// The driver that serves each device — the policy table. In a fuller system /// The driver that serves each device — the policy table. In a fuller system
/// this comes from the drivers describing what they bind (or a manifest under /// this comes from a manifest (docs/device-manager.md: the third bus type
/// /system/drivers); for now it is a small static map, which is enough to prove the /// triggers it); for now a static map. `null` = no driver for this class yet.
/// manager reads the tree and decides. `null` = no driver for this class yet.
fn driverFor(d: device.DeviceDescriptor) ?[]const u8 { fn driverFor(d: device.DeviceDescriptor) ?[]const u8 {
// detect device via DeviceClass // detect device via DeviceClass
if (d.class == @intFromEnum(device.DeviceClass.timer)) return "hpet"; if (d.class == @intFromEnum(device.DeviceClass.timer)) return "hpet";
@@ -47,10 +50,9 @@ fn driverFor(d: device.DeviceDescriptor) ?[]const u8 {
/// pci-class.zig decodes. /// pci-class.zig decodes.
const xhci_pci_class: u64 = 0x0C_03_30; const xhci_pci_class: u64 = 0x0C_03_30;
/// The bus driver that serves a PCI function, or null. Unlike the singleton drivers /// The bus driver that serves a PCI function, or null. A machine can carry
/// in `driverFor`, a machine can carry several identical controllers — so the caller /// several identical controllers — one driver instance per device, the id as
/// spawns one driver instance *per device*, passing the device id as argv[1] for the /// argv[1]. These drivers speak the protocol: a hello is expected.
/// instance to claim.
fn pciDriverFor(d: device.DeviceDescriptor) ?[]const u8 { fn pciDriverFor(d: device.DeviceDescriptor) ?[]const u8 {
if (d.class != @intFromEnum(device.DeviceClass.pci_device)) return null; if (d.class != @intFromEnum(device.DeviceClass.pci_device)) return null;
return switch (d.pci_class) { return switch (d.pci_class) {
@@ -59,24 +61,254 @@ fn pciDriverFor(d: device.DeviceDescriptor) ?[]const u8 {
}; };
} }
/// Spawn one instance of `driver_name` to serve the specific device `id` — the id // --- supervision -------------------------------------------------------------
/// arrives as argv[1]. No isProcessRunning gate here: the name alone cannot tell two
/// instances apart, and this manager is the sole spawner of drivers. /// How long a protocol driver has to hello after its spawn.
fn spawnForDevice(driver_name: []const u8, id: u64) void { const hello_deadline_ms: u64 = 3000;
var text: [20]u8 = undefined; /// Deaths faster than this count toward the crash loop; slower ones reset it.
const id_text = std.fmt.bufPrint(&text, "{d}", .{id}) catch return; const fast_death_ns: u64 = 2_000_000_000;
if (system.spawnWithArguments(driver_name, &.{id_text}) != null) { /// Consecutive fast deaths before the manager gives up on a driver.
writeLine("device-manager: spawned {s} for device {d}\n", .{ driver_name, id }); const crash_loop_cap: u32 = 3;
} else { /// Restart backoff: base << (restarts - 1), so 300 ms, 600 ms, 1200 ms.
writeLine("device-manager: failed to spawn {s} for device {d}\n", .{ driver_name, id }); const backoff_base_ms: u64 = 300;
const DriverState = enum {
awaiting_hello, // spawned; the deadline is armed (protocol drivers only)
running,
restarting, // dead; respawn due at restart_due_ns
stopped, // exited cleanly — it meant to; not restarted
failed, // crash loop, or unspawnable; the manager gave up
};
const Driver = struct {
used: bool = false,
name_buffer: [24]u8 = undefined,
name_len: usize = 0,
// The assigned device id (becomes argv[1]), or protocol.no_device.
device_id: u64 = protocol.no_device,
// Whether this driver speaks the protocol (hello expected, deadline
// enforced). Legacy drivers (hpet, ps2-bus) are supervised and restarted
// but not yet required to hello.
speaks_protocol: bool = false,
process_id: u32 = 0,
state: DriverState = .running,
restarts: u32 = 0,
spawn_ns: u64 = 0,
hello_deadline_ns: u64 = 0,
restart_due_ns: u64 = 0,
fn name(driver: *const Driver) []const u8 {
return driver.name_buffer[0..driver.name_len];
}
};
const maximum_drivers = 16;
var drivers: [maximum_drivers]Driver = .{Driver{}} ** maximum_drivers;
var manager_endpoint: runtime.ipc.Handle = 0;
var test_restart_mode = false;
var test_usb_restart_mode = false;
var test_usb_killed = false;
var test_kill_pid: u32 = 0;
var test_kill_due_ns: u64 = 0;
/// The application subscribers (M18.3, the input-service pattern): endpoints
/// handed over as capabilities, each receiving every child add/remove as a
/// buffered message. A subscriber whose endpoint stops accepting (it died) is
/// dropped on the failed send.
const maximum_subscribers = 8;
var subscribers: [maximum_subscribers]?runtime.ipc.Handle = .{null} ** maximum_subscribers;
/// Publish one event (a ChildAdded or ChildRemoved struct, the same encoding
/// the bus drivers send) to every subscriber.
fn publishEvent(event: []const u8) void {
for (&subscribers) |*slot| {
if (slot.*) |handle| {
if (!runtime.ipc.send(handle, event)) slot.* = null; // dead subscriber
}
} }
} }
pub fn main() void { /// The manager's mirror of what bus drivers report (docs/device-manager.md "the
/// tree"): the children, keyed by (parent, bus address), each remembering which
/// driver instance reported it — that is what death-pruning sweeps by.
const Child = struct {
used: bool = false,
parent: u64 = 0,
bus_address: u64 = 0,
identity: u64 = 0,
reporter: u32 = 0, // the reporting driver instance's process id
};
const maximum_children = 32;
var children: [maximum_children]Child = .{Child{}} ** maximum_children;
/// Record (or refresh) a reported child. Refreshing matters: a restarted bus
/// driver re-reports what it rediscovers, and the same (parent, port) must not
/// duplicate.
fn addChild(parent: u64, bus_address: u64, identity: u64, reporter: u32) bool {
var free: ?*Child = null;
for (&children) |*child| {
if (child.used and child.parent == parent and child.bus_address == bus_address) {
child.identity = identity;
child.reporter = reporter;
return true;
}
if (!child.used and free == null) free = child;
}
const slot = free orelse return false;
slot.* = .{ .used = true, .parent = parent, .bus_address = bus_address, .identity = identity, .reporter = reporter };
return true;
}
/// Prune every child a dead driver instance reported: the children describe
/// protocol state (slots, rings) that died with the process — keeping the nodes
/// would be keeping a lie. The restarted instance rediscovers and re-reports.
/// Watchers hear the honest story: removed now, added again on rediscovery.
fn pruneChildrenOf(reporter: u32) void {
for (&children) |*child| {
if (child.used and child.reporter == reporter) {
writeLine("device-manager: child removed (device {d} port {d})\n", .{ child.parent, child.bus_address });
child.used = false;
const event = protocol.ChildRemoved{ .parent = child.parent, .bus_address = child.bus_address };
publishEvent(std.mem.asBytes(&event));
}
}
}
/// How many children a driver instance has reported (the test-usb-restart
/// trigger counts these).
fn childCountOf(reporter: u32) u32 {
var n: u32 = 0;
for (&children) |*child| {
if (child.used and child.reporter == reporter) n += 1;
}
return n;
}
fn driverByProcess(process_id: u32) ?*Driver {
for (&drivers) |*driver| {
if (driver.used and driver.process_id == process_id) return driver;
}
return null;
}
/// Whether a singleton driver is already in the table (two ACPI nodes can both
/// map to ps2-bus; one instance serves both).
fn alreadySupervised(name: []const u8) bool {
for (&drivers) |*driver| {
if (driver.used and std.mem.eql(u8, driver.name(), name)) return true;
}
return false;
}
/// Record a driver in the table and spawn its first instance.
fn addDriver(name: []const u8, device_id: u64, speaks_protocol: bool) void {
for (&drivers) |*driver| {
if (driver.used) continue;
const n = @min(name.len, driver.name_buffer.len);
@memcpy(driver.name_buffer[0..n], name[0..n]);
driver.name_len = n;
driver.device_id = device_id;
driver.speaks_protocol = speaks_protocol;
driver.used = true;
spawnDriver(driver);
return;
}
writeLine("device-manager: driver table full; cannot supervise {s}\n", .{name});
}
/// (Re)spawn a driver instance: supervised on the manager's own endpoint, the
/// device id as argv[1] when it has one, the hello deadline armed when it
/// speaks the protocol.
fn spawnDriver(driver: *Driver) void {
var id_text: [20]u8 = undefined;
var arguments: [1][]const u8 = undefined;
var argument_count: usize = 0;
if (driver.device_id != protocol.no_device) {
arguments[0] = std.fmt.bufPrint(&id_text, "{d}", .{driver.device_id}) catch return;
argument_count = 1;
}
const child = system.spawnSupervised(driver.name(), arguments[0..argument_count], manager_endpoint) orelse {
writeLine("device-manager: failed to spawn {s}\n", .{driver.name()});
driver.state = .failed;
return;
};
driver.process_id = child;
driver.spawn_ns = system.clock();
if (driver.speaks_protocol) {
driver.state = .awaiting_hello;
driver.hello_deadline_ns = driver.spawn_ns + hello_deadline_ms * 1_000_000;
_ = system.timerOnce(manager_endpoint, hello_deadline_ms + 100);
} else {
driver.state = .running;
}
if (driver.device_id != protocol.no_device) {
writeLine("device-manager: spawned {s} for device {d}\n", .{ driver.name(), driver.device_id });
} else {
writeLine("device-manager: spawned {s}\n", .{driver.name()});
}
}
/// A driver died. Prune what it reported first — then the exit reason (M17.2)
/// is the whole restart decision: a clean exit meant to stop; anything else
/// restarts with backoff until the crash-loop cap.
fn onDriverExit(driver: *Driver) void {
pruneChildrenOf(driver.process_id);
const reason = runtime.process.exitReason(driver.process_id) orelse .fault;
if (reason == .exited) {
driver.state = .stopped;
writeLine("device-manager: {s} exited cleanly; not restarting\n", .{driver.name()});
return;
}
const now = system.clock();
const alive_ns = now - driver.spawn_ns;
driver.restarts = if (alive_ns < fast_death_ns) driver.restarts + 1 else 1;
if (driver.restarts >= crash_loop_cap) {
driver.state = .failed;
writeLine("device-manager: {s} is failing repeatedly (crash loop); giving up\n", .{driver.name()});
return;
}
const delay_ms = backoff_base_ms << @intCast(driver.restarts - 1);
driver.state = .restarting;
driver.restart_due_ns = now + delay_ms * 1_000_000;
writeLine("device-manager: restarting {s} in {d} ms (died: {s})\n", .{ driver.name(), delay_ms, @tagName(reason) });
_ = system.timerOnce(manager_endpoint, delay_ms + 50);
}
/// A timer landed: sweep every deadline. Overdue hellos are killed (the exit
/// notification then routes through the normal restart policy); due restarts
/// respawn. Timers carry no id on purpose — the table is the state, and one
/// sweep serves every armed deadline.
fn sweepDeadlines() void {
const now = system.clock();
if (test_kill_pid != 0 and now >= test_kill_due_ns) {
writeLine("device-manager: test mode: killing the reporter\n", .{});
_ = system.kill(test_kill_pid);
test_kill_pid = 0;
}
for (&drivers) |*driver| {
if (!driver.used) continue;
switch (driver.state) {
.awaiting_hello => if (now >= driver.hello_deadline_ns) {
writeLine("device-manager: {s} missed its hello deadline\n", .{driver.name()});
_ = system.kill(driver.process_id);
// The exit notification finishes the job via onDriverExit.
},
.restarting => if (now >= driver.restart_due_ns) spawnDriver(driver),
else => {},
}
}
}
// --- the harness callbacks -----------------------------------------------------
fn initialise(endpoint: runtime.ipc.Handle) bool {
manager_endpoint = endpoint;
// Enumerate into a heap buffer (too big for the one-page user stack). // Enumerate into a heap buffer (too big for the one-page user stack).
const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch { const buffer = runtime.allocator().alloc(device.DeviceDescriptor, 64) catch {
_ = runtime.system.write("device-manager: out of memory\n"); _ = runtime.system.write("device-manager: out of memory\n");
return; return false;
}; };
const total = device.enumerate(buffer); const total = device.enumerate(buffer);
const count = @min(total, buffer.len); const count = @min(total, buffer.len);
@@ -85,28 +317,162 @@ pub fn main() void {
for (buffer[0..count]) |descriptor| { for (buffer[0..count]) |descriptor| {
if (pciDriverFor(descriptor)) |driver_name| { if (pciDriverFor(descriptor)) |driver_name| {
matched += 1; matched += 1;
spawnForDevice(driver_name, descriptor.id); addDriver(driver_name, descriptor.id, true);
continue; continue;
} }
const driver_name = driverFor(descriptor) orelse continue; const driver_name = driverFor(descriptor) orelse continue;
matched += 1; matched += 1;
if (!system.isProcessRunning(driver_name)) { // Skip a singleton that is already alive (the initial-ramdisk sweep test
if (runtime.system.spawn(driver_name) != null) { // starts every bundled binary bare, this manager included) — spawning a
writeLine("device-manager: spawned {s}\n", .{driver_name}); // second instance would only lose the claim race and churn the log.
} else { if (!alreadySupervised(driver_name) and !system.isProcessRunning(driver_name)) {
writeLine("device-manager: failed to spawn {s}\n", .{driver_name}); addDriver(driver_name, protocol.no_device, false);
} }
} else {
writeLine("device-manager: already spawned {s}\n", .{driver_name});
} }
if (test_restart_mode) {
// The driver-restart scenario's fixture: claims device 0 (the tree
// root, otherwise unclaimed), hellos, then faults — driving backoff,
// re-claim-after-death, and the crash-loop cap deterministically.
addDriver("crash-test", 0, true);
} }
if (matched == 0) { if (matched == 0) {
_ = runtime.system.write("device-manager: no matchable devices\n"); _ = runtime.system.write("device-manager: no matchable devices\n");
} else {
_ = runtime.system.write("device-manager: ok\n");
}
return true;
}
fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
if (message.len < 1) return 0;
switch (message[0]) {
@intFromEnum(protocol.Operation.child_added) => return onChildAdded(message, reply, sender),
@intFromEnum(protocol.Operation.child_removed) => return onChildRemoved(message, reply, sender),
@intFromEnum(protocol.Operation.enumerate) => return onEnumerate(reply),
@intFromEnum(protocol.Operation.subscribe) => return onSubscribe(reply, capability),
@intFromEnum(protocol.Operation.hello) => {},
else => return 0,
}
if (message.len < protocol.hello_size) return 0;
const hello = std.mem.bytesToValue(protocol.Hello, message[0..protocol.hello_size]);
var status: i32 = 0;
if (hello.version != protocol.version) {
status = -1;
writeLine("device-manager: refused hello (version {d}) from process {d}\n", .{ hello.version, sender });
} else if (driverByProcess(sender)) |driver| {
driver.state = .running;
writeLine("device-manager: hello from {s} (device {d})\n", .{ driver.name(), hello.device_id });
} else {
status = -1;
writeLine("device-manager: hello from unknown process {d}\n", .{sender});
}
const hello_reply = protocol.HelloReply{ .status = status };
@memcpy(reply[0..protocol.reply_size], std.mem.asBytes(&hello_reply));
return protocol.reply_size;
}
/// A bus driver reported a discovered device: mirror it, and in
/// test-usb-restart mode kill the reporter once after its second child — the
/// deterministic trigger for prune -> backoff -> respawn -> re-report.
fn onChildAdded(message: []const u8, reply: []u8, sender: u32) usize {
if (message.len < protocol.child_added_size) return 0;
const report = std.mem.bytesToValue(protocol.ChildAdded, message[0..protocol.child_added_size]);
var status: i32 = 0;
if (driverByProcess(sender)) |driver| {
if (!addChild(report.parent, report.bus_address, report.identity, sender)) status = -1;
writeLine("device-manager: child added (device {d} port {d}, identity {d}) by {s}\n", .{ report.parent, report.bus_address, report.identity, driver.name() });
if (status == 0) publishEvent(message[0..protocol.child_added_size]);
} else {
status = -1;
}
const report_reply = protocol.ReportReply{ .status = status };
@memcpy(reply[0..@sizeOf(protocol.ReportReply)], std.mem.asBytes(&report_reply));
if (test_usb_restart_mode and !test_usb_killed and childCountOf(sender) >= 2) {
// Delayed, not immediate: the device-list scenario's subscriber needs a
// window to enumerate and subscribe before the events start.
test_usb_killed = true;
test_kill_pid = sender;
test_kill_due_ns = system.clock() + 2_000_000_000;
_ = system.timerOnce(manager_endpoint, 2100);
}
return @sizeOf(protocol.ReportReply);
}
/// A bus driver reported a device gone (hot-unplug; no sender exists yet, but
/// the handler is protocol-complete — death-pruning covers removal until then).
fn onChildRemoved(message: []const u8, reply: []u8, sender: u32) usize {
if (message.len < protocol.child_removed_size) return 0;
const report = std.mem.bytesToValue(protocol.ChildRemoved, message[0..protocol.child_removed_size]);
var status: i32 = -1;
for (&children) |*child| {
if (child.used and child.parent == report.parent and child.bus_address == report.bus_address and child.reporter == sender) {
writeLine("device-manager: child removed (device {d} port {d})\n", .{ child.parent, child.bus_address });
child.used = false;
status = 0;
}
}
const report_reply = protocol.ReportReply{ .status = status };
@memcpy(reply[0..@sizeOf(protocol.ReportReply)], std.mem.asBytes(&report_reply));
return @sizeOf(protocol.ReportReply);
}
/// An application asked for the tree: the mirror, as a header plus entries.
fn onEnumerate(reply: []u8) usize {
var count: u32 = 0;
var offset: usize = @sizeOf(protocol.EnumerateReply);
for (&children) |*child| {
if (!child.used) continue;
if (offset + @sizeOf(protocol.ChildEntry) > reply.len) break;
const entry = protocol.ChildEntry{ .parent = child.parent, .bus_address = child.bus_address, .identity = child.identity };
@memcpy(reply[offset..][0..@sizeOf(protocol.ChildEntry)], std.mem.asBytes(&entry));
offset += @sizeOf(protocol.ChildEntry);
count += 1;
}
const header = protocol.EnumerateReply{ .status = 0, .count = count };
@memcpy(reply[0..@sizeOf(protocol.EnumerateReply)], std.mem.asBytes(&header));
return offset;
}
/// An application subscribed: its endpoint arrived as the call's capability.
fn onSubscribe(reply: []u8, capability: ?runtime.ipc.Handle) usize {
var status: i32 = -1;
if (capability) |handle| {
for (&subscribers) |*slot| {
if (slot.* == null) {
slot.* = handle;
status = 0;
break;
}
}
}
const report_reply = protocol.ReportReply{ .status = status };
@memcpy(reply[0..@sizeOf(protocol.ReportReply)], std.mem.asBytes(&report_reply));
return @sizeOf(protocol.ReportReply);
}
fn onNotification(badge: u64) void {
if (badge & runtime.ipc.notify_exit_bit != 0) {
const dead: u32 = @intCast(badge & ~(runtime.ipc.notify_badge_bit | runtime.ipc.notify_exit_bit));
if (driverByProcess(dead)) |driver| onDriverExit(driver);
return; return;
} }
_ = runtime.system.write("device-manager: ok\n"); if (badge & runtime.ipc.notify_timer_bit != 0) sweepDeadlines();
while (true) runtime.system.sleep(1000); }
pub fn main(init: runtime.process.Init) void {
if (init.arguments.get(1)) |mode| {
test_restart_mode = std.mem.eql(u8, mode, "test-restart");
test_usb_restart_mode = std.mem.eql(u8, mode, "test-usb-restart");
}
runtime.service.run(protocol.message_maximum, .{
.service = .device_manager,
.init = initialise,
.on_message = onMessage,
.on_notification = onNotification,
});
} }
pub const panic = runtime.panic; pub const panic = runtime.panic;
+34
View File
@@ -0,0 +1,34 @@
//! /system/services/fdt — the devicetree discovery service: the ARM twin of the
//! acpi service (docs/m19-m20-plan.md decision 7). **Placeholder: not
//! implemented.** It exists so the build's `-Ddiscovery` option has both of its
//! values from day one; the implementation lands with the Raspberry Pi
//! bring-up (docs/arm.md).
//!
//! What it becomes: the per-firmware discoverer for boots that hand over a
//! flattened device tree instead of ACPI tables. It claims the
//! `devicetree-blob` node the kernel publishes (the FDT the loader received),
//! walks the tree — pure data, no bytecode, so unlike the acpi service it
//! needs no port grant and no interpreter — and, like any bus-shaped driver:
//! `device_register`s what it finds (containment against the blob node's
//! recorded apertures), reports each child to the device manager
//! (`child_added`, identity = the node's `compatible` string), and stays
//! resident under the manager's supervision (hello, restart, the usual
//! contract).
//!
//! Known prerequisite recorded in the plan: `DeviceDescriptor`'s 8-byte `hid`
//! cannot hold an FDT `compatible` string ("brcm,bcm2835-aux-uart") — identity
//! widens before this file grows a body.
const runtime = @import("runtime");
pub fn main() void {
// Not implemented: exit cleanly and silently (a bare spawn by the
// initial-ramdisk sweep must not derange other tests' markers). The
// supervisor reads a clean exit as "meant to stop" — correct for a
// placeholder.
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
@@ -48,11 +48,92 @@ fn awaitChildExit(endpoint: runtime.ipc.Handle) u32 {
return received.childProcessId(); return received.childProcessId();
} }
/// The harness-run child of the signals test: echoes requests, logs the two
/// signals it handles. Terminate makes run() return, and returning from main is
/// the clean exit the parent reads as ExitReason.exited.
fn echo(message: []const u8, reply: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = sender;
_ = capability;
const n = @min(message.len, reply.len);
@memcpy(reply[0..n], message[0..n]);
return n;
}
fn onReload() void {
_ = runtime.system.write("process-test: reloaded\n");
}
fn onTerminate() void {
_ = runtime.system.write("process-test: terminating\n");
}
/// The parent of the signals test: drives ping, echo, reload, the one-shot
/// timer, and both endings of the stop sequence (polite -> exited; deaf ->
/// killed at the deadline). Prints "process-test: signals ok" as the marker.
fn signalRun() void {
const endpoint = runtime.ipc.createIpcEndpoint() orelse fail("create exit endpoint");
const child = runtime.system.spawnSupervised("process-test", &.{"service"}, endpoint) orelse fail("spawn service child");
// Reach the child's endpoint through the registry (retry: it may not be up).
var service_handle: ?runtime.ipc.Handle = null;
var tries: u32 = 0;
while (service_handle == null and tries < 200) : (tries += 1) {
service_handle = runtime.ipc.lookup(.input);
if (service_handle == null) runtime.system.sleep(20);
}
const h = service_handle orelse fail("service child never registered");
// The universal ping: a zero-length call answered zero-length by the harness.
var reply: [16]u8 = undefined;
const pong = runtime.ipc.call(h, &.{}, &reply) catch fail("ping call failed");
if (pong != 0) fail("ping reply not empty");
// An ordinary request still reaches on_message.
const n = runtime.ipc.call(h, "echo!", &reply) catch fail("echo call failed");
if (n != 5 or !std.mem.eql(u8, reply[0..5], "echo!")) fail("echo mismatch");
// reload: a statement — the child logs it; the kernel test reads the serial.
if (!runtime.process.sendSignal(child, .reload)) fail("send reload");
runtime.system.sleep(200);
// The one-shot timer: armed on our endpoint, lands as isTimer.
if (!runtime.system.timerOnce(endpoint, 100)) fail("arm timer");
var scratch: [8]u8 = undefined;
const landing = runtime.ipc.replyWait(endpoint, scratch[0..0], &scratch, null);
if (!landing.isTimer()) fail("expected the timer landing");
// The stop sequence, polite path: terminate, clean exit inside the deadline.
runtime.process.stop(child, 2000, endpoint);
if ((runtime.process.exitReason(child) orelse .killed) != .exited) fail("service child reason not exited");
// The deaf child: binds nothing, hears nothing — the deadline kills it.
const deaf = runtime.system.spawnSupervised("process-test", &.{"sleeper"}, endpoint) orelse fail("spawn deaf child");
runtime.system.sleep(50); // let it reach its sleep
runtime.process.stop(deaf, 300, endpoint);
if ((runtime.process.exitReason(deaf) orelse .exited) != .killed) fail("deaf child reason not killed");
_ = runtime.system.write("process-test: signals ok\n");
}
pub fn main(init: runtime.process.Init) void { pub fn main(init: runtime.process.Init) void {
const role = init.arguments.get(1) orelse return; // spawned bare (ramdisk sweep): stay silent const role = init.arguments.get(1) orelse return; // spawned bare (ramdisk sweep): stay silent
if (std.mem.eql(u8, role, "sleeper")) { if (std.mem.eql(u8, role, "sleeper")) {
while (true) runtime.system.sleep(500); while (true) runtime.system.sleep(500);
} }
if (std.mem.eql(u8, role, "service")) {
// Borrowed well-known id: the input service is not part of this scenario.
runtime.service.run(64, .{
.service = .input,
.on_message = echo,
.on_reload = onReload,
.on_terminate = onTerminate,
});
return; // terminate arrived; returning is the clean exit
}
if (std.mem.eql(u8, role, "signal-run")) {
signalRun();
return;
}
if (std.mem.eql(u8, role, "spinner")) { if (std.mem.eql(u8, role, "spinner")) {
var beat: u64 = 0; var beat: u64 = 0;
const touch: *volatile u64 = &beat; const touch: *volatile u64 = &beat;
@@ -89,6 +170,12 @@ pub fn main(init: runtime.process.Init) void {
if (listed(sleeper, "process-test")) fail("sleeper still listed after kill"); if (listed(sleeper, "process-test")) fail("sleeper still listed after kill");
if (listed(spinner, "process-test")) fail("spinner still listed after kill"); if (listed(spinner, "process-test")) fail("spinner still listed after kill");
// M17.2: both children were killed by us, and the reason says so — the whole
// restart-policy input, read through the runtime like a real supervisor would.
if ((runtime.process.exitReason(sleeper) orelse .exited) != .killed) fail("sleeper reason not killed");
if ((runtime.process.exitReason(spinner) orelse .exited) != .killed) fail("spinner reason not killed");
if (runtime.process.exitReason(0xFFFF_FFF0) != null) fail("unknown id had a reason");
_ = runtime.system.write("process-test: ok\n"); _ = runtime.system.write("process-test: ok\n");
} }
+21 -1
View File
@@ -6,10 +6,30 @@
const std = @import("std"); const std = @import("std");
const runtime = @import("runtime"); const runtime = @import("runtime");
pub fn main() void { pub fn main(init: runtime.process.Init) void {
const u = @import("posix").unistd; const u = @import("posix").unistd;
const payload = "hello-vfs"; const payload = "hello-vfs";
// The "park" role (the vfs-client-death test): open a file, then hold the
// handle forever without closing — the kill and the VFS's release-on-death
// are the point.
if (init.arguments.count > 1) {
var fd: i32 = -1;
var tries: u32 = 0;
while (fd < 0 and tries < 200) : (tries += 1) {
fd = u.open("parked", u.O_CREAT);
if (fd < 0) runtime.system.sleep(20);
}
if (fd < 0) {
_ = runtime.system.write("vfstest: park open failed\n");
return;
}
while (true) {
_ = runtime.system.write("vfstest: parked\n");
runtime.system.sleep(500);
}
}
// The VFS server may not have registered yet — retry open until it's up. // The VFS server may not have registered yet — retry open until it's up.
var fd: i32 = -1; var fd: i32 = -1;
var tries: u32 = 0; var tries: u32 = 0;
+53 -18
View File
@@ -23,6 +23,10 @@ const Node = struct {
const OpenFile = struct { const OpenFile = struct {
used: bool = false, used: bool = false,
node: usize = 0, node: usize = 0,
// The client (task id — an IPC badge is one) that opened this handle. What
// release-on-death sweeps by: a service must never depend on its clients
// cleaning up after themselves (docs/process-lifecycle.md).
owner: u32 = 0,
}; };
var nodes = [_]Node{.{}} ** 8; var nodes = [_]Node{.{}} ** 8;
@@ -65,8 +69,30 @@ fn fail(out: []u8) usize {
return writeReply(out, .{ .status = -1 }, &.{}); return writeReply(out, .{ .status = -1 }, &.{});
} }
/// Handle one request; write the reply into `out`, return its length. /// Format one whole log line and emit it in a single `debug_write`, so lines from
fn handle(message: []const u8, out: []u8) usize { /// concurrent processes can never land in the middle of it.
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [96]u8 = undefined;
_ = runtime.system.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
/// Release every open handle `client` held — called on that client's published
/// exit event. The nodes (the files) stay: ramfs contents outlive their writers,
/// only the dead client's handles go.
fn releaseClientHandles(client: u32) void {
var released: u32 = 0;
for (&opens) |*o| {
if (o.used and o.owner == client) {
o.used = false;
released += 1;
}
}
if (released != 0) writeLine("vfs: released {d} handle(s) for dead client {d}\n", .{ released, client });
}
/// Handle one request from `sender`; write the reply into `out`, return its length.
fn handle(message: []const u8, out: []u8, sender: u32, capability: ?runtime.ipc.Handle) usize {
_ = capability;
if (message.len < protocol.request_size) return fail(out); if (message.len < protocol.request_size) return fail(out);
const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]); const request = std.mem.bytesToValue(protocol.Request, message[0..protocol.request_size]);
const payload = message[protocol.request_size..]; const payload = message[protocol.request_size..];
@@ -77,7 +103,7 @@ fn handle(message: []const u8, out: []u8) usize {
const ni = findNode(name) orelse createNode(name) orelse return fail(out); const ni = findNode(name) orelse createNode(name) orelse return fail(out);
for (&opens, 0..) |*o, i| { for (&opens, 0..) |*o, i| {
if (!o.used) { if (!o.used) {
o.* = .{ .used = true, .node = ni }; o.* = .{ .used = true, .node = ni, .owner = sender };
return writeReply(out, .{ .status = 0, .node = i }, &.{}); return writeReply(out, .{ .status = 0, .node = i }, &.{});
} }
} }
@@ -113,27 +139,36 @@ fn handle(message: []const u8, out: []u8) usize {
} }
} }
pub fn main() void { /// Startup, under the harness: subscribe to the published exit events — when a
const endpoint = runtime.ipc.createIpcEndpoint() orelse { /// client dies holding open handles, the exit notification is how the VFS learns
_ = runtime.system.write("vfs: no endpoint\n"); /// to release them (docs/process-lifecycle.md).
return; fn initialise(endpoint: runtime.ipc.Handle) bool {
}; if (!runtime.process.subscribeExits(endpoint)) {
if (!runtime.ipc.register(.vfs, endpoint)) { _ = runtime.system.write("vfs: exit subscription failed\n");
_ = runtime.system.write("vfs: register failed\n");
return;
} }
_ = runtime.system.write("vfs: ready\n"); _ = runtime.system.write("vfs: ready\n");
return true;
}
var reply_buffer: [protocol.message_maximum]u8 = undefined; /// A non-signal notification: the only kind the VFS subscribes to is exit events.
var reply_len: usize = 0; fn onNotification(badge: u64) void {
var receive: [protocol.message_maximum]u8 = undefined; if (badge & runtime.ipc.notify_exit_bit != 0) {
while (true) { releaseClientHandles(@intCast(badge & ~(runtime.ipc.notify_badge_bit | runtime.ipc.notify_exit_bit)));
const got = runtime.ipc.replyWait(endpoint, reply_buffer[0..reply_len], &receive, null);
// Ignore notifications (none expected here); handle a request.
reply_len = handle(receive[0..got.len], &reply_buffer);
} }
} }
pub fn main() void {
// The harness owns the loop: requests dispatch to handle(), exit events to
// onNotification(), ping and terminate are answered for free — this service
// gained the whole lifecycle contract by deleting its hand-rolled loop.
runtime.service.run(protocol.message_maximum, .{
.service = .vfs,
.init = initialise,
.on_message = handle,
.on_notification = onNotification,
});
}
pub const panic = runtime.panic; pub const panic = runtime.panic;
comptime { comptime {
_ = &runtime.start._start; _ = &runtime.start._start;
+62
View File
@@ -240,6 +240,68 @@ CASES = [
"smp": 4, "smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS", "expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# M17.1: a dead process's device claims are released by the reap — kill a child
# holding a claim, the device must be claimable again (process-lifecycle.md).
{"name": "claim-release",
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M17.3: published exit events — the VFS subscribes, a client dies holding an
# open handle, and the VFS releases it (process-lifecycle.md "Who learns of a death").
{"name": "vfs-client-death",
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M17.4: signals over IPC — ping, reload, terminate (clean exit), the one-shot
# timer, and the stop sequence's two endings, all driven from ring 3.
{"name": "signals",
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M18.2: bus tree reports — the xHCI driver scans its root-hub ports and
# reports both QEMU devices; the manager mirrors, prunes on the reporter's
# death, and the respawned driver re-reports (docs/device-manager.md).
{"name": "usb-report",
"smp": 4,
"timeout": 90,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"device-manager: child added[\s\S]*"
r"device-manager: child added[\s\S]*"
r"device-manager: test mode: killing the reporter[\s\S]*"
r"device-manager: child removed[\s\S]*"
r"device-manager: restarting usb-xhci-bus[\s\S]*"
r"device-manager: child added",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M18.3: the application surface — device-list enumerates the tree over IPC,
# subscribes (endpoint as capability), and observes the removed/added events
# the reporter's test-kill produces (docs/device-manager.md).
{"name": "device-list",
"smp": 4,
"timeout": 90,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"device-list: 2 devices[\s\S]*"
r"device-list: subscribed[\s\S]*"
r"device-manager: test mode: killing the reporter[\s\S]*"
r"device-list: removed \(device[\s\S]*"
r"device-list: added \(device",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# M18.1: the device manager's hello + restart policy — xHCI hellos clean and
# stays; crash-test faults, is restarted with backoff (re-claiming its device
# each time), and hits the crash-loop cap (docs/device-manager.md).
{"name": "driver-restart",
"smp": 4,
"timeout": 90,
"qemu_extra": ["-device", "qemu-xhci,id=xhci",
"-device", "usb-kbd,bus=xhci.0",
"-device", "usb-mouse,bus=xhci.0"],
"expect": r"usb-xhci-bus: hello acknowledged[\s\S]*"
r"device-manager: restarting crash-test[\s\S]*"
r"device-manager: crash-test is failing repeatedly",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# The initial_ramdisk: the loader ferries a bundle of user binaries; the kernel parses # The initial_ramdisk: the loader ferries a bundle of user binaries; the kernel parses
# it and spawns each as a ring-3 process (here the VFS-server stub heartbeats). # it and spawns each as a ring-3 process (here the VFS-server stub heartbeats).
{"name": "initial-ramdisk", {"name": "initial-ramdisk",