Files
danos/library/kernel/process.zig
T
Daniel Samson f23f073624 kernel: the device rides system_spawn, so a driver never runs without it
Delegation moves out of onHello and into the spawn itself. The manager holds
the hardware and names it in the call that creates the driver; the kernel
checks the device is the caller's to give, then hands it over as part of
making the child.

The reason is the window. A transfer after spawning always leaves an
interval in which the child is running and does not yet hold its device. It
would close on QEMU every time and open occasionally on a machine with
different core counts and timing — the exact failure shape this track exists
to delete, and not one worth introducing while removing the others. Fused
into the spawn there is no interval: the child does not exist until it holds
the device.

Ownership is checked BEFORE the child is created, so a refusal leaves
nothing running rather than a driver without the hardware it was spawned
for. The IOMMU confinement moves with the device, as it does on the transfer
path. systemCall6 is added for the sixth argument; r9 was free, and abi
gains a no_device sentinel matching the protocol's.

No driver had to change to receive a device, which is what makes this
better than requiring every driver to hello: ps2-bus keeps its legacy
status, and discovery — which has no assignment at all, since it is what
produces the device tree — is unaffected.

The attacker fixture now tries the spawn as a back door: name someone else's
device, and both the spawn and any child must be refused. Verifying that
assertion exposed a bug in the fixture itself. The kernel case's pass marker
was "device-authority: ok", which matches the FIRST per-assertion line, so
its wait loop exited before any failure was printed — the case would have
passed with failures in it, and had been able to since D2. The verdict lines
now carry a distinct VERDICT prefix, and with the ownership check removed
the case genuinely fails. A green test that cannot go red is worse than no
test.

Suite 118/118.
2026-08-08 19:39:32 +01:00

241 lines
10 KiB
Zig

//! Process-level runtime types: what a user program receives at entry (`Init`,
//! the argv contract) and the process end of the lifecycle
//! (docs/process-lifecycle.md) — today the exit reason a supervisor reads to
//! decide restart; signals and the stop sequence land here with M17.4. Mirrors
//! the spirit of `std.process.Init.Minimal` in danos terms — std's `Args` holds
//! no data on freestanding targets, so the type is danos's own.
const std = @import("std");
const abi = @import("abi");
const sc = @import("system-call");
const ipc = @import("ipc");
const time = @import("time");
/// Everything a program receives at entry. Passed to
/// `pub fn main(init: runtime.process.Init)`; programs that need nothing keep
/// `pub fn main() void`. An `environment` field is added here once the kernel
/// passes a non-empty envp (today it is always empty — see docs/sysv.md).
pub const Init = struct {
arguments: Arguments,
};
/// The process arguments (argc/argv), parsed from the kernel-built System V
/// entry block. The bytes live in the entry block at the top of the stack page,
/// NUL-terminated, valid for the process's lifetime.
pub const Arguments = struct {
/// argc — at least 1: argument 0 is the path or name this binary was
/// spawned as.
count: usize,
/// The argv pointers in the entry block (NULL-terminated after `count`
/// entries).
vector: [*]const [*:0]const u8,
/// Argument `index` (0 = the program's own path/name), or null if out of
/// range.
pub fn get(arguments: Arguments, index: usize) ?[:0]const u8 {
if (index >= arguments.count) return null;
return std.mem.span(arguments.vector[index]);
}
pub fn iterate(arguments: Arguments) Iterator {
return .{ .arguments = arguments };
}
pub const Iterator = struct {
arguments: Arguments,
index: usize = 0,
pub fn next(iterator: *Iterator) ?[:0]const u8 {
const argument = iterator.arguments.get(iterator.index) orelse return null;
iterator.index += 1;
return argument;
}
};
};
/// How a process ended — what a supervisor's restart policy reads: a clean exit
/// meant to stop, a fault wants a restart with backoff, killed means the
/// supervisor did it itself (docs/process-lifecycle.md).
pub const ExitReason = abi.ExitReason;
/// How dead child `id` ended. Ask after the exit notification arrives — the
/// kernel records the reason before it posts the notification, so this never
/// races it. Returns null for an id that never lived, is still alive, was
/// evicted from the kernel's bounded record, or is not this process's child
/// (the same authority gate as `kill`).
pub fn exitReason(id: u32) ?ExitReason {
const r = sc.systemCall1(.process_exit_reason, id);
if (r > ~@as(usize, 0) - 4095) return null; // a wrapped -errno
return @enumFromInt(r);
}
/// The signal vocabulary (docs/process-lifecycle.md): POSIX's concepts, danos's
/// names, message delivery. A signal is a one-way coalescing statement — never a
/// question (liveness is the zero-length ping call) and never kill (that is
/// `system.kill`, unhandleable by definition).
pub const Signal = abi.Signal;
/// The coalesced set of signals one notification delivered: two pending
/// terminates arrive as one. Decode a received badge with `signalsFrom`.
pub const SignalSet = struct {
pending: u32,
pub fn has(set: SignalSet, signal: Signal) bool {
return set.pending & (@as(u32, 1) << @intFromEnum(signal)) != 0;
}
};
/// Nominate `endpoint` as this process's signal endpoint. Signals posted while
/// unbound have pended; they are delivered immediately on bind, coalesced.
pub fn bindSignals(endpoint: usize) bool {
return sc.systemCall1(.signal_bind, endpoint) == 0;
}
/// Decode a received badge into the signals it delivered, or null if it is not
/// a signal notification.
pub fn signalsFrom(badge: u64) ?SignalSet {
if (badge & abi.notify_badge_bit == 0 or badge & abi.notify_signal_bit == 0) return null;
return .{ .pending = @truncate(badge & ~(abi.notify_badge_bit | abi.notify_signal_bit)) };
}
/// Post `signal` to child `id` (or to yourself). Supervisor-gated, like kill;
/// non-blocking, always — a statement, not a conversation.
pub fn sendSignal(id: u32, signal: Signal) bool {
return sc.systemCall2(.process_signal, id, @intFromEnum(signal)) == 0;
}
/// The standard stop sequence (docs/process-lifecycle.md): terminate, wait up to
/// `deadline_ms` for the exit notification on `exit_endpoint` (the endpoint the
/// child was spawned with), then kill. Any *other* notifications arriving on
/// that endpoint while stopping are consumed and dropped — a supervisor with
/// concurrent traffic implements the same sequence inside its own event loop
/// (arm `system.timerOnce`, keep serving) instead of calling this.
pub fn stop(id: u32, deadline_ms: u64, exit_endpoint: usize) void {
_ = sendSignal(id, .terminate);
_ = time.timerOnce(exit_endpoint, deadline_ms);
var receive: [8]u8 = undefined;
while (true) {
const got = ipc.replyWait(exit_endpoint, &.{}, &receive, null);
if (got.isChildExit() and got.childProcessId() == id) return;
if (got.isTimer()) break; // the deadline passed first — escalate
}
_ = kill(id);
while (true) {
const got = ipc.replyWait(exit_endpoint, &.{}, &receive, null);
if (got.isChildExit() and got.childProcessId() == id) return;
}
}
/// Subscribe `endpoint` to published exit events: every process death posts an
/// asynchronous notification with the same badge encoding as a supervisor's exit
/// notice (decode with `ipc.Received.isChildExit`/`childProcessId`). For stateful
/// services: release what the dead client held — file handles, subscriptions —
/// because a service must never depend on clients cleaning up after themselves
/// (docs/process-lifecycle.md). Ungated, like `system.processes`.
pub fn subscribeExits(endpoint: usize) bool {
return sc.systemCall1(.process_subscribe, endpoint) == 0;
}
// --- raw process syscalls, formerly in the system.zig dumping ground ---
/// One `processes` entry — re-exported from the shared ABI so a program can declare its
/// snapshot buffer without importing `abi` itself.
pub const ProcessDescriptor = abi.ProcessDescriptor;
/// The calling task's own kernel id — its row in the process table, and the value
/// every other process sees as this one's `supervisor` after it spawns them. For a
/// single-threaded program that is its process id; in a threaded one it is the
/// calling thread's id (`Thread.getCurrentId` is the same system call, named for
/// the threading vocabulary). Ids are monotonic and never reused
/// (system/kernel/process.zig), which is what makes comparing one an identity
/// test where comparing a *name* is only a resemblance test — the registrar in
/// init leans on exactly that.
pub fn taskId() u32 {
return @intCast(sc.systemCall0(.thread_self));
}
/// Give up the rest of this quantum.
pub fn yield() void {
_ = sc.systemCall0(.yield);
}
/// End the process. Never returns.
pub fn exit(code: usize) noreturn {
_ = sc.systemCall1(.exit, code);
unreachable; // the kernel never returns from exit
}
/// Start the binary bundled in the initial-ramdisk under `name` as a new ring-3 process,
/// returning the child's process id (or null). argv[0] is `name`, and the caller becomes
/// its **supervisor** — the only process allowed to `kill` it.
pub fn spawn(name: []const u8) ?u32 {
return spawnSupervised(name, &.{}, null);
}
/// Like `spawn`, but hands the child argv[1..] (argv[0] is still `name`).
pub fn spawnWithArguments(name: []const u8, arguments: []const []const u8) ?u32 {
return spawnSupervised(name, arguments, null);
}
/// The full spawn: argv[1..] for the child, and an optional endpoint the kernel notifies
/// when the child ends (any way — clean exit, fault, or `kill`), delivered via
/// `ipc.replyWait` as a child-exit badge (`ipc.Received.isChildExit`/`childProcessId`), so
/// one endpoint can supervise many children. Returns the child's process id, or null.
pub fn spawnSupervised(name: []const u8, arguments: []const []const u8, exit_endpoint: ?usize) ?u32 {
return spawnSupervisedWithDevice(name, arguments, exit_endpoint, abi.no_device);
}
/// Spawn a supervised child **and give it a device you hold**, in one call.
///
/// The device manager's path: it holds the hardware and the driver it starts must have
/// it. Fusing the handover into the spawn is what removes the window a separate
/// transfer would leave — the child cannot run without its device, because it does not
/// exist until it holds it (docs/os-development/device-authority.md). The kernel checks
/// only that the device is the caller's to give.
pub fn spawnSupervisedWithDevice(
name: []const u8,
arguments: []const []const u8,
exit_endpoint: ?usize,
device: u64,
) ?u32 {
var blob: [256]u8 = undefined;
var len: usize = 0;
for (arguments, 0..) |argument, i| {
if (i != 0) {
if (len >= blob.len) return null;
blob[len] = 0;
len += 1;
}
if (len + argument.len > blob.len) return null;
@memcpy(blob[len..][0..argument.len], argument);
len += argument.len;
}
const r = sc.systemCall6(.system_spawn, @intFromPtr(name.ptr), name.len, if (len == 0) 0 else @intFromPtr(&blob), len, exit_endpoint orelse abi.no_cap, device);
if (r > ~@as(usize, 0) - 4095) return null; // a wrapped -errno
return @intCast(r);
}
/// Snapshot the process table into `out` and return the total number of live processes
/// (which may exceed `out.len`; call again with a larger buffer). Kernel tasks are
/// included, with an empty name. The primitive `ps` is built on.
pub fn processes(out: []ProcessDescriptor) usize {
return sc.systemCall2(.process_enumerate, @intFromPtr(out.ptr), out.len);
}
/// Whether a process spawned under `name` (its argv[0]) is currently alive.
pub fn isProcessRunning(name: []const u8) bool {
var table: [32]ProcessDescriptor = undefined;
const total = processes(&table);
for (table[0..@min(total, table.len)]) |descriptor| {
if (std.mem.eql(u8, descriptor.name[0..descriptor.name_length], name)) return true;
}
return false;
}
/// End process `id`. Only its supervisor — the process that spawned it — may; anyone else
/// gets false, as does a stale or unknown id. Delivery is prompt but asynchronous, like a
/// signal. True means the kill is accepted and irrevocable.
pub fn kill(id: u32) bool {
return sc.systemCall1(.process_kill, id) == 0;
}