Add process management: enumerate, supervisor-gated kill, exit notifications

process_enumerate snapshots the task table (the device_enumerate shape, so
ps is a user program); system_spawn returns the child id, records the caller
as supervisor, and takes an exit endpoint; process_kill is allowed only for
the supervisor. Every death — exit, fault, or kill — posts a child-exit badge
to that endpoint (the IRQ-as-IPC pattern as SIGCHLD). A target caught off-CPU
is reaped in place; a running one is condemned and finished at its next
system call or tick, guarded so teardown never lands mid-kernel-operation.
Tested by process-list, process-kill, and supervision (a ring-3 supervisor
exercising the whole surface); design notes in docs/process-management.md.
This commit is contained in:
Daniel Samson
2026-07-11 09:32:25 +01:00
parent a5fe63c1dd
commit d218d93f79
13 changed files with 949 additions and 68 deletions
+3
View File
@@ -300,6 +300,7 @@ pub fn build(b: *std.Build) void {
const bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, "bus", "system/drivers/bus/bus.zig"); const bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, "bus", "system/drivers/bus/bus.zig");
const device_manager_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, "device-manager", "system/services/device-manager/device-manager.zig"); const device_manager_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, "device-manager", "system/services/device-manager/device-manager.zig");
const args_echo_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, "args-echo", "system/services/args-echo/args-echo.zig"); const args_echo_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, "args-echo", "system/services/args-echo/args-echo.zig");
const process_test_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, "process-test", "system/services/process-test/process-test.zig");
// Pack the user binaries into the initial_ramdisk image with the host-side Python tool // Pack the user binaries into the initial_ramdisk image with the host-side Python tool
// (the container format is trivial, and Python sidesteps std API churn). Args: // (the container format is trivial, and Python sidesteps std API churn). Args:
@@ -319,6 +320,8 @@ pub fn build(b: *std.Build) void {
mk_run.addFileArg(device_manager_exe.getEmittedBin()); mk_run.addFileArg(device_manager_exe.getEmittedBin());
mk_run.addArg("args-echo"); mk_run.addArg("args-echo");
mk_run.addFileArg(args_echo_exe.getEmittedBin()); mk_run.addFileArg(args_echo_exe.getEmittedBin());
mk_run.addArg("process-test");
mk_run.addFileArg(process_test_exe.getEmittedBin());
// Also install the packed binaries to their FHS homes, so zig-out is a true image // Also install the packed binaries to their FHS homes, so zig-out is a true image
// of the filesystem — even though at boot they arrive inside the initial-ramdisk. // of the filesystem — even though at boot they arrive inside the initial-ramdisk.
+5 -1
View File
@@ -48,7 +48,11 @@ rather than restate it. Roughly in the order things happen at runtime:
real driver stacks factor into three shapes, how families share code, and the real driver stacks factor into three shapes, how families share code, and the
proposed ABI for the three primitives still missing (capability passing, DMA + proposed ABI for the three primitives still missing (capability passing, DMA +
memory barriers, MSI). memory barriers, MSI).
15. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and 15. **[process-management.md](process-management.md) — process management.** The
microkernel's `ps`/`kill`/SIGCHLD: enumerate as a table snapshot, the
supervision link as the kill authority, and child-exit notifications over the
same endpoints IRQs arrive on.
16. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
how `while (true) hlt` parks the CPU safely once there's nothing left to do. how `while (true) hlt` parks the CPU safely once there's nothing left to do.
Start with the north star: Start with the north star:
+112
View File
@@ -0,0 +1,112 @@
# Process Management
How danos lists, supervises, and kills processes — the microkernel answer to
`ps`, `kill`, and `SIGCHLD`/`wait`.
## Why system calls, not `/proc`
Unix systems sit on a spectrum. Classic BSD/macOS list processes through
syscalls (`sysctl(KERN_PROC)`) and kill through `kill(2)`; Linux renders the
process table as `/proc` for *reading* but still kills through a syscall; Plan 9
made the file tree the whole interface (`echo kill > /proc/n/ctl`). Microkernels
mostly abandon ambient PIDs: Minix and QNX route everything through a user-space
process-manager server, and Fuchsia/seL4 control processes only through handles.
danos rules out `/proc` **as the primitive**: here a `/proc` would be served by
the VFS server — a user process — which would put the VFS in the path of process
control. If the VFS (or anything under it) hangs, nothing could be listed or
killed, *including the hung VFS*. The control plane for processes must not
depend on a process. So the primitives are kernel system calls; a read-only
`/proc` rendering can be layered on later, and a POSIX-style process-manager
server can be built *from* these primitives when one is needed.
## The three primitives
### `process_enumerate(buffer, maximum) -> total`
A snapshot of the task table into a caller buffer of `abi.ProcessDescriptor`
(id, supervisor, state, priority, name) — the exact shape of
`device_enumerate`, so `ps` is a user program over a snapshot, not a kernel
service. The total may exceed what fit; call again with a larger buffer. Kernel
tasks are included with an empty name — an honest listing shows the idle tasks
too. Ungated and read-only: what is running is not a secret between cooperating
bring-up processes.
### `system_spawn(..., exit_endpoint) -> child id`, and the supervision link
`system_spawn` records the caller as the child's **supervisor** and returns the
child's process id (ids are monotonic, never reused — a stale id can only miss).
That link is the kill authority: it answers "who may kill process 7?" without
inventing users or permissions, the same way a device *claim* is the capability
for `mmio_map`. It composes with the supervision hierarchy the device manager
already forms: init supervises the services it starts, the device manager
supervises the drivers it matches. (A transferable process *handle* — Fuchsia
style — can replace the id once the handle table grows types beyond endpoints.)
`exit_endpoint` (a handle, or `abi.no_cap`) is the supervisor's death-watch: when
the child ends — clean exit, CPU fault, or `process_kill` — the kernel posts an
asynchronous notification to that endpoint, exactly like a bound IRQ. The badge
carries `abi.notify_badge_bit | abi.notify_exit_bit | child_id`, so one endpoint
supervises many children and can even share with IRQ notifications. This is the
microkernel's SIGCHLD: no new mechanism, just the IRQ-as-IPC pattern reused, and
a supervisor's event loop (`ipc.replyWait`) already knows how to receive it. The
child holds a reference to the endpoint from birth, so the notification cannot
dangle even if the supervisor dies first.
### `process_kill(id) -> 0 / -ESRCH / -EPERM`
Only the supervisor may kill; kernel tasks are not killable processes. Like a
signal, delivery is prompt but asynchronous — 0 means the kill is accepted and
irrevocable; the exit notification confirms completion.
## How a kill lands (the kernel mechanics)
Everything below runs under the big kernel lock, where task states cannot move.
- **Target ready or blocked** (not on any core): reaped on the killer's own
call. The reap releases what death always releases (IRQ bindings first, then
a client the target still owed a reply to is failed with `-EPEER`, IPC handles
closed, the exit notification posted last) — plus the unlinking only a
*remote* death needs: out of the ready queue, out of an endpoint's sender FIFO
(`Task.ipc_wait_endpoint`), out of a receive wait queue (`Task.wait_queue`),
and out of any server's owed-reply slot, so nothing ever dequeues a dangling
pointer. Destroying the address space is safe because no core can have it
loaded: every switch away from a task loads the next task's tables.
- **Target running on another core**: it cannot be torn down mid-instruction,
so it is condemned (`Task.kill_pending`) and dies at whichever comes first:
- its next **system_call entry** — checked before dispatch, so a condemned
process cannot spawn, claim, or message anything on its way out;
- its core's next **timer tick** — but only when the task is not inside one
of its own system calls (`Task.in_system_call`): the tick may have
interrupted kernel code mid-operation, where teardown would leak whatever
the operation held. User-mode execution is always a safe kill point. The
tick-time terminate abandons the interrupt frame exactly like the fault
path (the LAPIC is acknowledged before the tick hook runs);
- any core's tick finding it **blocked or ready** (it entered a syscall and
parked after being condemned) — reaped by the same remote-reap path.
A pure user-mode spin loop that never makes a system call therefore dies
within one tick; nothing a process does can outrun the kill.
The scheduler stays below the process layer: finishing a kill (IRQ bindings,
handles, the notification) is called *up* through two hooks process.zig
registers at boot (`terminate_current_hook`, `reap_task_hook`), mirroring how
the architecture layer calls up into `tick`.
## Known gaps (bring-up honesty)
- Device **claims** are not released on death (pre-existing: the fault path has
the same gap) — a killed driver's device stays claimed until reboot.
- Kernel stacks of dead tasks are leaked, as on every exit path (no reaper yet).
- There is no exit *status* in the notification, only the id; a supervisor that
needs the code can grow a wait-style call later.
- Enumerate writes through the caller's raw pointer under the bring-up trust
model, like `device_enumerate` (an unmapped page is a self-DoS, not an
isolation break).
## Tests
`process-list` (enumerate), `process-kill` (kernel-level kill paths, refusals,
notifications), `supervision` (the whole user-side surface via the process-test
service: spawn supervised → enumerate → kill blocked and spinning children →
notifications → gone). See test/qemu_test.py.
+21 -3
View File
@@ -84,6 +84,11 @@ pub fn call(h: Handle, message: []const u8, reply: []u8) CallError!usize {
/// GSI. See `isNotification`. /// GSI. See `isNotification`.
pub const notify_badge_bit: u64 = abi.notify_badge_bit; pub const notify_badge_bit: u64 = abi.notify_badge_bit;
/// Set alongside `notify_badge_bit` when the notification is a **child-exit
/// notice** — a process this one spawned (with an exit endpoint) has ended —
/// rather than a device interrupt. The low bits carry the child's process id.
pub const notify_exit_bit: u64 = abi.notify_exit_bit;
/// The result of a `replyWait`: the request length, the sender's badge (a task id, or /// The result of a `replyWait`: the request length, the sender's badge (a task id, or
/// an IRQ notification if the high bit is set), and any capability the request carried. /// an IRQ notification if the high bit is set), and any capability the request carried.
pub const Received = struct { pub const Received = struct {
@@ -91,16 +96,29 @@ pub const Received = struct {
badge: u64, badge: u64,
cap: ?Handle, cap: ?Handle,
/// True if this wake-up was a device interrupt, not a client request. A driver's /// True if this wake-up was an asynchronous notification (a device interrupt
/// event loop branches on this; there is no reply owed on the notification path. /// or a child-exit notice), not a client request. An event loop branches on
/// this; there is no reply owed on the notification path.
pub fn isNotification(self: Received) bool { pub fn isNotification(self: Received) bool {
return self.badge & notify_badge_bit != 0; return self.badge & notify_badge_bit != 0;
} }
/// The interrupt source (a GSI), meaningful only when `isNotification`. /// True if this wake-up tells of a supervised child's end — the notification
/// requested by passing an exit endpoint to `system.spawnSupervised`.
pub fn isChildExit(self: Received) bool {
return self.isNotification() and self.badge & notify_exit_bit != 0;
}
/// The interrupt source (a GSI), meaningful only when `isNotification` and
/// not `isChildExit`.
pub fn source(self: Received) u64 { pub fn source(self: Received) u64 {
return self.badge & ~notify_badge_bit; return self.badge & ~notify_badge_bit;
} }
/// The ended child's process id, meaningful only when `isChildExit`.
pub fn childProcessId(self: Received) u32 {
return @intCast(self.badge & ~(notify_badge_bit | notify_exit_bit));
}
}; };
/// Server side of IPC_ReplyWait: deliver `reply` to the client last received (if any, /// Server side of IPC_ReplyWait: deliver `reply` to the client last received (if any,
+49 -12
View File
@@ -11,6 +11,10 @@ pub const PROT_READ: usize = abi.prot_read;
pub const PROT_WRITE: usize = abi.prot_write; pub const PROT_WRITE: usize = abi.prot_write;
pub const PROT_EXEC: usize = abi.prot_exec; pub const PROT_EXEC: usize = abi.prot_exec;
/// One `processes` entry — re-exported from the shared ABI so a user program can
/// declare its snapshot buffer without importing `abi` itself.
pub const ProcessDescriptor = abi.ProcessDescriptor;
/// Give up the rest of this quantum. /// Give up the rest of this quantum.
pub fn yield() void { pub fn yield() void {
_ = sc.systemCall0(.yield); _ = sc.systemCall0(.yield);
@@ -44,31 +48,64 @@ pub fn exit(code: usize) noreturn {
} }
/// Start the binary bundled in the initial-ramdisk under `name` as a new ring-3 /// Start the binary bundled in the initial-ramdisk under `name` as a new ring-3
/// process, returning true on success. The child's argv[0] is `name`. This is how /// process, returning the child's process id (or null on failure). The child's
/// a supervisor (the device manager) launches a driver it matched — danos-native, /// argv[0] is `name`, and the caller becomes its **supervisor** — the only process
/// not POSIX (a spawn/exec family comes with the process work later). /// allowed to `kill` it. This is how a supervisor (the device manager) launches a
pub fn spawn(name: []const u8) bool { /// driver it matched — danos-native, not POSIX (a spawn/exec family comes with the
return sc.systemCall4(.system_spawn, @intFromPtr(name.ptr), name.len, 0, 0) == 0; /// POSIX layer later).
pub fn spawn(name: []const u8) ?u32 {
return spawnSupervised(name, &.{}, null);
} }
/// Like `spawn`, but hands the child command-line arguments: they arrive as /// Like `spawn`, but hands the child command-line arguments: they arrive as
/// argv[1..] on its System V entry stack (argv[0] is still `name`). Marshalled to /// argv[1..] on its System V entry stack (argv[0] is still `name`).
/// the kernel as one NUL-separated blob; the combined arguments must fit pub fn spawnWithArguments(name: []const u8, arguments: []const []const u8) ?u32 {
/// `blob` (the kernel caps the blob at 256 bytes and argc at 8 anyway). return spawnSupervised(name, arguments, null);
pub fn spawnWithArguments(name: []const u8, arguments: []const []const u8) bool { }
/// The full spawn: command-line arguments for the child, and an optional endpoint
/// (a handle from `ipc.createIpcEndpoint`) the kernel notifies when the child ends
/// — any way it ends: clean exit, fault, or `kill`. The notification arrives via
/// `ipc.replyWait` as a badge with the child-exit bit set and the child's id in
/// the low bits (`ipc.Received.isChildExit`/`childProcessId`), so one endpoint can
/// supervise many children. Arguments are marshalled to the kernel as one
/// NUL-separated blob; the combined arguments must fit `blob` (the kernel caps the
/// blob at 256 bytes and argc at 8 anyway). Returns the child's process id, or
/// null on failure.
pub fn spawnSupervised(name: []const u8, arguments: []const []const u8, exit_endpoint: ?usize) ?u32 {
var blob: [256]u8 = undefined; var blob: [256]u8 = undefined;
var len: usize = 0; var len: usize = 0;
for (arguments, 0..) |argument, i| { for (arguments, 0..) |argument, i| {
if (i != 0) { if (i != 0) {
if (len >= blob.len) return false; if (len >= blob.len) return null;
blob[len] = 0; blob[len] = 0;
len += 1; len += 1;
} }
if (len + argument.len > blob.len) return false; if (len + argument.len > blob.len) return null;
@memcpy(blob[len..][0..argument.len], argument); @memcpy(blob[len..][0..argument.len], argument);
len += argument.len; len += argument.len;
} }
return sc.systemCall4(.system_spawn, @intFromPtr(name.ptr), name.len, @intFromPtr(&blob), len) == 0; const r = sc.systemCall5(.system_spawn, @intFromPtr(name.ptr), name.len, if (len == 0) 0 else @intFromPtr(&blob), len, exit_endpoint orelse abi.no_cap);
if (r > ~@as(usize, 0) - 4095) return null; // a wrapped -errno
return @intCast(r);
}
/// Snapshot the process table into `out` (up to its length) and return the total
/// number of live processes — which may exceed `out.len`; call again with a larger
/// buffer for the full listing. Kernel tasks are included, with an empty name.
/// The primitive `ps` is built on.
pub fn processes(out: []abi.ProcessDescriptor) usize {
return sc.systemCall2(.process_enumerate, @intFromPtr(out.ptr), out.len);
}
/// End process `id`. Only its supervisor — the process that spawned it — may;
/// anyone else gets false, as does a stale or unknown id (ids are never reused).
/// Delivery is prompt but asynchronous, like a signal: a target caught running on
/// another core dies at its next system call or timer tick. True means the kill
/// is accepted and irrevocable; the exit notification (if an endpoint was given
/// at spawn) confirms completion.
pub fn kill(id: u32) bool {
return sc.systemCall1(.process_kill, id) == 0;
} }
/// Grant `len` bytes (rounded up to whole pages) of fresh, zeroed, writable /// Grant `len` bytes (rounded up to whole pages) of fresh, zeroed, writable
+40 -5
View File
@@ -43,13 +43,15 @@ pub const SystemCall = enum(u64) {
irq_bind = 14, // irq_bind(id, resource_index, endpoint): deliver a device IRQ as an IPC notification irq_bind = 14, // irq_bind(id, resource_index, endpoint): deliver a device IRQ as an IPC notification
irq_ack = 15, // irq_ack(id, resource_index): re-arm a bound IRQ after servicing it irq_ack = 15, // irq_ack(id, resource_index): re-arm a bound IRQ after servicing it
device_register = 16, // device_register(parent_id, descriptor) -> id: publish a child of a device you claimed device_register = 16, // device_register(parent_id, descriptor) -> id: publish a child of a device you claimed
system_spawn = 17, // system_spawn(name_ptr, name_len) -> 0: start a named initial-ramdisk binary as a new ring-3 process system_spawn = 17, // system_spawn(name_ptr, name_len, arguments_ptr, arguments_len, exit_endpoint) -> child process id: start a named initial-ramdisk binary as a new ring-3 process
dma_alloc = 18, // dma_alloc(len, flags) -> vaddr (rax), paddr (rdx): contiguous, pinned, uncacheable DMA memory dma_alloc = 18, // dma_alloc(len, flags) -> vaddr (rax), paddr (rdx): contiguous, pinned, uncacheable DMA memory
dma_free = 19, // dma_free(vaddr, len) -> 0: release a prior dma_alloc dma_free = 19, // dma_free(vaddr, len) -> 0: release a prior dma_alloc
msi_bind = 20, // msi_bind(device_id, endpoint) -> address (rax), data (rdx): a per-device MSI vector for a claimed device msi_bind = 20, // msi_bind(device_id, endpoint) -> address (rax), data (rdx): a per-device MSI vector for a claimed device
io_read = 21, // io_read(device_id, resource_index, offset, width) -> value: read a port in a claimed device's io_port resource io_read = 21, // io_read(device_id, resource_index, offset, width) -> value: read a port in a claimed device's io_port resource
io_write = 22, // io_write(device_id, resource_index, offset, width, value) -> 0: write a port in a claimed device's io_port resource io_write = 22, // io_write(device_id, resource_index, offset, width, value) -> 0: write a port in a claimed device's io_port resource
clock = 23, // clock() -> nanoseconds since boot: a monotonic time source (for timeouts/delays) clock = 23, // clock() -> nanoseconds since boot: a monotonic time source (for timeouts/delays)
process_enumerate = 24, // process_enumerate(buffer, maximum) -> total: snapshot the task table
process_kill = 25, // process_kill(id) -> 0/-errno: end a process this process spawned
_, _,
}; };
@@ -67,12 +69,45 @@ pub const dma_write_combining: u64 = 2; // write-combining (framebuffers); needs
pub const dma_below_4g: u64 = 4; // physical address must fit 32 bits (legacy DMA engines) pub const dma_below_4g: u64 = 4; // physical address must fit 32 bits (legacy DMA engines)
/// Set in the badge returned by `ipc_reply_wait` when what arrived is an /// Set in the badge returned by `ipc_reply_wait` when what arrived is an
/// **asynchronous notification** (today: a device interrupt bound with `irq_bind`) /// **asynchronous notification** (a device interrupt bound with `irq_bind`, or a
/// rather than a message from a client. There is no payload and no reply owed; the /// child-exit notice — see `notify_exit_bit`) rather than a message from a client.
/// low bits carry the source, a GSI. Shared so the kernel's ISR and the driver's /// There is no payload and no reply owed; the low bits carry the source. Shared so
/// event loop can't disagree about which bit means "the hardware spoke". /// the kernel's ISR and the driver's event loop can't disagree about which bit
/// means "the hardware spoke".
pub const notify_badge_bit: u64 = 1 << 63; pub const notify_badge_bit: u64 = 1 << 63;
/// Set (alongside `notify_badge_bit`) in the badge of a **child-exit notification**:
/// posted to the endpoint a supervisor passed to `system_spawn` when that child ends
/// — by clean exit, by a fault, or by `process_kill`. The low bits carry the child's
/// process id, so one endpoint can supervise many children (and even share with IRQ
/// notifications, which never set this bit). The microkernel's SIGCHLD.
pub const notify_exit_bit: u64 = 1 << 62;
/// Capacity of `ProcessDescriptor.name` — matches the longest name `system_spawn`
/// accepts, so a process's recorded name (its argv[0]) is never truncated.
pub const maximum_process_name = 64;
/// What a process is doing right now, as reported by `process_enumerate`. Crosses
/// the system_call boundary as `ProcessDescriptor.state`.
pub const ProcessState = enum(u32) {
ready = 0, // runnable, waiting for a core
running = 1, // executing on a core right now
blocked = 2, // waiting (sleeping, or blocked in IPC)
};
/// One `process_enumerate` entry — the kernel's view of a live task, kernel tasks
/// included (they carry an empty name and id 0 is the boot task). Fixed layout
/// (extern) because it crosses the kernel↔user boundary by memory copy, like
/// `DeviceDescriptor` in the device ABI.
pub const ProcessDescriptor = extern struct {
id: u32, // kernel-assigned process id; never reused (monotonic)
supervisor: u32, // id of the process that spawned it (0 = the kernel)
state: u32, // a ProcessState value
priority: u32,
name_length: u32,
name: [maximum_process_name]u8, // argv[0] at spawn; empty for kernel tasks
};
/// Well-known IPC service ids for the bootstrap name registry (create_ipc_endpoint + /// Well-known IPC service ids for the bootstrap name registry (create_ipc_endpoint +
/// ipc_register/ipc_lookup). Small integers, so no string interning is needed /// ipc_register/ipc_lookup). Small integers, so no string interning is needed
/// during bring-up. The VFS server registers under `vfs`; clients look it up. /// during bring-up. The VFS server registers under `vfs`; clients look it up.
+27
View File
@@ -47,6 +47,8 @@ pub const ENOENT: i64 = 4; // no such registered service
pub const ENOSPC: i64 = 5; // handle table or registry full pub const ENOSPC: i64 = 5; // handle table or registry full
pub const ENOMEM: i64 = 6; // out of memory pub const ENOMEM: i64 = 6; // out of memory
pub const EPEER: i64 = 7; // peer died before replying (its process exited or was killed) pub const EPEER: i64 = 7; // peer died before replying (its process exited or was killed)
pub const ESRCH: i64 = 8; // no such process (process_kill of an unknown/dead id)
pub const EPERM: i64 = 9; // not permitted (process_kill by anyone but the supervisor)
/// A badge with this bit set is an asynchronous notification (e.g. an IRQ), not a /// A badge with this bit set is an asynchronous notification (e.g. an IRQ), not a
/// message from a client — there is no reply owed. The low bits carry the source /// message from a client — there is no reply owed. The low bits carry the source
@@ -93,6 +95,7 @@ pub fn dropRef(endpoint: *Endpoint) void {
// --- sender FIFO (endpoint-local, via Task.next) ---------------------------- // --- sender FIFO (endpoint-local, via Task.next) ----------------------------
fn enqueueSender(endpoint: *Endpoint, t: *Task) void { fn enqueueSender(endpoint: *Endpoint, t: *Task) void {
t.ipc_wait_endpoint = @ptrCast(endpoint); // so a kill can unlink a parked caller
t.next = null; t.next = null;
if (endpoint.sender_tail) |tail| tail.next = t else endpoint.sender_head = t; if (endpoint.sender_tail) |tail| tail.next = t else endpoint.sender_head = t;
endpoint.sender_tail = t; endpoint.sender_tail = t;
@@ -102,10 +105,34 @@ fn dequeueSender(endpoint: *Endpoint) ?*Task {
const t = endpoint.sender_head orelse return null; const t = endpoint.sender_head orelse return null;
endpoint.sender_head = t.next; endpoint.sender_head = t.next;
if (endpoint.sender_head == null) endpoint.sender_tail = null; if (endpoint.sender_head == null) endpoint.sender_tail = null;
t.ipc_wait_endpoint = null;
t.next = null; t.next = null;
return t; return t;
} }
/// Unlink `t` from the sender FIFO it queues in, if any — the kill path for a
/// client parked in `call` that no server has received yet. Without this, a dead
/// caller would later be dequeued as a dangling pointer. The endpoint is still
/// alive here: `t`'s own handle table holds a reference until closeHandles runs
/// (which the kill path does *after* this). Precondition: the big kernel lock is
/// held.
pub fn abandonSenderLocked(t: *Task) void {
const endpoint: *Endpoint = @ptrCast(@alignCast(t.ipc_wait_endpoint orelse return));
t.ipc_wait_endpoint = null;
var previous: ?*Task = null;
var node = endpoint.sender_head;
while (node) |n| : ({
previous = n;
node = n.next;
}) {
if (n != t) continue;
if (previous) |p| p.next = t.next else endpoint.sender_head = t.next;
if (endpoint.sender_tail == t) endpoint.sender_tail = previous;
t.next = null;
return;
}
}
// --- cross-address-space copy ---------------------------------------------- // --- cross-address-space copy ----------------------------------------------
/// Copy `len` bytes from `source_va` in address space `source_as` to `destination_va` in /// Copy `len` bytes from `source_va` in address space `source_as` to `destination_va` in
+170 -27
View File
@@ -129,9 +129,14 @@ pub fn setInitialRamdisk(image: []const u8) void {
/// written back into the trap frame, since the entry paths restore user registers /// written back into the trap frame, since the entry paths restore user registers
/// from it. One handler serves both the system_call/sysret and int-0x80 entry paths. /// from it. One handler serves both the system_call/sysret and int-0x80 entry paths.
/// ///
/// Install it once at boot (before any user code runs) via `init`. /// Install it once at boot (before any user code runs) via `init`. Also registers
/// the scheduler's kill hooks: the scheduler sits below this layer, so finishing a
/// deferred process_kill (IRQ bindings, IPC handles, the exit notification) is
/// called back up into here from the tick (see scheduler.reapKillPendingLocked).
pub fn init() void { pub fn init() void {
architecture.setSystemCallHandler(system_call); architecture.setSystemCallHandler(system_call);
scheduler.terminate_current_hook = terminateCurrentLocked;
scheduler.reap_task_hook = reapTaskLocked;
} }
/// Return -1 (as an unsigned bit pattern) in the system_call result register. /// Return -1 (as an unsigned bit pattern) in the system_call result register.
@@ -140,6 +145,19 @@ fn fail(state: *architecture.CpuState) void {
} }
fn system_call(state: *architecture.CpuState) void { fn system_call(state: *architecture.CpuState) void {
const t = scheduler.current();
const user = t.aspace != 0;
if (user) {
// A condemned process (process_kill caught it running) dies at its next
// kernel entry — before it can spawn, claim, or message anything else.
if (t.kill_pending) terminateCurrent();
// Mark the span of this call so the timer tick never tears the task down
// in the middle of a kernel operation (scheduler.reapKillPendingLocked).
t.in_system_call = true;
}
defer if (user) {
t.in_system_call = false;
};
switch (@as(SystemCall, @enumFromInt(architecture.systemCallNumber(state)))) { switch (@as(SystemCall, @enumFromInt(architecture.systemCallNumber(state)))) {
.exit => { .exit => {
exit_code = architecture.systemCallArg(state, 0); exit_code = architecture.systemCallArg(state, 0);
@@ -178,6 +196,8 @@ fn system_call(state: *architecture.CpuState) void {
.io_read => systemIoRead(state), .io_read => systemIoRead(state),
.io_write => systemIoWrite(state), .io_write => systemIoWrite(state),
.clock => systemClock(state), .clock => systemClock(state),
.process_enumerate => systemProcessEnumerate(state),
.process_kill => systemProcessKill(state),
_ => fail(state), _ => fail(state),
} }
} }
@@ -415,28 +435,40 @@ fn systemDeviceRegister(state: *architecture.CpuState) void {
architecture.setSystemCallResult(state, id); architecture.setSystemCallResult(state, id);
} }
/// system_spawn(name_ptr, name_len, arguments_ptr, arguments_len) -> 0 on success, /// system_spawn(name_ptr, name_len, arguments_ptr, arguments_len, exit_endpoint)
/// -1 on failure. Load the binary bundled in the initial-ramdisk under `name` as a /// -> the child's process id on success, -1 on failure. Load the binary bundled in
/// fresh ring-3 process. `name` becomes the child's argv[0] (and its task name, so /// the initial-ramdisk under `name` as a fresh ring-3 process. `name` becomes the
/// a fault report can say which binary died); `arguments` is an optional /// child's argv[0] (and its task name, so a fault report can say which binary
/// NUL-separated blob that becomes argv[1..] — how a supervisor parameterises what /// died); `arguments` is an optional NUL-separated blob that becomes argv[1..] —
/// it starts ("you are the driver for device 12"). 0/0 means no extra arguments. /// how a supervisor parameterises what it starts ("you are the driver for device
/// This is the mechanism a user-space supervisor (the device manager) uses to start /// 12"). 0/0 means no extra arguments. This is the mechanism a user-space
/// a driver it matched: discovery and policy stay in user space, the kernel only /// supervisor (the device manager) uses to start a driver it matched: discovery
/// spawns. /// and policy stay in user space, the kernel only spawns.
/// ///
/// Ungated for now — any process may spawn any bundled binary. A capability (only a /// The caller is recorded as the child's **supervisor** — the sole holder of the
/// supervisor holds the right to spawn) belongs here once the model grows one; see /// right to `process_kill` it (docs/process-management.md). `exit_endpoint` (a
/// docs/driver-model.md. Both buffers are bounds-checked into the user half exactly /// handle, or `abi.no_cap` for none) names an endpoint of the caller's to notify
/// like `debug_write`, and an unknown name or a load failure returns -1. /// when the child ends, any way it ends — the IRQ-as-IPC pattern reused as the
/// microkernel's SIGCHLD.
///
/// Spawning itself is still ungated — any process may spawn any bundled binary; a
/// spawn capability belongs here once the model grows one (docs/driver-model.md).
/// Both buffers are bounds-checked into the user half exactly like `debug_write`,
/// and an unknown name or a load failure returns -1.
fn systemSpawn(state: *architecture.CpuState) void { fn systemSpawn(state: *architecture.CpuState) void {
const ptr = architecture.systemCallArg(state, 0); const ptr = architecture.systemCallArg(state, 0);
const len = architecture.systemCallArg(state, 1); const len = architecture.systemCallArg(state, 1);
const arguments_ptr = architecture.systemCallArg(state, 2); const arguments_ptr = architecture.systemCallArg(state, 2);
const arguments_len = architecture.systemCallArg(state, 3); const arguments_len = architecture.systemCallArg(state, 3);
if (len == 0 or len > 64 or ptr >= user_half_end or ptr + len > user_half_end) return fail(state); const exit_handle = architecture.systemCallArg(state, 4);
const t = scheduler.current();
if (len == 0 or len > scheduler.maximum_task_name or ptr >= user_half_end or ptr + len > user_half_end) return fail(state);
if (arguments_len > maximum_argument_bytes) return fail(state); if (arguments_len > maximum_argument_bytes) return fail(state);
if (arguments_len != 0 and (arguments_ptr >= user_half_end or arguments_ptr + arguments_len > user_half_end)) return fail(state); if (arguments_len != 0 and (arguments_ptr >= user_half_end or arguments_ptr + arguments_len > user_half_end)) return fail(state);
const exit_endpoint: ?*ipc.Endpoint = if (exit_handle == abi.no_cap)
null
else
ipc.resolveHandle(t, exit_handle) orelse return failErr(state, ipc.EBADF);
const image = ramdisk_image orelse return fail(state); const image = ramdisk_image orelse return fail(state);
const rd = initial_ramdisk.Reader.init(image) orelse return fail(state); const rd = initial_ramdisk.Reader.init(image) orelse return fail(state);
@@ -458,19 +490,51 @@ fn systemSpawn(state: *architecture.CpuState) void {
while (i < rd.count) : (i += 1) { while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue; const item = rd.entry(i) orelse continue;
if (!std.mem.eql(u8, item.name, name)) continue; if (!std.mem.eql(u8, item.name, name)) continue;
spawnProcess(item.blob, 4, argv[0..argc]) catch return fail(state); const child = spawnProcessSupervised(item.blob, 4, argv[0..argc], t.id, exit_endpoint) catch return fail(state);
architecture.setSystemCallResult(state, 0); architecture.setSystemCallResult(state, child);
return; return;
} }
fail(state); // no bundled binary by that name fail(state); // no bundled binary by that name
} }
/// process_enumerate(buffer, maximum) -> total: snapshot the task table into the
/// caller's buffer (up to `maximum` `abi.ProcessDescriptor` entries), returning
/// the total live-task count — the exact shape of `device_enumerate`, so a `ps`
/// is a user program over a snapshot, not a kernel service. Read-only and
/// ungated: what is running is not a secret between cooperating bring-up
/// processes.
fn systemProcessEnumerate(state: *architecture.CpuState) void {
const buffer_ptr = architecture.systemCallArg(state, 0);
const maximum = architecture.systemCallArg(state, 1);
const t = scheduler.current();
if (t.aspace == 0 or buffer_ptr >= user_half_end) return fail(state);
const sz = @sizeOf(abi.ProcessDescriptor);
const cap = @min(maximum, (user_half_end - buffer_ptr) / sz); // clamp to the user half
const out: [*]abi.ProcessDescriptor = @ptrFromInt(buffer_ptr);
architecture.setSystemCallResult(state, scheduler.enumerate(out[0..@intCast(cap)]));
}
/// process_kill(id) -> 0 / -ESRCH / -EPERM: end the process `id`. Only its
/// supervisor — the process that spawned it — may do so; the supervision link is
/// the kill capability, so no user/permission model is needed and a stray id
/// cannot be a weapon (ids are never reused, so a stale one just misses).
fn systemProcessKill(state: *architecture.CpuState) void {
const t = scheduler.current();
if (t.aspace == 0) return fail(state);
const id = architecture.systemCallArg(state, 0);
if (id > std.math.maxInt(u32)) return failErr(state, ipc.ESRCH);
const r = killProcess(t.id, @intCast(id));
architecture.setSystemCallResult(state, @bitCast(r));
}
/// Processes killed by a CPU fault rather than a clean exit. Evidence for the /// Processes killed by a CPU fault rather than a clean exit. Evidence for the
/// fault-recovery test, and a health signal a supervisor can consult later. /// fault-recovery test, and a health signal a supervisor can consult later.
pub var fault_kill_count: u64 = 0; pub var fault_kill_count: u64 = 0;
/// Tear down the current user process and reschedule; never returns. Shared by the /// Release everything a dying task holds and tell its supervisor — the shared
/// exit system call and the fault path (`killCurrentProcess`). The order matters: /// half of every path out of a process: clean exit, fault kill, and process_kill
/// (both the immediate reap and the deferred tick-time terminate). The order
/// matters:
/// - IRQ bindings are dropped before the handle table closes: dropping the last /// - IRQ bindings are dropped before the handle table closes: dropping the last
/// endpoint reference destroys the Endpoint, and a still-bound GSI would have an /// endpoint reference destroys the Endpoint, and a still-bound GSI would have an
/// ISR call notifyFromIsr on freed memory the next time the device fired. /// ISR call notifyFromIsr on freed memory the next time the device fired.
@@ -479,20 +543,84 @@ pub var fault_kill_count: u64 = 0;
/// - A client this task still owes a reply to (it died between receive and reply) /// - A client this task still owes a reply to (it died between receive and reply)
/// is failed with -EPEER rather than left blocked forever — a dead server must /// is failed with -EPEER rather than left blocked forever — a dead server must
/// not hang its callers. /// not hang its callers.
pub fn terminateCurrent() noreturn { /// - The task is unlinked from wherever IPC parked it (an endpoint's sender FIFO,
const t = scheduler.current(); /// a receive wait queue, or a server's owed-reply slot) *before* the handles
{ /// close, so nothing ever dequeues a dangling pointer. These are no-ops for a
const flags = sync.enter(); /// running task ending itself; they matter when process_kill reaps a blocked one.
defer sync.leave(flags); /// - The exit notification is posted last, once the process can no longer act, so
/// a supervisor that receives it observes a fully-released child. The endpoint
/// reference taken at spawn is dropped with it.
/// Precondition: the big kernel lock is held.
fn releaseTaskResourcesLocked(t: *scheduler.Task) void {
irq.releaseOwner(t.id); irq.releaseOwner(t.id);
if (t.ipc_client) |client| { if (t.ipc_client) |client| {
t.ipc_client = null; t.ipc_client = null;
client.ipc_status = -ipc.EPEER; client.ipc_status = -ipc.EPEER;
scheduler.readyLocked(client); // its blocked `call` now returns the error scheduler.readyLocked(client); // its blocked `call` now returns the error
} }
ipc.abandonSenderLocked(t);
scheduler.removeFromWaitQueueLocked(t);
scheduler.forgetIpcClientLocked(t);
ipc.closeHandles(t); ipc.closeHandles(t);
if (t.exit_endpoint) |raw| {
const endpoint: *ipc.Endpoint = @ptrCast(@alignCast(raw));
t.exit_endpoint = null;
ipc.notifyLocked(endpoint, abi.notify_exit_bit | t.id);
ipc.dropRef(endpoint);
} }
scheduler.exitUser(); }
/// Tear down the current user process and reschedule; never returns. Shared by the
/// exit system call and the fault path (`killCurrentProcess`). See
/// `releaseTaskResourcesLocked` for what is released, and in what order.
pub fn terminateCurrent() noreturn {
_ = sync.enter(); // handed off through the exit switch, released by the resumed task
terminateCurrentLocked();
}
/// The body of `terminateCurrent` for a caller that already holds the big kernel
/// lock — the scheduler's tick calls this (via `terminate_current_hook`) to finish
/// a deferred process_kill on its own core's current task. Never returns; the
/// tick's abandoned interrupt frame is fine (the LAPIC was acknowledged before the
/// tick hook ran), exactly as on the fault path.
fn terminateCurrentLocked() noreturn {
releaseTaskResourcesLocked(scheduler.current());
scheduler.exitUserLocked();
}
/// Reap a condemned task that is NOT running on any core (ready or blocked — and
/// it cannot start running: state changes need the lock we hold). The other half
/// of a deferred process_kill, called by the scheduler's tick (via
/// `reap_task_hook`) and directly by `killProcess` for targets caught off-CPU.
/// Precondition: the big kernel lock is held.
fn reapTaskLocked(t: *scheduler.Task) void {
releaseTaskResourcesLocked(t);
scheduler.removeFromReadyQueueLocked(t); // no-op unless it was ready in a queue
scheduler.destroyTaskLocked(t);
}
/// Kill process `target_id` on behalf of `caller_id` — the kernel half of the
/// process_kill system call. Returns 0, -ESRCH (no such live process — kernel
/// tasks are not killable processes and stale ids miss, since ids are never
/// reused), or -EPERM (the caller is not the target's supervisor).
///
/// A target that is ready or blocked is reaped on the spot. One that is running
/// on another core cannot be torn down mid-instruction, so it is condemned
/// (`kill_pending`) and dies at its next system_call entry, block, or timer tick
/// — like a Unix signal, delivery is prompt but asynchronous. Either way the
/// call returns 0: the kill is accepted and irrevocable.
pub fn killProcess(caller_id: u32, target_id: u32) i64 {
const flags = sync.enter();
defer sync.leave(flags);
const target = scheduler.taskByIdLocked(target_id) orelse return -ipc.ESRCH;
if (target.aspace == 0) return -ipc.ESRCH; // kernel tasks are not processes
if (target.supervisor != caller_id) return -ipc.EPERM;
if (target.state == .running) {
target.kill_pending = true;
} else {
reapTaskLocked(target);
}
return 0;
} }
/// Kill the current user process in response to a CPU fault it raised in ring 3. /// Kill the current user process in response to a CPU fault it raised in ring 3.
@@ -865,11 +993,21 @@ fn entryStackBytes(argv: []const []const u8) usize {
/// convention (`buildEntryStack`). `argv[0]` is required — it names the process: /// convention (`buildEntryStack`). `argv[0]` is required — it names the process:
/// the path or initial-ramdisk name it was spawned as. It is also recorded on the /// the path or initial-ramdisk name it was spawned as. It is also recorded on the
/// task, so a fault report can say *which* binary died, not just its id. /// task, so a fault report can say *which* binary died, not just its id.
/// The kernel-internal spawn (init at boot, tests): supervisor 0, no exit
/// notification. `spawnProcessSupervised` is the full form.
pub fn spawnProcess(image: []const u8, priority: u3, argv: []const []const u8) InitError!void {
_ = try spawnProcessSupervised(image, priority, argv, 0, null);
}
/// `spawnProcess`, recording `supervisor` (the id of the process that asked — the
/// kill authority) and, if given, `exit_endpoint` to notify when the child ends
/// (a reference is taken here and dropped when the notification posts).
/// Returns the child's process id.
/// Returns immediately — the process runs preemptively on its own page tables /// Returns immediately — the process runs preemptively on its own page tables
/// alongside everything else, and its exit is handled by the system_call layer. /// alongside everything else, and its exit is handled by the system_call layer.
/// The whole build (address space + ELF load + task) runs under the kernel lock so /// The whole build (address space + ELF load + task) runs under the kernel lock so
/// it appears atomically and can't race pmm/heap on another core. /// it appears atomically and can't race pmm/heap on another core.
pub fn spawnProcess(image: []const u8, priority: u3, argv: []const []const u8) InitError!void { pub fn spawnProcessSupervised(image: []const u8, priority: u3, argv: []const []const u8, supervisor: u32, exit_endpoint: ?*ipc.Endpoint) InitError!u32 {
if (argv.len == 0 or argv.len > maximum_arguments) return error.BadArguments; if (argv.len == 0 or argv.len > maximum_arguments) return error.BadArguments;
// The entry block must leave most of the page as actual stack. // The entry block must leave most of the page as actual stack.
if (entryStackBytes(argv) > page_size / 2) return error.BadArguments; if (entryStackBytes(argv) > page_size / 2) return error.BadArguments;
@@ -901,8 +1039,13 @@ pub fn spawnProcess(image: []const u8, priority: u3, argv: []const []const u8) I
architecture.mapUserPageInto(aspace, page_virtual, stack_frame, true, false); // RW + NX architecture.mapUserPageInto(aspace, page_virtual, stack_frame, true, false); // RW + NX
} }
if (!scheduler.spawnUserLocked(aspace, parsed.entry, user_sp, priority, argv[0])) const child = scheduler.spawnUserLocked(aspace, parsed.entry, user_sp, priority, argv[0], supervisor, if (exit_endpoint) |endpoint| @ptrCast(endpoint) else null) orelse
return error.OutOfMemory; return error.OutOfMemory;
// The child holds a reference to its exit endpoint from birth to death. Taken
// only now, after nothing can fail; the lock is still held, so the child
// cannot run (let alone die) before the reference exists.
if (exit_endpoint) |endpoint| endpoint.refcount += 1;
return child;
} }
/// clock() -> nanoseconds since boot: a monotonic time source. The kernel already owns /// clock() -> nanoseconds since boot: a monotonic time source. The kernel already owns
+206 -11
View File
@@ -18,6 +18,7 @@
//! shared queues. //! shared queues.
const std = @import("std"); const std = @import("std");
const abi = @import("abi");
const parameters = @import("parameters"); const parameters = @import("parameters");
const architecture = @import("architecture"); const architecture = @import("architecture");
const heap = @import("heap.zig"); const heap = @import("heap.zig");
@@ -41,6 +42,26 @@ pub const Task = struct {
kstack_top: usize = 0, // top of `stack` (== TSS.rsp0 for a user task); 0 = none kstack_top: usize = 0, // top of `stack` (== TSS.rsp0 for a user task); 0 = none
wake_at: u64 = 0, // uptime (ms) to wake a sleeping task; 0 = not sleeping wake_at: u64 = 0, // uptime (ms) to wake a sleeping task; 0 = not sleeping
affinity: ?u32 = null, // null = runs on any core; else the index of its pinned core affinity: ?u32 = null, // null = runs on any core; else the index of its pinned core
// --- process management (process.zig) ---
// Id of the process that spawned this one (0 = the kernel). The supervision
// link is the kill authority: only the supervisor may process_kill a child.
supervisor: u32 = 0,
// Endpoint to notify when this process ends (any way: exit, fault, kill), or
// null. Holds its own reference, dropped when the notification is posted.
// Opaque here for the same reason as `handles` below.
exit_endpoint: ?*anyopaque = null,
// Set by process_kill on a task that is running on another core; the kernel
// finishes the kill at that task's next system call or timer tick.
kill_pending: bool = false,
// True while this task executes its own system call — the timer tick must not
// tear a task down in the middle of a kernel operation, only while it runs
// user code (or sits at a block point, where teardown is safe).
in_system_call: bool = false,
// Where this task is parked while blocked, so a kill can unlink it: the
// WaitQueue it waits on (maintained by waitLocked/wakeLocked), or the endpoint
// whose sender FIFO it queues in (maintained by the IPC layer; opaque here).
wait_queue: ?*WaitQueue = null,
ipc_wait_endpoint: ?*anyopaque = null,
// Physical root of this task's address space, or 0 for a kernel task (which // Physical root of this task's address space, or 0 for a kernel task (which
// runs on the shared kernel page tables). A user task carries its own. // runs on the shared kernel page tables). A user task carries its own.
aspace: u64 = 0, aspace: u64 = 0,
@@ -85,8 +106,9 @@ pub const Task = struct {
}; };
/// Capacity of `Task.name_buffer` — matches the longest name `system_spawn` /// Capacity of `Task.name_buffer` — matches the longest name `system_spawn`
/// accepts, so a spawned name is never truncated. /// accepts, so a spawned name is never truncated. Shared with the ABI's
pub const maximum_task_name = 64; /// ProcessDescriptor, so `enumerate` copies names without clipping.
pub const maximum_task_name = abi.maximum_process_name;
/// Size of each task's IPC handle table. Kept here (not in ipc_sync.zig) because /// Size of each task's IPC handle table. Kept here (not in ipc_sync.zig) because
/// it dimensions a field of `Task`; ipc_sync.zig re-exports it. /// it dimensions a field of `Task`; ipc_sync.zig re-exports it.
@@ -275,14 +297,18 @@ pub fn spawnOn(entry: *const fn () void, priority: Priority, cpu: u32) bool {
/// Spawn a **user** task: a task with its own address space (`aspace`) that starts /// Spawn a **user** task: a task with its own address space (`aspace`) that starts
/// in user mode at `entry` on `user_sp`, recorded under `name` (its argv[0]). /// in user mode at `entry` on `user_sp`, recorded under `name` (its argv[0]).
/// `supervisor` is the id of the spawning process (0 = the kernel) — the kill
/// authority — and `exit_endpoint` (an *ipc.Endpoint whose reference the caller
/// has already taken, or null) is notified when this process ends.
/// It gets a fresh kernel stack for syscalls/interrupts, and its first switch-in /// It gets a fresh kernel stack for syscalls/interrupts, and its first switch-in
/// lands in `user_task_trampoline`. /// lands in `user_task_trampoline`.
/// Returns false (creating nothing) if the table is full or out of memory. /// Returns the new process id, or null (creating nothing) if the table is full or
/// out of memory.
/// **Caller must hold the kernel lock** (the loader that builds `aspace` holds it /// **Caller must hold the kernel lock** (the loader that builds `aspace` holds it
/// across the whole spawn, so the address space and the task appear atomically). /// across the whole spawn, so the address space and the task appear atomically).
pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, priority: Priority, task_name: []const u8) bool { pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, priority: Priority, task_name: []const u8, supervisor: u32, exit_endpoint: ?*anyopaque) ?u32 {
const t = freeSlot() orelse return false; const t = freeSlot() orelse return null;
const stack = heap.allocator().alloc(u8, stack_size) catch return false; const stack = heap.allocator().alloc(u8, stack_size) catch return null;
t.* = .{ t.* = .{
.id = next_id, .id = next_id,
.state = .ready, .state = .ready,
@@ -291,6 +317,8 @@ pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, priority: Priority
.aspace = aspace, .aspace = aspace,
.user_ip = entry, .user_ip = entry,
.user_sp = user_sp, .user_sp = user_sp,
.supervisor = supervisor,
.exit_endpoint = exit_endpoint,
}; };
const name_length = @min(task_name.len, maximum_task_name); const name_length = @min(task_name.len, maximum_task_name);
@memcpy(t.name_buffer[0..name_length], task_name[0..name_length]); @memcpy(t.name_buffer[0..name_length], task_name[0..name_length]);
@@ -302,7 +330,7 @@ pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, priority: Priority
// the user entry/stack from the Task itself). // the user entry/stack from the Task itself).
t.sp = architecture.initTaskStack(top, @intFromPtr(&startUserTask)); t.sp = architecture.initTaskStack(top, @intFromPtr(&startUserTask));
enqueue(t); enqueue(t);
return true; return t.id;
} }
/// The first thing a fresh user task runs (in ring 0, via task_trampoline). It /// The first thing a fresh user task runs (in ring 0, via task_trampoline). It
@@ -414,6 +442,7 @@ pub const WaitQueue = struct {
pub fn waitLocked(wait_queue: *WaitQueue) void { pub fn waitLocked(wait_queue: *WaitQueue) void {
const t = current(); const t = current();
t.state = .blocked; t.state = .blocked;
t.wait_queue = wait_queue; // so a kill can unlink a parked waiter
t.next = wait_queue.head; t.next = wait_queue.head;
wait_queue.head = t; wait_queue.head = t;
schedule(); schedule();
@@ -438,10 +467,78 @@ pub fn wakeLocked(wait_queue: *WaitQueue) void {
} }
const t = best orelse return; const t = best orelse return;
if (best_previous) |p| p.next = t.next else wait_queue.head = t.next; if (best_previous) |p| p.next = t.next else wait_queue.head = t.next;
t.wait_queue = null;
t.state = .ready; t.state = .ready;
enqueue(t); enqueue(t);
} }
/// Unlink `t` from the wait queue it is parked on, if any (the kill path — a
/// killed waiter must not be woken later as a dangling pointer). Precondition:
/// the big kernel lock is held.
pub fn removeFromWaitQueueLocked(t: *Task) void {
const wait_queue = t.wait_queue orelse return;
t.wait_queue = null;
var previous: ?*Task = null;
var node = wait_queue.head;
while (node) |n| : ({
previous = n;
node = n.next;
}) {
if (n != t) continue;
if (previous) |p| p.next = t.next else wait_queue.head = t.next;
t.next = null;
return;
}
}
/// Unlink `t` from the ready queue it sits in (global, or its affinity core's
/// pinned queue) — the kill path for a task that is runnable but not running.
/// Precondition: the big kernel lock is held.
pub fn removeFromReadyQueueLocked(t: *Task) void {
if (t.affinity) |cpu| {
const pc = &cpus[cpu];
removeFrom(&pc.pinned_head, &pc.pinned_tail, &pc.pinned_bitmap, t);
} else {
removeFrom(&ready_head, &ready_tail, &ready_bitmap, t);
}
}
fn removeFrom(head: *[number_priorities]?*Task, tail: *[number_priorities]?*Task, bitmap: *u8, t: *Task) void {
const level: usize = t.priority;
var previous: ?*Task = null;
var node = head[level];
while (node) |n| : ({
previous = n;
node = n.next;
}) {
if (n != t) continue;
if (previous) |p| p.next = t.next else head[level] = t.next;
if (tail[level] == t) tail[level] = previous;
if (head[level] == null) bitmap.* &= ~(@as(u8, 1) << @intCast(level));
t.next = null;
return;
}
}
/// Find a live task by process id, or null. Ids are monotonic and never reused,
/// so a stale id misses cleanly rather than naming a recycled slot.
/// Precondition: the big kernel lock is held.
pub fn taskByIdLocked(id: u32) ?*Task {
for (&tasks) |*t| {
if (t.state != .free and t.id == id) return t;
}
return null;
}
/// Make every server that still holds `t` as the client it owes a reply to forget
/// it — the reply of a dead client is dropped, not delivered into freed state.
/// Precondition: the big kernel lock is held.
pub fn forgetIpcClientLocked(t: *Task) void {
for (&tasks) |*other| {
if (other.state != .free and other.ipc_client == t) other.ipc_client = null;
}
}
/// Block the current task and switch away, without putting it on any wait queue — /// Block the current task and switch away, without putting it on any wait queue —
/// the caller has already linked it wherever it belongs (e.g. an endpoint's sender /// the caller has already linked it wherever it belongs (e.g. an endpoint's sender
/// FIFO). Precondition: the big kernel lock is held; still held on return (when the /// FIFO). Precondition: the big kernel lock is held; still held on return (when the
@@ -501,14 +598,55 @@ fn wakeExpired() void {
} }
} }
// Process-teardown hooks, registered by process.zig at init — the scheduler sits
// below the process layer, so finishing a kill (IRQ bindings, IPC handles, exit
// notification) is called *up* through these, mirroring how the architecture
// layer calls up into `tick`.
//
// `terminate_current_hook` ends the task running on THIS core (lock held, never
// returns — it switches away like `exitUserLocked`). `reap_task_hook` tears down
// a task that is NOT running on any core (lock held).
pub var terminate_current_hook: ?*const fn () noreturn = null;
pub var reap_task_hook: ?*const fn (*Task) void = null;
/// Finish any pending kills this core can see (the deferred half of process_kill;
/// the immediate half runs in the killer's own call). Precondition: the big kernel
/// lock is held, from `tick`.
///
/// - This core's *current* task, if condemned, is terminated here — but only when
/// it is not inside one of its own system calls (`in_system_call`): the tick may
/// have interrupted kernel code mid-operation, where teardown would leak or
/// corrupt what that operation holds. User-mode execution (and the system_call
/// entry/exit stubs, which hold nothing) are safe termination points. A task
/// that *is* mid-call dies at its next block, tick, or system_call entry instead.
/// The hook never returns; abandoning the interrupt frame is fine — the LAPIC
/// was acknowledged before the tick hook ran (see apic.timerTick), exactly as on
/// the fault-kill path.
/// - Condemned tasks that are ready or blocked are not running anywhere (state
/// changes need the lock we hold), so they are reaped in place.
fn reapKillPendingLocked() void {
const pc = thisCpu();
const cur = pc.current;
if (cur.kill_pending and cur.aspace != 0 and !cur.in_system_call) {
if (terminate_current_hook) |hook| hook(); // noreturn
}
if (reap_task_hook) |hook| {
for (&tasks) |*t| {
if (!t.kill_pending) continue;
if (t.state == .ready or t.state == .blocked) hook(t);
}
}
}
/// Called from the timer interrupt (interrupts already disabled): wake due /// Called from the timer interrupt (interrupts already disabled): wake due
/// sleepers, then preempt. Takes the kernel lock like any other critical section, /// sleepers, finish pending kills, then preempt. Takes the kernel lock like any
/// but releases it *without* touching the interrupt flag — the handler's `iretq` /// other critical section, but releases it *without* touching the interrupt flag
/// restores the interrupted context's flags, so re-enabling here would open a /// — the handler's `iretq` restores the interrupted context's flags, so
/// nested-interrupt window before the return. /// re-enabling here would open a nested-interrupt window before the return.
pub fn tick() void { pub fn tick() void {
_ = sync.enter(); _ = sync.enter();
wakeExpired(); wakeExpired();
reapKillPendingLocked();
if (preemption_enabled) schedule(); if (preemption_enabled) schedule();
sync.leaveIsr(); sync.leaveIsr();
} }
@@ -540,6 +678,13 @@ pub fn exit() noreturn {
/// itself is leaked, as in `exit` (no reaper yet). Never returns. /// itself is leaked, as in `exit` (no reaper yet). Never returns.
pub fn exitUser() noreturn { pub fn exitUser() noreturn {
_ = sync.enter(); _ = sync.enter();
exitUserLocked();
}
/// The body of `exitUser` for callers that already hold the big kernel lock (the
/// tick-time terminate path, which enters with the lock held). The lock is handed
/// off through the switch and released by the task that resumes. Never returns.
pub fn exitUserLocked() noreturn {
const pc = thisCpu(); const pc = thisCpu();
const dying = pc.current; const dying = pc.current;
const as = dying.aspace; const as = dying.aspace;
@@ -551,6 +696,8 @@ pub fn exitUser() noreturn {
} }
dying.state = .free; dying.state = .free;
dying.aspace = 0; dying.aspace = 0;
dying.kill_pending = false;
dying.in_system_call = false;
const next = dequeueHighest(pc) orelse @panic("sched: no task left to run"); const next = dequeueHighest(pc) orelse @panic("sched: no task left to run");
next.state = .running; next.state = .running;
pc.current = next; pc.current = next;
@@ -559,6 +706,54 @@ pub fn exitUser() noreturn {
unreachable; unreachable;
} }
/// Free a task that is NOT running on any core (it is ready or blocked, and the
/// caller — the kill path — has already unlinked it from every queue and released
/// what it held). Destroys its address space: safe here because no core can have
/// it loaded (every switch away from a task loads the next task's tables, and the
/// task isn't running). The kernel stack is leaked, as in `exitUser` (no reaper
/// yet). Precondition: the big kernel lock is held.
pub fn destroyTaskLocked(t: *Task) void {
if (t.aspace != 0) architecture.destroyAddressSpace(t.aspace);
t.aspace = 0;
t.kill_pending = false;
t.in_system_call = false;
t.wake_at = 0;
t.state = .free;
}
/// Snapshot the task table into `out` (up to its length), returning the total
/// number of live tasks — the kernel half of `process_enumerate`, mirroring
/// devices_broker.enumerate. Kernel tasks are included (empty name, supervisor 0):
/// an honest `ps` shows the idle tasks too. `out` may be user memory: the caller's
/// address space is loaded during its system call, and the same bring-up trust
/// applies as for device_enumerate (an unmapped user page faults the kernel).
pub fn enumerate(out: []abi.ProcessDescriptor) u64 {
const flags = sync.enter();
defer sync.leave(flags);
var total: u64 = 0;
for (&tasks) |*t| {
if (t.state == .free) continue;
if (total < out.len) {
const d = &out[total];
d.* = .{
.id = t.id,
.supervisor = t.supervisor,
.state = @intFromEnum(@as(abi.ProcessState, switch (t.state) {
.ready => .ready,
.running => .running,
.blocked => .blocked,
.free => unreachable,
})),
.priority = t.priority,
.name_length = t.name_length,
.name = t.name_buffer,
};
}
total += 1;
}
return total;
}
/// Whether the running task is a user process (has its own address space). /// Whether the running task is a user process (has its own address space).
pub fn currentIsUserProcess() bool { pub fn currentIsUserProcess() bool {
return current().aspace != 0; return current().aspace != 0;
+190 -1
View File
@@ -126,6 +126,12 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
initTest(boot_information); initTest(boot_information);
} else if (eql(case, "process")) { } else if (eql(case, "process")) {
processTest(boot_information); processTest(boot_information);
} else if (eql(case, "process-list")) {
processListTest(boot_information);
} else if (eql(case, "process-kill")) {
processKillTest(boot_information);
} else if (eql(case, "supervision")) {
supervisionTest(boot_information);
} else if (eql(case, "initial-ramdisk")) { } else if (eql(case, "initial-ramdisk")) {
initialRamdiskTest(boot_information); initialRamdiskTest(boot_information);
} else if (eql(case, "vfs")) { } else if (eql(case, "vfs")) {
@@ -1222,7 +1228,7 @@ fn spawnFaultingProcess() bool {
}; };
architecture.mapUserPageInto(aspace, process.stack_base_virtual, stack_frame, true, false); // RW + NX architecture.mapUserPageInto(aspace, process.stack_base_virtual, stack_frame, true, false); // RW + NX
if (!scheduler.spawnUserLocked(aspace, process.code_virtual, process.stack_base_virtual + abi.page_size, 4, "fault-probe")) { if (scheduler.spawnUserLocked(aspace, process.code_virtual, process.stack_base_virtual + abi.page_size, 4, "fault-probe", 0, null) == null) {
architecture.destroyAddressSpace(aspace); architecture.destroyAddressSpace(aspace);
return false; return false;
} }
@@ -1311,6 +1317,189 @@ fn initTest(boot_information: *const BootInformation) void {
result(); result();
} }
/// process_enumerate's kernel half: spawn two init processes next to the kernel
/// tasks and snapshot the table. The snapshot must list both by name with distinct,
/// kernel-supervised ids, include the kernel tasks (id 0, empty name), and report
/// the same total through a too-small buffer (the truncation contract: the caller
/// learns how big a buffer to bring).
fn processListTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: process-list\n", .{});
check("bootloader handed over /system/services/init", boot_information.init_len != 0);
if (boot_information.init_len == 0) {
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len];
var spawned: u32 = 0;
if (process.spawnProcess(image, 4, &.{"/system/services/init"})) spawned += 1 else |_| {}
if (process.spawnProcess(image, 4, &.{"/system/services/init"})) spawned += 1 else |_| {}
check("two init processes spawned", spawned == 2);
var table: [32]abi.ProcessDescriptor = undefined;
const total = scheduler.enumerate(&table);
check("enumerate counts the boot task and both processes (>=3)", total >= 3);
var inits: u32 = 0;
var init_ids: [2]u32 = .{ 0, 0 };
var kernel_task_seen = false;
var states_sane = true;
for (table[0..@min(total, table.len)]) |descriptor| {
if (descriptor.state > @intFromEnum(abi.ProcessState.blocked)) states_sane = false;
if (descriptor.name_length == 0) kernel_task_seen = true;
if (eql(descriptor.name[0..descriptor.name_length], "/system/services/init")) {
if (inits < 2) init_ids[inits] = descriptor.id;
inits += 1;
check("init entry is kernel-supervised (supervisor 0)", descriptor.supervisor == 0);
}
}
check("both init processes listed by name", inits == 2);
check("listed processes carry distinct ids", init_ids[0] != init_ids[1]);
check("kernel tasks are listed too (empty name)", kernel_task_seen);
check("every state is a ProcessState value", states_sane);
var one: [1]abi.ProcessDescriptor = undefined;
check("a too-small buffer still learns the true total", scheduler.enumerate(&one) == total);
result();
}
/// process_kill + the exit notification, kernel half. Two victims, two paths:
/// - init, which heartbeats and sleeps: caught blocked, reaped on the killer's
/// own call — and its heartbeat must stop.
/// - process-test's spinner role (from the initial ramdisk), which loops in user
/// mode making no system calls: with more cores it is caught running, taking
/// the deferred path (kill_pending, finished by the victim core's next tick).
/// Each death must post one exit notification badge (exit bit + the child's id)
/// on the endpoint given at spawn; wrong-supervisor and unknown-id kills must be
/// refused. The waits block in replyWait, so a lost notification times the
/// harness out rather than passing vacuously.
fn processKillTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: process-kill\n", .{});
check("bootloader handed over /system/services/init", boot_information.init_len != 0);
if (boot_information.init_len == 0 or boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", boot_information.initial_ramdisk_len != 0);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len];
const ramdisk = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(ramdisk) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
const me = scheduler.currentId();
const endpoint = ipcsync.createIpcEndpoint() orelse {
check("exit endpoint allocated", false);
result();
return;
};
process.write_count = 0;
const sleeper = process.spawnProcessSupervised(image, 4, &.{"/system/services/init"}, me, endpoint) catch 0;
check("init spawned as the supervised sleeper victim", sleeper != 0);
// Let it reach its heartbeat loop (write, then a 1 s sleep) so the kill most
// likely catches it blocked.
scheduler.setPriority(1);
const deadline = architecture.millis() + 8000;
while (process.write_count < 1 and architecture.millis() < deadline) scheduler.yield();
scheduler.setPriority(4);
check("victim heartbeat before the kill", process.write_count >= 1);
// Kills that must be refused, before the one that must not be.
check("a non-supervisor may not kill (-EPERM)", process.killProcess(me + 12345, sleeper) == -ipcsync.EPERM);
check("an unknown id misses (-ESRCH)", process.killProcess(me, 0xFFFF_FF00) == -ipcsync.ESRCH);
check("a kernel task is not a killable process (-ESRCH)", process.killProcess(me, 0) == -ipcsync.ESRCH);
check("the supervisor's kill is accepted", process.killProcess(me, sleeper) == 0);
var badge: u64 = 0;
var received_cap: u64 = 0;
var r = ipcsync.replyWait(endpoint, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap);
check("the sleeper's exit notification arrived (length 0)", r == 0);
check("its badge carries the exit bit and the child id", badge == abi.notify_badge_bit | abi.notify_exit_bit | sleeper);
const beats_at_kill = process.write_count;
scheduler.sleep(1500); // more than one heartbeat period
check("the heartbeat stopped with the kill", process.write_count == beats_at_kill);
// The spinner: no system calls, so only the tick can deliver a deferred kill.
var spinner: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "process-test")) continue;
spinner = process.spawnProcessSupervised(item.blob, 4, &.{ "process-test", "spinner" }, me, endpoint) catch 0;
break;
}
check("process-test spawned as the supervised spinner victim", spinner != 0);
scheduler.sleep(100); // give another core a chance to be running it
check("the spinner's kill is accepted", process.killProcess(me, spinner) == 0);
r = ipcsync.replyWait(endpoint, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap);
check("the spinner's exit notification arrived (length 0)", r == 0);
check("its badge carries the exit bit and the child id", badge == abi.notify_badge_bit | abi.notify_exit_bit | spinner);
var table: [32]abi.ProcessDescriptor = undefined;
const total = scheduler.enumerate(&table);
var still_listed = false;
for (table[0..@min(total, table.len)]) |descriptor| {
if (descriptor.id == sleeper or descriptor.id == spinner) still_listed = true;
}
check("neither victim is listed after its kill", !still_listed);
check("a killed id stays dead (-ESRCH on a second kill)", process.killProcess(me, sleeper) == -ipcsync.ESRCH);
result();
}
/// The whole user-side surface at once: spawn process-test's supervisor role,
/// which — entirely from ring 3 — creates an exit endpoint, spawns its two
/// children supervised, sees them in process_enumerate, kills them (one blocked,
/// one spinning), collects both exit notifications, and confirms they are gone.
/// Its "process-test: ok" is the pass marker; any FAIL line is specific.
fn supervisionTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: supervision\n", .{});
check("bootloader handed over an initial_ramdisk", boot_information.initial_ramdisk_len != 0);
if (boot_information.initial_ramdisk_len == 0) {
result();
return;
}
const ramdisk = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(ramdisk) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(ramdisk); // the supervisor system_spawns its children by name
process.write_count = 0;
process.write_from_user = false;
var started = false;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(item.name, "process-test")) continue;
started = if (process.spawnProcess(item.blob, 4, &.{ "process-test", "run" })) true else |_| false;
break;
}
check("process-test spawned as the user-space supervisor", started);
const marker = "process-test: ok";
scheduler.setPriority(1);
const deadline = architecture.millis() + 10000;
while (architecture.millis() < deadline) {
if (process.write_len >= marker.len and eql(process.write_buffer[0..marker.len], marker)) break;
scheduler.yield();
}
scheduler.setPriority(4);
const ok = process.write_len >= marker.len and eql(process.write_buffer[0..marker.len], marker);
if (!ok and process.write_len > 0) log("DANOS-SUPERVISION: got \"{s}\"\n", .{process.write_buffer[0..process.write_len]});
check("the supervisor completed every step (spawn/list/kill/notify)", ok);
check("it ran in user mode (CPL 3)", process.write_from_user);
result();
}
/// The initial_ramdisk path: the bootloader handed over an image bundling extra user /// The initial_ramdisk path: the bootloader handed over an image bundling extra user
/// binaries; parse it, spawn every program, and confirm one (the vfs stub) /// binaries; parse it, spawn every program, and confirm one (the vfs stub)
/// reaches ring 3 and heartbeats — proving the whole ferry-parse-spawn pipeline. /// reaches ring 3 and heartbeats — proving the whole ferry-parse-spawn pipeline.
@@ -38,7 +38,7 @@ pub fn main() void {
for (buffer[0..count]) |descriptor| { for (buffer[0..count]) |descriptor| {
const driver_name = driverFor(descriptor.class) orelse continue; const driver_name = driverFor(descriptor.class) orelse continue;
matched += 1; matched += 1;
if (runtime.system.spawn(driver_name)) { if (runtime.system.spawn(driver_name) != null) {
_ = runtime.system.write("device-manager: spawned "); _ = runtime.system.write("device-manager: spawned ");
_ = runtime.system.write(driver_name); _ = runtime.system.write(driver_name);
_ = runtime.system.write("\n"); _ = runtime.system.write("\n");
@@ -0,0 +1,100 @@
//! process-test — a test fixture for process management (bundled in the
//! initial-ramdisk, driven by the `supervision` test case). One binary, three
//! roles picked by argv, so the whole user-side surface is exercised end to end:
//!
//! - `process-test run` — the supervisor: spawns the two children below with an
//! exit-notification endpoint, sees them in `process_enumerate`, kills them,
//! receives both exit notifications, and confirms they are gone. Prints
//! "process-test: ok" for the kernel test to match, or a FAIL line naming the
//! step that broke.
//! - `process-test sleeper` — a child that blocks in `sleep` forever: its kill
//! exercises the immediate reap of a blocked task.
//! - `process-test spinner` — a child that spins in user mode making no system
//! calls: its kill exercises the deferred path (kill_pending, finished by the
//! timer tick).
//!
//! Spawned with no arguments (the initial-ramdisk sweep test starts every bundled
//! binary bare), it exits silently so it cannot derange other tests' output.
const std = @import("std");
const runtime = @import("runtime");
fn fail(step: []const u8) noreturn {
_ = runtime.system.write("process-test: FAIL ");
_ = runtime.system.write(step);
_ = runtime.system.write("\n");
runtime.system.exit(1);
}
/// Whether process `id` appears in a fresh `process_enumerate` snapshot, named
/// `name` (an id present under the wrong name is a table mix-up, not a pass).
fn listed(id: u32, name: []const u8) bool {
var table: [32]runtime.system.ProcessDescriptor = undefined;
const total = runtime.system.processes(&table);
for (table[0..@min(total, table.len)]) |descriptor| {
if (descriptor.id != id) continue;
return std.mem.eql(u8, descriptor.name[0..descriptor.name_length], name);
}
return false;
}
/// Block on the exit endpoint until a child-exit notification arrives; returns
/// the ended child's id. A wrong wake-up (there should be none — nothing else
/// knows this endpoint) fails the test rather than looping forever.
fn awaitChildExit(endpoint: runtime.ipc.Handle) u32 {
var scratch: [8]u8 = undefined;
const received = runtime.ipc.replyWait(endpoint, scratch[0..0], &scratch, null);
if (!received.isChildExit()) fail("expected a child-exit notification");
return received.childProcessId();
}
pub fn main() void {
if (runtime.argumentCount() <= 1) return; // spawned bare (ramdisk sweep): stay silent
const role = runtime.argument(1);
if (std.mem.eql(u8, role, "sleeper")) {
while (true) runtime.system.sleep(500);
}
if (std.mem.eql(u8, role, "spinner")) {
var beat: u64 = 0;
const touch: *volatile u64 = &beat;
while (true) touch.* +%= 1; // user mode only — no system calls to die at
}
// The supervisor ("run").
const endpoint = runtime.ipc.createIpcEndpoint() orelse fail("create exit endpoint");
const sleeper = runtime.system.spawnSupervised("process-test", &.{"sleeper"}, endpoint) orelse fail("spawn sleeper");
const spinner = runtime.system.spawnSupervised("process-test", &.{"spinner"}, endpoint) orelse fail("spawn spinner");
runtime.system.sleep(100); // let the sleeper block and the spinner get a core
if (!listed(sleeper, "process-test")) fail("sleeper not in process_enumerate");
if (!listed(spinner, "process-test")) fail("spinner not in process_enumerate");
// Kills that must be refused: a kernel task (id 0), and an id that was never
// issued — both -ESRCH. (-EPERM needs a second supervisor; the kernel-level
// `process-kill` test covers it.)
if (runtime.system.kill(0)) fail("killing a kernel task was allowed");
if (runtime.system.kill(0xFFFF_FFF0)) fail("killing an unknown id was allowed");
// The blocked child: usually reaped on the spot (it sits in `sleep`). The
// notification is the fence — after it, the child is certainly gone, so the
// second kill must miss (its id is never reused).
if (!runtime.system.kill(sleeper)) fail("kill sleeper");
if (awaitChildExit(endpoint) != sleeper) fail("sleeper exit notification");
if (runtime.system.kill(sleeper)) fail("double kill was allowed");
// The running child: the deferred path — condemned now, dead by the next tick.
if (!runtime.system.kill(spinner)) fail("kill spinner");
if (awaitChildExit(endpoint) != spinner) fail("spinner exit notification");
if (listed(sleeper, "process-test")) fail("sleeper still listed after kill");
if (listed(spinner, "process-test")) fail("spinner still listed after kill");
_ = runtime.system.write("process-test: ok\n");
}
pub const panic = runtime.panic;
comptime {
_ = &runtime.start._start; // pull the runtime entry shim into the image
}
+18
View File
@@ -222,6 +222,24 @@ CASES = [
"smp": 4, "smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS", "expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# process_enumerate: a task-table snapshot lists spawned processes by name and
# id alongside the kernel tasks, and a too-small buffer still reports the total.
{"name": "process-list",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# process_kill + the exit notification: only the supervisor may kill; a blocked
# victim is reaped in place and a spinning one dies by the deferred (tick) path;
# each death posts one exit badge to the endpoint given at spawn.
{"name": "process-kill",
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# The user-side whole: process-test supervises two children from ring 3 —
# spawn with an exit endpoint, enumerate, kill, notification, gone.
{"name": "supervision",
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# The initial_ramdisk: the loader ferries a bundle of user binaries; the kernel parses # The initial_ramdisk: the loader ferries a bundle of user binaries; the kernel parses
# it and spawns each as a ring-3 process (here the VFS-server stub heartbeats). # it and spawns each as a ring-3 process (here the VFS-server stub heartbeats).
{"name": "initial-ramdisk", {"name": "initial-ramdisk",