Add process management: enumerate, supervisor-gated kill, exit notifications
process_enumerate snapshots the task table (the device_enumerate shape, so ps is a user program); system_spawn returns the child id, records the caller as supervisor, and takes an exit endpoint; process_kill is allowed only for the supervisor. Every death — exit, fault, or kill — posts a child-exit badge to that endpoint (the IRQ-as-IPC pattern as SIGCHLD). A target caught off-CPU is reaped in place; a running one is condemned and finished at its next system call or tick, guarded so teardown never lands mid-kernel-operation. Tested by process-list, process-kill, and supervision (a ring-3 supervisor exercising the whole surface); design notes in docs/process-management.md.
This commit is contained in:
@@ -300,6 +300,7 @@ pub fn build(b: *std.Build) void {
|
|||||||
const bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, "bus", "system/drivers/bus/bus.zig");
|
const bus_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, "bus", "system/drivers/bus/bus.zig");
|
||||||
const device_manager_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, "device-manager", "system/services/device-manager/device-manager.zig");
|
const device_manager_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, "device-manager", "system/services/device-manager/device-manager.zig");
|
||||||
const args_echo_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, "args-echo", "system/services/args-echo/args-echo.zig");
|
const args_echo_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, "args-echo", "system/services/args-echo/args-echo.zig");
|
||||||
|
const process_test_exe = addUserBinary(b, kernel_target, runtime_module, posix_module, mmio_module, "process-test", "system/services/process-test/process-test.zig");
|
||||||
|
|
||||||
// Pack the user binaries into the initial_ramdisk image with the host-side Python tool
|
// Pack the user binaries into the initial_ramdisk image with the host-side Python tool
|
||||||
// (the container format is trivial, and Python sidesteps std API churn). Args:
|
// (the container format is trivial, and Python sidesteps std API churn). Args:
|
||||||
@@ -319,6 +320,8 @@ pub fn build(b: *std.Build) void {
|
|||||||
mk_run.addFileArg(device_manager_exe.getEmittedBin());
|
mk_run.addFileArg(device_manager_exe.getEmittedBin());
|
||||||
mk_run.addArg("args-echo");
|
mk_run.addArg("args-echo");
|
||||||
mk_run.addFileArg(args_echo_exe.getEmittedBin());
|
mk_run.addFileArg(args_echo_exe.getEmittedBin());
|
||||||
|
mk_run.addArg("process-test");
|
||||||
|
mk_run.addFileArg(process_test_exe.getEmittedBin());
|
||||||
|
|
||||||
// Also install the packed binaries to their FHS homes, so zig-out is a true image
|
// Also install the packed binaries to their FHS homes, so zig-out is a true image
|
||||||
// of the filesystem — even though at boot they arrive inside the initial-ramdisk.
|
// of the filesystem — even though at boot they arrive inside the initial-ramdisk.
|
||||||
|
|||||||
+5
-1
@@ -48,7 +48,11 @@ rather than restate it. Roughly in the order things happen at runtime:
|
|||||||
real driver stacks factor into three shapes, how families share code, and the
|
real driver stacks factor into three shapes, how families share code, and the
|
||||||
proposed ABI for the three primitives still missing (capability passing, DMA +
|
proposed ABI for the three primitives still missing (capability passing, DMA +
|
||||||
memory barriers, MSI).
|
memory barriers, MSI).
|
||||||
15. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
|
15. **[process-management.md](process-management.md) — process management.** The
|
||||||
|
microkernel's `ps`/`kill`/SIGCHLD: enumerate as a table snapshot, the
|
||||||
|
supervision link as the kill authority, and child-exit notifications over the
|
||||||
|
same endpoints IRQs arrive on.
|
||||||
|
16. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
|
||||||
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
|
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
|
||||||
|
|
||||||
Start with the north star:
|
Start with the north star:
|
||||||
|
|||||||
@@ -0,0 +1,112 @@
|
|||||||
|
# Process Management
|
||||||
|
|
||||||
|
How danos lists, supervises, and kills processes — the microkernel answer to
|
||||||
|
`ps`, `kill`, and `SIGCHLD`/`wait`.
|
||||||
|
|
||||||
|
## Why system calls, not `/proc`
|
||||||
|
|
||||||
|
Unix systems sit on a spectrum. Classic BSD/macOS list processes through
|
||||||
|
syscalls (`sysctl(KERN_PROC)`) and kill through `kill(2)`; Linux renders the
|
||||||
|
process table as `/proc` for *reading* but still kills through a syscall; Plan 9
|
||||||
|
made the file tree the whole interface (`echo kill > /proc/n/ctl`). Microkernels
|
||||||
|
mostly abandon ambient PIDs: Minix and QNX route everything through a user-space
|
||||||
|
process-manager server, and Fuchsia/seL4 control processes only through handles.
|
||||||
|
|
||||||
|
danos rules out `/proc` **as the primitive**: here a `/proc` would be served by
|
||||||
|
the VFS server — a user process — which would put the VFS in the path of process
|
||||||
|
control. If the VFS (or anything under it) hangs, nothing could be listed or
|
||||||
|
killed, *including the hung VFS*. The control plane for processes must not
|
||||||
|
depend on a process. So the primitives are kernel system calls; a read-only
|
||||||
|
`/proc` rendering can be layered on later, and a POSIX-style process-manager
|
||||||
|
server can be built *from* these primitives when one is needed.
|
||||||
|
|
||||||
|
## The three primitives
|
||||||
|
|
||||||
|
### `process_enumerate(buffer, maximum) -> total`
|
||||||
|
|
||||||
|
A snapshot of the task table into a caller buffer of `abi.ProcessDescriptor`
|
||||||
|
(id, supervisor, state, priority, name) — the exact shape of
|
||||||
|
`device_enumerate`, so `ps` is a user program over a snapshot, not a kernel
|
||||||
|
service. The total may exceed what fit; call again with a larger buffer. Kernel
|
||||||
|
tasks are included with an empty name — an honest listing shows the idle tasks
|
||||||
|
too. Ungated and read-only: what is running is not a secret between cooperating
|
||||||
|
bring-up processes.
|
||||||
|
|
||||||
|
### `system_spawn(..., exit_endpoint) -> child id`, and the supervision link
|
||||||
|
|
||||||
|
`system_spawn` records the caller as the child's **supervisor** and returns the
|
||||||
|
child's process id (ids are monotonic, never reused — a stale id can only miss).
|
||||||
|
That link is the kill authority: it answers "who may kill process 7?" without
|
||||||
|
inventing users or permissions, the same way a device *claim* is the capability
|
||||||
|
for `mmio_map`. It composes with the supervision hierarchy the device manager
|
||||||
|
already forms: init supervises the services it starts, the device manager
|
||||||
|
supervises the drivers it matches. (A transferable process *handle* — Fuchsia
|
||||||
|
style — can replace the id once the handle table grows types beyond endpoints.)
|
||||||
|
|
||||||
|
`exit_endpoint` (a handle, or `abi.no_cap`) is the supervisor's death-watch: when
|
||||||
|
the child ends — clean exit, CPU fault, or `process_kill` — the kernel posts an
|
||||||
|
asynchronous notification to that endpoint, exactly like a bound IRQ. The badge
|
||||||
|
carries `abi.notify_badge_bit | abi.notify_exit_bit | child_id`, so one endpoint
|
||||||
|
supervises many children and can even share with IRQ notifications. This is the
|
||||||
|
microkernel's SIGCHLD: no new mechanism, just the IRQ-as-IPC pattern reused, and
|
||||||
|
a supervisor's event loop (`ipc.replyWait`) already knows how to receive it. The
|
||||||
|
child holds a reference to the endpoint from birth, so the notification cannot
|
||||||
|
dangle even if the supervisor dies first.
|
||||||
|
|
||||||
|
### `process_kill(id) -> 0 / -ESRCH / -EPERM`
|
||||||
|
|
||||||
|
Only the supervisor may kill; kernel tasks are not killable processes. Like a
|
||||||
|
signal, delivery is prompt but asynchronous — 0 means the kill is accepted and
|
||||||
|
irrevocable; the exit notification confirms completion.
|
||||||
|
|
||||||
|
## How a kill lands (the kernel mechanics)
|
||||||
|
|
||||||
|
Everything below runs under the big kernel lock, where task states cannot move.
|
||||||
|
|
||||||
|
- **Target ready or blocked** (not on any core): reaped on the killer's own
|
||||||
|
call. The reap releases what death always releases (IRQ bindings first, then
|
||||||
|
a client the target still owed a reply to is failed with `-EPEER`, IPC handles
|
||||||
|
closed, the exit notification posted last) — plus the unlinking only a
|
||||||
|
*remote* death needs: out of the ready queue, out of an endpoint's sender FIFO
|
||||||
|
(`Task.ipc_wait_endpoint`), out of a receive wait queue (`Task.wait_queue`),
|
||||||
|
and out of any server's owed-reply slot, so nothing ever dequeues a dangling
|
||||||
|
pointer. Destroying the address space is safe because no core can have it
|
||||||
|
loaded: every switch away from a task loads the next task's tables.
|
||||||
|
- **Target running on another core**: it cannot be torn down mid-instruction,
|
||||||
|
so it is condemned (`Task.kill_pending`) and dies at whichever comes first:
|
||||||
|
- its next **system_call entry** — checked before dispatch, so a condemned
|
||||||
|
process cannot spawn, claim, or message anything on its way out;
|
||||||
|
- its core's next **timer tick** — but only when the task is not inside one
|
||||||
|
of its own system calls (`Task.in_system_call`): the tick may have
|
||||||
|
interrupted kernel code mid-operation, where teardown would leak whatever
|
||||||
|
the operation held. User-mode execution is always a safe kill point. The
|
||||||
|
tick-time terminate abandons the interrupt frame exactly like the fault
|
||||||
|
path (the LAPIC is acknowledged before the tick hook runs);
|
||||||
|
- any core's tick finding it **blocked or ready** (it entered a syscall and
|
||||||
|
parked after being condemned) — reaped by the same remote-reap path.
|
||||||
|
|
||||||
|
A pure user-mode spin loop that never makes a system call therefore dies
|
||||||
|
within one tick; nothing a process does can outrun the kill.
|
||||||
|
|
||||||
|
The scheduler stays below the process layer: finishing a kill (IRQ bindings,
|
||||||
|
handles, the notification) is called *up* through two hooks process.zig
|
||||||
|
registers at boot (`terminate_current_hook`, `reap_task_hook`), mirroring how
|
||||||
|
the architecture layer calls up into `tick`.
|
||||||
|
|
||||||
|
## Known gaps (bring-up honesty)
|
||||||
|
|
||||||
|
- Device **claims** are not released on death (pre-existing: the fault path has
|
||||||
|
the same gap) — a killed driver's device stays claimed until reboot.
|
||||||
|
- Kernel stacks of dead tasks are leaked, as on every exit path (no reaper yet).
|
||||||
|
- There is no exit *status* in the notification, only the id; a supervisor that
|
||||||
|
needs the code can grow a wait-style call later.
|
||||||
|
- Enumerate writes through the caller's raw pointer under the bring-up trust
|
||||||
|
model, like `device_enumerate` (an unmapped page is a self-DoS, not an
|
||||||
|
isolation break).
|
||||||
|
|
||||||
|
## Tests
|
||||||
|
|
||||||
|
`process-list` (enumerate), `process-kill` (kernel-level kill paths, refusals,
|
||||||
|
notifications), `supervision` (the whole user-side surface via the process-test
|
||||||
|
service: spawn supervised → enumerate → kill blocked and spinning children →
|
||||||
|
notifications → gone). See test/qemu_test.py.
|
||||||
+21
-3
@@ -84,6 +84,11 @@ pub fn call(h: Handle, message: []const u8, reply: []u8) CallError!usize {
|
|||||||
/// GSI. See `isNotification`.
|
/// GSI. See `isNotification`.
|
||||||
pub const notify_badge_bit: u64 = abi.notify_badge_bit;
|
pub const notify_badge_bit: u64 = abi.notify_badge_bit;
|
||||||
|
|
||||||
|
/// Set alongside `notify_badge_bit` when the notification is a **child-exit
|
||||||
|
/// notice** — a process this one spawned (with an exit endpoint) has ended —
|
||||||
|
/// rather than a device interrupt. The low bits carry the child's process id.
|
||||||
|
pub const notify_exit_bit: u64 = abi.notify_exit_bit;
|
||||||
|
|
||||||
/// The result of a `replyWait`: the request length, the sender's badge (a task id, or
|
/// The result of a `replyWait`: the request length, the sender's badge (a task id, or
|
||||||
/// an IRQ notification if the high bit is set), and any capability the request carried.
|
/// an IRQ notification if the high bit is set), and any capability the request carried.
|
||||||
pub const Received = struct {
|
pub const Received = struct {
|
||||||
@@ -91,16 +96,29 @@ pub const Received = struct {
|
|||||||
badge: u64,
|
badge: u64,
|
||||||
cap: ?Handle,
|
cap: ?Handle,
|
||||||
|
|
||||||
/// True if this wake-up was a device interrupt, not a client request. A driver's
|
/// True if this wake-up was an asynchronous notification (a device interrupt
|
||||||
/// event loop branches on this; there is no reply owed on the notification path.
|
/// or a child-exit notice), not a client request. An event loop branches on
|
||||||
|
/// this; there is no reply owed on the notification path.
|
||||||
pub fn isNotification(self: Received) bool {
|
pub fn isNotification(self: Received) bool {
|
||||||
return self.badge & notify_badge_bit != 0;
|
return self.badge & notify_badge_bit != 0;
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The interrupt source (a GSI), meaningful only when `isNotification`.
|
/// True if this wake-up tells of a supervised child's end — the notification
|
||||||
|
/// requested by passing an exit endpoint to `system.spawnSupervised`.
|
||||||
|
pub fn isChildExit(self: Received) bool {
|
||||||
|
return self.isNotification() and self.badge & notify_exit_bit != 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The interrupt source (a GSI), meaningful only when `isNotification` and
|
||||||
|
/// not `isChildExit`.
|
||||||
pub fn source(self: Received) u64 {
|
pub fn source(self: Received) u64 {
|
||||||
return self.badge & ~notify_badge_bit;
|
return self.badge & ~notify_badge_bit;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// The ended child's process id, meaningful only when `isChildExit`.
|
||||||
|
pub fn childProcessId(self: Received) u32 {
|
||||||
|
return @intCast(self.badge & ~(notify_badge_bit | notify_exit_bit));
|
||||||
|
}
|
||||||
};
|
};
|
||||||
|
|
||||||
/// Server side of IPC_ReplyWait: deliver `reply` to the client last received (if any,
|
/// Server side of IPC_ReplyWait: deliver `reply` to the client last received (if any,
|
||||||
|
|||||||
+49
-12
@@ -11,6 +11,10 @@ pub const PROT_READ: usize = abi.prot_read;
|
|||||||
pub const PROT_WRITE: usize = abi.prot_write;
|
pub const PROT_WRITE: usize = abi.prot_write;
|
||||||
pub const PROT_EXEC: usize = abi.prot_exec;
|
pub const PROT_EXEC: usize = abi.prot_exec;
|
||||||
|
|
||||||
|
/// One `processes` entry — re-exported from the shared ABI so a user program can
|
||||||
|
/// declare its snapshot buffer without importing `abi` itself.
|
||||||
|
pub const ProcessDescriptor = abi.ProcessDescriptor;
|
||||||
|
|
||||||
/// Give up the rest of this quantum.
|
/// Give up the rest of this quantum.
|
||||||
pub fn yield() void {
|
pub fn yield() void {
|
||||||
_ = sc.systemCall0(.yield);
|
_ = sc.systemCall0(.yield);
|
||||||
@@ -44,31 +48,64 @@ pub fn exit(code: usize) noreturn {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/// Start the binary bundled in the initial-ramdisk under `name` as a new ring-3
|
/// Start the binary bundled in the initial-ramdisk under `name` as a new ring-3
|
||||||
/// process, returning true on success. The child's argv[0] is `name`. This is how
|
/// process, returning the child's process id (or null on failure). The child's
|
||||||
/// a supervisor (the device manager) launches a driver it matched — danos-native,
|
/// argv[0] is `name`, and the caller becomes its **supervisor** — the only process
|
||||||
/// not POSIX (a spawn/exec family comes with the process work later).
|
/// allowed to `kill` it. This is how a supervisor (the device manager) launches a
|
||||||
pub fn spawn(name: []const u8) bool {
|
/// driver it matched — danos-native, not POSIX (a spawn/exec family comes with the
|
||||||
return sc.systemCall4(.system_spawn, @intFromPtr(name.ptr), name.len, 0, 0) == 0;
|
/// POSIX layer later).
|
||||||
|
pub fn spawn(name: []const u8) ?u32 {
|
||||||
|
return spawnSupervised(name, &.{}, null);
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Like `spawn`, but hands the child command-line arguments: they arrive as
|
/// Like `spawn`, but hands the child command-line arguments: they arrive as
|
||||||
/// argv[1..] on its System V entry stack (argv[0] is still `name`). Marshalled to
|
/// argv[1..] on its System V entry stack (argv[0] is still `name`).
|
||||||
/// the kernel as one NUL-separated blob; the combined arguments must fit
|
pub fn spawnWithArguments(name: []const u8, arguments: []const []const u8) ?u32 {
|
||||||
/// `blob` (the kernel caps the blob at 256 bytes and argc at 8 anyway).
|
return spawnSupervised(name, arguments, null);
|
||||||
pub fn spawnWithArguments(name: []const u8, arguments: []const []const u8) bool {
|
}
|
||||||
|
|
||||||
|
/// The full spawn: command-line arguments for the child, and an optional endpoint
|
||||||
|
/// (a handle from `ipc.createIpcEndpoint`) the kernel notifies when the child ends
|
||||||
|
/// — any way it ends: clean exit, fault, or `kill`. The notification arrives via
|
||||||
|
/// `ipc.replyWait` as a badge with the child-exit bit set and the child's id in
|
||||||
|
/// the low bits (`ipc.Received.isChildExit`/`childProcessId`), so one endpoint can
|
||||||
|
/// supervise many children. Arguments are marshalled to the kernel as one
|
||||||
|
/// NUL-separated blob; the combined arguments must fit `blob` (the kernel caps the
|
||||||
|
/// blob at 256 bytes and argc at 8 anyway). Returns the child's process id, or
|
||||||
|
/// null on failure.
|
||||||
|
pub fn spawnSupervised(name: []const u8, arguments: []const []const u8, exit_endpoint: ?usize) ?u32 {
|
||||||
var blob: [256]u8 = undefined;
|
var blob: [256]u8 = undefined;
|
||||||
var len: usize = 0;
|
var len: usize = 0;
|
||||||
for (arguments, 0..) |argument, i| {
|
for (arguments, 0..) |argument, i| {
|
||||||
if (i != 0) {
|
if (i != 0) {
|
||||||
if (len >= blob.len) return false;
|
if (len >= blob.len) return null;
|
||||||
blob[len] = 0;
|
blob[len] = 0;
|
||||||
len += 1;
|
len += 1;
|
||||||
}
|
}
|
||||||
if (len + argument.len > blob.len) return false;
|
if (len + argument.len > blob.len) return null;
|
||||||
@memcpy(blob[len..][0..argument.len], argument);
|
@memcpy(blob[len..][0..argument.len], argument);
|
||||||
len += argument.len;
|
len += argument.len;
|
||||||
}
|
}
|
||||||
return sc.systemCall4(.system_spawn, @intFromPtr(name.ptr), name.len, @intFromPtr(&blob), len) == 0;
|
const r = sc.systemCall5(.system_spawn, @intFromPtr(name.ptr), name.len, if (len == 0) 0 else @intFromPtr(&blob), len, exit_endpoint orelse abi.no_cap);
|
||||||
|
if (r > ~@as(usize, 0) - 4095) return null; // a wrapped -errno
|
||||||
|
return @intCast(r);
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Snapshot the process table into `out` (up to its length) and return the total
|
||||||
|
/// number of live processes — which may exceed `out.len`; call again with a larger
|
||||||
|
/// buffer for the full listing. Kernel tasks are included, with an empty name.
|
||||||
|
/// The primitive `ps` is built on.
|
||||||
|
pub fn processes(out: []abi.ProcessDescriptor) usize {
|
||||||
|
return sc.systemCall2(.process_enumerate, @intFromPtr(out.ptr), out.len);
|
||||||
|
}
|
||||||
|
|
||||||
|
/// End process `id`. Only its supervisor — the process that spawned it — may;
|
||||||
|
/// anyone else gets false, as does a stale or unknown id (ids are never reused).
|
||||||
|
/// Delivery is prompt but asynchronous, like a signal: a target caught running on
|
||||||
|
/// another core dies at its next system call or timer tick. True means the kill
|
||||||
|
/// is accepted and irrevocable; the exit notification (if an endpoint was given
|
||||||
|
/// at spawn) confirms completion.
|
||||||
|
pub fn kill(id: u32) bool {
|
||||||
|
return sc.systemCall1(.process_kill, id) == 0;
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Grant `len` bytes (rounded up to whole pages) of fresh, zeroed, writable
|
/// Grant `len` bytes (rounded up to whole pages) of fresh, zeroed, writable
|
||||||
|
|||||||
+40
-5
@@ -43,13 +43,15 @@ pub const SystemCall = enum(u64) {
|
|||||||
irq_bind = 14, // irq_bind(id, resource_index, endpoint): deliver a device IRQ as an IPC notification
|
irq_bind = 14, // irq_bind(id, resource_index, endpoint): deliver a device IRQ as an IPC notification
|
||||||
irq_ack = 15, // irq_ack(id, resource_index): re-arm a bound IRQ after servicing it
|
irq_ack = 15, // irq_ack(id, resource_index): re-arm a bound IRQ after servicing it
|
||||||
device_register = 16, // device_register(parent_id, descriptor) -> id: publish a child of a device you claimed
|
device_register = 16, // device_register(parent_id, descriptor) -> id: publish a child of a device you claimed
|
||||||
system_spawn = 17, // system_spawn(name_ptr, name_len) -> 0: start a named initial-ramdisk binary as a new ring-3 process
|
system_spawn = 17, // system_spawn(name_ptr, name_len, arguments_ptr, arguments_len, exit_endpoint) -> child process id: start a named initial-ramdisk binary as a new ring-3 process
|
||||||
dma_alloc = 18, // dma_alloc(len, flags) -> vaddr (rax), paddr (rdx): contiguous, pinned, uncacheable DMA memory
|
dma_alloc = 18, // dma_alloc(len, flags) -> vaddr (rax), paddr (rdx): contiguous, pinned, uncacheable DMA memory
|
||||||
dma_free = 19, // dma_free(vaddr, len) -> 0: release a prior dma_alloc
|
dma_free = 19, // dma_free(vaddr, len) -> 0: release a prior dma_alloc
|
||||||
msi_bind = 20, // msi_bind(device_id, endpoint) -> address (rax), data (rdx): a per-device MSI vector for a claimed device
|
msi_bind = 20, // msi_bind(device_id, endpoint) -> address (rax), data (rdx): a per-device MSI vector for a claimed device
|
||||||
io_read = 21, // io_read(device_id, resource_index, offset, width) -> value: read a port in a claimed device's io_port resource
|
io_read = 21, // io_read(device_id, resource_index, offset, width) -> value: read a port in a claimed device's io_port resource
|
||||||
io_write = 22, // io_write(device_id, resource_index, offset, width, value) -> 0: write a port in a claimed device's io_port resource
|
io_write = 22, // io_write(device_id, resource_index, offset, width, value) -> 0: write a port in a claimed device's io_port resource
|
||||||
clock = 23, // clock() -> nanoseconds since boot: a monotonic time source (for timeouts/delays)
|
clock = 23, // clock() -> nanoseconds since boot: a monotonic time source (for timeouts/delays)
|
||||||
|
process_enumerate = 24, // process_enumerate(buffer, maximum) -> total: snapshot the task table
|
||||||
|
process_kill = 25, // process_kill(id) -> 0/-errno: end a process this process spawned
|
||||||
_,
|
_,
|
||||||
};
|
};
|
||||||
|
|
||||||
@@ -67,12 +69,45 @@ pub const dma_write_combining: u64 = 2; // write-combining (framebuffers); needs
|
|||||||
pub const dma_below_4g: u64 = 4; // physical address must fit 32 bits (legacy DMA engines)
|
pub const dma_below_4g: u64 = 4; // physical address must fit 32 bits (legacy DMA engines)
|
||||||
|
|
||||||
/// Set in the badge returned by `ipc_reply_wait` when what arrived is an
|
/// Set in the badge returned by `ipc_reply_wait` when what arrived is an
|
||||||
/// **asynchronous notification** (today: a device interrupt bound with `irq_bind`)
|
/// **asynchronous notification** (a device interrupt bound with `irq_bind`, or a
|
||||||
/// rather than a message from a client. There is no payload and no reply owed; the
|
/// child-exit notice — see `notify_exit_bit`) rather than a message from a client.
|
||||||
/// low bits carry the source, a GSI. Shared so the kernel's ISR and the driver's
|
/// There is no payload and no reply owed; the low bits carry the source. Shared so
|
||||||
/// event loop can't disagree about which bit means "the hardware spoke".
|
/// the kernel's ISR and the driver's event loop can't disagree about which bit
|
||||||
|
/// means "the hardware spoke".
|
||||||
pub const notify_badge_bit: u64 = 1 << 63;
|
pub const notify_badge_bit: u64 = 1 << 63;
|
||||||
|
|
||||||
|
/// Set (alongside `notify_badge_bit`) in the badge of a **child-exit notification**:
|
||||||
|
/// posted to the endpoint a supervisor passed to `system_spawn` when that child ends
|
||||||
|
/// — by clean exit, by a fault, or by `process_kill`. The low bits carry the child's
|
||||||
|
/// process id, so one endpoint can supervise many children (and even share with IRQ
|
||||||
|
/// notifications, which never set this bit). The microkernel's SIGCHLD.
|
||||||
|
pub const notify_exit_bit: u64 = 1 << 62;
|
||||||
|
|
||||||
|
/// Capacity of `ProcessDescriptor.name` — matches the longest name `system_spawn`
|
||||||
|
/// accepts, so a process's recorded name (its argv[0]) is never truncated.
|
||||||
|
pub const maximum_process_name = 64;
|
||||||
|
|
||||||
|
/// What a process is doing right now, as reported by `process_enumerate`. Crosses
|
||||||
|
/// the system_call boundary as `ProcessDescriptor.state`.
|
||||||
|
pub const ProcessState = enum(u32) {
|
||||||
|
ready = 0, // runnable, waiting for a core
|
||||||
|
running = 1, // executing on a core right now
|
||||||
|
blocked = 2, // waiting (sleeping, or blocked in IPC)
|
||||||
|
};
|
||||||
|
|
||||||
|
/// One `process_enumerate` entry — the kernel's view of a live task, kernel tasks
|
||||||
|
/// included (they carry an empty name and id 0 is the boot task). Fixed layout
|
||||||
|
/// (extern) because it crosses the kernel↔user boundary by memory copy, like
|
||||||
|
/// `DeviceDescriptor` in the device ABI.
|
||||||
|
pub const ProcessDescriptor = extern struct {
|
||||||
|
id: u32, // kernel-assigned process id; never reused (monotonic)
|
||||||
|
supervisor: u32, // id of the process that spawned it (0 = the kernel)
|
||||||
|
state: u32, // a ProcessState value
|
||||||
|
priority: u32,
|
||||||
|
name_length: u32,
|
||||||
|
name: [maximum_process_name]u8, // argv[0] at spawn; empty for kernel tasks
|
||||||
|
};
|
||||||
|
|
||||||
/// Well-known IPC service ids for the bootstrap name registry (create_ipc_endpoint +
|
/// Well-known IPC service ids for the bootstrap name registry (create_ipc_endpoint +
|
||||||
/// ipc_register/ipc_lookup). Small integers, so no string interning is needed
|
/// ipc_register/ipc_lookup). Small integers, so no string interning is needed
|
||||||
/// during bring-up. The VFS server registers under `vfs`; clients look it up.
|
/// during bring-up. The VFS server registers under `vfs`; clients look it up.
|
||||||
|
|||||||
@@ -47,6 +47,8 @@ pub const ENOENT: i64 = 4; // no such registered service
|
|||||||
pub const ENOSPC: i64 = 5; // handle table or registry full
|
pub const ENOSPC: i64 = 5; // handle table or registry full
|
||||||
pub const ENOMEM: i64 = 6; // out of memory
|
pub const ENOMEM: i64 = 6; // out of memory
|
||||||
pub const EPEER: i64 = 7; // peer died before replying (its process exited or was killed)
|
pub const EPEER: i64 = 7; // peer died before replying (its process exited or was killed)
|
||||||
|
pub const ESRCH: i64 = 8; // no such process (process_kill of an unknown/dead id)
|
||||||
|
pub const EPERM: i64 = 9; // not permitted (process_kill by anyone but the supervisor)
|
||||||
|
|
||||||
/// A badge with this bit set is an asynchronous notification (e.g. an IRQ), not a
|
/// A badge with this bit set is an asynchronous notification (e.g. an IRQ), not a
|
||||||
/// message from a client — there is no reply owed. The low bits carry the source
|
/// message from a client — there is no reply owed. The low bits carry the source
|
||||||
@@ -93,6 +95,7 @@ pub fn dropRef(endpoint: *Endpoint) void {
|
|||||||
// --- sender FIFO (endpoint-local, via Task.next) ----------------------------
|
// --- sender FIFO (endpoint-local, via Task.next) ----------------------------
|
||||||
|
|
||||||
fn enqueueSender(endpoint: *Endpoint, t: *Task) void {
|
fn enqueueSender(endpoint: *Endpoint, t: *Task) void {
|
||||||
|
t.ipc_wait_endpoint = @ptrCast(endpoint); // so a kill can unlink a parked caller
|
||||||
t.next = null;
|
t.next = null;
|
||||||
if (endpoint.sender_tail) |tail| tail.next = t else endpoint.sender_head = t;
|
if (endpoint.sender_tail) |tail| tail.next = t else endpoint.sender_head = t;
|
||||||
endpoint.sender_tail = t;
|
endpoint.sender_tail = t;
|
||||||
@@ -102,10 +105,34 @@ fn dequeueSender(endpoint: *Endpoint) ?*Task {
|
|||||||
const t = endpoint.sender_head orelse return null;
|
const t = endpoint.sender_head orelse return null;
|
||||||
endpoint.sender_head = t.next;
|
endpoint.sender_head = t.next;
|
||||||
if (endpoint.sender_head == null) endpoint.sender_tail = null;
|
if (endpoint.sender_head == null) endpoint.sender_tail = null;
|
||||||
|
t.ipc_wait_endpoint = null;
|
||||||
t.next = null;
|
t.next = null;
|
||||||
return t;
|
return t;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Unlink `t` from the sender FIFO it queues in, if any — the kill path for a
|
||||||
|
/// client parked in `call` that no server has received yet. Without this, a dead
|
||||||
|
/// caller would later be dequeued as a dangling pointer. The endpoint is still
|
||||||
|
/// alive here: `t`'s own handle table holds a reference until closeHandles runs
|
||||||
|
/// (which the kill path does *after* this). Precondition: the big kernel lock is
|
||||||
|
/// held.
|
||||||
|
pub fn abandonSenderLocked(t: *Task) void {
|
||||||
|
const endpoint: *Endpoint = @ptrCast(@alignCast(t.ipc_wait_endpoint orelse return));
|
||||||
|
t.ipc_wait_endpoint = null;
|
||||||
|
var previous: ?*Task = null;
|
||||||
|
var node = endpoint.sender_head;
|
||||||
|
while (node) |n| : ({
|
||||||
|
previous = n;
|
||||||
|
node = n.next;
|
||||||
|
}) {
|
||||||
|
if (n != t) continue;
|
||||||
|
if (previous) |p| p.next = t.next else endpoint.sender_head = t.next;
|
||||||
|
if (endpoint.sender_tail == t) endpoint.sender_tail = previous;
|
||||||
|
t.next = null;
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
// --- cross-address-space copy ----------------------------------------------
|
// --- cross-address-space copy ----------------------------------------------
|
||||||
|
|
||||||
/// Copy `len` bytes from `source_va` in address space `source_as` to `destination_va` in
|
/// Copy `len` bytes from `source_va` in address space `source_as` to `destination_va` in
|
||||||
|
|||||||
+170
-27
@@ -129,9 +129,14 @@ pub fn setInitialRamdisk(image: []const u8) void {
|
|||||||
/// written back into the trap frame, since the entry paths restore user registers
|
/// written back into the trap frame, since the entry paths restore user registers
|
||||||
/// from it. One handler serves both the system_call/sysret and int-0x80 entry paths.
|
/// from it. One handler serves both the system_call/sysret and int-0x80 entry paths.
|
||||||
///
|
///
|
||||||
/// Install it once at boot (before any user code runs) via `init`.
|
/// Install it once at boot (before any user code runs) via `init`. Also registers
|
||||||
|
/// the scheduler's kill hooks: the scheduler sits below this layer, so finishing a
|
||||||
|
/// deferred process_kill (IRQ bindings, IPC handles, the exit notification) is
|
||||||
|
/// called back up into here from the tick (see scheduler.reapKillPendingLocked).
|
||||||
pub fn init() void {
|
pub fn init() void {
|
||||||
architecture.setSystemCallHandler(system_call);
|
architecture.setSystemCallHandler(system_call);
|
||||||
|
scheduler.terminate_current_hook = terminateCurrentLocked;
|
||||||
|
scheduler.reap_task_hook = reapTaskLocked;
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Return -1 (as an unsigned bit pattern) in the system_call result register.
|
/// Return -1 (as an unsigned bit pattern) in the system_call result register.
|
||||||
@@ -140,6 +145,19 @@ fn fail(state: *architecture.CpuState) void {
|
|||||||
}
|
}
|
||||||
|
|
||||||
fn system_call(state: *architecture.CpuState) void {
|
fn system_call(state: *architecture.CpuState) void {
|
||||||
|
const t = scheduler.current();
|
||||||
|
const user = t.aspace != 0;
|
||||||
|
if (user) {
|
||||||
|
// A condemned process (process_kill caught it running) dies at its next
|
||||||
|
// kernel entry — before it can spawn, claim, or message anything else.
|
||||||
|
if (t.kill_pending) terminateCurrent();
|
||||||
|
// Mark the span of this call so the timer tick never tears the task down
|
||||||
|
// in the middle of a kernel operation (scheduler.reapKillPendingLocked).
|
||||||
|
t.in_system_call = true;
|
||||||
|
}
|
||||||
|
defer if (user) {
|
||||||
|
t.in_system_call = false;
|
||||||
|
};
|
||||||
switch (@as(SystemCall, @enumFromInt(architecture.systemCallNumber(state)))) {
|
switch (@as(SystemCall, @enumFromInt(architecture.systemCallNumber(state)))) {
|
||||||
.exit => {
|
.exit => {
|
||||||
exit_code = architecture.systemCallArg(state, 0);
|
exit_code = architecture.systemCallArg(state, 0);
|
||||||
@@ -178,6 +196,8 @@ fn system_call(state: *architecture.CpuState) void {
|
|||||||
.io_read => systemIoRead(state),
|
.io_read => systemIoRead(state),
|
||||||
.io_write => systemIoWrite(state),
|
.io_write => systemIoWrite(state),
|
||||||
.clock => systemClock(state),
|
.clock => systemClock(state),
|
||||||
|
.process_enumerate => systemProcessEnumerate(state),
|
||||||
|
.process_kill => systemProcessKill(state),
|
||||||
_ => fail(state),
|
_ => fail(state),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -415,28 +435,40 @@ fn systemDeviceRegister(state: *architecture.CpuState) void {
|
|||||||
architecture.setSystemCallResult(state, id);
|
architecture.setSystemCallResult(state, id);
|
||||||
}
|
}
|
||||||
|
|
||||||
/// system_spawn(name_ptr, name_len, arguments_ptr, arguments_len) -> 0 on success,
|
/// system_spawn(name_ptr, name_len, arguments_ptr, arguments_len, exit_endpoint)
|
||||||
/// -1 on failure. Load the binary bundled in the initial-ramdisk under `name` as a
|
/// -> the child's process id on success, -1 on failure. Load the binary bundled in
|
||||||
/// fresh ring-3 process. `name` becomes the child's argv[0] (and its task name, so
|
/// the initial-ramdisk under `name` as a fresh ring-3 process. `name` becomes the
|
||||||
/// a fault report can say which binary died); `arguments` is an optional
|
/// child's argv[0] (and its task name, so a fault report can say which binary
|
||||||
/// NUL-separated blob that becomes argv[1..] — how a supervisor parameterises what
|
/// died); `arguments` is an optional NUL-separated blob that becomes argv[1..] —
|
||||||
/// it starts ("you are the driver for device 12"). 0/0 means no extra arguments.
|
/// how a supervisor parameterises what it starts ("you are the driver for device
|
||||||
/// This is the mechanism a user-space supervisor (the device manager) uses to start
|
/// 12"). 0/0 means no extra arguments. This is the mechanism a user-space
|
||||||
/// a driver it matched: discovery and policy stay in user space, the kernel only
|
/// supervisor (the device manager) uses to start a driver it matched: discovery
|
||||||
/// spawns.
|
/// and policy stay in user space, the kernel only spawns.
|
||||||
///
|
///
|
||||||
/// Ungated for now — any process may spawn any bundled binary. A capability (only a
|
/// The caller is recorded as the child's **supervisor** — the sole holder of the
|
||||||
/// supervisor holds the right to spawn) belongs here once the model grows one; see
|
/// right to `process_kill` it (docs/process-management.md). `exit_endpoint` (a
|
||||||
/// docs/driver-model.md. Both buffers are bounds-checked into the user half exactly
|
/// handle, or `abi.no_cap` for none) names an endpoint of the caller's to notify
|
||||||
/// like `debug_write`, and an unknown name or a load failure returns -1.
|
/// when the child ends, any way it ends — the IRQ-as-IPC pattern reused as the
|
||||||
|
/// microkernel's SIGCHLD.
|
||||||
|
///
|
||||||
|
/// Spawning itself is still ungated — any process may spawn any bundled binary; a
|
||||||
|
/// spawn capability belongs here once the model grows one (docs/driver-model.md).
|
||||||
|
/// Both buffers are bounds-checked into the user half exactly like `debug_write`,
|
||||||
|
/// and an unknown name or a load failure returns -1.
|
||||||
fn systemSpawn(state: *architecture.CpuState) void {
|
fn systemSpawn(state: *architecture.CpuState) void {
|
||||||
const ptr = architecture.systemCallArg(state, 0);
|
const ptr = architecture.systemCallArg(state, 0);
|
||||||
const len = architecture.systemCallArg(state, 1);
|
const len = architecture.systemCallArg(state, 1);
|
||||||
const arguments_ptr = architecture.systemCallArg(state, 2);
|
const arguments_ptr = architecture.systemCallArg(state, 2);
|
||||||
const arguments_len = architecture.systemCallArg(state, 3);
|
const arguments_len = architecture.systemCallArg(state, 3);
|
||||||
if (len == 0 or len > 64 or ptr >= user_half_end or ptr + len > user_half_end) return fail(state);
|
const exit_handle = architecture.systemCallArg(state, 4);
|
||||||
|
const t = scheduler.current();
|
||||||
|
if (len == 0 or len > scheduler.maximum_task_name or ptr >= user_half_end or ptr + len > user_half_end) return fail(state);
|
||||||
if (arguments_len > maximum_argument_bytes) return fail(state);
|
if (arguments_len > maximum_argument_bytes) return fail(state);
|
||||||
if (arguments_len != 0 and (arguments_ptr >= user_half_end or arguments_ptr + arguments_len > user_half_end)) return fail(state);
|
if (arguments_len != 0 and (arguments_ptr >= user_half_end or arguments_ptr + arguments_len > user_half_end)) return fail(state);
|
||||||
|
const exit_endpoint: ?*ipc.Endpoint = if (exit_handle == abi.no_cap)
|
||||||
|
null
|
||||||
|
else
|
||||||
|
ipc.resolveHandle(t, exit_handle) orelse return failErr(state, ipc.EBADF);
|
||||||
const image = ramdisk_image orelse return fail(state);
|
const image = ramdisk_image orelse return fail(state);
|
||||||
const rd = initial_ramdisk.Reader.init(image) orelse return fail(state);
|
const rd = initial_ramdisk.Reader.init(image) orelse return fail(state);
|
||||||
|
|
||||||
@@ -458,19 +490,51 @@ fn systemSpawn(state: *architecture.CpuState) void {
|
|||||||
while (i < rd.count) : (i += 1) {
|
while (i < rd.count) : (i += 1) {
|
||||||
const item = rd.entry(i) orelse continue;
|
const item = rd.entry(i) orelse continue;
|
||||||
if (!std.mem.eql(u8, item.name, name)) continue;
|
if (!std.mem.eql(u8, item.name, name)) continue;
|
||||||
spawnProcess(item.blob, 4, argv[0..argc]) catch return fail(state);
|
const child = spawnProcessSupervised(item.blob, 4, argv[0..argc], t.id, exit_endpoint) catch return fail(state);
|
||||||
architecture.setSystemCallResult(state, 0);
|
architecture.setSystemCallResult(state, child);
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
fail(state); // no bundled binary by that name
|
fail(state); // no bundled binary by that name
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// process_enumerate(buffer, maximum) -> total: snapshot the task table into the
|
||||||
|
/// caller's buffer (up to `maximum` `abi.ProcessDescriptor` entries), returning
|
||||||
|
/// the total live-task count — the exact shape of `device_enumerate`, so a `ps`
|
||||||
|
/// is a user program over a snapshot, not a kernel service. Read-only and
|
||||||
|
/// ungated: what is running is not a secret between cooperating bring-up
|
||||||
|
/// processes.
|
||||||
|
fn systemProcessEnumerate(state: *architecture.CpuState) void {
|
||||||
|
const buffer_ptr = architecture.systemCallArg(state, 0);
|
||||||
|
const maximum = architecture.systemCallArg(state, 1);
|
||||||
|
const t = scheduler.current();
|
||||||
|
if (t.aspace == 0 or buffer_ptr >= user_half_end) return fail(state);
|
||||||
|
const sz = @sizeOf(abi.ProcessDescriptor);
|
||||||
|
const cap = @min(maximum, (user_half_end - buffer_ptr) / sz); // clamp to the user half
|
||||||
|
const out: [*]abi.ProcessDescriptor = @ptrFromInt(buffer_ptr);
|
||||||
|
architecture.setSystemCallResult(state, scheduler.enumerate(out[0..@intCast(cap)]));
|
||||||
|
}
|
||||||
|
|
||||||
|
/// process_kill(id) -> 0 / -ESRCH / -EPERM: end the process `id`. Only its
|
||||||
|
/// supervisor — the process that spawned it — may do so; the supervision link is
|
||||||
|
/// the kill capability, so no user/permission model is needed and a stray id
|
||||||
|
/// cannot be a weapon (ids are never reused, so a stale one just misses).
|
||||||
|
fn systemProcessKill(state: *architecture.CpuState) void {
|
||||||
|
const t = scheduler.current();
|
||||||
|
if (t.aspace == 0) return fail(state);
|
||||||
|
const id = architecture.systemCallArg(state, 0);
|
||||||
|
if (id > std.math.maxInt(u32)) return failErr(state, ipc.ESRCH);
|
||||||
|
const r = killProcess(t.id, @intCast(id));
|
||||||
|
architecture.setSystemCallResult(state, @bitCast(r));
|
||||||
|
}
|
||||||
|
|
||||||
/// Processes killed by a CPU fault rather than a clean exit. Evidence for the
|
/// Processes killed by a CPU fault rather than a clean exit. Evidence for the
|
||||||
/// fault-recovery test, and a health signal a supervisor can consult later.
|
/// fault-recovery test, and a health signal a supervisor can consult later.
|
||||||
pub var fault_kill_count: u64 = 0;
|
pub var fault_kill_count: u64 = 0;
|
||||||
|
|
||||||
/// Tear down the current user process and reschedule; never returns. Shared by the
|
/// Release everything a dying task holds and tell its supervisor — the shared
|
||||||
/// exit system call and the fault path (`killCurrentProcess`). The order matters:
|
/// half of every path out of a process: clean exit, fault kill, and process_kill
|
||||||
|
/// (both the immediate reap and the deferred tick-time terminate). The order
|
||||||
|
/// matters:
|
||||||
/// - IRQ bindings are dropped before the handle table closes: dropping the last
|
/// - IRQ bindings are dropped before the handle table closes: dropping the last
|
||||||
/// endpoint reference destroys the Endpoint, and a still-bound GSI would have an
|
/// endpoint reference destroys the Endpoint, and a still-bound GSI would have an
|
||||||
/// ISR call notifyFromIsr on freed memory the next time the device fired.
|
/// ISR call notifyFromIsr on freed memory the next time the device fired.
|
||||||
@@ -479,20 +543,84 @@ pub var fault_kill_count: u64 = 0;
|
|||||||
/// - A client this task still owes a reply to (it died between receive and reply)
|
/// - A client this task still owes a reply to (it died between receive and reply)
|
||||||
/// is failed with -EPEER rather than left blocked forever — a dead server must
|
/// is failed with -EPEER rather than left blocked forever — a dead server must
|
||||||
/// not hang its callers.
|
/// not hang its callers.
|
||||||
pub fn terminateCurrent() noreturn {
|
/// - The task is unlinked from wherever IPC parked it (an endpoint's sender FIFO,
|
||||||
const t = scheduler.current();
|
/// a receive wait queue, or a server's owed-reply slot) *before* the handles
|
||||||
{
|
/// close, so nothing ever dequeues a dangling pointer. These are no-ops for a
|
||||||
const flags = sync.enter();
|
/// running task ending itself; they matter when process_kill reaps a blocked one.
|
||||||
defer sync.leave(flags);
|
/// - The exit notification is posted last, once the process can no longer act, so
|
||||||
|
/// a supervisor that receives it observes a fully-released child. The endpoint
|
||||||
|
/// reference taken at spawn is dropped with it.
|
||||||
|
/// Precondition: the big kernel lock is held.
|
||||||
|
fn releaseTaskResourcesLocked(t: *scheduler.Task) void {
|
||||||
irq.releaseOwner(t.id);
|
irq.releaseOwner(t.id);
|
||||||
if (t.ipc_client) |client| {
|
if (t.ipc_client) |client| {
|
||||||
t.ipc_client = null;
|
t.ipc_client = null;
|
||||||
client.ipc_status = -ipc.EPEER;
|
client.ipc_status = -ipc.EPEER;
|
||||||
scheduler.readyLocked(client); // its blocked `call` now returns the error
|
scheduler.readyLocked(client); // its blocked `call` now returns the error
|
||||||
}
|
}
|
||||||
|
ipc.abandonSenderLocked(t);
|
||||||
|
scheduler.removeFromWaitQueueLocked(t);
|
||||||
|
scheduler.forgetIpcClientLocked(t);
|
||||||
ipc.closeHandles(t);
|
ipc.closeHandles(t);
|
||||||
|
if (t.exit_endpoint) |raw| {
|
||||||
|
const endpoint: *ipc.Endpoint = @ptrCast(@alignCast(raw));
|
||||||
|
t.exit_endpoint = null;
|
||||||
|
ipc.notifyLocked(endpoint, abi.notify_exit_bit | t.id);
|
||||||
|
ipc.dropRef(endpoint);
|
||||||
}
|
}
|
||||||
scheduler.exitUser();
|
}
|
||||||
|
|
||||||
|
/// Tear down the current user process and reschedule; never returns. Shared by the
|
||||||
|
/// exit system call and the fault path (`killCurrentProcess`). See
|
||||||
|
/// `releaseTaskResourcesLocked` for what is released, and in what order.
|
||||||
|
pub fn terminateCurrent() noreturn {
|
||||||
|
_ = sync.enter(); // handed off through the exit switch, released by the resumed task
|
||||||
|
terminateCurrentLocked();
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The body of `terminateCurrent` for a caller that already holds the big kernel
|
||||||
|
/// lock — the scheduler's tick calls this (via `terminate_current_hook`) to finish
|
||||||
|
/// a deferred process_kill on its own core's current task. Never returns; the
|
||||||
|
/// tick's abandoned interrupt frame is fine (the LAPIC was acknowledged before the
|
||||||
|
/// tick hook ran), exactly as on the fault path.
|
||||||
|
fn terminateCurrentLocked() noreturn {
|
||||||
|
releaseTaskResourcesLocked(scheduler.current());
|
||||||
|
scheduler.exitUserLocked();
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Reap a condemned task that is NOT running on any core (ready or blocked — and
|
||||||
|
/// it cannot start running: state changes need the lock we hold). The other half
|
||||||
|
/// of a deferred process_kill, called by the scheduler's tick (via
|
||||||
|
/// `reap_task_hook`) and directly by `killProcess` for targets caught off-CPU.
|
||||||
|
/// Precondition: the big kernel lock is held.
|
||||||
|
fn reapTaskLocked(t: *scheduler.Task) void {
|
||||||
|
releaseTaskResourcesLocked(t);
|
||||||
|
scheduler.removeFromReadyQueueLocked(t); // no-op unless it was ready in a queue
|
||||||
|
scheduler.destroyTaskLocked(t);
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Kill process `target_id` on behalf of `caller_id` — the kernel half of the
|
||||||
|
/// process_kill system call. Returns 0, -ESRCH (no such live process — kernel
|
||||||
|
/// tasks are not killable processes and stale ids miss, since ids are never
|
||||||
|
/// reused), or -EPERM (the caller is not the target's supervisor).
|
||||||
|
///
|
||||||
|
/// A target that is ready or blocked is reaped on the spot. One that is running
|
||||||
|
/// on another core cannot be torn down mid-instruction, so it is condemned
|
||||||
|
/// (`kill_pending`) and dies at its next system_call entry, block, or timer tick
|
||||||
|
/// — like a Unix signal, delivery is prompt but asynchronous. Either way the
|
||||||
|
/// call returns 0: the kill is accepted and irrevocable.
|
||||||
|
pub fn killProcess(caller_id: u32, target_id: u32) i64 {
|
||||||
|
const flags = sync.enter();
|
||||||
|
defer sync.leave(flags);
|
||||||
|
const target = scheduler.taskByIdLocked(target_id) orelse return -ipc.ESRCH;
|
||||||
|
if (target.aspace == 0) return -ipc.ESRCH; // kernel tasks are not processes
|
||||||
|
if (target.supervisor != caller_id) return -ipc.EPERM;
|
||||||
|
if (target.state == .running) {
|
||||||
|
target.kill_pending = true;
|
||||||
|
} else {
|
||||||
|
reapTaskLocked(target);
|
||||||
|
}
|
||||||
|
return 0;
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Kill the current user process in response to a CPU fault it raised in ring 3.
|
/// Kill the current user process in response to a CPU fault it raised in ring 3.
|
||||||
@@ -865,11 +993,21 @@ fn entryStackBytes(argv: []const []const u8) usize {
|
|||||||
/// convention (`buildEntryStack`). `argv[0]` is required — it names the process:
|
/// convention (`buildEntryStack`). `argv[0]` is required — it names the process:
|
||||||
/// the path or initial-ramdisk name it was spawned as. It is also recorded on the
|
/// the path or initial-ramdisk name it was spawned as. It is also recorded on the
|
||||||
/// task, so a fault report can say *which* binary died, not just its id.
|
/// task, so a fault report can say *which* binary died, not just its id.
|
||||||
|
/// The kernel-internal spawn (init at boot, tests): supervisor 0, no exit
|
||||||
|
/// notification. `spawnProcessSupervised` is the full form.
|
||||||
|
pub fn spawnProcess(image: []const u8, priority: u3, argv: []const []const u8) InitError!void {
|
||||||
|
_ = try spawnProcessSupervised(image, priority, argv, 0, null);
|
||||||
|
}
|
||||||
|
|
||||||
|
/// `spawnProcess`, recording `supervisor` (the id of the process that asked — the
|
||||||
|
/// kill authority) and, if given, `exit_endpoint` to notify when the child ends
|
||||||
|
/// (a reference is taken here and dropped when the notification posts).
|
||||||
|
/// Returns the child's process id.
|
||||||
/// Returns immediately — the process runs preemptively on its own page tables
|
/// Returns immediately — the process runs preemptively on its own page tables
|
||||||
/// alongside everything else, and its exit is handled by the system_call layer.
|
/// alongside everything else, and its exit is handled by the system_call layer.
|
||||||
/// The whole build (address space + ELF load + task) runs under the kernel lock so
|
/// The whole build (address space + ELF load + task) runs under the kernel lock so
|
||||||
/// it appears atomically and can't race pmm/heap on another core.
|
/// it appears atomically and can't race pmm/heap on another core.
|
||||||
pub fn spawnProcess(image: []const u8, priority: u3, argv: []const []const u8) InitError!void {
|
pub fn spawnProcessSupervised(image: []const u8, priority: u3, argv: []const []const u8, supervisor: u32, exit_endpoint: ?*ipc.Endpoint) InitError!u32 {
|
||||||
if (argv.len == 0 or argv.len > maximum_arguments) return error.BadArguments;
|
if (argv.len == 0 or argv.len > maximum_arguments) return error.BadArguments;
|
||||||
// The entry block must leave most of the page as actual stack.
|
// The entry block must leave most of the page as actual stack.
|
||||||
if (entryStackBytes(argv) > page_size / 2) return error.BadArguments;
|
if (entryStackBytes(argv) > page_size / 2) return error.BadArguments;
|
||||||
@@ -901,8 +1039,13 @@ pub fn spawnProcess(image: []const u8, priority: u3, argv: []const []const u8) I
|
|||||||
architecture.mapUserPageInto(aspace, page_virtual, stack_frame, true, false); // RW + NX
|
architecture.mapUserPageInto(aspace, page_virtual, stack_frame, true, false); // RW + NX
|
||||||
}
|
}
|
||||||
|
|
||||||
if (!scheduler.spawnUserLocked(aspace, parsed.entry, user_sp, priority, argv[0]))
|
const child = scheduler.spawnUserLocked(aspace, parsed.entry, user_sp, priority, argv[0], supervisor, if (exit_endpoint) |endpoint| @ptrCast(endpoint) else null) orelse
|
||||||
return error.OutOfMemory;
|
return error.OutOfMemory;
|
||||||
|
// The child holds a reference to its exit endpoint from birth to death. Taken
|
||||||
|
// only now, after nothing can fail; the lock is still held, so the child
|
||||||
|
// cannot run (let alone die) before the reference exists.
|
||||||
|
if (exit_endpoint) |endpoint| endpoint.refcount += 1;
|
||||||
|
return child;
|
||||||
}
|
}
|
||||||
|
|
||||||
/// clock() -> nanoseconds since boot: a monotonic time source. The kernel already owns
|
/// clock() -> nanoseconds since boot: a monotonic time source. The kernel already owns
|
||||||
|
|||||||
+206
-11
@@ -18,6 +18,7 @@
|
|||||||
//! shared queues.
|
//! shared queues.
|
||||||
|
|
||||||
const std = @import("std");
|
const std = @import("std");
|
||||||
|
const abi = @import("abi");
|
||||||
const parameters = @import("parameters");
|
const parameters = @import("parameters");
|
||||||
const architecture = @import("architecture");
|
const architecture = @import("architecture");
|
||||||
const heap = @import("heap.zig");
|
const heap = @import("heap.zig");
|
||||||
@@ -41,6 +42,26 @@ pub const Task = struct {
|
|||||||
kstack_top: usize = 0, // top of `stack` (== TSS.rsp0 for a user task); 0 = none
|
kstack_top: usize = 0, // top of `stack` (== TSS.rsp0 for a user task); 0 = none
|
||||||
wake_at: u64 = 0, // uptime (ms) to wake a sleeping task; 0 = not sleeping
|
wake_at: u64 = 0, // uptime (ms) to wake a sleeping task; 0 = not sleeping
|
||||||
affinity: ?u32 = null, // null = runs on any core; else the index of its pinned core
|
affinity: ?u32 = null, // null = runs on any core; else the index of its pinned core
|
||||||
|
// --- process management (process.zig) ---
|
||||||
|
// Id of the process that spawned this one (0 = the kernel). The supervision
|
||||||
|
// link is the kill authority: only the supervisor may process_kill a child.
|
||||||
|
supervisor: u32 = 0,
|
||||||
|
// Endpoint to notify when this process ends (any way: exit, fault, kill), or
|
||||||
|
// null. Holds its own reference, dropped when the notification is posted.
|
||||||
|
// Opaque here for the same reason as `handles` below.
|
||||||
|
exit_endpoint: ?*anyopaque = null,
|
||||||
|
// Set by process_kill on a task that is running on another core; the kernel
|
||||||
|
// finishes the kill at that task's next system call or timer tick.
|
||||||
|
kill_pending: bool = false,
|
||||||
|
// True while this task executes its own system call — the timer tick must not
|
||||||
|
// tear a task down in the middle of a kernel operation, only while it runs
|
||||||
|
// user code (or sits at a block point, where teardown is safe).
|
||||||
|
in_system_call: bool = false,
|
||||||
|
// Where this task is parked while blocked, so a kill can unlink it: the
|
||||||
|
// WaitQueue it waits on (maintained by waitLocked/wakeLocked), or the endpoint
|
||||||
|
// whose sender FIFO it queues in (maintained by the IPC layer; opaque here).
|
||||||
|
wait_queue: ?*WaitQueue = null,
|
||||||
|
ipc_wait_endpoint: ?*anyopaque = null,
|
||||||
// Physical root of this task's address space, or 0 for a kernel task (which
|
// Physical root of this task's address space, or 0 for a kernel task (which
|
||||||
// runs on the shared kernel page tables). A user task carries its own.
|
// runs on the shared kernel page tables). A user task carries its own.
|
||||||
aspace: u64 = 0,
|
aspace: u64 = 0,
|
||||||
@@ -85,8 +106,9 @@ pub const Task = struct {
|
|||||||
};
|
};
|
||||||
|
|
||||||
/// Capacity of `Task.name_buffer` — matches the longest name `system_spawn`
|
/// Capacity of `Task.name_buffer` — matches the longest name `system_spawn`
|
||||||
/// accepts, so a spawned name is never truncated.
|
/// accepts, so a spawned name is never truncated. Shared with the ABI's
|
||||||
pub const maximum_task_name = 64;
|
/// ProcessDescriptor, so `enumerate` copies names without clipping.
|
||||||
|
pub const maximum_task_name = abi.maximum_process_name;
|
||||||
|
|
||||||
/// Size of each task's IPC handle table. Kept here (not in ipc_sync.zig) because
|
/// Size of each task's IPC handle table. Kept here (not in ipc_sync.zig) because
|
||||||
/// it dimensions a field of `Task`; ipc_sync.zig re-exports it.
|
/// it dimensions a field of `Task`; ipc_sync.zig re-exports it.
|
||||||
@@ -275,14 +297,18 @@ pub fn spawnOn(entry: *const fn () void, priority: Priority, cpu: u32) bool {
|
|||||||
|
|
||||||
/// Spawn a **user** task: a task with its own address space (`aspace`) that starts
|
/// Spawn a **user** task: a task with its own address space (`aspace`) that starts
|
||||||
/// in user mode at `entry` on `user_sp`, recorded under `name` (its argv[0]).
|
/// in user mode at `entry` on `user_sp`, recorded under `name` (its argv[0]).
|
||||||
|
/// `supervisor` is the id of the spawning process (0 = the kernel) — the kill
|
||||||
|
/// authority — and `exit_endpoint` (an *ipc.Endpoint whose reference the caller
|
||||||
|
/// has already taken, or null) is notified when this process ends.
|
||||||
/// It gets a fresh kernel stack for syscalls/interrupts, and its first switch-in
|
/// It gets a fresh kernel stack for syscalls/interrupts, and its first switch-in
|
||||||
/// lands in `user_task_trampoline`.
|
/// lands in `user_task_trampoline`.
|
||||||
/// Returns false (creating nothing) if the table is full or out of memory.
|
/// Returns the new process id, or null (creating nothing) if the table is full or
|
||||||
|
/// out of memory.
|
||||||
/// **Caller must hold the kernel lock** (the loader that builds `aspace` holds it
|
/// **Caller must hold the kernel lock** (the loader that builds `aspace` holds it
|
||||||
/// across the whole spawn, so the address space and the task appear atomically).
|
/// across the whole spawn, so the address space and the task appear atomically).
|
||||||
pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, priority: Priority, task_name: []const u8) bool {
|
pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, priority: Priority, task_name: []const u8, supervisor: u32, exit_endpoint: ?*anyopaque) ?u32 {
|
||||||
const t = freeSlot() orelse return false;
|
const t = freeSlot() orelse return null;
|
||||||
const stack = heap.allocator().alloc(u8, stack_size) catch return false;
|
const stack = heap.allocator().alloc(u8, stack_size) catch return null;
|
||||||
t.* = .{
|
t.* = .{
|
||||||
.id = next_id,
|
.id = next_id,
|
||||||
.state = .ready,
|
.state = .ready,
|
||||||
@@ -291,6 +317,8 @@ pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, priority: Priority
|
|||||||
.aspace = aspace,
|
.aspace = aspace,
|
||||||
.user_ip = entry,
|
.user_ip = entry,
|
||||||
.user_sp = user_sp,
|
.user_sp = user_sp,
|
||||||
|
.supervisor = supervisor,
|
||||||
|
.exit_endpoint = exit_endpoint,
|
||||||
};
|
};
|
||||||
const name_length = @min(task_name.len, maximum_task_name);
|
const name_length = @min(task_name.len, maximum_task_name);
|
||||||
@memcpy(t.name_buffer[0..name_length], task_name[0..name_length]);
|
@memcpy(t.name_buffer[0..name_length], task_name[0..name_length]);
|
||||||
@@ -302,7 +330,7 @@ pub fn spawnUserLocked(aspace: u64, entry: u64, user_sp: u64, priority: Priority
|
|||||||
// the user entry/stack from the Task itself).
|
// the user entry/stack from the Task itself).
|
||||||
t.sp = architecture.initTaskStack(top, @intFromPtr(&startUserTask));
|
t.sp = architecture.initTaskStack(top, @intFromPtr(&startUserTask));
|
||||||
enqueue(t);
|
enqueue(t);
|
||||||
return true;
|
return t.id;
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The first thing a fresh user task runs (in ring 0, via task_trampoline). It
|
/// The first thing a fresh user task runs (in ring 0, via task_trampoline). It
|
||||||
@@ -414,6 +442,7 @@ pub const WaitQueue = struct {
|
|||||||
pub fn waitLocked(wait_queue: *WaitQueue) void {
|
pub fn waitLocked(wait_queue: *WaitQueue) void {
|
||||||
const t = current();
|
const t = current();
|
||||||
t.state = .blocked;
|
t.state = .blocked;
|
||||||
|
t.wait_queue = wait_queue; // so a kill can unlink a parked waiter
|
||||||
t.next = wait_queue.head;
|
t.next = wait_queue.head;
|
||||||
wait_queue.head = t;
|
wait_queue.head = t;
|
||||||
schedule();
|
schedule();
|
||||||
@@ -438,10 +467,78 @@ pub fn wakeLocked(wait_queue: *WaitQueue) void {
|
|||||||
}
|
}
|
||||||
const t = best orelse return;
|
const t = best orelse return;
|
||||||
if (best_previous) |p| p.next = t.next else wait_queue.head = t.next;
|
if (best_previous) |p| p.next = t.next else wait_queue.head = t.next;
|
||||||
|
t.wait_queue = null;
|
||||||
t.state = .ready;
|
t.state = .ready;
|
||||||
enqueue(t);
|
enqueue(t);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Unlink `t` from the wait queue it is parked on, if any (the kill path — a
|
||||||
|
/// killed waiter must not be woken later as a dangling pointer). Precondition:
|
||||||
|
/// the big kernel lock is held.
|
||||||
|
pub fn removeFromWaitQueueLocked(t: *Task) void {
|
||||||
|
const wait_queue = t.wait_queue orelse return;
|
||||||
|
t.wait_queue = null;
|
||||||
|
var previous: ?*Task = null;
|
||||||
|
var node = wait_queue.head;
|
||||||
|
while (node) |n| : ({
|
||||||
|
previous = n;
|
||||||
|
node = n.next;
|
||||||
|
}) {
|
||||||
|
if (n != t) continue;
|
||||||
|
if (previous) |p| p.next = t.next else wait_queue.head = t.next;
|
||||||
|
t.next = null;
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Unlink `t` from the ready queue it sits in (global, or its affinity core's
|
||||||
|
/// pinned queue) — the kill path for a task that is runnable but not running.
|
||||||
|
/// Precondition: the big kernel lock is held.
|
||||||
|
pub fn removeFromReadyQueueLocked(t: *Task) void {
|
||||||
|
if (t.affinity) |cpu| {
|
||||||
|
const pc = &cpus[cpu];
|
||||||
|
removeFrom(&pc.pinned_head, &pc.pinned_tail, &pc.pinned_bitmap, t);
|
||||||
|
} else {
|
||||||
|
removeFrom(&ready_head, &ready_tail, &ready_bitmap, t);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
fn removeFrom(head: *[number_priorities]?*Task, tail: *[number_priorities]?*Task, bitmap: *u8, t: *Task) void {
|
||||||
|
const level: usize = t.priority;
|
||||||
|
var previous: ?*Task = null;
|
||||||
|
var node = head[level];
|
||||||
|
while (node) |n| : ({
|
||||||
|
previous = n;
|
||||||
|
node = n.next;
|
||||||
|
}) {
|
||||||
|
if (n != t) continue;
|
||||||
|
if (previous) |p| p.next = t.next else head[level] = t.next;
|
||||||
|
if (tail[level] == t) tail[level] = previous;
|
||||||
|
if (head[level] == null) bitmap.* &= ~(@as(u8, 1) << @intCast(level));
|
||||||
|
t.next = null;
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Find a live task by process id, or null. Ids are monotonic and never reused,
|
||||||
|
/// so a stale id misses cleanly rather than naming a recycled slot.
|
||||||
|
/// Precondition: the big kernel lock is held.
|
||||||
|
pub fn taskByIdLocked(id: u32) ?*Task {
|
||||||
|
for (&tasks) |*t| {
|
||||||
|
if (t.state != .free and t.id == id) return t;
|
||||||
|
}
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Make every server that still holds `t` as the client it owes a reply to forget
|
||||||
|
/// it — the reply of a dead client is dropped, not delivered into freed state.
|
||||||
|
/// Precondition: the big kernel lock is held.
|
||||||
|
pub fn forgetIpcClientLocked(t: *Task) void {
|
||||||
|
for (&tasks) |*other| {
|
||||||
|
if (other.state != .free and other.ipc_client == t) other.ipc_client = null;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
/// Block the current task and switch away, without putting it on any wait queue —
|
/// Block the current task and switch away, without putting it on any wait queue —
|
||||||
/// the caller has already linked it wherever it belongs (e.g. an endpoint's sender
|
/// the caller has already linked it wherever it belongs (e.g. an endpoint's sender
|
||||||
/// FIFO). Precondition: the big kernel lock is held; still held on return (when the
|
/// FIFO). Precondition: the big kernel lock is held; still held on return (when the
|
||||||
@@ -501,14 +598,55 @@ fn wakeExpired() void {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Process-teardown hooks, registered by process.zig at init — the scheduler sits
|
||||||
|
// below the process layer, so finishing a kill (IRQ bindings, IPC handles, exit
|
||||||
|
// notification) is called *up* through these, mirroring how the architecture
|
||||||
|
// layer calls up into `tick`.
|
||||||
|
//
|
||||||
|
// `terminate_current_hook` ends the task running on THIS core (lock held, never
|
||||||
|
// returns — it switches away like `exitUserLocked`). `reap_task_hook` tears down
|
||||||
|
// a task that is NOT running on any core (lock held).
|
||||||
|
pub var terminate_current_hook: ?*const fn () noreturn = null;
|
||||||
|
pub var reap_task_hook: ?*const fn (*Task) void = null;
|
||||||
|
|
||||||
|
/// Finish any pending kills this core can see (the deferred half of process_kill;
|
||||||
|
/// the immediate half runs in the killer's own call). Precondition: the big kernel
|
||||||
|
/// lock is held, from `tick`.
|
||||||
|
///
|
||||||
|
/// - This core's *current* task, if condemned, is terminated here — but only when
|
||||||
|
/// it is not inside one of its own system calls (`in_system_call`): the tick may
|
||||||
|
/// have interrupted kernel code mid-operation, where teardown would leak or
|
||||||
|
/// corrupt what that operation holds. User-mode execution (and the system_call
|
||||||
|
/// entry/exit stubs, which hold nothing) are safe termination points. A task
|
||||||
|
/// that *is* mid-call dies at its next block, tick, or system_call entry instead.
|
||||||
|
/// The hook never returns; abandoning the interrupt frame is fine — the LAPIC
|
||||||
|
/// was acknowledged before the tick hook ran (see apic.timerTick), exactly as on
|
||||||
|
/// the fault-kill path.
|
||||||
|
/// - Condemned tasks that are ready or blocked are not running anywhere (state
|
||||||
|
/// changes need the lock we hold), so they are reaped in place.
|
||||||
|
fn reapKillPendingLocked() void {
|
||||||
|
const pc = thisCpu();
|
||||||
|
const cur = pc.current;
|
||||||
|
if (cur.kill_pending and cur.aspace != 0 and !cur.in_system_call) {
|
||||||
|
if (terminate_current_hook) |hook| hook(); // noreturn
|
||||||
|
}
|
||||||
|
if (reap_task_hook) |hook| {
|
||||||
|
for (&tasks) |*t| {
|
||||||
|
if (!t.kill_pending) continue;
|
||||||
|
if (t.state == .ready or t.state == .blocked) hook(t);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
/// Called from the timer interrupt (interrupts already disabled): wake due
|
/// Called from the timer interrupt (interrupts already disabled): wake due
|
||||||
/// sleepers, then preempt. Takes the kernel lock like any other critical section,
|
/// sleepers, finish pending kills, then preempt. Takes the kernel lock like any
|
||||||
/// but releases it *without* touching the interrupt flag — the handler's `iretq`
|
/// other critical section, but releases it *without* touching the interrupt flag
|
||||||
/// restores the interrupted context's flags, so re-enabling here would open a
|
/// — the handler's `iretq` restores the interrupted context's flags, so
|
||||||
/// nested-interrupt window before the return.
|
/// re-enabling here would open a nested-interrupt window before the return.
|
||||||
pub fn tick() void {
|
pub fn tick() void {
|
||||||
_ = sync.enter();
|
_ = sync.enter();
|
||||||
wakeExpired();
|
wakeExpired();
|
||||||
|
reapKillPendingLocked();
|
||||||
if (preemption_enabled) schedule();
|
if (preemption_enabled) schedule();
|
||||||
sync.leaveIsr();
|
sync.leaveIsr();
|
||||||
}
|
}
|
||||||
@@ -540,6 +678,13 @@ pub fn exit() noreturn {
|
|||||||
/// itself is leaked, as in `exit` (no reaper yet). Never returns.
|
/// itself is leaked, as in `exit` (no reaper yet). Never returns.
|
||||||
pub fn exitUser() noreturn {
|
pub fn exitUser() noreturn {
|
||||||
_ = sync.enter();
|
_ = sync.enter();
|
||||||
|
exitUserLocked();
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The body of `exitUser` for callers that already hold the big kernel lock (the
|
||||||
|
/// tick-time terminate path, which enters with the lock held). The lock is handed
|
||||||
|
/// off through the switch and released by the task that resumes. Never returns.
|
||||||
|
pub fn exitUserLocked() noreturn {
|
||||||
const pc = thisCpu();
|
const pc = thisCpu();
|
||||||
const dying = pc.current;
|
const dying = pc.current;
|
||||||
const as = dying.aspace;
|
const as = dying.aspace;
|
||||||
@@ -551,6 +696,8 @@ pub fn exitUser() noreturn {
|
|||||||
}
|
}
|
||||||
dying.state = .free;
|
dying.state = .free;
|
||||||
dying.aspace = 0;
|
dying.aspace = 0;
|
||||||
|
dying.kill_pending = false;
|
||||||
|
dying.in_system_call = false;
|
||||||
const next = dequeueHighest(pc) orelse @panic("sched: no task left to run");
|
const next = dequeueHighest(pc) orelse @panic("sched: no task left to run");
|
||||||
next.state = .running;
|
next.state = .running;
|
||||||
pc.current = next;
|
pc.current = next;
|
||||||
@@ -559,6 +706,54 @@ pub fn exitUser() noreturn {
|
|||||||
unreachable;
|
unreachable;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Free a task that is NOT running on any core (it is ready or blocked, and the
|
||||||
|
/// caller — the kill path — has already unlinked it from every queue and released
|
||||||
|
/// what it held). Destroys its address space: safe here because no core can have
|
||||||
|
/// it loaded (every switch away from a task loads the next task's tables, and the
|
||||||
|
/// task isn't running). The kernel stack is leaked, as in `exitUser` (no reaper
|
||||||
|
/// yet). Precondition: the big kernel lock is held.
|
||||||
|
pub fn destroyTaskLocked(t: *Task) void {
|
||||||
|
if (t.aspace != 0) architecture.destroyAddressSpace(t.aspace);
|
||||||
|
t.aspace = 0;
|
||||||
|
t.kill_pending = false;
|
||||||
|
t.in_system_call = false;
|
||||||
|
t.wake_at = 0;
|
||||||
|
t.state = .free;
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Snapshot the task table into `out` (up to its length), returning the total
|
||||||
|
/// number of live tasks — the kernel half of `process_enumerate`, mirroring
|
||||||
|
/// devices_broker.enumerate. Kernel tasks are included (empty name, supervisor 0):
|
||||||
|
/// an honest `ps` shows the idle tasks too. `out` may be user memory: the caller's
|
||||||
|
/// address space is loaded during its system call, and the same bring-up trust
|
||||||
|
/// applies as for device_enumerate (an unmapped user page faults the kernel).
|
||||||
|
pub fn enumerate(out: []abi.ProcessDescriptor) u64 {
|
||||||
|
const flags = sync.enter();
|
||||||
|
defer sync.leave(flags);
|
||||||
|
var total: u64 = 0;
|
||||||
|
for (&tasks) |*t| {
|
||||||
|
if (t.state == .free) continue;
|
||||||
|
if (total < out.len) {
|
||||||
|
const d = &out[total];
|
||||||
|
d.* = .{
|
||||||
|
.id = t.id,
|
||||||
|
.supervisor = t.supervisor,
|
||||||
|
.state = @intFromEnum(@as(abi.ProcessState, switch (t.state) {
|
||||||
|
.ready => .ready,
|
||||||
|
.running => .running,
|
||||||
|
.blocked => .blocked,
|
||||||
|
.free => unreachable,
|
||||||
|
})),
|
||||||
|
.priority = t.priority,
|
||||||
|
.name_length = t.name_length,
|
||||||
|
.name = t.name_buffer,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
total += 1;
|
||||||
|
}
|
||||||
|
return total;
|
||||||
|
}
|
||||||
|
|
||||||
/// Whether the running task is a user process (has its own address space).
|
/// Whether the running task is a user process (has its own address space).
|
||||||
pub fn currentIsUserProcess() bool {
|
pub fn currentIsUserProcess() bool {
|
||||||
return current().aspace != 0;
|
return current().aspace != 0;
|
||||||
|
|||||||
+190
-1
@@ -126,6 +126,12 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
|
|||||||
initTest(boot_information);
|
initTest(boot_information);
|
||||||
} else if (eql(case, "process")) {
|
} else if (eql(case, "process")) {
|
||||||
processTest(boot_information);
|
processTest(boot_information);
|
||||||
|
} else if (eql(case, "process-list")) {
|
||||||
|
processListTest(boot_information);
|
||||||
|
} else if (eql(case, "process-kill")) {
|
||||||
|
processKillTest(boot_information);
|
||||||
|
} else if (eql(case, "supervision")) {
|
||||||
|
supervisionTest(boot_information);
|
||||||
} else if (eql(case, "initial-ramdisk")) {
|
} else if (eql(case, "initial-ramdisk")) {
|
||||||
initialRamdiskTest(boot_information);
|
initialRamdiskTest(boot_information);
|
||||||
} else if (eql(case, "vfs")) {
|
} else if (eql(case, "vfs")) {
|
||||||
@@ -1222,7 +1228,7 @@ fn spawnFaultingProcess() bool {
|
|||||||
};
|
};
|
||||||
architecture.mapUserPageInto(aspace, process.stack_base_virtual, stack_frame, true, false); // RW + NX
|
architecture.mapUserPageInto(aspace, process.stack_base_virtual, stack_frame, true, false); // RW + NX
|
||||||
|
|
||||||
if (!scheduler.spawnUserLocked(aspace, process.code_virtual, process.stack_base_virtual + abi.page_size, 4, "fault-probe")) {
|
if (scheduler.spawnUserLocked(aspace, process.code_virtual, process.stack_base_virtual + abi.page_size, 4, "fault-probe", 0, null) == null) {
|
||||||
architecture.destroyAddressSpace(aspace);
|
architecture.destroyAddressSpace(aspace);
|
||||||
return false;
|
return false;
|
||||||
}
|
}
|
||||||
@@ -1311,6 +1317,189 @@ fn initTest(boot_information: *const BootInformation) void {
|
|||||||
result();
|
result();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// process_enumerate's kernel half: spawn two init processes next to the kernel
|
||||||
|
/// tasks and snapshot the table. The snapshot must list both by name with distinct,
|
||||||
|
/// kernel-supervised ids, include the kernel tasks (id 0, empty name), and report
|
||||||
|
/// the same total through a too-small buffer (the truncation contract: the caller
|
||||||
|
/// learns how big a buffer to bring).
|
||||||
|
fn processListTest(boot_information: *const BootInformation) void {
|
||||||
|
log("DANOS-TEST-BEGIN: process-list\n", .{});
|
||||||
|
check("bootloader handed over /system/services/init", boot_information.init_len != 0);
|
||||||
|
if (boot_information.init_len == 0) {
|
||||||
|
result();
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len];
|
||||||
|
|
||||||
|
var spawned: u32 = 0;
|
||||||
|
if (process.spawnProcess(image, 4, &.{"/system/services/init"})) spawned += 1 else |_| {}
|
||||||
|
if (process.spawnProcess(image, 4, &.{"/system/services/init"})) spawned += 1 else |_| {}
|
||||||
|
check("two init processes spawned", spawned == 2);
|
||||||
|
|
||||||
|
var table: [32]abi.ProcessDescriptor = undefined;
|
||||||
|
const total = scheduler.enumerate(&table);
|
||||||
|
check("enumerate counts the boot task and both processes (>=3)", total >= 3);
|
||||||
|
|
||||||
|
var inits: u32 = 0;
|
||||||
|
var init_ids: [2]u32 = .{ 0, 0 };
|
||||||
|
var kernel_task_seen = false;
|
||||||
|
var states_sane = true;
|
||||||
|
for (table[0..@min(total, table.len)]) |descriptor| {
|
||||||
|
if (descriptor.state > @intFromEnum(abi.ProcessState.blocked)) states_sane = false;
|
||||||
|
if (descriptor.name_length == 0) kernel_task_seen = true;
|
||||||
|
if (eql(descriptor.name[0..descriptor.name_length], "/system/services/init")) {
|
||||||
|
if (inits < 2) init_ids[inits] = descriptor.id;
|
||||||
|
inits += 1;
|
||||||
|
check("init entry is kernel-supervised (supervisor 0)", descriptor.supervisor == 0);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
check("both init processes listed by name", inits == 2);
|
||||||
|
check("listed processes carry distinct ids", init_ids[0] != init_ids[1]);
|
||||||
|
check("kernel tasks are listed too (empty name)", kernel_task_seen);
|
||||||
|
check("every state is a ProcessState value", states_sane);
|
||||||
|
|
||||||
|
var one: [1]abi.ProcessDescriptor = undefined;
|
||||||
|
check("a too-small buffer still learns the true total", scheduler.enumerate(&one) == total);
|
||||||
|
result();
|
||||||
|
}
|
||||||
|
|
||||||
|
/// process_kill + the exit notification, kernel half. Two victims, two paths:
|
||||||
|
/// - init, which heartbeats and sleeps: caught blocked, reaped on the killer's
|
||||||
|
/// own call — and its heartbeat must stop.
|
||||||
|
/// - process-test's spinner role (from the initial ramdisk), which loops in user
|
||||||
|
/// mode making no system calls: with more cores it is caught running, taking
|
||||||
|
/// the deferred path (kill_pending, finished by the victim core's next tick).
|
||||||
|
/// Each death must post one exit notification badge (exit bit + the child's id)
|
||||||
|
/// on the endpoint given at spawn; wrong-supervisor and unknown-id kills must be
|
||||||
|
/// refused. The waits block in replyWait, so a lost notification times the
|
||||||
|
/// harness out rather than passing vacuously.
|
||||||
|
fn processKillTest(boot_information: *const BootInformation) void {
|
||||||
|
log("DANOS-TEST-BEGIN: process-kill\n", .{});
|
||||||
|
check("bootloader handed over /system/services/init", boot_information.init_len != 0);
|
||||||
|
if (boot_information.init_len == 0 or boot_information.initial_ramdisk_len == 0) {
|
||||||
|
check("bootloader handed over an initial_ramdisk", boot_information.initial_ramdisk_len != 0);
|
||||||
|
result();
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.init_base)))[0..boot_information.init_len];
|
||||||
|
const ramdisk = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
|
||||||
|
const rd = initial_ramdisk.Reader.init(ramdisk) orelse {
|
||||||
|
check("initial_ramdisk image is valid", false);
|
||||||
|
result();
|
||||||
|
return;
|
||||||
|
};
|
||||||
|
|
||||||
|
const me = scheduler.currentId();
|
||||||
|
const endpoint = ipcsync.createIpcEndpoint() orelse {
|
||||||
|
check("exit endpoint allocated", false);
|
||||||
|
result();
|
||||||
|
return;
|
||||||
|
};
|
||||||
|
|
||||||
|
process.write_count = 0;
|
||||||
|
const sleeper = process.spawnProcessSupervised(image, 4, &.{"/system/services/init"}, me, endpoint) catch 0;
|
||||||
|
check("init spawned as the supervised sleeper victim", sleeper != 0);
|
||||||
|
|
||||||
|
// Let it reach its heartbeat loop (write, then a 1 s sleep) so the kill most
|
||||||
|
// likely catches it blocked.
|
||||||
|
scheduler.setPriority(1);
|
||||||
|
const deadline = architecture.millis() + 8000;
|
||||||
|
while (process.write_count < 1 and architecture.millis() < deadline) scheduler.yield();
|
||||||
|
scheduler.setPriority(4);
|
||||||
|
check("victim heartbeat before the kill", process.write_count >= 1);
|
||||||
|
|
||||||
|
// Kills that must be refused, before the one that must not be.
|
||||||
|
check("a non-supervisor may not kill (-EPERM)", process.killProcess(me + 12345, sleeper) == -ipcsync.EPERM);
|
||||||
|
check("an unknown id misses (-ESRCH)", process.killProcess(me, 0xFFFF_FF00) == -ipcsync.ESRCH);
|
||||||
|
check("a kernel task is not a killable process (-ESRCH)", process.killProcess(me, 0) == -ipcsync.ESRCH);
|
||||||
|
|
||||||
|
check("the supervisor's kill is accepted", process.killProcess(me, sleeper) == 0);
|
||||||
|
|
||||||
|
var badge: u64 = 0;
|
||||||
|
var received_cap: u64 = 0;
|
||||||
|
var r = ipcsync.replyWait(endpoint, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap);
|
||||||
|
check("the sleeper's exit notification arrived (length 0)", r == 0);
|
||||||
|
check("its badge carries the exit bit and the child id", badge == abi.notify_badge_bit | abi.notify_exit_bit | sleeper);
|
||||||
|
|
||||||
|
const beats_at_kill = process.write_count;
|
||||||
|
scheduler.sleep(1500); // more than one heartbeat period
|
||||||
|
check("the heartbeat stopped with the kill", process.write_count == beats_at_kill);
|
||||||
|
|
||||||
|
// The spinner: no system calls, so only the tick can deliver a deferred kill.
|
||||||
|
var spinner: u32 = 0;
|
||||||
|
var i: u32 = 0;
|
||||||
|
while (i < rd.count) : (i += 1) {
|
||||||
|
const item = rd.entry(i) orelse continue;
|
||||||
|
if (!eql(item.name, "process-test")) continue;
|
||||||
|
spinner = process.spawnProcessSupervised(item.blob, 4, &.{ "process-test", "spinner" }, me, endpoint) catch 0;
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
check("process-test spawned as the supervised spinner victim", spinner != 0);
|
||||||
|
scheduler.sleep(100); // give another core a chance to be running it
|
||||||
|
check("the spinner's kill is accepted", process.killProcess(me, spinner) == 0);
|
||||||
|
r = ipcsync.replyWait(endpoint, 0, 0, 0, 0, abi.no_cap, &badge, &received_cap);
|
||||||
|
check("the spinner's exit notification arrived (length 0)", r == 0);
|
||||||
|
check("its badge carries the exit bit and the child id", badge == abi.notify_badge_bit | abi.notify_exit_bit | spinner);
|
||||||
|
|
||||||
|
var table: [32]abi.ProcessDescriptor = undefined;
|
||||||
|
const total = scheduler.enumerate(&table);
|
||||||
|
var still_listed = false;
|
||||||
|
for (table[0..@min(total, table.len)]) |descriptor| {
|
||||||
|
if (descriptor.id == sleeper or descriptor.id == spinner) still_listed = true;
|
||||||
|
}
|
||||||
|
check("neither victim is listed after its kill", !still_listed);
|
||||||
|
check("a killed id stays dead (-ESRCH on a second kill)", process.killProcess(me, sleeper) == -ipcsync.ESRCH);
|
||||||
|
result();
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The whole user-side surface at once: spawn process-test's supervisor role,
|
||||||
|
/// which — entirely from ring 3 — creates an exit endpoint, spawns its two
|
||||||
|
/// children supervised, sees them in process_enumerate, kills them (one blocked,
|
||||||
|
/// one spinning), collects both exit notifications, and confirms they are gone.
|
||||||
|
/// Its "process-test: ok" is the pass marker; any FAIL line is specific.
|
||||||
|
fn supervisionTest(boot_information: *const BootInformation) void {
|
||||||
|
log("DANOS-TEST-BEGIN: supervision\n", .{});
|
||||||
|
check("bootloader handed over an initial_ramdisk", boot_information.initial_ramdisk_len != 0);
|
||||||
|
if (boot_information.initial_ramdisk_len == 0) {
|
||||||
|
result();
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
const ramdisk = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
|
||||||
|
const rd = initial_ramdisk.Reader.init(ramdisk) orelse {
|
||||||
|
check("initial_ramdisk image is valid", false);
|
||||||
|
result();
|
||||||
|
return;
|
||||||
|
};
|
||||||
|
process.setInitialRamdisk(ramdisk); // the supervisor system_spawns its children by name
|
||||||
|
|
||||||
|
process.write_count = 0;
|
||||||
|
process.write_from_user = false;
|
||||||
|
var started = false;
|
||||||
|
var i: u32 = 0;
|
||||||
|
while (i < rd.count) : (i += 1) {
|
||||||
|
const item = rd.entry(i) orelse continue;
|
||||||
|
if (!eql(item.name, "process-test")) continue;
|
||||||
|
started = if (process.spawnProcess(item.blob, 4, &.{ "process-test", "run" })) true else |_| false;
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
check("process-test spawned as the user-space supervisor", started);
|
||||||
|
|
||||||
|
const marker = "process-test: ok";
|
||||||
|
scheduler.setPriority(1);
|
||||||
|
const deadline = architecture.millis() + 10000;
|
||||||
|
while (architecture.millis() < deadline) {
|
||||||
|
if (process.write_len >= marker.len and eql(process.write_buffer[0..marker.len], marker)) break;
|
||||||
|
scheduler.yield();
|
||||||
|
}
|
||||||
|
scheduler.setPriority(4);
|
||||||
|
|
||||||
|
const ok = process.write_len >= marker.len and eql(process.write_buffer[0..marker.len], marker);
|
||||||
|
if (!ok and process.write_len > 0) log("DANOS-SUPERVISION: got \"{s}\"\n", .{process.write_buffer[0..process.write_len]});
|
||||||
|
check("the supervisor completed every step (spawn/list/kill/notify)", ok);
|
||||||
|
check("it ran in user mode (CPL 3)", process.write_from_user);
|
||||||
|
result();
|
||||||
|
}
|
||||||
|
|
||||||
/// The initial_ramdisk path: the bootloader handed over an image bundling extra user
|
/// The initial_ramdisk path: the bootloader handed over an image bundling extra user
|
||||||
/// binaries; parse it, spawn every program, and confirm one (the vfs stub)
|
/// binaries; parse it, spawn every program, and confirm one (the vfs stub)
|
||||||
/// reaches ring 3 and heartbeats — proving the whole ferry-parse-spawn pipeline.
|
/// reaches ring 3 and heartbeats — proving the whole ferry-parse-spawn pipeline.
|
||||||
|
|||||||
@@ -38,7 +38,7 @@ pub fn main() void {
|
|||||||
for (buffer[0..count]) |descriptor| {
|
for (buffer[0..count]) |descriptor| {
|
||||||
const driver_name = driverFor(descriptor.class) orelse continue;
|
const driver_name = driverFor(descriptor.class) orelse continue;
|
||||||
matched += 1;
|
matched += 1;
|
||||||
if (runtime.system.spawn(driver_name)) {
|
if (runtime.system.spawn(driver_name) != null) {
|
||||||
_ = runtime.system.write("device-manager: spawned ");
|
_ = runtime.system.write("device-manager: spawned ");
|
||||||
_ = runtime.system.write(driver_name);
|
_ = runtime.system.write(driver_name);
|
||||||
_ = runtime.system.write("\n");
|
_ = runtime.system.write("\n");
|
||||||
|
|||||||
@@ -0,0 +1,100 @@
|
|||||||
|
//! process-test — a test fixture for process management (bundled in the
|
||||||
|
//! initial-ramdisk, driven by the `supervision` test case). One binary, three
|
||||||
|
//! roles picked by argv, so the whole user-side surface is exercised end to end:
|
||||||
|
//!
|
||||||
|
//! - `process-test run` — the supervisor: spawns the two children below with an
|
||||||
|
//! exit-notification endpoint, sees them in `process_enumerate`, kills them,
|
||||||
|
//! receives both exit notifications, and confirms they are gone. Prints
|
||||||
|
//! "process-test: ok" for the kernel test to match, or a FAIL line naming the
|
||||||
|
//! step that broke.
|
||||||
|
//! - `process-test sleeper` — a child that blocks in `sleep` forever: its kill
|
||||||
|
//! exercises the immediate reap of a blocked task.
|
||||||
|
//! - `process-test spinner` — a child that spins in user mode making no system
|
||||||
|
//! calls: its kill exercises the deferred path (kill_pending, finished by the
|
||||||
|
//! timer tick).
|
||||||
|
//!
|
||||||
|
//! Spawned with no arguments (the initial-ramdisk sweep test starts every bundled
|
||||||
|
//! binary bare), it exits silently so it cannot derange other tests' output.
|
||||||
|
|
||||||
|
const std = @import("std");
|
||||||
|
const runtime = @import("runtime");
|
||||||
|
|
||||||
|
fn fail(step: []const u8) noreturn {
|
||||||
|
_ = runtime.system.write("process-test: FAIL ");
|
||||||
|
_ = runtime.system.write(step);
|
||||||
|
_ = runtime.system.write("\n");
|
||||||
|
runtime.system.exit(1);
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Whether process `id` appears in a fresh `process_enumerate` snapshot, named
|
||||||
|
/// `name` (an id present under the wrong name is a table mix-up, not a pass).
|
||||||
|
fn listed(id: u32, name: []const u8) bool {
|
||||||
|
var table: [32]runtime.system.ProcessDescriptor = undefined;
|
||||||
|
const total = runtime.system.processes(&table);
|
||||||
|
for (table[0..@min(total, table.len)]) |descriptor| {
|
||||||
|
if (descriptor.id != id) continue;
|
||||||
|
return std.mem.eql(u8, descriptor.name[0..descriptor.name_length], name);
|
||||||
|
}
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Block on the exit endpoint until a child-exit notification arrives; returns
|
||||||
|
/// the ended child's id. A wrong wake-up (there should be none — nothing else
|
||||||
|
/// knows this endpoint) fails the test rather than looping forever.
|
||||||
|
fn awaitChildExit(endpoint: runtime.ipc.Handle) u32 {
|
||||||
|
var scratch: [8]u8 = undefined;
|
||||||
|
const received = runtime.ipc.replyWait(endpoint, scratch[0..0], &scratch, null);
|
||||||
|
if (!received.isChildExit()) fail("expected a child-exit notification");
|
||||||
|
return received.childProcessId();
|
||||||
|
}
|
||||||
|
|
||||||
|
pub fn main() void {
|
||||||
|
if (runtime.argumentCount() <= 1) return; // spawned bare (ramdisk sweep): stay silent
|
||||||
|
|
||||||
|
const role = runtime.argument(1);
|
||||||
|
if (std.mem.eql(u8, role, "sleeper")) {
|
||||||
|
while (true) runtime.system.sleep(500);
|
||||||
|
}
|
||||||
|
if (std.mem.eql(u8, role, "spinner")) {
|
||||||
|
var beat: u64 = 0;
|
||||||
|
const touch: *volatile u64 = &beat;
|
||||||
|
while (true) touch.* +%= 1; // user mode only — no system calls to die at
|
||||||
|
}
|
||||||
|
|
||||||
|
// The supervisor ("run").
|
||||||
|
const endpoint = runtime.ipc.createIpcEndpoint() orelse fail("create exit endpoint");
|
||||||
|
|
||||||
|
const sleeper = runtime.system.spawnSupervised("process-test", &.{"sleeper"}, endpoint) orelse fail("spawn sleeper");
|
||||||
|
const spinner = runtime.system.spawnSupervised("process-test", &.{"spinner"}, endpoint) orelse fail("spawn spinner");
|
||||||
|
|
||||||
|
runtime.system.sleep(100); // let the sleeper block and the spinner get a core
|
||||||
|
if (!listed(sleeper, "process-test")) fail("sleeper not in process_enumerate");
|
||||||
|
if (!listed(spinner, "process-test")) fail("spinner not in process_enumerate");
|
||||||
|
|
||||||
|
// Kills that must be refused: a kernel task (id 0), and an id that was never
|
||||||
|
// issued — both -ESRCH. (-EPERM needs a second supervisor; the kernel-level
|
||||||
|
// `process-kill` test covers it.)
|
||||||
|
if (runtime.system.kill(0)) fail("killing a kernel task was allowed");
|
||||||
|
if (runtime.system.kill(0xFFFF_FFF0)) fail("killing an unknown id was allowed");
|
||||||
|
|
||||||
|
// The blocked child: usually reaped on the spot (it sits in `sleep`). The
|
||||||
|
// notification is the fence — after it, the child is certainly gone, so the
|
||||||
|
// second kill must miss (its id is never reused).
|
||||||
|
if (!runtime.system.kill(sleeper)) fail("kill sleeper");
|
||||||
|
if (awaitChildExit(endpoint) != sleeper) fail("sleeper exit notification");
|
||||||
|
if (runtime.system.kill(sleeper)) fail("double kill was allowed");
|
||||||
|
|
||||||
|
// The running child: the deferred path — condemned now, dead by the next tick.
|
||||||
|
if (!runtime.system.kill(spinner)) fail("kill spinner");
|
||||||
|
if (awaitChildExit(endpoint) != spinner) fail("spinner exit notification");
|
||||||
|
|
||||||
|
if (listed(sleeper, "process-test")) fail("sleeper still listed after kill");
|
||||||
|
if (listed(spinner, "process-test")) fail("spinner still listed after kill");
|
||||||
|
|
||||||
|
_ = runtime.system.write("process-test: ok\n");
|
||||||
|
}
|
||||||
|
|
||||||
|
pub const panic = runtime.panic;
|
||||||
|
comptime {
|
||||||
|
_ = &runtime.start._start; // pull the runtime entry shim into the image
|
||||||
|
}
|
||||||
@@ -222,6 +222,24 @@ CASES = [
|
|||||||
"smp": 4,
|
"smp": 4,
|
||||||
"expect": r"DANOS-TEST-RESULT: PASS",
|
"expect": r"DANOS-TEST-RESULT: PASS",
|
||||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||||
|
# process_enumerate: a task-table snapshot lists spawned processes by name and
|
||||||
|
# id alongside the kernel tasks, and a too-small buffer still reports the total.
|
||||||
|
{"name": "process-list",
|
||||||
|
"expect": r"DANOS-TEST-RESULT: PASS",
|
||||||
|
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||||
|
# process_kill + the exit notification: only the supervisor may kill; a blocked
|
||||||
|
# victim is reaped in place and a spinning one dies by the deferred (tick) path;
|
||||||
|
# each death posts one exit badge to the endpoint given at spawn.
|
||||||
|
{"name": "process-kill",
|
||||||
|
"smp": 4,
|
||||||
|
"expect": r"DANOS-TEST-RESULT: PASS",
|
||||||
|
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||||
|
# The user-side whole: process-test supervises two children from ring 3 —
|
||||||
|
# spawn with an exit endpoint, enumerate, kill, notification, gone.
|
||||||
|
{"name": "supervision",
|
||||||
|
"smp": 4,
|
||||||
|
"expect": r"DANOS-TEST-RESULT: PASS",
|
||||||
|
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||||
# The initial_ramdisk: the loader ferries a bundle of user binaries; the kernel parses
|
# The initial_ramdisk: the loader ferries a bundle of user binaries; the kernel parses
|
||||||
# it and spawns each as a ring-3 process (here the VFS-server stub heartbeats).
|
# it and spawns each as a ring-3 process (here the VFS-server stub heartbeats).
|
||||||
{"name": "initial-ramdisk",
|
{"name": "initial-ramdisk",
|
||||||
|
|||||||
Reference in New Issue
Block a user