Add process management: enumerate, supervisor-gated kill, exit notifications
process_enumerate snapshots the task table (the device_enumerate shape, so ps is a user program); system_spawn returns the child id, records the caller as supervisor, and takes an exit endpoint; process_kill is allowed only for the supervisor. Every death — exit, fault, or kill — posts a child-exit badge to that endpoint (the IRQ-as-IPC pattern as SIGCHLD). A target caught off-CPU is reaped in place; a running one is condemned and finished at its next system call or tick, guarded so teardown never lands mid-kernel-operation. Tested by process-list, process-kill, and supervision (a ring-3 supervisor exercising the whole surface); design notes in docs/process-management.md.
This commit is contained in:
@@ -0,0 +1,112 @@
|
||||
# Process Management
|
||||
|
||||
How danos lists, supervises, and kills processes — the microkernel answer to
|
||||
`ps`, `kill`, and `SIGCHLD`/`wait`.
|
||||
|
||||
## Why system calls, not `/proc`
|
||||
|
||||
Unix systems sit on a spectrum. Classic BSD/macOS list processes through
|
||||
syscalls (`sysctl(KERN_PROC)`) and kill through `kill(2)`; Linux renders the
|
||||
process table as `/proc` for *reading* but still kills through a syscall; Plan 9
|
||||
made the file tree the whole interface (`echo kill > /proc/n/ctl`). Microkernels
|
||||
mostly abandon ambient PIDs: Minix and QNX route everything through a user-space
|
||||
process-manager server, and Fuchsia/seL4 control processes only through handles.
|
||||
|
||||
danos rules out `/proc` **as the primitive**: here a `/proc` would be served by
|
||||
the VFS server — a user process — which would put the VFS in the path of process
|
||||
control. If the VFS (or anything under it) hangs, nothing could be listed or
|
||||
killed, *including the hung VFS*. The control plane for processes must not
|
||||
depend on a process. So the primitives are kernel system calls; a read-only
|
||||
`/proc` rendering can be layered on later, and a POSIX-style process-manager
|
||||
server can be built *from* these primitives when one is needed.
|
||||
|
||||
## The three primitives
|
||||
|
||||
### `process_enumerate(buffer, maximum) -> total`
|
||||
|
||||
A snapshot of the task table into a caller buffer of `abi.ProcessDescriptor`
|
||||
(id, supervisor, state, priority, name) — the exact shape of
|
||||
`device_enumerate`, so `ps` is a user program over a snapshot, not a kernel
|
||||
service. The total may exceed what fit; call again with a larger buffer. Kernel
|
||||
tasks are included with an empty name — an honest listing shows the idle tasks
|
||||
too. Ungated and read-only: what is running is not a secret between cooperating
|
||||
bring-up processes.
|
||||
|
||||
### `system_spawn(..., exit_endpoint) -> child id`, and the supervision link
|
||||
|
||||
`system_spawn` records the caller as the child's **supervisor** and returns the
|
||||
child's process id (ids are monotonic, never reused — a stale id can only miss).
|
||||
That link is the kill authority: it answers "who may kill process 7?" without
|
||||
inventing users or permissions, the same way a device *claim* is the capability
|
||||
for `mmio_map`. It composes with the supervision hierarchy the device manager
|
||||
already forms: init supervises the services it starts, the device manager
|
||||
supervises the drivers it matches. (A transferable process *handle* — Fuchsia
|
||||
style — can replace the id once the handle table grows types beyond endpoints.)
|
||||
|
||||
`exit_endpoint` (a handle, or `abi.no_cap`) is the supervisor's death-watch: when
|
||||
the child ends — clean exit, CPU fault, or `process_kill` — the kernel posts an
|
||||
asynchronous notification to that endpoint, exactly like a bound IRQ. The badge
|
||||
carries `abi.notify_badge_bit | abi.notify_exit_bit | child_id`, so one endpoint
|
||||
supervises many children and can even share with IRQ notifications. This is the
|
||||
microkernel's SIGCHLD: no new mechanism, just the IRQ-as-IPC pattern reused, and
|
||||
a supervisor's event loop (`ipc.replyWait`) already knows how to receive it. The
|
||||
child holds a reference to the endpoint from birth, so the notification cannot
|
||||
dangle even if the supervisor dies first.
|
||||
|
||||
### `process_kill(id) -> 0 / -ESRCH / -EPERM`
|
||||
|
||||
Only the supervisor may kill; kernel tasks are not killable processes. Like a
|
||||
signal, delivery is prompt but asynchronous — 0 means the kill is accepted and
|
||||
irrevocable; the exit notification confirms completion.
|
||||
|
||||
## How a kill lands (the kernel mechanics)
|
||||
|
||||
Everything below runs under the big kernel lock, where task states cannot move.
|
||||
|
||||
- **Target ready or blocked** (not on any core): reaped on the killer's own
|
||||
call. The reap releases what death always releases (IRQ bindings first, then
|
||||
a client the target still owed a reply to is failed with `-EPEER`, IPC handles
|
||||
closed, the exit notification posted last) — plus the unlinking only a
|
||||
*remote* death needs: out of the ready queue, out of an endpoint's sender FIFO
|
||||
(`Task.ipc_wait_endpoint`), out of a receive wait queue (`Task.wait_queue`),
|
||||
and out of any server's owed-reply slot, so nothing ever dequeues a dangling
|
||||
pointer. Destroying the address space is safe because no core can have it
|
||||
loaded: every switch away from a task loads the next task's tables.
|
||||
- **Target running on another core**: it cannot be torn down mid-instruction,
|
||||
so it is condemned (`Task.kill_pending`) and dies at whichever comes first:
|
||||
- its next **system_call entry** — checked before dispatch, so a condemned
|
||||
process cannot spawn, claim, or message anything on its way out;
|
||||
- its core's next **timer tick** — but only when the task is not inside one
|
||||
of its own system calls (`Task.in_system_call`): the tick may have
|
||||
interrupted kernel code mid-operation, where teardown would leak whatever
|
||||
the operation held. User-mode execution is always a safe kill point. The
|
||||
tick-time terminate abandons the interrupt frame exactly like the fault
|
||||
path (the LAPIC is acknowledged before the tick hook runs);
|
||||
- any core's tick finding it **blocked or ready** (it entered a syscall and
|
||||
parked after being condemned) — reaped by the same remote-reap path.
|
||||
|
||||
A pure user-mode spin loop that never makes a system call therefore dies
|
||||
within one tick; nothing a process does can outrun the kill.
|
||||
|
||||
The scheduler stays below the process layer: finishing a kill (IRQ bindings,
|
||||
handles, the notification) is called *up* through two hooks process.zig
|
||||
registers at boot (`terminate_current_hook`, `reap_task_hook`), mirroring how
|
||||
the architecture layer calls up into `tick`.
|
||||
|
||||
## Known gaps (bring-up honesty)
|
||||
|
||||
- Device **claims** are not released on death (pre-existing: the fault path has
|
||||
the same gap) — a killed driver's device stays claimed until reboot.
|
||||
- Kernel stacks of dead tasks are leaked, as on every exit path (no reaper yet).
|
||||
- There is no exit *status* in the notification, only the id; a supervisor that
|
||||
needs the code can grow a wait-style call later.
|
||||
- Enumerate writes through the caller's raw pointer under the bring-up trust
|
||||
model, like `device_enumerate` (an unmapped page is a self-DoS, not an
|
||||
isolation break).
|
||||
|
||||
## Tests
|
||||
|
||||
`process-list` (enumerate), `process-kill` (kernel-level kill paths, refusals,
|
||||
notifications), `supervision` (the whole user-side surface via the process-test
|
||||
service: spawn supervised → enumerate → kill blocked and spinning children →
|
||||
notifications → gone). See test/qemu_test.py.
|
||||
Reference in New Issue
Block a user