Kill a faulting user process instead of halting the machine

A CPU exception raised in ring 3 by a scheduled process now kills that
process - IRQ bindings, IPC handles, and address space reclaimed, a
client it owed a reply to failed with the new -EPEER instead of hung -
and the core reschedules (docs/resilience.md step 2). Kernel-mode
faults, NMI, double fault, and machine check stay terminal, as does the
borrowed-thread isolation probe. Proven by the new fault-recovery QEMU
test: init keeps heartbeating after a process page-faults to death.
This commit is contained in:
Daniel Samson
2026-07-11 04:57:02 +01:00
parent 59104dd988
commit 6b3ae0c997
8 changed files with 198 additions and 35 deletions
+9 -3
View File
@@ -1,6 +1,10 @@
# Resilience: fault isolation and live restart
A design/research note, not built yet. This is the property danos is really chasing:
Steps 1–2 of the ordering below are **built**: user-mode isolation, and fault →
kill the process → keep the core (`onException` in `system/kernel/kernel.zig`; the
`fault-recovery` test proves a crashing ring-3 process dies alone while the system
keeps running). The supervisor notification and restart policy (steps 3+) are
still design. This is the property danos is really chasing:
**if a part of the OS breaks, isolate it, and re-initialise it — without rebooting.**
A crashed driver gets restarted; a wedged service gets killed and brought back. It's
the reason the [microkernel](vision.md) shape was chosen, and it's a *separate* goal
@@ -111,9 +115,11 @@ Honest boundaries:
## Suggested ordering
1. **User mode + address-space isolation** — the shared prerequisite (also on the
path for everything else).
path for everything else). **Done.**
2. **Kernel: fault → kill process → notify.** Turn today's "halt on fault" into
"confine to the process and report it."
"confine to the process and report it." **Done** (the kill and reclaim; the
supervisor notification waits for step 3's supervisor). A killed server's
pending client is unblocked with `-EPEER` rather than hung.
3. **A minimal supervisor server** that can (re)start a process.
4. **Resource cleanup on death** — reclaim memory/MMIO/IPC/IRQ, via caps or a grant
table.