Kill a faulting user process instead of halting the machine

A CPU exception raised in ring 3 by a scheduled process now kills that
process - IRQ bindings, IPC handles, and address space reclaimed, a
client it owed a reply to failed with the new -EPEER instead of hung -
and the core reschedules (docs/resilience.md step 2). Kernel-mode
faults, NMI, double fault, and machine check stay terminal, as does the
borrowed-thread isolation probe. Proven by the new fault-recovery QEMU
test: init keeps heartbeating after a process page-faults to death.
This commit is contained in:
Daniel Samson
2026-07-11 04:57:02 +01:00
parent 59104dd988
commit 6b3ae0c997
8 changed files with 198 additions and 35 deletions
+15 -4
View File
@@ -80,11 +80,22 @@ inline). `build.zig` adds `isr.s` to the arch module.
## Reporting a fault
`isr_common` calls `exceptionHandler`, which forwards to a swappable `on_fault`
hook. The generic kernel installs a reporter (`onException` in `main.zig`) that
prints, in red, the exception name and vector, the error code, the faulting RIP
hook. The generic kernel installs a reporter (`onException` in `kernel.zig`) that
prints the exception name and vector, the error code, the faulting RIP
and RSP, and — for a page fault (#PF, vector 14) — the faulting address from
**CR2**. Then it halts. There's no fault *recovery* yet, so every exception is
terminal; the point is that it's now **visible** instead of a silent reset.
**CR2**. What happens next depends on where the fault came from:
- **User mode (CPL 3): kill the process, keep the machine.** The kernel is intact
(the CPU trapped onto the task's kernel stack), so the faulting process is
killed — address space, IRQ bindings, and IPC handles reclaimed; a client it
owed a reply to is failed with `-EPEER` — and the core reschedules. A crashing
driver takes itself down, never the OS. This is fault recovery step 2 of
[resilience.md](resilience.md). NMI, double fault, and machine check are
excluded: they report machine trouble regardless of what was running.
- **Kernel mode: halt this core.** The trusted base itself is broken, so there is
nothing safe to kill; the fault is still *contained* to the core (an
application-processor fault leaves the rest of the system running), and the
report makes it **visible** instead of a silent reset.
The hook is set before `arch.init()` in `kmain`, so a fault during setup is still
caught.
+9 -3
View File
@@ -1,6 +1,10 @@
# Resilience: fault isolation and live restart
A design/research note, not built yet. This is the property danos is really chasing:
Steps 1–2 of the ordering below are **built**: user-mode isolation, and fault →
kill the process → keep the core (`onException` in `system/kernel/kernel.zig`; the
`fault-recovery` test proves a crashing ring-3 process dies alone while the system
keeps running). The supervisor notification and restart policy (steps 3+) are
still design. This is the property danos is really chasing:
**if a part of the OS breaks, isolate it, and re-initialise it — without rebooting.**
A crashed driver gets restarted; a wedged service gets killed and brought back. It's
the reason the [microkernel](vision.md) shape was chosen, and it's a *separate* goal
@@ -111,9 +115,11 @@ Honest boundaries:
## Suggested ordering
1. **User mode + address-space isolation** — the shared prerequisite (also on the
path for everything else).
path for everything else). **Done.**
2. **Kernel: fault → kill process → notify.** Turn today's "halt on fault" into
"confine to the process and report it."
"confine to the process and report it." **Done** (the kill and reclaim; the
supervisor notification waits for step 3's supervisor). A killed server's
pending client is unblocked with `-EPEER` rather than hung.
3. **A minimal supervisor server** that can (re)start a process.
4. **Resource cleanup on death** — reclaim memory/MMIO/IPC/IRQ, via caps or a grant
table.