Kill a faulting user process instead of halting the machine
A CPU exception raised in ring 3 by a scheduled process now kills that process - IRQ bindings, IPC handles, and address space reclaimed, a client it owed a reply to failed with the new -EPEER instead of hung - and the core reschedules (docs/resilience.md step 2). Kernel-mode faults, NMI, double fault, and machine check stay terminal, as does the borrowed-thread isolation probe. Proven by the new fault-recovery QEMU test: init keeps heartbeating after a process page-faults to death.
This commit is contained in:
+15
-4
@@ -80,11 +80,22 @@ inline). `build.zig` adds `isr.s` to the arch module.
|
||||
## Reporting a fault
|
||||
|
||||
`isr_common` calls `exceptionHandler`, which forwards to a swappable `on_fault`
|
||||
hook. The generic kernel installs a reporter (`onException` in `main.zig`) that
|
||||
prints, in red, the exception name and vector, the error code, the faulting RIP
|
||||
hook. The generic kernel installs a reporter (`onException` in `kernel.zig`) that
|
||||
prints the exception name and vector, the error code, the faulting RIP
|
||||
and RSP, and — for a page fault (#PF, vector 14) — the faulting address from
|
||||
**CR2**. Then it halts. There's no fault *recovery* yet, so every exception is
|
||||
terminal; the point is that it's now **visible** instead of a silent reset.
|
||||
**CR2**. What happens next depends on where the fault came from:
|
||||
|
||||
- **User mode (CPL 3): kill the process, keep the machine.** The kernel is intact
|
||||
(the CPU trapped onto the task's kernel stack), so the faulting process is
|
||||
killed — address space, IRQ bindings, and IPC handles reclaimed; a client it
|
||||
owed a reply to is failed with `-EPEER` — and the core reschedules. A crashing
|
||||
driver takes itself down, never the OS. This is fault recovery step 2 of
|
||||
[resilience.md](resilience.md). NMI, double fault, and machine check are
|
||||
excluded: they report machine trouble regardless of what was running.
|
||||
- **Kernel mode: halt this core.** The trusted base itself is broken, so there is
|
||||
nothing safe to kill; the fault is still *contained* to the core (an
|
||||
application-processor fault leaves the rest of the system running), and the
|
||||
report makes it **visible** instead of a silent reset.
|
||||
|
||||
The hook is set before `arch.init()` in `kmain`, so a fault during setup is still
|
||||
caught.
|
||||
|
||||
Reference in New Issue
Block a user