Kill a faulting user process instead of halting the machine

A CPU exception raised in ring 3 by a scheduled process now kills that
process - IRQ bindings, IPC handles, and address space reclaimed, a
client it owed a reply to failed with the new -EPEER instead of hung -
and the core reschedules (docs/resilience.md step 2). Kernel-mode
faults, NMI, double fault, and machine check stay terminal, as does the
borrowed-thread isolation probe. Proven by the new fault-recovery QEMU
test: init keeps heartbeating after a process page-faults to death.
This commit is contained in:
Daniel Samson
2026-07-11 04:57:02 +01:00
parent 59104dd988
commit 6b3ae0c997
8 changed files with 198 additions and 35 deletions
+6
View File
@@ -199,6 +199,12 @@ CASES = [
{"name": "user-pf",
"expect": r"page fault \(vector 14\)[\s\S]*error code : 0x5[\s\S]*IP\s*: 0x00007000000000",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Fault recovery: a scheduled ring-3 process that faults is killed — resources
# reclaimed, core kept — while init's heartbeat proves the OS survived.
{"name": "fault-recovery",
"timeout": 60,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# The real user binary: the bootloader ships /system/services/init off the ESP, the
# kernel loads the ELF and runs it in ring 3, and it writes + exits cleanly.
{"name": "init",