Kill a faulting user process instead of halting the machine
A CPU exception raised in ring 3 by a scheduled process now kills that process - IRQ bindings, IPC handles, and address space reclaimed, a client it owed a reply to failed with the new -EPEER instead of hung - and the core reschedules (docs/resilience.md step 2). Kernel-mode faults, NMI, double fault, and machine check stay terminal, as does the borrowed-thread isolation probe. Proven by the new fault-recovery QEMU test: init keeps heartbeating after a process page-faults to death.
This commit is contained in:
@@ -199,6 +199,12 @@ CASES = [
|
||||
{"name": "user-pf",
|
||||
"expect": r"page fault \(vector 14\)[\s\S]*error code : 0x5[\s\S]*IP\s*: 0x00007000000000",
|
||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||
# Fault recovery: a scheduled ring-3 process that faults is killed — resources
|
||||
# reclaimed, core kept — while init's heartbeat proves the OS survived.
|
||||
{"name": "fault-recovery",
|
||||
"timeout": 60,
|
||||
"expect": r"DANOS-TEST-RESULT: PASS",
|
||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||
# The real user binary: the bootloader ships /system/services/init off the ESP, the
|
||||
# kernel loads the ELF and runs it in ring 3, and it writes + exits cleanly.
|
||||
{"name": "init",
|
||||
|
||||
Reference in New Issue
Block a user