Kill a faulting user process instead of halting the machine
A CPU exception raised in ring 3 by a scheduled process now kills that process - IRQ bindings, IPC handles, and address space reclaimed, a client it owed a reply to failed with the new -EPEER instead of hung - and the core reschedules (docs/resilience.md step 2). Kernel-mode faults, NMI, double fault, and machine check stay terminal, as does the borrowed-thread isolation probe. Proven by the new fault-recovery QEMU test: init keeps heartbeating after a process page-faults to death.
This commit is contained in:
+41
-17
@@ -120,18 +120,10 @@ fn system_call(state: *architecture.CpuState) void {
|
||||
switch (@as(SystemCall, @enumFromInt(architecture.systemCallNumber(state)))) {
|
||||
.exit => {
|
||||
exit_code = architecture.systemCallArg(state, 0);
|
||||
// A scheduled process drops its endpoint references, frees its address
|
||||
// space, and reschedules; a borrowed test thread unwinds back to the
|
||||
// kernel that entered it.
|
||||
// A scheduled process tears down fully (terminateCurrent); a borrowed
|
||||
// test thread unwinds back to the kernel that entered it.
|
||||
if (scheduler.currentIsUserProcess()) {
|
||||
// Unbind before closeHandles: dropping the last reference destroys the
|
||||
// Endpoint, and a still-bound GSI would have an ISR call
|
||||
// notifyFromIsr on freed memory the next time the device fired.
|
||||
// unbindAll also leaves the line masked, so a dead driver's device
|
||||
// goes quiet rather than storming.
|
||||
releaseIrqs(scheduler.current());
|
||||
ipc.closeHandles(scheduler.current());
|
||||
scheduler.exitUser();
|
||||
terminateCurrent();
|
||||
} else architecture.userExit();
|
||||
},
|
||||
.yield => {
|
||||
@@ -428,12 +420,44 @@ fn systemSpawn(state: *architecture.CpuState) void {
|
||||
fail(state); // no bundled binary by that name
|
||||
}
|
||||
|
||||
/// Drop every IRQ binding `t` made. Called on exit, before the handle table is closed
|
||||
/// (which is what frees the endpoints an ISR would otherwise notify into).
|
||||
fn releaseIrqs(t: *scheduler.Task) void {
|
||||
const flags = sync.enter();
|
||||
defer sync.leave(flags);
|
||||
irq.releaseOwner(t.id);
|
||||
/// Processes killed by a CPU fault rather than a clean exit. Evidence for the
|
||||
/// fault-recovery test, and a health signal a supervisor can consult later.
|
||||
pub var fault_kill_count: u64 = 0;
|
||||
|
||||
/// Tear down the current user process and reschedule; never returns. Shared by the
|
||||
/// exit system call and the fault path (`killCurrentProcess`). The order matters:
|
||||
/// - IRQ bindings are dropped before the handle table closes: dropping the last
|
||||
/// endpoint reference destroys the Endpoint, and a still-bound GSI would have an
|
||||
/// ISR call notifyFromIsr on freed memory the next time the device fired.
|
||||
/// `releaseOwner` also leaves the line masked, so a dead driver's device goes
|
||||
/// quiet rather than storming.
|
||||
/// - A client this task still owes a reply to (it died between receive and reply)
|
||||
/// is failed with -EPEER rather than left blocked forever — a dead server must
|
||||
/// not hang its callers.
|
||||
pub fn terminateCurrent() noreturn {
|
||||
const t = scheduler.current();
|
||||
{
|
||||
const flags = sync.enter();
|
||||
defer sync.leave(flags);
|
||||
irq.releaseOwner(t.id);
|
||||
if (t.ipc_client) |client| {
|
||||
t.ipc_client = null;
|
||||
client.ipc_status = -ipc.EPEER;
|
||||
scheduler.readyLocked(client); // its blocked `call` now returns the error
|
||||
}
|
||||
ipc.closeHandles(t);
|
||||
}
|
||||
scheduler.exitUser();
|
||||
}
|
||||
|
||||
/// Kill the current user process in response to a CPU fault it raised in ring 3.
|
||||
/// The fault is confined to the process — the kernel trapped it on the task's own
|
||||
/// kernel stack and is intact — so everything the process held is reclaimed and the
|
||||
/// core reschedules. The system keeps running; only the faulting process dies
|
||||
/// (docs/resilience.md: fault -> kill -> continue).
|
||||
pub fn killCurrentProcess() noreturn {
|
||||
fault_kill_count += 1;
|
||||
terminateCurrent();
|
||||
}
|
||||
|
||||
/// Resolve `(device_id, resource_index)` to a GSI this process is entitled to bind, or null.
|
||||
|
||||
Reference in New Issue
Block a user