kernel: a hostile return address cannot fault the kernel
SYSRETQ with a non-canonical RIP raises a general protection fault in ring 0 — on the kernel stack, an instruction after the swapgs that installed the user's GS base. It is one of the better-known escalation primitives, and ring 3 reaches it without any kernel bug at all: the processor saves the address of the instruction after SYSCALL, so a program whose SYSCALL is the last two bytes of the last canonical page returns to the first non-canonical address. The new test does exactly that. The exit path now sign-extends the return address from bit 47 and compares; if the value changed, it returns through IRETQ instead, which commits the privilege change before fetching the new address, so the fault arrives from ring 3 and the process dies like any other. Four register-only operations and a branch that a correct program can never take — it could not have executed at a non-canonical address in the first place. Bit 47 is the right pivot because danos builds four-level page tables and nothing sets the five-level bit; a future port must move the pivot, and the comment says so. SFMASK grows one bit while we are here. SYSCALL, unlike an interrupt gate, does not clear the nested-task flag, so the kernel had been running every system call with whatever ring 3 last chose — harmless while the only exit was SYSRETQ, and a question worth not having now that one exit is IRETQ. The kernel is never nested; ring 3 still gets its own flag back. Suite 113/113. The new case asserts the refusal counter rather than the dying process: the emulator we test on kills it either way, so only the counter distinguishes a guard that ran from one that did not.
This commit is contained in:
@@ -315,6 +315,20 @@ pub fn userExit() noreturn {
|
||||
user_exit_to_kernel();
|
||||
}
|
||||
|
||||
/// The counter the syscall exit stub bumps when it refuses to return the fast way
|
||||
/// (isr.s, `sysret_non_canonical_count`).
|
||||
const non_canonical_returns = @extern(*u64, .{ .name = "sysret_non_canonical_count" });
|
||||
|
||||
/// How many times this boot's system_call returns took the slow, safe exit because
|
||||
/// the return address was not a canonical address (the SYSRETQ canonical-RIP guard
|
||||
/// in isr.s; an architecture without the hazard reports 0 forever). A correct
|
||||
/// program cannot produce one — it could not have executed at a non-canonical
|
||||
/// address in the first place — so a nonzero count is exactly the number of times
|
||||
/// a process tried to make the kernel fault on its way out.
|
||||
pub fn nonCanonicalReturnCount() u64 {
|
||||
return @atomicLoad(u64, non_canonical_returns, .monotonic);
|
||||
}
|
||||
|
||||
/// Register the handler for the user system_call gate (int 0x80, vector 128). The
|
||||
/// handler may write the trap frame (see `setSystemCallResult`).
|
||||
pub fn setSystemCallHandler(handler: *const fn (*CpuState) void) void {
|
||||
|
||||
Reference in New Issue
Block a user