kernel: a hostile return address cannot fault the kernel
SYSRETQ with a non-canonical RIP raises a general protection fault in ring 0 — on the kernel stack, an instruction after the swapgs that installed the user's GS base. It is one of the better-known escalation primitives, and ring 3 reaches it without any kernel bug at all: the processor saves the address of the instruction after SYSCALL, so a program whose SYSCALL is the last two bytes of the last canonical page returns to the first non-canonical address. The new test does exactly that. The exit path now sign-extends the return address from bit 47 and compares; if the value changed, it returns through IRETQ instead, which commits the privilege change before fetching the new address, so the fault arrives from ring 3 and the process dies like any other. Four register-only operations and a branch that a correct program can never take — it could not have executed at a non-canonical address in the first place. Bit 47 is the right pivot because danos builds four-level page tables and nothing sets the five-level bit; a future port must move the pivot, and the comment says so. SFMASK grows one bit while we are here. SYSCALL, unlike an interrupt gate, does not clear the nested-task flag, so the kernel had been running every system call with whatever ring 3 last chose — harmless while the only exit was SYSRETQ, and a question worth not having now that one exit is IRETQ. The kernel is never nested; ring 3 still gets its own flag back. Suite 113/113. The new case asserts the refusal counter rather than the dying process: the emulator we test on kills it either way, so only the counter distinguishes a guard that ran from one that did not.
This commit is contained in:
@@ -80,7 +80,16 @@ pub fn initSystemCall() void {
|
||||
io.wrmsr(ia32_star, (@as(u64, 0x08) << 32) | (@as(u64, 0x10) << 48));
|
||||
const entry = @extern(*const anyopaque, .{ .name = "syscall_entry" });
|
||||
io.wrmsr(ia32_lstar, @intFromPtr(entry));
|
||||
io.wrmsr(ia32_sfmask, 0x4_0700); // clear IF, TF, DF, AC on entry
|
||||
// Clear IF, TF, DF, AC and NT on entry. The first four are the usual
|
||||
// hygiene; NT is here because SYSCALL, unlike an interrupt gate, does not
|
||||
// clear it for us, so without this the kernel runs every system call with
|
||||
// whatever nested-task bit ring 3 last chose — and the canonical-RIP guard's
|
||||
// cold path (isr.s) leaves through IRETQ, whose behaviour with NT set is a
|
||||
// corner of the manuals not worth depending on either way. Masking it costs
|
||||
// one bit and removes the question: the kernel is never nested, and ring 3
|
||||
// still gets its own NT back, from R11 on the fast path and from the frame
|
||||
// on the cold one.
|
||||
io.wrmsr(ia32_sfmask, 0x4_4700);
|
||||
}
|
||||
|
||||
// --- supervisor-mode hardening (CR4) ---------------------------------------
|
||||
|
||||
Reference in New Issue
Block a user