kernel: a hostile return address cannot fault the kernel
SYSRETQ with a non-canonical RIP raises a general protection fault in ring 0 — on the kernel stack, an instruction after the swapgs that installed the user's GS base. It is one of the better-known escalation primitives, and ring 3 reaches it without any kernel bug at all: the processor saves the address of the instruction after SYSCALL, so a program whose SYSCALL is the last two bytes of the last canonical page returns to the first non-canonical address. The new test does exactly that. The exit path now sign-extends the return address from bit 47 and compares; if the value changed, it returns through IRETQ instead, which commits the privilege change before fetching the new address, so the fault arrives from ring 3 and the process dies like any other. Four register-only operations and a branch that a correct program can never take — it could not have executed at a non-canonical address in the first place. Bit 47 is the right pivot because danos builds four-level page tables and nothing sets the five-level bit; a future port must move the pivot, and the comment says so. SFMASK grows one bit while we are here. SYSCALL, unlike an interrupt gate, does not clear the nested-task flag, so the kernel had been running every system call with whatever ring 3 last chose — harmless while the only exit was SYSRETQ, and a question worth not having now that one exit is IRETQ. The kernel is never nested; ring 3 still gets its own flag back. Suite 113/113. The new case asserts the refusal counter rather than the dying process: the emulator we test on kills it either way, so only the counter distinguishes a guard that ran from one that did not.
This commit is contained in:
@@ -124,15 +124,26 @@ pub const maximum_mount_rewrite = 32;
|
||||
const auxiliary_vector_null: u64 = 0; // AT_NULL — end of the vector
|
||||
const auxiliary_vector_page_size: u64 = 6; // AT_PAGESZ
|
||||
|
||||
// The hand-assembled user program blob (isr.s, .rodata) — the isolation probe.
|
||||
// The hand-assembled user program blobs (isr.s, .rodata) — the isolation probes.
|
||||
const pf_start = @extern([*]const u8, .{ .name = "user_pf_start" });
|
||||
const pf_end = @extern([*]const u8, .{ .name = "user_pf_end" });
|
||||
const sysret_start = @extern([*]const u8, .{ .name = "user_sysret_start" });
|
||||
const sysret_end = @extern([*]const u8, .{ .name = "user_sysret_end" });
|
||||
|
||||
/// The isolation-proof program: reads a kernel-only page, must #PF.
|
||||
pub fn pfBlob() []const u8 {
|
||||
return pf_start[0 .. @intFromPtr(pf_end) - @intFromPtr(pf_start)];
|
||||
}
|
||||
|
||||
/// The SYSRET-guard program: two system calls, the second of which must be copied
|
||||
/// so that it ends at the last executable byte of the user half — its return
|
||||
/// address is then the first non-canonical address. Ends with `syscall`, so the
|
||||
/// caller places it at `page_size - len` inside the page at `user_half_end -
|
||||
/// page_size`; anywhere else and it proves nothing.
|
||||
pub fn nonCanonicalReturnBlob() []const u8 {
|
||||
return sysret_start[0 .. @intFromPtr(sysret_end) - @intFromPtr(sysret_start)];
|
||||
}
|
||||
|
||||
/// What debug_write syscalls produced (accumulated), and the exit system_call's code.
|
||||
pub var write_buffer: [256]u8 = undefined;
|
||||
pub var write_len: usize = 0;
|
||||
|
||||
Reference in New Issue
Block a user