kernel: a hostile return address cannot fault the kernel
SYSRETQ with a non-canonical RIP raises a general protection fault in ring 0 — on the kernel stack, an instruction after the swapgs that installed the user's GS base. It is one of the better-known escalation primitives, and ring 3 reaches it without any kernel bug at all: the processor saves the address of the instruction after SYSCALL, so a program whose SYSCALL is the last two bytes of the last canonical page returns to the first non-canonical address. The new test does exactly that. The exit path now sign-extends the return address from bit 47 and compares; if the value changed, it returns through IRETQ instead, which commits the privilege change before fetching the new address, so the fault arrives from ring 3 and the process dies like any other. Four register-only operations and a branch that a correct program can never take — it could not have executed at a non-canonical address in the first place. Bit 47 is the right pivot because danos builds four-level page tables and nothing sets the five-level bit; a future port must move the pivot, and the comment says so. SFMASK grows one bit while we are here. SYSCALL, unlike an interrupt gate, does not clear the nested-task flag, so the kernel had been running every system call with whatever ring 3 last chose — harmless while the only exit was SYSRETQ, and a question worth not having now that one exit is IRETQ. The kernel is never nested; ring 3 still gets its own flag back. Suite 113/113. The new case asserts the refusal counter rather than the dying process: the emulator we test on kills it either way, so only the counter distinguishes a guard that ran from one that did not.
This commit is contained in:
@@ -355,6 +355,21 @@ CASES = [
|
||||
r"(?=.*ring 0 \(kernel\))(?=.*error code : 0x11)"
|
||||
r"(?=.*fault addr : 0x0000700000000000)",
|
||||
"fail": r"SMEP not enforced|SMEP not enabled|no frame for the probe page"},
|
||||
# SYSRET canonical-RIP guard: a ring-3 probe whose `syscall` sits on the last two
|
||||
# bytes of executable address space, so the return RIP the kernel must hand back
|
||||
# is 0x0000800000000000 — non-canonical, and a ring-0 #GP if it reached SYSRETQ.
|
||||
# The kernel must return through IRETQ instead, which faults the *process* in
|
||||
# ring 3: the kill line names a general protection fault at exactly that address,
|
||||
# and the probe's own result line reports the guard counter. Any "ring 0 (kernel)"
|
||||
# report, panic or reset here is the hazard itself, so they are failures.
|
||||
# TCG does not model the ring-0 #GP (measured: with the guard removed these two
|
||||
# lines are unchanged), so the case's teeth are the counter assertion inside
|
||||
# DANOS-TEST-RESULT, not the kill line.
|
||||
{"name": "sysret-canonical",
|
||||
"expect": r"(?s)(?=.*killed by general protection fault \(vector 13\))"
|
||||
r"(?=.*IP\s+: 0x0000800000000000)"
|
||||
r"(?=.*DANOS-TEST-RESULT: PASS)",
|
||||
"fail": r"DANOS-TEST-RESULT: FAIL|CPU EXCEPTION|KERNEL PANIC"},
|
||||
# Memory grants: the mmap/munmap path hands out user pages into a process's
|
||||
# arena, translate resolves them, munmap frees them, and no frames leak.
|
||||
{"name": "usermem",
|
||||
|
||||
Reference in New Issue
Block a user