kernel: a hostile return address cannot fault the kernel
SYSRETQ with a non-canonical RIP raises a general protection fault in ring 0 — on the kernel stack, an instruction after the swapgs that installed the user's GS base. It is one of the better-known escalation primitives, and ring 3 reaches it without any kernel bug at all: the processor saves the address of the instruction after SYSCALL, so a program whose SYSCALL is the last two bytes of the last canonical page returns to the first non-canonical address. The new test does exactly that. The exit path now sign-extends the return address from bit 47 and compares; if the value changed, it returns through IRETQ instead, which commits the privilege change before fetching the new address, so the fault arrives from ring 3 and the process dies like any other. Four register-only operations and a branch that a correct program can never take — it could not have executed at a non-canonical address in the first place. Bit 47 is the right pivot because danos builds four-level page tables and nothing sets the five-level bit; a future port must move the pivot, and the comment says so. SFMASK grows one bit while we are here. SYSCALL, unlike an interrupt gate, does not clear the nested-task flag, so the kernel had been running every system call with whatever ring 3 last chose — harmless while the only exit was SYSRETQ, and a question worth not having now that one exit is IRETQ. The kernel is never nested; ring 3 still gets its own flag back. Suite 113/113. The new case asserts the refusal counter rather than the dying process: the emulator we test on kills it either way, so only the counter distinguishes a guard that ran from one that did not.
This commit is contained in:
@@ -187,11 +187,9 @@ user_exit_to_kernel:
|
||||
# CS/SS from STAR, masks RFLAGS with SFMASK (so IF is already clear), and jumps
|
||||
# here with RSP still the *user* stack. We swap in the kernel GS, switch to the
|
||||
# task's kernel stack via the per-CPU block, build a CpuState frame identical to
|
||||
# the interrupt path's, and reuse interruptDispatch (vector 128) — then SYSRET.
|
||||
#
|
||||
# Hazard (acceptable while init is the only, trusted, user program): SYSRETQ #GPs
|
||||
# in ring 0 if the return RIP (RCX) is non-canonical. A hostile user could arrange
|
||||
# that; hardening (canonical check / iretq fallback) is a later security-track item.
|
||||
# the interrupt path's, and reuse interruptDispatch (vector 128) — then SYSRET,
|
||||
# unless the return RIP is non-canonical, in which case the canonical-RIP guard at
|
||||
# the exit below returns through IRETQ instead (docs/os-development/smep-smap.md).
|
||||
.global syscall_entry
|
||||
syscall_entry:
|
||||
swapgs # kernel GS base
|
||||
@@ -249,16 +247,79 @@ syscall_entry:
|
||||
pop %rax
|
||||
add $16, %rsp # drop vector + error_code -> rsp at rip
|
||||
popq %rcx # rip -> RCX (SYSRETQ restores RIP from RCX)
|
||||
# --- canonical-RIP guard ---
|
||||
# SYSRETQ with a non-canonical RCX raises #GP *in ring 0* on Intel: on the
|
||||
# kernel stack, after the swapgs below has already installed the user's GS
|
||||
# base — a fault in the trusted base, on user-influenced state, which is the
|
||||
# classic escalation primitive (CVE-2012-0217). Ring 3 gets to choose that RIP
|
||||
# without any kernel bug: the CPU saves the address of the instruction *after*
|
||||
# `syscall`, so a program executing `syscall` as the last two bytes of the last
|
||||
# canonical page returns to 0x0000_8000_0000_0000. Nothing between entry and
|
||||
# here rewrites the frame's rip (no signal or context-restore path exists that
|
||||
# could), so this is the whole attack surface — and one compare closes it.
|
||||
#
|
||||
# Canonical means bits 63:47 all equal bit 47, so sign-extending from bit 47
|
||||
# and comparing is the complete test. Bit 47 is the right pivot because danos
|
||||
# is 4-level only: paging.zig builds a PML4 and nothing anywhere sets CR4.LA57
|
||||
# (bit 12), so a linear address is 48-bit on every core, on every machine we
|
||||
# boot. A future 5-level port must pivot on bit 56 instead — and should patch
|
||||
# this shift pair at boot rather than branch on a feature flag, to keep the
|
||||
# return path of every syscall in the system free of loads.
|
||||
#
|
||||
# Cost: four register-only ALU ops and one forward branch the predictor sees
|
||||
# taken exactly never (a correct program cannot have a non-canonical return
|
||||
# address — it could not have executed there). R11 is free scratch: SYSCALL
|
||||
# already destroyed the user's copy, and the fast path overwrites it with the
|
||||
# saved RFLAGS two instructions further on.
|
||||
movq %rcx, %r11
|
||||
shlq $16, %r11
|
||||
sarq $16, %r11 # sign-extend from bit 47
|
||||
cmpq %rcx, %r11 # changed by the round trip => non-canonical
|
||||
jne .Lnon_canonical_return
|
||||
addq $8, %rsp # skip the cs slot (SYSRETQ loads CS from STAR)
|
||||
popq %r11 # rflags -> R11 (SYSRETQ restores RFLAGS from R11)
|
||||
popq %rsp # user rsp (the ss slot below is abandoned)
|
||||
swapgs # user GS base
|
||||
sysretq # -> ring 3: RIP=RCX, RFLAGS=R11, CS/SS from STAR
|
||||
|
||||
# The guard's cold path: return through IRETQ, which is safe where SYSRETQ is not.
|
||||
# IRETQ loads CS — committing the privilege change to ring 3 — before the new RIP
|
||||
# is fetched, so the #GP arrives *from ring 3*: through the IDT, onto this task's
|
||||
# kernel stack, with a user CS in the frame, where isr_common swaps GS back and the
|
||||
# kernel kills the process like any other user fault. That ordering is why IRETQ is
|
||||
# the standard fallback for this exact case; the sysret-canonical test asserts it
|
||||
# on the machine we run on (the process dies, the kernel does not).
|
||||
#
|
||||
# State handed to ring 3 is identical to what the fast path would have produced.
|
||||
# IRETQ consumes the same 5-word frame the CPU pushes for an interrupt — rip, cs,
|
||||
# rflags, rsp, ss — which is exactly the frame syscall_entry built and the fast
|
||||
# path is part-way through dismantling, so un-popping the rip slot makes it whole:
|
||||
# same user RIP, same user RSP, same RFLAGS, and CS/SS = 0x23/0x1B, the very
|
||||
# selectors SYSRETQ would have loaded from STAR. R11 is reloaded from the frame's
|
||||
# rflags slot so even the register SYSRET synthesizes matches. The swapgs sits in
|
||||
# the same place relative to the ring change as the fast path's, so the swapgs
|
||||
# discipline is untouched: kernel GS while we still touch kernel data, user GS for
|
||||
# the instant before ring 3.
|
||||
.Lnon_canonical_return:
|
||||
# Cold-path diagnostic: how many hostile return addresses this boot refused.
|
||||
# `lock` because every core shares the counter, and it costs nothing here — a
|
||||
# process that reaches this line is about to die.
|
||||
lock incq sysret_non_canonical_count(%rip)
|
||||
movq 8(%rsp), %r11 # rflags -> R11, exactly as the fast path leaves it
|
||||
subq $8, %rsp # un-pop the rip slot: rsp back at the iretq frame
|
||||
swapgs # user GS base
|
||||
iretq # -> ring 3, where the bad RIP faults harmlessly
|
||||
|
||||
.section .bss
|
||||
.balign 8
|
||||
user_saved_rsp:
|
||||
.skip 8
|
||||
# Times the canonical-RIP guard above refused a SYSRETQ this boot. Read through the
|
||||
# architecture layer (cpu.zig nonCanonicalReturnCount); zero on any machine no
|
||||
# process has attacked.
|
||||
.global sysret_non_canonical_count
|
||||
sysret_non_canonical_count:
|
||||
.skip 8
|
||||
.text
|
||||
|
||||
# --- user-mode test program --------------------------------------------------
|
||||
@@ -282,6 +343,24 @@ user_pf_start:
|
||||
1: jmp 1b
|
||||
user_pf_end:
|
||||
|
||||
# The SYSRET-guard program: two harmless system calls, the second placed so that
|
||||
# its *return address* is not canonical. The caller (tests.zig) copies these bytes
|
||||
# to the very end of the last canonical user page, so the final `syscall` occupies
|
||||
# the last two bytes of address space ring 3 can execute, and the RIP the CPU saves
|
||||
# into RCX for it is 0x0000_8000_0000_0000 — the first non-canonical address.
|
||||
# The first call proves the ordinary SYSRETQ path still works (the program only
|
||||
# reaches the second instruction pair by returning correctly from the first).
|
||||
# 39 is abi.SystemCall.current_core: no arguments, no side effects, always
|
||||
# succeeds; the test asserts the immediate below still matches that enum.
|
||||
.global user_sysret_start
|
||||
.global user_sysret_end
|
||||
user_sysret_start:
|
||||
mov $39, %eax # current_core
|
||||
syscall # canonical return address (mid-page): the fast path
|
||||
mov $39, %eax # current_core
|
||||
syscall # return address = the end of the page = non-canonical
|
||||
user_sysret_end:
|
||||
|
||||
.text
|
||||
|
||||
# Stub for a vector the CPU does NOT push an error code for: push a dummy 0.
|
||||
|
||||
Reference in New Issue
Block a user