kernel: a hostile return address cannot fault the kernel

SYSRETQ with a non-canonical RIP raises a general protection fault in ring
0 — on the kernel stack, an instruction after the swapgs that installed the
user's GS base. It is one of the better-known escalation primitives, and
ring 3 reaches it without any kernel bug at all: the processor saves the
address of the instruction after SYSCALL, so a program whose SYSCALL is the
last two bytes of the last canonical page returns to the first
non-canonical address. The new test does exactly that.

The exit path now sign-extends the return address from bit 47 and compares;
if the value changed, it returns through IRETQ instead, which commits the
privilege change before fetching the new address, so the fault arrives from
ring 3 and the process dies like any other. Four register-only operations
and a branch that a correct program can never take — it could not have
executed at a non-canonical address in the first place. Bit 47 is the right
pivot because danos builds four-level page tables and nothing sets the
five-level bit; a future port must move the pivot, and the comment says so.

SFMASK grows one bit while we are here. SYSCALL, unlike an interrupt gate,
does not clear the nested-task flag, so the kernel had been running every
system call with whatever ring 3 last chose — harmless while the only exit
was SYSRETQ, and a question worth not having now that one exit is IRETQ.
The kernel is never nested; ring 3 still gets its own flag back.

Suite 113/113. The new case asserts the refusal counter rather than the
dying process: the emulator we test on kills it either way, so only the
counter distinguishes a guard that ran from one that did not.
This commit is contained in:
Daniel Samson
2026-08-01 11:22:07 +01:00
parent 36a7cc5fe9
commit cb30faf15f
8 changed files with 257 additions and 13 deletions
+14
View File
@@ -315,6 +315,20 @@ pub fn userExit() noreturn {
user_exit_to_kernel();
}
/// The counter the syscall exit stub bumps when it refuses to return the fast way
/// (isr.s, `sysret_non_canonical_count`).
const non_canonical_returns = @extern(*u64, .{ .name = "sysret_non_canonical_count" });
/// How many times this boot's system_call returns took the slow, safe exit because
/// the return address was not a canonical address (the SYSRETQ canonical-RIP guard
/// in isr.s; an architecture without the hazard reports 0 forever). A correct
/// program cannot produce one — it could not have executed at a non-canonical
/// address in the first place — so a nonzero count is exactly the number of times
/// a process tried to make the kernel fault on its way out.
pub fn nonCanonicalReturnCount() u64 {
return @atomicLoad(u64, non_canonical_returns, .monotonic);
}
/// Register the handler for the user system_call gate (int 0x80, vector 128). The
/// handler may write the trap frame (see `setSystemCallResult`).
pub fn setSystemCallHandler(handler: *const fn (*CpuState) void) void {
+84 -5
View File
@@ -187,11 +187,9 @@ user_exit_to_kernel:
# CS/SS from STAR, masks RFLAGS with SFMASK (so IF is already clear), and jumps
# here with RSP still the *user* stack. We swap in the kernel GS, switch to the
# task's kernel stack via the per-CPU block, build a CpuState frame identical to
# the interrupt path's, and reuse interruptDispatch (vector 128) — then SYSRET.
#
# Hazard (acceptable while init is the only, trusted, user program): SYSRETQ #GPs
# in ring 0 if the return RIP (RCX) is non-canonical. A hostile user could arrange
# that; hardening (canonical check / iretq fallback) is a later security-track item.
# the interrupt path's, and reuse interruptDispatch (vector 128) — then SYSRET,
# unless the return RIP is non-canonical, in which case the canonical-RIP guard at
# the exit below returns through IRETQ instead (docs/os-development/smep-smap.md).
.global syscall_entry
syscall_entry:
swapgs # kernel GS base
@@ -249,16 +247,79 @@ syscall_entry:
pop %rax
add $16, %rsp # drop vector + error_code -> rsp at rip
popq %rcx # rip -> RCX (SYSRETQ restores RIP from RCX)
# --- canonical-RIP guard ---
# SYSRETQ with a non-canonical RCX raises #GP *in ring 0* on Intel: on the
# kernel stack, after the swapgs below has already installed the user's GS
# base — a fault in the trusted base, on user-influenced state, which is the
# classic escalation primitive (CVE-2012-0217). Ring 3 gets to choose that RIP
# without any kernel bug: the CPU saves the address of the instruction *after*
# `syscall`, so a program executing `syscall` as the last two bytes of the last
# canonical page returns to 0x0000_8000_0000_0000. Nothing between entry and
# here rewrites the frame's rip (no signal or context-restore path exists that
# could), so this is the whole attack surface — and one compare closes it.
#
# Canonical means bits 63:47 all equal bit 47, so sign-extending from bit 47
# and comparing is the complete test. Bit 47 is the right pivot because danos
# is 4-level only: paging.zig builds a PML4 and nothing anywhere sets CR4.LA57
# (bit 12), so a linear address is 48-bit on every core, on every machine we
# boot. A future 5-level port must pivot on bit 56 instead — and should patch
# this shift pair at boot rather than branch on a feature flag, to keep the
# return path of every syscall in the system free of loads.
#
# Cost: four register-only ALU ops and one forward branch the predictor sees
# taken exactly never (a correct program cannot have a non-canonical return
# address — it could not have executed there). R11 is free scratch: SYSCALL
# already destroyed the user's copy, and the fast path overwrites it with the
# saved RFLAGS two instructions further on.
movq %rcx, %r11
shlq $16, %r11
sarq $16, %r11 # sign-extend from bit 47
cmpq %rcx, %r11 # changed by the round trip => non-canonical
jne .Lnon_canonical_return
addq $8, %rsp # skip the cs slot (SYSRETQ loads CS from STAR)
popq %r11 # rflags -> R11 (SYSRETQ restores RFLAGS from R11)
popq %rsp # user rsp (the ss slot below is abandoned)
swapgs # user GS base
sysretq # -> ring 3: RIP=RCX, RFLAGS=R11, CS/SS from STAR
# The guard's cold path: return through IRETQ, which is safe where SYSRETQ is not.
# IRETQ loads CS — committing the privilege change to ring 3 — before the new RIP
# is fetched, so the #GP arrives *from ring 3*: through the IDT, onto this task's
# kernel stack, with a user CS in the frame, where isr_common swaps GS back and the
# kernel kills the process like any other user fault. That ordering is why IRETQ is
# the standard fallback for this exact case; the sysret-canonical test asserts it
# on the machine we run on (the process dies, the kernel does not).
#
# State handed to ring 3 is identical to what the fast path would have produced.
# IRETQ consumes the same 5-word frame the CPU pushes for an interrupt — rip, cs,
# rflags, rsp, ss — which is exactly the frame syscall_entry built and the fast
# path is part-way through dismantling, so un-popping the rip slot makes it whole:
# same user RIP, same user RSP, same RFLAGS, and CS/SS = 0x23/0x1B, the very
# selectors SYSRETQ would have loaded from STAR. R11 is reloaded from the frame's
# rflags slot so even the register SYSRET synthesizes matches. The swapgs sits in
# the same place relative to the ring change as the fast path's, so the swapgs
# discipline is untouched: kernel GS while we still touch kernel data, user GS for
# the instant before ring 3.
.Lnon_canonical_return:
# Cold-path diagnostic: how many hostile return addresses this boot refused.
# `lock` because every core shares the counter, and it costs nothing here — a
# process that reaches this line is about to die.
lock incq sysret_non_canonical_count(%rip)
movq 8(%rsp), %r11 # rflags -> R11, exactly as the fast path leaves it
subq $8, %rsp # un-pop the rip slot: rsp back at the iretq frame
swapgs # user GS base
iretq # -> ring 3, where the bad RIP faults harmlessly
.section .bss
.balign 8
user_saved_rsp:
.skip 8
# Times the canonical-RIP guard above refused a SYSRETQ this boot. Read through the
# architecture layer (cpu.zig nonCanonicalReturnCount); zero on any machine no
# process has attacked.
.global sysret_non_canonical_count
sysret_non_canonical_count:
.skip 8
.text
# --- user-mode test program --------------------------------------------------
@@ -282,6 +343,24 @@ user_pf_start:
1: jmp 1b
user_pf_end:
# The SYSRET-guard program: two harmless system calls, the second placed so that
# its *return address* is not canonical. The caller (tests.zig) copies these bytes
# to the very end of the last canonical user page, so the final `syscall` occupies
# the last two bytes of address space ring 3 can execute, and the RIP the CPU saves
# into RCX for it is 0x0000_8000_0000_0000 — the first non-canonical address.
# The first call proves the ordinary SYSRETQ path still works (the program only
# reaches the second instruction pair by returning correctly from the first).
# 39 is abi.SystemCall.current_core: no arguments, no side effects, always
# succeeds; the test asserts the immediate below still matches that enum.
.global user_sysret_start
.global user_sysret_end
user_sysret_start:
mov $39, %eax # current_core
syscall # canonical return address (mid-page): the fast path
mov $39, %eax # current_core
syscall # return address = the end of the page = non-canonical
user_sysret_end:
.text
# Stub for a vector the CPU does NOT push an error code for: push a dummy 0.
+10 -1
View File
@@ -80,7 +80,16 @@ pub fn initSystemCall() void {
io.wrmsr(ia32_star, (@as(u64, 0x08) << 32) | (@as(u64, 0x10) << 48));
const entry = @extern(*const anyopaque, .{ .name = "syscall_entry" });
io.wrmsr(ia32_lstar, @intFromPtr(entry));
io.wrmsr(ia32_sfmask, 0x4_0700); // clear IF, TF, DF, AC on entry
// Clear IF, TF, DF, AC and NT on entry. The first four are the usual
// hygiene; NT is here because SYSCALL, unlike an interrupt gate, does not
// clear it for us, so without this the kernel runs every system call with
// whatever nested-task bit ring 3 last chose — and the canonical-RIP guard's
// cold path (isr.s) leaves through IRETQ, whose behaviour with NT set is a
// corner of the manuals not worth depending on either way. Masking it costs
// one bit and removes the question: the kernel is never nested, and ring 3
// still gets its own NT back, from R11 on the fast path and from the frame
// on the cold one.
io.wrmsr(ia32_sfmask, 0x4_4700);
}
// --- supervisor-mode hardening (CR4) ---------------------------------------