kernel: a hostile return address cannot fault the kernel
SYSRETQ with a non-canonical RIP raises a general protection fault in ring 0 — on the kernel stack, an instruction after the swapgs that installed the user's GS base. It is one of the better-known escalation primitives, and ring 3 reaches it without any kernel bug at all: the processor saves the address of the instruction after SYSCALL, so a program whose SYSCALL is the last two bytes of the last canonical page returns to the first non-canonical address. The new test does exactly that. The exit path now sign-extends the return address from bit 47 and compares; if the value changed, it returns through IRETQ instead, which commits the privilege change before fetching the new address, so the fault arrives from ring 3 and the process dies like any other. Four register-only operations and a branch that a correct program can never take — it could not have executed at a non-canonical address in the first place. Bit 47 is the right pivot because danos builds four-level page tables and nothing sets the five-level bit; a future port must move the pivot, and the comment says so. SFMASK grows one bit while we are here. SYSCALL, unlike an interrupt gate, does not clear the nested-task flag, so the kernel had been running every system call with whatever ring 3 last chose — harmless while the only exit was SYSRETQ, and a question worth not having now that one exit is IRETQ. The kernel is never nested; ring 3 still gets its own flag back. Suite 113/113. The new case asserts the refusal counter rather than the dying process: the emulator we test on kills it either way, so only the counter distinguishes a guard that ran from one that did not.
This commit is contained in:
@@ -1,6 +1,7 @@
|
||||
# SMEP and SMAP — supervisor-mode hardening
|
||||
|
||||
*Design, 2026-07-31. H2 (SMEP) has landed; HS and H3 are the remaining work. Companion to
|
||||
*Design, 2026-07-31. H2 (SMEP) and HS (the SYSRET guard) have landed; H3 is the
|
||||
remaining work. Companion to
|
||||
[protocol-namespace.md](protocol-namespace.md) on the security track — this is
|
||||
the hardware half; that is the namespace half.*
|
||||
|
||||
@@ -136,6 +137,11 @@ Broadwell+ for SMAP; AMD Zen+ for both). The test images should run with
|
||||
hardening bucket, independent fix (validate RCX before `sysretq`, fall
|
||||
back to `iretq`), should ride the same branch as H2/H3 but is not
|
||||
SMEP/SMAP.
|
||||
**Status — landed 2026-08-01 (HS).** The syscall exit sign-extends the
|
||||
return RIP from bit 47 (danos is 4-level only; nothing sets CR4.LA57) and
|
||||
falls back to `iretq` when that changes it, counting each refusal for the
|
||||
`sysret-canonical` case. Ring 3 could reach it: `syscall` as the last two
|
||||
bytes of the last canonical page returns to `user_half_end`.
|
||||
- **KPTI / Meltdown-class leaks are out of scope.** SMEP/SMAP police
|
||||
architectural accesses, not speculative ones. danos runs one kernel
|
||||
mapping in every address space and accepts that on affected hardware;
|
||||
|
||||
Reference in New Issue
Block a user