kernel: a hostile return address cannot fault the kernel
SYSRETQ with a non-canonical RIP raises a general protection fault in ring 0 — on the kernel stack, an instruction after the swapgs that installed the user's GS base. It is one of the better-known escalation primitives, and ring 3 reaches it without any kernel bug at all: the processor saves the address of the instruction after SYSCALL, so a program whose SYSCALL is the last two bytes of the last canonical page returns to the first non-canonical address. The new test does exactly that. The exit path now sign-extends the return address from bit 47 and compares; if the value changed, it returns through IRETQ instead, which commits the privilege change before fetching the new address, so the fault arrives from ring 3 and the process dies like any other. Four register-only operations and a branch that a correct program can never take — it could not have executed at a non-canonical address in the first place. Bit 47 is the right pivot because danos builds four-level page tables and nothing sets the five-level bit; a future port must move the pivot, and the comment says so. SFMASK grows one bit while we are here. SYSCALL, unlike an interrupt gate, does not clear the nested-task flag, so the kernel had been running every system call with whatever ring 3 last chose — harmless while the only exit was SYSRETQ, and a question worth not having now that one exit is IRETQ. The kernel is never nested; ring 3 still gets its own flag back. Suite 113/113. The new case asserts the refusal counter rather than the dying process: the emulator we test on kills it either way, so only the counter distinguishes a guard that ran from one that did not.
This commit is contained in:
@@ -315,6 +315,20 @@ pub fn userExit() noreturn {
|
||||
user_exit_to_kernel();
|
||||
}
|
||||
|
||||
/// The counter the syscall exit stub bumps when it refuses to return the fast way
|
||||
/// (isr.s, `sysret_non_canonical_count`).
|
||||
const non_canonical_returns = @extern(*u64, .{ .name = "sysret_non_canonical_count" });
|
||||
|
||||
/// How many times this boot's system_call returns took the slow, safe exit because
|
||||
/// the return address was not a canonical address (the SYSRETQ canonical-RIP guard
|
||||
/// in isr.s; an architecture without the hazard reports 0 forever). A correct
|
||||
/// program cannot produce one — it could not have executed at a non-canonical
|
||||
/// address in the first place — so a nonzero count is exactly the number of times
|
||||
/// a process tried to make the kernel fault on its way out.
|
||||
pub fn nonCanonicalReturnCount() u64 {
|
||||
return @atomicLoad(u64, non_canonical_returns, .monotonic);
|
||||
}
|
||||
|
||||
/// Register the handler for the user system_call gate (int 0x80, vector 128). The
|
||||
/// handler may write the trap frame (see `setSystemCallResult`).
|
||||
pub fn setSystemCallHandler(handler: *const fn (*CpuState) void) void {
|
||||
|
||||
@@ -187,11 +187,9 @@ user_exit_to_kernel:
|
||||
# CS/SS from STAR, masks RFLAGS with SFMASK (so IF is already clear), and jumps
|
||||
# here with RSP still the *user* stack. We swap in the kernel GS, switch to the
|
||||
# task's kernel stack via the per-CPU block, build a CpuState frame identical to
|
||||
# the interrupt path's, and reuse interruptDispatch (vector 128) — then SYSRET.
|
||||
#
|
||||
# Hazard (acceptable while init is the only, trusted, user program): SYSRETQ #GPs
|
||||
# in ring 0 if the return RIP (RCX) is non-canonical. A hostile user could arrange
|
||||
# that; hardening (canonical check / iretq fallback) is a later security-track item.
|
||||
# the interrupt path's, and reuse interruptDispatch (vector 128) — then SYSRET,
|
||||
# unless the return RIP is non-canonical, in which case the canonical-RIP guard at
|
||||
# the exit below returns through IRETQ instead (docs/os-development/smep-smap.md).
|
||||
.global syscall_entry
|
||||
syscall_entry:
|
||||
swapgs # kernel GS base
|
||||
@@ -249,16 +247,79 @@ syscall_entry:
|
||||
pop %rax
|
||||
add $16, %rsp # drop vector + error_code -> rsp at rip
|
||||
popq %rcx # rip -> RCX (SYSRETQ restores RIP from RCX)
|
||||
# --- canonical-RIP guard ---
|
||||
# SYSRETQ with a non-canonical RCX raises #GP *in ring 0* on Intel: on the
|
||||
# kernel stack, after the swapgs below has already installed the user's GS
|
||||
# base — a fault in the trusted base, on user-influenced state, which is the
|
||||
# classic escalation primitive (CVE-2012-0217). Ring 3 gets to choose that RIP
|
||||
# without any kernel bug: the CPU saves the address of the instruction *after*
|
||||
# `syscall`, so a program executing `syscall` as the last two bytes of the last
|
||||
# canonical page returns to 0x0000_8000_0000_0000. Nothing between entry and
|
||||
# here rewrites the frame's rip (no signal or context-restore path exists that
|
||||
# could), so this is the whole attack surface — and one compare closes it.
|
||||
#
|
||||
# Canonical means bits 63:47 all equal bit 47, so sign-extending from bit 47
|
||||
# and comparing is the complete test. Bit 47 is the right pivot because danos
|
||||
# is 4-level only: paging.zig builds a PML4 and nothing anywhere sets CR4.LA57
|
||||
# (bit 12), so a linear address is 48-bit on every core, on every machine we
|
||||
# boot. A future 5-level port must pivot on bit 56 instead — and should patch
|
||||
# this shift pair at boot rather than branch on a feature flag, to keep the
|
||||
# return path of every syscall in the system free of loads.
|
||||
#
|
||||
# Cost: four register-only ALU ops and one forward branch the predictor sees
|
||||
# taken exactly never (a correct program cannot have a non-canonical return
|
||||
# address — it could not have executed there). R11 is free scratch: SYSCALL
|
||||
# already destroyed the user's copy, and the fast path overwrites it with the
|
||||
# saved RFLAGS two instructions further on.
|
||||
movq %rcx, %r11
|
||||
shlq $16, %r11
|
||||
sarq $16, %r11 # sign-extend from bit 47
|
||||
cmpq %rcx, %r11 # changed by the round trip => non-canonical
|
||||
jne .Lnon_canonical_return
|
||||
addq $8, %rsp # skip the cs slot (SYSRETQ loads CS from STAR)
|
||||
popq %r11 # rflags -> R11 (SYSRETQ restores RFLAGS from R11)
|
||||
popq %rsp # user rsp (the ss slot below is abandoned)
|
||||
swapgs # user GS base
|
||||
sysretq # -> ring 3: RIP=RCX, RFLAGS=R11, CS/SS from STAR
|
||||
|
||||
# The guard's cold path: return through IRETQ, which is safe where SYSRETQ is not.
|
||||
# IRETQ loads CS — committing the privilege change to ring 3 — before the new RIP
|
||||
# is fetched, so the #GP arrives *from ring 3*: through the IDT, onto this task's
|
||||
# kernel stack, with a user CS in the frame, where isr_common swaps GS back and the
|
||||
# kernel kills the process like any other user fault. That ordering is why IRETQ is
|
||||
# the standard fallback for this exact case; the sysret-canonical test asserts it
|
||||
# on the machine we run on (the process dies, the kernel does not).
|
||||
#
|
||||
# State handed to ring 3 is identical to what the fast path would have produced.
|
||||
# IRETQ consumes the same 5-word frame the CPU pushes for an interrupt — rip, cs,
|
||||
# rflags, rsp, ss — which is exactly the frame syscall_entry built and the fast
|
||||
# path is part-way through dismantling, so un-popping the rip slot makes it whole:
|
||||
# same user RIP, same user RSP, same RFLAGS, and CS/SS = 0x23/0x1B, the very
|
||||
# selectors SYSRETQ would have loaded from STAR. R11 is reloaded from the frame's
|
||||
# rflags slot so even the register SYSRET synthesizes matches. The swapgs sits in
|
||||
# the same place relative to the ring change as the fast path's, so the swapgs
|
||||
# discipline is untouched: kernel GS while we still touch kernel data, user GS for
|
||||
# the instant before ring 3.
|
||||
.Lnon_canonical_return:
|
||||
# Cold-path diagnostic: how many hostile return addresses this boot refused.
|
||||
# `lock` because every core shares the counter, and it costs nothing here — a
|
||||
# process that reaches this line is about to die.
|
||||
lock incq sysret_non_canonical_count(%rip)
|
||||
movq 8(%rsp), %r11 # rflags -> R11, exactly as the fast path leaves it
|
||||
subq $8, %rsp # un-pop the rip slot: rsp back at the iretq frame
|
||||
swapgs # user GS base
|
||||
iretq # -> ring 3, where the bad RIP faults harmlessly
|
||||
|
||||
.section .bss
|
||||
.balign 8
|
||||
user_saved_rsp:
|
||||
.skip 8
|
||||
# Times the canonical-RIP guard above refused a SYSRETQ this boot. Read through the
|
||||
# architecture layer (cpu.zig nonCanonicalReturnCount); zero on any machine no
|
||||
# process has attacked.
|
||||
.global sysret_non_canonical_count
|
||||
sysret_non_canonical_count:
|
||||
.skip 8
|
||||
.text
|
||||
|
||||
# --- user-mode test program --------------------------------------------------
|
||||
@@ -282,6 +343,24 @@ user_pf_start:
|
||||
1: jmp 1b
|
||||
user_pf_end:
|
||||
|
||||
# The SYSRET-guard program: two harmless system calls, the second placed so that
|
||||
# its *return address* is not canonical. The caller (tests.zig) copies these bytes
|
||||
# to the very end of the last canonical user page, so the final `syscall` occupies
|
||||
# the last two bytes of address space ring 3 can execute, and the RIP the CPU saves
|
||||
# into RCX for it is 0x0000_8000_0000_0000 — the first non-canonical address.
|
||||
# The first call proves the ordinary SYSRETQ path still works (the program only
|
||||
# reaches the second instruction pair by returning correctly from the first).
|
||||
# 39 is abi.SystemCall.current_core: no arguments, no side effects, always
|
||||
# succeeds; the test asserts the immediate below still matches that enum.
|
||||
.global user_sysret_start
|
||||
.global user_sysret_end
|
||||
user_sysret_start:
|
||||
mov $39, %eax # current_core
|
||||
syscall # canonical return address (mid-page): the fast path
|
||||
mov $39, %eax # current_core
|
||||
syscall # return address = the end of the page = non-canonical
|
||||
user_sysret_end:
|
||||
|
||||
.text
|
||||
|
||||
# Stub for a vector the CPU does NOT push an error code for: push a dummy 0.
|
||||
|
||||
@@ -80,7 +80,16 @@ pub fn initSystemCall() void {
|
||||
io.wrmsr(ia32_star, (@as(u64, 0x08) << 32) | (@as(u64, 0x10) << 48));
|
||||
const entry = @extern(*const anyopaque, .{ .name = "syscall_entry" });
|
||||
io.wrmsr(ia32_lstar, @intFromPtr(entry));
|
||||
io.wrmsr(ia32_sfmask, 0x4_0700); // clear IF, TF, DF, AC on entry
|
||||
// Clear IF, TF, DF, AC and NT on entry. The first four are the usual
|
||||
// hygiene; NT is here because SYSCALL, unlike an interrupt gate, does not
|
||||
// clear it for us, so without this the kernel runs every system call with
|
||||
// whatever nested-task bit ring 3 last chose — and the canonical-RIP guard's
|
||||
// cold path (isr.s) leaves through IRETQ, whose behaviour with NT set is a
|
||||
// corner of the manuals not worth depending on either way. Masking it costs
|
||||
// one bit and removes the question: the kernel is never nested, and ring 3
|
||||
// still gets its own NT back, from R11 on the fast path and from the frame
|
||||
// on the cold one.
|
||||
io.wrmsr(ia32_sfmask, 0x4_4700);
|
||||
}
|
||||
|
||||
// --- supervisor-mode hardening (CR4) ---------------------------------------
|
||||
|
||||
@@ -124,15 +124,26 @@ pub const maximum_mount_rewrite = 32;
|
||||
const auxiliary_vector_null: u64 = 0; // AT_NULL — end of the vector
|
||||
const auxiliary_vector_page_size: u64 = 6; // AT_PAGESZ
|
||||
|
||||
// The hand-assembled user program blob (isr.s, .rodata) — the isolation probe.
|
||||
// The hand-assembled user program blobs (isr.s, .rodata) — the isolation probes.
|
||||
const pf_start = @extern([*]const u8, .{ .name = "user_pf_start" });
|
||||
const pf_end = @extern([*]const u8, .{ .name = "user_pf_end" });
|
||||
const sysret_start = @extern([*]const u8, .{ .name = "user_sysret_start" });
|
||||
const sysret_end = @extern([*]const u8, .{ .name = "user_sysret_end" });
|
||||
|
||||
/// The isolation-proof program: reads a kernel-only page, must #PF.
|
||||
pub fn pfBlob() []const u8 {
|
||||
return pf_start[0 .. @intFromPtr(pf_end) - @intFromPtr(pf_start)];
|
||||
}
|
||||
|
||||
/// The SYSRET-guard program: two system calls, the second of which must be copied
|
||||
/// so that it ends at the last executable byte of the user half — its return
|
||||
/// address is then the first non-canonical address. Ends with `syscall`, so the
|
||||
/// caller places it at `page_size - len` inside the page at `user_half_end -
|
||||
/// page_size`; anywhere else and it proves nothing.
|
||||
pub fn nonCanonicalReturnBlob() []const u8 {
|
||||
return sysret_start[0 .. @intFromPtr(sysret_end) - @intFromPtr(sysret_start)];
|
||||
}
|
||||
|
||||
/// What debug_write syscalls produced (accumulated), and the exit system_call's code.
|
||||
pub var write_buffer: [256]u8 = undefined;
|
||||
pub var write_len: usize = 0;
|
||||
|
||||
@@ -142,6 +142,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
|
||||
faultNull();
|
||||
} else if (eql(case, "fault-smep")) {
|
||||
faultSmep();
|
||||
} else if (eql(case, "sysret-canonical")) {
|
||||
sysretCanonicalTest();
|
||||
} else if (eql(case, "usermem")) {
|
||||
userMemTest();
|
||||
} else if (eql(case, "user-memory")) {
|
||||
@@ -1667,6 +1669,99 @@ fn faultRecoveryTest(boot_information: *const BootInformation) void {
|
||||
result();
|
||||
}
|
||||
|
||||
/// Spawn a ring-3 process whose second system call has to return to a NON-canonical
|
||||
/// address — the SYSRET hazard, staged the way a hostile program would stage it.
|
||||
/// The blob is copied to the *end* of the last canonical page of the user half, so
|
||||
/// its final `syscall` occupies the last two bytes ring 3 can execute and the return
|
||||
/// RIP the CPU hands the kernel is `user_half_end` itself, the first non-canonical
|
||||
/// address. Hand-built (address space, code page RO+X, stack RW+NX) like
|
||||
/// `spawnFaultingProcess` — a raw blob is not an ELF `spawnProcess` could load.
|
||||
/// Returns the process id, or null if any allocation failed.
|
||||
fn spawnNonCanonicalReturnProcess() ?u32 {
|
||||
const blob = process.nonCanonicalReturnBlob();
|
||||
const code_virtual = process.user_half_end - abi.page_size; // the last canonical page
|
||||
const entry = code_virtual + abi.page_size - blob.len; // ends flush with the page
|
||||
const flags = sync.enter();
|
||||
defer sync.leave(flags);
|
||||
|
||||
const address_space = architecture.createAddressSpace() orelse return null;
|
||||
const code_frame = pmm.alloc() orelse {
|
||||
architecture.destroyAddressSpace(address_space);
|
||||
return null;
|
||||
};
|
||||
// Fill through the physmap (the user mapping is read-only); int3 everywhere the
|
||||
// blob doesn't cover, so a stray entry traps instead of sliding into it.
|
||||
const code: [*]u8 = @ptrFromInt(boot_handoff.physicalToVirtual(code_frame));
|
||||
@memset(code[0..abi.page_size], 0xCC);
|
||||
@memcpy(code[abi.page_size - blob.len .. abi.page_size], blob);
|
||||
architecture.mapUserPageInto(address_space, code_virtual, code_frame, false, true); // RO + X
|
||||
|
||||
const stack_frame = pmm.alloc() orelse {
|
||||
architecture.destroyAddressSpace(address_space); // frees code_frame too — it's mapped
|
||||
return null;
|
||||
};
|
||||
architecture.mapUserPageInto(address_space, process.stack_base_virtual, stack_frame, true, false); // RW + NX
|
||||
|
||||
// Supervised by the calling test task, so exitReasonOf can read the verdict.
|
||||
return scheduler.spawnUserLocked(address_space, entry, process.stack_base_virtual + abi.page_size, 0, 4, "sysret-probe", scheduler.currentId(), null, 0) orelse {
|
||||
architecture.destroyAddressSpace(address_space);
|
||||
return null;
|
||||
};
|
||||
}
|
||||
|
||||
/// The SYSRET canonical-RIP guard (docs/os-development/smep-smap.md, "Adjacent,
|
||||
/// deliberately separate"): a process must not be able to make the kernel fault on
|
||||
/// its own way out of a system call. `sysretq` with a non-canonical RIP in RCX #GPs
|
||||
/// *in ring 0* on Intel — on the kernel stack, with the user's GS base already
|
||||
/// installed — so the exit path checks the return address first and returns through
|
||||
/// `iretq` instead, which faults in ring 3 where a bad address is just a dead
|
||||
/// process.
|
||||
///
|
||||
/// The probe reaches the hazard the way an attacker would: its `syscall` is the last
|
||||
/// two bytes of executable address space, so the return address the CPU saves is the
|
||||
/// first non-canonical one. Three things then have to be true — the guard fired, the
|
||||
/// *process* died of a protection fault (so the fault landed in ring 3, not in the
|
||||
/// kernel), and this task is still here to say so.
|
||||
///
|
||||
/// The counter is what makes this a regression test rather than a decoration.
|
||||
/// Measured with the guard's branch commented out (2026-08-01): QEMU's TCG does not
|
||||
/// model Intel's ring-0 #GP — `sysretq` simply returns to the bad address and the
|
||||
/// process dies in ring 3 anyway, so every *outcome* check still passed and only the
|
||||
/// counter noticed. On real Intel silicon the same run takes the kernel down. Assert
|
||||
/// the mechanism, not just the outcome, whenever the emulator is the softer machine.
|
||||
/// The probe's first system call is deliberately ordinary: it only reaches the second
|
||||
/// one by returning correctly from the first, so the fast path is exercised too.
|
||||
fn sysretCanonicalTest() void {
|
||||
log("DANOS-TEST-BEGIN: sysret-canonical\n", .{});
|
||||
const blob = process.nonCanonicalReturnBlob();
|
||||
check(
|
||||
"the probe's system call number still matches the ABI",
|
||||
blob.len > 1 and blob[0] == 0xB8 and blob[1] == @intFromEnum(abi.SystemCall.current_core),
|
||||
);
|
||||
|
||||
const before = architecture.nonCanonicalReturnCount();
|
||||
check("no return has been refused yet this boot", before == 0);
|
||||
process.fault_kill_count = 0;
|
||||
const probe = spawnNonCanonicalReturnProcess() orelse 0;
|
||||
check("the probe process spawned", probe != 0);
|
||||
|
||||
// Drop below the probe so it gets the core, and wait for the kill.
|
||||
scheduler.setPriority(1);
|
||||
const deadline = architecture.millis() + 5000;
|
||||
while (process.fault_kill_count < 1 and architecture.millis() < deadline) scheduler.yield();
|
||||
scheduler.setPriority(4);
|
||||
|
||||
const refused = architecture.nonCanonicalReturnCount() - before;
|
||||
check("the guard refused exactly one sysretq", refused == 1);
|
||||
check("the probe was killed (not the machine)", process.fault_kill_count == 1);
|
||||
check(
|
||||
"the probe died of a protection fault — the iretq fallback faulted it in ring 3",
|
||||
process.exitReasonOf(scheduler.currentId(), probe) == @intFromEnum(abi.ExitReason.protection_fault),
|
||||
);
|
||||
log("sysret-canonical: {d} non-canonical return(s) refused; core {d} still running\n", .{ refused, scheduler.currentCpuIndex() });
|
||||
result();
|
||||
}
|
||||
|
||||
/// Address-space refcount (docs/threading-plan.md M1): every process holds exactly one
|
||||
/// reference to its address space, released when it dies, so `destroyAddressSpace` runs
|
||||
/// exactly once per space — no leak, no double-free. Spawn and kill several ring-3
|
||||
|
||||
Reference in New Issue
Block a user