kernel: a hostile return address cannot fault the kernel
SYSRETQ with a non-canonical RIP raises a general protection fault in ring 0 — on the kernel stack, an instruction after the swapgs that installed the user's GS base. It is one of the better-known escalation primitives, and ring 3 reaches it without any kernel bug at all: the processor saves the address of the instruction after SYSCALL, so a program whose SYSCALL is the last two bytes of the last canonical page returns to the first non-canonical address. The new test does exactly that. The exit path now sign-extends the return address from bit 47 and compares; if the value changed, it returns through IRETQ instead, which commits the privilege change before fetching the new address, so the fault arrives from ring 3 and the process dies like any other. Four register-only operations and a branch that a correct program can never take — it could not have executed at a non-canonical address in the first place. Bit 47 is the right pivot because danos builds four-level page tables and nothing sets the five-level bit; a future port must move the pivot, and the comment says so. SFMASK grows one bit while we are here. SYSCALL, unlike an interrupt gate, does not clear the nested-task flag, so the kernel had been running every system call with whatever ring 3 last chose — harmless while the only exit was SYSRETQ, and a question worth not having now that one exit is IRETQ. The kernel is never nested; ring 3 still gets its own flag back. Suite 113/113. The new case asserts the refusal counter rather than the dying process: the emulator we test on kills it either way, so only the counter distinguishes a guard that ran from one that did not.
This commit is contained in:
@@ -142,6 +142,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
|
||||
faultNull();
|
||||
} else if (eql(case, "fault-smep")) {
|
||||
faultSmep();
|
||||
} else if (eql(case, "sysret-canonical")) {
|
||||
sysretCanonicalTest();
|
||||
} else if (eql(case, "usermem")) {
|
||||
userMemTest();
|
||||
} else if (eql(case, "user-memory")) {
|
||||
@@ -1667,6 +1669,99 @@ fn faultRecoveryTest(boot_information: *const BootInformation) void {
|
||||
result();
|
||||
}
|
||||
|
||||
/// Spawn a ring-3 process whose second system call has to return to a NON-canonical
|
||||
/// address — the SYSRET hazard, staged the way a hostile program would stage it.
|
||||
/// The blob is copied to the *end* of the last canonical page of the user half, so
|
||||
/// its final `syscall` occupies the last two bytes ring 3 can execute and the return
|
||||
/// RIP the CPU hands the kernel is `user_half_end` itself, the first non-canonical
|
||||
/// address. Hand-built (address space, code page RO+X, stack RW+NX) like
|
||||
/// `spawnFaultingProcess` — a raw blob is not an ELF `spawnProcess` could load.
|
||||
/// Returns the process id, or null if any allocation failed.
|
||||
fn spawnNonCanonicalReturnProcess() ?u32 {
|
||||
const blob = process.nonCanonicalReturnBlob();
|
||||
const code_virtual = process.user_half_end - abi.page_size; // the last canonical page
|
||||
const entry = code_virtual + abi.page_size - blob.len; // ends flush with the page
|
||||
const flags = sync.enter();
|
||||
defer sync.leave(flags);
|
||||
|
||||
const address_space = architecture.createAddressSpace() orelse return null;
|
||||
const code_frame = pmm.alloc() orelse {
|
||||
architecture.destroyAddressSpace(address_space);
|
||||
return null;
|
||||
};
|
||||
// Fill through the physmap (the user mapping is read-only); int3 everywhere the
|
||||
// blob doesn't cover, so a stray entry traps instead of sliding into it.
|
||||
const code: [*]u8 = @ptrFromInt(boot_handoff.physicalToVirtual(code_frame));
|
||||
@memset(code[0..abi.page_size], 0xCC);
|
||||
@memcpy(code[abi.page_size - blob.len .. abi.page_size], blob);
|
||||
architecture.mapUserPageInto(address_space, code_virtual, code_frame, false, true); // RO + X
|
||||
|
||||
const stack_frame = pmm.alloc() orelse {
|
||||
architecture.destroyAddressSpace(address_space); // frees code_frame too — it's mapped
|
||||
return null;
|
||||
};
|
||||
architecture.mapUserPageInto(address_space, process.stack_base_virtual, stack_frame, true, false); // RW + NX
|
||||
|
||||
// Supervised by the calling test task, so exitReasonOf can read the verdict.
|
||||
return scheduler.spawnUserLocked(address_space, entry, process.stack_base_virtual + abi.page_size, 0, 4, "sysret-probe", scheduler.currentId(), null, 0) orelse {
|
||||
architecture.destroyAddressSpace(address_space);
|
||||
return null;
|
||||
};
|
||||
}
|
||||
|
||||
/// The SYSRET canonical-RIP guard (docs/os-development/smep-smap.md, "Adjacent,
|
||||
/// deliberately separate"): a process must not be able to make the kernel fault on
|
||||
/// its own way out of a system call. `sysretq` with a non-canonical RIP in RCX #GPs
|
||||
/// *in ring 0* on Intel — on the kernel stack, with the user's GS base already
|
||||
/// installed — so the exit path checks the return address first and returns through
|
||||
/// `iretq` instead, which faults in ring 3 where a bad address is just a dead
|
||||
/// process.
|
||||
///
|
||||
/// The probe reaches the hazard the way an attacker would: its `syscall` is the last
|
||||
/// two bytes of executable address space, so the return address the CPU saves is the
|
||||
/// first non-canonical one. Three things then have to be true — the guard fired, the
|
||||
/// *process* died of a protection fault (so the fault landed in ring 3, not in the
|
||||
/// kernel), and this task is still here to say so.
|
||||
///
|
||||
/// The counter is what makes this a regression test rather than a decoration.
|
||||
/// Measured with the guard's branch commented out (2026-08-01): QEMU's TCG does not
|
||||
/// model Intel's ring-0 #GP — `sysretq` simply returns to the bad address and the
|
||||
/// process dies in ring 3 anyway, so every *outcome* check still passed and only the
|
||||
/// counter noticed. On real Intel silicon the same run takes the kernel down. Assert
|
||||
/// the mechanism, not just the outcome, whenever the emulator is the softer machine.
|
||||
/// The probe's first system call is deliberately ordinary: it only reaches the second
|
||||
/// one by returning correctly from the first, so the fast path is exercised too.
|
||||
fn sysretCanonicalTest() void {
|
||||
log("DANOS-TEST-BEGIN: sysret-canonical\n", .{});
|
||||
const blob = process.nonCanonicalReturnBlob();
|
||||
check(
|
||||
"the probe's system call number still matches the ABI",
|
||||
blob.len > 1 and blob[0] == 0xB8 and blob[1] == @intFromEnum(abi.SystemCall.current_core),
|
||||
);
|
||||
|
||||
const before = architecture.nonCanonicalReturnCount();
|
||||
check("no return has been refused yet this boot", before == 0);
|
||||
process.fault_kill_count = 0;
|
||||
const probe = spawnNonCanonicalReturnProcess() orelse 0;
|
||||
check("the probe process spawned", probe != 0);
|
||||
|
||||
// Drop below the probe so it gets the core, and wait for the kill.
|
||||
scheduler.setPriority(1);
|
||||
const deadline = architecture.millis() + 5000;
|
||||
while (process.fault_kill_count < 1 and architecture.millis() < deadline) scheduler.yield();
|
||||
scheduler.setPriority(4);
|
||||
|
||||
const refused = architecture.nonCanonicalReturnCount() - before;
|
||||
check("the guard refused exactly one sysretq", refused == 1);
|
||||
check("the probe was killed (not the machine)", process.fault_kill_count == 1);
|
||||
check(
|
||||
"the probe died of a protection fault — the iretq fallback faulted it in ring 3",
|
||||
process.exitReasonOf(scheduler.currentId(), probe) == @intFromEnum(abi.ExitReason.protection_fault),
|
||||
);
|
||||
log("sysret-canonical: {d} non-canonical return(s) refused; core {d} still running\n", .{ refused, scheduler.currentCpuIndex() });
|
||||
result();
|
||||
}
|
||||
|
||||
/// Address-space refcount (docs/threading-plan.md M1): every process holds exactly one
|
||||
/// reference to its address space, released when it dies, so `destroyAddressSpace` runs
|
||||
/// exactly once per space — no leak, no double-free. Spawn and kill several ring-3
|
||||
|
||||
Reference in New Issue
Block a user