kernel: a hostile return address cannot fault the kernel

SYSRETQ with a non-canonical RIP raises a general protection fault in ring
0 — on the kernel stack, an instruction after the swapgs that installed the
user's GS base. It is one of the better-known escalation primitives, and
ring 3 reaches it without any kernel bug at all: the processor saves the
address of the instruction after SYSCALL, so a program whose SYSCALL is the
last two bytes of the last canonical page returns to the first
non-canonical address. The new test does exactly that.

The exit path now sign-extends the return address from bit 47 and compares;
if the value changed, it returns through IRETQ instead, which commits the
privilege change before fetching the new address, so the fault arrives from
ring 3 and the process dies like any other. Four register-only operations
and a branch that a correct program can never take — it could not have
executed at a non-canonical address in the first place. Bit 47 is the right
pivot because danos builds four-level page tables and nothing sets the
five-level bit; a future port must move the pivot, and the comment says so.

SFMASK grows one bit while we are here. SYSCALL, unlike an interrupt gate,
does not clear the nested-task flag, so the kernel had been running every
system call with whatever ring 3 last chose — harmless while the only exit
was SYSRETQ, and a question worth not having now that one exit is IRETQ.
The kernel is never nested; ring 3 still gets its own flag back.

Suite 113/113. The new case asserts the refusal counter rather than the
dying process: the emulator we test on kills it either way, so only the
counter distinguishes a guard that ran from one that did not.
This commit is contained in:
Daniel Samson
2026-08-01 11:22:07 +01:00
parent 36a7cc5fe9
commit cb30faf15f
8 changed files with 257 additions and 13 deletions
+95
View File
@@ -142,6 +142,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
faultNull();
} else if (eql(case, "fault-smep")) {
faultSmep();
} else if (eql(case, "sysret-canonical")) {
sysretCanonicalTest();
} else if (eql(case, "usermem")) {
userMemTest();
} else if (eql(case, "user-memory")) {
@@ -1667,6 +1669,99 @@ fn faultRecoveryTest(boot_information: *const BootInformation) void {
result();
}
/// Spawn a ring-3 process whose second system call has to return to a NON-canonical
/// address — the SYSRET hazard, staged the way a hostile program would stage it.
/// The blob is copied to the *end* of the last canonical page of the user half, so
/// its final `syscall` occupies the last two bytes ring 3 can execute and the return
/// RIP the CPU hands the kernel is `user_half_end` itself, the first non-canonical
/// address. Hand-built (address space, code page RO+X, stack RW+NX) like
/// `spawnFaultingProcess` — a raw blob is not an ELF `spawnProcess` could load.
/// Returns the process id, or null if any allocation failed.
fn spawnNonCanonicalReturnProcess() ?u32 {
const blob = process.nonCanonicalReturnBlob();
const code_virtual = process.user_half_end - abi.page_size; // the last canonical page
const entry = code_virtual + abi.page_size - blob.len; // ends flush with the page
const flags = sync.enter();
defer sync.leave(flags);
const address_space = architecture.createAddressSpace() orelse return null;
const code_frame = pmm.alloc() orelse {
architecture.destroyAddressSpace(address_space);
return null;
};
// Fill through the physmap (the user mapping is read-only); int3 everywhere the
// blob doesn't cover, so a stray entry traps instead of sliding into it.
const code: [*]u8 = @ptrFromInt(boot_handoff.physicalToVirtual(code_frame));
@memset(code[0..abi.page_size], 0xCC);
@memcpy(code[abi.page_size - blob.len .. abi.page_size], blob);
architecture.mapUserPageInto(address_space, code_virtual, code_frame, false, true); // RO + X
const stack_frame = pmm.alloc() orelse {
architecture.destroyAddressSpace(address_space); // frees code_frame too — it's mapped
return null;
};
architecture.mapUserPageInto(address_space, process.stack_base_virtual, stack_frame, true, false); // RW + NX
// Supervised by the calling test task, so exitReasonOf can read the verdict.
return scheduler.spawnUserLocked(address_space, entry, process.stack_base_virtual + abi.page_size, 0, 4, "sysret-probe", scheduler.currentId(), null, 0) orelse {
architecture.destroyAddressSpace(address_space);
return null;
};
}
/// The SYSRET canonical-RIP guard (docs/os-development/smep-smap.md, "Adjacent,
/// deliberately separate"): a process must not be able to make the kernel fault on
/// its own way out of a system call. `sysretq` with a non-canonical RIP in RCX #GPs
/// *in ring 0* on Intel — on the kernel stack, with the user's GS base already
/// installed — so the exit path checks the return address first and returns through
/// `iretq` instead, which faults in ring 3 where a bad address is just a dead
/// process.
///
/// The probe reaches the hazard the way an attacker would: its `syscall` is the last
/// two bytes of executable address space, so the return address the CPU saves is the
/// first non-canonical one. Three things then have to be true — the guard fired, the
/// *process* died of a protection fault (so the fault landed in ring 3, not in the
/// kernel), and this task is still here to say so.
///
/// The counter is what makes this a regression test rather than a decoration.
/// Measured with the guard's branch commented out (2026-08-01): QEMU's TCG does not
/// model Intel's ring-0 #GP — `sysretq` simply returns to the bad address and the
/// process dies in ring 3 anyway, so every *outcome* check still passed and only the
/// counter noticed. On real Intel silicon the same run takes the kernel down. Assert
/// the mechanism, not just the outcome, whenever the emulator is the softer machine.
/// The probe's first system call is deliberately ordinary: it only reaches the second
/// one by returning correctly from the first, so the fast path is exercised too.
fn sysretCanonicalTest() void {
log("DANOS-TEST-BEGIN: sysret-canonical\n", .{});
const blob = process.nonCanonicalReturnBlob();
check(
"the probe's system call number still matches the ABI",
blob.len > 1 and blob[0] == 0xB8 and blob[1] == @intFromEnum(abi.SystemCall.current_core),
);
const before = architecture.nonCanonicalReturnCount();
check("no return has been refused yet this boot", before == 0);
process.fault_kill_count = 0;
const probe = spawnNonCanonicalReturnProcess() orelse 0;
check("the probe process spawned", probe != 0);
// Drop below the probe so it gets the core, and wait for the kill.
scheduler.setPriority(1);
const deadline = architecture.millis() + 5000;
while (process.fault_kill_count < 1 and architecture.millis() < deadline) scheduler.yield();
scheduler.setPriority(4);
const refused = architecture.nonCanonicalReturnCount() - before;
check("the guard refused exactly one sysretq", refused == 1);
check("the probe was killed (not the machine)", process.fault_kill_count == 1);
check(
"the probe died of a protection fault — the iretq fallback faulted it in ring 3",
process.exitReasonOf(scheduler.currentId(), probe) == @intFromEnum(abi.ExitReason.protection_fault),
);
log("sysret-canonical: {d} non-canonical return(s) refused; core {d} still running\n", .{ refused, scheduler.currentCpuIndex() });
result();
}
/// Address-space refcount (docs/threading-plan.md M1): every process holds exactly one
/// reference to its address space, released when it dies, so `destroyAddressSpace` runs
/// exactly once per space — no leak, no double-free. Spawn and kill several ring-3