kernel: ring 0 cannot execute a user page
SMEP turns the classic escalation — divert kernel control flow into a page the attacker wrote — from a silent takeover into an immediate fault with the offending address in the log. The bit is per-core state, so it is set where the syscall MSRs already are: in the per-CPU bring-up both the boot processor and every application processor run on their way in. A core that climbed the trampoline without it would be a hole no boot log would show, which is why the SMP case now reads CR4 on each core it lands on and requires every one of them to be hardened, not just the one that printed the banner. Enabling it that early is only safe because nothing ring 0 executes is mapped for ring 3, and that had to be established rather than assumed: kernel text carries only its ELF flags, the physmap is no-execute, the trampoline page is mapped supervisor and the core running it has not enabled the bit yet, and the boot processor turns it on while still on the loader's tables — which map nothing user-accessible at all. The one indirect call in the kernel takes a kernel address. The CPUID probing that was scattered across the timer code becomes a small shared helper, since the feature question is now asked from two places and each wanted the same maximum-leaf guard. Absence is tolerated and reported, like the IOMMU: danos still boots on a machine without the feature, and says which one it is. The test harness starts asking QEMU for a CPU that has the bit at all — its default model has neither SMEP nor SMAP, so the code would otherwise have been unreachable in every run. No case behaved differently under the richer model. Suite 112/112, with a new case that maps an executable user page, calls into it from the kernel, and requires the fault the CPU is supposed to raise.
This commit is contained in:
@@ -11,9 +11,16 @@
|
||||
//! transition is always an exit (the kernel starts in ring 0), the swap pairs
|
||||
//! keep the invariant without seeding KERNEL_GS_BASE. `scheduler()` is therefore
|
||||
//! valid in any ring-0 context and never sees a user-controlled base.
|
||||
//!
|
||||
//! The file has since become the home of **per-core CPU state set at bring-up**
|
||||
//! generally, not just the GS block: the fast-system_call MSRs and the CR4
|
||||
//! hardening bits live here too, because each is state a core owns and must set
|
||||
//! for itself. Both bring-up paths — `cpu.init` on the boot processor and
|
||||
//! `smp.apEntry` on every application processor — call the same functions here.
|
||||
|
||||
const std = @import("std");
|
||||
const io = @import("io.zig");
|
||||
const cpuid = @import("cpuid.zig");
|
||||
const parameters = @import("parameters");
|
||||
|
||||
const ia32_gs_base = 0xC000_0101;
|
||||
@@ -75,3 +82,57 @@ pub fn initSystemCall() void {
|
||||
io.wrmsr(ia32_lstar, @intFromPtr(entry));
|
||||
io.wrmsr(ia32_sfmask, 0x4_0700); // clear IF, TF, DF, AC on entry
|
||||
}
|
||||
|
||||
// --- supervisor-mode hardening (CR4) ---------------------------------------
|
||||
//
|
||||
// CR4 is per-core state, so these bits are set during *every* core's bring-up —
|
||||
// the BSP in cpu.init, each AP in smp.apEntry — and not in the AP trampoline,
|
||||
// which stays minimal and would only cover the APs anyway.
|
||||
|
||||
/// CR4.SMEP: an instruction fetch in ring 0 from a page whose U/S bit says *user*
|
||||
/// raises #PF. This is what makes the classic ret2usr shape (a kernel bug steered
|
||||
/// into attacker-prepared user code) a loud, attributable fault instead of a
|
||||
/// silent compromise. danos never executes user-mapped memory in ring 0 — kernel
|
||||
/// text lives in the higher half, the ring-3 entry paths are kernel code, and the
|
||||
/// AP trampoline page is a supervisor mapping — so nothing legitimate is refused.
|
||||
const cr4_smep: u64 = 1 << 20;
|
||||
|
||||
/// SMEP's feature bit: CPUID leaf 7, sub-leaf 0, EBX bit 7.
|
||||
fn smepSupported() bool {
|
||||
if (!cpuid.supports(7)) return false;
|
||||
return cpuid.leaf(7).ebx & (1 << 7) != 0;
|
||||
}
|
||||
|
||||
fn readCr4() u64 {
|
||||
return asm volatile ("mov %%cr4, %[out]"
|
||||
: [out] "=r" (-> u64),
|
||||
);
|
||||
}
|
||||
|
||||
fn writeCr4(value: u64) void {
|
||||
asm volatile ("mov %[in], %%cr4"
|
||||
:
|
||||
: [in] "r" (value),
|
||||
: .{ .memory = true });
|
||||
}
|
||||
|
||||
/// Turn on the supervisor-mode hardening this CPU offers, on the calling core.
|
||||
/// Called once per core, next to `initSystemCall`, from both bring-up paths.
|
||||
///
|
||||
/// **Fail-open, like the IOMMU and the clocksource:** an absent feature is a
|
||||
/// machine that boots unhardened, not a machine that refuses to boot. danos has
|
||||
/// to run on any VM, on real Intel and on real AMD; the boot log states the
|
||||
/// posture either way (see the platform block in system/kernel/kernel.zig), so
|
||||
/// an unhardened boot is visible rather than assumed.
|
||||
pub fn initHardening() void {
|
||||
if (!smepSupported()) return;
|
||||
writeCr4(readCr4() | cr4_smep);
|
||||
}
|
||||
|
||||
/// Whether supervisor-mode execution prevention is live on *this* core. Read
|
||||
/// straight out of CR4 rather than a remembered probe result, so the answer is
|
||||
/// the state the hardware is actually in — which is what both the boot log and
|
||||
/// the `fault-smep` test case want to assert.
|
||||
pub fn supervisorExecutePreventionEnabled() bool {
|
||||
return readCr4() & cr4_smep != 0;
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user