kernel: ring 0 cannot execute a user page
SMEP turns the classic escalation — divert kernel control flow into a page the attacker wrote — from a silent takeover into an immediate fault with the offending address in the log. The bit is per-core state, so it is set where the syscall MSRs already are: in the per-CPU bring-up both the boot processor and every application processor run on their way in. A core that climbed the trampoline without it would be a hole no boot log would show, which is why the SMP case now reads CR4 on each core it lands on and requires every one of them to be hardened, not just the one that printed the banner. Enabling it that early is only safe because nothing ring 0 executes is mapped for ring 3, and that had to be established rather than assumed: kernel text carries only its ELF flags, the physmap is no-execute, the trampoline page is mapped supervisor and the core running it has not enabled the bit yet, and the boot processor turns it on while still on the loader's tables — which map nothing user-accessible at all. The one indirect call in the kernel takes a kernel address. The CPUID probing that was scattered across the timer code becomes a small shared helper, since the feature question is now asked from two places and each wanted the same maximum-leaf guard. Absence is tolerated and reported, like the IOMMU: danos still boots on a machine without the feature, and says which one it is. The test harness starts asking QEMU for a CPU that has the bit at all — its default model has neither SMEP nor SMAP, so the code would otherwise have been unreachable in every run. No case behaved differently under the richer model. Suite 112/112, with a new case that maps an executable user page, calls into it from the kernel, and requires the fault the CPU is supposed to raise.
This commit is contained in:
@@ -11,6 +11,7 @@
|
||||
|
||||
const boot_handoff = @import("boot-handoff");
|
||||
const io = @import("io.zig");
|
||||
const cpuid = @import("cpuid.zig");
|
||||
const paging = @import("paging.zig");
|
||||
|
||||
/// The ACPI PM timer, as a calibration reference: an I/O port or MMIO counter.
|
||||
@@ -331,8 +332,8 @@ fn calibratePit() void {
|
||||
/// TSC frequency from CPUID leaf 0x15 (crystal_hz * numerator / denominator), or
|
||||
/// null if the CPU doesn't enumerate it (common under QEMU, and on AMD).
|
||||
fn cpuidTscHz() ?u64 {
|
||||
if (cpuid(0).eax < 0x15) return null;
|
||||
const r = cpuid(0x15);
|
||||
if (!cpuid.supports(0x15)) return null;
|
||||
const r = cpuid.leaf(0x15);
|
||||
if (r.eax == 0 or r.ebx == 0 or r.ecx == 0) return null; // ratio/crystal not given
|
||||
return @as(u64, r.ecx) * r.ebx / r.eax;
|
||||
}
|
||||
@@ -342,26 +343,8 @@ fn cpuidTscHz() ?u64 {
|
||||
/// at a constant rate across P/C-states and never stops. Requires the extended-leaf
|
||||
/// range to reach 0x80000007 first.
|
||||
fn tscIsInvariant() bool {
|
||||
if (cpuid(0x80000000).eax < 0x80000007) return false;
|
||||
return (cpuid(0x80000007).edx & (1 << 8)) != 0;
|
||||
}
|
||||
|
||||
const CpuidRegs = struct { eax: u32, ebx: u32, ecx: u32, edx: u32 };
|
||||
|
||||
fn cpuid(leaf: u32) CpuidRegs {
|
||||
var a: u32 = undefined;
|
||||
var b: u32 = undefined;
|
||||
var c: u32 = undefined;
|
||||
var d: u32 = undefined;
|
||||
asm volatile ("cpuid"
|
||||
: [a] "={eax}" (a),
|
||||
[b] "={ebx}" (b),
|
||||
[c] "={ecx}" (c),
|
||||
[d] "={edx}" (d),
|
||||
: [leaf] "{eax}" (leaf),
|
||||
[sub] "{ecx}" (@as(u32, 0)),
|
||||
);
|
||||
return .{ .eax = a, .ebx = b, .ecx = c, .edx = d };
|
||||
if (!cpuid.supports(0x8000_0007)) return false;
|
||||
return (cpuid.leaf(0x8000_0007).edx & (1 << 8)) != 0;
|
||||
}
|
||||
|
||||
// HPET registers: capabilities at +0x00 (period in the high dword, in fs; bit 13 =
|
||||
|
||||
@@ -141,11 +141,26 @@ pub fn debugconWrite(bytes: []const u8) void {
|
||||
/// stack for double faults), then the IDT with exception handlers. After this a
|
||||
/// CPU fault is reported instead of triple-faulting. Install the fault handler
|
||||
/// (setFaultHandler) first so early faults are caught.
|
||||
///
|
||||
/// This is the **boot processor's** half of per-core bring-up; smp.apEntry is the
|
||||
/// other half and must keep the per-CPU steps in step with it (the system_call
|
||||
/// MSRs and the CR4 hardening bits are per-core state, so every core sets its own).
|
||||
pub fn init() void {
|
||||
gdt.init();
|
||||
tss.init();
|
||||
idt.init();
|
||||
pcpu.initSystemCall();
|
||||
// Safe this early, before the kernel is on its own page tables: the loader's
|
||||
// bootstrap tables (boot/efi.zig) map with present|writable and never set the
|
||||
// U/S bit, so no page the BSP executes from is user-accessible.
|
||||
pcpu.initHardening();
|
||||
}
|
||||
|
||||
/// Whether ring 0 is barred from executing user-mapped pages on this core (CR4.SMEP
|
||||
/// on x86_64; the privileged-execute-never behaviour elsewhere). False means the CPU
|
||||
/// doesn't offer it and the machine is running unhardened — see per-cpu.zig.
|
||||
pub fn supervisorExecutePreventionEnabled() bool {
|
||||
return pcpu.supervisorExecutePreventionEnabled();
|
||||
}
|
||||
|
||||
/// Build the kernel's own page tables (with real permissions) and switch onto
|
||||
|
||||
@@ -0,0 +1,42 @@
|
||||
//! CPUID — the CPU describing itself.
|
||||
//!
|
||||
//! One helper for the whole architecture layer (the timer's TSC leaves, the
|
||||
//! supervisor-hardening feature bits), rather than a private copy per module.
|
||||
//! `leaf` always executes with ECX = 0, which is what every leaf danos reads
|
||||
//! wants: leaf 7's feature words live in sub-leaf 0, and leaves that ignore ECX
|
||||
//! don't care. A sub-leaf-taking caller would add its own entry point here.
|
||||
//!
|
||||
//! **Always gate on `supports` first.** CPUID does not fault on an out-of-range
|
||||
//! leaf — it returns the data of the highest supported leaf instead, which would
|
||||
//! be read as a feature bit that isn't there. The maximum lives in leaf 0 (basic
|
||||
//! range) and leaf 0x80000000 (extended range).
|
||||
|
||||
pub const Registers = struct { eax: u32, ebx: u32, ecx: u32, edx: u32 };
|
||||
|
||||
/// Execute CPUID for `number` at sub-leaf 0.
|
||||
pub fn leaf(number: u32) Registers {
|
||||
var a: u32 = undefined;
|
||||
var b: u32 = undefined;
|
||||
var c: u32 = undefined;
|
||||
var d: u32 = undefined;
|
||||
asm volatile ("cpuid"
|
||||
: [a] "={eax}" (a),
|
||||
[b] "={ebx}" (b),
|
||||
[c] "={ecx}" (c),
|
||||
[d] "={edx}" (d),
|
||||
: [leaf] "{eax}" (number),
|
||||
[sub] "{ecx}" (@as(u32, 0)),
|
||||
);
|
||||
return .{ .eax = a, .ebx = b, .ecx = c, .edx = d };
|
||||
}
|
||||
|
||||
/// Whether `number` is inside the range this CPU actually enumerates — the basic
|
||||
/// range for a leaf below 0x80000000, the extended range above it. Every read of
|
||||
/// a leaf beyond 0 or 0x80000000 must pass through here first (see the module doc).
|
||||
pub fn supports(number: u32) bool {
|
||||
const maximum = if (number >= 0x8000_0000)
|
||||
leaf(0x8000_0000).eax
|
||||
else
|
||||
leaf(0).eax;
|
||||
return maximum >= number;
|
||||
}
|
||||
@@ -11,9 +11,16 @@
|
||||
//! transition is always an exit (the kernel starts in ring 0), the swap pairs
|
||||
//! keep the invariant without seeding KERNEL_GS_BASE. `scheduler()` is therefore
|
||||
//! valid in any ring-0 context and never sees a user-controlled base.
|
||||
//!
|
||||
//! The file has since become the home of **per-core CPU state set at bring-up**
|
||||
//! generally, not just the GS block: the fast-system_call MSRs and the CR4
|
||||
//! hardening bits live here too, because each is state a core owns and must set
|
||||
//! for itself. Both bring-up paths — `cpu.init` on the boot processor and
|
||||
//! `smp.apEntry` on every application processor — call the same functions here.
|
||||
|
||||
const std = @import("std");
|
||||
const io = @import("io.zig");
|
||||
const cpuid = @import("cpuid.zig");
|
||||
const parameters = @import("parameters");
|
||||
|
||||
const ia32_gs_base = 0xC000_0101;
|
||||
@@ -75,3 +82,57 @@ pub fn initSystemCall() void {
|
||||
io.wrmsr(ia32_lstar, @intFromPtr(entry));
|
||||
io.wrmsr(ia32_sfmask, 0x4_0700); // clear IF, TF, DF, AC on entry
|
||||
}
|
||||
|
||||
// --- supervisor-mode hardening (CR4) ---------------------------------------
|
||||
//
|
||||
// CR4 is per-core state, so these bits are set during *every* core's bring-up —
|
||||
// the BSP in cpu.init, each AP in smp.apEntry — and not in the AP trampoline,
|
||||
// which stays minimal and would only cover the APs anyway.
|
||||
|
||||
/// CR4.SMEP: an instruction fetch in ring 0 from a page whose U/S bit says *user*
|
||||
/// raises #PF. This is what makes the classic ret2usr shape (a kernel bug steered
|
||||
/// into attacker-prepared user code) a loud, attributable fault instead of a
|
||||
/// silent compromise. danos never executes user-mapped memory in ring 0 — kernel
|
||||
/// text lives in the higher half, the ring-3 entry paths are kernel code, and the
|
||||
/// AP trampoline page is a supervisor mapping — so nothing legitimate is refused.
|
||||
const cr4_smep: u64 = 1 << 20;
|
||||
|
||||
/// SMEP's feature bit: CPUID leaf 7, sub-leaf 0, EBX bit 7.
|
||||
fn smepSupported() bool {
|
||||
if (!cpuid.supports(7)) return false;
|
||||
return cpuid.leaf(7).ebx & (1 << 7) != 0;
|
||||
}
|
||||
|
||||
fn readCr4() u64 {
|
||||
return asm volatile ("mov %%cr4, %[out]"
|
||||
: [out] "=r" (-> u64),
|
||||
);
|
||||
}
|
||||
|
||||
fn writeCr4(value: u64) void {
|
||||
asm volatile ("mov %[in], %%cr4"
|
||||
:
|
||||
: [in] "r" (value),
|
||||
: .{ .memory = true });
|
||||
}
|
||||
|
||||
/// Turn on the supervisor-mode hardening this CPU offers, on the calling core.
|
||||
/// Called once per core, next to `initSystemCall`, from both bring-up paths.
|
||||
///
|
||||
/// **Fail-open, like the IOMMU and the clocksource:** an absent feature is a
|
||||
/// machine that boots unhardened, not a machine that refuses to boot. danos has
|
||||
/// to run on any VM, on real Intel and on real AMD; the boot log states the
|
||||
/// posture either way (see the platform block in system/kernel/kernel.zig), so
|
||||
/// an unhardened boot is visible rather than assumed.
|
||||
pub fn initHardening() void {
|
||||
if (!smepSupported()) return;
|
||||
writeCr4(readCr4() | cr4_smep);
|
||||
}
|
||||
|
||||
/// Whether supervisor-mode execution prevention is live on *this* core. Read
|
||||
/// straight out of CR4 rather than a remembered probe result, so the answer is
|
||||
/// the state the hardware is actually in — which is what both the boot log and
|
||||
/// the `fault-smep` test case want to assert.
|
||||
pub fn supervisorExecutePreventionEnabled() bool {
|
||||
return readCr4() & cr4_smep != 0;
|
||||
}
|
||||
|
||||
@@ -179,6 +179,10 @@ fn apEntry(percpu: usize) callconv(.c) noreturn {
|
||||
idt.loadOnThisCpu(); // the shared IDT
|
||||
pcpu.setLocal(cpu, percpu); // per-CPU block via GS base — *after* the GDT reload
|
||||
pcpu.initSystemCall(); // enable system_call/sysret on this core
|
||||
// CR4 is per-core: this core starts from the trampoline's CR4 (PAE + SSE only),
|
||||
// so it sets its own hardening bits here rather than in the trampoline — one
|
||||
// Zig code path shared with the BSP (cpu.init), and the trampoline stays minimal.
|
||||
pcpu.initHardening(); // CR4.SMEP on this core, if the CPU has it
|
||||
|
||||
apic.initSecondary(); // software-enable this core's LAPIC
|
||||
apic.initTimer(apic.frequencyHz()); // arm its timer (still masked: interrupts off)
|
||||
|
||||
Reference in New Issue
Block a user