kernel: ring 0 cannot execute a user page
SMEP turns the classic escalation — divert kernel control flow into a page the attacker wrote — from a silent takeover into an immediate fault with the offending address in the log. The bit is per-core state, so it is set where the syscall MSRs already are: in the per-CPU bring-up both the boot processor and every application processor run on their way in. A core that climbed the trampoline without it would be a hole no boot log would show, which is why the SMP case now reads CR4 on each core it lands on and requires every one of them to be hardened, not just the one that printed the banner. Enabling it that early is only safe because nothing ring 0 executes is mapped for ring 3, and that had to be established rather than assumed: kernel text carries only its ELF flags, the physmap is no-execute, the trampoline page is mapped supervisor and the core running it has not enabled the bit yet, and the boot processor turns it on while still on the loader's tables — which map nothing user-accessible at all. The one indirect call in the kernel takes a kernel address. The CPUID probing that was scattered across the timer code becomes a small shared helper, since the feature question is now asked from two places and each wanted the same maximum-leaf guard. Absence is tolerated and reported, like the IOMMU: danos still boots on a machine without the feature, and says which one it is. The test harness starts asking QEMU for a CPU that has the bit at all — its default model has neither SMEP nor SMAP, so the code would otherwise have been unreachable in every run. No case behaved differently under the richer model. Suite 112/112, with a new case that maps an executable user page, calls into it from the kernel, and requires the fault the CPU is supposed to raise.
This commit is contained in:
@@ -141,11 +141,26 @@ pub fn debugconWrite(bytes: []const u8) void {
|
||||
/// stack for double faults), then the IDT with exception handlers. After this a
|
||||
/// CPU fault is reported instead of triple-faulting. Install the fault handler
|
||||
/// (setFaultHandler) first so early faults are caught.
|
||||
///
|
||||
/// This is the **boot processor's** half of per-core bring-up; smp.apEntry is the
|
||||
/// other half and must keep the per-CPU steps in step with it (the system_call
|
||||
/// MSRs and the CR4 hardening bits are per-core state, so every core sets its own).
|
||||
pub fn init() void {
|
||||
gdt.init();
|
||||
tss.init();
|
||||
idt.init();
|
||||
pcpu.initSystemCall();
|
||||
// Safe this early, before the kernel is on its own page tables: the loader's
|
||||
// bootstrap tables (boot/efi.zig) map with present|writable and never set the
|
||||
// U/S bit, so no page the BSP executes from is user-accessible.
|
||||
pcpu.initHardening();
|
||||
}
|
||||
|
||||
/// Whether ring 0 is barred from executing user-mapped pages on this core (CR4.SMEP
|
||||
/// on x86_64; the privileged-execute-never behaviour elsewhere). False means the CPU
|
||||
/// doesn't offer it and the machine is running unhardened — see per-cpu.zig.
|
||||
pub fn supervisorExecutePreventionEnabled() bool {
|
||||
return pcpu.supervisorExecutePreventionEnabled();
|
||||
}
|
||||
|
||||
/// Build the kernel's own page tables (with real permissions) and switch onto
|
||||
|
||||
Reference in New Issue
Block a user