kernel: ring 0 cannot execute a user page

SMEP turns the classic escalation — divert kernel control flow into a page
the attacker wrote — from a silent takeover into an immediate fault with
the offending address in the log. The bit is per-core state, so it is set
where the syscall MSRs already are: in the per-CPU bring-up both the boot
processor and every application processor run on their way in. A core that
climbed the trampoline without it would be a hole no boot log would show,
which is why the SMP case now reads CR4 on each core it lands on and
requires every one of them to be hardened, not just the one that printed
the banner.

Enabling it that early is only safe because nothing ring 0 executes is
mapped for ring 3, and that had to be established rather than assumed:
kernel text carries only its ELF flags, the physmap is no-execute, the
trampoline page is mapped supervisor and the core running it has not
enabled the bit yet, and the boot processor turns it on while still on the
loader's tables — which map nothing user-accessible at all. The one
indirect call in the kernel takes a kernel address.

The CPUID probing that was scattered across the timer code becomes a small
shared helper, since the feature question is now asked from two places and
each wanted the same maximum-leaf guard. Absence is tolerated and reported,
like the IOMMU: danos still boots on a machine without the feature, and
says which one it is.

The test harness starts asking QEMU for a CPU that has the bit at all —
its default model has neither SMEP nor SMAP, so the code would otherwise
have been unreachable in every run. No case behaved differently under the
richer model.

Suite 112/112, with a new case that maps an executable user page, calls
into it from the kernel, and requires the fault the CPU is supposed to
raise.
This commit is contained in:
Daniel Samson
2026-08-01 10:05:58 +01:00
parent b3011fbb40
commit 36a7cc5fe9
10 changed files with 231 additions and 29 deletions
@@ -0,0 +1,42 @@
//! CPUID — the CPU describing itself.
//!
//! One helper for the whole architecture layer (the timer's TSC leaves, the
//! supervisor-hardening feature bits), rather than a private copy per module.
//! `leaf` always executes with ECX = 0, which is what every leaf danos reads
//! wants: leaf 7's feature words live in sub-leaf 0, and leaves that ignore ECX
//! don't care. A sub-leaf-taking caller would add its own entry point here.
//!
//! **Always gate on `supports` first.** CPUID does not fault on an out-of-range
//! leaf — it returns the data of the highest supported leaf instead, which would
//! be read as a feature bit that isn't there. The maximum lives in leaf 0 (basic
//! range) and leaf 0x80000000 (extended range).
pub const Registers = struct { eax: u32, ebx: u32, ecx: u32, edx: u32 };
/// Execute CPUID for `number` at sub-leaf 0.
pub fn leaf(number: u32) Registers {
var a: u32 = undefined;
var b: u32 = undefined;
var c: u32 = undefined;
var d: u32 = undefined;
asm volatile ("cpuid"
: [a] "={eax}" (a),
[b] "={ebx}" (b),
[c] "={ecx}" (c),
[d] "={edx}" (d),
: [leaf] "{eax}" (number),
[sub] "{ecx}" (@as(u32, 0)),
);
return .{ .eax = a, .ebx = b, .ecx = c, .edx = d };
}
/// Whether `number` is inside the range this CPU actually enumerates — the basic
/// range for a leaf below 0x80000000, the extended range above it. Every read of
/// a leaf beyond 0 or 0x80000000 must pass through here first (see the module doc).
pub fn supports(number: u32) bool {
const maximum = if (number >= 0x8000_0000)
leaf(0x8000_0000).eax
else
leaf(0).eax;
return maximum >= number;
}