kernel: ring 0 cannot execute a user page
SMEP turns the classic escalation — divert kernel control flow into a page the attacker wrote — from a silent takeover into an immediate fault with the offending address in the log. The bit is per-core state, so it is set where the syscall MSRs already are: in the per-CPU bring-up both the boot processor and every application processor run on their way in. A core that climbed the trampoline without it would be a hole no boot log would show, which is why the SMP case now reads CR4 on each core it lands on and requires every one of them to be hardened, not just the one that printed the banner. Enabling it that early is only safe because nothing ring 0 executes is mapped for ring 3, and that had to be established rather than assumed: kernel text carries only its ELF flags, the physmap is no-execute, the trampoline page is mapped supervisor and the core running it has not enabled the bit yet, and the boot processor turns it on while still on the loader's tables — which map nothing user-accessible at all. The one indirect call in the kernel takes a kernel address. The CPUID probing that was scattered across the timer code becomes a small shared helper, since the feature question is now asked from two places and each wanted the same maximum-leaf guard. Absence is tolerated and reported, like the IOMMU: danos still boots on a machine without the feature, and says which one it is. The test harness starts asking QEMU for a CPU that has the bit at all — its default model has neither SMEP nor SMAP, so the code would otherwise have been unreachable in every run. No case behaved differently under the richer model. Suite 112/112, with a new case that maps an executable user page, calls into it from the kernel, and requires the fault the CPU is supposed to raise.
This commit is contained in:
@@ -0,0 +1,42 @@
|
||||
//! CPUID — the CPU describing itself.
|
||||
//!
|
||||
//! One helper for the whole architecture layer (the timer's TSC leaves, the
|
||||
//! supervisor-hardening feature bits), rather than a private copy per module.
|
||||
//! `leaf` always executes with ECX = 0, which is what every leaf danos reads
|
||||
//! wants: leaf 7's feature words live in sub-leaf 0, and leaves that ignore ECX
|
||||
//! don't care. A sub-leaf-taking caller would add its own entry point here.
|
||||
//!
|
||||
//! **Always gate on `supports` first.** CPUID does not fault on an out-of-range
|
||||
//! leaf — it returns the data of the highest supported leaf instead, which would
|
||||
//! be read as a feature bit that isn't there. The maximum lives in leaf 0 (basic
|
||||
//! range) and leaf 0x80000000 (extended range).
|
||||
|
||||
pub const Registers = struct { eax: u32, ebx: u32, ecx: u32, edx: u32 };
|
||||
|
||||
/// Execute CPUID for `number` at sub-leaf 0.
|
||||
pub fn leaf(number: u32) Registers {
|
||||
var a: u32 = undefined;
|
||||
var b: u32 = undefined;
|
||||
var c: u32 = undefined;
|
||||
var d: u32 = undefined;
|
||||
asm volatile ("cpuid"
|
||||
: [a] "={eax}" (a),
|
||||
[b] "={ebx}" (b),
|
||||
[c] "={ecx}" (c),
|
||||
[d] "={edx}" (d),
|
||||
: [leaf] "{eax}" (number),
|
||||
[sub] "{ecx}" (@as(u32, 0)),
|
||||
);
|
||||
return .{ .eax = a, .ebx = b, .ecx = c, .edx = d };
|
||||
}
|
||||
|
||||
/// Whether `number` is inside the range this CPU actually enumerates — the basic
|
||||
/// range for a leaf below 0x80000000, the extended range above it. Every read of
|
||||
/// a leaf beyond 0 or 0x80000000 must pass through here first (see the module doc).
|
||||
pub fn supports(number: u32) bool {
|
||||
const maximum = if (number >= 0x8000_0000)
|
||||
leaf(0x8000_0000).eax
|
||||
else
|
||||
leaf(0).eax;
|
||||
return maximum >= number;
|
||||
}
|
||||
Reference in New Issue
Block a user