kernel: ring 0 reaches user memory only through the checked copy

SMAP makes the rule the copy layer has followed since it was written into a
rule the hardware keeps. A ring-0 read or write of a user page now faults,
so any code that reaches for a user pointer directly fails the first time it
runs rather than the first time someone attacks it — and the suite becomes
the enforcement test, because every case exercises the kernel with the bit
on. Nothing had to be fixed to turn it on, which is the retrospective proof
that the nine stragglers converted earlier were all of them.

The interrupt entry needed one instruction first. Hardware does not clear
the alignment-check flag on its way into a handler, and ring 3 sets that
flag freely, so a process could have taken an interrupt with SMAP suspended
for the duration. The system call path was already covered — its flag mask
clears it — but the interrupt path needed a `clac`, which cannot simply be
assembled in: it is an invalid instruction on a processor without SMAP, and
danos boots on those too. So the entry ships as a three-byte NOP and is
patched at boot, through the physmap, because the kernel maps its own text
read-only.

The ordering that makes that safe is enforced rather than described: the
patch sets a flag, and no core will set the SMAP bit until it is true. A
translation that fails, or bytes that read back wrong through the address
they will actually be fetched from, leave the machine unhardened and saying
so — which is the same posture the IOMMU takes, and better than enforcing
over an entry path that cannot comply. The patch runs before interrupts are
enabled and before any second core exists; a comment says so, because the
three bytes pass through an encoding that must never be executed and a
future change that moves this later has to deal with that first.

Suite 114/114, with a case that reads a user page from ring 0 and requires
the fault, and the multi-core case asserting every core that ran work had
the bit — the same shape SMEP got, for the same reason: CR4 is per-core, and
a hardening is only as wide as its narrowest core.
This commit is contained in:
Daniel Samson
2026-08-01 12:20:44 +01:00
parent cb30faf15f
commit 5a5146ec13
10 changed files with 301 additions and 13 deletions
+113 -2
View File
@@ -17,10 +17,18 @@
//! hardening bits live here too, because each is state a core owns and must set
//! for itself. Both bring-up paths — `cpu.init` on the boot processor and
//! `smp.apEntry` on every application processor — call the same functions here.
//!
//! One deliberate exception to "per-core": `armSupervisorAccessPrevention` patches
//! a machine-wide instruction into the shared interrupt entry, once, on the boot
//! processor. It lives here anyway because it is the other half of CR4.SMAP —
//! same feature probe, same fail-open posture — and splitting a hardening measure
//! across two files is how the halves drift apart.
const std = @import("std");
const io = @import("io.zig");
const cpuid = @import("cpuid.zig");
const paging = @import("paging.zig");
const boot_handoff = @import("boot-handoff");
const parameters = @import("parameters");
const ia32_gs_base = 0xC000_0101;
@@ -112,6 +120,95 @@ fn smepSupported() bool {
return cpuid.leaf(7).ebx & (1 << 7) != 0;
}
/// CR4.SMAP: a ring-0 data *read or write* to a page whose U/S bit says *user*
/// raises #PF, unless EFLAGS.AC is set. It is the standing enforcement behind
/// system/kernel/user-memory.zig: that layer never dereferences a user virtual
/// address — it walks the process's tables and moves bytes through the physmap,
/// kernel mappings throughout — so nothing legitimate in this kernel is refused,
/// and any future code that reaches for a user pointer directly faults the first
/// time it runs. danos therefore opens no `stac` window anywhere; there is no
/// correct reason to have one, and adding one is how the guarantee is lost.
const cr4_smap: u64 = 1 << 21;
/// SMAP's feature bit: CPUID leaf 7, sub-leaf 0, EBX bit 20.
fn smapSupported() bool {
if (!cpuid.supports(7)) return false;
return cpuid.leaf(7).ebx & (1 << 20) != 0;
}
/// `clac` — the three bytes that replace the NOP at `isr_smap_patch` once the
/// interrupt entry is allowed to execute them.
const clac_opcode = [_]u8{ 0x0f, 0x01, 0xca };
/// Set only once the interrupt entry really clears AC, and read by every core
/// before it turns SMAP on. Nothing turns SMAP on until this is true, so there is
/// no window — not even on the boot processor, not even before ring 3 exists —
/// in which the bit is live while an interrupt could still be taken with AC set.
var interrupt_entry_clears_ac = false;
fn readCr3() u64 {
return asm volatile ("mov %%cr3, %[out]"
: [out] "=r" (-> u64),
);
}
/// Patch the `clac` into the shared interrupt entry, and by doing so authorize
/// CR4.SMAP. **Boot processor only, once, before any core sets the bit and before
/// the application processors are woken** — `initHardening` refuses SMAP until
/// this has run, so the order is enforced rather than merely documented, and an AP
/// climbing the trampoline cannot get ahead of it.
///
/// The write goes through the physmap, not through the kernel's own view of its
/// text: once `paging.init` has run, kernel `.text` is mapped read-only under W^X,
/// and a store to it would fault (or, worse, silently need CR0.WP cleared). The
/// physmap alias of the same frame is an ordinary supervisor RW mapping — the same
/// door `smp.arm` uses to write the AP trampoline and `process.run` uses to fill a
/// read-only user code frame. Going through it also makes this correct under
/// *either* set of tables: the loader's bootstrap tables map the image writable,
/// the kernel's own do not, and this runs before the switch.
///
/// Three bytes, translated one at a time, so a patch site that straddles a page
/// boundary is not a special case. A translation that fails leaves the NOP in
/// place and returns false, and the machine then boots without SMAP rather than
/// with SMAP and an entry path that cannot clear AC.
pub fn armSupervisorAccessPrevention() bool {
if (!smapSupported()) return false;
const site = @intFromPtr(@extern([*]const u8, .{ .name = "isr_smap_patch" }));
const root = readCr3() & 0x000F_FFFF_FFFF_F000;
// MUST run with interrupts masked, and does: this is reached from cpu.init,
// long before the kernel's `sti`, and before any application processor exists.
// The reason is that the three bytes go in one at a time, and the middle state
// — 0f 01 00, once the second byte lands — is `sgdt (%rax)`, a ten-byte write
// to wherever RAX points, sitting at the first instruction of every interrupt
// entry. Nothing can take that entry here, so nothing can execute it. A future
// change that moves this after interrupts are enabled has to close that window
// first (one 16-bit store covering both changed bytes is the shape, but it
// needs the two to be physically contiguous and 2-byte aligned — a naive
// version of exactly that triple-faulted this kernel).
for (clac_opcode, 0..) |byte, i| {
const physical = paging.translateIn(root, site + i) orelse return false;
const alias: *volatile u8 = @ptrFromInt(boot_handoff.physicalToVirtual(physical));
alias.* = byte;
}
// The bytes were written through a different linear address than the one they
// will be fetched from, so serialize before anyone can execute them: CPUID is
// the architecturally sanctioned way to discard whatever the core prefetched or
// decoded of the old encoding.
_ = cpuid.leaf(0);
// Read back through the *text* address, not the alias just written: that is the
// view the CPU will fetch from, so this is what proves the two are the same
// physical page and the patch landed where it will actually execute. A mismatch
// means the translation lied, and SMAP stays off rather than being enabled over
// an entry path that cannot clear AC.
const installed: [*]const volatile u8 = @ptrFromInt(site);
for (clac_opcode, 0..) |byte, i| {
if (installed[i] != byte) return false;
}
interrupt_entry_clears_ac = true;
return true;
}
fn readCr4() u64 {
return asm volatile ("mov %%cr4, %[out]"
: [out] "=r" (-> u64),
@@ -134,8 +231,14 @@ fn writeCr4(value: u64) void {
/// posture either way (see the platform block in system/kernel/kernel.zig), so
/// an unhardened boot is visible rather than assumed.
pub fn initHardening() void {
if (!smepSupported()) return;
writeCr4(readCr4() | cr4_smep);
var bits: u64 = 0;
if (smepSupported()) bits |= cr4_smep;
// SMAP only after the boot processor has armed the interrupt entry: with the
// NOP still in place an interrupt inherits ring 3's AC and suspends SMAP for
// the length of the handler, which is worse than not claiming the bit at all.
if (smapSupported() and interrupt_entry_clears_ac) bits |= cr4_smap;
if (bits == 0) return;
writeCr4(readCr4() | bits);
}
/// Whether supervisor-mode execution prevention is live on *this* core. Read
@@ -145,3 +248,11 @@ pub fn initHardening() void {
pub fn supervisorExecutePreventionEnabled() bool {
return readCr4() & cr4_smep != 0;
}
/// Whether supervisor-mode *access* prevention is live on *this* core. Read from
/// CR4 for the same reason as its neighbour: the question the boot log and the
/// `fault-smap` case are asking is what the hardware is doing, not what a probe
/// once concluded.
pub fn supervisorAccessPreventionEnabled() bool {
return readCr4() & cr4_smap != 0;
}