kernel: ring 0 reaches user memory only through the checked copy
SMAP makes the rule the copy layer has followed since it was written into a rule the hardware keeps. A ring-0 read or write of a user page now faults, so any code that reaches for a user pointer directly fails the first time it runs rather than the first time someone attacks it — and the suite becomes the enforcement test, because every case exercises the kernel with the bit on. Nothing had to be fixed to turn it on, which is the retrospective proof that the nine stragglers converted earlier were all of them. The interrupt entry needed one instruction first. Hardware does not clear the alignment-check flag on its way into a handler, and ring 3 sets that flag freely, so a process could have taken an interrupt with SMAP suspended for the duration. The system call path was already covered — its flag mask clears it — but the interrupt path needed a `clac`, which cannot simply be assembled in: it is an invalid instruction on a processor without SMAP, and danos boots on those too. So the entry ships as a three-byte NOP and is patched at boot, through the physmap, because the kernel maps its own text read-only. The ordering that makes that safe is enforced rather than described: the patch sets a flag, and no core will set the SMAP bit until it is true. A translation that fails, or bytes that read back wrong through the address they will actually be fetched from, leave the machine unhardened and saying so — which is the same posture the IOMMU takes, and better than enforcing over an entry path that cannot comply. The patch runs before interrupts are enabled and before any second core exists; a comment says so, because the three bytes pass through an encoding that must never be executed and a future change that moves this later has to deal with that first. Suite 114/114, with a case that reads a user page from ring 0 and requires the fault, and the multi-core case asserting every core that ran work had the bit — the same shape SMEP got, for the same reason: CR4 is per-core, and a hardening is only as wide as its narrowest core.
This commit is contained in:
@@ -442,7 +442,27 @@ STUB_NOERR 128
|
||||
# If the interrupt came from ring 3 the GS base holds the user's value, so swap
|
||||
# in the kernel's before anything reads per-CPU data (swapgs discipline; see
|
||||
# percpu.zig). CS sits at offset 24 here (vector@0, error@8, RIP@16, CS@24).
|
||||
#
|
||||
# The first three bytes are the SMAP guard, and they come before everything —
|
||||
# before the CPL test, before the swapgs. Interrupt delivery does not clear
|
||||
# EFLAGS.AC (SYSCALL does, through SFMASK; an IDT gate does not), and ring 3 sets
|
||||
# AC freely with popfq, so without this a hostile process could take an interrupt
|
||||
# with AC=1 and have the whole handler run with SMAP suspended. Ahead of the CPL
|
||||
# test because ring 0 inherits AC just as readily: a fault or IRQ nested inside
|
||||
# kernel code carries whatever AC the interrupted context had, and that context
|
||||
# may itself be an entry that has not reached its own guard yet. Clearing first,
|
||||
# unconditionally, means no path into the kernel is ever a path in with AC set.
|
||||
#
|
||||
# `clac` is #UD on a CPU without SMAP, so the image ships the 3-byte canonical NOP
|
||||
# (`nopl (%rax)`) and per-cpu.zig overwrites it with `clac` (0f 01 ca) at boot,
|
||||
# on the boot processor, only when CPUID says the instruction exists — a one-time
|
||||
# patch rather than a branch in the hottest path in the kernel. Neither encoding
|
||||
# touches the flags the `testb` below sets, so the guard is invisible to the code
|
||||
# that follows it either way.
|
||||
.global isr_smap_patch
|
||||
isr_common:
|
||||
isr_smap_patch:
|
||||
.byte 0x0f, 0x1f, 0x00 # nopl (%rax) -> patched to `clac` when the CPU has SMAP
|
||||
testb $3, 24(%rsp)
|
||||
jz 1f
|
||||
swapgs
|
||||
|
||||
Reference in New Issue
Block a user