kernel: ring 0 reaches user memory only through the checked copy

SMAP makes the rule the copy layer has followed since it was written into a
rule the hardware keeps. A ring-0 read or write of a user page now faults,
so any code that reaches for a user pointer directly fails the first time it
runs rather than the first time someone attacks it — and the suite becomes
the enforcement test, because every case exercises the kernel with the bit
on. Nothing had to be fixed to turn it on, which is the retrospective proof
that the nine stragglers converted earlier were all of them.

The interrupt entry needed one instruction first. Hardware does not clear
the alignment-check flag on its way into a handler, and ring 3 sets that
flag freely, so a process could have taken an interrupt with SMAP suspended
for the duration. The system call path was already covered — its flag mask
clears it — but the interrupt path needed a `clac`, which cannot simply be
assembled in: it is an invalid instruction on a processor without SMAP, and
danos boots on those too. So the entry ships as a three-byte NOP and is
patched at boot, through the physmap, because the kernel maps its own text
read-only.

The ordering that makes that safe is enforced rather than described: the
patch sets a flag, and no core will set the SMAP bit until it is true. A
translation that fails, or bytes that read back wrong through the address
they will actually be fetched from, leave the machine unhardened and saying
so — which is the same posture the IOMMU takes, and better than enforcing
over an entry path that cannot comply. The patch runs before interrupts are
enabled and before any second core exists; a comment says so, because the
three bytes pass through an encoding that must never be executed and a
future change that moves this later has to deal with that first.

Suite 114/114, with a case that reads a user page from ring 0 and requires
the fault, and the multi-core case asserting every core that ran work had
the bit — the same shape SMEP got, for the same reason: CR4 is per-core, and
a hardening is only as wide as its narrowest core.
This commit is contained in:
Daniel Samson
2026-08-01 12:20:44 +01:00
parent cb30faf15f
commit 5a5146ec13
10 changed files with 301 additions and 13 deletions
+20
View File
@@ -182,6 +182,26 @@ test for "is this an abbreviation I must expand" is simply: *is there a longer w
is a clipped form of?* If yes, write the word. If it's an initialism standing in for a
phrase, leave it.
## Kernel code touches user memory only through `user-memory`
A syscall argument is an attacker-controlled integer. Kernel code never
dereferences one: every read of a process's memory goes through
`copyFromUser` and every write through `copyToUser`
(`system/kernel/user-memory.zig`), which walk that address space's page tables
and move the bytes through the physmap, with the permissions ring 3 itself
would face. A bad pointer then fails the call instead of faulting the kernel,
and a struct pulled in once cannot change underneath the checks that follow it.
This is not a review convention — CR4.SMAP enforces it in hardware
(`docs/os-development/smep-smap.md`), so a raw dereference of a user address is
a #PF with a kernel instruction pointer the first time the QEMU suite reaches
it. Which is also why **there is no `stac` in this tree, and never should be**:
`stac` suspends exactly that enforcement, the copy layer needs no such window
by construction, and a change that adds one has removed the guarantee rather
than worked around a limitation. The same goes for the boot-time `clac` patch
at the interrupt entry — it exists so that ring 3 cannot suspend SMAP either,
by taking an interrupt with `EFLAGS.AC` set.
## Zen of Zig
* Communicate intent precisely.