kernel: ring 0 reaches user memory only through the checked copy

SMAP makes the rule the copy layer has followed since it was written into a
rule the hardware keeps. A ring-0 read or write of a user page now faults,
so any code that reaches for a user pointer directly fails the first time it
runs rather than the first time someone attacks it — and the suite becomes
the enforcement test, because every case exercises the kernel with the bit
on. Nothing had to be fixed to turn it on, which is the retrospective proof
that the nine stragglers converted earlier were all of them.

The interrupt entry needed one instruction first. Hardware does not clear
the alignment-check flag on its way into a handler, and ring 3 sets that
flag freely, so a process could have taken an interrupt with SMAP suspended
for the duration. The system call path was already covered — its flag mask
clears it — but the interrupt path needed a `clac`, which cannot simply be
assembled in: it is an invalid instruction on a processor without SMAP, and
danos boots on those too. So the entry ships as a three-byte NOP and is
patched at boot, through the physmap, because the kernel maps its own text
read-only.

The ordering that makes that safe is enforced rather than described: the
patch sets a flag, and no core will set the SMAP bit until it is true. A
translation that fails, or bytes that read back wrong through the address
they will actually be fetched from, leave the machine unhardened and saying
so — which is the same posture the IOMMU takes, and better than enforcing
over an entry path that cannot comply. The patch runs before interrupts are
enabled and before any second core exists; a comment says so, because the
three bytes pass through an encoding that must never be executed and a
future change that moves this later has to deal with that first.

Suite 114/114, with a case that reads a user page from ring 0 and requires
the fault, and the multi-core case asserting every core that ran work had
the bit — the same shape SMEP got, for the same reason: CR4 is per-core, and
a hardening is only as wide as its narrowest core.
This commit is contained in:
Daniel Samson
2026-08-01 12:20:44 +01:00
parent cb30faf15f
commit 5a5146ec13
10 changed files with 301 additions and 13 deletions
+31 -7
View File
@@ -36,11 +36,11 @@ plain `main` checkout always tells the truth about where the work is.**
| | |
|---|---|
| Working on | **H3** — SMAP (the last phase) |
| Branch carrying it | `feat/security-group-4` (pushed to origin) |
| On `main` | everything through P4c — groups 1, 2 and 3 merged |
| Awaiting merge | H2, HS — land with the group 4 merge |
| Suite | 113 cases, all passing |
| Working on | nothing — the track is complete |
| Branch carrying it | — |
| On `main` | every phase — groups 1, 2, 3 and 4 merged |
| Awaiting merge | nothing |
| Suite | 114 cases, all passing |
| Last updated | 2026-08-01 |
A checkbox below means the phase met its definition of green and was
@@ -95,8 +95,27 @@ group boundary.
counter, not the outcome: measured with the guard's branch removed, TCG does
not model Intel's ring-0 #GP and every outcome check still passed. Suite
113/113)
- [ ] **H3** — SMAP + boot-patched `clac`; `-cpu max` in the harness; negative tests
- [ ] **merge** group 4 → main, push
- [x] **H3** — SMAP on every core, and the boot-patched `clac` that makes it
hold. Hardware does not clear `EFLAGS.AC` on interrupt delivery and ring 3
sets it with `popfq`, so without the patch a hostile process suspends SMAP
for the length of any handler it can provoke; `clac` is #UD without the
feature, so the image ships the 3-byte canonical NOP at the first byte of
`isr_common` — ahead of the CPL test, because a fault nested in ring 0
inherits AC just as readily — and the boot processor overwrites it through
the physmap, the same door `smp.arm` and `process.run` already use to write
a page the executing mapping holds read-only. The two halves are wired
together rather than merely ordered: `initHardening` refuses CR4 bit 21
until the patch has been written *and* read back through the text address it
will be fetched from, so no core can turn SMAP on ahead of it and a
translation that lied leaves the machine unhardened and saying so instead of
enforcing over an entry path that cannot clear AC. The `smp` case now
requires SMAP on every core it lands on, which is the ordering checked from
the far end. New `fault-smap` case reads a mapped user page from ring 0 and
requires the fault: error code `0x1` — present, and nothing else, since a
SMAP violation has no bit of its own and U/S reports the ring of the access.
The suite is now the standing enforcement test, and it found nothing left to
find: H1's conversion of the nine stragglers was complete. Suite 114/114
- [x] **merge** group 4 → main, push
---
@@ -547,6 +566,11 @@ review-level proof plus the existing fault cases regression. Suite 112.
**Test:** new QEMU case `fault-smap`: ring-0 deliberate read of a mapped
user page; expect vector 14 + `error code : 0x1` + kernel IP. And the
whole suite becomes the tripwire — any missed straggler now fails loudly.
*(Landed: the patch site sits at the very first byte of `isr_common`, and
CR4.SMAP is refused on every core until the patch has been read back through
the text mapping it will execute from — so the ordering is enforced, not
merely documented. Error code observed: `0x1` exactly, present and nothing
else. The tripwire found no missed straggler: H1 had converted them all.)*
Suite 114 (HS added one).
---