Hardened paging into a real VMM

This commit is contained in:
2026-07-03 13:29:34 +01:00
parent d910e6261a
commit 20b4661ff3
10 changed files with 328 additions and 120 deletions
+78 -55
View File
@@ -1,78 +1,101 @@
# Paging: the kernel's own page tables
# Paging: the kernel's page tables and VMM
Every memory access the CPU makes goes through the **page tables**: hardware walks
them to translate a virtual address into a physical one, and faults if there's no
valid mapping. Up to now danos ran on the *firmware's* page tables — which live in
memory we'd like to reclaim and which we don't control. This step builds the
kernel's own tables and switches onto them.
valid mapping. danos builds its own tables (rather than staying on the firmware's,
which live in memory we'd like to reclaim and don't control), switches CR3 onto
them, and — crucially — maps with **real permissions**.
It's x86_64-specific (the 4-level table format is an Intel/AMD thing), so it lives
behind the [arch](arch.md) boundary in `src/arch/x86_64/paging.zig`.
## What we map, and why identity
## The format
x86_64 uses **4 levels**: PML4 → PDPT → PD → PT, each a 512-entry table, with 9
bits of the virtual address indexing each level. A leaf can be a 4 KiB page (at
the PT level) or, with the "huge" bit, a 2 MiB page straight from the PD.
bits of the virtual address indexing each level and the low 12 bits the offset into
the final 4 KiB page. Each entry holds a physical address plus flag bits —
present, writable, and (bit 63) **no-execute**. danos maps everything with 4 KiB
pages: precise, and the extra table memory is negligible against available RAM.
For this first set of tables we **identity-map the low 4 GiB** — virtual address
equals physical address. Identity mapping is the pragmatic bootstrap: it keeps
everything that's already running valid across the CR3 switch without relocating a
single thing. The kernel image (at 1 MiB), the stack, the frame allocator's
bitmap, the framebuffer (at 2 GiB), and device MMIO all sit in the low 4 GiB, so
one flat identity range covers the lot.
## What gets mapped, and with what permissions
Using 2 MiB pages makes the whole map tiny: 4 GiB ÷ 2 MiB = 2048 leaves, which is
one PML4, one PDPT, and four page directories — six frames total, from the
[frame allocator](frame-allocator.md).
The address space is built in three passes (`init`):
## Building and switching
1. **All RAM, identity-mapped RW + NX.** Every non-MMIO region from the
[memory map](memory-map.md) is mapped virtual == physical, read-write and
*non-executable*. Identity mapping keeps everything already running valid across
the CR3 switch (the frame allocator addresses frames by physical address, page
tables are reached the same way, the stack stays put).
2. **The framebuffer and the Local APIC**, the device memory we actually touch,
also RW + NX. Everything else — unbacked address space, other MMIO — is simply
left unmapped, so a stray access faults instead of silently succeeding.
3. **The kernel's own segments, overlaid with their true ELF permissions.** This is
the interesting part.
`paging.init` takes a frame allocator and:
### W^X from the ELF program headers
1. Allocates and zeroes a PML4. **Zeroing matters**: the frame is recycled memory,
and any stale non-zero entry would map a bogus region.
2. Walks PML4 → PDPT → PD for each 2 MiB address, creating intermediate tables on
demand (`descend`) and writing the leaf with present + writable + huge.
3. Loads the PML4's physical address into **CR3**, which both switches address
spaces and flushes the TLB in one instruction.
Blanket RW+NX is fine for data but wrong for the kernel's own code, which must be
executable — and its code must *not* be writable (W^X: no page is both). We get the
right permissions per region straight from the kernel ELF: the **loader already
parses the program headers**, so `efi.zig` records each `PT_LOAD` segment's
address, size and R/W/X flags into `BootInfo`. Pass 3 re-maps those ranges with
flags derived from the ELF flags:
This all works because the firmware's identity map is still active *while we
build*, so a freshly allocated frame's physical address is usable directly as a
pointer. After the CR3 load we're on our own tables — and since those page-table
frames came from low RAM, they remain mapped (and thus editable) for later.
| segment | ELF flags | mapped as |
|---------|-----------|-----------|
| `.text` | R + X | present, **not** writable, **not** NX |
| `.rodata` | R | present, not writable, NX |
| `.data`/`.bss` | R + W | present, writable, NX |
So code can execute but not be written, and data can be written but not executed.
(Intermediate table entries are left writable and executable so the *leaf's* bits
govern — a page is writable only if every level is, and non-executable if any level
is.) NX itself has to be switched on first via `EFER.NXE`, or the NX bit would be a
reserved bit and fault.
### The null guard
Page 0 is deliberately left unmapped. A null (or near-null) pointer dereference now
takes a page fault instead of quietly reading or writing real memory — turning a
whole class of silent bugs into an immediate, located crash.
## Switching on, and the on-demand API
Loading the PML4's physical address into **CR3** switches address spaces and
flushes the TLB in one step. This works because the firmware's identity map is
still active *while we build*, so freshly allocated table frames are reachable by
physical address; afterwards they're covered by pass 1.
`init` keeps the PML4 and the frame allocator around and exposes `map(virt, phys,
writable)` / `unmap(virt)` (with `invlpg` TLB invalidation) — the primitive the
kernel heap will build on to map pages on demand.
## Verifying it
Two temporary tests confirmed both that we switched tables and that faults are
caught with the right detail:
Four tests (see [testing.md](testing.md)) pin down the guarantees:
- **Touch an address above 4 GiB** (`0xdeadbeef000`). Under the firmware's tables
this *didn't* fault (they map a huge range); under ours it does:
- **`vmm`** — map a fresh frame at an unused virtual address, write and read it
back. Proves `map` works end to end.
- **`fault-pf`** — an access far above all mapped RAM faults, with the address in
CR2. Proves we're on our own (deliberately sparse) map.
- **`fault-nx`** — calling into a data page (NX) faults on the instruction fetch.
Proves NX is enforced.
- **`fault-null`** — writing to address 0 faults. Proves the null guard.
```
CPU EXCEPTION: page fault (vector 14)
error code : 0x2 (write, page not present)
CR2 (addr) : 0x00000deadbeef000 (the faulting address)
```
A #PF at exactly the unmapped address is proof we're on our own, deliberately
smaller, map — and that CR2 reporting works (see [interrupts.md](interrupts.md)).
- The console keeps working *after* the switch, confirming the framebuffer, kernel
code, and stack are all still mapped.
> Toolchain note: a volatile store to a *compile-time-constant* address tripped a
> codegen bug in the Zig self-hosted x86_64 backend ("no encoding for mov moffs").
> Computing the address in a runtime variable sidesteps it (register-relative
> store). Worth remembering when poking fixed MMIO addresses.
> Toolchain notes, both hit while writing the tests: a volatile access to a
> compile-time-*constant* address either trips the self-hosted backend's
> `mov moffs` gap or (for address 0) Zig's null-pointer safety check — so the
> null-guard test launders the address through empty asm and uses an `allowzero`
> pointer to force a real hardware access. And `invlpg`, like `lgdt`, needs its
> operand staged through a register in inline asm.
## What's next (not done here)
- **A higher-half kernel**: map the kernel at a high virtual base (e.g.
`0xffffffff80000000`) so user address space can own the low half later.
- **Real permissions**: today every page is writable and executable. Map code
read-execute, data read-write + no-execute (needs setting EFER.NXE and the NX
bit), and leave a guard page unmapped to catch null-ish dereferences.
- **4 KiB pages / a general `map(virt, phys, flags)`** for fine-grained mappings,
and unmapping with TLB invalidation (`invlpg`).
- **Per-address-space tables** once there are user processes.
- **A kernel heap** — the first real user of `map`, giving the kernel dynamic
allocation. This is the natural next milestone.
- **A higher-half kernel**: relink the kernel at a high virtual base so a future
user address space can own the low half.
- **Per-address-space tables** once there are user processes, and shared/copy-on-
write mappings.
- **Uncacheable MMIO**: the APIC/framebuffer pages are mapped writeback-cacheable;
real hardware wants MMIO marked uncacheable.
+4
View File
@@ -46,9 +46,13 @@ Current cases:
| Case | What it checks | How the harness confirms it |
|------|----------------|-----------------------------|
| `smoke` | memory map has usable RAM; frame alloc/free; paging active | `DANOS-TEST-RESULT: PASS` |
| `timer` | device interrupts fire and return (tick count advances) | `DANOS-TEST-RESULT: PASS` |
| `vmm` | on-demand `map` works: a mapped page is writable and reads back | `DANOS-TEST-RESULT: PASS` |
| `fault-ud` | invalid-opcode exception is caught | serial shows `invalid opcode (vector 6)` |
| `fault-pf` | page fault caught with CR2 | `page fault (vector 14)` |
| `fault-df` | double fault caught on IST1 (not a triple-fault reset) | `double fault (vector 8)` |
| `fault-nx` | executing a data page (NX) faults | `page fault (vector 14)` |
| `fault-null` | dereferencing the unmapped page 0 faults | `page fault (vector 14)` |
The faulting cases don't print a result line — they deliberately raise a CPU
exception, and the harness asserts on the [exception report](interrupts.md) the