Built paging / the kernel's own page tables (with a TSS+IST)

This commit is contained in:
2026-07-03 12:23:55 +01:00
parent 0cc71ec8aa
commit 9cf135302d
10 changed files with 283 additions and 13 deletions
+9 -5
View File
@@ -19,10 +19,13 @@ rather than restate it. Roughly in the order things happen at runtime:
5. **[frame-allocator.md](frame-allocator.md) — the physical frame allocator.** The
bitmap allocator that hands out and reclaims 4 KiB physical frames from that
map — the primitive page tables and the heap will be built on.
6. **[interrupts.md](interrupts.md) — interrupts and exceptions.** The GDT and IDT,
the exception stubs, and the handler that reports a CPU fault in red instead of
letting it triple-fault into a silent reset.
7. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
6. **[interrupts.md](interrupts.md) — interrupts and exceptions.** The GDT, IDT and
TSS, the exception stubs, and the handler that reports a CPU fault in red instead
of letting it triple-fault into a silent reset.
7. **[paging.md](paging.md) — the kernel's page tables.** Building our own 4-level
page tables, identity-mapping the low 4 GiB, and switching CR3 off the firmware's
tables onto ours.
8. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
Cutting across all of these:
@@ -39,6 +42,7 @@ queries the **GOP** to pick a graphics mode ([gop.md](gop.md)), hands the kernel
map** of physical RAM ([memory-map.md](memory-map.md)); the kernel turns that map
into a **frame allocator** ([frame-allocator.md](frame-allocator.md)), installs
its **descriptor tables** so CPU faults are caught ([interrupts.md](interrupts.md)),
builds its own **page tables** and switches onto them ([paging.md](paging.md)),
runs — its CPU-specific bits behind the [arch](arch.md) boundary — and when it has
finished, or panics, it **halts** ([halting.md](halting.md)).
@@ -51,5 +55,5 @@ finished, or panics, it **halts** ([halting.md](halting.md)).
| Shared loader↔kernel contract (`BootInfo`, `Framebuffer`, `MemoryMap`, ABI) | `src/root.zig` |
| Physical frame allocator | `src/pmm.zig` |
| Framebuffer text console | `src/console.zig` |
| Arch-specific kernel code (`halt`, GDT/IDT, exception stubs, linker script) | `src/arch/x86_64/` |
| Arch-specific kernel code (`halt`, GDT/IDT/TSS, exception stubs, page tables, linker script) | `src/arch/x86_64/` |
| Build + `run-efi` (QEMU/OVMF) | `build.zig` |
+24 -3
View File
@@ -37,6 +37,25 @@ means present, ring 0, 64-bit interrupt gate. `src/arch/x86_64/idt.zig` builds t
table, points the first 32 vectors at their stubs, and loads it with `lidt`
(`idt_flush`).
## The TSS and the double-fault stack
There's one more table, the **Task State Segment**. In long mode its main
remaining job is the **Interrupt Stack Table (IST)**: an IDT gate can name an IST
slot, and when that vector fires the CPU switches to the stack recorded there —
*regardless* of what the interrupted stack looked like.
This matters most for the **double fault** (#DF, vector 8). A #DF means the CPU
hit a fault *while trying to deliver another fault* — very often because the
current stack pointer is bad, so pushing the exception frame itself faulted. If
the #DF handler then tried to push onto that same bad stack, it would fault a
third time and **triple-fault** — an instant reset. So the #DF gate is pointed at
**IST1**, a small dedicated stack (`src/arch/x86_64/tss.zig`) that's always valid.
Bringing it up: fill in the TSS's IST1 pointer, publish the TSS through a
descriptor in the GDT (`gdt.setTss`), and load it into the task register with
`ltr`. The TSS descriptor is a 16-byte system descriptor spanning two GDT slots,
which is why the GDT grew from three entries to five.
## The stubs and the trap frame
On an exception the CPU pushes a small frame (SS, RSP, RFLAGS, CS, RIP) and, for
@@ -87,11 +106,13 @@ together confirm the whole path: the GDT is active (we're still executing), the
IDT vectored to the right stub, the stub built a correct `CpuState`, and the Zig
handler read it and reported instead of triple-faulting.
Separately, pointing RSP at an unmapped address and faulting forced a **double
fault** — reported cleanly (`double fault (vector 8)`) rather than triple-faulting
into a reset, which only works because #DF ran on IST1. That's the proof the
TSS/IST is wired up: the handler survived a completely broken stack.
## What's next (not done here)
- **A TSS with an IST** (interrupt stack table) so the double-fault handler runs
on a known-good stack — important because a double fault often means the current
stack is unusable, and without an IST the handler would itself fault.
- **Device interrupts**: program the local APIC and IO-APIC, wire a timer and the
keyboard onto vectors ≥ 32, and (unlike exceptions) actually *return* from them
with `iretq` — which `isr_common` already does.
+78
View File
@@ -0,0 +1,78 @@
# Paging: the kernel's own page tables
Every memory access the CPU makes goes through the **page tables**: hardware walks
them to translate a virtual address into a physical one, and faults if there's no
valid mapping. Up to now danos ran on the *firmware's* page tables — which live in
memory we'd like to reclaim and which we don't control. This step builds the
kernel's own tables and switches onto them.
It's x86_64-specific (the 4-level table format is an Intel/AMD thing), so it lives
behind the [arch](arch.md) boundary in `src/arch/x86_64/paging.zig`.
## What we map, and why identity
x86_64 uses **4 levels**: PML4 → PDPT → PD → PT, each a 512-entry table, with 9
bits of the virtual address indexing each level. A leaf can be a 4 KiB page (at
the PT level) or, with the "huge" bit, a 2 MiB page straight from the PD.
For this first set of tables we **identity-map the low 4 GiB** — virtual address
equals physical address. Identity mapping is the pragmatic bootstrap: it keeps
everything that's already running valid across the CR3 switch without relocating a
single thing. The kernel image (at 1 MiB), the stack, the frame allocator's
bitmap, the framebuffer (at 2 GiB), and device MMIO all sit in the low 4 GiB, so
one flat identity range covers the lot.
Using 2 MiB pages makes the whole map tiny: 4 GiB ÷ 2 MiB = 2048 leaves, which is
one PML4, one PDPT, and four page directories — six frames total, from the
[frame allocator](frame-allocator.md).
## Building and switching
`paging.init` takes a frame allocator and:
1. Allocates and zeroes a PML4. **Zeroing matters**: the frame is recycled memory,
and any stale non-zero entry would map a bogus region.
2. Walks PML4 → PDPT → PD for each 2 MiB address, creating intermediate tables on
demand (`descend`) and writing the leaf with present + writable + huge.
3. Loads the PML4's physical address into **CR3**, which both switches address
spaces and flushes the TLB in one instruction.
This all works because the firmware's identity map is still active *while we
build*, so a freshly allocated frame's physical address is usable directly as a
pointer. After the CR3 load we're on our own tables — and since those page-table
frames came from low RAM, they remain mapped (and thus editable) for later.
## Verifying it
Two temporary tests confirmed both that we switched tables and that faults are
caught with the right detail:
- **Touch an address above 4 GiB** (`0xdeadbeef000`). Under the firmware's tables
this *didn't* fault (they map a huge range); under ours it does:
```
CPU EXCEPTION: page fault (vector 14)
error code : 0x2 (write, page not present)
CR2 (addr) : 0x00000deadbeef000 (the faulting address)
```
A #PF at exactly the unmapped address is proof we're on our own, deliberately
smaller, map — and that CR2 reporting works (see [interrupts.md](interrupts.md)).
- The console keeps working *after* the switch, confirming the framebuffer, kernel
code, and stack are all still mapped.
> Toolchain note: a volatile store to a *compile-time-constant* address tripped a
> codegen bug in the Zig self-hosted x86_64 backend ("no encoding for mov moffs").
> Computing the address in a runtime variable sidesteps it (register-relative
> store). Worth remembering when poking fixed MMIO addresses.
## What's next (not done here)
- **A higher-half kernel**: map the kernel at a high virtual base (e.g.
`0xffffffff80000000`) so user address space can own the low half later.
- **Real permissions**: today every page is writable and executable. Map code
read-execute, data read-write + no-execute (needs setting EFER.NXE and the NX
bit), and leave a guard page unmapped to catch null-ish dereferences.
- **4 KiB pages / a general `map(virt, phys, flags)`** for fine-grained mappings,
and unmapping with TLB invalidation (`invlpg`).
- **Per-address-space tables** once there are user processes.