Built GDT + IDT + exception handlers

This commit is contained in:
2026-07-03 12:05:41 +01:00
parent 6312e84262
commit 0cc71ec8aa
9 changed files with 458 additions and 23 deletions
+9 -5
View File
@@ -19,7 +19,10 @@ rather than restate it. Roughly in the order things happen at runtime:
5. **[frame-allocator.md](frame-allocator.md) — the physical frame allocator.** The
bitmap allocator that hands out and reclaims 4 KiB physical frames from that
map — the primitive page tables and the heap will be built on.
6. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
6. **[interrupts.md](interrupts.md) — interrupts and exceptions.** The GDT and IDT,
the exception stubs, and the handler that reports a CPU fault in red instead of
letting it triple-fault into a silent reset.
7. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
Cutting across all of these:
@@ -34,9 +37,10 @@ The boot flow ties them together: UEFI runs the loader ([efi.md](efi.md)), which
queries the **GOP** to pick a graphics mode ([gop.md](gop.md)), hands the kernel a
**framebuffer** to draw into ([framebuffer.md](framebuffer.md)) and a **memory
map** of physical RAM ([memory-map.md](memory-map.md)); the kernel turns that map
into a **frame allocator** ([frame-allocator.md](frame-allocator.md)), runs — its
CPU-specific bits behind the [arch](arch.md) boundary — and when it has finished,
or panics, it **halts** ([halting.md](halting.md)).
into a **frame allocator** ([frame-allocator.md](frame-allocator.md)), installs
its **descriptor tables** so CPU faults are caught ([interrupts.md](interrupts.md)),
runs — its CPU-specific bits behind the [arch](arch.md) boundary — and when it has
finished, or panics, it **halts** ([halting.md](halting.md)).
## Source map
@@ -47,5 +51,5 @@ or panics, it **halts** ([halting.md](halting.md)).
| Shared loader↔kernel contract (`BootInfo`, `Framebuffer`, `MemoryMap`, ABI) | `src/root.zig` |
| Physical frame allocator | `src/pmm.zig` |
| Framebuffer text console | `src/console.zig` |
| Arch-specific kernel code (`halt`, linker script) | `src/arch/x86_64/` |
| Arch-specific kernel code (`halt`, GDT/IDT, exception stubs, linker script) | `src/arch/x86_64/` |
| Build + `run-efi` (QEMU/OVMF) | `build.zig` |
+7 -2
View File
@@ -62,8 +62,13 @@ There are really two independent questions, and it's worth not conflating them:
## Current x86_64 contents
- **`src/arch/x86_64/cpu.zig`** — the `arch` module root. Exposes `halt()` (see
[halting.md](halting.md)); GDT, IDT and paging will join it here as the kernel
grows.
[halting.md](halting.md)), `init()` (bring up the descriptor tables),
`setFaultHandler`, `readCr2`, and the `CpuState` trap frame. Paging will join it
here as the kernel grows.
- **`src/arch/x86_64/gdt.zig`** / **`idt.zig`** — the GDT and IDT plus CPU-exception
handling (see [interrupts.md](interrupts.md)).
- **`src/arch/x86_64/isr.s`** — the exception stubs and the `lgdt`/`lidt` load
helpers, in real assembly because Zig inline asm can't express them.
- **`src/arch/x86_64/linker.ld`** — the kernel link layout (fixed low load
address, one PT_LOAD per permission set).
+103
View File
@@ -0,0 +1,103 @@
# Interrupts and exceptions
When something goes wrong on the CPU — a bad pointer, a divide by zero, a
malformed page table — the processor raises an **exception**. If nothing is set
up to catch it, the fault escalates: the CPU tries to invoke a handler, finds
none, faults again trying to handle *that*, and on the third strike triple-faults,
which on real hardware and in QEMU means a silent reset. Debugging by spontaneous
reboot is miserable.
This is the machinery that catches those faults and prints what happened instead.
It's all x86_64-specific, so it lives behind the [arch](arch.md) boundary in
`src/arch/x86_64/`. Only the 32 CPU-defined exception vectors are wired up so far;
device interrupts (timer, keyboard, via the APIC) come later, on the same IDT.
## First the GDT
In 64-bit long mode, segmentation is mostly switched off — but the CPU still
requires valid **segment descriptors** for code and data, and, crucially, every
IDT gate names a code-segment *selector* that must resolve in the current GDT. The
firmware left a GDT in place, but we don't control it, so we install our own with
known selectors: `0x08` kernel code, `0x10` kernel data.
`src/arch/x86_64/gdt.zig` holds three flat descriptors — a required null entry,
plus code and data — where the only bits that matter in long mode are the access
byte and the code segment's long-mode (`L`) flag. Loading it (`gdt_flush` in
`isr.s`) does two things: `lgdt`, then reload the segment registers. The data
registers take a plain `mov`, but **CS can't** — so we reload it with a far
return, pushing the new selector and a return address and letting `lretq` pop them
into CS:RIP.
## Then the IDT
The **Interrupt Descriptor Table** maps each of 256 vectors to a handler. Each
entry is a 16-byte *gate* holding the handler's address (split across three
fields, a quirk of the format), the code selector (`0x08`), and flags: `0x8E`
means present, ring 0, 64-bit interrupt gate. `src/arch/x86_64/idt.zig` builds the
table, points the first 32 vectors at their stubs, and loads it with `lidt`
(`idt_flush`).
## The stubs and the trap frame
On an exception the CPU pushes a small frame (SS, RSP, RFLAGS, CS, RIP) and, for
*some* vectors, an **error code**. That inconsistency is a nuisance, so each stub
in `src/arch/x86_64/isr.s` normalises it: vectors that don't get a hardware error
code push a dummy `0`, then every stub pushes its **vector number** and jumps to a
shared tail, `isr_common`. The tail pushes all the general registers and calls the
Zig handler with a pointer to the whole thing.
The result on the stack is a uniform **`CpuState`** — register block, then vector
and error code, then the CPU's frame. Its field order in `idt.zig` is exactly the
push order in `isr.s`; the two must stay in sync.
### Why a separate `.s` file
The stubs and table-loads are real assembly rather than Zig inline asm because
they need things inline asm on this toolchain can't express: cross-symbol
`jmp`/`call` (a stub jumping to `isr_common`, which calls the exported
`exceptionHandler`), and the `lgdt`/`lidt` memory operands (which LLVM rejects
inline). `build.zig` adds `isr.s` to the arch module.
## Reporting a fault
`isr_common` calls `exceptionHandler`, which forwards to a swappable `on_fault`
hook. The generic kernel installs a reporter (`onException` in `main.zig`) that
prints, in red, the exception name and vector, the error code, the faulting RIP
and RSP, and — for a page fault (#PF, vector 14) — the faulting address from
**CR2**. Then it halts. There's no fault *recovery* yet, so every exception is
terminal; the point is that it's now **visible** instead of a silent reset.
The hook is set before `arch.init()` in `kmain`, so a fault during setup is still
caught.
## Verifying it
A temporary `ud2` (unconditional invalid-opcode instruction) in `kmain` produced,
in red:
```
CPU EXCEPTION: invalid opcode (vector 6)
error code : 0x0
RIP : 0x000000000010dadd <- the ud2, in the kernel image at 0x100000+
RSP : 0x0000000007e8aed0
```
Vector 6 with no error code, a RIP inside the loaded kernel, and a sane RSP
together confirm the whole path: the GDT is active (we're still executing), the
IDT vectored to the right stub, the stub built a correct `CpuState`, and the Zig
handler read it and reported instead of triple-faulting.
## What's next (not done here)
- **A TSS with an IST** (interrupt stack table) so the double-fault handler runs
on a known-good stack — important because a double fault often means the current
stack is unusable, and without an IST the handler would itself fault.
- **Device interrupts**: program the local APIC and IO-APIC, wire a timer and the
keyboard onto vectors ≥ 32, and (unlike exceptions) actually *return* from them
with `iretq` — which `isr_common` already does.
- **SSE state**: the stubs save general registers but not the vector registers, so
recoverable interrupts that return to SSE-using code will need that added. Fine
for now, since exceptions here don't return.
With faults now debuggable, the paging work that comes next — where a wrong
page-table entry means an instant #PF — is far less painful.