Built device interrupts

This commit is contained in:
2026-07-03 12:57:40 +01:00
parent c8e89e8115
commit 5ea521d054
12 changed files with 367 additions and 23 deletions
+6 -2
View File
@@ -25,7 +25,10 @@ rather than restate it. Roughly in the order things happen at runtime:
7. **[paging.md](paging.md) — the kernel's page tables.** Building our own 4-level
page tables, identity-mapping the low 4 GiB, and switching CR3 off the firmware's
tables onto ours.
8. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
8. **[device-interrupts.md](device-interrupts.md) — device interrupts.** The Local
APIC and its timer — the kernel's first interrupt that is *handled and returned
from*, giving it a heartbeat.
9. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
Cutting across all of these:
@@ -46,6 +49,7 @@ map** of physical RAM ([memory-map.md](memory-map.md)); the kernel turns that ma
into a **frame allocator** ([frame-allocator.md](frame-allocator.md)), installs
its **descriptor tables** so CPU faults are caught ([interrupts.md](interrupts.md)),
builds its own **page tables** and switches onto them ([paging.md](paging.md)),
starts the **timer** so it has a heartbeat ([device-interrupts.md](device-interrupts.md)),
runs — its CPU-specific bits behind the [arch](arch.md) boundary — and when it has
finished, or panics, it **halts** ([halting.md](halting.md)).
@@ -59,6 +63,6 @@ finished, or panics, it **halts** ([halting.md](halting.md)).
| Physical frame allocator | `src/pmm.zig` |
| Framebuffer text console (mirrors to serial) | `src/console.zig` |
| In-kernel test cases | `src/tests.zig` |
| Arch-specific kernel code (`halt`, GDT/IDT/TSS, exception stubs, page tables, serial, linker script) | `src/arch/x86_64/` |
| Arch-specific kernel code (`halt`, GDT/IDT/TSS, exception + interrupt stubs, page tables, APIC/timer, serial, linker script) | `src/arch/x86_64/` |
| Build + `run-efi` (QEMU/OVMF) | `build.zig` |
| QEMU integration test harness | `test/qemu_test.py` |
+5 -2
View File
@@ -69,8 +69,11 @@ There are really two independent questions, and it's worth not conflating them:
TSS plus CPU-exception handling (see [interrupts.md](interrupts.md)).
- **`src/arch/x86_64/paging.zig`** — the kernel's page tables (see
[paging.md](paging.md)).
- **`src/arch/x86_64/serial.zig`** — the COM1 UART, the kernel's machine-readable
log channel (see [testing.md](testing.md)).
- **`src/arch/x86_64/apic.zig`** — the Local APIC and its timer, the source of
device interrupts (see [device-interrupts.md](device-interrupts.md)).
- **`src/arch/x86_64/serial.zig`** / **`io.zig`** — the COM1 UART (the kernel's
machine-readable log channel, see [testing.md](testing.md)) and the shared
port-I/O + MSR primitives.
- **`src/arch/x86_64/isr.s`** — the exception stubs and the `lgdt`/`lidt`/`ltr`
load helpers, in real assembly because Zig inline asm can't express them.
- **`src/arch/x86_64/linker.ld`** — the kernel link layout (fixed low load
+109
View File
@@ -0,0 +1,109 @@
# Device interrupts
CPU exceptions ([interrupts.md](interrupts.md)) are the kernel reacting to its own
mistakes. **Device interrupts** are the opposite: hardware asking for attention —
a timer firing, a key pressed, a packet arriving. They share the IDT, but differ
in one fundamental way: an exception here is terminal (we report and halt), while a
device interrupt is *handled and returned from*, so the interrupted code resumes as
if nothing happened. This is danos's first code that takes an interrupt and comes
back — the same mechanism a scheduler will later use to preempt tasks.
The first device we bring up is the **timer**, because it's the simplest: it lives
entirely on the CPU's local interrupt controller, needing no external routing.
It's all x86_64-specific, behind the [arch](arch.md) boundary.
## The APIC, not the PIC
Interrupt delivery on modern x86 goes through the **APIC**, not the legacy 8259
PIC. There are two halves; we only need one so far:
- The **Local APIC** (per-CPU, memory-mapped at physical `0xFEE00000`) handles the
CPU's own timer and receives interrupts routed to it. `src/arch/x86_64/apic.zig`.
- The **IO-APIC** routes *external* device lines (keyboard, etc.) to LAPIC vectors.
Not needed for the timer — it'll arrive with the keyboard.
The old PIC has to be dealt with first, though: left alone it would deliver
interrupts on vectors `0x08-0x0F`, which **collide with the CPU exception
vectors** — a spurious IRQ would look like a double fault. So `init` remaps the
PIC's vectors to `0x20-0x2F` and masks every line, taking it out of the picture.
Then the LAPIC is enabled in two places: the `IA32_APIC_BASE` MSR's global-enable
bit, and the LAPIC's own spurious-vector register (bit 8 = software enable). The
spurious vector is `0x2F` — low nibble `F` by convention, and inside our gate
range so a stray spurious interrupt lands on a valid no-op.
## The timer
The LAPIC timer is three register writes (`initTimer`): a divide setting, then the
LVT-timer entry giving it a **vector** (32) and **periodic** mode, then an initial
count that becomes the reload value. From then on it fires vector 32 repeatedly, on
its own, forever.
> The count isn't calibrated to real time yet — the tick *rate* is arbitrary
> (bus-clock dependent). Turning it into a known frequency (say 100 Hz) needs a
> reference clock to measure against (the PIT, HPET, or the TSC). That's a later
> step; for now it just needs to tick.
## Two kinds of vector, one dispatch
The IDT now installs gates `0-47`: the 32 exceptions plus the device range. Every
gate still funnels through the same stub tail (`isr_common`), which calls one
dispatcher that branches on the vector (`interruptDispatch` in `idt.zig`):
```zig
if (state.vector < 32) {
on_fault(state); // exception: report and halt (never returns)
} else if (handlers[state.vector]) |handler| {
handler(); // device: run the registered handler
apic.eoi(); // ...acknowledge the LAPIC
}
// else: spurious/unhandled — deliberately no EOI
```
Two things make device interrupts *return* where exceptions don't:
1. **The handler returns.** The timer handler just bumps a tick counter. Control
flows back to `isr_common`, which restores every register it saved and executes
`iretq` — resuming the interrupted instruction exactly. (This is why the stub
saves *all* the general registers.)
2. **End-of-interrupt.** After handling, we write the LAPIC's EOI register. Miss
this and the LAPIC thinks we're still busy and never delivers the next
interrupt. It's the single most common "my timer fired once and stopped" bug.
A device handler is a plain `fn () void` — a timer or keyboard handler doesn't need
the interrupted registers. (Note: the stubs don't save the SSE/vector registers, so
a handler must not use them; ours don't.)
## Turning them on
Exceptions can't be masked, which is why they worked all along. Maskable device
interrupts don't fire until the CPU's interrupt flag is set — so the final step is
`sti` (`arch.enableInterrupts()`), after the APIC and timer are configured. From
that instant the kernel has a heartbeat, and its idle `hlt` loop
([halting.md](halting.md)) wakes on every tick and dozes off again.
## Verifying it
The `timer` test (see [testing.md](testing.md)) is the proof that an interrupt both
*fires* and *returns*: it records the tick count, busy-waits, and checks the count
advanced on its own.
```
$ python3 test/qemu_test.py timer
timer ... PASS (matched 'DANOS-TEST-RESULT: PASS')
```
If the APIC weren't enabled, or `sti` were missing, or EOI were forgotten, the
count would stay put and the test would fail. That it advances — while the CPU was
spinning in unrelated code — is the whole mechanism working end to end.
## What's next (not done here)
- **The keyboard**: bring up the IO-APIC, route its IRQ to a vector, and read
scancodes from the PS/2 controller — the first *input* device.
- **A calibrated timer** at a known frequency, and a monotonic clock.
- **Uncacheable MMIO**: the LAPIC page is currently mapped writeback-cacheable like
the rest of the identity map. QEMU tolerates it, but real hardware wants MMIO
marked uncacheable (via the page's cache bits or an MTRR).
- **Preemption**: once there are tasks, the timer handler is where the scheduler
decides to switch — the reason a *returning* interrupt matters.
+6 -6
View File
@@ -113,12 +113,12 @@ TSS/IST is wired up: the handler survived a completely broken stack.
## What's next (not done here)
- **Device interrupts**: program the local APIC and IO-APIC, wire a timer and the
keyboard onto vectors ≥ 32, and (unlike exceptions) actually *return* from them
with `iretq` — which `isr_common` already does.
- **The IO-APIC and the keyboard**: the timer (a local-APIC device interrupt) is
covered in [device-interrupts.md](device-interrupts.md); external devices like
the keyboard also need the IO-APIC to route their lines onto vectors.
- **SSE state**: the stubs save general registers but not the vector registers, so
recoverable interrupts that return to SSE-using code will need that added. Fine
for now, since exceptions here don't return.
a returning interrupt whose handler uses SSE will need that added. Fine for now,
since our handlers don't.
With faults now debuggable, the paging work that comes next — where a wrong
With faults now debuggable, the paging work that follows — where a wrong
page-table entry means an instant #PF — is far less painful.