diff --git a/docs/README.md b/docs/README.md index b3339c1..b488e15 100644 --- a/docs/README.md +++ b/docs/README.md @@ -25,7 +25,10 @@ rather than restate it. Roughly in the order things happen at runtime: 7. **[paging.md](paging.md) — the kernel's page tables.** Building our own 4-level page tables, identity-mapping the low 4 GiB, and switching CR3 off the firmware's tables onto ours. -8. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and +8. **[device-interrupts.md](device-interrupts.md) — device interrupts.** The Local + APIC and its timer — the kernel's first interrupt that is *handled and returned + from*, giving it a heartbeat. +9. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and how `while (true) hlt` parks the CPU safely once there's nothing left to do. Cutting across all of these: @@ -46,6 +49,7 @@ map** of physical RAM ([memory-map.md](memory-map.md)); the kernel turns that ma into a **frame allocator** ([frame-allocator.md](frame-allocator.md)), installs its **descriptor tables** so CPU faults are caught ([interrupts.md](interrupts.md)), builds its own **page tables** and switches onto them ([paging.md](paging.md)), +starts the **timer** so it has a heartbeat ([device-interrupts.md](device-interrupts.md)), runs — its CPU-specific bits behind the [arch](arch.md) boundary — and when it has finished, or panics, it **halts** ([halting.md](halting.md)). @@ -59,6 +63,6 @@ finished, or panics, it **halts** ([halting.md](halting.md)). | Physical frame allocator | `src/pmm.zig` | | Framebuffer text console (mirrors to serial) | `src/console.zig` | | In-kernel test cases | `src/tests.zig` | -| Arch-specific kernel code (`halt`, GDT/IDT/TSS, exception stubs, page tables, serial, linker script) | `src/arch/x86_64/` | +| Arch-specific kernel code (`halt`, GDT/IDT/TSS, exception + interrupt stubs, page tables, APIC/timer, serial, linker script) | `src/arch/x86_64/` | | Build + `run-efi` (QEMU/OVMF) | `build.zig` | | QEMU integration test harness | `test/qemu_test.py` | diff --git a/docs/arch.md b/docs/arch.md index 86b38a8..8e0d8d9 100644 --- a/docs/arch.md +++ b/docs/arch.md @@ -69,8 +69,11 @@ There are really two independent questions, and it's worth not conflating them: TSS plus CPU-exception handling (see [interrupts.md](interrupts.md)). - **`src/arch/x86_64/paging.zig`** — the kernel's page tables (see [paging.md](paging.md)). -- **`src/arch/x86_64/serial.zig`** — the COM1 UART, the kernel's machine-readable - log channel (see [testing.md](testing.md)). +- **`src/arch/x86_64/apic.zig`** — the Local APIC and its timer, the source of + device interrupts (see [device-interrupts.md](device-interrupts.md)). +- **`src/arch/x86_64/serial.zig`** / **`io.zig`** — the COM1 UART (the kernel's + machine-readable log channel, see [testing.md](testing.md)) and the shared + port-I/O + MSR primitives. - **`src/arch/x86_64/isr.s`** — the exception stubs and the `lgdt`/`lidt`/`ltr` load helpers, in real assembly because Zig inline asm can't express them. - **`src/arch/x86_64/linker.ld`** — the kernel link layout (fixed low load diff --git a/docs/device-interrupts.md b/docs/device-interrupts.md new file mode 100644 index 0000000..4fb4f08 --- /dev/null +++ b/docs/device-interrupts.md @@ -0,0 +1,109 @@ +# Device interrupts + +CPU exceptions ([interrupts.md](interrupts.md)) are the kernel reacting to its own +mistakes. **Device interrupts** are the opposite: hardware asking for attention — +a timer firing, a key pressed, a packet arriving. They share the IDT, but differ +in one fundamental way: an exception here is terminal (we report and halt), while a +device interrupt is *handled and returned from*, so the interrupted code resumes as +if nothing happened. This is danos's first code that takes an interrupt and comes +back — the same mechanism a scheduler will later use to preempt tasks. + +The first device we bring up is the **timer**, because it's the simplest: it lives +entirely on the CPU's local interrupt controller, needing no external routing. +It's all x86_64-specific, behind the [arch](arch.md) boundary. + +## The APIC, not the PIC + +Interrupt delivery on modern x86 goes through the **APIC**, not the legacy 8259 +PIC. There are two halves; we only need one so far: + +- The **Local APIC** (per-CPU, memory-mapped at physical `0xFEE00000`) handles the + CPU's own timer and receives interrupts routed to it. `src/arch/x86_64/apic.zig`. +- The **IO-APIC** routes *external* device lines (keyboard, etc.) to LAPIC vectors. + Not needed for the timer — it'll arrive with the keyboard. + +The old PIC has to be dealt with first, though: left alone it would deliver +interrupts on vectors `0x08-0x0F`, which **collide with the CPU exception +vectors** — a spurious IRQ would look like a double fault. So `init` remaps the +PIC's vectors to `0x20-0x2F` and masks every line, taking it out of the picture. + +Then the LAPIC is enabled in two places: the `IA32_APIC_BASE` MSR's global-enable +bit, and the LAPIC's own spurious-vector register (bit 8 = software enable). The +spurious vector is `0x2F` — low nibble `F` by convention, and inside our gate +range so a stray spurious interrupt lands on a valid no-op. + +## The timer + +The LAPIC timer is three register writes (`initTimer`): a divide setting, then the +LVT-timer entry giving it a **vector** (32) and **periodic** mode, then an initial +count that becomes the reload value. From then on it fires vector 32 repeatedly, on +its own, forever. + +> The count isn't calibrated to real time yet — the tick *rate* is arbitrary +> (bus-clock dependent). Turning it into a known frequency (say 100 Hz) needs a +> reference clock to measure against (the PIT, HPET, or the TSC). That's a later +> step; for now it just needs to tick. + +## Two kinds of vector, one dispatch + +The IDT now installs gates `0-47`: the 32 exceptions plus the device range. Every +gate still funnels through the same stub tail (`isr_common`), which calls one +dispatcher that branches on the vector (`interruptDispatch` in `idt.zig`): + +```zig +if (state.vector < 32) { + on_fault(state); // exception: report and halt (never returns) +} else if (handlers[state.vector]) |handler| { + handler(); // device: run the registered handler + apic.eoi(); // ...acknowledge the LAPIC +} +// else: spurious/unhandled — deliberately no EOI +``` + +Two things make device interrupts *return* where exceptions don't: + +1. **The handler returns.** The timer handler just bumps a tick counter. Control + flows back to `isr_common`, which restores every register it saved and executes + `iretq` — resuming the interrupted instruction exactly. (This is why the stub + saves *all* the general registers.) +2. **End-of-interrupt.** After handling, we write the LAPIC's EOI register. Miss + this and the LAPIC thinks we're still busy and never delivers the next + interrupt. It's the single most common "my timer fired once and stopped" bug. + +A device handler is a plain `fn () void` — a timer or keyboard handler doesn't need +the interrupted registers. (Note: the stubs don't save the SSE/vector registers, so +a handler must not use them; ours don't.) + +## Turning them on + +Exceptions can't be masked, which is why they worked all along. Maskable device +interrupts don't fire until the CPU's interrupt flag is set — so the final step is +`sti` (`arch.enableInterrupts()`), after the APIC and timer are configured. From +that instant the kernel has a heartbeat, and its idle `hlt` loop +([halting.md](halting.md)) wakes on every tick and dozes off again. + +## Verifying it + +The `timer` test (see [testing.md](testing.md)) is the proof that an interrupt both +*fires* and *returns*: it records the tick count, busy-waits, and checks the count +advanced on its own. + +``` +$ python3 test/qemu_test.py timer + timer ... PASS (matched 'DANOS-TEST-RESULT: PASS') +``` + +If the APIC weren't enabled, or `sti` were missing, or EOI were forgotten, the +count would stay put and the test would fail. That it advances — while the CPU was +spinning in unrelated code — is the whole mechanism working end to end. + +## What's next (not done here) + +- **The keyboard**: bring up the IO-APIC, route its IRQ to a vector, and read + scancodes from the PS/2 controller — the first *input* device. +- **A calibrated timer** at a known frequency, and a monotonic clock. +- **Uncacheable MMIO**: the LAPIC page is currently mapped writeback-cacheable like + the rest of the identity map. QEMU tolerates it, but real hardware wants MMIO + marked uncacheable (via the page's cache bits or an MTRR). +- **Preemption**: once there are tasks, the timer handler is where the scheduler + decides to switch — the reason a *returning* interrupt matters. diff --git a/docs/interrupts.md b/docs/interrupts.md index f30db52..71be04a 100644 --- a/docs/interrupts.md +++ b/docs/interrupts.md @@ -113,12 +113,12 @@ TSS/IST is wired up: the handler survived a completely broken stack. ## What's next (not done here) -- **Device interrupts**: program the local APIC and IO-APIC, wire a timer and the - keyboard onto vectors ≥ 32, and (unlike exceptions) actually *return* from them - with `iretq` — which `isr_common` already does. +- **The IO-APIC and the keyboard**: the timer (a local-APIC device interrupt) is + covered in [device-interrupts.md](device-interrupts.md); external devices like + the keyboard also need the IO-APIC to route their lines onto vectors. - **SSE state**: the stubs save general registers but not the vector registers, so - recoverable interrupts that return to SSE-using code will need that added. Fine - for now, since exceptions here don't return. + a returning interrupt whose handler uses SSE will need that added. Fine for now, + since our handlers don't. -With faults now debuggable, the paging work that comes next — where a wrong +With faults now debuggable, the paging work that follows — where a wrong page-table entry means an instant #PF — is far less painful. diff --git a/src/arch/x86_64/apic.zig b/src/arch/x86_64/apic.zig new file mode 100644 index 0000000..30701cb --- /dev/null +++ b/src/arch/x86_64/apic.zig @@ -0,0 +1,91 @@ +//! Local APIC and its timer — the source of device interrupts. +//! +//! Modern x86 routes interrupts through the per-CPU Local APIC (the legacy 8259 +//! PIC is remapped out of the way and masked). The LAPIC also has a built-in +//! timer, which is the simplest device interrupt to bring up: it needs no +//! external routing, just a vector and a count. We use it as danos's heartbeat. +//! +//! The LAPIC is memory-mapped (default physical 0xFEE00000, inside our identity +//! map). Every interrupt must be acknowledged with an end-of-interrupt write, or +//! the LAPIC won't deliver the next one. + +const io = @import("io.zig"); + +/// IDT vector the timer fires on (in the device range, >= 32). +pub const timer_vector = 32; +/// Spurious-interrupt vector. Low nibble 0xF by convention; also in our gate +/// range so a stray spurious interrupt lands on a valid (no-op) handler. +const spurious_vector = 47; + +// LAPIC register offsets. +const reg_spurious = 0x0F0; +const reg_eoi = 0x0B0; +const reg_lvt_timer = 0x320; +const reg_timer_initial = 0x380; +const reg_timer_divide = 0x3E0; + +const ia32_apic_base_msr = 0x1B; + +/// LAPIC MMIO base. A runtime var (not a constant) both because we read it from +/// the MSR and so register writes compile to normal stores rather than a +/// `mov moffs`, which the self-hosted backend can't encode. +var base: usize = 0xFEE00000; + +var tick_count: u64 = 0; + +fn read(reg: u32) u32 { + return @as(*volatile u32, @ptrFromInt(base + reg)).*; +} +fn write(reg: u32, value: u32) void { + @as(*volatile u32, @ptrFromInt(base + reg)).* = value; +} + +/// Move the legacy 8259 PIC's vectors to 0x20-0x2F (clear of the CPU exception +/// vectors) and mask every line, so it can't deliver interrupts behind the APIC. +fn remapAndMaskPic() void { + io.outb(0x20, 0x11); // start init (cascade mode) + io.outb(0xA0, 0x11); + io.outb(0x21, 0x20); // master offset 0x20 + io.outb(0xA1, 0x28); // slave offset 0x28 + io.outb(0x21, 0x04); // tell master about slave on IRQ2 + io.outb(0xA1, 0x02); + io.outb(0x21, 0x01); // 8086 mode + io.outb(0xA1, 0x01); + io.outb(0x21, 0xFF); // mask all + io.outb(0xA1, 0xFF); +} + +/// Enable the Local APIC: mask the PIC, set the global-enable MSR bit, and +/// software-enable the APIC via its spurious-vector register. +pub fn init() void { + remapAndMaskPic(); + + const msr = io.rdmsr(ia32_apic_base_msr); + base = @intCast(msr & 0xFFFFF000); // physical base is bits 12+ + io.wrmsr(ia32_apic_base_msr, msr | (1 << 11)); // global enable + + write(reg_spurious, 0x100 | spurious_vector); // bit 8 = software enable +} + +/// Arm the LAPIC timer in periodic mode on `timer_vector`. +pub fn initTimer() void { + write(reg_timer_divide, 0x3); // divide bus clock by 16 + write(reg_lvt_timer, timer_vector | (1 << 17)); // periodic mode + write(reg_timer_initial, 1_000_000); // reload count -> periodic ticks +} + +/// Acknowledge the current interrupt so the LAPIC will deliver the next one. +pub fn eoi() void { + write(reg_eoi, 0); +} + +/// The timer interrupt handler: just count ticks for now. +pub fn timerTick() void { + tick_count +%= 1; +} + +/// Number of timer ticks so far. Volatile load: the count is bumped +/// asynchronously by the interrupt handler, so callers must re-read memory. +pub fn ticks() u64 { + return @as(*const volatile u64, &tick_count).*; +} diff --git a/src/arch/x86_64/cpu.zig b/src/arch/x86_64/cpu.zig index 996687b..fd891e5 100644 --- a/src/arch/x86_64/cpu.zig +++ b/src/arch/x86_64/cpu.zig @@ -9,6 +9,7 @@ const tss = @import("tss.zig"); const idt = @import("idt.zig"); const paging = @import("paging.zig"); const serial = @import("serial.zig"); +const apic = @import("apic.zig"); /// The saved register/trap frame passed to a fault handler. pub const CpuState = idt.CpuState; @@ -47,6 +48,29 @@ pub fn readCr3() u64 { ); } +/// Enable the Local APIC and start its periodic timer, the kernel's heartbeat. +/// Interrupts still have to be unmasked with enableInterrupts() to be delivered. +pub fn startTimer() void { + apic.init(); + idt.setHandler(apic.timer_vector, apic.timerTick); + apic.initTimer(); +} + +/// Number of timer ticks since startTimer(). +pub fn ticks() u64 { + return apic.ticks(); +} + +/// Unmask maskable interrupts (`sti`) so device interrupts get delivered. +pub fn enableInterrupts() void { + asm volatile ("sti"); +} + +/// Mask maskable interrupts (`cli`). +pub fn disableInterrupts() void { + asm volatile ("cli"); +} + /// Route CPU exceptions to `handler`, which receives the trap frame and does not /// return. Until set, faults just halt the core. pub fn setFaultHandler(handler: *const fn (*const CpuState) noreturn) void { diff --git a/src/arch/x86_64/idt.zig b/src/arch/x86_64/idt.zig index 8dbd171..954828c 100644 --- a/src/arch/x86_64/idt.zig +++ b/src/arch/x86_64/idt.zig @@ -1,13 +1,30 @@ -//! Interrupt Descriptor Table and the CPU-exception handlers. Without this, any -//! fault (a stray pointer, a bad page-table entry) triple-faults and silently -//! resets the machine. With it, the CPU vectors into our stubs, which capture the -//! register state and hand it to a reporter that prints what went wrong. +//! Interrupt Descriptor Table, CPU-exception handlers, and device-interrupt +//! dispatch. Without this, any fault (a stray pointer, a bad page-table entry) +//! triple-faults and silently resets the machine. With it, the CPU vectors into +//! our stubs, which capture the register state and hand it to a dispatcher. //! -//! Only the 32 architecture-defined exception vectors are wired up here; device -//! interrupts (the APIC, timer, keyboard) come later. +//! Vectors split in two: 0-31 are CPU exceptions (terminal — reported and +//! halted); 32+ are device interrupts (a registered handler runs, the APIC is +//! acknowledged, and we return to the interrupted code). const gdt = @import("gdt.zig"); const tss = @import("tss.zig"); +const apic = @import("apic.zig"); + +/// Highest vector we install a gate/stub for (exceptions 0-31 plus the device +/// range 32-47, which covers the timer and the spurious vector). +const gate_count = 48; + +/// A device-interrupt handler. It doesn't get the trap frame (a timer or keyboard +/// handler doesn't need the interrupted registers); add that if one ever does. +pub const Handler = *const fn () void; + +var handlers = [_]?Handler{null} ** 256; + +/// Register `handler` for a device-interrupt `vector` (>= 32). +pub fn setHandler(vector: usize, handler: Handler) void { + handlers[vector] = handler; +} /// The register + trap frame the ISR stubs build on the stack, laid out so the /// lowest address (where RSP points when we call the handler) is the first field. @@ -101,9 +118,10 @@ fn setGate(vector: usize, handler: u64) void { }; } -/// Point the first 32 vectors at the stubs defined in isr.s and load the IDT. +/// Point every installed vector at its stub (isr.s) and load the IDT. pub fn init() void { - inline for (0..32) |vector| { + @setEvalBranchQuota(20000); // comptimePrint across all the gates adds up + inline for (0..gate_count) |vector| { const stub = @extern(*const anyopaque, .{ .name = std.fmt.comptimePrint("isr{d}", .{vector}) }); setGate(vector, @intFromPtr(stub)); } @@ -118,9 +136,16 @@ pub fn init() void { } /// Called by isr_common (isr.s) with a pointer to the trap frame. Exported so the -/// assembly stubs can `call` it by name. -export fn exceptionHandler(state: *const CpuState) callconv(.c) void { - on_fault(state); +/// assembly stubs can `call` it by name. Exceptions are terminal; device +/// interrupts run their handler, get acknowledged, and return. +export fn interruptDispatch(state: *const CpuState) callconv(.c) void { + if (state.vector < 32) { + on_fault(state); // CPU exception — never returns + } else if (handlers[state.vector]) |handler| { + handler(); + apic.eoi(); + } + // else: spurious/unhandled device interrupt — don't acknowledge it } const std = @import("std"); diff --git a/src/arch/x86_64/io.zig b/src/arch/x86_64/io.zig new file mode 100644 index 0000000..2951a5c --- /dev/null +++ b/src/arch/x86_64/io.zig @@ -0,0 +1,38 @@ +//! x86 port I/O and model-specific registers — the low-level primitives the +//! serial port and the APIC talk to hardware through. + +pub fn outb(port: u16, value: u8) void { + asm volatile ("outb %[value], %[port]" + : + : [value] "{al}" (value), + [port] "{dx}" (port), + ); +} + +pub fn inb(port: u16) u8 { + return asm volatile ("inb %[port], %[value]" + : [value] "={al}" (-> u8), + : [port] "{dx}" (port), + ); +} + +/// Read a model-specific register (returns edx:eax combined). +pub fn rdmsr(msr: u32) u64 { + var low: u32 = undefined; + var high: u32 = undefined; + asm volatile ("rdmsr" + : [low] "={eax}" (low), + [high] "={edx}" (high), + : [msr] "{ecx}" (msr), + ); + return (@as(u64, high) << 32) | low; +} + +pub fn wrmsr(msr: u32, value: u64) void { + asm volatile ("wrmsr" + : + : [msr] "{ecx}" (msr), + [low] "{eax}" (@as(u32, @truncate(value))), + [high] "{edx}" (@as(u32, @truncate(value >> 32))), + ); +} diff --git a/src/arch/x86_64/isr.s b/src/arch/x86_64/isr.s index 909d4d8..0167a2b 100644 --- a/src/arch/x86_64/isr.s +++ b/src/arch/x86_64/isr.s @@ -89,7 +89,26 @@ STUB_NOERR 29 STUB_NOERR 30 STUB_NOERR 31 -.extern exceptionHandler +# Device-interrupt vectors (timer, spurious, room for more). None push an error +# code, so they all use the dummy-zero form. +STUB_NOERR 32 +STUB_NOERR 33 +STUB_NOERR 34 +STUB_NOERR 35 +STUB_NOERR 36 +STUB_NOERR 37 +STUB_NOERR 38 +STUB_NOERR 39 +STUB_NOERR 40 +STUB_NOERR 41 +STUB_NOERR 42 +STUB_NOERR 43 +STUB_NOERR 44 +STUB_NOERR 45 +STUB_NOERR 46 +STUB_NOERR 47 + +.extern interruptDispatch # Shared tail. Register push order here defines the CpuState field order. isr_common: @@ -109,7 +128,7 @@ isr_common: push %r14 push %r15 mov %rsp, %rdi # first argument: pointer to the trap frame - call exceptionHandler + call interruptDispatch pop %r15 pop %r14 pop %r13 diff --git a/src/main.zig b/src/main.zig index fe7ba74..5dba7bc 100644 --- a/src/main.zig +++ b/src/main.zig @@ -91,6 +91,11 @@ fn kmain(boot_info: *const BootInfo) noreturn { con.print("\ndanos: paging enabled\n", .{}); con.print(" page tables: CR3 = 0x{x:0>16}\n", .{arch.readCr3()}); + // Start the timer and unmask interrupts — the kernel now has a heartbeat. + arch.startTimer(); + arch.enableInterrupts(); + con.write("\ndanos: timer interrupts enabled\n"); + // In a test build (`zig build -Dtest-case=`), run that case and stop. // Normal builds fall through to the idle halt. if (build_options.test_case) |case| { diff --git a/src/tests.zig b/src/tests.zig index 998282b..83fd77a 100644 --- a/src/tests.zig +++ b/src/tests.zig @@ -36,6 +36,8 @@ fn check(name: []const u8, ok: bool) void { pub fn run(case: []const u8, boot_info: *const BootInfo) void { if (eql(case, "smoke")) { smoke(boot_info); + } else if (eql(case, "timer")) { + timer(); } else if (eql(case, "fault-ud")) { faultInvalidOpcode(); } else if (eql(case, "fault-pf")) { @@ -91,6 +93,26 @@ fn smoke(boot_info: *const BootInfo) void { log("DANOS-TEST-DONE\n", .{}); } +/// Verify device interrupts fire and return: the timer tick counter must advance +/// on its own. Interrupts are already enabled by kmain before tests run. +fn timer() void { + log("DANOS-TEST-BEGIN: timer\n", .{}); + const start = arch.ticks(); + // Busy-wait for the counter to advance. arch.ticks() is a volatile load, so + // the compiler re-reads it each iteration and sees the interrupt's update. + // The cap is only a safety net; the harness timeout is the real backstop. + var spins: u64 = 0; + while (arch.ticks() == start and spins < 5_000_000_000) spins +%= 1; + check("timer interrupts advance the tick count", arch.ticks() > start); + + log("DANOS-TEST-RESULT: {s} ({d} passed, {d} failed)\n", .{ + if (failed == 0) "PASS" else "FAIL", + passed, + failed, + }); + log("DANOS-TEST-DONE\n", .{}); +} + fn faultInvalidOpcode() void { log("DANOS-TEST-BEGIN: fault-ud\n", .{}); asm volatile ("ud2"); @@ -108,6 +130,7 @@ fn faultPageFault() void { fn faultDoubleFault() void { log("DANOS-TEST-BEGIN: fault-df\n", .{}); + arch.disableInterrupts(); // so only the ud2 delivery (not a timer tick) triggers the #DF // Point RSP at unmapped memory, then fault: the CPU can't push the fault // frame, which escalates to #DF — survivable only because #DF runs on IST1. var bad_sp: u64 = 0x5000000000; diff --git a/test/qemu_test.py b/test/qemu_test.py index 5656e62..1fec8cb 100644 --- a/test/qemu_test.py +++ b/test/qemu_test.py @@ -63,6 +63,9 @@ CASES = [ {"name": "smoke", "expect": r"DANOS-TEST-RESULT: PASS", "fail": r"DANOS-TEST-RESULT: FAIL"}, + {"name": "timer", + "expect": r"DANOS-TEST-RESULT: PASS", + "fail": r"DANOS-TEST-RESULT: FAIL"}, {"name": "fault-ud", "expect": r"invalid opcode \(vector 6\)"}, {"name": "fault-pf", "expect": r"page fault \(vector 14\)"}, {"name": "fault-df", "expect": r"double fault \(vector 8\)"},