Re-organize the source tree as a monorepo mirroring the FHS

The source layout now mirrors the runtime filesystem hierarchy
(docs/danos-file-system-hierarchy-FSH.md): what lives under system/ in the
source is what a running danos represents under /system. Each service and
driver is a sub-project directory that is its own Zig module — cross-project
references go by module name, never by a path into another project's files.

Moves (all git mv, history preserved):
- src/            -> system/            (danos internals; the self-representation)
    root.zig      -> danos.zig          (the kernel<->user contract module)
    kernel/arch/  -> kernel/architecture/   (arch -> architecture)
    device/       -> devices/           (what /system/devices reflects)
    boot/         -> /boot              (the loaders, top level)
- sbin/           -> split by role:
    init, vfs     -> system/services/<name>/<name>.zig
    hpetd, busd   -> system/drivers/<name>/<name>.zig
    vfs-test      -> system/services/vfs/vfs-test.zig  (inside the vfs project)
- lib/            -> library/runtime/   (room for other libraries beside runtime)

The VFS wire protocol becomes its own module, system/services/vfs/protocol.zig
("vfs-protocol"): the vfs sub-project exposes its interface, and the runtime's
file layer imports it by name. First instance of the "protocol module" pattern
(docs/driver-model.md); usb/block will expose theirs the same way.

Also: fix a naming-standard violation in the protocol — Op -> Operation (and
req -> request, _pad -> _padding). Docs updated: /system/services added to the
FHS doc, a repository-layout section added to the docs index, and stale source
paths swept across comments and docs.

Runtime boot paths are unchanged (the bootloader still loads /sbin/init);
aligning the runtime filesystem to the FHS is a separate follow-up. Suite 35/35
plus host tests green.
This commit is contained in:
Daniel Samson
2026-07-10 12:55:56 +01:00
parent 15b70856c9
commit 8754d4e46a
83 changed files with 334 additions and 177 deletions
+413
View File
@@ -0,0 +1,413 @@
//! Local APIC and its timer — the source of device interrupts.
//!
//! Modern x86 routes interrupts through the per-CPU Local APIC (the legacy 8259
//! PIC is remapped out of the way and masked). The LAPIC also has a built-in
//! timer, which is the simplest device interrupt to bring up: it needs no
//! external routing, just a vector and a count. We use it as danos's heartbeat.
//!
//! The LAPIC is memory-mapped (default physical 0xFEE00000, inside our identity
//! map). Every interrupt must be acknowledged with an end-of-interrupt write, or
//! the LAPIC won't deliver the next one.
const danos = @import("danos");
const io = @import("io.zig");
const paging = @import("paging.zig");
/// The ACPI PM timer, as a calibration reference: an I/O port or MMIO counter.
pub const PmTimer = struct { mmio: bool, address: u64, is_32bit: bool };
// Platform facts from discovery (set by `configure` before bring-up). Defaults are
// the legacy-safe assumptions so the code still works if discovery never ran.
var configuration_pic_present: bool = true;
var configuration_hpet_base: u64 = 0; // 0 = no HPET discovered
var configuration_pm_timer: ?PmTimer = null;
/// Which reference the last calibration used, for logging.
var cal_source: []const u8 = "none";
/// Hand the LAPIC bring-up the discovered platform facts. Call before `init`.
pub fn configure(pic_present: bool, hpet_base: u64, pm_timer: ?PmTimer) void {
configuration_pic_present = pic_present;
configuration_hpet_base = hpet_base;
configuration_pm_timer = pm_timer;
}
/// The calibration reference the timer was measured against ("cpuid"/"hpet"/…).
pub fn calibrationSource() []const u8 {
return cal_source;
}
/// IDT vector the timer fires on (in the device range, >= 32).
pub const timer_vector = 32;
/// Spurious-interrupt vector. Low nibble 0xF by convention; also in our gate
/// range so a stray spurious interrupt lands on a valid (no-op) handler.
const spurious_vector = 47;
// LAPIC register offsets.
const register_spurious = 0x0F0;
const register_eoi = 0x0B0;
const register_id = 0x020; // this core's LAPIC id, in bits 24-31
const register_icr_low = 0x300; // interrupt command register, low dword (writing it sends)
const register_icr_high = 0x310; // ICR high dword (destination APIC id in bits 24-31)
const register_lvt_timer = 0x320;
const register_timer_initial = 0x380;
const register_timer_current = 0x390;
const register_timer_divide = 0x3E0;
const icr_delivery_pending = 1 << 12; // ICR low bit 12: a previous IPI is still in flight
const lvt_masked = 1 << 16;
const lvt_periodic = 1 << 17;
const timer_divide_16 = 0x3;
const ia32_apic_base_msr = 0x1B;
/// LAPIC MMIO base. A runtime var (not a constant) both because we read it from
/// the MSR and so register writes compile to normal stores rather than a
/// `mov moffs`, which the self-hosted backend can't encode.
var base: usize = 0xFEE00000;
var tick_count: u64 = 0;
/// LAPIC timer counts per millisecond, measured against the PIT (see calibrate).
/// At divide-by-16, this is the effective counting rate.
var ticks_per_ms: u32 = 0;
/// The periodic-interrupt frequency the timer is armed at, once initTimer runs.
var timer_hz: u32 = 0;
/// TSC (Time Stamp Counter) calibration: cycles per second, and the count at boot.
/// The TSC is a per-core cycle counter, giving a ~nanosecond high-resolution
/// monotonic clock — far finer than the millisecond timer tick.
var tsc_hz: u64 = 0;
var tsc_base: u64 = 0;
/// Read the 64-bit Time Stamp Counter.
fn rdtsc() u64 {
var low: u32 = undefined;
var high: u32 = undefined;
asm volatile ("rdtsc"
: [low] "={eax}" (low),
[high] "={edx}" (high),
);
return (@as(u64, high) << 32) | low;
}
fn read(register: u32) u32 {
return @as(*volatile u32, @ptrFromInt(base + register)).*;
}
fn write(register: u32, value: u32) void {
@as(*volatile u32, @ptrFromInt(base + register)).* = value;
}
/// Move the legacy 8259 PIC's vectors to 0x20-0x2F (clear of the CPU exception
/// vectors) and mask every line, so it can't deliver interrupts behind the APIC.
fn remapAndMaskPic() void {
io.outb(0x20, 0x11); // start init (cascade mode)
io.outb(0xA0, 0x11);
io.outb(0x21, 0x20); // master offset 0x20
io.outb(0xA1, 0x28); // slave offset 0x28
io.outb(0x21, 0x04); // tell master about slave on IRQ2
io.outb(0xA1, 0x02);
io.outb(0x21, 0x01); // 8086 mode
io.outb(0xA1, 0x01);
io.outb(0x21, 0xFF); // mask all
io.outb(0xA1, 0xFF);
}
/// Enable the Local APIC: mask the PIC (only if one is present — a legacy-free
/// UEFI Class 3 machine may have none), set the global-enable MSR bit, and
/// software-enable the APIC via its spurious-vector register.
pub fn init() void {
if (configuration_pic_present) remapAndMaskPic();
const msr = io.rdmsr(ia32_apic_base_msr);
// Reach the LAPIC through the physmap (paging.init maps its page there).
base = @intCast(danos.physicalToVirtual(msr & 0xFFFFF000)); // physical base is bits 12+
io.wrmsr(ia32_apic_base_msr, msr | (1 << 11)); // global enable
write(register_spurious, 0x100 | spurious_vector); // bit 8 = software enable
}
/// Software-enable *this* core's Local APIC — the application-processor counterpart
/// of `init`, minus the one-time PIC remap (the BSP already masked it) and minus
/// calibration (the timer rate is a shared hardware constant, measured once). Each
/// core has its own LAPIC at the same MMIO address, so no per-core base is needed.
pub fn initSecondary() void {
const msr = io.rdmsr(ia32_apic_base_msr);
io.wrmsr(ia32_apic_base_msr, msr | (1 << 11)); // global enable
write(register_spurious, 0x100 | spurious_vector); // software enable
}
// --- application-processor wakeup (INIT–SIPI–SIPI) --------------------------
/// Send an INIT IPI to the core with Local APIC id `apic_id` — the first step of
/// the wake sequence. Blocks until the LAPIC reports the IPI was delivered.
pub fn sendInit(apic_id: u32) void {
write(register_icr_high, apic_id << 24);
write(register_icr_low, 0x4500); // INIT, physical destination, assert, edge-triggered
waitIcrIdle();
}
/// Send a STARTUP IPI (SIPI) telling the target core to begin executing at physical
/// address `vector << 12` (in real mode). Per the Intel bring-up protocol this is
/// sent twice after the INIT; both calls block until delivery completes.
pub fn sendStartup(apic_id: u32, vector: u8) void {
write(register_icr_high, apic_id << 24);
write(register_icr_low, 0x4600 | @as(u32, vector)); // STARTUP with the page vector
waitIcrIdle();
}
fn waitIcrIdle() void {
while (read(register_icr_low) & icr_delivery_pending != 0) {}
}
/// The calibration window: we time everything against a 10 ms reference interval.
const calib_ms = 10;
/// Measure the LAPIC timer's and the TSC's rates. The PIT (legacy 8254) can be
/// absent on UEFI Class 3 firmware — and polling it would hang — so we pick a
/// reference clock in order of preference: the CPU's own TSC frequency (CPUID leaf
/// 0x15, no external timer needed), then the discovered HPET, then the ACPI PM
/// timer, and only the PIT as a last resort. Each path yields the same two rates.
pub fn calibrate() void {
var done = false;
// 1. CPUID leaf 0x15 gives the TSC frequency directly — measure the LAPIC
// against the TSC itself, needing no external timer at all.
if (cpuidTscHz()) |hz| {
measure(hz, ~@as(u64, 0), rdtsc);
tsc_hz = hz; // keep the exact enumerated value
cal_source = "cpuid";
done = true;
}
// 2. The discovered HPET.
if (!done and configuration_hpet_base != 0) {
if (hpetHz()) |hpet_hz| {
measure(hpet_hz, hpetMask(), readHpet);
cal_source = "hpet";
done = true;
}
}
// 3. The ACPI PM timer (fixed 3.579545 MHz).
if (!done) {
if (configuration_pm_timer) |pt| {
measure(3_579_545, if (pt.is_32bit) 0xFFFF_FFFF else 0xFF_FFFF, readPmTimer);
cal_source = "pm-timer";
done = true;
}
}
// 4. The legacy PIT, last resort.
if (!done) {
calibratePit();
cal_source = "pit";
}
// A bad measurement (no reference actually ticked) leaves nonsense; fall back.
if (ticks_per_ms == 0 or tsc_hz == 0) {
calibratePit();
cal_source = "pit";
}
tsc_base = rdtsc(); // the clock's zero point (boot)
}
/// Run the LAPIC timer one-shot from its maximum count while a monotonic reference
/// clock (frequency `ref_hz`, counter width `ref_mask`) counts out `calib_ms`, and
/// snapshot the TSC across the same window. Yields `ticks_per_ms` and `tsc_hz`.
fn measure(ref_hz: u64, ref_mask: u64, refNow: *const fn () u64) void {
const calib_ticks = ref_hz / (1000 / calib_ms); // reference ticks in calib_ms
write(register_timer_divide, timer_divide_16);
write(register_lvt_timer, lvt_masked);
write(register_timer_initial, 0xFFFFFFFF);
const ref0 = refNow();
const tsc0 = rdtsc();
while (((refNow() -% ref0) & ref_mask) < calib_ticks) {}
const tsc1 = rdtsc();
const elapsed = 0xFFFFFFFF - read(register_timer_current);
write(register_timer_initial, 0);
ticks_per_ms = elapsed / calib_ms;
tsc_hz = (tsc1 -% tsc0) * (1000 / calib_ms);
}
/// The PIT fallback (legacy 8254 channel 2, polled). Only reached when no better
/// reference exists — on a legacy-free machine this path isn't taken.
fn calibratePit() void {
const pit_hz = 1_193_182;
const pit_count: u16 = @intCast(pit_hz / 1000 * calib_ms);
write(register_timer_divide, timer_divide_16);
write(register_lvt_timer, lvt_masked);
write(register_timer_initial, 0xFFFFFFFF);
io.outb(0x61, io.inb(0x61) & 0xFC); // speaker off, gate low
io.outb(0x43, 0xB0); // channel 2, lo/hi byte, mode 0
io.outb(0x42, @truncate(pit_count));
io.outb(0x42, @truncate(pit_count >> 8));
const tsc_start = rdtsc();
io.outb(0x61, (io.inb(0x61) & 0xFC) | 0x01); // gate high -> start
var guard: u64 = 0;
while (io.inb(0x61) & 0x20 == 0 and guard < 100_000_000) : (guard += 1) {} // bounded
const tsc_end = rdtsc();
const elapsed = 0xFFFFFFFF - read(register_timer_current);
write(register_timer_initial, 0);
ticks_per_ms = elapsed / calib_ms;
tsc_hz = (tsc_end -% tsc_start) * (1000 / calib_ms);
}
// --- reference clocks ------------------------------------------------------
/// TSC frequency from CPUID leaf 0x15 (crystal_hz * numerator / denominator), or
/// null if the CPU doesn't enumerate it (common under QEMU).
fn cpuidTscHz() ?u64 {
if (cpuid(0).eax < 0x15) return null;
const r = cpuid(0x15);
if (r.eax == 0 or r.ebx == 0 or r.ecx == 0) return null; // ratio/crystal not given
return @as(u64, r.ecx) * r.ebx / r.eax;
}
const CpuidRegs = struct { eax: u32, ebx: u32, ecx: u32, edx: u32 };
fn cpuid(leaf: u32) CpuidRegs {
var a: u32 = undefined;
var b: u32 = undefined;
var c: u32 = undefined;
var d: u32 = undefined;
asm volatile ("cpuid"
: [a] "={eax}" (a),
[b] "={ebx}" (b),
[c] "={ecx}" (c),
[d] "={edx}" (d),
: [leaf] "{eax}" (leaf),
[sub] "{ecx}" (@as(u32, 0)),
);
return .{ .eax = a, .ebx = b, .ecx = c, .edx = d };
}
// HPET registers: capabilities at +0x00 (period in the high dword, in fs; bit 13 =
// 64-bit-counter capable), general configuration at +0x10, main counter at +0xF0.
fn hpetRead64(off: usize) u64 {
return @as(*volatile u64, @ptrFromInt(configuration_hpet_base + off)).*;
}
fn hpetWrite64(off: usize, value: u64) void {
@as(*volatile u64, @ptrFromInt(configuration_hpet_base + off)).* = value;
}
/// Map + enable the HPET and return its tick frequency, or null if unusable.
/// Maps the HPET into the physmap and switches configuration_hpet_base to that virtual
/// address, so the register accessors reach it without the identity map.
fn hpetHz() ?u64 {
configuration_hpet_base = paging.mapMmio(configuration_hpet_base, 0x400, true);
const caps = hpetRead64(0x00);
const period_fs = caps >> 32; // femtoseconds per tick
if (period_fs == 0) return null;
hpetWrite64(0x10, hpetRead64(0x10) | 1); // ENABLE_CNF: start the main counter
return 1_000_000_000_000_000 / period_fs; // 1e15 fs/s ÷ fs/tick
}
/// The HPET counter width mask (64- or 32-bit, per caps bit 13).
fn hpetMask() u64 {
return if (hpetRead64(0x00) & (1 << 13) != 0) ~@as(u64, 0) else 0xFFFF_FFFF;
}
fn readHpet() u64 {
return hpetRead64(0xF0);
}
fn readPmTimer() u64 {
const pt = configuration_pm_timer.?;
// MMIO PM timer via the physmap (mapMmio is idempotent); the common case is
// a legacy I/O port.
if (pt.mmio) return @as(*volatile u32, @ptrFromInt(paging.mapMmio(pt.address, 4, false))).*;
return io.inl(@intCast(pt.address));
}
/// Arm the LAPIC timer to fire on `timer_vector` at `hz` (periodic). Requires
/// calibrate() to have run.
pub fn initTimer(hz: u32) void {
timer_hz = hz;
const count = @as(u64, ticks_per_ms) * 1000 / hz; // counts per (1/hz) second
write(register_timer_divide, timer_divide_16);
write(register_lvt_timer, timer_vector | lvt_periodic);
write(register_timer_initial, @intCast(count));
}
/// Configured periodic-interrupt frequency (Hz).
pub fn frequencyHz() u32 {
return timer_hz;
}
/// Measured LAPIC timer frequency (Hz), for reporting/sanity checks.
pub fn lapicHz() u64 {
return @as(u64, ticks_per_ms) * 1000;
}
/// Measured TSC frequency (Hz).
pub fn tscHz() u64 {
return tsc_hz;
}
// Monotonic high-resolution clock, from the TSC. A function per resolution, each
// scaling the cycle delta directly at its unit (the 128-bit intermediate avoids
// overflow across a long uptime). nanos() resolves to a few ns; millis() is what
// the scheduler uses for sleep deadlines.
pub fn nanos() u64 {
if (tsc_hz == 0) return 0;
return @intCast(@as(u128, rdtsc() -% tsc_base) * 1_000_000_000 / tsc_hz);
}
pub fn micros() u64 {
if (tsc_hz == 0) return 0;
return @intCast(@as(u128, rdtsc() -% tsc_base) * 1_000_000 / tsc_hz);
}
pub fn millis() u64 {
if (tsc_hz == 0) return 0;
return @intCast(@as(u128, rdtsc() -% tsc_base) * 1_000 / tsc_hz);
}
/// Acknowledge the current interrupt so the LAPIC will deliver the next one.
pub fn eoi() void {
write(register_eoi, 0);
}
/// Optional callback run each tick (the scheduler registers it for preemption).
var on_tick: ?*const fn () void = null;
pub fn setTickHook(hook: *const fn () void) void {
on_tick = hook;
}
/// The timer interrupt handler: advance the monotonic tick count, then run the
/// tick hook (which may switch tasks). The interrupt is already acknowledged by
/// the dispatcher before we get here, so a task switch here doesn't stall it.
pub fn timerTick() void {
// Acknowledge before the tick hook: `on_tick` is the scheduler, which may switch
// tasks and not return promptly, and the LAPIC mustn't wait on it to deliver the
// next interrupt. (Each device handler now owns its own EOI — see
// `idt.interruptDispatch` — because a *routed* interrupt must be masked at the
// I/O APIC before it is acknowledged, an ordering the dispatcher can't impose.)
eoi();
tick_count +%= 1;
if (on_tick) |hook| hook();
}
/// This core's Local APIC id — the interrupt destination for `routeGsi`.
pub fn localId() u8 {
return @truncate(read(register_id) >> 24);
}
/// Number of timer ticks so far. Volatile load: the count is bumped
/// asynchronously by the interrupt handler, so callers must re-read memory.
pub fn ticks() u64 {
return @as(*const volatile u64, &tick_count).*;
}
+578
View File
@@ -0,0 +1,578 @@
//! x86_64 CPU operations. This is the "architecture" module: the generic kernel imports
//! it as `@import("architecture")` and never names x86_64 directly, so a second
//! architecture is added by pointing that module at a different directory in
//! build.zig — no change to the generic code. Keep everything CPU-specific here
//! (halt, the descriptor tables, later paging), and nothing generic.
const danos = @import("danos");
const parameters = @import("parameters");
const gdt = @import("gdt.zig");
const tss = @import("tss.zig");
const idt = @import("idt.zig");
const paging = @import("paging.zig");
const serial = @import("serial.zig");
const apic = @import("apic.zig");
const ioapic = @import("ioapic.zig");
const io = @import("io.zig");
const smp = @import("smp.zig");
const pcpu = @import("per-cpu.zig");
/// The saved register/trap frame passed to a fault handler.
pub const CpuState = idt.CpuState;
// --- trap-frame accessors ---------------------------------------------------
// The frame's fields are x86_64 registers; the generic kernel reads it through
// these accessors so it never names one.
/// The interrupted/faulting instruction address (RIP here; ELR_EL1 on aarch64,
/// sepc on riscv64).
pub fn instructionPointer(state: *const CpuState) u64 {
return state.rip;
}
/// The interrupted stack pointer (RSP here).
pub fn stackPointer(state: *const CpuState) u64 {
return state.rsp;
}
/// Whether the trap came from user mode (CPL 3 here; EL0 on aarch64, U-mode on
/// riscv64).
pub fn fromUser(state: *const CpuState) bool {
return state.cs & 3 == 3;
}
/// The faulting virtual address, if this trap is a page fault (CR2 here;
/// FAR_EL1 on aarch64, stval on riscv64). Null for any other exception.
pub fn faultAddress(state: *const CpuState) ?u64 {
if (state.vector != 14) return null;
return asm volatile ("mov %%cr2, %[out]"
: [out] "=r" (-> u64),
);
}
// --- system_call ABI ------------------------------------------------------------
// The System V-style register convention (number in rax, arguments in
// rdi/rsi/rdx/r10/r8/r9, result in rax), exposed positionally so the generic
// dispatcher never names a register.
/// The system_call number the user program passed.
pub fn systemCallNumber(state: *const CpuState) u64 {
return state.rax;
}
/// Positional system_call argument `n`.
pub fn systemCallArg(state: *const CpuState, n: u8) u64 {
return switch (n) {
0 => state.rdi,
1 => state.rsi,
2 => state.rdx,
3 => state.r10,
4 => state.r8,
5 => state.r9,
else => 0,
};
}
/// Write the system_call's return value into the frame — the entry paths restore
/// user registers from it.
pub fn setSystemCallResult(state: *CpuState, value: u64) void {
state.rax = value;
}
/// Write a *second* system_call return value (rdx here — restored by both the
/// system_call/sysret and int-0x80 entry paths; unlike rcx/r11 it is not consumed by
/// sysretq). Used by IPC_ReplyWait to hand back the sender's badge alongside the
/// message length in rax.
pub fn setSystemCallResult2(state: *CpuState, value: u64) void {
state.rdx = value;
}
/// Bring up the serial port (the kernel's machine-readable log). No dependencies,
/// so it can be the very first thing called.
pub fn serialInit() void {
serial.init();
}
/// Write bytes to the serial port.
pub fn serialWrite(bytes: []const u8) void {
serial.write(bytes);
}
/// Emit a one-byte progress checkpoint to whatever hardware debug sink the
/// platform has — here the POST diagnostic port (0x80), which a POST card or BMC
/// displays. The last-resort progress signal when there's no text output at all.
/// Writing 0x80 is universally safe (it's the legacy I/O-delay port).
pub fn checkpoint(code: u8) void {
io.outb(0x80, code);
}
/// Whether a Bochs/QEMU-style debug console is on port 0xE9 (it returns 0xE9 when
/// read). On real hardware the port reads back 0xFF, so this stays false — a safe
/// probe before we write to it.
pub fn debugconPresent() bool {
return io.inb(0xE9) == 0xE9;
}
/// Output sink: write bytes to the 0xE9 debug console (see `debugconPresent`).
pub fn debugconWrite(bytes: []const u8) void {
for (bytes) |b| io.outb(0xE9, b);
}
/// Set up the CPU's descriptor tables: our own GDT, the TSS (with an interrupt
/// stack for double faults), then the IDT with exception handlers. After this a
/// CPU fault is reported instead of triple-faulting. Install the fault handler
/// (setFaultHandler) first so early faults are caught.
pub fn init() void {
gdt.init();
tss.init();
idt.init();
pcpu.initSystemCall();
}
/// Build the kernel's own page tables (with real permissions) and switch onto
/// them. Needs the frame allocator and the boot info (for the memory map and the
/// kernel's segment layout). Call once the frame allocator is up.
pub fn enablePaging(allocFrame: *const fn () ?u64, freeFrame: *const fn (u64) void, boot_information: *const danos.BootInformation) void {
paging.init(allocFrame, freeFrame, boot_information);
}
/// Create a new address space (returns the physical address of its root table —
/// the PML4 here — or null). Shares the kernel's higher half; the user (low)
/// half starts empty.
pub fn createAddressSpace() ?u64 {
return paging.createAddressSpace();
}
/// Free an address space and everything mapped in its user half. Caller must not
/// be running on it.
pub fn destroyAddressSpace(root: u64) void {
paging.destroyAddressSpace(root);
}
/// Map a user page into address space `root` (W^X is the caller's contract).
pub fn mapUserPageInto(root: u64, virtual: u64, physical: u64, writable: bool, executable: bool) void {
paging.mapUserInto(root, virtual, physical, writable, executable);
}
/// Map a device MMIO window into address space `root`: strong-uncacheable, RW+NX,
/// and marked so teardown won't free the MMIO frames as RAM. For IO passthrough.
pub fn mapUserDeviceInto(root: u64, virtual: u64, physical: u64, len: u64) void {
paging.mapUserDeviceInto(root, virtual, physical, len);
}
/// Map a page into the kernel address space (non-executable). For the heap, etc.
pub fn mapPage(virtual: u64, physical: u64, writable: bool) void {
paging.map(virtual, physical, writable);
}
/// Map a device MMIO range and return the virtual address to reach it at. This
/// is the device layer's `Hal.mapMmio` — it hands back a physmap pointer and
/// never exposes how the mapping is placed.
pub fn mapMmio(physical: u64, len: u64, writable: bool) u64 {
return paging.mapMmio(physical, len, writable);
}
/// The kernel's page-table root (physical), shared into every address space.
pub fn kernelPageTable() u64 {
return paging.kernelPml4();
}
/// Switch the active address space (load CR3 with a physical root table).
pub fn loadPageTable(root: u64) void {
paging.loadCr3(root);
}
/// The physical root of the currently active page tables (CR3 here; TTBR0/satp
/// elsewhere).
pub fn activePageTable() u64 {
return asm volatile ("mov %%cr3, %[out]"
: [out] "=r" (-> u64),
);
}
/// Set core `cpu`'s kernel stack pointer for ring-3 -> ring-0 transitions:
/// TSS.rsp0 (for interrupts/exceptions, which switch stacks in hardware) and the
/// per-CPU `kernel_rsp` (for the system_call stub, which switches by hand). Updated
/// by the scheduler when it switches to a user task.
pub fn setKernelStack(cpu: usize, top: usize) void {
tss.rsp0Ptr(cpu).* = top;
pcpu.setKernelRsp(cpu, top);
}
/// Remove a kernel mapping.
pub fn unmapPage(virtual: u64) void {
paging.unmap(virtual);
}
/// Remove a page mapping from address space `root` (for munmap of user pages).
/// Clears the leaf entry only; freeing the underlying frame is the caller's job.
pub fn unmapUserPageInto(root: u64, virtual: u64) void {
paging.unmapInto(root, virtual);
}
/// Resolve `virtual` to its physical address in the address space rooted at `root`
/// (any address space, not just the live one), or null if unmapped. Used to find
/// the frame behind a user page for munmap, and for cross-address-space copies.
pub fn translate(root: u64, virtual: u64) ?u64 {
return paging.translateIn(root, virtual);
}
/// Map a page accessible from ring 3 (U/S bit at every level). The caller keeps
/// W^X: code read-only + executable, data writable + no-execute.
pub fn mapUserPage(virtual: u64, physical: u64, writable: bool, executable: bool) void {
paging.mapUser(virtual, physical, writable, executable);
}
// --- ring 3 entry/exit -----------------------------------------------------
/// Drop to ring 3 at `rip` on `rsp` (defined in isr.s). Saves the kernel context,
/// publishes the kernel stack pointer through `rsp0_slot` (this core's TSS.rsp0,
/// so ring-3 interrupts land on a good stack), builds an iretq frame with the
/// user selectors, and iretq's. "Returns" only when the user program triggers
/// the exit path (user_exit_to_kernel).
extern fn enter_user(rip: u64, rsp: u64, rsp0_slot: *align(4) u64) callconv(.c) void;
/// Abandon the in-flight ring-3 trap context and resume the kernel as if
/// `enter_user` had returned (defined in isr.s). Called by the exit system_call.
extern fn user_exit_to_kernel() callconv(.c) noreturn;
/// Run user code at `entry` with stack `stack_top` on this core (`cpu` = the
/// caller's CPU index; the architecture layer can't ask the scheduler). Returns after the
/// user program exits via system_call. Interrupts are disabled on return (the exit
/// arrives through an interrupt gate) — the caller re-enables.
pub fn enterUser(cpu: usize, entry: u64, stack_top: u64) void {
enter_user(entry, stack_top, tss.rsp0Ptr(cpu));
}
/// Never returns to the user program: unwind to the kernel context that called
/// `enterUser`. For the exit system_call's handler.
pub fn userExit() noreturn {
user_exit_to_kernel();
}
/// Register the handler for the user system_call gate (int 0x80, vector 128). The
/// handler may write the trap frame (see `setSystemCallResult`).
pub fn setSystemCallHandler(handler: *const fn (*CpuState) void) void {
idt.setSystemCallHandler(handler);
}
/// Publish core `cpu`'s scheduler pointer via its per-CPU block (GS base). Each
/// core calls this once, after its GDT is in place (a GS *selector* reload would
/// clobber the base). See percpu.zig for the swapgs discipline.
pub fn setCpuLocal(cpu: usize, ptr: usize) void {
pcpu.setLocal(cpu, ptr);
}
/// This core's scheduler pointer (via the GS base) — a per-core register, so each
/// core sees its own without locking. Valid in any ring-0 context.
pub fn cpuLocal() usize {
return pcpu.scheduler();
}
// --- SMP: application-processor bring-up ----------------------------------
/// Record the low (<1 MiB) frame reserved for the AP trampoline. Run once at boot.
/// The frame stays inert (zeroed, non-executable) between wakes and is armed only
/// while a core is climbing — so a core can be (re)woken at any time (retry, or a
/// future power manager) without leaving an executable page resident. See smp.zig.
pub fn setTrampolinePage(physical: u64) void {
smp.setTrampolinePage(physical);
}
/// Wake the core with hardware id `hw_id` (its Local APIC id here; MPIDR on
/// aarch64, hart id on riscv64) as dense CPU `index`, giving it `stack_top` and
/// its per-CPU pointer `percpu`; it adopts the kernel page tables. Returns false
/// if it doesn't come online within the timeout. Blocks until the core reports in.
pub fn startSecondary(hw_id: u32, stack_top: usize, percpu: usize, index: usize) bool {
// The AP adopts the kernel page tables explicitly — never the caller's live
// CR3, which a future re-wake from a core running a process would make a
// process address space.
return smp.startAp(hw_id, stack_top, percpu, index, paging.kernelPml4());
}
/// Register the generic entry a woken AP jumps to once its architecture state is up (its own
/// descriptor tables, LAPIC, and timer). The kernel passes its scheduler entry here.
pub fn setSecondaryEntry(entry: *const fn () callconv(.c) noreturn) void {
smp.setSecondaryEntry(entry);
}
/// Bytes the kernel should allocate for a secondary core's dedicated fault stack
/// (the IST double-fault stack here), and where to record its top before waking
/// the core. The stack is heap-allocated per online core (the boot CPU's is
/// static — it's needed before the allocator exists). See tss.zig.
pub const fault_stack_size = tss.ist_stack_size;
pub fn setFaultStack(cpu: usize, top: usize) void {
tss.setApIstStack(cpu, top);
}
/// Test hook: force the next `n` AP wake attempts to fail, so the retry path can be
/// exercised deterministically (see the smp-retry test). No effect when `n` is 0.
pub fn testFailNextWakes(n: u32) void {
smp.testFailNextWakes(n);
}
/// The reserved AP-trampoline frame (0 if none). For tests that check it's inert.
pub fn trampolinePage() u64 {
return smp.trampolinePage();
}
/// Whether the page at `virtual` is currently mapped executable (present, NX clear).
pub fn pageExecutable(virtual: u64) bool {
return paging.isExecutable(virtual);
}
/// Kernel tick rate (the scheduler's time quantum), from configuration.
pub const timer_hz = parameters.timer_hz;
/// The ACPI PM timer, as a calibration reference (re-exported for the configuration).
pub const PmTimer = apic.PmTimer;
/// A MADT interrupt-source override (re-exported for the configuration).
pub const IsoEntry = ioapic.IsoEntry;
/// Discovered platform facts the architecture layer needs so it makes no legacy
/// assumptions — sourced from the device tree + ACPI, passed in by the kernel.
pub const PlatformConfiguration = struct {
/// Whether the legacy 8259 PIC is present (skip programming it if not).
pic_present: bool = true,
/// HPET MMIO base (0 = none) — a calibration reference for the timer.
hpet_base: u64 = 0,
/// The ACPI PM timer, another calibration reference.
pm_timer: ?PmTimer = null,
/// I/O APIC MMIO base + its first global system interrupt (0 = none).
ioapic_base: u64 = 0,
ioapic_gsi_base: u32 = 0,
/// MADT ISA-IRQ overrides, for I/O APIC routing.
overrides: []const IsoEntry = &.{},
};
/// Apply the discovered platform configuration. Must run before `startTimer` (the timer
/// calibration reads `hpet_base`/`pm_timer`) and before any interrupt routing.
/// Maps + masks the I/O APIC immediately.
pub fn configurePlatform(configuration: PlatformConfiguration) void {
apic.configure(configuration.pic_present, configuration.hpet_base, configuration.pm_timer);
ioapic.configure(configuration.ioapic_base, configuration.ioapic_gsi_base, configuration.overrides);
ioapic.init();
}
/// Point the serial console at the UART ACPI's SPCR table named (MMIO or I/O port).
pub fn serialReconfigure(is_mmio: bool, address: u64) void {
serial.reconfigure(is_mmio, address);
}
/// The reference clock the timer was calibrated against ("cpuid"/"hpet"/…).
pub fn timerCalibrationSource() []const u8 {
return apic.calibrationSource();
}
/// External-interrupt-router diagnostics, for boot logging / verification (the
/// I/O APIC's redirection entries here; a GIC distributor or PLIC elsewhere).
pub fn irqRouteCount() u32 {
return ioapic.entryCount();
}
pub fn irqRouteRaw(n: u32) u32 {
return ioapic.entryLow(n);
}
// --- device-IRQ plumbing, for system/kernel/irq.zig -----------------------------
//
// The generic IRQ layer speaks GSIs and vectors; everything below hides the fact
// that on x86_64 those mean "I/O APIC redirection entry" and "IDT gate". The
// vector window is bounded by the stubs isr.s actually emits: `gate_count` = 48,
// vector 32 is the LAPIC timer and 47 is the spurious vector, leaving 33..46.
pub const irq_vector_base: u8 = 33;
pub const irq_vector_count: u8 = 14; // 33..46 inclusive
/// True if `gsi` is one this machine's interrupt router can deliver.
pub fn irqOwnsGsi(gsi: u32) bool {
return ioapic.ownsGsi(gsi);
}
/// Install `handler` on `vector` (an absolute IDT gate index).
pub fn irqSetHandler(vector: u8, handler: *const fn () void) void {
idt.setHandler(vector, handler);
}
/// Route `gsi` to `vector` on *this* core, masked. Unmask with `irqUnmask` once bound.
pub fn irqRoute(gsi: u32, vector: u8, level: bool, active_low: bool) void {
ioapic.routeGsi(gsi, vector, apic.localId(), level, active_low);
}
pub fn irqMask(gsi: u32) void {
ioapic.maskGsi(gsi);
}
pub fn irqUnmask(gsi: u32) void {
ioapic.unmaskGsi(gsi);
}
/// Acknowledge the interrupt currently in service on this core's LAPIC.
pub fn irqEoi() void {
apic.eoi();
}
/// Enable the Local APIC, calibrate its timer against the best available reference
/// (see apic.calibrate — no longer the PIT by default), and start it firing at
/// `timer_hz` — the kernel's real-time heartbeat. Interrupts still have to be
/// unmasked with enableInterrupts() to be delivered. Run `configurePlatform` first.
pub fn startTimer() void {
apic.init();
apic.calibrate();
idt.setHandler(apic.timer_vector, apic.timerTick);
apic.initTimer(timer_hz);
}
/// Number of timer ticks since startTimer().
pub fn ticks() u64 {
return apic.ticks();
}
// Monotonic high-resolution clock (from the TSC), one function per resolution.
pub fn nanos() u64 {
return apic.nanos();
}
pub fn micros() u64 {
return apic.micros();
}
pub fn millis() u64 {
return apic.millis();
}
/// Measured frequency of the tick timer's input clock (the LAPIC timer here), in
/// Hz, from calibration.
pub fn timerClockHz() u64 {
return apic.lapicHz();
}
/// Measured frequency of the monotonic clock's underlying counter (the TSC here;
/// CNTVCT on aarch64, `time` on riscv64), in Hz.
pub fn clockHz() u64 {
return apic.tscHz();
}
/// Unmask maskable interrupts (`sti`) so device interrupts get delivered.
pub fn enableInterrupts() void {
asm volatile ("sti");
}
/// Mask maskable interrupts (`cli`).
pub fn disableInterrupts() void {
asm volatile ("cli");
}
/// Disable interrupts and return the previous flags, so a nested critical section
/// can restore the caller's state rather than blindly re-enabling. Pairs with
/// restoreInterrupts.
pub fn saveInterrupts() u64 {
var flags: u64 = undefined;
asm volatile (
\\pushfq
\\pop %[f]
\\cli
: [f] "=r" (flags),
:
: .{ .memory = true }
);
return flags;
}
/// Re-enable interrupts only if they were enabled when `flags` was captured.
pub fn restoreInterrupts(flags: u64) void {
if (flags & 0x200 != 0) asm volatile ("sti" ::: .{ .memory = true }); // bit 9 = IF
}
/// Register a callback the timer interrupt invokes each tick (e.g. the scheduler).
pub fn setTickHook(hook: *const fn () void) void {
apic.setTickHook(hook);
}
// --- context switching (for the scheduler) -------------------------------
/// Save the current task's registers/stack and resume `new_rsp`; the old stack
/// pointer is written to `old_rsp`. Defined in isr.s.
extern fn switch_context(old_rsp: *usize, new_rsp: usize) callconv(.c) void;
pub fn switchContext(old_sp: *usize, new_sp: usize) void {
switch_context(old_sp, new_sp);
}
/// Build the initial stack for a new task so that switching to it lands in
/// `task_trampoline`, which then calls `entry`. Returns the saved stack pointer.
/// The layout must match switch_context's push order (callee-saved, then the
/// return address on top); `entry` is smuggled in via the r15 slot.
pub fn initTaskStack(stack_top: usize, entry: usize) usize {
const trampoline = @extern(*const anyopaque, .{ .name = "task_trampoline" });
var sp = stack_top;
const push = struct {
fn f(p: *usize, value: usize) void {
p.* -= @sizeOf(usize);
@as(*usize, @ptrFromInt(p.*)).* = value;
}
}.f;
push(&sp, @intFromPtr(trampoline)); // return address for switch_context's `ret`
push(&sp, 0); // rbx
push(&sp, 0); // rbp
push(&sp, 0); // r12
push(&sp, 0); // r13
push(&sp, 0); // r14
push(&sp, entry); // r15 -> task entry, read by task_trampoline
return sp;
}
/// Drop the current (kernel-context) task to ring 3 at `rip` on `rsp`, never
/// returning (defined in isr.s). Used by the scheduler's user-task trampoline
/// once it has switched onto the task and read its entry/stack. Interrupts are
/// disabled across the swapgs+iretq so no interrupt observes the user GS base in
/// ring 0; the pushed RFLAGS re-enables them in ring 3.
extern fn jump_to_user(rip: u64, rsp: u64) callconv(.c) noreturn;
pub fn jumpToUser(entry: u64, stack_top: u64) noreturn {
jump_to_user(entry, stack_top);
}
/// Route CPU exceptions to `handler`, which receives the trap frame and does not
/// return. Until set, faults just halt the core.
pub fn setFaultHandler(handler: *const fn (*const CpuState) noreturn) void {
idt.on_fault = handler;
}
/// A human-readable name for a CPU exception vector.
pub fn exceptionName(vector: u64) []const u8 {
return idt.vectorName(vector);
}
/// Read `width` bytes (1/2/4) from an I/O port. The generic device layer drives
/// ACPI registers through this rather than naming x86 port instructions; on an
/// MMIO-only architecture this would be implemented differently.
pub fn pioRead(width: u8, port: u16) u32 {
return switch (width) {
1 => io.inb(port),
2 => io.inw(port),
4 => io.inl(port),
else => 0,
};
}
/// Write `width` bytes (1/2/4) to an I/O port.
pub fn pioWrite(width: u8, port: u16, value: u32) void {
switch (width) {
1 => io.outb(port, @truncate(value)),
2 => io.outw(port, @truncate(value)),
4 => io.outl(port, value),
else => {},
}
}
/// Park the core forever. `hlt` drops it into a low-power idle until the next
/// interrupt; the loop re-halts on every wake so the stop is permanent. See
/// docs/halting.md for the full reasoning.
pub fn halt() noreturn {
while (true) asm volatile ("hlt");
}
/// Spin-wait hint (`pause`). Emitted in the body of a spinlock's busy-wait: it
/// relaxes the core while it polls a contended lock — yielding pipeline resources
/// to a hyperthread sibling and easing the cache-coherency traffic on the lock
/// line. Purely a performance/power hint; correct to omit, but kinder on the bus.
pub fn cpuRelax() void {
asm volatile ("pause");
}
+89
View File
@@ -0,0 +1,89 @@
//! Global Descriptor Table. In long mode segmentation is mostly vestigial, but
//! the CPU still needs valid code/data segment descriptors, and the IDT's gates
//! reference a code selector — so we install our own flat GDT with known
//! selectors (0x08/0x10 kernel code/data, 0x18/0x20 user data/code for ring 3)
//! rather than trusting whatever the firmware left in place.
//!
//! The code/data descriptors are identical on every core, but the **TSS descriptor
//! is per-core** (each core needs its own TSS — its own interrupt/fault stacks; see
//! tss.zig). Two cores can't share one TSS descriptor slot, so each core gets its
//! own copy of the table with its own TSS descriptor. Slot 0 is the BSP.
const parameters = @import("parameters");
/// Selectors into the table (index * 8). Same on every core's GDT.
pub const kernel_code = 0x08;
pub const kernel_data = 0x10;
pub const user_data = 0x18;
pub const user_code = 0x20;
pub const tss_selector = 0x28;
/// Ring-3 selectors as loaded from user mode: RPL 3 or'd in. The user *data*
/// descriptor is load-bearing even in long mode — iretq to CPL 3 with a null SS
/// raises #GP(0).
pub const user_code_rpl3 = user_code | 3;
pub const user_data_rpl3 = user_data | 3;
const maximum_cpus = parameters.maximum_cpus;
const entries = 7; // null, kcode, kdata, udata, ucode, TSS-low, TSS-high
/// The shared descriptors (slots 0-4); slots 5-6 hold this core's TSS descriptor,
/// filled in per core by `setTssFor`.
/// kernel code: present, ring 0, executable, readable, L=1 -> 0x00AF9A00_0000FFFF
/// kernel data: present, ring 0, writable -> 0x00CF9200_0000FFFF
/// user data: present, ring 3, writable -> 0x00CFF200_0000FFFF
/// user code: present, ring 3, executable, readable, L=1 -> 0x00AFFA00_0000FFFF
/// User data sits below user code so a future SYSRET works unchanged: it loads
/// CS = STAR.SYSRET_CS + 16 and SS = STAR.SYSRET_CS + 8, so with SYSRET_CS = 0x10
/// those land on 0x20 (user code) and 0x18 (user data).
const template = [entries]u64{
0, // null descriptor (required)
0x00AF9A000000FFFF, // kernel code (0x08)
0x00CF92000000FFFF, // kernel data (0x10)
0x00CFF2000000FFFF, // user data (0x18)
0x00AFFA000000FFFF, // user code (0x20)
0, // TSS descriptor low (0x28)
0, // TSS descriptor high
};
/// One GDT per core (each a copy of the template, differing only in its TSS slot).
var gdts = [_][entries]u64{template} ** maximum_cpus;
/// Fill core `cpu`'s 64-bit TSS system descriptor (two GDT slots) so its task
/// register can point at its own TSS. Type 0x89 = present, ring 0, available 64-bit
/// TSS. Write it into that core's GDT before it loads the TSS selector.
pub fn setTssFor(cpu: usize, base: u64, limit: u64) void {
gdts[cpu][5] = (limit & 0xFFFF) |
((base & 0xFFFF) << 16) |
(((base >> 16) & 0xFF) << 32) |
(@as(u64, 0x89) << 40) |
(((limit >> 16) & 0xF) << 48) |
(((base >> 24) & 0xFF) << 56);
gdts[cpu][6] = (base >> 32) & 0xFFFFFFFF;
}
/// The operand `lgdt` wants: table byte-length minus one, then its address.
const Descriptor = packed struct {
limit: u16,
base: u64,
};
/// Loads the GDT and reloads the segment registers (including CS). Defined in
/// isr.s — it uses the selectors 0x08 (code) and 0x10 (data) that match the table.
extern fn gdt_flush(descriptor: *const Descriptor) callconv(.c) void;
/// Load core `cpu`'s GDT and switch onto its segments. Note this reloads the segment
/// registers, which zeroes the GS base — so a core must publish its per-CPU pointer
/// (setCpuLocal) *after* calling this.
pub fn loadOnThisCpu(cpu: usize) void {
const descriptor = Descriptor{
.limit = @sizeOf([entries]u64) - 1,
.base = @intFromPtr(&gdts[cpu]),
};
gdt_flush(&descriptor);
}
/// Install the bootstrap processor's GDT (slot 0) and switch onto its segments.
pub fn init() void {
loadOnThisCpu(0);
}
+186
View File
@@ -0,0 +1,186 @@
//! Interrupt Descriptor Table, CPU-exception handlers, and device-interrupt
//! dispatch. Without this, any fault (a stray pointer, a bad page-table entry)
//! triple-faults and silently resets the machine. With it, the CPU vectors into
//! our stubs, which capture the register state and hand it to a dispatcher.
//!
//! Vectors split in two: 0-31 are CPU exceptions (terminal — reported and
//! halted); 32+ are device interrupts (a registered handler runs, the APIC is
//! acknowledged, and we return to the interrupted code).
const gdt = @import("gdt.zig");
const tss = @import("tss.zig");
/// Highest vector we install a gate/stub for (exceptions 0-31 plus the device
/// range 32-47, which covers the timer and the spurious vector).
const gate_count = 48;
/// A device-interrupt handler. It doesn't get the trap frame (a timer or keyboard
/// handler doesn't need the interrupted registers); add that if one ever does.
pub const Handler = *const fn () void;
var handlers = [_]?Handler{null} ** 256;
/// Register `handler` for a device-interrupt `vector` (>= 32).
pub fn setHandler(vector: usize, handler: Handler) void {
handlers[vector] = handler;
}
/// The ring-3 system_call gate's vector (`int $0x80`, the classic choice — well away
/// from the device range) and its handler. Unlike device handlers, a system_call
/// handler gets the (mutable) trap frame: it reads its arguments from the saved
/// user registers and writes rax as the return value, which isr_common then
/// restores into the user context.
pub const system_call_vector = 128;
var system_call_handler: ?*const fn (*CpuState) void = null;
pub fn setSystemCallHandler(handler: *const fn (*CpuState) void) void {
system_call_handler = handler;
}
/// The register + trap frame the ISR stubs build on the stack, laid out so the
/// lowest address (where RSP points when we call the handler) is the first field.
/// See the push order in `isrCommon` below.
pub const CpuState = extern struct {
r15: u64,
r14: u64,
r13: u64,
r12: u64,
r11: u64,
r10: u64,
r9: u64,
r8: u64,
rbp: u64,
rdi: u64,
rsi: u64,
rdx: u64,
rcx: u64,
rbx: u64,
rax: u64,
vector: u64, // pushed by the per-vector stub
error_code: u64, // real one from the CPU, or 0 pushed by the stub
rip: u64, // from here down: pushed by the CPU on entry
cs: u64,
rflags: u64,
rsp: u64,
ss: u64,
};
/// Where a fault is reported. The kernel overrides this (see setFaultHandler) with
/// something that prints to the console; until then, just stop.
pub var on_fault: *const fn (*const CpuState) noreturn = defaultFault;
fn defaultFault(_: *const CpuState) noreturn {
while (true) asm volatile ("hlt");
}
/// Names for the 32 defined exception vectors, for readable output.
const names = [_][]const u8{
"divide error", "debug",
"NMI", "breakpoint",
"overflow", "bound range exceeded",
"invalid opcode", "device not available",
"double fault", "coprocessor segment overrun",
"invalid TSS", "segment not present",
"stack-segment fault", "general protection fault",
"page fault", "reserved (15)",
"x87 floating-point", "alignment check",
"machine check", "SIMD floating-point",
"virtualization", "control protection",
"reserved (22)", "reserved (23)",
"reserved (24)", "reserved (25)",
"reserved (26)", "reserved (27)",
"hypervisor injection", "VMM communication",
"security exception", "reserved (31)",
};
pub fn vectorName(vector: u64) []const u8 {
return if (vector < names.len) names[vector] else "unknown";
}
/// A 64-bit IDT gate descriptor (16 bytes).
const Gate = packed struct {
offset_low: u16,
selector: u16,
ist: u8, // interrupt-stack-table index; 0 = use the current stack
flags: u8, // present, DPL, gate type
offset_mid: u16,
offset_high: u32,
reserved: u32 = 0,
};
var idt = [_]Gate{std.mem.zeroes(Gate)} ** 256;
const Descriptor = packed struct {
limit: u16,
base: u64,
};
/// Loads the IDT (`lidt`). Defined in isr.s.
extern fn idt_flush(descriptor: *const Descriptor) callconv(.c) void;
fn setGate(vector: usize, handler: u64) void {
idt[vector] = .{
.offset_low = @truncate(handler),
.selector = gdt.kernel_code,
.ist = 0,
.flags = 0x8E, // present, ring 0, 64-bit interrupt gate
.offset_mid = @truncate(handler >> 16),
.offset_high = @truncate(handler >> 32),
};
}
/// Point every installed vector at its stub (isr.s) and load the IDT.
pub fn init() void {
@setEvalBranchQuota(20000); // comptimePrint across all the gates adds up
inline for (0..gate_count) |vector| {
const stub = @extern(*const anyopaque, .{ .name = std.fmt.comptimePrint("isr{d}", .{vector}) });
setGate(vector, @intFromPtr(stub));
}
// Run the double-fault handler (vector 8) on IST1: a #DF usually means the
// current stack is unusable, so it needs a guaranteed-good one. See tss.zig.
idt[8].ist = tss.double_fault_ist;
// The system_call gate. Installed outside the 0..gate_count loop (stubs 48-127
// don't exist) and with DPL 3 — without it, `int $0x80` from ring 3 is a
// #GP. An interrupt gate (not trap): IF is cleared for the handler, which
// the ring-3 exit path relies on.
const system_call_stub = @extern(*const anyopaque, .{ .name = "isr128" });
setGate(system_call_vector, @intFromPtr(system_call_stub));
idt[system_call_vector].flags = 0xEE; // present, DPL 3, 64-bit interrupt gate
loadOnThisCpu();
}
/// Load the (shared, already-populated) IDT on the current core. The gate table is
/// read-only after `init`, so every core points its IDTR at the same one. Called by
/// the BSP via `init` and by each AP during bring-up.
pub fn loadOnThisCpu() void {
const descriptor = Descriptor{
.limit = @sizeOf(@TypeOf(idt)) - 1,
.base = @intFromPtr(&idt),
};
idt_flush(&descriptor);
}
/// Called by isr_common (isr.s) with a pointer to the trap frame. Exported so the
/// assembly stubs can `call` it by name. Exceptions are terminal; device
/// interrupts run their handler, get acknowledged, and return.
export fn interruptDispatch(state: *CpuState) callconv(.c) void {
if (state.vector < 32) {
on_fault(state); // CPU exception — never returns
} else if (state.vector == system_call_vector) {
// Software interrupt from ring 3 — no LAPIC ISR bit is set, so no EOI.
if (system_call_handler) |handler| handler(state);
} else if (handlers[state.vector]) |handler| {
// The handler owns its EOI. It used to be issued here, before the call —
// correct for the LAPIC timer, but impossible to reconcile with a
// level-triggered device line, which must be **masked at the I/O APIC
// before** it is acknowledged or it redelivers instantly and storms
// (the driver that would quiet it lives in ring 3 and hasn't run yet).
// Only the handler knows which discipline its source needs, so only the
// handler can sequence it. See apic.timerTick and irq.dispatch.
handler();
}
// else: spurious/unhandled device interrupt — don't acknowledge it
}
const std = @import("std");
+68
View File
@@ -0,0 +1,68 @@
//! x86 port I/O and model-specific registers — the low-level primitives the
//! serial port and the APIC talk to hardware through.
pub fn outb(port: u16, value: u8) void {
asm volatile ("outb %[value], %[port]"
:
: [value] "{al}" (value),
[port] "{dx}" (port),
);
}
pub fn inb(port: u16) u8 {
return asm volatile ("inb %[port], %[value]"
: [value] "={al}" (-> u8),
: [port] "{dx}" (port),
);
}
pub fn outw(port: u16, value: u16) void {
asm volatile ("outw %[value], %[port]"
:
: [value] "{ax}" (value),
[port] "{dx}" (port),
);
}
pub fn inw(port: u16) u16 {
return asm volatile ("inw %[port], %[value]"
: [value] "={ax}" (-> u16),
: [port] "{dx}" (port),
);
}
pub fn outl(port: u16, value: u32) void {
asm volatile ("outl %[value], %[port]"
:
: [value] "{eax}" (value),
[port] "{dx}" (port),
);
}
pub fn inl(port: u16) u32 {
return asm volatile ("inl %[port], %[value]"
: [value] "={eax}" (-> u32),
: [port] "{dx}" (port),
);
}
/// Read a model-specific register (returns edx:eax combined).
pub fn rdmsr(msr: u32) u64 {
var low: u32 = undefined;
var high: u32 = undefined;
asm volatile ("rdmsr"
: [low] "={eax}" (low),
[high] "={edx}" (high),
: [msr] "{ecx}" (msr),
);
return (@as(u64, high) << 32) | low;
}
pub fn wrmsr(msr: u32, value: u64) void {
asm volatile ("wrmsr"
:
: [msr] "{ecx}" (msr),
[low] "{eax}" (@as(u32, @truncate(value))),
[high] "{edx}" (@as(u32, @truncate(value >> 32))),
);
}
@@ -0,0 +1,152 @@
//! I/O APIC — routes external device interrupts (a device's line) to a LAPIC
//! vector on a chosen CPU. Its address and the ISA-IRQ-to-GSI remappings come from
//! ACPI's MADT (via discovery), never assumed.
//!
//! `init` maps the I/O APIC and **masks every input** — the correct quiescent state
//! on a legacy-free machine. Lines are then unmasked one at a time, as user-space
//! drivers bind them (`routeGsi`/`unmaskGsi`, driven by system/kernel/irq.zig).
//!
//! Two entry points, for two kinds of caller. `routeIrq` takes a legacy **ISA IRQ**
//! and resolves it through the MADT overrides — for in-kernel use, and still without
//! a caller. `routeGsi` takes a **GSI** directly, which is what a device's own
//! routing capability names (e.g. the HPET's `Tn_INT_ROUTE_CAP`), and is the path a
//! bound driver interrupt takes.
const paging = @import("paging.zig");
/// A MADT Interrupt Source Override: an ISA IRQ that appears at a different global
/// system interrupt, with its own polarity/trigger (MPS INTI `flags`).
pub const IsoEntry = struct { source: u8, gsi: u32, flags: u16 };
var base: u64 = 0; // 0 = no I/O APIC discovered
var gsi_base: u32 = 0;
var maximum_entries: u32 = 0;
var overrides: [16]IsoEntry = undefined;
var override_count: usize = 0;
// The I/O APIC exposes an index register (IOREGSEL) and a data window (IOWIN).
const register_ioregsel = 0x00;
const register_iowin = 0x10;
const register_version = 0x01;
const redir_base = 0x10; // redirection table: two 32-bit regs per entry
const redir_mask = 1 << 16; // mask bit in the low dword
/// Supply the discovered I/O APIC location + the MADT IRQ overrides. Call before `init`.
pub fn configure(ioapic_base: u64, ioapic_gsi_base: u32, isos: []const IsoEntry) void {
base = ioapic_base;
gsi_base = ioapic_gsi_base;
override_count = @min(isos.len, overrides.len);
for (isos[0..override_count], 0..) |iso, i| overrides[i] = iso;
}
fn registerRead(index: u32) u32 {
@as(*volatile u32, @ptrFromInt(base + register_ioregsel)).* = index;
return @as(*volatile u32, @ptrFromInt(base + register_iowin)).*;
}
fn registerWrite(index: u32, value: u32) void {
@as(*volatile u32, @ptrFromInt(base + register_ioregsel)).* = index;
@as(*volatile u32, @ptrFromInt(base + register_iowin)).* = value;
}
fn writeEntry(n: u32, low: u32, high: u32) void {
registerWrite(redir_base + 2 * n, low);
registerWrite(redir_base + 2 * n + 1, high);
}
/// Map the I/O APIC and mask every redirection entry — the safe quiescent state.
pub fn init() void {
if (base == 0) return;
// Reach the I/O APIC through the physmap; switch `base` to that virtual
// address so the register accessors work without the identity map.
base = paging.mapMmio(base, 0x1000, true);
maximum_entries = ((registerRead(register_version) >> 16) & 0xFF) + 1;
var n: u32 = 0;
while (n < maximum_entries) : (n += 1) writeEntry(n, redir_mask, 0);
}
/// Route ISA `irq` to `vector` on the LAPIC `apic_id`, honouring a MADT override
/// for its GSI/polarity/trigger, and unmask it. No caller yet — groundwork for the
/// first device driver.
pub fn routeIrq(irq: u8, vector: u8, apic_id: u8) void {
if (base == 0) return;
var gsi: u32 = irq;
var flags: u16 = 0;
for (overrides[0..override_count]) |o| {
if (o.source == irq) {
gsi = o.gsi;
flags = o.flags;
}
}
if (gsi < gsi_base) return;
const n = gsi - gsi_base;
if (n >= maximum_entries) return;
// Low dword: vector + delivery mode fixed(0) + physical dest(0), unmasked.
// MPS INTI flags: bits [1:0] polarity (3 = active low), [3:2] trigger (3 = level).
var low: u32 = vector;
if (flags & 0x3 == 3) low |= (1 << 13);
if ((flags >> 2) & 0x3 == 3) low |= (1 << 15);
const high: u32 = @as(u32, apic_id) << 24; // destination APIC ID
writeEntry(n, low, high);
}
// --- GSI-level control (the user-space driver path) --------------------------
//
// `routeIrq` above takes an *ISA IRQ* and resolves it through the MADT overrides.
// A driver-bound interrupt is already a **GSI** (the device told us so, e.g. the
// HPET's `Tn_INT_ROUTE_CAP`), so it needs no override lookup — just the redirection
// entry. These three are what `system/kernel/irq.zig` drives.
//
// Callers must serialise: the I/O APIC is reached through an index/data register
// pair, so two cores interleaving `registerWrite` would corrupt each other. The kernel
// holds the big lock across these.
/// Redirection-entry index for `gsi`, or null if this I/O APIC doesn't own it.
fn entryFor(gsi: u32) ?u32 {
if (base == 0 or gsi < gsi_base) return null;
const n = gsi - gsi_base;
return if (n < maximum_entries) n else null;
}
/// True if `gsi` lands on this I/O APIC — the kernel's validity check before binding.
pub fn ownsGsi(gsi: u32) bool {
return entryFor(gsi) != null;
}
/// Point `gsi` at `vector` on the LAPIC `apic_id`, with explicit polarity/trigger,
/// and leave it **masked**. The caller unmasks once a handler is bound — otherwise a
/// device asserting between route and bind would fire into a null handler.
pub fn routeGsi(gsi: u32, vector: u8, apic_id: u8, level: bool, active_low: bool) void {
const n = entryFor(gsi) orelse return;
var low: u32 = @as(u32, vector) | redir_mask; // masked until bound
if (active_low) low |= (1 << 13);
if (level) low |= (1 << 15);
writeEntry(n, low, @as(u32, apic_id) << 24);
}
/// Stop `gsi` reaching any CPU. Called from the ISR *before* the LAPIC EOI: a
/// level-triggered line is still asserted at that point, so an unmasked entry would
/// redeliver immediately and storm before the user-space driver ever runs.
pub fn maskGsi(gsi: u32) void {
const n = entryFor(gsi) orelse return;
registerWrite(redir_base + 2 * n, registerRead(redir_base + 2 * n) | redir_mask);
}
/// Let `gsi` through again — the tail of `irq_ack`, once the driver has quieted the
/// device (so the line is deasserted and this can't immediately refire).
pub fn unmaskGsi(gsi: u32) void {
const n = entryFor(gsi) orelse return;
registerWrite(redir_base + 2 * n, registerRead(redir_base + 2 * n) & ~@as(u32, redir_mask));
}
/// Number of redirection entries the I/O APIC advertises (0 until `init`).
pub fn entryCount() u32 {
return maximum_entries;
}
/// The low dword of redirection entry `n` — for diagnostics/read-back.
pub fn entryLow(n: u32) u32 {
if (base == 0) return 0;
return registerRead(redir_base + 2 * n);
}
+382
View File
@@ -0,0 +1,382 @@
# x86_64 low-level entry code: the CPU-exception stubs, plus the GDT/IDT load
# helpers. Kept in a dedicated assembly file rather than inline asm because these
# need real labels and cross-symbol jumps/calls (isr_common, exceptionHandler),
# and because `lgdt`/`lidt` memory operands aren't expressible in Zig inline asm.
#
# Each exception vector normalises the stack to a uniform trap frame — a dummy
# error code where the CPU pushes none, then the vector number — and jumps to the
# shared tail, which saves the general registers and calls the Zig handler with a
# pointer to the frame (matching src/arch/x86_64/idt.zig's CpuState).
.text
# _start: the kernel entry. The loader jumps here (higher-half address) with
# boot_info in RDI, still on the loader's low stack. Switch to a kernel-owned
# stack in .bss (the loader stack is a low address that goes away once the low
# half is dropped), keeping RDI, then call the Zig entry. kmainEntry never
# returns; the hlt loop is a belt-and-braces backstop.
.global _start
_start:
leaq bootstrap_stack_top(%rip), %rsp
call kmainEntry
1: hlt
jmp 1b
# The kernel's initial stack (used until the scheduler hands each task its own).
# 64 KiB: kmain's discovery path includes the recursive AML interpreter, so it
# needs more than a token stack. Lives in .bss (zeroed, higher-half).
.section .bss
.balign 16
bootstrap_stack:
.skip 65536
bootstrap_stack_top:
.text
# gdt_flush(rdi = *GDT descriptor): load the GDT, reload the data segment
# registers to the data selector, and reload CS to the code selector. CS can't be
# set with mov, so we far-return through the caller's own return address.
.global gdt_flush
gdt_flush:
lgdt (%rdi)
mov $0x10, %ax # kernel data selector
mov %ax, %ds
mov %ax, %es
mov %ax, %ss
mov %ax, %fs
mov %ax, %gs
pop %rax # caller's return address
push $0x08 # kernel code selector (new CS)
push %rax # return address (new RIP)
lretq
# idt_flush(rdi = *IDT descriptor): load the IDT.
.global idt_flush
idt_flush:
lidt (%rdi)
ret
# load_tr(di = TSS selector): load the task register.
.global load_tr
load_tr:
ltr %di
ret
# switch_context(rdi = &old_task.rsp, rsi = new_task.rsp)
# Cooperative context switch: save the callee-saved registers on the current
# stack, stash the stack pointer in the old task, load the new task's stack
# pointer, restore its callee-saved registers, and return into it. Caller-saved
# registers are the compiler's responsibility (this looks like a normal call).
.global switch_context
switch_context:
push %rbx
push %rbp
push %r12
push %r13
push %r14
push %r15
mov %rsp, (%rdi) # save old stack pointer into old_task.rsp
mov %rsi, %rsp # switch to the new task's stack
pop %r15
pop %r14
pop %r13
pop %r12
pop %rbp
pop %rbx
ret # return into the new task's saved instruction pointer
# task_trampoline: the first thing a freshly-spawned task runs. init_task_stack
# leaves its entry function in r15. A fresh task is switched to with the big kernel
# lock held (the hand-off rule in sync.zig) but has no enter/leave frame of its own,
# so it releases the lock here before running its body. r15 survives the call (it's
# callee-saved). New tasks then start with interrupts enabled.
.extern releaseForFreshTask
.global task_trampoline
task_trampoline:
call releaseForFreshTask # drop the kernel lock we inherited across the switch
sti
call *%r15 # call the task entry (fn() void)
1: hlt # if the entry returns, idle (still preemptible)
jmp 1b
# user_task_trampoline: the first thing a freshly-spawned *user* task runs.
# init_user_task_stack leaves the user entry in r15 and the user stack in r14
# (both callee-saved, so they survive the lock-release call). Like task_trampoline
# it drops the inherited kernel lock, then — instead of calling a kernel fn — it
# builds an iretq frame and drops to ring 3. The scheduler's switchTo already
# loaded this task's address space (CR3) and published its kernel stack
# (TSS.rsp0 + gs kernel_rsp) before switching here, so interrupts and syscalls
# from ring 3 land correctly. IF is set in the pushed RFLAGS (no sti needed).
# jump_to_user(rdi = user rip, rsi = user rsp): drop the current kernel context
# to ring 3, never returning. The scheduler calls this from a fresh user task's
# trampoline (after the lock is released and the entry/stack read from the Task).
# cli guards the swapgs..iretq window: an interrupt there would run in ring 0
# with the user GS base and mis-read per-CPU data. The pushed RFLAGS (IF set)
# re-enables interrupts on the drop to ring 3.
.global jump_to_user
jump_to_user:
cli
push $0x1B # user SS (0x18 | RPL 3)
push %rsi # user RSP
push $0x202 # RFLAGS: IF | reserved-1
push $0x23 # user CS (0x20 | RPL 3)
push %rdi # user RIP
swapgs # user GS base (isr_common/syscall swap back on entry)
iretq
# --- ring 3 entry/exit ------------------------------------------------------
# enter_user(rdi = user rip, rsi = user rsp, rdx = &TSS.rsp0)
# Drop to ring 3. Saves the callee-saved registers and the kernel stack pointer
# (so user_exit_to_kernel can unwind back here), publishes that stack pointer as
# this core's TSS.rsp0 — everything below it is dead, so ring-3 interrupt frames
# grow safely into it — then builds the 5-word iretq frame with the user
# selectors (RPL 3) and drops privilege. IF is set in the pushed RFLAGS so the
# timer keeps running in user mode.
.global enter_user
enter_user:
push %rbx
push %rbp
push %r12
push %r13
push %r14
push %r15
mov %rsp, user_saved_rsp(%rip) # where user_exit_to_kernel unwinds to
mov %rsp, (%rdx) # TSS.rsp0: ring-3 interrupts stack here
mov %rsp, %gs:0 # kernel_rsp: the syscall stub stacks here too
push $0x1B # user SS (0x18 | RPL 3)
push %rsi # user RSP
push $0x202 # RFLAGS: IF | reserved-1
push $0x23 # user CS (0x20 | RPL 3)
push %rdi # user RIP
swapgs # user GS base for ring 3 (isr_common swaps back)
iretq
# user_exit_to_kernel: abandon the in-flight ring-3 trap frame (it lives in the
# dead zone below user_saved_rsp) and return as if enter_user's call completed.
# Reached from the exit syscall's handler; interrupts are off (interrupt gate)
# and stay off — the Zig caller re-enables.
.global user_exit_to_kernel
user_exit_to_kernel:
mov user_saved_rsp(%rip), %rsp
pop %r15
pop %r14
pop %r13
pop %r12
pop %rbp
pop %rbx
ret
# syscall_entry: the target of the SYSCALL instruction (LSTAR). The CPU does NOT
# switch stacks — it puts the return RIP in RCX, the saved RFLAGS in R11, loads
# CS/SS from STAR, masks RFLAGS with SFMASK (so IF is already clear), and jumps
# here with RSP still the *user* stack. We swap in the kernel GS, switch to the
# task's kernel stack via the per-CPU block, build a CpuState frame identical to
# the interrupt path's, and reuse interruptDispatch (vector 128) — then SYSRET.
#
# Hazard (acceptable while init is the only, trusted, user program): SYSRETQ #GPs
# in ring 0 if the return RIP (RCX) is non-canonical. A hostile user could arrange
# that; hardening (canonical check / iretq fallback) is a later security-track item.
.global syscall_entry
syscall_entry:
swapgs # kernel GS base
movq %rsp, %gs:8 # stash user rsp in the scratch slot
movq %gs:0, %rsp # switch to this task's kernel stack
# Build the trap frame (same field order as isr_common), highest field first.
pushq $0x1B # ss (user data | 3)
pushq %gs:8 # rsp (user, from scratch)
pushq %r11 # rflags (saved by syscall)
pushq $0x23 # cs (user code | 3)
pushq %rcx # rip (saved by syscall)
pushq $0 # error_code (none for a syscall)
pushq $128 # vector (same as the int 0x80 gate)
push %rax
push %rbx
push %rcx
push %rdx
push %rsi
push %rdi
push %rbp
push %r8
push %r9
push %r10
push %r11
push %r12
push %r13
push %r14
push %r15
mov %rsp, %rdi # trap-frame pointer
call interruptDispatch
pop %r15
pop %r14
pop %r13
pop %r12
pop %r11
pop %r10
pop %r9
pop %r8
pop %rbp
pop %rdi
pop %rsi
pop %rdx
pop %rcx
pop %rbx
pop %rax
add $16, %rsp # drop vector + error_code -> rsp at rip
popq %rcx # rip -> RCX (SYSRETQ restores RIP from RCX)
addq $8, %rsp # skip the cs slot (SYSRETQ loads CS from STAR)
popq %r11 # rflags -> R11 (SYSRETQ restores RFLAGS from R11)
popq %rsp # user rsp (the ss slot below is abandoned)
swapgs # user GS base
sysretq # -> ring 3: RIP=RCX, RFLAGS=R11, CS/SS from STAR
.section .bss
.balign 8
user_saved_rsp:
.skip 8
.text
# --- user-mode test program --------------------------------------------------
# A hand-assembled ring-3 blob, copied by the kernel onto a user-mapped page and
# entered via enter_user. Position-independent (immediates and short jumps only).
# In .rodata: these bytes are data to the kernel — they only execute at CPL 3
# from the user mapping. (The old hello/ping blob was retired once /sbin/init
# became the real ring-3 exerciser; only the isolation proof remains.)
.section .rodata
# The isolation-proof program: read a kernel-only page from ring 3. The LAPIC
# lives in the kernel's physmap (physmap_base + 0xFEE00000) as a supervisor
# page, so this must take a #PF with error code 0x5 (present | user) before any
# access happens. movabs loads the full 64-bit higher-half address (a disp32
# would sign-extend and miss).
.global user_pf_start
.global user_pf_end
user_pf_start:
movabs $0xFFFF8800FEE00000, %rcx
mov (%rcx), %rax
1: jmp 1b
user_pf_end:
.text
# Stub for a vector the CPU does NOT push an error code for: push a dummy 0.
.macro STUB_NOERR vec
.global isr\vec
isr\vec:
pushq $0
pushq $\vec
jmp isr_common
.endm
# Stub for a vector the CPU DOES push an error code for: leave it in place.
.macro STUB_ERR vec
.global isr\vec
isr\vec:
pushq $\vec
jmp isr_common
.endm
STUB_NOERR 0
STUB_NOERR 1
STUB_NOERR 2
STUB_NOERR 3
STUB_NOERR 4
STUB_NOERR 5
STUB_NOERR 6
STUB_NOERR 7
STUB_ERR 8
STUB_NOERR 9
STUB_ERR 10
STUB_ERR 11
STUB_ERR 12
STUB_ERR 13
STUB_ERR 14
STUB_NOERR 15
STUB_NOERR 16
STUB_ERR 17
STUB_NOERR 18
STUB_NOERR 19
STUB_NOERR 20
STUB_ERR 21
STUB_NOERR 22
STUB_NOERR 23
STUB_NOERR 24
STUB_NOERR 25
STUB_NOERR 26
STUB_NOERR 27
STUB_NOERR 28
STUB_NOERR 29
STUB_NOERR 30
STUB_NOERR 31
# Device-interrupt vectors (timer, spurious, room for more). None push an error
# code, so they all use the dummy-zero form.
STUB_NOERR 32
STUB_NOERR 33
STUB_NOERR 34
STUB_NOERR 35
STUB_NOERR 36
STUB_NOERR 37
STUB_NOERR 38
STUB_NOERR 39
STUB_NOERR 40
STUB_NOERR 41
STUB_NOERR 42
STUB_NOERR 43
STUB_NOERR 44
STUB_NOERR 45
STUB_NOERR 46
STUB_NOERR 47
# The syscall gate (int $0x80 from ring 3). Same frame shape as every other
# vector; dispatched specially in interruptDispatch.
STUB_NOERR 128
.extern interruptDispatch
# Shared tail. Register push order here defines the CpuState field order.
# If the interrupt came from ring 3 the GS base holds the user's value, so swap
# in the kernel's before anything reads per-CPU data (swapgs discipline; see
# percpu.zig). CS sits at offset 24 here (vector@0, error@8, RIP@16, CS@24).
isr_common:
testb $3, 24(%rsp)
jz 1f
swapgs
1: push %rax
push %rbx
push %rcx
push %rdx
push %rsi
push %rdi
push %rbp
push %r8
push %r9
push %r10
push %r11
push %r12
push %r13
push %r14
push %r15
mov %rsp, %rdi # first argument: pointer to the trap frame
call interruptDispatch
pop %r15
pop %r14
pop %r13
pop %r12
pop %r11
pop %r10
pop %r9
pop %r8
pop %rbp
pop %rdi
pop %rsi
pop %rdx
pop %rcx
pop %rbx
pop %rax
add $16, %rsp # drop the vector and error code
# Symmetric to entry: if returning to ring 3, restore the user GS base. CS is
# now at offset 8 (RIP@0, CS@8).
testb $3, 8(%rsp)
jz 1f
swapgs
1: iretq
@@ -0,0 +1,52 @@
/* Kernel link layout — higher half.
*
* The kernel is linked to *run* in the higher half (virtual base
* 0xFFFF_FFFF_8000_0000, matching danos.kernel_virt_base and build.zig's
* image_base) but is *loaded* low. Each section's load address (LMA) is its
* virtual address minus KERNEL_VIRT_BASE via AT(), so the ELF's p_paddr lands
* at a low physical address (.text at 1 MiB) that the loader can allocate and
* copy into. The loader maps p_vaddr (high) -> p_paddr (low) in its bootstrap
* tables and jumps to the high entry; the kernel then builds its own tables
* with the physmap and abandons the identity map. Requires LLD (build.zig pins
* it) — the self-hosted linker ignores PHDRS/AT()/section order.
*/
KERNEL_VIRT_BASE = 0xFFFFFFFF80000000;
ENTRY(_start)
/* One loadable segment per permission set, so the loader can map .text as R+X,
* .rodata as R, and .data/.bss as R+W. FLAGS bits: 1=X, 2=W, 4=R. */
PHDRS {
text PT_LOAD FLAGS(5); /* R + X */
rodata PT_LOAD FLAGS(4); /* R */
data PT_LOAD FLAGS(6); /* R + W */
}
SECTIONS {
.text ALIGN(4K) : AT(ADDR(.text) - KERNEL_VIRT_BASE) {
*(.text .text.*)
} :text
.rodata ALIGN(4K) : AT(ADDR(.rodata) - KERNEL_VIRT_BASE) {
*(.rodata .rodata.*)
} :rodata
.data ALIGN(4K) : AT(ADDR(.data) - KERNEL_VIRT_BASE) {
*(.data .data.*)
} :data
/* .bss occupies memory but not file space. The loader zeroes it via the
* gap between each PT_LOAD segment's file size and memory size, so no
* boundary symbols are needed here. */
.bss ALIGN(4K) : AT(ADDR(.bss) - KERNEL_VIRT_BASE) {
*(.bss .bss.*)
*(COMMON)
} :data
/DISCARD/ : {
*(.comment)
*(.note .note.*)
*(.eh_frame .eh_frame_hdr)
}
}
@@ -0,0 +1,389 @@
//! The kernel's page tables and virtual memory manager.
//!
//! Builds our own 4-level page tables and switches CR3 onto them, replacing the
//! firmware's. Unlike the earlier bootstrap this maps with real permissions:
//! RAM is identity-mapped read-write + no-execute, the kernel's own segments get
//! their ELF permissions (code R+X, rodata R, data R+W+NX), and page 0 is left
//! unmapped as a null guard. It also exposes map/unmap for on-demand mapping,
//! which the kernel heap will build on.
//!
//! Everything is 4 KiB pages — precise and simple; the extra table memory is
//! negligible against available RAM.
const danos = @import("danos");
const io = @import("io.zig");
const page_size = danos.page_size;
// Page-table entry bits.
const present: u64 = 1 << 0;
const writable: u64 = 1 << 1;
const user: u64 = 1 << 2; // U/S: accessible from ring 3 (must be set at every level)
const pwt: u64 = 1 << 3; // page write-through
const pcd: u64 = 1 << 4; // page cache disable (with PWT: strong-uncacheable under the default PAT)
const device_grant: u64 = 1 << 9; // available bit: this leaf maps device MMIO, not RAM — do not reclaim
const no_execute: u64 = 1 << 63;
const address_mask: u64 = 0x000F_FFFF_FFFF_F000;
// ELF segment flags (p_flags).
const pf_x: u32 = 1;
const pf_w: u32 = 2;
// State kept after init so map()/unmap() can serve later callers (e.g. the heap).
var kernel_pml4: u64 = 0;
var alloc_frame: *const fn () ?u64 = undefined;
var free_frame: *const fn (u64) void = undefined; // for tearing down address spaces
/// Set once the kernel is running on its own tables (past the CR3 load in
/// `init`). Before that, the kernel reaches page-table frames through the
/// *loader's* bootstrap physmap, which only covers the low 4 GiB — so every
/// frame allocated for a table during that window must be below 4 GiB. Both the
/// frame allocator and this code scan from low addresses up, so it holds
/// naturally; the assertion in `allocTable` makes a violation loud rather than
/// a silent fault. After the switch the kernel's own physmap covers all RAM.
var on_own_tables = false;
/// Set at the end of `init`. Guards against a new *higher-half* PML4 entry being
/// created afterward: the kernel half is pre-populated at init and then shared
/// by copying PML4[256..512) into every process address space (M3), so a late
/// top-half entry would be invisible to already-created address spaces.
var init_done = false;
const bootstrap_physmap_limit: u64 = 4 << 30;
/// Dereference a page-table frame by its physical address, via the physmap.
/// This is the single hinge for the higher-half move: page tables hold physical
/// frame addresses (pmm gives out physical frames, and CR3/PTEs must be
/// physical), but the kernel reaches them at `physmap_base + physical`. Valid under
/// both the loader's bootstrap tables and the kernel's own, which share the
/// physmap base.
fn tableAt(physical: u64) *[512]u64 {
return @ptrFromInt(danos.physicalToVirtual(physical));
}
fn allocTable() u64 {
const frame = alloc_frame() orelse @panic("paging: out of memory building page tables");
if (!on_own_tables and frame >= bootstrap_physmap_limit)
@panic("paging: table frame above the 4 GiB bootstrap physmap");
@memset(tableAt(frame)[0..], 0);
return frame;
}
/// Return the table an entry points at, creating it if empty. Intermediate
/// entries are writable and executable so the leaf's bits govern (a page is
/// writable only if every level is; non-executable if any level is).
fn descend(entry: *u64) u64 {
if (entry.* & present != 0) return entry.* & address_mask;
const frame = allocTable();
entry.* = frame | present | writable;
return frame;
}
/// Map one 4 KiB page `virtual` -> `physical` with `flags` (present is added).
fn mapPage(pml4: u64, virtual: u64, physical: u64, flags: u64) void {
const pml4e = &tableAt(pml4)[(virtual >> 39) & 0x1FF];
// The kernel half is fixed after init: every top-half PML4 entry is
// pre-created so address spaces can share it by copying these slots. A new
// one here would be invisible to address spaces already made.
if (init_done and (virtual >> 63) == 1 and pml4e.* & present == 0)
@panic("paging: new higher-half PML4 entry after init");
const pdpt = descend(pml4e);
const pdpte = &tableAt(pdpt)[(virtual >> 30) & 0x1FF];
const pd = descend(pdpte);
const pde = &tableAt(pd)[(virtual >> 21) & 0x1FF];
const pt = descend(pde);
tableAt(pt)[(virtual >> 12) & 0x1FF] = (physical & address_mask) | flags | present;
}
/// Map [physical_base, physical_base+len) into the physmap (at physicalToVirtual(physical)) with
/// `flags`, rounded out to whole pages. This is how the kernel keeps a permanent
/// window onto physical memory once the low identity map goes away.
fn mapRangePhysmap(pml4: u64, physical_base: u64, len: u64, flags: u64) void {
var address = physical_base & ~@as(u64, page_size - 1);
const end = physical_base + len;
while (address < end) : (address += page_size) {
mapPage(pml4, danos.physicalToVirtual(address), address, flags);
}
}
fn regions(mm: danos.MemoryMap) []const danos.MemoryRegion {
return @as([*]const danos.MemoryRegion, @ptrFromInt(danos.physicalToVirtual(mm.regions)))[0..mm.len];
}
/// Enable the NX bit in the page-table format (EFER.NXE). Must happen before we
/// load a CR3 whose entries set the NX bit, or those bits are reserved and fault.
fn enableNx() void {
const efer_msr = 0xC0000080;
io.wrmsr(efer_msr, io.rdmsr(efer_msr) | (1 << 11));
}
/// Build the address space and switch onto it.
pub fn init(allocFrame: *const fn () ?u64, freeFrame: *const fn (u64) void, boot_information: *const danos.BootInformation) void {
alloc_frame = allocFrame;
free_frame = freeFrame;
enableNx();
const pml4 = allocTable();
// 1. All RAM in the physmap (physicalToVirtual(physical)) RW + NX. No identity/low-half
// mapping: the low half belongs to user space. MMIO is skipped here and
// mapped on demand (mapMmio) or explicitly below.
for (regions(boot_information.memory_map)) |r| {
if (r.kind == .mmio) continue;
mapRangePhysmap(pml4, r.base, r.pages * page_size, present | writable | no_execute);
}
// 2. Physmap windows for the framebuffer and the Local APIC (device memory
// the kernel touches directly), RW + NX.
const fb = boot_information.framebuffer;
mapRangePhysmap(pml4, fb.base, @as(u64, fb.height) * fb.pitch, present | writable | no_execute);
mapPage(pml4, danos.physicalToVirtual(0xFEE00000), 0xFEE00000, present | writable | no_execute);
// 3. The kernel's own segments at their higher-half link addresses, mapped
// to their low physical load addresses with real ELF permissions: code
// R+X, rodata R, data R+W+NX. This is the W^X guarantee.
for (boot_information.kernel_segments[0..boot_information.kernel_segment_count]) |seg| {
var flags: u64 = present;
if (seg.flags & pf_w != 0) flags |= writable;
if (seg.flags & pf_x == 0) flags |= no_execute;
var off: u64 = 0;
while (off < seg.pages * page_size) : (off += page_size) {
mapPage(pml4, seg.virtual + off, seg.physical + off, flags);
}
}
// 4. Pre-create every higher-half PML4 entry (an empty PDPT where none
// exists yet), so the whole kernel half is a fixed set of top-level
// slots. A process address space (M3) then shares the kernel half simply
// by copying PML4[256..512) — growth beneath these slots (heap, on-demand
// MMIO) propagates to every address space because they share the PDPTs.
for (256..512) |i| {
const e = &tableAt(pml4)[i];
if (e.* & present == 0) e.* = allocTable() | present | writable;
}
kernel_pml4 = pml4;
asm volatile ("mov %[pml4], %%cr3"
:
: [pml4] "r" (pml4),
: .{ .memory = true }
);
on_own_tables = true; // now on the kernel's physmap (covers all RAM)
init_done = true; // the kernel half is fixed from here
}
/// The kernel's own top-level page table (physical). Every kernel task and every
/// per-process address space shares this table's higher half.
pub fn kernelPml4() u64 {
return kernel_pml4;
}
/// Load CR3 (switch the active address space). `pml4` is a physical frame.
pub fn loadCr3(pml4: u64) void {
asm volatile ("mov %[pml4], %%cr3"
:
: [pml4] "r" (pml4),
: .{ .memory = true });
}
/// Map a page into the kernel address space on demand (for the heap, etc.).
/// `writable_page` controls W; pages are always mapped non-executable.
pub fn map(virtual: u64, physical: u64, writable_page: bool) void {
var flags: u64 = present | no_execute;
if (writable_page) flags |= writable;
mapPage(kernel_pml4, virtual, physical, flags);
invalidate(virtual);
}
/// Map a device MMIO range into the physmap and return the virtual address to
/// use for it (physicalToVirtual(physical)). The single way the kernel (and the device
/// layer, via the HAL) reaches memory-mapped registers once the identity map is
/// gone: physmap pages are RW + NX, so a driver never executes device memory.
/// Idempotent for already-mapped ranges. `len` 0 maps one page.
pub fn mapMmio(physical: u64, len: u64, writable_page: bool) u64 {
var flags: u64 = present | no_execute;
if (writable_page) flags |= writable;
const first = physical & ~@as(u64, page_size - 1);
const last = physical + (if (len == 0) 1 else len) - 1;
var address = first;
while (address <= (last & ~@as(u64, page_size - 1))) : (address += page_size) {
const virtual = danos.physicalToVirtual(address);
mapPage(kernel_pml4, virtual, address, flags);
invalidate(virtual);
}
return danos.physicalToVirtual(physical);
}
/// Like `descend`, but also sets the U/S bit on the intermediate entry (new or
/// pre-existing): ring-3 access requires U at *every* level, and `descend` leaves
/// existing entries untouched. Only used under user-exclusive virtual ranges, so
/// no kernel mapping's protection is widened (the leaf still governs).
fn descendUser(entry: *u64) u64 {
const table = descend(entry);
entry.* |= user;
return table;
}
/// Map one 4 KiB page `virtual` -> `physical` accessible from ring 3. W^X is the
/// caller's contract: code pages are read-only + executable, data pages are
/// writable + no-execute. `virtual` must lie in a user-exclusive region (see
/// `descendUser`).
pub fn mapUser(virtual: u64, physical: u64, writable_page: bool, executable: bool) void {
mapUserInto(kernel_pml4, virtual, physical, writable_page, executable);
}
/// Map a ring-3-accessible page into the address space rooted at `pml4` (which
/// may be a process's own table or the kernel's). W^X is the caller's contract.
pub fn mapUserInto(pml4: u64, virtual: u64, physical: u64, writable_page: bool, executable: bool) void {
var flags: u64 = present | user;
if (writable_page) flags |= writable;
if (!executable) flags |= no_execute;
const pml4e = &tableAt(pml4)[(virtual >> 39) & 0x1FF];
const pdpt = descendUser(pml4e);
const pdpte = &tableAt(pdpt)[(virtual >> 30) & 0x1FF];
const pd = descendUser(pdpte);
const pde = &tableAt(pd)[(virtual >> 21) & 0x1FF];
const pt = descendUser(pde);
tableAt(pt)[(virtual >> 12) & 0x1FF] = (physical & address_mask) | flags;
invalidate(virtual);
}
/// Map a device MMIO window `[physical, physical+len)` into the user (low) half of the
/// address space rooted at `pml4`, page by page. Unlike `mapUserInto` these pages
/// are **strong-uncacheable** (PCD|PWT — device registers must not be cached) and
/// carry the `device_grant` bit so teardown does not return the MMIO frames to the
/// RAM allocator (`freeSubtree`). RW + NX; the caller places `virtual` in a
/// user-exclusive range (PML4[225]). Both `virtual` and `physical` are page-aligned by
/// the caller; a sub-page `physical` offset is the caller's to re-apply.
pub fn mapUserDeviceInto(pml4: u64, virtual: u64, physical: u64, len: u64) void {
const flags: u64 = present | user | writable | no_execute | pcd | pwt | device_grant;
const first = physical & ~@as(u64, page_size - 1);
const last = (physical + (if (len == 0) 1 else len) - 1) & ~@as(u64, page_size - 1);
var off: u64 = 0;
while (first + off <= last) : (off += page_size) {
const v = virtual + off;
const pml4e = &tableAt(pml4)[(v >> 39) & 0x1FF];
const pdpt = descendUser(pml4e);
const pdpte = &tableAt(pdpt)[(v >> 30) & 0x1FF];
const pd = descendUser(pdpte);
const pde = &tableAt(pd)[(v >> 21) & 0x1FF];
const pt = descendUser(pde);
tableAt(pt)[(v >> 12) & 0x1FF] = ((first + off) & address_mask) | flags;
invalidate(v);
}
}
/// Create a new address space: a fresh PML4 with an empty user half and the
/// kernel's higher half shared in (copying PML4[256..512), whose entries point
/// at the kernel's PDPTs — pre-created at init and never restaled, so growth in
/// the kernel half propagates to every address space). Returns the physical
/// PML4, or null if out of frames.
pub fn createAddressSpace() ?u64 {
const pml4 = alloc_frame() orelse return null;
const t = tableAt(pml4);
@memset(t[0..256], 0); // empty user half
@memcpy(t[256..512], tableAt(kernel_pml4)[256..512]); // shared kernel half
return pml4;
}
/// Tear down an address space created by `createAddressSpace`: free every frame
/// and table in the user half [0..256), then the PML4 itself. The shared kernel
/// half [256..512) is never touched. The caller must not be running on `pml4`.
pub fn destroyAddressSpace(pml4: u64) void {
const t = tableAt(pml4);
for (0..256) |i| {
if (t[i] & present != 0) freeSubtree(t[i] & address_mask, 3); // PDPT level
}
free_frame(pml4);
}
/// Recursively free a page-table subtree: `level` 3 = PDPT, 2 = PD, 1 = PT. At
/// level 1 the entries are leaf data frames; above, they are child tables.
fn freeSubtree(physical: u64, level: u32) void {
const t = tableAt(physical);
for (t) |e| {
if (e & present == 0) continue;
if (level > 1) {
freeSubtree(e & address_mask, level - 1);
} else if (e & device_grant == 0) {
// A device-grant leaf points at MMIO, not RAM — returning it to the
// frame allocator would corrupt the pool. Only reclaim real RAM.
free_frame(e & address_mask);
}
}
free_frame(physical); // page-table frames are always real RAM
}
/// Whether `virtual` is currently mapped **executable** — present with the NX bit
/// clear. Walks the 4-level tables (all danos mappings are 4 KiB, so no huge-page
/// case). Returns false if unmapped. Used for W^X checks in tests.
pub fn isExecutable(virtual: u64) bool {
const pml4e = tableAt(kernel_pml4)[(virtual >> 39) & 0x1FF];
if (pml4e & present == 0) return false;
const pdpte = tableAt(pml4e & address_mask)[(virtual >> 30) & 0x1FF];
if (pdpte & present == 0) return false;
const pde = tableAt(pdpte & address_mask)[(virtual >> 21) & 0x1FF];
if (pde & present == 0) return false;
const pte = tableAt(pde & address_mask)[(virtual >> 12) & 0x1FF];
if (pte & present == 0) return false;
return pte & no_execute == 0;
}
/// Make an already-identity-mapped RAM page **executable** (clear its NX bit),
/// leaving it present and writable. The blanket RAM mapping is NX for W^X, but the
/// application processors fetch the AP trampoline from a low RAM page under paging —
/// so that one page must be executable. A deliberate, temporary W^X exception for a
/// single bring-up page; the caller frees it once every AP is up.
pub fn setExecutable(physical: u64) void {
mapPage(kernel_pml4, physical, physical, present | writable); // note: no no_execute
invalidate(physical);
}
/// Remove a mapping and flush it from the TLB.
pub fn unmap(virtual: u64) void {
unmapInto(kernel_pml4, virtual);
}
/// Remove a mapping from the address space rooted at `pml4` (a process's own
/// table or the kernel's) and flush it from the TLB. Clears only the leaf PTE —
/// the intermediate tables and any frame the PTE pointed at are left to the
/// caller (munmap frees the frame; `destroyAddressSpace` reclaims the tables).
pub fn unmapInto(pml4: u64, virtual: u64) void {
const pml4e = tableAt(pml4)[(virtual >> 39) & 0x1FF];
if (pml4e & present == 0) return;
const pdpte = tableAt(pml4e & address_mask)[(virtual >> 30) & 0x1FF];
if (pdpte & present == 0) return;
const pde = tableAt(pdpte & address_mask)[(virtual >> 21) & 0x1FF];
if (pde & present == 0) return;
tableAt(pde & address_mask)[(virtual >> 12) & 0x1FF] = 0;
invalidate(virtual);
}
/// Resolve a virtual address to a physical one in the address space rooted at
/// `pml4`, walking the tables through the physmap (CR3-independent — works for
/// any address space, not just the live one). Returns null if `virtual` is not
/// mapped at any level. All danos mappings are 4 KiB, so there is no huge-page
/// case. The foundation for cross-address-space copies and for munmap (which
/// needs the frame behind a user vaddr to free it).
pub fn translateIn(pml4: u64, virtual: u64) ?u64 {
const pml4e = tableAt(pml4)[(virtual >> 39) & 0x1FF];
if (pml4e & present == 0) return null;
const pdpte = tableAt(pml4e & address_mask)[(virtual >> 30) & 0x1FF];
if (pdpte & present == 0) return null;
const pde = tableAt(pdpte & address_mask)[(virtual >> 21) & 0x1FF];
if (pde & present == 0) return null;
const pte = tableAt(pde & address_mask)[(virtual >> 12) & 0x1FF];
if (pte & present == 0) return null;
return (pte & address_mask) | (virtual & (page_size - 1));
}
fn invalidate(virtual: u64) void {
// invlpg needs its operand via a register-indirect memory reference that Zig
// inline asm won't form directly, so stage the address in a register first.
asm volatile (
\\mov %[v], %%rax
\\invlpg (%%rax)
:
: [v] "r" (virtual),
: .{ .rax = true, .memory = true }
);
}
@@ -0,0 +1,77 @@
//! Per-CPU data reached through the GS segment base. The GS base holds a pointer
//! to this core's `ArchitecturePerCpu`, so kernel code gets the running core's block with
//! a single MSR read (`scheduler()`) and the system_call entry stub gets its kernel stack
//! with a `%gs`-relative load (no usable stack yet at that point).
//!
//! **swapgs discipline.** In ring 0 the GS base points here; in ring 3 it holds
//! the user's own GS (which ring 3 may set freely), and this pointer lives in the
//! KERNEL_GS_BASE MSR instead. Every ring-3 -> ring-0 entry (`swapgs` in the
//! system_call stub and the conditional swapgs in isr_common) brings it back, and
//! every ring-0 -> ring-3 exit swaps it away. Because the very first ring
//! transition is always an exit (the kernel starts in ring 0), the swap pairs
//! keep the invariant without seeding KERNEL_GS_BASE. `scheduler()` is therefore
//! valid in any ring-0 context and never sees a user-controlled base.
const std = @import("std");
const io = @import("io.zig");
const parameters = @import("parameters");
const ia32_gs_base = 0xC000_0101;
/// Layout is load-bearing: the system_call entry stub in isr.s reaches `kernel_rsp`
/// at `%gs:0` and `scratch` at `%gs:8`. Keep those two first; the asserts below
/// pin the offsets.
pub const ArchitecturePerCpu = extern struct {
kernel_rsp: u64 = 0, // %gs:0 — kernel stack top for system_call entry (== TSS.rsp0)
scratch: u64 = 0, // %gs:8 — stashes the user rsp during system_call entry
scheduler: usize = 0, // the scheduler's PerCpu pointer (what `cpuLocal` returns)
};
comptime {
std.debug.assert(@offsetOf(ArchitecturePerCpu, "kernel_rsp") == 0);
std.debug.assert(@offsetOf(ArchitecturePerCpu, "scratch") == 8);
}
var blocks = [_]ArchitecturePerCpu{.{}} ** parameters.maximum_cpus;
/// Publish core `index`'s per-CPU block: record the scheduler pointer and point
/// the GS base at the block. Called once per core during bring-up, after the GDT
/// is loaded (a GS *selector* reload would clobber the base).
pub fn setLocal(index: usize, scheduler_ptr: usize) void {
blocks[index].scheduler = scheduler_ptr;
io.wrmsr(ia32_gs_base, @intFromPtr(&blocks[index]));
}
/// The scheduler pointer for the running core (via the GS base). Valid in any
/// ring-0 context under the swapgs discipline.
pub fn scheduler() usize {
return @as(*const ArchitecturePerCpu, @ptrFromInt(io.rdmsr(ia32_gs_base))).scheduler;
}
/// Record core `index`'s kernel stack top, used by the system_call entry stub to
/// switch off the user stack. The scheduler sets this (and TSS.rsp0) whenever it
/// switches to a user task.
pub fn setKernelRsp(index: usize, top: usize) void {
blocks[index].kernel_rsp = top;
}
// Fast-system_call MSRs.
const ia32_efer = 0xC000_0080;
const ia32_star = 0xC000_0081;
const ia32_lstar = 0xC000_0082;
const ia32_sfmask = 0xC000_0084;
/// Enable the `system_call`/`sysret` fast path on this core (BSP and each AP). EFER.SCE
/// turns the instructions on; STAR sets the selectors system_call/sysret load; LSTAR
/// is the entry stub (isr.s); SFMASK clears RFLAGS bits on entry (notably IF —
/// the handler runs with interrupts off, like the int-gate path). The GDT is laid
/// out (kernel code 0x08, then user data 0x18 / code 0x20) precisely so these line
/// up: system_call loads CS 0x08 / SS 0x10; sysret loads CS = base+16 and SS = base+8
/// with RPL forced to 3, so base 0x10 gives CS 0x23 (user code|3) and SS 0x1B.
pub fn initSystemCall() void {
io.wrmsr(ia32_efer, io.rdmsr(ia32_efer) | 1); // SCE
io.wrmsr(ia32_star, (@as(u64, 0x08) << 32) | (@as(u64, 0x10) << 48));
const entry = @extern(*const anyopaque, .{ .name = "syscall_entry" });
io.wrmsr(ia32_lstar, @intFromPtr(entry));
io.wrmsr(ia32_sfmask, 0x4_0700); // clear IF, TF, DF, AC on entry
}
@@ -0,0 +1,86 @@
//! Serial console (16550-compatible UART) — the kernel's machine-readable output
//! channel. Unlike the framebuffer console, serial text can be captured to a file
//! by QEMU (`-serial file:...`), which is what the test harness asserts on.
//!
//! The UART defaults to the legacy PC COM1 at I/O port `0x3F8`, but a UEFI Class 3
//! (legacy-free) machine may have no COM1 — or its debug UART somewhere else, and
//! reachable via MMIO rather than port I/O. So the location is a runtime value:
//! `reconfigure` repoints it once ACPI's SPCR table has been read. Early boot logs
//! optimistically to COM1 (harmless if absent); the framebuffer console is the
//! always-present log.
const paging = @import("paging.zig");
/// How the UART registers are reached: legacy I/O ports or memory-mapped.
const Access = enum { port, mmio };
var access: Access = .port;
var base: u64 = 0x3F8; // COM1
fn portOut(p: u16, value: u8) void {
asm volatile ("outb %[value], %[p]"
:
: [value] "{al}" (value),
[p] "{dx}" (p),
);
}
fn portIn(p: u16) u8 {
return asm volatile ("inb %[p], %[value]"
: [value] "={al}" (-> u8),
: [p] "{dx}" (p),
);
}
/// Read UART register `off` through the active access method.
fn register(off: u64) u8 {
if (access == .mmio) return @as(*volatile u8, @ptrFromInt(base + off)).*;
return portIn(@intCast(base + off));
}
/// Write UART register `off` through the active access method.
fn setRegister(off: u64, value: u8) void {
if (access == .mmio) {
@as(*volatile u8, @ptrFromInt(base + off)).* = value;
} else {
portOut(@intCast(base + off), value);
}
}
/// Configure the UART: 38400 baud, 8N1, FIFO on. Safe to call before anything
/// else; it has no dependencies, and is a harmless no-op if the port is absent.
pub fn init() void {
setRegister(1, 0x00); // disable interrupts
setRegister(3, 0x80); // enable DLAB (set baud divisor)
setRegister(0, 0x03); // divisor low: 38400 baud
setRegister(1, 0x00); // divisor high
setRegister(3, 0x03); // 8 bits, no parity, one stop bit; DLAB off
setRegister(2, 0xC7); // enable + clear FIFO, 14-byte threshold
setRegister(4, 0x0B); // RTS/DSR set
}
/// Point the console at the UART ACPI's SPCR table names (MMIO or I/O port) and
/// re-run the UART setup there. Called after discovery when an SPCR entry exists.
pub fn reconfigure(is_mmio: bool, address: u64) void {
access = if (is_mmio) .mmio else .port;
// An MMIO UART is reached through the physmap; an I/O-port UART keeps its
// port number unchanged.
base = if (is_mmio) paging.mapMmio(address, 0x100, true) else address;
init();
}
fn writeByte(c: u8) void {
// Wait for the transmit-holding register to empty — but bounded, so an absent
// UART (whose line-status register reads back as 0x00) can't hang the kernel.
var guard: u32 = 0;
while (register(5) & 0x20 == 0 and guard < 100_000) : (guard += 1) {}
setRegister(0, c);
}
/// Write bytes, translating LF to CRLF so terminals and logs line up.
pub fn write(bytes: []const u8) void {
for (bytes) |c| {
if (c == '\n') writeByte('\r');
writeByte(c);
}
}
+182
View File
@@ -0,0 +1,182 @@
//! Application-processor (AP) bring-up: waking the cores the firmware left parked.
//!
//! The firmware starts only the bootstrap processor (BSP); the others sit idle until
//! the kernel wakes them with an INIT–SIPI–SIPI sequence (Intel SDM Vol.3, "MP
//! Initialization"). A woken core begins in 16-bit real mode at a low physical page,
//! runs the [trampoline](trampoline.s) up into 64-bit long mode, and lands in
//! `apEntry` here. This module copies the trampoline into place, patches its
//! per-AP parameters, drives the wake IPIs, and waits for each core to report in.
//!
//! Cores are brought up **one at a time**: a single trampoline page and parameter
//! block are reused, so the BSP patches, wakes, and waits for one AP before the
//! next. That also lets `apEntry` pick up its dense CPU index from a plain global.
//! Once a core has its own descriptor tables, LAPIC, and timer, it calls the generic
//! scheduler entry and joins the run loop — mechanism here, policy there.
const danos = @import("danos");
const io = @import("io.zig");
const gdt = @import("gdt.zig");
const tss = @import("tss.zig");
const idt = @import("idt.zig");
const apic = @import("apic.zig");
const paging = @import("paging.zig");
const pcpu = @import("per-cpu.zig");
/// IA32_GS_BASE — the per-CPU data pointer (see cpu.zig; kept in sync here so the AP
/// path doesn't depend on cpu.zig and risk an import cycle).
const ia32_gs_base = 0xC000_0101;
const page_size = 0x1000;
/// Physical address of the low (<1 MiB) frame reserved for the trampoline. Held for
/// the life of the system so any core can be (re)woken on demand — a retry, or a
/// future power manager bringing a core back online. The frame is kept **inert**
/// between wakes (zeroed and non-executable) and only armed for the brief moment a
/// core is actually climbing. Its low 20 bits are zero, so `physical >> 12` is the SIPI
/// vector.
var tramp_physical: u64 = 0;
/// Set to 1 by a freshly-woken AP once it reaches `apEntry` and finishes its own
/// bring-up. The BSP clears it before each wake and polls it afterwards — a simple
/// one-at-a-time handshake (only one AP is being started at any moment).
var ap_alive: u32 = 0;
/// The dense CPU index of the AP currently being started. Set by the BSP before the
/// wake, read by `apEntry` (safe because bring-up is strictly one core at a time).
var boot_index: usize = 0;
/// The generic scheduler entry a woken core jumps to once its architecture state is up. Set
/// by the kernel via `setSecondaryEntry`; never returns.
var secondary_entry: ?*const fn () callconv(.c) noreturn = null;
/// Register the generic entry an AP calls once its per-CPU tables/LAPIC/timer are up.
pub fn setSecondaryEntry(entry: *const fn () callconv(.c) noreturn) void {
secondary_entry = entry;
}
/// Test hook: force the next `n` wake attempts to fail (skipping the actual
/// INIT-SIPI-SIPI), so the retry path can be exercised deterministically. Zero in
/// normal operation — the smp-retry test arms it via `architecture.testFailNextWakes`.
var fail_next_wakes: u32 = 0;
pub fn testFailNextWakes(n: u32) void {
fail_next_wakes = n;
}
/// Record the reserved low frame the trampoline uses. Call once at boot. The frame
/// starts inert (identity-mapped RW+NX like all RAM); each wake arms it and disarms
/// it again, so it's only ever executable while a core is climbing.
pub fn setTrampolinePage(physical: u64) void {
tramp_physical = physical;
}
/// The reserved trampoline frame (0 if SMP bring-up never ran). Exposed so a test
/// can verify it's inert — zeroed and non-executable — when dormant.
pub fn trampolinePage() u64 {
return tramp_physical;
}
/// Arm the trampoline for a wake: make its page executable (W^X exception for the
/// duration of the climb) and copy the blob in.
fn arm() void {
// The AP executes this page at its physical address (identity) while it
// climbs from real to long mode, so it needs a low identity mapping that is
// executable — the one deliberate, transient W^X exception. The BSP writes
// the blob into the frame through the physmap.
paging.setExecutable(tramp_physical);
const start = @extern([*]const u8, .{ .name = "ap_trampoline_start" });
const end = @extern([*]const u8, .{ .name = "ap_trampoline_end" });
const len = @intFromPtr(end) - @intFromPtr(start);
const destination: [*]u8 = @ptrFromInt(danos.physicalToVirtual(tramp_physical));
@memcpy(destination[0..len], start[0..len]);
}
/// Disarm after a wake: wipe the page through the physmap and remove its low
/// identity mapping, so no executable code (nor any stale bytes, nor any
/// low-half mapping) lingers between wakes. Safe once the woken core has
/// reported in — it's long past the trampoline by then, in the kernel image; a
/// core that never answered is dead and can't be mid-climb.
fn disarm() void {
const destination: [*]u8 = @ptrFromInt(danos.physicalToVirtual(tramp_physical));
@memset(destination[0..page_size], 0);
paging.unmap(tramp_physical); // drop the transient low identity mapping
}
/// Address of a patchable trampoline parameter, by symbol name: the copied blob's
/// base plus the field's offset within it (a same-section symbol difference). The
/// pointer is `align(1)` — the fields aren't 8-aligned within the blob, and x86
/// tolerates unaligned stores, so we don't force layout constraints on the asm.
fn param(comptime name: []const u8) *align(1) volatile u64 {
const start = @intFromPtr(@extern([*]const u8, .{ .name = "ap_trampoline_start" }));
const sym = @intFromPtr(@extern([*]const u8, .{ .name = name }));
return @ptrFromInt(danos.physicalToVirtual(tramp_physical + (sym - start)));
}
/// Wake the core with Local APIC id `apic_id` as dense CPU `index`, hand it
/// `stack_top` and its per-CPU pointer `percpu`, and wait for it to come alive. This
/// is one self-contained attempt: it arms the trampoline, drives INIT–SIPI–SIPI, and
/// disarms again before returning — so it's safe to call repeatedly (a retry, or a
/// power manager re-waking a core; the INIT resets a core that was wedged). Returns
/// false if the core doesn't report in within the timeout (left parked, no harm to
/// the running system). `cr3` is the kernel page tables the AP adopts. Precondition:
/// `setTrampolinePage` has run.
pub fn startAp(apic_id: u32, stack_top: usize, percpu: usize, index: usize, cr3: u64) bool {
// The trampoline loads CR3 with a 32-bit `movl` before it reaches long mode,
// so the page-table root must be addressable in 32 bits.
if (cr3 >= (1 << 32)) @panic("smp: kernel page tables above 4 GiB");
arm();
defer disarm();
if (fail_next_wakes > 0) { // test hook: simulate a core missing this attempt
fail_next_wakes -= 1;
return false;
}
boot_index = index;
param("ap_tramp_cr3").* = cr3;
param("ap_tramp_stack").* = stack_top;
param("ap_tramp_entry").* = @intFromPtr(&apEntry);
param("ap_tramp_percpu").* = percpu;
@atomicStore(u32, &ap_alive, 0, .seq_cst);
const vector: u8 = @intCast(tramp_physical >> 12);
apic.sendInit(apic_id);
delayMicros(10_000); // 10 ms INIT settle
apic.sendStartup(apic_id, vector);
delayMicros(200);
apic.sendStartup(apic_id, vector);
// Wait up to 100 ms for the AP to reach apEntry and set the flag.
const deadline = apic.millis() + 100;
while (apic.millis() < deadline) {
if (@atomicLoad(u32, &ap_alive, .acquire) != 0) return true;
asm volatile ("pause");
}
return false;
}
/// Busy-wait `us` microseconds against the calibrated TSC clock (the AP wake happens
/// after the timer is up, so the clock is available).
fn delayMicros(us: u64) void {
const start = apic.micros();
while (apic.micros() - start < us) asm volatile ("pause");
}
/// The 64-bit entry every AP lands on, called from the trampoline with its per-CPU
/// pointer in RDI. Brings up this core's own descriptor tables, LAPIC and timer,
/// signals the BSP, then jumps to the generic scheduler entry. Never returns.
fn apEntry(percpu: usize) callconv(.c) noreturn {
const cpu = boot_index;
gdt.loadOnThisCpu(cpu); // this core's GDT (with its own TSS slot)
tss.setupThisCpu(cpu); // this core's TSS + IST stack, loaded into TR
idt.loadOnThisCpu(); // the shared IDT
pcpu.setLocal(cpu, percpu); // per-CPU block via GS base — *after* the GDT reload
pcpu.initSystemCall(); // enable system_call/sysret on this core
apic.initSecondary(); // software-enable this core's LAPIC
apic.initTimer(apic.frequencyHz()); // arm its timer (still masked: interrupts off)
@atomicStore(u32, &ap_alive, 1, .release); // "architecture state up" — BSP is polling this
if (secondary_entry) |enterScheduler| enterScheduler(); // joins the run loop
while (true) asm volatile ("hlt"); // (only if no entry was registered)
}
@@ -0,0 +1,150 @@
# AP trampoline: brings a waking application processor from the 16-bit real mode it
# starts in (after INIT-SIPI-SIPI) up through protected mode into 64-bit long mode,
# then jumps to the Zig AP entry (arch/x86_64/smp.zig:apEntry).
#
# A STARTUP IPI vectors a core to physical address `vector << 12` in real mode, so
# this blob is copied to a low (<1 MiB) page and started there; at entry CS = that
# page >> 4 and IP = 0. It is fully **position-independent**: it derives its own
# linear base (CS << 4) into EBX and addresses every internal datum as
# `(label - ap_trampoline_start)(%ebx)` — a difference of two symbols in the same
# section, which the assembler folds to a constant page offset no matter where the
# blob was linked or copied to. The BSP patches the parameter block (CR3, stack,
# entry, per-CPU pointer) before each wake; see arch/x86_64/smp.zig.
#
# It lives in .rodata (not .text): it is data to be copied out and executed
# elsewhere, never run at its link address, so it must not be a normal code segment.
.section .rodata
.balign 16
.code16
.global ap_trampoline_start
ap_trampoline_start:
cli
cld
# Linear base of this page (CS << 4) into EBX; all data is addressed off it.
xorl %eax, %eax
mov %cs, %ax
shll $4, %eax
movl %eax, %ebx
mov %cs, %ax # DS = CS, so we address our data as DS:(label - start):
mov %ax, %ds # the segment base (CS<<4) already supplies the page base,
# so data operands use the page *offset*, not EBX.
# Relocate the pointers whose absolute (linear) targets depend on where we were
# copied: the GDT base and the two far-jump targets = EBX + their page offsets.
# EBX supplies the base for the *value* (via leal); the store address is DS-rel.
leal (gdt32 - ap_trampoline_start)(%ebx), %eax
movl %eax, gdtr32_base - ap_trampoline_start
leal (prot_entry - ap_trampoline_start)(%ebx), %eax
movl %eax, jmp32_off - ap_trampoline_start
leal (long_entry - ap_trampoline_start)(%ebx), %eax
movl %eax, jmp64_off - ap_trampoline_start
lgdtl gdtr32 - ap_trampoline_start
movl %cr0, %eax # enter protected mode (CR0.PE)
orl $1, %eax
movl %eax, %cr0
ljmpl *(jmp32_ptr - ap_trampoline_start) # -> prot_entry, CS = 0x08
.code32
prot_entry:
movw $0x10, %ax # flat 32-bit data segments
movw %ax, %ds
movw %ax, %es
movw %ax, %ss
movw %ax, %fs
movw %ax, %gs
# CR4: PAE (required for long mode) + OSFXSR/OSXMMEXCPT. The kernel is built with
# SSE (part of the x86_64 baseline), and the compiler emits SSE for things as
# ordinary as a struct copy — without OSFXSR those instructions #UD. The BSP got
# these bits from UEFI; an AP starts fresh, so we must set them ourselves.
movl %cr4, %eax
orl $((1 << 5) | (1 << 9) | (1 << 10)), %eax
movl %eax, %cr4
# CR0: clear EM (no x87 emulation) and set MP, so SSE/x87 don't fault.
movl %cr0, %eax
andl $~(1 << 2), %eax # ~EM
orl $(1 << 1), %eax # MP
movl %eax, %cr0
movl (param_cr3 - ap_trampoline_start)(%ebx), %eax # kernel page tables
movl %eax, %cr3
movl $0xC0000080, %ecx # EFER: long mode enable (LME) + NX enable (NXE, since
rdmsr # the kernel's PTEs set the NX bit)
orl $((1 << 8) | (1 << 11)), %eax
wrmsr
movl %cr0, %eax # paging on (CR0.PG) — now in long mode (compat sub-mode)
orl $(1 << 31), %eax
movl %eax, %cr0
ljmpl *(jmp64_ptr - ap_trampoline_start)(%ebx) # -> long_entry, CS = 0x18 (L=1)
.code64
long_entry:
movw $0x10, %ax # sane flat data segments
movw %ax, %ds
movw %ax, %es
movw %ax, %ss
# RBX = EBX (zero-extended) = page base. Load our stack and per-CPU pointer, then
# call the Zig entry — which runs from the kernel image and never returns.
movq (param_stack - ap_trampoline_start)(%rbx), %rsp
movq (param_percpu - ap_trampoline_start)(%rbx), %rdi # SysV arg 0
movq (param_entry - ap_trampoline_start)(%rbx), %rax
callq *%rax
1: hlt # unreachable; guard against a stray return
jmp 1b
# --- data: GDT, far pointers, and the BSP-patched parameter block -----------
.balign 8
gdt32:
.quad 0x0000000000000000 # 0x00 null
.quad 0x00CF9A000000FFFF # 0x08 32-bit code (G, D, present, exec/read)
.quad 0x00CF92000000FFFF # 0x10 data (valid in 32- and 64-bit)
.quad 0x00AF9A000000FFFF # 0x18 64-bit code (L=1)
gdt32_end:
gdtr32:
.word gdt32_end - gdt32 - 1
gdtr32_base:
.long 0 # patched (16-bit code): linear base of gdt32
jmp32_ptr: # indirect far-jump operand: offset then selector
jmp32_off:
.long 0 # patched: linear address of prot_entry
.word 0x08 # 32-bit code selector
jmp64_ptr:
jmp64_off:
.long 0 # patched: linear address of long_entry
.word 0x18 # 64-bit code selector
# The parameter block, filled in by the BSP (smp.zig) before each STARTUP IPI. Global
# so the Zig side can locate each field as (symbol - ap_trampoline_start).
.global ap_tramp_cr3
.global ap_tramp_stack
.global ap_tramp_entry
.global ap_tramp_percpu
param_cr3:
ap_tramp_cr3:
.quad 0 # kernel PML4 physical address (CR3)
param_stack:
ap_tramp_stack:
.quad 0 # top of this AP's kernel stack
param_entry:
ap_tramp_entry:
.quad 0 # address of apEntry (the Zig AP entry)
param_percpu:
ap_tramp_percpu:
.quad 0 # this AP's per-CPU pointer (goes in GS base)
.global ap_trampoline_end
ap_trampoline_end:
+87
View File
@@ -0,0 +1,87 @@
//! Task State Segment and its interrupt stacks. In long mode the TSS has two
//! jobs. First, the Interrupt Stack Table: an IDT gate can name an IST entry,
//! and the CPU switches to that stack when the exception fires — no matter how
//! broken the interrupted stack was. We use IST1 for the double-fault handler,
//! so a fault that happens *because* the current stack is unusable still lands
//! on solid ground instead of triple-faulting. Second, rsp0: the kernel stack
//! the CPU switches to when an interrupt arrives from ring 3 (published by the
//! user-mode entry path via `rsp0Ptr`).
//!
//! Each core needs **its own TSS** (its own IST stack): two cores taking a fault at
//! once can't share one fault stack. So the TSS and its IST stack are per-core,
//! indexed by CPU number; slot 0 is the BSP.
const parameters = @import("parameters");
const gdt = @import("gdt.zig");
/// x86_64 TSS. `packed` because several 64-bit fields sit at 4-byte-unaligned
/// offsets (rsp0 at byte 4), which a normal struct would pad away.
const Tss = packed struct {
reserved0: u32 = 0,
rsp0: u64 = 0,
rsp1: u64 = 0,
rsp2: u64 = 0,
reserved1: u64 = 0,
ist1: u64 = 0,
ist2: u64 = 0,
ist3: u64 = 0,
ist4: u64 = 0,
ist5: u64 = 0,
ist6: u64 = 0,
ist7: u64 = 0,
reserved2: u64 = 0,
reserved3: u16 = 0,
iomap_base: u16 = 0,
};
/// The IST slot (1-based, as the IDT gate encodes it) used for critical faults.
pub const double_fault_ist = 1;
const maximum_cpus = parameters.maximum_cpus;
pub const ist_stack_size = parameters.ist_stack_size;
/// One TSS per core (small — kept static). The IST stacks are 16 KiB each, so only
/// the **BSP's** is static: it must exist before the frame allocator does, to catch a
/// fault during early boot. Each **AP** gets a heap-allocated IST stack at bring-up
/// (after the heap is up), the top of which the BSP records here before waking it —
/// so we reserve big stacks only for cores that actually come online.
var tss_table = [_]Tss{.{}} ** maximum_cpus;
var bsp_ist_stack: [ist_stack_size]u8 align(16) = undefined;
var ap_ist_top = [_]usize{0} ** maximum_cpus; // per-AP IST stack top (0 = BSP / not set)
/// Loads the task register with the TSS selector. Defined in isr.s.
extern fn load_tr(selector: u16) callconv(.c) void;
/// Address of core `cpu`'s rsp0 slot — the kernel stack the CPU switches to on a
/// ring-3 -> ring-0 interrupt. Computed as base + 4 (rsp0's architectural offset,
/// which is why the pointer is only 4-aligned) rather than `&t.rsp0`, which on a
/// packed struct would be an unaligned bit-pointer type. The ring-3 entry path
/// (enter_user in isr.s) writes the current kernel stack pointer through this
/// before dropping to user mode.
pub fn rsp0Ptr(cpu: usize) *align(4) u64 {
return @ptrFromInt(@intFromPtr(&tss_table[cpu]) + 4);
}
/// Record the top of the IST stack the kernel allocated for AP `cpu`. Called on the
/// BSP before waking that core; read by the core's own `setupThisCpu`.
pub fn setApIstStack(cpu: usize, top: usize) void {
ap_ist_top[cpu] = top;
}
/// Set up core `cpu`'s TSS: point IST1 at its stack (the BSP's static one for core 0,
/// the allocated one recorded via `setApIstStack` for an AP), install the TSS
/// descriptor into that core's GDT, and load it into the task register. Requires the
/// core's GDT to already be loaded (gdt.loadOnThisCpu first).
pub fn setupThisCpu(cpu: usize) void {
const t = &tss_table[cpu];
t.* = .{};
t.ist1 = if (cpu == 0) @intFromPtr(&bsp_ist_stack) + ist_stack_size else ap_ist_top[cpu];
t.iomap_base = @sizeOf(Tss); // == limit: no I/O permission bitmap
gdt.setTssFor(cpu, @intFromPtr(t), @sizeOf(Tss) - 1);
load_tr(gdt.tss_selector);
}
/// Set up the bootstrap processor's TSS (slot 0). Requires gdt.init first.
pub fn init() void {
setupThisCpu(0);
}