diff --git a/docs/README.md b/docs/README.md index 96663c8..6beed55 100644 --- a/docs/README.md +++ b/docs/README.md @@ -129,7 +129,7 @@ Cutting across all of these: (MBR ESP partition + El Torito EFI entry, one embedded image) that Etcher/dd flash to USB or a burner writes to disc — built by an in-repo pure-Python tool, like the FAT image itself. -- **[arch.md](arch.md) — the architecture split.** How CPU-specific code is kept +- **[architecture.md](architecture.md) — the architecture split.** How CPU-specific code is kept behind a build-time `arch` module so the generic kernel never names x86_64, leaving room for other systems (e.g. an AArch64 Raspberry Pi) later. - **[arm.md](arm.md) — ARM targets.** The Raspberry Pi landscape the arch split is @@ -180,7 +180,7 @@ brings up the **heap** for dynamic allocation ([heap.md](heap.md)), starts the **scheduler** ([scheduling.md](scheduling.md)) and the **timer** that preempts it ([device-interrupts.md](device-interrupts.md)) — with tasks blocking, sleeping and passing messages over **[IPC](ipc.md)** channels — runs, its CPU-specific bits -behind the [arch](arch.md) boundary, and when idle, or on a panic, it **halts** +behind the [architecture](architecture.md) boundary, and when idle, or on a panic, it **halts** ([halting.md](halting.md)). Above that line the microkernel proper begins: **discovery** ([discovery.md](discovery.md), diff --git a/docs/acpi.md b/docs/acpi.md index 76c5adf..c43d3ae 100644 --- a/docs/acpi.md +++ b/docs/acpi.md @@ -5,7 +5,7 @@ PCIe config window, the timer, the power registers — in a set of **system description tables** (SDTs). But before it can read any of them, danos has to *find* them, and they aren't at a fixed address. Getting there is a short chain of pointers, and this note explains it — in particular the question it's easy to trip on: **how -does the [platform / device module](arch.md) know where the RSDT is?** +does the [platform / device module](architecture.md) know where the RSDT is?** Short answer: it doesn't receive the RSDT. The firmware hands over the **RSDP**, and the RSDT's address is a field *inside* the RSDP. The platform follows that pointer. @@ -168,5 +168,5 @@ vocabulary, and orderly shutdown — is the power service, [power.md](power.md). ACPI enumeration and events moved to the ring-3 acpi service. - [power.md](power.md) — the domain-named power service the ACPI event side publishes to (button, lid, battery) and its orderly-shutdown path into S5. -- [arch.md](arch.md) — why the kernel reaches the device code through a `platform` +- [architecture.md](architecture.md) — why the kernel reaches the device code through a `platform` module and never names ACPI directly. diff --git a/docs/arch.md b/docs/arch.md deleted file mode 100644 index 693134f..0000000 --- a/docs/arch.md +++ /dev/null @@ -1,105 +0,0 @@ -# Architecture split - -danos targets x86_64 today, but is meant to grow onto other systems later — a -Raspberry Pi, say, which is AArch64 and has no UEFI. To keep that possible without -a rewrite, CPU-specific kernel code lives behind a boundary: the generic kernel -never names an architecture, and each architecture plugs in behind it. - -## The seam is a build-time module named `arch` - -The mechanism is deliberately boring — no vtables, no function-pointer tables, no -runtime dispatch. `build.zig` exposes one architecture's code as a module called -`arch`: - -```zig -const arch_mod = b.addModule("arch", .{ - .root_source_file = b.path("system/kernel/architecture/x86_64/cpu.zig"), -}); -``` - -and the generic kernel imports it by that name: - -```zig -const arch = @import("arch"); -// ... -arch.halt(); // never says "x86_64" -``` - -Adding a second architecture is then a build-time choice: create -`system/kernel/arch/aarch64/`, and point the `arch` module at it when the target CPU is -AArch64. `main.zig` and `console.zig` don't change. **That compiler-checked module -boundary _is_ the architecture interface** — when a new arch is missing a function -the generic kernel calls, the build fails and names exactly what's missing. - -## What's arch-specific vs generic - -The split follows a simple test: does it name a CPU instruction, a hardware -register, or a memory-management structure? If so, it's arch-specific. - -| Arch-specific — `system/kernel/architecture/x86_64/` | Generic — kernel core | -|---|---| -| `cpu.zig`: `halt()` (`hlt`), later GDT/IDT/paging | `console.zig` — pure pixel math, works anywhere | -| `linker.ld` — link layout, load address | `main.zig` — `kmain` orchestration, panic handler | -| (future) interrupt controller, MMU setup | `root.zig` — the neutral handoff contract | - -Notice the framebuffer console is *generic*: it just writes pixels into whatever -framebuffer it's handed, so it needs no per-arch version. Most of the kernel -should end up on the generic side; the arch module stays small. - -## Two axes, kept separate - -There are really two independent questions, and it's worth not conflating them: - -- **CPU architecture** (x86_64 vs AArch64): instructions, MMU, interrupts → - `system/kernel/arch//`. -- **Boot protocol** (UEFI vs Raspberry Pi firmware + device tree): handled - *separately*, because loaders are their own binaries. `boot/efi.zig` builds - `BOOTX64.efi`, a distinct executable from the kernel ELF. On a Pi there is no - separate loader at all — the firmware jumps straight into the kernel with a - device-tree pointer, so that entry work would live in the AArch64 arch code. - Either path converges on the same neutral [`BootInfo`](memory-map.md). - -## Current x86_64 contents - -- **`system/kernel/architecture/x86_64/cpu.zig`** — the `arch` module root. Exposes `halt()` (see - [halting.md](halting.md)), `init()` (bring up the descriptor tables), - `enablePaging()`, `setFaultHandler`, `readCr2`/`readCr3`, and the `CpuState` - trap frame. -- **`system/kernel/architecture/x86_64/gdt.zig`** / **`idt.zig`** / **`tss.zig`** — the GDT, IDT and - TSS plus CPU-exception handling (see [interrupts.md](interrupts.md)). -- **`system/kernel/architecture/x86_64/paging.zig`** — the kernel's page tables (see - [paging.md](paging.md)). -- **`system/kernel/architecture/x86_64/apic.zig`** — the Local APIC and its timer, the source of - device interrupts (see [device-interrupts.md](device-interrupts.md)). -- **`system/kernel/architecture/x86_64/serial.zig`** / **`io.zig`** — the COM1 UART (the kernel's - machine-readable log channel, see [testing.md](testing.md)) and the shared - port-I/O + MSR primitives. -- **`system/kernel/architecture/x86_64/isr.s`** — the exception stubs, the `lgdt`/`lidt`/`ltr` load - helpers, and the context switch (`switch_context` / `task_trampoline`, see - [scheduling.md](scheduling.md)) — real assembly, since Zig inline asm can't - express them. -- **`system/kernel/architecture/x86_64/linker.ld`** — the kernel link layout (fixed low load - address, one PT_LOAD per permission set). - -The kernel entry point `_start` currently still lives in the generic `main.zig` as -a thin trampoline into `kmain`. It's arch-adjacent (its calling convention is -x86_64 [SysV](sysv.md), via the shared `system.kernel_abi`), but it's three lines -and mostly generic, so it stays put for now. When AArch64 arrives — where entry means setting -up a stack and reading a device-tree pointer from a register — the entry work will -be substantial and per-arch, and *that* is when we extract an entry interface into -the arch modules. - -## The discipline - -The thing that makes this help rather than hurt: **only extract what's provably -architecture-specific, and let the interface emerge with the second -implementation.** With a single architecture you're guessing at the seam, and a -wrong guess encoded as elaborate abstraction is expensive to undo. So: - -- Move code into `arch/` only when it genuinely names CPU-specific machinery. -- Grow the `arch` surface one function at a time, as steps need it. -- Don't pre-design the interrupt or paging interfaces before writing them. - -Directory hygiene is cheap and reversible; premature abstraction is neither. When -arch #2 lands and something doesn't fit, reshaping a few hundred lines is nothing — -unwinding an abstraction empire is not. diff --git a/docs/architecture.md b/docs/architecture.md new file mode 100644 index 0000000..8702ff8 --- /dev/null +++ b/docs/architecture.md @@ -0,0 +1,109 @@ +# Architecture split + +danos targets x86_64 today, but is meant to grow onto other systems later — a +Raspberry Pi, say, which is AArch64 and has no UEFI. To keep that possible without +a rewrite, CPU-specific kernel code lives behind a boundary: the generic kernel +never names an architecture, and each architecture plugs in behind it. + +## The seam is a build-time module named `architecture` + +The mechanism is deliberately boring — no vtables, no function-pointer tables, no +runtime dispatch. `build.zig` exposes one architecture's code as a module called +`architecture`: + +```zig +const architecture_module = b.addModule("architecture", .{ + .root_source_file = b.path("system/kernel/architecture/x86_64/cpu.zig"), +}); +``` + +and the generic kernel imports it by that name: + +```zig +const architecture = @import("architecture"); +// ... +architecture.halt(); // never says "x86_64" +``` + +Adding a second architecture is then a build-time choice: create +`system/kernel/architecture/aarch64/`, and point the `architecture` module at it when the target CPU is +AArch64. `kernel.zig` and `console.zig` don't change. **That compiler-checked module +boundary _is_ the architecture interface** — when a new architecture is missing a function +the generic kernel calls, the build fails and names exactly what's missing. + +## What's arch-specific vs generic + +The split follows a simple test: does it name a CPU instruction, a hardware +register, or a memory-management structure? If so, it's arch-specific. + +| Arch-specific — `system/kernel/architecture/x86_64/` | Generic — kernel core | +|---|---| +| `cpu.zig`: CPU state, trap-frame accessors, paging, SMP | `console.zig` — pure pixel math, framebuffer drawing | +| `gdt.zig`, `idt.zig`, `tss.zig` — descriptor tables | `kernel.zig` — kernel orchestration, scheduler, IPC | +| `paging.zig` — page-table setup and management | `process.zig` — process lifecycle, address spaces | +| `apic.zig`, `ioapic.zig` — interrupt controllers | `scheduler.zig` — task scheduling and context switch | +| `serial.zig`, `io.zig` — UART, I/O primitives | `vfs.zig` — filesystem abstraction | +| `isr.s`, `smp.zig` — exceptions, AP bring-up, context switch | `irq.zig`, `ipc*.zig` — interrupt dispatch, messaging | +| `linker.ld` — kernel link layout, load address | | + +Notice the framebuffer console is *generic*: it just writes pixels into whatever +framebuffer it's handed, so it needs no per-arch version. Most of the kernel +should end up on the generic side; the architecture module stays small. + +## Two axes, kept separate + +There are really two independent questions, and it's worth not conflating them: + +- **CPU architecture** (x86_64 vs AArch64): instructions, MMU, interrupts → + `system/kernel/architecture//`. +- **Boot protocol** (UEFI vs Raspberry Pi firmware + device tree): handled + *separately*, because loaders are their own binaries. `boot/efi.zig` builds + `BOOTX64.efi`, a distinct executable from the kernel ELF. On a Pi there is no + separate loader at all — the firmware jumps straight into the kernel with a + device-tree pointer, so that entry work would live in the AArch64 architecture code. + Either path converges on the same neutral [`BootInfo`](memory-map.md). + +## Current x86_64 contents + +- **`system/kernel/architecture/x86_64/cpu.zig`** — the `architecture` module root. Exposes the trap-frame + `CpuState` and accessors, `init()` (bring up the descriptor tables), `enablePaging()`, + `enterUser()`/`userExit()` for ring-0 ↔ ring-3 transitions, address-space management, and SMP + entry points (see [halting.md](halting.md), [interrupts.md](interrupts.md), [paging.md](paging.md), + [scheduling.md](scheduling.md)). +- **`system/kernel/architecture/x86_64/gdt.zig`** / **`idt.zig`** / **`tss.zig`** — the GDT, IDT and + TSS plus CPU-exception handling (see [interrupts.md](interrupts.md)). +- **`system/kernel/architecture/x86_64/paging.zig`** — the kernel's page tables and address-space + management (see [paging.md](paging.md)). +- **`system/kernel/architecture/x86_64/apic.zig`** / **`ioapic.zig`** — the Local APIC, its timer, + and the I/O APIC for device interrupts (see [device-interrupts.md](device-interrupts.md)). +- **`system/kernel/architecture/x86_64/serial.zig`** / **`io.zig`** — the COM1 UART (the kernel's + machine-readable log channel, see [testing.md](testing.md)) and the shared port-I/O + MSR primitives. +- **`system/kernel/architecture/x86_64/smp.zig`** / **`per-cpu.zig`** — application-processor bring-up + and per-CPU state (GS base, system-call entry point, see [scheduling.md](scheduling.md)). +- **`system/kernel/architecture/x86_64/isr.s`** — the exception stubs, the `lgdt`/`lidt`/`ltr` load + helpers, ring-0 ↔ ring-3 transitions, and the context switch — real assembly, since Zig inline asm can't + express them (see [scheduling.md](scheduling.md)). +- **`system/kernel/architecture/x86_64/linker.ld`** — the kernel link layout (fixed low load + address, one PT_LOAD per permission set). + +The kernel entry point `_start` lives in the architecture-specific `isr.s` (x86_64 here). +On x86_64 it sets up the kernel stack in BSS and jumps to `kmain()` in `kernel.zig`. +This is already per-architecture — an AArch64 port would have its own `isr.s` entry +that parses the device-tree pointer from a register and jumps to the same `kmain()`. +The entry interface is minimal and emerges naturally from the [boot-handoff](memory-map.md) +contract both share. + +## The discipline + +The thing that makes this help rather than hurt: **only extract what's provably +architecture-specific, and let the interface emerge with the second +implementation.** With a single architecture you're guessing at the seam, and a +wrong guess encoded as elaborate abstraction is expensive to undo. So: + +- Move code into `architecture/` only when it genuinely names CPU-specific machinery. +- Grow the `architecture` surface one function at a time, as steps need it. +- Don't pre-design the interrupt or paging interfaces before writing them. + +Directory hygiene is cheap and reversible; premature abstraction is neither. When +architecture #2 lands and something doesn't fit, reshaping a few hundred lines is nothing — +unwinding an abstraction empire is not. diff --git a/docs/arm.md b/docs/arm.md index ea87229..db0df85 100644 --- a/docs/arm.md +++ b/docs/arm.md @@ -5,7 +5,7 @@ though — the Pis span **two different CPU architectures** (32-bit `arm` and 64 `aarch64`) and (stock) a different boot protocol from x86-64's UEFI. **danos targets `aarch64` only** (see the decision below); the `arm`/`aarch64` distinction still matters for understanding why. This page maps the landscape so the -[arch split](arch.md) and build system can be planned for it. +[architecture split](architecture.md) and build system can be planned for it. ## `arm` vs `aarch64` — 32-bit vs 64-bit @@ -39,7 +39,7 @@ page-table formats, and calling conventions. Each needs its own `system/kernel/a ## Booting: UEFI is not x86-only -The boot protocol is a **separate axis** from the CPU (see [arch.md](arch.md)): +The boot protocol is a **separate axis** from the CPU (see [architecture.md](architecture.md)): - **UEFI** exists for ARM too — ARM servers require it (SBSA/SBBR), QEMU boots it with **AAVMF** (the AArch64 build of the same EDK2 firmware as x86's OVMF), and @@ -114,7 +114,7 @@ Two routes, mirroring how we test x86-64 with OVMF: ## Related -- [arch.md](arch.md) — the arch-module boundary these targets plug into, and the +- [architecture.md](architecture.md) — the arch-module boundary these targets plug into, and the CPU-arch vs boot-protocol "two axes". - [efi.md](efi.md) — the UEFI loader that carries over to aarch64-UEFI. - [vision.md](vision.md) — why isolated, portable-across-architectures is the goal. diff --git a/docs/device-interrupts.md b/docs/device-interrupts.md index 874d63e..547cc0a 100644 --- a/docs/device-interrupts.md +++ b/docs/device-interrupts.md @@ -10,7 +10,7 @@ back — the same mechanism a scheduler will later use to preempt tasks. The first device we bring up is the **timer**, because it's the simplest: it lives entirely on the CPU's local interrupt controller, needing no external routing. -It's all x86_64-specific, behind the [arch](arch.md) boundary. +It's all x86_64-specific, behind the [architecture](architecture.md) boundary. ## The APIC, not the PIC diff --git a/docs/discovery.md b/docs/discovery.md index 0f1298d..554f75b 100644 --- a/docs/discovery.md +++ b/docs/discovery.md @@ -5,7 +5,7 @@ MMIO addresses, IRQ numbers, CPU count, and interrupt controller it can't just assume. On x86 that description comes from **ACPI** tables; on ARM from a **device tree** (DTB). This note is a design plan, not built yet: *when* danos should tackle it, and *how* to keep it architecture-agnostic — the same discipline the -[memory map](memory-map.md) and [arch split](arch.md) already follow. +[memory map](memory-map.md) and [architecture split](architecture.md) already follow. ## What the kernel assumes today diff --git a/docs/frame-allocator.md b/docs/frame-allocator.md index 5c73194..69567b3 100644 --- a/docs/frame-allocator.md +++ b/docs/frame-allocator.md @@ -10,7 +10,7 @@ per-process memory all ultimately ask the frame allocator for pages. It's **generic kernel code**: it operates on the neutral `system.MemoryRegion` array, so there's no UEFI in it and nothing architecture-specific beyond the 4 KiB -page. (Contrast [arch.md](arch.md), which is where CPU-specific code lives.) +page. (Contrast [architecture.md](architecture.md), which is where CPU-specific code lives.) ## Why a bitmap diff --git a/docs/halting.md b/docs/halting.md index 37b2cc5..9592973 100644 --- a/docs/halting.md +++ b/docs/halting.md @@ -16,7 +16,7 @@ safely, until the machine is reset or powered off. ## The core of it: `hlt` Everything comes down to one x86 instruction. It's CPU-specific, so it lives in -the arch module, `system/kernel/architecture/x86_64/cpu.zig` (see [arch.md](arch.md)), and the +the arch module, `system/kernel/architecture/x86_64/cpu.zig` (see [architecture.md](architecture.md)), and the generic kernel calls it as `arch.halt()`: ```zig diff --git a/docs/interrupts.md b/docs/interrupts.md index 12809d1..a2a5193 100644 --- a/docs/interrupts.md +++ b/docs/interrupts.md @@ -8,7 +8,7 @@ which on real hardware and in QEMU means a silent reset. Debugging by spontaneou reboot is miserable. This is the machinery that catches those faults and prints what happened instead. -It's all x86_64-specific, so it lives behind the [arch](arch.md) boundary in +It's all x86_64-specific, so it lives behind the [architecture](architecture.md) boundary in `system/kernel/architecture/x86_64/`. Only the 32 CPU-defined exception vectors are wired up so far; device interrupts (timer, keyboard, via the APIC) come later, on the same IDT. diff --git a/docs/memory-map.md b/docs/memory-map.md index bd1d2b7..1b5771e 100644 --- a/docs/memory-map.md +++ b/docs/memory-map.md @@ -16,7 +16,7 @@ forward it to the kernel as-is. We don't, for two reasons: memory-type numbers and walk the array using UEFI's variable descriptor stride. That's UEFI vocabulary bleeding across the handoff — and danos wants to boot on systems that have no UEFI at all (a Raspberry Pi describes its memory with a - *device tree* instead). See [arch.md](arch.md) for the same "keep the kernel + *device tree* instead). See [architecture.md](architecture.md) for the same "keep the kernel platform-agnostic" principle applied to CPU code. 2. **We already established the better pattern.** The loader doesn't hand the kernel a raw UEFI GOP either — [`queryFramebuffer`](gop.md) converts it to diff --git a/docs/paging.md b/docs/paging.md index 5ecbc47..5fbb4e0 100644 --- a/docs/paging.md +++ b/docs/paging.md @@ -7,7 +7,7 @@ which live in memory we'd like to reclaim and don't control), switches CR3 onto them, and — crucially — maps with **real permissions**. It's x86_64-specific (the 4-level table format is an Intel/AMD thing), so it lives -behind the [arch](arch.md) boundary in `system/kernel/architecture/x86_64/paging.zig`. +behind the [architecture](architecture.md) boundary in `system/kernel/architecture/x86_64/paging.zig`. ## The format diff --git a/docs/scheduling.md b/docs/scheduling.md index 6aa6d5d..9d63fcc 100644 --- a/docs/scheduling.md +++ b/docs/scheduling.md @@ -7,7 +7,7 @@ chosen for [real-time](vision.md) — it's predictable (you can reason about whi task runs when) and its decisions are O(1), unlike a fair-share scheduler. The scheduler proper (`system/kernel/sched.zig`) is generic; the context switch and new-task -stack setup are architecture-specific (`system/kernel/architecture/x86_64/`, see [arch](arch.md)). +stack setup are architecture-specific (`system/kernel/architecture/x86_64/`, see [architecture](architecture.md)). ## Tasks diff --git a/docs/sysv.md b/docs/sysv.md index da8bf7c..2f2d377 100644 --- a/docs/sysv.md +++ b/docs/sysv.md @@ -116,7 +116,7 @@ has to speak that same convention at the point of the call. ## A note on other architectures -This is x86-64-specific. An AArch64 port ([arch.md](arch.md)) has its own calling +This is x86-64-specific. An AArch64 port ([architecture.md](architecture.md)) has its own calling convention (arguments in X0–X7, and so on) — a different ABI entirely. `kernel_abi` would be set per-architecture, but the *principle* is the same: the loader/entry boundary and the kernel must agree on how arguments are passed. diff --git a/docs/testing.md b/docs/testing.md index d3187a8..8d1a6e7 100644 --- a/docs/testing.md +++ b/docs/testing.md @@ -24,7 +24,7 @@ appears on serial as plain text. QEMU captures that with `-serial file:serial.log`, giving a machine-readable transcript. Serial is per-architecture (x86 uses port I/O; an ARM board uses a -memory-mapped UART), so it lives behind the [arch](arch.md) boundary — and adding +memory-mapped UART), so it lives behind the [architecture](architecture.md) boundary — and adding a new architecture's UART is what makes the same tests run there. The serial log sink is **compiled in only under `-Dserial`** (off by default). diff --git a/docs/vdso.md b/docs/vdso.md index de478bb..c536f41 100644 --- a/docs/vdso.md +++ b/docs/vdso.md @@ -72,7 +72,7 @@ into every process's address space. The blob is: - **Architecture-specific.** The x86-64 blob wraps `syscall`; an aarch64 blob wraps `svc #0`. It lives beside the other per-architecture kernel sources (`system/kernel/architecture//`), selected the same way the - `architecture` module is (docs/arch.md). + `architecture` module is (docs/architecture.md). ### Shape: a function table, not an ELF diff --git a/docs/vision.md b/docs/vision.md index 31a70c9..41dd62d 100644 --- a/docs/vision.md +++ b/docs/vision.md @@ -99,7 +99,7 @@ own page tables — plus a [test harness](testing.md). resource cleanup on death, then a restartable driver as proof. Needs isolation. See [resilience.md](resilience.md). - **ARM track** — the `aarch64` port so danos runs on the Zero 2 W and Pi 5. Largely - independent of the others (it's the [arch layer](arch.md)); directly serves the win + independent of the others (it's the [architecture layer](architecture.md)); directly serves the win condition. Likely via aarch64-UEFI first (QEMU `virt` + AAVMF), then real boards. See [arm.md](arm.md), and [discovery.md](discovery.md) for the device tree it needs. - **GUI track** — a framebuffer-based windowing/compositor, and the input + display