Move to multi arch support and added memory map

This commit is contained in:
2026-07-03 11:11:19 +01:00
parent c2435760c4
commit 628c4f6d57
11 changed files with 423 additions and 41 deletions
+18 -6
View File
@@ -13,22 +13,34 @@ rather than restate it. Roughly in the order things happen at runtime:
3. **[framebuffer.md](framebuffer.md) — the framebuffer.** What the linear
framebuffer the loader hands over actually is, and what **pitch** (stride)
means versus width — the detail you have to get right to avoid a skewed image.
4. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
4. **[memory-map.md](memory-map.md) — the memory map.** How the loader learns what
physical RAM exists and hands it to the kernel in danos's own neutral format,
rather than leaking UEFI's memory descriptors across the boundary.
5. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
Cutting across all of these:
- **[arch.md](arch.md) — the architecture split.** How CPU-specific code is kept
behind a build-time `arch` module so the generic kernel never names x86_64,
leaving room for other systems (e.g. an AArch64 Raspberry Pi) later.
## How the pieces relate
The boot flow ties them together: UEFI runs the loader ([efi.md](efi.md)), which
queries the **GOP** to pick a graphics mode ([gop.md](gop.md)) and hands the
kernel a **framebuffer** to draw into ([framebuffer.md](framebuffer.md)); when the
kernel has finished — or panics — it **halts** ([halting.md](halting.md)).
queries the **GOP** to pick a graphics mode ([gop.md](gop.md)), hands the kernel a
**framebuffer** to draw into ([framebuffer.md](framebuffer.md)) and a **memory
map** of physical RAM ([memory-map.md](memory-map.md)); the kernel runs — its
CPU-specific bits behind the [arch](arch.md) boundary — and when it has finished,
or panics, it **halts** ([halting.md](halting.md)).
## Source map
| Area | Code |
|------|------|
| Bootloader (UEFI app) | `src/efi.zig` → `BOOTX64.efi` |
| Kernel entry, panic, halt | `src/main.zig` |
| Shared loader↔kernel contract (`BootInfo`, `Framebuffer`, ABI) | `src/root.zig` |
| Kernel entry, panic, memory-map read-out | `src/main.zig` |
| Shared loader↔kernel contract (`BootInfo`, `Framebuffer`, `MemoryMap`, ABI) | `src/root.zig` |
| Framebuffer text console | `src/console.zig` |
| Arch-specific kernel code (`halt`, linker script) | `src/arch/x86_64/` |
| Build + `run-efi` (QEMU/OVMF) | `build.zig` |
+91
View File
@@ -0,0 +1,91 @@
# Architecture split
danos targets x86_64 today, but is meant to grow onto other systems later — a
Raspberry Pi, say, which is AArch64 and has no UEFI. To keep that possible without
a rewrite, CPU-specific kernel code lives behind a boundary: the generic kernel
never names an architecture, and each architecture plugs in behind it.
## The seam is a build-time module named `arch`
The mechanism is deliberately boring — no vtables, no function-pointer tables, no
runtime dispatch. `build.zig` exposes one architecture's code as a module called
`arch`:
```zig
const arch_mod = b.addModule("arch", .{
.root_source_file = b.path("src/arch/x86_64/cpu.zig"),
});
```
and the generic kernel imports it by that name:
```zig
const arch = @import("arch");
// ...
arch.halt(); // never says "x86_64"
```
Adding a second architecture is then a build-time choice: create
`src/arch/aarch64/`, and point the `arch` module at it when the target CPU is
AArch64. `main.zig` and `console.zig` don't change. **That compiler-checked module
boundary _is_ the architecture interface** — when a new arch is missing a function
the generic kernel calls, the build fails and names exactly what's missing.
## What's arch-specific vs generic
The split follows a simple test: does it name a CPU instruction, a hardware
register, or a memory-management structure? If so, it's arch-specific.
| Arch-specific — `src/arch/x86_64/` | Generic — kernel core |
|---|---|
| `cpu.zig`: `halt()` (`hlt`), later GDT/IDT/paging | `console.zig` — pure pixel math, works anywhere |
| `linker.ld` — link layout, load address | `main.zig` — `kmain` orchestration, panic handler |
| (future) interrupt controller, MMU setup | `root.zig` — the neutral handoff contract |
Notice the framebuffer console is *generic*: it just writes pixels into whatever
framebuffer it's handed, so it needs no per-arch version. Most of the kernel
should end up on the generic side; the arch module stays small.
## Two axes, kept separate
There are really two independent questions, and it's worth not conflating them:
- **CPU architecture** (x86_64 vs AArch64): instructions, MMU, interrupts →
`src/arch/<cpu>/`.
- **Boot protocol** (UEFI vs Raspberry Pi firmware + device tree): handled
*separately*, because loaders are their own binaries. `src/efi.zig` builds
`BOOTX64.efi`, a distinct executable from the kernel ELF. On a Pi there is no
separate loader at all — the firmware jumps straight into the kernel with a
device-tree pointer, so that entry work would live in the AArch64 arch code.
Either path converges on the same neutral [`BootInfo`](memory-map.md).
## Current x86_64 contents
- **`src/arch/x86_64/cpu.zig`** — the `arch` module root. Exposes `halt()` (see
[halting.md](halting.md)); GDT, IDT and paging will join it here as the kernel
grows.
- **`src/arch/x86_64/linker.ld`** — the kernel link layout (fixed low load
address, one PT_LOAD per permission set).
The kernel entry point `_start` currently still lives in the generic `main.zig` as
a thin trampoline into `kmain`. It's arch-adjacent (its calling convention is
x86_64 SysV, via the shared `danos.kernel_abi`), but it's three lines and mostly
generic, so it stays put for now. When AArch64 arrives — where entry means setting
up a stack and reading a device-tree pointer from a register — the entry work will
be substantial and per-arch, and *that* is when we extract an entry interface into
the arch modules.
## The discipline
The thing that makes this help rather than hurt: **only extract what's provably
architecture-specific, and let the interface emerge with the second
implementation.** With a single architecture you're guessing at the seam, and a
wrong guess encoded as elaborate abstraction is expensive to undo. So:
- Move code into `arch/` only when it genuinely names CPU-specific machinery.
- Grow the `arch` surface one function at a time, as steps need it.
- Don't pre-design the interrupt or paging interfaces before writing them.
Directory hygiene is cheap and reversible; premature abstraction is neither. When
arch #2 lands and something doesn't fit, reshaping a few hundred lines is nothing —
unwinding an abstraction empire is not.
+16 -13
View File
@@ -15,12 +15,14 @@ safely, until the machine is reset or powered off.
## The core of it: `hlt`
Everything comes down to one x86 instruction. In `src/main.zig`:
Everything comes down to one x86 instruction. It's CPU-specific, so it lives in
the arch module, `src/arch/x86_64/cpu.zig` (see [arch.md](arch.md)), and the
generic kernel calls it as `arch.halt()`:
```zig
/// Stop the CPU. `hlt` in a loop parks the core at near-zero power until the
/// next interrupt; we loop because `hlt` returns when one arrives.
fn hang() noreturn {
/// Park the core forever. `hlt` drops it into a low-power idle until the next
/// interrupt; the loop re-halts on every wake so the stop is permanent.
pub fn halt() noreturn {
while (true) asm volatile ("hlt");
}
```
@@ -71,7 +73,7 @@ makes us robust to all of them.)
## `noreturn`: telling the compiler it's the end
`hang()` is typed `noreturn` — a real Zig type meaning "this function never gives
`halt()` is typed `noreturn` — a real Zig type meaning "this function never gives
control back to its caller." That isn't decoration; it changes how the compiler
treats the call:
@@ -82,26 +84,27 @@ treats the call:
jumps back out.
You can see the chain in `src/main.zig`: `_start` is `noreturn`, it calls
`kmain` which is `noreturn`, which ends by calling `hang()` which is `noreturn`.
The "never returns" property is threaded all the way down.
`kmain` which is `noreturn`, which ends by calling `arch.halt()` which is
`noreturn`. The "never returns" property is threaded all the way down.
## Where danos halts
There are three halt sites, and they're all the same idea:
1. **Normal end of kernel work** — `kmain` prints its status, then calls `hang()`:
1. **Normal end of kernel work** — `kmain` prints its status, then calls
`arch.halt()`:
```zig
con.write("\nkernel initialised; nothing left to do, halting.\n");
hang();
arch.halt();
```
There's genuinely nothing more to do yet, so the kernel parks itself.
2. **Kernel panic** — the freestanding panic handler has no OS to report to, so
it prints the message in red (if the console is up) and halts via the same
`hang()`. A panic is unrecoverable here, so stopping the machine — rather than
limping on with corrupted state — is the safe response.
`arch.halt()`. A panic is unrecoverable here, so stopping the machine — rather
than limping on with corrupted state — is the safe response.
3. **Bootloader failure** — in `src/efi.zig`, if `boot()` fails *before* handing
off to the kernel, `main` logs the error and parks the machine with the same
@@ -116,8 +119,8 @@ There are three halt sites, and they're all the same idea:
};
```
(Here it's an inline loop rather than `hang()` because `hang` lives in the
kernel, which the loader is a separate binary from.)
(Here it's an inline loop rather than `arch.halt()` because that lives in the
kernel's arch module, and the loader is a separate binary from the kernel.)
## Summary
+144
View File
@@ -0,0 +1,144 @@
# The memory map
Before a kernel can manage memory, it has to *know what memory exists*: which
physical address ranges are real RAM it may use, and which are firmware, hardware
registers, or already occupied. That inventory is the **memory map**, and the
firmware is the only thing that knows it. This page covers how danos gets that map
from the firmware and hands it to the kernel — deliberately without dragging UEFI
into the kernel.
## Why not just pass UEFI's map through?
UEFI hands the loader a perfectly good memory map. The tempting shortcut is to
forward it to the kernel as-is. We don't, for two reasons:
1. **It would tie the kernel to UEFI.** The kernel would compare against UEFI's
memory-type numbers and walk the array using UEFI's variable descriptor stride.
That's UEFI vocabulary bleeding across the handoff — and danos wants to boot on
systems that have no UEFI at all (a Raspberry Pi describes its memory with a
*device tree* instead). See [arch.md](arch.md) for the same "keep the kernel
platform-agnostic" principle applied to CPU code.
2. **We already established the better pattern.** The loader doesn't hand the
kernel a raw UEFI GOP either — [`queryFramebuffer`](gop.md) converts it to
danos's own `Framebuffer`. The memory map follows the same discipline.
So the boundary is: **each boot path translates its native memory description into
danos's own neutral format, and the kernel only ever sees that.**
## The neutral format
Defined in `src/root.zig`, the shared loader↔kernel contract:
```zig
pub const MemoryKind = enum(u32) {
usable, // free RAM the kernel may allocate
reserved, // firmware / MMIO / kernel image — never hand out
reclaimable, // usable once boot-time structures are done with
acpi_tables, // parse, then reclaim
acpi_nvs, // preserve across sleep
};
pub const MemoryRegion = extern struct {
base: u64, // physical start
pages: u64, // length in page_size (4 KiB) units
kind: MemoryKind,
_pad: u32 = 0,
};
pub const MemoryMap = extern struct {
regions: usize, // pointer to a [len]MemoryRegion
len: usize,
};
```
`MemoryKind` is danos's *own* vocabulary — not UEFI's ~15 types, just the
distinctions the kernel actually acts on. And because danos defines `MemoryRegion`
itself, `@sizeOf` is authoritative: the kernel walks a plain `[]MemoryRegion` with
no variable-stride subtlety (that stride problem is a UEFI-ism, and it stays in the
loader).
`BootInfo` carries it alongside the framebuffer:
```zig
pub const BootInfo = extern struct {
framebuffer: Framebuffer,
memory_map: MemoryMap,
};
```
## The loader side (UEFI)
Two functions in `src/efi.zig`, called from `exitBootServices`:
- **`classify`** maps each UEFI memory type to a `MemoryKind`:
`conventional_memory → usable`; `boot_services_code`/`boot_services_data →
reclaimable` (free once we've exited); `acpi_reclaim_memory → acpi_tables`;
`acpi_memory_nvs → acpi_nvs`; **everything else → reserved** (the safe default).
Our own `loader_data` — the kernel image and these buffers — falls into
`reserved`, so it won't be handed out until the kernel deliberately reclaims it.
- **`convertMemoryMap`** walks the UEFI descriptors (striding by
`descriptor_size`, *not* `@sizeOf`), classifies each, and writes danos
`MemoryRegion`s into an output buffer, coalescing adjacent same-kind regions.
### The ordering that makes it correct
This is the fiddly part, dictated by two UEFI rules: you can only allocate memory
*before* `ExitBootServices`, and the memory map is only final *at* the moment you
exit (its "key" proves you've seen the latest state). So `exitBootServices` does,
per attempt:
1. `getMemoryMapInfo` to size things, then `allocatePool` **two** LoaderData
buffers — one for the raw UEFI map, one for the converted regions. Allocating
now, before exit, is mandatory.
2. `getMemoryMap` then `exitBootServices(key)`. If either fails (allocating can
perturb the map and invalidate the key), free both buffers and retry.
3. **After** the exit succeeds, convert. Conversion is pure computation on memory
we already hold — no boot-services calls — so it's safe once services are gone.
Both buffers are `LoaderData`, which survives `ExitBootServices`, so the converted
array the kernel is pointed at stays valid. (The raw UEFI buffer is just scratch
for the conversion.)
## The kernel side
The kernel receives a plain array and reads it with zero UEFI knowledge:
```zig
const mm = boot_info.memory_map;
const regions = @as([*]const danos.MemoryRegion, @ptrFromInt(mm.regions))[0..mm.len];
for (regions) |r| {
if (r.kind == .usable) usable_pages += r.pages;
}
```
Today `kmain` just prints the region count and total usable RAM — enough to prove
the handoff works. Booted in QEMU with 128 MiB, it reports something like:
```
mem regions: 35
usable RAM : 77 MiB
```
with the balance being `reclaimable` boot-services memory (~44 MiB) and a large
`reserved` span that is mostly MMIO address space, not RAM. Those figures summing
back to ~128 MiB is the sanity check that nothing was dropped.
## How Raspberry Pi will fit
No UEFI there, but the boundary is unchanged. The Pi's firmware jumps into the
kernel with a **device-tree blob**; the AArch64 entry code will parse its
`/memory` and `/reserved-memory` nodes and produce the *same* `MemoryRegion`
array. The kernel's memory code — the frame allocator and everything above it —
never knows the difference.
## What's next (not done here)
This is plumbing plus classification only. Still to come:
- A **physical frame allocator** that consumes `usable` regions and hands out
4 KiB frames — the foundation everything else stands on.
- Reclaiming `reclaimable` regions, and carefully freeing `reserved` `loader_data`
(kernel image, these buffers) once the kernel is done reading them.
- Paging / the kernel's own page tables, then a heap.
See the roadmap in [efi.md](efi.md) for where this sits in the boot flow.