Move to multi arch support and added memory map
This commit is contained in:
+18
-6
@@ -13,22 +13,34 @@ rather than restate it. Roughly in the order things happen at runtime:
|
||||
3. **[framebuffer.md](framebuffer.md) — the framebuffer.** What the linear
|
||||
framebuffer the loader hands over actually is, and what **pitch** (stride)
|
||||
means versus width — the detail you have to get right to avoid a skewed image.
|
||||
4. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
|
||||
4. **[memory-map.md](memory-map.md) — the memory map.** How the loader learns what
|
||||
physical RAM exists and hands it to the kernel in danos's own neutral format,
|
||||
rather than leaking UEFI's memory descriptors across the boundary.
|
||||
5. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
|
||||
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
|
||||
|
||||
Cutting across all of these:
|
||||
|
||||
- **[arch.md](arch.md) — the architecture split.** How CPU-specific code is kept
|
||||
behind a build-time `arch` module so the generic kernel never names x86_64,
|
||||
leaving room for other systems (e.g. an AArch64 Raspberry Pi) later.
|
||||
|
||||
## How the pieces relate
|
||||
|
||||
The boot flow ties them together: UEFI runs the loader ([efi.md](efi.md)), which
|
||||
queries the **GOP** to pick a graphics mode ([gop.md](gop.md)) and hands the
|
||||
kernel a **framebuffer** to draw into ([framebuffer.md](framebuffer.md)); when the
|
||||
kernel has finished — or panics — it **halts** ([halting.md](halting.md)).
|
||||
queries the **GOP** to pick a graphics mode ([gop.md](gop.md)), hands the kernel a
|
||||
**framebuffer** to draw into ([framebuffer.md](framebuffer.md)) and a **memory
|
||||
map** of physical RAM ([memory-map.md](memory-map.md)); the kernel runs — its
|
||||
CPU-specific bits behind the [arch](arch.md) boundary — and when it has finished,
|
||||
or panics, it **halts** ([halting.md](halting.md)).
|
||||
|
||||
## Source map
|
||||
|
||||
| Area | Code |
|
||||
|------|------|
|
||||
| Bootloader (UEFI app) | `src/efi.zig` → `BOOTX64.efi` |
|
||||
| Kernel entry, panic, halt | `src/main.zig` |
|
||||
| Shared loader↔kernel contract (`BootInfo`, `Framebuffer`, ABI) | `src/root.zig` |
|
||||
| Kernel entry, panic, memory-map read-out | `src/main.zig` |
|
||||
| Shared loader↔kernel contract (`BootInfo`, `Framebuffer`, `MemoryMap`, ABI) | `src/root.zig` |
|
||||
| Framebuffer text console | `src/console.zig` |
|
||||
| Arch-specific kernel code (`halt`, linker script) | `src/arch/x86_64/` |
|
||||
| Build + `run-efi` (QEMU/OVMF) | `build.zig` |
|
||||
|
||||
@@ -0,0 +1,91 @@
|
||||
# Architecture split
|
||||
|
||||
danos targets x86_64 today, but is meant to grow onto other systems later — a
|
||||
Raspberry Pi, say, which is AArch64 and has no UEFI. To keep that possible without
|
||||
a rewrite, CPU-specific kernel code lives behind a boundary: the generic kernel
|
||||
never names an architecture, and each architecture plugs in behind it.
|
||||
|
||||
## The seam is a build-time module named `arch`
|
||||
|
||||
The mechanism is deliberately boring — no vtables, no function-pointer tables, no
|
||||
runtime dispatch. `build.zig` exposes one architecture's code as a module called
|
||||
`arch`:
|
||||
|
||||
```zig
|
||||
const arch_mod = b.addModule("arch", .{
|
||||
.root_source_file = b.path("src/arch/x86_64/cpu.zig"),
|
||||
});
|
||||
```
|
||||
|
||||
and the generic kernel imports it by that name:
|
||||
|
||||
```zig
|
||||
const arch = @import("arch");
|
||||
// ...
|
||||
arch.halt(); // never says "x86_64"
|
||||
```
|
||||
|
||||
Adding a second architecture is then a build-time choice: create
|
||||
`src/arch/aarch64/`, and point the `arch` module at it when the target CPU is
|
||||
AArch64. `main.zig` and `console.zig` don't change. **That compiler-checked module
|
||||
boundary _is_ the architecture interface** — when a new arch is missing a function
|
||||
the generic kernel calls, the build fails and names exactly what's missing.
|
||||
|
||||
## What's arch-specific vs generic
|
||||
|
||||
The split follows a simple test: does it name a CPU instruction, a hardware
|
||||
register, or a memory-management structure? If so, it's arch-specific.
|
||||
|
||||
| Arch-specific — `src/arch/x86_64/` | Generic — kernel core |
|
||||
|---|---|
|
||||
| `cpu.zig`: `halt()` (`hlt`), later GDT/IDT/paging | `console.zig` — pure pixel math, works anywhere |
|
||||
| `linker.ld` — link layout, load address | `main.zig` — `kmain` orchestration, panic handler |
|
||||
| (future) interrupt controller, MMU setup | `root.zig` — the neutral handoff contract |
|
||||
|
||||
Notice the framebuffer console is *generic*: it just writes pixels into whatever
|
||||
framebuffer it's handed, so it needs no per-arch version. Most of the kernel
|
||||
should end up on the generic side; the arch module stays small.
|
||||
|
||||
## Two axes, kept separate
|
||||
|
||||
There are really two independent questions, and it's worth not conflating them:
|
||||
|
||||
- **CPU architecture** (x86_64 vs AArch64): instructions, MMU, interrupts →
|
||||
`src/arch/<cpu>/`.
|
||||
- **Boot protocol** (UEFI vs Raspberry Pi firmware + device tree): handled
|
||||
*separately*, because loaders are their own binaries. `src/efi.zig` builds
|
||||
`BOOTX64.efi`, a distinct executable from the kernel ELF. On a Pi there is no
|
||||
separate loader at all — the firmware jumps straight into the kernel with a
|
||||
device-tree pointer, so that entry work would live in the AArch64 arch code.
|
||||
Either path converges on the same neutral [`BootInfo`](memory-map.md).
|
||||
|
||||
## Current x86_64 contents
|
||||
|
||||
- **`src/arch/x86_64/cpu.zig`** — the `arch` module root. Exposes `halt()` (see
|
||||
[halting.md](halting.md)); GDT, IDT and paging will join it here as the kernel
|
||||
grows.
|
||||
- **`src/arch/x86_64/linker.ld`** — the kernel link layout (fixed low load
|
||||
address, one PT_LOAD per permission set).
|
||||
|
||||
The kernel entry point `_start` currently still lives in the generic `main.zig` as
|
||||
a thin trampoline into `kmain`. It's arch-adjacent (its calling convention is
|
||||
x86_64 SysV, via the shared `danos.kernel_abi`), but it's three lines and mostly
|
||||
generic, so it stays put for now. When AArch64 arrives — where entry means setting
|
||||
up a stack and reading a device-tree pointer from a register — the entry work will
|
||||
be substantial and per-arch, and *that* is when we extract an entry interface into
|
||||
the arch modules.
|
||||
|
||||
## The discipline
|
||||
|
||||
The thing that makes this help rather than hurt: **only extract what's provably
|
||||
architecture-specific, and let the interface emerge with the second
|
||||
implementation.** With a single architecture you're guessing at the seam, and a
|
||||
wrong guess encoded as elaborate abstraction is expensive to undo. So:
|
||||
|
||||
- Move code into `arch/` only when it genuinely names CPU-specific machinery.
|
||||
- Grow the `arch` surface one function at a time, as steps need it.
|
||||
- Don't pre-design the interrupt or paging interfaces before writing them.
|
||||
|
||||
Directory hygiene is cheap and reversible; premature abstraction is neither. When
|
||||
arch #2 lands and something doesn't fit, reshaping a few hundred lines is nothing —
|
||||
unwinding an abstraction empire is not.
|
||||
+16
-13
@@ -15,12 +15,14 @@ safely, until the machine is reset or powered off.
|
||||
|
||||
## The core of it: `hlt`
|
||||
|
||||
Everything comes down to one x86 instruction. In `src/main.zig`:
|
||||
Everything comes down to one x86 instruction. It's CPU-specific, so it lives in
|
||||
the arch module, `src/arch/x86_64/cpu.zig` (see [arch.md](arch.md)), and the
|
||||
generic kernel calls it as `arch.halt()`:
|
||||
|
||||
```zig
|
||||
/// Stop the CPU. `hlt` in a loop parks the core at near-zero power until the
|
||||
/// next interrupt; we loop because `hlt` returns when one arrives.
|
||||
fn hang() noreturn {
|
||||
/// Park the core forever. `hlt` drops it into a low-power idle until the next
|
||||
/// interrupt; the loop re-halts on every wake so the stop is permanent.
|
||||
pub fn halt() noreturn {
|
||||
while (true) asm volatile ("hlt");
|
||||
}
|
||||
```
|
||||
@@ -71,7 +73,7 @@ makes us robust to all of them.)
|
||||
|
||||
## `noreturn`: telling the compiler it's the end
|
||||
|
||||
`hang()` is typed `noreturn` — a real Zig type meaning "this function never gives
|
||||
`halt()` is typed `noreturn` — a real Zig type meaning "this function never gives
|
||||
control back to its caller." That isn't decoration; it changes how the compiler
|
||||
treats the call:
|
||||
|
||||
@@ -82,26 +84,27 @@ treats the call:
|
||||
jumps back out.
|
||||
|
||||
You can see the chain in `src/main.zig`: `_start` is `noreturn`, it calls
|
||||
`kmain` which is `noreturn`, which ends by calling `hang()` which is `noreturn`.
|
||||
The "never returns" property is threaded all the way down.
|
||||
`kmain` which is `noreturn`, which ends by calling `arch.halt()` which is
|
||||
`noreturn`. The "never returns" property is threaded all the way down.
|
||||
|
||||
## Where danos halts
|
||||
|
||||
There are three halt sites, and they're all the same idea:
|
||||
|
||||
1. **Normal end of kernel work** — `kmain` prints its status, then calls `hang()`:
|
||||
1. **Normal end of kernel work** — `kmain` prints its status, then calls
|
||||
`arch.halt()`:
|
||||
|
||||
```zig
|
||||
con.write("\nkernel initialised; nothing left to do, halting.\n");
|
||||
hang();
|
||||
arch.halt();
|
||||
```
|
||||
|
||||
There's genuinely nothing more to do yet, so the kernel parks itself.
|
||||
|
||||
2. **Kernel panic** — the freestanding panic handler has no OS to report to, so
|
||||
it prints the message in red (if the console is up) and halts via the same
|
||||
`hang()`. A panic is unrecoverable here, so stopping the machine — rather than
|
||||
limping on with corrupted state — is the safe response.
|
||||
`arch.halt()`. A panic is unrecoverable here, so stopping the machine — rather
|
||||
than limping on with corrupted state — is the safe response.
|
||||
|
||||
3. **Bootloader failure** — in `src/efi.zig`, if `boot()` fails *before* handing
|
||||
off to the kernel, `main` logs the error and parks the machine with the same
|
||||
@@ -116,8 +119,8 @@ There are three halt sites, and they're all the same idea:
|
||||
};
|
||||
```
|
||||
|
||||
(Here it's an inline loop rather than `hang()` because `hang` lives in the
|
||||
kernel, which the loader is a separate binary from.)
|
||||
(Here it's an inline loop rather than `arch.halt()` because that lives in the
|
||||
kernel's arch module, and the loader is a separate binary from the kernel.)
|
||||
|
||||
## Summary
|
||||
|
||||
|
||||
@@ -0,0 +1,144 @@
|
||||
# The memory map
|
||||
|
||||
Before a kernel can manage memory, it has to *know what memory exists*: which
|
||||
physical address ranges are real RAM it may use, and which are firmware, hardware
|
||||
registers, or already occupied. That inventory is the **memory map**, and the
|
||||
firmware is the only thing that knows it. This page covers how danos gets that map
|
||||
from the firmware and hands it to the kernel — deliberately without dragging UEFI
|
||||
into the kernel.
|
||||
|
||||
## Why not just pass UEFI's map through?
|
||||
|
||||
UEFI hands the loader a perfectly good memory map. The tempting shortcut is to
|
||||
forward it to the kernel as-is. We don't, for two reasons:
|
||||
|
||||
1. **It would tie the kernel to UEFI.** The kernel would compare against UEFI's
|
||||
memory-type numbers and walk the array using UEFI's variable descriptor stride.
|
||||
That's UEFI vocabulary bleeding across the handoff — and danos wants to boot on
|
||||
systems that have no UEFI at all (a Raspberry Pi describes its memory with a
|
||||
*device tree* instead). See [arch.md](arch.md) for the same "keep the kernel
|
||||
platform-agnostic" principle applied to CPU code.
|
||||
2. **We already established the better pattern.** The loader doesn't hand the
|
||||
kernel a raw UEFI GOP either — [`queryFramebuffer`](gop.md) converts it to
|
||||
danos's own `Framebuffer`. The memory map follows the same discipline.
|
||||
|
||||
So the boundary is: **each boot path translates its native memory description into
|
||||
danos's own neutral format, and the kernel only ever sees that.**
|
||||
|
||||
## The neutral format
|
||||
|
||||
Defined in `src/root.zig`, the shared loader↔kernel contract:
|
||||
|
||||
```zig
|
||||
pub const MemoryKind = enum(u32) {
|
||||
usable, // free RAM the kernel may allocate
|
||||
reserved, // firmware / MMIO / kernel image — never hand out
|
||||
reclaimable, // usable once boot-time structures are done with
|
||||
acpi_tables, // parse, then reclaim
|
||||
acpi_nvs, // preserve across sleep
|
||||
};
|
||||
|
||||
pub const MemoryRegion = extern struct {
|
||||
base: u64, // physical start
|
||||
pages: u64, // length in page_size (4 KiB) units
|
||||
kind: MemoryKind,
|
||||
_pad: u32 = 0,
|
||||
};
|
||||
|
||||
pub const MemoryMap = extern struct {
|
||||
regions: usize, // pointer to a [len]MemoryRegion
|
||||
len: usize,
|
||||
};
|
||||
```
|
||||
|
||||
`MemoryKind` is danos's *own* vocabulary — not UEFI's ~15 types, just the
|
||||
distinctions the kernel actually acts on. And because danos defines `MemoryRegion`
|
||||
itself, `@sizeOf` is authoritative: the kernel walks a plain `[]MemoryRegion` with
|
||||
no variable-stride subtlety (that stride problem is a UEFI-ism, and it stays in the
|
||||
loader).
|
||||
|
||||
`BootInfo` carries it alongside the framebuffer:
|
||||
|
||||
```zig
|
||||
pub const BootInfo = extern struct {
|
||||
framebuffer: Framebuffer,
|
||||
memory_map: MemoryMap,
|
||||
};
|
||||
```
|
||||
|
||||
## The loader side (UEFI)
|
||||
|
||||
Two functions in `src/efi.zig`, called from `exitBootServices`:
|
||||
|
||||
- **`classify`** maps each UEFI memory type to a `MemoryKind`:
|
||||
`conventional_memory → usable`; `boot_services_code`/`boot_services_data →
|
||||
reclaimable` (free once we've exited); `acpi_reclaim_memory → acpi_tables`;
|
||||
`acpi_memory_nvs → acpi_nvs`; **everything else → reserved** (the safe default).
|
||||
Our own `loader_data` — the kernel image and these buffers — falls into
|
||||
`reserved`, so it won't be handed out until the kernel deliberately reclaims it.
|
||||
- **`convertMemoryMap`** walks the UEFI descriptors (striding by
|
||||
`descriptor_size`, *not* `@sizeOf`), classifies each, and writes danos
|
||||
`MemoryRegion`s into an output buffer, coalescing adjacent same-kind regions.
|
||||
|
||||
### The ordering that makes it correct
|
||||
|
||||
This is the fiddly part, dictated by two UEFI rules: you can only allocate memory
|
||||
*before* `ExitBootServices`, and the memory map is only final *at* the moment you
|
||||
exit (its "key" proves you've seen the latest state). So `exitBootServices` does,
|
||||
per attempt:
|
||||
|
||||
1. `getMemoryMapInfo` to size things, then `allocatePool` **two** LoaderData
|
||||
buffers — one for the raw UEFI map, one for the converted regions. Allocating
|
||||
now, before exit, is mandatory.
|
||||
2. `getMemoryMap` then `exitBootServices(key)`. If either fails (allocating can
|
||||
perturb the map and invalidate the key), free both buffers and retry.
|
||||
3. **After** the exit succeeds, convert. Conversion is pure computation on memory
|
||||
we already hold — no boot-services calls — so it's safe once services are gone.
|
||||
|
||||
Both buffers are `LoaderData`, which survives `ExitBootServices`, so the converted
|
||||
array the kernel is pointed at stays valid. (The raw UEFI buffer is just scratch
|
||||
for the conversion.)
|
||||
|
||||
## The kernel side
|
||||
|
||||
The kernel receives a plain array and reads it with zero UEFI knowledge:
|
||||
|
||||
```zig
|
||||
const mm = boot_info.memory_map;
|
||||
const regions = @as([*]const danos.MemoryRegion, @ptrFromInt(mm.regions))[0..mm.len];
|
||||
for (regions) |r| {
|
||||
if (r.kind == .usable) usable_pages += r.pages;
|
||||
}
|
||||
```
|
||||
|
||||
Today `kmain` just prints the region count and total usable RAM — enough to prove
|
||||
the handoff works. Booted in QEMU with 128 MiB, it reports something like:
|
||||
|
||||
```
|
||||
mem regions: 35
|
||||
usable RAM : 77 MiB
|
||||
```
|
||||
|
||||
with the balance being `reclaimable` boot-services memory (~44 MiB) and a large
|
||||
`reserved` span that is mostly MMIO address space, not RAM. Those figures summing
|
||||
back to ~128 MiB is the sanity check that nothing was dropped.
|
||||
|
||||
## How Raspberry Pi will fit
|
||||
|
||||
No UEFI there, but the boundary is unchanged. The Pi's firmware jumps into the
|
||||
kernel with a **device-tree blob**; the AArch64 entry code will parse its
|
||||
`/memory` and `/reserved-memory` nodes and produce the *same* `MemoryRegion`
|
||||
array. The kernel's memory code — the frame allocator and everything above it —
|
||||
never knows the difference.
|
||||
|
||||
## What's next (not done here)
|
||||
|
||||
This is plumbing plus classification only. Still to come:
|
||||
|
||||
- A **physical frame allocator** that consumes `usable` regions and hands out
|
||||
4 KiB frames — the foundation everything else stands on.
|
||||
- Reclaiming `reclaimable` regions, and carefully freeing `reserved` `loader_data`
|
||||
(kernel image, these buffers) once the kernel is done reading them.
|
||||
- Paging / the kernel's own page tables, then a heap.
|
||||
|
||||
See the roadmap in [efi.md](efi.md) for where this sits in the boot flow.
|
||||
Reference in New Issue
Block a user