Move to multi arch support and added memory map
This commit is contained in:
@@ -1,4 +1,9 @@
|
|||||||
DanOS
|
# DanOS
|
||||||
====
|
Codename: Shodan
|
||||||
|
Version: 1
|
||||||
|
|
||||||
Operating System written for me
|
Operating System written for me
|
||||||
|
|
||||||
|
## Logo
|
||||||
|
|
||||||
|
San Serif Text "Dan OS" with a black karate belt around it.
|
||||||
|
|||||||
@@ -11,6 +11,13 @@ pub fn build(b: *std.Build) void {
|
|||||||
.root_source_file = b.path("src/root.zig"),
|
.root_source_file = b.path("src/root.zig"),
|
||||||
});
|
});
|
||||||
|
|
||||||
|
// Architecture-specific kernel code (CPU ops, entry, later GDT/IDT/paging).
|
||||||
|
// The generic kernel imports this as "arch" and never names x86_64, so a new
|
||||||
|
// architecture is a matter of pointing this module at a different directory.
|
||||||
|
const arch_mod = b.addModule("arch", .{
|
||||||
|
.root_source_file = b.path("src/arch/x86_64/cpu.zig"),
|
||||||
|
});
|
||||||
|
|
||||||
// --- Kernel: freestanding x86_64 ELF, jumped to by the bootloader ---
|
// --- Kernel: freestanding x86_64 ELF, jumped to by the bootloader ---
|
||||||
// SSE2 is part of the x86_64 baseline and UEFI leaves it enabled at handoff,
|
// SSE2 is part of the x86_64 baseline and UEFI leaves it enabled at handoff,
|
||||||
// so we keep it: disabling it forces soft-float and makes the compiler unable
|
// so we keep it: disabling it forces soft-float and makes the compiler unable
|
||||||
@@ -35,10 +42,11 @@ pub fn build(b: *std.Build) void {
|
|||||||
.stack_protector = false,
|
.stack_protector = false,
|
||||||
.imports = &.{
|
.imports = &.{
|
||||||
.{ .name = "danos", .module = mod },
|
.{ .name = "danos", .module = mod },
|
||||||
|
.{ .name = "arch", .module = arch_mod },
|
||||||
},
|
},
|
||||||
}),
|
}),
|
||||||
});
|
});
|
||||||
exe.setLinkerScript(b.path("src/linker.ld"));
|
exe.setLinkerScript(b.path("src/arch/x86_64/linker.ld"));
|
||||||
exe.entry = .{ .symbol_name = "_start" };
|
exe.entry = .{ .symbol_name = "_start" };
|
||||||
// Physical address the bootloader loads the kernel to (identity-mapped under
|
// Physical address the bootloader loads the kernel to (identity-mapped under
|
||||||
// UEFI). Overrides Zig's default image base so the linker script's layout is
|
// UEFI). Overrides Zig's default image base so the linker script's layout is
|
||||||
|
|||||||
+18
-6
@@ -13,22 +13,34 @@ rather than restate it. Roughly in the order things happen at runtime:
|
|||||||
3. **[framebuffer.md](framebuffer.md) — the framebuffer.** What the linear
|
3. **[framebuffer.md](framebuffer.md) — the framebuffer.** What the linear
|
||||||
framebuffer the loader hands over actually is, and what **pitch** (stride)
|
framebuffer the loader hands over actually is, and what **pitch** (stride)
|
||||||
means versus width — the detail you have to get right to avoid a skewed image.
|
means versus width — the detail you have to get right to avoid a skewed image.
|
||||||
4. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
|
4. **[memory-map.md](memory-map.md) — the memory map.** How the loader learns what
|
||||||
|
physical RAM exists and hands it to the kernel in danos's own neutral format,
|
||||||
|
rather than leaking UEFI's memory descriptors across the boundary.
|
||||||
|
5. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
|
||||||
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
|
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
|
||||||
|
|
||||||
|
Cutting across all of these:
|
||||||
|
|
||||||
|
- **[arch.md](arch.md) — the architecture split.** How CPU-specific code is kept
|
||||||
|
behind a build-time `arch` module so the generic kernel never names x86_64,
|
||||||
|
leaving room for other systems (e.g. an AArch64 Raspberry Pi) later.
|
||||||
|
|
||||||
## How the pieces relate
|
## How the pieces relate
|
||||||
|
|
||||||
The boot flow ties them together: UEFI runs the loader ([efi.md](efi.md)), which
|
The boot flow ties them together: UEFI runs the loader ([efi.md](efi.md)), which
|
||||||
queries the **GOP** to pick a graphics mode ([gop.md](gop.md)) and hands the
|
queries the **GOP** to pick a graphics mode ([gop.md](gop.md)), hands the kernel a
|
||||||
kernel a **framebuffer** to draw into ([framebuffer.md](framebuffer.md)); when the
|
**framebuffer** to draw into ([framebuffer.md](framebuffer.md)) and a **memory
|
||||||
kernel has finished — or panics — it **halts** ([halting.md](halting.md)).
|
map** of physical RAM ([memory-map.md](memory-map.md)); the kernel runs — its
|
||||||
|
CPU-specific bits behind the [arch](arch.md) boundary — and when it has finished,
|
||||||
|
or panics, it **halts** ([halting.md](halting.md)).
|
||||||
|
|
||||||
## Source map
|
## Source map
|
||||||
|
|
||||||
| Area | Code |
|
| Area | Code |
|
||||||
|------|------|
|
|------|------|
|
||||||
| Bootloader (UEFI app) | `src/efi.zig` → `BOOTX64.efi` |
|
| Bootloader (UEFI app) | `src/efi.zig` → `BOOTX64.efi` |
|
||||||
| Kernel entry, panic, halt | `src/main.zig` |
|
| Kernel entry, panic, memory-map read-out | `src/main.zig` |
|
||||||
| Shared loader↔kernel contract (`BootInfo`, `Framebuffer`, ABI) | `src/root.zig` |
|
| Shared loader↔kernel contract (`BootInfo`, `Framebuffer`, `MemoryMap`, ABI) | `src/root.zig` |
|
||||||
| Framebuffer text console | `src/console.zig` |
|
| Framebuffer text console | `src/console.zig` |
|
||||||
|
| Arch-specific kernel code (`halt`, linker script) | `src/arch/x86_64/` |
|
||||||
| Build + `run-efi` (QEMU/OVMF) | `build.zig` |
|
| Build + `run-efi` (QEMU/OVMF) | `build.zig` |
|
||||||
|
|||||||
@@ -0,0 +1,91 @@
|
|||||||
|
# Architecture split
|
||||||
|
|
||||||
|
danos targets x86_64 today, but is meant to grow onto other systems later — a
|
||||||
|
Raspberry Pi, say, which is AArch64 and has no UEFI. To keep that possible without
|
||||||
|
a rewrite, CPU-specific kernel code lives behind a boundary: the generic kernel
|
||||||
|
never names an architecture, and each architecture plugs in behind it.
|
||||||
|
|
||||||
|
## The seam is a build-time module named `arch`
|
||||||
|
|
||||||
|
The mechanism is deliberately boring — no vtables, no function-pointer tables, no
|
||||||
|
runtime dispatch. `build.zig` exposes one architecture's code as a module called
|
||||||
|
`arch`:
|
||||||
|
|
||||||
|
```zig
|
||||||
|
const arch_mod = b.addModule("arch", .{
|
||||||
|
.root_source_file = b.path("src/arch/x86_64/cpu.zig"),
|
||||||
|
});
|
||||||
|
```
|
||||||
|
|
||||||
|
and the generic kernel imports it by that name:
|
||||||
|
|
||||||
|
```zig
|
||||||
|
const arch = @import("arch");
|
||||||
|
// ...
|
||||||
|
arch.halt(); // never says "x86_64"
|
||||||
|
```
|
||||||
|
|
||||||
|
Adding a second architecture is then a build-time choice: create
|
||||||
|
`src/arch/aarch64/`, and point the `arch` module at it when the target CPU is
|
||||||
|
AArch64. `main.zig` and `console.zig` don't change. **That compiler-checked module
|
||||||
|
boundary _is_ the architecture interface** — when a new arch is missing a function
|
||||||
|
the generic kernel calls, the build fails and names exactly what's missing.
|
||||||
|
|
||||||
|
## What's arch-specific vs generic
|
||||||
|
|
||||||
|
The split follows a simple test: does it name a CPU instruction, a hardware
|
||||||
|
register, or a memory-management structure? If so, it's arch-specific.
|
||||||
|
|
||||||
|
| Arch-specific — `src/arch/x86_64/` | Generic — kernel core |
|
||||||
|
|---|---|
|
||||||
|
| `cpu.zig`: `halt()` (`hlt`), later GDT/IDT/paging | `console.zig` — pure pixel math, works anywhere |
|
||||||
|
| `linker.ld` — link layout, load address | `main.zig` — `kmain` orchestration, panic handler |
|
||||||
|
| (future) interrupt controller, MMU setup | `root.zig` — the neutral handoff contract |
|
||||||
|
|
||||||
|
Notice the framebuffer console is *generic*: it just writes pixels into whatever
|
||||||
|
framebuffer it's handed, so it needs no per-arch version. Most of the kernel
|
||||||
|
should end up on the generic side; the arch module stays small.
|
||||||
|
|
||||||
|
## Two axes, kept separate
|
||||||
|
|
||||||
|
There are really two independent questions, and it's worth not conflating them:
|
||||||
|
|
||||||
|
- **CPU architecture** (x86_64 vs AArch64): instructions, MMU, interrupts →
|
||||||
|
`src/arch/<cpu>/`.
|
||||||
|
- **Boot protocol** (UEFI vs Raspberry Pi firmware + device tree): handled
|
||||||
|
*separately*, because loaders are their own binaries. `src/efi.zig` builds
|
||||||
|
`BOOTX64.efi`, a distinct executable from the kernel ELF. On a Pi there is no
|
||||||
|
separate loader at all — the firmware jumps straight into the kernel with a
|
||||||
|
device-tree pointer, so that entry work would live in the AArch64 arch code.
|
||||||
|
Either path converges on the same neutral [`BootInfo`](memory-map.md).
|
||||||
|
|
||||||
|
## Current x86_64 contents
|
||||||
|
|
||||||
|
- **`src/arch/x86_64/cpu.zig`** — the `arch` module root. Exposes `halt()` (see
|
||||||
|
[halting.md](halting.md)); GDT, IDT and paging will join it here as the kernel
|
||||||
|
grows.
|
||||||
|
- **`src/arch/x86_64/linker.ld`** — the kernel link layout (fixed low load
|
||||||
|
address, one PT_LOAD per permission set).
|
||||||
|
|
||||||
|
The kernel entry point `_start` currently still lives in the generic `main.zig` as
|
||||||
|
a thin trampoline into `kmain`. It's arch-adjacent (its calling convention is
|
||||||
|
x86_64 SysV, via the shared `danos.kernel_abi`), but it's three lines and mostly
|
||||||
|
generic, so it stays put for now. When AArch64 arrives — where entry means setting
|
||||||
|
up a stack and reading a device-tree pointer from a register — the entry work will
|
||||||
|
be substantial and per-arch, and *that* is when we extract an entry interface into
|
||||||
|
the arch modules.
|
||||||
|
|
||||||
|
## The discipline
|
||||||
|
|
||||||
|
The thing that makes this help rather than hurt: **only extract what's provably
|
||||||
|
architecture-specific, and let the interface emerge with the second
|
||||||
|
implementation.** With a single architecture you're guessing at the seam, and a
|
||||||
|
wrong guess encoded as elaborate abstraction is expensive to undo. So:
|
||||||
|
|
||||||
|
- Move code into `arch/` only when it genuinely names CPU-specific machinery.
|
||||||
|
- Grow the `arch` surface one function at a time, as steps need it.
|
||||||
|
- Don't pre-design the interrupt or paging interfaces before writing them.
|
||||||
|
|
||||||
|
Directory hygiene is cheap and reversible; premature abstraction is neither. When
|
||||||
|
arch #2 lands and something doesn't fit, reshaping a few hundred lines is nothing —
|
||||||
|
unwinding an abstraction empire is not.
|
||||||
+16
-13
@@ -15,12 +15,14 @@ safely, until the machine is reset or powered off.
|
|||||||
|
|
||||||
## The core of it: `hlt`
|
## The core of it: `hlt`
|
||||||
|
|
||||||
Everything comes down to one x86 instruction. In `src/main.zig`:
|
Everything comes down to one x86 instruction. It's CPU-specific, so it lives in
|
||||||
|
the arch module, `src/arch/x86_64/cpu.zig` (see [arch.md](arch.md)), and the
|
||||||
|
generic kernel calls it as `arch.halt()`:
|
||||||
|
|
||||||
```zig
|
```zig
|
||||||
/// Stop the CPU. `hlt` in a loop parks the core at near-zero power until the
|
/// Park the core forever. `hlt` drops it into a low-power idle until the next
|
||||||
/// next interrupt; we loop because `hlt` returns when one arrives.
|
/// interrupt; the loop re-halts on every wake so the stop is permanent.
|
||||||
fn hang() noreturn {
|
pub fn halt() noreturn {
|
||||||
while (true) asm volatile ("hlt");
|
while (true) asm volatile ("hlt");
|
||||||
}
|
}
|
||||||
```
|
```
|
||||||
@@ -71,7 +73,7 @@ makes us robust to all of them.)
|
|||||||
|
|
||||||
## `noreturn`: telling the compiler it's the end
|
## `noreturn`: telling the compiler it's the end
|
||||||
|
|
||||||
`hang()` is typed `noreturn` — a real Zig type meaning "this function never gives
|
`halt()` is typed `noreturn` — a real Zig type meaning "this function never gives
|
||||||
control back to its caller." That isn't decoration; it changes how the compiler
|
control back to its caller." That isn't decoration; it changes how the compiler
|
||||||
treats the call:
|
treats the call:
|
||||||
|
|
||||||
@@ -82,26 +84,27 @@ treats the call:
|
|||||||
jumps back out.
|
jumps back out.
|
||||||
|
|
||||||
You can see the chain in `src/main.zig`: `_start` is `noreturn`, it calls
|
You can see the chain in `src/main.zig`: `_start` is `noreturn`, it calls
|
||||||
`kmain` which is `noreturn`, which ends by calling `hang()` which is `noreturn`.
|
`kmain` which is `noreturn`, which ends by calling `arch.halt()` which is
|
||||||
The "never returns" property is threaded all the way down.
|
`noreturn`. The "never returns" property is threaded all the way down.
|
||||||
|
|
||||||
## Where danos halts
|
## Where danos halts
|
||||||
|
|
||||||
There are three halt sites, and they're all the same idea:
|
There are three halt sites, and they're all the same idea:
|
||||||
|
|
||||||
1. **Normal end of kernel work** — `kmain` prints its status, then calls `hang()`:
|
1. **Normal end of kernel work** — `kmain` prints its status, then calls
|
||||||
|
`arch.halt()`:
|
||||||
|
|
||||||
```zig
|
```zig
|
||||||
con.write("\nkernel initialised; nothing left to do, halting.\n");
|
con.write("\nkernel initialised; nothing left to do, halting.\n");
|
||||||
hang();
|
arch.halt();
|
||||||
```
|
```
|
||||||
|
|
||||||
There's genuinely nothing more to do yet, so the kernel parks itself.
|
There's genuinely nothing more to do yet, so the kernel parks itself.
|
||||||
|
|
||||||
2. **Kernel panic** — the freestanding panic handler has no OS to report to, so
|
2. **Kernel panic** — the freestanding panic handler has no OS to report to, so
|
||||||
it prints the message in red (if the console is up) and halts via the same
|
it prints the message in red (if the console is up) and halts via the same
|
||||||
`hang()`. A panic is unrecoverable here, so stopping the machine — rather than
|
`arch.halt()`. A panic is unrecoverable here, so stopping the machine — rather
|
||||||
limping on with corrupted state — is the safe response.
|
than limping on with corrupted state — is the safe response.
|
||||||
|
|
||||||
3. **Bootloader failure** — in `src/efi.zig`, if `boot()` fails *before* handing
|
3. **Bootloader failure** — in `src/efi.zig`, if `boot()` fails *before* handing
|
||||||
off to the kernel, `main` logs the error and parks the machine with the same
|
off to the kernel, `main` logs the error and parks the machine with the same
|
||||||
@@ -116,8 +119,8 @@ There are three halt sites, and they're all the same idea:
|
|||||||
};
|
};
|
||||||
```
|
```
|
||||||
|
|
||||||
(Here it's an inline loop rather than `hang()` because `hang` lives in the
|
(Here it's an inline loop rather than `arch.halt()` because that lives in the
|
||||||
kernel, which the loader is a separate binary from.)
|
kernel's arch module, and the loader is a separate binary from the kernel.)
|
||||||
|
|
||||||
## Summary
|
## Summary
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,144 @@
|
|||||||
|
# The memory map
|
||||||
|
|
||||||
|
Before a kernel can manage memory, it has to *know what memory exists*: which
|
||||||
|
physical address ranges are real RAM it may use, and which are firmware, hardware
|
||||||
|
registers, or already occupied. That inventory is the **memory map**, and the
|
||||||
|
firmware is the only thing that knows it. This page covers how danos gets that map
|
||||||
|
from the firmware and hands it to the kernel — deliberately without dragging UEFI
|
||||||
|
into the kernel.
|
||||||
|
|
||||||
|
## Why not just pass UEFI's map through?
|
||||||
|
|
||||||
|
UEFI hands the loader a perfectly good memory map. The tempting shortcut is to
|
||||||
|
forward it to the kernel as-is. We don't, for two reasons:
|
||||||
|
|
||||||
|
1. **It would tie the kernel to UEFI.** The kernel would compare against UEFI's
|
||||||
|
memory-type numbers and walk the array using UEFI's variable descriptor stride.
|
||||||
|
That's UEFI vocabulary bleeding across the handoff — and danos wants to boot on
|
||||||
|
systems that have no UEFI at all (a Raspberry Pi describes its memory with a
|
||||||
|
*device tree* instead). See [arch.md](arch.md) for the same "keep the kernel
|
||||||
|
platform-agnostic" principle applied to CPU code.
|
||||||
|
2. **We already established the better pattern.** The loader doesn't hand the
|
||||||
|
kernel a raw UEFI GOP either — [`queryFramebuffer`](gop.md) converts it to
|
||||||
|
danos's own `Framebuffer`. The memory map follows the same discipline.
|
||||||
|
|
||||||
|
So the boundary is: **each boot path translates its native memory description into
|
||||||
|
danos's own neutral format, and the kernel only ever sees that.**
|
||||||
|
|
||||||
|
## The neutral format
|
||||||
|
|
||||||
|
Defined in `src/root.zig`, the shared loader↔kernel contract:
|
||||||
|
|
||||||
|
```zig
|
||||||
|
pub const MemoryKind = enum(u32) {
|
||||||
|
usable, // free RAM the kernel may allocate
|
||||||
|
reserved, // firmware / MMIO / kernel image — never hand out
|
||||||
|
reclaimable, // usable once boot-time structures are done with
|
||||||
|
acpi_tables, // parse, then reclaim
|
||||||
|
acpi_nvs, // preserve across sleep
|
||||||
|
};
|
||||||
|
|
||||||
|
pub const MemoryRegion = extern struct {
|
||||||
|
base: u64, // physical start
|
||||||
|
pages: u64, // length in page_size (4 KiB) units
|
||||||
|
kind: MemoryKind,
|
||||||
|
_pad: u32 = 0,
|
||||||
|
};
|
||||||
|
|
||||||
|
pub const MemoryMap = extern struct {
|
||||||
|
regions: usize, // pointer to a [len]MemoryRegion
|
||||||
|
len: usize,
|
||||||
|
};
|
||||||
|
```
|
||||||
|
|
||||||
|
`MemoryKind` is danos's *own* vocabulary — not UEFI's ~15 types, just the
|
||||||
|
distinctions the kernel actually acts on. And because danos defines `MemoryRegion`
|
||||||
|
itself, `@sizeOf` is authoritative: the kernel walks a plain `[]MemoryRegion` with
|
||||||
|
no variable-stride subtlety (that stride problem is a UEFI-ism, and it stays in the
|
||||||
|
loader).
|
||||||
|
|
||||||
|
`BootInfo` carries it alongside the framebuffer:
|
||||||
|
|
||||||
|
```zig
|
||||||
|
pub const BootInfo = extern struct {
|
||||||
|
framebuffer: Framebuffer,
|
||||||
|
memory_map: MemoryMap,
|
||||||
|
};
|
||||||
|
```
|
||||||
|
|
||||||
|
## The loader side (UEFI)
|
||||||
|
|
||||||
|
Two functions in `src/efi.zig`, called from `exitBootServices`:
|
||||||
|
|
||||||
|
- **`classify`** maps each UEFI memory type to a `MemoryKind`:
|
||||||
|
`conventional_memory → usable`; `boot_services_code`/`boot_services_data →
|
||||||
|
reclaimable` (free once we've exited); `acpi_reclaim_memory → acpi_tables`;
|
||||||
|
`acpi_memory_nvs → acpi_nvs`; **everything else → reserved** (the safe default).
|
||||||
|
Our own `loader_data` — the kernel image and these buffers — falls into
|
||||||
|
`reserved`, so it won't be handed out until the kernel deliberately reclaims it.
|
||||||
|
- **`convertMemoryMap`** walks the UEFI descriptors (striding by
|
||||||
|
`descriptor_size`, *not* `@sizeOf`), classifies each, and writes danos
|
||||||
|
`MemoryRegion`s into an output buffer, coalescing adjacent same-kind regions.
|
||||||
|
|
||||||
|
### The ordering that makes it correct
|
||||||
|
|
||||||
|
This is the fiddly part, dictated by two UEFI rules: you can only allocate memory
|
||||||
|
*before* `ExitBootServices`, and the memory map is only final *at* the moment you
|
||||||
|
exit (its "key" proves you've seen the latest state). So `exitBootServices` does,
|
||||||
|
per attempt:
|
||||||
|
|
||||||
|
1. `getMemoryMapInfo` to size things, then `allocatePool` **two** LoaderData
|
||||||
|
buffers — one for the raw UEFI map, one for the converted regions. Allocating
|
||||||
|
now, before exit, is mandatory.
|
||||||
|
2. `getMemoryMap` then `exitBootServices(key)`. If either fails (allocating can
|
||||||
|
perturb the map and invalidate the key), free both buffers and retry.
|
||||||
|
3. **After** the exit succeeds, convert. Conversion is pure computation on memory
|
||||||
|
we already hold — no boot-services calls — so it's safe once services are gone.
|
||||||
|
|
||||||
|
Both buffers are `LoaderData`, which survives `ExitBootServices`, so the converted
|
||||||
|
array the kernel is pointed at stays valid. (The raw UEFI buffer is just scratch
|
||||||
|
for the conversion.)
|
||||||
|
|
||||||
|
## The kernel side
|
||||||
|
|
||||||
|
The kernel receives a plain array and reads it with zero UEFI knowledge:
|
||||||
|
|
||||||
|
```zig
|
||||||
|
const mm = boot_info.memory_map;
|
||||||
|
const regions = @as([*]const danos.MemoryRegion, @ptrFromInt(mm.regions))[0..mm.len];
|
||||||
|
for (regions) |r| {
|
||||||
|
if (r.kind == .usable) usable_pages += r.pages;
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Today `kmain` just prints the region count and total usable RAM — enough to prove
|
||||||
|
the handoff works. Booted in QEMU with 128 MiB, it reports something like:
|
||||||
|
|
||||||
|
```
|
||||||
|
mem regions: 35
|
||||||
|
usable RAM : 77 MiB
|
||||||
|
```
|
||||||
|
|
||||||
|
with the balance being `reclaimable` boot-services memory (~44 MiB) and a large
|
||||||
|
`reserved` span that is mostly MMIO address space, not RAM. Those figures summing
|
||||||
|
back to ~128 MiB is the sanity check that nothing was dropped.
|
||||||
|
|
||||||
|
## How Raspberry Pi will fit
|
||||||
|
|
||||||
|
No UEFI there, but the boundary is unchanged. The Pi's firmware jumps into the
|
||||||
|
kernel with a **device-tree blob**; the AArch64 entry code will parse its
|
||||||
|
`/memory` and `/reserved-memory` nodes and produce the *same* `MemoryRegion`
|
||||||
|
array. The kernel's memory code — the frame allocator and everything above it —
|
||||||
|
never knows the difference.
|
||||||
|
|
||||||
|
## What's next (not done here)
|
||||||
|
|
||||||
|
This is plumbing plus classification only. Still to come:
|
||||||
|
|
||||||
|
- A **physical frame allocator** that consumes `usable` regions and hands out
|
||||||
|
4 KiB frames — the foundation everything else stands on.
|
||||||
|
- Reclaiming `reclaimable` regions, and carefully freeing `reserved` `loader_data`
|
||||||
|
(kernel image, these buffers) once the kernel is done reading them.
|
||||||
|
- Paging / the kernel's own page tables, then a heap.
|
||||||
|
|
||||||
|
See the roadmap in [efi.md](efi.md) for where this sits in the boot flow.
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
//! x86_64 CPU operations. This is the "arch" module: the generic kernel imports
|
||||||
|
//! it as `@import("arch")` and never names x86_64 directly, so a second
|
||||||
|
//! architecture is added by pointing that module at a different directory in
|
||||||
|
//! build.zig — no change to the generic code. Keep everything CPU-specific here
|
||||||
|
//! (halt now; GDT, IDT and paging will join it), and nothing generic.
|
||||||
|
|
||||||
|
/// Park the core forever. `hlt` drops it into a low-power idle until the next
|
||||||
|
/// interrupt; the loop re-halts on every wake so the stop is permanent. See
|
||||||
|
/// docs/halting.md for the full reasoning.
|
||||||
|
pub fn halt() noreturn {
|
||||||
|
while (true) asm volatile ("hlt");
|
||||||
|
}
|
||||||
+70
-10
@@ -5,6 +5,7 @@ const danos = @import("danos");
|
|||||||
const BootInfo = danos.BootInfo;
|
const BootInfo = danos.BootInfo;
|
||||||
const GraphicsOutput = uefi.protocol.GraphicsOutput;
|
const GraphicsOutput = uefi.protocol.GraphicsOutput;
|
||||||
const EdidActive = uefi.protocol.edid.Active;
|
const EdidActive = uefi.protocol.edid.Active;
|
||||||
|
const MemoryMapSlice = uefi.tables.MemoryMapSlice;
|
||||||
|
|
||||||
/// Name of the kernel ELF on the boot volume (installed to the ESP root by
|
/// Name of the kernel ELF on the boot volume (installed to the ESP root by
|
||||||
/// build.zig). UEFI wants a UTF-16, null-terminated path.
|
/// build.zig). UEFI wants a UTF-16, null-terminated path.
|
||||||
@@ -34,12 +35,13 @@ fn boot() !noreturn {
|
|||||||
// services, since afterwards none of these calls are usable.
|
// services, since afterwards none of these calls are usable.
|
||||||
var boot_info: BootInfo = .{
|
var boot_info: BootInfo = .{
|
||||||
.framebuffer = try queryFramebuffer(bs),
|
.framebuffer = try queryFramebuffer(bs),
|
||||||
|
.memory_map = undefined, // filled by exitBootServices, just below
|
||||||
};
|
};
|
||||||
|
|
||||||
const entry = try loadKernel(bs);
|
const entry = try loadKernel(bs);
|
||||||
|
|
||||||
log("danos: kernel loaded, exiting boot services\r\n");
|
log("danos: kernel loaded, exiting boot services\r\n");
|
||||||
try exitBootServices(bs);
|
boot_info.memory_map = try exitBootServices(bs);
|
||||||
|
|
||||||
// Hand control to the kernel. `danos.kernel_abi` is SysV, so the pointer is
|
// Hand control to the kernel. `danos.kernel_abi` is SysV, so the pointer is
|
||||||
// passed in RDI as the kernel expects — not RCX, which this UEFI binary's
|
// passed in RDI as the kernel expects — not RCX, which this UEFI binary's
|
||||||
@@ -217,27 +219,85 @@ fn loadElf(bs: *uefi.tables.BootServices, image: []u8) !usize {
|
|||||||
return @intCast(ehdr.e_entry);
|
return @intCast(ehdr.e_entry);
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Fetch the memory map and exit boot services. Allocating the map buffer can
|
/// Fetch the memory map, exit boot services, and hand back the map in danos's
|
||||||
/// itself change the map (invalidating the key), so retry until it takes.
|
/// neutral form. Allocating the buffers can itself change the map (invalidating
|
||||||
fn exitBootServices(bs: *uefi.tables.BootServices) !void {
|
/// the key), so retry until it takes. Both buffers are LoaderData, which survives
|
||||||
|
/// the exit, so the returned map stays valid for the kernel.
|
||||||
|
fn exitBootServices(bs: *uefi.tables.BootServices) !danos.MemoryMap {
|
||||||
var attempts: usize = 0;
|
var attempts: usize = 0;
|
||||||
while (attempts < 8) : (attempts += 1) {
|
while (attempts < 8) : (attempts += 1) {
|
||||||
const info = try bs.getMemoryMapInfo();
|
const info = try bs.getMemoryMapInfo();
|
||||||
// Spare descriptors to absorb the growth from the allocatePool below.
|
// Spare descriptors to absorb the growth from the allocations below.
|
||||||
const buf = try bs.allocatePool(.loader_data, (info.len + 8) * info.descriptor_size);
|
const cap = info.len + 8;
|
||||||
const map = bs.getMemoryMap(buf) catch {
|
const map_buf = try bs.allocatePool(.loader_data, cap * info.descriptor_size);
|
||||||
_ = bs.freePool(buf.ptr) catch {};
|
const regions_buf = try bs.allocatePool(.loader_data, cap * @sizeOf(danos.MemoryRegion));
|
||||||
|
const map = bs.getMemoryMap(map_buf) catch {
|
||||||
|
_ = bs.freePool(map_buf.ptr) catch {};
|
||||||
|
_ = bs.freePool(regions_buf.ptr) catch {};
|
||||||
continue;
|
continue;
|
||||||
};
|
};
|
||||||
bs.exitBootServices(uefi.handle, map.info.key) catch {
|
bs.exitBootServices(uefi.handle, map.info.key) catch {
|
||||||
_ = bs.freePool(buf.ptr) catch {};
|
_ = bs.freePool(map_buf.ptr) catch {};
|
||||||
|
_ = bs.freePool(regions_buf.ptr) catch {};
|
||||||
continue;
|
continue;
|
||||||
};
|
};
|
||||||
return; // Boot services are gone; do not touch `bs` again.
|
// Boot services are gone; do not touch `bs` again. Converting the map is
|
||||||
|
// pure computation on memory we already hold, so it's safe here.
|
||||||
|
return convertMemoryMap(map, regions_buf);
|
||||||
}
|
}
|
||||||
return error.ExitBootServicesFailed;
|
return error.ExitBootServicesFailed;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Translate UEFI's memory map into danos's neutral `MemoryRegion` array, written
|
||||||
|
/// into `out` (sized for at least `map.info.len` regions). Adjacent regions of
|
||||||
|
/// the same kind are coalesced. This is the loader's job precisely so the kernel
|
||||||
|
/// never sees UEFI's vocabulary — the same seam the framebuffer already uses.
|
||||||
|
fn convertMemoryMap(map: MemoryMapSlice, out: []u8) danos.MemoryMap {
|
||||||
|
const regions: [*]danos.MemoryRegion = @ptrCast(@alignCast(out.ptr));
|
||||||
|
var count: usize = 0;
|
||||||
|
var i: usize = 0;
|
||||||
|
while (i < map.info.len) : (i += 1) {
|
||||||
|
// Stride by descriptor_size, NOT @sizeOf — firmware descriptors may be
|
||||||
|
// larger than the struct.
|
||||||
|
const d: *const uefi.tables.MemoryDescriptor =
|
||||||
|
@ptrCast(@alignCast(map.ptr + i * map.info.descriptor_size));
|
||||||
|
if (d.number_of_pages == 0) continue;
|
||||||
|
const kind = classify(d.@"type");
|
||||||
|
|
||||||
|
// Coalesce with the previous region if it's the same kind and contiguous.
|
||||||
|
if (count > 0) {
|
||||||
|
const prev = ®ions[count - 1];
|
||||||
|
if (prev.kind == kind and
|
||||||
|
prev.base + prev.pages * danos.page_size == d.physical_start)
|
||||||
|
{
|
||||||
|
prev.pages += d.number_of_pages;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
regions[count] = .{
|
||||||
|
.base = d.physical_start,
|
||||||
|
.pages = d.number_of_pages,
|
||||||
|
.kind = kind,
|
||||||
|
};
|
||||||
|
count += 1;
|
||||||
|
}
|
||||||
|
return .{ .regions = @intFromPtr(regions), .len = count };
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Map a UEFI memory type to danos's neutral kind. Anything we don't explicitly
|
||||||
|
/// recognise is treated as `reserved` — the safe default. Our own loader data
|
||||||
|
/// (the kernel image, these buffers) is LoaderData, which falls here too and so
|
||||||
|
/// stays reserved until the kernel decides to reclaim it.
|
||||||
|
fn classify(t: uefi.tables.MemoryType) danos.MemoryKind {
|
||||||
|
return switch (t) {
|
||||||
|
.conventional_memory => .usable,
|
||||||
|
.boot_services_code, .boot_services_data => .reclaimable,
|
||||||
|
.acpi_reclaim_memory => .acpi_tables,
|
||||||
|
.acpi_memory_nvs => .acpi_nvs,
|
||||||
|
else => .reserved,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
/// Write a compile-time string to the console (best effort).
|
/// Write a compile-time string to the console (best effort).
|
||||||
fn log(comptime msg: []const u8) void {
|
fn log(comptime msg: []const u8) void {
|
||||||
const out = uefi.system_table.con_out orelse return;
|
const out = uefi.system_table.con_out orelse return;
|
||||||
|
|||||||
+14
-8
@@ -1,5 +1,6 @@
|
|||||||
const std = @import("std");
|
const std = @import("std");
|
||||||
const danos = @import("danos");
|
const danos = @import("danos");
|
||||||
|
const arch = @import("arch");
|
||||||
const console = @import("console.zig");
|
const console = @import("console.zig");
|
||||||
const BootInfo = danos.BootInfo;
|
const BootInfo = danos.BootInfo;
|
||||||
|
|
||||||
@@ -33,15 +34,20 @@ fn kmain(boot_info: *const BootInfo) noreturn {
|
|||||||
con.print(" pitch : {d} bytes\n", .{fb.pitch});
|
con.print(" pitch : {d} bytes\n", .{fb.pitch});
|
||||||
con.print(" format : {s}\n", .{@tagName(fb.format)});
|
con.print(" format : {s}\n", .{@tagName(fb.format)});
|
||||||
con.print(" framebuffer: 0x{x:0>16}\n", .{fb.base});
|
con.print(" framebuffer: 0x{x:0>16}\n", .{fb.base});
|
||||||
|
|
||||||
|
// Summarise the physical memory the loader handed us. The array is danos's
|
||||||
|
// own MemoryRegion, so this is a plain slice — no firmware layout in sight.
|
||||||
|
const regions = @as([*]const danos.MemoryRegion, @ptrFromInt(boot_info.memory_map.regions))[0..boot_info.memory_map.len];
|
||||||
|
var usable_pages: u64 = 0;
|
||||||
|
for (regions) |r| {
|
||||||
|
if (r.kind == .usable) usable_pages += r.pages;
|
||||||
|
}
|
||||||
|
con.print(" mem regions: {d}\n", .{regions.len});
|
||||||
|
con.print(" usable RAM : {d} MiB\n", .{usable_pages * danos.page_size / (1024 * 1024)});
|
||||||
|
|
||||||
con.write("\nkernel initialised; nothing left to do, halting.\n");
|
con.write("\nkernel initialised; nothing left to do, halting.\n");
|
||||||
|
|
||||||
hang();
|
arch.halt();
|
||||||
}
|
|
||||||
|
|
||||||
/// Stop the CPU. `hlt` in a loop parks the core at near-zero power until the
|
|
||||||
/// next interrupt; we loop because `hlt` returns when one arrives.
|
|
||||||
fn hang() noreturn {
|
|
||||||
while (true) asm volatile ("hlt");
|
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Freestanding has no OS to receive a panic. Print it to the console (if it is
|
/// Freestanding has no OS to receive a panic. Print it to the console (if it is
|
||||||
@@ -55,6 +61,6 @@ pub const panic = std.debug.FullPanic(struct {
|
|||||||
con.write(msg);
|
con.write(msg);
|
||||||
con.write("\n");
|
con.write("\n");
|
||||||
}
|
}
|
||||||
hang();
|
arch.halt();
|
||||||
}
|
}
|
||||||
}.panic);
|
}.panic);
|
||||||
|
|||||||
@@ -32,8 +32,49 @@ pub const Framebuffer = extern struct {
|
|||||||
format: PixelFormat,
|
format: PixelFormat,
|
||||||
};
|
};
|
||||||
|
|
||||||
|
/// Page size the memory map is measured in. 4 KiB on every architecture danos
|
||||||
|
/// targets so far.
|
||||||
|
pub const page_size = 4096;
|
||||||
|
|
||||||
|
/// danos's own classification of a span of physical memory — deliberately not
|
||||||
|
/// UEFI's vocabulary. Each boot path (UEFI now, device tree later) translates its
|
||||||
|
/// native memory description into these kinds, so the kernel never learns what
|
||||||
|
/// booted it. [[arch]] keeps the same discipline for CPU code.
|
||||||
|
pub const MemoryKind = enum(u32) {
|
||||||
|
/// Free RAM the kernel may allocate.
|
||||||
|
usable,
|
||||||
|
/// Firmware, MMIO, the kernel image, our own boot buffers — never hand out.
|
||||||
|
reserved,
|
||||||
|
/// Usable once the kernel is done with boot-time structures (e.g. UEFI boot
|
||||||
|
/// services memory, which is free after ExitBootServices).
|
||||||
|
reclaimable,
|
||||||
|
/// ACPI tables: parse, then reclaim.
|
||||||
|
acpi_tables,
|
||||||
|
/// ACPI non-volatile storage: preserve across sleep, do not allocate.
|
||||||
|
acpi_nvs,
|
||||||
|
};
|
||||||
|
|
||||||
|
/// One contiguous span of physical memory. Because danos defines this layout
|
||||||
|
/// itself (unlike the UEFI descriptor it's built from), `@sizeOf` is
|
||||||
|
/// authoritative — the kernel walks a plain `[]MemoryRegion`, with none of the
|
||||||
|
/// firmware's variable descriptor-stride to worry about.
|
||||||
|
pub const MemoryRegion = extern struct {
|
||||||
|
base: u64, // physical start address
|
||||||
|
pages: u64, // length in `page_size` units
|
||||||
|
kind: MemoryKind,
|
||||||
|
_pad: u32 = 0,
|
||||||
|
};
|
||||||
|
|
||||||
|
/// The physical memory layout handed to the kernel: a pointer to an array of
|
||||||
|
/// `len` `MemoryRegion`s, in a buffer that outlives the loader.
|
||||||
|
pub const MemoryMap = extern struct {
|
||||||
|
regions: usize, // address of a `[len]MemoryRegion`
|
||||||
|
len: usize,
|
||||||
|
};
|
||||||
|
|
||||||
/// Handoff structure the bootloader fills in and passes to the kernel's
|
/// Handoff structure the bootloader fills in and passes to the kernel's
|
||||||
/// `_start` in RDI (the first argument under the SysV AMD64 C ABI).
|
/// `_start` in RDI (the first argument under the SysV AMD64 C ABI).
|
||||||
pub const BootInfo = extern struct {
|
pub const BootInfo = extern struct {
|
||||||
framebuffer: Framebuffer,
|
framebuffer: Framebuffer,
|
||||||
|
memory_map: MemoryMap,
|
||||||
};
|
};
|
||||||
|
|||||||
Reference in New Issue
Block a user