setup framebuffer with console

This commit is contained in:
2026-07-03 09:32:15 +01:00
parent ca3c3d131c
commit 21cc0c87e4
6 changed files with 415 additions and 16 deletions
+162
View File
@@ -0,0 +1,162 @@
# EFI / The Boot Process
## What EFI is
**UEFI** (Unified Extensible Firmware Interface) is the software baked into your
machine's flash chip that runs the instant it powers on — the modern successor
to the legacy BIOS. Its job is to bring the hardware up to a sane state and then
find and launch an operating system. From our point of view it's a small runtime
that hands us a working CPU, a memory map, and a screen, and then gets out of the
way.
The key thing to understand: **UEFI is not our OS, it's a stepping stone.** It
exists to load *us*. Our `src/efi.zig` is a UEFI *application* — a normal program
that the firmware runs — and its entire purpose is to gather what the kernel needs
and then jump into the kernel.
## How the firmware finds us
UEFI boots by looking for a FAT-formatted partition called the **EFI System
Partition (ESP)** and running a file at a well-known fallback path:
```
esp/EFI/BOOT/BOOTX64.efi <- the "removable media" default for x86-64
```
That's exactly the layout `build.zig` assembles. It builds `src/efi.zig` for the
`uefi` target, installs it to `esp/EFI/BOOT/BOOTX64.efi`, and drops the kernel ELF
at `esp/danos`. The `run-efi` step then points QEMU at OVMF (UEFI firmware for
virtual machines) and presents that `esp/` directory to the guest as a FAT drive.
The firmware finds `BOOTX64.efi` and runs it — that's our `main()`.
## Boot services: the firmware's API
While a UEFI app runs, it has access to **boot services** — a table of function
pointers the firmware provides for allocating memory, reading files, locating
hardware protocols, and so on. In `boot()` this is the very first thing we grab:
```zig
const bs = uefi.system_table.boot_services orelse return error.NoBootServices;
```
Everything the firmware offers hangs off tables reachable from
`uefi.system_table`: `boot_services`, `con_out` (the text console we `log()` to),
and the various *protocols* (GOP for graphics, SimpleFileSystem for disk access).
**The critical rule:** boot services are *temporary*. They stop existing the
moment we call `ExitBootServices`. So the loader's structure is dictated by one
constraint — **gather everything the kernel could ever need first, then exit.**
The comment in `boot()` says exactly this:
> Everything the kernel needs must be gathered *before* we exit boot services,
> since afterwards none of these calls are usable.
## What our loader actually does
`boot()` runs four steps in order:
### 1. Query the framebuffer (`queryFramebuffer`)
We ask the firmware for the **Graphics Output Protocol (GOP)**, which describes
the linear framebuffer — its address, resolution, pitch, and pixel format. We
copy those facts into our own `Framebuffer` struct. This *must* happen now,
because after exit there's no GOP to ask. (See [framebuffer.md](framebuffer.md)
for what those fields mean.)
### 2. Load the kernel (`loadKernel` + `loadElf`)
- Use the **LoadedImage** protocol to discover which device we booted from, then
**SimpleFileSystem** to open that volume.
- Open the file named `danos`, seek to the end to learn its size, rewind, and read
the whole ELF into a firmware-allocated pool buffer. (`read` may return short, so
we loop.)
- Parse the ELF: validate the `\x7fELF` magic and the `x86_64` machine type, then
walk the program headers. For every `PT_LOAD` segment we:
- reserve the exact physical pages it's linked at (`p_paddr`) via
`allocatePages`,
- `@memcpy` the file-backed bytes to that address,
- `@memset` the `.bss` tail (the part where `p_memsz > p_filesz`) to zero.
The kernel is linked to load at physical `0x100000` (1 MiB) — set by
`exe.image_base` in `build.zig` and the linker script. UEFI identity-maps memory,
so the physical address the ELF asks for is the address it actually runs at. If a
segment's `p_paddr` collided with firmware-reserved memory, `allocatePages` would
fail and we'd need to move `image_base`.
`loadElf` returns `e_entry`, the kernel's entry-point address.
### 3. Exit boot services (`exitBootServices`)
This is the handoff's trickiest step. To exit, the firmware demands the current
**memory map** and its *key* — proof that we've seen the latest state of memory.
But allocating the buffer to hold the memory map can itself *change* the map,
invalidating the key. So it's a retry loop:
```
get map info -> allocate buffer (+ spare descriptors) -> get map
-> try exit with map.key
-> if it failed, the map moved: free, retry
```
Once `exitBootServices` succeeds, **the firmware's services are gone for good** —
we must never touch `bs`, `con_out`, or any protocol again. The machine is now
entirely ours.
### 4. Jump to the kernel
```zig
const kernel: *const fn (*const BootInfo) callconv(danos.kernel_abi) noreturn =
@ptrFromInt(entry);
kernel(&boot_info);
```
We cast the entry address to a function pointer and call it, passing a pointer to
the `BootInfo` we filled in. This never returns.
## The ABI subtlety: RCX vs RDI
There's a deliberate detail worth calling out. A UEFI binary is compiled with the
**Microsoft x64** calling convention (first argument in register **RCX**). Our
kernel is freestanding and uses the **SysV AMD64** convention (first argument in
**RDI**). If we let each side use its target's default, the loader would place
`boot_info` in RCX while the kernel looked for it in RDI — and the kernel would
read garbage.
So both sides pin the convention explicitly to SysV via the shared
`danos.kernel_abi` (defined in `src/root.zig`). The loader's function-pointer type
and the kernel's `_start` both reference it, so the pointer lands in the register
the kernel expects. This is the whole reason `kernel_abi` lives in the shared
`danos` module: it's a contract both binaries must agree on.
## The handoff contract
The loader and kernel are two *separate* binaries built for two different targets,
so everything they exchange must have an identically-defined memory layout. That's
what `src/root.zig` provides — imported by both as the `danos` module:
- `BootInfo` — the top-level struct passed to the kernel (currently just the
framebuffer; this is where future handoff data like the memory map will go).
- `Framebuffer`, `PixelFormat`, `kernel_abi` — the shared field layouts and the
calling convention.
Both are `extern struct`, giving them a stable, C-compatible layout so the bytes
the loader writes are the bytes the kernel reads.
## The whole flow at a glance
```
power on
-> UEFI firmware initialises hardware
-> finds esp/EFI/BOOT/BOOTX64.efi, runs it (our efi.zig main)
-> grab boot services
-> queryFramebuffer (via GOP)
-> loadKernel (read danos ELF, load PT_LOAD segments to 0x100000)
-> exitBootServices (retry until the memory-map key holds)
-> jump to e_entry, boot_info pointer in RDI
-> kernel _start (src/main.zig: framebuffer console, then halt)
```
Bottom line: **UEFI's job is to give us a CPU, memory, and a framebuffer, then
disappear.** `src/efi.zig` is the thin bridge that collects those gifts into a
`BootInfo`, tears down the firmware, and jumps into the kernel — after which we're
on our own.
+106
View File
@@ -0,0 +1,106 @@
# The Framebuffer
## What a framebuffer is
A **framebuffer** is just a big region of memory where each element is one
pixel's color. The display hardware continuously scans this memory and turns
each value into light on the screen. There's no drawing API involved — you
write a 32-bit value to the right address, and a pixel changes color. That's
exactly what `Console.pixel` does:
```zig
self.rowPtr(y)[x] = color; // src/console.zig
```
Our `Framebuffer` struct (`src/root.zig`) is the four facts you need to
address it:
| Field | Meaning |
|----------|---------|
| `base` | the memory address where pixel data starts |
| `width` | visible pixels per row (e.g. 1920) |
| `height` | visible rows (e.g. 1080) |
| `pitch` | **bytes** from the start of one row to the start of the next |
The bootloader (UEFI GOP, in our case) sets all this up and hands it over. The
kernel just writes into it: no firmware, no driver — just pixels.
## The mental model: it's 1D memory pretending to be 2D
The screen is a grid, but memory is a flat line of bytes. So the pixels are
stored row after row, laid end to end:
```
row 0: [px0][px1][px2]...[width-1] <padding?>
row 1: [px0][px1][px2]...[width-1] <padding?>
row 2: ...
```
To find pixel `(x, y)` you compute:
```
address = base + y * (bytes per row) + x * (bytes per pixel)
```
## So what is pitch?
**Pitch is "bytes per row"** — sometimes called *stride*. The obvious guess
would be `pitch = width * 4` (4 bytes = 32 bits per pixel). And often it is.
**But not always** — and that's the whole reason the field exists.
Hardware frequently wants each row to start at a nicely aligned address (a
multiple of 32, 64, or a page). If `width` doesn't land on that boundary, the
firmware pads the end of every row with a few extra unused bytes. That padding
is invisible — it's never shown — but it's physically there in memory between
the last pixel of one row and the first pixel of the next.
Example: a 1366-pixel-wide display at 32bpp:
- `width * 4` = 1366 × 4 = **5464 bytes** of actual pixels
- but `pitch` might be **5504 bytes** (padded up to a multiple of 64)
- those extra 40 bytes per row are dead space
This is exactly why `rowPtr` uses `pitch`, not `width`, to step between rows:
```zig
inline fn rowPtr(self: *Console, y: u32) [*]volatile u32 {
const base: [*]volatile u8 = @ptrFromInt(self.fb.base);
return @ptrCast(@alignCast(base + y * self.fb.pitch)); // <- pitch, not width*4
}
```
Note the deliberate detail: `base` is cast to a **byte** pointer
(`[*]volatile u8`) *before* adding `y * pitch`, because pitch is measured in
bytes. Then it's cast to a `u32` pointer so that `[x]` indexes whole pixels. If
you'd done the arithmetic on a `u32` pointer, `+ pitch` would step `pitch`
*pixels* (4× too far).
### Why you must use pitch, not `width * 4`
If you assumed rows were `width * 4` apart on a display where
`pitch > width * 4`, every row would start a little too early. The error
accumulates: row 0 is fine, row 1 is off by (pitch − width×4) bytes, row 2 by
twice that, and so on. The image ends up **skewed diagonally** — a slanted,
sheared picture — because each row creeps sideways relative to where the
hardware actually reads it.
Using `pitch` is what keeps each row landing exactly where the scanout expects
it.
## Two subtleties worth noting
1. **`width` vs `pitch` in the loops.** In `fillRow`/`copyRow` we iterate `x`
up to `self.fb.width` — the *visible* count — but jump between rows with
`pitch`. That's the correct pairing: touch only real pixels, but skip the
full stride (including padding) to reach the next row. We never write into
the padding, which is right.
2. **`volatile`.** The pointer is `volatile` because this memory is special —
it's watched by the display hardware. `volatile` tells the compiler *"don't
optimize these writes away or reorder/coalesce them"*; every store must
actually hit memory, because something outside the CPU's knowledge (the
scanout engine) is reading it.
Bottom line: **width is how wide the picture is; pitch is how wide the memory
rows are.** They're usually equal (×4) but not guaranteed to be, so always
advance rows by pitch.