selecting the displays native resolution
This commit is contained in:
+21
-5
@@ -58,10 +58,26 @@ The comment in `boot()` says exactly this:
|
||||
### 1. Query the framebuffer (`queryFramebuffer`)
|
||||
|
||||
We ask the firmware for the **Graphics Output Protocol (GOP)**, which describes
|
||||
the linear framebuffer — its address, resolution, pitch, and pixel format. We
|
||||
copy those facts into our own `Framebuffer` struct. This *must* happen now,
|
||||
because after exit there's no GOP to ask. (See [framebuffer.md](framebuffer.md)
|
||||
for what those fields mean.)
|
||||
the linear framebuffer — its address, resolution, pitch, and pixel format —
|
||||
and, when we can, switch the display to its native resolution first:
|
||||
|
||||
- Locate GOP via its *handle* (not `locateProtocol`), because the same handle
|
||||
also carries the display's **EDID**.
|
||||
- Read the EDID (trying the `EDID_ACTIVE` then `EDID_DISCOVERED` protocol on each
|
||||
GOP handle) and parse the first Detailed Timing Descriptor — by convention the
|
||||
panel's preferred (native) resolution. This is best-effort: firmware installs
|
||||
these protocols inconsistently, and OVMF with QEMU's stdvga doesn't expose them
|
||||
at all.
|
||||
- If we got a native resolution, enumerate the GOP modes with `queryMode` and
|
||||
`setMode` to the one that matches it (must be a linear 32bpp layout we can
|
||||
paint into). If EDID gave us nothing, we **keep the firmware's current default
|
||||
mode** rather than guess — with a valid EDID the firmware normally defaults to
|
||||
the native mode itself, so its choice beats second-guessing it with, say, the
|
||||
largest advertised mode.
|
||||
- Copy the resulting address/resolution/pitch/format into our own `Framebuffer`.
|
||||
|
||||
All of this *must* happen now, because after exit there's no GOP to ask. (See
|
||||
[framebuffer.md](framebuffer.md) for what those fields mean.)
|
||||
|
||||
### 2. Load the kernel (`loadKernel` + `loadElf`)
|
||||
|
||||
@@ -149,7 +165,7 @@ power on
|
||||
-> UEFI firmware initialises hardware
|
||||
-> finds esp/EFI/BOOT/BOOTX64.efi, runs it (our efi.zig main)
|
||||
-> grab boot services
|
||||
-> queryFramebuffer (via GOP)
|
||||
-> queryFramebuffer (via GOP: EDID native res, setMode, describe fb)
|
||||
-> loadKernel (read danos ELF, load PT_LOAD segments to 0x100000)
|
||||
-> exitBootServices (retry until the memory-map key holds)
|
||||
-> jump to e_entry, boot_info pointer in RDI
|
||||
|
||||
+71
@@ -0,0 +1,71 @@
|
||||
# GOP
|
||||
|
||||
Graphics Output Protocol (GOP) is a UEFI driver interface that replaces legacy VGA BIOS functions to provide graphics console output in the pre-OS phase.
|
||||
|
||||
It allows the firmware to display boot screens and setup menus by providing direct access to the hardware frame buffer, enabling multiple GPUs to function equally without proprietary INT15 handshaking.
|
||||
GOP has no standardized mode numbers. Instead, the firmware (really the GPU's GOP driver) builds a list of modes at boot, and each is just an opaque index 0 .. MaxMode-1. Mode 0 might be 1920×1080 on your laptop and 800×600 on someone else's — the index carries no fixed meaning. To learn what a mode actually is, you have to ask:
|
||||
|
||||
- gop.mode.max_mode — how many modes exist
|
||||
- gop.query_mode(n) — returns the info (resolution, pixel format, pixels_per_scan_line) for mode n
|
||||
- gop.set_mode(n) — switch to it
|
||||
- gop.mode.info — the info for the currently active mode
|
||||
|
||||
queryFramebuffer used to read gop.mode.info directly and never call set_mode — taking whatever mode the firmware selected as its default. It now works out the monitor's native resolution from EDID and, if a matching mode exists, calls set_mode to switch to it. When EDID is unavailable it keeps the firmware's default mode rather than guessing — with a valid EDID present the firmware normally defaults to the native mode itself (this is exactly how the QEMU run boots at 1280x720: OVMF's stdvga driver reads the emulated EDID and defaults to the preferred mode, even though it never exposes the EDID protocols to us). See the "picks the native resolution" note at the bottom.
|
||||
|
||||
Does EFI detect native resolution?
|
||||
|
||||
Sometimes, but it's not guaranteed by the spec. Here's the actual chain:
|
||||
|
||||
1. The GOP driver reads the monitor's EDID over the DDC/I²C wire — a data blob the display publishes describing its supported resolutions, including its preferred (native) timing.
|
||||
2. The firmware then picks a default GOP mode. Many modern firmwares (and OVMF in a VM, driven by the emulated display) do default to the native/preferred resolution. But plenty of firmware defaults to a safe fallback like 1024×768 or 800×600 regardless of what the panel can do.
|
||||
|
||||
So gop.mode.info giving you native res is a common outcome, not a promise. Two more caveats:
|
||||
|
||||
- The mode list itself may not even contain the true native resolution — a GOP driver can expose only a handful of modes.
|
||||
- On a headless VM (your QEMU/OVMF setup), there's no real EDID; the "native" resolution is whatever the emulated GPU advertises. OVMF's default is typically 800×600 or 1024×768 unless you configure it (e.g. QEMU's -device virtio-vga with a set resolution, or the OVMF Platform config).
|
||||
|
||||
One note worth flagging: this picks the native resolution but doesn't force a particular pixel format — a native mode is only chosen if it's a paintable linear 32bpp layout (RGBX/BGRX); a bit_mask/blt_only native mode is skipped and we keep the firmware's default instead. That's the right trade-off for now since your console assumes linear 32bpp.
|
||||
|
||||
## Pixel formats
|
||||
|
||||
Every mode also carries a pixel format, and queryFramebuffer only accepts two of the four GOP formats. The full enum:
|
||||
|
||||
- red_green_blue_reserved_8_bit_per_color (RGBX) — accepted
|
||||
- blue_green_red_reserved_8_bit_per_color (BGRX) — accepted
|
||||
- bit_mask — rejected
|
||||
- blt_only — rejected
|
||||
|
||||
RGBX/BGRX tell you each pixel is a 32-bit value with a fixed byte order. That's what the console needs: a known layout it can write directly with rowPtr(y)[x] = color. The other two break one of those assumptions.
|
||||
|
||||
### bit_mask (spec: PixelBitMask)
|
||||
|
||||
The framebuffer is still linear memory you can write to directly — but the bits aren't in a standard RGBX/BGRX arrangement. Instead the firmware hands you a PixelBitmask struct describing where each channel lives:
|
||||
|
||||
```zig
|
||||
pub const PixelBitmask = extern struct {
|
||||
red_mask: u32,
|
||||
green_mask: u32,
|
||||
blue_mask: u32,
|
||||
reserved_mask: u32,
|
||||
};
|
||||
```
|
||||
|
||||
Each mask marks which bits of the pixel word belong to that channel. This is how the format expresses non-standard layouts — for example 16-bit RGB565 (red = 0xF800, green = 0x07E0, blue = 0x001F: 5/6/5 bits, only 16 bits per pixel), or an odd 32-bit order. To draw a color you'd have to read the masks, work out each channel's bit position and width, shift/scale your 8-bit R/G/B into place, and OR them together — per pixel. It's fully drawable, just not with a hardcoded 32-bit write, so we reject it rather than carry that machinery.
|
||||
|
||||
(RGBX/BGRX are really just two hardcoded special cases of a bitmask. The spec names them separately precisely so simple loaders can skip mask-decoding in the common case.) In practice bit_mask is rare on modern PC firmware — you'll almost always get RGBX or BGRX — so rejecting it costs basically nothing.
|
||||
|
||||
### blt_only (spec: PixelBltOnly)
|
||||
|
||||
This one is more fundamental: there is no linear framebuffer you can address at all. gop.mode.frame_buffer_base is meaningless — you have no pointer to pixel memory. The only way to put pixels on screen is through GOP's Blt ("block transfer") service, the third function pointer in the protocol: you build pixels in your own buffer and ask the firmware to copy ("blit") a rectangle onto the display. The firmware owns the actual scanout memory, wherever it lives (across a bus, behind a GPU command interface, in a layout the CPU can't map directly).
|
||||
|
||||
The catch that matters for us: Blt is a boot service. It stops working the instant you call ExitBootServices — which is exactly when the kernel runs. So a blt_only display gives a post-exit kernel no way to draw pixels at all, and there's genuinely nothing our framebuffer console could do with it. Rejecting it is the only correct response.
|
||||
|
||||
### Summary
|
||||
|
||||
| Format | Linear memory? | Layout | Console |
|
||||
| --- | --- | --- | --- |
|
||||
| rgbx / bgrx | yes | fixed 32bpp byte order | works — direct writes |
|
||||
| bit_mask | yes | arbitrary, described by masks | rejected — would need per-pixel mask decoding |
|
||||
| blt_only | no | no CPU-visible framebuffer; Blt service only | rejected — and Blt is gone after ExitBootServices anyway |
|
||||
|
||||
So the two-case accept list is the right line to draw: RGBX/BGRX are the only formats that give a post-ExitBootServices kernel a flat block of pixel memory it can write to without help from firmware that no longer exists.
|
||||
+133
@@ -0,0 +1,133 @@
|
||||
# Halting
|
||||
|
||||
## Why a kernel needs to halt
|
||||
|
||||
An ordinary program ends by *returning* — `main` finishes, the C runtime calls
|
||||
`exit`, and the OS reclaims the process. A kernel has none of that. There is no
|
||||
OS underneath it, no runtime to return to, and no caller waiting. `_start` is the
|
||||
end of the line. So when the kernel has nothing left to do — whether it finished
|
||||
its work or hit a fatal error — it can't "quit". It has to explicitly park the
|
||||
CPU forever, because if execution ever ran off the end it would just keep fetching
|
||||
whatever bytes follow in memory and execute garbage.
|
||||
|
||||
That's what halting is: deliberately stopping the processor so it does nothing,
|
||||
safely, until the machine is reset or powered off.
|
||||
|
||||
## The core of it: `hlt`
|
||||
|
||||
Everything comes down to one x86 instruction. In `src/main.zig`:
|
||||
|
||||
```zig
|
||||
/// Stop the CPU. `hlt` in a loop parks the core at near-zero power until the
|
||||
/// next interrupt; we loop because `hlt` returns when one arrives.
|
||||
fn hang() noreturn {
|
||||
while (true) asm volatile ("hlt");
|
||||
}
|
||||
```
|
||||
|
||||
**`hlt`** ("halt") tells the CPU core to stop executing instructions and drop into
|
||||
a low-power idle state. It's not a busy-wait — the core genuinely stops, drawing
|
||||
almost no power and generating almost no heat, until something wakes it.
|
||||
|
||||
This is much better than the naïve alternative, a spin loop:
|
||||
|
||||
```zig
|
||||
while (true) {} // "busy-wait" — DON'T do this to idle
|
||||
```
|
||||
|
||||
A bare `while (true) {}` keeps the core running flat out, executing the jump
|
||||
back to the top of the loop billions of times a second — 100% CPU, hot, and (on a
|
||||
laptop) draining the battery, all to accomplish nothing. `hlt` achieves the same
|
||||
"do nothing" outcome while letting the core sleep.
|
||||
|
||||
## Why the loop around it?
|
||||
|
||||
Here's the subtlety the comment points at: **`hlt` is not permanent.** It halts
|
||||
the core only until the *next interrupt* arrives. An interrupt is a signal — from
|
||||
a timer, a keypress, a device — that wakes the CPU so it can respond. When one
|
||||
fires, the core comes out of `hlt` and executes the next instruction.
|
||||
|
||||
If we wrote just a single `hlt`, the very first stray interrupt would wake the
|
||||
core and execution would continue past it — falling off the end of the function
|
||||
into whatever comes next in memory. Wrapping it in `while (true)` closes that
|
||||
door: every time an interrupt wakes the core, the loop immediately runs `hlt`
|
||||
again and it goes back to sleep. The net effect is a permanent halt that still
|
||||
sleeps between the interrupts it can't prevent.
|
||||
|
||||
(At this stage danos hasn't set up an interrupt descriptor table, so most
|
||||
interrupts aren't even something we handle — but non-maskable interrupts and
|
||||
system-management interrupts can still wake a halted core regardless. The loop
|
||||
makes us robust to all of them.)
|
||||
|
||||
## The `asm volatile` part
|
||||
|
||||
`hlt` has no equivalent in plain Zig, so we drop to inline assembly:
|
||||
|
||||
- **`asm`** emits the raw instruction directly into the function.
|
||||
- **`volatile`** tells the compiler *"this has side effects you can't see — do
|
||||
not optimize it away or reorder it."* Without it, an optimizing compiler might
|
||||
reason that the assembly produces no value anyone uses and delete it, or hoist
|
||||
it somewhere wrong. `volatile` pins it exactly where we wrote it.
|
||||
|
||||
## `noreturn`: telling the compiler it's the end
|
||||
|
||||
`hang()` is typed `noreturn` — a real Zig type meaning "this function never gives
|
||||
control back to its caller." That isn't decoration; it changes how the compiler
|
||||
treats the call:
|
||||
|
||||
- Code *after* a `noreturn` call is unreachable, so the compiler needn't emit a
|
||||
return sequence, and won't warn about "missing return value" in the callers.
|
||||
- It lets `kmain` and `_start` themselves be `noreturn`, which is the honest
|
||||
signature for a kernel entry point — the bootloader jumps in and nothing ever
|
||||
jumps back out.
|
||||
|
||||
You can see the chain in `src/main.zig`: `_start` is `noreturn`, it calls
|
||||
`kmain` which is `noreturn`, which ends by calling `hang()` which is `noreturn`.
|
||||
The "never returns" property is threaded all the way down.
|
||||
|
||||
## Where danos halts
|
||||
|
||||
There are three halt sites, and they're all the same idea:
|
||||
|
||||
1. **Normal end of kernel work** — `kmain` prints its status, then calls `hang()`:
|
||||
|
||||
```zig
|
||||
con.write("\nkernel initialised; nothing left to do, halting.\n");
|
||||
hang();
|
||||
```
|
||||
|
||||
There's genuinely nothing more to do yet, so the kernel parks itself.
|
||||
|
||||
2. **Kernel panic** — the freestanding panic handler has no OS to report to, so
|
||||
it prints the message in red (if the console is up) and halts via the same
|
||||
`hang()`. A panic is unrecoverable here, so stopping the machine — rather than
|
||||
limping on with corrupted state — is the safe response.
|
||||
|
||||
3. **Bootloader failure** — in `src/efi.zig`, if `boot()` fails *before* handing
|
||||
off to the kernel, `main` logs the error and parks the machine with the same
|
||||
loop so the message stays on screen:
|
||||
|
||||
```zig
|
||||
boot() catch |err| {
|
||||
log("\r\ndanos: boot failed: ");
|
||||
logBytes(@errorName(err));
|
||||
log("\r\n");
|
||||
while (true) asm volatile ("hlt");
|
||||
};
|
||||
```
|
||||
|
||||
(Here it's an inline loop rather than `hang()` because `hang` lives in the
|
||||
kernel, which the loader is a separate binary from.)
|
||||
|
||||
## Summary
|
||||
|
||||
- A kernel can't "exit" — it must explicitly stop the CPU or it runs off into
|
||||
garbage.
|
||||
- **`hlt`** parks the core in a low-power idle until the next interrupt — far
|
||||
better than a 100%-CPU spin loop.
|
||||
- **`while (true) hlt`** makes that halt permanent, since any interrupt would
|
||||
otherwise wake the core and let execution continue.
|
||||
- **`asm volatile`** emits the instruction and forbids the compiler from removing
|
||||
it; **`noreturn`** encodes "control never comes back" into the type system.
|
||||
- danos halts on normal completion, on a kernel panic, and on a bootloader error
|
||||
— the same "stop safely and stay stopped" in all three.
|
||||
Reference in New Issue
Block a user