docs: full docs-vs-code audit — fix every stale claim across 40 docs
Every doc verified claim-by-claim against the code by parallel audit agents, then fixed and adversarially re-verified. Two waves of staleness corrected: the originally audited findings (higher-half boot handoff, kernel VFS takeover, fault isolation + claim release + driver restart, AML/S5 moving to ring 3, threading's shipped design, USB+FAT landing) and a second pass of adjacent claims the verifiers caught (smp.md 'not built yet' intro, system-requirements' PS/2-only and no-storage claims, halting.md's red-panic and no-IDT text, testing.md's serial mirroring, router-era vfs-protocol wording, capsule-first boot loading). threading.md now documents the shared-fate gap explicitly: the design says a process dies whole, the kernel today kills only the offending thread. Also fixes three stale code comments (isr.s exceptionHandler, acpi.zig sleepValue, build.zig boot-volume) — comments only, no behavior change.
This commit is contained in:
+64
-39
@@ -23,14 +23,19 @@ Partition (ESP)** and running a file at a well-known fallback path:
|
||||
EFI/BOOT/BOOTX64.efi <- the "removable media" default for x86-64
|
||||
```
|
||||
|
||||
The boot volume is the **FHS-shaped `zig-out`** itself (see the repository-layout note
|
||||
in [README.md](README.md)): `build.zig` installs `boot/efi.zig` (built for the `uefi`
|
||||
target) to `zig-out/EFI/BOOT/BOOTX64.efi` — the one path UEFI firmware fixes — and lays
|
||||
the rest out by FHS path: the kernel at `zig-out/system/kernel`, init at
|
||||
`zig-out/system/services/init`, the initial-ramdisk at `zig-out/boot/`. The
|
||||
`run-x86-64` step points QEMU at OVMF (UEFI firmware for virtual machines) and presents
|
||||
`zig-out` to the guest as a FAT drive. The firmware finds `BOOTX64.efi` and runs it —
|
||||
that's our `main()`, which then loads the kernel and init from their FHS paths.
|
||||
The boot volume is **FHS-shaped** (see the repository-layout note in
|
||||
[README.md](README.md)): `build.zig` installs `boot/efi.zig` (built for the `uefi`
|
||||
target) at `EFI/BOOT/BOOTX64.efi` — the one path UEFI firmware fixes — and lays
|
||||
the rest out by FHS path: the kernel at `system/kernel`, init at
|
||||
`system/services/init`, the pre-packed boot capsule at `boot/system.img`.
|
||||
`zig-out` mirrors that tree, but what a machine actually boots is the
|
||||
self-contained FAT32 image `tools/make-fat-image.py` builds from the same files
|
||||
(`danos-usb.img`). The `run-x86-64` step points QEMU at OVMF (UEFI firmware for
|
||||
virtual machines) and attaches that image (its serial-logging twin, built the
|
||||
same way) as a USB mass-storage device on the xHCI bus — the guest never sees
|
||||
`zig-out`. The firmware finds `BOOTX64.efi` on the image and runs it — that's
|
||||
our `main()`, which then loads the kernel and the system binaries from their
|
||||
FHS paths.
|
||||
|
||||
## Boot services: the firmware's API
|
||||
|
||||
@@ -56,7 +61,15 @@ The comment in `boot()` says exactly this:
|
||||
|
||||
## What our loader actually does
|
||||
|
||||
`boot()` runs four steps in order:
|
||||
The four milestones below are the spine of `boot()`. Along the way it also
|
||||
captures the **ACPI RSDP** from the UEFI configuration table (while boot
|
||||
services are still up), loads the system binaries into an in-RAM
|
||||
**initial ramdisk** (`loadSystemTree` — normally a single read of the pre-packed
|
||||
`boot\system.img` capsule, which already *is* the ramdisk wire format; it falls
|
||||
back to opening each manifest-listed path, and walks the `/system` tree only as
|
||||
a last resort for hand-assembled sticks. Best-effort either way — a kernel-only
|
||||
volume still boots), and builds the **bootstrap page tables** the kernel starts
|
||||
life on (`buildBootstrapTables`), all before the jump:
|
||||
|
||||
### 1. Query the framebuffer (`queryFramebuffer`)
|
||||
|
||||
@@ -91,16 +104,20 @@ All of this *must* happen now, because after exit there's no GOP to ask. (See
|
||||
may return short, so we loop.)
|
||||
- Parse the ELF: validate the `\x7fELF` magic and the `x86_64` machine type, then
|
||||
walk the program headers. For every `PT_LOAD` segment we:
|
||||
- reserve the exact physical pages it's linked at (`p_paddr`) via
|
||||
- reserve the exact physical pages it asks to be loaded at (`p_paddr`) via
|
||||
`allocatePages`,
|
||||
- `@memcpy` the file-backed bytes to that address,
|
||||
- `@memset` the `.bss` tail (the part where `p_memsz > p_filesz`) to zero.
|
||||
|
||||
The kernel is linked to load at physical `0x100000` (1 MiB) — set by
|
||||
`exe.image_base` in `build.zig` and the linker script. UEFI identity-maps memory,
|
||||
so the physical address the ELF asks for is the address it actually runs at. If a
|
||||
segment's `p_paddr` collided with firmware-reserved memory, `allocatePages` would
|
||||
fail and we'd need to move `image_base`.
|
||||
The kernel is linked to *run* in the higher half (virtual base
|
||||
`0xFFFFFFFF80000000`; `exe.image_base` in `build.zig` is the *virtual* address
|
||||
`0xFFFFFFFF80100000`) but is *loaded* low: the linker script's `AT()` clauses
|
||||
give every segment a low physical load address (`p_paddr`, with `.text` at
|
||||
`0x100000`, 1 MiB), which is what the loader allocates and copies into. The
|
||||
bootstrap page tables built before the jump map the high link addresses onto
|
||||
those low physical pages. If a segment's `p_paddr` collided with
|
||||
firmware-reserved memory, `allocatePages` would fail and we'd need to move the
|
||||
load addresses.
|
||||
|
||||
`loadElf` returns `e_entry`, the kernel's entry-point address.
|
||||
|
||||
@@ -124,13 +141,12 @@ entirely ours.
|
||||
### 4. Jump to the kernel
|
||||
|
||||
```zig
|
||||
const kernel: *const fn (*const BootInfo) callconv(boot_handoff.kernel_abi) noreturn =
|
||||
@ptrFromInt(entry);
|
||||
kernel(&boot_info);
|
||||
handoff(cr3, entry, &boot_information);
|
||||
```
|
||||
|
||||
We cast the entry address to a function pointer and call it, passing a pointer to
|
||||
the `BootInfo` we filled in. This never returns.
|
||||
`handoff` is a single inline-asm block — `cli`, load the bootstrap page tables'
|
||||
`cr3`, place the `boot_information` pointer in RDI, then `callq *entry` — so
|
||||
nothing runs between the CR3 load and the jump. This never returns.
|
||||
|
||||
## The ABI subtlety: RCX vs RDI
|
||||
|
||||
@@ -138,14 +154,15 @@ There's a deliberate detail worth calling out. A UEFI binary is compiled with th
|
||||
**Microsoft x64** calling convention (first argument in register **RCX**). Our
|
||||
kernel is freestanding and uses the **SysV AMD64** convention (first argument in
|
||||
**RDI**). If we let each side use its target's default, the loader would place
|
||||
`boot_info` in RCX while the kernel looked for it in RDI — and the kernel would
|
||||
read garbage.
|
||||
`boot_information` in RCX while the kernel looked for it in RDI — and the kernel
|
||||
would read garbage.
|
||||
|
||||
So both sides pin the convention explicitly to SysV via the shared
|
||||
`boot_handoff.kernel_abi` (defined in `system/boot-handoff.zig`). The loader's
|
||||
function-pointer type and the kernel's `_start` both reference it, so the pointer lands
|
||||
in the register the kernel expects. This is the whole reason `kernel_abi` lives in the
|
||||
shared `boot-handoff` module: it's a contract both binaries must agree on. See
|
||||
So the convention is pinned explicitly to SysV via the shared
|
||||
`boot_handoff.kernel_abi` (defined in `system/boot-handoff.zig`). The kernel's
|
||||
`kmainEntry` declares `callconv(kernel_abi)`; the loader honours the same contract
|
||||
by loading RDI by hand in `handoff`'s inline asm rather than trusting its own
|
||||
Microsoft-x64 default. `kernel_abi` lives in the shared `boot-handoff` module
|
||||
because it's a contract both binaries must agree on. See
|
||||
[sysv.md](sysv.md) for what "SysV" means and where else it shows up.
|
||||
|
||||
## The handoff contract
|
||||
@@ -156,13 +173,16 @@ what `system/boot-handoff.zig` provides — imported by both as the `boot-handof
|
||||
It is *only* the handoff: the kernel↔user ABI (`system/abi.zig`) and the device types
|
||||
(`system/devices/device-abi.zig`) are separate contracts the bootloader never sees.
|
||||
|
||||
- `BootInfo` — the top-level struct passed to the kernel (currently just the
|
||||
framebuffer; this is where future handoff data like the memory map will go).
|
||||
- `Framebuffer`, `PixelFormat`, `kernel_abi` — the shared field layouts and the
|
||||
calling convention.
|
||||
- `BootInformation` — the top-level struct passed to the kernel: the
|
||||
framebuffer, the memory map, the kernel's own `PT_LOAD` segments
|
||||
(`kernel_segments` + `kernel_segment_count`, so the kernel can re-map itself
|
||||
with correct permissions), the ACPI RSDP address, and the initial-ramdisk
|
||||
base/length.
|
||||
- `Framebuffer`, `PixelFormat`, `MemoryMap`/`MemoryRegion`, `KernelSegment`,
|
||||
`kernel_abi` — the shared field layouts and the calling convention.
|
||||
|
||||
Both are `extern struct`, giving them a stable, C-compatible layout so the bytes
|
||||
the loader writes are the bytes the kernel reads.
|
||||
The structs are `extern struct`, giving them a stable, C-compatible layout so the
|
||||
bytes the loader writes are the bytes the kernel reads.
|
||||
|
||||
## The whole flow at a glance
|
||||
|
||||
@@ -170,15 +190,20 @@ the loader writes are the bytes the kernel reads.
|
||||
power on
|
||||
-> UEFI firmware initialises hardware
|
||||
-> finds EFI/BOOT/BOOTX64.efi on the FHS volume, runs it (our efi.zig main)
|
||||
-> grab boot services
|
||||
-> grab boot services (+ the ACPI RSDP from the configuration table)
|
||||
-> queryFramebuffer (via GOP: EDID native res, setMode, describe fb)
|
||||
-> loadKernel (read system/kernel ELF, load PT_LOAD segments to 0x100000)
|
||||
-> loadKernel (read system/kernel ELF, load PT_LOAD segments low, .text at 0x100000)
|
||||
-> loadSystemTree (read the boot\system.img capsule as the in-RAM initial ramdisk;
|
||||
fallbacks: manifest-listed paths, then a /system tree walk)
|
||||
-> buildBootstrapTables (identity + physmap + higher-half kernel mappings)
|
||||
-> exitBootServices (retry until the memory-map key holds)
|
||||
-> jump to e_entry, boot_info pointer in RDI
|
||||
-> kernel _start (system/kernel/kernel.zig: framebuffer console, then halt)
|
||||
-> handoff: load bootstrap CR3, jump to e_entry, boot_information pointer in RDI
|
||||
-> kernel _start (architecture/x86_64/isr.s: switch to a kernel-owned stack,
|
||||
call kmainEntry -> paging, heap, device discovery,
|
||||
scheduler, SMP, user space)
|
||||
```
|
||||
|
||||
Bottom line: **UEFI's job is to give us a CPU, memory, and a framebuffer, then
|
||||
disappear.** `boot/efi.zig` is the thin bridge that collects those gifts into a
|
||||
`BootInfo`, tears down the firmware, and jumps into the kernel — after which we're
|
||||
on our own.
|
||||
`BootInformation`, tears down the firmware, and jumps into the kernel — after
|
||||
which we're on our own.
|
||||
|
||||
Reference in New Issue
Block a user