docs: full docs-vs-code audit — fix every stale claim across 40 docs

Every doc verified claim-by-claim against the code by parallel audit agents,
then fixed and adversarially re-verified. Two waves of staleness corrected:
the originally audited findings (higher-half boot handoff, kernel VFS
takeover, fault isolation + claim release + driver restart, AML/S5 moving to
ring 3, threading's shipped design, USB+FAT landing) and a second pass of
adjacent claims the verifiers caught (smp.md 'not built yet' intro,
system-requirements' PS/2-only and no-storage claims, halting.md's red-panic
and no-IDT text, testing.md's serial mirroring, router-era vfs-protocol
wording, capsule-first boot loading).

threading.md now documents the shared-fate gap explicitly: the design says a
process dies whole, the kernel today kills only the offending thread.

Also fixes three stale code comments (isr.s exceptionHandler, acpi.zig
sleepValue, build.zig boot-volume) — comments only, no behavior change.
This commit is contained in:
Daniel Samson
2026-07-22 09:09:53 +01:00
parent 52df2ba6f6
commit e854f65623
43 changed files with 821 additions and 546 deletions
+31 -23
View File
@@ -17,7 +17,7 @@ safely, until the machine is reset or powered off.
Everything comes down to one x86 instruction. It's CPU-specific, so it lives in
the arch module, `system/kernel/architecture/x86_64/cpu.zig` (see [architecture.md](architecture.md)), and the
generic kernel calls it as `arch.halt()`:
generic kernel calls it as `architecture.halt()`:
```zig
/// Park the core forever. `hlt` drops it into a low-power idle until the next
@@ -56,10 +56,10 @@ door: every time an interrupt wakes the core, the loop immediately runs `hlt`
again and it goes back to sleep. The net effect is a permanent halt that still
sleeps between the interrupts it can't prevent.
(At this stage danos hasn't set up an interrupt descriptor table, so most
interrupts aren't even something we handle — but non-maskable interrupts and
system-management interrupts can still wake a halted core regardless. The loop
makes us robust to all of them.)
(danos handles plenty of interrupts through its interrupt descriptor table —
the timer tick waking a halted idle core is exactly how scheduling works — and
non-maskable and system-management interrupts can wake a halted core regardless
of what we handle. The loop makes the halt robust to all of them.)
## The `asm volatile` part
@@ -79,32 +79,40 @@ treats the call:
- Code *after* a `noreturn` call is unreachable, so the compiler needn't emit a
return sequence, and won't warn about "missing return value" in the callers.
- It lets `kmain` and `_start` themselves be `noreturn`, which is the honest
signature for a kernel entry point — the bootloader jumps in and nothing ever
jumps back out.
- It lets `kmain` and the exported entry shim `kmainEntry` themselves be
`noreturn`, which is the honest signature for a kernel entry point — the
bootloader jumps in and nothing ever jumps back out.
You can see the chain in `system/kernel/kernel.zig`: `_start` is `noreturn`, it calls
`kmain` which is `noreturn`, which ends by calling `arch.halt()` which is
`noreturn`. The "never returns" property is threaded all the way down.
You can see the chain in the code: `_start` (an assembly stub in
`system/kernel/architecture/x86_64/isr.s`) installs a kernel-owned stack and calls
`kmainEntry` in `system/kernel/kernel.zig`, which is `noreturn`; it calls `kmain`,
also `noreturn`, which ends by calling `architecture.halt()`, again `noreturn`.
The "never returns" property is threaded all the way down.
## Where danos halts
There are three halt sites, and they're all the same idea:
1. **Normal end of kernel work** — `kmain` prints its status, then calls
`arch.halt()`:
1. **The BSP's idle loop** — `kmain` no longer runs out of work: it spawns
`/system/services/init` as PID 1, drops itself to priority 0, and ends as the
bootstrap core's idle task — still by calling `architecture.halt()`:
```zig
con.write("\nkernel initialised; nothing left to do, halting.\n");
arch.halt();
scheduler.setPriority(0);
status("\n/system/kernel: kernel idle; user space is running.\n");
architecture.halt();
```
There's genuinely nothing more to do yet, so the kernel parks itself.
The timer keeps preempting the idle context into init and whatever else is
ready; between those interrupts, the halt loop is exactly the low-power park
described above.
2. **Kernel panic** — the freestanding panic handler has no OS to report to, so
it prints the message in red (if the console is up) and halts via the same
`arch.halt()`. A panic is unrecoverable here, so stopping the machine — rather
than limping on with corrupted state — is the safe response.
it prints the message to the diagnostic log and the on-screen console —
forcing the console back on even if a display service had it suppressed — and
halts via the same `architecture.halt()`. A panic is unrecoverable here, so
stopping the machine — rather than limping on with corrupted state — is the
safe response.
3. **Bootloader failure** — in `boot/efi.zig`, if `boot()` fails *before* handing
off to the kernel, `main` logs the error and parks the machine with the same
@@ -112,14 +120,14 @@ There are three halt sites, and they're all the same idea:
```zig
boot() catch |err| {
log("\r\ndanos: boot failed: ");
log("\r\nEFI: boot failed: ");
logBytes(@errorName(err));
log("\r\n");
while (true) asm volatile ("hlt");
};
```
(Here it's an inline loop rather than `arch.halt()` because that lives in the
(Here it's an inline loop rather than `architecture.halt()` because that lives in the
kernel's arch module, and the loader is a separate binary from the kernel.)
## Summary
@@ -132,5 +140,5 @@ There are three halt sites, and they're all the same idea:
otherwise wake the core and let execution continue.
- **`asm volatile`** emits the instruction and forbids the compiler from removing
it; **`noreturn`** encodes "control never comes back" into the type system.
- danos halts on normal completion, on a kernel panic, and on a bootloader error
— the same "stop safely and stay stopped" in all three.
- danos halts in the kernel's idle loop, on a kernel panic, and on a bootloader
error — the same "park the core safely" in all three.