docs: full docs-vs-code audit — fix every stale claim across 40 docs

Every doc verified claim-by-claim against the code by parallel audit agents,
then fixed and adversarially re-verified. Two waves of staleness corrected:
the originally audited findings (higher-half boot handoff, kernel VFS
takeover, fault isolation + claim release + driver restart, AML/S5 moving to
ring 3, threading's shipped design, USB+FAT landing) and a second pass of
adjacent claims the verifiers caught (smp.md 'not built yet' intro,
system-requirements' PS/2-only and no-storage claims, halting.md's red-panic
and no-IDT text, testing.md's serial mirroring, router-era vfs-protocol
wording, capsule-first boot loading).

threading.md now documents the shared-fate gap explicitly: the design says a
process dies whole, the kernel today kills only the offending thread.

Also fixes three stale code comments (isr.s exceptionHandler, acpi.zig
sleepValue, build.zig boot-volume) — comments only, no behavior change.
This commit is contained in:
Daniel Samson
2026-07-22 09:09:53 +01:00
parent 52df2ba6f6
commit e854f65623
43 changed files with 821 additions and 546 deletions
+17 -11
View File
@@ -1,10 +1,11 @@
# SMP: multiple cores, the microkernel way
A design/research note, not built yet. danos runs on **one core** today (see
[scheduling.md](scheduling.md)); this maps how microkernels — especially the L4
family and seL4 — handle **symmetric multiprocessing (SMP)**, so the eventual port
has a plan and a reading list. It also flags where those choices depend on whether
danos is chasing **real-time** or **resilience** (see the note at the end).
A design/research note that predates the build — danos now runs on **multiple
cores** by default (see [Implementation status](#implementation-status) and
[scheduling.md](scheduling.md)). This maps how microkernels — especially the L4
family and seL4 — handle **symmetric multiprocessing (SMP)**, the plan and
reading list the port followed. It also flags where those choices depend on
whether danos is chasing **real-time** or **resilience** (see the note at the end).
## First, the vocabulary
@@ -158,11 +159,15 @@ next lands.
`ipc.zig` run every critical section under it. Uncontended on one core, so behaviour
is identical to the old interrupt-flag model.
- **Per-CPU state** — a `PerCpu` struct (running task, idle task, APIC id) per core,
its pointer kept in the x86 **GS base** (`IA32_GS_BASE`; no `swapgs`, since there's
no user mode yet). The old global `current` is now `thisCpu().current`. The ready
its pointer kept in the x86 **GS base** (`IA32_GS_BASE`). No `swapgs` was needed at
the time — there was no user mode yet; with ring 3 in place, every ring transition
now swaps it against the user's own GS base under the `swapgs` discipline (see
`system/kernel/architecture/x86_64/per-cpu.zig`), and each AP enables the fast
system-call path (`initSystemCall`) for itself at bring-up. The old global
`current` is now `thisCpu().current`. The ready
queues stay **global** under the lock — work-conserving, so any idle core will pull
the highest-priority ready task; per-core queues are a later optimisation.
- **AP wake to long mode** — `arch.startSecondary` drives INIT–SIPI–SIPI (via the
- **AP wake to long mode** — `architecture.startSecondary` drives INIT–SIPI–SIPI (via the
LAPIC ICR) to wake each parked core one at a time. A woken core starts in 16-bit
real mode at a low page and runs the [trampoline](../system/kernel/architecture/x86_64/trampoline.s)
up through protected mode into 64-bit long mode, then lands in `smp.zig:apEntry`,
@@ -203,8 +208,9 @@ next lands.
[scheduling.md](scheduling.md#affinity-pinning-a-task-to-a-core)). The `affinity`
test confirms a pinned task never migrates. This is the mechanism the fault-on-AP
test rides on, and the *explicit-affinity* real-time-predictable model.
- **Right-sized footprint** — the per-CPU ceiling (`system.max_cpus`, one constant
shared by discovery, the scheduler, and the per-core GDT/TSS) is generous (128), but
- **Right-sized footprint** — the per-CPU ceiling (`parameters.maximum_cpus`, one
constant shared by discovery, the scheduler, and the per-core GDT/TSS) is generous
(128), but
the *large* per-core resources — the kernel and IST (double-fault) stacks — are
**heap-allocated at bring-up**, only for cores that actually come online. Only the
BSP's IST stack is static, because it must exist before the frame allocator does.
@@ -217,7 +223,7 @@ next lands.
life, but kept **inert between wakes**: zeroed and non-executable, armed (blob
copied in, page made executable) only for the moment a core is actually climbing,
then disarmed again. So there's never a dormant executable page, and a core can be
(re)woken at any time — `arch.startSecondary` is one self-contained attempt (arm →
(re)woken at any time — `architecture.startSecondary` is one self-contained attempt (arm →
INIT–SIPI–SIPI → disarm), and its `INIT` resets a wedged core, so retrying just
works. Boot retries a non-responding core up to three times; the same primitive is
the groundwork a future **power manager** would drive to bring cores up (and,