docs: full docs-vs-code audit — fix every stale claim across 40 docs

Every doc verified claim-by-claim against the code by parallel audit agents,
then fixed and adversarially re-verified. Two waves of staleness corrected:
the originally audited findings (higher-half boot handoff, kernel VFS
takeover, fault isolation + claim release + driver restart, AML/S5 moving to
ring 3, threading's shipped design, USB+FAT landing) and a second pass of
adjacent claims the verifiers caught (smp.md 'not built yet' intro,
system-requirements' PS/2-only and no-storage claims, halting.md's red-panic
and no-IDT text, testing.md's serial mirroring, router-era vfs-protocol
wording, capsule-first boot loading).

threading.md now documents the shared-fate gap explicitly: the design says a
process dies whole, the kernel today kills only the offending thread.

Also fixes three stale code comments (isr.s exceptionHandler, acpi.zig
sleepValue, build.zig boot-volume) — comments only, no behavior change.
This commit is contained in:
Daniel Samson
2026-07-22 09:09:53 +01:00
parent 52df2ba6f6
commit e854f65623
43 changed files with 821 additions and 546 deletions
+22 -14
View File
@@ -29,7 +29,7 @@ The compositor itself (layers, back buffer, damage) does not change. Only the la
│
└─ VirtioGpuBackend talks to a virtio-gpu driver process over a `scanout`
service: present via a shared resource + fenced flush,
EDID mode list, runtime mode-set.
a fixed mode list, runtime mode-set, EDID refresh rate.
```
A **backend** is a small interface the compositor calls:
@@ -53,10 +53,12 @@ manager brings it up after boot), and because danos is meant to be resilient:
blank screen while drivers load — the exact v1 behaviour.
2. **Upgrade on announce.** When the virtio-gpu driver has claimed its device and set up a
scanout, it **announces itself to the display service** (a `push`: the driver looks up
`.display` and sends an *attach-scanout* message carrying its `scanout` endpoint as a
capability). The compositor switches to `VirtioGpuBackend` and re-presents the current
frame full-screen. Push beats polling — the compositor doesn't know a priori which
driver, if any, exists, and danos has no service-registration pub/sub.
`.display` and sends an *attach-scanout* message carrying the shared scanout **surface**
as a capability; the compositor maps it and reaches the driver's present/mode channel by
looking up the registered `scanout` service). The compositor switches to
`VirtioGpuBackend` and re-presents the current frame full-screen. Push beats polling —
the compositor doesn't know a priori which driver, if any, exists, and danos has no
service-registration pub/sub.
3. **Native is restartable, not fallback-on-crash.** Once a native driver has reprogrammed
the device, the firmware's GOP framebuffer is **stale** — "native → GOP" is not a clean
fall-back. So a native driver that **crashes** is *restarted* by its supervisor (the
@@ -64,9 +66,10 @@ manager brings it up after boot), and because danos is meant to be resilient:
(native → native). The screen freezes on the last frame during the gap — acceptable.
4. **GOP is the floor for "no driver was ever there."** On a real GPU (NVIDIA/AMD/Intel)
the class-0x03 device matches nothing in the driver table, no `scanout` is ever
announced, and the compositor stays on GOP forever — no special-casing. Only if a
native driver *permanently* gives up (crash-loop cap) does the compositor attempt GOP
again, and even then only if the LFB is still mappable.
announced, and the compositor stays on GOP forever — no special-casing. If a native
driver *permanently* gives up (the device manager's crash-loop cap), nothing tells the
compositor and there is no path back to GOP — the screen stays frozen on the last
frame. A GOP revert (sensible only while the LFB is still mappable) is not built.
## The shared-memory primitive this needs
@@ -91,11 +94,15 @@ reference instead of drawing by command). One piece of kernel work, two features
A new ring-3 driver process (the topology v1 anticipated — "split the driver from the
compositor when a second backend arrives"). It claims the virtio-gpu PCI function, and:
- sets up the **virtqueues** (control + cursor) and the device's config space,
- sets up the **control virtqueue** (queue 0 — the only queue it uses; no cursor queue)
and the device's config space,
- creates a **2D scanout resource** backed by a shared-memory region, `attach_backing`s it,
`set_scanout`s it to a CRTC, and `resource_flush`es damaged rectangles,
- reads **EDID** (the `GET_EDID` control command) for the mode list, and `set_scanout`
at a chosen mode for **runtime mode-setting**,
`set_scanout`s it to a CRTC, and on each present transfers and `resource_flush`es the
full current-mode rectangle (damage-narrowed flushes are a later refinement),
- reads **EDID** (the `GET_EDID` control command) to log the monitor's preferred timing
and derive the refresh rate it announces (the compositor's frame-clock seed); the mode
list it offers is a fixed pair — 640×480 and 800×600 — and `set_scanout` at a chosen
mode gives **runtime mode-setting**,
- registers a `scanout` service and announces to the display service.
Its `resource_flush` is the real **present** — and gives a **fenced, tear-free** path a
@@ -111,8 +118,9 @@ the existing IRQ-as-IPC path) or the compositor's own frame clock.
## What v2 unlocks — and its honest scope
Behind the abstraction, a native backend gives runtime **mode-setting** (resolution /
refresh / bpp), **EDID** enumeration, and **fenced presents**. But only on devices we have a driver
Behind the abstraction, a native backend gives runtime **mode-setting** (resolution only —
the scanout protocol carries neither refresh nor bpp), an **EDID**-derived refresh rate, and
**fenced presents**. But only on devices we have a driver
for — realistically **VMs** (virtio-gpu, and later maybe Bochs DISPI). Real discrete GPUs
need per-vendor KMS-class drivers that aren't getting written, so they **stay on GOP** —
which is genuinely fine (v1 on the NVIDIA box is smooth). So v2's real value is twofold: