docs: full docs-vs-code audit — fix every stale claim across 40 docs

Every doc verified claim-by-claim against the code by parallel audit agents,
then fixed and adversarially re-verified. Two waves of staleness corrected:
the originally audited findings (higher-half boot handoff, kernel VFS
takeover, fault isolation + claim release + driver restart, AML/S5 moving to
ring 3, threading's shipped design, USB+FAT landing) and a second pass of
adjacent claims the verifiers caught (smp.md 'not built yet' intro,
system-requirements' PS/2-only and no-storage claims, halting.md's red-panic
and no-IDT text, testing.md's serial mirroring, router-era vfs-protocol
wording, capsule-first boot loading).

threading.md now documents the shared-fate gap explicitly: the design says a
process dies whole, the kernel today kills only the offending thread.

Also fixes three stale code comments (isr.s exceptionHandler, acpi.zig
sleepValue, build.zig boot-volume) — comments only, no behavior change.
This commit is contained in:
Daniel Samson
2026-07-22 09:09:53 +01:00
parent 52df2ba6f6
commit e854f65623
43 changed files with 821 additions and 546 deletions
+41 -29
View File
@@ -31,8 +31,8 @@ kernel ──spawns──► init (PID 1) ──spawns──► device-manag
| | |
spawns only init, the service supervisor: the driver supervisor: enumerates
publishes the starts the system /system/devices, matches each device
initial-ramdisk services (vfs, the to a driver, and system_spawn's it
so user space can device-manager). Its
initial-ramdisk services (device-manager, to a driver, and system_spawn's it
so user space can fat, logger, ...). Its
system_spawn from it list is init policy.
```
@@ -42,15 +42,22 @@ in the initial-ramdisk as a fresh ring-3 process — `name` becoming its argv[0]
the optional arguments its argv[1..], on a SysV entry stack, see sysv.md). Everything else is a user-space decision:
- **init** ([system/services/init](system/services/init/init.zig)) is the **service
supervisor**. It spawns the system services danos brings up at boot — today `vfs` and
the `device-manager` — from a small list. Drivers are deliberately *not* its job.
supervisor**. It spawns the system services danos brings up at boot — today `input`,
the `device-manager`, `fat`, `display`, `display-demo`, and the `logger` — from a
small list. Drivers are deliberately *not* its job. (An earlier draft listed a `vfs`
service here; that service is retired — the router moved into the kernel as
`fs_resolve`.)
- **device-manager** ([system/services/device-manager](system/services/device-manager/device-manager.zig))
is the **driver supervisor**. It does the three steps a monolithic kernel would do in
its probe path, entirely from ring 3:
1. **Discover** — `device_enumerate` snapshots the device table the kernel built from
ACPI/PCI ([discovery](discovery.md)).
2. **Match** — for each device it looks up a driver by `DeviceClass`. The match policy
is a table (`driverFor`): today a static `timer → hpet` map; a fuller system reads
2. **Match** — for each device it looks up a driver. The match policy is code, a few
small per-bus tables: from the boot snapshot only the PCI host bridge matches
(→ `pci-bus`); everything else arrives later as bus reports and matches on
identity — `pciDriverForIdentity` (xHCI → `usb-xhci-bus`, virtio-gpu →
`virtio-gpu`), `hidDriverFor` (PNP0303/PNP0F13 → `ps2-bus`), and
`usbDriverForIdentity` (USB keyboard, mouse, storage). A fuller system reads
what each driver *binds* (a manifest under `/system/drivers`, or the driver
describing its own match).
3. **Spawn** — `system_spawn(driver_name, arguments)` starts the matched driver (the
@@ -84,7 +91,7 @@ names a device by id and a resource by index. That indirection is the entire sec
model. If `mmio_map` took a physical address, any process could map the kernel's
memory; if `irq_bind` took a GSI, any process could bind the keyboard's line and
silently intercept it. Instead the kernel checks two things (`process.ownedGsi`, and
the same check at the top of `sysMmioMap`):
the same check at the top of `systemMmioMap`):
- `devices_broker.ownerOf(dev_id) == me` — you claimed it, and claims are exclusive
- the resource at `res_idx` is of the right *kind* — `memory` for `mmio_map`, `irq`
@@ -95,11 +102,15 @@ The claim is the capability. Everything else follows from it.
## Registers: `mmio_map`
`mmio_map` walks the caller's page tables and installs the device's physical frames
with `present | user | writable | nx | pcd | pwt`
(`arch/x86_64/paging.zig:mapUserDeviceInto`). Two of those bits are load-bearing:
with `present | user | writable | nx | device_grant` plus a cache mode
(`system/kernel/architecture/x86_64/paging.zig:mapUserDeviceInto`). Two of those bits
are load-bearing:
- **`pcd | pwt`** — strong-uncacheable. A device register is not memory; a cached read
would return a stale value and a write might never leave the CPU.
- **`pcd | pwt`** — strong-uncacheable, the default cache mode. A device register is
not memory; a cached read would return a stale value and a write might never leave
the CPU. The one exception: a resource flagged write-combining
(`resource_flag_write_combining` — today the kernel-seeded display framebuffer)
gets the PAT bit instead, so pixel writes batch into bursts.
- **`device_grant`** (bit 9, one of the PTE's available bits) — marks the leaf as MMIO
rather than RAM, so `freeSubtree` skips `pmm.free` on it when the address space is
destroyed. Without this, killing a driver would hand the HPET's registers back to
@@ -143,7 +154,7 @@ interrupt fires exactly once, ever; call it before the device is quiet and you g
interrupt storm. That single fact explains why `irq_bind` and `irq_ack` are two
syscalls and not one.
This is also why `interruptDispatch` (`arch/x86_64/idt.zig`) no longer issues the EOI
This is also why `interruptDispatch` (`system/kernel/architecture/x86_64/idt.zig`) no longer issues the EOI
itself. It used to, before running the handler — correct for the LAPIC timer, and
impossible for a routed device line. Each handler now owns its EOI, because only the
handler knows which discipline its source needs.
@@ -234,10 +245,10 @@ static capability.
A device that *contains other devices* — a PCI bridge, a USB hub, or the HPET's block
of comparators — needs a driver that enumerates it and tells the kernel what it found.
That's `device_register`, and it makes the device table a tree rather than a list
(`DeviceDesc.parent`).
(`DeviceDescriptor.parent`).
```zig
var child = std.mem.zeroes(dev.DeviceDesc);
var child = std.mem.zeroes(dev.DeviceDescriptor);
child.class = @intFromEnum(dev.DeviceClass.timer);
child.resource_count = 1;
child.resources[0] = .{ .kind = memory, .start = bus_base + 0x100, .len = 0x20 };
@@ -248,10 +259,12 @@ The child is left **unclaimed**, which is the whole point: another process claim
`mmio_map`s it, and sees only that 0x20-byte window.
The rule the kernel enforces is **containment**: every resource of a child must lie
inside a resource of the same kind on its parent. Ranges must nest; an IRQ must match
exactly. This isn't bureaucracy — a `DeviceDesc` is a licence to map physical memory, so
without containment `device_register` would be a syscall for mapping any page you like. A
bus driver may only ever subdivide what it already owns.
inside a resource of the same kind on its parent. Ranges must nest; a child's IRQ —
still exactly one line — must fall within the parent's IRQ range (a length-1 parent
range is the old exact-match rule). This isn't bureaucracy — a `DeviceDescriptor` is a
licence to map physical memory, so without containment `device_register` would be a
syscall for mapping any page you like. A bus driver may only ever subdivide what it
already owns.
A device with **no resources** is legal and common. A USB device is reached through its
controller, not by MMIO, so it gets `resource_count = 0`.
@@ -277,8 +290,13 @@ Several things this list used to warn about are now available (see
[driver-model.md](driver-model.md)): **port I/O** (`io_read`/`io_write`, claim-gated by
the device's `io_port` resource — direct ring-3 `in`/`out` is still a #GP, so a PS/2 or
16550 driver goes through these), **DMA memory** (`dma_alloc`: contiguous, pinned,
uncacheable, physical address exposed), and **memory barriers** (`/lib/mmio`'s
`mb`/`rmb`/`wmb`). What remains:
uncacheable, physical address exposed), **memory barriers** (`library/mmio`'s
`mb`/`rmb`/`wmb`, imported as the `mmio` module), **fault isolation** (a ring-3 fault kills only the faulting
process — `killCurrentProcess` — and the machine keeps running,
[resilience](resilience.md)), and **reclaim + restart on death** (every path out of a
process releases its claims and IRQ/MSI bindings — `releaseAllOwnedBy`,
`irq.releaseOwner` — and the device manager respawns the driver with backoff,
[device-manager.md](device-manager.md)). What remains:
- **Page granularity.** `mmio_map` rounds to 4 KiB. Two devices sharing a page means
granting one grants the other. A `device_register`ed child's *resource* can be narrower
@@ -289,8 +307,9 @@ uncacheable, physical address exposed), and **memory barriers** (`/lib/mmio`'s
are programmed, so `device_claim` on a DMA-capable device is still effectively
equivalent to granting ring 0. This is the largest gap between the design's promise and
what it delivers; enforcement lands with the first DMA driver.
- **No `dev_release`.** A claim is never dropped (only IRQ/MSI bindings are, on exit), so
a device stays owned for the life of its driver — which blocks restart.
- **No voluntary `dev_release`.** A *live* driver can't drop a claim — only exit
releases it (any path out of a process runs `releaseAllOwnedBy`) — so handing a
device between running drivers still means exiting.
- **One endpoint per GSI**, so shared legacy PCI INTx lines can't be split between two
drivers. MSI/MSI-X — one vector per device, edge-triggered, unshared — is the real
answer, and QEMU's HPET doesn't offer it (`Tn_FSB_INT_DEL_CAP = 0`).
@@ -303,13 +322,6 @@ uncacheable, physical address exposed), and **memory barriers** (`/lib/mmio`'s
the ISR until `irq_ack`, so at most one badge is ever outstanding. Bind nine devices
to one endpoint, though, and a dropped badge leaves that line masked with nobody
left to ack it.
- **A faulting driver still kills the machine.** There is no per-process kill path: a
ring-3 page fault halts the kernel, so `releaseIrqs` runs only on a voluntary
`exit`. Fault isolation is the whole premise ([vision](vision.md)) and it is
[not built yet](resilience.md).
- **A dead driver's device is not reclaimed.** `releaseIrqs` unbinds and masks the
line on exit, but the claim is never released — restart is
[not built](resilience.md).
- **On real hardware, the mask/EOI cycle may need a remote-IRR flush.** Masking a
level-triggered redirection entry with remote-IRR set doesn't clear it on some
chipsets, and the line never fires again. QEMU clears it on EOI regardless, so the