docs: flip stale status markers across the tracks (audit found 23)

An all-tracks docs-vs-code audit (the same method that caught the storage
drift) found 23 confirmed inaccuracies where a doc's build-status claim no
longer matches the source — status markers that were never flipped after a
track landed, and a few paths left over from completed flag-days. All verified
against the code before editing; docs only, no behavior change.

The systemic ones:
- IOMMU enforcement (driver-model.md, drivers.md): docs said enforcement was
  not built and "device_claim = ring 0" / "memory-safe is not true yet". It is
  built (per-device VT-d/AMD-Vi domains programmed at device_claim, -ECONFINE
  rollback, dma_alloc buffers bound and torn down at death; fail-open only with
  no IOMMU). Restated; M16 marker flipped to done.
- The FHS flag-day paths: /etc/devices.csv -> /system/configuration/devices.csv
  (devices-csv.md, new-driver-checklist.md, device-manager.md), /var/log ->
  /system/logs (logging.md, new-driver-checklist.md), /mnt/usb -> /volumes/usb
  (process-management.md). Following the old paths silently breaks driver match.
- protocol-namespace P4 "remaining" -> landed (only P5 remains); shared-fate
  fan-out "not yet enforced" -> enforced; wall_clock "not built" -> built;
  SMP affinity + fault-recovery "left" -> built; process_enumerate raw-pointer
  trust model -> checked copyToUser/EFAULT; bounds.md maximum_devices static
  hole -> dynamic per-registrar quota; init spawns fat -> volume-manager;
  config "hardcoded, move to /etc" -> already CSV data files; vdso.md three-
  value call; zig-self-hosting library/ layout; python argv "new" -> built.

Found and fixed by a multi-agent audit across 12 doc clusters, each finding
adversarially verified against the source.
This commit is contained in:
Daniel Samson
2026-08-09 22:04:17 +01:00
parent 8216be991d
commit bf0595763e
17 changed files with 134 additions and 105 deletions
+10 -5
View File
@@ -89,14 +89,19 @@ comptime {
}
```
`maximum_domains = 64` and `maximum_devices = 64` agree today only by a sentence in a
comment, and the agreement fails open. This is the clause with a live hole behind it,
and the reason raising `maximum_devices` alone would be a privilege escalation rather
than a fix.
`maximum_domains = 64` and `maximum_devices = 64` once agreed only by a sentence in a
comment, and that agreement failed open — the clause with the live hole behind it, where
raising `maximum_devices` alone left every device id past the end of `iommu.confined`
unconfined while `confineDevice` still reported success, a privilege escalation rather
than a fix. The assert closed that: it held the two together while both stayed fixed, and
when the device table was later made dynamic — no `maximum_devices` any more, only a
per-registrar quota — that forced them apart, the assert having done its job. `confined`
now grows to cover every id the broker mints, and `confineDevice` refuses when it cannot
record a confinement rather than failing open.
## The worked bad case
`devices_broker.maximum_devices`, which had no comment at all:
`devices_broker.maximum_devices`, which had no comment at all, before it was made dynamic:
```zig
/// bound: device nodes for the whole machine — firmware-discovered plus registered
+10 -5
View File
@@ -118,17 +118,22 @@ regions `init` frees. Keeping that boot-protocol knowledge on the loader side is
deliberate — the kernel has no notion of "reclaimable" or of UEFI at all.
The one live piece in that memory is the boot stack the kernel starts on; the loader
leaves the single region containing it `reserved`, so `init` won't hand it out. A
later step will move task 0 onto a kernel-owned stack, freeing that last ~1 MiB
region too (and giving user mode the clean stack it wants).
leaves the single region containing it `reserved`, so `init` won't hand it out. The
kernel is only on it for an instant, though — `_start`'s first instruction switches
to a kernel-owned 64 KiB stack in `.bss` (that context becomes task 0). The region
stays `reserved` because the loader's own `convertMemoryMap` was executing on that
stack when it reclassified the RAM, and, like the map buffers below, nothing frees it
yet.
## What's next (partly done since)
- **Contiguous allocation** — done: `allocContiguous` scans for a run of clear
bits, with an optional physical ceiling for DMA (`dma_alloc` is its user), and
`allocBelow` serves the SMP trampoline.
- **A kernel stack for task 0** — still open: the boot processor's idle task runs
on the boot stack to this day, so that region can't be freed.
- **A kernel stack for task 0** — done: `_start`'s first instruction switches `rsp`
to a kernel-owned 64 KiB stack in `.bss` (`bootstrap_stack`), and `scheduler.init`
registers that running context as task 0 — the kernel is on the loader's boot
stack for that one instruction and never again.
- **Freeing the `reserved` `loader_data`** (the boot-time map buffers) — still
open: the bitmap deliberately tracks those frames so they *can* be freed, but
nothing frees them yet.
+6 -6
View File
@@ -14,7 +14,7 @@ are mirrored to it explicitly (`system/kernel/kernel.zig`).
## The pipeline
```
process std.log ──▶ debug_write(level) ──▶ tagged kernel ring ──▶ logger service ──▶ /var/log/<boot-stamp>/<binary-path>.log
process std.log ──▶ debug_write(level) ──▶ tagged kernel ring ──▶ logger service ──▶ /system/logs/<boot-stamp>/<binary-path>.log
kernel log.print ─┘ │
└▶ serial / 0xE9 sinks (QEMU, -Dserial)
```
@@ -49,12 +49,12 @@ kernel log.print ─┘ │
5. **Persist.** The **logger service** (`system/services/logger`) drains the
ring every 250 ms and demultiplexes records into one file per source under
`/var/log/<boot-stamp>/`, e.g.
`/system/logs/<boot-stamp>/`, e.g.
```
/var/log/2026-07-21T150434Z/kernel.log
/var/log/2026-07-21T150434Z/system/services/fat.log
/var/log/2026-07-21T150434Z/system/drivers/usb-storage.log
/system/logs/2026-07-21T150434Z/kernel.log
/system/logs/2026-07-21T150434Z/system/services/fat.log
/system/logs/2026-07-21T150434Z/system/drivers/usb-storage.log
```
The boot stamp is the RTC anchor from `klog_status` (FAT-safe: no colons; a
@@ -95,6 +95,6 @@ written last so a reader only trusts a complete record).
- A write-spamming process can evict other processes' unread records from the
ring (a per-process quota is future work); the loss is at least visible via
sequence gaps in every affected file.
- `/var/log` files have no privacy until the VFS grows permissions.
- `/system/logs` files have no privacy until the VFS grows permissions.
- Records emitted after the logger's final shutdown drain reach serial and the
ring but not the files.
+6 -4
View File
@@ -14,7 +14,7 @@ process-manager server, and Fuchsia/seL4 control processes only through handles.
danos rules out `/proc` **as the primitive**: the path router lives in the
kernel (`fs_resolve`), but what is mounted under a path is served by a
user-process filesystem server (the way FAT serves `/mnt/usb`) — a `/proc`
user-process filesystem server (the way FAT serves `/volumes/usb`) — a `/proc`
would be one more such server, which would put a user process in the path of
process control. If that server (or anything under it) hangs, nothing could be
listed or killed, *including the hung server*. The control plane for processes
@@ -117,9 +117,11 @@ the architecture layer calls up into `tick`.
`process_exit_reason` (`process.exitReason`). This is the input to
restart policy ([process-lifecycle.md](process-lifecycle.md)); an exit *code*
for the clean case can still ride alongside later.
- Enumerate writes through the caller's raw pointer under the bring-up trust
model, like `device_enumerate` (an unmapped page is a self-DoS, not an
isolation break).
- ~~Enumerate writes through the caller's raw pointer under the bring-up trust
model, like `device_enumerate`~~ Closed (8d4a7cf): both `process_enumerate`
and `device_enumerate` describe a chunk into a kernel buffer and place it with
`copyToUser`, which validates the range and resolves each page — an unmapped
page returns `-EFAULT`, and the kernel never stores through the user pointer.
## Tests
+3 -3
View File
@@ -1,9 +1,9 @@
# The protocol namespace
*Design, agreed 2026-07-31. Supersedes the `ServiceId` registry. P1–P3 of the
*Design, agreed 2026-07-31. Supersedes the `ServiceId` registry. P1–P4 of the
migration plan at the end have landed (the envelope, the registry and the
`ServiceId` flag-day, and restriction stage one); P4 and P5 are the remaining
work list.*
`ServiceId` flag-day, restriction stage one, and the protocol rebase); P5
(restriction stage two) is the remaining work.*
How a program finds, connects to, and is restricted from the things it talks to.
Three ideas, kept deliberately separate:
+8 -6
View File
@@ -23,8 +23,8 @@ INIT–SIPI–SIPI, brings each up into 64-bit long mode with its own descriptor
LAPIC and timer, and drops it into the scheduler. Tasks run **genuinely in parallel** —
the `smp` self-test confirms worker tasks executing on all four cores at once under
QEMU `-smp 4`. Shared kernel state (scheduler queues, IPC) is serialised behind a big
kernel lock. What's left is refinement, not first-light: per-core run queues, IPIs,
and thread-to-core affinity (see [Implementation status](#implementation-status)).
kernel lock. What's left is refinement, not first-light: per-core run queues and IPIs
(see [Implementation status](#implementation-status)).
## The common microkernel instinct: don't share kernel state
@@ -238,10 +238,12 @@ next lands.
- **Per-core run queues** — the Fiasco.OC direction, if the single global queue's lock
contention ever bites. (Thread *affinity* already exists — see above; this is the
further step of giving each core its own primary run queue for load distribution.)
- **Fault recovery** — today a fault halts (only) the faulting core. Turning that into
"kill the task, keep the core running" is the [resilience](resilience.md) track (it
needs the task's lock/resource state handled), and for taking a core fully offline,
its tasks migrated first.
- **Fault recovery** — a ring-3 fault already **kills the faulting process and keeps the
core (and the rest of the system) running**: the kernel trapped it on the task's own
kernel stack, reclaims what the process held, and reschedules
(`process.killCurrentProcess`; the `fault-recovery` test proves init keeps heartbeating
through the kill) — the [resilience](resilience.md) track. What's still open here is
taking a core fully **offline**, which additionally needs its tasks migrated off first.
## Further reading
+1 -1
View File
@@ -84,7 +84,7 @@ address space. Threads deliberately remove that boundary *within* a process:
there is no isolation **between** threads.
- Threads share fate — by contract: a fault in any thread, or a "kill the process"
decision, takes down **all** of them, so restartability lives at the process level,
not the thread level. (The kernel does not yet enforce this fan-out — see the
not the thread level. (The kernel enforces this fan-out — see the
Lifecycle note under
[Interaction with the rest of the kernel](#interaction-with-the-rest-of-the-kernel).)
- Shared mutable state reintroduces data races — the failure class the
+13 -10
View File
@@ -36,9 +36,10 @@ and inside a VM alike; only the source behind it differs. The mechanism is in
So the timer hardware lives in the kernel, and there is **no `hpet` driver and no time
server** to consume. (An earlier HPET driver existed only to *demonstrate* the driver
model; that role now lives in [drivers.md](../device-driver-development/drivers.md), as documentation.) The one place
a user-space time service *is* justified — **wall-clock / calendar time** — is discussed
at the end; it is deliberately not built yet.
model; that role now lives in [drivers.md](../device-driver-development/drivers.md), as documentation.) The one part
left to user space — **calendar policy** over wall-clock time (time zones, formatting) —
is discussed at the end; the wall-clock *seconds* it builds on are a kernel syscall
(`wall_clock`), like the monotonic clock.
## The three system calls
@@ -95,15 +96,17 @@ The raw wrappers (`clock`, `sleepMillis`, `timerOnce`) and the ergonomic
`Instant`/`Duration` layer both live in the `time` module
(`library/kernel/time.zig`); the latter is what everyday code uses.
## Wall-clock time (not built)
## Wall-clock time
Everything above is **monotonic**: elapsed time since boot, perfect for timeouts and
measurement, useless for "what is the date?" Calendar time — a real-time clock, time
zones, leap seconds — is genuinely a **user-space** concern, and it *is* the case a time
service is for. It would be backed by an **RTC** driver (the CMOS real-time clock), not
the HPET, and exposed as a `CLOCK_REALTIME`-style service alongside the monotonic
syscall. It is deferred until something needs it; the monotonic clock the kernel already
owns covers every current use.
measurement, useless for "what is the date?" Calendar time needs a **real-time clock**.
The kernel owns wall-clock *seconds* as mechanism, exactly like the monotonic clock: the
`wall_clock` syscall (#33) returns Unix epoch seconds (UTC). The CMOS **RTC** is read
once at boot and anchored to the monotonic clock (`system/kernel/wall-clock.zig`), so a
query is a cheap arithmetic offset rather than a per-call CMOS poll; `time`'s
`wallClock()` (`library/kernel/time.zig`) wraps it. Reading the hardware's value is not
policy — time zones, leap seconds, calendars, and formatting layer on top in user space.
It exists because the filesystem needs real timestamps (mtime).
## Verifying it
+9 -6
View File
@@ -123,12 +123,15 @@ The calls that return two values in `rax:rdx` today — `dma_alloc`
`fs_resolve` (route tag + node token / backend handle) —
become functions returning a two-`u64` struct. The System V ABI returns a
16-byte struct in `rax:rdx`, so the stub is a plain `syscall; ret` — the
C-ABI spelling of the existing convention, at zero cost. The one call that
returns *three* values — `ipc_reply_wait` (receive_len in `rax`, badge in
`rdx`, received capability in `r8`) — exceeds the two-register return: its
function returns a three-`u64` struct, which the ABI passes via a hidden
result pointer, so that one stub stores `rax`/`rdx`/`r8` through the pointer
after the `syscall` — a few instructions rather than one.
C-ABI spelling of the existing convention, at zero cost. The calls that
return *three* values — `ipc_reply_wait` (receive_len in `rax`, badge in
`rdx`, received capability in `r8`) and `dma_alloc` when the region is
`shareable` (virtual in `rax`, physical in `rdx`, handle in `r8`) — exceed the
two-register return: their functions return a three-`u64` struct, which the
ABI passes via a hidden result pointer, so each stub stores `rax`/`rdx`/`r8`
through the pointer after the `syscall` — a few instructions rather than one.
(`ipc_call` likewise carries a received capability in `r8` alongside its `rax`
result.)
Grouped as `abi.zig` groups them: