docs: full docs-vs-code audit — fix every stale claim across 40 docs

Every doc verified claim-by-claim against the code by parallel audit agents,
then fixed and adversarially re-verified. Two waves of staleness corrected:
the originally audited findings (higher-half boot handoff, kernel VFS
takeover, fault isolation + claim release + driver restart, AML/S5 moving to
ring 3, threading's shipped design, USB+FAT landing) and a second pass of
adjacent claims the verifiers caught (smp.md 'not built yet' intro,
system-requirements' PS/2-only and no-storage claims, halting.md's red-panic
and no-IDT text, testing.md's serial mirroring, router-era vfs-protocol
wording, capsule-first boot loading).

threading.md now documents the shared-fate gap explicitly: the design says a
process dies whole, the kernel today kills only the offending thread.

Also fixes three stale code comments (isr.s exceptionHandler, acpi.zig
sleepValue, build.zig boot-volume) — comments only, no behavior change.
This commit is contained in:
Daniel Samson
2026-07-22 09:09:53 +01:00
parent 52df2ba6f6
commit e854f65623
43 changed files with 821 additions and 546 deletions
+91 -72
View File
@@ -81,9 +81,11 @@ address space. Threads deliberately remove that boundary *within* a process:
- Threads share one address space, so one thread's stray write corrupts them all —
there is no isolation **between** threads.
- Threads share fate: a fault in any thread, or a "kill the process" decision, takes
down **all** of them. Restartability lives at the process level, not the thread
level.
- Threads share fate — by contract: a fault in any thread, or a "kill the process"
decision, takes down **all** of them, so restartability lives at the process level,
not the thread level. (The kernel does not yet enforce this fan-out — see the
Lifecycle note under
[Interaction with the rest of the kernel](#interaction-with-the-rest-of-the-kernel).)
- Shared mutable state reintroduces data races — the failure class the
isolate-and-message model was chosen to avoid.
@@ -103,16 +105,15 @@ Lives in `library/runtime/thread.zig`, re-exported as `runtime.Thread`.
pub const Thread = struct {
pub const Id = u32; // the kernel task id
pub const SpawnConfig = struct {
stack_size: usize = default_stack_size,
allocator: ?std.mem.Allocator = null, // for the closure + stack bookkeeping
stack_size: usize = default_stack_size, // no allocator: the closure lives at the top of the thread's own stack
};
pub const SpawnError = error{ OutOfMemory, ThreadQuotaExceeded, SystemResources };
pub const SpawnError = error{SystemResources};
pub fn spawn(config: SpawnConfig, comptime function: anytype, args: anytype) SpawnError!Thread;
pub fn join(self: Thread) void; // block until the thread ends, reclaim its stack
pub fn detach(self: Thread) void; // give up the right to join; kernel reclaims on exit
pub fn detach(self: Thread) void; // give up the right to join; stack reclaimed at process exit
pub fn getCurrentId() Id;
pub fn yield() void; // -> existing `yield` syscall
pub fn currentCore() Id; // danos extension: the calling core's dense index
pub const Mutex = struct { pub fn lock(*Mutex) void; pub fn tryLock(*Mutex) bool; pub fn unlock(*Mutex) void; };
pub const Condition = struct { pub fn wait(*Condition, *Mutex) void; pub fn timedWait(*Condition, *Mutex, u64) error{Timeout}!void; pub fn signal(*Condition) void; pub fn broadcast(*Condition) void; };
@@ -127,18 +128,21 @@ Deviations from `std.Thread`, called out honestly:
- **The thread function's return value is discarded** (as `std.Thread.join` returns
`void`). Return data through shared state or a `Semaphore`/`Condition`, not the
return.
- `getCpuCount()` maps to the existing SMP core count ([smp.md](smp.md)); a service
rarely needs it.
- No `getCpuCount()` (a service rarely needs it) and no `Thread.yield()` — `yield`
lives in `runtime.system`. Instead `currentCore()` exposes the calling core's dense
index ([smp.md](smp.md)), used to observe genuine cross-core parallelism.
## Kernel primitives (new private syscalls)
Four new entries extend [abi.zig](../system/abi.zig) `SystemCall` after
`shared_memory_physical = 36`, each with a `library/runtime` wrapper:
Five core entries extend [abi.zig](../system/abi.zig) `SystemCall` after
`shared_memory_physical = 36` (plus small helpers `current_core`, `thread_self`, and
`set_thread_pointer`), each with a `library/runtime` wrapper:
| Syscall | Signature | Purpose |
|---|---|---|
| `thread_spawn` | `(entry, stack_top, arg) -> tid` | create a task sharing the **caller's** address space |
| `thread_exit` | `(stack_base, stack_len)` | end the calling thread; hand back its stack range for reclaim |
| `thread_spawn` | `(entry, stack_top, arg, exit_endpoint) -> tid` | create a task sharing the **caller's** address space; the runtime passes `no_cap` for `exit_endpoint` (join is a syscall, not an endpoint) |
| `thread_exit` | `()` | end the calling thread; its stack is reclaimed by the joiner's `munmap`, not the kernel |
| `thread_join` | `(tid) -> 0` | block until the thread with id `tid` has exited |
| `futex_wait` | `(addr, expected, timeout_ns) -> status` | block if `*addr == expected`, until woken or timeout |
| `futex_wake` | `(addr, count) -> woken` | wake up to `count` waiters on `addr` |
@@ -148,56 +152,56 @@ Plus one invariant change with no new syscall: **address-space reference countin
### Address-space reference counting
Today an address space is 1:1 with a task: `spawnUserLocked` records `address_space` on the
Task, and teardown does `destroyAddressSpace(t.address_space)` when **any** user task exits
Before this work an address space was 1:1 with a task: `spawnUserLocked` records
`address_space` on the Task (as it still does), and teardown did
`destroyAddressSpace(t.address_space)` when **any** user task exited
([scheduler.zig](../system/kernel/scheduler.zig)). With threads, several tasks share
one `address_space`, so the first to exit would rip the address space out from under its
siblings.
Fix: a small refcount keyed by the address-space root (`createAddressSpace` in
[process.zig](../system/kernel/process.zig) sets it to 1). `thread_spawn` increments
it; task teardown decrements and only calls `destroyAddressSpace` at **zero**. All of
Fix: a small refcount keyed by the address-space root, kept in
[scheduler.zig](../system/kernel/scheduler.zig): `retainAddressSpace` takes a
reference for every user task `spawnUserLocked` starts (count 1 on the first take, so
a thread sharing the caller's space increments it); task teardown calls
`releaseAddressSpace`, which only calls `destroyAddressSpace` at **zero**. All of
this is already under the big kernel lock, so no new locking. This is the one piece
that must land and be proven before anything shares an address space.
### `thread_spawn` and the trampoline
The scheduler already accepts an arbitrary `address_space` and does **not** smuggle values
through registers — `startUserTask` reads the entry/stack from the Task and
`jumpToUser`s ([scheduler.zig](../system/kernel/scheduler.zig)). That makes the thread
path clean:
The scheduler already accepts an arbitrary `address_space` and does **not** smuggle
values through scratch registers — `startUserTask` reads the entry/stack (and the
thread's closure arg, delivered in `rdi` via `jumpToUserArg`) from the Task
([scheduler.zig](../system/kernel/scheduler.zig)). That makes the thread path clean:
1. The runtime's `spawn` `mmap`s a stack (syscall `4`), heap-allocates a closure —
`{ fn_ptr, args_tuple, completion }`, the std "Instance" pattern — and writes the
closure pointer to the **top word of the new stack**.
2. It calls `thread_spawn(entry = &threadTrampoline, stack_top, arg = closure_ptr)`.
The kernel calls the same `spawnUserLocked` path with the **caller's address space**
(refcount++), `entry`, and `user_sp = stack_top`.
3. `threadTrampoline` (a small runtime shim) reads the closure off its stack, calls
the user function, then calls `thread_exit`. No new register ABI — the closure
pointer rides the stack the runtime set up, mirroring how `startUserTask` avoids
register smuggling.
1. The runtime's `spawn` `mmap`s a stack (syscall `4`) and writes the closure —
`{ tls_base, args }`, the std "Instance" pattern — at the **top of the new stack
itself** (no heap allocation), with a small per-thread TLS block just below it.
2. It calls `thread_spawn(entry = &Closure.entry, stack_top, arg = closure_ptr,
exit_endpoint = no_cap)`. The kernel calls the same `spawnUserLocked` path with the
**caller's address space** (refcount++), `entry`, and `user_sp = stack_top`.
3. `Closure.entry` (a small runtime shim) receives the closure pointer in `rdi` — the
kernel delivers `arg` as the entry's first C-ABI argument — sets the thread
pointer, calls the user function, then calls `thread_exit`.
Unlike a process start, there is **no** System V argc/argv/auxv block
([sysv.md](sysv.md)) — a thread stack carries only the closure pointer.
([sysv.md](sysv.md)) — a thread stack carries only the closure and its TLS block.
### Lifetime: exit, join, detach, stack reclaim
- **`thread_exit`** marks the task dead and hands the kernel the thread's user-stack
range. The kernel reaps the task on the scheduler (already running on a *kernel*
stack, so it can safely unmap the user stack), decrements the address-space refcount, and
frees the task slot.
- **`join` — Stage 1** reuses the existing exit-notification machinery
([process-lifecycle.md](process-lifecycle.md)): `spawn` passes a per-thread
`exit_endpoint`, and `join` blocks in `ipc_reply_wait` until the child-exit
notification for that `tid` arrives, then `munmap`s the stack. No futex needed to
land spawn/join.
- **`join` — Stage 2 refinement** migrates to the std shape: a `completion` word in
the closure that `thread_exit`'s trampoline `futex_wake`s and `join` `futex_wait`s
on — dropping the per-thread endpoint. Kept as a refinement so Stage 1 ships first.
- **`detach`** relinquishes the join right; the kernel reclaims the stack and slot on
`thread_exit` (a detached thread's stack range is unmapped by the reaper, since no
joiner will).
- **`thread_exit`** (no arguments) marks the task dead. The kernel releases the
task's resources, decrements the address-space refcount, and frees the task slot —
the user stack is not the kernel's to unmap; the joiner reclaims it.
- **`join`** is a dedicated `thread_join(tid)` syscall: the caller blocks in the
kernel (`joinThreadLocked`, woken by `wakeJoinersLocked` when the thread exits),
then `munmap`s the stack. (The plan staged join over a per-thread `exit_endpoint`
first, with a futex `completion` word as a Stage-2 refinement; neither shipped — the
dedicated syscall replaced both. `thread_spawn` still accepts an `exit_endpoint`
argument, which the runtime passes as `no_cap`.)
- **`detach`** relinquishes the join right: no one waits for the thread, and its
stack is reclaimed at process exit — kernel-side reclaim of a detached thread's
user stack stays deferred (as the intro notes), since `thread_exit` passes no stack
range.
### Futex, and the sync primitives on top
@@ -219,21 +223,24 @@ decision, not a "maybe later."
### TLS and `getCurrentId`
danos sets up no thread-pointer TLS today (fine under `single_threaded`). Two scoped needs:
Per-thread thread-pointer TLS is in place (the `threadlocal` *compiler* layer is not —
see the intro). Two scoped pieces, as built:
- **`getCurrentId`** returns the kernel task id — either a trivial syscall or, better,
a value the runtime stashes in a per-thread control block.
- **`threadlocal` variables** need a real per-thread TLS block and the thread pointer set per
thread. `thread_spawn` sets the thread pointer to a runtime-allocated per-thread block; full
`threadlocal` support is Stage 3, only if a consumer needs it. Nothing in the core
- **`getCurrentId`** returns the kernel task id via the trivial `thread_self` syscall.
- **The thread pointer** is per-thread: `spawn` carves a small TLS block (an `fs:0`
self-pointer plus scratch) from the top of the thread's own stack, the trampoline
calls `set_thread_pointer` before any user code runs, and the scheduler saves and
restores the pointer per task across context switches. Full `threadlocal` support is
runtime+linker work on top of this, only if a consumer needs it. Nothing in the core
spawn/join/mutex path requires `threadlocal`.
### Build: multi-threaded codegen, opt-in
`addUserBinary` gains a `threaded: bool = false` parameter; when set it builds that
binary `single_threaded = false` so atomics and (later) TLS are real. Threads and
atomics are unsound in a `single_threaded` image, so a binary must opt in **before**
it may call `runtime.Thread.spawn`. Everyone else stays single-threaded and lean.
A binary opts in by being added with `addThreadedUserBinary` — as `addUserBinary`,
but the shared implementation builds it `single_threaded = false` — so atomics and
(later) TLS are real. Threads and atomics are unsound in a `single_threaded` image,
so a binary must opt in **before** it may call `runtime.Thread.spawn`. Everyone else
stays single-threaded and lean.
## Interaction with the rest of the kernel
@@ -243,12 +250,22 @@ it may call `runtime.Thread.spawn`. Everyone else stays single-threaded and lean
process can run on different cores simultaneously — that is the point.
- **Halting** ([halting.md](halting.md)): futex-parked waiters keep the "idle core
halts" property intact under lock contention — no busy-wait.
- **Lifecycle** ([process-lifecycle.md](process-lifecycle.md)): killing a process
must kill *all* its threads and only then drop the last address-space ref. The kill path
already targets a process; it fans out to every task on that address space.
- **Resilience** ([resilience.md](resilience.md)): a faulting thread kills its whole
process (shared fate). The supervisor restarts the **process**, which respawns its
threads from a known-good state — restart granularity stays the process.
- **Lifecycle** ([process-lifecycle.md](process-lifecycle.md)): the contract is that
killing a process kills *all* its threads and only then drops the last address-space
ref. **The kernel does not implement that fan-out yet**: `process_kill` reaps only
the one task it resolves, and no death path loops over the tasks sharing an address
space — the refcount keeps the space (and the sibling threads) alive and running.
The gap is hit in practice: of the only threaded binaries (the `display` service and
the `thread-test` harness), `display` is a boot service that handles no `.terminate`
signal, so init's stop sequence escalates to `process_kill` on every orderly
shutdown — benign only because poweroff follows. Whether to implement the fan-out or
amend the contract is a decision still to be made.
- **Resilience** ([resilience.md](resilience.md)): by the same contract, a faulting
thread kills its whole process (shared fate); the supervisor restarts the
**process**, which respawns its threads from a known-good state — restart
granularity stays the process. Today a CPU fault kills only the faulting task
(`killCurrentProcess` tears down a single task), so sibling threads keep running —
the same implementation gap as above.
- **IPC — two consequences threads forced ([ipc.md](ipc.md)):**
- *Handles do not cross threads.* The handle table lives on the `Task`
([scheduler.zig](../system/kernel/scheduler.zig)), so a handle number is meaningful
@@ -262,8 +279,8 @@ it may call `runtime.Thread.spawn`. Everyone else stays single-threaded and lean
mutate the global service registry, endpoint refcounts, and handle tables. Those paths
were unlocked because a single-threaded process could not race itself; a multi-threaded
one can, from two cores at once. They now take `sync.enter()` like `call`/`reply_wait`/
`send` already did — the kernel heap has no lock of its own yet (heap.zig: "a lock comes
with threads/SMP"), so the big lock is what keeps its callers serialized.
`send` already did — the kernel heap has no lock of its own (heap.zig: "every kernel
entry takes the big kernel lock"), so the big lock is what keeps its callers serialized.
## Build-out plan (staged, each gate serial-checkable)
@@ -278,11 +295,13 @@ a verifiable gate (`python3 test/qemu_test.py <case>`, asserting serial markers;
*Gate:* the full QEMU suite stays green (no regression) — proves the reframing is
invisible until used.
- **Stage 1 — spawn / join / detach.** `thread_spawn` + `thread_exit`, the trampoline,
stacks via `mmap`, join over the exit-endpoint, the `threaded` build flag.
*Gate:* `-Dtest-case=thread-spawn` — a threaded test service spawns N threads that
each `@atomicRmw`-increment a shared counter, the parent joins all N, and asserts
the total is exactly N × iterations. Runs `smp` (multi-core) to prove real
parallelism.
stacks via `mmap`, join over the exit-endpoint (as built, join became the dedicated
`thread_join` syscall instead), the `addThreadedUserBinary` build opt-in.
*Gate:* two cases as built — `-Dtest-case=thread-spawn`, where a worker thread runs
in the caller's address space (a shared-memory write, observed by the main thread),
and `-Dtest-case=thread-join`, where N workers each atomically increment a shared
counter K times, the parent joins all N and asserts the total is exactly N × K —
the join case running multi-core (`smp` 4) to prove real parallelism.
- **Stage 2 — blocking synchronization.** `futex_wait`/`futex_wake` + `Futex`,
`Mutex`, `Condition`, `Semaphore`; optionally migrate join to a futex completion
word. *Gate:* `-Dtest-case=thread-mutex` — a bounded producer/consumer over a