docs: full docs-vs-code audit — fix every stale claim across 40 docs
Every doc verified claim-by-claim against the code by parallel audit agents, then fixed and adversarially re-verified. Two waves of staleness corrected: the originally audited findings (higher-half boot handoff, kernel VFS takeover, fault isolation + claim release + driver restart, AML/S5 moving to ring 3, threading's shipped design, USB+FAT landing) and a second pass of adjacent claims the verifiers caught (smp.md 'not built yet' intro, system-requirements' PS/2-only and no-storage claims, halting.md's red-panic and no-IDT text, testing.md's serial mirroring, router-era vfs-protocol wording, capsule-first boot loading). threading.md now documents the shared-fate gap explicitly: the design says a process dies whole, the kernel today kills only the offending thread. Also fixes three stale code comments (isr.s exceptionHandler, acpi.zig sleepValue, build.zig boot-volume) — comments only, no behavior change.
This commit is contained in:
+91
-72
@@ -81,9 +81,11 @@ address space. Threads deliberately remove that boundary *within* a process:
|
||||
|
||||
- Threads share one address space, so one thread's stray write corrupts them all —
|
||||
there is no isolation **between** threads.
|
||||
- Threads share fate: a fault in any thread, or a "kill the process" decision, takes
|
||||
down **all** of them. Restartability lives at the process level, not the thread
|
||||
level.
|
||||
- Threads share fate — by contract: a fault in any thread, or a "kill the process"
|
||||
decision, takes down **all** of them, so restartability lives at the process level,
|
||||
not the thread level. (The kernel does not yet enforce this fan-out — see the
|
||||
Lifecycle note under
|
||||
[Interaction with the rest of the kernel](#interaction-with-the-rest-of-the-kernel).)
|
||||
- Shared mutable state reintroduces data races — the failure class the
|
||||
isolate-and-message model was chosen to avoid.
|
||||
|
||||
@@ -103,16 +105,15 @@ Lives in `library/runtime/thread.zig`, re-exported as `runtime.Thread`.
|
||||
pub const Thread = struct {
|
||||
pub const Id = u32; // the kernel task id
|
||||
pub const SpawnConfig = struct {
|
||||
stack_size: usize = default_stack_size,
|
||||
allocator: ?std.mem.Allocator = null, // for the closure + stack bookkeeping
|
||||
stack_size: usize = default_stack_size, // no allocator: the closure lives at the top of the thread's own stack
|
||||
};
|
||||
pub const SpawnError = error{ OutOfMemory, ThreadQuotaExceeded, SystemResources };
|
||||
pub const SpawnError = error{SystemResources};
|
||||
|
||||
pub fn spawn(config: SpawnConfig, comptime function: anytype, args: anytype) SpawnError!Thread;
|
||||
pub fn join(self: Thread) void; // block until the thread ends, reclaim its stack
|
||||
pub fn detach(self: Thread) void; // give up the right to join; kernel reclaims on exit
|
||||
pub fn detach(self: Thread) void; // give up the right to join; stack reclaimed at process exit
|
||||
pub fn getCurrentId() Id;
|
||||
pub fn yield() void; // -> existing `yield` syscall
|
||||
pub fn currentCore() Id; // danos extension: the calling core's dense index
|
||||
|
||||
pub const Mutex = struct { pub fn lock(*Mutex) void; pub fn tryLock(*Mutex) bool; pub fn unlock(*Mutex) void; };
|
||||
pub const Condition = struct { pub fn wait(*Condition, *Mutex) void; pub fn timedWait(*Condition, *Mutex, u64) error{Timeout}!void; pub fn signal(*Condition) void; pub fn broadcast(*Condition) void; };
|
||||
@@ -127,18 +128,21 @@ Deviations from `std.Thread`, called out honestly:
|
||||
- **The thread function's return value is discarded** (as `std.Thread.join` returns
|
||||
`void`). Return data through shared state or a `Semaphore`/`Condition`, not the
|
||||
return.
|
||||
- `getCpuCount()` maps to the existing SMP core count ([smp.md](smp.md)); a service
|
||||
rarely needs it.
|
||||
- No `getCpuCount()` (a service rarely needs it) and no `Thread.yield()` — `yield`
|
||||
lives in `runtime.system`. Instead `currentCore()` exposes the calling core's dense
|
||||
index ([smp.md](smp.md)), used to observe genuine cross-core parallelism.
|
||||
|
||||
## Kernel primitives (new private syscalls)
|
||||
|
||||
Four new entries extend [abi.zig](../system/abi.zig) `SystemCall` after
|
||||
`shared_memory_physical = 36`, each with a `library/runtime` wrapper:
|
||||
Five core entries extend [abi.zig](../system/abi.zig) `SystemCall` after
|
||||
`shared_memory_physical = 36` (plus small helpers `current_core`, `thread_self`, and
|
||||
`set_thread_pointer`), each with a `library/runtime` wrapper:
|
||||
|
||||
| Syscall | Signature | Purpose |
|
||||
|---|---|---|
|
||||
| `thread_spawn` | `(entry, stack_top, arg) -> tid` | create a task sharing the **caller's** address space |
|
||||
| `thread_exit` | `(stack_base, stack_len)` | end the calling thread; hand back its stack range for reclaim |
|
||||
| `thread_spawn` | `(entry, stack_top, arg, exit_endpoint) -> tid` | create a task sharing the **caller's** address space; the runtime passes `no_cap` for `exit_endpoint` (join is a syscall, not an endpoint) |
|
||||
| `thread_exit` | `()` | end the calling thread; its stack is reclaimed by the joiner's `munmap`, not the kernel |
|
||||
| `thread_join` | `(tid) -> 0` | block until the thread with id `tid` has exited |
|
||||
| `futex_wait` | `(addr, expected, timeout_ns) -> status` | block if `*addr == expected`, until woken or timeout |
|
||||
| `futex_wake` | `(addr, count) -> woken` | wake up to `count` waiters on `addr` |
|
||||
|
||||
@@ -148,56 +152,56 @@ Plus one invariant change with no new syscall: **address-space reference countin
|
||||
|
||||
### Address-space reference counting
|
||||
|
||||
Today an address space is 1:1 with a task: `spawnUserLocked` records `address_space` on the
|
||||
Task, and teardown does `destroyAddressSpace(t.address_space)` when **any** user task exits
|
||||
Before this work an address space was 1:1 with a task: `spawnUserLocked` records
|
||||
`address_space` on the Task (as it still does), and teardown did
|
||||
`destroyAddressSpace(t.address_space)` when **any** user task exited
|
||||
([scheduler.zig](../system/kernel/scheduler.zig)). With threads, several tasks share
|
||||
one `address_space`, so the first to exit would rip the address space out from under its
|
||||
siblings.
|
||||
|
||||
Fix: a small refcount keyed by the address-space root (`createAddressSpace` in
|
||||
[process.zig](../system/kernel/process.zig) sets it to 1). `thread_spawn` increments
|
||||
it; task teardown decrements and only calls `destroyAddressSpace` at **zero**. All of
|
||||
Fix: a small refcount keyed by the address-space root, kept in
|
||||
[scheduler.zig](../system/kernel/scheduler.zig): `retainAddressSpace` takes a
|
||||
reference for every user task `spawnUserLocked` starts (count 1 on the first take, so
|
||||
a thread sharing the caller's space increments it); task teardown calls
|
||||
`releaseAddressSpace`, which only calls `destroyAddressSpace` at **zero**. All of
|
||||
this is already under the big kernel lock, so no new locking. This is the one piece
|
||||
that must land and be proven before anything shares an address space.
|
||||
|
||||
### `thread_spawn` and the trampoline
|
||||
|
||||
The scheduler already accepts an arbitrary `address_space` and does **not** smuggle values
|
||||
through registers — `startUserTask` reads the entry/stack from the Task and
|
||||
`jumpToUser`s ([scheduler.zig](../system/kernel/scheduler.zig)). That makes the thread
|
||||
path clean:
|
||||
The scheduler already accepts an arbitrary `address_space` and does **not** smuggle
|
||||
values through scratch registers — `startUserTask` reads the entry/stack (and the
|
||||
thread's closure arg, delivered in `rdi` via `jumpToUserArg`) from the Task
|
||||
([scheduler.zig](../system/kernel/scheduler.zig)). That makes the thread path clean:
|
||||
|
||||
1. The runtime's `spawn` `mmap`s a stack (syscall `4`), heap-allocates a closure —
|
||||
`{ fn_ptr, args_tuple, completion }`, the std "Instance" pattern — and writes the
|
||||
closure pointer to the **top word of the new stack**.
|
||||
2. It calls `thread_spawn(entry = &threadTrampoline, stack_top, arg = closure_ptr)`.
|
||||
The kernel calls the same `spawnUserLocked` path with the **caller's address space**
|
||||
(refcount++), `entry`, and `user_sp = stack_top`.
|
||||
3. `threadTrampoline` (a small runtime shim) reads the closure off its stack, calls
|
||||
the user function, then calls `thread_exit`. No new register ABI — the closure
|
||||
pointer rides the stack the runtime set up, mirroring how `startUserTask` avoids
|
||||
register smuggling.
|
||||
1. The runtime's `spawn` `mmap`s a stack (syscall `4`) and writes the closure —
|
||||
`{ tls_base, args }`, the std "Instance" pattern — at the **top of the new stack
|
||||
itself** (no heap allocation), with a small per-thread TLS block just below it.
|
||||
2. It calls `thread_spawn(entry = &Closure.entry, stack_top, arg = closure_ptr,
|
||||
exit_endpoint = no_cap)`. The kernel calls the same `spawnUserLocked` path with the
|
||||
**caller's address space** (refcount++), `entry`, and `user_sp = stack_top`.
|
||||
3. `Closure.entry` (a small runtime shim) receives the closure pointer in `rdi` — the
|
||||
kernel delivers `arg` as the entry's first C-ABI argument — sets the thread
|
||||
pointer, calls the user function, then calls `thread_exit`.
|
||||
|
||||
Unlike a process start, there is **no** System V argc/argv/auxv block
|
||||
([sysv.md](sysv.md)) — a thread stack carries only the closure pointer.
|
||||
([sysv.md](sysv.md)) — a thread stack carries only the closure and its TLS block.
|
||||
|
||||
### Lifetime: exit, join, detach, stack reclaim
|
||||
|
||||
- **`thread_exit`** marks the task dead and hands the kernel the thread's user-stack
|
||||
range. The kernel reaps the task on the scheduler (already running on a *kernel*
|
||||
stack, so it can safely unmap the user stack), decrements the address-space refcount, and
|
||||
frees the task slot.
|
||||
- **`join` — Stage 1** reuses the existing exit-notification machinery
|
||||
([process-lifecycle.md](process-lifecycle.md)): `spawn` passes a per-thread
|
||||
`exit_endpoint`, and `join` blocks in `ipc_reply_wait` until the child-exit
|
||||
notification for that `tid` arrives, then `munmap`s the stack. No futex needed to
|
||||
land spawn/join.
|
||||
- **`join` — Stage 2 refinement** migrates to the std shape: a `completion` word in
|
||||
the closure that `thread_exit`'s trampoline `futex_wake`s and `join` `futex_wait`s
|
||||
on — dropping the per-thread endpoint. Kept as a refinement so Stage 1 ships first.
|
||||
- **`detach`** relinquishes the join right; the kernel reclaims the stack and slot on
|
||||
`thread_exit` (a detached thread's stack range is unmapped by the reaper, since no
|
||||
joiner will).
|
||||
- **`thread_exit`** (no arguments) marks the task dead. The kernel releases the
|
||||
task's resources, decrements the address-space refcount, and frees the task slot —
|
||||
the user stack is not the kernel's to unmap; the joiner reclaims it.
|
||||
- **`join`** is a dedicated `thread_join(tid)` syscall: the caller blocks in the
|
||||
kernel (`joinThreadLocked`, woken by `wakeJoinersLocked` when the thread exits),
|
||||
then `munmap`s the stack. (The plan staged join over a per-thread `exit_endpoint`
|
||||
first, with a futex `completion` word as a Stage-2 refinement; neither shipped — the
|
||||
dedicated syscall replaced both. `thread_spawn` still accepts an `exit_endpoint`
|
||||
argument, which the runtime passes as `no_cap`.)
|
||||
- **`detach`** relinquishes the join right: no one waits for the thread, and its
|
||||
stack is reclaimed at process exit — kernel-side reclaim of a detached thread's
|
||||
user stack stays deferred (as the intro notes), since `thread_exit` passes no stack
|
||||
range.
|
||||
|
||||
### Futex, and the sync primitives on top
|
||||
|
||||
@@ -219,21 +223,24 @@ decision, not a "maybe later."
|
||||
|
||||
### TLS and `getCurrentId`
|
||||
|
||||
danos sets up no thread-pointer TLS today (fine under `single_threaded`). Two scoped needs:
|
||||
Per-thread thread-pointer TLS is in place (the `threadlocal` *compiler* layer is not —
|
||||
see the intro). Two scoped pieces, as built:
|
||||
|
||||
- **`getCurrentId`** returns the kernel task id — either a trivial syscall or, better,
|
||||
a value the runtime stashes in a per-thread control block.
|
||||
- **`threadlocal` variables** need a real per-thread TLS block and the thread pointer set per
|
||||
thread. `thread_spawn` sets the thread pointer to a runtime-allocated per-thread block; full
|
||||
`threadlocal` support is Stage 3, only if a consumer needs it. Nothing in the core
|
||||
- **`getCurrentId`** returns the kernel task id via the trivial `thread_self` syscall.
|
||||
- **The thread pointer** is per-thread: `spawn` carves a small TLS block (an `fs:0`
|
||||
self-pointer plus scratch) from the top of the thread's own stack, the trampoline
|
||||
calls `set_thread_pointer` before any user code runs, and the scheduler saves and
|
||||
restores the pointer per task across context switches. Full `threadlocal` support is
|
||||
runtime+linker work on top of this, only if a consumer needs it. Nothing in the core
|
||||
spawn/join/mutex path requires `threadlocal`.
|
||||
|
||||
### Build: multi-threaded codegen, opt-in
|
||||
|
||||
`addUserBinary` gains a `threaded: bool = false` parameter; when set it builds that
|
||||
binary `single_threaded = false` so atomics and (later) TLS are real. Threads and
|
||||
atomics are unsound in a `single_threaded` image, so a binary must opt in **before**
|
||||
it may call `runtime.Thread.spawn`. Everyone else stays single-threaded and lean.
|
||||
A binary opts in by being added with `addThreadedUserBinary` — as `addUserBinary`,
|
||||
but the shared implementation builds it `single_threaded = false` — so atomics and
|
||||
(later) TLS are real. Threads and atomics are unsound in a `single_threaded` image,
|
||||
so a binary must opt in **before** it may call `runtime.Thread.spawn`. Everyone else
|
||||
stays single-threaded and lean.
|
||||
|
||||
## Interaction with the rest of the kernel
|
||||
|
||||
@@ -243,12 +250,22 @@ it may call `runtime.Thread.spawn`. Everyone else stays single-threaded and lean
|
||||
process can run on different cores simultaneously — that is the point.
|
||||
- **Halting** ([halting.md](halting.md)): futex-parked waiters keep the "idle core
|
||||
halts" property intact under lock contention — no busy-wait.
|
||||
- **Lifecycle** ([process-lifecycle.md](process-lifecycle.md)): killing a process
|
||||
must kill *all* its threads and only then drop the last address-space ref. The kill path
|
||||
already targets a process; it fans out to every task on that address space.
|
||||
- **Resilience** ([resilience.md](resilience.md)): a faulting thread kills its whole
|
||||
process (shared fate). The supervisor restarts the **process**, which respawns its
|
||||
threads from a known-good state — restart granularity stays the process.
|
||||
- **Lifecycle** ([process-lifecycle.md](process-lifecycle.md)): the contract is that
|
||||
killing a process kills *all* its threads and only then drops the last address-space
|
||||
ref. **The kernel does not implement that fan-out yet**: `process_kill` reaps only
|
||||
the one task it resolves, and no death path loops over the tasks sharing an address
|
||||
space — the refcount keeps the space (and the sibling threads) alive and running.
|
||||
The gap is hit in practice: of the only threaded binaries (the `display` service and
|
||||
the `thread-test` harness), `display` is a boot service that handles no `.terminate`
|
||||
signal, so init's stop sequence escalates to `process_kill` on every orderly
|
||||
shutdown — benign only because poweroff follows. Whether to implement the fan-out or
|
||||
amend the contract is a decision still to be made.
|
||||
- **Resilience** ([resilience.md](resilience.md)): by the same contract, a faulting
|
||||
thread kills its whole process (shared fate); the supervisor restarts the
|
||||
**process**, which respawns its threads from a known-good state — restart
|
||||
granularity stays the process. Today a CPU fault kills only the faulting task
|
||||
(`killCurrentProcess` tears down a single task), so sibling threads keep running —
|
||||
the same implementation gap as above.
|
||||
- **IPC — two consequences threads forced ([ipc.md](ipc.md)):**
|
||||
- *Handles do not cross threads.* The handle table lives on the `Task`
|
||||
([scheduler.zig](../system/kernel/scheduler.zig)), so a handle number is meaningful
|
||||
@@ -262,8 +279,8 @@ it may call `runtime.Thread.spawn`. Everyone else stays single-threaded and lean
|
||||
mutate the global service registry, endpoint refcounts, and handle tables. Those paths
|
||||
were unlocked because a single-threaded process could not race itself; a multi-threaded
|
||||
one can, from two cores at once. They now take `sync.enter()` like `call`/`reply_wait`/
|
||||
`send` already did — the kernel heap has no lock of its own yet (heap.zig: "a lock comes
|
||||
with threads/SMP"), so the big lock is what keeps its callers serialized.
|
||||
`send` already did — the kernel heap has no lock of its own (heap.zig: "every kernel
|
||||
entry takes the big kernel lock"), so the big lock is what keeps its callers serialized.
|
||||
|
||||
## Build-out plan (staged, each gate serial-checkable)
|
||||
|
||||
@@ -278,11 +295,13 @@ a verifiable gate (`python3 test/qemu_test.py <case>`, asserting serial markers;
|
||||
*Gate:* the full QEMU suite stays green (no regression) — proves the reframing is
|
||||
invisible until used.
|
||||
- **Stage 1 — spawn / join / detach.** `thread_spawn` + `thread_exit`, the trampoline,
|
||||
stacks via `mmap`, join over the exit-endpoint, the `threaded` build flag.
|
||||
*Gate:* `-Dtest-case=thread-spawn` — a threaded test service spawns N threads that
|
||||
each `@atomicRmw`-increment a shared counter, the parent joins all N, and asserts
|
||||
the total is exactly N × iterations. Runs `smp` (multi-core) to prove real
|
||||
parallelism.
|
||||
stacks via `mmap`, join over the exit-endpoint (as built, join became the dedicated
|
||||
`thread_join` syscall instead), the `addThreadedUserBinary` build opt-in.
|
||||
*Gate:* two cases as built — `-Dtest-case=thread-spawn`, where a worker thread runs
|
||||
in the caller's address space (a shared-memory write, observed by the main thread),
|
||||
and `-Dtest-case=thread-join`, where N workers each atomically increment a shared
|
||||
counter K times, the parent joins all N and asserts the total is exactly N × K —
|
||||
the join case running multi-core (`smp` 4) to prove real parallelism.
|
||||
- **Stage 2 — blocking synchronization.** `futex_wait`/`futex_wake` + `Futex`,
|
||||
`Mutex`, `Condition`, `Semaphore`; optionally migrate join to a futex completion
|
||||
word. *Gate:* `-Dtest-case=thread-mutex` — a bounded producer/consumer over a
|
||||
|
||||
Reference in New Issue
Block a user