# Threading — build plan (`runtime.Thread` over a private thread ABI) The ordered, checkpointable build-out for [threading.md](threading.md). Each milestone lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run, like [display-v2-plan.md](display-v2-plan.md). Read threading.md first for the *why*. ## Locked decisions (do not relitigate) - **`runtime.Thread` mirrors `std.Thread`'s API; the implementation is danos-native.** Not literal `std.Thread` — that would break the [private ABI](syscall.md). - **Threads are a narrow, per-binary opt-in.** Default concurrency stays process + IPC ([resilience.md](resilience.md)); only a service that asks is built `single_threaded = false`. - **Blocking is futex-backed, never spin-backed** — waiters park in the kernel so an idle core still halts ([halting.md](halting.md)). - **New syscalls are private**: extend [abi.zig](../system/abi.zig) `SystemCall` after `shm_physical = 36` (`thread_spawn = 37`, `thread_exit = 38`, `current_core = 39`, `futex_wait = 40`, `futex_wake = 41`) + a `library/runtime` wrapper; user code never names a number. - **Restart granularity stays the process** — a faulting thread kills its process; the supervisor restarts the process, which respawns its threads. ## Conventions Follow [coding-standards.md](coding-standards.md): spell out non-acronym abbreviations, kebab-case file names, no `Co-Authored-By` trailers. New user binaries go through `addUserBinary` (with the new `threaded` flag where a binary spawns threads) and get packed into the initial-ramdisk; new syscalls extend [abi.zig](../system/abi.zig) `SystemCall` + a `library/runtime` wrapper; test services live beside the code they exercise and register a `ServiceId` if they must be looked up. ## How to verify along the way **Every gate is serial-checkable — no screenshots** (this plan runs unattended). A thread proves it ran by writing to **shared memory** the parent reads back, and proves parallelism by stamping the **core index** it ran on (like the `smp`/`affinity` cases). - `zig build test` — host unit tests (closure packing, mutex state machine, futex wrapper encodings). - `python3 test/qemu_test.py ` — boots the kernel in QEMU; asserts on serial markers. Thread cases set `smp: true` (real parallelism) and bump `mem` (they boot the process/scheduler stack); each milestone **adds its case to `CASES`** so its gate is runnable. - **Guardrail every milestone:** the concurrency-sensitive existing cases stay green — `smoke`, `sched`, `priority`, `smp`, `affinity`, `process`, `process-kill`, `supervision`, `fault-recovery`, `vfs-client-death`, `ipc`/`ipc-cap`, `display-service`. A threading change that regresses those is rejected. ## Unattended execution (the loop contract) This plan runs to completion **without human input**. Every design choice is already fixed in *Locked decisions*; the checkboxes are the only state. A loop iteration must: 1. **Resume** at the first milestone that still has an unchecked `- [ ]`. (All earlier milestones are done — do not revisit them.) 2. **Work on a branch.** On the first iteration, branch off the current `main` into a new branch (e.g. `threading-phase2` — Phase 1's `threading` is already merged); never commit to `main` directly. Push that **branch** to `origin` after each milestone (step 5) so progress is backed up remotely; **do not push `main`** — merging Phase 2 into `main` stays a human step. 3. **Implement** every unchecked item in that milestone, including adding its `-Dtest-case` to `CASES` in [test/qemu_test.py](../test/qemu_test.py) (with `smp: true` / a `mem` bump where noted) so the gate is runnable. 4. **Run the gate**: `python3 test/qemu_test.py `, then the full **guardrail set**, then `zig build` (clean) and `zig build test` (green). 5. **Decide, do not ask:** - **Green** = the milestone's case prints its stated marker(s) and reports `PASS`, the whole guardrail set passes, `zig build` is clean, and host tests are green. → tick this milestone's boxes **and** its `**Gate:**`-referenced case, `git commit` (`threads(M): `, no `Co-Authored-By` trailer per [coding-standards.md](coding-standards.md)), then **`git push` the working branch to `origin`** (use `-u` on the first push to set upstream). Continue to the next milestone in the same iteration if budget remains; otherwise let the loop re-fire. - **Red** = anything above fails. Diagnose from the captured serial log (`zig-out/qemu-test/-failed-serial.log`) and fix in place, then re-run — up to **3 fix attempts** for that gate. A concurrency case that fails then passes on a bare re-run is **flaky, not green**: re-run it **twice more** and treat green only if it passes all; otherwise fix the race (a real threading bug), don't paper over it. 6. **A genuinely ambiguous fork is not a stop.** Pick the option most consistent with [threading.md](threading.md)'s *Locked decisions*, note the choice in the commit message, and continue. Do not pause for confirmation on in-scope, reversible work — this plan is that authorization. **The only stop conditions:** - **Done** — every milestone box **in this plan** is checked (M1 through M11), `zig build` clean, the whole `thread-*` suite + guardrail green. Phase 1 (M1–M6) is *already* checked, so do **not** read that as Done: the loop's real work is the first plan section that still has unchecked boxes — Phase 2 (M7–M11). Only stop when M7–M11 are all checked too. Update threading.md's status line, push the final branch state to `origin`, and stop. The branch is on `origin` for review; **merging Phase 2 into `main` is the user's step**, not the loop's. - **Blocked** — a gate is still red after 3 fix attempts, or a step needs something outside the repo (a toolchain change, new hardware, a decision no locked decision covers). Append `> **BLOCKED (M):** ` under that milestone, commit **and push** the WIP on the branch, and stop. Do not thrash further and do not silently skip the milestone. Nothing else warrants stopping — not "should I proceed?", not "is this right?". The checkboxes + git history are the resumable record; the next iteration picks up from the first unchecked box. --- ## M1 — Address-space refcount (kernel foundation, no API, no behaviour change) ✅ The one invariant change threads require, landed and proven **before** anything shares an address space. Today aspace is 1:1 with a task and teardown destroys it on any user task's exit; make destruction happen on the **last** exit. - [x] A refcount keyed by the address-space root, held in `scheduler.zig` (`aspace_refs`): `retainAspace` takes a reference in `spawnUserLocked` (on the success path, after the slot + stack are secured), all under the big kernel lock. - [x] Both task-teardown paths ([scheduler.zig](../system/kernel/scheduler.zig): `exitUserLocked` and `destroyTaskLocked`) call `releaseAspace`, which decrements and only `destroyAddressSpace`s at **zero**; an unretained space (hand-built test spaces) is destroyed directly, preserving prior behaviour. - [x] `-Dtest-case=aspace-refcount`: spawn and reap several ring-3 processes in sequence and assert (via test-observable `liveAspaceCount`/`aspaceDestroyCount`) that the live-space count returns to **baseline** and destructions advance by exactly that many — each space destroyed exactly once, no leak, no double-free. (Refcount observables, not raw frame counts, since kernel stacks are still leaked on exit.) **Gate (met):** `python3 test/qemu_test.py aspace-refcount` passes (`aspace-refcount: spaces released to baseline ok` → `DANOS-TEST-RESULT: PASS`), and the full guardrail set passes unchanged — 13/13 (`smoke`, `sched`, `priority`, `smp`, `affinity`, `process`, `process-kill`, `supervision`, `fault-recovery`, `vfs-client-death`, `ipc`, `ipc-cap`, `display-service`); default `zig build` clean, `zig build test` green. The reframing is invisible until an aspace is actually shared. ## M2 — `thread_spawn` + `thread_exit`: a thread runs in the shared address space ✅ Spawn only — no join yet. Prove a second task executes in the **caller's** address space and exits cleanly. - [x] [abi.zig](../system/abi.zig): `thread_spawn = 37`, `thread_exit = 38`. Handlers in process.zig; `thread_spawn` calls `scheduler.spawnThread` (shares the caller's aspace, `retainAspace`); `thread_exit` ends the task like a process `exit(0)` (`terminateCurrent` → `releaseAspace`). The closure pointer is delivered in the new thread's **rdi** via a new `jump_to_user_arg` asm path (`t.user_arg`, 0 for a process) — no naked runtime asm. - [x] `library/runtime/thread.zig` (barrel-exported as `runtime.Thread`): `spawn` maps a stack (`mmap`), heap-allocates the `{args}` closure, and calls `thread_spawn(&Closure.entry, stack_top, closure)`; `Closure.entry` (a plain C-ABI Zig fn, closure in rdi) runs the function and calls `thread_exit`. Stack top is 16-aligned-minus-8 for the C entry. - [x] A `threaded` flag on the user-binary recipe (`addThreadedUserBinary` → `single_threaded = false`); `thread-test` is the first opt-in binary. - [x] `-Dtest-case=thread-spawn`: `thread-test` spawns a worker that writes a sentinel to a **shared** global and release-stores `done`; the main thread acquire-polls `done` and asserts the shared global holds the sentinel — proof the worker ran in the same address space. **Gate (met):** `python3 test/qemu_test.py thread-spawn` passes (`thread-test: child ran in shared aspace ok` → `DANOS-TEST-RESULT: PASS`); guardrail set 16/16 green (incl. `args`/`init`/`process`, which exercise the new `jump_to_user_arg` process path with arg 0) plus `aspace-refcount`; `zig build` clean, `zig build test` green. > **Note (deferred to M3+):** the mmap arena is per-*task* (`heap_next`), so two threads > in one aspace that both `mmap` would collide. Fine for M2 (only the parent maps, for the > child's stack); make the arena per-aspace and the runtime heap thread-safe alongside the > `Mutex` work (M5). ## M3 — `join` + `detach` + real parallelism ✅ - [x] `join` over the existing exit-notification path ([process-lifecycle.md](process-lifecycle.md)): `thread_spawn` gained a 4th arg, an `exit_endpoint` handle (resolved + refcounted like `spawnProcessSupervised`, via `spawnThreadSupervised`); `join` blocks in `ipc_reply_wait` on that endpoint until the child-exit notice for its `tid`, then `munmap`s the stack. `detach` relinquishes the join right (its stack is reclaimed at process exit — kernel-reaper reclaim for detached threads is deferred; see note). - [x] `runtime.Thread.join` / `detach`, plus `Thread.currentCore()` (a new `current_core` = 39 syscall) for the parallelism proof. `getCurrentId` deferred to M6 (TLS), where a lighter self-id fits. The closure now rides the **thread's own stack** (not the heap) — private per thread, so spawn/join touch no shared heap. - [x] `-Dtest-case=thread-join` (`smp: 4`): `thread-test` join mode spawns N=4 workers that each do K=100k `@atomicRmw`-increments on a shared counter and stamp the core they ran on; the main thread joins all N and asserts `counter == N*K` **and** `@popCount(cores_seen) > 1` (genuine cross-core parallelism), then a detached worker proves `detach` runs without a join. **Gate (met):** `python3 test/qemu_test.py thread-join` passes (`thread-test: join ok` → `DANOS-TEST-RESULT: PASS`), robust across 4 runs; guardrail 17/17 green (incl. `smp`, `affinity`, `process-kill`, and `args`/`init`/`process` on the exit-endpoint spawn path) plus `aspace-refcount`/`thread-spawn`; `zig build` clean, `zig build test` green. > **Note (deferred):** a detached thread's stack is freed only at process exit (not by the > reaper on thread exit) — kernel user-stack tracking + reclaim is a later refinement. And > the runtime heap is still not thread-safe: threads that both allocate concurrently would > race (the thread *machinery* avoids the heap, but worker code sharing an allocator does > not). Both fold into the M5 `Mutex`/allocator work. ## M4 — Futex: the one blocking primitive ✅ - [x] [abi.zig](../system/abi.zig): `futex_wait = 40`, `futex_wake = 41`. A waiter is a `.blocked` task tagged with `Task.futex_addr` (no queue linkage); `futex_wait(addr, expected, timeout_ns)` reads the user word under the big lock, parks iff `*addr == expected`, and returns on wake or timeout; `futex_wake(addr, count)` scans the task table and readies up to `count` matching waiters (same address space). No spinning — a parked waiter leaves its core free to `hlt`. A timed wait also sets `wake_at`, so the timer's `wakeExpired` wakes it; `futex_addr` staying non-zero (only `futex_wake` clears it) is how the waiter tells timeout from a real wake. - [x] `runtime.Thread.Futex` (`wait` / `timedWait` / `wake`) over the syscall wrappers. - [x] `-Dtest-case=thread-futex` (`smp: 4`): a waiter thread prints `waiting` and `futex_wait`s on a word; the main thread publishes it, prints `waking`, and `futex_wake`s; the waiter prints `woke`. Then a `timedWait` on an unwoken word reports `error.Timeout`. **Gate (met):** `python3 test/qemu_test.py thread-futex` passes, robust across 3 runs — the case's **ordered** regex asserts `waiting → waking → woke → PASS` on the serial stream (the handoff proof), and `thread-futex: timeout ok` confirms the timeout. Guardrail 18/18 green (incl. `sleep`/`event`/`ipc` blocking paths) + `aspace-refcount`, `thread-spawn`, `thread-join`; `zig build` clean, `zig build test` green. > **Note:** the kernel test checks only the freshest verdict marker via `bufferHas` (the > in-memory log ring buffer evicts older lines); ordering is asserted against the full > serial stream by the qemu regex instead. ## M5 — `Mutex` + `Condition` + `Semaphore` ✅ - [x] `runtime.Thread.Mutex` (three-state futex mutex: CAS fast path, `futex_wait`/`wake` slow path), `Condition` (`wait`/`timedWait`/`signal`/`broadcast`, a futex sequence counter), `Semaphore` (permits over `Mutex`+`Condition`) — the same state machines `std.Thread` uses, ported onto our `Futex`. - [x] `-Dtest-case=thread-mutex` (`smp: 4`): a bounded producer/consumer — 2 producers + 2 consumers over one `Mutex` and two `Condition`s move N=2000 unique items through an 8-slot ring; the consumed checksum and tally match exactly (no lost/duplicated item, no overrun) under real cross-core contention. The small ring forces producers to block on full and consumers on empty, exercising `Condition.wait`. **Gate (met):** `python3 test/qemu_test.py thread-mutex` passes (`thread-mutex: ok` → `DANOS-TEST-RESULT: PASS`), robust across 3 runs; guardrail 17/17 green (incl. `sleep`/`event`/`ipc`) + all M1–M4 thread cases; `zig build` clean, `zig build test` green. > **Deferred (with rationale):** > - **`join` → futex completion word** — the exit-endpoint join (M3) is correct and > tested. A futex-completion join needs the *kernel* to clear+wake a word after the > thread is fully off its stack (a CLONE_CHILD_CLEARTID-style mechanism); doing it in > the thread's own trampoline would let `join` `munmap` the stack while the thread still > runs on it (use-after-free). Left on the exit-endpoint path; the kernel clear-on-exit > is a later, separate refinement. > - **Host unit tests for the state machines** — `Mutex`/`Condition` bottom out in the > `futex_*` syscalls, unavailable on the host without a mockable `Futex` seam. The QEMU > `thread-mutex` gate exercises them under real concurrency instead; a host-side mock is > future work. ## M6 — `getCurrentId`, docs, and CI wiring ✅ - [x] `getCurrentId` via a small `thread_self = 42` syscall (`runtime.Thread.getCurrentId` returns the kernel task id). **Per-thread `threadlocal` TLS is deferred** — no consumer needs it, and it would require context-switching `fs.base` per task (real kernel + per-switch cost) for an unused feature; threaded binaries have run fine without it through M2–M5. threading.md's TLS reasoning already scoped it as deferred-unless-needed. When a consumer appears, the shape is: `thread_spawn` allocates a per-thread TLS block, sets `fs.base`, and the context switch saves/ restores it. - [x] `RwLock` / `WaitGroup` deferred (no consumer yet); they slot onto the same `Futex`/`Mutex`/`Condition` when wanted. - [x] All `thread-*` cases wired into [test/qemu_test.py](../test/qemu_test.py) (`thread-spawn`/`-join`/`-futex`/`-mutex`/`-id`); threading.md + docs/README.md status updated to **built**; the worked example is threading.md's win-condition. - [x] `-Dtest-case=thread-id` (`smp: 4`): two workers read `getCurrentId`; the main thread confirms all three ids are non-zero and distinct — each thread has its own kernel identity. (Renamed from `thread-tls`, which implied `threadlocal`.) **Gate (met):** `python3 test/qemu_test.py thread-id` passes; the whole `thread-*` suite (`thread-spawn`/`-join`/`-futex`/`-mutex`/`-id`) plus the full guardrail set pass; default `zig build` clean, `zig build test` green. --- ## Status **Phase 1 (M1–M6): built.** danos has `runtime.Thread` — `spawn`/`join`/`detach`, cross-core parallelism, futex, and `Mutex`/`Condition`/`Semaphore`, all over a private thread ABI behind the runtime. **Phase 2 (M7–M11): planned below** — hardening the deferred parts so threads are safe for real workloads and reclaimed like everything else danos owns. --- ## Phase 2 — hardening (M7–M11) The organising principle, so Phase 2 reinforces danos's goals rather than eroding them: - **Everything a thread owns is reclaimed on process death.** Thread stacks, TLS blocks, and futex words live in the process's **address space**, and the kernel's per-process state is keyed by the aspace root — so the M1 refcount + `destroyAddressSpace` already free all of it when the last thread exits. A crashed or killed threaded process leaves **nothing** behind. Phase 2 closes the one thing that is *not* aspace-owned — the per-task **kernel** stack (kernel heap) — with a reaper (M8). This is the [resilience](resilience.md) restart guarantee, extended to threads. - **Kernel owns mechanism; the runtime owns policy.** The kernel maps pages, saves/ restores `fs.base`, and reaps dead tasks; the runtime decides allocation, TLS layout, and lock algorithms. Every new kernel entry stays a private syscall behind the runtime ([syscall.md](syscall.md)) — the ABI stays renumberable. - **The process is still the isolation and restart boundary.** Threads share fate within one process; Phase 2 never adds a way for one process to reach into another (the cross-process futex stays explicitly out of scope, below). ### M7 — Thread-safe allocation (the correctness gap) ✅ Today the mmap arena cursor is per-*task* and the runtime heap is unlocked, so two threads in one process that both allocate corrupt each other. The thread *machinery* avoids this (closure on the stack, stacks mmap'd only by the spawner), but real multi-threaded code would hit it. Closed it: - [x] **Kernel — per-address-space mmap arena.** Grew M1's `aspace_refs` entry into the per-address-space object holding the `mmap`/`mmio` arena cursors (moved off `Task`); `scheduler.aspaceMmapNextPtr`/`aspaceDeviceMapNextPtr` expose them. `systemMmap` reserves a disjoint range under a *brief* lock, then maps **per page** under a short-held lock — not the whole grant — because the big lock is held with interrupts disabled, so pinning it across a multi-MiB memset+map froze other cores (it timed the `affinity` scenario out mid-bring-up). Freed at refcount zero, so the cursors vanish with the process. - [x] **Runtime — thread-safe heap.** The allocator's two free-list mutators (`rawAlloc`/`rawFree`) take a `Thread.Mutex`, gated on `!@import("builtin").single_threaded` so single-threaded binaries compile it out and pay nothing. Uncontended acquisition is a single CAS (no syscall). - [x] `-Dtest-case=thread-alloc` (`smp: 4`): 4 threads each do 500 `alloc`/fill/verify/ `free` cycles of varied sizes; each block is filled with a per-thread pattern and verified before free, so any overlap between concurrent allocations is caught. **Gate (met):** `thread-alloc` passes (3× non-flaky); full guardrail 23/23 green, `zig build`/`zig build test` clean. > **Also fixed here:** the `affinity` guardrail's fixed-count busy-loop (`while (spins < > 3e9)`) had codegen-dependent wall-time — adding a function to `tests.zig` flipped how > the optimiser compiled it, swinging affinity from ~4 s to ~63 s and timing it out. > Reworked it (and the settle loop) to wait on the wall clock instead, so its duration is > independent of unrelated code changes. ### M8 — The task reaper (cleanup + resilience) A dead task's **kernel** stack is currently leaked ("no reaper yet") — every process *and* thread death loses one, so a crash loop bleeds kernel memory. A reaper fixes it and serves the [resilience](resilience.md) restart goal directly: - [ ] A dying task cannot free the kernel stack it runs on, so it hands itself to a **reap list** and switches away; the kernel stack (and, for a detached thread, its user stack) is reclaimed from another context — a low-priority reaper step drained on the scheduler tick and when a core goes idle. Extends the existing `reap_task_hook`/`destroyTaskLocked` path rather than inventing a parallel one. - [ ] `-Dtest-case=task-reap`: spawn and exit many threads and processes; assert the kernel-heap free bytes (a new test observable) return to **baseline** — kernel stacks reclaimed, no leak — and that the `fault-recovery`/kill paths reclaim too. **Gate:** `task-reap` passes; `fault-recovery`, `supervision`, `process-kill`, `aspace-refcount` still green. ### M9 — Futex-completion join (retire the per-thread endpoint) With the reaper (M8) able to act *after* a thread is fully off its stack, migrate `join` to the std shape and drop M3's per-thread exit endpoint: - [ ] `thread_spawn` takes a user **completion word** (in the `Thread` handle's memory) and a joinable/detached flag. On reap the kernel writes 0 to that word and `futex_wake`s it (a `CLONE_CHILD_CLEARTID` equivalent — safe now the thread is off its stack). `join` = `futex_wait` on the word, then `munmap` the stack; a **detached** thread's stack is `munmap`ped by the reaper instead. No IPC endpoint per thread. - [ ] `thread-join` passes on the new path; a check confirms joining N threads creates no per-thread endpoints (handle count stable). **Gate:** `thread-join`/`thread-mutex` green on futex-completion join; guardrail green. ### M10 — Per-thread TLS (`threadlocal`) Give each thread its own `threadlocal` storage — the piece self-hosting Zig ([zig-self-hosting.md](zig-self-hosting.md)) will force: - [ ] **Runtime** allocates a per-thread TLS block from the binary's `PT_TLS` template (linker symbols: copy `.tdata`, zero `.tbss`, variant-II TCB self-pointer) and hands its thread pointer to `thread_spawn`; the main thread sets its own via a new `set_thread_pointer` syscall in `_start`. The block is aspace memory → reclaimed on teardown. - [ ] **Kernel** stores `fs_base` on `Task`, loads it at first entry and restores it on context switch only when it changes (the same conditional-load pattern as CR3). `getCurrentId` can then read a TLS self-slot instead of a syscall. - [ ] `-Dtest-case=thread-tls`: two threads each write and read their own `threadlocal` slot with no cross-talk, and observe distinct `getCurrentId`. **Gate:** `thread-tls` passes; full `thread-*` suite + guardrail green. ### M11 — `RwLock`, `WaitGroup`, and host-testable sync - [ ] `runtime.Thread.RwLock` and `WaitGroup` on the existing `Futex`/`Mutex`/ `Condition`. - [ ] A compile-time `Futex` seam: syscalls on the danos target, a host-backed impl under `zig build test`, so the `Mutex`/`Condition`/`RwLock` state machines run as host unit tests (fast iteration, no QEMU). - [ ] `-Dtest-case=thread-rwlock` (`smp: 4`): many readers + writers over an `RwLock` keep an invariant (a reader never observes a half-written value); host tests cover the lock transitions. **Gate:** host `zig build test` covers the sync primitives; `thread-rwlock` passes; guardrail green. --- ## Deferred (explicitly not in this plan) - **Cross-process shared-memory futex** — the `(aspace, vaddr)` key can become a physical-address key so two processes share a futex through an [shm](display-v2.md) region. Not needed for intra-process threads. - **Per-thread priorities / affinity distinct from the process** — threads inherit the process priority ([scheduling.md](scheduling.md)); revisit only if it earns its keep. - **Per-thread signal delivery** — signals stay process-scoped ([process-lifecycle.md](process-lifecycle.md)). - **A `pthread`/POSIX surface** — the API is `std.Thread`-shaped Zig, nothing more. - **A real `std.Thread` backend** — arrives with self-hosting ([zig-self-hosting.md](zig-self-hosting.md)); it sits on these same primitives, so it swaps the impl under `runtime.Thread`, not the call sites.