# Threading — build plan (`runtime.Thread` over a private thread ABI) The ordered, checkpointable build-out for [threading.md](threading.md). Each milestone lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run, like [display-v2-plan.md](display-v2-plan.md). Read threading.md first for the *why*. ## Locked decisions (do not relitigate) - **`runtime.Thread` mirrors `std.Thread`'s API; the implementation is danos-native.** Not literal `std.Thread` — that would break the [private ABI](syscall.md). - **Threads are a narrow, per-binary opt-in.** Default concurrency stays process + IPC ([resilience.md](resilience.md)); only a service that asks is built `single_threaded = false`. - **Blocking is futex-backed, never spin-backed** — waiters park in the kernel so an idle core still halts ([halting.md](halting.md)). - **New syscalls are private**: extend [abi.zig](../system/abi.zig) `SystemCall` after `shared_memory_physical = 36` (`thread_spawn = 37`, `thread_exit = 38`, `current_core = 39`, `futex_wait = 40`, `futex_wake = 41`) + a `library/runtime` wrapper; user code never names a number. - **Restart granularity stays the process** — a faulting thread kills its process; the supervisor restarts the process, which respawns its threads. ## Conventions Follow [coding-standards.md](coding-standards.md): spell out non-acronym abbreviations, kebab-case file names, no `Co-Authored-By` trailers. New user binaries go through `addUserBinary` (with the new `threaded` flag where a binary spawns threads) and get packed into the initial-ramdisk; new syscalls extend [abi.zig](../system/abi.zig) `SystemCall` + a `library/runtime` wrapper; test services live beside the code they exercise and register a `ServiceId` if they must be looked up. ## How to verify along the way **Every gate is serial-checkable — no screenshots** (this plan runs unattended). A thread proves it ran by writing to **shared memory** the parent reads back, and proves parallelism by stamping the **core index** it ran on (like the `smp`/`affinity` cases). - `zig build test` — host unit tests (closure packing, mutex state machine, futex wrapper encodings). - `python3 test/qemu_test.py ` — boots the kernel in QEMU; asserts on serial markers. Thread cases set `smp: true` (real parallelism) and bump `mem` (they boot the process/scheduler stack); each milestone **adds its case to `CASES`** so its gate is runnable. - **Guardrail every milestone:** the concurrency-sensitive existing cases stay green — `smoke`, `sched`, `priority`, `smp`, `affinity`, `process`, `process-kill`, `supervision`, `fault-recovery`, `vfs-client-death`, `ipc`/`ipc-cap`, `display-service`. A threading change that regresses those is rejected. ## Unattended execution (the loop contract) This plan runs to completion **without human input**. Every design choice is already fixed in *Locked decisions*; the checkboxes are the only state. A loop iteration must: 1. **Resume** at the first milestone that still has an unchecked `- [ ]`. (All earlier milestones are done — do not revisit them.) 2. **Work on a branch.** On the first iteration, branch off the current `main` into a new branch (e.g. `threading-phase2` — Phase 1's `threading` is already merged); never commit to `main` directly. Push that **branch** to `origin` after each milestone (step 5) so progress is backed up remotely; **do not push `main`** — merging Phase 2 into `main` stays a human step. 3. **Implement** every unchecked item in that milestone, including adding its `-Dtest-case` to `CASES` in [test/qemu_test.py](../test/qemu_test.py) (with `smp: true` / a `mem` bump where noted) so the gate is runnable. 4. **Run the gate**: `python3 test/qemu_test.py `, then the full **guardrail set**, then `zig build` (clean) and `zig build test` (green). 5. **Decide, do not ask:** - **Green** = the milestone's case prints its stated marker(s) and reports `PASS`, the whole guardrail set passes, `zig build` is clean, and host tests are green. → tick this milestone's boxes **and** its `**Gate:**`-referenced case, `git commit` (`threads(M): `, no `Co-Authored-By` trailer per [coding-standards.md](coding-standards.md)), then **`git push` the working branch to `origin`** (use `-u` on the first push to set upstream). Continue to the next milestone in the same iteration if budget remains; otherwise let the loop re-fire. - **Red** = anything above fails. Diagnose from the captured serial log (`zig-out/qemu-test/-failed-serial.log`) and fix in place, then re-run — up to **3 fix attempts** for that gate. A concurrency case that fails then passes on a bare re-run is **flaky, not green**: re-run it **twice more** and treat green only if it passes all; otherwise fix the race (a real threading bug), don't paper over it. 6. **A genuinely ambiguous fork is not a stop.** Pick the option most consistent with [threading.md](threading.md)'s *Locked decisions*, note the choice in the commit message, and continue. Do not pause for confirmation on in-scope, reversible work — this plan is that authorization. **The only stop conditions:** - **Done** — every milestone box **in this plan** is checked (M1 through M11), `zig build` clean, the whole `thread-*` suite + guardrail green. Phase 1 (M1–M6) is *already* checked, so do **not** read that as Done: the loop's real work is the first plan section that still has unchecked boxes — Phase 2 (M7–M11). Only stop when M7–M11 are all checked too. Update threading.md's status line, push the final branch state to `origin`, and stop. The branch is on `origin` for review; **merging Phase 2 into `main` is the user's step**, not the loop's. - **Blocked** — a gate is still red after 3 fix attempts, or a step needs something outside the repo (a toolchain change, new hardware, a decision no locked decision covers). Append `> **BLOCKED (M):** ` under that milestone, commit **and push** the WIP on the branch, and stop. Do not thrash further and do not silently skip the milestone. Nothing else warrants stopping — not "should I proceed?", not "is this right?". The checkboxes + git history are the resumable record; the next iteration picks up from the first unchecked box. --- ## M1 — Address-space refcount (kernel foundation, no API, no behaviour change) ✅ The one invariant change threads require, landed and proven **before** anything shares an address space. Today address space is 1:1 with a task and teardown destroys it on any user task's exit; make destruction happen on the **last** exit. - [x] A refcount keyed by the address-space root, held in `scheduler.zig` (`address_space_refs`): `retainAddressSpace` takes a reference in `spawnUserLocked` (on the success path, after the slot + stack are secured), all under the big kernel lock. - [x] Both task-teardown paths ([scheduler.zig](../system/kernel/scheduler.zig): `exitUserLocked` and `destroyTaskLocked`) call `releaseAspace`, which decrements and only `destroyAddressSpace`s at **zero**; an unretained space (hand-built test spaces) is destroyed directly, preserving prior behaviour. - [x] `-Dtest-case=address-space-refcount`: spawn and reap several ring-3 processes in sequence and assert (via test-observable `liveAddressSpaceCount`/`addressSpaceDestroyCount`) that the live-space count returns to **baseline** and destructions advance by exactly that many — each space destroyed exactly once, no leak, no double-free. (Refcount observables, not raw frame counts, since kernel stacks are still leaked on exit.) **Gate (met):** `python3 test/qemu_test.py address-space-refcount` passes (`address-space-refcount: spaces released to baseline ok` → `DANOS-TEST-RESULT: PASS`), and the full guardrail set passes unchanged — 13/13 (`smoke`, `sched`, `priority`, `smp`, `affinity`, `process`, `process-kill`, `supervision`, `fault-recovery`, `vfs-client-death`, `ipc`, `ipc-cap`, `display-service`); default `zig build` clean, `zig build test` green. The reframing is invisible until an address space is actually shared. ## M2 — `thread_spawn` + `thread_exit`: a thread runs in the shared address space ✅ Spawn only — no join yet. Prove a second task executes in the **caller's** address space and exits cleanly. - [x] [abi.zig](../system/abi.zig): `thread_spawn = 37`, `thread_exit = 38`. Handlers in process.zig; `thread_spawn` calls `scheduler.spawnThread` (shares the caller's address space, `retainAddressSpace`); `thread_exit` ends the task like a process `exit(0)` (`terminateCurrent` → `releaseAspace`). The closure pointer is delivered in the new thread's **rdi** via a new `jump_to_user_arg` asm path (`t.user_arg`, 0 for a process) — no naked runtime asm. - [x] `library/runtime/thread.zig` (barrel-exported as `runtime.Thread`): `spawn` maps a stack (`mmap`), heap-allocates the `{args}` closure, and calls `thread_spawn(&Closure.entry, stack_top, closure)`; `Closure.entry` (a plain C-ABI Zig fn, closure in rdi) runs the function and calls `thread_exit`. Stack top is 16-aligned-minus-8 for the C entry. - [x] A `threaded` flag on the user-binary recipe (`addThreadedUserBinary` → `single_threaded = false`); `thread-test` is the first opt-in binary. - [x] `-Dtest-case=thread-spawn`: `thread-test` spawns a worker that writes a sentinel to a **shared** global and release-stores `done`; the main thread acquire-polls `done` and asserts the shared global holds the sentinel — proof the worker ran in the same address space. **Gate (met):** `python3 test/qemu_test.py thread-spawn` passes (`thread-test: child ran in shared address space ok` → `DANOS-TEST-RESULT: PASS`); guardrail set 16/16 green (incl. `args`/`init`/`process`, which exercise the new `jump_to_user_arg` process path with arg 0) plus `address-space-refcount`; `zig build` clean, `zig build test` green. > **Note (deferred to M3+):** the mmap arena is per-*task* (`heap_next`), so two threads > in one address space that both `mmap` would collide. Fine for M2 (only the parent maps, for the > child's stack); make the arena per-address-space and the runtime heap thread-safe alongside the > `Mutex` work (M5). ## M3 — `join` + `detach` + real parallelism ✅ - [x] `join` over the existing exit-notification path ([process-lifecycle.md](process-lifecycle.md)): `thread_spawn` gained a 4th arg, an `exit_endpoint` handle (resolved + refcounted like `spawnProcessSupervised`, via `spawnThreadSupervised`); `join` blocks in `ipc_reply_wait` on that endpoint until the child-exit notice for its `tid`, then `munmap`s the stack. `detach` relinquishes the join right (its stack is reclaimed at process exit — kernel-reaper reclaim for detached threads is deferred; see note). - [x] `runtime.Thread.join` / `detach`, plus `Thread.currentCore()` (a new `current_core` = 39 syscall) for the parallelism proof. `getCurrentId` deferred to M6 (TLS), where a lighter self-id fits. The closure now rides the **thread's own stack** (not the heap) — private per thread, so spawn/join touch no shared heap. - [x] `-Dtest-case=thread-join` (`smp: 4`): `thread-test` join mode spawns N=4 workers that each do K=100k `@atomicRmw`-increments on a shared counter and stamp the core they ran on; the main thread joins all N and asserts `counter == N*K` **and** `@popCount(cores_seen) > 1` (genuine cross-core parallelism), then a detached worker proves `detach` runs without a join. **Gate (met):** `python3 test/qemu_test.py thread-join` passes (`thread-test: join ok` → `DANOS-TEST-RESULT: PASS`), robust across 4 runs; guardrail 17/17 green (incl. `smp`, `affinity`, `process-kill`, and `args`/`init`/`process` on the exit-endpoint spawn path) plus `address-space-refcount`/`thread-spawn`; `zig build` clean, `zig build test` green. > **Note (deferred):** a detached thread's stack is freed only at process exit (not by the > reaper on thread exit) — kernel user-stack tracking + reclaim is a later refinement. And > the runtime heap is still not thread-safe: threads that both allocate concurrently would > race (the thread *machinery* avoids the heap, but worker code sharing an allocator does > not). Both fold into the M5 `Mutex`/allocator work. ## M4 — Futex: the one blocking primitive ✅ - [x] [abi.zig](../system/abi.zig): `futex_wait = 40`, `futex_wake = 41`. A waiter is a `.blocked` task tagged with `Task.futex_addr` (no queue linkage); `futex_wait(addr, expected, timeout_ns)` reads the user word under the big lock, parks iff `*addr == expected`, and returns on wake or timeout; `futex_wake(addr, count)` scans the task table and readies up to `count` matching waiters (same address space). No spinning — a parked waiter leaves its core free to `hlt`. A timed wait also sets `wake_at`, so the timer's `wakeExpired` wakes it; `futex_addr` staying non-zero (only `futex_wake` clears it) is how the waiter tells timeout from a real wake. - [x] `runtime.Thread.Futex` (`wait` / `timedWait` / `wake`) over the syscall wrappers. - [x] `-Dtest-case=thread-futex` (`smp: 4`): a waiter thread prints `waiting` and `futex_wait`s on a word; the main thread publishes it, prints `waking`, and `futex_wake`s; the waiter prints `woke`. Then a `timedWait` on an unwoken word reports `error.Timeout`. **Gate (met):** `python3 test/qemu_test.py thread-futex` passes, robust across 3 runs — the case's **ordered** regex asserts `waiting → waking → woke → PASS` on the serial stream (the handoff proof), and `thread-futex: timeout ok` confirms the timeout. Guardrail 18/18 green (incl. `sleep`/`event`/`ipc` blocking paths) + `address-space-refcount`, `thread-spawn`, `thread-join`; `zig build` clean, `zig build test` green. > **Note:** the kernel test checks only the freshest verdict marker via `bufferHas` (the > in-memory log ring buffer evicts older lines); ordering is asserted against the full > serial stream by the qemu regex instead. ## M5 — `Mutex` + `Condition` + `Semaphore` ✅ - [x] `runtime.Thread.Mutex` (three-state futex mutex: CAS fast path, `futex_wait`/`wake` slow path), `Condition` (`wait`/`timedWait`/`signal`/`broadcast`, a futex sequence counter), `Semaphore` (permits over `Mutex`+`Condition`) — the same state machines `std.Thread` uses, ported onto our `Futex`. - [x] `-Dtest-case=thread-mutex` (`smp: 4`): a bounded producer/consumer — 2 producers + 2 consumers over one `Mutex` and two `Condition`s move N=2000 unique items through an 8-slot ring; the consumed checksum and tally match exactly (no lost/duplicated item, no overrun) under real cross-core contention. The small ring forces producers to block on full and consumers on empty, exercising `Condition.wait`. **Gate (met):** `python3 test/qemu_test.py thread-mutex` passes (`thread-mutex: ok` → `DANOS-TEST-RESULT: PASS`), robust across 3 runs; guardrail 17/17 green (incl. `sleep`/`event`/`ipc`) + all M1–M4 thread cases; `zig build` clean, `zig build test` green. > **Deferred (with rationale):** > - **`join` → futex completion word** — the exit-endpoint join (M3) is correct and > tested. A futex-completion join needs the *kernel* to clear+wake a word after the > thread is fully off its stack (a CLONE_CHILD_CLEARTID-style mechanism); doing it in > the thread's own trampoline would let `join` `munmap` the stack while the thread still > runs on it (use-after-free). Left on the exit-endpoint path; the kernel clear-on-exit > is a later, separate refinement. > - **Host unit tests for the state machines** — `Mutex`/`Condition` bottom out in the > `futex_*` syscalls, unavailable on the host without a mockable `Futex` seam. The QEMU > `thread-mutex` gate exercises them under real concurrency instead; a host-side mock is > future work. ## M6 — `getCurrentId`, docs, and CI wiring ✅ - [x] `getCurrentId` via a small `thread_self = 42` syscall (`runtime.Thread.getCurrentId` returns the kernel task id). **Per-thread `threadlocal` TLS is deferred** — no consumer needs it, and it would require context-switching the thread pointer per task (real kernel + per-switch cost) for an unused feature; threaded binaries have run fine without it through M2–M5. threading.md's TLS reasoning already scoped it as deferred-unless-needed. When a consumer appears, the shape is: `thread_spawn` allocates a per-thread TLS block, sets the thread pointer, and the context switch saves/ restores it. - [x] `RwLock` / `WaitGroup` deferred (no consumer yet); they slot onto the same `Futex`/`Mutex`/`Condition` when wanted. - [x] All `thread-*` cases wired into [test/qemu_test.py](../test/qemu_test.py) (`thread-spawn`/`-join`/`-futex`/`-mutex`/`-id`); threading.md + docs/README.md status updated to **built**; the worked example is threading.md's win-condition. - [x] `-Dtest-case=thread-id` (`smp: 4`): two workers read `getCurrentId`; the main thread confirms all three ids are non-zero and distinct — each thread has its own kernel identity. (Renamed from `thread-tls`, which implied `threadlocal`.) **Gate (met):** `python3 test/qemu_test.py thread-id` passes; the whole `thread-*` suite (`thread-spawn`/`-join`/`-futex`/`-mutex`/`-id`) plus the full guardrail set pass; default `zig build` clean, `zig build test` green. --- ## Status **Phase 1 (M1–M6): built.** danos has `runtime.Thread` — `spawn`/`join`/`detach`, cross-core parallelism, futex, and `Mutex`/`Condition`/`Semaphore`, all over a private thread ABI behind the runtime. **Phase 2 (M7–M11): built.** Thread-safe allocation (M7), a task reaper that reclaims dead tasks' kernel stacks (M8), endpoint-free `thread_join` (M9), the per-thread thread pointer (M10), and `RwLock`/`WaitGroup` + host-testable sync (M11). Two things stay deferred by design (no consumer): the Zig `threadlocal` *compiler* layer (M10) and detached-thread user-stack reclaim (M9) — both noted in place. --- ## Phase 2 — hardening (M7–M11) The organising principle, so Phase 2 reinforces danos's goals rather than eroding them: - **Everything a thread owns is reclaimed on process death.** Thread stacks, TLS blocks, and futex words live in the process's **address space**, and the kernel's per-process state is keyed by the address-space root — so the M1 refcount + `destroyAddressSpace` already free all of it when the last thread exits. A crashed or killed threaded process leaves **nothing** behind. Phase 2 closes the one thing that is *not* address-space-owned — the per-task **kernel** stack (kernel heap) — with a reaper (M8). This is the [resilience](resilience.md) restart guarantee, extended to threads. - **Kernel owns mechanism; the runtime owns policy.** The kernel maps pages, saves/ restores the thread pointer, and reaps dead tasks; the runtime decides allocation, TLS layout, and lock algorithms. Every new kernel entry stays a private syscall behind the runtime ([syscall.md](syscall.md)) — the ABI stays renumberable. - **The process is still the isolation and restart boundary.** Threads share fate within one process; Phase 2 never adds a way for one process to reach into another (the cross-process futex stays explicitly out of scope, below). ### M7 — Thread-safe allocation (the correctness gap) ✅ Today the mmap arena cursor is per-*task* and the runtime heap is unlocked, so two threads in one process that both allocate corrupt each other. The thread *machinery* avoids this (closure on the stack, stacks mmap'd only by the spawner), but real multi-threaded code would hit it. Closed it: - [x] **Kernel — per-address-space mmap arena.** Grew M1's `address_space_refs` entry into the per-address-space object holding the `mmap`/`mmio` arena cursors (moved off `Task`); `scheduler.addressSpaceMmapNextPtr`/`addressSpaceDeviceMapNextPtr` expose them. `systemMmap` reserves a disjoint range under a *brief* lock, then maps **per page** under a short-held lock — not the whole grant — because the big lock is held with interrupts disabled, so pinning it across a multi-MiB memset+map froze other cores (it timed the `affinity` scenario out mid-bring-up). Freed at refcount zero, so the cursors vanish with the process. - [x] **Runtime — thread-safe heap.** The allocator's two free-list mutators (`rawAlloc`/`rawFree`) take a `Thread.Mutex`, gated on `!@import("builtin").single_threaded` so single-threaded binaries compile it out and pay nothing. Uncontended acquisition is a single CAS (no syscall). - [x] `-Dtest-case=thread-alloc` (`smp: 4`): 4 threads each do 500 `alloc`/fill/verify/ `free` cycles of varied sizes; each block is filled with a per-thread pattern and verified before free, so any overlap between concurrent allocations is caught. **Gate (met):** `thread-alloc` passes (3× non-flaky); full guardrail 23/23 green, `zig build`/`zig build test` clean. > **Also fixed here:** the `affinity` guardrail's fixed-count busy-loop (`while (spins < > 3e9)`) had codegen-dependent wall-time — adding a function to `tests.zig` flipped how > the optimiser compiled it, swinging affinity from ~4 s to ~63 s and timing it out. > Reworked it (and the settle loop) to wait on the wall clock instead, so its duration is > independent of unrelated code changes. ### M8 — The task reaper (cleanup + resilience) ✅ A dead task's **kernel** stack was leaked ("no reaper yet") — every process *and* thread death lost one, so a crash loop bled kernel memory. The reaper fixes it and serves the [resilience](resilience.md) restart goal directly: - [x] A dying task cannot free the kernel stack it runs on, so `exit()`/`exitUserLocked` record it in a **per-core `reap_after_switch` slot** and switch away; the task that resumes on that core frees the stack in `switchTo`'s tail (it's on its own stack, the big lock is still held so the slot can't have been reused). A **tick-time drain** (`reapKillPendingLocked`) is the safety net for the case where the next task is *fresh* (enters via the trampoline, bypassing `switchTo`'s tail). A task killed while *not* running is freed immediately in `destroyTaskLocked`. A `live_stack_bytes` counter is the observable. *(Detached-thread user-stack reclaim moves to M9, which adds the joinable/detached flag.)* - [x] `-Dtest-case=task-reap` (`smp: 4`): spawn and kill 12 processes; poll the test-observable `scheduler.liveStackBytes()` until it returns to **baseline** (a correct reaper gets there in a few ms; a genuine leak times out) — every kernel stack reclaimed, no leak. Threads exit through the same `exitUserLocked`, so covered. **Gate (met):** `task-reap` passes (5× isolated + 2× in the full batch); `fault-recovery`, `supervision`, `process-kill`, `address-space-refcount`, `smp`, `affinity` all still green (24/24 full guardrail); `zig build`/`zig build test` clean. > **Bug found + fixed here (touches every context switch):** the post-`switchContext` reap > first read the `pc` **parameter**, but a task that migrated cores carries a *stale* `pc` > in its saved `switchTo` frame — so it read the wrong core's slot and freed a live stack > (a #GP under SMP). Fixed to re-fetch `thisCpu()` after the switch (the switch only swaps > stacks on the current core). ### M9 — Futex-completion join (retire the per-thread endpoint) With the reaper (M8) able to act *after* a thread is fully off its stack, migrate `join` to the std shape and drop M3's per-thread exit endpoint: - [x] A **`thread_join(tid)` syscall** (not a user futex word): it blocks the caller until the task with id `tid` exits, and the exit paths call `wakeJoinersLocked`. `join` only reclaims the joined thread's **user** stack, which the thread vacates the moment it enters the kernel to exit — so waking at *exit* time (not reap time) is safe, and no reaper/address-space juggling or user-memory write is needed. This is equally std-shaped (like `pthread_join`) and much simpler/safer than the planned reaper-written completion word. `thread_spawn` no longer takes an exit endpoint (the runtime passes `no_cap`); the per-thread IPC endpoint is gone. - [x] `thread-join` passes on the new path, and its join mode now runs **40 spawn+join cycles** — under the old per-thread-endpoint scheme those leaked handles would exhaust the 16-slot handle table; here they all succeed, proving join is endpoint-free. **Gate (met):** `thread-join` passes (3× isolated) on the `thread_join` path; full guardrail 26/26 (incl. `process-kill`, `supervision`, `fault-recovery`, `task-reap`); `zig build`/`zig build test` clean. > **Reaper hardened here (fixes an M8 flake).** M8's single per-core reap slot could be > *overwritten* by a second death on that core before the first drained (a fresh-task/SMP > timing window) — an intermittent one-stack leak (`task-reap` flaked ~20%). Replaced it > with a per-core reap **list** plus a `.reaping` task state so a pending slot can't be > reused before its stack is freed. `task-reap` now 11/11 isolated + 2× in the batch. > **Deferred:** detached-thread **user-stack** reclaim (still freed at process exit, as in > M3). Doing it in the reaper needs the saved address space + stack range and a > translate/unmap in a not-currently-loaded address space — real complexity for a bounded leak. > A follow-up when a consumer needs it. ### M10 — Per-thread TLS: the thread-pointer mechanism ✅ Give each thread its own thread pointer and private TLS storage — the foundation self-hosting Zig ([zig-self-hosting.md](zig-self-hosting.md)) will build `threadlocal` on. - [x] **Kernel** stores `thread_pointer` on `Task` and restores it on every context switch **only when it changes** (the same conditional-load discipline as CR3; `architecture.setThreadPointer` → `wrmsr IA32_FS_BASE` on x86_64). A `set_thread_pointer(addr)` = 44 syscall sets the caller's `thread_pointer` and loads it now. The kernel never touches FS, so there is no swapgs complication. - [x] **Runtime** lays a small per-thread TLS block at the top of each thread's stack (self-pointer at `%fs:0` + scratch slots) and the thread trampoline calls `set_thread_pointer` before any user code — so every spawned thread has a private, switch-stable thread pointer. Reclaimed with the stack. - [x] `-Dtest-case=thread-tls` (`smp: 4`): two threads each write a unique marker to their own `%fs:8` slot and — after both have written — read it back; a shared (non-per-thread) FS base would clobber one and cause cross-talk. Both read their own marker → pass. **Gate (met):** `thread-tls` passes (3×); full guardrail 25/25 (the switch-time thread-pointer restore touches every context switch); `zig build`/`zig build test` clean. > **Deferred: the Zig `threadlocal` *compiler* layer.** Real `threadlocal` variables need > the ELF **variant-II TLS** surface — `.tdata`/`.tbss` sections + a `PT_TLS` program header > in `user.ld`, a runtime that copies the template with exact negative-offset layout, and > the `.large`-code-model TLS section names — a high-uncertainty lift for a feature with > **no consumer today** (threading.md scopes it "only if a consumer needs it"). What lands > here is the load-bearing piece — the per-thread thread pointer, context-switched — so adding the > compiler layer later is purely runtime+linker work on top, no kernel change. `getCurrentId` > stays the `thread_self` syscall (M6) rather than an fs self-slot (which would need the > main thread's TLS set up in `_start` too). **Gate:** `thread-tls` passes; full `thread-*` suite + guardrail green. ### M11 — `RwLock`, `WaitGroup`, and host-testable sync ✅ - [x] `runtime.Thread.RwLock` (reader-preferring: `>0` readers / `-1` writer / `0` free, with `lock`/`tryLock`/`unlock` + `lockShared`/`tryLockShared`/`unlockShared`) and `WaitGroup` (`start`/`finish`/`wait`), both on the existing `Mutex`/`Condition`. - [x] A compile-time `Futex` seam gated on `builtin.os.tag == .freestanding`: the futex syscalls on danos, a spin+yield mock off-target (Zig 0.16 has no `std.Thread.Futex`; `wake` is a no-op since the state machines re-check). `thread.zig` is wired into `zig build test`, so `Mutex`/`RwLock`/`WaitGroup` run as **host unit tests** with real `std.Thread` threads (`test` blocks only compile under test). - [x] `-Dtest-case=thread-rwlock` (`smp: 4`): 2 writers set both halves of a value under the exclusive lock while 3 readers check the halves match under the shared lock — zero half-write observations across ~150k reads. Host tests cover the Mutex, RwLock, and WaitGroup state machines. **Gate (met):** `zig build test` covers the sync primitives (host threads); `thread-rwlock` passes (3×); full Done gate **26/26** (whole `thread-*` suite + guardrail); `zig build` clean. --- ## Deferred (explicitly not in this plan) - **Cross-process shared-memory futex** — the `(address_space, virtual_address)` key can become a physical-address key so two processes share a futex through a [shared-memory](display-v2.md) region. Not needed for intra-process threads. - **Per-thread priorities / affinity distinct from the process** — threads inherit the process priority ([scheduling.md](scheduling.md)); revisit only if it earns its keep. - **Per-thread signal delivery** — signals stay process-scoped ([process-lifecycle.md](process-lifecycle.md)). - **A `pthread`/POSIX surface** — the API is `std.Thread`-shaped Zig, nothing more. - **A real `std.Thread` backend** — arrives with self-hosting ([zig-self-hosting.md](zig-self-hosting.md)); it sits on these same primitives, so it swaps the impl under `runtime.Thread`, not the call sites.