re-org docs

This commit is contained in:
Daniel Samson
2026-07-23 00:25:34 +01:00
parent 52d6e372fd
commit 757c6f14c3
52 changed files with 156 additions and 156 deletions
+471
View File
@@ -0,0 +1,471 @@
# Threading — build plan (`runtime.Thread` over a private thread ABI)
The ordered, checkpointable build-out for [threading.md](threading.md). Each milestone
lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run, like
[display-v2-plan.md](../device-driver-development/display-v2-plan.md). Read threading.md first for the *why*.
## Locked decisions (do not relitigate)
- **`runtime.Thread` mirrors `std.Thread`'s API; the implementation is danos-native.**
Not literal `std.Thread` — that would break the [private ABI](syscall.md).
- **Threads are a narrow, per-binary opt-in.** Default concurrency stays process + IPC
([resilience.md](resilience.md)); only a service that asks is built
`single_threaded = false`.
- **Blocking is futex-backed, never spin-backed** — waiters park in the kernel so an
idle core still halts ([halting.md](halting.md)).
- **New syscalls are private**: extend [abi.zig](../../system/abi.zig) `SystemCall` after
`shared_memory_physical = 36` (`thread_spawn = 37`, `thread_exit = 38`, `current_core = 39`,
`futex_wait = 40`, `futex_wake = 41`) + a `library/runtime` wrapper; user code never names a number.
- **Restart granularity stays the process** — a faulting thread kills its process; the
supervisor restarts the process, which respawns its threads.
## Conventions
Follow [coding-standards.md](../coding-standards.md): spell out non-acronym abbreviations,
kebab-case file names, no `Co-Authored-By` trailers. New user binaries go through
`addUserBinary` (with the new `threaded` flag where a binary spawns threads) and get
packed into the initial-ramdisk; new syscalls extend [abi.zig](../../system/abi.zig)
`SystemCall` + a `library/runtime` wrapper; test services live beside the code they
exercise and register a `ServiceId` if they must be looked up.
## How to verify along the way
**Every gate is serial-checkable — no screenshots** (this plan runs unattended). A
thread proves it ran by writing to **shared memory** the parent reads back, and proves
parallelism by stamping the **core index** it ran on (like the `smp`/`affinity` cases).
- `zig build test` — host unit tests (closure packing, mutex state machine, futex
wrapper encodings).
- `python3 test/qemu_test.py <case>` — boots the kernel in QEMU; asserts on serial
markers. Thread cases set `smp: true` (real parallelism) and bump `mem` (they boot
the process/scheduler stack); each milestone **adds its case to `CASES`** so its gate
is runnable.
- **Guardrail every milestone:** the concurrency-sensitive existing cases stay green —
`smoke`, `sched`, `priority`, `smp`, `affinity`, `process`, `process-kill`,
`supervision`, `fault-recovery`, `vfs-client-death`, `ipc`/`ipc-cap`,
`display-service`. A threading change that regresses those is rejected.
## Unattended execution (the loop contract)
This plan runs to completion **without human input**. Every design choice is already
fixed in *Locked decisions*; the checkboxes are the only state. A loop iteration must:
1. **Resume** at the first milestone that still has an unchecked `- [ ]`. (All earlier
milestones are done — do not revisit them.)
2. **Work on a branch.** On the first iteration, branch off the current `main` into a new
branch (e.g. `threading-phase2` — Phase 1's `threading` is already merged); never
commit to `main` directly. Push that **branch** to `origin` after each milestone (step
5) so progress is backed up remotely; **do not push `main`** — merging Phase 2 into
`main` stays a human step.
3. **Implement** every unchecked item in that milestone, including adding its
`-Dtest-case` to `CASES` in [test/qemu_test.py](../../test/qemu_test.py) (with
`smp: true` / a `mem` bump where noted) so the gate is runnable.
4. **Run the gate**: `python3 test/qemu_test.py <case>`, then the full **guardrail
set**, then `zig build` (clean) and `zig build test` (green).
5. **Decide, do not ask:**
- **Green** = the milestone's case prints its stated marker(s) and reports `PASS`,
the whole guardrail set passes, `zig build` is clean, and host tests are green.
→ tick this milestone's boxes **and** its `**Gate:**`-referenced case, `git commit`
(`threads(M<n>): <summary>`, no `Co-Authored-By` trailer per
[coding-standards.md](../coding-standards.md)), then **`git push` the working branch to
`origin`** (use `-u` on the first push to set upstream). Continue to the next
milestone in the same iteration if budget remains; otherwise let the loop re-fire.
- **Red** = anything above fails. Diagnose from the captured serial log
(`zig-out/qemu-test/<case>-failed-serial.log`) and fix in place, then re-run — up to
**3 fix attempts** for that gate. A concurrency case that fails then passes on a
bare re-run is **flaky, not green**: re-run it **twice more** and treat green only
if it passes all; otherwise fix the race (a real threading bug), don't paper over
it.
6. **A genuinely ambiguous fork is not a stop.** Pick the option most consistent with
[threading.md](threading.md)'s *Locked decisions*, note the choice in the commit
message, and continue. Do not pause for confirmation on in-scope, reversible work —
this plan is that authorization.
**The only stop conditions:**
- **Done** — every milestone box **in this plan** is checked (M1 through M11), `zig build`
clean, the whole `thread-*` suite + guardrail green. Phase 1 (M1–M6) is *already*
checked, so do **not** read that as Done: the loop's real work is the first plan section
that still has unchecked boxes — Phase 2 (M7–M11). Only stop when M7–M11 are all checked
too. Update threading.md's status line, push the final branch state to `origin`, and
stop. The branch is on `origin` for review; **merging Phase 2 into `main` is the user's
step**, not the loop's.
- **Blocked** — a gate is still red after 3 fix attempts, or a step needs something
outside the repo (a toolchain change, new hardware, a decision no locked decision
covers). Append `> **BLOCKED (M<n>):** <what failed, what was tried, the serial
marker missing>` under that milestone, commit **and push** the WIP on the branch, and
stop. Do not thrash further and do not silently skip the milestone.
Nothing else warrants stopping — not "should I proceed?", not "is this right?". The
checkboxes + git history are the resumable record; the next iteration picks up from the
first unchecked box.
---
## M1 — Address-space refcount (kernel foundation, no API, no behaviour change) ✅
The one invariant change threads require, landed and proven **before** anything shares
an address space. Today address space is 1:1 with a task and teardown destroys it on any user
task's exit; make destruction happen on the **last** exit.
- [x] A refcount keyed by the address-space root, held in `scheduler.zig`
(`address_space_refs`): `retainAddressSpace` takes a reference in `spawnUserLocked` (on the
success path, after the slot + stack are secured), all under the big kernel lock.
- [x] Both task-teardown paths ([scheduler.zig](../../system/kernel/scheduler.zig):
`exitUserLocked` and `destroyTaskLocked`) call `releaseAddressSpace`, which decrements
and only `destroyAddressSpace`s at **zero**; an unretained space (hand-built test
spaces) is destroyed directly, preserving prior behaviour.
- [x] `-Dtest-case=address-space-refcount`: spawn and reap several ring-3 processes in sequence
and assert (via test-observable `liveAddressSpaceCount`/`addressSpaceDestroyCount`) that the
live-space count returns to **baseline** and destructions advance by exactly that
many — each space destroyed exactly once, no leak, no double-free. (Refcount
observables, not raw frame counts, since kernel stacks are still leaked on exit.)
**Gate (met):** `python3 test/qemu_test.py address-space-refcount` passes
(`address-space-refcount: spaces released to baseline ok` → `DANOS-TEST-RESULT: PASS`), and the
full guardrail set passes unchanged — 13/13 (`smoke`, `sched`, `priority`, `smp`,
`affinity`, `process`, `process-kill`, `supervision`, `fault-recovery`,
`vfs-client-death`, `ipc`, `ipc-cap`, `display-service`); default `zig build` clean,
`zig build test` green. The reframing is invisible until an address space is actually shared.
## M2 — `thread_spawn` + `thread_exit`: a thread runs in the shared address space ✅
Spawn only — no join yet. Prove a second task executes in the **caller's** address
space and exits cleanly.
- [x] [abi.zig](../../system/abi.zig): `thread_spawn = 37`, `thread_exit = 38`. Handlers in
process.zig; `thread_spawn` calls `scheduler.spawnThread` (today, after M3, the
handler goes `spawnThreadSupervised` → `scheduler.spawnUserLocked`; shares the caller's
address space, `retainAddressSpace`); `thread_exit` ends the task like a process `exit(0)`
(`terminateCurrent` → `releaseAddressSpace`). The closure pointer is delivered in the new
thread's **rdi** via a new `jump_to_user_arg` asm path (`t.user_arg`, 0 for a
process) — no naked runtime asm.
- [x] `library/runtime/thread.zig` (barrel-exported as `runtime.Thread`): `spawn` maps a
stack (`mmap`), heap-allocates the `{args}` closure, and calls
`thread_spawn(&Closure.entry, stack_top, closure)`; `Closure.entry` (a plain C-ABI
Zig fn, closure in rdi) runs the function and calls `thread_exit`. Stack top is
16-aligned-minus-8 for the C entry.
- [x] A `threaded` flag on the user-binary recipe (`addThreadedUserBinary` →
`single_threaded = false`); `thread-test` is the first opt-in binary.
- [x] `-Dtest-case=thread-spawn`: `thread-test` spawns a worker that writes a sentinel to
a **shared** global and release-stores `done`; the main thread acquire-polls `done`
and asserts the shared global holds the sentinel — proof the worker ran in the same
address space.
**Gate (met):** `python3 test/qemu_test.py thread-spawn` passes
(`thread-test: child ran in shared address space ok` → `DANOS-TEST-RESULT: PASS`); guardrail set
16/16 green (incl. `args`/`init`/`process`, which exercise the new `jump_to_user_arg`
process path with arg 0) plus `address-space-refcount`; `zig build` clean, `zig build test`
green.
> **Note (deferred to M3+):** the mmap arena is per-*task* (`heap_next`), so two threads
> in one address space that both `mmap` would collide. Fine for M2 (only the parent maps, for the
> child's stack); make the arena per-address-space and the runtime heap thread-safe alongside the
> `Mutex` work (M5).
## M3 — `join` + `detach` + real parallelism ✅
- [x] `join` over the existing exit-notification path
([process-lifecycle.md](process-lifecycle.md)): `thread_spawn` gained a 4th arg, an
`exit_endpoint` handle (resolved + refcounted like `spawnProcessSupervised`, via
`spawnThreadSupervised`); `join` blocks in `ipc_reply_wait` on that endpoint until
the child-exit notice for its `tid`, then `munmap`s the stack. `detach` relinquishes
the join right (its stack is reclaimed at process exit — kernel-reaper reclaim for
detached threads is deferred; see note).
- [x] `runtime.Thread.join` / `detach`, plus `Thread.currentCore()` (a new `current_core`
= 39 syscall) for the parallelism proof. `getCurrentId` deferred to M6 (TLS), where
a lighter self-id fits. The closure now rides the **thread's own stack** (not the
heap) — private per thread, so spawn/join touch no shared heap.
- [x] `-Dtest-case=thread-join` (`smp: 4`): `thread-test` join mode spawns N=4 workers
that each do K=100k `@atomicRmw`-increments on a shared counter and stamp the core
they ran on; the main thread joins all N and asserts `counter == N*K` **and**
`@popCount(cores_seen) > 1` (genuine cross-core parallelism), then a detached worker
proves `detach` runs without a join.
**Gate (met):** `python3 test/qemu_test.py thread-join` passes (`thread-test: join ok` →
`DANOS-TEST-RESULT: PASS`), robust across 4 runs; guardrail 17/17 green (incl. `smp`,
`affinity`, `process-kill`, and `args`/`init`/`process` on the exit-endpoint spawn path)
plus `address-space-refcount`/`thread-spawn`; `zig build` clean, `zig build test` green.
> **Note (deferred):** a detached thread's stack is freed only at process exit (not by the
> reaper on thread exit) — kernel user-stack tracking + reclaim is a later refinement. And
> the runtime heap is still not thread-safe: threads that both allocate concurrently would
> race (the thread *machinery* avoids the heap, but worker code sharing an allocator does
> not). Both fold into the M5 `Mutex`/allocator work.
## M4 — Futex: the one blocking primitive ✅
- [x] [abi.zig](../../system/abi.zig): `futex_wait = 40`, `futex_wake = 41`. A waiter is a
`.blocked` task tagged with `Task.futex_addr` (no queue linkage);
`futex_wait(addr, expected, timeout_ns)` reads the user word under the big lock,
parks iff `*addr == expected`, and returns on wake or timeout; `futex_wake(addr,
count)` scans the task table and readies up to `count` matching waiters (same
address space). No spinning — a parked waiter leaves its core free to `hlt`. A
timed wait also sets `wake_at`, so the timer's `wakeExpired` wakes it; `futex_addr`
staying non-zero (only `futex_wake` clears it) is how the waiter tells timeout from
a real wake.
- [x] `runtime.Thread.Futex` (`wait` / `timedWait` / `wake`) over the syscall wrappers.
- [x] `-Dtest-case=thread-futex` (`smp: 4`): a waiter thread prints `waiting` and
`futex_wait`s on a word; the main thread publishes it, prints `waking`, and
`futex_wake`s; the waiter prints `woke`. Then a `timedWait` on an unwoken word
reports `error.Timeout`.
**Gate (met):** `python3 test/qemu_test.py thread-futex` passes, robust across 3 runs —
the case's **ordered** regex asserts `waiting → waking → woke → PASS` on the serial
stream (the handoff proof), and `thread-futex: timeout ok` confirms the timeout.
Guardrail 18/18 green (incl. `sleep`/`event`/`ipc` blocking paths) + `address-space-refcount`,
`thread-spawn`, `thread-join`; `zig build` clean, `zig build test` green.
> **Note:** the kernel test checks only the freshest verdict marker via `bufferHas` (the
> in-memory log ring buffer evicts older lines); ordering is asserted against the full
> serial stream by the qemu regex instead.
## M5 — `Mutex` + `Condition` + `Semaphore` ✅
- [x] `runtime.Thread.Mutex` (three-state futex mutex: CAS fast path, `futex_wait`/`wake`
slow path), `Condition` (`wait`/`timedWait`/`signal`/`broadcast`, a futex sequence
counter), `Semaphore` (permits over `Mutex`+`Condition`) — the same state machines
`std.Thread` uses, ported onto our `Futex`.
- [x] `-Dtest-case=thread-mutex` (`smp: 4`): a bounded producer/consumer — 2 producers +
2 consumers over one `Mutex` and two `Condition`s move N=2000 unique items through
an 8-slot ring; the consumed checksum and tally match exactly (no lost/duplicated
item, no overrun) under real cross-core contention. The small ring forces producers
to block on full and consumers on empty, exercising `Condition.wait`.
**Gate (met):** `python3 test/qemu_test.py thread-mutex` passes (`thread-mutex: ok` →
`DANOS-TEST-RESULT: PASS`), robust across 3 runs; guardrail 17/17 green (incl.
`sleep`/`event`/`ipc`) + all M1–M4 thread cases; `zig build` clean, `zig build test`
green.
> **Deferred (with rationale):**
> - **`join` → futex completion word** — the exit-endpoint join (M3) is correct and
> tested. A futex-completion join needs the *kernel* to clear+wake a word after the
> thread is fully off its stack (a CLONE_CHILD_CLEARTID-style mechanism); doing it in
> the thread's own trampoline would let `join` `munmap` the stack while the thread still
> runs on it (use-after-free). Left on the exit-endpoint path; the kernel clear-on-exit
> is a later, separate refinement.
> - **Host unit tests for the state machines** — `Mutex`/`Condition` bottom out in the
> `futex_*` syscalls, unavailable on the host without a mockable `Futex` seam. The QEMU
> `thread-mutex` gate exercises them under real concurrency instead; a host-side mock is
> future work.
## M6 — `getCurrentId`, docs, and CI wiring ✅
- [x] `getCurrentId` via a small `thread_self = 42` syscall (`runtime.Thread.getCurrentId`
returns the kernel task id). **Per-thread `threadlocal` TLS is deferred** — no
consumer needs it, and it would require context-switching the thread pointer per task
(real kernel + per-switch cost) for an unused feature; threaded binaries have run fine
without it through M2–M5. threading.md's TLS reasoning already scoped it as
deferred-unless-needed. When a consumer appears, the shape is: `thread_spawn`
allocates a per-thread TLS block, sets the thread pointer, and the context switch saves/
restores it.
- [x] `RwLock` / `WaitGroup` deferred (no consumer yet); they slot onto the same
`Futex`/`Mutex`/`Condition` when wanted.
- [x] All `thread-*` cases wired into [test/qemu_test.py](../../test/qemu_test.py)
(`thread-spawn`/`-join`/`-futex`/`-mutex`/`-id`); threading.md + docs/README.md
status updated to **built**; the worked example is threading.md's win-condition.
- [x] `-Dtest-case=thread-id` (`smp: 4`): two workers read `getCurrentId`; the main
thread confirms all three ids are non-zero and distinct — each thread has its own
kernel identity. (Renamed from `thread-tls`, which implied `threadlocal`.)
**Gate (met):** `python3 test/qemu_test.py thread-id` passes; the whole `thread-*` suite
(`thread-spawn`/`-join`/`-futex`/`-mutex`/`-id`) plus the full guardrail set pass; default
`zig build` clean, `zig build test` green.
---
## Status
**Phase 1 (M1–M6): built.** danos has `runtime.Thread` — `spawn`/`join`/`detach`,
cross-core parallelism, futex, and `Mutex`/`Condition`/`Semaphore`, all over a private
thread ABI behind the runtime.
**Phase 2 (M7–M11): built.** Thread-safe allocation (M7), a task reaper that reclaims dead
tasks' kernel stacks (M8), endpoint-free `thread_join` (M9), the per-thread thread pointer (M10),
and `RwLock`/`WaitGroup` + host-testable sync (M11). Two things stay deferred by design
(no consumer): the Zig `threadlocal` *compiler* layer (M10) and detached-thread user-stack
reclaim (M9) — both noted in place.
---
## Phase 2 — hardening (M7–M11)
The organising principle, so Phase 2 reinforces danos's goals rather than eroding them:
- **Everything a thread owns is reclaimed on process death.** Thread stacks, TLS blocks,
and futex words live in the process's **address space**, and the kernel's per-process
state is keyed by the address-space root — so the M1 refcount + `destroyAddressSpace` already
free all of it when the last thread exits. A crashed or killed threaded process leaves
**nothing** behind. Phase 2 closes the one thing that is *not* address-space-owned — the
per-task **kernel** stack (kernel heap) — with a reaper (M8). This is the
[resilience](resilience.md) restart guarantee, extended to threads.
- **Kernel owns mechanism; the runtime owns policy.** The kernel maps pages, saves/
restores the thread pointer, and reaps dead tasks; the runtime decides allocation, TLS layout,
and lock algorithms. Every new kernel entry stays a private syscall behind the runtime
([syscall.md](syscall.md)) — the ABI stays renumberable.
- **The process is still the isolation and restart boundary.** Threads share fate within
one process; Phase 2 never adds a way for one process to reach into another (the
cross-process futex stays explicitly out of scope, below).
### M7 — Thread-safe allocation (the correctness gap) ✅
Today the mmap arena cursor is per-*task* and the runtime heap is unlocked, so two
threads in one process that both allocate corrupt each other. The thread *machinery*
avoids this (closure on the stack, stacks mmap'd only by the spawner), but real
multi-threaded code would hit it. Closed it:
- [x] **Kernel — per-address-space mmap arena.** Grew M1's `address_space_refs` entry into the
per-address-space object holding the `mmap`/`mmio` arena cursors (moved off `Task`);
`scheduler.addressSpaceMmapNextPtr`/`addressSpaceDeviceMapNextPtr` expose them. `systemMmap`
reserves a disjoint range under a *brief* lock, then maps **per page** under a
short-held lock — not the whole grant — because the big lock is held with interrupts
disabled, so pinning it across a multi-MiB memset+map froze other cores (it timed
the `affinity` scenario out mid-bring-up). Freed at refcount zero, so the cursors
vanish with the process.
- [x] **Runtime — thread-safe heap.** The allocator's two free-list mutators
(`rawAlloc`/`rawFree`) take a `Thread.Mutex`, gated on
`!@import("builtin").single_threaded` so single-threaded binaries compile it out and
pay nothing. Uncontended acquisition is a single CAS (no syscall).
- [x] `-Dtest-case=thread-alloc` (`smp: 4`): 4 threads each do 500 `alloc`/fill/verify/
`free` cycles of varied sizes; each block is filled with a per-thread pattern and
verified before free, so any overlap between concurrent allocations is caught.
**Gate (met):** `thread-alloc` passes (3× non-flaky); full guardrail 23/23 green,
`zig build`/`zig build test` clean.
> **Also fixed here:** the `affinity` guardrail's fixed-count busy-loop (`while (spins <
> 3e9)`) had codegen-dependent wall-time — adding a function to `tests.zig` flipped how
> the optimiser compiled it, swinging affinity from ~4 s to ~63 s and timing it out.
> Reworked it (and the settle loop) to wait on the wall clock instead, so its duration is
> independent of unrelated code changes.
### M8 — The task reaper (cleanup + resilience) ✅
A dead task's **kernel** stack was leaked ("no reaper yet") — every process *and* thread
death lost one, so a crash loop bled kernel memory. The reaper fixes it and serves the
[resilience](resilience.md) restart goal directly:
- [x] A dying task cannot free the kernel stack it runs on, so `exit()`/`exitUserLocked`
record it in a **per-core `reap_after_switch` slot** and switch away; the task that
resumes on that core frees the stack in `switchTo`'s tail (it's on its own stack, the
big lock is still held so the slot can't have been reused). A **tick-time drain**
(`reapKillPendingLocked`) is the safety net for the case where the next task is
*fresh* (enters via the trampoline, bypassing `switchTo`'s tail). A task killed while
*not* running is freed immediately in `destroyTaskLocked`. A `live_stack_bytes`
counter is the observable. *(Detached-thread user-stack reclaim moves to M9, which
adds the joinable/detached flag.)*
- [x] `-Dtest-case=task-reap` (`smp: 4`): spawn and kill 12 processes; poll the
test-observable `scheduler.liveStackBytes()` until it returns to **baseline** (a
correct reaper gets there in a few ms; a genuine leak times out) — every kernel
stack reclaimed, no leak. Threads exit through the same `exitUserLocked`, so covered.
**Gate (met):** `task-reap` passes (5× isolated + 2× in the full batch); `fault-recovery`,
`supervision`, `process-kill`, `address-space-refcount`, `smp`, `affinity` all still green (24/24
full guardrail); `zig build`/`zig build test` clean.
> **Bug found + fixed here (touches every context switch):** the post-`switchContext` reap
> first read the `pc` **parameter**, but a task that migrated cores carries a *stale* `pc`
> in its saved `switchTo` frame — so it read the wrong core's slot and freed a live stack
> (a #GP under SMP). Fixed to re-fetch `thisCpu()` after the switch (the switch only swaps
> stacks on the current core).
### M9 — Futex-completion join (retire the per-thread endpoint)
With the reaper (M8) able to act *after* a thread is fully off its stack, migrate `join`
to the std shape and drop M3's per-thread exit endpoint:
- [x] A **`thread_join(tid)` syscall** (not a user futex word): it blocks the caller until
the task with id `tid` exits, and the exit paths call `wakeJoinersLocked`. `join`
only reclaims the joined thread's **user** stack, which the thread vacates the moment
it enters the kernel to exit — so waking at *exit* time (not reap time) is safe, and
no reaper/address-space juggling or user-memory write is needed. This is equally
std-shaped (like `pthread_join`) and much simpler/safer than the planned
reaper-written completion word. The runtime no longer passes `thread_spawn` an exit
endpoint (it passes `no_cap`; the kernel's 4th `exit_endpoint` arg remains and is
still honored); the runtime's per-thread IPC endpoint is gone.
- [x] `thread-join` passes on the new path, and its join mode now runs **40 spawn+join
cycles** — under the old per-thread-endpoint scheme those leaked handles would
exhaust the 16-slot handle table; here they all succeed, proving join is endpoint-free.
**Gate (met):** `thread-join` passes (3× isolated) on the `thread_join` path; full
guardrail 26/26 (incl. `process-kill`, `supervision`, `fault-recovery`, `task-reap`);
`zig build`/`zig build test` clean.
> **Reaper hardened here (fixes an M8 flake).** M8's single per-core reap slot could be
> *overwritten* by a second death on that core before the first drained (a fresh-task/SMP
> timing window) — an intermittent one-stack leak (`task-reap` flaked ~20%). Replaced it
> with a per-core reap **list** plus a `.reaping` task state so a pending slot can't be
> reused before its stack is freed. `task-reap` now 11/11 isolated + 2× in the batch.
> **Deferred:** detached-thread **user-stack** reclaim (still freed at process exit, as in
> M3). Doing it in the reaper needs the saved address space + stack range and a
> translate/unmap in a not-currently-loaded address space — real complexity for a bounded leak.
> A follow-up when a consumer needs it.
### M10 — Per-thread TLS: the thread-pointer mechanism ✅
Give each thread its own thread pointer and private TLS storage — the foundation
self-hosting Zig ([zig-self-hosting.md](../zig-self-hosting.md)) will build `threadlocal` on.
- [x] **Kernel** stores `thread_pointer` on `Task` and restores it on every context switch
**only when it changes** (the same conditional-load discipline as CR3;
`architecture.setThreadPointer` → `wrmsr IA32_FS_BASE` on x86_64). A
`set_thread_pointer(addr)` = 44 syscall sets the caller's `thread_pointer` and loads it
now. The kernel never touches FS, so there is no swapgs complication.
- [x] **Runtime** lays a small per-thread TLS block at the top of each thread's stack
(self-pointer at `%fs:0` + scratch slots) and the thread trampoline calls
`set_thread_pointer` before any user code — so every spawned thread has a private,
switch-stable thread pointer. Reclaimed with the stack.
- [x] `-Dtest-case=thread-tls` (`smp: 4`): two threads each write a unique marker to their
own `%fs:8` slot and — after both have written — read it back; a shared (non-per-thread)
FS base would clobber one and cause cross-talk. Both read their own marker → pass.
**Gate (met):** `thread-tls` passes (3×); full guardrail 25/25 (the switch-time thread-pointer
restore touches every context switch); `zig build`/`zig build test` clean.
> **Deferred: the Zig `threadlocal` *compiler* layer.** Real `threadlocal` variables need
> the ELF **variant-II TLS** surface — `.tdata`/`.tbss` sections + a `PT_TLS` program header
> in `user.ld`, a runtime that copies the template with exact negative-offset layout, and
> the `.large`-code-model TLS section names — a high-uncertainty lift for a feature with
> **no consumer today** (threading.md scopes it "only if a consumer needs it"). What lands
> here is the load-bearing piece — the per-thread thread pointer, context-switched — so adding the
> compiler layer later is purely runtime+linker work on top, no kernel change. `getCurrentId`
> stays the `thread_self` syscall (M6) rather than an fs self-slot (which would need the
> main thread's TLS set up in `_start` too).
**Gate:** `thread-tls` passes; full `thread-*` suite + guardrail green.
### M11 — `RwLock`, `WaitGroup`, and host-testable sync ✅
- [x] `runtime.Thread.RwLock` (reader-preferring: `>0` readers / `-1` writer / `0` free,
with `lock`/`tryLock`/`unlock` + `lockShared`/`tryLockShared`/`unlockShared`) and
`WaitGroup` (`start`/`finish`/`wait`), both on the existing `Mutex`/`Condition`.
- [x] A compile-time `Futex` seam gated on `builtin.os.tag == .freestanding`: the futex
syscalls on danos, a spin+yield mock off-target (Zig 0.16 has no `std.Thread.Futex`;
`wake` is a no-op since the state machines re-check). `thread.zig` is wired into
`zig build test`, so `Mutex`/`RwLock`/`WaitGroup` run as **host unit tests** with real
`std.Thread` threads (`test` blocks only compile under test).
- [x] `-Dtest-case=thread-rwlock` (`smp: 4`): 2 writers set both halves of a value under
the exclusive lock while 3 readers check the halves match under the shared lock —
zero half-write observations across ~150k reads. Host tests cover the Mutex,
RwLock, and WaitGroup state machines.
**Gate (met):** `zig build test` covers the sync primitives (host threads); `thread-rwlock`
passes (3×); full Done gate **26/26** (whole `thread-*` suite + guardrail); `zig build`
clean.
---
## Deferred (explicitly not in this plan)
- **Cross-process shared-memory futex** — the `(address_space, virtual_address)` key can become a
physical-address key so two processes share a futex through a [shared-memory](../device-driver-development/display-v2.md)
region. Not needed for intra-process threads.
- **Per-thread priorities / affinity distinct from the process** — threads inherit the
process priority ([scheduling.md](scheduling.md)); revisit only if it earns its keep.
- **Per-thread signal delivery** — signals stay process-scoped
([process-lifecycle.md](process-lifecycle.md)).
- **A `pthread`/POSIX surface** — the API is `std.Thread`-shaped Zig, nothing more.
- **A real `std.Thread` backend** — arrives with self-hosting
([zig-self-hosting.md](../zig-self-hosting.md)); it sits on these same primitives, so it
swaps the impl under `runtime.Thread`, not the call sites.