473 lines
30 KiB
Markdown
473 lines
30 KiB
Markdown
# Threading — build plan (`runtime.Thread` over a private thread ABI)
|
||
|
||
The ordered, checkpointable build-out for [threading.md](threading.md). Each milestone
|
||
lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run, like
|
||
[display-v2-plan.md](../device-driver-development/display-v2-plan.md). Read threading.md first for the *why*.
|
||
|
||
## Locked decisions (do not relitigate)
|
||
|
||
- **`runtime.Thread` mirrors `std.Thread`'s API; the implementation is danos-native.**
|
||
Not literal `std.Thread` — that would break the [private ABI](syscall.md).
|
||
- **Threads are a narrow, per-binary opt-in.** Default concurrency stays process + IPC
|
||
([resilience.md](resilience.md)); only a service that asks is built
|
||
`single_threaded = false`.
|
||
- **Blocking is futex-backed, never spin-backed** — waiters park in the kernel so an
|
||
idle core still halts ([halting.md](halting.md)).
|
||
- **New syscalls are private**: extend [abi.zig](../../system/abi.zig) `SystemCall` after
|
||
`shared_memory_physical = 36` (`thread_spawn = 37`, `thread_exit = 38`, `current_core = 39`,
|
||
`futex_wait = 40`, `futex_wake = 41`) + a `library/runtime` wrapper; user code never names a number.
|
||
- **Restart granularity stays the process** — a faulting thread kills its process; the
|
||
supervisor restarts the process, which respawns its threads.
|
||
|
||
## Conventions
|
||
|
||
Follow [coding-standards.md](../coding-standards.md): spell out non-acronym abbreviations,
|
||
kebab-case file names, no `Co-Authored-By` trailers. New user binaries are
|
||
packages whose build.zig calls `build_support.userBinary` (with `.threaded =
|
||
true` where a binary spawns threads) and get packed into the initial-ramdisk;
|
||
new syscalls extend [abi.zig](../../system/abi.zig) `SystemCall` + a
|
||
`library/kernel` wrapper; test services live beside the code they exercise and
|
||
bind a `/protocol/test/...` name if they must be reachable.
|
||
|
||
## How to verify along the way
|
||
|
||
**Every gate is serial-checkable — no screenshots** (this plan runs unattended). A
|
||
thread proves it ran by writing to **shared memory** the parent reads back, and proves
|
||
parallelism by stamping the **core index** it ran on (like the `smp`/`affinity` cases).
|
||
|
||
- `zig build test` — host unit tests (closure packing, mutex state machine, futex
|
||
wrapper encodings).
|
||
- `python3 test/qemu_test.py <case>` — boots the kernel in QEMU; asserts on serial
|
||
markers. Thread cases set `smp: true` (real parallelism) and bump `mem` (they boot
|
||
the process/scheduler stack); each milestone **adds its case to `CASES`** so its gate
|
||
is runnable.
|
||
- **Guardrail every milestone:** the concurrency-sensitive existing cases stay green —
|
||
`smoke`, `sched`, `priority`, `smp`, `affinity`, `process`, `process-kill`,
|
||
`supervision`, `fault-recovery`, `vfs-client-death`, `ipc`/`ipc-cap`,
|
||
`display-service`. A threading change that regresses those is rejected.
|
||
|
||
## Unattended execution (the loop contract)
|
||
|
||
This plan runs to completion **without human input**. Every design choice is already
|
||
fixed in *Locked decisions*; the checkboxes are the only state. A loop iteration must:
|
||
|
||
1. **Resume** at the first milestone that still has an unchecked `- [ ]`. (All earlier
|
||
milestones are done — do not revisit them.)
|
||
2. **Work on a branch.** On the first iteration, branch off the current `main` into a new
|
||
branch (e.g. `threading-phase2` — Phase 1's `threading` is already merged); never
|
||
commit to `main` directly. Push that **branch** to `origin` after each milestone (step
|
||
5) so progress is backed up remotely; **do not push `main`** — merging Phase 2 into
|
||
`main` stays a human step.
|
||
3. **Implement** every unchecked item in that milestone, including adding its
|
||
`-Dtest-case` to `CASES` in [test/qemu_test.py](../../test/qemu_test.py) (with
|
||
`smp: true` / a `mem` bump where noted) so the gate is runnable.
|
||
4. **Run the gate**: `python3 test/qemu_test.py <case>`, then the full **guardrail
|
||
set**, then `zig build` (clean) and `zig build test` (green).
|
||
5. **Decide, do not ask:**
|
||
- **Green** = the milestone's case prints its stated marker(s) and reports `PASS`,
|
||
the whole guardrail set passes, `zig build` is clean, and host tests are green.
|
||
→ tick this milestone's boxes **and** its `**Gate:**`-referenced case, `git commit`
|
||
(`threads(M<n>): <summary>`, no `Co-Authored-By` trailer per
|
||
[coding-standards.md](../coding-standards.md)), then **`git push` the working branch to
|
||
`origin`** (use `-u` on the first push to set upstream). Continue to the next
|
||
milestone in the same iteration if budget remains; otherwise let the loop re-fire.
|
||
- **Red** = anything above fails. Diagnose from the captured serial log
|
||
(`zig-out/qemu-test/<case>-failed-serial.log`) and fix in place, then re-run — up to
|
||
**3 fix attempts** for that gate. A concurrency case that fails then passes on a
|
||
bare re-run is **flaky, not green**: re-run it **twice more** and treat green only
|
||
if it passes all; otherwise fix the race (a real threading bug), don't paper over
|
||
it.
|
||
6. **A genuinely ambiguous fork is not a stop.** Pick the option most consistent with
|
||
[threading.md](threading.md)'s *Locked decisions*, note the choice in the commit
|
||
message, and continue. Do not pause for confirmation on in-scope, reversible work —
|
||
this plan is that authorization.
|
||
|
||
**The only stop conditions:**
|
||
|
||
- **Done** — every milestone box **in this plan** is checked (M1 through M11), `zig build`
|
||
clean, the whole `thread-*` suite + guardrail green. Phase 1 (M1–M6) is *already*
|
||
checked, so do **not** read that as Done: the loop's real work is the first plan section
|
||
that still has unchecked boxes — Phase 2 (M7–M11). Only stop when M7–M11 are all checked
|
||
too. Update threading.md's status line, push the final branch state to `origin`, and
|
||
stop. The branch is on `origin` for review; **merging Phase 2 into `main` is the user's
|
||
step**, not the loop's.
|
||
- **Blocked** — a gate is still red after 3 fix attempts, or a step needs something
|
||
outside the repo (a toolchain change, new hardware, a decision no locked decision
|
||
covers). Append `> **BLOCKED (M<n>):** <what failed, what was tried, the serial
|
||
marker missing>` under that milestone, commit **and push** the WIP on the branch, and
|
||
stop. Do not thrash further and do not silently skip the milestone.
|
||
|
||
Nothing else warrants stopping — not "should I proceed?", not "is this right?". The
|
||
checkboxes + git history are the resumable record; the next iteration picks up from the
|
||
first unchecked box.
|
||
|
||
---
|
||
|
||
## M1 — Address-space refcount (kernel foundation, no API, no behaviour change) ✅
|
||
|
||
The one invariant change threads require, landed and proven **before** anything shares
|
||
an address space. Today address space is 1:1 with a task and teardown destroys it on any user
|
||
task's exit; make destruction happen on the **last** exit.
|
||
|
||
- [x] A refcount keyed by the address-space root, held in `scheduler.zig`
|
||
(`address_space_refs`): `retainAddressSpace` takes a reference in `spawnUserLocked` (on the
|
||
success path, after the slot + stack are secured), all under the big kernel lock.
|
||
- [x] Both task-teardown paths ([scheduler.zig](../../system/kernel/scheduler.zig):
|
||
`exitUserLocked` and `destroyTaskLocked`) call `releaseAddressSpace`, which decrements
|
||
and only `destroyAddressSpace`s at **zero**; an unretained space (hand-built test
|
||
spaces) is destroyed directly, preserving prior behaviour.
|
||
- [x] `-Dtest-case=address-space-refcount`: spawn and reap several ring-3 processes in sequence
|
||
and assert (via test-observable `liveAddressSpaceCount`/`addressSpaceDestroyCount`) that the
|
||
live-space count returns to **baseline** and destructions advance by exactly that
|
||
many — each space destroyed exactly once, no leak, no double-free. (Refcount
|
||
observables, not raw frame counts, since kernel stacks are still leaked on exit.)
|
||
|
||
**Gate (met):** `python3 test/qemu_test.py address-space-refcount` passes
|
||
(`address-space-refcount: spaces released to baseline ok` → `DANOS-TEST-RESULT: PASS`), and the
|
||
full guardrail set passes unchanged — 13/13 (`smoke`, `sched`, `priority`, `smp`,
|
||
`affinity`, `process`, `process-kill`, `supervision`, `fault-recovery`,
|
||
`vfs-client-death`, `ipc`, `ipc-cap`, `display-service`); default `zig build` clean,
|
||
`zig build test` green. The reframing is invisible until an address space is actually shared.
|
||
|
||
## M2 — `thread_spawn` + `thread_exit`: a thread runs in the shared address space ✅
|
||
|
||
Spawn only — no join yet. Prove a second task executes in the **caller's** address
|
||
space and exits cleanly.
|
||
|
||
- [x] [abi.zig](../../system/abi.zig): `thread_spawn = 37`, `thread_exit = 38`. Handlers in
|
||
process.zig; `thread_spawn` calls `scheduler.spawnThread` (today, after M3, the
|
||
handler goes `spawnThreadSupervised` → `scheduler.spawnUserLocked`; shares the caller's
|
||
address space, `retainAddressSpace`); `thread_exit` ends the task like a process `exit(0)`
|
||
(`terminateCurrent` → `releaseAddressSpace`). The closure pointer is delivered in the new
|
||
thread's **rdi** via a new `jump_to_user_arg` asm path (`t.user_arg`, 0 for a
|
||
process) — no naked runtime asm.
|
||
- [x] `library/runtime/thread.zig` (barrel-exported as `runtime.Thread`): `spawn` maps a
|
||
stack (`mmap`), heap-allocates the `{args}` closure, and calls
|
||
`thread_spawn(&Closure.entry, stack_top, closure)`; `Closure.entry` (a plain C-ABI
|
||
Zig fn, closure in rdi) runs the function and calls `thread_exit`. Stack top is
|
||
16-aligned-minus-8 for the C entry.
|
||
- [x] A `threaded` flag on the user-binary recipe (`addThreadedUserBinary` →
|
||
`single_threaded = false`); `thread-test` is the first opt-in binary.
|
||
- [x] `-Dtest-case=thread-spawn`: `thread-test` spawns a worker that writes a sentinel to
|
||
a **shared** global and release-stores `done`; the main thread acquire-polls `done`
|
||
and asserts the shared global holds the sentinel — proof the worker ran in the same
|
||
address space.
|
||
|
||
**Gate (met):** `python3 test/qemu_test.py thread-spawn` passes
|
||
(`thread-test: child ran in shared address space ok` → `DANOS-TEST-RESULT: PASS`); guardrail set
|
||
16/16 green (incl. `args`/`init`/`process`, which exercise the new `jump_to_user_arg`
|
||
process path with arg 0) plus `address-space-refcount`; `zig build` clean, `zig build test`
|
||
green.
|
||
|
||
> **Note (deferred to M3+):** the mmap arena is per-*task* (`heap_next`), so two threads
|
||
> in one address space that both `mmap` would collide. Fine for M2 (only the parent maps, for the
|
||
> child's stack); make the arena per-address-space and the runtime heap thread-safe alongside the
|
||
> `Mutex` work (M5).
|
||
|
||
## M3 — `join` + `detach` + real parallelism ✅
|
||
|
||
- [x] `join` over the existing exit-notification path
|
||
([process-lifecycle.md](process-lifecycle.md)): `thread_spawn` gained a 4th arg, an
|
||
`exit_endpoint` handle (resolved + refcounted like `spawnProcessSupervised`, via
|
||
`spawnThreadSupervised`); `join` blocks in `ipc_reply_wait` on that endpoint until
|
||
the child-exit notice for its `tid`, then `munmap`s the stack. `detach` relinquishes
|
||
the join right (its stack is reclaimed at process exit — kernel-reaper reclaim for
|
||
detached threads is deferred; see note).
|
||
- [x] `runtime.Thread.join` / `detach`, plus `Thread.currentCore()` (a new `current_core`
|
||
= 39 syscall) for the parallelism proof. `getCurrentId` deferred to M6 (TLS), where
|
||
a lighter self-id fits. The closure now rides the **thread's own stack** (not the
|
||
heap) — private per thread, so spawn/join touch no shared heap.
|
||
- [x] `-Dtest-case=thread-join` (`smp: 4`): `thread-test` join mode spawns N=4 workers
|
||
that each do K=100k `@atomicRmw`-increments on a shared counter and stamp the core
|
||
they ran on; the main thread joins all N and asserts `counter == N*K` **and**
|
||
`@popCount(cores_seen) > 1` (genuine cross-core parallelism), then a detached worker
|
||
proves `detach` runs without a join.
|
||
|
||
**Gate (met):** `python3 test/qemu_test.py thread-join` passes (`thread-test: join ok` →
|
||
`DANOS-TEST-RESULT: PASS`), robust across 4 runs; guardrail 17/17 green (incl. `smp`,
|
||
`affinity`, `process-kill`, and `args`/`init`/`process` on the exit-endpoint spawn path)
|
||
plus `address-space-refcount`/`thread-spawn`; `zig build` clean, `zig build test` green.
|
||
|
||
> **Note (deferred):** a detached thread's stack is freed only at process exit (not by the
|
||
> reaper on thread exit) — kernel user-stack tracking + reclaim is a later refinement. And
|
||
> the runtime heap is still not thread-safe: threads that both allocate concurrently would
|
||
> race (the thread *machinery* avoids the heap, but worker code sharing an allocator does
|
||
> not). Both fold into the M5 `Mutex`/allocator work.
|
||
|
||
## M4 — Futex: the one blocking primitive ✅
|
||
|
||
- [x] [abi.zig](../../system/abi.zig): `futex_wait = 40`, `futex_wake = 41`. A waiter is a
|
||
`.blocked` task tagged with `Task.futex_addr` (no queue linkage);
|
||
`futex_wait(addr, expected, timeout_ns)` reads the user word under the big lock,
|
||
parks iff `*addr == expected`, and returns on wake or timeout; `futex_wake(addr,
|
||
count)` scans the task table and readies up to `count` matching waiters (same
|
||
address space). No spinning — a parked waiter leaves its core free to `hlt`. A
|
||
timed wait also sets `wake_at`, so the timer's `wakeExpired` wakes it; `futex_addr`
|
||
staying non-zero (only `futex_wake` clears it) is how the waiter tells timeout from
|
||
a real wake.
|
||
- [x] `runtime.Thread.Futex` (`wait` / `timedWait` / `wake`) over the syscall wrappers.
|
||
- [x] `-Dtest-case=thread-futex` (`smp: 4`): a waiter thread prints `waiting` and
|
||
`futex_wait`s on a word; the main thread publishes it, prints `waking`, and
|
||
`futex_wake`s; the waiter prints `woke`. Then a `timedWait` on an unwoken word
|
||
reports `error.Timeout`.
|
||
|
||
**Gate (met):** `python3 test/qemu_test.py thread-futex` passes, robust across 3 runs —
|
||
the case's **ordered** regex asserts `waiting → waking → woke → PASS` on the serial
|
||
stream (the handoff proof), and `thread-futex: timeout ok` confirms the timeout.
|
||
Guardrail 18/18 green (incl. `sleep`/`event`/`ipc` blocking paths) + `address-space-refcount`,
|
||
`thread-spawn`, `thread-join`; `zig build` clean, `zig build test` green.
|
||
|
||
> **Note:** the kernel test checks only the freshest verdict marker via `bufferHas` (the
|
||
> in-memory log ring buffer evicts older lines); ordering is asserted against the full
|
||
> serial stream by the qemu regex instead.
|
||
|
||
## M5 — `Mutex` + `Condition` + `Semaphore` ✅
|
||
|
||
- [x] `runtime.Thread.Mutex` (three-state futex mutex: CAS fast path, `futex_wait`/`wake`
|
||
slow path), `Condition` (`wait`/`timedWait`/`signal`/`broadcast`, a futex sequence
|
||
counter), `Semaphore` (permits over `Mutex`+`Condition`) — the same state machines
|
||
`std.Thread` uses, ported onto our `Futex`.
|
||
- [x] `-Dtest-case=thread-mutex` (`smp: 4`): a bounded producer/consumer — 2 producers +
|
||
2 consumers over one `Mutex` and two `Condition`s move N=2000 unique items through
|
||
an 8-slot ring; the consumed checksum and tally match exactly (no lost/duplicated
|
||
item, no overrun) under real cross-core contention. The small ring forces producers
|
||
to block on full and consumers on empty, exercising `Condition.wait`.
|
||
|
||
**Gate (met):** `python3 test/qemu_test.py thread-mutex` passes (`thread-mutex: ok` →
|
||
`DANOS-TEST-RESULT: PASS`), robust across 3 runs; guardrail 17/17 green (incl.
|
||
`sleep`/`event`/`ipc`) + all M1–M4 thread cases; `zig build` clean, `zig build test`
|
||
green.
|
||
|
||
> **Deferred (with rationale):**
|
||
> - **`join` → futex completion word** — the exit-endpoint join (M3) is correct and
|
||
> tested. A futex-completion join needs the *kernel* to clear+wake a word after the
|
||
> thread is fully off its stack (a CLONE_CHILD_CLEARTID-style mechanism); doing it in
|
||
> the thread's own trampoline would let `join` `munmap` the stack while the thread still
|
||
> runs on it (use-after-free). Left on the exit-endpoint path; the kernel clear-on-exit
|
||
> is a later, separate refinement.
|
||
> - **Host unit tests for the state machines** — `Mutex`/`Condition` bottom out in the
|
||
> `futex_*` syscalls, unavailable on the host without a mockable `Futex` seam. The QEMU
|
||
> `thread-mutex` gate exercises them under real concurrency instead; a host-side mock is
|
||
> future work.
|
||
|
||
## M6 — `getCurrentId`, docs, and CI wiring ✅
|
||
|
||
- [x] `getCurrentId` via a small `thread_self = 42` syscall (`runtime.Thread.getCurrentId`
|
||
returns the kernel task id). **Per-thread `threadlocal` TLS is deferred** — no
|
||
consumer needs it, and it would require context-switching the thread pointer per task
|
||
(real kernel + per-switch cost) for an unused feature; threaded binaries have run fine
|
||
without it through M2–M5. threading.md's TLS reasoning already scoped it as
|
||
deferred-unless-needed. When a consumer appears, the shape is: `thread_spawn`
|
||
allocates a per-thread TLS block, sets the thread pointer, and the context switch saves/
|
||
restores it.
|
||
- [x] `RwLock` / `WaitGroup` deferred (no consumer yet); they slot onto the same
|
||
`Futex`/`Mutex`/`Condition` when wanted.
|
||
- [x] All `thread-*` cases wired into [test/qemu_test.py](../../test/qemu_test.py)
|
||
(`thread-spawn`/`-join`/`-futex`/`-mutex`/`-id`); threading.md + docs/README.md
|
||
status updated to **built**; the worked example is threading.md's win-condition.
|
||
- [x] `-Dtest-case=thread-id` (`smp: 4`): two workers read `getCurrentId`; the main
|
||
thread confirms all three ids are non-zero and distinct — each thread has its own
|
||
kernel identity. (Renamed from `thread-tls`, which implied `threadlocal`.)
|
||
|
||
**Gate (met):** `python3 test/qemu_test.py thread-id` passes; the whole `thread-*` suite
|
||
(`thread-spawn`/`-join`/`-futex`/`-mutex`/`-id`) plus the full guardrail set pass; default
|
||
`zig build` clean, `zig build test` green.
|
||
|
||
---
|
||
|
||
## Status
|
||
|
||
**Phase 1 (M1–M6): built.** danos has `runtime.Thread` — `spawn`/`join`/`detach`,
|
||
cross-core parallelism, futex, and `Mutex`/`Condition`/`Semaphore`, all over a private
|
||
thread ABI behind the runtime.
|
||
|
||
**Phase 2 (M7–M11): built.** Thread-safe allocation (M7), a task reaper that reclaims dead
|
||
tasks' kernel stacks (M8), endpoint-free `thread_join` (M9), the per-thread thread pointer (M10),
|
||
and `RwLock`/`WaitGroup` + host-testable sync (M11). Two things stay deferred by design
|
||
(no consumer): the Zig `threadlocal` *compiler* layer (M10) and detached-thread user-stack
|
||
reclaim (M9) — both noted in place.
|
||
|
||
---
|
||
|
||
## Phase 2 — hardening (M7–M11)
|
||
|
||
The organising principle, so Phase 2 reinforces danos's goals rather than eroding them:
|
||
|
||
- **Everything a thread owns is reclaimed on process death.** Thread stacks, TLS blocks,
|
||
and futex words live in the process's **address space**, and the kernel's per-process
|
||
state is keyed by the address-space root — so the M1 refcount + `destroyAddressSpace` already
|
||
free all of it when the last thread exits. A crashed or killed threaded process leaves
|
||
**nothing** behind. Phase 2 closes the one thing that is *not* address-space-owned — the
|
||
per-task **kernel** stack (kernel heap) — with a reaper (M8). This is the
|
||
[resilience](resilience.md) restart guarantee, extended to threads.
|
||
- **Kernel owns mechanism; the runtime owns policy.** The kernel maps pages, saves/
|
||
restores the thread pointer, and reaps dead tasks; the runtime decides allocation, TLS layout,
|
||
and lock algorithms. Every new kernel entry stays a private syscall behind the runtime
|
||
([syscall.md](syscall.md)) — the ABI stays renumberable.
|
||
- **The process is still the isolation and restart boundary.** Threads share fate within
|
||
one process; Phase 2 never adds a way for one process to reach into another (the
|
||
cross-process futex stays explicitly out of scope, below).
|
||
|
||
### M7 — Thread-safe allocation (the correctness gap) ✅
|
||
|
||
Today the mmap arena cursor is per-*task* and the runtime heap is unlocked, so two
|
||
threads in one process that both allocate corrupt each other. The thread *machinery*
|
||
avoids this (closure on the stack, stacks mmap'd only by the spawner), but real
|
||
multi-threaded code would hit it. Closed it:
|
||
|
||
- [x] **Kernel — per-address-space mmap arena.** Grew M1's `address_space_refs` entry into the
|
||
per-address-space object holding the `mmap`/`mmio` arena cursors (moved off `Task`);
|
||
`scheduler.addressSpaceMmapNextPtr`/`addressSpaceDeviceMapNextPtr` expose them. `systemMmap`
|
||
reserves a disjoint range under a *brief* lock, then maps **per page** under a
|
||
short-held lock — not the whole grant — because the big lock is held with interrupts
|
||
disabled, so pinning it across a multi-MiB memset+map froze other cores (it timed
|
||
the `affinity` scenario out mid-bring-up). Freed at refcount zero, so the cursors
|
||
vanish with the process.
|
||
- [x] **Runtime — thread-safe heap.** The allocator's two free-list mutators
|
||
(`rawAlloc`/`rawFree`) take a `Thread.Mutex`, gated on
|
||
`!@import("builtin").single_threaded` so single-threaded binaries compile it out and
|
||
pay nothing. Uncontended acquisition is a single CAS (no syscall).
|
||
- [x] `-Dtest-case=thread-alloc` (`smp: 4`): 4 threads each do 500 `alloc`/fill/verify/
|
||
`free` cycles of varied sizes; each block is filled with a per-thread pattern and
|
||
verified before free, so any overlap between concurrent allocations is caught.
|
||
|
||
**Gate (met):** `thread-alloc` passes (3× non-flaky); full guardrail 23/23 green,
|
||
`zig build`/`zig build test` clean.
|
||
|
||
> **Also fixed here:** the `affinity` guardrail's fixed-count busy-loop (`while (spins <
|
||
> 3e9)`) had codegen-dependent wall-time — adding a function to `tests.zig` flipped how
|
||
> the optimiser compiled it, swinging affinity from ~4 s to ~63 s and timing it out.
|
||
> Reworked it (and the settle loop) to wait on the wall clock instead, so its duration is
|
||
> independent of unrelated code changes.
|
||
|
||
### M8 — The task reaper (cleanup + resilience) ✅
|
||
|
||
A dead task's **kernel** stack was leaked ("no reaper yet") — every process *and* thread
|
||
death lost one, so a crash loop bled kernel memory. The reaper fixes it and serves the
|
||
[resilience](resilience.md) restart goal directly:
|
||
|
||
- [x] A dying task cannot free the kernel stack it runs on, so `exit()`/`exitUserLocked`
|
||
record it in a **per-core `reap_after_switch` slot** and switch away; the task that
|
||
resumes on that core frees the stack in `switchTo`'s tail (it's on its own stack, the
|
||
big lock is still held so the slot can't have been reused). A **tick-time drain**
|
||
(`reapKillPendingLocked`) is the safety net for the case where the next task is
|
||
*fresh* (enters via the trampoline, bypassing `switchTo`'s tail). A task killed while
|
||
*not* running is freed immediately in `destroyTaskLocked`. A `live_stack_bytes`
|
||
counter is the observable. *(Detached-thread user-stack reclaim moves to M9, which
|
||
adds the joinable/detached flag.)*
|
||
- [x] `-Dtest-case=task-reap` (`smp: 4`): spawn and kill 12 processes; poll the
|
||
test-observable `scheduler.liveStackBytes()` until it returns to **baseline** (a
|
||
correct reaper gets there in a few ms; a genuine leak times out) — every kernel
|
||
stack reclaimed, no leak. Threads exit through the same `exitUserLocked`, so covered.
|
||
|
||
**Gate (met):** `task-reap` passes (5× isolated + 2× in the full batch); `fault-recovery`,
|
||
`supervision`, `process-kill`, `address-space-refcount`, `smp`, `affinity` all still green (24/24
|
||
full guardrail); `zig build`/`zig build test` clean.
|
||
|
||
> **Bug found + fixed here (touches every context switch):** the post-`switchContext` reap
|
||
> first read the `pc` **parameter**, but a task that migrated cores carries a *stale* `pc`
|
||
> in its saved `switchTo` frame — so it read the wrong core's slot and freed a live stack
|
||
> (a #GP under SMP). Fixed to re-fetch `thisCpu()` after the switch (the switch only swaps
|
||
> stacks on the current core).
|
||
|
||
### M9 — Futex-completion join (retire the per-thread endpoint)
|
||
|
||
With the reaper (M8) able to act *after* a thread is fully off its stack, migrate `join`
|
||
to the std shape and drop M3's per-thread exit endpoint:
|
||
|
||
- [x] A **`thread_join(tid)` syscall** (not a user futex word): it blocks the caller until
|
||
the task with id `tid` exits, and the exit paths call `wakeJoinersLocked`. `join`
|
||
only reclaims the joined thread's **user** stack, which the thread vacates the moment
|
||
it enters the kernel to exit — so waking at *exit* time (not reap time) is safe, and
|
||
no reaper/address-space juggling or user-memory write is needed. This is equally
|
||
std-shaped (like `pthread_join`) and much simpler/safer than the planned
|
||
reaper-written completion word. The runtime no longer passes `thread_spawn` an exit
|
||
endpoint (it passes `no_cap`; the kernel's 4th `exit_endpoint` arg remains and is
|
||
still honored); the runtime's per-thread IPC endpoint is gone.
|
||
- [x] `thread-join` passes on the new path, and its join mode now runs **40 spawn+join
|
||
cycles** — under the old per-thread-endpoint scheme those leaked handles would
|
||
exhaust the 16-slot handle table; here they all succeed, proving join is endpoint-free.
|
||
|
||
**Gate (met):** `thread-join` passes (3× isolated) on the `thread_join` path; full
|
||
guardrail 26/26 (incl. `process-kill`, `supervision`, `fault-recovery`, `task-reap`);
|
||
`zig build`/`zig build test` clean.
|
||
|
||
> **Reaper hardened here (fixes an M8 flake).** M8's single per-core reap slot could be
|
||
> *overwritten* by a second death on that core before the first drained (a fresh-task/SMP
|
||
> timing window) — an intermittent one-stack leak (`task-reap` flaked ~20%). Replaced it
|
||
> with a per-core reap **list** plus a `.reaping` task state so a pending slot can't be
|
||
> reused before its stack is freed. `task-reap` now 11/11 isolated + 2× in the batch.
|
||
|
||
> **Deferred:** detached-thread **user-stack** reclaim (still freed at process exit, as in
|
||
> M3). Doing it in the reaper needs the saved address space + stack range and a
|
||
> translate/unmap in a not-currently-loaded address space — real complexity for a bounded leak.
|
||
> A follow-up when a consumer needs it.
|
||
|
||
### M10 — Per-thread TLS: the thread-pointer mechanism ✅
|
||
|
||
Give each thread its own thread pointer and private TLS storage — the foundation
|
||
self-hosting Zig ([zig-self-hosting.md](../zig-self-hosting.md)) will build `threadlocal` on.
|
||
|
||
- [x] **Kernel** stores `thread_pointer` on `Task` and restores it on every context switch
|
||
**only when it changes** (the same conditional-load discipline as CR3;
|
||
`architecture.setThreadPointer` → `wrmsr IA32_FS_BASE` on x86_64). A
|
||
`set_thread_pointer(addr)` = 44 syscall sets the caller's `thread_pointer` and loads it
|
||
now. The kernel never touches FS, so there is no swapgs complication.
|
||
- [x] **Runtime** lays a small per-thread TLS block at the top of each thread's stack
|
||
(self-pointer at `%fs:0` + scratch slots) and the thread trampoline calls
|
||
`set_thread_pointer` before any user code — so every spawned thread has a private,
|
||
switch-stable thread pointer. Reclaimed with the stack.
|
||
- [x] `-Dtest-case=thread-tls` (`smp: 4`): two threads each write a unique marker to their
|
||
own `%fs:8` slot and — after both have written — read it back; a shared (non-per-thread)
|
||
FS base would clobber one and cause cross-talk. Both read their own marker → pass.
|
||
|
||
**Gate (met):** `thread-tls` passes (3×); full guardrail 25/25 (the switch-time thread-pointer
|
||
restore touches every context switch); `zig build`/`zig build test` clean.
|
||
|
||
> **Deferred: the Zig `threadlocal` *compiler* layer.** Real `threadlocal` variables need
|
||
> the ELF **variant-II TLS** surface — `.tdata`/`.tbss` sections + a `PT_TLS` program header
|
||
> in `user.ld`, a runtime that copies the template with exact negative-offset layout, and
|
||
> the `.large`-code-model TLS section names — a high-uncertainty lift for a feature with
|
||
> **no consumer today** (threading.md scopes it "only if a consumer needs it"). What lands
|
||
> here is the load-bearing piece — the per-thread thread pointer, context-switched — so adding the
|
||
> compiler layer later is purely runtime+linker work on top, no kernel change. `getCurrentId`
|
||
> stays the `thread_self` syscall (M6) rather than an fs self-slot (which would need the
|
||
> main thread's TLS set up in `_start` too).
|
||
|
||
**Gate:** `thread-tls` passes; full `thread-*` suite + guardrail green.
|
||
|
||
### M11 — `RwLock`, `WaitGroup`, and host-testable sync ✅
|
||
|
||
- [x] `runtime.Thread.RwLock` (reader-preferring: `>0` readers / `-1` writer / `0` free,
|
||
with `lock`/`tryLock`/`unlock` + `lockShared`/`tryLockShared`/`unlockShared`) and
|
||
`WaitGroup` (`start`/`finish`/`wait`), both on the existing `Mutex`/`Condition`.
|
||
- [x] A compile-time `Futex` seam gated on `builtin.os.tag == .freestanding`: the futex
|
||
syscalls on danos, a spin+yield mock off-target (Zig 0.16 has no `std.Thread.Futex`;
|
||
`wake` is a no-op since the state machines re-check). `thread.zig` is wired into
|
||
`zig build test`, so `Mutex`/`RwLock`/`WaitGroup` run as **host unit tests** with real
|
||
`std.Thread` threads (`test` blocks only compile under test).
|
||
- [x] `-Dtest-case=thread-rwlock` (`smp: 4`): 2 writers set both halves of a value under
|
||
the exclusive lock while 3 readers check the halves match under the shared lock —
|
||
zero half-write observations across ~150k reads. Host tests cover the Mutex,
|
||
RwLock, and WaitGroup state machines.
|
||
|
||
**Gate (met):** `zig build test` covers the sync primitives (host threads); `thread-rwlock`
|
||
passes (3×); full Done gate **26/26** (whole `thread-*` suite + guardrail); `zig build`
|
||
clean.
|
||
|
||
---
|
||
|
||
## Deferred (explicitly not in this plan)
|
||
|
||
- **Cross-process shared-memory futex** — the `(address_space, virtual_address)` key can become a
|
||
physical-address key so two processes share a futex through a [shared-memory](../device-driver-development/display-v2.md)
|
||
region. Not needed for intra-process threads.
|
||
- **Per-thread priorities / affinity distinct from the process** — threads inherit the
|
||
process priority ([scheduling.md](scheduling.md)); revisit only if it earns its keep.
|
||
- **Per-thread signal delivery** — signals stay process-scoped
|
||
([process-lifecycle.md](process-lifecycle.md)).
|
||
- **A `pthread`/POSIX surface** — the API is `std.Thread`-shaped Zig, nothing more.
|
||
- **A real `std.Thread` backend** — arrives with self-hosting
|
||
([zig-self-hosting.md](../zig-self-hosting.md)); it sits on these same primitives, so it
|
||
swaps the impl under `runtime.Thread`, not the call sites.
|