Files
danos/docs/threading-plan.md
T
daniel 8b7f1d009c threads: plan Phase 2 (M7-M11) — hardening the deferred parts
Add M7-M11 to docs/threading-plan.md, designed so threading reinforces danos's
goals: everything a thread owns lives in the address space (reclaimed on process
death via the M1 refcount), the kernel owns mechanism while the runtime owns
policy, and the process stays the isolation/restart boundary.

M7 thread-safe allocation (per-aspace mmap arena + locked runtime heap);
M8 task reaper (reclaim kernel stacks + detached user stacks — the resilience gap);
M9 futex-completion join (retire the per-thread endpoint, built on M8);
M10 per-thread TLS (threadlocal + fs.base, for self-hosting);
M11 RwLock/WaitGroup + host-testable sync.
2026-07-20 22:10:32 +01:00

24 KiB
Raw Blame History

Threading — build plan (runtime.Thread over a private thread ABI)

The ordered, checkpointable build-out for threading.md. Each milestone lands on its own and ends in a verifiable gate — shaped for a /loop run, like display-v2-plan.md. Read threading.md first for the why.

Locked decisions (do not relitigate)

  • runtime.Thread mirrors std.Thread's API; the implementation is danos-native. Not literal std.Thread — that would break the private ABI.
  • Threads are a narrow, per-binary opt-in. Default concurrency stays process + IPC (resilience.md); only a service that asks is built single_threaded = false.
  • Blocking is futex-backed, never spin-backed — waiters park in the kernel so an idle core still halts (halting.md).
  • New syscalls are private: extend abi.zig SystemCall after shm_physical = 36 (thread_spawn = 37, thread_exit = 38, current_core = 39, futex_wait = 40, futex_wake = 41) + a library/runtime wrapper; user code never names a number.
  • Restart granularity stays the process — a faulting thread kills its process; the supervisor restarts the process, which respawns its threads.

Conventions

Follow coding-standards.md: spell out non-acronym abbreviations, kebab-case file names, no Co-Authored-By trailers. New user binaries go through addUserBinary (with the new threaded flag where a binary spawns threads) and get packed into the initial-ramdisk; new syscalls extend abi.zig SystemCall + a library/runtime wrapper; test services live beside the code they exercise and register a ServiceId if they must be looked up.

How to verify along the way

Every gate is serial-checkable — no screenshots (this plan runs unattended). A thread proves it ran by writing to shared memory the parent reads back, and proves parallelism by stamping the core index it ran on (like the smp/affinity cases).

  • zig build test — host unit tests (closure packing, mutex state machine, futex wrapper encodings).
  • python3 test/qemu_test.py <case> — boots the kernel in QEMU; asserts on serial markers. Thread cases set smp: true (real parallelism) and bump mem (they boot the process/scheduler stack); each milestone adds its case to CASES so its gate is runnable.
  • Guardrail every milestone: the concurrency-sensitive existing cases stay green — smoke, sched, priority, smp, affinity, process, process-kill, supervision, fault-recovery, vfs-client-death, ipc/ipc-cap, display-service. A threading change that regresses those is rejected.

Unattended execution (the loop contract)

This plan runs to completion without human input. Every design choice is already fixed in Locked decisions; the checkboxes are the only state. A loop iteration must:

  1. Resume at the first milestone that still has an unchecked - [ ]. (All earlier milestones are done — do not revisit them.)
  2. Work on a branch. On the first iteration, branch off main (e.g. threading); never commit threading work to main. All work stays local — do not push.
  3. Implement every unchecked item in that milestone, including adding its -Dtest-case to CASES in test/qemu_test.py (with smp: true / a mem bump where noted) so the gate is runnable.
  4. Run the gate: python3 test/qemu_test.py <case>, then the full guardrail set, then zig build (clean) and zig build test (green).
  5. Decide, do not ask:
    • Green = the milestone's case prints its stated marker(s) and reports PASS, the whole guardrail set passes, zig build is clean, and host tests are green. → tick this milestone's boxes and its **Gate:**-referenced case, git commit (threads(M<n>): <summary>, no Co-Authored-By trailer per coding-standards.md), and continue to the next milestone in the same iteration if budget remains; otherwise let the loop re-fire.
    • Red = anything above fails. Diagnose from the captured serial log (zig-out/qemu-test/<case>-failed-serial.log) and fix in place, then re-run — up to 3 fix attempts for that gate. A concurrency case that fails then passes on a bare re-run is flaky, not green: re-run it twice more and treat green only if it passes all; otherwise fix the race (a real threading bug), don't paper over it.
  6. A genuinely ambiguous fork is not a stop. Pick the option most consistent with threading.md's Locked decisions, note the choice in the commit message, and continue. Do not pause for confirmation on in-scope, reversible work — this plan is that authorization.

The only stop conditions:

  • Done — every milestone box is checked (M1–M6), zig build clean, whole thread-* suite + guardrail green. Update threading.md's status line to "built" (that is M6's own task) and stop.
  • Blocked — a gate is still red after 3 fix attempts, or a step needs something outside the repo (a toolchain change, new hardware, a decision no locked decision covers). Append > **BLOCKED (M<n>):** <what failed, what was tried, the serial marker missing> under that milestone, commit the WIP on the branch, and stop. Do not thrash further and do not silently skip the milestone.

Nothing else warrants stopping — not "should I proceed?", not "is this right?". The checkboxes + git history are the resumable record; the next iteration picks up from the first unchecked box.


M1 — Address-space refcount (kernel foundation, no API, no behaviour change) ✅

The one invariant change threads require, landed and proven before anything shares an address space. Today aspace is 1:1 with a task and teardown destroys it on any user task's exit; make destruction happen on the last exit.

  • A refcount keyed by the address-space root, held in scheduler.zig (aspace_refs): retainAspace takes a reference in spawnUserLocked (on the success path, after the slot + stack are secured), all under the big kernel lock.
  • Both task-teardown paths (scheduler.zig: exitUserLocked and destroyTaskLocked) call releaseAspace, which decrements and only destroyAddressSpaces at zero; an unretained space (hand-built test spaces) is destroyed directly, preserving prior behaviour.
  • -Dtest-case=aspace-refcount: spawn and reap several ring-3 processes in sequence and assert (via test-observable liveAspaceCount/aspaceDestroyCount) that the live-space count returns to baseline and destructions advance by exactly that many — each space destroyed exactly once, no leak, no double-free. (Refcount observables, not raw frame counts, since kernel stacks are still leaked on exit.)

Gate (met): python3 test/qemu_test.py aspace-refcount passes (aspace-refcount: spaces released to baseline ok → DANOS-TEST-RESULT: PASS), and the full guardrail set passes unchanged — 13/13 (smoke, sched, priority, smp, affinity, process, process-kill, supervision, fault-recovery, vfs-client-death, ipc, ipc-cap, display-service); default zig build clean, zig build test green. The reframing is invisible until an aspace is actually shared.

M2 — thread_spawn + thread_exit: a thread runs in the shared address space ✅

Spawn only — no join yet. Prove a second task executes in the caller's address space and exits cleanly.

  • abi.zig: thread_spawn = 37, thread_exit = 38. Handlers in process.zig; thread_spawn calls scheduler.spawnThread (shares the caller's aspace, retainAspace); thread_exit ends the task like a process exit(0) (terminateCurrent → releaseAspace). The closure pointer is delivered in the new thread's rdi via a new jump_to_user_arg asm path (t.user_arg, 0 for a process) — no naked runtime asm.
  • library/runtime/thread.zig (barrel-exported as runtime.Thread): spawn maps a stack (mmap), heap-allocates the {args} closure, and calls thread_spawn(&Closure.entry, stack_top, closure); Closure.entry (a plain C-ABI Zig fn, closure in rdi) runs the function and calls thread_exit. Stack top is 16-aligned-minus-8 for the C entry.
  • A threaded flag on the user-binary recipe (addThreadedUserBinary → single_threaded = false); thread-test is the first opt-in binary.
  • -Dtest-case=thread-spawn: thread-test spawns a worker that writes a sentinel to a shared global and release-stores done; the main thread acquire-polls done and asserts the shared global holds the sentinel — proof the worker ran in the same address space.

Gate (met): python3 test/qemu_test.py thread-spawn passes (thread-test: child ran in shared aspace ok → DANOS-TEST-RESULT: PASS); guardrail set 16/16 green (incl. args/init/process, which exercise the new jump_to_user_arg process path with arg 0) plus aspace-refcount; zig build clean, zig build test green.

Note (deferred to M3+): the mmap arena is per-task (heap_next), so two threads in one aspace that both mmap would collide. Fine for M2 (only the parent maps, for the child's stack); make the arena per-aspace and the runtime heap thread-safe alongside the Mutex work (M5).

M3 — join + detach + real parallelism ✅

  • join over the existing exit-notification path (process-lifecycle.md): thread_spawn gained a 4th arg, an exit_endpoint handle (resolved + refcounted like spawnProcessSupervised, via spawnThreadSupervised); join blocks in ipc_reply_wait on that endpoint until the child-exit notice for its tid, then munmaps the stack. detach relinquishes the join right (its stack is reclaimed at process exit — kernel-reaper reclaim for detached threads is deferred; see note).
  • runtime.Thread.join / detach, plus Thread.currentCore() (a new current_core = 39 syscall) for the parallelism proof. getCurrentId deferred to M6 (TLS), where a lighter self-id fits. The closure now rides the thread's own stack (not the heap) — private per thread, so spawn/join touch no shared heap.
  • -Dtest-case=thread-join (smp: 4): thread-test join mode spawns N=4 workers that each do K=100k @atomicRmw-increments on a shared counter and stamp the core they ran on; the main thread joins all N and asserts counter == N*K and @popCount(cores_seen) > 1 (genuine cross-core parallelism), then a detached worker proves detach runs without a join.

Gate (met): python3 test/qemu_test.py thread-join passes (thread-test: join ok → DANOS-TEST-RESULT: PASS), robust across 4 runs; guardrail 17/17 green (incl. smp, affinity, process-kill, and args/init/process on the exit-endpoint spawn path) plus aspace-refcount/thread-spawn; zig build clean, zig build test green.

Note (deferred): a detached thread's stack is freed only at process exit (not by the reaper on thread exit) — kernel user-stack tracking + reclaim is a later refinement. And the runtime heap is still not thread-safe: threads that both allocate concurrently would race (the thread machinery avoids the heap, but worker code sharing an allocator does not). Both fold into the M5 Mutex/allocator work.

M4 — Futex: the one blocking primitive ✅

  • abi.zig: futex_wait = 40, futex_wake = 41. A waiter is a .blocked task tagged with Task.futex_addr (no queue linkage); futex_wait(addr, expected, timeout_ns) reads the user word under the big lock, parks iff *addr == expected, and returns on wake or timeout; futex_wake(addr, count) scans the task table and readies up to count matching waiters (same address space). No spinning — a parked waiter leaves its core free to hlt. A timed wait also sets wake_at, so the timer's wakeExpired wakes it; futex_addr staying non-zero (only futex_wake clears it) is how the waiter tells timeout from a real wake.
  • runtime.Thread.Futex (wait / timedWait / wake) over the syscall wrappers.
  • -Dtest-case=thread-futex (smp: 4): a waiter thread prints waiting and futex_waits on a word; the main thread publishes it, prints waking, and futex_wakes; the waiter prints woke. Then a timedWait on an unwoken word reports error.Timeout.

Gate (met): python3 test/qemu_test.py thread-futex passes, robust across 3 runs — the case's ordered regex asserts waiting → waking → woke → PASS on the serial stream (the handoff proof), and thread-futex: timeout ok confirms the timeout. Guardrail 18/18 green (incl. sleep/event/ipc blocking paths) + aspace-refcount, thread-spawn, thread-join; zig build clean, zig build test green.

Note: the kernel test checks only the freshest verdict marker via bufferHas (the in-memory log ring buffer evicts older lines); ordering is asserted against the full serial stream by the qemu regex instead.

M5 — Mutex + Condition + Semaphore ✅

  • runtime.Thread.Mutex (three-state futex mutex: CAS fast path, futex_wait/wake slow path), Condition (wait/timedWait/signal/broadcast, a futex sequence counter), Semaphore (permits over Mutex+Condition) — the same state machines std.Thread uses, ported onto our Futex.
  • -Dtest-case=thread-mutex (smp: 4): a bounded producer/consumer — 2 producers + 2 consumers over one Mutex and two Conditions move N=2000 unique items through an 8-slot ring; the consumed checksum and tally match exactly (no lost/duplicated item, no overrun) under real cross-core contention. The small ring forces producers to block on full and consumers on empty, exercising Condition.wait.

Gate (met): python3 test/qemu_test.py thread-mutex passes (thread-mutex: ok → DANOS-TEST-RESULT: PASS), robust across 3 runs; guardrail 17/17 green (incl. sleep/event/ipc) + all M1–M4 thread cases; zig build clean, zig build test green.

Deferred (with rationale):

  • join → futex completion word — the exit-endpoint join (M3) is correct and tested. A futex-completion join needs the kernel to clear+wake a word after the thread is fully off its stack (a CLONE_CHILD_CLEARTID-style mechanism); doing it in the thread's own trampoline would let join munmap the stack while the thread still runs on it (use-after-free). Left on the exit-endpoint path; the kernel clear-on-exit is a later, separate refinement.
  • Host unit tests for the state machines — Mutex/Condition bottom out in the futex_* syscalls, unavailable on the host without a mockable Futex seam. The QEMU thread-mutex gate exercises them under real concurrency instead; a host-side mock is future work.

M6 — getCurrentId, docs, and CI wiring ✅

  • getCurrentId via a small thread_self = 42 syscall (runtime.Thread.getCurrentId returns the kernel task id). Per-thread threadlocal TLS is deferred — no consumer needs it, and it would require context-switching fs.base per task (real kernel + per-switch cost) for an unused feature; threaded binaries have run fine without it through M2–M5. threading.md's TLS reasoning already scoped it as deferred-unless-needed. When a consumer appears, the shape is: thread_spawn allocates a per-thread TLS block, sets fs.base, and the context switch saves/ restores it.
  • RwLock / WaitGroup deferred (no consumer yet); they slot onto the same Futex/Mutex/Condition when wanted.
  • All thread-* cases wired into test/qemu_test.py (thread-spawn/-join/-futex/-mutex/-id); threading.md + docs/README.md status updated to built; the worked example is threading.md's win-condition.
  • -Dtest-case=thread-id (smp: 4): two workers read getCurrentId; the main thread confirms all three ids are non-zero and distinct — each thread has its own kernel identity. (Renamed from thread-tls, which implied threadlocal.)

Gate (met): python3 test/qemu_test.py thread-id passes; the whole thread-* suite (thread-spawn/-join/-futex/-mutex/-id) plus the full guardrail set pass; default zig build clean, zig build test green.


Status

Phase 1 (M1–M6): built. danos has runtime.Thread — spawn/join/detach, cross-core parallelism, futex, and Mutex/Condition/Semaphore, all over a private thread ABI behind the runtime.

Phase 2 (M7–M11): planned below — hardening the deferred parts so threads are safe for real workloads and reclaimed like everything else danos owns.


Phase 2 — hardening (M7–M11)

The organising principle, so Phase 2 reinforces danos's goals rather than eroding them:

  • Everything a thread owns is reclaimed on process death. Thread stacks, TLS blocks, and futex words live in the process's address space, and the kernel's per-process state is keyed by the aspace root — so the M1 refcount + destroyAddressSpace already free all of it when the last thread exits. A crashed or killed threaded process leaves nothing behind. Phase 2 closes the one thing that is not aspace-owned — the per-task kernel stack (kernel heap) — with a reaper (M8). This is the resilience restart guarantee, extended to threads.
  • Kernel owns mechanism; the runtime owns policy. The kernel maps pages, saves/ restores fs.base, and reaps dead tasks; the runtime decides allocation, TLS layout, and lock algorithms. Every new kernel entry stays a private syscall behind the runtime (syscall.md) — the ABI stays renumberable.
  • The process is still the isolation and restart boundary. Threads share fate within one process; Phase 2 never adds a way for one process to reach into another (the cross-process futex stays explicitly out of scope, below).

M7 — Thread-safe allocation (the correctness gap)

Today the mmap arena cursor is per-task and the runtime heap is unlocked, so two threads in one process that both allocate corrupt each other. The thread machinery avoids this (closure on the stack, stacks mmap'd only by the spawner), but real multi-threaded code would hit it. Close it:

  • Kernel — per-address-space mmap arena. Grow M1's aspace_refs entry into a small address-space object holding the mmap/mmio arena cursors (moved off Task); systemMmap/mmio_map bump the aspace's cursor under the big lock, so sibling threads get disjoint, serialized grants. Freed at refcount zero, so the cursors vanish with the process.
  • Runtime — thread-safe heap. Guard the allocator with a Thread.Mutex, gated on !@import("builtin").single_threaded so single-threaded binaries compile it out and pay nothing. (The heap grows via mmap, now safe per above.)
  • -Dtest-case=thread-alloc (smp: 4): N threads each do many alloc/free of varied sizes, write a per-thread pattern, verify it, and free; assert every block round-trips intact and all memory returns — no corruption under concurrent allocation. A direct check confirms two threads' concurrent mmaps are disjoint.

Gate: thread-alloc passes; guardrail + all thread-* cases green.

M8 — The task reaper (cleanup + resilience)

A dead task's kernel stack is currently leaked ("no reaper yet") — every process and thread death loses one, so a crash loop bleeds kernel memory. A reaper fixes it and serves the resilience restart goal directly:

  • A dying task cannot free the kernel stack it runs on, so it hands itself to a reap list and switches away; the kernel stack (and, for a detached thread, its user stack) is reclaimed from another context — a low-priority reaper step drained on the scheduler tick and when a core goes idle. Extends the existing reap_task_hook/destroyTaskLocked path rather than inventing a parallel one.
  • -Dtest-case=task-reap: spawn and exit many threads and processes; assert the kernel-heap free bytes (a new test observable) return to baseline — kernel stacks reclaimed, no leak — and that the fault-recovery/kill paths reclaim too.

Gate: task-reap passes; fault-recovery, supervision, process-kill, aspace-refcount still green.

M9 — Futex-completion join (retire the per-thread endpoint)

With the reaper (M8) able to act after a thread is fully off its stack, migrate join to the std shape and drop M3's per-thread exit endpoint:

  • thread_spawn takes a user completion word (in the Thread handle's memory) and a joinable/detached flag. On reap the kernel writes 0 to that word and futex_wakes it (a CLONE_CHILD_CLEARTID equivalent — safe now the thread is off its stack). join = futex_wait on the word, then munmap the stack; a detached thread's stack is munmapped by the reaper instead. No IPC endpoint per thread.
  • thread-join passes on the new path; a check confirms joining N threads creates no per-thread endpoints (handle count stable).

Gate: thread-join/thread-mutex green on futex-completion join; guardrail green.

M10 — Per-thread TLS (threadlocal)

Give each thread its own threadlocal storage — the piece self-hosting Zig (zig-self-hosting.md) will force:

  • Runtime allocates a per-thread TLS block from the binary's PT_TLS template (linker symbols: copy .tdata, zero .tbss, variant-II TCB self-pointer) and hands its thread pointer to thread_spawn; the main thread sets its own via a new set_thread_pointer syscall in _start. The block is aspace memory → reclaimed on teardown.
  • Kernel stores fs_base on Task, loads it at first entry and restores it on context switch only when it changes (the same conditional-load pattern as CR3). getCurrentId can then read a TLS self-slot instead of a syscall.
  • -Dtest-case=thread-tls: two threads each write and read their own threadlocal slot with no cross-talk, and observe distinct getCurrentId.

Gate: thread-tls passes; full thread-* suite + guardrail green.

M11 — RwLock, WaitGroup, and host-testable sync

  • runtime.Thread.RwLock and WaitGroup on the existing Futex/Mutex/ Condition.
  • A compile-time Futex seam: syscalls on the danos target, a host-backed impl under zig build test, so the Mutex/Condition/RwLock state machines run as host unit tests (fast iteration, no QEMU).
  • -Dtest-case=thread-rwlock (smp: 4): many readers + writers over an RwLock keep an invariant (a reader never observes a half-written value); host tests cover the lock transitions.

Gate: host zig build test covers the sync primitives; thread-rwlock passes; guardrail green.


Deferred (explicitly not in this plan)

  • Cross-process shared-memory futex — the (aspace, vaddr) key can become a physical-address key so two processes share a futex through an shm region. Not needed for intra-process threads.
  • Per-thread priorities / affinity distinct from the process — threads inherit the process priority (scheduling.md); revisit only if it earns its keep.
  • Per-thread signal delivery — signals stay process-scoped (process-lifecycle.md).
  • A pthread/POSIX surface — the API is std.Thread-shaped Zig, nothing more.
  • A real std.Thread backend — arrives with self-hosting (zig-self-hosting.md); it sits on these same primitives, so it swaps the impl under runtime.Thread, not the call sites.