2 Commits
Author SHA1 Message Date
daniel f4813c8e99 threads: make the Phase 2 plan loop-runnable
Adjust the unattended loop contract for Phase 2: the Done condition targets M1
through M11 (M1-M6 being checked is no longer Done), the loop branches off the
current main into a new branch (Phase 1's threading is merged), and each green
milestone pushes the working branch to origin (main stays a human merge).
2026-07-20 22:22:33 +01:00
daniel 8b7f1d009c threads: plan Phase 2 (M7-M11) — hardening the deferred parts
Add M7-M11 to docs/threading-plan.md, designed so threading reinforces danos's
goals: everything a thread owns lives in the address space (reclaimed on process
death via the M1 refcount), the kernel owns mechanism while the runtime owns
policy, and the process stays the isolation/restart boundary.

M7 thread-safe allocation (per-aspace mmap arena + locked runtime heap);
M8 task reaper (reclaim kernel stacks + detached user stacks — the resilience gap);
M9 futex-completion join (retire the per-thread endpoint, built on M8);
M10 per-thread TLS (threadlocal + fs.base, for self-hosting);
M11 RwLock/WaitGroup + host-testable sync.
2026-07-20 22:10:32 +01:00
+133 -15
View File
@@ -52,8 +52,11 @@ fixed in *Locked decisions*; the checkboxes are the only state. A loop iteration
1. **Resume** at the first milestone that still has an unchecked `- [ ]`. (All earlier 1. **Resume** at the first milestone that still has an unchecked `- [ ]`. (All earlier
milestones are done — do not revisit them.) milestones are done — do not revisit them.)
2. **Work on a branch.** On the first iteration, branch off `main` (e.g. `threading`); 2. **Work on a branch.** On the first iteration, branch off the current `main` into a new
never commit threading work to `main`. All work stays local — **do not push**. branch (e.g. `threading-phase2` — Phase 1's `threading` is already merged); never
commit to `main` directly. Push that **branch** to `origin` after each milestone (step
5) so progress is backed up remotely; **do not push `main`** — merging Phase 2 into
`main` stays a human step.
3. **Implement** every unchecked item in that milestone, including adding its 3. **Implement** every unchecked item in that milestone, including adding its
`-Dtest-case` to `CASES` in [test/qemu_test.py](../test/qemu_test.py) (with `-Dtest-case` to `CASES` in [test/qemu_test.py](../test/qemu_test.py) (with
`smp: true` / a `mem` bump where noted) so the gate is runnable. `smp: true` / a `mem` bump where noted) so the gate is runnable.
@@ -64,8 +67,9 @@ fixed in *Locked decisions*; the checkboxes are the only state. A loop iteration
the whole guardrail set passes, `zig build` is clean, and host tests are green. the whole guardrail set passes, `zig build` is clean, and host tests are green.
→ tick this milestone's boxes **and** its `**Gate:**`-referenced case, `git commit` → tick this milestone's boxes **and** its `**Gate:**`-referenced case, `git commit`
(`threads(M<n>): <summary>`, no `Co-Authored-By` trailer per (`threads(M<n>): <summary>`, no `Co-Authored-By` trailer per
[coding-standards.md](coding-standards.md)), and continue to the next milestone in [coding-standards.md](coding-standards.md)), then **`git push` the working branch to
the same iteration if budget remains; otherwise let the loop re-fire. `origin`** (use `-u` on the first push to set upstream). Continue to the next
milestone in the same iteration if budget remains; otherwise let the loop re-fire.
- **Red** = anything above fails. Diagnose from the captured serial log - **Red** = anything above fails. Diagnose from the captured serial log
(`zig-out/qemu-test/<case>-failed-serial.log`) and fix in place, then re-run — up to (`zig-out/qemu-test/<case>-failed-serial.log`) and fix in place, then re-run — up to
**3 fix attempts** for that gate. A concurrency case that fails then passes on a **3 fix attempts** for that gate. A concurrency case that fails then passes on a
@@ -79,14 +83,18 @@ fixed in *Locked decisions*; the checkboxes are the only state. A loop iteration
**The only stop conditions:** **The only stop conditions:**
- **Done** — every milestone box is checked (M1–M6), `zig build` clean, whole - **Done** — every milestone box **in this plan** is checked (M1 through M11), `zig build`
`thread-*` suite + guardrail green. Update threading.md's status line to "built" (that clean, the whole `thread-*` suite + guardrail green. Phase 1 (M1–M6) is *already*
is M6's own task) and stop. checked, so do **not** read that as Done: the loop's real work is the first plan section
that still has unchecked boxes — Phase 2 (M7–M11). Only stop when M7–M11 are all checked
too. Update threading.md's status line, push the final branch state to `origin`, and
stop. The branch is on `origin` for review; **merging Phase 2 into `main` is the user's
step**, not the loop's.
- **Blocked** — a gate is still red after 3 fix attempts, or a step needs something - **Blocked** — a gate is still red after 3 fix attempts, or a step needs something
outside the repo (a toolchain change, new hardware, a decision no locked decision outside the repo (a toolchain change, new hardware, a decision no locked decision
covers). Append `> **BLOCKED (M<n>):** <what failed, what was tried, the serial covers). Append `> **BLOCKED (M<n>):** <what failed, what was tried, the serial
marker missing>` under that milestone, commit the WIP on the branch, and stop. Do not marker missing>` under that milestone, commit **and push** the WIP on the branch, and
thrash further and do not silently skip the milestone. stop. Do not thrash further and do not silently skip the milestone.
Nothing else warrants stopping — not "should I proceed?", not "is this right?". The Nothing else warrants stopping — not "should I proceed?", not "is this right?". The
checkboxes + git history are the resumable record; the next iteration picks up from the checkboxes + git history are the resumable record; the next iteration picks up from the
@@ -265,13 +273,123 @@ green.
--- ---
## Status: built ## Status
M1–M6 complete. danos has `runtime.Thread` — `spawn`/`join`/`detach`, cross-core **Phase 1 (M1–M6): built.** danos has `runtime.Thread` — `spawn`/`join`/`detach`,
parallelism, futex, and `Mutex`/`Condition`/`Semaphore`, all over a private thread ABI cross-core parallelism, futex, and `Mutex`/`Condition`/`Semaphore`, all over a private
behind the runtime. Deferred (with rationale, no consumer yet): `threadlocal` TLS, thread ABI behind the runtime.
`RwLock`/`WaitGroup`, kernel clear-on-exit for a futex-completion `join`, a per-aspace
mmap arena / thread-safe runtime heap, and host-side unit tests via a mockable `Futex`. **Phase 2 (M7–M11): planned below** — hardening the deferred parts so threads are safe
for real workloads and reclaimed like everything else danos owns.
---
## Phase 2 — hardening (M7–M11)
The organising principle, so Phase 2 reinforces danos's goals rather than eroding them:
- **Everything a thread owns is reclaimed on process death.** Thread stacks, TLS blocks,
and futex words live in the process's **address space**, and the kernel's per-process
state is keyed by the aspace root — so the M1 refcount + `destroyAddressSpace` already
free all of it when the last thread exits. A crashed or killed threaded process leaves
**nothing** behind. Phase 2 closes the one thing that is *not* aspace-owned — the
per-task **kernel** stack (kernel heap) — with a reaper (M8). This is the
[resilience](resilience.md) restart guarantee, extended to threads.
- **Kernel owns mechanism; the runtime owns policy.** The kernel maps pages, saves/
restores `fs.base`, and reaps dead tasks; the runtime decides allocation, TLS layout,
and lock algorithms. Every new kernel entry stays a private syscall behind the runtime
([syscall.md](syscall.md)) — the ABI stays renumberable.
- **The process is still the isolation and restart boundary.** Threads share fate within
one process; Phase 2 never adds a way for one process to reach into another (the
cross-process futex stays explicitly out of scope, below).
### M7 — Thread-safe allocation (the correctness gap)
Today the mmap arena cursor is per-*task* and the runtime heap is unlocked, so two
threads in one process that both allocate corrupt each other. The thread *machinery*
avoids this (closure on the stack, stacks mmap'd only by the spawner), but real
multi-threaded code would hit it. Close it:
- [ ] **Kernel — per-address-space mmap arena.** Grow M1's `aspace_refs` entry into a
small address-space object holding the `mmap`/`mmio` arena cursors (moved off
`Task`); `systemMmap`/`mmio_map` bump the *aspace's* cursor under the big lock, so
sibling threads get disjoint, serialized grants. Freed at refcount zero, so the
cursors vanish with the process.
- [ ] **Runtime — thread-safe heap.** Guard the allocator with a `Thread.Mutex`, gated on
`!@import("builtin").single_threaded` so single-threaded binaries compile it out and
pay nothing. (The heap grows via mmap, now safe per above.)
- [ ] `-Dtest-case=thread-alloc` (`smp: 4`): N threads each do many `alloc`/`free` of
varied sizes, write a per-thread pattern, verify it, and free; assert every block
round-trips intact and all memory returns — no corruption under concurrent
allocation. A direct check confirms two threads' concurrent `mmap`s are disjoint.
**Gate:** `thread-alloc` passes; guardrail + all `thread-*` cases green.
### M8 — The task reaper (cleanup + resilience)
A dead task's **kernel** stack is currently leaked ("no reaper yet") — every process
*and* thread death loses one, so a crash loop bleeds kernel memory. A reaper fixes it and
serves the [resilience](resilience.md) restart goal directly:
- [ ] A dying task cannot free the kernel stack it runs on, so it hands itself to a
**reap list** and switches away; the kernel stack (and, for a detached thread, its
user stack) is reclaimed from another context — a low-priority reaper step drained
on the scheduler tick and when a core goes idle. Extends the existing
`reap_task_hook`/`destroyTaskLocked` path rather than inventing a parallel one.
- [ ] `-Dtest-case=task-reap`: spawn and exit many threads and processes; assert the
kernel-heap free bytes (a new test observable) return to **baseline** — kernel
stacks reclaimed, no leak — and that the `fault-recovery`/kill paths reclaim too.
**Gate:** `task-reap` passes; `fault-recovery`, `supervision`, `process-kill`,
`aspace-refcount` still green.
### M9 — Futex-completion join (retire the per-thread endpoint)
With the reaper (M8) able to act *after* a thread is fully off its stack, migrate `join`
to the std shape and drop M3's per-thread exit endpoint:
- [ ] `thread_spawn` takes a user **completion word** (in the `Thread` handle's memory)
and a joinable/detached flag. On reap the kernel writes 0 to that word and
`futex_wake`s it (a `CLONE_CHILD_CLEARTID` equivalent — safe now the thread is off
its stack). `join` = `futex_wait` on the word, then `munmap` the stack; a
**detached** thread's stack is `munmap`ped by the reaper instead. No IPC endpoint
per thread.
- [ ] `thread-join` passes on the new path; a check confirms joining N threads creates no
per-thread endpoints (handle count stable).
**Gate:** `thread-join`/`thread-mutex` green on futex-completion join; guardrail green.
### M10 — Per-thread TLS (`threadlocal`)
Give each thread its own `threadlocal` storage — the piece self-hosting Zig
([zig-self-hosting.md](zig-self-hosting.md)) will force:
- [ ] **Runtime** allocates a per-thread TLS block from the binary's `PT_TLS` template
(linker symbols: copy `.tdata`, zero `.tbss`, variant-II TCB self-pointer) and hands
its thread pointer to `thread_spawn`; the main thread sets its own via a new
`set_thread_pointer` syscall in `_start`. The block is aspace memory → reclaimed on
teardown.
- [ ] **Kernel** stores `fs_base` on `Task`, loads it at first entry and restores it on
context switch only when it changes (the same conditional-load pattern as CR3).
`getCurrentId` can then read a TLS self-slot instead of a syscall.
- [ ] `-Dtest-case=thread-tls`: two threads each write and read their own `threadlocal`
slot with no cross-talk, and observe distinct `getCurrentId`.
**Gate:** `thread-tls` passes; full `thread-*` suite + guardrail green.
### M11 — `RwLock`, `WaitGroup`, and host-testable sync
- [ ] `runtime.Thread.RwLock` and `WaitGroup` on the existing `Futex`/`Mutex`/
`Condition`.
- [ ] A compile-time `Futex` seam: syscalls on the danos target, a host-backed impl under
`zig build test`, so the `Mutex`/`Condition`/`RwLock` state machines run as host
unit tests (fast iteration, no QEMU).
- [ ] `-Dtest-case=thread-rwlock` (`smp: 4`): many readers + writers over an `RwLock` keep
an invariant (a reader never observes a half-written value); host tests cover the
lock transitions.
**Gate:** host `zig build test` covers the sync primitives; `thread-rwlock` passes;
guardrail green.
--- ---