threads: plan Phase 2 (M7-M11) — hardening the deferred parts
Add M7-M11 to docs/threading-plan.md, designed so threading reinforces danos's goals: everything a thread owns lives in the address space (reclaimed on process death via the M1 refcount), the kernel owns mechanism while the runtime owns policy, and the process stays the isolation/restart boundary. M7 thread-safe allocation (per-aspace mmap arena + locked runtime heap); M8 task reaper (reclaim kernel stacks + detached user stacks — the resilience gap); M9 futex-completion join (retire the per-thread endpoint, built on M8); M10 per-thread TLS (threadlocal + fs.base, for self-hosting); M11 RwLock/WaitGroup + host-testable sync.
This commit is contained in:
+116
-6
@@ -265,13 +265,123 @@ green.
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Status: built
|
## Status
|
||||||
|
|
||||||
M1–M6 complete. danos has `runtime.Thread` — `spawn`/`join`/`detach`, cross-core
|
**Phase 1 (M1–M6): built.** danos has `runtime.Thread` — `spawn`/`join`/`detach`,
|
||||||
parallelism, futex, and `Mutex`/`Condition`/`Semaphore`, all over a private thread ABI
|
cross-core parallelism, futex, and `Mutex`/`Condition`/`Semaphore`, all over a private
|
||||||
behind the runtime. Deferred (with rationale, no consumer yet): `threadlocal` TLS,
|
thread ABI behind the runtime.
|
||||||
`RwLock`/`WaitGroup`, kernel clear-on-exit for a futex-completion `join`, a per-aspace
|
|
||||||
mmap arena / thread-safe runtime heap, and host-side unit tests via a mockable `Futex`.
|
**Phase 2 (M7–M11): planned below** — hardening the deferred parts so threads are safe
|
||||||
|
for real workloads and reclaimed like everything else danos owns.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Phase 2 — hardening (M7–M11)
|
||||||
|
|
||||||
|
The organising principle, so Phase 2 reinforces danos's goals rather than eroding them:
|
||||||
|
|
||||||
|
- **Everything a thread owns is reclaimed on process death.** Thread stacks, TLS blocks,
|
||||||
|
and futex words live in the process's **address space**, and the kernel's per-process
|
||||||
|
state is keyed by the aspace root — so the M1 refcount + `destroyAddressSpace` already
|
||||||
|
free all of it when the last thread exits. A crashed or killed threaded process leaves
|
||||||
|
**nothing** behind. Phase 2 closes the one thing that is *not* aspace-owned — the
|
||||||
|
per-task **kernel** stack (kernel heap) — with a reaper (M8). This is the
|
||||||
|
[resilience](resilience.md) restart guarantee, extended to threads.
|
||||||
|
- **Kernel owns mechanism; the runtime owns policy.** The kernel maps pages, saves/
|
||||||
|
restores `fs.base`, and reaps dead tasks; the runtime decides allocation, TLS layout,
|
||||||
|
and lock algorithms. Every new kernel entry stays a private syscall behind the runtime
|
||||||
|
([syscall.md](syscall.md)) — the ABI stays renumberable.
|
||||||
|
- **The process is still the isolation and restart boundary.** Threads share fate within
|
||||||
|
one process; Phase 2 never adds a way for one process to reach into another (the
|
||||||
|
cross-process futex stays explicitly out of scope, below).
|
||||||
|
|
||||||
|
### M7 — Thread-safe allocation (the correctness gap)
|
||||||
|
|
||||||
|
Today the mmap arena cursor is per-*task* and the runtime heap is unlocked, so two
|
||||||
|
threads in one process that both allocate corrupt each other. The thread *machinery*
|
||||||
|
avoids this (closure on the stack, stacks mmap'd only by the spawner), but real
|
||||||
|
multi-threaded code would hit it. Close it:
|
||||||
|
|
||||||
|
- [ ] **Kernel — per-address-space mmap arena.** Grow M1's `aspace_refs` entry into a
|
||||||
|
small address-space object holding the `mmap`/`mmio` arena cursors (moved off
|
||||||
|
`Task`); `systemMmap`/`mmio_map` bump the *aspace's* cursor under the big lock, so
|
||||||
|
sibling threads get disjoint, serialized grants. Freed at refcount zero, so the
|
||||||
|
cursors vanish with the process.
|
||||||
|
- [ ] **Runtime — thread-safe heap.** Guard the allocator with a `Thread.Mutex`, gated on
|
||||||
|
`!@import("builtin").single_threaded` so single-threaded binaries compile it out and
|
||||||
|
pay nothing. (The heap grows via mmap, now safe per above.)
|
||||||
|
- [ ] `-Dtest-case=thread-alloc` (`smp: 4`): N threads each do many `alloc`/`free` of
|
||||||
|
varied sizes, write a per-thread pattern, verify it, and free; assert every block
|
||||||
|
round-trips intact and all memory returns — no corruption under concurrent
|
||||||
|
allocation. A direct check confirms two threads' concurrent `mmap`s are disjoint.
|
||||||
|
|
||||||
|
**Gate:** `thread-alloc` passes; guardrail + all `thread-*` cases green.
|
||||||
|
|
||||||
|
### M8 — The task reaper (cleanup + resilience)
|
||||||
|
|
||||||
|
A dead task's **kernel** stack is currently leaked ("no reaper yet") — every process
|
||||||
|
*and* thread death loses one, so a crash loop bleeds kernel memory. A reaper fixes it and
|
||||||
|
serves the [resilience](resilience.md) restart goal directly:
|
||||||
|
|
||||||
|
- [ ] A dying task cannot free the kernel stack it runs on, so it hands itself to a
|
||||||
|
**reap list** and switches away; the kernel stack (and, for a detached thread, its
|
||||||
|
user stack) is reclaimed from another context — a low-priority reaper step drained
|
||||||
|
on the scheduler tick and when a core goes idle. Extends the existing
|
||||||
|
`reap_task_hook`/`destroyTaskLocked` path rather than inventing a parallel one.
|
||||||
|
- [ ] `-Dtest-case=task-reap`: spawn and exit many threads and processes; assert the
|
||||||
|
kernel-heap free bytes (a new test observable) return to **baseline** — kernel
|
||||||
|
stacks reclaimed, no leak — and that the `fault-recovery`/kill paths reclaim too.
|
||||||
|
|
||||||
|
**Gate:** `task-reap` passes; `fault-recovery`, `supervision`, `process-kill`,
|
||||||
|
`aspace-refcount` still green.
|
||||||
|
|
||||||
|
### M9 — Futex-completion join (retire the per-thread endpoint)
|
||||||
|
|
||||||
|
With the reaper (M8) able to act *after* a thread is fully off its stack, migrate `join`
|
||||||
|
to the std shape and drop M3's per-thread exit endpoint:
|
||||||
|
|
||||||
|
- [ ] `thread_spawn` takes a user **completion word** (in the `Thread` handle's memory)
|
||||||
|
and a joinable/detached flag. On reap the kernel writes 0 to that word and
|
||||||
|
`futex_wake`s it (a `CLONE_CHILD_CLEARTID` equivalent — safe now the thread is off
|
||||||
|
its stack). `join` = `futex_wait` on the word, then `munmap` the stack; a
|
||||||
|
**detached** thread's stack is `munmap`ped by the reaper instead. No IPC endpoint
|
||||||
|
per thread.
|
||||||
|
- [ ] `thread-join` passes on the new path; a check confirms joining N threads creates no
|
||||||
|
per-thread endpoints (handle count stable).
|
||||||
|
|
||||||
|
**Gate:** `thread-join`/`thread-mutex` green on futex-completion join; guardrail green.
|
||||||
|
|
||||||
|
### M10 — Per-thread TLS (`threadlocal`)
|
||||||
|
|
||||||
|
Give each thread its own `threadlocal` storage — the piece self-hosting Zig
|
||||||
|
([zig-self-hosting.md](zig-self-hosting.md)) will force:
|
||||||
|
|
||||||
|
- [ ] **Runtime** allocates a per-thread TLS block from the binary's `PT_TLS` template
|
||||||
|
(linker symbols: copy `.tdata`, zero `.tbss`, variant-II TCB self-pointer) and hands
|
||||||
|
its thread pointer to `thread_spawn`; the main thread sets its own via a new
|
||||||
|
`set_thread_pointer` syscall in `_start`. The block is aspace memory → reclaimed on
|
||||||
|
teardown.
|
||||||
|
- [ ] **Kernel** stores `fs_base` on `Task`, loads it at first entry and restores it on
|
||||||
|
context switch only when it changes (the same conditional-load pattern as CR3).
|
||||||
|
`getCurrentId` can then read a TLS self-slot instead of a syscall.
|
||||||
|
- [ ] `-Dtest-case=thread-tls`: two threads each write and read their own `threadlocal`
|
||||||
|
slot with no cross-talk, and observe distinct `getCurrentId`.
|
||||||
|
|
||||||
|
**Gate:** `thread-tls` passes; full `thread-*` suite + guardrail green.
|
||||||
|
|
||||||
|
### M11 — `RwLock`, `WaitGroup`, and host-testable sync
|
||||||
|
|
||||||
|
- [ ] `runtime.Thread.RwLock` and `WaitGroup` on the existing `Futex`/`Mutex`/
|
||||||
|
`Condition`.
|
||||||
|
- [ ] A compile-time `Futex` seam: syscalls on the danos target, a host-backed impl under
|
||||||
|
`zig build test`, so the `Mutex`/`Condition`/`RwLock` state machines run as host
|
||||||
|
unit tests (fast iteration, no QEMU).
|
||||||
|
- [ ] `-Dtest-case=thread-rwlock` (`smp: 4`): many readers + writers over an `RwLock` keep
|
||||||
|
an invariant (a reader never observes a half-written value); host tests cover the
|
||||||
|
lock transitions.
|
||||||
|
|
||||||
|
**Gate:** host `zig build test` covers the sync primitives; `thread-rwlock` passes;
|
||||||
|
guardrail green.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user