Completes the std.Thread-shaped runtime.Thread. Phase 1 (M1-M6: address-space
refcount, thread_spawn/exit, join/detach, futex + Mutex/Condition/Semaphore,
getCurrentId) was already on main; this brings the Phase 2 hardening and the
coding-standards/arch-neutrality passes on top:
M7 thread-safe allocation (per-address-space mmap arena + locked heap)
M8 the task reaper — reclaim dead tasks' kernel stacks
M9 thread_join syscall — retire the per-thread endpoint
M10 per-thread thread pointer — the TLS mechanism (x86_64 IA32_FS_BASE)
M11 RwLock, WaitGroup, and host-testable sync
Plus: the TLS thread pointer named arch-neutrally (not fs.base) so the kernel
stays architecture-agnostic; aspace/vaddr/paddr spelled out per
docs/coding-standards.md across kernel, runtime, ABI, tests, and docs; and the
misleading fs.base dot-notation dropped in favour of "thread pointer"
(arch-neutral) / "FS base" (x86-specific).
Deferred by design (no consumer yet): the Zig threadlocal *compiler* layer
(M10) and detached-thread user-stack reclaim (M9) — both noted in place.
Verified: zig build, zig build test, and the full 25-case QEMU guardrail suite
all green.