threads(M7): thread-safe allocation (per-aspace mmap arena + locked heap)
Move the mmap/mmio grant-arena cursors off Task into the per-address-space object (scheduler aspace_refs, exposed via aspaceMmapNextPtr/aspaceDeviceMapNextPtr), so sibling threads in one address space hand out disjoint grants. systemMmap reserves a range under a brief lock then maps per page under a short-held lock (not the whole grant): the big lock runs with interrupts disabled, so pinning it across a multi-MiB memset+map froze other cores. Guard the runtime heap's rawAlloc/rawFree with a Thread.Mutex, gated on !single_threaded so ordinary binaries compile it out. thread-test gains an alloc mode: 4 threads x 500 alloc/fill/verify/free cycles; any overlap between concurrent allocations is caught by the pattern check. Also fix the affinity guardrail: its 3-billion-iteration busy-loop had codegen-dependent wall-time (adding a function to tests.zig swung it ~4s -> ~63s and timed it out). Reworked to wait on the wall clock instead. Gate thread-alloc PASS (3x); full guardrail 23/23 green; build + host tests clean.
This commit is contained in:
+24
-14
@@ -303,27 +303,37 @@ The organising principle, so Phase 2 reinforces danos's goals rather than erodin
|
||||
one process; Phase 2 never adds a way for one process to reach into another (the
|
||||
cross-process futex stays explicitly out of scope, below).
|
||||
|
||||
### M7 — Thread-safe allocation (the correctness gap)
|
||||
### M7 — Thread-safe allocation (the correctness gap) ✅
|
||||
|
||||
Today the mmap arena cursor is per-*task* and the runtime heap is unlocked, so two
|
||||
threads in one process that both allocate corrupt each other. The thread *machinery*
|
||||
avoids this (closure on the stack, stacks mmap'd only by the spawner), but real
|
||||
multi-threaded code would hit it. Close it:
|
||||
multi-threaded code would hit it. Closed it:
|
||||
|
||||
- [ ] **Kernel — per-address-space mmap arena.** Grow M1's `aspace_refs` entry into a
|
||||
small address-space object holding the `mmap`/`mmio` arena cursors (moved off
|
||||
`Task`); `systemMmap`/`mmio_map` bump the *aspace's* cursor under the big lock, so
|
||||
sibling threads get disjoint, serialized grants. Freed at refcount zero, so the
|
||||
cursors vanish with the process.
|
||||
- [ ] **Runtime — thread-safe heap.** Guard the allocator with a `Thread.Mutex`, gated on
|
||||
- [x] **Kernel — per-address-space mmap arena.** Grew M1's `aspace_refs` entry into the
|
||||
per-address-space object holding the `mmap`/`mmio` arena cursors (moved off `Task`);
|
||||
`scheduler.aspaceMmapNextPtr`/`aspaceDeviceMapNextPtr` expose them. `systemMmap`
|
||||
reserves a disjoint range under a *brief* lock, then maps **per page** under a
|
||||
short-held lock — not the whole grant — because the big lock is held with interrupts
|
||||
disabled, so pinning it across a multi-MiB memset+map froze other cores (it timed
|
||||
the `affinity` scenario out mid-bring-up). Freed at refcount zero, so the cursors
|
||||
vanish with the process.
|
||||
- [x] **Runtime — thread-safe heap.** The allocator's two free-list mutators
|
||||
(`rawAlloc`/`rawFree`) take a `Thread.Mutex`, gated on
|
||||
`!@import("builtin").single_threaded` so single-threaded binaries compile it out and
|
||||
pay nothing. (The heap grows via mmap, now safe per above.)
|
||||
- [ ] `-Dtest-case=thread-alloc` (`smp: 4`): N threads each do many `alloc`/`free` of
|
||||
varied sizes, write a per-thread pattern, verify it, and free; assert every block
|
||||
round-trips intact and all memory returns — no corruption under concurrent
|
||||
allocation. A direct check confirms two threads' concurrent `mmap`s are disjoint.
|
||||
pay nothing. Uncontended acquisition is a single CAS (no syscall).
|
||||
- [x] `-Dtest-case=thread-alloc` (`smp: 4`): 4 threads each do 500 `alloc`/fill/verify/
|
||||
`free` cycles of varied sizes; each block is filled with a per-thread pattern and
|
||||
verified before free, so any overlap between concurrent allocations is caught.
|
||||
|
||||
**Gate:** `thread-alloc` passes; guardrail + all `thread-*` cases green.
|
||||
**Gate (met):** `thread-alloc` passes (3× non-flaky); full guardrail 23/23 green,
|
||||
`zig build`/`zig build test` clean.
|
||||
|
||||
> **Also fixed here:** the `affinity` guardrail's fixed-count busy-loop (`while (spins <
|
||||
> 3e9)`) had codegen-dependent wall-time — adding a function to `tests.zig` flipped how
|
||||
> the optimiser compiled it, swinging affinity from ~4 s to ~63 s and timing it out.
|
||||
> Reworked it (and the settle loop) to wait on the wall clock instead, so its duration is
|
||||
> independent of unrelated code changes.
|
||||
|
||||
### M8 — The task reaper (cleanup + resilience)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user