threads(M7): thread-safe allocation (per-aspace mmap arena + locked heap)

Move the mmap/mmio grant-arena cursors off Task into the per-address-space object
(scheduler aspace_refs, exposed via aspaceMmapNextPtr/aspaceDeviceMapNextPtr), so
sibling threads in one address space hand out disjoint grants. systemMmap reserves
a range under a brief lock then maps per page under a short-held lock (not the
whole grant): the big lock runs with interrupts disabled, so pinning it across a
multi-MiB memset+map froze other cores. Guard the runtime heap's rawAlloc/rawFree
with a Thread.Mutex, gated on !single_threaded so ordinary binaries compile it out.

thread-test gains an alloc mode: 4 threads x 500 alloc/fill/verify/free cycles;
any overlap between concurrent allocations is caught by the pattern check.

Also fix the affinity guardrail: its 3-billion-iteration busy-loop had
codegen-dependent wall-time (adding a function to tests.zig swung it ~4s -> ~63s
and timed it out). Reworked to wait on the wall clock instead.

Gate thread-alloc PASS (3x); full guardrail 23/23 green; build + host tests clean.
This commit is contained in:
2026-07-20 22:57:11 +01:00
parent f4813c8e99
commit 8259678f0a
7 changed files with 228 additions and 47 deletions
+25 -3
View File
@@ -9,15 +9,32 @@
//! — `grow` asks the kernel for pages via `mmap` instead of mapping frames
//! itself, and the kernel picks the base address.
//!
//! Single-threaded and 16-byte maximum alignment, exactly like the kernel heap; a
//! lock and larger alignments come when user programs gain threads.
//! 16-byte maximum alignment, exactly like the kernel heap. The free list is guarded by
//! a `Thread.Mutex` **only in multi-threaded binaries** (`addThreadedUserBinary`): the
//! guard is gated on `builtin.single_threaded`, so an ordinary single-threaded binary
//! compiles it out and pays nothing, while a threaded one can allocate safely from
//! several threads at once (docs/threading-plan.md M7). The lock lives at the two
//! free-list mutators — `rawAlloc`/`rawFree` — which every entry point funnels through.
const std = @import("std");
const builtin = @import("builtin");
const abi = @import("abi");
const system_calls = @import("system.zig");
const Mutex = @import("thread.zig").Thread.Mutex;
const page_size = abi.page_size;
/// Guards `free_list`. A no-op in single-threaded builds (compiled out); a real futex
/// mutex in threaded ones. Uncontended acquisition is a single CAS — no syscall.
var heap_mutex: Mutex = .{};
inline fn lockHeap() void {
if (comptime !builtin.single_threaded) heap_mutex.lock();
}
inline fn unlockHeap() void {
if (comptime !builtin.single_threaded) heap_mutex.unlock();
}
/// A block header, at the start of every block; while free it also links the
/// free list via `next`.
const Block = extern struct {
@@ -84,8 +101,11 @@ fn insertFree(block: *Block) void {
}
}
/// Allocate `len` bytes (16-byte aligned), or null if out of memory.
/// Allocate `len` bytes (16-byte aligned), or null if out of memory. Holds the heap lock
/// across the free-list search and any `grow` (which also touches the free list).
fn rawAlloc(len: usize) ?[*]u8 {
lockHeap();
defer unlockHeap();
const need = alignUp(header_size + len, 16);
var attempts: u32 = 0;
@@ -119,6 +139,8 @@ fn rawAlloc(len: usize) ?[*]u8 {
}
fn rawFree(ptr: [*]u8) void {
lockHeap();
defer unlockHeap();
const block: *Block = @ptrFromInt(@intFromPtr(ptr) - header_size);
insertFree(block);
}