docs+code: drop misleading fs.base notation for the thread pointer

The "fs.base" spelling read like a field/submodule access, but no such
identifier exists — it meant the x86_64 FS segment base (the IA32_FS_BASE
MSR). Two problems fixed:

- Arch-neutrality: in generic docs, the runtime, and the plan, the mechanism
  is now named by its arch-neutral concept — the "thread pointer" — matching
  the already-renamed `thread_pointer` Task field, `set_thread_pointer`
  syscall, and `architecture.setThreadPointer` fn. x86-specific spots keep
  the precise names: `IA32_FS_BASE` (the MSR), `%fs:8`/`%fs:0`, variant-II.

- Stale identifiers: the M10 section in threading-plan.md still referenced
  `fs_base` on Task and `architecture.setFsBase` — both renamed away in the
  arch-neutral pass. Corrected to `thread_pointer` / `setThreadPointer`.

The x86-only `thread-tls` test (which really does write `%fs:8`) now says
"FS base" (no dot) consistently, matching the established form already in
tests.zig. The matched serial markers ("thread-tls: ok" / "thread-tls:
FAIL") are unchanged; only a non-load-bearing FAIL parenthetical was
reworded.

Verified: zig build clean, thread-tls passes.
This commit is contained in:
2026-07-21 01:52:10 +01:00
parent 28b3635979
commit e2dddc941f
5 changed files with 28 additions and 28 deletions
+13 -13
View File
@@ -252,11 +252,11 @@ green.
- [x] `getCurrentId` via a small `thread_self = 42` syscall (`runtime.Thread.getCurrentId`
returns the kernel task id). **Per-thread `threadlocal` TLS is deferred** — no
consumer needs it, and it would require context-switching `fs.base` per task (real
kernel + per-switch cost) for an unused feature; threaded binaries have run fine
consumer needs it, and it would require context-switching the thread pointer per task
(real kernel + per-switch cost) for an unused feature; threaded binaries have run fine
without it through M2–M5. threading.md's TLS reasoning already scoped it as
deferred-unless-needed. When a consumer appears, the shape is: `thread_spawn`
allocates a per-thread TLS block, sets `fs.base`, and the context switch saves/
allocates a per-thread TLS block, sets the thread pointer, and the context switch saves/
restores it.
- [x] `RwLock` / `WaitGroup` deferred (no consumer yet); they slot onto the same
`Futex`/`Mutex`/`Condition` when wanted.
@@ -280,7 +280,7 @@ cross-core parallelism, futex, and `Mutex`/`Condition`/`Semaphore`, all over a p
thread ABI behind the runtime.
**Phase 2 (M7–M11): built.** Thread-safe allocation (M7), a task reaper that reclaims dead
tasks' kernel stacks (M8), endpoint-free `thread_join` (M9), per-thread `fs.base` (M10),
tasks' kernel stacks (M8), endpoint-free `thread_join` (M9), the per-thread thread pointer (M10),
and `RwLock`/`WaitGroup` + host-testable sync (M11). Two things stay deferred by design
(no consumer): the Zig `threadlocal` *compiler* layer (M10) and detached-thread user-stack
reclaim (M9) — both noted in place.
@@ -299,7 +299,7 @@ The organising principle, so Phase 2 reinforces danos's goals rather than erodin
per-task **kernel** stack (kernel heap) — with a reaper (M8). This is the
[resilience](resilience.md) restart guarantee, extended to threads.
- **Kernel owns mechanism; the runtime owns policy.** The kernel maps pages, saves/
restores `fs.base`, and reaps dead tasks; the runtime decides allocation, TLS layout,
restores the thread pointer, and reaps dead tasks; the runtime decides allocation, TLS layout,
and lock algorithms. Every new kernel entry stays a private syscall behind the runtime
([syscall.md](syscall.md)) — the ABI stays renumberable.
- **The process is still the isolation and restart boundary.** Threads share fate within
@@ -400,25 +400,25 @@ guardrail 26/26 (incl. `process-kill`, `supervision`, `fault-recovery`, `task-re
> translate/unmap in a not-currently-loaded address space — real complexity for a bounded leak.
> A follow-up when a consumer needs it.
### M10 — Per-thread TLS: the `fs.base` mechanism ✅
### M10 — Per-thread TLS: the thread-pointer mechanism ✅
Give each thread its own thread pointer and private TLS storage — the foundation
self-hosting Zig ([zig-self-hosting.md](zig-self-hosting.md)) will build `threadlocal` on.
- [x] **Kernel** stores `fs_base` on `Task` and restores it on every context switch
- [x] **Kernel** stores `thread_pointer` on `Task` and restores it on every context switch
**only when it changes** (the same conditional-load discipline as CR3;
`architecture.setFsBase` → `wrmsr IA32_FS_BASE`). A `set_thread_pointer(addr)` = 44
syscall sets the caller's `fs_base` and loads it now. The kernel never touches FS, so
there is no swapgs complication.
`architecture.setThreadPointer` → `wrmsr IA32_FS_BASE` on x86_64). A
`set_thread_pointer(addr)` = 44 syscall sets the caller's `thread_pointer` and loads it
now. The kernel never touches FS, so there is no swapgs complication.
- [x] **Runtime** lays a small per-thread TLS block at the top of each thread's stack
(self-pointer at `%fs:0` + scratch slots) and the thread trampoline calls
`set_thread_pointer` before any user code — so every spawned thread has a private,
switch-stable thread pointer. Reclaimed with the stack.
- [x] `-Dtest-case=thread-tls` (`smp: 4`): two threads each write a unique marker to their
own `%fs:8` slot and — after both have written — read it back; a shared (non-per-thread)
fs.base would clobber one and cause cross-talk. Both read their own marker → pass.
FS base would clobber one and cause cross-talk. Both read their own marker → pass.
**Gate (met):** `thread-tls` passes (3×); full guardrail 25/25 (the switch-time `fs.base`
**Gate (met):** `thread-tls` passes (3×); full guardrail 25/25 (the switch-time thread-pointer
restore touches every context switch); `zig build`/`zig build test` clean.
> **Deferred: the Zig `threadlocal` *compiler* layer.** Real `threadlocal` variables need
@@ -426,7 +426,7 @@ restore touches every context switch); `zig build`/`zig build test` clean.
> in `user.ld`, a runtime that copies the template with exact negative-offset layout, and
> the `.large`-code-model TLS section names — a high-uncertainty lift for a feature with
> **no consumer today** (threading.md scopes it "only if a consumer needs it"). What lands
> here is the load-bearing piece — per-thread `fs.base`, context-switched — so adding the
> here is the load-bearing piece — the per-thread thread pointer, context-switched — so adding the
> compiler layer later is purely runtime+linker work on top, no kernel change. `getCurrentId`
> stays the `thread_self` syscall (M6) rather than an fs self-slot (which would need the
> main thread's TLS set up in `_start` too).
+6 -6
View File
@@ -5,9 +5,9 @@ A note on danos **threads** — several tasks sharing one address space — prov
kernel entry behind the [runtime](../library/runtime). **Built** (M1–M11, see
[threading-plan.md](threading-plan.md)): `spawn`/`join`/`detach`, cross-core parallelism,
a futex, `Mutex`/`Condition`/`Semaphore`/`RwLock`/`WaitGroup`, `getCurrentId`/`currentCore`,
per-thread `fs.base` TLS, thread-safe allocation, and a task reaper that reclaims dead
per-thread thread-pointer TLS, thread-safe allocation, and a task reaper that reclaims dead
tasks' kernel stacks. Deferred by design (no consumer yet): the Zig `threadlocal`
*compiler* layer (per-thread `fs.base` is in place, so it's runtime+linker work on top) and
*compiler* layer (the per-thread thread pointer is in place, so it's runtime+linker work on top) and
detached-thread user-stack reclaim — see the plan's M9/M10 notes. The analysis is against
**Zig 0.16** (the pinned toolchain); `std.Thread`'s internals move between releases, so
treat upstream shapes as "0.16.x."
@@ -219,12 +219,12 @@ decision, not a "maybe later."
### TLS and `getCurrentId`
danos sets up no `fs.base` TLS today (fine under `single_threaded`). Two scoped needs:
danos sets up no thread-pointer TLS today (fine under `single_threaded`). Two scoped needs:
- **`getCurrentId`** returns the kernel task id — either a trivial syscall or, better,
a value the runtime stashes in a per-thread control block.
- **`threadlocal` variables** need a real per-thread TLS block and `fs.base` set per
thread. `thread_spawn` sets `fs.base` to a runtime-allocated per-thread block; full
- **`threadlocal` variables** need a real per-thread TLS block and the thread pointer set per
thread. `thread_spawn` sets the thread pointer to a runtime-allocated per-thread block; full
`threadlocal` support is Stage 3, only if a consumer needs it. Nothing in the core
spawn/join/mutex path requires `threadlocal`.
@@ -273,7 +273,7 @@ a verifiable gate (`python3 test/qemu_test.py <case>`, asserting serial markers;
word. *Gate:* `-Dtest-case=thread-mutex` — a bounded producer/consumer over a
`Mutex` + `Condition` moves K items with no lost wakeups and no busy-wait (assert
the consumer blocked, e.g. via a low idle tick count).
- **Stage 3 — polish.** Per-thread TLS / `fs.base` and `threadlocal` (only if a
- **Stage 3 — polish.** Per-thread TLS / thread pointer and `threadlocal` (only if a
consumer needs it), `RwLock`/`WaitGroup` as demanded, and this doc's cases wired
into [test/qemu_test.py](../test/qemu_test.py).
+2 -2
View File
@@ -60,7 +60,7 @@ pub const Thread = struct {
/// thread's TLS pointer, runs the user function, then ends the thread.
fn entry(self_addr: usize) callconv(.c) noreturn {
const self: *@This() = @ptrFromInt(self_addr);
setThreadPointer(self.tls_base); // per-thread FS base before any user code
setThreadPointer(self.tls_base); // per-thread thread pointer before any user code
@call(.auto, function, self.args);
exitThread();
}
@@ -70,7 +70,7 @@ pub const Thread = struct {
if (system.mmapFailed(base)) return error.SystemResources;
// Top of the thread's own stack, downward: the closure, then a small per-thread TLS
// block (fs.base points here; slot 0 is the variant-II self-pointer, the rest is
// block (the thread pointer points here; slot 0 is the variant-II self-pointer, the rest is
// scratch for user TLS), then the stack proper (rsp starts below the TLS block, so
// the growing stack never overwrites either).
var closure_addr = (base + config.stack_size) - @sizeOf(Closure);
+3 -3
View File
@@ -1744,9 +1744,9 @@ fn threadAllocTest(boot_information: *const BootInformation) void {
result();
}
/// Per-thread TLS / fs.base (docs/threading-plan.md M10): `thread-test` in tls mode has two
/// Per-thread TLS / FS base (docs/threading-plan.md M10): `thread-test` in tls mode has two
/// threads each set their own FS base and write a unique marker to `%fs:8`, then — after
/// both have written — read it back. If fs.base were not per-thread and restored across
/// both have written — read it back. If the FS base were not per-thread and restored across
/// context switches, the second write would clobber the first and a thread would read the
/// wrong marker. The verdict marker means both read their own value (no cross-talk).
fn threadTlsTest(boot_information: *const BootInformation) void {
@@ -1783,7 +1783,7 @@ fn threadTlsTest(boot_information: *const BootInformation) void {
}
scheduler.setPriority(4);
check("each thread has its own fs.base TLS slot (no cross-talk across switches)", bufferHas(ok_marker) and !bufferHas(fail_marker));
check("each thread has its own FS-base TLS slot (no cross-talk across switches)", bufferHas(ok_marker) and !bufferHas(fail_marker));
result();
}
+4 -4
View File
@@ -357,7 +357,7 @@ fn runAllocMode() void {
write("thread-alloc: ok\n"); // the M7 verdict marker
}
// --- M10: tls mode (per-thread fs.base storage) -----------------------------
// --- M10: tls mode (per-thread FS base storage) -----------------------------
fn writeTlsSlot(value: u64) void {
asm volatile ("movq %[v], %%fs:8"
@@ -379,9 +379,9 @@ var tls_ok = std.atomic.Value(u32).init(0);
fn tlsWorker(marker: u64) void {
writeTlsSlot(marker);
_ = tls_written.fetchAdd(1, .release);
// Wait until both threads have written their own slot. If fs.base were shared, the
// Wait until both threads have written their own slot. If the FS base were shared, the
// second write would clobber the first, and the read below would return the wrong
// marker — cross-talk. Per-thread fs.base keeps each thread's slot private.
// marker — cross-talk. A per-thread FS base keeps each thread's slot private.
var spins: usize = 0;
while (tls_written.load(.acquire) < 2 and spins < 50_000_000) : (spins += 1) {
runtime.system.yield();
@@ -406,7 +406,7 @@ fn runTlsMode() void {
if (tls_ok.load(.acquire) == 2) {
write("thread-tls: ok\n"); // the M10 verdict marker
} else {
write("thread-tls: FAIL cross-talk (fs.base not per-thread)\n");
write("thread-tls: FAIL cross-talk (FS base not per-thread)\n");
}
}