wake application processors to long mode

INIT-SIPI-SIPI plus a self-relocating real-mode trampoline.
This commit is contained in:
Daniel Samson
2026-07-08 12:35:30 +01:00
parent 36c29d2d6d
commit ed7f542006
12 changed files with 482 additions and 8 deletions
+66 -6
View File
@@ -16,10 +16,13 @@ danos specifics):
- **Real-time** — whether timing is *predictable*. Comes from bounded operations
(our O(1) scheduler), not from core count.
danos is uniprocessor today: one global `current` task, one set of ready queues, one
timer. Even on an 8-core CPU, the firmware starts only the **bootstrap processor
(BSP)**; the other cores (**application processors**, APs) sit parked until the
kernel wakes them, which it doesn't yet.
danos is mid-transition. The firmware starts only the **bootstrap processor (BSP)**;
the other cores (**application processors**, APs) sit parked until the kernel wakes
them. As of the SMP work in progress (see [Implementation status](#implementation-status)
below), danos now *does* wake the APs — each climbs to 64-bit long mode and reports
in — and the shared kernel state (scheduler queues, IPC) is already serialised behind
a big kernel lock. What's not done yet is letting the woken APs actually run tasks;
`current` is per-CPU but the run loop is still BSP-only.
## The common microkernel instinct: don't share kernel state
@@ -119,12 +122,15 @@ Whatever the top goal, the *sequence* is the same and seL4 validates starting si
and `platform.cpus()` returns the list (see [discovery.md](discovery.md)). The boot
log reports the count; the ARM (device-tree) path still needs it.
2. **Wake the APs** — INIT–SIPI–SIPI on x86; PSCI/spin-tables on ARM. Each core brings
up its own tables, timer, and idle task.
up its own tables, timer, and idle task. **In progress on x86** — the cores reach
long mode and park; the per-core tables/timer/scheduler entry is the next step
([status](#implementation-status)).
3. **Start with a big kernel lock.** It's a legitimate first design, not a shortcut —
philosophically aligned with a tiny kernel, and it lets the single-core correctness
model you already have (the interrupt-flag discipline in
[scheduling.md](scheduling.md)) stay largely intact: one lock around kernel entry
instead of rethinking every critical section.
instead of rethinking every critical section. **Done** — see
`src/kernel/sync.zig`.
4. **Later, if contention bites,** evolve toward **per-core run queues + explicit
affinity** (the Fiasco.OC direction) — also the more real-time-predictable model.
5. **Placement stays a user-space policy** — the kernel runs a thread on the core it's
@@ -133,6 +139,60 @@ Whatever the top goal, the *sequence* is the same and seL4 validates starting si
Big-lock-first → per-core-later. The affinity/MCS depth is only worth it if real-time
turns out to be the actual goal.
## Implementation status
The "wake + schedule" build (real parallel task execution) is going in as a sequence
of green checkpoints — each step keeps the single-core test suite passing before the
next lands.
**Done:**
- **Core enumeration** — the MADT parse records every usable Local APIC (with its
`apic_id`, which an AP wake targets); `platform.cpus()` returns the list. See
[discovery.md](discovery.md).
- **The big kernel lock** (`src/kernel/sync.zig`) — one coarse spinlock guarding the
scheduler queues and IPC, always held with local interrupts disabled. It is held
*across* a context switch and released by whichever task resumes (the hand-off
rule); `task_trampoline` releases it for a freshly-spawned task. `scheduler.zig` and
`ipc.zig` run every critical section under it. Uncontended on one core, so behaviour
is identical to the old interrupt-flag model.
- **Per-CPU state** — a `PerCpu` struct (running task, idle task, APIC id) per core,
its pointer kept in the x86 **GS base** (`IA32_GS_BASE`; no `swapgs`, since there's
no user mode yet). The old global `current` is now `thisCpu().current`. The ready
queues stay **global** under the lock — work-conserving, so any idle core will pull
the highest-priority ready task; per-core queues are a later optimisation.
- **AP wake to long mode** — `arch.startSecondary` drives INIT–SIPI–SIPI (via the
LAPIC ICR) to wake each parked core one at a time. A woken core starts in 16-bit
real mode at a low page and runs the [trampoline](../src/kernel/arch/x86_64/trampoline.s)
up through protected mode into 64-bit long mode, then lands in `smp.zig:apEntry`,
publishes its per-CPU pointer, and reports in. Verified in QEMU with `-smp 4`:
all four cores report `online`.
The trampoline earns its complexity from three hardware facts:
- a STARTUP IPI vectors a core to physical `vector << 12` (a *byte* vector), so the
trampoline must live **below 1 MiB** — the kernel reserves that page from the frame
allocator at boot, before paging/heap draw down the scarce low frames;
- the blanket RAM identity map is **NX** (W^X), but the AP fetches the trampoline
from it under paging, so that one page is made executable for bring-up;
- the blob is copied to a page whose address isn't known at link time, so it is
**position-independent**: it derives its own base from `CS` and, crucially,
addresses data *segment-relative in real mode* (where the segment base already
supplies the page base) but *base-register-relative in protected/long mode* (flat
segments, base 0). Getting that distinction wrong was the first bug found.
**Next:**
- **Per-core descriptor tables + scheduler entry** — each AP needs its own TSS (its
own IST/`rsp0` stack) and to load the kernel GDT/IDT, enable its LAPIC timer, and
enter the scheduler run loop under the big lock. (The 3a checkpoint deliberately
parks the APs on the trampoline's tables with interrupts off; loading the shared
kernel GDT on an AP faulted, and the per-core-TSS work is where that's resolved.)
- **A parallelism test** — a case where N cores drive N counters at once, proving work
runs truly in parallel rather than just that the APs booted.
- **IPIs** (deferred) — cross-core wake/preempt. Not needed for correctness: an idle
core wakes on its own timer tick and pulls ready work then; IPIs only cut that
latency from ≤1 ms to near-instant.
## Further reading
**Microkernel SMP & scheduling**