wake application processors to long mode
INIT-SIPI-SIPI plus a self-relocating real-mode trampoline.
This commit is contained in:
+66
-6
@@ -16,10 +16,13 @@ danos specifics):
|
||||
- **Real-time** — whether timing is *predictable*. Comes from bounded operations
|
||||
(our O(1) scheduler), not from core count.
|
||||
|
||||
danos is uniprocessor today: one global `current` task, one set of ready queues, one
|
||||
timer. Even on an 8-core CPU, the firmware starts only the **bootstrap processor
|
||||
(BSP)**; the other cores (**application processors**, APs) sit parked until the
|
||||
kernel wakes them, which it doesn't yet.
|
||||
danos is mid-transition. The firmware starts only the **bootstrap processor (BSP)**;
|
||||
the other cores (**application processors**, APs) sit parked until the kernel wakes
|
||||
them. As of the SMP work in progress (see [Implementation status](#implementation-status)
|
||||
below), danos now *does* wake the APs — each climbs to 64-bit long mode and reports
|
||||
in — and the shared kernel state (scheduler queues, IPC) is already serialised behind
|
||||
a big kernel lock. What's not done yet is letting the woken APs actually run tasks;
|
||||
`current` is per-CPU but the run loop is still BSP-only.
|
||||
|
||||
## The common microkernel instinct: don't share kernel state
|
||||
|
||||
@@ -119,12 +122,15 @@ Whatever the top goal, the *sequence* is the same and seL4 validates starting si
|
||||
and `platform.cpus()` returns the list (see [discovery.md](discovery.md)). The boot
|
||||
log reports the count; the ARM (device-tree) path still needs it.
|
||||
2. **Wake the APs** — INIT–SIPI–SIPI on x86; PSCI/spin-tables on ARM. Each core brings
|
||||
up its own tables, timer, and idle task.
|
||||
up its own tables, timer, and idle task. **In progress on x86** — the cores reach
|
||||
long mode and park; the per-core tables/timer/scheduler entry is the next step
|
||||
([status](#implementation-status)).
|
||||
3. **Start with a big kernel lock.** It's a legitimate first design, not a shortcut —
|
||||
philosophically aligned with a tiny kernel, and it lets the single-core correctness
|
||||
model you already have (the interrupt-flag discipline in
|
||||
[scheduling.md](scheduling.md)) stay largely intact: one lock around kernel entry
|
||||
instead of rethinking every critical section.
|
||||
instead of rethinking every critical section. **Done** — see
|
||||
`src/kernel/sync.zig`.
|
||||
4. **Later, if contention bites,** evolve toward **per-core run queues + explicit
|
||||
affinity** (the Fiasco.OC direction) — also the more real-time-predictable model.
|
||||
5. **Placement stays a user-space policy** — the kernel runs a thread on the core it's
|
||||
@@ -133,6 +139,60 @@ Whatever the top goal, the *sequence* is the same and seL4 validates starting si
|
||||
Big-lock-first → per-core-later. The affinity/MCS depth is only worth it if real-time
|
||||
turns out to be the actual goal.
|
||||
|
||||
## Implementation status
|
||||
|
||||
The "wake + schedule" build (real parallel task execution) is going in as a sequence
|
||||
of green checkpoints — each step keeps the single-core test suite passing before the
|
||||
next lands.
|
||||
|
||||
**Done:**
|
||||
|
||||
- **Core enumeration** — the MADT parse records every usable Local APIC (with its
|
||||
`apic_id`, which an AP wake targets); `platform.cpus()` returns the list. See
|
||||
[discovery.md](discovery.md).
|
||||
- **The big kernel lock** (`src/kernel/sync.zig`) — one coarse spinlock guarding the
|
||||
scheduler queues and IPC, always held with local interrupts disabled. It is held
|
||||
*across* a context switch and released by whichever task resumes (the hand-off
|
||||
rule); `task_trampoline` releases it for a freshly-spawned task. `scheduler.zig` and
|
||||
`ipc.zig` run every critical section under it. Uncontended on one core, so behaviour
|
||||
is identical to the old interrupt-flag model.
|
||||
- **Per-CPU state** — a `PerCpu` struct (running task, idle task, APIC id) per core,
|
||||
its pointer kept in the x86 **GS base** (`IA32_GS_BASE`; no `swapgs`, since there's
|
||||
no user mode yet). The old global `current` is now `thisCpu().current`. The ready
|
||||
queues stay **global** under the lock — work-conserving, so any idle core will pull
|
||||
the highest-priority ready task; per-core queues are a later optimisation.
|
||||
- **AP wake to long mode** — `arch.startSecondary` drives INIT–SIPI–SIPI (via the
|
||||
LAPIC ICR) to wake each parked core one at a time. A woken core starts in 16-bit
|
||||
real mode at a low page and runs the [trampoline](../src/kernel/arch/x86_64/trampoline.s)
|
||||
up through protected mode into 64-bit long mode, then lands in `smp.zig:apEntry`,
|
||||
publishes its per-CPU pointer, and reports in. Verified in QEMU with `-smp 4`:
|
||||
all four cores report `online`.
|
||||
|
||||
The trampoline earns its complexity from three hardware facts:
|
||||
- a STARTUP IPI vectors a core to physical `vector << 12` (a *byte* vector), so the
|
||||
trampoline must live **below 1 MiB** — the kernel reserves that page from the frame
|
||||
allocator at boot, before paging/heap draw down the scarce low frames;
|
||||
- the blanket RAM identity map is **NX** (W^X), but the AP fetches the trampoline
|
||||
from it under paging, so that one page is made executable for bring-up;
|
||||
- the blob is copied to a page whose address isn't known at link time, so it is
|
||||
**position-independent**: it derives its own base from `CS` and, crucially,
|
||||
addresses data *segment-relative in real mode* (where the segment base already
|
||||
supplies the page base) but *base-register-relative in protected/long mode* (flat
|
||||
segments, base 0). Getting that distinction wrong was the first bug found.
|
||||
|
||||
**Next:**
|
||||
|
||||
- **Per-core descriptor tables + scheduler entry** — each AP needs its own TSS (its
|
||||
own IST/`rsp0` stack) and to load the kernel GDT/IDT, enable its LAPIC timer, and
|
||||
enter the scheduler run loop under the big lock. (The 3a checkpoint deliberately
|
||||
parks the APs on the trampoline's tables with interrupts off; loading the shared
|
||||
kernel GDT on an AP faulted, and the per-core-TSS work is where that's resolved.)
|
||||
- **A parallelism test** — a case where N cores drive N counters at once, proving work
|
||||
runs truly in parallel rather than just that the APs booted.
|
||||
- **IPIs** (deferred) — cross-core wake/preempt. Not needed for correctness: an idle
|
||||
core wakes on its own timer tick and pulls ready work then; IPIs only cut that
|
||||
latency from ≤1 ms to near-instant.
|
||||
|
||||
## Further reading
|
||||
|
||||
**Microkernel SMP & scheduling**
|
||||
|
||||
Reference in New Issue
Block a user