schedule tasks across all cores

Per-core GDT/TSS and AP scheduler entry; fix AP SSE + single_threaded.
This commit is contained in:
Daniel Samson
2026-07-08 12:35:30 +01:00
parent ed7f542006
commit 43afe6bf2e
11 changed files with 239 additions and 81 deletions
+37 -23
View File
@@ -16,13 +16,14 @@ danos specifics):
- **Real-time** — whether timing is *predictable*. Comes from bounded operations
(our O(1) scheduler), not from core count.
danos is mid-transition. The firmware starts only the **bootstrap processor (BSP)**;
the other cores (**application processors**, APs) sit parked until the kernel wakes
them. As of the SMP work in progress (see [Implementation status](#implementation-status)
below), danos now *does* wake the APs — each climbs to 64-bit long mode and reports
in — and the shared kernel state (scheduler queues, IPC) is already serialised behind
a big kernel lock. What's not done yet is letting the woken APs actually run tasks;
`current` is per-CPU but the run loop is still BSP-only.
danos now runs on multiple cores. The firmware starts only the **bootstrap processor
(BSP)**; the kernel wakes the other cores (**application processors**, APs) with
INIT–SIPI–SIPI, brings each up into 64-bit long mode with its own descriptor tables,
LAPIC and timer, and drops it into the scheduler. Tasks run **genuinely in parallel** —
the `smp` self-test confirms worker tasks executing on all four cores at once under
QEMU `-smp 4`. Shared kernel state (scheduler queues, IPC) is serialised behind a big
kernel lock. What's left is refinement, not first-light: per-core run queues, IPIs,
and thread-to-core affinity (see [Implementation status](#implementation-status)).
## The common microkernel instinct: don't share kernel state
@@ -122,9 +123,9 @@ Whatever the top goal, the *sequence* is the same and seL4 validates starting si
and `platform.cpus()` returns the list (see [discovery.md](discovery.md)). The boot
log reports the count; the ARM (device-tree) path still needs it.
2. **Wake the APs** — INIT–SIPI–SIPI on x86; PSCI/spin-tables on ARM. Each core brings
up its own tables, timer, and idle task. **In progress on x86** — the cores reach
long mode and park; the per-core tables/timer/scheduler entry is the next step
([status](#implementation-status)).
up its own tables, timer, and idle task. **Done on x86** — cores climb to long mode,
set up their own GDT/TSS, and enter the scheduler; tasks run in parallel across all
cores ([status](#implementation-status)).
3. **Start with a big kernel lock.** It's a legitimate first design, not a shortcut —
philosophically aligned with a tiny kernel, and it lets the single-core correctness
model you already have (the interrupt-flag discipline in
@@ -168,7 +169,7 @@ next lands.
publishes its per-CPU pointer, and reports in. Verified in QEMU with `-smp 4`:
all four cores report `online`.
The trampoline earns its complexity from three hardware facts:
The trampoline earns its complexity from four hardware facts:
- a STARTUP IPI vectors a core to physical `vector << 12` (a *byte* vector), so the
trampoline must live **below 1 MiB** — the kernel reserves that page from the frame
allocator at boot, before paging/heap draw down the scarce low frames;
@@ -178,20 +179,33 @@ next lands.
**position-independent**: it derives its own base from `CS` and, crucially,
addresses data *segment-relative in real mode* (where the segment base already
supplies the page base) but *base-register-relative in protected/long mode* (flat
segments, base 0). Getting that distinction wrong was the first bug found.
segments, base 0). Getting that distinction wrong was the first bug found;
- an AP starts with a bare `CR0`/`CR4`, but the kernel is built **with SSE** (the
x86_64 baseline) and the compiler emits SSE for things as ordinary as a struct
copy — so the trampoline must set `CR4.OSFXSR`/`OSXMMEXCPT` and fix `CR0.EM`/`MP`,
or the first SSE instruction on the AP `#UD`s. The BSP inherited those bits from
UEFI; the AP has to set them itself. This was the second bug — it masqueraded as a
fault in `lgdt` (the first kernel code after entry that the compiler vectorised).
**Next:**
- **Per-core tables + scheduler entry** — each AP loads **its own GDT** (with its own
TSS descriptor) and **its own TSS** (its own IST/`rsp0` stack), loads the shared
IDT, enables its LAPIC and timer, then calls the generic `secondaryMain`: it turns
its bring-up context into the core's idle task (as task 0 is for the BSP), marks the
core online, and enters the run loop. With interrupts on, each core's own timer tick
preempts its idle context into whatever the global ready queue offers — so all cores
pull real work in parallel. The `smp` test spawns CPU-bound workers and confirms they
execute on all four cores at once.
- **`single_threaded` off** — the kernel was built `single_threaded = true`, which
compiles `std.atomic` down to plain non-atomic ops. Harmless on one core, but it
quietly breaks the big kernel lock across cores; it's now `false`.
- **Per-core descriptor tables + scheduler entry** — each AP needs its own TSS (its
own IST/`rsp0` stack) and to load the kernel GDT/IDT, enable its LAPIC timer, and
enter the scheduler run loop under the big lock. (The 3a checkpoint deliberately
parks the APs on the trampoline's tables with interrupts off; loading the shared
kernel GDT on an AP faulted, and the per-core-TSS work is where that's resolved.)
- **A parallelism test** — a case where N cores drive N counters at once, proving work
runs truly in parallel rather than just that the APs booted.
- **IPIs** (deferred) — cross-core wake/preempt. Not needed for correctness: an idle
core wakes on its own timer tick and pulls ready work then; IPIs only cut that
latency from ≤1 ms to near-instant.
**Next (refinement, not first-light):**
- **IPIs** — cross-core wake/preempt. Not needed for correctness: an idle core wakes
on its own timer tick and pulls ready work then; IPIs only cut that latency from
≤1 ms to near-instant.
- **Per-core run queues + thread affinity** — the Fiasco.OC direction, if the single
global queue's lock contention ever bites (and the more real-time-predictable model).
## Further reading