schedule tasks across all cores
Per-core GDT/TSS and AP scheduler entry; fix AP SSE + single_threaded.
This commit is contained in:
+37
-23
@@ -16,13 +16,14 @@ danos specifics):
|
||||
- **Real-time** — whether timing is *predictable*. Comes from bounded operations
|
||||
(our O(1) scheduler), not from core count.
|
||||
|
||||
danos is mid-transition. The firmware starts only the **bootstrap processor (BSP)**;
|
||||
the other cores (**application processors**, APs) sit parked until the kernel wakes
|
||||
them. As of the SMP work in progress (see [Implementation status](#implementation-status)
|
||||
below), danos now *does* wake the APs — each climbs to 64-bit long mode and reports
|
||||
in — and the shared kernel state (scheduler queues, IPC) is already serialised behind
|
||||
a big kernel lock. What's not done yet is letting the woken APs actually run tasks;
|
||||
`current` is per-CPU but the run loop is still BSP-only.
|
||||
danos now runs on multiple cores. The firmware starts only the **bootstrap processor
|
||||
(BSP)**; the kernel wakes the other cores (**application processors**, APs) with
|
||||
INIT–SIPI–SIPI, brings each up into 64-bit long mode with its own descriptor tables,
|
||||
LAPIC and timer, and drops it into the scheduler. Tasks run **genuinely in parallel** —
|
||||
the `smp` self-test confirms worker tasks executing on all four cores at once under
|
||||
QEMU `-smp 4`. Shared kernel state (scheduler queues, IPC) is serialised behind a big
|
||||
kernel lock. What's left is refinement, not first-light: per-core run queues, IPIs,
|
||||
and thread-to-core affinity (see [Implementation status](#implementation-status)).
|
||||
|
||||
## The common microkernel instinct: don't share kernel state
|
||||
|
||||
@@ -122,9 +123,9 @@ Whatever the top goal, the *sequence* is the same and seL4 validates starting si
|
||||
and `platform.cpus()` returns the list (see [discovery.md](discovery.md)). The boot
|
||||
log reports the count; the ARM (device-tree) path still needs it.
|
||||
2. **Wake the APs** — INIT–SIPI–SIPI on x86; PSCI/spin-tables on ARM. Each core brings
|
||||
up its own tables, timer, and idle task. **In progress on x86** — the cores reach
|
||||
long mode and park; the per-core tables/timer/scheduler entry is the next step
|
||||
([status](#implementation-status)).
|
||||
up its own tables, timer, and idle task. **Done on x86** — cores climb to long mode,
|
||||
set up their own GDT/TSS, and enter the scheduler; tasks run in parallel across all
|
||||
cores ([status](#implementation-status)).
|
||||
3. **Start with a big kernel lock.** It's a legitimate first design, not a shortcut —
|
||||
philosophically aligned with a tiny kernel, and it lets the single-core correctness
|
||||
model you already have (the interrupt-flag discipline in
|
||||
@@ -168,7 +169,7 @@ next lands.
|
||||
publishes its per-CPU pointer, and reports in. Verified in QEMU with `-smp 4`:
|
||||
all four cores report `online`.
|
||||
|
||||
The trampoline earns its complexity from three hardware facts:
|
||||
The trampoline earns its complexity from four hardware facts:
|
||||
- a STARTUP IPI vectors a core to physical `vector << 12` (a *byte* vector), so the
|
||||
trampoline must live **below 1 MiB** — the kernel reserves that page from the frame
|
||||
allocator at boot, before paging/heap draw down the scarce low frames;
|
||||
@@ -178,20 +179,33 @@ next lands.
|
||||
**position-independent**: it derives its own base from `CS` and, crucially,
|
||||
addresses data *segment-relative in real mode* (where the segment base already
|
||||
supplies the page base) but *base-register-relative in protected/long mode* (flat
|
||||
segments, base 0). Getting that distinction wrong was the first bug found.
|
||||
segments, base 0). Getting that distinction wrong was the first bug found;
|
||||
- an AP starts with a bare `CR0`/`CR4`, but the kernel is built **with SSE** (the
|
||||
x86_64 baseline) and the compiler emits SSE for things as ordinary as a struct
|
||||
copy — so the trampoline must set `CR4.OSFXSR`/`OSXMMEXCPT` and fix `CR0.EM`/`MP`,
|
||||
or the first SSE instruction on the AP `#UD`s. The BSP inherited those bits from
|
||||
UEFI; the AP has to set them itself. This was the second bug — it masqueraded as a
|
||||
fault in `lgdt` (the first kernel code after entry that the compiler vectorised).
|
||||
|
||||
**Next:**
|
||||
- **Per-core tables + scheduler entry** — each AP loads **its own GDT** (with its own
|
||||
TSS descriptor) and **its own TSS** (its own IST/`rsp0` stack), loads the shared
|
||||
IDT, enables its LAPIC and timer, then calls the generic `secondaryMain`: it turns
|
||||
its bring-up context into the core's idle task (as task 0 is for the BSP), marks the
|
||||
core online, and enters the run loop. With interrupts on, each core's own timer tick
|
||||
preempts its idle context into whatever the global ready queue offers — so all cores
|
||||
pull real work in parallel. The `smp` test spawns CPU-bound workers and confirms they
|
||||
execute on all four cores at once.
|
||||
- **`single_threaded` off** — the kernel was built `single_threaded = true`, which
|
||||
compiles `std.atomic` down to plain non-atomic ops. Harmless on one core, but it
|
||||
quietly breaks the big kernel lock across cores; it's now `false`.
|
||||
|
||||
- **Per-core descriptor tables + scheduler entry** — each AP needs its own TSS (its
|
||||
own IST/`rsp0` stack) and to load the kernel GDT/IDT, enable its LAPIC timer, and
|
||||
enter the scheduler run loop under the big lock. (The 3a checkpoint deliberately
|
||||
parks the APs on the trampoline's tables with interrupts off; loading the shared
|
||||
kernel GDT on an AP faulted, and the per-core-TSS work is where that's resolved.)
|
||||
- **A parallelism test** — a case where N cores drive N counters at once, proving work
|
||||
runs truly in parallel rather than just that the APs booted.
|
||||
- **IPIs** (deferred) — cross-core wake/preempt. Not needed for correctness: an idle
|
||||
core wakes on its own timer tick and pulls ready work then; IPIs only cut that
|
||||
latency from ≤1 ms to near-instant.
|
||||
**Next (refinement, not first-light):**
|
||||
|
||||
- **IPIs** — cross-core wake/preempt. Not needed for correctness: an idle core wakes
|
||||
on its own timer tick and pulls ready work then; IPIs only cut that latency from
|
||||
≤1 ms to near-instant.
|
||||
- **Per-core run queues + thread affinity** — the Fiasco.OC direction, if the single
|
||||
global queue's lock contention ever bites (and the more real-time-predictable model).
|
||||
|
||||
## Further reading
|
||||
|
||||
|
||||
Reference in New Issue
Block a user