Compare commits
12
Commits
26d2f5259c
...
26ac97df31
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
26ac97df31 | ||
|
|
37f72a4df8 | ||
|
|
f02259cae0 | ||
|
|
5940864958 | ||
|
|
91f2cfa17b | ||
|
|
debe815a5c | ||
|
|
dba3939a0f | ||
|
|
43afe6bf2e | ||
|
|
ed7f542006 | ||
|
|
36c29d2d6d | ||
|
|
941ab091db | ||
|
|
cf7c6df41c |
@@ -73,6 +73,9 @@ pub fn build(b: *std.Build) void {
|
||||
// CPU-exception stubs — real assembly, since they need cross-symbol
|
||||
// jumps/calls that Zig inline asm can't express (see the file's header).
|
||||
arch_mod.addAssemblyFile(b.path("src/kernel/arch/x86_64/isr.s"));
|
||||
// The AP bring-up trampoline: 16-/32-/64-bit mode-switch code that can't be
|
||||
// inline asm (it runs relocated to a low page, not at its link address).
|
||||
arch_mod.addAssemblyFile(b.path("src/kernel/arch/x86_64/trampoline.s"));
|
||||
|
||||
// Firmware-agnostic device discovery. The generic kernel imports this as
|
||||
// "platform" and asks it to enumerate hardware into a backend-neutral device
|
||||
@@ -111,7 +114,7 @@ pub fn build(b: *std.Build) void {
|
||||
.optimize = optimize,
|
||||
.code_model = .small, // kernel is linked in the low 2 GiB (see image_base)
|
||||
.red_zone = false, // interrupts would corrupt the SysV red zone
|
||||
.single_threaded = true, // no scheduler yet; avoids pulling in TLS/atomics
|
||||
.single_threaded = false, // SMP: the big kernel lock's atomics must be real across cores
|
||||
.sanitize_c = .off, // the UBSan runtime needs f128/SSE support we don't provide
|
||||
.stack_check = false, // stack-probe calls have no runtime to land in
|
||||
.stack_protector = false,
|
||||
|
||||
@@ -67,6 +67,27 @@ exist, which is what a real-time scheduler needs.
|
||||
- **Round-robin within a level.** When a task is descheduled it goes to the *back*
|
||||
of its level's queue, so equal-priority tasks share the CPU fairly.
|
||||
|
||||
## Affinity: pinning a task to a core
|
||||
|
||||
By default a task runs on **any** core — the ready queue above is global, and any
|
||||
idle core pulls the highest-priority task from it (work-conserving; see
|
||||
[smp.md](smp.md)). A task can instead be **pinned** to one core with
|
||||
`spawnOn(entry, priority, cpu)`, giving it an *affinity*: it will only ever run
|
||||
there, never migrating.
|
||||
|
||||
Mechanically, each core has its **own** pinned queue (same 8-level FIFO + bitmap)
|
||||
alongside the global one. A pinned task is enqueued only into its core's pinned
|
||||
queue; selection compares the top of the global queue and the running core's pinned
|
||||
queue and takes the higher priority (still O(1) — two bit-scans and a compare), with
|
||||
a pinned task winning an equal-priority tie so it can't be starved by global work.
|
||||
Because every queue is mutated under the [big kernel lock](smp.md), one core enqueuing
|
||||
into another core's pinned queue is safe.
|
||||
|
||||
This is the *explicit-affinity* model (no surprise migration mid-deadline), which is
|
||||
the more real-time-predictable direction. `spawnOn` refuses to pin to an offline or
|
||||
out-of-range core — it creates the task unpinned instead, so it still runs somewhere
|
||||
rather than stranding in a queue no core services, and returns whether the pin took.
|
||||
|
||||
## Sleeping and the idle task
|
||||
|
||||
A task can **block** — give up the CPU until an event, rather than busy-wait
|
||||
|
||||
+107
-7
@@ -16,10 +16,14 @@ danos specifics):
|
||||
- **Real-time** — whether timing is *predictable*. Comes from bounded operations
|
||||
(our O(1) scheduler), not from core count.
|
||||
|
||||
danos is uniprocessor today: one global `current` task, one set of ready queues, one
|
||||
timer. Even on an 8-core CPU, the firmware starts only the **bootstrap processor
|
||||
(BSP)**; the other cores (**application processors**, APs) sit parked until the
|
||||
kernel wakes them, which it doesn't yet.
|
||||
danos now runs on multiple cores. The firmware starts only the **bootstrap processor
|
||||
(BSP)**; the kernel wakes the other cores (**application processors**, APs) with
|
||||
INIT–SIPI–SIPI, brings each up into 64-bit long mode with its own descriptor tables,
|
||||
LAPIC and timer, and drops it into the scheduler. Tasks run **genuinely in parallel** —
|
||||
the `smp` self-test confirms worker tasks executing on all four cores at once under
|
||||
QEMU `-smp 4`. Shared kernel state (scheduler queues, IPC) is serialised behind a big
|
||||
kernel lock. What's left is refinement, not first-light: per-core run queues, IPIs,
|
||||
and thread-to-core affinity (see [Implementation status](#implementation-status)).
|
||||
|
||||
## The common microkernel instinct: don't share kernel state
|
||||
|
||||
@@ -114,14 +118,20 @@ active reconsideration in favour of resilience — see [vision.md](vision.md).)
|
||||
Whatever the top goal, the *sequence* is the same and seL4 validates starting simple:
|
||||
|
||||
1. **Enumerate cores** — needs [device discovery](discovery.md) (ACPI MADT on x86,
|
||||
device tree on ARM). SMP is a concrete consumer of that work.
|
||||
device tree on ARM). SMP is a concrete consumer of that work. **Done on x86:** the
|
||||
MADT parse records every usable Local APIC — with the `apic_id` an AP wake targets —
|
||||
and `platform.cpus()` returns the list (see [discovery.md](discovery.md)). The boot
|
||||
log reports the count; the ARM (device-tree) path still needs it.
|
||||
2. **Wake the APs** — INIT–SIPI–SIPI on x86; PSCI/spin-tables on ARM. Each core brings
|
||||
up its own tables, timer, and idle task.
|
||||
up its own tables, timer, and idle task. **Done on x86** — cores climb to long mode,
|
||||
set up their own GDT/TSS, and enter the scheduler; tasks run in parallel across all
|
||||
cores ([status](#implementation-status)).
|
||||
3. **Start with a big kernel lock.** It's a legitimate first design, not a shortcut —
|
||||
philosophically aligned with a tiny kernel, and it lets the single-core correctness
|
||||
model you already have (the interrupt-flag discipline in
|
||||
[scheduling.md](scheduling.md)) stay largely intact: one lock around kernel entry
|
||||
instead of rethinking every critical section.
|
||||
instead of rethinking every critical section. **Done** — see
|
||||
`src/kernel/sync.zig`.
|
||||
4. **Later, if contention bites,** evolve toward **per-core run queues + explicit
|
||||
affinity** (the Fiasco.OC direction) — also the more real-time-predictable model.
|
||||
5. **Placement stays a user-space policy** — the kernel runs a thread on the core it's
|
||||
@@ -130,6 +140,96 @@ Whatever the top goal, the *sequence* is the same and seL4 validates starting si
|
||||
Big-lock-first → per-core-later. The affinity/MCS depth is only worth it if real-time
|
||||
turns out to be the actual goal.
|
||||
|
||||
## Implementation status
|
||||
|
||||
The "wake + schedule" build (real parallel task execution) is going in as a sequence
|
||||
of green checkpoints — each step keeps the single-core test suite passing before the
|
||||
next lands.
|
||||
|
||||
**Done:**
|
||||
|
||||
- **Core enumeration** — the MADT parse records every usable Local APIC (with its
|
||||
`apic_id`, which an AP wake targets); `platform.cpus()` returns the list. See
|
||||
[discovery.md](discovery.md).
|
||||
- **The big kernel lock** (`src/kernel/sync.zig`) — one coarse spinlock guarding the
|
||||
scheduler queues and IPC, always held with local interrupts disabled. It is held
|
||||
*across* a context switch and released by whichever task resumes (the hand-off
|
||||
rule); `task_trampoline` releases it for a freshly-spawned task. `scheduler.zig` and
|
||||
`ipc.zig` run every critical section under it. Uncontended on one core, so behaviour
|
||||
is identical to the old interrupt-flag model.
|
||||
- **Per-CPU state** — a `PerCpu` struct (running task, idle task, APIC id) per core,
|
||||
its pointer kept in the x86 **GS base** (`IA32_GS_BASE`; no `swapgs`, since there's
|
||||
no user mode yet). The old global `current` is now `thisCpu().current`. The ready
|
||||
queues stay **global** under the lock — work-conserving, so any idle core will pull
|
||||
the highest-priority ready task; per-core queues are a later optimisation.
|
||||
- **AP wake to long mode** — `arch.startSecondary` drives INIT–SIPI–SIPI (via the
|
||||
LAPIC ICR) to wake each parked core one at a time. A woken core starts in 16-bit
|
||||
real mode at a low page and runs the [trampoline](../src/kernel/arch/x86_64/trampoline.s)
|
||||
up through protected mode into 64-bit long mode, then lands in `smp.zig:apEntry`,
|
||||
publishes its per-CPU pointer, and reports in. Verified in QEMU with `-smp 4`:
|
||||
all four cores report `online`.
|
||||
|
||||
The trampoline earns its complexity from four hardware facts:
|
||||
- a STARTUP IPI vectors a core to physical `vector << 12` (a *byte* vector), so the
|
||||
trampoline must live **below 1 MiB** — the kernel reserves that page from the frame
|
||||
allocator at boot, before paging/heap draw down the scarce low frames;
|
||||
- the blanket RAM identity map is **NX** (W^X), but the AP fetches the trampoline
|
||||
from it under paging, so that one page is made executable for bring-up;
|
||||
- the blob is copied to a page whose address isn't known at link time, so it is
|
||||
**position-independent**: it derives its own base from `CS` and, crucially,
|
||||
addresses data *segment-relative in real mode* (where the segment base already
|
||||
supplies the page base) but *base-register-relative in protected/long mode* (flat
|
||||
segments, base 0). Getting that distinction wrong was the first bug found;
|
||||
- an AP starts with a bare `CR0`/`CR4`, but the kernel is built **with SSE** (the
|
||||
x86_64 baseline) and the compiler emits SSE for things as ordinary as a struct
|
||||
copy — so the trampoline must set `CR4.OSFXSR`/`OSXMMEXCPT` and fix `CR0.EM`/`MP`,
|
||||
or the first SSE instruction on the AP `#UD`s. The BSP inherited those bits from
|
||||
UEFI; the AP has to set them itself. This was the second bug — it masqueraded as a
|
||||
fault in `lgdt` (the first kernel code after entry that the compiler vectorised).
|
||||
|
||||
- **Per-core tables + scheduler entry** — each AP loads **its own GDT** (with its own
|
||||
TSS descriptor) and **its own TSS** (its own IST/`rsp0` stack), loads the shared
|
||||
IDT, enables its LAPIC and timer, then calls the generic `secondaryMain`: it turns
|
||||
its bring-up context into the core's idle task (as task 0 is for the BSP), marks the
|
||||
core online, and enters the run loop. With interrupts on, each core's own timer tick
|
||||
preempts its idle context into whatever the global ready queue offers — so all cores
|
||||
pull real work in parallel. The `smp` test spawns CPU-bound workers and confirms they
|
||||
execute on all four cores at once, and `fault-ap-df` pins a #DF to an AP and checks
|
||||
that core catches it on **its own** IST (a broken per-core TSS would triple-fault) —
|
||||
reported as "core N: …", so a fault is always attributed to the core it happened on,
|
||||
and is contained to that core (the rest of the system keeps running).
|
||||
- **Thread affinity** — `spawnOn(entry, priority, cpu)` pins a task to a core (its own
|
||||
per-core pinned queue, merged with the global queue at selection; see
|
||||
[scheduling.md](scheduling.md#affinity-pinning-a-task-to-a-core)). The `affinity`
|
||||
test confirms a pinned task never migrates. This is the mechanism the fault-on-AP
|
||||
test rides on, and the *explicit-affinity* real-time-predictable model.
|
||||
- **`single_threaded` off** — the kernel was built `single_threaded = true`, which
|
||||
compiles `std.atomic` down to plain non-atomic ops. Harmless on one core, but it
|
||||
quietly breaks the big kernel lock across cores; it's now `false`.
|
||||
- **Re-armable wake + retry** — the trampoline frame is reserved for the system's
|
||||
life, but kept **inert between wakes**: zeroed and non-executable, armed (blob
|
||||
copied in, page made executable) only for the moment a core is actually climbing,
|
||||
then disarmed again. So there's never a dormant executable page, and a core can be
|
||||
(re)woken at any time — `arch.startSecondary` is one self-contained attempt (arm →
|
||||
INIT–SIPI–SIPI → disarm), and its `INIT` resets a wedged core, so retrying just
|
||||
works. Boot retries a non-responding core up to three times; the same primitive is
|
||||
the groundwork a future **power manager** would drive to bring cores up (and,
|
||||
eventually, its counterpart to take them offline — which additionally needs the
|
||||
core's tasks migrated off first).
|
||||
|
||||
**Next (refinement, not first-light):**
|
||||
|
||||
- **IPIs** — cross-core wake/preempt. Not needed for correctness: an idle core wakes
|
||||
on its own timer tick and pulls ready work then; IPIs only cut that latency from
|
||||
≤1 ms to near-instant.
|
||||
- **Per-core run queues** — the Fiasco.OC direction, if the single global queue's lock
|
||||
contention ever bites. (Thread *affinity* already exists — see above; this is the
|
||||
further step of giving each core its own primary run queue for load distribution.)
|
||||
- **Fault recovery** — today a fault halts (only) the faulting core. Turning that into
|
||||
"kill the task, keep the core running" is the [resilience](resilience.md) track (it
|
||||
needs the task's lock/resource state handled), and for taking a core fully offline,
|
||||
its tasks migrated first.
|
||||
|
||||
## Further reading
|
||||
|
||||
**Microkernel SMP & scheduling**
|
||||
|
||||
@@ -91,6 +91,35 @@ pub const PlatformInfo = struct {
|
||||
/// Filled in by `discover`; the arch layer reads it during bring-up.
|
||||
pub var platform_info: PlatformInfo = .{};
|
||||
|
||||
/// One usable logical processor, from a MADT type-0 (Local APIC) record. The
|
||||
/// `apic_id` is the Local APIC ID that SMP bring-up targets to wake this core
|
||||
/// (INIT–SIPI–SIPI); `processor_id` is the ACPI namespace handle. Only processors
|
||||
/// the firmware marks *enabled* are recorded — a disabled one can't be started.
|
||||
pub const Cpu = struct {
|
||||
processor_id: u8,
|
||||
apic_id: u8,
|
||||
/// MADT flags bit 1: usable but firmware-started offline (hot-plug / deferred
|
||||
/// bring-up), as opposed to already available. Informational for now.
|
||||
online_capable: bool,
|
||||
};
|
||||
|
||||
/// The set of usable logical processors the MADT listed — the hardware's degree of
|
||||
/// parallelism. Includes the bootstrap processor danos already runs on; the rest
|
||||
/// are the application processors SMP bring-up would start (see docs/smp.md).
|
||||
pub const CpuInfo = struct {
|
||||
/// A static pool sized well above any danos target (a desktop, two 4-core Pis).
|
||||
/// If the MADT ever lists more, the surplus is dropped and counted in `dropped`
|
||||
/// so the truncation is never silent.
|
||||
cpus: [max_cpus]Cpu = undefined,
|
||||
count: usize = 0,
|
||||
dropped: usize = 0,
|
||||
};
|
||||
|
||||
const max_cpus = 64;
|
||||
|
||||
/// Filled in by `discover` (from the MADT); SMP bring-up reads it to wake the APs.
|
||||
pub var cpu_info: CpuInfo = .{};
|
||||
|
||||
/// Integrity/diagnostics for the AML parse. `consumed == total` means the parser
|
||||
/// walked every byte of the DSDT/SSDTs without desyncing.
|
||||
pub const AmlStats = struct {
|
||||
@@ -438,6 +467,18 @@ fn parseMadt(dt: *DeviceTree, header: *const SystemDescriptorTableHeader) !void
|
||||
var nb: [24]u8 = undefined;
|
||||
const nm = std.fmt.bufPrint(&nb, "cpu{d}", .{la.processor_id}) catch "cpu";
|
||||
_ = try dt.addChild(dt.root, .processor, nm);
|
||||
// Also record it as a schedulable core (with the APIC ID an AP
|
||||
// wake needs, which the device node name doesn't preserve).
|
||||
if (cpu_info.count < cpu_info.cpus.len) {
|
||||
cpu_info.cpus[cpu_info.count] = .{
|
||||
.processor_id = la.processor_id,
|
||||
.apic_id = la.apic_id,
|
||||
.online_capable = la.flags & 2 != 0,
|
||||
};
|
||||
cpu_info.count += 1;
|
||||
} else {
|
||||
cpu_info.dropped += 1;
|
||||
}
|
||||
}
|
||||
},
|
||||
1 => {
|
||||
|
||||
@@ -24,6 +24,7 @@ pub const AmlStats = acpi.AmlStats;
|
||||
pub const PlatformInfo = acpi.PlatformInfo;
|
||||
pub const RegAccess = acpi.RegAccess;
|
||||
pub const IsoEntry = acpi.IsoEntry;
|
||||
pub const Cpu = acpi.Cpu;
|
||||
|
||||
/// The register map + sleep types discovery extracted, for logging/diagnostics.
|
||||
pub fn powerInfo() PowerInfo {
|
||||
@@ -41,6 +42,22 @@ pub fn amlStats() AmlStats {
|
||||
return acpi.aml_stats;
|
||||
}
|
||||
|
||||
/// The usable logical processors discovered during enumeration — one entry per
|
||||
/// core danos may schedule on, each carrying the Local APIC ID an SMP wake targets.
|
||||
/// `len` is the hardware's degree of parallelism: how many tasks *could* run at the
|
||||
/// same instant once the application processors are started. Today only the
|
||||
/// bootstrap processor is actually running, so starting the rest is the pending SMP
|
||||
/// step (see docs/smp.md). Borrowed from static storage populated by `discover`.
|
||||
pub fn cpus() []const Cpu {
|
||||
return acpi.cpu_info.cpus[0..acpi.cpu_info.count];
|
||||
}
|
||||
|
||||
/// Non-zero only if enumeration found more processors than the static pool holds
|
||||
/// (the surplus were dropped from `cpus()`); surfaced so the cap is never silent.
|
||||
pub fn cpusDropped() usize {
|
||||
return acpi.cpu_info.dropped;
|
||||
}
|
||||
|
||||
/// Enumerate hardware into a fresh device tree. `hal` supplies the hardware
|
||||
/// primitives the backend needs (MMIO mapping for PCIe config space, port I/O for
|
||||
/// ACPI registers); pass the arch implementation. Errors leave nothing to clean up
|
||||
|
||||
@@ -44,11 +44,15 @@ const spurious_vector = 47;
|
||||
// LAPIC register offsets.
|
||||
const reg_spurious = 0x0F0;
|
||||
const reg_eoi = 0x0B0;
|
||||
const reg_icr_low = 0x300; // interrupt command register, low dword (writing it sends)
|
||||
const reg_icr_high = 0x310; // ICR high dword (destination APIC id in bits 24-31)
|
||||
const reg_lvt_timer = 0x320;
|
||||
const reg_timer_initial = 0x380;
|
||||
const reg_timer_current = 0x390;
|
||||
const reg_timer_divide = 0x3E0;
|
||||
|
||||
const icr_delivery_pending = 1 << 12; // ICR low bit 12: a previous IPI is still in flight
|
||||
|
||||
const lvt_masked = 1 << 16;
|
||||
const lvt_periodic = 1 << 17;
|
||||
const timer_divide_16 = 0x3;
|
||||
@@ -120,6 +124,39 @@ pub fn init() void {
|
||||
write(reg_spurious, 0x100 | spurious_vector); // bit 8 = software enable
|
||||
}
|
||||
|
||||
/// Software-enable *this* core's Local APIC — the application-processor counterpart
|
||||
/// of `init`, minus the one-time PIC remap (the BSP already masked it) and minus
|
||||
/// calibration (the timer rate is a shared hardware constant, measured once). Each
|
||||
/// core has its own LAPIC at the same MMIO address, so no per-core base is needed.
|
||||
pub fn initSecondary() void {
|
||||
const msr = io.rdmsr(ia32_apic_base_msr);
|
||||
io.wrmsr(ia32_apic_base_msr, msr | (1 << 11)); // global enable
|
||||
write(reg_spurious, 0x100 | spurious_vector); // software enable
|
||||
}
|
||||
|
||||
// --- application-processor wakeup (INIT–SIPI–SIPI) --------------------------
|
||||
|
||||
/// Send an INIT IPI to the core with Local APIC id `apic_id` — the first step of
|
||||
/// the wake sequence. Blocks until the LAPIC reports the IPI was delivered.
|
||||
pub fn sendInit(apic_id: u32) void {
|
||||
write(reg_icr_high, apic_id << 24);
|
||||
write(reg_icr_low, 0x4500); // INIT, physical destination, assert, edge-triggered
|
||||
waitIcrIdle();
|
||||
}
|
||||
|
||||
/// Send a STARTUP IPI (SIPI) telling the target core to begin executing at physical
|
||||
/// address `vector << 12` (in real mode). Per the Intel bring-up protocol this is
|
||||
/// sent twice after the INIT; both calls block until delivery completes.
|
||||
pub fn sendStartup(apic_id: u32, vector: u8) void {
|
||||
write(reg_icr_high, apic_id << 24);
|
||||
write(reg_icr_low, 0x4600 | @as(u32, vector)); // STARTUP with the page vector
|
||||
waitIcrIdle();
|
||||
}
|
||||
|
||||
fn waitIcrIdle() void {
|
||||
while (read(reg_icr_low) & icr_delivery_pending != 0) {}
|
||||
}
|
||||
|
||||
/// The calibration window: we time everything against a 10 ms reference interval.
|
||||
const calib_ms = 10;
|
||||
|
||||
|
||||
@@ -13,6 +13,7 @@ const serial = @import("serial.zig");
|
||||
const apic = @import("apic.zig");
|
||||
const ioapic = @import("ioapic.zig");
|
||||
const io = @import("io.zig");
|
||||
const smp = @import("smp.zig");
|
||||
|
||||
/// The saved register/trap frame passed to a fault handler.
|
||||
pub const CpuState = idt.CpuState;
|
||||
@@ -81,6 +82,64 @@ pub fn readCr3() u64 {
|
||||
);
|
||||
}
|
||||
|
||||
/// IA32_GS_BASE: the hidden base of the GS segment. We repurpose it as the per-CPU
|
||||
/// data pointer (there's no user mode yet, so no `swapgs` dance — GS base is always
|
||||
/// the running core's per-CPU block). Set once per core during bring-up, after the
|
||||
/// GDT is loaded (loading a GS *selector* would otherwise clobber this base).
|
||||
const ia32_gs_base = 0xC000_0101;
|
||||
|
||||
/// Publish this core's per-CPU data pointer so `cpuLocal` can retrieve it. Each
|
||||
/// core calls this once, after its GDT is in place.
|
||||
pub fn setCpuLocal(ptr: usize) void {
|
||||
io.wrmsr(ia32_gs_base, ptr);
|
||||
}
|
||||
|
||||
/// This core's per-CPU data pointer (the value `setCpuLocal` stored). Reads the GS
|
||||
/// base MSR — a per-core register, so each core sees its own without any locking.
|
||||
pub fn cpuLocal() usize {
|
||||
return io.rdmsr(ia32_gs_base);
|
||||
}
|
||||
|
||||
// --- SMP: application-processor bring-up ----------------------------------
|
||||
|
||||
/// Record the low (<1 MiB) frame reserved for the AP trampoline. Run once at boot.
|
||||
/// The frame stays inert (zeroed, non-executable) between wakes and is armed only
|
||||
/// while a core is climbing — so a core can be (re)woken at any time (retry, or a
|
||||
/// future power manager) without leaving an executable page resident. See smp.zig.
|
||||
pub fn setTrampolinePage(phys: u64) void {
|
||||
smp.setTrampolinePage(phys);
|
||||
}
|
||||
|
||||
/// Wake the core with Local APIC id `apic_id` as dense CPU `index`, giving it
|
||||
/// `stack_top` and its per-CPU pointer `percpu`; it adopts the current (kernel) page
|
||||
/// tables. Returns false if it doesn't come online within the timeout. Blocks until
|
||||
/// the core reports in.
|
||||
pub fn startSecondary(apic_id: u32, stack_top: usize, percpu: usize, index: usize) bool {
|
||||
return smp.startAp(apic_id, stack_top, percpu, index, readCr3());
|
||||
}
|
||||
|
||||
/// Register the generic entry a woken AP jumps to once its arch state is up (its own
|
||||
/// descriptor tables, LAPIC, and timer). The kernel passes its scheduler entry here.
|
||||
pub fn setSecondaryEntry(entry: *const fn () callconv(.c) noreturn) void {
|
||||
smp.setSecondaryEntry(entry);
|
||||
}
|
||||
|
||||
/// Test hook: force the next `n` AP wake attempts to fail, so the retry path can be
|
||||
/// exercised deterministically (see the smp-retry test). No effect when `n` is 0.
|
||||
pub fn testFailNextWakes(n: u32) void {
|
||||
smp.testFailNextWakes(n);
|
||||
}
|
||||
|
||||
/// The reserved AP-trampoline frame (0 if none). For tests that check it's inert.
|
||||
pub fn trampolinePage() u64 {
|
||||
return smp.trampolinePage();
|
||||
}
|
||||
|
||||
/// Whether the page at `virt` is currently mapped executable (present, NX clear).
|
||||
pub fn pageExecutable(virt: u64) bool {
|
||||
return paging.isExecutable(virt);
|
||||
}
|
||||
|
||||
/// Kernel tick rate: 1000 Hz (1 ms), the scheduler's time quantum.
|
||||
pub const timer_hz = 1000;
|
||||
|
||||
@@ -282,3 +341,11 @@ pub fn readCr2() u64 {
|
||||
pub fn halt() noreturn {
|
||||
while (true) asm volatile ("hlt");
|
||||
}
|
||||
|
||||
/// Spin-wait hint (`pause`). Emitted in the body of a spinlock's busy-wait: it
|
||||
/// relaxes the core while it polls a contended lock — yielding pipeline resources
|
||||
/// to a hyperthread sibling and easing the cache-coherency traffic on the lock
|
||||
/// line. Purely a performance/power hint; correct to omit, but kinder on the bus.
|
||||
pub fn cpuRelax() void {
|
||||
asm volatile ("pause");
|
||||
}
|
||||
|
||||
@@ -3,18 +3,25 @@
|
||||
//! reference a code selector — so we install our own flat GDT with known
|
||||
//! selectors (0x08 kernel code, 0x10 kernel data) rather than trusting whatever
|
||||
//! the firmware left in place.
|
||||
//!
|
||||
//! The code/data descriptors are identical on every core, but the **TSS descriptor
|
||||
//! is per-core** (each core needs its own TSS — its own interrupt/fault stacks; see
|
||||
//! tss.zig). Two cores can't share one TSS descriptor slot, so each core gets its
|
||||
//! own copy of the table with its own TSS descriptor. Slot 0 is the BSP.
|
||||
|
||||
/// Selectors into the table below (index * 8).
|
||||
/// Selectors into the table (index * 8). Same on every core's GDT.
|
||||
pub const kernel_code = 0x08;
|
||||
pub const kernel_data = 0x10;
|
||||
pub const tss_selector = 0x18;
|
||||
|
||||
/// Flat 64-bit descriptors. Base/limit are ignored in long mode; what matters is
|
||||
/// the access byte and, for code, the long-mode (L) flag.
|
||||
const max_cpus = 64; // matches the scheduler / discovery pool
|
||||
const entries = 5; // null, code, data, TSS-low, TSS-high
|
||||
|
||||
/// The shared descriptors (slots 0-2); slots 3-4 hold this core's TSS descriptor,
|
||||
/// filled in per core by `setTssFor`.
|
||||
/// code: present, ring 0, executable, readable, L=1 -> 0x00AF9A00_0000FFFF
|
||||
/// data: present, ring 0, writable -> 0x00CF9200_0000FFFF
|
||||
/// The last two slots hold one 16-byte TSS descriptor, filled in by setTss.
|
||||
var table = [_]u64{
|
||||
const template = [entries]u64{
|
||||
0, // null descriptor (required)
|
||||
0x00AF9A000000FFFF, // kernel code (0x08)
|
||||
0x00CF92000000FFFF, // kernel data (0x10)
|
||||
@@ -22,16 +29,20 @@ var table = [_]u64{
|
||||
0, // TSS descriptor high
|
||||
};
|
||||
|
||||
/// Fill the 64-bit TSS system descriptor (two GDT slots) so the task register can
|
||||
/// point at our TSS. Type 0x89 = present, ring 0, available 64-bit TSS.
|
||||
pub fn setTss(base: u64, limit: u64) void {
|
||||
table[3] = (limit & 0xFFFF) |
|
||||
/// One GDT per core (each a copy of the template, differing only in its TSS slot).
|
||||
var gdts = [_][entries]u64{template} ** max_cpus;
|
||||
|
||||
/// Fill core `cpu`'s 64-bit TSS system descriptor (two GDT slots) so its task
|
||||
/// register can point at its own TSS. Type 0x89 = present, ring 0, available 64-bit
|
||||
/// TSS. Write it into that core's GDT before it loads the TSS selector.
|
||||
pub fn setTssFor(cpu: usize, base: u64, limit: u64) void {
|
||||
gdts[cpu][3] = (limit & 0xFFFF) |
|
||||
((base & 0xFFFF) << 16) |
|
||||
(((base >> 16) & 0xFF) << 32) |
|
||||
(@as(u64, 0x89) << 40) |
|
||||
(((limit >> 16) & 0xF) << 48) |
|
||||
(((base >> 24) & 0xFF) << 56);
|
||||
table[4] = (base >> 32) & 0xFFFFFFFF;
|
||||
gdts[cpu][4] = (base >> 32) & 0xFFFFFFFF;
|
||||
}
|
||||
|
||||
/// The operand `lgdt` wants: table byte-length minus one, then its address.
|
||||
@@ -41,14 +52,21 @@ const Descriptor = packed struct {
|
||||
};
|
||||
|
||||
/// Loads the GDT and reloads the segment registers (including CS). Defined in
|
||||
/// isr.s — it uses the selectors 0x08 (code) and 0x10 (data) that match `table`.
|
||||
/// isr.s — it uses the selectors 0x08 (code) and 0x10 (data) that match the table.
|
||||
extern fn gdt_flush(descriptor: *const Descriptor) callconv(.c) void;
|
||||
|
||||
/// Install our GDT and switch onto its segments.
|
||||
pub fn init() void {
|
||||
/// Load core `cpu`'s GDT and switch onto its segments. Note this reloads the segment
|
||||
/// registers, which zeroes the GS base — so a core must publish its per-CPU pointer
|
||||
/// (setCpuLocal) *after* calling this.
|
||||
pub fn loadOnThisCpu(cpu: usize) void {
|
||||
const descriptor = Descriptor{
|
||||
.limit = @sizeOf(@TypeOf(table)) - 1,
|
||||
.base = @intFromPtr(&table),
|
||||
.limit = @sizeOf([entries]u64) - 1,
|
||||
.base = @intFromPtr(&gdts[cpu]),
|
||||
};
|
||||
gdt_flush(&descriptor);
|
||||
}
|
||||
|
||||
/// Install the bootstrap processor's GDT (slot 0) and switch onto its segments.
|
||||
pub fn init() void {
|
||||
loadOnThisCpu(0);
|
||||
}
|
||||
|
||||
@@ -128,6 +128,13 @@ pub fn init() void {
|
||||
// Run the double-fault handler (vector 8) on IST1: a #DF usually means the
|
||||
// current stack is unusable, so it needs a guaranteed-good one. See tss.zig.
|
||||
idt[8].ist = tss.double_fault_ist;
|
||||
loadOnThisCpu();
|
||||
}
|
||||
|
||||
/// Load the (shared, already-populated) IDT on the current core. The gate table is
|
||||
/// read-only after `init`, so every core points its IDTR at the same one. Called by
|
||||
/// the BSP via `init` and by each AP during bring-up.
|
||||
pub fn loadOnThisCpu() void {
|
||||
const descriptor = Descriptor{
|
||||
.limit = @sizeOf(@TypeOf(idt)) - 1,
|
||||
.base = @intFromPtr(&idt),
|
||||
|
||||
@@ -63,9 +63,14 @@ switch_context:
|
||||
ret # return into the new task's saved instruction pointer
|
||||
|
||||
# task_trampoline: the first thing a freshly-spawned task runs. init_task_stack
|
||||
# leaves its entry function in r15. New tasks start with interrupts enabled.
|
||||
# leaves its entry function in r15. A fresh task is switched to with the big kernel
|
||||
# lock held (the hand-off rule in sync.zig) but has no enter/leave frame of its own,
|
||||
# so it releases the lock here before running its body. r15 survives the call (it's
|
||||
# callee-saved). New tasks then start with interrupts enabled.
|
||||
.extern releaseForFreshTask
|
||||
.global task_trampoline
|
||||
task_trampoline:
|
||||
call releaseForFreshTask # drop the kernel lock we inherited across the switch
|
||||
sti
|
||||
call *%r15 # call the task entry (fn() void)
|
||||
1: hlt # if the entry returns, idle (still preemptible)
|
||||
|
||||
@@ -128,6 +128,31 @@ pub fn map(virt: u64, phys: u64, writable_page: bool) void {
|
||||
invalidate(virt);
|
||||
}
|
||||
|
||||
/// Whether `virt` is currently mapped **executable** — present with the NX bit
|
||||
/// clear. Walks the 4-level tables (all danos mappings are 4 KiB, so no huge-page
|
||||
/// case). Returns false if unmapped. Used for W^X checks in tests.
|
||||
pub fn isExecutable(virt: u64) bool {
|
||||
const pml4e = tableAt(kernel_pml4)[(virt >> 39) & 0x1FF];
|
||||
if (pml4e & present == 0) return false;
|
||||
const pdpte = tableAt(pml4e & addr_mask)[(virt >> 30) & 0x1FF];
|
||||
if (pdpte & present == 0) return false;
|
||||
const pde = tableAt(pdpte & addr_mask)[(virt >> 21) & 0x1FF];
|
||||
if (pde & present == 0) return false;
|
||||
const pte = tableAt(pde & addr_mask)[(virt >> 12) & 0x1FF];
|
||||
if (pte & present == 0) return false;
|
||||
return pte & no_execute == 0;
|
||||
}
|
||||
|
||||
/// Make an already-identity-mapped RAM page **executable** (clear its NX bit),
|
||||
/// leaving it present and writable. The blanket RAM mapping is NX for W^X, but the
|
||||
/// application processors fetch the AP trampoline from a low RAM page under paging —
|
||||
/// so that one page must be executable. A deliberate, temporary W^X exception for a
|
||||
/// single bring-up page; the caller frees it once every AP is up.
|
||||
pub fn setExecutable(phys: u64) void {
|
||||
mapPage(kernel_pml4, phys, phys, present | writable); // note: no no_execute
|
||||
invalidate(phys);
|
||||
}
|
||||
|
||||
/// Remove a mapping and flush it from the TLB.
|
||||
pub fn unmap(virt: u64) void {
|
||||
const pml4e = tableAt(kernel_pml4)[(virt >> 39) & 0x1FF];
|
||||
|
||||
@@ -0,0 +1,171 @@
|
||||
//! Application-processor (AP) bring-up: waking the cores the firmware left parked.
|
||||
//!
|
||||
//! The firmware starts only the bootstrap processor (BSP); the others sit idle until
|
||||
//! the kernel wakes them with an INIT–SIPI–SIPI sequence (Intel SDM Vol.3, "MP
|
||||
//! Initialization"). A woken core begins in 16-bit real mode at a low physical page,
|
||||
//! runs the [trampoline](trampoline.s) up into 64-bit long mode, and lands in
|
||||
//! `apEntry` here. This module copies the trampoline into place, patches its
|
||||
//! per-AP parameters, drives the wake IPIs, and waits for each core to report in.
|
||||
//!
|
||||
//! Cores are brought up **one at a time**: a single trampoline page and parameter
|
||||
//! block are reused, so the BSP patches, wakes, and waits for one AP before the
|
||||
//! next. That also lets `apEntry` pick up its dense CPU index from a plain global.
|
||||
//! Once a core has its own descriptor tables, LAPIC, and timer, it calls the generic
|
||||
//! scheduler entry and joins the run loop — mechanism here, policy there.
|
||||
|
||||
const io = @import("io.zig");
|
||||
const gdt = @import("gdt.zig");
|
||||
const tss = @import("tss.zig");
|
||||
const idt = @import("idt.zig");
|
||||
const apic = @import("apic.zig");
|
||||
const paging = @import("paging.zig");
|
||||
|
||||
/// IA32_GS_BASE — the per-CPU data pointer (see cpu.zig; kept in sync here so the AP
|
||||
/// path doesn't depend on cpu.zig and risk an import cycle).
|
||||
const ia32_gs_base = 0xC000_0101;
|
||||
const page_size = 0x1000;
|
||||
|
||||
/// Physical address of the low (<1 MiB) frame reserved for the trampoline. Held for
|
||||
/// the life of the system so any core can be (re)woken on demand — a retry, or a
|
||||
/// future power manager bringing a core back online. The frame is kept **inert**
|
||||
/// between wakes (zeroed and non-executable) and only armed for the brief moment a
|
||||
/// core is actually climbing. Its low 20 bits are zero, so `phys >> 12` is the SIPI
|
||||
/// vector.
|
||||
var tramp_phys: u64 = 0;
|
||||
|
||||
/// Set to 1 by a freshly-woken AP once it reaches `apEntry` and finishes its own
|
||||
/// bring-up. The BSP clears it before each wake and polls it afterwards — a simple
|
||||
/// one-at-a-time handshake (only one AP is being started at any moment).
|
||||
var ap_alive: u32 = 0;
|
||||
|
||||
/// The dense CPU index of the AP currently being started. Set by the BSP before the
|
||||
/// wake, read by `apEntry` (safe because bring-up is strictly one core at a time).
|
||||
var boot_index: usize = 0;
|
||||
|
||||
/// The generic scheduler entry a woken core jumps to once its arch state is up. Set
|
||||
/// by the kernel via `setSecondaryEntry`; never returns.
|
||||
var secondary_entry: ?*const fn () callconv(.c) noreturn = null;
|
||||
|
||||
/// Register the generic entry an AP calls once its per-CPU tables/LAPIC/timer are up.
|
||||
pub fn setSecondaryEntry(entry: *const fn () callconv(.c) noreturn) void {
|
||||
secondary_entry = entry;
|
||||
}
|
||||
|
||||
/// Test hook: force the next `n` wake attempts to fail (skipping the actual
|
||||
/// INIT-SIPI-SIPI), so the retry path can be exercised deterministically. Zero in
|
||||
/// normal operation — the smp-retry test arms it via `arch.testFailNextWakes`.
|
||||
var fail_next_wakes: u32 = 0;
|
||||
pub fn testFailNextWakes(n: u32) void {
|
||||
fail_next_wakes = n;
|
||||
}
|
||||
|
||||
/// Record the reserved low frame the trampoline uses. Call once at boot. The frame
|
||||
/// starts inert (identity-mapped RW+NX like all RAM); each wake arms it and disarms
|
||||
/// it again, so it's only ever executable while a core is climbing.
|
||||
pub fn setTrampolinePage(phys: u64) void {
|
||||
tramp_phys = phys;
|
||||
}
|
||||
|
||||
/// The reserved trampoline frame (0 if SMP bring-up never ran). Exposed so a test
|
||||
/// can verify it's inert — zeroed and non-executable — when dormant.
|
||||
pub fn trampolinePage() u64 {
|
||||
return tramp_phys;
|
||||
}
|
||||
|
||||
/// Arm the trampoline for a wake: make its page executable (W^X exception for the
|
||||
/// duration of the climb) and copy the blob in.
|
||||
fn arm() void {
|
||||
paging.setExecutable(tramp_phys);
|
||||
const start = @extern([*]const u8, .{ .name = "ap_trampoline_start" });
|
||||
const end = @extern([*]const u8, .{ .name = "ap_trampoline_end" });
|
||||
const len = @intFromPtr(end) - @intFromPtr(start);
|
||||
const dst: [*]u8 = @ptrFromInt(tramp_phys);
|
||||
@memcpy(dst[0..len], start[0..len]);
|
||||
}
|
||||
|
||||
/// Disarm after a wake: wipe the page and restore it to inert RW+NX, so no
|
||||
/// executable code (nor any stale bytes) lingers between wakes. Safe to run once the
|
||||
/// woken core has reported in — it's long past the trampoline by then, in the kernel
|
||||
/// image; a core that never answered is dead and can't be mid-climb.
|
||||
fn disarm() void {
|
||||
const dst: [*]u8 = @ptrFromInt(tramp_phys);
|
||||
@memset(dst[0..page_size], 0);
|
||||
paging.map(tramp_phys, tramp_phys, true); // RW + NX, like every other RAM frame
|
||||
}
|
||||
|
||||
/// Address of a patchable trampoline parameter, by symbol name: the copied blob's
|
||||
/// base plus the field's offset within it (a same-section symbol difference). The
|
||||
/// pointer is `align(1)` — the fields aren't 8-aligned within the blob, and x86
|
||||
/// tolerates unaligned stores, so we don't force layout constraints on the asm.
|
||||
fn param(comptime name: []const u8) *align(1) volatile u64 {
|
||||
const start = @intFromPtr(@extern([*]const u8, .{ .name = "ap_trampoline_start" }));
|
||||
const sym = @intFromPtr(@extern([*]const u8, .{ .name = name }));
|
||||
return @ptrFromInt(tramp_phys + (sym - start));
|
||||
}
|
||||
|
||||
/// Wake the core with Local APIC id `apic_id` as dense CPU `index`, hand it
|
||||
/// `stack_top` and its per-CPU pointer `percpu`, and wait for it to come alive. This
|
||||
/// is one self-contained attempt: it arms the trampoline, drives INIT–SIPI–SIPI, and
|
||||
/// disarms again before returning — so it's safe to call repeatedly (a retry, or a
|
||||
/// power manager re-waking a core; the INIT resets a core that was wedged). Returns
|
||||
/// false if the core doesn't report in within the timeout (left parked, no harm to
|
||||
/// the running system). `cr3` is the kernel page tables the AP adopts. Precondition:
|
||||
/// `setTrampolinePage` has run.
|
||||
pub fn startAp(apic_id: u32, stack_top: usize, percpu: usize, index: usize, cr3: u64) bool {
|
||||
arm();
|
||||
defer disarm();
|
||||
|
||||
if (fail_next_wakes > 0) { // test hook: simulate a core missing this attempt
|
||||
fail_next_wakes -= 1;
|
||||
return false;
|
||||
}
|
||||
|
||||
boot_index = index;
|
||||
param("ap_tramp_cr3").* = cr3;
|
||||
param("ap_tramp_stack").* = stack_top;
|
||||
param("ap_tramp_entry").* = @intFromPtr(&apEntry);
|
||||
param("ap_tramp_percpu").* = percpu;
|
||||
|
||||
@atomicStore(u32, &ap_alive, 0, .seq_cst);
|
||||
|
||||
const vector: u8 = @intCast(tramp_phys >> 12);
|
||||
apic.sendInit(apic_id);
|
||||
delayMicros(10_000); // 10 ms INIT settle
|
||||
apic.sendStartup(apic_id, vector);
|
||||
delayMicros(200);
|
||||
apic.sendStartup(apic_id, vector);
|
||||
|
||||
// Wait up to 100 ms for the AP to reach apEntry and set the flag.
|
||||
const deadline = apic.millis() + 100;
|
||||
while (apic.millis() < deadline) {
|
||||
if (@atomicLoad(u32, &ap_alive, .acquire) != 0) return true;
|
||||
asm volatile ("pause");
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
/// Busy-wait `us` microseconds against the calibrated TSC clock (the AP wake happens
|
||||
/// after the timer is up, so the clock is available).
|
||||
fn delayMicros(us: u64) void {
|
||||
const start = apic.micros();
|
||||
while (apic.micros() - start < us) asm volatile ("pause");
|
||||
}
|
||||
|
||||
/// The 64-bit entry every AP lands on, called from the trampoline with its per-CPU
|
||||
/// pointer in RDI. Brings up this core's own descriptor tables, LAPIC and timer,
|
||||
/// signals the BSP, then jumps to the generic scheduler entry. Never returns.
|
||||
fn apEntry(percpu: usize) callconv(.c) noreturn {
|
||||
const cpu = boot_index;
|
||||
gdt.loadOnThisCpu(cpu); // this core's GDT (with its own TSS slot)
|
||||
tss.setupThisCpu(cpu); // this core's TSS + IST stack, loaded into TR
|
||||
idt.loadOnThisCpu(); // the shared IDT
|
||||
io.wrmsr(ia32_gs_base, percpu); // per-CPU pointer — *after* the GDT reload
|
||||
|
||||
apic.initSecondary(); // software-enable this core's LAPIC
|
||||
apic.initTimer(apic.frequencyHz()); // arm its timer (still masked: interrupts off)
|
||||
|
||||
@atomicStore(u32, &ap_alive, 1, .release); // "arch state up" — BSP is polling this
|
||||
|
||||
if (secondary_entry) |enterScheduler| enterScheduler(); // joins the run loop
|
||||
while (true) asm volatile ("hlt"); // (only if no entry was registered)
|
||||
}
|
||||
@@ -0,0 +1,150 @@
|
||||
# AP trampoline: brings a waking application processor from the 16-bit real mode it
|
||||
# starts in (after INIT-SIPI-SIPI) up through protected mode into 64-bit long mode,
|
||||
# then jumps to the Zig AP entry (arch/x86_64/smp.zig:apEntry).
|
||||
#
|
||||
# A STARTUP IPI vectors a core to physical address `vector << 12` in real mode, so
|
||||
# this blob is copied to a low (<1 MiB) page and started there; at entry CS = that
|
||||
# page >> 4 and IP = 0. It is fully **position-independent**: it derives its own
|
||||
# linear base (CS << 4) into EBX and addresses every internal datum as
|
||||
# `(label - ap_trampoline_start)(%ebx)` — a difference of two symbols in the same
|
||||
# section, which the assembler folds to a constant page offset no matter where the
|
||||
# blob was linked or copied to. The BSP patches the parameter block (CR3, stack,
|
||||
# entry, per-CPU pointer) before each wake; see arch/x86_64/smp.zig.
|
||||
#
|
||||
# It lives in .rodata (not .text): it is data to be copied out and executed
|
||||
# elsewhere, never run at its link address, so it must not be a normal code segment.
|
||||
|
||||
.section .rodata
|
||||
.balign 16
|
||||
.code16
|
||||
.global ap_trampoline_start
|
||||
ap_trampoline_start:
|
||||
cli
|
||||
cld
|
||||
|
||||
# Linear base of this page (CS << 4) into EBX; all data is addressed off it.
|
||||
xorl %eax, %eax
|
||||
mov %cs, %ax
|
||||
shll $4, %eax
|
||||
movl %eax, %ebx
|
||||
|
||||
mov %cs, %ax # DS = CS, so we address our data as DS:(label - start):
|
||||
mov %ax, %ds # the segment base (CS<<4) already supplies the page base,
|
||||
# so data operands use the page *offset*, not EBX.
|
||||
|
||||
# Relocate the pointers whose absolute (linear) targets depend on where we were
|
||||
# copied: the GDT base and the two far-jump targets = EBX + their page offsets.
|
||||
# EBX supplies the base for the *value* (via leal); the store address is DS-rel.
|
||||
leal (gdt32 - ap_trampoline_start)(%ebx), %eax
|
||||
movl %eax, gdtr32_base - ap_trampoline_start
|
||||
leal (prot_entry - ap_trampoline_start)(%ebx), %eax
|
||||
movl %eax, jmp32_off - ap_trampoline_start
|
||||
leal (long_entry - ap_trampoline_start)(%ebx), %eax
|
||||
movl %eax, jmp64_off - ap_trampoline_start
|
||||
|
||||
lgdtl gdtr32 - ap_trampoline_start
|
||||
|
||||
movl %cr0, %eax # enter protected mode (CR0.PE)
|
||||
orl $1, %eax
|
||||
movl %eax, %cr0
|
||||
|
||||
ljmpl *(jmp32_ptr - ap_trampoline_start) # -> prot_entry, CS = 0x08
|
||||
|
||||
.code32
|
||||
prot_entry:
|
||||
movw $0x10, %ax # flat 32-bit data segments
|
||||
movw %ax, %ds
|
||||
movw %ax, %es
|
||||
movw %ax, %ss
|
||||
movw %ax, %fs
|
||||
movw %ax, %gs
|
||||
|
||||
# CR4: PAE (required for long mode) + OSFXSR/OSXMMEXCPT. The kernel is built with
|
||||
# SSE (part of the x86_64 baseline), and the compiler emits SSE for things as
|
||||
# ordinary as a struct copy — without OSFXSR those instructions #UD. The BSP got
|
||||
# these bits from UEFI; an AP starts fresh, so we must set them ourselves.
|
||||
movl %cr4, %eax
|
||||
orl $((1 << 5) | (1 << 9) | (1 << 10)), %eax
|
||||
movl %eax, %cr4
|
||||
|
||||
# CR0: clear EM (no x87 emulation) and set MP, so SSE/x87 don't fault.
|
||||
movl %cr0, %eax
|
||||
andl $~(1 << 2), %eax # ~EM
|
||||
orl $(1 << 1), %eax # MP
|
||||
movl %eax, %cr0
|
||||
|
||||
movl (param_cr3 - ap_trampoline_start)(%ebx), %eax # kernel page tables
|
||||
movl %eax, %cr3
|
||||
|
||||
movl $0xC0000080, %ecx # EFER: long mode enable (LME) + NX enable (NXE, since
|
||||
rdmsr # the kernel's PTEs set the NX bit)
|
||||
orl $((1 << 8) | (1 << 11)), %eax
|
||||
wrmsr
|
||||
|
||||
movl %cr0, %eax # paging on (CR0.PG) — now in long mode (compat sub-mode)
|
||||
orl $(1 << 31), %eax
|
||||
movl %eax, %cr0
|
||||
|
||||
ljmpl *(jmp64_ptr - ap_trampoline_start)(%ebx) # -> long_entry, CS = 0x18 (L=1)
|
||||
|
||||
.code64
|
||||
long_entry:
|
||||
movw $0x10, %ax # sane flat data segments
|
||||
movw %ax, %ds
|
||||
movw %ax, %es
|
||||
movw %ax, %ss
|
||||
|
||||
# RBX = EBX (zero-extended) = page base. Load our stack and per-CPU pointer, then
|
||||
# call the Zig entry — which runs from the kernel image and never returns.
|
||||
movq (param_stack - ap_trampoline_start)(%rbx), %rsp
|
||||
movq (param_percpu - ap_trampoline_start)(%rbx), %rdi # SysV arg 0
|
||||
movq (param_entry - ap_trampoline_start)(%rbx), %rax
|
||||
callq *%rax
|
||||
1: hlt # unreachable; guard against a stray return
|
||||
jmp 1b
|
||||
|
||||
# --- data: GDT, far pointers, and the BSP-patched parameter block -----------
|
||||
.balign 8
|
||||
gdt32:
|
||||
.quad 0x0000000000000000 # 0x00 null
|
||||
.quad 0x00CF9A000000FFFF # 0x08 32-bit code (G, D, present, exec/read)
|
||||
.quad 0x00CF92000000FFFF # 0x10 data (valid in 32- and 64-bit)
|
||||
.quad 0x00AF9A000000FFFF # 0x18 64-bit code (L=1)
|
||||
gdt32_end:
|
||||
|
||||
gdtr32:
|
||||
.word gdt32_end - gdt32 - 1
|
||||
gdtr32_base:
|
||||
.long 0 # patched (16-bit code): linear base of gdt32
|
||||
|
||||
jmp32_ptr: # indirect far-jump operand: offset then selector
|
||||
jmp32_off:
|
||||
.long 0 # patched: linear address of prot_entry
|
||||
.word 0x08 # 32-bit code selector
|
||||
|
||||
jmp64_ptr:
|
||||
jmp64_off:
|
||||
.long 0 # patched: linear address of long_entry
|
||||
.word 0x18 # 64-bit code selector
|
||||
|
||||
# The parameter block, filled in by the BSP (smp.zig) before each STARTUP IPI. Global
|
||||
# so the Zig side can locate each field as (symbol - ap_trampoline_start).
|
||||
.global ap_tramp_cr3
|
||||
.global ap_tramp_stack
|
||||
.global ap_tramp_entry
|
||||
.global ap_tramp_percpu
|
||||
param_cr3:
|
||||
ap_tramp_cr3:
|
||||
.quad 0 # kernel PML4 physical address (CR3)
|
||||
param_stack:
|
||||
ap_tramp_stack:
|
||||
.quad 0 # top of this AP's kernel stack
|
||||
param_entry:
|
||||
ap_tramp_entry:
|
||||
.quad 0 # address of apEntry (the Zig AP entry)
|
||||
param_percpu:
|
||||
ap_tramp_percpu:
|
||||
.quad 0 # this AP's per-CPU pointer (goes in GS base)
|
||||
|
||||
.global ap_trampoline_end
|
||||
ap_trampoline_end:
|
||||
@@ -4,6 +4,10 @@
|
||||
//! interrupted stack was. We use IST1 for the double-fault handler, so a fault
|
||||
//! that happens *because* the current stack is unusable still lands on solid
|
||||
//! ground instead of triple-faulting.
|
||||
//!
|
||||
//! Each core needs **its own TSS** (its own IST stack): two cores taking a fault at
|
||||
//! once can't share one fault stack. So the TSS and its IST stack are per-core,
|
||||
//! indexed by CPU number; slot 0 is the BSP.
|
||||
|
||||
const gdt = @import("gdt.zig");
|
||||
|
||||
@@ -30,19 +34,30 @@ const Tss = packed struct {
|
||||
/// The IST slot (1-based, as the IDT gate encodes it) used for critical faults.
|
||||
pub const double_fault_ist = 1;
|
||||
|
||||
var tss: Tss align(16) = .{};
|
||||
const max_cpus = 64; // matches gdt.zig / the scheduler
|
||||
const ist_stack_size = 16 * 1024;
|
||||
|
||||
/// Dedicated stack for IST1. Static so it needs no allocator and is always valid.
|
||||
var ist1_stack: [16 * 1024]u8 align(16) = undefined;
|
||||
/// One TSS per core, and one IST1 stack per core. Static, so they need no allocator
|
||||
/// and are always valid. (max_cpus × 16 KiB of BSS for the IST stacks.)
|
||||
var tss_table = [_]Tss{.{}} ** max_cpus;
|
||||
var ist_stacks: [max_cpus][ist_stack_size]u8 align(16) = undefined;
|
||||
|
||||
/// Loads the task register with the TSS selector. Defined in isr.s.
|
||||
extern fn load_tr(selector: u16) callconv(.c) void;
|
||||
|
||||
/// Point IST1 at its stack, publish the TSS through the GDT, and load it into the
|
||||
/// task register. Requires the GDT to already be loaded (gdt.init first).
|
||||
pub fn init() void {
|
||||
tss.ist1 = @intFromPtr(&ist1_stack) + ist1_stack.len; // stacks grow down
|
||||
tss.iomap_base = @sizeOf(Tss); // == limit: no I/O permission bitmap
|
||||
gdt.setTss(@intFromPtr(&tss), @sizeOf(Tss) - 1);
|
||||
/// Set up core `cpu`'s TSS: point IST1 at that core's stack, install the TSS
|
||||
/// descriptor into that core's GDT, and load it into the task register. Requires the
|
||||
/// core's GDT to already be loaded (gdt.loadOnThisCpu first).
|
||||
pub fn setupThisCpu(cpu: usize) void {
|
||||
const t = &tss_table[cpu];
|
||||
t.* = .{};
|
||||
t.ist1 = @intFromPtr(&ist_stacks[cpu]) + ist_stack_size; // stacks grow down
|
||||
t.iomap_base = @sizeOf(Tss); // == limit: no I/O permission bitmap
|
||||
gdt.setTssFor(cpu, @intFromPtr(t), @sizeOf(Tss) - 1);
|
||||
load_tr(gdt.tss_selector);
|
||||
}
|
||||
|
||||
/// Set up the bootstrap processor's TSS (slot 0). Requires gdt.init first.
|
||||
pub fn init() void {
|
||||
setupThisCpu(0);
|
||||
}
|
||||
|
||||
+5
-5
@@ -11,8 +11,8 @@
|
||||
//! When user mode arrives, the same primitive carries messages across the
|
||||
//! isolation boundary (with the payload copied between address spaces).
|
||||
|
||||
const arch = @import("arch");
|
||||
const sched = @import("scheduler.zig");
|
||||
const sync = @import("sync.zig");
|
||||
|
||||
/// A bounded blocking channel of `capacity` messages of type `T`.
|
||||
pub fn Channel(comptime T: type, comptime capacity: usize) type {
|
||||
@@ -28,7 +28,7 @@ pub fn Channel(comptime T: type, comptime capacity: usize) type {
|
||||
|
||||
/// Send a message, blocking while the channel is full.
|
||||
pub fn send(self: *Self, msg: T) void {
|
||||
const flags = arch.saveInterrupts();
|
||||
const flags = sync.enter();
|
||||
// Recheck the condition in a loop: a wakeup only means "try again"
|
||||
// (another waiter may have taken the slot first).
|
||||
while (self.count == capacity) sched.waitLocked(&self.not_full);
|
||||
@@ -36,18 +36,18 @@ pub fn Channel(comptime T: type, comptime capacity: usize) type {
|
||||
self.tail = (self.tail + 1) % capacity;
|
||||
self.count += 1;
|
||||
sched.wakeLocked(&self.not_empty); // a receiver can now proceed
|
||||
arch.restoreInterrupts(flags);
|
||||
sync.leave(flags);
|
||||
}
|
||||
|
||||
/// Receive a message, blocking while the channel is empty.
|
||||
pub fn recv(self: *Self) T {
|
||||
const flags = arch.saveInterrupts();
|
||||
const flags = sync.enter();
|
||||
while (self.count == 0) sched.waitLocked(&self.not_empty);
|
||||
const msg = self.buffer[self.head];
|
||||
self.head = (self.head + 1) % capacity;
|
||||
self.count -= 1;
|
||||
sched.wakeLocked(&self.not_full); // a sender can now proceed
|
||||
arch.restoreInterrupts(flags);
|
||||
sync.leave(flags);
|
||||
return msg;
|
||||
}
|
||||
};
|
||||
|
||||
+73
-5
@@ -30,6 +30,11 @@ const cp_running = 0x70;
|
||||
const cp_exception = 0xE0;
|
||||
const cp_panic = 0xEE;
|
||||
|
||||
/// Physical address of the low page reserved at boot for the AP trampoline (0 = none
|
||||
/// was available). Claimed right after the frame allocator comes up, before paging
|
||||
/// and the heap consume the scarce sub-1 MiB frames.
|
||||
var ap_trampoline_page: u64 = 0;
|
||||
|
||||
/// Kernel entry point. The bootloader jumps here after `ExitBootServices` with a
|
||||
/// pointer to the handoff data. There is no runtime, no stack unwinding, and no
|
||||
/// caller to return to, so this never returns.
|
||||
@@ -98,6 +103,10 @@ fn kmain(boot_info: *const BootInfo) noreturn {
|
||||
// Bring up the physical frame allocator over that map, and prove it works:
|
||||
// allocate three frames, then hand them back.
|
||||
pmm.init(boot_info.memory_map);
|
||||
// Claim the AP trampoline's low (<1 MiB) page *now*, before paging and the heap
|
||||
// draw down sub-1 MiB frames (the allocator scans upward from frame 0). Held
|
||||
// until SMP bring-up; 0 means none was available (we stay uniprocessor).
|
||||
ap_trampoline_page = pmm.allocBelow(0x100000) orelse 0;
|
||||
const s1 = pmm.stats();
|
||||
log.print("\ndanos: frame allocator online\n", .{});
|
||||
log.print(" free frames: {d} ({d} MiB)\n", .{ s1.free_frames, mib(s1.free_frames) });
|
||||
@@ -200,6 +209,10 @@ fn kmain(boot_info: *const BootInfo) noreturn {
|
||||
log.write(" console UART: none in SPCR -> legacy COM1\n");
|
||||
}
|
||||
log.print(" ioapic : base 0x{x}, {d} inputs (masked); entry0 low 0x{x}\n", .{ ioapic_base, arch.ioapicEntryCount(), arch.ioapicEntryLow(0) });
|
||||
const cores = platform.cpus();
|
||||
log.print(" cpus : {d} usable core(s); 1 running (BSP), {d} AP(s) parked (SMP bring-up pending)\n", .{ cores.len, if (cores.len > 0) cores.len - 1 else 0 });
|
||||
if (platform.cpusDropped() > 0)
|
||||
log.print(" cpus : WARNING {d} core(s) beyond pool cap dropped\n", .{platform.cpusDropped()});
|
||||
} else |err| {
|
||||
log.print("\ndanos: device discovery failed: {s}\n", .{@errorName(err)});
|
||||
}
|
||||
@@ -217,6 +230,10 @@ fn kmain(boot_info: *const BootInfo) noreturn {
|
||||
log.checkpoint(cp_timer);
|
||||
log.print("danos: timer online ({d} Hz tick; LAPIC {d} MHz, TSC {d} MHz; calibrated via {s})\n", .{ arch.timer_hz, arch.lapicHz() / 1_000_000, arch.tscHz() / 1_000_000, arch.timerCalibrationSource() });
|
||||
|
||||
// Wake the other cores (application processors). A no-op on a single-core
|
||||
// machine; on SMP each AP climbs to long mode and reports in (docs/smp.md).
|
||||
bringUpSecondaries();
|
||||
|
||||
// In a test build (`zig build -Dtest-case=<name>`), run that case and stop.
|
||||
// Normal builds fall through to the idle halt.
|
||||
if (build_options.test_case) |case| {
|
||||
@@ -234,6 +251,53 @@ fn kmain(boot_info: *const BootInfo) noreturn {
|
||||
arch.halt();
|
||||
}
|
||||
|
||||
/// Wake the application processors the firmware left parked. Allocates the low
|
||||
/// trampoline page (and makes it executable), then wakes each non-boot core in turn,
|
||||
/// handing it a fresh kernel stack and its per-CPU slot. Cores that don't report in
|
||||
/// are left parked — the running system is unaffected. See docs/smp.md.
|
||||
fn bringUpSecondaries() void {
|
||||
const cores = platform.cpus();
|
||||
if (cores.len <= 1) return;
|
||||
|
||||
// A low (<1 MiB) frame was reserved at boot for the real-mode trampoline (a SIPI
|
||||
// vector addresses it). It's kept for the system's life — armed only during a
|
||||
// wake, inert (zeroed, non-executable) otherwise — so cores can be re-woken later.
|
||||
if (ap_trampoline_page == 0) {
|
||||
log.write("danos: smp: no low page for the AP trampoline; staying uniprocessor\n");
|
||||
return;
|
||||
}
|
||||
arch.setTrampolinePage(ap_trampoline_page);
|
||||
arch.setSecondaryEntry(scheduler.secondaryMain); // where a woken core joins the run loop
|
||||
|
||||
// Test hook: the smp-retry case forces the first wake to fail, so the retry below
|
||||
// must still bring every core online. Inert in a normal build (test_case is null).
|
||||
if (build_options.test_case) |tc| {
|
||||
if (std.mem.eql(u8, tc, "smp-retry")) arch.testFailNextWakes(1);
|
||||
}
|
||||
|
||||
log.print("\ndanos: bringing up {d} application processor(s)\n", .{cores.len - 1});
|
||||
const max_wake_attempts = 3; // a core that misses the first INIT-SIPI-SIPI gets retried
|
||||
for (cores[1..], 1..) |core, index| {
|
||||
const stack = heap.allocator().alloc(u8, 16 * 1024) catch {
|
||||
log.print(" cpu apic_id {d}: no stack; skipped\n", .{core.apic_id});
|
||||
continue;
|
||||
};
|
||||
const stack_top = (@intFromPtr(stack.ptr) + stack.len) & ~@as(usize, 15);
|
||||
const pc = scheduler.prepareSecondary(index, core.apic_id);
|
||||
var attempt: u32 = 1;
|
||||
while (attempt <= max_wake_attempts) : (attempt += 1) {
|
||||
if (arch.startSecondary(core.apic_id, stack_top, @intFromPtr(pc), index)) {
|
||||
pc.online = true;
|
||||
log.print(" cpu apic_id {d}: online (attempt {d})\n", .{ core.apic_id, attempt });
|
||||
break;
|
||||
}
|
||||
if (attempt == max_wake_attempts)
|
||||
log.print(" cpu apic_id {d}: no response after {d} attempts (parked)\n", .{ core.apic_id, max_wake_attempts });
|
||||
}
|
||||
}
|
||||
log.print("danos: {d}/{d} cores online\n", .{ scheduler.onlineCount(), cores.len });
|
||||
}
|
||||
|
||||
/// A user-facing status line: to the diagnostic `log` *and* the on-screen console
|
||||
/// (if a framebuffer is present). The verbose log uses `log.*` directly and never
|
||||
/// touches the framebuffer.
|
||||
@@ -256,21 +320,25 @@ fn kib(frames: u64) u64 {
|
||||
return frames * danos.page_size / (1024);
|
||||
}
|
||||
|
||||
/// Report a CPU exception and halt. There's no fault recovery yet, so any
|
||||
/// exception is terminal — but it reports what and where (to every output sink,
|
||||
/// plus a POST code and a persistent breadcrumb) instead of silently resetting.
|
||||
/// Report a CPU exception and halt **this core**. There's no fault recovery yet, so
|
||||
/// the faulting core is terminal — but the fault is *contained* to it: on an
|
||||
/// application processor only that core stops, and the rest of the system keeps
|
||||
/// running (full recovery — kill the task, keep the core — is the resilience track,
|
||||
/// see docs/resilience.md). The report names the core so an AP fault is attributed,
|
||||
/// and goes to every output sink plus a POST code and a persistent breadcrumb.
|
||||
fn onException(state: *const arch.CpuState) noreturn {
|
||||
log.checkpoint(cp_exception);
|
||||
const core = scheduler.currentCpuIndex();
|
||||
// A fault is user-facing enough to paint on screen too (via statusPrint), on
|
||||
// top of the diagnostic log.
|
||||
statusPrint("\nCPU EXCEPTION: {s} (vector {d})\n", .{ arch.vectorName(state.vector), state.vector });
|
||||
statusPrint("\nCPU EXCEPTION on core {d}: {s} (vector {d})\n", .{ core, arch.vectorName(state.vector), state.vector });
|
||||
statusPrint(" error code : 0x{x}\n", .{state.error_code});
|
||||
statusPrint(" RIP : 0x{x:0>16}\n", .{state.rip});
|
||||
statusPrint(" RSP : 0x{x:0>16}\n", .{state.rsp});
|
||||
if (state.vector == 14) statusPrint(" CR2 (addr) : 0x{x:0>16}\n", .{arch.readCr2()});
|
||||
|
||||
var buf: [128]u8 = undefined;
|
||||
log.recordPanic(std.fmt.bufPrint(&buf, "CPU exception {s} (vector {d}) at RIP 0x{x}", .{ arch.vectorName(state.vector), state.vector, state.rip }) catch "cpu exception");
|
||||
log.recordPanic(std.fmt.bufPrint(&buf, "CPU exception {s} (vector {d}) on core {d} at RIP 0x{x}", .{ arch.vectorName(state.vector), state.vector, core, state.rip }) catch "cpu exception");
|
||||
arch.halt();
|
||||
}
|
||||
|
||||
|
||||
@@ -141,6 +141,24 @@ pub fn alloc() ?u64 {
|
||||
return null; // out of physical memory
|
||||
}
|
||||
|
||||
/// Allocate one free frame whose physical address is below `limit`, or null if
|
||||
/// none is free down there. The AP trampoline needs this: an x86 STARTUP IPI vectors
|
||||
/// a waking core to physical `vector << 12`, and `vector` is a byte — so the
|
||||
/// trampoline must live under 1 MiB. A short linear scan of the low frames; only run
|
||||
/// a handful of times at boot, so it needn't be fast.
|
||||
pub fn allocBelow(limit: u64) ?u64 {
|
||||
const cap = @min(total_frames, @as(usize, @intCast(limit / page_size)));
|
||||
var f: usize = 1; // frame 0 stays reserved as the "none" address
|
||||
while (f < cap) : (f += 1) {
|
||||
if (!isUsed(f)) {
|
||||
setUsed(f);
|
||||
used_frames += 1;
|
||||
return @as(u64, f) * page_size;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/// Return a frame obtained from alloc() to the pool. Bogus or double frees are
|
||||
/// ignored rather than corrupting the count.
|
||||
pub fn free(addr: u64) void {
|
||||
|
||||
+234
-64
@@ -9,10 +9,18 @@
|
||||
//! Switching happens both cooperatively (`yield`) and preemptively (the timer
|
||||
//! calls `tick`). See docs/scheduling.md for the interrupt-flag discipline that
|
||||
//! makes those two paths coexist.
|
||||
//!
|
||||
//! Cross-core safety is the **big kernel lock** (`sync.zig`): every critical
|
||||
//! section here runs under it, and it is held across a context switch and released
|
||||
//! by the task that resumes (see sync.zig's hand-off rule). On a single core the
|
||||
//! lock is never contended, so the behaviour is exactly the old interrupt-flag
|
||||
//! model; it's what lets a second core enter `schedule()` without corrupting the
|
||||
//! shared queues.
|
||||
|
||||
const std = @import("std");
|
||||
const arch = @import("arch");
|
||||
const heap = @import("heap.zig");
|
||||
const sync = @import("sync.zig");
|
||||
|
||||
/// Priority level: 0 (lowest) .. 7 (highest). 8 levels total.
|
||||
pub const Priority = u3;
|
||||
@@ -30,26 +38,70 @@ const Task = struct {
|
||||
rsp: usize = 0, // saved stack pointer, valid while not running
|
||||
stack: []u8 = &.{},
|
||||
wake_at: u64 = 0, // uptime (ms) to wake a sleeping task; 0 = not sleeping
|
||||
affinity: ?u32 = null, // null = runs on any core; else the index of its pinned core
|
||||
next: ?*Task = null, // ready-queue link
|
||||
};
|
||||
|
||||
var tasks = [_]Task{.{}} ** max_tasks;
|
||||
var current: *Task = undefined;
|
||||
var next_id: u32 = 1;
|
||||
|
||||
// Per-priority FIFO ready queues, and a bitmap of which levels are non-empty.
|
||||
/// Per-CPU scheduler state: the task each core is running, its own idle task, and a
|
||||
/// queue of tasks **pinned** to it. One entry per core; the arch layer stashes a
|
||||
/// pointer to the *running* core's entry in the GS base, so `thisCpu()` fetches it
|
||||
/// with a single read and no lock.
|
||||
///
|
||||
/// Most work stays in the **global** ready queue (below), which any idle core pulls
|
||||
/// from — work-conserving. A task given an *affinity* instead goes to that core's
|
||||
/// `pinned_*` queue and is only ever run there (no surprise migration — the more
|
||||
/// real-time-predictable model, docs/smp.md). The two queues are merged at selection
|
||||
/// time. Both are still mutated only under the big kernel lock, so one core enqueuing
|
||||
/// into another core's pinned queue is safe.
|
||||
pub const PerCpu = struct {
|
||||
current: *Task = undefined, // the task running on this core
|
||||
idle: *Task = undefined, // this core's idle task (always ready, lowest priority)
|
||||
apic_id: u32 = 0, // the core's Local APIC id
|
||||
index: u32 = 0, // dense 0-based core index
|
||||
online: bool = false, // has this core finished bring-up?
|
||||
// Tasks pinned to this core (affinity == index), per priority level + bitmap.
|
||||
pinned_head: [num_priorities]?*Task = .{null} ** num_priorities,
|
||||
pinned_tail: [num_priorities]?*Task = .{null} ** num_priorities,
|
||||
pinned_bitmap: u8 = 0,
|
||||
};
|
||||
|
||||
const max_cpus = 64; // matches the discovery pool (src/device/acpi.zig)
|
||||
var cpus = [_]PerCpu{.{}} ** max_cpus;
|
||||
|
||||
/// This core's per-CPU state, via the arch layer's GS-base pointer. Valid only once
|
||||
/// this core has run its scheduler bring-up (BSP in `init`, AP in `secondaryInit`).
|
||||
inline fn thisCpu() *PerCpu {
|
||||
return @ptrFromInt(arch.cpuLocal());
|
||||
}
|
||||
|
||||
/// The task running on this core — the per-CPU replacement for the old global
|
||||
/// `current`. A convenience reader; writes go through `thisCpu().current`.
|
||||
inline fn cur() *Task {
|
||||
return thisCpu().current;
|
||||
}
|
||||
|
||||
// Per-priority FIFO ready queues, and a bitmap of which levels are non-empty. These
|
||||
// are shared across all cores and mutated only under the big kernel lock.
|
||||
var ready_head: [num_priorities]?*Task = .{null} ** num_priorities;
|
||||
var ready_tail: [num_priorities]?*Task = .{null} ** num_priorities;
|
||||
var ready_bitmap: u8 = 0;
|
||||
|
||||
var preemption_enabled = true;
|
||||
|
||||
/// Register the currently-running kernel context as the first task, spawn the
|
||||
/// idle task, and hook the timer for preemption.
|
||||
/// Bring up scheduling on the bootstrap processor: register the currently-running
|
||||
/// kernel context as task 0, publish this core's per-CPU state (via the GS base),
|
||||
/// give the core an idle task, and hook the timer for preemption. Runs once, at
|
||||
/// boot, before interrupts are enabled — so no lock is needed here.
|
||||
pub fn init(boot_priority: Priority) void {
|
||||
const pc = &cpus[0];
|
||||
pc.* = .{ .index = 0, .online = true };
|
||||
arch.setCpuLocal(@intFromPtr(pc));
|
||||
tasks[0] = .{ .id = 0, .state = .running, .priority = boot_priority };
|
||||
current = &tasks[0];
|
||||
spawn(idle, 0); // lowest priority, always runnable — runs when nothing else is
|
||||
pc.current = &tasks[0];
|
||||
pc.idle = create(idle, 0, null); // this core's idle task: always ready, lowest priority
|
||||
arch.setTickHook(tick);
|
||||
}
|
||||
|
||||
@@ -59,40 +111,128 @@ fn idle() void {
|
||||
while (true) asm volatile ("hlt");
|
||||
}
|
||||
|
||||
fn enqueue(t: *Task) void {
|
||||
t.next = null;
|
||||
const p: usize = t.priority;
|
||||
if (ready_tail[p]) |tail| tail.next = t else ready_head[p] = t;
|
||||
ready_tail[p] = t;
|
||||
ready_bitmap |= levelBit(t.priority);
|
||||
/// Reserve and initialise the per-CPU slot for an application processor at dense
|
||||
/// `index` (1-based; 0 is the BSP) with Local APIC id `apic_id`, and return a
|
||||
/// pointer the arch bring-up hands to the core (it publishes it in its GS base).
|
||||
/// Called on the BSP before waking each AP; the AP marks itself `online`.
|
||||
pub fn prepareSecondary(index: usize, apic_id: u32) *PerCpu {
|
||||
const pc = &cpus[index];
|
||||
pc.* = .{ .index = @intCast(index), .apic_id = apic_id, .online = false };
|
||||
return pc;
|
||||
}
|
||||
|
||||
fn dequeueHighest() ?*Task {
|
||||
if (ready_bitmap == 0) return null;
|
||||
const level: Priority = @intCast(num_priorities - 1 - @clz(ready_bitmap));
|
||||
const t = ready_head[level].?;
|
||||
ready_head[level] = t.next;
|
||||
if (ready_head[level] == null) {
|
||||
ready_tail[level] = null;
|
||||
ready_bitmap &= ~levelBit(level);
|
||||
/// Entry for an application processor once the arch layer has set up its per-CPU
|
||||
/// tables, LAPIC, and timer. It turns this bring-up context into the core's idle task
|
||||
/// (as task 0 is for the BSP), marks the core online, and enters the run loop: with
|
||||
/// interrupts enabled the timer preempts this idle context into whatever the global
|
||||
/// ready queue offers, so the core runs real work in parallel with the others. The
|
||||
/// `.c` calling convention lets the arch trampoline path jump here. Never returns.
|
||||
pub fn secondaryMain() callconv(.c) noreturn {
|
||||
const flags = sync.enter();
|
||||
const pc = thisCpu();
|
||||
const t = freeSlot() orelse @panic("sched: task table full (AP idle task)");
|
||||
t.* = .{ .id = next_id, .state = .running, .priority = 0 };
|
||||
next_id += 1;
|
||||
pc.current = t;
|
||||
pc.idle = t;
|
||||
pc.online = true;
|
||||
sync.leave(flags);
|
||||
|
||||
arch.enableInterrupts(); // the timer now preempts this idle context into work
|
||||
while (true) asm volatile ("hlt"); // idle when this core has nothing ready
|
||||
}
|
||||
|
||||
/// Number of cores that have finished bring-up (the BSP plus every online AP).
|
||||
pub fn onlineCount() usize {
|
||||
var n: usize = 0;
|
||||
for (&cpus) |*pc| {
|
||||
if (pc.online) n += 1;
|
||||
}
|
||||
return n;
|
||||
}
|
||||
|
||||
/// Make `t` ready. A pinned task (affinity set) goes to that core's pinned queue;
|
||||
/// everything else goes to the shared global queue.
|
||||
fn enqueue(t: *Task) void {
|
||||
if (t.affinity) |cpu| {
|
||||
const pc = &cpus[cpu];
|
||||
enqueueTo(&pc.pinned_head, &pc.pinned_tail, &pc.pinned_bitmap, t);
|
||||
} else {
|
||||
enqueueTo(&ready_head, &ready_tail, &ready_bitmap, t);
|
||||
}
|
||||
}
|
||||
|
||||
fn enqueueTo(head: *[num_priorities]?*Task, tail: *[num_priorities]?*Task, bitmap: *u8, t: *Task) void {
|
||||
t.next = null;
|
||||
const p: usize = t.priority;
|
||||
if (tail[p]) |tl| tl.next = t else head[p] = t;
|
||||
tail[p] = t;
|
||||
bitmap.* |= @as(u8, 1) << t.priority;
|
||||
}
|
||||
|
||||
/// The highest non-empty priority level in a bitmap, or -1 if empty.
|
||||
fn topLevel(bitmap: u8) i32 {
|
||||
if (bitmap == 0) return -1;
|
||||
return @as(i32, num_priorities - 1) - @as(i32, @clz(bitmap));
|
||||
}
|
||||
|
||||
/// Pick the highest-priority ready task for core `pc`: the better of the global queue
|
||||
/// and this core's pinned queue. Still O(1) (two `clz` and a compare). A pinned task
|
||||
/// wins an equal-priority tie, so it can't be starved by global work at its level.
|
||||
fn dequeueHighest(pc: *PerCpu) ?*Task {
|
||||
const g = topLevel(ready_bitmap);
|
||||
const p = topLevel(pc.pinned_bitmap);
|
||||
if (g < 0 and p < 0) return null;
|
||||
if (p >= g) return dequeueFrom(&pc.pinned_head, &pc.pinned_tail, &pc.pinned_bitmap, @intCast(p));
|
||||
return dequeueFrom(&ready_head, &ready_tail, &ready_bitmap, @intCast(g));
|
||||
}
|
||||
|
||||
fn dequeueFrom(head: *[num_priorities]?*Task, tail: *[num_priorities]?*Task, bitmap: *u8, level: usize) ?*Task {
|
||||
const t = head[level].?;
|
||||
head[level] = t.next;
|
||||
if (head[level] == null) {
|
||||
tail[level] = null;
|
||||
bitmap.* &= ~(@as(u8, 1) << @intCast(level));
|
||||
}
|
||||
t.next = null;
|
||||
return t;
|
||||
}
|
||||
|
||||
fn levelBit(p: Priority) u8 {
|
||||
return @as(u8, 1) << p;
|
||||
/// Create a task that runs `entry` at `priority`, runnable on any core. It becomes
|
||||
/// ready immediately. Takes the kernel lock: it mutates the shared task table and
|
||||
/// ready queues and allocates from the (non-thread-safe) heap, so on SMP it must be
|
||||
/// serialised.
|
||||
pub fn spawn(entry: *const fn () void, priority: Priority) void {
|
||||
const flags = sync.enter();
|
||||
_ = create(entry, priority, null);
|
||||
sync.leave(flags);
|
||||
}
|
||||
|
||||
/// Create a task that runs `entry` at `priority`. It becomes ready immediately.
|
||||
pub fn spawn(entry: *const fn () void, priority: Priority) void {
|
||||
/// Like `spawn`, but **pins** the task to core `cpu` — it will only ever run there.
|
||||
/// Returns true if pinned; false if `cpu` isn't a valid, online core, in which case
|
||||
/// the task is still created but left unpinned (so it runs *somewhere* rather than
|
||||
/// stranding in a queue no core services). Callers that require the pin (e.g. tests)
|
||||
/// should check the result.
|
||||
pub fn spawnOn(entry: *const fn () void, priority: Priority, cpu: u32) bool {
|
||||
const flags = sync.enter();
|
||||
defer sync.leave(flags);
|
||||
const ok = cpu < max_cpus and cpus[cpu].online;
|
||||
_ = create(entry, priority, if (ok) cpu else null);
|
||||
return ok;
|
||||
}
|
||||
|
||||
/// The unlocked task-creation primitive. Caller must hold the kernel lock (or be the
|
||||
/// single-threaded boot path). `affinity` pins the task to a core (null = any).
|
||||
/// Returns the new task so a core can keep a handle to its idle task.
|
||||
fn create(entry: *const fn () void, priority: Priority, affinity: ?u32) *Task {
|
||||
const t = freeSlot() orelse @panic("sched: task table full");
|
||||
const stack = heap.allocator().alloc(u8, stack_size) catch @panic("sched: no memory for task stack");
|
||||
t.* = .{ .id = next_id, .state = .ready, .priority = priority, .stack = stack };
|
||||
t.* = .{ .id = next_id, .state = .ready, .priority = priority, .stack = stack, .affinity = affinity };
|
||||
next_id += 1;
|
||||
const top = @intFromPtr(stack.ptr) + stack.len;
|
||||
t.rsp = arch.initTaskStack(top, @intFromPtr(entry));
|
||||
enqueue(t);
|
||||
return t;
|
||||
}
|
||||
|
||||
fn freeSlot() ?*Task {
|
||||
@@ -102,38 +242,43 @@ fn freeSlot() ?*Task {
|
||||
return null;
|
||||
}
|
||||
|
||||
/// Pick the highest-priority ready task and switch to it. Interrupts must be
|
||||
/// disabled by the caller.
|
||||
/// Pick the highest-priority ready task and switch this core to it. The big kernel
|
||||
/// lock must be held by the caller (which also keeps local interrupts disabled);
|
||||
/// it serialises every core's scheduling, so no other core can touch the shared
|
||||
/// queues while we requeue `prev` and dequeue `next`. A dequeued task is `.ready`,
|
||||
/// never running elsewhere, so two cores never run the same task.
|
||||
fn schedule() void {
|
||||
const prev = current;
|
||||
const pc = thisCpu();
|
||||
const prev = pc.current;
|
||||
if (prev.state == .running) {
|
||||
prev.state = .ready;
|
||||
enqueue(prev); // back of its level's queue (round-robin)
|
||||
}
|
||||
const next = dequeueHighest() orelse {
|
||||
const next = dequeueHighest(pc) orelse {
|
||||
prev.state = .running; // nothing else ready — keep running
|
||||
return;
|
||||
};
|
||||
next.state = .running;
|
||||
current = next;
|
||||
pc.current = next;
|
||||
if (next != prev) arch.switchContext(&prev.rsp, next.rsp);
|
||||
}
|
||||
|
||||
/// Voluntarily give up the CPU to the next ready task.
|
||||
pub fn yield() void {
|
||||
const flags = arch.saveInterrupts();
|
||||
const flags = sync.enter();
|
||||
schedule();
|
||||
arch.restoreInterrupts(flags);
|
||||
sync.leave(flags);
|
||||
}
|
||||
|
||||
/// Block the current task for `ms` milliseconds, then let it become runnable
|
||||
/// again. The idle task (or other work) runs in the meantime.
|
||||
pub fn sleep(ms: u64) void {
|
||||
const flags = arch.saveInterrupts();
|
||||
current.wake_at = arch.millis() + ms;
|
||||
current.state = .blocked;
|
||||
const flags = sync.enter();
|
||||
const t = cur();
|
||||
t.wake_at = arch.millis() + ms;
|
||||
t.state = .blocked;
|
||||
schedule(); // current is blocked, so schedule() won't re-enqueue it
|
||||
arch.restoreInterrupts(flags);
|
||||
sync.leave(flags);
|
||||
}
|
||||
|
||||
// --- event-based blocking -------------------------------------------------
|
||||
@@ -147,27 +292,29 @@ pub const WaitQueue = struct {
|
||||
head: ?*Task = null,
|
||||
};
|
||||
|
||||
/// Block the current task on `wq` and switch away. Precondition: interrupts are
|
||||
/// disabled (the caller holds them, so a condition can be checked and the block
|
||||
/// committed atomically). On return — when woken — interrupts are still disabled.
|
||||
/// Block the current task on `wq` and switch away. Precondition: the big kernel
|
||||
/// lock is held (so a condition can be checked and the block committed atomically;
|
||||
/// it also keeps local interrupts disabled). On return — when woken — the lock is
|
||||
/// still held.
|
||||
pub fn waitLocked(wq: *WaitQueue) void {
|
||||
current.state = .blocked;
|
||||
current.next = wq.head;
|
||||
wq.head = current;
|
||||
const t = cur();
|
||||
t.state = .blocked;
|
||||
t.next = wq.head;
|
||||
wq.head = t;
|
||||
schedule();
|
||||
}
|
||||
|
||||
/// Move the highest-priority waiter on `wq` (if any) to the ready queue.
|
||||
/// Precondition: interrupts disabled. Does not preempt — the caller decides.
|
||||
/// Precondition: the big kernel lock is held. Does not preempt — the caller decides.
|
||||
pub fn wakeLocked(wq: *WaitQueue) void {
|
||||
// Find the highest-priority waiter (bounded scan) and unlink it.
|
||||
var best_prev: ?*Task = null;
|
||||
var best: ?*Task = null;
|
||||
var prev: ?*Task = null;
|
||||
var cur = wq.head;
|
||||
while (cur) |t| : ({
|
||||
var node = wq.head;
|
||||
while (node) |t| : ({
|
||||
prev = t;
|
||||
cur = t.next;
|
||||
node = t.next;
|
||||
}) {
|
||||
if (best == null or t.priority > best.?.priority) {
|
||||
best = t;
|
||||
@@ -182,25 +329,31 @@ pub fn wakeLocked(wq: *WaitQueue) void {
|
||||
|
||||
/// Block on `wq` (a self-contained critical section).
|
||||
pub fn wait(wq: *WaitQueue) void {
|
||||
const flags = arch.saveInterrupts();
|
||||
const flags = sync.enter();
|
||||
waitLocked(wq);
|
||||
arch.restoreInterrupts(flags);
|
||||
sync.leave(flags);
|
||||
}
|
||||
|
||||
/// Wake the highest-priority waiter on `wq`, preempting if it outranks us.
|
||||
pub fn wake(wq: *WaitQueue) void {
|
||||
const flags = arch.saveInterrupts();
|
||||
const flags = sync.enter();
|
||||
const pc = thisCpu();
|
||||
wakeLocked(wq);
|
||||
// If a higher-priority task is now ready, run it immediately.
|
||||
if (highestReadyPriority()) |p| {
|
||||
if (p > current.priority) schedule();
|
||||
// If a task this core would now pick outranks the running one, run it at once.
|
||||
// (A waiter pinned to *another* core isn't counted — that core picks it up on its
|
||||
// next tick; this core doesn't preempt for work it can't run.)
|
||||
if (highestReadyPriority(pc)) |p| {
|
||||
if (p > pc.current.priority) schedule();
|
||||
}
|
||||
arch.restoreInterrupts(flags);
|
||||
sync.leave(flags);
|
||||
}
|
||||
|
||||
fn highestReadyPriority() ?Priority {
|
||||
if (ready_bitmap == 0) return null;
|
||||
return @intCast(num_priorities - 1 - @clz(ready_bitmap));
|
||||
/// The highest-priority task core `pc` could run right now — the better of the global
|
||||
/// queue and this core's pinned queue — or null if it would fall back to idle.
|
||||
fn highestReadyPriority(pc: *PerCpu) ?Priority {
|
||||
const top = @max(topLevel(ready_bitmap), topLevel(pc.pinned_bitmap));
|
||||
if (top < 0) return null;
|
||||
return @intCast(top);
|
||||
}
|
||||
|
||||
/// Wake any sleeping task whose deadline has passed. Bounded by the task count,
|
||||
@@ -217,10 +370,15 @@ fn wakeExpired() void {
|
||||
}
|
||||
|
||||
/// Called from the timer interrupt (interrupts already disabled): wake due
|
||||
/// sleepers, then preempt.
|
||||
/// sleepers, then preempt. Takes the kernel lock like any other critical section,
|
||||
/// but releases it *without* touching the interrupt flag — the handler's `iretq`
|
||||
/// restores the interrupted context's flags, so re-enabling here would open a
|
||||
/// nested-interrupt window before the return.
|
||||
pub fn tick() void {
|
||||
_ = sync.enter();
|
||||
wakeExpired();
|
||||
if (preemption_enabled) schedule();
|
||||
sync.leaveIsr();
|
||||
}
|
||||
|
||||
/// Enable or disable timer-driven preemption (cooperative-only when off).
|
||||
@@ -229,23 +387,35 @@ pub fn setPreemption(enabled: bool) void {
|
||||
}
|
||||
|
||||
/// End the current task and switch away for good; never returns. The task's stack
|
||||
/// is leaked for now (no reaper yet).
|
||||
/// is leaked for now (no reaper yet). Acquires the kernel lock and hands it off to
|
||||
/// the task we switch into (which releases it) — this frame never returns to leave.
|
||||
pub fn exit() noreturn {
|
||||
arch.disableInterrupts();
|
||||
current.state = .free;
|
||||
const next = dequeueHighest() orelse @panic("sched: no task left to run");
|
||||
_ = sync.enter();
|
||||
const pc = thisCpu();
|
||||
pc.current.state = .free;
|
||||
const next = dequeueHighest(pc) orelse @panic("sched: no task left to run");
|
||||
next.state = .running;
|
||||
current = next;
|
||||
pc.current = next;
|
||||
var discard: usize = 0;
|
||||
arch.switchContext(&discard, next.rsp);
|
||||
unreachable;
|
||||
}
|
||||
|
||||
pub fn currentId() u32 {
|
||||
return current.id;
|
||||
return cur().id;
|
||||
}
|
||||
|
||||
/// The dense index of the core this task is currently running on (0 = BSP). Reads
|
||||
/// per-CPU state, so a task calling it on different cores sees different values —
|
||||
/// which is how a test can prove work is running in parallel. Returns 0 if the GS
|
||||
/// base isn't published yet (a fault in very early boot, before `init`), so a fault
|
||||
/// reporter can call it unconditionally without a second fault.
|
||||
pub fn currentCpuIndex() u32 {
|
||||
if (arch.cpuLocal() == 0) return 0;
|
||||
return thisCpu().index;
|
||||
}
|
||||
|
||||
/// Change the running task's priority (takes effect next time it's enqueued).
|
||||
pub fn setPriority(p: Priority) void {
|
||||
current.priority = p;
|
||||
cur().priority = p;
|
||||
}
|
||||
|
||||
@@ -0,0 +1,80 @@
|
||||
//! The big kernel lock (BKL) — the coarse mutual exclusion that lets more than one
|
||||
//! CPU run kernel code safely.
|
||||
//!
|
||||
//! Until SMP, the kernel's mutual exclusion *was* the interrupt flag: a critical
|
||||
//! section did `cli`, and since only one core existed, nothing else could touch
|
||||
//! kernel state (the discipline in docs/scheduling.md). That invariant dies the
|
||||
//! instant a second core runs kernel code — `cli` on one core does nothing to
|
||||
//! another. So the kernel's shared state (the scheduler queues, IPC channels) is
|
||||
//! guarded by a spinlock, and the lock is **always held with local interrupts
|
||||
//! disabled**, so a core's own timer interrupt can't re-enter the kernel and
|
||||
//! deadlock against the lock it already holds.
|
||||
//!
|
||||
//! This is deliberately *one coarse lock*, not many fine ones: it's philosophically
|
||||
//! aligned with a tiny kernel and it keeps the single-core correctness model
|
||||
//! (docs/scheduling.md) largely intact — one lock around kernel entry instead of
|
||||
//! rethinking every critical section. It's the first-design choice seL4 makes and
|
||||
//! docs/smp.md endorses; per-core run queues + fine-grained locking come later, if
|
||||
//! contention ever bites. Because the kernel does little, the lock is held briefly.
|
||||
//!
|
||||
//! **The hand-off rule.** The lock is held *across* a context switch and released
|
||||
//! by whichever task resumes, not by the one that switched away. A task that blocks
|
||||
//! or yields calls `enter`, mutates the queues, `schedule()`s — switching to another
|
||||
//! task *with the lock still held* — and only calls `leave` once it is eventually
|
||||
//! resumed and its critical section runs to the end. So every call into `schedule()`
|
||||
//! (and thus `switch_context`) happens with the lock held, and every task resumes
|
||||
//! from a switch holding it. A freshly-spawned task has no `enter`/`leave` frame to
|
||||
//! resume into, so `task_trampoline` releases the lock explicitly on its behalf via
|
||||
//! `releaseForFreshTask` before running the task body.
|
||||
|
||||
const std = @import("std");
|
||||
const arch = @import("arch");
|
||||
|
||||
/// 0 = free, 1 = held. A single global lock for the whole kernel.
|
||||
var held = std.atomic.Value(u32).init(0);
|
||||
|
||||
/// Enter the kernel: disable interrupts on this core, then spin until we own the
|
||||
/// lock. Returns the caller's prior interrupt flags for `leave` to restore.
|
||||
/// Interrupts stay off for the whole critical section so this core's timer tick
|
||||
/// can't try to re-acquire the lock we're holding.
|
||||
pub fn enter() u64 {
|
||||
const flags = arch.saveInterrupts();
|
||||
acquire();
|
||||
return flags;
|
||||
}
|
||||
|
||||
/// Release the lock and restore the interrupt flags `enter` returned (re-enabling
|
||||
/// interrupts only if they were on beforehand). The normal exit for a critical
|
||||
/// section reached from task context (`yield`, `sleep`, `wait`, `wake`, IPC).
|
||||
pub fn leave(flags: u64) void {
|
||||
release();
|
||||
arch.restoreInterrupts(flags);
|
||||
}
|
||||
|
||||
/// Release the lock but leave interrupts as they are. The exit for a critical
|
||||
/// section running inside an interrupt handler (the timer `tick`): the handler's
|
||||
/// `iretq` is what restores the interrupted context's flags, so restoring them
|
||||
/// here too would open a nested-interrupt window before the return. Release only.
|
||||
pub fn leaveIsr() void {
|
||||
release();
|
||||
}
|
||||
|
||||
/// Release the lock on behalf of a freshly-spawned task. Such a task is switched to
|
||||
/// (with the lock held) but has no `enter`/`leave` frame of its own to release
|
||||
/// through — `task_trampoline` calls this before running the task body. Interrupts
|
||||
/// are enabled separately by the trampoline. Exported for the assembly trampoline.
|
||||
export fn releaseForFreshTask() callconv(.c) void {
|
||||
release();
|
||||
}
|
||||
|
||||
fn acquire() void {
|
||||
// Test-and-test-and-set: try once, then spin read-only until the lock looks
|
||||
// free before retrying the (bus-locked) swap — cheaper on the coherency fabric.
|
||||
while (held.swap(1, .acquire) != 0) {
|
||||
while (held.load(.monotonic) != 0) arch.cpuRelax();
|
||||
}
|
||||
}
|
||||
|
||||
fn release() void {
|
||||
held.store(0, .release);
|
||||
}
|
||||
@@ -50,6 +50,10 @@ fn result() void {
|
||||
pub fn run(case: []const u8, boot_info: *const BootInfo) void {
|
||||
if (eql(case, "smoke")) {
|
||||
smoke(boot_info);
|
||||
} else if (eql(case, "discovery")) {
|
||||
discoveryTest();
|
||||
} else if (eql(case, "wx")) {
|
||||
wxTest();
|
||||
} else if (eql(case, "timer")) {
|
||||
timer();
|
||||
} else if (eql(case, "clock")) {
|
||||
@@ -68,12 +72,22 @@ pub fn run(case: []const u8, boot_info: *const BootInfo) void {
|
||||
eventTest();
|
||||
} else if (eql(case, "ipc")) {
|
||||
ipcTest();
|
||||
} else if (eql(case, "smp")) {
|
||||
smpTest();
|
||||
} else if (eql(case, "affinity")) {
|
||||
affinityTest();
|
||||
} else if (eql(case, "smp-stress")) {
|
||||
stressTest();
|
||||
} else if (eql(case, "smp-retry")) {
|
||||
smpRetryTest();
|
||||
} else if (eql(case, "fault-ud")) {
|
||||
faultInvalidOpcode();
|
||||
} else if (eql(case, "fault-pf")) {
|
||||
faultPageFault();
|
||||
} else if (eql(case, "fault-df")) {
|
||||
faultDoubleFault();
|
||||
} else if (eql(case, "fault-ap-df")) {
|
||||
faultApTest();
|
||||
} else if (eql(case, "fault-nx")) {
|
||||
faultNoExecute();
|
||||
} else if (eql(case, "fault-null")) {
|
||||
@@ -165,6 +179,54 @@ fn timer() void {
|
||||
result();
|
||||
}
|
||||
|
||||
/// Verify device discovery populated the platform facts the rest of the kernel
|
||||
/// depends on — the results ACPI parsing stashed in globals at boot. These are
|
||||
/// stable for the QEMU q35 + OVMF machine the harness runs, and span the tables:
|
||||
/// MADT (LAPIC base, CPU count), FADT (PM/reset registers), and the AML parse
|
||||
/// (the sleep type, plus the integrity check that every byte was consumed).
|
||||
fn discoveryTest() void {
|
||||
log("DANOS-TEST-BEGIN: discovery\n", .{});
|
||||
const pinfo = platform.platformInfo();
|
||||
const pw = platform.powerInfo();
|
||||
const am = platform.amlStats();
|
||||
|
||||
check("LAPIC base discovered (MADT)", pinfo.lapic_base == 0xFEE00000);
|
||||
check("ACPI PM timer found (FADT)", pinfo.pm_timer.present());
|
||||
check("PM1a control register found (FADT)", pw.pm1a_cnt.present());
|
||||
check("reset register supported (FADT)", pw.reset_supported);
|
||||
check("S5 sleep type found (AML)", pw.s5 != null);
|
||||
check("AML parsed completely (consumed == total)", am.total > 0 and am.consumed == am.total);
|
||||
check("at least one CPU enumerated (MADT)", platform.cpus().len >= 1);
|
||||
|
||||
result();
|
||||
}
|
||||
|
||||
/// Audit the W^X invariant across the memory classes: kernel code must be
|
||||
/// executable, everything else must not be. `arch.pageExecutable` reads the leaf
|
||||
/// page-table entry's NX bit, so this guards the permission overlay in paging.zig —
|
||||
/// a broader check than `fault-nx`, which only exercises one data page.
|
||||
fn wxTest() void {
|
||||
log("DANOS-TEST-BEGIN: wx\n", .{});
|
||||
|
||||
check("kernel code is executable (R+X)", arch.pageExecutable(@intFromPtr(&wxTest)));
|
||||
|
||||
const ro = "danos-wx-probe"; // string literal -> .rodata
|
||||
check("rodata is non-executable (NX)", !arch.pageExecutable(@intFromPtr(ro.ptr)));
|
||||
|
||||
check("kernel data is non-executable (NX)", !arch.pageExecutable(@intFromPtr(&passed)));
|
||||
|
||||
if (heap.allocator().alloc(u8, 64) catch null) |h| {
|
||||
check("heap is non-executable (NX)", !arch.pageExecutable(@intFromPtr(h.ptr)));
|
||||
heap.allocator().free(h);
|
||||
}
|
||||
|
||||
var local: u64 = 0;
|
||||
_ = &local;
|
||||
check("stack is non-executable (NX)", !arch.pageExecutable(@intFromPtr(&local)));
|
||||
|
||||
result();
|
||||
}
|
||||
|
||||
/// Verify the on-demand VMM: map a fresh frame at an unused virtual address, and
|
||||
/// check it's writable and reads back.
|
||||
fn vmm() void {
|
||||
@@ -437,6 +499,209 @@ fn sleepTest() void {
|
||||
result();
|
||||
}
|
||||
|
||||
// --- SMP parallelism ------------------------------------------------------
|
||||
|
||||
var seen_core = [_]bool{false} ** 8;
|
||||
var smp_running: bool = true;
|
||||
|
||||
/// A worker that, while running, records which core it's executing on. Spread across
|
||||
/// spawned workers and idle APs, these should land on more than one core.
|
||||
fn smpWorker() void {
|
||||
const p: *volatile bool = &smp_running;
|
||||
while (p.*) {
|
||||
const c = sched.currentCpuIndex();
|
||||
if (c < seen_core.len) seen_core[c] = true;
|
||||
}
|
||||
sched.exit();
|
||||
}
|
||||
|
||||
/// Prove tasks run **in parallel** on multiple cores (not just interleaved on one).
|
||||
/// Spawn several CPU-bound workers; each stamps the core it runs on into `seen_core`.
|
||||
/// With the application processors online, more than one core should show up — which
|
||||
/// can only happen if work is genuinely running at the same time on different cores.
|
||||
/// (Run with QEMU `-smp N`; on a single core this would see just one and fail.)
|
||||
fn smpTest() void {
|
||||
log("DANOS-TEST-BEGIN: smp\n", .{});
|
||||
seen_core = .{false} ** 8;
|
||||
smp_running = true;
|
||||
|
||||
var i: usize = 0;
|
||||
while (i < 4) : (i += 1) sched.spawn(smpWorker, 4);
|
||||
|
||||
// Let the workers run across cores for a stretch of real time.
|
||||
var spins: u64 = 0;
|
||||
while (spins < 2_000_000_000) spins +%= 1;
|
||||
smp_running = false;
|
||||
|
||||
var cores_seen: u32 = 0;
|
||||
for (seen_core) |s| {
|
||||
if (s) cores_seen += 1;
|
||||
}
|
||||
log("DANOS-SMP: workers ran on {d} distinct core(s)\n", .{cores_seen});
|
||||
check("tasks ran on multiple cores in parallel", cores_seen >= 2);
|
||||
|
||||
// Bring-up is done, so the trampoline frame must be inert: zeroed (no stale code)
|
||||
// and non-executable (W^X restored). It's armed only while a core is climbing.
|
||||
const tramp = arch.trampolinePage();
|
||||
check("trampoline frame reserved", tramp != 0);
|
||||
if (tramp != 0) {
|
||||
const bytes: [*]const u8 = @ptrFromInt(tramp);
|
||||
var zeroed = true;
|
||||
for (0..4096) |b| {
|
||||
if (bytes[b] != 0) zeroed = false;
|
||||
}
|
||||
check("trampoline page zeroed when dormant", zeroed);
|
||||
check("trampoline page non-executable when dormant", !arch.pageExecutable(tramp));
|
||||
}
|
||||
result();
|
||||
}
|
||||
|
||||
// --- affinity: a pinned task never migrates -------------------------------
|
||||
|
||||
var affinity_cores = [_]bool{false} ** 8;
|
||||
var affinity_running: bool = true;
|
||||
|
||||
fn affinityWorker() void {
|
||||
const p: *volatile bool = &affinity_running;
|
||||
while (p.*) {
|
||||
const c = sched.currentCpuIndex();
|
||||
if (c < affinity_cores.len) affinity_cores[c] = true;
|
||||
}
|
||||
sched.exit();
|
||||
}
|
||||
|
||||
/// A task pinned to a core must run **only** on that core. Pin a busy worker to
|
||||
/// core 1 and let it run through many preemptions; it must have stamped core 1 and no
|
||||
/// other. An *unpinned* task scatters across cores (that's what the smp test shows),
|
||||
/// so a broken pin fails this deterministically — over this many time slices a
|
||||
/// free-floating task will land on some other core.
|
||||
fn affinityTest() void {
|
||||
log("DANOS-TEST-BEGIN: affinity\n", .{});
|
||||
affinity_cores = .{false} ** 8;
|
||||
affinity_running = true;
|
||||
|
||||
if (!sched.spawnOn(affinityWorker, 4, 1)) {
|
||||
check("worker pinned to core 1 (run with -smp)", false);
|
||||
result();
|
||||
return;
|
||||
}
|
||||
|
||||
var spins: u64 = 0;
|
||||
while (spins < 3_000_000_000) spins +%= 1; // many time slices across the cores
|
||||
affinity_running = false;
|
||||
var settle: u64 = 0;
|
||||
while (settle < 200_000_000) settle +%= 1; // let the worker see the flag and exit
|
||||
|
||||
var others: u32 = 0;
|
||||
for (affinity_cores, 0..) |seen, c| {
|
||||
if (seen and c != 1) others += 1;
|
||||
}
|
||||
log("DANOS-AFFINITY: pinned worker touched core 1={}, other cores={d}\n", .{ affinity_cores[1], others });
|
||||
check("pinned task ran on its core (1)", affinity_cores[1]);
|
||||
check("pinned task never migrated to another core", others == 0);
|
||||
result();
|
||||
}
|
||||
|
||||
// --- SMP stress: hammer the big kernel lock across cores ------------------
|
||||
|
||||
const stress_pairs = 4; // producer/consumer pairs (8 tasks; fits the 16-task pool)
|
||||
const stress_msgs = 100_000; // messages per pair
|
||||
const stress_cap = 4; // small channel -> constant block/wake, more lock churn
|
||||
|
||||
var stress_chan = [_]ipc.Channel(u64, stress_cap){.{}} ** stress_pairs;
|
||||
var stress_recv = [_]u64{0} ** stress_pairs; // messages received per pair
|
||||
var stress_order_ok = [_]bool{true} ** stress_pairs; // FIFO order held per pair
|
||||
var stress_cores = [_]bool{false} ** 8; // cores that ran a consumer
|
||||
var stress_prod_claim: usize = 0;
|
||||
var stress_cons_claim: usize = 0;
|
||||
|
||||
fn stressProducer() void {
|
||||
// Claim a unique pair index (atomic: producers start on different cores).
|
||||
const idx = @atomicRmw(usize, &stress_prod_claim, .Add, 1, .monotonic);
|
||||
var v: u64 = 1;
|
||||
while (v <= stress_msgs) : (v += 1) stress_chan[idx].send(v);
|
||||
sched.exit();
|
||||
}
|
||||
|
||||
fn stressConsumer() void {
|
||||
const idx = @atomicRmw(usize, &stress_cons_claim, .Add, 1, .monotonic);
|
||||
var expected: u64 = 1;
|
||||
while (expected <= stress_msgs) : (expected += 1) {
|
||||
const got = stress_chan[idx].recv();
|
||||
if (got != expected) stress_order_ok[idx] = false; // lost/reordered => lock broke
|
||||
const c = sched.currentCpuIndex();
|
||||
if (c < stress_cores.len) stress_cores[c] = true;
|
||||
stress_recv[idx] = expected;
|
||||
}
|
||||
sched.exit();
|
||||
}
|
||||
|
||||
/// Stress the big kernel lock under sustained cross-core contention. Each pair drives
|
||||
/// `stress_msgs` sequenced messages through a 4-slot channel — every send and recv
|
||||
/// takes the lock, and the small buffer forces constant block/wake (so the scheduler
|
||||
/// churns too). A single-producer/single-consumer channel must deliver in strict FIFO
|
||||
/// order; if the lock let two cores into a critical section at once, the ring buffer
|
||||
/// corrupts and the consumer sees a wrong or out-of-order value (or the run hangs /
|
||||
/// faults). Passing means ~320k lock acquisitions across the cores stayed consistent.
|
||||
fn stressTest() void {
|
||||
log("DANOS-TEST-BEGIN: smp-stress\n", .{});
|
||||
stress_chan = [_]ipc.Channel(u64, stress_cap){.{}} ** stress_pairs;
|
||||
stress_recv = [_]u64{0} ** stress_pairs;
|
||||
stress_order_ok = [_]bool{true} ** stress_pairs;
|
||||
stress_cores = [_]bool{false} ** 8;
|
||||
stress_prod_claim = 0;
|
||||
stress_cons_claim = 0;
|
||||
|
||||
var i: usize = 0;
|
||||
while (i < stress_pairs) : (i += 1) sched.spawn(stressConsumer, 4);
|
||||
i = 0;
|
||||
while (i < stress_pairs) : (i += 1) sched.spawn(stressProducer, 4);
|
||||
|
||||
// Drop below the workers so they get the cores; wake periodically to check for
|
||||
// completion. A broken lock instead hangs here (harness timeout) or faults.
|
||||
sched.setPriority(1);
|
||||
var spins: u64 = 0;
|
||||
while (spins < 40_000_000_000) : (spins += 1) {
|
||||
var done = true;
|
||||
for (stress_recv) |n| {
|
||||
if (n < stress_msgs) done = false;
|
||||
}
|
||||
if (done) break;
|
||||
}
|
||||
sched.setPriority(4);
|
||||
|
||||
var total: u64 = 0;
|
||||
for (stress_recv) |n| total += n;
|
||||
var order_ok = true;
|
||||
for (stress_order_ok) |ok| {
|
||||
if (!ok) order_ok = false;
|
||||
}
|
||||
var cores: u32 = 0;
|
||||
for (stress_cores) |s| {
|
||||
if (s) cores += 1;
|
||||
}
|
||||
|
||||
log("DANOS-STRESS: {d}/{d} pairs complete on {d} cores\n", .{ total, @as(u64, stress_pairs) * stress_msgs, cores });
|
||||
check("every message delivered", total == @as(u64, stress_pairs) * stress_msgs);
|
||||
check("strict FIFO order held (no lock corruption)", order_ok);
|
||||
check("contention was genuinely cross-core", cores >= 2);
|
||||
result();
|
||||
}
|
||||
|
||||
/// Retry: `main` forced the first AP wake attempt to fail (arch.testFailNextWakes),
|
||||
/// so a core missed its first INIT-SIPI-SIPI. The boot retry must have brought it back
|
||||
/// anyway — every enumerated core should be online. If retry were broken, that core
|
||||
/// would be parked and the count would fall short.
|
||||
fn smpRetryTest() void {
|
||||
log("DANOS-TEST-BEGIN: smp-retry\n", .{});
|
||||
const total = platform.cpus().len;
|
||||
const online = sched.onlineCount();
|
||||
log("DANOS-RETRY: {d}/{d} cores online after a forced first-wake failure\n", .{ online, total });
|
||||
check("multiple cores enumerated (run with -smp)", total >= 2);
|
||||
check("retry brought every core online despite a failed first wake", online == total);
|
||||
result();
|
||||
}
|
||||
|
||||
fn faultInvalidOpcode() void {
|
||||
log("DANOS-TEST-BEGIN: fault-ud\n", .{});
|
||||
asm volatile ("ud2");
|
||||
@@ -489,3 +754,44 @@ fn faultDoubleFault() void {
|
||||
);
|
||||
bad_sp += 0;
|
||||
}
|
||||
|
||||
var ap_reached_fault: bool = false;
|
||||
|
||||
/// A task that faults with a #DF *on whatever core it's pinned to*. Announces the
|
||||
/// core, then triggers the same double fault as `faultDoubleFault` — which is only
|
||||
/// survivable on IST1, so it exercises that core's own TSS.
|
||||
fn apDoubleFaultTask() void {
|
||||
log("DANOS-AP: task running on core {d}, triggering #DF\n", .{sched.currentCpuIndex()});
|
||||
@atomicStore(bool, &ap_reached_fault, true, .release);
|
||||
arch.disableInterrupts();
|
||||
var bad_sp: u64 = 0x5000000000;
|
||||
asm volatile (
|
||||
\\mov %[sp], %%rsp
|
||||
\\ud2
|
||||
:
|
||||
: [sp] "r" (bad_sp),
|
||||
: .{ .memory = true }
|
||||
);
|
||||
bad_sp += 0;
|
||||
}
|
||||
|
||||
/// Fault on an application processor. Pins a double-faulting task to core 1, so the
|
||||
/// fault is taken and handled by *that core's own* IDT and TSS/IST — not the BSP's.
|
||||
/// The harness matches "core N: double fault (vector 8)" with N ≥ 1, which can only
|
||||
/// appear if the AP caught the #DF on its IST1 (a broken per-core TSS would
|
||||
/// triple-fault and reset instead). We then show the BSP still runs afterwards, so
|
||||
/// the fault was *contained* to the AP, not fatal to the system.
|
||||
fn faultApTest() void {
|
||||
log("DANOS-TEST-BEGIN: fault-ap-df\n", .{});
|
||||
if (!sched.spawnOn(apDoubleFaultTask, 6, 1)) {
|
||||
log("DANOS-AP: could not pin to core 1 (run with -smp) - FAIL\n", .{});
|
||||
arch.halt();
|
||||
}
|
||||
// Wait until the AP is about to fault, then keep running to prove containment.
|
||||
var spins: u64 = 0;
|
||||
while (!@atomicLoad(bool, &ap_reached_fault, .acquire) and spins < 5_000_000_000) spins +%= 1;
|
||||
var settle: u64 = 0;
|
||||
while (settle < 500_000_000) settle +%= 1; // let the AP take + report the fault
|
||||
log("DANOS-BSP: core {d} still running after the AP fault (contained)\n", .{sched.currentCpuIndex()});
|
||||
arch.halt();
|
||||
}
|
||||
|
||||
+46
-2
@@ -81,6 +81,12 @@ CASES = [
|
||||
{"name": "smoke",
|
||||
"expect": r"DANOS-TEST-RESULT: PASS",
|
||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||
{"name": "discovery",
|
||||
"expect": r"DANOS-TEST-RESULT: PASS",
|
||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||
{"name": "wx",
|
||||
"expect": r"DANOS-TEST-RESULT: PASS",
|
||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||
{"name": "timer",
|
||||
"expect": r"DANOS-TEST-RESULT: PASS",
|
||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||
@@ -108,13 +114,48 @@ CASES = [
|
||||
{"name": "ipc",
|
||||
"expect": r"DANOS-TEST-RESULT: PASS",
|
||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||
# Parallelism: needs more than one core, so this case boots with -smp 4.
|
||||
{"name": "smp",
|
||||
"smp": 4,
|
||||
"expect": r"DANOS-TEST-RESULT: PASS",
|
||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||
# Affinity: a pinned task must never migrate off its core.
|
||||
{"name": "affinity",
|
||||
"smp": 4,
|
||||
"expect": r"DANOS-TEST-RESULT: PASS",
|
||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||
# Stress the big kernel lock across cores; heavier, so a longer timeout.
|
||||
{"name": "smp-stress",
|
||||
"smp": 4,
|
||||
"timeout": 90,
|
||||
"expect": r"DANOS-TEST-RESULT: PASS",
|
||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||
# Retry: a forced first-wake failure must still bring every core online.
|
||||
{"name": "smp-retry",
|
||||
"smp": 4,
|
||||
"expect": r"DANOS-TEST-RESULT: PASS",
|
||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||
{"name": "fault-ud", "expect": r"invalid opcode \(vector 6\)"},
|
||||
{"name": "fault-pf", "expect": r"page fault \(vector 14\)"},
|
||||
{"name": "fault-df", "expect": r"double fault \(vector 8\)"},
|
||||
# Fault on an application processor: a #DF pinned to core 1 must be caught by that
|
||||
# core's own IST (N >= 1), proving per-core TSS works; a broken one triple-faults.
|
||||
{"name": "fault-ap-df",
|
||||
"smp": 4,
|
||||
"expect": r"core [1-9]\d*: double fault \(vector 8\)",
|
||||
"fail": r"could not pin"},
|
||||
{"name": "fault-nx",
|
||||
"expect": r"page fault \(vector 14\)",
|
||||
"fail": r"NX not enforced"},
|
||||
{"name": "fault-null", "expect": r"page fault \(vector 14\)"},
|
||||
# The ACPI power path succeeds by QEMU *exiting* (S5 off / reset), so match the
|
||||
# pre-transition marker; the FAIL line only appears if the transition didn't take.
|
||||
{"name": "poweroff",
|
||||
"expect": r"DANOS-POWER: attempting poweroff",
|
||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||
{"name": "reboot",
|
||||
"expect": r"DANOS-POWER: attempting reboot",
|
||||
"fail": r"DANOS-TEST-RESULT: FAIL"},
|
||||
]
|
||||
|
||||
TIMEOUT = 30 # seconds per case
|
||||
@@ -175,9 +216,12 @@ def run_case(arch, case):
|
||||
fail = re.compile(case["fail"]) if case.get("fail") else None
|
||||
|
||||
cmd = [arch["qemu"]] + arch["qemu_args"](arch, esp, vars_fd, serial)
|
||||
if case.get("smp"): # some cases need more than one core (e.g. parallelism)
|
||||
cmd += ["-smp", str(case["smp"])]
|
||||
qemu = subprocess.Popen(cmd, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
|
||||
try:
|
||||
deadline = time.monotonic() + TIMEOUT
|
||||
timeout = case.get("timeout", TIMEOUT)
|
||||
deadline = time.monotonic() + timeout
|
||||
while time.monotonic() < deadline:
|
||||
time.sleep(0.2)
|
||||
text = ""
|
||||
@@ -192,7 +236,7 @@ def run_case(arch, case):
|
||||
if expect.search(text):
|
||||
return True, "matched " + repr(case["expect"])
|
||||
return False, "QEMU exited before matching (triple fault?)"
|
||||
return False, f"timed out after {TIMEOUT}s without matching {case['expect']!r}"
|
||||
return False, f"timed out after {timeout}s without matching {case['expect']!r}"
|
||||
finally:
|
||||
if qemu.poll() is None:
|
||||
qemu.terminate()
|
||||
|
||||
Reference in New Issue
Block a user