12 Commits
Author SHA1 Message Date
Daniel Samson 26ac97df31 test fault-on-AP and affinity; report the faulting core
fault-ap-df pins a #DF to an AP (its own IST must catch it); affinity checks a pinned task never migrates. onException now names the core, so an AP fault is attributed and shown contained. Both teeth-checked.
2026-07-08 14:45:28 +01:00
Daniel Samson 37f72a4df8 add thread affinity: pin a task to a core
spawnOn(entry, priority, cpu) routes to a per-core pinned queue, merged with the global queue at O(1) selection. Falls back to unpinned for an offline/invalid core.
2026-07-08 14:45:28 +01:00
Daniel Samson f02259cae0 close audited test gaps: discovery, W^X, power
Add a discovery case (asserts stable ACPI/MADT/FADT/AML facts) and a wx case (audits R+X code vs NX data/rodata/heap/stack via arch.pageExecutable). Wire the existing poweroff/reboot cases into the harness. Suite: 18 -> 22.
2026-07-08 14:17:54 +01:00
Daniel Samson 5940864958 test the trampoline is inert (zeroed + NX) when dormant
Adds paging.isExecutable and arch.trampolinePage; the smp case now asserts the frame is zeroed and non-executable after bring-up. Teeth-checked against a no-op disarm.
2026-07-08 13:46:53 +01:00
Daniel Samson 91f2cfa17b add smp-retry test for the AP wake retry path
Test hook forces the first wake to fail; the case asserts every core still comes online. Verified it fails when retry is disabled.
2026-07-08 13:41:14 +01:00
Daniel Samson debe815a5c keep the AP trampoline inert between wakes, and retry failed cores
Arm the low frame (copy blob, make executable) only while a core climbs, then zero it and restore RW+NX; the frame stays reserved so cores can be re-woken. startSecondary is one re-runnable attempt; boot retries a non-responding core 3x. Validated the retry path by forcing a first-attempt failure.
2026-07-08 13:32:34 +01:00
Daniel Samson dba3939a0f add an SMP lock-stress test case
Four pairs push 400k sequenced messages through small channels; checks FIFO order and cross-core execution. Verified to fail with the lock disabled.
2026-07-08 12:45:02 +01:00
Daniel Samson 43afe6bf2e schedule tasks across all cores
Per-core GDT/TSS and AP scheduler entry; fix AP SSE + single_threaded.
2026-07-08 12:35:30 +01:00
Daniel Samson ed7f542006 wake application processors to long mode
INIT-SIPI-SIPI plus a self-relocating real-mode trampoline.
2026-07-08 12:35:30 +01:00
Daniel Samson 36c29d2d6d move the running task into per-CPU state
PerCpu.current via the GS base; ready queues stay global.
2026-07-08 12:35:30 +01:00
Daniel Samson 941ab091db add a big kernel lock for SMP
Guards the scheduler and IPC; held across the context switch.
2026-07-08 12:35:30 +01:00
Daniel Samson cf7c6df41c enumerate usable CPU cores from the MADT
Keep each Local APIC's id and expose platform.cpus().
2026-07-08 12:35:30 +01:00
21 changed files with 1472 additions and 109 deletions
+4 -1
View File
@@ -73,6 +73,9 @@ pub fn build(b: *std.Build) void {
// CPU-exception stubs — real assembly, since they need cross-symbol
// jumps/calls that Zig inline asm can't express (see the file's header).
arch_mod.addAssemblyFile(b.path("src/kernel/arch/x86_64/isr.s"));
// The AP bring-up trampoline: 16-/32-/64-bit mode-switch code that can't be
// inline asm (it runs relocated to a low page, not at its link address).
arch_mod.addAssemblyFile(b.path("src/kernel/arch/x86_64/trampoline.s"));
// Firmware-agnostic device discovery. The generic kernel imports this as
// "platform" and asks it to enumerate hardware into a backend-neutral device
@@ -111,7 +114,7 @@ pub fn build(b: *std.Build) void {
.optimize = optimize,
.code_model = .small, // kernel is linked in the low 2 GiB (see image_base)
.red_zone = false, // interrupts would corrupt the SysV red zone
.single_threaded = true, // no scheduler yet; avoids pulling in TLS/atomics
.single_threaded = false, // SMP: the big kernel lock's atomics must be real across cores
.sanitize_c = .off, // the UBSan runtime needs f128/SSE support we don't provide
.stack_check = false, // stack-probe calls have no runtime to land in
.stack_protector = false,
+21
View File
@@ -67,6 +67,27 @@ exist, which is what a real-time scheduler needs.
- **Round-robin within a level.** When a task is descheduled it goes to the *back*
of its level's queue, so equal-priority tasks share the CPU fairly.
## Affinity: pinning a task to a core
By default a task runs on **any** core — the ready queue above is global, and any
idle core pulls the highest-priority task from it (work-conserving; see
[smp.md](smp.md)). A task can instead be **pinned** to one core with
`spawnOn(entry, priority, cpu)`, giving it an *affinity*: it will only ever run
there, never migrating.
Mechanically, each core has its **own** pinned queue (same 8-level FIFO + bitmap)
alongside the global one. A pinned task is enqueued only into its core's pinned
queue; selection compares the top of the global queue and the running core's pinned
queue and takes the higher priority (still O(1) — two bit-scans and a compare), with
a pinned task winning an equal-priority tie so it can't be starved by global work.
Because every queue is mutated under the [big kernel lock](smp.md), one core enqueuing
into another core's pinned queue is safe.
This is the *explicit-affinity* model (no surprise migration mid-deadline), which is
the more real-time-predictable direction. `spawnOn` refuses to pin to an offline or
out-of-range core — it creates the task unpinned instead, so it still runs somewhere
rather than stranding in a queue no core services, and returns whether the pin took.
## Sleeping and the idle task
A task can **block** — give up the CPU until an event, rather than busy-wait
+107 -7
View File
@@ -16,10 +16,14 @@ danos specifics):
- **Real-time** — whether timing is *predictable*. Comes from bounded operations
(our O(1) scheduler), not from core count.
danos is uniprocessor today: one global `current` task, one set of ready queues, one
timer. Even on an 8-core CPU, the firmware starts only the **bootstrap processor
(BSP)**; the other cores (**application processors**, APs) sit parked until the
kernel wakes them, which it doesn't yet.
danos now runs on multiple cores. The firmware starts only the **bootstrap processor
(BSP)**; the kernel wakes the other cores (**application processors**, APs) with
INIT–SIPI–SIPI, brings each up into 64-bit long mode with its own descriptor tables,
LAPIC and timer, and drops it into the scheduler. Tasks run **genuinely in parallel** —
the `smp` self-test confirms worker tasks executing on all four cores at once under
QEMU `-smp 4`. Shared kernel state (scheduler queues, IPC) is serialised behind a big
kernel lock. What's left is refinement, not first-light: per-core run queues, IPIs,
and thread-to-core affinity (see [Implementation status](#implementation-status)).
## The common microkernel instinct: don't share kernel state
@@ -114,14 +118,20 @@ active reconsideration in favour of resilience — see [vision.md](vision.md).)
Whatever the top goal, the *sequence* is the same and seL4 validates starting simple:
1. **Enumerate cores** — needs [device discovery](discovery.md) (ACPI MADT on x86,
device tree on ARM). SMP is a concrete consumer of that work.
device tree on ARM). SMP is a concrete consumer of that work. **Done on x86:** the
MADT parse records every usable Local APIC — with the `apic_id` an AP wake targets —
and `platform.cpus()` returns the list (see [discovery.md](discovery.md)). The boot
log reports the count; the ARM (device-tree) path still needs it.
2. **Wake the APs** — INIT–SIPI–SIPI on x86; PSCI/spin-tables on ARM. Each core brings
up its own tables, timer, and idle task.
up its own tables, timer, and idle task. **Done on x86** — cores climb to long mode,
set up their own GDT/TSS, and enter the scheduler; tasks run in parallel across all
cores ([status](#implementation-status)).
3. **Start with a big kernel lock.** It's a legitimate first design, not a shortcut —
philosophically aligned with a tiny kernel, and it lets the single-core correctness
model you already have (the interrupt-flag discipline in
[scheduling.md](scheduling.md)) stay largely intact: one lock around kernel entry
instead of rethinking every critical section.
instead of rethinking every critical section. **Done** — see
`src/kernel/sync.zig`.
4. **Later, if contention bites,** evolve toward **per-core run queues + explicit
affinity** (the Fiasco.OC direction) — also the more real-time-predictable model.
5. **Placement stays a user-space policy** — the kernel runs a thread on the core it's
@@ -130,6 +140,96 @@ Whatever the top goal, the *sequence* is the same and seL4 validates starting si
Big-lock-first → per-core-later. The affinity/MCS depth is only worth it if real-time
turns out to be the actual goal.
## Implementation status
The "wake + schedule" build (real parallel task execution) is going in as a sequence
of green checkpoints — each step keeps the single-core test suite passing before the
next lands.
**Done:**
- **Core enumeration** — the MADT parse records every usable Local APIC (with its
`apic_id`, which an AP wake targets); `platform.cpus()` returns the list. See
[discovery.md](discovery.md).
- **The big kernel lock** (`src/kernel/sync.zig`) — one coarse spinlock guarding the
scheduler queues and IPC, always held with local interrupts disabled. It is held
*across* a context switch and released by whichever task resumes (the hand-off
rule); `task_trampoline` releases it for a freshly-spawned task. `scheduler.zig` and
`ipc.zig` run every critical section under it. Uncontended on one core, so behaviour
is identical to the old interrupt-flag model.
- **Per-CPU state** — a `PerCpu` struct (running task, idle task, APIC id) per core,
its pointer kept in the x86 **GS base** (`IA32_GS_BASE`; no `swapgs`, since there's
no user mode yet). The old global `current` is now `thisCpu().current`. The ready
queues stay **global** under the lock — work-conserving, so any idle core will pull
the highest-priority ready task; per-core queues are a later optimisation.
- **AP wake to long mode** — `arch.startSecondary` drives INIT–SIPI–SIPI (via the
LAPIC ICR) to wake each parked core one at a time. A woken core starts in 16-bit
real mode at a low page and runs the [trampoline](../src/kernel/arch/x86_64/trampoline.s)
up through protected mode into 64-bit long mode, then lands in `smp.zig:apEntry`,
publishes its per-CPU pointer, and reports in. Verified in QEMU with `-smp 4`:
all four cores report `online`.
The trampoline earns its complexity from four hardware facts:
- a STARTUP IPI vectors a core to physical `vector << 12` (a *byte* vector), so the
trampoline must live **below 1 MiB** — the kernel reserves that page from the frame
allocator at boot, before paging/heap draw down the scarce low frames;
- the blanket RAM identity map is **NX** (W^X), but the AP fetches the trampoline
from it under paging, so that one page is made executable for bring-up;
- the blob is copied to a page whose address isn't known at link time, so it is
**position-independent**: it derives its own base from `CS` and, crucially,
addresses data *segment-relative in real mode* (where the segment base already
supplies the page base) but *base-register-relative in protected/long mode* (flat
segments, base 0). Getting that distinction wrong was the first bug found;
- an AP starts with a bare `CR0`/`CR4`, but the kernel is built **with SSE** (the
x86_64 baseline) and the compiler emits SSE for things as ordinary as a struct
copy — so the trampoline must set `CR4.OSFXSR`/`OSXMMEXCPT` and fix `CR0.EM`/`MP`,
or the first SSE instruction on the AP `#UD`s. The BSP inherited those bits from
UEFI; the AP has to set them itself. This was the second bug — it masqueraded as a
fault in `lgdt` (the first kernel code after entry that the compiler vectorised).
- **Per-core tables + scheduler entry** — each AP loads **its own GDT** (with its own
TSS descriptor) and **its own TSS** (its own IST/`rsp0` stack), loads the shared
IDT, enables its LAPIC and timer, then calls the generic `secondaryMain`: it turns
its bring-up context into the core's idle task (as task 0 is for the BSP), marks the
core online, and enters the run loop. With interrupts on, each core's own timer tick
preempts its idle context into whatever the global ready queue offers — so all cores
pull real work in parallel. The `smp` test spawns CPU-bound workers and confirms they
execute on all four cores at once, and `fault-ap-df` pins a #DF to an AP and checks
that core catches it on **its own** IST (a broken per-core TSS would triple-fault) —
reported as "core N: …", so a fault is always attributed to the core it happened on,
and is contained to that core (the rest of the system keeps running).
- **Thread affinity** — `spawnOn(entry, priority, cpu)` pins a task to a core (its own
per-core pinned queue, merged with the global queue at selection; see
[scheduling.md](scheduling.md#affinity-pinning-a-task-to-a-core)). The `affinity`
test confirms a pinned task never migrates. This is the mechanism the fault-on-AP
test rides on, and the *explicit-affinity* real-time-predictable model.
- **`single_threaded` off** — the kernel was built `single_threaded = true`, which
compiles `std.atomic` down to plain non-atomic ops. Harmless on one core, but it
quietly breaks the big kernel lock across cores; it's now `false`.
- **Re-armable wake + retry** — the trampoline frame is reserved for the system's
life, but kept **inert between wakes**: zeroed and non-executable, armed (blob
copied in, page made executable) only for the moment a core is actually climbing,
then disarmed again. So there's never a dormant executable page, and a core can be
(re)woken at any time — `arch.startSecondary` is one self-contained attempt (arm →
INIT–SIPI–SIPI → disarm), and its `INIT` resets a wedged core, so retrying just
works. Boot retries a non-responding core up to three times; the same primitive is
the groundwork a future **power manager** would drive to bring cores up (and,
eventually, its counterpart to take them offline — which additionally needs the
core's tasks migrated off first).
**Next (refinement, not first-light):**
- **IPIs** — cross-core wake/preempt. Not needed for correctness: an idle core wakes
on its own timer tick and pulls ready work then; IPIs only cut that latency from
≤1 ms to near-instant.
- **Per-core run queues** — the Fiasco.OC direction, if the single global queue's lock
contention ever bites. (Thread *affinity* already exists — see above; this is the
further step of giving each core its own primary run queue for load distribution.)
- **Fault recovery** — today a fault halts (only) the faulting core. Turning that into
"kill the task, keep the core running" is the [resilience](resilience.md) track (it
needs the task's lock/resource state handled), and for taking a core fully offline,
its tasks migrated first.
## Further reading
**Microkernel SMP & scheduling**
+41
View File
@@ -91,6 +91,35 @@ pub const PlatformInfo = struct {
/// Filled in by `discover`; the arch layer reads it during bring-up.
pub var platform_info: PlatformInfo = .{};
/// One usable logical processor, from a MADT type-0 (Local APIC) record. The
/// `apic_id` is the Local APIC ID that SMP bring-up targets to wake this core
/// (INIT–SIPI–SIPI); `processor_id` is the ACPI namespace handle. Only processors
/// the firmware marks *enabled* are recorded — a disabled one can't be started.
pub const Cpu = struct {
processor_id: u8,
apic_id: u8,
/// MADT flags bit 1: usable but firmware-started offline (hot-plug / deferred
/// bring-up), as opposed to already available. Informational for now.
online_capable: bool,
};
/// The set of usable logical processors the MADT listed — the hardware's degree of
/// parallelism. Includes the bootstrap processor danos already runs on; the rest
/// are the application processors SMP bring-up would start (see docs/smp.md).
pub const CpuInfo = struct {
/// A static pool sized well above any danos target (a desktop, two 4-core Pis).
/// If the MADT ever lists more, the surplus is dropped and counted in `dropped`
/// so the truncation is never silent.
cpus: [max_cpus]Cpu = undefined,
count: usize = 0,
dropped: usize = 0,
};
const max_cpus = 64;
/// Filled in by `discover` (from the MADT); SMP bring-up reads it to wake the APs.
pub var cpu_info: CpuInfo = .{};
/// Integrity/diagnostics for the AML parse. `consumed == total` means the parser
/// walked every byte of the DSDT/SSDTs without desyncing.
pub const AmlStats = struct {
@@ -438,6 +467,18 @@ fn parseMadt(dt: *DeviceTree, header: *const SystemDescriptorTableHeader) !void
var nb: [24]u8 = undefined;
const nm = std.fmt.bufPrint(&nb, "cpu{d}", .{la.processor_id}) catch "cpu";
_ = try dt.addChild(dt.root, .processor, nm);
// Also record it as a schedulable core (with the APIC ID an AP
// wake needs, which the device node name doesn't preserve).
if (cpu_info.count < cpu_info.cpus.len) {
cpu_info.cpus[cpu_info.count] = .{
.processor_id = la.processor_id,
.apic_id = la.apic_id,
.online_capable = la.flags & 2 != 0,
};
cpu_info.count += 1;
} else {
cpu_info.dropped += 1;
}
}
},
1 => {
+17
View File
@@ -24,6 +24,7 @@ pub const AmlStats = acpi.AmlStats;
pub const PlatformInfo = acpi.PlatformInfo;
pub const RegAccess = acpi.RegAccess;
pub const IsoEntry = acpi.IsoEntry;
pub const Cpu = acpi.Cpu;
/// The register map + sleep types discovery extracted, for logging/diagnostics.
pub fn powerInfo() PowerInfo {
@@ -41,6 +42,22 @@ pub fn amlStats() AmlStats {
return acpi.aml_stats;
}
/// The usable logical processors discovered during enumeration — one entry per
/// core danos may schedule on, each carrying the Local APIC ID an SMP wake targets.
/// `len` is the hardware's degree of parallelism: how many tasks *could* run at the
/// same instant once the application processors are started. Today only the
/// bootstrap processor is actually running, so starting the rest is the pending SMP
/// step (see docs/smp.md). Borrowed from static storage populated by `discover`.
pub fn cpus() []const Cpu {
return acpi.cpu_info.cpus[0..acpi.cpu_info.count];
}
/// Non-zero only if enumeration found more processors than the static pool holds
/// (the surplus were dropped from `cpus()`); surfaced so the cap is never silent.
pub fn cpusDropped() usize {
return acpi.cpu_info.dropped;
}
/// Enumerate hardware into a fresh device tree. `hal` supplies the hardware
/// primitives the backend needs (MMIO mapping for PCIe config space, port I/O for
/// ACPI registers); pass the arch implementation. Errors leave nothing to clean up
+37
View File
@@ -44,11 +44,15 @@ const spurious_vector = 47;
// LAPIC register offsets.
const reg_spurious = 0x0F0;
const reg_eoi = 0x0B0;
const reg_icr_low = 0x300; // interrupt command register, low dword (writing it sends)
const reg_icr_high = 0x310; // ICR high dword (destination APIC id in bits 24-31)
const reg_lvt_timer = 0x320;
const reg_timer_initial = 0x380;
const reg_timer_current = 0x390;
const reg_timer_divide = 0x3E0;
const icr_delivery_pending = 1 << 12; // ICR low bit 12: a previous IPI is still in flight
const lvt_masked = 1 << 16;
const lvt_periodic = 1 << 17;
const timer_divide_16 = 0x3;
@@ -120,6 +124,39 @@ pub fn init() void {
write(reg_spurious, 0x100 | spurious_vector); // bit 8 = software enable
}
/// Software-enable *this* core's Local APIC — the application-processor counterpart
/// of `init`, minus the one-time PIC remap (the BSP already masked it) and minus
/// calibration (the timer rate is a shared hardware constant, measured once). Each
/// core has its own LAPIC at the same MMIO address, so no per-core base is needed.
pub fn initSecondary() void {
const msr = io.rdmsr(ia32_apic_base_msr);
io.wrmsr(ia32_apic_base_msr, msr | (1 << 11)); // global enable
write(reg_spurious, 0x100 | spurious_vector); // software enable
}
// --- application-processor wakeup (INIT–SIPI–SIPI) --------------------------
/// Send an INIT IPI to the core with Local APIC id `apic_id` — the first step of
/// the wake sequence. Blocks until the LAPIC reports the IPI was delivered.
pub fn sendInit(apic_id: u32) void {
write(reg_icr_high, apic_id << 24);
write(reg_icr_low, 0x4500); // INIT, physical destination, assert, edge-triggered
waitIcrIdle();
}
/// Send a STARTUP IPI (SIPI) telling the target core to begin executing at physical
/// address `vector << 12` (in real mode). Per the Intel bring-up protocol this is
/// sent twice after the INIT; both calls block until delivery completes.
pub fn sendStartup(apic_id: u32, vector: u8) void {
write(reg_icr_high, apic_id << 24);
write(reg_icr_low, 0x4600 | @as(u32, vector)); // STARTUP with the page vector
waitIcrIdle();
}
fn waitIcrIdle() void {
while (read(reg_icr_low) & icr_delivery_pending != 0) {}
}
/// The calibration window: we time everything against a 10 ms reference interval.
const calib_ms = 10;
+67
View File
@@ -13,6 +13,7 @@ const serial = @import("serial.zig");
const apic = @import("apic.zig");
const ioapic = @import("ioapic.zig");
const io = @import("io.zig");
const smp = @import("smp.zig");
/// The saved register/trap frame passed to a fault handler.
pub const CpuState = idt.CpuState;
@@ -81,6 +82,64 @@ pub fn readCr3() u64 {
);
}
/// IA32_GS_BASE: the hidden base of the GS segment. We repurpose it as the per-CPU
/// data pointer (there's no user mode yet, so no `swapgs` dance — GS base is always
/// the running core's per-CPU block). Set once per core during bring-up, after the
/// GDT is loaded (loading a GS *selector* would otherwise clobber this base).
const ia32_gs_base = 0xC000_0101;
/// Publish this core's per-CPU data pointer so `cpuLocal` can retrieve it. Each
/// core calls this once, after its GDT is in place.
pub fn setCpuLocal(ptr: usize) void {
io.wrmsr(ia32_gs_base, ptr);
}
/// This core's per-CPU data pointer (the value `setCpuLocal` stored). Reads the GS
/// base MSR — a per-core register, so each core sees its own without any locking.
pub fn cpuLocal() usize {
return io.rdmsr(ia32_gs_base);
}
// --- SMP: application-processor bring-up ----------------------------------
/// Record the low (<1 MiB) frame reserved for the AP trampoline. Run once at boot.
/// The frame stays inert (zeroed, non-executable) between wakes and is armed only
/// while a core is climbing — so a core can be (re)woken at any time (retry, or a
/// future power manager) without leaving an executable page resident. See smp.zig.
pub fn setTrampolinePage(phys: u64) void {
smp.setTrampolinePage(phys);
}
/// Wake the core with Local APIC id `apic_id` as dense CPU `index`, giving it
/// `stack_top` and its per-CPU pointer `percpu`; it adopts the current (kernel) page
/// tables. Returns false if it doesn't come online within the timeout. Blocks until
/// the core reports in.
pub fn startSecondary(apic_id: u32, stack_top: usize, percpu: usize, index: usize) bool {
return smp.startAp(apic_id, stack_top, percpu, index, readCr3());
}
/// Register the generic entry a woken AP jumps to once its arch state is up (its own
/// descriptor tables, LAPIC, and timer). The kernel passes its scheduler entry here.
pub fn setSecondaryEntry(entry: *const fn () callconv(.c) noreturn) void {
smp.setSecondaryEntry(entry);
}
/// Test hook: force the next `n` AP wake attempts to fail, so the retry path can be
/// exercised deterministically (see the smp-retry test). No effect when `n` is 0.
pub fn testFailNextWakes(n: u32) void {
smp.testFailNextWakes(n);
}
/// The reserved AP-trampoline frame (0 if none). For tests that check it's inert.
pub fn trampolinePage() u64 {
return smp.trampolinePage();
}
/// Whether the page at `virt` is currently mapped executable (present, NX clear).
pub fn pageExecutable(virt: u64) bool {
return paging.isExecutable(virt);
}
/// Kernel tick rate: 1000 Hz (1 ms), the scheduler's time quantum.
pub const timer_hz = 1000;
@@ -282,3 +341,11 @@ pub fn readCr2() u64 {
pub fn halt() noreturn {
while (true) asm volatile ("hlt");
}
/// Spin-wait hint (`pause`). Emitted in the body of a spinlock's busy-wait: it
/// relaxes the core while it polls a contended lock — yielding pipeline resources
/// to a hyperthread sibling and easing the cache-coherency traffic on the lock
/// line. Purely a performance/power hint; correct to omit, but kinder on the bus.
pub fn cpuRelax() void {
asm volatile ("pause");
}
+33 -15
View File
@@ -3,18 +3,25 @@
//! reference a code selector — so we install our own flat GDT with known
//! selectors (0x08 kernel code, 0x10 kernel data) rather than trusting whatever
//! the firmware left in place.
//!
//! The code/data descriptors are identical on every core, but the **TSS descriptor
//! is per-core** (each core needs its own TSS — its own interrupt/fault stacks; see
//! tss.zig). Two cores can't share one TSS descriptor slot, so each core gets its
//! own copy of the table with its own TSS descriptor. Slot 0 is the BSP.
/// Selectors into the table below (index * 8).
/// Selectors into the table (index * 8). Same on every core's GDT.
pub const kernel_code = 0x08;
pub const kernel_data = 0x10;
pub const tss_selector = 0x18;
/// Flat 64-bit descriptors. Base/limit are ignored in long mode; what matters is
/// the access byte and, for code, the long-mode (L) flag.
const max_cpus = 64; // matches the scheduler / discovery pool
const entries = 5; // null, code, data, TSS-low, TSS-high
/// The shared descriptors (slots 0-2); slots 3-4 hold this core's TSS descriptor,
/// filled in per core by `setTssFor`.
/// code: present, ring 0, executable, readable, L=1 -> 0x00AF9A00_0000FFFF
/// data: present, ring 0, writable -> 0x00CF9200_0000FFFF
/// The last two slots hold one 16-byte TSS descriptor, filled in by setTss.
var table = [_]u64{
const template = [entries]u64{
0, // null descriptor (required)
0x00AF9A000000FFFF, // kernel code (0x08)
0x00CF92000000FFFF, // kernel data (0x10)
@@ -22,16 +29,20 @@ var table = [_]u64{
0, // TSS descriptor high
};
/// Fill the 64-bit TSS system descriptor (two GDT slots) so the task register can
/// point at our TSS. Type 0x89 = present, ring 0, available 64-bit TSS.
pub fn setTss(base: u64, limit: u64) void {
table[3] = (limit & 0xFFFF) |
/// One GDT per core (each a copy of the template, differing only in its TSS slot).
var gdts = [_][entries]u64{template} ** max_cpus;
/// Fill core `cpu`'s 64-bit TSS system descriptor (two GDT slots) so its task
/// register can point at its own TSS. Type 0x89 = present, ring 0, available 64-bit
/// TSS. Write it into that core's GDT before it loads the TSS selector.
pub fn setTssFor(cpu: usize, base: u64, limit: u64) void {
gdts[cpu][3] = (limit & 0xFFFF) |
((base & 0xFFFF) << 16) |
(((base >> 16) & 0xFF) << 32) |
(@as(u64, 0x89) << 40) |
(((limit >> 16) & 0xF) << 48) |
(((base >> 24) & 0xFF) << 56);
table[4] = (base >> 32) & 0xFFFFFFFF;
gdts[cpu][4] = (base >> 32) & 0xFFFFFFFF;
}
/// The operand `lgdt` wants: table byte-length minus one, then its address.
@@ -41,14 +52,21 @@ const Descriptor = packed struct {
};
/// Loads the GDT and reloads the segment registers (including CS). Defined in
/// isr.s — it uses the selectors 0x08 (code) and 0x10 (data) that match `table`.
/// isr.s — it uses the selectors 0x08 (code) and 0x10 (data) that match the table.
extern fn gdt_flush(descriptor: *const Descriptor) callconv(.c) void;
/// Install our GDT and switch onto its segments.
pub fn init() void {
/// Load core `cpu`'s GDT and switch onto its segments. Note this reloads the segment
/// registers, which zeroes the GS base — so a core must publish its per-CPU pointer
/// (setCpuLocal) *after* calling this.
pub fn loadOnThisCpu(cpu: usize) void {
const descriptor = Descriptor{
.limit = @sizeOf(@TypeOf(table)) - 1,
.base = @intFromPtr(&table),
.limit = @sizeOf([entries]u64) - 1,
.base = @intFromPtr(&gdts[cpu]),
};
gdt_flush(&descriptor);
}
/// Install the bootstrap processor's GDT (slot 0) and switch onto its segments.
pub fn init() void {
loadOnThisCpu(0);
}
+7
View File
@@ -128,6 +128,13 @@ pub fn init() void {
// Run the double-fault handler (vector 8) on IST1: a #DF usually means the
// current stack is unusable, so it needs a guaranteed-good one. See tss.zig.
idt[8].ist = tss.double_fault_ist;
loadOnThisCpu();
}
/// Load the (shared, already-populated) IDT on the current core. The gate table is
/// read-only after `init`, so every core points its IDTR at the same one. Called by
/// the BSP via `init` and by each AP during bring-up.
pub fn loadOnThisCpu() void {
const descriptor = Descriptor{
.limit = @sizeOf(@TypeOf(idt)) - 1,
.base = @intFromPtr(&idt),
+6 -1
View File
@@ -63,9 +63,14 @@ switch_context:
ret # return into the new task's saved instruction pointer
# task_trampoline: the first thing a freshly-spawned task runs. init_task_stack
# leaves its entry function in r15. New tasks start with interrupts enabled.
# leaves its entry function in r15. A fresh task is switched to with the big kernel
# lock held (the hand-off rule in sync.zig) but has no enter/leave frame of its own,
# so it releases the lock here before running its body. r15 survives the call (it's
# callee-saved). New tasks then start with interrupts enabled.
.extern releaseForFreshTask
.global task_trampoline
task_trampoline:
call releaseForFreshTask # drop the kernel lock we inherited across the switch
sti
call *%r15 # call the task entry (fn() void)
1: hlt # if the entry returns, idle (still preemptible)
+25
View File
@@ -128,6 +128,31 @@ pub fn map(virt: u64, phys: u64, writable_page: bool) void {
invalidate(virt);
}
/// Whether `virt` is currently mapped **executable** — present with the NX bit
/// clear. Walks the 4-level tables (all danos mappings are 4 KiB, so no huge-page
/// case). Returns false if unmapped. Used for W^X checks in tests.
pub fn isExecutable(virt: u64) bool {
const pml4e = tableAt(kernel_pml4)[(virt >> 39) & 0x1FF];
if (pml4e & present == 0) return false;
const pdpte = tableAt(pml4e & addr_mask)[(virt >> 30) & 0x1FF];
if (pdpte & present == 0) return false;
const pde = tableAt(pdpte & addr_mask)[(virt >> 21) & 0x1FF];
if (pde & present == 0) return false;
const pte = tableAt(pde & addr_mask)[(virt >> 12) & 0x1FF];
if (pte & present == 0) return false;
return pte & no_execute == 0;
}
/// Make an already-identity-mapped RAM page **executable** (clear its NX bit),
/// leaving it present and writable. The blanket RAM mapping is NX for W^X, but the
/// application processors fetch the AP trampoline from a low RAM page under paging —
/// so that one page must be executable. A deliberate, temporary W^X exception for a
/// single bring-up page; the caller frees it once every AP is up.
pub fn setExecutable(phys: u64) void {
mapPage(kernel_pml4, phys, phys, present | writable); // note: no no_execute
invalidate(phys);
}
/// Remove a mapping and flush it from the TLB.
pub fn unmap(virt: u64) void {
const pml4e = tableAt(kernel_pml4)[(virt >> 39) & 0x1FF];
+171
View File
@@ -0,0 +1,171 @@
//! Application-processor (AP) bring-up: waking the cores the firmware left parked.
//!
//! The firmware starts only the bootstrap processor (BSP); the others sit idle until
//! the kernel wakes them with an INIT–SIPI–SIPI sequence (Intel SDM Vol.3, "MP
//! Initialization"). A woken core begins in 16-bit real mode at a low physical page,
//! runs the [trampoline](trampoline.s) up into 64-bit long mode, and lands in
//! `apEntry` here. This module copies the trampoline into place, patches its
//! per-AP parameters, drives the wake IPIs, and waits for each core to report in.
//!
//! Cores are brought up **one at a time**: a single trampoline page and parameter
//! block are reused, so the BSP patches, wakes, and waits for one AP before the
//! next. That also lets `apEntry` pick up its dense CPU index from a plain global.
//! Once a core has its own descriptor tables, LAPIC, and timer, it calls the generic
//! scheduler entry and joins the run loop — mechanism here, policy there.
const io = @import("io.zig");
const gdt = @import("gdt.zig");
const tss = @import("tss.zig");
const idt = @import("idt.zig");
const apic = @import("apic.zig");
const paging = @import("paging.zig");
/// IA32_GS_BASE — the per-CPU data pointer (see cpu.zig; kept in sync here so the AP
/// path doesn't depend on cpu.zig and risk an import cycle).
const ia32_gs_base = 0xC000_0101;
const page_size = 0x1000;
/// Physical address of the low (<1 MiB) frame reserved for the trampoline. Held for
/// the life of the system so any core can be (re)woken on demand — a retry, or a
/// future power manager bringing a core back online. The frame is kept **inert**
/// between wakes (zeroed and non-executable) and only armed for the brief moment a
/// core is actually climbing. Its low 20 bits are zero, so `phys >> 12` is the SIPI
/// vector.
var tramp_phys: u64 = 0;
/// Set to 1 by a freshly-woken AP once it reaches `apEntry` and finishes its own
/// bring-up. The BSP clears it before each wake and polls it afterwards — a simple
/// one-at-a-time handshake (only one AP is being started at any moment).
var ap_alive: u32 = 0;
/// The dense CPU index of the AP currently being started. Set by the BSP before the
/// wake, read by `apEntry` (safe because bring-up is strictly one core at a time).
var boot_index: usize = 0;
/// The generic scheduler entry a woken core jumps to once its arch state is up. Set
/// by the kernel via `setSecondaryEntry`; never returns.
var secondary_entry: ?*const fn () callconv(.c) noreturn = null;
/// Register the generic entry an AP calls once its per-CPU tables/LAPIC/timer are up.
pub fn setSecondaryEntry(entry: *const fn () callconv(.c) noreturn) void {
secondary_entry = entry;
}
/// Test hook: force the next `n` wake attempts to fail (skipping the actual
/// INIT-SIPI-SIPI), so the retry path can be exercised deterministically. Zero in
/// normal operation — the smp-retry test arms it via `arch.testFailNextWakes`.
var fail_next_wakes: u32 = 0;
pub fn testFailNextWakes(n: u32) void {
fail_next_wakes = n;
}
/// Record the reserved low frame the trampoline uses. Call once at boot. The frame
/// starts inert (identity-mapped RW+NX like all RAM); each wake arms it and disarms
/// it again, so it's only ever executable while a core is climbing.
pub fn setTrampolinePage(phys: u64) void {
tramp_phys = phys;
}
/// The reserved trampoline frame (0 if SMP bring-up never ran). Exposed so a test
/// can verify it's inert — zeroed and non-executable — when dormant.
pub fn trampolinePage() u64 {
return tramp_phys;
}
/// Arm the trampoline for a wake: make its page executable (W^X exception for the
/// duration of the climb) and copy the blob in.
fn arm() void {
paging.setExecutable(tramp_phys);
const start = @extern([*]const u8, .{ .name = "ap_trampoline_start" });
const end = @extern([*]const u8, .{ .name = "ap_trampoline_end" });
const len = @intFromPtr(end) - @intFromPtr(start);
const dst: [*]u8 = @ptrFromInt(tramp_phys);
@memcpy(dst[0..len], start[0..len]);
}
/// Disarm after a wake: wipe the page and restore it to inert RW+NX, so no
/// executable code (nor any stale bytes) lingers between wakes. Safe to run once the
/// woken core has reported in — it's long past the trampoline by then, in the kernel
/// image; a core that never answered is dead and can't be mid-climb.
fn disarm() void {
const dst: [*]u8 = @ptrFromInt(tramp_phys);
@memset(dst[0..page_size], 0);
paging.map(tramp_phys, tramp_phys, true); // RW + NX, like every other RAM frame
}
/// Address of a patchable trampoline parameter, by symbol name: the copied blob's
/// base plus the field's offset within it (a same-section symbol difference). The
/// pointer is `align(1)` — the fields aren't 8-aligned within the blob, and x86
/// tolerates unaligned stores, so we don't force layout constraints on the asm.
fn param(comptime name: []const u8) *align(1) volatile u64 {
const start = @intFromPtr(@extern([*]const u8, .{ .name = "ap_trampoline_start" }));
const sym = @intFromPtr(@extern([*]const u8, .{ .name = name }));
return @ptrFromInt(tramp_phys + (sym - start));
}
/// Wake the core with Local APIC id `apic_id` as dense CPU `index`, hand it
/// `stack_top` and its per-CPU pointer `percpu`, and wait for it to come alive. This
/// is one self-contained attempt: it arms the trampoline, drives INIT–SIPI–SIPI, and
/// disarms again before returning — so it's safe to call repeatedly (a retry, or a
/// power manager re-waking a core; the INIT resets a core that was wedged). Returns
/// false if the core doesn't report in within the timeout (left parked, no harm to
/// the running system). `cr3` is the kernel page tables the AP adopts. Precondition:
/// `setTrampolinePage` has run.
pub fn startAp(apic_id: u32, stack_top: usize, percpu: usize, index: usize, cr3: u64) bool {
arm();
defer disarm();
if (fail_next_wakes > 0) { // test hook: simulate a core missing this attempt
fail_next_wakes -= 1;
return false;
}
boot_index = index;
param("ap_tramp_cr3").* = cr3;
param("ap_tramp_stack").* = stack_top;
param("ap_tramp_entry").* = @intFromPtr(&apEntry);
param("ap_tramp_percpu").* = percpu;
@atomicStore(u32, &ap_alive, 0, .seq_cst);
const vector: u8 = @intCast(tramp_phys >> 12);
apic.sendInit(apic_id);
delayMicros(10_000); // 10 ms INIT settle
apic.sendStartup(apic_id, vector);
delayMicros(200);
apic.sendStartup(apic_id, vector);
// Wait up to 100 ms for the AP to reach apEntry and set the flag.
const deadline = apic.millis() + 100;
while (apic.millis() < deadline) {
if (@atomicLoad(u32, &ap_alive, .acquire) != 0) return true;
asm volatile ("pause");
}
return false;
}
/// Busy-wait `us` microseconds against the calibrated TSC clock (the AP wake happens
/// after the timer is up, so the clock is available).
fn delayMicros(us: u64) void {
const start = apic.micros();
while (apic.micros() - start < us) asm volatile ("pause");
}
/// The 64-bit entry every AP lands on, called from the trampoline with its per-CPU
/// pointer in RDI. Brings up this core's own descriptor tables, LAPIC and timer,
/// signals the BSP, then jumps to the generic scheduler entry. Never returns.
fn apEntry(percpu: usize) callconv(.c) noreturn {
const cpu = boot_index;
gdt.loadOnThisCpu(cpu); // this core's GDT (with its own TSS slot)
tss.setupThisCpu(cpu); // this core's TSS + IST stack, loaded into TR
idt.loadOnThisCpu(); // the shared IDT
io.wrmsr(ia32_gs_base, percpu); // per-CPU pointer — *after* the GDT reload
apic.initSecondary(); // software-enable this core's LAPIC
apic.initTimer(apic.frequencyHz()); // arm its timer (still masked: interrupts off)
@atomicStore(u32, &ap_alive, 1, .release); // "arch state up" — BSP is polling this
if (secondary_entry) |enterScheduler| enterScheduler(); // joins the run loop
while (true) asm volatile ("hlt"); // (only if no entry was registered)
}
+150
View File
@@ -0,0 +1,150 @@
# AP trampoline: brings a waking application processor from the 16-bit real mode it
# starts in (after INIT-SIPI-SIPI) up through protected mode into 64-bit long mode,
# then jumps to the Zig AP entry (arch/x86_64/smp.zig:apEntry).
#
# A STARTUP IPI vectors a core to physical address `vector << 12` in real mode, so
# this blob is copied to a low (<1 MiB) page and started there; at entry CS = that
# page >> 4 and IP = 0. It is fully **position-independent**: it derives its own
# linear base (CS << 4) into EBX and addresses every internal datum as
# `(label - ap_trampoline_start)(%ebx)` — a difference of two symbols in the same
# section, which the assembler folds to a constant page offset no matter where the
# blob was linked or copied to. The BSP patches the parameter block (CR3, stack,
# entry, per-CPU pointer) before each wake; see arch/x86_64/smp.zig.
#
# It lives in .rodata (not .text): it is data to be copied out and executed
# elsewhere, never run at its link address, so it must not be a normal code segment.
.section .rodata
.balign 16
.code16
.global ap_trampoline_start
ap_trampoline_start:
cli
cld
# Linear base of this page (CS << 4) into EBX; all data is addressed off it.
xorl %eax, %eax
mov %cs, %ax
shll $4, %eax
movl %eax, %ebx
mov %cs, %ax # DS = CS, so we address our data as DS:(label - start):
mov %ax, %ds # the segment base (CS<<4) already supplies the page base,
# so data operands use the page *offset*, not EBX.
# Relocate the pointers whose absolute (linear) targets depend on where we were
# copied: the GDT base and the two far-jump targets = EBX + their page offsets.
# EBX supplies the base for the *value* (via leal); the store address is DS-rel.
leal (gdt32 - ap_trampoline_start)(%ebx), %eax
movl %eax, gdtr32_base - ap_trampoline_start
leal (prot_entry - ap_trampoline_start)(%ebx), %eax
movl %eax, jmp32_off - ap_trampoline_start
leal (long_entry - ap_trampoline_start)(%ebx), %eax
movl %eax, jmp64_off - ap_trampoline_start
lgdtl gdtr32 - ap_trampoline_start
movl %cr0, %eax # enter protected mode (CR0.PE)
orl $1, %eax
movl %eax, %cr0
ljmpl *(jmp32_ptr - ap_trampoline_start) # -> prot_entry, CS = 0x08
.code32
prot_entry:
movw $0x10, %ax # flat 32-bit data segments
movw %ax, %ds
movw %ax, %es
movw %ax, %ss
movw %ax, %fs
movw %ax, %gs
# CR4: PAE (required for long mode) + OSFXSR/OSXMMEXCPT. The kernel is built with
# SSE (part of the x86_64 baseline), and the compiler emits SSE for things as
# ordinary as a struct copy — without OSFXSR those instructions #UD. The BSP got
# these bits from UEFI; an AP starts fresh, so we must set them ourselves.
movl %cr4, %eax
orl $((1 << 5) | (1 << 9) | (1 << 10)), %eax
movl %eax, %cr4
# CR0: clear EM (no x87 emulation) and set MP, so SSE/x87 don't fault.
movl %cr0, %eax
andl $~(1 << 2), %eax # ~EM
orl $(1 << 1), %eax # MP
movl %eax, %cr0
movl (param_cr3 - ap_trampoline_start)(%ebx), %eax # kernel page tables
movl %eax, %cr3
movl $0xC0000080, %ecx # EFER: long mode enable (LME) + NX enable (NXE, since
rdmsr # the kernel's PTEs set the NX bit)
orl $((1 << 8) | (1 << 11)), %eax
wrmsr
movl %cr0, %eax # paging on (CR0.PG) — now in long mode (compat sub-mode)
orl $(1 << 31), %eax
movl %eax, %cr0
ljmpl *(jmp64_ptr - ap_trampoline_start)(%ebx) # -> long_entry, CS = 0x18 (L=1)
.code64
long_entry:
movw $0x10, %ax # sane flat data segments
movw %ax, %ds
movw %ax, %es
movw %ax, %ss
# RBX = EBX (zero-extended) = page base. Load our stack and per-CPU pointer, then
# call the Zig entry — which runs from the kernel image and never returns.
movq (param_stack - ap_trampoline_start)(%rbx), %rsp
movq (param_percpu - ap_trampoline_start)(%rbx), %rdi # SysV arg 0
movq (param_entry - ap_trampoline_start)(%rbx), %rax
callq *%rax
1: hlt # unreachable; guard against a stray return
jmp 1b
# --- data: GDT, far pointers, and the BSP-patched parameter block -----------
.balign 8
gdt32:
.quad 0x0000000000000000 # 0x00 null
.quad 0x00CF9A000000FFFF # 0x08 32-bit code (G, D, present, exec/read)
.quad 0x00CF92000000FFFF # 0x10 data (valid in 32- and 64-bit)
.quad 0x00AF9A000000FFFF # 0x18 64-bit code (L=1)
gdt32_end:
gdtr32:
.word gdt32_end - gdt32 - 1
gdtr32_base:
.long 0 # patched (16-bit code): linear base of gdt32
jmp32_ptr: # indirect far-jump operand: offset then selector
jmp32_off:
.long 0 # patched: linear address of prot_entry
.word 0x08 # 32-bit code selector
jmp64_ptr:
jmp64_off:
.long 0 # patched: linear address of long_entry
.word 0x18 # 64-bit code selector
# The parameter block, filled in by the BSP (smp.zig) before each STARTUP IPI. Global
# so the Zig side can locate each field as (symbol - ap_trampoline_start).
.global ap_tramp_cr3
.global ap_tramp_stack
.global ap_tramp_entry
.global ap_tramp_percpu
param_cr3:
ap_tramp_cr3:
.quad 0 # kernel PML4 physical address (CR3)
param_stack:
ap_tramp_stack:
.quad 0 # top of this AP's kernel stack
param_entry:
ap_tramp_entry:
.quad 0 # address of apEntry (the Zig AP entry)
param_percpu:
ap_tramp_percpu:
.quad 0 # this AP's per-CPU pointer (goes in GS base)
.global ap_trampoline_end
ap_trampoline_end:
+24 -9
View File
@@ -4,6 +4,10 @@
//! interrupted stack was. We use IST1 for the double-fault handler, so a fault
//! that happens *because* the current stack is unusable still lands on solid
//! ground instead of triple-faulting.
//!
//! Each core needs **its own TSS** (its own IST stack): two cores taking a fault at
//! once can't share one fault stack. So the TSS and its IST stack are per-core,
//! indexed by CPU number; slot 0 is the BSP.
const gdt = @import("gdt.zig");
@@ -30,19 +34,30 @@ const Tss = packed struct {
/// The IST slot (1-based, as the IDT gate encodes it) used for critical faults.
pub const double_fault_ist = 1;
var tss: Tss align(16) = .{};
const max_cpus = 64; // matches gdt.zig / the scheduler
const ist_stack_size = 16 * 1024;
/// Dedicated stack for IST1. Static so it needs no allocator and is always valid.
var ist1_stack: [16 * 1024]u8 align(16) = undefined;
/// One TSS per core, and one IST1 stack per core. Static, so they need no allocator
/// and are always valid. (max_cpus × 16 KiB of BSS for the IST stacks.)
var tss_table = [_]Tss{.{}} ** max_cpus;
var ist_stacks: [max_cpus][ist_stack_size]u8 align(16) = undefined;
/// Loads the task register with the TSS selector. Defined in isr.s.
extern fn load_tr(selector: u16) callconv(.c) void;
/// Point IST1 at its stack, publish the TSS through the GDT, and load it into the
/// task register. Requires the GDT to already be loaded (gdt.init first).
pub fn init() void {
tss.ist1 = @intFromPtr(&ist1_stack) + ist1_stack.len; // stacks grow down
tss.iomap_base = @sizeOf(Tss); // == limit: no I/O permission bitmap
gdt.setTss(@intFromPtr(&tss), @sizeOf(Tss) - 1);
/// Set up core `cpu`'s TSS: point IST1 at that core's stack, install the TSS
/// descriptor into that core's GDT, and load it into the task register. Requires the
/// core's GDT to already be loaded (gdt.loadOnThisCpu first).
pub fn setupThisCpu(cpu: usize) void {
const t = &tss_table[cpu];
t.* = .{};
t.ist1 = @intFromPtr(&ist_stacks[cpu]) + ist_stack_size; // stacks grow down
t.iomap_base = @sizeOf(Tss); // == limit: no I/O permission bitmap
gdt.setTssFor(cpu, @intFromPtr(t), @sizeOf(Tss) - 1);
load_tr(gdt.tss_selector);
}
/// Set up the bootstrap processor's TSS (slot 0). Requires gdt.init first.
pub fn init() void {
setupThisCpu(0);
}
+5 -5
View File
@@ -11,8 +11,8 @@
//! When user mode arrives, the same primitive carries messages across the
//! isolation boundary (with the payload copied between address spaces).
const arch = @import("arch");
const sched = @import("scheduler.zig");
const sync = @import("sync.zig");
/// A bounded blocking channel of `capacity` messages of type `T`.
pub fn Channel(comptime T: type, comptime capacity: usize) type {
@@ -28,7 +28,7 @@ pub fn Channel(comptime T: type, comptime capacity: usize) type {
/// Send a message, blocking while the channel is full.
pub fn send(self: *Self, msg: T) void {
const flags = arch.saveInterrupts();
const flags = sync.enter();
// Recheck the condition in a loop: a wakeup only means "try again"
// (another waiter may have taken the slot first).
while (self.count == capacity) sched.waitLocked(&self.not_full);
@@ -36,18 +36,18 @@ pub fn Channel(comptime T: type, comptime capacity: usize) type {
self.tail = (self.tail + 1) % capacity;
self.count += 1;
sched.wakeLocked(&self.not_empty); // a receiver can now proceed
arch.restoreInterrupts(flags);
sync.leave(flags);
}
/// Receive a message, blocking while the channel is empty.
pub fn recv(self: *Self) T {
const flags = arch.saveInterrupts();
const flags = sync.enter();
while (self.count == 0) sched.waitLocked(&self.not_empty);
const msg = self.buffer[self.head];
self.head = (self.head + 1) % capacity;
self.count -= 1;
sched.wakeLocked(&self.not_full); // a sender can now proceed
arch.restoreInterrupts(flags);
sync.leave(flags);
return msg;
}
};
+73 -5
View File
@@ -30,6 +30,11 @@ const cp_running = 0x70;
const cp_exception = 0xE0;
const cp_panic = 0xEE;
/// Physical address of the low page reserved at boot for the AP trampoline (0 = none
/// was available). Claimed right after the frame allocator comes up, before paging
/// and the heap consume the scarce sub-1 MiB frames.
var ap_trampoline_page: u64 = 0;
/// Kernel entry point. The bootloader jumps here after `ExitBootServices` with a
/// pointer to the handoff data. There is no runtime, no stack unwinding, and no
/// caller to return to, so this never returns.
@@ -98,6 +103,10 @@ fn kmain(boot_info: *const BootInfo) noreturn {
// Bring up the physical frame allocator over that map, and prove it works:
// allocate three frames, then hand them back.
pmm.init(boot_info.memory_map);
// Claim the AP trampoline's low (<1 MiB) page *now*, before paging and the heap
// draw down sub-1 MiB frames (the allocator scans upward from frame 0). Held
// until SMP bring-up; 0 means none was available (we stay uniprocessor).
ap_trampoline_page = pmm.allocBelow(0x100000) orelse 0;
const s1 = pmm.stats();
log.print("\ndanos: frame allocator online\n", .{});
log.print(" free frames: {d} ({d} MiB)\n", .{ s1.free_frames, mib(s1.free_frames) });
@@ -200,6 +209,10 @@ fn kmain(boot_info: *const BootInfo) noreturn {
log.write(" console UART: none in SPCR -> legacy COM1\n");
}
log.print(" ioapic : base 0x{x}, {d} inputs (masked); entry0 low 0x{x}\n", .{ ioapic_base, arch.ioapicEntryCount(), arch.ioapicEntryLow(0) });
const cores = platform.cpus();
log.print(" cpus : {d} usable core(s); 1 running (BSP), {d} AP(s) parked (SMP bring-up pending)\n", .{ cores.len, if (cores.len > 0) cores.len - 1 else 0 });
if (platform.cpusDropped() > 0)
log.print(" cpus : WARNING {d} core(s) beyond pool cap dropped\n", .{platform.cpusDropped()});
} else |err| {
log.print("\ndanos: device discovery failed: {s}\n", .{@errorName(err)});
}
@@ -217,6 +230,10 @@ fn kmain(boot_info: *const BootInfo) noreturn {
log.checkpoint(cp_timer);
log.print("danos: timer online ({d} Hz tick; LAPIC {d} MHz, TSC {d} MHz; calibrated via {s})\n", .{ arch.timer_hz, arch.lapicHz() / 1_000_000, arch.tscHz() / 1_000_000, arch.timerCalibrationSource() });
// Wake the other cores (application processors). A no-op on a single-core
// machine; on SMP each AP climbs to long mode and reports in (docs/smp.md).
bringUpSecondaries();
// In a test build (`zig build -Dtest-case=<name>`), run that case and stop.
// Normal builds fall through to the idle halt.
if (build_options.test_case) |case| {
@@ -234,6 +251,53 @@ fn kmain(boot_info: *const BootInfo) noreturn {
arch.halt();
}
/// Wake the application processors the firmware left parked. Allocates the low
/// trampoline page (and makes it executable), then wakes each non-boot core in turn,
/// handing it a fresh kernel stack and its per-CPU slot. Cores that don't report in
/// are left parked — the running system is unaffected. See docs/smp.md.
fn bringUpSecondaries() void {
const cores = platform.cpus();
if (cores.len <= 1) return;
// A low (<1 MiB) frame was reserved at boot for the real-mode trampoline (a SIPI
// vector addresses it). It's kept for the system's life — armed only during a
// wake, inert (zeroed, non-executable) otherwise — so cores can be re-woken later.
if (ap_trampoline_page == 0) {
log.write("danos: smp: no low page for the AP trampoline; staying uniprocessor\n");
return;
}
arch.setTrampolinePage(ap_trampoline_page);
arch.setSecondaryEntry(scheduler.secondaryMain); // where a woken core joins the run loop
// Test hook: the smp-retry case forces the first wake to fail, so the retry below
// must still bring every core online. Inert in a normal build (test_case is null).
if (build_options.test_case) |tc| {
if (std.mem.eql(u8, tc, "smp-retry")) arch.testFailNextWakes(1);
}
log.print("\ndanos: bringing up {d} application processor(s)\n", .{cores.len - 1});
const max_wake_attempts = 3; // a core that misses the first INIT-SIPI-SIPI gets retried
for (cores[1..], 1..) |core, index| {
const stack = heap.allocator().alloc(u8, 16 * 1024) catch {
log.print(" cpu apic_id {d}: no stack; skipped\n", .{core.apic_id});
continue;
};
const stack_top = (@intFromPtr(stack.ptr) + stack.len) & ~@as(usize, 15);
const pc = scheduler.prepareSecondary(index, core.apic_id);
var attempt: u32 = 1;
while (attempt <= max_wake_attempts) : (attempt += 1) {
if (arch.startSecondary(core.apic_id, stack_top, @intFromPtr(pc), index)) {
pc.online = true;
log.print(" cpu apic_id {d}: online (attempt {d})\n", .{ core.apic_id, attempt });
break;
}
if (attempt == max_wake_attempts)
log.print(" cpu apic_id {d}: no response after {d} attempts (parked)\n", .{ core.apic_id, max_wake_attempts });
}
}
log.print("danos: {d}/{d} cores online\n", .{ scheduler.onlineCount(), cores.len });
}
/// A user-facing status line: to the diagnostic `log` *and* the on-screen console
/// (if a framebuffer is present). The verbose log uses `log.*` directly and never
/// touches the framebuffer.
@@ -256,21 +320,25 @@ fn kib(frames: u64) u64 {
return frames * danos.page_size / (1024);
}
/// Report a CPU exception and halt. There's no fault recovery yet, so any
/// exception is terminal — but it reports what and where (to every output sink,
/// plus a POST code and a persistent breadcrumb) instead of silently resetting.
/// Report a CPU exception and halt **this core**. There's no fault recovery yet, so
/// the faulting core is terminal — but the fault is *contained* to it: on an
/// application processor only that core stops, and the rest of the system keeps
/// running (full recovery — kill the task, keep the core — is the resilience track,
/// see docs/resilience.md). The report names the core so an AP fault is attributed,
/// and goes to every output sink plus a POST code and a persistent breadcrumb.
fn onException(state: *const arch.CpuState) noreturn {
log.checkpoint(cp_exception);
const core = scheduler.currentCpuIndex();
// A fault is user-facing enough to paint on screen too (via statusPrint), on
// top of the diagnostic log.
statusPrint("\nCPU EXCEPTION: {s} (vector {d})\n", .{ arch.vectorName(state.vector), state.vector });
statusPrint("\nCPU EXCEPTION on core {d}: {s} (vector {d})\n", .{ core, arch.vectorName(state.vector), state.vector });
statusPrint(" error code : 0x{x}\n", .{state.error_code});
statusPrint(" RIP : 0x{x:0>16}\n", .{state.rip});
statusPrint(" RSP : 0x{x:0>16}\n", .{state.rsp});
if (state.vector == 14) statusPrint(" CR2 (addr) : 0x{x:0>16}\n", .{arch.readCr2()});
var buf: [128]u8 = undefined;
log.recordPanic(std.fmt.bufPrint(&buf, "CPU exception {s} (vector {d}) at RIP 0x{x}", .{ arch.vectorName(state.vector), state.vector, state.rip }) catch "cpu exception");
log.recordPanic(std.fmt.bufPrint(&buf, "CPU exception {s} (vector {d}) on core {d} at RIP 0x{x}", .{ arch.vectorName(state.vector), state.vector, core, state.rip }) catch "cpu exception");
arch.halt();
}
+18
View File
@@ -141,6 +141,24 @@ pub fn alloc() ?u64 {
return null; // out of physical memory
}
/// Allocate one free frame whose physical address is below `limit`, or null if
/// none is free down there. The AP trampoline needs this: an x86 STARTUP IPI vectors
/// a waking core to physical `vector << 12`, and `vector` is a byte — so the
/// trampoline must live under 1 MiB. A short linear scan of the low frames; only run
/// a handful of times at boot, so it needn't be fast.
pub fn allocBelow(limit: u64) ?u64 {
const cap = @min(total_frames, @as(usize, @intCast(limit / page_size)));
var f: usize = 1; // frame 0 stays reserved as the "none" address
while (f < cap) : (f += 1) {
if (!isUsed(f)) {
setUsed(f);
used_frames += 1;
return @as(u64, f) * page_size;
}
}
return null;
}
/// Return a frame obtained from alloc() to the pool. Bogus or double frees are
/// ignored rather than corrupting the count.
pub fn free(addr: u64) void {
+234 -64
View File
@@ -9,10 +9,18 @@
//! Switching happens both cooperatively (`yield`) and preemptively (the timer
//! calls `tick`). See docs/scheduling.md for the interrupt-flag discipline that
//! makes those two paths coexist.
//!
//! Cross-core safety is the **big kernel lock** (`sync.zig`): every critical
//! section here runs under it, and it is held across a context switch and released
//! by the task that resumes (see sync.zig's hand-off rule). On a single core the
//! lock is never contended, so the behaviour is exactly the old interrupt-flag
//! model; it's what lets a second core enter `schedule()` without corrupting the
//! shared queues.
const std = @import("std");
const arch = @import("arch");
const heap = @import("heap.zig");
const sync = @import("sync.zig");
/// Priority level: 0 (lowest) .. 7 (highest). 8 levels total.
pub const Priority = u3;
@@ -30,26 +38,70 @@ const Task = struct {
rsp: usize = 0, // saved stack pointer, valid while not running
stack: []u8 = &.{},
wake_at: u64 = 0, // uptime (ms) to wake a sleeping task; 0 = not sleeping
affinity: ?u32 = null, // null = runs on any core; else the index of its pinned core
next: ?*Task = null, // ready-queue link
};
var tasks = [_]Task{.{}} ** max_tasks;
var current: *Task = undefined;
var next_id: u32 = 1;
// Per-priority FIFO ready queues, and a bitmap of which levels are non-empty.
/// Per-CPU scheduler state: the task each core is running, its own idle task, and a
/// queue of tasks **pinned** to it. One entry per core; the arch layer stashes a
/// pointer to the *running* core's entry in the GS base, so `thisCpu()` fetches it
/// with a single read and no lock.
///
/// Most work stays in the **global** ready queue (below), which any idle core pulls
/// from — work-conserving. A task given an *affinity* instead goes to that core's
/// `pinned_*` queue and is only ever run there (no surprise migration — the more
/// real-time-predictable model, docs/smp.md). The two queues are merged at selection
/// time. Both are still mutated only under the big kernel lock, so one core enqueuing
/// into another core's pinned queue is safe.
pub const PerCpu = struct {
current: *Task = undefined, // the task running on this core
idle: *Task = undefined, // this core's idle task (always ready, lowest priority)
apic_id: u32 = 0, // the core's Local APIC id
index: u32 = 0, // dense 0-based core index
online: bool = false, // has this core finished bring-up?
// Tasks pinned to this core (affinity == index), per priority level + bitmap.
pinned_head: [num_priorities]?*Task = .{null} ** num_priorities,
pinned_tail: [num_priorities]?*Task = .{null} ** num_priorities,
pinned_bitmap: u8 = 0,
};
const max_cpus = 64; // matches the discovery pool (src/device/acpi.zig)
var cpus = [_]PerCpu{.{}} ** max_cpus;
/// This core's per-CPU state, via the arch layer's GS-base pointer. Valid only once
/// this core has run its scheduler bring-up (BSP in `init`, AP in `secondaryInit`).
inline fn thisCpu() *PerCpu {
return @ptrFromInt(arch.cpuLocal());
}
/// The task running on this core — the per-CPU replacement for the old global
/// `current`. A convenience reader; writes go through `thisCpu().current`.
inline fn cur() *Task {
return thisCpu().current;
}
// Per-priority FIFO ready queues, and a bitmap of which levels are non-empty. These
// are shared across all cores and mutated only under the big kernel lock.
var ready_head: [num_priorities]?*Task = .{null} ** num_priorities;
var ready_tail: [num_priorities]?*Task = .{null} ** num_priorities;
var ready_bitmap: u8 = 0;
var preemption_enabled = true;
/// Register the currently-running kernel context as the first task, spawn the
/// idle task, and hook the timer for preemption.
/// Bring up scheduling on the bootstrap processor: register the currently-running
/// kernel context as task 0, publish this core's per-CPU state (via the GS base),
/// give the core an idle task, and hook the timer for preemption. Runs once, at
/// boot, before interrupts are enabled — so no lock is needed here.
pub fn init(boot_priority: Priority) void {
const pc = &cpus[0];
pc.* = .{ .index = 0, .online = true };
arch.setCpuLocal(@intFromPtr(pc));
tasks[0] = .{ .id = 0, .state = .running, .priority = boot_priority };
current = &tasks[0];
spawn(idle, 0); // lowest priority, always runnable — runs when nothing else is
pc.current = &tasks[0];
pc.idle = create(idle, 0, null); // this core's idle task: always ready, lowest priority
arch.setTickHook(tick);
}
@@ -59,40 +111,128 @@ fn idle() void {
while (true) asm volatile ("hlt");
}
fn enqueue(t: *Task) void {
t.next = null;
const p: usize = t.priority;
if (ready_tail[p]) |tail| tail.next = t else ready_head[p] = t;
ready_tail[p] = t;
ready_bitmap |= levelBit(t.priority);
/// Reserve and initialise the per-CPU slot for an application processor at dense
/// `index` (1-based; 0 is the BSP) with Local APIC id `apic_id`, and return a
/// pointer the arch bring-up hands to the core (it publishes it in its GS base).
/// Called on the BSP before waking each AP; the AP marks itself `online`.
pub fn prepareSecondary(index: usize, apic_id: u32) *PerCpu {
const pc = &cpus[index];
pc.* = .{ .index = @intCast(index), .apic_id = apic_id, .online = false };
return pc;
}
fn dequeueHighest() ?*Task {
if (ready_bitmap == 0) return null;
const level: Priority = @intCast(num_priorities - 1 - @clz(ready_bitmap));
const t = ready_head[level].?;
ready_head[level] = t.next;
if (ready_head[level] == null) {
ready_tail[level] = null;
ready_bitmap &= ~levelBit(level);
/// Entry for an application processor once the arch layer has set up its per-CPU
/// tables, LAPIC, and timer. It turns this bring-up context into the core's idle task
/// (as task 0 is for the BSP), marks the core online, and enters the run loop: with
/// interrupts enabled the timer preempts this idle context into whatever the global
/// ready queue offers, so the core runs real work in parallel with the others. The
/// `.c` calling convention lets the arch trampoline path jump here. Never returns.
pub fn secondaryMain() callconv(.c) noreturn {
const flags = sync.enter();
const pc = thisCpu();
const t = freeSlot() orelse @panic("sched: task table full (AP idle task)");
t.* = .{ .id = next_id, .state = .running, .priority = 0 };
next_id += 1;
pc.current = t;
pc.idle = t;
pc.online = true;
sync.leave(flags);
arch.enableInterrupts(); // the timer now preempts this idle context into work
while (true) asm volatile ("hlt"); // idle when this core has nothing ready
}
/// Number of cores that have finished bring-up (the BSP plus every online AP).
pub fn onlineCount() usize {
var n: usize = 0;
for (&cpus) |*pc| {
if (pc.online) n += 1;
}
return n;
}
/// Make `t` ready. A pinned task (affinity set) goes to that core's pinned queue;
/// everything else goes to the shared global queue.
fn enqueue(t: *Task) void {
if (t.affinity) |cpu| {
const pc = &cpus[cpu];
enqueueTo(&pc.pinned_head, &pc.pinned_tail, &pc.pinned_bitmap, t);
} else {
enqueueTo(&ready_head, &ready_tail, &ready_bitmap, t);
}
}
fn enqueueTo(head: *[num_priorities]?*Task, tail: *[num_priorities]?*Task, bitmap: *u8, t: *Task) void {
t.next = null;
const p: usize = t.priority;
if (tail[p]) |tl| tl.next = t else head[p] = t;
tail[p] = t;
bitmap.* |= @as(u8, 1) << t.priority;
}
/// The highest non-empty priority level in a bitmap, or -1 if empty.
fn topLevel(bitmap: u8) i32 {
if (bitmap == 0) return -1;
return @as(i32, num_priorities - 1) - @as(i32, @clz(bitmap));
}
/// Pick the highest-priority ready task for core `pc`: the better of the global queue
/// and this core's pinned queue. Still O(1) (two `clz` and a compare). A pinned task
/// wins an equal-priority tie, so it can't be starved by global work at its level.
fn dequeueHighest(pc: *PerCpu) ?*Task {
const g = topLevel(ready_bitmap);
const p = topLevel(pc.pinned_bitmap);
if (g < 0 and p < 0) return null;
if (p >= g) return dequeueFrom(&pc.pinned_head, &pc.pinned_tail, &pc.pinned_bitmap, @intCast(p));
return dequeueFrom(&ready_head, &ready_tail, &ready_bitmap, @intCast(g));
}
fn dequeueFrom(head: *[num_priorities]?*Task, tail: *[num_priorities]?*Task, bitmap: *u8, level: usize) ?*Task {
const t = head[level].?;
head[level] = t.next;
if (head[level] == null) {
tail[level] = null;
bitmap.* &= ~(@as(u8, 1) << @intCast(level));
}
t.next = null;
return t;
}
fn levelBit(p: Priority) u8 {
return @as(u8, 1) << p;
/// Create a task that runs `entry` at `priority`, runnable on any core. It becomes
/// ready immediately. Takes the kernel lock: it mutates the shared task table and
/// ready queues and allocates from the (non-thread-safe) heap, so on SMP it must be
/// serialised.
pub fn spawn(entry: *const fn () void, priority: Priority) void {
const flags = sync.enter();
_ = create(entry, priority, null);
sync.leave(flags);
}
/// Create a task that runs `entry` at `priority`. It becomes ready immediately.
pub fn spawn(entry: *const fn () void, priority: Priority) void {
/// Like `spawn`, but **pins** the task to core `cpu` — it will only ever run there.
/// Returns true if pinned; false if `cpu` isn't a valid, online core, in which case
/// the task is still created but left unpinned (so it runs *somewhere* rather than
/// stranding in a queue no core services). Callers that require the pin (e.g. tests)
/// should check the result.
pub fn spawnOn(entry: *const fn () void, priority: Priority, cpu: u32) bool {
const flags = sync.enter();
defer sync.leave(flags);
const ok = cpu < max_cpus and cpus[cpu].online;
_ = create(entry, priority, if (ok) cpu else null);
return ok;
}
/// The unlocked task-creation primitive. Caller must hold the kernel lock (or be the
/// single-threaded boot path). `affinity` pins the task to a core (null = any).
/// Returns the new task so a core can keep a handle to its idle task.
fn create(entry: *const fn () void, priority: Priority, affinity: ?u32) *Task {
const t = freeSlot() orelse @panic("sched: task table full");
const stack = heap.allocator().alloc(u8, stack_size) catch @panic("sched: no memory for task stack");
t.* = .{ .id = next_id, .state = .ready, .priority = priority, .stack = stack };
t.* = .{ .id = next_id, .state = .ready, .priority = priority, .stack = stack, .affinity = affinity };
next_id += 1;
const top = @intFromPtr(stack.ptr) + stack.len;
t.rsp = arch.initTaskStack(top, @intFromPtr(entry));
enqueue(t);
return t;
}
fn freeSlot() ?*Task {
@@ -102,38 +242,43 @@ fn freeSlot() ?*Task {
return null;
}
/// Pick the highest-priority ready task and switch to it. Interrupts must be
/// disabled by the caller.
/// Pick the highest-priority ready task and switch this core to it. The big kernel
/// lock must be held by the caller (which also keeps local interrupts disabled);
/// it serialises every core's scheduling, so no other core can touch the shared
/// queues while we requeue `prev` and dequeue `next`. A dequeued task is `.ready`,
/// never running elsewhere, so two cores never run the same task.
fn schedule() void {
const prev = current;
const pc = thisCpu();
const prev = pc.current;
if (prev.state == .running) {
prev.state = .ready;
enqueue(prev); // back of its level's queue (round-robin)
}
const next = dequeueHighest() orelse {
const next = dequeueHighest(pc) orelse {
prev.state = .running; // nothing else ready — keep running
return;
};
next.state = .running;
current = next;
pc.current = next;
if (next != prev) arch.switchContext(&prev.rsp, next.rsp);
}
/// Voluntarily give up the CPU to the next ready task.
pub fn yield() void {
const flags = arch.saveInterrupts();
const flags = sync.enter();
schedule();
arch.restoreInterrupts(flags);
sync.leave(flags);
}
/// Block the current task for `ms` milliseconds, then let it become runnable
/// again. The idle task (or other work) runs in the meantime.
pub fn sleep(ms: u64) void {
const flags = arch.saveInterrupts();
current.wake_at = arch.millis() + ms;
current.state = .blocked;
const flags = sync.enter();
const t = cur();
t.wake_at = arch.millis() + ms;
t.state = .blocked;
schedule(); // current is blocked, so schedule() won't re-enqueue it
arch.restoreInterrupts(flags);
sync.leave(flags);
}
// --- event-based blocking -------------------------------------------------
@@ -147,27 +292,29 @@ pub const WaitQueue = struct {
head: ?*Task = null,
};
/// Block the current task on `wq` and switch away. Precondition: interrupts are
/// disabled (the caller holds them, so a condition can be checked and the block
/// committed atomically). On return — when woken — interrupts are still disabled.
/// Block the current task on `wq` and switch away. Precondition: the big kernel
/// lock is held (so a condition can be checked and the block committed atomically;
/// it also keeps local interrupts disabled). On return — when woken — the lock is
/// still held.
pub fn waitLocked(wq: *WaitQueue) void {
current.state = .blocked;
current.next = wq.head;
wq.head = current;
const t = cur();
t.state = .blocked;
t.next = wq.head;
wq.head = t;
schedule();
}
/// Move the highest-priority waiter on `wq` (if any) to the ready queue.
/// Precondition: interrupts disabled. Does not preempt — the caller decides.
/// Precondition: the big kernel lock is held. Does not preempt — the caller decides.
pub fn wakeLocked(wq: *WaitQueue) void {
// Find the highest-priority waiter (bounded scan) and unlink it.
var best_prev: ?*Task = null;
var best: ?*Task = null;
var prev: ?*Task = null;
var cur = wq.head;
while (cur) |t| : ({
var node = wq.head;
while (node) |t| : ({
prev = t;
cur = t.next;
node = t.next;
}) {
if (best == null or t.priority > best.?.priority) {
best = t;
@@ -182,25 +329,31 @@ pub fn wakeLocked(wq: *WaitQueue) void {
/// Block on `wq` (a self-contained critical section).
pub fn wait(wq: *WaitQueue) void {
const flags = arch.saveInterrupts();
const flags = sync.enter();
waitLocked(wq);
arch.restoreInterrupts(flags);
sync.leave(flags);
}
/// Wake the highest-priority waiter on `wq`, preempting if it outranks us.
pub fn wake(wq: *WaitQueue) void {
const flags = arch.saveInterrupts();
const flags = sync.enter();
const pc = thisCpu();
wakeLocked(wq);
// If a higher-priority task is now ready, run it immediately.
if (highestReadyPriority()) |p| {
if (p > current.priority) schedule();
// If a task this core would now pick outranks the running one, run it at once.
// (A waiter pinned to *another* core isn't counted — that core picks it up on its
// next tick; this core doesn't preempt for work it can't run.)
if (highestReadyPriority(pc)) |p| {
if (p > pc.current.priority) schedule();
}
arch.restoreInterrupts(flags);
sync.leave(flags);
}
fn highestReadyPriority() ?Priority {
if (ready_bitmap == 0) return null;
return @intCast(num_priorities - 1 - @clz(ready_bitmap));
/// The highest-priority task core `pc` could run right now — the better of the global
/// queue and this core's pinned queue — or null if it would fall back to idle.
fn highestReadyPriority(pc: *PerCpu) ?Priority {
const top = @max(topLevel(ready_bitmap), topLevel(pc.pinned_bitmap));
if (top < 0) return null;
return @intCast(top);
}
/// Wake any sleeping task whose deadline has passed. Bounded by the task count,
@@ -217,10 +370,15 @@ fn wakeExpired() void {
}
/// Called from the timer interrupt (interrupts already disabled): wake due
/// sleepers, then preempt.
/// sleepers, then preempt. Takes the kernel lock like any other critical section,
/// but releases it *without* touching the interrupt flag — the handler's `iretq`
/// restores the interrupted context's flags, so re-enabling here would open a
/// nested-interrupt window before the return.
pub fn tick() void {
_ = sync.enter();
wakeExpired();
if (preemption_enabled) schedule();
sync.leaveIsr();
}
/// Enable or disable timer-driven preemption (cooperative-only when off).
@@ -229,23 +387,35 @@ pub fn setPreemption(enabled: bool) void {
}
/// End the current task and switch away for good; never returns. The task's stack
/// is leaked for now (no reaper yet).
/// is leaked for now (no reaper yet). Acquires the kernel lock and hands it off to
/// the task we switch into (which releases it) — this frame never returns to leave.
pub fn exit() noreturn {
arch.disableInterrupts();
current.state = .free;
const next = dequeueHighest() orelse @panic("sched: no task left to run");
_ = sync.enter();
const pc = thisCpu();
pc.current.state = .free;
const next = dequeueHighest(pc) orelse @panic("sched: no task left to run");
next.state = .running;
current = next;
pc.current = next;
var discard: usize = 0;
arch.switchContext(&discard, next.rsp);
unreachable;
}
pub fn currentId() u32 {
return current.id;
return cur().id;
}
/// The dense index of the core this task is currently running on (0 = BSP). Reads
/// per-CPU state, so a task calling it on different cores sees different values —
/// which is how a test can prove work is running in parallel. Returns 0 if the GS
/// base isn't published yet (a fault in very early boot, before `init`), so a fault
/// reporter can call it unconditionally without a second fault.
pub fn currentCpuIndex() u32 {
if (arch.cpuLocal() == 0) return 0;
return thisCpu().index;
}
/// Change the running task's priority (takes effect next time it's enqueued).
pub fn setPriority(p: Priority) void {
current.priority = p;
cur().priority = p;
}
+80
View File
@@ -0,0 +1,80 @@
//! The big kernel lock (BKL) — the coarse mutual exclusion that lets more than one
//! CPU run kernel code safely.
//!
//! Until SMP, the kernel's mutual exclusion *was* the interrupt flag: a critical
//! section did `cli`, and since only one core existed, nothing else could touch
//! kernel state (the discipline in docs/scheduling.md). That invariant dies the
//! instant a second core runs kernel code — `cli` on one core does nothing to
//! another. So the kernel's shared state (the scheduler queues, IPC channels) is
//! guarded by a spinlock, and the lock is **always held with local interrupts
//! disabled**, so a core's own timer interrupt can't re-enter the kernel and
//! deadlock against the lock it already holds.
//!
//! This is deliberately *one coarse lock*, not many fine ones: it's philosophically
//! aligned with a tiny kernel and it keeps the single-core correctness model
//! (docs/scheduling.md) largely intact — one lock around kernel entry instead of
//! rethinking every critical section. It's the first-design choice seL4 makes and
//! docs/smp.md endorses; per-core run queues + fine-grained locking come later, if
//! contention ever bites. Because the kernel does little, the lock is held briefly.
//!
//! **The hand-off rule.** The lock is held *across* a context switch and released
//! by whichever task resumes, not by the one that switched away. A task that blocks
//! or yields calls `enter`, mutates the queues, `schedule()`s — switching to another
//! task *with the lock still held* — and only calls `leave` once it is eventually
//! resumed and its critical section runs to the end. So every call into `schedule()`
//! (and thus `switch_context`) happens with the lock held, and every task resumes
//! from a switch holding it. A freshly-spawned task has no `enter`/`leave` frame to
//! resume into, so `task_trampoline` releases the lock explicitly on its behalf via
//! `releaseForFreshTask` before running the task body.
const std = @import("std");
const arch = @import("arch");
/// 0 = free, 1 = held. A single global lock for the whole kernel.
var held = std.atomic.Value(u32).init(0);
/// Enter the kernel: disable interrupts on this core, then spin until we own the
/// lock. Returns the caller's prior interrupt flags for `leave` to restore.
/// Interrupts stay off for the whole critical section so this core's timer tick
/// can't try to re-acquire the lock we're holding.
pub fn enter() u64 {
const flags = arch.saveInterrupts();
acquire();
return flags;
}
/// Release the lock and restore the interrupt flags `enter` returned (re-enabling
/// interrupts only if they were on beforehand). The normal exit for a critical
/// section reached from task context (`yield`, `sleep`, `wait`, `wake`, IPC).
pub fn leave(flags: u64) void {
release();
arch.restoreInterrupts(flags);
}
/// Release the lock but leave interrupts as they are. The exit for a critical
/// section running inside an interrupt handler (the timer `tick`): the handler's
/// `iretq` is what restores the interrupted context's flags, so restoring them
/// here too would open a nested-interrupt window before the return. Release only.
pub fn leaveIsr() void {
release();
}
/// Release the lock on behalf of a freshly-spawned task. Such a task is switched to
/// (with the lock held) but has no `enter`/`leave` frame of its own to release
/// through — `task_trampoline` calls this before running the task body. Interrupts
/// are enabled separately by the trampoline. Exported for the assembly trampoline.
export fn releaseForFreshTask() callconv(.c) void {
release();
}
fn acquire() void {
// Test-and-test-and-set: try once, then spin read-only until the lock looks
// free before retrying the (bus-locked) swap — cheaper on the coherency fabric.
while (held.swap(1, .acquire) != 0) {
while (held.load(.monotonic) != 0) arch.cpuRelax();
}
}
fn release() void {
held.store(0, .release);
}
+306
View File
@@ -50,6 +50,10 @@ fn result() void {
pub fn run(case: []const u8, boot_info: *const BootInfo) void {
if (eql(case, "smoke")) {
smoke(boot_info);
} else if (eql(case, "discovery")) {
discoveryTest();
} else if (eql(case, "wx")) {
wxTest();
} else if (eql(case, "timer")) {
timer();
} else if (eql(case, "clock")) {
@@ -68,12 +72,22 @@ pub fn run(case: []const u8, boot_info: *const BootInfo) void {
eventTest();
} else if (eql(case, "ipc")) {
ipcTest();
} else if (eql(case, "smp")) {
smpTest();
} else if (eql(case, "affinity")) {
affinityTest();
} else if (eql(case, "smp-stress")) {
stressTest();
} else if (eql(case, "smp-retry")) {
smpRetryTest();
} else if (eql(case, "fault-ud")) {
faultInvalidOpcode();
} else if (eql(case, "fault-pf")) {
faultPageFault();
} else if (eql(case, "fault-df")) {
faultDoubleFault();
} else if (eql(case, "fault-ap-df")) {
faultApTest();
} else if (eql(case, "fault-nx")) {
faultNoExecute();
} else if (eql(case, "fault-null")) {
@@ -165,6 +179,54 @@ fn timer() void {
result();
}
/// Verify device discovery populated the platform facts the rest of the kernel
/// depends on — the results ACPI parsing stashed in globals at boot. These are
/// stable for the QEMU q35 + OVMF machine the harness runs, and span the tables:
/// MADT (LAPIC base, CPU count), FADT (PM/reset registers), and the AML parse
/// (the sleep type, plus the integrity check that every byte was consumed).
fn discoveryTest() void {
log("DANOS-TEST-BEGIN: discovery\n", .{});
const pinfo = platform.platformInfo();
const pw = platform.powerInfo();
const am = platform.amlStats();
check("LAPIC base discovered (MADT)", pinfo.lapic_base == 0xFEE00000);
check("ACPI PM timer found (FADT)", pinfo.pm_timer.present());
check("PM1a control register found (FADT)", pw.pm1a_cnt.present());
check("reset register supported (FADT)", pw.reset_supported);
check("S5 sleep type found (AML)", pw.s5 != null);
check("AML parsed completely (consumed == total)", am.total > 0 and am.consumed == am.total);
check("at least one CPU enumerated (MADT)", platform.cpus().len >= 1);
result();
}
/// Audit the W^X invariant across the memory classes: kernel code must be
/// executable, everything else must not be. `arch.pageExecutable` reads the leaf
/// page-table entry's NX bit, so this guards the permission overlay in paging.zig —
/// a broader check than `fault-nx`, which only exercises one data page.
fn wxTest() void {
log("DANOS-TEST-BEGIN: wx\n", .{});
check("kernel code is executable (R+X)", arch.pageExecutable(@intFromPtr(&wxTest)));
const ro = "danos-wx-probe"; // string literal -> .rodata
check("rodata is non-executable (NX)", !arch.pageExecutable(@intFromPtr(ro.ptr)));
check("kernel data is non-executable (NX)", !arch.pageExecutable(@intFromPtr(&passed)));
if (heap.allocator().alloc(u8, 64) catch null) |h| {
check("heap is non-executable (NX)", !arch.pageExecutable(@intFromPtr(h.ptr)));
heap.allocator().free(h);
}
var local: u64 = 0;
_ = &local;
check("stack is non-executable (NX)", !arch.pageExecutable(@intFromPtr(&local)));
result();
}
/// Verify the on-demand VMM: map a fresh frame at an unused virtual address, and
/// check it's writable and reads back.
fn vmm() void {
@@ -437,6 +499,209 @@ fn sleepTest() void {
result();
}
// --- SMP parallelism ------------------------------------------------------
var seen_core = [_]bool{false} ** 8;
var smp_running: bool = true;
/// A worker that, while running, records which core it's executing on. Spread across
/// spawned workers and idle APs, these should land on more than one core.
fn smpWorker() void {
const p: *volatile bool = &smp_running;
while (p.*) {
const c = sched.currentCpuIndex();
if (c < seen_core.len) seen_core[c] = true;
}
sched.exit();
}
/// Prove tasks run **in parallel** on multiple cores (not just interleaved on one).
/// Spawn several CPU-bound workers; each stamps the core it runs on into `seen_core`.
/// With the application processors online, more than one core should show up — which
/// can only happen if work is genuinely running at the same time on different cores.
/// (Run with QEMU `-smp N`; on a single core this would see just one and fail.)
fn smpTest() void {
log("DANOS-TEST-BEGIN: smp\n", .{});
seen_core = .{false} ** 8;
smp_running = true;
var i: usize = 0;
while (i < 4) : (i += 1) sched.spawn(smpWorker, 4);
// Let the workers run across cores for a stretch of real time.
var spins: u64 = 0;
while (spins < 2_000_000_000) spins +%= 1;
smp_running = false;
var cores_seen: u32 = 0;
for (seen_core) |s| {
if (s) cores_seen += 1;
}
log("DANOS-SMP: workers ran on {d} distinct core(s)\n", .{cores_seen});
check("tasks ran on multiple cores in parallel", cores_seen >= 2);
// Bring-up is done, so the trampoline frame must be inert: zeroed (no stale code)
// and non-executable (W^X restored). It's armed only while a core is climbing.
const tramp = arch.trampolinePage();
check("trampoline frame reserved", tramp != 0);
if (tramp != 0) {
const bytes: [*]const u8 = @ptrFromInt(tramp);
var zeroed = true;
for (0..4096) |b| {
if (bytes[b] != 0) zeroed = false;
}
check("trampoline page zeroed when dormant", zeroed);
check("trampoline page non-executable when dormant", !arch.pageExecutable(tramp));
}
result();
}
// --- affinity: a pinned task never migrates -------------------------------
var affinity_cores = [_]bool{false} ** 8;
var affinity_running: bool = true;
fn affinityWorker() void {
const p: *volatile bool = &affinity_running;
while (p.*) {
const c = sched.currentCpuIndex();
if (c < affinity_cores.len) affinity_cores[c] = true;
}
sched.exit();
}
/// A task pinned to a core must run **only** on that core. Pin a busy worker to
/// core 1 and let it run through many preemptions; it must have stamped core 1 and no
/// other. An *unpinned* task scatters across cores (that's what the smp test shows),
/// so a broken pin fails this deterministically — over this many time slices a
/// free-floating task will land on some other core.
fn affinityTest() void {
log("DANOS-TEST-BEGIN: affinity\n", .{});
affinity_cores = .{false} ** 8;
affinity_running = true;
if (!sched.spawnOn(affinityWorker, 4, 1)) {
check("worker pinned to core 1 (run with -smp)", false);
result();
return;
}
var spins: u64 = 0;
while (spins < 3_000_000_000) spins +%= 1; // many time slices across the cores
affinity_running = false;
var settle: u64 = 0;
while (settle < 200_000_000) settle +%= 1; // let the worker see the flag and exit
var others: u32 = 0;
for (affinity_cores, 0..) |seen, c| {
if (seen and c != 1) others += 1;
}
log("DANOS-AFFINITY: pinned worker touched core 1={}, other cores={d}\n", .{ affinity_cores[1], others });
check("pinned task ran on its core (1)", affinity_cores[1]);
check("pinned task never migrated to another core", others == 0);
result();
}
// --- SMP stress: hammer the big kernel lock across cores ------------------
const stress_pairs = 4; // producer/consumer pairs (8 tasks; fits the 16-task pool)
const stress_msgs = 100_000; // messages per pair
const stress_cap = 4; // small channel -> constant block/wake, more lock churn
var stress_chan = [_]ipc.Channel(u64, stress_cap){.{}} ** stress_pairs;
var stress_recv = [_]u64{0} ** stress_pairs; // messages received per pair
var stress_order_ok = [_]bool{true} ** stress_pairs; // FIFO order held per pair
var stress_cores = [_]bool{false} ** 8; // cores that ran a consumer
var stress_prod_claim: usize = 0;
var stress_cons_claim: usize = 0;
fn stressProducer() void {
// Claim a unique pair index (atomic: producers start on different cores).
const idx = @atomicRmw(usize, &stress_prod_claim, .Add, 1, .monotonic);
var v: u64 = 1;
while (v <= stress_msgs) : (v += 1) stress_chan[idx].send(v);
sched.exit();
}
fn stressConsumer() void {
const idx = @atomicRmw(usize, &stress_cons_claim, .Add, 1, .monotonic);
var expected: u64 = 1;
while (expected <= stress_msgs) : (expected += 1) {
const got = stress_chan[idx].recv();
if (got != expected) stress_order_ok[idx] = false; // lost/reordered => lock broke
const c = sched.currentCpuIndex();
if (c < stress_cores.len) stress_cores[c] = true;
stress_recv[idx] = expected;
}
sched.exit();
}
/// Stress the big kernel lock under sustained cross-core contention. Each pair drives
/// `stress_msgs` sequenced messages through a 4-slot channel — every send and recv
/// takes the lock, and the small buffer forces constant block/wake (so the scheduler
/// churns too). A single-producer/single-consumer channel must deliver in strict FIFO
/// order; if the lock let two cores into a critical section at once, the ring buffer
/// corrupts and the consumer sees a wrong or out-of-order value (or the run hangs /
/// faults). Passing means ~320k lock acquisitions across the cores stayed consistent.
fn stressTest() void {
log("DANOS-TEST-BEGIN: smp-stress\n", .{});
stress_chan = [_]ipc.Channel(u64, stress_cap){.{}} ** stress_pairs;
stress_recv = [_]u64{0} ** stress_pairs;
stress_order_ok = [_]bool{true} ** stress_pairs;
stress_cores = [_]bool{false} ** 8;
stress_prod_claim = 0;
stress_cons_claim = 0;
var i: usize = 0;
while (i < stress_pairs) : (i += 1) sched.spawn(stressConsumer, 4);
i = 0;
while (i < stress_pairs) : (i += 1) sched.spawn(stressProducer, 4);
// Drop below the workers so they get the cores; wake periodically to check for
// completion. A broken lock instead hangs here (harness timeout) or faults.
sched.setPriority(1);
var spins: u64 = 0;
while (spins < 40_000_000_000) : (spins += 1) {
var done = true;
for (stress_recv) |n| {
if (n < stress_msgs) done = false;
}
if (done) break;
}
sched.setPriority(4);
var total: u64 = 0;
for (stress_recv) |n| total += n;
var order_ok = true;
for (stress_order_ok) |ok| {
if (!ok) order_ok = false;
}
var cores: u32 = 0;
for (stress_cores) |s| {
if (s) cores += 1;
}
log("DANOS-STRESS: {d}/{d} pairs complete on {d} cores\n", .{ total, @as(u64, stress_pairs) * stress_msgs, cores });
check("every message delivered", total == @as(u64, stress_pairs) * stress_msgs);
check("strict FIFO order held (no lock corruption)", order_ok);
check("contention was genuinely cross-core", cores >= 2);
result();
}
/// Retry: `main` forced the first AP wake attempt to fail (arch.testFailNextWakes),
/// so a core missed its first INIT-SIPI-SIPI. The boot retry must have brought it back
/// anyway — every enumerated core should be online. If retry were broken, that core
/// would be parked and the count would fall short.
fn smpRetryTest() void {
log("DANOS-TEST-BEGIN: smp-retry\n", .{});
const total = platform.cpus().len;
const online = sched.onlineCount();
log("DANOS-RETRY: {d}/{d} cores online after a forced first-wake failure\n", .{ online, total });
check("multiple cores enumerated (run with -smp)", total >= 2);
check("retry brought every core online despite a failed first wake", online == total);
result();
}
fn faultInvalidOpcode() void {
log("DANOS-TEST-BEGIN: fault-ud\n", .{});
asm volatile ("ud2");
@@ -489,3 +754,44 @@ fn faultDoubleFault() void {
);
bad_sp += 0;
}
var ap_reached_fault: bool = false;
/// A task that faults with a #DF *on whatever core it's pinned to*. Announces the
/// core, then triggers the same double fault as `faultDoubleFault` — which is only
/// survivable on IST1, so it exercises that core's own TSS.
fn apDoubleFaultTask() void {
log("DANOS-AP: task running on core {d}, triggering #DF\n", .{sched.currentCpuIndex()});
@atomicStore(bool, &ap_reached_fault, true, .release);
arch.disableInterrupts();
var bad_sp: u64 = 0x5000000000;
asm volatile (
\\mov %[sp], %%rsp
\\ud2
:
: [sp] "r" (bad_sp),
: .{ .memory = true }
);
bad_sp += 0;
}
/// Fault on an application processor. Pins a double-faulting task to core 1, so the
/// fault is taken and handled by *that core's own* IDT and TSS/IST — not the BSP's.
/// The harness matches "core N: double fault (vector 8)" with N ≥ 1, which can only
/// appear if the AP caught the #DF on its IST1 (a broken per-core TSS would
/// triple-fault and reset instead). We then show the BSP still runs afterwards, so
/// the fault was *contained* to the AP, not fatal to the system.
fn faultApTest() void {
log("DANOS-TEST-BEGIN: fault-ap-df\n", .{});
if (!sched.spawnOn(apDoubleFaultTask, 6, 1)) {
log("DANOS-AP: could not pin to core 1 (run with -smp) - FAIL\n", .{});
arch.halt();
}
// Wait until the AP is about to fault, then keep running to prove containment.
var spins: u64 = 0;
while (!@atomicLoad(bool, &ap_reached_fault, .acquire) and spins < 5_000_000_000) spins +%= 1;
var settle: u64 = 0;
while (settle < 500_000_000) settle +%= 1; // let the AP take + report the fault
log("DANOS-BSP: core {d} still running after the AP fault (contained)\n", .{sched.currentCpuIndex()});
arch.halt();
}
+46 -2
View File
@@ -81,6 +81,12 @@ CASES = [
{"name": "smoke",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
{"name": "discovery",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
{"name": "wx",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
{"name": "timer",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
@@ -108,13 +114,48 @@ CASES = [
{"name": "ipc",
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Parallelism: needs more than one core, so this case boots with -smp 4.
{"name": "smp",
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Affinity: a pinned task must never migrate off its core.
{"name": "affinity",
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Stress the big kernel lock across cores; heavier, so a longer timeout.
{"name": "smp-stress",
"smp": 4,
"timeout": 90,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# Retry: a forced first-wake failure must still bring every core online.
{"name": "smp-retry",
"smp": 4,
"expect": r"DANOS-TEST-RESULT: PASS",
"fail": r"DANOS-TEST-RESULT: FAIL"},
{"name": "fault-ud", "expect": r"invalid opcode \(vector 6\)"},
{"name": "fault-pf", "expect": r"page fault \(vector 14\)"},
{"name": "fault-df", "expect": r"double fault \(vector 8\)"},
# Fault on an application processor: a #DF pinned to core 1 must be caught by that
# core's own IST (N >= 1), proving per-core TSS works; a broken one triple-faults.
{"name": "fault-ap-df",
"smp": 4,
"expect": r"core [1-9]\d*: double fault \(vector 8\)",
"fail": r"could not pin"},
{"name": "fault-nx",
"expect": r"page fault \(vector 14\)",
"fail": r"NX not enforced"},
{"name": "fault-null", "expect": r"page fault \(vector 14\)"},
# The ACPI power path succeeds by QEMU *exiting* (S5 off / reset), so match the
# pre-transition marker; the FAIL line only appears if the transition didn't take.
{"name": "poweroff",
"expect": r"DANOS-POWER: attempting poweroff",
"fail": r"DANOS-TEST-RESULT: FAIL"},
{"name": "reboot",
"expect": r"DANOS-POWER: attempting reboot",
"fail": r"DANOS-TEST-RESULT: FAIL"},
]
TIMEOUT = 30 # seconds per case
@@ -175,9 +216,12 @@ def run_case(arch, case):
fail = re.compile(case["fail"]) if case.get("fail") else None
cmd = [arch["qemu"]] + arch["qemu_args"](arch, esp, vars_fd, serial)
if case.get("smp"): # some cases need more than one core (e.g. parallelism)
cmd += ["-smp", str(case["smp"])]
qemu = subprocess.Popen(cmd, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
try:
deadline = time.monotonic() + TIMEOUT
timeout = case.get("timeout", TIMEOUT)
deadline = time.monotonic() + timeout
while time.monotonic() < deadline:
time.sleep(0.2)
text = ""
@@ -192,7 +236,7 @@ def run_case(arch, case):
if expect.search(text):
return True, "matched " + repr(case["expect"])
return False, "QEMU exited before matching (triple fault?)"
return False, f"timed out after {TIMEOUT}s without matching {case['expect']!r}"
return False, f"timed out after {timeout}s without matching {case['expect']!r}"
finally:
if qemu.poll() is None:
qemu.terminate()