Fixed-priority preemptive scheduler
This commit is contained in:
+9
-4
@@ -31,7 +31,10 @@ rather than restate it. Roughly in the order things happen at runtime:
|
||||
9. **[heap.md](heap.md) — the kernel heap.** A growable free-list allocator built on
|
||||
the VMM, exposed as a `std.mem.Allocator` so std containers work — dynamic
|
||||
allocation for the kernel.
|
||||
10. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
|
||||
10. **[scheduling.md](scheduling.md) — the scheduler.** Fixed-priority preemptive
|
||||
multitasking: kernel threads, the context switch, and O(1) priority selection —
|
||||
the leap to a running system.
|
||||
11. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
|
||||
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
|
||||
|
||||
Start with the north star:
|
||||
@@ -61,9 +64,10 @@ into a **frame allocator** ([frame-allocator.md](frame-allocator.md)), installs
|
||||
its **descriptor tables** so CPU faults are caught ([interrupts.md](interrupts.md)),
|
||||
builds its own **page tables** and switches onto them ([paging.md](paging.md)),
|
||||
brings up the **heap** for dynamic allocation ([heap.md](heap.md)), starts the
|
||||
**timer** so it has a heartbeat ([device-interrupts.md](device-interrupts.md)),
|
||||
runs — its CPU-specific bits behind the [arch](arch.md) boundary — and when it has
|
||||
finished, or panics, it **halts** ([halting.md](halting.md)).
|
||||
**scheduler** ([scheduling.md](scheduling.md)) and the **timer** that preempts it
|
||||
([device-interrupts.md](device-interrupts.md)), runs — its CPU-specific bits behind
|
||||
the [arch](arch.md) boundary — and when idle, or on a panic, it **halts**
|
||||
([halting.md](halting.md)).
|
||||
|
||||
## Source map
|
||||
|
||||
@@ -74,6 +78,7 @@ finished, or panics, it **halts** ([halting.md](halting.md)).
|
||||
| Shared loader↔kernel contract (`BootInfo`, `Framebuffer`, `MemoryMap`, ABI) | `src/root.zig` |
|
||||
| Physical frame allocator | `src/pmm.zig` |
|
||||
| Kernel heap (`std.mem.Allocator`) | `src/heap.zig` |
|
||||
| Scheduler (fixed-priority preemptive) | `src/sched.zig` |
|
||||
| Framebuffer text console (mirrors to serial) | `src/console.zig` |
|
||||
| In-kernel test cases | `src/tests.zig` |
|
||||
| Arch-specific kernel code (`halt`, GDT/IDT/TSS, exception + interrupt stubs, page tables, APIC/timer, serial, linker script) | `src/arch/x86_64/` |
|
||||
|
||||
+4
-2
@@ -74,8 +74,10 @@ There are really two independent questions, and it's worth not conflating them:
|
||||
- **`src/arch/x86_64/serial.zig`** / **`io.zig`** — the COM1 UART (the kernel's
|
||||
machine-readable log channel, see [testing.md](testing.md)) and the shared
|
||||
port-I/O + MSR primitives.
|
||||
- **`src/arch/x86_64/isr.s`** — the exception stubs and the `lgdt`/`lidt`/`ltr`
|
||||
load helpers, in real assembly because Zig inline asm can't express them.
|
||||
- **`src/arch/x86_64/isr.s`** — the exception stubs, the `lgdt`/`lidt`/`ltr` load
|
||||
helpers, and the context switch (`switch_context` / `task_trampoline`, see
|
||||
[scheduling.md](scheduling.md)) — real assembly, since Zig inline asm can't
|
||||
express them.
|
||||
- **`src/arch/x86_64/linker.ld`** — the kernel link layout (fixed low load
|
||||
address, one PT_LOAD per permission set).
|
||||
|
||||
|
||||
@@ -0,0 +1,92 @@
|
||||
# Scheduling
|
||||
|
||||
The scheduler turns danos from a linear "boot then halt" kernel into a **running
|
||||
multitasking system**. It's **fixed-priority preemptive**: the highest-priority
|
||||
ready task always runs, and tasks at the same priority take turns. That model is
|
||||
chosen for [real-time](vision.md) — it's predictable (you can reason about which
|
||||
task runs when) and its decisions are O(1), unlike a fair-share scheduler.
|
||||
|
||||
The scheduler proper (`src/sched.zig`) is generic; the context switch and new-task
|
||||
stack setup are architecture-specific (`src/arch/x86_64/`, see [arch](arch.md)).
|
||||
|
||||
## Tasks
|
||||
|
||||
A **task** is a kernel thread: ring-0 code with its own 16 KiB stack (allocated
|
||||
from the [heap](heap.md)). A task struct holds its saved stack pointer, priority,
|
||||
state, and a ready-queue link. The currently-running kernel context (kmain)
|
||||
registers itself as task 0, so there's always something to switch *from*.
|
||||
|
||||
## The context switch
|
||||
|
||||
Switching tasks means swapping stacks. `switch_context(old, new)` (in `isr.s`)
|
||||
saves the **callee-saved** registers on the current stack, stores the stack pointer
|
||||
into the old task, loads the new task's stack pointer, restores *its* callee-saved
|
||||
registers, and `ret`s — landing wherever the new task was last suspended. Only
|
||||
callee-saved registers are handled explicitly: to the compiler this looks like a
|
||||
normal function call, so it already preserves the caller-saved ones itself (this is
|
||||
the [SysV](sysv.md) convention doing the work).
|
||||
|
||||
A **freshly spawned** task has never run, so there's nothing to restore. Its stack
|
||||
is faked to look as if it had just called `switch_context`: `init_task_stack` lays
|
||||
down a return address pointing at `task_trampoline` and zeroed callee-saved slots
|
||||
(smuggling the entry function in via the `r15` slot). When first switched to, the
|
||||
`ret` lands in the trampoline, which enables interrupts and calls the entry.
|
||||
|
||||
## Two ways to switch, one flag discipline
|
||||
|
||||
`schedule()` — pick the best task and switch — runs from two places:
|
||||
|
||||
- **`yield()`** — a task voluntarily gives up the CPU.
|
||||
- **`tick()`** — the 1000 Hz [timer](device-interrupts.md) preempts the running
|
||||
task. This is what lets a task that never yields still share the CPU.
|
||||
|
||||
The subtlety in mixing them is the **interrupt flag (IF)**. The rule: `switch_context`
|
||||
is always entered with interrupts *disabled* — naturally so inside the timer ISR,
|
||||
and explicitly (`cli`) in `yield`. Then every task ends up with interrupts enabled
|
||||
again through whichever path resumes it:
|
||||
|
||||
- a task suspended in `yield` re-enables them (`sti`) right after `schedule` returns;
|
||||
- a task suspended mid-ISR resumes through the interrupt return (`iretq`), which
|
||||
restores the `RFLAGS` it had when it was preempted (IF set);
|
||||
- a brand-new task enables them in the trampoline.
|
||||
|
||||
One related detail: the timer interrupt is **acknowledged (EOI) before** its handler
|
||||
runs, so a handler that switches tasks and doesn't return promptly can't stall the
|
||||
LAPIC from delivering the next tick.
|
||||
|
||||
## Priority selection, in O(1)
|
||||
|
||||
Ready tasks live in a **FIFO queue per priority level** (8 levels), plus a
|
||||
**bitmap** with one bit per non-empty level. Picking the next task is: find the
|
||||
highest set bit (one instruction), take the front of that level's queue. No list
|
||||
walking, no scanning — the decision cost is constant regardless of how many tasks
|
||||
exist, which is what a real-time scheduler needs.
|
||||
|
||||
- **Highest priority wins.** A ready high-priority task always runs before a
|
||||
lower-priority one.
|
||||
- **Round-robin within a level.** When a task is descheduled it goes to the *back*
|
||||
of its level's queue, so equal-priority tasks share the CPU fairly.
|
||||
|
||||
## Verifying it
|
||||
|
||||
Two tests (see [testing.md](testing.md)) prove the two guarantees:
|
||||
|
||||
- **`sched`** spawns three tasks that busy-loop *without ever yielding*. They all
|
||||
make progress — which can only happen if the timer is **preempting** between them
|
||||
and the context switch is correct (nothing yields voluntarily).
|
||||
- **`priority`** (with preemption off, for determinism) spawns tasks at three
|
||||
priorities; they run and exit **highest-priority first** — `[6, 4, 2]`.
|
||||
|
||||
## What's next (not done here)
|
||||
|
||||
- **Blocking and sleep.** Right now a task can only yield or exit; it can't wait for
|
||||
a condition or a duration. `sleep(ms)` (on the calibrated clock) and blocking
|
||||
come next, and are what a real-time task really needs.
|
||||
- **Priority inheritance.** Once tasks block on shared resources (locks, IPC),
|
||||
danos will need it to bound priority inversion — a [real-time](vision.md)
|
||||
requirement.
|
||||
- **Task exit / a reaper.** `exit` currently leaks the task's stack; nothing frees
|
||||
finished tasks' memory yet.
|
||||
- **Per-address-space tasks.** Today all tasks share the kernel address space. User
|
||||
processes will each get their own, switching page tables (CR3) on the context
|
||||
switch.
|
||||
@@ -50,6 +50,8 @@ Current cases:
|
||||
| `clock` | calibrated LAPIC frequency is sane; monotonic uptime advances | `DANOS-TEST-RESULT: PASS` |
|
||||
| `vmm` | on-demand `map` works: a mapped page is writable and reads back | `DANOS-TEST-RESULT: PASS` |
|
||||
| `heap` | kernel heap: alloc/free, block reuse, growth, and a std container on it | `DANOS-TEST-RESULT: PASS` |
|
||||
| `sched` | preemption: three non-yielding tasks all make progress | `DANOS-TEST-RESULT: PASS` |
|
||||
| `priority` | fixed-priority tasks run highest-first | `DANOS-TEST-RESULT: PASS` |
|
||||
| `fault-ud` | invalid-opcode exception is caught | serial shows `invalid opcode (vector 6)` |
|
||||
| `fault-pf` | page fault caught with CR2 | `page fault (vector 14)` |
|
||||
| `fault-df` | double fault caught on IST1 (not a triple-fault reset) | `double fault (vector 8)` |
|
||||
|
||||
Reference in New Issue
Block a user