This commit is contained in:
2026-07-03 15:09:16 +01:00
parent df774691c8
commit 31bf86af56
12 changed files with 492 additions and 41 deletions
+11 -6
View File
@@ -32,9 +32,12 @@ rather than restate it. Roughly in the order things happen at runtime:
the VMM, exposed as a `std.mem.Allocator` so std containers work — dynamic
allocation for the kernel.
10. **[scheduling.md](scheduling.md) — the scheduler.** Fixed-priority preemptive
multitasking: kernel threads, the context switch, and O(1) priority selection —
the leap to a running system.
11. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
multitasking: kernel threads, the context switch, O(1) priority selection, and
blocking (sleep, wait queues) — the leap to a running system.
11. **[ipc.md](ipc.md) — inter-process communication.** Bounded blocking
message-passing channels — the backbone the microkernel's isolated servers will
talk over.
12. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
Start with the north star:
@@ -65,8 +68,9 @@ its **descriptor tables** so CPU faults are caught ([interrupts.md](interrupts.m
builds its own **page tables** and switches onto them ([paging.md](paging.md)),
brings up the **heap** for dynamic allocation ([heap.md](heap.md)), starts the
**scheduler** ([scheduling.md](scheduling.md)) and the **timer** that preempts it
([device-interrupts.md](device-interrupts.md)), runs — its CPU-specific bits behind
the [arch](arch.md) boundary — and when idle, or on a panic, it **halts**
([device-interrupts.md](device-interrupts.md)) — with tasks blocking, sleeping and
passing messages over **[IPC](ipc.md)** channels — runs, its CPU-specific bits
behind the [arch](arch.md) boundary, and when idle, or on a panic, it **halts**
([halting.md](halting.md)).
## Source map
@@ -78,7 +82,8 @@ the [arch](arch.md) boundary — and when idle, or on a panic, it **halts**
| Shared loader↔kernel contract (`BootInfo`, `Framebuffer`, `MemoryMap`, ABI) | `src/root.zig` |
| Physical frame allocator | `src/pmm.zig` |
| Kernel heap (`std.mem.Allocator`) | `src/heap.zig` |
| Scheduler (fixed-priority preemptive) | `src/sched.zig` |
| Scheduler (fixed-priority preemptive; blocking, wait queues) | `src/sched.zig` |
| IPC channels (message passing) | `src/ipc.zig` |
| Framebuffer text console (mirrors to serial) | `src/console.zig` |
| In-kernel test cases | `src/tests.zig` |
| Arch-specific kernel code (`halt`, GDT/IDT/TSS, exception + interrupt stubs, page tables, APIC/timer, serial, linker script) | `src/arch/x86_64/` |
+19 -2
View File
@@ -46,8 +46,25 @@ it against the **PIT** (the legacy 8254, whose 1.193182 MHz is fixed): run the
LAPIC timer one-shot from its maximum count while the PIT counts out a known 10 ms
(polling channel 2, no interrupt needed), then see how far the LAPIC got. That
yields its counts-per-millisecond, from which `initTimer(hz)` computes the reload
count for any target frequency. danos runs it at **1000 Hz** (a 1 ms tick), and the
tick count times the known period gives a monotonic `uptimeMs()`.
count for any target frequency. danos runs it at **1000 Hz** (a 1 ms tick).
## The high-resolution clock (TSC)
The timer tick gives *scheduling* — a 1 ms quantum — but 1 ms is coarse for a
real-time system to *measure* with (interrupt latency, jitter, timeouts). So the
same calibration also measures the **TSC** (Time Stamp Counter): a per-core cycle
counter read with `rdtsc` in a couple of cycles, giving roughly **nanosecond**
resolution — a million times finer than the tick. We snapshot the TSC across the
same 10 ms PIT window to get its frequency (measured ~3.6 GHz on the test host).
The monotonic clock is exposed as one function per resolution — `nanos()`,
`micros()`, `millis()` — each scaling the cycle delta directly at its unit (with a
128-bit intermediate so a long uptime doesn't overflow) rather than chaining
divisions. `millis()` is what the scheduler uses for `sleep` deadlines; `nanos()`
is there for fine measurement. Note the two clocks are distinct: the **tick** drives
preemption and wakeups (1 ms granularity); the **TSC** is the resolution you read
time at. Making `sleep` itself sub-millisecond would take a tickless one-shot
timer — a later step.
## Two kinds of vector, one dispatch
+57
View File
@@ -0,0 +1,57 @@
# IPC: message-passing channels
Inter-process communication is the **backbone of a microkernel**. Once drivers and
services run isolated in their own address spaces ([vision](vision.md)), they can't
just call each other — a request becomes a **message**. In a microkernel, whatever
was a function call across a monolithic kernel is IPC, so it's a first-class
concern, not an afterthought.
This first form is a **bounded blocking channel** (`src/ipc.zig`): a fixed-size
ring buffer of messages with a producer/consumer rendezvous, built on the
scheduler's [wait queues](scheduling.md).
## The channel
`Channel(T, capacity)` is generic over the message type and buffer size. It holds a
ring buffer, a count, and two wait queues:
- **`send(msg)`** — if the channel is full, block on the *not-full* queue; otherwise
write the message, bump the count, and wake a waiting receiver.
- **`recv()`** — if the channel is empty, block on the *not-empty* queue; otherwise
take a message, drop the count, and wake a waiting sender.
Neither side busy-waits: a full channel parks the sender, an empty one parks the
receiver, and each operation wakes the other side when it makes progress possible.
Two details make it correct:
- **Recheck in a loop.** A woken task re-tests the condition (`while (full) wait`)
rather than assuming the slot is still available — another waiter may have taken
it first. This is the standard guard against spurious or racing wakeups.
- **One critical section.** `send`/`recv` run under `saveInterrupts` /
`restoreInterrupts` (the composable form, see [scheduling.md](scheduling.md)), so
checking the condition and committing the block/enqueue happen atomically with
respect to the timer preempting mid-operation. `waitLocked` / `wakeLocked` are the
variants that assume the caller already holds that critical section.
## Verifying it
The `ipc` test (see [testing.md](testing.md)) runs a producer and a consumer passing
**100 messages through a 4-slot channel**. The small buffer means the channel goes
full and empty over and over, so both the blocking-send and blocking-recv paths are
exercised heavily. The messages arrive intact and in order (their sum is the
expected `5050`), and neither task busy-waits — they block and wake each other.
## What's next (not done here)
- **Across address spaces.** Today both endpoints are kernel threads sharing the
kernel's memory, so the message is copied within one address space. When user
mode arrives, the same channel carries messages between *isolated* processes,
copying the payload across the boundary — which is where IPC earns its place as
the microkernel's backbone.
- **Synchronous call/reply.** A request/response pattern (send-and-wait-for-reply)
on top of channels, the shape most driver/service calls take.
- **Interrupts as messages.** A hardware interrupt delivered to the driver task
that owns the device, as an IPC message.
- **Priority inheritance** through IPC, so a high-priority client blocked on a
low-priority server doesn't suffer unbounded priority inversion.
+39 -4
View File
@@ -67,21 +67,56 @@ exist, which is what a real-time scheduler needs.
- **Round-robin within a level.** When a task is descheduled it goes to the *back*
of its level's queue, so equal-priority tasks share the CPU fairly.
## Sleeping and the idle task
A task can **block** — give up the CPU until an event, rather than busy-wait
(busy-waiting is the enemy of a real-time system: it wastes cycles a
higher-priority task should get). The first form is time-based: **`sleep(ms)`**
marks the task blocked with a wake deadline and switches away. On every tick the
timer wakes any task whose deadline has passed (a bounded scan, so it stays
deterministic), which makes it ready again; the scheduler then runs it when its
priority comes up. `sleep` measures its deadline on the [calibrated
clock](device-interrupts.md), so it's real time.
When *every* task is blocked, something still has to run — so there's an **idle
task** at the lowest priority that just `hlt`s until the next interrupt (see
[halting.md](halting.md)). Because it's always runnable, the scheduler always has a
task to pick, and the "nothing to run" case never arises.
## Event-based blocking
The other form of blocking is waiting for an **event** rather than a duration. A
**wait queue** is a set of tasks parked until something happens: `wait(wq)` blocks
the caller on it, `wake(wq)` moves the highest-priority waiter back to ready
(preempting if it now outranks the running task). A task links into a wait queue
through the same field the ready queues use — it's in exactly one queue at a time.
These are the primitives locks, semaphores and [IPC](ipc.md) are built on.
Blocking safely needs **composable critical sections**. A blanket `cli`/`sti` pair
doesn't nest: an IPC channel that `cli`s and then calls `wait` would have `wait`'s
`sti` re-enable interrupts too early, mid-operation. So the blocking primitives use
`saveInterrupts` / `restoreInterrupts` — capture the interrupt flag, disable, and
later restore *only if it was set* — which nests correctly. The invariant that
makes it all work: `schedule()` is always entered with interrupts disabled, so a
task always resumes from a switch with interrupts disabled and can restore its
caller's state.
## Verifying it
Two tests (see [testing.md](testing.md)) prove the two guarantees:
Three tests (see [testing.md](testing.md)) prove the guarantees:
- **`sched`** spawns three tasks that busy-loop *without ever yielding*. They all
make progress — which can only happen if the timer is **preempting** between them
and the context switch is correct (nothing yields voluntarily).
- **`priority`** (with preemption off, for determinism) spawns tasks at three
priorities; they run and exit **highest-priority first** — `[6, 4, 2]`.
- **`sleep`** blocks a task for 50 ms and checks the elapsed time on the clock — a
real block (the idle task runs meanwhile), not a busy-wait.
- **`event`** blocks a task on a wait queue; waking it (from another task) resumes
it, and since it's higher priority it preempts immediately.
## What's next (not done here)
- **Blocking and sleep.** Right now a task can only yield or exit; it can't wait for
a condition or a duration. `sleep(ms)` (on the calibrated clock) and blocking
come next, and are what a real-time task really needs.
- **Priority inheritance.** Once tasks block on shared resources (locks, IPC),
danos will need it to bound priority inversion — a [real-time](vision.md)
requirement.
+4 -1
View File
@@ -47,11 +47,14 @@ Current cases:
|------|----------------|-----------------------------|
| `smoke` | memory map has usable RAM; frame alloc/free; paging active | `DANOS-TEST-RESULT: PASS` |
| `timer` | device interrupts fire and return (tick count advances) | `DANOS-TEST-RESULT: PASS` |
| `clock` | calibrated LAPIC frequency is sane; monotonic uptime advances | `DANOS-TEST-RESULT: PASS` |
| `clock` | LAPIC + TSC calibrated; monotonic uptime advances; `nanos()` has sub-ms resolution | `DANOS-TEST-RESULT: PASS` |
| `vmm` | on-demand `map` works: a mapped page is writable and reads back | `DANOS-TEST-RESULT: PASS` |
| `heap` | kernel heap: alloc/free, block reuse, growth, and a std container on it | `DANOS-TEST-RESULT: PASS` |
| `sched` | preemption: three non-yielding tasks all make progress | `DANOS-TEST-RESULT: PASS` |
| `priority` | fixed-priority tasks run highest-first | `DANOS-TEST-RESULT: PASS` |
| `sleep` | a task blocks for ~50 ms (real block, not a busy-wait) | `DANOS-TEST-RESULT: PASS` |
| `event` | a task blocks on a wait queue and is woken (preempting) | `DANOS-TEST-RESULT: PASS` |
| `ipc` | producer/consumer pass 100 messages through a 4-slot channel intact | `DANOS-TEST-RESULT: PASS` |
| `fault-ud` | invalid-opcode exception is caught | serial shows `invalid opcode (vector 6)` |
| `fault-pf` | page fault caught with CR2 | `page fault (vector 14)` |
| `fault-df` | double fault caught on IST1 (not a triple-fault reset) | `double fault (vector 8)` |