IPC
This commit is contained in:
+11
-6
@@ -32,9 +32,12 @@ rather than restate it. Roughly in the order things happen at runtime:
|
||||
the VMM, exposed as a `std.mem.Allocator` so std containers work — dynamic
|
||||
allocation for the kernel.
|
||||
10. **[scheduling.md](scheduling.md) — the scheduler.** Fixed-priority preemptive
|
||||
multitasking: kernel threads, the context switch, and O(1) priority selection —
|
||||
the leap to a running system.
|
||||
11. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
|
||||
multitasking: kernel threads, the context switch, O(1) priority selection, and
|
||||
blocking (sleep, wait queues) — the leap to a running system.
|
||||
11. **[ipc.md](ipc.md) — inter-process communication.** Bounded blocking
|
||||
message-passing channels — the backbone the microkernel's isolated servers will
|
||||
talk over.
|
||||
12. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
|
||||
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
|
||||
|
||||
Start with the north star:
|
||||
@@ -65,8 +68,9 @@ its **descriptor tables** so CPU faults are caught ([interrupts.md](interrupts.m
|
||||
builds its own **page tables** and switches onto them ([paging.md](paging.md)),
|
||||
brings up the **heap** for dynamic allocation ([heap.md](heap.md)), starts the
|
||||
**scheduler** ([scheduling.md](scheduling.md)) and the **timer** that preempts it
|
||||
([device-interrupts.md](device-interrupts.md)), runs — its CPU-specific bits behind
|
||||
the [arch](arch.md) boundary — and when idle, or on a panic, it **halts**
|
||||
([device-interrupts.md](device-interrupts.md)) — with tasks blocking, sleeping and
|
||||
passing messages over **[IPC](ipc.md)** channels — runs, its CPU-specific bits
|
||||
behind the [arch](arch.md) boundary, and when idle, or on a panic, it **halts**
|
||||
([halting.md](halting.md)).
|
||||
|
||||
## Source map
|
||||
@@ -78,7 +82,8 @@ the [arch](arch.md) boundary — and when idle, or on a panic, it **halts**
|
||||
| Shared loader↔kernel contract (`BootInfo`, `Framebuffer`, `MemoryMap`, ABI) | `src/root.zig` |
|
||||
| Physical frame allocator | `src/pmm.zig` |
|
||||
| Kernel heap (`std.mem.Allocator`) | `src/heap.zig` |
|
||||
| Scheduler (fixed-priority preemptive) | `src/sched.zig` |
|
||||
| Scheduler (fixed-priority preemptive; blocking, wait queues) | `src/sched.zig` |
|
||||
| IPC channels (message passing) | `src/ipc.zig` |
|
||||
| Framebuffer text console (mirrors to serial) | `src/console.zig` |
|
||||
| In-kernel test cases | `src/tests.zig` |
|
||||
| Arch-specific kernel code (`halt`, GDT/IDT/TSS, exception + interrupt stubs, page tables, APIC/timer, serial, linker script) | `src/arch/x86_64/` |
|
||||
|
||||
@@ -46,8 +46,25 @@ it against the **PIT** (the legacy 8254, whose 1.193182 MHz is fixed): run the
|
||||
LAPIC timer one-shot from its maximum count while the PIT counts out a known 10 ms
|
||||
(polling channel 2, no interrupt needed), then see how far the LAPIC got. That
|
||||
yields its counts-per-millisecond, from which `initTimer(hz)` computes the reload
|
||||
count for any target frequency. danos runs it at **1000 Hz** (a 1 ms tick), and the
|
||||
tick count times the known period gives a monotonic `uptimeMs()`.
|
||||
count for any target frequency. danos runs it at **1000 Hz** (a 1 ms tick).
|
||||
|
||||
## The high-resolution clock (TSC)
|
||||
|
||||
The timer tick gives *scheduling* — a 1 ms quantum — but 1 ms is coarse for a
|
||||
real-time system to *measure* with (interrupt latency, jitter, timeouts). So the
|
||||
same calibration also measures the **TSC** (Time Stamp Counter): a per-core cycle
|
||||
counter read with `rdtsc` in a couple of cycles, giving roughly **nanosecond**
|
||||
resolution — a million times finer than the tick. We snapshot the TSC across the
|
||||
same 10 ms PIT window to get its frequency (measured ~3.6 GHz on the test host).
|
||||
|
||||
The monotonic clock is exposed as one function per resolution — `nanos()`,
|
||||
`micros()`, `millis()` — each scaling the cycle delta directly at its unit (with a
|
||||
128-bit intermediate so a long uptime doesn't overflow) rather than chaining
|
||||
divisions. `millis()` is what the scheduler uses for `sleep` deadlines; `nanos()`
|
||||
is there for fine measurement. Note the two clocks are distinct: the **tick** drives
|
||||
preemption and wakeups (1 ms granularity); the **TSC** is the resolution you read
|
||||
time at. Making `sleep` itself sub-millisecond would take a tickless one-shot
|
||||
timer — a later step.
|
||||
|
||||
## Two kinds of vector, one dispatch
|
||||
|
||||
|
||||
+57
@@ -0,0 +1,57 @@
|
||||
# IPC: message-passing channels
|
||||
|
||||
Inter-process communication is the **backbone of a microkernel**. Once drivers and
|
||||
services run isolated in their own address spaces ([vision](vision.md)), they can't
|
||||
just call each other — a request becomes a **message**. In a microkernel, whatever
|
||||
was a function call across a monolithic kernel is IPC, so it's a first-class
|
||||
concern, not an afterthought.
|
||||
|
||||
This first form is a **bounded blocking channel** (`src/ipc.zig`): a fixed-size
|
||||
ring buffer of messages with a producer/consumer rendezvous, built on the
|
||||
scheduler's [wait queues](scheduling.md).
|
||||
|
||||
## The channel
|
||||
|
||||
`Channel(T, capacity)` is generic over the message type and buffer size. It holds a
|
||||
ring buffer, a count, and two wait queues:
|
||||
|
||||
- **`send(msg)`** — if the channel is full, block on the *not-full* queue; otherwise
|
||||
write the message, bump the count, and wake a waiting receiver.
|
||||
- **`recv()`** — if the channel is empty, block on the *not-empty* queue; otherwise
|
||||
take a message, drop the count, and wake a waiting sender.
|
||||
|
||||
Neither side busy-waits: a full channel parks the sender, an empty one parks the
|
||||
receiver, and each operation wakes the other side when it makes progress possible.
|
||||
|
||||
Two details make it correct:
|
||||
|
||||
- **Recheck in a loop.** A woken task re-tests the condition (`while (full) wait`)
|
||||
rather than assuming the slot is still available — another waiter may have taken
|
||||
it first. This is the standard guard against spurious or racing wakeups.
|
||||
- **One critical section.** `send`/`recv` run under `saveInterrupts` /
|
||||
`restoreInterrupts` (the composable form, see [scheduling.md](scheduling.md)), so
|
||||
checking the condition and committing the block/enqueue happen atomically with
|
||||
respect to the timer preempting mid-operation. `waitLocked` / `wakeLocked` are the
|
||||
variants that assume the caller already holds that critical section.
|
||||
|
||||
## Verifying it
|
||||
|
||||
The `ipc` test (see [testing.md](testing.md)) runs a producer and a consumer passing
|
||||
**100 messages through a 4-slot channel**. The small buffer means the channel goes
|
||||
full and empty over and over, so both the blocking-send and blocking-recv paths are
|
||||
exercised heavily. The messages arrive intact and in order (their sum is the
|
||||
expected `5050`), and neither task busy-waits — they block and wake each other.
|
||||
|
||||
## What's next (not done here)
|
||||
|
||||
- **Across address spaces.** Today both endpoints are kernel threads sharing the
|
||||
kernel's memory, so the message is copied within one address space. When user
|
||||
mode arrives, the same channel carries messages between *isolated* processes,
|
||||
copying the payload across the boundary — which is where IPC earns its place as
|
||||
the microkernel's backbone.
|
||||
- **Synchronous call/reply.** A request/response pattern (send-and-wait-for-reply)
|
||||
on top of channels, the shape most driver/service calls take.
|
||||
- **Interrupts as messages.** A hardware interrupt delivered to the driver task
|
||||
that owns the device, as an IPC message.
|
||||
- **Priority inheritance** through IPC, so a high-priority client blocked on a
|
||||
low-priority server doesn't suffer unbounded priority inversion.
|
||||
+39
-4
@@ -67,21 +67,56 @@ exist, which is what a real-time scheduler needs.
|
||||
- **Round-robin within a level.** When a task is descheduled it goes to the *back*
|
||||
of its level's queue, so equal-priority tasks share the CPU fairly.
|
||||
|
||||
## Sleeping and the idle task
|
||||
|
||||
A task can **block** — give up the CPU until an event, rather than busy-wait
|
||||
(busy-waiting is the enemy of a real-time system: it wastes cycles a
|
||||
higher-priority task should get). The first form is time-based: **`sleep(ms)`**
|
||||
marks the task blocked with a wake deadline and switches away. On every tick the
|
||||
timer wakes any task whose deadline has passed (a bounded scan, so it stays
|
||||
deterministic), which makes it ready again; the scheduler then runs it when its
|
||||
priority comes up. `sleep` measures its deadline on the [calibrated
|
||||
clock](device-interrupts.md), so it's real time.
|
||||
|
||||
When *every* task is blocked, something still has to run — so there's an **idle
|
||||
task** at the lowest priority that just `hlt`s until the next interrupt (see
|
||||
[halting.md](halting.md)). Because it's always runnable, the scheduler always has a
|
||||
task to pick, and the "nothing to run" case never arises.
|
||||
|
||||
## Event-based blocking
|
||||
|
||||
The other form of blocking is waiting for an **event** rather than a duration. A
|
||||
**wait queue** is a set of tasks parked until something happens: `wait(wq)` blocks
|
||||
the caller on it, `wake(wq)` moves the highest-priority waiter back to ready
|
||||
(preempting if it now outranks the running task). A task links into a wait queue
|
||||
through the same field the ready queues use — it's in exactly one queue at a time.
|
||||
These are the primitives locks, semaphores and [IPC](ipc.md) are built on.
|
||||
|
||||
Blocking safely needs **composable critical sections**. A blanket `cli`/`sti` pair
|
||||
doesn't nest: an IPC channel that `cli`s and then calls `wait` would have `wait`'s
|
||||
`sti` re-enable interrupts too early, mid-operation. So the blocking primitives use
|
||||
`saveInterrupts` / `restoreInterrupts` — capture the interrupt flag, disable, and
|
||||
later restore *only if it was set* — which nests correctly. The invariant that
|
||||
makes it all work: `schedule()` is always entered with interrupts disabled, so a
|
||||
task always resumes from a switch with interrupts disabled and can restore its
|
||||
caller's state.
|
||||
|
||||
## Verifying it
|
||||
|
||||
Two tests (see [testing.md](testing.md)) prove the two guarantees:
|
||||
Three tests (see [testing.md](testing.md)) prove the guarantees:
|
||||
|
||||
- **`sched`** spawns three tasks that busy-loop *without ever yielding*. They all
|
||||
make progress — which can only happen if the timer is **preempting** between them
|
||||
and the context switch is correct (nothing yields voluntarily).
|
||||
- **`priority`** (with preemption off, for determinism) spawns tasks at three
|
||||
priorities; they run and exit **highest-priority first** — `[6, 4, 2]`.
|
||||
- **`sleep`** blocks a task for 50 ms and checks the elapsed time on the clock — a
|
||||
real block (the idle task runs meanwhile), not a busy-wait.
|
||||
- **`event`** blocks a task on a wait queue; waking it (from another task) resumes
|
||||
it, and since it's higher priority it preempts immediately.
|
||||
|
||||
## What's next (not done here)
|
||||
|
||||
- **Blocking and sleep.** Right now a task can only yield or exit; it can't wait for
|
||||
a condition or a duration. `sleep(ms)` (on the calibrated clock) and blocking
|
||||
come next, and are what a real-time task really needs.
|
||||
- **Priority inheritance.** Once tasks block on shared resources (locks, IPC),
|
||||
danos will need it to bound priority inversion — a [real-time](vision.md)
|
||||
requirement.
|
||||
|
||||
+4
-1
@@ -47,11 +47,14 @@ Current cases:
|
||||
|------|----------------|-----------------------------|
|
||||
| `smoke` | memory map has usable RAM; frame alloc/free; paging active | `DANOS-TEST-RESULT: PASS` |
|
||||
| `timer` | device interrupts fire and return (tick count advances) | `DANOS-TEST-RESULT: PASS` |
|
||||
| `clock` | calibrated LAPIC frequency is sane; monotonic uptime advances | `DANOS-TEST-RESULT: PASS` |
|
||||
| `clock` | LAPIC + TSC calibrated; monotonic uptime advances; `nanos()` has sub-ms resolution | `DANOS-TEST-RESULT: PASS` |
|
||||
| `vmm` | on-demand `map` works: a mapped page is writable and reads back | `DANOS-TEST-RESULT: PASS` |
|
||||
| `heap` | kernel heap: alloc/free, block reuse, growth, and a std container on it | `DANOS-TEST-RESULT: PASS` |
|
||||
| `sched` | preemption: three non-yielding tasks all make progress | `DANOS-TEST-RESULT: PASS` |
|
||||
| `priority` | fixed-priority tasks run highest-first | `DANOS-TEST-RESULT: PASS` |
|
||||
| `sleep` | a task blocks for ~50 ms (real block, not a busy-wait) | `DANOS-TEST-RESULT: PASS` |
|
||||
| `event` | a task blocks on a wait queue and is woken (preempting) | `DANOS-TEST-RESULT: PASS` |
|
||||
| `ipc` | producer/consumer pass 100 messages through a 4-slot channel intact | `DANOS-TEST-RESULT: PASS` |
|
||||
| `fault-ud` | invalid-opcode exception is caught | serial shows `invalid opcode (vector 6)` |
|
||||
| `fault-pf` | page fault caught with CR2 | `page fault (vector 14)` |
|
||||
| `fault-df` | double fault caught on IST1 (not a triple-fault reset) | `double fault (vector 8)` |
|
||||
|
||||
Reference in New Issue
Block a user