M11–M12: IRQ-as-IPC and bus drivers; expand names tree-wide

Two driver-model milestones plus a tree-wide naming pass. Suite 35/35
(QEMU) + host tests green.

M11 — IRQ-as-IPC. A ring-3 driver now sleeps until its device interrupts
it. New src/kernel/irq.zig: per-GSI endpoint bindings, comptime per-vector
trampolines, dispatch = mask GSI -> LAPIC EOI -> notifyLocked, all under one
lock region. irq_bind/irq_ack syscalls, gated by the device claim like
mmio_map. interruptDispatch no longer EOIs — each handler owns its EOI,
because a level line must be masked before it is acknowledged (irq_ack is
the unmask). Bindings are keyed on the owning task and released on exit
(a shared endpoint's siblings survive). hpetd rewritten interrupt-driven.
Tests: hpet (rewritten, reads back the I/O APIC routing) and irqfree.

M12 — bus drivers. DeviceDesc gains a parent, making the device table a
tree. dev_register (device_register) lets a process publish children below
a device it claimed; the kernel enforces resource containment (a child's
resources must nest in its parent's), so a descriptor can't fabricate a
window over kernel RAM. Descriptor copied in via copyFromUser (physmap
walk — an unmapped user pointer fails the call instead of faulting the
kernel). Per-parent child cap bounds table exhaustion. sbin/busd.zig is a
worked bus driver. Test: bus.

Naming — per docs/coding-standards.md: non-acronym abbreviations spelled
out (message, descriptor, device_service, scheduler, runtime, physical,
interpreter, ...); acronyms kept (IPC, MMIO, DMA, HCD, ...); files are
kebab-case (ipc-synchronous.zig, device-service.zig, vfs-protocol.zig, ...).
Exceptions: POSIX/C ABI names and Zig idioms (init/len/ptr) kept. Module
collisions resolved by specific naming (config -> parameters, device.zig
alias -> device_model). AML op/Op disambiguated: op = opcode, Op =
operation; per-opcode parse handlers renamed opX -> parseX.

New driver docs: drivers.md, driver-model.md (bus/class/HCD shapes + the
proposed M13–M16 ABI), coding-standards.md.
This commit is contained in:
Daniel Samson
2026-07-10 11:39:56 +01:00
parent 83881641ca
commit 15b70856c9
63 changed files with 4722 additions and 2690 deletions
+37 -7
View File
@@ -35,9 +35,20 @@ rather than restate it. Roughly in the order things happen at runtime:
multitasking: kernel threads, the context switch, O(1) priority selection, and
blocking (sleep, wait queues) — the leap to a running system.
11. **[ipc.md](ipc.md) — inter-process communication.** Bounded blocking
message-passing channels — the backbone the microkernel's isolated servers will
talk over.
12. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
message-passing channels, then synchronous call/reply between *processes* over
endpoints — the backbone the microkernel's isolated servers talk over.
12. **[syscall.md](syscall.md) — system calls.** How ring 3 asks the kernel for
something: the `syscall`/`sysret` fast path, the trap frame, and why the table is
deliberately tiny.
13. **[drivers.md](drivers.md) — writing a driver.** The payoff: a driver is an
ordinary ring-3 process that claims a device, maps its registers, and **sleeps
until its hardware interrupts it**. The claim is the capability; `irq_ack` is the
unmask.
14. **[driver-model.md](driver-model.md) — buses, classes and host controllers.** How
real driver stacks factor into three shapes, how families share code, and the
proposed ABI for the three primitives still missing (capability passing, DMA +
memory barriers, MSI).
15. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
Start with the north star:
@@ -68,6 +79,10 @@ Cutting across all of these:
- **[smp.md](smp.md) — multiple cores.** A design/research note on how microkernels
(L4, seL4) handle SMP — big kernel lock vs per-CPU vs multikernel — and how the
right choice depends on whether danos is chasing real-time or resilience.
- **[coding-standards.md](coding-standards.md) — coding standards.** The naming rule the
tree follows: non-acronyms are spelled out in full (`message`, not `msg`), files are
`kebab-case`, code follows Zig's case conventions, and the handful of exceptions
(POSIX/C ABI names, `init`/`len`/`ptr`, acronyms).
- **[sysv.md](sysv.md) — the calling convention.** What "the kernel is SysV" means,
and why the loader→kernel boundary has to pin it (the RDI-vs-RCX handoff).
- **[testing.md](testing.md) — testing.** How the kernel is tested by booting it in
@@ -94,19 +109,34 @@ passing messages over **[IPC](ipc.md)** channels — runs, its CPU-specific bits
behind the [arch](arch.md) boundary, and when idle, or on a panic, it **halts**
([halting.md](halting.md)).
Above that line the microkernel proper begins: **discovery** ([discovery.md](discovery.md),
[acpi.md](acpi.md)) learns what hardware exists, ring-3 processes ask the kernel for
things through the small **[syscall](syscall.md)** table, isolated servers reach each
other over IPC **endpoints** ([ipc.md](ipc.md)), and a **[driver](drivers.md)** claims
a device, maps its registers, and sleeps until the hardware interrupts it — which is
the whole reason for the arrangement ([vision.md](vision.md)).
## Source map
| Area | Code |
|------|------|
| Boot methods (one per way of booting the kernel) | `src/boot/` — `efi.zig` (UEFI) → `BOOTX64.efi` |
| Kernel entry, panic, bring-up | `src/kernel/main.zig` |
| Shared loader↔kernel contract (`BootInfo`, `Framebuffer`, `MemoryMap`, ABI) | `src/root.zig` |
| Shared loader↔kernel contract (`BootInfo`, `Framebuffer`, `MemoryMap`, `Syscall`, ABI) | `src/root.zig` |
| Physical frame allocator | `src/kernel/pmm.zig` |
| Kernel heap (`std.mem.Allocator`) | `src/kernel/heap.zig` |
| Scheduler (fixed-priority preemptive; blocking, wait queues) | `src/kernel/sched.zig` |
| IPC channels (message passing) | `src/kernel/ipc.zig` |
| Scheduler (fixed-priority preemptive; blocking, wait queues) | `src/kernel/scheduler.zig` |
| Big kernel lock + interrupt-safe critical sections | `src/kernel/sync.zig` |
| IPC channels between kernel threads (message passing) | `src/kernel/ipc.zig` |
| IPC endpoints: cross-address-space call/reply, handles, notifications | `src/kernel/ipc-synchronous.zig` |
| User processes: ELF loading, address spaces, the syscall table | `src/kernel/process.zig` |
| Device tree + claim capability + `device_register` containment | `src/kernel/device-service.zig` |
| IRQ-as-IPC: routing a device interrupt to a driver's endpoint | `src/kernel/irq.zig` |
| Hardware discovery (ACPI/device tree) behind one neutral device model | `src/device/` |
| Framebuffer text console (mirrors to serial) | `src/kernel/console.zig` |
| In-kernel test cases | `src/kernel/tests.zig` |
| Arch-specific kernel code (`halt`, GDT/IDT/TSS, exception + interrupt stubs, page tables, APIC/timer, serial, linker script) | `src/kernel/arch/x86_64/` |
| Arch-specific kernel code (`halt`, GDT/IDT/TSS, exception + interrupt stubs, page tables, APIC/IO-APIC/timer, serial, linker script) | `src/kernel/arch/x86_64/` |
| User runtime library (`rt`): syscalls, heap, stdio, IPC, device access | `lib/` |
| User-space programs shipped in the initrd (`init`, `vfs`, `hpetd` leaf driver, `busd` bus driver) | `sbin/` |
| Build + `run-x86-64` (QEMU/OVMF) | `build.zig` |
| QEMU integration test harness | `test/qemu_test.py` |
+137
View File
@@ -0,0 +1,137 @@
# Coding standards
Conventions for danos source. The overriding one, from which most of the rest follows:
> **Names are spelled out in full. An identifier is not abbreviated unless the
> abbreviation is an acronym.**
`interruptDispatch`, not `intDisp`. `message_len`, not `message_len` (`msg` expands, `len`
is a Zig idiom — see the exceptions). `device_service`, not `device_service`. `scheduler`, not
`sched`. The cost of a longer name is paid once, at the keyboard; the cost of a
cryptic one is paid every time the code is read, by everyone who reads it. In a
microkernel whose whole argument is that a human can hold each piece in their head,
that trade is not close.
## The rule, precisely
**Acronyms and initialisms stay.** They *are* the full name — expanding them would make
the code worse, not better. `IPC`, `MMIO`, `DMA`, `IRQ`, `TSS`, `GDT`, `IDT`, `APIC`,
`GSI`, `HPET`, `ACPI`, `PCI`, `EOI`, `BAR`, `ECAM`, `MSI`, `CPU`, `ELF`, `ABI`, `UEFI`,
`MMU`, `TLB`, `ISR`, `ISA`, `GAS`, `HAL`, `PMM`, `VMM`, `VFS`, `HID`, `HCD`, `SMP`,
`AML`, `MADT`, `MCFG`, `FADT`, `RSDP`, `XSDT`, `RSDT`, `GOP`, `EDID`, `TSC`, `PIT`,
`RTC`, `LAPIC`, `SIPI`. In code they carry whatever case the surrounding convention
demands: `Hal` the type, `hal` the variable, `mapMmio` the function.
**Everything else is spelled out.** If it's a word with letters removed, restore them:
| Abbreviation | Full |
|---|---|
| `proto` | `protocol` |
| `msg` | `message` |
| `desc` | `descriptor` |
| `res` | `resource` |
| `recv` | `receive` |
| `buf` | `buffer` |
| `cur` | `current` |
| `src` / `dst` | `source` / `destination` |
| `idx` | `index` |
| `addr` | `address` |
| `reg` | `register` |
| `prev` | `previous` |
| `cfg` / `config` | `configuration` |
| `arch` | `architecture` |
| `sched` | `scheduler` |
| `dev` | `device` |
| `sys` / `syscall` | `system` / `system_call` |
| `info` | `information` |
| `dt` | `device_tree` |
| `ep` | `endpoint` |
| `rt` | `runtime` |
| `func` | `function` |
| `phys` / `virt` | `physical` / `virtual` |
| `wq` | `wait_queue` |
This list is illustrative, not exhaustive. The rule is the rule; when you meet a new
abbreviation, expand it.
## Exceptions
Three, and only three.
1. **Foreign ABI names are spelled exactly as the ABI spells them.** A function that
*is* the C or POSIX interface keeps its name: `fopen`, `fwrite`, `fread`, `malloc`,
`calloc`, `realloc`, `free`, `memcpy`, `mmap`, `munmap`, `open`, `read`, `write`,
`close`, `lseek`, `stat`, `errno`. We don't get to rename `fwrite` to
`fileWrite` — it wouldn't be `fwrite` any more. This also covers the syscall
*wrappers* that exist to match those names. It does **not** license inventing new
abbreviated names in that style.
2. **Zig idioms are spelled the way Zig spells them.** Three names are the language's,
not ours, and are left alone:
- **`init` / `deinit`** — the constructor convention (`std.ArrayList.init`), not a
shortening of "initialize".
- **`len` / `ptr`** — the slice field names (`slice.len`, `slice.ptr`). Our own
structs use bare `len`/`ptr` fields to mirror them, so a reader carries one
mental model. (Compounds still expand: a field is `message_len`, not
`message_length` — `len` is kept, `msg` is not.)
- The builtins (`@min`, `@max`, `@memcpy`) and `allocator.alloc` / `.create` are
Zig's spelling.
The rule governs the names *we* coin.
3. **Single-letter variables in a trivial local scope.** `for (items) |item, i|` may
keep `i`; a coordinate may be `x`, `y`. The moment the scope is big enough that the
letter's meaning isn't obvious on sight, give it a real name. When in doubt, name it.
4. **Established Unix filesystem and program conventions.** Top-level directories keep
their conventional names — `src`, `lib`, `sbin`, `bin`, `docs` — as do daemon
programs by their `d` suffix (`hpetd`, `busd`, following `sshd`/`httpd`). These are
names a Unix reader already knows; expanding them fights the convention rather than
serving it.
## A note on collisions
Two identifiers can legitimately expand to the same word. When they do, keep both
meaningful by renaming one to its *specific* identity rather than the generic
expansion. Two cases resolved this way:
- The `config` module (compile-time tunables — `maximum_cpus`, `timer_hz`) would
collide with `cfg` (a `PlatformConfiguration` value) at `configuration`. The module
became **`parameters`**, which is what it holds.
- The kernel `device.zig` module would collide with `dev` (a device value) at
`device`. The module alias became **`device_model`**, which is what it is — the
device data model (`Device`, `DeviceTree`, `ResourceKind`).
- The `Namespace` module alias (`ns`/`nsp` across the AML files) collides with a
`Namespace` **instance**. Resolved by dropping the module alias entirely — the two
types it provided are imported directly (`const Node = @import("namespace.zig").Node;`)
— which frees `namespace` for the instance.
A related case is one abbreviation with two meanings. In the AML code, `op` means
**opcode** (`opcodes.zig`, the `*_opcode` constants) but `Op` in `BinaryOperation` /
`LogicOperation` means **operation** — distinguished by case. The per-opcode parser
handlers, formerly `opName`/`opField`, are `parseName`/`parseField`: they *parse* the
opcode's structure, which says what they do without overloading "op".
## Case and file names
Within those spelling rules, follow Zig's own conventions:
- **Types** — `PascalCase`: `DeviceDescriptor`, `Endpoint`, `WaitQueue`.
- **Functions** — `camelCase`: `mapUserDeviceInto`, `notifyFromIsr`.
- **Variables, fields, constants** — `snake_case`: `message_length`, `device_service`,
`notify_badge_bit`.
**File names are `kebab-case`.** A file named for a multi-word thing hyphenates it:
`device-tree.zig`, `ipc-synchronous.zig`, `vfs-protocol.zig`, `device-service.zig`. A
single word or acronym needs no hyphen: `scheduler.zig`, `paging.zig`, `apic.zig`,
`idt.zig`. (The module *alias* a file is imported under still follows the code
conventions above — `snake_case` — because it's an identifier, not a filename.)
## Why acronyms are the line
Because an acronym has no letters to restore. `MMIO` doesn't become "memory mapped
input output" in code — that expansion is what the acronym *is for*. But `msg` is just
`message` with three letters stolen, and stealing them buys nothing a reader wants. The
test for "is this an abbreviation I must expand" is simply: *is there a longer word this
is a clipped form of?* If yes, write the word. If it's an initialism standing in for a
phrase, leave it.
+32 -11
View File
@@ -89,7 +89,6 @@ if (state.vector < 32) {
on_fault(state); // exception: report and halt (never returns)
} else if (handlers[state.vector]) |handler| {
handler(); // device: run the registered handler
apic.eoi(); // ...acknowledge the LAPIC
}
// else: spurious/unhandled — deliberately no EOI
```
@@ -100,10 +99,23 @@ Two things make device interrupts *return* where exceptions don't:
flows back to `isr_common`, which restores every register it saved and executes
`iretq` — resuming the interrupted instruction exactly. (This is why the stub
saves *all* the general registers.)
2. **End-of-interrupt.** After handling, we write the LAPIC's EOI register. Miss
2. **End-of-interrupt.** Somewhere in there we write the LAPIC's EOI register. Miss
this and the LAPIC thinks we're still busy and never delivers the next
interrupt. It's the single most common "my timer fired once and stopped" bug.
**Each handler issues its own EOI**, rather than the dispatcher doing it around the
call. That looks like a needless devolution while the timer is the only device, and
`apic.timerTick` indeed does nothing but `eoi()` before bumping its counter (early,
because the tick hook is the scheduler, which may switch tasks and not return
promptly — the LAPIC mustn't wait on it).
It stops looking needless with the second device. A *routed* interrupt — one arriving
through the I/O APIC from a real device line — must be **masked before it is
acknowledged**, because a level-triggered line is still asserted at EOI time and would
redeliver instantly, forever. Only the handler knows which discipline its source
needs, so only the handler can sequence it. See [drivers.md](drivers.md), where the
device is quieted by a driver in ring 3, long after the ISR has returned.
A device handler is a plain `fn () void` — a timer or keyboard handler doesn't need
the interrupted registers. (Note: the stubs don't save the SSE/vector registers, so
a handler must not use them; ours don't.)
@@ -131,14 +143,23 @@ If the APIC weren't enabled, or `sti` were missing, or EOI were forgotten, the
count would stay put and the test would fail. That it advances — while the CPU was
spinning in unrelated code — is the whole mechanism working end to end.
## Since (done elsewhere)
- **Preemption**: the timer handler is where the scheduler decides to switch — the
reason a *returning* interrupt matters. See [scheduling.md](scheduling.md).
- **`sleep()` / timeouts** built on the calibrated clock.
- **The I/O APIC, routed**: external device lines now reach a vector, and the
interrupt is delivered onward to a *user-space* driver as an IPC message. See
[drivers.md](drivers.md).
- **Uncacheable MMIO**: device grants are mapped `PCD|PWT` (strong-uncacheable) for
user drivers — see [paging.md](paging.md).
## What's next (not done here)
- **The keyboard**: bring up the IO-APIC, route its IRQ to a vector, and read
scancodes from the PS/2 controller — the first *input* device.
- **`sleep()` / timeouts** built on the calibrated clock (the monotonic
`uptimeMs()` is in place).
- **Uncacheable MMIO**: the LAPIC page is currently mapped writeback-cacheable like
the rest of the identity map. QEMU tolerates it, but real hardware wants MMIO
marked uncacheable (via the page's cache bits or an MTRR).
- **Preemption**: once there are tasks, the timer handler is where the scheduler
decides to switch — the reason a *returning* interrupt matters.
- **The keyboard**: the PS/2 controller is port-mapped (`0x60`/`0x64`), and ring 3
has no port I/O yet, so the first *input* device is blocked on either an I/O
permission bitmap or `io_in`/`io_out` syscalls ([drivers.md](drivers.md)).
- **MSI/MSI-X**: per-device vectors, edge-triggered and unshared, which retire the
I/O APIC's mask/ack cycle and its 24-GSI ceiling.
- **The LAPIC's own page** is still mapped writeback-cacheable like the rest of the
identity map. QEMU tolerates it; real hardware wants it uncacheable.
+305
View File
@@ -0,0 +1,305 @@
# The driver model: buses, classes, and host controllers
[drivers.md](drivers.md) shows how to write *a* driver — claim a device, map its
registers, sleep on its interrupt. That's enough for a leaf device like the HPET. It is
not enough for a disk, a keyboard, or a network card, because those hang off a
*controller*, on a *bus*, speaking a *protocol*, and no single process should have to
know all three.
Real driver stacks factor into three shapes. This document is about what each one is,
what the kernel must give it, how they share code — and precisely which primitive each
is still blocked on.
## Three shapes
| Shape | Owns | Reaches hardware by | Talks to |
|---|---|---|---|
| **Host controller driver** (HCD) | a controller — an xHCI PCI function, an AHCI port block | `mmio_map` + `irq_bind` + DMA | the devices behind it, in its bus's language |
| **Bus driver** | a bus — a PCI bridge, a USB hub | `device_register`, to publish what it finds | class drivers, over IPC |
| **Class / protocol driver** | *nothing* | *nothing* | its bus driver, over IPC |
The last row is the surprising one and the whole point. A USB keyboard driver touches
no registers, takes no interrupts, and maps no memory. It sends HID protocol messages
to whatever published the device, and it works identically whether the controller
below is xHCI, EHCI, or a Raspberry Pi's DWC2. That is what buys you drivers that
outlive the hardware they were written for.
In practice **HCD and bus driver are usually the same process**. An xHCI driver is a
host controller driver (it owns the PCI function, its BARs, its interrupt, its DMA
rings) *and* a bus driver (it enumerates USB devices and publishes them). Splitting
them is a fiction; what matters is that both *roles* have kernel support, because a
plain bus driver with no controller — a USB hub — is also a real thing.
## The device table is the spine
danos already has the right central structure. `src/kernel/device-service.zig` holds a table of
`DeviceDesc`, each with a parent, a class, and a set of resources. Firmware discovery
seeds it ([discovery.md](discovery.md)); `device_register` grows it.
Three invariants make it a capability system rather than a directory:
1. **A claim is exclusive.** `device_claim(id)` succeeds once. Everything downstream —
`mmio_map`, `irq_bind`, `device_register` — checks `device_service.ownerOf(id) == me`.
2. **A descriptor is a licence to map physical memory.** Whoever claims a device may
map its `.memory` resources and bind its `.irq` resources. This is why
`device_register` cannot be a free-for-all.
3. **Therefore: containment.** Every resource of a registered child must lie inside a
resource of the same kind on its parent (`device_service.contains`). A bus driver can only
ever *subdivide* what it already holds. Without this, `device_register` would be a
syscall named "map any physical page you like."
Containment is transitive by construction: a grandchild is contained in its child,
which is contained in the bus. Nothing can be laundered through a chain.
Note that firmware topology does **not** obey containment, and isn't asked to — a PCI
function's BAR is not inside its host bridge's `bus_range`, because a bus-number range
is not an address window. Discovery is trusted; user space is not.
### What a bus driver looks like
`sbin/busd.zig` is the smallest honest one. Its "bus" is the HPET's register block and
its "devices" are the block's comparators:
```zig
_ = dev.claim(bus.id); // 1. own the bus
const base = dev.mmioMap(bus.id, 0).?; // 2. enumerate it — from the hardware
const n = ((cap.* >> 8) & 0x1F) + 1; // GENERAL_CAP says how many children
for (0..n) |i| { // 3. publish each child
var child = std.mem.zeroes(dev.DeviceDesc);
child.class = @intFromEnum(dev.DeviceClass.timer);
child.resource_count = 1;
child.resources[0] = .{ .kind = memory,
.start = bus_mmio.start + 0x100 + 0x20 * i,
.len = 0x20 };
_ = dev.register(bus.id, &child).?; // kernel checks containment
}
```
Each child is left **unclaimed**, which is the handoff: a comparator driver can now
`device_claim` one and `mmio_map` it, and will see only its own 0x20-byte window. A child
whose window escapes the bus is refused — `busd` asserts that, and the `bus` test
asserts the kernel's table upholds it.
A USB device has *no* resources at all: `resource_count = 0`, because it's addressed
through its controller, not by MMIO. That case is allowed and is the common one.
## Families: sharing code between drivers
A "family" is two modules, not one:
- **A logic module** — the parts of the bus that every driver on it re-derives. Config
space walking and BAR decode for PCI. Descriptor parsing, control transfers, and hub
protocol for USB.
- **A protocol module** — the IPC message types that let a class driver talk to
*whatever* published its device. This is the part that makes class drivers portable.
danos already has one of each: `lib/device.zig` is a logic module,
[`lib/vfs-protocol.zig`](lib/vfs-protocol.zig) is a protocol module shared by `sbin/vfs.zig`
and its clients. The pattern generalises directly:
```
lib/
rt.zig module "rt" — syscalls, heap, ipc, dev, stdio
mmio.zig module "mmio" — volatile register access + barriers [M14]
bus/
pci.zig module "pci" — ECAM, BAR decode, capability walk
usb.zig module "usb" — descriptors, control transfers, hubs
proto/
vfs.zig module "proto.vfs" (today: lib/vfs-protocol.zig)
block.zig module "proto.block"
hid.zig module "proto.hid"
sbin/
xhcid.zig HCD + bus driver imports rt, pci, usb, mmio
usbhid.zig class driver imports rt, usb, proto.hid
blockd.zig class driver imports rt, proto.block
```
The only build change needed: [`addUserBinary`](build.zig) currently takes exactly one
module (`rt_mod`) and injects it. It should take a slice of modules. That's a
five-line change, and it's the *entire* mechanism — Zig modules already give you
everything else.
The discipline that makes this work: **a class driver must not import a bus's logic
module.** `usbhid` imports `proto.hid` and `usb` (for descriptor types), never `pci`.
If a class driver needs `mmio`, it has become an HCD and should be one.
## What exists today
- **M10** — `device_enumerate`, `device_claim`, `mmio_map`. Strong-uncacheable device
grants, `device_grant` teardown.
- **M11** — `irq_bind` / `irq_ack`. IRQ delivered as an IPC notification; mask before
EOI; `irq_ack` is the unmask.
- **M12** — `parent` in `DeviceDesc`, `device_register` with resource containment.
So: **bus drivers work now.** HCDs and class drivers do not. Here is exactly why, and
exactly what would fix it.
---
# Proposed ABI
## M13 — capability passing, for class drivers
**The blocker.** A class driver has to reach *its* device. Today the only way to find
an endpoint is the name registry: `ipc_register(service_id, h)` / `ipc_lookup(id)`,
where `ServiceId` is a global integer namespace with `max_services = 8`. You cannot
mint one endpoint per USB device that way, and there is no way for a bus driver to
*hand* a class driver an endpoint. M7 deferred this deliberately.
**The fix.** Let a message carry one handle. Sender names a handle in its own table;
the kernel installs the endpoint into the receiver's table (bumping `refcount`) and
tells the receiver the index it landed at.
```
ipc_call(h, msg, message_len, reply, reply_cap, send_cap) -> reply_len
ipc_reply_wait(h, reply, reply_len, recv, recv_cap, send_cap)
-> recv_len (rax), badge (rdx), received_cap (r8)
```
`send_cap` is a handle or `no_cap` (`~0`). `received_cap` is the index the transferred
endpoint was installed at in the receiver's table, or `no_cap`.
- Both calls grow from 5 args to 6, which fits: `syscall5` uses `rdi/rsi/rdx/r10/r8`,
leaving `r9`. `ipc_reply_wait` already returns two values via `setSyscallResult2`;
this needs a third (`setSyscallResult3`).
- If the receiver's handle table is full, the call fails `-ENOSPC` and **the message is
not delivered** — a half-delivered capability is worse than a failed send.
- `closeHandles` already drops references on exit, so the lifetime story is unchanged.
That single primitive gives you the standard `open` pattern:
```zig
// class driver // bus driver
const h = ipc.lookup(.usb).?; const r = ipc.replyWait(ep, ...);
const dev_ep = ipc.callCap(h, // ... mint a per-device endpoint,
.{ .op = .open, .id = dev_id }); // reply with it as send_cap
// now dev_ep is a private channel to that one device
```
## M14 — DMA memory and the memory-ordering contract, for HCDs
**The blocker.** An HCD is a DMA-engine programmer. It needs a descriptor ring the
device can read, which means memory that is (a) physically contiguous, (b) at a
physical address the driver knows, (c) of the right cacheability, and (d) pinned.
[`sysMmap`](src/kernel/process.zig) gives you *none* of the four: it calls `pmm.alloc()`
once per page, maps writeback-cached, and never reveals a physical address.
**The fix.**
```
dma_alloc(len, flags) -> vaddr (rax), paddr (rdx)
dma_free(vaddr, len) -> 0
flags: dma_coherent (1) uncacheable; the default and the only one that's portable
dma_wc (2) write-combining — needs PAT programmed; for framebuffers
dma_below_4g (4) for devices with 32-bit DMA addressing
```
Guarantees: page-aligned, physically contiguous, zeroed, pinned for the life of the
mapping, and the physical address is stable. It needs one thing the kernel lacks —
`pmm.allocContiguous(n, max_phys)`; today `pmm.alloc()` hands out one frame at a time
with no adjacency guarantee.
**The memory-ordering contract.** danos has, at the time of writing, **zero memory
barriers anywhere in the tree.** That is currently correct-by-accident and won't
survive the first DMA driver, or the first ARM boot.
`volatile` is not a barrier. In Zig it means: don't elide this access, and don't
reorder it against *other volatile* accesses. It says nothing about your *ordinary*
stores — the descriptor you just filled in normal WB memory — which LLVM may freely
sink past a volatile MMIO write. The canonical bug:
```zig
ring[i] = descriptor; // ordinary store to WB RAM
doorbell.* = i; // volatile store to UC MMIO
// nothing stops the compiler reordering these; the device reads a stale descriptor
```
So the rules, which belong in `lib/mmio.zig` and behind `arch`:
| Situation | Required |
|---|---|
| MMIO register read/write | `mmio.read` / `mmio.write` (volatile) |
| Fill DMA descriptor, then ring doorbell | `wmb()` between them |
| Woken by IRQ, then read what the device wrote | `rmb()` before the read |
| MMIO write that must complete before the next read | `mb()` |
And the per-arch lowering — the reason this must be an `arch` primitive and not a
sprinkling of `asm volatile`:
| | x86_64 | aarch64 |
|---|---|---|
| `mb()` | `mfence` | `dsb sy` |
| `rmb()` | `lfence` | `dsb ld` |
| `wmb()` | `sfence` | `dsb st` |
| DMA cache coherency | coherent; nothing to do | **not guaranteed**; needs non-cacheable buffers or cache maintenance |
x86 is forgiving here — TSO plus strong-uncacheable MMIO means you usually get away
with a compiler barrier alone. ARM is not, and [vision.md](vision.md) makes ARM the win
condition. Build the abstraction while there is one caller to fix.
(Zig note: `@fence` was **removed in 0.16**. Use `@atomicRmw(..., .seq_cst)` for a full
barrier, or per-arch inline asm — which is what `lib/mmio.zig` should hide.)
## M15 — interrupts for PCI devices
**The blocker, and it's a hard one.** No PCI device can take an interrupt today.
[`addBars`](src/device/acpi.zig) records `.memory` and `.io_port` BARs and never an
`.irq`; there is no `_PRT` parsing anywhere in the tree. `hpetd` only works because the
HPET advertises its own routing options in its own registers — a privilege no ordinary
device has.
**The fix, in two halves.**
*Legacy INTx*: parse `_PRT` from the DSDT to map (device, INTA–D) → GSI, and record it
as an `.irq` resource. Then `irq_bind` works unchanged. But INTx lines are **shared**,
and `irq.bound[gsi]` holds one endpoint. Sharing needs a list, and every driver on the
line must be polled on each interrupt — the reason everyone left INTx behind.
*MSI/MSI-X*, which is the real answer: per-device vectors, edge-triggered, unshared, no
mask/ack cycle, no 24-GSI ceiling. The kernel allocates a vector and hands the driver
the (address, data) pair to program into its own MSI capability:
```
msi_bind(dev_id, endpoint, out) -> 0 // out: extern struct { addr: u64, data: u32 }
```
The driver writes those into config space itself — which means it needs config space,
which means **discovery should give each `pci_device` a `.memory` resource for its
4 KiB ECAM slot**. That's a small change to `parseMcfg` and it unblocks the whole
capability walk (MSI, MSI-X, PCIe extended caps) without any new syscall.
Note QEMU's HPET reports `Tn_FSB_INT_DEL_CAP = 0` — no MSI — so `hpetd` can never
exercise this path. The first MSI driver will be the first PCI driver.
## M16 — the IOMMU, and the honest caveat
Everything above is capability-gated at the *CPU*. None of it is gated at the *device*.
A driver that can program a bus-mastering engine can make that device write to any
physical address, because page tables sit between the CPU and RAM, not between a device
and RAM. Until VT-d/DMAR (or SMMU on ARM) is programmed from the DMAR table, **`device_claim`
on any DMA-capable device is equivalent to granting ring 0.**
This does not make the model useless — it's the same position Linux is in with the
IOMMU off, and every other guarantee (crash isolation, restart, no shared address
space) still holds. But "user-space drivers are memory-safe" is not true yet, and the
gap should be named rather than implied.
## Ordering
`M13` (capability passing) is independent of `M14`/`M15` and is the cheapest. It
unlocks class drivers, which are the shape with no hardware requirements at all — you
could write a real one against `busd`'s comparators tomorrow.
`M14` and `M15` together unlock the first HCD. `M14`'s barrier layer is worth landing
on its own regardless: it's small, obviously correct, and stops every future driver
from hand-rolling `*volatile` and getting ARM wrong.
## See also
- [drivers.md](drivers.md) — how to write one, concretely.
- [discovery.md](discovery.md) / [acpi.md](acpi.md) — where the device table comes from.
- [ipc.md](ipc.md) — endpoints, badges, and the notification path an IRQ arrives on.
- [resilience.md](resilience.md) — restart, the reason any of this is worth the trouble.
+319
View File
@@ -0,0 +1,319 @@
# Writing a driver
In a monolithic kernel a driver is a function call away from everything: it runs in
ring 0, dereferences any physical address, and its interrupt handler *is* the ISR. In
danos a driver is **an ordinary ring-3 process**. It has its own address space, it
can crash without taking the kernel with it, and — the point of this document — it
can be restarted ([resilience](resilience.md)).
That leaves three questions the kernel has to answer, because a process can't answer
them for itself:
1. **What hardware exists?** → `device_enumerate`, over the device table discovery built
([discovery](discovery.md), [acpi](acpi.md)).
2. **How do I touch its registers?** → `device_claim` + `mmio_map`: the kernel maps the
device's physical MMIO window into your address space, and from then on it's plain
memory. No syscall per register access.
3. **How do I find out it wants something?** → `irq_bind`: the interrupt is delivered
to you as an IPC notification. You block; the hardware wakes you.
A driver is, in one sentence, *a process that sleeps until its device has something to
say.*
## The capability: claim before touch
The five driver syscalls (`src/root.zig`, dispatched in `src/kernel/process.zig`):
| # | Call | Meaning |
|---|------|---------|
| 11 | `device_enumerate(buf, max) -> total` | Snapshot the device table |
| 12 | `device_claim(id) -> ok` | Take **exclusive** ownership |
| 13 | `mmio_map(id, res_idx) -> vaddr` | Map a claimed device's register window |
| 14 | `irq_bind(id, res_idx, endpoint)` | Deliver that device's IRQ as a notification |
| 15 | `irq_ack(id, res_idx)` | Re-arm the IRQ after servicing the device |
| 16 | `device_register(parent_id, desc) -> id` | Publish a child of a device you claimed |
Notice that **nothing takes a physical address or an interrupt number.** Every call
names a device by id and a resource by index. That indirection is the entire security
model. If `mmio_map` took a physical address, any process could map the kernel's
memory; if `irq_bind` took a GSI, any process could bind the keyboard's line and
silently intercept it. Instead the kernel checks two things (`process.ownedGsi`, and
the same check at the top of `sysMmioMap`):
- `device_service.ownerOf(dev_id) == me` — you claimed it, and claims are exclusive
- the resource at `res_idx` is of the right *kind* — `memory` for `mmio_map`, `irq`
for `irq_bind`
The claim is the capability. Everything else follows from it.
## Registers: `mmio_map`
`mmio_map` walks the caller's page tables and installs the device's physical frames
with `present | user | writable | nx | pcd | pwt`
(`arch/x86_64/paging.zig:mapUserDeviceInto`). Two of those bits are load-bearing:
- **`pcd | pwt`** — strong-uncacheable. A device register is not memory; a cached read
would return a stale value and a write might never leave the CPU.
- **`device_grant`** (bit 9, one of the PTE's available bits) — marks the leaf as MMIO
rather than RAM, so `freeSubtree` skips `pmm.free` on it when the address space is
destroyed. Without this, killing a driver would hand the HPET's registers back to
the frame allocator as if they were free RAM. The `iopass` test guards it.
Grants land in their own arena, `0x0000_7100_0000_0000` (PML4[226]), so device pages
never widen an existing mapping.
Then you just… use it:
```zig
const base = dev.mmioMap(dev_id, mmio_res) orelse return;
const counter: *volatile u64 = @ptrFromInt(base + 0xF0);
const now = counter.*; // a load, straight to the hardware. no kernel involved.
```
## Interrupts: the cycle, and why it has that shape
An interrupt handler in a microkernel has a problem. The code that knows how to quiet
the device is in ring 3, in another address space, and it will not run for
microseconds or milliseconds — after a context switch, when the scheduler gets to it.
But the CPU wants an EOI *now*, and a **level-triggered** line stays asserted until
the device is quieted. EOI a still-asserted line and the I/O APIC redelivers
immediately. Forever. The driver never gets to run at all.
The way out is to mask the line before acknowledging it:
```
kernel ISR irqMask(gsi) // line still asserted; stop it reaching a CPU
irqEoi() // now safe to tell the LAPIC we're done
notifyFromIsr() // wake the driver — it runs much later
driver replyWait() -> badge with the notify bit set
<clear the device's status register> // NOW the line deasserts
irq_ack(dev, res) // kernel unmasks: quiet, so it can't refire
```
`irq_ack` is not bookkeeping you could skip. **It is the unmask.** Forget it and the
interrupt fires exactly once, ever; call it before the device is quiet and you get an
interrupt storm. That single fact explains why `irq_bind` and `irq_ack` are two
syscalls and not one.
This is also why `interruptDispatch` (`arch/x86_64/idt.zig`) no longer issues the EOI
itself. It used to, before running the handler — correct for the LAPIC timer, and
impossible for a routed device line. Each handler now owns its EOI, because only the
handler knows which discipline its source needs.
### The driver side is an event loop, not a callback
`IPC_ReplyWait` returns *either* a client request *or* a notification, told apart by
the top bit of the badge (`ipc_sync.notify_badge_bit`). So a driver is one
single-threaded loop over both of its event sources:
```zig
while (true) {
const r = ipc.replyWait(endpoint, reply, &recv);
if (r.isNotification()) { // r.source() is the GSI
service_device(); // clear the status register
_ = dev.irqAck(id, irq_res); // re-arm
} else {
handle_client_request(recv[0..r.len]);
}
}
```
No reentrancy, no "what am I allowed to call from an interrupt handler", no shared
state between ISR and task context. The interrupt is just a message.
Two properties worth knowing:
- **An interrupt taken while you're elsewhere is not lost.** If the driver is off in
an `ipc_call` to another server when the IRQ fires, `wakeLocked` finds nobody
waiting, but the badge is already on the endpoint's notify ring. The next
`replyWait` pops it (`ipc_sync.replyWait` checks `popNotify` before the sender FIFO).
- **Notifications coalesce, they don't count.** The ring is 8 deep and drops on
overflow. That's correct: an IRQ notification is a *level* ("the device wants
attention"), not a tally. Re-read the device's status register; never assume one
notification means exactly one event. Because the ISR masks the line until you
`irq_ack`, at most one badge per GSI can be outstanding — so the ring can only
overflow if you bind more than eight GSIs to a single endpoint. Don't.
## A whole driver
`sbin/hpetd.zig` is ~150 lines and does all of it. The shape:
```zig
const hpet = findHpet(buf) orelse return; // device_enumerate, look for
// class=timer with memory + irq
_ = dev.claim(hpet.dev_id); // the capability
const base = dev.mmioMap(hpet.dev_id, hpet.mmio).?;
const endpoint = ipc.createEndpoint().?;
// program the hardware over the mapping we were just handed
reg(base, 0x100).* = level | int_enb | (hpet.gsi << 9); // timer 0 config
reg(base, 0x108).* = reg(base, 0xF0).* + period; // comparator
reg(base, 0x010).* |= 1; // ENABLE
_ = dev.irqBind(hpet.dev_id, hpet.irq, endpoint);
while (...) {
const r = ipc.replyWait(endpoint, &.{}, &recv); // blocked. not polling.
if (r.badge & notify_bit == 0) continue;
reg(base, 0x020).* = 1; // clear status -> deassert
reg(base, 0x108).* = reg(base, 0xF0).* + period; // re-arm
_ = dev.irqAck(hpet.dev_id, hpet.irq); // unmask
}
```
The HPET is a good first driver for a reason that isn't obvious. Its *counter* is a
clocksource — the only way to use it is to read it, so it proved `mmio_map` without
needing interrupts at all. Its *comparators* are a clockevent, and can be configured
**level-triggered** (`Tn_INT_TYPE_CNF`), which asserts a bit in `GENERAL_INT_STATUS`
that the driver must write-1-to-clear. That's a genuine deassert step, so the full
mask/ack cycle above is exercised for real rather than being decoration on an
edge-triggered line that would have been fine without it.
One wrinkle it also demonstrates: the ACPI HPET table carries **no interrupt number**.
Which I/O APIC inputs a comparator may drive is a bitmask in `Tn_INT_ROUTE_CAP`, in
the device's own registers. So discovery (`acpi.parseHpet`) maps the block, reads the
mask, and records one concrete GSI as an `irq` resource. The driver then programs
`Tn_INT_ROUTE_CNF` to raise exactly that line — and the kernel will only bind the one
it recorded. Hardware that describes itself at runtime still has to fit through a
static capability.
## Publishing children: `device_register`
A device that *contains other devices* — a PCI bridge, a USB hub, or the HPET's block
of comparators — needs a driver that enumerates it and tells the kernel what it found.
That's `device_register`, and it makes the device table a tree rather than a list
(`DeviceDesc.parent`).
```zig
var child = std.mem.zeroes(dev.DeviceDesc);
child.class = @intFromEnum(dev.DeviceClass.timer);
child.resource_count = 1;
child.resources[0] = .{ .kind = memory, .start = bus_base + 0x100, .len = 0x20 };
const child_id = dev.register(bus_id, &child).?;
```
The child is left **unclaimed**, which is the whole point: another process claims it and
`mmio_map`s it, and sees only that 0x20-byte window.
The rule the kernel enforces is **containment**: every resource of a child must lie
inside a resource of the same kind on its parent. Ranges must nest; an IRQ must match
exactly. This isn't bureaucracy — a `DeviceDesc` is a licence to map physical memory, so
without containment `device_register` would be a syscall for mapping any page you like. A
bus driver may only ever subdivide what it already owns.
A device with **no resources** is legal and common. A USB device is reached through its
controller, not by MMIO, so it gets `resource_count = 0`.
See [`sbin/busd.zig`](../sbin/busd.zig) for a complete one, and
[driver-model.md](driver-model.md) for how bus drivers, class drivers and host
controller drivers fit together.
## What the kernel does not do for you
- **It does not quiet your device.** That's the whole reason `irq_ack` exists.
- **It does not know your registers.** `mmio_map` hands you a base address; every
offset in this document came from the HPET spec, not from danos.
- **It does not serialise your driver.** Two clients calling one driver endpoint are
serialised by `replyWait`, but nothing stops your driver from being preempted.
## Limits, today
Worth knowing before you write the second driver:
- **Ring 3 has no port I/O.** The TSS I/O permission bitmap is absent
(`tss.zig`: `iomap_base = @sizeOf(Tss)`), and IOPL is never raised, so `in`/`out`
from a driver is a #GP. That rules out a user-space 16550 UART (`0x3F8`), PS/2
(`0x60`/`0x64`), and legacy PCI config (`0xCF8`/`0xCFC`). Everything must be MMIO.
`io_port` resources are recorded by discovery and then ignored.
- **Page granularity.** `mmio_map` rounds to 4 KiB. Two devices sharing a page means
granting one grants the other. A `device_register`ed child's *resource* can be narrower
than a page, but its *mapping* can't.
- **No DMA memory.** `mmap` gives you writeback-cached, non-contiguous pages and never
tells you their physical address, so you cannot build a descriptor ring. Any driver
for a bus-mastering device is blocked on this.
- **No memory barriers.** There are none in the tree, and `volatile` is not one — it
won't stop the compiler sinking an ordinary store (your DMA descriptor) past a
volatile MMIO store (your doorbell). On x86 you mostly get away with it; on ARM you
will not. See [driver-model.md](driver-model.md#m14).
- **DMA is not contained.** A driver that can program a bus-mastering device can make
that device write to *any* physical address — page tables don't sit between a device
and RAM; an IOMMU does. Until VT-d/DMAR is programmed, `device_claim` on a DMA-capable
device is effectively equivalent to granting ring 0. This is the largest gap between
the design's promise and what it delivers.
- **One endpoint per GSI**, so shared legacy PCI INTx lines can't be split between two
drivers. MSI/MSI-X — one vector per device, edge-triggered, unshared — is the real
answer, and QEMU's HPET doesn't offer it (`Tn_FSB_INT_DEL_CAP = 0`).
- **Polarity is hardcoded** active-high in `irq.bind`. A device whose MADT override
says active-low needs that threaded through from discovery.
- **14 device vectors** (33–46) and **24 GSIs**, bounded by the stubs `isr.s` emits and
by a single I/O APIC.
- **Don't bind more than 8 GSIs to one endpoint.** The notify ring is 8 deep and drops
on overflow. With one GSI per endpoint that's unreachable — the line is masked from
the ISR until `irq_ack`, so at most one badge is ever outstanding. Bind nine devices
to one endpoint, though, and a dropped badge leaves that line masked with nobody
left to ack it.
- **A faulting driver still kills the machine.** There is no per-process kill path: a
ring-3 page fault halts the kernel, so `releaseIrqs` runs only on a voluntary
`exit`. Fault isolation is the whole premise ([vision](vision.md)) and it is
[not built yet](resilience.md).
- **A dead driver's device is not reclaimed.** `releaseIrqs` unbinds and masks the
line on exit, but the claim is never released — restart is
[not built](resilience.md).
- **On real hardware, the mask/EOI cycle may need a remote-IRR flush.** Masking a
level-triggered redirection entry with remote-IRR set doesn't clear it on some
chipsets, and the line never fires again. QEMU clears it on EOI regardless, so the
tests can't see this. Linux flushes remote-IRR by toggling the entry to edge and
back. See the note at the top of `src/kernel/irq.zig`.
## Verifying it
The `hpet` test spawns `hpetd` from the initrd and watches the serial log. The driver
prints `hpetd: ok` only after being woken five times, and its loop's only exit is
through `replyWait` returning a notification — it cannot reach that line by polling.
The last check doesn't trust the driver's self-report at all: the kernel reads the I/O
APIC redirection entry back and asserts the line really is routed to a device vector,
really is level-triggered, and really was left unmasked by the driver's final
`irq_ack`.
```
$ python3 test/qemu_test.py hpet irqfree iopass
hpet ... PASS (matched 'DANOS-TEST-RESULT: PASS')
irqfree ... PASS (matched 'DANOS-TEST-RESULT: PASS')
iopass ... PASS (matched 'DANOS-TEST-RESULT: PASS')
```
Two companions cover what `hpetd` can't, because it never exits:
- **`irqfree`** — the teardown path. Binds two owners to one shared endpoint, releases
one, and reads the I/O APIC back: the departing owner's line is masked, the sibling's
is not. That second half is why bindings are keyed on the owning *task* and not on
the endpoint pointer — endpoints are shared, so releasing "everything pointing at
this endpoint" would silently mask a live driver's device.
- **`iopass`** — the `device_grant` teardown rule, so destroying a driver's address
space never returns MMIO frames to the RAM pool.
## What's next (not done here)
The big ones — capability passing (class drivers), DMA + barriers and MSI (host
controller drivers), and the IOMMU — have proposed signatures in
[driver-model.md](driver-model.md). Smaller items:
- **Port I/O grants**, so a PS/2 or 16550 driver is possible: either a per-device TSS
I/O permission bitmap swapped on context switch, or `io_in`/`io_out` syscalls gated
by the same claim. The legacy devices that need it are all low-rate, so the syscall
is likely fast enough.
- **Releasing a claim.** There is no `dev_release`, and `device_service` never drops a claim on
exit — only IRQ bindings are released. A dead driver's device stays owned forever,
which blocks restart.
- **Unregistering children.** `device_register` only appends. A USB device that is
unplugged cannot be removed, and a bus driver in a loop can exhaust the 64-entry
table.
- **Restart.** A driver that dies should release its claim, have its device quiesced,
and be respawned by a supervisor. Some pieces (`releaseIrqs`, `device_grant`
teardown, the claim table) exist; the policy doesn't.
- **Interrupt priority / threaded IRQ latency.** `notifyFromIsr` enqueues the woken
driver but doesn't preempt (`wakeLocked` deliberately leaves that to the caller), so
a woken driver waits for the next scheduling point.
+55 -12
View File
@@ -6,12 +6,20 @@ just call each other — a request becomes a **message**. In a microkernel, what
was a function call across a monolithic kernel is IPC, so it's a first-class
concern, not an afterthought.
This first form is a **bounded blocking channel** (`src/kernel/ipc.zig`): a fixed-size
ring buffer of messages with a producer/consumer rendezvous, built on the
scheduler's [wait queues](scheduling.md).
There are two layers, built a milestone apart:
- **`src/kernel/ipc.zig`** — a bounded blocking channel between *kernel threads*,
described below. The primitive, and where the blocking discipline was worked out.
- **`src/kernel/ipc-synchronous.zig`** — synchronous call/reply between *processes*, across
address spaces. What user-space servers and drivers actually talk over. It's the
second half of this document.
## The channel
The first form is a **bounded blocking channel** (`src/kernel/ipc.zig`): a fixed-size
ring buffer of messages with a producer/consumer rendezvous, built on the
scheduler's [wait queues](scheduling.md).
`Channel(T, capacity)` is generic over the message type and buffer size. It holds a
ring buffer, a count, and two wait queues:
@@ -42,16 +50,51 @@ full and empty over and over, so both the blocking-send and blocking-recv paths
exercised heavily. The messages arrive intact and in order (their sum is the
expected `5050`), and neither task busy-waits — they block and wake each other.
## Endpoints: call/reply across address spaces
A channel connects two kernel threads sharing one address space. Real servers are
*processes*, so the payload has to cross an address-space boundary. That's
`src/kernel/ipc-synchronous.zig`, and its shape is L4's: a synchronous **rendezvous** at an
`Endpoint`, with the message copied directly from the sender's pages to the receiver's
(`copyAcross` walks both sets of page tables through the physmap — no CR3 switch, no
bounce buffer).
Two syscalls carry it:
- **`ipc_call(h, msg, reply)`** — copy `msg` to the server, block until it replies.
- **`ipc_reply_wait(h, reply, recv)`** — reply to the client you're still holding (if
any), then block for the next request. One syscall, because a server's steady state
is *always* "finish the last one, wait for the next".
An endpoint is reached by **handle** — a small integer index into the process's handle
table (`Task.handles`), exactly like a file descriptor, and just as unforgeable. The
bootstrap problem (how do you get the first handle?) is solved by a tiny name registry:
a server calls `ipc_register(service_id, h)` under a well-known small integer, and a
client calls `ipc_lookup(service_id)`.
The server never learns the client's identity beyond a **badge**, delivered alongside
the message: the caller's task id.
### Interrupts are messages too
`notifyFromIsr` posts an *asynchronous* notification to an endpoint — no payload, no
reply owed — and wakes whoever is blocked in `reply_wait`. Its badge has the top bit
set (`notify_badge_bit`), which is how a driver's single event loop distinguishes "a
client wants something" from "the hardware wants something". Notifications sit in a
small coalescing ring on the endpoint, so an interrupt taken while the driver was busy
elsewhere is not lost.
This is what makes a user-space driver possible at all, and it's the subject of
[drivers.md](drivers.md).
## What's next (not done here)
- **Across address spaces.** Today both endpoints are kernel threads sharing the
kernel's memory, so the message is copied within one address space. When user
mode arrives, the same channel carries messages between *isolated* processes,
copying the payload across the boundary — which is where IPC earns its place as
the microkernel's backbone.
- **Synchronous call/reply.** A request/response pattern (send-and-wait-for-reply)
on top of channels, the shape most driver/service calls take.
- **Interrupts as messages.** A hardware interrupt delivered to the driver task
that owns the device, as an IPC message.
- **Priority inheritance** through IPC, so a high-priority client blocked on a
low-priority server doesn't suffer unbounded priority inversion.
- **Handle transfer.** A server can't hand a client a handle to a third endpoint, so
every capability is either well-known (the registry) or inherited — there's no way
to delegate one.
- **Asynchronous / buffered send** for the cases where a rendezvous is the wrong
shape (logging, notifications between servers).
- **A bounded reply.** `MSG_MAX` is 256 bytes and the copy runs under the big kernel
lock; a bulk transfer wants shared pages, not a copy.