M11–M12: IRQ-as-IPC and bus drivers; expand names tree-wide
Two driver-model milestones plus a tree-wide naming pass. Suite 35/35 (QEMU) + host tests green. M11 — IRQ-as-IPC. A ring-3 driver now sleeps until its device interrupts it. New src/kernel/irq.zig: per-GSI endpoint bindings, comptime per-vector trampolines, dispatch = mask GSI -> LAPIC EOI -> notifyLocked, all under one lock region. irq_bind/irq_ack syscalls, gated by the device claim like mmio_map. interruptDispatch no longer EOIs — each handler owns its EOI, because a level line must be masked before it is acknowledged (irq_ack is the unmask). Bindings are keyed on the owning task and released on exit (a shared endpoint's siblings survive). hpetd rewritten interrupt-driven. Tests: hpet (rewritten, reads back the I/O APIC routing) and irqfree. M12 — bus drivers. DeviceDesc gains a parent, making the device table a tree. dev_register (device_register) lets a process publish children below a device it claimed; the kernel enforces resource containment (a child's resources must nest in its parent's), so a descriptor can't fabricate a window over kernel RAM. Descriptor copied in via copyFromUser (physmap walk — an unmapped user pointer fails the call instead of faulting the kernel). Per-parent child cap bounds table exhaustion. sbin/busd.zig is a worked bus driver. Test: bus. Naming — per docs/coding-standards.md: non-acronym abbreviations spelled out (message, descriptor, device_service, scheduler, runtime, physical, interpreter, ...); acronyms kept (IPC, MMIO, DMA, HCD, ...); files are kebab-case (ipc-synchronous.zig, device-service.zig, vfs-protocol.zig, ...). Exceptions: POSIX/C ABI names and Zig idioms (init/len/ptr) kept. Module collisions resolved by specific naming (config -> parameters, device.zig alias -> device_model). AML op/Op disambiguated: op = opcode, Op = operation; per-opcode parse handlers renamed opX -> parseX. New driver docs: drivers.md, driver-model.md (bus/class/HCD shapes + the proposed M13–M16 ABI), coding-standards.md.
This commit is contained in:
+37
-7
@@ -35,9 +35,20 @@ rather than restate it. Roughly in the order things happen at runtime:
|
||||
multitasking: kernel threads, the context switch, O(1) priority selection, and
|
||||
blocking (sleep, wait queues) — the leap to a running system.
|
||||
11. **[ipc.md](ipc.md) — inter-process communication.** Bounded blocking
|
||||
message-passing channels — the backbone the microkernel's isolated servers will
|
||||
talk over.
|
||||
12. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
|
||||
message-passing channels, then synchronous call/reply between *processes* over
|
||||
endpoints — the backbone the microkernel's isolated servers talk over.
|
||||
12. **[syscall.md](syscall.md) — system calls.** How ring 3 asks the kernel for
|
||||
something: the `syscall`/`sysret` fast path, the trap frame, and why the table is
|
||||
deliberately tiny.
|
||||
13. **[drivers.md](drivers.md) — writing a driver.** The payoff: a driver is an
|
||||
ordinary ring-3 process that claims a device, maps its registers, and **sleeps
|
||||
until its hardware interrupts it**. The claim is the capability; `irq_ack` is the
|
||||
unmask.
|
||||
14. **[driver-model.md](driver-model.md) — buses, classes and host controllers.** How
|
||||
real driver stacks factor into three shapes, how families share code, and the
|
||||
proposed ABI for the three primitives still missing (capability passing, DMA +
|
||||
memory barriers, MSI).
|
||||
15. **[halting.md](halting.md) — halting.** Why a kernel can't just "exit", and
|
||||
how `while (true) hlt` parks the CPU safely once there's nothing left to do.
|
||||
|
||||
Start with the north star:
|
||||
@@ -68,6 +79,10 @@ Cutting across all of these:
|
||||
- **[smp.md](smp.md) — multiple cores.** A design/research note on how microkernels
|
||||
(L4, seL4) handle SMP — big kernel lock vs per-CPU vs multikernel — and how the
|
||||
right choice depends on whether danos is chasing real-time or resilience.
|
||||
- **[coding-standards.md](coding-standards.md) — coding standards.** The naming rule the
|
||||
tree follows: non-acronyms are spelled out in full (`message`, not `msg`), files are
|
||||
`kebab-case`, code follows Zig's case conventions, and the handful of exceptions
|
||||
(POSIX/C ABI names, `init`/`len`/`ptr`, acronyms).
|
||||
- **[sysv.md](sysv.md) — the calling convention.** What "the kernel is SysV" means,
|
||||
and why the loader→kernel boundary has to pin it (the RDI-vs-RCX handoff).
|
||||
- **[testing.md](testing.md) — testing.** How the kernel is tested by booting it in
|
||||
@@ -94,19 +109,34 @@ passing messages over **[IPC](ipc.md)** channels — runs, its CPU-specific bits
|
||||
behind the [arch](arch.md) boundary, and when idle, or on a panic, it **halts**
|
||||
([halting.md](halting.md)).
|
||||
|
||||
Above that line the microkernel proper begins: **discovery** ([discovery.md](discovery.md),
|
||||
[acpi.md](acpi.md)) learns what hardware exists, ring-3 processes ask the kernel for
|
||||
things through the small **[syscall](syscall.md)** table, isolated servers reach each
|
||||
other over IPC **endpoints** ([ipc.md](ipc.md)), and a **[driver](drivers.md)** claims
|
||||
a device, maps its registers, and sleeps until the hardware interrupts it — which is
|
||||
the whole reason for the arrangement ([vision.md](vision.md)).
|
||||
|
||||
## Source map
|
||||
|
||||
| Area | Code |
|
||||
|------|------|
|
||||
| Boot methods (one per way of booting the kernel) | `src/boot/` — `efi.zig` (UEFI) → `BOOTX64.efi` |
|
||||
| Kernel entry, panic, bring-up | `src/kernel/main.zig` |
|
||||
| Shared loader↔kernel contract (`BootInfo`, `Framebuffer`, `MemoryMap`, ABI) | `src/root.zig` |
|
||||
| Shared loader↔kernel contract (`BootInfo`, `Framebuffer`, `MemoryMap`, `Syscall`, ABI) | `src/root.zig` |
|
||||
| Physical frame allocator | `src/kernel/pmm.zig` |
|
||||
| Kernel heap (`std.mem.Allocator`) | `src/kernel/heap.zig` |
|
||||
| Scheduler (fixed-priority preemptive; blocking, wait queues) | `src/kernel/sched.zig` |
|
||||
| IPC channels (message passing) | `src/kernel/ipc.zig` |
|
||||
| Scheduler (fixed-priority preemptive; blocking, wait queues) | `src/kernel/scheduler.zig` |
|
||||
| Big kernel lock + interrupt-safe critical sections | `src/kernel/sync.zig` |
|
||||
| IPC channels between kernel threads (message passing) | `src/kernel/ipc.zig` |
|
||||
| IPC endpoints: cross-address-space call/reply, handles, notifications | `src/kernel/ipc-synchronous.zig` |
|
||||
| User processes: ELF loading, address spaces, the syscall table | `src/kernel/process.zig` |
|
||||
| Device tree + claim capability + `device_register` containment | `src/kernel/device-service.zig` |
|
||||
| IRQ-as-IPC: routing a device interrupt to a driver's endpoint | `src/kernel/irq.zig` |
|
||||
| Hardware discovery (ACPI/device tree) behind one neutral device model | `src/device/` |
|
||||
| Framebuffer text console (mirrors to serial) | `src/kernel/console.zig` |
|
||||
| In-kernel test cases | `src/kernel/tests.zig` |
|
||||
| Arch-specific kernel code (`halt`, GDT/IDT/TSS, exception + interrupt stubs, page tables, APIC/timer, serial, linker script) | `src/kernel/arch/x86_64/` |
|
||||
| Arch-specific kernel code (`halt`, GDT/IDT/TSS, exception + interrupt stubs, page tables, APIC/IO-APIC/timer, serial, linker script) | `src/kernel/arch/x86_64/` |
|
||||
| User runtime library (`rt`): syscalls, heap, stdio, IPC, device access | `lib/` |
|
||||
| User-space programs shipped in the initrd (`init`, `vfs`, `hpetd` leaf driver, `busd` bus driver) | `sbin/` |
|
||||
| Build + `run-x86-64` (QEMU/OVMF) | `build.zig` |
|
||||
| QEMU integration test harness | `test/qemu_test.py` |
|
||||
|
||||
@@ -0,0 +1,137 @@
|
||||
# Coding standards
|
||||
|
||||
Conventions for danos source. The overriding one, from which most of the rest follows:
|
||||
|
||||
> **Names are spelled out in full. An identifier is not abbreviated unless the
|
||||
> abbreviation is an acronym.**
|
||||
|
||||
`interruptDispatch`, not `intDisp`. `message_len`, not `message_len` (`msg` expands, `len`
|
||||
is a Zig idiom — see the exceptions). `device_service`, not `device_service`. `scheduler`, not
|
||||
`sched`. The cost of a longer name is paid once, at the keyboard; the cost of a
|
||||
cryptic one is paid every time the code is read, by everyone who reads it. In a
|
||||
microkernel whose whole argument is that a human can hold each piece in their head,
|
||||
that trade is not close.
|
||||
|
||||
## The rule, precisely
|
||||
|
||||
**Acronyms and initialisms stay.** They *are* the full name — expanding them would make
|
||||
the code worse, not better. `IPC`, `MMIO`, `DMA`, `IRQ`, `TSS`, `GDT`, `IDT`, `APIC`,
|
||||
`GSI`, `HPET`, `ACPI`, `PCI`, `EOI`, `BAR`, `ECAM`, `MSI`, `CPU`, `ELF`, `ABI`, `UEFI`,
|
||||
`MMU`, `TLB`, `ISR`, `ISA`, `GAS`, `HAL`, `PMM`, `VMM`, `VFS`, `HID`, `HCD`, `SMP`,
|
||||
`AML`, `MADT`, `MCFG`, `FADT`, `RSDP`, `XSDT`, `RSDT`, `GOP`, `EDID`, `TSC`, `PIT`,
|
||||
`RTC`, `LAPIC`, `SIPI`. In code they carry whatever case the surrounding convention
|
||||
demands: `Hal` the type, `hal` the variable, `mapMmio` the function.
|
||||
|
||||
**Everything else is spelled out.** If it's a word with letters removed, restore them:
|
||||
|
||||
| Abbreviation | Full |
|
||||
|---|---|
|
||||
| `proto` | `protocol` |
|
||||
| `msg` | `message` |
|
||||
| `desc` | `descriptor` |
|
||||
| `res` | `resource` |
|
||||
| `recv` | `receive` |
|
||||
| `buf` | `buffer` |
|
||||
| `cur` | `current` |
|
||||
| `src` / `dst` | `source` / `destination` |
|
||||
| `idx` | `index` |
|
||||
| `addr` | `address` |
|
||||
| `reg` | `register` |
|
||||
| `prev` | `previous` |
|
||||
| `cfg` / `config` | `configuration` |
|
||||
| `arch` | `architecture` |
|
||||
| `sched` | `scheduler` |
|
||||
| `dev` | `device` |
|
||||
| `sys` / `syscall` | `system` / `system_call` |
|
||||
| `info` | `information` |
|
||||
| `dt` | `device_tree` |
|
||||
| `ep` | `endpoint` |
|
||||
| `rt` | `runtime` |
|
||||
| `func` | `function` |
|
||||
| `phys` / `virt` | `physical` / `virtual` |
|
||||
| `wq` | `wait_queue` |
|
||||
|
||||
This list is illustrative, not exhaustive. The rule is the rule; when you meet a new
|
||||
abbreviation, expand it.
|
||||
|
||||
## Exceptions
|
||||
|
||||
Three, and only three.
|
||||
|
||||
1. **Foreign ABI names are spelled exactly as the ABI spells them.** A function that
|
||||
*is* the C or POSIX interface keeps its name: `fopen`, `fwrite`, `fread`, `malloc`,
|
||||
`calloc`, `realloc`, `free`, `memcpy`, `mmap`, `munmap`, `open`, `read`, `write`,
|
||||
`close`, `lseek`, `stat`, `errno`. We don't get to rename `fwrite` to
|
||||
`fileWrite` — it wouldn't be `fwrite` any more. This also covers the syscall
|
||||
*wrappers* that exist to match those names. It does **not** license inventing new
|
||||
abbreviated names in that style.
|
||||
|
||||
2. **Zig idioms are spelled the way Zig spells them.** Three names are the language's,
|
||||
not ours, and are left alone:
|
||||
- **`init` / `deinit`** — the constructor convention (`std.ArrayList.init`), not a
|
||||
shortening of "initialize".
|
||||
- **`len` / `ptr`** — the slice field names (`slice.len`, `slice.ptr`). Our own
|
||||
structs use bare `len`/`ptr` fields to mirror them, so a reader carries one
|
||||
mental model. (Compounds still expand: a field is `message_len`, not
|
||||
`message_length` — `len` is kept, `msg` is not.)
|
||||
- The builtins (`@min`, `@max`, `@memcpy`) and `allocator.alloc` / `.create` are
|
||||
Zig's spelling.
|
||||
|
||||
The rule governs the names *we* coin.
|
||||
|
||||
3. **Single-letter variables in a trivial local scope.** `for (items) |item, i|` may
|
||||
keep `i`; a coordinate may be `x`, `y`. The moment the scope is big enough that the
|
||||
letter's meaning isn't obvious on sight, give it a real name. When in doubt, name it.
|
||||
|
||||
4. **Established Unix filesystem and program conventions.** Top-level directories keep
|
||||
their conventional names — `src`, `lib`, `sbin`, `bin`, `docs` — as do daemon
|
||||
programs by their `d` suffix (`hpetd`, `busd`, following `sshd`/`httpd`). These are
|
||||
names a Unix reader already knows; expanding them fights the convention rather than
|
||||
serving it.
|
||||
|
||||
## A note on collisions
|
||||
|
||||
Two identifiers can legitimately expand to the same word. When they do, keep both
|
||||
meaningful by renaming one to its *specific* identity rather than the generic
|
||||
expansion. Two cases resolved this way:
|
||||
|
||||
- The `config` module (compile-time tunables — `maximum_cpus`, `timer_hz`) would
|
||||
collide with `cfg` (a `PlatformConfiguration` value) at `configuration`. The module
|
||||
became **`parameters`**, which is what it holds.
|
||||
- The kernel `device.zig` module would collide with `dev` (a device value) at
|
||||
`device`. The module alias became **`device_model`**, which is what it is — the
|
||||
device data model (`Device`, `DeviceTree`, `ResourceKind`).
|
||||
- The `Namespace` module alias (`ns`/`nsp` across the AML files) collides with a
|
||||
`Namespace` **instance**. Resolved by dropping the module alias entirely — the two
|
||||
types it provided are imported directly (`const Node = @import("namespace.zig").Node;`)
|
||||
— which frees `namespace` for the instance.
|
||||
|
||||
A related case is one abbreviation with two meanings. In the AML code, `op` means
|
||||
**opcode** (`opcodes.zig`, the `*_opcode` constants) but `Op` in `BinaryOperation` /
|
||||
`LogicOperation` means **operation** — distinguished by case. The per-opcode parser
|
||||
handlers, formerly `opName`/`opField`, are `parseName`/`parseField`: they *parse* the
|
||||
opcode's structure, which says what they do without overloading "op".
|
||||
|
||||
## Case and file names
|
||||
|
||||
Within those spelling rules, follow Zig's own conventions:
|
||||
|
||||
- **Types** — `PascalCase`: `DeviceDescriptor`, `Endpoint`, `WaitQueue`.
|
||||
- **Functions** — `camelCase`: `mapUserDeviceInto`, `notifyFromIsr`.
|
||||
- **Variables, fields, constants** — `snake_case`: `message_length`, `device_service`,
|
||||
`notify_badge_bit`.
|
||||
|
||||
**File names are `kebab-case`.** A file named for a multi-word thing hyphenates it:
|
||||
`device-tree.zig`, `ipc-synchronous.zig`, `vfs-protocol.zig`, `device-service.zig`. A
|
||||
single word or acronym needs no hyphen: `scheduler.zig`, `paging.zig`, `apic.zig`,
|
||||
`idt.zig`. (The module *alias* a file is imported under still follows the code
|
||||
conventions above — `snake_case` — because it's an identifier, not a filename.)
|
||||
|
||||
## Why acronyms are the line
|
||||
|
||||
Because an acronym has no letters to restore. `MMIO` doesn't become "memory mapped
|
||||
input output" in code — that expansion is what the acronym *is for*. But `msg` is just
|
||||
`message` with three letters stolen, and stealing them buys nothing a reader wants. The
|
||||
test for "is this an abbreviation I must expand" is simply: *is there a longer word this
|
||||
is a clipped form of?* If yes, write the word. If it's an initialism standing in for a
|
||||
phrase, leave it.
|
||||
+32
-11
@@ -89,7 +89,6 @@ if (state.vector < 32) {
|
||||
on_fault(state); // exception: report and halt (never returns)
|
||||
} else if (handlers[state.vector]) |handler| {
|
||||
handler(); // device: run the registered handler
|
||||
apic.eoi(); // ...acknowledge the LAPIC
|
||||
}
|
||||
// else: spurious/unhandled — deliberately no EOI
|
||||
```
|
||||
@@ -100,10 +99,23 @@ Two things make device interrupts *return* where exceptions don't:
|
||||
flows back to `isr_common`, which restores every register it saved and executes
|
||||
`iretq` — resuming the interrupted instruction exactly. (This is why the stub
|
||||
saves *all* the general registers.)
|
||||
2. **End-of-interrupt.** After handling, we write the LAPIC's EOI register. Miss
|
||||
2. **End-of-interrupt.** Somewhere in there we write the LAPIC's EOI register. Miss
|
||||
this and the LAPIC thinks we're still busy and never delivers the next
|
||||
interrupt. It's the single most common "my timer fired once and stopped" bug.
|
||||
|
||||
**Each handler issues its own EOI**, rather than the dispatcher doing it around the
|
||||
call. That looks like a needless devolution while the timer is the only device, and
|
||||
`apic.timerTick` indeed does nothing but `eoi()` before bumping its counter (early,
|
||||
because the tick hook is the scheduler, which may switch tasks and not return
|
||||
promptly — the LAPIC mustn't wait on it).
|
||||
|
||||
It stops looking needless with the second device. A *routed* interrupt — one arriving
|
||||
through the I/O APIC from a real device line — must be **masked before it is
|
||||
acknowledged**, because a level-triggered line is still asserted at EOI time and would
|
||||
redeliver instantly, forever. Only the handler knows which discipline its source
|
||||
needs, so only the handler can sequence it. See [drivers.md](drivers.md), where the
|
||||
device is quieted by a driver in ring 3, long after the ISR has returned.
|
||||
|
||||
A device handler is a plain `fn () void` — a timer or keyboard handler doesn't need
|
||||
the interrupted registers. (Note: the stubs don't save the SSE/vector registers, so
|
||||
a handler must not use them; ours don't.)
|
||||
@@ -131,14 +143,23 @@ If the APIC weren't enabled, or `sti` were missing, or EOI were forgotten, the
|
||||
count would stay put and the test would fail. That it advances — while the CPU was
|
||||
spinning in unrelated code — is the whole mechanism working end to end.
|
||||
|
||||
## Since (done elsewhere)
|
||||
|
||||
- **Preemption**: the timer handler is where the scheduler decides to switch — the
|
||||
reason a *returning* interrupt matters. See [scheduling.md](scheduling.md).
|
||||
- **`sleep()` / timeouts** built on the calibrated clock.
|
||||
- **The I/O APIC, routed**: external device lines now reach a vector, and the
|
||||
interrupt is delivered onward to a *user-space* driver as an IPC message. See
|
||||
[drivers.md](drivers.md).
|
||||
- **Uncacheable MMIO**: device grants are mapped `PCD|PWT` (strong-uncacheable) for
|
||||
user drivers — see [paging.md](paging.md).
|
||||
|
||||
## What's next (not done here)
|
||||
|
||||
- **The keyboard**: bring up the IO-APIC, route its IRQ to a vector, and read
|
||||
scancodes from the PS/2 controller — the first *input* device.
|
||||
- **`sleep()` / timeouts** built on the calibrated clock (the monotonic
|
||||
`uptimeMs()` is in place).
|
||||
- **Uncacheable MMIO**: the LAPIC page is currently mapped writeback-cacheable like
|
||||
the rest of the identity map. QEMU tolerates it, but real hardware wants MMIO
|
||||
marked uncacheable (via the page's cache bits or an MTRR).
|
||||
- **Preemption**: once there are tasks, the timer handler is where the scheduler
|
||||
decides to switch — the reason a *returning* interrupt matters.
|
||||
- **The keyboard**: the PS/2 controller is port-mapped (`0x60`/`0x64`), and ring 3
|
||||
has no port I/O yet, so the first *input* device is blocked on either an I/O
|
||||
permission bitmap or `io_in`/`io_out` syscalls ([drivers.md](drivers.md)).
|
||||
- **MSI/MSI-X**: per-device vectors, edge-triggered and unshared, which retire the
|
||||
I/O APIC's mask/ack cycle and its 24-GSI ceiling.
|
||||
- **The LAPIC's own page** is still mapped writeback-cacheable like the rest of the
|
||||
identity map. QEMU tolerates it; real hardware wants it uncacheable.
|
||||
|
||||
@@ -0,0 +1,305 @@
|
||||
# The driver model: buses, classes, and host controllers
|
||||
|
||||
[drivers.md](drivers.md) shows how to write *a* driver — claim a device, map its
|
||||
registers, sleep on its interrupt. That's enough for a leaf device like the HPET. It is
|
||||
not enough for a disk, a keyboard, or a network card, because those hang off a
|
||||
*controller*, on a *bus*, speaking a *protocol*, and no single process should have to
|
||||
know all three.
|
||||
|
||||
Real driver stacks factor into three shapes. This document is about what each one is,
|
||||
what the kernel must give it, how they share code — and precisely which primitive each
|
||||
is still blocked on.
|
||||
|
||||
## Three shapes
|
||||
|
||||
| Shape | Owns | Reaches hardware by | Talks to |
|
||||
|---|---|---|---|
|
||||
| **Host controller driver** (HCD) | a controller — an xHCI PCI function, an AHCI port block | `mmio_map` + `irq_bind` + DMA | the devices behind it, in its bus's language |
|
||||
| **Bus driver** | a bus — a PCI bridge, a USB hub | `device_register`, to publish what it finds | class drivers, over IPC |
|
||||
| **Class / protocol driver** | *nothing* | *nothing* | its bus driver, over IPC |
|
||||
|
||||
The last row is the surprising one and the whole point. A USB keyboard driver touches
|
||||
no registers, takes no interrupts, and maps no memory. It sends HID protocol messages
|
||||
to whatever published the device, and it works identically whether the controller
|
||||
below is xHCI, EHCI, or a Raspberry Pi's DWC2. That is what buys you drivers that
|
||||
outlive the hardware they were written for.
|
||||
|
||||
In practice **HCD and bus driver are usually the same process**. An xHCI driver is a
|
||||
host controller driver (it owns the PCI function, its BARs, its interrupt, its DMA
|
||||
rings) *and* a bus driver (it enumerates USB devices and publishes them). Splitting
|
||||
them is a fiction; what matters is that both *roles* have kernel support, because a
|
||||
plain bus driver with no controller — a USB hub — is also a real thing.
|
||||
|
||||
## The device table is the spine
|
||||
|
||||
danos already has the right central structure. `src/kernel/device-service.zig` holds a table of
|
||||
`DeviceDesc`, each with a parent, a class, and a set of resources. Firmware discovery
|
||||
seeds it ([discovery.md](discovery.md)); `device_register` grows it.
|
||||
|
||||
Three invariants make it a capability system rather than a directory:
|
||||
|
||||
1. **A claim is exclusive.** `device_claim(id)` succeeds once. Everything downstream —
|
||||
`mmio_map`, `irq_bind`, `device_register` — checks `device_service.ownerOf(id) == me`.
|
||||
2. **A descriptor is a licence to map physical memory.** Whoever claims a device may
|
||||
map its `.memory` resources and bind its `.irq` resources. This is why
|
||||
`device_register` cannot be a free-for-all.
|
||||
3. **Therefore: containment.** Every resource of a registered child must lie inside a
|
||||
resource of the same kind on its parent (`device_service.contains`). A bus driver can only
|
||||
ever *subdivide* what it already holds. Without this, `device_register` would be a
|
||||
syscall named "map any physical page you like."
|
||||
|
||||
Containment is transitive by construction: a grandchild is contained in its child,
|
||||
which is contained in the bus. Nothing can be laundered through a chain.
|
||||
|
||||
Note that firmware topology does **not** obey containment, and isn't asked to — a PCI
|
||||
function's BAR is not inside its host bridge's `bus_range`, because a bus-number range
|
||||
is not an address window. Discovery is trusted; user space is not.
|
||||
|
||||
### What a bus driver looks like
|
||||
|
||||
`sbin/busd.zig` is the smallest honest one. Its "bus" is the HPET's register block and
|
||||
its "devices" are the block's comparators:
|
||||
|
||||
```zig
|
||||
_ = dev.claim(bus.id); // 1. own the bus
|
||||
const base = dev.mmioMap(bus.id, 0).?; // 2. enumerate it — from the hardware
|
||||
const n = ((cap.* >> 8) & 0x1F) + 1; // GENERAL_CAP says how many children
|
||||
|
||||
for (0..n) |i| { // 3. publish each child
|
||||
var child = std.mem.zeroes(dev.DeviceDesc);
|
||||
child.class = @intFromEnum(dev.DeviceClass.timer);
|
||||
child.resource_count = 1;
|
||||
child.resources[0] = .{ .kind = memory,
|
||||
.start = bus_mmio.start + 0x100 + 0x20 * i,
|
||||
.len = 0x20 };
|
||||
_ = dev.register(bus.id, &child).?; // kernel checks containment
|
||||
}
|
||||
```
|
||||
|
||||
Each child is left **unclaimed**, which is the handoff: a comparator driver can now
|
||||
`device_claim` one and `mmio_map` it, and will see only its own 0x20-byte window. A child
|
||||
whose window escapes the bus is refused — `busd` asserts that, and the `bus` test
|
||||
asserts the kernel's table upholds it.
|
||||
|
||||
A USB device has *no* resources at all: `resource_count = 0`, because it's addressed
|
||||
through its controller, not by MMIO. That case is allowed and is the common one.
|
||||
|
||||
## Families: sharing code between drivers
|
||||
|
||||
A "family" is two modules, not one:
|
||||
|
||||
- **A logic module** — the parts of the bus that every driver on it re-derives. Config
|
||||
space walking and BAR decode for PCI. Descriptor parsing, control transfers, and hub
|
||||
protocol for USB.
|
||||
- **A protocol module** — the IPC message types that let a class driver talk to
|
||||
*whatever* published its device. This is the part that makes class drivers portable.
|
||||
|
||||
danos already has one of each: `lib/device.zig` is a logic module,
|
||||
[`lib/vfs-protocol.zig`](lib/vfs-protocol.zig) is a protocol module shared by `sbin/vfs.zig`
|
||||
and its clients. The pattern generalises directly:
|
||||
|
||||
```
|
||||
lib/
|
||||
rt.zig module "rt" — syscalls, heap, ipc, dev, stdio
|
||||
mmio.zig module "mmio" — volatile register access + barriers [M14]
|
||||
bus/
|
||||
pci.zig module "pci" — ECAM, BAR decode, capability walk
|
||||
usb.zig module "usb" — descriptors, control transfers, hubs
|
||||
proto/
|
||||
vfs.zig module "proto.vfs" (today: lib/vfs-protocol.zig)
|
||||
block.zig module "proto.block"
|
||||
hid.zig module "proto.hid"
|
||||
|
||||
sbin/
|
||||
xhcid.zig HCD + bus driver imports rt, pci, usb, mmio
|
||||
usbhid.zig class driver imports rt, usb, proto.hid
|
||||
blockd.zig class driver imports rt, proto.block
|
||||
```
|
||||
|
||||
The only build change needed: [`addUserBinary`](build.zig) currently takes exactly one
|
||||
module (`rt_mod`) and injects it. It should take a slice of modules. That's a
|
||||
five-line change, and it's the *entire* mechanism — Zig modules already give you
|
||||
everything else.
|
||||
|
||||
The discipline that makes this work: **a class driver must not import a bus's logic
|
||||
module.** `usbhid` imports `proto.hid` and `usb` (for descriptor types), never `pci`.
|
||||
If a class driver needs `mmio`, it has become an HCD and should be one.
|
||||
|
||||
## What exists today
|
||||
|
||||
- **M10** — `device_enumerate`, `device_claim`, `mmio_map`. Strong-uncacheable device
|
||||
grants, `device_grant` teardown.
|
||||
- **M11** — `irq_bind` / `irq_ack`. IRQ delivered as an IPC notification; mask before
|
||||
EOI; `irq_ack` is the unmask.
|
||||
- **M12** — `parent` in `DeviceDesc`, `device_register` with resource containment.
|
||||
|
||||
So: **bus drivers work now.** HCDs and class drivers do not. Here is exactly why, and
|
||||
exactly what would fix it.
|
||||
|
||||
---
|
||||
|
||||
# Proposed ABI
|
||||
|
||||
## M13 — capability passing, for class drivers
|
||||
|
||||
**The blocker.** A class driver has to reach *its* device. Today the only way to find
|
||||
an endpoint is the name registry: `ipc_register(service_id, h)` / `ipc_lookup(id)`,
|
||||
where `ServiceId` is a global integer namespace with `max_services = 8`. You cannot
|
||||
mint one endpoint per USB device that way, and there is no way for a bus driver to
|
||||
*hand* a class driver an endpoint. M7 deferred this deliberately.
|
||||
|
||||
**The fix.** Let a message carry one handle. Sender names a handle in its own table;
|
||||
the kernel installs the endpoint into the receiver's table (bumping `refcount`) and
|
||||
tells the receiver the index it landed at.
|
||||
|
||||
```
|
||||
ipc_call(h, msg, message_len, reply, reply_cap, send_cap) -> reply_len
|
||||
ipc_reply_wait(h, reply, reply_len, recv, recv_cap, send_cap)
|
||||
-> recv_len (rax), badge (rdx), received_cap (r8)
|
||||
```
|
||||
|
||||
`send_cap` is a handle or `no_cap` (`~0`). `received_cap` is the index the transferred
|
||||
endpoint was installed at in the receiver's table, or `no_cap`.
|
||||
|
||||
- Both calls grow from 5 args to 6, which fits: `syscall5` uses `rdi/rsi/rdx/r10/r8`,
|
||||
leaving `r9`. `ipc_reply_wait` already returns two values via `setSyscallResult2`;
|
||||
this needs a third (`setSyscallResult3`).
|
||||
- If the receiver's handle table is full, the call fails `-ENOSPC` and **the message is
|
||||
not delivered** — a half-delivered capability is worse than a failed send.
|
||||
- `closeHandles` already drops references on exit, so the lifetime story is unchanged.
|
||||
|
||||
That single primitive gives you the standard `open` pattern:
|
||||
|
||||
```zig
|
||||
// class driver // bus driver
|
||||
const h = ipc.lookup(.usb).?; const r = ipc.replyWait(ep, ...);
|
||||
const dev_ep = ipc.callCap(h, // ... mint a per-device endpoint,
|
||||
.{ .op = .open, .id = dev_id }); // reply with it as send_cap
|
||||
// now dev_ep is a private channel to that one device
|
||||
```
|
||||
|
||||
## M14 — DMA memory and the memory-ordering contract, for HCDs
|
||||
|
||||
**The blocker.** An HCD is a DMA-engine programmer. It needs a descriptor ring the
|
||||
device can read, which means memory that is (a) physically contiguous, (b) at a
|
||||
physical address the driver knows, (c) of the right cacheability, and (d) pinned.
|
||||
[`sysMmap`](src/kernel/process.zig) gives you *none* of the four: it calls `pmm.alloc()`
|
||||
once per page, maps writeback-cached, and never reveals a physical address.
|
||||
|
||||
**The fix.**
|
||||
|
||||
```
|
||||
dma_alloc(len, flags) -> vaddr (rax), paddr (rdx)
|
||||
dma_free(vaddr, len) -> 0
|
||||
|
||||
flags: dma_coherent (1) uncacheable; the default and the only one that's portable
|
||||
dma_wc (2) write-combining — needs PAT programmed; for framebuffers
|
||||
dma_below_4g (4) for devices with 32-bit DMA addressing
|
||||
```
|
||||
|
||||
Guarantees: page-aligned, physically contiguous, zeroed, pinned for the life of the
|
||||
mapping, and the physical address is stable. It needs one thing the kernel lacks —
|
||||
`pmm.allocContiguous(n, max_phys)`; today `pmm.alloc()` hands out one frame at a time
|
||||
with no adjacency guarantee.
|
||||
|
||||
**The memory-ordering contract.** danos has, at the time of writing, **zero memory
|
||||
barriers anywhere in the tree.** That is currently correct-by-accident and won't
|
||||
survive the first DMA driver, or the first ARM boot.
|
||||
|
||||
`volatile` is not a barrier. In Zig it means: don't elide this access, and don't
|
||||
reorder it against *other volatile* accesses. It says nothing about your *ordinary*
|
||||
stores — the descriptor you just filled in normal WB memory — which LLVM may freely
|
||||
sink past a volatile MMIO write. The canonical bug:
|
||||
|
||||
```zig
|
||||
ring[i] = descriptor; // ordinary store to WB RAM
|
||||
doorbell.* = i; // volatile store to UC MMIO
|
||||
// nothing stops the compiler reordering these; the device reads a stale descriptor
|
||||
```
|
||||
|
||||
So the rules, which belong in `lib/mmio.zig` and behind `arch`:
|
||||
|
||||
| Situation | Required |
|
||||
|---|---|
|
||||
| MMIO register read/write | `mmio.read` / `mmio.write` (volatile) |
|
||||
| Fill DMA descriptor, then ring doorbell | `wmb()` between them |
|
||||
| Woken by IRQ, then read what the device wrote | `rmb()` before the read |
|
||||
| MMIO write that must complete before the next read | `mb()` |
|
||||
|
||||
And the per-arch lowering — the reason this must be an `arch` primitive and not a
|
||||
sprinkling of `asm volatile`:
|
||||
|
||||
| | x86_64 | aarch64 |
|
||||
|---|---|---|
|
||||
| `mb()` | `mfence` | `dsb sy` |
|
||||
| `rmb()` | `lfence` | `dsb ld` |
|
||||
| `wmb()` | `sfence` | `dsb st` |
|
||||
| DMA cache coherency | coherent; nothing to do | **not guaranteed**; needs non-cacheable buffers or cache maintenance |
|
||||
|
||||
x86 is forgiving here — TSO plus strong-uncacheable MMIO means you usually get away
|
||||
with a compiler barrier alone. ARM is not, and [vision.md](vision.md) makes ARM the win
|
||||
condition. Build the abstraction while there is one caller to fix.
|
||||
|
||||
(Zig note: `@fence` was **removed in 0.16**. Use `@atomicRmw(..., .seq_cst)` for a full
|
||||
barrier, or per-arch inline asm — which is what `lib/mmio.zig` should hide.)
|
||||
|
||||
## M15 — interrupts for PCI devices
|
||||
|
||||
**The blocker, and it's a hard one.** No PCI device can take an interrupt today.
|
||||
[`addBars`](src/device/acpi.zig) records `.memory` and `.io_port` BARs and never an
|
||||
`.irq`; there is no `_PRT` parsing anywhere in the tree. `hpetd` only works because the
|
||||
HPET advertises its own routing options in its own registers — a privilege no ordinary
|
||||
device has.
|
||||
|
||||
**The fix, in two halves.**
|
||||
|
||||
*Legacy INTx*: parse `_PRT` from the DSDT to map (device, INTA–D) → GSI, and record it
|
||||
as an `.irq` resource. Then `irq_bind` works unchanged. But INTx lines are **shared**,
|
||||
and `irq.bound[gsi]` holds one endpoint. Sharing needs a list, and every driver on the
|
||||
line must be polled on each interrupt — the reason everyone left INTx behind.
|
||||
|
||||
*MSI/MSI-X*, which is the real answer: per-device vectors, edge-triggered, unshared, no
|
||||
mask/ack cycle, no 24-GSI ceiling. The kernel allocates a vector and hands the driver
|
||||
the (address, data) pair to program into its own MSI capability:
|
||||
|
||||
```
|
||||
msi_bind(dev_id, endpoint, out) -> 0 // out: extern struct { addr: u64, data: u32 }
|
||||
```
|
||||
|
||||
The driver writes those into config space itself — which means it needs config space,
|
||||
which means **discovery should give each `pci_device` a `.memory` resource for its
|
||||
4 KiB ECAM slot**. That's a small change to `parseMcfg` and it unblocks the whole
|
||||
capability walk (MSI, MSI-X, PCIe extended caps) without any new syscall.
|
||||
|
||||
Note QEMU's HPET reports `Tn_FSB_INT_DEL_CAP = 0` — no MSI — so `hpetd` can never
|
||||
exercise this path. The first MSI driver will be the first PCI driver.
|
||||
|
||||
## M16 — the IOMMU, and the honest caveat
|
||||
|
||||
Everything above is capability-gated at the *CPU*. None of it is gated at the *device*.
|
||||
A driver that can program a bus-mastering engine can make that device write to any
|
||||
physical address, because page tables sit between the CPU and RAM, not between a device
|
||||
and RAM. Until VT-d/DMAR (or SMMU on ARM) is programmed from the DMAR table, **`device_claim`
|
||||
on any DMA-capable device is equivalent to granting ring 0.**
|
||||
|
||||
This does not make the model useless — it's the same position Linux is in with the
|
||||
IOMMU off, and every other guarantee (crash isolation, restart, no shared address
|
||||
space) still holds. But "user-space drivers are memory-safe" is not true yet, and the
|
||||
gap should be named rather than implied.
|
||||
|
||||
## Ordering
|
||||
|
||||
`M13` (capability passing) is independent of `M14`/`M15` and is the cheapest. It
|
||||
unlocks class drivers, which are the shape with no hardware requirements at all — you
|
||||
could write a real one against `busd`'s comparators tomorrow.
|
||||
|
||||
`M14` and `M15` together unlock the first HCD. `M14`'s barrier layer is worth landing
|
||||
on its own regardless: it's small, obviously correct, and stops every future driver
|
||||
from hand-rolling `*volatile` and getting ARM wrong.
|
||||
|
||||
## See also
|
||||
|
||||
- [drivers.md](drivers.md) — how to write one, concretely.
|
||||
- [discovery.md](discovery.md) / [acpi.md](acpi.md) — where the device table comes from.
|
||||
- [ipc.md](ipc.md) — endpoints, badges, and the notification path an IRQ arrives on.
|
||||
- [resilience.md](resilience.md) — restart, the reason any of this is worth the trouble.
|
||||
+319
@@ -0,0 +1,319 @@
|
||||
# Writing a driver
|
||||
|
||||
In a monolithic kernel a driver is a function call away from everything: it runs in
|
||||
ring 0, dereferences any physical address, and its interrupt handler *is* the ISR. In
|
||||
danos a driver is **an ordinary ring-3 process**. It has its own address space, it
|
||||
can crash without taking the kernel with it, and — the point of this document — it
|
||||
can be restarted ([resilience](resilience.md)).
|
||||
|
||||
That leaves three questions the kernel has to answer, because a process can't answer
|
||||
them for itself:
|
||||
|
||||
1. **What hardware exists?** → `device_enumerate`, over the device table discovery built
|
||||
([discovery](discovery.md), [acpi](acpi.md)).
|
||||
2. **How do I touch its registers?** → `device_claim` + `mmio_map`: the kernel maps the
|
||||
device's physical MMIO window into your address space, and from then on it's plain
|
||||
memory. No syscall per register access.
|
||||
3. **How do I find out it wants something?** → `irq_bind`: the interrupt is delivered
|
||||
to you as an IPC notification. You block; the hardware wakes you.
|
||||
|
||||
A driver is, in one sentence, *a process that sleeps until its device has something to
|
||||
say.*
|
||||
|
||||
## The capability: claim before touch
|
||||
|
||||
The five driver syscalls (`src/root.zig`, dispatched in `src/kernel/process.zig`):
|
||||
|
||||
| # | Call | Meaning |
|
||||
|---|------|---------|
|
||||
| 11 | `device_enumerate(buf, max) -> total` | Snapshot the device table |
|
||||
| 12 | `device_claim(id) -> ok` | Take **exclusive** ownership |
|
||||
| 13 | `mmio_map(id, res_idx) -> vaddr` | Map a claimed device's register window |
|
||||
| 14 | `irq_bind(id, res_idx, endpoint)` | Deliver that device's IRQ as a notification |
|
||||
| 15 | `irq_ack(id, res_idx)` | Re-arm the IRQ after servicing the device |
|
||||
| 16 | `device_register(parent_id, desc) -> id` | Publish a child of a device you claimed |
|
||||
|
||||
Notice that **nothing takes a physical address or an interrupt number.** Every call
|
||||
names a device by id and a resource by index. That indirection is the entire security
|
||||
model. If `mmio_map` took a physical address, any process could map the kernel's
|
||||
memory; if `irq_bind` took a GSI, any process could bind the keyboard's line and
|
||||
silently intercept it. Instead the kernel checks two things (`process.ownedGsi`, and
|
||||
the same check at the top of `sysMmioMap`):
|
||||
|
||||
- `device_service.ownerOf(dev_id) == me` — you claimed it, and claims are exclusive
|
||||
- the resource at `res_idx` is of the right *kind* — `memory` for `mmio_map`, `irq`
|
||||
for `irq_bind`
|
||||
|
||||
The claim is the capability. Everything else follows from it.
|
||||
|
||||
## Registers: `mmio_map`
|
||||
|
||||
`mmio_map` walks the caller's page tables and installs the device's physical frames
|
||||
with `present | user | writable | nx | pcd | pwt`
|
||||
(`arch/x86_64/paging.zig:mapUserDeviceInto`). Two of those bits are load-bearing:
|
||||
|
||||
- **`pcd | pwt`** — strong-uncacheable. A device register is not memory; a cached read
|
||||
would return a stale value and a write might never leave the CPU.
|
||||
- **`device_grant`** (bit 9, one of the PTE's available bits) — marks the leaf as MMIO
|
||||
rather than RAM, so `freeSubtree` skips `pmm.free` on it when the address space is
|
||||
destroyed. Without this, killing a driver would hand the HPET's registers back to
|
||||
the frame allocator as if they were free RAM. The `iopass` test guards it.
|
||||
|
||||
Grants land in their own arena, `0x0000_7100_0000_0000` (PML4[226]), so device pages
|
||||
never widen an existing mapping.
|
||||
|
||||
Then you just… use it:
|
||||
|
||||
```zig
|
||||
const base = dev.mmioMap(dev_id, mmio_res) orelse return;
|
||||
const counter: *volatile u64 = @ptrFromInt(base + 0xF0);
|
||||
const now = counter.*; // a load, straight to the hardware. no kernel involved.
|
||||
```
|
||||
|
||||
## Interrupts: the cycle, and why it has that shape
|
||||
|
||||
An interrupt handler in a microkernel has a problem. The code that knows how to quiet
|
||||
the device is in ring 3, in another address space, and it will not run for
|
||||
microseconds or milliseconds — after a context switch, when the scheduler gets to it.
|
||||
But the CPU wants an EOI *now*, and a **level-triggered** line stays asserted until
|
||||
the device is quieted. EOI a still-asserted line and the I/O APIC redelivers
|
||||
immediately. Forever. The driver never gets to run at all.
|
||||
|
||||
The way out is to mask the line before acknowledging it:
|
||||
|
||||
```
|
||||
kernel ISR irqMask(gsi) // line still asserted; stop it reaching a CPU
|
||||
irqEoi() // now safe to tell the LAPIC we're done
|
||||
notifyFromIsr() // wake the driver — it runs much later
|
||||
|
||||
driver replyWait() -> badge with the notify bit set
|
||||
<clear the device's status register> // NOW the line deasserts
|
||||
irq_ack(dev, res) // kernel unmasks: quiet, so it can't refire
|
||||
|
||||
```
|
||||
|
||||
`irq_ack` is not bookkeeping you could skip. **It is the unmask.** Forget it and the
|
||||
interrupt fires exactly once, ever; call it before the device is quiet and you get an
|
||||
interrupt storm. That single fact explains why `irq_bind` and `irq_ack` are two
|
||||
syscalls and not one.
|
||||
|
||||
This is also why `interruptDispatch` (`arch/x86_64/idt.zig`) no longer issues the EOI
|
||||
itself. It used to, before running the handler — correct for the LAPIC timer, and
|
||||
impossible for a routed device line. Each handler now owns its EOI, because only the
|
||||
handler knows which discipline its source needs.
|
||||
|
||||
### The driver side is an event loop, not a callback
|
||||
|
||||
`IPC_ReplyWait` returns *either* a client request *or* a notification, told apart by
|
||||
the top bit of the badge (`ipc_sync.notify_badge_bit`). So a driver is one
|
||||
single-threaded loop over both of its event sources:
|
||||
|
||||
```zig
|
||||
while (true) {
|
||||
const r = ipc.replyWait(endpoint, reply, &recv);
|
||||
if (r.isNotification()) { // r.source() is the GSI
|
||||
service_device(); // clear the status register
|
||||
_ = dev.irqAck(id, irq_res); // re-arm
|
||||
} else {
|
||||
handle_client_request(recv[0..r.len]);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
No reentrancy, no "what am I allowed to call from an interrupt handler", no shared
|
||||
state between ISR and task context. The interrupt is just a message.
|
||||
|
||||
Two properties worth knowing:
|
||||
|
||||
- **An interrupt taken while you're elsewhere is not lost.** If the driver is off in
|
||||
an `ipc_call` to another server when the IRQ fires, `wakeLocked` finds nobody
|
||||
waiting, but the badge is already on the endpoint's notify ring. The next
|
||||
`replyWait` pops it (`ipc_sync.replyWait` checks `popNotify` before the sender FIFO).
|
||||
- **Notifications coalesce, they don't count.** The ring is 8 deep and drops on
|
||||
overflow. That's correct: an IRQ notification is a *level* ("the device wants
|
||||
attention"), not a tally. Re-read the device's status register; never assume one
|
||||
notification means exactly one event. Because the ISR masks the line until you
|
||||
`irq_ack`, at most one badge per GSI can be outstanding — so the ring can only
|
||||
overflow if you bind more than eight GSIs to a single endpoint. Don't.
|
||||
|
||||
## A whole driver
|
||||
|
||||
`sbin/hpetd.zig` is ~150 lines and does all of it. The shape:
|
||||
|
||||
```zig
|
||||
const hpet = findHpet(buf) orelse return; // device_enumerate, look for
|
||||
// class=timer with memory + irq
|
||||
_ = dev.claim(hpet.dev_id); // the capability
|
||||
const base = dev.mmioMap(hpet.dev_id, hpet.mmio).?;
|
||||
const endpoint = ipc.createEndpoint().?;
|
||||
|
||||
// program the hardware over the mapping we were just handed
|
||||
reg(base, 0x100).* = level | int_enb | (hpet.gsi << 9); // timer 0 config
|
||||
reg(base, 0x108).* = reg(base, 0xF0).* + period; // comparator
|
||||
reg(base, 0x010).* |= 1; // ENABLE
|
||||
|
||||
_ = dev.irqBind(hpet.dev_id, hpet.irq, endpoint);
|
||||
|
||||
while (...) {
|
||||
const r = ipc.replyWait(endpoint, &.{}, &recv); // blocked. not polling.
|
||||
if (r.badge & notify_bit == 0) continue;
|
||||
reg(base, 0x020).* = 1; // clear status -> deassert
|
||||
reg(base, 0x108).* = reg(base, 0xF0).* + period; // re-arm
|
||||
_ = dev.irqAck(hpet.dev_id, hpet.irq); // unmask
|
||||
}
|
||||
```
|
||||
|
||||
The HPET is a good first driver for a reason that isn't obvious. Its *counter* is a
|
||||
clocksource — the only way to use it is to read it, so it proved `mmio_map` without
|
||||
needing interrupts at all. Its *comparators* are a clockevent, and can be configured
|
||||
**level-triggered** (`Tn_INT_TYPE_CNF`), which asserts a bit in `GENERAL_INT_STATUS`
|
||||
that the driver must write-1-to-clear. That's a genuine deassert step, so the full
|
||||
mask/ack cycle above is exercised for real rather than being decoration on an
|
||||
edge-triggered line that would have been fine without it.
|
||||
|
||||
One wrinkle it also demonstrates: the ACPI HPET table carries **no interrupt number**.
|
||||
Which I/O APIC inputs a comparator may drive is a bitmask in `Tn_INT_ROUTE_CAP`, in
|
||||
the device's own registers. So discovery (`acpi.parseHpet`) maps the block, reads the
|
||||
mask, and records one concrete GSI as an `irq` resource. The driver then programs
|
||||
`Tn_INT_ROUTE_CNF` to raise exactly that line — and the kernel will only bind the one
|
||||
it recorded. Hardware that describes itself at runtime still has to fit through a
|
||||
static capability.
|
||||
|
||||
## Publishing children: `device_register`
|
||||
|
||||
A device that *contains other devices* — a PCI bridge, a USB hub, or the HPET's block
|
||||
of comparators — needs a driver that enumerates it and tells the kernel what it found.
|
||||
That's `device_register`, and it makes the device table a tree rather than a list
|
||||
(`DeviceDesc.parent`).
|
||||
|
||||
```zig
|
||||
var child = std.mem.zeroes(dev.DeviceDesc);
|
||||
child.class = @intFromEnum(dev.DeviceClass.timer);
|
||||
child.resource_count = 1;
|
||||
child.resources[0] = .{ .kind = memory, .start = bus_base + 0x100, .len = 0x20 };
|
||||
const child_id = dev.register(bus_id, &child).?;
|
||||
```
|
||||
|
||||
The child is left **unclaimed**, which is the whole point: another process claims it and
|
||||
`mmio_map`s it, and sees only that 0x20-byte window.
|
||||
|
||||
The rule the kernel enforces is **containment**: every resource of a child must lie
|
||||
inside a resource of the same kind on its parent. Ranges must nest; an IRQ must match
|
||||
exactly. This isn't bureaucracy — a `DeviceDesc` is a licence to map physical memory, so
|
||||
without containment `device_register` would be a syscall for mapping any page you like. A
|
||||
bus driver may only ever subdivide what it already owns.
|
||||
|
||||
A device with **no resources** is legal and common. A USB device is reached through its
|
||||
controller, not by MMIO, so it gets `resource_count = 0`.
|
||||
|
||||
See [`sbin/busd.zig`](../sbin/busd.zig) for a complete one, and
|
||||
[driver-model.md](driver-model.md) for how bus drivers, class drivers and host
|
||||
controller drivers fit together.
|
||||
|
||||
## What the kernel does not do for you
|
||||
|
||||
- **It does not quiet your device.** That's the whole reason `irq_ack` exists.
|
||||
- **It does not know your registers.** `mmio_map` hands you a base address; every
|
||||
offset in this document came from the HPET spec, not from danos.
|
||||
- **It does not serialise your driver.** Two clients calling one driver endpoint are
|
||||
serialised by `replyWait`, but nothing stops your driver from being preempted.
|
||||
|
||||
## Limits, today
|
||||
|
||||
Worth knowing before you write the second driver:
|
||||
|
||||
- **Ring 3 has no port I/O.** The TSS I/O permission bitmap is absent
|
||||
(`tss.zig`: `iomap_base = @sizeOf(Tss)`), and IOPL is never raised, so `in`/`out`
|
||||
from a driver is a #GP. That rules out a user-space 16550 UART (`0x3F8`), PS/2
|
||||
(`0x60`/`0x64`), and legacy PCI config (`0xCF8`/`0xCFC`). Everything must be MMIO.
|
||||
`io_port` resources are recorded by discovery and then ignored.
|
||||
- **Page granularity.** `mmio_map` rounds to 4 KiB. Two devices sharing a page means
|
||||
granting one grants the other. A `device_register`ed child's *resource* can be narrower
|
||||
than a page, but its *mapping* can't.
|
||||
- **No DMA memory.** `mmap` gives you writeback-cached, non-contiguous pages and never
|
||||
tells you their physical address, so you cannot build a descriptor ring. Any driver
|
||||
for a bus-mastering device is blocked on this.
|
||||
- **No memory barriers.** There are none in the tree, and `volatile` is not one — it
|
||||
won't stop the compiler sinking an ordinary store (your DMA descriptor) past a
|
||||
volatile MMIO store (your doorbell). On x86 you mostly get away with it; on ARM you
|
||||
will not. See [driver-model.md](driver-model.md#m14).
|
||||
- **DMA is not contained.** A driver that can program a bus-mastering device can make
|
||||
that device write to *any* physical address — page tables don't sit between a device
|
||||
and RAM; an IOMMU does. Until VT-d/DMAR is programmed, `device_claim` on a DMA-capable
|
||||
device is effectively equivalent to granting ring 0. This is the largest gap between
|
||||
the design's promise and what it delivers.
|
||||
- **One endpoint per GSI**, so shared legacy PCI INTx lines can't be split between two
|
||||
drivers. MSI/MSI-X — one vector per device, edge-triggered, unshared — is the real
|
||||
answer, and QEMU's HPET doesn't offer it (`Tn_FSB_INT_DEL_CAP = 0`).
|
||||
- **Polarity is hardcoded** active-high in `irq.bind`. A device whose MADT override
|
||||
says active-low needs that threaded through from discovery.
|
||||
- **14 device vectors** (33–46) and **24 GSIs**, bounded by the stubs `isr.s` emits and
|
||||
by a single I/O APIC.
|
||||
- **Don't bind more than 8 GSIs to one endpoint.** The notify ring is 8 deep and drops
|
||||
on overflow. With one GSI per endpoint that's unreachable — the line is masked from
|
||||
the ISR until `irq_ack`, so at most one badge is ever outstanding. Bind nine devices
|
||||
to one endpoint, though, and a dropped badge leaves that line masked with nobody
|
||||
left to ack it.
|
||||
- **A faulting driver still kills the machine.** There is no per-process kill path: a
|
||||
ring-3 page fault halts the kernel, so `releaseIrqs` runs only on a voluntary
|
||||
`exit`. Fault isolation is the whole premise ([vision](vision.md)) and it is
|
||||
[not built yet](resilience.md).
|
||||
- **A dead driver's device is not reclaimed.** `releaseIrqs` unbinds and masks the
|
||||
line on exit, but the claim is never released — restart is
|
||||
[not built](resilience.md).
|
||||
- **On real hardware, the mask/EOI cycle may need a remote-IRR flush.** Masking a
|
||||
level-triggered redirection entry with remote-IRR set doesn't clear it on some
|
||||
chipsets, and the line never fires again. QEMU clears it on EOI regardless, so the
|
||||
tests can't see this. Linux flushes remote-IRR by toggling the entry to edge and
|
||||
back. See the note at the top of `src/kernel/irq.zig`.
|
||||
|
||||
## Verifying it
|
||||
|
||||
The `hpet` test spawns `hpetd` from the initrd and watches the serial log. The driver
|
||||
prints `hpetd: ok` only after being woken five times, and its loop's only exit is
|
||||
through `replyWait` returning a notification — it cannot reach that line by polling.
|
||||
|
||||
The last check doesn't trust the driver's self-report at all: the kernel reads the I/O
|
||||
APIC redirection entry back and asserts the line really is routed to a device vector,
|
||||
really is level-triggered, and really was left unmasked by the driver's final
|
||||
`irq_ack`.
|
||||
|
||||
```
|
||||
$ python3 test/qemu_test.py hpet irqfree iopass
|
||||
hpet ... PASS (matched 'DANOS-TEST-RESULT: PASS')
|
||||
irqfree ... PASS (matched 'DANOS-TEST-RESULT: PASS')
|
||||
iopass ... PASS (matched 'DANOS-TEST-RESULT: PASS')
|
||||
```
|
||||
|
||||
Two companions cover what `hpetd` can't, because it never exits:
|
||||
|
||||
- **`irqfree`** — the teardown path. Binds two owners to one shared endpoint, releases
|
||||
one, and reads the I/O APIC back: the departing owner's line is masked, the sibling's
|
||||
is not. That second half is why bindings are keyed on the owning *task* and not on
|
||||
the endpoint pointer — endpoints are shared, so releasing "everything pointing at
|
||||
this endpoint" would silently mask a live driver's device.
|
||||
- **`iopass`** — the `device_grant` teardown rule, so destroying a driver's address
|
||||
space never returns MMIO frames to the RAM pool.
|
||||
|
||||
## What's next (not done here)
|
||||
|
||||
The big ones — capability passing (class drivers), DMA + barriers and MSI (host
|
||||
controller drivers), and the IOMMU — have proposed signatures in
|
||||
[driver-model.md](driver-model.md). Smaller items:
|
||||
|
||||
- **Port I/O grants**, so a PS/2 or 16550 driver is possible: either a per-device TSS
|
||||
I/O permission bitmap swapped on context switch, or `io_in`/`io_out` syscalls gated
|
||||
by the same claim. The legacy devices that need it are all low-rate, so the syscall
|
||||
is likely fast enough.
|
||||
- **Releasing a claim.** There is no `dev_release`, and `device_service` never drops a claim on
|
||||
exit — only IRQ bindings are released. A dead driver's device stays owned forever,
|
||||
which blocks restart.
|
||||
- **Unregistering children.** `device_register` only appends. A USB device that is
|
||||
unplugged cannot be removed, and a bus driver in a loop can exhaust the 64-entry
|
||||
table.
|
||||
- **Restart.** A driver that dies should release its claim, have its device quiesced,
|
||||
and be respawned by a supervisor. Some pieces (`releaseIrqs`, `device_grant`
|
||||
teardown, the claim table) exist; the policy doesn't.
|
||||
- **Interrupt priority / threaded IRQ latency.** `notifyFromIsr` enqueues the woken
|
||||
driver but doesn't preempt (`wakeLocked` deliberately leaves that to the caller), so
|
||||
a woken driver waits for the next scheduling point.
|
||||
+55
-12
@@ -6,12 +6,20 @@ just call each other — a request becomes a **message**. In a microkernel, what
|
||||
was a function call across a monolithic kernel is IPC, so it's a first-class
|
||||
concern, not an afterthought.
|
||||
|
||||
This first form is a **bounded blocking channel** (`src/kernel/ipc.zig`): a fixed-size
|
||||
ring buffer of messages with a producer/consumer rendezvous, built on the
|
||||
scheduler's [wait queues](scheduling.md).
|
||||
There are two layers, built a milestone apart:
|
||||
|
||||
- **`src/kernel/ipc.zig`** — a bounded blocking channel between *kernel threads*,
|
||||
described below. The primitive, and where the blocking discipline was worked out.
|
||||
- **`src/kernel/ipc-synchronous.zig`** — synchronous call/reply between *processes*, across
|
||||
address spaces. What user-space servers and drivers actually talk over. It's the
|
||||
second half of this document.
|
||||
|
||||
## The channel
|
||||
|
||||
The first form is a **bounded blocking channel** (`src/kernel/ipc.zig`): a fixed-size
|
||||
ring buffer of messages with a producer/consumer rendezvous, built on the
|
||||
scheduler's [wait queues](scheduling.md).
|
||||
|
||||
`Channel(T, capacity)` is generic over the message type and buffer size. It holds a
|
||||
ring buffer, a count, and two wait queues:
|
||||
|
||||
@@ -42,16 +50,51 @@ full and empty over and over, so both the blocking-send and blocking-recv paths
|
||||
exercised heavily. The messages arrive intact and in order (their sum is the
|
||||
expected `5050`), and neither task busy-waits — they block and wake each other.
|
||||
|
||||
## Endpoints: call/reply across address spaces
|
||||
|
||||
A channel connects two kernel threads sharing one address space. Real servers are
|
||||
*processes*, so the payload has to cross an address-space boundary. That's
|
||||
`src/kernel/ipc-synchronous.zig`, and its shape is L4's: a synchronous **rendezvous** at an
|
||||
`Endpoint`, with the message copied directly from the sender's pages to the receiver's
|
||||
(`copyAcross` walks both sets of page tables through the physmap — no CR3 switch, no
|
||||
bounce buffer).
|
||||
|
||||
Two syscalls carry it:
|
||||
|
||||
- **`ipc_call(h, msg, reply)`** — copy `msg` to the server, block until it replies.
|
||||
- **`ipc_reply_wait(h, reply, recv)`** — reply to the client you're still holding (if
|
||||
any), then block for the next request. One syscall, because a server's steady state
|
||||
is *always* "finish the last one, wait for the next".
|
||||
|
||||
An endpoint is reached by **handle** — a small integer index into the process's handle
|
||||
table (`Task.handles`), exactly like a file descriptor, and just as unforgeable. The
|
||||
bootstrap problem (how do you get the first handle?) is solved by a tiny name registry:
|
||||
a server calls `ipc_register(service_id, h)` under a well-known small integer, and a
|
||||
client calls `ipc_lookup(service_id)`.
|
||||
|
||||
The server never learns the client's identity beyond a **badge**, delivered alongside
|
||||
the message: the caller's task id.
|
||||
|
||||
### Interrupts are messages too
|
||||
|
||||
`notifyFromIsr` posts an *asynchronous* notification to an endpoint — no payload, no
|
||||
reply owed — and wakes whoever is blocked in `reply_wait`. Its badge has the top bit
|
||||
set (`notify_badge_bit`), which is how a driver's single event loop distinguishes "a
|
||||
client wants something" from "the hardware wants something". Notifications sit in a
|
||||
small coalescing ring on the endpoint, so an interrupt taken while the driver was busy
|
||||
elsewhere is not lost.
|
||||
|
||||
This is what makes a user-space driver possible at all, and it's the subject of
|
||||
[drivers.md](drivers.md).
|
||||
|
||||
## What's next (not done here)
|
||||
|
||||
- **Across address spaces.** Today both endpoints are kernel threads sharing the
|
||||
kernel's memory, so the message is copied within one address space. When user
|
||||
mode arrives, the same channel carries messages between *isolated* processes,
|
||||
copying the payload across the boundary — which is where IPC earns its place as
|
||||
the microkernel's backbone.
|
||||
- **Synchronous call/reply.** A request/response pattern (send-and-wait-for-reply)
|
||||
on top of channels, the shape most driver/service calls take.
|
||||
- **Interrupts as messages.** A hardware interrupt delivered to the driver task
|
||||
that owns the device, as an IPC message.
|
||||
- **Priority inheritance** through IPC, so a high-priority client blocked on a
|
||||
low-priority server doesn't suffer unbounded priority inversion.
|
||||
- **Handle transfer.** A server can't hand a client a handle to a third endpoint, so
|
||||
every capability is either well-known (the registry) or inherited — there's no way
|
||||
to delegate one.
|
||||
- **Asynchronous / buffered send** for the cases where a rendezvous is the wrong
|
||||
shape (logging, notifications between servers).
|
||||
- **A bounded reply.** `MSG_MAX` is 256 bytes and the copy runs under the big kernel
|
||||
lock; a bulk transfer wants shared pages, not a copy.
|
||||
|
||||
Reference in New Issue
Block a user