Two driver-model milestones plus a tree-wide naming pass. Suite 35/35 (QEMU) + host tests green. M11 — IRQ-as-IPC. A ring-3 driver now sleeps until its device interrupts it. New src/kernel/irq.zig: per-GSI endpoint bindings, comptime per-vector trampolines, dispatch = mask GSI -> LAPIC EOI -> notifyLocked, all under one lock region. irq_bind/irq_ack syscalls, gated by the device claim like mmio_map. interruptDispatch no longer EOIs — each handler owns its EOI, because a level line must be masked before it is acknowledged (irq_ack is the unmask). Bindings are keyed on the owning task and released on exit (a shared endpoint's siblings survive). hpetd rewritten interrupt-driven. Tests: hpet (rewritten, reads back the I/O APIC routing) and irqfree. M12 — bus drivers. DeviceDesc gains a parent, making the device table a tree. dev_register (device_register) lets a process publish children below a device it claimed; the kernel enforces resource containment (a child's resources must nest in its parent's), so a descriptor can't fabricate a window over kernel RAM. Descriptor copied in via copyFromUser (physmap walk — an unmapped user pointer fails the call instead of faulting the kernel). Per-parent child cap bounds table exhaustion. sbin/busd.zig is a worked bus driver. Test: bus. Naming — per docs/coding-standards.md: non-acronym abbreviations spelled out (message, descriptor, device_service, scheduler, runtime, physical, interpreter, ...); acronyms kept (IPC, MMIO, DMA, HCD, ...); files are kebab-case (ipc-synchronous.zig, device-service.zig, vfs-protocol.zig, ...). Exceptions: POSIX/C ABI names and Zig idioms (init/len/ptr) kept. Module collisions resolved by specific naming (config -> parameters, device.zig alias -> device_model). AML op/Op disambiguated: op = opcode, Op = operation; per-opcode parse handlers renamed opX -> parseX. New driver docs: drivers.md, driver-model.md (bus/class/HCD shapes + the proposed M13–M16 ABI), coding-standards.md.
15 KiB
The driver model: buses, classes, and host controllers
drivers.md shows how to write a driver — claim a device, map its registers, sleep on its interrupt. That's enough for a leaf device like the HPET. It is not enough for a disk, a keyboard, or a network card, because those hang off a controller, on a bus, speaking a protocol, and no single process should have to know all three.
Real driver stacks factor into three shapes. This document is about what each one is, what the kernel must give it, how they share code — and precisely which primitive each is still blocked on.
Three shapes
| Shape | Owns | Reaches hardware by | Talks to |
|---|---|---|---|
| Host controller driver (HCD) | a controller — an xHCI PCI function, an AHCI port block | mmio_map + irq_bind + DMA |
the devices behind it, in its bus's language |
| Bus driver | a bus — a PCI bridge, a USB hub | device_register, to publish what it finds |
class drivers, over IPC |
| Class / protocol driver | nothing | nothing | its bus driver, over IPC |
The last row is the surprising one and the whole point. A USB keyboard driver touches no registers, takes no interrupts, and maps no memory. It sends HID protocol messages to whatever published the device, and it works identically whether the controller below is xHCI, EHCI, or a Raspberry Pi's DWC2. That is what buys you drivers that outlive the hardware they were written for.
In practice HCD and bus driver are usually the same process. An xHCI driver is a host controller driver (it owns the PCI function, its BARs, its interrupt, its DMA rings) and a bus driver (it enumerates USB devices and publishes them). Splitting them is a fiction; what matters is that both roles have kernel support, because a plain bus driver with no controller — a USB hub — is also a real thing.
The device table is the spine
danos already has the right central structure. src/kernel/device-service.zig holds a table of
DeviceDesc, each with a parent, a class, and a set of resources. Firmware discovery
seeds it (discovery.md); device_register grows it.
Three invariants make it a capability system rather than a directory:
- A claim is exclusive.
device_claim(id)succeeds once. Everything downstream —mmio_map,irq_bind,device_register— checksdevice_service.ownerOf(id) == me. - A descriptor is a licence to map physical memory. Whoever claims a device may
map its
.memoryresources and bind its.irqresources. This is whydevice_registercannot be a free-for-all. - Therefore: containment. Every resource of a registered child must lie inside a
resource of the same kind on its parent (
device_service.contains). A bus driver can only ever subdivide what it already holds. Without this,device_registerwould be a syscall named "map any physical page you like."
Containment is transitive by construction: a grandchild is contained in its child, which is contained in the bus. Nothing can be laundered through a chain.
Note that firmware topology does not obey containment, and isn't asked to — a PCI
function's BAR is not inside its host bridge's bus_range, because a bus-number range
is not an address window. Discovery is trusted; user space is not.
What a bus driver looks like
sbin/busd.zig is the smallest honest one. Its "bus" is the HPET's register block and
its "devices" are the block's comparators:
_ = dev.claim(bus.id); // 1. own the bus
const base = dev.mmioMap(bus.id, 0).?; // 2. enumerate it — from the hardware
const n = ((cap.* >> 8) & 0x1F) + 1; // GENERAL_CAP says how many children
for (0..n) |i| { // 3. publish each child
var child = std.mem.zeroes(dev.DeviceDesc);
child.class = @intFromEnum(dev.DeviceClass.timer);
child.resource_count = 1;
child.resources[0] = .{ .kind = memory,
.start = bus_mmio.start + 0x100 + 0x20 * i,
.len = 0x20 };
_ = dev.register(bus.id, &child).?; // kernel checks containment
}
Each child is left unclaimed, which is the handoff: a comparator driver can now
device_claim one and mmio_map it, and will see only its own 0x20-byte window. A child
whose window escapes the bus is refused — busd asserts that, and the bus test
asserts the kernel's table upholds it.
A USB device has no resources at all: resource_count = 0, because it's addressed
through its controller, not by MMIO. That case is allowed and is the common one.
Families: sharing code between drivers
A "family" is two modules, not one:
- A logic module — the parts of the bus that every driver on it re-derives. Config space walking and BAR decode for PCI. Descriptor parsing, control transfers, and hub protocol for USB.
- A protocol module — the IPC message types that let a class driver talk to whatever published its device. This is the part that makes class drivers portable.
danos already has one of each: lib/device.zig is a logic module,
lib/vfs-protocol.zig is a protocol module shared by sbin/vfs.zig
and its clients. The pattern generalises directly:
lib/
rt.zig module "rt" — syscalls, heap, ipc, dev, stdio
mmio.zig module "mmio" — volatile register access + barriers [M14]
bus/
pci.zig module "pci" — ECAM, BAR decode, capability walk
usb.zig module "usb" — descriptors, control transfers, hubs
proto/
vfs.zig module "proto.vfs" (today: lib/vfs-protocol.zig)
block.zig module "proto.block"
hid.zig module "proto.hid"
sbin/
xhcid.zig HCD + bus driver imports rt, pci, usb, mmio
usbhid.zig class driver imports rt, usb, proto.hid
blockd.zig class driver imports rt, proto.block
The only build change needed: addUserBinary currently takes exactly one
module (rt_mod) and injects it. It should take a slice of modules. That's a
five-line change, and it's the entire mechanism — Zig modules already give you
everything else.
The discipline that makes this work: a class driver must not import a bus's logic
module. usbhid imports proto.hid and usb (for descriptor types), never pci.
If a class driver needs mmio, it has become an HCD and should be one.
What exists today
- M10 —
device_enumerate,device_claim,mmio_map. Strong-uncacheable device grants,device_grantteardown. - M11 —
irq_bind/irq_ack. IRQ delivered as an IPC notification; mask before EOI;irq_ackis the unmask. - M12 —
parentinDeviceDesc,device_registerwith resource containment.
So: bus drivers work now. HCDs and class drivers do not. Here is exactly why, and exactly what would fix it.
Proposed ABI
M13 — capability passing, for class drivers
The blocker. A class driver has to reach its device. Today the only way to find
an endpoint is the name registry: ipc_register(service_id, h) / ipc_lookup(id),
where ServiceId is a global integer namespace with max_services = 8. You cannot
mint one endpoint per USB device that way, and there is no way for a bus driver to
hand a class driver an endpoint. M7 deferred this deliberately.
The fix. Let a message carry one handle. Sender names a handle in its own table;
the kernel installs the endpoint into the receiver's table (bumping refcount) and
tells the receiver the index it landed at.
ipc_call(h, msg, message_len, reply, reply_cap, send_cap) -> reply_len
ipc_reply_wait(h, reply, reply_len, recv, recv_cap, send_cap)
-> recv_len (rax), badge (rdx), received_cap (r8)
send_cap is a handle or no_cap (~0). received_cap is the index the transferred
endpoint was installed at in the receiver's table, or no_cap.
- Both calls grow from 5 args to 6, which fits:
syscall5usesrdi/rsi/rdx/r10/r8, leavingr9.ipc_reply_waitalready returns two values viasetSyscallResult2; this needs a third (setSyscallResult3). - If the receiver's handle table is full, the call fails
-ENOSPCand the message is not delivered — a half-delivered capability is worse than a failed send. closeHandlesalready drops references on exit, so the lifetime story is unchanged.
That single primitive gives you the standard open pattern:
// class driver // bus driver
const h = ipc.lookup(.usb).?; const r = ipc.replyWait(ep, ...);
const dev_ep = ipc.callCap(h, // ... mint a per-device endpoint,
.{ .op = .open, .id = dev_id }); // reply with it as send_cap
// now dev_ep is a private channel to that one device
M14 — DMA memory and the memory-ordering contract, for HCDs
The blocker. An HCD is a DMA-engine programmer. It needs a descriptor ring the
device can read, which means memory that is (a) physically contiguous, (b) at a
physical address the driver knows, (c) of the right cacheability, and (d) pinned.
sysMmap gives you none of the four: it calls pmm.alloc()
once per page, maps writeback-cached, and never reveals a physical address.
The fix.
dma_alloc(len, flags) -> vaddr (rax), paddr (rdx)
dma_free(vaddr, len) -> 0
flags: dma_coherent (1) uncacheable; the default and the only one that's portable
dma_wc (2) write-combining — needs PAT programmed; for framebuffers
dma_below_4g (4) for devices with 32-bit DMA addressing
Guarantees: page-aligned, physically contiguous, zeroed, pinned for the life of the
mapping, and the physical address is stable. It needs one thing the kernel lacks —
pmm.allocContiguous(n, max_phys); today pmm.alloc() hands out one frame at a time
with no adjacency guarantee.
The memory-ordering contract. danos has, at the time of writing, zero memory barriers anywhere in the tree. That is currently correct-by-accident and won't survive the first DMA driver, or the first ARM boot.
volatile is not a barrier. In Zig it means: don't elide this access, and don't
reorder it against other volatile accesses. It says nothing about your ordinary
stores — the descriptor you just filled in normal WB memory — which LLVM may freely
sink past a volatile MMIO write. The canonical bug:
ring[i] = descriptor; // ordinary store to WB RAM
doorbell.* = i; // volatile store to UC MMIO
// nothing stops the compiler reordering these; the device reads a stale descriptor
So the rules, which belong in lib/mmio.zig and behind arch:
| Situation | Required |
|---|---|
| MMIO register read/write | mmio.read / mmio.write (volatile) |
| Fill DMA descriptor, then ring doorbell | wmb() between them |
| Woken by IRQ, then read what the device wrote | rmb() before the read |
| MMIO write that must complete before the next read | mb() |
And the per-arch lowering — the reason this must be an arch primitive and not a
sprinkling of asm volatile:
| x86_64 | aarch64 | |
|---|---|---|
mb() |
mfence |
dsb sy |
rmb() |
lfence |
dsb ld |
wmb() |
sfence |
dsb st |
| DMA cache coherency | coherent; nothing to do | not guaranteed; needs non-cacheable buffers or cache maintenance |
x86 is forgiving here — TSO plus strong-uncacheable MMIO means you usually get away with a compiler barrier alone. ARM is not, and vision.md makes ARM the win condition. Build the abstraction while there is one caller to fix.
(Zig note: @fence was removed in 0.16. Use @atomicRmw(..., .seq_cst) for a full
barrier, or per-arch inline asm — which is what lib/mmio.zig should hide.)
M15 — interrupts for PCI devices
The blocker, and it's a hard one. No PCI device can take an interrupt today.
addBars records .memory and .io_port BARs and never an
.irq; there is no _PRT parsing anywhere in the tree. hpetd only works because the
HPET advertises its own routing options in its own registers — a privilege no ordinary
device has.
The fix, in two halves.
Legacy INTx: parse _PRT from the DSDT to map (device, INTA–D) → GSI, and record it
as an .irq resource. Then irq_bind works unchanged. But INTx lines are shared,
and irq.bound[gsi] holds one endpoint. Sharing needs a list, and every driver on the
line must be polled on each interrupt — the reason everyone left INTx behind.
MSI/MSI-X, which is the real answer: per-device vectors, edge-triggered, unshared, no mask/ack cycle, no 24-GSI ceiling. The kernel allocates a vector and hands the driver the (address, data) pair to program into its own MSI capability:
msi_bind(dev_id, endpoint, out) -> 0 // out: extern struct { addr: u64, data: u32 }
The driver writes those into config space itself — which means it needs config space,
which means discovery should give each pci_device a .memory resource for its
4 KiB ECAM slot. That's a small change to parseMcfg and it unblocks the whole
capability walk (MSI, MSI-X, PCIe extended caps) without any new syscall.
Note QEMU's HPET reports Tn_FSB_INT_DEL_CAP = 0 — no MSI — so hpetd can never
exercise this path. The first MSI driver will be the first PCI driver.
M16 — the IOMMU, and the honest caveat
Everything above is capability-gated at the CPU. None of it is gated at the device.
A driver that can program a bus-mastering engine can make that device write to any
physical address, because page tables sit between the CPU and RAM, not between a device
and RAM. Until VT-d/DMAR (or SMMU on ARM) is programmed from the DMAR table, device_claim
on any DMA-capable device is equivalent to granting ring 0.
This does not make the model useless — it's the same position Linux is in with the IOMMU off, and every other guarantee (crash isolation, restart, no shared address space) still holds. But "user-space drivers are memory-safe" is not true yet, and the gap should be named rather than implied.
Ordering
M13 (capability passing) is independent of M14/M15 and is the cheapest. It
unlocks class drivers, which are the shape with no hardware requirements at all — you
could write a real one against busd's comparators tomorrow.
M14 and M15 together unlock the first HCD. M14's barrier layer is worth landing
on its own regardless: it's small, obviously correct, and stops every future driver
from hand-rolling *volatile and getting ARM wrong.
See also
- drivers.md — how to write one, concretely.
- discovery.md / acpi.md — where the device table comes from.
- ipc.md — endpoints, badges, and the notification path an IRQ arrives on.
- resilience.md — restart, the reason any of this is worth the trouble.