22 KiB
Writing a driver
In a monolithic kernel a driver is a function call away from everything: it runs in ring 0, dereferences any physical address, and its interrupt handler is the ISR. In danos a driver is an ordinary ring-3 process. It has its own address space, it can crash without taking the kernel with it, and — the point of this document — it can be restarted (resilience).
That leaves three questions the kernel has to answer, because a process can't answer them for itself:
- What hardware exists? →
device_enumerate, over the device table discovery built (discovery, acpi). - How do I touch its registers? →
device_claim+mmio_map: the kernel maps the device's physical MMIO window into your address space, and from then on it's plain memory. No syscall per register access. - How do I find out it wants something? →
irq_bind: the interrupt is delivered to you as an IPC notification. You block; the hardware wakes you.
A driver is, in one sentence, a process that sleeps until its device has something to say.
How a driver gets started: discover, match, spawn
Nothing in the kernel decides that the PCI host bridge needs the pci-bus driver — that
is policy, and policy lives in user space. Boot brings user space up as a three-level
supervision hierarchy, each level owning one job:
kernel ──spawns──► init (PID 1) ──spawns──► device-manager ──spawns──► pci-bus
| | |
spawns only init, the service supervisor: the driver supervisor: enumerates
publishes the starts the system /system/devices, matches each device
initial-ramdisk services (vfs, the to a driver, and system_spawn's it
so user space can device-manager). Its
system_spawn from it list is init policy.
The kernel launches exactly one process — init — and hands it nothing but the raw
ability to start more (system_spawn(name, arguments), which loads a binary bundled
in the initial-ramdisk as a fresh ring-3 process — name becoming its argv[0],
the optional arguments its argv[1..], on a SysV entry stack, see sysv.md). Everything else is a user-space decision:
- init (system/services/init) is the service
supervisor. It spawns the system services danos brings up at boot — today
vfsand thedevice-manager— from a small list. Drivers are deliberately not its job. - device-manager (system/services/device-manager)
is the driver supervisor. It does the three steps a monolithic kernel would do in
its probe path, entirely from ring 3:
- Discover —
device_enumeratesnapshots the device table the kernel built from ACPI/PCI (discovery). - Match — for each device it looks up a driver by
DeviceClass. The match policy is a table (driverFor): today a statictimer → hpetmap; a fuller system reads what each driver binds (a manifest under/system/drivers, or the driver describing its own match). - Spawn —
system_spawn(driver_name, arguments)starts the matched driver (the arguments can carry which device it matched), which then claims its device and runs the event loop below.
- Discover —
So "how is a driver discovered and configured" has two halves: discovery is the
kernel's device table, read by anyone; configuration is two user-space policies —
init's service list and the device-manager's match table. Both are hardcoded in their
respective programs today; the natural next step is to move them into /etc (see the
milestone notes in driver-model.md). system_spawn is currently
ungated — any process may spawn any bundled binary — because there is no spawn
capability yet.
The capability: claim before touch
The driver syscall numbers (system/abi.zig) with the device types they carry
(system/devices/device-abi.zig), dispatched in system/kernel/process.zig:
| # | Call | Meaning |
|---|---|---|
| 11 | device_enumerate(buf, max) -> total |
Snapshot the device table |
| 12 | device_claim(id) -> ok |
Take exclusive ownership |
| 13 | mmio_map(id, res_idx) -> vaddr |
Map a claimed device's register window |
| 14 | irq_bind(id, res_idx, endpoint) |
Deliver that device's IRQ as a notification |
| 15 | irq_ack(id, res_idx) |
Re-arm the IRQ after servicing the device |
| 16 | device_register(parent_id, desc) -> id |
Publish a child of a device you claimed |
Notice that nothing takes a physical address or an interrupt number. Every call
names a device by id and a resource by index. That indirection is the entire security
model. If mmio_map took a physical address, any process could map the kernel's
memory; if irq_bind took a GSI, any process could bind the keyboard's line and
silently intercept it. Instead the kernel checks two things (process.ownedGsi, and
the same check at the top of sysMmioMap):
devices_broker.ownerOf(dev_id) == me— you claimed it, and claims are exclusive- the resource at
res_idxis of the right kind —memoryformmio_map,irqforirq_bind
The claim is the capability. Everything else follows from it.
Registers: mmio_map
mmio_map walks the caller's page tables and installs the device's physical frames
with present | user | writable | nx | pcd | pwt
(arch/x86_64/paging.zig:mapUserDeviceInto). Two of those bits are load-bearing:
pcd | pwt— strong-uncacheable. A device register is not memory; a cached read would return a stale value and a write might never leave the CPU.device_grant(bit 9, one of the PTE's available bits) — marks the leaf as MMIO rather than RAM, sofreeSubtreeskipspmm.freeon it when the address space is destroyed. Without this, killing a driver would hand the HPET's registers back to the frame allocator as if they were free RAM. Theiopasstest guards it.
Grants land in their own arena, 0x0000_7100_0000_0000 (PML4[226]), so device pages
never widen an existing mapping.
Then you just… use it:
const base = dev.mmioMap(dev_id, mmio_res) orelse return;
const counter: *volatile u64 = @ptrFromInt(base + 0xF0);
const now = counter.*; // a load, straight to the hardware. no kernel involved.
Interrupts: the cycle, and why it has that shape
An interrupt handler in a microkernel has a problem. The code that knows how to quiet the device is in ring 3, in another address space, and it will not run for microseconds or milliseconds — after a context switch, when the scheduler gets to it. But the CPU wants an EOI now, and a level-triggered line stays asserted until the device is quieted. EOI a still-asserted line and the I/O APIC redelivers immediately. Forever. The driver never gets to run at all.
The way out is to mask the line before acknowledging it:
kernel ISR irqMask(gsi) // line still asserted; stop it reaching a CPU
irqEoi() // now safe to tell the LAPIC we're done
notifyFromIsr() // wake the driver — it runs much later
driver replyWait() -> badge with the notify bit set
<clear the device's status register> // NOW the line deasserts
irq_ack(dev, res) // kernel unmasks: quiet, so it can't refire
irq_ack is not bookkeeping you could skip. It is the unmask. Forget it and the
interrupt fires exactly once, ever; call it before the device is quiet and you get an
interrupt storm. That single fact explains why irq_bind and irq_ack are two
syscalls and not one.
This is also why interruptDispatch (arch/x86_64/idt.zig) no longer issues the EOI
itself. It used to, before running the handler — correct for the LAPIC timer, and
impossible for a routed device line. Each handler now owns its EOI, because only the
handler knows which discipline its source needs.
The driver side is an event loop, not a callback
IPC_ReplyWait returns either a client request or a notification, told apart by
the top bit of the badge (ipc_sync.notify_badge_bit). So a driver is one
single-threaded loop over both of its event sources:
while (true) {
const r = ipc.replyWait(endpoint, reply, &recv);
if (r.isNotification()) { // r.source() is the GSI
service_device(); // clear the status register
_ = dev.irqAck(id, irq_res); // re-arm
} else {
handle_client_request(recv[0..r.len]);
}
}
No reentrancy, no "what am I allowed to call from an interrupt handler", no shared state between ISR and task context. The interrupt is just a message.
Two properties worth knowing:
- An interrupt taken while you're elsewhere is not lost. If the driver is off in
an
ipc_callto another server when the IRQ fires,wakeLockedfinds nobody waiting, but the badge is already on the endpoint's notify ring. The nextreplyWaitpops it (ipc_sync.replyWaitcheckspopNotifybefore the sender FIFO). - Notifications coalesce, they don't count. The ring is 8 deep and drops on
overflow. That's correct: an IRQ notification is a level ("the device wants
attention"), not a tally. Re-read the device's status register; never assume one
notification means exactly one event. Because the ISR masks the line until you
irq_ack, at most one badge per GSI can be outstanding — so the ring can only overflow if you bind more than eight GSIs to a single endpoint. Don't.
A whole driver
A minimal leaf driver is only ~150 lines and does all of it. danos ships no such
example binary — the driver model is proven by the real drivers (pci-bus, ps2-bus,
usb-xhci-bus), and a teaching example belongs here, in the docs, rather than as a
compiled program nobody runs. Illustrated with a hypothetical HPET timer driver, the
shape is:
const hpet = findHpet(buf) orelse return; // device_enumerate, look for
// class=timer with memory + irq
_ = dev.claim(hpet.dev_id); // the capability
const base = dev.mmioMap(hpet.dev_id, hpet.mmio).?;
const endpoint = ipc.createIpcEndpoint().?;
// program the hardware over the mapping we were just handed
reg(base, 0x100).* = level | int_enb | (hpet.gsi << 9); // timer 0 config
reg(base, 0x108).* = reg(base, 0xF0).* + period; // comparator
reg(base, 0x010).* |= 1; // ENABLE
_ = dev.irqBind(hpet.dev_id, hpet.irq, endpoint);
while (...) {
const r = ipc.replyWait(endpoint, &.{}, &recv); // blocked. not polling.
if (r.badge & notify_bit == 0) continue;
reg(base, 0x020).* = 1; // clear status -> deassert
reg(base, 0x108).* = reg(base, 0xF0).* + period; // re-arm
_ = dev.irqAck(hpet.dev_id, hpet.irq); // unmask
}
The HPET makes a good illustration for a reason that isn't obvious. Its counter is a
clocksource — the only way to use it is to read it, so it exercises mmio_map without
needing interrupts at all. Its comparators are a clockevent, and can be configured
level-triggered (Tn_INT_TYPE_CNF), which asserts a bit in GENERAL_INT_STATUS
that the driver must write-1-to-clear. That's a genuine deassert step, so the full
mask/ack cycle above is exercised for real rather than being decoration on an
edge-triggered line that would have been fine without it.
One wrinkle it also demonstrates: the ACPI HPET table carries no interrupt number.
Which I/O APIC inputs a comparator may drive is a bitmask in Tn_INT_ROUTE_CAP, in
the device's own registers. So discovery (acpi.parseHpet) maps the block, reads the
mask, and records one concrete GSI as an irq resource. The driver then programs
Tn_INT_ROUTE_CNF to raise exactly that line — and the kernel will only bind the one
it recorded. Hardware that describes itself at runtime still has to fit through a
static capability.
Publishing children: device_register
A device that contains other devices — a PCI bridge, a USB hub, or the HPET's block
of comparators — needs a driver that enumerates it and tells the kernel what it found.
That's device_register, and it makes the device table a tree rather than a list
(DeviceDesc.parent).
var child = std.mem.zeroes(dev.DeviceDesc);
child.class = @intFromEnum(dev.DeviceClass.timer);
child.resource_count = 1;
child.resources[0] = .{ .kind = memory, .start = bus_base + 0x100, .len = 0x20 };
const child_id = dev.register(bus_id, &child).?;
The child is left unclaimed, which is the whole point: another process claims it and
mmio_maps it, and sees only that 0x20-byte window.
The rule the kernel enforces is containment: every resource of a child must lie
inside a resource of the same kind on its parent. Ranges must nest; an IRQ must match
exactly. This isn't bureaucracy — a DeviceDesc is a licence to map physical memory, so
without containment device_register would be a syscall for mapping any page you like. A
bus driver may only ever subdivide what it already owns.
A device with no resources is legal and common. A USB device is reached through its
controller, not by MMIO, so it gets resource_count = 0.
See system/drivers/pci-bus/pci-bus.zig for a
real one — it claims a PCI host bridge, maps its ECAM window, and publishes each function
it finds as a child — and driver-model.md for how bus drivers, class
drivers and host controller drivers fit together.
What the kernel does not do for you
- It does not quiet your device. That's the whole reason
irq_ackexists. - It does not know your registers.
mmio_maphands you a base address; every offset in this document came from the HPET spec, not from danos. - It does not serialise your driver. Two clients calling one driver endpoint are
serialised by
replyWait, but nothing stops your driver from being preempted.
Limits, today
Worth knowing before you write the second driver:
Several things this list used to warn about are now available (see
driver-model.md): port I/O (io_read/io_write, claim-gated by
the device's io_port resource — direct ring-3 in/out is still a #GP, so a PS/2 or
16550 driver goes through these), DMA memory (dma_alloc: contiguous, pinned,
uncacheable, physical address exposed), and memory barriers (/lib/mmio's
mb/rmb/wmb). What remains:
- Page granularity.
mmio_maprounds to 4 KiB. Two devices sharing a page means granting one grants the other. Adevice_registered child's resource can be narrower than a page, but its mapping can't. - DMA is not contained. A driver that can program a bus-mastering device can make
that device write to any physical address — page tables don't sit between a device
and RAM; an IOMMU does. The IOMMU is now detected (M16), but no translation domains
are programmed, so
device_claimon a DMA-capable device is still effectively equivalent to granting ring 0. This is the largest gap between the design's promise and what it delivers; enforcement lands with the first DMA driver. - No
dev_release. A claim is never dropped (only IRQ/MSI bindings are, on exit), so a device stays owned for the life of its driver — which blocks restart. - One endpoint per GSI, so shared legacy PCI INTx lines can't be split between two
drivers. MSI/MSI-X — one vector per device, edge-triggered, unshared — is the real
answer, and QEMU's HPET doesn't offer it (
Tn_FSB_INT_DEL_CAP = 0). - Polarity is hardcoded active-high in
irq.bind. A device whose MADT override says active-low needs that threaded through from discovery. - 14 device vectors (33–46) and 24 GSIs, bounded by the stubs
isr.semits and by a single I/O APIC. - Don't bind more than 8 GSIs to one endpoint. The notify ring is 8 deep and drops
on overflow. With one GSI per endpoint that's unreachable — the line is masked from
the ISR until
irq_ack, so at most one badge is ever outstanding. Bind nine devices to one endpoint, though, and a dropped badge leaves that line masked with nobody left to ack it. - A faulting driver still kills the machine. There is no per-process kill path: a
ring-3 page fault halts the kernel, so
releaseIrqsruns only on a voluntaryexit. Fault isolation is the whole premise (vision) and it is not built yet. - A dead driver's device is not reclaimed.
releaseIrqsunbinds and masks the line on exit, but the claim is never released — restart is not built. - On real hardware, the mask/EOI cycle may need a remote-IRR flush. Masking a
level-triggered redirection entry with remote-IRR set doesn't clear it on some
chipsets, and the line never fires again. QEMU clears it on EOI regardless, so the
tests can't see this. Linux flushes remote-IRR by toggling the entry to edge and
back. See the note at the top of
system/kernel/irq.zig.
Verifying it
No demo driver ships to prove this end to end; the real drivers do, so the tests target them and the kernel primitives directly:
device-manager— boots only the device manager, which discovers the PCI host bridge, matchespci-bus, andsystem_spawns it. The test reads kernel state — the process table and the device tree — to confirm pci-bus came up and registered the functions it enumerated: the whole discover → match → spawn → driver-up chain.acpi-ps2— a user-space driver (ps2-bus) is woken by its device's IRQ, delivered as an IPC notification, and attaches the keyboard: IRQ-as-IPC, end to end.pci-scan— a user-space driver (pci-bus) maps its device's MMIO (the ECAM window) and walks it:mmio_map, end to end.containment— the kernel refuses adevice_registerwhose child window escapes the parent's grant (else it would be a syscall for mapping arbitrary memory), while an identical re-register stays idempotent. Asserted in-kernel, straight against the broker.irqfree— the teardown path. Binds two owners to one shared endpoint, releases one, and reads the I/O APIC back: the departing owner's line is masked, the sibling's is not. That second half is why bindings are keyed on the owning task and not on the endpoint pointer — endpoints are shared, so releasing "everything pointing at this endpoint" would silently mask a live driver's device.iopass— thedevice_grantteardown rule, so destroying a driver's address space never returns MMIO frames to the RAM pool.
$ python3 test/qemu_test.py device-manager acpi-ps2 pci-scan containment irqfree iopass
device-manager ... PASS (matched 'DANOS-TEST-RESULT: PASS')
acpi-ps2 ... PASS
pci-scan ... PASS (matched 'DANOS-TEST-RESULT: PASS')
containment ... PASS (matched 'DANOS-TEST-RESULT: PASS')
irqfree ... PASS (matched 'DANOS-TEST-RESULT: PASS')
iopass ... PASS (matched 'DANOS-TEST-RESULT: PASS')
What's next (not done here)
The big driver-model pieces — capability passing (class drivers), DMA + barriers, MSI,
and IOMMU detection — are now done (driver-model.md, M13–M16), as
is port I/O (io_read/io_write, the claim-gated syscalls that make a PS/2 or 16550
driver possible). What's left is IOMMU enforcement (per-device domains — it waits on
the first DMA driver to protect and test against) and these smaller items:
- Releasing a claim. There is no
dev_release, anddevices_brokernever drops a claim on exit — only IRQ bindings are released. A dead driver's device stays owned forever, which blocks restart. - Unregistering children.
device_registeronly appends. A USB device that is unplugged cannot be removed, and a bus driver in a loop can exhaust the 64-entry table. - Restart. A supervisor that spawns drivers now exists — the device-manager starts
them with
system_spawn— but a supervisor that restarts them does not. A driver that dies should release its claim, have its device quiesced, and be respawned; today nothing notices the death. Some pieces (releaseIrqs,device_grantteardown, the claim table) exist, anddev_release(below) is the missing mechanism; the restart policy is the resilience track (resilience.md). - Interrupt priority / threaded IRQ latency.
notifyFromIsrenqueues the woken driver but doesn't preempt (wakeLockeddeliberately leaves that to the caller), so a woken driver waits for the next scheduling point.
The driver contract (M17–M18)
Claiming and mapping is half of being a danos driver; the other half is the lifecycle and protocol contract, and the runtime makes it nearly free:
- Build on
runtime.service.run— one replyWait loop folding protocol requests, signals, and notifications into callbacks. The harness answers the universal zero-length ping and turnsterminateinto a clean exit for you (process-lifecycle.md). - A driver spawned with an assignment (its device id as argv[1]) sends the
versioned
helloto the device manager inside the deadline, and a bus driver reports what it discovers withchild_added(device-manager.md; usb-xhci-bus is the reference implementation). - Crash freely — that is the design. The kernel releases your claims, IRQ bindings, and MSI vectors at death; the manager reads your exit reason, prunes what you reported, restarts you with backoff, and your fresh instance re-claims and re-reports. Never depend on your own cleanup running (iron rule 1).