danos/docs/device-driver-development/drivers.md

23 KiB
Raw Blame History

Writing a driver

In a monolithic kernel a driver is a function call away from everything: it runs in ring 0, dereferences any physical address, and its interrupt handler is the ISR. In danos a driver is an ordinary ring-3 process. It has its own address space, it can crash without taking the kernel with it, and — the point of this document — it can be restarted (resilience).

That leaves three questions the kernel has to answer, because a process can't answer them for itself:

  1. What hardware exists?device_enumerate, over the device table discovery built (discovery, acpi).
  2. How do I touch its registers?device_claim + mmio_map: the kernel maps the device's physical MMIO window into your address space, and from then on it's plain memory. No syscall per register access.
  3. How do I find out it wants something?irq_bind: the interrupt is delivered to you as an IPC notification. You block; the hardware wakes you.

A driver is, in one sentence, a process that sleeps until its device has something to say.

How a driver gets started: discover, match, spawn

Nothing in the kernel decides that the PCI host bridge needs the pci-bus driver — that is policy, and policy lives in user space. Boot brings user space up as a three-level supervision hierarchy, each level owning one job:

kernel  ──spawns──►  init (PID 1)  ──spawns──►  device-manager  ──spawns──►  pci-bus
  |                     |                          |
  spawns only init,     the service supervisor:    the driver supervisor: enumerates
  publishes the         starts the system          /system/devices, matches each device
  initial-ramdisk       services (device-manager,  to a driver, and system_spawn's it
  so user space can     fat, logger, ...). Its
  system_spawn from it  list is init policy.

The kernel launches exactly one process — init — and hands it nothing but the raw ability to start more (system_spawn(name, arguments), which loads a binary bundled in the initial-ramdisk as a fresh ring-3 process — name becoming its argv[0], the optional arguments its argv[1..], on a SysV entry stack, see sysv.md). Everything else is a user-space decision:

  • init (system/services/init) is the service supervisor. It spawns the system services danos brings up at boot — today input, the device-manager, fat, display, display-demo, and the logger — from a small list. Drivers are deliberately not its job. (An earlier draft listed a vfs service here; that service is retired — the router moved into the kernel as fs_resolve.)
  • device-manager (system/services/device-manager) is the driver supervisor. It does the three steps a monolithic kernel would do in its probe path, entirely from ring 3:
    1. Discoverdevice_enumerate snapshots the device table the kernel built from ACPI/PCI (discovery).
    2. Match — for each device it looks up a driver. The match policy is code, a few small per-bus tables: from the boot snapshot only the PCI host bridge matches (→ pci-bus); everything else arrives later as bus reports and matches on identity — pciDriverForIdentity (xHCI → usb-xhci-bus, virtio-gpu → virtio-gpu), hidDriverFor (PNP0303/PNP0F13 → ps2-bus), and usbDriverForIdentity (USB keyboard, mouse, storage). A fuller system reads what each driver binds (a manifest under /system/drivers, or the driver describing its own match).
    3. Spawnsystem_spawn(driver_name, arguments) starts the matched driver (the arguments can carry which device it matched), which then claims its device and runs the event loop below.

So "how is a driver discovered and configured" has two halves: discovery is the kernel's device table, read by anyone; configuration is two user-space policies — init's service list and the device-manager's match table. Both are hardcoded in their respective programs today; the natural next step is to move them into /etc (see the milestone notes in driver-model.md). system_spawn is currently ungated — any process may spawn any bundled binary — because there is no spawn capability yet.

The capability: claim before touch

The driver syscall numbers (system/abi.zig) with the device types they carry (library/device/model/device-abi.zig), dispatched in system/kernel/process.zig:

# Call Meaning
11 device_enumerate(buf, max) -> total Snapshot the device table
12 device_claim(id) -> ok Take exclusive ownership
13 mmio_map(id, res_idx) -> virtual_address Map a claimed device's register window
14 irq_bind(id, res_idx, endpoint) Deliver that device's IRQ as a notification
15 irq_ack(id, res_idx) Re-arm the IRQ after servicing the device
16 device_register(parent_id, desc) -> id Publish a child of a device you claimed

Notice that nothing takes a physical address or an interrupt number. Every call names a device by id and a resource by index. That indirection is the entire security model. If mmio_map took a physical address, any process could map the kernel's memory; if irq_bind took a GSI, any process could bind the keyboard's line and silently intercept it. Instead the kernel checks two things (process.ownedGsi, and the same check at the top of systemMmioMap):

  • devices_broker.ownerOf(dev_id) == me — you claimed it, and claims are exclusive
  • the resource at res_idx is of the right kindmemory for mmio_map, irq for irq_bind

The claim is the capability. Everything else follows from it.

Registers: mmio_map

mmio_map walks the caller's page tables and installs the device's physical frames with present | user | writable | nx | device_grant plus a cache mode (system/kernel/architecture/x86_64/paging.zig:mapUserDeviceInto). Two of those bits are load-bearing:

  • pcd | pwt — strong-uncacheable, the default cache mode. A device register is not memory; a cached read would return a stale value and a write might never leave the CPU. The one exception: a resource flagged write-combining (resource_flag_write_combining — today the kernel-seeded display framebuffer) gets the PAT bit instead, so pixel writes batch into bursts.
  • device_grant (bit 9, one of the PTE's available bits) — marks the leaf as MMIO rather than RAM, so freeSubtree skips pmm.free on it when the address space is destroyed. Without this, killing a driver would hand the HPET's registers back to the frame allocator as if they were free RAM. The iopass test guards it.

Grants land in their own arena, 0x0000_7100_0000_0000 (PML4[226]), so device pages never widen an existing mapping.

Then you just… use it:

const base = dev.mmioMap(dev_id, mmio_res) orelse return;
const counter: *volatile u64 = @ptrFromInt(base + 0xF0);
const now = counter.*;   // a load, straight to the hardware. no kernel involved.

Interrupts: the cycle, and why it has that shape

An interrupt handler in a microkernel has a problem. The code that knows how to quiet the device is in ring 3, in another address space, and it will not run for microseconds or milliseconds — after a context switch, when the scheduler gets to it. But the CPU wants an EOI now, and a level-triggered line stays asserted until the device is quieted. EOI a still-asserted line and the I/O APIC redelivers immediately. Forever. The driver never gets to run at all.

The way out is to mask the line before acknowledging it:

kernel ISR   irqMask(gsi)      // line still asserted; stop it reaching a CPU
             irqEoi()          // now safe to tell the LAPIC we're done
             notifyFromIsr()   // wake the driver — it runs much later

driver       replyWait() -> badge with the notify bit set
             <clear the device's status register>   // NOW the line deasserts
             irq_ack(dev, res) // kernel unmasks: quiet, so it can't refire

irq_ack is not bookkeeping you could skip. It is the unmask. Forget it and the interrupt fires exactly once, ever; call it before the device is quiet and you get an interrupt storm. That single fact explains why irq_bind and irq_ack are two syscalls and not one.

This is also why interruptDispatch (system/kernel/architecture/x86_64/idt.zig) no longer issues the EOI itself. It used to, before running the handler — correct for the LAPIC timer, and impossible for a routed device line. Each handler now owns its EOI, because only the handler knows which discipline its source needs.

The driver side is an event loop, not a callback

IPC_ReplyWait returns either a client request or a notification, told apart by the top bit of the badge (ipc_sync.notify_badge_bit). So a driver is one single-threaded loop over both of its event sources:

while (true) {
    const r = ipc.replyWait(endpoint, reply, &recv);
    if (r.isNotification()) {            // r.source() is the GSI
        service_device();                // clear the status register
        _ = dev.irqAck(id, irq_res);     // re-arm
    } else {
        handle_client_request(recv[0..r.len]);
    }
}

No reentrancy, no "what am I allowed to call from an interrupt handler", no shared state between ISR and task context. The interrupt is just a message.

Two properties worth knowing:

  • An interrupt taken while you're elsewhere is not lost. If the driver is off in an ipc_call to another server when the IRQ fires, wakeLocked finds nobody waiting, but the badge is already on the endpoint's notify ring. The next replyWait pops it (ipc_sync.replyWait checks popNotify before the sender FIFO).
  • Notifications coalesce, they don't count. The ring is 8 deep and drops on overflow. That's correct: an IRQ notification is a level ("the device wants attention"), not a tally. Re-read the device's status register; never assume one notification means exactly one event. Because the ISR masks the line until you irq_ack, at most one badge per GSI can be outstanding — so the ring can only overflow if you bind more than eight GSIs to a single endpoint. Don't.

A whole driver

A minimal leaf driver is only ~150 lines and does all of it. danos ships no such example binary — the driver model is proven by the real drivers (pci-bus, ps2-bus, usb-xhci-bus), and a teaching example belongs here, in the docs, rather than as a compiled program nobody runs. Illustrated with a hypothetical HPET timer driver, the shape is:

const hpet = findHpet(buf) orelse return;      // device_enumerate, look for
                                               //   class=timer with memory + irq
_ = dev.claim(hpet.dev_id);                    // the capability
const base = dev.mmioMap(hpet.dev_id, hpet.mmio).?;
const endpoint = ipc.createIpcEndpoint().?;

// program the hardware over the mapping we were just handed
reg(base, 0x100).* = level | int_enb | (hpet.gsi << 9);   // timer 0 config
reg(base, 0x108).* = reg(base, 0xF0).* + period;          // comparator
reg(base, 0x010).* |= 1;                                  // ENABLE

_ = dev.irqBind(hpet.dev_id, hpet.irq, endpoint);

while (...) {
    const r = ipc.replyWait(endpoint, &.{}, &recv);       // blocked. not polling.
    if (r.badge & notify_bit == 0) continue;
    reg(base, 0x020).* = 1;                               // clear status -> deassert
    reg(base, 0x108).* = reg(base, 0xF0).* + period;      // re-arm
    _ = dev.irqAck(hpet.dev_id, hpet.irq);                // unmask
}

The HPET makes a good illustration for a reason that isn't obvious. Its counter is a clocksource — the only way to use it is to read it, so it exercises mmio_map without needing interrupts at all. Its comparators are a clockevent, and can be configured level-triggered (Tn_INT_TYPE_CNF), which asserts a bit in GENERAL_INT_STATUS that the driver must write-1-to-clear. That's a genuine deassert step, so the full mask/ack cycle above is exercised for real rather than being decoration on an edge-triggered line that would have been fine without it.

One wrinkle it also demonstrates: the ACPI HPET table carries no interrupt number. Which I/O APIC inputs a comparator may drive is a bitmask in Tn_INT_ROUTE_CAP, in the device's own registers. So discovery (acpi.parseHpet) maps the block, reads the mask, and records one concrete GSI as an irq resource. The driver then programs Tn_INT_ROUTE_CNF to raise exactly that line — and the kernel will only bind the one it recorded. Hardware that describes itself at runtime still has to fit through a static capability.

Publishing children: device_register

A device that contains other devices — a PCI bridge, a USB hub, or the HPET's block of comparators — needs a driver that enumerates it and tells the kernel what it found. That's device_register, and it makes the device table a tree rather than a list (DeviceDescriptor.parent).

var child = std.mem.zeroes(dev.DeviceDescriptor);
child.class = @intFromEnum(dev.DeviceClass.timer);
child.resource_count = 1;
child.resources[0] = .{ .kind = memory, .start = bus_base + 0x100, .len = 0x20 };
const child_id = dev.register(bus_id, &child).?;

The child is left unclaimed, which is the whole point: another process claims it and mmio_maps it, and sees only that 0x20-byte window.

The rule the kernel enforces is containment: every resource of a child must lie inside a resource of the same kind on its parent. Ranges must nest; a child's IRQ — still exactly one line — must fall within the parent's IRQ range (a length-1 parent range is the old exact-match rule). This isn't bureaucracy — a DeviceDescriptor is a licence to map physical memory, so without containment device_register would be a syscall for mapping any page you like. A bus driver may only ever subdivide what it already owns.

A device with no resources is legal and common. A USB device is reached through its controller, not by MMIO, so it gets resource_count = 0.

See system/drivers/pci-bus/pci-bus.zig for a real one — it claims a PCI host bridge, maps its ECAM window, and publishes each function it finds as a child — and driver-model.md for how bus drivers, class drivers and host controller drivers fit together.

What the kernel does not do for you

  • It does not quiet your device. That's the whole reason irq_ack exists.
  • It does not know your registers. mmio_map hands you a base address; every offset in this document came from the HPET spec, not from danos.
  • It does not serialise your driver. Two clients calling one driver endpoint are serialised by replyWait, but nothing stops your driver from being preempted.

Limits, today

Worth knowing before you write the second driver:

Several things this list used to warn about are now available (see driver-model.md): port I/O (io_read/io_write, claim-gated by the device's io_port resource — direct ring-3 in/out is still a #GP, so a PS/2 or 16550 driver goes through these), DMA memory (dma_alloc: contiguous, pinned, uncacheable, physical address exposed), memory barriers (library/device/mmio's memoryBarrier/readMemoryBarrier/writeMemoryBarrier, imported as the mmio module), fault isolation (a ring-3 fault kills only the faulting process — killCurrentProcess — and the machine keeps running, resilience), and reclaim + restart on death (every path out of a process releases its claims and IRQ/MSI bindings — releaseAllOwnedBy, irq.releaseOwner — and the device manager respawns the driver with backoff, device-manager.md). What remains:

  • Page granularity. mmio_map rounds to 4 KiB. Two devices sharing a page means granting one grants the other. A device_registered child's resource can be narrower than a page, but its mapping can't.
  • DMA is not contained. A driver that can program a bus-mastering device can make that device write to any physical address — page tables don't sit between a device and RAM; an IOMMU does. The IOMMU is now detected (M16), but no translation domains are programmed, so device_claim on a DMA-capable device is still effectively equivalent to granting ring 0. This is the largest gap between the design's promise and what it delivers; enforcement lands with the first DMA driver.
  • No voluntary dev_release. A live driver can't drop a claim — only exit releases it (any path out of a process runs releaseAllOwnedBy) — so handing a device between running drivers still means exiting.
  • One endpoint per GSI, so shared legacy PCI INTx lines can't be split between two drivers. MSI/MSI-X — one vector per device, edge-triggered, unshared — is the real answer, and QEMU's HPET doesn't offer it (Tn_FSB_INT_DEL_CAP = 0).
  • Polarity is hardcoded active-high in irq.bind. A device whose MADT override says active-low needs that threaded through from discovery.
  • 14 device vectors (3346) and 24 GSIs, bounded by the stubs isr.s emits and by a single I/O APIC.
  • Don't bind more than 8 GSIs to one endpoint. The notify ring is 8 deep and drops on overflow. With one GSI per endpoint that's unreachable — the line is masked from the ISR until irq_ack, so at most one badge is ever outstanding. Bind nine devices to one endpoint, though, and a dropped badge leaves that line masked with nobody left to ack it.
  • On real hardware, the mask/EOI cycle may need a remote-IRR flush. Masking a level-triggered redirection entry with remote-IRR set doesn't clear it on some chipsets, and the line never fires again. QEMU clears it on EOI regardless, so the tests can't see this. Linux flushes remote-IRR by toggling the entry to edge and back. See the note at the top of system/kernel/irq.zig.

Verifying it

No demo driver ships to prove this end to end; the real drivers do, so the tests target them and the kernel primitives directly:

  • device-manager — boots only the device manager, which discovers the PCI host bridge, matches pci-bus, and system_spawns it. The test reads kernel state — the process table and the device tree — to confirm pci-bus came up and registered the functions it enumerated: the whole discover → match → spawn → driver-up chain.
  • acpi-ps2 — a user-space driver (ps2-bus) is woken by its device's IRQ, delivered as an IPC notification, and attaches the keyboard: IRQ-as-IPC, end to end.
  • pci-scan — a user-space driver (pci-bus) maps its device's MMIO (the ECAM window) and walks it: mmio_map, end to end.
  • containment — the kernel refuses a device_register whose child window escapes the parent's grant (else it would be a syscall for mapping arbitrary memory), while an identical re-register stays idempotent. Asserted in-kernel, straight against the broker.
  • irqfree — the teardown path. Binds two owners to one shared endpoint, releases one, and reads the I/O APIC back: the departing owner's line is masked, the sibling's is not. That second half is why bindings are keyed on the owning task and not on the endpoint pointer — endpoints are shared, so releasing "everything pointing at this endpoint" would silently mask a live driver's device.
  • iopass — the device_grant teardown rule, so destroying a driver's address space never returns MMIO frames to the RAM pool.
$ python3 test/qemu_test.py device-manager acpi-ps2 pci-scan containment irqfree iopass
  device-manager ... PASS  (matched 'DANOS-TEST-RESULT: PASS')
  acpi-ps2     ... PASS
  pci-scan     ... PASS  (matched 'DANOS-TEST-RESULT: PASS')
  containment  ... PASS  (matched 'DANOS-TEST-RESULT: PASS')
  irqfree      ... PASS  (matched 'DANOS-TEST-RESULT: PASS')
  iopass       ... PASS  (matched 'DANOS-TEST-RESULT: PASS')

What's next (not done here)

The big driver-model pieces — capability passing (class drivers), DMA + barriers, MSI, and IOMMU detection — are now done (driver-model.md, M13M16), as is port I/O (io_read/io_write, the claim-gated syscalls that make a PS/2 or 16550 driver possible). What's left is IOMMU enforcement (per-device domains — it waits on the first DMA driver to protect and test against) and these smaller items:

  • Releasing a claim — half done. The kernel now drops all of a dead driver's claims on every path out of a process (releaseAllOwnedBy, called from process teardown), which unblocked restart. A voluntary dev_release for a live driver still doesn't exist.
  • Unregistering children — half done. Hot-remove works at the manager layer: the xHCI bus reports child_removed on unplug and the device manager prunes its tree. The kernel's own device table is still append-only, so a bus driver in a loop can still exhaust the 64-entry table.
  • Restart — done. The device manager notices a driver's death, reads its exit reason, prunes the children it reported, and respawns it with exponential backoff — with a crash-loop cap that marks a repeat offender failed instead of respawning forever (device-manager.md).
  • Interrupt priority / threaded IRQ latency — still open. notifyFromIsr enqueues the woken driver but doesn't preempt (wakeLocked deliberately leaves that to the caller), so a woken driver waits for the next scheduling point.

The driver contract (M17M18)

Claiming and mapping is half of being a danos driver; the other half is the lifecycle and protocol contract, and the runtime makes it nearly free:

  • Build on service.run — one replyWait loop folding protocol requests, signals, and notifications into callbacks. The harness answers the universal zero-length ping and turns terminate into a clean exit for you (process-lifecycle.md).
  • A driver spawned with an assignment (its device id as argv[1]) sends the versioned hello to the device manager inside the deadline, and a bus driver reports what it discovers with child_added (device-manager.md; usb-xhci-bus is the reference implementation).
  • Crash freely — that is the design. The kernel releases your claims, IRQ bindings, and MSI vectors at death; the manager reads your exit reason, prunes what you reported, restarts you with backoff, and your fresh instance re-claims and re-reports. Never depend on your own cleanup running (iron rule 1).