396 lines
22 KiB
Markdown
396 lines
22 KiB
Markdown
# Writing a driver
|
||
|
||
In a monolithic kernel a driver is a function call away from everything: it runs in
|
||
ring 0, dereferences any physical address, and its interrupt handler *is* the ISR. In
|
||
danos a driver is **an ordinary ring-3 process**. It has its own address space, it
|
||
can crash without taking the kernel with it, and — the point of this document — it
|
||
can be restarted ([resilience](resilience.md)).
|
||
|
||
That leaves three questions the kernel has to answer, because a process can't answer
|
||
them for itself:
|
||
|
||
1. **What hardware exists?** → `device_enumerate`, over the device table discovery built
|
||
([discovery](discovery.md), [acpi](acpi.md)).
|
||
2. **How do I touch its registers?** → `device_claim` + `mmio_map`: the kernel maps the
|
||
device's physical MMIO window into your address space, and from then on it's plain
|
||
memory. No syscall per register access.
|
||
3. **How do I find out it wants something?** → `irq_bind`: the interrupt is delivered
|
||
to you as an IPC notification. You block; the hardware wakes you.
|
||
|
||
A driver is, in one sentence, *a process that sleeps until its device has something to
|
||
say.*
|
||
|
||
## How a driver gets started: discover, match, spawn
|
||
|
||
Nothing in the kernel decides that the PCI host bridge needs the `pci-bus` driver — that
|
||
is policy, and policy lives in user space. Boot brings user space up as a three-level
|
||
supervision hierarchy, each level owning one job:
|
||
|
||
```
|
||
kernel ──spawns──► init (PID 1) ──spawns──► device-manager ──spawns──► pci-bus
|
||
| | |
|
||
spawns only init, the service supervisor: the driver supervisor: enumerates
|
||
publishes the starts the system /system/devices, matches each device
|
||
initial-ramdisk services (vfs, the to a driver, and system_spawn's it
|
||
so user space can device-manager). Its
|
||
system_spawn from it list is init policy.
|
||
```
|
||
|
||
The kernel launches exactly one process — `init` — and hands it nothing but the raw
|
||
ability to start more (`system_spawn(name, arguments)`, which loads a binary bundled
|
||
in the initial-ramdisk as a fresh ring-3 process — `name` becoming its argv[0],
|
||
the optional arguments its argv[1..], on a SysV entry stack, see sysv.md). Everything else is a user-space decision:
|
||
|
||
- **init** ([system/services/init](system/services/init/init.zig)) is the **service
|
||
supervisor**. It spawns the system services danos brings up at boot — today `vfs` and
|
||
the `device-manager` — from a small list. Drivers are deliberately *not* its job.
|
||
- **device-manager** ([system/services/device-manager](system/services/device-manager/device-manager.zig))
|
||
is the **driver supervisor**. It does the three steps a monolithic kernel would do in
|
||
its probe path, entirely from ring 3:
|
||
1. **Discover** — `device_enumerate` snapshots the device table the kernel built from
|
||
ACPI/PCI ([discovery](discovery.md)).
|
||
2. **Match** — for each device it looks up a driver by `DeviceClass`. The match policy
|
||
is a table (`driverFor`): today a static `timer → hpet` map; a fuller system reads
|
||
what each driver *binds* (a manifest under `/system/drivers`, or the driver
|
||
describing its own match).
|
||
3. **Spawn** — `system_spawn(driver_name, arguments)` starts the matched driver (the
|
||
arguments can carry *which* device it matched), which then claims
|
||
its device and runs the event loop below.
|
||
|
||
So "how is a driver discovered and configured" has two halves: **discovery** is the
|
||
kernel's device table, read by anyone; **configuration** is two user-space policies —
|
||
init's service list and the device-manager's match table. Both are hardcoded in their
|
||
respective programs today; the natural next step is to move them into `/etc` (see the
|
||
milestone notes in [driver-model.md](driver-model.md)). `system_spawn` is currently
|
||
ungated — any process may spawn any bundled binary — because there is no spawn
|
||
capability yet.
|
||
|
||
## The capability: claim before touch
|
||
|
||
The driver syscall numbers (`system/abi.zig`) with the device types they carry
|
||
(`system/devices/device-abi.zig`), dispatched in `system/kernel/process.zig`:
|
||
|
||
| # | Call | Meaning |
|
||
|---|------|---------|
|
||
| 11 | `device_enumerate(buf, max) -> total` | Snapshot the device table |
|
||
| 12 | `device_claim(id) -> ok` | Take **exclusive** ownership |
|
||
| 13 | `mmio_map(id, res_idx) -> virtual_address` | Map a claimed device's register window |
|
||
| 14 | `irq_bind(id, res_idx, endpoint)` | Deliver that device's IRQ as a notification |
|
||
| 15 | `irq_ack(id, res_idx)` | Re-arm the IRQ after servicing the device |
|
||
| 16 | `device_register(parent_id, desc) -> id` | Publish a child of a device you claimed |
|
||
|
||
Notice that **nothing takes a physical address or an interrupt number.** Every call
|
||
names a device by id and a resource by index. That indirection is the entire security
|
||
model. If `mmio_map` took a physical address, any process could map the kernel's
|
||
memory; if `irq_bind` took a GSI, any process could bind the keyboard's line and
|
||
silently intercept it. Instead the kernel checks two things (`process.ownedGsi`, and
|
||
the same check at the top of `sysMmioMap`):
|
||
|
||
- `devices_broker.ownerOf(dev_id) == me` — you claimed it, and claims are exclusive
|
||
- the resource at `res_idx` is of the right *kind* — `memory` for `mmio_map`, `irq`
|
||
for `irq_bind`
|
||
|
||
The claim is the capability. Everything else follows from it.
|
||
|
||
## Registers: `mmio_map`
|
||
|
||
`mmio_map` walks the caller's page tables and installs the device's physical frames
|
||
with `present | user | writable | nx | pcd | pwt`
|
||
(`arch/x86_64/paging.zig:mapUserDeviceInto`). Two of those bits are load-bearing:
|
||
|
||
- **`pcd | pwt`** — strong-uncacheable. A device register is not memory; a cached read
|
||
would return a stale value and a write might never leave the CPU.
|
||
- **`device_grant`** (bit 9, one of the PTE's available bits) — marks the leaf as MMIO
|
||
rather than RAM, so `freeSubtree` skips `pmm.free` on it when the address space is
|
||
destroyed. Without this, killing a driver would hand the HPET's registers back to
|
||
the frame allocator as if they were free RAM. The `iopass` test guards it.
|
||
|
||
Grants land in their own arena, `0x0000_7100_0000_0000` (PML4[226]), so device pages
|
||
never widen an existing mapping.
|
||
|
||
Then you just… use it:
|
||
|
||
```zig
|
||
const base = dev.mmioMap(dev_id, mmio_res) orelse return;
|
||
const counter: *volatile u64 = @ptrFromInt(base + 0xF0);
|
||
const now = counter.*; // a load, straight to the hardware. no kernel involved.
|
||
```
|
||
|
||
## Interrupts: the cycle, and why it has that shape
|
||
|
||
An interrupt handler in a microkernel has a problem. The code that knows how to quiet
|
||
the device is in ring 3, in another address space, and it will not run for
|
||
microseconds or milliseconds — after a context switch, when the scheduler gets to it.
|
||
But the CPU wants an EOI *now*, and a **level-triggered** line stays asserted until
|
||
the device is quieted. EOI a still-asserted line and the I/O APIC redelivers
|
||
immediately. Forever. The driver never gets to run at all.
|
||
|
||
The way out is to mask the line before acknowledging it:
|
||
|
||
```
|
||
kernel ISR irqMask(gsi) // line still asserted; stop it reaching a CPU
|
||
irqEoi() // now safe to tell the LAPIC we're done
|
||
notifyFromIsr() // wake the driver — it runs much later
|
||
|
||
driver replyWait() -> badge with the notify bit set
|
||
<clear the device's status register> // NOW the line deasserts
|
||
irq_ack(dev, res) // kernel unmasks: quiet, so it can't refire
|
||
|
||
```
|
||
|
||
`irq_ack` is not bookkeeping you could skip. **It is the unmask.** Forget it and the
|
||
interrupt fires exactly once, ever; call it before the device is quiet and you get an
|
||
interrupt storm. That single fact explains why `irq_bind` and `irq_ack` are two
|
||
syscalls and not one.
|
||
|
||
This is also why `interruptDispatch` (`arch/x86_64/idt.zig`) no longer issues the EOI
|
||
itself. It used to, before running the handler — correct for the LAPIC timer, and
|
||
impossible for a routed device line. Each handler now owns its EOI, because only the
|
||
handler knows which discipline its source needs.
|
||
|
||
### The driver side is an event loop, not a callback
|
||
|
||
`IPC_ReplyWait` returns *either* a client request *or* a notification, told apart by
|
||
the top bit of the badge (`ipc_sync.notify_badge_bit`). So a driver is one
|
||
single-threaded loop over both of its event sources:
|
||
|
||
```zig
|
||
while (true) {
|
||
const r = ipc.replyWait(endpoint, reply, &recv);
|
||
if (r.isNotification()) { // r.source() is the GSI
|
||
service_device(); // clear the status register
|
||
_ = dev.irqAck(id, irq_res); // re-arm
|
||
} else {
|
||
handle_client_request(recv[0..r.len]);
|
||
}
|
||
}
|
||
```
|
||
|
||
No reentrancy, no "what am I allowed to call from an interrupt handler", no shared
|
||
state between ISR and task context. The interrupt is just a message.
|
||
|
||
Two properties worth knowing:
|
||
|
||
- **An interrupt taken while you're elsewhere is not lost.** If the driver is off in
|
||
an `ipc_call` to another server when the IRQ fires, `wakeLocked` finds nobody
|
||
waiting, but the badge is already on the endpoint's notify ring. The next
|
||
`replyWait` pops it (`ipc_sync.replyWait` checks `popNotify` before the sender FIFO).
|
||
- **Notifications coalesce, they don't count.** The ring is 8 deep and drops on
|
||
overflow. That's correct: an IRQ notification is a *level* ("the device wants
|
||
attention"), not a tally. Re-read the device's status register; never assume one
|
||
notification means exactly one event. Because the ISR masks the line until you
|
||
`irq_ack`, at most one badge per GSI can be outstanding — so the ring can only
|
||
overflow if you bind more than eight GSIs to a single endpoint. Don't.
|
||
|
||
## A whole driver
|
||
|
||
A minimal leaf driver is only ~150 lines and does all of it. danos ships **no such
|
||
example binary** — the driver model is proven by the real drivers (`pci-bus`, `ps2-bus`,
|
||
`usb-xhci-bus`), and a teaching example belongs here, in the docs, rather than as a
|
||
compiled program nobody runs. Illustrated with a hypothetical HPET timer driver, the
|
||
shape is:
|
||
|
||
```zig
|
||
const hpet = findHpet(buf) orelse return; // device_enumerate, look for
|
||
// class=timer with memory + irq
|
||
_ = dev.claim(hpet.dev_id); // the capability
|
||
const base = dev.mmioMap(hpet.dev_id, hpet.mmio).?;
|
||
const endpoint = ipc.createIpcEndpoint().?;
|
||
|
||
// program the hardware over the mapping we were just handed
|
||
reg(base, 0x100).* = level | int_enb | (hpet.gsi << 9); // timer 0 config
|
||
reg(base, 0x108).* = reg(base, 0xF0).* + period; // comparator
|
||
reg(base, 0x010).* |= 1; // ENABLE
|
||
|
||
_ = dev.irqBind(hpet.dev_id, hpet.irq, endpoint);
|
||
|
||
while (...) {
|
||
const r = ipc.replyWait(endpoint, &.{}, &recv); // blocked. not polling.
|
||
if (r.badge & notify_bit == 0) continue;
|
||
reg(base, 0x020).* = 1; // clear status -> deassert
|
||
reg(base, 0x108).* = reg(base, 0xF0).* + period; // re-arm
|
||
_ = dev.irqAck(hpet.dev_id, hpet.irq); // unmask
|
||
}
|
||
```
|
||
|
||
The HPET makes a good illustration for a reason that isn't obvious. Its *counter* is a
|
||
clocksource — the only way to use it is to read it, so it exercises `mmio_map` without
|
||
needing interrupts at all. Its *comparators* are a clockevent, and can be configured
|
||
**level-triggered** (`Tn_INT_TYPE_CNF`), which asserts a bit in `GENERAL_INT_STATUS`
|
||
that the driver must write-1-to-clear. That's a genuine deassert step, so the full
|
||
mask/ack cycle above is exercised for real rather than being decoration on an
|
||
edge-triggered line that would have been fine without it.
|
||
|
||
One wrinkle it also demonstrates: the ACPI HPET table carries **no interrupt number**.
|
||
Which I/O APIC inputs a comparator may drive is a bitmask in `Tn_INT_ROUTE_CAP`, in
|
||
the device's own registers. So discovery (`acpi.parseHpet`) maps the block, reads the
|
||
mask, and records one concrete GSI as an `irq` resource. The driver then programs
|
||
`Tn_INT_ROUTE_CNF` to raise exactly that line — and the kernel will only bind the one
|
||
it recorded. Hardware that describes itself at runtime still has to fit through a
|
||
static capability.
|
||
|
||
## Publishing children: `device_register`
|
||
|
||
A device that *contains other devices* — a PCI bridge, a USB hub, or the HPET's block
|
||
of comparators — needs a driver that enumerates it and tells the kernel what it found.
|
||
That's `device_register`, and it makes the device table a tree rather than a list
|
||
(`DeviceDesc.parent`).
|
||
|
||
```zig
|
||
var child = std.mem.zeroes(dev.DeviceDesc);
|
||
child.class = @intFromEnum(dev.DeviceClass.timer);
|
||
child.resource_count = 1;
|
||
child.resources[0] = .{ .kind = memory, .start = bus_base + 0x100, .len = 0x20 };
|
||
const child_id = dev.register(bus_id, &child).?;
|
||
```
|
||
|
||
The child is left **unclaimed**, which is the whole point: another process claims it and
|
||
`mmio_map`s it, and sees only that 0x20-byte window.
|
||
|
||
The rule the kernel enforces is **containment**: every resource of a child must lie
|
||
inside a resource of the same kind on its parent. Ranges must nest; an IRQ must match
|
||
exactly. This isn't bureaucracy — a `DeviceDesc` is a licence to map physical memory, so
|
||
without containment `device_register` would be a syscall for mapping any page you like. A
|
||
bus driver may only ever subdivide what it already owns.
|
||
|
||
A device with **no resources** is legal and common. A USB device is reached through its
|
||
controller, not by MMIO, so it gets `resource_count = 0`.
|
||
|
||
See [`system/drivers/pci-bus/pci-bus.zig`](../system/drivers/pci-bus/pci-bus.zig) for a
|
||
real one — it claims a PCI host bridge, maps its ECAM window, and publishes each function
|
||
it finds as a child — and [driver-model.md](driver-model.md) for how bus drivers, class
|
||
drivers and host controller drivers fit together.
|
||
|
||
## What the kernel does not do for you
|
||
|
||
- **It does not quiet your device.** That's the whole reason `irq_ack` exists.
|
||
- **It does not know your registers.** `mmio_map` hands you a base address; every
|
||
offset in this document came from the HPET spec, not from danos.
|
||
- **It does not serialise your driver.** Two clients calling one driver endpoint are
|
||
serialised by `replyWait`, but nothing stops your driver from being preempted.
|
||
|
||
## Limits, today
|
||
|
||
Worth knowing before you write the second driver:
|
||
|
||
Several things this list used to warn about are now available (see
|
||
[driver-model.md](driver-model.md)): **port I/O** (`io_read`/`io_write`, claim-gated by
|
||
the device's `io_port` resource — direct ring-3 `in`/`out` is still a #GP, so a PS/2 or
|
||
16550 driver goes through these), **DMA memory** (`dma_alloc`: contiguous, pinned,
|
||
uncacheable, physical address exposed), and **memory barriers** (`/lib/mmio`'s
|
||
`mb`/`rmb`/`wmb`). What remains:
|
||
|
||
- **Page granularity.** `mmio_map` rounds to 4 KiB. Two devices sharing a page means
|
||
granting one grants the other. A `device_register`ed child's *resource* can be narrower
|
||
than a page, but its *mapping* can't.
|
||
- **DMA is not contained.** A driver that can program a bus-mastering device can make
|
||
that device write to *any* physical address — page tables don't sit between a device
|
||
and RAM; an IOMMU does. The IOMMU is now *detected* (M16), but no translation domains
|
||
are programmed, so `device_claim` on a DMA-capable device is still effectively
|
||
equivalent to granting ring 0. This is the largest gap between the design's promise and
|
||
what it delivers; enforcement lands with the first DMA driver.
|
||
- **No `dev_release`.** A claim is never dropped (only IRQ/MSI bindings are, on exit), so
|
||
a device stays owned for the life of its driver — which blocks restart.
|
||
- **One endpoint per GSI**, so shared legacy PCI INTx lines can't be split between two
|
||
drivers. MSI/MSI-X — one vector per device, edge-triggered, unshared — is the real
|
||
answer, and QEMU's HPET doesn't offer it (`Tn_FSB_INT_DEL_CAP = 0`).
|
||
- **Polarity is hardcoded** active-high in `irq.bind`. A device whose MADT override
|
||
says active-low needs that threaded through from discovery.
|
||
- **14 device vectors** (33–46) and **24 GSIs**, bounded by the stubs `isr.s` emits and
|
||
by a single I/O APIC.
|
||
- **Don't bind more than 8 GSIs to one endpoint.** The notify ring is 8 deep and drops
|
||
on overflow. With one GSI per endpoint that's unreachable — the line is masked from
|
||
the ISR until `irq_ack`, so at most one badge is ever outstanding. Bind nine devices
|
||
to one endpoint, though, and a dropped badge leaves that line masked with nobody
|
||
left to ack it.
|
||
- **A faulting driver still kills the machine.** There is no per-process kill path: a
|
||
ring-3 page fault halts the kernel, so `releaseIrqs` runs only on a voluntary
|
||
`exit`. Fault isolation is the whole premise ([vision](vision.md)) and it is
|
||
[not built yet](resilience.md).
|
||
- **A dead driver's device is not reclaimed.** `releaseIrqs` unbinds and masks the
|
||
line on exit, but the claim is never released — restart is
|
||
[not built](resilience.md).
|
||
- **On real hardware, the mask/EOI cycle may need a remote-IRR flush.** Masking a
|
||
level-triggered redirection entry with remote-IRR set doesn't clear it on some
|
||
chipsets, and the line never fires again. QEMU clears it on EOI regardless, so the
|
||
tests can't see this. Linux flushes remote-IRR by toggling the entry to edge and
|
||
back. See the note at the top of `system/kernel/irq.zig`.
|
||
|
||
## Verifying it
|
||
|
||
No demo driver ships to prove this end to end; the *real* drivers do, so the tests
|
||
target them and the kernel primitives directly:
|
||
|
||
- **`device-manager`** — boots only the device manager, which discovers the PCI host
|
||
bridge, matches `pci-bus`, and `system_spawn`s it. The test reads kernel state — the
|
||
process table and the device tree — to confirm pci-bus came up and registered the
|
||
functions it enumerated: the whole discover → match → spawn → driver-up chain.
|
||
- **`acpi-ps2`** — a user-space driver (`ps2-bus`) is woken by its device's IRQ,
|
||
delivered as an IPC notification, and attaches the keyboard: IRQ-as-IPC, end to end.
|
||
- **`pci-scan`** — a user-space driver (`pci-bus`) maps its device's MMIO (the ECAM
|
||
window) and walks it: `mmio_map`, end to end.
|
||
- **`containment`** — the kernel refuses a `device_register` whose child window escapes
|
||
the parent's grant (else it would be a syscall for mapping arbitrary memory), while an
|
||
identical re-register stays idempotent. Asserted in-kernel, straight against the broker.
|
||
- **`irqfree`** — the teardown path. Binds two owners to one shared endpoint, releases
|
||
one, and reads the I/O APIC back: the departing owner's line is masked, the sibling's
|
||
is not. That second half is why bindings are keyed on the owning *task* and not on the
|
||
endpoint pointer — endpoints are shared, so releasing "everything pointing at this
|
||
endpoint" would silently mask a live driver's device.
|
||
- **`iopass`** — the `device_grant` teardown rule, so destroying a driver's address
|
||
space never returns MMIO frames to the RAM pool.
|
||
|
||
```
|
||
$ python3 test/qemu_test.py device-manager acpi-ps2 pci-scan containment irqfree iopass
|
||
device-manager ... PASS (matched 'DANOS-TEST-RESULT: PASS')
|
||
acpi-ps2 ... PASS
|
||
pci-scan ... PASS (matched 'DANOS-TEST-RESULT: PASS')
|
||
containment ... PASS (matched 'DANOS-TEST-RESULT: PASS')
|
||
irqfree ... PASS (matched 'DANOS-TEST-RESULT: PASS')
|
||
iopass ... PASS (matched 'DANOS-TEST-RESULT: PASS')
|
||
```
|
||
|
||
## What's next (not done here)
|
||
|
||
The big driver-model pieces — capability passing (class drivers), DMA + barriers, MSI,
|
||
and IOMMU detection — are **now done** ([driver-model.md](driver-model.md), M13–M16), as
|
||
is **port I/O** (`io_read`/`io_write`, the claim-gated syscalls that make a PS/2 or 16550
|
||
driver possible). What's left is IOMMU *enforcement* (per-device domains — it waits on
|
||
the first DMA driver to protect and test against) and these smaller items:
|
||
|
||
- **Releasing a claim.** There is no `dev_release`, and `devices_broker` never drops a claim on
|
||
exit — only IRQ bindings are released. A dead driver's device stays owned forever,
|
||
which blocks restart.
|
||
- **Unregistering children.** `device_register` only appends. A USB device that is
|
||
unplugged cannot be removed, and a bus driver in a loop can exhaust the 64-entry
|
||
table.
|
||
- **Restart.** A supervisor that *spawns* drivers now exists — the device-manager starts
|
||
them with `system_spawn` — but a supervisor that *restarts* them does not. A driver that
|
||
dies should release its claim, have its device quiesced, and be respawned; today nothing
|
||
notices the death. Some pieces (`releaseIrqs`, `device_grant` teardown, the claim table)
|
||
exist, and `dev_release` (below) is the missing mechanism; the restart policy is the
|
||
resilience track ([resilience.md](resilience.md)).
|
||
- **Interrupt priority / threaded IRQ latency.** `notifyFromIsr` enqueues the woken
|
||
driver but doesn't preempt (`wakeLocked` deliberately leaves that to the caller), so
|
||
a woken driver waits for the next scheduling point.
|
||
|
||
## The driver contract (M17–M18)
|
||
|
||
Claiming and mapping is half of being a danos driver; the other half is the
|
||
**lifecycle and protocol contract**, and the runtime makes it nearly free:
|
||
|
||
- Build on `runtime.service.run` — one replyWait loop folding protocol
|
||
requests, signals, and notifications into callbacks. The harness answers the
|
||
universal zero-length ping and turns `terminate` into a clean exit for you
|
||
([process-lifecycle.md](process-lifecycle.md)).
|
||
- A driver spawned with an assignment (its device id as argv[1]) sends the
|
||
versioned `hello` to the device manager inside the deadline, and a **bus**
|
||
driver reports what it discovers with `child_added`
|
||
([device-manager.md](device-manager.md); usb-xhci-bus is the reference
|
||
implementation).
|
||
- Crash freely — that is the design. The kernel releases your claims, IRQ
|
||
bindings, and MSI vectors at death; the manager reads your exit reason,
|
||
prunes what you reported, restarts you with backoff, and your fresh instance
|
||
re-claims and re-reports. Never depend on your own cleanup running
|
||
(iron rule 1).
|