ipc_call and ipc_reply_wait grow a `send_cap` argument (r9) and a `received_cap` return (r8): an endpoint travels alongside a message, installed into the receiver's handle table. The transfer is a share, not a move — the endpoint's refcount is bumped and the sender keeps its handle. If the receiver's table is full the call fails -ENOSPC and the message is NOT delivered (a half-delivered capability is worse than a failed send); a bad handle fails -EBADF. Both directions carry a cap: a client's call hands one to the server (seen in the server's replyWait), and the server's reply hands one back (seen in the client's call return). This is the "open" primitive the driver model was blocked on: a bus driver mints a per-device endpoint and hands it to a class driver, giving it a private channel to one device without the 8-slot global name registry. Kernel: shareCapability in ipc-synchronous.zig at both copy points; new setSystemCallResult3 (r8, saved/restored by the syscall stub); Task gains ipc_send_cap / ipc_received_cap. Runtime: callCap + Reply, replyWait gains send_cap and Received.cap; plain call/replyWait delegate with no_cap. New abi.no_cap. New ipc-cap test (two kernel tasks exercise both directions, each verifying the endpoint it received is the same object shared, refcount bumped to 2). No class driver consumes callCap yet — it lands with the first one. Suite 37/37 plus host tests.
322 lines
16 KiB
Markdown
322 lines
16 KiB
Markdown
# The driver model: buses, classes, and host controllers
|
||
|
||
[drivers.md](drivers.md) shows how to write *a* driver — claim a device, map its
|
||
registers, sleep on its interrupt. That's enough for a leaf device like the HPET. It is
|
||
not enough for a disk, a keyboard, or a network card, because those hang off a
|
||
*controller*, on a *bus*, speaking a *protocol*, and no single process should have to
|
||
know all three.
|
||
|
||
Real driver stacks factor into three shapes. This document is about what each one is,
|
||
what the kernel must give it, how they share code — and precisely which primitive each
|
||
is still blocked on.
|
||
|
||
## Three shapes
|
||
|
||
| Shape | Owns | Reaches hardware by | Talks to |
|
||
|---|---|---|---|
|
||
| **Host controller driver** (HCD) | a controller — an xHCI PCI function, an AHCI port block | `mmio_map` + `irq_bind` + DMA | the devices behind it, in its bus's language |
|
||
| **Bus driver** | a bus — a PCI bridge, a USB hub | `device_register`, to publish what it finds | class drivers, over IPC |
|
||
| **Class / protocol driver** | *nothing* | *nothing* | its bus driver, over IPC |
|
||
|
||
The last row is the surprising one and the whole point. A USB keyboard driver touches
|
||
no registers, takes no interrupts, and maps no memory. It sends HID protocol messages
|
||
to whatever published the device, and it works identically whether the controller
|
||
below is xHCI, EHCI, or a Raspberry Pi's DWC2. That is what buys you drivers that
|
||
outlive the hardware they were written for.
|
||
|
||
In practice **HCD and bus driver are usually the same process**. An xHCI driver is a
|
||
host controller driver (it owns the PCI function, its BARs, its interrupt, its DMA
|
||
rings) *and* a bus driver (it enumerates USB devices and publishes them). Splitting
|
||
them is a fiction; what matters is that both *roles* have kernel support, because a
|
||
plain bus driver with no controller — a USB hub — is also a real thing.
|
||
|
||
## The device table is the spine
|
||
|
||
danos already has the right central structure. `system/kernel/devices-broker.zig` holds a table of
|
||
`DeviceDesc`, each with a parent, a class, and a set of resources. Firmware discovery
|
||
seeds it ([discovery.md](discovery.md)); `device_register` grows it.
|
||
|
||
Three invariants make it a capability system rather than a directory:
|
||
|
||
1. **A claim is exclusive.** `device_claim(id)` succeeds once. Everything downstream —
|
||
`mmio_map`, `irq_bind`, `device_register` — checks `devices_broker.ownerOf(id) == me`.
|
||
2. **A descriptor is a licence to map physical memory.** Whoever claims a device may
|
||
map its `.memory` resources and bind its `.irq` resources. This is why
|
||
`device_register` cannot be a free-for-all.
|
||
3. **Therefore: containment.** Every resource of a registered child must lie inside a
|
||
resource of the same kind on its parent (`devices_broker.contains`). A bus driver can only
|
||
ever *subdivide* what it already holds. Without this, `device_register` would be a
|
||
syscall named "map any physical page you like."
|
||
|
||
Containment is transitive by construction: a grandchild is contained in its child,
|
||
which is contained in the bus. Nothing can be laundered through a chain.
|
||
|
||
Note that firmware topology does **not** obey containment, and isn't asked to — a PCI
|
||
function's BAR is not inside its host bridge's `bus_range`, because a bus-number range
|
||
is not an address window. Discovery is trusted; user space is not.
|
||
|
||
### What a bus driver looks like
|
||
|
||
`system/drivers/bus/bus.zig` is the smallest honest one. Its "bus" is the HPET's register block and
|
||
its "devices" are the block's comparators:
|
||
|
||
```zig
|
||
_ = dev.claim(bus.id); // 1. own the bus
|
||
const base = dev.mmioMap(bus.id, 0).?; // 2. enumerate it — from the hardware
|
||
const n = ((cap.* >> 8) & 0x1F) + 1; // GENERAL_CAP says how many children
|
||
|
||
for (0..n) |i| { // 3. publish each child
|
||
var child = std.mem.zeroes(dev.DeviceDesc);
|
||
child.class = @intFromEnum(dev.DeviceClass.timer);
|
||
child.resource_count = 1;
|
||
child.resources[0] = .{ .kind = memory,
|
||
.start = bus_mmio.start + 0x100 + 0x20 * i,
|
||
.len = 0x20 };
|
||
_ = dev.register(bus.id, &child).?; // kernel checks containment
|
||
}
|
||
```
|
||
|
||
Each child is left **unclaimed**, which is the handoff: a comparator driver can now
|
||
`device_claim` one and `mmio_map` it, and will see only its own 0x20-byte window. A child
|
||
whose window escapes the bus is refused — `bus` asserts that, and the `bus` test
|
||
asserts the kernel's table upholds it.
|
||
|
||
A USB device has *no* resources at all: `resource_count = 0`, because it's addressed
|
||
through its controller, not by MMIO. That case is allowed and is the common one.
|
||
|
||
## Families: sharing code between drivers
|
||
|
||
A "family" is two modules, not one:
|
||
|
||
- **A logic module** — the parts of the bus that every driver on it re-derives. Config
|
||
space walking and BAR decode for PCI. Descriptor parsing, control transfers, and hub
|
||
protocol for USB.
|
||
- **A protocol module** — the IPC message types that let a class driver talk to
|
||
*whatever* published its device. This is the part that makes class drivers portable.
|
||
|
||
danos already has one of each: `library/runtime/device.zig` is a logic module,
|
||
[`system/services/vfs/protocol.zig`](system/services/vfs/protocol.zig) is a protocol module shared by `system/services/vfs/vfs.zig`
|
||
and its clients. The pattern generalises directly:
|
||
|
||
```
|
||
library/
|
||
runtime/ module "runtime" — syscalls, heap, ipc, device, stdio
|
||
mmio/ module "mmio" — volatile register access + barriers [M14]
|
||
bus/
|
||
pci/ module "pci" — ECAM, BAR decode, capability walk
|
||
usb/ module "usb" — descriptors, control transfers, hubs
|
||
proto/
|
||
vfs/ module "vfs-protocol" (today: system/services/vfs/protocol.zig)
|
||
block/ module "block-protocol"
|
||
hid/ module "hid-protocol"
|
||
|
||
system/drivers/ one sub-project each → /system/drivers (no `d` suffix)
|
||
xhci/ HCD + bus driver imports runtime, pci, usb, mmio
|
||
usb-hid/ class driver imports runtime, usb, hid-protocol
|
||
block/ class driver imports runtime, block-protocol
|
||
```
|
||
|
||
The only build change needed: [`addUserBinary`](build.zig) currently takes exactly one
|
||
module (`rt_mod`) and injects it. It should take a slice of modules. That's a
|
||
five-line change, and it's the *entire* mechanism — Zig modules already give you
|
||
everything else.
|
||
|
||
The discipline that makes this work: **a class driver must not import a bus's logic
|
||
module.** `usbhid` imports `proto.hid` and `usb` (for descriptor types), never `pci`.
|
||
If a class driver needs `mmio`, it has become an HCD and should be one.
|
||
|
||
## What exists today
|
||
|
||
- **M10** — `device_enumerate`, `device_claim`, `mmio_map`. Strong-uncacheable device
|
||
grants, `device_grant` teardown.
|
||
- **M11** — `irq_bind` / `irq_ack`. IRQ delivered as an IPC notification; mask before
|
||
EOI; `irq_ack` is the unmask.
|
||
- **M12** — `parent` in `DeviceDesc`, `device_register` with resource containment.
|
||
- **M13** — capability passing. `ipc_call` / `ipc_reply_wait` grew a `send_cap` argument
|
||
and a `received_cap` return (r8): an endpoint travels with a message, installed into
|
||
the receiver's handle table (shared, refcount-bumped — a copy, not a move). A full
|
||
table fails `-ENOSPC` and does not half-deliver. This is the "open" primitive — a bus
|
||
driver mints a per-device endpoint and hands it to a class driver. The runtime exposes
|
||
`callCap` and `replyWait(..., send_cap)`; no class driver consumes it yet.
|
||
- **`system_spawn`** — a user-space supervisor starts a driver: `system_spawn(name)`
|
||
loads a binary bundled in the initial-ramdisk as a fresh ring-3 process. This is what
|
||
turned the device manager from "log the match" into "run the driver": the kernel now
|
||
spawns only `init`, `init` spawns the services, and the **device-manager** discovers
|
||
the hardware and spawns each driver ([drivers.md](drivers.md)). Ungated for now — a
|
||
spawn capability is future work.
|
||
|
||
So: **bus drivers work now, and they're started by the device manager, not the kernel.**
|
||
HCDs and class drivers do not work yet. Here is exactly why, and exactly what would fix it.
|
||
|
||
---
|
||
|
||
# Proposed ABI
|
||
|
||
## M13 — capability passing, for class drivers ✅ done
|
||
|
||
*Implemented as described below (see "What exists today"). The signatures landed
|
||
verbatim: `send_cap` in r9, `received_cap` returned in r8, `-ENOSPC` on a full receiver
|
||
table with no delivery. The rest of this section is the original design note.*
|
||
|
||
**The blocker.** A class driver has to reach *its* device. Today the only way to find
|
||
an endpoint is the name registry: `ipc_register(service_id, h)` / `ipc_lookup(id)`,
|
||
where `ServiceId` is a global integer namespace with `max_services = 8`. You cannot
|
||
mint one endpoint per USB device that way, and there is no way for a bus driver to
|
||
*hand* a class driver an endpoint. M7 deferred this deliberately.
|
||
|
||
**The fix.** Let a message carry one handle. Sender names a handle in its own table;
|
||
the kernel installs the endpoint into the receiver's table (bumping `refcount`) and
|
||
tells the receiver the index it landed at.
|
||
|
||
```
|
||
ipc_call(h, msg, message_len, reply, reply_cap, send_cap) -> reply_len
|
||
ipc_reply_wait(h, reply, reply_len, recv, recv_cap, send_cap)
|
||
-> recv_len (rax), badge (rdx), received_cap (r8)
|
||
```
|
||
|
||
`send_cap` is a handle or `no_cap` (`~0`). `received_cap` is the index the transferred
|
||
endpoint was installed at in the receiver's table, or `no_cap`.
|
||
|
||
- Both calls grow from 5 args to 6, which fits: `syscall5` uses `rdi/rsi/rdx/r10/r8`,
|
||
leaving `r9`. `ipc_reply_wait` already returns two values via `setSyscallResult2`;
|
||
this needs a third (`setSyscallResult3`).
|
||
- If the receiver's handle table is full, the call fails `-ENOSPC` and **the message is
|
||
not delivered** — a half-delivered capability is worse than a failed send.
|
||
- `closeHandles` already drops references on exit, so the lifetime story is unchanged.
|
||
|
||
That single primitive gives you the standard `open` pattern:
|
||
|
||
```zig
|
||
// class driver // bus driver
|
||
const h = ipc.lookup(.usb).?; const r = ipc.replyWait(ep, ...);
|
||
const dev_ep = ipc.callCap(h, // ... mint a per-device endpoint,
|
||
.{ .op = .open, .id = dev_id }); // reply with it as send_cap
|
||
// now dev_ep is a private channel to that one device
|
||
```
|
||
|
||
## M14 — DMA memory and the memory-ordering contract, for HCDs
|
||
|
||
**The blocker.** An HCD is a DMA-engine programmer. It needs a descriptor ring the
|
||
device can read, which means memory that is (a) physically contiguous, (b) at a
|
||
physical address the driver knows, (c) of the right cacheability, and (d) pinned.
|
||
[`sysMmap`](system/kernel/process.zig) gives you *none* of the four: it calls `pmm.alloc()`
|
||
once per page, maps writeback-cached, and never reveals a physical address.
|
||
|
||
**The fix.**
|
||
|
||
```
|
||
dma_alloc(len, flags) -> vaddr (rax), paddr (rdx)
|
||
dma_free(vaddr, len) -> 0
|
||
|
||
flags: dma_coherent (1) uncacheable; the default and the only one that's portable
|
||
dma_wc (2) write-combining — needs PAT programmed; for framebuffers
|
||
dma_below_4g (4) for devices with 32-bit DMA addressing
|
||
```
|
||
|
||
Guarantees: page-aligned, physically contiguous, zeroed, pinned for the life of the
|
||
mapping, and the physical address is stable. It needs one thing the kernel lacks —
|
||
`pmm.allocContiguous(n, max_phys)`; today `pmm.alloc()` hands out one frame at a time
|
||
with no adjacency guarantee.
|
||
|
||
**The memory-ordering contract.** danos has, at the time of writing, **zero memory
|
||
barriers anywhere in the tree.** That is currently correct-by-accident and won't
|
||
survive the first DMA driver, or the first ARM boot.
|
||
|
||
`volatile` is not a barrier. In Zig it means: don't elide this access, and don't
|
||
reorder it against *other volatile* accesses. It says nothing about your *ordinary*
|
||
stores — the descriptor you just filled in normal WB memory — which LLVM may freely
|
||
sink past a volatile MMIO write. The canonical bug:
|
||
|
||
```zig
|
||
ring[i] = descriptor; // ordinary store to WB RAM
|
||
doorbell.* = i; // volatile store to UC MMIO
|
||
// nothing stops the compiler reordering these; the device reads a stale descriptor
|
||
```
|
||
|
||
So the rules, which belong in `library/mmio.zig` and behind `arch`:
|
||
|
||
| Situation | Required |
|
||
|---|---|
|
||
| MMIO register read/write | `mmio.read` / `mmio.write` (volatile) |
|
||
| Fill DMA descriptor, then ring doorbell | `wmb()` between them |
|
||
| Woken by IRQ, then read what the device wrote | `rmb()` before the read |
|
||
| MMIO write that must complete before the next read | `mb()` |
|
||
|
||
And the per-arch lowering — the reason this must be an `arch` primitive and not a
|
||
sprinkling of `asm volatile`:
|
||
|
||
| | x86_64 | aarch64 |
|
||
|---|---|---|
|
||
| `mb()` | `mfence` | `dsb sy` |
|
||
| `rmb()` | `lfence` | `dsb ld` |
|
||
| `wmb()` | `sfence` | `dsb st` |
|
||
| DMA cache coherency | coherent; nothing to do | **not guaranteed**; needs non-cacheable buffers or cache maintenance |
|
||
|
||
x86 is forgiving here — TSO plus strong-uncacheable MMIO means you usually get away
|
||
with a compiler barrier alone. ARM is not, and [vision.md](vision.md) makes ARM the win
|
||
condition. Build the abstraction while there is one caller to fix.
|
||
|
||
(Zig note: `@fence` was **removed in 0.16**. Use `@atomicRmw(..., .seq_cst)` for a full
|
||
barrier, or per-arch inline asm — which is what `library/mmio.zig` should hide.)
|
||
|
||
## M15 — interrupts for PCI devices
|
||
|
||
**The blocker, and it's a hard one.** No PCI device can take an interrupt today.
|
||
[`addBars`](system/devices/acpi.zig) records `.memory` and `.io_port` BARs and never an
|
||
`.irq`; there is no `_PRT` parsing anywhere in the tree. `hpet` only works because the
|
||
HPET advertises its own routing options in its own registers — a privilege no ordinary
|
||
device has.
|
||
|
||
**The fix, in two halves.**
|
||
|
||
*Legacy INTx*: parse `_PRT` from the DSDT to map (device, INTA–D) → GSI, and record it
|
||
as an `.irq` resource. Then `irq_bind` works unchanged. But INTx lines are **shared**,
|
||
and `irq.bound[gsi]` holds one endpoint. Sharing needs a list, and every driver on the
|
||
line must be polled on each interrupt — the reason everyone left INTx behind.
|
||
|
||
*MSI/MSI-X*, which is the real answer: per-device vectors, edge-triggered, unshared, no
|
||
mask/ack cycle, no 24-GSI ceiling. The kernel allocates a vector and hands the driver
|
||
the (address, data) pair to program into its own MSI capability:
|
||
|
||
```
|
||
msi_bind(dev_id, endpoint, out) -> 0 // out: extern struct { addr: u64, data: u32 }
|
||
```
|
||
|
||
The driver writes those into config space itself — which means it needs config space,
|
||
which means **discovery should give each `pci_device` a `.memory` resource for its
|
||
4 KiB ECAM slot**. That's a small change to `parseMcfg` and it unblocks the whole
|
||
capability walk (MSI, MSI-X, PCIe extended caps) without any new syscall.
|
||
|
||
Note QEMU's HPET reports `Tn_FSB_INT_DEL_CAP = 0` — no MSI — so `hpet` can never
|
||
exercise this path. The first MSI driver will be the first PCI driver.
|
||
|
||
## M16 — the IOMMU, and the honest caveat
|
||
|
||
Everything above is capability-gated at the *CPU*. None of it is gated at the *device*.
|
||
A driver that can program a bus-mastering engine can make that device write to any
|
||
physical address, because page tables sit between the CPU and RAM, not between a device
|
||
and RAM. Until VT-d/DMAR (or SMMU on ARM) is programmed from the DMAR table, **`device_claim`
|
||
on any DMA-capable device is equivalent to granting ring 0.**
|
||
|
||
This does not make the model useless — it's the same position Linux is in with the
|
||
IOMMU off, and every other guarantee (crash isolation, restart, no shared address
|
||
space) still holds. But "user-space drivers are memory-safe" is not true yet, and the
|
||
gap should be named rather than implied.
|
||
|
||
## Ordering
|
||
|
||
`M13` (capability passing) is independent of `M14`/`M15` and is the cheapest. It
|
||
unlocks class drivers, which are the shape with no hardware requirements at all — you
|
||
could write a real one against `bus`'s comparators tomorrow.
|
||
|
||
`M14` and `M15` together unlock the first HCD. `M14`'s barrier layer is worth landing
|
||
on its own regardless: it's small, obviously correct, and stops every future driver
|
||
from hand-rolling `*volatile` and getting ARM wrong.
|
||
|
||
## See also
|
||
|
||
- [drivers.md](drivers.md) — how to write one, concretely.
|
||
- [discovery.md](discovery.md) / [acpi.md](acpi.md) — where the device table comes from.
|
||
- [ipc.md](ipc.md) — endpoints, badges, and the notification path an IRQ arrives on.
|
||
- [resilience.md](resilience.md) — restart, the reason any of this is worth the trouble.
|