//! Device service: the kernel side of user-space driver access. At boot it //! flattens the discovered device tree (src/device) into a stable, id-indexed //! snapshot and a per-device claim table. User drivers enumerate the snapshot, //! claim the device they own, and map its MMIO — the claim is the capability that //! gates `mmio_map`/`irq_bind`, so a process can only ever touch hardware the //! firmware-neutral device tree says it owns. //! //! The table is a **tree**: each entry carries its parent's id. Firmware discovery //! seeds it, and a **bus driver** grows it — a process that has claimed a bus can //! `register` children below it as it enumerates them (USB devices behind a hub, PCI //! functions behind a bridge, comparators inside a timer block). //! //! Registration is where the capability model earns its keep. A `DeviceDescriptor` is, in //! effect, a licence to map physical memory: whoever claims it may `mmio_map` its //! `.memory` resources and `irq_bind` its `.irq` resources. If a bus driver could //! invent arbitrary resources, it would invent one covering the kernel's RAM, claim //! it, and map it. So `register` enforces **containment**: every resource of a child //! must lie inside a resource of the same kind on its parent. A bus driver can only //! ever subdivide what it was already given. const std = @import("std"); const abi = @import("abi"); const platform = @import("platform"); const device_abi = @import("device-abi"); const heap = @import("heap.zig"); /// How many devices a single registrar may put in the table. /// /// **This is a runaway detector, not a security boundary**, and the difference matters. /// It cannot stop a malicious driver — a quota generous enough never to bite a real /// machine is still generous enough to be unpleasant — and it is not trying to. What /// stops malice is that only a driver the manager handed a device can register children /// under it (docs/os-development/device-authority.md). What this catches is a /// *legitimate* driver in a loop, early, and attributably: the driver that did it is /// refused, named in its own log, and restarted, while every other driver is untouched. /// /// The shared ceilings it replaces could not do that. `maximum_devices = 64` was a /// guess about someone else's computer and one driver's enumeration starved every /// other — which is how an AMD Ryzen came to boot with a working display, no USB and no /// storage. A per-registrar allowance is the same protection charged to whoever caused /// it, which is the microkernel property rather than a workaround for it. /// /// The number: a machine's whole PCI segment tops out at 65536 functions, and the /// biggest real registrar seen is pci-bus at a few dozen. 4096 is far above anything a /// real machine produces and around 1.4 MiB of descriptors, well under the kernel heap. /// **Reaching it is a bug report, not a tuning request** — no correct driver gets near. /// /// bound: devices one task may register /// decided-by: ours /// protects: the kernel heap, against a driver looping device_register /// at-limit: refuse — ECHILDREN to the registrar; every other driver is unaffected /// observed-by: the bus driver's line naming the reason (pci-bus reconciles functions /// found against registered) const maximum_devices_per_registrar = 4096; /// The table, grown on demand from the kernel heap. **There is no ceiling**: how many /// devices a machine has is the machine's business, and no specification bounds it, so /// nothing here should. `devices_broker.init` runs after `heap.init` (kernel.zig), so /// there was never a reason for this to be static beyond it having been written that /// way first. var devices: []device_abi.DeviceDescriptor = &.{}; /// Owner task id per device, or null. Parallel to `devices` and grown with it. var claimed: []?u32 = &.{}; /// The task that called `register` for each device, so the per-registrar allowance can /// be charged to whoever caused the entry. Firmware-discovered nodes carry `no_registrar` /// — they are the kernel's own, not anybody's doing. /// Optional, not a sentinel: **task 0 is a real task** (the kernel's own), so any /// "none" value inside the id space is a device belonging to somebody reading back as /// belonging to nobody. Found the moment the first test asserted a giver, because the /// task doing the giving was task 0. var registrar: []?u32 = &.{}; /// **Who gave each device away**, or `no_giver` if nobody ever did. /// /// One field, and the whole authority rule follows from it: a device that was *given* /// to someone is delegated hardware, so it may only be handed on, never taken /// (`claim` refuses it); and when its holder dies it goes back to whoever lent it, /// rather than becoming free for anyone to grab. /// /// It also settles the framebuffer without mentioning it. Nobody delegates the /// loader's framebuffer, so it has no giver, so the display service claims it exactly /// as it always has — no exemption, no special case, no `display` anywhere in the rule. var giver: []?u32 = &.{}; var count: usize = 0; /// Grow the three parallel arrays so at least one more device fits. False if the heap /// cannot satisfy it, which the callers report rather than swallow. fn reserve() bool { if (count < devices.len) return true; // Double from a deliberately SMALL first block. Sizing it for a typical machine // would mean the growth path never ran on the hardware we test on, and only woke up // on someone else's larger machine — which is the exact shape of the failure this // whole track exists to stop. At 8, every boot grows the table several times, so // the path is exercised constantly and the suite asserts it. const wanted = if (devices.len == 0) 8 else devices.len * 2; const allocator = heap.allocator(); const grown_devices = allocator.realloc(devices, wanted) catch return false; devices = grown_devices; const grown_claimed = allocator.realloc(claimed, wanted) catch return false; claimed = grown_claimed; const grown_registrar = allocator.realloc(registrar, wanted) catch return false; registrar = grown_registrar; const grown_giver = allocator.realloc(giver, wanted) catch return false; giver = grown_giver; for (claimed[count..], registrar[count..], giver[count..]) |*slot, *who, *lender| { slot.* = null; who.* = null; lender.* = null; } return true; } /// How many devices `task` has registered — the allowance is charged per registrar, so /// a driver in a loop exhausts its own and no one else's. fn registeredBy(task: u32) usize { var n: usize = 0; for (registrar[0..count]) |who| { if (who != null and who.? == task) n += 1; } return n; } /// The id of the seeded framebuffer node (`seedDisplay`), or null when the machine /// handed over no framebuffer. Lets the process layer recognise the display claim /// (to quiesce the bootstrap console) without threading the id through every caller. var display_device: ?u64 = null; /// Devices discovery found but could not record. The table grows on demand, so this is /// no longer "the machine is bigger than our guess" — it means the kernel heap could not /// satisfy the growth, which would otherwise be an entirely silent failure. Logged at /// boot. pub var dropped: usize = 0; /// Snapshot the device tree into the flat table. Run once, right after discovery. pub fn init(device_tree: *const platform.DeviceTree) void { count = 0; dropped = 0; display_device = null; for (claimed) |*c| c.* = null; for (registrar) |*r| r.* = null; for (giver) |*g| g.* = null; walk(device_tree.root, device_abi.no_parent); } /// Publish the loader's framebuffer as a `display` device — a root-level node with one /// write-combining `memory` resource over the linear framebuffer and its geometry in /// `.display`. The framebuffer is *not* firmware-discovered (it rides the /// [[boot-handoff]], not the device tree), so it is seeded explicitly, after `init`. /// Returns the new device id, or null when there is no framebuffer (headless) or the /// table is full. Idempotent-ish: only ever call once per boot. pub fn seedDisplay(base: u64, width: u32, height: u32, pitch: u32, format: u32, refresh_hz: u32) ?u64 { if (base == 0 or width == 0 or height == 0) return null; // headless if (!reserve()) { dropped += 1; return null; } var d = std.mem.zeroes(device_abi.DeviceDescriptor); d.id = count; d.parent = device_abi.no_parent; d.class = @intFromEnum(device_abi.DeviceClass.display); d.pci_class = device_abi.no_pci_class; d.resource_count = 1; d.resources[0] = .{ .kind = @intFromEnum(device_abi.ResourceKind.memory), .start = base, .len = @as(u64, height) * pitch, .flags = device_abi.resource_flag_write_combining, }; d.display = .{ .width = width, .height = height, .pitch = pitch, .format = format, .refresh_hz = refresh_hz }; devices[count] = d; display_device = d.id; count += 1; return d.id; } /// The id of the seeded framebuffer device, or null when none was seeded. pub fn displayDevice() ?u64 { return display_device; } /// Whether the framebuffer device is currently claimed by some process. The bootstrap /// console uses this (via the process layer) to fall silent while a display service /// owns the screen, and to resume if that service dies and its claim is released. pub fn displayClaimed() bool { const id = display_device orelse return false; return ownerOf(id) != null; } /// Record `node` (unless it's the synthetic root) and recurse, threading the id we /// assigned it down to its children as their parent. fn walk(node: *platform.Device, parent_id: u64) void { const id = if (node.class == .root) device_abi.no_parent else record(node, parent_id); var child = node.first_child; while (child) |c| : (child = c.next_sibling) walk(c, id); } fn record(node: *platform.Device, parent_id: u64) u64 { if (!reserve()) { dropped += 1; return device_abi.no_parent; // children of a dropped node become roots, not orphans } var d = std.mem.zeroes(device_abi.DeviceDescriptor); d.id = count; d.parent = parent_id; d.class = @intFromEnum(node.class); d.pci_class = if (node.ids.pci_class) |code| code else device_abi.no_pci_class; const h = node.hid(); d.hid_len = @min(h.len, d.hid.len); @memcpy(d.hid[0..d.hid_len], h[0..d.hid_len]); const rc = @min(node.resource_count, device_abi.maximum_device_resources); d.resource_count = rc; for (0..rc) |i| { const r = node.resources[i]; d.resources[i] = .{ .kind = @intFromEnum(r.kind), .start = r.start, .len = r.len }; } devices[count] = d; count += 1; return d.id; } /// Copy up to `out.len` device descriptors into `out`; returns the total count /// available (which may exceed `out.len`). For kernel callers with a buffer big /// enough to take the whole table in one go. pub fn enumerate(out: []device_abi.DeviceDescriptor) usize { _ = enumerateFrom(0, out); return count; } /// How many devices the table holds — the total `device_enumerate` reports back /// however few of them fit in the caller's buffer. pub fn deviceCount() usize { return count; } /// Copy up to `out.len` descriptors starting at table index `start`, returning how /// many were filled (0 once `start` reaches the end). The chunked form: the /// `device_enumerate` system call bounces the table out through a small kernel /// buffer, one chunk at a time, because a descriptor is far too big to stage a /// whole user-requested array of them on a 16 KiB kernel stack. pub fn enumerateFrom(start: usize, out: []device_abi.DeviceDescriptor) usize { if (start >= count) return 0; const n = @min(count - start, out.len); @memcpy(out[0..n], devices[start..][0..n]); return n; } /// Take exclusive ownership of device `id` for task `owner`. The two ways this can /// fail want different responses from a driver — a stale id means re-enumerate, a /// live claimant means back off — so they are distinguishable (`ClaimError`). pub fn claim(id: u64, owner: u32) ClaimError!void { if (id >= count) return error.NoSuchDevice; if (claimed[@intCast(id)] != null) return error.AlreadyClaimed; // **Delegated hardware may be handed on, never taken.** // // Belt and braces, and worth being honest about: with the loan rule above this is // **currently unreachable**. A device that was given to someone is held, so it is // refused as `AlreadyClaimed` before reaching here; and when the holder dies the // device goes back to its lender (or, if the lender is gone, has its giver cleared // with its claim), so there is no state where a device is unheld *and* still on // loan. The window a stranger could have used simply stops existing. // // It stays because it is one comparison and it fails closed: any future path that // frees a device without clearing its giver would otherwise hand delegated // hardware to whoever asked first, which is exactly the hole this run closed. // // Note it leaves the loader's framebuffer alone without naming it: nobody delegates // the framebuffer, so it has no giver, so the display service claims it as always. if (giver[@intCast(id)] != null) return error.NotYours; claimed[@intCast(id)] = owner; } /// The task that gave device `id` away, or null if nobody ever did. A device with a /// giver is delegated hardware: it may be handed on, never taken. pub fn giverOf(id: u64) ?u32 { if (id >= count) return null; return giver[@intCast(id)]; } /// The task that owns device `id`, or null. pub fn ownerOf(id: u64) ?u32 { if (id >= count) return null; return claimed[@intCast(id)]; } /// Release every claim held by `owner` — called by the process layer on every /// path out of a process (exit, fault, kill), so a restarted driver can claim its /// hardware again (docs/process-lifecycle.md iron rule 1: cleanup is the kernel's /// job). The devices stay in the table — they describe hardware, which did not go /// away — only their ownership clears. /// Whether a task is still alive, injected by the process layer (which owns the task /// table) the same way the scheduler's other hooks are. Null means "assume not", so a /// kernel built without it clears claims rather than handing them to a ghost. pub var task_alive_hook: ?*const fn (u32) bool = null; fn alive(task: u32) bool { const hook = task_alive_hook orelse return false; return hook(task); } pub fn releaseAllOwnedBy(owner: u32) void { for (claimed[0..count], 0..) |*slot, id| { const holder = slot.* orelse continue; if (holder != owner) continue; // **A grant is a loan.** A device this task was *given* goes back to whoever // lent it, not to nobody — so the device manager gets its hardware back the // instant a driver dies, and hands it to the replacement. // // Without this the kernel released the claim to no one and the manager // re-claimed first-come, so every driver restart reopened the window this // rule closes. And once `claim` refuses a device that has a giver, releasing // to nobody would strand it: no one could ever take it again. // // A dead lender is no lender: clear the claim and the giver together, so the // device is genuinely free rather than owed to a ghost. if (giver[id]) |lender| { if (alive(lender)) { slot.* = lender; giver[id] = null; // returned; it is the lender's own again, not on loan continue; } giver[id] = null; } slot.* = null; } } /// Release the claim on `id` iff `owner` holds it — the rollback for a claim that /// cannot be confined (the IOMMU domain could not be created/attached). Returns true /// when a claim was actually cleared. pub fn unclaim(id: u64, owner: u32) bool { if (id >= count) return false; if (claimed[@intCast(id)]) |o| { if (o == owner) { claimed[@intCast(id)] = null; return true; } } return false; } /// Resource `index` of device `id`, or null if out of range. pub fn resourceOf(id: u64, index: u64) ?device_abi.ResourceDescriptor { if (id >= count) return null; const d = &devices[@intCast(id)]; if (index >= d.resource_count) return null; return d.resources[@intCast(index)]; } /// The PCI requester id (bus<<8 | device<<3 | function) of device `id`, derived from /// its config-space slice against its host bridge's ECAM window — the identity a VT-d /// context entry / AMD-Vi DTE is keyed by. null when `id` is not a PCI function or the /// geometry doesn't decode. The kernel never stored the BDF (the descriptor has no such /// field); pci-bus encodes it into resource 0's physical base as /// `ecam_base + ((bus - start_bus) << 20 | device << 15 | function << 12)`, and the /// requester id the device emits uses the absolute bus, so we add `start_bus << 8` back. pub fn pciAddressOf(id: u64) ?u16 { if (id >= count) return null; const d = &devices[@intCast(id)]; if (d.class != @intFromEnum(device_abi.DeviceClass.pci_device)) return null; if (d.resource_count == 0) return null; const config = d.resources[0]; if (config.kind != @intFromEnum(device_abi.ResourceKind.memory) or config.len != 4096) return null; // Walk up to the host bridge, whose resource 0 is the segment's ECAM window and // resource 1 the bus_range (start_bus, bus_count). var parent = d.parent; while (parent != device_abi.no_parent and parent < count) { const p = &devices[@intCast(parent)]; if (p.class == @intFromEnum(device_abi.DeviceClass.pci_host_bridge)) { if (p.resource_count < 2) return null; const ecam = p.resources[0]; const bus_range = p.resources[1]; if (config.start < ecam.start or config.start >= ecam.start + ecam.len) return null; const offset = config.start - ecam.start; const start_bus: u16 = @intCast(bus_range.start & 0xFF); return @intCast((offset >> 12) + (@as(u64, start_bus) << 8)); } parent = p.parent; } return null; } /// Call `visit(id, bdf)` for every PCI function in the table — the IOMMU core's boot /// sweep to place every device under a domain. Only functions whose BDF decodes are /// visited. pub fn forEachPciFunction(visit: *const fn (id: u64, bdf: u16) void) void { var id: u64 = 0; while (id < count) : (id += 1) { if (pciAddressOf(id)) |bdf| visit(id, bdf); } } /// Is `child` wholly inside `parent`? For a range (memory, io_port, bus_range) that's /// interval containment; for an irq it's equality, since an interrupt line is not /// divisible. Zero-length child ranges are refused — an empty window is meaningless /// and would otherwise vacuously "fit" anywhere. fn contains(parent: device_abi.ResourceDescriptor, child: device_abi.ResourceDescriptor) bool { if (parent.kind != child.kind) return false; if (child.kind == @intFromEnum(device_abi.ResourceKind.irq)) { // Range containment: an interrupt line is still indivisible (a child owns // exactly one GSI), but a parent may own a *range* of lines so a broad // owner — the acpi-tables node, whose firmware names any legacy IRQ — // can contain its children's specific lines. A length-1 parent range is // exactly the old equality rule, so existing single-IRQ parents are // unaffected. const span = if (parent.len == 0) 1 else parent.len; return child.start >= parent.start and child.start < parent.start + span; } if (child.len == 0 or parent.len == 0) return false; // No overflow: a resource that wraps the address space is not containable. const child_end = std.math.add(u64, child.start, child.len) catch return false; const parent_end = std.math.add(u64, parent.start, parent.len) catch return false; return child.start >= parent.start and child_end <= parent_end; } /// Why a `register` was refused. Each variant maps to its own errno (`errnoOf`), so /// a bus driver's log line can name the rule that stopped it — "this parent is at /// its child cap" and "the table is full" want different fixes, and telling them /// apart from a bare -1 cost a debugging session (docs/fixed-bounds-audit.md). pub const RegisterError = error{ NoSpace, // the device table is full NoSuchParent, // no device with that id NotYourParent, // that device exists but this task has not claimed it TooManyResources, // the descriptor declares more resources than one device may hold TooManyChildren, // the caller is at its per-registrar allowance NotContained, // a child resource escapes its parent's window }; /// The errno a refused `register` returns to ring 3. pub fn errnoOf(e: RegisterError) i64 { return switch (e) { error.NoSpace => abi.ENOSPC, error.NoSuchParent => abi.ENODEV, error.NotYourParent => abi.EPERM, error.TooManyResources => abi.E2BIG, error.TooManyChildren => abi.ECHILDREN, error.NotContained => abi.ERANGE, }; } /// Why a `claim` was refused. pub const ClaimError = error{ NoSuchDevice, // no device with that id AlreadyClaimed, // a live task already owns it NotYours, // delegated hardware: it has a giver, so it must be handed on, not taken }; /// Why a `transfer` was refused. pub const TransferError = error{ NoSuchDevice, // no device with that id NotHeld, // the caller does not hold it — you may only give away what you have }; /// The errno a refused `transfer` returns to ring 3. (`ESRCH` — no such recipient — is /// raised by the caller in system/kernel/process.zig, which is what can see the task /// table.) pub fn transferErrnoOf(e: TransferError) i64 { return switch (e) { error.NoSuchDevice => abi.ENODEV, error.NotHeld => abi.EPERM, }; } /// Move device `id` from `from` to `to`. **A move, not a copy**: a claim is exclusive /// (driver-model.md, invariant 1), so the giver stops holding it the moment the /// receiver starts. /// /// This is the mechanism behind delegation — the device manager claims what firmware /// discovery seeded and passes each device to the driver it matched, which replaces /// first-come-first-served `device_claim` with policy /// (docs/device-driver-development/device-manager.md). The kernel checks only that the /// caller holds the device: *you may give away what you have*. It knows nothing about /// which task is the manager, and needs to know nothing. /// /// Note this is deliberately NOT the M13 capability-passing path, which shares a handle /// refcounted — a copy. Exclusivity cannot be expressed that way. pub fn transfer(id: u64, from: u32, to: u32) TransferError!void { try canTransfer(id, from); claimed[@intCast(id)] = to; giver[@intCast(id)] = from; } /// The checks `transfer` will make, without the move. The syscall layer runs them /// first — under the same lock hold that the transfer itself will run under — so it /// can refuse, or arrange the IOMMU confinement the move needs, while nothing has /// mutated yet and there is nothing to roll back. pub fn canTransfer(id: u64, from: u32) TransferError!void { if (id >= count) return error.NoSuchDevice; const holder = claimed[@intCast(id)] orelse return error.NotHeld; if (holder != from) return error.NotHeld; } /// The errno a refused `claim` returns to ring 3. (`ECONFINE` — the claim stood but /// the IOMMU would not confine the device — is raised by the caller in /// system/kernel/process.zig, which is what rolls the claim back.) pub fn claimErrnoOf(e: ClaimError) i64 { return switch (e) { error.NoSuchDevice => abi.ENODEV, error.AlreadyClaimed => abi.EBUSY, error.NotYours => abi.EPERM, }; } /// The id of a child of `parent_id` already identical to `descriptor`, or null. /// Exact on class, identity and every resource — anything less would let a bus /// silently adopt an entry that is not the device it just found. The caller must /// have bounded `descriptor.resource_count` first. fn existingChild(parent_id: u64, descriptor: *const device_abi.DeviceDescriptor) ?u64 { for (devices[0..count]) |*existing| { if (existing.parent != parent_id) continue; if (existing.class != descriptor.class) continue; if (existing.pci_class != descriptor.pci_class) continue; if (existing.hid_len != descriptor.hid_len) continue; if (!std.mem.eql(u8, existing.hid[0..@intCast(existing.hid_len)], descriptor.hid[0..@intCast(descriptor.hid_len)])) continue; if (existing.resource_count != descriptor.resource_count) continue; var same = true; for (0..@intCast(descriptor.resource_count)) |i| { const a = existing.resources[i]; const b = descriptor.resources[i]; if (a.kind != b.kind or a.start != b.start or a.len != b.len) same = false; } if (same) return existing.id; } return null; } /// Publish `descriptor` as a child of `parent_id`, on behalf of `owner`. Returns the new /// device id. The child is left **unclaimed**, so another process (a class driver) /// can claim it — that is how a bus hands a device to its driver. /// /// `owner` must have claimed `parent_id`, and every resource in `descriptor` must be /// contained in a parent resource of the same kind. A device with no resources is /// fine and common: a USB device is addressed through its controller, not by MMIO. pub fn register(parent_id: u64, owner: u32, descriptor: *const device_abi.DeviceDescriptor) RegisterError!u64 { if (parent_id >= count) return error.NoSuchParent; const parent_owner = ownerOf(parent_id) orelse return error.NotYourParent; if (parent_owner != owner) return error.NotYourParent; // Bounds every `descriptor.resources` read below, including the match scan's. if (descriptor.resource_count > device_abi.maximum_device_resources) return error.TooManyResources; // Idempotent on exact match (docs/device-manager.md): a restarted registering // bus re-registers what it rediscovers, and the table has no unregister — an // identical child under the same parent returns the existing id instead of // appending a duplicate. // // Checked **before the caps**, because a re-registration consumes no slot. // Charging it against the child cap refused a restarted bus its own devices the // second time it started, which turned the supervision restart this system leans // on into a one-way ratchet toward a degraded machine. The match is exact — class, // identity, and every resource — so an entry returned this way was contained when // it was first admitted, and a stored descriptor's resources never change // afterwards (they are written only by `record`, `seedDisplay` and the append // below). if (existingChild(parent_id, descriptor)) |existing_id| return existing_id; // The allowance is charged to whoever is registering, so a driver in a loop // exhausts its own and every other driver carries on. There is no machine-wide // ceiling any more: the table grows. if (registeredBy(owner) >= maximum_devices_per_registrar) return error.TooManyChildren; if (!reserve()) return error.NoSpace; const parent = &devices[@intCast(parent_id)]; for (0..@intCast(descriptor.resource_count)) |i| { const r = descriptor.resources[i]; var ok = false; for (0..@intCast(parent.resource_count)) |j| { if (contains(parent.resources[j], r)) ok = true; } if (!ok) return error.NotContained; } var d = std.mem.zeroes(device_abi.DeviceDescriptor); d.id = count; d.parent = parent_id; d.class = descriptor.class; d.pci_class = descriptor.pci_class; d.hid_len = @min(descriptor.hid_len, d.hid.len); @memcpy(d.hid[0..@intCast(d.hid_len)], descriptor.hid[0..@intCast(d.hid_len)]); d.resource_count = descriptor.resource_count; for (0..@intCast(descriptor.resource_count)) |i| d.resources[i] = descriptor.resources[i]; devices[count] = d; registrar[count] = owner; count += 1; return d.id; }