Commit Graph
9 Commits
Author SHA1 Message Date
Daniel Samson f7151ed577 kernel: the device table has no ceiling; a runaway is charged to whoever caused it
maximum_devices = 64 is gone. It was a guess about someone else's computer,
and because it was shared, one driver's enumeration starved every other —
which is how an AMD Ryzen booted with a working display, no USB and no
storage. The table now grows from the kernel heap. It was always built after
heap.init; nothing ever prevented this except it having been written static
first.

What replaces it is an allowance charged to the registrar, so a driver
looping device_register exhausts its own and every other driver carries on.
It is declared as what it is — a runaway detector, NOT a security boundary.
A quota generous enough never to bite a real machine is still generous
enough to be unpleasant, and it is not trying to be the defence; delegation
is. What this catches is a legitimate driver in a loop, early, attributably,
and without collateral. Reaching 4096 is a bug report, not a tuning request.

The initial block is 8, deliberately small. Sizing it for a typical machine
would mean the growth path never ran on the hardware we test on and only
woke up on someone else's larger machine — the exact failure shape this
track exists to stop. At 8 it grows several times every boot; disabling
growth now fails the suite with the HPET not fitting, which is the Ryzen
failure in miniature.

The comptime coupling assert added earlier fired, and was right to. confined
(one slot per device id) and domains (the IOMMU's own translation pool) were
sized by the same constant only because device ids happened to stop at 64
too. Two unrelated quantities: confined now grows with the device table,
while maximum_domains stays as the hardware's number — both VT-d and AMD-Vi
report how many domains they support, and reading it is phase 4. The assert
existed for exactly this and did its job.

Suite 118/118.
2026-08-08 18:40:20 +01:00
Daniel Samson a2d6056772 usb: the xHCI controller arrives by delegation, not by claiming
The first driver to stop claiming its own hardware. The device manager holds
the controller and transfers it in the hello reply, so its matching becomes
authoritative instead of advisory — until now the driver claimed the id it
found in argv[1], and any process could have claimed the same integer first.

The manager claims before it spawns, so there is no window in which anything
else could take the device, and transfers in onHello using invocation.sender
— the kernel-stamped task id, which cannot be forged by the caller. hello is
synchronous, so the transfer has completed before the reply lands: no gap
between being told yes and holding the thing.

usb-xhci-bus's hello moves from after controller bring-up to before anything
that needs the device, which is the bring-up reorder the design predicted.
It is the first member of an explicit delegated set, so every unconverted
driver keeps claiming exactly as before and the suite stays green; the set
and device_claim both go at D6. D3 and D4 could not be separated and the
plan records why: the moment the manager claims, any driver still calling
device_claim is refused, and D3 applied to nothing changes no behaviour and
cannot be tested.

This step introduced a regression and the incremental conversion is what
caught it. confineDevice runs inside systemDeviceClaim, so a device arriving
by transfer was never confined for its new owner. Three IOMMU+USB cases
failed on the driver's DMA rings going unbound, and two worse consequences
were latent: a manager death would have torn down a domain a live driver was
using, and a driver death would have leaked one. iommu.reassign now moves
the confinement with the device, keeping the domain and its attachment
intact so it never translates through nothing. Converting all five drivers
at once would have produced the same three failures with five suspects.

A log line of mine claimed "holding controller device N" before anything
verified it — it printed even in the failure case, where the driver held
nothing. Reworded to state only what is known there: where the registers
are.

usb-hid asserts the delegation with the device id backreferenced, so the id
delegated and the id the driver ends up with must match. Emptying the
delegated set fails it with "hello acknowledged" then "mmio_map failed".

usb-hub failed once in a full run and has passed six times since (four
isolated, two full) — recorded in the plan as a suspected instance of the
known intermittent AP fault, not dismissed, since this step did shift boot
timing.

Suite 118/118.
2026-08-08 17:55:59 +01:00
Daniel Samson f4eb88e7d2 build: a new compile-time ceiling declares itself or does not land
The convention that tunables live in system/parameters.zig with their
reasoning attached predates this and got 2% compliance — 5 of 235. A
convention with no teeth is how a bare `const maximum_devices = 64` reached
an AMD desktop and cost it USB and storage. This is the same rule with a
gate behind it.

tools/check-bounds.py finds every bound-shaped declaration — a `maximum_*`
const with a literal value, or a type with a literal array length — and
requires the five-field block above it: what it counts, who decides its
size, what it protects, what happens at the limit, and how anyone finds out.

The at-limit vocabulary is closed: refuse, degrade, truncate, grow. There is
deliberately no way to spell "silent", no way to spell "drop", and nothing
meaning "allow", so the behaviours that did the damage cannot be written
down. Truncation is legal only carrying a marker the reader can see, which
is why klog_maximum_message qualifies and a USB descriptor cut at 512 bytes
does not.

An array length that names a declared bound is not itself a bound; only
literal lengths are flagged, which pushes ceilings toward having names.

The 273 that predate the rule are allowlisted so this lands without a
tree-wide sweep in front of it, and that list may only shrink: declaring a
bound means deleting its line, and the check fails on a stale entry too.
Nothing may be added.

Wired into `zig build test` and available alone as `zig build bounds`. Not
in the default build — it reads the whole tree, and a red bounds check
should not stop you booting a kernel.

Five are now declared rather than allowlisted. Writing them out is its own
argument: maximum_devices reads "protects: nothing — this is a sizing guess
about someone else's computer", and maximum_tasks now carries the fact that
it has been raised twice, each time by something that outgrew it.

Verified the gate refuses an undeclared bound, a declared one using
forbidden vocabulary, and an allowlist entry that has since been declared.
Suite 115/115.
2026-08-08 11:27:27 +01:00
Daniel Samson a86559648e kernel: a refusal names its rule, and two bounds stop failing open
An AMD Ryzen booted to a working compositor with no USB and no storage,
and the log said only "register refused". A tree-wide audit of every
compile-time ceiling followed: 235 of them, 139 on quantities the machine
or a file decides rather than us, 5 documented anywhere, 171 silent when
reached. docs/fixed-bounds-audit.md has the inventory.

Errno attribution. The errno space was split between the kernel and the
envelope, free to drift; it is now one list in system/abi.zig, restated on
both sides, with a comptime check in library/device/driver where the two
halves are visible. device_register's six refusals and device_claim's three
are distinct codes, so a bus driver can say which rule stopped it, and
BadParent splits into NoSuchParent and NotYourParent. pci-bus reconciles
found against registered instead of counting refused functions as found.

Idempotency ordering. The child cap was checked before the identity match,
so a restarted bus was refused its own devices — the supervision restart the
system leans on ratcheted toward a degraded machine. A re-registration
consumes no slot and is now admitted first.

IOMMU fail-closed. confineDevice returned success for a device id past the
confinement table, leaving the device outside every domain while the caller
believed it confined — unreachable only while ids stop at 64, which both the
inventory move and a hardware-reported domain count would change. It refuses
now, and the coupling to the broker's device cap is a comptime assert rather
than a sentence in a comment.

PCI apertures. The bridge's MMIO apertures are derived from the holes in the
firmware memory map, and the derivation copied sub-4 GiB entries into a
fixed [64] array and skipped the rest. A skipped region is not merely lost:
the gap finder concludes it is free, so a real machine's 60-200 entry map
yields an aperture over live RAM, and containment then admits a child BAR
covering kernel memory. Rewritten to walk the map in place, with the hole
finder extracted as a pure function and driven by a synthetic 100-entry map
in a new test case. Both new tests were verified to fail on the old code.

parameters.zig gains the rationale it was missing and loses a stale sentence
pointing at the wrong file; vdso.md documents the errno space, including
EPEER, which had no written meaning anywhere.

docs/os-development/bounds.md is how a ceiling is declared from here.
docs/bounds-track-plan.md is the plan to remove the ones we invented.

Suite 114 -> 115.
2026-08-08 11:09:54 +01:00
Daniel Samson fa8203cdba kernel: the IOMMU backends move behind the architecture boundary
VT-d and AMD-Vi are x86 hardware, but lived in the architecture-neutral
kernel tree and leaked further: the core's public Kind enum named both
vendors, and the ACPI parser read the VT-d version/capability registers
(raw volatile MMIO inside table discovery). Now the vendor backends
live in architecture/x86_64/ behind architecture.iommu — the core hands
over the discovery facts plus an injected environment (frame allocation
+ the log sink, the same pattern enablePaging uses) and receives the
hardware vtable back, so the backends never import kernel internals and
an ARM port supplies its SMMU with no core change. Discovery keeps
table facts only; the live-unit register check moved into VT-d detect
(version reading zero now stays fail-open). The unused kindOf() is
gone. Log shapes the harness pins (iommu online, DANOS-IOMMU-FAULT)
are unchanged; all five IOMMU QEMU cases pass.
2026-07-30 07:16:08 +01:00
Daniel Samson 6a687fbc2b iommu: AMD-Vi backend behind the vendor-neutral core
The second hardware backend. The IOMMU core, DMA-region capabilities, and
per-device enforcement are unchanged; this adds AMD-Vi (IVRS) as an
alternative to Intel VT-d (DMAR) under the same Backend vtable.

- parseIvrs records the IOMMU control-register base from the first IVHD;
  the platform layer gains iommu_is_amd, and the core picks the backend by
  vendor at init. VT-d and AMD-Vi are mutually exclusive on real hardware.
- iommu-amd.zig: a 2 MiB device table (every DTE zeroed = deny-all until a
  device is claimed), AMD native-format page tables (4 KiB leaves), a
  command buffer (INVALIDATE_DEVTAB_ENTRY / INVALIDATE_IOMMU_PAGES /
  COMPLETION_WAIT) and an event log for faults. The DTE forwards
  interrupts unmapped, so MSI passthrough works exactly as on VT-d.
- The boot log and the iommu self-test are now vendor-aware.

**UNTESTED on real AMD hardware** — danos is developed on Intel, so this
is validated only against QEMU's amd-iommu, and every log line and doc
says so. QEMU quirk handled: its amd-iommu does not observe the
COMPLETION_WAIT store form, but consumes the command ring synchronously on
the tail-register write, so invalidations are already applied by the time
we poll — the backend warns once and proceeds.

Cases: amd-iommu (detection + scratch-domain walker) and
amd-iommu-usb-storage (full storage stack through AMD device-table
translation with per-grant capabilities), both green. 106/106.

This completes the IOVA/IOMMU-enforcement track: per-device DMA domains on
both vendors, with buffers reachable only through delegated capabilities.
2026-07-26 18:31:27 +01:00
Daniel Samson 4e7cbc9792 iommu: DMA-region capabilities — per-grant reachability, protocol flag-day
Replaces L2's interim DMA pool (every buffer reachable by every claimed
device) with true per-grant confinement: a device reaches only buffers
whose capability was delegated to its driver.

Kernel:
- DmaRegionObject (handle kind 2): a delegation token naming a dma_alloc'd
  region, passable across processes on the IPC cap slot like an endpoint
  or shared-memory object. Frames stay owned by the allocating address
  space (freed on dma_free/teardown as before); the token carries a `dead`
  flag so a stale downstream handle can no longer bind a freed region.
- dma_alloc gains the dma_shareable flag: it returns a capability handle
  in r8 and every region is tracked in a registry. A task's own regions
  auto-bind into the devices it claims (its rings just work); foreign
  buffers are bound explicitly.
- dma_bind / dma_unbind / handle_close syscalls (51-53). dma_bind maps a
  held region (or shared-memory) capability into a claimed device's domain;
  it is idempotent. handle_close reclaims a table slot (raised 16 -> 32).
- dma_free and task death unmap a region from every domain and invalidate
  BEFORE its frames return to the allocator — the stale-IOTLB use-after-
  free window, closed structurally.

Protocols (flag-day): block gains attach, usb-transfer gains dma_attach —
each carries a region capability on the cap slot. fat allocates its bounce
buffer shareable and attaches it; usb-storage allocates its transport
buffers shareable, attaches them to the controller, and forwards fat's
capability downstream; usb-xhci-bus binds and closes; virtio-gpu binds its
shared scanout surface. The physical addresses on the wire are unchanged
(identity IOVA), so no register-programming code moved.

Cross-process DMA (fat -> usb-storage -> xHC) now flows only through
delegated capabilities. iommu-usb-storage / iommu-usb-hid / iommu-fault
all green under per-grant enforcement; 104/104 overall (fail-open paths
unchanged).
2026-07-26 18:31:27 +01:00
Daniel Samson e94adcfc02 iommu: per-device domains with interim DMA-pool enforcement
Replaces L1's shared blanket identity domain with a private translation
domain per claimed PCI function. A device now reaches only:
  - the DMA pool: every dma_alloc'd region, mapped into every claimed
    device's domain (poolAdd/poolRemove, driven from the dma_alloc and
    dma_free syscalls). This keeps the cross-process buffer handoff
    working (fat's bounce buffer reaches the xHC) while blocking the
    kernel, page tables, process heaps, MMIO, and unallocated RAM.
  - its own firmware reserved region (RMRR), seeded at confine time.
The pool is the honest interim: devices can still reach one another's
DMA buffers. The DMA-region capability layer (next) narrows it to
per-grant reachability.

dma_free unmaps from every domain and invalidates BEFORE the frames
return to the allocator, closing the stale-IOTLB use-after-free window.
Driver death tears down its domains (detach + free tables) before the
broker claims and DMA frames are released.

New iommu_fault_drain syscall (+ driver.iommuFaultDrain) forces pending
fault records to the log on demand. The new iommu-fault case proves it:
a claimed e1000e is programmed to DMA-fetch its TX ring from an unmapped
page; VT-d faults the access (bdf 00:03.0 addr 0x1000 reason 0x6) and the
system stays alive. 104/104.
2026-07-26 18:31:27 +01:00
Daniel Samson f477ef7d9f iommu: enable Intel VT-d translation with per-claim device confinement
First enforcement step of the IOVA track. A vendor-neutral IOMMU core
(iommu.zig) drives an Intel VT-d backend (iommu-intel.zig) to give DMA a
real translation layer instead of the fail-open free-for-all M16 left.

- Boot posture is now stated explicitly: "iommu online (Intel VT-d)"
  with version/agaw/rmrr, or "none present - DMA fail-open (unisolated)".
- DMAR parsing extended to select the INCLUDE_PCI_ALL unit (real Intel
  PCs put an iGPU-scoped unit first) and record single-path-endpoint
  RMRRs; multi-hop scopes and extra DRHDs are counted and warned, never
  silently dropped.
- Translation is enabled at boot into a blanket identity domain (all RAM
  + RMRRs, 2 MiB leaves). PCI functions are enumerated post-boot by the
  ring-3 pci-bus driver, so a device is attached to the domain when its
  driver claims it (confineDevice, with claim rollback if confinement
  fails) and detached on driver death, before broker release and DMA
  frame teardown. Unclaimed devices are non-present: their DMA faults.
- Interrupt remapping stays off, so MSI writes to 0xFEE00000 bypass
  translation and the interrupt-driven xHC keeps working.
- devices-broker gains pciAddressOf (derives BDF from the config-space
  ECAM offset), unclaim, and forEachPciFunction.

Faults are drained and logged rate-limited as DANOS-IOMMU-FAULT.

Cases: iommu extended (translation on, scratch-domain map/resolve/unmap,
zero idle faults); new iommu-usb-storage and iommu-usb-hid run the full
storage + input stacks through translated DMA with MSI intact. 103/103.
2026-07-26 18:31:27 +01:00