a5840789ddf1cb356b0170a177d2c776fbf72534
9
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f7151ed577 |
kernel: the device table has no ceiling; a runaway is charged to whoever caused it
maximum_devices = 64 is gone. It was a guess about someone else's computer, and because it was shared, one driver's enumeration starved every other — which is how an AMD Ryzen booted with a working display, no USB and no storage. The table now grows from the kernel heap. It was always built after heap.init; nothing ever prevented this except it having been written static first. What replaces it is an allowance charged to the registrar, so a driver looping device_register exhausts its own and every other driver carries on. It is declared as what it is — a runaway detector, NOT a security boundary. A quota generous enough never to bite a real machine is still generous enough to be unpleasant, and it is not trying to be the defence; delegation is. What this catches is a legitimate driver in a loop, early, attributably, and without collateral. Reaching 4096 is a bug report, not a tuning request. The initial block is 8, deliberately small. Sizing it for a typical machine would mean the growth path never ran on the hardware we test on and only woke up on someone else's larger machine — the exact failure shape this track exists to stop. At 8 it grows several times every boot; disabling growth now fails the suite with the HPET not fitting, which is the Ryzen failure in miniature. The comptime coupling assert added earlier fired, and was right to. confined (one slot per device id) and domains (the IOMMU's own translation pool) were sized by the same constant only because device ids happened to stop at 64 too. Two unrelated quantities: confined now grows with the device table, while maximum_domains stays as the hardware's number — both VT-d and AMD-Vi report how many domains they support, and reading it is phase 4. The assert existed for exactly this and did its job. Suite 118/118. |
||
|
|
a2d6056772 |
usb: the xHCI controller arrives by delegation, not by claiming
The first driver to stop claiming its own hardware. The device manager holds the controller and transfers it in the hello reply, so its matching becomes authoritative instead of advisory — until now the driver claimed the id it found in argv[1], and any process could have claimed the same integer first. The manager claims before it spawns, so there is no window in which anything else could take the device, and transfers in onHello using invocation.sender — the kernel-stamped task id, which cannot be forged by the caller. hello is synchronous, so the transfer has completed before the reply lands: no gap between being told yes and holding the thing. usb-xhci-bus's hello moves from after controller bring-up to before anything that needs the device, which is the bring-up reorder the design predicted. It is the first member of an explicit delegated set, so every unconverted driver keeps claiming exactly as before and the suite stays green; the set and device_claim both go at D6. D3 and D4 could not be separated and the plan records why: the moment the manager claims, any driver still calling device_claim is refused, and D3 applied to nothing changes no behaviour and cannot be tested. This step introduced a regression and the incremental conversion is what caught it. confineDevice runs inside systemDeviceClaim, so a device arriving by transfer was never confined for its new owner. Three IOMMU+USB cases failed on the driver's DMA rings going unbound, and two worse consequences were latent: a manager death would have torn down a domain a live driver was using, and a driver death would have leaked one. iommu.reassign now moves the confinement with the device, keeping the domain and its attachment intact so it never translates through nothing. Converting all five drivers at once would have produced the same three failures with five suspects. A log line of mine claimed "holding controller device N" before anything verified it — it printed even in the failure case, where the driver held nothing. Reworded to state only what is known there: where the registers are. usb-hid asserts the delegation with the device id backreferenced, so the id delegated and the id the driver ends up with must match. Emptying the delegated set fails it with "hello acknowledged" then "mmio_map failed". usb-hub failed once in a full run and has passed six times since (four isolated, two full) — recorded in the plan as a suspected instance of the known intermittent AP fault, not dismissed, since this step did shift boot timing. Suite 118/118. |
||
|
|
f4eb88e7d2 |
build: a new compile-time ceiling declares itself or does not land
The convention that tunables live in system/parameters.zig with their reasoning attached predates this and got 2% compliance — 5 of 235. A convention with no teeth is how a bare `const maximum_devices = 64` reached an AMD desktop and cost it USB and storage. This is the same rule with a gate behind it. tools/check-bounds.py finds every bound-shaped declaration — a `maximum_*` const with a literal value, or a type with a literal array length — and requires the five-field block above it: what it counts, who decides its size, what it protects, what happens at the limit, and how anyone finds out. The at-limit vocabulary is closed: refuse, degrade, truncate, grow. There is deliberately no way to spell "silent", no way to spell "drop", and nothing meaning "allow", so the behaviours that did the damage cannot be written down. Truncation is legal only carrying a marker the reader can see, which is why klog_maximum_message qualifies and a USB descriptor cut at 512 bytes does not. An array length that names a declared bound is not itself a bound; only literal lengths are flagged, which pushes ceilings toward having names. The 273 that predate the rule are allowlisted so this lands without a tree-wide sweep in front of it, and that list may only shrink: declaring a bound means deleting its line, and the check fails on a stale entry too. Nothing may be added. Wired into `zig build test` and available alone as `zig build bounds`. Not in the default build — it reads the whole tree, and a red bounds check should not stop you booting a kernel. Five are now declared rather than allowlisted. Writing them out is its own argument: maximum_devices reads "protects: nothing — this is a sizing guess about someone else's computer", and maximum_tasks now carries the fact that it has been raised twice, each time by something that outgrew it. Verified the gate refuses an undeclared bound, a declared one using forbidden vocabulary, and an allowlist entry that has since been declared. Suite 115/115. |
||
|
|
a86559648e |
kernel: a refusal names its rule, and two bounds stop failing open
An AMD Ryzen booted to a working compositor with no USB and no storage, and the log said only "register refused". A tree-wide audit of every compile-time ceiling followed: 235 of them, 139 on quantities the machine or a file decides rather than us, 5 documented anywhere, 171 silent when reached. docs/fixed-bounds-audit.md has the inventory. Errno attribution. The errno space was split between the kernel and the envelope, free to drift; it is now one list in system/abi.zig, restated on both sides, with a comptime check in library/device/driver where the two halves are visible. device_register's six refusals and device_claim's three are distinct codes, so a bus driver can say which rule stopped it, and BadParent splits into NoSuchParent and NotYourParent. pci-bus reconciles found against registered instead of counting refused functions as found. Idempotency ordering. The child cap was checked before the identity match, so a restarted bus was refused its own devices — the supervision restart the system leans on ratcheted toward a degraded machine. A re-registration consumes no slot and is now admitted first. IOMMU fail-closed. confineDevice returned success for a device id past the confinement table, leaving the device outside every domain while the caller believed it confined — unreachable only while ids stop at 64, which both the inventory move and a hardware-reported domain count would change. It refuses now, and the coupling to the broker's device cap is a comptime assert rather than a sentence in a comment. PCI apertures. The bridge's MMIO apertures are derived from the holes in the firmware memory map, and the derivation copied sub-4 GiB entries into a fixed [64] array and skipped the rest. A skipped region is not merely lost: the gap finder concludes it is free, so a real machine's 60-200 entry map yields an aperture over live RAM, and containment then admits a child BAR covering kernel memory. Rewritten to walk the map in place, with the hole finder extracted as a pure function and driven by a synthetic 100-entry map in a new test case. Both new tests were verified to fail on the old code. parameters.zig gains the rationale it was missing and loses a stale sentence pointing at the wrong file; vdso.md documents the errno space, including EPEER, which had no written meaning anywhere. docs/os-development/bounds.md is how a ceiling is declared from here. docs/bounds-track-plan.md is the plan to remove the ones we invented. Suite 114 -> 115. |
||
|
|
fa8203cdba |
kernel: the IOMMU backends move behind the architecture boundary
VT-d and AMD-Vi are x86 hardware, but lived in the architecture-neutral kernel tree and leaked further: the core's public Kind enum named both vendors, and the ACPI parser read the VT-d version/capability registers (raw volatile MMIO inside table discovery). Now the vendor backends live in architecture/x86_64/ behind architecture.iommu — the core hands over the discovery facts plus an injected environment (frame allocation + the log sink, the same pattern enablePaging uses) and receives the hardware vtable back, so the backends never import kernel internals and an ARM port supplies its SMMU with no core change. Discovery keeps table facts only; the live-unit register check moved into VT-d detect (version reading zero now stays fail-open). The unused kindOf() is gone. Log shapes the harness pins (iommu online, DANOS-IOMMU-FAULT) are unchanged; all five IOMMU QEMU cases pass. |
||
|
|
6a687fbc2b |
iommu: AMD-Vi backend behind the vendor-neutral core
The second hardware backend. The IOMMU core, DMA-region capabilities, and per-device enforcement are unchanged; this adds AMD-Vi (IVRS) as an alternative to Intel VT-d (DMAR) under the same Backend vtable. - parseIvrs records the IOMMU control-register base from the first IVHD; the platform layer gains iommu_is_amd, and the core picks the backend by vendor at init. VT-d and AMD-Vi are mutually exclusive on real hardware. - iommu-amd.zig: a 2 MiB device table (every DTE zeroed = deny-all until a device is claimed), AMD native-format page tables (4 KiB leaves), a command buffer (INVALIDATE_DEVTAB_ENTRY / INVALIDATE_IOMMU_PAGES / COMPLETION_WAIT) and an event log for faults. The DTE forwards interrupts unmapped, so MSI passthrough works exactly as on VT-d. - The boot log and the iommu self-test are now vendor-aware. **UNTESTED on real AMD hardware** — danos is developed on Intel, so this is validated only against QEMU's amd-iommu, and every log line and doc says so. QEMU quirk handled: its amd-iommu does not observe the COMPLETION_WAIT store form, but consumes the command ring synchronously on the tail-register write, so invalidations are already applied by the time we poll — the backend warns once and proceeds. Cases: amd-iommu (detection + scratch-domain walker) and amd-iommu-usb-storage (full storage stack through AMD device-table translation with per-grant capabilities), both green. 106/106. This completes the IOVA/IOMMU-enforcement track: per-device DMA domains on both vendors, with buffers reachable only through delegated capabilities. |
||
|
|
4e7cbc9792 |
iommu: DMA-region capabilities — per-grant reachability, protocol flag-day
Replaces L2's interim DMA pool (every buffer reachable by every claimed device) with true per-grant confinement: a device reaches only buffers whose capability was delegated to its driver. Kernel: - DmaRegionObject (handle kind 2): a delegation token naming a dma_alloc'd region, passable across processes on the IPC cap slot like an endpoint or shared-memory object. Frames stay owned by the allocating address space (freed on dma_free/teardown as before); the token carries a `dead` flag so a stale downstream handle can no longer bind a freed region. - dma_alloc gains the dma_shareable flag: it returns a capability handle in r8 and every region is tracked in a registry. A task's own regions auto-bind into the devices it claims (its rings just work); foreign buffers are bound explicitly. - dma_bind / dma_unbind / handle_close syscalls (51-53). dma_bind maps a held region (or shared-memory) capability into a claimed device's domain; it is idempotent. handle_close reclaims a table slot (raised 16 -> 32). - dma_free and task death unmap a region from every domain and invalidate BEFORE its frames return to the allocator — the stale-IOTLB use-after- free window, closed structurally. Protocols (flag-day): block gains attach, usb-transfer gains dma_attach — each carries a region capability on the cap slot. fat allocates its bounce buffer shareable and attaches it; usb-storage allocates its transport buffers shareable, attaches them to the controller, and forwards fat's capability downstream; usb-xhci-bus binds and closes; virtio-gpu binds its shared scanout surface. The physical addresses on the wire are unchanged (identity IOVA), so no register-programming code moved. Cross-process DMA (fat -> usb-storage -> xHC) now flows only through delegated capabilities. iommu-usb-storage / iommu-usb-hid / iommu-fault all green under per-grant enforcement; 104/104 overall (fail-open paths unchanged). |
||
|
|
e94adcfc02 |
iommu: per-device domains with interim DMA-pool enforcement
Replaces L1's shared blanket identity domain with a private translation
domain per claimed PCI function. A device now reaches only:
- the DMA pool: every dma_alloc'd region, mapped into every claimed
device's domain (poolAdd/poolRemove, driven from the dma_alloc and
dma_free syscalls). This keeps the cross-process buffer handoff
working (fat's bounce buffer reaches the xHC) while blocking the
kernel, page tables, process heaps, MMIO, and unallocated RAM.
- its own firmware reserved region (RMRR), seeded at confine time.
The pool is the honest interim: devices can still reach one another's
DMA buffers. The DMA-region capability layer (next) narrows it to
per-grant reachability.
dma_free unmaps from every domain and invalidates BEFORE the frames
return to the allocator, closing the stale-IOTLB use-after-free window.
Driver death tears down its domains (detach + free tables) before the
broker claims and DMA frames are released.
New iommu_fault_drain syscall (+ driver.iommuFaultDrain) forces pending
fault records to the log on demand. The new iommu-fault case proves it:
a claimed e1000e is programmed to DMA-fetch its TX ring from an unmapped
page; VT-d faults the access (bdf 00:03.0 addr 0x1000 reason 0x6) and the
system stays alive. 104/104.
|
||
|
|
f477ef7d9f |
iommu: enable Intel VT-d translation with per-claim device confinement
First enforcement step of the IOVA track. A vendor-neutral IOMMU core (iommu.zig) drives an Intel VT-d backend (iommu-intel.zig) to give DMA a real translation layer instead of the fail-open free-for-all M16 left. - Boot posture is now stated explicitly: "iommu online (Intel VT-d)" with version/agaw/rmrr, or "none present - DMA fail-open (unisolated)". - DMAR parsing extended to select the INCLUDE_PCI_ALL unit (real Intel PCs put an iGPU-scoped unit first) and record single-path-endpoint RMRRs; multi-hop scopes and extra DRHDs are counted and warned, never silently dropped. - Translation is enabled at boot into a blanket identity domain (all RAM + RMRRs, 2 MiB leaves). PCI functions are enumerated post-boot by the ring-3 pci-bus driver, so a device is attached to the domain when its driver claims it (confineDevice, with claim rollback if confinement fails) and detached on driver death, before broker release and DMA frame teardown. Unclaimed devices are non-present: their DMA faults. - Interrupt remapping stays off, so MSI writes to 0xFEE00000 bypass translation and the interrupt-driven xHC keeps working. - devices-broker gains pciAddressOf (derives BDF from the config-space ECAM offset), unclaim, and forEachPciFunction. Faults are drained and logged rate-limited as DANOS-IOMMU-FAULT. Cases: iommu extended (translation on, scratch-domain map/resolve/unmap, zero idle faults); new iommu-usb-storage and iommu-usb-hid run the full storage + input stacks through translated DMA with MSI intact. 103/103. |