3111c7c5e6bc6e2029da6bfa26b0abcb497e4105
7
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
3111c7c5e6 |
kernel: device_transfer — you may give away what you hold
The mechanism behind delegation, which device-manager.md named as the step after hello: the device manager claims what discovery seeded and hands each device to the driver it matched, so assignment stops being first-come-first-served. It is a MOVE, not a copy. A claim is exclusive (driver-model.md, invariant 1), so the giver stops holding the device the instant the receiver starts. That is why this is a new syscall rather than the M13 capability path, where a passed handle is shared refcounted — exclusivity cannot be expressed that way. The kernel's whole rule is that you may give away what you hold. It has no notion of which task is the device manager and deliberately gains none: a binary name inside the kernel is not something that cannot safely live in user space. A recipient that does not exist is refused, because a device moved to nobody would be unreachable for the rest of the boot — nothing un-holds a device but task death. Three errnos, each naming its own rule: ENODEV no such device, EPERM you do not hold it, ESRCH no such recipient. Nothing uses it yet. The five claimants move across one at a time in D4-D5, so the suite stays green throughout and a regression names the driver that caused it. Ten assertions, verified to discriminate: removing the ownership check flips four of them, including the giveaway that an illegal transfer then blocks the legitimate claim behind it. Suite 116 -> 117. |
||
|
|
a86559648e |
kernel: a refusal names its rule, and two bounds stop failing open
An AMD Ryzen booted to a working compositor with no USB and no storage, and the log said only "register refused". A tree-wide audit of every compile-time ceiling followed: 235 of them, 139 on quantities the machine or a file decides rather than us, 5 documented anywhere, 171 silent when reached. docs/fixed-bounds-audit.md has the inventory. Errno attribution. The errno space was split between the kernel and the envelope, free to drift; it is now one list in system/abi.zig, restated on both sides, with a comptime check in library/device/driver where the two halves are visible. device_register's six refusals and device_claim's three are distinct codes, so a bus driver can say which rule stopped it, and BadParent splits into NoSuchParent and NotYourParent. pci-bus reconciles found against registered instead of counting refused functions as found. Idempotency ordering. The child cap was checked before the identity match, so a restarted bus was refused its own devices — the supervision restart the system leans on ratcheted toward a degraded machine. A re-registration consumes no slot and is now admitted first. IOMMU fail-closed. confineDevice returned success for a device id past the confinement table, leaving the device outside every domain while the caller believed it confined — unreachable only while ids stop at 64, which both the inventory move and a hardware-reported domain count would change. It refuses now, and the coupling to the broker's device cap is a comptime assert rather than a sentence in a comment. PCI apertures. The bridge's MMIO apertures are derived from the holes in the firmware memory map, and the derivation copied sub-4 GiB entries into a fixed [64] array and skipped the rest. A skipped region is not merely lost: the gap finder concludes it is free, so a real machine's 60-200 entry map yields an aperture over live RAM, and containment then admits a child BAR covering kernel memory. Rewritten to walk the map in place, with the hole finder extracted as a pure function and driven by a synthetic 100-entry map in a new test case. Both new tests were verified to fail on the old code. parameters.zig gains the rationale it was missing and loses a stale sentence pointing at the wrong file; vdso.md documents the errno space, including EPEER, which had no written meaning anywhere. docs/os-development/bounds.md is how a ceiling is declared from here. docs/bounds-track-plan.md is the plan to remove the ones we invented. Suite 114 -> 115. |
||
|
|
2719b93530 |
library: the last three protocols speak the envelope
These were the awkward ones. Each began with an operation packed into a single byte — two of them with a version wedged in beside it — so there was no wrapping them: the layouts had to be rebuilt. The device manager's own enumerate and subscribe become the reserved verbs that mean the same thing everywhere, its replies lose three status structs the envelope already carries, and a device id becomes the packet's target. Power drops the version it repeated on every request, because describe is the handshake, and stops claiming a 64-byte ceiling it never needed for calls. USB moves a control transfer's data to the packet tail in both directions, which makes the status length the transferred length and retires a field that had been saying the same thing twice. The danger in this one was not the protocols but their readers. Init recognised a power button by two bytes at the head of a message, the ACPI service dispatched on the first byte, the xHCI driver read its operation with a raw integer load, and the HID drivers reinterpreted a report wholesale — none of which would have failed to compile once the layouts moved. They would simply have stopped: no shutdown on the power button, no reports from the keyboard. Every one of them now reads through the generated types, and the shutdown gate that answers only a subscriber is the same code it was. Two sizes were decided by measuring rather than assuming. The child-added message is both a request and the event broadcast to subscribers, and alignment rounds it to 48 bytes, which puts its packet exactly on the 64-byte push floor — a test pins that, because a field added carelessly would now overflow it. The interrupt report gives up eight bytes of inline room to make space for the header; the two drivers that produce reports send eight and four. Suite 110/110. |
||
|
|
1379b699f3 |
init: /protocol replaces the ServiceId registry
A protocol is reached by name now, not by a compile-time integer. Init is PID 1 and already knows which binary it started, so init serves /protocol as a vfs backend: bind claims a contract with the provider's endpoint attached, open answers with that endpoint as the reply's capability, and readdir lists what is bound with the task and binary behind it. The kernel reserves the prefix — nothing may mount over it, under it, or unmount it — and ServiceId, ipc_register and ipc_lookup are gone, their syscall numbers left vacant. A bind is authorized by who the caller *is*: the kernel-stamped binary together with the supervising task's identity, matched against /system/configuration/protocol.csv. Identity, not spelling — spawn is ungated, so an attacker can run any bundled binary, and a name-only rule would have let it launder grants through an init of its own making. A name a live process holds is refused to everyone else; a dead one's is released. Three review rounds against a hostile ring-3 process found what 108 green tests could not, because the suite contains no attacker. Publishing init's supervision endpoint as the registry put PID 1's mailbox in every process's hands, where two forged bytes reached the shutdown path: privileged traffic is now believed only from the task that holds the contract it speaks for. A capability arriving on a request outlived every path that ignored it, one handle per call until the table was full — in init, and in the harness ten services share — so the arriving capability is owned by the turn and released unless a handler says otherwise. And the kernel let anyone holding an endpoint handle aim signals, timers, exit notices and interrupts at it: binding now requires having created it. Suite 108/108. The new protocol-registry case asserts eleven properties, each one an attack that must fail. |
||
|
|
4e7cbc9792 |
iommu: DMA-region capabilities — per-grant reachability, protocol flag-day
Replaces L2's interim DMA pool (every buffer reachable by every claimed device) with true per-grant confinement: a device reaches only buffers whose capability was delegated to its driver. Kernel: - DmaRegionObject (handle kind 2): a delegation token naming a dma_alloc'd region, passable across processes on the IPC cap slot like an endpoint or shared-memory object. Frames stay owned by the allocating address space (freed on dma_free/teardown as before); the token carries a `dead` flag so a stale downstream handle can no longer bind a freed region. - dma_alloc gains the dma_shareable flag: it returns a capability handle in r8 and every region is tracked in a registry. A task's own regions auto-bind into the devices it claims (its rings just work); foreign buffers are bound explicitly. - dma_bind / dma_unbind / handle_close syscalls (51-53). dma_bind maps a held region (or shared-memory) capability into a claimed device's domain; it is idempotent. handle_close reclaims a table slot (raised 16 -> 32). - dma_free and task death unmap a region from every domain and invalidate BEFORE its frames return to the allocator — the stale-IOTLB use-after- free window, closed structurally. Protocols (flag-day): block gains attach, usb-transfer gains dma_attach — each carries a region capability on the cap slot. fat allocates its bounce buffer shareable and attaches it; usb-storage allocates its transport buffers shareable, attaches them to the controller, and forwards fat's capability downstream; usb-xhci-bus binds and closes; virtio-gpu binds its shared scanout surface. The physical addresses on the wire are unchanged (identity IOVA), so no register-programming code moved. Cross-process DMA (fat -> usb-storage -> xHC) now flows only through delegated capabilities. iommu-usb-storage / iommu-usb-hid / iommu-fault all green under per-grant enforcement; 104/104 overall (fail-open paths unchanged). |
||
|
|
e94adcfc02 |
iommu: per-device domains with interim DMA-pool enforcement
Replaces L1's shared blanket identity domain with a private translation
domain per claimed PCI function. A device now reaches only:
- the DMA pool: every dma_alloc'd region, mapped into every claimed
device's domain (poolAdd/poolRemove, driven from the dma_alloc and
dma_free syscalls). This keeps the cross-process buffer handoff
working (fat's bounce buffer reaches the xHC) while blocking the
kernel, page tables, process heaps, MMIO, and unallocated RAM.
- its own firmware reserved region (RMRR), seeded at confine time.
The pool is the honest interim: devices can still reach one another's
DMA buffers. The DMA-region capability layer (next) narrows it to
per-grant reachability.
dma_free unmaps from every domain and invalidates BEFORE the frames
return to the allocator, closing the stale-IOTLB use-after-free window.
Driver death tears down its domains (detach + free tables) before the
broker claims and DMA frames are released.
New iommu_fault_drain syscall (+ driver.iommuFaultDrain) forces pending
fault records to the log on demand. The new iommu-fault case proves it:
a claimed e1000e is programmed to DMA-fetch its TX ring from an unmapped
page; VT-d faults the access (bdf 00:03.0 addr 0x1000 reason 0x6) and the
system stays alive. 104/104.
|
||
|
|
dded46726b |
reorg: relocate device/service clients (C3/C4)
- device.zig + device-manager.zig -> library/device/driver/driver.zig (module
"driver"): the driver author's whole interface — device access (claim/mmioMap/
irqBind/...) plus the device-manager hello() handshake, folded into one import.
- block.zig -> library/device/block/block.zig (a device type).
- display.zig/input.zig -> library/client/{display,input}/ (userspace-service
clients — they talk to services, not the kernel).
build.zig module graph updated; the runtime shim now maps runtime.device and
runtime.device_manager onto "driver", so consumers stay untouched (migrated in C2).
zig build green; driver-restart, usb-storage, display-native, input pass.
|