docs: flip stale status markers across the tracks (audit found 23)
An all-tracks docs-vs-code audit (the same method that caught the storage drift) found 23 confirmed inaccuracies where a doc's build-status claim no longer matches the source — status markers that were never flipped after a track landed, and a few paths left over from completed flag-days. All verified against the code before editing; docs only, no behavior change. The systemic ones: - IOMMU enforcement (driver-model.md, drivers.md): docs said enforcement was not built and "device_claim = ring 0" / "memory-safe is not true yet". It is built (per-device VT-d/AMD-Vi domains programmed at device_claim, -ECONFINE rollback, dma_alloc buffers bound and torn down at death; fail-open only with no IOMMU). Restated; M16 marker flipped to done. - The FHS flag-day paths: /etc/devices.csv -> /system/configuration/devices.csv (devices-csv.md, new-driver-checklist.md, device-manager.md), /var/log -> /system/logs (logging.md, new-driver-checklist.md), /mnt/usb -> /volumes/usb (process-management.md). Following the old paths silently breaks driver match. - protocol-namespace P4 "remaining" -> landed (only P5 remains); shared-fate fan-out "not yet enforced" -> enforced; wall_clock "not built" -> built; SMP affinity + fault-recovery "left" -> built; process_enumerate raw-pointer trust model -> checked copyToUser/EFAULT; bounds.md maximum_devices static hole -> dynamic per-registrar quota; init spawns fat -> volume-manager; config "hardcoded, move to /etc" -> already CSV data files; vdso.md three- value call; zig-self-hosting library/ layout; python argv "new" -> built. Found and fixed by a multi-agent audit across 12 doc clusters, each finding adversarially verified against the source.
This commit is contained in:
@@ -193,12 +193,14 @@ class driver, the device manager, or the kernel may share them freely.
|
||||
a PS/2 or 16550 driver possible; the low-rate legacy hardware that needs it is fine with
|
||||
a syscall per access. `io_port` resources were recorded by discovery and ignored — now
|
||||
they're used.
|
||||
- **M16 (detection)** — the IOMMU is now *found*: discovery parses the ACPI DMAR table,
|
||||
maps the first VT-d unit, and reads its version + capabilities (`iommu_present` in the
|
||||
platform info). This is detection only — **no translation domains are programmed, so
|
||||
DMA is still unprotected** (the caveat below). Enforcement lands with the first DMA
|
||||
driver, which is what there is to protect and test against. Proven in the `iommu` test,
|
||||
booted with an emulated `intel-iommu`.
|
||||
- **M16 (enforcement)** — the IOMMU is now *found and used*: discovery parses the ACPI
|
||||
DMAR table, maps the first VT-d unit, and reads its version + capabilities
|
||||
(`iommu_present` in the platform info), and **`device_claim` programs a private
|
||||
per-device translation domain** for the claimed function — rolling the claim back with
|
||||
`-ECONFINE` if it cannot confine it — so `dma_alloc` buffers are bound into that domain
|
||||
and torn down at process death. DMA is protected on any IOMMU-equipped machine; the
|
||||
system fails open only when there is no IOMMU at all. Proven in the `iommu` test, booted
|
||||
with an emulated `intel-iommu`.
|
||||
- **`system_spawn`** — a user-space supervisor starts a driver:
|
||||
`system_spawn(name, arguments)` loads a binary bundled in the initial-ramdisk as a
|
||||
fresh ring-3 process; `name` becomes the child's argv[0] and the optional
|
||||
@@ -369,25 +371,29 @@ capability walk (MSI, MSI-X, PCIe extended caps) without any new syscall.
|
||||
Note QEMU's HPET reports `Tn_FSB_INT_DEL_CAP = 0` — no MSI — so an HPET timer could never
|
||||
exercise this path. The first MSI driver will be the first PCI driver.
|
||||
|
||||
## M16 — the IOMMU, and the honest caveat ◑ detection done, enforcement pending
|
||||
## M16 — the IOMMU, and the honest caveat ✅ done
|
||||
|
||||
*The IOMMU is now detected (DMAR parsed, VT-d unit mapped and read — see the `iommu`
|
||||
test), but **enforcement is not built**: no translation domains are programmed, so the
|
||||
caveat below still holds in full. Detection can't be taken further usefully until there
|
||||
is a DMA driver to protect and QEMU's `intel-iommu` to test the protection against —
|
||||
building the per-device domains alongside that first driver is both the natural order
|
||||
and the only way to verify them. The rest of this section is the original caveat.*
|
||||
test) **and enforced**: `device_claim` programs a private per-device VT-d/AMD-Vi
|
||||
translation domain for the claimed function and rolls the claim back with `-ECONFINE` if
|
||||
it cannot confine it, `dma_alloc` buffers are bound into that domain and torn down at
|
||||
process death, and the machine fails open only when it has no IOMMU at all. So the caveat
|
||||
below no longer holds except on IOMMU-less hardware. The rest of this section is the
|
||||
original caveat, kept for the reasoning.*
|
||||
|
||||
Everything above is capability-gated at the *CPU*. None of it is gated at the *device*.
|
||||
A driver that can program a bus-mastering engine can make that device write to any
|
||||
physical address, because page tables sit between the CPU and RAM, not between a device
|
||||
and RAM. Until VT-d/DMAR (or SMMU on ARM) is programmed from the DMAR table, **`device_claim`
|
||||
on any DMA-capable device is equivalent to granting ring 0.**
|
||||
Everything above is capability-gated at the *CPU*. CPU page tables alone do not gate the
|
||||
*device*: a driver that can program a bus-mastering engine could make that device write to
|
||||
any physical address, because those page tables sit between the CPU and RAM, not between a
|
||||
device and RAM. That is exactly what the IOMMU closes. Now that VT-d/DMAR (and AMD-Vi; SMMU
|
||||
on ARM) is programmed, **`device_claim` confines the function into a private translation
|
||||
domain** and rolls the claim back with `-ECONFINE` if it cannot — so a claimed DMA-capable
|
||||
device is no longer equivalent to granting ring 0.
|
||||
|
||||
This does not make the model useless — it's the same position Linux is in with the
|
||||
IOMMU off, and every other guarantee (crash isolation, restart, no shared address
|
||||
space) still holds. But "user-space drivers are memory-safe" is not true yet, and the
|
||||
gap should be named rather than implied.
|
||||
This puts the model ahead of Linux-with-the-IOMMU-off: with an IOMMU present,
|
||||
"user-space drivers are memory-safe" now holds, alongside every other guarantee (crash
|
||||
isolation, restart, no shared address space). The one remaining gap — a machine with no
|
||||
IOMMU at all, where the system deliberately fails open — should be named rather than
|
||||
implied.
|
||||
|
||||
## Ordering
|
||||
|
||||
|
||||
Reference in New Issue
Block a user