re-org docs
This commit is contained in:
@@ -0,0 +1,344 @@
|
||||
# Native AMD GPU support — feasibility and roadmap
|
||||
|
||||
**Status: research snapshot, not implemented.** This records what a *minimal, display-only*
|
||||
native driver for a real discrete AMD GPU — specifically an **RX 6600-class card (Navi 23,
|
||||
RDNA2, DCN 3.0.2)**, the market analog of the RTX 3060 — would take, and how it slots into
|
||||
danos's pluggable scanout architecture. It is a survey of primary sources (the Linux
|
||||
[amdgpu Display Core](https://github.com/torvalds/linux/tree/master/drivers/gpu/drm/amd/display)
|
||||
driver and its [kernel documentation](https://docs.kernel.org/gpu/amdgpu/display/index.html),
|
||||
the AtomBIOS interpreter in
|
||||
[drivers/gpu/drm/amd](https://github.com/torvalds/linux/tree/master/drivers/gpu/drm/amd),
|
||||
linux-firmware's `LICENSE.amdgpu`, and Haiku's
|
||||
[radeon_hd](https://github.com/haiku/haiku/tree/master/src/add-ons/kernel/drivers/graphics/radeon_hd)),
|
||||
not an implementation. It completes the trilogy with [nvidia-gpus.md](nvidia-gpus.md) and
|
||||
[intel-igpu.md](intel-igpu.md) and should be read against both — AMD lands *between* them:
|
||||
NVIDIA-class discrete-card mechanics, but Intel-class (better, in one way) reference material.
|
||||
|
||||
Read [display.md](display.md) and [display-v2.md](display-v2.md) first — this doc assumes the
|
||||
v2 model where scanout is a **pluggable backend** and a native driver is just another `.scanout`
|
||||
service (like the virtio-gpu one), announcing to the compositor over `attach_scanout`.
|
||||
|
||||
## TL;DR
|
||||
|
||||
- **AMD's decisive advantage is that the vendor's own display driver is the register manual, and
|
||||
it's MIT-licensed.** The entire Display Core (DC) — hardware sequencer, per-block code for
|
||||
OTG/OPTC, HUBP, DPP, MPC, DIO/link encoders, plus the `asic_reg` register headers — ships in
|
||||
the Linux tree under MIT/X11, deliberately written OS-agnostic because AMD shares it across
|
||||
operating systems. You can study it, port it, even copy from it into a danos driver without
|
||||
license contamination. NVIDIA has no analog (nouveau is GPL); Intel has PRM prose but you
|
||||
still write the code yourself.
|
||||
- **The firmware wall is one small blob, not a GSP.** The only display-side firmware the Linux
|
||||
driver hard-requires is **DMCUB** (the display microcontroller), and only on **DCN 2.1
|
||||
through 4.x** — which includes Navi 23. It is redistributable from linux-firmware, and it is
|
||||
a display helper, not a full-card resource manager: on DCN 3.0.x, hardware init
|
||||
(`dcn30_init_hw`) is **host-driven direct register programming** — the only DMUB call in it
|
||||
is a capability query. All DCE generations, DCN 1.0 (Raven), and DCN 2.0 (Navi 10/12/14) run
|
||||
display with **no display firmware at all**.
|
||||
- **Whether the *silicon* (vs. the Linux driver) needs DMCUB for a bare GOP-inheriting modeset
|
||||
is unproven** — Linux fails init with `-EINVAL` if the blob is missing on a DMUB ASIC, but
|
||||
what it's *used for* at minimum scope (vs. PSR/ABM/offloaded DP link training) isn't
|
||||
documented. The safe plan ships the blob; it's legally and practically cheap to do so.
|
||||
- **Programming model is direct MMIO, not channel DMA.** DCN mode-set is ordered register-write
|
||||
sequences (the DC "hardware sequencer") against named, header-documented registers — no
|
||||
pushbuffers, no method streams, no RAMHT, no supervisor-interrupt handshake. This deletes the
|
||||
hardest structural layer of the NVIDIA path.
|
||||
- **Scanout is VRAM-only on discrete cards** — the claim that DCN can scan out of GTT/system
|
||||
memory was checked and *refuted* for dGPUs (Linux allows GTT scanout only on select APUs). So
|
||||
a small VRAM allocator + BAR CPU mapping is required, same as NVIDIA. Pitch-linear surfaces
|
||||
are supported; no DCC/tiling needed.
|
||||
- **danos's GOP boot helps here too, with a caveat.** DC explicitly models taking over a
|
||||
VBIOS/GOP-lit pipe (`dc_validate_boot_timing` reads back live DIG/OTG/pixel-clock state), so
|
||||
"repoint the surface on the running pipe" is demonstrably hardware-feasible — but Linux's
|
||||
seamless-boot path is **eDP-only and default-off on discrete cards**, so plan on a full
|
||||
self-owned modeset (including DP retrain) right after first light rather than living on the
|
||||
inherited link.
|
||||
- **AMD has real non-Linux prior art — but only for the old hardware.** Haiku's MIT `radeon_hd`
|
||||
mode-sets by executing VBIOS **AtomBIOS command tables** through AMD's own MIT interpreter;
|
||||
its compiled-in ceiling is **DCE 8.5 (Hawaii, ~2013)** — every Polaris/Vega/Navi entry sits
|
||||
in a `#if 0` block. There is zero non-Linux DCN precedent; a danos DCN driver would be first.
|
||||
- **Effort tier ≈ high-3 to 4** for a native DCN 3.0.x display-only driver on Navi 23 — the raw
|
||||
register surface is GA106-class (tier 4), but the MIT vendor reference, the one-blob firmware
|
||||
wall, and the absence of channel-DMA plumbing pull real risk out. The AtomBIOS-interpreter
|
||||
route is tier ≈ 3 but dead-ends at pre-2016 silicon.
|
||||
- **Recommendation:** the same sober conclusion as the other two docs — GOP already gives
|
||||
native-res scanout for zero code — but if danos ever does drive real discrete silicon
|
||||
natively, **an RDNA2 card is the best target of the three**: modern, mainstream, in-warranty
|
||||
hardware with a legally clean, vendor-authored reference. That combination exists nowhere
|
||||
else.
|
||||
|
||||
## The firmware wall (a fence, next to NVIDIA's wall)
|
||||
|
||||
AMD GPUs carry a zoo of firmware: PSP (security processor), SMU (power/clock management), CP/RLC
|
||||
(graphics), SDMA, VCN (media) — and, on the display side, DMCU (legacy) then **DMCUB**
|
||||
("Display Micro-Controller Unit, version B"), a per-generation blob in linux-firmware
|
||||
(`navi23_dmcub.bin` etc.). The display-only question is: which of these does a scanout driver
|
||||
actually need?
|
||||
|
||||
The Linux answer is precise and readable in
|
||||
[`amdgpu_dm.c`](https://github.com/torvalds/linux/blob/master/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm.c):
|
||||
`dm_init_microcode()` switches on the display IP version — **DCN 2.1 (Renoir) through DCN
|
||||
3.0.x / 3.1.x / 3.2 / 3.5 / 4.x** request a DMCUB blob as `AMDGPU_UCODE_REQUIRED`, and
|
||||
`dm_dmub_hw_init()` fails driver init with `-EINVAL` if it's absent. Everything earlier — **all
|
||||
of DCE (Southern Islands through Vega), DCN 1.0 (Raven), and DCN 2.0 (Navi 10/12/14)** — hits
|
||||
the `default:` case, `dmub_srv` stays NULL, and display runs with no display firmware at all
|
||||
([kernel display-manager doc](https://docs.kernel.org/6.2/gpu/amdgpu/display/display-manager.html)).
|
||||
|
||||
Two nuances survive verification:
|
||||
|
||||
- **The requirement is Linux-driver enforcement backed by real functional need, and it's
|
||||
version-sensitive.** AMD force-switched all Renoir ASICs to DMUB to fix a USB-C/resume bug
|
||||
(kernel commit `652de07addd2`, "with new dmub f/w dmcu is superseded"), which regressed users
|
||||
on old blobs and had to be patched with explicit `dmcub_fw_version` gating (`91adec9e0709`).
|
||||
What DMUB is *used for* varies with blob version. Ship a current blob.
|
||||
- **DMCUB is an architectural fixture, not a bolt-on** — the DCN hardware itself contains a DMU
|
||||
block housing the microcontroller
|
||||
([DCN overview](https://docs.kernel.org/gpu/amdgpu/display/dcn-overview.html)) — but it is
|
||||
**not a mediator of the programming model** on DCN 3.0: `dcn30_init_hw()` initializes clocks,
|
||||
disables power gating, and powers up link encoders via direct register writes; its sole DMUB
|
||||
interaction is `dc_dmub_srv_query_caps_cmd`. Firmware-*assisted* PHY/link bring-up appears
|
||||
from **DCN 3.1** onward — one more reason to target 3.0.x. Features like PSR and ABM are
|
||||
DMUB-offloaded on all generations; a minimal driver simply doesn't enable them.
|
||||
|
||||
**Contrast with NVIDIA's GSP:** the GSP is a full resource manager with a signed multi-stage
|
||||
boot chain and a firmware ABI that breaks every driver release. DMCUB is a display helper blob
|
||||
you copy onto the boot image once, load into a reserved buffer, and mostly ignore. There is no
|
||||
signature fuse-matching, no WPR carve-out, no RPC-only register access. The one genuinely open
|
||||
question — whether a GOP-inheriting minimal modeset could skip DMCUB entirely on DCN 3.0.2 —
|
||||
doesn't need answering, because shipping the blob costs nothing (see [Licensing](#licensing)).
|
||||
|
||||
**PSP/SMU remain the flagged risk.** Nothing display-only touches CP/RLC/SDMA (those gate the
|
||||
graphics rings, exactly like NVIDIA's PGRAPH — irrelevant here). But `dcn30_init_hw` calls into
|
||||
the clock manager, and on discrete cards the clock manager may message the SMU to change display
|
||||
clocks (DISPCLK/DPPCLK). Whether inherited GOP boot clocks suffice for a same-or-lower mode —
|
||||
avoiding SMU (and hence PSP firmware-load) entirely — is the largest unverified assumption in
|
||||
the milestone list below. The survey produced no confirmed claim either way.
|
||||
|
||||
## The display engine landscape
|
||||
|
||||
Two eras, one boundary that matters:
|
||||
|
||||
| Generation | Display IP | Cards | Display firmware | Route |
|
||||
|---|---|---|---|---|
|
||||
| GCN 1–4 (SI→Polaris) | DCE 6/8/10/11 | HD 7000 → RX 580 | none | AtomBIOS tables or direct DCE registers |
|
||||
| Vega / Raven | DCE 12 / DCN 1.0 | Vega 56/64, APUs | none | DC code (first DCN) |
|
||||
| Navi 1x (RDNA1) | DCN 2.0 | RX 5500–5700 | none | DC code |
|
||||
| Renoir APU | DCN 2.1 | 4000-series APUs | **DMCUB required** | DC code |
|
||||
| **Navi 2x (RDNA2)** | **DCN 3.0.x** | **RX 6600–6900** | **DMCUB required** | **DC code, host-driven init** |
|
||||
| RDNA3/RDNA4+ | DCN 3.1+/3.2/3.5/4.x | RX 7000/9000 | DMCUB required, fw-assisted PHY | DC code, more DMUB offload |
|
||||
|
||||
The best modern first-pixel target is **DCN 3.0.x**: it has the full MIT block stack from the
|
||||
June 2020 Sienna Cichlid patch series (207 patches, Linux 5.9; Navi 23 reuses the dcn30
|
||||
sequencer), host-driven hardware init, and sits *before* the DCN 3.1 shift toward
|
||||
firmware-assisted link management. Older DCE cards are even simpler (no firmware at all, plus
|
||||
the AtomBIOS escape hatch) but are 2013–2016 hardware; newer DCN 3.5/4.x pushes more into DMUB.
|
||||
|
||||
The DCN pipe, in one line each (the vocabulary the DC code speaks —
|
||||
[programming model](https://docs.kernel.org/next/gpu/amdgpu/display/programming-model-dcn.html)):
|
||||
**HUBP** fetches and unpacks the surface from memory (this is where the scanout address and
|
||||
pitch live), **DPP** scales/converts colors, **MPC** blends planes (bypassable for one plane),
|
||||
**OPP** packs output, **OTG/OPTC** generates raster timings (the CRTC), and the **DIO** block's
|
||||
DIG encoders + PHY drive the connector. Mode-set is the DC *hardware sequencer* walking these
|
||||
blocks with ordered register writes — plain MMIO with polling, no pushbuffer channels, no
|
||||
supervisor interrupts. Structurally this is Intel-shaped, not NVIDIA-shaped.
|
||||
|
||||
## Two routes: AtomBIOS interpreter vs. native DC-derived registers
|
||||
|
||||
**AtomBIOS** is AMD's VBIOS bytecode: every card's ROM carries *data tables* (connector
|
||||
topology, clock limits — the DCB equivalent) and *command tables* (`SetPixelClock`,
|
||||
`SetCRTC_Timing`, `EnableCRTC`, DIG encoder/transmitter control), executed by a small
|
||||
interpreter the driver embeds (`atom.c`, ~1.5k lines). The classic radeon driver and Haiku's
|
||||
`radeon_hd` mode-set this way: parse the tables, execute them, and the VBIOS does the
|
||||
register-level work for you — inherently per-board correct, since the tables come from the
|
||||
card's own ROM.
|
||||
|
||||
- **Where it's proven:** through DCE 8.5 (Haiku's ceiling, below) and in Linux's pre-DC code
|
||||
through Polaris (DCE 11.2). AMD's interpreter itself is MIT (Haiku ships AMD's own
|
||||
`atom.cpp`, "Copyright 2008 Advanced Micro Devices").
|
||||
- **Where it's unproven:** DCN. amdgpu's DC still *uses* AtomBIOS for init sub-steps
|
||||
(`bios_golden_init` in `dcn30_init_hw` executes host-interpreted tables) and reads the data
|
||||
tables for connector topology — but nobody drives a full DCN modeset from command tables, and
|
||||
whether RDNA2 VBIOSes still carry a complete modeset path or vestigial init-only tables is an
|
||||
open question no source answers. Do not bet on it.
|
||||
|
||||
**The native route** is: port the relevant slice of DC. Not wholesale — in-tree DC has
|
||||
accumulated Linux-isms (kernel-FPU guards around DML, the bandwidth-calculation library, which a
|
||||
single-plane fixed-mode driver can largely sidestep) — but the DC core is *designed* to be
|
||||
retargeted: the kernel docs state outright that DC "is shared with other OSes" and holds the
|
||||
OS-agnostic hardware programming behind a `dm_services` shim (register access, memory, delays,
|
||||
firmware loading). Dave Airlie initially rejected the DAL/DC merge in 2016 *because* it was
|
||||
AMD's cross-OS codebase — hostile-witness confirmation that this exact code runs outside Linux.
|
||||
Reimplement the shim in Zig, and the dcn30 sequences sit on top.
|
||||
|
||||
## The memory floor
|
||||
|
||||
Same shape as NVIDIA's, and the survey *hardened* one assumption:
|
||||
|
||||
- **VRAM-only scanout on discrete cards.** The documented DCN fetch path is VRAM → Data Fabric
|
||||
(SDP) → DCHUB → HUBP; the claim that display buffers can live in GTT/system memory was
|
||||
refuted for dGPUs in verification — Linux permits GTT scanout only on select APUs. The NVIDIA
|
||||
doc's "maybe sysmem ctxdma?" hope has a firm *no* here. Budget for a small VRAM allocator.
|
||||
- **Pitch-linear is fine.** HUBP programs a surface address + pitch; linear (untiled, no DCC)
|
||||
surfaces are first-class for scanout. No tiling math.
|
||||
- **CPU access via the VRAM BAR.** Compositing writes go through the PCI VRAM aperture;
|
||||
resizable BAR helps but isn't needed — one pitch-linear surface fits comfortably in a
|
||||
fixed 256 MB small-BAR window.
|
||||
- **No GPU VMM.** Display addresses are physical VRAM addresses programmed into HUBP; no page
|
||||
tables, no GEM/TTM, no eviction.
|
||||
|
||||
**Net:** (1) a contiguous aligned VRAM allocator, (2) a BAR CPU mapping, (3) a small reserved
|
||||
buffer for the DMCUB firmware regions. That's the whole memory story.
|
||||
|
||||
## Inheriting GOP state
|
||||
|
||||
danos's GOP boot pays off again, with sharper edges than on NVIDIA:
|
||||
|
||||
- **DC models pipe takeover explicitly.** `dc_validate_boot_timing()`
|
||||
([dc.c](https://github.com/torvalds/linux/blob/master/drivers/gpu/drm/amd/display/dc/core/dc.c))
|
||||
reads back *live* hardware — `is_dig_enabled` on the link encoder, OTG timing registers,
|
||||
pixel clock within tolerance — and keeps the VBIOS/GOP-lit pipe running until first flip.
|
||||
This is a vendor-blessed recipe for milestone 2 below: the exact register set that tells you
|
||||
which pipe is alive and how it's configured.
|
||||
- **But Linux's seamless path is eDP-only and APU-gated.** The code comment is blunt: "Support
|
||||
seamless boot on EDP displays only", and the enabling check requires an APU with DCN ≥ 3.0
|
||||
unless forced with `amdgpu.seamless=1`. On a discrete RX 6600 with DP/HDMI, Linux does a full
|
||||
modeset at takeover. Read that as a warning, not a prohibition: repointing HUBP at your own
|
||||
surface on the live pipe should still work (the readback code proves the state is
|
||||
inspectable), but plan the full self-owned modeset — **including DP retraining** — as the
|
||||
immediate next step, not a someday.
|
||||
- **No supervisor handshake exists to re-learn.** The NVIDIA doc's SV1/SV2/SV3 open question has
|
||||
no AMD counterpart; commit sequencing is ordered register writes + vblank/lock waits in the
|
||||
hwseq, all visible in MIT source.
|
||||
|
||||
## Licensing
|
||||
|
||||
The inverse of the NVIDIA situation, and the single strongest argument for AMD:
|
||||
|
||||
- **The reference code is MIT.** `amdgpu_dm.c` carries `SPDX-License-Identifier: MIT`;
|
||||
`dc/core/dc.c` and `dmub_srv.h` carry the full X11-style grant (use, copy, modify, merge,
|
||||
publish, distribute, sell). The `asic_reg` register headers ship under the same terms. One
|
||||
diligence note: SPDX tagging isn't uniform across the tree, so header-check each file before
|
||||
copying from it — but no GPL files are known inside `dc/`. Where nouveau forces a
|
||||
GPL-or-clean-room choice, here the *easy technical path and the permissive path are the same
|
||||
path*.
|
||||
- **Firmware redistribution is a solved problem.** linux-firmware's `LICENSE.amdgpu` grants
|
||||
anyone a royalty-free right to reproduce and distribute the blobs, binary-only, with the
|
||||
license text attached — no OSI-license gate like NVIDIA's, no AMD agreement needed. danos can
|
||||
ship `navi23_dmcub.bin` (and PSP/SMU blobs if ever needed) on its boot image today. The same
|
||||
license **prohibits reverse-engineering the blobs** — all programming knowledge must come
|
||||
from the MIT source, never from blob disassembly. (VBIOS images aren't in linux-firmware;
|
||||
they're read from the card's own ROM, as Haiku does.)
|
||||
- **Prose register docs are a DCE-era artifact.** AMD's classic X.Org-hosted PDFs cover the old
|
||||
families — and the famous `R6xx_3D_Registers.pdf` turns out to be 3D-only (verified: zero
|
||||
display content; the display material lives in the separate per-ASIC Register Reference
|
||||
Guides). For DCN there is **no prose display spec at all**: the MIT DC source plus the
|
||||
`asic_reg` headers *are* the register manual. Plan accordingly.
|
||||
|
||||
## Prior art
|
||||
|
||||
AMD, unlike NVIDIA, has genuine working non-Linux precedent — with a hard generational ceiling:
|
||||
|
||||
- **Haiku `radeon_hd`** (MIT, still in the tree): a real, shipping, from-scratch display driver
|
||||
that executes AtomBIOS command tables via AMD's own MIT interpreter. Verified ceiling:
|
||||
the last *enabled* device entry is **Hawaii (DCE 8.5, 0x67be)**; everything newer —
|
||||
Tonga/Fiji, Carrizo/Polaris, Vega/Raven, and every Navi/RDNA2 entry up to the RX 6900 XT —
|
||||
sits inside one `#if 0 /* disabled for R1/beta5 */` block under the comment "WARN: DCE
|
||||
versions below here get sketchy."
|
||||
- **AmigaOS/MorphOS RadeonHD drivers** (hdrlab): commercial non-Linux Radeon display drivers,
|
||||
again for the DCE era.
|
||||
- **FreeBSD** `drm-kmod`: a port of Linux amdgpu (DC and all), not independent prior art — but
|
||||
proof the DC codebase transplants.
|
||||
|
||||
**Nobody has driven DCN outside Linux-derived code.** A danos DCN 3.0.x driver would be a
|
||||
first — but a first with the vendor's MIT code as its map, which is a different proposition
|
||||
from nouveau-as-only-reference.
|
||||
|
||||
## Alternatives
|
||||
|
||||
| Option | What you get | The tradeoff |
|
||||
|---|---|---|
|
||||
| **Stay on GOP** (working today) | Native-res scanout, zero GPU code/firmware/maintenance | Resolution frozen at ExitBootServices; no runtime mode change, no hardware vsync, no multihead |
|
||||
| **AtomBIOS interpreter on an old DCE card** | Proven end-to-end (Haiku); interpreter is small + MIT; board-correct by construction | 2013–2016 hardware ceiling; tier ≈ 3; teaches AtomBIOS, not modern DCN |
|
||||
| **Native DCN 3.0.x on RX 6600** (this doc) | Runtime modeset, vsync, multihead on modern silicon; MIT vendor reference; one redistributable blob | Tier ≈ high-3–4; SMU/clock question open; no non-Linux precedent |
|
||||
| **Port DC wholesale** (reimplement `dm_services`) | Vendor-maintained sequences verbatim; designed-for-porting seam | Big codebase to carry (DML, abstractions); Linux-isms to shear off; overkill for one plane |
|
||||
| **RDNA3+/DCN 3.5+** | Newer cards | More DMUB offload (fw-assisted PHY from DCN 3.1); strictly harder than 3.0.x for no display-only gain |
|
||||
|
||||
## "First light" milestones (native DCN 3.0.x path)
|
||||
|
||||
Framed as a danos `.scanout` service, inheriting the GOP-initialized display:
|
||||
|
||||
1. **PCI/BAR bring-up** — enumerate Navi 23, map the register BAR and the VRAM BAR via danos
|
||||
MMIO grants; prove the pipe is GOP-live by writing pixels into the *existing* GOP
|
||||
framebuffer through the VRAM BAR.
|
||||
2. **Read back the live pipe** — port the `dc_validate_boot_timing` register set: which OTG is
|
||||
running, its timings, which DIG/link encoder is enabled, current HUBP surface address/pitch.
|
||||
This is pure reads — zero risk, high information.
|
||||
3. **Repoint the surface** — allocate a danos-owned pitch-linear VRAM surface, program the HUBP
|
||||
surface address/pitch on the live pipe at vblank. First self-owned pixel with **no modeset,
|
||||
no firmware, no clock changes**.
|
||||
4. **DMCUB bring-up** — load `navi23_dmcub.bin` (redistributed per `LICENSE.amdgpu`) into its
|
||||
reserved regions, minimal `dmub_srv` init, verify the caps query answers.
|
||||
5. **Full owned modeset** — port the dcn30 hwseq slice: OTG timing programming, MPC bypass
|
||||
(single plane), DIG/PHY enable, **DP link retrain** (or start on HDMI to defer it, exactly
|
||||
as the NVIDIA doc advises). This is where the SMU/clock question lands — first attempt:
|
||||
reuse inherited boot clocks for a same-or-lower mode.
|
||||
6. **EDID** — AUX (DP) / DDC (HDMI) over the DCN AUX engine registers; parse and build the mode
|
||||
list; connector topology from the VBIOS AtomBIOS data tables.
|
||||
7. **Wire into the compositor** — `attach_scanout`, vsync from the vblank/pageflip interrupt,
|
||||
then multihead.
|
||||
|
||||
Keep the GOP backend as the fallback the whole way — a stall at any step still leaves danos with
|
||||
a working display (the resilience v2 already provides via re-attach).
|
||||
|
||||
## Reading list
|
||||
|
||||
**The DC core (MIT — the register manual for DCN):**
|
||||
- `drivers/gpu/drm/amd/display/dc/hwss/dcn30/dcn30_hwseq.c` — hardware init + the modeset
|
||||
sequencer for the target generation (host-driven; one DMUB caps query).
|
||||
- `dc/dcn30/` + `dc/dcn302/` blocks: `dcn30_hubp.c` (surface address/pitch — milestone 3),
|
||||
`dcn30_optc.c` (OTG timings), `dcn30_dio_link_encoder.c` (DIG/PHY), `dcn30_mpc.c` (bypass),
|
||||
`clk_mgr/dcn30/` (the SMU question, read before milestone 5).
|
||||
- `dc/core/dc.c` — `dc_validate_boot_timing()`: the GOP-takeover readback recipe.
|
||||
- `asic_reg/dcn/dcn_3_0_0_{offset,sh_mask}.h` — every register name and bitfield.
|
||||
- `dmub/` (`dmub_srv.h`, `src/dmub_dcn30.c`) — firmware regions + bring-up for milestone 4.
|
||||
|
||||
**The Linux glue (for logic, not porting):**
|
||||
[`amdgpu_dm.c`](https://github.com/torvalds/linux/blob/master/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm.c)
|
||||
— `dm_init_microcode` / `dm_dmub_hw_init` (the firmware-wall switch), seamless-boot gating.
|
||||
|
||||
**Kernel docs (read first):**
|
||||
[DCN overview](https://docs.kernel.org/gpu/amdgpu/display/dcn-overview.html) (the block diagram
|
||||
+ VRAM→DF→DCHUB fetch path) ·
|
||||
[DC programming model](https://docs.kernel.org/next/gpu/amdgpu/display/programming-model-dcn.html)
|
||||
(dc_plane/dc_stream/dc_link objects, hwseq, block APIs) ·
|
||||
[display manager](https://docs.kernel.org/6.2/gpu/amdgpu/display/display-manager.html).
|
||||
|
||||
**AtomBIOS:** `drivers/gpu/drm/amd/amdgpu/atom.c` (the interpreter), `atombios.h` (table
|
||||
formats), [osdev AMD AtomBIOS](https://wiki.osdev.org/AMD_Atombios) (hobby-OS orientation),
|
||||
Haiku [`radeon_hd`](https://github.com/haiku/haiku/tree/master/src/add-ons/kernel/drivers/graphics/radeon_hd)
|
||||
(a complete worked example, MIT, through DCE 8.5).
|
||||
|
||||
**Licensing:** linux-firmware
|
||||
[`LICENSE.amdgpu`](https://github.com/endlessm/linux-firmware/blob/master/LICENSE.amdgpu);
|
||||
DCE-era prose specs at [x.org/docs/AMD](https://www.x.org/docs/AMD/) (Register Reference
|
||||
Guides — display; note `R6xx_3D_Registers.pdf` is 3D-only).
|
||||
|
||||
## Open questions (unresolved by the survey)
|
||||
|
||||
- **Does DCN 3.0.2 silicon need DMCUB for a bare inherit-and-modeset path**, or only for
|
||||
PSR/ABM/offloaded features? (Moot if the blob is shipped regardless — but it decides whether
|
||||
milestone 3 can precede milestone 4.)
|
||||
- **Can a display-only driver avoid PSP and SMU entirely** by inheriting GOP boot clocks — what
|
||||
does the dcn30 clock manager actually require from SMU messaging on Navi 23 for a
|
||||
same-or-lower mode? *The largest open risk in the plan.*
|
||||
- **Do RDNA2 VBIOS command tables still carry a complete modeset path**, or are they vestigial
|
||||
init-only tables? (Would open a Haiku-style route on modern cards; no source answers it.)
|
||||
- **Is the DCN 3.0 AUX/DDC engine and DP retrain fully host-drivable without DMUB**, as
|
||||
`dcn30_hwseq` implies?
|
||||
- Exact HUBP surface alignment/pitch constraints for linear scanout on Navi 23 (in the headers;
|
||||
not captured verbatim in the survey).
|
||||
|
||||
---
|
||||
|
||||
*Research snapshot (2026-07); findings pinned to Linux master and Haiku master as of the survey
|
||||
date. DMUB coverage only grows with new DCN generations — re-verify the firmware-wall switch in
|
||||
`amdgpu_dm.c` against current source before building.*
|
||||
@@ -0,0 +1,210 @@
|
||||
# Device interrupts
|
||||
|
||||
CPU exceptions ([interrupts.md](../os-development/interrupts.md)) are the kernel reacting to its own
|
||||
mistakes. **Device interrupts** are the opposite: hardware asking for attention —
|
||||
a timer firing, a key pressed, a packet arriving. They share the IDT, but differ
|
||||
in one fundamental way: an exception here is terminal (we report and halt), while a
|
||||
device interrupt is *handled and returned from*, so the interrupted code resumes as
|
||||
if nothing happened. This is danos's first code that takes an interrupt and comes
|
||||
back — the same mechanism a scheduler will later use to preempt tasks.
|
||||
|
||||
The first device we bring up is the **timer**, because it's the simplest: it lives
|
||||
entirely on the CPU's local interrupt controller, needing no external routing.
|
||||
It's all x86_64-specific, behind the [architecture](../os-development/architecture.md) boundary.
|
||||
|
||||
## The APIC, not the PIC
|
||||
|
||||
Interrupt delivery on modern x86 goes through the **APIC**, not the legacy 8259
|
||||
PIC. There are two halves; we only need one so far:
|
||||
|
||||
- The **Local APIC** (per-CPU, memory-mapped at physical `0xFEE00000`) handles the
|
||||
CPU's own timer and receives interrupts routed to it. `system/kernel/architecture/x86_64/apic.zig`.
|
||||
- The **IO-APIC** routes *external* device lines (keyboard, etc.) to LAPIC vectors.
|
||||
Not needed for the timer — it'll arrive with the keyboard.
|
||||
|
||||
The old PIC has to be dealt with first, though: left alone it would deliver
|
||||
interrupts on vectors `0x08-0x0F`, which **collide with the CPU exception
|
||||
vectors** — a spurious IRQ would look like a double fault. So `init` remaps the
|
||||
PIC's vectors to `0x20-0x2F` and masks every line, taking it out of the picture.
|
||||
|
||||
Then the LAPIC is enabled in two places: the `IA32_APIC_BASE` MSR's global-enable
|
||||
bit, and the LAPIC's own spurious-vector register (bit 8 = software enable). The
|
||||
spurious vector is `0x2F` — low nibble `F` by convention, and inside our gate
|
||||
range so a stray spurious interrupt lands on a valid no-op.
|
||||
|
||||
## The timer
|
||||
|
||||
The LAPIC timer is three register writes (`initTimer`): a divide setting, then the
|
||||
LVT-timer entry giving it a **vector** (32) and **periodic** mode, then an initial
|
||||
count that becomes the reload value. From then on it fires vector 32 repeatedly, on
|
||||
its own, forever.
|
||||
|
||||
The reload count isn't picked arbitrarily — it's **calibrated to real time**,
|
||||
which the [real-time](../vision.md) scheduling guarantees depend on. Since the LAPIC
|
||||
timer's raw rate is bus-clock dependent and unknown up front, `calibrate` runs the
|
||||
LAPIC timer one-shot from its maximum count while a **reference clock** counts out a
|
||||
known 10 ms, then sees how far the LAPIC got — its counts-per-millisecond, from which
|
||||
`initTimer(hz)` computes the reload count for any target frequency. danos runs it at
|
||||
**1000 Hz** (a 1 ms tick).
|
||||
|
||||
The reference clock is chosen in order of preference, so danos calibrates on
|
||||
legacy-free **UEFI Class 3** hardware where the old 8254 PIT may be *absent* (polling
|
||||
a missing PIT would hang the boot):
|
||||
|
||||
1. **CPUID leaf 0x15** — the CPU's TSC frequency directly, needing no external timer
|
||||
at all (the LAPIC is then measured against the TSC).
|
||||
2. The **HPET**, discovered via ACPI (see [discovery](../os-development/discovery.md) / [acpi](../os-development/acpi.md)).
|
||||
3. The **ACPI PM timer** (a fixed 3.579545 MHz counter from the FADT).
|
||||
4. The **PIT** (legacy 8254, 1.193182 MHz) — last resort, and bounded so it can't hang.
|
||||
|
||||
All four yield the same rate; on QEMU (no CPUID crystal enumeration) it lands on the
|
||||
HPET, matching the PIT numbers to within measurement jitter.
|
||||
|
||||
## The high-resolution clock (TSC)
|
||||
|
||||
The timer tick gives *scheduling* — a 1 ms quantum — but 1 ms is coarse for a
|
||||
real-time system to *measure* with (interrupt latency, jitter, timeouts). So the
|
||||
same calibration also measures the **TSC** (Time Stamp Counter): a per-core cycle
|
||||
counter read with `rdtsc` in a couple of cycles, giving roughly **nanosecond**
|
||||
resolution — a million times finer than the tick. We snapshot the TSC across the
|
||||
same 10 ms calibration window to get its frequency (measured ~1 GHz under QEMU).
|
||||
|
||||
The monotonic clock is exposed as one function per resolution — `nanos()`,
|
||||
`micros()`, `millis()` — each scaling the cycle delta directly at its unit (with a
|
||||
128-bit intermediate so a long uptime doesn't overflow) rather than chaining
|
||||
divisions. `millis()` is what the scheduler uses for `sleep` deadlines; `nanos()`
|
||||
is there for fine measurement. Note the two clocks are distinct: the **tick** drives
|
||||
preemption and wakeups (1 ms granularity); the **TSC** is the resolution you read
|
||||
time at. Making `sleep` itself sub-millisecond would take a tickless one-shot
|
||||
timer — a later step.
|
||||
|
||||
### Is the TSC trustworthy? Invariant, and synchronized
|
||||
|
||||
A cycle counter is only a valid *clock* if two things hold, and danos checks both,
|
||||
because they decide whether we read time with a cheap `rdtsc` or fall back to the HPET.
|
||||
|
||||
**Invariant.** An old TSC counted core clock cycles, so it sped up and slowed down with
|
||||
frequency scaling — useless as wall time. Modern CPUs (all of danos's targets) provide an
|
||||
**invariant TSC**: a constant rate across P/C-states that never stops. The guarantee is a
|
||||
CPUID bit — leaf `0x80000007`, EDX bit 8 — on both Intel *and* AMD. danos reads it in
|
||||
`calibrate`, and a TSC that doesn't advertise it is demoted to the HPET clocksource —
|
||||
provided a usable HPET exists (64-bit; a 32-bit one wraps too fast to stay monotonic).
|
||||
With no such fallback the TSC stays, there being nothing steadier to switch to. AMD is
|
||||
why this matters in practice: it doesn't populate the Intel leaf `0x15` that enumerates
|
||||
the TSC *frequency*, so danos already measures AMD's rate against the HPET — but a
|
||||
measured frequency without the invariance guarantee is not enough.
|
||||
|
||||
**Synchronized.** Each core has its own TSC. Even invariant ones can start at different
|
||||
values (a second socket, some firmware), so a thread migrating from a core reading
|
||||
`1_000_000` to one reading `999_000` would see time jump *backward*. danos runs a **warp
|
||||
check** as each application processor comes online (`checkWarpSource`, adapted from
|
||||
Linux's): the waking core and the BSP hammer a shared "highest seen" TSC under a lock,
|
||||
and if either ever reads below it, the cores' TSCs are skewed. It's pairwise because APs
|
||||
come up one at a time ([smp.md](../os-development/smp.md)).
|
||||
|
||||
**The fallback.** When the TSC fails either test — non-invariant (a bare VM such as the
|
||||
default qemu64), or warped between cores — danos moves the monotonic clock onto the
|
||||
**HPET** main counter: one fixed-rate counter, so it can neither skew between cores nor
|
||||
drift with frequency. It costs a memory-mapped read instead of a register read, but it
|
||||
keeps time *accurate*, which is the whole point. The switch preserves the current value,
|
||||
so the clock never jumps. The boot log names the outcome:
|
||||
|
||||
```
|
||||
/system/kernel: clocksource tsc (TSC invariant: yes, synchronized: yes) # real Intel/AMD
|
||||
/system/kernel: clocksource hpet (TSC invariant: no, synchronized: yes) # a bare VM (TCG)
|
||||
```
|
||||
|
||||
## Two kinds of vector, one dispatch
|
||||
|
||||
The IDT now installs gates `0-47`: the 32 exceptions plus the device range. Every
|
||||
gate still funnels through the same stub tail (`isr_common`), which calls one
|
||||
dispatcher that branches on the vector (`interruptDispatch` in `idt.zig`):
|
||||
|
||||
```zig
|
||||
if (state.vector < 32) {
|
||||
on_fault(state); // exception: report and halt (never returns)
|
||||
} else if (handlers[state.vector]) |handler| {
|
||||
handler(); // device: run the registered handler
|
||||
}
|
||||
// else: spurious/unhandled — deliberately no EOI
|
||||
```
|
||||
|
||||
(A third branch has since joined for user mode, elided here:
|
||||
`state.vector == system_call_vector` (128) hands the trap frame to the ring-3
|
||||
syscall handler.)
|
||||
|
||||
Two things make device interrupts *return* where exceptions don't:
|
||||
|
||||
1. **The handler returns.** The timer handler just bumps a tick counter. Control
|
||||
flows back to `isr_common`, which restores every register it saved and executes
|
||||
`iretq` — resuming the interrupted instruction exactly. (This is why the stub
|
||||
saves *all* the general registers.)
|
||||
2. **End-of-interrupt.** Somewhere in there we write the LAPIC's EOI register. Miss
|
||||
this and the LAPIC thinks we're still busy and never delivers the next
|
||||
interrupt. It's the single most common "my timer fired once and stopped" bug.
|
||||
|
||||
**Each handler issues its own EOI**, rather than the dispatcher doing it around the
|
||||
call. That looks like a needless devolution while the timer is the only device, and
|
||||
`apic.timerTick` indeed does nothing but `eoi()` before bumping its counter (early,
|
||||
because the tick hook is the scheduler, which may switch tasks and not return
|
||||
promptly — the LAPIC mustn't wait on it).
|
||||
|
||||
It stops looking needless with the second device. A *routed* interrupt — one arriving
|
||||
through the I/O APIC from a real device line — must be **masked before it is
|
||||
acknowledged**, because a level-triggered line is still asserted at EOI time and would
|
||||
redeliver instantly, forever. Only the handler knows which discipline its source
|
||||
needs, so only the handler can sequence it. See [drivers.md](drivers.md), where the
|
||||
device is quieted by a driver in ring 3, long after the ISR has returned.
|
||||
|
||||
A device handler is a plain `fn () void` — a timer or keyboard handler doesn't need
|
||||
the interrupted registers. (The stubs originally didn't save the SSE/vector
|
||||
registers, so a handler couldn't use them; `isr_common` now does an
|
||||
`fxsave`/`fxrstor` of the full SSE/x87 state around dispatch — see
|
||||
[interrupts.md](../os-development/interrupts.md).)
|
||||
|
||||
## Turning them on
|
||||
|
||||
Exceptions can't be masked, which is why they worked all along. Maskable device
|
||||
interrupts don't fire until the CPU's interrupt flag is set — so the final step is
|
||||
`sti` (`arch.enableInterrupts()`), after the APIC and timer are configured. From
|
||||
that instant the kernel has a heartbeat, and its idle `hlt` loop
|
||||
([halting.md](../os-development/halting.md)) wakes on every tick and dozes off again.
|
||||
|
||||
## Verifying it
|
||||
|
||||
The `timer` test (see [testing.md](../testing.md)) is the proof that an interrupt both
|
||||
*fires* and *returns*: it records the tick count, busy-waits, and checks the count
|
||||
advanced on its own.
|
||||
|
||||
```
|
||||
$ python3 test/qemu_test.py timer
|
||||
timer ... PASS (matched 'DANOS-TEST-RESULT: PASS')
|
||||
```
|
||||
|
||||
If the APIC weren't enabled, or `sti` were missing, or EOI were forgotten, the
|
||||
count would stay put and the test would fail. That it advances — while the CPU was
|
||||
spinning in unrelated code — is the whole mechanism working end to end.
|
||||
|
||||
## Since (done elsewhere)
|
||||
|
||||
- **Preemption**: the timer handler is where the scheduler decides to switch — the
|
||||
reason a *returning* interrupt matters. See [scheduling.md](../os-development/scheduling.md).
|
||||
- **`sleep()` / timeouts** built on the calibrated clock.
|
||||
- **The I/O APIC, routed**: external device lines now reach a vector, and the
|
||||
interrupt is delivered onward to a *user-space* driver as an IPC message. See
|
||||
[drivers.md](drivers.md).
|
||||
- **Uncacheable MMIO**: device grants are mapped `PCD|PWT` (strong-uncacheable) for
|
||||
user drivers — see [paging.md](../os-development/paging.md).
|
||||
|
||||
## What's next (partly done since)
|
||||
|
||||
- **The keyboard** — done, exactly as sketched: the PS/2 bus driver
|
||||
(`system/drivers/ps2-bus/`) claims the port-mapped 8042 controller through the
|
||||
claim-gated `io_read`/`io_write` syscalls ([drivers.md](drivers.md)), binds
|
||||
IRQ 1 (and the aux mouse's IRQ 12), reads scancodes from `0x60`, and decodes
|
||||
them into HID events for the [input service](input.md).
|
||||
- **MSI-X** — still open: `msi_bind` gives one per-device edge-triggered vector
|
||||
(M15); MSI-X's multi-vector table (many queues per device, e.g. NVMe) is the
|
||||
remaining extension.
|
||||
- **The LAPIC's own page** — still mapped writeback-cacheable like the rest of
|
||||
the identity map. QEMU tolerates it; real hardware wants it uncacheable.
|
||||
@@ -0,0 +1,177 @@
|
||||
# The device manager
|
||||
|
||||
**Status: the protocol and supervision are built** (M18.1, 2026-07-13): `hello`
|
||||
with its deadline, supervised spawn, restart with backoff, and the crash-loop
|
||||
cap are in — usb-xhci-bus is the first conforming driver, and the
|
||||
`driver-restart` scenario proves fault → backoff → re-claim → cap end to end.
|
||||
Tree reports are built too (M18.2, 2026-07-13): the xHCI driver scans its
|
||||
root-hub ports and reports each connected device (`child_added`); the manager
|
||||
mirrors them and prunes a dead reporter's children, and the `usb-report`
|
||||
scenario proves report → prune → respawn → re-report. The application surface is built (M18.3, 2026-07-13):
|
||||
`enumerate` and `subscribe` over IPC, with `device-list` as the first client —
|
||||
the manager is now the one answer to "what devices exist" for applications.
|
||||
The primitives underneath are real ([process-management.md](../os-development/process-management.md):
|
||||
spawn/supervise/kill/exit-notification; [driver-model.md](driver-model.md): the device
|
||||
table as a capability system; [drivers.md](drivers.md): claim/map/IRQ), and the first
|
||||
per-device driver spawn works (the device manager matches the xHCI controller by PCI
|
||||
class and spawns `usb-xhci-bus` with the device id as argv[1]). This document designs
|
||||
the rest: the device manager as **the tree, the matcher, and the supervisor** — the
|
||||
policy process that turns [resilience.md](../os-development/resilience.md)'s restart goal into practice
|
||||
for drivers.
|
||||
|
||||
How processes stop, reload, and report their deaths is deliberately **not** in this
|
||||
document: that is the universal lifecycle every danos process speaks —
|
||||
[process-lifecycle.md](../os-development/process-lifecycle.md), signals over IPC and the stable
|
||||
`process` interface. The device manager is that design's first serious
|
||||
customer, not its owner. Its own protocol contains nothing lifecycle-shaped; a
|
||||
driver is stopped, health-checked, and buried exactly like any other process.
|
||||
|
||||
## The tree: structure in the manager, authority in the kernel
|
||||
|
||||
The device tree is two things fused: *information* (what exists, how it nests) and
|
||||
*authority* (a descriptor is a licence to map physical memory). They separate:
|
||||
|
||||
- The **kernel keeps the capability system** — device, I/O-port, and interrupt
|
||||
claims, resource containment on `device_register`, the
|
||||
`mmio_map`/`irq_bind`/`msi_bind` gates — and **cleans all of it up when a process
|
||||
dies** (settled; it is increment 1 of
|
||||
[process-lifecycle.md](../os-development/process-lifecycle.md)). The three invariants in
|
||||
[driver-model.md](driver-model.md) stay exactly where they are. A device manager
|
||||
that could mint MMIO mappings by its own say-so would be a second kernel, and a
|
||||
buggy one would un-earn everything the microkernel bought.
|
||||
- The **device manager owns the tree as data** — identity, topology, naming, driver
|
||||
matching, hotplug events, and being the one process everything else asks about
|
||||
devices. Firmware discovery seeds it (today via the kernel's snapshot); **bus
|
||||
drivers grow it** by reporting what they see; applications query and watch it.
|
||||
`device_enumerate` fades to a manager-internal (then deleted) seam.
|
||||
|
||||
Long-term, discovery itself leaves the kernel — but not *into* the manager. PCI
|
||||
enumeration is a **pci-bus driver**: the manager spawns it against the host bridge
|
||||
(already a device with the ECAM window as a resource), it scans, it reports functions
|
||||
like any bus reports children. ACPI becomes an **acpi service** that interprets the
|
||||
tables and reports the namespace. The manager only orchestrates and merges. Moving
|
||||
AML interpretation out of ring 0 is its own project on its own track; nothing here
|
||||
depends on when it lands. (It landed: [discovery.md](../os-development/discovery.md), M19–M20.)
|
||||
|
||||
`device_register` is **idempotent on exact match**: a re-registration with an
|
||||
identical (parent, class, identity, resources) tuple returns the existing id
|
||||
instead of appending a duplicate. The kernel table has no unregister, so without
|
||||
this a restarted registering bus would re-report its children as fresh nodes on
|
||||
every respawn. Idempotence is what makes restart-and-re-report sound for *every*
|
||||
reporting bus — pci-bus, the acpi service, a future fdt service — not just one,
|
||||
and it is why supervision (below) can prune a dead bus's subtree and trust the
|
||||
restarted instance to rebuild exactly the same ids.
|
||||
|
||||
## The protocol
|
||||
|
||||
A `device-manager-protocol` module (the vfs-protocol pattern): extern-struct
|
||||
messages, a version in the handshake, reserved fields everywhere. The manager is a
|
||||
well-known endpoint (`ipc.register(.device_manager)`); the badge tells it who is
|
||||
talking; the same endpoint receives its children's exit notifications — one loop,
|
||||
one world.
|
||||
|
||||
| Direction | Message | Purpose |
|
||||
|---|---|---|
|
||||
| driver → manager | `hello { version, role, device_id }` | confirms the argv assignment, starts the deadline clock |
|
||||
| bus → manager | `child_added { parent, bus_address, identity, device_id, hid }` | one node the bus discovered |
|
||||
| bus → manager | `child_removed { parent, bus_address }` | unplug, or the bus lost it |
|
||||
| app → manager | `enumerate` | snapshot of the tree (read-only) |
|
||||
| app → manager | `subscribe` | receive published add/remove events |
|
||||
|
||||
`hello` is the one deadline the manager enforces itself: spawned and silent past the
|
||||
deadline means wrong binary, wrong protocol version, or wedged before main — apply
|
||||
the stop sequence and the restart policy. Everything else lifecycle-shaped
|
||||
(terminate, the common `ping` liveness call, exit reasons) arrives through
|
||||
[process-lifecycle.md](../os-development/process-lifecycle.md)'s vocabulary, not this protocol.
|
||||
|
||||
Assignment stays argv (`usb-xhci-bus <device id>`) for now — simple, and it works.
|
||||
The step after `hello` exists is delegation: the manager claims (or is granted) the
|
||||
devices and passes the claim to the driver over IPC (the M13 capability-transfer
|
||||
mechanism), replacing first-come-first-served `device_claim` with policy. Identity in
|
||||
`child_added` is per-bus: PCI children carry the class triple (`pci_class`, as the
|
||||
xHCI match already uses); USB children carry the (class, subclass, protocol) triple
|
||||
from usb-ids.zig — each bus's native language, decoded by the shared ids modules.
|
||||
|
||||
## Supervision and restart
|
||||
|
||||
Every driver is spawned with the manager's exit endpoint (`spawnSupervised` — built).
|
||||
On a death notification:
|
||||
|
||||
1. **Read the reason** ([process-lifecycle.md](../os-development/process-lifecycle.md) increment 2).
|
||||
Clean exit → it meant to; don't restart. Fault or missed `hello` deadline →
|
||||
restart with **backoff**, and a crash-loop cap (three fast deaths → mark failed,
|
||||
stop respawning, log loudly; a later `reload` to the manager can retry).
|
||||
2. **Prune the subtree** the dead bus driver reported. Its children describe
|
||||
protocol state (xHCI slot ids, transfer rings) that died with the process;
|
||||
keeping the nodes would be keeping a lie. Watchers receive `child_removed` — the
|
||||
input service losing, then regaining, a keyboard is the *honest* description of
|
||||
what happened. The restarted instance rediscovers and re-reports.
|
||||
3. **The claim is already free** because the kernel released it at death — the
|
||||
restarted instance claims the same controller and comes up.
|
||||
|
||||
Who supervises the supervisor: **init** (PID 1), which already supervises the
|
||||
services it starts. If the manager dies, drivers keep running (they hold their
|
||||
claims; the kernel doesn't care who their supervisor was — though their exit
|
||||
notifications now dangle harmlessly). The restarted manager re-learns the world:
|
||||
kernel snapshot, then a re-`hello` round — drivers answer a broadcast or are stopped
|
||||
and respawned. Full state handoff is deliberately not attempted.
|
||||
|
||||
## Thin drivers, class protocols
|
||||
|
||||
The [driver-model.md](driver-model.md) three-shape split, restated as processes:
|
||||
|
||||
- A **bus driver** (usb-xhci-bus) owns its controller — claim, MMIO, IRQ/MSI, DMA
|
||||
rings — and offers a *transfer* protocol ("submit a control transfer to device N",
|
||||
built from the usb-abi request constructors) plus tree reports to the manager.
|
||||
- A **class driver** (usb-hid, usb-storage) owns nothing: it is matched to a reported
|
||||
child by its identity triple, speaks the bus's transfer protocol downward and its
|
||||
service's protocol upward — HID reports to the input service, blocks to the block
|
||||
service. It works unchanged over any controller.
|
||||
- **Services** (input, display, block) aggregate class drivers and face applications.
|
||||
|
||||
Each arrow is a protocol module. The manager routes none of the data plane — it
|
||||
introduces the parties (matching), supervises them (lifecycle), and gets out of the
|
||||
way.
|
||||
|
||||
## Increments
|
||||
|
||||
Increments 1–4 are the lifecycle prerequisites and live in
|
||||
[process-lifecycle.md](../os-development/process-lifecycle.md) (claim cleanup on death, exit reasons,
|
||||
published exit events, signals + `process`). On top of those:
|
||||
|
||||
5. **device-manager-protocol**: `hello`, supervised spawn with restart policy;
|
||||
usb-xhci-bus becomes the first conforming driver.
|
||||
6. **Tree reports**: `child_added`/`child_removed`; the manager mirrors; xHCI reports
|
||||
the mouse and keyboard QEMU already hangs off it.
|
||||
7. **App surface**: `enumerate`/`subscribe` over IPC; `device_enumerate` retreats
|
||||
to a manager-internal seam.
|
||||
8. **Discovery migration** — DONE (M19–M20, 2026-07-13): enumeration moved to
|
||||
ring 3 as swappable per-firmware discoverers — the pci-bus driver (M19) then
|
||||
the acpi service (M20), see [discovery.md](../os-development/discovery.md); of the enumerable
|
||||
devices, the kernel seeds only the host bridge and the acpi-tables node (the
|
||||
non-enumerable platform nodes — processors, interrupt controllers, the HPET,
|
||||
the loader's framebuffer — stay kernel-seeded too). Matching moved with it:
|
||||
`child_added` grew a `device_id` (the kernel-registered id, `no_device` for
|
||||
unregistered leaves like USB ports) and a firmware `hid`, and the manager now
|
||||
matches drivers from those **reports** rather than its boot-time snapshot. The
|
||||
PCI arm flipped in M19.3, the ACPI arm (ps2-bus matched from `_HID`) in M20.3
|
||||
— each in a single phase so no device is ever matched from both sources at
|
||||
once. The acpi service reports only the non-PCI `_HID` devices, since pci-bus
|
||||
already reports PCI functions (M20.2).
|
||||
|
||||
## Settled questions (2026-07-12)
|
||||
|
||||
- **Stateful buses**: pruning the subtree on bus-driver death is right for USB. A
|
||||
future storage bus with in-flight writes wants drain-before-terminate — which is
|
||||
exactly the `deadline_ms` parameter `stop()` already has; a per-driver deadline
|
||||
is one value in the manager's policy table when such a bus arrives. No design
|
||||
change.
|
||||
- **Manager death**: drivers survive the manager; the restarted manager re-learns
|
||||
the world (above). Checkpointing driver state with the manager is deferred until
|
||||
something demonstrates the need.
|
||||
- **Matching stays code until the third bus.** `driverFor`/`pciDriverFor` were
|
||||
honest at two bus types; the third was expected to trigger the manifest (a driver
|
||||
declares what it binds: a PCI class triple, a USB class triple, an ACPI `_HID`).
|
||||
(Since then: the third bus — USB — arrived and is matched in code too. Today's
|
||||
matchers are `pciDriverForIdentity`, `hidDriverFor`, and `usbDriverForIdentity`;
|
||||
the manifest waits until code matching actually hurts.)
|
||||
@@ -0,0 +1,182 @@
|
||||
# Display service — build plan (v1: the dumb-framebuffer compositor)
|
||||
|
||||
The ordered, checkpointable build-out for [display.md](display.md). Each milestone is
|
||||
small, lands on its own, and ends in a **verifiable gate** — shaped for a `/loop` run.
|
||||
Read [display.md](display.md) first for the *why*; this is the *what* and the *order*.
|
||||
|
||||
## Locked decisions (do not relitigate)
|
||||
|
||||
- **Handoff = device node + write-combining `mmio_map`.** The kernel seeds a synthetic
|
||||
display node (found by class, not name) from `BootInformation.framebuffer`; the
|
||||
service claims + WC-maps it.
|
||||
(Not a bespoke `framebuffer_map` syscall — the device route inherits ownership,
|
||||
release-on-death, and re-claim-on-restart.)
|
||||
- **v1 = the full compositor pipeline on the dumb framebuffer.** One `display` service
|
||||
owns the LFB + a cacheable back buffer + a layer stack; double-buffer + damage-driven
|
||||
present; clients draw via server-side commands. **No** runtime mode-setting, **no**
|
||||
shared-memory surfaces — both deferred (see display.md, "What v1 does not do").
|
||||
|
||||
## Conventions
|
||||
|
||||
Follow [coding-standards.md](../coding-standards.md): spell out non-acronym abbreviations in
|
||||
full, kebab-case file names, no `Co-Authored-By` trailers on commits. New user binaries
|
||||
go through `addUserBinary` in [build.zig](../../build.zig) and get packed into the
|
||||
initial-ramdisk; protocols are `b.addModule("…-protocol", …)` and imported into the
|
||||
`runtime` module.
|
||||
|
||||
## How to verify along the way
|
||||
|
||||
- `zig build test` — host unit tests (compositor math: layer clipping, damage merge,
|
||||
pitch/format blits are all host-testable with a fake framebuffer).
|
||||
- `python3 test/qemu_test.py <case>` — boots the real kernel in QEMU; assert on the
|
||||
serial log ([tests.zig](../../system/kernel/tests.zig) is the registry).
|
||||
- The `run-efi` target renders to QEMU's display (`-device VGA,edid=on,xres=1280,yres=720`)
|
||||
— a screenshot confirms pixels for the milestones whose gate is visual.
|
||||
|
||||
---
|
||||
|
||||
## D1 — The handoff primitive (kernel) ✅
|
||||
|
||||
Make the boot framebuffer reachable and mappable **write-combining** from user space.
|
||||
|
||||
- [x] [device-abi.zig](../../library/device/model/device-abi.zig): added `DeviceClass.display`; a
|
||||
`DisplayInfo{ width, height, pitch, format }` carried on the descriptor; a
|
||||
`flags` field on `ResourceDescriptor` + `resource_flag_write_combining`.
|
||||
- [x] [devices-broker.zig](../../system/kernel/devices-broker.zig): `seedDisplay(base, w, h,
|
||||
pitch, format)` publishes a root-level `display` node with one WC-flagged `memory`
|
||||
resource `[base, height*pitch]` + the `DisplayInfo`; `displayDevice()` /
|
||||
`displayClaimed()`. Seeded from `kmain` after `devices_broker.init`.
|
||||
- [x] [process.zig](../../system/kernel/process.zig) `systemMmioMap` + paging
|
||||
(`mapUserDeviceInto` gains a `write_combining` bool): a resource's WC flag maps it
|
||||
through the WC PAT slot (`setupPat`) instead of strong-uncacheable.
|
||||
- [x] [console.zig](../../system/kernel/console.zig): `setSuppressed` quiesces `write` while
|
||||
the display device is claimed (driven from `systemDeviceClaim` / release); the
|
||||
terminal panic + exception paths clear it first so a dying machine still draws.
|
||||
|
||||
**Gate (met, automated):** the `display` kernel test (`python3 test/qemu_test.py display`,
|
||||
`displayTest` in [tests.zig](../../system/kernel/tests.zig)) asserts the seeded node's shape
|
||||
and geometry, then walks the real claim + `mmio_map` path into a throwaway address space
|
||||
and verifies the leaf is **write-combining** (PAT entry 4: PAT bit set, PCD/PWT clear) —
|
||||
with an uncacheable-still-uncacheable regression guard. Chosen over the original
|
||||
screenshot-of-a-fill gate because it proves the *actual* WC property headlessly; the
|
||||
visible fill folds into D2's gate (the service clears the screen through the back buffer).
|
||||
Regression-checked: `discovery`, `ioport`, `claim-release`, `supervision`, `device-list`,
|
||||
`device-manager` all still pass with the +1 device in the table.
|
||||
|
||||
## D2 — Service skeleton, protocol, runtime module ✅
|
||||
|
||||
Stand up the named service and the double-buffer, no layers yet.
|
||||
|
||||
- [x] `library/protocol/display/display-protocol.zig`: `Operation{ info, create_layer,
|
||||
configure_layer, destroy_layer, fill_rect, blit_tile, damage, present }`; `extern`
|
||||
`Request`/`Reply`; size + `maximum_payload` consts. (Model: block/protocol.zig.)
|
||||
- [x] [abi.zig](../../system/abi.zig): `ServiceId.display = 9`.
|
||||
- [x] `system/services/display/display.zig`: `main` → enumerate + claim + WC-map the LFB
|
||||
(front) → `mmap` a cacheable back buffer of `height*pitch` → `runtime.service.run`.
|
||||
`info` and a whole-screen `present` (back → front) are live; layer ops fail-stub
|
||||
until D3. Init clears the back buffer and presents it — the double-buffer path.
|
||||
- [x] [library/runtime/display.zig](../library/runtime/runtime.zig) (+ barrel export of
|
||||
`display` and `display_protocol`): `info()` and `present()`, cached `.display`
|
||||
lookup with retry (model: block.zig).
|
||||
- [x] [init.zig](../../system/services/init/init.zig): `"display"` added to `boot_services`.
|
||||
- [x] [build.zig](../../build.zig): `display-protocol` module on the runtime; `display` exe
|
||||
via `addUserBinary`; packed into the initial-ramdisk; installed to
|
||||
`/system/services/display`.
|
||||
- [x] **Kernel fix the back buffer surfaced:** `mmap` was capped at 256 pages (1 MiB) by
|
||||
a fixed kernel-stack `frames` array. Rewrote `systemMmap` to map page-by-page with
|
||||
rollback (no scratch array) and raised the cap to 8192 pages (32 MiB) — enough for a
|
||||
4K back buffer. A real limitation met, exactly the kind this project chases.
|
||||
|
||||
**Gate (met, automated):** `python3 test/qemu_test.py display-service` spawns the
|
||||
compositor and matches its own serial heartbeats — `display: online {w}x{h} pitch …`
|
||||
followed by `display: presented frame 0` — which it prints only after the whole
|
||||
claim → WC-map → back-buffer → clear → present chain succeeds (matched on serial like the
|
||||
fault cases, since a lone blocking service can't reschedule the in-kernel test context to
|
||||
poll). Regression-checked: `usermem`, `heap` (the `mmap` rewrite), `init` (the boot-list
|
||||
addition), and D1's `display` all still pass.
|
||||
|
||||
## D3 — Layer stack + compositor + damage present ✅
|
||||
|
||||
The heart: composite an ordered layer stack, present only what changed.
|
||||
|
||||
- [x] A layer table (16 slots): each `Layer` = position, z, visible, a server-owned
|
||||
`mmap`'d surface (freed on `destroy_layer`). `damage` accumulates the dirty screen
|
||||
region since the last present.
|
||||
- [x] `create_layer` / `configure_layer` (damages old + new footprints) / `destroy_layer`,
|
||||
`fill_rect`, `blit_tile` (reads the inline tile from the IPC payload, unaligned-safe),
|
||||
`damage`, `present`.
|
||||
- [x] Pure, host-tested [compositor.zig](../../system/services/display/compositor.zig): `Rect`
|
||||
(intersect/unite), `Surface`, `fillRect`, `composite` (opaque, clipped to a damage
|
||||
rect), `blitTile`. `present` clears the damaged region to the wallpaper, paints the
|
||||
visible layers bottom-to-top (z-sorted), and flushes just that rect back → front (WC).
|
||||
Colour packing (rgbx/bgrx) is `protocol.pack`, also host-tested.
|
||||
- [x] Host tests (`zig build test`, green): rect intersect/unite, `fillRect` clipping +
|
||||
`stride > width` padding, `composite` overlap-shows-top + damage clipping, `blitTile`
|
||||
unaligned read + clipping, and `pack` for both pixel formats.
|
||||
|
||||
**Gate (met):** `zig build test` green for the compositor + pack unit tests, **and** the
|
||||
`display-service` case's startup self-check composites two overlapping layers on the real
|
||||
framebuffer and reads back the composited pixels — overlap = top layer, outside = bottom
|
||||
layer — logging `display: compositor self-check ok` (matched by the harness).
|
||||
|
||||
## D4 — Client API + the demo client ✅
|
||||
|
||||
Prove the pipeline end-to-end from a separate process.
|
||||
|
||||
- [x] Finished [runtime/display.zig](../library/runtime/runtime.zig): a `Layer` handle with
|
||||
`fill` / `blitTile` (inline tile) / `configure` (move/restack/show) / `damage` /
|
||||
`destroy`, `createLayer`, and a `color(r,g,b)` helper (caches the mode, packs via
|
||||
`protocol.pack`). Coordinates are signed over the wire (`@bitCast` both ways).
|
||||
- [x] `system/services/display-demo/`: a hardware-free client (the `input-source` analog)
|
||||
— a full-screen wallpaper layer, a rectangle that slides back and forth (moved by
|
||||
`configure` each frame, so the compositor repaints old + new), and a cursor layer;
|
||||
presents in a loop paced by `runtime.time`. Wired into build + initial-ramdisk.
|
||||
- [x] **Bug this surfaced:** `protocol.message_maximum` was 4096, but the kernel caps
|
||||
every IPC message at `MESSAGE_MAXIMUM` = 256 — so `replyWait` rejected the oversized
|
||||
receive buffer with `-E2BIG` and the serve loop had been *spinning* since D2 (unseen,
|
||||
as D2/D3 matched init-time heartbeats). Set it to 256; `blit_tile` is now explicitly
|
||||
a small-tile path (≤ 54 px inline), larger bitmaps being the deferred shared-memory surface.
|
||||
|
||||
**Gate (met):** `python3 test/qemu_test.py display-demo` spawns the service + `display-demo`;
|
||||
the demo drives a run of frames of motion through the layer client API and logs
|
||||
`display-demo: ok` (the visible motion is a screenshot via `zig build run-x86-64`).
|
||||
Regression-checked: `zig build test`, `display` (D1), and `display-service` (D2/D3) all
|
||||
still pass, and the default `zig build` is clean.
|
||||
|
||||
## D5 — Test cases + docs ✅
|
||||
|
||||
- [x] The three integration cases exist and pass: `display` (D1 handoff, kernel),
|
||||
`display-service` (D2/D3 compositor + self-check), and `display-demo` (D4 full
|
||||
pipeline: spawn `display` + `display-demo`, match `display-demo: ok`) —
|
||||
[tests.zig](../../system/kernel/tests.zig) + [qemu_test.py](../../test/qemu_test.py). Plus
|
||||
the pure host tests (`zig build test`).
|
||||
- [x] [display.md](display.md) updated to the built state (the "Verifying it" section names
|
||||
the real cases); [README index](../README.md) entry present (#19); the `display-track`
|
||||
memory marked DONE with the commits.
|
||||
|
||||
**Gate (met):** `python3 test/qemu_test.py display display-service display-demo` all pass,
|
||||
`zig build test` is green, and the default `zig build` is clean.
|
||||
|
||||
---
|
||||
|
||||
## v1 status: complete
|
||||
|
||||
D1–D5 done. The display service is a working framebuffer compositor: it owns the
|
||||
framebuffer (write-combining), composites a z-ordered layer stack into a cacheable back
|
||||
buffer, presents only the damaged region, and is driven over IPC by the `runtime.display`
|
||||
client — proven end-to-end by a separate demo process. Two limitations are deliberate and
|
||||
documented (docs/display.md): no runtime mode-setting (native backend) and no true vsync
|
||||
(no vblank on a dumb framebuffer). Next steps are the Deferred items below.
|
||||
|
||||
---
|
||||
|
||||
## Deferred (explicitly not in this plan)
|
||||
|
||||
- **Shared-memory surfaces** — generalize M13 capability passing to memory objects
|
||||
(`shared_memory_create`/`shared_memory_map`), so bitmap clients hand the compositor a rendered surface
|
||||
instead of drawing commands. The compositor's layer model already anticipates it.
|
||||
- **Native backend (Bochs DISPI, then virtio-gpu)** — behind the same internal backend
|
||||
interface as the dumb framebuffer: EDID mode list + runtime resolution/bpp change +
|
||||
(eventually) a vblank/flip path for true vsync.
|
||||
- **Driver/compositor process split** — only when a second backend or a second head makes
|
||||
the abstraction pay for itself.
|
||||
@@ -0,0 +1,175 @@
|
||||
# Display v2 — build plan (pluggable scanout: GOP floor + virtio-gpu native)
|
||||
|
||||
The ordered, checkpointable build-out for [display-v2.md](display-v2.md). Each milestone
|
||||
lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run, like
|
||||
[display-plan.md](display-plan.md). Read display-v2.md first for the *why*.
|
||||
|
||||
## Locked decisions (do not relitigate)
|
||||
|
||||
- **First native backend = virtio-gpu** (VM standard: mode-set + fenced present/flush).
|
||||
- **Dynamic hot-attach**: boot on GOP, upgrade to native when the driver **announces**
|
||||
(push, not polling); re-attach across driver restarts; GOP is the floor for "no driver
|
||||
ever," not a live fall-back after a reprogram.
|
||||
- **v2 builds the shared-memory capability** (endpoints → memory objects), shared with the future
|
||||
client-surface path.
|
||||
- The compositor's layers/back-buffer/damage are **unchanged**; only scanout is pluggable.
|
||||
|
||||
## Conventions
|
||||
|
||||
Follow [coding-standards.md](../coding-standards.md): spell out non-acronym abbreviations,
|
||||
kebab-case file names, no `Co-Authored-By` trailers. New user binaries go through
|
||||
`addUserBinary` and get packed into the initial-ramdisk; protocols are
|
||||
`b.addModule("…-protocol", …)` imported into `runtime`; new syscalls extend
|
||||
[abi.zig](../../system/abi.zig) `SystemCall` + a `library/runtime` wrapper.
|
||||
|
||||
## How to verify along the way
|
||||
|
||||
**Every gate is serial-checkable — no screenshots** (this plan is built to run unattended).
|
||||
Where "does it actually display" would otherwise need a human eyeball, the code **reads its
|
||||
own pixels back**: the scanout resource is CPU-visible RAM (shared-memory-backed) and the back buffer
|
||||
is cacheable, so a driver/compositor can write a known value, read it back, and log a
|
||||
pass/fail — and a virtio `resource_flush` is confirmed by the device **acking it on the
|
||||
used ring**. Those two together (pixel-readback + flush-ack) are the automated stand-in for
|
||||
"it's on screen."
|
||||
|
||||
- `zig build test` — host unit tests (backend selection, virtio struct sizes/encodings,
|
||||
pixel-check helpers).
|
||||
- `python3 test/qemu_test.py <case>` — boots the kernel in QEMU; asserts on serial markers.
|
||||
The virtio cases boot with `-device virtio-gpu` (a per-case `qemu_extra`).
|
||||
- `run-x86-64` renders to a window — for the human's own satisfaction, **not** a gate.
|
||||
|
||||
---
|
||||
|
||||
## V1 — The scanout backend seam (refactor, no behaviour change) ✅
|
||||
|
||||
Extract scanout from the compositor so today's path becomes one backend among future ones.
|
||||
|
||||
- [x] `system/services/display/backend.zig`: a `Backend` tagged union with `info()`,
|
||||
`surface()` (the cacheable compose target), `present(damage)`, and capability flags
|
||||
(`canModeSet`/`hasFencedPresent`, both false for GOP).
|
||||
- [x] The v1 GOP path is now `backend.Gop` (claims the `display` node, WC-maps the LFB,
|
||||
keeps the cacheable back buffer, `present` = the damage-rect WC copy). display.zig
|
||||
composes into `backend.surface()` and calls `backend.present(damage)` — no LFB or
|
||||
framebuffer geometry left in the compositor core.
|
||||
- [x] The selection decision is the pure `chooseKind(native_available)` (gop unless a
|
||||
native driver announced), split from the syscall-bound `select()`/`Gop.init()`.
|
||||
|
||||
**Gate (met):** `display-service` + `display-demo` pass **unchanged** (pure refactor; GOP
|
||||
is the only backend), and `zig build test` stays green.
|
||||
|
||||
## V2 — The shared-memory cross-process capability (kernel) ✅
|
||||
|
||||
- [x] [abi.zig](../../system/abi.zig): `shared_memory_create` (34) / `shared_memory_map` (35) syscalls + a
|
||||
`shared_memory_test` service id. Handlers in process.zig: `shared_memory_create(len)` allocates contiguous,
|
||||
zeroed, **cacheable** frames, wraps them in a refcounted object, installs a capability
|
||||
handle, maps them into the caller's shared-memory arena → returns virtual_address + handle; `shared_memory_map(cap)`
|
||||
maps the same physical pages into the receiver. Reclaimed on death (see below).
|
||||
- [x] The capability core (ipc-synchronous.zig) is now **kind-tagged**: `scheduler.Task`'s
|
||||
handle table holds `HandleObject{kind, ptr}`; `closeHandles` and `shareCapability`
|
||||
dispatch by kind, so a `SharedMemoryObject` rides an `ipc_call` `send_cap` exactly like an
|
||||
endpoint and frees only when its last capability drops. `mapUserSharedInto` (paging)
|
||||
maps WB-cacheable + `device_grant`, so a sharer's teardown never frees the shared
|
||||
frames — the object owns them.
|
||||
- [x] `library/runtime/shared-memory.zig` (+ barrel export): `create(len) -> Region{ptr, handle, len}`,
|
||||
`map(handle) -> ptr`.
|
||||
|
||||
**Gate (met):** `python3 test/qemu_test.py shared-memory` — `shared-memory-client` creates a region, writes a
|
||||
pattern, and passes its capability to `shared-memory-server` as an `ipc_call` send_cap; the server
|
||||
`shared_memory_map`s it and reads the **same bytes** back → `shared-memory: shared 4096 bytes ok`. Guardrail:
|
||||
`ipc`/`ipc-call`/`ipc-cap`, `supervision`, `dma`, `usermem`, `display-service`, and host
|
||||
tests all still pass — the handle-table change broke no existing IPC.
|
||||
|
||||
## V3 — The virtio-gpu driver: bring-up + a frame on screen ✅
|
||||
|
||||
- [x] `system/drivers/virtio-gpu/`: claim the virtio-gpu PCI function (device-manager
|
||||
match on the display/other class triple, driver self-confirms vendor 0x1AF4/device
|
||||
0x1050 from config space), enable memory-space + bus-master, walk the vendor
|
||||
capabilities in config space to find common-config + notify, map the BAR, negotiate
|
||||
VERSION_1, and stand up the control virtqueue in coherent DMA. `virtio-gpu-protocol.zig`
|
||||
+ `virtio-pci.zig` for the control/transport structs (host-tested sizes).
|
||||
- [x] Create a 2D scanout resource backed by a coherent DMA region (V4 swaps this for the
|
||||
shared-memory surface), `attach_backing`, `set_scanout` to scanout 0, `transfer_to_host_2d`
|
||||
+ `resource_flush` of a test pattern, and wait on the used ring.
|
||||
- [x] Register a `scanout` service (`ServiceId.scanout` = 11).
|
||||
|
||||
**Gate (met):** the `virtio-gpu` case (QEMU `-device virtio-gpu-pci`) boots the
|
||||
device-manager stack, which discovers the function and spawns the driver; the driver writes
|
||||
a known test pattern into the scanout backing, `transfer_to_host_2d` + `resource_flush`es
|
||||
it, and **waits for the device's used-ring ack**, then reads the backing back and checks the
|
||||
pattern — logging `virtio-gpu: scanout 640x480 online` and `virtio-gpu: flush acked, pixel
|
||||
check ok`. That proves virtqueue + resource + attach + set_scanout + transfer + flush end to
|
||||
end without a screenshot (the used-ring ack is the device confirming it consumed the frame).
|
||||
|
||||
## V4 — The native backend + hot-attach ✅
|
||||
|
||||
- [x] `backend.VirtioGpu` in the compositor: `surface()` = the shared-memory scanout surface
|
||||
(the compositor composes straight into the device's resource backing; x86 DMA is
|
||||
coherent, so the cacheable shared pages need no flush), `present(damage)` = a `present`
|
||||
request over the driver's `.scanout` endpoint (→ transfer-to-host + resource flush).
|
||||
- [x] The driver **announces** to `.display` after bring-up (looks it up with a bounded retry,
|
||||
sends `attach_scanout` with the geometry + the shared surface as an `ipc_call` send_cap).
|
||||
The compositor maps it, looks up `.scanout` itself (no need to pass the endpoint — the
|
||||
driver registered it), switches backend, and re-composites the current frame full-screen.
|
||||
The present is deferred to a one-shot timer so it runs *after* the reply unblocks the
|
||||
driver and it serves `.scanout` — presenting inline would deadlock.
|
||||
- [x] Boot still starts on `backend.Gop`; the upgrade happens on announce. `shared_memory_physical` (a
|
||||
new syscall) gives the driver the guest-physical of the shared surface for `attach_backing`.
|
||||
|
||||
**Gate (met):** the `display-native` case (QEMU `-device virtio-gpu-pci`, `mem` bumped since it
|
||||
boots the whole system) starts the compositor + `display-demo` + device-manager; the driver
|
||||
announces, the compositor logs `display: scanout upgraded to virtio-gpu`, drives frames through
|
||||
the native backend, and **reads a pixel back** from the shared surface after a present to
|
||||
confirm the composited frame landed (`display: native present verified`), while `display-demo:
|
||||
ok` still fires — checked order-independently. Without `-device virtio-gpu-pci` nothing is
|
||||
announced and it stays on GOP: the v1 `display-service`/`display-demo` gates pass unchanged.
|
||||
|
||||
## V5 — Mode-setting, EDID, and fenced presents ✅
|
||||
|
||||
- [x] The driver negotiates `VIRTIO_GPU_F_EDID` (when offered) and reads the monitor's EDID,
|
||||
logging its preferred mode; it offers a small mode list over `.scanout` `get_modes`. The
|
||||
resource + shared surface are sized to the largest mode, so `set_mode` just re-points the
|
||||
scanout rectangle (no resource/surface churn) — a runtime resolution change. `runtime.display`
|
||||
gains `modes()` / `setMode()` (display-protocol `get_modes`/`set_mode`, forwarded to the backend).
|
||||
- [x] Every `resource_flush` is issued fenced (`VIRTIO_GPU_FLAG_FENCE`); the device signals the
|
||||
fence when it has consumed the frame, which the used-ring ack the synchronous present waits
|
||||
on already gates — a tear-free present. (Completion feedback, **not vblank**: base
|
||||
virtio-gpu 2D has no display-refresh event, so nothing paces presents to the monitor —
|
||||
see the "Fenced is not vsync" note in [display-v2.md](display-v2.md).)
|
||||
- [x] `backend.VirtioGpu` reports `canModeSet` / `hasFencedPresent` = true.
|
||||
|
||||
**Gate (met):** the `display-modeset` case (reusing the display-native boot) upgrades to
|
||||
virtio-gpu, queries the driver's modes, `setMode`s to a different resolution, and confirms the
|
||||
change by reading the backend's geometry back (`display: mode set to {w}x{h}, verified`); the
|
||||
fenced present path is exercised and confirmed (`display: fenced present ok`) — both from serial,
|
||||
passing 3/3. The driver also logs the EDID preferred mode (`virtio-gpu: EDID preferred mode …`).
|
||||
|
||||
## V6 — Resilience (restart + re-attach) + tests + docs ✅
|
||||
|
||||
- [x] The virtio-gpu driver now **hellos** the device manager (role: bus) so it is properly
|
||||
supervised — no longer stopped at the hello deadline — and is restarted on death. On
|
||||
driver loss the compositor keeps the last frame (its `.scanout` calls now return
|
||||
`-EPEER` instead of hanging — a kernel fix: an endpoint is marked dead when its owner
|
||||
dies) and **re-attaches** when the restarted driver re-announces. A permanent give-up
|
||||
(crash-loop cap) leaves the frozen frame; GOP is not re-taken.
|
||||
- [x] `test/qemu_test.py`: the `virtio-gpu`, `display-native` (hot-attach), `display-modeset`,
|
||||
and `display-reattach` (driver-kill/re-attach) cases. display-v2.md status updated.
|
||||
|
||||
**Gate (met):** the `display-reattach` case — device-manager (in `test-scanout-restart` mode)
|
||||
kills the virtio-gpu driver once after it hellos; the restart policy respawns it, it
|
||||
re-announces, and the compositor logs `display: scanout re-attached` after the initial
|
||||
`display: scanout upgraded to virtio-gpu`, with no CPU exception / panic (the compositor
|
||||
survives) — passing 3/3. All v1 + v2 cases (host tests, `ipc`/`ipc-call`/`ipc-cap`,
|
||||
`supervision`, `shared-memory`, `display-service`, `display-demo`, `virtio-gpu`, `display-native`,
|
||||
`display-modeset`) pass; default `zig build` is clean.
|
||||
|
||||
---
|
||||
|
||||
## Deferred (explicitly not in this plan)
|
||||
|
||||
- **Client-rendered surfaces** — now unblocked by the shared-memory capability (V2): an app renders
|
||||
its own bitmap and hands the compositor a reference. A natural follow-on.
|
||||
- **Bochs DISPI backend** — a simpler second native backend (mode-set only, dumb scanout);
|
||||
slots behind the same interface if wanted.
|
||||
- **Real-GPU (NVIDIA/AMD/Intel) drivers** — out of scope; those devices stay on the GOP
|
||||
floor by design.
|
||||
- **Hardware-accelerated compositing / multiple heads** — future.
|
||||
@@ -0,0 +1,147 @@
|
||||
# The display service v2: a pluggable scanout backend
|
||||
|
||||
**Status: complete (V1–V6).** The compositor boots on the GOP framebuffer and, when a
|
||||
virtio-gpu driver announces itself, hot-attaches a native backend over the shared-memory
|
||||
scanout surface — with runtime mode-setting, EDID, and fenced presents, and it
|
||||
re-attaches across driver restarts. All serial-gated (see [display-v2-plan.md](display-v2-plan.md)).
|
||||
|
||||
v1 ([display.md](display.md)) is a compositor that owns the **GOP framebuffer** — it
|
||||
composites a layer stack into a cacheable back buffer and streams damage to the linear
|
||||
framebuffer the firmware handed over. That path is portable and good: it drives any GPU,
|
||||
including a real NVIDIA card at an ultrawide's native resolution, with zero GPU-specific
|
||||
code. v2 keeps it as the **floor** and makes *scanout* — how a finished frame reaches the
|
||||
panel — a **pluggable backend**, so the compositor can **upgrade to a real GPU driver when
|
||||
one is present** and fall back to the framebuffer when it isn't.
|
||||
|
||||
The compositor itself (layers, back buffer, damage) does not change. Only the last step —
|
||||
"put this frame on screen" — becomes swappable.
|
||||
|
||||
## The shape
|
||||
|
||||
```
|
||||
compositor (display service) ── layer stack + back buffer + damage (unchanged)
|
||||
│ composites a frame, then: backend.present(damage)
|
||||
▼
|
||||
scanout backend (selected at runtime — GOP by default, native when it appears)
|
||||
│
|
||||
├─ GopBackend the v1 path: WC copy back→front to the firmware LFB.
|
||||
│ Always available. No mode-set, no present fence. THE FLOOR.
|
||||
│
|
||||
└─ VirtioGpuBackend talks to a virtio-gpu driver process over a `scanout`
|
||||
service: present via a shared resource + fenced flush,
|
||||
a fixed mode list, runtime mode-set, EDID refresh rate.
|
||||
```
|
||||
|
||||
A **backend** is a small interface the compositor calls:
|
||||
|
||||
- `surface()` → the pixels to compose into and their geometry `{ptr, pitch, format, w, h}`
|
||||
(the LFB for GOP; a shared scanout resource for virtio-gpu),
|
||||
- `present(damage: Rect)` → make the damaged region visible (a no-op-ish WC copy for GOP;
|
||||
a fenced virtio flush for the native path),
|
||||
- capability queries — `canModeSet`, `hasFencedPresent` — and, when supported, `modes()` /
|
||||
`setMode(m)`.
|
||||
|
||||
The compositor composes into `surface()` and calls `present(damage)` exactly as it does
|
||||
today; everything device-specific lives behind the interface.
|
||||
|
||||
## Selection and hot-attach
|
||||
|
||||
The choice is **dynamic**, because a GPU driver is spawned asynchronously (the device
|
||||
manager brings it up after boot), and because danos is meant to be resilient:
|
||||
|
||||
1. **Boot on GOP.** The compositor starts on `GopBackend` immediately, so there is never a
|
||||
blank screen while drivers load — the exact v1 behaviour.
|
||||
2. **Upgrade on announce.** When the virtio-gpu driver has claimed its device and set up a
|
||||
scanout, it **announces itself to the display service** (a `push`: the driver looks up
|
||||
`.display` and sends an *attach-scanout* message carrying the shared scanout **surface**
|
||||
as a capability; the compositor maps it and reaches the driver's present/mode channel by
|
||||
looking up the registered `scanout` service). The compositor switches to
|
||||
`VirtioGpuBackend` and re-presents the current frame full-screen. Push beats polling —
|
||||
the compositor doesn't know a priori which driver, if any, exists, and danos has no
|
||||
service-registration pub/sub.
|
||||
3. **Native is restartable, not fallback-on-crash.** Once a native driver has reprogrammed
|
||||
the device, the firmware's GOP framebuffer is **stale** — "native → GOP" is not a clean
|
||||
fall-back. So a native driver that **crashes** is *restarted* by its supervisor (the
|
||||
resilience work already merged), re-announces, and the compositor **re-attaches**
|
||||
(native → native). The screen freezes on the last frame during the gap — acceptable.
|
||||
4. **GOP is the floor for "no driver was ever there."** On a real GPU (NVIDIA/AMD/Intel)
|
||||
the class-0x03 device matches nothing in the driver table, no `scanout` is ever
|
||||
announced, and the compositor stays on GOP forever — no special-casing. If a native
|
||||
driver *permanently* gives up (the device manager's crash-loop cap), nothing tells the
|
||||
compositor and there is no path back to GOP — the screen stays frozen on the last
|
||||
frame. A GOP revert (sensible only while the LFB is still mappable) is not built.
|
||||
|
||||
## The shared-memory primitive this needs
|
||||
|
||||
virtio-gpu's scanout resource is **guest RAM** — the driver allocates it and attaches it
|
||||
to a virtio resource, and the compositor composes into it. That means the compositor
|
||||
writing into the driver's buffer is **cross-process memory sharing**, the primitive v1
|
||||
deferred (docs/display.md, "What v1 does not do"). v2 builds it: the natural generalization
|
||||
of M13 capability-passing from *endpoints* to *memory objects* —
|
||||
|
||||
```
|
||||
shared_memory_create(len) -> {handle, virtual_address} // a shareable, page-aligned RAM region
|
||||
… pass `handle` as the send_cap on an ipc_call …
|
||||
shared_memory_map(cap) -> virtual_address // the receiver maps the same physical pages
|
||||
```
|
||||
|
||||
The payoff is leverage: the **same** primitive unlocks **both** native GPU drivers *and*
|
||||
client-rendered surfaces (an app composing its own bitmap and handing the compositor a
|
||||
reference instead of drawing by command). One piece of kernel work, two features.
|
||||
|
||||
## The virtio-gpu driver
|
||||
|
||||
A new ring-3 driver process (the topology v1 anticipated — "split the driver from the
|
||||
compositor when a second backend arrives"). It claims the virtio-gpu PCI function, and:
|
||||
|
||||
- sets up the **control virtqueue** (queue 0 — the only queue it uses; no cursor queue)
|
||||
and the device's config space,
|
||||
- creates a **2D scanout resource** backed by a shared-memory region, `attach_backing`s it,
|
||||
`set_scanout`s it to a CRTC, and on each present transfers and `resource_flush`es the
|
||||
full current-mode rectangle (damage-narrowed flushes are a later refinement),
|
||||
- reads **EDID** (the `GET_EDID` control command) to log the monitor's preferred timing
|
||||
and derive the refresh rate it announces (the compositor's frame-clock seed); the mode
|
||||
list it offers is a fixed pair — 640×480 and 800×600 — and `set_scanout` at a chosen
|
||||
mode gives **runtime mode-setting**,
|
||||
- registers a `scanout` service and announces to the display service.
|
||||
|
||||
Its `resource_flush` is the real **present** — and gives a **fenced, tear-free** path a
|
||||
dumb GOP framebuffer can't.
|
||||
|
||||
**Fenced is not vsync.** The fence completes when the device has *consumed* the frame:
|
||||
real completion feedback, and tear-freedom by snapshot semantics (the host displays
|
||||
discrete transferred frames, never a half-written surface). It is **not** a vblank —
|
||||
base virtio-gpu 2D has no display-refresh event at all (Linux's driver for this device
|
||||
fakes one with a software timer), so nothing paces presents to the monitor's refresh.
|
||||
Refresh-paced presents need either a native driver's vblank interrupt (delivered over
|
||||
the existing IRQ-as-IPC path) or the compositor's own frame clock.
|
||||
|
||||
## What v2 unlocks — and its honest scope
|
||||
|
||||
Behind the abstraction, a native backend gives runtime **mode-setting** (resolution only —
|
||||
the scanout protocol carries neither refresh nor bpp), an **EDID**-derived refresh rate, and
|
||||
**fenced presents**. But only on devices we have a driver
|
||||
for — realistically **VMs** (virtio-gpu, and later maybe Bochs DISPI). Real discrete GPUs
|
||||
need per-vendor KMS-class drivers that aren't getting written, so they **stay on GOP** —
|
||||
which is genuinely fine (v1 on the NVIDIA box is smooth). So v2's real value is twofold:
|
||||
the **pluggable architecture** (a driver slots in when one exists) and a **rich, fenced
|
||||
path in VMs**, where danos development happens. The framebuffer floor never goes away.
|
||||
|
||||
## Locked decisions
|
||||
|
||||
- **First native backend: virtio-gpu** — the VM standard; gives mode-set + a real
|
||||
present/flush (fenced), and exercises the whole pluggable design. Tested with QEMU
|
||||
`-device virtio-gpu`.
|
||||
- **Dynamic hot-attach** — boot on GOP, upgrade to native on the driver's announce,
|
||||
re-attach across driver restarts; GOP is the floor for "no driver ever," not a live
|
||||
fall-back after a reprogram.
|
||||
- **Detection = push** (the driver announces to `.display`), not compositor polling.
|
||||
- **v2 builds the shared-memory capability** (endpoints → memory objects), shared with the future
|
||||
client-surface path.
|
||||
|
||||
## See also
|
||||
|
||||
- [display.md](display.md) — v1: the compositor, the GOP-vs-device split, the WC discipline.
|
||||
- [display-v2-plan.md](display-v2-plan.md) — the ordered build-out.
|
||||
- [driver-model.md](driver-model.md) — claim / `mmio_map` / MSI / capability passing (M13).
|
||||
- [resilience.md](../os-development/resilience.md) — the restart machinery the hot-attach leans on.
|
||||
@@ -0,0 +1,325 @@
|
||||
# The display service: a framebuffer compositor
|
||||
|
||||
The [framebuffer](../os-development/framebuffer.md) the loader hands over is a flat block of pixel
|
||||
memory, and the kernel's [bootstrap console](../../system/kernel/console.zig) draws text
|
||||
into it directly. That console is a stop-gap. The **display service**
|
||||
(`system/services/display/`) is the real thing: an ordinary ring-3 process that *owns*
|
||||
the framebuffer, composes a stack of **layers** into an off-screen back buffer, and
|
||||
**presents** finished frames to the screen — the display half of the GUI track
|
||||
([vision.md](../vision.md)), the sibling of the [input service](input.md).
|
||||
|
||||
This note is the architecture and the reasoning behind it. The concrete build order
|
||||
lives in [display-plan.md](display-plan.md).
|
||||
|
||||
## First, a distinction that shapes everything: GOP vs. the PCI device
|
||||
|
||||
It is tempting to think "the GOP framebuffer" and "the VGA-compatible display
|
||||
controller in the PCIe tree" are two different things. They are not — they are **two
|
||||
interfaces to the same silicon, at different times and different levels**, and knowing
|
||||
which one you're holding decides what you can do.
|
||||
|
||||
- **GOP is firmware's *temporary* driver** for the display controller. It gives you a
|
||||
linear framebuffer pointer and can set video modes — but only until
|
||||
`ExitBootServices`. The loader already leans on this: [`queryFramebuffer`](../../boot/efi.zig)
|
||||
reads the monitor's EDID, picks the native mode, and calls `set_mode` **before**
|
||||
exiting ([gop.md](../os-development/gop.md)). Once the kernel runs, GOP is **gone** — no `set_mode`, no
|
||||
mode list, no EDID. What survives is the frozen snapshot in
|
||||
[`BootInformation.framebuffer`](../../system/boot-handoff.zig): `{base, width, height,
|
||||
pitch, format, refresh_hz}`, and nothing more.
|
||||
|
||||
- **The PCI class-0x03 device is the raw controller** — BARs, config space, registers,
|
||||
IO ports. It is what you actually *own* after boot. On QEMU's emulated adapter
|
||||
([`-device VGA,edid=on`](../../build.zig), the Bochs VBE/DISPI model) the `base` GOP handed
|
||||
you *is* that device's linear-framebuffer BAR — the same physical memory, seen through
|
||||
a different door. On a real discrete GPU, GOP's `base` is an aperture inside the GPU's
|
||||
VRAM BAR. danos already decodes this device
|
||||
([pci-class.zig](../../library/device/pci/pci-class.zig) has the full `display` namespace, and
|
||||
`pci-bus` already reports it to the [device manager](device-manager.md) with its class
|
||||
triple) — but nothing binds it yet.
|
||||
|
||||
What that difference costs you, concretely:
|
||||
|
||||
| You want to… | Dumb GOP framebuffer (boot handoff) | Native device driver (PCI 0x03) |
|
||||
|-------------------------------------------|-------------------------------------|------------------------------------------|
|
||||
| **Report** the current mode | ✅ from the handoff | ✅ |
|
||||
| **Change resolution / bpp at runtime** | ❌ GOP is gone | ✅ program DISPI regs / virtio-gpu queue |
|
||||
| **Re-read EDID, enumerate monitor modes** | ❌ | ✅ the device exposes an EDID block |
|
||||
| **Refresh rate** | ❌ (virtual anyway) | only a real KMS driver — far future |
|
||||
| **vblank / tear-free present** | ❌ no vblank signal | ✅ vblank IRQ + page-flip (real GPUs) |
|
||||
| **Works on the Pi (no PCI VGA)** | ✅ VideoCore hands a simple FB | ✗ per-device |
|
||||
|
||||
The lesson: the **portable base for the whole GUI stack is the GOP / boot-handoff linear
|
||||
framebuffer**. Runtime mode-setting is a *per-device upgrade* layered on top — and on
|
||||
the Raspberry Pis there is no PCI VGA at all, so the neutral framebuffer is the only
|
||||
thing all three target machines share. That is why the display service is built on the
|
||||
dumb framebuffer first, with the native backend as an optional module behind the same
|
||||
interface.
|
||||
|
||||
## Two constraints this service exists to meet
|
||||
|
||||
Like the input service — which existed partly to motivate the asynchronous
|
||||
[`ipc_send`](ipc.md) primitive — the display service runs straight into two limits the
|
||||
rest of the system hasn't had to face:
|
||||
|
||||
1. **The framebuffer is kernel-only today.** It arrives through the boot handoff, is
|
||||
mapped into the kernel's physmap, and is touched only by
|
||||
[`console.zig`](../../system/kernel/console.zig). It is *not* a
|
||||
[devices-broker](../../system/kernel/devices-broker.zig) node, so
|
||||
`device.claim`/`mmio_map` cannot reach it, and there is no framebuffer
|
||||
[syscall](../os-development/syscall.md). A user-space display service needs a **new mechanism just to
|
||||
touch the pixels**. (See "The handoff" below — this is built.)
|
||||
|
||||
2. **danos had no cross-process shared memory.** At v1 the memory syscalls were `mmap`
|
||||
(private, zeroed), `mmio_map` (a *claimed device's* MMIO), and `dma_alloc` (new
|
||||
pinned physical). The block driver's "pass a buffer by physical address" trick
|
||||
([block/protocol.zig](../../library/protocol/block/block-protocol.zig)) works *only because its
|
||||
consumer is DMA hardware*. A compositor that CPU-reads and blends client layers can't
|
||||
use it — it would have to *map* another process's memory, which nothing allowed. v1
|
||||
sidesteps it entirely (see "What v1 does not do"); v2 has since built the primitive
|
||||
(`shared_memory_create` / `shared_memory_map` / `shared_memory_physical` —
|
||||
[display-v2.md](display-v2.md)).
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
kernel ── owns the boot framebuffer; bootstrap console only
|
||||
│ seeds a display-class device node from BootInformation.framebuffer
|
||||
│ (ResourceKind.memory = [base, height*pitch], write-combining hint,
|
||||
│ plus DisplayInfo{width, height, pitch, format, refresh_hz})
|
||||
▼
|
||||
display service (system/services/display/, ServiceId.display) ← the compositor
|
||||
│ device.claim(display node) → mmio_map(WRITE-COMBINING) = FRONT buffer (the LFB)
|
||||
│ mmap(cacheable) a BACK buffer of the same geometry
|
||||
│ owns: an ordered LAYER STACK + a per-frame DAMAGE tracker (rect list or tile grid)
|
||||
│ loop: composite dirty layers → back buffer → present dirty rects → front
|
||||
│ backend is an INTERNAL interface: {gop-fb} at boot; {virtio-gpu} on hot-attach (v2)
|
||||
▼ reached by name (ipc_lookup); clients drive it over the display protocol
|
||||
┌────────────────────────────────────┬──────────────────────────────────────┐
|
||||
drawing clients (v1) surface clients (deferred)
|
||||
display commands: display surfaces:
|
||||
create_layer / configure_layer shared_memory_create → pass as a capability →
|
||||
fill_rect / blit_tile / damage the compositor maps & composites the
|
||||
present client-rendered bitmap directly
|
||||
```
|
||||
|
||||
The bring-up sequence mirrors a hardware driver's — it is the
|
||||
[`usb-xhci-bus` `initialise`](../../system/drivers/usb-xhci-bus/usb-xhci-bus.zig) shape
|
||||
(claim → `mmio_map` → run loop) — and the request/reply service shell is the
|
||||
[FAT](../../system/services/fat/fat.zig) / [input](../../system/services/input/input.zig) shape
|
||||
([`service.run`](../../library/kernel/service.zig) with a `protocol.zig` of
|
||||
`extern struct` messages and an `Operation` tag).
|
||||
|
||||
**One process, for now.** v1 is a *single* service that both owns the framebuffer and
|
||||
composites — it does not split a "framebuffer driver" from a "compositor" the way input
|
||||
splits `ps2-bus` from the input service. The backend (dumb FB vs. a native GPU) is an
|
||||
*internal* interface, not a process boundary. That boundary earns its keep only when a
|
||||
second backend or a second monitor appears; until then it is complexity with no payoff.
|
||||
|
||||
## The handoff: a device node + a write-combining map
|
||||
|
||||
The framebuffer crosses into user space through the machinery that already exists for
|
||||
every other device, rather than a bespoke syscall — so it inherits ownership,
|
||||
release-on-death, and re-claim-on-restart for free (the [resilience](../os-development/resilience.md)
|
||||
story: a crashed display service returns the LFB to the kernel, and its restart
|
||||
re-claims it).
|
||||
|
||||
- The kernel seeds a synthetic **display-class** node into the
|
||||
[devices-broker](../../system/kernel/devices-broker.zig) at init (`seedDisplay`), from
|
||||
`BootInformation.framebuffer`: one `ResourceKind.memory` resource spanning
|
||||
`[base, height*pitch]`, tagged **write-combining**, plus a small
|
||||
`DisplayInfo{width, height, pitch, format, refresh_hz}` (the memory resource says *where*
|
||||
and *how big*; `DisplayInfo` says how to *interpret* the bytes — and `refresh_hz`, the
|
||||
panel refresh the loader computed from EDID before `ExitBootServices`, seeds the
|
||||
compositor's frame clock). The node carries no name or index; it is identified purely by
|
||||
its `display` device class.
|
||||
- The service `device.claim`s it and `mmio_map`s the resource. The map is
|
||||
**write-combining**, not the strong-uncacheable that `mmio_map` uses for register
|
||||
MMIO. The kernel already programs a WC PAT slot for its own console
|
||||
([`setupPat`](../../system/kernel/architecture/x86_64/paging.zig)); this reaches it from
|
||||
the user mapping path. **This matters:** an uncacheable framebuffer makes the
|
||||
back→front blit unusably slow.
|
||||
- On `claim`, the kernel's bootstrap console goes quiet, so the two never fight over the
|
||||
LFB. A panic is the one exception — by then the service is likely dead anyway, and a
|
||||
panic on screen wins.
|
||||
|
||||
The display service is a **named boot service**: `init` spawns it by name alongside
|
||||
`input`/`device-manager`/`fat` ([init.zig](../../system/services/init/init.zig)), and it
|
||||
self-discovers the display node with `device.enumerate` (matching on `DeviceClass.display`). The [device manager](device-manager.md)
|
||||
matching path (PCI class 0x03 → a driver) is reserved for the future *native* backend, not
|
||||
this singleton synthetic node.
|
||||
|
||||
## Double buffering and the write-combining discipline
|
||||
|
||||
Two buffers, with deliberately different memory types:
|
||||
|
||||
- The **front buffer** is the LFB — **write-combining**: fast to *write*, slow to
|
||||
*read*. The rule is therefore **never read the front buffer**. Only ever stream into
|
||||
it, sequentially.
|
||||
- The **back buffer** is ordinary **cacheable** RAM (`mmap`), the same geometry. All
|
||||
compositing happens here, where reads and read-modify-write blends are cheap.
|
||||
|
||||
So a frame is: compose every dirty layer into the cacheable back buffer, then **present**
|
||||
— copy the changed regions back→front in sequential, WC-friendly writes. Two details the
|
||||
[framebuffer](../os-development/framebuffer.md) note already establishes carry over: step rows by `pitch`,
|
||||
not `width*4`; and handle both `rgbx` and `bgrx` [pixel formats](../os-development/gop.md).
|
||||
|
||||
## Flicker vs. tearing — what double buffering does and doesn't buy
|
||||
|
||||
These are two different artifacts, and the dumb framebuffer fixes exactly one of them:
|
||||
|
||||
- **Flicker** is the user seeing intermediate, half-drawn states (a clear-then-redraw
|
||||
flash). Double buffering **eliminates it completely** — the screen only ever receives
|
||||
whole, finished frames.
|
||||
- **Tearing** is a present landing while the display's scanout beam is mid-frame, so the
|
||||
top of the screen shows the new frame and the bottom the old. Avoiding it requires
|
||||
presenting during the vertical blank (**vsync**) — which needs a vblank signal. **A
|
||||
dumb GOP framebuffer has no vblank.**
|
||||
|
||||
So v1 is **flicker-free**, and it *minimizes* the tear window by presenting only damaged
|
||||
rectangles (less to copy → a smaller window in which the beam can catch a half-updated
|
||||
frame), but it is **not tear-free**. Genuine vsync waits for a backend with a vblank IRQ
|
||||
or a flush/flip path — a native-device capability, not something the firmware
|
||||
framebuffer can offer. Stated plainly here so the limitation is understood, not
|
||||
discovered.
|
||||
|
||||
## Layers and the client protocol
|
||||
|
||||
The compositor holds an **ordered stack of layers**. Each layer has a rectangle, a
|
||||
z-order, a visibility flag, and a surface. Presenting walks the stack bottom-to-top,
|
||||
painting each dirty layer into the back buffer, then flushes the damage to the front.
|
||||
Damage is tracked by one of two interchangeable trackers behind a compile-time
|
||||
`damage_mode` A/B switch ([display.zig](../../system/services/display/display.zig)): a
|
||||
free-form dirty-rectangle **list** (tight bounds, heuristic merging) or a fixed 64-px
|
||||
**tile grid** (exact O(1) merging, tile-quantized repaints) — the grid is the default;
|
||||
[compositor.zig](../../system/services/display/compositor.zig) has both, with the trade-off
|
||||
discussion.
|
||||
|
||||
In v1 the surfaces are **server-owned**, and clients draw into them with a small
|
||||
immediate-mode command protocol — essentially the model early X used, and enough for a
|
||||
shell, a terminal, a cursor, and a wallpaper:
|
||||
|
||||
| Operation | Meaning |
|
||||
|--------------------|---------------------------------------------------------------|
|
||||
| `info` | report `{width, height, pitch, format}` of the display |
|
||||
| `create_layer` | allocate a server-owned surface, return a layer handle |
|
||||
| `configure_layer` | set a layer's rect, z-order, visibility |
|
||||
| `destroy_layer` | release a layer |
|
||||
| `fill_rect` | fill a rectangle of a layer with a colour |
|
||||
| `blit_tile` | copy a small client-supplied pixel tile into a layer (inline) |
|
||||
| `damage` | mark a region of a layer dirty |
|
||||
| `present` | request a repaint: composited at the next frame-clock tick |
|
||||
|
||||
Text is intentionally *not* an operation — a client renders glyphs by blitting tiles
|
||||
(the [PSF font](../../system/kernel/font.psf) path the console already uses can move into a
|
||||
client). Keeping the protocol to rectangles and tiles keeps the compositor small and the
|
||||
policy in the client.
|
||||
|
||||
`present` is a *request*, not an immediate flush: the compositor runs a ~60 Hz **frame
|
||||
clock** (a one-shot kernel timer re-armed on demand), and each tick composites all the
|
||||
damage accumulated since the last one. Any number of client presents and cursor moves
|
||||
inside one interval coalesce into a single repaint — the software stand-in for vblank
|
||||
pacing on backends that have none (all of them today; see
|
||||
[display-v2.md](display-v2.md), "Fenced is not vsync"). Bring-up paths that must put
|
||||
pixels on screen synchronously (initialisation, the self-checks) bypass the clock.
|
||||
|
||||
## `display`
|
||||
|
||||
Clients speak the protocol through a new [`library/client/display/display.zig`](../../library/client/display/display.zig),
|
||||
the [`block`](../../library/device/block/block.zig) shape (a cached `.display` lookup
|
||||
with a boot-race retry): `display.info()`, a `Layer` handle with `fill` / `blitTile` /
|
||||
`damage`, and `present()`. Application code never issues the raw syscalls — it calls the
|
||||
client module, as with every other danos service.
|
||||
|
||||
## The cursor: a mouse-listener thread feeding the compositor
|
||||
|
||||
The compositor is the single owner of the framebuffer — only the main `service.run` loop
|
||||
touches the backend and the layer stack. Tracking the mouse without breaking that
|
||||
ownership is the display's first use of [threads](../os-development/threading.md): the service is built
|
||||
multi-threaded (`addThreadedUserBinary`) and, at startup, spawns a **mouse-listener
|
||||
thread** beside the compositor loop.
|
||||
|
||||
- **Listener thread.** Blocks on the input service's mouse stream
|
||||
(`input.subscribeMouse()`), accumulates the relative `dx`/`dy` motion into an absolute
|
||||
cursor position clamped to the screen, and hands it to the compositor. It never touches
|
||||
the compositor — so no lock guards the framebuffer. A parked `next()` leaves its core
|
||||
free to halt ([halting.md](../os-development/halting.md)).
|
||||
- **The channel.** A single-slot *latest-value* cell (`CursorChannel`) guarded by a
|
||||
`Thread.Mutex`: the renderer wants where the cursor *is now*, not a replay of
|
||||
every delta, so a new position overwrites the old. The listener also **pokes** the
|
||||
compositor awake — the main loop is parked in `replyWait`, so the listener posts a
|
||||
zero-payload `ipc.send` to the compositor's endpoint, which arrives as a
|
||||
message-notification ([ipc.md](ipc.md)). The poke is *coalesced*: at most one is queued
|
||||
while the main loop has not drained the last, so a fast mouse cannot flood the endpoint.
|
||||
- **Render.** On the poke, the main loop takes the latest position and moves the cursor —
|
||||
which is just a top-z compositor layer — with the existing `configure` + `present` path
|
||||
(it damages the old and new footprints, so only those two rectangles repaint).
|
||||
|
||||
Two threading facts shape this (both in [threading.md](../os-development/threading.md)). IPC **handles do
|
||||
not cross threads**, so the listener can't reuse the main loop's endpoint handle — it
|
||||
`ipc.lookup(.display)`s its *own* handle to the same endpoint to poke through. And a
|
||||
multi-threaded service doing concurrent IPC is why the kernel's endpoint-create / register
|
||||
/ lookup syscalls now serialize under the big kernel lock. Shared fate applies: a fault in
|
||||
the listener takes the whole display down, and the supervisor restarts the process
|
||||
([resilience.md](../os-development/resilience.md)).
|
||||
|
||||
## What v1 does not do (and why that's fine)
|
||||
|
||||
Two capabilities are deliberately out of the first cut. Neither reshapes anything above;
|
||||
both are clean additions behind the interfaces v1 establishes.
|
||||
|
||||
- **Client-rendered surfaces (shared memory).** The fast path for a bitmap-heavy app is
|
||||
to render into its *own* buffer and hand the compositor a *reference*, not a stream of
|
||||
commands. That needs a cross-process shared-memory primitive — the natural
|
||||
generalization of the existing M13 [capability passing](driver-model.md)
|
||||
from *endpoints* to *memory objects* (`shared_memory_create(len) → {cap, virtual_address}`, pass `cap` on
|
||||
an `ipc_call`, receiver `shared_memory_map(cap) → virtual_address`). v1 avoids it because server-owned
|
||||
surfaces already prove the whole pipeline; v2 has since built exactly that primitive
|
||||
([display-v2.md](display-v2.md)) — the client-surface path on top of it is still open.
|
||||
|
||||
- **Runtime mode-setting (a native backend).** Detecting the EDID mode list and changing
|
||||
resolution / bpp at runtime needs the raw PCI device. The first native backend — since
|
||||
built by v2 ([display-v2.md](display-v2.md)) — is virtio-gpu, behind the same
|
||||
internal backend interface the dumb framebuffer sits behind. Refresh-rate and colour
|
||||
management (a gamma LUT) are real-GPU-KMS territory, far beyond this.
|
||||
|
||||
## Verifying it
|
||||
|
||||
Four QEMU test cases ([tests.zig](../../system/kernel/tests.zig), `python3
|
||||
test/qemu_test.py <case>`), each layering on the last:
|
||||
|
||||
- **`display`** — the kernel handoff: the seeded `display` device is shaped correctly and
|
||||
the claim → `mmio_map` leaf is genuinely **write-combining** (PAT entry 4), asserted at
|
||||
the page-table level.
|
||||
- **`display-service`** — the compositor comes up: it claims the framebuffer, allocates
|
||||
the cacheable back buffer, presents a cleared frame through the double-buffer path
|
||||
(`display: online … / presented frame 0`), and a startup **self-check** composites two
|
||||
overlapping layers on the real framebuffer and reads them back — overlap = the top
|
||||
layer — logging `display: compositor self-check ok`.
|
||||
- **`display-demo`** — the full pipeline from a separate process: the hardware-free
|
||||
[`display-demo`](../system/services/display-demo/) client (the
|
||||
[`input-source`](../test/system/services/input-source/) analog) drives layers — a wallpaper and
|
||||
a sliding rectangle — through the layer client API and heartbeats
|
||||
`display-demo: ok`, proving a frame travelled client → compositor → screen, exactly as
|
||||
the [input test](input.md) proves an event travels source → service → subscriber. It draws
|
||||
no cursor and reads no input — the cursor is the service's own (below), and the demo
|
||||
animates on its own frame timer, independent of the mouse (the test spawns `input`
|
||||
alongside it to keep that independence honest). The visible motion itself is a screenshot
|
||||
away via `zig build run-x86-64`.
|
||||
- **`display-cursor`** — the mouse-listener thread end to end: with the `input` service up,
|
||||
`input-source mouse` publishes pure motion, and the display's listener thread accumulates
|
||||
it into a cursor position handed to the render loop over the `CursorChannel`. Once the
|
||||
cursor has tracked a run of that motion, the service logs
|
||||
`display: cursor tracking mouse ok`. Runs `smp: 4` — the compositor and listener threads
|
||||
execute on different cores, which is what surfaced the IPC-under-lock requirement above.
|
||||
|
||||
The compositor's pixel math (rectangle clipping, fill, composite, tile blit) and colour
|
||||
packing are additionally covered by pure host unit tests under `zig build test`.
|
||||
|
||||
## See also
|
||||
|
||||
- [framebuffer.md](../os-development/framebuffer.md) — the linear framebuffer, pitch vs. width, `volatile`.
|
||||
- [gop.md](../os-development/gop.md) — GOP, and why only linear RGBX/BGRX modes are paintable.
|
||||
- [input.md](input.md) — the sibling service; the async `ipc_send` fan-out.
|
||||
- [driver-model.md](driver-model.md) — claim / `mmio_map`, capability passing, the trust model.
|
||||
- [device-manager.md](device-manager.md) — matching and supervision (the native backend's route).
|
||||
- [display-plan.md](display-plan.md) — the ordered build-out.
|
||||
@@ -0,0 +1,403 @@
|
||||
# The driver model: buses, classes, and host controllers
|
||||
|
||||
[drivers.md](drivers.md) shows how to write *a* driver — claim a device, map its
|
||||
registers, sleep on its interrupt. That's enough for a leaf device like the HPET. It is
|
||||
not enough for a disk, a keyboard, or a network card, because those hang off a
|
||||
*controller*, on a *bus*, speaking a *protocol*, and no single process should have to
|
||||
know all three.
|
||||
|
||||
Real driver stacks factor into three shapes. This document is about what each one is,
|
||||
what the kernel must give it, how they share code — and precisely which primitive each
|
||||
is still blocked on.
|
||||
|
||||
## Three shapes
|
||||
|
||||
| Shape | Owns | Reaches hardware by | Talks to |
|
||||
|---|---|---|---|
|
||||
| **Host controller driver** (HCD) | a controller — an xHCI PCI function, an AHCI port block | `mmio_map` + `irq_bind` + DMA | the devices behind it, in its bus's language |
|
||||
| **Bus driver** | a bus — a PCI bridge, a USB hub | `device_register`, to publish what it finds | class drivers, over IPC |
|
||||
| **Class / protocol driver** | *nothing* | *nothing* | its bus driver, over IPC |
|
||||
|
||||
The last row is the surprising one and the whole point. A USB keyboard driver touches
|
||||
no registers, takes no interrupts, and maps no memory. It sends HID protocol messages
|
||||
to whatever published the device, and it works identically whether the controller
|
||||
below is xHCI, EHCI, or a Raspberry Pi's DWC2. That is what buys you drivers that
|
||||
outlive the hardware they were written for.
|
||||
|
||||
In practice **HCD and bus driver are usually the same process**. An xHCI driver is a
|
||||
host controller driver (it owns the PCI function, its BARs, its interrupt, its DMA
|
||||
rings) *and* a bus driver (it enumerates USB devices and publishes them). Splitting
|
||||
them is a fiction; what matters is that both *roles* have kernel support, because a
|
||||
plain bus driver with no controller — a USB hub — is also a real thing.
|
||||
|
||||
## The device table is the spine
|
||||
|
||||
danos already has the right central structure. `system/kernel/devices-broker.zig` holds a table of
|
||||
`DeviceDescriptor`, each with a parent, a class, and a set of resources. Firmware discovery
|
||||
seeds it ([discovery.md](../os-development/discovery.md)); `device_register` grows it.
|
||||
|
||||
Three invariants make it a capability system rather than a directory:
|
||||
|
||||
1. **A claim is exclusive.** `device_claim(id)` succeeds once. Everything downstream —
|
||||
`mmio_map`, `irq_bind`, `device_register` — checks `devices_broker.ownerOf(id) == me`.
|
||||
2. **A descriptor is a licence to map physical memory.** Whoever claims a device may
|
||||
map its `.memory` resources and bind its `.irq` resources. This is why
|
||||
`device_register` cannot be a free-for-all.
|
||||
3. **Therefore: containment.** Every resource of a registered child must lie inside a
|
||||
resource of the same kind on its parent (`devices_broker.contains`). A bus driver can only
|
||||
ever *subdivide* what it already holds. Without this, `device_register` would be a
|
||||
syscall named "map any physical page you like."
|
||||
|
||||
Containment is transitive by construction: a grandchild is contained in its child,
|
||||
which is contained in the bus. Nothing can be laundered through a chain.
|
||||
|
||||
Note that firmware topology does **not** obey containment, and isn't asked to — a PCI
|
||||
function's BAR is not inside its host bridge's `bus_range`, because a bus-number range
|
||||
is not an address window. Discovery is trusted; user space is not.
|
||||
|
||||
### What a bus driver looks like
|
||||
|
||||
danos ships no demo bus driver — the real ones are `pci-bus`, `ps2-bus`, and
|
||||
`usb-xhci-bus`. The smallest *honest* shape, illustrated here with an HPET register block
|
||||
as the "bus" and its comparators as the "devices", is:
|
||||
|
||||
```zig
|
||||
_ = dev.claim(bus.id); // 1. own the bus
|
||||
const base = dev.mmioMap(bus.id, 0).?; // 2. enumerate it — from the hardware
|
||||
const n = ((cap.* >> 8) & 0x1F) + 1; // GENERAL_CAP says how many children
|
||||
|
||||
for (0..n) |i| { // 3. publish each child
|
||||
var child = std.mem.zeroes(dev.DeviceDescriptor);
|
||||
child.class = @intFromEnum(dev.DeviceClass.timer);
|
||||
child.resource_count = 1;
|
||||
child.resources[0] = .{ .kind = memory,
|
||||
.start = bus_mmio.start + 0x100 + 0x20 * i,
|
||||
.len = 0x20 };
|
||||
_ = dev.register(bus.id, &child).?; // kernel checks containment
|
||||
}
|
||||
```
|
||||
|
||||
Each child is left **unclaimed**, which is the handoff: a comparator driver can now
|
||||
`device_claim` one and `mmio_map` it, and will see only its own 0x20-byte window. A child
|
||||
whose window escapes the bus is refused; the in-kernel `containment` test asserts the
|
||||
kernel's table upholds that ([drivers.md](drivers.md)).
|
||||
|
||||
A USB device has *no* resources at all: `resource_count = 0`, because it's addressed
|
||||
through its controller, not by MMIO. That case is allowed and is the common one.
|
||||
|
||||
## Families: sharing code between drivers
|
||||
|
||||
A "family" is two modules, not one:
|
||||
|
||||
- **A logic module** — the parts of the bus that every driver on it re-derives. Config
|
||||
space walking and BAR decode for PCI. Descriptor parsing, control transfers, and hub
|
||||
protocol for USB.
|
||||
- **A protocol module** — the IPC message types that let a class driver talk to
|
||||
*whatever* published its device. This is the part that makes class drivers portable.
|
||||
|
||||
danos already has one of each: `library/device/pci/pci.zig` is a logic module (the
|
||||
`Function` view of a claimed PCI function),
|
||||
[`library/protocol/vfs/vfs-protocol.zig`](../../library/protocol/vfs/vfs-protocol.zig) is a
|
||||
protocol module shared by the mount backends (today the fat server) and their clients.
|
||||
(The user-space VFS server it was originally written against has since retired — path
|
||||
routing moved into the kernel, `system/kernel/vfs.zig`'s `fs_resolve` — but the protocol
|
||||
module outlived it, which is rather the point.) The pattern generalises directly:
|
||||
|
||||
```
|
||||
library/
|
||||
kernel/ the system library (kernel32-style): the syscall surface split by concern
|
||||
— ipc, memory (heap/dma/shared-memory), process, time, logging,
|
||||
file-system, thread, service, plus system-call stubs + start/root
|
||||
device/ device code grouped by domain; each domain splits into a shareable
|
||||
data module (enums/wire types, std-only) and a logic module (mmio/IPC)
|
||||
mmio/ module "mmio" — typed volatile register access + barriers [M14]
|
||||
model/ module "device-abi" — DeviceDescriptor, DeviceClass, ResourceKind
|
||||
pci/ "pci-class" (data) + "pci" — config/BAR/capability walk (Function)
|
||||
usb/ "usb-abi" + "usb-ids" (data) + "usb" — descriptors, control/interrupt/bulk client
|
||||
acpi/ "acpi-ids" (data) + "aml" — _HID names, the AML interpreter
|
||||
driver/ module "driver" — device-access syscalls + device-manager hello
|
||||
block/ module "block" — the block-device client (a device type)
|
||||
client/ userspace service clients — display, input (a program's view of a service)
|
||||
protocol/ driver <-> service wire contracts, one module per directory
|
||||
vfs/ block/ display/ scanout/ input/ power/ device-manager/ usb-transfer/
|
||||
|
||||
system/drivers/ one sub-project each → /system/drivers (no `d` suffix)
|
||||
usb-xhci-bus/ HCD + bus driver imports usb, mmio, usb-transfer-protocol (+ kernel modules)
|
||||
usb-hid/ class driver imports usb, input-protocol (+ kernel modules)
|
||||
virtio-gpu/ scanout driver imports pci, mmio, display-/scanout-protocol (+ kernel modules)
|
||||
```
|
||||
|
||||
The split by *dependency weight* is what lets the microkernel stay out of device
|
||||
business: it imports only the `device-abi` data module (the descriptor types its broker
|
||||
marshals across the syscall boundary) — never a logic module, never a taxonomy. That one
|
||||
pure-data import is the only edge from `system/kernel/` into `library/`; decoding a class
|
||||
code or `_HID` to a name is user space's job (the device manager owns those taxonomies).
|
||||
|
||||
A protocol lives in `library/protocol/` when it is the seam between a low-level driver and
|
||||
a higher-level service (block ↔ filesystem, a scanout driver ↔ the compositor). A driver's
|
||||
private wire to its *hardware* — virtio-gpu's command set — is not that; it stays a
|
||||
driver-private file, like the virtio-pci transport beside it.
|
||||
|
||||
The build side of this has since landed: [`addUserBinary`](build.zig) injects the
|
||||
default modules — the library/kernel concern modules (`ipc`, `memory`, `process`, `time`,
|
||||
`logging`, `file-system`, `thread`, `service`), the device/service clients (`driver`,
|
||||
`block`, `display`, `input`), plus `mmio`, `xkeyboard-config`, `acpi-ids` — into every user
|
||||
binary, and per-binary extras — protocol modules, bus logic — are added with
|
||||
`programModule(exe).addImport(...)`. That's the *entire* mechanism — Zig modules
|
||||
already give you everything else.
|
||||
|
||||
The discipline that makes this work: **a class driver must not import a bus's *hardware*
|
||||
logic module.** `usb-hid` imports `usb` (the transfer client) and `input-protocol`, never
|
||||
`pci` and never `mmio`. If a class driver needs `mmio`, it has become an HCD and should be
|
||||
one. The domain data modules (`usb-abi`, `usb-ids`, `pci-class`) carry no such weight — a
|
||||
class driver, the device manager, or the kernel may share them freely.
|
||||
|
||||
## What exists today
|
||||
|
||||
- **M10** — `device_enumerate`, `device_claim`, `mmio_map`. Strong-uncacheable device
|
||||
grants, `device_grant` teardown.
|
||||
- **M11** — `irq_bind` / `irq_ack`. IRQ delivered as an IPC notification; mask before
|
||||
EOI; `irq_ack` is the unmask.
|
||||
- **M12** — `parent` in `DeviceDescriptor`, `device_register` with resource containment.
|
||||
- **M13** — capability passing. `ipc_call` / `ipc_reply_wait` grew a `send_cap` argument
|
||||
and a `received_cap` return (r8): an endpoint travels with a message, installed into
|
||||
the receiver's handle table (shared, refcount-bumped — a copy, not a move). A full
|
||||
table fails `-ENOSPC` and does not half-deliver. This is the "open" primitive — a bus
|
||||
driver mints a per-device endpoint and hands it to a class driver. The `ipc` module
|
||||
exposes `callCap` and `replyWait(..., send_cap)`, and class drivers consume them now: the
|
||||
PS/2 keyboard and mouse drivers attach to ps2-bus this way, and the `usb` / `input`
|
||||
client modules open their per-device and subscription channels with `callCap`.
|
||||
- **M14** — DMA memory + the memory-ordering layer. `/lib/device/mmio` gives drivers typed
|
||||
volatile access and `memoryBarrier`/`readMemoryBarrier`/`writeMemoryBarrier` (per-arch); `dma_alloc`/`dma_free` grant
|
||||
physically-contiguous, pinned, uncacheable, reclaim-on-teardown buffers with the
|
||||
physical address exposed (`pmm.allocContiguous`, a DMA arena, `mapUserDmaInto`).
|
||||
`dma_below_4g` caps the address for legacy engines; `dma_write_combining` is accepted
|
||||
but falls back to coherent until PAT is programmed. The bus drivers use `/lib/device/mmio`,
|
||||
and `dma_alloc` has real consumers now: the xHCI driver's rings and contexts,
|
||||
usb-storage's command/status wrappers, virtio-gpu's virtqueue, and the fat
|
||||
service's bounce buffer.
|
||||
- **M15** — interrupts for PCI devices, the MSI half. Discovery now gives every PCI
|
||||
function its 4 KiB ECAM config space as resource 0 (unblocking the capability walk
|
||||
with no new syscall), and `msi_bind(device_id, endpoint) -> address, data` allocates a
|
||||
per-device edge-triggered vector, delivered as an IPC notification with no mask and no
|
||||
ack cycle. Legacy INTx (`_PRT` parsing + shared lines) is deliberately skipped — MSI
|
||||
is the real answer. QEMU's HPET has no MSI, so delivery is proven with a self-IPI; the
|
||||
first PCI driver is the first real consumer.
|
||||
- **Port I/O** — `io_read`/`io_write(device_id, resource_index, offset, width[, value])`:
|
||||
a claimed device's `io_port` resource lets a driver read/write its ports, gated exactly
|
||||
like `mmio_map` gates memory (direct ring-3 `in`/`out` stays a #GP). This is what makes
|
||||
a PS/2 or 16550 driver possible; the low-rate legacy hardware that needs it is fine with
|
||||
a syscall per access. `io_port` resources were recorded by discovery and ignored — now
|
||||
they're used.
|
||||
- **M16 (detection)** — the IOMMU is now *found*: discovery parses the ACPI DMAR table,
|
||||
maps the first VT-d unit, and reads its version + capabilities (`iommu_present` in the
|
||||
platform info). This is detection only — **no translation domains are programmed, so
|
||||
DMA is still unprotected** (the caveat below). Enforcement lands with the first DMA
|
||||
driver, which is what there is to protect and test against. Proven in the `iommu` test,
|
||||
booted with an emulated `intel-iommu`.
|
||||
- **`system_spawn`** — a user-space supervisor starts a driver:
|
||||
`system_spawn(name, arguments)` loads a binary bundled in the initial-ramdisk as a
|
||||
fresh ring-3 process; `name` becomes the child's argv[0] and the optional
|
||||
NUL-separated `arguments` blob its argv[1..], delivered on a SysV entry stack
|
||||
([sysv.md](../os-development/sysv.md)). This is what
|
||||
turned the device manager from "log the match" into "run the driver": the kernel now
|
||||
spawns only `init`, `init` spawns the services, and the **device-manager** discovers
|
||||
the hardware and spawns each driver ([drivers.md](drivers.md)). Ungated for now — a
|
||||
spawn capability is future work.
|
||||
|
||||
So: **all three shapes work now, and they're started by the device manager, not the
|
||||
kernel.** The xHCI driver is the HCD-and-bus proof; the USB HID/storage and PS/2 class
|
||||
drivers reach their devices purely over IPC. What follows are the original design notes
|
||||
for the primitives that unblocked each shape — exactly why each was the blocker, and
|
||||
exactly what fixed it.
|
||||
|
||||
---
|
||||
|
||||
# Proposed ABI
|
||||
|
||||
## M13 — capability passing, for class drivers ✅ done
|
||||
|
||||
*Implemented as described below (see "What exists today"). The signatures landed
|
||||
verbatim: `send_cap` in r9, `received_cap` returned in r8, `-ENOSPC` on a full receiver
|
||||
table with no delivery. The rest of this section is the original design note.*
|
||||
|
||||
**The blocker.** A class driver has to reach *its* device. Today the only way to find
|
||||
an endpoint is the name registry: `ipc_register(service_id, h)` / `ipc_lookup(id)`,
|
||||
where `ServiceId` is a global integer namespace with `max_services = 8`. You cannot
|
||||
mint one endpoint per USB device that way, and there is no way for a bus driver to
|
||||
*hand* a class driver an endpoint. M7 deferred this deliberately.
|
||||
|
||||
**The fix.** Let a message carry one handle. Sender names a handle in its own table;
|
||||
the kernel installs the endpoint into the receiver's table (bumping `refcount`) and
|
||||
tells the receiver the index it landed at.
|
||||
|
||||
```
|
||||
ipc_call(h, msg, message_len, reply, reply_cap, send_cap) -> reply_len
|
||||
ipc_reply_wait(h, reply, reply_len, recv, recv_cap, send_cap)
|
||||
-> recv_len (rax), badge (rdx), received_cap (r8)
|
||||
```
|
||||
|
||||
`send_cap` is a handle or `no_cap` (`~0`). `received_cap` is the index the transferred
|
||||
endpoint was installed at in the receiver's table, or `no_cap`.
|
||||
|
||||
- Both calls grow from 5 args to 6, which fits: `syscall5` uses `rdi/rsi/rdx/r10/r8`,
|
||||
leaving `r9`. `ipc_reply_wait` already returns two values via `setSyscallResult2`;
|
||||
this needs a third (`setSyscallResult3`).
|
||||
- If the receiver's handle table is full, the call fails `-ENOSPC` and **the message is
|
||||
not delivered** — a half-delivered capability is worse than a failed send.
|
||||
- `closeHandles` already drops references on exit, so the lifetime story is unchanged.
|
||||
|
||||
That single primitive gives you the standard `open` pattern:
|
||||
|
||||
```zig
|
||||
// class driver // bus driver
|
||||
const h = ipc.lookup(.usb).?; const r = ipc.replyWait(ep, ...);
|
||||
const dev_ep = ipc.callCap(h, // ... mint a per-device endpoint,
|
||||
.{ .op = .open, .id = dev_id }); // reply with it as send_cap
|
||||
// now dev_ep is a private channel to that one device
|
||||
```
|
||||
|
||||
## M14 — DMA memory and the memory-ordering contract, for HCDs ✅ done
|
||||
|
||||
*Implemented: `/lib/device/mmio` (typed volatile access + `memoryBarrier`/`readMemoryBarrier`/`writeMemoryBarrier`, per-arch) and
|
||||
`dma_alloc`/`dma_free` (contiguous, pinned, uncacheable, reclaim-on-teardown, physical
|
||||
address exposed). `dma_write_combining` still falls back to coherent — real WC needs
|
||||
PAT, a small follow-up. The rest of this section is the original design note.*
|
||||
|
||||
**The blocker.** An HCD is a DMA-engine programmer. It needs a descriptor ring the
|
||||
device can read, which means memory that is (a) physically contiguous, (b) at a
|
||||
physical address the driver knows, (c) of the right cacheability, and (d) pinned.
|
||||
[`sysMmap`](system/kernel/process.zig) gives you *none* of the four: it calls `pmm.alloc()`
|
||||
once per page, maps writeback-cached, and never reveals a physical address.
|
||||
|
||||
**The fix.**
|
||||
|
||||
```
|
||||
dma_alloc(len, flags) -> virtual_address (rax), physical_address (rdx)
|
||||
dma_free(virtual_address, len) -> 0
|
||||
|
||||
flags: dma_coherent (1) uncacheable; the default and the only one that's portable
|
||||
dma_wc (2) write-combining — needs PAT programmed; for framebuffers
|
||||
dma_below_4g (4) for devices with 32-bit DMA addressing
|
||||
```
|
||||
|
||||
Guarantees: page-aligned, physically contiguous, zeroed, pinned for the life of the
|
||||
mapping, and the physical address is stable. It needs one thing the kernel lacks —
|
||||
`pmm.allocContiguous(n, max_phys)`; today `pmm.alloc()` hands out one frame at a time
|
||||
with no adjacency guarantee.
|
||||
|
||||
**The memory-ordering contract.** danos has, at the time of writing, **zero memory
|
||||
barriers anywhere in the tree.** That is currently correct-by-accident and won't
|
||||
survive the first DMA driver, or the first ARM boot.
|
||||
|
||||
`volatile` is not a barrier. In Zig it means: don't elide this access, and don't
|
||||
reorder it against *other volatile* accesses. It says nothing about your *ordinary*
|
||||
stores — the descriptor you just filled in normal WB memory — which LLVM may freely
|
||||
sink past a volatile MMIO write. The canonical bug:
|
||||
|
||||
```zig
|
||||
ring[i] = descriptor; // ordinary store to WB RAM
|
||||
doorbell.* = i; // volatile store to UC MMIO
|
||||
// nothing stops the compiler reordering these; the device reads a stale descriptor
|
||||
```
|
||||
|
||||
So the rules, which belong in `library/device/mmio/mmio.zig` and behind `arch`:
|
||||
|
||||
| Situation | Required |
|
||||
|---|---|
|
||||
| MMIO register read/write | `mmio.read` / `mmio.write` (volatile) |
|
||||
| Fill DMA descriptor, then ring doorbell | `writeMemoryBarrier()` between them |
|
||||
| Woken by IRQ, then read what the device wrote | `readMemoryBarrier()` before the read |
|
||||
| MMIO write that must complete before the next read | `memoryBarrier()` |
|
||||
|
||||
And the per-arch lowering — the reason this must be an `arch` primitive and not a
|
||||
sprinkling of `asm volatile`:
|
||||
|
||||
| | x86_64 | aarch64 |
|
||||
|---|---|---|
|
||||
| `memoryBarrier()` | `mfence` | `dsb sy` |
|
||||
| `readMemoryBarrier()` | `lfence` | `dsb ld` |
|
||||
| `writeMemoryBarrier()` | `sfence` | `dsb st` |
|
||||
| DMA cache coherency | coherent; nothing to do | **not guaranteed**; needs non-cacheable buffers or cache maintenance |
|
||||
|
||||
x86 is forgiving here — TSO plus strong-uncacheable MMIO means you usually get away
|
||||
with a compiler barrier alone. ARM is not, and [vision.md](../vision.md) makes ARM the win
|
||||
condition. Build the abstraction while there is one caller to fix.
|
||||
|
||||
(Zig note: `@fence` was **removed in 0.16**. Use `@atomicRmw(..., .seq_cst)` for a full
|
||||
barrier, or per-arch inline asm — which is what `library/device/mmio/mmio.zig` should hide.)
|
||||
|
||||
## M15 — interrupts for PCI devices ✅ done (MSI)
|
||||
|
||||
*Implemented the MSI half: ECAM config space per PCI function (resource 0) and
|
||||
`msi_bind` (per-device edge-triggered vector, delivered as a notification). Legacy INTx
|
||||
`_PRT` parsing is skipped on purpose. `msi_bind` returns (address, data) as two values
|
||||
rather than an out-struct. The rest of this section is the original design note.*
|
||||
|
||||
**The blocker, and it's a hard one.** No PCI device can take an interrupt today.
|
||||
`addBars` (then in the kernel's ACPI discovery; BAR decode has since moved to the
|
||||
ring-3 pci-bus driver, `system/drivers/pci-bus/pci-bus.zig`) records `.memory` and
|
||||
`.io_port` BARs and never an `.irq`; there is no `_PRT` parsing anywhere in the tree.
|
||||
The HPET is the one exception —
|
||||
it advertises its own interrupt routing in its own registers, a privilege no ordinary
|
||||
device has.
|
||||
|
||||
**The fix, in two halves.**
|
||||
|
||||
*Legacy INTx*: parse `_PRT` from the DSDT to map (device, INTA–D) → GSI, and record it
|
||||
as an `.irq` resource. Then `irq_bind` works unchanged. But INTx lines are **shared**,
|
||||
and `irq.bound[gsi]` holds one endpoint. Sharing needs a list, and every driver on the
|
||||
line must be polled on each interrupt — the reason everyone left INTx behind.
|
||||
|
||||
*MSI/MSI-X*, which is the real answer: per-device vectors, edge-triggered, unshared, no
|
||||
mask/ack cycle, no 24-GSI ceiling. The kernel allocates a vector and hands the driver
|
||||
the (address, data) pair to program into its own MSI capability:
|
||||
|
||||
```
|
||||
msi_bind(dev_id, endpoint, out) -> 0 // out: extern struct { addr: u64, data: u32 }
|
||||
```
|
||||
|
||||
The driver writes those into config space itself — which means it needs config space,
|
||||
which means **discovery should give each `pci_device` a `.memory` resource for its
|
||||
4 KiB ECAM slot**. That's a small change to `parseMcfg` and it unblocks the whole
|
||||
capability walk (MSI, MSI-X, PCIe extended caps) without any new syscall.
|
||||
|
||||
Note QEMU's HPET reports `Tn_FSB_INT_DEL_CAP = 0` — no MSI — so an HPET timer could never
|
||||
exercise this path. The first MSI driver will be the first PCI driver.
|
||||
|
||||
## M16 — the IOMMU, and the honest caveat ◑ detection done, enforcement pending
|
||||
|
||||
*The IOMMU is now detected (DMAR parsed, VT-d unit mapped and read — see the `iommu`
|
||||
test), but **enforcement is not built**: no translation domains are programmed, so the
|
||||
caveat below still holds in full. Detection can't be taken further usefully until there
|
||||
is a DMA driver to protect and QEMU's `intel-iommu` to test the protection against —
|
||||
building the per-device domains alongside that first driver is both the natural order
|
||||
and the only way to verify them. The rest of this section is the original caveat.*
|
||||
|
||||
Everything above is capability-gated at the *CPU*. None of it is gated at the *device*.
|
||||
A driver that can program a bus-mastering engine can make that device write to any
|
||||
physical address, because page tables sit between the CPU and RAM, not between a device
|
||||
and RAM. Until VT-d/DMAR (or SMMU on ARM) is programmed from the DMAR table, **`device_claim`
|
||||
on any DMA-capable device is equivalent to granting ring 0.**
|
||||
|
||||
This does not make the model useless — it's the same position Linux is in with the
|
||||
IOMMU off, and every other guarantee (crash isolation, restart, no shared address
|
||||
space) still holds. But "user-space drivers are memory-safe" is not true yet, and the
|
||||
gap should be named rather than implied.
|
||||
|
||||
## Ordering
|
||||
|
||||
`M13` (capability passing) is independent of `M14`/`M15` and is the cheapest. It
|
||||
unlocks class drivers, which are the shape with no hardware requirements at all — you
|
||||
could write a real one against any device a bus driver publishes tomorrow.
|
||||
|
||||
`M14` and `M15` together unlock the first HCD. `M14`'s barrier layer is worth landing
|
||||
on its own regardless: it's small, obviously correct, and stops every future driver
|
||||
from hand-rolling `*volatile` and getting ARM wrong.
|
||||
|
||||
## See also
|
||||
|
||||
- [drivers.md](drivers.md) — how to write one, concretely.
|
||||
- [discovery.md](../os-development/discovery.md) / [acpi.md](../os-development/acpi.md) — where the device table comes from.
|
||||
- [ipc.md](ipc.md) — endpoints, badges, and the notification path an IRQ arrives on.
|
||||
- [resilience.md](../os-development/resilience.md) — restart, the reason any of this is worth the trouble.
|
||||
@@ -0,0 +1,407 @@
|
||||
# Writing a driver
|
||||
|
||||
In a monolithic kernel a driver is a function call away from everything: it runs in
|
||||
ring 0, dereferences any physical address, and its interrupt handler *is* the ISR. In
|
||||
danos a driver is **an ordinary ring-3 process**. It has its own address space, it
|
||||
can crash without taking the kernel with it, and — the point of this document — it
|
||||
can be restarted ([resilience](../os-development/resilience.md)).
|
||||
|
||||
That leaves three questions the kernel has to answer, because a process can't answer
|
||||
them for itself:
|
||||
|
||||
1. **What hardware exists?** → `device_enumerate`, over the device table discovery built
|
||||
([discovery](../os-development/discovery.md), [acpi](../os-development/acpi.md)).
|
||||
2. **How do I touch its registers?** → `device_claim` + `mmio_map`: the kernel maps the
|
||||
device's physical MMIO window into your address space, and from then on it's plain
|
||||
memory. No syscall per register access.
|
||||
3. **How do I find out it wants something?** → `irq_bind`: the interrupt is delivered
|
||||
to you as an IPC notification. You block; the hardware wakes you.
|
||||
|
||||
A driver is, in one sentence, *a process that sleeps until its device has something to
|
||||
say.*
|
||||
|
||||
## How a driver gets started: discover, match, spawn
|
||||
|
||||
Nothing in the kernel decides that the PCI host bridge needs the `pci-bus` driver — that
|
||||
is policy, and policy lives in user space. Boot brings user space up as a three-level
|
||||
supervision hierarchy, each level owning one job:
|
||||
|
||||
```
|
||||
kernel ──spawns──► init (PID 1) ──spawns──► device-manager ──spawns──► pci-bus
|
||||
| | |
|
||||
spawns only init, the service supervisor: the driver supervisor: enumerates
|
||||
publishes the starts the system /system/devices, matches each device
|
||||
initial-ramdisk services (device-manager, to a driver, and system_spawn's it
|
||||
so user space can fat, logger, ...). Its
|
||||
system_spawn from it list is init policy.
|
||||
```
|
||||
|
||||
The kernel launches exactly one process — `init` — and hands it nothing but the raw
|
||||
ability to start more (`system_spawn(name, arguments)`, which loads a binary bundled
|
||||
in the initial-ramdisk as a fresh ring-3 process — `name` becoming its argv[0],
|
||||
the optional arguments its argv[1..], on a SysV entry stack, see sysv.md). Everything else is a user-space decision:
|
||||
|
||||
- **init** ([system/services/init](system/services/init/init.zig)) is the **service
|
||||
supervisor**. It spawns the system services danos brings up at boot — today `input`,
|
||||
the `device-manager`, `fat`, `display`, `display-demo`, and the `logger` — from a
|
||||
small list. Drivers are deliberately *not* its job. (An earlier draft listed a `vfs`
|
||||
service here; that service is retired — the router moved into the kernel as
|
||||
`fs_resolve`.)
|
||||
- **device-manager** ([system/services/device-manager](system/services/device-manager/device-manager.zig))
|
||||
is the **driver supervisor**. It does the three steps a monolithic kernel would do in
|
||||
its probe path, entirely from ring 3:
|
||||
1. **Discover** — `device_enumerate` snapshots the device table the kernel built from
|
||||
ACPI/PCI ([discovery](../os-development/discovery.md)).
|
||||
2. **Match** — for each device it looks up a driver. The match policy is code, a few
|
||||
small per-bus tables: from the boot snapshot only the PCI host bridge matches
|
||||
(→ `pci-bus`); everything else arrives later as bus reports and matches on
|
||||
identity — `pciDriverForIdentity` (xHCI → `usb-xhci-bus`, virtio-gpu →
|
||||
`virtio-gpu`), `hidDriverFor` (PNP0303/PNP0F13 → `ps2-bus`), and
|
||||
`usbDriverForIdentity` (USB keyboard, mouse, storage). A fuller system reads
|
||||
what each driver *binds* (a manifest under `/system/drivers`, or the driver
|
||||
describing its own match).
|
||||
3. **Spawn** — `system_spawn(driver_name, arguments)` starts the matched driver (the
|
||||
arguments can carry *which* device it matched), which then claims
|
||||
its device and runs the event loop below.
|
||||
|
||||
So "how is a driver discovered and configured" has two halves: **discovery** is the
|
||||
kernel's device table, read by anyone; **configuration** is two user-space policies —
|
||||
init's service list and the device-manager's match table. Both are hardcoded in their
|
||||
respective programs today; the natural next step is to move them into `/etc` (see the
|
||||
milestone notes in [driver-model.md](driver-model.md)). `system_spawn` is currently
|
||||
ungated — any process may spawn any bundled binary — because there is no spawn
|
||||
capability yet.
|
||||
|
||||
## The capability: claim before touch
|
||||
|
||||
The driver syscall numbers (`system/abi.zig`) with the device types they carry
|
||||
(`library/device/model/device-abi.zig`), dispatched in `system/kernel/process.zig`:
|
||||
|
||||
| # | Call | Meaning |
|
||||
|---|------|---------|
|
||||
| 11 | `device_enumerate(buf, max) -> total` | Snapshot the device table |
|
||||
| 12 | `device_claim(id) -> ok` | Take **exclusive** ownership |
|
||||
| 13 | `mmio_map(id, res_idx) -> virtual_address` | Map a claimed device's register window |
|
||||
| 14 | `irq_bind(id, res_idx, endpoint)` | Deliver that device's IRQ as a notification |
|
||||
| 15 | `irq_ack(id, res_idx)` | Re-arm the IRQ after servicing the device |
|
||||
| 16 | `device_register(parent_id, desc) -> id` | Publish a child of a device you claimed |
|
||||
|
||||
Notice that **nothing takes a physical address or an interrupt number.** Every call
|
||||
names a device by id and a resource by index. That indirection is the entire security
|
||||
model. If `mmio_map` took a physical address, any process could map the kernel's
|
||||
memory; if `irq_bind` took a GSI, any process could bind the keyboard's line and
|
||||
silently intercept it. Instead the kernel checks two things (`process.ownedGsi`, and
|
||||
the same check at the top of `systemMmioMap`):
|
||||
|
||||
- `devices_broker.ownerOf(dev_id) == me` — you claimed it, and claims are exclusive
|
||||
- the resource at `res_idx` is of the right *kind* — `memory` for `mmio_map`, `irq`
|
||||
for `irq_bind`
|
||||
|
||||
The claim is the capability. Everything else follows from it.
|
||||
|
||||
## Registers: `mmio_map`
|
||||
|
||||
`mmio_map` walks the caller's page tables and installs the device's physical frames
|
||||
with `present | user | writable | nx | device_grant` plus a cache mode
|
||||
(`system/kernel/architecture/x86_64/paging.zig:mapUserDeviceInto`). Two of those bits
|
||||
are load-bearing:
|
||||
|
||||
- **`pcd | pwt`** — strong-uncacheable, the default cache mode. A device register is
|
||||
not memory; a cached read would return a stale value and a write might never leave
|
||||
the CPU. The one exception: a resource flagged write-combining
|
||||
(`resource_flag_write_combining` — today the kernel-seeded display framebuffer)
|
||||
gets the PAT bit instead, so pixel writes batch into bursts.
|
||||
- **`device_grant`** (bit 9, one of the PTE's available bits) — marks the leaf as MMIO
|
||||
rather than RAM, so `freeSubtree` skips `pmm.free` on it when the address space is
|
||||
destroyed. Without this, killing a driver would hand the HPET's registers back to
|
||||
the frame allocator as if they were free RAM. The `iopass` test guards it.
|
||||
|
||||
Grants land in their own arena, `0x0000_7100_0000_0000` (PML4[226]), so device pages
|
||||
never widen an existing mapping.
|
||||
|
||||
Then you just… use it:
|
||||
|
||||
```zig
|
||||
const base = dev.mmioMap(dev_id, mmio_res) orelse return;
|
||||
const counter: *volatile u64 = @ptrFromInt(base + 0xF0);
|
||||
const now = counter.*; // a load, straight to the hardware. no kernel involved.
|
||||
```
|
||||
|
||||
## Interrupts: the cycle, and why it has that shape
|
||||
|
||||
An interrupt handler in a microkernel has a problem. The code that knows how to quiet
|
||||
the device is in ring 3, in another address space, and it will not run for
|
||||
microseconds or milliseconds — after a context switch, when the scheduler gets to it.
|
||||
But the CPU wants an EOI *now*, and a **level-triggered** line stays asserted until
|
||||
the device is quieted. EOI a still-asserted line and the I/O APIC redelivers
|
||||
immediately. Forever. The driver never gets to run at all.
|
||||
|
||||
The way out is to mask the line before acknowledging it:
|
||||
|
||||
```
|
||||
kernel ISR irqMask(gsi) // line still asserted; stop it reaching a CPU
|
||||
irqEoi() // now safe to tell the LAPIC we're done
|
||||
notifyFromIsr() // wake the driver — it runs much later
|
||||
|
||||
driver replyWait() -> badge with the notify bit set
|
||||
<clear the device's status register> // NOW the line deasserts
|
||||
irq_ack(dev, res) // kernel unmasks: quiet, so it can't refire
|
||||
|
||||
```
|
||||
|
||||
`irq_ack` is not bookkeeping you could skip. **It is the unmask.** Forget it and the
|
||||
interrupt fires exactly once, ever; call it before the device is quiet and you get an
|
||||
interrupt storm. That single fact explains why `irq_bind` and `irq_ack` are two
|
||||
syscalls and not one.
|
||||
|
||||
This is also why `interruptDispatch` (`system/kernel/architecture/x86_64/idt.zig`) no longer issues the EOI
|
||||
itself. It used to, before running the handler — correct for the LAPIC timer, and
|
||||
impossible for a routed device line. Each handler now owns its EOI, because only the
|
||||
handler knows which discipline its source needs.
|
||||
|
||||
### The driver side is an event loop, not a callback
|
||||
|
||||
`IPC_ReplyWait` returns *either* a client request *or* a notification, told apart by
|
||||
the top bit of the badge (`ipc_sync.notify_badge_bit`). So a driver is one
|
||||
single-threaded loop over both of its event sources:
|
||||
|
||||
```zig
|
||||
while (true) {
|
||||
const r = ipc.replyWait(endpoint, reply, &recv);
|
||||
if (r.isNotification()) { // r.source() is the GSI
|
||||
service_device(); // clear the status register
|
||||
_ = dev.irqAck(id, irq_res); // re-arm
|
||||
} else {
|
||||
handle_client_request(recv[0..r.len]);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
No reentrancy, no "what am I allowed to call from an interrupt handler", no shared
|
||||
state between ISR and task context. The interrupt is just a message.
|
||||
|
||||
Two properties worth knowing:
|
||||
|
||||
- **An interrupt taken while you're elsewhere is not lost.** If the driver is off in
|
||||
an `ipc_call` to another server when the IRQ fires, `wakeLocked` finds nobody
|
||||
waiting, but the badge is already on the endpoint's notify ring. The next
|
||||
`replyWait` pops it (`ipc_sync.replyWait` checks `popNotify` before the sender FIFO).
|
||||
- **Notifications coalesce, they don't count.** The ring is 8 deep and drops on
|
||||
overflow. That's correct: an IRQ notification is a *level* ("the device wants
|
||||
attention"), not a tally. Re-read the device's status register; never assume one
|
||||
notification means exactly one event. Because the ISR masks the line until you
|
||||
`irq_ack`, at most one badge per GSI can be outstanding — so the ring can only
|
||||
overflow if you bind more than eight GSIs to a single endpoint. Don't.
|
||||
|
||||
## A whole driver
|
||||
|
||||
A minimal leaf driver is only ~150 lines and does all of it. danos ships **no such
|
||||
example binary** — the driver model is proven by the real drivers (`pci-bus`, `ps2-bus`,
|
||||
`usb-xhci-bus`), and a teaching example belongs here, in the docs, rather than as a
|
||||
compiled program nobody runs. Illustrated with a hypothetical HPET timer driver, the
|
||||
shape is:
|
||||
|
||||
```zig
|
||||
const hpet = findHpet(buf) orelse return; // device_enumerate, look for
|
||||
// class=timer with memory + irq
|
||||
_ = dev.claim(hpet.dev_id); // the capability
|
||||
const base = dev.mmioMap(hpet.dev_id, hpet.mmio).?;
|
||||
const endpoint = ipc.createIpcEndpoint().?;
|
||||
|
||||
// program the hardware over the mapping we were just handed
|
||||
reg(base, 0x100).* = level | int_enb | (hpet.gsi << 9); // timer 0 config
|
||||
reg(base, 0x108).* = reg(base, 0xF0).* + period; // comparator
|
||||
reg(base, 0x010).* |= 1; // ENABLE
|
||||
|
||||
_ = dev.irqBind(hpet.dev_id, hpet.irq, endpoint);
|
||||
|
||||
while (...) {
|
||||
const r = ipc.replyWait(endpoint, &.{}, &recv); // blocked. not polling.
|
||||
if (r.badge & notify_bit == 0) continue;
|
||||
reg(base, 0x020).* = 1; // clear status -> deassert
|
||||
reg(base, 0x108).* = reg(base, 0xF0).* + period; // re-arm
|
||||
_ = dev.irqAck(hpet.dev_id, hpet.irq); // unmask
|
||||
}
|
||||
```
|
||||
|
||||
The HPET makes a good illustration for a reason that isn't obvious. Its *counter* is a
|
||||
clocksource — the only way to use it is to read it, so it exercises `mmio_map` without
|
||||
needing interrupts at all. Its *comparators* are a clockevent, and can be configured
|
||||
**level-triggered** (`Tn_INT_TYPE_CNF`), which asserts a bit in `GENERAL_INT_STATUS`
|
||||
that the driver must write-1-to-clear. That's a genuine deassert step, so the full
|
||||
mask/ack cycle above is exercised for real rather than being decoration on an
|
||||
edge-triggered line that would have been fine without it.
|
||||
|
||||
One wrinkle it also demonstrates: the ACPI HPET table carries **no interrupt number**.
|
||||
Which I/O APIC inputs a comparator may drive is a bitmask in `Tn_INT_ROUTE_CAP`, in
|
||||
the device's own registers. So discovery (`acpi.parseHpet`) maps the block, reads the
|
||||
mask, and records one concrete GSI as an `irq` resource. The driver then programs
|
||||
`Tn_INT_ROUTE_CNF` to raise exactly that line — and the kernel will only bind the one
|
||||
it recorded. Hardware that describes itself at runtime still has to fit through a
|
||||
static capability.
|
||||
|
||||
## Publishing children: `device_register`
|
||||
|
||||
A device that *contains other devices* — a PCI bridge, a USB hub, or the HPET's block
|
||||
of comparators — needs a driver that enumerates it and tells the kernel what it found.
|
||||
That's `device_register`, and it makes the device table a tree rather than a list
|
||||
(`DeviceDescriptor.parent`).
|
||||
|
||||
```zig
|
||||
var child = std.mem.zeroes(dev.DeviceDescriptor);
|
||||
child.class = @intFromEnum(dev.DeviceClass.timer);
|
||||
child.resource_count = 1;
|
||||
child.resources[0] = .{ .kind = memory, .start = bus_base + 0x100, .len = 0x20 };
|
||||
const child_id = dev.register(bus_id, &child).?;
|
||||
```
|
||||
|
||||
The child is left **unclaimed**, which is the whole point: another process claims it and
|
||||
`mmio_map`s it, and sees only that 0x20-byte window.
|
||||
|
||||
The rule the kernel enforces is **containment**: every resource of a child must lie
|
||||
inside a resource of the same kind on its parent. Ranges must nest; a child's IRQ —
|
||||
still exactly one line — must fall within the parent's IRQ range (a length-1 parent
|
||||
range is the old exact-match rule). This isn't bureaucracy — a `DeviceDescriptor` is a
|
||||
licence to map physical memory, so without containment `device_register` would be a
|
||||
syscall for mapping any page you like. A bus driver may only ever subdivide what it
|
||||
already owns.
|
||||
|
||||
A device with **no resources** is legal and common. A USB device is reached through its
|
||||
controller, not by MMIO, so it gets `resource_count = 0`.
|
||||
|
||||
See [`system/drivers/pci-bus/pci-bus.zig`](../../system/drivers/pci-bus/pci-bus.zig) for a
|
||||
real one — it claims a PCI host bridge, maps its ECAM window, and publishes each function
|
||||
it finds as a child — and [driver-model.md](driver-model.md) for how bus drivers, class
|
||||
drivers and host controller drivers fit together.
|
||||
|
||||
## What the kernel does not do for you
|
||||
|
||||
- **It does not quiet your device.** That's the whole reason `irq_ack` exists.
|
||||
- **It does not know your registers.** `mmio_map` hands you a base address; every
|
||||
offset in this document came from the HPET spec, not from danos.
|
||||
- **It does not serialise your driver.** Two clients calling one driver endpoint are
|
||||
serialised by `replyWait`, but nothing stops your driver from being preempted.
|
||||
|
||||
## Limits, today
|
||||
|
||||
Worth knowing before you write the second driver:
|
||||
|
||||
Several things this list used to warn about are now available (see
|
||||
[driver-model.md](driver-model.md)): **port I/O** (`io_read`/`io_write`, claim-gated by
|
||||
the device's `io_port` resource — direct ring-3 `in`/`out` is still a #GP, so a PS/2 or
|
||||
16550 driver goes through these), **DMA memory** (`dma_alloc`: contiguous, pinned,
|
||||
uncacheable, physical address exposed), **memory barriers** (`library/device/mmio`'s
|
||||
`memoryBarrier`/`readMemoryBarrier`/`writeMemoryBarrier`, imported as the `mmio` module), **fault isolation** (a ring-3 fault kills only the faulting
|
||||
process — `killCurrentProcess` — and the machine keeps running,
|
||||
[resilience](../os-development/resilience.md)), and **reclaim + restart on death** (every path out of a
|
||||
process releases its claims and IRQ/MSI bindings — `releaseAllOwnedBy`,
|
||||
`irq.releaseOwner` — and the device manager respawns the driver with backoff,
|
||||
[device-manager.md](device-manager.md)). What remains:
|
||||
|
||||
- **Page granularity.** `mmio_map` rounds to 4 KiB. Two devices sharing a page means
|
||||
granting one grants the other. A `device_register`ed child's *resource* can be narrower
|
||||
than a page, but its *mapping* can't.
|
||||
- **DMA is not contained.** A driver that can program a bus-mastering device can make
|
||||
that device write to *any* physical address — page tables don't sit between a device
|
||||
and RAM; an IOMMU does. The IOMMU is now *detected* (M16), but no translation domains
|
||||
are programmed, so `device_claim` on a DMA-capable device is still effectively
|
||||
equivalent to granting ring 0. This is the largest gap between the design's promise and
|
||||
what it delivers; enforcement lands with the first DMA driver.
|
||||
- **No voluntary `dev_release`.** A *live* driver can't drop a claim — only exit
|
||||
releases it (any path out of a process runs `releaseAllOwnedBy`) — so handing a
|
||||
device between running drivers still means exiting.
|
||||
- **One endpoint per GSI**, so shared legacy PCI INTx lines can't be split between two
|
||||
drivers. MSI/MSI-X — one vector per device, edge-triggered, unshared — is the real
|
||||
answer, and QEMU's HPET doesn't offer it (`Tn_FSB_INT_DEL_CAP = 0`).
|
||||
- **Polarity is hardcoded** active-high in `irq.bind`. A device whose MADT override
|
||||
says active-low needs that threaded through from discovery.
|
||||
- **14 device vectors** (33–46) and **24 GSIs**, bounded by the stubs `isr.s` emits and
|
||||
by a single I/O APIC.
|
||||
- **Don't bind more than 8 GSIs to one endpoint.** The notify ring is 8 deep and drops
|
||||
on overflow. With one GSI per endpoint that's unreachable — the line is masked from
|
||||
the ISR until `irq_ack`, so at most one badge is ever outstanding. Bind nine devices
|
||||
to one endpoint, though, and a dropped badge leaves that line masked with nobody
|
||||
left to ack it.
|
||||
- **On real hardware, the mask/EOI cycle may need a remote-IRR flush.** Masking a
|
||||
level-triggered redirection entry with remote-IRR set doesn't clear it on some
|
||||
chipsets, and the line never fires again. QEMU clears it on EOI regardless, so the
|
||||
tests can't see this. Linux flushes remote-IRR by toggling the entry to edge and
|
||||
back. See the note at the top of `system/kernel/irq.zig`.
|
||||
|
||||
## Verifying it
|
||||
|
||||
No demo driver ships to prove this end to end; the *real* drivers do, so the tests
|
||||
target them and the kernel primitives directly:
|
||||
|
||||
- **`device-manager`** — boots only the device manager, which discovers the PCI host
|
||||
bridge, matches `pci-bus`, and `system_spawn`s it. The test reads kernel state — the
|
||||
process table and the device tree — to confirm pci-bus came up and registered the
|
||||
functions it enumerated: the whole discover → match → spawn → driver-up chain.
|
||||
- **`acpi-ps2`** — a user-space driver (`ps2-bus`) is woken by its device's IRQ,
|
||||
delivered as an IPC notification, and attaches the keyboard: IRQ-as-IPC, end to end.
|
||||
- **`pci-scan`** — a user-space driver (`pci-bus`) maps its device's MMIO (the ECAM
|
||||
window) and walks it: `mmio_map`, end to end.
|
||||
- **`containment`** — the kernel refuses a `device_register` whose child window escapes
|
||||
the parent's grant (else it would be a syscall for mapping arbitrary memory), while an
|
||||
identical re-register stays idempotent. Asserted in-kernel, straight against the broker.
|
||||
- **`irqfree`** — the teardown path. Binds two owners to one shared endpoint, releases
|
||||
one, and reads the I/O APIC back: the departing owner's line is masked, the sibling's
|
||||
is not. That second half is why bindings are keyed on the owning *task* and not on the
|
||||
endpoint pointer — endpoints are shared, so releasing "everything pointing at this
|
||||
endpoint" would silently mask a live driver's device.
|
||||
- **`iopass`** — the `device_grant` teardown rule, so destroying a driver's address
|
||||
space never returns MMIO frames to the RAM pool.
|
||||
|
||||
```
|
||||
$ python3 test/qemu_test.py device-manager acpi-ps2 pci-scan containment irqfree iopass
|
||||
device-manager ... PASS (matched 'DANOS-TEST-RESULT: PASS')
|
||||
acpi-ps2 ... PASS
|
||||
pci-scan ... PASS (matched 'DANOS-TEST-RESULT: PASS')
|
||||
containment ... PASS (matched 'DANOS-TEST-RESULT: PASS')
|
||||
irqfree ... PASS (matched 'DANOS-TEST-RESULT: PASS')
|
||||
iopass ... PASS (matched 'DANOS-TEST-RESULT: PASS')
|
||||
```
|
||||
|
||||
## What's next (not done here)
|
||||
|
||||
The big driver-model pieces — capability passing (class drivers), DMA + barriers, MSI,
|
||||
and IOMMU detection — are **now done** ([driver-model.md](driver-model.md), M13–M16), as
|
||||
is **port I/O** (`io_read`/`io_write`, the claim-gated syscalls that make a PS/2 or 16550
|
||||
driver possible). What's left is IOMMU *enforcement* (per-device domains — it waits on
|
||||
the first DMA driver to protect and test against) and these smaller items:
|
||||
|
||||
- **Releasing a claim** — half done. The kernel now drops *all* of a dead driver's
|
||||
claims on every path out of a process (`releaseAllOwnedBy`, called from process
|
||||
teardown), which unblocked restart. A voluntary `dev_release` for a live driver
|
||||
still doesn't exist.
|
||||
- **Unregistering children** — half done. Hot-remove works at the manager layer:
|
||||
the xHCI bus reports `child_removed` on unplug and the device manager prunes its
|
||||
tree. The kernel's own device table is still append-only, so a bus driver in a
|
||||
loop can still exhaust the 64-entry table.
|
||||
- **Restart** — done. The device manager notices a driver's death, reads its exit
|
||||
reason, prunes the children it reported, and respawns it with exponential
|
||||
backoff — with a crash-loop cap that marks a repeat offender `failed` instead
|
||||
of respawning forever ([device-manager.md](device-manager.md)).
|
||||
- **Interrupt priority / threaded IRQ latency** — still open. `notifyFromIsr`
|
||||
enqueues the woken driver but doesn't preempt (`wakeLocked` deliberately leaves
|
||||
that to the caller), so a woken driver waits for the next scheduling point.
|
||||
|
||||
## The driver contract (M17–M18)
|
||||
|
||||
Claiming and mapping is half of being a danos driver; the other half is the
|
||||
**lifecycle and protocol contract**, and the runtime makes it nearly free:
|
||||
|
||||
- Build on `service.run` — one replyWait loop folding protocol
|
||||
requests, signals, and notifications into callbacks. The harness answers the
|
||||
universal zero-length ping and turns `terminate` into a clean exit for you
|
||||
([process-lifecycle.md](../os-development/process-lifecycle.md)).
|
||||
- A driver spawned with an assignment (its device id as argv[1]) sends the
|
||||
versioned `hello` to the device manager inside the deadline, and a **bus**
|
||||
driver reports what it discovers with `child_added`
|
||||
([device-manager.md](device-manager.md); usb-xhci-bus is the reference
|
||||
implementation).
|
||||
- Crash freely — that is the design. The kernel releases your claims, IRQ
|
||||
bindings, and MSI vectors at death; the manager reads your exit reason,
|
||||
prunes what you reported, restarts you with backoff, and your fresh instance
|
||||
re-claims and re-reports. Never depend on your own cleanup running
|
||||
(iron rule 1).
|
||||
@@ -0,0 +1,164 @@
|
||||
# The input module: broadcasting input events
|
||||
|
||||
A keyboard driver has one keystroke and *many* programs that might want it — a shell, a
|
||||
window server, a logger. None of them owns the hardware, and the driver should not know
|
||||
who is listening. So between the drivers and the listeners sits the **input service**
|
||||
(`system/services/input/`): drivers **publish** events to it, programs **subscribe**, and
|
||||
it fans each event out to every interested subscriber. It is an ordinary ring-3 process
|
||||
reached over IPC, like the [FAT server](../../system/services/fat/fat.zig) — no kernel knows
|
||||
what a key is.
|
||||
|
||||
## One service, several device classes
|
||||
|
||||
The service carries three device classes today — **keyboard**, **mouse**, and
|
||||
**joystick/gamepad** — and is built to take more
|
||||
([protocol.zig](../../library/protocol/input/input-protocol.zig)). Each class has its own typed
|
||||
event:
|
||||
|
||||
- `KeyEvent` — `key_down`/`key_up` (physical make/break) and `key_press` (a character was
|
||||
produced, carrying the Unicode scalar); plus a layout-independent `keycode` and a
|
||||
`modifiers` bitmask.
|
||||
- `MouseEvent` — relative `motion` (`dx`/`dy`), `button_down`/`button_up`, and `scroll`.
|
||||
- `JoystickEvent` — `axis` moves (a signed value on a `control` index) and
|
||||
`button_down`/`button_up`.
|
||||
|
||||
All three travel in one **`InputEvent` envelope** tagged with a `DeviceKind`, so the
|
||||
fan-out is a single code path and a subscriber can take a mix of classes on one stream.
|
||||
Decode an envelope with `asKeyboard()` / `asMouse()` / `asJoystick()` (each returns null
|
||||
unless the tag matches). A subscriber names the classes it wants with a **`device_mask`**,
|
||||
and the service routes each event only to subscribers whose mask includes its class — so a
|
||||
mouse-only listener never wakes for keystrokes.
|
||||
|
||||
## Why this needed a new kernel primitive
|
||||
|
||||
The interesting part is delivery, and it runs straight into the shape of danos IPC.
|
||||
[ipc.md](ipc.md) describes a **synchronous rendezvous**: a server holds exactly one
|
||||
pending reply (`Task.ipc_client`) and *must* answer it on its next `replyWait`. Two
|
||||
consequences decide the whole design:
|
||||
|
||||
1. **You cannot block N subscribers waiting for "the next event".** A server can hold only
|
||||
one caller at a time, so the natural "subscriber calls `next_event()` and blocks" API
|
||||
is impossible for more than one subscriber. Delivery therefore has to be **push** — the
|
||||
service reaching out to subscribers — not pull.
|
||||
|
||||
2. **A synchronous push can hang the whole service.** If the service delivered with
|
||||
`ipc_call`, it would block until each subscriber replied. `ipc_call` has no timeout, and
|
||||
a subscriber's endpoint is an *unregistered* capability the kernel's death path cannot
|
||||
reach (since display v2's V6, `killOwnedEndpointsLocked` in
|
||||
[ipc-synchronous.zig](../../system/kernel/ipc-synchronous.zig) marks a dead owner's
|
||||
*registered* endpoints dead and wakes parked callers with `-EPEER` — but unregistered
|
||||
ones just drop with the task's handle table). One subscriber that exits mid-delivery
|
||||
would wedge input for everyone. That is the opposite of the resilience the microkernel
|
||||
is for.
|
||||
|
||||
The fix is the asynchronous send that [ipc.md](ipc.md) had already earmarked as future
|
||||
work ("asynchronous / buffered send … for notifications between servers"):
|
||||
|
||||
```
|
||||
ipc_send(handle, message_ptr, message_len) -> 0 / -errno
|
||||
```
|
||||
|
||||
`ipc_send` copies a small payload into the endpoint's **bounded queue** and wakes a
|
||||
receiver, then returns immediately — it never blocks and so can never hang on a dead or
|
||||
slow subscriber. The receiver picks it up through the same `replyWait` it already runs:
|
||||
the wake arrives as a **buffered message** — `notify_badge_bit | notify_message_bit` set in
|
||||
the badge (distinguishing it from a bare IRQ/child-exit notification), the sender's task id
|
||||
in the low bits, and the payload in the receive buffer, with no reply owed. The queue holds
|
||||
16 messages per endpoint; a full queue **drops the oldest**, because a buffered message is
|
||||
discrete data, not a coalescing "level" like an interrupt. See
|
||||
[ipc-synchronous.zig](../../system/kernel/ipc-synchronous.zig) (`sendLocked`, `popPost`, and
|
||||
the `replyWait` receive loop).
|
||||
|
||||
This is the async counterpart of `ipc_call`, and the input service is its first consumer.
|
||||
|
||||
## How the pieces fit
|
||||
|
||||
```
|
||||
keyboard/mouse driver, input-source input service subscriber(s)
|
||||
----------------------------------- ------------- -------------
|
||||
connectSource(); loop: replyWait: subscribeKeyboard()/…All:
|
||||
publishKeyboardEvent(k) ─ ipc_call ─▶ publish → broadcast: createIpcEndpoint()
|
||||
publishMouseEvent(m) for each sub whose callCap(subscribe,
|
||||
publishJoystickEvent(j) mask matches event.device: send_cap = ep,
|
||||
ipc_send(sub_ep) ──────▶ device_mask)
|
||||
reply ok loop: next()
|
||||
subscribe → store {ep cap, └─ replyWait(ep)
|
||||
task id, device_mask} → InputEvent
|
||||
```
|
||||
|
||||
- A **subscriber** calls `input.subscribe(mask)` — or a typed helper: `subscribeKeyboard()`,
|
||||
`subscribeMouse()`, `subscribeJoystick()` (one class, `next()` returns the decoded event),
|
||||
or `subscribeAll()` (every class, `next()` returns a tagged `InputEvent`)
|
||||
([library/client/input/input.zig](../../library/client/input/input.zig)). It creates its own endpoint
|
||||
and hands it to the service as a **capability** (M13 capability passing — the input
|
||||
service is that feature's first real user), along with its `device_mask`. Then it loops on
|
||||
`next()`, a `replyWait` on that endpoint returning each pushed event.
|
||||
- A **source** (a keyboard, mouse, or joystick driver) calls `input.connectSource()` and the
|
||||
method for its class: `publishKeyboardEvent`, `publishMouseEvent`, or
|
||||
`publishJoystickEvent`. Publishing is a short synchronous `ipc_call` the service answers at
|
||||
once; the service's own fan-out is asynchronous, so publishing never blocks on a slow
|
||||
subscriber.
|
||||
- The **service** ([input.zig](../../system/services/input/input.zig)) keeps a small subscriber
|
||||
table (endpoint handle + owning task id + `device_mask`). On `publish` it `ipc_send`s the
|
||||
event to every subscriber whose mask includes the event's device class. On `subscribe` it
|
||||
stores the passed capability and mask and, as housekeeping, prunes any slot whose owning
|
||||
process has exited (checked against `process_enumerate`) — not for correctness (an async
|
||||
send to an orphaned endpoint is harmless) but to reclaim the slot.
|
||||
|
||||
Publisher and subscriber must be **separate processes**: a single thread that both
|
||||
published and serviced its own subscription would deadlock (its `publish` call blocks until
|
||||
the service delivers to its endpoint, which only the same thread could receive).
|
||||
|
||||
## Status and follow-ups
|
||||
|
||||
- **The keyboard is real.** The `ps2-bus` driver owns PNP0303, which carries *both* the
|
||||
0x60/0x64 ports and IRQ1, so reading the hardware lives in the bus, not in
|
||||
[keyboard.zig](../../system/drivers/ps2-bus/keyboard.zig): the bus binds IRQ1 and, on each
|
||||
interrupt, drains port 0x60, routing every byte by the status register's
|
||||
auxiliary-output bit to whichever child driver **attached** for that device (an
|
||||
`AttachRequest` to the well-known `ps2_bus` service, carrying the child's endpoint as a
|
||||
capability; the bytes then arrive as asynchronous `ForwardedByte` messages, so the IRQ
|
||||
path never blocks on a child). The keyboard driver decodes the stream — scancode **set 2**,
|
||||
what the keyboard sends with the 8042's legacy translation off, decoded by
|
||||
[scancode.zig](../../system/drivers/ps2-bus/scancode.zig) into USB HID usage keycodes with
|
||||
make/break, typematic-repeat, and modifier tracking (host-tested under `zig build test`) —
|
||||
and publishes real `key_down`/`key_press`/`key_up` events.
|
||||
- **Keycode → character** is wired in: the keyboard driver fills a `key_press` event's
|
||||
`character` through [`library/xkeyboard-config`](../../library/xkeyboard-config/README.md)
|
||||
(`xkb.map(layout, keycode, mods)` → keysym + Unicode character), synthesizing the ASCII
|
||||
control characters for Enter/Tab/Backspace/Escape, whose keysyms map to no Unicode. The
|
||||
layout defaults to `us`; the bus can pass another as the driver's argv[2] — the seam for
|
||||
a future settings source.
|
||||
- **The mouse is real too.** IRQ12 is enumerated on the auxiliary device's own ACPI node
|
||||
(PNP0F13), so the bus claims that node alongside the controller and routes both IRQs to
|
||||
its one endpoint, acking whichever line the notification's badge names.
|
||||
[mouse.zig](../../system/drivers/ps2-bus/mouse.zig) attaches the way the keyboard does and
|
||||
assembles the forwarded bytes with
|
||||
[mouse-packet.zig](../../system/drivers/ps2-bus/mouse-packet.zig) (three-byte stream-mode
|
||||
packets: sync/overflow handling, nine-bit movement, screen-convention `dy` — host-tested
|
||||
under `zig build test`) into `button_down`/`button_up` transitions and `motion` events.
|
||||
**Follow-up:** the IntelliMouse magic-knock for a scroll wheel (four-byte packets) and
|
||||
`scroll` events. The hardware-free `input-source` still rotates through all three classes
|
||||
synthetically (including a joystick, which has no driver yet) via the
|
||||
`input.synthetic*Event` helpers.
|
||||
- **Drop-oldest under overflow** is a defined loss; the 16-slot ring absorbs normal bursts.
|
||||
Real backpressure/flow-control is future work.
|
||||
- **`publish` is unauthenticated** — any process may publish, consistent with the current
|
||||
bring-up trust model (see [driver-model.md](driver-model.md)). A source capability is
|
||||
future work.
|
||||
|
||||
## Verifying it
|
||||
|
||||
The `input` case (`python3 test/qemu_test.py input`, in
|
||||
[tests.zig](../../system/kernel/tests.zig) `inputTest`) boots the real kernel and spawns the
|
||||
service, the synthetic source (which cycles keyboard, mouse, and joystick events), and a
|
||||
subscriber that took all three classes. It passes only when the subscriber heartbeats
|
||||
`input-test: ok` — proof that an event travelled source → service → subscriber over IPC,
|
||||
exercising `ipc_send`, capability-passing subscription, and per-device routing. Each
|
||||
serial line names the class received, so the log shows all three arriving on one stream.
|
||||
|
||||
## See also
|
||||
|
||||
- [ipc.md](ipc.md) — the synchronous rendezvous and the notification path `ipc_send` extends.
|
||||
- [syscall.md](../os-development/syscall.md) — the system-call surface, including `ipc_send`.
|
||||
- [driver-model.md](driver-model.md) — class drivers, capability passing (M13), the trust model.
|
||||
@@ -0,0 +1,556 @@
|
||||
# Native Intel iGPU display support — feasibility and roadmap
|
||||
|
||||
**Status: research snapshot, not implemented.** This records what a *minimal, display-only*
|
||||
native driver for an **Intel integrated GPU** — EDID read + mode-set + framebuffer scanout, with
|
||||
**no** 3D/media/compute — would take, and how it slots into danos's pluggable scanout
|
||||
architecture. It is a survey of primary sources (Intel's open-source
|
||||
[Programmer's Reference Manuals](https://www.intel.com/content/www/us/en/docs/graphics-for-linux/developer-reference/1-0/overview.html),
|
||||
coreboot's [libgfxinit](https://doc.coreboot.org/gfx/libgfxinit.html), the Linux
|
||||
[i915 display](https://github.com/torvalds/linux/tree/master/drivers/gpu/drm/i915/display) driver,
|
||||
and Haiku's [intel_extreme](https://github.com/haiku/haiku/tree/master/src/add-ons/kernel/drivers/graphics/intel_extreme/)),
|
||||
not an implementation. It is the companion to [nvidia-gpus.md](nvidia-gpus.md) and should be read
|
||||
against it — the two answer the same question for opposite silicon.
|
||||
|
||||
Read [display.md](display.md) and [display-v2.md](display-v2.md) first — this doc assumes the v2
|
||||
model where scanout is a **pluggable backend** and a native driver is just another `.scanout`
|
||||
service (like the virtio-gpu one), announcing to the compositor over `attach_scanout`.
|
||||
|
||||
## TL;DR
|
||||
|
||||
- **Intel is a materially easier, lower-tier target than the NVIDIA RTX 3060 — and the reason is
|
||||
documentation, not silicon.** Intel publishes official, register-level, per-platform **Display
|
||||
Engine** PRMs with named registers, bitfields, and numbered enable sequences; NVIDIA publishes
|
||||
no display PRM and forces reverse-engineering against GPL nouveau. A minimal Intel display-only
|
||||
driver is roughly **tier 2 to low-tier 3** for well-covered generations (Skylake / Kaby Lake /
|
||||
Coffee Lake), versus NVIDIA's **tier 4** for GA106. This is the load-bearing conclusion.
|
||||
- **The display block is a genuinely separable register domain.** Mode-set + scanout touch only
|
||||
display registers (pipes, planes, transcoders, DDI buffers, PLLs, power wells, GMBUS/AUX) — **no
|
||||
render engine, no command streamer, no GEM/3D, no signed microcode.** Two small carve-outs, both
|
||||
trivial pokes that do *not* pull in the render engine: a real CDCLK frequency change writes the
|
||||
shared GT PCODE mailbox, and the plane's surface register is a GGTT (memory-interface) address.
|
||||
- **There is no firmware wall on the display path.** The only display microcontroller (DMC / "CSR",
|
||||
Skylake+) is **optional** — its sole job is saving/restoring display state across DC5/DC6
|
||||
low-power idle. Without it, i915 prints "Disabling runtime power management" and mode-sets and
|
||||
scans out normally. GuC/HuC are render/media coprocessors, never touched by a display driver.
|
||||
Pre-Skylake parts have no display microcontroller at all yet mode-set fine. There is **nothing
|
||||
analogous to NVIDIA's GSP**.
|
||||
- **The scanout memory model is dramatically simpler than a discrete GPU.** Intel iGPUs have **no
|
||||
VRAM**: the display scans out of ordinary system RAM addressed through the Global GTT (GGTT), a
|
||||
flat single-level page table. Linear (untiled) framebuffers are first-class. You need **no
|
||||
GEM/TTM, no VMM, no VRAM allocator, no BAR1 aperture juggling** — the exact machinery the NVIDIA
|
||||
path forces on you.
|
||||
- **coreboot libgfxinit is a compact, complete, display-only reference** doing precisely this scope
|
||||
(EDID + PLL/mode-set + scanout, zero 3D) in ~22k lines of formally-analysed SPARK/Ada — versus
|
||||
i915's ~400k lines. It is a *read-and-reimplement* reference, not drop-in code (GPL-2.0-or-later,
|
||||
and Ada, not Zig).
|
||||
- **The clean-room, permissively-licensed path is real** — you can implement from the PRM without
|
||||
reading GPL code, and Haiku's MIT `intel_extreme` is a permissive precedent. This is the decisive
|
||||
contrast with NVIDIA, where no vendor register spec exists.
|
||||
- **The practical catch is hardware, not software.** On a desktop with an RTX 3060, the monitor is
|
||||
almost certainly cabled to the *card*, so an iGPU driver would light a dark motherboard port; the
|
||||
CPU may be an **F-SKU with the iGPU fused off entirely**; and every clean-room reference targets
|
||||
*older* Intel. Intel is the right target to **learn** display bring-up — "run it on my machine"
|
||||
is a separate, machine-dependent question that may not resolve in the reader's favour.
|
||||
- **Recommendation:** as with the NVIDIA doc, GOP already gives native-resolution scanout with zero
|
||||
GPU code. A native Intel driver buys runtime mode changes, hardware vsync, and multihead — and it
|
||||
reaches "first pixel" far faster than the NVIDIA path *if* the target machine actually has a
|
||||
usable, cable-attached iGPU of a documented generation.
|
||||
|
||||
## Display engine architecture, and why it's separable
|
||||
|
||||
For the common single-display path (SST DisplayPort / HDMI / eDP), the Intel display data flow is a
|
||||
small, fully documented, essentially fixed sequence:
|
||||
|
||||
```
|
||||
memory surface → PLANE(s) → PIPE → TRANSCODER → DDI (drives IO/PHY) → connector
|
||||
```
|
||||
|
||||
The Tiger Lake PRM Vol 12 states it verbatim: *"The front end of the display contains the pipes.
|
||||
The pipes connect to the transcoders. The transcoders, except for wireless, connect to the DDIs to
|
||||
drive the IO/PHY."* A **pipe** blends planes (primary/sprite/cursor) into one raster stream; the
|
||||
**transcoder** wraps it in port-protocol timing (DP/HDMI/eDP/DSI); the **DDI** is the physical port
|
||||
and PHY. Pipe, Planes, Transcoder, and Digital Display Interface are each first-class PRM chapters
|
||||
with per-object files in libgfxinit
|
||||
([TGL PRM Vol 12](https://cdrdv2-public.intel.com/705833/intel-gfx-prm-osrc-tgl-vol-12-display-engine.pdf)).
|
||||
|
||||
**Two honest qualifications** the raw research overstated (per verification):
|
||||
|
||||
- The pipeline is *not* strictly linear in all cases — the same PRM pages document optional branches
|
||||
a minimal driver simply ignores (wireless writeback to memory, MIPI DSI, DisplayPort multistream
|
||||
many-to-one, DSC/tiled pipe-joining). Ignoring them does not weaken feasibility.
|
||||
- The four-object model *as named* is **Haswell-onward** (DDI introduced ~2013), not "every gen."
|
||||
Pre-Haswell used FDI + PCH transcoders + port-specific encoders. Within the modern iGPU range
|
||||
danos would realistically target (Skylake → Meteor/Lunar Lake) the model is stable.
|
||||
|
||||
**The DPLL/clock block is a separate, per-port programmable clock source** and is one of the harder,
|
||||
most gen-specific pieces: pick/enable a PLL, route its output to the DDI, then bring up the port.
|
||||
The register layout and divider math change substantially per generation — pre-SKL SPLL/WRPLL/LCPLL,
|
||||
Skylake+ shared DPLL0–3, Gen11+ combo-PHY plus Type-C MG/DKL PLLs. Pixel-clock computation is a
|
||||
classic per-gen rewrite.
|
||||
|
||||
### Separable from render — the single most important enabler
|
||||
|
||||
The display is a distinct register domain from render/media, and this is confirmed at the primary
|
||||
level: the TGL PRM ships display as its own volume (Vol 12), separate from Render Engine (Vol 9) and
|
||||
Media (Vol 11); Linux's KMS "is provided by Intel Display Driver, and **shared with drm/xe**"
|
||||
([kernel.org i915](https://docs.kernel.org/gpu/i915.html)) — i.e. the display module is
|
||||
reused across two different GPU drivers. A full mode-set lights a display end-to-end using only power
|
||||
wells, PLL/port-clock, DDI-buffer/PHY, transcoder and pipe registers — **zero render commands, zero
|
||||
GEM objects, zero command-streamer.** libgfxinit is decisive proof: complete EDID + modeset +
|
||||
framebuffer with no render/3D code at all.
|
||||
|
||||
Two carve-outs the "touches ONLY display registers" phrasing needs (per verification), **neither of
|
||||
which drags in the render engine**:
|
||||
|
||||
1. A mode-set that changes the **Core Display Clock (CDCLK)** frequency/voltage pokes the shared **GT
|
||||
Driver Mailbox** (PCODE/PCU power-controller interface), per Vol 12's own "Display Voltage
|
||||
Frequency Switching" step. A trivial register handshake, documented alongside the display sequence.
|
||||
2. The primary plane's surface register (`PLANE_SURF`) holds a **GGTT graphics address** (a
|
||||
memory-interface concept, not covered in Vol 12). Using pre-mapped stolen memory — as libgfxinit
|
||||
does — sidesteps any active GGTT programming. See [Memory and scanout](#memory-and-scanout).
|
||||
|
||||
### Per-gen churn: what's stable, what you rewrite
|
||||
|
||||
The **object model** (pipes/planes/transcoders/DDIs, GMBUS-for-EDID, double-buffered plane registers
|
||||
armed atomically) is conceptually stable from Ironlake/Haswell through Tiger Lake. What you rewrite
|
||||
per generation is:
|
||||
|
||||
1. the **CPU-vs-PCH split and interconnect**,
|
||||
2. the **port/PHY + DPLL** programming,
|
||||
3. **register offsets + power-well / CDCLK topology**, and
|
||||
4. the **mode-set enable sequence itself** (power-well ordering, PLL lock, DDI-buffer enable,
|
||||
transcoder clock-select) — an effective fourth axis the raw research folded into (1)/(2).
|
||||
|
||||
Interconnect eras, with the timeline **corrected** (the cited Haiku doc was chronologically loose):
|
||||
|
||||
- **Gen5 Ironlake (2010) → Ivy Bridge:** FDI (Flexible Display Interface) links the CPU display
|
||||
engine to PCH-resident ports. The FDI/PCH-split era begins at **Ironlake**, not Gen7.
|
||||
- **Haswell (Gen7.5):** the main digital outputs come **back onto the CPU die as DDIs** (DDI A = eDP)
|
||||
— the *opposite* of "moving output to the PCH," and it collapses the FDI/PCH dance **for the
|
||||
digital ports only**. FDI is **retained** for the legacy VGA/CRT path (DDI E → PCH CRT DAC), so a
|
||||
driver gets the single DDI code path only by omitting analog VGA (which a minimal driver does).
|
||||
- **Skylake (Gen9):** reworks clock/PLL, CDCLK, and the power-well model; introduces the optional DMC.
|
||||
- **Gen11 Ice Lake / Gen12 Tiger Lake:** add combo-PHY + USB-Type-C/Thunderbolt MG/DKL PHYs — the
|
||||
single biggest cost increase, and the reason "newest silicon" is *not* the easiest target. (DSC is
|
||||
documented per-**pipe**; MSO is an eDP feature — not "per-transcoder" as the raw research said.)
|
||||
|
||||
### The tractable sweet spot
|
||||
|
||||
The documented, tractable sweet spot for a from-scratch display-only driver is the
|
||||
**Haswell (Gen7.5) / Broadwell (Gen8) DDI family, with Skylake (Gen9) as the modern-hardware pick**
|
||||
since it shares the same DDI object model. Rationale:
|
||||
|
||||
- Broadwell has a complete, freely downloadable
|
||||
[PRM Vol 11 Display](https://cdrdv2-public.intel.com/690828/intel-gfx-prm-osrc-bdw-vol-11-display.pdf);
|
||||
its engine (3 pipes A/B/C, 4 transcoders incl. transcoder-EDP that floats onto any pipe, DDI A–E,
|
||||
WRPLL/SPLL/LCPLL) is the classic "DDI + transcoder + WRPLL" model.
|
||||
- It predates the combo-PHY / Type-C / MG-DKL complexity of Ice Lake / Tiger Lake.
|
||||
- libgfxinit's DDI **connector/EDID/DP layer is uniform from Haswell through Coffee Lake**, so the
|
||||
hardest-to-get-right port logic generalises widely.
|
||||
|
||||
Two supporting claims from the raw research are **wrong and corrected here (verification):**
|
||||
|
||||
- **The BDW and SKL PRMs are NOT 0BSD-licensed.** Both carry a Creative Commons
|
||||
**Attribution-NoDerivatives** notice. Only the *newer* OSRC PRMs (Tiger Lake 2021 onward) put their
|
||||
embedded code samples under **Zero-Clause BSD**. So for the recommended Haswell/Broadwell/Skylake
|
||||
generations there are no "copy-pasteable 0BSD code samples" — the legal basis is *reimplementation
|
||||
from a CC-BY-ND spec* (register facts are not copyrightable), not copying.
|
||||
- **FDI+PCH is not fully eliminated on Haswell/Broadwell.** The BDW PRM keeps FDI for the DDI E → PCH
|
||||
CRT DAC. The "one DDI code path" holds only for the digital outputs a minimal driver targets.
|
||||
|
||||
Sandy/Ivy Bridge (Gen6/7) is where the hobby-doc walkthroughs concentrate (the OSDev GMBUS/EDID
|
||||
material) but carries the FDI+PCH split cost. *(Low confidence on the OSDev specifics — the wiki
|
||||
returns 403 to automated fetches and its "guaranteed to work" phrasing is a hobby assertion, not a
|
||||
silicon guarantee.)*
|
||||
|
||||
## Documentation — and the clean-room question
|
||||
|
||||
This is the crux of the whole comparison. **Intel hands you the register spec that NVIDIA withholds.**
|
||||
|
||||
- The Tiger Lake **"Vol 12: Display Engine"** PRM is a real, first-party, open-source document —
|
||||
**433 pages, verified by direct download** — with named registers + addresses + bitfield tables
|
||||
(`TRANS_DDI_FUNC_CTL`, `DDI_BUF_CTL`, `DP_TP_CTL`, `PLANE_STRIDE`, `DPLL_CFGCR0/1`, `CDCLK_CTL`,
|
||||
`PWR_WELL_CTL_DDI`, …) and **numbered, step-by-step enable sequences** with explicit writes, wait
|
||||
conditions, and microsecond timeouts. It even includes the "magic value" tables older PRMs deferred
|
||||
to the driver (DisplayPort PLL DCO/divider values; voltage-swing/de-emphasis in mV). *"A spec you
|
||||
could write a driver from directly"* is well-supported, not hyperbole
|
||||
([TGL Vol 12](https://cdrdv2-public.intel.com/705833/intel-gfx-prm-osrc-tgl-vol-12-display-engine.pdf)).
|
||||
- **Clean-room, permissively-licensed implementation is legally and practically feasible from the
|
||||
PRM alone.** CC-BY-ND governs redistribution of the *document*; register addresses and bit
|
||||
definitions are functional facts, and original code implementing a described hardware interface is
|
||||
not a derivative of the PDF. *(This is standard copyright reasoning, not adjudicated case law —
|
||||
treat it as well-grounded, not settled.)* Two independent implementations already exist built
|
||||
essentially from these docs (libgfxinit, Haiku), so the spec is demonstrably sufficient.
|
||||
|
||||
**The documentation ceiling — corrected.** The raw research said public PRMs stop "roughly at Ice
|
||||
Lake / Tiger Lake." Verification refuted this: full public **"Vol 12 Display Engine"** PRMs exist for
|
||||
Ice Lake, Lakefield, Tiger Lake, Rocket Lake, DG1, **and DG2/Arc "Alchemist" (Gen12.5, 2022)** —
|
||||
[the ACM display PRM is public](https://www.x.org/docs/intel/ACM/intel-gfx-prm-osrc-acm-vol12-displayengine.pdf).
|
||||
The genuine cliff is **Meteor Lake (2023) and newer**: those have only a high-level architecture
|
||||
overview, no register-level display PRM, and i915 references their display registers by opaque
|
||||
internal **Bspec numeric IDs**. Alder Lake and Raptor Lake iGPUs are Gen12 Xe-LP display — the same
|
||||
IP as Tiger Lake — so despite lacking a dedicated PRM they are effectively covered by the TGL PRM.
|
||||
|
||||
Net: a from-docs driver can confidently target **Skylake through DG2/Arc**, which is essentially the
|
||||
entire current laptop/NUC installed base; only Meteor Lake and later slide back toward the NVIDIA
|
||||
situation (reverse-engineering or reading GPL i915). The PRMs also survived 01.org's shutdown and are
|
||||
mirrored in several stable places (Intel's cdrdv2 host, the
|
||||
[Igalia CC-BY-ND archive](https://github.com/Igalia/intel-osrc-gfx-prm) for Gen4–Gen9.5,
|
||||
[kiwitree](https://kiwitree.net/~lina/intel-gfx-docs/prm/), x.org) — not a single point of failure.
|
||||
*(Note: the Igalia archive stops at Kaby Lake and contains no Display Engine volume; the TGL/DG2
|
||||
display PRMs are separate Intel/x.org downloads.)*
|
||||
|
||||
## coreboot libgfxinit — the native reference
|
||||
|
||||
[libgfxinit](https://doc.coreboot.org/gfx/libgfxinit.html) is the closest thing to a template danos
|
||||
could ask for: a self-contained **native modeset library** (no VBIOS/int10, no firmware blobs) that
|
||||
probes displays via EDID over DDC/I²C and DP AUX, and drives LVDS, eDP, DP1–3, HDMI1–3, analog VGA,
|
||||
plus USB-C DP/HDMI alt-mode on Tiger Lake. It sets up pipes (Primary/Secondary/Tertiary), planes,
|
||||
transcoders, PLLs, panel power/backlight, the GTT, and framebuffer scanout — **display-only, zero
|
||||
3D/media/compute**, which is exactly danos's scope. Its public entry is essentially
|
||||
`Initialize()` then `Update_Outputs(Pipe_Configs)`, where each `Pipe_Config` carries
|
||||
`{Port, Framebuffer, Cursor, Mode}` — a near-perfect fit for a pluggable scanout backend.
|
||||
|
||||
Why it beats i915 as a reference (**verified by measurement**): **131 Ada source files, ~818 KB,
|
||||
~22k code lines** across *all* generations, factored precisely along the axes you care about (`edid`,
|
||||
`dp_aux`, `dp_training`, `pipe_setup`, `transcoder`, `plls`, `connectors`, `port_detect`), with
|
||||
**none** of the DRM/KMS/GEM/TTM, GT/3D, RC6/RPS, or GuC/HuC machinery that makes
|
||||
`drivers/gpu/drm/i915` **~419k lines / 900 files / 12 MB**. (A grep confirms *zero* gem/ttm/guc/huc/
|
||||
execbuf identifiers in the tree.) It depends only on a small HW-access shim, `libhwbase`
|
||||
(`HW.PCI`, `HW.Port_IO`, `HW.MMIO`, `HW.Time`), which maps naturally onto danos's MMIO-grant + IPC
|
||||
primitives — you provide Zig equivalents and the modeset logic sits on top. *(Correction to the raw
|
||||
research: the widely-quoted "~13–14k LOC" is only the generic `common/` layer; the eight
|
||||
per-generation subdirs roughly double it.)*
|
||||
|
||||
**It is a read-and-reimplement reference, not drop-in code.** Two hard constraints:
|
||||
|
||||
- **License is GPL-2.0-or-later** (the COPYING file is GPLv2; per-file headers add "or any later
|
||||
version"). The CC-BY-4.0 on the docs *site* is a footer, not the source license. Copyleft applies
|
||||
to ported code.
|
||||
- **It is SPARK/Ada, and designed to run as coreboot boot-firmware**, not a runtime OS driver. A
|
||||
danos port means either an Ada/GNAT toolchain in the build or hand-transliteration into Zig; the
|
||||
SPARK "absence of runtime errors" proof does **not** carry over to your reimplementation (and note
|
||||
it proves absence of runtime errors, **not** functional modeset correctness).
|
||||
|
||||
Two more caveats worth knowing: its **error handling is limited** — "only the case that no display
|
||||
could be found counts as failure"; a later DP link-training failure is *not* propagated. And its
|
||||
**verified-in-coreboot** hardware list stops at **Coffee Lake + Apollo Lake**, even though the tree
|
||||
contains a `tigerlake/` directory (Ice Lake has no directory at all, and Alder Lake support is only
|
||||
"begun"). So treat Haswell..Coffee Lake as the trustworthy transliteration window and TGL as
|
||||
present-but-less-proven.
|
||||
|
||||
The orchestration reads as a clean state machine (`hw-gfx-gma.adb` `Enable_Output`):
|
||||
`Fill_Port_Config → Preferred_Link_Setting → PLLs.Alloc → [retry] Connectors.Pre_On →
|
||||
Display_Controller.On → Connectors.Post_On`, with a literal *"try each DP-lane configuration twice"*
|
||||
inner retry and an outer link-setting step-down. `hw-gfx-dp_training.adb` (398 lines) is a complete,
|
||||
generic DP link-training implementation (TP1/TP2/TP3, CR + EQ loops, swing/pre-emphasis adjust from
|
||||
sink status). Per-generation buffer translations plug in underneath via
|
||||
`Program_Buffer_Translations`, gated on `Config.Has_DDI_Buffer_Trans`. All of this was confirmed
|
||||
against the source line-by-line.
|
||||
|
||||
## The EDID + mode-set path (Haswell/Broadwell target)
|
||||
|
||||
The whole path is memory-mapped register programming with polled status bits — no command ring, no
|
||||
microcode, no DMA channel.
|
||||
|
||||
**EDID over DDC (GMBUS).** Pure MMIO poking of the GMBUS I²C controller (`GMBUS0`–`GMBUS5`): `GMBUS0`
|
||||
selects pin-pair/port + clock; `GMBUS1` carries slave address (`0x50` for EDID), byte count,
|
||||
direction, SW-ready; `GMBUS2` exposes HW-ready/NAK/ACTIVE to poll; `GMBUS3` is a 4-byte data FIFO;
|
||||
`GMBUS5` gives the 2-byte segment index for E-DDC. A read is: write `GMBUS0`, write `GMBUS1`
|
||||
(`CYCLE_WAIT | count | SLAVE_READ | SW_RDY | slave<<addr`), loop {poll `HW_RDY`, read 4 bytes}, then
|
||||
STOP ([i915 intel_gmbus.c](https://github.com/torvalds/linux/blob/master/drivers/gpu/drm/i915/display/intel_gmbus.c)).
|
||||
|
||||
**EDID + DPCD over DP AUX.** For DisplayPort/eDP, EDID (as I²C-over-AUX to `0x50`) and all DPCD
|
||||
capability/link-status registers are read over the AUX channel: per-DDI `DDI_AUX_CTL` + 5×
|
||||
`DDI_AUX_DATA`. Build a 3–5 byte header + payload, set SEND_BUSY, poll it clear, read
|
||||
DONE/TIMEOUT/RECEIVE_ERROR. Message size 1–20 bytes; spec requires ≥3 retries. On Haswell/BDW the AUX
|
||||
clock divider is programmed explicitly; SKL+ derive it automatically
|
||||
([i915 intel_dp_aux.c](https://github.com/torvalds/linux/blob/master/drivers/gpu/drm/i915/display/intel_dp_aux.c)).
|
||||
Both GMBUS and DP-AUX live in libgfxinit's shared `common/` — cheap and nearly gen-invariant.
|
||||
|
||||
**The mode-set is a fixed, documented register sequence.** The Broadwell DisplayPort enable order
|
||||
(verbatim from BDW PRM Vol 11, pp.98–99): (1) DDI lane capability; (2) panel power sequencing if
|
||||
needed; (3) enable the CPU display PLL (WRPLL/SPLL) and wait ~20 µs; (4) Port Clock Select → DDI,
|
||||
enable `DP_TP_CTL` with training pattern 1, configure `DDI_BUF_TRANS`, enable `DDI_BUF_CTL`, wait
|
||||
>518 µs, run link training, set `DP_TP_CTL` to Normal (Idle first for eDP); (5) Transcoder Clock
|
||||
Select, enable the plane, panel fitter if needed, program transcoder timings + M/N/TU, enable
|
||||
`TRANS_DDI_FUNC_CTL`, enable `TRANS_CONF`, then backlight. Disable is the exact reverse — a bounded
|
||||
checklist.
|
||||
|
||||
**DisplayPort/eDP link training is driver-driven in software over AUX** — the CPU runs the
|
||||
clock-recovery and channel-equalization state machines by hand; it is **not** offloaded to a hardware
|
||||
sequencer or firmware. The source side exposes only primitives: `DP_TP_CTL` selects the training
|
||||
pattern the port emits; `DDI_BUF_CTL`/`DDI_BUF_TRANS` set voltage-swing/pre-emphasis. The driver
|
||||
loops: emit pattern + set source levels → write `TRAINING_PATTERN_SET` (DPCD 0x102) + `TRAINING_LANEx_SET`
|
||||
(0x103) over AUX → delay (100 µs CR / 400 µs EQ) → read `LANE_STATUS` → on failure adjust to the
|
||||
sink's `ADJUST_REQUEST` values and retry. A few hundred lines of ordinary CPU/AUX code (libgfxinit
|
||||
`Train_DP`: CR loop 1..32, EQ loop 1..6). **This is the single fiddliest, most fragile piece** — a
|
||||
TMDS/HDMI panel avoids it entirely, and targeting an already-lit eDP panel avoids most of it.
|
||||
|
||||
**The clock (WRPLL) is documented divider math, not a magic table.** On Haswell/BDW the WRPLL derives
|
||||
the symbol clock from a 2700 MHz LCPLL reference through R2/N2/P dividers with VCO 2400–4800 MHz —
|
||||
small integer arithmetic. DP is *easier* than HDMI because it runs at a few fixed link rates (1.62 /
|
||||
2.7 / 5.4 GHz), so a DP/eDP-only minimal driver can often use fixed rates and skip most of the search.
|
||||
|
||||
**Plane/scanout programming is trivial for a compositor.** The primary plane is `PRI_CTL`
|
||||
(enable + pixel format), `PRI_STRIDE`, `PRI_SURF` (surface base — writing it triggers the atomic
|
||||
update), `PRI_OFFSET`; formats include 32-bit BGRX 8:8:8 and 16-bit BGRX 5:6:5 — a direct match for a
|
||||
linear XRGB compositor buffer. Plane registers are double-buffered and latch at vblank via an
|
||||
**arming** write — so a page-flip is "write base + stride + size, then the arming write." This is
|
||||
*exactly* the primitive danos's damage-driven compositor already expresses over GOP/virtio-gpu; the
|
||||
incremental work is "program these display-domain registers," not a new scanout model. The panel
|
||||
fitter (`PF_WIN_POS`/`PF_WIN_SZ`/`PF_CTRL`) can be left disabled for native-resolution scanout;
|
||||
Skylake+ replaces it with a shared pipe-scaler (`PS_CTRL`).
|
||||
|
||||
**Smallest useful target:** eDP (DDI A / transcoder-EDP) or a single DP output at native resolution,
|
||||
panel fitter off, plane in 32bpp XRGB. That is: GMBUS + I²C-over-AUX EDID/DPCD, one fixed-rate or
|
||||
WRPLL config, the ~20-step enable sequence, the software CR/EQ loop, and `PRI_*` plane setup with
|
||||
`PRI_SURF`-write flips. Out of scope: 3D, media, tiling, RC6/power-gating, PSR, audio.
|
||||
|
||||
## Memory and scanout
|
||||
|
||||
This is where Intel's *architecture* — not just its docs — makes the job smaller, and it is the
|
||||
biggest single simplification versus a discrete GPU.
|
||||
|
||||
- **No VRAM.** Intel iGPUs have a unified memory architecture; the display scans out of ordinary
|
||||
**system RAM** addressed through the **Global GTT (GGTT)**. The only way to give the GPU memory is
|
||||
to bind system pages into the GGTT
|
||||
([i915/GEM crashcourse](https://blog.ffwll.ch/2012/10/i915gem-crashcourse.html)).
|
||||
- **The plane surface register is a GGTT offset**, not a raw physical address — the display walks the
|
||||
GGTT to fetch pixels, so a scanout buffer must be GGTT-mapped (global, not per-process). libgfxinit
|
||||
writes the framebuffer offset straight into `DSPSURF`/`PLANE_SURF` masked to 4 KB.
|
||||
- **Linear (untiled) scanout is a first-class supported mode** — the plane's tiling field value 0 is
|
||||
Linear. No X/Y/Yf tiling engine is needed for a display-only driver. (UEFI GOP itself hands off a
|
||||
linear framebuffer the plane is already scanning.)
|
||||
- **No memory manager.** You need only (1) some contiguous-ish system pages and (2) GGTT PTEs
|
||||
pointing at them (`physical_addr | valid_bit` — the GGTT is a flat single-level array of PTEs in
|
||||
the `GTTMMADR` MMIO BAR), then program the plane. **No GEM/TTM/PPGTT/GuC.** coreboot's native-init
|
||||
literally does `for(i…) WRITE32(base + i*inc | 1, (i*4) | 1)`.
|
||||
- **"Stolen memory"** (GSM/DSM) is firmware-reserved system RAM where the firmware places the GGTT
|
||||
itself and the boot framebuffer. A driver is not obligated to keep scanout there — it can rebind
|
||||
GGTT entries to its own pages. Stolen memory matters mainly for *inheriting* the GOP framebuffer at
|
||||
handoff.
|
||||
|
||||
**The contrast with NVIDIA is stark.** On a discrete GPU the scanout surface must live in **VRAM**
|
||||
(nouveau always pins scanout to VRAM), CPU access goes through the **BAR1** aperture (which on
|
||||
consumer cards can be far smaller than total VRAM unless Resizable BAR is on), and you need a
|
||||
contiguous aligned VRAM allocator plus a BAR1 mapping. The Intel iGPU path **eliminates all of that**
|
||||
— scanout is plain system RAM, and a userspace compositor can write the framebuffer pages directly
|
||||
(as danos already does with the GOP WC framebuffer).
|
||||
|
||||
Because danos boots via GOP, an Intel driver attaches to a display whose **GGTT is already populated
|
||||
and whose plane is already scanning a linear framebuffer at native resolution.** A minimal driver can
|
||||
reuse that live mapping and reprogram the running plane rather than come up from cold — the same
|
||||
"attach to a live display" advantage the NVIDIA doc identifies, but with a far smaller register
|
||||
surface and no firmware wall. *(Low-confidence, per-target details to pin from the specific gen's
|
||||
PRM: GGTT PTE size — 4-byte pre-gen8 vs 8-byte gen8+ — the `GTTMMADR`/aperture BAR layout, surface
|
||||
alignment — 4 KB floor but some gens/tilings want 256 KB — and whether the display's GGTT-mediated
|
||||
DMA sits before or after danos's M16 IOMMU on the target platform.)*
|
||||
|
||||
## Firmware
|
||||
|
||||
A minimal display-only Intel driver is **effectively firmware-free — more so than NVIDIA.**
|
||||
|
||||
- **DMC (Display Microcontroller, "CSR", Skylake+) is NOT required for mode-set or scanout.** Its
|
||||
sole job is saving/restoring display-engine registers across DC5/DC6 low-power idle. Absent, i915
|
||||
prints *"Failed to load DMC firmware … Disabling runtime power management"* and the display
|
||||
mode-sets and scans out normally — you lose only the deep display idle states, not output
|
||||
([intel_dmc.c](https://github.com/torvalds/linux/blob/master/drivers/gpu/drm/i915/display/intel_dmc.c);
|
||||
corroborated by multiple distro bug threads). *(A source-level `HAS_DMC` early-return citation would
|
||||
strengthen this beyond distro testimony, but the conclusion is well-supported.)*
|
||||
- **Pre-Skylake parts have no display microcontroller at all** yet perform full mode-set (and even
|
||||
Panel Self Refresh). This confirms the display engine is fundamentally CPU/MMIO-driven; the
|
||||
microcontroller is an add-on for autonomous idling, not a prerequisite for lighting a panel.
|
||||
Targeting a pre-Skylake or DMC-optional generation sidesteps the question entirely.
|
||||
- **GuC and HuC are render/media microcontrollers on the GT side** — GuC schedules the render engines,
|
||||
HuC assists HEVC/H.265 codec (plus later HDCP/PXP/GSC). Neither is in the scanout path; a
|
||||
display-only driver never loads them
|
||||
([kernel.org microcontrollers](https://docs.kernel.org/gpu/i915.html)).
|
||||
- **PSR firmware lives on the panel**, not in the OS — a minimal driver simply doesn't enable PSR.
|
||||
- **Type-C/TCSS (Ice Lake+) firmware** (PMC/IOM/PHY) is part of platform BIOS/coreboot init and the
|
||||
hardware, *not* a signed blob the display driver loads at runtime. A driver attaching to an
|
||||
already-lit GOP connector, or targeting classic DDI ports, avoids it. *(Cold DP-alt-mode changes
|
||||
from a userspace driver on modern TCSS platforms were not traced to primary source — flagged.)*
|
||||
|
||||
There is **no signed-firmware wall over the Intel GPU at all** on the display path. This is the
|
||||
architectural opposite of NVIDIA's mandatory, unsignable, ABI-unstable GSP — which even on the
|
||||
near-side "direct" display path is a permanent maintenance liability for anything beyond scanout.
|
||||
|
||||
## Licensing
|
||||
|
||||
The situation is *better* than NVIDIA's but still nuanced.
|
||||
|
||||
- **The two best code references are both GPL** — Linux i915 (GPL-2.0) and coreboot libgfxinit
|
||||
(GPL-2.0-or-later). You cannot copy either into a permissively-licensed danos. libgfxinit's WRPLL
|
||||
divider math is itself copied from i915, so it carries the same encumbrance.
|
||||
- **But you don't need to copy code.** The Intel PRM is a *specification*, and a clean-room Zig
|
||||
implementation written from the PRM (using libgfxinit/i915 only to understand behaviour, never to
|
||||
copy) is legitimate — register numbers and bit definitions are functional facts, not copyrightable
|
||||
expression. This is the exact inverse of the NVIDIA case, where no such spec exists and the only
|
||||
guide is the GPL/RE'd code itself.
|
||||
- **A permissive precedent exists: Haiku's `intel_extreme` is MIT-licensed** and was built from
|
||||
Intel's public docs. So if danos wants a permissive license, the model is: implement from the PRM,
|
||||
optionally read MIT Haiku for structure, treat GPL libgfxinit/i915 as documentation-of-last-resort.
|
||||
- **A licensing nuance on the recommended generations:** the "copy the 0BSD PRM code samples" shortcut
|
||||
only applies to Tiger-Lake-era (2021+) PRMs. The Haswell/Broadwell/Skylake PRMs are CC-BY-ND, so
|
||||
their register *facts* are free to implement but there are no code samples to lift.
|
||||
|
||||
As with the NVIDIA doc: danos's userspace-driver-over-IPC model (a driver is a separate process behind
|
||||
a defined protocol) is the cleanest possible license boundary if the project ever chooses to ship a
|
||||
GPL display-driver binary and keep the rest of danos permissive — but that is a boundary judgement
|
||||
wanting real diligence, not a settled fact. The clean-room-from-PRM route avoids the question.
|
||||
|
||||
## Prior art outside Linux
|
||||
|
||||
This is a **real contrast with NVIDIA**, where no one has built a from-scratch native driver outside
|
||||
Linux. For Intel there are **multiple independent, non-Linux, clean-room native modeset
|
||||
implementations** to learn from:
|
||||
|
||||
- **coreboot libgfxinit** — SPARK/Ada, G45/GM45 and Arrandale → Coffee Lake + Apollo Lake (TGL
|
||||
in-tree), the strongest structural reference.
|
||||
- **Haiku `intel_extreme`** — modeset-only (no 2D/3D accel), **MIT-licensed**, i845 through Sandy
|
||||
Bridge solid, newer Gemini/Ice/Tiger Lake in progress but "hit or miss, as the driver lags behind
|
||||
the specs" ([Haiku generations](https://www.haiku-os.org/docs/develop/drivers/intel_extreme/generations.html),
|
||||
[Phoronix Sept 2024](https://www.phoronix.com/news/Haiku-OS-September-2024)).
|
||||
- **SerenityOS** — added basic native Intel graphics ([PR #6277](https://github.com/SerenityOS/serenity/pull/6277)),
|
||||
though only for very old ICH7-class hardware.
|
||||
- **managarm** — native Intel G45 support.
|
||||
|
||||
The catch: **every clean-room non-Linux implementation targets old hardware.** A modern Gen12 "Xe"
|
||||
desktop iGPU is beyond all of them; for the very newest parts only GPL i915 covers the registers. So
|
||||
the wealth of prior art is real but concentrated below Tiger Lake.
|
||||
|
||||
## The practical desktop caveat
|
||||
|
||||
Before any effort estimate is trusted, three hardware realities — the honest reason "Intel is easier"
|
||||
does **not** automatically mean "it'll light up the reader's monitor":
|
||||
|
||||
1. **Muxing / cabling.** On a desktop with a discrete RTX 3060, the monitor is almost certainly
|
||||
plugged into the *card's* outputs, not the motherboard's. An iGPU driver would light a
|
||||
**different, currently-dark** output. To see danos on Intel the reader would have to physically
|
||||
move the cable to a motherboard video port **and** likely enable the iGPU / "IGD Multi-Monitor" in
|
||||
BIOS. Intel-first probably does **not** light the current display without re-cabling.
|
||||
2. **No iGPU at all.** Intel **F-SKU** desktop chips (i5-9400F, i5-12400F, i5-13400F, i7-13700KF, …)
|
||||
ship the graphics **fused off** and cannot be re-enabled. These are extremely common in
|
||||
budget/mid gaming builds paired with an RTX 3060. On an F-SKU (or an X-series HEDT part) the
|
||||
Intel-iGPU path is a **non-starter** regardless of cabling.
|
||||
3. **Generation coverage.** If the CPU *is* a recent non-F part, its iGPU may be Gen12 Xe (Alder/
|
||||
Raptor Lake), beyond libgfxinit's verified set and beyond most non-Linux prior art — leaving GPL
|
||||
i915 (or the TGL-class PRM, which covers Alder/Raptor display IP) as the only reference.
|
||||
|
||||
A cleaner path for *learning* without the hardware lottery: an older bare-metal Intel box (Haswell/
|
||||
Skylake NUC or laptop) whose panel is natively on the iGPU. Note QEMU does **not** emulate an Intel
|
||||
iGPU display engine, so a VM cannot exercise a real Intel modeset path — virtio-gpu (already working)
|
||||
is the VM answer.
|
||||
|
||||
## Alternatives, and the honest Intel-vs-NVIDIA verdict
|
||||
|
||||
| Option | What you get | The tradeoff |
|
||||
|---|---|---|
|
||||
| **Stay on GOP** (working today) | Native-res scanout, zero GPU code/firmware/maintenance | Resolution frozen at ExitBootServices; no runtime mode change, no hardware vsync, no multihead |
|
||||
| **Intel iGPU, reuse-GOP** | EDID read + plane page-flips on the GOP-set mode | Still bounded to GOP's resolution; but real driver-owned scanout |
|
||||
| **Intel iGPU, full modeset** (this doc) | Runtime modeset, vsync, multihead, from public docs | Tier 2–3 effort; DP link training; per-gen churn; **needs a cable-attached, documented iGPU** |
|
||||
| **Native NVIDIA GA106 direct** ([nvidia-gpus.md](nvidia-gpus.md)) | Same, on the RTX 3060 the monitor is actually plugged into | **Tier 4**; GPL-only reference; DMA channel modeset; de-emphasised legacy path |
|
||||
| **GA106 via GSP/OGKM** | Also unlocks 3D later | Tier 5; unstable version-pinned firmware ABI |
|
||||
|
||||
**The verdict for *this reader* (RTX 3060 box):** For pure "see danos on my screen," **NVIDIA-direct
|
||||
is paradoxically the more relevant path**, because the monitor is already cabled to the 3060 and GOP
|
||||
already drives it — a native NVIDIA driver reprograms *that* live display. An Intel driver, however
|
||||
much easier to *write*, likely lights a dark motherboard port the reader isn't looking at, or hits an
|
||||
F-SKU with no iGPU.
|
||||
|
||||
**The verdict for *learning display bring-up*:** **Intel wins decisively.** Public register PRMs, four
|
||||
independent open reference drivers, an MIT precedent (Haiku), a compact formally-analysed blueprint
|
||||
(libgfxinit), no signed-firmware wall, no VRAM/BAR memory manager, and a legitimate permissive
|
||||
clean-room path. It reaches "first pixel" far faster than the NVIDIA native path — *on hardware that
|
||||
actually has a cable-attached, documented Intel iGPU.* Those two goals — "run on my machine" and
|
||||
"learn the craft" — point at different silicon, and that is the honest bottom line.
|
||||
|
||||
## "First light" milestones — a danos `.scanout` service
|
||||
|
||||
Framed as a danos `.scanout` service (like the virtio-gpu and proposed NVIDIA ones), inheriting the
|
||||
GOP-initialized display — no firmware, no cold POST:
|
||||
|
||||
1. **PCI/BAR bring-up** — enumerate the iGPU, map its MMIO BAR (`GTTMMADR` + register block) and the
|
||||
aperture BAR via danos MMIO grants; confirm the display engine is GOP-live.
|
||||
2. **EDID** — implement GMBUS DDC (`0x50`) and DP AUX; read + parse the panel EDID and DPCD caps.
|
||||
*(Smallest self-contained, gen-invariant milestone — a good first commit.)*
|
||||
3. **First pixel = reprogram, don't re-modeset** — with GOP's mode and GGTT mapping inherited,
|
||||
reprogram the running plane (`PRI_CTL`/`PRI_STRIDE`/`PRI_SURF`, linear, 32bpp XRGB) to point at a
|
||||
danos-owned system-RAM buffer; prove a page-flip via the `PRI_SURF` arming write on the *current*
|
||||
mode before changing timings. This defers the entire DPLL/DDI/transcoder/link-training surface —
|
||||
the hardest, most gen-specific ~70% of the work.
|
||||
4. **GGTT ownership** — write your own GGTT PTEs (via an MMIO grant to `GTTMMADR`) pointing at
|
||||
compositor-owned pages, for double-buffered damage-driven present.
|
||||
5. **Wire into the compositor `.scanout` backend** (`attach_scanout`); add vsync via the display
|
||||
vblank interrupt (IRQ-as-IPC).
|
||||
6. **Full mode-set** (the hard, gen-specific step) — for one chosen generation (Haswell/Broadwell or
|
||||
Skylake): WRPLL/DPLL programming, the ~20-step DDI/transcoder/pipe enable sequence, panel power
|
||||
sequencing for eDP (`PP_CONTROL`/`PP_ON_DELAYS`/`PP_OFF_DELAYS` — a common black-screen pitfall).
|
||||
7. **DisplayPort link training** — only if the panel is DP and GOP's link can't be reused; the
|
||||
software CR/EQ state machine over AUX. TMDS/HDMI avoids it; a live eDP panel avoids most of it.
|
||||
8. **Multihead**, then optionally a second generation once one is solid.
|
||||
|
||||
Keep the GOP backend as the fallback the whole way — a stall at any step still leaves danos with a
|
||||
working display, exactly the resilience v2 already provides via re-attach.
|
||||
|
||||
## Reading list
|
||||
|
||||
**Native reference — coreboot libgfxinit (GPL-2.0-or-later, SPARK/Ada):**
|
||||
- `common/hw-gfx-gma.adb` — `Enable_Output`, the end-to-end modeset state machine.
|
||||
- `common/hw-gfx-dp_training.adb` — the complete generic DP link-training CR/EQ loops.
|
||||
- `common/hw-gfx-gma-pipe_setup.adb` — plane/pipe/scaler + `DSPSURF`/`DSPSTRIDE`/`DSPCNTR` scanout.
|
||||
- `common/hw-gfx-gma-transcoder.adb` — timing generator; `common/hw-gfx-edid.adb`,
|
||||
`hw-gfx-gma-i2c.adb`, `hw-gfx-dp_aux_ch.adb` — EDID/DDC/AUX; `hw-gfx-gma-registers.ads` — offsets.
|
||||
- `common/haswell*/`, `skylake/`, `tigerlake/` — the per-gen PLL/PHY/buffer-translation backends.
|
||||
|
||||
**Vendor register specs — Intel OSRC PRMs:**
|
||||
- [Broadwell Vol 11: Display](https://cdrdv2-public.intel.com/690828/intel-gfx-prm-osrc-bdw-vol-11-display.pdf)
|
||||
(CC-BY-ND) — the recommended Haswell/Broadwell-class enable sequences, plane, panel fitter.
|
||||
- [Tiger Lake Vol 12: Display Engine](https://cdrdv2-public.intel.com/705833/intel-gfx-prm-osrc-tgl-vol-12-display-engine.pdf)
|
||||
(code samples 0BSD) — the most complete modern reference incl. PLL/voltage-swing value tables.
|
||||
- [DG2/Arc Vol 12: Display Engine](https://www.x.org/docs/intel/ACM/intel-gfx-prm-osrc-acm-vol12-displayengine.pdf)
|
||||
— the newest public display PRM (Gen12.5, 2022).
|
||||
- [Igalia CC-BY-ND archive](https://github.com/Igalia/intel-osrc-gfx-prm) (Gen4–Gen9.5) and the
|
||||
[kiwitree mirror](https://kiwitree.net/~lina/intel-gfx-docs/prm/) — stable mirrors.
|
||||
|
||||
**GPL reference-of-last-resort — Linux i915 display:**
|
||||
- `intel_gmbus.c`, `intel_dp_aux.c` — the concrete EDID/DDC and DP-AUX register sequences.
|
||||
- `intel_ddi.c` / `intel_ddi_buf_trans.c`, `intel_cdclk.c`, `intel_dpll_mgr.c` — DDI/CDCLK/PLL;
|
||||
`i9xx_plane.c`, `intel_crtc.c` — plane/pipe; `intel_dp.c` — link training. Huge and modular; a
|
||||
reference to confirm undocumented quirks, not a template.
|
||||
|
||||
**Permissive prior art — Haiku `intel_extreme` (MIT):**
|
||||
- [`src/add-ons/kernel/drivers/graphics/intel_extreme/`](https://github.com/haiku/haiku/tree/master/src/add-ons/kernel/drivers/graphics/intel_extreme/)
|
||||
— a second independent modeset-only driver; MIT, so structurally readable for a permissive danos.
|
||||
- [generations.html](https://www.haiku-os.org/docs/develop/drivers/intel_extreme/generations.html)
|
||||
— the best plain-English per-generation fault-line map.
|
||||
|
||||
## Open questions (unresolved by the survey)
|
||||
|
||||
- **Does the target machine have a usable, cable-attached iGPU at all?** F-SKU check, CPU generation,
|
||||
and monitor cabling must be resolved before any effort estimate is trusted (see
|
||||
[practical caveat](#the-practical-desktop-caveat)).
|
||||
- **Does danos even need native mode-*setting*, or only plane/scanout control on the GOP-set mode?**
|
||||
If runtime mode changes aren't required, the driver collapses to EDID + plane page-flips, dropping
|
||||
the DPLL/DDI/link-training ~70% of the work.
|
||||
- **GGTT vs raw physical:** confirm from the exact target-gen PRM that `PLANE_SURF` is interpreted as
|
||||
a GGTT graphics address (well-established, but per-gen confirmation advisable), and the PTE size /
|
||||
`GTTMMADR` / aperture layout for writing GGTT entries.
|
||||
- **Reuse the firmware/GOP GGTT + framebuffer, or install your own GGTT entries?** The latter (needed
|
||||
for double-buffering) means writing GGTT PTEs from the userspace driver via an MMIO grant.
|
||||
- **eDP panel power sequencing** (`PP_*`, T1–T12 delays) — not covered in this pass and a common
|
||||
black-screen source.
|
||||
- **IOMMU interaction** — whether the display's GGTT-mediated DMA needs IOMMU passthrough for the
|
||||
framebuffer pages under danos's M16 IOMMU, or sits before the IOMMU on the target platform.
|
||||
- **DP link-training / AUX robustness and per-generation register drift** are the dominant *risks* —
|
||||
not documentation scarcity.
|
||||
- **Exact Haswell/BDW MMIO offsets** (commonly cited: GMBUS ~`0xC5100`, `DDI_AUX_CTL_A` ~`0x64010`,
|
||||
`DDI_BUF_CTL_A` ~`0x64000`, `DP_TP_CTL_A` ~`0x64040`) were not extracted verbatim from the PRM —
|
||||
confirm against `i915_reg.h` before coding.
|
||||
|
||||
---
|
||||
|
||||
*Research snapshot; verify against current libgfxinit / i915 source and the specific target
|
||||
generation's PRM before building. Intel's public-PRM coverage and the muxing/F-SKU realities of a
|
||||
given machine both change what is actually achievable.*
|
||||
@@ -0,0 +1,132 @@
|
||||
# IPC: message-passing channels
|
||||
|
||||
Inter-process communication is the **backbone of a microkernel**. Once drivers and
|
||||
services run isolated in their own address spaces ([vision](../vision.md)), they can't
|
||||
just call each other — a request becomes a **message**. In a microkernel, whatever
|
||||
was a function call across a monolithic kernel is IPC, so it's a first-class
|
||||
concern, not an afterthought.
|
||||
|
||||
There are two layers, built a milestone apart:
|
||||
|
||||
- **`system/kernel/ipc.zig`** — a bounded blocking channel between *kernel threads*,
|
||||
described below. The primitive, and where the blocking discipline was worked out.
|
||||
- **`system/kernel/ipc-synchronous.zig`** — synchronous call/reply between *processes*, across
|
||||
address spaces. What user-space servers and drivers actually talk over. It's the
|
||||
second half of this document.
|
||||
|
||||
## The channel
|
||||
|
||||
The first form is a **bounded blocking channel** (`system/kernel/ipc.zig`): a fixed-size
|
||||
ring buffer of messages with a producer/consumer rendezvous, built on the
|
||||
scheduler's [wait queues](../os-development/scheduling.md).
|
||||
|
||||
`Channel(T, capacity)` is generic over the message type and buffer size. It holds a
|
||||
ring buffer, a count, and two wait queues:
|
||||
|
||||
- **`send(msg)`** — if the channel is full, block on the *not-full* queue; otherwise
|
||||
write the message, bump the count, and wake a waiting receiver.
|
||||
- **`receive()`** — if the channel is empty, block on the *not-empty* queue; otherwise
|
||||
take a message, drop the count, and wake a waiting sender.
|
||||
|
||||
Neither side busy-waits: a full channel parks the sender, an empty one parks the
|
||||
receiver, and each operation wakes the other side when it makes progress possible.
|
||||
|
||||
Two details make it correct:
|
||||
|
||||
- **Recheck in a loop.** A woken task re-tests the condition (`while (full) wait`)
|
||||
rather than assuming the slot is still available — another waiter may have taken
|
||||
it first. This is the standard guard against spurious or racing wakeups.
|
||||
- **One critical section.** `send`/`receive` run under the [big kernel
|
||||
lock](../os-development/smp.md) (`sync.enter` / `sync.leave`), which disables interrupts on this
|
||||
core *and* takes the kernel's one spinlock — since SMP, the interrupt flag alone
|
||||
is not atomicity, because `cli` on one core does nothing to another. So checking
|
||||
the condition and committing the block/enqueue happen atomically both with respect
|
||||
to the timer preempting mid-operation and to the other side running on another
|
||||
CPU. `waitLocked` / `wakeLocked` are the variants that assume the caller already
|
||||
holds that critical section.
|
||||
|
||||
## Verifying it
|
||||
|
||||
The `ipc` test (see [testing.md](../testing.md)) runs a producer and a consumer passing
|
||||
**100 messages through a 4-slot channel**. The small buffer means the channel goes
|
||||
full and empty over and over, so both the blocking-send and blocking-receive paths are
|
||||
exercised heavily. The messages arrive intact and in order (their sum is the
|
||||
expected `5050`), and neither task busy-waits — they block and wake each other.
|
||||
|
||||
## Endpoints: call/reply across address spaces
|
||||
|
||||
A channel connects two kernel threads sharing one address space. Real servers are
|
||||
*processes*, so the payload has to cross an address-space boundary. That's
|
||||
`system/kernel/ipc-synchronous.zig`, and its shape is L4's: a synchronous **rendezvous** at an
|
||||
`Endpoint`, with the message copied directly from the sender's pages to the receiver's
|
||||
(`copyAcross` walks both sets of page tables through the physmap — no CR3 switch, no
|
||||
bounce buffer).
|
||||
|
||||
Two syscalls carry it:
|
||||
|
||||
- **`ipc_call(h, msg, reply)`** — copy `msg` to the server, block until it replies.
|
||||
- **`ipc_reply_wait(h, reply, recv)`** — reply to the client you're still holding (if
|
||||
any), then block for the next request. One syscall, because a server's steady state
|
||||
is *always* "finish the last one, wait for the next".
|
||||
|
||||
An endpoint is reached by **handle** — a small integer index into the process's handle
|
||||
table (`Task.handles`), exactly like a file descriptor, and just as unforgeable. The
|
||||
bootstrap problem (how do you get the first handle?) is solved by a tiny name registry:
|
||||
a server calls `ipc_register(service_id, h)` under a well-known small integer, and a
|
||||
client calls `ipc_lookup(service_id)`.
|
||||
|
||||
The server never learns the client's identity beyond a **badge**, delivered alongside
|
||||
the message: the caller's task id.
|
||||
|
||||
### Interrupts are messages too
|
||||
|
||||
`notifyFromIsr` posts an *asynchronous* notification to an endpoint — no payload, no
|
||||
reply owed — and wakes whoever is blocked in `reply_wait`. Its badge has the top bit
|
||||
set (`notify_badge_bit`), which is how a driver's single event loop distinguishes "a
|
||||
client wants something" from "the hardware wants something". Notifications sit in a
|
||||
small coalescing ring on the endpoint, so an interrupt taken while the driver was busy
|
||||
elsewhere is not lost.
|
||||
|
||||
This is what makes a user-space driver possible at all, and it's the subject of
|
||||
[drivers.md](drivers.md).
|
||||
|
||||
## What's next (partly done since)
|
||||
|
||||
- **Priority inheritance** through IPC — still open: a high-priority client
|
||||
blocked on a low-priority server suffers unbounded priority inversion.
|
||||
- **Handle transfer.** *Landed as cap-passing (M13)*: `ipc_call` and
|
||||
`ipc_reply_wait` carry an optional capability alongside the bytes (`send_cap`),
|
||||
copying an endpoint or shared-memory handle into the peer's table. First user:
|
||||
[input](input.md) subscribers register by handing over their own endpoint, and
|
||||
class drivers get a private channel to one device.
|
||||
- **Asynchronous / buffered send** for the cases where a rendezvous is the wrong
|
||||
shape (logging, notifications between servers). *Landed as `ipc_send`* — a
|
||||
non-blocking post to an endpoint's bounded payload queue, delivered through
|
||||
`reply_wait` as a buffered message (badge bit `notify_message_bit`). Built for, and
|
||||
first used by, the [input service](input.md)'s keyboard-event broadcast, where a
|
||||
synchronous push would let one dead subscriber hang the fan-out. A full queue drops
|
||||
the oldest (discrete messages, not a coalescing level like the notification ring).
|
||||
- **A bounded reply** — half landed. The copy is still 256 bytes
|
||||
(`MESSAGE_MAXIMUM`) under the big kernel lock, but bulk transfer got its shared
|
||||
pages: `shared_memory_create`/`map`/`physical`, the region handle delegated as
|
||||
a capability (above). virtio-gpu's scanout surface is the first user
|
||||
([display-v2.md](display-v2.md)).
|
||||
|
||||
## Lifecycle conventions over IPC (M17)
|
||||
|
||||
Three conventions from [process-lifecycle.md](../os-development/process-lifecycle.md) ride the
|
||||
notification mechanism:
|
||||
|
||||
- **Signals** arrive as notifications on the endpoint a process nominated with
|
||||
`signal_bind` (`process.bindSignals`): badge = the signal bit plus the
|
||||
coalesced pending mask (`process.signalsFrom` decodes). Statements,
|
||||
never questions; no payload, no reply.
|
||||
- **One-shot timers** (`timer_bind`, `time.timerOnce`) land as a
|
||||
timer-bit notification — the timed wait: a service arms a deadline and keeps
|
||||
serving, instead of blocking in sleep.
|
||||
- **The universal ping**: a **zero-length request is the liveness probe**,
|
||||
answered with a zero-length reply by the service harness itself
|
||||
(`service.run`). No protocol's requests start at length zero, so the
|
||||
encoding cannot collide, and a wedged service simply fails to answer — which
|
||||
is the diagnosis. Deep health ("can I reach my hardware?") stays a per-service
|
||||
protocol message.
|
||||
@@ -0,0 +1,246 @@
|
||||
# Native NVIDIA GPU support — feasibility and roadmap
|
||||
|
||||
**Status: research snapshot, not implemented.** This records what a *native* display driver for a
|
||||
real discrete NVIDIA GPU — specifically an **RTX 3060 (Ampere GA106)** — would take, and how it
|
||||
would slot into danos's pluggable scanout architecture. It is a survey of primary sources
|
||||
(NVIDIA's [open-gpu-kernel-modules](https://github.com/NVIDIA/open-gpu-kernel-modules), the Linux
|
||||
[nouveau/nvkm](https://github.com/torvalds/linux/tree/master/drivers/gpu/drm/nouveau) driver,
|
||||
NVIDIA's [open-gpu-doc](https://nvidia.github.io/open-gpu-doc/), and
|
||||
[linux-firmware](https://github.com/NVIDIA/linux-firmware)), not an implementation. The NVIDIA
|
||||
driver landscape moves quickly (GSP defaults, firmware ABIs); treat specifics as a mid-decade
|
||||
snapshot and re-verify against current source before building.
|
||||
|
||||
Read [display.md](display.md) and [display-v2.md](display-v2.md) first — this doc assumes the
|
||||
v2 model where scanout is a **pluggable backend** and a native driver is just another `.scanout`
|
||||
service (like the virtio-gpu one), announcing to the compositor over `attach_scanout`.
|
||||
|
||||
## TL;DR
|
||||
|
||||
- A **minimal display-only driver** (EDID + mode-set + framebuffer scanout, **no** 3D/compute)
|
||||
for the RTX 3060 **can and should avoid the GSP entirely**. nouveau has a register-level,
|
||||
CPU-driven display path for Ampere (`nvkm/engine/disp/ga102.c`) that lights up GA106 with no
|
||||
external firmware; the signed-firmware wall gates the **compute/graphics** engines (PGRAPH),
|
||||
**not** the display controller. "GSP is mandatory on Ampere" is true only for NVIDIA's own
|
||||
RM-object route.
|
||||
- **danos's UEFI GOP boot is the single biggest thing in its favour.** The VBIOS/GOP has already
|
||||
run devinit and brought up the display PLLs, so a driver attaches to a **live, initialized**
|
||||
GA106 — no firmware load, no cold-boot POST, no devinit interpreter. You reprogram a running
|
||||
display rather than bring one up from cold.
|
||||
- It is still a **hard, multi-week-to-months expert effort** (effort tier ≈ 4/5) dominated by
|
||||
NVDisplay channel-DMA programming, SOR/head routing, DisplayPort AUX + link training, and the
|
||||
display supervisor handshake. The GSP/RM route is tier 5 (near-infeasible solo).
|
||||
- The **licensing tension is counterintuitive**: the permissively-licensed reference (NVIDIA
|
||||
open-gpu-kernel-modules, MIT/GPLv2) is the **hard GSP path**; the register-level display code
|
||||
you actually want lives in **GPL nouveau**. See [Licensing](#licensing).
|
||||
- The **window is closing**: GA10x (Ampere) is the *last* NVIDIA family with a register-level
|
||||
display path — Ada (RTX 40) deleted its non-GSP display HAL. Targeting Ampere specifically
|
||||
matters.
|
||||
- **Recommendation:** for *this card*, GOP already gives native-resolution scanout with zero GPU
|
||||
code and zero maintenance. A native driver buys only runtime mode changes, hardware
|
||||
vsync/vblank, and multihead. It is justified if that runtime control is a danos goal, or to
|
||||
*learn the craft* — for which an Intel iGPU or a pre-Turing NVIDIA card reaches "first pixel"
|
||||
far faster.
|
||||
|
||||
## The GSP wall, and why display sits on the near side of it
|
||||
|
||||
On Turing and later, NVIDIA split its driver's Resource Manager into a host **CPU-RM** and a
|
||||
**GSP-RM** running on an on-die RISC-V core ("Peregrine"), talking over RPC
|
||||
([LWN 953144](https://lwn.net/Articles/953144/)). The GSP is a *full resource manager*, not a
|
||||
display coprocessor — there is no "display-only" GSP image and no small display RPC subset. Its
|
||||
boot chain is entirely signed and mandatory: a VBIOS-resident **FWSEC-FRTS** app carves a
|
||||
write-protected region (WPR2), a signed **Booter** on the SEC2 falcon loads the GSP bootloader,
|
||||
and that loads **GSP-RM** inside WPR. The firmware ships pre-computed signatures and the driver
|
||||
picks one by an on-chip fuse-version register — **you cannot self-sign**, and there is **no stable
|
||||
firmware ABI** (it is revised every driver release; nouveau and the Rust nova-core driver each pin
|
||||
exactly one version). A GSP driver is a permanent maintenance liability, not a one-time build
|
||||
([LWN 1037379](https://lwn.net/Articles/1037379/),
|
||||
[nova-core cover letter](https://lore.freedesktop.org/nouveau/20250826-nova_firmware-v2-7-93566252fe3a@nvidia.com/T/)).
|
||||
|
||||
**But display doesn't need any of that on Ampere.** `nvkm/engine/disp/ga102.c` dual-dispatches:
|
||||
|
||||
```
|
||||
if (nvkm_gsp_rm(device->gsp)) return r535_disp_new(&ga102_disp, ...); // GSP RPC path
|
||||
return nvkm_disp_new_(&ga102_disp, ...); // direct register path
|
||||
```
|
||||
|
||||
Both branches use the same `ga102_disp` HAL and the same `GA102_DISP_*` class IDs; GSP merely
|
||||
swaps register programming for RPC. GA106 (chipset `0x176`) is wired to `ga102_disp_new` in the
|
||||
device table, identical to GA102/103/104/107. Ampere lit up displays via the **direct** path in
|
||||
Linux 5.11/5.17 — two years before GSP-RM landed (6.7, 2023)
|
||||
([ga102.c](https://raw.githubusercontent.com/torvalds/linux/master/drivers/gpu/drm/nouveau/nvkm/engine/disp/ga102.c),
|
||||
[Phoronix GA106](https://www.phoronix.com/news/Nouveau-NVIDIA-GA106)).
|
||||
|
||||
**Caveat — this is now the legacy path.** As of Linux 6.18, nouveau defaults to GSP on
|
||||
Turing/Ampere; the direct path is a retained, forceable fallback (`nouveau.config=NvGspRm=0`, and
|
||||
automatic when GSP firmware is absent). It is stable and proven, but NVIDIA and nova-core are
|
||||
moving to GSP-only, and **Ada already deleted its non-GSP display HAL**. GA10x is the last family
|
||||
that keeps a register-level display path.
|
||||
|
||||
## What "direct" actually entails
|
||||
|
||||
"Direct" is not "plain register pokes." Only SOR / PLL / DP-link / clock setup is bare MMIO. The
|
||||
**mode-set and scanout themselves flow through the NVDisplay channels — a DMA pushbuffer**:
|
||||
|
||||
- Display classes for Ampere (the C670 family): core `GA102_DISP_CORE_CHANNEL_DMA` (`0xc67d`),
|
||||
window `0xc67e`, window-immediate `0xc67b`, cursor `0xc67a` (headers `clc67d.h` / `clc67e.h` /
|
||||
`clc67a.h` in [open-gpu-doc `classes/display/`](https://github.com/NVIDIA/open-gpu-doc/tree/master/classes/display)).
|
||||
- The core channel needs **instance memory, a RAMHT, DMA objects, and a channel user-MMIO
|
||||
region** ([disp/chan.c](https://raw.githubusercontent.com/torvalds/linux/master/drivers/gpu/drm/nouveau/nvkm/engine/disp/chan.c)).
|
||||
The register-level "plumbing" to allocate/kick a channel is in NVIDIA's GA102 display register
|
||||
manual: `NV_PDISP_FE_CHNCTL_CORE/WIN/CURS`, `NV_PDISP_FE_PBBASE/PBBASEHI`
|
||||
([dev_display_withoffset.ref.txt](https://raw.githubusercontent.com/NVIDIA/open-gpu-doc/master/manuals/ampere/ga102/dev_display_withoffset.ref.txt)).
|
||||
- **Mode-set is a method stream** on the core channel: `HEAD_SET_RASTER_*`,
|
||||
`HEAD_SET_PIXEL_CLOCK_FREQUENCY`, `HEAD_SET_CONTROL_OUTPUT_RESOURCE`, `SOR_SET_CONTROL`
|
||||
(protocol select), viewport/scaler, then `UPDATE`. The window channel points at the scanout
|
||||
surface (`SET_CONTEXT_DMA_ISO`, `SET_STORAGE`, `SET_OFFSET`).
|
||||
- After `UPDATE` you must complete the display **supervisor** interrupt handshake (SV1/SV2/SV3).
|
||||
|
||||
**EDID and DisplayPort are a separate subdev you must port.** open-gpu-doc documents *none* of
|
||||
EDID/DDC/AUX. On the direct path you read EDID in-driver via nouveau's `nvkm/subdev/i2c`: bit-bang
|
||||
**DDC/I²C at address `0x50`** (E-DDC `0x30`) for TMDS/HDMI, or native **DP AUX** in `i2c/aux.c`
|
||||
for DisplayPort. DisplayPort **link training** (the `dp.c` `train_cr` / `train_eq` state machine
|
||||
over AUX — clock recovery, lane/rate, voltage-swing/pre-emphasis) is the single hardest and most
|
||||
fragile piece; a DVI/HDMI (TMDS) panel avoids it entirely.
|
||||
|
||||
## The memory floor (smaller than you'd fear)
|
||||
|
||||
Neither route hands you a framebuffer allocator — even GSP-RM does not manage the scanout
|
||||
framebuffer; the driver owns VRAM and merely tells GSP where its page directory is. But
|
||||
display-only is a small fraction of a full GEM/TTM stack:
|
||||
|
||||
- **Pitch-linear (untiled) scanout is allowed** on nv50→Ampere — the window's storage method has a
|
||||
`PITCH` layout mode, so you skip block-linear tiling math
|
||||
([wndwc37e.c](https://raw.githubusercontent.com/torvalds/linux/master/drivers/gpu/drm/nouveau/dispnv50/wndwc37e.c)).
|
||||
- The window references its surface through a simple **display context-DMA**
|
||||
(`SET_CONTEXT_DMA_ISO` + a 256-byte-granular `SET_OFFSET = addr>>8`) — a base/limit descriptor,
|
||||
**not** the GPU's 5-level compute page tables. **No full GPU VMM is needed** for scanout.
|
||||
- The surface must live in **VRAM** in practice (nouveau always pins scanout to VRAM). *Open
|
||||
question:* whether GA10x can scan out from a system-memory (GART) surface via a sysmem-target
|
||||
ctxdma — which would let danos skip a VRAM allocator. No source forbids it; nouveau never does
|
||||
it (confidence: medium).
|
||||
- **CPU access** to the framebuffer for compositing goes through **BAR1** (a VRAM aperture); BAR0
|
||||
is the 16 MB register window. BAR1 can be smaller than 12 GB of VRAM unless Resizable BAR maps
|
||||
it all.
|
||||
|
||||
**Net:** you need (1) a contiguous aligned VRAM allocator (256-byte base, pitch a multiple of
|
||||
64 bytes — confirm against the Ampere display refs), (2) a little instmem for the channel
|
||||
pushbuffers + iso ctxdma, (3) a BAR1 CPU mapping. You do **not** need the 5-level VMM, GEM/TTM
|
||||
eviction, or tiling.
|
||||
|
||||
## Licensing
|
||||
|
||||
The tension is the opposite of convenient:
|
||||
|
||||
- **NVIDIA open-gpu-kernel-modules is dual MIT/GPLv2** — usable under MIT, no copyleft on your
|
||||
other code — **but its display logic is the GSP/RM-object route.** Its class headers
|
||||
(`cl0073.h`, `cl2080.h`, `ctrl0073*.h`) are useful, permissive references.
|
||||
- **nouveau is GPLv2**, and the **register-level display sequences you actually want live in
|
||||
nouveau**, not in the MIT code. So the *easy technical path is the GPL-licensed one.* Reading
|
||||
GPL nouveau and reimplementing it in Zig is a derivative-work risk proportional to how closely
|
||||
your code tracks its structure/constants.
|
||||
|
||||
Options: **(a)** accept that the danos NVIDIA display driver is a **GPL component**. danos's
|
||||
userspace-driver-over-IPC model (a driver is a separate process behind a defined protocol, not
|
||||
linked into the kernel) is about the cleanest possible GPL boundary, so the GPL would be contained
|
||||
to that one binary and the rest of danos could keep its own license — but this is a
|
||||
licensing-boundary judgement that wants real diligence, not a settled fact. **(b)** clean-room
|
||||
from *specification* rather than *code*: [envytools](https://envytools.readthedocs.io) + NVIDIA's
|
||||
open-gpu-doc register manuals + the MIT OGKM class headers, treating nouveau as
|
||||
documentation-of-last-resort.
|
||||
|
||||
**Firmware licensing is moot for the direct path** (no firmware is loaded). For completeness: the
|
||||
GSP blobs are marked redistributable under `LICENCE.nvidia`, which permits use by **any
|
||||
OSI-approved open-source OS** (not just Linux), on NVIDIA GPUs, **unmodified**, with **no
|
||||
reverse-engineering of the firmware binary**. The one gate — is danos released under an OSI
|
||||
license? — is only reached on the GSP route, which this doc recommends against for this card.
|
||||
|
||||
## Prior art
|
||||
|
||||
**No one has built a from-scratch native NVIDIA driver outside Linux.** FreeBSD ships
|
||||
`nvidia-drm-kmod`, a *port of NVIDIA's own closed `nvidia-drm.ko`* loading the GSP blob (its old
|
||||
nouveau port was removed). Haiku's NVIDIA support is likewise a *port of OGKM* (GSP, Turing+, very
|
||||
alpha). OpenBSD / DragonFly have neither. Every non-Linux OS that supports modern NVIDIA chose to
|
||||
**wrap NVIDIA's GSP stack** rather than write a native driver. A danos direct-register driver
|
||||
would have exactly one reference implementation — GPL nouveau — and no non-Linux precedent.
|
||||
|
||||
## Alternatives
|
||||
|
||||
| Option | What you get | The tradeoff |
|
||||
|---|---|---|
|
||||
| **Stay on GOP** (working today) | Native-res scanout, zero GPU code/firmware/maintenance | Resolution frozen at ExitBootServices; **no runtime mode change, no hardware vsync, no multihead** |
|
||||
| **Pre-Turing NVIDIA** (Kepler / early Maxwell) | Direct EVO/disp-core + CRTC/PLL modeset, **no signed firmware, no coprocessor**; mature nouveau reference | Older display class; not this card; only reclocking is firmware-gated |
|
||||
| **Intel iGPU** | **Publicly documented** register interfaces (Intel PRMs); no coprocessor mediating modeset | i915 is huge + generation-specific; write one generation from the PRM |
|
||||
| **Native GA106 direct** (this doc) | Runtime modeset, vsync, multihead on the actual card | Tier-4 effort; GPL reference; DP link training; legacy/de-emphasized path |
|
||||
| **GA106 via GSP/OGKM** | Also unlocks 3D / reclocking later | Tier-5; ~14k-line ante; unstable version-pinned ABI; unprecedented outside Linux |
|
||||
|
||||
## "First light" milestones (direct path, inheriting GOP state)
|
||||
|
||||
Framed as a danos `.scanout` service (like the virtio-gpu driver), taking the direct register path
|
||||
and inheriting the GOP-initialized display — no signed firmware, no devinit, no GSP:
|
||||
|
||||
1. **PCI/BAR bring-up** — enumerate GA106 (`0x176`), map **BAR0** (registers) and **BAR1** (VRAM
|
||||
aperture) via danos MMIO grants; confirm the display engine is GOP-live.
|
||||
2. **VRAM + instmem allocator** — contiguous aligned VRAM for the scanout surface (256-byte base)
|
||||
+ small instmem for pushbuffers / RAMHT / iso ctxdma. No VMM, no TTM.
|
||||
3. **EDID** — port `nvkm/subdev/i2c` DDC (`0x50`) + DP-AUX (`aux.c`); read + parse the panel EDID.
|
||||
4. **Core channel up** — allocate the `0xc67d` core channel as a DMA pushbuffer; stand up the
|
||||
SV1/SV2/SV3 supervisor-interrupt handshake.
|
||||
5. **First pixel = reprogram, don't re-POST** — bind a window (`0xc67e`) at the existing WC
|
||||
framebuffer via `SET_CONTEXT_DMA_ISO` + `SET_OFFSET`, pitch-linear, `UPDATE`; prove you can
|
||||
drive the *current* GOP mode from your own channel before changing anything.
|
||||
6. **Modeset** — push raster timings on a head, route head→SOR→connector, program the pixel-clock
|
||||
PLL, switch to an EDID mode (needs the `clc67d/e` method opcodes from the OGKM headers + the
|
||||
supervisor timing from nouveau `head.c`).
|
||||
7. **DisplayPort link training** — only if the panel is DP and GOP's link can't be reused; the
|
||||
`dp.c` `train_cr`/`train_eq` state machine. TMDS/HDMI is far simpler.
|
||||
8. **Wire into the compositor `.scanout` backend** (`attach_scanout`), add vsync via the display
|
||||
interrupt, then multihead.
|
||||
|
||||
Keep the GOP backend as the fallback the whole way — a stall at any step still leaves danos with a
|
||||
working display (exactly the resilience v2 already provides via re-attach).
|
||||
|
||||
## Reading list
|
||||
|
||||
**Direct path — nouveau (GPLv2):**
|
||||
- `nvkm/engine/disp/ga102.c` — the GA10x display HAL + the GSP/non-GSP dispatch.
|
||||
- `nvkm/engine/disp/{head.c, ior.c, dp.c, hdmi.c, chan.c}` — head/SOR routing, DP AUX + link
|
||||
training, channel-DMA plumbing.
|
||||
- `dispnv50/{corec37d.c, corec57d.c, wndwc37e.c, wndwc57e.c, wndwc67e.c, headc37d.c, cursc37a.c}`.
|
||||
- `nvkm/subdev/i2c` (DDC + `aux.c`) for EDID; `nvkm/subdev/bios/init.c` + `devinit/` **only** if
|
||||
you ever have to re-POST (danos's GOP handoff means you shouldn't).
|
||||
|
||||
**Object model / GSP path — NVIDIA OGKM (MIT/GPLv2):** class headers `cl0073.h`, `cl2080.h`,
|
||||
`ctrl0073system.h`, `ctrl0073specific.h`; `src/nvidia/` for RM control sequences.
|
||||
`nvidia-modeset.ko` (NVKMS) is a *policy* layer over RM and can be bypassed entirely.
|
||||
[nova-core](https://lore.freedesktop.org/nouveau/) (Rust) is the forward-looking reference for GSP
|
||||
boot mechanics (falcon signing, queue rings, RPC).
|
||||
|
||||
**Register / method specs — NVIDIA open-gpu-doc:**
|
||||
- [`classes/display/README.txt`](https://raw.githubusercontent.com/NVIDIA/open-gpu-doc/master/classes/display/README.txt)
|
||||
— the channel model + class-to-GPU map (read first).
|
||||
- [`classes/display/clc67d.h`](https://raw.githubusercontent.com/NVIDIA/open-gpu-doc/master/classes/display/clc67d.h)
|
||||
+ `clc67e.h` / `clc67a.h` — the Ampere core/window/cursor mode-set method vocabulary.
|
||||
- [`manuals/ampere/ga102/dev_display_withoffset.ref.txt`](https://raw.githubusercontent.com/NVIDIA/open-gpu-doc/master/manuals/ampere/ga102/dev_display_withoffset.ref.txt)
|
||||
— `NV_PDISP_FE_*` channel/pushbuffer registers + SOR.
|
||||
- [`DCB`](https://github.com/NVIDIA/open-gpu-doc/tree/master/DCB) — connector→output-resource
|
||||
routing; [`Devinit`](https://github.com/NVIDIA/open-gpu-doc/tree/master/Devinit) +
|
||||
[`BIOS-Information-Table`](https://github.com/NVIDIA/open-gpu-doc/tree/master/BIOS-Information-Table)
|
||||
— VBIOS parsing (bring-up reference; not needed if inheriting GOP).
|
||||
- The 632 KB Volta [`dev_display.ref`](https://download.nvidia.com/open-gpu-doc/Display-Ref-Manuals/1/gv100/dev_display.ref)
|
||||
is the best shot at SOR-DP/AUX register detail the smaller Ampere file omits.
|
||||
|
||||
## Open questions (unresolved by the survey)
|
||||
|
||||
Each needs a direct read of the named nouveau file or experimentation on the actual card:
|
||||
|
||||
- Exact GA106 register/method offsets and PADLINK→SOR→connector wiring (can vary by board vendor).
|
||||
- Whether *any* PLL/devinit re-run is unavoidable vs. fully inherited from GOP.
|
||||
- Whether DisplayPort needs full retraining on takeover, or the GOP-established link can be reused.
|
||||
- The precise SV1/SV2/SV3 supervisor sequence.
|
||||
- Whether a system-memory-target scanout ctxdma could eliminate the VRAM allocator.
|
||||
- The exact `clc67d.h`/`clc67e.h` method opcode numbers (not captured verbatim in the survey).
|
||||
|
||||
---
|
||||
|
||||
*Research snapshot; verify against current nouveau / open-gpu-kernel-modules source before
|
||||
building — NVIDIA's GSP defaults and firmware ABIs change per release.*
|
||||
@@ -0,0 +1,114 @@
|
||||
# USB hubs (M22)
|
||||
|
||||
A hub is USB **bus infrastructure**, not an application peripheral, so hub
|
||||
topology is handled **inside the `usb-xhci-bus` driver** — the process that owns
|
||||
the controller's device slots and contexts. A device behind a hub is not reached
|
||||
by any hub-specific software path: it is reached by the **controller**,
|
||||
programmed with a *route string* in its slot context. Route strings and slot
|
||||
contexts are xHCI hardware concepts that only exist inside the controller driver,
|
||||
so that is where hub handling belongs. Class drivers (HID, storage) stay separate
|
||||
and unaware — the hub is transparent to them; a keyboard behind a hub reaches the
|
||||
same `usb-hid-keyboard` driver as one on a root port.
|
||||
|
||||
This is a deliberate scoping choice, not a microkernel compromise: the USB *bus*
|
||||
driver handles USB *bus* topology. The alternative — a separate `usb-hub`
|
||||
class-driver process plus a cross-process enumeration protocol — would only
|
||||
shuttle the bus's own topology state (slot ids, route strings, TT linkage) out to
|
||||
another process and back, since the hub driver cannot build a slot context
|
||||
itself.
|
||||
|
||||
## The compound-hub reality
|
||||
|
||||
A USB 3.0 hub is physically **two hubs** sharing each connector: a SuperSpeed hub
|
||||
and a USB 2.0 companion hub, enumerated as **separate devices on separate root
|
||||
ports**. A full- or low-speed device plugged into a USB 3.0 hub attaches to the
|
||||
**USB 2.0 companion**, not the SuperSpeed hub. So supporting full-speed devices
|
||||
(keyboards, mice) behind a hub means driving the USB 2.0 companion and handling
|
||||
**transaction translators** — there is no SuperSpeed-only shortcut that reaches a
|
||||
full-speed keyboard.
|
||||
|
||||
## Slot-context fields for a downstream device
|
||||
|
||||
`buildAddressInputContext` fills the Slot Context from the device record: a
|
||||
root-port device carries route 0 and its own root-hub port. A downstream device
|
||||
additionally carries:
|
||||
|
||||
- **Route String** (Slot Context dword 0, bits 19:0) — 5 tiers × 4 bits, each
|
||||
tier the downstream hub-port number. Composed as
|
||||
`route = (parent_route << 4) | hub_port`, capped at the xHCI 5-tier max.
|
||||
- **Root Hub Port Number** (dword 1, bits 23:16) — the *root* port the whole hub
|
||||
chain hangs off, inherited from the parent hub (not the hub's own port number).
|
||||
- **Speed** (dword 0, bits 23:20) — read from the hub's downstream port status
|
||||
after reset, not assumed.
|
||||
- **Parent Hub Slot ID** (dword 2, bits 7:0) + **Parent Port Number** (dword 2,
|
||||
bits 13:8) — the **transaction translator**: set when a full/low-speed device
|
||||
sits behind a high-speed hub, so the controller routes split transactions
|
||||
through that hub's TT. For a multi-TT hub, **MTT** (Slot Context dword 0 bit
|
||||
25) is set and the TT port is the device's own hub port.
|
||||
|
||||
## Detection: the status-change interrupt endpoint
|
||||
|
||||
A hub has one interrupt IN endpoint that returns a **port-status-change bitmap**
|
||||
(bit N set = port N changed). The bus arms an interrupt transfer on it (reusing
|
||||
the controller's existing interrupt-endpoint machinery, but serviced
|
||||
**in-process** — no class-driver subscription IPC), and on each report:
|
||||
|
||||
1. For each changed port, `GET_STATUS` (hub class request) reads the port's
|
||||
connect/enable/reset state and speed, and `CLEAR_FEATURE(C_PORT_*)`
|
||||
acknowledges the change.
|
||||
2. On a **connect**: `SET_FEATURE(PORT_RESET)`, wait for reset-complete via a
|
||||
later status-change report, read the enabled speed, then `setupDevice` with
|
||||
the composed route string / root port / TT fields, `enumerate`, and register
|
||||
the interfaces — exactly the existing path, recursing if the new device is
|
||||
itself a hub.
|
||||
3. On a **disconnect**: tear down the downstream device (report each interface
|
||||
`ChildRemoved`, Disable Slot) — the B3 teardown path, keyed by the device's
|
||||
route rather than a root port.
|
||||
|
||||
## Hub setup (once, when the hub enumerates)
|
||||
|
||||
When the bus scan (or a hot-plug bring-up) finds a device of class 9:
|
||||
|
||||
1. Read the **hub descriptor** (class GET_DESCRIPTOR, type 0x2A for a USB 3.0
|
||||
hub / 0x29 for USB 2.0) → downstream port count, characteristics.
|
||||
2. For a USB 3.0 hub, `SET_FEATURE(BH_PORT_RESET)` semantics and the depth
|
||||
(`SET_HUB_DEPTH`) so the hub knows its tier for route-string forwarding.
|
||||
3. `SET_FEATURE(PORT_POWER)` each downstream port.
|
||||
4. Configure the hub's slot as a hub: **Hub** bit (Slot Context dword 0 bit 26),
|
||||
**Number of Ports** (dword 1, bits 31:24), **TT Think Time** and **MTT** for a
|
||||
USB 2.0 multi-TT hub — via an Evaluate/Configure Endpoint on the hub's slot.
|
||||
5. Arm the status-change interrupt endpoint.
|
||||
|
||||
## Testing
|
||||
|
||||
QEMU's `usb-hub` is a USB 2.0 single-TT hub. A **static boot topology** on a
|
||||
dedicated second controller (`-device qemu-xhci,id=xhci2 -device
|
||||
usb-hub,bus=xhci2.0,port=1 -device usb-kbd,bus=xhci2.0,port=1.1` — isolated from
|
||||
the boot controller's auto-assigned devices, whose ports the hub would collide
|
||||
with) presents the downstream device connected from the start, so the bus reads
|
||||
it on the first status-change report — exercising the full path (hub setup, TT slot context,
|
||||
downstream enumerate, class-driver bind) without needing a hot-plug event. A new
|
||||
`usb-hub` QEMU case asserts the hub enumerates, the downstream keyboard
|
||||
enumerates behind it, and `usb-hid-keyboard` binds.
|
||||
|
||||
Real-hardware validation (the user's SuperSpeed Genesys hub + full-speed
|
||||
keyboard/mouse on its USB 2.0 companion) is flagged separately — the compound
|
||||
USB 3.0 hub path is not modelled by QEMU's USB 2.0 hub.
|
||||
|
||||
## Milestones (all complete)
|
||||
|
||||
- **B4a** ✓ — hub recognition + setup: detect class 9 in the scan, read the hub
|
||||
descriptor, configure the slot as a hub, power downstream ports, log the
|
||||
topology.
|
||||
- **B4b** ✓ — downstream enumeration: the in-process status-change subscription,
|
||||
port reset, Address Device with route string + root port + TT fields,
|
||||
enumerate + register. A full-speed keyboard behind a USB2 hub binds
|
||||
`usb-hid-keyboard` in QEMU.
|
||||
- **B4c** ✓ — disconnect teardown (recursive: a hub takes its subtree with it)
|
||||
and hub-behind-hub recursion (route strings compose across tiers). QEMU's hub
|
||||
*does* raise downstream status changes, so both connect and disconnect are
|
||||
harness-tested (`usb-hub`, `usb-hub-nested`, `usb-hub-unplug`).
|
||||
|
||||
Real-hardware validation of the user's SuperSpeed Genesys hub with full-speed
|
||||
devices on its USB 2.0 companion remains pending — QEMU's USB 2.0 hub does not
|
||||
model the compound USB 3.0 hub.
|
||||
Reference in New Issue
Block a user