re-org docs

This commit is contained in:
Daniel Samson
2026-07-23 00:25:34 +01:00
parent 52d6e372fd
commit 757c6f14c3
52 changed files with 156 additions and 156 deletions
+344
View File
@@ -0,0 +1,344 @@
# Native AMD GPU support — feasibility and roadmap
**Status: research snapshot, not implemented.** This records what a *minimal, display-only*
native driver for a real discrete AMD GPU — specifically an **RX 6600-class card (Navi 23,
RDNA2, DCN 3.0.2)**, the market analog of the RTX 3060 — would take, and how it slots into
danos's pluggable scanout architecture. It is a survey of primary sources (the Linux
[amdgpu Display Core](https://github.com/torvalds/linux/tree/master/drivers/gpu/drm/amd/display)
driver and its [kernel documentation](https://docs.kernel.org/gpu/amdgpu/display/index.html),
the AtomBIOS interpreter in
[drivers/gpu/drm/amd](https://github.com/torvalds/linux/tree/master/drivers/gpu/drm/amd),
linux-firmware's `LICENSE.amdgpu`, and Haiku's
[radeon_hd](https://github.com/haiku/haiku/tree/master/src/add-ons/kernel/drivers/graphics/radeon_hd)),
not an implementation. It completes the trilogy with [nvidia-gpus.md](nvidia-gpus.md) and
[intel-igpu.md](intel-igpu.md) and should be read against both — AMD lands *between* them:
NVIDIA-class discrete-card mechanics, but Intel-class (better, in one way) reference material.
Read [display.md](display.md) and [display-v2.md](display-v2.md) first — this doc assumes the
v2 model where scanout is a **pluggable backend** and a native driver is just another `.scanout`
service (like the virtio-gpu one), announcing to the compositor over `attach_scanout`.
## TL;DR
- **AMD's decisive advantage is that the vendor's own display driver is the register manual, and
it's MIT-licensed.** The entire Display Core (DC) — hardware sequencer, per-block code for
OTG/OPTC, HUBP, DPP, MPC, DIO/link encoders, plus the `asic_reg` register headers — ships in
the Linux tree under MIT/X11, deliberately written OS-agnostic because AMD shares it across
operating systems. You can study it, port it, even copy from it into a danos driver without
license contamination. NVIDIA has no analog (nouveau is GPL); Intel has PRM prose but you
still write the code yourself.
- **The firmware wall is one small blob, not a GSP.** The only display-side firmware the Linux
driver hard-requires is **DMCUB** (the display microcontroller), and only on **DCN 2.1
through 4.x** — which includes Navi 23. It is redistributable from linux-firmware, and it is
a display helper, not a full-card resource manager: on DCN 3.0.x, hardware init
(`dcn30_init_hw`) is **host-driven direct register programming** — the only DMUB call in it
is a capability query. All DCE generations, DCN 1.0 (Raven), and DCN 2.0 (Navi 10/12/14) run
display with **no display firmware at all**.
- **Whether the *silicon* (vs. the Linux driver) needs DMCUB for a bare GOP-inheriting modeset
is unproven** — Linux fails init with `-EINVAL` if the blob is missing on a DMUB ASIC, but
what it's *used for* at minimum scope (vs. PSR/ABM/offloaded DP link training) isn't
documented. The safe plan ships the blob; it's legally and practically cheap to do so.
- **Programming model is direct MMIO, not channel DMA.** DCN mode-set is ordered register-write
sequences (the DC "hardware sequencer") against named, header-documented registers — no
pushbuffers, no method streams, no RAMHT, no supervisor-interrupt handshake. This deletes the
hardest structural layer of the NVIDIA path.
- **Scanout is VRAM-only on discrete cards** — the claim that DCN can scan out of GTT/system
memory was checked and *refuted* for dGPUs (Linux allows GTT scanout only on select APUs). So
a small VRAM allocator + BAR CPU mapping is required, same as NVIDIA. Pitch-linear surfaces
are supported; no DCC/tiling needed.
- **danos's GOP boot helps here too, with a caveat.** DC explicitly models taking over a
VBIOS/GOP-lit pipe (`dc_validate_boot_timing` reads back live DIG/OTG/pixel-clock state), so
"repoint the surface on the running pipe" is demonstrably hardware-feasible — but Linux's
seamless-boot path is **eDP-only and default-off on discrete cards**, so plan on a full
self-owned modeset (including DP retrain) right after first light rather than living on the
inherited link.
- **AMD has real non-Linux prior art — but only for the old hardware.** Haiku's MIT `radeon_hd`
mode-sets by executing VBIOS **AtomBIOS command tables** through AMD's own MIT interpreter;
its compiled-in ceiling is **DCE 8.5 (Hawaii, ~2013)** — every Polaris/Vega/Navi entry sits
in a `#if 0` block. There is zero non-Linux DCN precedent; a danos DCN driver would be first.
- **Effort tier ≈ high-3 to 4** for a native DCN 3.0.x display-only driver on Navi 23 — the raw
register surface is GA106-class (tier 4), but the MIT vendor reference, the one-blob firmware
wall, and the absence of channel-DMA plumbing pull real risk out. The AtomBIOS-interpreter
route is tier ≈ 3 but dead-ends at pre-2016 silicon.
- **Recommendation:** the same sober conclusion as the other two docs — GOP already gives
native-res scanout for zero code — but if danos ever does drive real discrete silicon
natively, **an RDNA2 card is the best target of the three**: modern, mainstream, in-warranty
hardware with a legally clean, vendor-authored reference. That combination exists nowhere
else.
## The firmware wall (a fence, next to NVIDIA's wall)
AMD GPUs carry a zoo of firmware: PSP (security processor), SMU (power/clock management), CP/RLC
(graphics), SDMA, VCN (media) — and, on the display side, DMCU (legacy) then **DMCUB**
("Display Micro-Controller Unit, version B"), a per-generation blob in linux-firmware
(`navi23_dmcub.bin` etc.). The display-only question is: which of these does a scanout driver
actually need?
The Linux answer is precise and readable in
[`amdgpu_dm.c`](https://github.com/torvalds/linux/blob/master/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm.c):
`dm_init_microcode()` switches on the display IP version — **DCN 2.1 (Renoir) through DCN
3.0.x / 3.1.x / 3.2 / 3.5 / 4.x** request a DMCUB blob as `AMDGPU_UCODE_REQUIRED`, and
`dm_dmub_hw_init()` fails driver init with `-EINVAL` if it's absent. Everything earlier — **all
of DCE (Southern Islands through Vega), DCN 1.0 (Raven), and DCN 2.0 (Navi 10/12/14)** — hits
the `default:` case, `dmub_srv` stays NULL, and display runs with no display firmware at all
([kernel display-manager doc](https://docs.kernel.org/6.2/gpu/amdgpu/display/display-manager.html)).
Two nuances survive verification:
- **The requirement is Linux-driver enforcement backed by real functional need, and it's
version-sensitive.** AMD force-switched all Renoir ASICs to DMUB to fix a USB-C/resume bug
(kernel commit `652de07addd2`, "with new dmub f/w dmcu is superseded"), which regressed users
on old blobs and had to be patched with explicit `dmcub_fw_version` gating (`91adec9e0709`).
What DMUB is *used for* varies with blob version. Ship a current blob.
- **DMCUB is an architectural fixture, not a bolt-on** — the DCN hardware itself contains a DMU
block housing the microcontroller
([DCN overview](https://docs.kernel.org/gpu/amdgpu/display/dcn-overview.html)) — but it is
**not a mediator of the programming model** on DCN 3.0: `dcn30_init_hw()` initializes clocks,
disables power gating, and powers up link encoders via direct register writes; its sole DMUB
interaction is `dc_dmub_srv_query_caps_cmd`. Firmware-*assisted* PHY/link bring-up appears
from **DCN 3.1** onward — one more reason to target 3.0.x. Features like PSR and ABM are
DMUB-offloaded on all generations; a minimal driver simply doesn't enable them.
**Contrast with NVIDIA's GSP:** the GSP is a full resource manager with a signed multi-stage
boot chain and a firmware ABI that breaks every driver release. DMCUB is a display helper blob
you copy onto the boot image once, load into a reserved buffer, and mostly ignore. There is no
signature fuse-matching, no WPR carve-out, no RPC-only register access. The one genuinely open
question — whether a GOP-inheriting minimal modeset could skip DMCUB entirely on DCN 3.0.2 —
doesn't need answering, because shipping the blob costs nothing (see [Licensing](#licensing)).
**PSP/SMU remain the flagged risk.** Nothing display-only touches CP/RLC/SDMA (those gate the
graphics rings, exactly like NVIDIA's PGRAPH — irrelevant here). But `dcn30_init_hw` calls into
the clock manager, and on discrete cards the clock manager may message the SMU to change display
clocks (DISPCLK/DPPCLK). Whether inherited GOP boot clocks suffice for a same-or-lower mode —
avoiding SMU (and hence PSP firmware-load) entirely — is the largest unverified assumption in
the milestone list below. The survey produced no confirmed claim either way.
## The display engine landscape
Two eras, one boundary that matters:
| Generation | Display IP | Cards | Display firmware | Route |
|---|---|---|---|---|
| GCN 1–4 (SI→Polaris) | DCE 6/8/10/11 | HD 7000 → RX 580 | none | AtomBIOS tables or direct DCE registers |
| Vega / Raven | DCE 12 / DCN 1.0 | Vega 56/64, APUs | none | DC code (first DCN) |
| Navi 1x (RDNA1) | DCN 2.0 | RX 5500–5700 | none | DC code |
| Renoir APU | DCN 2.1 | 4000-series APUs | **DMCUB required** | DC code |
| **Navi 2x (RDNA2)** | **DCN 3.0.x** | **RX 6600–6900** | **DMCUB required** | **DC code, host-driven init** |
| RDNA3/RDNA4+ | DCN 3.1+/3.2/3.5/4.x | RX 7000/9000 | DMCUB required, fw-assisted PHY | DC code, more DMUB offload |
The best modern first-pixel target is **DCN 3.0.x**: it has the full MIT block stack from the
June 2020 Sienna Cichlid patch series (207 patches, Linux 5.9; Navi 23 reuses the dcn30
sequencer), host-driven hardware init, and sits *before* the DCN 3.1 shift toward
firmware-assisted link management. Older DCE cards are even simpler (no firmware at all, plus
the AtomBIOS escape hatch) but are 2013–2016 hardware; newer DCN 3.5/4.x pushes more into DMUB.
The DCN pipe, in one line each (the vocabulary the DC code speaks —
[programming model](https://docs.kernel.org/next/gpu/amdgpu/display/programming-model-dcn.html)):
**HUBP** fetches and unpacks the surface from memory (this is where the scanout address and
pitch live), **DPP** scales/converts colors, **MPC** blends planes (bypassable for one plane),
**OPP** packs output, **OTG/OPTC** generates raster timings (the CRTC), and the **DIO** block's
DIG encoders + PHY drive the connector. Mode-set is the DC *hardware sequencer* walking these
blocks with ordered register writes — plain MMIO with polling, no pushbuffer channels, no
supervisor interrupts. Structurally this is Intel-shaped, not NVIDIA-shaped.
## Two routes: AtomBIOS interpreter vs. native DC-derived registers
**AtomBIOS** is AMD's VBIOS bytecode: every card's ROM carries *data tables* (connector
topology, clock limits — the DCB equivalent) and *command tables* (`SetPixelClock`,
`SetCRTC_Timing`, `EnableCRTC`, DIG encoder/transmitter control), executed by a small
interpreter the driver embeds (`atom.c`, ~1.5k lines). The classic radeon driver and Haiku's
`radeon_hd` mode-set this way: parse the tables, execute them, and the VBIOS does the
register-level work for you — inherently per-board correct, since the tables come from the
card's own ROM.
- **Where it's proven:** through DCE 8.5 (Haiku's ceiling, below) and in Linux's pre-DC code
through Polaris (DCE 11.2). AMD's interpreter itself is MIT (Haiku ships AMD's own
`atom.cpp`, "Copyright 2008 Advanced Micro Devices").
- **Where it's unproven:** DCN. amdgpu's DC still *uses* AtomBIOS for init sub-steps
(`bios_golden_init` in `dcn30_init_hw` executes host-interpreted tables) and reads the data
tables for connector topology — but nobody drives a full DCN modeset from command tables, and
whether RDNA2 VBIOSes still carry a complete modeset path or vestigial init-only tables is an
open question no source answers. Do not bet on it.
**The native route** is: port the relevant slice of DC. Not wholesale — in-tree DC has
accumulated Linux-isms (kernel-FPU guards around DML, the bandwidth-calculation library, which a
single-plane fixed-mode driver can largely sidestep) — but the DC core is *designed* to be
retargeted: the kernel docs state outright that DC "is shared with other OSes" and holds the
OS-agnostic hardware programming behind a `dm_services` shim (register access, memory, delays,
firmware loading). Dave Airlie initially rejected the DAL/DC merge in 2016 *because* it was
AMD's cross-OS codebase — hostile-witness confirmation that this exact code runs outside Linux.
Reimplement the shim in Zig, and the dcn30 sequences sit on top.
## The memory floor
Same shape as NVIDIA's, and the survey *hardened* one assumption:
- **VRAM-only scanout on discrete cards.** The documented DCN fetch path is VRAM → Data Fabric
(SDP) → DCHUB → HUBP; the claim that display buffers can live in GTT/system memory was
refuted for dGPUs in verification — Linux permits GTT scanout only on select APUs. The NVIDIA
doc's "maybe sysmem ctxdma?" hope has a firm *no* here. Budget for a small VRAM allocator.
- **Pitch-linear is fine.** HUBP programs a surface address + pitch; linear (untiled, no DCC)
surfaces are first-class for scanout. No tiling math.
- **CPU access via the VRAM BAR.** Compositing writes go through the PCI VRAM aperture;
resizable BAR helps but isn't needed — one pitch-linear surface fits comfortably in a
fixed 256 MB small-BAR window.
- **No GPU VMM.** Display addresses are physical VRAM addresses programmed into HUBP; no page
tables, no GEM/TTM, no eviction.
**Net:** (1) a contiguous aligned VRAM allocator, (2) a BAR CPU mapping, (3) a small reserved
buffer for the DMCUB firmware regions. That's the whole memory story.
## Inheriting GOP state
danos's GOP boot pays off again, with sharper edges than on NVIDIA:
- **DC models pipe takeover explicitly.** `dc_validate_boot_timing()`
([dc.c](https://github.com/torvalds/linux/blob/master/drivers/gpu/drm/amd/display/dc/core/dc.c))
reads back *live* hardware — `is_dig_enabled` on the link encoder, OTG timing registers,
pixel clock within tolerance — and keeps the VBIOS/GOP-lit pipe running until first flip.
This is a vendor-blessed recipe for milestone 2 below: the exact register set that tells you
which pipe is alive and how it's configured.
- **But Linux's seamless path is eDP-only and APU-gated.** The code comment is blunt: "Support
seamless boot on EDP displays only", and the enabling check requires an APU with DCN ≥ 3.0
unless forced with `amdgpu.seamless=1`. On a discrete RX 6600 with DP/HDMI, Linux does a full
modeset at takeover. Read that as a warning, not a prohibition: repointing HUBP at your own
surface on the live pipe should still work (the readback code proves the state is
inspectable), but plan the full self-owned modeset — **including DP retraining** — as the
immediate next step, not a someday.
- **No supervisor handshake exists to re-learn.** The NVIDIA doc's SV1/SV2/SV3 open question has
no AMD counterpart; commit sequencing is ordered register writes + vblank/lock waits in the
hwseq, all visible in MIT source.
## Licensing
The inverse of the NVIDIA situation, and the single strongest argument for AMD:
- **The reference code is MIT.** `amdgpu_dm.c` carries `SPDX-License-Identifier: MIT`;
`dc/core/dc.c` and `dmub_srv.h` carry the full X11-style grant (use, copy, modify, merge,
publish, distribute, sell). The `asic_reg` register headers ship under the same terms. One
diligence note: SPDX tagging isn't uniform across the tree, so header-check each file before
copying from it — but no GPL files are known inside `dc/`. Where nouveau forces a
GPL-or-clean-room choice, here the *easy technical path and the permissive path are the same
path*.
- **Firmware redistribution is a solved problem.** linux-firmware's `LICENSE.amdgpu` grants
anyone a royalty-free right to reproduce and distribute the blobs, binary-only, with the
license text attached — no OSI-license gate like NVIDIA's, no AMD agreement needed. danos can
ship `navi23_dmcub.bin` (and PSP/SMU blobs if ever needed) on its boot image today. The same
license **prohibits reverse-engineering the blobs** — all programming knowledge must come
from the MIT source, never from blob disassembly. (VBIOS images aren't in linux-firmware;
they're read from the card's own ROM, as Haiku does.)
- **Prose register docs are a DCE-era artifact.** AMD's classic X.Org-hosted PDFs cover the old
families — and the famous `R6xx_3D_Registers.pdf` turns out to be 3D-only (verified: zero
display content; the display material lives in the separate per-ASIC Register Reference
Guides). For DCN there is **no prose display spec at all**: the MIT DC source plus the
`asic_reg` headers *are* the register manual. Plan accordingly.
## Prior art
AMD, unlike NVIDIA, has genuine working non-Linux precedent — with a hard generational ceiling:
- **Haiku `radeon_hd`** (MIT, still in the tree): a real, shipping, from-scratch display driver
that executes AtomBIOS command tables via AMD's own MIT interpreter. Verified ceiling:
the last *enabled* device entry is **Hawaii (DCE 8.5, 0x67be)**; everything newer —
Tonga/Fiji, Carrizo/Polaris, Vega/Raven, and every Navi/RDNA2 entry up to the RX 6900 XT —
sits inside one `#if 0 /* disabled for R1/beta5 */` block under the comment "WARN: DCE
versions below here get sketchy."
- **AmigaOS/MorphOS RadeonHD drivers** (hdrlab): commercial non-Linux Radeon display drivers,
again for the DCE era.
- **FreeBSD** `drm-kmod`: a port of Linux amdgpu (DC and all), not independent prior art — but
proof the DC codebase transplants.
**Nobody has driven DCN outside Linux-derived code.** A danos DCN 3.0.x driver would be a
first — but a first with the vendor's MIT code as its map, which is a different proposition
from nouveau-as-only-reference.
## Alternatives
| Option | What you get | The tradeoff |
|---|---|---|
| **Stay on GOP** (working today) | Native-res scanout, zero GPU code/firmware/maintenance | Resolution frozen at ExitBootServices; no runtime mode change, no hardware vsync, no multihead |
| **AtomBIOS interpreter on an old DCE card** | Proven end-to-end (Haiku); interpreter is small + MIT; board-correct by construction | 2013–2016 hardware ceiling; tier ≈ 3; teaches AtomBIOS, not modern DCN |
| **Native DCN 3.0.x on RX 6600** (this doc) | Runtime modeset, vsync, multihead on modern silicon; MIT vendor reference; one redistributable blob | Tier ≈ high-3–4; SMU/clock question open; no non-Linux precedent |
| **Port DC wholesale** (reimplement `dm_services`) | Vendor-maintained sequences verbatim; designed-for-porting seam | Big codebase to carry (DML, abstractions); Linux-isms to shear off; overkill for one plane |
| **RDNA3+/DCN 3.5+** | Newer cards | More DMUB offload (fw-assisted PHY from DCN 3.1); strictly harder than 3.0.x for no display-only gain |
## "First light" milestones (native DCN 3.0.x path)
Framed as a danos `.scanout` service, inheriting the GOP-initialized display:
1. **PCI/BAR bring-up** — enumerate Navi 23, map the register BAR and the VRAM BAR via danos
MMIO grants; prove the pipe is GOP-live by writing pixels into the *existing* GOP
framebuffer through the VRAM BAR.
2. **Read back the live pipe** — port the `dc_validate_boot_timing` register set: which OTG is
running, its timings, which DIG/link encoder is enabled, current HUBP surface address/pitch.
This is pure reads — zero risk, high information.
3. **Repoint the surface** — allocate a danos-owned pitch-linear VRAM surface, program the HUBP
surface address/pitch on the live pipe at vblank. First self-owned pixel with **no modeset,
no firmware, no clock changes**.
4. **DMCUB bring-up** — load `navi23_dmcub.bin` (redistributed per `LICENSE.amdgpu`) into its
reserved regions, minimal `dmub_srv` init, verify the caps query answers.
5. **Full owned modeset** — port the dcn30 hwseq slice: OTG timing programming, MPC bypass
(single plane), DIG/PHY enable, **DP link retrain** (or start on HDMI to defer it, exactly
as the NVIDIA doc advises). This is where the SMU/clock question lands — first attempt:
reuse inherited boot clocks for a same-or-lower mode.
6. **EDID** — AUX (DP) / DDC (HDMI) over the DCN AUX engine registers; parse and build the mode
list; connector topology from the VBIOS AtomBIOS data tables.
7. **Wire into the compositor** — `attach_scanout`, vsync from the vblank/pageflip interrupt,
then multihead.
Keep the GOP backend as the fallback the whole way — a stall at any step still leaves danos with
a working display (the resilience v2 already provides via re-attach).
## Reading list
**The DC core (MIT — the register manual for DCN):**
- `drivers/gpu/drm/amd/display/dc/hwss/dcn30/dcn30_hwseq.c` — hardware init + the modeset
sequencer for the target generation (host-driven; one DMUB caps query).
- `dc/dcn30/` + `dc/dcn302/` blocks: `dcn30_hubp.c` (surface address/pitch — milestone 3),
`dcn30_optc.c` (OTG timings), `dcn30_dio_link_encoder.c` (DIG/PHY), `dcn30_mpc.c` (bypass),
`clk_mgr/dcn30/` (the SMU question, read before milestone 5).
- `dc/core/dc.c` — `dc_validate_boot_timing()`: the GOP-takeover readback recipe.
- `asic_reg/dcn/dcn_3_0_0_{offset,sh_mask}.h` — every register name and bitfield.
- `dmub/` (`dmub_srv.h`, `src/dmub_dcn30.c`) — firmware regions + bring-up for milestone 4.
**The Linux glue (for logic, not porting):**
[`amdgpu_dm.c`](https://github.com/torvalds/linux/blob/master/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm.c)
— `dm_init_microcode` / `dm_dmub_hw_init` (the firmware-wall switch), seamless-boot gating.
**Kernel docs (read first):**
[DCN overview](https://docs.kernel.org/gpu/amdgpu/display/dcn-overview.html) (the block diagram
+ VRAM→DF→DCHUB fetch path) ·
[DC programming model](https://docs.kernel.org/next/gpu/amdgpu/display/programming-model-dcn.html)
(dc_plane/dc_stream/dc_link objects, hwseq, block APIs) ·
[display manager](https://docs.kernel.org/6.2/gpu/amdgpu/display/display-manager.html).
**AtomBIOS:** `drivers/gpu/drm/amd/amdgpu/atom.c` (the interpreter), `atombios.h` (table
formats), [osdev AMD AtomBIOS](https://wiki.osdev.org/AMD_Atombios) (hobby-OS orientation),
Haiku [`radeon_hd`](https://github.com/haiku/haiku/tree/master/src/add-ons/kernel/drivers/graphics/radeon_hd)
(a complete worked example, MIT, through DCE 8.5).
**Licensing:** linux-firmware
[`LICENSE.amdgpu`](https://github.com/endlessm/linux-firmware/blob/master/LICENSE.amdgpu);
DCE-era prose specs at [x.org/docs/AMD](https://www.x.org/docs/AMD/) (Register Reference
Guides — display; note `R6xx_3D_Registers.pdf` is 3D-only).
## Open questions (unresolved by the survey)
- **Does DCN 3.0.2 silicon need DMCUB for a bare inherit-and-modeset path**, or only for
PSR/ABM/offloaded features? (Moot if the blob is shipped regardless — but it decides whether
milestone 3 can precede milestone 4.)
- **Can a display-only driver avoid PSP and SMU entirely** by inheriting GOP boot clocks — what
does the dcn30 clock manager actually require from SMU messaging on Navi 23 for a
same-or-lower mode? *The largest open risk in the plan.*
- **Do RDNA2 VBIOS command tables still carry a complete modeset path**, or are they vestigial
init-only tables? (Would open a Haiku-style route on modern cards; no source answers it.)
- **Is the DCN 3.0 AUX/DDC engine and DP retrain fully host-drivable without DMUB**, as
`dcn30_hwseq` implies?
- Exact HUBP surface alignment/pitch constraints for linear scanout on Navi 23 (in the headers;
not captured verbatim in the survey).
---
*Research snapshot (2026-07); findings pinned to Linux master and Haiku master as of the survey
date. DMUB coverage only grows with new DCN generations — re-verify the firmware-wall switch in
`amdgpu_dm.c` against current source before building.*
@@ -0,0 +1,210 @@
# Device interrupts
CPU exceptions ([interrupts.md](../os-development/interrupts.md)) are the kernel reacting to its own
mistakes. **Device interrupts** are the opposite: hardware asking for attention —
a timer firing, a key pressed, a packet arriving. They share the IDT, but differ
in one fundamental way: an exception here is terminal (we report and halt), while a
device interrupt is *handled and returned from*, so the interrupted code resumes as
if nothing happened. This is danos's first code that takes an interrupt and comes
back — the same mechanism a scheduler will later use to preempt tasks.
The first device we bring up is the **timer**, because it's the simplest: it lives
entirely on the CPU's local interrupt controller, needing no external routing.
It's all x86_64-specific, behind the [architecture](../os-development/architecture.md) boundary.
## The APIC, not the PIC
Interrupt delivery on modern x86 goes through the **APIC**, not the legacy 8259
PIC. There are two halves; we only need one so far:
- The **Local APIC** (per-CPU, memory-mapped at physical `0xFEE00000`) handles the
CPU's own timer and receives interrupts routed to it. `system/kernel/architecture/x86_64/apic.zig`.
- The **IO-APIC** routes *external* device lines (keyboard, etc.) to LAPIC vectors.
Not needed for the timer — it'll arrive with the keyboard.
The old PIC has to be dealt with first, though: left alone it would deliver
interrupts on vectors `0x08-0x0F`, which **collide with the CPU exception
vectors** — a spurious IRQ would look like a double fault. So `init` remaps the
PIC's vectors to `0x20-0x2F` and masks every line, taking it out of the picture.
Then the LAPIC is enabled in two places: the `IA32_APIC_BASE` MSR's global-enable
bit, and the LAPIC's own spurious-vector register (bit 8 = software enable). The
spurious vector is `0x2F` — low nibble `F` by convention, and inside our gate
range so a stray spurious interrupt lands on a valid no-op.
## The timer
The LAPIC timer is three register writes (`initTimer`): a divide setting, then the
LVT-timer entry giving it a **vector** (32) and **periodic** mode, then an initial
count that becomes the reload value. From then on it fires vector 32 repeatedly, on
its own, forever.
The reload count isn't picked arbitrarily — it's **calibrated to real time**,
which the [real-time](../vision.md) scheduling guarantees depend on. Since the LAPIC
timer's raw rate is bus-clock dependent and unknown up front, `calibrate` runs the
LAPIC timer one-shot from its maximum count while a **reference clock** counts out a
known 10 ms, then sees how far the LAPIC got — its counts-per-millisecond, from which
`initTimer(hz)` computes the reload count for any target frequency. danos runs it at
**1000 Hz** (a 1 ms tick).
The reference clock is chosen in order of preference, so danos calibrates on
legacy-free **UEFI Class 3** hardware where the old 8254 PIT may be *absent* (polling
a missing PIT would hang the boot):
1. **CPUID leaf 0x15** — the CPU's TSC frequency directly, needing no external timer
at all (the LAPIC is then measured against the TSC).
2. The **HPET**, discovered via ACPI (see [discovery](../os-development/discovery.md) / [acpi](../os-development/acpi.md)).
3. The **ACPI PM timer** (a fixed 3.579545 MHz counter from the FADT).
4. The **PIT** (legacy 8254, 1.193182 MHz) — last resort, and bounded so it can't hang.
All four yield the same rate; on QEMU (no CPUID crystal enumeration) it lands on the
HPET, matching the PIT numbers to within measurement jitter.
## The high-resolution clock (TSC)
The timer tick gives *scheduling* — a 1 ms quantum — but 1 ms is coarse for a
real-time system to *measure* with (interrupt latency, jitter, timeouts). So the
same calibration also measures the **TSC** (Time Stamp Counter): a per-core cycle
counter read with `rdtsc` in a couple of cycles, giving roughly **nanosecond**
resolution — a million times finer than the tick. We snapshot the TSC across the
same 10 ms calibration window to get its frequency (measured ~1 GHz under QEMU).
The monotonic clock is exposed as one function per resolution — `nanos()`,
`micros()`, `millis()` — each scaling the cycle delta directly at its unit (with a
128-bit intermediate so a long uptime doesn't overflow) rather than chaining
divisions. `millis()` is what the scheduler uses for `sleep` deadlines; `nanos()`
is there for fine measurement. Note the two clocks are distinct: the **tick** drives
preemption and wakeups (1 ms granularity); the **TSC** is the resolution you read
time at. Making `sleep` itself sub-millisecond would take a tickless one-shot
timer — a later step.
### Is the TSC trustworthy? Invariant, and synchronized
A cycle counter is only a valid *clock* if two things hold, and danos checks both,
because they decide whether we read time with a cheap `rdtsc` or fall back to the HPET.
**Invariant.** An old TSC counted core clock cycles, so it sped up and slowed down with
frequency scaling — useless as wall time. Modern CPUs (all of danos's targets) provide an
**invariant TSC**: a constant rate across P/C-states that never stops. The guarantee is a
CPUID bit — leaf `0x80000007`, EDX bit 8 — on both Intel *and* AMD. danos reads it in
`calibrate`, and a TSC that doesn't advertise it is demoted to the HPET clocksource —
provided a usable HPET exists (64-bit; a 32-bit one wraps too fast to stay monotonic).
With no such fallback the TSC stays, there being nothing steadier to switch to. AMD is
why this matters in practice: it doesn't populate the Intel leaf `0x15` that enumerates
the TSC *frequency*, so danos already measures AMD's rate against the HPET — but a
measured frequency without the invariance guarantee is not enough.
**Synchronized.** Each core has its own TSC. Even invariant ones can start at different
values (a second socket, some firmware), so a thread migrating from a core reading
`1_000_000` to one reading `999_000` would see time jump *backward*. danos runs a **warp
check** as each application processor comes online (`checkWarpSource`, adapted from
Linux's): the waking core and the BSP hammer a shared "highest seen" TSC under a lock,
and if either ever reads below it, the cores' TSCs are skewed. It's pairwise because APs
come up one at a time ([smp.md](../os-development/smp.md)).
**The fallback.** When the TSC fails either test — non-invariant (a bare VM such as the
default qemu64), or warped between cores — danos moves the monotonic clock onto the
**HPET** main counter: one fixed-rate counter, so it can neither skew between cores nor
drift with frequency. It costs a memory-mapped read instead of a register read, but it
keeps time *accurate*, which is the whole point. The switch preserves the current value,
so the clock never jumps. The boot log names the outcome:
```
/system/kernel: clocksource tsc (TSC invariant: yes, synchronized: yes) # real Intel/AMD
/system/kernel: clocksource hpet (TSC invariant: no, synchronized: yes) # a bare VM (TCG)
```
## Two kinds of vector, one dispatch
The IDT now installs gates `0-47`: the 32 exceptions plus the device range. Every
gate still funnels through the same stub tail (`isr_common`), which calls one
dispatcher that branches on the vector (`interruptDispatch` in `idt.zig`):
```zig
if (state.vector < 32) {
on_fault(state); // exception: report and halt (never returns)
} else if (handlers[state.vector]) |handler| {
handler(); // device: run the registered handler
}
// else: spurious/unhandled — deliberately no EOI
```
(A third branch has since joined for user mode, elided here:
`state.vector == system_call_vector` (128) hands the trap frame to the ring-3
syscall handler.)
Two things make device interrupts *return* where exceptions don't:
1. **The handler returns.** The timer handler just bumps a tick counter. Control
flows back to `isr_common`, which restores every register it saved and executes
`iretq` — resuming the interrupted instruction exactly. (This is why the stub
saves *all* the general registers.)
2. **End-of-interrupt.** Somewhere in there we write the LAPIC's EOI register. Miss
this and the LAPIC thinks we're still busy and never delivers the next
interrupt. It's the single most common "my timer fired once and stopped" bug.
**Each handler issues its own EOI**, rather than the dispatcher doing it around the
call. That looks like a needless devolution while the timer is the only device, and
`apic.timerTick` indeed does nothing but `eoi()` before bumping its counter (early,
because the tick hook is the scheduler, which may switch tasks and not return
promptly — the LAPIC mustn't wait on it).
It stops looking needless with the second device. A *routed* interrupt — one arriving
through the I/O APIC from a real device line — must be **masked before it is
acknowledged**, because a level-triggered line is still asserted at EOI time and would
redeliver instantly, forever. Only the handler knows which discipline its source
needs, so only the handler can sequence it. See [drivers.md](drivers.md), where the
device is quieted by a driver in ring 3, long after the ISR has returned.
A device handler is a plain `fn () void` — a timer or keyboard handler doesn't need
the interrupted registers. (The stubs originally didn't save the SSE/vector
registers, so a handler couldn't use them; `isr_common` now does an
`fxsave`/`fxrstor` of the full SSE/x87 state around dispatch — see
[interrupts.md](../os-development/interrupts.md).)
## Turning them on
Exceptions can't be masked, which is why they worked all along. Maskable device
interrupts don't fire until the CPU's interrupt flag is set — so the final step is
`sti` (`arch.enableInterrupts()`), after the APIC and timer are configured. From
that instant the kernel has a heartbeat, and its idle `hlt` loop
([halting.md](../os-development/halting.md)) wakes on every tick and dozes off again.
## Verifying it
The `timer` test (see [testing.md](../testing.md)) is the proof that an interrupt both
*fires* and *returns*: it records the tick count, busy-waits, and checks the count
advanced on its own.
```
$ python3 test/qemu_test.py timer
timer ... PASS (matched 'DANOS-TEST-RESULT: PASS')
```
If the APIC weren't enabled, or `sti` were missing, or EOI were forgotten, the
count would stay put and the test would fail. That it advances — while the CPU was
spinning in unrelated code — is the whole mechanism working end to end.
## Since (done elsewhere)
- **Preemption**: the timer handler is where the scheduler decides to switch — the
reason a *returning* interrupt matters. See [scheduling.md](../os-development/scheduling.md).
- **`sleep()` / timeouts** built on the calibrated clock.
- **The I/O APIC, routed**: external device lines now reach a vector, and the
interrupt is delivered onward to a *user-space* driver as an IPC message. See
[drivers.md](drivers.md).
- **Uncacheable MMIO**: device grants are mapped `PCD|PWT` (strong-uncacheable) for
user drivers — see [paging.md](../os-development/paging.md).
## What's next (partly done since)
- **The keyboard** — done, exactly as sketched: the PS/2 bus driver
(`system/drivers/ps2-bus/`) claims the port-mapped 8042 controller through the
claim-gated `io_read`/`io_write` syscalls ([drivers.md](drivers.md)), binds
IRQ 1 (and the aux mouse's IRQ 12), reads scancodes from `0x60`, and decodes
them into HID events for the [input service](input.md).
- **MSI-X** — still open: `msi_bind` gives one per-device edge-triggered vector
(M15); MSI-X's multi-vector table (many queues per device, e.g. NVMe) is the
remaining extension.
- **The LAPIC's own page** — still mapped writeback-cacheable like the rest of
the identity map. QEMU tolerates it; real hardware wants it uncacheable.
@@ -0,0 +1,177 @@
# The device manager
**Status: the protocol and supervision are built** (M18.1, 2026-07-13): `hello`
with its deadline, supervised spawn, restart with backoff, and the crash-loop
cap are in — usb-xhci-bus is the first conforming driver, and the
`driver-restart` scenario proves fault → backoff → re-claim → cap end to end.
Tree reports are built too (M18.2, 2026-07-13): the xHCI driver scans its
root-hub ports and reports each connected device (`child_added`); the manager
mirrors them and prunes a dead reporter's children, and the `usb-report`
scenario proves report → prune → respawn → re-report. The application surface is built (M18.3, 2026-07-13):
`enumerate` and `subscribe` over IPC, with `device-list` as the first client —
the manager is now the one answer to "what devices exist" for applications.
The primitives underneath are real ([process-management.md](../os-development/process-management.md):
spawn/supervise/kill/exit-notification; [driver-model.md](driver-model.md): the device
table as a capability system; [drivers.md](drivers.md): claim/map/IRQ), and the first
per-device driver spawn works (the device manager matches the xHCI controller by PCI
class and spawns `usb-xhci-bus` with the device id as argv[1]). This document designs
the rest: the device manager as **the tree, the matcher, and the supervisor** — the
policy process that turns [resilience.md](../os-development/resilience.md)'s restart goal into practice
for drivers.
How processes stop, reload, and report their deaths is deliberately **not** in this
document: that is the universal lifecycle every danos process speaks —
[process-lifecycle.md](../os-development/process-lifecycle.md), signals over IPC and the stable
`process` interface. The device manager is that design's first serious
customer, not its owner. Its own protocol contains nothing lifecycle-shaped; a
driver is stopped, health-checked, and buried exactly like any other process.
## The tree: structure in the manager, authority in the kernel
The device tree is two things fused: *information* (what exists, how it nests) and
*authority* (a descriptor is a licence to map physical memory). They separate:
- The **kernel keeps the capability system** — device, I/O-port, and interrupt
claims, resource containment on `device_register`, the
`mmio_map`/`irq_bind`/`msi_bind` gates — and **cleans all of it up when a process
dies** (settled; it is increment 1 of
[process-lifecycle.md](../os-development/process-lifecycle.md)). The three invariants in
[driver-model.md](driver-model.md) stay exactly where they are. A device manager
that could mint MMIO mappings by its own say-so would be a second kernel, and a
buggy one would un-earn everything the microkernel bought.
- The **device manager owns the tree as data** — identity, topology, naming, driver
matching, hotplug events, and being the one process everything else asks about
devices. Firmware discovery seeds it (today via the kernel's snapshot); **bus
drivers grow it** by reporting what they see; applications query and watch it.
`device_enumerate` fades to a manager-internal (then deleted) seam.
Long-term, discovery itself leaves the kernel — but not *into* the manager. PCI
enumeration is a **pci-bus driver**: the manager spawns it against the host bridge
(already a device with the ECAM window as a resource), it scans, it reports functions
like any bus reports children. ACPI becomes an **acpi service** that interprets the
tables and reports the namespace. The manager only orchestrates and merges. Moving
AML interpretation out of ring 0 is its own project on its own track; nothing here
depends on when it lands. (It landed: [discovery.md](../os-development/discovery.md), M19–M20.)
`device_register` is **idempotent on exact match**: a re-registration with an
identical (parent, class, identity, resources) tuple returns the existing id
instead of appending a duplicate. The kernel table has no unregister, so without
this a restarted registering bus would re-report its children as fresh nodes on
every respawn. Idempotence is what makes restart-and-re-report sound for *every*
reporting bus — pci-bus, the acpi service, a future fdt service — not just one,
and it is why supervision (below) can prune a dead bus's subtree and trust the
restarted instance to rebuild exactly the same ids.
## The protocol
A `device-manager-protocol` module (the vfs-protocol pattern): extern-struct
messages, a version in the handshake, reserved fields everywhere. The manager is a
well-known endpoint (`ipc.register(.device_manager)`); the badge tells it who is
talking; the same endpoint receives its children's exit notifications — one loop,
one world.
| Direction | Message | Purpose |
|---|---|---|
| driver → manager | `hello { version, role, device_id }` | confirms the argv assignment, starts the deadline clock |
| bus → manager | `child_added { parent, bus_address, identity, device_id, hid }` | one node the bus discovered |
| bus → manager | `child_removed { parent, bus_address }` | unplug, or the bus lost it |
| app → manager | `enumerate` | snapshot of the tree (read-only) |
| app → manager | `subscribe` | receive published add/remove events |
`hello` is the one deadline the manager enforces itself: spawned and silent past the
deadline means wrong binary, wrong protocol version, or wedged before main — apply
the stop sequence and the restart policy. Everything else lifecycle-shaped
(terminate, the common `ping` liveness call, exit reasons) arrives through
[process-lifecycle.md](../os-development/process-lifecycle.md)'s vocabulary, not this protocol.
Assignment stays argv (`usb-xhci-bus <device id>`) for now — simple, and it works.
The step after `hello` exists is delegation: the manager claims (or is granted) the
devices and passes the claim to the driver over IPC (the M13 capability-transfer
mechanism), replacing first-come-first-served `device_claim` with policy. Identity in
`child_added` is per-bus: PCI children carry the class triple (`pci_class`, as the
xHCI match already uses); USB children carry the (class, subclass, protocol) triple
from usb-ids.zig — each bus's native language, decoded by the shared ids modules.
## Supervision and restart
Every driver is spawned with the manager's exit endpoint (`spawnSupervised` — built).
On a death notification:
1. **Read the reason** ([process-lifecycle.md](../os-development/process-lifecycle.md) increment 2).
Clean exit → it meant to; don't restart. Fault or missed `hello` deadline →
restart with **backoff**, and a crash-loop cap (three fast deaths → mark failed,
stop respawning, log loudly; a later `reload` to the manager can retry).
2. **Prune the subtree** the dead bus driver reported. Its children describe
protocol state (xHCI slot ids, transfer rings) that died with the process;
keeping the nodes would be keeping a lie. Watchers receive `child_removed` — the
input service losing, then regaining, a keyboard is the *honest* description of
what happened. The restarted instance rediscovers and re-reports.
3. **The claim is already free** because the kernel released it at death — the
restarted instance claims the same controller and comes up.
Who supervises the supervisor: **init** (PID 1), which already supervises the
services it starts. If the manager dies, drivers keep running (they hold their
claims; the kernel doesn't care who their supervisor was — though their exit
notifications now dangle harmlessly). The restarted manager re-learns the world:
kernel snapshot, then a re-`hello` round — drivers answer a broadcast or are stopped
and respawned. Full state handoff is deliberately not attempted.
## Thin drivers, class protocols
The [driver-model.md](driver-model.md) three-shape split, restated as processes:
- A **bus driver** (usb-xhci-bus) owns its controller — claim, MMIO, IRQ/MSI, DMA
rings — and offers a *transfer* protocol ("submit a control transfer to device N",
built from the usb-abi request constructors) plus tree reports to the manager.
- A **class driver** (usb-hid, usb-storage) owns nothing: it is matched to a reported
child by its identity triple, speaks the bus's transfer protocol downward and its
service's protocol upward — HID reports to the input service, blocks to the block
service. It works unchanged over any controller.
- **Services** (input, display, block) aggregate class drivers and face applications.
Each arrow is a protocol module. The manager routes none of the data plane — it
introduces the parties (matching), supervises them (lifecycle), and gets out of the
way.
## Increments
Increments 1–4 are the lifecycle prerequisites and live in
[process-lifecycle.md](../os-development/process-lifecycle.md) (claim cleanup on death, exit reasons,
published exit events, signals + `process`). On top of those:
5. **device-manager-protocol**: `hello`, supervised spawn with restart policy;
usb-xhci-bus becomes the first conforming driver.
6. **Tree reports**: `child_added`/`child_removed`; the manager mirrors; xHCI reports
the mouse and keyboard QEMU already hangs off it.
7. **App surface**: `enumerate`/`subscribe` over IPC; `device_enumerate` retreats
to a manager-internal seam.
8. **Discovery migration** — DONE (M19–M20, 2026-07-13): enumeration moved to
ring 3 as swappable per-firmware discoverers — the pci-bus driver (M19) then
the acpi service (M20), see [discovery.md](../os-development/discovery.md); of the enumerable
devices, the kernel seeds only the host bridge and the acpi-tables node (the
non-enumerable platform nodes — processors, interrupt controllers, the HPET,
the loader's framebuffer — stay kernel-seeded too). Matching moved with it:
`child_added` grew a `device_id` (the kernel-registered id, `no_device` for
unregistered leaves like USB ports) and a firmware `hid`, and the manager now
matches drivers from those **reports** rather than its boot-time snapshot. The
PCI arm flipped in M19.3, the ACPI arm (ps2-bus matched from `_HID`) in M20.3
— each in a single phase so no device is ever matched from both sources at
once. The acpi service reports only the non-PCI `_HID` devices, since pci-bus
already reports PCI functions (M20.2).
## Settled questions (2026-07-12)
- **Stateful buses**: pruning the subtree on bus-driver death is right for USB. A
future storage bus with in-flight writes wants drain-before-terminate — which is
exactly the `deadline_ms` parameter `stop()` already has; a per-driver deadline
is one value in the manager's policy table when such a bus arrives. No design
change.
- **Manager death**: drivers survive the manager; the restarted manager re-learns
the world (above). Checkpointing driver state with the manager is deferred until
something demonstrates the need.
- **Matching stays code until the third bus.** `driverFor`/`pciDriverFor` were
honest at two bus types; the third was expected to trigger the manifest (a driver
declares what it binds: a PCI class triple, a USB class triple, an ACPI `_HID`).
(Since then: the third bus — USB — arrived and is matched in code too. Today's
matchers are `pciDriverForIdentity`, `hidDriverFor`, and `usbDriverForIdentity`;
the manifest waits until code matching actually hurts.)
@@ -0,0 +1,182 @@
# Display service — build plan (v1: the dumb-framebuffer compositor)
The ordered, checkpointable build-out for [display.md](display.md). Each milestone is
small, lands on its own, and ends in a **verifiable gate** — shaped for a `/loop` run.
Read [display.md](display.md) first for the *why*; this is the *what* and the *order*.
## Locked decisions (do not relitigate)
- **Handoff = device node + write-combining `mmio_map`.** The kernel seeds a synthetic
display node (found by class, not name) from `BootInformation.framebuffer`; the
service claims + WC-maps it.
(Not a bespoke `framebuffer_map` syscall — the device route inherits ownership,
release-on-death, and re-claim-on-restart.)
- **v1 = the full compositor pipeline on the dumb framebuffer.** One `display` service
owns the LFB + a cacheable back buffer + a layer stack; double-buffer + damage-driven
present; clients draw via server-side commands. **No** runtime mode-setting, **no**
shared-memory surfaces — both deferred (see display.md, "What v1 does not do").
## Conventions
Follow [coding-standards.md](../coding-standards.md): spell out non-acronym abbreviations in
full, kebab-case file names, no `Co-Authored-By` trailers on commits. New user binaries
go through `addUserBinary` in [build.zig](../../build.zig) and get packed into the
initial-ramdisk; protocols are `b.addModule("…-protocol", …)` and imported into the
`runtime` module.
## How to verify along the way
- `zig build test` — host unit tests (compositor math: layer clipping, damage merge,
pitch/format blits are all host-testable with a fake framebuffer).
- `python3 test/qemu_test.py <case>` — boots the real kernel in QEMU; assert on the
serial log ([tests.zig](../../system/kernel/tests.zig) is the registry).
- The `run-efi` target renders to QEMU's display (`-device VGA,edid=on,xres=1280,yres=720`)
— a screenshot confirms pixels for the milestones whose gate is visual.
---
## D1 — The handoff primitive (kernel) ✅
Make the boot framebuffer reachable and mappable **write-combining** from user space.
- [x] [device-abi.zig](../../library/device/model/device-abi.zig): added `DeviceClass.display`; a
`DisplayInfo{ width, height, pitch, format }` carried on the descriptor; a
`flags` field on `ResourceDescriptor` + `resource_flag_write_combining`.
- [x] [devices-broker.zig](../../system/kernel/devices-broker.zig): `seedDisplay(base, w, h,
pitch, format)` publishes a root-level `display` node with one WC-flagged `memory`
resource `[base, height*pitch]` + the `DisplayInfo`; `displayDevice()` /
`displayClaimed()`. Seeded from `kmain` after `devices_broker.init`.
- [x] [process.zig](../../system/kernel/process.zig) `systemMmioMap` + paging
(`mapUserDeviceInto` gains a `write_combining` bool): a resource's WC flag maps it
through the WC PAT slot (`setupPat`) instead of strong-uncacheable.
- [x] [console.zig](../../system/kernel/console.zig): `setSuppressed` quiesces `write` while
the display device is claimed (driven from `systemDeviceClaim` / release); the
terminal panic + exception paths clear it first so a dying machine still draws.
**Gate (met, automated):** the `display` kernel test (`python3 test/qemu_test.py display`,
`displayTest` in [tests.zig](../../system/kernel/tests.zig)) asserts the seeded node's shape
and geometry, then walks the real claim + `mmio_map` path into a throwaway address space
and verifies the leaf is **write-combining** (PAT entry 4: PAT bit set, PCD/PWT clear) —
with an uncacheable-still-uncacheable regression guard. Chosen over the original
screenshot-of-a-fill gate because it proves the *actual* WC property headlessly; the
visible fill folds into D2's gate (the service clears the screen through the back buffer).
Regression-checked: `discovery`, `ioport`, `claim-release`, `supervision`, `device-list`,
`device-manager` all still pass with the +1 device in the table.
## D2 — Service skeleton, protocol, runtime module ✅
Stand up the named service and the double-buffer, no layers yet.
- [x] `library/protocol/display/display-protocol.zig`: `Operation{ info, create_layer,
configure_layer, destroy_layer, fill_rect, blit_tile, damage, present }`; `extern`
`Request`/`Reply`; size + `maximum_payload` consts. (Model: block/protocol.zig.)
- [x] [abi.zig](../../system/abi.zig): `ServiceId.display = 9`.
- [x] `system/services/display/display.zig`: `main` → enumerate + claim + WC-map the LFB
(front) → `mmap` a cacheable back buffer of `height*pitch` → `runtime.service.run`.
`info` and a whole-screen `present` (back → front) are live; layer ops fail-stub
until D3. Init clears the back buffer and presents it — the double-buffer path.
- [x] [library/runtime/display.zig](../library/runtime/runtime.zig) (+ barrel export of
`display` and `display_protocol`): `info()` and `present()`, cached `.display`
lookup with retry (model: block.zig).
- [x] [init.zig](../../system/services/init/init.zig): `"display"` added to `boot_services`.
- [x] [build.zig](../../build.zig): `display-protocol` module on the runtime; `display` exe
via `addUserBinary`; packed into the initial-ramdisk; installed to
`/system/services/display`.
- [x] **Kernel fix the back buffer surfaced:** `mmap` was capped at 256 pages (1 MiB) by
a fixed kernel-stack `frames` array. Rewrote `systemMmap` to map page-by-page with
rollback (no scratch array) and raised the cap to 8192 pages (32 MiB) — enough for a
4K back buffer. A real limitation met, exactly the kind this project chases.
**Gate (met, automated):** `python3 test/qemu_test.py display-service` spawns the
compositor and matches its own serial heartbeats — `display: online {w}x{h} pitch …`
followed by `display: presented frame 0` — which it prints only after the whole
claim → WC-map → back-buffer → clear → present chain succeeds (matched on serial like the
fault cases, since a lone blocking service can't reschedule the in-kernel test context to
poll). Regression-checked: `usermem`, `heap` (the `mmap` rewrite), `init` (the boot-list
addition), and D1's `display` all still pass.
## D3 — Layer stack + compositor + damage present ✅
The heart: composite an ordered layer stack, present only what changed.
- [x] A layer table (16 slots): each `Layer` = position, z, visible, a server-owned
`mmap`'d surface (freed on `destroy_layer`). `damage` accumulates the dirty screen
region since the last present.
- [x] `create_layer` / `configure_layer` (damages old + new footprints) / `destroy_layer`,
`fill_rect`, `blit_tile` (reads the inline tile from the IPC payload, unaligned-safe),
`damage`, `present`.
- [x] Pure, host-tested [compositor.zig](../../system/services/display/compositor.zig): `Rect`
(intersect/unite), `Surface`, `fillRect`, `composite` (opaque, clipped to a damage
rect), `blitTile`. `present` clears the damaged region to the wallpaper, paints the
visible layers bottom-to-top (z-sorted), and flushes just that rect back → front (WC).
Colour packing (rgbx/bgrx) is `protocol.pack`, also host-tested.
- [x] Host tests (`zig build test`, green): rect intersect/unite, `fillRect` clipping +
`stride > width` padding, `composite` overlap-shows-top + damage clipping, `blitTile`
unaligned read + clipping, and `pack` for both pixel formats.
**Gate (met):** `zig build test` green for the compositor + pack unit tests, **and** the
`display-service` case's startup self-check composites two overlapping layers on the real
framebuffer and reads back the composited pixels — overlap = top layer, outside = bottom
layer — logging `display: compositor self-check ok` (matched by the harness).
## D4 — Client API + the demo client ✅
Prove the pipeline end-to-end from a separate process.
- [x] Finished [runtime/display.zig](../library/runtime/runtime.zig): a `Layer` handle with
`fill` / `blitTile` (inline tile) / `configure` (move/restack/show) / `damage` /
`destroy`, `createLayer`, and a `color(r,g,b)` helper (caches the mode, packs via
`protocol.pack`). Coordinates are signed over the wire (`@bitCast` both ways).
- [x] `system/services/display-demo/`: a hardware-free client (the `input-source` analog)
— a full-screen wallpaper layer, a rectangle that slides back and forth (moved by
`configure` each frame, so the compositor repaints old + new), and a cursor layer;
presents in a loop paced by `runtime.time`. Wired into build + initial-ramdisk.
- [x] **Bug this surfaced:** `protocol.message_maximum` was 4096, but the kernel caps
every IPC message at `MESSAGE_MAXIMUM` = 256 — so `replyWait` rejected the oversized
receive buffer with `-E2BIG` and the serve loop had been *spinning* since D2 (unseen,
as D2/D3 matched init-time heartbeats). Set it to 256; `blit_tile` is now explicitly
a small-tile path (≤ 54 px inline), larger bitmaps being the deferred shared-memory surface.
**Gate (met):** `python3 test/qemu_test.py display-demo` spawns the service + `display-demo`;
the demo drives a run of frames of motion through the layer client API and logs
`display-demo: ok` (the visible motion is a screenshot via `zig build run-x86-64`).
Regression-checked: `zig build test`, `display` (D1), and `display-service` (D2/D3) all
still pass, and the default `zig build` is clean.
## D5 — Test cases + docs ✅
- [x] The three integration cases exist and pass: `display` (D1 handoff, kernel),
`display-service` (D2/D3 compositor + self-check), and `display-demo` (D4 full
pipeline: spawn `display` + `display-demo`, match `display-demo: ok`) —
[tests.zig](../../system/kernel/tests.zig) + [qemu_test.py](../../test/qemu_test.py). Plus
the pure host tests (`zig build test`).
- [x] [display.md](display.md) updated to the built state (the "Verifying it" section names
the real cases); [README index](../README.md) entry present (#19); the `display-track`
memory marked DONE with the commits.
**Gate (met):** `python3 test/qemu_test.py display display-service display-demo` all pass,
`zig build test` is green, and the default `zig build` is clean.
---
## v1 status: complete
D1–D5 done. The display service is a working framebuffer compositor: it owns the
framebuffer (write-combining), composites a z-ordered layer stack into a cacheable back
buffer, presents only the damaged region, and is driven over IPC by the `runtime.display`
client — proven end-to-end by a separate demo process. Two limitations are deliberate and
documented (docs/display.md): no runtime mode-setting (native backend) and no true vsync
(no vblank on a dumb framebuffer). Next steps are the Deferred items below.
---
## Deferred (explicitly not in this plan)
- **Shared-memory surfaces** — generalize M13 capability passing to memory objects
(`shared_memory_create`/`shared_memory_map`), so bitmap clients hand the compositor a rendered surface
instead of drawing commands. The compositor's layer model already anticipates it.
- **Native backend (Bochs DISPI, then virtio-gpu)** — behind the same internal backend
interface as the dumb framebuffer: EDID mode list + runtime resolution/bpp change +
(eventually) a vblank/flip path for true vsync.
- **Driver/compositor process split** — only when a second backend or a second head makes
the abstraction pay for itself.
@@ -0,0 +1,175 @@
# Display v2 — build plan (pluggable scanout: GOP floor + virtio-gpu native)
The ordered, checkpointable build-out for [display-v2.md](display-v2.md). Each milestone
lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run, like
[display-plan.md](display-plan.md). Read display-v2.md first for the *why*.
## Locked decisions (do not relitigate)
- **First native backend = virtio-gpu** (VM standard: mode-set + fenced present/flush).
- **Dynamic hot-attach**: boot on GOP, upgrade to native when the driver **announces**
(push, not polling); re-attach across driver restarts; GOP is the floor for "no driver
ever," not a live fall-back after a reprogram.
- **v2 builds the shared-memory capability** (endpoints → memory objects), shared with the future
client-surface path.
- The compositor's layers/back-buffer/damage are **unchanged**; only scanout is pluggable.
## Conventions
Follow [coding-standards.md](../coding-standards.md): spell out non-acronym abbreviations,
kebab-case file names, no `Co-Authored-By` trailers. New user binaries go through
`addUserBinary` and get packed into the initial-ramdisk; protocols are
`b.addModule("…-protocol", …)` imported into `runtime`; new syscalls extend
[abi.zig](../../system/abi.zig) `SystemCall` + a `library/runtime` wrapper.
## How to verify along the way
**Every gate is serial-checkable — no screenshots** (this plan is built to run unattended).
Where "does it actually display" would otherwise need a human eyeball, the code **reads its
own pixels back**: the scanout resource is CPU-visible RAM (shared-memory-backed) and the back buffer
is cacheable, so a driver/compositor can write a known value, read it back, and log a
pass/fail — and a virtio `resource_flush` is confirmed by the device **acking it on the
used ring**. Those two together (pixel-readback + flush-ack) are the automated stand-in for
"it's on screen."
- `zig build test` — host unit tests (backend selection, virtio struct sizes/encodings,
pixel-check helpers).
- `python3 test/qemu_test.py <case>` — boots the kernel in QEMU; asserts on serial markers.
The virtio cases boot with `-device virtio-gpu` (a per-case `qemu_extra`).
- `run-x86-64` renders to a window — for the human's own satisfaction, **not** a gate.
---
## V1 — The scanout backend seam (refactor, no behaviour change) ✅
Extract scanout from the compositor so today's path becomes one backend among future ones.
- [x] `system/services/display/backend.zig`: a `Backend` tagged union with `info()`,
`surface()` (the cacheable compose target), `present(damage)`, and capability flags
(`canModeSet`/`hasFencedPresent`, both false for GOP).
- [x] The v1 GOP path is now `backend.Gop` (claims the `display` node, WC-maps the LFB,
keeps the cacheable back buffer, `present` = the damage-rect WC copy). display.zig
composes into `backend.surface()` and calls `backend.present(damage)` — no LFB or
framebuffer geometry left in the compositor core.
- [x] The selection decision is the pure `chooseKind(native_available)` (gop unless a
native driver announced), split from the syscall-bound `select()`/`Gop.init()`.
**Gate (met):** `display-service` + `display-demo` pass **unchanged** (pure refactor; GOP
is the only backend), and `zig build test` stays green.
## V2 — The shared-memory cross-process capability (kernel) ✅
- [x] [abi.zig](../../system/abi.zig): `shared_memory_create` (34) / `shared_memory_map` (35) syscalls + a
`shared_memory_test` service id. Handlers in process.zig: `shared_memory_create(len)` allocates contiguous,
zeroed, **cacheable** frames, wraps them in a refcounted object, installs a capability
handle, maps them into the caller's shared-memory arena → returns virtual_address + handle; `shared_memory_map(cap)`
maps the same physical pages into the receiver. Reclaimed on death (see below).
- [x] The capability core (ipc-synchronous.zig) is now **kind-tagged**: `scheduler.Task`'s
handle table holds `HandleObject{kind, ptr}`; `closeHandles` and `shareCapability`
dispatch by kind, so a `SharedMemoryObject` rides an `ipc_call` `send_cap` exactly like an
endpoint and frees only when its last capability drops. `mapUserSharedInto` (paging)
maps WB-cacheable + `device_grant`, so a sharer's teardown never frees the shared
frames — the object owns them.
- [x] `library/runtime/shared-memory.zig` (+ barrel export): `create(len) -> Region{ptr, handle, len}`,
`map(handle) -> ptr`.
**Gate (met):** `python3 test/qemu_test.py shared-memory` — `shared-memory-client` creates a region, writes a
pattern, and passes its capability to `shared-memory-server` as an `ipc_call` send_cap; the server
`shared_memory_map`s it and reads the **same bytes** back → `shared-memory: shared 4096 bytes ok`. Guardrail:
`ipc`/`ipc-call`/`ipc-cap`, `supervision`, `dma`, `usermem`, `display-service`, and host
tests all still pass — the handle-table change broke no existing IPC.
## V3 — The virtio-gpu driver: bring-up + a frame on screen ✅
- [x] `system/drivers/virtio-gpu/`: claim the virtio-gpu PCI function (device-manager
match on the display/other class triple, driver self-confirms vendor 0x1AF4/device
0x1050 from config space), enable memory-space + bus-master, walk the vendor
capabilities in config space to find common-config + notify, map the BAR, negotiate
VERSION_1, and stand up the control virtqueue in coherent DMA. `virtio-gpu-protocol.zig`
+ `virtio-pci.zig` for the control/transport structs (host-tested sizes).
- [x] Create a 2D scanout resource backed by a coherent DMA region (V4 swaps this for the
shared-memory surface), `attach_backing`, `set_scanout` to scanout 0, `transfer_to_host_2d`
+ `resource_flush` of a test pattern, and wait on the used ring.
- [x] Register a `scanout` service (`ServiceId.scanout` = 11).
**Gate (met):** the `virtio-gpu` case (QEMU `-device virtio-gpu-pci`) boots the
device-manager stack, which discovers the function and spawns the driver; the driver writes
a known test pattern into the scanout backing, `transfer_to_host_2d` + `resource_flush`es
it, and **waits for the device's used-ring ack**, then reads the backing back and checks the
pattern — logging `virtio-gpu: scanout 640x480 online` and `virtio-gpu: flush acked, pixel
check ok`. That proves virtqueue + resource + attach + set_scanout + transfer + flush end to
end without a screenshot (the used-ring ack is the device confirming it consumed the frame).
## V4 — The native backend + hot-attach ✅
- [x] `backend.VirtioGpu` in the compositor: `surface()` = the shared-memory scanout surface
(the compositor composes straight into the device's resource backing; x86 DMA is
coherent, so the cacheable shared pages need no flush), `present(damage)` = a `present`
request over the driver's `.scanout` endpoint (→ transfer-to-host + resource flush).
- [x] The driver **announces** to `.display` after bring-up (looks it up with a bounded retry,
sends `attach_scanout` with the geometry + the shared surface as an `ipc_call` send_cap).
The compositor maps it, looks up `.scanout` itself (no need to pass the endpoint — the
driver registered it), switches backend, and re-composites the current frame full-screen.
The present is deferred to a one-shot timer so it runs *after* the reply unblocks the
driver and it serves `.scanout` — presenting inline would deadlock.
- [x] Boot still starts on `backend.Gop`; the upgrade happens on announce. `shared_memory_physical` (a
new syscall) gives the driver the guest-physical of the shared surface for `attach_backing`.
**Gate (met):** the `display-native` case (QEMU `-device virtio-gpu-pci`, `mem` bumped since it
boots the whole system) starts the compositor + `display-demo` + device-manager; the driver
announces, the compositor logs `display: scanout upgraded to virtio-gpu`, drives frames through
the native backend, and **reads a pixel back** from the shared surface after a present to
confirm the composited frame landed (`display: native present verified`), while `display-demo:
ok` still fires — checked order-independently. Without `-device virtio-gpu-pci` nothing is
announced and it stays on GOP: the v1 `display-service`/`display-demo` gates pass unchanged.
## V5 — Mode-setting, EDID, and fenced presents ✅
- [x] The driver negotiates `VIRTIO_GPU_F_EDID` (when offered) and reads the monitor's EDID,
logging its preferred mode; it offers a small mode list over `.scanout` `get_modes`. The
resource + shared surface are sized to the largest mode, so `set_mode` just re-points the
scanout rectangle (no resource/surface churn) — a runtime resolution change. `runtime.display`
gains `modes()` / `setMode()` (display-protocol `get_modes`/`set_mode`, forwarded to the backend).
- [x] Every `resource_flush` is issued fenced (`VIRTIO_GPU_FLAG_FENCE`); the device signals the
fence when it has consumed the frame, which the used-ring ack the synchronous present waits
on already gates — a tear-free present. (Completion feedback, **not vblank**: base
virtio-gpu 2D has no display-refresh event, so nothing paces presents to the monitor —
see the "Fenced is not vsync" note in [display-v2.md](display-v2.md).)
- [x] `backend.VirtioGpu` reports `canModeSet` / `hasFencedPresent` = true.
**Gate (met):** the `display-modeset` case (reusing the display-native boot) upgrades to
virtio-gpu, queries the driver's modes, `setMode`s to a different resolution, and confirms the
change by reading the backend's geometry back (`display: mode set to {w}x{h}, verified`); the
fenced present path is exercised and confirmed (`display: fenced present ok`) — both from serial,
passing 3/3. The driver also logs the EDID preferred mode (`virtio-gpu: EDID preferred mode …`).
## V6 — Resilience (restart + re-attach) + tests + docs ✅
- [x] The virtio-gpu driver now **hellos** the device manager (role: bus) so it is properly
supervised — no longer stopped at the hello deadline — and is restarted on death. On
driver loss the compositor keeps the last frame (its `.scanout` calls now return
`-EPEER` instead of hanging — a kernel fix: an endpoint is marked dead when its owner
dies) and **re-attaches** when the restarted driver re-announces. A permanent give-up
(crash-loop cap) leaves the frozen frame; GOP is not re-taken.
- [x] `test/qemu_test.py`: the `virtio-gpu`, `display-native` (hot-attach), `display-modeset`,
and `display-reattach` (driver-kill/re-attach) cases. display-v2.md status updated.
**Gate (met):** the `display-reattach` case — device-manager (in `test-scanout-restart` mode)
kills the virtio-gpu driver once after it hellos; the restart policy respawns it, it
re-announces, and the compositor logs `display: scanout re-attached` after the initial
`display: scanout upgraded to virtio-gpu`, with no CPU exception / panic (the compositor
survives) — passing 3/3. All v1 + v2 cases (host tests, `ipc`/`ipc-call`/`ipc-cap`,
`supervision`, `shared-memory`, `display-service`, `display-demo`, `virtio-gpu`, `display-native`,
`display-modeset`) pass; default `zig build` is clean.
---
## Deferred (explicitly not in this plan)
- **Client-rendered surfaces** — now unblocked by the shared-memory capability (V2): an app renders
its own bitmap and hands the compositor a reference. A natural follow-on.
- **Bochs DISPI backend** — a simpler second native backend (mode-set only, dumb scanout);
slots behind the same interface if wanted.
- **Real-GPU (NVIDIA/AMD/Intel) drivers** — out of scope; those devices stay on the GOP
floor by design.
- **Hardware-accelerated compositing / multiple heads** — future.
@@ -0,0 +1,147 @@
# The display service v2: a pluggable scanout backend
**Status: complete (V1–V6).** The compositor boots on the GOP framebuffer and, when a
virtio-gpu driver announces itself, hot-attaches a native backend over the shared-memory
scanout surface — with runtime mode-setting, EDID, and fenced presents, and it
re-attaches across driver restarts. All serial-gated (see [display-v2-plan.md](display-v2-plan.md)).
v1 ([display.md](display.md)) is a compositor that owns the **GOP framebuffer** — it
composites a layer stack into a cacheable back buffer and streams damage to the linear
framebuffer the firmware handed over. That path is portable and good: it drives any GPU,
including a real NVIDIA card at an ultrawide's native resolution, with zero GPU-specific
code. v2 keeps it as the **floor** and makes *scanout* — how a finished frame reaches the
panel — a **pluggable backend**, so the compositor can **upgrade to a real GPU driver when
one is present** and fall back to the framebuffer when it isn't.
The compositor itself (layers, back buffer, damage) does not change. Only the last step —
"put this frame on screen" — becomes swappable.
## The shape
```
compositor (display service) ── layer stack + back buffer + damage (unchanged)
│ composites a frame, then: backend.present(damage)
▼
scanout backend (selected at runtime — GOP by default, native when it appears)
│
├─ GopBackend the v1 path: WC copy back→front to the firmware LFB.
│ Always available. No mode-set, no present fence. THE FLOOR.
│
└─ VirtioGpuBackend talks to a virtio-gpu driver process over a `scanout`
service: present via a shared resource + fenced flush,
a fixed mode list, runtime mode-set, EDID refresh rate.
```
A **backend** is a small interface the compositor calls:
- `surface()` → the pixels to compose into and their geometry `{ptr, pitch, format, w, h}`
(the LFB for GOP; a shared scanout resource for virtio-gpu),
- `present(damage: Rect)` → make the damaged region visible (a no-op-ish WC copy for GOP;
a fenced virtio flush for the native path),
- capability queries — `canModeSet`, `hasFencedPresent` — and, when supported, `modes()` /
`setMode(m)`.
The compositor composes into `surface()` and calls `present(damage)` exactly as it does
today; everything device-specific lives behind the interface.
## Selection and hot-attach
The choice is **dynamic**, because a GPU driver is spawned asynchronously (the device
manager brings it up after boot), and because danos is meant to be resilient:
1. **Boot on GOP.** The compositor starts on `GopBackend` immediately, so there is never a
blank screen while drivers load — the exact v1 behaviour.
2. **Upgrade on announce.** When the virtio-gpu driver has claimed its device and set up a
scanout, it **announces itself to the display service** (a `push`: the driver looks up
`.display` and sends an *attach-scanout* message carrying the shared scanout **surface**
as a capability; the compositor maps it and reaches the driver's present/mode channel by
looking up the registered `scanout` service). The compositor switches to
`VirtioGpuBackend` and re-presents the current frame full-screen. Push beats polling —
the compositor doesn't know a priori which driver, if any, exists, and danos has no
service-registration pub/sub.
3. **Native is restartable, not fallback-on-crash.** Once a native driver has reprogrammed
the device, the firmware's GOP framebuffer is **stale** — "native → GOP" is not a clean
fall-back. So a native driver that **crashes** is *restarted* by its supervisor (the
resilience work already merged), re-announces, and the compositor **re-attaches**
(native → native). The screen freezes on the last frame during the gap — acceptable.
4. **GOP is the floor for "no driver was ever there."** On a real GPU (NVIDIA/AMD/Intel)
the class-0x03 device matches nothing in the driver table, no `scanout` is ever
announced, and the compositor stays on GOP forever — no special-casing. If a native
driver *permanently* gives up (the device manager's crash-loop cap), nothing tells the
compositor and there is no path back to GOP — the screen stays frozen on the last
frame. A GOP revert (sensible only while the LFB is still mappable) is not built.
## The shared-memory primitive this needs
virtio-gpu's scanout resource is **guest RAM** — the driver allocates it and attaches it
to a virtio resource, and the compositor composes into it. That means the compositor
writing into the driver's buffer is **cross-process memory sharing**, the primitive v1
deferred (docs/display.md, "What v1 does not do"). v2 builds it: the natural generalization
of M13 capability-passing from *endpoints* to *memory objects* —
```
shared_memory_create(len) -> {handle, virtual_address} // a shareable, page-aligned RAM region
… pass `handle` as the send_cap on an ipc_call …
shared_memory_map(cap) -> virtual_address // the receiver maps the same physical pages
```
The payoff is leverage: the **same** primitive unlocks **both** native GPU drivers *and*
client-rendered surfaces (an app composing its own bitmap and handing the compositor a
reference instead of drawing by command). One piece of kernel work, two features.
## The virtio-gpu driver
A new ring-3 driver process (the topology v1 anticipated — "split the driver from the
compositor when a second backend arrives"). It claims the virtio-gpu PCI function, and:
- sets up the **control virtqueue** (queue 0 — the only queue it uses; no cursor queue)
and the device's config space,
- creates a **2D scanout resource** backed by a shared-memory region, `attach_backing`s it,
`set_scanout`s it to a CRTC, and on each present transfers and `resource_flush`es the
full current-mode rectangle (damage-narrowed flushes are a later refinement),
- reads **EDID** (the `GET_EDID` control command) to log the monitor's preferred timing
and derive the refresh rate it announces (the compositor's frame-clock seed); the mode
list it offers is a fixed pair — 640×480 and 800×600 — and `set_scanout` at a chosen
mode gives **runtime mode-setting**,
- registers a `scanout` service and announces to the display service.
Its `resource_flush` is the real **present** — and gives a **fenced, tear-free** path a
dumb GOP framebuffer can't.
**Fenced is not vsync.** The fence completes when the device has *consumed* the frame:
real completion feedback, and tear-freedom by snapshot semantics (the host displays
discrete transferred frames, never a half-written surface). It is **not** a vblank —
base virtio-gpu 2D has no display-refresh event at all (Linux's driver for this device
fakes one with a software timer), so nothing paces presents to the monitor's refresh.
Refresh-paced presents need either a native driver's vblank interrupt (delivered over
the existing IRQ-as-IPC path) or the compositor's own frame clock.
## What v2 unlocks — and its honest scope
Behind the abstraction, a native backend gives runtime **mode-setting** (resolution only —
the scanout protocol carries neither refresh nor bpp), an **EDID**-derived refresh rate, and
**fenced presents**. But only on devices we have a driver
for — realistically **VMs** (virtio-gpu, and later maybe Bochs DISPI). Real discrete GPUs
need per-vendor KMS-class drivers that aren't getting written, so they **stay on GOP** —
which is genuinely fine (v1 on the NVIDIA box is smooth). So v2's real value is twofold:
the **pluggable architecture** (a driver slots in when one exists) and a **rich, fenced
path in VMs**, where danos development happens. The framebuffer floor never goes away.
## Locked decisions
- **First native backend: virtio-gpu** — the VM standard; gives mode-set + a real
present/flush (fenced), and exercises the whole pluggable design. Tested with QEMU
`-device virtio-gpu`.
- **Dynamic hot-attach** — boot on GOP, upgrade to native on the driver's announce,
re-attach across driver restarts; GOP is the floor for "no driver ever," not a live
fall-back after a reprogram.
- **Detection = push** (the driver announces to `.display`), not compositor polling.
- **v2 builds the shared-memory capability** (endpoints → memory objects), shared with the future
client-surface path.
## See also
- [display.md](display.md) — v1: the compositor, the GOP-vs-device split, the WC discipline.
- [display-v2-plan.md](display-v2-plan.md) — the ordered build-out.
- [driver-model.md](driver-model.md) — claim / `mmio_map` / MSI / capability passing (M13).
- [resilience.md](../os-development/resilience.md) — the restart machinery the hot-attach leans on.
+325
View File
@@ -0,0 +1,325 @@
# The display service: a framebuffer compositor
The [framebuffer](../os-development/framebuffer.md) the loader hands over is a flat block of pixel
memory, and the kernel's [bootstrap console](../../system/kernel/console.zig) draws text
into it directly. That console is a stop-gap. The **display service**
(`system/services/display/`) is the real thing: an ordinary ring-3 process that *owns*
the framebuffer, composes a stack of **layers** into an off-screen back buffer, and
**presents** finished frames to the screen — the display half of the GUI track
([vision.md](../vision.md)), the sibling of the [input service](input.md).
This note is the architecture and the reasoning behind it. The concrete build order
lives in [display-plan.md](display-plan.md).
## First, a distinction that shapes everything: GOP vs. the PCI device
It is tempting to think "the GOP framebuffer" and "the VGA-compatible display
controller in the PCIe tree" are two different things. They are not — they are **two
interfaces to the same silicon, at different times and different levels**, and knowing
which one you're holding decides what you can do.
- **GOP is firmware's *temporary* driver** for the display controller. It gives you a
linear framebuffer pointer and can set video modes — but only until
`ExitBootServices`. The loader already leans on this: [`queryFramebuffer`](../../boot/efi.zig)
reads the monitor's EDID, picks the native mode, and calls `set_mode` **before**
exiting ([gop.md](../os-development/gop.md)). Once the kernel runs, GOP is **gone** — no `set_mode`, no
mode list, no EDID. What survives is the frozen snapshot in
[`BootInformation.framebuffer`](../../system/boot-handoff.zig): `{base, width, height,
pitch, format, refresh_hz}`, and nothing more.
- **The PCI class-0x03 device is the raw controller** — BARs, config space, registers,
IO ports. It is what you actually *own* after boot. On QEMU's emulated adapter
([`-device VGA,edid=on`](../../build.zig), the Bochs VBE/DISPI model) the `base` GOP handed
you *is* that device's linear-framebuffer BAR — the same physical memory, seen through
a different door. On a real discrete GPU, GOP's `base` is an aperture inside the GPU's
VRAM BAR. danos already decodes this device
([pci-class.zig](../../library/device/pci/pci-class.zig) has the full `display` namespace, and
`pci-bus` already reports it to the [device manager](device-manager.md) with its class
triple) — but nothing binds it yet.
What that difference costs you, concretely:
| You want to… | Dumb GOP framebuffer (boot handoff) | Native device driver (PCI 0x03) |
|-------------------------------------------|-------------------------------------|------------------------------------------|
| **Report** the current mode | ✅ from the handoff | ✅ |
| **Change resolution / bpp at runtime** | ❌ GOP is gone | ✅ program DISPI regs / virtio-gpu queue |
| **Re-read EDID, enumerate monitor modes** | ❌ | ✅ the device exposes an EDID block |
| **Refresh rate** | ❌ (virtual anyway) | only a real KMS driver — far future |
| **vblank / tear-free present** | ❌ no vblank signal | ✅ vblank IRQ + page-flip (real GPUs) |
| **Works on the Pi (no PCI VGA)** | ✅ VideoCore hands a simple FB | ✗ per-device |
The lesson: the **portable base for the whole GUI stack is the GOP / boot-handoff linear
framebuffer**. Runtime mode-setting is a *per-device upgrade* layered on top — and on
the Raspberry Pis there is no PCI VGA at all, so the neutral framebuffer is the only
thing all three target machines share. That is why the display service is built on the
dumb framebuffer first, with the native backend as an optional module behind the same
interface.
## Two constraints this service exists to meet
Like the input service — which existed partly to motivate the asynchronous
[`ipc_send`](ipc.md) primitive — the display service runs straight into two limits the
rest of the system hasn't had to face:
1. **The framebuffer is kernel-only today.** It arrives through the boot handoff, is
mapped into the kernel's physmap, and is touched only by
[`console.zig`](../../system/kernel/console.zig). It is *not* a
[devices-broker](../../system/kernel/devices-broker.zig) node, so
`device.claim`/`mmio_map` cannot reach it, and there is no framebuffer
[syscall](../os-development/syscall.md). A user-space display service needs a **new mechanism just to
touch the pixels**. (See "The handoff" below — this is built.)
2. **danos had no cross-process shared memory.** At v1 the memory syscalls were `mmap`
(private, zeroed), `mmio_map` (a *claimed device's* MMIO), and `dma_alloc` (new
pinned physical). The block driver's "pass a buffer by physical address" trick
([block/protocol.zig](../../library/protocol/block/block-protocol.zig)) works *only because its
consumer is DMA hardware*. A compositor that CPU-reads and blends client layers can't
use it — it would have to *map* another process's memory, which nothing allowed. v1
sidesteps it entirely (see "What v1 does not do"); v2 has since built the primitive
(`shared_memory_create` / `shared_memory_map` / `shared_memory_physical` —
[display-v2.md](display-v2.md)).
## Architecture
```
kernel ── owns the boot framebuffer; bootstrap console only
│ seeds a display-class device node from BootInformation.framebuffer
│ (ResourceKind.memory = [base, height*pitch], write-combining hint,
│ plus DisplayInfo{width, height, pitch, format, refresh_hz})
▼
display service (system/services/display/, ServiceId.display) ← the compositor
│ device.claim(display node) → mmio_map(WRITE-COMBINING) = FRONT buffer (the LFB)
│ mmap(cacheable) a BACK buffer of the same geometry
│ owns: an ordered LAYER STACK + a per-frame DAMAGE tracker (rect list or tile grid)
│ loop: composite dirty layers → back buffer → present dirty rects → front
│ backend is an INTERNAL interface: {gop-fb} at boot; {virtio-gpu} on hot-attach (v2)
▼ reached by name (ipc_lookup); clients drive it over the display protocol
┌────────────────────────────────────┬──────────────────────────────────────┐
drawing clients (v1) surface clients (deferred)
display commands: display surfaces:
create_layer / configure_layer shared_memory_create → pass as a capability →
fill_rect / blit_tile / damage the compositor maps & composites the
present client-rendered bitmap directly
```
The bring-up sequence mirrors a hardware driver's — it is the
[`usb-xhci-bus` `initialise`](../../system/drivers/usb-xhci-bus/usb-xhci-bus.zig) shape
(claim → `mmio_map` → run loop) — and the request/reply service shell is the
[FAT](../../system/services/fat/fat.zig) / [input](../../system/services/input/input.zig) shape
([`service.run`](../../library/kernel/service.zig) with a `protocol.zig` of
`extern struct` messages and an `Operation` tag).
**One process, for now.** v1 is a *single* service that both owns the framebuffer and
composites — it does not split a "framebuffer driver" from a "compositor" the way input
splits `ps2-bus` from the input service. The backend (dumb FB vs. a native GPU) is an
*internal* interface, not a process boundary. That boundary earns its keep only when a
second backend or a second monitor appears; until then it is complexity with no payoff.
## The handoff: a device node + a write-combining map
The framebuffer crosses into user space through the machinery that already exists for
every other device, rather than a bespoke syscall — so it inherits ownership,
release-on-death, and re-claim-on-restart for free (the [resilience](../os-development/resilience.md)
story: a crashed display service returns the LFB to the kernel, and its restart
re-claims it).
- The kernel seeds a synthetic **display-class** node into the
[devices-broker](../../system/kernel/devices-broker.zig) at init (`seedDisplay`), from
`BootInformation.framebuffer`: one `ResourceKind.memory` resource spanning
`[base, height*pitch]`, tagged **write-combining**, plus a small
`DisplayInfo{width, height, pitch, format, refresh_hz}` (the memory resource says *where*
and *how big*; `DisplayInfo` says how to *interpret* the bytes — and `refresh_hz`, the
panel refresh the loader computed from EDID before `ExitBootServices`, seeds the
compositor's frame clock). The node carries no name or index; it is identified purely by
its `display` device class.
- The service `device.claim`s it and `mmio_map`s the resource. The map is
**write-combining**, not the strong-uncacheable that `mmio_map` uses for register
MMIO. The kernel already programs a WC PAT slot for its own console
([`setupPat`](../../system/kernel/architecture/x86_64/paging.zig)); this reaches it from
the user mapping path. **This matters:** an uncacheable framebuffer makes the
back→front blit unusably slow.
- On `claim`, the kernel's bootstrap console goes quiet, so the two never fight over the
LFB. A panic is the one exception — by then the service is likely dead anyway, and a
panic on screen wins.
The display service is a **named boot service**: `init` spawns it by name alongside
`input`/`device-manager`/`fat` ([init.zig](../../system/services/init/init.zig)), and it
self-discovers the display node with `device.enumerate` (matching on `DeviceClass.display`). The [device manager](device-manager.md)
matching path (PCI class 0x03 → a driver) is reserved for the future *native* backend, not
this singleton synthetic node.
## Double buffering and the write-combining discipline
Two buffers, with deliberately different memory types:
- The **front buffer** is the LFB — **write-combining**: fast to *write*, slow to
*read*. The rule is therefore **never read the front buffer**. Only ever stream into
it, sequentially.
- The **back buffer** is ordinary **cacheable** RAM (`mmap`), the same geometry. All
compositing happens here, where reads and read-modify-write blends are cheap.
So a frame is: compose every dirty layer into the cacheable back buffer, then **present**
— copy the changed regions back→front in sequential, WC-friendly writes. Two details the
[framebuffer](../os-development/framebuffer.md) note already establishes carry over: step rows by `pitch`,
not `width*4`; and handle both `rgbx` and `bgrx` [pixel formats](../os-development/gop.md).
## Flicker vs. tearing — what double buffering does and doesn't buy
These are two different artifacts, and the dumb framebuffer fixes exactly one of them:
- **Flicker** is the user seeing intermediate, half-drawn states (a clear-then-redraw
flash). Double buffering **eliminates it completely** — the screen only ever receives
whole, finished frames.
- **Tearing** is a present landing while the display's scanout beam is mid-frame, so the
top of the screen shows the new frame and the bottom the old. Avoiding it requires
presenting during the vertical blank (**vsync**) — which needs a vblank signal. **A
dumb GOP framebuffer has no vblank.**
So v1 is **flicker-free**, and it *minimizes* the tear window by presenting only damaged
rectangles (less to copy → a smaller window in which the beam can catch a half-updated
frame), but it is **not tear-free**. Genuine vsync waits for a backend with a vblank IRQ
or a flush/flip path — a native-device capability, not something the firmware
framebuffer can offer. Stated plainly here so the limitation is understood, not
discovered.
## Layers and the client protocol
The compositor holds an **ordered stack of layers**. Each layer has a rectangle, a
z-order, a visibility flag, and a surface. Presenting walks the stack bottom-to-top,
painting each dirty layer into the back buffer, then flushes the damage to the front.
Damage is tracked by one of two interchangeable trackers behind a compile-time
`damage_mode` A/B switch ([display.zig](../../system/services/display/display.zig)): a
free-form dirty-rectangle **list** (tight bounds, heuristic merging) or a fixed 64-px
**tile grid** (exact O(1) merging, tile-quantized repaints) — the grid is the default;
[compositor.zig](../../system/services/display/compositor.zig) has both, with the trade-off
discussion.
In v1 the surfaces are **server-owned**, and clients draw into them with a small
immediate-mode command protocol — essentially the model early X used, and enough for a
shell, a terminal, a cursor, and a wallpaper:
| Operation | Meaning |
|--------------------|---------------------------------------------------------------|
| `info` | report `{width, height, pitch, format}` of the display |
| `create_layer` | allocate a server-owned surface, return a layer handle |
| `configure_layer` | set a layer's rect, z-order, visibility |
| `destroy_layer` | release a layer |
| `fill_rect` | fill a rectangle of a layer with a colour |
| `blit_tile` | copy a small client-supplied pixel tile into a layer (inline) |
| `damage` | mark a region of a layer dirty |
| `present` | request a repaint: composited at the next frame-clock tick |
Text is intentionally *not* an operation — a client renders glyphs by blitting tiles
(the [PSF font](../../system/kernel/font.psf) path the console already uses can move into a
client). Keeping the protocol to rectangles and tiles keeps the compositor small and the
policy in the client.
`present` is a *request*, not an immediate flush: the compositor runs a ~60 Hz **frame
clock** (a one-shot kernel timer re-armed on demand), and each tick composites all the
damage accumulated since the last one. Any number of client presents and cursor moves
inside one interval coalesce into a single repaint — the software stand-in for vblank
pacing on backends that have none (all of them today; see
[display-v2.md](display-v2.md), "Fenced is not vsync"). Bring-up paths that must put
pixels on screen synchronously (initialisation, the self-checks) bypass the clock.
## `display`
Clients speak the protocol through a new [`library/client/display/display.zig`](../../library/client/display/display.zig),
the [`block`](../../library/device/block/block.zig) shape (a cached `.display` lookup
with a boot-race retry): `display.info()`, a `Layer` handle with `fill` / `blitTile` /
`damage`, and `present()`. Application code never issues the raw syscalls — it calls the
client module, as with every other danos service.
## The cursor: a mouse-listener thread feeding the compositor
The compositor is the single owner of the framebuffer — only the main `service.run` loop
touches the backend and the layer stack. Tracking the mouse without breaking that
ownership is the display's first use of [threads](../os-development/threading.md): the service is built
multi-threaded (`addThreadedUserBinary`) and, at startup, spawns a **mouse-listener
thread** beside the compositor loop.
- **Listener thread.** Blocks on the input service's mouse stream
(`input.subscribeMouse()`), accumulates the relative `dx`/`dy` motion into an absolute
cursor position clamped to the screen, and hands it to the compositor. It never touches
the compositor — so no lock guards the framebuffer. A parked `next()` leaves its core
free to halt ([halting.md](../os-development/halting.md)).
- **The channel.** A single-slot *latest-value* cell (`CursorChannel`) guarded by a
`Thread.Mutex`: the renderer wants where the cursor *is now*, not a replay of
every delta, so a new position overwrites the old. The listener also **pokes** the
compositor awake — the main loop is parked in `replyWait`, so the listener posts a
zero-payload `ipc.send` to the compositor's endpoint, which arrives as a
message-notification ([ipc.md](ipc.md)). The poke is *coalesced*: at most one is queued
while the main loop has not drained the last, so a fast mouse cannot flood the endpoint.
- **Render.** On the poke, the main loop takes the latest position and moves the cursor —
which is just a top-z compositor layer — with the existing `configure` + `present` path
(it damages the old and new footprints, so only those two rectangles repaint).
Two threading facts shape this (both in [threading.md](../os-development/threading.md)). IPC **handles do
not cross threads**, so the listener can't reuse the main loop's endpoint handle — it
`ipc.lookup(.display)`s its *own* handle to the same endpoint to poke through. And a
multi-threaded service doing concurrent IPC is why the kernel's endpoint-create / register
/ lookup syscalls now serialize under the big kernel lock. Shared fate applies: a fault in
the listener takes the whole display down, and the supervisor restarts the process
([resilience.md](../os-development/resilience.md)).
## What v1 does not do (and why that's fine)
Two capabilities are deliberately out of the first cut. Neither reshapes anything above;
both are clean additions behind the interfaces v1 establishes.
- **Client-rendered surfaces (shared memory).** The fast path for a bitmap-heavy app is
to render into its *own* buffer and hand the compositor a *reference*, not a stream of
commands. That needs a cross-process shared-memory primitive — the natural
generalization of the existing M13 [capability passing](driver-model.md)
from *endpoints* to *memory objects* (`shared_memory_create(len) → {cap, virtual_address}`, pass `cap` on
an `ipc_call`, receiver `shared_memory_map(cap) → virtual_address`). v1 avoids it because server-owned
surfaces already prove the whole pipeline; v2 has since built exactly that primitive
([display-v2.md](display-v2.md)) — the client-surface path on top of it is still open.
- **Runtime mode-setting (a native backend).** Detecting the EDID mode list and changing
resolution / bpp at runtime needs the raw PCI device. The first native backend — since
built by v2 ([display-v2.md](display-v2.md)) — is virtio-gpu, behind the same
internal backend interface the dumb framebuffer sits behind. Refresh-rate and colour
management (a gamma LUT) are real-GPU-KMS territory, far beyond this.
## Verifying it
Four QEMU test cases ([tests.zig](../../system/kernel/tests.zig), `python3
test/qemu_test.py <case>`), each layering on the last:
- **`display`** — the kernel handoff: the seeded `display` device is shaped correctly and
the claim → `mmio_map` leaf is genuinely **write-combining** (PAT entry 4), asserted at
the page-table level.
- **`display-service`** — the compositor comes up: it claims the framebuffer, allocates
the cacheable back buffer, presents a cleared frame through the double-buffer path
(`display: online … / presented frame 0`), and a startup **self-check** composites two
overlapping layers on the real framebuffer and reads them back — overlap = the top
layer — logging `display: compositor self-check ok`.
- **`display-demo`** — the full pipeline from a separate process: the hardware-free
[`display-demo`](../system/services/display-demo/) client (the
[`input-source`](../test/system/services/input-source/) analog) drives layers — a wallpaper and
a sliding rectangle — through the layer client API and heartbeats
`display-demo: ok`, proving a frame travelled client → compositor → screen, exactly as
the [input test](input.md) proves an event travels source → service → subscriber. It draws
no cursor and reads no input — the cursor is the service's own (below), and the demo
animates on its own frame timer, independent of the mouse (the test spawns `input`
alongside it to keep that independence honest). The visible motion itself is a screenshot
away via `zig build run-x86-64`.
- **`display-cursor`** — the mouse-listener thread end to end: with the `input` service up,
`input-source mouse` publishes pure motion, and the display's listener thread accumulates
it into a cursor position handed to the render loop over the `CursorChannel`. Once the
cursor has tracked a run of that motion, the service logs
`display: cursor tracking mouse ok`. Runs `smp: 4` — the compositor and listener threads
execute on different cores, which is what surfaced the IPC-under-lock requirement above.
The compositor's pixel math (rectangle clipping, fill, composite, tile blit) and colour
packing are additionally covered by pure host unit tests under `zig build test`.
## See also
- [framebuffer.md](../os-development/framebuffer.md) — the linear framebuffer, pitch vs. width, `volatile`.
- [gop.md](../os-development/gop.md) — GOP, and why only linear RGBX/BGRX modes are paintable.
- [input.md](input.md) — the sibling service; the async `ipc_send` fan-out.
- [driver-model.md](driver-model.md) — claim / `mmio_map`, capability passing, the trust model.
- [device-manager.md](device-manager.md) — matching and supervision (the native backend's route).
- [display-plan.md](display-plan.md) — the ordered build-out.
@@ -0,0 +1,403 @@
# The driver model: buses, classes, and host controllers
[drivers.md](drivers.md) shows how to write *a* driver — claim a device, map its
registers, sleep on its interrupt. That's enough for a leaf device like the HPET. It is
not enough for a disk, a keyboard, or a network card, because those hang off a
*controller*, on a *bus*, speaking a *protocol*, and no single process should have to
know all three.
Real driver stacks factor into three shapes. This document is about what each one is,
what the kernel must give it, how they share code — and precisely which primitive each
is still blocked on.
## Three shapes
| Shape | Owns | Reaches hardware by | Talks to |
|---|---|---|---|
| **Host controller driver** (HCD) | a controller — an xHCI PCI function, an AHCI port block | `mmio_map` + `irq_bind` + DMA | the devices behind it, in its bus's language |
| **Bus driver** | a bus — a PCI bridge, a USB hub | `device_register`, to publish what it finds | class drivers, over IPC |
| **Class / protocol driver** | *nothing* | *nothing* | its bus driver, over IPC |
The last row is the surprising one and the whole point. A USB keyboard driver touches
no registers, takes no interrupts, and maps no memory. It sends HID protocol messages
to whatever published the device, and it works identically whether the controller
below is xHCI, EHCI, or a Raspberry Pi's DWC2. That is what buys you drivers that
outlive the hardware they were written for.
In practice **HCD and bus driver are usually the same process**. An xHCI driver is a
host controller driver (it owns the PCI function, its BARs, its interrupt, its DMA
rings) *and* a bus driver (it enumerates USB devices and publishes them). Splitting
them is a fiction; what matters is that both *roles* have kernel support, because a
plain bus driver with no controller — a USB hub — is also a real thing.
## The device table is the spine
danos already has the right central structure. `system/kernel/devices-broker.zig` holds a table of
`DeviceDescriptor`, each with a parent, a class, and a set of resources. Firmware discovery
seeds it ([discovery.md](../os-development/discovery.md)); `device_register` grows it.
Three invariants make it a capability system rather than a directory:
1. **A claim is exclusive.** `device_claim(id)` succeeds once. Everything downstream —
`mmio_map`, `irq_bind`, `device_register` — checks `devices_broker.ownerOf(id) == me`.
2. **A descriptor is a licence to map physical memory.** Whoever claims a device may
map its `.memory` resources and bind its `.irq` resources. This is why
`device_register` cannot be a free-for-all.
3. **Therefore: containment.** Every resource of a registered child must lie inside a
resource of the same kind on its parent (`devices_broker.contains`). A bus driver can only
ever *subdivide* what it already holds. Without this, `device_register` would be a
syscall named "map any physical page you like."
Containment is transitive by construction: a grandchild is contained in its child,
which is contained in the bus. Nothing can be laundered through a chain.
Note that firmware topology does **not** obey containment, and isn't asked to — a PCI
function's BAR is not inside its host bridge's `bus_range`, because a bus-number range
is not an address window. Discovery is trusted; user space is not.
### What a bus driver looks like
danos ships no demo bus driver — the real ones are `pci-bus`, `ps2-bus`, and
`usb-xhci-bus`. The smallest *honest* shape, illustrated here with an HPET register block
as the "bus" and its comparators as the "devices", is:
```zig
_ = dev.claim(bus.id); // 1. own the bus
const base = dev.mmioMap(bus.id, 0).?; // 2. enumerate it — from the hardware
const n = ((cap.* >> 8) & 0x1F) + 1; // GENERAL_CAP says how many children
for (0..n) |i| { // 3. publish each child
var child = std.mem.zeroes(dev.DeviceDescriptor);
child.class = @intFromEnum(dev.DeviceClass.timer);
child.resource_count = 1;
child.resources[0] = .{ .kind = memory,
.start = bus_mmio.start + 0x100 + 0x20 * i,
.len = 0x20 };
_ = dev.register(bus.id, &child).?; // kernel checks containment
}
```
Each child is left **unclaimed**, which is the handoff: a comparator driver can now
`device_claim` one and `mmio_map` it, and will see only its own 0x20-byte window. A child
whose window escapes the bus is refused; the in-kernel `containment` test asserts the
kernel's table upholds that ([drivers.md](drivers.md)).
A USB device has *no* resources at all: `resource_count = 0`, because it's addressed
through its controller, not by MMIO. That case is allowed and is the common one.
## Families: sharing code between drivers
A "family" is two modules, not one:
- **A logic module** — the parts of the bus that every driver on it re-derives. Config
space walking and BAR decode for PCI. Descriptor parsing, control transfers, and hub
protocol for USB.
- **A protocol module** — the IPC message types that let a class driver talk to
*whatever* published its device. This is the part that makes class drivers portable.
danos already has one of each: `library/device/pci/pci.zig` is a logic module (the
`Function` view of a claimed PCI function),
[`library/protocol/vfs/vfs-protocol.zig`](../../library/protocol/vfs/vfs-protocol.zig) is a
protocol module shared by the mount backends (today the fat server) and their clients.
(The user-space VFS server it was originally written against has since retired — path
routing moved into the kernel, `system/kernel/vfs.zig`'s `fs_resolve` — but the protocol
module outlived it, which is rather the point.) The pattern generalises directly:
```
library/
kernel/ the system library (kernel32-style): the syscall surface split by concern
— ipc, memory (heap/dma/shared-memory), process, time, logging,
file-system, thread, service, plus system-call stubs + start/root
device/ device code grouped by domain; each domain splits into a shareable
data module (enums/wire types, std-only) and a logic module (mmio/IPC)
mmio/ module "mmio" — typed volatile register access + barriers [M14]
model/ module "device-abi" — DeviceDescriptor, DeviceClass, ResourceKind
pci/ "pci-class" (data) + "pci" — config/BAR/capability walk (Function)
usb/ "usb-abi" + "usb-ids" (data) + "usb" — descriptors, control/interrupt/bulk client
acpi/ "acpi-ids" (data) + "aml" — _HID names, the AML interpreter
driver/ module "driver" — device-access syscalls + device-manager hello
block/ module "block" — the block-device client (a device type)
client/ userspace service clients — display, input (a program's view of a service)
protocol/ driver <-> service wire contracts, one module per directory
vfs/ block/ display/ scanout/ input/ power/ device-manager/ usb-transfer/
system/drivers/ one sub-project each → /system/drivers (no `d` suffix)
usb-xhci-bus/ HCD + bus driver imports usb, mmio, usb-transfer-protocol (+ kernel modules)
usb-hid/ class driver imports usb, input-protocol (+ kernel modules)
virtio-gpu/ scanout driver imports pci, mmio, display-/scanout-protocol (+ kernel modules)
```
The split by *dependency weight* is what lets the microkernel stay out of device
business: it imports only the `device-abi` data module (the descriptor types its broker
marshals across the syscall boundary) — never a logic module, never a taxonomy. That one
pure-data import is the only edge from `system/kernel/` into `library/`; decoding a class
code or `_HID` to a name is user space's job (the device manager owns those taxonomies).
A protocol lives in `library/protocol/` when it is the seam between a low-level driver and
a higher-level service (block ↔ filesystem, a scanout driver ↔ the compositor). A driver's
private wire to its *hardware* — virtio-gpu's command set — is not that; it stays a
driver-private file, like the virtio-pci transport beside it.
The build side of this has since landed: [`addUserBinary`](build.zig) injects the
default modules — the library/kernel concern modules (`ipc`, `memory`, `process`, `time`,
`logging`, `file-system`, `thread`, `service`), the device/service clients (`driver`,
`block`, `display`, `input`), plus `mmio`, `xkeyboard-config`, `acpi-ids` — into every user
binary, and per-binary extras — protocol modules, bus logic — are added with
`programModule(exe).addImport(...)`. That's the *entire* mechanism — Zig modules
already give you everything else.
The discipline that makes this work: **a class driver must not import a bus's *hardware*
logic module.** `usb-hid` imports `usb` (the transfer client) and `input-protocol`, never
`pci` and never `mmio`. If a class driver needs `mmio`, it has become an HCD and should be
one. The domain data modules (`usb-abi`, `usb-ids`, `pci-class`) carry no such weight — a
class driver, the device manager, or the kernel may share them freely.
## What exists today
- **M10** — `device_enumerate`, `device_claim`, `mmio_map`. Strong-uncacheable device
grants, `device_grant` teardown.
- **M11** — `irq_bind` / `irq_ack`. IRQ delivered as an IPC notification; mask before
EOI; `irq_ack` is the unmask.
- **M12** — `parent` in `DeviceDescriptor`, `device_register` with resource containment.
- **M13** — capability passing. `ipc_call` / `ipc_reply_wait` grew a `send_cap` argument
and a `received_cap` return (r8): an endpoint travels with a message, installed into
the receiver's handle table (shared, refcount-bumped — a copy, not a move). A full
table fails `-ENOSPC` and does not half-deliver. This is the "open" primitive — a bus
driver mints a per-device endpoint and hands it to a class driver. The `ipc` module
exposes `callCap` and `replyWait(..., send_cap)`, and class drivers consume them now: the
PS/2 keyboard and mouse drivers attach to ps2-bus this way, and the `usb` / `input`
client modules open their per-device and subscription channels with `callCap`.
- **M14** — DMA memory + the memory-ordering layer. `/lib/device/mmio` gives drivers typed
volatile access and `memoryBarrier`/`readMemoryBarrier`/`writeMemoryBarrier` (per-arch); `dma_alloc`/`dma_free` grant
physically-contiguous, pinned, uncacheable, reclaim-on-teardown buffers with the
physical address exposed (`pmm.allocContiguous`, a DMA arena, `mapUserDmaInto`).
`dma_below_4g` caps the address for legacy engines; `dma_write_combining` is accepted
but falls back to coherent until PAT is programmed. The bus drivers use `/lib/device/mmio`,
and `dma_alloc` has real consumers now: the xHCI driver's rings and contexts,
usb-storage's command/status wrappers, virtio-gpu's virtqueue, and the fat
service's bounce buffer.
- **M15** — interrupts for PCI devices, the MSI half. Discovery now gives every PCI
function its 4 KiB ECAM config space as resource 0 (unblocking the capability walk
with no new syscall), and `msi_bind(device_id, endpoint) -> address, data` allocates a
per-device edge-triggered vector, delivered as an IPC notification with no mask and no
ack cycle. Legacy INTx (`_PRT` parsing + shared lines) is deliberately skipped — MSI
is the real answer. QEMU's HPET has no MSI, so delivery is proven with a self-IPI; the
first PCI driver is the first real consumer.
- **Port I/O** — `io_read`/`io_write(device_id, resource_index, offset, width[, value])`:
a claimed device's `io_port` resource lets a driver read/write its ports, gated exactly
like `mmio_map` gates memory (direct ring-3 `in`/`out` stays a #GP). This is what makes
a PS/2 or 16550 driver possible; the low-rate legacy hardware that needs it is fine with
a syscall per access. `io_port` resources were recorded by discovery and ignored — now
they're used.
- **M16 (detection)** — the IOMMU is now *found*: discovery parses the ACPI DMAR table,
maps the first VT-d unit, and reads its version + capabilities (`iommu_present` in the
platform info). This is detection only — **no translation domains are programmed, so
DMA is still unprotected** (the caveat below). Enforcement lands with the first DMA
driver, which is what there is to protect and test against. Proven in the `iommu` test,
booted with an emulated `intel-iommu`.
- **`system_spawn`** — a user-space supervisor starts a driver:
`system_spawn(name, arguments)` loads a binary bundled in the initial-ramdisk as a
fresh ring-3 process; `name` becomes the child's argv[0] and the optional
NUL-separated `arguments` blob its argv[1..], delivered on a SysV entry stack
([sysv.md](../os-development/sysv.md)). This is what
turned the device manager from "log the match" into "run the driver": the kernel now
spawns only `init`, `init` spawns the services, and the **device-manager** discovers
the hardware and spawns each driver ([drivers.md](drivers.md)). Ungated for now — a
spawn capability is future work.
So: **all three shapes work now, and they're started by the device manager, not the
kernel.** The xHCI driver is the HCD-and-bus proof; the USB HID/storage and PS/2 class
drivers reach their devices purely over IPC. What follows are the original design notes
for the primitives that unblocked each shape — exactly why each was the blocker, and
exactly what fixed it.
---
# Proposed ABI
## M13 — capability passing, for class drivers ✅ done
*Implemented as described below (see "What exists today"). The signatures landed
verbatim: `send_cap` in r9, `received_cap` returned in r8, `-ENOSPC` on a full receiver
table with no delivery. The rest of this section is the original design note.*
**The blocker.** A class driver has to reach *its* device. Today the only way to find
an endpoint is the name registry: `ipc_register(service_id, h)` / `ipc_lookup(id)`,
where `ServiceId` is a global integer namespace with `max_services = 8`. You cannot
mint one endpoint per USB device that way, and there is no way for a bus driver to
*hand* a class driver an endpoint. M7 deferred this deliberately.
**The fix.** Let a message carry one handle. Sender names a handle in its own table;
the kernel installs the endpoint into the receiver's table (bumping `refcount`) and
tells the receiver the index it landed at.
```
ipc_call(h, msg, message_len, reply, reply_cap, send_cap) -> reply_len
ipc_reply_wait(h, reply, reply_len, recv, recv_cap, send_cap)
-> recv_len (rax), badge (rdx), received_cap (r8)
```
`send_cap` is a handle or `no_cap` (`~0`). `received_cap` is the index the transferred
endpoint was installed at in the receiver's table, or `no_cap`.
- Both calls grow from 5 args to 6, which fits: `syscall5` uses `rdi/rsi/rdx/r10/r8`,
leaving `r9`. `ipc_reply_wait` already returns two values via `setSyscallResult2`;
this needs a third (`setSyscallResult3`).
- If the receiver's handle table is full, the call fails `-ENOSPC` and **the message is
not delivered** — a half-delivered capability is worse than a failed send.
- `closeHandles` already drops references on exit, so the lifetime story is unchanged.
That single primitive gives you the standard `open` pattern:
```zig
// class driver // bus driver
const h = ipc.lookup(.usb).?; const r = ipc.replyWait(ep, ...);
const dev_ep = ipc.callCap(h, // ... mint a per-device endpoint,
.{ .op = .open, .id = dev_id }); // reply with it as send_cap
// now dev_ep is a private channel to that one device
```
## M14 — DMA memory and the memory-ordering contract, for HCDs ✅ done
*Implemented: `/lib/device/mmio` (typed volatile access + `memoryBarrier`/`readMemoryBarrier`/`writeMemoryBarrier`, per-arch) and
`dma_alloc`/`dma_free` (contiguous, pinned, uncacheable, reclaim-on-teardown, physical
address exposed). `dma_write_combining` still falls back to coherent — real WC needs
PAT, a small follow-up. The rest of this section is the original design note.*
**The blocker.** An HCD is a DMA-engine programmer. It needs a descriptor ring the
device can read, which means memory that is (a) physically contiguous, (b) at a
physical address the driver knows, (c) of the right cacheability, and (d) pinned.
[`sysMmap`](system/kernel/process.zig) gives you *none* of the four: it calls `pmm.alloc()`
once per page, maps writeback-cached, and never reveals a physical address.
**The fix.**
```
dma_alloc(len, flags) -> virtual_address (rax), physical_address (rdx)
dma_free(virtual_address, len) -> 0
flags: dma_coherent (1) uncacheable; the default and the only one that's portable
dma_wc (2) write-combining — needs PAT programmed; for framebuffers
dma_below_4g (4) for devices with 32-bit DMA addressing
```
Guarantees: page-aligned, physically contiguous, zeroed, pinned for the life of the
mapping, and the physical address is stable. It needs one thing the kernel lacks —
`pmm.allocContiguous(n, max_phys)`; today `pmm.alloc()` hands out one frame at a time
with no adjacency guarantee.
**The memory-ordering contract.** danos has, at the time of writing, **zero memory
barriers anywhere in the tree.** That is currently correct-by-accident and won't
survive the first DMA driver, or the first ARM boot.
`volatile` is not a barrier. In Zig it means: don't elide this access, and don't
reorder it against *other volatile* accesses. It says nothing about your *ordinary*
stores — the descriptor you just filled in normal WB memory — which LLVM may freely
sink past a volatile MMIO write. The canonical bug:
```zig
ring[i] = descriptor; // ordinary store to WB RAM
doorbell.* = i; // volatile store to UC MMIO
// nothing stops the compiler reordering these; the device reads a stale descriptor
```
So the rules, which belong in `library/device/mmio/mmio.zig` and behind `arch`:
| Situation | Required |
|---|---|
| MMIO register read/write | `mmio.read` / `mmio.write` (volatile) |
| Fill DMA descriptor, then ring doorbell | `writeMemoryBarrier()` between them |
| Woken by IRQ, then read what the device wrote | `readMemoryBarrier()` before the read |
| MMIO write that must complete before the next read | `memoryBarrier()` |
And the per-arch lowering — the reason this must be an `arch` primitive and not a
sprinkling of `asm volatile`:
| | x86_64 | aarch64 |
|---|---|---|
| `memoryBarrier()` | `mfence` | `dsb sy` |
| `readMemoryBarrier()` | `lfence` | `dsb ld` |
| `writeMemoryBarrier()` | `sfence` | `dsb st` |
| DMA cache coherency | coherent; nothing to do | **not guaranteed**; needs non-cacheable buffers or cache maintenance |
x86 is forgiving here — TSO plus strong-uncacheable MMIO means you usually get away
with a compiler barrier alone. ARM is not, and [vision.md](../vision.md) makes ARM the win
condition. Build the abstraction while there is one caller to fix.
(Zig note: `@fence` was **removed in 0.16**. Use `@atomicRmw(..., .seq_cst)` for a full
barrier, or per-arch inline asm — which is what `library/device/mmio/mmio.zig` should hide.)
## M15 — interrupts for PCI devices ✅ done (MSI)
*Implemented the MSI half: ECAM config space per PCI function (resource 0) and
`msi_bind` (per-device edge-triggered vector, delivered as a notification). Legacy INTx
`_PRT` parsing is skipped on purpose. `msi_bind` returns (address, data) as two values
rather than an out-struct. The rest of this section is the original design note.*
**The blocker, and it's a hard one.** No PCI device can take an interrupt today.
`addBars` (then in the kernel's ACPI discovery; BAR decode has since moved to the
ring-3 pci-bus driver, `system/drivers/pci-bus/pci-bus.zig`) records `.memory` and
`.io_port` BARs and never an `.irq`; there is no `_PRT` parsing anywhere in the tree.
The HPET is the one exception —
it advertises its own interrupt routing in its own registers, a privilege no ordinary
device has.
**The fix, in two halves.**
*Legacy INTx*: parse `_PRT` from the DSDT to map (device, INTA–D) → GSI, and record it
as an `.irq` resource. Then `irq_bind` works unchanged. But INTx lines are **shared**,
and `irq.bound[gsi]` holds one endpoint. Sharing needs a list, and every driver on the
line must be polled on each interrupt — the reason everyone left INTx behind.
*MSI/MSI-X*, which is the real answer: per-device vectors, edge-triggered, unshared, no
mask/ack cycle, no 24-GSI ceiling. The kernel allocates a vector and hands the driver
the (address, data) pair to program into its own MSI capability:
```
msi_bind(dev_id, endpoint, out) -> 0 // out: extern struct { addr: u64, data: u32 }
```
The driver writes those into config space itself — which means it needs config space,
which means **discovery should give each `pci_device` a `.memory` resource for its
4 KiB ECAM slot**. That's a small change to `parseMcfg` and it unblocks the whole
capability walk (MSI, MSI-X, PCIe extended caps) without any new syscall.
Note QEMU's HPET reports `Tn_FSB_INT_DEL_CAP = 0` — no MSI — so an HPET timer could never
exercise this path. The first MSI driver will be the first PCI driver.
## M16 — the IOMMU, and the honest caveat ◑ detection done, enforcement pending
*The IOMMU is now detected (DMAR parsed, VT-d unit mapped and read — see the `iommu`
test), but **enforcement is not built**: no translation domains are programmed, so the
caveat below still holds in full. Detection can't be taken further usefully until there
is a DMA driver to protect and QEMU's `intel-iommu` to test the protection against —
building the per-device domains alongside that first driver is both the natural order
and the only way to verify them. The rest of this section is the original caveat.*
Everything above is capability-gated at the *CPU*. None of it is gated at the *device*.
A driver that can program a bus-mastering engine can make that device write to any
physical address, because page tables sit between the CPU and RAM, not between a device
and RAM. Until VT-d/DMAR (or SMMU on ARM) is programmed from the DMAR table, **`device_claim`
on any DMA-capable device is equivalent to granting ring 0.**
This does not make the model useless — it's the same position Linux is in with the
IOMMU off, and every other guarantee (crash isolation, restart, no shared address
space) still holds. But "user-space drivers are memory-safe" is not true yet, and the
gap should be named rather than implied.
## Ordering
`M13` (capability passing) is independent of `M14`/`M15` and is the cheapest. It
unlocks class drivers, which are the shape with no hardware requirements at all — you
could write a real one against any device a bus driver publishes tomorrow.
`M14` and `M15` together unlock the first HCD. `M14`'s barrier layer is worth landing
on its own regardless: it's small, obviously correct, and stops every future driver
from hand-rolling `*volatile` and getting ARM wrong.
## See also
- [drivers.md](drivers.md) — how to write one, concretely.
- [discovery.md](../os-development/discovery.md) / [acpi.md](../os-development/acpi.md) — where the device table comes from.
- [ipc.md](ipc.md) — endpoints, badges, and the notification path an IRQ arrives on.
- [resilience.md](../os-development/resilience.md) — restart, the reason any of this is worth the trouble.
+407
View File
@@ -0,0 +1,407 @@
# Writing a driver
In a monolithic kernel a driver is a function call away from everything: it runs in
ring 0, dereferences any physical address, and its interrupt handler *is* the ISR. In
danos a driver is **an ordinary ring-3 process**. It has its own address space, it
can crash without taking the kernel with it, and — the point of this document — it
can be restarted ([resilience](../os-development/resilience.md)).
That leaves three questions the kernel has to answer, because a process can't answer
them for itself:
1. **What hardware exists?** → `device_enumerate`, over the device table discovery built
([discovery](../os-development/discovery.md), [acpi](../os-development/acpi.md)).
2. **How do I touch its registers?** → `device_claim` + `mmio_map`: the kernel maps the
device's physical MMIO window into your address space, and from then on it's plain
memory. No syscall per register access.
3. **How do I find out it wants something?** → `irq_bind`: the interrupt is delivered
to you as an IPC notification. You block; the hardware wakes you.
A driver is, in one sentence, *a process that sleeps until its device has something to
say.*
## How a driver gets started: discover, match, spawn
Nothing in the kernel decides that the PCI host bridge needs the `pci-bus` driver — that
is policy, and policy lives in user space. Boot brings user space up as a three-level
supervision hierarchy, each level owning one job:
```
kernel ──spawns──► init (PID 1) ──spawns──► device-manager ──spawns──► pci-bus
| | |
spawns only init, the service supervisor: the driver supervisor: enumerates
publishes the starts the system /system/devices, matches each device
initial-ramdisk services (device-manager, to a driver, and system_spawn's it
so user space can fat, logger, ...). Its
system_spawn from it list is init policy.
```
The kernel launches exactly one process — `init` — and hands it nothing but the raw
ability to start more (`system_spawn(name, arguments)`, which loads a binary bundled
in the initial-ramdisk as a fresh ring-3 process — `name` becoming its argv[0],
the optional arguments its argv[1..], on a SysV entry stack, see sysv.md). Everything else is a user-space decision:
- **init** ([system/services/init](system/services/init/init.zig)) is the **service
supervisor**. It spawns the system services danos brings up at boot — today `input`,
the `device-manager`, `fat`, `display`, `display-demo`, and the `logger` — from a
small list. Drivers are deliberately *not* its job. (An earlier draft listed a `vfs`
service here; that service is retired — the router moved into the kernel as
`fs_resolve`.)
- **device-manager** ([system/services/device-manager](system/services/device-manager/device-manager.zig))
is the **driver supervisor**. It does the three steps a monolithic kernel would do in
its probe path, entirely from ring 3:
1. **Discover** — `device_enumerate` snapshots the device table the kernel built from
ACPI/PCI ([discovery](../os-development/discovery.md)).
2. **Match** — for each device it looks up a driver. The match policy is code, a few
small per-bus tables: from the boot snapshot only the PCI host bridge matches
(→ `pci-bus`); everything else arrives later as bus reports and matches on
identity — `pciDriverForIdentity` (xHCI → `usb-xhci-bus`, virtio-gpu →
`virtio-gpu`), `hidDriverFor` (PNP0303/PNP0F13 → `ps2-bus`), and
`usbDriverForIdentity` (USB keyboard, mouse, storage). A fuller system reads
what each driver *binds* (a manifest under `/system/drivers`, or the driver
describing its own match).
3. **Spawn** — `system_spawn(driver_name, arguments)` starts the matched driver (the
arguments can carry *which* device it matched), which then claims
its device and runs the event loop below.
So "how is a driver discovered and configured" has two halves: **discovery** is the
kernel's device table, read by anyone; **configuration** is two user-space policies —
init's service list and the device-manager's match table. Both are hardcoded in their
respective programs today; the natural next step is to move them into `/etc` (see the
milestone notes in [driver-model.md](driver-model.md)). `system_spawn` is currently
ungated — any process may spawn any bundled binary — because there is no spawn
capability yet.
## The capability: claim before touch
The driver syscall numbers (`system/abi.zig`) with the device types they carry
(`library/device/model/device-abi.zig`), dispatched in `system/kernel/process.zig`:
| # | Call | Meaning |
|---|------|---------|
| 11 | `device_enumerate(buf, max) -> total` | Snapshot the device table |
| 12 | `device_claim(id) -> ok` | Take **exclusive** ownership |
| 13 | `mmio_map(id, res_idx) -> virtual_address` | Map a claimed device's register window |
| 14 | `irq_bind(id, res_idx, endpoint)` | Deliver that device's IRQ as a notification |
| 15 | `irq_ack(id, res_idx)` | Re-arm the IRQ after servicing the device |
| 16 | `device_register(parent_id, desc) -> id` | Publish a child of a device you claimed |
Notice that **nothing takes a physical address or an interrupt number.** Every call
names a device by id and a resource by index. That indirection is the entire security
model. If `mmio_map` took a physical address, any process could map the kernel's
memory; if `irq_bind` took a GSI, any process could bind the keyboard's line and
silently intercept it. Instead the kernel checks two things (`process.ownedGsi`, and
the same check at the top of `systemMmioMap`):
- `devices_broker.ownerOf(dev_id) == me` — you claimed it, and claims are exclusive
- the resource at `res_idx` is of the right *kind* — `memory` for `mmio_map`, `irq`
for `irq_bind`
The claim is the capability. Everything else follows from it.
## Registers: `mmio_map`
`mmio_map` walks the caller's page tables and installs the device's physical frames
with `present | user | writable | nx | device_grant` plus a cache mode
(`system/kernel/architecture/x86_64/paging.zig:mapUserDeviceInto`). Two of those bits
are load-bearing:
- **`pcd | pwt`** — strong-uncacheable, the default cache mode. A device register is
not memory; a cached read would return a stale value and a write might never leave
the CPU. The one exception: a resource flagged write-combining
(`resource_flag_write_combining` — today the kernel-seeded display framebuffer)
gets the PAT bit instead, so pixel writes batch into bursts.
- **`device_grant`** (bit 9, one of the PTE's available bits) — marks the leaf as MMIO
rather than RAM, so `freeSubtree` skips `pmm.free` on it when the address space is
destroyed. Without this, killing a driver would hand the HPET's registers back to
the frame allocator as if they were free RAM. The `iopass` test guards it.
Grants land in their own arena, `0x0000_7100_0000_0000` (PML4[226]), so device pages
never widen an existing mapping.
Then you just… use it:
```zig
const base = dev.mmioMap(dev_id, mmio_res) orelse return;
const counter: *volatile u64 = @ptrFromInt(base + 0xF0);
const now = counter.*; // a load, straight to the hardware. no kernel involved.
```
## Interrupts: the cycle, and why it has that shape
An interrupt handler in a microkernel has a problem. The code that knows how to quiet
the device is in ring 3, in another address space, and it will not run for
microseconds or milliseconds — after a context switch, when the scheduler gets to it.
But the CPU wants an EOI *now*, and a **level-triggered** line stays asserted until
the device is quieted. EOI a still-asserted line and the I/O APIC redelivers
immediately. Forever. The driver never gets to run at all.
The way out is to mask the line before acknowledging it:
```
kernel ISR irqMask(gsi) // line still asserted; stop it reaching a CPU
irqEoi() // now safe to tell the LAPIC we're done
notifyFromIsr() // wake the driver — it runs much later
driver replyWait() -> badge with the notify bit set
<clear the device's status register> // NOW the line deasserts
irq_ack(dev, res) // kernel unmasks: quiet, so it can't refire
```
`irq_ack` is not bookkeeping you could skip. **It is the unmask.** Forget it and the
interrupt fires exactly once, ever; call it before the device is quiet and you get an
interrupt storm. That single fact explains why `irq_bind` and `irq_ack` are two
syscalls and not one.
This is also why `interruptDispatch` (`system/kernel/architecture/x86_64/idt.zig`) no longer issues the EOI
itself. It used to, before running the handler — correct for the LAPIC timer, and
impossible for a routed device line. Each handler now owns its EOI, because only the
handler knows which discipline its source needs.
### The driver side is an event loop, not a callback
`IPC_ReplyWait` returns *either* a client request *or* a notification, told apart by
the top bit of the badge (`ipc_sync.notify_badge_bit`). So a driver is one
single-threaded loop over both of its event sources:
```zig
while (true) {
const r = ipc.replyWait(endpoint, reply, &recv);
if (r.isNotification()) { // r.source() is the GSI
service_device(); // clear the status register
_ = dev.irqAck(id, irq_res); // re-arm
} else {
handle_client_request(recv[0..r.len]);
}
}
```
No reentrancy, no "what am I allowed to call from an interrupt handler", no shared
state between ISR and task context. The interrupt is just a message.
Two properties worth knowing:
- **An interrupt taken while you're elsewhere is not lost.** If the driver is off in
an `ipc_call` to another server when the IRQ fires, `wakeLocked` finds nobody
waiting, but the badge is already on the endpoint's notify ring. The next
`replyWait` pops it (`ipc_sync.replyWait` checks `popNotify` before the sender FIFO).
- **Notifications coalesce, they don't count.** The ring is 8 deep and drops on
overflow. That's correct: an IRQ notification is a *level* ("the device wants
attention"), not a tally. Re-read the device's status register; never assume one
notification means exactly one event. Because the ISR masks the line until you
`irq_ack`, at most one badge per GSI can be outstanding — so the ring can only
overflow if you bind more than eight GSIs to a single endpoint. Don't.
## A whole driver
A minimal leaf driver is only ~150 lines and does all of it. danos ships **no such
example binary** — the driver model is proven by the real drivers (`pci-bus`, `ps2-bus`,
`usb-xhci-bus`), and a teaching example belongs here, in the docs, rather than as a
compiled program nobody runs. Illustrated with a hypothetical HPET timer driver, the
shape is:
```zig
const hpet = findHpet(buf) orelse return; // device_enumerate, look for
// class=timer with memory + irq
_ = dev.claim(hpet.dev_id); // the capability
const base = dev.mmioMap(hpet.dev_id, hpet.mmio).?;
const endpoint = ipc.createIpcEndpoint().?;
// program the hardware over the mapping we were just handed
reg(base, 0x100).* = level | int_enb | (hpet.gsi << 9); // timer 0 config
reg(base, 0x108).* = reg(base, 0xF0).* + period; // comparator
reg(base, 0x010).* |= 1; // ENABLE
_ = dev.irqBind(hpet.dev_id, hpet.irq, endpoint);
while (...) {
const r = ipc.replyWait(endpoint, &.{}, &recv); // blocked. not polling.
if (r.badge & notify_bit == 0) continue;
reg(base, 0x020).* = 1; // clear status -> deassert
reg(base, 0x108).* = reg(base, 0xF0).* + period; // re-arm
_ = dev.irqAck(hpet.dev_id, hpet.irq); // unmask
}
```
The HPET makes a good illustration for a reason that isn't obvious. Its *counter* is a
clocksource — the only way to use it is to read it, so it exercises `mmio_map` without
needing interrupts at all. Its *comparators* are a clockevent, and can be configured
**level-triggered** (`Tn_INT_TYPE_CNF`), which asserts a bit in `GENERAL_INT_STATUS`
that the driver must write-1-to-clear. That's a genuine deassert step, so the full
mask/ack cycle above is exercised for real rather than being decoration on an
edge-triggered line that would have been fine without it.
One wrinkle it also demonstrates: the ACPI HPET table carries **no interrupt number**.
Which I/O APIC inputs a comparator may drive is a bitmask in `Tn_INT_ROUTE_CAP`, in
the device's own registers. So discovery (`acpi.parseHpet`) maps the block, reads the
mask, and records one concrete GSI as an `irq` resource. The driver then programs
`Tn_INT_ROUTE_CNF` to raise exactly that line — and the kernel will only bind the one
it recorded. Hardware that describes itself at runtime still has to fit through a
static capability.
## Publishing children: `device_register`
A device that *contains other devices* — a PCI bridge, a USB hub, or the HPET's block
of comparators — needs a driver that enumerates it and tells the kernel what it found.
That's `device_register`, and it makes the device table a tree rather than a list
(`DeviceDescriptor.parent`).
```zig
var child = std.mem.zeroes(dev.DeviceDescriptor);
child.class = @intFromEnum(dev.DeviceClass.timer);
child.resource_count = 1;
child.resources[0] = .{ .kind = memory, .start = bus_base + 0x100, .len = 0x20 };
const child_id = dev.register(bus_id, &child).?;
```
The child is left **unclaimed**, which is the whole point: another process claims it and
`mmio_map`s it, and sees only that 0x20-byte window.
The rule the kernel enforces is **containment**: every resource of a child must lie
inside a resource of the same kind on its parent. Ranges must nest; a child's IRQ —
still exactly one line — must fall within the parent's IRQ range (a length-1 parent
range is the old exact-match rule). This isn't bureaucracy — a `DeviceDescriptor` is a
licence to map physical memory, so without containment `device_register` would be a
syscall for mapping any page you like. A bus driver may only ever subdivide what it
already owns.
A device with **no resources** is legal and common. A USB device is reached through its
controller, not by MMIO, so it gets `resource_count = 0`.
See [`system/drivers/pci-bus/pci-bus.zig`](../../system/drivers/pci-bus/pci-bus.zig) for a
real one — it claims a PCI host bridge, maps its ECAM window, and publishes each function
it finds as a child — and [driver-model.md](driver-model.md) for how bus drivers, class
drivers and host controller drivers fit together.
## What the kernel does not do for you
- **It does not quiet your device.** That's the whole reason `irq_ack` exists.
- **It does not know your registers.** `mmio_map` hands you a base address; every
offset in this document came from the HPET spec, not from danos.
- **It does not serialise your driver.** Two clients calling one driver endpoint are
serialised by `replyWait`, but nothing stops your driver from being preempted.
## Limits, today
Worth knowing before you write the second driver:
Several things this list used to warn about are now available (see
[driver-model.md](driver-model.md)): **port I/O** (`io_read`/`io_write`, claim-gated by
the device's `io_port` resource — direct ring-3 `in`/`out` is still a #GP, so a PS/2 or
16550 driver goes through these), **DMA memory** (`dma_alloc`: contiguous, pinned,
uncacheable, physical address exposed), **memory barriers** (`library/device/mmio`'s
`memoryBarrier`/`readMemoryBarrier`/`writeMemoryBarrier`, imported as the `mmio` module), **fault isolation** (a ring-3 fault kills only the faulting
process — `killCurrentProcess` — and the machine keeps running,
[resilience](../os-development/resilience.md)), and **reclaim + restart on death** (every path out of a
process releases its claims and IRQ/MSI bindings — `releaseAllOwnedBy`,
`irq.releaseOwner` — and the device manager respawns the driver with backoff,
[device-manager.md](device-manager.md)). What remains:
- **Page granularity.** `mmio_map` rounds to 4 KiB. Two devices sharing a page means
granting one grants the other. A `device_register`ed child's *resource* can be narrower
than a page, but its *mapping* can't.
- **DMA is not contained.** A driver that can program a bus-mastering device can make
that device write to *any* physical address — page tables don't sit between a device
and RAM; an IOMMU does. The IOMMU is now *detected* (M16), but no translation domains
are programmed, so `device_claim` on a DMA-capable device is still effectively
equivalent to granting ring 0. This is the largest gap between the design's promise and
what it delivers; enforcement lands with the first DMA driver.
- **No voluntary `dev_release`.** A *live* driver can't drop a claim — only exit
releases it (any path out of a process runs `releaseAllOwnedBy`) — so handing a
device between running drivers still means exiting.
- **One endpoint per GSI**, so shared legacy PCI INTx lines can't be split between two
drivers. MSI/MSI-X — one vector per device, edge-triggered, unshared — is the real
answer, and QEMU's HPET doesn't offer it (`Tn_FSB_INT_DEL_CAP = 0`).
- **Polarity is hardcoded** active-high in `irq.bind`. A device whose MADT override
says active-low needs that threaded through from discovery.
- **14 device vectors** (33–46) and **24 GSIs**, bounded by the stubs `isr.s` emits and
by a single I/O APIC.
- **Don't bind more than 8 GSIs to one endpoint.** The notify ring is 8 deep and drops
on overflow. With one GSI per endpoint that's unreachable — the line is masked from
the ISR until `irq_ack`, so at most one badge is ever outstanding. Bind nine devices
to one endpoint, though, and a dropped badge leaves that line masked with nobody
left to ack it.
- **On real hardware, the mask/EOI cycle may need a remote-IRR flush.** Masking a
level-triggered redirection entry with remote-IRR set doesn't clear it on some
chipsets, and the line never fires again. QEMU clears it on EOI regardless, so the
tests can't see this. Linux flushes remote-IRR by toggling the entry to edge and
back. See the note at the top of `system/kernel/irq.zig`.
## Verifying it
No demo driver ships to prove this end to end; the *real* drivers do, so the tests
target them and the kernel primitives directly:
- **`device-manager`** — boots only the device manager, which discovers the PCI host
bridge, matches `pci-bus`, and `system_spawn`s it. The test reads kernel state — the
process table and the device tree — to confirm pci-bus came up and registered the
functions it enumerated: the whole discover → match → spawn → driver-up chain.
- **`acpi-ps2`** — a user-space driver (`ps2-bus`) is woken by its device's IRQ,
delivered as an IPC notification, and attaches the keyboard: IRQ-as-IPC, end to end.
- **`pci-scan`** — a user-space driver (`pci-bus`) maps its device's MMIO (the ECAM
window) and walks it: `mmio_map`, end to end.
- **`containment`** — the kernel refuses a `device_register` whose child window escapes
the parent's grant (else it would be a syscall for mapping arbitrary memory), while an
identical re-register stays idempotent. Asserted in-kernel, straight against the broker.
- **`irqfree`** — the teardown path. Binds two owners to one shared endpoint, releases
one, and reads the I/O APIC back: the departing owner's line is masked, the sibling's
is not. That second half is why bindings are keyed on the owning *task* and not on the
endpoint pointer — endpoints are shared, so releasing "everything pointing at this
endpoint" would silently mask a live driver's device.
- **`iopass`** — the `device_grant` teardown rule, so destroying a driver's address
space never returns MMIO frames to the RAM pool.
```
$ python3 test/qemu_test.py device-manager acpi-ps2 pci-scan containment irqfree iopass
device-manager ... PASS (matched 'DANOS-TEST-RESULT: PASS')
acpi-ps2 ... PASS
pci-scan ... PASS (matched 'DANOS-TEST-RESULT: PASS')
containment ... PASS (matched 'DANOS-TEST-RESULT: PASS')
irqfree ... PASS (matched 'DANOS-TEST-RESULT: PASS')
iopass ... PASS (matched 'DANOS-TEST-RESULT: PASS')
```
## What's next (not done here)
The big driver-model pieces — capability passing (class drivers), DMA + barriers, MSI,
and IOMMU detection — are **now done** ([driver-model.md](driver-model.md), M13–M16), as
is **port I/O** (`io_read`/`io_write`, the claim-gated syscalls that make a PS/2 or 16550
driver possible). What's left is IOMMU *enforcement* (per-device domains — it waits on
the first DMA driver to protect and test against) and these smaller items:
- **Releasing a claim** — half done. The kernel now drops *all* of a dead driver's
claims on every path out of a process (`releaseAllOwnedBy`, called from process
teardown), which unblocked restart. A voluntary `dev_release` for a live driver
still doesn't exist.
- **Unregistering children** — half done. Hot-remove works at the manager layer:
the xHCI bus reports `child_removed` on unplug and the device manager prunes its
tree. The kernel's own device table is still append-only, so a bus driver in a
loop can still exhaust the 64-entry table.
- **Restart** — done. The device manager notices a driver's death, reads its exit
reason, prunes the children it reported, and respawns it with exponential
backoff — with a crash-loop cap that marks a repeat offender `failed` instead
of respawning forever ([device-manager.md](device-manager.md)).
- **Interrupt priority / threaded IRQ latency** — still open. `notifyFromIsr`
enqueues the woken driver but doesn't preempt (`wakeLocked` deliberately leaves
that to the caller), so a woken driver waits for the next scheduling point.
## The driver contract (M17–M18)
Claiming and mapping is half of being a danos driver; the other half is the
**lifecycle and protocol contract**, and the runtime makes it nearly free:
- Build on `service.run` — one replyWait loop folding protocol
requests, signals, and notifications into callbacks. The harness answers the
universal zero-length ping and turns `terminate` into a clean exit for you
([process-lifecycle.md](../os-development/process-lifecycle.md)).
- A driver spawned with an assignment (its device id as argv[1]) sends the
versioned `hello` to the device manager inside the deadline, and a **bus**
driver reports what it discovers with `child_added`
([device-manager.md](device-manager.md); usb-xhci-bus is the reference
implementation).
- Crash freely — that is the design. The kernel releases your claims, IRQ
bindings, and MSI vectors at death; the manager reads your exit reason,
prunes what you reported, restarts you with backoff, and your fresh instance
re-claims and re-reports. Never depend on your own cleanup running
(iron rule 1).
+164
View File
@@ -0,0 +1,164 @@
# The input module: broadcasting input events
A keyboard driver has one keystroke and *many* programs that might want it — a shell, a
window server, a logger. None of them owns the hardware, and the driver should not know
who is listening. So between the drivers and the listeners sits the **input service**
(`system/services/input/`): drivers **publish** events to it, programs **subscribe**, and
it fans each event out to every interested subscriber. It is an ordinary ring-3 process
reached over IPC, like the [FAT server](../../system/services/fat/fat.zig) — no kernel knows
what a key is.
## One service, several device classes
The service carries three device classes today — **keyboard**, **mouse**, and
**joystick/gamepad** — and is built to take more
([protocol.zig](../../library/protocol/input/input-protocol.zig)). Each class has its own typed
event:
- `KeyEvent` — `key_down`/`key_up` (physical make/break) and `key_press` (a character was
produced, carrying the Unicode scalar); plus a layout-independent `keycode` and a
`modifiers` bitmask.
- `MouseEvent` — relative `motion` (`dx`/`dy`), `button_down`/`button_up`, and `scroll`.
- `JoystickEvent` — `axis` moves (a signed value on a `control` index) and
`button_down`/`button_up`.
All three travel in one **`InputEvent` envelope** tagged with a `DeviceKind`, so the
fan-out is a single code path and a subscriber can take a mix of classes on one stream.
Decode an envelope with `asKeyboard()` / `asMouse()` / `asJoystick()` (each returns null
unless the tag matches). A subscriber names the classes it wants with a **`device_mask`**,
and the service routes each event only to subscribers whose mask includes its class — so a
mouse-only listener never wakes for keystrokes.
## Why this needed a new kernel primitive
The interesting part is delivery, and it runs straight into the shape of danos IPC.
[ipc.md](ipc.md) describes a **synchronous rendezvous**: a server holds exactly one
pending reply (`Task.ipc_client`) and *must* answer it on its next `replyWait`. Two
consequences decide the whole design:
1. **You cannot block N subscribers waiting for "the next event".** A server can hold only
one caller at a time, so the natural "subscriber calls `next_event()` and blocks" API
is impossible for more than one subscriber. Delivery therefore has to be **push** — the
service reaching out to subscribers — not pull.
2. **A synchronous push can hang the whole service.** If the service delivered with
`ipc_call`, it would block until each subscriber replied. `ipc_call` has no timeout, and
a subscriber's endpoint is an *unregistered* capability the kernel's death path cannot
reach (since display v2's V6, `killOwnedEndpointsLocked` in
[ipc-synchronous.zig](../../system/kernel/ipc-synchronous.zig) marks a dead owner's
*registered* endpoints dead and wakes parked callers with `-EPEER` — but unregistered
ones just drop with the task's handle table). One subscriber that exits mid-delivery
would wedge input for everyone. That is the opposite of the resilience the microkernel
is for.
The fix is the asynchronous send that [ipc.md](ipc.md) had already earmarked as future
work ("asynchronous / buffered send … for notifications between servers"):
```
ipc_send(handle, message_ptr, message_len) -> 0 / -errno
```
`ipc_send` copies a small payload into the endpoint's **bounded queue** and wakes a
receiver, then returns immediately — it never blocks and so can never hang on a dead or
slow subscriber. The receiver picks it up through the same `replyWait` it already runs:
the wake arrives as a **buffered message** — `notify_badge_bit | notify_message_bit` set in
the badge (distinguishing it from a bare IRQ/child-exit notification), the sender's task id
in the low bits, and the payload in the receive buffer, with no reply owed. The queue holds
16 messages per endpoint; a full queue **drops the oldest**, because a buffered message is
discrete data, not a coalescing "level" like an interrupt. See
[ipc-synchronous.zig](../../system/kernel/ipc-synchronous.zig) (`sendLocked`, `popPost`, and
the `replyWait` receive loop).
This is the async counterpart of `ipc_call`, and the input service is its first consumer.
## How the pieces fit
```
keyboard/mouse driver, input-source input service subscriber(s)
----------------------------------- ------------- -------------
connectSource(); loop: replyWait: subscribeKeyboard()/…All:
publishKeyboardEvent(k) ─ ipc_call ─▶ publish → broadcast: createIpcEndpoint()
publishMouseEvent(m) for each sub whose callCap(subscribe,
publishJoystickEvent(j) mask matches event.device: send_cap = ep,
ipc_send(sub_ep) ──────▶ device_mask)
reply ok loop: next()
subscribe → store {ep cap, └─ replyWait(ep)
task id, device_mask} → InputEvent
```
- A **subscriber** calls `input.subscribe(mask)` — or a typed helper: `subscribeKeyboard()`,
`subscribeMouse()`, `subscribeJoystick()` (one class, `next()` returns the decoded event),
or `subscribeAll()` (every class, `next()` returns a tagged `InputEvent`)
([library/client/input/input.zig](../../library/client/input/input.zig)). It creates its own endpoint
and hands it to the service as a **capability** (M13 capability passing — the input
service is that feature's first real user), along with its `device_mask`. Then it loops on
`next()`, a `replyWait` on that endpoint returning each pushed event.
- A **source** (a keyboard, mouse, or joystick driver) calls `input.connectSource()` and the
method for its class: `publishKeyboardEvent`, `publishMouseEvent`, or
`publishJoystickEvent`. Publishing is a short synchronous `ipc_call` the service answers at
once; the service's own fan-out is asynchronous, so publishing never blocks on a slow
subscriber.
- The **service** ([input.zig](../../system/services/input/input.zig)) keeps a small subscriber
table (endpoint handle + owning task id + `device_mask`). On `publish` it `ipc_send`s the
event to every subscriber whose mask includes the event's device class. On `subscribe` it
stores the passed capability and mask and, as housekeeping, prunes any slot whose owning
process has exited (checked against `process_enumerate`) — not for correctness (an async
send to an orphaned endpoint is harmless) but to reclaim the slot.
Publisher and subscriber must be **separate processes**: a single thread that both
published and serviced its own subscription would deadlock (its `publish` call blocks until
the service delivers to its endpoint, which only the same thread could receive).
## Status and follow-ups
- **The keyboard is real.** The `ps2-bus` driver owns PNP0303, which carries *both* the
0x60/0x64 ports and IRQ1, so reading the hardware lives in the bus, not in
[keyboard.zig](../../system/drivers/ps2-bus/keyboard.zig): the bus binds IRQ1 and, on each
interrupt, drains port 0x60, routing every byte by the status register's
auxiliary-output bit to whichever child driver **attached** for that device (an
`AttachRequest` to the well-known `ps2_bus` service, carrying the child's endpoint as a
capability; the bytes then arrive as asynchronous `ForwardedByte` messages, so the IRQ
path never blocks on a child). The keyboard driver decodes the stream — scancode **set 2**,
what the keyboard sends with the 8042's legacy translation off, decoded by
[scancode.zig](../../system/drivers/ps2-bus/scancode.zig) into USB HID usage keycodes with
make/break, typematic-repeat, and modifier tracking (host-tested under `zig build test`) —
and publishes real `key_down`/`key_press`/`key_up` events.
- **Keycode → character** is wired in: the keyboard driver fills a `key_press` event's
`character` through [`library/xkeyboard-config`](../../library/xkeyboard-config/README.md)
(`xkb.map(layout, keycode, mods)` → keysym + Unicode character), synthesizing the ASCII
control characters for Enter/Tab/Backspace/Escape, whose keysyms map to no Unicode. The
layout defaults to `us`; the bus can pass another as the driver's argv[2] — the seam for
a future settings source.
- **The mouse is real too.** IRQ12 is enumerated on the auxiliary device's own ACPI node
(PNP0F13), so the bus claims that node alongside the controller and routes both IRQs to
its one endpoint, acking whichever line the notification's badge names.
[mouse.zig](../../system/drivers/ps2-bus/mouse.zig) attaches the way the keyboard does and
assembles the forwarded bytes with
[mouse-packet.zig](../../system/drivers/ps2-bus/mouse-packet.zig) (three-byte stream-mode
packets: sync/overflow handling, nine-bit movement, screen-convention `dy` — host-tested
under `zig build test`) into `button_down`/`button_up` transitions and `motion` events.
**Follow-up:** the IntelliMouse magic-knock for a scroll wheel (four-byte packets) and
`scroll` events. The hardware-free `input-source` still rotates through all three classes
synthetically (including a joystick, which has no driver yet) via the
`input.synthetic*Event` helpers.
- **Drop-oldest under overflow** is a defined loss; the 16-slot ring absorbs normal bursts.
Real backpressure/flow-control is future work.
- **`publish` is unauthenticated** — any process may publish, consistent with the current
bring-up trust model (see [driver-model.md](driver-model.md)). A source capability is
future work.
## Verifying it
The `input` case (`python3 test/qemu_test.py input`, in
[tests.zig](../../system/kernel/tests.zig) `inputTest`) boots the real kernel and spawns the
service, the synthetic source (which cycles keyboard, mouse, and joystick events), and a
subscriber that took all three classes. It passes only when the subscriber heartbeats
`input-test: ok` — proof that an event travelled source → service → subscriber over IPC,
exercising `ipc_send`, capability-passing subscription, and per-device routing. Each
serial line names the class received, so the log shows all three arriving on one stream.
## See also
- [ipc.md](ipc.md) — the synchronous rendezvous and the notification path `ipc_send` extends.
- [syscall.md](../os-development/syscall.md) — the system-call surface, including `ipc_send`.
- [driver-model.md](driver-model.md) — class drivers, capability passing (M13), the trust model.
@@ -0,0 +1,556 @@
# Native Intel iGPU display support — feasibility and roadmap
**Status: research snapshot, not implemented.** This records what a *minimal, display-only*
native driver for an **Intel integrated GPU** — EDID read + mode-set + framebuffer scanout, with
**no** 3D/media/compute — would take, and how it slots into danos's pluggable scanout
architecture. It is a survey of primary sources (Intel's open-source
[Programmer's Reference Manuals](https://www.intel.com/content/www/us/en/docs/graphics-for-linux/developer-reference/1-0/overview.html),
coreboot's [libgfxinit](https://doc.coreboot.org/gfx/libgfxinit.html), the Linux
[i915 display](https://github.com/torvalds/linux/tree/master/drivers/gpu/drm/i915/display) driver,
and Haiku's [intel_extreme](https://github.com/haiku/haiku/tree/master/src/add-ons/kernel/drivers/graphics/intel_extreme/)),
not an implementation. It is the companion to [nvidia-gpus.md](nvidia-gpus.md) and should be read
against it — the two answer the same question for opposite silicon.
Read [display.md](display.md) and [display-v2.md](display-v2.md) first — this doc assumes the v2
model where scanout is a **pluggable backend** and a native driver is just another `.scanout`
service (like the virtio-gpu one), announcing to the compositor over `attach_scanout`.
## TL;DR
- **Intel is a materially easier, lower-tier target than the NVIDIA RTX 3060 — and the reason is
documentation, not silicon.** Intel publishes official, register-level, per-platform **Display
Engine** PRMs with named registers, bitfields, and numbered enable sequences; NVIDIA publishes
no display PRM and forces reverse-engineering against GPL nouveau. A minimal Intel display-only
driver is roughly **tier 2 to low-tier 3** for well-covered generations (Skylake / Kaby Lake /
Coffee Lake), versus NVIDIA's **tier 4** for GA106. This is the load-bearing conclusion.
- **The display block is a genuinely separable register domain.** Mode-set + scanout touch only
display registers (pipes, planes, transcoders, DDI buffers, PLLs, power wells, GMBUS/AUX) — **no
render engine, no command streamer, no GEM/3D, no signed microcode.** Two small carve-outs, both
trivial pokes that do *not* pull in the render engine: a real CDCLK frequency change writes the
shared GT PCODE mailbox, and the plane's surface register is a GGTT (memory-interface) address.
- **There is no firmware wall on the display path.** The only display microcontroller (DMC / "CSR",
Skylake+) is **optional** — its sole job is saving/restoring display state across DC5/DC6
low-power idle. Without it, i915 prints "Disabling runtime power management" and mode-sets and
scans out normally. GuC/HuC are render/media coprocessors, never touched by a display driver.
Pre-Skylake parts have no display microcontroller at all yet mode-set fine. There is **nothing
analogous to NVIDIA's GSP**.
- **The scanout memory model is dramatically simpler than a discrete GPU.** Intel iGPUs have **no
VRAM**: the display scans out of ordinary system RAM addressed through the Global GTT (GGTT), a
flat single-level page table. Linear (untiled) framebuffers are first-class. You need **no
GEM/TTM, no VMM, no VRAM allocator, no BAR1 aperture juggling** — the exact machinery the NVIDIA
path forces on you.
- **coreboot libgfxinit is a compact, complete, display-only reference** doing precisely this scope
(EDID + PLL/mode-set + scanout, zero 3D) in ~22k lines of formally-analysed SPARK/Ada — versus
i915's ~400k lines. It is a *read-and-reimplement* reference, not drop-in code (GPL-2.0-or-later,
and Ada, not Zig).
- **The clean-room, permissively-licensed path is real** — you can implement from the PRM without
reading GPL code, and Haiku's MIT `intel_extreme` is a permissive precedent. This is the decisive
contrast with NVIDIA, where no vendor register spec exists.
- **The practical catch is hardware, not software.** On a desktop with an RTX 3060, the monitor is
almost certainly cabled to the *card*, so an iGPU driver would light a dark motherboard port; the
CPU may be an **F-SKU with the iGPU fused off entirely**; and every clean-room reference targets
*older* Intel. Intel is the right target to **learn** display bring-up — "run it on my machine"
is a separate, machine-dependent question that may not resolve in the reader's favour.
- **Recommendation:** as with the NVIDIA doc, GOP already gives native-resolution scanout with zero
GPU code. A native Intel driver buys runtime mode changes, hardware vsync, and multihead — and it
reaches "first pixel" far faster than the NVIDIA path *if* the target machine actually has a
usable, cable-attached iGPU of a documented generation.
## Display engine architecture, and why it's separable
For the common single-display path (SST DisplayPort / HDMI / eDP), the Intel display data flow is a
small, fully documented, essentially fixed sequence:
```
memory surface → PLANE(s) → PIPE → TRANSCODER → DDI (drives IO/PHY) → connector
```
The Tiger Lake PRM Vol 12 states it verbatim: *"The front end of the display contains the pipes.
The pipes connect to the transcoders. The transcoders, except for wireless, connect to the DDIs to
drive the IO/PHY."* A **pipe** blends planes (primary/sprite/cursor) into one raster stream; the
**transcoder** wraps it in port-protocol timing (DP/HDMI/eDP/DSI); the **DDI** is the physical port
and PHY. Pipe, Planes, Transcoder, and Digital Display Interface are each first-class PRM chapters
with per-object files in libgfxinit
([TGL PRM Vol 12](https://cdrdv2-public.intel.com/705833/intel-gfx-prm-osrc-tgl-vol-12-display-engine.pdf)).
**Two honest qualifications** the raw research overstated (per verification):
- The pipeline is *not* strictly linear in all cases — the same PRM pages document optional branches
a minimal driver simply ignores (wireless writeback to memory, MIPI DSI, DisplayPort multistream
many-to-one, DSC/tiled pipe-joining). Ignoring them does not weaken feasibility.
- The four-object model *as named* is **Haswell-onward** (DDI introduced ~2013), not "every gen."
Pre-Haswell used FDI + PCH transcoders + port-specific encoders. Within the modern iGPU range
danos would realistically target (Skylake → Meteor/Lunar Lake) the model is stable.
**The DPLL/clock block is a separate, per-port programmable clock source** and is one of the harder,
most gen-specific pieces: pick/enable a PLL, route its output to the DDI, then bring up the port.
The register layout and divider math change substantially per generation — pre-SKL SPLL/WRPLL/LCPLL,
Skylake+ shared DPLL0–3, Gen11+ combo-PHY plus Type-C MG/DKL PLLs. Pixel-clock computation is a
classic per-gen rewrite.
### Separable from render — the single most important enabler
The display is a distinct register domain from render/media, and this is confirmed at the primary
level: the TGL PRM ships display as its own volume (Vol 12), separate from Render Engine (Vol 9) and
Media (Vol 11); Linux's KMS "is provided by Intel Display Driver, and **shared with drm/xe**"
([kernel.org i915](https://docs.kernel.org/gpu/i915.html)) — i.e. the display module is
reused across two different GPU drivers. A full mode-set lights a display end-to-end using only power
wells, PLL/port-clock, DDI-buffer/PHY, transcoder and pipe registers — **zero render commands, zero
GEM objects, zero command-streamer.** libgfxinit is decisive proof: complete EDID + modeset +
framebuffer with no render/3D code at all.
Two carve-outs the "touches ONLY display registers" phrasing needs (per verification), **neither of
which drags in the render engine**:
1. A mode-set that changes the **Core Display Clock (CDCLK)** frequency/voltage pokes the shared **GT
Driver Mailbox** (PCODE/PCU power-controller interface), per Vol 12's own "Display Voltage
Frequency Switching" step. A trivial register handshake, documented alongside the display sequence.
2. The primary plane's surface register (`PLANE_SURF`) holds a **GGTT graphics address** (a
memory-interface concept, not covered in Vol 12). Using pre-mapped stolen memory — as libgfxinit
does — sidesteps any active GGTT programming. See [Memory and scanout](#memory-and-scanout).
### Per-gen churn: what's stable, what you rewrite
The **object model** (pipes/planes/transcoders/DDIs, GMBUS-for-EDID, double-buffered plane registers
armed atomically) is conceptually stable from Ironlake/Haswell through Tiger Lake. What you rewrite
per generation is:
1. the **CPU-vs-PCH split and interconnect**,
2. the **port/PHY + DPLL** programming,
3. **register offsets + power-well / CDCLK topology**, and
4. the **mode-set enable sequence itself** (power-well ordering, PLL lock, DDI-buffer enable,
transcoder clock-select) — an effective fourth axis the raw research folded into (1)/(2).
Interconnect eras, with the timeline **corrected** (the cited Haiku doc was chronologically loose):
- **Gen5 Ironlake (2010) → Ivy Bridge:** FDI (Flexible Display Interface) links the CPU display
engine to PCH-resident ports. The FDI/PCH-split era begins at **Ironlake**, not Gen7.
- **Haswell (Gen7.5):** the main digital outputs come **back onto the CPU die as DDIs** (DDI A = eDP)
— the *opposite* of "moving output to the PCH," and it collapses the FDI/PCH dance **for the
digital ports only**. FDI is **retained** for the legacy VGA/CRT path (DDI E → PCH CRT DAC), so a
driver gets the single DDI code path only by omitting analog VGA (which a minimal driver does).
- **Skylake (Gen9):** reworks clock/PLL, CDCLK, and the power-well model; introduces the optional DMC.
- **Gen11 Ice Lake / Gen12 Tiger Lake:** add combo-PHY + USB-Type-C/Thunderbolt MG/DKL PHYs — the
single biggest cost increase, and the reason "newest silicon" is *not* the easiest target. (DSC is
documented per-**pipe**; MSO is an eDP feature — not "per-transcoder" as the raw research said.)
### The tractable sweet spot
The documented, tractable sweet spot for a from-scratch display-only driver is the
**Haswell (Gen7.5) / Broadwell (Gen8) DDI family, with Skylake (Gen9) as the modern-hardware pick**
since it shares the same DDI object model. Rationale:
- Broadwell has a complete, freely downloadable
[PRM Vol 11 Display](https://cdrdv2-public.intel.com/690828/intel-gfx-prm-osrc-bdw-vol-11-display.pdf);
its engine (3 pipes A/B/C, 4 transcoders incl. transcoder-EDP that floats onto any pipe, DDI A–E,
WRPLL/SPLL/LCPLL) is the classic "DDI + transcoder + WRPLL" model.
- It predates the combo-PHY / Type-C / MG-DKL complexity of Ice Lake / Tiger Lake.
- libgfxinit's DDI **connector/EDID/DP layer is uniform from Haswell through Coffee Lake**, so the
hardest-to-get-right port logic generalises widely.
Two supporting claims from the raw research are **wrong and corrected here (verification):**
- **The BDW and SKL PRMs are NOT 0BSD-licensed.** Both carry a Creative Commons
**Attribution-NoDerivatives** notice. Only the *newer* OSRC PRMs (Tiger Lake 2021 onward) put their
embedded code samples under **Zero-Clause BSD**. So for the recommended Haswell/Broadwell/Skylake
generations there are no "copy-pasteable 0BSD code samples" — the legal basis is *reimplementation
from a CC-BY-ND spec* (register facts are not copyrightable), not copying.
- **FDI+PCH is not fully eliminated on Haswell/Broadwell.** The BDW PRM keeps FDI for the DDI E → PCH
CRT DAC. The "one DDI code path" holds only for the digital outputs a minimal driver targets.
Sandy/Ivy Bridge (Gen6/7) is where the hobby-doc walkthroughs concentrate (the OSDev GMBUS/EDID
material) but carries the FDI+PCH split cost. *(Low confidence on the OSDev specifics — the wiki
returns 403 to automated fetches and its "guaranteed to work" phrasing is a hobby assertion, not a
silicon guarantee.)*
## Documentation — and the clean-room question
This is the crux of the whole comparison. **Intel hands you the register spec that NVIDIA withholds.**
- The Tiger Lake **"Vol 12: Display Engine"** PRM is a real, first-party, open-source document —
**433 pages, verified by direct download** — with named registers + addresses + bitfield tables
(`TRANS_DDI_FUNC_CTL`, `DDI_BUF_CTL`, `DP_TP_CTL`, `PLANE_STRIDE`, `DPLL_CFGCR0/1`, `CDCLK_CTL`,
`PWR_WELL_CTL_DDI`, …) and **numbered, step-by-step enable sequences** with explicit writes, wait
conditions, and microsecond timeouts. It even includes the "magic value" tables older PRMs deferred
to the driver (DisplayPort PLL DCO/divider values; voltage-swing/de-emphasis in mV). *"A spec you
could write a driver from directly"* is well-supported, not hyperbole
([TGL Vol 12](https://cdrdv2-public.intel.com/705833/intel-gfx-prm-osrc-tgl-vol-12-display-engine.pdf)).
- **Clean-room, permissively-licensed implementation is legally and practically feasible from the
PRM alone.** CC-BY-ND governs redistribution of the *document*; register addresses and bit
definitions are functional facts, and original code implementing a described hardware interface is
not a derivative of the PDF. *(This is standard copyright reasoning, not adjudicated case law —
treat it as well-grounded, not settled.)* Two independent implementations already exist built
essentially from these docs (libgfxinit, Haiku), so the spec is demonstrably sufficient.
**The documentation ceiling — corrected.** The raw research said public PRMs stop "roughly at Ice
Lake / Tiger Lake." Verification refuted this: full public **"Vol 12 Display Engine"** PRMs exist for
Ice Lake, Lakefield, Tiger Lake, Rocket Lake, DG1, **and DG2/Arc "Alchemist" (Gen12.5, 2022)** —
[the ACM display PRM is public](https://www.x.org/docs/intel/ACM/intel-gfx-prm-osrc-acm-vol12-displayengine.pdf).
The genuine cliff is **Meteor Lake (2023) and newer**: those have only a high-level architecture
overview, no register-level display PRM, and i915 references their display registers by opaque
internal **Bspec numeric IDs**. Alder Lake and Raptor Lake iGPUs are Gen12 Xe-LP display — the same
IP as Tiger Lake — so despite lacking a dedicated PRM they are effectively covered by the TGL PRM.
Net: a from-docs driver can confidently target **Skylake through DG2/Arc**, which is essentially the
entire current laptop/NUC installed base; only Meteor Lake and later slide back toward the NVIDIA
situation (reverse-engineering or reading GPL i915). The PRMs also survived 01.org's shutdown and are
mirrored in several stable places (Intel's cdrdv2 host, the
[Igalia CC-BY-ND archive](https://github.com/Igalia/intel-osrc-gfx-prm) for Gen4–Gen9.5,
[kiwitree](https://kiwitree.net/~lina/intel-gfx-docs/prm/), x.org) — not a single point of failure.
*(Note: the Igalia archive stops at Kaby Lake and contains no Display Engine volume; the TGL/DG2
display PRMs are separate Intel/x.org downloads.)*
## coreboot libgfxinit — the native reference
[libgfxinit](https://doc.coreboot.org/gfx/libgfxinit.html) is the closest thing to a template danos
could ask for: a self-contained **native modeset library** (no VBIOS/int10, no firmware blobs) that
probes displays via EDID over DDC/I²C and DP AUX, and drives LVDS, eDP, DP1–3, HDMI1–3, analog VGA,
plus USB-C DP/HDMI alt-mode on Tiger Lake. It sets up pipes (Primary/Secondary/Tertiary), planes,
transcoders, PLLs, panel power/backlight, the GTT, and framebuffer scanout — **display-only, zero
3D/media/compute**, which is exactly danos's scope. Its public entry is essentially
`Initialize()` then `Update_Outputs(Pipe_Configs)`, where each `Pipe_Config` carries
`{Port, Framebuffer, Cursor, Mode}` — a near-perfect fit for a pluggable scanout backend.
Why it beats i915 as a reference (**verified by measurement**): **131 Ada source files, ~818 KB,
~22k code lines** across *all* generations, factored precisely along the axes you care about (`edid`,
`dp_aux`, `dp_training`, `pipe_setup`, `transcoder`, `plls`, `connectors`, `port_detect`), with
**none** of the DRM/KMS/GEM/TTM, GT/3D, RC6/RPS, or GuC/HuC machinery that makes
`drivers/gpu/drm/i915` **~419k lines / 900 files / 12 MB**. (A grep confirms *zero* gem/ttm/guc/huc/
execbuf identifiers in the tree.) It depends only on a small HW-access shim, `libhwbase`
(`HW.PCI`, `HW.Port_IO`, `HW.MMIO`, `HW.Time`), which maps naturally onto danos's MMIO-grant + IPC
primitives — you provide Zig equivalents and the modeset logic sits on top. *(Correction to the raw
research: the widely-quoted "~13–14k LOC" is only the generic `common/` layer; the eight
per-generation subdirs roughly double it.)*
**It is a read-and-reimplement reference, not drop-in code.** Two hard constraints:
- **License is GPL-2.0-or-later** (the COPYING file is GPLv2; per-file headers add "or any later
version"). The CC-BY-4.0 on the docs *site* is a footer, not the source license. Copyleft applies
to ported code.
- **It is SPARK/Ada, and designed to run as coreboot boot-firmware**, not a runtime OS driver. A
danos port means either an Ada/GNAT toolchain in the build or hand-transliteration into Zig; the
SPARK "absence of runtime errors" proof does **not** carry over to your reimplementation (and note
it proves absence of runtime errors, **not** functional modeset correctness).
Two more caveats worth knowing: its **error handling is limited** — "only the case that no display
could be found counts as failure"; a later DP link-training failure is *not* propagated. And its
**verified-in-coreboot** hardware list stops at **Coffee Lake + Apollo Lake**, even though the tree
contains a `tigerlake/` directory (Ice Lake has no directory at all, and Alder Lake support is only
"begun"). So treat Haswell..Coffee Lake as the trustworthy transliteration window and TGL as
present-but-less-proven.
The orchestration reads as a clean state machine (`hw-gfx-gma.adb` `Enable_Output`):
`Fill_Port_Config → Preferred_Link_Setting → PLLs.Alloc → [retry] Connectors.Pre_On →
Display_Controller.On → Connectors.Post_On`, with a literal *"try each DP-lane configuration twice"*
inner retry and an outer link-setting step-down. `hw-gfx-dp_training.adb` (398 lines) is a complete,
generic DP link-training implementation (TP1/TP2/TP3, CR + EQ loops, swing/pre-emphasis adjust from
sink status). Per-generation buffer translations plug in underneath via
`Program_Buffer_Translations`, gated on `Config.Has_DDI_Buffer_Trans`. All of this was confirmed
against the source line-by-line.
## The EDID + mode-set path (Haswell/Broadwell target)
The whole path is memory-mapped register programming with polled status bits — no command ring, no
microcode, no DMA channel.
**EDID over DDC (GMBUS).** Pure MMIO poking of the GMBUS I²C controller (`GMBUS0`–`GMBUS5`): `GMBUS0`
selects pin-pair/port + clock; `GMBUS1` carries slave address (`0x50` for EDID), byte count,
direction, SW-ready; `GMBUS2` exposes HW-ready/NAK/ACTIVE to poll; `GMBUS3` is a 4-byte data FIFO;
`GMBUS5` gives the 2-byte segment index for E-DDC. A read is: write `GMBUS0`, write `GMBUS1`
(`CYCLE_WAIT | count | SLAVE_READ | SW_RDY | slave<<addr`), loop {poll `HW_RDY`, read 4 bytes}, then
STOP ([i915 intel_gmbus.c](https://github.com/torvalds/linux/blob/master/drivers/gpu/drm/i915/display/intel_gmbus.c)).
**EDID + DPCD over DP AUX.** For DisplayPort/eDP, EDID (as I²C-over-AUX to `0x50`) and all DPCD
capability/link-status registers are read over the AUX channel: per-DDI `DDI_AUX_CTL` + 5×
`DDI_AUX_DATA`. Build a 3–5 byte header + payload, set SEND_BUSY, poll it clear, read
DONE/TIMEOUT/RECEIVE_ERROR. Message size 1–20 bytes; spec requires ≥3 retries. On Haswell/BDW the AUX
clock divider is programmed explicitly; SKL+ derive it automatically
([i915 intel_dp_aux.c](https://github.com/torvalds/linux/blob/master/drivers/gpu/drm/i915/display/intel_dp_aux.c)).
Both GMBUS and DP-AUX live in libgfxinit's shared `common/` — cheap and nearly gen-invariant.
**The mode-set is a fixed, documented register sequence.** The Broadwell DisplayPort enable order
(verbatim from BDW PRM Vol 11, pp.98–99): (1) DDI lane capability; (2) panel power sequencing if
needed; (3) enable the CPU display PLL (WRPLL/SPLL) and wait ~20 µs; (4) Port Clock Select → DDI,
enable `DP_TP_CTL` with training pattern 1, configure `DDI_BUF_TRANS`, enable `DDI_BUF_CTL`, wait
>518 µs, run link training, set `DP_TP_CTL` to Normal (Idle first for eDP); (5) Transcoder Clock
Select, enable the plane, panel fitter if needed, program transcoder timings + M/N/TU, enable
`TRANS_DDI_FUNC_CTL`, enable `TRANS_CONF`, then backlight. Disable is the exact reverse — a bounded
checklist.
**DisplayPort/eDP link training is driver-driven in software over AUX** — the CPU runs the
clock-recovery and channel-equalization state machines by hand; it is **not** offloaded to a hardware
sequencer or firmware. The source side exposes only primitives: `DP_TP_CTL` selects the training
pattern the port emits; `DDI_BUF_CTL`/`DDI_BUF_TRANS` set voltage-swing/pre-emphasis. The driver
loops: emit pattern + set source levels → write `TRAINING_PATTERN_SET` (DPCD 0x102) + `TRAINING_LANEx_SET`
(0x103) over AUX → delay (100 µs CR / 400 µs EQ) → read `LANE_STATUS` → on failure adjust to the
sink's `ADJUST_REQUEST` values and retry. A few hundred lines of ordinary CPU/AUX code (libgfxinit
`Train_DP`: CR loop 1..32, EQ loop 1..6). **This is the single fiddliest, most fragile piece** — a
TMDS/HDMI panel avoids it entirely, and targeting an already-lit eDP panel avoids most of it.
**The clock (WRPLL) is documented divider math, not a magic table.** On Haswell/BDW the WRPLL derives
the symbol clock from a 2700 MHz LCPLL reference through R2/N2/P dividers with VCO 2400–4800 MHz —
small integer arithmetic. DP is *easier* than HDMI because it runs at a few fixed link rates (1.62 /
2.7 / 5.4 GHz), so a DP/eDP-only minimal driver can often use fixed rates and skip most of the search.
**Plane/scanout programming is trivial for a compositor.** The primary plane is `PRI_CTL`
(enable + pixel format), `PRI_STRIDE`, `PRI_SURF` (surface base — writing it triggers the atomic
update), `PRI_OFFSET`; formats include 32-bit BGRX 8:8:8 and 16-bit BGRX 5:6:5 — a direct match for a
linear XRGB compositor buffer. Plane registers are double-buffered and latch at vblank via an
**arming** write — so a page-flip is "write base + stride + size, then the arming write." This is
*exactly* the primitive danos's damage-driven compositor already expresses over GOP/virtio-gpu; the
incremental work is "program these display-domain registers," not a new scanout model. The panel
fitter (`PF_WIN_POS`/`PF_WIN_SZ`/`PF_CTRL`) can be left disabled for native-resolution scanout;
Skylake+ replaces it with a shared pipe-scaler (`PS_CTRL`).
**Smallest useful target:** eDP (DDI A / transcoder-EDP) or a single DP output at native resolution,
panel fitter off, plane in 32bpp XRGB. That is: GMBUS + I²C-over-AUX EDID/DPCD, one fixed-rate or
WRPLL config, the ~20-step enable sequence, the software CR/EQ loop, and `PRI_*` plane setup with
`PRI_SURF`-write flips. Out of scope: 3D, media, tiling, RC6/power-gating, PSR, audio.
## Memory and scanout
This is where Intel's *architecture* — not just its docs — makes the job smaller, and it is the
biggest single simplification versus a discrete GPU.
- **No VRAM.** Intel iGPUs have a unified memory architecture; the display scans out of ordinary
**system RAM** addressed through the **Global GTT (GGTT)**. The only way to give the GPU memory is
to bind system pages into the GGTT
([i915/GEM crashcourse](https://blog.ffwll.ch/2012/10/i915gem-crashcourse.html)).
- **The plane surface register is a GGTT offset**, not a raw physical address — the display walks the
GGTT to fetch pixels, so a scanout buffer must be GGTT-mapped (global, not per-process). libgfxinit
writes the framebuffer offset straight into `DSPSURF`/`PLANE_SURF` masked to 4 KB.
- **Linear (untiled) scanout is a first-class supported mode** — the plane's tiling field value 0 is
Linear. No X/Y/Yf tiling engine is needed for a display-only driver. (UEFI GOP itself hands off a
linear framebuffer the plane is already scanning.)
- **No memory manager.** You need only (1) some contiguous-ish system pages and (2) GGTT PTEs
pointing at them (`physical_addr | valid_bit` — the GGTT is a flat single-level array of PTEs in
the `GTTMMADR` MMIO BAR), then program the plane. **No GEM/TTM/PPGTT/GuC.** coreboot's native-init
literally does `for(i…) WRITE32(base + i*inc | 1, (i*4) | 1)`.
- **"Stolen memory"** (GSM/DSM) is firmware-reserved system RAM where the firmware places the GGTT
itself and the boot framebuffer. A driver is not obligated to keep scanout there — it can rebind
GGTT entries to its own pages. Stolen memory matters mainly for *inheriting* the GOP framebuffer at
handoff.
**The contrast with NVIDIA is stark.** On a discrete GPU the scanout surface must live in **VRAM**
(nouveau always pins scanout to VRAM), CPU access goes through the **BAR1** aperture (which on
consumer cards can be far smaller than total VRAM unless Resizable BAR is on), and you need a
contiguous aligned VRAM allocator plus a BAR1 mapping. The Intel iGPU path **eliminates all of that**
— scanout is plain system RAM, and a userspace compositor can write the framebuffer pages directly
(as danos already does with the GOP WC framebuffer).
Because danos boots via GOP, an Intel driver attaches to a display whose **GGTT is already populated
and whose plane is already scanning a linear framebuffer at native resolution.** A minimal driver can
reuse that live mapping and reprogram the running plane rather than come up from cold — the same
"attach to a live display" advantage the NVIDIA doc identifies, but with a far smaller register
surface and no firmware wall. *(Low-confidence, per-target details to pin from the specific gen's
PRM: GGTT PTE size — 4-byte pre-gen8 vs 8-byte gen8+ — the `GTTMMADR`/aperture BAR layout, surface
alignment — 4 KB floor but some gens/tilings want 256 KB — and whether the display's GGTT-mediated
DMA sits before or after danos's M16 IOMMU on the target platform.)*
## Firmware
A minimal display-only Intel driver is **effectively firmware-free — more so than NVIDIA.**
- **DMC (Display Microcontroller, "CSR", Skylake+) is NOT required for mode-set or scanout.** Its
sole job is saving/restoring display-engine registers across DC5/DC6 low-power idle. Absent, i915
prints *"Failed to load DMC firmware … Disabling runtime power management"* and the display
mode-sets and scans out normally — you lose only the deep display idle states, not output
([intel_dmc.c](https://github.com/torvalds/linux/blob/master/drivers/gpu/drm/i915/display/intel_dmc.c);
corroborated by multiple distro bug threads). *(A source-level `HAS_DMC` early-return citation would
strengthen this beyond distro testimony, but the conclusion is well-supported.)*
- **Pre-Skylake parts have no display microcontroller at all** yet perform full mode-set (and even
Panel Self Refresh). This confirms the display engine is fundamentally CPU/MMIO-driven; the
microcontroller is an add-on for autonomous idling, not a prerequisite for lighting a panel.
Targeting a pre-Skylake or DMC-optional generation sidesteps the question entirely.
- **GuC and HuC are render/media microcontrollers on the GT side** — GuC schedules the render engines,
HuC assists HEVC/H.265 codec (plus later HDCP/PXP/GSC). Neither is in the scanout path; a
display-only driver never loads them
([kernel.org microcontrollers](https://docs.kernel.org/gpu/i915.html)).
- **PSR firmware lives on the panel**, not in the OS — a minimal driver simply doesn't enable PSR.
- **Type-C/TCSS (Ice Lake+) firmware** (PMC/IOM/PHY) is part of platform BIOS/coreboot init and the
hardware, *not* a signed blob the display driver loads at runtime. A driver attaching to an
already-lit GOP connector, or targeting classic DDI ports, avoids it. *(Cold DP-alt-mode changes
from a userspace driver on modern TCSS platforms were not traced to primary source — flagged.)*
There is **no signed-firmware wall over the Intel GPU at all** on the display path. This is the
architectural opposite of NVIDIA's mandatory, unsignable, ABI-unstable GSP — which even on the
near-side "direct" display path is a permanent maintenance liability for anything beyond scanout.
## Licensing
The situation is *better* than NVIDIA's but still nuanced.
- **The two best code references are both GPL** — Linux i915 (GPL-2.0) and coreboot libgfxinit
(GPL-2.0-or-later). You cannot copy either into a permissively-licensed danos. libgfxinit's WRPLL
divider math is itself copied from i915, so it carries the same encumbrance.
- **But you don't need to copy code.** The Intel PRM is a *specification*, and a clean-room Zig
implementation written from the PRM (using libgfxinit/i915 only to understand behaviour, never to
copy) is legitimate — register numbers and bit definitions are functional facts, not copyrightable
expression. This is the exact inverse of the NVIDIA case, where no such spec exists and the only
guide is the GPL/RE'd code itself.
- **A permissive precedent exists: Haiku's `intel_extreme` is MIT-licensed** and was built from
Intel's public docs. So if danos wants a permissive license, the model is: implement from the PRM,
optionally read MIT Haiku for structure, treat GPL libgfxinit/i915 as documentation-of-last-resort.
- **A licensing nuance on the recommended generations:** the "copy the 0BSD PRM code samples" shortcut
only applies to Tiger-Lake-era (2021+) PRMs. The Haswell/Broadwell/Skylake PRMs are CC-BY-ND, so
their register *facts* are free to implement but there are no code samples to lift.
As with the NVIDIA doc: danos's userspace-driver-over-IPC model (a driver is a separate process behind
a defined protocol) is the cleanest possible license boundary if the project ever chooses to ship a
GPL display-driver binary and keep the rest of danos permissive — but that is a boundary judgement
wanting real diligence, not a settled fact. The clean-room-from-PRM route avoids the question.
## Prior art outside Linux
This is a **real contrast with NVIDIA**, where no one has built a from-scratch native driver outside
Linux. For Intel there are **multiple independent, non-Linux, clean-room native modeset
implementations** to learn from:
- **coreboot libgfxinit** — SPARK/Ada, G45/GM45 and Arrandale → Coffee Lake + Apollo Lake (TGL
in-tree), the strongest structural reference.
- **Haiku `intel_extreme`** — modeset-only (no 2D/3D accel), **MIT-licensed**, i845 through Sandy
Bridge solid, newer Gemini/Ice/Tiger Lake in progress but "hit or miss, as the driver lags behind
the specs" ([Haiku generations](https://www.haiku-os.org/docs/develop/drivers/intel_extreme/generations.html),
[Phoronix Sept 2024](https://www.phoronix.com/news/Haiku-OS-September-2024)).
- **SerenityOS** — added basic native Intel graphics ([PR #6277](https://github.com/SerenityOS/serenity/pull/6277)),
though only for very old ICH7-class hardware.
- **managarm** — native Intel G45 support.
The catch: **every clean-room non-Linux implementation targets old hardware.** A modern Gen12 "Xe"
desktop iGPU is beyond all of them; for the very newest parts only GPL i915 covers the registers. So
the wealth of prior art is real but concentrated below Tiger Lake.
## The practical desktop caveat
Before any effort estimate is trusted, three hardware realities — the honest reason "Intel is easier"
does **not** automatically mean "it'll light up the reader's monitor":
1. **Muxing / cabling.** On a desktop with a discrete RTX 3060, the monitor is almost certainly
plugged into the *card's* outputs, not the motherboard's. An iGPU driver would light a
**different, currently-dark** output. To see danos on Intel the reader would have to physically
move the cable to a motherboard video port **and** likely enable the iGPU / "IGD Multi-Monitor" in
BIOS. Intel-first probably does **not** light the current display without re-cabling.
2. **No iGPU at all.** Intel **F-SKU** desktop chips (i5-9400F, i5-12400F, i5-13400F, i7-13700KF, …)
ship the graphics **fused off** and cannot be re-enabled. These are extremely common in
budget/mid gaming builds paired with an RTX 3060. On an F-SKU (or an X-series HEDT part) the
Intel-iGPU path is a **non-starter** regardless of cabling.
3. **Generation coverage.** If the CPU *is* a recent non-F part, its iGPU may be Gen12 Xe (Alder/
Raptor Lake), beyond libgfxinit's verified set and beyond most non-Linux prior art — leaving GPL
i915 (or the TGL-class PRM, which covers Alder/Raptor display IP) as the only reference.
A cleaner path for *learning* without the hardware lottery: an older bare-metal Intel box (Haswell/
Skylake NUC or laptop) whose panel is natively on the iGPU. Note QEMU does **not** emulate an Intel
iGPU display engine, so a VM cannot exercise a real Intel modeset path — virtio-gpu (already working)
is the VM answer.
## Alternatives, and the honest Intel-vs-NVIDIA verdict
| Option | What you get | The tradeoff |
|---|---|---|
| **Stay on GOP** (working today) | Native-res scanout, zero GPU code/firmware/maintenance | Resolution frozen at ExitBootServices; no runtime mode change, no hardware vsync, no multihead |
| **Intel iGPU, reuse-GOP** | EDID read + plane page-flips on the GOP-set mode | Still bounded to GOP's resolution; but real driver-owned scanout |
| **Intel iGPU, full modeset** (this doc) | Runtime modeset, vsync, multihead, from public docs | Tier 2–3 effort; DP link training; per-gen churn; **needs a cable-attached, documented iGPU** |
| **Native NVIDIA GA106 direct** ([nvidia-gpus.md](nvidia-gpus.md)) | Same, on the RTX 3060 the monitor is actually plugged into | **Tier 4**; GPL-only reference; DMA channel modeset; de-emphasised legacy path |
| **GA106 via GSP/OGKM** | Also unlocks 3D later | Tier 5; unstable version-pinned firmware ABI |
**The verdict for *this reader* (RTX 3060 box):** For pure "see danos on my screen," **NVIDIA-direct
is paradoxically the more relevant path**, because the monitor is already cabled to the 3060 and GOP
already drives it — a native NVIDIA driver reprograms *that* live display. An Intel driver, however
much easier to *write*, likely lights a dark motherboard port the reader isn't looking at, or hits an
F-SKU with no iGPU.
**The verdict for *learning display bring-up*:** **Intel wins decisively.** Public register PRMs, four
independent open reference drivers, an MIT precedent (Haiku), a compact formally-analysed blueprint
(libgfxinit), no signed-firmware wall, no VRAM/BAR memory manager, and a legitimate permissive
clean-room path. It reaches "first pixel" far faster than the NVIDIA native path — *on hardware that
actually has a cable-attached, documented Intel iGPU.* Those two goals — "run on my machine" and
"learn the craft" — point at different silicon, and that is the honest bottom line.
## "First light" milestones — a danos `.scanout` service
Framed as a danos `.scanout` service (like the virtio-gpu and proposed NVIDIA ones), inheriting the
GOP-initialized display — no firmware, no cold POST:
1. **PCI/BAR bring-up** — enumerate the iGPU, map its MMIO BAR (`GTTMMADR` + register block) and the
aperture BAR via danos MMIO grants; confirm the display engine is GOP-live.
2. **EDID** — implement GMBUS DDC (`0x50`) and DP AUX; read + parse the panel EDID and DPCD caps.
*(Smallest self-contained, gen-invariant milestone — a good first commit.)*
3. **First pixel = reprogram, don't re-modeset** — with GOP's mode and GGTT mapping inherited,
reprogram the running plane (`PRI_CTL`/`PRI_STRIDE`/`PRI_SURF`, linear, 32bpp XRGB) to point at a
danos-owned system-RAM buffer; prove a page-flip via the `PRI_SURF` arming write on the *current*
mode before changing timings. This defers the entire DPLL/DDI/transcoder/link-training surface —
the hardest, most gen-specific ~70% of the work.
4. **GGTT ownership** — write your own GGTT PTEs (via an MMIO grant to `GTTMMADR`) pointing at
compositor-owned pages, for double-buffered damage-driven present.
5. **Wire into the compositor `.scanout` backend** (`attach_scanout`); add vsync via the display
vblank interrupt (IRQ-as-IPC).
6. **Full mode-set** (the hard, gen-specific step) — for one chosen generation (Haswell/Broadwell or
Skylake): WRPLL/DPLL programming, the ~20-step DDI/transcoder/pipe enable sequence, panel power
sequencing for eDP (`PP_CONTROL`/`PP_ON_DELAYS`/`PP_OFF_DELAYS` — a common black-screen pitfall).
7. **DisplayPort link training** — only if the panel is DP and GOP's link can't be reused; the
software CR/EQ state machine over AUX. TMDS/HDMI avoids it; a live eDP panel avoids most of it.
8. **Multihead**, then optionally a second generation once one is solid.
Keep the GOP backend as the fallback the whole way — a stall at any step still leaves danos with a
working display, exactly the resilience v2 already provides via re-attach.
## Reading list
**Native reference — coreboot libgfxinit (GPL-2.0-or-later, SPARK/Ada):**
- `common/hw-gfx-gma.adb` — `Enable_Output`, the end-to-end modeset state machine.
- `common/hw-gfx-dp_training.adb` — the complete generic DP link-training CR/EQ loops.
- `common/hw-gfx-gma-pipe_setup.adb` — plane/pipe/scaler + `DSPSURF`/`DSPSTRIDE`/`DSPCNTR` scanout.
- `common/hw-gfx-gma-transcoder.adb` — timing generator; `common/hw-gfx-edid.adb`,
`hw-gfx-gma-i2c.adb`, `hw-gfx-dp_aux_ch.adb` — EDID/DDC/AUX; `hw-gfx-gma-registers.ads` — offsets.
- `common/haswell*/`, `skylake/`, `tigerlake/` — the per-gen PLL/PHY/buffer-translation backends.
**Vendor register specs — Intel OSRC PRMs:**
- [Broadwell Vol 11: Display](https://cdrdv2-public.intel.com/690828/intel-gfx-prm-osrc-bdw-vol-11-display.pdf)
(CC-BY-ND) — the recommended Haswell/Broadwell-class enable sequences, plane, panel fitter.
- [Tiger Lake Vol 12: Display Engine](https://cdrdv2-public.intel.com/705833/intel-gfx-prm-osrc-tgl-vol-12-display-engine.pdf)
(code samples 0BSD) — the most complete modern reference incl. PLL/voltage-swing value tables.
- [DG2/Arc Vol 12: Display Engine](https://www.x.org/docs/intel/ACM/intel-gfx-prm-osrc-acm-vol12-displayengine.pdf)
— the newest public display PRM (Gen12.5, 2022).
- [Igalia CC-BY-ND archive](https://github.com/Igalia/intel-osrc-gfx-prm) (Gen4–Gen9.5) and the
[kiwitree mirror](https://kiwitree.net/~lina/intel-gfx-docs/prm/) — stable mirrors.
**GPL reference-of-last-resort — Linux i915 display:**
- `intel_gmbus.c`, `intel_dp_aux.c` — the concrete EDID/DDC and DP-AUX register sequences.
- `intel_ddi.c` / `intel_ddi_buf_trans.c`, `intel_cdclk.c`, `intel_dpll_mgr.c` — DDI/CDCLK/PLL;
`i9xx_plane.c`, `intel_crtc.c` — plane/pipe; `intel_dp.c` — link training. Huge and modular; a
reference to confirm undocumented quirks, not a template.
**Permissive prior art — Haiku `intel_extreme` (MIT):**
- [`src/add-ons/kernel/drivers/graphics/intel_extreme/`](https://github.com/haiku/haiku/tree/master/src/add-ons/kernel/drivers/graphics/intel_extreme/)
— a second independent modeset-only driver; MIT, so structurally readable for a permissive danos.
- [generations.html](https://www.haiku-os.org/docs/develop/drivers/intel_extreme/generations.html)
— the best plain-English per-generation fault-line map.
## Open questions (unresolved by the survey)
- **Does the target machine have a usable, cable-attached iGPU at all?** F-SKU check, CPU generation,
and monitor cabling must be resolved before any effort estimate is trusted (see
[practical caveat](#the-practical-desktop-caveat)).
- **Does danos even need native mode-*setting*, or only plane/scanout control on the GOP-set mode?**
If runtime mode changes aren't required, the driver collapses to EDID + plane page-flips, dropping
the DPLL/DDI/link-training ~70% of the work.
- **GGTT vs raw physical:** confirm from the exact target-gen PRM that `PLANE_SURF` is interpreted as
a GGTT graphics address (well-established, but per-gen confirmation advisable), and the PTE size /
`GTTMMADR` / aperture layout for writing GGTT entries.
- **Reuse the firmware/GOP GGTT + framebuffer, or install your own GGTT entries?** The latter (needed
for double-buffering) means writing GGTT PTEs from the userspace driver via an MMIO grant.
- **eDP panel power sequencing** (`PP_*`, T1–T12 delays) — not covered in this pass and a common
black-screen source.
- **IOMMU interaction** — whether the display's GGTT-mediated DMA needs IOMMU passthrough for the
framebuffer pages under danos's M16 IOMMU, or sits before the IOMMU on the target platform.
- **DP link-training / AUX robustness and per-generation register drift** are the dominant *risks* —
not documentation scarcity.
- **Exact Haswell/BDW MMIO offsets** (commonly cited: GMBUS ~`0xC5100`, `DDI_AUX_CTL_A` ~`0x64010`,
`DDI_BUF_CTL_A` ~`0x64000`, `DP_TP_CTL_A` ~`0x64040`) were not extracted verbatim from the PRM —
confirm against `i915_reg.h` before coding.
---
*Research snapshot; verify against current libgfxinit / i915 source and the specific target
generation's PRM before building. Intel's public-PRM coverage and the muxing/F-SKU realities of a
given machine both change what is actually achievable.*
+132
View File
@@ -0,0 +1,132 @@
# IPC: message-passing channels
Inter-process communication is the **backbone of a microkernel**. Once drivers and
services run isolated in their own address spaces ([vision](../vision.md)), they can't
just call each other — a request becomes a **message**. In a microkernel, whatever
was a function call across a monolithic kernel is IPC, so it's a first-class
concern, not an afterthought.
There are two layers, built a milestone apart:
- **`system/kernel/ipc.zig`** — a bounded blocking channel between *kernel threads*,
described below. The primitive, and where the blocking discipline was worked out.
- **`system/kernel/ipc-synchronous.zig`** — synchronous call/reply between *processes*, across
address spaces. What user-space servers and drivers actually talk over. It's the
second half of this document.
## The channel
The first form is a **bounded blocking channel** (`system/kernel/ipc.zig`): a fixed-size
ring buffer of messages with a producer/consumer rendezvous, built on the
scheduler's [wait queues](../os-development/scheduling.md).
`Channel(T, capacity)` is generic over the message type and buffer size. It holds a
ring buffer, a count, and two wait queues:
- **`send(msg)`** — if the channel is full, block on the *not-full* queue; otherwise
write the message, bump the count, and wake a waiting receiver.
- **`receive()`** — if the channel is empty, block on the *not-empty* queue; otherwise
take a message, drop the count, and wake a waiting sender.
Neither side busy-waits: a full channel parks the sender, an empty one parks the
receiver, and each operation wakes the other side when it makes progress possible.
Two details make it correct:
- **Recheck in a loop.** A woken task re-tests the condition (`while (full) wait`)
rather than assuming the slot is still available — another waiter may have taken
it first. This is the standard guard against spurious or racing wakeups.
- **One critical section.** `send`/`receive` run under the [big kernel
lock](../os-development/smp.md) (`sync.enter` / `sync.leave`), which disables interrupts on this
core *and* takes the kernel's one spinlock — since SMP, the interrupt flag alone
is not atomicity, because `cli` on one core does nothing to another. So checking
the condition and committing the block/enqueue happen atomically both with respect
to the timer preempting mid-operation and to the other side running on another
CPU. `waitLocked` / `wakeLocked` are the variants that assume the caller already
holds that critical section.
## Verifying it
The `ipc` test (see [testing.md](../testing.md)) runs a producer and a consumer passing
**100 messages through a 4-slot channel**. The small buffer means the channel goes
full and empty over and over, so both the blocking-send and blocking-receive paths are
exercised heavily. The messages arrive intact and in order (their sum is the
expected `5050`), and neither task busy-waits — they block and wake each other.
## Endpoints: call/reply across address spaces
A channel connects two kernel threads sharing one address space. Real servers are
*processes*, so the payload has to cross an address-space boundary. That's
`system/kernel/ipc-synchronous.zig`, and its shape is L4's: a synchronous **rendezvous** at an
`Endpoint`, with the message copied directly from the sender's pages to the receiver's
(`copyAcross` walks both sets of page tables through the physmap — no CR3 switch, no
bounce buffer).
Two syscalls carry it:
- **`ipc_call(h, msg, reply)`** — copy `msg` to the server, block until it replies.
- **`ipc_reply_wait(h, reply, recv)`** — reply to the client you're still holding (if
any), then block for the next request. One syscall, because a server's steady state
is *always* "finish the last one, wait for the next".
An endpoint is reached by **handle** — a small integer index into the process's handle
table (`Task.handles`), exactly like a file descriptor, and just as unforgeable. The
bootstrap problem (how do you get the first handle?) is solved by a tiny name registry:
a server calls `ipc_register(service_id, h)` under a well-known small integer, and a
client calls `ipc_lookup(service_id)`.
The server never learns the client's identity beyond a **badge**, delivered alongside
the message: the caller's task id.
### Interrupts are messages too
`notifyFromIsr` posts an *asynchronous* notification to an endpoint — no payload, no
reply owed — and wakes whoever is blocked in `reply_wait`. Its badge has the top bit
set (`notify_badge_bit`), which is how a driver's single event loop distinguishes "a
client wants something" from "the hardware wants something". Notifications sit in a
small coalescing ring on the endpoint, so an interrupt taken while the driver was busy
elsewhere is not lost.
This is what makes a user-space driver possible at all, and it's the subject of
[drivers.md](drivers.md).
## What's next (partly done since)
- **Priority inheritance** through IPC — still open: a high-priority client
blocked on a low-priority server suffers unbounded priority inversion.
- **Handle transfer.** *Landed as cap-passing (M13)*: `ipc_call` and
`ipc_reply_wait` carry an optional capability alongside the bytes (`send_cap`),
copying an endpoint or shared-memory handle into the peer's table. First user:
[input](input.md) subscribers register by handing over their own endpoint, and
class drivers get a private channel to one device.
- **Asynchronous / buffered send** for the cases where a rendezvous is the wrong
shape (logging, notifications between servers). *Landed as `ipc_send`* — a
non-blocking post to an endpoint's bounded payload queue, delivered through
`reply_wait` as a buffered message (badge bit `notify_message_bit`). Built for, and
first used by, the [input service](input.md)'s keyboard-event broadcast, where a
synchronous push would let one dead subscriber hang the fan-out. A full queue drops
the oldest (discrete messages, not a coalescing level like the notification ring).
- **A bounded reply** — half landed. The copy is still 256 bytes
(`MESSAGE_MAXIMUM`) under the big kernel lock, but bulk transfer got its shared
pages: `shared_memory_create`/`map`/`physical`, the region handle delegated as
a capability (above). virtio-gpu's scanout surface is the first user
([display-v2.md](display-v2.md)).
## Lifecycle conventions over IPC (M17)
Three conventions from [process-lifecycle.md](../os-development/process-lifecycle.md) ride the
notification mechanism:
- **Signals** arrive as notifications on the endpoint a process nominated with
`signal_bind` (`process.bindSignals`): badge = the signal bit plus the
coalesced pending mask (`process.signalsFrom` decodes). Statements,
never questions; no payload, no reply.
- **One-shot timers** (`timer_bind`, `time.timerOnce`) land as a
timer-bit notification — the timed wait: a service arms a deadline and keeps
serving, instead of blocking in sleep.
- **The universal ping**: a **zero-length request is the liveness probe**,
answered with a zero-length reply by the service harness itself
(`service.run`). No protocol's requests start at length zero, so the
encoding cannot collide, and a wedged service simply fails to answer — which
is the diagnosis. Deep health ("can I reach my hardware?") stays a per-service
protocol message.
@@ -0,0 +1,246 @@
# Native NVIDIA GPU support — feasibility and roadmap
**Status: research snapshot, not implemented.** This records what a *native* display driver for a
real discrete NVIDIA GPU — specifically an **RTX 3060 (Ampere GA106)** — would take, and how it
would slot into danos's pluggable scanout architecture. It is a survey of primary sources
(NVIDIA's [open-gpu-kernel-modules](https://github.com/NVIDIA/open-gpu-kernel-modules), the Linux
[nouveau/nvkm](https://github.com/torvalds/linux/tree/master/drivers/gpu/drm/nouveau) driver,
NVIDIA's [open-gpu-doc](https://nvidia.github.io/open-gpu-doc/), and
[linux-firmware](https://github.com/NVIDIA/linux-firmware)), not an implementation. The NVIDIA
driver landscape moves quickly (GSP defaults, firmware ABIs); treat specifics as a mid-decade
snapshot and re-verify against current source before building.
Read [display.md](display.md) and [display-v2.md](display-v2.md) first — this doc assumes the
v2 model where scanout is a **pluggable backend** and a native driver is just another `.scanout`
service (like the virtio-gpu one), announcing to the compositor over `attach_scanout`.
## TL;DR
- A **minimal display-only driver** (EDID + mode-set + framebuffer scanout, **no** 3D/compute)
for the RTX 3060 **can and should avoid the GSP entirely**. nouveau has a register-level,
CPU-driven display path for Ampere (`nvkm/engine/disp/ga102.c`) that lights up GA106 with no
external firmware; the signed-firmware wall gates the **compute/graphics** engines (PGRAPH),
**not** the display controller. "GSP is mandatory on Ampere" is true only for NVIDIA's own
RM-object route.
- **danos's UEFI GOP boot is the single biggest thing in its favour.** The VBIOS/GOP has already
run devinit and brought up the display PLLs, so a driver attaches to a **live, initialized**
GA106 — no firmware load, no cold-boot POST, no devinit interpreter. You reprogram a running
display rather than bring one up from cold.
- It is still a **hard, multi-week-to-months expert effort** (effort tier ≈ 4/5) dominated by
NVDisplay channel-DMA programming, SOR/head routing, DisplayPort AUX + link training, and the
display supervisor handshake. The GSP/RM route is tier 5 (near-infeasible solo).
- The **licensing tension is counterintuitive**: the permissively-licensed reference (NVIDIA
open-gpu-kernel-modules, MIT/GPLv2) is the **hard GSP path**; the register-level display code
you actually want lives in **GPL nouveau**. See [Licensing](#licensing).
- The **window is closing**: GA10x (Ampere) is the *last* NVIDIA family with a register-level
display path — Ada (RTX 40) deleted its non-GSP display HAL. Targeting Ampere specifically
matters.
- **Recommendation:** for *this card*, GOP already gives native-resolution scanout with zero GPU
code and zero maintenance. A native driver buys only runtime mode changes, hardware
vsync/vblank, and multihead. It is justified if that runtime control is a danos goal, or to
*learn the craft* — for which an Intel iGPU or a pre-Turing NVIDIA card reaches "first pixel"
far faster.
## The GSP wall, and why display sits on the near side of it
On Turing and later, NVIDIA split its driver's Resource Manager into a host **CPU-RM** and a
**GSP-RM** running on an on-die RISC-V core ("Peregrine"), talking over RPC
([LWN 953144](https://lwn.net/Articles/953144/)). The GSP is a *full resource manager*, not a
display coprocessor — there is no "display-only" GSP image and no small display RPC subset. Its
boot chain is entirely signed and mandatory: a VBIOS-resident **FWSEC-FRTS** app carves a
write-protected region (WPR2), a signed **Booter** on the SEC2 falcon loads the GSP bootloader,
and that loads **GSP-RM** inside WPR. The firmware ships pre-computed signatures and the driver
picks one by an on-chip fuse-version register — **you cannot self-sign**, and there is **no stable
firmware ABI** (it is revised every driver release; nouveau and the Rust nova-core driver each pin
exactly one version). A GSP driver is a permanent maintenance liability, not a one-time build
([LWN 1037379](https://lwn.net/Articles/1037379/),
[nova-core cover letter](https://lore.freedesktop.org/nouveau/20250826-nova_firmware-v2-7-93566252fe3a@nvidia.com/T/)).
**But display doesn't need any of that on Ampere.** `nvkm/engine/disp/ga102.c` dual-dispatches:
```
if (nvkm_gsp_rm(device->gsp)) return r535_disp_new(&ga102_disp, ...); // GSP RPC path
return nvkm_disp_new_(&ga102_disp, ...); // direct register path
```
Both branches use the same `ga102_disp` HAL and the same `GA102_DISP_*` class IDs; GSP merely
swaps register programming for RPC. GA106 (chipset `0x176`) is wired to `ga102_disp_new` in the
device table, identical to GA102/103/104/107. Ampere lit up displays via the **direct** path in
Linux 5.11/5.17 — two years before GSP-RM landed (6.7, 2023)
([ga102.c](https://raw.githubusercontent.com/torvalds/linux/master/drivers/gpu/drm/nouveau/nvkm/engine/disp/ga102.c),
[Phoronix GA106](https://www.phoronix.com/news/Nouveau-NVIDIA-GA106)).
**Caveat — this is now the legacy path.** As of Linux 6.18, nouveau defaults to GSP on
Turing/Ampere; the direct path is a retained, forceable fallback (`nouveau.config=NvGspRm=0`, and
automatic when GSP firmware is absent). It is stable and proven, but NVIDIA and nova-core are
moving to GSP-only, and **Ada already deleted its non-GSP display HAL**. GA10x is the last family
that keeps a register-level display path.
## What "direct" actually entails
"Direct" is not "plain register pokes." Only SOR / PLL / DP-link / clock setup is bare MMIO. The
**mode-set and scanout themselves flow through the NVDisplay channels — a DMA pushbuffer**:
- Display classes for Ampere (the C670 family): core `GA102_DISP_CORE_CHANNEL_DMA` (`0xc67d`),
window `0xc67e`, window-immediate `0xc67b`, cursor `0xc67a` (headers `clc67d.h` / `clc67e.h` /
`clc67a.h` in [open-gpu-doc `classes/display/`](https://github.com/NVIDIA/open-gpu-doc/tree/master/classes/display)).
- The core channel needs **instance memory, a RAMHT, DMA objects, and a channel user-MMIO
region** ([disp/chan.c](https://raw.githubusercontent.com/torvalds/linux/master/drivers/gpu/drm/nouveau/nvkm/engine/disp/chan.c)).
The register-level "plumbing" to allocate/kick a channel is in NVIDIA's GA102 display register
manual: `NV_PDISP_FE_CHNCTL_CORE/WIN/CURS`, `NV_PDISP_FE_PBBASE/PBBASEHI`
([dev_display_withoffset.ref.txt](https://raw.githubusercontent.com/NVIDIA/open-gpu-doc/master/manuals/ampere/ga102/dev_display_withoffset.ref.txt)).
- **Mode-set is a method stream** on the core channel: `HEAD_SET_RASTER_*`,
`HEAD_SET_PIXEL_CLOCK_FREQUENCY`, `HEAD_SET_CONTROL_OUTPUT_RESOURCE`, `SOR_SET_CONTROL`
(protocol select), viewport/scaler, then `UPDATE`. The window channel points at the scanout
surface (`SET_CONTEXT_DMA_ISO`, `SET_STORAGE`, `SET_OFFSET`).
- After `UPDATE` you must complete the display **supervisor** interrupt handshake (SV1/SV2/SV3).
**EDID and DisplayPort are a separate subdev you must port.** open-gpu-doc documents *none* of
EDID/DDC/AUX. On the direct path you read EDID in-driver via nouveau's `nvkm/subdev/i2c`: bit-bang
**DDC/I²C at address `0x50`** (E-DDC `0x30`) for TMDS/HDMI, or native **DP AUX** in `i2c/aux.c`
for DisplayPort. DisplayPort **link training** (the `dp.c` `train_cr` / `train_eq` state machine
over AUX — clock recovery, lane/rate, voltage-swing/pre-emphasis) is the single hardest and most
fragile piece; a DVI/HDMI (TMDS) panel avoids it entirely.
## The memory floor (smaller than you'd fear)
Neither route hands you a framebuffer allocator — even GSP-RM does not manage the scanout
framebuffer; the driver owns VRAM and merely tells GSP where its page directory is. But
display-only is a small fraction of a full GEM/TTM stack:
- **Pitch-linear (untiled) scanout is allowed** on nv50→Ampere — the window's storage method has a
`PITCH` layout mode, so you skip block-linear tiling math
([wndwc37e.c](https://raw.githubusercontent.com/torvalds/linux/master/drivers/gpu/drm/nouveau/dispnv50/wndwc37e.c)).
- The window references its surface through a simple **display context-DMA**
(`SET_CONTEXT_DMA_ISO` + a 256-byte-granular `SET_OFFSET = addr>>8`) — a base/limit descriptor,
**not** the GPU's 5-level compute page tables. **No full GPU VMM is needed** for scanout.
- The surface must live in **VRAM** in practice (nouveau always pins scanout to VRAM). *Open
question:* whether GA10x can scan out from a system-memory (GART) surface via a sysmem-target
ctxdma — which would let danos skip a VRAM allocator. No source forbids it; nouveau never does
it (confidence: medium).
- **CPU access** to the framebuffer for compositing goes through **BAR1** (a VRAM aperture); BAR0
is the 16 MB register window. BAR1 can be smaller than 12 GB of VRAM unless Resizable BAR maps
it all.
**Net:** you need (1) a contiguous aligned VRAM allocator (256-byte base, pitch a multiple of
64 bytes — confirm against the Ampere display refs), (2) a little instmem for the channel
pushbuffers + iso ctxdma, (3) a BAR1 CPU mapping. You do **not** need the 5-level VMM, GEM/TTM
eviction, or tiling.
## Licensing
The tension is the opposite of convenient:
- **NVIDIA open-gpu-kernel-modules is dual MIT/GPLv2** — usable under MIT, no copyleft on your
other code — **but its display logic is the GSP/RM-object route.** Its class headers
(`cl0073.h`, `cl2080.h`, `ctrl0073*.h`) are useful, permissive references.
- **nouveau is GPLv2**, and the **register-level display sequences you actually want live in
nouveau**, not in the MIT code. So the *easy technical path is the GPL-licensed one.* Reading
GPL nouveau and reimplementing it in Zig is a derivative-work risk proportional to how closely
your code tracks its structure/constants.
Options: **(a)** accept that the danos NVIDIA display driver is a **GPL component**. danos's
userspace-driver-over-IPC model (a driver is a separate process behind a defined protocol, not
linked into the kernel) is about the cleanest possible GPL boundary, so the GPL would be contained
to that one binary and the rest of danos could keep its own license — but this is a
licensing-boundary judgement that wants real diligence, not a settled fact. **(b)** clean-room
from *specification* rather than *code*: [envytools](https://envytools.readthedocs.io) + NVIDIA's
open-gpu-doc register manuals + the MIT OGKM class headers, treating nouveau as
documentation-of-last-resort.
**Firmware licensing is moot for the direct path** (no firmware is loaded). For completeness: the
GSP blobs are marked redistributable under `LICENCE.nvidia`, which permits use by **any
OSI-approved open-source OS** (not just Linux), on NVIDIA GPUs, **unmodified**, with **no
reverse-engineering of the firmware binary**. The one gate — is danos released under an OSI
license? — is only reached on the GSP route, which this doc recommends against for this card.
## Prior art
**No one has built a from-scratch native NVIDIA driver outside Linux.** FreeBSD ships
`nvidia-drm-kmod`, a *port of NVIDIA's own closed `nvidia-drm.ko`* loading the GSP blob (its old
nouveau port was removed). Haiku's NVIDIA support is likewise a *port of OGKM* (GSP, Turing+, very
alpha). OpenBSD / DragonFly have neither. Every non-Linux OS that supports modern NVIDIA chose to
**wrap NVIDIA's GSP stack** rather than write a native driver. A danos direct-register driver
would have exactly one reference implementation — GPL nouveau — and no non-Linux precedent.
## Alternatives
| Option | What you get | The tradeoff |
|---|---|---|
| **Stay on GOP** (working today) | Native-res scanout, zero GPU code/firmware/maintenance | Resolution frozen at ExitBootServices; **no runtime mode change, no hardware vsync, no multihead** |
| **Pre-Turing NVIDIA** (Kepler / early Maxwell) | Direct EVO/disp-core + CRTC/PLL modeset, **no signed firmware, no coprocessor**; mature nouveau reference | Older display class; not this card; only reclocking is firmware-gated |
| **Intel iGPU** | **Publicly documented** register interfaces (Intel PRMs); no coprocessor mediating modeset | i915 is huge + generation-specific; write one generation from the PRM |
| **Native GA106 direct** (this doc) | Runtime modeset, vsync, multihead on the actual card | Tier-4 effort; GPL reference; DP link training; legacy/de-emphasized path |
| **GA106 via GSP/OGKM** | Also unlocks 3D / reclocking later | Tier-5; ~14k-line ante; unstable version-pinned ABI; unprecedented outside Linux |
## "First light" milestones (direct path, inheriting GOP state)
Framed as a danos `.scanout` service (like the virtio-gpu driver), taking the direct register path
and inheriting the GOP-initialized display — no signed firmware, no devinit, no GSP:
1. **PCI/BAR bring-up** — enumerate GA106 (`0x176`), map **BAR0** (registers) and **BAR1** (VRAM
aperture) via danos MMIO grants; confirm the display engine is GOP-live.
2. **VRAM + instmem allocator** — contiguous aligned VRAM for the scanout surface (256-byte base)
+ small instmem for pushbuffers / RAMHT / iso ctxdma. No VMM, no TTM.
3. **EDID** — port `nvkm/subdev/i2c` DDC (`0x50`) + DP-AUX (`aux.c`); read + parse the panel EDID.
4. **Core channel up** — allocate the `0xc67d` core channel as a DMA pushbuffer; stand up the
SV1/SV2/SV3 supervisor-interrupt handshake.
5. **First pixel = reprogram, don't re-POST** — bind a window (`0xc67e`) at the existing WC
framebuffer via `SET_CONTEXT_DMA_ISO` + `SET_OFFSET`, pitch-linear, `UPDATE`; prove you can
drive the *current* GOP mode from your own channel before changing anything.
6. **Modeset** — push raster timings on a head, route head→SOR→connector, program the pixel-clock
PLL, switch to an EDID mode (needs the `clc67d/e` method opcodes from the OGKM headers + the
supervisor timing from nouveau `head.c`).
7. **DisplayPort link training** — only if the panel is DP and GOP's link can't be reused; the
`dp.c` `train_cr`/`train_eq` state machine. TMDS/HDMI is far simpler.
8. **Wire into the compositor `.scanout` backend** (`attach_scanout`), add vsync via the display
interrupt, then multihead.
Keep the GOP backend as the fallback the whole way — a stall at any step still leaves danos with a
working display (exactly the resilience v2 already provides via re-attach).
## Reading list
**Direct path — nouveau (GPLv2):**
- `nvkm/engine/disp/ga102.c` — the GA10x display HAL + the GSP/non-GSP dispatch.
- `nvkm/engine/disp/{head.c, ior.c, dp.c, hdmi.c, chan.c}` — head/SOR routing, DP AUX + link
training, channel-DMA plumbing.
- `dispnv50/{corec37d.c, corec57d.c, wndwc37e.c, wndwc57e.c, wndwc67e.c, headc37d.c, cursc37a.c}`.
- `nvkm/subdev/i2c` (DDC + `aux.c`) for EDID; `nvkm/subdev/bios/init.c` + `devinit/` **only** if
you ever have to re-POST (danos's GOP handoff means you shouldn't).
**Object model / GSP path — NVIDIA OGKM (MIT/GPLv2):** class headers `cl0073.h`, `cl2080.h`,
`ctrl0073system.h`, `ctrl0073specific.h`; `src/nvidia/` for RM control sequences.
`nvidia-modeset.ko` (NVKMS) is a *policy* layer over RM and can be bypassed entirely.
[nova-core](https://lore.freedesktop.org/nouveau/) (Rust) is the forward-looking reference for GSP
boot mechanics (falcon signing, queue rings, RPC).
**Register / method specs — NVIDIA open-gpu-doc:**
- [`classes/display/README.txt`](https://raw.githubusercontent.com/NVIDIA/open-gpu-doc/master/classes/display/README.txt)
— the channel model + class-to-GPU map (read first).
- [`classes/display/clc67d.h`](https://raw.githubusercontent.com/NVIDIA/open-gpu-doc/master/classes/display/clc67d.h)
+ `clc67e.h` / `clc67a.h` — the Ampere core/window/cursor mode-set method vocabulary.
- [`manuals/ampere/ga102/dev_display_withoffset.ref.txt`](https://raw.githubusercontent.com/NVIDIA/open-gpu-doc/master/manuals/ampere/ga102/dev_display_withoffset.ref.txt)
— `NV_PDISP_FE_*` channel/pushbuffer registers + SOR.
- [`DCB`](https://github.com/NVIDIA/open-gpu-doc/tree/master/DCB) — connector→output-resource
routing; [`Devinit`](https://github.com/NVIDIA/open-gpu-doc/tree/master/Devinit) +
[`BIOS-Information-Table`](https://github.com/NVIDIA/open-gpu-doc/tree/master/BIOS-Information-Table)
— VBIOS parsing (bring-up reference; not needed if inheriting GOP).
- The 632 KB Volta [`dev_display.ref`](https://download.nvidia.com/open-gpu-doc/Display-Ref-Manuals/1/gv100/dev_display.ref)
is the best shot at SOR-DP/AUX register detail the smaller Ampere file omits.
## Open questions (unresolved by the survey)
Each needs a direct read of the named nouveau file or experimentation on the actual card:
- Exact GA106 register/method offsets and PADLINK→SOR→connector wiring (can vary by board vendor).
- Whether *any* PLL/devinit re-run is unavoidable vs. fully inherited from GOP.
- Whether DisplayPort needs full retraining on takeover, or the GOP-established link can be reused.
- The precise SV1/SV2/SV3 supervisor sequence.
- Whether a system-memory-target scanout ctxdma could eliminate the VRAM allocator.
- The exact `clc67d.h`/`clc67e.h` method opcode numbers (not captured verbatim in the survey).
---
*Research snapshot; verify against current nouveau / open-gpu-kernel-modules source before
building — NVIDIA's GSP defaults and firmware ABIs change per release.*
+114
View File
@@ -0,0 +1,114 @@
# USB hubs (M22)
A hub is USB **bus infrastructure**, not an application peripheral, so hub
topology is handled **inside the `usb-xhci-bus` driver** — the process that owns
the controller's device slots and contexts. A device behind a hub is not reached
by any hub-specific software path: it is reached by the **controller**,
programmed with a *route string* in its slot context. Route strings and slot
contexts are xHCI hardware concepts that only exist inside the controller driver,
so that is where hub handling belongs. Class drivers (HID, storage) stay separate
and unaware — the hub is transparent to them; a keyboard behind a hub reaches the
same `usb-hid-keyboard` driver as one on a root port.
This is a deliberate scoping choice, not a microkernel compromise: the USB *bus*
driver handles USB *bus* topology. The alternative — a separate `usb-hub`
class-driver process plus a cross-process enumeration protocol — would only
shuttle the bus's own topology state (slot ids, route strings, TT linkage) out to
another process and back, since the hub driver cannot build a slot context
itself.
## The compound-hub reality
A USB 3.0 hub is physically **two hubs** sharing each connector: a SuperSpeed hub
and a USB 2.0 companion hub, enumerated as **separate devices on separate root
ports**. A full- or low-speed device plugged into a USB 3.0 hub attaches to the
**USB 2.0 companion**, not the SuperSpeed hub. So supporting full-speed devices
(keyboards, mice) behind a hub means driving the USB 2.0 companion and handling
**transaction translators** — there is no SuperSpeed-only shortcut that reaches a
full-speed keyboard.
## Slot-context fields for a downstream device
`buildAddressInputContext` fills the Slot Context from the device record: a
root-port device carries route 0 and its own root-hub port. A downstream device
additionally carries:
- **Route String** (Slot Context dword 0, bits 19:0) — 5 tiers × 4 bits, each
tier the downstream hub-port number. Composed as
`route = (parent_route << 4) | hub_port`, capped at the xHCI 5-tier max.
- **Root Hub Port Number** (dword 1, bits 23:16) — the *root* port the whole hub
chain hangs off, inherited from the parent hub (not the hub's own port number).
- **Speed** (dword 0, bits 23:20) — read from the hub's downstream port status
after reset, not assumed.
- **Parent Hub Slot ID** (dword 2, bits 7:0) + **Parent Port Number** (dword 2,
bits 13:8) — the **transaction translator**: set when a full/low-speed device
sits behind a high-speed hub, so the controller routes split transactions
through that hub's TT. For a multi-TT hub, **MTT** (Slot Context dword 0 bit
25) is set and the TT port is the device's own hub port.
## Detection: the status-change interrupt endpoint
A hub has one interrupt IN endpoint that returns a **port-status-change bitmap**
(bit N set = port N changed). The bus arms an interrupt transfer on it (reusing
the controller's existing interrupt-endpoint machinery, but serviced
**in-process** — no class-driver subscription IPC), and on each report:
1. For each changed port, `GET_STATUS` (hub class request) reads the port's
connect/enable/reset state and speed, and `CLEAR_FEATURE(C_PORT_*)`
acknowledges the change.
2. On a **connect**: `SET_FEATURE(PORT_RESET)`, wait for reset-complete via a
later status-change report, read the enabled speed, then `setupDevice` with
the composed route string / root port / TT fields, `enumerate`, and register
the interfaces — exactly the existing path, recursing if the new device is
itself a hub.
3. On a **disconnect**: tear down the downstream device (report each interface
`ChildRemoved`, Disable Slot) — the B3 teardown path, keyed by the device's
route rather than a root port.
## Hub setup (once, when the hub enumerates)
When the bus scan (or a hot-plug bring-up) finds a device of class 9:
1. Read the **hub descriptor** (class GET_DESCRIPTOR, type 0x2A for a USB 3.0
hub / 0x29 for USB 2.0) → downstream port count, characteristics.
2. For a USB 3.0 hub, `SET_FEATURE(BH_PORT_RESET)` semantics and the depth
(`SET_HUB_DEPTH`) so the hub knows its tier for route-string forwarding.
3. `SET_FEATURE(PORT_POWER)` each downstream port.
4. Configure the hub's slot as a hub: **Hub** bit (Slot Context dword 0 bit 26),
**Number of Ports** (dword 1, bits 31:24), **TT Think Time** and **MTT** for a
USB 2.0 multi-TT hub — via an Evaluate/Configure Endpoint on the hub's slot.
5. Arm the status-change interrupt endpoint.
## Testing
QEMU's `usb-hub` is a USB 2.0 single-TT hub. A **static boot topology** on a
dedicated second controller (`-device qemu-xhci,id=xhci2 -device
usb-hub,bus=xhci2.0,port=1 -device usb-kbd,bus=xhci2.0,port=1.1` — isolated from
the boot controller's auto-assigned devices, whose ports the hub would collide
with) presents the downstream device connected from the start, so the bus reads
it on the first status-change report — exercising the full path (hub setup, TT slot context,
downstream enumerate, class-driver bind) without needing a hot-plug event. A new
`usb-hub` QEMU case asserts the hub enumerates, the downstream keyboard
enumerates behind it, and `usb-hid-keyboard` binds.
Real-hardware validation (the user's SuperSpeed Genesys hub + full-speed
keyboard/mouse on its USB 2.0 companion) is flagged separately — the compound
USB 3.0 hub path is not modelled by QEMU's USB 2.0 hub.
## Milestones (all complete)
- **B4a** ✓ — hub recognition + setup: detect class 9 in the scan, read the hub
descriptor, configure the slot as a hub, power downstream ports, log the
topology.
- **B4b** ✓ — downstream enumeration: the in-process status-change subscription,
port reset, Address Device with route string + root port + TT fields,
enumerate + register. A full-speed keyboard behind a USB2 hub binds
`usb-hid-keyboard` in QEMU.
- **B4c** ✓ — disconnect teardown (recursive: a hub takes its subtree with it)
and hub-behind-hub recursion (route strings compose across tiers). QEMU's hub
*does* raise downstream status changes, so both connect and disconnect are
harness-tested (`usb-hub`, `usb-hub-nested`, `usb-hub-unplug`).
Real-hardware validation of the user's SuperSpeed Genesys hub with full-speed
devices on its USB 2.0 companion remains pending — QEMU's USB 2.0 hub does not
model the compound USB 3.0 hub.