Files
danos/docs/display-v2.md
T
Daniel Samson 0217662808 display: resilient scanout — supervised driver, survive loss, re-attach (v2 V6)
The compositor now survives the virtio-gpu driver dying and re-attaches when device-manager
restarts it — the last piece of display v2.

Three parts:

- The driver hellos the device manager (role: bus). It never did, so the manager — which
  spawns it supervised and expects a hello — was stopping it at the 3s hello deadline every
  run (the gate markers just printed first). Now it is properly supervised: not stopped for
  silence, and restarted on death.

- A kernel IPC fix so a call to a dead service errors instead of hanging. An endpoint records
  its owner; when that task dies, its registered endpoints are marked dead (and any parked
  senders woken with -EPEER), so ipc_call returns -EPEER rather than blocking on a reply that
  will never come. Without this the compositor's first present after the driver died blocked
  forever. General robustness — any client of any service benefits.

- The compositor re-attaches. Its .scanout calls now fail cleanly (caught), freezing the last
  frame; when the restarted driver re-announces, attach_scanout detects the backend is already
  virtio and logs "scanout re-attached", mapping the fresh shared surface and re-looking-up
  .scanout. (The previous shm mapping leaks — no shm_unmap syscall yet — but its frames are the
  dead driver's, reclaimed on exit.)

- device-manager gains a "test-scanout-restart" mode (like test-usb-restart) that kills the
  virtio-gpu driver once after it hellos; the displayReattachTest kernel scenario drives it.

Gate: python3 test/qemu_test.py display-reattach — "scanout upgraded to virtio-gpu" then
"scanout re-attached", no CPU exception, passing 3/3. host tests, ipc/ipc-call/ipc-cap,
supervision, shm, display-service, display-demo, virtio-gpu, display-native, and
display-modeset all pass; default zig build clean. v2 (V1-V6) complete.
2026-07-14 13:10:21 +01:00

7.4 KiB
Raw Blame History

The display service v2: a pluggable scanout backend

Status: complete (V1–V6). The compositor boots on the GOP framebuffer and, when a virtio-gpu driver announces itself, hot-attaches a native backend over the shared shm scanout surface — with runtime mode-setting, EDID, and fenced (vsync) presents, and it re-attaches across driver restarts. All serial-gated (see display-v2-plan.md).

v1 (display.md) is a compositor that owns the GOP framebuffer — it composites a layer stack into a cacheable back buffer and streams damage to the linear framebuffer the firmware handed over. That path is portable and good: it drives any GPU, including a real NVIDIA card at an ultrawide's native resolution, with zero GPU-specific code. v2 keeps it as the floor and makes scanout — how a finished frame reaches the panel — a pluggable backend, so the compositor can upgrade to a real GPU driver when one is present and fall back to the framebuffer when it isn't.

The compositor itself (layers, back buffer, damage) does not change. Only the last step — "put this frame on screen" — becomes swappable.

The shape

   compositor (display service)  ── layer stack + back buffer + damage   (unchanged)
        │  composites a frame, then: backend.present(damage)
        ▼
   scanout backend  (selected at runtime — GOP by default, native when it appears)
        │
        ├─ GopBackend        the v1 path: WC copy back→front to the firmware LFB.
        │                    Always available. No mode-set, no vsync. THE FLOOR.
        │
        └─ VirtioGpuBackend  talks to a virtio-gpu driver process over a `scanout`
                             service: present via a shared resource + flush (real vsync),
                             EDID mode list, runtime mode-set.

A backend is a small interface the compositor calls:

  • surface() → the pixels to compose into and their geometry {ptr, pitch, format, w, h} (the LFB for GOP; a shared scanout resource for virtio-gpu),
  • present(damage: Rect) → make the damaged region visible (a no-op-ish WC copy for GOP; a virtio flush, optionally vsync-fenced, for the native path),
  • capability queries — canModeSet, hasVsync — and, when supported, modes() / setMode(m).

The compositor composes into surface() and calls present(damage) exactly as it does today; everything device-specific lives behind the interface.

Selection and hot-attach

The choice is dynamic, because a GPU driver is spawned asynchronously (the device manager brings it up after boot), and because danos is meant to be resilient:

  1. Boot on GOP. The compositor starts on GopBackend immediately, so there is never a blank screen while drivers load — the exact v1 behaviour.
  2. Upgrade on announce. When the virtio-gpu driver has claimed its device and set up a scanout, it announces itself to the display service (a push: the driver looks up .display and sends an attach-scanout message carrying its scanout endpoint as a capability). The compositor switches to VirtioGpuBackend and re-presents the current frame full-screen. Push beats polling — the compositor doesn't know a priori which driver, if any, exists, and danos has no service-registration pub/sub.
  3. Native is restartable, not fallback-on-crash. Once a native driver has reprogrammed the device, the firmware's GOP framebuffer is stale — "native → GOP" is not a clean fall-back. So a native driver that crashes is restarted by its supervisor (the resilience work already merged), re-announces, and the compositor re-attaches (native → native). The screen freezes on the last frame during the gap — acceptable.
  4. GOP is the floor for "no driver was ever there." On a real GPU (NVIDIA/AMD/Intel) the class-0x03 device matches nothing in the driver table, no scanout is ever announced, and the compositor stays on GOP forever — no special-casing. Only if a native driver permanently gives up (crash-loop cap) does the compositor attempt GOP again, and even then only if the LFB is still mappable.

The shared-memory primitive this needs

virtio-gpu's scanout resource is guest RAM — the driver allocates it and attaches it to a virtio resource, and the compositor composes into it. That means the compositor writing into the driver's buffer is cross-process memory sharing, the primitive v1 deferred (docs/display.md, "What v1 does not do"). v2 builds it: the natural generalization of M13 capability-passing from endpoints to memory objects —

shm_create(len) -> {handle, vaddr}     // a shareable, page-aligned RAM region
… pass `handle` as the send_cap on an ipc_call …
shm_map(cap)    -> vaddr                // the receiver maps the same physical pages

The payoff is leverage: the same primitive unlocks both native GPU drivers and client-rendered surfaces (an app composing its own bitmap and handing the compositor a reference instead of drawing by command). One piece of kernel work, two features.

The virtio-gpu driver

A new ring-3 driver process (the topology v1 anticipated — "split the driver from the compositor when a second backend arrives"). It claims the virtio-gpu PCI function, and:

  • sets up the virtqueues (control + cursor) and the device's config space,
  • creates a 2D scanout resource backed by an shm region, attach_backings it, set_scanouts it to a CRTC, and resource_flushes damaged rectangles,
  • reads EDID (the GET_EDID control command) for the mode list, and set_scanout at a chosen mode for runtime mode-setting,
  • registers a scanout service and announces to the display service.

Its resource_flush is the real present — and gives a genuine vsync/tear-free path a dumb GOP framebuffer can't.

What v2 unlocks — and its honest scope

Behind the abstraction, a native backend gives runtime mode-setting (resolution / refresh / bpp), EDID enumeration, and vsync. But only on devices we have a driver for — realistically VMs (virtio-gpu, and later maybe Bochs DISPI). Real discrete GPUs need per-vendor KMS-class drivers that aren't getting written, so they stay on GOP — which is genuinely fine (v1 on the NVIDIA box is smooth). So v2's real value is twofold: the pluggable architecture (a driver slots in when one exists) and a rich, vsync'd path in VMs, where danos development happens. The framebuffer floor never goes away.

Locked decisions

  • First native backend: virtio-gpu — the VM standard; gives mode-set + a real present/flush (and vsync), and exercises the whole pluggable design. Tested with QEMU -device virtio-gpu.
  • Dynamic hot-attach — boot on GOP, upgrade to native on the driver's announce, re-attach across driver restarts; GOP is the floor for "no driver ever," not a live fall-back after a reprogram.
  • Detection = push (the driver announces to .display), not compositor polling.
  • v2 builds the shm capability (endpoints → memory objects), shared with the future client-surface path.

See also