Author SHA1 Message Date
Daniel Samson ec6e888076 display: pace the frame clock by the panel's EDID refresh rate
Both EDID moments the system has are now captured and carried to the
compositor's frame clock:

- EFI: the loader derives refresh from the preferred detailed timing
  (pixel clock / total pixels) while GOP is still alive — the only moment
  it is readable — and hands it through the boot handoff into the
  display0 node's DisplayInfo (new refresh_hz field, 0 = unknown).
- GPU: the virtio-gpu driver derives the same figure from its own EDID
  read and carries it in the attach_scanout announce (request.y).

updateFrameClock() re-derives the interval from the active backend's
info at bring-up and again on every backend change — the boot
framebuffer's clock dies with the GOP floor at upgrade, replaced by the
GPU's rate. Unknown rate defaults to 60 Hz; the result is clamped to
[30, 120] Hz so a mis-parsed EDID can neither starve nor flood the
compositor. Rate only, never phase: without vblank, presents still
free-run (docs/display-v2.md, 'Fenced is not vsync').

Observed in QEMU: OVMF exposes no EDID for the VGA adapter, so the GOP
floor logs 'frame clock 62 Hz (default)' (real firmware does expose it);
the virtio-gpu EDID advertises 75 Hz and the upgrade logs 'frame clock
76 Hz (panel EDID)'. The display kernel test asserts refresh_hz rides
the seeded node; all 7 display QEMU cases pass.
2026-07-21 11:37:29 +01:00
Daniel Samson 4f02f75602 display: a ~60 Hz frame clock — presents are scheduled, not immediate
Client 'present' requests and cursor pokes no longer repaint on the spot:
they accumulate damage and arm a one-shot 16 ms timer, and the tick
composites everything pending as one frame. A fast mouse previously
turned every input event into a full present (100+/s); now any number of
draws and moves inside one interval coalesce into a single repaint.

No backend has a real vblank to pace by (docs/display-v2.md, 'Fenced is
not vsync'), so this is the software stand-in — the same strategy Linux
uses atop virtio-gpu. Bring-up paths that need pixels synchronously
(initialise, self-checks) still present directly.

onNotification now handles the message and timer badge bits
independently: one coalesced badge can carry both, and the old
either/or dispatch would have dropped a tick.

All six display QEMU cases pass.
2026-07-21 11:26:49 +01:00
Daniel Samson 16618d2cdc display: stop calling the fenced present 'vsync' — it isn't
The virtio-gpu present fence completes when the device has consumed the
frame: real completion feedback, and tear-freedom by snapshot semantics.
It is not a vblank — base virtio-gpu 2D has no display-refresh event at
all (Linux fakes one with a timer), so nothing paces presents to the
monitor. The code and docs claimed vsync anyway; now they don't.

- backend.hasVsync -> hasFencedPresent, with an honest doc comment
- marker 'display: vsync present ok' -> 'display: fenced present ok'
  (display-modeset test expectation updated, passes)
- display-v2.md gains a 'Fenced is not vsync' note; the vsync claims in
  both v2 docs are corrected
- true vsync arrives with a native driver's vblank IRQ, or approximated
  by a compositor frame clock
2026-07-21 11:24:23 +01:00
Daniel Samson 15107f54be build: run-x86-64-gpu — the interactive twin of the display-native test
Boots the usual VGA/GOP floor plus a virtio-gpu-pci adapter, so the
device-manager stack spawns the driver and the compositor upgrades to the
native fenced-present backend interactively. QEMU shows one head per
adapter: the live output is the virtio-gpu console in the View menu (the
VGA head freezes at the moment of upgrade). 512M, matching display-native.
2026-07-21 11:22:07 +01:00
Daniel Samson 23f915c593 display: damage-rect list + tile-grid trackers, vectorizable pixel loops, wide WC stores
Tearing mitigation for the GOP floor, attacking the copy window from three sides:

- Damage is no longer one bounding box. Two trackers, A/B-switchable at
  compile time (display.zig damage_mode): DamageList (free-form dirty rects,
  overlap-merged) and TileGrid (fixed 64-px tiles, exact O(1) marking, runs
  coalesced back into rects). Far-apart changes — the cursor here, an
  animating layer there — no longer unite into one huge repaint.
- fillRect/composite/blitTile now work in row spans (@memset/@memcpy), so
  the compiler vectorizes them and ReleaseSafe bounds checks drop to per-row.
- The back->front present streams 8-byte volatile stores (presentSpan);
  Backend.present takes the rect list, so each present copies only what
  changed, faster.

Host tests cover both trackers; the display QEMU cases all pass.
2026-07-21 11:22:00 +01:00
daniel 981a4af7e0 Merge display-mouse-thread: the compositor tracks the mouse (Shape A)
The display becomes danos's first multi-threaded service. A mouse-listener
thread runs beside the compositor loop; the compositor stays the single owner
of the framebuffer, and the cursor it draws is a top-z layer moved from the
listener's accumulated position over a single-slot latest-value CursorChannel
(Thread.Mutex) with a coalesced self-ipc.send poke that wakes the parked
replyWait (docs/display.md, docs/threading.md).

Two kernel-level findings this surfaced, both fixed:
  - IPC handles are per-Task, not per-address-space — a thread reaches another
    thread's endpoint via ipc.lookup() to install its own handle.
  - Concurrent IPC from two threads raced unlocked kernel state (flaky #GP);
    create_ipc_endpoint / ipc_register / ipc_lookup now take the big kernel
    lock like call/reply_wait/send already did (the kernel heap has no lock of
    its own yet — heap.zig: "a lock comes with threads/SMP").

Also decouples display-demo: it no longer draws a cursor (the service owns it)
or blocks its animation loop on a mouse read, so a normal boot shows one
responsive cursor and animation that runs independent of the mouse. The
display-demo test now spawns the input service to keep that independence honest.

New test display-cursor (smp:4) drives synthetic motion end to end. Verified:
zig build, zig build test, and the full 29-case QEMU guardrail suite all green.
2026-07-21 02:41:45 +01:00
daniel 5ab7263c9c display-demo: stop drawing a cursor and stop blocking on the mouse
Two real bugs visible in a normal `zig build run-x86-64` boot (but not in the
display-demo test, which spawns no input service):

  - Two cursors. The demo drew its own cursor layer while the display service
    now draws one too (its mouse-listener thread). The demo's went through the
    client IPC protocol and lagged; the service's is in-process and tracks
    tightly — "one responds better than the other."

  - Animation frozen until the mouse moves. The demo called a *blocking*
    `mouse.next()` (ipc_reply_wait) inside its animation loop, so the sliding
    box advanced only one frame per mouse event. In the display-demo test
    there is no input service, so subscribeMouse() returned null and the loop
    ran free on its 30ms timer — which is exactly why the test passed while
    the real boot was broken.

The compositor now owns the cursor (docs/display.md), so the demo should draw
none and read no input: it becomes a pure client-animation proof whose loop is
independent of the mouse. Removes its cursor layer, mouse subscription, and the
blocking read.

Also harden displayDemoTest to spawn the `input` service alongside the demo
(matching real boot): a client that blocks its animation loop on a mouse read
would now stall before `display-demo: ok` and fail the test, instead of
passing because no input service happened to be present.

Verified: display / display-service / display-demo / display-cursor all green.
2026-07-21 02:37:47 +01:00
daniel 7f415e724f display: track the mouse with a listener thread (Shape A) + cursor
The display's first use of threads (docs/threading.md, docs/display.md). The
compositor stays the single owner of the framebuffer — only the main
service.run loop touches the backend and layer stack — and a dedicated
mouse-listener thread runs beside it:

  - Listener: blocks on input.subscribeMouse(), accumulates relative dx/dy
    into an absolute cursor position clamped to the screen, and hands it to
    the compositor. It never touches the compositor, so no lock guards the
    framebuffer; a parked next() lets the core halt.
  - CursorChannel: a single-slot latest-value cell under a Thread.Mutex (the
    renderer wants where the cursor is now, not a replay of deltas), with a
    coalesced self-ipc.send poke that wakes the main loop — parked in
    replyWait — as a message-notification. At most one poke is queued while
    the last is undrained, so a fast mouse can't flood the endpoint.
  - Render: the cursor is a top-z compositor layer; on the poke the main loop
    moves it via configure + present (which damages old + new footprints).

The display binary opts into threads (addThreadedUserBinary), and
input-source gains a "mouse" mode that publishes pure motion to drive it.

Two kernel-level findings this surfaced, both fixed:

  1. IPC handles do not cross threads. The handle table lives on the Task, so
     the listener can't reuse the main loop's endpoint handle — it
     ipc.lookup(.display)s its own handle to the same endpoint to poke through.

  2. Concurrent IPC from two threads raced unlocked kernel state. The display
     is the first process issuing IPC syscalls from two threads at once, which
     exposed a data race (flaky #GP in installEntry): create_ipc_endpoint /
     ipc_register / ipc_lookup allocate from the kernel heap and mutate the
     global registry, endpoint refcounts, and handle tables without the big
     kernel lock. They were safe only while a process couldn't race itself.
     They now sync.enter() like call/reply_wait/send already did (the kernel
     heap has no lock of its own yet — heap.zig: "a lock comes with
     threads/SMP" — so the big lock keeps its callers serialized).

Test: -Dtest-case=display-cursor (smp:4) spawns the input service, the
threaded display, and input-source in mouse mode; asserts the display's
"cursor tracking mouse ok" marker once the cursor has tracked a run of motion
end to end. Verified green 6/6 under stress (the race hit ~1-in-4 before the
lock fix) and in the full 29-case QEMU guardrail suite; zig build test clean.
2026-07-21 02:31:29 +01:00
daniel cf140eb772 Merge threading-phase2: threading Phase 2 (M7-M11) + naming cleanups
Completes the std.Thread-shaped runtime.Thread. Phase 1 (M1-M6: address-space
refcount, thread_spawn/exit, join/detach, futex + Mutex/Condition/Semaphore,
getCurrentId) was already on main; this brings the Phase 2 hardening and the
coding-standards/arch-neutrality passes on top:

  M7  thread-safe allocation (per-address-space mmap arena + locked heap)
  M8  the task reaper — reclaim dead tasks' kernel stacks
  M9  thread_join syscall — retire the per-thread endpoint
  M10 per-thread thread pointer — the TLS mechanism (x86_64 IA32_FS_BASE)
  M11 RwLock, WaitGroup, and host-testable sync

Plus: the TLS thread pointer named arch-neutrally (not fs.base) so the kernel
stays architecture-agnostic; aspace/vaddr/paddr spelled out per
docs/coding-standards.md across kernel, runtime, ABI, tests, and docs; and the
misleading fs.base dot-notation dropped in favour of "thread pointer"
(arch-neutral) / "FS base" (x86-specific).

Deferred by design (no consumer yet): the Zig threadlocal *compiler* layer
(M10) and detached-thread user-stack reclaim (M9) — both noted in place.

Verified: zig build, zig build test, and the full 25-case QEMU guardrail suite
all green.
2026-07-21 01:55:23 +01:00
20 changed files with 945 additions and 161 deletions
+18 -6
View File
@@ -97,7 +97,7 @@ fn boot() !noreturn {
}
/// A display resolution in pixels.
const Resolution = struct { width: u32, height: u32 };
const Resolution = struct { width: u32, height: u32, refresh_hz: u32 };
/// Switch the GPU to the monitor's native resolution (when we can determine it)
/// and read the resulting graphics mode into our own framebuffer description.
@@ -128,6 +128,10 @@ fn queryFramebuffer(bs: *uefi.tables.BootServices) !boot_handoff.Framebuffer {
// Each pixel is 32 bits, so the byte pitch is 4 * pixels-per-row.
.pitch = info.pixels_per_scan_line * 4,
.format = try pixelFormat(info.pixel_format),
// The refresh rate rides the EDID preferred timing. If the firmware kept a
// non-native mode it may not describe that mode exactly — but it is the panel's
// own clock, a far better frame-clock seed than a hardcoded 60 Hz.
.refresh_hz = if (native) |n| n.refresh_hz else 0,
};
}
@@ -176,10 +180,12 @@ fn nativeResolution(bs: *uefi.tables.BootServices, handles: []uefi.Handle) ?Reso
return null;
}
/// Parse the native resolution from a raw EDID block. The first Detailed Timing
/// Descriptor (at byte 54) is the preferred — i.e. native — mode by convention;
/// its active pixel counts are split across low bytes and the high nibbles of
/// later bytes.
/// Parse the native resolution and refresh rate from a raw EDID block. The first
/// Detailed Timing Descriptor (at byte 54) is the preferred — i.e. native — mode by
/// convention; its active pixel counts are split across low bytes and the high nibbles
/// of later bytes. The refresh rate is derived, not stored: the descriptor carries the
/// pixel clock (10 kHz units) and the active+blanking extents, and
/// refresh = clock / (horizontal total × vertical total).
fn edidNative(edid: []const u8) ?Resolution {
if (edid.len < 128) return null;
// Every EDID begins with this fixed 8-byte header.
@@ -193,7 +199,13 @@ fn edidNative(edid: []const u8) ?Resolution {
const w = @as(u32, dtd[2]) | (@as(u32, dtd[4] & 0xf0) << 4);
const h = @as(u32, dtd[5]) | (@as(u32, dtd[7] & 0xf0) << 4);
if (w == 0 or h == 0) return null;
return .{ .width = w, .height = h };
const clock_hz = (@as(u64, dtd[0]) | (@as(u64, dtd[1]) << 8)) * 10_000;
const h_blank = @as(u64, dtd[3]) | (@as(u64, dtd[4] & 0x0f) << 8);
const v_blank = @as(u64, dtd[6]) | (@as(u64, dtd[7] & 0x0f) << 8);
const total = (@as(u64, w) + h_blank) * (@as(u64, h) + v_blank);
const refresh: u32 = if (total == 0) 0 else @intCast((clock_hz + total / 2) / total);
return .{ .width = w, .height = h, .refresh_hz = refresh };
}
/// Open the kernel on the volume we booted from, read it into a pool buffer,
+50 -1
View File
@@ -535,7 +535,9 @@ pub fn build(b: *std.Build) void {
// The FAT filesystem server: mounts the block device and serves it into the VFS
// at /mnt/usb. Its engine (engine.zig / on-disk.zig) is imported relatively.
const fat_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "fat", "system/services/fat/fat.zig");
const display_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "display", "system/services/display/display.zig");
// Threaded: the display runs a mouse-listener thread alongside its compositor loop
// (docs/threading.md, docs/display.md), so it opts into real atomics/TLS.
const display_exe = addThreadedUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "display", "system/services/display/display.zig");
const display_demo_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "display-demo", "system/services/display-demo/display-demo.zig");
const virtio_gpu_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "virtio-gpu", "system/drivers/virtio-gpu/virtio-gpu.zig");
const shm_server_exe = addUserBinary(b, kernel_target, runtime_module, mmio_module, xkeyboard_config_module, acpi_ids_module, "shm-server", "system/services/shm-server/shm-server.zig");
@@ -848,6 +850,53 @@ pub fn build(b: *std.Build) void {
const run_efi_step = b.step("run-x86-64", "Boot the x86-64 kernel in QEMU (UEFI/OVMF); serial0 is logged to zig-out/qemu-test/run-x86-64-serial0-<timestamp>.log");
run_efi_step.dependOn(&run_efi.step);
// --- run-x86-64-gpu: the same boot plus a virtio-gpu adapter ---
// The VGA device still supplies the boot (GOP) framebuffer the compositor starts
// on; the virtio-gpu function is discovered by the device-manager stack, its
// driver announces a shared scanout, and the compositor upgrades off the GOP
// floor to fenced, tear-free native presents (docs/display-v2.md).
// This is the interactive twin of the `display-native` test case, and 512M
// matches it (the whole driver stack + the compositor's surfaces at once).
// QEMU shows one head per adapter: pick the virtio-gpu head in the View menu
// to watch the native output.
const run_gpu = b.addSystemCommand(&.{
"qemu-system-x86_64",
"-device",
"qemu-xhci,id=xhci",
"-device",
"usb-mouse,bus=xhci.0",
"-device",
"usb-kbd,bus=xhci.0",
"-machine",
"q35",
"-m",
"512M",
"-drive",
b.fmt("if=pflash,format=raw,readonly=on,file={s}", .{ovmf_code}),
});
run_gpu.addArg("-drive");
run_gpu.addPrefixedFileArg("if=pflash,format=raw,file=", vars_out);
run_gpu.addArg("-drive");
run_gpu.addPrefixedFileArg("if=none,id=bootusb,format=raw,file=", fat_image_serial);
run_gpu.addArgs(&.{
"-device",
"usb-storage,bus=xhci.0,drive=bootusb,removable=on,bootindex=0",
"-net",
"none",
"-vga",
"none",
"-device",
"VGA,edid=on,xres=1280,yres=720",
"-device",
"virtio-gpu-pci",
});
const gpu_serial_log = b.fmt("{s}/run-x86-64-gpu-serial0-{s}.log", .{ log_dir, timestamp(b) });
run_gpu.addArgs(&.{ "-serial", b.fmt("file:{s}", .{gpu_serial_log}) });
run_gpu.step.dependOn(&make_log_dir.step);
const run_gpu_step = b.step("run-x86-64-gpu", "Boot in QEMU with a virtio-gpu adapter: the compositor upgrades to fenced (tear-free) native presents; watch the virtio-gpu head in QEMU's View menu");
run_gpu_step.dependOn(&run_gpu.step);
// const run_cmd = b.addRunArtifact(exe);
// const run_step = b.step("run", "Run the app");
// run_step.dependOn(&run_cmd.step);
+9 -7
View File
@@ -6,7 +6,7 @@ lands on its own and ends in a **verifiable gate** — shaped for a `/loop` run,
## Locked decisions (do not relitigate)
- **First native backend = virtio-gpu** (VM standard: mode-set + present/flush + vsync).
- **First native backend = virtio-gpu** (VM standard: mode-set + fenced present/flush).
- **Dynamic hot-attach**: boot on GOP, upgrade to native when the driver **announces**
(push, not polling); re-attach across driver restarts; GOP is the floor for "no driver
ever," not a live fall-back after a reprogram.
@@ -46,7 +46,7 @@ Extract scanout from the compositor so today's path becomes one backend among fu
- [x] `system/services/display/backend.zig`: a `Backend` tagged union with `info()`,
`surface()` (the cacheable compose target), `present(damage)`, and capability flags
(`canModeSet`/`hasVsync`, both false for GOP).
(`canModeSet`/`hasFencedPresent`, both false for GOP).
- [x] The v1 GOP path is now `backend.Gop` (claims the `display` node, WC-maps the LFB,
keeps the cacheable back buffer, `present` = the damage-rect WC copy). display.zig
composes into `backend.surface()` and calls `backend.present(damage)` — no LFB or
@@ -123,7 +123,7 @@ confirm the composited frame landed (`display: native present verified`), while
ok` still fires — checked order-independently. Without `-device virtio-gpu-pci` nothing is
announced and it stays on GOP: the v1 `display-service`/`display-demo` gates pass unchanged.
## V5 — Mode-setting, EDID, and vsync ✅
## V5 — Mode-setting, EDID, and fenced presents ✅
- [x] The driver negotiates `VIRTIO_GPU_F_EDID` (when offered) and reads the monitor's EDID,
logging its preferred mode; it offers a small mode list over `.scanout` `get_modes`. The
@@ -131,14 +131,16 @@ announced and it stays on GOP: the v1 `display-service`/`display-demo` gates pas
scanout rectangle (no resource/surface churn) — a runtime resolution change. `runtime.display`
gains `modes()` / `setMode()` (display-protocol `get_modes`/`set_mode`, forwarded to the backend).
- [x] Every `resource_flush` is issued fenced (`VIRTIO_GPU_FLAG_FENCE`); the device signals the
fence when the frame is on screen, which the used-ring ack the synchronous present waits on
already gates — a tear-free present.
- [x] `backend.VirtioGpu` reports `canModeSet` / `hasVsync` = true.
fence when it has consumed the frame, which the used-ring ack the synchronous present waits
on already gates — a tear-free present. (Completion feedback, **not vblank**: base
virtio-gpu 2D has no display-refresh event, so nothing paces presents to the monitor —
see the "Fenced is not vsync" note in [display-v2.md](display-v2.md).)
- [x] `backend.VirtioGpu` reports `canModeSet` / `hasFencedPresent` = true.
**Gate (met):** the `display-modeset` case (reusing the display-native boot) upgrades to
virtio-gpu, queries the driver's modes, `setMode`s to a different resolution, and confirms the
change by reading the backend's geometry back (`display: mode set to {w}x{h}, verified`); the
fenced present path is exercised and confirmed (`display: vsync present ok`) — both from serial,
fenced present path is exercised and confirmed (`display: fenced present ok`) — both from serial,
passing 3/3. The driver also logs the EDID preferred mode (`virtio-gpu: EDID preferred mode …`).
## V6 — Resilience (restart + re-attach) + tests + docs ✅
+18 -10
View File
@@ -2,7 +2,7 @@
**Status: complete (V1–V6).** The compositor boots on the GOP framebuffer and, when a
virtio-gpu driver announces itself, hot-attaches a native backend over the shared `shm`
scanout surface — with runtime mode-setting, EDID, and fenced (vsync) presents, and it
scanout surface — with runtime mode-setting, EDID, and fenced presents, and it
re-attaches across driver restarts. All serial-gated (see [display-v2-plan.md](display-v2-plan.md)).
v1 ([display.md](display.md)) is a compositor that owns the **GOP framebuffer** — it
@@ -25,10 +25,10 @@ The compositor itself (layers, back buffer, damage) does not change. Only the la
scanout backend (selected at runtime — GOP by default, native when it appears)
│
├─ GopBackend the v1 path: WC copy back→front to the firmware LFB.
│ Always available. No mode-set, no vsync. THE FLOOR.
│ Always available. No mode-set, no present fence. THE FLOOR.
│
└─ VirtioGpuBackend talks to a virtio-gpu driver process over a `scanout`
service: present via a shared resource + flush (real vsync),
service: present via a shared resource + fenced flush,
EDID mode list, runtime mode-set.
```
@@ -37,8 +37,8 @@ A **backend** is a small interface the compositor calls:
- `surface()` → the pixels to compose into and their geometry `{ptr, pitch, format, w, h}`
(the LFB for GOP; a shared scanout resource for virtio-gpu),
- `present(damage: Rect)` → make the damaged region visible (a no-op-ish WC copy for GOP;
a virtio flush, optionally vsync-fenced, for the native path),
- capability queries — `canModeSet`, `hasVsync` — and, when supported, `modes()` /
a fenced virtio flush for the native path),
- capability queries — `canModeSet`, `hasFencedPresent` — and, when supported, `modes()` /
`setMode(m)`.
The compositor composes into `surface()` and calls `present(damage)` exactly as it does
@@ -98,23 +98,31 @@ compositor when a second backend arrives"). It claims the virtio-gpu PCI functio
at a chosen mode for **runtime mode-setting**,
- registers a `scanout` service and announces to the display service.
Its `resource_flush` is the real **present** — and gives a genuine **vsync/tear-free**
path a dumb GOP framebuffer can't.
Its `resource_flush` is the real **present** — and gives a **fenced, tear-free** path a
dumb GOP framebuffer can't.
**Fenced is not vsync.** The fence completes when the device has *consumed* the frame:
real completion feedback, and tear-freedom by snapshot semantics (the host displays
discrete transferred frames, never a half-written surface). It is **not** a vblank —
base virtio-gpu 2D has no display-refresh event at all (Linux's driver for this device
fakes one with a software timer), so nothing paces presents to the monitor's refresh.
Refresh-paced presents need either a native driver's vblank interrupt (delivered over
the existing IRQ-as-IPC path) or the compositor's own frame clock.
## What v2 unlocks — and its honest scope
Behind the abstraction, a native backend gives runtime **mode-setting** (resolution /
refresh / bpp), **EDID** enumeration, and **vsync**. But only on devices we have a driver
refresh / bpp), **EDID** enumeration, and **fenced presents**. But only on devices we have a driver
for — realistically **VMs** (virtio-gpu, and later maybe Bochs DISPI). Real discrete GPUs
need per-vendor KMS-class drivers that aren't getting written, so they **stay on GOP** —
which is genuinely fine (v1 on the NVIDIA box is smooth). So v2's real value is twofold:
the **pluggable architecture** (a driver slots in when one exists) and a **rich, vsync'd
the **pluggable architecture** (a driver slots in when one exists) and a **rich, fenced
path in VMs**, where danos development happens. The framebuffer floor never goes away.
## Locked decisions
- **First native backend: virtio-gpu** — the VM standard; gives mode-set + a real
present/flush (and vsync), and exercises the whole pluggable design. Tested with QEMU
present/flush (fenced), and exercises the whole pluggable design. Tested with QEMU
`-device virtio-gpu`.
- **Dynamic hot-attach** — boot on GOP, upgrade to native on the driver's announce,
re-attach across driver restarts; GOP is the floor for "no driver ever," not a live
+55 -6
View File
@@ -196,13 +196,21 @@ shell, a terminal, a cursor, and a wallpaper:
| `fill_rect` | fill a rectangle of a layer with a colour |
| `blit_tile` | copy a small client-supplied pixel tile into a layer (inline) |
| `damage` | mark a region of a layer dirty |
| `present` | composite dirty layers and flush to the screen |
| `present` | request a repaint: composited at the next frame-clock tick |
Text is intentionally *not* an operation — a client renders glyphs by blitting tiles
(the [PSF font](../system/kernel/font.psf) path the console already uses can move into a
client). Keeping the protocol to rectangles and tiles keeps the compositor small and the
policy in the client.
`present` is a *request*, not an immediate flush: the compositor runs a ~60 Hz **frame
clock** (a one-shot kernel timer re-armed on demand), and each tick composites all the
damage accumulated since the last one. Any number of client presents and cursor moves
inside one interval coalesce into a single repaint — the software stand-in for vblank
pacing on backends that have none (all of them today; see
[display-v2.md](display-v2.md), "Fenced is not vsync"). Bring-up paths that must put
pixels on screen synchronously (initialisation, the self-checks) bypass the clock.
## `runtime.display`
Clients speak the protocol through a new [`library/runtime/display.zig`](../library/runtime/runtime.zig),
@@ -211,6 +219,38 @@ with a boot-race retry): `display.info()`, a `Layer` handle with `fill` / `blitT
`damage`, and `present()`. Application code never issues the raw syscalls — it calls the
runtime, as with every other danos service.
## The cursor: a mouse-listener thread feeding the compositor
The compositor is the single owner of the framebuffer — only the main `service.run` loop
touches the backend and the layer stack. Tracking the mouse without breaking that
ownership is the display's first use of [threads](threading.md): the service is built
multi-threaded (`addThreadedUserBinary`) and, at startup, spawns a **mouse-listener
thread** beside the compositor loop.
- **Listener thread.** Blocks on the input service's mouse stream
(`input.subscribeMouse()`), accumulates the relative `dx`/`dy` motion into an absolute
cursor position clamped to the screen, and hands it to the compositor. It never touches
the compositor — so no lock guards the framebuffer. A parked `next()` leaves its core
free to halt ([halting.md](halting.md)).
- **The channel.** A single-slot *latest-value* cell (`CursorChannel`) guarded by a
`runtime.Thread.Mutex`: the renderer wants where the cursor *is now*, not a replay of
every delta, so a new position overwrites the old. The listener also **pokes** the
compositor awake — the main loop is parked in `replyWait`, so the listener posts a
zero-payload `ipc.send` to the compositor's endpoint, which arrives as a
message-notification ([ipc.md](ipc.md)). The poke is *coalesced*: at most one is queued
while the main loop has not drained the last, so a fast mouse cannot flood the endpoint.
- **Render.** On the poke, the main loop takes the latest position and moves the cursor —
which is just a top-z compositor layer — with the existing `configure` + `present` path
(it damages the old and new footprints, so only those two rectangles repaint).
Two threading facts shape this (both in [threading.md](threading.md)). IPC **handles do
not cross threads**, so the listener can't reuse the main loop's endpoint handle — it
`ipc.lookup(.display)`s its *own* handle to the same endpoint to poke through. And a
multi-threaded service doing concurrent IPC is why the kernel's endpoint-create / register
/ lookup syscalls now serialize under the big kernel lock. Shared fate applies: a fault in
the listener takes the whole display down, and the supervisor restarts the process
([resilience.md](resilience.md)).
## What v1 does not do (and why that's fine)
Two capabilities are deliberately out of the first cut. Neither reshapes anything above;
@@ -232,7 +272,7 @@ both are clean additions behind the interfaces v1 establishes.
## Verifying it
Three QEMU test cases ([tests.zig](../system/kernel/tests.zig), `python3
Four QEMU test cases ([tests.zig](../system/kernel/tests.zig), `python3
test/qemu_test.py <case>`), each layering on the last:
- **`display`** — the kernel handoff: the seeded `display` device is shaped correctly and
@@ -245,11 +285,20 @@ test/qemu_test.py <case>`), each layering on the last:
layer — logging `display: compositor self-check ok`.
- **`display-demo`** — the full pipeline from a separate process: the hardware-free
[`display-demo`](../system/services/display-demo/) client (the
[`input-source`](../system/services/input-source/) analog) drives layers — a wallpaper, a
sliding rectangle, a cursor — through the layer client API and heartbeats
[`input-source`](../system/services/input-source/) analog) drives layers — a wallpaper and
a sliding rectangle — through the layer client API and heartbeats
`display-demo: ok`, proving a frame travelled client → compositor → screen, exactly as
the [input test](input.md) proves an event travels source → service → subscriber. The
visible motion itself is a screenshot away via `zig build run-x86-64`.
the [input test](input.md) proves an event travels source → service → subscriber. It draws
no cursor and reads no input — the cursor is the service's own (below), and the demo
animates on its own frame timer, independent of the mouse (the test spawns `input`
alongside it to keep that independence honest). The visible motion itself is a screenshot
away via `zig build run-x86-64`.
- **`display-cursor`** — the mouse-listener thread end to end: with the `input` service up,
`input-source mouse` publishes pure motion, and the display's listener thread accumulates
it into a cursor position handed to the render loop over the `CursorChannel`. Once the
cursor has tracked a run of that motion, the service logs
`display: cursor tracking mouse ok`. Runs `smp: 4` — the compositor and listener threads
execute on different cores, which is what surfaced the IPC-under-lock requirement above.
The compositor's pixel math (rectangle clipping, fill, composite, tile blit) and colour
packing are additionally covered by pure host unit tests under `zig build test`.
+15
View File
@@ -249,6 +249,21 @@ it may call `runtime.Thread.spawn`. Everyone else stays single-threaded and lean
- **Resilience** ([resilience.md](resilience.md)): a faulting thread kills its whole
process (shared fate). The supervisor restarts the **process**, which respawns its
threads from a known-good state — restart granularity stays the process.
- **IPC — two consequences threads forced ([ipc.md](ipc.md)):**
- *Handles do not cross threads.* The handle table lives on the `Task`
([scheduler.zig](../system/kernel/scheduler.zig)), so a handle number is meaningful
only to the thread that created it — thread A's endpoint handle `3` is not thread B's.
A thread that needs to reach an endpoint another thread owns looks it up
(`ipc.lookup(service)`) to install its **own** handle to the same underlying endpoint.
This is how the display's mouse-listener thread reaches the compositor loop's endpoint
to poke it awake (docs/display.md).
- *IPC syscalls that touch shared kernel state now serialize under the big kernel lock.*
`create_ipc_endpoint`/`ipc_register`/`ipc_lookup` allocate from the kernel heap and
mutate the global service registry, endpoint refcounts, and handle tables. Those paths
were unlocked because a single-threaded process could not race itself; a multi-threaded
one can, from two cores at once. They now take `sync.enter()` like `call`/`reply_wait`/
`send` already did — the kernel heap has no lock of its own yet (heap.zig: "a lock comes
with threads/SMP"), so the big lock is what keeps its callers serialized.
## Build-out plan (staged, each gate serial-checkable)
+5
View File
@@ -37,6 +37,11 @@ pub const Framebuffer = extern struct {
height: u32, // visible rows (e.g. 1080)
pitch: u32, // bytes from the start of one row to the start of the next
format: PixelFormat,
/// The panel's refresh rate in Hz, computed from its EDID preferred timing (pixel
/// clock / total pixels per frame) while GOP was still alive — the one moment it is
/// readable (docs/gop.md). 0 = unknown (no EDID). The display service paces its
/// frame clock by it; without vblank this fixes the *rate*, never the *phase*.
refresh_hz: u32 = 0,
/// Whether a usable framebuffer was handed over.
pub fn present(self: Framebuffer) bool {
+1
View File
@@ -97,6 +97,7 @@ pub const DisplayInfo = extern struct {
height: u32 = 0, // visible rows
pitch: u32 = 0, // bytes from one row's start to the next
format: u32 = 0, // a DisplayFormat value
refresh_hz: u32 = 0, // panel refresh rate from EDID (0 = unknown); see boot-handoff
};
/// `DeviceDescriptor.parent` for a device with no parent — a root of the device tree.
+20 -5
View File
@@ -53,13 +53,18 @@ const offered_modes = [_]Mode{ .{ .width = 640, .height = 480 }, .{ .width = 800
var current_width: u32 = offered_modes[0].width;
var current_height: u32 = offered_modes[0].height;
/// Monotonic fence id for fenced (vsync) flushes; the device signals the fence when the flush
/// is complete, which its used-ring ack already gates our synchronous present on.
/// Monotonic fence id for fenced flushes; the device signals the fence when the flush is
/// complete, which its used-ring ack already gates our synchronous present on. Completion
/// feedback, not vblank — nothing here is paced to the display's refresh.
var fence_next: u64 = 1;
/// Whether the device offered VIRTIO_GPU_F_EDID, so `get_edid` is worth issuing.
var edid_available = false;
/// The panel refresh rate parsed from the EDID preferred timing (0 = unknown). Carried to
/// the compositor in the announce so its frame clock paces to the panel, not a guess.
var edid_refresh_hz: u32 = 0;
/// The control virtqueue. We drive it synchronously — one command, notify, poll the used
/// ring — so a depth of 16 is ample; we ask the device to shrink to it (virtio 1.0 lets the
/// driver reduce queue_size), keeping the whole ring inside one page.
@@ -467,10 +472,18 @@ fn readEdid() void {
}
// The first detailed timing descriptor (EDID base-block offset 54) is the preferred mode:
// active pixels are 12-bit, low byte + high nibble (bytes 2/4 horizontal, 5/7 vertical).
// The refresh rate is derived from the same descriptor: pixel clock (bytes 0-1, 10 kHz
// units) over total (active + blanking) pixels per frame — the loader does the identical
// computation for the boot framebuffer (boot/efi.zig edidNative).
const e = &response.edid;
const h_active = @as(u32, e[56]) | (@as(u32, e[58] & 0xF0) << 4);
const v_active = @as(u32, e[59]) | (@as(u32, e[61] & 0xF0) << 4);
log("virtio-gpu: EDID preferred mode {d}x{d}\n", .{ h_active, v_active });
const clock_hz = (@as(u64, e[54]) | (@as(u64, e[55]) << 8)) * 10_000;
const h_blank = @as(u64, e[57]) | (@as(u64, e[58] & 0x0F) << 8);
const v_blank = @as(u64, e[60]) | (@as(u64, e[61] & 0x0F) << 8);
const total = (@as(u64, h_active) + h_blank) * (@as(u64, v_active) + v_blank);
if (total != 0) edid_refresh_hz = @intCast((clock_hz + total / 2) / total);
log("virtio-gpu: EDID preferred mode {d}x{d} @ {d} Hz\n", .{ h_active, v_active, edid_refresh_hz });
}
/// Present the whole surface: copy the guest backing into the host resource, then flush it to
@@ -492,8 +505,9 @@ fn presentFull() bool {
if (command_nodata(@sizeOf(vg.TransferToHost2d)) != ok_nodata) return false;
}
{
// A fenced flush (vsync): the device signals the fence when the frame is actually on
// screen — which its used-ring ack, what our synchronous submit waits on, already gates.
// A fenced flush: the device signals the fence once it has consumed the frame — which
// its used-ring ack, what our synchronous submit waits on, already gates. Completion
// feedback and a tear-free snapshot, not vblank pacing.
const request = requestAt(vg.ResourceFlush);
request.* = .{
.hdr = .{ .type = @intFromEnum(vg.CmdType.resource_flush), .flags = vg.flag_fence, .fence_id = fence_next },
@@ -548,6 +562,7 @@ fn announce() void {
var request = dp.Request{
.operation = @intFromEnum(dp.Operation.attach_scanout),
.x = max_width, // the shared surface's row stride in pixels (it is sized to the max mode)
.y = edid_refresh_hz, // the panel refresh from EDID (0 = unknown) — the frame-clock seed
.width = current_width,
.height = current_height,
.colour = display_format_bgrx,
+2 -2
View File
@@ -61,7 +61,7 @@ pub fn init(device_tree: *const platform.DeviceTree) void {
/// [[boot-handoff]], not the device tree), so it is seeded explicitly, after `init`.
/// Returns the new device id, or null when there is no framebuffer (headless) or the
/// table is full. Idempotent-ish: only ever call once per boot.
pub fn seedDisplay(base: u64, width: u32, height: u32, pitch: u32, format: u32) ?u64 {
pub fn seedDisplay(base: u64, width: u32, height: u32, pitch: u32, format: u32, refresh_hz: u32) ?u64 {
if (base == 0 or width == 0 or height == 0) return null; // headless
if (count >= maximum_devices) {
dropped += 1;
@@ -79,7 +79,7 @@ pub fn seedDisplay(base: u64, width: u32, height: u32, pitch: u32, format: u32)
.len = @as(u64, height) * pitch,
.flags = device_abi.resource_flag_write_combining,
};
d.display = .{ .width = width, .height = height, .pitch = pitch, .format = format };
d.display = .{ .width = width, .height = height, .pitch = pitch, .format = format, .refresh_hz = refresh_hz };
devices[count] = d;
display_device = d.id;
count += 1;
+2 -2
View File
@@ -200,8 +200,8 @@ fn kmain(boot_information: *const BootInformation) noreturn {
// Publish the loader's framebuffer as a claimable `display` device, so a
// user-space display service can take it over the same claim + mmio_map path as
// any other hardware (it is not firmware-discovered; it rides the boot handoff).
if (devices_broker.seedDisplay(fb.base, fb.width, fb.height, fb.pitch, @intFromEnum(fb.format))) |display_id| {
log.print("/system/kernel: framebuffer device {d} seeded ({d}x{d}, pitch {d}, write-combining)\n", .{ display_id, fb.width, fb.height, fb.pitch });
if (devices_broker.seedDisplay(fb.base, fb.width, fb.height, fb.pitch, @intFromEnum(fb.format), fb.refresh_hz)) |display_id| {
log.print("/system/kernel: framebuffer device {d} seeded ({d}x{d}, pitch {d}, {d} Hz, write-combining)\n", .{ display_id, fb.width, fb.height, fb.pitch, fb.refresh_hz });
}
// Install the device-IRQ trampolines, so a driver's irq_bind has vectors to
+16
View File
@@ -256,6 +256,13 @@ fn failErr(state: *architecture.CpuState, errno: i64) void {
/// create_ipc_endpoint() -> handle: allocate an endpoint and install it in the
/// caller's handle table.
fn systemCreateIpcEndpoint(state: *architecture.CpuState) void {
// Under the big kernel lock: this allocates from the kernel heap and mutates the
// caller's handle table. A multi-threaded process (e.g. the display's compositor +
// mouse-listener threads) can drive this concurrently from two cores, so the endpoint
// allocation and every other lock holder must serialize (heap.zig: "a lock comes with
// threads/SMP").
const flags = sync.enter();
defer sync.leave(flags);
const endpoint = ipc.createIpcEndpoint() orelse return failErr(state, ipc.ENOMEM);
const h = ipc.installHandle(scheduler.current(), endpoint);
if (h < 0) {
@@ -268,6 +275,10 @@ fn systemCreateIpcEndpoint(state: *architecture.CpuState) void {
/// ipc_register(service_id, handle): publish the caller's endpoint under a
/// well-known id so other processes can find it.
fn systemIpcRegister(state: *architecture.CpuState) void {
// Under the big kernel lock: mutates the global service registry and endpoint
// refcounts, which threads of the same (or another) process can race.
const flags = sync.enter();
defer sync.leave(flags);
const id: u32 = @truncate(architecture.systemCallArg(state, 0));
const endpoint = ipc.resolveHandle(scheduler.current(), architecture.systemCallArg(state, 1)) orelse return failErr(state, ipc.EBADF);
architecture.setSystemCallResult(state, @bitCast(ipc.register(id, endpoint)));
@@ -276,6 +287,11 @@ fn systemIpcRegister(state: *architecture.CpuState) void {
/// ipc_lookup(service_id) -> handle: find a published endpoint and install a
/// handle to it in the caller.
fn systemIpcLookup(state: *architecture.CpuState) void {
// Under the big kernel lock: reads the global registry, takes an endpoint reference,
// and installs a handle — all racy against concurrent threads (this is the path the
// display's mouse-listener thread takes to reach the compositor endpoint).
const flags = sync.enter();
defer sync.leave(flags);
const id: u32 = @truncate(architecture.systemCallArg(state, 0));
const endpoint = ipc.lookup(id) orelse return failErr(state, ipc.ENOENT);
const h = ipc.installHandle(scheduler.current(), endpoint);
+61
View File
@@ -101,6 +101,8 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
displayServiceTest(boot_information);
} else if (eql(case, "display-demo")) {
displayDemoTest(boot_information);
} else if (eql(case, "display-cursor")) {
displayCursorTest(boot_information);
} else if (eql(case, "shm")) {
shmTest(boot_information);
} else if (eql(case, "virtio-gpu")) {
@@ -2782,6 +2784,46 @@ fn displayServiceTest(boot_information: *const BootInformation) void {
while (true) scheduler.yield();
}
/// The threaded compositor tracks a mouse (docs/threading.md, docs/display.md). Spawn the
/// `input` fan-out service, the display (which runs a mouse-listener thread alongside its
/// compositor loop and draws a top-z cursor), and `input-source` in `mouse` mode — a
/// synthetic source publishing pure motion. The display's own marker,
/// `display: cursor tracking mouse ok`, is printed once the cursor has tracked a run of
/// motion end to end (source -> input service -> listener thread -> channel -> render), so
/// like the other display cases we match on serial rather than poll in-kernel.
fn displayCursorTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: display-cursor\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
if (!spawnNamed(rd, "input")) {
log("display-cursor: could not spawn the input service\n", .{});
result();
return;
}
if (!spawnNamed(rd, "display")) {
log("display-cursor: could not spawn the display service\n", .{});
result();
return;
}
if (!spawnNamedWithArg(rd, "input-source", "mouse")) {
log("display-cursor: could not spawn the mouse source\n", .{});
result();
return;
}
scheduler.setPriority(1); // below the services, so they run
while (true) scheduler.yield();
}
/// D4 — a separate process drives the compositor. Spawn the display service and the
/// hardware-free `display-demo` client, which creates a wallpaper, a moving rectangle,
/// and a cursor and presents a run of frames. Its `display-demo: ok` heartbeat — printed
@@ -2808,6 +2850,11 @@ fn displayDemoTest(boot_information: *const BootInformation) void {
result();
return;
}
// Spawn the input service too — real boot has it, and it guards the demo's
// independence from input: the demo must animate to `display-demo: ok` on its own
// frame timer even with the input service available (a client that blocks its
// animation loop on a mouse read would stall here, never reaching the marker).
_ = spawnNamed(rd, "input");
_ = spawnNamed(rd, "display-demo");
scheduler.setPriority(1); // below the service + demo, so they run
while (true) scheduler.yield();
@@ -3029,6 +3076,19 @@ fn spawnNamed(rd: initial_ramdisk.Reader, name: []const u8) bool {
return false;
}
/// As `spawnNamed`, but passes one extra argv entry (argv[1]) — e.g. a mode selector like
/// `input-source mouse`.
fn spawnNamedWithArg(rd: initial_ramdisk.Reader, name: []const u8, arg: []const u8) bool {
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (eql(item.name, name)) {
return if (process.spawnProcess(item.blob, 4, &.{ item.name, arg })) true else |_| false;
}
}
return false;
}
/// The GSI discovery recorded for the HPET, from the same device table drivers see.
fn hpetGsi() ?u32 {
var buffer: [16]device_abi.DeviceDescriptor = undefined;
@@ -3344,6 +3404,7 @@ fn displayTest(boot_information: *const BootInformation) void {
check("the node is class display", d.class == @intFromEnum(device_abi.DeviceClass.display));
check("it carries the framebuffer geometry", d.display.width == fb.width and d.display.height == fb.height and d.display.pitch == fb.pitch);
check("it carries the panel refresh rate", d.display.refresh_hz == fb.refresh_hz);
check("it has exactly one resource", d.resource_count == 1);
const r = d.resources[0];
check("that resource is a memory window", r.kind == @intFromEnum(device_abi.ResourceKind.memory));
+9 -32
View File
@@ -1,15 +1,19 @@
//! system/services/display-demo — a hardware-free client of the display service, the
//! `input-source` analog for the compositor. It creates a wallpaper, a rectangle it moves
//! each frame, and a small cursor, then drives the compositor in a present loop — proof
//! that a *separate process* can compose a moving scene through the display service over
//! IPC, exercising the layer client API and damage-driven present end to end
//! `input-source` analog for the compositor. It creates a wallpaper and a rectangle it
//! slides each frame, then drives the compositor in a present loop — proof that a
//! *separate process* can compose a moving scene through the display service over IPC,
//! exercising the layer client API and damage-driven present end to end
//! (docs/display.md). It logs `display-demo: ok` once it has driven a run of frames.
//!
//! It draws no cursor and reads no input: the on-screen cursor is the display service's
//! own, tracked by the service's mouse-listener thread (docs/display.md). The demo's job
//! is only to prove client-driven animation, so its loop runs on its own frame timer and
//! is deliberately independent of the mouse.
const runtime = @import("runtime");
const display = runtime.display;
const system = runtime.system;
const time = runtime.time;
const input = runtime.input;
pub fn main() void {
const mode = display.info() orelse {
@@ -28,15 +32,6 @@ pub fn main() void {
const box = display.createLayer(0, box_y, box_w, box_h, 1) orelse return createFailed();
_ = box.fill(0, 0, box_w, box_h, display.color(0xE0, 0x60, 0x40));
// A little cursor on top. Its position is signed (the layer API is i32) and clamped to
// the screen; mouse motion arrives as relative deltas we accumulate below.
var cursor_x: i32 = @intCast(mode.width / 2);
var cursor_y: i32 = @intCast(mode.height / 2);
const cursor_max_x: i32 = @as(i32, @intCast(mode.width)) - 12;
const cursor_max_y: i32 = @as(i32, @intCast(mode.height)) - 12;
const cursor = display.createLayer(cursor_x, cursor_y, 12, 12, 2) orelse return createFailed();
_ = cursor.fill(0, 0, 12, 12, display.color(0xF0, 0xF0, 0xF0));
_ = display.present();
_ = system.write("display-demo: scene up; animating\n");
@@ -45,18 +40,7 @@ pub fn main() void {
var dx: i32 = 8;
var frame: u32 = 0;
var mouse = input.subscribeMouse(); // type: ?input.MouseSubscriber
if (mouse == null) _ = system.write("display-demo: no mouse; animating without it\n");
while (true) : (frame += 1) {
if (mouse) |*ms| {
if (ms.next()) |event| {
cursor_x = clamp(cursor_x + event.dx, 0, cursor_max_x);
cursor_y = clamp(cursor_y + event.dy, 0, cursor_max_y);
_ = cursor.configure(cursor_x, cursor_y, 2, true);
}
}
x += dx;
if (x <= 0) {
x = 0;
@@ -74,13 +58,6 @@ pub fn main() void {
}
}
/// Clamp `v` to the inclusive range [lo, hi].
fn clamp(v: i32, lo: i32, hi: i32) i32 {
if (v < lo) return lo;
if (v > hi) return hi;
return v;
}
fn createFailed() void {
_ = system.write("display-demo: create failed\n");
}
+59 -23
View File
@@ -16,8 +16,10 @@ const scanout_protocol = runtime.scanout_protocol;
const Rect = compositor.Rect;
const Surface = compositor.Surface;
/// The current display mode, as a backend reports it.
pub const Info = struct { width: u32, height: u32, pitch: u32, format: u32 };
/// The current display mode, as a backend reports it. `refresh_hz` is the panel's
/// refresh rate from EDID (0 = unknown) — the frame clock's pacing seed; without vblank
/// it fixes the rate, never the phase (docs/display-v2.md, "Fenced is not vsync").
pub const Info = struct { width: u32, height: u32, pitch: u32, format: u32, refresh_hz: u32 };
/// Enumeration scratch — a `DeviceDescriptor` is large, and only one scan is ever needed.
var device_table: [64]device.DeviceDescriptor = undefined;
@@ -26,7 +28,7 @@ var device_table: [64]device.DeviceDescriptor = undefined;
/// framebuffer write-combining as the front buffer, and keeps a cacheable back buffer of
/// the same geometry as the compose target. `present` streams the damaged rectangle from
/// the back buffer to the LFB (sequential WC writes; the LFB is never read). No mode-set,
/// no vsync — the portable floor (docs/display-v2.md).
/// no present fence — the portable floor (docs/display-v2.md).
pub const Gop = struct {
device_id: u64,
front: [*]volatile u8, // the LFB (write-combining)
@@ -35,13 +37,14 @@ pub const Gop = struct {
height: u32,
pitch: u32,
format: u32,
refresh_hz: u32, // from the boot EDID via the display0 node (0 = unknown)
/// The framebuffer's id and geometry, captured together. `findDisplay` reads these out of
/// the enumeration table and returns them by value, so the caller never re-reads the table
/// across later syscalls (`device_enumerate` writes the whole table straight into this
/// process's memory; reading a descriptor's tail again after other syscalls have run is a
/// window we simply avoid by copying the few fields we need up front).
const Found = struct { id: u64, width: u32, height: u32, pitch: u32, format: u32 };
const Found = struct { id: u64, width: u32, height: u32, pitch: u32, format: u32, refresh_hz: u32 };
/// The first `display`-class device with a *valid* (non-zero) geometry, or null. A zero
/// geometry is treated as "not ready yet" so the caller retries — a real framebuffer always
@@ -52,7 +55,7 @@ pub const Gop = struct {
for (device_table[0..n]) |*d| {
if (d.class != @intFromEnum(device.DeviceClass.display)) continue;
if (d.display.width == 0 or d.display.height == 0 or d.display.pitch == 0) continue;
return .{ .id = d.id, .width = d.display.width, .height = d.display.height, .pitch = d.display.pitch, .format = d.display.format };
return .{ .id = d.id, .width = d.display.width, .height = d.display.height, .pitch = d.display.pitch, .format = d.display.format, .refresh_hz = d.display.refresh_hz };
}
return null;
}
@@ -93,11 +96,12 @@ pub const Gop = struct {
.height = found.height,
.pitch = found.pitch,
.format = found.format,
.refresh_hz = found.refresh_hz,
};
}
pub fn info(self: *const Gop) Info {
return .{ .width = self.width, .height = self.height, .pitch = self.pitch, .format = self.format };
return .{ .width = self.width, .height = self.height, .pitch = self.pitch, .format = self.format, .refresh_hz = self.refresh_hz };
}
/// The cacheable compose target (the back buffer).
@@ -110,22 +114,48 @@ pub const Gop = struct {
};
}
/// Stream the damaged rectangle from the back buffer to the write-combining LFB, row by
/// row (sequential writes — what WC memory wants; the LFB is never read).
pub fn present(self: *const Gop, damage: Rect) void {
const c = damage.intersect(.{ .x = 0, .y = 0, .w = @intCast(self.width), .h = @intCast(self.height) });
if (c.isEmpty()) return;
/// Stream each damaged rectangle from the back buffer to the write-combining LFB, row
/// by row (sequential writes — what WC memory wants; the LFB is never read). The rows
/// are copied by `presentSpan` below, which widens the stores by hand: `volatile`
/// keeps the compiler from eliding or reordering framebuffer writes, but it also
/// forbids it from merging them, so a naive per-pixel loop is stuck at one 4-byte
/// store per iteration. Keeping each copy small (the damage list) and each store wide
/// shrinks the window in which scanout can sample a half-written frame.
pub fn present(self: *const Gop, damage: []const Rect) void {
const bounds = Rect{ .x = 0, .y = 0, .w = @intCast(self.width), .h = @intCast(self.height) };
for (damage) |rect| {
const c = rect.intersect(bounds);
if (c.isEmpty()) continue;
const span: usize = @intCast(c.w);
var y: i32 = c.y;
while (y < c.bottom()) : (y += 1) {
const off = @as(usize, @intCast(y)) * self.pitch;
const src: [*]const u32 = @ptrCast(@alignCast(self.back + off));
const dst: [*]volatile u32 = @ptrCast(@alignCast(self.front + off));
var x: i32 = c.x;
while (x < c.right()) : (x += 1) dst[@intCast(x)] = src[@intCast(x)];
const offset = @as(usize, @intCast(y)) * self.pitch + @as(usize, @intCast(c.x)) * 4;
const source: [*]const u32 = @ptrCast(@alignCast(self.back + offset));
const front_row: [*]volatile u32 = @ptrCast(@alignCast(self.front + offset));
presentSpan(front_row, source, span);
}
}
}
};
/// Copy `count` pixels into the write-combining front buffer with 8-byte volatile stores
/// (plus a 4-byte head/tail where the span isn't 8-aligned — pixel spans are always
/// 4-aligned). The loads come from the cacheable back buffer and are assembled into a
/// `u64` in registers, so nothing here reads the front buffer.
fn presentSpan(destination: [*]volatile u32, source: [*]const u32, count: usize) void {
var i: usize = 0;
if (i < count and (@intFromPtr(destination) & 7) != 0) {
destination[0] = source[0];
i = 1;
}
while (i + 2 <= count) : (i += 2) {
const pair = @as(u64, source[i]) | (@as(u64, source[i + 1]) << 32);
const wide: *volatile u64 = @ptrCast(@alignCast(destination + i));
wide.* = pair;
}
if (i < count) destination[i] = source[i];
}
/// A display mode the native backend can switch to.
pub const Mode = scanout_protocol.Mode;
@@ -143,17 +173,19 @@ pub const VirtioGpu = struct {
width: u32, // the active mode
height: u32,
format: u32,
refresh_hz: u32, // from the driver's EDID read, carried in the announce (0 = unknown)
scanout: ipc.Handle, // the driver's present + mode channel (looked up on `.scanout`)
pub fn info(self: *const VirtioGpu) Info {
return .{ .width = self.width, .height = self.height, .pitch = self.stride * 4, .format = self.format };
return .{ .width = self.width, .height = self.height, .pitch = self.stride * 4, .format = self.format, .refresh_hz = self.refresh_hz };
}
pub fn surface(self: *const VirtioGpu) Surface {
return .{ .pixels = self.pixels, .stride = self.stride, .width = self.width, .height = self.height };
}
/// Ask the driver to present. The composited pixels are already in the shared surface, so
/// this is a single request over `.scanout`; the driver transfers + fenced-flushes.
pub fn present(self: *const VirtioGpu, damage: Rect) void {
/// this is a single request over `.scanout` regardless of how many damage rectangles
/// accumulated; the driver transfers + fenced-flushes the whole frame.
pub fn present(self: *const VirtioGpu, damage: []const Rect) void {
_ = damage;
var request = scanout_protocol.Request{
.operation = @intFromEnum(scanout_protocol.Operation.present),
@@ -210,7 +242,7 @@ pub const Backend = union(enum) {
inline else => |*b| b.surface(),
};
}
pub fn present(self: *const Backend, damage: Rect) void {
pub fn present(self: *const Backend, damage: []const Rect) void {
switch (self.*) {
inline else => |*b| b.present(damage),
}
@@ -236,9 +268,13 @@ pub const Backend = union(enum) {
.virtio => true,
};
}
/// Whether this backend has a vblank/fence for tear-free present (virtio-gpu: yes, V5 — every
/// flush is fenced, so the device signals completion when the frame is actually on screen).
pub fn hasVsync(self: *const Backend) bool {
/// Whether this backend's present is **fenced** — it completes only once the device has
/// consumed the frame (virtio-gpu: every flush carries a fence the used-ring ack waits on).
/// A fence gives completion feedback and tear-free snapshot presents; it is *not* vblank —
/// nothing paces presents to the display's refresh (base virtio-gpu 2D has no vblank event
/// at all). True vsync needs a native driver's vblank interrupt. See docs/display-v2.md,
/// "Fenced is not vsync".
pub fn hasFencedPresent(self: *const Backend) bool {
return switch (self.*) {
.gop => false,
.virtio => true,
+285 -25
View File
@@ -57,6 +57,177 @@ pub const Rect = struct {
}
};
/// The dirty screen regions accumulated between presents. Kept as a *list* of rectangles,
/// not one bounding box: when two small things move far apart — the cursor on one side of
/// the screen, an animating layer on the other — a single bounding box unites them into a
/// huge region, and presenting it streams megabytes to the framebuffer for a few thousand
/// changed pixels. The long copy widens the window in which scanout (or QEMU's display
/// refresh) samples a half-written frame — visible as tearing and cursor trails. Small
/// separate rectangles keep each copy, and that window, tight.
///
/// A new rectangle that overlaps an existing entry is united into it (repainting a modest
/// superset is harmless — compositing is idempotent); the grown entry is *not* re-merged
/// against the rest, so entries may overlap, which costs only a duplicate repaint. When
/// the table is full the newcomer folds into the last entry — degrading toward the old
/// bounding-box behaviour instead of dropping damage.
pub const DamageList = struct {
pub const capacity = 16;
rects: [capacity]Rect = [_]Rect{Rect.empty} ** capacity,
count: usize = 0,
pub fn add(self: *DamageList, r: Rect) void {
if (r.isEmpty()) return;
for (self.rects[0..self.count]) |*existing| {
if (!existing.intersect(r).isEmpty()) {
existing.* = existing.unite(r);
return;
}
}
if (self.count < capacity) {
self.rects[self.count] = r;
self.count += 1;
return;
}
self.rects[capacity - 1] = self.rects[capacity - 1].unite(r);
}
pub fn isEmpty(self: *const DamageList) bool {
return self.count == 0;
}
pub fn slice(self: *const DamageList) []const Rect {
return self.rects[0..self.count];
}
pub fn clear(self: *DamageList) void {
self.count = 0;
}
};
/// The alternative damage tracker: a **fixed tile grid**, the scheme browser compositors
/// and tile-based GPUs use. The screen is divided into `tile_size`-pixel tiles up front;
/// `add` marks the tiles a rectangle touches (a bit per tile — merging is free and exact,
/// no heuristics), and `collect` walks the grid turning runs of adjacent dirty tiles into
/// repaint rectangles (horizontal runs, then equal-span rows merged vertically, so
/// full-screen damage collapses back to a single rectangle).
///
/// Trade-off against `DamageList`: tracking is O(1) with a strictly bounded worst case
/// (never more than the dirty tiles), but repaints are quantized — a 1-pixel change
/// repaints a whole tile. Which wins depends on the workload; the display service has a
/// compile-time switch (`damage_mode`) to compare them.
pub const TileGrid = struct {
pub const tile_size = 64;
pub const maximum_columns = 128; // supports screens up to 8192 px wide…
pub const maximum_rows = 128; // …and 8192 px tall (beyond that, edge tiles stretch)
pub const maximum_tiles = maximum_columns * maximum_rows;
/// The most rectangles `collect` produces; extras fold into the last (never dropped).
pub const maximum_rects = 64;
width: u32 = 0,
height: u32 = 0,
columns: u32 = 0,
rows: u32 = 0,
dirty_count: u32 = 0,
dirty: [maximum_tiles]bool = [_]bool{false} ** maximum_tiles,
/// Size the grid for a screen. Also clears it — callers reset on a geometry change,
/// where the mode-set paths damage the whole new screen anyway.
pub fn reset(self: *TileGrid, width: u32, height: u32) void {
self.width = width;
self.height = height;
self.columns = @min((width + tile_size - 1) / tile_size, maximum_columns);
self.rows = @min((height + tile_size - 1) / tile_size, maximum_rows);
self.clear();
}
pub fn matches(self: *const TileGrid, width: u32, height: u32) bool {
return self.width == width and self.height == height;
}
pub fn isEmpty(self: *const TileGrid) bool {
return self.dirty_count == 0;
}
pub fn clear(self: *TileGrid) void {
@memset(&self.dirty, false);
self.dirty_count = 0;
}
/// Mark every tile `r` touches. Clips to the screen first, so out-of-range
/// rectangles are harmless.
pub fn add(self: *TileGrid, r: Rect) void {
const screen = Rect{ .x = 0, .y = 0, .w = @intCast(self.width), .h = @intCast(self.height) };
const c = r.intersect(screen);
if (c.isEmpty()) return;
const column_first: u32 = @intCast(@divTrunc(c.x, tile_size));
const row_first: u32 = @intCast(@divTrunc(c.y, tile_size));
const column_last: u32 = @min(@as(u32, @intCast(@divTrunc(c.right() - 1, tile_size))), self.columns - 1);
const row_last: u32 = @min(@as(u32, @intCast(@divTrunc(c.bottom() - 1, tile_size))), self.rows - 1);
var row = row_first;
while (row <= row_last) : (row += 1) {
var column = column_first;
while (column <= column_last) : (column += 1) {
const index = row * self.columns + column;
if (!self.dirty[index]) {
self.dirty[index] = true;
self.dirty_count += 1;
}
}
}
}
/// The screen rectangle covered by tiles [column_first, column_end) of `row`. Edge
/// tiles clamp to the true screen size (the last column/row may be partial — or, on a
/// screen wider than the grid supports, stretched to cover the remainder).
fn tileSpanRect(self: *const TileGrid, column_first: u32, column_end: u32, row: u32) Rect {
const x: i32 = @intCast(column_first * tile_size);
const y: i32 = @intCast(row * tile_size);
const right: i32 = if (column_end >= self.columns) @intCast(self.width) else @intCast(column_end * tile_size);
const bottom: i32 = if (row + 1 >= self.rows) @intCast(self.height) else @intCast((row + 1) * tile_size);
return .{ .x = x, .y = y, .w = right - x, .h = bottom - y };
}
/// Turn the dirty tiles into repaint rectangles in `out`: coalesce each row's runs of
/// adjacent dirty tiles, then merge a run into the rectangle directly above it when
/// the spans match — so a dirty block of tiles becomes one rectangle. Returns the
/// filled prefix of `out`.
pub fn collect(self: *const TileGrid, out: []Rect) []Rect {
var count: usize = 0;
var row: u32 = 0;
while (row < self.rows) : (row += 1) {
var column: u32 = 0;
while (column < self.columns) {
if (!self.dirty[row * self.columns + column]) {
column += 1;
continue;
}
var run_end = column + 1;
while (run_end < self.columns and self.dirty[row * self.columns + run_end]) run_end += 1;
const rect = self.tileSpanRect(column, run_end, row);
column = run_end;
var merged = false;
for (out[0..count]) |*existing| {
if (existing.x == rect.x and existing.w == rect.w and existing.bottom() == rect.y) {
existing.h += rect.h;
merged = true;
break;
}
}
if (merged) continue;
if (count < out.len) {
out[count] = rect;
count += 1;
} else {
out[count - 1] = out[count - 1].unite(rect);
}
}
}
return out[0..count];
}
};
/// A block of 32-bit pixels: `pixels` addressed row-major with `stride` pixels between
/// row starts (≥ width — the framebuffer's stride is pitch/4, a layer's is its width).
pub const Surface = struct {
@@ -74,15 +245,17 @@ pub const Surface = struct {
}
};
/// Fill `rect` of `s` with the native pixel `colour`, clipped to `s`'s bounds.
/// Fill `rect` of `s` with the native pixel `colour`, clipped to `s`'s bounds. Each row is
/// one `@memset` over the clipped span, so the compiler vectorizes it and the bounds check
/// runs once per row, not once per pixel.
pub fn fillRect(s: Surface, rect: Rect, colour: u32) void {
const c = rect.intersect(s.bounds());
if (c.isEmpty()) return;
const x0: usize = @intCast(c.x);
const span: usize = @intCast(c.w);
var y: i32 = c.y;
while (y < c.bottom()) : (y += 1) {
const r = s.row(@intCast(y));
var x: i32 = c.x;
while (x < c.right()) : (x += 1) r[@intCast(x)] = colour;
@memset((s.row(@intCast(y)) + x0)[0..span], colour);
}
}
@@ -94,35 +267,36 @@ pub fn composite(dst: Surface, dx: i32, dy: i32, layer: Surface, clip: Rect) voi
const on_screen = Rect{ .x = dx, .y = dy, .w = @intCast(layer.width), .h = @intCast(layer.height) };
const region = on_screen.intersect(clip).intersect(dst.bounds());
if (region.isEmpty()) return;
const span: usize = @intCast(region.w);
const dst_x: usize = @intCast(region.x);
const src_x: usize = @intCast(region.x - dx);
var y: i32 = region.y;
while (y < region.bottom()) : (y += 1) {
const src = layer.row(@intCast(y - dy));
const d = dst.row(@intCast(y));
var x: i32 = region.x;
while (x < region.right()) : (x += 1) {
d[@intCast(x)] = src[@intCast(x - dx)];
}
const source_row = layer.row(@intCast(y - dy)) + src_x;
const destination_row = dst.row(@intCast(y)) + dst_x;
@memcpy(destination_row[0..span], source_row[0..span]);
}
}
/// Copy a `w`×`h` tile of native pixels from `src` (raw little-endian bytes, row-major,
/// tightly packed) into `dst` at (`dx`, `dy`), clipped to `dst`'s bounds. `src` is read
/// with `readInt` because it comes straight out of an IPC message buffer and carries no
/// alignment guarantee. Returns without touching anything if `src` is short.
/// tightly packed) into `dst` at (`dx`, `dy`), clipped to `dst`'s bounds. `src` comes
/// straight out of an IPC message buffer and carries no alignment guarantee, so each
/// clipped row is a byte-wise `@memcpy` — which equals the old per-pixel little-endian
/// `readInt` on every danos target (all little-endian) without the alignment concern.
/// Returns without touching anything if `src` is short.
pub fn blitTile(dst: Surface, dx: i32, dy: i32, src: []const u8, w: u32, h: u32) void {
if (src.len < @as(usize, w) * h * 4) return;
var ty: u32 = 0;
while (ty < h) : (ty += 1) {
const yy = dy + @as(i32, @intCast(ty));
if (yy < 0 or yy >= dst.height) continue;
const drow = dst.row(@intCast(yy));
var tx: u32 = 0;
while (tx < w) : (tx += 1) {
const xx = dx + @as(i32, @intCast(tx));
if (xx < 0 or xx >= dst.width) continue;
const off = (@as(usize, ty) * w + tx) * 4;
drow[@intCast(xx)] = std.mem.readInt(u32, src[off..][0..4], .little);
}
const region = Rect.init(dx, dy, @intCast(w), @intCast(h)).intersect(dst.bounds());
if (region.isEmpty()) return;
const span: usize = @intCast(region.w);
const tile_x: usize = @intCast(region.x - dx);
const dst_x: usize = @intCast(region.x);
var y: i32 = region.y;
while (y < region.bottom()) : (y += 1) {
const tile_y: usize = @intCast(y - dy);
const offset = (tile_y * w + tile_x) * 4;
const destination_row = dst.row(@intCast(y)) + dst_x;
@memcpy(std.mem.sliceAsBytes(destination_row[0..span]), src[offset..][0 .. span * 4]);
}
}
@@ -179,6 +353,92 @@ test "composite honours the damage rectangle" {
try std.testing.expectEqual(@as(u32, 0), back[4 * 8 + 4]); // outside damage
}
test "damage list keeps disjoint rectangles separate and merges overlap" {
var list = DamageList{};
list.add(Rect.init(0, 0, 10, 10));
list.add(Rect.init(100, 100, 10, 10)); // far away: its own entry
try std.testing.expectEqual(@as(usize, 2), list.slice().len);
list.add(Rect.init(5, 5, 10, 10)); // overlaps the first: united into it
try std.testing.expectEqual(@as(usize, 2), list.slice().len);
try std.testing.expectEqual(Rect.init(0, 0, 15, 15), list.slice()[0]);
try std.testing.expect(!list.isEmpty());
list.clear();
try std.testing.expect(list.isEmpty());
}
test "damage list folds overflow into the last entry instead of dropping it" {
var list = DamageList{};
var i: i32 = 0;
while (i < DamageList.capacity) : (i += 1) {
list.add(Rect.init(i * 100, 0, 10, 10)); // disjoint: fills every slot
}
try std.testing.expectEqual(@as(usize, DamageList.capacity), list.slice().len);
const overflow = Rect.init(0, 5000, 10, 10);
list.add(overflow);
try std.testing.expectEqual(@as(usize, DamageList.capacity), list.slice().len);
const last = list.slice()[DamageList.capacity - 1];
try std.testing.expect(!last.intersect(overflow).isEmpty()); // still covered
}
test "damage list ignores empty rectangles" {
var list = DamageList{};
list.add(Rect.empty);
try std.testing.expect(list.isEmpty());
}
test "tile grid coalesces a run of adjacent tiles into one rectangle" {
var grid = TileGrid{};
grid.reset(256, 128); // 4×2 tiles of 64 px
grid.add(Rect.init(10, 10, 100, 10)); // spans tiles (0,0) and (1,0)
var scratch: [TileGrid.maximum_rects]Rect = undefined;
const rects = grid.collect(&scratch);
try std.testing.expectEqual(@as(usize, 1), rects.len);
try std.testing.expectEqual(Rect.init(0, 0, 128, 64), rects[0]);
}
test "tile grid: full-screen damage collapses back to a single rectangle" {
var grid = TileGrid{};
grid.reset(1280, 720); // 20×12 tiles; the bottom row is partial (720 = 11*64 + 16)
grid.add(Rect.init(0, 0, 1280, 720));
var scratch: [TileGrid.maximum_rects]Rect = undefined;
const rects = grid.collect(&scratch);
try std.testing.expectEqual(@as(usize, 1), rects.len);
try std.testing.expectEqual(Rect.init(0, 0, 1280, 720), rects[0]);
}
test "tile grid keeps far-apart damage as separate rectangles" {
var grid = TileGrid{};
grid.reset(1280, 720);
grid.add(Rect.init(0, 0, 10, 10)); // top-left tile
grid.add(Rect.init(1000, 600, 10, 10)); // a far-away tile
var scratch: [TileGrid.maximum_rects]Rect = undefined;
const rects = grid.collect(&scratch);
try std.testing.expectEqual(@as(usize, 2), rects.len);
}
test "tile grid clamps edge tiles to the true screen size" {
var grid = TileGrid{};
grid.reset(100, 100); // 2×2 tiles, both partial in each axis
grid.add(Rect.init(0, 0, 100, 100));
var scratch: [TileGrid.maximum_rects]Rect = undefined;
const rects = grid.collect(&scratch);
try std.testing.expectEqual(@as(usize, 1), rects.len);
try std.testing.expectEqual(Rect.init(0, 0, 100, 100), rects[0]);
}
test "tile grid clear empties it and reset resizes it" {
var grid = TileGrid{};
grid.reset(256, 256);
grid.add(Rect.init(0, 0, 256, 256));
try std.testing.expect(!grid.isEmpty());
grid.clear();
try std.testing.expect(grid.isEmpty());
try std.testing.expect(grid.matches(256, 256));
grid.reset(512, 512);
try std.testing.expect(!grid.matches(256, 256));
try std.testing.expect(grid.isEmpty());
}
test "blitTile copies a packed tile, clipping and reading unaligned bytes" {
var back = [_]u32{0} ** (4 * 4);
const dst = Surface{ .pixels = &back, .stride = 4, .width = 4, .height = 4 };
+272 -28
View File
@@ -10,7 +10,9 @@
//! z-order, and visibility. Clients create layers, draw into them by command (`fill_rect`,
//! `blit_tile`), mark `damage`, and ask for a `present`; the compositor repaints only the
//! damaged region — clear it, paint the visible layers bottom-to-top into the backend's
//! surface, then `backend.present(damage)`. Shared-memory client surfaces are later
//! surface, then `backend.present(damage)`. Presents are paced by a ~60 Hz **frame clock**
//! (see `schedulePresent`), so any number of client presents and cursor moves inside one
//! interval coalesce into a single frame. Shared-memory client surfaces are later
//! (docs/display-v2.md).
const std = @import("std");
@@ -21,6 +23,8 @@ const backend_mod = @import("backend.zig");
const protocol = runtime.display_protocol;
const ipc = runtime.ipc;
const system = runtime.system;
const input = runtime.input;
const Thread = runtime.Thread;
const Rect = compositor.Rect;
const Surface = compositor.Surface;
@@ -48,8 +52,10 @@ var pending_modeset_check: bool = false;
var background: u32 = 0;
/// The layer stack. A fixed table (a compositor has few top-level surfaces during
/// bring-up); each used slot owns an mmap'd surface. `damage` accumulates the dirty
/// screen region since the last `present`, so a present touches only what changed.
/// bring-up); each used slot owns an mmap'd surface. `damage_list` accumulates the dirty
/// screen rectangles since the last `present`, so a present touches only what changed —
/// and keeps far-apart changes (the cursor here, an animating layer there) as *separate*
/// small copies rather than one huge bounding box (see compositor.DamageList).
const maximum_layers = 16;
const Layer = struct {
@@ -63,7 +69,67 @@ const Layer = struct {
};
var layers: [maximum_layers]Layer = [_]Layer{.{}} ** maximum_layers;
var damage: Rect = Rect.empty;
/// Which damage tracker drives `present` — a compile-time A/B switch (both are in
/// compositor.zig with the trade-off discussion):
/// .list — free-form dirty rectangles (tight bounds, heuristic merging)
/// .grid — a fixed 64-px tile grid (exact O(1) merging, tile-quantized repaints)
const DamageMode = enum { list, grid };
const damage_mode: DamageMode = .grid;
var damage_list: compositor.DamageList = .{};
var damage_grid: compositor.TileGrid = .{};
/// The **frame clock**: client `present` requests and cursor motion don't repaint
/// immediately — they accumulate damage and arm a one-shot timer, and the tick composites
/// everything pending as one frame. That paces presents to ~60 Hz no matter how fast
/// clients draw or the mouse moves (previously every mouse event became a full present).
/// No backend has a real vblank to pace by (docs/display-v2.md, "Fenced is not vsync");
/// this is the software stand-in, the same strategy Linux uses atop virtio-gpu. Bring-up
/// paths that need pixels on screen *now* (initialise, the self-checks) still call
/// `present()` directly.
///
/// The interval comes from the *active backend's* panel refresh rate (EDID: the loader
/// captures it for the GOP floor while firmware still runs; the native driver reads its
/// own and carries it in the announce). `updateFrameClock` re-derives it whenever the
/// backend changes — the boot framebuffer's clock dies with the GOP floor at upgrade.
/// Without a rate the clock defaults to 60 Hz, and it is clamped to [30, 120] Hz so a
/// mis-parsed EDID can neither starve nor flood the compositor.
var frame_interval_milliseconds: u64 = 16;
var frame_timer_armed = false;
/// Derive the frame-clock interval from the active backend's refresh rate and log what
/// the clock is now pacing to. Called at bring-up and again on every backend change.
fn updateFrameClock() void {
const reported = backend.info().refresh_hz;
const rate: u64 = if (reported == 0) 60 else @min(@max(reported, 30), 120);
frame_interval_milliseconds = @max(1000 / rate, 1);
var line: [96]u8 = undefined;
_ = system.write(std.fmt.bufPrint(&line, "display: frame clock {d} Hz ({s})\n", .{
1000 / frame_interval_milliseconds,
if (reported == 0) "default" else "panel EDID",
}) catch return);
}
/// Arm the frame clock unless a tick is already pending: any number of requests inside
/// one interval coalesce into that single tick's present.
fn schedulePresent() void {
if (frame_timer_armed) return;
frame_timer_armed = true;
_ = system.timerOnce(service_endpoint, frame_interval_milliseconds);
}
/// A timer landing — the frame clock, or the deferred first native present armed by
/// `attach_scanout`: present the accumulated damage, then run the one-shot mode-set
/// self-check if the native upgrade queued it.
fn frameTick() void {
frame_timer_armed = false;
present();
if (pending_modeset_check) {
pending_modeset_check = false;
modesetSelfCheck();
}
}
// --- geometry helpers -------------------------------------------------------
@@ -76,9 +142,20 @@ fn layerScreenRect(l: *const Layer) Rect {
return .{ .x = l.x, .y = l.y, .w = @intCast(l.surface.width), .h = @intCast(l.surface.height) };
}
/// Add `r` (screen coordinates) to the pending damage, clipped to the screen.
/// Add `r` (screen coordinates) to the pending damage, clipped to the screen. In grid
/// mode the grid re-sizes itself lazily when the screen geometry changes — every
/// geometry-changing path (`attach_scanout`, `set_mode`) damages the whole new screen
/// right after, so damage pending from the old geometry is safely superseded.
fn addDamage(r: Rect) void {
damage = damage.unite(r.intersect(screenRect()));
const clipped = r.intersect(screenRect());
switch (damage_mode) {
.list => damage_list.add(clipped),
.grid => {
const mode = backend.info();
if (!damage_grid.matches(mode.width, mode.height)) damage_grid.reset(mode.width, mode.height);
damage_grid.add(clipped);
},
}
}
// --- layer operations (called from onMessage and the self-check) ------------
@@ -181,21 +258,29 @@ fn compositeInto(clip: Rect) void {
}
}
/// Composite the accumulated damage into the backend's surface, hand it to the backend to
/// put on screen, then clear the damage. A no-op when nothing is dirty. The frame counter
/// advances regardless, so callers can name frames.
/// Composite each accumulated damage rectangle into the backend's surface, hand the list
/// to the backend to put on screen, then clear the damage. A no-op when nothing is dirty.
/// The frame counter advances regardless, so callers can name frames.
fn present() void {
const dirty = damage.intersect(screenRect());
if (!dirty.isEmpty()) {
compositeInto(dirty);
var scratch: [compositor.TileGrid.maximum_rects]Rect = undefined;
const dirty: []const Rect = switch (damage_mode) {
.list => damage_list.slice(),
.grid => damage_grid.collect(&scratch),
};
const had_damage = dirty.len != 0;
if (had_damage) {
for (dirty) |region| compositeInto(region);
backend.present(dirty);
}
damage = Rect.empty;
switch (damage_mode) {
.list => damage_list.clear(),
.grid => damage_grid.clear(),
}
frames += 1;
// The first present after a native upgrade confirms the composited frame actually reached
// the shared scanout surface (the automated stand-in for "it's on screen").
if (pending_native_verify and !dirty.isEmpty()) {
if (pending_native_verify and had_damage) {
pending_native_verify = false;
verifyNativePresent();
}
@@ -219,7 +304,7 @@ fn verifyNativePresent() void {
/// present channel, switch the backend to virtio-gpu, and queue a full-screen repaint. The
/// present is deferred to a timer (see `service_endpoint`) so it happens after this reply
/// unblocks the driver and it starts serving `.scanout`.
fn attachScanout(stride: u32, width: u32, height: u32, format: u32, capability: ?ipc.Handle, reply: []u8) usize {
fn attachScanout(stride: u32, width: u32, height: u32, format: u32, refresh_hz: u32, capability: ?ipc.Handle, reply: []u8) usize {
const cap = capability orelse return fail(reply);
if (width == 0 or height == 0 or stride < width) return fail(reply);
const mapped = runtime.shm.map(cap) orelse return fail(reply);
@@ -238,9 +323,11 @@ fn attachScanout(stride: u32, width: u32, height: u32, format: u32, capability:
.width = width,
.height = height,
.format = format,
.refresh_hz = refresh_hz,
.scanout = scanout,
} };
background = protocol.pack(format, 0x20, 0x30, 0x48); // re-pack the wallpaper for the mode
updateFrameClock(); // the GOP floor's clock dies here — pace by the GPU's EDID now
addDamage(screenRect()); // the whole new surface must be painted
pending_native_verify = true;
if (!reattach) pending_modeset_check = true; // the mode-set self-check runs once, on first upgrade
@@ -255,7 +342,8 @@ fn attachScanout(stride: u32, width: u32, height: u32, format: u32, capability:
/// After the native upgrade is verified, prove the runtime-resolution-change and fenced-present
/// paths: query the driver's modes, switch to one that differs from the current, re-composite
/// the whole screen at the new size, and confirm the backend now reports that geometry. The
/// present goes through the driver's fenced flush, so a clean present is a vsync present.
/// present goes through the driver's fenced flush, so a clean present is a *fenced* present —
/// completion-acknowledged and tear-free, not vblank-paced (docs/display-v2.md).
fn modesetSelfCheck() void {
if (!backend.canModeSet()) return;
var mode_list: [4]backend_mod.Mode = undefined;
@@ -287,7 +375,7 @@ fn modesetSelfCheck() void {
if (now.width == wanted.width and now.height == wanted.height) {
var line: [80]u8 = undefined;
_ = system.write(std.fmt.bufPrint(&line, "display: mode set to {d}x{d}, verified\n", .{ now.width, now.height }) catch "display: mode set, verified\n");
if (backend.hasVsync()) _ = system.write("display: vsync present ok\n");
if (backend.hasFencedPresent()) _ = system.write("display: fenced present ok\n");
} else {
_ = system.write("display: mode set FAILED (geometry unchanged)\n");
}
@@ -328,6 +416,157 @@ fn fail_check(_: []const u8) void {
_ = system.write("display: compositor self-check FAILED (setup)\n");
}
// --- cursor + mouse-input thread --------------------------------------------
//
// The compositor is the single owner of the framebuffer: only the main service
// loop touches `backend` and the layer stack. A dedicated listener thread (spawned
// in `initialise`) blocks on the input service's mouse stream, accumulates relative
// motion into an absolute cursor position, and hands that position to the main loop
// through `cursor_channel` — a single-slot latest-value cell (the renderer wants
// where the cursor *is*, not a replay of every delta). The listener never touches
// the compositor; it only writes the channel and pokes the main loop awake with a
// self-directed `ipc.send`, which arrives as a message-notification in the service
// loop (docs/threading.md, docs/display.md). Shared fate: a fault in the listener
// takes the whole display down and the supervisor restarts it (docs/resilience.md).
const cursor_size = 10; // a small square sprite — enough to prove tracking
const cursor_z = 0xFFFF_FFFF; // always above client layers
const cursor_report_threshold = 5; // px of travel before the tracking marker latches
var cursor_layer: ?u32 = null;
var cursor_origin_x: i32 = 0;
var cursor_origin_y: i32 = 0;
/// Latched once the cursor has demonstrably tracked a run of motion end to end
/// (source -> input service -> listener -> channel -> render): the `display-cursor`
/// test's success marker.
var cursor_tracking_reported: bool = false;
const poke_byte = [_]u8{0}; // the poke carries no payload; the value lives in the channel
/// Shared between the listener thread (producer) and the main loop (consumer).
/// Latest-value semantics with a coalesced wake: at most one poke is queued while
/// the main loop has not drained the last one, so a fast mouse cannot flood the
/// service endpoint.
const CursorChannel = struct {
lock: Thread.Mutex = .{},
poke_endpoint: ipc.Handle = 0,
x: i32 = 0,
y: i32 = 0,
buttons: u32 = 0,
dirty: bool = false,
poke_pending: bool = false,
const Snapshot = struct { x: i32, y: i32, buttons: u32 };
/// Producer (listener thread): record the newest position and, unless a wake is
/// already queued, poke the main loop awake.
fn publish(self: *CursorChannel, x: i32, y: i32, buttons: u32) void {
self.lock.lock();
self.x = x;
self.y = y;
self.buttons = buttons;
self.dirty = true;
const need_poke = !self.poke_pending;
if (need_poke) self.poke_pending = true;
self.lock.unlock();
if (need_poke) _ = ipc.send(self.poke_endpoint, &poke_byte);
}
/// Consumer (main loop): take the latest position, or null if nothing changed
/// since the last take. Clears the wake latch so the next publish pokes again.
fn take(self: *CursorChannel) ?Snapshot {
self.lock.lock();
defer self.lock.unlock();
self.poke_pending = false;
if (!self.dirty) return null;
self.dirty = false;
return .{ .x = self.x, .y = self.y, .buttons = self.buttons };
}
};
var cursor_channel: CursorChannel = .{};
fn clampAxis(value: i32, max: i32) i32 {
if (value < 0) return 0;
if (value > max) return max;
return value;
}
/// The mouse-listener thread. Blocks on the input service's mouse stream, accumulates
/// relative motion into an absolute position clamped to the screen, and publishes each
/// update. Runs for the life of the process; a parked `next()` leaves the core free to
/// halt (docs/halting.md). It reads only its own state and the channel — never the
/// compositor — so no lock guards the framebuffer.
fn mouseListener(width: u32, height: u32) void {
var mouse = input.subscribeMouse() orelse {
_ = system.write("display: mouse subscribe failed\n");
return;
};
// Our own handle to the compositor's endpoint. IPC handles are per-thread, so we
// cannot reuse the main thread's service handle — we look the service up to install a
// handle in this thread's table. A poke posted here wakes the compositor loop parked
// in replyWait (docs/threading.md: handles do not cross threads).
cursor_channel.poke_endpoint = ipc.lookup(.display) orelse {
_ = system.write("display: mouse listener could not reach the compositor endpoint\n");
return;
};
const max_x: i32 = @as(i32, @intCast(width)) - 1;
const max_y: i32 = @as(i32, @intCast(height)) - 1;
var x: i32 = @divTrunc(max_x, 2);
var y: i32 = @divTrunc(max_y, 2);
var buttons: u32 = 0;
while (true) {
const event = mouse.next() orelse continue;
// Switch on the raw kind (not @enumFromInt, which would panic on a scroll or
// future kind): motion moves the cursor, anything else just updates buttons.
if (event.kind == @intFromEnum(input.MouseEventKind.motion)) {
x = clampAxis(x + event.dx, max_x);
y = clampAxis(y + event.dy, max_y);
} else {
buttons = event.buttons;
}
cursor_channel.publish(x, y, buttons);
}
}
/// Consume the latest cursor position from the channel and move the cursor layer to it.
/// Runs on the main loop (the compositor owner) in response to a listener poke.
/// `configureLayer` damages both the old and new footprints; the frame clock presents
/// them at the next tick, so a fast mouse coalesces to at most ~60 repaints a second.
fn renderCursor() void {
const snapshot = cursor_channel.take() orelse return;
const id = cursor_layer orelse return;
_ = configureLayer(id, snapshot.x, snapshot.y, cursor_z, true);
schedulePresent();
if (!cursor_tracking_reported and
@abs(snapshot.x - cursor_origin_x) >= cursor_report_threshold and
@abs(snapshot.y - cursor_origin_y) >= cursor_report_threshold)
{
cursor_tracking_reported = true;
_ = system.write("display: cursor tracking mouse ok\n");
}
}
/// Create the cursor sprite (a top-z square) at screen centre and spawn the listener
/// thread. Called from `initialise` once the backend is up. If either step fails the
/// display still serves drawing clients — it just has no cursor.
fn startCursorTracking() void {
const mode = backend.info();
cursor_origin_x = @divTrunc(@as(i32, @intCast(mode.width)), 2);
cursor_origin_y = @divTrunc(@as(i32, @intCast(mode.height)), 2);
const id = createLayer(cursor_origin_x, cursor_origin_y, cursor_size, cursor_size, cursor_z, true) orelse {
_ = system.write("display: could not create cursor layer\n");
return;
};
cursor_layer = id;
_ = fillLayer(id, Rect.init(0, 0, cursor_size, cursor_size), protocol.pack(mode.format, 0xF0, 0xF0, 0xF0));
present(); // show the cursor at its start position
_ = Thread.spawn(.{}, mouseListener, .{ mode.width, mode.height }) catch {
_ = system.write("display: could not spawn mouse listener\n");
};
}
// --- service ----------------------------------------------------------------
fn initialise(endpoint: ipc.Handle) bool {
@@ -347,9 +586,13 @@ fn initialise(endpoint: ipc.Handle) bool {
_ = system.write(std.fmt.bufPrint(&line, "display: online {d}x{d} pitch {d} format {d}\n", .{
mode.width, mode.height, mode.pitch, mode.format,
}) catch "display: online\n");
updateFrameClock();
_ = system.write("display: presented frame 0\n");
selfCheck();
// Bring up the cursor and the mouse-listener thread now that the backend is live.
startCursorTracking();
return true;
}
@@ -405,11 +648,13 @@ fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?ipc.Han
return ok(reply);
},
@intFromEnum(protocol.Operation.present) => {
present();
// Scheduled, not immediate: the frame clock composites the accumulated damage
// at the next tick, so back-to-back client presents coalesce into one frame.
schedulePresent();
return ok(reply);
},
@intFromEnum(protocol.Operation.attach_scanout) => {
return attachScanout(request.x, request.width, request.height, request.colour, capability, reply);
return attachScanout(request.x, request.width, request.height, request.colour, request.y, capability, reply);
},
@intFromEnum(protocol.Operation.set_mode) => {
if (!backend.setMode(request.width, request.height)) return fail(reply);
@@ -435,16 +680,15 @@ fn onMessage(message: []const u8, reply: []u8, sender: u32, capability: ?ipc.Han
}
}
/// The only notification the compositor arms is the post-attach present timer: repaint the
/// screen into the freshly attached native surface, verify the frame landed, then run the
/// one-shot mode-set self-check (V5).
/// Two notification sources reach the compositor, and one coalesced badge can carry
/// both, so each bit is handled independently. A **message-notification** is a poke from
/// the mouse-listener thread (a buffered self-`ipc.send`, `notify_message_bit`): fold the
/// newest cursor position into the scene. A **timer** (`notify_timer_bit`) is the frame
/// clock — or the deferred first native present after `attach_scanout` — either way,
/// present the accumulated damage.
fn onNotification(badge: u64) void {
_ = badge;
present(); // native present + verify (first timer fire after the upgrade)
if (pending_modeset_check) {
pending_modeset_check = false;
modesetSelfCheck();
}
if (badge & ipc.notify_message_bit != 0) renderCursor();
if (badge & ipc.notify_timer_bit != 0) frameTick();
}
pub fn main() void {
+7 -5
View File
@@ -24,11 +24,13 @@ pub const Operation = enum(u32) {
damage = 6,
/// present(): composite the dirty layers and flush to the screen.
present = 7,
/// attach_scanout(x=stride, width, height, colour=format) + <surface capability>: a native
/// scanout driver announces itself, handing over the shared scanout surface as an `ipc_call`
/// send_cap. The compositor maps it, looks up the driver's `.scanout` present channel, and
/// upgrades off the GOP floor (docs/display-v2.md V4). `x` is the surface's row stride in
/// pixels, `colour` the DisplayFormat.
/// attach_scanout(x=stride, y=refresh_hz, width, height, colour=format) + <surface
/// capability>: a native scanout driver announces itself, handing over the shared scanout
/// surface as an `ipc_call` send_cap. The compositor maps it, looks up the driver's
/// `.scanout` present channel, and upgrades off the GOP floor (docs/display-v2.md V4).
/// `x` is the surface's row stride in pixels, `y` the panel refresh rate from the
/// driver's EDID read (0 = unknown; paces the compositor's frame clock), `colour` the
/// DisplayFormat.
attach_scanout = 8,
/// set_mode(width, height): change the display resolution — only a native backend that
/// reports `canModeSet` honours it; on the GOP floor it fails (docs/display-v2.md V5).
+23 -2
View File
@@ -10,17 +10,38 @@
//! keyboard and mouse drivers publish their own synthetic streams today; swapping in
//! decoded hardware is a follow-up (see docs/input.md).
const std = @import("std");
const runtime = @import("runtime");
const input = runtime.input;
const system = runtime.system;
pub fn main() void {
pub fn main(init: runtime.process.Init) void {
var source = input.connectSource() orelse {
_ = system.write("input-source: input service unavailable\n");
return;
};
_ = system.write("input-source: publishing synthetic input events\n");
// "mouse" mode publishes a steady stream of pure motion (dx=dy=+1), for driving a
// cursor (the `display-cursor` test). The default "rotate" mode cycles all device
// classes to exercise the service's per-device routing (the `input` test).
const mode = init.arguments.get(1) orelse "rotate";
if (std.mem.eql(u8, mode, "mouse")) {
_ = system.write("input-source: publishing synthetic mouse motion\n");
while (true) {
_ = source.publishMouseEvent(.{
.kind = @intFromEnum(input.MouseEventKind.motion),
.button = 0,
.dx = 1,
.dy = 1,
.scroll_x = 0,
.scroll_y = 0,
.buttons = 0,
});
system.sleep(20); // ~50 events/sec: moves the cursor briskly
}
}
_ = system.write("input-source: publishing synthetic input events\n");
var step: usize = 0;
while (true) : (step +%= 1) {
// Rotate across the device classes so every publish path (and the service's
+16 -5
View File
@@ -185,6 +185,16 @@ CASES = [
{"name": "display-demo",
"expect": r"display-demo: scene up[\s\S]*display-demo: ok",
"fail": r"display-demo: (no display|create failed)|display: could not|CPU EXCEPTION|KERNEL PANIC"},
# Threaded compositor tracks a mouse (docs/threading.md, docs/display.md): the display
# runs a mouse-listener thread alongside its compositor loop. `input-source mouse`
# publishes pure motion -> the input service fans it to the display's listener -> the
# listener accumulates it into a cursor position handed to the render loop over a
# single-slot channel. `display: cursor tracking mouse ok` latches once the cursor has
# tracked a run of that motion end to end.
{"name": "display-cursor",
"smp": 4,
"expect": r"display: online \d+x\d+[\s\S]*display: cursor tracking mouse ok",
"fail": r"display: (could not|mouse subscribe failed)|CPU EXCEPTION|KERNEL PANIC"},
# Shared memory (v2 V2): shm-client creates a region, writes a pattern, and passes its
# capability to shm-server, which maps it and confirms the same bytes — proving
# cross-process shared pages over the extended capability passing.
@@ -211,15 +221,16 @@ CASES = [
# require all three markers to appear somewhere rather than in a fixed order.
"expect": r"(?s)(?=.*display: scanout upgraded to virtio-gpu)(?=.*display: native present verified)(?=.*display-demo: ok)",
"fail": r"display: native present FAILED|display: could not|display-demo: (no display|create failed)|CPU EXCEPTION|KERNEL PANIC"},
# Mode-set + EDID + vsync (v2 V5): same boot as display-native. After upgrading, the
# compositor queries the driver's modes, switches to a different resolution, and confirms the
# backend now reports it; the fenced present path makes it a vsync present. (The driver also
# logs the EDID preferred mode during bring-up.) Reuses the display-native kernel scenario.
# Mode-set + EDID + fenced presents (v2 V5): same boot as display-native. After upgrading,
# the compositor queries the driver's modes, switches to a different resolution, and confirms
# the backend now reports it; each present is fenced — completion-acknowledged and tear-free,
# not vblank-paced (docs/display-v2.md, "Fenced is not vsync"). (The driver also logs the
# EDID preferred mode during bring-up.) Reuses the display-native kernel scenario.
{"name": "display-modeset",
"build_case": "display-native",
"qemu_extra": ["-device", "virtio-gpu-pci"],
"mem": "512M",
"expect": r"(?s)(?=.*display: mode set to \d+x\d+, verified)(?=.*display: vsync present ok)",
"expect": r"(?s)(?=.*display: mode set to \d+x\d+, verified)(?=.*display: fenced present ok)",
"fail": r"display: mode set FAILED|display: mode-set self-check: |display: native present FAILED|CPU EXCEPTION|KERNEL PANIC"},
# Resilience: driver restart + re-attach (v2 V6). device-manager (in test-scanout-restart
# mode) kills the virtio-gpu driver once after it hellos; the restart policy respawns it, it